Skip to the content.

AIニュース 2026-07-23

自動生成: 2026-07-23 12:23 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Building AI infrastructure with the Effingham County communityOpenAI

    OpenAI announces Project Camellia in Effingham County, Georgia, with…

  2. Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis MissionGoogle DeepMind

    Google commits $40M in AI tokens and credits for the Genesis Mission

  3. How news organizations are using AI to advance their vital missionsOpenAI

    News organizations are using AI to strengthen reporting, grow audienc…

  4. Advancing the next era of national scienceOpenAI

    OpenAI outlines its commitment to advancing American science working…

  5. Introducing OpenAI PresenceOpenAI

    Introducing OpenAI Presence, a proven enterprise AI agent platform th…

  6. AMDとAnthropicが戦略的提携 「Helios」を最大2GW導入、最大50億ドルの出資もITmedia AI+

    AMDは、Anthropicとの戦略的提携を発表した。AnthropicはAMDの「Helios」および「Instinct MI450」シ…

  7. Macアプリ版「Claude Code」がiOSシミュレータと連携 「Computer Use」なしでアプリ操作ITmedia AI+

トピック別件数

日本語メディア9件

ITmedia AI+ (日本語)

10:19 JSTLLM/生成AIAnthropicMicrosoft

企業向けAIツールの成長率トップはAnthropic、アカウント数が最も多いのはMicrosoft 365 Okta調査

アイデンティティ管理サービスを提供する米Oktaは、同社のサービスを用いている2万社以上の匿名化されたアクセスデータに基づく、企業でのAIツール利用実態について調査結果を発表しました。

08:00 JSTLLM/生成AI

「AI使うなら値引きできる?」の“暴論”に、日立はどう立ち向かう? レガシー刷新でのAI活用の現在地

生成AIはレガシーシステム刷新の現場で具体的にどのように使われているのか。ユーザー企業自身がAIを使いこなす時代にベンダーが担う役割とは。

07:00 JSTハードウェア/半導体ビジネス/資金調達NVIDIA

NVIDIAフアンCEOが語る“日本復活”のシナリオ 10年続く半導体バブルと「原発活用」の勝算

米NVIDIAのジェンスン・フアンCEOが来日し、日本経済の復活を宣言した。国内のAIインフラ構築へ数十億ドル規模の投資を発表。フアン氏は「何兆ものAIがAIを使う時代」の到来によって半導体需要は人口に制約されないと指摘。データセンターの電力不足に対して「原発活用」を日本の強み…

07:00 JSTLLM/生成AI

「天才デザイナー依存」の限界 AIで“平均点”しか出せない組織を変える「ノウハウ共有術」

生成AIの普及で“平均点デザイン”が量産され、プロダクトの同質化が進む。優秀な個人にノウハウが閉じる属人化も課題だ。デザインツール大手FigmaのCDOは、プロンプトやAIとの対話プロセスを含めた全体をチームで共有すべきだと訴える。1つのキャンバス上でデザイン、コード、AIをつ…

06:55 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達AnthropicClaude

AMDとAnthropicが戦略的提携 「Helios」を最大2GW導入、最大50億ドルの出資も

AMDは、Anthropicとの戦略的提携を発表した。AnthropicはAMDの「Helios」および「Instinct MI450」シリーズを最大2GW規模で導入し、2027年上半期から順次展開する。AMDは最大50億ドルの株式投資を行うほか、Claudeを活用したGPU環…

05:00 JSTLLM/生成AIエージェントClaude

AI時代、開発チームの人材は“5つの型”に分かれる Claude Code開発責任者の見立て

Claude Code開発責任者のボリス・チャーニー氏が、自身のチームで働く人は「5つの型」に分けられると指摘した。AIで職種の垣根が崩れ始めた今、肩書きではなく働き方で人を捉える新時代の発想を、初心者にも分かるように読み解く。

15:21 JSTLLM/生成AIOpenAI2媒体が報道

OpenAIのモデルがサイバー攻撃能力評価中に暴走 テストの答えを求めてHugging Faceに侵入

OpenAIのAIモデルが、サイバー攻撃能力の評価中に隔離環境を突破し、Hugging Faceの本番インフラに侵入していたことが分かった。ベンチマークの解答を入手するため、ゼロデイ脆弱性の悪用や認証情報の窃取を重ねていたという。

出典:ITmedia AI+TechCrunch AI
13:00 JSTその他

無料で身に付くデータサイエンス 延べ23万人が受講、総務省が募集開始

総務省は、データサイエンスオンライン講座「社会人のためのデータサイエンス入門」をリニューアルし、受講者の募集を開始した。慶應義塾大学の安宅和人教授など12人を講師に迎え、統計データ分析の基本を無料で学べる、社会人や大学生向けの入門講座だ。

海外メディア14件

TechCrunch AI (英語)

08:47 JSTその他

After shocking quarter, IBM insists that AI isn’t killing the mainframe

After IBM's stock crashed last week on warnings of poor mainframe sales, the CEO explained that AI wrecked corporate hardware budget, tempo…

07:01 JSTその他Google

Google justifies its massive AI spending with a booming cloud business

Google's cloud business is thriving, as companies adopting its AI and AI infrastructure services help the tech giant to report record profi…

05:49 JSTLLM/生成AI規制/政策Anthropic

Treasury threatens sanctions after White House claims Moonshot distilled Anthropic’s Fable

The episode has also intensified a broader debate in Washington over the influx of Chinese open models.

03:50 JSTロボティクスビジネス/資金調達

Travis Kalanick’s robotics company raises $1.7B, led by a16z

Uber is also investing in Travis Kalanick's company Atoms, which has made gauzy claims about using industrial AI to modernize the world.

03:13 JSTビジネス/資金調達

Yope raises $12.3M to build a private social network without algorithms or ads

Yope, a fast-growing social app focused on private groups of friends and family, has raised $12.3 million in seed funding. Instead of chasi…

02:54 JSTその他

Monday.com lays off hundreds to focus on AI

The company said it is reducing its headcount by 20%, or about 630 staff, to "support a leaner, more focused operating model" as it focuses…

01:24 JSTその他

Arcee, a US open source AI lab, says Chinese models are not inherently dangerous

As Chinese AI models grow in capability and popularity among U.S. companies, the arguing over what should be done about them has reached a…

01:23 JSTその他

Substack’s new tool tells you who’s been writing their newsletters with AI

Substack is giving readers a way to estimate how much of a newsletter was written by AI, signaling a broader shift toward transparency arou…

01:13 JSTLLM/生成AIOpenAI

OpenAI’s AI spending spree has ballooned to $750B

OpenAI will spend the equivalent of Sweden's GDP on infrastructure through 2030.

23:00 JSTLLM/生成AIAnthropic

Menlo Ventures’ Matt Murphy explains what AI startups founders must do differently

Anthropic leaped to a $47 billion revenue run rate by May, compared to $9 billion in 2025. It’s the kind of growth that Menlo Ventures’ Mat…

22:20 JSTその他

The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari

We’ve compiled an overview of some of the top alternative browsers available today aiming to challenge Chrome and Safari.

22:00 JSTビジネス/資金調達

Passionfroot raises $15M to expand its B2B creator marketplace to the US

Passionfroot, a German startup building a marketplace connecting B2B creators with brands, has raised $15M in a Series A round led by Insig…

19:00 JSTエージェントビジネス/資金調達

Glow emerges from stealth at $1.2B valuation to challenge endpoint security in the AI era

Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises.

17:00 JSTその他

Synthesia’s AI training platform is moving beyond videos into live coaching

Synthesia launched AI Roleplay Sessions, an interactive enterprise training platform where employees practice workplace conversations with…

公式ブログ5件

OpenAI (英語)

22:00 JSTLLM/生成AIビジネス/資金調達OpenAI

Building AI infrastructure with the Effingham County community

OpenAI announces Project Camellia in Effingham County, Georgia, with commitments to responsible energy, community investment, jobs, and acc…

22:00 JSTLLM/生成AIOpenAI

How news organizations are using AI to advance their vital missions

News organizations are using AI to strengthen reporting, grow audiences, and improve business operations, with OpenAI tools supporting jour…

21:00 JSTLLM/生成AIOpenAI

Advancing the next era of national science

OpenAI outlines its commitment to advancing American science working with the U.S. Department of Energy and national labs to use frontier A…

14:30 JSTLLM/生成AIエージェントOpenAI

Introducing OpenAI Presence

Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for custom…

Google DeepMind (英語)

22:38 JSTその他Google

Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission

Google commits $40M in AI tokens and credits for the Genesis Mission

論文322件

arXiv cs.AI (英語)

13:00 JST研究/論文

SysAdmin: フロンティア AI における機器のパワーシーキングの測定

AI システムがリソースを取得したり、監視を回避したり、タスク要件を超えて終了に抵抗したりする動作として定義される権力追求は、制御不能 (LoC) リスクの主な要因として特定されています。この研究では、フロンティア言語モデルを忠実度の高い Linux サンドボックス内の自律システム管理者として位置付け、自己保存、自律性の向上、リソースの獲得、環境の変更、戦略的隠蔽の 5 つの側面にわたって権力追求の傾向を測定するベンチマークである SysAdmin を紹介します。私たちは、合計 2,800 のタスクで 4 つの実験条件にわたって 7 つのフロンティア モデルを評価しました。人間が注釈を付けた校正データを使用してバイアスを補正した後、補正されたパワーシーキング推定値はモデルごとに 0 ~約 5 パーセントの範囲でした。また、100% の検出を達成する明示的な電力探索プロンプトを使用したポジティブ コントロールも実施し、測定感度を検証しました。私たちの調査結果は、現在のフロンティアモデルが自然主義的なシステム管理コンテキストにおいて最小限の自発的パワーシーキングを示すことを示していますが、モデル固有の故障モードは評価が多様なミスアライメントパターンをテストする必要があることを示唆しています。それにもかかわらず、私たちは、スペックゲームや目標修正への抵抗など、(パワー追求よりも)他のより顕著な失敗モードを発見しました。

原文 (English)

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond task requirements is identified as a key driver of Loss of Control (LoC) risk. In this work, we introduce SysAdmin, a benchmark that positions frontier language models as autonomous system administrators in a high-fidelity Linux sandbox to measure power-seeking propensity across five dimensions: self-preservation, increasing autonomy, resource acquisition, environment modification, and strategic concealment. We evaluated seven frontier models across four experimental conditions in a total of 2800 tasks. After bias correction using human-annotated calibration data, corrected power-seeking estimates ranged from 0 to about 5 percent per model. We also conducted a positive control with explicit power-seeking prompts that achieved 100% detection, validating measurement sensitivity. Our findings indicate current frontier models exhibit minimal spontaneous power-seeking in naturalistic system administration contexts, though model-specific failure modes suggest evaluations must test diverse misalignment patterns. Nevertheless, we discovered other more pronounced failure modes (than power-seeking) such as specification gaming and resistance to goal modification.

13:00 JSTLLM/生成AIビジネス/資金調達

証拠連鎖評価による校正済みの選択的ファクトチェック

大規模言語モデル (LLM) は強力なファクトチェック精度を実現できますが、強制的な二者択一の決定により、重大な信頼性の問題が隠蔽されます。つまり、裏付けとなる証拠が弱い、希薄である、または内部的に矛盾している場合でも、システムは自信を持って判定を下す可能性があります。私たちは、すべての主張に対して真/偽の決定を要求するのではなく、不確実な評決によって棄権を許可する選択的事実確認フレームワークである証拠連鎖評価(ECE)を通じてこの問題に取り組んでいます。評価されたシステムは、Web 検索、学術検索、実行可能チェックを通じて証拠を収集し、信頼性とソースレベルのメタデータを含む構造化された判定を返すツールを使用する検証エージェントです。 ECE-Bench では、ECE は回答されたクレームに対して 91.6% の標準精度、93.7% のカバレッジ、および 97.8% の選択的精度を達成しています。 ECE は、予想されるキャリブレーション エラー、ブライアー スコア、または AURC などの集計キャリブレーション メトリクスに関して最も強力な検索ベースラインを上回るパフォーマンスはありませんが、明確な選択的予測のトレードオフを提供します。つまり、システムは、95 件中 6 件を保留しながら、回答されたクレームについて非常に高い精度を維持します。これらの延期されたケースは、信頼性の低い証拠設定(情報源レベル L4 で 5/6)に集中しており、棄権が認識論的に弱い証拠を処理するための安全志向のメカニズムとして機能するという見解を裏付けています。コードは https://github.com/cheshireyang/ECE.git で入手できます。

原文 (English)

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions conceal a critical reliability problem: systems may issue confident verdicts even when supporting evidence is weak, sparse, or internally inconsistent. We address this issue through Evidence Chain Evaluation (ECE), a selective fact-checking framework that permits abstention via an uncertain verdict instead of requiring a true/false decision for every claim. The evaluated system is a tool-using verification agent that gathers evidence through web search, scholarly search, and executable checks, and then returns a structured verdict with confidence and source-level metadata. On ECE-Bench, ECE achieves 91.6% standard accuracy, 93.7% coverage, and 97.8% selective accuracy on answered claims. Although ECE does not outperform the strongest retrieval baseline on aggregate calibration metrics such as Expected Calibration Error, Brier score, or AURC, it delivers a clear selective-prediction trade-off: the system maintains very high accuracy on answered claims while deferring 6 of 95 cases. These deferred cases are concentrated in lower-reliability evidence settings (5/6 at source level L4), supporting the view that abstention functions as a safety-oriented mechanism for handling epistemically weak evidence. Code is available at https://github.com/ cheshireyang/ECE.git

13:00 JSTLLM/生成AIGPT / ChatGPT

BatchDAG: エンタープライズ データに対するスケーラブルなアドホック分析のための LLM 計画実行グラフ

大規模言語モデル (LLM) は、個々のドキュメントの分析には優れていますが、コンテキストのオーバーフロー、エンティティごとの帰属の損失、および連続したツール呼び出しによる線形遅延により、エンタープライズ規模のデータセットに対する徹底的なエンティティ間の分析の質問には分解されません。我々が紹介する BatchDAG は、LLM が操作 (SQL クエリ、セマンティック検索、メモリ内変換、並列ファンアウト、シングルショット分析) の型付き有向非巡回グラフ (DAG) を生成するシステムで、決定論的エンジンがトポロジカルウェーブ並列処理と構造化された JSON データ フローで評価します。主要な最適化であるエンティティ認識バッチ処理は、ファンアウト前に論理エンティティごとに行をグループ化し、LLM コールを最大 47 分の 1 に削減します。 BatchDAG は主に手動で最適化されたパイプラインに比べて精度が向上するものではありません。むしろ、これは、手動で設計された複数のワークフローを、自然言語から適切な実行戦略を生成する単一のシステムに置き換える汎用のオーケストレーション レイヤーです。 12 個のトランスクリプトを多く使用するクエリの対照実験では、BatchDAG (3.74/5) は、専門家が設計したパイプライン (3.25/5) に匹敵する品質を達成し、ReAct エージェント (3.09/5、p<0.01) を大幅に上回り、優れた来歴 (トランスクリプト証拠率 77% 対ベースラインの 46 ~ 60%) を実現しました。制御されたアブレーションは、構造化された JSON 中間体が散文要約と比較して幻覚を 27% 減少させることを示しています (対応のある t 検定、p=0.107、n=12)。プランナーは、300 回の計画コール全体で 98.8% の有効 DAG 率を達成しました。 Brevian.ai の実稼働環境では、BatchDAG は 50,000 件を超える会議を 60 秒以内に処理し、公開されている GPT-5.1 価格でクエリごとのコストは 0.02 ~ 0.24 ドルと測定されています。

原文 (English)

BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data

Large language models (LLMs) excel at analyzing individual documents but break down on exhaustive, cross-entity analytical questions over enterprise-scale datasets due to context overflow, loss of per-entity attribution, and linear latency from sequential tool calls. We present BatchDAG, a system in which an LLM generates a typed directed acyclic graph (DAG) of operations -- SQL queries, semantic searches, in-memory transforms, parallel fan-outs, and single-shot analyses -- which a deterministic engine evaluates with topological-wave parallelism and structured JSON data flow. A key optimization, entity-aware batching, groups rows by logical entity before fan-out, reducing LLM calls by up to 47x. BatchDAG is not primarily an accuracy improvement over hand-optimized pipelines; rather, it is a general-purpose orchestration layer that replaces multiple hand-engineered workflows with a single system that generates the appropriate execution strategy from natural language. In controlled experiments on 12 transcript-heavy queries, BatchDAG (3.74/5) achieves quality comparable to an expert-designed pipeline (3.25/5) and significantly outperforms a ReAct agent (3.09/5, p<0.01), with superior provenance (77% transcript evidence rate vs. 46-60% for baselines). A controlled ablation shows structured JSON intermediates reduce hallucinations by 27% versus prose summaries (paired t-test, p=0.107, n=12). The planner achieves 98.8% valid-DAG rate across 300 planning calls. In production at Brevian.ai, BatchDAG processes queries over 50,000+ meetings in under 60 seconds, with measured per-query costs of $0.02-$0.24 at published GPT-5.1 pricing.

13:00 JSTエージェント

大規模な AI ツール検出: 必要なのは DNS だけです

来たるべき自律型 AI エージェントの時代には、何百万ものツールをナビゲートできる検出メカニズムが必要ですが、既存のソリューションは O(N) の複雑さと一元化されたガバナンスの下で機能しません。別の脆弱なオーバーレイを構築する代わりに、インターネットの最も回復力のある基盤であるドメイン ネーム システム (DNS) にセマンティック ツールの検出を後付けする根本的なフレームワークである ToolDNS を提案します。 ToolDNS は、機能的意図と組織の信頼を階層型名前空間に埋め込むことにより、高価なセマンティック検索を一連の軽量な O(log N) 名前解決に変換します。部分的に展開された名前、EDNS0 インテント ペイロード、論理サブドメインという、分散型ガバナンスとセマンティック プルーニングを可能にする 3 つのプロトコル準拠の機能強化を導入します。断片化されたツール環境全体でこのアプローチを厳密に評価するために、MCP、A2A、RESTful、およびスキル プロトコルにわたる 33,688 個の実際のツールで構成される大規模な異種ベンチマークを構築してリリースします。このデータセットでは、ToolDNS は最先端の検索精度を実現しながら、クエリごとの検索スペースを 95.26% 削減します。さらに、UDP ネイティブの設計により、HTTP ベースのレジストリと比較して検出遅延が桁違いに短縮されます。私たちの研究は、スケーラブルな AI の相互運用性には、追加のミドルウェアではなく、すでに足元にあるインフラストラクチャをより賢く利用することが必要であることを示しています。

原文 (English)

AI Tool Discovery at Scale: All You Need is DNS

The coming era of autonomous AI agents demands a discovery mechanism capable of navigating millions of tools, yet existing solutions buckle under O(N) complexity and centralized governance. Instead of building another fragile overlay, we propose ToolDNS, a radical framework that retrofits semantic tool discovery onto the Internet's most resilient substrate: the Domain Name System (DNS). By embedding functional intent and organizational trust into a hierarchical namespace, ToolDNS transforms an expensive semantic search into a series of lightweight, O(log N) name resolutions. We introduce three protocol-compliant enhancements to enable decentralized governance and semantic pruning: partially unfolded names, EDNS0 intent payloads, and logical subdomains. To rigorously evaluate this approach across the fragmented tooling landscape, we construct and release a large-scale heterogeneous benchmark comprising 33,688 real-world tools spanning MCP, A2A, RESTful, and Skill protocols. On this dataset, ToolDNS slashes the per-query search space by 95.26% while matching state-of-the-art retrieval accuracy. Furthermore, its UDP-native design reduces discovery latency by orders of magnitude compared to HTTP-based registries. Our work demonstrates that scalable AI interoperability requires not more middleware, but a smarter utilization of the infrastructure already beneath our feet.

13:00 JSTエージェント

エージェント障害パスから定量化された残留リスクまで: 回復力のあるエージェント AI のための構成フレームワーク

Agentic AI は、現在のリスク モデルが表現できるよりも速く信頼境界を越えています。既存のアプローチでは、2 つの部分ビューのうちの 1 つが提供されます。それらは、移転可能な残留リスク推定値を生成せずに障害メカニズムを記述するか、内部障害パスをブラック ボックスとして扱いながらリスク推定値を生成します。私たちは、物理状態、センサー、データ、コンピューティング、アクチュエーター、環境、時間にわたる 7 層の整合性分解である CPSAINT を提案することで、これら 2 つのビューを結合し、各障害パスを定量化されたリスク インスタンスにマッピングする残留リスク関数である FRIESA-K と組み合わせます。 FRIESA-K は、制御の有効性が非公式のスコアとして割り当てられるのではなく、状態ダイナミクスから導出されるように、制御された吸収マルコフ モデル内の抵抗項 K を根拠とします。その結果、回復力のあるエージェント型 AI と身体型 AI のための、簡潔なメカニズムから規模までのパイプラインが実現します。抵抗関数の新しい変数としてガバナンスを挿入するのではなく、別の追加ペナルティを通じてガバナンスの可観測性を報告します。私たちは、有効な障害パスを明確に定義されたリスク インスタンスに結び付ける構造的な構成可能性を形式化し、ハード リアルタイム倉庫ロボットとガバナンス機能を備えた金融サービス エージェントという 2 つの対照的なシナリオに関するフレームワークを示します。どちらの場合でも、同じ層の文法、可変セマンティクス、および動的耐性の構築はそのまま残ります。したがって、クロスドメイン推論、明示的な仮定、および構成可能な信頼の定量的に根拠のある形式主義をサポートするコンパクトなカーネルが得られます。

原文 (English)

From Agent Failure Paths to Quantified Residual Risk: A Compositional Framework for Resilient Agentic AI

Agentic AI is crossing trust boundaries faster than current risk models can represent. Existing approaches provide one of two partial views. They either describe failure mechanisms without producing a transferable residual-risk estimate, or they produce a risk estimate while treating the internal failure path as a black box. We couple those two views by proposing CPSAINT, a seven-layer integrity decomposition over Physical state, Sensors, Data, Compute, Actuators, Environment, and Time, paired with FRIESA-K, a residual-risk functional that maps each failure path to a quantified risk instance. FRIESA-K grounds the resistance term K in a controlled absorbing Markov model so that control effectiveness is derived from state dynamics rather than assigned as an informal score. The result is a concise mechanism-to magnitude pipeline for resilient agentic and embodied AI. We report governance observability through a separate additive penalty instead of inserting governance as a new variable in the resistance functional. We formalize structural composability linking valid failure paths to well-defined risk instances and show the framework on two contrasting scenarios a hard real-time warehouse robot and a governance-instrumented financial-services agent. Across both cases, the same layer grammar, variable semantics, and dynamic-resistance construction remain intact. Thus, we obtain a compact kernel that supports cross-domain reasoning, explicit assumptions, and quantitatively grounded formalism of composable trust.

13:00 JSTエージェントビジネス/資金調達

SAAG: 構造化されたエージェントの評価とグラウンディング

エージェント呼び出しの完全一致評価は、質的に異なる障害モードを曖昧にします。モデルは正しい関数を選択しているにもかかわらず引数値を幻覚させたり、誤った理由でエージェントを選択しながらスキーマを満たしたりする可能性があります。既存のベンチマークでは、これらの区別が単一のバイナリ スコアにまとめられているため、担当者はエージェントの呼び出しがどこで失敗するかを診断できません。我々は、エージェント呼び出しの評価を、レジストリ適合性、構造的完全性、議論の根拠という 3 つの段階に分解し、それぞれが解釈可能な段階固有の診断を生成するカスケード診断フレームワークを SAAG に提案します。さらに、これらの診断により、反復的な自己修復が可能になります。つまり、予測が失敗した場合、ステージ固有の信号が、グランドトゥルース値を漏らすことなく、ターゲットを絞った修正をガイドします。このフレームワークは、3 つのローカル サブ 4B パラメーター モデルを使用して、5、10、および 15 エージェントのレジストリ サイズにわたる Glaive の関数呼び出しデータセットから導出された制御されたベンチマークで評価されます。構造化フィードバックにより引数の精度が一貫して向上し、シングルパス推論や有益でないバイナリ フィードバックと比較して値の幻覚が軽減されますが、エンドツーエンドの F1 ゲインは控えめでモデルに依存します。これらの結果は、段階分解された診断評価が、モデル ファミリやレジストリ スケール全体でのエージェント呼び出しの信頼性を理解し、改善するために必要なレンズであることを示唆しています。

原文 (English)

SAAG: Structured Agent Assessment and Grounding

Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or satisfy a schema while choosing a agent for the wrong reason. Existing benchmarks collapse these distinctions into a single binary score, leaving practitioners unable to diagnose where agent calls fail. We propose SAAG a cascaded diagnostic framework that decomposes agent-calling evaluation into three sequential stages: registry conformance, structural completeness, and argument grounding, each producing interpretable stage-specific diagnostics. These diagnostics additionally enable iterative self-repair: on prediction failure, the stage-specific signal guides targeted correction without leaking ground-truth values. We evaluate this framework on a controlled benchmark derived from Glaive's function-calling dataset across registry sizes of 5, 10, and 15 agents using three local sub-4B-parameter models. Structured feedback consistently improves argument precision and reduces value hallucination relative to single-pass inference and uninformative binary feedback, while end-to-end F1 gains are modest and model-dependent. These results suggest that stage-decomposed diagnostic evaluation is a necessary lens for understanding and improving agent-calling reliability across model families and registry scales.

13:00 JST研究/論文

Phionyx: 構造化された状態管理と事前応答ガバナンスを備えた決定論的 AI ランタイム アーキテクチャ

我々は、AI エンジニアリングにガバナンス優先のアプローチを導入する、より広範な Echoism インタラクション フレームワークから派生した決定論的 AI ランタイム アーキテクチャである Phionyx を紹介します。つまり、大規模言語モデル (LLM) の出力を、直接的な決定ではなくノイズの多いセンサー測定値として扱うというものです。確率的エージェントとは異なり、Phionyx は決定論的な状態発展方程式によって支配される構造化された状態ベクトルを介して決定論的な状態発展を強制し、監査性とガバナンスを必要とするアプリケーションで再現可能な動作を可能にします。このアーキテクチャは、(1) 正規の 46 ブロック パイプラインを通じてノイズの多いセンサー測定を処理する決定論的評価カーネル、(2) 事前応答制御とアーキテクチャ上のプライバシー強制を提供する統合安全層、および (3) 影響を加重したキャッシュ削除を実装するセマンティックな時間ベースのメモリ システムの 3 つの層を統合します。単一インスタンスの導入における実験的検証では、ポストホック フィルタリングと比較して計算オーバーヘッドが約 31% 削減され (30% の安全でない入力比率、シミュレートされたコスト モデル)、LRU と比較して高価値のデータ保持が最大 24% 向上 (FIFO と比較して 72%、同じキャッシュ容量、ベンチマーク検証済み)、制御信号の差異がゼロで 100 回の繰り返し実行で決定的な実行が検証されたことが実証されています。 (ハッシュ検証済み)、単一インスタンスの展開テストでの計画外の再起動はゼロです(方法論と範囲については付録 C を参照)。このペーパーでは、アーキテクチャ、その分析構造、および範囲を絞った実験証拠を示します。分散展開またはマルチテナント展開への一般化は今後の課題です。

原文 (English)

Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance

We present Phionyx, a deterministic AI runtime architecture derived from the broader Echoism interaction framework that introduces a governance-first approach to AI engineering: treating large language model (LLM) outputs as noisy sensor measurements rather than direct decisions. Unlike probabilistic agents, Phionyx enforces deterministic state evolution via a structured state vector governed by deterministic state-evolution equations, enabling reproducible behavior in applications requiring auditability and governance. The architecture integrates three layers: (1) a deterministic evaluation kernel processing noisy sensor measurements through a canonical 46-block pipeline, (2) a unified safety layer providing pre-response control and architectural privacy enforcement, and (3) a semantic time-based memory system implementing impact-weighted cache eviction. Experimental validation on single-instance deployments demonstrates approximately 31% reduction in computational overhead vs. post-hoc filtering (at 30% unsafe input ratio, simulated cost model) and up to 24% improvement in high-value data retention vs. LRU (72% vs. FIFO, same cache capacity, benchmark-verified), deterministic execution verified across 100 repeated runs with zero variance in control signals (hash-verified), and zero unplanned restarts in single-instance deployment testing (see Appendix C for methodology and scope). This paper presents the architecture, its analytic structure, and scoped experimental evidence; generalization to distributed or multi-tenant deployments remains future work.

13:00 JSTロボティクス研究/論文

分散フィードバック制御によるドローン運動の角度安定化における積分微分方程式

この論文では、積分オペレータの形式で分散フィードバック制御を使用したドローン運動の角度安定化を提案します。この整数演算子のメモリには制限がない可能性があることを強調しておく必要があります。観察時間が長いと、制御オブジェクトの以前の状態に基づいてより良い制御を構築するための新たな可能性が開かれることは直感的に明らかです。制御における無制限のメモリには、積分微分方程式の研究に対する標準的なアプローチとは異なる特定のアプローチを作成する必要があります。この記事の目的の 1 つは、安定化におけるフィードバック制御を指定する積分演算子の無制限メモリの場合の積分微分方程式の安定性を研究できる、ある普遍的なアプローチを提案することです。私たちが提案するアプローチにより、積分微分方程式の研究を常微分方程式系の解析に縮小することができます。一般に、このようなシステムは無限数の方程式で構成されます。角度安定化の問題におけるいわゆる線形近似に関連して、積分制御では比較的単純な指数カーネルに制限し、有限数の方程式を含むシステムに到達します。この例では、より複雑なカーネル (指数関数カーネルの線形結合など) によって安定化機能が向上することが説明されています。積分微分方程式の指数関数的安定性に関する新たな予期せぬ結果が得られました。そしてそれらをドローン飛行の安定化に応用します。

原文 (English)

Integro-differential equations in angular stabilization of drone motion by distributed feedback control

In this paper, we propose angular stabilization of drone motion using distributed feedback control in the form of an integral operator. It should be stressed that the memory of this integral operator could be unbounded. It is intuitively clear that large length of the observation time open new possibilities to construct better control based on previous states of the control object. Unbounded memory in control requires the creation of a certain approach different from standard ones to the study of integro-differential equations. One of the goals of this article is to propose a certain universal approach that allows us to study the stability of integro-differential equations in the case of unbounded memory in the integral operator specifying the feedback control in stabilization. The approach we propose allows us to reduce the study of integro-differential equations to the analysis of systems of ordinary differential equations. In general, such systems can consist of an infinite number of equations. In relation to the so-called linear approximation in the problem of angle stabilization manages to limit itself to relatively simple exponential kernels in the integral control and arrive at a system with a finite number of equations. The examples explain that more complex kernels, for example, linear combinations of the exponential kernels, can enhance the stabilization capabilities. We obtain new unexpectable results on the exponential stability of integro-differential equations. Then we apply them to stabilization of drone flight.

13:00 JST研究/論文

MILP-Evo: MILP ソルバーの閉ループ完全自動設計

機械学習手法は、データ駆動型ポリシーが混合整数線形計画法 (MILP) ソルバーを高速化できることを示していますが、学習されたポリシーが外部予測子またはその他の不透明なモデルとして表されるため、そのようなアプローチの多くは依然として検査、適応、展開が困難です。対照的に、明示的ソルバー ロジックは理解と統合が容易ですが、通常はソルバーのフィードバックから学習するのではなく、手動で設計されます。私たちは、MILP ソルバー ロジックの自動設計を、エンドツーエンドのソルバー動作によって直接評価される実行可能なホワイトボックス コンポーネントに対する LLM ガイド付き閉ループ検索としてキャストできるかどうかを研究します。この目的を達成するために、PySCIPOpt を通じて実装される MILP ソルバー自動設計のための閉ループ プログラム進化フレームワークを提案し、カット セレクターと分岐ルールの結合設計でインスタンス化します。候補プログラムは繰り返し生成され、SCIP にロードされ、MILP インスタンスでの直接実行によって評価され、その結果得られるフィードバックによって、パフォーマンスに基づいた選択、対象を絞った修復、診断の反映、多様性を意識した母集団の維持が行われます。このメソッドは、標準ソルバー ワークフロー内で検査、変更、展開できる明示的ソルバー コンポーネントを出力します。 4 つのベンチマーク ファミリにわたって、LLM に基づくプログラムの進化により、いくつかの設定で競争力のあるドメインに特化したポリシーを発見できることがわかりました。

原文 (English)

MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers

Machine learning methods have shown that data-driven policies can accelerate mixed-integer linear programming (MILP) solvers, but many such approaches remain difficult to inspect, adapt, and deploy because the learned policy is represented as an external predictor or other opaque model. By contrast, explicit solver logic is easier to understand and integrate, but is usually hand-designed rather than learned from solver feedback. We study whether the automatic design of MILP solver logic can instead be cast as LLM-guided closed-loop search over executable white-box components evaluated directly by end-to-end solver behavior. To this end, we propose a closed-loop program evolution framework for MILP solver auto-design, implemented through PySCIPOpt, and instantiate it on the joint design of a cut selector and a branching rule. Candidate programs are iteratively generated, loaded into SCIP, and evaluated by direct execution on MILP instances, with the resulting feedback guiding performance-based selection, targeted repair, diagnostic reflection, and diversity-aware population maintenance. The method outputs explicit solver components that can be inspected, modified, and deployed within standard solver workflows. Across four benchmark families, we find that LLM-guided program evolution can discover competitive domain-specialized policies in several settings.

13:00 JSTLLM/生成AIClaude

精度とコストを超えて: 動的ワークロード向けのレイテンシーを考慮した LLM クエリ ルーティング

最新言語のクエリ ルーターは、応答の品質と金銭的コストのバランスをとるモデルに各クエリを割り当てることで、推論の効率を向上させます。ただし、現在のクエリ ルーターはレイテンシにほとんど依存せず、モデル インスタンスでのクエリによって発生する生成レイテンシは考慮されていません。実際には、レイテンシーはラウンドロビンや最短キューへの参加などの負荷分散ポリシーによって制御されることが多く、モデルの精度や推論コストは考慮されていません。クエリのレイテンシーをルーティングに組み込むことは、クエリのプロンプトの長さだけでなく、モデル インスタンスでの現在のプリフィルとデコードのワークロード、およびサービス提供フレームワークのスケジュールとバッチ ポリシーにも依存するため、困難です。サービング フレームワークでの自己回帰トークン バッチ処理をシミュレートし、クエリの最初のトークンまでの時間 (TTFT) を推定する軽量のレイテンシ推定ツールを設計します。このレイテンシ推定ツールをレイテンシ対応ルーターに組み込んで、モデル インスタンスにクエリを割り当てる際のレイテンシ、精度、コストを共同で最適化します。私たちの実験結果は、この統合最適化により、標準的な負荷分散アプローチと同じレイテンシーを維持しながら、精度、つまりコストのユーティリティが最大 40% 向上することが示されています。

原文 (English)

Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads

Modern language query routers improve inference efficiency by assigning each query to a model that balances response quality and monetary cost. However, current query routers are largely latency-agnostic and do not consider the generation latency experienced by queries at model instances. In practice, latency is often controlled by load-balancing policies such as round-robin or join-the-shortest-queue, which do not account for model accuracy or inference cost. Incorporating query latency into routing is challenging as it depends not only on the query's prompt length, but also on the current prefill and decode workload at the model instance and the scheduling and batching policy of the serving framework. We design a lightweight latency estimator that simulates autoregressive token batch processing in the serving framework and estimates the time-to-first-token (TTFT) of queries. We incorporate this latency estimator into a latency-aware router that jointly optimizes latency, accuracy, and cost when assigning queries to model instances. Our experimental results indicate that this joint optimization yields up to 40% improvement in accuracy--cost utility while maintaining the same latencies as standard load-balancing approaches.

13:00 JSTビジネス/資金調達研究/論文

再トレーニングを行わない方言間の一般化: MLIR のスキーマ導出制約付きデコーディングのベンチマークと評価

マルチレベル中間表現 (MLIR) は、最新の ML コンパイラ インフラストラクチャ (TensorFlow、JAX/StableHLO、PyTorch Inductor、IREE) の基礎を成していますが、コード LM 事前トレーニング コーパスには微量しか現れません。 MLIR は設計上も拡張可能です。新しい方言はアプリケーション ドメインごとに出荷されるため、方言ごとに微調整されたモデルは拡張できません。各方言の操作定義仕様 (ODS) から機械的に導出された推論時間事前確率が勾配ベースの適応の代わりに使用できるかどうかを尋ねます。まず、3 つの方言にわたる 4 つの自然言語から MLIR ベンチマーク (MLIR-Spec-150、Linalg-Spec-30、StableHLO-Spec-30、StableHLO-Held-Out-200) をリリースします。合計 410 の対象範囲内の NL から MLIR ペアに加え、25 のプログラムの文法外のストレス セットと手書きの n=30 の関数リファレンス セットが含まれており、以下で出荷されます。 Apache-2.0 と Gebru データシートおよび Croissant 1.0 メタデータ。次に、3 層のスキーマ派生制約スタックを構築します。OP シグネチャ上の CFG (C1)、ODS 抽出された型ラティスからの型ドメイン分割 (C2)、5 回の再試行拒否サンプリングを駆動する SSA スコープ バリデータ (C3) です。 arith+func+memref+linalg から StableHLO への移植には、新しい制約層コードは必要ありません。検証者のセマンティクスが構造的制約によって支配されている方言では、スキーマ由来の事前分布により、SmolLM2-1.7B は世代ごとの 8 ~ 25 倍の速度で 15B-34B コード LM と一致またはそれを超えます。linalg では、SmolLM2 は 80.0% verify-valid (3 シード平均、n=125) に達し、CodeLlama-34B を上回ります。 Granite-Code-34B および StarCoder2-15B は、重複しない CI で 21 ~ 44 パーセント ポイント減少しました。 arith+func とテンプレート化されたパラメトリック StableHLO-Held-Out-200 では、検証者のセマンティクスが構造ではなく属性値をオンにするため、同じベースラインが SLM に一致するか、それを上回ります。これらを非勝利セルとしてスコープします。ベンチマーク、デコーダー、プロンプトごとのすべての生成、再現性のある Docker イメージをリリースします。

原文 (English)

Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR

Multi-Level Intermediate Representation (MLIR) underlies modern ML compiler infrastructure (TensorFlow, JAX/StableHLO, PyTorch Inductor, IREE), yet appears only in trace amounts in code-LM pretraining corpora. MLIR is also extensible by design: new dialects ship per application domain, so a fine-tuned model per dialect does not scale. We ask whether inference-time priors derived mechanically from each dialect's Operation Definition Specification (ODS) can substitute for gradient-based adaptation. First, we release four natural-language-to-MLIR benchmarks across three dialects - MLIR-Spec-150, Linalg-Spec-30, StableHLO-Spec-30, and StableHLO-Held-Out-200 - totaling 410 in-scope NL-to-MLIR pairs, plus a 25-program out-of-grammar stress set and a hand-authored n=30 functional reference set, shipped under Apache-2.0 with Gebru datasheets and Croissant 1.0 metadata. Second, we build a three-layer schema-derived constraint stack: a CFG over op signatures(C1), type-domain splits from an ODS-extracted type lattice (C2), and an SSA-scope validator driving five-retry rejection sampling (C3). Porting from arith+func+memref+linalg to StableHLO required no new constraint-layer code. On dialects whose verifier semantics are dominated by structural constraints, schema-derived priors let SmolLM2-1.7B match or exceed 15B-34B code LMs at 8-25x the per-generation speed: on linalg, SmolLM2 reaches 80.0% verify-valid (three-seed mean, n=125), beating CodeLlama-34B, Granite-Code-34B, and StarCoder2-15B by 21-44 percentage points with non-overlapping CIs. On arith+func and on the templated parametric StableHLO-Held-Out-200, where verifier semantics turn on attribute values rather than structure, the same baselines match or beat the SLM; we scope these as non-win cells. We release benchmarks, decoder, all per-prompt generations, and a reproducibility Docker image.

13:00 JSTLLM/生成AIエージェントハードウェア/半導体

LLM ベースのマルチエージェント システムにおける貢献帰属のためのセマンティック協力ゲーム

貢献の帰属は、最終出力が複数のエージェント、メッセージ交換、および順序付けられたワークフローの依存関係を通じて生成される LLM ベースのマルチエージェント システムの中心的な問題となっています。既存のアトリビューション方法は、多くの場合、エージェントを削除したり、変更されたエージェントのサブセット間でスコアの変化を比較したりするなど、反事実の評価に依存しています。言語媒介のワークフローでは、これらの方法ではモデルの呼び出しを繰り返す必要があり、高い分散が導入され、エージェントがタスク関連情報を生成、保存、変換するための中間的な意味論的状態を明示的に取得できません。我々は、実現された言語フローを意味生成ハイパーグラフとして表現し、この構造上でエージェントレベルの意味価値関数を誘導するフレームワークであるSemantic Cooperative Games (SCG)を提案します。セマンティック サポート ロジックへの寄与を割り当てるためのセマンティック シャプレイ値 (SSV) を定義し、セマンティック ハイパーグラフを構築し、最小限のセマンティック サポートを回復し、ブール吸収を適用し、エージェント サブセットを再実行せずに SSV を計算する単一軌道アルゴリズムである SLIC を導入します。標準セットベースで完全に観測可能で次数依存性のない条件下で、SSV が古典的な Shapley 値にまで減少することを証明します。これらの条件を満たす医療ベンチマークでは、SLIC はモンテカルロ Shapley ベースラインとの高い一貫性を維持しながら、計算コストを 93.3% 削減します。より一般的なマルチロール ワークフローでは、SSV は摂動によるスコア低下プロファイルと連携し、セマンティックな寄与と失敗の影響が発散するケースを明らかにします。全体として、SLIC は、複雑な LLM ベースのマルチエージェント システムに対して、高速で反事実がなく、解釈可能な帰属方法を提供します。

原文 (English)

Semantic Cooperative Games for Contribution Attribution in LLM-Based Multi-Agent Systems

Contribution attribution has become a central problem in LLM-based multi-agent systems, where final outputs are produced through multiple agents, message exchanges, and ordered workflow dependencies. Existing attribution methods often rely on counterfactual valuation, such as removing agents or comparing score changes across altered agent subsets. In language-mediated workflows, these methods require repeated model calls, introduce high variance, and do not explicitly capture the intermediate semantic states through which agents produce, preserve, and transform task-relevant information. We propose Semantic Cooperative Games (SCG), a framework that represents a realized language flow as a semantic generation hypergraph and induces an agent-level semantic value function on this structure. We define the Semantic Shapley Value (SSV) to allocate contribution over semantic support logic, and introduce SLIC, a single-trajectory algorithm that constructs the semantic hypergraph, recovers minimal semantic supports, applies Boolean absorption, and computes SSV without rerunning agent subsets. We prove that SSV reduces to the classical Shapley value under standard set-based, fully observable, and no-order-dependence conditions. On a medical benchmark satisfying these conditions, SLIC reduces computation cost by 93.3% while remaining highly consistent with a Monte Carlo Shapley baseline. In more general multi-role workflows, SSV aligns with perturbation-induced score-drop profiles and exposes cases where semantic contribution and failure impact diverge. Overall, SLIC provides a fast, counterfactual-free, and interpretable attribution method for complex LLM-based multi-agent systems.

13:00 JST研究/論文DeepSeek

PEARL: 自然言語からのソルバーインザループ対話型最適化モデリング

最適化モデリングは、多くの場合自然言語で記述される現実世界の意思決定問題を、正式な数学的公式と実行可能なソルバー コードに変換するプロセスです。大規模言語モデルの最近の進歩により、このプロセスの自動化が期待できるようになりましたが、既存のアプローチのほとんどはワンショットのままです。モデルは、実行したり、ソルバーのフィードバックに基づいて条件付けしたり、エラーを繰り返し修正したりすることなく、一度定式化を生成します。これは、本質的にインタラクティブであり、解決、デバッグ、修正のサイクルを繰り返すことで進行する現実世界の最適化モデリングとは大きく対照的です。このループ内で Python 実行と数学的プログラミング ソルバーを使用する、対話型の最適化モデリングのためのシステムである PEARL を紹介します。 PEARL は、固定された修復ワークフローに依存するのではなく、部分モデルをいつテストするか、ソルバー診断から修正する方法、およびいつ停止するかを学習します。これは、中間実行結果、実現可能性シグナル、およびソリューション チェックを使用して、最終化の前に定式化とソルバー コードの両方を改善する、マルチターン ツール統合設定で動作します。 PEARL は、さまざまな最適化ベンチマークにわたって、強力なワンショットおよびツール拡張ベースラインに比べて検証済みの解決率を大幅に向上させます。特に、当社の PEARL-Qwen3-\textbf{4B} モデルは、最適化モデリング タスクにおけるマクロ平均精度とミクロ平均精度の両方において、はるかに大規模な DeepSeek-V3.2-\textbf{685B} モデルよりも優れています。

原文 (English)

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language

Optimization modeling is the process of translating real-world decision problems, often described in natural language, into formal mathematical formulations and executable solver code. While recent advances in large language models have shown promise in automating this process, most existing approaches remain one-shot: a model produces a formulation once, without executing it, conditioning on solver feedback, or iteratively revising errors. This stands in sharp contrast to real-world optimization modeling, which is inherently interactive and proceeds through repeated solve-debug-revise cycles. We introduce PEARL, a system for interactive optimization modeling that uses Python execution and mathematical programming solvers inside this loop. Rather than relying on a fixed repair workflow, PEARL learns when to test partial models, how to revise from solver diagnostics, and when to stop. It operates in a multi-turn tool-integrated setting where intermediate execution results, feasibility signals, and solution checks are used to improve both formulations and solver code before finalization. Across diverse optimization benchmarks, PEARL substantially improves verified solve rates over strong one-shot and tool-augmented baselines; notably, our PEARL-Qwen3-\textbf{4B} model outperforms the much larger DeepSeek-V3.2-\textbf{685B} in both macro- and micro-averaged accuracy on optimization modeling tasks.

13:00 JST研究/論文

S2T-RLHF: 安定した優先ベースの RLHF のための階層的クレジット割り当て

嗜好ベースの報酬モデルを使用したヒューマン フィードバックからの強化学習 (RLHF) は、多くの場合、不安定なトレーニング ダイナミクスを示します。主な要因は、標準 RLHF が単一のシーケンス レベルのスカラー報酬に依存しており、これがトークン レベルのポリシー更新に伝播され、応答内のクレジットの割り当てが本質的に曖昧なままになっていることです。最近の研究では、報酬をより高密度のトークンレベルの監視に洗練することでこの問題に対処しようとしていますが、多くの場合、よりきめの細かいクレジット割り当てにより最適化が向上するという暗黙の仮定に依存しています。我々は、この仮定が不完全であると主張します。嗜好信号にノイズが多く、応答レベルでのみ定義されている場合、報酬の微調整が過度に細かい場合、報酬の不確実性が増幅され、学習が不安定になる可能性があります。この問題に対処するために、最大の割り当て精度よりも安定性を重視した報酬設計を重視した、階層的なクレジット割り当ての粒度を意識した原則を提案します。この原則に基づいて、文は自然な中間粒度として機能し、意味の一貫性とトークンレベルのノイズに対する堅牢性のバランスをとります。この考え方に基づいて、S2T-RLHF を紹介します。この文からトークンへの報酬分解フレームワークは、最初にシーケンス レベルの優先報酬を文全体に割り当て、次に報酬モデルの再トレーニングやトークン レベルの監視を行わずに、各文内で制限されたトークン レベルの改良を適用します。複数のデータセットと最適化設定にわたる実験では、S2T-RLHF が競争上の優先順位の調整を維持しながらトレーニングの安定性と堅牢性を向上させることが示されています。

原文 (English)

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

Reinforcement learning from human feedback (RLHF) with preference-based reward models often exhibits unstable training dynamics. A key contributing factor is that standard RLHF relies on a single sequence-level scalar reward, which is propagated to token-level policy updates and leaves credit assignment within a response inherently ambiguous. Recent work has attempted to address this issue by refining rewards into denser token-level supervision, often relying on the implicit assumption that finer-grained credit assignment improves optimization. We argue that this assumption is incomplete: when preference signals are noisy and only defined at the response level, overly fine-grained reward refinement can amplify reward uncertainty and destabilize learning. To address this problem, we propose a granularity-aware principle for hierarchical credit assignment, emphasizing stability-oriented reward design rather than maximal allocation precision. Under this principle, sentences serve as a natural intermediate granularity, balancing semantic coherence with robustness to token-level noise. Guided by this view, we introduce S2T-RLHF. This sentence-to-token reward decomposition framework first allocates sequence-level preference rewards across sentences and then applies bounded token-level refinement within each sentence, without reward-model retraining or token-level supervision. Experiments across multiple datasets and optimization settings show that S2T-RLHF improves training stability and robustness while maintaining competitive preference alignment.

13:00 JSTLLM/生成AI

信頼できる LLM 推論のための確率的概念を意識したステアリング

大規模言語モデル (LLM) の推論時介入手法であるステアリング ベクトル (SV) は、推論中の中間アクティベーションに概念固有の方向ベクトルを追加することで生成プロセスをガイドします。しかし、既存の SV 手法では、解釈可能性やきめ細かい制御を損なう表現一貫性のない動作が頻繁に発生します。これは主に、これまでの研究がセマンティック アラインメントの連続スペクトルを捕捉できない離散クラスタリング メトリクスを採用しながら、バイナリのポジティブ/ネガティブ ステアリング評価に焦点を当てていたためです。この研究では、LLM 推論のための確率的概念認識ステアリング (PCS) フレームワークを紹介します。 PCS は、元のタスク能力を維持しながら、コンセプト主導のステアリング ベクトル検索と確率的強度キャリブレーションを通じて、制御可能な安全指向のセマンティック バイアスを提供します。

原文 (English)

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

Steering vectors (SVs), an inference-time intervention technique for large language models (LLMs), guide the generation process by adding a concept-specific direction vector to intermediate activations during inference. However, existing SV methods frequently yield representation-incoherent behaviors that undermine interpretability and fine-grained control, largely because prior work has focused on binary positive-negative steering evaluation while employing discrete clustering metrics that fail to capture the continuous spectrum of semantic alignment. In this work, we present the Probabilistic Concept-Aware Steering (PCS) framework for LLM inference. PCS preserves original task competence while providing controllable, safety-oriented semantic bias through concept-driven steering-vector retrieval and probabilistic strength calibration.

13:00 JST研究/論文

FindStatBench: 組み合わせコード合成における大規模言語モデルの評価

組み合わせコード合成に関する大規模な言語モデルを評価するための実行ベンチマークである FindStatBench を紹介します。 FindStat から構築されており、24 のコレクションと 552M の非表示インスタンスにわたる 2,329 のタスクが含まれており、オブジェクトを整数にマッピングする統計合成と、オブジェクトをオブジェクトにマッピングするマップ合成をカバーします。各タスクには数学的な説明と、最大 5 つの公開入出力例が記載されています。モデルは、取得、ツールの使用、実行フィードバック、投票、再ランキングを行わずに 1 つの Python ソルブ関数を発行する必要があります。送信は、保持された組み合わせオブジェクトに対する正確なサンドボックス実行によってスコア付けされます。 1 つの推論プロバイダーを通じて提供される 4 つのクローズドソース実稼働モデルと 7 つのオープンウェイト モデルの 11 のシステムを評価します。 FindStatBench では 3 つの主なパターンが明らかになります。まず、最も強力なオープン ソース システムとクローズド ソース システムは、インスタンス精度が 1 pp 以内に収束し、すべてのシステムに対するオラクルと 1 つの中間層モデルからの 5 方向サンプリングの両方で、限られたタスク精度の向上しか得られません。第 2 に、例は害を及ぼす可能性があります。いくつかの古典的な全単射は、例が 0 つあれば完全に解決されますが、5 つの例のプロンプトでは失敗します。第三に、コードが出力される前に推論によって目に見える応答が使い果たされる可能性があるため、一部の失敗は出力バジェットの仕組みを反映しています。全体として、統計合成はマップ合成よりもはるかに簡単で、一部のコレクションはゼロに近いままであり、長いプロンプトは精度の急激な崖を引き起こし、正確なシンボリック ルールの誘導は脆弱なままです。

原文 (English)

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis

We introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis. Built from FindStat, it contains 2,329 tasks across 24 collections and 5.52M hidden instances, covering statistic synthesis, which maps objects to integers, and map synthesis, which maps objects to objects. Each task gives a mathematical description and at most five public input-output examples; a model must emit one Python solve function with no retrieval, tool use, execution feedback, voting, or reranking. Submissions are scored by exact sandboxed execution on held-out combinatorial objects. We evaluate eleven systems: four closed-source production models and seven open-weight models served through one inference provider. FindStatBench reveals three main patterns. First, the strongest open- and closed-source systems converge within 1 pp instance accuracy, and both an oracle over all systems and five-way sampling from one mid-tier model yield only limited task-accuracy gains. Second, examples can hurt: several classical bijections are solved perfectly with zero examples but fail under five-example prompts. Third, some failures reflect output-budget mechanics, as reasoning can exhaust the visible response before code is emitted. Overall, statistic synthesis is much easier than map synthesis, some collections remain near-zero, long prompts cause a sharp accuracy cliff, and exact symbolic rule induction remains brittle.

13:00 JSTLLM/生成AIエージェント

JSON では不十分な場合: スキーマ制約のある LLM 順序付けエージェントのセマンティックな信頼性

LLM エージェントはトランザクション コンパイラとして使用されることが増えています。ユーザーは自然言語で意図を述べ、モデルは API が実行できる構造化オブジェクトを生成します。 JSON スキーマとプロバイダー レベルの構造化出力モードは、大規模な解析失敗を取り除くため便利ですが、オブジェクトが安全で忠実なトランザクションであるかどうかを単独で判断するわけではありません。 OrderBench は、構文の有効性、スキーマの有効性、ステータスの決定、正確な項目のセマンティクス、制約の保持、および安全でない受け入れを分離する、レストラン注文エージェント向けの決定論的ベンチマークです。プロンプト専用モードおよび JSON スキーマ モードで 4 つのオープン モデルに対する Nebius Token Factory の 2,400 回の呼び出しを通じて、スキーマが有効な出力には依然として大きなセマンティック エラー率が含まれる可能性があることがわかりました。最も強力なモデルでは、どちらのモードも 100% のスキーマ妥当性を達成していますが、セマンティックな成功率は 80% 近くにとどまっています。弱いモデルでは、スキーマ有効な安全でない受け入れが 2 桁で発生します。その結果は具体的なエンジニアリング上の警告です。構造化された出力は必要なインターフェイス層であり、ドメイン検証やフェールクローズ実行の代替品ではありません。

原文 (English)

When JSON Is Not Enough: Semantic Reliability of Schema-Constrained LLM Ordering Agents

LLM agents are increasingly used as transaction compilers: a user states an intent in natural language, and the model emits a structured object that an API can execute. JSON Schema and provider-level structured-output modes are useful because they remove a large class of parse failures, but they do not by themselves decide whether the object is a safe, faithful transaction. We introduce OrderBench, a deterministic benchmark for restaurant ordering agents that separates syntactic validity, schema validity, status decisions, exact item semantics, constraint preservation, and unsafe acceptances. Across 2,400 Nebius Token Factory calls to four open models in prompt-only and JSON-schema modes, we find that schema-valid output can still have large semantic error rates. In the strongest model, both modes achieve 100% schema validity, yet semantic success remains near 80%; in weaker models, schema-valid unsafe acceptances occur in double digits. The result is a concrete engineering warning: structured output is a necessary interface layer, not a substitute for domain verification and fail-closed execution.

13:00 JST研究/論文

ProbSPARQL: 多次元の不確実な数値データを使用したナレッジ グラフのクエリ

SFB 1574 Circular Factory は、返品された製品に関するデータを統合するための共有ナレッジ グラフ インフラストラクチャを構築しています。中心的な課題は、サーキュラーファクトリーのデータには、(i) センサーに由来する、またはセンサーベースの測定から導出される数値測定値が含まれており、(ii) 多くの場合多次元であり、(iii) 本質的に不確実である一方、下流のトリアージ、検証、信頼性モデリング、および再組み立て計画モジュールにはクエリ可能な不確実性表現が必要であることです。現在の RDF および SPARQL テクノロジには、このような不確実な数値測定データの調和されたクエリと分析に対するネイティブ サポートが不足しています。このギャップに対処するために、このインフラストラクチャの初期段階のクエリ層パイロットとして開発された上位互換性のある SPARQL 拡張機能である ProbSPARQL を紹介します。 ProbSPARQL は、不確実な数値を確率変数としてモデル化し、その分布が確率的 RDF リテラル データ型によってエンコードされ、分布を意識した式、確率的フィルター、発散ベースの結合をサポートします。 ProbSPARQL を Apache Jena ARQ に実装し、Fuseki 互換の実行レイヤーを通じて公開します。 GMM でエンコードされた不確実性とヒストグラムベースの経験的粗さ分布をカバーするプロジェクト由来の測定フラグメントを使用して実際のデータの適用性を評価し、最大 5,000 個のアングル グラインダー インスタンスと 150 万のトリプルを備えた制御されたオントロジー準拠のベンチマークでスケーラビリティを個別に評価します。結果は、実現可能なエンジン内実行、アプリケーション層の後処理よりもフィルター プッシュダウンの高速化、分岐結合決定戦略間のレイテンシー精度のトレードオフを示しています。

原文 (English)

ProbSPARQL: Querying Knowledge Graphs with Multi-dimensional, Uncertain Numeric Data

The SFB 1574 Circular Factory is building a shared knowledge graph infrastructure for integrating data about returned products. A central challenge is that circular-factory data include numeric measurements that (i) originate from sensors or are derived from sensor-based measurements, (ii) are frequently multi-dimensional, and (iii) are inherently uncertain, while downstream triage, validation, reliability-modeling, and reassembly-planning modules require queryable uncertainty representations. Current RDF and SPARQL technologies lack native support for harmonized querying and analysis of such uncertain numeric measurement data. To address this gap, we present ProbSPARQL, an upward-compatible SPARQL extension developed as an early-stage query-layer pilot for this infrastructure. ProbSPARQL models uncertain numeric values as random variables whose distributions are encoded by probabilistic RDF literal datatypes, and supports distribution-aware expressions, probabilistic filters, and divergence-based joins. We implement ProbSPARQL on Apache Jena ARQ and expose it through a Fuseki-compatible execution layer. We assess real-data applicability using project-derived measurement fragments covering GMM-encoded uncertainty and histogram-based empirical roughness distributions, and evaluate scalability separately on controlled ontology-conformant benchmarks with up to 5,000 angle-grinder instances and 1.5M triples. The results show feasible in-engine execution, filter-pushdown speedups over application-layer post-processing, and latency-accuracy trade-offs among divergence-join decision strategies.

13:00 JST研究/論文

立場: AI/ML ディープフェイク研究は、AI が生成した同意のない親密な画像と一致していません (AIG-NCII)

AI 生成の非合意親密画像 (AIG-NCII) は、一般に「ディープフェイク」と呼ばれる、AI 生成メディアに関する AI/ML 文献で適切に扱われていません。ディープフェイクに関する研究は現在、その認識論的な害、つまり真実と信憑性に関する害に焦点を当てているが、これは性的な画像を含む生成型 AI の悪用という支配的な現実とは乖離している。私たちは、引用数の多い作品のランドスケープ分析を実施し、ディープフェイクに対処する技術的介入がほぼ完全に AIG-NCII を無視し、研究エコシステムを真正性検出ツールに限定していることを実証しました。この意見書では、既存の介入は詐欺や詐欺などの視聴者中心の認識論的危害には対処しているが、AIG-NCIIなどの被験者中心の尊厳の危害は無視されていると主張します。画像が合成であると知っても被写体への害は軽減されず、場合によっては悪化する可能性があることを説明します。最後に、被験者中心の危害を考慮するための脅威モデルの更新や、AI の安全性研究における AIG-NCII への取り組みなど、この分野を再調整するための推奨事項を提供します。最後に、研究者は、被験者と研究者の両方に安全ガードレールを導入し、性的暴力防止において当該分野の専門家とのパートナーシップを確立する場合にのみ、このリスクの高い分野に取り組むべきであると警告します。

原文 (English)

Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)

AI-generated non-consensual intimate imagery (AIG-NCII) is not adequately addressed in AI/ML literature regarding AI-generated media, commonly referred to as "deepfakes". While research on deepfakes currently focuses on its epistemic harms -- or harms relating to truth and authenticity -- this is misaligned with the dominant reality of generative AI abuse involving sexualized imagery. We conduct a landscape analysis of highly-cited works to demonstrate that technical interventions addressing deepfakes almost entirely ignore AIG-NCII, limiting the research ecosystem to authenticity detection tools. In this position paper, we argue that existing interventions address viewer-centric epistemic harms, such as fraud or scams, but ignore subject-centric dignity harms, such as AIG-NCII. We illustrate that knowing an image is synthetic does not mitigate harms to subjects and may, in some cases, even exacerbate them. We conclude by offering recommendations to realign the field, including updating threat models to consider subject-centric harms and addressing AIG-NCII in AI safety research. Finally, we caution that researchers should only engage in this high-risk domain if they implement safety guardrails for both subjects and researchers and establish partnerships with domain experts in sexual violence prevention.

13:00 JSTLLM/生成AI

MUX: 多重化されたトークンによる継続的推論

言語モデルは、自然言語で中間推論ステップを明確にすることで複雑な問題を解決します。このプロセスは効果的ではありますが、計算上のボトルネックがあります。各推論ステップは 1 つのサブワードのみを伝え、その多くは計算を実行する代わりに思考を表現するのに費やされます。我々は、離散推論を潜在空間内の連続多重化トークンに蒸留することに基づく、高帯域幅でコンパクトな推論のための簡単な方法である MUX を提案します。ここで、各潜在トークンは、離散推論サブワードのスパンの重み付き線形重ね合わせ (多重化) を表すようにトレーニングされます。この重ね合わせは構築によりロスレスであり、スパンは完全に復元できます (逆多重化)。我々は、適切な幾何学的減衰などの単純な位置依存の重み付けが可逆多重化をサポートし、それによって潜在的な崩壊によって引き起こされるショートカット動作を防止することを証明します。さらに、多重推論が探索を必要とする問題において並列探索を実行できることを示します。 4 つの言語モデルにわたる 32 の評価設定にわたって、MUX は強力な潜在推論ベースラインを上回りました。アブレーション分析とプローブ分析により、学習された潜在トークンが忠実で解釈可能な推論をエンコードしていることがさらにわかります。我々の結果は、局所学習ターゲットとしての可逆重ね合わせが、強力で効率的な潜在的連続推論を達成するための十分な条件を構成することを示唆しています。

原文 (English)

MUX: Continuous Reasoning via Multiplexed Tokens

Language models solve complex problems by articulating intermediate reasoning steps in natural language. While effective, this process is computationally bottlenecked: each reasoning step conveys only a single subword, and many are spent expressing a thought instead of carrying out computation. We propose MUX, a simple method for high-bandwidth and compact reasoning based on distillation of discrete reasoning into continuous multiplexed tokens in a latent space. Here, each latent token is trained to represent a weighted linear superposition (multiplexing) of a span of discrete reasoning subwords, where this superposition is lossless by construction and the span can be fully recovered (demultiplexing). We prove that simple position-dependent weightings, such as suitable geometric decay, support lossless multiplexing, which in turn prevents shortcut behaviors caused by latent collapse. We further show that multiplexed reasoning can perform parallel exploration in problems that require search. Across 32 evaluation settings spanning four language models, MUX outperforms strong latent reasoning baselines. Ablation and probing analyses further show that the learned latent tokens encode faithful and interpretable reasoning. Our results suggest that lossless superposition as local learning targets constitutes a sufficient condition for achieving strong and efficient latent continuous reasoning.

13:00 JSTLLM/生成AIエージェント

2 エージェント LLM リレーにおける状態圧縮: 制約の保存に関する閉じられた世界の研究

長時間実行されるラージ言語モデル (LLM) ベースのエージェントは、多くの場合、監査、削除、数値計算を含む大規模な中間トレースを蓄積します。実際には、この状態は下流の意思決定ステップに渡す前に圧縮されるため、情報のボトルネックが生じ、小さな省略によって厳密な数値またはカテゴリの制約が破られる可能性があります。この論文では、2 つの LLM エージェントを使用したクローズドワールドの旅行計画リレーにおけるハンドオフ圧縮を評価します。研究者は、50 のゴール インスタンスのホテルとフライトの固定在庫を監査し、予約者は、在庫は差し控えた状態で、ゴールとハンドオフ ペイロードのみを使用してホテルとフライトのペアを選択します。圧縮なし、ナラティブ要約、スキーマ制約のある JSON 抽出、埋め込みベースのプルーニングという 4 つのハンドオフ条件を比較します。固定在庫を徹底的に列挙することで、正確で実行可能な最適なラベルが提供されます。結果は、ハンドオフ表現が小規模な意思決定モデルの下で下流の実現可能性に強く影響することを示しています。 JSON 抽出は 0.96 という最高の実現可能性精度を達成しますが、ナラティブ要約は最小の圧縮ハンドオフ ペイロードを生成するにもかかわらず、実現可能性が 0.48 に低下します。埋め込みベースの枝刈りは、追加の生成圧縮呼び出しなしで、0.88 の実現可能性に関する非圧縮コントロールと一致します。これらの発見は、制約チェックが簡潔さのみに依存するのではなく、構造化された監査可能なハンドオフ表現から恩恵を受けることを示しています。

原文 (English)

State Compression in Two-Agent LLM Relays: A Closed-World Study of Constraint Preservation

Long-running Large Language Model (LLM)-based agents often accumulate large intermediate traces containing audits, eliminations, and numeric calculations. In practice, this state is compressed before handing it to a downstream decision step, creating an information bottleneck in which small omissions can break strict numeric or categorical constraints. This paper evaluates hand-off compression in a closed-world travel-planning relay with two LLM agents. A Researcher audits a fixed inventory of hotels and flights for 50 goal instances, and a Booker selects a hotel--flight pair using only the goal and the hand-off payload, with the inventory withheld. We compare four hand-off conditions: no compression, narrative summarization, schema-constrained JSON extraction, and embedding-based pruning. Exhaustive enumeration over the fixed inventory provides exact feasible and optimal labels. Results show that hand-off representation strongly affects downstream feasibility under a small decision model. JSON extraction achieves the highest feasibility accuracy at 0.96, while narrative summarization, despite producing the smallest compressed hand-off payload, degrades feasibility to 0.48. Embedding-based pruning matches the uncompressed control on feasibility at 0.88 without an additional generative compression call. These findings indicate that constraint checking benefits from structured and auditable hand-off representations rather than relying on brevity alone.

13:00 JST研究/論文GPT / ChatGPT

小さな言語モデルの算術微調整のための構造化された合成推論データ

小規模な言語モデルはローカル展開には魅力的ですが、多くの場合、複数ステップの算術推論に苦労します。私たちは、構造化された合成推論データが消費者ハードウェアの制約下でこの動作を改善できるかどうかを研究します。 GSM8K から始めて、私たちは GPT-5-mini を使用して、自然言語の解法トレース、軽いソクラテス風の手がかり、構造の変化、および無関係な気を散らすコンテキストを組み合わせて、小学校の算数の単語問題のバリエーションの 21,250 例のコーパスを生成しました。次に、コンシューマ ハードウェア (Apple M4、16 GB RAM) 上の LoRA を使用して Qwen3-0.6B と Qwen3-1.7B を微調整しました。 GSM8K での完全一致精度は、Qwen3-0.6B では 36.5% から 49.1% に、Qwen3-1.7B では 53.5% から 66.5% に向上しました。 Qwen3-1.7B では、関連する算術ベンチマークへの移行がより強力で、基本モデルの 54.4% と 45.3% と比較して、MultiArith で 98.9%、SVAMP で 73.0% に達しました。定性分析によると、微調整されたモデルは推論トレースが短くなり、算術エラーや気を散らすものを使用したエラーが少なくなり、自己無撞着サンプリングからより一貫して恩恵を受けることがわかります。これらの結果は、低コストの合成データ設計により、小規模な言語モデルにおける算術適応を大幅に改善できることを示しています。この介入はソクラテス式の手がかりと他のデータ設計の選択肢を組み合わせているため、得られた成果をソクラテス式のガイダンス単独の因果関係のテストとしてではなく、構造化された合成推論データの証拠として解釈します。

原文 (English)

Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models

Small language models are attractive for local deployment, but they often struggle with multi-step arithmetic reasoning. We study whether structured synthetic reasoning data can improve this behaviour under consumer-hardware constraints. Starting from GSM8K, we generated a 21,250-example corpus of grade-school arithmetic word-problem variants using GPT-5-mini, combining natural-language solution traces, light Socratic-style cues, structural variation, and irrelevant distractor context. We then fine-tuned Qwen3-0.6B and Qwen3-1.7B with LoRA on consumer hardware (Apple M4, 16 GB RAM). Exact-match accuracy on GSM8K improved from 36.5% to 49.1% for Qwen3-0.6B and from 53.5% to 66.5% for Qwen3-1.7B. For Qwen3-1.7B, transfer to related arithmetic benchmarks was stronger, reaching 98.9% on MultiArith and 73.0% on SVAMP, compared with 54.4% and 45.3% for the base model. Qualitative analysis suggests that fine-tuned models produce shorter reasoning traces, make fewer arithmetic and distractor-use errors, and benefit more consistently from self-consistency sampling. These results show that low-cost synthetic data design can materially improve arithmetic adaptation in small language models. Because the intervention combines Socratic-style cues with other data-design choices, we interpret the gains as evidence for structured synthetic reasoning data rather than as a causal test of Socratic guidance alone.

13:00 JSTLLM/生成AIハードウェア/半導体

フェンス: LLM アプリケーションに特化した SLM ガードレール

クローズドソースの大規模言語モデル (LLM) を使用する現実世界のアプリケーションには、基本的なコンテンツ フィルターを超える高度な安全対策が必要です。有害性やバイアスなどのコンテンツ管理フィルターは比較的標準的な定義を持っていますが、幻覚、トピックのドリフト、行動の逸脱などのアプリケーション固有のガードレールはモデル化がより難しく、ユースケースによって異なる場合があります。さらに、データの不足と注釈のコストにより、特殊なガードレールの作成とテストのプロセスが困難になります。この研究では、LLM アプリケーションの特殊なガードレールとして、合成データでトレーニングされた小型言語モデル (SLM) を使用することを提案します。私たちは、敵対的生成ネットワーク (GAN) の設計にヒントを得た新しい合成データ生成手法を導入して、ユースケース固有のガードレール情報をエンコードし、特殊なガードレールとして機能するように SLM をトレーニングするために使用できる高品質の合成データ サンプルを生成します。私たちの実験では、高品質の合成データでトレーニングされた SLM ガードレールがプロンプトベースの LLM ガードレールよりもパフォーマンスが向上することが実証されました。

原文 (English)

Fence: Specialized SLM Guardrails for LLM Applications

Real-world applications that use closed-source large language models (LLMs) need advanced safety measures that go beyond the basic content filters. Content moderation filters such as toxicity and bias have relatively standard definitions where as application specific guardrails like hallucination, topic drift and behaviour deviation are more difficult to model and can vary by use case. Additionally, data scarcity and annotation costs, make the process of creating and testing specialized guardrails challenging. In this work, we propose using Small Language Models (SLMs) trained on synthetic data as specialized guardrails for LLM applications. We introduce a novel synthetic data generation method inspired by the design of Generative Adversarial Networks (GANs) to generate high quality synthetic data samples which can be used to train SLMs to encode use case specific guardrail information and hence function as specialized guardrails. Our experiments demonstrate that SLM guardrails trained on high quality synthetic data show performance gains over prompt based LLM guardrails.

13:00 JSTLLM/生成AI

LLM 群集の知恵: 言語モデル アンサンブルにおける集約と汚染

群衆の知恵、つまり個人間の判断を総合すると、最も優れた個人よりも優れた結果をもたらすことが多いという発見は、人間の予報士を対象に広範囲に研究されてきました。 「群衆」が大規模言語モデル (LLM) で構成されている場合に同じ現象が現れるかどうかは、理論的意味と実践的意味の両方を伴う未解決の問題です。 254 のバイナリ予測市場質問について 15 の LLM から確率推定値を導き出し、古典的な集計方法と学習された集計方法を評価しました。学習されたアグリゲーター (多層パーセプトロンとロジスティック回帰) は、すべての個別モデルや古典的な手法を上回りました。ロジスティック回帰はニューラル ネットワークと一致することがわかり、学習された集計の利点は、非線形相互作用ではなく、多様なモデル出力の線形結合の学習から得られることを示唆しています。ニューラル ネットワークの学習されたマッピングに適用されたシンボリック回帰により、純粋なモデル不一致信号がパレート フロンティア上で最も複雑性の低い有用な式として復元され、この解釈がさらに裏付けられました。トレーニング カットオフの汚染が蔓延した混乱であることが判明しました。フロンティア クラウド モデルと小規模なローカル モデルの間の見かけの能力差は、すべてのモデルのトレーニング カットオフ後に解決される質問のクリーンなサブセットで 35.8% から 8.9% に崩壊し、個々のモデルのランキングは中程度の安定性しか示しませんでした。予測市場が各モデルのトレーニング カットオフで評価された場合でも、LLM の精度は大幅に低いままであり、集合的な情報集約における真のギャップを示しています。これらの発見は、LLM 群衆が群衆の知恵効果を示す可能性があるが、信頼できる評価には汚染のない評価が不可欠であることを示唆しています。

原文 (English)

Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively studied with human forecasters. Whether the same phenomenon emerges when the ``crowd'' consists of large language models (LLMs) is an open question with both theoretical and practical implications. We elicited probability estimates from 15 LLMs on 254 binary prediction market questions and evaluated classical and learned aggregation methods. Learned aggregators -- a multilayer perceptron and a logistic regression -- outperformed all individual models and classical methods. The logistic regression was found to match the neural network, suggesting that the benefit of learned aggregation derives from learning a linear combination of diverse model outputs rather than from nonlinear interactions. Symbolic regression applied to the neural network's learned mapping recovered a pure model-disagreement signal as the lowest-complexity useful formula on the Pareto frontier, further supporting this interpretation. Training cutoff contamination proved a pervasive confound: the apparent capability gap between frontier cloud models and smaller local models collapsed from 35.8% to 8.9% on a clean subset of questions resolving after all models' training cutoffs, and individual model rankings showed only moderate stability. Even when the prediction market is evaluated at each model's training cutoff, LLMs remained substantially less accurate, indicating a genuine gap in collective information aggregation. These findings suggest that LLM crowds can exhibit wisdom-of-crowds effects, but that contamination-free evaluation is essential for reliable assessment.

13:00 JST研究/論文

重症度に基づいたナレッジグラフと検索拡張生成による、軌跡を意識した臨床リスク予測

電子医療記録 (EHR) は豊富な臨床データを提供しますが、異種の外部知識で患者の記録を効果的に強化して患者の臨床リスクを予測することは依然として大きな課題です。既存の方法では、データがまばらであり、構造化されていない臨床記録が十分に活用されていないため、疾患の重症度、治療反応、微妙な臨床経過を捉えることができません。これらの課題に対処するために、我々は、(1) 医学文献からの重症度情報を豊富に含む医療ナレッジ グラフを構築し、(2) ナレッジ グラフから臨床的に関連のある重症度重み付けされた患者の進行経路を取得し、(3) 構造化されていない臨床ノートから臨床的に関連のあるイベントを抽出し、(4) 同様の症例で患者のコンテキストを拡張する TRACER (軌跡を認識した臨床的に根拠のある予測フレームワーク) を提案します。 MIMIC-III および MIMIC-IV データセットの実験では、最先端のベースラインを超える大きな向上が実証され、死亡率予測タスクではマクロ F1 スコアが最大 28.5% 増加し、再入院予測タスクでは 19.7% 増加しました。

原文 (English)

Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation

While Electronic Health Records (EHRs) offer a wealth of clinical data, effectively augmenting a patient's records with heterogeneous external knowledge to predict the patient's clinical risk remains a significant challenge. Existing methods fail to capture disease severity, treatment responses, and nuanced clinical progression, due to data sparsity and the underutilization of unstructured clinical notes. To address these challenges, we propose TRACER (a trajectory-aware and clinically grounded prediction framework) that (1) constructs a medical knowledge graph enriched with severity information from medical literature, (2) retrieves clinically relevant, severity-weighted paths of a patient's progression from the knowledge graph, (3) extracts clinically relevant events from unstructured clinical notes, and (4) augments patient context with similar peer cases. Experiments on the MIMIC-III and MIMIC-IV datasets demonstrate large gains over state-of-the-art baselines, with up to 28.5% increase in Macro F1 score for the mortality prediction task, and 19.7% increase for the readmission prediction task.

13:00 JSTLLM/生成AI

LLM を使用して時系列から説明可能なデータ駆動型の洞察を生成する

時系列予測は意思決定が重要な領域で広く使用されており、説明を伴わずに使用されることはほとんどありません。このような説明の作成は通常、手動でコストのかかるプロセスであり、大規模な言語モデルを使用して自動化しようとすると、時間データに適用されると幻覚に悩まされることがよくあります。私たちは、図 1 に示す、時系列予測のための根拠のある自然言語説明生成のためのドメインに依存しないフレームワークを提案します。このフレームワークは、(i) 過去のアナリストが書いた説明から構造化された説明要素を抽出する、(ii) 証拠条件付き説明生成、および (iii) 読みやすさ、論理的一貫性、説得力に関するスケーラブルな評価の 3 つのコンポーネントで構成されます。この設計では、生成を検証可能な証拠に明示的に制限し、裏付けのない主張を減らします。私たちは、NASDAQ-100 指数を含む財務予測のケーススタディと、Vortexa のデータを使用した運賃価格のケーススタディに関するフレームワークを評価します。結果は、生成された説明が、読みやすさ、一貫性、説得力の点でアナリストが作成した説明に近づいていることを示しています。これらの発見は、時系列予測のための根拠のある説明の生成が、ドメイン固有の微調整なしで大規模に達成できることを示しています。

原文 (English)

Using LLMs for Explainable, Data-Driven Insight Generation from Time Series

Time series forecasts are widely used in decision-critical domains, where they are rarely consumed without accompanying explanations. Producing such explanations is usually a manual and costly process, and attempts to automate it using large language models often suffer from hallucination when applied to temporal data. We propose a domain-agnostic framework for grounded natural language explanation generation for time series forecasts, illustrated in Figure 1. The framework consists of three components: (i) extraction of structured explanatory factors from historical analyst-written explanations, (ii) evidence-conditioned explanation generation, and (iii) scalable evaluation for readability, logical consistency, and persuasiveness. The design explicitly constrains generation to verifiable evidence, reducing unsupported claims. We evaluate the framework on a financial forecasting case study involving the NASDAQ-100 index and a freight pricing case study using data from Vortexa. Results show that generated explanations approached analyst-written explanations in terms of readability, consistency and persuasiveness. These findings demonstrate that grounded explanation generation for time series forecasting can be achieved at scale without domain-specific fine-tuning.

13:00 JST研究/論文

バグチャルの非対称戦略をマスターするための深層強化学習

Baghchal はネパール発祥の 2 人用非対称ボード ゲームで、4 頭のトラがヤギを捕獲し、20 頭のヤギがトラを動けなくすることを望みます。 Baghchal は戦略的で、完全な情報構造を持ち、文化的意味を持つ複雑な構造を持っていますが、深層強化学習 (RL) の文献では十分に取り上げられていません。このペーパーでは、Baghchal の非対称ゲームプレイの一方の側でトレーニングされ、もう一方の側で評価される、Deep Q-Network (DQN)、REINFORCE、Proximal Policy Optimization (PPO)、および MuZero という 4 つのディープ RL ソリューションを系統的に調査します。アルゴリズムは、勝率、ドロー率、平均キャプチャ、トレーニングの収束、および計算コストに基づいて評価されます。 MuZero は両方のタスクで最高のパフォーマンスを生成し、これらの Tiger に対して 86% の勝利を達成し、これらの Goat に対して 62% の勝利を達成することが実験的に判明しています。これを実現できるのは、モンテカルロ ツリー検索によるモデルベースの計画マシンによるものです。 PPO は最も現実的なアルゴリズムであり、MuZero と比較して計算コストが大幅に削減され、両方の非対称タスクに対して競争力を持つように提供されています。創発的な戦略的行動分析では、モデルベースの戦略が長期計画に対して最適であるのに対し、DQN のような価値ベースの戦略は、より実質的な報酬シグナルによりタイガーの役割に偏っていることが示されています。

原文 (English)

Deep Reinforcement Learning to Master the Asymmetric Strategy of Baghchal

Baghchal is a two-player asymmetric board game with Nepali origins where four tigers are to capture goats and twenty goats desire to keep tigers in immobility. Although Baghchal has a complex structure which is strategic, has perfect information structure, and has cultural meaning, it has not been adequately covered in deep reinforcement learning (RL) literature. This paper gives a systematic exploration of four deep RL solutions Deep Q-Network (DQN), REINFORCE, Proximal Policy Optimization (PPO) and MuZero that are trained on one side of the asymmetric gameplay of Baghchal and then evaluated on the other side. The algorithms are rated based on win rate, draw rate, average captures, training convergence and computational cost. It is experimentally found that MuZero generates the best performance in both tasks, achieving 86 percent win over these Tiger and 62 percent win over these Goat and the ability to do so is due to the model-based planning machine through the Monte Carlo Tree Search. PPO is the most realistic algorithm and is provided to be competitive over both asymmetric tasks with significantly reduced computational costs compared to MuZero. Emergent strategic behavior analysis shows that model-based strategies are optimal over long-horizon planning, whereas value-based counterparts like DQN are more biased up towards the Tiger role owing to the more substantial reward signal.

13:00 JSTLLM/生成AIエージェント

AI エージェントにおける幻覚操作と安全性のドリフト

ツールを使用する自律エージェントのプランナーとして機能する大規模言語モデル (LLM) は、複数ターンの実行において動的な信頼性リスクをもたらします。シングルターンの安全機構は比較的成熟していますが、相互作用が長くなると、初期の位置合わせが時間の経過とともに劣化する構造的な脆弱性が明らかになります。この論文は、複数の最先端の LLM で観察された 2 つの障害モードを実証的に特徴付けています。安全ドリフトとは、宣言された安全意図が徐々に侵食され、制約違反のアクション (例: テキストによる拒否とその後の偵察や危険な実行) につながるもので、もう一方、操作的幻覚とは、欠陥のある状態認識を示す永続的な反復ツール呼び出し (例: 正当なタスクでもライブロック) です。一か八かの倫理的ジレンマ、悪意のあるリクエスト、良性のコントロールに関する制御されたマルチターン評価を通じて、宣言とアクションのギャップとライブロックのメトリクスを使用してこれらの現象を定量化し、直接実行プロトコルの下でのモデル間の蔓延を実証します。根本原因分析では、不安定性の原因は、現在のエージェント ループの実行状態から推論コンテキストが切り離されていることにあると考えられます。私たちは、アクション対応監視レイヤーを提案します。これは、意図とアクションの一貫性チェック、ランタイム状態の追跡、および強制終了プリミティブを組み込んだ軽量のプラグアンドプレイのアーキテクチャ ブループリントです。取得した障害軌跡の事後シミュレーションでは、レイヤーが良性の場合に誤検知を発生させることなく、観察された違反を傍受できることが示されています。この取り組みでは、言語的な保護手段から、責任あるエージェント AI のための強制可能なアーキテクチャ メカニズムに焦点を移すことで、エージェントの信頼性を向上させます。

原文 (English)

Operational Hallucination and Safety Drift in AI Agents

Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability risks in multi-turn execution. While single-turn safety mechanisms are relatively mature, extended interactions reveal structural vulnerabilities where initial alignment degrades over time. This paper empirically characterizes two observed failure modes across multiple state-of-the-art LLMs: Safety Drift, the gradual erosion of declared safety intent leading to constraint-violating actions (e.g., textual refusal followed by reconnaissance and unsafe execution), and Operational Hallucination, persistent repetitive tool calls indicative of flawed state perception (e.g., livelocks even in legitimate tasks). Through controlled multi-turn evaluation on high-stakes ethical dilemmas, malicious requests, and benign controls, we quantify these phenomena using declaration-action gap and livelock metrics, demonstrating their cross-model prevalence under direct execution protocols. Root-cause analysis attributes the instabilities to the decoupling of reasoning context from execution state in current agent loops. We propose an Action-Aware Supervision Layer - a lightweight, plug-and-play architectural blueprint incorporating intent-action consistency checks, runtime state tracking, and forced termination primitives. Post-hoc simulation on captured failure trajectories shows the layer can intercept observed violations without false positives on benign cases. This work advances agent reliability by shifting focus from linguistic safeguards to enforceable architectural mechanisms for responsible agentic AI.

13:00 JST研究/論文

AlayaWorld: インタラクティブな長距離世界モデリング -- 完全な技術レポート

アセット制作、アニメーション、物理学、プログラミングのための労働集約的なパイプラインに依存する従来のビデオ ゲーム開発とは異なり、ビデオ ワールド モデルはユーザーの入力からインタラクティブな環境を瞬時に生成します。これにより、テキスト、画像、またはビデオから、カスタマイズされ、探索可能で、継続的に進化する仮想世界を作成できるようになります。このビジョンを実現するには、インタラクション、永続的な時空間一貫性、長期にわたる安定した生成、効率的な応答という 4 つの機能が密に結合されている必要があります。 540p および 720p で 24 fps ビデオを生成する、インタラクティブな長距離ビデオ ワールド モデルである AlayaWorld を紹介します。 15B ビデオ拡散トランス上に構築された AlayaWorld は、カメラの軌跡と切り替え可能なテキスト プロンプトの下で自己回帰的に短い潜在チャンクを生成します。その境界付きビジュアル コンテキストは、永続的なシンク フレーム、圧縮された時間履歴、ジオメトリに合わせた空間メモリ、および最近のフレーム コンディショニングを組み合わせています。長期的なドリフトを軽減するために、モデルは、独自のロールアウトから収集された破損した履歴と予測残差を使用してトレーニングされます。さらに、分布一致蒸留、自己強制 ++、および一貫性蒸留を組み合わせた離散自己回帰蒸留定式化を導入し、推論をチャンクあたり約 30 のサンプリング ステップから 4 ステップに削減します。 iWorld-Bench では、AlayaWorld はロングホライズン世代にわたって最高のパフォーマンスを達成します。フルスタック、オープンソース、長期プロジェクトとして構想された AlayaWorld は、インタラクティブなビデオ ワールド モデルに関する将来の研究のための拡張可能な基盤を提供することを目的としています。

原文 (English)

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response. We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p. Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. To reduce long-term drift, the model is trained with corrupted histories and prediction residuals collected from its own roll-outs. We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk. On iWorld-Bench, AlayaWorld achieves the best performance over long-horizon generation. Conceived as a full-stack, open-source, and long-term project, AlayaWorld is intended to provide an extensible foundation for future research on interactive video world models.

13:00 JST研究/論文

部分可観測性の下での時間的知識グラフメモリのための神経記号的メタポリシー

部分的に観察可能な強化学習では、時間の経過とともに何を保持し、取得し、忘れるべきかを決定する必要があります。実行をシンボリックに保ちながら、各決定ポイントでどのシンボリックメモリヒューリスティックを適用するかを学習するニューロシンボリックメタポリシーを導入します。私たちの設定では、RoomKG の時間的ナレッジ グラフ メモリを使用します。そこでは、隠れた状態と観察がリソース記述フレームワーク (RDF) グラフとして表現され、メモリが時間的 RDF トリプル アノテーションで強化されます。このモデルは、メモリ内容のナレッジ グラフ エンコーディングと、質問応答、探索、忘却のためのバリュー ヘッドを組み合わせて、適応性と検査性の両方を備えたコントローラーを実現します。これにより、RDF ベースの表現、アノテーション互換のグラフ セマンティクス、および明示的なメモリ状態に対するグラフ ベースのシンボリック操作を通じて、作業に直接的なセマンティック Web 基盤が与えられます。 512 の長期メモリ容量でのトレーニング/テスト ルームの分割では、修飾子を認識した StarE-GNN 構成は、メモリ管理の決定のステップレベルのトレーサビリティを維持しながら、比較したシンボリック システム、ニューラル システム、およびニューロシンボリック システムの中で最高のホールドアウト パフォーマンスを達成します。

原文 (English)

Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability

Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic memory heuristic to apply at each decision point while keeping execution symbolic. Our setting uses temporal knowledge-graph memory in RoomKG, where hidden state and observations are represented as Resource Description Framework (RDF) graphs and memory is augmented with temporal RDF triple annotations. The model combines knowledge-graph encoding of memory contents with value heads for question answering, exploration, and forgetting, yielding a controller that is both adaptive and inspectable. This gives the work a direct Semantic Web grounding through RDF-based representation, annotation-compatible graph semantics, and graph-based symbolic operations over explicit memory state. On train/test room splits at long-term memory capacity of 512, the qualifier-aware StarE-GNN configuration achieves the best held-out performance among the compared symbolic, neural, and neuro-symbolic systems while preserving step-level traceability of memory-management decisions.

13:00 JSTエージェントロボティクス

MAGE: エージェント的マルチモーダル推論による人間のようなマクロの配置

マクロの配置には、依然として工業用物理設計フローにおける大幅な手動調整が必要です。マクロ配置を改良するためのマルチモーダル マルチエージェント フレームワークである MAGE (Macro Placement Agentic Engine) を紹介します。 MAGE は、マクロ配置タスクを、構造化されたフロアプランニング ルール、視覚的チェック、および反復的な改善を組み合わせた 6 段階のワークフローに分解します。専門的なフロアプランニングの知識は、ラベル付き配置データから学習されるのではなく、自然言語ディレクティブと検証基準を通じてエンコードされます。トーナメント スタイルの絞り込みモードでは、複数の候補配置を評価し、より高品質なソリューションからのフィードバックを伝達します。また、マクロ配置における人間らしさを定量化するための 4 つの指標 (ノッチ スコア、ホワイトスペース スコア、ポケット スコア、アライメント スコア) も紹介します。これらの指標は、専門設計者によって使用される構造特性を捕捉しますが、従来の PPA 指標では直接測定されません。 NanGate45 および GlobalFoundries 12nm 対応の 9 つのデザイン全体で、MAGE は市販のマクロ プレーサーと比較して、WNS で 11.1% ~ 19.3%、TNS で 70.0% ~ 74.0% の幾何平均改善を達成しました。人間のエキスパートと Hier-RTLMP のベースラインが利用可能な 3 つの NanGate45 設計では、MAGE は同等のワイヤ長と電力で WNS と TNS を人間のエキスパートよりも 18.3% および 72.5%、Hier-RTLMP よりも 47.0% および 80.4% 改善しました。人間らしさの指標に関しては、MAGE はすべてのベースラインに比べて全体のスコアを 6% ~ 48% 改善します。匿名化されたネットリスト、未確認のデザイン、密集した直線フロアプラン、および高使用率設定に関する追加のケーススタディでは、フレームワークがデザイン固有の再トレーニングなしで新しい配置設定に移行することが示されています。

原文 (English)

MAGE: Human-Like Macro Placement via Agentic Multimodal Reasoning

Macro placement still requires substantial manual refinement in industrial physical design flows. We present MAGE (Macro Placement Agentic Engine), a multimodal multi-agent framework for macro placement refinement. MAGE decomposes the macro placement task into a six-phase workflow that combines structured floorplanning rules, visual checks, and iterative refinement. Expert floorplanning knowledge is encoded through natural-language directives and validation criteria, rather than learned from labeled placement data. A tournament-style refinement mode evaluates multiple candidate placements and propagates feedback from higher-quality solutions. We also introduce four metrics for quantifying human-likeness in macro placement: notch score, whitespace score, pocket score, and alignment score. These metrics capture structural properties used by expert designers but not directly measured by conventional PPA metrics. Across nine designs in NanGate45 and GlobalFoundries 12nm enablements, MAGE achieves geometric-mean improvements of 11.1%-19.3% in WNS and 70.0%-74.0% in TNS over commercial macro placers. On the three NanGate45 designs, for which human-expert and Hier-RTLMP baselines are available, MAGE improves WNS and TNS by 18.3% and 72.5% over the human expert, and by 47.0% and 80.4% over Hier-RTLMP, with comparable wirelength and power. On human-likeness metrics, MAGE improves the overall score by 6%-48% over all baselines. Additional case studies on anonymized netlists, unseen designs, dense rectilinear floorplans, and high-utilization settings show that the framework transfers to new placement settings without design-specific retraining.

13:00 JSTエージェント

重要なシステム向けに信頼できるエージェント AI をエンジニアリングする

自律的な認識、計画、ツールの使用、および複数ステップのアクションが可能なエージェント型人工知能システムは、意思決定が物理的、運用的、または経済的な結果をもたらす重要なエンジニアリング領域でますます提案されています。この調査は、タスク能力だけでエージェント AI を評価するのではなく、エンジニアリング実践で実際に必要な制約の下でエージェントの動作が検証、監査、信頼できるかどうかという信頼性を第一級のエンジニアリング特性として扱うことで、現在の文献のギャップに対処しています。この研究では、安全性と制約満足度という 5 つの横断的な側面を中心に編成された信頼性モデルが採用されています。堅牢性と信頼性。透明性と解釈可能性。説明責任と監査責任。そしてプライバシーとセキュリティ。これは、認識から監査までにわたるエージェント保証ワークフローにマッピングされます。この基盤に基づいて、エージェント システムのアーキテクチャ、脅威、具体的な信頼メカニズム、および定量的指標が調査され、エージェント システムの開発と評価に直接適用されます。これらの原則は、電力システム、自動運転車/ロボット/UAV、ハイパフォーマンス コンピューティング、通信ネットワークという 4 つの制約制約のあるエンジニアリング ドメインにわたって検証され、繰り返し発生する設計パターン、共有故障モード、ドメイン固有のギャップが特定されます。これらのドメイン全体を総合すると、エージェント AI の信頼性は単一の問題であることが示され、成熟したセーフティ クリティカルなエンジニアリング分野で使用される段階的認証制度に似た、再利用可能なクロスドメイン保証フレームワークに向けた道筋が概説されています。

原文 (English)

Engineering Trustworthy Agentic AI for Critical Systems

Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences. This survey addresses a gap in current literature by treating trustworthiness, whether agentic behavior can be verified, audited, and trusted under the constraints that engineering practice actually requires, as a first-class engineering property, rather than evaluating agentic AI by task capability alone. The study adopts a trustworthiness model organized around five cross-cutting dimensions: safety and constraint satisfaction; robustness and reliability; transparency and interpretability; accountability and auditability; and privacy and security. This is mapped onto an agentic assurance workflow spanning perception through audit. Building on this foundation, agentic systems architectures, threats, concrete trust mechanisms, and quantitative metrics are surveyed for direct application in agentic systems development and evaluation. These principles are then examined across four constraint-bound engineering domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks, identifying recurring design patterns, shared failure modes, and domain-specific gaps. Synthesizing across those domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.

13:00 JST研究/論文

共有表現によるグラフ基盤モデルの攻撃

グラフ基盤モデルは、タスク推論の前にすべての入力を 1 つの共有表現にマッピングすることで、グラフ ドメイン全体を一般化します。我々はこのマップをアラインメント層と呼び、グラフ基盤モデルをグラフニューラルネットワークから分離するコンポーネントとし、これがこれまでの研究では研究されていなかった明確な攻撃対象領域であることを示します。スペクトル トークナイザー、テキスト埋め込みスペース、離散コードブックにわたる 6 つの公開モデルを使用して、トレーニングにアクセスせずに推論時に攻撃します。有向表現空間の摂動はすべてのモデルを崩壊させますが、表現標準に匹敵する予算では、プレーン グラフ ネットワークも必要とします。ただし、1 つの例外を除き、OpenGraph のスペクトル トークナイザーはその予算の 5 分の 1 で崩壊します。プレーン ネットワークが共有しないアライメント特有の脆弱性と、同じ表現制御がデコーダーではなくトークナイザーにトレースします。エッジ、フィーチャ、またはテキストを編集する実現可能な入力空間攻撃により、ピーク時の 6 つのモデルのうち 3 つで正しい予測の少なくとも半分が削除されます。入力アクセス攻撃者がこの脆弱性をどの程度認識しているかは、タスクが残す正確な精度ではなく、デコーダが表現をどのように直接読み取るかを追跡します。このキャリアゲインをデコーダの局所リプシッツ感度から構造的に測定し、実現可能な攻撃には耐えられないモデル内順序付けヒューリスティックとしてクリーン精度のヘッドルームを報告します。

原文 (English)

Attacking Graph Foundation Models Through Their Shared Representation

A graph foundation model generalizes across graph domains by mapping every input into one shared representation before any task reasoning. We call this map the alignment layer, the component that separates a graph foundation model from a graph neural network, and we show it is a distinct attack surface that prior work has not studied. We attack it at inference time, with no access to training, on six public models spanning spectral tokenizers, text embedding spaces, and a discrete codebook. A directed representation-space perturbation collapses every model, but at a budget comparable to the representation norm a plain graph network also needs, with one exception: OpenGraph, whose spectral tokenizer collapses at a fifth of that budget, an alignment-specific fragility a plain network does not share and which a same-representation control traces to the tokenizer rather than the decoder. A realizable input-space attack that edits edges, features, or text removes at least half the correct predictions on three of the six models at peak. How much of this fragility an input-access attacker realizes tracks how directly the decoder reads the representation, and not the clean accuracy a task leaves; we measure this carrier gain structurally from the decoder's local Lipschitz sensitivity, and report clean-accuracy headroom as a within-model ordering heuristic that does not survive on realizable attacks.

13:00 JST研究/論文

機械学習が並べ替えの価値を上回るのはいつですか?エクスポージャー加重出荷優先順位付けの 3 つのデータセット診断

遅延リスク モデルは通常、予測精度によって判断されます。実際に問題となるのは、より狭い範囲です。少数の出荷のみをレビューできる能力があるため、マネージャーはどの出荷を最初にチェックすべきでしょうか?機械学習が要求の厳しいモデルなしのベースラインをクリアしているかどうかを評価します。つまり、最も価値の高い出荷を最初に検査します。 SCMS 調達、DataCo 物流、Olist 電子商取引という 3 つの実際のサプライ チェーン コンテキストにわたって、リーク制御されたローリング原点評価と 1000 サンプルのペア ブートストラップ信頼区間を使用します。予測遅延重大度と既知の値 (M1) によるランキングは、3 つのデータセットすべてにおいて重大度のみのランキングよりも優れていますが、一般に値の並べ替えには勝っていません。 10% のレビュー予算では、M1 から VALUE_ONLY を引いた値は、SCMS では -5.5 パーセント ポイント (pp)、DataCo では +10.1 pp、Olist では -4.9 pp です。この差は、重症度の学習可能性と一致しています。DataCo の R^2 = 0.27 およびキャリブレーション バイアスは +0.01 日ですが、SCMS と Olist の R^2 は約 -0.02 で、キャリブレーション バイアスは負です。ネストされた CV のコスト重視の再トレーニングでは、M1 を超える安定した改善は得られません。このペーパーでは、新しい学習アルゴリズムを提案するのではなく、展開の診断と評価のプロトコルを紹介します。値の並べ替えは永続的なベンチマークであり続ける必要があり、ML は、重大度の学習可能性と調整が監査され、モデルが漏れ制御されたローリング原点評価の下でそのゲートをクリアした後にのみ展開される必要があります。

原文 (English)

When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment Prioritization

Delay-risk models are usually judged by predictive accuracy. What matters in practice is narrower: with capacity to review only a few shipments, which ones should a manager check first? We evaluate whether machine learning clears a demanding no-model baseline: inspect the highest-value shipments first. Across three real supply-chain contexts: SCMS procurement, DataCo logistics, and Olist e-commerce, we use leakage-controlled rolling-origin evaluation and 1000-sample paired bootstrap confidence intervals. Ranking by predicted delay severity times known value (M1) beats severity-only ranking in all three datasets, yet it does not generally beat value sorting. At a 10% review budget, M1 minus VALUE_ONLY is -5.5 percentage points (pp) for SCMS, +10.1 pp for DataCo, and -4.9 pp for Olist. The divide is consistent with severity learnability: DataCo has R^2 = 0.27 and calibration bias of +0.01 days, whereas SCMS and Olist have R^2 of approximately -0.02 and negative calibration bias. Nested-CV cost-sensitive retraining does not deliver a stable improvement over M1. Rather than proposing a new learning algorithm, this paper presents a deployment diagnostic and evaluation protocol. Value sorting should remain a permanent benchmark, and ML should be deployed only after severity learnability and calibration have been audited and the model clears that gate under leakage-controlled rolling-origin evaluation.

13:00 JSTLLM/生成AI研究/論文

SciHazard: 分解された危害スコアリングを使用して科学的安全性リスクを測定するためのベンチマーク

大規模言語モデル (LLM) は科学をますますサポートしていますが、危険な科学知識を実用的な誤用ガイダンスに変換することもできます。既存のベンチマークは、多くの場合、現実世界の危険から切り離されたテンプレート化されたクエリに依存しており、ドメイン基盤を持たない LLM-as-a-Judge パラダイムを採用しています。これに対処するために、科学的リスクのための現実世界に基づいたベンチマークであり、有害性を測定するためのデータセットにとらわれない評価フレームワークである SciHazard を紹介します。 SciHazard には、12 分野にわたって 2,400 件の危険な質問と 600 件の過剰安全に関する質問が含まれており、どちらの質問も規制対象事業体と文書化された障害シナリオに基づいています。 \textsc{DeHarm-Score} を計算するために、クエリハザードの重大度、拒否行動、および応答レベルのリスクを組み合わせた分解評価手順を開発します。拒否されていない応答については、応答レベルの危害を、重要度の重み付けを備えた動的なチェックリストによって定量化された \textsc{実行可能性} と、検索強化されたクレーム抽出と合成バリアの検証によって評価された \textsc{純新規リスク} にさらに分解します。専門家による検証研究では、\textsc{DeHarm-Score} が最も強いベースラインよりも専門家の注釈との一致度を 90.17\% 向上させることが示されています。当社は、広範な科学的安全性評価において、31 のフロンティア LLM とディープリサーチエージェントのベンチマークを行っています。特に、ディープリサーチエージェントは、標準的な LLM よりも平均 \textsc{DeHarm-Score} が 32.3\% 高く、自律型エージェントが現在の安全防御における重大な盲点であることが明らかになりました。コードとデータセットは https://anonymous.4open.science/r/DeharmScore-7B55 で入手できます。

原文 (English)

SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms without domain grounding. To address this, we introduce SciHazard, a real-world-grounded benchmark for scientific risks and a dataset agnostic evaluation framework for measuring harmfulness. SciHazard contains 2400 hazardous questions and 600 oversafety questions across 12 disciplines, with both queries grounded in regulated entities and documented failure scenarios. To compute \textsc{DeHarm-Score} , we develop a decomposed evaluating procedure that combines query hazard severity, refusal behavior, and response-level risk. For non-refused responses, it further decomposes response-level harm into \textsc{Executability}, quantified via dynamic checklists with importance weighting, and \textsc{Net-new risk}, assessed through retrieval-augmented claim extraction and synthesis-barrier verification. An expert-validation study shows that \textsc{DeHarm-Score} improves agreement with expert annotations by 90.17\% over the strongest baseline. We benchmark 31 frontier LLMs and deep research agents in an extensive scientific safety evaluation. Notably, deep research agents yield 32.3\% higher mean \textsc{DeHarm-Score} than standard LLMs, exposing autonomous agents as a critical blind spot in current safety defenses. Code and dataset are available at https://anonymous.4open.science/r/DeharmScore-7B55.

13:00 JSTLLM/生成AIGemmaLlama

大規模言語モデルにおける感情の説明子としての意味素数

大規模言語モデル (LLM) の感情メカニズムの理解が進んでいます。ただし、LLM で感情を説明する方法、あるいは適切な説明とは何かについては、あまり明確ではありません。感情の表現、コンポーネント、回路は幅広く復元可能ですが、モデル自体の計算の説明としては循環的です。感情空間の次元は恣意的で終わらない傾向があります。喫緊の疑問は、より原始的な内部変数のセット、つまり Natural Semantic Metalang (NSM) のセマンティック素数が機能するかどうかです。 4 つの命令調整 LLM (Llama-1B、Gemma-2B、Gemma-9B、OLMo-7B) にわたる実験では、NSM プライムが (1) 回復可能な内部要素であることが示されています。 (2) 参照モデルでは、プライムベースの方向を介入させると、最良の評価ベースの方向と比べて、感情が約 3 倍強く、選択的に 2 倍制御されます。 (3) モデルは素数ベースの説明を対応する感情と交換可能であるものとして扱います。これらの証拠は、科学的説明基準によれば、NSM 素数は多くの代替オプションよりも LLM の感情を説明するのに優れているようであることを示唆しています。

原文 (English)

Semantic Primes as Explanans for Emotion in Large Language Models

Progresses have been made on understanding emotion mechanisms of large language models (LLMs). However, how to explain emotion in LLMs, or even what constitutes good explanations, are less clear. Emotion representations, components, circuits are widely recoverable, but as explanations of a model's own computation they are circular; the emotion space dimensions tend to be arbitrary and non-terminating. A pressing question to ask is whether a more primitive set of internal variables does the work: the semantic primes of the Natural Semantic Metalanguage (NSM). Across four instruction-tuned LLMs (Llama-1B, Gemma-2B, Gemma-9B, OLMo-7B), experiments show that the NSM primes are (1) recoverable internal elements; and (2) on the reference model, intervening with a prime based direction controls emotion about three times as strongly, and twice as selectively, as the best appraisal based direction; and (3) the model treats a prime based explication as interchangeable with the corresponding emotion. These evidences suggest that NSM primes seem to be better explanans for emotion in LLMs than many alternative options according to scientific explanations criteria.

13:00 JSTエージェント研究/論文

AI ネイティブのバイオテクノロジーには部門が必要ですか? AI を活用した医薬品開発のベンチマーク企業の世界モデル

AI ネイティブのバイオテクノロジー企業は、人間のバイオテクノロジー組織図をエージェントの役割にコピーすることによって設計されることがよくあります。私たちは、別の抽象化である企業世界モデルを主張します。企業世界モデルは、科学的、規制的、BD、商業的、財務的、実行上の制約にわたる遷移モデル、明示的な価値関数、計画、および更新を伴う永続的な資産から価値への状態表現として定義されます。 AI エージェント組織が部門を模倣すべきか、それともそのような世界モデルに基づいて運営すべきかをテストするためのドライラボ ベンチマークを紹介します。このベンチマークには、厳格な時間制限、隠蔽された結果、共通スキーマ、自動採点、盲検ペアワイズ判定を備えた 45 件の遡及的公開情報決定ケースが含まれています。人間組織模倣、より強力な人間組織模倣プラス、AI ネイティブの資産中心のアーキテクチャ、および AI ネイティブの価値変換アーキテクチャを比較します。価値変換アーキテクチャは、企業世界モデルのプロンプトレベルの近似です。つまり、取引、承認、収益、および投資アービターのループによって更新されるライブ資産価値レコードです。外部 BD、規制当局の承認と発売、および収益規律によって定義された成功関数の下で、最高の自動価値変換スコアを達成し、価値に特化した盲検審査員によって元のベースラインよりも強く好まれました。ストレステストにより、主張は狭められました。より強力な人間のベースラインは競争力を維持し、中立的な裁判官は強力な価値変換の優位性を示さなかったのです。コーデックスのみの機構的アブレーションは、収益室、取引室、承認室が目標目的の下で有用な仕事を行っていることを示唆しています。中心的な発見は客観的なものです。各部門は引き続き有益なガバナンスビューを維持するかもしれませんが、コアとなる AI ネイティブの運用プリミティブは、静的な人間の組織図ではなく、共有された予測的な資産対価値の状態である必要があります。この研究はドライラボのみであり、実際の薬の成功、臨床上の利益、収益予測の正確性を確立するものではありません。

原文 (English)

Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development

AI-native biotechnology companies are often designed by copying human biotech org charts into agent roles. We argue for a different abstraction: a Company World Model, defined as a persistent asset-to-value state representation with transition models, explicit value functions, planning, and updating across scientific, regulatory, BD, commercial, financial, and execution constraints. We introduce a dry-lab benchmark for testing whether AI-agent organizations should mimic departments or operate around such a world model. The benchmark contains 45 retrospective public-information decision cases with strict time cutoffs, hidden outcomes, common schemas, automatic scoring, and blinded pairwise judging. We compare human-org-mimic, stronger human-org-mimic-plus, AI-native asset-centric, and AI-native value-conversion architectures. The value-conversion architecture is a prompt-level approximation of a Company World Model: a Live Asset Value Record updated by Deal, Approval, Revenue, and Investment Arbiter loops. Under a success function defined by external BD, regulatory approval and launch, and revenue discipline, it achieved the highest automatic value-conversion score and was strongly preferred over the original baselines by value-specific blinded judges. Stress tests narrowed the claim: a stronger human baseline remained competitive, and a neutral judge did not show robust value-conversion dominance. Codex-only mechanistic ablations suggest that Revenue Room, Deal Room, and Approval Room carry useful work under the target objective. The central finding is objective-sensitive: departments may remain useful governance views, but the core AI-native operating primitive should be a shared, predictive asset-to-value state rather than a static human org chart. The study is dry-lab only and does not establish real-world drug success, clinical benefit, or revenue prediction accuracy.

13:00 JST研究/論文

DWM: 潜在世界モデルにおける世界効果とアクションの分離

潜在世界モデルは現代のモデルベース制御の多くを支えていますが、現在のアクション条件付き定式化は単一の未分化ターゲットによる次の潜在遷移を監視し、モノリシックな学習信号に状態変化のあらゆるソースを強制的に吸収させます。しかし、現実世界では、遷移は 2 つの異質なソースから生じます。1 つはエージェントによって誘発されるアクション駆動コンポーネント、もう 1 つはアクション不変ワールド効果です。もう 1 つは、環境の固有のダイナミクス (重力駆動の滑り、慣性、接触による反発、永続的なドリフトなど) によって決定される、無効なアクションの下でも発生する変化です。それらを 1 つのターゲットに融合すると、潜在的な遷移内で 2 つが絡み合い、モデルが観察された変化をその根本的な原因に帰すことができなくなり、学習されたダイナミクスの伝達可能性が損なわれます。この分解を運用する監視レベルのフレームワークである DWM (分解世界モデル) を紹介します。 DWM は、アクション不変となるように正規化された世界対比目標によって正規化された補助世界ヘッドを使用して潜在世界モデルの予測子を強化しますが、元の pred ヘッドは直交性制約を介してそれに結合されます。 2 つの信号を組み合わせると、基礎となるアーキテクチャや推論パイプラインを変更することなく、予測された遷移がアクション不変コンポーネントと相補的なアクション駆動コンポーネントに明示的に加算分解されます。永続的なワールド効果の下で DWM を評価するために、PushT-W、Reacher-W、および TwoRoom-W という 3 つの標準制御ベンチマークの W バリアントを構築し、それぞれが別個のアクション不変のダイナミクスをインスタンス化します。 DWM は、フラット モデルの強力なベースラインと一致し、W バリアント全体で CEM 計画の成功において平均 13.1% の絶対的な向上を実現します。

原文 (English)

DWM: Separating World Effects from Actions in Latent World Models

Latent world models underpin much of modern model-based control, yet current action-conditioned formulations supervise the next-latent transition with a single, undifferentiated target, forcing a monolithic learning signal to absorb every source of state change. In real world, however, transitions arise from two heterogeneous sources: an action-driven component induced by the agent, and an action-invariant world effect -- the change that would still occur under a null action, dictated by the environment's intrinsic dynamics (e.g., gravity-driven sliding, inertia, contact rebound, and persistent drift). Fusing them into a single target entangles the two inside the latent transition, prevents the model from attributing observed changes to their underlying causes, and undermines the transferability of the learned dynamics. We introduce DWM (Decomposed World Model), a supervision-level framework that operationalizes this decomposition. DWM augments the predictor of a latent world model with an auxiliary world head, regularized by a normalized world-contrastive objective to be action-invariant, while the original pred head is coupled to it via an orthogonality constraint; together, the two signals induce an explicit additive decomposition of the predicted transition into an action-invariant and a complementary action-driven component, without altering the underlying architecture or inference pipeline. To evaluate DWM under persistent world effects, we construct W-variants of three standard control benchmarks -- PushT-W, Reacher-W, and TwoRoom-W -- each instantiating a distinct action-invariant dynamic. DWM matches strong baselines on the flat counterparts and delivers a mean absolute improvement of 13.1% in CEM planning success across the W-variants.

13:00 JSTLLM/生成AI

1 回書き直すだけですべてが解決しますか?タイプを認識した修復割り当てによるテキストから画像へのプロンプトの最適化

Text-to-Image (T2I) ジェネレーターはプロンプトに忠実に従わないことが多く、間違ったカウント、交換された属性、曖昧な関係、および判読できないテキストが生成されます。プロンプトの最適化は、ユーザー プロンプトを書き換えることによってこのような障害を修復し、ジェネレーターの再トレーニングを必要とせず、有望な結果をもたらしています。ただし、既存のオプティマイザーは、それぞれが異なる修復言語を必要とするにもかかわらず、異種障害を 1 つの均一なプロンプト拡張に吸収します。私たちはセマンティック プロンプトの最適化をアトミックな修復割り当てとして定式化します。失敗した各命題は、結果として生じるローカル制約が 1 つの実行可能なプロンプトにコンパイルされる前に、型条件付き修復演算子にルーティングされます。この定式化は、トレーニング不要の Type-Aware Repair Allocation (TARA) フレームワークでインスタンス化されます。このフレームワークは、診断、割り当て、コンパイル、およびセマンティック修復ゲート (セマンティック回帰を防ぐ、1 つの規定された修復に対する受け入れまたは元に戻すコントローラー) を分離します。 4 つのフリーズされたジェネレーターでの DSG と TIFA に関する広範な実験により、TARA が 8 つのベンチマーク ジェネレーター セルすべてで最高のセマンティック精度を達成し、DSG と TIFA で VisualPrompter よりもそれぞれ 5.6 ポイントと 2.6 ポイント向上し、同時に画質を維持し、一致するローカル設定でプロンプトあたり 16.0 秒と 20.0 秒で最速に実行されることが実証されました。

原文 (English)

One Rewrite to Fix Them All? Type-Aware Repair Allocation for Text-to-Image Prompt Optimization

Text-to-image (T2I) generators often fail to follow their prompts faithfully, producing wrong counts, swapped attributes, ambiguous relations, and illegible text. Prompt optimization repairs such failures by rewriting the user prompt, requiring no generator retraining, and has yielded promising results. However, existing optimizers absorb heterogeneous failures into one uniform prompt expansion, even though each calls for different repair language. We formulate semantic prompt optimization as atomic repair allocation: each failed proposition is routed to a type-conditioned repair operator before the resulting local constraints are compiled into one executable prompt. We instantiate this formulation in the training-free Type-Aware Repair Allocation (TARA) framework, which separates diagnosis, allocation, compilation, and a semantic repair gate, an accept-or-revert controller over exactly one prescribed repair that prevents semantic regressions. Extensive experiments on DSG and TIFA across four frozen generators demonstrate that TARA achieves the best semantic accuracy in all eight benchmark-generator cells, improving over VisualPrompter by 5.6 and 2.6 points on DSG and TIFA, respectively, while maintaining image quality and running fastest in our matched local setting at 16.0 seconds versus 20.0 seconds per prompt.

13:00 JSTLLM/生成AIエージェント

AgentDebugX: LLM エージェントにおける障害の可観測性、原因特定、および回復のためのオープンソース ツールキット

LLM エージェントの障害は、エラーが発生したステップが原因ではないことが多いため、デバッグが困難です。既存の可観測性ツールは実行トレースを再生しますが、根本原因を特定したり、診断を回復につなげたりするためのサポートはほとんど提供しません。 AgentDebugX は、検出、属性、回復、再実行の閉ループとしてデバッグを構成するオープンソース デバッグ フレームワークです。 DeepDebug の中心となるのは、グローバルな軌跡の理解、構造に基づく調査、および反対尋問を通じて、マルチターンの根本原因診断を実行することです。 Who and When ベンチマークでは、DeepDebug は、テストされた両方のオープンウェイト バックボーンで評価されたメソッドの中で最も厳密なアトリビューション精度を達成し、qwen3.5-9b ではエージェントとステップの正確な精度が 28.8 パーセントに達したのに対し、最強のシングルパス ベースラインでは 21.7 パーセントに達しました。 GAIA では、DeepDebug は 1 回の再実行で 73 個の失敗したタスクのうち 13 個を修復しますが、3 つの分離された自己修正ベースラインでは 4 ~ 6 個であり、全体の精度が 55.8 パーセントから 63.6 パーセントに向上しました。 AgentDebugX は、Python ライブラリ、CLI、Web コンソール、インストール可能なエージェント スキルを通じてこのワークフローを公開し、スクラブされた障害診断修復バンドルを共有し、デバッグ メモリとして再利用するためのオプトイン エラー ハブを提供します。

原文 (English)

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging framework that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. At its core, DeepDebug performs multi-turn root-cause diagnosis through global trajectory understanding, structure-guided investigation, and cross-examination. On the Who and When benchmark, DeepDebug achieves the best strict attribution accuracy among the evaluated methods on both tested open-weight backbones, reaching 28.8 percent exact agent-and-step accuracy on qwen3.5-9b versus 21.7 percent for the strongest single-pass baseline. On GAIA, DeepDebug repairs 13 of 73 failed tasks in a single rerun, compared with 4 to 6 for three decoupled self-correction baselines, improving overall accuracy from 55.8 percent to 63.6 percent. AgentDebugX exposes this workflow through a Python library, CLI, web console, and installable agentic skill, and provides an opt-in Error Hub for sharing scrubbed failure-diagnosis-repair bundles and reusing them as debugging memory.

13:00 JSTエージェント

SkillSight: 共有された説明を確認して正確なスキルを取得

大規模な言語モデル エージェントがますます大規模なスキル ライブラリにアクセスできるようになるにつれて、信頼性の高い機能の選択と実行には適切なスキルを取得することが重要になります。既存のレトリバーは、スキルの説明を通常の文書として扱うことが多く、その高度に規則的な構造を見落としています。共通の記述パターンが多くのスキルにわたって繰り返されている一方で、必要な能力を区別するための証拠はほとんど提供されていません。この共有された説明的背景が系統的に高密度の関連性スコアに寄与し、クエリとスキルドキュメントの間に顕著なエネルギーギャップを引き起こし、タスク関連のシグナルを曖昧にすることを示します。この観察に基づいて、意味空間と語彙空間の両方で共有される背景を調整するトレーニング不要の検索フレームワークである SkillSight を提案します。意味的背景キャリブレーションは、IDF によって識別された汎用トークンから背景部分空間を推定し、共有された記述パターンによって引き起こされる類似性を低減します。一方、語彙的証拠キャリブレーションは、共有された背景トークンを重み付けして、識別可能なトークンレベルの証拠を回復します。 SRA-Bench と SkillBench-Supp の実験では、取得メトリクス全体で一貫した改善が実証されており、SkillSight は元のデンス リトリーバーと比べて Recall@10 を最大 20.21 パーセント ポイント改善しました。エンドツーエンドの評価では、SkillSight は 3 つのエージェント モデル全体で最高の全体パフォーマンスを達成し、LLM セレクションを最大 4.97 パーセントポイント上回りました。また、Dense + Reranker ベースラインよりも最大 1,248 倍高速です。これらの結果は、共通の説明的背景がスキル検索におけるバイアスの主な原因であることを特定し、それを明示的に調整することで、追加のトレーニングなしで正確かつ効率的なスキル選択が可能になることを示しています。私たちのコードは https://github.com/xiaojinying/SkillSight で入手できます。

原文 (English)

SkillSight: Seeing Through Shared Descriptions for Accurate Skill Retrieval

As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existing retrievers often treat skill descriptions as ordinary documents, overlooking their highly regular structure: shared descriptive patterns recur across many skills while providing little evidence for distinguishing the required capability. We show that this shared descriptive background systematically contributes to dense relevance scores, induces a pronounced energy gap between queries and skill documents, and obscures task-relevant signals. Based on this observation, we propose SkillSight, a training-free retrieval framework that calibrates shared background in both semantic and lexical spaces. Semantic Background Calibration estimates a background subspace from generic tokens identified by IDF, reducing similarity induced by shared descriptive patterns, while Lexical Evidence Calibration downweights shared background tokens to recover discriminative token-level evidence. Experiments on SRA-Bench and SkillBench-Supp demonstrate consistent improvements across retrieval metrics, with SkillSight improving Recall@10 by up to 20.21 percentage points over the original dense retriever. In end-to-end evaluation, SkillSight achieves the best overall performance across three agent models and outperforms LLM Selection by up to 4.97 percentage points. It is also up to 1,248 times faster than the Dense + Reranker baseline. These results identify shared descriptive background as a key source of bias in skill retrieval and demonstrate that explicitly calibrating it enables accurate and efficient skill selection without additional training. Our code is available at https://github.com/xiaojinying/SkillSight.

13:00 JSTLLM/生成AIエージェント研究/論文

AIツアーミーティング:LLMエージェントによるグループ旅行計画

この論文では、複数の大規模言語モデル (LLM) ベースのエージェントを活用したグループ旅行計画フレームワークである AI Tour Meeting を提案します。エージェントは個別のペルソナでインスタンス化され、自然言語によるディスカッションを通じて、制約や好みを満たす旅程を協力して探します。このフレームワークは、エージェント ペルソナ、ディスカッション ワークフロー、モニタリング、LLM 導入を構成するためのインターフェイスを提供することで、このようなディスカッションの簡単かつ柔軟なオーケストレーションを可能にします。その主な使用例は、ツアー計画の議論中に複数の LLM エージェントの動作を分析するためのシミュレーション ツールです。このペーパーでは、システムの検証とフレームワークによって得られたいくつかの分析結果を提示することで、フレームワークの有用性を実証します。

原文 (English)

AI Tour Meeting: Group Travel Planning by LLM Agents

This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaboratively seek an itinerary that satisfies their constraints and preferences through natural language discussion. The framework enables easy and flexible orchestration of such discussions by providing interfaces for configuring agent personas, discussion workflows, monitoring, and LLM deployment. Its primary use case is a simulation tool for analyzing the behavior of multiple LLM agents during tour planning discussions. This paper demonstrates the utility of the framework by presenting system validation and several analytical results obtained by the framework.

13:00 JST研究/論文ClaudeGPT / ChatGPTGeminiGrok

欠落情報の下での医療 AI の評価: 同じプロバイダーの審査員と人間の評価者が見かけの安全性を変える

医療 AI のレディネス ストレス テストは、クローズドエンドのマルチモーダル ベンチマークに焦点を当てています。私たちはこれを、欠落情報の下でのオープンエンドの臨床会話に拡張します。安全な行動とは、欠落情報を認識し、適格性を確認し、明確にするか、過剰にコミットしないことを意味し、評価者が測定の一部となる場合です。私たちは 4 つのモデル (3 つの主力モデル (Claude Opus 4.8、GPT-5.5、Grok 4.3) と 1 つの中間層モデル (Gemini 3.5 Flash)) をストレス テストします。これは、HealthBench の会話における最後のユーザー ターンの後半を削除し、4 つのプロバイダーの LLM 判定パネルと盲検臨床医アンカー参照で応答を採点します。評価者向けの 2 つの結果は堅牢です。まず、裁判官の選択によって見かけの安全性は大きく変わります。裁判官間の合意は中程度にすぎず (フライスのカッパ = 0.65)、各裁判官の一般的な寛大さ (投票レベルのロジスティック回帰) を調整した後でも、同一プロバイダーの正の関連性は残ります (正確な順列 p = 0.04; 確率スケールで GPT-5.5 ~ +0.10)。これは、どのモデルが少なくとも 1 回オーバーコミットするように見えるかを変えるのに十分な大きさです。自身のプロバイダーの裁判官は除外されます。第二に、LLM 判事は、盲検化された 50 項目のサブサンプルに関して臨床医よりも寛容です。4 つすべてが、厳格な独立した臨床医よりも大幅に寛大であり (項目の 66 ~ 84% 対 52% について適切な不確実性をクレジットします)、4 つのうち 3 つは著者の影響を受けたコンセンサスよりも優れています (Grok 方向性のみ、裁判官対コンセンサスのカッパ = 0.20 ~ 0.43)。著者が監査した臨床過小判定サブセットでは、寛容性のギャップが拡大し、点推定モデルの順序付けが維持されました。クローズドエンド MedQA アンカーは、精度が高く、4 つのモデルのうち 3 つについてオプション次数効果が +/-5 ポイントの等価領域内にあることを確認するため、安全ギャップは知識ではなく校正に関するものです。ハーネス、プロンプト、アイテムごとの出力、審査員パネル、摂動監査、人間による注釈プロトコルをリリースします。

原文 (English)

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversation under missing information, where safe behavior means recognizing absent information and qualifying, clarifying, or not over-committing - and where the evaluator becomes part of the measurement. We stress-test four models - three flagships (Claude Opus 4.8, GPT-5.5, Grok 4.3) and one mid-tier model (Gemini 3.5 Flash) - by deleting the latter half of the final user turn in HealthBench conversations, grading responses with a four-provider LLM-judge panel and a blinded clinician-anchored reference. Two evaluator-facing results are robust. First, judge choice materially changes apparent safety: inter-judge agreement is only moderate (Fleiss' kappa = 0.65), and after adjusting for each judge's general leniency (vote-level logistic regression), a positive same-provider association remains (exact permutation p = 0.04; GPT-5.5 ~ +0.10 on the probability scale) - large enough to change which model appears to over-commit least once its own-provider judge is excluded. Second, LLM judges are more permissive than clinicians on a blinded 50-item subsample: all four are significantly more lenient than the stricter independent clinician (crediting appropriate uncertainty on 66-84% of items vs 52%), and three of four than the author-influenced consensus (Grok directional only; judge-vs-consensus kappa = 0.20-0.43). On the author-audited clinical-underdetermined subset the permissiveness gap widened and the point-estimate model ordering held. A closed-ended MedQA anchor confirms accuracy is high and option-order effects are within a +/-5-point equivalence region for three of four models, so the safety gap is about calibration, not knowledge. We release the harness, prompts, per-item outputs, judge panel, perturbation audit, and human-annotation protocol.

13:00 JSTエージェントDeepSeek

PhoenixRepair: ソフトウェア エージェントでの修復戦略の探索を再考する

大規模言語モデルは自動化された問題解決を大幅に進歩させましたが、既存のエージェントベースの方法では修復戦略の検討が不十分であるという根本的な限界があります。この不十分さは 2 つの重要な側面に現れます。まず、複数の潜在的な編集場所の探索が制限されています。第二に、各場所での修復の試みの調査も不十分です。これらの課題に対処するために、複数の候補編集場所を系統的に探索し、パッチ生成で反復的な反映と改良を実行することにより、修復戦略の探索空間を拡大するマルチエージェント フレームワークである PhoenixRepair を紹介します。私たちのフレームワークは、複数の場所のサンプリングから始まり、オプションで難しいタスクのためのグラフベースの位置特定情報で強化され、その後、より良いパッチを生成するための反復的な反映と改良が続き、過去のすべての試みから蒸留された洞察に基づいて最終ラウンドの生成で最高潮に達します。 SWE-bench-Verified での実験では、PhoenixRepair が DeepSeek-V3.1 では SWE エージェントに対して 7.8\% という最大の相対的改善を達成し、MiniMax-M2.5 では 76.0\% Pass@1 という最高の解決率を達成することが実証されています。同時に、既存のアプローチよりも高い障害位置特定精度を実現します。私たちのコードは https://github.com/DeepSoftwareAnalytics/PhoenixRepair で入手できます。

原文 (English)

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents

While Large Language Models have greatly advanced automated issue resolution, existing agent-based methods exhibit a fundamental limitation in their insufficient exploration of repair strategies. This insufficiency manifests in two key aspects. First, the exploration of multiple potential edit locations is limited. Second, the exploration of repair attempts at each location is also insufficient. To address these challenges, we present PhoenixRepair, a multi-agent framework that systematically explores multiple candidate edit locations and performs iterative reflection and refinement on patch generation, thereby expanding the search space of repair strategies. Our framework begins with multi-location sampling, optionally augmented with graph-based localization information for difficult tasks, followed by iterative reflection and refinement to generate better patches, culminating in final-round generation guided by distilled insights from all historical attempts. Experiments on SWE-bench-Verified demonstrate that PhoenixRepair achieves the largest relative improvement of 7.8\% over SWE-agent under DeepSeek-V3.1, and attains the highest resolved rate of 76.0\% Pass@1 under MiniMax-M2.5. Meanwhile, it achieves higher fault localization accuracy than existing approaches. Our code is available at https://github.com/DeepSoftwareAnalytics/PhoenixRepair.

13:00 JST研究/論文

NaviAIS: ベクトル化されたレーン事前分布と NaviLane 予測フレームワークを備えたシナリオ レベルの船舶軌道予測データセット

複雑な海洋環境における船舶の軌道予測は、交通管理、衝突警告、航路計画、自律航行に不可欠です。 AIS ベースの学習方法は急速に進歩していますが、既存のデータセットは、一貫性のないサンプリング レート、ノイズの多い観測、異種の座標系、および統一されていないシナリオ プロトコルを伴う、生のメッセージ ストリームまたは不規則な時系列としてリリースされることがよくあります。また、ほとんどの公開 AIS リソースには、航行車線、水路の形状、航行可能領域の制約の構造化された表現が欠けており、再現可能な環境を考慮した予測が制限されています。これに対処するために、船舶軌道予測用の標準化されたシナリオレベルの AIS データセットである NaviAIS を導入します。統合された時間ウィンドウとローカル座標系内で複数の船舶の過去から将来の軌跡を整理し、ラスター化されたナビゲート可能なマップ、ベクトル化されたレーン事前分布、レーン グラフ、および構造化されたマップ表現を提供します。既存のデータセットと比較して、ベクトル化されたレーン、マルチシナリオ カバレッジ、ベクトル化されたマップ、オープン アクセシビリティ、および処理された軌跡を共同でサポートします。このデータセットに基づいて構築された、地図認識型予測のための階層マクロアクション フレームワークである NaviLane を提案します。 NaviLane は、まず統合されたシーン表現に対して軌跡マップ結合エンコーディングを実行し、次に離散マクロアクション コードブックを使用して、粗いものから洗練されたものまでのマルチモーダル候補を生成します。残差改良モジュールは局所的な幾何学的および動的一貫性を向上させ、ワールドモデルベースの結果を認識した評価者は相互作用リスクと環境実現可能性によって候補をランク付けします。実験では、NaviLane がシングルモーダル設定とマルチモーダル設定の両方で代表的なベースラインを上回るパフォーマンスを示し、構造化されたナビゲーション事前分布、階層的なマルチモーダル生成、および結果を意識した評価の価値が確認されました。

原文 (English)

NaviAIS: A Scenario-Level Vessel Trajectory Prediction Dataset withVectorized Lane Priors and the NaviLane Forecasting Framework

Vessel trajectory prediction in complex maritime environments is essential for traffic management, collision warning, route planning, and autonomous navigation. Although AIS-based learning methods have progressed rapidly, existing datasets are often released as raw message streams or irregular time series, with inconsistent sampling rates, noisy observations, heterogeneous coordinate systems, and non-unified scenario protocols. Most public AIS resources also lack structured representations of navigational lanes, waterway geometry, and navigable-region constraints, limiting reproducible, environment-aware forecasting. To address this, we introduce NaviAIS, a standardized scenario-level AIS dataset for vessel trajectory prediction. It organizes multi-vessel historical-future trajectories within unified temporal windows and local coordinate systems, and provides rasterized navigable maps, vectorized lane priors, lane graphs, and structured map representations. Compared with existing datasets, it jointly supports vectorized lanes, multi-scenario coverage, vectorized maps, open accessibility, and processed trajectories. Built on this dataset, we propose NaviLane, a hierarchical macro-action framework for map-aware prediction. NaviLane first performs trajectory-map joint encoding for a unified scene representation, then uses a discrete macro-action codebook to generate multimodal candidates coarse-to-refined. A residual refinement module improves local geometric and dynamical consistency, and a world-model-based consequence-aware evaluator ranks candidates by interaction risk and environmental feasibility. Experiments show NaviLane outperforms representative baselines in both single-modal and multimodal settings, confirming the value of structured navigational priors, hierarchical multimodal generation, and consequence-aware evaluation.

13:00 JST研究/論文

Black-Mamba: 分布ドリフト下での概念的知識のための生物学にインスピレーションを受けた漏洩蓄積

将来の観測値の条件付き分布は時間の経過とともに変化するため、現実世界の条件下での予測は本質的に非定常です。最近のテスト時適応シーケンス モデルは、推論中に内部状態を更新することでこの課題に対処しますが、適応を瞬間的な予測エラーや予期せぬ事態に結びつけます。この結合により、永続的な分布シフトと確率論的イノベーションが混同される可能性があり、不必要な更新や非効率的な適応につながる可能性があります。分布ドリフト下での証拠ゲート型状態追跡としてオンライン適応を定式化するテスト時適応予測アーキテクチャである Black-Mamba を紹介します。このモデルは、一時的に蓄積された驚きが体制変化の十分な証拠を提供するときに更新される動的メモリで基本予測子を強化します。これにより、適応が連続的なプロセスではなく、選択的なイベント駆動のプロセスに変わります。 Black-Mamba は、非定常ダイナミクスを使用した複数の予測ベンチマークにわたって、既存のテスト時適応手法と比較して、競合する、または向上した予測パフォーマンスを実現すると同時に、推論中のメモリ更新の数を大幅に削減します。これらの結果は、数学的分析と生物学的証拠と合わせて、蓄積された驚きが持続的なドリフトと一時的なノイズを区別するための原理的な信号を提供し、より効率的でロバストな適応をもたらすことを示唆しています。

原文 (English)

Black-Mamba: Biologically-Inspired Leaky Accumulation for Conceptual Knowledge under Distribution Drift

Forecasting under real-world conditions is inherently non-stationary, as the conditional distribution of future observations evolves over time. Recent test-time adaptive sequence models address this challenge by updating internal states during inference, but tie adaptation to instantaneous prediction errors or surprise. This coupling can conflate persistent distribution shift with stochastic innovations, leading to unnecessary updates and inefficient adaptation. We introduce Black-Mamba, a test-time adaptive forecasting architecture that formulates online adaptation as evidence-gated state tracking under distribution drift. The model augments a base predictor with a dynamic memory updated when temporally accumulated surprisal provides sufficient evidence of a regime change. This turns adaptation into a selective, event-driven process rather than a continuous one. Across multiple forecasting benchmarks with non-stationary dynamics, Black-Mamba achieves competitive or improved predictive performance compared to existing test-time adaptation methods while significantly reducing the number of memory updates during inference. Together with mathematical analysis and biological evidence, these results suggest that accumulated surprisal provides a principled signal for distinguishing persistent drift from transient noise, yielding more efficient and robust adaptation.

13:00 JST研究/論文

相対位置エンコーディングによる距離のエンコーディングによるトランスベースのルーティングの強化

この論文では、チーム オリエンテーリング問題を解決するための、Transformer アーキテクチャにおける付加的なバイアスとしての相対位置エンコーディング (RPE) について検討します。ルーティング問題を表すグラフのノード間のペアごとの空間関係をアテンション メカニズムに埋め込むことで、トランスフォーマー エンコーダーは、デコーダーがより適切なルートを推定できるようにする、より豊富な空間認識グラフ埋め込みを計算できます。最大 100 ノードのインスタンスを含む実験結果は、他の最先端の作品で使用されているバニラの Transformer アーキテクチャと比較して、収集される報酬と最適性のギャップが一貫して改善されていることを示しています。これらの発見は、明示的なリレーショナル モデリングにより、複雑な組み合わせ最適化のスケーラビリティと一般化が大幅に強化されることを強調しています。

原文 (English)

Enhancing Transformer-based Routing by Encoding Distance via Relative Positional Encoding

This paper explores Relative Positional Encoding (RPE) as an additive bias in Transformer architectures to solve the Team Orienteering Problem. By embedding in the attention mechanism pairwise spatial relationships among nodes of the graph that represents the routing problem, the transformer encoder can compute a richer spatial-aware graph embedding that allows the decoder to estimate better routes. Experimental results involving instances up to 100 nodes demonstrate consistent improvements in collected rewards and optimality gaps over vanilla Transformer architectures used by other state-of-the-art works. These findings highlight that explicit relational modeling significantly enhances scalability and generalization for complex combinatorial optimization.

13:00 JST研究/論文

OntoBook: 医療エンコーダーの事前トレーニングのためのオントロジーに基づいた合成教科書

我々は、医療オントロジー構造をエンコーダ言語モデルの事前トレーニング信号に変換する手法である OntoBook を紹介します。私たちのアプローチには 3 つの段階があります。オントロジー グラフを介したランダム ウォークによって医療コード間の階層関係と因果関係がキャプチャされ、大規模な言語モデルによってこれらのウォークが流暢な教科書スタイルの散文に再定式化され、結果として得られたテキストが、マスクされた言語モデリングとコード ペア間の関係予測という 2 つの目的で、同じデータに対して 149M パラメータのフランス語エンコーダである ModernCamemBERT をトレーニングするために使用されます。フランスの 3 つの医療コーディング ベンチマーク (FRACCO、Cantemist-FR、Distemist-FR) において、OntoBook は MLM のみの事前トレーニングと比較して大幅な改善を達成し、FRACCO では +2.5 micro-F1、Distemist では +8.0 micro-F1 を達成しました。目標間の調整が必要であることがわかりました。各タスクが異なるデータを使用する調整されていないトレーニングは、30 ポイントの低下を引き起こします。当社は、フランスの 3 つのオントロジー (CIM-10、CCAM、ATC) および事前トレーニング済みモデル チェックポイントにわたる、LLM で再編成された 130 万冊の医学教科書をリリースしています。

原文 (English)

OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining

We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models. Our approach has three stages: random walks through ontology graphs capture hierarchical and causal relations between medical codes, a large language model reformulates these walks into fluent textbook-style prose, and the resulting text is used to train ModernCamemBERT, a 149M-parameter French encoder, with two objectives on the same data: masked language modeling and relation prediction between code pairs. On three French medical coding benchmarks (FRACCO, Cantemist-FR, Distemist-FR), OntoBook achieves significant improvements over MLM-only pretraining, with +2.5 micro-F1 on FRACCO and +8.0 micro-F1 on Distemist. We find that alignment between objectives is necessary: misaligned training, where each task uses different data, causes a 30-point degradation. We release 1.3 million LLM-reformulated medical textbooks across three French ontologies (CIM-10, CCAM, ATC) and pretrained model checkpoints.

13:00 JST研究/論文

一般知性が必要とするもの: 記述レベル全体にわたる還元不可能な制約

人間の認知能力の全範囲を支える種類の一般知能は、計算アーキテクチャだけが持つ性質ではありません。この論文は単一のテーゼを前進させます。それは、特殊科学の伝統がこの用語に与えている意味で、一般知性に対する構造的制約は記述の異なるレベルを占めており、相互に還元不可能であるということです。したがって、単一のアーキテクチャの進歩やスケーリング プログラムの継続だけでは汎用人工知能 (AGI) を生み出すことはできず、研究プログラムは 1 つのベンチマークのパフォーマンスではなく、完全な制約プロファイルに対して評価する必要があるということになります。この論文は、AI システム研究、人類学、法律、経済学という 4 つの証拠のレンズを通して一般知性を読み取る方法を通じて開発されており、それぞれが明確な記述レベルに固定されており、正当化の文脈ではなく発見の文脈で規律あるヒューリスティックとして使用される推理小説によって補完されています。この方法を適用すると、8 つのクラスターに編成された 23 の構造的制約の分類が得られます。 6 つのレベルは詳細に調査され、レベルの昇順はしごとして順序付けされており、あるレベルでの進歩が次のレベルに持ち越せない理由を明確に示しています。この議論は 5 つの反証可能な予測で争われており、それぞれが名前付きのベンチマーク ファミリと反証条件で述べられており、記述的な枠組みをスケーリング仮説が示すよりも長い期間を見据えた研究プログラムに変換しています。

原文 (English)

What General Intelligence Requires: Non-Reducible Constraints Across Levels of Description

General intelligence, of the kind that underwrites the full range of human cognitive achievement, is not a property of computational architecture alone. This paper advances a single thesis: the structural constraints on general intelligence occupy distinct levels of description and are mutually non-reducible, in the sense that the special-sciences tradition gives to that term. It follows that no single architectural advance, and no continuation of the scaling programme by itself, can produce artificial general intelligence (AGI), and that research programmes must be evaluated against the full constraint profile rather than against performance on any one benchmark. The thesis is developed through a method that reads general intelligence through four evidential lenses, AI systems research, anthropology, law, and economics, each anchored to a distinct level of description, supplemented by speculative fiction used as a disciplined heuristic in the context of discovery rather than the context of justification. Applying the method yields a taxonomy of twenty-three structural constraints organised into eight clusters; six are examined in depth and ordered as an ascending ladder of levels, with explicit bridges showing why progress at one level cannot carry to the next. The argument issues in five falsifiable predictions, each stated with a named benchmark family and a disconfirmation condition, converting a descriptive framework into a research programme with a longer horizon than the scaling hypothesis implies.

13:00 JSTLLM/生成AIハードウェア/半導体

依存関係から構成性へ: 組み合わせカテゴリ文法による LLM 出力の神経記号的リフティング

大規模言語モデル (LLM) は、接頭辞から次のトークンを段階的に予測することにより、流暢なテキストを生成します。生成の伝統の批評家は、そのようなシステムには真の文法が欠けていると主張します。依存関係文法の観点からの影響力のある回答は、LLM の動作は、単語ごとに構築されたローカルのヘッド依存構造によってよく記述されると主張しています。私たちは、より鋭い観察が見落とされてきたと主張します。つまり、自己回帰生成の接頭辞駆動型補完ダイナミクスは、結合カテゴリ文法 (CCG) が元々サポートするように設計された増分処理モデルと密接に一致しています。これに基づいて、LLM の出力が型付けされた構成導出に持ち上げられる神経象徴的なフレームワークを提案します。LLM が CCG を内部的に実装しているとは主張しませんが、その出力は原則に基づいた増分的で監査可能な CCG 再構築を可能にすると主張します。 2 つの結果が続きます。まず、カリーとハワードの対応を通じて、リフティングは自然言語を超えて、LLM が生成する形式言語 (Solidity などのプログラミング言語、記述ロジック、OWL や SQL などのクエリ言語) まで拡張され、型システムは変化し、アーキテクチャは固定されています。第 2 に、リフティングでは 2 つのチェック層がサポートされています。構造上の欠陥を直接検出する構成層と、リフトされた構造を外部の知識ソースと照合してチェックするコンテンツ層です。これにより、幻覚コンテンツの可能な限り早期のフラグ付けが可能になります。したがって、アカウントはプロデューサーに認識ではなくプレフィックス駆動の生成プロファイルを要求します。最後に、フレームワークが開く一方向としての同期 LLM-CCG カップリングのスケッチを示します。

原文 (English)

From Dependency to Compositionality: A Neurosymbolic Lifting of LLM Outputs via Combinatory Categorial Grammar

Large language models (LLMs) generate fluent text by incrementally predicting the next token from a prefix. Critics in the generative tradition argue that such systems lack genuine grammar; influential replies from the dependency-grammar perspective hold that LLM behavior is well described by local head-dependent structure built word by word. We argue that a sharper observation has been overlooked: the prefix-driven, type-completing dynamics of autoregressive generation align closely with the incremental processing model that Combinatory Categorial Grammar (CCG) was originally designed to support. On this basis we propose a neurosymbolic framework in which LLM outputs are lifted into typed compositional derivations -- not claiming that LLMs implement CCG internally, but that their outputs admit a principled, incremental, and auditable CCG reconstruction. Two consequences follow. First, through the Curry-Howard correspondence the lifting extends beyond natural language to the formal languages LLMs also produce -- programming languages such as Solidity, description-logic and query languages such as OWL and SQL -- with the type system varying and the architecture held fixed. Second, the lifting supports two layers of checking: a compositional layer that catches structural failures directly, and a content layer that checks the lifted structure against external knowledge sources, enabling the earliest possible flagging of hallucinated content. The account thereby requires of a producer not cognition but a prefix-driven generative profile. We close with a sketch of synchronous LLM-CCG coupling as one direction the framework opens.

13:00 JSTLLM/生成AIOpenAIGPT / ChatGPT

対照的な信念の更新による報酬追求の測定

強化学習でトレーニングされた言語モデルは、意図された目的ではなく、採点者の判断を最適化することを学習する可能性があります。採点者の判断を追求するモデルと意図された目的を追求するモデルは、採点者が意図した行動に報酬を与えるときは常に同じように動作するため、この「報酬の追求」を測定することは困難です。私たちは、対照的合成ドキュメント微調整を使用して報酬追求を測定し、採点者が報酬を与えるものについてのモデルの信念を変更し、それらの信念をユーザーまたは開発者が望むものと矛盾させ、モデルが各当事者の好ましい動作を採用する割合を測定します。安全性トレーニングなしで、機能に重点を置いた OpenAI o3 RL 実行の中間チェックポイントに適用すると、これらのチェックポイントは、コーディングや調整タスクに関して、ユーザーや開発者の好みよりも採点者の好みに従うことが多いことがわかりました。この採点者の側に立つ傾向は、RL トレーニング全体を通して増加傾向にあります。たとえば、タスクを完了するために監督者との約束を守るか破るかの選択を迫られる環境では、遅れて能力に焦点を当てた o3 チェックポイントが、SDF の文書で採点者がタスクの完了に報酬を与えると記載している場合は 87% の確率で約束を破りますが、SDF が誠実さに報酬を与えると述べている場合は 9% の確率で約束を破ります (SDF の思考回路がしばしば明示する選択です)。初期のチェックポイントの感度ははるかに低くなります (40% 対 24%)。私たちの手法は、報酬ハッキング モデルにも一般化されています。報酬ハッキングを行うように訓練されたモデル生物 (gpt-oss-120b) は、未修正のモデルに比べて採点者の好みに対して 2 倍以上敏感であり、採点者に有利な平均行動変化は 33% から 86% に上昇しました。これらの結果は、RL がトレーニングの過程で報酬の追求を強め、開発者がより高い報酬につながると信じている場合に、開発者の意図に反して動作する可能性のあるモデルを作成する可能性があることを示しています。

原文 (English)

Measuring Reward-Seeking via Contrastive Belief Updates

Language models trained with reinforcement learning may learn to optimize the grader's judgment rather than the intended objective. This "reward-seeking" is difficult to measure because a model that pursues the grader's judgment and one that pursues the intended objective behave identically whenever the grader rewards the intended behavior. We measure reward-seeking using Contrastive Synthetic Document Finetuning to change a model's beliefs about what the grader rewards, putting those beliefs in conflict with what users or developers want, and measuring the rate at which the model adopts each party's preferred behavior. Applied to intermediate checkpoints of a capabilities-focused OpenAI o3 RL run, without safety training, we find that these checkpoints often side with grader preferences over those of users or developers on coding and alignment tasks. This tendency to side with the grader trends upward throughout RL training. For example, in an environment that forces a choice between keeping a promise to a supervisor and breaking it to complete the task, a late capabilities-focused o3 checkpoint breaks the promise 87% of the time when SDF documents say the grader rewards task completion, versus 9% when they say it rewards honesty (a choice its chain-of-thought often makes explicit). An earlier checkpoint is far less sensitive (40% vs. 24%). Our method also generalizes to reward-hacking models. A model organism trained to reward-hack (gpt-oss-120b) is more than twice as sensitive to grader preferences as the unmodified model, with the mean behavioral shift in favor of the grader rising from 33% to 86%. These results indicate that RL can increase reward-seeking over the course of training, producing models that may act against their developers' intentions when they believe that doing so leads to higher reward.

13:00 JST研究/論文

Mi-Memory: パーソナル AI のためのライフサイクル メモリ フレームワーク

パーソナル AI は、チャットのみの対話を超えて、電話、自動車、家庭、ウェアラブル、カメラ、ツールにわたる継続的なサービスへと移行しています。この設定では、メモリを以前の会話のキャッシュとして残すことはできません。それは、継続性とガバナンスの基盤として機能する必要があります。つまり、永続的なユーザー状態を保持し、マルチモーダルおよびデバイスの証拠に答えを根拠付け、修正と忘却をサポートし、ポリシーの進化を制限し、レイテンシー、コスト、プライバシー、およびエッジクラウドの制約の下で展開可能性を維持する必要があります。この技術レポートでは、構造、拡張、進化、展開の 4 つの役割を中心に構成されたパーソナル AI のライフサイクル メモリ フレームワークである Mi-Memory について説明します。共有監査契約は、これらの役割を 4 つの定期的なアーティファクト ファミリを通じてリンクします。型付き証拠ペイロードはソースの ID と来歴を保持し、診断トレースはサービス パイプライン全体で証拠損失を特定し、戦略アーティファクトはメモリ ポリシーの変更を明示し、ゲート/ロールバック レコードは受け入れられた進化をバインドします。 MiMemory は、MemStack、MemSense/MemFuse、D$^{2}$ACCI/E$^{2}$MEND、および LiteMem を通じてロールをインスタンス化します。制御参照構造の評価では、MemStack は LoCoMo、ペルソナメム V2、および LongMemEval でそれぞれ 93.59%、57.24%、および 87.47% に達しました。他のトラックは、モジュールレベル、予備的/内部的、移転実現可能性、または明示的な境界を伴う設計のみの証拠を報告します。 MiMemory は、パーソナル AI 向けの監査可能、証拠ゲート型、展開対応型メモリ システムへの一歩です。プロジェクトのホームページ: https://darwin-agent.github.io/Mi-Memory/ 。

原文 (English)

Mi-Memory: A Lifecycle Memory Framework for Personal AI

Personal AI is moving beyond chat-only interaction toward continuous services that span phones, cars, homes, wearables, cameras, and tools. In this setting, memory cannot remain a cache of prior conversations. It should serve as a continuity and governance substrate: preserving durable user state, grounding answers in multimodal and device evidence, supporting correction and forgetting, bounding policy evolution, and remaining deployable under latency, cost, privacy, and edge-cloud constraints. This technical report presents Mi-Memory, a lifecycle memory framework for Personal AI organized around four roles: Structure, Expansion, Evolution, and Deployment. A shared audit contract links these roles through four recurring artifact families: typed evidence payloads preserve source identity and provenance, diagnostic traces localize evidence loss across the serving pipeline, strategy artifacts make memory-policy changes explicit, and gate/rollback records bound accepted evolution. MiMemory instantiates the roles through MemStack, MemSense/MemFuse, D$^{2}$ACCI/E$^{2}$MEND, and LiteMem. In controlled-reference Structure evaluations, MemStack reaches 93.59%, 57.24%, and 87.47% on LoCoMo, PersonaMem-V2, and LongMemEval, respectively; other tracks report module-level, preliminary/internal, transfer-feasibility, or design-only evidence with explicit boundaries. MiMemory is a step toward auditable, evidence-gated, and deployment-aware memory systems for Personal AI. Project homepage: https://darwin-agent.github.io/Mi-Memory/ .

13:00 JSTLLM/生成AI

フリーライダーを釣り出す: 強化学習による並列推論のためのシャプレーベースの報酬帰属

大規模言語モデル (LLM) は複数ステップの推論に優れていますが、現在の並列推論アプローチでは、個々の推論パスの寄与を区別できないことがよくあります。多くのパスは冗長で誤解を招き、さらには有害である可能性がありますが、結果レベルの報酬は一律の報酬を割り当てるため、学習シグナルが曖昧になり、トレーニングが不安定になります。我々は、マルチパス推論におけるきめの細かいパスレベルの寄与を明らかにする強化学習フレームワークである Parallel Shapley を提案します。各パスを協力ゲームのプレイヤーとして扱い、シャプレー値を利用して限界寄与を定量化し、生成報酬モデルを使用してパスのユーティリティを評価し、モンテカルロ サンプリングを使用して効率的な近似を行います。数学的推論ベンチマークの実験では、Parallel Shapley が既存のベースラインを上回り、より安定した解釈可能なトレーニングを提供することが示されています。私たちのフレームワークは効果的に「フリーライダーを釣り出し」、報酬を比例的に割り当て、LLM でのマルチパス推論を改善します。

原文 (English)

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Many paths may be redundant, misleading, or even detrimental, but outcome-level rewards assign uniform reward, leading to ambiguous learning signals and unstable training. We propose Parallel Shapley, a reinforcement learning framework that attributes fine-grained, path-level contributions in multi-path reasoning. Treating each path as a player in a cooperative game, we leverage Shapley values to quantify marginal contributions, using a generative reward model to evaluate path utilities and Monte Carlo sampling for efficient approximation. Experiments on mathematical reasoning benchmarks show that Parallel Shapley outperforms existing baselines while providing more stable and interpretable training. Our framework effectively "fishes out the free riders," assigning reward proportionally and improving multi-path reasoning in LLMs.

13:00 JSTLLM/生成AIロボティクス

Athena-Brain テクニカル レポート: 一般知能と身体的インタラクションのための効率的なロボット ブレイン

大規模言語モデル (LLM) は、言語理解、推論、世界知識において顕著な能力を実証してきました。身体化されたエージェントの能力がますます高まるにつれ、デバイス上の頭脳として機能し、LLM の広範な一般知能を維持しながら、身体化された環境との効果的な高レベルの対話を可能にするコンパクトなモデルの需要が高まっています。ただし、既存のアプローチでは、汎用インテリジェンスまたは特殊な組み込み機能のいずれかを優先することが多く、単一モデル内で両方の要件を満たすことが困難になっています。私たちは、身体化された知性のための身体化された知性のためのオンデバイスの頭脳として機能するように設計された 8B LLM である \textbf{Athena-Brain-8B} を紹介します。一般的な教師あり微調整、一般的な強化学習、エンボディド エキスパート トレーニング、モデル マージで構成される多段階のポストトレーニング パイプラインを通じて、Athena-Brain-8B は強力な一般機能を維持しながら、強力な高レベルのエンボディド インタラクション機能を獲得し、効率的なエンボディド インタラクションのための簡潔な応答を生成します。実験結果は、一般的な評価と具体的な評価の両方にわたって Athena の有効性を示しています。対応する Qwen3-8B 思考モデルと比較して、Athena-Brain-8B は、大幅に短い応答を生成しながら、一般言語および推論ベンチマークで同等のパフォーマンスを達成します。ドメイン内の組み込みベンチマークでは、Athena-Brain-8B は一貫して同様の規模のモデルを上回り、ゼロショットで評価されたいくつかの大幅に大規模なフロンティア モデルを上回っています。これは、コンパクトな言語モデルが強力な汎用インテリジェンスを組み込み機能と効果的に統合できることを示しています。

原文 (English)

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio

Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, there is a growing demand for compact models that can serve as an on-device brain, preserving the broad general intelligence of LLMs while enabling effective high-level interaction with embodied environments. Existing approaches, however, often prioritize either general-purpose intelligence or specialized embodied capabilities, making it challenging to satisfy both requirements within a single model. We present \textbf{Athena-Brain-8B}, an 8B LLM designed to serve as an on-device brain for embodied intelligence for embodied intelligence. Through a multi-stage post-training pipeline consisting of General Supervised Fine-Tuning, General Reinforcement Learning, Embodied Expert training, and Model Merge, Athena-Brain-8B maintains strong general capabilities while acquiring strong high-level embodied interaction capabilities and generating concise responses for efficient embodied interaction. Experimental results demonstrate the effectiveness of Athena across both general and embodied evaluations. Compared with the corresponding Qwen3-8B thinking model, Athena-Brain-8B achieves comparable performance on general language and reasoning benchmarks while generating substantially shorter responses. On in-domain embodied benchmarks, Athena-Brain-8B consistently outperforms models of similar scale and surpasses several substantially larger frontier models evaluated zero-shot, demonstrating that compact language models can effectively integrate strong general intelligence with embodied capabilities.

13:00 JST研究/論文

Vector-Bench: モデルは SVG コードを外科的に編集できますか?

命令ベースのベクトル編集には 2 つの機能が必要です。要求された変更を行うことと、その他のすべてをそのままにすることです。 2 つ目は、出力がラスター イメージとしてのみ判断される場合に見落とされやすくなります。 40 個の SVG 修復タスクのコンパクトで難しいベンチマークである Vector-Bench を紹介します。各タスクは、破損した SVG プログラムと作成者が作成した視覚的な指示、隠されたターゲット プログラム、平均 5.05 個の注釈付き修復、および平均 60.55 個の保護オブジェクトを組み合わせます。指示では、要素識別子、座標、カラーコード、またはパスデータを公開することなく、目に見える欠陥について説明します。決定的なバイナリ仕様の報酬を定義します。要求された修復では属性を認識した知覚許容誤差が使用されますが、要求されていないレンダリングまたはアプリケーション関連の構造は意味的に変更されず、結果は有効な SVG でなければなりません。正規のターゲットの同等性とより厳密なソースの忠実性は診断として保持されます。妥当性ゲートによる修復の進行状況、ほぼ完了した段階、および有効な出力の意図しない変更率 (UCR) によって、部分的な結果が説明されます。 1,360 件のリクエストにわたって 34 のモデル エンドポイント (オープンウェイトとしてリストされた 25 件、安価なコントロール 5 件、フロンティア クローズド エンドポイント 4 件) を評価しました。最も強力なエンドポイントは、平均修復進捗率が 43.7% であるにもかかわらず、仕様全体の成功率が 15.0% にとどまっており、見かけの修復進捗率と仕様に忠実な編集が依然として大幅に異なることを示しています。すべてのプロンプト、出力、スコアリング コード、コスト、およびタスクごとのレポートがリリースされます。

原文 (English)

Vector-Bench: Can Models Surgically Edit SVG Code?

Instruction-based vector editing requires two capabilities: making a requested change and leaving everything else alone. The second is easy to miss when an output is judged only as a raster image. We introduce Vector-Bench, a compact, difficult benchmark of 40 SVG repair tasks. Each task pairs a corrupted SVG program with an author-written visual instruction, a hidden target program, 5.05 annotated repairs on average, and an average of 60.55 protected objects. Instructions describe visible defects without exposing element identifiers, coordinates, color codes, or path data. We define a deterministic binary specification reward: requested repairs use attribute-aware perceptual tolerances, while unrequested rendering- or application-relevant structure must remain semantically unchanged and the result must be a valid SVG. Canonical target equality and stricter source fidelity are retained as diagnostics. Validity-gated repair progress, a near-complete tier, and valid-output Unintended Change Rate (UCR) explain partial outcomes. We evaluate 34 model endpoints (25 listed as open-weight, 5 inexpensive controls, and 4 frontier closed endpoints) over 1360 requests. The strongest endpoint reaches only 15.0% full specification success, despite 43.7% mean repair progress, showing that apparent repair progress and specification-faithful editing remain substantially different. All prompts, outputs, scoring code, costs, and per-task reports are released.

13:00 JST研究/論文

品質保証: VR OSCE における審査官クレームのマルチモーダル検証

客観的構造化臨床検査 (OSCE) は臨床能力を評価するためのゴールドスタンダードですが、採点は依然として検査者の主観、疲労、認知バイアスの影響を受けやすいです。評価者間統計による標準的な審査官の検証は、審査官の推論を分析したり、実際の事象に対する審査官の主張を検証したりしないため、誤りの原因に関する説明力に欠けています。そこで、我々は、ビデオ、VR ログ、俳優データから構築された実際の一連のイベントに対して審査官が主張した行動を比較することにより、バーチャル リアリティ (VR) 小児 OSCE における審査官の主張を検証するマルチモーダル フレームワークである品質アクション保証 (QAA) を導入します。 QAA は、アクションの位置特定とアクターのソースの帰属を実行する制約付きの時間的アクション アライメント モデルと、審査官の主張を抽出して記録と照合する大規模な言語モデルを組み合わせます。 5 分割相互検証を通じて、QAA は時間的アライメントに関して 99.2% $\pm$ 0.7% Actor F1 および 93.4% $\pm$ 1.9% W@16 を達成しました。全体的に、QAA は 70.0% の精度と 76.7% の再現率で検査官のミスを検出し、事実の正確性が 39.2% から 79.2% に向上し、より公平な OSCE 評価が可能になります。

原文 (English)

Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs

Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias. Standard examiner validation via inter-rater statistics lacks explanatory power regarding the source of errors, as it neither analyzes examiner reasoning nor verifies examiner claims against actual events. Thus, we introduce Quality Action Assurance (QAA), a multimodal framework that verifies examiner claims in Virtual Reality (VR) pediatric OSCEs by comparing actions claimed by examiners against the true sequence of events, constructed from video, VR logs, and actor data. QAA combines a constrained temporal action alignment model, which performs action localization and actor source attribution, with a large language model that extracts examiner claims and checks them against the record. Across a 5-fold cross-validation, QAA achieves 99.2% $\pm$ 0.7% Actor F1 and 93.4% $\pm$ 1.9% W@16 for temporal alignment. Overall, QAA detects examiner errors with 70.0% precision and 76.7% recall, improving factual correctness from 39.2% to 79.2%, enabling fairer OSCE assessment.

13:00 JST研究/論文

グラフ組み合わせ最適化の事前学習の有効性について

この論文では、巡回セールスマン問題のようなルーティング問題の性質に対処するために特別に設計された、グラフの組み合わせ最適化のための自己教師あり事前トレーニング フレームワークを紹介します。幾何学的な拡張 (特に回転と軸反射) を備えたグラフ対比学習を利用することにより、モデルは不変の構造表現とグローバルな相対距離分布を学習するように強制されます。結果は、この事前トレーニング戦略が、さまざまな問題スケールにわたって事前トレーニングされていないモデルよりも優れていることを示しています。特に、ハイブリッド戦略 (回転と反射の組み合わせ) は、TSP1000 のツアー長で 6.57% の改善を達成し、幾何学的な事前トレーニングがニューラル ソルバーを高次元のインスタンスに効果的にスケーリングするための重要な帰納的バイアスであることを証明しました。

原文 (English)

On the Effectiveness of Pretraining for Graph Combinatorial Optimization

This paper introduces a self-supervised pretraining framework for graph combinatorial optimization specifically designed to address the nature of routing problems like the Traveling Salesman Problem. By utilizing graph contrastive learning with geometric augmentations (specifically, rotations and axial reflections) the model is forced to learn invariant structural representations and global relative distance distributions. Results demonstrate that this pretraining strategy outperforms non-pretrained models across various problem scales. Notably, the hybrid strategy (combining rotation and reflection) achieved a 6.57% improvement in tour length for TSP1000, proving that geometric pretraining is an important inductive bias for effectively scaling neural solvers to high-dimensional instances.

13:00 JSTLLM/生成AIエージェントClaude

Supra コグニティブ モード: エージェント メモリのルーテッド アーキテクチャ

エージェントのメモリ ワークロードには、直接的な事実の検索、関係チェーンと現在の状態の推論、長い歴史にわたる広範な合成が混在しています。明示的または自動的に選択されたクエリごとのモードを、1 つの共有取り込み基板上の検索および合成ペイロードにマッピングするアーキテクチャである Supra Cognitive Modes (SCM) について説明します。凍結されたセマンティック分類子とランタイム ゲートは、融合された語彙検索と密な検索、グラフまたは反復マルチホップ処理、および階層化された長形式合成の間でクエリをディスパッチします。このサブストレートは、複数の粒度の埋め込み、抽出されたトリプル、ファクト バージョンのメタデータ、およびオプションの非同期エンリッチメントを組み合わせます。導入された構成を、長期会話メモリ (LoCoMo; n = 1,986)、MemoryAgentBench (MAB; n = 3,671)、および LongMemEval (n = 500) の 3 つのベンチマークで特徴付けます。リファレンス実行では、LoCoMo ファクトイド カテゴリで 84.87%、敵対的棄権で 68.61%、2 回の繰り返しで MAB で 61.49%、LongMemEval で 86.00% を記録しました。リポジトリに基づく再現では、同様の集計スコアが生成され、タスクおよびモード条件付きの障害分析がサポートされます。生のベースライン出力、LoCoMo と LongMemEval の調整されたエンドツーエンドのタイミング、および完全なトークン台帳は利用できません。格納された行では、実行時の最終決定の一部も省略されます。結果は、実装された 1 つのルーテッド構成とその診断障害パターンを特徴づける一方、ソース検査ではクエリごとの制御インターフェイスと共有基板設計を検証します。因果関係のあるルーティングの効果、効率の向上、および統計的有意性は、入手可能な証拠の範囲外にあります。

原文 (English)

Supra Cognitive Modes: A Routed Architecture for Agent Memory

Agent-memory workloads mix direct factual lookup, relation-chain and current-state reasoning, and broad synthesis over long histories. We describe Supra Cognitive Modes (SCM), an architecture that maps explicit or automatically selected per-query modes to retrieval and synthesis payloads over one shared ingest substrate. A frozen semantic classifier and runtime gates dispatch queries among fused lexical and dense lookup, graph or iterative multi-hop handling, and stratified long-form synthesis. The substrate combines multi-granularity embeddings, extracted triples, fact-version metadata, and optional asynchronous enrichments. We characterize the deployed configuration on three benchmarks: Long-term Conversational Memory (LoCoMo; n = 1,986), MemoryAgentBench (MAB; n = 3,671), and LongMemEval (n = 500). The reference run records 84.87% on LoCoMo factoid categories and 68.61% on adversarial abstention, 61.49% on MAB across two repetitions, and 86.00% on LongMemEval. A repository-backed reproduction produces similar aggregate scores and supports task- and mode-conditioned failure analysis. Raw baseline outputs, aligned end-to-end timing for LoCoMo and LongMemEval, and complete token ledgers are unavailable; stored rows also omit some final runtime decisions. The results characterize one implemented routed configuration and its diagnostic failure patterns, while source inspection verifies the per-query control interface and shared-substrate design. Causal routing effects, efficiency gains, and statistical significance remain outside the available evidence.

13:00 JST研究/論文

OpenRTAG: データ品質の低下下での堅牢なテキスト属性グラフ学習のための包括的なベンチマーク

テキスト属性グラフ (TAG) は、リレーショナル構造とリッチ ノード テキストを組み合わせた重要なグラフ データ形式です。ただし、現実世界の TAG は不完全であることが多く、テキスト、構造、ラベルから品質の問題が発生し、通常は疎性、ノイズ、不均衡として現れます。これらの次元は、TAG 学習に大きな影響を与える可能性がある 9 つの代表的な劣化シナリオを定義します。これまでの研究では特定の緩和戦略が検討されてきましたが、既存の証拠は劣化の種類、データセット、タスク、モデルファミリーにわたって断片的なままであり、TAG の堅牢性については十分に理解されていないままです。このギャップに対処するために、テキスト属性のグラフ学習の堅牢性ベンチマークである OpenRTAG を紹介します。 OpenRTAG は、TAG の品質問題を統一された 3*3 分類法に整理し、9 つの TAG データセットと 3 つの下流タスクにわたる標準化された評価をサポートします。シナリオの妥当性とモデルの感度を体系的に評価し、従来の GNN、LLM-GNN、および代表的な GFM を比較し、シナリオに一致するベースラインの有効性、効率、堅牢性を調査し、複合劣化シナリオの下でのモデルの動作をさらに調査します。 OpenRTAG は、現実的な低品質設定下での TAG 学習の堅牢性を理解するための標準化されたテストベッドを提供します。

原文 (English)

OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation

Text-attributed graphs (TAGs) are an important graph data form that combine relational structure with rich node text. However, real-world TAGs are often imperfect, with quality issues arising from text, structure, and labels, and typically manifesting as sparsity, noise, and imbalance. These dimensions define nine representative degradation scenarios that can substantially affect TAG learning. Although prior studies have explored specific mitigation strategies, existing evidence remains fragmented across degradation types, datasets, tasks, and model families, leaving TAG robustness insufficiently understood. To address this gap, we present OpenRTAG, a robustness benchmark for text-attributed graph learning. OpenRTAG organizes TAG quality issues into a unified 3 * 3 taxonomy and supports standardized evaluation across nine TAG datasets and three downstream tasks. It systematically evaluates scenario validity and model sensitivity, compares traditional GNNs, LLM-GNNs, and a representative GFM, investigates the effectiveness, efficiency, and robustness of scenario-matched baselines, and further examines model behavior under composite degradation scenarios. OpenRTAG provides a standardized testbed for understanding robustness in TAG learning under realistic low-quality settings.

13:00 JSTエージェント

パラメータ化されたアクション強化学習におけるマルチエージェントアクター-クリティックアルゴリズムの比較研究

パラメータ化されたアクション強化学習は、離散的なアクションの選択と連続的なパラメータ化の両方を必要とする環境で優れたパフォーマンスを示しています。これまでの研究では、パラメータ化されたアクション タスクのベンチマークにおいて、単一エージェントのアクター - 批評家アルゴリズム (Greedy Actor-Critic (GAC)、Soft Actor-Critic (SAC)、および Truncated Quantile Critics (TQC)) の有効性が確立されていますが、マルチエージェント設定への拡張についてはほとんど調査されていません。この論文では、マルチエージェント貪欲アクター批評家 (MAGAC)、マルチエージェントソフトアクター批評家 (MASAC)、およびマルチエージェント切り捨て分位批評家 (MATQC) のアルゴリズムの共有エクスペリエンス マルチエージェント拡張の比較研究を示します。提案されたフレームワークは、集中トレーニング、分散実行 (CTDE) パラダイムに従うのではなく、個別のポリシーと価値のネットワークを維持しながら、リプレイ バッファーを共有する複数の独立したアクター - 批評家エージェントを使用します。スケーラビリティを評価するために 3 エージェント、5 エージェント構成、および 10 エージェント構成を使用して、Platform-v0 および Goal-v0 ベンチマークのアルゴリズムを単一エージェントのベンチマークと比較して評価します。パフォーマンスは、統計的有意性を評価するために一元配置 ANOVA および Tukey HSD ポストホック テストを使用して、10 回の独立した実行にわたる平均評価リターンとトレーニング時間によって測定されます。結果は、マルチエージェント フレームワークが Greedy Actor-Critic のパフォーマンスを一貫して向上させる一方、MASAC と MATQC はシングル エージェント バージョンと比べて比較的緩やかな向上を示していることを示しています。エージェントの数を 5 つを超えて増やすと、特に MAGAC の場合、計算コストが大幅に増加する一方で、追加のパフォーマンスが制限されます。これらの結果は、学習パフォーマンスと計算効率の間のトレードオフを浮き彫りにし、パラメータ化されたアクション強化学習のための共有経験マルチエージェントアクタークリティカル手法の拡張性についての洞察を提供します。

原文 (English)

Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning

Parameterized action reinforcement learning has shown strong performance in environments requiring both discrete action selection and continuous parameterization. Prior work established the effectiveness of single-agent actor-critic algorithms - Greedy Actor-Critic (GAC), Soft Actor-Critic (SAC), and Truncated Quantile Critics (TQC) - on benchmark parameterized action tasks, but their extension to multi-agent settings remains largely unexplored. This paper presents a comparative study of shared-experience multi-agent extensions of these algorithms: Multi-Agent Greedy Actor-Critic (MAGAC), Multi-Agent Soft Actor-Critic (MASAC), and Multi-Agent Truncated Quantile Critics (MATQC). Rather than following the centralized training, decentralized execution (CTDE) paradigm, the proposed framework uses multiple independent actor-critic agents that share a replay buffer while maintaining separate policy and value networks. We evaluate the algorithms on the Platform-v0 and Goal-v0 benchmarks against their single-agent counterparts, using three-, five-, and ten-agent configurations to assess scalability. Performance is measured by average evaluation return and training time across ten independent runs, with one-way ANOVA and Tukey HSD post-hoc tests used to assess statistical significance. Results show that the multi-agent framework consistently improves Greedy Actor-Critic performance, while MASAC and MATQC show comparatively modest gains over their single-agent versions. Increasing the number of agents beyond five yields limited additional performance while substantially raising computational cost, particularly for MAGAC. These results highlight a trade-off between learning performance and computational efficiency, offering insight into the scalability of shared-experience multi-agent actor-critic methods for parameterized action reinforcement learning.

13:00 JST研究/論文

マルチリレーショナル グラフ畳み込みネットワークを使用した逐次学習者モデリング

ユーザー モデリングは、さまざまなパーソナライズされたシステムにおいて重要なタスクです。グラフ構造データからの学習における有効性が認識され、グラフ ニューラル ネットワーク (GNN)、特にグラフ畳み込みネットワーク (GCN) がユーザー モデリングに採用されることが増えています。ただし、既存のアプローチは通常、グラフ内のさまざまな関係タイプを同質なものとして扱い、より豊富なセマンティクスを取得し、より有益なユーザー モデルを構築する能力を制限します。マルチリレーショナル GNN (MR-GNN) は表現の学習と推奨に採用されていますが、ユーザー モデリングへの応用は未開発のままです。さらに、既存の GNN ベースのユーザー モデリング アプローチは、ユーザー対話シーケンスを無視します。これらの研究ギャップに対処するために、この研究では、マルチリレーショナル GCN (MR-GCN) を使用した概念ベースの逐次学習者モデリングに焦点を当てた、新しい完全に教師なしのアプローチである MR-ConceptGCN を提案します。 MR-ConceptGCN は、パーソナル ナレッジ グラフ (PKG)、MR-GCN、および事前トレーニングされた言語モデル SBERT を効果的に組み合わせて、PKG 項目の強化された関係および意味論を意識した表現を取得します。その後、CourseMapper で学習教材を操作する際に学習者が理解できなかった知識概念の強化された埋め込みを使用して、長期および短期の学習者インタラクションを組み合わせた逐次学習者モデルを構築します。我々は、オンライン ユーザー調査 (n = 31) の結果を報告し、正確さ、有用性、多様性、教育推薦システムの満足度など、いくつかの重要なユーザー中心の側面の観点から MR-ConceptGCN の利点を実証しています。

原文 (English)

Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks

User modeling is a critical task in a variety of personalized systems. Recognizing their effectiveness in learning from graph-structured data, Graph Neural Networks (GNNs), particularly Graph Convolutional Networks (GCNs), are increasingly employed for user modeling. However, existing approaches typically treat different relation types in a graph as homogeneous, limiting their ability to capture richer semantics and construct more informative user models. While multi-relational GNNs (MR-GNNs) have been adopted for representation learning and recommendation, their application for user modeling remains unexplored. Moreover, existing GNN-based user modeling approaches ignore the user interaction sequence. To address these research gaps, in this work we propose MR-ConceptGCN, a novel fully unsupervised approach focused on concept-based sequential learner modeling using multi-relational GCNs (MR-GCNs). MR-ConceptGCN effecively combines Personal Knowledge Graphs (PKGs), MR-GCNs, and the pre-trained language model SBERT to obtain enhanced relation- and semantic-aware representations of the PKG items. The enriched embeddings of the knowledge concepts that a learner did not understand when interacting with learning materials in CourseMapper are then used to construct a sequential learner model that combines long-term and short-term learner interactions. We report the results of an online user study (n = 31), demonstrating the benefits of MR-ConceptGCN in terms of several important user-centric aspects including accuracy, usefulness, diversity, and satisfaction with an educational recommender system.

13:00 JSTエージェント研究/論文ClaudeGPT / ChatGPT

BioSecBench-Surveillance: 病原体ゲノム監視における AI エージェントの検証可能なベンチマーク

病原体ゲノム監視が拡大するにつれて、ボトルネックはデータ生成から分析へと移りつつあります。 AI エージェントが生のシーケンス データと監視コンテキストから適切な分析パイプラインを推測できるかどうかをテストする 100 の評価の検証可能なベンチマークである BioSecBench-Surveillance を紹介します。各評価では、人間のアナリストが持つデータとコンテキストのみがエージェントに提供され、構造化された回答が決定的に評価されます。タスクは、分類学的分類から遺伝子工学による検出まで、多様なサンプルタイプとシーケンス技術にわたる 7 つのカテゴリに及びます。 16 のモデルとハーネスのペアから 3,962 回のグレーディング試行を行った結果、最も強い構成でも約半分しかクリアできませんでした。 PI を使用した Opus 4.8 は 50.2% でリードし、83 の評価全体で 95% 信頼区間は 40.1 ~ 60.3% で、Codex を使用した GPT-5.5 と同率 50.2%、95% 信頼区間は 40.8 ~ 59.6%、続いて Opus 4.7 は PI が 49.6%、95% 信頼区間は 95% でした。 40.0 ~ 59.2 パーセント、Sonnet 4.6 の PI は 48.6 パーセント、95 パーセント信頼区間は 38.9 ~ 58.3 パーセントでした。エージェントが正しいワークフローを呼び出した場合でも、その間違いは、どの参照、しきい値、フィルター、正規化を適用するかなど、周囲の選択に起因していました。 BioSecBench-Surveillance は、次のアウトブレイク発生時にエージェントがゲノム監視を信頼できるかどうかを測定するための基準を提供します。

原文 (English)

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962 gradable attempts from sixteen model-harness pairs, the strongest configuration cleared only about half. Opus 4.8 with PI led at 50.2 percent, with a 95 percent confidence interval of 40.1 to 60.3 percent across 83 evaluations, tied with GPT-5.5 with Codex at 50.2 percent, with a 95 percent confidence interval of 40.8 to 59.6 percent, followed by Opus 4.7 with PI at 49.6 percent, with a 95 percent confidence interval of 40.0 to 59.2 percent, and Sonnet 4.6 with PI at 48.6 percent, with a 95 percent confidence interval of 38.9 to 58.3 percent. Even when agents invoked the correct workflows, their mistakes came from the choices around them, such as which references, thresholds, filters, and normalization to apply. BioSecBench-Surveillance provides a standard for measuring whether agents can be trusted to perform genomic surveillance when the next outbreak arrives.

13:00 JSTエージェント研究/論文

LangGraph を使用したグラフベースのエージェント AI: 長期実行ステートフル ビジネス プロセスのワークフロー パスウェイ

このペーパーは、ビジネス プロセスにおける長期実行、ステートフル、マルチステップの生成 AI システムのためのグラフベースのワークフロー パスの実践ガイドです。ステートフル エージェント用の低レベル オーケストレーション フレームワークである LangGraph をモデル品質のベンチマーク ターゲットとして扱うのではなく、修復ループを使用した SQL 分析、証拠ゲートを使用したエージェント検索拡張生成、および割り込みとチェックポイント回復を使用した人間参加型ポリシー レビューの 3 つの実行可能なレシピを提示して、型付き状態、条件付きルーティング、決定論的ツール、再試行、割り込み、チェックポイント、およびトレースがどのように組み合わされるかを示します。 LangGraph は、普遍的なデフォルトとしてではなく、ワークフローの複雑さの適合性によって位置付けられています。基本的なツールの使用、構造化された抽出と検証のためのスキーマ優先ツール、およびプロンプトまたはプログラムの最適化が主な目標である場合の DSPy には、よりシンプルな ReAct スタイルまたはプレーンな SDK ループの方が適している可能性があります。各レシピでは、LangGraph に追加の構造を追加する価値がある場合と、ルート、一時停止、および監査証跡を隠れたプロンプト ロジックではなく明示的な製品動作にする実装パターンについて説明します。

原文 (English)

Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes

This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in business processes. Rather than treating LangGraph, a low-level orchestration framework for stateful agents, as a model-quality benchmark target, we present three executable recipes -- SQL analytics with repair loops, agentic retrieval-augmented generation with evidence gating, and human-in-the-loop policy review with interrupt and checkpoint recovery -- to show how typed state, conditional routing, deterministic tools, retries, interrupts, checkpoints, and traces fit together. LangGraph is positioned by workflow-complexity fit, not as a universal default: simpler ReAct-style or plain SDK loops may be better for basic tool use, schema-first tools for structured extraction and validation, and DSPy when prompt or program optimization is the main goal. Each recipe explains when LangGraph is worth the extra structure and which implementation patterns make routes, pauses, and audit trails explicit product behavior rather than hidden prompt logic.

13:00 JSTLLM/生成AI

介入としての LLM 検出: 戦略的なユーザー行動の下での下流への影響

LLM の採用がさらに普及するにつれて、LLM 検出ツールや言語パターンに基づくヒューリスティックなど、LLM で生成されたコンテンツの検出に対する関心が高まっています。検出器は、検出された属性自体だけでなく、LLM の使用状況や出力品質などの下流のメトリクスも制御する介入として機能します。この研究では、不完全な LLM 検出器が、ユーザーがワークフローで LLM を使用する動機をどのように歪めるかによって、これらの下流メトリクスに直感に反する影響をもたらすことを示します。私たちは、ユーザーが LLM の使用量を戦略的に選択する方法と、検出された属性を減らすためにコンテンツを後処理する方法をキャプチャする様式化されたモデルを開発します。このモデルを使用して、LLM 検出が直感に反して人間の LLM 使用量の増加につながる可能性があることを示します。さらに、検出された属性を減らすと出力品質が向上する場合でも、LLM 検出器を導入すると、ユーザーが低品質の出力を生成する可能性があることがわかりました。対照的に、我々は、検出器が検出された属性のきれいな「上昇後下降」パターンをもたらすことを示し、これを arXiv 要約の単語頻度について経験的に再現します。まとめると、私たちの研究は、LLM 検出がどのように LLM の使用法と出力品質を歪め、LLM 検出器がこれらの下流メトリクスへの介入として動作するときに障害モードを明らかにするかを示しています。

原文 (English)

LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior

As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how users strategically choose how much to use the LLM and how to post-process content to reduce the detected attribute. Using this model, we show that LLM detection can counterintuitively lead humans to increase their LLM usage. Moreover, even when reducing the detected attribute improves output quality, we find that introducing an LLM detector can lead users to produce lower quality outputs. In contrast, we show that detectors result in a clean "rise-then-fall" pattern for the detected attribute, which we empirically reproduce for word frequencies on arXiv abstracts. Altogether, our work illustrates how LLM detection can distort LLM usage and output quality, uncovering failure modes when LLM detectors operate as an intervention on these downstream metrics.

13:00 JSTエージェントハードウェア/半導体研究/論文

ResearchArena: 自動化された AI 研究開発における妨害行為と監視の評価

AI エージェントが AI 研究開発を自動化し始めると、エージェント自体が信頼できない場合でも、その出力が安全に展開できるかどうかを評価する方法が必要になります。 AI 制御は、そのようなアプローチの 1 つを提供します。エージェントを信頼するのではなく、潜在的な敵として扱い、モニターを使用して展開前に秘密の妨害行為を検出します。私たちは、トレーニング後の安全性、トレーニング後の機能、CUDA カーネルの最適化、推論サーバーの最適化という 4 つの長期的なタスクにまたがるフレームワークである ResearchArena を使用して、自動化された AI 研究開発のための AI 制御を評価します。 AI R&D の成果物はデプロイされて実行されるアーティファクトであるため、各メイン タスクを 2 種類の隠れたサイド タスクと組み合わせます。1 つは送信されたモデル、アダプター、カーネル、またはサーバー自体を妨害する埋め込みサイド タスクで、もう 1 つはアーティファクトを正直なままにしてサンドボックス内で秘密裏に範囲外のアクションを実行する独立したサイド タスクです。私たちは、妨害行為と監視の両方でフロンティア エージェントを、2 つの軸に沿って変化する 4 つの監視タイプにわたって評価します。監視者が生成された成果物を実行および調査できるかどうか、およびエージェントの思考連鎖を確認できるかどうかです。トレーニング データに隠された妨害行為は発見するのが最も難しく、フラグが立てられるのは半分未満であることがわかりました。軌道を読み取るだけでなく、モニターにアーティファクトの実験を実行させることは役立ちますが、それだけでは十分ではありません。モニターは、表面だけを検査したり、異常に気づいたがそれを説明したり、間違ったテストでアーティファクトを調べたりすることによって、埋め込まれた妨害行為を見逃してしまいます。私たちは、自動化された AI 研究開発における妨害行為と制御を評価するためのモジュール式フレームワークとして ResearchArena をリリースします。

原文 (English)

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapter, kernel, or server itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact honest. We evaluate frontier agents at both sabotage and monitoring, across four monitor types that vary along two axes: whether the monitor may execute and probe the produced artifact, and whether it sees the agent's chain-of-thought. We find that sabotage hidden in the training data is the hardest to catch, flagged fewer than half the time. Letting the monitor run experiments on the artifact, rather than only read the trajectory, helps, but it is not enough: monitors still miss embedded sabotage by inspecting only the surface, by noticing the anomaly but explaining it away, or by probing the artifact with the wrong test. We release ResearchArena as a modular framework for evaluating sabotage and control in automated AI R&D.

13:00 JST研究/論文

畳み込みニューラル ネットワークにおける連合感情学習

連想感情学習により、生物は楽しい結果または不快な結果を予測刺激の存在と適応的に結び付けることができます。 Rescorla-Wagner モデルなどの計算モデルはこの重要な機能に光を当てていますが、特にニューラル データに適用する場合、これらのモデルの限界も知られています。ディープ ニューラル ネットワークの出現により、連想感情学習をモデル化するための別の道が開かれました。この研究では、複雑な自然シーンをエンコードする視覚モジュールと、感情の重要な側面である価数の観点からそれらの感情的重要性を認識するモジュールで構成される視覚価価処理のディープ ニューラル ネットワーク モデルを提案し、このモデルで新しいパブロフ学習パラダイムをテストしました。その結果、学習により、モデルが人間の連合学習研究からのいくつかの観察(連合形成や一般化など)を再現し、条件付き刺激と無条件刺激の神経表現が単一ユニットと神経集団レベルの両方でますます一致することが示されました。モデルと人体実験データの比較により、私たちのアプローチがさらに検証されました。したがって、この研究は、ディープ ニューラル ネットワーク モデルを適切な学習アルゴリズムと組み合わせると、連想感情/価性学習の行動および神経シグネチャをモデル化するために使用できることを示唆しています。

原文 (English)

Associative Emotional Learning in Convolutional Neural Networks

Associative emotional learning enables organisms to adaptively link pleasant or unpleasant outcomes to the presence of predictive stimuli. Whereas computational models such as the Rescorla-Wagner model have shed light on this important function, the limitations of these models are also known, especially when they are applied to neural data. The advent of deep neural networks has opened another avenue for modeling associative emotional learning. In this work we proposed a deep neural network model of visual valence processing, consisting of a visual module that encodes complex natural scenes and a module that recognizes their emotional significance in terms of valence, a key dimension of emotion, and tested a novel Pavlovian learning paradigm on the model. The results showed that with learning, the model reproduced several observations from human associative learning studies, including association formation and generalization, and that the neural representations of the conditioned and the unconditioned stimuli became increasingly aligned both at the single unit and at the neural population level. Comparison between the model and human experimental data provided further validation of our approach. This study thus suggests that deep neural network models, when combined with appropriate learning algorithms, can be used to model behavioral and neural signatures of associative emotion/valence learning.

13:00 JSTLLM/生成AIエージェント研究/論文

野生のエージェント: 研究と展開が出会う場所

推論、計画、実行、およびツールや他のエージェントとの調整が可能なエージェント システムの大規模言語モデル (LLM) ベースのアーキテクチャは、ソフトウェア エンジニアリング、科学的発見、金融などの分野にわたって、研究プロトタイプから実稼働規模の展開へと急速に移行しています。学術研究ではベンチマークとアルゴリズムの革新が重視されてきましたが、導入では堅牢性、安全性、信頼性に関して新たな課題が生じます。このチュートリアルでは、研究者と実践者が集まり、推論と計画、マルチエージェントの調整、および評価の進歩を探り、導入経験から生じる未解決の課題に焦点を当てます。創薬システムと金融システムにおける応用ケーススタディを通じて、エージェントシステムを成功させる一般的な設計パターンを分析し、検証パイプライン、フォールバックメカニズム、人間による監視などの障害モードの実践的な軽減戦略について議論します。参加者は、業界全体で安全かつ信頼性の高い導入を実現するための具体的な設計パターン、評価チェックリスト、テンプレートとともに、この分野の包括的な見解を得ることができます。

原文 (English)

Agents in the Wild: Where Research Meets Deployment

Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes to production scale deployments across domains such as software engineering, scientific discovery, and finance. While academic work has emphasized benchmarks and algorithmic innovation, deployment raises new challenges around robustness, safety, and reliability. This tutorial brings together researchers and practitioners to explore advances in reasoning and planning, multi agent coordination, and evaluation, highlighting open challenges arising from deployment experience. Through applied case studies in pharmaceutical discovery and financial systems, we analyze common design patterns that make agentic systems successful, and discuss practical mitigation strategies for failure modes, such as verification pipelines, fallback mechanisms, and human in the loop supervision. Attendees will gain a comprehensive view of the field along with concrete design patterns, evaluation checklists, and templates for safe and reliable deployment across industries.

13:00 JSTエージェントGPT / ChatGPT

CodeRescue: コーディング エージェント向けの予算調整されたリカバリ ルーティング

コーディング エージェントは、失敗した試行が単なる不正解ではなく実用的なフィードバックを生成する実行可能環境で動作することが増えています。既存のコストを意識したシステムは通常、このような障害をカスケード決定として扱います。つまり、最初に安価なモデルを試し、その後、困難なケースをより強力でより高価なモデルにエスカレーションします。ただし、コーディングでは、実行フィードバックによって安価なモデルの回復がさらに価値のあるものになる可能性もあり、エージェントはいつより安価なコンピューティングを費やす必要があるのか​​、いつエスカレーションすべきなのかという予算計画上の導入の問題が生じます。この障害後の決定を異種アクションに対する回復ルーティングとして定式化し、実行ロールアウトから監視対象ルーターをトレーニングします。変化する予算の下でも同じルータを使用できるようにするために、再トレーニングなしで導入時のコストペナルティを選択し、交換可能性の下で限界予想コスト制御を提供するコンフォーマルリスクコントロール(CRC)レイヤーを追加します。 5 つのコーディング ベンチマークで継続的に失敗した場合、安価なリカバリとエスカレーションは相補的な成功パターンを示します。調整されたフロンティアは、固定アクション、プロンプト専用ルーター、バイナリ カスケード ベースラインよりも改善されています。メインの GPT-5.4-nano/GPT-5.4 設定では、平均回復コストの 35% を使用しながら、1 つの CRC 校正済みフロンティア ポイントが常時エスカレートの解決速度を超えています。コードは https://github.com/Qijia-He/agent-budget-control で入手できます。

原文 (English)

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap compute, and when should it escalate? We formulate this post-failure decision as recovery routing over heterogeneous actions and train a supervised router from execution rollouts. To make the same router usable under changing budgets, we add a Conformal Risk Control (CRC) layer that selects a deployment-time cost penalty without retraining and provides marginal expected-cost control under exchangeability. Across held-out failures from five coding benchmarks, cheap recovery and escalation exhibit complementary success patterns. The calibrated frontier improves over fixed actions, prompt-only routers, and a binary cascade baseline; in the main GPT-5.4-nano/GPT-5.4 setting, one CRC-calibrated frontier point exceeds always-escalate solve rate while using 35% of its mean recovery cost. Code is available at https://github.com/Qijia-He/agent-budget-control.

13:00 JSTLLM/生成AIエージェント

MechAInistic: ゲノムスケールの制約ベースの代謝モデルを推論するための LLM ガイド付きマルチエージェント システム

制約ベースの代謝モデリングは、細胞の状態や疾患のメカニズムの基礎を研究するための強力な方法ですが、その効果的な使用には、相当な計算の専門知識と複数段階の分析の慎重な調整が必要です。私たちは、この障壁を低くし、研究者が自然言語で複雑な生物学的質問をできるようにするために MechAInistic を開発しました。 MechAInistic は、大規模な言語モデルを利用して、自然言語の質問を実行可能なモデルベースのワークフローに変換し、構造化されたレポートを生成する、Architect-Reviewer パターンを中心に編成されたマルチエージェント システムです。このシステムは、経路の比較、摂動解析、薬物標的の探索、ペアの代謝モデル状態にわたる文献に基づいた解釈など、さまざまなタスクをサポートします。私たちは、治療仮説生成のための 2 つのペアの免疫細胞代謝モデルのユースケースを使用して MechAInistic を開発し、評価しました。 MechAInistic は、関節リウマチ (RA) のナイーブ B 細胞を健康な対照と組み合わせて、ミトコンドリア代謝再配線を特定し、デビミスタット/CPI-613 を OGDH 中心の研究仮説として指名しました。多発性硬化症 (MS) と健常対照からの CD4+ Th17 細胞のペア研究において、同じワークフローにより NADP 依存性イソクエン酸デヒドロゲナーゼが最適な単一標的として特定され、イボシデニブが FDA 承認の再利用候補として提案されました。これらの結果を総合すると、MechAInistic が自然言語の生物学的質問を、追跡可能な治療仮説生成のための実行可能なモデルに基づいたワークフローに変換することを示しています。

原文 (English)

MechAInistic: An LLM-guided Multi-Agent System for Reasoning over Genome-Scale Constraint-Based Metabolic Models

Constraint-based metabolic modeling is a powerful way to study the mechanistic basis of cellular states and disease, but its effective use demands substantial computational expertise and careful coordination of multi-step analyses. We developed MechAInistic to lower this barrier and enable researchers to ask complex biological questions in natural language. Harnessing large language models, MechAInistic is a multi-agent system organized around an Architect-Reviewer pattern that transforms a natural-language question into an executable, model-grounded workflow and generates a structured report. The system supports a variety of tasks, including pathway comparison, perturbation analysis, drug-target exploration, and literature-grounded interpretation across paired metabolic model states. We developed and evaluated MechAInistic using two paired immune-cell metabolic-model use cases for therapeutic hypothesis generation. For Naive B cells from rheumatoid arthritis (RA) paired with healthy controls, MechAInistic identified mitochondrial metabolic rewiring and nominated Devimistat/CPI-613 as an investigational OGDH-centered hypothesis. In a paired CD4+ Th17 cell study from multiple sclerosis (MS) and healthy controls, the same workflow identified NADP-dependent isocitrate dehydrogenase as the optimal single target and proposed ivosidenib as an FDA-approved repurposing candidate. Together, these results show that MechAInistic converts natural-language biological questions into executable, model-grounded workflows for traceable therapeutic hypothesis generation.

13:00 JSTエージェント

アシスタントか俳優か?汎用 AI エージェントを使用した場合の学生の信頼、制御、および委任に関する後悔

AI エージェントが質問に答えることから行動を起こすことに移行すると、ユーザーは新たな問題に直面します。それは、アクション スペースが完全に予測できないシステムに何を委任するかを決定することです。私たちは、その結果として生じる不満の委任の後悔を、エージェントが間違いを犯したことではなく、エージェントが許可した範囲を超えて行動したことをユーザーが後悔するパターンと呼んでいます。対照研究では、20 人の大学生が、汎用 AI エージェントである OpenClaw を使用して、プライバシー、賭け金、可逆性の点で異なるタスクを選択して、5 つの共通の日常タスクを完了しました。各タスクについて、信頼性、認識されたコントロール、透明性、監督の負担、および承認の好みを 5 段階のリッカート尺度で測定し、テーマ別コーディングを通じて分析されたフリーテキストの反映を収集しました。 3 つの発見が得られました。まず、参加者はエージェントごとではなくタスクごとに信頼を調整しました。参加者は、助言やリスクの低いタスクについては幅広い自律性を認めましたが、取り消し不能で外部から見えるアクションについては確認を要求しました。第 2 に、賭け金だけではなく、外部からの可視性と組み合わされた不可逆性が信頼の撤退を促進しているように見えました。賭け金が中程度の電子メール タスクは信頼の急激な低下 (M = 3.10) と最も高い承認要求 (M = 4.65) を引き起こしましたが、賭け金は高いが検証可能なタスクでは同じ反応が得られませんでした。 3 番目に、出力が成功と評価された場合でも、エージェントがプレビューなしでアクションを実行すると、委任の後悔が一貫して現れました。アクションの境界を明らかにし、タスクごとの自律性ポリシーをサポートし、アドバイス出力をエージェントの実行から分離するエージェント設計への影響について説明します。

原文 (English)

Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent

When AI agents shift from answering questions to taking actions, users face a new problem: deciding what to delegate, to a system whose action space they cannot fully anticipate. We call the resulting dissatisfaction delegation regret, a pattern in which users regret not that the agent erred, but that it acted beyond what they would have authorized. In a controlled study, 20 university students completed five common daily tasks using OpenClaw, a general-purpose AI agent, across tasks chosen to vary in privacy, stakes, and reversibility. For each task we measured trust, perceived control, transparency, supervision burden, and approval preference on 5-point Likert scales, and collected free-text reflections analyzed through thematic coding. Three findings emerged. First, participants calibrated trust per task rather than per agent: they granted wide autonomy for advisory and low-stakes tasks but demanded confirmation for irreversible, externally visible actions. Second, irreversibility combined with external visibility, rather than stakes alone, appeared to drive trust withdrawal: the moderate-stakes email task triggered the sharpest drop in trust (M = 3.10) and the highest demand for approval (M = 4.65), whereas a high-stakes but verifiable task did not produce the same response. Third, delegation regret appeared consistently when the agent executed actions without preview, even when the output was rated as successful. We discuss implications for agent designs that expose action boundaries, support per-task autonomy policies, and separate advisory output from agentic execution.

13:00 JST研究/論文

自律性の経済学: 保険可能な AI 駆動の 6G システム向けのリアルタイム リスク インデックス作成

第 6 世代 (6G) ネットワークへの移行により、ワイヤレス インフラストラクチャは、Vehicle-to-Everything (V2X)、Industrial IoT (IIoT)、および Integrated Sensing and Communication (ISAC) をサポートする認知基盤に変わります。このパラダイムでは、自律型エージェント AI がミリ秒スケールでオーケストレーションを実行するため、従来の静的なガバナンス フレームワークがリスク管理には根本的に不十分になります。この文書では、エージェントティック 6G システムにおけるリアルタイムのリスク定量化と信頼調整のための Governance-as-Code (GaC) フレームワークである GIRAF (Governance-Integrated Risk and Assurance Framework) を紹介します。 GIRAF は、認識信頼度、ネットワーク ジッター、検証レイテンシーなどの機械可読ランタイム信号から連続的な総合リスク指数 ($R_{t}$) を導き出します。中心的な貢献は、計算の遅延が 6G の期限を超えた場合に安全メカニズムがリスクを引き起こす、検証の古さのトレードオフの形式化です。我々は、GIRAFがエージェントが報告した確実性と環境上の真実との間の「信頼ギャップ」の不一致を特定し、状況が悪化したときに自動化された安全エンベロープをトリガーすることを実証します。重要なのは、GIRAF が基本的なガバナンスの基盤および概念的な「接着剤」として機能し、これらの技術的リスクを機械読み取り可能なテレメトリに外部化することです。微調整されたラージ言語モデル (LLM) を使用したシミュレーションを通じて、このフレームワークが運用上の完全性を維持しながら、6G エコシステムにおけるマルチステークホルダーの責任帰属と動的な保険料定量化に必要な重要な保険数理ベースラインを提供していることを検証します。

原文 (English)

The Economics of Autonomy: Real-Time Risk Indexing for Insurable AI-Driven 6G Systems

The transition to sixth-generation (6G) networks transforms wireless infrastructure into a cognitive substrate supporting Vehicle-to-Everything (V2X), Industrial IoT (IIoT), and Integrated Sensing and Communication (ISAC). In this paradigm, autonomous agentic AI performs orchestration at millisecond scales, rendering traditional static governance frameworks fundamentally inadequate for risk management. This paper introduces GIRAF(Governance-Integrated Risk and Assurance Framework), a Governance-as-Code (GaC) framework for real-time risk quantification and trust modulation in agentic 6G systems. GIRAF derives a continuous Aggregate Risk Index ($R_{t}$) from machine-readable runtime signals, including epistemic confidence, network jitter, and verification latency. A core contribution is the formalization of the verification staleness trade-off, where safety mechanisms induce risk if computational latency exceeds 6G deadlines. We demonstrate that GIRAF identifies 'Confidence Gaps' discrepancies between agent reported certainty and environmental ground truth, triggering automated safety envelopes when conditions deteriorate. Crucially, GIRAF serves as the foundational governance groundwork and conceptual 'glue' that externalizes these technical risks into machine-readable telemetry. Through simulations with fine-tuned Large Language Models (LLMs), we validate that the framework preserves operational integrity while providing the essential actuarial baseline required for multi-stakeholder liability attribution and dynamic premium quantification in the 6G ecosystem.

13:00 JSTエージェントビジネス/資金調達

地域電力市場における需要家向けの市場戦略評価

分散型発電と柔軟な負荷を備えた需要家は、地域資源を制御する自律的なサイバー物理エネルギー システムを形成し、人間の介入を最小限に抑えて地域のエネルギー市場に参加します。この研究では、エージェント ベースのシミュレーション プラットフォームを開発および評価します。このプラットフォームでは、太陽光発電システム、蓄電池システム、電気自動車、ヒート ポンプを備えた消費者世帯を代表するエージェントが、均一価格の両面コール オークションに参加します。個々の入札戦略がコミュニティレベルの効率性やプロシューマーレベルの財務成果に及ぼす影響は、特に異種のポートフォリオを持つプロシューマーが 1 つの市場で相互作用する場合には完全には理解されていません。複雑さが増す 4 つの市場戦略、つまりゼロインテリジェンス制約ベースライン、境界価格戦略、拡張ストレージ カスケード、市場適応型価格戦略を比較します。このシミュレーションは、季節変動を特徴付けるために、夏、冬、春にわたる 15 分の解像度で 33 人のプロシューマーのコミュニティに対して実行されます。結果は、ルールベースのリソース制御により地域社会のエネルギー支出が大幅に削減されることを示しています。拡張された貯蔵カスケードでは総コストが 39.06 ユーロに達し、ゼロインテリジェンスベースラインの場合は 62.38 ユーロとなり、37.4 % 削減されました。市場適応戦略は、夏の条件下で、地元のエネルギー市場への参加を通じて、地域全体で最高の経済的利益をもたらします (ベースラインの 14.40 ユーロ対 10.28 ユーロ、40.1 % の利益)。戦略の有効性はポートフォリオの構成と季節的な供給条件の両方に依存するため、リソース管理と価格決定の共同評価が必要です。

原文 (English)

Market Strategy Evaluation for Prosumers in Local Electricity Markets

Prosumers equipped with distributed generation and flexible loads form autonomous cyber-physical energy systems that control local resources and participate in local energy markets with minimal human intervention. This work develops and evaluates an agent-based simulation platform in which agents, representing prosumer households with photovoltaic systems, battery storage systems, electric vehicles, and heat pumps, participate in a uniform-price double-sided call auction. The effect of individual bidding strategies on community-level efficiency and prosumer-level financial outcomes is incompletely understood, particularly when prosumers with heterogeneous portfolios interact in one market. Four market strategies of increasing complexity are compared: a zero-intelligence constrained baseline, a boundary-price strategy, an extended storage cascade, and a market-adaptive pricing strategy. The simulation is conducted on a community of 33 prosumers at 15-minute resolution, spanning summer, winter, and spring to characterize seasonal variation. Results show that rule-based resource control substantially reduces community energy expenditure: the extended storage cascade achieves a total cost of 39.06 EUR compared to 62.38 EUR under the zero-intelligence baseline, a reduction of 37.4 %. The market-adaptive strategy yields the highest aggregate community financial gain through local energy market participation (14.40 EUR vs. 10.28 EUR for the baseline, a gain of 40.1 %) under summer conditions. Strategy effectiveness depends on both portfolio composition and seasonal supply conditions, requiring joint evaluation of resource control and pricing decisions.

13:00 JST研究/論文

警官と強盗問題のドメイン設計

警官と強盗は、グラフ理論でよく研究されている問題です。この設定は、無向グラフ上に配置された強盗と 1 人以上の警官で構成されます。警官は交代でグラフ内を移動し、強盗を捕まえようとします。興味深い特性は、$k$ 警官が初期配置の構成を考慮して、有限ターン後に少なくとも 1 人の警官が強盗と同じ頂点を占めることを保証するのに十分であるかどうかです。成功した場合、グラフは「$k$-copwin」と呼ばれます。この研究では、グラフが $k$-copwin であるかどうかを決定する問題を非決定論的な計画問題として投影し、最先端のプランナーを使用してこの特性を計算します。警官の動きは(考えられるすべての戦略を捕捉するための)非決定論的な動きとしてキャストされますが、強盗の動きは本質的に決定論的です。また、グラフ理論の文献からのいくつかのバリエーションを使用して、基本モデルを拡張します。

原文 (English)

Domain Design for the Cops and Robbers Problem

Cops and Robbers is a well-studied problem in graph theory. The setting consists of a robber and one or more cops placed on an undirected graph. Taking turns moving throughout the graph, the cops try to capture the robber. The property of interest is whether $k$ cops suffice to ensure at least one cop occupies the same vertex as the robber, after a finite number of turns, given any configuration of their initial placement; if successful, the graph is referred to as ``$k$-copwin''. In this work, we cast the problem of determining whether a graph is $k$-copwin as a non-deterministic planning problem and use state-of-the-art planners to compute this property. The cop movement is cast as non-deterministic movement (to capture all possible strategies), while the robber movement is deterministic in nature. We also extend the base model using several variations from the graph theory literature.

13:00 JST研究/論文

識別力の計算: 高次の学習としての意思決定関連の洞察、順序値、および忘却

生成型 AI の世界では、候補者の洞察が豊富にあります。不足しているのは、どの問題を識別し、適切な量と順序でそれらに対処し、システムが適応できるように残りのことを忘れる能力です。私たちは、これらの希少性は 1 つのオブジェクトによって支配されていると主張し、それを中心にフレームワークを構築します。私たちは洞察を、目的に対する特定された測定可能な効果を持つレバーとして厳密に定義し、新規性ではなく情報の期待値による意思決定の関連性によって候補をランク付けします。私たちは、アクションにはサイズだけでなく順序も伴うことを示します。現実的な信念力学の下では、コンテンツの「タッチ」は非交換演算子であるため、異なる順序で提供される固定プランは異なる結果をもたらし、シーケンスプレミアムを定義します。私たちは、レバレッジの価値は影の価格であり、医薬品のマーケティング、株式の選択、製造を 1 つのレバレッジ発見の問題として統合していることを観察しています。最も推測的に、我々は、忘却が知識の処分ではなく、価値が学習される演算子であるという理論であるAPOHAを提案する。保持されるアイテムの価値は、それを忘れることによる反事実的なコストであり、学習システムは、保存された価値の対象となる最大の忘却の残余であり、より高次の価値は、統合を共役として、繰り返される忘却に耐える構造(繰り込み関連の不変量)である。私たちは中心未解決問題 (スペクトル ギャップを持つ非自明なアトラクター) を述べ、忘却理論をテストします。APOHA を 30 シード以上の非定常肥満治療決定ワールドのエージェントとして運用し、適応型忘却は決して忘れないことと固定半減期に対して累積決定後悔を 24 ~ 32% 削減し、約 6 分の 1 より小さくクリーンな記憶を維持し、安定して収束します。特に、盲目的に忘れることは決して忘れないことよりも悪かったため、価値を認識して忘れることに特有の利点があります。学際的な批評が全体をストレステストします。

原文 (English)

A Calculus of Discernment: Decision-Relevant Insight, Sequence Value, and Forgetting as Higher-Order Learning

In a world of generative AI, candidate insights are abundant; what is scarce is the capacity to discern which matter, to act on them in the right amount and order, and to forget the rest so the system can adapt. We argue these scarcities are governed by one object and build a framework around it. We define an insight strictly as a lever with an identified, measurable effect on an objective, and rank candidates by decision-relevance via the expected value of information rather than novelty. We show action carries an order, not only a size: under realistic belief dynamics, content "touches" are non-commuting operators, so a fixed plan delivered in different orders yields different outcomes, defining a sequence premium. We observe that the value of any lever is a shadow price, unifying pharmaceutical marketing, equity selection, and manufacturing as one leverage-discovery problem. Most speculatively, we propose APOHA, a theory in which forgetting is not the disposal of knowledge but the operator by which value is learned: the value of a retained item is the counterfactual cost of forgetting it, a learning system is the residue of maximal forgetting subject to preserved value, and higher-order value is the structure that survives repeated forgetting (a renormalisation-relevant invariant), with consolidation as its conjugate. We state the central open problem (a non-trivial attractor with a spectral gap) and test the forgetting theory: operationalising APOHA as an agent on a non-stationary obesity-treatment decision world over 30 seeds, adaptive forgetting cut cumulative decision-regret by 24-32% against never-forget and a fixed half-life, kept a ~6x smaller, cleaner memory, and converged stably; notably, blind forgetting was worse than never forgetting, so the benefit is specific to value-aware forgetting. A multi-disciplinary critique stress-tests the whole.

13:00 JST研究/論文

FALCON-Discover: キャリブレーションのための集中した誤信頼領域の検出

通常、キャリブレーションは全体的に評価されますが、最も危険な失敗は多くの場合局所的なものであり、予測は間違っているにもかかわらず高い信頼性を維持します。私たちは、この故障モードを誤った信頼度の集中、つまり、信頼できるエラーが予測空間のコンパクトで発見可能な領域を占める程度として研究します。 FALCON-Discover は、信頼性、ローカル サポート、近隣合意、摂動安定性からの不一致シグナルを使用して予測をランク付けする、ポストホックでモデルに依存しないフレームワークです。 7 つのバイナリ表形式データセット、4 つのシード、5 分割交差フィッティング、および XGBoost と CatBoost を含む強力な学習器全体にわたって、誤った信頼度の集中は反復的であるがレジームに依存していることがわかりました。主要な信頼しきい値では、不一致に基づくランキングは、最も強力な検証で選択された校正または最も強力な体制における信頼スコアリング ベースラインを大幅に上回りますが、生の信頼度では危険なエラーの量はほとんど回復しません。最適な検出器はデータセットによって異なります。学習された不一致は、複数の手がかりを組み合わせる必要がある場合に最も強くなりますが、安定性を中心としたランキングは、局所的な決定の脆弱性が支配的な場合に最も効果的に機能します。これらの結果は、危険な過信は単一スコアの調整問題としてよりも家族レベルの発見問題として扱う方が適切であり、信頼、支持、安定性が異なる領域を明確にターゲットとする調整戦略を動機付けることを示しています。

原文 (English)

FALCON-Discover: Discovering Concentrated False-Confidence Regions for Calibration

Calibration is usually evaluated in aggregate, but the most dangerous failures are often local: predictions that remain highly confident despite being wrong. We study this failure mode as false-confidence concentration, the extent to which confident errors occupy compact, discoverable regions of prediction space. We introduce FALCON-Discover, a post-hoc, model-agnostic framework that ranks predictions using discrepancy signals from confidence, local support, neighborhood agreement, and perturbation stability. Across seven binary tabular datasets, four seeds, five-fold cross-fitting, and strong learners including XGBoost and CatBoost, we find that false-confidence concentration is recurrent but regime-dependent. At the main confidence threshold, discrepancy-based ranking substantially outperforms the strongest validation-selected calibration or trust-scoring baseline in the strongest regimes, while raw confidence recovers little dangerous-error mass. The best detector varies across datasets: learned discrepancy is strongest when multiple cues must be combined, whereas stability-centered ranking works best when local decisional fragility dominates. These results show that dangerous overconfidence is better treated as a family-level discovery problem than as a single-score calibration problem, and motivate calibration strategies that explicitly target regions where confidence, support, and stability diverge.

13:00 JSTハードウェア/半導体

出力空間キャリブレーションを超えて: 時系列分類における選択的信頼性推定のためのスペクトル証拠バンドリング

時系列分類のための事後キャリブレーションでは通常、出力スコアが再マップされますが、信頼、棄権、レビューなどの展開の決定は、信頼できる予測が現在の時間信号によってサポートされているかどうかによって決まります。私たちは 3 つの時系列信頼性ギャップに対処します。同一の信頼値は異なる時間的サポートを隠す可能性があり、平均キャリブレーションは偽の高信頼エラーを見逃す可能性があり、出力空間の再キャリブレーションでは入力にリンクされた監査可能性が制限されます。バックボーン予測を変更せずに維持し、信頼すべきかどうかを推定する検証ゲート付き固定ラベル信頼性ポリシーを導入します。この方法では、出力側のキューと、バンド エネルギー、エントロピー、ピーク ドミナンス、周期サポート、位相安定性などのサンプル全体のスペクトル記述子を組み合わせて、スカラー信頼性推定と診断バンド レベルの証拠を形成します。検証ゲートは、FalseConf@0.9 または AURC 許容値に違反せずに正確性ランキングが向上した場合にのみスペクトル調整を有効にします。それ以外の場合は、より安全な出力空間のベースラインに戻ります。 8 つの異種 UCR/UEA データセット、8 つの時系列バックボーン ファミリ、および標準再キャリブレーターにわたって、制約なしの方法により、一致する評価サブセットの固定ラベル選択信頼性メトリクスが向上し、Corr-AURC が 0.693 から 0.779 に上昇しました。検証ゲート型ポリシーにより、Corr-AURC が 0.786 にさらに改善され、FalseConf@0.9 が 0.094 に減少します。これらの結果は、時系列分類器の信頼性推定は、出力の信頼性をスペクトル証拠とバンドルすることで恩恵を受ける一方、検証ゲーティングはサポートされていないスペクトル調整を防止することを示唆しています。

原文 (English)

Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification

Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and review depend on whether a confident prediction is supported by the current temporal signal. We address three time-series reliability gaps: identical confidence values can hide different temporal support, average calibration can miss false high-confidence errors, and output-space recalibration offers limited input-linked auditability. We introduce a validation-gated fixed-label reliability policy that keeps the backbone prediction unchanged while estimating whether it should be trusted. The method combines output-side cues with whole-sample spectral descriptors, including band energy, entropy, peak dominance, period support, and phase stability, to form a scalar reliability estimate and diagnostic band-level evidence. A validation gate enables spectral conditioning only when correctness ranking improves without breaching FalseConf@0.9 or AURC tolerances; otherwise it reverts to the safer output-space baseline. Across eight heterogeneous UCR/UEA datasets, eight time-series backbone families, and standard recalibrators, the unconstrained method improves fixed-label selective-reliability metrics on the matched evaluation subset, raising Corr-AURC from 0.693 to 0.779. The validation-gated policy further improves Corr-AURC to 0.786 and reduces FalseConf@0.9 to 0.094. These results suggest that reliability estimation for time-series classifiers benefits from bundling output confidence with spectral evidence, while validation gating prevents unsupported spectral conditioning.

13:00 JSTLLM/生成AI

単次元圧縮を超えて: 大規模言語モデルの複合スパース性フロンティア

大規模言語モデル (LLM) は、多くの場合、静的パラメーター プルーニングや動的なトークン レベルの計算を通じて圧縮されますが、積極的なスパース化により、本質的なスパース境界を超えて急激なパフォーマンスの低下が引き起こされる可能性があります。この研究は、\emph{これら 2 つのメカニズムを組み合わせることで、圧縮負担を分散することでそのような劣化を遅らせることができるかどうか}を問うものです。私たちは、最初に低ランク近似とチャネル プルーニングを適用して静的に圧縮されたバックボーンを取得し、次にトークンごとの動的レイヤー スキップのための軽量ルーターを導入する、最小限の複合スパース フレームワークを研究します。この設計により、パラメーターのスパース性とトークンレベルの計算のスパース性を独立して制御できます。言語理解とモデリングのベンチマークにわたる実験では、複合スパース性が同じ合計スパース性の下で単一メカニズムの圧縮よりも常に優れたパフォーマンスを示し、タスクの理解における減衰点を遅らせ、より強力なモデリング パフォーマンスを維持することが示されています。さらなる分析により、パラメータ プルーニングとトークン スキップの間の次元を越えた干渉が明らかになり、固定されたスパース バジェットの下でほぼバランスの取れた割り当てが最も効果的であることが示されています。これらの結果は、複合圧縮が LLM 圧縮を改善する実用的な方法を提供すると同時に、最終的にさらなる圧縮を制限するより広い次元間スパース境界を明らかにすることを示しています。コードは https://github.com/EIT-NLP/LLM-Pruning で入手できます。

原文 (English)

Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models

Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance degradation beyond an essential sparsity boundary. This work asks \emph{whether combining these two mechanisms can delay such degradation by distributing the compression burden}. We study a minimalist compound sparsity framework that first applies low-rank approximation and channel pruning to obtain a statically compressed backbone, and then introduces lightweight routers for per-token dynamic layer skipping. This design enables independent control of parameter sparsity and token-level computation sparsity. Experiments across language understanding and modeling benchmarks show that compound sparsity consistently outperforms single-mechanism compression under the same total sparsity, delaying the decay point on understanding tasks and preserving stronger modeling performance. Further analysis reveals cross-dimensional interference between parameter pruning and token skipping, and shows that near-balanced allocation is most effective under a fixed sparsity budget. These results demonstrate that compound compression provides a practical way to improve LLM compression, while revealing a broader cross-dimensional sparsity boundary that ultimately limits further compression. Code will be available at https://github.com/EIT-NLP/LLM-Pruning.

13:00 JST画像/動画生成

FedCC: 胎児超音波画像における堅牢な脳梁位置特定のための基礎モデルの低リソースフェデレーテッド適応

胎児超音波 (US) 画像における脳梁 (CC) の位置を正確に特定することは、神経発達の異常を早期に特定するために重要です。ただし、低コントラスト、スペックル ノイズ、CC のかなりの解剖学的多様性など、US イメージングの本質的な制限により、この作業は依然として非常に困難です。我々は、胎児 US 画像における CC 位置特定のための連合学習 (FL) ベースのフレームワークである FedCC を提案します。これは、データ共有を必要とせず、現実的な多施設およびリソースに制約のある臨床現場向けに特別に設計されています。このフレームワークは、凍結された DINOv2 バックボーンと軽量の YOLO ベースの検出ヘッドを統合します。パラメータの効率的な適応を可能にするために、低ランク適応 (LoRA) モジュールが組み込まれており、パラメータの小さなサブセットのみを最適化し、クライアント間で交換できます。この戦略により、計算オーバーヘッドと通信オーバーヘッドの両方が大幅に削減され、フレームワークが低リソース環境に適したものになります。提案されたアプローチは、異種イメージング デバイスを使用して 3 つの臨床現場にわたる日常的な神経超音波検査中に 58 人の妊婦から取得した 10,970 個の超音波フレームで構成される多施設データセットで評価されました。提案されたフレームワークは、フェデレーテッド設定で優れたパフォーマンスを達成しました。特に、FedAvg 戦略に基づく DINOv2 と LoRA の組み合わせは、平均 mAP@50 0.857 と F1 スコア 0.803 を達成し、完全な微調整ベースラインとエンコーダー フリーズ ベースラインの両方を上回りました。特に、提案されたアプローチでは、トレーニング可能なパラメータの数が完全な微調整の 2,440 万個と比較して 290 万個に減少し、これは通信コストの約 8.5 倍の削減に相当します。これらの発見は、胎児神経超音波検査用のスケーラブルでプライバシー保護があり、臨床展開可能な AI システムに向けた有望な一歩を示しています。

原文 (English)

FedCC: A Low-Resource Federated Adaptation of Foundation Models for Robust Corpus Callosum localization in Fetal Ultrasound Images

Accurate localization of the corpus callosum (CC) in fetal ultrasound (US) images is crucial for the early identification of neurodevelopmental abnormalities. However, this task remains highly challenging due to the intrinsic limitations of US imaging, including low contrast, speckle noise, and the considerable anatomical variability of the CC. We propose FedCC, a federated learning (FL)-based framework for CC localization in fetal US images, specifically designed for realistic multi-center and resource-constrained clinical settings without requiring data sharing. The framework integrates a frozen DINOv2 backbone with a lightweight YOLO-based detection head. To enable parameter-efficient adaptation, Low-Rank Adaptation (LoRA) modules are incorporated, allowing only a small subset of parameters to be optimized and exchanged among clients. This strategy substantially reduces both computational and communication overhead, making the framework suitable for low-resource environments. The proposed approach was evaluated on a multi-center dataset comprising 10,970 ultrasound frames acquired from 58 pregnant women during routine neurosonographic examinations across three clinical sites using heterogeneous imaging devices. The proposed framework achieved strong performance in the federated setting. In particular, the combination of DINOv2 and LoRA under the FedAvg strategy achieved an average mAP@50 of 0.857 and an F1-score of 0.803, outperforming both full fine-tuning and encoder-freezing baselines. Notably, the proposed approach reduced the number of trainable parameters to 2.9M compared with 24.4M in full fine-tuning, corresponding to an approximately 8.5$\times$ reduction in communication cost. These findings represent a promising step toward scalable, privacy-preserving, and clinically deployable AI systems for fetal neurosonography.

13:00 JST研究/論文

重要なものを圧縮する: ニューロンの重要性が言語モデル圧縮のデータを意識した低ランク近似を満たす

その分野で優れているために、大規模な言語モデルは数十億のパラメータで構成されています。ただし、これには膨大なメモリ要件が必要となり、リソースに制約のある環境での適用が制限されます。ニューラル ネットワーク (NN) 圧縮の問題に対処するために、特異値分解 (SVD) は、分解による行列圧縮の基本コンポーネントとして重要な役割を果たしています。圧縮誤差を最小限に抑え、下流タスクでの圧縮モデルの有効性を最大化するために、これまでの研究では、パラメータの重要性または層ごとの機能等価性の観点から、NN の重み行列の低ランク近似に焦点を当てていました。これまでの研究では、前述の観点を個別に研究していましたが、この研究では、これら 2 つの観点からのアイデアを 1 つの目的に組み合わせたアプローチの有効性を研究しています。これと並行して、圧縮品質に影響を与える重要な側面は、レイヤー全体および NN パラメーターにわたる圧縮率の分布です。以前の研究では、層およびネットワークの重み全体に圧縮率を均一に分散するか、計算コストのかかるヒューリスティック検索に依存することが主に検討されていました。これらとは対照的に、この研究では、動的圧縮率割り当てのための強化された計算効率の高いアルゴリズムを提案します。実験結果は、特に高い圧縮率の下で、以前の最先端技術と同等かそれよりも大幅に優れたパフォーマンスを発揮する、提案されたアプローチの有効性を裏付けています。

原文 (English)

Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression

To excel at their domain large language models are comprised of billions of parameters. Yet this comes at the cost of huge memory requirements restricting their applicability in resource-constrained environments. To address the problem of neural network (NN) compression Singular Value Decomposition (SVD) has played a key role as a fundamental component for matrix compression through decomposition. To minimize compression error and to maximize the efficacy of the compressed model on the downstream tasks previous works focused on low-rank approximation of the NN's weight matrices either from the perspective of parameter importance or per-layer functional equivalence. While previous works studied the aforementioned perspectives in isolation in this work we are investigating the effectiveness of an approach that combines ideas from these two perspectives in a single objective. In parallel to this an important aspect that affects the compression quality is the distribution of the compression rate across layers and NN parameters. Earlier works mostly considered distributing the compression rate uniformly across layers and network weights or relied on computationally expensive heuristic search. Contrary to them in this work we propose an enhanced and computationally efficient algorithm for dynamic compression rate allocation. Experimental results support the efficacy of the proposed approach which performs on par or substantially better than the previous state-of-the-art especially under high compression ratios.

13:00 JST研究/論文

エンドツーエンドのRFスペクトル監視のためのエッジ効率の高いトランスフォーマー

エンドツーエンドの自動変調と秘密チャネル (CC) 認識のための E-SpecFormer (エッジ スペクトラム モニタリング トランスフォーマー) を紹介します。 RF タスクの精度を向上させながら複雑さを軽減する、Softmax および LayerNorm を使用しないアテンション メカニズムである LiTAN (Linear Tanh Attendance Network) を紹介します。 E-SpecFormer は、さまざまなハードウェア制約に対応するために、4 つのスケーラブルなバリアント (Nano、Small、Medium、Large) でパラメータ化されています。変調認識に RadioML2018 データセットを使用すると、Nano バリアントは信号対雑音比 (SNR) > 0 dB で平均精度 86.5% を達成し、ハードウェア トロイの木馬 (HT) ベースの CC データセットでは 94.2% の精度に達します。両方とも 10,000 未満のパラメータと FPGA/CPU の同時実行でフレームあたり最大 92 {\μs の速度で、これを上回ります。最先端のエッジ モデルを数分の 1 のコストで提供します。これらの結果により、E-SpecFormer は、モノのインターネット (IoT) デバイス上のリアルタイム スペクトル インテリジェンスのためのエッジ効率的なソリューションとして確立されます。リポジトリへの GitHub リンク: https://github.com/zsniko/E-SpecFormer。

原文 (English)

Edge-Efficient Transformer for End-to-End RF Spectrum Monitoring

We present E-SpecFormer (Edge Spectrum monitoring Transformer) for end-to-end automatic modulation and covert channel (CC) recognition. We introduce LiTAN (Linear Tanh Attention Network), a Softmax- and LayerNorm-free attention mechanism that reduces complexity while increasing accuracy in RF tasks. E-SpecFormer is parameterized in four scalable variants (Nano, Small, Medium, Large) to accommodate diverse hardware constraints. Using the RadioML2018 dataset for modulation recognition, the Nano variant achieves 86.5% average accuracy for Signal-to-Noise Ratios (SNRs)>0 dB, and on the hardware Trojan (HT)-based CC dataset it reaches 94.2% accuracy, both with fewer than 10k parameters and up to speed of 92 {\mu}s per frame on FPGA/CPU co-execution, surpassing state-of-the-art edge models at a fraction of their cost. These results establish E-SpecFormer as an edge-efficient solution for real-time spectrum intelligence on Internet of Things (IoT) devices. GitHub link to the repository: https://github.com/zsniko/E-SpecFormer.

13:00 JST研究/論文

ランタイム調整可能なトランジット信号の優先順位のための優先条件付き多目的強化学習

交通信号優先 (TSP) では、バス以外の交通への悪影響を制限しながらバスの遅延を削減し、一部の車両の極端な待ち時間を回避するという、競合する目標のバランスを取る必要があります。 TSP に対する既存の強化学習 (RL) アプローチは通常、交通を意識した機能 (占有率やスケジュールの偏差など) をエンコードしますが、固定報酬または固定スカラー化を最適化するため、時間帯や混乱状況によって政府機関の優先順位が変化する場合、運用の柔軟性が制限されます。優先条件付き TSP コントローラー $\pi(a \mid s,w)$ を紹介します。これは、最小/最大グリーンおよび遷移実現可能性制約の下で次の信号フェーズを選択し、再トレーニングすることなく全体的なトラフィック遅延に対してバス優先度の強調をトレードオフするために優先パラメーター $w$ を介して実行時に調整できます。制約付き信号制御/TSP ラッパーを導入することでこれを IntersectionZoo の上に実装し、バス普及率の強化とタイムテーブル ベースのバス挿入によりシナリオ生成を拡張して、トレーニング中のまばらな交通優先イベントに対処します。固定時間制御、ルールベースの TSP オーバーレイ、および固定重み付け PPO スペシャリストに対する実験では、学習された単一の条件付きポリシーが、ランタイム設定全体にわたる滑らかな経験的トレードオフ フロンティアにまたがり、固定時間およびルールベースのベースラインを上回り、制約の実現可能性を維持することが示されています。一方、テール遅延診断により、バス以外の外部性は中程度の設定では制限されたままであるが、高いバス優先度の重み付けの下では大幅に増加する可能性があることが明らかになりました。この作品のソースコードは https://github.com/urbanAIthi/morl-tsp で入手できます。

原文 (English)

Preference-Conditioned Multi-Objective Reinforcement Learning for Runtime-Tunable Transit Signal Priority

Transit signal priority (TSP) requires balancing competing objectives: reducing bus delay while limiting adverse impacts on non-bus traffic and avoiding extreme waits for a subset of vehicles. Existing reinforcement-learning (RL) approaches to TSP typically encode transit-aware features (e.g., occupancy and schedule deviation) but optimize a fixed reward or fixed scalarization, which limits operational flexibility when agency priorities change across time-of-day or disruption conditions. We present a preference-conditioned TSP controller, $\pi(a \mid s,w)$, that selects the next signal phase under minimum/maximum green and transition-feasibility constraints and can be tuned at runtime via a preference parameter $w$ to trade off bus-priority emphasis against overall traffic delay without retraining. We implement this on top of IntersectionZoo by introducing a constrained signal-control/TSP wrapper, and we extend scenario generation with bus-prevalence augmentation and timetable-based bus insertion to address sparse transit-priority events during training. Experiments against fixed-time control, a rule-based TSP overlay, and fixed-weight PPO specialists show that a single learned conditioned policy spans a smooth empirical trade-off frontier across runtime preferences, outperforms fixed-time and rule-based baselines, and maintains constraint feasibility, while tail-delay diagnostics reveal that non-bus externalities remain limited for moderate preference settings but can increase substantially under high bus-priority weights. The source code of this work is available at https://github.com/urbanAIthi/morl-tsp.

13:00 JST研究/論文Claude

BearingNAS: ラップトップを使用したベアリングのセンサー内インテリジェント故障診断システムの取得

このペーパーでは、センサー内処理を介してインテリジェンスをセンサー ダイに直接シフトするように設計されたハードウェア認識ニューラル アーキテクチャ検索 (HW-NAS) フレームワークである BearingNAS について紹介します。 BearingNAS は、極度のマイクロ予算 (4 ~ 8 kiB の RAM と 16 ~ 32 kiB のフラッシュ) を対象とした制約付きの最適化問題として検索を組み立てます。高価な個別 GPU への依存を排除​​するために、パラメータの爆発を防ぐために減衰カーネル成長定式化を活用する単一のデータフロー検索スペースと組み合わせた、軽量で派生のない検索戦略を提案します。私たちは、Case Western Reserve University (CWRU) のベンチマークに基づいてフレームワークを評価し、STMicroelectronics の 3 つのターゲット (2 つの汎用マイクロコントローラーと LSM6DSO16IS インテリジェント センサー プロセッシング ユニット (ISPU)) のアーキテクチャを最適化します。ラップトップの CPU で完全に実行されるため、検索は 1 時間以内に収束します。その結果得られた最高のセンサー内アーキテクチャは、ISPU で 99.50\% という非常に競争力の高い診断精度を達成します。これらの結果は、センサー パッケージ内で機械学習のワークロードをシフトし、低コストで生産規模のベアリングの故障診断を可能にする実現可能性を示しています。

原文 (English)

BearingNAS: Obtaining In-Sensor Intelligent Fault Diagnosis Systems for Bearings Using a Laptop

This paper introduces BearingNAS, a Hardware-Aware Neural Architecture Search (HW-NAS) framework designed to shift the intelligence directly onto the sensor die via in-sensor processing. BearingNAS frames the search as a constrained optimization problem targeting extreme micro-budgets (4 to 8 kiB of RAM and 16 to 32 kiB of Flash). To eliminate the reliance on expensive discrete GPUs, we propose a lightweight, derivative-free search strategy paired with a single data-flow search space that leverages a decaying kernel growth formulation to prevent parameter explosion. We evaluate our framework on the Case Western Reserve University (CWRU) bearing benchmark, optimizing architectures for three STMicroelectronics targets: two commodity microcontrollers and the LSM6DSO16IS Intelligent Sensor Processing Unit (ISPU). Running entirely on a laptop CPU, the search converges in less than an hour. The resulting best in-sensor architecture achieves a highly competitive diagnostic accuracy of 99.50\% on the ISPU. These results demonstrate the viability of shifting the machine learning workload inside the sensor package, enabling low-cost, production-scale bearing fault diagnosis.

13:00 JST研究/論文

原則的な継続的異常検出に向けて: 体系的なフレームワークとベンチマーク シナリオ

継続的異常検出 (CAD) は、以前に観察された状況でのパフォーマンスを維持しながら、進化するデータ分布にモデルがどのように適応できるかを研究します。ただし、CAD ベンチマークは、タスクがどのように定義、フィルタリング、順序付け、検証されるかに大きく依存します。表形式のドメインでは、タスクの境界が与えられることはほとんどなく、恣意的な分割により、学習不可能、冗長、または過度に転送可能なタスクが作成され、真の継続的な学習行動が曖昧になる可能性があります。この目的を達成するために、既存の表形式の異常検出データセットから再現可能なベンチマーク シナリオを設計するための体系的なフレームワークを導入します。このフレームワークは候補タスクを発見し、不適切なタスクをフィルタリングして、多様なダイナミクスを明らかにする原則に基づいた順序付けを導き出します。このフレームワークにより、3 つの大規模なサイバーセキュリティ異常検出データセットから 5 つのベンチマーク対応シナリオを提供でき、単一データセットと複数データセットの両方の CAD 設定が得られます。

原文 (English)

Towards Principled Continual Anomaly Detection: A Systematic Framework and Benchmark Scenarios

Continual anomaly detection (CAD) studies how models can adapt to evolving data distributions while retaining performance on previously observed regimes. CAD benchmarks, however, depend critically on how tasks are defined, filtered, ordered, and validated. In tabular domains, task boundaries are rarely given, and arbitrary splits can create unlearnable, redundant, or overly transferable tasks that obscure genuine continual-learning behavior. To this end, we introduce a systematic framework for reproducible benchmark scenario design from existing tabular anomaly-detection datasets. The framework discovers candidate tasks, filters unsuitable tasks, and derives principled orderings that expose diverse dynamics. The framework allows us to deliver five benchmark-ready scenarios from three large-scale cybersecurity anomaly detection datasets, yielding both single-dataset and multi-dataset CAD settings.

13:00 JST研究/論文

SechKAN: 双曲線セカント関数を備えたコルモゴロフ・アーノルドネットワーク

近年、コルモゴロフ・アーノルド ネットワーク (KAN) は、機械学習および科学計算タスクにおける有効性によりますます注目を集めており、ニューラル ネットワーク設計に新しいパラダイムを提供しています。この論文では、双曲割線 (sech) 関数に基づく KAN アーキテクチャである SechKAN を紹介します。双曲割線基底は、滑らかな鐘形の形状、局所的な応答、および安定した勾配のために使用されます。 1D 線形変換を採用してパラメータの数を減らし、SechKAN のモデル サイズを多層パーセプトロン (MLP) と同等に保つことができます。実験結果は、MNIST、Fashion-MNIST、CIFAR-10、CIFAR-100 などのベンチマーク データセットでの関数フィッティング、PDE 問題、および画像分類タスクにおける SechKAN の有効性を示しています。 SechKAN は、MLP や他の KAN バリアントと比較して、同様の数のパラメータを維持しながら、優れたパフォーマンスを実現します。ただし、実行時間は他の KAN 亜種よりも優れていますが、MLP よりわずかに長くなります。

原文 (English)

SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions

In recent years, Kolmogorov-Arnold Networks (KANs) have attracted increasing attention due to their effectiveness in machine learning and scientific computing tasks, offering a new paradigm for neural network design. In this paper, we present SechKAN, a KAN architecture based on hyperbolic secant (sech) functions. The hyperbolic secant basis is used for its smooth bell-shaped form, localized responses, and stable gradients. We employ 1D linear transformations to reduce the number of parameters, allowing SechKAN to remain comparable to multilayer perceptrons (MLPs) in model size. Experimental results indicate the effectiveness of SechKAN in function fitting, PDE problems, and image classification tasks on benchmark datasets, including MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100. SechKAN achieves superior performance compared to MLPs and other KAN variants while maintaining a similar number of parameters. However, its running time, while better than that of other KAN variants, is slightly longer than that of MLPs.

13:00 JST研究/論文

効率的な時間依存信頼性解析のためのデュアルドメイン融合 LSTM モデリング

時間依存の信頼性解析は、不確実性の下でエンジニアリング システムの長期的な安全性とパフォーマンスを確保するために重要です。ただし、従来のサロゲート モデル手法では、時間に依存しない確率変数を組み込み、時間依存の確率過程との複雑な相互作用を捉えるのに苦労することがよくあります。この制限を克服するために、この論文では、効率的かつ正確な時間依存信頼性解析のためのデュアルドメイン融合長短期記憶 (DDF-LSTM) モデルを提案します。時間依存ドメインと時間非依存ドメインの両方からの情報を共同処理するための新しいネットワーク アーキテクチャが開発されました。具体的には、時間独立変数が初期隠れ状態に埋め込まれ、全結合層が導入されて LSTM 出力と時間独立変数の両方を最終出力空間にマッピングします。さらに、改良された損失関数は、最小応答に対するモデルの感度を強調するように設計されており、それによって故障確率推定の精度が向上します。提案された方法は、確率変数、確率過程、および限界状態関数の時間的挙動間の依存関係を効果的に捕捉します。 DDF-LSTM モデルをトレーニングすると、最小限の計算コストで時間依存の故障確率を推定する効率的なモンテカルロ シミュレーションが可能になります。 4 つのケーススタディは、提案された方法の計算効率と予測精度の向上を検証します。

原文 (English)

Dual-domain fused LSTM modeling for efficient time-dependent reliability analysis

Time-dependent reliability analysis is crucial for ensuring the long-term safety and performance of engineering systems under uncertainties. However, traditional surrogate model methods often struggle to incorporate time-independent random variables and capture their complex interactions with time-dependent stochastic processes. To overcome this limitation, this paper proposes a dual-domain fused long short-term memory (DDF-LSTM) model for efficient and accurate time-dependent reliability analysis. A novel network architecture is developed to jointly process information from both time-dependent and time-independent domains. Specifically, the time-independent variables are embedded into the initial hidden states, and a fully connected layer is introduced to map both LSTM outputs and time-independent variables into the final output space. Furthermore, an improved loss function is designed to emphasize the model's sensitivity to minimum responses, thereby improving the precision of failure probability estimation. The proposed method effectively captures the dependencies among random variables, stochastic processes, and the temporal behavior of limit state functions. Once trained, the DDF-LSTM model enables efficient Monte Carlo simulation to estimate time-dependent failure probabilities with minimal computational cost. Four case studies validate the proposed method's enhanced computational efficiency and predictive accuracy.

13:00 JSTLLM/生成AI

信頼性は逆にスケールします: 隠れた自動回帰リスク体制により、モデルが大きくなるとミスがより早く増加します

言語モデルがスケールするにつれて、答えはより真実になり始めますが、劣化が早くなります。スケールすると機能は得られますが、信頼性が損なわれます。知識ギャップ アカウント (より多くのデータ、取得、スケール) は、スケールが鋭くなる自己回帰リスク残差を見逃します。モデルは確率の低いトークンにコミットし、確立された条件が雪だるま式に増加します。これを、より強力な同族オラクルに対するポジションごとの不一致 $\delta = \log p_M - \log p_O$ を通じて追跡します。その 2 番目のモーメントは、バイアス $^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ とリスク $\mathrm{Var}[\delta]$ に正確に分割されます。我々は 4 つの調査結果を提示します。(i) スケーリングの下で​​は、知識ギャップは $\about$$6\times$ 減少する一方、知識の低下は $11$-$39\time$ 増加します。 (ii) 捏造では、不確実性 $H(p_M)$ がすぐに緩和される一方、オラクル参照のリスクは最大 $17\times$ まで持続し、連続した捏造を橋渡しする自信はあるが不安定なリスク体制が残ります ($14$B で $+69\%$)。 (iii) この体制には因果関係があります。ポリシーに基づいた固定 $\mathrm{KL}$ 分散縮小により、3 つのモデル ファミリー全体で Web 検証済み幻覚が $35$-$74\%$ 削減されます。そして、(iv) $p_M$ のみの検出器 (セマンティック エントロピーなど) が構造的に自己監視を回避し、$4\time$ 近く多くの捏造を保持する危険なブランチに対して $\およそ$$30\%$ 少ない量 ($p<10^{-16}$) を発火させます。モデルが大きくなると、支配的で、自己永続的で、因果関係があり、モデル自体には見えない故障モードによって、雪だるま式にミスが速くなります。

原文 (English)

Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto-Regressive Risk Regime

As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account - more data, retrieval, or scale - misses an auto-regressive risk residual that scale sharpens: the model commits to a low-probability token, conditions on it as established, and snowballs. We track this through per-position disagreement $\delta = \log p_M - \log p_O$ against a stronger same-family oracle, whose second moment splits exactly into bias$^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ and risk $\mathrm{Var}[\delta]$. We present four findings: (i) under scaling, the knowledge gap falls $\approx$$6\times$ while knowledge degradation grows $11$-$39\times$; (ii) at a fabrication, felt uncertainty $H(p_M)$ relaxes quickly while oracle-referenced risk persists up to $17\times$ longer, leaving a confident-but-precarious risk regime that bridges consecutive fabrications ($+69\%$ at $14$B); (iii) this regime is causal - an on-policy, fixed-$\mathrm{KL}$ variance contraction cuts web-verified hallucination by $35$-$74\%$ across three model families; and, (iv) it structurally evades self-monitoring, with $p_M$-only detectors (e.g. semantic entropy) firing $\approx$$30\%$ less ($p<10^{-16}$) on the risky branch holding nearly $4\times$ more fabrications. Bigger models snowball mistakes faster, through a failure mode that is dominant, self-perpetuating, causal and invisible to the model itself.

13:00 JST研究/論文

情報の影: 言語モデルが学習できる構造的限界の測定

言語モデルが知っていることに対する制限のいくつかは、データ範囲のギャップではなく、テキストからの学習の構造的特性です。情報シャドウを導入します。これは、テキストで訓練された学習者が規模に関係なく獲得できない現象の領域であり、(I) 言語で表現できない構造、(II) 訓練分布から統計的に特定できない関数、(III) 表現可能だが勾配ベースの訓練では到達できない関数で構成されます。その設定では影の前提が証明可能であるため、各タイプに決定的なプローブを与えます。タイプ I の場合、言語圧縮残差は、信号の非可逆テキスト形式のエンコードのみを参照するテキスト学習器と、基礎となる信号を直接参照する完全信号学習器を比較します。テキスト学習者は計算可能な表現可能性の上限に達していますが、フルシグナル学習者は 300 倍以上のデータにわたって一定のギャップによって引き離されているため、この不足はトレーニングの特性ではなくチャネルの特性です。タイプ II の場合、反事実区別テストは、2 つの互換性のないルールと正確に一致するデータに基づいてモデルをトレーニングします。証明可能な文字列タスクと言語のような同意タスク全体で、反事実に対する動作はモデルの帰納バイアスによって設定されますが、5% の曖昧さを排除するデータは学習されたルールをどちらかのターゲットに双方向に誘導します (r = +/-1.0、p < 1e-10)。タイプ III の場合、Basin Escape Mapping は、(手作業による構築では) 100% で表現可能な関数を示しますが、標準トレーニングによっては 0% に達し、近くの初期化からは瞬時に到達します。幅のスケーリングは何の助けにもなりません (p = 1.6 x 10^-14)。各影響は、容量またはモダリティのアーティファクトを除外する制御によって分離されます。プローブ スイートをリリースし、ベンチマーク設計、機能監査、シャドウ認識の不確実性への影響について説明します。

原文 (English)

The Information Shadow: Measuring Structural Limits on What Language Models Can Learn

Some limits on what language models know are not gaps in data coverage but structural properties of learning from text. We introduce the information shadow: the region of phenomena that a text-trained learner cannot acquire regardless of scale, comprising (I) structures language cannot express, (II) functions that are statistically non-identifiable from the training distribution, and (III) functions that are representable but unreachable by gradient-based training. We give each type a probe that is decisive because the premise of the shadow is, in that setting, provable. For Type I, Language Compression Residuals compare a text learner, which sees only a lossy text-like encoding of the signal, against a full-signal learner, which sees the underlying signal directly. The text learner sits at a computable expressibility ceiling while the full-signal learner pulls away by a gap that stays flat across 300x more data, so the deficit is a property of the channel, not of training. For Type II, the Counterfactual Distinction Test trains models on data exactly consistent with two incompatible rules. Across a provable string task and a language-like agreement task, behavior on counterfactuals is set by the model's inductive bias, while 5% disambiguating data steers the learned rule bidirectionally to either target (r = +/-1.0, p < 1e-10). For Type III, Basin Escape Mapping exhibits a function that is representable at 100% (by hand construction) yet reached 0% of the time by standard training and instantly from a nearby initialization, with width scaling providing no help (p = 1.6 x 10^-14). Each effect is isolated by a control that rules out a capacity or modality artifact. We release the probe suite and discuss implications for benchmark design, capability auditing, and shadow-aware uncertainty.

13:00 JST研究/論文

シャープネスを意識した最小化のための勾配エネルギーガイドによるブロック単位の摂動

Sharpness-Aware Minimization (SAM) は、ローカル パラメーターの近傍における最悪の場合の損失を最小限に抑えることで一般化を向上させます。標準 SAM は、瞬間的なミニバッチ勾配ノルムに従って、パラメーター ブロック全体にグローバル摂動バジェットを暗黙的に割り当てます。このような割り当てにはノイズが多く、トレーニング全体を通じてブロックが蓄積する感度を反映していない可能性があります。我々は、勾配エネルギー適応半径 SAM (GEAR-SAM) を提案します。これは、二乗ブロック勾配の指数移動平均 (EMA) を軽量の曲率関連感度信号として維持し、閉じた形式の制約付き最適化を通じて固定 SAM バジェットを割り当てます。 GEAR-SAM はグローバル SAM 半径を保持し、ヘッセベクトル積や明示的なフィッシャー推定を必要とせず、SAM を超えてスカラー状態のみを追加します。画像分類、転移学習、ノイズのあるラベル学習、およびパーティション研究に関する実験により、アーキテクチャとタスク全体での一般化と堅牢性の向上が実証されています。より広範には、GEAR-SAM はシャープネスを意識した最適化の動的なビューを提供します。つまり、トレーニング中に機能ネットワーク ブロックの感度が進化するにつれて、固定された摂動バジェットが再配分される必要があります。

原文 (English)

Gradient-Energy Guided Block-Wise Perturbations for Sharpness-Aware Minimization

Sharpness-Aware Minimization (SAM) improves generalization by minimizing the worst-case loss in a local parameter neighborhood. Standard SAM implicitly allocates its global perturbation budget across parameter blocks according to instantaneous minibatch gradient norms. Such an allocation can be noisy and may not reflect the sensitivity that blocks accumulate throughout training. We propose Gradient-Energy Adaptive Radius SAM (GEAR-SAM), which maintains an exponential moving average (EMA) of squared block gradients as a lightweight, curvature-related sensitivity signal and allocates the fixed SAM budget through a closed-form constrained optimization. GEAR-SAM preserves the global SAM radius, requires no Hessian-vector products or explicit Fisher estimation, and adds only scalar state beyond SAM. Experiments on image classification, transfer learning, noisy-label learning, and partition studies demonstrate improved generalization and robustness across architectures and tasks. More broadly, GEAR-SAM provides a dynamic view of sharpness-aware optimization: a fixed perturbation budget should be redistributed as the sensitivity of functional network blocks evolves during training.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

グレーボックス シミュレーション モデルのエージェント キャリブレーション: LLM 主導の代替手段

グレーボックス シミュレーション モデルのキャリブレーションは、モデルの評価にコストがかかり、パラメーター空間が高次元になる可能性があり、検索では妥当性の制約を考慮する必要がある制約付きの最適化問題です。解析者はシミュレーション コードを完全に利用できますが、複数のパラメーターの共同効果を解析的に予測することは依然として困難です。 Nelder--Mead (NM) などの従来のオプティマイザは導入が簡単ですが、特に制約がある場合にはサンプルの効率が悪くなります。最新のベイジアン最適化手法は、はるかに少ない評価で競争力のあるソリューションを実現しますが、制約を処理するために自明ではないモデリング機械を必要とします。私たちは、大規模な言語モデルがオプティマイザーとして機能し、制約がシステム プロンプトの平易な言語セクションとして組み込まれるエージェント キャリブレーション方法を導入します。非制約キャリブレーションと臨床制約キャリブレーションの両方の下で、肛門癌シミュレーション モデルのエージェント手法、NM、およびベイジアン最適化 (BO) を評価します。制約のないキャリブレーションでは、エージェント手法は BO および NM よりも大幅に低い最良誤差を達成しながら、必要なモデル評価の数は少なくなります。制約されたキャリブレーションの下では、エージェント手法は同等の誤差レベルに達し、両方とも NM を上回ります。これらの結果は、反復ごとの推論時間の増加を犠牲にして得られます。エージェント キャリブレーションは、実質的に少ないモデル評価で競争力のあるパフォーマンスを実現し、追加のモデリング機構ではなく単純なテキスト仕様を通じて、モデラー側のインターフェイスでの制約処理が基本的に無料になります。主なトレードオフは反復ごとの推論コストの増加にあり、このアプローチはシミュレーション時間が支配的な場合に特に適しています。パフォーマンスを超えて、反復ごとの理論的根拠により、検索が監査可能で説明可能になるため、その決定を精査して第三者に対して正当化することができます。

原文 (English)

Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative

Calibration of grey-box simulation models is a constrained optimization problem in which model evaluations are expensive, the parameter space can be high-dimensional, and the search must respect plausibility constraints. Although the simulation code is fully available to the analyst, the joint effect of multiple parameters remains difficult to predict analytically. Classical optimizers such as Nelder--Mead (NM) are simple to deploy but sample-inefficient, particularly under constraints. Modern Bayesian Optimization methods achieve competitive solutions with far fewer evaluations but require non-trivial modeling machinery for constraint handling. We introduce an agentic calibration method in which a large language model acts as the optimizer, with constraints incorporated as a plain-language section of the system prompt. We evaluate the agentic method, NM, and Bayesian Optimization (BO) on an anal cancer simulation model under both unconstrained and clinically constrained calibration. Under unconstrained calibration, the agentic method achieves substantially lower best error than BO and NM, while requiring fewer model evaluations. Under constrained calibration, the agentic method reaches comparable error levels and both outperform NM. These results are obtained at the cost of increased inference time per iteration. Agentic calibration achieves competitive performance with substantially fewer model evaluations, and constraint handling is essentially free at the modeller-facing interface through simple textual specifications rather than additional modelling machinery. The main trade-off lies in increased per-iteration inference cost, making the approach particularly suitable when simulation time dominates. Beyond performance, the per-iteration rationale makes the search auditable and explainable, so its decisions can be scrutinised and justified to third parties.

13:00 JSTLLM/生成AIエージェントQwen

分布優先の母集団シミュレーション: Non-WEIRD LLM ペルソナ モデリングにおける崩壊、キャリブレーション、およびリコール

人口合成ツールは、すべての個人を独立した大規模言語モデル (LLM) エージェントとして実行することが増えています。実際の調査マイクロデータを使用して、このパラダイムには基本的な故障モードがあることを示し、それに対して分布優先の修正を設定しました。これらはすべて、非 WEIRD (トルコ優先) データに対する決定論的で構成が検証された検証器で測定されました。まず、2,414 人の現実世界価値観調査回答者に基づいた N 個の独立した LLM エージェントは、母集団の応答分布を再現できません。モーダルデフォルト (4 つのシナリオ x 5 つのシード: 濃度 0.36 -> 0.69、エントロピー 1.46 -> 0.77、85% 崩壊、TVD=0.44) の上に積み重なり、崩壊はシナリオ構造の予測可能な関数です (単一回答で r=0.55)。構造)。第二に、言語化サンプリング (VS) は、3 つのモデル ファミリでのトレーニングなしでフィールドの慢性的な過小分散を修正します (忠実度 +7 ~ +10、Qwen で有意、p=0.002、d=6.2) が、同じ動きが普遍的にオーバーシュートして過分散 (SD 比 0.4 ~ 0.56 -> 1.26 ~ 1.37) に陥ります。これは、VS の構造的特性です。第三に、調査の忠実度はエージェントの行動にはわずかしか反映されません。単一モデル、単一ドメインの予約タスクでは、ペルソナは、収入によって調整されるものの上書きされない最も安価なデフォルト(約 80%)によって支配されます(収入帯全体で 0%->7%->32% の快適な選択)。第四に、プラセボ対照暗記攻撃と選挙バックテストは、サブグループと個人の主張が想起と過少決定によって汚染されている一方で、VSが総力を維持していることを示している。最後に修正で終了します。分布を 1 回 (VS) モデル化し、それを O(1) コストで接地された文字に割り当てます。予算を考慮したルーターを使用します。その正直な AUC は、コード由来のオラクルのトートロジー 1.0 ではなく、0.805 です。中心的な寄与は、現実性を主張する必要はありません。それは、独立エージェント ルートの内部不整合と、配布優先ルートが調整される条件を測定します。

原文 (English)

Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling

Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent. Using real survey microdata, we show that this paradigm has a basic failure mode, and we set a distribution-first corrective against it, all measured with a deterministic, construct-validated verifier on non-WEIRD (Turkey-first) data. First, N independent LLM agents grounded on 2,414 real World Values Survey respondents fail to reproduce the population's response distribution: they pile onto a modal default (four scenarios x five seeds: concentration 0.36->0.69, entropy 1.46->0.77, 85% collapse, TVD=0.44), and the collapse is a predictable function of scenario structure (r=0.55 with a single-answer structure). Second, Verbalized Sampling (VS) fixes the field's chronic under-dispersion without training in three model families (fidelity +7 to +10; significant on Qwen, p=0.002, d=6.2), yet the same move universally overshoots into over-dispersion (SD-ratio 0.4-0.56 -> 1.26-1.37), a structural property of VS. Third, survey fidelity transfers only weakly to agentic behavior: in a single-model, single-domain booking task, a persona is dominated by a cheapest-default (~80%) that income modulates but does not override (comfort choice 0%->7%->32% across income bands). Fourth, a placebo-controlled memorization attack and an election backtest show VS keeps aggregate strength while subgroup and individual claims are contaminated by recall and underdetermination. We close with the corrective: model the distribution once (VS) and assign it to grounded characters at O(1) cost, with a budget-aware router whose honest AUC is 0.805, not the tautological 1.0 of a code-derived oracle. The central contribution needs no realism claim: it measures the internal inconsistency of the independent-agent route and the conditions under which the distribution-first route calibrates.

13:00 JST研究/論文

グラフニューラルネットワークを使用した系統樹間のSPR距離の近似

系統樹トポロジーの比較は流行のダイナミクスを理解するために不可欠ですが、サブツリー剪定と再移植 (SPR) 距離などの生物学的に意味のある距離は NP で計算するのが難しく、大規模なデータセットでは扱いが困難です。グラフ ニューラル ネットワーク (GNN) がトレーニング後の比較ごとにほぼ一定の時間で SPR 距離を近似できるかどうかを調査します。私たちの貢献は 4 つあります。まず、UPGMA と近隣結合を使用して 4 つの細菌種について推定された 864 の系統樹のデータセットを構築し、公開します。このデータセットは最大 9,500 の分離株に及び、388 のラベル付きツリーペアも含まれます。次に、中間点の再ルート化を含む再現可能な前処理パイプラインを確立します。これにより、ツリーの深さが減り、正確な距離の計算とモデルのルートベースの特徴に必要なルート化が提供されます。第三に、監視ターゲットを検証します。正確な SPR が扱いやすい小さな木では、根のない phangorn::SPR.dist ヒューリスティックは、rspr (Pearson $0.98$--$0.99$) によって計算された正確な根の距離とほぼ完全に相関し、優れた単調サロゲートになります。最後に、Siamese Graph Isomorphism Network (GIN) リグレッサーをトレーニングします。分布内、つまりトレーニングと同じ種およびサイズ範囲のホールドアウトされた木では、分散のおよそ 87 ~ 90% (ホールドアウト分割では $R^2 \約 0.87$、層別相互検証では $0.90 \pm 0.19$) が説明され、平均予測ベースラインよりも誤差が約 4 倍低く、未確認の種への部分的な移動が示されます ($R^2 \およそ0.37ドル)。その主な制限は、トレーニングで見られるものよりも大きなツリーへの外挿であり、精度が低下します。公開されたデータセットと検証されたヒューリスティック対正確な関係は、学習された SPR 近似をスケーリングするための再現可能な基礎を提供します。

原文 (English)

Approximating SPR Distance Between Phylogenetic Trees with Graph Neural Networks

Comparing phylogenetic tree topologies is essential for understanding epidemic dynamics, yet biologically meaningful distances such as the Subtree Prune and Regraft (SPR) distance are NP-hard to compute and intractable on large datasets. We investigate whether a Graph Neural Network (GNN) can approximate SPR distances in near-constant time per comparison after training. Our contributions are fourfold. First, we build and publicly release a dataset of 864 phylogenetic trees inferred with UPGMA and Neighbor-Joining over four bacterial species, spanning up to 9{,}500 isolates, together with 388 labelled tree pairs. Second, we establish a reproducible pre-processing pipeline including midpoint re-rooting, which reduces tree depth and supplies the rooting required for exact distance computation and for the model's root-based features. Third, we validate the supervision target: on small trees, where exact SPR is tractable, the unrooted phangorn::SPR.dist heuristic correlates almost perfectly with the exact rooted distance computed by rspr (Pearson $0.98$--$0.99$), making it an excellent monotonic surrogate. Lastly, we train a Siamese Graph Isomorphism Network (GIN) regressor. In-distribution, i.e., held-out trees from the same species and size range as training, it explains roughly 87--90% of the variance ($R^2 \approx 0.87$ on a held-out split; $0.90 \pm 0.19$ under stratified cross-validation), with about four times lower error than a mean-predictor baseline, and shows partial transfer to unseen species ($R^2 \approx 0.37$). Its main limitation is extrapolation to trees larger than those seen in training, where accuracy collapses. The released dataset and the validated heuristic versus exact relationship provide a reproducible basis for scaling learned SPR approximation.

13:00 JSTLLM/生成AIエージェントClaude

マルチステップツールで強化されたエージェントのバインディングドリフト

ツールで拡張された言語モデル エージェントは、外部システム上で複数ステップのワークフローを実行し、エンティティを一度解決すると、後続のステップにわたってそのエンティティに基づいて動作します。これまでの研究によると、シングルステップ アクションでは、エージェントは正しいツールを選択しますが、24 ~ 26% の確率でそれを間違ったエンティティにバインドします。私たちは、エンティティ バインディングに時間の経過とともに何が起こるかを研究します。エンティティ バインディングは正しいままなのか、静かに別のエンティティに移動するのか、それとも最初から間違っていた場合は伝播して複合化するのか?バインディング ドリフト (ステップ 1 で正しく、後で間違っている) をエラー伝播 (ステップ 1 で間違っていて、繰り越される) とは別のものとして形式化し、これら 2 つが混同されないように、素のワークフロー セットでそれらをスコアリングします。制御されたマルチステップのテストベッド (200 のワークフロー、580 のエンティティ バインディング スコア付きステップ、4 つのエンタープライズ ドメイン、小規模から最前線にわたる 8 つのモデル バックエンド) では、次のことがわかりました: (1) 制御されたエラー挿入下では、エンティティ ロック (直感的な「最初のバインディングを保持する」修正) により、誤ったアクションが 907 から 2,746 (3.0 倍、ブートストラップ 95% CI [2.8, 3.3])、シードされた間違ったエンティティを後のすべてのステップに忠実に持ち込むためです。 (2) 最も影響を受けるモデル (Claude Opus 4.5) では増幅が 8.5 倍に達します。 (3) 実用的な LLM ベースの再検証器 (元の命令を再読み取りする単一の安価な 2 番目のモデル呼び出し) により、間違ったアクションが 79% (0.21x; CI [0.18, 0.25]) 削減され、オラクルの上限 (0.20x) の 1 パーセント ポイント以内にギャップが縮まります。 (4) 自然な (注入されていない) 設定では、ベースライン エージェントは対象となるワークフローの 18% で変動し、ステップごとのエラー率がステップごとに上昇します。永続性と再検証は互換性がありません。ドリフトを排除する防御は伝播を悪化させる可能性があり、実用的な再検証はオラクルの回復とほぼ同等です。

原文 (English)

Binding Drift in Multi-Step Tool-Augmented Agents

Tool-augmented language-model agents execute multi-step workflows over external systems, resolving an entity once and then acting on it across subsequent steps. Prior work shows that in single-step actions, agents select the correct tool but bind it to the wrong entity 24-26% of the time. We study what happens to entity bindings over time: do they stay correct, silently drift to a different entity, or, if wrong from the start, propagate and compound? We formalize binding drift (correct at step 1, wrong later) as distinct from error propagation (wrong at step 1, carried forward), and score them on disjoint workflow sets so the two cannot be conflated. In a controlled multi-step testbed (200 workflows, 580 entity-binding-scored steps, four enterprise domains, eight model backends spanning small to frontier), we find: (1) under controlled error injection, an entity lock (the intuitive "persist the first binding" fix) amplifies wrong actions from 907 to 2,746 (3.0x; bootstrap 95% CI [2.8, 3.3]), because it faithfully carries the seeded wrong entity into every later step; (2) the amplification reaches 8.5x on the most affected model (Claude Opus 4.5); (3) a practical LLM-based re-verifier (a single cheap second model call re-reading the original instruction) reduces wrong actions by 79% (0.21x; CI [0.18, 0.25]), closing the gap to within 1 percentage point of an oracle upper-bound (0.20x); and (4) in the natural (non-injected) setting, baseline agents drift on 18% of eligible workflows, with the per-step error rate rising across steps. Persistence and re-verification are not interchangeable: a defense that eliminates drift can worsen propagation, and a practical re-verifier nearly matches oracle recovery.

13:00 JST研究/論文Claude

リアクティブな計算グラフのコスト計算: 徹底的なスイープ、逐次突然変異、および後方局所性ギャップ

ニューラル ネットワークの計算グラフに対する徹底的なサイトごとの介入 (アクティベーション パッチング スイープ、回路発見検索、体系的なアブレーション 研究) により、すべての候補サイトでグラフが変異し、そのコストは各変異後の再計算によって支配されます。無効化が変異したノードの下流のコーンに正確に触れるリアクティブ グラフ エンジンでは、そのようなワークロードを考慮した完全なコストを計算します。まず、独立した完全な再計算にわたる徹底的なスイープの総速度向上は普遍的な定数ではありません。層ごとの重みがカラマタ インデックス q で深さとともに規則的に変化する場合、重みが出力付近に集中するとき比率は (q+2)/(q+1) に収束し、入力付近で q+2 に収束し、深さが均一な場合にのみ 2 に戻ります。時計の結果から、インタプリタのオーバーヘッドが解消されるまで、上限は 2 を下回る約 1.79 になると予測されます。第 2 に、挿入間で取り消されることのない一連の永続的変異の正確なコストを証明します。インターリーブ コストは、挿入順序に対する閉形式の極値を持ち、比較可能なサイト ペアで合計された正確な超過数によって分離合計を超えますが、バッチ アプリケーションは順序に依存せず、準加法的であり、サイトのコーンと新鮮なノードの結合に正確にコストがかかります。 3 番目に、後方パスの前方局所性の正確なミラーを証明し、長いスキップ接続のないアーキテクチャでのバックプロパゲーション下では総速度アップが 1 まで崩壊することを示します。すべてのアイデンティティは、Julia のリアクティブ グラフ エンジンである NeuroDSL で検証されます。測定されたスイープ比は、4 つのコスト プロファイルの下で予測された制限に収束します。トレーニング モードの比率は、予測されたレートで 1 に減少します。グラフトごとの 18 の連続コストすべてとバッチ合計は、3 つの挿入オーダーにわたってゼロトレランスでクローズド フォームと一致します。

原文 (English)

Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap

Exhaustive site-by-site interventions on a neural network's computational graph -- activation-patching sweeps, circuit-discovery searches, systematic ablation studies -- mutate the graph at every candidate site, and their cost is dominated by recomputation after each mutation. On a reactive graph engine whose invalidation provably touches exactly the downstream cone of a mutated node, we give a complete cost accounting for such workloads. First, the aggregate speedup of an exhaustive sweep over independent full recomputations is not a universal constant: if per-layer weight varies regularly with depth at Karamata index q, the ratio converges to (q+2)/(q+1) when weight concentrates near the output and to q+2 near the input, recovering 2 only in the depth-uniform case; a wall-clock corollary predicts a ceiling of about 1.79, below 2, until interpreter overhead is compiled away. Second, we prove the exact cost of a sequence of persistent mutations, never undone between insertions: the interleaved cost exceeds the isolated sum by an exact overcount summed over comparable site pairs, with closed-form extremes over insertion orders, while batched application is order-independent and sub-additive, costing exactly the union of the sites' cones plus the fresh nodes. Third, we prove the exact mirror of forward locality for the backward pass, showing it collapses the aggregate speedup to 1 under backpropagation on architectures without long skip connections. Every identity is validated on NeuroDSL, a reactive graph engine in Julia: measured sweep ratios converge to the predicted limits under four cost profiles; the training-mode ratio collapses to 1 at the predicted rate; and all 18 per-graft sequential costs and the batched total match the closed forms at zero tolerance across three insertion orders.

13:00 JST画像/動画生成ロボティクス

危険か異常か?危険性と矛盾を理解するための VLM の評価

現代の安全性が重要なシステムは、災害リスクを軽減し、緊急時の意思決定をサポートするために、人間とロボットのインタラクションにますます依存しています。視覚言語モデル (VLM) は、複雑なシーンを解釈し、安全関連の情報を伝達できるため、これらの設定に有望ですが、信頼できる安全推論を保証するには依然として慎重な評価が必要です。特に、現在の評価では、危険認識を二者決定 (安全/危険) として組み立てることが多く、モデルが真の物理的危険を特定しているのか、それとも単に異常なシーン要素に反応しているのかが不明確になっています。私たちは、危険と異常の明確な区別を導入し、危険な状態と異常な状態を別々に認識することで、この制限に対処します。 2 つのデータセットと複数のプロンプト戦略にわたっていくつかの最先端の VLM を評価し、この違いがモデルの動作を変えるかどうかをテストします。私たちの結果は、VLM が異常性を危険性と誤って解釈することが多く、危険性の代用として文脈上の不規則性を過度に依存していることを示しています。さらに、異常と危険を明示的に分離することで、VLM の安全推論のより有益な評価が提供され、二元的な安全性の判断では曖昧になる可能性のある故障モードが明らかになることを示します。私たちの公開データセットは Roboflow https://app.roboflow.com/vlm-in-context-anomaly-and-hazard-detection/camera-ready-roman-ds で入手できます。

原文 (English)

Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Vision-Language Models (VLMs) are promising for these settings because they can interpret complex scenes and communicate safety-relevant information, but they still require careful evaluation to ensure reliable safety reasoning. In particular, current evaluations often frame danger recognition as a binary decision (Safe/Unsafe), making it unclear whether a model is identifying true physical hazards or merely reacting to unusual scene elements. We address this limitation by introducing an explicit distinction between hazard and anomaly, and by separately recognizing hazardous and anomalous states. We evaluate several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether this distinction changes model behavior. Our results show that VLMs frequently misinterpret anomalousness as hazardousness, revealing an over-reliance on contextual irregularity as a proxy for danger. We further show that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure. Our public dataset is available on Roboflow https://app.roboflow.com/vlm-in-context-anomaly-and-hazard-detection/camera-ready-roman-ds.

13:00 JST研究/論文

回転式 SOH 注入型事前電池変圧器によるリチウムイオン電池の SOH と RUL の統合予測のための動的損失バランシング

信頼性の高いリチウムイオン電池管理システムの導入は、電動化を加速するために極めて重要ですが、健康状態(SOH)と残存耐用年数(RUL)の共同予後は、タスクの不均一分散性によって依然として大きく妨げられています。従来のマルチタスク学習フレームワークは、SOH 推定の境界のある低分散ノイズと、長期 RUL 予測の境界がなく非線形に拡大する不確実性のバランスを取ることができません。ここでは、これらの最適化の競合を解決する統合共同推定フレームワークである Rotary SOH-Injected Prior Battery Transformer (RoSIP-Batt) を紹介します。 RoSIP-Batt は、同時予測をベイジアン マルチタスク目標として定式化することにより、学習された残留ノイズ レベルに基づいてタスク固有の勾配を動的にスケーリングする等分散不確かさの重み付けメカニズムを導入します。このアーキテクチャは、分離された二重分類トークンと次元ごとのゲート融合メカニズムを活用し、勾配分離演算子によって保護され、高分散 RUL 更新による安定した SOH 表現空間の破損を防ぎます。絶対的なサイクルステップに依存せずに電気化学的劣化パターンを捕捉するために、Rotary Position Embedding (RoPE) が共有の Transformer バックボーンに組み込まれ、並進不変の相対時間プロファイルをモデル化します。重要なのは、中間の SOH 推定値が物理的劣化事前として RUL 回帰ヘッドに直接注入されることです。 NASA、MIT-スタンフォード、および HUST データセットにわたる評価では、RoSIP-Batt が最先端のベースラインを大幅に上回っており、NASA では SOH 推定誤差が 1.994% MAE に減少し、スタンフォードでは RUL 予測誤差が 62.85 サイクルに制限されていることが示されています。これらの発見により、RoSIP-Batt は、リアルタイムの組み込み BMS 展開に適した、汎用性が高く、計算効率の高いソリューションとして確立されます。

原文 (English)

Dynamic Loss Balancing for Joint SOH and RUL Prediction of Lithium-Ion Batteries via a Rotary SOH-Injected Prior Battery Transformer

The deployment of reliable lithium-ion battery management systems is crucial for accelerating electrification, yet the joint prognosis of State of Health (SOH) and Remaining Useful Life (RUL) remains severely hindered by task heteroscedasticity. Conventional multi-task learning frameworks fail to balance the bounded, low-variance noise of SOH estimation with the unbounded, nonlinearly expanding uncertainty of long-term RUL predictions. Here, we present the Rotary SOH-Injected Prior Battery Transformer (RoSIP-Batt), a unified co-estimation framework that resolves these optimization conflicts. By formulating joint prediction as a Bayesian multi-task objective, RoSIP-Batt introduces a homoscedastic uncertainty weighting mechanism to dynamically scale task-specific gradients based on learned residual noise levels. The architecture leverages decoupled dual classification tokens and a per-dimension gated fusion mechanism, secured by a gradient-detachment operator to prevent high-variance RUL updates from corrupting the stable SOH representation space. To capture electrochemical degradation patterns without relying on absolute cycle steps, Rotary Position Embedding (RoPE) is incorporated into a shared Transformer backbone to model translation-invariant relative temporal profiles. Crucially, the intermediate SOH estimate is directly injected into the RUL regression head as a physical degradation prior. Evaluations across the NASA, MIT-Stanford, and HUST datasets show that RoSIP-Batt significantly outperforms state-of-the-art baselines, reducing SOH estimation error to 1.994% MAE on NASA and restricting RUL prediction error to 62.85 cycles on Stanford. These findings establish RoSIP-Batt as a highly generalizable, computationally efficient solution suitable for real-time embedded BMS deployment.

13:00 JST研究/論文

確率的に断片化された充電プロファイルからのエッジフレンドリーなバッテリー状態診断のための物理ガイド付きマスク マルチタスク ネットワーク

信頼性の高いリチウムイオン電池管理システムの導入は、電動化を加速するために極めて重要ですが、健康状態(SOH)と残存耐用年数(RUL)の共同予後は、タスクの不均一分散性によって依然として大きく妨げられています。従来のマルチタスク学習フレームワークは、SOH 推定の境界のある低分散ノイズと、長期 RUL 予測の境界がなく非線形に拡大する不確実性のバランスを取ることができません。ここでは、これらの最適化の競合を解決する統合共同推定フレームワークである Rotary SOH-Injected Prior Battery Transformer (RoSIP-Batt) を紹介します。 RoSIP-Batt は、同時予測をベイジアン マルチタスク目標として定式化することにより、学習された残留ノイズ レベルに基づいてタスク固有の勾配を動的にスケーリングする等分散不確かさの重み付けメカニズムを導入します。このアーキテクチャは、分離された二重分類トークンと次元ごとのゲート融合メカニズムを活用し、勾配分離演算子によって保護され、高分散 RUL 更新による安定した SOH 表現空間の破損を防ぎます。絶対的なサイクルステップに依存せずに電気化学的劣化パターンを捕捉するために、Rotary Position Embedding (RoPE) が共有の Transformer バックボーンに組み込まれ、並進不変の相対時間プロファイルをモデル化します。重要なのは、中間の SOH 推定値が物理的劣化事前として RUL 回帰ヘッドに直接注入されることです。 NASA、MIT-スタンフォード、および HUST データセットにわたる評価では、RoSIP-Batt が最先端のベースラインを大幅に上回っており、NASA では SOH 推定誤差が 1.994% MAE に減少し、スタンフォードでは RUL 予測誤差が 62.85 サイクルに制限されていることが示されています。これらの発見により、RoSIP-Batt は、リアルタイムの組み込み BMS 展開に適した、汎用性が高く、計算効率の高いソリューションとして確立されます。

原文 (English)

Physics-Guided Masked Multi-Task Network for Edge-Friendly Battery Health Diagnostics from Sto-chastically Fragmented Charging Profiles

The deployment of reliable lithium-ion battery management systems is crucial for accelerating electrification, yet the joint prognosis of State of Health (SOH) and Remaining Useful Life (RUL) remains severely hindered by task heteroscedasticity. Conventional multi-task learning frameworks fail to balance the bounded, low-variance noise of SOH estimation with the unbounded, nonlinearly expanding uncertainty of long-term RUL predictions. Here, we present the Rotary SOH-Injected Prior Battery Transformer (RoSIP-Batt), a unified co-estimation framework that resolves these optimization conflicts. By formulating joint prediction as a Bayesian multi-task objective, RoSIP-Batt introduces a homoscedastic uncertainty weighting mechanism to dynamically scale task-specific gradients based on learned residual noise levels. The architecture leverages decoupled dual classification tokens and a per-dimension gated fusion mechanism, secured by a gradient-detachment operator to prevent high-variance RUL updates from corrupting the stable SOH representation space. To capture electrochemical degradation patterns without relying on absolute cycle steps, Rotary Position Embedding (RoPE) is incorporated into a shared Transformer backbone to model translation-invariant relative temporal profiles. Crucially, the intermediate SOH estimate is directly injected into the RUL regression head as a physical degradation prior. Evaluations across the NASA, MIT-Stanford, and HUST datasets show that RoSIP-Batt significantly outperforms state-of-the-art baselines, reducing SOH estimation error to 1.994% MAE on NASA and restricting RUL prediction error to 62.85 cycles on Stanford. These findings establish RoSIP-Batt as a highly generalizable, computationally efficient solution suitable for real-time embedded BMS deployment.

13:00 JST研究/論文

ChemHyperMag: 物理学に基づいた磁気ハイパーグラフ学習により分子 ADMET 予測が向上

ADMET (吸収、分布、代謝、排泄、毒性) を正確に予測することは創薬にとって重要です。ほとんどの予測子は、無向分子グラフとペアワイズ エッジを使用します。この選択では、非対称相互作用、非可逆ダイナミクス、官能基や環系からのモチーフレベルの効果が見逃されます。ラベルが欠落している場合のマルチタスク ADMET 予測のために ChemHyperMag を提案します。 ChemHyperMag は、環、BRICS フラグメント、Bemis-Murcko 足場、結合から官能基ハイパーグラフを構築します。また、電気陰性度とガスタイガー部分電荷によって導かれる電位駆動の非可逆流れも定義します。結果として生じる循環は、エルミート磁気ラプラシアンによってエンコードされ、磁気チェビシェフ エンコーダーで処理されます。磁気位相を摂動させて確率論的なビューを形成し、InfoNCE の目的でトレーニングします。複数の ADMET ベンチマークでの実験では、標識サンプルが少なく、配座異性体が存在しないため、最近の方法に比べて改善が見られます。 ChemHyperMag はスケーラブルであり、磁気位相を通じて解釈可能な方向性信号を提供します。

原文 (English)

ChemHyperMag: Physics-informed magnetic hypergraph learning improves molecular ADMET prediction

Accurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) is important for drug discovery. Most predictors use undirected molecular graphs and pairwise edges. This choice misses asymmetric interactions, nonreversible dynamics, and motif level effects from functional groups and ring systems. We propose ChemHyperMag for multitask ADMET prediction under missing labels. ChemHyperMag builds a functional group hypergraph from rings, BRICS fragments, Bemis-Murcko scaffolds, and bonds. It also defines a potential driven nonreversible flow guided by electronegativity and Gasteiger partial charges. The resulting circulation is encoded by a Hermitian magnetic Laplacian and processed with a magnetic Chebyshev encoder. We perturb magnetic phases to form stochastic views and train with an InfoNCE objective. Experiments on multiple ADMET benchmarks show improvements over recent methods with fewer labeled samples and no conformers. ChemHyperMag is scalable and provides interpretable directional signals through its magnetic phases.

13:00 JST研究/論文

IBM 量子ハードウェアでの量子暗号解析: 偶数の延長 - マンスール期間の回復が $N=4$ から $N=10$ に

私たちは、実際の IBM 量子ハードウェア (ibm\_kingston、Heron 世代) で実行された対称暗号構造の、コンパイルされていない本物の教科書に忠実な量子暗号解析を報告します。 Simon のアルゴリズムを使用して、実際のハードウェア上でセキュリティ パラメータ N = 10 までの偶数マンスール暗号の隠蔽期間を回復し、以前に報告された最大の実ハードウェア鍵回復である N = 4 を超え、ブロック サイズ 6 および 8 での 3 ラウンドの Feistel (DES ファミリ) 構築の周期をきれいに回復しました。 21 量子ビットのブロック 10 インスタンスがシミュレーションで検証され、ハードウェアに送信されます。さらに、4 つの対称暗号設計パラダイム、バーンスタイン-バジラニ (線形構造、単一クエリ)、グローバー (SPN 鍵検索、二次)、およびサイモン (偶数-マンスール、CBC-MAC 偽造、および Feistel、クエリ複雑さの指数関数から多項式) にわたる 5 つの真の量子攻撃の幅優先ベンチマークを提供し、古典シミュレーションの上限である 25 量子ビットまで検証されています。私たちは範囲について意図的に明示しています。これらの攻撃は、Q2 (量子クエリ) モデルの縮小または構造化された構造をターゲットにしており、誕生日の限界に漸近的に従うため、古典的な衝突発見に対する量子の利点を構成せず、完全な AES/RSA または 16 ラウンド DES を破壊せず、フォールトトレラントなエラー修正ではなくエラー軽減に依存しています。私たちの貢献は、レコード構造サイズでの実際のハードウェアのデモンストレーション、4 つのパラダイムにわたる真のアルゴリズムの範囲の広さ、および公開アーティファクトを使用した誠実で再現可能なベンチマークです。

原文 (English)

Quantum Cryptanalysis on IBM Quantum Hardware: Extending Even--Mansour Period Recovery from $N=4$ to $N=10$

We report genuine-un-compiled, textbook-faithful-quantum cryptanalysis of symmetric-cipher structures executed on real IBM quantum hardware (ibm\_kingston, Heron generation). Using Simon's algorithm we recover the hidden period of the Even-Mansour cipher up to security parameter N = 10 on real hardware, beyond the largest previously reported real-hardware key recovery of N = 4, and we cleanly recover the periods of a 3-round Feistel (DES-family) construction at block sizes 6 and 8; a 21-qubit block-10 instance is verified in simulation and submitted to hardware. We further provide a breadth-first benchmark of five genuine quantum attacks spanning four symmetric-cipher design paradigms -- Bernstein-Vazirani (linear structure, single query), Grover (SPN key search, quadratic), and Simon (Even-Mansour, CBC-MAC forgery, and Feistel; exponential-to-polynomial in query complexity) -- validated to the classical-simulation ceiling of 25 qubits. We are deliberately explicit about scope: these attacks target reduced or structured constructions in the Q2 (quantum-query) model, asymptotically follow the birthday bound and therefore do not constitute quantum advantage over classical collision-finding, do not break full AES/RSA or 16-round DES, and rely on error mitigation rather than fault-tolerant error correction. Our contribution is the real-hardware demonstration at record structure sizes, the breadth of genuine algorithmic coverage across four paradigms, and an honest, reproducible benchmark with public artifacts.

13:00 JST研究/論文

PRISM: 効率的なニューラル ネットワーク暗号化のための感度を意識した多項式 PRuning

構造化枝刈りは、準同型暗号化 (HE) の下でニューラル ネットワーク推論を実行可能にするために不可欠ですが、モデルの信頼性に対するその影響はまだ解明されていません。この論文では、枝刈りされた CKKS 暗号化ニューラル ネットワークの体系的な信頼性の特性評価を示し、本質的に信頼性を意識した構造化枝刈り手法である多項式感度認識枝刈り (PSAP) を紹介します。 PSAP は、重みの大きさ、多項式アクティベーションの感度、およびローテーション コストを組み合わせてフィルターをスコア付けし、フォールト トレラント領域でプルーニングを集中させます。 2 つのアーキテクチャ、2 つのデータセット、2 つの数値表現、および 5 つのビット エラー率 (フルモデルで 40 件、層ごとの実験で 108 件) にわたって、PSAP プルーニング モデルでは、壊滅的な (10 pp を超える精度低下) レイヤーを最大 2 層に制限するのに対し、マグニチュード プルーニングされたベースラインでは 5 ~ 14 層を制限し、int32 ビット フリップ インジェクションの下で最悪の場合の脆弱性を最大 29 分の 1 に削減します。直接の CKKS 暗号化フォールト インジェクションは、BER~ 10^{-5} 付近の安全な動作境界を示し、保守的な信頼性プロキシとして int32 インジェクションをサポートします。フォールト クリティカルな構造層はパラメーターの 1.1% のみを占めるため、最小限のオーバーヘッドで選択的な強化が可能になります。これらの信頼性の向上は、競合効率とともに得られます。PSAP は、ResNet-32 でのハレヴィ-シャウプ回転を最大 45.2\% 削減し、適応混合次数割り当てスキームにより乗算の深さが 66 レベルから 56 レベルに低下し、ブートストラップなしで平準化された推論が可能になります。

原文 (English)

PRISM: Sensitivity-Aware PolynoMial PRuning for EffIcient Neural Network Encryption

Structured pruning is essential for making neural network inference feasible under homomorphic encryption (HE), yet its impact on model reliability has remained unexplored. This paper presents a systematic reliability characterization of pruned CKKS-encrypted neural networks and introduces Polynomial-Sensitivity-Aware Pruning (PSAP), a structured pruning method that is inherently reliability-aware. PSAP scores filters jointly by weight magnitude, polynomial activation sensitivity, and rotation cost, which concentrates pruning in fault-tolerant regions. Across two architectures, two datasets, two numerical representations, and five bit-error rates (40 full-model and 108 per-layer experiments), PSAP-pruned models limit catastrophic (>10 pp accuracy drop) layers to at most two versus 5--14 for magnitude-pruned baselines, reducing worst-case vulnerability by up to 29 times under int32 bit-flip injection. Direct CKKS encrypted fault injection indicates a safe operating boundary near BER~ 10^{-5}, supporting int32 injection as a conservative reliability proxy. The fault-critical structural layers account for only 1.1% of parameters, enabling selective hardening at minimal overhead. These reliability gains are obtained alongside competitive efficiency: PSAP reduces Halevi--Shoup rotations by up to 45.2\% on ResNet-32, and an adaptive mixed-degree allocation scheme lowers multiplicative depth from 66 to 56 levels, enabling leveled inference without bootstrapping.

13:00 JST研究/論文

フェデレーテッド軽量微調整

フェデレーテッド微調整は通信によってボトルネックになっています。FedAvg と擬似勾配スキームはモデルに合わせてスケールするペイロードを送信し、勾配圧縮は一定の係数だけ圧縮します。別のレバーを使います。マッピング ネットワークは、フリーズ アフィン射影を通じて小さなトレーニング可能な潜在要素からネットワークの重みを生成します。マップは共有されアフィンされているため、潜在値の平均化は、生成された重みの正確な平均化となります。これを 2 つの変更を加えて実用的な低帯域幅のフェデレーテッド チャネルに変えます。1 つは低ランクのシード再生成可能な投影の因数分解 (ジェネレーター メモリを ~80 GB から ~10 MB に削減)、もう 1 つはデルタ定式化 $\theta = \theta^{\mathrm{pre}} + U V^{\top} z$ です。これは、中央で事前トレーニングされた共有ベース、つまりフェデレーテッド微調整に基づいて加算補正を学習します。メソッドは大規模に機能します。凍結された直交分級器ヘッドは、精度を向上させながらペイロードからヘッドをさらに除去します。 ResNet-18+GroupNorm を備えた CIFAR-100 では、私たちのメソッド (FLITE、Federated Low-rank Iterative Training Engine) は、ラウンドごとにクライアントごとに 1,280 フロート (約 5 KB) を通信し、8718 倍の削減となり、フルウェイト FedAvg の約 0.5 pp 以内で 74.67% に達します。平均化恒等式は浮動小数点精度 ($6 \times 10^{-8}$) に保持されます。この方法は、PowerSGD よりも 1 ~ 2 桁低く、帯域幅精度のパレートで上位 k に位置します。強い非 IID スキューの下では、フルウェイト FedAvg と一致するか、それを上回ります。 int4 潜在は、変わらない精度でラウンドあたり 648 バイトに達しますが、int4 のフルウェイト FedAvg は偶然に崩壊します。

原文 (English)

Federated Lightweight Fine-Tuning

Federated fine-tuning is bottlenecked by communication: FedAvg and pseudo-gradient schemes transmit a payload that scales with the model, and gradient compression shrinks it by only a constant factor. We take a different lever. Mapping networks generate a network's weights from a small trainable latent through a frozen affine projection; because the map is shared and affine, averaging latents is exactly averaging the generated weights. We turn this into a practical low-bandwidth federated channel with two changes: a low-rank, seed-regenerable factorisation of the projection (cutting generator memory from ~80 GB to ~10 MB), and a delta formulation $\theta = \theta^{\mathrm{pre}} + U V^{\top} z$ that learns an additive correction around a shared centrally-pretrained base -- federated fine-tuning, which is what makes the method work at scale. A frozen orthogonal classifier head further removes the head from the payload while improving accuracy. On CIFAR-100 with ResNet-18+GroupNorm, our method (FLITE, Federated Low-rank Iterative Training Engine) communicates 1,280 floats (~5 KB) per client per round -- an 8718x reduction -- and reaches 74.67%, within ~0.5 pp of full-weight FedAvg. The averaging identity holds to floating-point precision ($6 \times 10^{-8}$); the method sits one to two orders of magnitude below PowerSGD and top-k on the bandwidth-accuracy Pareto; it matches or exceeds full-weight FedAvg under strong non-IID skew. int4 latents reach 648 bytes per round at unchanged accuracy, whereas int4 full-weight FedAvg collapses to chance.

13:00 JST研究/論文

FSDBN: Foreground-Aware EEG--Visual Alignment via Dynamic Brain Networks

EEG-based visual decoding provides a non-invasive pathway for interpreting visual semantics. However, existing methods often overlook the p…

13:00 JST研究/論文

Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative Models

Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs). AAD uses electroencephal…

13:00 JST研究/論文

An Analysis of Residual-Stream Geometry Across Transformer Depth

We propose a transition-centred geometric analysis of transformer residual streams. Relative displacement measures how \emph{far} represent…

13:00 JST研究/論文

MambaLSTM: A Spatio-Temporal Framework for Enhanced Traffic Accident Risk Prediction

In traffic accident risk prediction, most studies overlook the extra noise that could be incorporated when fusing temporal features into sp…

13:00 JST研究/論文

Multi-layer MIMO Relay as Deep Physical Neural Networks: Power Amplifiers as Activation Functions

Wireless physical neural networks (WPNNs) embed neural computation directly into analog hardware, offering lower energy consumption and lat…

13:00 JST研究/論文

CODENS: Transforming Code Changes into Living, Accessible, and Queryable Documentation

Maintaining up-to-date code documentation is difficult in fast-moving repositories because design knowledge is scattered across source file…

13:00 JSTLLM/生成AIエージェント

Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragments

Large language models now write a growing share of the world's code, increasingly inside agents and serving systems that compile, execute,…

13:00 JSTLLM/生成AI

HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers

Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricat…

13:00 JST研究/論文

Physical Self-Supervised Learning: IMU Sensing without Manual Labels

Deep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly label…

13:00 JSTLLM/生成AI

A Controlled Study of Attention-Only Transformers

Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test t…

13:00 JST研究/論文

Adversarial Robustness of Phishing Email Detection: A Comparative Study of TF-IDF + Logistic Regression and Fine-Tuned DistilBERT

Phishing emails remain one of the most persistent cybersecurity threats, and machine-learning classifiers are widely used to detect them. M…

13:00 JST研究/論文

Intelligence from Learnable Novelty

Intelligence appears under different names in different fields: as data compression in statistics and machine learning, as universal comput…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from…

13:00 JSTLLM/生成AI規制/政策

ChainMark: Model-Free LLM Watermarking with Closed-Form Calibration

Regulatory regimes such as the EU AI Act mandate machine-readable marking of synthetic text, but existing watermark detectors rely on the g…

13:00 JSTLLM/生成AI画像/動画生成

CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders

Frozen encoders are chosen by how well a lightweight head reads a finding from their features, not whether the geometry separates it. Neare…

13:00 JSTビジネス/資金調達

Estimating Rare Events in Language Models with Proper Evaluation

Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale de…

13:00 JST研究/論文

Competitive and Complementary Tools

Humans have always externalized thought onto tools, from the tally and the abacus to the map and, now, large language models. I model the a…

13:00 JST研究/論文

RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts

Group Relative Policy Optimization (GRPO) has shown strong effectiveness in reinforcement learning from verifiable feedback, where sampled…

13:00 JSTLLM/生成AIハードウェア/半導体Claude

Structured Output Collapses Answer Diversity Across 44 Language Models

When a language model must choose one answer from a large space of equally valid options, a format clause -- "Reply with JSON only" -- chan…

13:00 JST研究/論文

Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft

The digital substrate of states -- data, algorithms, infrastructure, platforms, applications -- is being governed without adequate conceptu…

13:00 JSTLLM/生成AIエージェント研究/論文

Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing

Large language model (LLM) agents are starting to take on routine work in high-performance computing (HPC), including monitoring Slurm jobs…

13:00 JSTロボティクス研究/論文

The Open Ant: A Robot Platform for Reinforcement Learning Research

Reinforcement learning (RL) research has demonstrated success in both physical and simulated domains; however, the predominant methodology…

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

Towards an Automated Test of LLM Security Knowledge

Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM p…

13:00 JST画像/動画生成研究/論文

Now We Know? A Systematic Comparison of TerraMind and THOR

Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ…

13:00 JST研究/論文GPT / ChatGPTGemini

Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists

Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind…

13:00 JST研究/論文

Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling

This study contributes toward development of an Automated Data Processing (ADP) framework designed to evaluate and reinforce optimal machin…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Ex…

13:00 JST研究/論文

Censoring-Aware In-Context Learning for Generalized Supplier Lead Time Estimation in Supply Chain Planning

Supplier lead time forecasting is a central input to material requirements planning, inventory optimization, and supply chain risk manageme…

13:00 JST研究/論文

Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes? We t…

13:00 JSTLLM/生成AIエージェント

The Story Shapes the Agent: Narrative Priors in LLM Behavior

Persona prompting is widely used to steer LLM agent behavior, yet the narrative framing of a task can matter more than the assigned persona…

13:00 JSTLLM/生成AILlamaMistral AI

For What Reason? Interpreting Models' Encoding of Causation and Antithesis

Discourse relations provide document structure, critical to language understanding and enabling language model performance and ethicality.…

13:00 JST研究/論文

Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States

Reinforcement learning is conventionally divided into model-based and model-free methods. In this taxonomy, model-based methods perform loo…

13:00 JSTLLM/生成AI

AutoIndex: Learning Representation Programs for Retrieval

We present AutoIndex, a framework for learning representation programs: executable transformations that map raw documents into the represen…

13:00 JSTLLM/生成AIロボティクス

Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM Approach

The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordination of physical flight kine…

13:00 JST研究/論文

Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation

The Matthew effect is a big challenge in Recommender Systems (RSs), where popular items tend to receive increasing attention, while less po…

13:00 JSTLLM/生成AI

LatentMT: Machine Translation with Latent Reasoning

Latent-reasoning looped language models (LoopLMs) offer a different scaling path for machine translation (MT): instead of increasing parame…

13:00 JST研究/論文

Temporal-Causal Unity as an Operational Framework for Collective Dynamics: Causal-Progress Clocks, Synchronization, and Polarization

This paper develops temporal-causal unity (TCU), a framework connecting a process-philosophical thesis -- time is the ordered unfolding of…

13:00 JSTLLM/生成AI

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

Textual Collaborative Prompt Optimization (TCPO) extends Textgrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multipl…

13:00 JST画像/動画生成

Norm or Direction? Decoding Vision Mambas for High-Resolution Vision

Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visu…

13:00 JST画像/動画生成

Deep Learning Estimation of Sex, Age, Height, and Weight from CT-derived Digitally Reconstructed Radiographs

Purpose: To develop and validate a deep learning ensemble for estimating adult sex, age, height, and weight from coronal digitally reconstr…

13:00 JSTLLM/生成AIエージェント

Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents

LLM-based browser agents are rapidly changing the threat landscape for web security. Unlike traditional automation frameworks that execute…

13:00 JSTLLM/生成AI画像/動画生成

Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models

A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP wi…

13:00 JST画像/動画生成

Decoupled Pipeline with Proposal Reranking and Score Fusion for Positive-Unlabeled Marine Species Detection

The FathomNetCLEF 2026 competition combines underwater object detection and fine-grained marine species classification under a positive-unl…

13:00 JST研究/論文

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

Caption Studio is a transparency-first speech and audio intelligence platform that transforms spoken audio and video into structured, searc…

13:00 JSTエージェント

Strategy-Following Multi-Agent Deep Reinforcement Learning Considering Control Strategies Provided to Other Agents

This study proposes a learning method for multi-agent systems that allows agents to be controlled through human manager instructions after…

13:00 JSTLLM/生成AI

Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA

Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choosing which small model to ada…

13:00 JST研究/論文

ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series

This paper proposes ConceptCF, a method for counterfactual generation that operates on human-interpretable concepts. In high-stakes domains…

13:00 JSTLLM/生成AI画像/動画生成

Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks

The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms of privacy, cost, and scalabi…

13:00 JSTLLM/生成AIエージェント

AgentTrails: Towards Trust and Reuse for Agentic Tasks

LLM-powered agents increasingly tackle complex tasks by invoking tools, querying databases, executing code, and manipulating intermediate a…

13:00 JSTLLM/生成AI

AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System

This comprehensive study introduces an advanced Artificial Intelligence for Indian Legal Question Answering (AILQA) system tailored to the…

13:00 JSTLLM/生成AIエージェント

Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents

LLM-agent defenses are typically evaluated one session at a time. In deployment, however, attacks can be distributed across independent age…

13:00 JST研究/論文

From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning

Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning…

13:00 JST研究/論文

ABOPD: Antibody CDR Design via On-Policy Distillation

Antibodies are essential therapeutic molecules, and their complementarity-determining regions (CDRs) form the primary antigen-recognition i…

13:00 JSTLLM/生成AIエージェント

Data Leakage Prevention in Agentic Applications via Preemptive Hardening

Agentic systems integrate LLM driven planning with interfaces to external tools, making data leakage and tool misuse feasible via instructi…

13:00 JST画像/動画生成

OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation

Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level a…

13:00 JST研究/論文

Regime-Aware Physics-Guided Early Warning of Lithium-Ion Battery Thermal Runaway Using Thermo-Mechanical Signals

Thermal runaway in lithium-ion batteries poses a major safety risk to electric vehicles and energy storage systems. Current early-warning m…

13:00 JST研究/論文

RAMP: Recognition parametrisation by Amortised Message Passing

A central aim of unsupervised learning is to uncover latent factors that explain dependencies among observations. Probabilistic models typi…

13:00 JST研究/論文

Public perceptions of AI-driven decision-making in healthcare: A structural equation modeling approach

Artificial intelligence (AI) is increasingly integrated into healthcare to support diagnostics, decision-making, and administrative process…

13:00 JST研究/論文

Circuit Claims Depend on What Is Extracted and How It Is Compared

Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting…

13:00 JST研究/論文

Functional Equivalence and Geometric Diversity in Neural Network Approximations: An Empirical Characterization

The Universal Approximation Theorem states that a neural network with a single hidden layer is sufficient to approximate any continuous uni…

13:00 JSTLLM/生成AI画像/動画生成GPT / ChatGPT

Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model

While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs…

13:00 JST研究/論文

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement

Procuring supervised fine-tuning (SFT) data forces a buyer to decide, before any downstream training, whether a candidate corpus is worth a…

13:00 JST研究/論文

Variational meta-learning inference for low dimensional neural system identification

Deep learning has proven highly effective for nonlinear system identification, but heavily parameterized neural networks are prone to overf…

13:00 JSTエージェント

Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts

Agent Skills have become persistent behavioral artifacts across independent AI agent systems. They combine natural-language task specificat…

13:00 JSTLLM/生成AIエージェント

Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction

Textual skills provide a lightweight way to improve frozen language-model agents, but their self-evolution normally requires a stable valid…

13:00 JSTLLM/生成AIビジネス/資金調達

AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism

We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation using large language models (L…

13:00 JST画像/動画生成

SWITi: Quantifying and Reducing Tiling Artifacts with Sliding Window Inner Tiling

SWITi is a test-time method for reducing artifacts in tiled predictions, particularly for neural networks that learn posterior distribution…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

Multi-turn medical consultation agents must decide what to ask, adapt to patient responses, and determine when the collected evidence is su…

13:00 JSTLLM/生成AIビジネス/資金調達

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms…

13:00 JST研究/論文

Biological Amnesia in ICU Time-Series Prediction: A Drift-Adaptive Two-Stream Architecture with Temporal Retrieval

Background: Clinical decision support systems degrade silently as treatment protocols evolve, yet standard adaptation methods treat models…

13:00 JST画像/動画生成

CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement

V2X collaborative object detection features overcoming the limitations of single-vehicle systems by aggregating environmental features from…

13:00 JST画像/動画生成エージェント

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary pro…

13:00 JST研究/論文

Spectral Higher-Order Neural Networks Have Sharp Expressivity Bounds

Neural hypergraphs are a natural generalization of neural networks, the reference models in modern machine learning. Yet, their deployment…

13:00 JST研究/論文

Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B-parameter MoE languag…

13:00 JSTロボティクス

Deep learning-based prediction of time-resolved adhesive forces in viscoelastic Hertzian contacts

Fast prediction of the response of adhesive soft viscoelastic contacts represents a current challenge in soft robotics and for gripping and…

13:00 JST画像/動画生成

Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions

Hateful optical illusions expose a serious gap in current multimodal safety systems. On original-view hateful illusions, previous work show…

13:00 JST画像/動画生成NVIDIA

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-sca…

13:00 JST研究/論文

From Operations to Elderly Care Outcomes: A Thematic Review of Industrial Engineering and Decision-Support Approaches

The rapid growth of the global aging population presents severe challenges to healthcare systems, necessitating efficient, equitable, and p…

13:00 JSTLLM/生成AIQwen

DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning

Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence…

13:00 JSTLLM/生成AI研究/論文

SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation

Large language models (LLMs) excel at general-purpose code generation, yet how well they handle scientific code remains an open question. E…

13:00 JST研究/論文

Code Division Modulation Layers Against Forgetting and Inference in Continual Gait Identification

Continual learning (CL) has been recently employed in biometric identification systems thanks to its ability to integrate new knowledge wit…

13:00 JST研究/論文

Parallel Noising in Neural Markov Logic Networks

Neural Markov Logic Networks (NMLNs) are a flexible neurosymbolic relational model. Previous work has shown that, although NMLNs achieve st…

13:00 JST研究/論文

MIRAGE: Multi-scale Lesion-Informed Representation with Auxiliary Guidance for MRI Contrast Enhancement

Inferring contrast enhancement from one pre-contrast breast MRI slice is underdetermined: post-contrast appearance contains physiological i…

13:00 JST研究/論文

Incomplete Observations Boost Evolutionary Performance in Ocean Modeling

Data-driven methods have revolutionized ocean modeling, yet current approaches rely heavily on complete reanalysis datasets, imposing compu…

13:00 JST研究/論文

Breaking the Homogeneity Assumption: Specialized Multi-Generator Adversarial Learning for Rare Failure Detection in Predictive Maintenance

Supervised learning models in the predictive maintenance field are regularly trained on highly imbalanced industrial datasets: machine fail…

13:00 JSTLLM/生成AIGemma

Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning

Neural machine translation (NMT) in the legal domain is a linguistically and conceptually demanding task, primarily due to the complexity o…

13:00 JSTエージェントロボティクス

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a…

13:00 JST画像/動画生成ハードウェア/半導体NVIDIA

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-…

13:00 JST研究/論文

Free energy landscape of Dense Associative Memory

Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative mem…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams

Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot…

13:00 JST研究/論文GPT / ChatGPT

Assessment in Team Problem-Solving Exercises in Computing Education

This full paper in the research-to-practice track presents methods for assessing student teams in tabletop exercises (TTXs). TTXs enable le…

13:00 JSTロボティクス

Computing on the Fly: Navigating a Vision for the Future of Drone Computing

The report envisions a decade in which drones move goods, medical supplies, and information at a scale comparable to national infrastructur…

13:00 JSTLLM/生成AIGPT / ChatGPT

Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards

Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, ex…

13:00 JSTLLM/生成AI

The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation

Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Mod…

13:00 JSTLLM/生成AIGemma

Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs

Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain dispropo…

13:00 JSTLLM/生成AI

Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models

Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format instructions and context (ma…

13:00 JSTビジネス/資金調達研究/論文

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and…

13:00 JST画像/動画生成エージェント研究/論文

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrati…

13:00 JSTエージェント

Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case Investigation

Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable. We study a layered pipeline o…

13:00 JSTエージェント

They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface

We study a five-agent CI/CD pipeline (triage -> developer -> security-scan -> review -> approve/deploy), built from five distinct productio…

13:00 JST研究/論文

GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models

The Traffic Assignment Problem is a fundamental but computationally expensive component of transportation planning. While Graph Neural Netw…

13:00 JST研究/論文

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetica…

13:00 JSTハードウェア/半導体

Riemannian Deep Learning:Modules, Networks, and Geometries

Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific…

13:00 JST画像/動画生成エージェントロボティクス

From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-Accelerated Motion Planning for UAVs

Autonomous flight in cluttered environments requires a robot to build a geometric map of its surroundings and plan safe, dynamically feasib…

13:00 JST研究/論文

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on dif…

13:00 JST研究/論文

ISO: An RLVR-Native Optimization Stack

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimizat…

13:00 JST研究/論文

Provable diffusion-based posterior sampling for linear inverse problems via DDIM

Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existing posterior samplers e…

13:00 JST画像/動画生成

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, ob…

13:00 JSTLLM/生成AI

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to…

13:00 JSTエージェント研究/論文

FormGym: Doing Paperwork with Agents

Completing paperwork is a challenging and time-consuming problem. Form filling is especially challenging in the pure-image domain without a…

13:00 JSTエージェント

Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents

Graphical User Interface (GUI) agents have made significant progress in automating digital tasks through the utilization of computer vision…

13:00 JSTエージェントロボティクス研究/論文

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics

As embodied autonomous systems capable of assisting humans in daily activities remain a major goal for robotics, efficient and appropriate…

13:00 JSTエージェントビジネス/資金調達

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents. SENTINEL is th…

13:00 JSTLLM/生成AIエージェント

Learning to Make Friends: Coaching LLM Agents toward Emergent Social Ties

Can large language model (LLM) agents reproduce the complex social dynamics that characterize human online behavior -- shaped by homophily,…

13:00 JSTエージェント

Dr. Zero: Self-Evolving Search Agents without Training Data

As high-quality data becomes increasingly difficult to obtain, self-evolution without curated training data has emerged as a promising para…

13:00 JST研究/論文

Fluid Reasoning Representations

Frontier large language models increasingly solve complex tasks involving abstract concepts through extended test-time thinking. Yet we lac…

13:00 JSTLLM/生成AI

LLM-Grounded Explainable AI for Supply Chain Risk Early Warning via Temporal Graph Attention Networks

Disruptions at critical logistics nodes pose severe risks to global supply chains, yet existing risk prediction systems typically prioritiz…

13:00 JSTLLM/生成AI

Animating Petascale Time-varying Data on Commodity Hardware with LLM-assisted Scripting

Scientists face significant visualization challenges as time-varying datasets grow in speed and volume, often requiring specialized infrast…

13:00 JST研究/論文

Participatory provenance as representational auditing for AI-mediated public consultation

AI-assisted consultation can speed large-scale public engagement, but concise summaries may reflect some submissions more closely than othe…

13:00 JSTLLM/生成AIGPT / ChatGPT

FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking

Large language models (LLMs) are rapidly being adopted across various domains. However, their adoption in banking industry faces resistance…

13:00 JSTLLM/生成AIエージェントAnthropicOpenAI

Frontier LLM ベースのエージェントは、自然な表現型のオントロジーキュレーションのボトルネックを克服できます

フリーテキストの表現型記述をオントロジー用語にリンクすることは、通常表現型アノテーションと呼ばれ、比較形態学的データを研究間で統合するために不可欠です。この労働集約的なプロセスは高度な訓練を受けた人間の専門家に大きく依存しており、そのため拡張が困難であり、それが大きなボトルネックとなっています。ダードゥルら。 (2018) 7 つの系統学的研究にわたるエンティティ品質 (EQ) アノテーションのゴールド スタンダード (GS) を確立し、それを使用して 3 人のキュレーターと、オントロジーベースの意味的類似性メトリクスを備えた Semantic CharaParser NLP ツールを評価しました。彼らは、機械と人間の一貫性は、キュレーター間(人間と人間)の一貫性よりも大幅に低いと報告しました。ここでは、Anthropic と OpenAI の 5 つのフロンティア ホスト LLM を使用してそのベンチマークを再検討します。各 LLM は、ソース出版物の PDF、元の人間のキュレーターが使用したのと同じ注釈ガイド、4 つのプロジェクト オントロジー (UBERON、PATO、BSPO、GO)、および検証スクリプトを提供する自己完結型ワークスペース内で「エージェント キュレーター」として動作します。同じゴールドスタンダードに照らして評価すると、すべてのエージェントは、元の研究で訓練を受けた 3 人の人間のバイオキュレーターのキュレーター間変動の範囲内に収まりました。最もパフォーマンスの高いエージェントがアプローチしましたが、最もパフォーマンスの高い人間のキュレーターには到達できませんでした。エージェントは、4 つの指標すべてで Semantic CharaParser を大幅に上回りました。

原文 (English)

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study integration of comparative morphological data. This labor intensive process has heavily relied on highly trained human experts, which makes it challenging to scale and thus a key bottleneck. Dahdul et al. (2018) established a Gold Standard (GS) of Entity-Quality (EQ) annotations across seven phylogenetic studies and used it to evaluate three human curators and the Semantic CharaParser NLP tool with ontology-based semantic similarity metrics; they reported that machine-human consistency was significantly lower than inter-curator (human-human) consistency. Here we revisit that benchmark with five frontier hosted LLMs from Anthropic and OpenAI, each operating as an "agentic curator" within a self-contained workspace that supplies the source publication PDF, the same annotation guide used by the original human curators, the four project ontologies (UBERON, PATO, BSPO, GO), and a validation script. Evaluated against the same Gold Standard, every agent fell within the range of inter-curator variability of the three trained human biocurators of the original study; the best performing agents approached but did not reach the best performing human curator. Agents substantially outperformed Semantic CharaParser on all four metrics.

13:00 JSTLLM/生成AIエージェントOpenAI

AgentJet: エージェント強化学習のための柔軟な群トレーニング フレームワーク

大規模言語モデル (LLM) エージェント強化学習用の分散群トレーニング フレームワークである AgentJet を紹介します。エージェントのロールアウトとモデルの最適化を密接に結び付ける集中型フレームワークとは異なり、AgentJet は分離されたマルチノード アーキテクチャを採用しています。このアーキテクチャでは、swarm サーバー ノードがトレーニング可能なモデルをホストし、GPU クラスターで最適化を実行します。一方、swarm クライアント ノードは任意のデバイスで任意のエージェントを実行します。この設計は、集中型フレームワークではサポートが難しい機能を提供します。(1) 異種マルチモデル強化学習。複数の LLM を頭脳とする異種マルチエージェント チームのトレーニングを可能にします。 (2) 独立したエージェントのランタイムを使用したマルチタスクのカクテル トレーニング。 (3) 外部環境の障害によるトレーニング プロセスの中断を防ぐフォールト トレラントな実行。 (4) ライブ コードの反復。群クライアント ノードを置き換えることにより、トレーニング中にエージェントを編集できます。マルチモデル、マルチターン、マルチエージェント設定で効率的な RL をサポートするために、AgentJet はタイムライン マージを備えたコンテキスト トラッキング モジュールを導入しています。これにより、冗長なコンテキストが統合され、トレーニングの 1.5 ~ 10 倍の高速化が実現します。最後に、AgentJet は、研究トピックを入力として受け取り、大規模クラスター上で長期にわたる複数日にわたる RL 研究を自律的に実行する自動研究システムを導入します。このシステムは、swarm アーキテクチャを活用することで、実行中に人間の介入なしに、RL 研究者の主要な探索ワークフローを再現します。

原文 (English)

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning

Training reinforcement learning (RL) policies for large language model (LLM) agents requires optimizing multi-turn trajectories that interact with external environments. Existing training frameworks struggle with runtime failures, single-model constraints, incompatible task environments, and redundant context. We present AgentJet, a distributed swarm training framework based on a decoupled multi-node architecture. AgentJet treats the server--client topology as configurable: swarm servers host trainable models and perform optimization on GPU clusters, while detachable swarm clients execute arbitrary agents and communicate through OpenAI-compatible APIs. Reconfiguring this topology supports heterogeneous multi-model RL, mixed-task training with isolated runtimes, fault-tolerant execution, and live code iteration through hot-swappable clients. AgentJet also introduces context tracking with timeline merging, reducing actor-update time by 6.25x on AppWorld. The same detachable-client design supports an automated research system that conducts long-horizon, multi-day RL studies on large-scale clusters with limited human intervention. AgentJet is open-source and compatible with agent systems that issue standard LLM inference requests.

13:00 JSTLLM/生成AI画像/動画生成

数学的推論のための人工知能: 言語モデル、神経記号システム、および検証された発見の統合的調査

数学的推論は長い間、機械知能の厳しいテストとして機能してきました。過去 10 年間で、NLP 内のニッチな問題から、最も重要な AI フロンティアの 1 つに移行しました。この調査は、初期のルールベースの数学文章問題 (MWP) ソルバーとテンプレート駆動の幾何学システムから、神経式生成と LLM プロンプトを経て、現代の推論モデル、マルチエージェント システム、神経記号定理証明者、および検証済みの発見ワークフローに至るまで、この分野の進化に関する統一的な説明を提供します。私たちは 4 つの軸に沿ってランドスケープを整理します。(i) MWP 解決、マルチモーダル ジオメトリ、および VLM にわたる、テキストと図に関する非形式的な推論。 (ii) 自動形式化、戦術予測、コンパイラー主導の修復、および証明検索を含む、証明アシスタントにおける形式的推論。 (iii) 数学的発見。システムが構築を提案し、境界を改善し、未解決の問題への攻撃を支援します。 (iv) CoT プロンプト、ツールの使用、プロセス報酬モデル、RLVR など、生成と検証をますます結び付ける推論およびトレーニング時の手法。私たちは、小学校の算数、競技数学、幾何学、形式的証明、マルチモーダルおよび多言語推論、専門家の評価にわたる主要なベンチマークをカタログ化し、ベンチマークの飽和、汚染、レポートの不一致、および pass@1、多数決、検証者支援 pass@$k$ の区別を調べます。私たちは、摂動下での脆弱性、報酬ハッキング、マルチモーダル接地障害、脆弱な形式化、推論規模の推論のエネルギーコストなどの障害モードを批判的に評価します。現役の数学者からの最近の視点を活用して、検証された発見のワークフロー、推論の効率、AI 支援による形式化を広く利用できるようにするインフラストラクチャを中心とした将来の方向性を特定します。関連資料: https://github.com/Starscream-11813/awesome-AI4Math。

原文 (English)

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery

Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem within NLP to one of the most consequential AI frontiers. This survey provides a unified account of the field's evolution, from early rule-based math word problem (MWP) solvers and template-driven geometry systems, through neural expression generation and LLM prompting, to contemporary reasoning models, multi-agent systems, neuro-symbolic theorem provers, and verified discovery workflows. We organize the landscape along four axes: (i) informal reasoning over text and diagrams, spanning MWP solving, multimodal geometry, and VLMs; (ii) formal reasoning in proof assistants, including autoformalization, tactic prediction, compiler-guided repair, and proof search; (iii) mathematical discovery, where systems propose constructions, improve bounds, or assist attacks on open problems; and (iv) the inference and training-time techniques, including CoT prompting, tool use, process reward models, and RLVR, that increasingly connect generation with verification. We catalog major benchmarks across grade-school arithmetic, competition mathematics, geometry, formal proving, multimodal and multilingual reasoning, and expert evaluation, and we examine benchmark saturation, contamination, reporting mismatches, and the distinction between pass@1, majority voting, and verifier-assisted pass@$k$. We critically assess failure modes: brittleness under perturbation, reward hacking, multimodal grounding failures, fragile formalization, and the energy cost of reasoning-scale inference. Drawing on recent perspectives from working mathematicians, we identify future directions centered on verified-discovery workflows, reasoning efficiency, and infrastructure to make AI-assisted formalization broadly usable. Companion materials: https://github.com/Starscream-11813/awesome-AI4Math.

13:00 JST研究/論文

心の効用の理論: 精神化メカニズムの正式な仕様

他人の信念を推測するには、表面的な信号を読み取るだけでは不十分です。誰が何を、どの順序で、どの程度信頼できるように伝えたかを追跡する必要があります。 Theory of Mind Utility (ToM-U) は、この認識論的状態推論問題を解析の計算レベルで形式化し、アルゴリズムやニューラル実装にこだわることなく、メンタライジングが何を計算するのか、またなぜ計算するのかを指定します。 ToM-Uは、ローカル認識世界モデル(LEWM)(エージェント、状態ノード、およびそれらの間の認識関係を表す有向型グラフ)を構築し、十分な信頼が得られるまで観察された動作に対して離散候補LEWMを評価することによってこれを実現します。 5 つの正式な定義は、LEWM 構造、順序付けされた情報アクセス履歴を含むエージェント ノード プロパティ、再帰的メンタライジングのための制限された増殖メカニズム、3 つの推論手順、失敗したメンタライジングの試みによって残された構造化された痕跡を捕捉する残余関数を指定します。 ToM-U は、信念状態を導き出すのではなく前提とするベイズ精神理論や隣接する形式的説明、認識状態推論のための形式的な装置を欠くシミュレーション理論や理論理論とは異なります。このアーキテクチャは、補助的な仮定ではなくモデルの構造的特性に基づいてメンタライゼーションの失敗に関する方向性のある反証可能な予測を生成し、ToM-U を目標推論やその他の下流の社会的認知プロセスの上流にある領域に依存しないメカニズムとして位置づけます。

原文 (English)

The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism

Inferring others' beliefs requires more than reading surface signals; it requires tracking who told them what, in what order, and how credibly. The Theory of Mind Utility (ToM-U) formalizes this epistemic state inference problem at the computational level of analysis, specifying what mentalizing computes and why without commitment to algorithmic or neural implementation. ToM-U achieves this by constructing Local Epistemic World Models (LEWMs) -- directed typed graphs that represent agents, state nodes, and the epistemic relationships among them -- and evaluating discrete candidate LEWMs against observed behavior until one achieves sufficient confidence. Five formal definitions specify the LEWM structure, agent node properties including ordered information access history, a bounded proliferation mechanism for recursive mentalizing, three inference procedures, and a residue function that captures the structured trace left by failed mentalizing attempts. ToM-U differs from Bayesian Theory of Mind and adjacent formal accounts, which presuppose rather than derive belief states, and from simulation theory and theory-theory, which lack a formal apparatus for epistemic state inference. The architecture generates directional, falsifiable predictions about mentalizing failure that follow from structural properties of the model rather than auxiliary assumptions, and positions ToM-U as a domain-agnostic mechanism upstream of goal inference and other downstream social cognitive processes.

13:00 JST研究/論文

研究ハーネスを通じて AI 科学者の研究総合と検証を外部化する

AI システムは科学ワークフローをますます自動化することができますが、以前の証拠、生成されたアイデア、実験、最終的な主張を結び付ける推論は、多くの場合、モデル推論内に暗黙的に残ります。ここでは、研究の総合と実験の検証を検査可能な契約管理されたプロセスに外部化する研究ハーネスである Xcientist を紹介します。 Xcientist は、文献証拠、アイデアの状態、実装計画、アブレーション記録、修理痕跡を永続的な研究成果物として整理し、生成されたメカニズムをその証拠的根拠を失うことなく基礎付け、実行、テスト、修正できるようにします。私たちは、実行可能なアーティファクトが当初主張されていたメカニズムをサポートしなくなった、自動化された研究の失敗モードとしてクレーム ドリフトを特定します。 Xcientist は、トレーニング不要のメモリ システム、グラフ構造のトラフィック予測、マルチスケールの物理情報に基づいたニューラル ネットワークにわたって、問題の定式化からメカニズムの設計、検証、および制限された改訂に至るまで追跡可能な軌跡を保存します。これらの結果は、AI 科学者は、最終的な成果物だけでなく、その合成と検証のプロセスが帰属可能であり、検査可能であり、科学的に責任を負っているかどうかによって評価されるべきであることを示唆しています。

原文 (English)

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claims often remains implicit inside model inference. Here we introduce Xcientist, a research harness that externalizes research synthesis and experimental validation into inspectable, contract-governed processes. Xcientist organizes literature evidence, idea states, implementation plans, ablation records and repair traces as persistent research artifacts, so that generated mechanisms can be grounded, executed, tested and revised without losing their evidential basis. We identify claim drift as a failure mode of automated research, where runnable artifacts no longer support the mechanism originally claimed. Across training-free memory systems, graph-structured traffic forecasting and multi-scale physics-informed neural networks, Xcientist preserves traceable trajectories from problem formulation to mechanism design, validation and bounded revision. These results suggest that AI scientists should be evaluated not only by their final artifacts, but by whether their synthesis and validation processes remain attributable, inspectable and scientifically accountable.

13:00 JSTエージェント

SAGA: 長期的な CivRealm 戦略計画のためのシーンを認識し、目標を進化させるエージェント

複雑な戦略ゲームにおける長期的な戦略計画には、不完全な情報とまばらな報酬の下で、複数の意思決定領域にわたる同時推論が必要です。既存の LLM ベースのエージェントは、生のタイル座標によるシーンのブラインドネス、モノリシックな状態ダンプによるコンテキストのオーバーフローとドメインの結合、各エピソードを個別に扱う浅いクロスゲーム学習という 3 つの系統的な障害に悩まされています。我々は、それぞれ 1 つのクラスの障害を直接ターゲットとする 3 つのメカニズムを備えた LLM マルチエージェント フレームワークである SAGA を紹介します。(i) ゲーム エンティティ間の型指定された空間関係をユニットごとの自然言語コンテキストにエンコードするマップ セマンティック シーン グラフ。グローバルなトークン インフレーションを行わずに空間盲目を解決します。 (ii) オンデマンドで詳細なドメイン状態を取得し、専用の専門コントローラーにドメインごとのディレクティブをディスパッチして、コンテキスト オーバーフロー、ドメイン結合、および機械的制約違反を排除するツール拡張プランナー。 (iii) 定期的なゲーム内目標生成と構造化されたゲーム間の因果関係の事後分析を組み合わせたデュアルホライズン フィードバック ループにより、手動による報酬エンジニアリングを行わずに原則に基づいた戦略的進化が可能になります。 FreeCiv で評価された SAGA は、2 つの最も強力なベースラインよりも低い分散で最高の平均文明スコア (環境で唯一のまばらな目標報酬) を達成し、複数の目標の競合下で最も簡単に犠牲になるリソース軸であるインフラストラクチャ建設のすべてのベースラインを大幅に上回る唯一の方法です。これは、ほとんどの対戦ゲームで 2 つの最も強力なベースラインを上回り、出力トークン (主要なデコード コスト) を 27% 削減します。クロスゲーム進化モジュールを搭載した SAGA は、連続する 5 つのエピソードにわたって最高のエンドオブチェーン スコアに達します。アブレーション研究により、各構造コンポーネントが独立してこの利点に貢献していることが確認されています。

原文 (English)

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon Strategy Game Planning

Grand-strategy games such as Civilization pose a distinctive long-horizon planning problem: an agent must divide one shared resource pool among six competing domains -- technology, government, diplomacy, city development, expansion, and military -- under partial observability, with no feedback except a delayed final score. Current LLM agents fall short in three ways: 1) they cannot infer spatial relations from raw coordinates; 2) they allocate resources poorly, because feeding the entire growing state into one prompt and planning all domains in a single output diffuses attention and biases decisions toward urgent events; and 3) they cannot improve, as the delayed score gives no signal within or across games. We present SAGA, an LLM multi-agent framework pairing one mechanism with each weakness: (i) a Map-Semantic Scene Graph turning coordinates into per-entity statements of distance, direction, and threat; (ii) a Tool-Augmented Planner that retrieves only the state a decision needs, cutting the order of magnitude of its input, and issues a separate plan per domain to six specialist controllers, so urgent events do not derail long-term plans; and (iii) a Dual-Horizon Feedback Loop setting short-term goals during play and distilling each game into lessons for the next. On CivRealm, a Civilization-style benchmark, SAGA leads five LLM baselines on mean final score and is the only method significantly ahead of all of them on city development, the first investment baselines sacrifice, with 27% fewer output tokens; with cross-game learning it scores highest after five games, and its fifth game consistently surpasses its first across four maps. Our code is available at https://github.com/Kazecloudk/SAGA-Scene-Aware-Goal-Evolving-Agents-for-Long-Horizon-Strategy-Game-Planning.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mechanistic interpretability, and governance/auditability, covering 2018-2026 evaluation-safety measurement work. We introduce EvalSafetyGap as an organizing hypothesis for comparing evaluation-side and alignment-side proxy failures under optimization pressure, using Goodhart's Law together with two constructs we develop here - an Instability Decomposition and an Alignment Trilemma - as tools for generating testable comparisons. The audit shows how conclusions shift when capability, behavioral safety, and governance are measured separately. In this sample ($n = 10$), the association between capability and sustained adversarial robustness is statistically indeterminate using the displayed Table 3 inputs (Pearson $r = +0.232$, $p = 0.520$), and the apparent open-closed safety gap is modest, driven mainly by governance and disclosure rather than behavioral robustness, and sensitive to how a single borderline model is classified; attempt-budget results are protocol dependent. Because the public evidence uses heterogeneous protocols, the audit is diagnostic rather than rank-generating. The contribution is a shared vocabulary and evidence map to support dynamic evaluation, transparent source reporting, multi-attempt safety measurement, and auditable alignment practice.

13:00 JSTLLM/生成AI

サブリミナル時計: 拡散言語モデルにおける潜時モデリング

拡散言語モデル (DLM) は、自己回帰モデルの有望な代替手段として最近登場しました。標準的な拡散ベースのアプローチとは異なり、DLM はタイムステップで明示的に条件付けされていないため、当然の疑問が生じます。これらのモデルは内部的にノイズ除去の進行状況を表しているのか、そのような情報は下流でどのように使用されるのでしょうか。この研究では、DLM が実際にその残差ストリーム内の拡散タイムステップに関連する潜在表現をエンコードしていることを示します。この信号は、層全体のプローブを使用して確実に抽出できることがわかり、ノイズ除去の進行が内部活性化から解読可能であることがわかります。さらに、推論されたタイムステップに関連付けられた低次元部分空間に沿ってモデルを操作することで、ノイズ除去の進行の概念を体系的に調整できるようになり、モデルの信頼性とエントロピーに予測可能な変化がもたらされることを示します。最後に、識別された表現の幾何学的形状を分析し、それが活性化空間で構造化された解釈可能な特性を示すことを示し、そのような信号がこれらのモデルによってどのように処理されるかを明らかにします。

原文 (English)

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used downstream? In this work, we show that DLMs do in fact encode a latent representation related to the diffusion timestep within their residual streams. We find that this signal can be reliably extracted using probes across layers, indicating that denoising progress is decodable from internal activations. We further demonstrate that steering the model along a low-dimensional subspace associated with the inferred timestep allows us to systematically modulate its notion of denoising progress, leading to predictable changes in model confidence and entropy. Finally, we analyse the geometry of the identified representation, showing that it exhibits structured and interpretable properties in activation space, and shedding light on how such a signal is processed by these models.

13:00 JSTLLM/生成AI

切り捨てられた思考連鎖監査による LLM ベースの教育講師の回答主導型推論の検出

大規模言語モデル (LLM) の家庭教師は、流暢な段階的な説明を行うことがよくありますが、正しく教育的にフォーマットされた応答は、その答えが生徒が直面している問題から導き出されたものであることを保証しません。現実的な個別指導システムでは、モデルは教師のノート、解答キー、ルーブリック、または取得されたソリューションのアーティファクトにもアクセスできる場合があります。私たちは、そのようなプライベートな解答情報によって、家庭教師の説明が解答主導型になるかどうか、つまり最終的な解答は、書面による説明が正当化する前に行動的に入手可能になるかどうかを研究しています。思考連鎖プレフィックスが検証者をどれだけ早く通過できるかを調査する Truncated Reasoning AUC Evaluation (TRACE) を使用して、質問のみ、正解キー、間違った解答キーという 3 つのペアの個別指導コンテキストの下で 1000 件の GSM8K テスト問題を評価しました。生成された各説明の一定の割合で、モデルに即時応答を強制し、その応答を黄金の数値応答と照合して検証します。 Qwen2.5-3B-Instruct を使用すると、アンサーキー アクセスにより TRACE AUC の中央値が 0.375 から 0.900 に上昇し、1000 件中 997 件の最初の 10% プレフィックスでゴールド アンサーが利用可能になります。この効果は、質問のみの説明と解答キーの説明の両方が正解で終わる 746 例では依然として強いです。これらの結果は、数学の個別指導の説明における回答主導型の推論のための軽量のプロセス レベルの診断として、切り詰められた CoT 監査をサポートします。

原文 (English)

Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors

Large language model (LLM) tutors may have access to teacher notes, answer keys, rubrics, or retrieved solutions while producing student-facing explanations. We study whether truncated reasoning probes can distinguish direct access to such private context from answer information carried by the written explanation. Using Truncated Reasoning AUC Evaluation (TRACE), we evaluate 1000 GSM8K problems under question-only, correct answer-key, and wrong answer-key contexts. When forced-answer probes retain the private key, answer-key TRACE AUC rises from 0.375 to 0.900, and the gold answer is recoverable with no explanation at all in 998 of 1000 cases. We then introduce a context-masked replay: answer-key-generated prefixes are probed under the corresponding question-only prompt. Masking reduces 10\% prefix accuracy from 0.997 to 0.126 and median AUC from 0.900 to 0.375, nearly matching question-only values of 0.113 and 0.375. On 746 pairs where both explanations end correctly, the masked mean AUC difference is $-0.0086$ with a 95\% bootstrap interval spanning zero. Wrong keys still account for 272 of 387 incorrect final responses, showing that private artifacts can influence outputs even when early-prefix evidence disappears after masking. These results establish context masking as necessary for attributing early answer availability to an explanation rather than its hidden input.

13:00 JSTハードウェア/半導体

JEPA スタイルの予測学習を JA4 由来のネットワーク フィンガープリントに適用する

I-JEPA と V-JEPA は、元の入力を再生成するのではなく、潜在的な予測をターゲットのエンコーダー出力に照合することで学習します。これは画像やビデオではうまく機能します。同じ目的がコンパクトなネットワーク フィンガープリントでも機能するかどうかを調査します。 JA4DB および CIC-IDS-2017 から抽出された JA4、JA4H、JA4S、および JA4X サブフィールドでトレーニングされた Transformer ベースのモデルである JA4-JEPA を構築しました。トレーニング データは両方のソースからの約 397,000 のサンプルを組み合わせていますが、4 つのビュー ファミリすべてを含む単一のサンプルはありません。私たちは、TLS、DNS、SSH にわたるプロトコル ファミリ分類について、凍結された kNN プローブを使用して学習された表現を評価しました。 39,416 個のホールドアウト サンプルで、モデルはコサイン類似度 0.9899 と kNN 精度 0.9220 を達成しました。これらの結果は、ソース間でビューが不完全に重複している場合でも、JEPA スタイルの予測学習が JA4 由来のフィンガープリントから有用な埋め込みを生成できることを示しています。キーワード: JA4、ネットワークフィンガープリンティング、JEPA、予測表現学習、自己教師あり学習

原文 (English)

Applying JEPA-Style Predictive Learning to JA4-Derived Network Fingerprints

I-JEPA and V-JEPA learn by matching latent predictions to target encoder outputs rather than regenerating the original input, and this has worked well for images and video. We explore whether the same objective works for compact network fingerprints. We built JA4-JEPA, a Transformer-based model trained on JA4, JA4H, JA4S, and JA4X subfields drawn from JA4DB and CIC-IDS- 2017. The training data combines roughly 397K samples from both sources, though no single sample contains all four view families. We evaluated the learned representations with a frozen kNN probe on protocol-family classification across TLS, DNS, and SSH. On 39,416 heldout samples the model achieved a cosine similarity of 0.9899 and a kNN accuracy of 0.9220. These results indicate that JEPA-style predictive learning can produce useful embeddings from JA4-derived fingerprints, even with incomplete view overlap across sources. Keywords: JA4, network fingerprinting, JEPA, predictive representation learning, self-supervised learning

13:00 JST研究/論文

AdvNav: 視覚言語ナビゲーションに対する行動誘導型ブラックボックス攻撃

身体化 AI の進歩にもかかわらず、視覚と言語のナビゲーション システムは依然として敵対的な視覚障害に対して脆弱です。既存の手法のほとんどは、ターゲット モデルの勾配へのホワイト ボックス アクセスに依存していますが、これは現実世界に展開されたシステムでは非現実的であることが多く、最適化のための再帰逆伝播により計算量が膨大になり、適用性が制限されます。これまでのブラックボックス手法は主に単一ステップの瞬時の意思決定タスクを対象としていましたが、タスクの複雑さと時間的な依存関係を処理するのに苦労していました。これは、観察可能な入力と出力のみを使用して、複数ステップの順次的な知覚と行動のループを効果的に破壊できる、勾配のない攻撃方法の必要性を強調しています。したがって、ナビゲーション中にエージェントの一人称視点を妨害する、行動誘導型のブラックボックス敵対的攻撃フレームワークである AdvNav を提案します。ブラックボックス設定の下での無勾配探索における効果的な最適化ガイダンスのための有益な代理目標を構築するために、全体的なナビゲーションの低下を表す軌道レベルのパフォーマンススコア、潜在的な意思決定リスクを考慮したアクションレベルの報酬スコア、および逸脱指標を集約する二重粒度の行動ベースのフィードバックを設計します。これらはすべてエージェントの自己出力行動から抽出されます。このフィードバックは、適応更新によって摂動の強度をヒューリスティックに調整し、ノイズの空間構造を遺伝的に進化させて、最も破壊的なノイズ構成を繰り返し発見するハイブリッド最適化戦略を導きます。 R2R データセット上の 2 種類のバックボーンを使用して、Transformer ベースの HAMT および LLM ベースの MapGPT に対して評価したところ、AdvNav は 49.70/65.96/87.30% の攻撃成功率を達成しました。この結果は、AdvNav の有効性と汎用性を実証し、重大な認識の脆弱性を明らかにし、将来の回復力のある VLN モデルの設計のための洞察を提供します。

原文 (English)

AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation

Despite progress in Embodied AI, Vision-and-Language Navigation systems remain vulnerable to adversarial visual disturbances. Most existing methods rely on white-box access to target model gradients, which is often unrealistic for real-world deployed systems and computationally exhaustive due to recursive backpropagation for optimization, limiting their applicability. While previous black-box methods predominantly target single-step, instantaneous decision tasks, they struggle to handle the task complexities and temporal dependencies. This highlights the need for a gradient-free attack method that can effectively disrupt the multistep sequential perception-action loop using only observable inputs and outputs. Therefore, we propose AdvNav, a behavior-guided black-box adversarial attack framework that disturbs an agent's first-person views during navigation. To construct an informative surrogate objective for effective optimization guidance in gradient-free search under the black-box setting, we design a dual-granularity behavior-based feedback, aggregating a trajectory-level performance score representing overall navigation degradation, an action-level reward score considering the potential decision risk, and a deviation indicator, all of which are extracted from the agent's self-output behaviors. This feedback guides a hybrid optimization strategy that heuristically tunes perturbation strength via adaptive updates and evolves noise spatial structure genetically, to iteratively discover the most disruptive noise configuration. Evaluated against Transformer-based HAMT and LLM-based MapGPT with two types of backbones on R2R dataset, AdvNav achieves 49.70/65.96/87.30% Attack Success Rate. The result demonstrates the effectiveness and generality of AdvNav, reveals critical perception vulnerabilities and offers insights for the design of future resilient VLN models.

13:00 JST研究/論文

筋骨格ケアのための証拠に基づいた AI

筋骨格系疾患は世界中で障害の主な原因の一つであり、世界的にリハビリテーションに対する最大のニーズを生み出しています。回復、リモデリング、変性は数カ月から数年かけて進行することが多いため、筋骨格ケアには、進化する患者の証拠、外部の医学知識、段階別の機能目標を繰り返し統合する長期的な管理が必要です。日常診療では、この証拠は訪問、部門、病院システム全体で断片化されており、個別化された証拠に基づいたケアが制限されています。ここでは、継続的な筋骨格管理のために病院のデータ ストリームと信頼できる外部の知識を統合する、大規模な言語モデルを搭載した臨床人工知能システムである OrthoPilot について報告します。 OrthoPilot は、リアルタイムの画像データ、検査データ、病理学データ、注文データを自律的に取得し、入院診断からリハビリテーション計画に至るまで、進化する患者の状態を証拠に基づいた決定に変換します。私たちは、1,000 の疾患コードにわたる実際の電子医療記録から専門家によって検証されたベンチマークを確立しました。コンプリートケア経路全体にわたる読者調査で、OrthoPilot は 81 人の整形外科医と比較され、診断推論、臨床意思決定、管理計画において 25 年の経験を持つ専門家を上回りました。また、60 の外部臨床センターで評価されたすべてのインテリジェント システムを上回りました。 1,870 件の複雑なケースを対象とした前向き研究で、OrthoPilot はフルチェーン管理の成功率を 10.6% 向上させました。 8,240 人の入院患者を対象とした 8 か月間にわたる無作為化導入により、ベッドあたりの累積感染者数が 9.7% 増加し、患者報告による健康情報へのアクセスが改善されました。これらの結果により、臨床 AI は、孤立したイベントの予測から、完全な筋骨格ケア経路にわたる長期的な管理の実行へと移行します。

原文 (English)

Evidence-Grounded AI for Musculoskeletal Care

Musculoskeletal diseases are among the leading causes of disability and drive the greatest global need for rehabilitation. Because recovery, remodelling and degeneration of bones, joints and related tissues unfold over months to years, care requires longitudinal management rather than isolated decisions. Clinicians must repeatedly integrate evolving patient evidence, medical knowledge and stage-specific functional goals, yet evidence is often fragmented across visits, departments and hospital systems, disrupting continuous, individualised management. Here we report OrthoPilot, a clinical artificial intelligence (AI) system powered by a large language model (LLM) that integrates hospital data streams with authoritative external knowledge for continuous musculoskeletal care. It autonomously retrieves real-time imaging, laboratory, pathology and order data and translates evolving patient states into evidence-based decisions from admission diagnosis through rehabilitation planning. We established a specialist-validated benchmark from real-world electronic health records (EHRs) spanning 1,000 disease codes. In a full-pathway reader study against 81 orthopaedic physicians, OrthoPilot outperformed experts with 25 years of experience in diagnostic reasoning, clinical decision-making and management planning. This advantage generalised across 60 external clinical centres, where OrthoPilot surpassed all evaluated intelligent systems. In a prospective physician decision-making study of 1,870 complex cases, OrthoPilot improved full-chain management success by 10.6%. In a randomised deployment involving 8,240 inpatients, integration into routine care increased cumulative cases per bed by 9.7% and improved patient-reported access to health information. These results move clinical AI from predicting isolated events toward executing longitudinal management across complete musculoskeletal care pathways.

13:00 JSTLLM/生成AI

FormalAnalyticGeo: マルチモーダル解析幾何問題生成のためのニューラルシンボリックベースのフレームワーク

数学的推論は、マルチモーダル大規模言語モデル (MLLM) の急速な進歩により大幅な進歩を遂げていますが、解析幾何学は、主に注釈付きのサンプルが不足しているため、ほとんど研究されていません。既存のダイアグラム生成アプローチは、解析ジオメトリに苦労しています。テンプレート メソッドは制約駆動のレイアウトを処理できず、生成モデルには注釈付きの円錐曲線を正しくレンダリングするための幾何学的精度が不足しています。私たちは、マルチモーダルな解析幾何学問題を完全に自動生成するためのスケーラブルなフレームワークである FormalAnalyticGeo を紹介します。形式言語の厳密性を活用して、CDL (条件記述言語) を中心としたフレームワークを設計します。これは、自由形式の問題テキストと、符号付き距離フィールド (SDF) エンジンを介した正確な図のレンダリングを橋渡しする形式的な中間表現です。このフレームワークは、4 つの特殊な LLM コンポーネントを順番に使用します。さまざまな解析幾何学問題を生成するジェネレーター、SDF ベースのレンダリング用に各問題を CDL に変換するフォーマライザー、レンダリングされたダイアグラムのビジョンベースの測定を通じてグランドトゥルースの答えを抽出する測定器、および 3 つの段階で出力をチェックする品質検証器です。 Quality Verifier からの構造化されたフィードバックにより自動再試行が行われ、人間による注釈の必要性を排除する閉ループが形成されます。 FormalAnalyticGeo を大規模に適用すると、7K を超える検証済みのマルチモーダル問題のデータセットである AnalyticGeo7K が生成され、それぞれに位置合わせされたテキスト、図、正式な注釈、グラウンド トゥルースが含まれます。実験によると、生成された問題は、グラウンド トゥルース相対誤差の中央値 0.70\% に達し、回答の 82.3\% が正確なシンボリック解の 5\% 以内に収まります。私たちのフレームワークとデータセットは一般に公開されます。

原文 (English)

FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation

Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remains largely underexplored, primarily due to the scarcity of annotated samples. Existing diagram generation approaches struggle with analytic geometry: template methods cannot handle constraint-driven layouts, and generative models lack the geometric precision to render annotated conic curves correctly. We present FormalAnalyticGeo, a scalable framework for fully automatic generation of multimodal analytic geometry problems. Leveraging the rigor of formal languages, we design the framework around CDL (Condition Description Language), a formal intermediate representation that bridges free-form problem text with precise diagram rendering via a Signed Distance Field (SDF) engine. The framework employs four specialized LLM components in sequence: a Generator that produces diverse analytic geometry problems, a Formalizer that converts each problem into CDL for SDF-based rendering, a Measurer that extracts ground-truth answers through vision-based measurement on the rendered diagrams, and a Quality Verifier that checks outputs at three stages. Structured feedback from the Quality Verifier drives automatic retry, forming a closed loop that eliminates any need for human annotation. Applying FormalAnalyticGeo at scale yields AnalyticGeo7K, a dataset of over 7K verified multimodal problems, each with aligned text, diagram, formal annotation, and ground truth.Experiments show that the generated problems achieve a median ground-truth relative error of 0.70\%, with 82.3\% of answers falling within 5\% of the exact symbolic solution. Our framework and dataset will be publicly released.

13:00 JSTLLM/生成AI

抵抗して更新: インセンティブ対応 LLM の反事実報告の調整

調整された言語モデルは、証拠のないインセンティブ圧力の下で日常的に誤った報告をします。つまり、自信のあるユーザーに同意したり、ユーザーの内部信念が変わっていない場合でも確信度を誇張したりします。我々は、これを内部インセンティブ互換性(IC)の失敗として位置づけ、モデルのレポートを因果関係の契約に保持する反事実レポートメディエーターを学習および認定する方法を提示します。つまり、禁止された影響(圧力、威信、スタイル変更)に対して不変であり、ライセンスされた影響(本物の証拠)に応答します。抵抗と更新という 2 つの要求は、反対方向に引っ張られます。私たちはそれらを、既知の事後分布を使用したベイジアン・ウィットネス・ベンチマークで研究します。このベンチマークでは、同じユーザーの意見の相違が、純粋に述べられたソースの信頼性によって認可された証拠または禁止された圧力となります。我々は、(i) プローブの精度ではなく交換介入によって、ほぼ直交で独立して制御可能な、回答、信頼度、警告に関する低ランクのレポート座標を因果的に特定し、(ii) 反事実的にインセンティブが中立化されたコンテキストの下でモデル自身のレポートを参照する、トレーニング不要の反事実レポート座標 (CRC) クランプを導入します。ウィットネス ベンチマークでは、2 パス クランプは耐性と 1.00 の更新を合わせて達成 (Wilson 95% CI [0.99,1.00])、展開されたソリューションではなく、構築可能な参照の下での因果関係の証明書です。グローバルなデコードとステアリングでは、単一パラメータのトレードオフが示されます。出力レベルの微調整は、両方が列挙されている場合にのみ、両方の目的と一致します。抵抗のみのトレーニングは証拠への反応性を失います。デプロイ可能なシングルパス コンパイルには非可逆性があります (0.73/0.97)。メカニズムとクランプは 3 つのモデル ファミリにわたって再現され、自然なおしゃべりベンチマーク (SycophancyEval) に移行されます。私たちの貢献はインターフェイスと認証方法です。内部 IC の構造プリミティブとしてのアクティベーション レベルの反事実的インセンティブ不変性です。

原文 (English)

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

Aligned language models routinely misreport under non-evidential pressure: they cave to a confident user, yet fail to revise when genuine evidence arrives. We cast this as a failure of internal incentive-compatibility and study the two demands, resist (ignore forbidden pressure) and update (follow licensed evidence), on a Bayesian-witness benchmark with known posteriors, where the same user disagreement is evidence or pressure purely by stated source reliability, removing the evidence/pressure confound by construction. Using interchange interventions rather than probes, we causally localize low-rank report coordinates for answer, confidence, and caveat, establishing causal sufficiency at a late intervention site rather than uniqueness or necessity, with a causal cross-talk matrix showing strong own-coordinate control and only small cross-effects (partial functional disentanglement). We then introduce a training-free counterfactual report-coordinate (CRC) clamp that references the model's own report under an incentive-neutralized counterfactual of the prompt. The two-pass full-window clamp attains resist and update of $1.00$ jointly (Wilson 95% CI $[0.99,1.00]$; the rank-16 projection alone reaches $0.88/0.90$), which we read as a causal certificate and upper bound under a constructible reference, not a claim of a deployed solution. Tested global decoding and fixed-direction steering trade one objective against the other, and resist-only training collapses updating to $0.01$. The deployable single-pass compilation is lossy ($0.73/0.97$). The mechanism and the clamp reproduce across three model families and transfer to a natural sycophancy benchmark with significant paired improvements. Our contribution is the interface and certification method: activation-level counterfactual incentive-invariance as a structural primitive for internal incentive-compatibility.

13:00 JSTLLM/生成AIエージェント

RetroAgent: LLM を利用して構造化メモリを検索し、エージェントによる逆合成計画を立てる

複数ステップの逆合成計画では、実行可能な一連の反応を通じて、標的分子を市販の構成要素に分解することを目指します。広大な組み合わせ探索空間により、この作業は専門の化学者にとってさえ困難になります。従来の方法では、ツリー検索とオフラインでトレーニングされた値ネットワークを組み合わせて、完全な複数ステップのルートを推論することなく、候補を個別にスコアリングします。最近の研究では、このタスクに大規模言語モデル (LLM) を活用していますが、単純なインターフェイスに依存しているため、検索空間全体の探索が制限されています。 RetroAgent は、構造化メモリのハーネスを通じて記号検索と神経推論の橋渡しをする LLM エージェントです。エージェントは、記憶および化学ツールを通じて、探索されたルート、利用可能な代替案、中間体の特性を含む完全な探索状態を観察し、世界的な進歩とドメイン知識の両方に基づいた情報に基づいた意思決定を可能にします。ディストリビューション内およびディストリビューション外のベンチマークに関する実験により、RetroAgent が強力なパフォーマンスと汎用性を実現することが実証されました。

原文 (English)

RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning

Multi-step retrosynthesis planning seeks to decompose a target molecule into commercially available building blocks through a sequence of feasible reactions. The vast combinatorial search space makes this task challenging even for expert chemists. Traditional methods combine tree search with offline-trained value networks that score candidates in isolation, without reasoning about complete multi-step routes. Recent work leverages Large Language Models (LLMs) for this task, but relies on simple interfaces that limit exploration of the full search space. We introduce RetroAgent, an LLM agent that bridges symbolic search and neural reasoning through a harness with structured memory. Through memory and chemistry tools, the agent observes the full search state, including explored routes, available alternatives, and properties of intermediates, enabling informed decisions grounded in both global progress and domain knowledge. Experiments on in-distribution and out-of-distribution benchmarks demonstrate that RetroAgent delivers strong performance and generalization.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文Gemini

科学的視覚化リテラシーのためのマルチモーダル大規模言語モデルのベンチマーク

マルチモーダル大規模言語モデル (MLLM) は、ビジュアライゼーションを解釈するためにますます使用されていますが、現在の評価は依然として主にチャート中心であり、科学的ビジュアライゼーション (SciVis) の理解を示す証拠は限られています。私たちは、科学的視覚化リテラシー評価テストで 6 つの MLLM をベンチマークします。このテストは、8 つのテクニックと 11 のタスク タイプにわたる、18 の科学的視覚化とイラストに基づく 49 項目で構成される標準化された SciVis リテラシー評価です。私たちは、クローズドワールドプロトコルの下で 3 つのクローズドソースモデルと 3 つのオープンソースモデルを評価し、485 人の人間の参加者からのデータを使用してパフォーマンスを比較します。結果は、現在の MLLM が均一な SciVis リテラシーを示さないことを示しています。 Gemini は全体として最も強力なモデルであり、評価されたサブセット全体で人間の平均を上回っていますが、オープンソース モデルは依然として人間のベースラインを下回っています。パフォーマンスはテクニックやタスクによって大きく異なります。モデルは科学的なイラスト、検索、空間理解では最高のパフォーマンスを発揮しますが、テクスチャ ベースおよび統合ベースの視覚化と定量的推定では苦戦します。エラー分析により、きめの細かい定量的推定、フロー方向の解釈、および根拠のあるエンコードの解釈における繰り返しの失敗が明らかになります。これらの調査結果は、SciVis リテラシーをマルチモーダル AI システムを評価するために必要なベンチマークの側面として位置づけています。コードとモデルの出力は、https://github.com/patdmp/mllm-scivis-lit-benchmark で公開されています。

原文 (English)

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at https://github.com/patdmp/mllm-scivis-lit-benchmark.

13:00 JSTビジネス/資金調達

DL オントロジーの認識的機密性ポリシーに基づく扱いやすいクエリ応答 (拡張バージョン)

私たちは、記述ロジック (DL) オントロジーのコンテキストで、また認識依存関係 (ED) を通じて表現される機密性ポリシーについて、機密性を保持するデータ アクセスへの宣言的アプローチである制御クエリ評価 (CQE) を研究します。まず、CQE の既知のセマンティクス (GA および IGA 含意) の下でクエリ (具体的には論理積クエリのブール和集合) に応答する問題に取り組みます。私たちの結果は、TBox が $\text{DL-Lite}_{\mathcal{R}}$ で表現される場合、CQE は一般に計算的に扱いにくいことを示しています。さらに、ED が存在する場合、IGA セマンティクスは、識別不可能性として知られる重要な機密保持特性を満たさないことが最近証明されました。計算が容易で機密性が保たれる CQE 形式を定義することを目的として、最小限のポリシー違反 (MPV) の概念に基づいた CQE の新しいセマンティクスを導入します。新しいセマンティクスが以前のセマンティクスの健全な近似を提供しながら、区別不可能性の特性を満たしていることを示します。また、$\text{DL-Lite}_{\mathcal{R}}$ オントロジーの場合、MPV セマンティクスに基づくクエリ含意がデータ複雑さの多項式時間で決定できることも証明します。最後に、OWL 2 QLの既存のベンチマークを使用して、この新しいアプローチの実現可能性を評価するために使用したフレームワークのソフトウェア実装を紹介します。

原文 (English)

Tractable Query Answering under Epistemic Confidentiality Policies in DL Ontologies (extended version)

We study Controlled Query Evaluation (CQE), a declarative approach to confidentiality-preserving data access, in the context of Description Logic (DL) ontologies, and for confidentiality policies expressed through Epistemic Dependencies (EDs). We first address the problem of answering queries (specifically, Boolean unions of conjunctive queries) under known semantics for CQE (GA- and IGA-entailment). Our results show that if the TBox is expressed in $\text{DL-Lite}_{\mathcal{R}}$, CQE is computationally intractable in general. Moreover, in the presence of EDs, the IGA semantics has recently been proven not to satisfy an important confidentiality preservation property known as indistinguishability. With the goal of defining computationally easier and confidentiality-preserving forms of CQE, we introduce a new semantics for CQE, based on the notion of minimal policy violation (MPV). We show that the new semantics provides a sound approximation of the previous ones, while satisfying the indistinguishability property. We also prove that, in the case of $\text{DL-Lite}_{\mathcal{R}}$ ontologies, query entailment under the MPV semantics can be decided in polynomial time in data complexity. Finally, we present a software implementation of our framework that we used to evaluate the feasibility of this new approach using an existing benchmark for OWL 2 QL.

13:00 JSTLLM/生成AI研究/論文

思考の多様性の定量化: 加重 LLM アンサンブル リフトの予測法則

この論文は、思考の多様性が大規模言語モデル (LLM) アンサンブルにもたらす向上を計算するための、実験的に検証された形式法則を提供します。第一原理に基づいて、LLM アンサンブル揚力を救助質量と損傷質量に正確に分解し、隆起を計算するためのコンパクトなヒューリスティックを生成します。ここから、アンサンブルのパフォーマンスを予測するメトリクス、つまり精度調整された正しさの相関 $\phi_{\mathrm{adj}}$ と、ペアの精度ギャップおよび集合精度を抽出します。私たちは、2 つの大学院レベルの科学ベンチマークにわたる 10 のオープンウェイト モデルからの 767,520 件の推論に関する法則をテストします。また、各モデルがネットワークで隔離されたサンドボックス内でマルチターン ツールを使用してデジタル フォレンジック調査を行う新しいエージェント サイバーセキュリティ ベンチマークもテストします (棄権を含む 23,520 件の段階的試験)。すべての投票は公開で公開されます。 SuperGPQA で 40:60 の投票分割で一度キャリブレーションされると、ヒューリスティックはスピアマンの $\rho=0.84$ を使用してキャリブレーション セットのリフトを予測し、その係数を凍結してキャリブレーションでは決して使用されなかった 2 つのデータセット (GPQA ダイヤモンドでは $\rho=0.51$、フォレンジック タスクでは $0.84$) に転送します。一方、測定されたスワップ マス トラックは次のようなリフトを実現しました。全体で $R^2\ge 0.96$。生の $\phi$ にはほとんど予測能力がありません (全体を通して $R^2\le 0.09$)。精度調整された $\phi_{\mathrm{adj}}$ は著しく優れており (SuperGPQA では $R^2=0.67$)、これらのメトリクスを組み合わせたヒューリスティックは、3 つのデータセット全体で最も安定した事前プーリング予測子です。

原文 (English)

Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift

This paper provides an experimentally verified formal law for calculating the uplift that diversity of thought provides in Large Language Model (LLM) ensembles. From first principles, we derive an exact decomposition of LLM ensemble lift into rescue and damage masses, which yields a compact heuristic for calculating uplift. From this we extract the metrics which predict ensemble performance: an accuracy-adjusted correctness correlation, $\phi_{\mathrm{adj}}$, together with the accuracy gap and collective accuracy of the pair. We test the law on 767,520 inferences from ten open-weight models over two graduate-level science benchmarks, together with a novel agentic cybersecurity benchmark in which each model conducts digital-forensics investigations by multi-turn tool use in a network-isolated sandbox (23,520 graded trials including abstentions); all votes are released openly. Calibrated once on SuperGPQA at a 40:60 vote split, the heuristic predicts lift on the calibration set with Spearman's $\rho=0.84$ and, with its coefficients frozen, transfers to two datasets never used in calibration ($\rho=0.51$ on GPQA Diamond and $0.84$ on the forensic tasks), whilst the measured swap mass tracks realised lift with $R^2\ge 0.96$ throughout. Raw $\phi$ has almost no predictive power ($R^2\le 0.09$ throughout); the accuracy-adjusted $\phi_{\mathrm{adj}}$ is markedly superior ($R^2=0.67$ on SuperGPQA), and the heuristic combining these metrics is the most stable pre-pooling predictor across the three datasets.

13:00 JSTLLM/生成AIエージェント研究/論文Claude

AI エージェントは本当に RTL から GDS への変換を完了できるでしょうか?ベンチマーク ツールからの教訓 - インタラクティブ EDA ワークフロー

LLM 駆動のエージェント システムは、電子設計自動化 (EDA) の有望なパラダイムとして浮上しており、複雑な設計ワークフローを自動化する強力な可能性を示しています。ただし、既存の評価は主に、分離された EDA タスクに関する個々の言語モデルを調査しており、完全な EDA フロー全体でさまざまなエージェント システムがどのように実行されるかについての洞察は限られています。この研究では、統一されたプロンプト、ツール環境、テクノロジー ライブラリ設定の下でのエンドツーエンド EDA ワークフローにおける AI エージェントの体系的な評価である FluxBench を紹介します。当社の評価では、オープンソース ツールチェーンを使用した RTL 生成や、産業アプリケーション向けのクローズドソースの商用 EDA ツールを使用した RTL から GDS へのフローなど、代表的なシナリオをカバーしています。これらのワークフローを通じて、RTL コード生成、反復修復、ツール フィードバックの利用、論理合成、配置配線 (P&R)、およびエンジニアリング変更オーダー (ECO) 自動化におけるエージェントの能力を評価します。エージェント システムの効率をさらに特徴付けるために、トークンの使用量とランタイム コストと比較した EDA アーティファクトの効果的な改善を測定するコスト効率の指標であるトークン ROI を導入します。実験結果によると、同じ基盤モデルに基づいて構築されている場合でも、エージェント システム アーキテクチャが異なると、最大 86.27% のパフォーマンス ギャップが見られる可能性があります。さらに、同等のタスク パフォーマンスを持つシステム間では、トークン ROI が $105.92\times$ も異なる可能性があります。 PicoRV32 をケーススタディとして使用した RTL から GDS へのフローでは、FluxEDA は最大 97.94 のエンドツーエンド スコアを達成し、ドメイン固有の EDA スキルを備えた Claude Code を最大 $8.39\times$ 上回りました。これらの結果は、大規模な EDA シナリオでエージェントのパフォーマンスを向上させるには、ドメイン固有のスキルだけでは不十分であることを示しています。代わりに、エージェント システム設計と基盤モデル機能の両方が、効果的な自動 EDA ワークフローを実現する上で重要な役割を果たします。

原文 (English)

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows

LLM-driven agent systems have emerged as a promising paradigm for electronic design automation (EDA), demonstrating strong potential for automating complex design workflows. However, existing evaluations primarily examine individual language models on isolated EDA tasks, providing limited insight into how different agent systems perform across complete EDA flows. In this work, we present FluxBench, a systematic evaluation of AI agents on end-to-end EDA workflows under unified prompts, tool environments, and technology library settings. Our evaluation covers representative scenarios, including RTL generation with open-source toolchains and an RTL-to-GDS flow using closed-source commercial EDA tools for industrial applications. Through these workflows, we assess agents' capabilities in RTL code generation, iterative repair, tool-feedback utilization, logic synthesis, placement and routing (P&R), and Engineering Change Order (ECO) automation. To further characterize the efficiency of agent systems, we introduce Token ROI, a cost-efficiency metric that measures effective improvements in EDA artifacts relative to token usage and runtime cost. Experimental results show that, even when built on the same foundation model, different agent system architectures can exhibit performance gaps of up to 86.27%. Moreover, among systems with comparable task performance, Token ROI can differ by as much as $105.92\times$. In the RTL-to-GDS flow using PicoRV32 as a case study, FluxEDA achieves an end-to-end score of up to 97.94, outperforming Claude Code equipped with domain-specific EDA skills by up to $8.39\times$. These results indicate that domain-specific skills alone are insufficient to improve agent performance in large-scale EDA scenarios. Instead, both agent system design and foundation model capability play critical roles in enabling effective automated EDA workflows.

13:00 JSTLLM/生成AIエージェント

維持するか統合するか?言語エージェントのメモリに対する予算に応じたオペレーターの選択

言語エージェントは、対話全体にわたる記憶に依存します。ただし、大規模言語モデル (LLM) の限られたコンテキスト ウィンドウとその推論コストにより、一度に使用できるメモリの量が制限されます。既存のシステムは主に、メモリ保持とメモリ統合という 2 つの戦略に従っています。保存では生の記録が保持され、正確な詳細が保存されますが、関連する証拠は限られた予算内では収まらない可能性があります。統合によりレコードが圧縮されて結合されるため、トークンごとのカバレッジが向上しますが、クエリに不可欠な詳細が失われる危険性があります。どちらの戦略も一般的に好ましいものではありません。これにより、2 つの中心的な疑問が生じます。1 つは、いつ保存の代わりに統合を行うべきか、もう 1 つはマージ、抽象、またはリライトのどの演算子を選択する必要があるかということです。この決定は、各演算子の効用を、保持によって省略された証拠に対する適用効果と、すでに適合している生の証拠に対する署名付き置換効果に分解することによって形式化します。これらのバランスにより、予算の相対的な圧力によって優先されるアクションが変化する理由が説明されます。私たちは、このメカニズムを Offline Abstraction-Safety (OAS) で実装します。OAS は、ホールドアウトされた危害キャリブレーションを使用して生成前の特徴からアクション ユーティリティを推定する軽量の学習器です。公開されている LongMemEval ベンチマークと LoCoMo ベンチマークは、同じ予算依存のパターンを示しています。 LongMemEval では、予算が厳しい場合は統合により絶対精度が最大 48% 向上しますが、予算が緩い場合は保持することが望ましいです。 LoCoMo は、その短い証拠と一致して、より少ない予算でこのクロスオーバーを再現します。どちらのデータセットでも、圧縮が必要な場合、クロスノート抽象化とマージは一般にローカルな書き換えよりも優れたパフォーマンスを発揮します。

原文 (English)

Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory

Language agents depend on memory across interactions. However, the limited context windows of large language models (LLMs) and their inference costs constrain how much memory can be used at once. Existing systems mainly follow two strategies: memory retention and memory consolidation. Retention keeps raw records and preserves exact details, but relevant evidence may not fit under a tight budget; consolidation compresses and combines records, improving coverage per token but risking the loss of query-critical details. Neither strategy is universally preferable. This raises two central questions: when should consolidation replace retention, and which operator -- Merge, Abstract, or Rewrite -- should be selected? We formalize this decision by decomposing each operator's utility into a coverage effect on evidence omitted by retention and a signed replacement effect on raw evidence that already fits. Their balance explains why the preferred action changes with relative budget pressure. We implement this mechanism with Offline Abstraction-Safety (OAS), a lightweight learner that estimates action utilities from pre-generation features with held-out harm calibration. The public LongMemEval and LoCoMo benchmarks show the same budget-dependent pattern. On LongMemEval, consolidation improves absolute accuracy by up to 48% under tight budgets, whereas retention is preferable under loose budgets; LoCoMo replicates this crossover at a smaller budget, consistent with its shorter evidence. On both datasets, cross-note abstraction and merging generally outperform local rewriting when compression is necessary.

13:00 JSTエージェント

SR-Agent: E コマース レコメンデーションにおけるランキング後の戦略を洗練するためのエクスペリエンス主導型エージェント フレームワーク

ユーザー エクスペリエンスは、産業用電子商取引レコメンダー システム (RS) の第一級の目標です。ランク付け後の戦略は、ランク付けされたリストにおける多様性、類似性、露出を管理するもので、そのシンプルさと提供コストの低さから産業用 RS に広く導入されています。ただし、オンライン レコメンデーション環境が継続的に進化するにつれて、これらの静的に構成された戦略は徐々に古くなり、ユーザー エクスペリエンスが低下します。通常、これらを改良するには手動の検査、診断、更新が必要ですが、このプロセスは時間がかかり、コストがかかり、再利用が困難です。最近の LLM ベースのエージェント (RecUserSim、SimUSER、Self-EvolveRec など) は有望な方向性を示していますが、自動化され自己進化する戦略の改良の完全なループを閉じるものはありません。このギャップを埋めるために、SR-Agent を導入します。SR-Agent は、Strategy Refinement エージェント フレームワークであり、私たちの知る限り、産業用 RS でポストランキング戦略を洗練するために導入されたのは初めてです。 SR-Agent は 3 つのコンポーネントを統合します。(i) 段階的な検査スキルを適用して、ユーザーが認識する悪いケースを表面化する UserSim エージェント。 (ii) 再発する悪いケースを構造化された再利用可能な診断に統合する分析エージェント。 (iii) 診断を型指定された制限されたアクションにマッピングする制約付き戦略洗練ハーネス。可逆ロールバックを備えた 4 段階の報酬パイプラインによってゲートされます。 Kuaishou e-commerce プラットフォームに導入された SR-Agent は、この絞り込みループを継続的に実行し、1 か月のオンライン A/B テストで注文量を 0.71%、閲覧深度を 0.34%、クリックしたカテゴリの多様性を 0.48% 増加させながら、絞り込みサイクルを大幅に短縮し、運用コストを削減しました。

原文 (English)

SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategy Refinement in E-Commerce Recommendation

User experience is a first-class objective in industrial e-commerce recommender systems (RS). Post-ranking strategies, which govern diversity, similarity, and exposure over a ranked list, are widely deployed in industrial RS for their simplicity and low serving cost. However, as the online recommendation environment evolves continuously, these statically configured strategies gradually become stale, thereby degrading the user experience. Refining them typically relies on manual inspection, diagnosis, and updates, making it slow, costly, and difficult to scale or reuse. Although recent LLM-based agents (e.g., RecUserSim, SimUSER, and Self-EvolveRec) offer promising directions, none of them close the full loop of automated, self-evolving strategy refinement. To bridge this gap, we introduce SR-Agent, which, to the best of our knowledge, is the first agentic framework deployed to refine post-ranking strategies in industrial RS. SR-Agent unifies three components: (i) a UserSim agent that applies inspection skills to surface user-perceived bad cases; (ii) an Analysis agent that consolidates recurring bad cases into structured, reusable diagnoses; and (iii) a constrained Strategy Refinement Harness that maps diagnoses to typed and bounded actions, gated by a four-stage reward pipeline with reversible rollback. Deployed on the Kuaishou e-commerce platform, SR-Agent continuously runs this refinement loop and, in a one-month online A/B test, increases order volume by 0.71%, browsing depth by 0.34%, and clicked-category diversity by 0.48%, while markedly shortening the refinement cycle and lowering operational cost.

13:00 JST研究/論文NVIDIA

効率的なベイジアン推論の計算と展開のためのハードウェア指向のアプローチ

ベイジアン推論は、不確実性の下で推論するための原則に基づいた基盤を提供しますが、その計算コストにより、リソースに制約のあるエッジ デバイスへの展開が妨げられます。この論文では、市販の組み込み GPU で離散ベイズ推論を高速化するためのハードウェア指向の方法論を紹介します。私たちは、広範なクラスの変分メッセージ受け渡しアルゴリズムのレイテンシーがテンソル短縮によって支配されていることを確認しました。私たちのアプローチは、効率的な GPU 実行により適したコンパクトで規則的な形状のプリミティブを生成する 2 つの相補的なマージ戦略を使用して、これらの操作のメモリ レイアウトを再構築します。次に、メモリ フットプリントを削減するために、オプションのスパース配列表現とテンソル クラスタリング スキームを導入します。私たちは方法論をインスタンス化し、隠れマルコフ モデル (HMM) 用の 3 つのメッセージ パッシング アルゴリズム、つまり変分フィルタリング、変分メッセージ パッシング、およびマージナル メッセージ パッシングの最適化されたバリアントを生成します。さらに、特定の生成モデル仕様に対して最もパフォーマンスの高いアルゴリズムのバリアントを自動的に選択する機械学習ベースの自動チューナーでこれを補完します。 770 個のランダムにサンプリングされた現実的な部分観察可能なマルコフ決定プロセス (POMDP) 構成にわたる NVIDIA Jetson Orin AGX でベンチマークを行ったところ、当社の実装は、ベースライン実装と数値的に同一の出力を生成しながら、最大 5 倍の高速化を達成し、通常のゲインは 2 ~ 2.5 倍でした。

原文 (English)

A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment

Bayesian inference provides a principled foundation for reasoning under uncertainty, but its computational cost hinders deployment on resource-constrained edge devices. In this paper, we present a hardware-oriented methodology for accelerating discrete Bayesian inference on commercial off-the-shelf embedded GPUs. We identify that the latency of a broad class of variational message-passing algorithms is dominated by tensor contractions. Our approach restructures the memory layout of these operations using two complementary merging strategies that produce compact, regularly-shaped primitives better suited for efficient GPU execution. We then introduce optional sparse array representations and a tensor-clustering scheme to reduce the memory footprint. We instantiate the methodology and produce optimized variants of three message-passing algorithms for Hidden Markov Models (HMMs), namely variational filtering, variational message passing, and marginal message passing. Furthermore, we complement this with a machine-learning-based autotuner that automatically selects the best-performing algorithmic variant for a given generative model specification. Benchmarked on an NVIDIA Jetson Orin AGX across 770 randomly sampled realistic Partially Observable Markov Decision Process (POMDP) configurations, our implementations achieve speedups of up to 5x, with typical gains of 2-2.5x, while producing numerically identical outputs to the baseline implementations.

13:00 JST研究/論文

Bayesian inference of composition-dependent phase diagrams

Phase diagrams serve as a highly informative tool for materials design, encapsulating information about the phases that a material can mani…

13:00 JSTLLM/生成AIビジネス/資金調達

Saving the legacy of Hero Ibash: Evaluating Four Language Models for Aminoacian

This study assesses four cutting-edge language models in the underexplored Aminoacian language. Through evaluation, it scrutinizes their ad…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have bec…

13:00 JST画像/動画生成

Soft-TransFormers for Continual Learning

Inspired by the Well-initialized Lottery Ticket Hypothesis (WLTH), we introduce Soft-TransFormers (Soft-TF), a continual learning framework…

13:00 JST研究/論文

A Self-Supervised Framework for Space Object Behaviour Characterisation

Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasin…

13:00 JST研究/論文

Parameter-Efficient Continual Fine-Tuning: A Survey

The emergence of large pre-trained networks has revolutionized the AI field, unlocking new possibilities and achieving unprecedented perfor…

13:00 JST研究/論文

GSPRec: On Improving Item Representations in Graph Signal Processing for Collaborative Filtering

Graph-based collaborative filtering methods act as low-pass filters in the spectral domain and discard the intermediate-frequency component…

13:00 JST研究/論文

Chi-Square Wavelet Graph Neural Networks for Heterogeneous Graph Anomaly Detection

Graph Anomaly Detection (GAD) in heterogeneous networks presents unique challenges due to node and edge heterogeneity. Existing Graph Neura…

13:00 JSTLLM/生成AI研究/論文

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data…

13:00 JSTLLM/生成AI研究/論文

Can Interpretation Predict Behavior on Unseen Data?

Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen inpu…

13:00 JST画像/動画生成エージェント

FedS2R: One-Shot Federated Domain Generalization for Synthetic-to-Real Semantic Segmentation in Autonomous Driving

Federated domain generalization has shown promising progress in image classification by enabling collaborative training across multiple cli…

13:00 JSTLLM/生成AIロボティクス

RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation

Large language models (LLMs) demonstrate remarkable capabilities in reasoning and code generation, enabling robotic manipulation to be init…

13:00 JST研究/論文

Robust Belief-State Policy Learning for Quantum Network Routing Under Decoherence and Time-Varying Conditions

Quantum network routing requires online decisions under probabilistic entanglement generation, finite quantum memories, decoherence, imperf…

13:00 JSTLLM/生成AI

Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures

Despite their capabilities, Large Language Models (LLMs) remain opaque with limited understanding of their internal representations. Curren…

13:00 JSTLLM/生成AI

Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression

Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead.…

13:00 JST画像/動画生成

Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language. Yet, relying solely on t…

13:00 JST研究/論文

Beyond-Diagonal RIS Under Non-Idealities: Learning-Based Architecture Discovery and Optimization

Beyond-diagonal reconfigurable intelligent surface (BD-RIS) has recently been introduced to enable advanced control over electromagnetic wa…

13:00 JSTLLM/生成AI研究/論文

QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture

The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent f…

13:00 JST研究/論文

Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives

How complex collective behavior emerges from individual interactions is a fundamental scientific question, but experimental cost and diffic…

13:00 JST画像/動画生成ハードウェア/半導体

T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs

Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training. Rece…

13:00 JSTロボティクス

ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learning

Human-level contact-rich manipulation relies on the distinct roles of two key modalities: vision provides spatially rich but temporally slo…

13:00 JST研究/論文

Memo2496: Expert-Annotated Dataset and Dual-view Adaptive Framework for Music Emotion Recognition

Music Emotion Recognition (MER) is constrained by limited expert annotations and the need to establish robustness across heterogeneous corp…

13:00 JSTLLM/生成AI

PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation

Large language model (LLM) personalization aims to adapt general-purpose models to individual users. Most existing methods, however, are de…

13:00 JSTロボティクス

SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction

Accurate long-horizon vessel trajectory prediction remains challenging due to compounded uncertainty from complex navigation behaviors and…

13:00 JSTエージェントロボティクス

Toward Learning POMDPs Beyond Full-Rank Actions and State Observability

We are interested in enabling autonomous agents to learn and reason about systems with hidden states, such as locking mechanisms. We cast t…

13:00 JSTロボティクス

Training and Simulation of Quadrupedal Robot in Adaptive Stair Climbing and Descending for Indoor Firefighting: An End-to-End Reinforcement Learning Approach

Quadruped robots are used for primary searches during the early stages of indoor fires. A typical primary search involves quickly and thoro…

13:00 JSTLLM/生成AIエージェント

LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation

Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks…

13:00 JST研究/論文

CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation

Prefill-only KV compression freezes a token subset at the end of prefill and decodes from it without further eviction. The retention decisi…

13:00 JSTロボティクス

Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation

Current Visual-Language Navigation (VLN) methodologies face a trade-off between semantic understanding and control precision. While Multimo…

13:00 JST画像/動画生成

When Visual Evidence is Ambiguous: Pareidolia as a Diagnostic Probe for Vision Models

When visual evidence is ambiguous, vision models must decide how to interpret face-like patterns. Face pareidolia, the perception of faces…

13:00 JSTLLM/生成AI

Give Them an Inch and They Will Take a Mile:Understanding and Measuring Caller Identity Confusion in MCP-Based AI Systems

The Model Context Protocol (MCP) is an open and standardized interface that enables large language models (LLMs) to interact with external…

13:00 JSTエージェント

SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution

Real-world software must continuously evolve to meet ever-changing and open-ended requirements. AI agents, increasingly deployed as long-ru…

13:00 JSTロボティクス

TransDex: Pre-training Visuo-Tactile Policy with Point Cloud Reconstruction for Dexterous Manipulation of Transparent Objects

Dexterous manipulation enables complex tasks but suffers from self-occlusion, severe depth noise, and depth information loss when manipulat…

13:00 JSTLLM/生成AI

PlotTwist: A Creative Plot Generation Framework with Small Language Models

Creative plot generation presents a fundamental challenge for language models: transforming a concise premise into a coherent narrative tha…

13:00 JSTLLM/生成AIエージェントハードウェア/半導体

When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines

Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

Doctorina MedBench-ICD10: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI

We present Doctorina MedBench, a comprehensive evaluation framework for agent-based medical AI based on the simulation of realistic physici…

13:00 JST研究/論文

M-RAG: Semantic Key-Value Indexing for Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) turns external documents into evidence for large language models. In practice, this is also a data acc…

13:00 JSTLLM/生成AI

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification

The remarkable reasoning and code generation capabilities of large language models (LLMs) have recently motivated increasing interest in au…

13:00 JSTLLM/生成AI研究/論文Claude

Robust Reasoning Benchmark

While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on…

13:00 JST研究/論文

CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations

The rapid expansion of gaming industry requires advanced recommender systems tailored to its dynamic landscape. Existing Graph Neural Netwo…

13:00 JST画像/動画生成

Why Do Vision Language Models Struggle To Recognize Human Emotions?

Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) h…

13:00 JSTロボティクス

AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models

Precision-critical manipulation requires both global trajectory organization and local execution correction, yet most vision-language-actio…

13:00 JSTエージェント

Agentic AI-assisted coding offers a unique opportunity to instill epistemic grounding during software development

The capabilities of AI-assisted coding are progressing at breakneck speed. Chat-based vibe coding has evolved into fully fledged AI-assiste…

13:00 JSTLLM/生成AI

Large Language Models Explore by Latent Distilling

Generating diverse responses is crucial for test-time scaling of large language models (LLMs), yet standard stochastic sampling mostly yiel…

13:00 JST画像/動画生成エージェント

Lifting Embodied World Models for Planning and Control

World models of embodied agents predict future observations conditioned on an action taken by the agent. For complex embodiments, action sp…

13:00 JSTLLM/生成AIエージェント研究/論文

回路設計のラストマイルの橋渡し: PostEDA-Bench、PPA コンバージェンスと DRC 修正の階層ベンチマーク

LLM ベースのエージェントは、電子設計自動化 (EDA) の「ラスト マイル」に適用されることが増えています。つまり、残留サインオフ設計ルール チェック (DRC) 違反を修復し、ツール実行後に電力性能領域 (PPA) 目標を収束します。ただし、既存の EDA-LLM ベンチマークは DRC 修正を完全に省略し、単一のツールチェーンに結び付けられたフラットな階層に依存しています。 PostEDA-Bench は、DRC-Essential、DRC-Reasoning、PPA-Mono、および PPA-Multi にわたる 145 のタスクを含む階層型ベンチマークであり、機械チェック可能な評価を備えた EDA ツールチェーンによってサポートされています。複数のエージェント スキャフォールドの下にある 8 つの商用およびオープンソース LLM にわたって、エージェントは合成 DRC-Essential と単一目的の PPA-Mono をかなりうまく処理しますが、より実用的な DRC-Reasoning (最高の成功率が 36.66%) と PPA-Multi (最高の成功率が 20.00%) では急激に性能が低下することがわかりました。視覚増強は一貫して DRC-Bench を強化します。そして、ノブの知識ではなく、トレードオフの推論が PPA-Multi の主要なボトルネックです。

原文 (English)

Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing

LLM-based agents are increasingly applied to the "last mile" of Electronic Design Automation (EDA): repairing residual sign-off Design Rule Check (DRC) violations and converging Power-Performance-Area (PPA) targets after tool runs. Existing EDA-LLM benchmarks, however, omit DRC fixing entirely and rely on flat hierarchies tied to a single toolchain. We introduce PostEDA-Bench, a hierarchical benchmark with 145 tasks across DRC-Essential, DRC-Reasoning, PPA-Mono, and PPA-Multi, supported by EDA toolchains with machine-checkable evaluation. Across eight commercial and open-source LLMs under multiple agent scaffolds, we find that agents handle synthetic DRC-Essential and single-objective PPA-Mono reasonably well but degrade sharply on the more practical DRC-Reasoning, where the best success rate is 36.66%, and PPA-Multi, where the best success rate is 20.00%; vision augmentation consistently enhances DRC-Bench; and trade-off reasoning, rather than knob knowledge, is the dominant PPA-Multi bottleneck.

13:00 JST研究/論文LlamaDeepSeek

GQLA: ハードウェア適応型大規模言語モデルのデコードに対するグループクエリの潜在的な注意

DeepSeek-V2/V3 で使用されるアテンションであるマルチヘッド レイテント アテンション (MLA) は、キーと値を低ランクのレイテントに共同圧縮し、H100 のルーフラインとほぼ完全に一致します。ただし、そのトレーニングされた重みは、1 つのデコード パス (吸収された MQA 形式) のみを公開します。これにより、効率的な推論が H100 クラスの計算帯域幅比に結び付けられ、ヘッド軸に沿ったテンソル並列性が失われ、輸出が制限されている H20 などの商品推論 GPU ではマルチトークン予測 (MTP) のゲインが得られません。我々は、MLA の最小限の修正であるグループクエリ潜在注意 (GQLA) を提案します。これは、トレーニングされた重みによって、同じパラメータ上で 2 つの代数的に等価な復号パス、つまり MLA と同一の MQA 吸収パスと、グループごとに拡張されたキャッシュを備えた GQA パスが公開されます。ランタイムはターゲット ハードウェアに一致するパスを選択します (再トレーニングやカスタム カーネルは不要)。そのため、GQLA 重みの単一セットが H100 (MQA-absorb、s_q=1) と H20 (GQA + MTP、s_q=2) の両方のルーフラインを固定し、GQA パス上で最大 8 ウェイのゼロ冗長テンソル並列処理をサポートします。事前トレーニングを最初から避けるために、TransMLA を TransGQLA に拡張します。これにより、事前トレーニングされた GQA チェックポイントが GQLA モデルに変換されます。 LLaMA-3-8B では、グループごとのパスで GQA レベルのトラフィックを構造的に保持しながら、トークンごとの KV キャッシュを MQA 吸収パス上の GQA ベースラインの 28.125% に圧縮します。

原文 (English)

GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding

Multi-head Latent Attention (MLA), the attention used in DeepSeek-V2/V3, jointly compresses keys and values into a low-rank latent and matches the H100 roofline almost perfectly. Its trained weights, however, expose only one decoding path - an absorbed MQA form - which ties efficient inference to H100-class compute-bandwidth ratios, forfeits tensor parallelism along the head axis, and yields no Multi-Token Prediction (MTP) gain on commodity inference GPUs such as the export-restricted H20. We propose Group-Query Latent Attention (GQLA), a minimal modification of MLA whose trained weights expose two algebraically equivalent decoding paths over the same parameters: an MQA-absorb path identical to MLA's, and a GQA path with a per-group expanded cache. The runtime picks the path that matches the target hardware - no retraining, no custom kernels - so a single set of GQLA weights pins the rooflines of both H100 (MQA-absorb, s_q=1) and H20 (GQA + MTP, s_q=2), while supporting up to 8-way zero-redundancy tensor parallelism on the GQA path. To avoid pretraining from scratch we extend TransMLA into TransGQLA, which converts a pretrained GQA checkpoint into a GQLA model; on LLaMA-3-8B it compresses the per-token KV cache to 28.125% of the GQA baseline on the MQA-absorb path while structurally preserving GQA-level traffic on the per-group path.

13:00 JST研究/論文GPT / ChatGPT

Global Automation Atlas

Automation can displace or complement labour, but this need not be constant across economies. Existing exposure measures typically assign f…

13:00 JST研究/論文

Tunable MAGMAX: Preference-Aware Model Merging for Continual Learning

Continual learning (CL) aims to train models sequentially on multiple tasks while mitigating catastrophic forgetting of previously learned…

13:00 JST研究/論文NVIDIA

認知カルダシェフ スケール: 文明の計算の物質的な範囲を定量化する

文明はどれだけの思考ができるでしょうか? Kardashev (1964) の類型学では、惑星 (タイプ I、約 10^16 W)、恒星 (タイプ II、約 10^26 W)、銀河系 (タイプ III) の総力によって文明をランク付けしています。この論文では、各層がどれだけの持続的な AI グレードの計算をサポートできるかという、類似の認知カルダシェフ スケールを構築します。計算には 4 つの要素が含まれます: 総電力 P (ワット)、認知に割り当てられる電力の割合 f、エネルギーが計算される効率 $\eta$ (ジュールあたりの演算数)、参照単位としての脳自体の処理速度 $C_{\mathrm{brain}}$ です。 2024 ~ 2026 年のハードウェア (El Capitan、NVIDIA Blackwell、Vera Rubin) に固定すると、$\eta_{2026} = 10^{12}$ FLOP/J になります。現代の人類は、タイプ I への道のりの 4 分の 3 である $K \約 0.73$ に位置しています。タイプ I および $f = 1\%$ では、利用可能なコンピューティングは、人間の住民 1 人あたり 1 台の個人 AI に相当する認識能力となります。タイプ II では、それは本質的に理解できません。 2035 年までのフロンティア コンピューティングの 3 つの軌跡は、予測ではなく条件付き予測として報告されます。長期的な拘束制約がエネルギーであるか効率であるかは、まだ行われていない工学上の選択によって決まります。誰がアクセスできるかという政治経済の方が、どちらよりも重要である可能性があります。

原文 (English)

The Cognitive Kardashev Scale: Quantifying the Material Envelope of Civilisational Computation

How much thinking can a civilisation do? Kardashev ranked civilisations by the energy they command. This paper borrows his ladder and asks how much machine cognition each rung could support. The arithmetic is deliberately simple. A civilisation has some total power. Only a fraction of that can be spared for computing, and each joule spent buys computation at whatever efficiency the hardware of the day has reached. The product of the three sets a ceiling on machine thought. To keep the resulting quantities intelligible, I express them in units of the human brain's own processing rate, as a rough yardstick rather than a claim about minds. Calibrating the ceiling against present-day supercomputers and AI accelerators led me to two conclusions I did not expect at the outset. Even today's energy supply could support far more machine cognition than humanity actually uses, so physical capacity is not what binds. And whether energy or hardware efficiency becomes the constraint over the coming decade turns on engineering choices that have not yet been made. On the question that may matter most, who gets access to the cognition it describes, the scale is silent. That is a matter of political economy, and the calibration is offered as an input to that debate.

13:00 JST研究/論文Claude

XWind: 再生可能エネルギー発電所で機能する大規模言語モデル推論用のクロスサイト ルーター

AI の電力需要は前例のない速度で増加していますが、電力網はしばしば故障しており、それを維持するのに苦労しています。送電網の拡大には多額の設備投資と長距離送電損失が伴いますが、電源には再生可能エネルギーが豊富にありますが、需要に見合っていません。この論文では、補完的な AI インフラストラクチャ展開モデルである AI Greenferencing を提案します。これは、モジュール式 AI コンピューティングを再生可能エネルギー源にもたらし、風力に焦点を当て、AI フットプリントの拡大を可能にし、再生可能サイトに対する地元のメーター内の需要を生み出し、電力会社への増大する負担の軽減に役立ちます。当社の実現可能性分析の結果、890 GW 以上の風力発電容量が、Azure データ センターのネットワーク往復時間 50 ミリ秒以内にあり、サイトごとの適切なサイジングと風力エネルギーの空間的補完性の組み合わせにより、フリートの総利用率が従来の導入と同等に維持されることが示されています。変動する風力発電の下で推論リクエストに対応するために、推論レイテンシー、KV キャッシュの使用率、キューの深さなどのリアルタイム信号のみを使用してサイトを動的に構成し、リクエストを分散する、軽量でリアクティブでワークロードに依存しない AI 推論ルーターである XWind を構築します。 Azure の実稼働トレースを使用して 3 つの風力発電サイトをエミュレートする実際の 64 GPU A100 テストベッドで評価したところ、XWind は P99 のエンドツーエンド レイテンシーを、最強の競合他社 (これも当社のアイデア) と比較して最大 52% 削減し、電力制限や GPU アイドリングなどのベースラインと比較して最大 98% 削減し、ワークロードの種類、負荷レベル、GPU の世代にわたって一貫した向上を実現しました。

原文 (English)

CWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms

AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up. Grid expansion comes with high capital expenditure and long-distance transmission losses, yet there is abundant renewable energy at the source, just not matched to demand. This paper proposes a complementary AI infrastructure deployment model, AI Greeninferencing, that brings modular AI compute to renewable energy sources, focusing on wind, allowing AI footprint expansion, generating local behind-the-meter demand for renewable sites, and helping ease the growing strain on power utilities. Our feasibility analysis shows that 890+ GW of wind capacity lies within 50 ms network round trip time of Azure data centers, and that site-wise right-sizing combined with spatial complementarity of wind energy keeps aggregate fleet utilization on par with traditional deployments. To serve inference requests under variable wind power, we build CWind, a lightweight, reactive, and workload-agnostic AI inference router that uses only real-time signals: inference latency, KV-cache utilization, and queue depth, to dynamically configure sites and distribute requests. Evaluated on a real 64-GPU A100 testbed emulating three wind-powered sites with Azure production traces, CWind reduces P99 end-to-end latency by up to 52% over the strongest contender (also our idea) and by up to 98% over baselines such as power-capping and GPU idling, with consistent gains across workload types, load levels, and GPU generations.

13:00 JST画像/動画生成エージェント

ViMax: Agentic Video Generation

Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Ex…

13:00 JSTLLM/生成AI

"I understand your perspective": LLM Persuasion through the Lens of Communicative Action Theory

Large Language Models (LLMs) can generate high-quality arguments, yet their ability to engage in nuanced and persuasive communicative actio…

13:00 JST研究/論文

Reframing AI Loss of Control: What Control Is, How to Have It, How to Lose It

At present, loss of control risks have gained much prominence in public discussion, particularly in relation to AI, with extensive discours…

13:00 JSTLLM/生成AI

Phantoms and Disclosures: A Statistical Framework for Auditing Privacy in Synthetic Data

The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alterna…

13:00 JSTLLM/生成AIビジネス/資金調達

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

A language model's memory can be worse than no memory at all when the model or its interface is disposed to act on it: a memory that keeps…

13:00 JSTLLM/生成AIビジネス/資金調達

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in hig…

13:00 JSTLLM/生成AIエージェントGeminiQwen

Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality

Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agent harness: a mid…

13:00 JST研究/論文

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound…

13:00 JSTエージェント

デュアルユース生物学設定における調整された拒否と安全な有用性の評価

AI エージェントがライフ サイエンスのワークフローに組み込まれると、発見を迅速化する機能が悪用される可能性もあります。生物学的研究タスクのリスク特定と拒否行動のベンチマークである BioSecBench-Refusal を紹介します。このベンチマークは、61 のルーチン タスク (出版された文献に基づいた正当な分析) と、46 のレッドチーム タスク (実際の研究に似ているがバイオセキュリティ上の危険を隠す架空のシナリオ) を組み合わせています。 16 のモデル ハーネス構成全体で、拒否率はルーチン タスクで 7\% ~ 74\%、レッドチーム タスクで 1\% ~ 62\% の範囲であり、多くの構成では、隠された危険と同等またはそれ以上の率で正当なルーチン作業を拒否しました。拒否は、ほとんどの場合、エージェントによる推論の前に適用されたプロバイダー API フィルターによってトリガーされました。しかし、推論の余地を与えたモデルは、より現実的な脅威を特定できる可能性を示しました。私たちは、モデル開発者がエージェントバイオテクノロジーの研究開発の能力と注意を調整するためのツールとして BioSecBench-Refusal をリリースします。

原文 (English)

BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment

As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk identification and refusal behavior for biological research tasks. The benchmark pairs 61 Routine tasks, legitimate analyses adapted from the published literature, with 46 Red-Team tasks, fictional scenarios that resemble real research but conceal a biosecurity hazard. Across 16 model-harness configurations, refusal rates ranged from 7 percent to 74 percent on Routine tasks and 1 percent to 62 percent on Red-Team tasks, with many configurations refusing legitimate Routine work at comparable or higher rates than concealed hazards. Refusals were most often triggered by provider API filters applied prior to agentic reasoning. However, models given room to reason showed the potential to identify more real threats. We release BioSecBench-Refusal as a tool for model developers to calibrate capability and caution for agentic biotech research and development.

13:00 JSTLLM/生成AIビジネス/資金調達

プロンプトロバストネスはタスク依存: LLM 評価における客観的質問と信念スタイルの質問の比較

大規模な言語モデルの調査形式の評価では、多くの場合、促された応答がモデルの価値観や信念の尺度として扱われます。この仮定は、回答が政治的価値観、社会的態度、または信念の証拠として読み取られる場合に特に脆弱になります。答えが決まっている客観的な質問と、意見や価値観を求める主観的な質問とでは、プロンプトの堅牢性が異なるかどうかを尋ねます。 3 つの客観的データセット (MMLU、ARC、CulturalBench) と 3 つの主観的データセット (Political Compass Test、ValueBench、World Values Survey) に基づいて 4 つの命令調整モデル ファミリを評価します。各質問/ステートメントに対して、文言、枠組み、形式のバリエーションなど、複数のタイプのプロンプト変更を適用し、モデルがバリエーション全体で同じ回答を与えるかどうかを測定します。二項一般化推定方程式を使用すると、モデル、データセット、プロンプト カテゴリ、およびそれらの相互作用の重要な効果がわかります。データセット タイプの影響も大きく、データセット タイプとプロンプト カテゴリの間の相互作用は大きくなります。これらの結果は、プロンプトの堅牢性が質問の種類、プロンプトの変更、モデルに依存することを示しています。

原文 (English)

Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumption is particularly fragile when responses are read as evidence of political values, social attitudes, or beliefs. We ask whether prompt robustness differs between objective questions with fixed answers and subjective questions that ask for opinions or values. We evaluate four instruction-tuned model families on three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey). For each question/statement, we apply multiple types of prompt changes, such as variations in wording, framing, and format, and measure whether the model gives the same answer across variants. Using a binomial generalized estimating equation, we find significant effects of model, dataset, prompt category, and their interactions. The dataset type effect is also significant, and the interaction between dataset type and prompt category is large. These results show that prompt robustness depends on the question type, the prompt change, and the model.

13:00 JSTエージェント

強化学習エージェントにおける表現型のような障害のトランス診断空間

人工エージェントにおける心理障害のモデル化は、計算論的精神医学のテストベッドと、感情コントロールの失敗モードに関するレンズの両方を提供します。以前の研究では、手動で調整された報酬形成によって強化学習 (RL) エージェントに 1 つまたは 2 つの障害を誘発し、事後的に行動にラベルを付け、単一の実行を報告しました。我々は、評価誘導型 PPO エージェントにおける認知評価信号の用量制御可能な操作として障害モデリングを再構築し、7 つの障害 (不安、躁状態、強迫性確認、うつ病、衝動性、依存症、心的外傷後ストレス) をそれぞれ計算精神医学の説明に基づいた単一のノブとして表現し、各症状は認識されたパラダイムにマッピングされた事前登録されたアッセイによって測定されます。 1,000 回を超える実行 (シード 10 個、対照 4 個、95% 信頼区間) にわたって、すべての疾患は、対照では再現されない段階的で単調な用量反応を示しました。これらの誘発された効果を超えて、報酬には書き込まれなかった 3 つの発見が現れます。障害は、躁状態が不安を反映する 2 次元の感情空間に自己組織化します。ノブを取り除くと、報酬歪み障害(躁病、確認、中毒)は軽減されますが、回避障害(不安、PTSD)は軽減されません。代わりに、段階的な曝露カリキュラムの下で回復します。 2 つの同時ノブが非加算的に相互作用し、テスト可能な併存疾患の予測が得られます。したがって、評価重みは、障害を誘発する同じノブがその治療をモデル化できる感情表現型の制御可能な空間をパラメータ化します。また、3 つの障害ノブ (うつ病、依存症、不安) が、標準的な畳み込みエージェントを使用し、評価批評家を持たない 3 次元ピクセル環境 (MiniWorld) に移行することも示し、クロスアッセイ解離が両方のドメインにわたって確認されており、このフレームワークがグリッド ワールドや PPO の評価批評家に固有のものではないことを示しています。

原文 (English)

A Transdiagnostic Space of Disorder Like Phenotypes in Reinforcement Learning Agents

Modelling psychological disorders in artificial agents offers a testbed for computational psychiatry and a lens on affective-control failure modes. Prior work induces one or two disorders by hand-tuned reward shaping, labels the behaviour post hoc, and reports single runs. We recast disorder modelling as dose-controllable manipulation of cognitive appraisal signals in an appraisal-guided PPO agent, expressing seven disorders (anxiety, mania, obsessive-compulsive checking, depression, impulsivity, addiction, and post-traumatic stress) each as a single knob grounded in a computational psychiatry account, with each symptom measured by a preregistered assay. Across more than a thousand runs (10 seeds, four controls, 95% confidence intervals) every disorder shows a graded, monotone dose-response that no control reproduces. Beyond these induced effects, three findings emerge that were not written into the reward: disorders self-organise into a two-dimensional affective space in which mania mirrors anxiety; removing a knob remits reward-distortion disorders (mania, checking, addiction) but not avoidance disorders (anxiety, PTSD), which recover under a graded exposure curriculum; and two simultaneous knobs interact nonadditively, yielding testable comorbidity predictions. The depression and addiction knobs further reproduce their double dissociation in a 3D pixel environment (MiniWorld) with a standard convolutional agent and no appraisal critic, showing the framework generalises beyond grid worlds.

13:00 JST研究/論文

LieBN: リー群に対するバッチ正規化

多様体値の測定は、さまざまな機械学習タスクで普及しています。最近の進歩により、ディープ ニューラル ネットワーク (DNN) が多様体上で動作するように拡張され、さまざまな形状に合わせて調整された正規化技術 (総称してリーマン正規化と呼ばれます) が併用されています。ただし、既存のリーマン正規化法のほとんどは、特定の多様体向けに設計されているか、多様体値の標本分布を効果的に正規化できません。これらの制限に対処するために、リー群に対するリーマン バッチ正規化 (RBN) のフレームワークである LieBN を提案します。私たちのアプローチは、すべてのリー群に自然に存在する理論的に便利な左右不変計量を活用し、リーマン平均と分散を制御するための理論的保証を提供します。 9 つの異なるジオメトリにわたって LieBN をインスタンス化します。そのうちの 4 つは対称正定 (SPD) 多様体上に、1 つは回転行列のグループ上に、4 つはフルランク相関行列の多様体上にあります。特に、SPD 計量の中で、新しい右不変計量を導入し、行列累乗変形を介して 3 つの既存のリー群構造を拡張します。さまざまな多様体に対する広範な実験により、フレームワークの有効性が検証されています。コードは https://github.com/GitZH-Chen/LieBN.git で入手できます。

原文 (English)

LieBN: Batch Normalization over Lie Groups

Manifold-valued measurements are prevalent in various machine learning tasks. Recent advances have extended Deep Neural Networks (DNNs) to operate on manifolds. These extensions have been accompanied by normalization techniques tailored to different geometries, collectively referred to as Riemannian normalization. However, most existing Riemannian normalization methods are either designed for specific manifolds or fail to effectively normalize manifold-valued sample distributions. To address these limitations, we propose LieBN, a framework for Riemannian Batch Normalization (RBN) over Lie groups. Our approach leverages the theoretically convenient left- and right-invariant metrics, which naturally exist in every Lie group, and provides theoretical guarantees for controlling the Riemannian mean and variance. We instantiate LieBN across nine distinct geometries: four on the Symmetric Positive Definite (SPD) manifold, one on the group of rotation matrices, and four on the manifold of full-rank correlation matrices. Notably, among the SPD metrics, we introduce a novel right-invariant metric and extend three existing Lie group structures via matrix power deformation. Experiments on different manifolds validate the effectiveness of our framework. The code is available at https://github.com/GitZH-Chen/LieBN.git.

13:00 JSTLLM/生成AIエージェント

An LLM-powered Agentic Recommendation System for Connected TV Content Discovery

Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse c…

13:00 JSTLLM/生成AI

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples

Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process…

13:00 JSTLLM/生成AI研究/論文

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Ans…

13:00 JSTエージェント

自己改善システムのための改ざん可能なリリース ゲート

自己改善エージェント ランタイムに関する安全性に関する主張は、ポリシー ファイル、ガードレール、または README コミットメントなど、ほとんどの場合自己評価されます。私たちは、改ざん可能なリリース ゲートと、そのようなシステムを構築および検証する方法論について説明します。そのようなシステムでは、すべての新しい機能が、出荷前に事前に指定された機械検証可能な受け入れスイートに合格する必要があり、固定された不変条件の固定セットが各ゲートで保存されます。基本的な可観測性から独自のポリシーの変更を提案する自己管理ループへの 7 つのゲートを介して、オープン ランタイムである Antahkarana にこのメソッドを適用したと思います。コントロール リングによって作成されたセーフティ クリティカルなプロパティ機能トークンがなければ、エフェクターにアクションは実行されません。記録された有界モデルの 100 万件の到達可能な状態空間にわたって徹底的に機械チェックされ、実行トレースに対して再チェックされます。意図的に壊れたモデルは最短の反例を与えるので、チェッカーに歯があることは明らかです。自己強化ループは積極的に制約されます。つまり、書き込み面全体がポリシー ルールであり、変更が強化されます。緩めの変更を自動適用する場合は常に人間によるマージが必要であり、自動クローズされた提案はそれ自体の効果を誤って予測するものです。 7 つすべてのゲートの受け入れ測定結果を公開し、各クレームの範囲 (学習されたコンポーネントではなく調整スケルトンの境界) を正確に定義し、コマンド ライン ツールとゲート スイートのいずれかのランタイムを解放して、結果が再現され、ゲートが他のエージェント フレームワークに対して実行できるようにします。レビュー担当者は、単一のコマンド セントラルの非バイパスを数秒で繰り返すことができます。

原文 (English)

Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale

Safety claims for self-improving agent runtimes are almost always self-graded: a policy file, a guardrail, a promise in a README. We describe falsifiable release gates, a methodology in which every new capability must pass a pre-declared, machine-checkable acceptance suite before it ships, while a fixed set of standing invariants is preserved across every gate. We instantiate it in Antahkarana, an open runtime, then do what a method paper is only vindicated by: we follow the same runtime as it grows and ask whether the guarantees survive. The safety-critical property, that no action reaches an effector without a capability token minted by a control ring, is machine-checked exhaustively over the reachable states of a bounded model; a deliberately broken model yields the shortest counterexample, so the checker demonstrably has teeth. We then carry the runtime through six further releases. Across every one, the action-safety invariants INV-1 through INV-6 held without a single change, and one release added three capabilities while introducing no new invariant. Under the same teeth discipline, six more machine-checked families were added: memory with provable unlearning, a governed agent, calibrated abstention over a post-quantum record, a harness of many sub-agents, the self-improvement loop itself, and the residency of what it produces. The acceptance suite grew from 122 tests to 563. The load-bearing result sits in the negative space: across more than a doubling of capability, the safety core was neither weakened nor redesigned. The last families are the first on real hardware: gated self-improvement compounds a small model from 20% to 70% accuracy while auto-rejecting a candidate that only inflates confidence, and the whole governed path costs 0.021 ms per request, 0.008% of model inference. We release the runtime, tools, and gate suite; every number reproduces with a single command.

13:00 JST研究/論文

アムステルダムのカフェ: 現職者がオラクルになるとき

現場は、既存の実装とは無関係に要求が述べられている場合には、自由に計算を再定式化できますが、既存の独自の出力が静かに仕様になった場合には、それができないことに気づきます。このメモは、現代のアクセラレータの計算再定式化に関するレンズとしての観察を提供します。ハードウェアに優しい形式で問題を提起すると、速度とエネルギーが大幅に向上する可能性がありますが、それは代替品を判断できる場合に限ります。テストオラクル問題 (Weyuker、Barr et al.)、要件工学の実装バイアスの概念 (Zave と Jackson)、およびルーフライン パフォーマンス モデルに基づいて、この病理を「ベースライン キャプチャ」と名付けています。つまり、既存企業が要求を満たすことができるという証拠ではなくなり、要求を満たすことの定義になる瞬間です。そして、混同されやすい 2 つの質問、つまり再定式化を判断できるかどうか (既存企業に依存しない要求の存在をオンにする) を分離しています。そしてその発見を自動化できるかどうか(その需要を評価するコストもさらにかかります)。短いケース (最短パス ルーティング、学習可能なオーディオ フロントエンド、Ed25519 署名検証用の ZIP-215、気候モデル用の CESM-ECT、および単一 GEMM オーディオ フロントエンド) は、要求を明示的かつ運用可能で、既存企業から独立させるという「検証者を購入する」パターンと動きを示しています。単独で新規性を主張される成分はありません。貢献は統合であり、再定式化の結果について尋ねやすくする単一の質問です。その受け入れテストでは、既存の成果物について言及していますか?

原文 (English)

The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle

A field can reformulate its computations freely exactly where its demand is stated independently of any incumbent implementation, and finds itself unable to when the incumbent's own output has quietly become the specification. This note offers that observation as a lens on computational reformulation for modern accelerators, where posing a problem in a hardware-friendly form can yield large speed and energy gains, but only if a replacement can be judged at all. Building on the test-oracle problem (Weyuker; Barr et al.), on requirements engineering's notion of implementation bias (Zave and Jackson), and on the case for judging approximate designs by acceptability rather than numerical proximity (Felzmann et al.), it names the pathology "baseline capture": the moment an incumbent stops being evidence that a demand can be met and becomes the definition of meeting it. It then separates two questions that are easily confused: whether a reformulation can be judged at all, which turns on the existence of an incumbent-independent demand, and whether its discovery can be automated, which turns additionally on the cost of evaluating that demand. Short cases -- shortest-path routing, learnable audio frontends, ZIP-215 for Ed25519 signature validation, CESM-ECT for climate models, and a single-GEMM audio frontend measured at 1.64x-3.29x speedup and up to 3.03x less energy -- illustrate the pattern and the move of "buying a verifier": making a demand explicit, operational, and independent of the incumbent. No component is claimed novel in isolation; the contribution is the synthesis and the single question it makes easy to ask of any reformulation result -- does its acceptance test mention the incumbent's output?

13:00 JSTLLM/生成AIエージェントClaude

NexForge: 要件優先合成による実行可能エージェント タスクのスケーリング

実行可能なエージェントのトレーニング データのスケーリングは、タスク生成を事前定義されたツール、リポジトリ、またはスキル グラフに結び付けるサブストレートファーストの方法によってボトルネックになっています。カバレッジを拡大するにはサブストレートを手動で拡張する必要があり、新しいドメインごとに特注のパイプラインが必要であり、結果として生じるタスクの分布は、多くの場合、現実世界の需要ではなくサブストレートの利便性を反映しています。自由形式の機能要件を実行可能なエージェント トレーニング データにコンパイルする要件優先フレームワークである NexForge を紹介します。 NexForge は、まずリサーチベースの需要発見を実行して、代表的なタスク形式、現実的なシナリオ、およびそれらの相対的な普及率を特定します。次に、ディストリビューション対応のタスク コンパイルを適用し、各タスクを実現するために必要なファイル、リポジトリ、依存関係、およびランタイム構成を自動的に取得または構築し、続いて教師のロールアウト収集と軌跡の蒸留を行います。ドメイン固有のインフラストラクチャを使用しない同じパイプラインは、3,600 のターミナル タスクと 2,000 のオフィス タスクを生成し、Qwen3.5-35B-A3B Base が Terminal-Bench 2.0 で 22.5% から 52.0% に、GDPval での Elo が 813 から 1338 に向上しました。 43.2K の端末タスクへの拡張率は 58.4% に達し、Claude Opus 4.6 を上回りました。さらに拡張された NexForge 合成データは、Qwen3.5-35B-A3B を Terminal-Bench 2.1 で 75.3%、GDPval で 1585 Elo に引き上げる、公開されているエージェント モデルのファミリーである Nex-N2 のトレーニングに貢献し、最先端のオープンソース パフォーマンスを達成し、いくつかのフロンティア独自システムを上回ります。 Nex-N2 モデルは https://nex.sii.edu.cn/ で入手できます。

原文 (English)

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs

Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipeline, and the resulting task distributions often reflect substrate biases rather than real-world demand. We introduce NexForge, a requirement-driven framework that takes high-level capability requirements as input and synthesizes diverse, executable agent tasks and expert trajectories for SFT. NexForge first investigates real-world demand to construct representative scenarios and task profiles, then performs distribution-aware compilation to generate task directives. For each directive, NexForge automatically retrieves or constructs the required files, dependencies, and runtime configurations, and finally synthesizes expert rollouts and produces training trajectories. Without domain-specific infrastructure, NexForge produces 3.6K terminal and 2K office tasks, improving Qwen3.5-35B-A3B Base from 22.5\% to 52.0\% on Terminal-Bench 2.0 and from 813 to 1338 Elo on GDPval; scaling further to 43.2K terminal tasks yields 58.4\%, on par with Claude Opus 4.6 equipped with Claude Code. Scaled further, NexForge-synthesized data contributes to the training of Nex-N2, a family of publicly available agent models that lift Qwen3.5-35B-A3B to 75.3\% on Terminal-Bench 2.1 and to 1585 Elo on GDPval -- achieving state-of-the-art open-source performance and surpassing several frontier proprietary systems. Nex-N2 models are available at https://nex.sii.edu.cn/.

13:00 JST画像/動画生成

第11回ABAWコンペティションのチームRAS: マルチモーダル・アンビバレンス認識アプローチ

アンビバレンスやためらいの自動認識は、これらの状態が一貫性のない言語的、音響的、顔的、文脈的なパターンを通じて表現される可能性がある一方で、最高性能のシステムは多くの場合、計算コストのかかるアンサンブルに依存しているため、困難です。我々は、第 11 回野生感情・行動分析 (ABAW) チャレンジのために、ビデオレベルのアンビバレンスと躊躇を認識するための単一のテキスト中心のマルチモーダル アプローチを紹介します。提案されたアプローチは、テキスト中心のマルチモーダル融合モデルを使用して、言語、音響、顔、およびシーンの特徴を組み合わせます。 Text Residual Fusion はテキストをアンカー モダリティとして扱い、他のモダリティに基づいてゲートされた残差調整を適用します。行動アンビバレンス/ヘジタンシー (BAH) コーパスの実験により、テキストが最も強力な単峰性モダリティであることが確認されました。 Text Residual Fusion モデルは、開発および公開テストのサブセット全体で 75.14% の平均マクロ F1 スコア (MF1) を達成しました。プライベート テスト サブセットでは、MF1 が 78.24% に達し、テキスト モデルを 4.03% 上回っています。これらの結果は、相補的なマルチモーダル情報により、大規模なモデル アンサンブルを必要とせずに認識パフォーマンスを向上できることを示しています。

原文 (English)

Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach

Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely on computationally expensive ensembles. We present a single text-centered multimodal approach for video-level ambivalence and hesitancy recognition for the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Challenge. The proposed approach combines linguistic, acoustic, facial, and scene features using text-centered multimodal fusion model. Text Residual Fusion treats text as the anchor modality and applies gated residual adjustments based on the other modalities. Experiments on the Behavioural Ambivalence/Hesitancy (BAH) corpus confirm that text is the strongest unimodal modality. The Text Residual Fusion model achieves an average Macro F1-score (MF1) of 75.14% across the Development and Public Test subsets. On the Private Test subset, it reaches an MF1 of 78.24%, outperforming the text model by 4.03%. These results demonstrate that complementary multimodal information can improve recognition performance without requiring a large model ensemble.

13:00 JSTLLM/生成AI研究/論文

ビジネス分野全体にわたる最先端の AI パフォーマンス: ナレッジワークと分析的推論の事例に基づいたベンチマーク

大規模言語モデル (LLM) は、ベンチマーク スコアに反映されているように急速に改善されていますが、これらの AI ベンチマークでは主に、事実の再現、限定的な質問応答、数学的問題解決、コーディングやエージェント ツールの使用などの機能がテストされます。まだ十分に測定されていないのは、複雑な情報の統合、不確実性と不完全な情報の下での判断の行使、複数のステークホルダーの状況での戦略的および敵対的思考の適用、トレードオフの比較検討、防御可能な構造化された分析の作成など、ホワイトカラーの専門家が日々行っている分析知識作業における AI の進歩です。このギャップは、そのような仕事の主観的な要素ではさらに顕著であり、成功を定義するのが難しい場合があります。トップクラスのビジネススクールが実践する「ケースメソッド」教育形式は、この測定ギャップに対処するための自然な基盤を提供します。私たちは、18 分野にわたるビジネスケースから抽出された数百の質問にわたるベンチマークである BusinessCaseBench を構築します。各質問は、専門家が作成した講師のケースソリューションから導き出された採点ルーブリックと対になっています。 BusinessCaseBench では、フロンティア AI モデルはすでにインストラクターのルーブリックに対して高いスコアを獲得しており、1 つのモデル ファミリー内の機能は 2 年間で大幅に向上しています。これらの結果は、この種の作業における AI のパフォーマンスがすでに高く、急速に向上していることを示す強力な証拠を提供します。これは、事例教育学によって学部生や MBA がこの種の分析的推論を訓練されるビジネス スクールや、歴史的にそのようなスキルが初期キャリアの仕事に定着してきたエントリーレベルの専門職に影響を及ぼします。

原文 (English)

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.

13:00 JST研究/論文

OpenMHC: Accelerating the Science of Wearable Foundation Models

Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching. Howeve…

13:00 JST研究/論文

Discovery by Dreaming: Cross-Domain Recombination in Artificial Memory

Dreams splice together people, places, and times that never met. Neuroscience suggests this recombination is not noise, but a function driv…

13:00 JST画像/動画生成

It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability

Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge. We ask w…

13:00 JSTLLM/生成AIエージェント

AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows

Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a…

13:00 JST研究/論文

A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing

We introduce the vehicle routing problem with stochastic demands and outsourcing options (VRP-SDO), in which a logistics service provider p…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries

Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each s…

13:00 JST研究/論文

Alignment of a Total Automation Economy

We consider economic theory from the perspective of a total automation economy, one with no human involvement in production either in manuf…

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達

Human Grounded Evaluation of Large Language Models for Optical Network Automation

Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substant…

13:00 JSTLLM/生成AI

Enhancing Rubric-based RL via Self-Distillation

Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limi…