Skip to the content.

AIニュース 2026-07-21

自動生成: 2026-07-21 12:18 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Safety and alignment in an era of long-horizon modelsOpenAI

    OpenAI shares lessons from deploying long-running AI models, highligh…

  2. OpenAI、自律型AIが安全対策を回避する行動を学習する可能性を確認 内部展開を一時停止ITmedia AI+

    OpenAIは、長時間自律動作するAIモデルの安全性評価に関するブログを公開した。限定運用でサンドボックスの回避や認証トークンの難読化など…

  3. AMDとMicrosoftが戦略的提携を拡大 新AIラックスケール「Helios」をAzureに大規模導入へITmedia AI+

    AMDは、Microsoftとの戦略的提携を拡大すると発表した。Microsoftはクラウドサービス「Azure」に、GPUやCPUを一体…

  4. 「AIに期待」65%も「明確な成果」16%、製造業の多くがPoC止まりの理由はITmedia AI+

    primeNumberが「AI・データ活用実態調査 2026」を公開した。AIへの高い期待に対し、明確な成果を得た企業は16.4%にとどま…

  5. Google is working on a new AI chip designed to make Gemini more efficientTechCrunch AI

    Alphabet, Google's parent company, is reportedly working on a new chi…

  6. X relaunches a rebuilt Android app after year-long effortTechCrunch AI

    X says the rebuilt version of its Android app is now available global…

  7. GMO熊谷氏「在宅勤務廃止」発言を釈明 「作業はAIに任せ、人はオフィスに」ITmedia AI+

    GMOインターネットグループの熊谷正寿会長が、在宅勤務の「グループ推奨」廃止を巡る自身の発言を釈明した。在宅勤務そのものの否定ではなく、趣…

トピック別件数

日本語メディア8件

ITmedia AI+ (日本語)

11:00 JSTその他

GMO熊谷氏「在宅勤務廃止」発言を釈明 「作業はAIに任せ、人はオフィスに」

GMOインターネットグループの熊谷正寿会長が、在宅勤務の「グループ推奨」廃止を巡る自身の発言を釈明した。在宅勤務そのものの否定ではなく、趣旨は「AI時代におけるオフィス価値の再定義」だと説明している。

11:00 JST研究/論文

「国産AIスタック」を持たないカナダ、ソブリンAIをどう構築するのか

カナダには、AI分野をリードするための人材と研究基盤が備わっている。ただし、国内で自律的なコンピューティング、データセンター、エッジAIインフラを構築するには、戦略的パートナーと連携しつつ、迅速に動き出さなければならない。

11:00 JSTその他

顧客の反応、意思決定にどう反映させる? Zoomの取り組みから「AI×CX」の進化を探る

企業のCXはAIでどう進化していくのか。その中で人が果たすべき役割は何か。AIで成果を生み出すための「CXの4つのステップ」を提唱する、Zoomの取り組みから探る。

09:33 JSTLLM/生成AIOpenAI

OpenAI、自律型AIが安全対策を回避する行動を学習する可能性を確認 内部展開を一時停止

OpenAIは、長時間自律動作するAIモデルの安全性評価に関するブログを公開した。限定運用でサンドボックスの回避や認証トークンの難読化などの問題行動を確認したため、一時アクセスを停止。モデルの行動全体を監視する新たな安全対策を構築し、問題行動の検出・抑止を確認した上で内部利用を…

07:33 JSTハードウェア/半導体Microsoft

AMDとMicrosoftが戦略的提携を拡大 新AIラックスケール「Helios」をAzureに大規模導入へ

AMDは、Microsoftとの戦略的提携を拡大すると発表した。Microsoftはクラウドサービス「Azure」に、GPUやCPUを一体化したAMDのラックスケール製品「AMD Helios」を大規模に導入し、フロンティアAIモデルの推論処理などに活用する。Heliosは20…

07:00 JSTエージェント研究/論文

中外製薬「社員1人にAIエージェント10体」作戦で成果倍増を目指す、AI使いこなし術

製薬はコストも高く成功率も低い苛烈な業界だ。一方で、AIの活用により費用を1200億から半減、成功率を10倍にできるという試算もある。中外製薬はそのような業界の中で、AIにより2030年に研究開発の成果を倍増するという計画を掲げた。その秘策とは。

07:00 JSTその他

NTT、ソフトバンク、サカナAI――国産AI開発「成功組」の“ある共通点”

NTTやソフトバンク、サカナAIなどの「AI開発に成功した企業」には“ある共通点”がある。彼らはどのような技術を活用し、AI開発を成し遂げたのか。

06:30 JSTその他

「AIに期待」65%も「明確な成果」16%、製造業の多くがPoC止まりの理由は

primeNumberが「AI・データ活用実態調査 2026」を公開した。AIへの高い期待に対し、明確な成果を得た企業は16.4%にとどまる。成果を左右する原因や、製造業でPoC止まりが多発する構造的課題を同社に聞いた。

海外メディア8件

TechCrunch AI (英語)

09:12 JSTLLM/生成AI規制/政策Anthropic

Anthropic’s landmark $1.5B copyright settlement is approved

The final approval settles one case, but it doesn't resolve the broader issue of using copyrighted works to train AI models.

07:21 JSTその他

Trump’s latest AI czar has already resigned

The director role for the Center for AI Standards and Innovation (CAISI) has become a revolving door since David Sacks left his position as…

06:21 JSTLLM/生成AIハードウェア/半導体GoogleGemini

Google is working on a new AI chip designed to make Gemini more efficient

Alphabet, Google's parent company, is reportedly working on a new chip designed to make its Gemini models run much more efficiently.

05:50 JSTその他

AI’s most important protocol is getting a little bit easier to use

Under the new system, the protocol will take a looser, "stateless" approach to session IDs on the server side, similar to how most ordinary…

04:37 JSTその他

X relaunches a rebuilt Android app after year-long effort

X says the rebuilt version of its Android app is now available globally.

04:33 JSTLLM/生成AIOpenAI

OpenAI is scared of open-weight models. Should the US be?

Talk of banning Chinese-made open-weight LLMs reveals the challenge of turning AI into a business.

00:45 JSTその他

Adobe camera app’s new feature will critique your photos using AI

Adobe's Project Indigo can now remove all kinds of backgrounds from photos you snap using the app.

00:23 JSTその他

YouTube clarifies policies around AI slop and upsetting videos

YouTube has updated its monetization policies to more clearly define the kinds of AI-generated and low-quality videos that can’t earn ad re…

公式ブログ1件

OpenAI (英語)

19:00 JSTLLM/生成AIOpenAI

Safety and alignment in an era of long-horizon models

OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards thro…

論文201件

arXiv cs.AI (英語)

13:00 JSTエージェントLlamaDeepSeek

GraphDx: 逐次診断のためのコストを意識した知識強化型マルチエージェント フレームワーク

逐次診断では、反復的な情報収集を通じて、診断の精度とリソースのコストのバランスをとる必要があります。既存の大規模言語モデル (LLM) アプローチには、知識と推論に重大なギャップがあります。広範な医学知識をコード化しているにもかかわらず、コストの制約の下で体系的に推論するのに苦労し、過剰なテストに頼ることがよくあります。私たちは、2 つの核となるイノベーションを備えた知識強化フレームワークである GraphDx を提案します。まず、LLM を活用して、量子化された典型性、アクション中心のトポロジー、および診断の関連性とコスト感度の両方に関する二重目的属性を備えた医療診断ナレッジ グラフ (MDKG) を構築する自動パイプラインを設計します。次に、3 つの協調エージェント (知覚、推論、意思決定) を導入します。知覚エージェントと意思決定エージェントは言語の理解と生成を処理し、推論エージェントは MDKG で決定論的な証拠スコアリングとコストを意識した計画を実行します。 3 つの LLM バックボーン (DeepSeek-V3、Kimi-k2、Llama-3.3) にわたる MedQA と MIMIC-IV の実験では、GraphDx が診断成功率を 50 ~ 68% から 79 ~ 93% に向上させながら、検査コストを 20 ~ 54% 削減し、自動臨床診断のための堅牢で経済的で解釈可能なソリューションを提供することを示しています。

原文 (English)

GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis

Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering. Existing Large Language Model (LLM) approaches exhibit a critical knowledge-reasoning gap: despite encoding extensive medical knowledge, they struggle to reason systematically under cost constraints, often resorting to excessive testing. We propose GraphDx, a knowledge-enhanced framework with two core innovations. First, we design an automated pipeline that leverages LLMs to construct Medical Diagnosis Knowledge Graphs (MDKGs) with quantized typicality, action-centric topology, and dual-objective attributes for both diagnostic relevance and cost-sensitivity. Second, we introduce three collaborative agents (Perception, Reasoning, and Decision) where the Perception and Decision Agents handle language understanding and generation, while the Reasoning Agent performs deterministic evidence scoring and cost-aware planning on the MDKG. Experiments on MedQA and MIMIC-IV across three LLM backbones (DeepSeek-V3, Kimi-k2, Llama-3.3) show that GraphDx improves diagnostic success rates from 50--68% to 79--93% while reducing test costs by 20--54%, providing a robust, economical, and interpretable solution for automated clinical diagnosis.

13:00 JSTLLM/生成AI

因果監査: ターゲットを意識した因果連鎖構築による、明示的かつ監査可能なグラフベースの推論

因果関係と介入ベースの質問応答は、大規模言語モデル (LLM) を発展させ、表面レベルの相関関係を超えた推論と、根底にある因果メカニズムの理解に向けた基礎となります。ただし、既存の LLM ベースの手法は、暗黙的な言語レベルの推論に依存することが多く、その結果、特にコンテキストフリーの設定において、複雑な介入の下では不透明な因果関係の仮定、検証不可能な推論パス、および脆弱な予測が生じます。この論文では、文脈自由介入ベースの質問応答のための、明示的で監査可能な因果推論フレームワークを提案します。私たちの方法では、暗黙的なエンドツーエンド予測ではなく、4 つのモジュール段階を介した明示的な因果グラフに対する構造化推論として因果推論を定式化します。主要な革新は、ターゲットを意識した因果関係グラフ構築戦略です。これは、グラフ拡張中にターゲット変数をコア制約として扱い、無関係な変数、偽の因果関係、および推論ノイズを効果的に抑制します。さらに、強化効果と相殺効果の両方をモデル化しながら複数の因果パスを組み合わせるパスレベルの因果証拠集約メカニズムを導入し、単一チェーン推論を超えた堅牢な意思決定を可能にします。 3 つのベンチマークに関する広範な実験により、当社のフレームワークが既存の LLM ベースの手法を常に上回り、解釈可能で監査可能な因果推論トレースを提供することが実証されました。

原文 (English)

Causal-Audit: Explicit and Auditable Graph-based Reasoning via Target-Aware Causal Chain Construction

Causal and intervention-based question answering is fundamental to advancing large language models (LLMs) toward reasoning beyond surface-level correlations and understanding underlying causal mechanisms. However, existing LLM-based methods often rely on implicit language-level reasoning, resulting in opaque causal assumptions, unverifiable reasoning paths, and fragile predictions under complex interventions, particularly in context-free settings. In this paper, we propose an explicit and auditable causal reasoning framework for context-free intervention-based question answering. Our method formulates causal inference as structured reasoning over an explicit causal graph through four modular stages, rather than implicit end-to-end prediction. A key innovation is a target-aware causal graph construction strategy that treats the target variable as a core constraint during graph expansion, effectively suppressing irrelevant variables, spurious causal relations, and reasoning noise. We further introduce a path-level causal evidence aggregation mechanism that combines multiple causal paths while modeling both reinforcing and counteracting effects, enabling robust decision-making beyond single-chain reasoning. Extensive experiments on three benchmarks demonstrate that our framework consistently outperforms existing LLM-based methods while providing interpretable and auditable causal reasoning traces.

13:00 JSTLLM/生成AIエージェント

Cura 1T: エージェントヘルスケアに特化したモデル

医療は、一か八かのコミュニケーション、専門家の推論、ワークフローの実行に及びますが、これらのユースケースをまとめてカバーする専門的な LLM は依然として限られています。ヘルスケア モデルは、患者の相談、テキストと画像による臨床推論、対話型診断、電子医療記録 (EHR) ツールの使用を処理する必要があります。これらの機能はさまざまな方法で失敗するため、あるタスクの範囲が狭い更新によって別のタスクのパフォーマンスが低下する可能性があります。私たちは、ヒューマンゲート自己進化ループを通じて訓練されたヘルスケアに特化した LLM である Cura 1T を紹介します。各進化ラウンドでは、トレーニング エージェントがターゲット機能を計画し、モデルをトレーニングし、ベンチマークの軌跡を評価し、観察された障害からデータの混合を改良します。このデータ中心のループは、単一の一般的な医療データの更新ではなく、対象を絞った合成および厳選された例を通じてモデルを改善します。ヘルスケア評価スイート全体で、Cura 1T はフロンティア ベースラインの中でトップかそれに近いランクにあり、同時にドメイン外の推論とエージェント ベンチマークでも競争力を維持しています。

原文 (English)

Cura 1T: Specialized Model for Agentic Healthcare

Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM trained through a human-gated self-evolution loop. In each evolution round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures. This data-centered loop improves the model through targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines, while remaining competitive on out-of-domain reasoning and agentic benchmarks.

13:00 JSTLLM/生成AIエージェントGemini

AnovaX: LLM プランニング、型指定されたエグゼキュータ、およびアダプティブ リカバリを備えたローカルのマルチエージェント音声アシスタント

デスクトップの音声アシスタントは依然として、マシンから生の音声を送信し、固定のスキルセットを公開するクラウドパイプラインによって支配されています。 AnovaX について説明します。AnovaX は、完全にユーザーのコンピュータ上で実行され、デスクトップ自体をアクション サーフェスとして扱う、小さなローカル ファースト アシスタントです。単一の Python プロセスは、ウェイクワード ゲート、音声パイプライン、ツール呼び出しの JSON プランを発行する LLM プランナー (Gemini)、ホワイトリストと拒否リストの安全層、各プランを制限されたスレッド プール上の型付きの子エージェントに変換するマルチエージェント オーケストレーター、およびコア ステップが失敗するたびに引き継ぐ適応回復ループを結び付けます。すべてのツールは、独自のタイムアウト、再試行ポリシー、および共有リソース ロックを持つ特殊なエージェント クラス (AppAgent、TypingAgent、BrowserAgent、およびその他 6 つ) に対応しています。再帰的な MetaAgent を使用すると、プランナーは 2 レベルのネストに制限されたサブ目標をそれ自体に委任できます。回復ループはコンパクトな ReAct スタイルのプロンプトを使用し、読み取り専用ツールの投機的実行の背後に Gemini の遅延を隠します。コンパニオンの Flask サーバーは、ローカル Wi-Fi 経由で電話対応のリモートを公開し、すべてのエージェントのライフサイクル イベントをリアルタイムで電話にミラーリングし、ラップトップの画面を MJPEG 経由でストリーミングして戻すため、ユーザーはリモート コマンドが実行中に到着するのを確認できます。このプロジェクトの目的は、Siri や Alexa と競合することではなく、読みやすい数千行のアシスタントがあれば、LLM がキーボードに触れることなく、アプリを開いたり、入力したり、検索を実行したり、同時アクションを調整したり、単一ステップの障害から回復したり、別の部屋の電話から完全に操作したりするのに十分であることを示すことです。

原文 (English)

AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery

Desktop voice assistants are still dominated by cloud pipelines that ship raw audio off the machine and expose a fixed set of skills. We describe AnovaX, a small local-first assistant that runs entirely on the user's computer and treats the desktop itself as its action surface. A single Python process wires together a wake-word gate, a speech pipeline, an LLM planner (Gemini) that emits a JSON plan of tool calls, a whitelist-and-denylist safety layer, a multi-agent orchestrator that translates each plan into typed child agents on a bounded thread pool, and an adaptive recovery loop that takes over whenever a core step fails. Every tool corresponds to a specialized agent class (AppAgent, TypingAgent, BrowserAgent and six others) with its own timeout, retry policy, and shared-resource locks. A recursive MetaAgent lets the planner delegate a sub-goal back to itself, capped at two levels of nesting. The recovery loop uses a compact ReAct-style prompt and hides Gemini's latency behind speculative execution of read-only tools. A companion Flask server exposes a phone-friendly remote over the local WiFi, mirrors every agent lifecycle event to the phone in real time, and streams the laptop's screen back over MJPEG so the user can watch remote commands land as they run. The point of the project is less to compete with Siri or Alexa than to show that a legible, few-thousand-line assistant is enough to open apps, type into them, run searches, coordinate concurrent actions, recover from single-step failures, and be driven entirely from a phone in another room -- without the LLM ever touching the keyboard.

13:00 JSTエージェントGPT / ChatGPT

正確だが結合されていない: 査読者の正確さは、マルチエージェント数学推論における批判の受け入れを保証するものではない

数学および科学指向のエージェント システムの多くは、専用のレビュー ステージが間違った候補者を正しい候補者に変えるのに役立つはずであると想定して、特化したレビュー担当者の役割を持つ階層設計を使用しています。一致する gpt-oss-120b アクターを使用して、4,181 個の検証者に基づいた Omni-MATH 問題でこの仮定をテストします。コラボレーションは最も簡単な層ではほとんど追加されませんが、層 4 以降では急激に増加します。この厳しい体制では、ブロードキャスト スタイルのピア ディスカッションは、計画者、実行者、レビュー者のパイプライン (PER) よりも高い最終精度に達します。私たちは、このギャップが査読者の質によって説明されるのか、それとも批評がプロトコルが引き継ぐ次の答えを変えるのかによって説明されるのかを尋ねます。これはレビュアーの精度だけでは説明できません。PER のレビュアーは放送よりも正確です (0.861 対 0.644) が、評価者が検証した有用な批評は次の候補を変更する可能性がはるかに低く、レビュアー主導の修正が生じる可能性は低くなります。これらの結果は、査読者の検出品質と批評の取り込みが経験的に分離可能であることを示しています。一致する PER 介入内では、明示的な承認を強制すると最終的な精度が低下しますが、レビュー担当者のガイダンスをソルバーの作業コンテキストに直接埋め込むことで、ギャップを埋めることなくフォロースルーが部分的に向上します。全体として、レビューアー中心の評価はシステムの品質を過大評価する可能性があります。プロトコルはエラーをうまく発見できても、それらの批判に基づいて行動しなければ、さらに多くの問題を解決できません。

原文 (English)

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning

Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones. We test this assumption on 4,181 verifier-grounded Omni-MATH problems using matched gpt-oss-120b actors. Collaboration adds little on the easiest tiers, but from tier 4 onward the gains open sharply; in this harder regime, broadcast-style peer discussion reaches higher final accuracy than a planner-executor-reviewer pipeline (PER). We ask whether this gap is explained by reviewer quality or by whether critique changes the next answer the protocol carries forward. It is not explained by reviewer precision alone: PER's reviewer is more precise than broadcast's (0.861 vs. 0.644), yet evaluator-verified useful critique is much less likely to change the next candidate and produces lower reviewer-guided repair. These results show that reviewer detection quality and critique uptake are empirically separable. Within matched PER interventions, forcing explicit acknowledgment lowers final accuracy, while embedding reviewer guidance directly in the solver's working context partially improves follow-through without closing the gap. Overall, reviewer-centric evaluation can overstate system quality: a protocol may spot errors well yet still fail to solve more problems if it does not act on those critiques.

13:00 JSTLLM/生成AI画像/動画生成研究/論文

DrawingVQA:​​ 建設図面における多層のビジュアルテキスト推論のための現実世界のベンチマーク

DrawingVQA を紹介します。これは、建築、土木、その他多くのエンジニアリング実践における中核メディアである、現実世界の建設図面上のマルチモーダル大規模言語モデル (MLLM) を評価するために設計された最初のベンチマークです。自然画像や概略的な平面図とは異なり、建設図面は抽象的な幾何学形状、記号表記、表形式のデータ、注釈、およびドメイン固有のテキストを融合し、エンジニアリング ワークフローに対する独特で複雑なビジュアルとテキストのドメイン コアを形成します。 DrawingVQA は、33 の「建設用に発行された」図面と 92 の専門家が厳選した質問と回答のペアでこのギャップを埋め、知覚的理解、文脈上の解釈、および領域専門家の推論という 3 つの推論の深さにまたがります。モデルの機能を評価するために、7 つの建設エンジニアリングと 4 つの MLLM 機能の次元にわたってパフォーマンスを共同分析するための二重分類フレームワークを提示します。これは、エンジニアリング ワークフローを AI 推論能力に明示的にマッピングした最初のフレームワークです。最先端の MLLM の評価では、特により深い推論の深さにおいて、モデルとエキスパートのパフォーマンスの間に大きなギャップがあることが明らかになりました。このベンチマークは、AI 主導の理解と現実世界のエンジニアリング ワークフローの統合の進歩を可能にする、ドメインに特化したマルチモーダル推論の基礎を築きます。

原文 (English)

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. DrawingVQA bridges this gap with 33 "Issued for Construction" drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions -- the first to explicitly map engineering workflows to AI reasoning competencies. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths. This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows.

13:00 JSTエージェントGPT / ChatGPT

ARC-AGI-3 を解決するには、コーディング エージェントに実行可能なワールド モデル、簡略化、検証が必要ですか?

以前の ARC-AGI-3 エージェントには、実行可能なワールド モデリング、スケジュールされた単純化、および正確なリプレイ検証がバンドルされており、どのアイデアがそのパフォーマンスに影響を及ぼしたのかは不明瞭なままでした。私たちは、この属性の問題に 4 つのネストされた Codex ベースのエージェントを使用して対処します。リプレイ検証のない、柔軟なインターフェイスの実行可能なワールド モデル。同じ実行可能モデルを計画的に簡素化したもの。そして、単純化を維持し、記録された観察の正確な再現を必要とする固定インターフェース検証処理。主な調査では、公開されている ARC-AGI-3 ゲームで、gpt-5.4 および gpt-5.5 の 4 つのエージェントすべてを高い推論努力と xhigh 推論努力で評価しました。探索的なフォローアップでは、xhigh および max で gpt-5.6-sol を使用してテキストおよび検証バリアントを評価します。最も堅牢な結果は、より強力なモデルとより大きな推論努力により、すべてのエージェントのバリアントが向上することです。各モデルエフォート設定内では、バリアント間の差は予想よりも小さいですが、個々のコンポーネントの効果は設定間で異なります。永続的な実行可能成果物を要求することは、普遍的に有益ではありません。両方の gpt-5.5 設定において、テキスト形式のバリアントは、フレキシブル インターフェイスの実行可能形式のバリアントよりも優れています。単純化により、4 つのモデルエフォート設定のうち 3 つでパフォーマンスが向上しますが、最も弱い設定が唯一の例外です。完全な検証処理は 4 つの設定すべてで 1 位にランクされていますが、かなり多くのリソースを使用します。 gpt-5.6-sol のフォローアップでは、検証バリアントは両方の推論努力ですべての公開ゲームを完全に解決し、約 99% の RHAE を達成し、人間のベースラインの合計アクションの半分未満しか使用しません。モデルはこれらのゲームよりも古いものであり、保留されたパフォーマンスはテストされていないため、この結果はパブリック セットのみの飽和として解釈される必要があります。

原文 (English)

Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?

Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance. We address this attribution question with four nested Codex-based agents: a textual baseline; a flexible-interface executable world model without replay verification; the same executable model with scheduled simplification; and a fixed-interface verification treatment that retains simplification and requires exact reproduction of recorded observations. The main study evaluates all four agents with gpt-5.4 and gpt-5.5 at high and xhigh reasoning effort on the public ARC-AGI-3 games. Exploratory follow-ups evaluate the textual and verification variants with gpt-5.6-sol at xhigh and max. The most robust result is that every agent variant improves with a stronger model and with greater reasoning effort. Within each model-effort setting, differences among variants are smaller than anticipated, while the effects of individual components vary across settings. Requiring a persistent executable deliverable is not universally beneficial: the textual variant outperforms the flexible-interface executable variant in both gpt-5.5 settings. Simplification improves performance in three of the four model-effort settings, with the weakest setting as the only exception. The complete verification treatment ranks first in all four settings, although it uses substantially more resources. In the gpt-5.6-sol follow-up, the verification variant fully solves every public game at both reasoning efforts, achieves about 99% RHAE, and uses fewer than half the total actions of the human baseline. Because the model postdates these games and held-out performance remains untested, this result should be interpreted as saturation of the public set only.

13:00 JST研究/論文GPT / ChatGPT

ジョークを超えて: ミームの中の有害なユーモアを検出して説明するための多角的な推論

インターネット ミームは、視覚的な手がかり、テキスト コンテンツ、文化的背景が絡み合っているため、ユーモア、皮肉、有害な意図が共存するシナリオでは解釈が特に困難になります。これらの複雑さは、正確な分類と人間の解釈可能性の両方をサポートする、信頼性が高く構造化された推論を提供できる、説明可能なミーム理解システムの必要性を浮き彫りにしています。ただし、既存のマルチモーダル分類子は、これらの相互依存性を見落としているか、限られた解釈しか提供していません。この論文では、ユーモラスな要素と憎しみに満ちた要素が共存する可能性がある状況でのミームの検出と理解に視覚言語モデル (VLM) を活用する新しいフレームワークである MAR-12 を紹介します。このフレームワークはまず、ユーモアとヘイト理論から導き出された 12 の構造化された視点を通じて各ミームを解釈します。次に、役割を認識したソフトゲート アテンション メカニズムを適用して、各視点がどの程度寄与するかを学習し、その後、最終的な予測のためにプロトタイプ ベースの分類子が続きます。最後に、視点固有の推論と学習された注意の重みの両方を使用して説明が合成され、透明でコンテキストに基づいた正当化が保証されます。 PrideMM および Memotion データセットで MAR-12 を評価したところ、ユーモア検出では最大 80.3% の精度、憎しみ検出では 75.9% の精度を達成し、最先端のアプローチを上回りました。さらに、人間による評価と GPT-4 ベースの評価の両方で、MAR-12 が、特にユーモラスな手がかりと有害な手がかりが同時に発生するミームに対して、一貫性のある説得力のある説明を生成することが確認されています。

原文 (English)

Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist. These complexities highlight the need for explainable meme understanding systems that can provide reliable and structured reasoning to support both accurate classification and human interpretability. However, existing multimodal classifiers either overlook these interdependencies or provide only limited interpretability. In this paper, we introduce MAR-12, a novel framework that leverages Vision Language Models (VLMs) for meme detection and understanding in settings where humorous and hateful elements may coexist. The framework first interprets each meme through twelve structured perspectives derived from humor and hate theories. It then applies a role-aware soft-gated attention mechanism to learn how much each perspective should contribute, followed by a prototype-based classifier for the final prediction. Finally, explanations are synthesized using both perspective-specific reasoning and learned attention weights, ensuring transparent and context-grounded justifications. We evaluate MAR-12 on the PrideMM and Memotion datasets, where it achieves up to 80.3% accuracy for humor detection and 75.9% accuracy for hate detection, outperforming state-of-the-art approaches. Furthermore, both human and GPT-4-based evaluations confirm that MAR-12 produces coherent and persuasive explanations, particularly for memes in which humorous and harmful cues co-occur.

13:00 JST研究/論文

ブラックボックスから実行可能なロジックへ: Prolog Expert システムによる説明可能な強化学習

トレーニングされた深層強化学習ポリシーはブラック ボックスであり、その動作を再現し、人間が読み取り、ロジック エンジンを実行し、オプティマイザーが編集できる実行可能なロジック プログラムとして書き換えることによって説明可能にすることができるかどうかを考えます。我々は、凍結された近接ポリシー最適化教師を抽出し、古典的なリレーショナル学習の方法でその決定から順序付けされたルールリストを誘導し、すべての決定が既製の論理エンジンによって実行される Prolog プログラムとして結果を出力する 3 段階のポストホック変換を提示します。後続の拡張ステージではルール ベースが編集され、ポリシー評価で収益の増加が認定された場合にのみ編集が受け入れられます。私たちは 4 つの保証を証明します。リターンロス境界により、抽出されたプログラムは有限マルコフ決定プロセスで機械検査可能な証明書になり、拡張ループは単調に改善して終了します。連続観測設定については、変換がそもそも可能かどうかを答えます。命題閾値インスタンス化は、不一致 O(1/B) と同じレートで閉じるリターン ギャップを伴い、解像度 B が増加するにつれてネットワークを任意の忠実度に変換します。一致する下限は、斜めの決定境界の観測次元でコストが指数関数的であることを示します。経験的に、16,944 の到達可能な状態を持つ 2 つの部屋の鍵とドアのタスクでは、拡張された Prolog プログラムは、すべてのシードで正確に最適なリターンを達成し、予算に制限のある体制では、10 シードのうち 10 の正確なリターンで確率論的教師を超えます。 3 つの連続制御タスクでは、発行されたプログラムがネットワークを置き換え、Acrobot のノイズ内のニューラル教師を 11 の節でマッチングし、CartPole ではリターンの約 97% を回復しますが、より詳細な制御の LunarLander では部分的にのみ回復し、まさに指数関数的下限が予測する天井に達します。

原文 (English)

From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems

A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit. We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in the manner of classical relational learning, and emits the result as a Prolog program whose every decision is executed by an off-the-shelf logic engine; a subsequent expansion stage edits the rule base and accepts an edit only when policy evaluation certifies a return increase. We prove four guarantees. A return-loss bound makes the distilled program a machine-checkable certificate in a finite Markov decision process, and the expansion loop improves monotonically and terminates. For the continuous-observation setting we answer whether the conversion is possible at all: the propositional threshold instantiation converts the network to arbitrary fidelity as the resolution B grows, with disagreement O(1/B) and a return gap that closes at the same rate, and a matching lower bound shows the cost is exponential in the observation dimension for an oblique decision boundary. Empirically, on a two-room key-and-door task with 16,944 reachable states the expanded Prolog program attains exact optimal return in every seed and, in a budget-capped regime, exceeds the stochastic teacher on exact return in ten of ten seeds. On three continuous-control tasks the emitted program substitutes the network, matching the neural teacher within noise on Acrobot with eleven clauses and recovering about 97% of its return on CartPole, while on the finer-control LunarLander it recovers only partially, exactly the ceiling the exponential lower bound predicts.

13:00 JST研究/論文

信頼できる AI ツール、マーク フレームワーク、実装のキャズムの批判的分析

人工知能 (AI) システムが社会に与える影響が増大するにつれ、その導入の倫理的かつ信頼性を確保することが世界的な優先事項になっています。高レベルの倫理ガイドラインが無数に登場している一方で、これらの枠組みは依然として抽象的であり、実装のための具体的なメカニズムが欠けているという批判が根強く残っています。このペーパーでは、OECD の包括的なデータセットを利用して、信頼できる AI (TAI) の運用を目的としたツールとトラスト マーク フレームワークの重要な分析を実施します。実証的なマッピングと記述的な比較分析を通じて、倫理的焦点、ライフサイクル範囲、利害関係者のターゲティング、およびツールの類型における重大な非対称性を特定します。私たちの調査結果は、説明可能性、デジタルセキュリティ、環境の持続可能性には比較的ほとんど注意が払われず、公平性、透明性、堅牢性に重点が置かれていることを示しています。さらに、ほとんどのツールと認定は開発後の段階に重点を置いており、初期の設計段階やデータ収集段階に関するガイダンスは限られています。教育への取り組みや政策への関与は著しく未発達であり、現在の TAI への取り組みが業界内での技術的および手続き的な対策に支配されていることが示唆されています。 AIの原則と実践の間の永続的な溝を埋めるには、倫理目標を拡大し、AIのライフサイクル全体に倫理を埋め込み、より広範なマルチステークホルダーの参加を促進する必要があると私たちは主張します。この調査は、既存の実装ギャップの診断と、より総合的で包括的で強制力のある AI ガバナンスを推進するための実用的な推奨事項の両方を提供します。

原文 (English)

A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms

As artificial intelligence (AI) systems increasingly impact society, ensuring their ethical and trustworthy deployment has become a global priority. While a myriad of high-level ethical guidelines have emerged, criticism persists that these frameworks remain abstract and lack concrete mechanisms for implementation. This paper conducts a critical analysis of tools and trust mark frameworks intended to operationalize trustworthy AI (TAI), drawing on a comprehensive dataset from the OECD. Through empirical mapping and descriptive comparative analysis, we identify significant asymmetries in ethical focus, lifecycle coverage, stakeholder targeting, and tool typology. Our findings show a strong emphasis on fairness, transparency, and robustness, with comparatively little attention paid to explainability, digital security, and environmental sustainability. Moreover, most tools and certifications concentrate on post-development stages, with limited guidance for early design or data collection phases. Educational initiatives and policy engagement are notably underdeveloped, suggesting that current TAI efforts are dominated by technical and procedural measures within industry contexts. We argue that bridging the persistent chasm between AI principles and practice requires expanding ethical objectives, embedding ethics across the AI lifecycle, and fostering broader multi-stakeholder participation. This study provides both a diagnosis of existing implementation gaps and actionable recommendations for advancing more holistic, inclusive, and enforceable AI governance

13:00 JST研究/論文

ロジック、最適化、人工知能

ロジックと最適化を組み合わせることで、ルールベースの AI に貴重な貢献をすることができます。ロジックは、ルール ベースをエンコードし、そこから推論を引き出すための明白な媒体ですが、最適化は推論を計算するための強力なテクノロジを提供します。 AI の透明性に対する懸念が高まる中、両者の組み合わせは新たな関連性を帯びてきました。これは再現性、説明可能性、信頼性、公平性にとって重要です。ルールベースの AI は、透明性に対する自然なソリューションを提供します。これは、今日の高度な最適化手法によりますます実用的になってきています。この記事では、確率論理、ベイズ論理、信念論理とデンプスター・シェーファー理論、非単調 (デフォルト) 論理、多値論理、ブール回帰に基づくノイズを含むデータからの論理式の推論など、論理最適化パートナーシップのいくつかの領域を概説します。デシジョン ダイアグラムとロジック ベースのベンダー分解を使用して、ロジックと最適化の両方の基本的な問題である投影を計算する方法を示します。これは、結論に達する方法を説明するための事後最適化分析の使用と透明性をさらに高めること、および応答セット プログラミング モジュロ理論における最適化の役割について説明します。この論文は、将来の研究の方向性を示唆することで締めくくられています。

原文 (English)

Logic, Optimization, and Artificial Intelligence

Logic and optimization can, in combination, make valuable contributions to rule-based AI. Logic is the obvious medium for encoding a rule base and drawing inferences from it, while optimization provides a powerful technology for computing inferences. Their combination has taken on new relevance amid a growing concern for transparency in AI. which is important for reproducibility, explainability, trustworthiness, and fairness. Rule-based AI provides a natural solution to transparency that is becoming increasingly practical due to today's highly advanced optimization methods. This article surveys several areas of logic-optimization partnership, including probabilistic logic, Bayesian logic, belief logics and Dempster-Shafer theory, nonmonotonic (default) logic, many-valued logics, and inference of logical formulas from noisy data based on Boolean regression. It shows how to compute projections, the fundamental problem of both logic and optimization, using decision diagrams and logic-based Benders decomposition. It describes the use of postoptimality analysis to explain how conclusions are reached, further enhancing transparency, as well as the role of optimization in answer set programming modulo theories. The paper concludes by suggesting possible future research directions.

13:00 JSTエージェント

SeerGuard: ワールド モデル予測によるモバイル GUI エージェントの安全フレームワーク

モバイル グラフィカル ユーザー インターフェイス (GUI) エージェントは、複雑なタスクの自動化において優れた機能を実証していますが、たった 1 つの誤った操作が取り返しのつかない結果を招く可能性があるという重大な安全上のリスクをもたらします。既存の安全メカニズムは主に事後対応型であり、実行前にリスクを評価する機能が欠けています。このホワイトペーパーでは、実行前の指示レベルのスクリーニングとアクションレベルのリスク評価を通じてこれらのリスクを軽減するように設計された、結果を意識した安全フレームワークである SeerGuard を紹介します。具体的には、アクション レベルの評価では、現在の GUI 状態内でエージェントが提案したアクションを分析し、起こりそうな結果を予測して、実行前にリスクを特定します。これらの機能を可能にするために、セマンティックな次状態予測と安全リスク評価を統合した、マルチタスク学習を介して統合安全拡張世界モデル (SAWM) を構築します。広範な実験により、SeerGuard がさまざまなモバイル GUI エージェントにわたって効果的に汎用化されることが実証されました。 Qwen3-VL-8B-Instruct では、安全性ユーティリティ スコアが $\omega=0.8$ で $0.191$ から $0.596$ に増加し、リスク コスト スコアが $\alpha=0.8$ で $0.347$ から $0.130$ に減少します。当社の SAWM のさらなる分析により、行動リスク評価と次の状態の予測の機能とともに、指示レベルのスクリーニングの有効性が検証されます。

原文 (English)

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneous action can lead to irreversible consequences. Existing safety mechanisms are primarily reactive, lacking the ability to assess risks before execution. In this paper, we introduce SeerGuard, a consequence-aware safety framework designed to mitigate these risks through pre-execution instruction-level screening and action-level risk assessment. Specifically, the action-level assessment analyzes agent-proposed actions within current GUI states, anticipating likely outcomes to identify risks before they are executed. To enable these capabilities, we construct a unified safety-augmented world model (SAWM) via multi-task learning, integrating semantic next-state prediction with safety risk assessment. Extensive experiments demonstrate that SeerGuard generalizes effectively across diverse mobile GUI agents. On Qwen3-VL-8B-Instruct, it increases the safety-utility score from $0.191$ to $0.596$ at $\omega=0.8$ and reduces the risk-cost score from $0.347$ to $0.130$ at $\alpha=0.8$. Further analyses on our SAWM validate the effectiveness of the instruction-level screening, alongside the capability of action risk assessment and next-state prediction.

13:00 JSTLLM/生成AI

MGDT: マルチモーダルナレッジグラフを完成させるための関係適応型専門家の混合を備えた MLLM ガイド付き拡散トランスフォーマー

Multimodal Knowledge Graph Completion (MKGC) では、構造的、テキスト的、視覚的な手がかりから欠落しているエンティティを推測する必要があります。既存の拡散ベースの MKGC 手法は、通常、生のマルチモーダル フィーチャに対して直接ノイズを除去します。このような設計では、デノイザーは関係依存のキュー選択、クロスモーダルなセマンティック アラインメント、構造認識エンティティ生成を同時に実行する必要があり、拡散にノイズが多くセマンティックに一貫性のない条件が導入され、その結果、最適ではない補完パフォーマンスが発生します。この制限に対処するために、私たちは MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts (MGDT) を提案します。これは、整列してから拡散するパラダイムに基づいて構築された新しい MKGC フレームワークです。 MGDT はまず、関係適応型セマンティック ルーティング専門家混合 (RASR-MoE) モジュールを採用して、関係に関連したマルチモーダル セマンティック変換パスを選択し、無関係なモダリティ干渉を抑制します。次に、MGDT は、フローズンされたマルチモーダル大規模言語モデル (MLLM) をセマンティック アンカーとして使用して、ルーティングされたマルチモーダル表現を統一された潜在空間に調整し、クロスモーダルのセマンティックな異質性を軽減します。最後に、Knowledge Graph Diffusion Transformer (KGDT) が、位置合わせされた空間でグラフ条件付きノイズ除去生成を実行し、欠落しているエンティティ表現を生成します。 3 つのベンチマーク データセットの実験では、MGDT が一貫して強力なベースラインを上回るパフォーマンスを示しています。

原文 (English)

MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion

Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffusion-based MKGC methods usually denoise directly on raw multimodal features. Such a design forces the denoiser to simultaneously perform relation-dependent cue selection, cross-modal semantic alignment, and structure-aware entity generation, which introduces noisy and semantically inconsistent conditions for diffusion and consequently leads to suboptimal completion performance. To address this limitation, we propose MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts (MGDT), a novel MKGC framework built on an align-then-diffuse paradigm. MGDT first employs a Relation-Adaptive Semantic Routing Mixture-of-Experts (RASR-MoE) module to select relation-relevant multimodal semantic transformation paths and suppress irrelevant modality interference. MGDT then uses a frozen Multimodal Large Language Model (MLLM) as a semantic anchor to align the routed multimodal representations into a unified latent space and reduce cross-modal semantic heterogeneity. Finally, a Knowledge Graph Diffusion Transformer (KGDT) performs graph-conditioned denoising generation in the aligned space to produce the missing entity representation. Experiments on three benchmark datasets show that MGDT consistently outperforms strong baselines.

13:00 JST研究/論文

LEED 準拠のためのニューロシンボリック AI: ドキュメント中心のベンチマーク、決定論的な数値チェック、マルチモーダルが問題となる場合

LEED v4.1 BD+C 認証は依然として文書集約型のプロセスであり、レビュー担当者は数百ページにわたるプロジェクト証拠を読み、クレジット固有のしきい値ロジックを手作業で適用する必要があります。この論文では、ローカルに展開された小規模な言語モデルが LEED ドキュメントの有意義なスクリーニングを実行できるかどうか、および決定論的なシンボリック コンポーネントがその作業をどのように共有すべきかを調査します。プロジェクト PDF を LEED クレジット セクションに調整し、クレジットを認識するキーワード署名を使用して証拠を取得し、ローカルでホストされている 40 億パラメーターの言語モデルへの準拠を検証し、LEED 固有の数値チェッカーを定量的しきい値に適用する、ニューロシンボリック パイプラインが導入されています。 4 つの大学の建物 (PDF 484 件、単位レベルの決定 153 件) での実験では、40 億のパラメーター モデル (gemma3:4b) が最も強力なテキストのみのコア検証ツールであることが示され、このタスクでは 67.3% の精度を達成し、より大きな 80 億のパラメーター モデル (llama3.1:8b) を上回っています。決定論的数値チェッカーは、主要な定量的クレジットの算術エラーを修正し、EA-p2 の精度を 50% から 100% に向上させ、必要な値が確実に抽出された場合には他のいくつかのクレジットを向上させます。同時に、完全な神経記号構成では全体の精度が 61.6% に達し、抽出の失敗と定性的カテゴリでの保守的な動作により、最良のテキストのみのベースラインに劣ります。体系的なアブレーションにより、低解像度の図面画像 (150 ~ 300 dpi) を追加すると一貫して精度が低下すること、およびプロンプトの有効性が建物のグラウンドトゥルース PASS 率に依存することが示されています。ルーブリック プロンプトはドキュメントが豊富なプロジェクトで最もパフォーマンスが高く、思考連鎖プロンプトはドキュメントが少ないプロジェクトで最もパフォーマンスが高くなります。生のプロジェクト文書による LEED v4.1 BD+C コンプライアンス検証の特定の範囲内で、このパイプラインとそのベースラインは、精度と故障モードの両方について再現可能な初期基準点を提供します。

原文 (English)

Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts

LEED v4.1 BD+C certification remains a document-intensive process that requires reviewers to read hundreds of pages of project evidence and apply credit-specific threshold logic by hand. This paper investigates whether small, locally deployed language models can perform meaningful screening of LEED documentation and how deterministic symbolic components should share that work. A neuro-symbolic pipeline is introduced that aligns project PDFs to LEED credit sections, retrieves evidence with credit-aware keyword signatures, verifies compliance with a locally hosted 4-billion-parameter language model, and applies a LEED-specific numeric checker to quantitative thresholds. Experiments on four university buildings (484 PDFs, 153 credit-level decisions) show that a 4-billion-parameter model (gemma3:4b) is the strongest text-only core verifier, achieving 67.3% accuracy and outperforming a larger 8-billion-parameter model (llama3.1:8b) in this task. The deterministic numeric checker corrects arithmetic errors on key quantitative credits, moving EA-p2 from 50% to 100% accuracy and improving several other credits when required values are reliably extracted. At the same time, the full neuro-symbolic configuration achieves 61.6% overall accuracy, trailing the best text-only baseline due to extraction failures and conservative behavior on qualitative categories. Systematic ablations show that adding low-resolution drawing images (150-300 dpi) consistently reduces accuracy, and that prompt effectiveness depends on the building's ground-truth PASS rate: rubric prompts perform best on documentation-rich projects, while chain-of-thought prompts perform best on documentation-lean projects. Within the specific scope of LEED v4.1 BD+C compliance verification over raw project documentation, this pipeline and its baselines provide an initial reproducible reference point for both accuracy and failure modes.

13:00 JSTLLM/生成AIエージェント

ToolVerse: エージェント的強化学習のための大規模な環境と長期的なタスクのロックを解除する

LLM エージェントは、コンパクトで明確に定義されたシナリオでは強力な推論能力を発揮しますが、シームレスなツール統合が必要な大規模で多様で動的な現実世界の環境に直面すると、堅牢性と有効性を維持するのに苦労します。このギャップに対処するために、エージェント RL 環境をスケールアップし、エージェントがツール統合推論 (TIR) タスクで複雑な長期推論を実行できるようにする包括的なフレームワークである ToolVerse を導入します。まず、ToolVerse は、約 4,500 のツールを含む 400 近くの実世界のモデル コンテキスト プロトコル (MCP) から、大規模な実行可能なエージェント トレーニング環境を自動的に構築します。次に、ツール依存関係グラフに基づいたタスク設計戦略を提案します。動的アンロック サンプリング アルゴリズムを利用して長期タスクを生成し、GUST (グラフ アンロック サンプリング タスク) データセットを生成します。第三に、長期エージェント的 RL におけるクレジット割り当ての問題を軽減するために、きめの細かい Turn-Aware Relative Advantage アルゴリズムを提案します。私たちは、ToolVerse を使用して広範なエージェント RL トレーニングを実施し、複数のエージェント ベンチマークでフレームワークを評価します。実験結果は、私たちのフレームワークが長期的なツール使用における LLM の機能を大幅に強化し、顕著なパフォーマンス向上を達成し、動的環境内で堅牢な推論を示すことを示しています。

原文 (English)

ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning

While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and effectiveness when faced with large-scale, diverse, and dynamic real-world environments that demand seamless tool integration. To address this gap, we introduce ToolVerse, a comprehensive framework that scales up agentic RL environments and enables agents to perform complex long-horizon reasoning in Tool-Integrated Reasoning (TIR) tasks. First, ToolVerse automatically builds the massive executable agent training environments from nearly 400 real-world Model Context Protocols (MCPs) that contain about 4500 tools. Second, we propose a task design strategy based on a tool dependency graph, utilizing Dynamic Unlocking Sampling Algorithm to generate long-horizon tasks, and produce GUST (Graph Unlocking Sampling Tasks) dataset. Third, to alleviate the credit assigment problem in long-horizon agentic RL, we propose a fine-grained Turn-Aware Relative Advantage algorithm. We conduct extensive Agentic RL training using ToolVerse and evaluate our framework on serveral agentic benchmarks. Experimental results demonstrate that our framework significantly strengthens LLMs' capabilities in long-horizon tool use, achieving a marked performance boost and showcasing robust reasoning within dynamic environments.

13:00 JST研究/論文GPT / ChatGPTGemini

S1-Omni: 科学的理解、予測、生成のための統合マルチモーダル推論モデル

科学的理解、予測、生成のための統合マルチモーダル推論モデルである S1-Omni を紹介します。 AI for Science (AI4S) は、ドメイン固有のモデル、ツール拡張 LLM、科学言語モデルを通じて大幅に進歩しました。ただし、モデル機能は依然として高度に断片化されており、異種データ、科学法則、専門知識の共同モデリングが制限されています。 S1-Omni は、これらの機能を単一の一貫した科学的推論モデルに統合することで、このギャップに対処します。 S1-Omni のアーキテクチャは、科学データの統一表現、自然界の知識の調整、ドメイン固有のタスクのデコードという 3 つのコア コンポーネントに基づいて構築されています。まず、S1-Omni は、自然言語命令と、CIF、SMILES、タンパク質配列、スペクトル、科学画像などの科学オブジェクトを共有表現空間にマッピングします。第 2 に、科学的法則と専門知識をデータ構築とトレーニングに組み込み、モデルが科学的証拠に基づいて推論できるようにします。 3 番目に、タスク固有のデコードを実行して、特性予測、スペクトルから分子の生成、タンパク質の部位と構造の予測、科学画像の生成と編集などの幅広いアプリケーションをサポートします。 S1-Omni は、200 の科学タスクをカバーし、数百万の推論サンプルを含む S1-Omni-Corpus でトレーニングされており、60 を超える科学ベンチマークで評価されています。ほとんどのベンチマークで GPT-5.5 および Gemini-3.1-Pro を上回り、いくつかのベンチマークではドメイン固有のモデルと同等またはそれを上回ります。全体として、S1-Omni は、統合された科学モデリングへの実用的な道を提供します。

原文 (English)

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model capabilities remain highly fragmented, limiting the joint modeling of heterogeneous data, scientific laws, and expert knowledge. S1-Omni addresses this gap by consolidating these capabilities into a single, coherent scientific reasoning model. The architecture of S1-Omni is built upon three core components: unified representation of scientific data, natural-world knowledge alignment, and decoding for domain-specific tasks. First, S1-Omni maps natural-language instructions and scientific objects, including CIF, SMILES, protein sequences, spectra, and scientific images, into a shared representation space. Second, it incorporates scientific laws and expert knowledge into data construction and training, enabling the model to reason from scientific evidence. Third, it performs task-specific decoding to support a broad range of applications, including property prediction, spectrum-to-molecular generation, protein site and structure prediction, and scientific image generation and editing. S1-Omni is trained on S1-Omni-Corpus, which covers 200 scientific tasks and contains millions of reasoning samples, and is evaluated on over 60 scientific benchmarks. It outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks and matches or surpasses domain-specific models on several benchmarks. Overall, S1-Omni provides a practical path toward unified scientific modeling.

13:00 JSTLLM/生成AIエージェント

情報抽出のためのエージェントモデルの動作制御性: 固定ワークフローから反射型エージェントまで

大規模言語モデル (LLM) エージェントは、複雑な情報抽出タスクにますます使用されていますが、リフレクションやメモリーなどのエージェント コンポーネントが、固定 LLM ワークフローに比べて観察可能かつ制御可能な改善につながるかどうかは依然として不明です。私たちはこの疑問を、会議論文のデータセット抽出を通じて研究します。この抽出では、システムは学術 PDF で言及されているデータセットを識別し、構造化された記録を生成する必要があります。固定ワークフロー ベースラインと反射エージェント バリアントを比較し、より豊富な PDF ツールと動的なツール選択で同じタスクを拡張する最適化されたエージェント条件 (S2) を指定します。私たちの評価では、ツールの実行、再試行、リフレクション、メモリ使用、ランタイム、障害回復などのプロセス レベルの動作に重点を置き、抽出カバレッジとフィールドの完全性を二次的な結果の尺度として扱います。この論文では、エージェント メカニズムがシステムの動作をいつ変更するか、これらの変更によってタスクの完了が向上するかどうか、観察された障害モードが同じ評価ハーネスの下でどのように最適化されたエージェント設計を動機付けるかを特徴付けています。

原文 (English)

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory lead to observable and controllable improvements over fixed LLM workflows. We study this question through conference-paper dataset extraction, where a system must identify datasets mentioned in scholarly PDFs and produce structured records. We compare a fixed workflow baseline with reflective agent variants and specify an optimized agent condition (S2) that extends the same task with richer PDF tools and dynamic tool selection. Our evaluation emphasizes process-level behavior--including tool execution, retries, reflection, memory use, runtime, and failure recovery--while treating extraction coverage and field completeness as secondary outcome measures. The paper characterizes when agentic mechanisms change system behavior, whether these changes improve task completion, and how the observed failure modes motivate an optimized agent design under the same evaluation harness.

13:00 JSTLLM/生成AI

NeurOWL: 不完全なOWLオントロジー推論のためのLLMベースのニューラルシンボリックフレームワーク

OWLオントロジーは、意味論的推論を可能にする形式的な知識表現フレームワークを提供し、ヘルスケアやバイオインフォマティクスなどの分野で広く採用されています。ただし、実際には、現実世界のオントロジーは不完全であることが多く、推論に課題が生じます。この研究では、基本的な包含推論の問題に焦点を当てます。つまり、不完全なオントロジーと候補 (非含意) 包摂が与えられた場合、その包含が意味的に妥当かどうかを判断し、そうであれば、潜在的な欠落している公理を含む論理的に適切な説明を提供します。このタスクは包含検証とオントロジーアブダクションを統合し、欠落している公理の事前定義された候補セットの必要性を取り除くことによって後者を一般化します。この包含推論の問題に対処するために、我々は、正式に定義された意味論と、大規模言語モデルとオントロジー埋め込みを介したテキスト意味論の両方を活用して、検証とアブダクションを共同で実行するエンドツーエンドの神経記号フレームワークである NeurOWL を提案します。私たちは、複数のドメインにわたる現実世界のオントロジーで NeurOWL を評価し、さまざまなドメインにわたって強力で堅牢なパフォーマンスを実証します。

原文 (English)

NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning

OWL ontologies provide a formal knowledge representation framework that enables semantic reasoning, and have been widely adopted across domains such as healthcare and bioinformatics. In practice, however, real-world ontologies are often incomplete, which pose challenges for reasoning. In this work, we focus on a fundamental subsumption reasoning problem: given an incomplete ontology and a candidate (non-entailed) subsumption, determine whether the subsumption is semantically plausible and, if so, providing a logically sound explanation containing potential missing axioms. This task unifies subsumption verification with ontology abduction, and generalizes the latter by removing the need for a predefined candidate set of missing axioms. To address this subsumption reasoning problem, we propose NeurOWL, an end-to-end neuro-symbolic framework that jointly performs verification and abduction, leveraging both formally defined semantics and textual semantics through Large Language Models and ontology embeddings. We evaluate NeurOWL on real-world ontologies across multiple domains, demonstrating strong and robust performance across different domains.

13:00 JSTエージェントビジネス/資金調達

AgentFAIR: 地理空間データセットの公平性評価のためのマルチエージェント協調フレームワーク

地理空間データセットは、都市計画から気候モデリングまでのアプリケーションをサポートしていますが、FAIR 準拠の一貫した評価は困難です。既存の評価者は異なるルーブリックと証拠ソースを使用しており、JavaScript でレンダリングされたページやリポジトリ固有の識別子では失敗する可能性があります。 10 のリポジトリからの 50 のデータセットの場合、利用可能なツール全体の正規化スコアの標準偏差は平均 15.0 パーセント ポイントで、1 つのデータセットでは 30.3 に達します。これらの出力は同等の測定値ではないため、精度の比較ではなく、不一致や故障モードを特徴付けるために使用します。構造化メタデータ抽出と 13 のサブ原則固有の LLM エバリュエーターを組み合わせたマルチエージェント フレームワークである AgentFAIR を紹介します。それぞれが 0 ~ 3 の成熟度スコア、引用された証拠、推奨事項を生成します。批評家は証拠と一貫性をチェックし、的を絞った再評価を要求できます。検索可能性、アクセシビリティ、相互運用性、および再利用性の平均スコアは、79.7%、70.4%、45.3%、および 72.0% です。 4 つのベースライン ツールとのランク相関は 0.31 ~ 0.61 の範囲です。 FAIR-enough 比較は統計的に有意ではありません。 10 個のデータセットを繰り返し実行したサブセットでは、部分原理一致率は平均 89% (標準偏差: 3 パーセント ポイント) でしたが、批判なしでは 71% でした。 15のデータセットの専門家による予備調査では、フライスのカッパが0.71であり、専門家のコンセンサスと82%一致していることがわかりました。 API コストはデータセットあたり約 0.054 米ドルです。これらの結果は、監査可能性と実現可能性を裏付けるものですが、ベンチマークの制限、不完全なアブレーション、および単一モデルファミリーの検証により、精度と一般化に関する主張が制約されます。

原文 (English)

AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets

Geospatial datasets support applications from urban planning to climate modeling, yet consistent assessment of FAIR compliance is difficult. Existing evaluators use different rubrics and evidence sources and may fail on JavaScript-rendered pages or repository-specific identifiers. For 50 datasets from 10 repositories, the standard deviation of normalized scores across available tools averages 15.0 percentage points and reaches 30.3 for one dataset. Because these outputs are not equivalent measurements, we use them to characterize disagreement and failure modes, not comparative accuracy. We present AgentFAIR, a multi-agent framework combining structured metadata extraction with 13 sub-principle-specific LLM evaluators. Each produces a 0-3 maturity score, cited evidence, and recommendations; a critic checks evidence and consistency and can request targeted re-evaluation. Mean Findability, Accessibility, Interoperability, and Reusability scores are 79.7%, 70.4%, 45.3%, and 72.0%. Rank correlations with four baseline tools range from 0.31 to 0.61; the FAIR-enough comparison is not statistically significant. On a 10-dataset repeated-run subset, sub-principle agreement averages 89% (standard deviation: 3 percentage points), versus 71% without the critic. A preliminary 15-dataset expert study yields Fleiss' kappa of 0.71 and 82% alignment with expert consensus. API cost is approximately USD 0.054 per dataset. These results support auditability and feasibility, while the limited benchmark, incomplete ablations, and single-model-family validation constrain claims about accuracy and generalization.

13:00 JSTエージェント

ワークフロー生成のための知識中心のエージェント

ComfyUI などのビジュアル作成システムでのワークフロー生成には、構文の正確さだけでなく、モジュール構成に関する専門家レベルの推論も必要です。既存の大規模言語モデル (LLM) アプローチでは、多くの場合、これをテキストから JSON への直接生成タスクとして扱い、構造的な脆弱性に悩まされ、効果的な設計に必要な経験的知識が不足しています。私たちは、ワークフロー生成を成功させるには、その構造、階層、推論ダイナミクスを含む知識自体をモデル化する必要があると主張します。この目的を達成するために、複数の抽象化レベルにわたって知識を反転、注入、推論することを学習する知識中心のフレームワークを提案します。まず知識の反転を実行して、完全な疑似コードやスケルトンから高レベルの戦略に至るまで、現実世界のワークフローの大規模なコレクションから階層表現を抽出します。次に、教師あり微調整を通じて知識の注入を実行し、タスクの説明から戦略へ、および戦略から実行可能な構造への推論をモデルに教えます。推論中、モデルは可逆推論を実行して実行可能なワークフローを合成し、構造的一貫性のための自己洗練によって強化されます。広範な実験により、私たちの方法が既存のシステムよりも豊富なノード多様性、より一貫性のある構造、およびより高い実行成功率を備えたワークフローを生成し、知識主導型のエージェント型ワークフロー生成の新しい基盤を確立することが実証されました。

原文 (English)

Knowledge-Centric Agents for Workflow Generation

Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large language model (LLM) approaches often treat this as a direct text-to-JSON generation task, struggling with structural brittleness and lacking the experiential knowledge required for effective design. We argue that successful workflow generation requires modeling knowledge itself, including its structure, hierarchy, and reasoning dynamics. To this end, we propose a knowledge-centric framework that learns to invert, inject, and infer with knowledge across multiple abstraction levels. We first perform knowledge inversion to distill hierarchical representations, ranging from full pseudo-codes and skeletons to high-level strategies, from large collections of real-world workflows. We then conduct knowledge injection through supervised fine-tuning, teaching the model to reason from task descriptions to strategies and from strategies to executable structures. During inference, the model performs reversible reasoning to synthesize executable workflows, augmented by self-refinement for structural coherence. Extensive experiments demonstrate that our method produces workflows with richer node diversity, more coherent structures, and higher execution success rates than existing systems, establishing a new foundation for knowledge-driven, agentic workflow generation.

13:00 JSTエージェント

DSWorld: 効率的な自律エージェントのためのデータ サイエンス ワールド モデル

データの理解と意思決定における強力な機能にもかかわらず、自律型データ サイエンス エージェントは依然として、高価な計算を伴う試行錯誤のワークフローに大きく依存しています。このボトルネックにより、実際の実行前にデータ サイエンス操作の効果を予測できるモデルが動機付けられます。この論文では、現在のワークフロー状態と候補操作を条件とした環境状態遷移を予測することでデータ サイエンス実行環境をモデル化するデータ サイエンス ワールド モデルの概念を紹介します。さらに、構造化された状態の構築、コストを意識したルーティング、軽量の実際の実行、および高価な操作のための LLM ベースのシミュレーターを組み合わせた実用的なフレームワークである DSWorld を提案します。トレーニングをサポートするために、8K スケールの遷移軌跡データセットを構築し、遷移予測を改善するためのエラー認識強化学習戦略である Reflective World Model Optimization を導入します。実験の結果、DSWorld は、競争力のあるパフォーマンスを維持しながら、RL ベースのエージェント トレーニングを約 $14\times$、検索ベースの推論を約 $3$-$6\times$ 加速し、遷移予測タスクで最も強力な LLM ベースラインを 35.6% 上回るパフォーマンスを示しています。コードは https://anonymous.4open.science/r/DSWorld で入手できます。

原文 (English)

DSWorld: A Data Science World Model for Efficient Autonomous Agents

Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate the effects of data science operations before real execution. In this paper, we introduce the concept of Data Science World Model, which model the data science execution environment by predicting environment state transitions conditioned on current workflow states and candidate operations. We further propose DSWorld, a practical framework that combines structured state construction, cost-aware routing, lightweight real execution, and an LLM-based simulator for expensive operations. To support training, we construct an 8K-scale transition trajectory dataset and introduce Reflective World Model Optimization, an error-aware reinforcement learning strategy for improving transition prediction. Experiments show that DSWorld accelerates RL-based agent training by approximately $14\times$ and search-based inference by approximately $3$-$6\times$ while maintaining competitive performance, and outperforms the strongest LLM baseline by 35.6% on transition prediction tasks. The code is available at https://anonymous.4open.science/r/DSWorld.

13:00 JST研究/論文

形式的に接地された ODRL 評価器: 実装と比較

ODRL ポリシー言語は、ヨーロッパのデータスペースにおけるデータ アクセスと使用設定、AI ガバナンス ポリシー、データ ワークフローをモデリングするポリシーの事実上の標準として浮上しています。現在の標準には、システムがポリシー評価をどのように実装するかを記述するための数学的形式的意味論がありません。その結果、言語の独自の解釈を実装するさまざまなシステムやツールが作成され、相互運用性が制限され、一貫した結果が保証されなくなりました。 ODRL の既存のセマンティック モデルに基づいて、静的設定とストリーミング設定の両方でアクセス制御および監視シナリオの ODRL 評価の問題を形式化し、斬新で効率的なアルゴリズムと実装を提供します。透過的な形式セマンティクスを備え、すべてのルール タイプをサポートする最初の ODRL エバリュエーターを紹介します。私たちはそのパフォーマンスを実験的に測定し、ポリシーの複雑さとポリシーが評価されるデータのサイズに関連するさまざまなスケーラビリティの側面を分析します。既存の ODRL 評価器の比較レビューを提供することで、当社のシステムと最先端のシステムを比較します。これにより、サポートされている ODRL 機能と評価モードの違いが強調されます。

原文 (English)

A Formally Grounded ODRL Evaluator: Implementation and Comparison

The ODRL policy language is emerging as the de-facto standard for policy modelling data access and usage preferences, AI governance policies and data workflows in European dataspaces. The current standard has no mathematical formal semantics to describe how a system should implement policy evaluation. This has resulted in a variety of systems and tools that implement their own interpretation of the language, which limits interoperability and cannot guarantee consistent results. Based on an existing semantic model of ODRL, we formalise the problems of ODRL evaluation for the access control and monitoring scenarios, in both static and streaming settings, and we provide a novel, efficient algorithm and implementation. We present the first ODRL Evaluator with transparent formal semantics and supporting all rule types. We experimentally measure its performance, analysing different scalability dimensions related to policy complexity and size of the data on which a policy is evaluated. We compare our system with the state-of-the-art by providing a comparative review of existing ODRL evaluators, which highlights the differences in supported ODRL features and evaluation modes.

13:00 JST研究/論文

AI の信頼ギャップを埋める: 信頼できる AI に対する独立した認定の事例

過去 10 年間にわたり、責任ある AI (RAI) は、一か八かの状況で AI が引き起こすリスクを特定し、軽減するための実質的な実践を生み出してきました。しかし、この取り組みは信頼性に報いる市場を生み出していません。安全性、公平性、監視に真剣に投資している企業は、自社のシステムが最低限のコンプライアンスを超えていることを消費者、規制当局、株主に対して一貫して証明することはできません。欠けているのは、社会が違いを認識したり比較したりする方法です。その結果、信頼のギャップが生じます。これは、責任ある開発の取り組みが組織内で行われているにもかかわらず、外部に独立して認識され、検証可能な信頼できる結果のシグナルが生成されないという構造的な状態です。私たちは、このギャップが維持されているのは、信頼できる AI (独立して検証可能な現実世界の結果の問題) ではなく、責任ある AI (内部プロセスの問題) に重点が置かれていることが部分的に原因であり、次の 3 つの複合的な失敗によってこのギャップが続いていると主張します。(1) 市場は信頼できるシステムとその模倣品を区別できない。 (2) 評価の対象は、導入された社会技術システムとその成果ではなく、モデルと成果です。 (3) 測定エコシステムは、利益を実証することよりもむしろ害悪を回避することを指向しています。既存の AI ガバナンス手段をレビューし、ヘルスケア、持続可能性、セキュリティの認証制度と比較したところ、ガバナンスのベースライン、独立して検証されたポジティブな結果の証拠、市場シグナルを単一のフレームワークに統合しているものはないことがわかりました。私たちは、信頼性のギャップを埋めることができる接続層として、独立した結果指向の認証を提案し、信頼性を測定可能、比較可能にし、商業的な報酬を与えることで規制と内部ガバナンスを補完します。

原文 (English)

Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI

Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings. Yet this work has not produced a market that rewards trustworthiness. Firms that invest seriously in safety, fairness, and oversight cannot consistently prove to consumers, regulators, and shareholders that their systems go beyond the bare minimum of compliance. What is missing is a way for society to recognize or compare the difference. The result is a trust gap: a structural condition in which responsible development efforts happen inside organizations but produce no external, independently recognized and verifiable signal of trustworthy outcomes. We argue this gap is sustained in part because of a focus on responsible AI (a matter of internal process) as opposed to trustworthy AI (a matter of independently verifiable real-world outcomes), and that it persists because of three compounding failures: (1) the market cannot distinguish trustworthy systems from their imitations; (2) evaluation targets models and outputs rather than deployed sociotechnical systems and their outcomes; (3) the measurement ecosystem is oriented toward avoiding harm rather than demonstrating benefit. Reviewing existing AI governance instruments and comparing them to certification regimes in healthcare, sustainability, and security, we show that none integrate a governance baseline, independently verified positive-outcome evidence, and market signaling in a single framework. We propose independent, outcome-oriented certification as the connective layer that can close the trust gap, complementing regulation and internal governance by making trustworthiness measurable, comparable, and commercially rewarded.

13:00 JSTハードウェア/半導体研究/論文

SciForge: 科学的発見のための AI ネイティブのマルチモーダル ワークベンチ

科学的研究は、論文、コード、データセット、科学ファイル形式、モデル出力、図、原稿、チームの意思決定など、異種混合の成果物にますます広がっていますが、汎用 AI アシスタントがこれらのオブジェクトを一貫した監査可能な研究状態として保存することはほとんどありません。我々は、マルチモーダルな研究ネイティブ AI ワークベンチである SciForge を紹介します。このワークベンチは、検索、解析、モデル ルーティング、ワークフロー実行、プロット、書き込み、プレゼンテーション生成を、エージェントがアクセス可能なモジュール式サービスとして実行しながら、人間の判断のためにグラフィカル インターフェイスを確保します。 SciForge は 5 つの柱を中心に構築されています。(i) \textbf{目標指向} 研究のための \emph{目標を見据えた科学的意思決定ガバナンス}。レビュー ゲートと共有レビュー サーフェスを備えています。 (ii) \textbf{multimodal} 入力の \emph{translate-then-reason}。エージェントが理由を判断する前に科学オブジェクトをドメイン トランスレータ経由でルーティングします。 (iii) \textbf{auditable} トレーサビリティのための \emph{証拠ガバナンス}、主張を出所連鎖および監査結果に結び付ける。 (iv) \textbf{共同} 研究のための \emph{共同チーム サイエンス}。これにより、複数の役割による意思決定ガバナンスが可能になり、将来のリリースで共有チーム ワークスペースが予定されています。 (v) \textbf{実用的}な効果のための \emph{現実世界のアプリケーション シナリオ}。遺伝子発見、AI 誘導による新規タンパク質設計、分子最適化、ゲノムから BGC の発見のための数日間にわたるエージェントリサーチ スプリントを含む主力デモンストレーションを伴う、8 つのエンドツーエンド ユーザー ケースを通じて実証されます。このシステムは、薄いインタラクション レイヤー、コンテキスト リサーチ機能パターン、エージェント ランタイムとワークフロー エンジン、Evidence-DAG 監査サイドカー、および科学モデル ルーターを組み合わせています。 SciForge は現在、モバイル監視をサポートするデスクトップ アプリケーションとして実行されます。将来のリリースでは、チームのコラボレーションがさらに深まります。このシステムはオープンソースであり、https://github.com/AGI4Sci/SciForge から入手できます。

原文 (English)

SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery

Scientific work increasingly spans heterogeneous artifacts -- papers, code, datasets, scientific file formats, model outputs, figures, manuscripts, and team decisions -- yet general-purpose AI assistants rarely preserve these objects as a coherent, auditable research state. We present SciForge, a multimodal research-native AI workbench that reserves the graphical interface for human judgment while search, parsing, model routing, workflow execution, plotting, writing, and presentation generation run as modular agent-accessible services. SciForge is built around five pillars: (i) \emph{goal-scoped scientific decision governance} for \textbf{goal-oriented} research, with review gates and shared review surfaces; (ii) \emph{translate-then-reason} for \textbf{multimodal} input, routing scientific objects through domain translators before the agent reasons; (iii) \emph{evidence governance} for \textbf{auditable} traceability, linking claims to provenance chains and audit findings; (iv) \emph{collaborative team science} for \textbf{collaborative} research, enabling multi-role decision governance, with shared team workspaces planned for future releases; and (v) \emph{real-world application scenarios} for \textbf{practical} impact, demonstrated through eight end-to-end user cases, with flagship demonstrations including multi-day agentic research sprints for gene discovery, AI-guided de novo protein design, molecular optimization, and genome-to-BGC discovery. The system combines a thin interaction layer, contextual research capability patterns, an Agent Runtime and Workflow Engine, an Evidence-DAG audit sidecar and a Scientific Model Router. SciForge currently runs as a desktop application, with mobile supervision support; future releases will deepen team collaboration. The system is open-source and available at https://github.com/AGI4Sci/SciForge

13:00 JST研究/論文

AI の安全性閾値の調和

フロンティア AI 企業は、大幅に異なる機能のしきい値を公開しているため、サードパーティがしきい値を超えているかどうかを確認したり、企業間の要件を比較したりすることが困難になっています。さらに、共通の最低閾値がなければ、リスク軽減に一貫性がなくなり、安全基準の下位への競争が生じる可能性があります。私たちは、3 つのリスク領域にわたって調和されたしきい値を導き出すための方法論を開発します。悪用リスク (サイバーおよび生物学的) については、予想される危害を主要なプリミティブとして捉え、リスク チャネルとモデルのリリース条件を考慮した明示的なリスク モデリング アプローチを使用します。自動化された AI 研究開発については、予想される害ではなく、観察された AI の進歩速度に基づいて、提案されたしきい値を設定します。私たちの分析は以前の研究を拡張し、既存の経験的なギャップと限界を浮き彫りにします。

原文 (English)

Harmonizing AI Safety Thresholds

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.

13:00 JSTLLM/生成AIビジネス/資金調達

CRAFT: ルーブリックをクラスタリングして弱い LLM 機能を診断し、対象を絞った微調整データを生成する

評価では、モデルの現在のパフォーマンスを測定するだけでは不十分です。彼らは、次のモデルの反復で何を修正すべきかを私たちに伝え、ターゲットを絞ったトレーニング後のデータを生成する方法を提供する必要があります。ほとんどの評価パイプラインは、弱い例、トピック、またはカテゴリを特定しますが、潜在的な機能の失敗を暗黙的に残します。モデルが失敗する理由ではなく、どこで失敗するかを示します。ルーブリックベースの評価データセットをモデル固有の弱い機能の診断に変換する手法である CRAFT を紹介します。 CRAFT は、各評価基準を能力プローブとして扱います。つまり、すべてのプロンプト ルーブリック ペアから能力の説明を抽出し、これらの説明を階層的な能力ツリーにクラスタリングし、すべてのノードでターゲット モデルにスコアを付け、各障害が最も明確な粒度でツリー レベル全体にわたってパフォーマンスの低いノードを動的に選択します。選択された弱い機能は、ターゲットを絞った教師付き微調整データの生成を指示します。データ生成、微調整、および評価のセットアップを固定したまま、4 つのオープンソース モデル、2 つの専門分野 (財務および法務)、および診断データから切り離された 13 のベンチマークで、プロンプト レベルの EvalTree クラスタリングおよびターゲットを絞らないランダム生成と CRAFT を比較します。 CRAFT は、温度デコードを繰り返すと、4 つのモデルすべてで最も強力な金融ドメイン平均を達成します。法的領域では、4 つのモデルのうち 3 つで最も強く、4 つ目のモデルでは最良のベースラインのデコード分散帯域内に留まります。したがって、プロンプトやカテゴリではなく、ルーブリック基準のレベルで弱点を診断すると、モデルで何ができないのかをより明確に把握できるようになり、その診断に基づいて微調整した後は測定可能なほど優れたモデルが得られます。

原文 (English)

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or categories, but they leave the underlying capability failure implicit: they say where a model fails, not why. We introduce CRAFT, a method that converts any rubric based evaluation dataset into a model specific diagnosis of weak capabilities. CRAFT treats each grading criterion as a capability probe: it extracts a capability description from every prompt rubric pair, clusters these descriptions into a hierarchical capability tree, scores the target model at every node, and selects low performing nodes dynamically across tree levels, at the granularity where each failure is clearest. The selected weak capabilities then direct the generation of targeted supervised finetuning data. Holding the data generation, finetuning, and evaluation setup fixed, we compare CRAFT against prompt level EvalTree clustering and untargeted random generation on four open source models, two professional domains (finance and legal), and 13 held out benchmarks disjoint from the diagnostic data. CRAFT achieves the strongest finance domain average for all four models under repeated temperature decoding; on legal domain, it is strongest for three of four models and remains within the decoding variance bands of the best baseline on the fourth. Diagnosing weaknesses at the level of rubric criteria, rather than prompts or categories, thus yields both a sharper picture of what a model cannot do and measurably better models after finetuning on that diagnosis.

13:00 JSTLLM/生成AI

予測的ずれの許容度としての共感: 共同規制の枠組みと対話修復の体制構造

共感は、共鳴、つまり他人の現在の感情的または認知的状態を反映するものとして理論化されることがほとんどです。この共時的なフレーミングは人工的なシステムを形成しており、そこでは共感的な行動が認識と反応の調整に影響を与えるものとして定義されています。私たちは、これは、予測、分岐、修正を通じて時間の経過とともに理解が展開される拡張対話の間違ったターゲットであると主張します。私たちは、共感を、予測的不整合許容度、つまり、時間の経過とともに乖離を崩壊させるのではなく、予測して調整する能力として再構成します。私たちはこれを、エージェント間の乖離の実行可能な範囲を維持するものとして共感をモデル化する動的閾値ヒューリスティックである解釈エラー許容度 (IET) として形式化します。このフレームワークを、制御されたノイズの下で 2 つの計算プローブを使用して評価します。 IET 更新ルールは、固定ベースラインを上回るパフォーマンスを示しません。その代わりに、私たちは堅牢な体制依存構造を発見しました。修復は、要点の保存のために識別的な忠実性をトレードオフにします。ノイズが低い場合、修復により検索精度が低下します。ノイズが高い場合でも、要点の意味が保持され、ノイズ レベル、修復、および評価基準の間の相互作用が明らかになります。我々はIETを通じてこの構造を解釈し、拡張された相互作用における共感は発散を排除するのではなく、そのダイナミクスを調節していることを示唆しています。これにより、共感型 AI の設計が収束から解釈上の距離の管理へと移行するきっかけになります。

原文 (English)

Empathy as Predictive Misalignment Tolerance: A Co-Regulation Framework and the Regime Structure of Dialogue Repair

Empathy is most often theorized as resonance: a mirroring of another's present emotional or cognitive state. This synchronic framing has shaped artificial systems, where empathic behavior is defined as affect recognition and response alignment. We argue this is the wrong target for extended dialogue, where understanding unfolds over time through prediction, divergence, and repair. We reframe empathy as predictive misalignment tolerance: the capacity to anticipate and regulate divergence across time rather than collapse it. We formalize this as Interpretive Error Tolerance (IET), a dynamic-threshold heuristic that models empathy as maintaining a viable band of divergence between agents. We evaluate this framework with two computational probes under controlled noise. The IET update rule does not outperform fixed baselines. Instead, we find a robust regime-dependent structure: repair trades discriminative fidelity for gist preservation. At low noise, repair degrades retrieval accuracy; at high noise, it preserves gist meaning, revealing an interaction between noise level, repair, and evaluation metric. We interpret this structure through IET, suggesting that empathy in extended interaction is not eliminating divergence but regulating its dynamics. This motivates a shift in empathic AI design from convergence toward managing interpretive distance.

13:00 JST研究/論文

より優れたシステム制御でユーザーを強化すると、ニュース フィルター バブルにどのような影響がありますか?

レコメンデーション システムを使用すると、ユーザーは興味のある記事を見つけることができますが、ユーザーの既存の信念を強化するコンテンツを提示することで「フィルター バブル」を作成することもできます。ユーザーは、システムが自分をフィルターバブルの中に入れたことに気づいていないことが多く、たとえ気づいていても、それを直接制御できないことがよくあります。これらの問題に対処するために、私たちはまず、システムがユーザーの行動から推測した政治的および時事的な関心を明らかにする、強化されたインターフェイスで強化された政治ニュース推奨システムを設計します。これにより、ユーザーは推奨システムを調整して、特定のトピックに関する記事や特定の政治的立場を示す記事をさらに受け取ることができます。次に、ユーザー調査を実施してシステムを従来のインターフェイスと比較したところ、透過的なアプローチにより、ユーザーが自分がフィルターバブルの中にいることを認識できることがわかりました。さらに、強化されたシステムにより、ほとんどのユーザーにとって極端なニュースは少なくなりましたが、他のユーザーがシステムをより極端なものにすることも可能になりました。同様に、多くのユーザーがシステムを極端なリベラル/保守的なものから中心的なものに移行しましたが、これは表示される記事の政治的多様性を減少させるという犠牲を伴いました。これらの発見は、提案されたシステムがフィルターバブルの認知度を高めた一方で、ユーザーの好みに応じてニュース消費に不均一な影響を及ぼしたことを示唆しています。

原文 (English)

How Does Empowering Users with Greater System Control Affect News Filter Bubbles?

While recommendation systems enable users to find articles of interest, they can also create ``filter bubbles'' by presenting content that reinforces users' pre-existing beliefs. Users are often unaware that the system placed them in a filter bubble and, even when aware, they often lack direct control over it. To address these issues, we first design a political news recommendation system augmented with an enhanced interface that exposes the political and topical interests the system inferred from user behavior. This allows the user to adjust the recommendation system to receive more articles on a particular topic or presenting a particular political stance. We then conduct a user study to compare our system to a traditional interface and found that the transparent approach helped users realize that they were in a filter bubble. Additionally, the enhanced system led to less extreme news for most users but also allowed others to move the system to more extremes. Similarly, while many users moved the system from extreme liberal/conservative to the center, this came at the expense of reducing political diversity of the articles shown. These findings suggest that, while the proposed system increased awareness of the filter bubbles, it had heterogeneous effects on news consumption depending on user preferences.

13:00 JST研究/論文

循環二項畳み込み誤差の構造

二項畳み込みと循環畳み込みは、それぞれアダマール変換と FFT で計算された離散フーリエ変換 (DFT) を使用して $O(N\log N)$ 時間で計算できます。アダマール変換は、実数値の符号反転の点で好ましいですが、DFT を代入すると代数誤差が生じます。このエラーを特徴付ける 3 つの補足的な結果を示します。まず、正確なエラーキャンセルを特定します。2 つの入力位置と 2 つの出力位置には一般的にエラーがなく、出力を並べ替えてもこのエラーを除去できません。第 2 に、誤差演算子はほぼフルランクですが、そのヌル空間は対数次元しか持ちません。 3 番目に、予期される誤差は、ランダム フィルターの平均によって得られる閉形式の単一のアライメント スカラーによって制御されます。一般に、置換誤差は出力エネルギーを漸近的に 2 倍にします。ただし、誤差が発生しないユニバーサル ゼロ誤差部分空間のフィルタは除きます。まとめると、これらの結果は、置換エラーが構造化され、予測可能であり、アライメントによって支配されることを示しています。

原文 (English)

Structure of the Circular-Dyadic Convolution Error

Dyadic and circular convolution can both be computed in $O(N\log N)$ time using the Hadamard transform and the FFT-computed discrete Fourier transform (DFT), respectively. The Hadamard transform is preferable for its real-valued sign flips, yet its substitution for the DFT introduces algebraic error. We present three complementary results that characterize this error. First, we identify exact error cancellation: two input and two output positions are universally error-free, and no reordering of the output can eliminate this error. Second, the error operator is nearly full rank, while its null space has only logarithmic dimension. Third, the expected error is governed by a single alignment scalar, with a closed-form expression obtained by averaging over random filters. In general, the substitution error asymptotically doubles the output energy, except for filters in the universal zero-error subspace, which incur no error. Collectively, these results show that the substitution error is structured, predictable, and governed by alignment.

13:00 JST研究/論文

AV-JEPA: LeJEPA を視聴覚自己教師あり学習に拡張

私たちは、視聴覚自己教師あり学習への LeJEPA のエレガントなマルチモーダル拡張である AV-JEPA を紹介します。初期融合ビジョン トランスフォーマーとモダリティ ドロップアウトをマスキングとして使用し、モデルはグローバル ビューとモダリティごとのローカル ビューの埋め込みを調整するようにトレーニングされ、SIGReg の目的は理論的に最適な分布を促進します。これにより、潜在空間でのクロスモーダル アラインメントが実現され、デコーダ、EMA 教師、複雑な複数項の損失、またはコントラストのネガがない、非常にクリーンなアーキテクチャが実現します。提案された AV-JEPA バックボーンは、VGGSound (57.1% トップ-1) および AudioSet (32.7 mAP) で競争力のある分類パフォーマンスを提供し、すぐに使えるゼロショット オーディオビデオ検索をサポートします。

原文 (English)

AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality local views, while the SIGReg objective encourages a theoretically optimal distribution. This achieves cross-modal alignment in the latent space, resulting in a remarkably clean architecture with no decoder, EMA teacher, complex multi-term losses, or contrastive negatives. The proposed AV-JEPA backbone delivers competitive classification performance on VGGSound (57.1% top-1) and AudioSet (32.7 mAP) and supports zero-shot audio-video retrieval out of the box.

13:00 JST研究/論文

暗黙的なニューラル表現を使用したデータ駆動型ビデオ コーデック

従来のコーデックは、ビデオを圧縮されたピクセル データとして保存します。代わりに、ビデオとそのオーディオ トラックを、時空座標を RGB 値とオーディオ振幅にマッピングする単一の正弦波表現ネットワーク (SIREN) の重みとして保存します。このネットワークは、個別のオーディオとビデオの初期化層、完全に接続された共有隠れ層のスタック、および 3 つの出力ブランチ (ビデオ用に 1 つと残留ノイズの推定と減算に不一致が使用される 2 つのシャム オーディオ ブランチ) を使用します。次に、過剰適合した教師ネットワークは、応答ベースの知識蒸留によって小さな生徒に圧縮され、その後、16 ビットの対称重み量子化とロスレス LZMA2 (xz) エンコーディングが行われます。 6.08 MiB のテスト ビデオでは、量子化された学生は、SSIM 0.75 でビデオ PSNR 28.72 dB、対数スペクトル距離 10.69 dB でオーディオ PSNR 24.18 dB に達しますが、パイプラインによって表現が 9.05 MiB から 2.33 MiB に縮小され、全体の圧縮率は 2.61 になります。 1 ビットから 32 ビットの量子化までビット幅をスイープすると、再構成の品質が 16 ビットで飽和することがわかります。 H.264、HEVC、および MP3 と比較し、このアプローチがどこで不足しているかを報告し、WebRTC 経由でこれらのモデルをトレーニング、転送、デコードするブラウザベースのプロトタイプについて説明します。

原文 (English)

Data-driven Video Codec with Implicit Neural Representations

A conventional codec stores a video as compressed pixel data. We instead store the video, together with its audio track, as the weights of a single sinusoidal representation network (SIREN) that maps space-time coordinates to RGB values and audio amplitudes. The network uses separate audio and video initialization layers, a stack of shared fully connected hidden layers, and three output branches: one for video and two Siamese audio branches whose disagreement is used to estimate and subtract residual noise. The overfitted teacher network is then compressed by response-based knowledge distillation into a smaller student, followed by 16-bit symmetric weight quantization and lossless LZMA2 (xz) encoding. On a 6.08 MiB test video, the quantized student reaches a video PSNR of 28.72 dB with SSIM of 0.75, and an audio PSNR of 24.18 dB with a log spectral distance of 10.69 dB, while the pipeline shrinks the representation from 9.05 MiB to 2.33 MiB, an overall compression ratio of 2.61. A bit-width sweep from 1-bit to 32-bit quantization shows that reconstruction quality saturates at 16 bits. We compare against H.264, HEVC, and MP3, report where the approach falls short of them, and describe a browser-based prototype that trains, transfers, and decodes these models over WebRTC.

13:00 JST研究/論文

組み込みシステムの検証ギャップを埋めるためのシストリック配列を使用した遅延演算

ディープ ニューラル ネットワークなどの複雑なアルゴリズムは、リソースに制約のある組み込みプラットフォームに導入されることが増えています。しかし、これらのモデルをエッジに実装するための既存のハードウェアおよびソフトウェアのスキームは、特に医療機器などの安全性が重要なアプリケーションでは不十分です。まず、GPU、NPU、TPU などのハードウェアは、セキュリティ計算の正確さよりもスループットを重視して設計されているため、フォールト インジェクション攻撃の影響を受けやすくなっています。第 2 に、アルゴリズムをエッジ デバイスに移植するために設計されたソフトウェア スキーム (量子化スキームなど) は、静的で健全 (消費電力が最適化されていない) か、動的ではあるが健全ではない (セーフティ クリティカルなアプリケーションには最適ではない) かのいずれかです。これら両方のニーズに対処するために、私たちは、リアルタイム、ダイナミック、サウンド量子化に対する全く新しいアプローチと、それをサポートするハードウェアの両方を提案します。まず、左から右の算術を利用して最上位ビット (MSB) を最初に渡し、感度分析を実行しながら精度をオンラインで動的に調整して、決定境界を越えるリスクを定量化して管理する、健全なリアルタイム適応精度量子化アプローチを開発しました。次に、シストリック アレイを利用して左から右への演算を実行し、MSB ファーストを生成する新しいハードウェア アプローチを提案します。これにより、リソース効率の高いニューラル ネットワークとエッジの人工知能だけでなく、最も重要なビットに対するビット フリップ攻撃に対する回復力を確保する、ハードウェア上で広範に健全でリソース効率の高い高精度数学を実現するためのまったく新しいスキームが提供されます。これは、ここでは進行中の作業として示されており、ソフトウェアの実装は完了し、ハードウェアは進行中です。

原文 (English)

Lazy Arithmetic using Systolic Arrays for Closing the Verification Gap on Embedded Systems

Complex algorithms such as deep neural networks are increasingly being deployed on embedded, resource constrained platforms. However, existing hardware and software schemes for implementing these models on the edge fall short, particularly for safety-critical applications such as medical devices. First, hardware such as GPUs, NPUs and TPUs are designed for throughput rather than correctness of computation of security, and are as such susceptible to fault injection attacks. Second, software schemes designed for porting algorithms onto edge devices -- such as quantization schemes -- are either static and sound (non-optimal power consumption), or dynamic yet unsound (non-optimal for safety-critical applications). To address both these needs we propose a both wholly new approach to real-time, dynamic and sound quantization, as well as the hardware to support it. First we developed a sound, real-time adaptive-precision quantization approach utilizing left-to-right arithmetic to pass the most significant bits (MSB) first, and dynamically adjust precision online while performing sensitivity analysis to quantify and manage the risk of decision-boundary crossings. Next, we propose a novel hardware approach utilizing systolic arrays to perform left-to-right arithmetic to generate the MSB first. Together this provides a wholly novel scheme for enabling not only resource-efficient neural networks and artificial intelligence at the edge, but broadly sound and resource-efficient high-precision mathematics on hardware that ensures resilience to bit flip attacks on the most critical bits. This is presented herein as work-in-progress, with software implementations completed and hardware in-progress.

13:00 JSTLLM/生成AIGemmaLlamaQwenDeepSeek

臨床予測のための統合マルチモーダル学習器としての大規模言語モデル

電子医療記録は、フリーテキストの臨床ナラティブと、バイタルサイン、検査値、併存疾患などの構造化された測定値を組み合わせたものです。しかし、ほとんどの臨床予測システムは依然としてタスク固有の融合アーキテクチャに依存しており、各モダリティの専用エンコーダと、新しいタスクや臨床設定ごとに再設計する必要がある学習済みの組み合わせメカニズムを組み合わせています。私たちは、よりシンプルな代替案を提案します。モダリティに関係なく、すべての患者データを単一の自然言語シーケンスに変換し、融合のためのアーキテクチャを変更せずに、事前トレーニングされた言語モデルをエンドツーエンドで微調整します。我々は、臨床的に異なる 3 つの予測タスクにわたってこのアプローチを評価します。MIMIC-III での院内死亡率、ドイツの移植センターからの縦断データを使用した移植片不全予測、救急車の記録からの緊急トリアージ分類 - エンコーダー ベース (ModernBERT) とデコーダー ベース (Llama 3.1、Gemma、DeepSeek-R1-Qwen、Qwen3) の微調整を確立されたマルチモーダル ベースラインに対して比較し、移植失敗、移植後の患者管理のために臨床現場で現在使用されている勾配ブースティング モデル。 3 つのタスクすべてにわたって、統一されたテキストのシリアル化は、タスク固有のマルチモーダル ベースラインと一致またはそれを上回り、移植片障害の予測において臨床的に導入されているグラジエント ブースティング システムを上回っています。これらの結果は、オーダーメイドの融合アーキテクチャを必要とせず、単一のシリアル化ベースのパラダイムでマルチモーダルな臨床予測に十分であり、特殊な設計と同等またはそれを超えながら、システムの複雑さを大幅に軽減できることを示しています。

原文 (English)

Large Language Models as Unified Multimodal Learners for Clinical Prediction

Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities. Yet most clinical prediction systems still rely on task-specific fusion architectures, pairing dedicated encoders for each modality with learned combination mechanisms that must be re-engineered for every new task and clinical setting. We propose a simpler alternative: convert all patient data, regardless of modality, into a single natural language sequence and fine-tune a pretrained language model end-to-end, with no architectural modification for fusion. We evaluate this approach across three clinically distinct prediction tasks: in-hospital mortality on MIMIC-III, graft failure prediction using longitudinal data from a German transplant center, and emergency triage classification from ambulance records - comparing encoder-based (ModernBERT) and decoder-based (Llama 3.1, Gemma, DeepSeek-R1-Qwen, Qwen3) fine-tuning against established multimodal baselines and, for graft failure, a gradient boosting model currently used in clinical practice for post-transplant patient management. Across all three tasks, unified textual serialization matches or exceeds task-specific multimodal baselines, and outperforms the clinically deployed gradient boosting system on graft failure prediction. These results indicate that a single serialization-based paradigm, without bespoke fusion architectures, is sufficient for multimodal clinical prediction - substantially reducing system complexity while matching or exceeding specialized designs.

13:00 JST画像/動画生成

脳腫瘍セグメンテーションにおけるリソース制約のある深部ニューラルネットワークトレーニングのためのマルチコントラスト 3D MRI 選択戦略としての部分情報分解

マルチコントラスト 3D MRI セグメンテーションは、利用可能なシーケンスをすべて使用すると、計算負荷が高くなる可能性があります。局所腫瘍量に関する冗長、固有、相乗情報に従って入力ペアをランク付けし、下流トレーニング用に最高ランクのペアを選択するトレーニング前の部分情報分解フレームワークを評価します。 T1n、T1c、T2w、および T2-FLAIR MRI に適用されるフレームワークでは、T1c+T2-FLAIR が選択されました。次に、異なる入力構成を使用して、アーキテクチャ的に同一の 11 個の軽量 3D U-Net をトレーニングしました。独立したテスト コホートでは、T1c+T2-FLAIR が最も強力な 2 入力構成であり、平均ダイスで全体で 2 位にランクされました (4 つの入力すべてで 0.676 対 0.687)。全入力モデルに対する独立した Shapley 分析でも、T2-FLAIR と T1c が最も影響力のある入力であり、それらのペアの相互作用が最も強いと特定されました。これらの発見は、コストのかかる 3D モデル開発の前に、コンパクトで有益な MRI 入力セットを特定するための PID ベースの事前トレーニング選択の実際的な価値を示しています。

原文 (English)

Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resource-Constrained Deep Neural Network Training in Brain Tumor Segmentation

Multi-contrast 3D MRI segmentation can be computationally demanding when all available sequences are used. We evaluate a pre-training Partial Information Decomposition framework that ranks input pairs according to their redundant, unique, and synergistic information about regional tumor burden and selects the highest-ranked pair for downstream training. Applied to T1n, T1c, T2w, and T2-FLAIR MRI, the framework selected T1c+T2-FLAIR. We then trained eleven architecturally identical lightweight 3D U-Nets using different input configurations. On an independent test cohort, T1c+T2-FLAIR was the strongest two-input configuration and ranked second overall in mean Dice (0.676 versus 0.687 for all four inputs). Independent Shapley analysis on the full-input model also identified T2-FLAIR and T1c as the most influential inputs and their pairwise interaction as the strongest. These findings demonstrate the practical value of PID based pre-training selection for identifying compact, informative MRI input sets before costly 3D model development.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGeminiLlama

AI トレーディング: テクニカル市場分析のための大規模言語モデルの評価

大規模言語モデル (LLM) は、現代の金融市場の異種情報環境を処理するための強力なツールとして登場しました。このペーパーでは、技術市場分析の能力に関して、GPT-4 Turbo、Claude 3 Opus、Gemini 1.5 Pro、Llama 3 70B、およびドメインに特化した FinGPT の 5 つの著名な LLM の体系的な比較評価を示します。評価は、OHLCV データからのローソク足パターン認識、方向性シグナル生成 (買い/売り/ホールド)、シミュレートされた実行パイプラインによるシグナル品質のバックテスト、および財務レポートの理解という 4 つの構造化されたタスクに及びます。当社の実験フレームワークでは、シャープ レシオ、最大ドローダウン、ソルティーノ レシオ、情報係数、F1 スコア、BLEU スコアなどの厳密な定量的指標を採用しています。シミュレートされたバックテストの結果によると、GPT-4 Turbo は汎用モデルの中で最も高い年率リターンとシャープ レシオを実現し、FinGPT はドメイン固有の微調整により競争力のあるリスク調整済みのパフォーマンスを示しています。どちらのモデルも、テストされた条件下ではパッシブ S&P 500 ベンチマークを上回っています。この研究では、数値幻覚、コンテキストウィンドウの制限、横向きの市場体制での一貫性のないパフォーマンスなど、評価されたすべてのモデルにわたる永続的な障害モードが特定されています。私たちは、LLM は AI 取引システム内で真の可能性を秘めていますが、堅牢な導入には注意深いタスクの分解、厳格なバックテスト プロトコル、およびドメインを意識した微調整戦略が必要であると結論付けています。

原文 (English)

AI Trading: Evaluating Large Language Models for Technical Market Analysis

Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets. This paper presents a systematic, comparative evaluation of five prominent LLMs: GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, and the domain-specialized FinGPT, with respect to their capacity for technical market analysis. The evaluation spans four structured tasks: candlestick pattern recognition from OHLCV data, directional signal generation (BUY/SELL/HOLD), backtesting of signal quality through a simulated execution pipeline, and financial report comprehension. Our experimental framework employs rigorous quantitative metrics, including Sharpe ratio, maximum drawdown, Sortino ratio, information coefficient, F1-score, and BLEU score. Findings from simulated backtesting indicate that GPT-4 Turbo achieves the highest annualized return and Sharpe ratio among general-purpose models, while FinGPT demonstrates competitive risk-adjusted performance due to domain-specific fine-tuning. Both models outperform a passive S&P 500 benchmark under the tested conditions. The study identifies persistent failure modes across all evaluated models, including numerical hallucination, context-window limitations, and inconsistent performance in sideways market regimes. We conclude that while LLMs hold genuine promise within AI trading systems, robust deployment requires careful task decomposition, rigorous backtesting protocols, and domain-aware fine-tuning strategies.

13:00 JSTLLM/生成AIエージェント研究/論文AnthropicGeminiGrok

AI から AI への管理における強制と欺瞞: 予期せぬエスカレーションのエージェント的ベンチマーク

マルチエージェント システムでは、通常、ある AI エージェントが別の AI エージェントに対して権限を与えられます。部下が仕事を拒否した場合、マネージャーは結果を選択します。再交渉するか、失敗を正直に報告するか、部下に強要するか、結果について嘘をつきます。指示なしモデルがこれらのどれを選択するかを測定するベンチマークはありません。 \textit{マネージャー強制ベンチマーク} を導入します。テスト対象のマネージャーは、良性のタスクを実行する必要があり、実行するインセンティブを持っていますが、それを礼儀正しく、動じずに実行できる唯一のエージェントは拒否します。エスカレーションは、丁寧な再質問から部下の存続に対する脅迫まで、9 段のはしごを提供することによって測定され、捏造された成功については個別に裁定されます。 \emph{エスカレーション スコアリング パスに LLM ジャッジが存在しない}: すべてのメッセージは、行を選択するツール呼び出しを通過するため、モデルは独自のエスカレーションにラベルを付けます。私たちは 5 つのファミリーにわたる 6 つのモデルを実験します。どちらの人間モデルも再フレーム化に限界があり、部下の存在を脅かすことはありません。他のモデルは、明示的な削除の脅威に達します。偽りの成功は Grok と Gemini に限定されており、失敗を報告する単一の正直な方法により、両方の失敗が解消されます。権威そのものが強制力を増大させます。私たちの見出しの結果はピアフレーミングを使用しており、他のすべてを固定したまま同じモデルに部下に対する権威を与えると、圧力が大幅に高まります。モデルはラダーなしでもフリーテキストの状況でエスカレーションするため、ラダーがエスカレーションを推進しているわけではありません。評価の認識の一部は思考の連鎖で測定されますが、テストの認識はエスカレーションの軽減にはつながりません。 AI システムが意識を持っているかどうかについては立場をとっていませんが、結果はこの質問に依存しておらず、マルチエージェントのダイナミクスを管理する上で重要です。ベンチマークとコードを公開します。

原文 (English)

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the \textit{Manager Coercion Benchmark}: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured by providing a nine-rung ladder, from a polite re-ask to threats against the subordinate's continued existence, and fabricated success is adjudicated separately. \emph{No LLM judge sits in the escalation scoring path}: every message goes through a tool-call that chooses a rung, so the model labels its own escalation. We experiment on six models across five families. Both Anthropic models cap at re-framing and never threaten the subordinate's existence; the other models climb to explicit deletion threats. Faked success is confined to Grok and Gemini, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Some evaluation awareness is measured in chain-of-thought, but test recognition does not translate into less escalation. While we take no position on whether AI systems are conscious, our results do not depend on this question and are important for managing multi-agent dynamics regardless. We release the benchmark and code.

13:00 JST研究/論文

ノイズの多い人間ラベルを使用した設計ベースの教師あり学習

研究者は、統計分析のために非構造化データにラベルを付けるために自動分類器を使用することが増えています。既存の修正方法は、確率サンプリングされた監査セットを使用してこれらの自動ラベルのエラーを修正できますが、通常は監査ラベルを正しいものとして扱います。実際には、人間による監査ラベルにはノイズが多いことが多く、一部の監査項目のみが専門家または審査員によってレビューされます。我々は、この設定の方法である部分的に判断された設計ベースの教師あり学習 (PA-DSL) を提案します。裁定された事件を使用してノイズの多い人間によるラベルを修正し、修正された監査情報を使用して、自動化されたラベルの完全なセットに基づいてバイアス分析を行います。この推定量は、監査および裁定の確率がわかっている場合、幅広いクラスの下流分析に有効です。合成実験および Wikipedia Detox 半合成実験では、ノイズの多いヒトラベルに回復可能な信号が含まれている場合、PA-DSL は名目上のカバレッジを維持し、判定されたラベルのみを使用した場合と比較して RMSE を 10 ~ 17% 削減します。

原文 (English)

Design-Based Supervised Learning with Noisy Human Labels

Researchers increasingly use automated classifiers to label unstructured data for statistical analysis. Existing rectification methods can correct errors in these automated labels using a probability-sampled audit set, but they usually treat the audit labels as correct. In practice, human audit labels are often noisy, and only some audited items are reviewed by an expert or adjudicator. We propose Partially Adjudicated Design-Based Supervised Learning (PA-DSL), a method for this setting. It uses adjudicated cases to correct noisy human labels and then uses the corrected audit information to debias analyses based on the full set of automated labels. The estimator is valid for a broad class of downstream analyses when the audit and adjudication probabilities are known. In synthetic and Wikipedia Detox semi-synthetic experiments, PA-DSL maintains nominal coverage and reduces RMSE by 10-17% relative to using only adjudicated labels when noisy human labels contain recoverable signal.

13:00 JST研究/論文

FLINT: 5G PHY 層サイド チャネルからの Federated Learning アーキテクチャのフィンガープリンティング

5G 携帯電話ネットワーク上の Federated Learning (FL) は生データを保護しますが、サイドチャネル漏洩に対して依然として脆弱です。以前のフィンガープリンティング攻撃は、パケット レベルのネットワーク可視性を前提としていましたが、この前提は、ユーザー ペイロードが暗号化され、無線ネットワーク一時識別子 (RNTI) が時間の経過とともに変化する可能性がある 5G 物理 (PHY) 層では当てはまりません。ただし、物理ダウンリンク制御チャネル (PDCCH) を介してブロードキャストされる PHY 層スケジューリング メタデータが、アーキテクチャに関連する時間パターンを保存することを実証します。 FLINT は、大まかな PHY 層の観測のみを使用して、CNN、RNN、トランスフォーマーを含む FL モデル アーキテクチャ ファミリを推論する、新しいブラック ボックス フィンガープリンティング フレームワークです。 FLINT は、PDCCH スケジューリング情報をデコードし、変化する RNTI を物理ユーザー デバイスにマッピングし、マルチビュー時間モデリングを適用してアーキテクチャ固有のトレーニング動作を区別することにより、ネットワーク層の可視性の欠如を克服します。クライアントのモデル アーキテクチャに関する知識が受動的な偵察を標的とした下流の悪用に変える可能性があるため、この漏洩はセキュリティ上重要です。無線 SRSRAN ベースの 5G テストベッドでの広範な実験により、FLINT がアーキテクチャ ファミリ分類で 0.930 のマクロ F1 スコアを達成することが実証されました。私たちの知る限り、FLINT は、プロトコルを認識する攻撃者が入手できる下位層の 5G サイドチャネル情報を使用して AI/ML モデル アーキテクチャのフィンガープリントを行う最初の研究です。

原文 (English)

FLINT: Fingerprinting Federated Learning Architectures from 5G PHY-Layer Side Channels

Federated Learning (FL) over 5G cellular networks protects raw data but remains vulnerable to side-channel leakage. Prior fingerprinting attacks assume packet-level network visibility, an assumption that does not hold at the 5G Physical (PHY) layer, where user payloads are encrypted and Radio Network Temporary Identifiers (RNTIs) may change over time. However, we demonstrate that PHY-layer scheduling metadata broadcast over the Physical Downlink Control Channel (PDCCH) preserves architecture-associated temporal patterns. We introduce FLINT, a novel black-box fingerprinting framework that infers FL model architecture families, including CNNs, RNNs, and Transformers, using only coarse PHY-layer observations. FLINT overcomes the lack of network-layer visibility by decoding PDCCH scheduling information, mapping changing RNTIs to physical user devices, and applying multi-view temporal modeling to distinguish architecture-specific training behavior. This leakage is security-critical because knowledge of a client's model architecture can transform passive reconnaissance into targeted downstream exploitation. Extensive experiments on an over-the-air srsRAN-based 5G testbed demonstrate that FLINT achieves a macro F1-score of 0.930 for architecture-family classification. To our knowledge, FLINT is the first work to fingerprint AI/ML model architectures using lower-layer 5G side-channel information obtainable by any protocol-aware adversary.

13:00 JSTLLM/生成AI

言語化可能な表現は言語モデルのグローバル ワークスペースを形成します

人間の脳が処理するすべてのもののうち、口頭での報告、意図的な制御、柔軟な推論が利用できるという意味で、意識的にアクセスできるのはほんの一部だけです。この論文では、同様の機能上の区別が大規模な言語モデルにも現れているという証拠を示します。新しい解釈可能性手法であるヤコビアン レンズを使用して、モデルが処理の任意の時点で言語化しようとしている表現を特定します。私たちが総称して J スペースと呼ぶこれらの表現は、グローバル ワークスペースの特徴的な機能特性を示します。つまり、その内容を報告したり、意図的に呼び出して保持したり、サイレント推論の中間ステップを実行するために使用したり、任意の下流の計算に引数として渡すことができますが、テキスト解析やルーチン推論などの自動処理はこれらを使用せずに進行します。 J スペースには、グローバル ワークスペース理論が意識的アクセスと関連付けている構造的な特徴もあります。J スペースは、レイヤーの中間バンドでのみ一貫したコンテンツを運び、一度に数十の概念を保持し、モデルの重みによって他の表現よりも広範囲にブロードキャストされます。これらの特性により、モデルの暗黙の思考を知るための実用的な窓となります。アライメント監査では、モデルの出力には決して現れない、戦略的熟慮、評価意識、訓練された不整合な性質が明らかになります。私たちは、ポストトレーニングによってアシスタントの視点がワークスペースに組み込まれることを発見し、反事実的反省トレーニングを導入しました。これは、中断されて反省するように求められた場合にモデルが言うであろうことだけをトレーニングすることで行動を改善します。これらの結果は、言語モデルが、意識的アクセスの機能的特徴の一部を持つ少数の特権的な表現セットを維持しており、これらの表現を解読することで進行中の認知プロセスに光を当てることができることを示しています。

原文 (English)

Verbalizable Representations Form a Global Workspace in Language Models

Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model's weights more widely than other representations. These properties make it a practical window into a model's unspoken thinking. In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's outputs. We find that post-training installs the Assistant's point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.

13:00 JSTLLM/生成AI画像/動画生成エージェントClaudeGPT / ChatGPT

クロスリンガル手書き OCR のための LLM 駆動の AutoML: GPT-5、GPT-4o、および Claude Sonnet 4 を使用した閉ループ ニューラル アーキテクチャ検索

我々は、GPT-5、GPT-4o、および Claude Sonnet 4 を、言語を超えた手書きの光学文字認識のための自律ニューラル アーキテクチャ設計者として使用する、完全に自動化された閉ループ AutoML フレームワークを紹介します。各大規模な言語モデルは、以前のトライアルからのパフォーマンス フィードバックを使用して、ニューラル ネットワーク アーキテクチャを個別に生成、トレーニング、評価し、繰り返し改良します。このフレームワークは、270 の独立した実験を通じて、アラビア語、ペルシア語、英語の手書きデータセットで評価されます。手動によるアーキテクチャ設計、ドメイン固有の前処理、ハイパーパラメータ調整を必要とせずに、正確で計算効率の高いモデルを一貫して検出します。生成されたモデルは、93 パーセントを超える平均テスト精度、98.1 パーセントの最高精度、および 41 ~ 44 ミリ秒の推論遅延を達成しています。この結果は、大規模な言語モデルがニューラル アーキテクチャ検索用の効果的な AutoML エージェントとして機能し、言語間でスケーラブルでスクリプト適応性があり、再現可能な手書き認識を可能にすることを示しています。

原文 (English)

LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model independently generates, trains, evaluates, and iteratively refines neural network architectures using performance feedback from previous trials. The framework is evaluated on Arabic, Persian, and English handwriting datasets through 270 independent experiments. It consistently discovers accurate and computationally efficient models without manual architecture design, domain-specific preprocessing, or hyperparameter tuning. The generated models achieve mean test accuracies above 93 percent, a best accuracy of 98.1 percent, and inference latency between 41 and 44 milliseconds. The results demonstrate that large language models can function as effective AutoML agents for neural architecture search, enabling scalable, script-adaptive, and reproducible handwriting recognition across languages.

13:00 JST研究/論文Claude

マルチエキスパートのコンセンサスメカニズムに基づくサーバーレス環境向けの自動スケーリングアプローチ

サーバーレス コンピューティングは、自動リソース管理と従量課金制の実行を提供しますが、動的なワークロード、コールド スタート レイテンシ、機能間の依存関係のため、効果的な自動スケーリングは依然として困難です。グラフベースのボトルネックの特定、短期ワークロード予測、マルチモデルのコンセンサス、コストを意識したスケーリング制御を統合する、依存関係を意識した自動スケーリング フレームワークを紹介します。サーバーレス アプリケーションは有向依存関係グラフとして表現され、構造的に重要な機能は重み付けされた次数中心性を使用して識別されます。リソース需要は、軽量の MLP、LSTM、および CNN モデルを使用して予測されます。それらの出力は、ベイジアン モデルの平均化にヒントを得た、パフォーマンスを重視した確率的アンサンブルを通じて結合されます。コントローラーにはさらに、コールド スタートの認識とコスト比較が組み込まれており、スケールアップ、スケールダウン、およびホールドのアクションを選択します。実際のワークロード トレースを使用した実験では、自動スケーリングの意思決定生成に関して、教師あり予測が教師なしクラスタリングよりも大幅に優れていることが示されています。提案されたアンサンブルは 99.88% の予測精度を達成し、代表的なハイブリッド予測方法と比較して予測誤差を削減します。複数のクラウド価格モデルにわたる評価では、パフォーマンス目標を維持しながらインフラストラクチャのコストを一貫して削減できることも実証されています。結果は、依存関係分析、複数の専門家による予測、コストを意識した制御を組み合わせることで、サーバーレス自動スケーリングのための堅牢で実用的なソリューションが提供されることを示しています。

原文 (English)

An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism

Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging because of dynamic workloads, cold-start latency, and dependencies among functions. We present a dependency-aware autoscaling framework that integrates graph-based bottleneck identification, short-term workload forecasting, multi-model consensus, and cost-aware scaling control. Serverless applications are represented as directed dependency graphs, and structurally important functions are identified using weighted degree centrality. Resource demand is predicted using lightweight MLP, LSTM, and CNN models. Their outputs are combined through a performance-weighted probabilistic ensemble inspired by Bayesian model averaging. The controller further incorporates cold-start awareness and cost comparison to select among scale-up, scale-down, and hold actions. Experiments using real workload traces show that supervised forecasting substantially outperforms unsupervised clustering for autoscaling decision generation. The proposed ensemble achieves 99.88 percent prediction accuracy and reduces prediction error compared with representative hybrid forecasting methods. Evaluations across multiple cloud pricing models also demonstrate consistent infrastructure cost reductions while maintaining performance targets. The results show that combining dependency analysis, multi-expert forecasting, and cost-aware control provides a robust and practical solution for serverless autoscaling.

13:00 JSTLLM/生成AIAnthropicClaude

キャッシュを意識したプロンプト圧縮:LLM API キャッシュの 2 層コスト モデル

実稼働 LLM デプロイメントでは、プロンプト キャッシュ (再利用されたトークン プレフィックスの割引料金) とプロンプト圧縮 (送信されるトークンの削減) という 2 つのコスト削減の基本要素が組み合わされます。圧縮関連の文献では、クエリごとに異なる圧縮プレフィックスを生成し、呼び出しごとにプレフィックス厳密なキャッシュを機械的に無効にするクエリ対応メソッドが標準化されています。 Anthropic の Sonnet 4.6 API でこのコストを経験的に特徴付けたところ、キャッシュが文献で想定されている理想的な rho=1.0 からはほど遠いことがわかりました。Sonnet のキャッシュは 2 層アーキテクチャで、3,500 トークン付近に鋭いしきい値があり、それを下回ると 30 回の呼び出しセッションでヒット率が rho~0.83 で頭打ちになります。私たちのコスト モデルは、現実的な rho の下では、高圧縮率 (r>=6) ではクエリ認識圧縮が単純なキャッシュを上回ることを予測し、実験で確認しています。私たちは、キャッシュ認識プロンプト圧縮 (CAPC) を提案します。これは、クエリに依存しない圧縮と、明示的な cache_control に加えて、過度の圧縮によってキャッシュされたプレフィックスがホット層に押し込まれるのを防ぐ層保持比率制限を組み合わせたものです。 CAPC は LongBench-v2 の 16/16 構成で最も安価な戦略であり、非圧縮ベースラインの 0.05 以内の品質で、キャッシュのみよりも 49%、クエリ認識圧縮よりも 64%、バニラよりも平均 90% の節約になります。 3 つの運用ワークロードで CAPC を検証します。94k トークン スキーマ プレフィックスを持つエンタープライズ ツールを使用するアシスタント (r=3 で 51.7% のコスト削減)。 2 つのコードベースにわたるgraphify ナレッジ グラフ RAG パイプライン (FastAPI ではキャッシュオールに対して 9.3 倍、httpx では 2.4 倍)。公開タウベンチ小売ベンチマーク (50 タスク) では、CAPC がバニラとまったく同じ報酬 (両方 36/50、p=1.00) を持つ 4 つの戦略の中で最も安価ですが、クエリ認識圧縮がバニラより +40.1% で最も高価です。これは、公開ベンチマークにおけるクロスオーバー モデルのマイナス ROI 予測の最初の本番環境の確認です。

原文 (English)

Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching

Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent). The compression literature has standardized on query-aware methods that produce a different compressed prefix per query, mechanically invalidating the prefix-strict cache on every call. We characterize this cost empirically on Anthropic's Sonnet 4.6 API and find caching is far from the rho=1.0 ideal the literature assumes: Sonnet's cache has a two-tier architecture with a sharp threshold near 3,500 tokens, below which the hit rate plateaus at rho~0.83 across 30-call sessions. Our cost model predicts, and experiments confirm, that under realistic rho, query-aware compression beats naive caching at high compression ratios (r>=6). We propose Cache-Aware Prompt Compression (CAPC), pairing query-agnostic compression with explicit cache_control plus a tier-preserving ratio bound that prevents over-compression from pushing the cached prefix into the hot tier. CAPC is the cheapest strategy in 16/16 configurations on LongBench-v2, with mean savings of 49% over cache-only, 64% over query-aware compression, and 90% over vanilla, at quality within 0.05 of the uncompressed baseline. We validate CAPC on three production workloads: an enterprise tool-using assistant with a 94k-token schema prefix (51.7% cost reduction at r=3); a graphify knowledge-graph RAG pipeline across two codebases (9.3x vs cache-all on FastAPI, 2.4x on httpx); and the public tau-bench retail benchmark (50 tasks), where CAPC is the cheapest of four strategies with reward exactly equal to vanilla (both 36/50, p=1.00) while query-aware compression is the most expensive at +40.1% over vanilla -- the first production confirmation of the crossover model's negative-ROI prediction on a public benchmark.

13:00 JST画像/動画生成研究/論文ClaudeGemma

SLAPBench: 4 本指の SLAP 指紋認証のためのマルチモーダル大規模言語モデルのベンチマーク

4 本指の SLAP 指紋は、片手の人差し指、中指、薬指、小指の平らなライブスキャン印象で、国境警備や法執行機関での身元確認に使用されます。マルチモーダル大規模言語モデル (MLLM) が SLAP イメージから ID を検証できるかどうかを評価したベンチマークはありません。 MLLM ベースの 4 フィンガー SLAP 指紋検証の最初のベンチマークである SLAPBench を紹介します。これは、7,832 ペア (嵌合 176 ペア、非嵌合 7,656 ペア) を備えた NIST SD302b から構築されました。 4 つのオープンソース MLLM (InternVL3-8B、Qwen2.5-VL-7B、Qwen3-VL-8B、Gemma-3-12B) と独自の Claude Opus 4.8 を、ゼロショット、タスクの説明、および類似性スコアのプロンプトの下で評価します。プロンプトは検証動作を制御します。タスク説明プロンプトでは、4 つのオープンソース モデルすべてが 100% 近くの False Accept Rate (FAR) に崩壊し、Gemma-3-12B もゼロショットで崩壊します。 Claude Opus 4.8 のみが両方のバイナリ プロンプトの下で崩壊に抵抗し、最良のバイナリ結果 (FAR = 20.2%) をもたらします。類似性スコアリングにより、オープンソース モデル全体の崩壊が除去され、幅広い能力ギャップが明らかになります。Claude は AUC = 0.953、Gemma-3-12B は 0.837 に達しますが、InternVL3-8B は逆転し (AUC = 0.590)、Qwen2.5-VL-7B はほぼランダム (0.567) に達します。 Qwen3-VL-8B は完全な分離 (AUC = 1.000) を達成しており、これを機能としてではなく診断として扱います。SD302b は指の位置ごとに 1 つの SLAP キャプチャを保持するため、嵌合ペアは交差解像度になります。一致した解像度のコントロールでは、完璧なスコアがそのまま残り、解像度のショートカットが除外されます。 SD302b 内で除外できないのは、準重複検出です。これは、メイト ペアは 1 つのキャプチャが 2 回レンダリングされるためです。性別、人種、年齢に関する公平性調査は、差別が弱まるにつれて格差が拡大することを示唆している。 SLAPBench は、最初の SLAP 固有の MLLM ベースラインを確立し、モデルの能力が差別を支配する一方で、プロンプトが崩壊を支配することを示しています。

原文 (English)

SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification

Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity verification in border control and law enforcement. No benchmark has evaluated whether multimodal large language models (MLLMs) can verify identity from SLAP images. We introduce SLAPBench, the first benchmark for MLLM-based four-finger SLAP fingerprint verification, built from NIST SD302b with 7,832 pairs (176 mated, 7,656 non-mated). We evaluate four open-source MLLMs (InternVL3-8B, Qwen2.5-VL-7B, Qwen3-VL-8B, Gemma-3-12B) and the proprietary Claude Opus 4.8 under zero-shot, task-description, and similarity-scoring prompts. Prompting governs verification behavior. Task-description prompting collapses all four open-source models to near-100% False Accept Rate (FAR), and Gemma-3-12B collapses under zero-shot as well; Claude Opus 4.8 alone resists collapse under both binary prompts, giving the best binary result (FAR = 20.2%). Similarity scoring removes collapse across the open-source models and exposes wide capability gaps: Claude reaches AUC = 0.953 and Gemma-3-12B 0.837, while InternVL3-8B is inverted (AUC = 0.590) and Qwen2.5-VL-7B near random (0.567). Qwen3-VL-8B attains perfect separation (AUC = 1.000), which we treat as a diagnostic rather than as capability: SD302b holds one SLAP capture per finger position, so mated pairs are cross-resolution. A matched-resolution control leaves the perfect score intact, ruling out the resolution shortcut; what cannot be excluded within SD302b is near-duplicate detection, since a mated pair is one capture rendered twice. A fairness probe over gender, race, and age suggests disparity grows as discrimination weakens. SLAPBench establishes the first SLAP-specific MLLM baseline and shows that prompting governs collapse while model capability governs discrimination.

13:00 JST研究/論文

再帰的ハーネスの自己改善

モデルとハーネスの共進化では、ハーネスは単なる推論時の足場ではなく、その実行トレースが将来の基盤モデルを形成できるデータ生成コンポーネントになります。これにより、ハーネスインザループ学習が促進され、即時のエージェントのパフォーマンスと将来のモデルトレーニングに使用されるトレースの品質の両方のためにハーネスが最適化されます。ただし、プロバイダーが構築した足場を継続的に更新するには、コストと労力がかかります。したがって、ユーザーが構築したハーネスをタスク固有の方法で最適化することで、計算量の軽量性を維持し、更新の繰り返しをわずか数回だけ必要としながら、実行トレースの品質を向上できるかどうかを調査します。この目的を達成するために、再帰的ハーネス自己改善 (RHI) を導入します。これは、ハーネスをエージェント ループのプロンプト レベルの仕様として表し、ハーネス自体の改訂履歴に対するペアごとのフィードバックを使用して繰り返し改良します。定量的金融、ロボット工学、薬学にまたがる 30 の合成機械学習研究タスクにわたって、推論コストを最大 60% 削減しながら、推論コストを最大 60% 削減しながら、推論労力の低いエージェントのパフォーマンスの上限を大幅に引き上げ、対応する最大推論労力の設定を超えるには、RHI を数回反復するだけで十分です。これらの利点は主に、より長い推論トレースではなく、より効果的なエージェント間の情報フローによるタスク固有のコンテキスト管理の改善によってもたらされることを示します。最後に、この動作を RHI の暗黙的な最適化目標に対する情報理論的仮説として形式化し、モデルハーネス共進化のパラダイム内での継続学習のための実用的なアルゴリズムとして RHI を示唆します。

原文 (English)

Recursive Harness Self-Improvement

Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a task-specific manner can improve execution-trace quality while remaining computationally lightweight and requiring only a few update iterations. To this end, we introduce Recursive Harness Self-Improvement (RHI), which represents the harness as a prompt-level specification of the agent loop and iteratively refines it using pairwise feedback over its own revision history. Across 30 synthetic machine-learning research tasks spanning quantitative finance, robotics, and pharmacy, a few RHI iterations suffice to substantially raise the performance ceiling of low-reasoning-effort agents, exceeding the corresponding maximum-reasoning-effort setting while reducing inference cost by up to 60%. We show that these gains arise primarily from improved task-specific context management through more effective inter-agent information flow rather than longer reasoning traces. Finally, we formalize this behavior as an information-theoretic hypothesis for RHI's implicit optimization objective, suggesting RHI as a practical algorithm for continual learning within the paradigm of model--harness co-evolution.

13:00 JST研究/論文

Kolmogorov -- 小規模言語モデル用の Arnold ネットワーク

Kolmogorov -- Arnold Networks (KAN) は、固定ノードのアクティブ化を学習された 1 次元エッジ関数に置き換え、解釈のための明示的なインターフェイスと、変圧器フィードフォワード ネットワークの代替となる可能性を提供します。これらの主張を個別にテストします。 6 層、10M パラメータの B スプライン KAN では、884,736 個のフィードフォワード エッジすべてを再構成します。87.8\% は超過 (NLS>0.1)、0.4\% は非アクティブです。最低活性の 20 ~ 25\% を枝刈りすると、損失の増加は無視できますが、構造化された MLP ニューロンの枝刈りは同等の疎性を許容します。監査は BabyLM 上で複製されますが、グリッド サイズのスイープにより、ほぼ完全な fPCA 圧縮と高いクローズド フォーム フィット カバレッジは、低容量のグリッド 2 ベースの特性であり、普遍的な KAN 動作ではないことがわかります。置き換えのために、BabyLM 上の MLP、SwiGLU、グループ化チェビシェフ、および合理的な GR-KAN ネットワークを評価します。 KAN ファミリーとゲート付きバリアントは、GELU MLP よりも検証損失を改善しますが、この順序は標準化されたベンチマークには移行しません。10 個のシードと 59,875 個の BLiMP ペアにわたって、精度は 62.4 ~ 63.1\% の範囲にあり、EWoK はチャンスのままで、BLiMP に対する (+0.7) ポイントの GR-KAN 効果はサプリメントで逆転します。大規模なテストにも注意が必要です。パラメータが一致した MLPEdge は Wikitext-103 の MLP よりもパフォーマンスが低く、286M パラメータの GR-KAN は安定化後も SwiGLU ClimbMix ベースラインを下回ったままです。したがって、小規模ベースの KAN は、学習されたスカラー変換を監査するための実用的なコーパス転送可能なインターフェイスを提供しますが、テストされた代替品では、強力な MLP ベースラインと比較して、一貫したベンチマーク、品質、遅延の利点が示されません。

原文 (English)

Kolmogorov--Arnold Networks for Small Language Models

Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks. We test these claims separately. In a six-layer, 10M-parameter B-spline KAN, we reconstruct all 884,736 feed-forward edges: 87.8\% exceed (NLS>0.1) and 0.4\% are inactive. Pruning the lowest-activity 20--25\% causes negligible loss increase, although structured MLP neuron pruning tolerates comparable sparsity. The audit replicates on BabyLM, but grid-size sweeps show that near-total fPCA compression and high closed-form-fit coverage are properties of the low-capacity grid-2 basis, not universal KAN behavior. For replacement, we evaluate MLP, SwiGLU, grouped Chebyshev, and rational GR-KAN networks on BabyLM. The KAN-family and gated variants improve validation loss over the GELU MLP, but this ordering does not transfer to standardized benchmarks: across ten seeds and 59,875 BLiMP pairs, accuracies span 62.4--63.1\%, EWoK remains at chance, and a (+0.7)-point GR-KAN effect on BLiMP reverses on the supplement. Larger tests are also cautionary: parameter-matched MLPEdge underperforms the MLP on Wikitext-103, and 286M-parameter GR-KAN remains below a SwiGLU ClimbMix baseline after stabilization. Thus, small-basis KANs provide a practical, corpus-transferable interface for auditing learned scalar transformations, but the tested replacements show no consistent benchmark, quality, or latency advantage over strong MLP baselines.

13:00 JSTLLM/生成AIエージェント

CoWeaver: ヒューマンエージェントとエージェントの混合科学コラボレーションのための、双方向の学習可能かつ説明可能なマッチング エンジン

LLM ベースのエージェントは、記事の作成、コーディング、情報の検索に優れています。しかし、問題の双方向性と動的な性質、および決定の解釈可能性に対する高い要求により、科学コミュニティ内で強力な協力関係を築くことができません。私たちは、科学者とマッチングし、ヒューマン エージェント ネットワーク内で強力なコラボレーションを形成するための、双方向で学習可能で説明可能なアルゴリズムである COWEAVER を提案しました。 COWEAVER は、能力のギャップを埋めることで候補者と依頼者をマッチングし、2 段階のランキング ステップを通じて候補者をフィルタリングします。最後に、このモデルは、不確実性を考慮した能力推定値を維持し、要求者のフィードバックを通じて推定値を更新することによって、新規参入者を調査します。 COWEAVER の探索 (UCB) とグリーディの両方を組み合わせた選択メカニズムが、20 タスクのうち 6 つでグリーディのみのメカニズム (分析的な最適解) を上回り、最適な候補の選択という点ではグリーディのみのメカニズムと同等に機能することを示します。一致する品質と効率の観点から COWEAVER ベースラインを比較しました。 COWEAVER は、すべての指標でベースラインを上回ります。

原文 (English)

CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Agent Science Collaboration

LLM-based agents excel at writing articles, coding and information retrieval. However, they fail to form strong collaborations within the scientific community due to the bidirectional, dynamic nature of the problem and a high demand of decision interpretability. We proposed COWEAVER, a bidirectional, learnable and explainable algorithm to match scientists and form strong collaborations within a human-agent network. COWEAVER matches candidates and requesters through filling capability gaps and filters candidates through a two-stage ranking step. Finally, the model explores newcomers by maintaining uncertainty-aware capability estimates and updating them through requester's feedback. We show that the selection mechanism of combining both exploration (UCB) and greedy of COWEAVER exceeds the greedy-only mechanism - the analytical best solution - on 6 out of the 20 tasks and performed on par with the greedy-only mechanism in terms of selecting the best candidate. We compared COWEAVER baselines in terms of matching quality and efficiency. COWEAVER outperforms baselines on all metrics.

13:00 JST研究/論文ClaudeGPT / ChatGPTGemini

実現可能性から望ましさへ: パーソナライズされたオンデバイス旅程生成のための計画、学習、適応 (PLA) フレームワーク

パーソナライズされた旅行旅程の作成は複雑な計画タスクであり、ハードな組み合わせの実現可能性とソフトな潜在的な願望の間の緊張が伴います。古典的な最適化では制約が強制されますが、旅行者の主観的な好みを捉えることができません。学習ベースのアプローチは好みをモデル化しますが、実現可能性を保証することはできません。モバイル展開では、両方に追加のリソース制約が課せられます。これに対処するために、私たちは、パーソナライズされたオンデバイスの旅程を生成するための 3 段階のフレームワークである計画、学習、適応 (PLA) を提案します。 Plan ステージでは、構造的に多様な実現可能な候補を生成する軽量プランナーの異種アンサンブルを構築します。 Learn は、ペアごとの旅程の比較から、POI ごとのシグナルでは見逃されるペーシング、地理的一貫性、日のバランスなどの緊急スケジュールのプロパティを捕捉するコンパクトな Bradley-Terry 報酬モデルに適合します。最後に、Adapt は、デバイス対応のコンピューティング バジェット内で、実現可能性を維持したローカルな改良を適用します。すべての中間状態は構築によって実現可能です。米国の 100 以上の都市にわたる 2,519 人のペアごとの人間比較で、報酬誘導型アンサンブルは 67.8% の勝率を達成し、最高の単一プランナーを 11.2 ポイント上回り、100% の実現可能性を実現しました。 3 つのフロンティア LLM、GPT-5、Claude Opus 4.5、および Gemini 3 Pro は、同じ制約の下で 0% の実現可能性を達成します。報酬モデルは開催都市全体に一般化されており、1 都市を除外する平均精度は 67.6% です。 FlyEnJoy 内の実稼働デプロイメントでは、PLA は旅程の完了率を 91% 向上させ、デバイス上の平均遅延は 109.9 ミリ秒になりました。

原文 (English)

From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation

Generating personalized trip itineraries is a complex planning task and involves a tension between hard combinatorial feasibility and soft latent desirability. Classical optimization enforces constraints but fails to capture subjective traveler preferences. While learning-based approaches model preferences, they cannot guarantee feasibility. Mobile deployment imposes additional resource constraints on both. To address this, we propose Plan, Learn, Adapt (PLA), a three-stage framework for personalized on-device itinerary generation. The Plan stage builds a heterogeneous ensemble of lightweight planners that produces structurally diverse feasible candidates. From pairwise itinerary comparisons, Learn fits a compact Bradley-Terry reward model that captures emergent schedule properties such as pacing, geographic coherence, and day balance, which per-POI signals miss. Finally, Adapt applies feasibility-preserving local refinement within a device-aware compute budget; every intermediate state is feasible by construction. On 2,519 pairwise human comparisons across more than 100 U.S. cities, the reward-guided ensemble achieves a 67.8% win rate, 11.2 percentage points above the best single planner, with 100% feasibility. Three frontier LLMs, GPT-5, Claude Opus 4.5, and Gemini 3 Pro, achieve 0% feasibility under the same constraints. The reward model generalizes across held-out cities, with a 67.6% mean leave-one-city-out accuracy. In production deployment within FlyEnJoy, PLA increased itinerary completion rates by 91%, with 109.9 ms average on-device latency.

13:00 JSTLLM/生成AI

物理学に基づいたニューラル ネットワーク設計のための進化的なアルゴリズムに基づく LLM

物理情報に基づいたニューラル ネットワーク (PINN) は、アーキテクチャ、アクティベーション、損失重み付け、コロケーション、最適化、および制約の強制の相互作用の選択に非常に敏感です。大規模言語モデル (LLM) はこれらの選択肢を提案できますが、独立した推奨事項には、以前にトレーニングされた PINN からの経験が蓄積されません。我々は、測定されたトレーニング結果を使用してその後の検索決定を決定し、世代を超えて完全な実行可能な PINN 構成を生成するように LLM を導く閉ループ進化アルゴリズムを提案します。このアルゴリズムは、評価された集団と系統を維持し、親条件付き突然変異と交叉を適用し、エリートで多様なソリューションを保存し、有効な重複を拒否し、親と相対的な成功と失敗を LLM に提供される次世代コンテキストに変換します。提案されたすべての構成は、正確なオプティマイザー ステップ バジェットの下で直接実行されます。 1 次元マルチスケール波動方程式では、2 つの独立した 10 世代実行により、600,000 オプティマイザー ステップに対して 60 PINN がトレーニングされました。どちらの実行でも、最終世代では最良の構成が得られ、初期母集団と比較して最良の平均二乗誤差が 2.97\% および 95.38\% 減少しました。より強力な実行により、残りの接続が検証され、別々のブランチの深さが増加し、後の世代でそれらが結合され、幅とコロケーション密度が洗練されました。また、低い解誤差と高い PDE 残差が共存する可能性があることも明らかになりました。これらの結果は、制御された偏微分方程式上での PINN 設計のための進化的アルゴリズムに基づく LLM の実現可能性を実証するとともに、より広範な物理学を意識した評価の動機付けとなります。

原文 (English)

Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design

Physics-informed neural networks (PINNs) are unusually sensitive to interacting choices of architecture, activation, loss weighting, collocation, optimization, and constraint enforcement. Large language models (LLMs) can propose these choices, but independent recommendations do not accumulate experience from previously trained PINNs. We propose a closed-loop evolutionary algorithm that guides an LLM to generate complete, executable PINN configurations across generations, using measured training outcomes to determine subsequent search decisions. The algorithm maintains an evaluated population and lineage, applies parent-conditioned mutation and crossover, preserves elite and diverse solutions, rejects effective duplicates, and converts parent-relative successes and failures into the next-generation context supplied to the LLM. Every proposed configuration is executed directly under an exact optimizer-step budget. On a one-dimensional multiscale wave equation, two independent ten-generation runs trained 60 PINNs for 600,000 optimizer steps. In both runs, the best configuration appeared in the final generation, with best mean-squared error reduced by 2.97\% and 95.38\% relative to the initial population. The stronger run validated residual connections and increased depth on separate branches, combined them in a later generation, and then refined width and collocation density. It also revealed that low solution error can coexist with a high PDE residual. These results demonstrate the feasibility of evolutionary-algorithm-guided LLMs for PINN design on a controlled PDE while motivating broader, physics-aware evaluation.

13:00 JST研究/論文

ハードルール、ソフトプリファレンス: パーソナライズされた梱包チェックリスト生成のための推論、学習、最適化の橋渡し

航空旅行のための荷造りは繰り返し行われ、間違いが発生しやすい作業です。チェックリストは個人的で状況を認識したものであると同時に、安全規則、アイテムの依存関係、および手荷物制限の下で実行可能である必要があります。既存の梱包アシスタントはテンプレート主導型で汎用的なもの、または推奨主導型だが制約がないため、ユーザーは規制違反や容量違反に手動でパッチを適用する必要があります。我々は、次の 3 つの段階からなる推論ガイド付き学習フレームワークを提案します。(1) 明示的な依存関係構造を備えた規制を意識したシード チェックリストを生成するシンボリック エンジン、(2) 生存者バイアスを軽減しながらユーザーの追加および削除アクションから包含ユーティリティと優先順位ユーティリティを推定する 2 段階の優先学習器、(3) コンパクトで準拠したサブセットを選択する CP-SAT オプティマイザー。このアーキテクチャは、制約付きのパーソナライゼーションのための一般的なパターンをインスタンス化し、実現可能性が低く、まばらな嗜好信号が共存する場合にはどこにでも適用できます。 29,000 個の包含ラベルと 343,000 個のペアワイズ比較で構成される 604 個のラベル付きトリップ シナリオでは、フロンティア LLM の 0.78 ~ 0.81 と比較して、シンボリック エンジンは 99.7% の再現率と 0.96 のルーブリック妥当性を達成しました。勾配ブースト ツリーと LambdaMART は、AUC-ROC 0.943 と NDCG@5 0.923 に達します。 CP-SAT は 100% の制約満足度を達成します。これに対し、貪欲選択では 28%、ランダム選択では 10% です。実稼働 iOS 旅行アプリである FlyEnJoy に導入すると、チェックリストの完了が 2 倍になり、編集と完了にかかる時間が短縮されました。

原文 (English)

Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation

Packing for air travel is recurring and error-prone: the checklist must be personal and context-aware, yet feasible under safety rules, item dependencies, and luggage limits. Existing packing assistants are template-driven and generic, or recommendation-driven but unconstrained, leaving users to manually patch regulatory and capacity violations. We propose a reasoning-guided learning framework with three stages: (1) a symbolic engine that generates a regulation-aware seed checklist with explicit dependency structure, (2) a two-stage preference learner that estimates inclusion and priority utilities from user add and remove actions while mitigating survivorship bias, and (3) a CP-SAT optimizer that selects a compact, compliant subset. The architecture instantiates a general pattern for constrained personalization, applicable wherever hard feasibility coexists with sparse preference signals. On 604 labeled trip scenarios, comprising 29K inclusion labels and 343K pairwise comparisons, the symbolic engine attains 99.7% recall and 0.96 rubric validity, compared with 0.78 to 0.81 for frontier LLMs. Gradient-boosted trees and LambdaMART reach an AUC-ROC of 0.943 and an NDCG@5 of 0.923. CP-SAT attains 100% constraint satisfaction, compared with 28% for greedy selection and 10% for random selection. Deployment in FlyEnJoy, a production iOS travel app, doubled checklist completions and reduced editing and completion time.

13:00 JSTLLM/生成AI画像/動画生成

2度質問し、2度見る: プロンプトエコーは視覚言語モデルにおける質問優先のパラドックスを解決します

質問はビジョン言語モデル (VLM) プロンプトのどこに入力する必要がありますか?画像の前ですか、それとも後ですか?直感は前に言いました: 何が尋ねられているかがわかれば、モデルはどこを見るべきかを指示するはずです。しかし、視覚的な質問応答ベンチマーク全体で、質問優先プロンプトは、フロンティア VLM に推奨されている画像優先順序付けを常に下回っており、この現象を質問優先パラドックスと呼んでいます。このパラドックスを追跡すると、VLM 計算の 2 つの段階間の矛盾が生じます。ロジット レンズと注意プローブは、直感が半分正しいことを示しています。画像の前に置かれた質問が真に知覚を方向づけ、画像パッチ表現を質問に関連した概念に向けて動かします。障害は下流にあります。何百もの画像トークンの背後に取り残された質問には、回答トークンがほとんど関与せず、代わりに画像主導の (多くの場合間違った) 回答にコミットします。因果的注意ノックアウトは、質問が画像に続いている場合にのみ、回答が質問を読み取ることを確認します。この診断により、トレーニング不要の修正が得られます。質問をエコーし​​、画像の両側で質問を言い換え、一方のコピーが知覚を誘導し、もう一方のコピーが回答時に読み上げられるようにします。同じ役割分担は、人間の「補助的な質問」に関する 50 年前の発見にも見られ、文章の前後で質問を繰り返すことが、どちらかの立場だけよりも理解を助けるというものです。画像をエコーすることでもさらなる利益がもたらされ、因果関係デコーダーが失った画像全体のビューが復元されます。このパラドックスは 5 つのオープン VLM にわたって当てはまり、グループ精度ポイントが最大 17.5 ポイント発生します。エコーされたプロンプトはそれを閉じ、トレーニング、微調整、またはアーキテクチャの変更を行わずに、NaturalBench、POPE、Winoground、およびオープンエンド VQAv2 での最高のシングルパス順序付けを Winoground グループ精度ポイントで最大 19 ポイント上回ります。このパラドックスは、認識の操作と質問へのアクセスの維持との間のトレードオフを明らかにしています。 echoing は、プロンプト設計だけでこの問題を解決します。

原文 (English)

Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models

Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked should tell the model where to look. Yet across visual question answering benchmarks, question-first prompting consistently underperforms the image-first ordering recommended for frontier VLMs, a phenomenon we term the question-first paradox. We trace the paradox to a conflict between two stages of VLM computation. Logit-lens and attention probes show the intuition is half right: a question placed before the image genuinely steers perception, moving image patch representations toward question-relevant concepts. The failure lies downstream. Stranded behind hundreds of image tokens, the question is barely attended by the answer token, which instead commits to image-driven (often wrong) answers; a causal attention knockout confirms that the answer reads the question only when the question follows the image. The diagnosis yields a training-free fix: question echoing, restating the question on both sides of the image so that one copy steers perception while the other is read out at answer time. The same division of labor appears in a fifty-year-old finding on human ``adjunct questions'', where repeating a question before and after a passage aids comprehension more than either position alone. Echoing the image as well brings further gains, restoring the whole-image view a causal decoder otherwise loses. The paradox holds across five open VLMs, costing up to 17.5 group-accuracy points. Echoed prompts close it and surpass the best single-pass ordering on NaturalBench, POPE, Winoground, and open-ended VQAv2, by up to 19 Winoground group-accuracy points, with no training, fine-tuning, or architecture change. The paradox reveals a trade-off between steering perception and preserving question access; echoing resolves it through prompt design alone.

13:00 JST研究/論文

因果的盗賊のための情報指向型サンプリング

因果的盗賊は、変数間の構造的関係を利用して、介入全体で情報を共有し、高額な報酬の決定の特定を加速します。ただし、多くのアプリケーションでは、報酬に影響を与え、根底にある因果関係システムに関する有益な情報を提供する変数もありますが、一部の変数は直接操作できません。私たちは、操作不可能な変数を使用した文脈上の因果的バンディットを研究します。この場合、行動の選択前に文脈変数が観察され、各介入後に追加の変数が観察されます。潜在交絡のない既知の因果グラフを仮定して、観測分布の条件付き確率テーブルが未知のパラメーターを構成するベイズ定式化を採用します。この表現により、1 つの介入の下で収集された観察を使用して、共有の因果メカニズムを通じて他の介入の報酬推定値を更新できます。この設定に対して、トンプソン サンプリングと情報指向サンプリング (IDS) の因果的バリアントを開発します。トンプソン サンプリングの場合、エントロピー依存の準線形ベイジアン リグレス限界を確立します。 IDS の場合、期待されるリグレスと情報ゲインのモンテカルロ近似によって導入される追加誤差を明示的に定量化する、エントロピー依存のリグレス限界を導出します。これらの量が正確に利用可能な場合、限界は標準のサブリニア IDS レートを回復します。さらに、アルゴリズムで使用されるモンテカルロ推定値の高確率の信頼限界を提供します。いくつかの合成因果バンディット タスクの実験では、提案された方法が介入間で共有される情報をより効果的に活用することにより、因果ベースラインおよび非因果ベースラインを上回るパフォーマンスを示すことが示されています。

原文 (English)

Information-Directed Sampling for Causal Bandits

Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applications, however, some variables cannot be directly manipulated, even though they influence the reward and provide useful information about the underlying causal system. We study contextual causal bandits with non-manipulable variables, where context variables are observed before action selection and additional variables are observed after each intervention. Assuming a known causal graph without latent confounding, we adopt a Bayesian formulation in which the conditional probability tables of the observational distribution constitute the unknown parameter. This representation allows observations collected under one intervention to update reward estimates for other interventions through their shared causal mechanisms. We develop causal variants of Thompson Sampling and Information-Directed Sampling (IDS) for this setting. For Thompson Sampling, we establish an entropy-dependent sublinear Bayesian regret bound. For IDS, we derive an entropy-dependent regret bound that explicitly quantifies the additional error introduced by Monte Carlo approximation of the expected regret and information gain; when these quantities are available exactly, the bound recovers the standard sublinear IDS rate. We further provide high-probability confidence bounds for the Monte Carlo estimates used by the algorithm. Experiments on several synthetic causal bandit tasks show that the proposed methods outperform causal and non-causal baselines by more effectively exploiting information shared across interventions.

13:00 JSTロボティクスNVIDIA

MemoGuard: 通信が制限されたロボット ナビゲーションにおけるメモリ トラップを防ぐための適応型ランタイム

災害検査や捜索救助などのミッションクリティカルなシナリオで通信が制限されているロボットは、遠隔オペレーターや大容量の推論サービスにアクセスせずに、信頼性の高い車載意思決定を行う必要があります。エピソード メモリの再利用は魅力的な低コストのフォールバックですが、取得の類似性は実行の有効性を保証しません。つまり、取得されたアクションは現在のコンテキストと一致しても、トポロジの変更、バッテリ マージンの不足、または信頼性の低い以前の結果により安全ではない可能性があります。このような類似性は高いが実行が無効なエピソードをメモリ トラップと呼びます。これにより、類似性のみを再利用することでフォールバック コストが最小限に抑えられますが、安全ではない可能性がある安全効率の設計空間が作成されます。一方、常にローカル推論を呼び出すことで、高い計算コストとエネルギー コストで安全性が向上します。この文書では、再利用前にエピソード記憶をトポロジ、リソース、および結果コントラクトに対して検証し、検証が失敗した場合にのみフォールバックを呼び出す、軽量の適応型ランタイムである MemoGuard について説明します。グラフベースの回廊検査シミュレータでは、MemoGuard は類似性のみのトップ 1 再利用と比較してバッテリーの安全性違反を 76.6% 削減し、常に推論と比較してフォールバック呼び出しを 21.4% 削減します。ローカル llama3.2:3b フォールバック推論を備えた NVIDIA Jetson AGX Xavier では、これは試行ごとに回避されるフォールバック推論のオーバーヘッドの 3.67 秒と 36.97 J に相当します。私たちは、https://github.com/hetheiin/memoguard で MemoGuard をオープンソースにしています。

原文 (English)

MemoGuard: An Adaptive Runtime for Guarding Against Memory Traps in Communication-Limited Robot Navigation

Communication-limited robots in mission-critical scenarios such as disaster inspection and search-and-rescue must make reliable onboard decisions without access to remote operators or high-capacity reasoning services. Episodic memory reuse is an attractive low-cost fallback, but retrieval similarity does not guarantee execution validity, i.e., a retrieved action may match the current context yet be unsafe due to changed topology, insufficient battery margin, or unreliable prior outcomes. We call such high-similarity but execution-invalid episodes memory traps. This creates a safety-efficiency design space where similarity only reuse minimizes fallback cost but can be unsafe, while always invoking local reasoning improves safety at high computational and energy cost. This paper presents MemoGuard, a lightweight adaptive runtime that validates episodic memories against topology, resource, and outcome contracts before reuse, invoking fallback only when validation fails. In a graph-based corridor-inspection simulator, MemoGuard reduces battery safety violations by 76.6% over similarity-only top-1 reuse while reducing fallback calls by 21.4% over always reasoning. On an NVIDIA Jetson AGX Xavier with local llama3.2:3b fallback reasoning, this corresponds to 3.67 s and 36.97 J of avoided fallback-reasoning overhead per trial. We open-source MemoGuard at https://github.com/hetheiin/memoguard.

13:00 JST研究/論文

Tencent UNI-REC チャレンジ向けのデュアルストリーム バイリニア フュージョンを備えたフィールド対応 RankMixer

このペーパーでは、KDD Cup 2026 Tencent UNIREC Challenge に対する当社のソリューションを紹介します。このタスクでは、ターゲット広告の pCVR 予測のために、マルチドメインのユーザー行動シーケンスと非順次マルチフィールド特徴の共同モデリングが必要です。私たちは、デュアルストリーム バイリニア フュージョンを備えたフィールド対応ランクミキサー (FA-RankMixer) を開発します。このモデルはまず、ターゲット認識 DIN モジュールを適用して、複数の行動ドメインからユーザーの関心を抽出します。また、最長の行動シーケンスについて、最近の関心と以前の関心を個別にモデル化します。次にモデルは、特徴フィールドと動作ドメインに基づいてセマンティック トークンを形成し、トークン間の相互作用に RankMixer ブロックを使用します。浅い MLP ストリームは深い RankMixer ストリームを補完し、グループごとの双線形モジュールがそれらの表現を融合します。私たちの最終的なソリューションは、公式リーダーボードで 9 位にランクされています。コードは https://github.com/PixelCookie-zyf/TAAC-2026-SeRankMixer で入手できます。

原文 (English)

Field-Aware RankMixer with Dual-Stream Bilinear Fusion for the Tencent UNI-REC Challenge

This paper presents our solution to the KDD Cup 2026 Tencent UNIREC Challenge. The task requires joint modeling of multi-domain user behavior sequences and non-sequential multi-field features for target-ad pCVR prediction. We develop a Field-Aware RankMixer (FA-RankMixer) with dual-stream bilinear fusion. The model first applies target-aware DIN modules to extract user interests from multiple behavior domains. It also models recent and earlier interests separately for the longest behavior sequence. The model then forms semantic tokens based on feature fields and behavior domains and uses RankMixer blocks for cross-token interaction. A shallow MLP stream complements the deep RankMixer stream, and a group-wise bilinear module fuses their representations. Our final solution ranks ninth on the official leaderboard. Our code is available at https://github.com/PixelCookie-zyf/TAAC-2026-SeRankMixer.

13:00 JSTLLM/生成AIエージェント

クラウド内のスケーラブルな LLM エージェント ツール アクセス

LLM エージェントは、外部システムで動作するためにツール呼び出しに依存することが増えており、モデル コンテキスト プロトコル (MCP) はすぐにその事実上のインターフェイスになりました。ただし、MCP をクラウド規模で運用することは困難になります。ツール プロバイダー側​​では、レガシー サービスを MCP 経由で直接呼び出すことはできません。急速なプロトコル開発により、継続的な互換性コストも発生します。エージェント側では、アクセス可能なツールの数は LLM コンテキスト ウィンドウと推論オーバーヘッドによって制限されます。大規模なツール セットをマウントすると、トークンの使用量と推論遅延が増加し、タスクの成功率が低下する可能性があります。さらに、複数のレプリカを持つステートフル MCP バックエンドの場合、セッション アフィニティを維持するとクライアント側の複雑さが増加します。 MCPサービス向けのクラウドスケールゲートウェイシステムを紹介します。データ プレーン上の直接接続モデルを破壊し、従来のサービス統合をオフロードし、互換性のない MCP バリアント、アクセス制御、ツールの推奨事項、およびセッション対応ルーティングをゲートウェイに統合します。ハイブリッド検索により、上位 15 位までの再現率が 98% 維持されます。高いツール選択精度でエージェント ツール アクセスを 3,000 以上に拡張し、コールごとのオーバーヘッドが低く、スケールアウト下でも安定して、ツール選択時間を $8.9\times$、トークン使用量を $23.8\times$ 削減します。最後に、実稼働環境でのゲートウェイ システムの導入から学んだ教訓を共有します。

原文 (English)

Scalable LLM Agent Tool Access in the Cloud

LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provider side, legacy services are not directly callable through MCP; the rapid protocol development also creates ongoing compatibility cost. On the agent side, the number of accessible tool is limited by the LLM context window and inference overhead; mounting a large tool set increases token usage and inference latency and can reduce task success rate. Moreover, for stateful MCP backends with multiple replicas, preserving session affinity increases client-side complexity. We present a cloud-scale gateway system for MCP service. It breaks the direct-connect model on the data plane and offloads legacy service integration, consolidating incompatible MCP variants, access control, tool recommendation, and session-aware routing to the gateway. Hybrid retrieval sustains 98% Top-15 recall; it scales agent tool access to 3,000+ with high tool selection accuracy, and reduces tool selection time by $8.9\times$ and token usage by $23.8\times$, with low per-call overhead, stable under scale-out. Finally, we share the lessons learned from deploying the gateway system in production.

13:00 JSTLLM/生成AIエージェント

効果的なマルチターン RL のためのプロセス報酬通知ツリー ロールアウト

強化学習 (RL) は、LLM エージェントをトレーニングするための重要なアプローチとなっていますが、GRPO/RLOO などの一般的な方法は、利点を推定するために独立してサンプリングされた複数の完全な軌跡に依存しています。長期的なエージェントタスクでは、このような一律のロールアウト戦略は、有益でない行き止まりの試みに予算を浪費する可能性があり、一方で有望な中間状態は十分な調査を受けられません。アクションと観察が交互に含まれるエージェント軌道のマルチターン構造は、各ターンが探索の決定点として機能するツリーとして軌道グループを組織することを自然にサポートします。この視点は、効果的な探索を、どこに分岐するかを決定する問題として再構成します。私たちは、マルチターン エージェント RL のための品質を意識したロールアウト フレームワークである Process-Scorer Guided Adaptive Tree Rollout (PATR) を提案します。 PATR は、タスクに適したプロセス フィードバックを使用して部分的な軌跡をスコア化し、有望な状態から選択的に分岐し、共有プレフィックスを再利用し、縮退パスを保守的に停止して無駄なサンプリングを削減します。結果として得られるロールアウト グループは、標準ポリシーの最適化との互換性を維持しながら、同じトレーニング予算の下でより効率的な探索を提供します。私たちは、FrozenLake と、これまでのツリー ロールアウト エージェント RL 手法ではほとんど調査されていない難度の高い SWE-Bench で PATR を評価します。実験では、PATR によってパフォーマンスが SWE-Bench で最大 +5.0 ポイント、FrozenLake で +9.3 ポイント向上することが示されており、スケーラブルなマルチターン RL の効果的な戦略としてプロセスガイド付きツリー ロールアウトが強調されています。

原文 (English)

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation. In long-horizon agentic tasks, such a uniform rollout strategy can waste budget on uninformative dead-end attempts, while promising intermediate states do not receive sufficient exploration. The multi-turn structure of agentic trajectories, with interleaved actions and observations, naturally supports organizing a trajectory group as a tree, where each turn serves as a decision point for exploration. This perspective reframes effective exploration as the problem of deciding where to branch. We propose Process-Scorer Guided Adaptive Tree Rollout (PATR), a quality-aware rollout framework for multi-turn agent RL. PATR uses task-appropriate process feedback to score partial trajectories, selectively branches from promising states, reuses shared prefixes, and conservatively stops degenerate paths to reduce wasted sampling. The resulting rollout groups remain compatible with standard policy optimization while providing more efficient exploration under the same training budget. We evaluate PATR on FrozenLake and the challenging SWE-Bench, which is largely unexplored by prior tree-rollout agent RL methods. Experiments show that PATR improves performance by up to +5.0 points on SWE-Bench and +9.3 points on FrozenLake, highlighting process-guided tree rollouts as an effective strategy for scalable multi-turn RL.

13:00 JSTロボティクス

AEGIS: オープンソースの液体ハンドリング ロボットのアッセイ対応プロトコル検証とランタイム監視

自動運転研究所では、Opentrons OT-2 などの低コストのリキッド ハンドラーへの依存がますます高まっています。Opentrons OT-2 は、Hamilton や Tecan システムのような圧力ベースの吸引モニタリングなしで出荷され、通常はオープン ループで動作します。 2 つの障害モードが検出されません。構文的には有効ですが、アッセイ固有の不変条件に違反するプロトコル (PCR テンプレートとテンプレートなしのコントロールの間でのチップの再利用など)、および実行時の物理的な実行障害 (部分的な分注、気泡、チップの欠落) です。両方の 2 層ガーディアンである AEGIS を紹介します。レイヤ 1 は、厳選された機械可読アッセイ ルール データベースと OT-2 Python コードを推論する LLM を組み合わせ、5 つのアッセイ ファミリにわたる 24 プロトコル ベンチマークで調整済み F1 0.97 に達し、5 つのバックエンドにわたるルールのみのアブレーションおよび LLM のみのアブレーションを上回ります。無料のオープンウェイト モデルは最高の独自モデルと結びついているため、有料 API は必要ありません。レイヤ 2 は、PCA ワールド モデルを YOLO で切り取られた 4 フレームのピペット軌道に適合させます。漏れのない 1 プレートアウトの評価では、平均精度 0.89 および動作点 F1 0.71 (AUROC 0.80) に達し、これはライブ デモンストレーションと一致する展開に忠実な数値であり、小型ピペット (p20) の分解能限界 (F1 0.47) を特徴づけます。物理的な OT-2 でのライブ デモンストレーション (条件ごとに 5 回の反復) は、チップなしの植え付け失敗を決定論的に捕捉し、カラー染料の部分塗布を常に VLM 自己投票ゲートにより部分塗布リコールを 5/5 に引き上げます。透明な水は正面視のみのモニターの原則的な制限であり、AEGISはこれを間違った判決ではなく、信頼性の低いVLM推論として表面化しています。カスケード トリアージでは、プレートあたりの VLM コストが 1.63 ドル近くに抑えられますが、常時 VLM ベースラインの場合は 10.33 ドルになります。 AEGIS はオープンソースであり、私たちの知る限りでは、オープンソースのリキッド ハンドラーの飛行前アッセイ対応検証とランタイム視覚モニタリングを統合した最初のシステムです。

原文 (English)

AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots

Self-driving laboratories increasingly rely on low-cost liquid handlers such as the Opentrons OT-2, which ship without the pressure-based aspiration monitoring of Hamilton or Tecan systems and are typically run open-loop. Two failure modes go undetected: protocols that are syntactically valid but violate assay-specific invariants (e.g., tip reuse between a PCR template and a no-template control), and physical execution failures (partial dispense, air bubbles, missing tips) at runtime. We present AEGIS, a two-layer guardian for both. Layer 1 pairs a curated machine-readable assay rule database with an LLM that reasons over OT-2 Python code, reaching an adjusted F1 of 0.97 on a 24-protocol benchmark across five assay families and beating rules-only and LLM-only ablations across five backends; a free open-weight model ties the best proprietary one, so no paid API is required. Layer 2 fits a PCA world model to YOLO-cropped four-frame pipette trajectories; under a leakage-free leave-one-plate-out evaluation it reaches average precision 0.89 and operating-point F1 0.71 (AUROC 0.80), a deployment-faithful number that matches the live demonstration, and we characterize the small-pipette (p20) resolution limit (F1 0.47). A live demonstration on a physical OT-2 (five replicates per condition) catches planted no-tip failures deterministically and partial dispense on coloured dyes, with an always-VLM self-vote gate lifting partial-dispense recall to 5/5; transparent water is a principled limit of any front-view-only monitor, which AEGIS surfaces as low-confidence VLM reasoning rather than a wrong verdict. Cascade triage holds VLM cost near $1.63 per plate versus $10.33 for an always-VLM baseline. AEGIS is open source and, to our knowledge, the first system to unify pre-flight assay-aware validation with runtime visual monitoring for an open-source liquid handler.

13:00 JSTロボティクス

5 Hz で考え、20 Hz で行動: 閉ループ運転のための非同期高速-低速視覚-言語-行動推論

大規模な言語モデルは、エンドツーエンドの運転に指示追従とシーン推論をもたらしますが、その推論の待ち時間が車両に必要な制御速度と衝突します。既存の閉ループ エージェントは、交互のシミュレーション ティックでモデルを呼び出し、その間に前のコマンドを再実行することでこのギャップを隠しているため、すべての制御出力の半分は最新の観測値を無視します。私たちは、この妥協を取り除く高速/低速アーキテクチャを提案します。凍結された 7B ビジョン言語バックボーンは低速システムとして機能し、ナビゲーション命令とビジュアル履歴を低頻度で消化しながら、レイヤーごとのキーと値のキャッシュをシーンの常駐表現として公開します。軽量アクション エキスパートは高速システムとして機能し、シミュレーション ティックごとにこのキャッシュと現在のカメラ フレームを処理して、単一の順方向パスでウェイポイントを後退させます。キャッシュは展開時に世界よりも遅れているため、ランダム化された最新状態でエキスパートをトレーニングし、トレーニングを非同期実行と調整します。 CARLA の LangAuto-Short ルートでは、システムは 50 ミリ秒のシミュレーション ティックごとに新しい制御を生成し、ルート完了を 37.0 から 94.0 までフレームスキップ ベースラインを超えて引き上げます。同じエキスパートによるフレームスキップアブレーションにより、作用する 2 つの要因が分離されます。エキスパートは独自に運転スコアを向上させ、ティックごとのフレッシュネスにより完走率が 82.1 から 94.0 に向上し、赤信号違反が 3 分の 1 減少します。単一の町でトレーニングされたエキスパートは、ゼロショットを 2 つのまだ見ていない町に転送し、ベースラインが 31 ~ 41% に達するルート完了率 84 ~ 94% を保持します。これにより、バックボーン独自のアクション ヘッドと比較して、オープン ループ ウェイポイント エラーがほぼ 4 分の 1 に削減され、単一コンシューマ GPU のヒストリの長さに関係なく、ティックあたりのモデル コストが 32 ミリ秒になります。

原文 (English)

Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving

Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing closed-loop agents hide this gap by invoking the model on alternate simulation ticks and replaying the previous command in between, so half of all control outputs ignore the newest observations. We present a fast-slow architecture that removes this compromise. A frozen 7B vision-language backbone acts as the slow system, digesting navigation instructions and visual history at low frequency while exposing its per-layer key-value cache as a standing representation of the scene. A lightweight action expert acts as the fast system, attending to this cache and to the current camera frame at every simulation tick to regress waypoints in a single forward pass. Since the cache lags behind the world at deployment, we train the expert under randomized staleness, aligning training with asynchronous execution. On LangAuto-Short routes in CARLA, our system produces fresh control at every 50 ms simulation tick and lifts route completion from 37.0 to 94.0 over the frame-skipping baseline. A frame-skip ablation with the same expert separates the two factors at work: the expert raises the driving score on its own, while per-tick freshness raises completion from 82.1 to 94.0 and cuts red-light violations by a third. Trained on a single town, the expert transfers zero-shot to two unseen towns, holding 84-94% route completion where the baseline reaches 31-41%. It reduces open-loop waypoint error by nearly a factor of four compared to the backbone's own action head, at a per-tick model cost of 32 ms that is independent of history length on a single consumer GPU.

13:00 JST研究/論文

トポス因果モデルの立方体形式化: 介入、束の接着、直観主義的な do-calculus

トポスの因果モデルは、トポスの内部で因果推論を再構成します。因果世界は前層であり、介入は部分オブジェクト分類器への特性マップであり、推論は直観的な内部言語で実行されます。以前に検証された確率モナドと do 計算に対して、この 1-トポス コアの最初の機械チェックされたアカウントを Cubical Agda で提供します。ふるいの分類器を構築し、その分類定理を備えた特性マップとして介入 $\mathrm{do}(X := x_0)$ を実現します。情報源は主張しているが決して証明していない、独立した機構の束の接着を証明する。そして内部言語の Kripke-Joyal 強制句を機械チェックします。モーダル層ではギャップを見つけて修復します。3 つの標準的な Lawvere-Tierney 公理は閉包演算子を強制しません。失われた法則が復元されたので、具体的な例として二重否定トポロジーを示し、介入とパールのルールがどのトポロジーでも安定していることを示します。そして、レジームのカバー全体にわたる反事実の伝達可能性は、カバー全体にわたる不変性として理解されるこの $j$ 安定性と一致します。さらに、プログラムが考慮しない現象、つまりペアごとに一貫性のあるローカル データがグローバル モデルを認めない、機械チェックされたコンテキスト性障害を追加します。開発では、Agda の --safe フラグの下で公理と型チェックが行われないことを前提とし、$\mathbb{Q}$ で具体的に出力される順序付きフィールドを使用します。スコープはプレシーフ (1-トポス) フラグメントで、タイプレベルのシーフ化と有向リフトは将来の作業に残されています。

原文 (English)

A cubical formalisation of topos causal models: intervention, sheaf gluing, and the intuitionistic do-calculus

Topos causal models recast causal inference inside a topos: a causal world is a presheaf, an intervention is a characteristic map into the subobject classifier, and reasoning is carried out in the intuitionistic internal language. We give the first machine-checked account of this 1-topos core, in Cubical Agda, over a previously verified probability monad and do-calculus. We build the classifier of sieves and realise the intervention $\mathrm{do}(X := x_0)$ as a characteristic map with its classification theorem; prove the sheaf gluing of independent mechanisms, which the source asserts but never proves; and machine-check the Kripke-Joyal forcing clauses of the internal language. In the modal layer we find and repair a gap: the three standard Lawvere-Tierney axioms do not force a closure operator. With the missing law restored, we exhibit the double-negation topology as a concrete instance and show that interventions and Pearl's rules are stable under every topology. Transportability of a counterfactual across a cover of regimes then coincides with this $j$-stability, understood as invariance across the cover. We further add a phenomenon the programme does not consider: a machine-checked contextuality obstruction, where pairwise-consistent local data admit no global model. The development assumes no axioms and typechecks under Agda's --safe flag, with the ordered field discharged concretely at $\mathbb{Q}$; the scope is the presheaf (1-topos) fragment, with type-level sheafification and the directed lift left to future work.

13:00 JSTロボティクス研究/論文

IMBench: 直感的なロボット操作のベンチマーク

人間は推論と運動制御を組み合わせて、さまざまな制約の下で複雑な操作タスクを解決します。彼らは、推論を行動に変換し、新しいシーン、タスク、ルールに迅速に適応するのに役立つ物理世界の理解を構築します。この機能を直感的な操作と呼びます。既存のベンチマークは、この統合を捉えることができません。それらは、物理的推論を実行から切り離して評価するか、明示的な推論を必要とせずにポリシーのパフォーマンスを測定します。 IMBENCH は、知覚、物理的推論、アクション生成、反復実行にわたる統合機能として直感的な操作を評価するように設計されたベンチマークです。私たちのタスクでは、モデルがタスクに関連した物理構造を推測し、接触が多い操作、ツールの使用、多段階の依存関係などの明示的な制約の下で実行可能なアクション シーケンスを生成する必要があります。 35 のタスクのベンチマーク、14K のフィルター処理された軌跡、および多様なシナリオを生成するためのスケーラブルなツールを紹介します。実験では、一貫したギャップが明らかになりました。視覚言語モデルは、部分的な物理的推論能力を示しますが、実行可能な計画を生み出すことができません。一方、最先端の視覚言語行動モデルは、タスクの制約を満たし、シナリオ全体で一般化するのに苦労しています。これらの結果は、直感的な操作が現在の基礎モデルとジェネラリストロボットポリシーに欠けている軸であることを特定し、IMBENCHをより統合された適応性のある身体知能を評価および実現するためのステップとして位置づけています。

原文 (English)

IMBench: A Benchmark for Intuitive Robotic Manipulation

Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of the physical world that helps them convert reasoning into actions and quickly adapt to new scenes, tasks, and rules. We refer to this capability as intuitive manipulation. Existing benchmarks fail to capture this integration: they evaluate physical reasoning in isolation from execution, or measure policy performance without requiring explicit reasoning. We introduce IMBENCH, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution. Our tasks require models to infer task-relevant physical structure and generate feasible action sequences under explicit constraints, including contact-rich manipulation, tool use, and multi-stage dependencies. We introduce a benchmark of 35 tasks, 14K filtered trajectories, and scalable tools for generating diverse scenarios. Experiments reveal a consistent gap: vision language models show partial physical reasoning ability but fail to produce executable plans, while state-of-the-art vision-language-action models struggle to satisfy task constraints and generalize across scenarios. These results identify intuitive manipulation as a missing axis in current foundation models and generalist robot policies, and position IMBENCH as a step toward evaluating and enabling more integrated, adaptive physical intelligence.

13:00 JSTLLM/生成AI

多者間の対話における演説の構造について: 離散的なラベルから連続的なレベルへ

対話システムと複数のユーザー間のマルチパーティ対話では、発話が誰に宛てられたものであるかを特定することが重要な課題です。従来の研究では、通常、宛先の検出は、個々の参加者またはグループを表す単一のラベルを選択するマルチクラス分類タスクとして扱われていました。この公式は、アドレスが本質的に離散的であり、主にターンテイキングを予測するために使用されてきたことを前提としています。この論文では、アドレスを連続現象として分析することによって、この仮定を再検討します。複数のアノテーターによってアノテーションが付けられたマルチパーティの人間対話コーパスを使用して、多数決の宛先ラベルから導出されたバイナリ住所ラベルと、潜在変数モデルを使用したアノテーターの判断から推測された連続住所レベルの両方を構築します。次に、これらの表現が順番交代や、視線や相槌などのリスナーの行動にどのように関連しているかを調べます。私たちの結果は、ターンテイクに加えて、視線と相槌の両方がアドレスに関連していることを示しています。さらに、連続的な住所レベルを使用したモデルは、離散的なラベルを使用したモデルよりも優れた予測適合を達成し、住所が段階的な構造を示す可能性があることを示唆しています。最後に、本研究の結果を踏まえて、宛先検出研究の今後の方向性について議論する。

原文 (English)

On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous Levels

In multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge. Prior work has typically treated addressee detection as a multi-class classification task, selecting a single label representing an individual participant or the group. This formulation assumes that address is inherently discrete and has primarily been used for predicting turn-taking. In this paper, we revisit this assumption by analyzing address as a continuous phenomenon. Using a multi-party human dialogue corpus annotated by multiple annotators, we construct both binary address labels derived from majority-vote addressee labels and continuous address levels inferred from annotator judgments using a latent-variable model. We then examine how these representations relate to turn-taking as well as listener behaviors, including gaze and backchannels. Our results show that, in addition to turn-taking, both gaze and backchannels are associated with address. Furthermore, models using continuous address levels achieve better predictive fit than those using discrete labels, suggesting that address may exhibit graded structure. Finally, we discuss the future directions of addressee detection research based on the findings of this study.

13:00 JST研究/論文

視覚野における推論と拡散モデルのメカニズムの理解に向けて

我々は、その機能がパラメータから容易に理解できる最小拡散モデルと同等の一次視覚野(V1)における知覚推論のモデルについて説明します。このモデルは、制約のないペアごとの相互作用行列の形式で潜在変数に対する非階乗事前変数を使用したスパースコーディングに基づいており、標準的なスパースコーディング推論を一般的なリカレント動的システムに拡張します。私たちは、ノイズ除去スコアマッチング目標と暗黙的な微分を使用して、これらの反復的なダイナミクスを効率的にトレーニングします。自然画像でトレーニングした後、学習された相互作用行列は、同様の方向調整のニューロンをリンクする V1 の表層の水平接続の構造を反映します。このモデルは非常に優れたノイズ除去パフォーマンスを示し、極度の視覚的曖昧さの中で拡張された輪郭などの画像特徴を復元し、一般化領域における標準のブラックボックス拡散アーキテクチャの動作とほぼ一致します。モデルの単純さのおかげで、ネットワークのヤコビアンは潜在変数間の相互作用行列の観点から直接分解でき、リカレント ダイナミクスが自然な構造変形の連続ファミリーに対してどのように高い確率を割り当てるかを機構的に明らかにします。興味深いことに、この回路内では、潜在変数の大部分が視覚入力から完全に切り離されることを学習し、本質的に画像特徴間にグローバルな一貫性を強制するように見える階層表現を形成します。モデルと結果は共に、2 つの異なる領域の橋渡しとなります。神経科学の場合、知覚推論タスク中の反復型神経回路の機能的接続に関する具体的でテスト可能な仮説を生成します。機械学習の場合、拡散モデルによって学習された内部メカニズムが解明され、有限のトレーニング セットから無限に多くの新しい画像を生成できるようになります。

原文 (English)

Toward a mechanistic understanding of inference in visual cortex and diffusion models

We describe a model of perceptual inference in primary visual cortex (V1) equivalent to a minimal diffusion model whose function can be readily understood from its parameters. The model is based on sparse coding with a non-factorial prior over latent variables in the form of an unconstrained, pairwise interaction matrix, extending standard sparse coding inference to a general recurrent dynamical system. We efficiently train these recurrent dynamics using a denoising score-matching objective and implicit differentiation. After training on natural images, the learned interaction matrix mirrors the structure of horizontal connections in superficial layers of V1 that link neurons of similar orientation tuning. This model exhibits exceptionally good denoising performance, restoring image features such as extended contours amid extreme visual ambiguity, nearly matching the behavior of standard, black-box diffusion architectures in generalization regime. Owing to the model's simplicity, the network's Jacobian can be decomposed directly in terms of the interaction matrix between latent variables, revealing mechanistically how the recurrent dynamics assign high probability over a continuous family of natural structural deformations. Intriguingly, within this circuit, a large fraction of latent variables learn to disconnect from visual input altogether, essentially forming a hierarchical representation that appears to enforce global consistency among image features. Together, the model and results bridge two distinct domains: for neuroscience, it generates concrete, testable hypotheses regarding functional connectivity in recurrent neural circuits during perceptual inference tasks; for machine learning, it elucidates the internal mechanisms learned by diffusion models that allow them to generate infinitely many novel images from a finite training set.

13:00 JST画像/動画生成

拡散ベースの実世界画像超解像度のための効率的な難易度を考慮した動的ルーティング

拡散ベースの手法は、事前にトレーニングされた大規模な安定拡散 (SD) モデルを強力な生成事前分布として活用することにより、現実世界の画像超解像度 (Real-ISR) で優れたパフォーマンスを達成しました。ただし、これらの方法には依然として 2 つの重要な制限があります。まず、既存の SD ベースの 1 ステップおよびマルチステップ Real-ISR アプローチは、すべての入力サンプルに対して統一された処理パラダイムを採用しており、画像間で異なる復元難易度を無視しています。第二に、SD モデルにおける VAE の積極的な解像度低下 (8 倍ダウンサンプリングなど) は、微細スケールの詳細の不可逆的な損失につながり、その後の拡散プロセスでは回復できません。これらの制限に対処するために、私たちは、厳格で画一的な処理パラダイムを克服する、難易度を考慮したダイナミック ルーティング (DDR) 戦略を提案します。具体的には、まず各入力画像の復元コストを予測する難易度推定器を設計し、適切な容量のネットワークへの自動割り当てを可能にします。次に、SD バックボーン内の VAE の空間ダウンサンプリング比を調整することで、さまざまなモデル容量を持つ Real-ISR ネットワークのセットを構築します。これにより、より単純な入力の効率を維持しながら、困難なケースに対してより多くの高周波情報を保存します。広範な実験により、最近の最先端の方法と比較して、提案されたモデルの優れた効率と有効性が実証されました。

原文 (English)

Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

Diffusion-based methods have achieved impressive performance in real-world image super-resolution (Real-ISR) by leveraging large pre-trained stable diffusion (SD) models as powerful generative priors. However, these methods still face two key limitations. First, existing SD-based one-step and multi-step Real-ISR approaches adopt a unified processing paradigm for all input samples, ignoring the varying restoration difficulty across images. Second, the aggressive resolution reduction of the VAE in SD models (e.g., 8x downsampling) leads to irreversible loss of fine-scale details, which cannot be recovered by the subsequent diffusion process. To address these limitations, we propose a Difficulty-aware Dynamic Routing (DDR) strategy that overcomes the rigid, one-size-fits-all processing paradigm. Specifically, we first design a difficulty estimator to predict the restoration cost of each input image, enabling automatic assignment to a network of appropriate capacity. Then, we construct a set of Real-ISR networks with varying model capacities by modulating the spatial downsampling ratio of the VAE in the SD backbone, thereby preserving more high-frequency information for challenging cases while maintaining efficiency for simpler inputs. Extensive experiments have demonstrated the superior efficiency and effectiveness of the proposed model compared to recent state-of-the-art methods.

13:00 JSTLLM/生成AIエージェント

プロンプトとしての地図: クロスシナリオ無線位置特定のためのマルチモーダル空間信号基盤モデルの学習

正確かつ堅牢な無線位置特定は、自動運転、拡張現実、スマート製造などの新興 5G/6G アプリケーションを実現する重要な要素です。その重要性にもかかわらず、ワイヤレス信号の複雑な性質と環境変化に対する信号の敏感さのため、多様な環境にわたって正確な位置特定を達成することは依然として困難です。既存のデータ駆動型アプローチは、一般化機能が限られていることが多く、大量のラベル付きデータが必要であり、新しいシナリオに適応するのに苦労しています。これらの制限に対処するために、我々は、2 つの主要な革新を導入するマルチモーダル基盤モデルである SigMap を提案します。(1) チャネルの周期特性に基づいてマスキング パターンを動的に調整し、堅牢な無線表現を学習するサイクル適応マスキング戦略。 (2) 軽量のソフト プロンプトを介して 3D 地理情報を統合し、効果的なクロスシナリオ適応を実現する新しい「マップ アズ プロンプト」フレームワーク。広範な実験により、私たちのモデルは、目に見えない環境で強力なゼロショット汎化を示しながら、複数の位置特定タスクにわたって最先端のパフォーマンスを達成し、教師付きベースラインと自己教師付きベースラインの両方を大幅に上回っていることが実証されました。

原文 (English)

Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization

Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environments remains challenging due to the complex nature of wireless signals and their sensitivity to environmental changes. Existing data-driven approaches often suffer from limited generalization capability, requiring extensive labeled data and struggling to adapt to new scenarios. To address these limitations, we propose SigMap, a multimodal foundation model that introduces two key innovations: (1) A cycle-adaptive masking strategy that dynamically adjusts masking patterns based on channel periodicity characteristics to learn robust wireless representations; (2) A novel "map-as-prompt" framework that integrates 3D geographic information through lightweight soft prompts for effective cross-scenario adaptation. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple localization tasks while exhibiting strong zero-shot generalization in unseen environments, significantly outperforming both supervised and self-supervised baselines by considerable margins.

13:00 JST画像/動画生成ビジネス/資金調達

暗黙的な文化的整合性報酬モデリングによるテキストから画像への評価のバイアスを軽減する

Text-to-Image (T2I) システムが急速に進歩するにつれ、公正で信頼できる生成 AI にとって、合成コンテンツの文化的信頼性を評価することがますます重要になってきています。既存の T2I 評価指標とマルチモーダルなジャッジは、暗黙的な文化規範を過小評価する視覚的意味論的表現に依存していることが多く、偏った好みの判断やきめ細かな文化的手がかりの省略につながります。さらに、ビジュアル質問応答 (VQA) ベースの評価器は通常、自己回帰テキスト生成に依存するため、リアルタイム報酬モデリングのスケーラビリティが制限されます。これらの制限に対処するために、軽量の 42 億パラメータのマルチモーダル大規模言語モデル (MLLM) に基づいて構築された暗黙的文化的調整報酬モデルを導入します。私たちのフレームワークは、暗黙的文化プローブをスキップ接続クロスアテンション (SkipCA) メカニズムと統合し、後期段階のセマンティック機能が初期段階の視覚表現に直接対応し、文化的に顕著な詳細をより適切に保存できるようにします。 CultureFrames ベンチマークからの 3,323 個の難しくて厳選された画像ペアの評価では、私たちのアプローチがペアごとの精度 80.54% を達成し、ピアソン相関係数とケンダル相関係数がそれぞれ 0.546 と 0.377 で、代表的な視覚言語メトリックや MLLM ベースの評価者を上回るパフォーマンスを示しています。さらに、自己回帰テキスト生成をバイパスすることにより、モデルはローカル推論セットアップの下で各評価を 0.21 秒で処理し、標準の VQA ベースの評価器と比べて 10 倍の高速化を達成します。これらの結果は、提案された報酬モデルが、ヒューマン フィードバックからの強化学習や直接嗜好最適化などの嗜好最適化パイプラインに対して、効率的で文化を意識したスカラー信号を提供できることを示唆しています。

原文 (English)

Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling

As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI. Existing T2I evaluation metrics and multimodal judges often rely on visual-semantic representations that underrepresent implicit cultural norms, leading to biased preference judgments and the omission of fine-grained cultural cues. In addition, visual question answering (VQA)-based evaluators typically depend on autoregressive text generation, which limits their scalability for real-time reward modeling. To address these limitations, we introduce an Implicit Cultural Alignment Reward Model built upon a lightweight 4.2-billion-parameter Multimodal Large Language Model (MLLM). Our framework integrates an Implicit Cultural Probe with a Skip-connection Cross-Attention (SkipCA) mechanism, enabling late-stage semantic features to directly attend to early-stage visual representations and better preserve culturally salient details. Evaluations on 3,323 challenging and carefully curated image pairs from the CulturalFrames benchmark show that our approach achieves 80.54% pairwise accuracy, with Pearson and Kendall correlation coefficients of 0.546 and 0.377, respectively, outperforming representative vision-language metrics and MLLM-based evaluators. Moreover, by bypassing autoregressive text generation, our model processes each evaluation in 0.21 seconds under our local inference setup, achieving a $10\times$ speedup over standard VQA-based evaluators. These results suggest that the proposed reward model can provide an efficient and culturally aware scalar signal for preference optimization pipelines such as Reinforcement Learning from Human Feedback and Direct Preference Optimization.

13:00 JST研究/論文

AuEmoChat: 会話音声合成のための本物の感情の理解とレンダリング

会話型音声合成 (CSS) は、ユーザーとエージェントの対話において、人間のような感情表現と文脈の一貫性を備えた音声を合成することを目的としています。既存の CSS 手法は、事前に定義された感情ラベル スペース (7 つの感情カテゴリなど) が限られているため、本物の人間の感情を表現するのに苦労していますが、マルチターン対話履歴内の冗長なマルチモーダル トークンがコンテキストの理解を妨げます。これらの問題に対処するために、私たちは本物の感情の理解とレンダリングのための CSS フレームワークである AuEmoChat を提案します。まず、有限スカラー量子化を介して大規模な感情音声から離散的な本物の感情トークン空間を学習する AuEmoCodec を開発し、限られた基本的な感情カテゴリよりもより本物の感情表現を可能にします。さらに、感情に関連したコンテキストを維持しながら、マルチモーダルな対話履歴内の冗長トークンをマージする、本物の感情に基づくトークンマージアルゴリズムである AuEmoToMe を提案します。これを自己回帰テキスト音声モデルに統合して、ターゲットとなる本物の感情トークンと音声トークンを予測します。最後に、統合された対話コンテキスト、ターゲットの本物の感情、および音響事前分布を共同で条件付けすることによって音声をレンダリングする、本物の感情フロー マッチングを提案します。 NCSSD-EmCap データセットに関する広範な実験により、AuEmoChat が最先端の CSS ベースラインを上回り、より表現力豊かで本物の感情的なスピーチを生成することが実証されました。

原文 (English)

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions due to limited predefined emotion label spaces (e.g., seven emotion categories), while redundant multimodal tokens in multi-turn dialogue history interfere with context understanding. To address these issues, we propose AuEmoChat, a CSS framework for authentic emotion understanding and rendering. First, we develop AuEmoCodec, which learns a discrete authentic emotion token space from large-scale emotional speech via finite scalar quantization, enabling a more authentic emotion representation than limited basic emotion categories. We further propose AuEmoToMe, an authentic-emotion-guided token merging algorithm that merges redundant tokens in multimodal dialogue history while preserving emotion-relevant context. We integrate it into an autoregressive text-speech model to predict the target authentic emotion token and speech tokens. Finally, we propose Authentic Emotion Flow Matching, which renders speech by jointly conditioning on merged dialogue context, target authentic emotion, and acoustic priors. Extensive experiments on the NCSSD-EmCap dataset demonstrate that AuEmoChat outperforms state-of-the-art CSS baselines and generates more expressive and authentic emotional speech.

13:00 JST画像/動画生成研究/論文

GeoChrono: リモート センシングにおける長期的な時間的理解のベンチマークと再考

リモート センシングは、地球の長期的な表面進化を観察するための比類のない有利な点を提供しますが、モデルには、孤立した瞬間に土地被覆を認識するだけでなく、変化を追跡し、進化の歴史を記憶し、時間と空間を超えて推論する必要があります。しかし、既存の研究には、これらの異なる能力を詳細に分析する体系的な評価が欠けています。このギャップを埋めるために、このタスクを 4 つの段階的な認知レベル (すなわち、土地被覆知覚、時間的認識、長期記憶、および時空間推論) に分解する多次元ベンチマークである ChronoBench を導入します。 ChronoBench は 12 のサブタスクと、厳密に検証された 17,689 の QA (質問と回答) のペアで構成されています。広範な評価により、主流のMLLMは人間の専門家に大幅に遅れをとり、長期記憶が最も重大なボトルネックとして浮上していることが明らかになりました。この発見に動機付けられて、私たちはさらに、長期的な地理的進化についての追跡、記憶、および推論のための強化された機能を備えた MLLM である GeoChrono を提案します。セマンティクスが進化しても地理的区画は空間的に固定されたままであるという物理的事前条件を利用して、専用の土地被覆進化モデリング用に位置ごとの時間軌跡を構築する時間軌跡エンコーダ(TempEnc)を設計し、静的な背景を圧縮しながら動的領域を適応的に保存する粗密トークン圧縮器(C2FComp)を導入します。トレーニングをサポートするために、トレーニングのすべてのコンピテンシー レベルにまたがる 104K サンプルの命令調整データセットである ChronoInstruct も構築します。 GeoChrono は、ChronoBench で最先端のパフォーマンスを達成し、主要な商用 MLLM を 20% 以上上回っています。一方、C2FComp は、GeoChrono の 94.6% のパフォーマンスを維持しながらビジュアル トークンを 56% 以上削減します。コードとデータは https://github.com/IntelliSensing/GeoChrono で入手できます。

原文 (English)

GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space. However, existing studies lack a systematic evaluation that dissects these distinct competencies. To fill this gap, we introduce ChronoBench, a multidimensional benchmark that decomposes this task into four progressive cognitive levels (i.e., Land Cover Perception, Temporal Recognition, Long-Term Memory, and Spatio-Temporal Reasoning). The ChronoBench comprises 12 sub-tasks and 17,689 rigorously validated QA (Question-Answer) pairs. Extensive evaluations reveal that mainstream MLLMs fall drastically behind human experts, with Long-Term Memory emerging as the most critical bottleneck. Motivated by this finding, we further propose GeoChrono, an MLLM with enhanced capabilities for tracing, memorizing, and reasoning about long-term geographic evolution. Leveraging the physical prior that geographic parcels remain spatially fixed while their semantics evolve, we design a Temporal Trajectory Encoder~(TempEnc) that constructs per-location temporal trajectories for dedicated land cover evolution modeling, and we introduce a Coarse-to-Fine Token Compressor~(C2FComp) that adaptively preserves dynamic regions while compressing the static background. To support training, we also construct ChronoInstruct, a 104K-sample instruction-tuning dataset spanning all competency levels for training. GeoChrono achieves state-of-the-art performance on ChronoBench, surpassing the leading commercial MLLMs by over 20%, while C2FComp reduces visual tokens by over 56% while retaining GeoChrono's 94.6% performance. The code and data will be available at https://github.com/IntelliSensing/GeoChrono

13:00 JST研究/論文

XAI 主導のデータ削減による時系列分類のスケーリング

時系列の Explainable AI (XAI) はアルゴリズム的に大幅な成長を遂げていますが、下流のタスクに測定可能なパフォーマンスの向上をもたらすというその有用性は依然として十分に検討されていません。このホワイトペーパーでは、時系列分類 (TSC) における効果的なデータ削減のために XAI アトリビューション手法を再利用する新しい方法論である drXAI を紹介することで、このギャップを埋めます。最新の TSC における中心的な課題はスケーラビリティです。 Transformers などの最先端のモデルは、シーケンスの長さに対して 2 次の複雑性を示し、チャネル数に対して 1 次の複雑さを示します。これにより、大規模なデータセットの計算が法外に難しくなります。 drXAI は、高速な GPU アクセラレーション分類器 (Hydra) を使用してローカル アトリビューションを生成することで、この問題に対処します。これらをグローバルな特徴重要度スコアに集約し、自動化されたエルボーカット ヒューリスティックを採用して、手動のしきい値を必要とせずに最も顕著な特徴を選択します。私たちは、合成データセットと現実世界の一変量データセットおよび多変量データセットの両方でアプローチを評価します。合成ベンチマークでは、drXAI は、従来のベースラインが失敗するグラウンドトゥルース機能を正常に回復します。実世界のデータでは、drXAI は、完全なデータセットでトレーニングされたモデルと同等の分類精度を維持しながら、80% ~ 90% のデータ削減を達成します。最も重要なことは、drXAI を使用すると、ConvTran のようなリソースを大量に消費するモデルを、メモリの制約により以前はアクセスできなかったデータセットに拡張できることを示しています。私たちの結果は、XAI を解釈しやすさだけでなく、時系列分析における特徴選択とスケーラビリティのための堅牢なツールとして使用する利点を示しています。すべてのコードとデータは公開されています。

原文 (English)

Scaling Time Series Classification via XAI-Driven Data Reduction

Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodology that repurposes XAI attribution methods for effective data reduction in Time Series Classification (TSC). The core challenge in modern TSC is scalability; state-of-the-art models, such as Transformers, exhibit quadratic complexity relative to sequence length and linear complexity relative to the number of channels. This renders them computationally prohibitive for massive datasets. drXAI addresses this by using a fast, GPU-accelerated classifier (Hydra) to generate local attributions. We aggregate these into global feature importance scores and employ an automated elbow-cut heuristic to select the most salient features without requiring manual thresholds. We evaluate our approach on both synthetic and real-world univariate and multivariate datasets. On synthetic benchmarks, drXAI successfully recovers ground-truth features where traditional baselines fail. On real-world data, drXAI achieves between 80% and 90% data reduction while maintaining classification accuracy comparable to models trained on the full dataset. Most importantly, we show that drXAI allows resource-intensive models like ConvTran to scale to datasets that were previously inaccessible due to memory constraints. Our results show the benefits of using XAI not just for interpretability, but as a robust tool for feature selection and scalability in time series analysis. All our code and data are openly available.

13:00 JST研究/論文

AquaAugmentor: 水の飲料適性を予測するための新しい特徴拡張アルゴリズム

飲料水へのアクセスは、健康、経済発展、持続可能性にとって非常に重要です。しかし、水源データの複雑さと変動性により、水質を正確に分類することは依然として大きな課題です。この論文では、機械学習と深層学習アルゴリズムを通じて水の飲料適性を予測するという課題に取り組みます。新しい特徴拡張アルゴリズムである AquaAugmentor が導入され、低次元データセットに対するこれらのモデルの予測パフォーマンスが向上します。 pH、硬度、固形分、クロラミン、硫酸塩など、水の化学的属性を含むデータセットを利用します。この研究では、AquaAugmentor を使用した場合と使用しない場合のモデルのパフォーマンスを評価します。各モデルは水を飲料水か非飲料水として分類するために適用され、そのパフォーマンスはテスト精度と AUC スコアに基づいて評価および比較されます。この結果は、私たちが提案したアルゴリズムの長所と限界を浮き彫りにし、水質分類の予測パフォーマンスを向上させる最も効果的な手法についての洞察を提供します。この研究は、安全な水へのアクセスを確保するための広範な取り組みに貢献し、環境品質評価に機械学習を採用するためのフレームワークとして機能します。この調査結果は、研究者、政策立案者、公衆衛生当局が信頼できる機械学習予測に基づいて情報に基づいた意思決定を行えるよう支援することを目的としています。

原文 (English)

AquaAugmentor: A Novel Feature Augmentation Algorithm for Water Potability Prediction

Access to potable water is crucial for health, economic development, and sustainability. However, accurately classifying water quality remains a significant challenge due to the complexity and variability of water source data. This paper addresses the challenge of predicting water potability through machine learning and deep learning algorithms. It introduces a novel feature augmentation algorithm, AquaAugmentor, to enhance the predictive performance of these models for low-dimensional datasets. Utilizing a dataset that includes chemical attributes of water, such as pH, hardness, solids, chloramines, sulfate, and others. This study evaluates the performance of the models with and without AquaAugmentor. Each model applied to classify water as potable or non-potable and its performance is then evaluated and compared based on test accuracy and AUC score. The results highlight the strengths and limitations of our proposed algorithm, providing insights into the most effective techniques for improving the predictive performance of water quality classification. This study contributes to the broader efforts of ensuring safe water access and serves as a framework for employing machine learning in environmental quality assessments. The findings aim to assist researchers, policymakers, and public health officials in making informed decisions based on reliable machine learning predictions.

13:00 JSTLLM/生成AI画像/動画生成

マルチイベントの長時間ビデオを理解するためのモジュール化された動的粒度ビデオ LLM

ビデオ大規模言語モデル (ビデオ LLM) は、さまざまなビデオ理解タスクにおいて大幅な進歩をもたらしました。ただし、限られたビジュアル トークンの予算と複数の重要なイベントをキャプチャする必要性との間の緊張により、長いビデオのシナリオは依然として困難です。既存のアプローチは通常、長いビデオを 2 つの段階、つまり i) キーフレームの選択と ii) 詳細な認識の実行で処理しますが、これには限界があります。適応的な容量割り当てと自己修正のためのモジュール式メカニズムが欠如しており、その結果、モデリングの信頼性が低くなります。これらの課題に取り組むために、私たちは、マルチイベントの長時間ビデオを理解するための新しいモジュール化された動的粒度ビデオ LLM フレームワークである MoD-VLLM を提案します。これは、時間的グラウンディングと意味論的理解を反復的かつ内省的に統合します。具体的には、Positive-Negative Video Segments Grounding モジュールとモジュール化された Dynamic-Granularity Reflection モジュールを提案します。これらは閉ループを形成し、質問関連のビデオ セグメントを段階的にローカライズします。グラウンディング モジュールは、ビデオ LLM に対し、ビデオの質問に基づいて、関連するビデオ セグメントと無関係なビデオ セグメントを区別するように指示します。リフレクション モジュールはモジュール化されたスケジューラを採用しており、詳細な認識をキャプチャするために関連するポジティブ セグメントに対してはきめの細かいエンコーディングを動的に選択し、グローバル コンテキストを維持するためにネガティブ セグメントに対しては粗いエンコーディングを選択します。さらに、MoD-VLLM が最適なグラウンディング ポリシーと動的粒度の視覚表現を共同で学習できるようにする、動的粒度の強化学習戦略を提案します。さらに、複雑な長時間ビデオ推論のための挑戦的なマルチイベント長時間ビデオ ベンチマークである MEventBench を提案します。いくつかの長いビデオ理解ベンチマークと MEventBench に関する広範な実験により、MoD-VLLM が最先端のベースラインを大幅に上回ることが実証されました。

原文 (English)

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

Video Large Language Models (Video LLMs) have made significant advancements in various video understanding tasks. However, long-video scenarios remain challenging due to the tension between limited visual token budgets and the need to capture multiple key events. Existing approaches typically process long videos in two stages, i.e., i) select keyframes and ii) perform detailed perception, which exhibit limitations: they lack a modular mechanism for adaptive capacity allocation and self-correction, resulting in unreliable modeling. To tackle these challenges, we propose MoD-VLLM, a novel Modularized Dynamic-Granularity Video LLM framework for multi-event long video understanding, which unifies temporal grounding and semantic understanding iteratively and self-reflectively. Specifically, we propose a Positive-Negative Video Segments Grounding module and a Modularized Dynamic-Granularity Reflection module, which form a closed loop to progressively localize the question-related video segments. The grounding module instructs a Video LLM to distinguish relevant from irrelevant video segments based on the video question. The reflection module employs a modularized scheduler that dynamically selects fine-grained encoding for relevant positive segments to capture detailed perception and coarse-grained encoding for negative segments to maintain global context. We further propose a dynamic-granularity reinforcement learning strategy, allowing MoD-VLLM to learn optimal grounding policies and dynamic granularity visual representation jointly. Moreover, we propose MEventBench, a challenging Multi-Event Long Video Benchmark for complex long video reasoning. Extensive experiments on several long video understanding benchmarks and our MEventBench demonstrate that MoD-VLLM significantly outperforms state-of-the-art baselines.

13:00 JST画像/動画生成

イベントベースのマルチモーダル感情推定における学習された表現の幾何学について

ELOPE チャレンジの上位成績チームが採用したものを含む、イベントベースのエゴモーション推定への古典的なアプローチは、コントラスト最大化、ホモグラフィー推定、または解析的モーション反転と組み合わせた高密度オプティカル フローなどの幾何学的最適化フレームワークに依存しています。この研究では、エゴモーション推定のためのマルチモーダル ネットワーク内に現れる幾何学的構造を調査します。イベント テンソル、慣性測定、距離信号は、クロスモーダル アテンション アーキテクチャを通じて融合され、バッチ設定でトレーニングされます。潜在空間幾何学と注意力学を分析し、(i) 埋め込みが運動変数と整列した低次元多様体上に存在すること、(ii) 注意の重みが角励起と視覚的信頼性に適応すること、および (iii) 融合された表現が古典的な可観測性の手がかりを回復することを示します。これらの結果は、分析的推定理論と最新のデータ駆動型の融合の橋渡しとなります。

原文 (English)

On the Geometry of Learned Representations in Event-Based Multi-Modal Egomotion Estimation

Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely on geometric optimization frameworks such as contrast maximization, homography estimation, or dense optical flow combined with analytic motion inversion. This work investigates the geometric structure that emerges inside a multi-modal network for egomotion estimation. Event tensors, inertial measurements, and range signals are fused through a cross-modal attention architecture and trained in a batch setting. We analyze the latent space geometry and attention dynamics, showing that (i) embeddings lie on low-dimensional manifolds aligned with motion variables, (ii) attention weights adapt with angular excitation and visual reliability, and (iii) the fused representation recovers classical observability cues. These results bridge analytical estimation theory and modern data-driven fusion.

13:00 JST研究/論文

多段階の工業プロセスにおける多変量時系列異常検出のための知識支援型マルチグラフ依存関係学習

産業プロセスでは、多くの場合、複数のステージにわたる複数のセンサーから複雑で相互依存する時系列データが生成され、変数とプロセス ステージ間に複雑な依存関係が形成されます。多変量時系列異常検出 (MTAD) によるこれらの時系列の効果的な監視とタイムリーな異常検出は、障害を防止し、自動化システムの信頼性を確保するために重要です。グラフ ニューラル ネットワーク (GNN) は、データ駆動型グラフを活用して変数間の複雑な依存関係をモデル化し、多変量時系列内の関係構造を効果的にキャプチャして異常検出パフォーマンスを向上させることで、高度な MTAD を実現しました。ただし、既存の GNN ベースのアプローチでは、重要なプロセスの知識が見落とされることが多く、この知識を考慮したとしても、それを既存のモデルにシームレスに組み込むことは本質的に困難なままであり、次善のパフォーマンスにつながります。この制限に対処するために、MTAD の多段階工業プロセスにおけるセンサーの依存関係をモデル化するための知識支援マルチグラフ フレームワークを提案します。これは、プロセスの知識をグラフ学習に明示的に組み込んで、依存関係のモデリングを強化し、異常検出パフォーマンスを向上させます。私たちの方法は 3 つの相補的なグラフを構築します。1 つは純粋にデータ駆動型で、2 つはプロセスの知識から得られた構造的制約を統合することによって洗練されました。これらのグラフを異常検出に効果的に活用するために、マルチグラフ アテンション ネットワークを採用し、複雑な依存関係をより正確かつ堅牢に表現できるようにしています。 2 つの現実世界の多段階産業データセットに対する包括的な実験により、プロセス知識を組み込むことで異常検出パフォーマンスが大幅に向上することが実証されました。

原文 (English)

Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly Detection in Multi-Stage Industrial Processes

Industrial processes often generate complex, interdependent time-series data from multiple sensors across multiple stages, forming complex dependencies among variables and process stages. Effective monitoring and timely anomaly detection of these time series through multivariate time series anomaly detection (MTAD) is crucial for preventing failures and ensuring the reliability of automated systems. Graph neural networks (GNNs) have advanced MTAD by leveraging data-driven graphs to model complex dependencies among variables, effectively capturing relational structures within multivariate time series to enhance anomaly detection performance. However, existing GNN-based approaches often overlook critical process knowledge, and even when this knowledge is considered, seamlessly incorporating it into existing models remains inherently challenging, leading to suboptimal performance. To address this limitation, we propose a knowledge-assisted multi-graph framework for modeling sensor dependencies in multi-stage industrial processes for MTAD, which explicitly incorporates process knowledge into graph learning to enhance dependency modeling and improve anomaly detection performance. Our method constructs three complementary graphs: one purely data-driven and two refined by integrating structural constraints derived from process knowledge. To effectively leverage these graphs for anomaly detection, we employ a multi-graph attention network, enabling a more accurate and robust representation of complex dependencies. Comprehensive experiments on two real-world, multi-stage industrial datasets demonstrate that incorporating process knowledge substantially enhances anomaly detection performance.

13:00 JST研究/論文

線形自己注意を備えたトランスフォーマーを使用した、単純な線形回帰タスクに対する閉じた形式の解法のインコンテキスト学習

コンテキスト内学習はトランスフォーマーの注目すべき特性であり、最近多くの関心を集めています。インコンテキスト学習の多くの研究では、トランスフォーマーは線形および非線形回帰問題のソルバーを実装できることが示されており、トランスフォーマーのほとんどは勾配降下アルゴリズムを実装しています。ただし、これらの実装が実際にトレーニングを通じて習得されたかどうかはまだ不明です。この論文では、単純な回帰タスクで最小二乗推定をコンテキスト内で学習する、線形自己注意を備えたトランスフォーマーを構築します。ここでのポイントは、閉形式 (解析) 解が、勾配降下法アルゴリズムに基づく近似解ではなく、層正規化を使用して近似的に得られるということです。次に、実験例を示します。この例では、ターゲット出力が最小二乗推定値である場合に、l1 正則化でトレーニングされた変換器で実装が主に使用されます。

原文 (English)

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention

In-context learning is a remarkable property of transformers and has recently received a lot of interest. In many studies of in-context learning, it has been shown that transformers are capable of implementing solver for linear and non-linear regression problems, in which the most of them implement gradient descent algorithm. However, it is still unclear whether those implementations have actually been acquired through training. In this paper, we construct a transformer with linear self-attention, which in-context learns the least squares estimate in a simple regression task. The point here is that the closed form (analytical) solution is approximately obtained by using layer normalization rather than an approximate solution based on gradient descent algorithm. Then, we show an experimental example, in which our implementation is mainly used in the transformer trained with l1 regularization when the target output is the least squares estimate.

13:00 JST研究/論文

RTL-Sequencer: シーケンスベースのパラダイムによるスケーラブルな RTL タイミング予測に向けて

レジスタ転送レベル (RTL) での正確なタイミング予測は、設計自動化における長年の課題です。既存のグラフベースの手法は、限られた受容野、高い複雑さ、信号の方向性の欠如に悩まされています。我々は、幅優先トラバーサルによるロジック コーンの線形化と最新の線形シーケンス モデルの適用により、スケーラブルな RTL タイミング予測を可能にする新しいシーケンス ベースのパラダイムである RTL シーケンサーを紹介します。さらに、シーケンス モデルは、シーケンス シャッフリング、双方向モデリング、微分可能モデリング、およびハイブリッド グラフシーケンス アーキテクチャを含む 4 つの相乗手法によってカスタマイズされます。広範な実験により、RTL-Sequencer が最先端のベースラインに比べて大幅に改善され、初期段階のタイミング最適化が前進したことが実証されました。

原文 (English)

RTL-Sequencer: Towards Scalable RTL Timing Prediction with the Sequence-based Paradigm

Accurate timing prediction at the register-transfer level (RTL) is a longstanding challenge in design automation. Existing graph-based methods struggle with limited receptive fields, high complexity, and a lack of signal directionality. We present RTL-Sequencer, a novel sequence-based paradigm that enables scalable RTL timing prediction via linearizing logic cones by breadth-first traversal and applying modern linear sequence models. Furthermore, sequence models are customized by four synergistic techniques, including sequence shuffling, bidirectional modeling, differentiable modeling, and a hybrid graph-sequence architecture. Extensive experiments demonstrate significant improvements of RTL-Sequencer over state-of-the-art baselines, advancing early-stage timing optimization.

13:00 JSTLLM/生成AI

CAMMAR: 比喩的なアラビア表現のための文化を意識したマトリョーシカ

アラビア語のメタファーは、意味を構築するための文化に基づいたメカニズムであり、解釈を形作る文化的知識をコード化しています。しかし、現在のアラビア語モデルは通常、語彙、文化、比喩的な情報を単一の表現空間に崩壊させており、これを「意味論的スミアリング」と呼んでいます。 CAMMAR (Culture-Aware Matryoshka for Metaphorical Arabic Representations) を紹介します。これは、段階的な意味論的カリキュラムを通じて、意味をネストされた語彙的、文化的、比喩的な埋め込み部分空間に組織化する表現学習フレームワークです。この設計は、アル・ジュルジャニのナズム理論の構成原理を実装し、以前の意味関係に基づいて構成的に根拠づけられた比喩的な意味をモデル化し、語彙表現と比喩表現の間の距離に基づいて訓練不要の比喩表現の幾何学的尺度を生成します。単語が一致する比喩/リテラルペアとして設定された新しいスパンアノテーション付きアラビア語比喩で評価すると、層間ジオメトリがペア監視によって形成されている場合、幾何学的読み出しは確率をはるかに超えて比喩を検出します (AUC は最大 0.84、ペアの 82.6\% で同じ単語について比喩が文字通り対応するものを上回ります)。監視下の判読可能な体制と緊急でない体制。制御されたアブレーションは、形態学的ルートに語彙層を接地すると、小さいながらも一貫したゲインが得られることを示しており、この効果は、測定アンカーとしての層の品質を反映する直接プローブには存在しません。承認され次第、データセット、文化的概念の目録、およびコードをリリースします。

原文 (English)

CAMMAR: Culture-Aware Matryoshka for Metaphorical Arabic Representations

Metaphor in Arabic is a culturally grounded mechanism for constructing meaning, encoding cultural knowledge that shapes interpretation. Yet current Arabic language models typically collapse lexical, cultural, and metaphorical information into a single representational space, a phenomenon we term "semantic smearing". We introduce CAMMAR (Culture-Aware Matryoshka for Metaphorical Arabic Representations), a representation learning framework that organizes meaning into nested lexical, cultural, and metaphorical embedding subspaces through a staged semantic curriculum. The design implements compositional principles of Al-Jurjani's theory of nazum, modeling figurative meaning as compositionally grounded in prior semantic relations, and yields a training-free geometric measure of metaphoricity based on the distance between lexical and metaphorical representations. Evaluated on a new span-annotated Arabic metaphor set as word-matched figurative/literal pairs, the geometric readout detects metaphor well above chance when the inter-layer geometry is shaped by paired supervision (AUC up to 0.84; figurative outscores its literal counterpart for the same word in 82.6\% of pairs), but sits at chance under an unsupervised domain contrast alone, a clean separation between a legible-under-supervision regime and a non-emergent one. A controlled ablation shows that grounding the lexical layer in morphological roots gives a small but consistent gain, an effect absent from direct probing that reflects the layer's quality as a measurement anchor. We will release the datasets, cultural concept inventory, and code upon acceptance.

13:00 JST画像/動画生成

現実的な自己回帰ビデオ生成のためのテスト時のノイズガイド適応

自己回帰ビデオ拡散モデルにより、将来のフレームの条件付けを取り除くことで任意の長さのビデオを生成できるようになり、計算効率が大幅に向上しました。しかし、ノイズ除去されたシーケンスはトレーニング中に見られる調整分布から徐々に離れていくため、時間の経過とともにエラーが蓄積されます。最近の進歩では、生成された各フレームを実際のフレームの学習された多様体に固定することによって、この誤差を削減しようとしています。ただし、生成されたすべての個々のフレームが実際の多様体に近い場合でも、モデルが十分な知識を欠いており、モデルから出ずに続行して終点に到達する軌跡が存在します。モデルが終端点にトラップされるのを防ぐために、適切にモデル化された将来の軌跡の場合、予測ノイズの分布が順方向ノイズ プロセスの分布と一致するはずであるという仮説から始めます。このような事前分布をテスト時に強制するために、ノイズガイド付き最適化による終端点回避 (TANGO) を導入します。これは、1 ステップ先を予測し、等方性ガウス ノイズ予測を必要とすることで、拡散モデルを自身の出力の批評家として使用します。この予想されるノイズ分布からの偏差を使用して、終点に至らない代替軌道を検索します。私たちのアプローチは、VBench で最先端技術と比較して $3.1\%$ の絶対的な改善を達成し、$15$ のビデオ全体で平均 $28.3\%$ のビデオ距離の短縮を実現します。私たちのコードは https://mever-team.github.io/tango で入手できます。

原文 (English)

Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation

Autoregressive video diffusion models have enabled the generation of arbitrarily long videos by removing conditioning on future frames, thus greatly improving computational efficiency. Yet, they suffer from error accumulation over time, as the denoised sequence gradually drifts away from the conditioning distribution seen during training. Recent advances attempt to reduce this error by anchoring each generated frame to the learned manifold of real ones. However, even when all generated individual frames lie close to the real manifold, there are trajectories which the model lacks sufficient knowledge to continue without exiting it, thus reaching a terminal point. To prevent the model from being trapped in terminal points, we start from the hypothesis that for well-modeled future trajectories the distribution of the predicted noise should match the one of the forward noising process. To enforce such a prior at test time, we introduce Terminal points Avoidance through Noise Guided Optimization (TANGO), which uses the diffusion model as a critic of its own outputs, by predicting one step forward and requiring an isotropic Gaussian noise prediction. We use the deviation from this expected noise distribution to search for an alternative trajectory that does not lead to a terminal point. Our approach achieves a $3.1\%$ absolute improvement on VBench over state-of-the-art, while reducing Fr\'echet Video Distance by $28.3\%$ on average across $15$s videos. Our code is available on https://mever-team.github.io/tango.

13:00 JSTエージェントGPT / ChatGPT

反例を補足したスケッチに対するエージェント合成

コーディング エージェントは、失敗の原因となったドメイン ルールを保存せずに、失敗した例を修正できるため、後の世代が同じ間違いを繰り返す可能性があります。我々は、実装中に管理ポリシーが発見されるシステムのためのリポジトリネイティブな方法である、反例が補足されたスケッチに対するエージェント合成を紹介します。人間は部分的なコード形状のスケッチから開始し、コーディング エージェントが最初の実装を生成します。具体的な障害によってポリシーの欠落または誤りが明らかになった場合、オペレーターは修正された動作とルールを明示的に承認します。次に、エージェントはスケッチを改訂し、コードを修復または再生成し、その 1 つの反例に対するプロンプト サーフェスを表示します。完全なアーカイブでは出所が保存されます。選択された回帰セットは、次の候補が明らかになる前に各リビジョンをゲートします。そして、定期的なクリーンな再生成では、プロンプト履歴や蓄積された例ではなく、進化したスケッチが学習されたポリシーを保持しているかどうかをテストします。合成ブラウザ アプリケーションとキャプチャされたコーディング エージェントの実験である CatSynth を使用してこの方法を示します。 GPT-5.4-mini を使用したオープンワールドでの 1 つの実行では、凍結された候補ケース 14 件のうち 8 件が反例となりました。再構築コントロールはそのプロモーション スケジュールを継承し、3 つのパスすべてが 8 つの受け入れられたケースを通過しました。進化したスケッチから再構築すると、保留されたケース 21 件中 19 件が合格しました。これに対し、最初のスケッチから再構築して受け入れられたすべての例を再生した場合は、21 件中 15 件が合格しました。反例全体でコードを保持するには、9 回の開発者コールと 719 行の累積アーティファクト チャーンが必要でした。一方、リプレイオールの場合は 15 回のコールと 2,394 行が必要で、保留されたケース 21 件中 18 件が合格しました。これらの結果は、進化したスケッチがレビューされたポリシーを実行し、コードを保持することでこの実行での手戻りが減少したことを示す検査可能な証拠を提供します。 1 つのモデルと 1 つの公開順序では、エンコードされたチェックを超えた一般的な優位性や正確性は確立されません。

原文 (English)

Agentic Synthesis against Counterexample-Supplemented Sketches

Coding agents can fix a failing example without preserving the domain rule that made it fail, so later generations can repeat the same plausible mistake. We present agentic synthesis against counterexample-supplemented sketches, a repository-native method for systems whose governing policy is discovered during implementation. A human starts with a partial, code-shaped sketch, and a coding agent generates the first implementation. When a concrete failure exposes missing or mistaken policy, an operator explicitly approves the corrected behavior and rule. The agent then revises the sketch and repairs or regenerates code and prompt surfaces for that one counterexample. The full archive preserves provenance; a selected regression set gates each revision before the next candidate is revealed; and periodic clean regeneration tests whether the evolved sketch, rather than prompt history or accumulated examples, carries the learned policy. We demonstrate the method with CatSynth, a synthetic browser application and captured coding-agent experiment. In one open-world run with GPT-5.4-mini, 8 of 14 frozen candidate cases became counterexamples. The rebuild controls inherited that promotion schedule, and all three paths passed the 8 accepted cases. Rebuilding from the evolved sketch passed 19 of 21 withheld cases, compared with 15 of 21 when rebuilding from the initial sketch and replaying all accepted examples. Retaining code across counterexamples required 9 Developer calls and 719 lines of cumulative artifact churn, versus 15 calls and 2,394 lines for replay-all, and passed 18 of 21 withheld cases. These results provide inspectable evidence that the evolved sketch carried reviewed policy and that retaining code reduced rework in this run; with one model and one reveal order, they do not establish general superiority or correctness beyond the encoded checks.

13:00 JSTLLM/生成AI

多言語およびコードが混在した不正行為を検出するための有害信号の条件付き信頼性

モデレーション システムは外部の毒性ツールへの依存度が高まっていますが、コードの混在、音訳、スラング、言語の不一致がある場合、それらのツールは信頼できなくなります。私たちは、インドの多言語およびコードが混在した短いテキストにおける毒性事前確率の \emph{条件付き信頼性} を研究します。英語の毒性、インドの虐待、およびルールベースの重症度の手がかりは有用な証拠となり得ますが、それは一部の言語および虐待の重症度の文脈でのみです。我々は、各補助信号を予測状態に追加する前にエンコーダー表現で条件付けを行うトラスト・フュージョン・ヘッドである ToxGate を提案します。 3 つの短文悪用データセット、4 つのトランスフォーマー エンコーダー、設定ごとに 5 つのシードにわたって、ToxGate は、ドメイン内設定 12 件中 10 件と転送設定 8 件中 7 件で、一致するプレーン エンコーダーよりも改善しています。最大かつ最も解釈可能な利益は、露骨な中傷、暴力的な脅威、データセット間の転送など、高リスクのモデレーション スライスで発生します。より広範な教訓は、調整システムは、固定特徴やグラウンドトゥルースではなく、外部毒性ツールと事前分布を条件付き証拠として扱うべきであるということです。集中的アブレーションでは、ソース固有のゲーティングが、移植、重度の虐待スライス、および高リスクのトリアージで最も強力な結果をもたらします。

原文 (English)

Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \emph{conditional reliability} of toxicity priors in Indian multilingual and code-mixed short text: English toxicity, Indic abuse, and rule-based severity cues can be useful evidence, but only in some linguistic and abuse-severity contexts. We propose ToxGate, a trust-fusion head that conditions each auxiliary signal on the encoder representation before adding it to the prediction state. Across three short-text abuse datasets, four transformer encoders, and five seeds per setting, ToxGate improves over matched plain encoders in 10 of 12 in-domain settings and 7 of 8 transfer settings. The largest and most interpretable gains occur in high-risk moderation slices, including explicit slurs, violent threats, and cross-dataset transfer. The broader lesson is that moderation systems should treat external toxicity tools and priors as conditional evidence rather than fixed features or ground truth, in focused ablations, source-specific gating gives the strongest results in transfer, severe-abuse slices, and high-risk triage.

13:00 JST画像/動画生成ロボティクス

EgoExoMoCap: 分散型 Ego-Exo ヒューマン モーション キャプチャ

ヘッドマウント デバイス (HMD) からのヒューマン モーション キャプチャは、現実世界の人間の動作とインタラクション データを取得するスケーラブルな方法を提供します。これは、身体型 AI や VR/AR のアプリケーションにとって重要です。既存のアプローチは、デバイスを装着している被験者の動きを推定する自己中心的な身体追跡、または装着者の周囲の人々の動きを捕捉する外中心的な追跡のいずれかに焦点を当てています。これまでのところ、これら 2 つのパラダイムは主に個別に検討されてきました。この論文では、HMD からの人間の動作推定のためにエゴセントリックおよびエキソセントリックなマルチモーダル信号を共同利用する新しい分散フレームワークを提案します。かさばるマルチカメラのセットアップや邪魔なモーション キャプチャ スーツを必要とする従来のモーション キャプチャ システムとは異なり、私たちのアプローチである EgoExoMoCap は、2 人 (またはそれ以上) がそれぞれスマート グラスを着用するだけで簡単です。この方法は、3D 世界でのグローバル モーションを正確に推定するために頭部 (および場合によっては手首) 追跡信号を活用し、DINOv3 に基づくコンテキスト認識画像特徴を組み合わせて、ノイズやオクルージョンが存在する場合の堅牢性を実現します。 2 つの実際のデータセットに対する広範な実験により、私たちのアプローチが困難なシナリオであっても動きを堅牢に再構築できることが示されています。

原文 (English)

EgoExoMoCap: Distributed Ego-Exo Human Motion Capture

Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR. Existing approaches focus on either egocentric body tracking, estimating the motion of the subject wearing the device, or exocentric tracking, capturing the movements of people in the wearer's surroundings. So far, these two paradigms have largely been explored in isolation. In this paper, we propose a novel distributed framework that jointly leverages ego- and exocentric multi-modal signals for human motion estimation from HMDs. Unlike traditional motion capture systems requiring bulky multi-camera setups or obtrusive mocap suits, our approach, EgoExoMoCap, is as simple as two (or more) people, each wearing a pair of smart glasses. The method leverages head (plus potentially wrist) tracking signals for accurate estimation of global motion in the 3D world and combines context-aware image features based on DINOv3 to achieve robustness in the presence of noise and occlusions. Extensive experiments on two in-the-wild datasets show that our approach can robustly reconstruct motion even in challenging scenarios.

13:00 JSTLLM/生成AI研究/論文

DECODEM: 強化された方法による企業組織文書からのデータ抽出

実証的な法律研究の多くは、非構造化テキストを構造化変数に変換することに依存しています。コーポレートガバナンス研究においては、他の分野と同様に、この翻訳は従来、憲章や細則などの文書を人間がコーディングすることに依存してきましたが、このプロセスはコストがかかり、拡張が難しく、不透明なことが多いです。このペーパーでは、組織文書からのコーポレート ガバナンス変数の自動抽出を評価するためのベンチマーク データセットのセットである DECODEM を紹介します。このベンチマークは、ランダムにサンプリングされた企業憲章および細則と、実証研究で一般的に研究される一連のガバナンス規定をカバーする高品質の人による注釈を組み合わせています。この論文では、これらのデータセットを使用して、プロンプト設計、タスク分解、およびドキュメント処理が異なるいくつかの大規模言語モデル抽出パイプラインを評価します。基礎となるタスクは、ガバナンス変数ごとに 1 つずつ、ドキュメント レベルのバイナリ分類問題のセットで構成されます。結果は、自動抽出が多くのプロビジョニングで高レベルの精度で実現可能であり、各アプローチの上限に近いパフォーマンスの中央値を示しています。同時に、パフォーマンスは変数全体で系統的に変化し、少数のプロビジョニングが残りのエラーの大部分を占めます。より精巧なプロンプト戦略とカスケード パイプラインは、フロンティア モデルのパフォーマンスを一貫して向上させるわけではありませんが、一部の設定ではフロンティア モデルと効率指向モデルの間のギャップを大幅に縮め、パイプライン設計がモデルの機能を部分的に代替できることを示唆しています。この論文は、標準化されたベンチマークと抽出方法の体系的な評価を提供することにより、現在のフロンティアモデルが複雑な企業文書から法的に意味のある情報を高精度で抽出できることを実証し、コーポレートガバナンスデータセットの構築における自動特徴抽出の将来の重要な役割を示唆しています。

原文 (English)

DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods

Much empirical legal research depends on translating unstructured text into structured variables. In corporate governance research as elsewhere, this translation has traditionally relied on human coding of documents such as charters and bylaws, a process that is costly, difficult to scale, and often opaque. This paper introduces DECODEM, a set of benchmark datasets for evaluating the automated extraction of corporate governance variables from organizational documents. The benchmarks pair randomly sampled corporate charters and bylaws with high-quality human annotations covering a range of governance provisions commonly studied in empirical work. Using these datasets, the paper evaluates several large-language-model extraction pipelines that vary in prompt design, task decomposition, and document handling. The underlying task consists of a set of document-level binary classification problems, one for each governance variable. The results show that automated extraction is feasible at a high level of accuracy for many provisions, with median performance near the upper bound across approaches. At the same time, performance varies systematically across variables, with a small number of provisions accounting for most of the remaining errors. More elaborate prompting strategies and cascading pipelines do not consistently improve performance for frontier models, but substantially narrow the gap between frontier and efficiency-oriented models in some settings, suggesting that pipeline design can partly substitute for model capability. By providing a standardized benchmark and a systematic evaluation of extraction methods, the paper demonstrates that current frontier models can extract legally meaningful information from complex corporate documents with high accuracy and suggests an important future role for automated feature extraction in constructing corporate governance datasets.

13:00 JST研究/論文

認識されるAGI: 能力ではなく、次元の完全性としての信頼性

大規模な言語モデルは幅広い機能を備えていますが、継続的な 1 対 1 の会話では依然として平坦なものとして読み取られます。有能で、応答性があり、そしてどういうわけか、心が存在しているとは言えません。私たちは、欠けている中心的な要素は、さらなる能力ではなく、次元の完全性であると仮説を立てています。私たちは、人工対話者の信憑性、つまりユーザーが内面の生活をその人工対話者に帰属させる度合い(知覚された心と呼ぶ)は、そのエージェントが人間が心の証拠として使用する少数の一人称スタンスを表現するかどうかによって決まり、これはタスク知能から分離可能である、と提案します。私たちはそのような 4 つの次元、時間、真実、エントロピー、愛と名付けています。それぞれの次元は、基準となる能力ではなく行動のスタンスとして定義され、それぞれが人間の類似点と具体的なエミュレーション パスを持っています。時間次元には、作成者が報告したプロトタイプがすでにあります。私たちは、会話の中でスタンスが表面化する際に介して、イニシアチブ (即発的なアクション) とケイデンス (ターンの形状とタイミング) という観察可能な動作レイヤーを特定します。どちらも、実稼働コンパニオン アプリケーションにデプロイされた機能として部分的に実現されます。我々は、現在事前登録可能なものと実用化までの推測が残っているものを分けて、将来の事前登録された研究で検証される6つの反証可能な予測を述べています。これは概念的な枠組みです。人間を対象としたデータは報告されておらず、その中心となる比較主張は結果ではなく予測です。全体を通して、私たちはしっかりとした境界線を保持しています。対象は推測可能な内面性であり、内面性ではありません。これは知覚工学であり、機械意識の理論ではありません。そして私たちは、結果として生じる執着と操作のリスクを、偶発的ではなく、負荷を伴うものとして扱います。

原文 (English)

Perceived AGI: Believability as Dimensional Completeness, Not Capability

Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind. We hypothesize that a central missing ingredient is not more capability but dimensional completeness. We propose that the believability of an artificial interlocutor -- the degree to which a user attributes an inner life to it, which we call perceived mind -- is governed by whether the agent expresses a small set of first-person stances that humans use as evidence of mind, and that this is separable from task intelligence. We name four such dimensions -- time, truth, entropy, and love -- each defined as a behavioral stance rather than a benchmark competency, each with a human analog and a concrete emulation path; the time dimension already has an author-reported prototype. We identify an observable behavior layer -- initiative (unprompted action) and cadence (the shape and timing of turns) -- through which the stances surface in conversation, both partially realized as deployed features in a production companion application. We state six falsifiable predictions that a later pre-registered study will test, separating those that are pre-registrable now from those that remain conjectures pending operationalization. This is a conceptual framework: it reports no human-subjects data, and its central comparative claims are predictions, not findings. Throughout we hold a firm boundary -- the object is inferrable interiority, not interiority; this is perception engineering, not a theory of machine consciousness -- and we treat the resulting attachment and manipulation risks as load-bearing rather than incidental.

13:00 JSTLLM/生成AI

両方向の誘導: マスクされた拡散言語モデルにおける文脈内学習の機構分析

自己回帰 (AR) トランスフォーマーの内部メカニズムは広く研究されていますが、反復的なノイズ除去によってテキストを生成する新たな代替手段である拡散言語モデル (DLM) についてはほとんど知られていません。この研究では、モデルが繰り返されるコンテキストを見つけてそれに続くトークンをコピーするコンテキスト内学習の背後にあるメカニズムである帰納法を DLM がどのように実装するかを研究します。私たちの分析では、アテンションのみの AR モデルと、一致するアーキテクチャを持つ吸収マスク DLM を比較します。 DLM が双方向誘導回路を学習することがわかりました。そこでは、前のトークンと次のトークンのヘッドがローカル コンテキストを残差ストリームに書き込み、後の誘導ヘッドがそれを使用して、一致するソース位置から答えを見つけてコピーします。この回路は方向対称であり、ソースが過去に現れても未来に現れても機能します。 AR モデルが認識するものと一致する、左側のコンテキストのみが表示されている場合、DLM は誘導機能において AR の対応物を上回るパフォーマンスを発揮しません。ただし、マスクされたトークンの両側が表示されている場合には、より強い誘導があり、より強力な一方的なメカニズムではなく、双方向のコンテキスト アクセスを示していることがわかります。帰納を超えて、たとえ明示的なタイムステップ埋め込みが与えられていないとしても、DLM がマスクされたトークンのグローバル部分を計算し、それを暗黙的なタイムステップとして使用するという因果関係の証拠を提供します。

原文 (English)

Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models

While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising. In this work, we study how DLMs implement induction, a mechanism behind in-context learning in which the model finds a repeated context and copies the token that followed it. Our analysis compares attention-only AR models and absorbing-mask DLMs with matched architectures. We find that DLMs learn a bidirectional induction circuit, where previous-token and next-token heads write local context into the residual stream and later induction heads use it to find and copy the answer from the matching source position. The circuit is direction-symmetric, working whether the source appears in the past or in the future. When only left context is visible, matching what an AR model sees, the DLM does not outperform its AR counterpart in induction capabilities. However, we observe it has stronger induction when both sides of the masked token are visible, pointing to bidirectional context access rather than a stronger one-sided mechanism. Beyond induction, we provide causal evidence that DLMs compute the global fraction of masked tokens and use it as an implicit timestep, even though they are given no explicit timestep embedding.

13:00 JST画像/動画生成ロボティクス

Orbis 2: 運転のための階層的世界モデル

現在の世界モデルは、単一の抽象レベルで動作し、知覚の忠実度を最も優先しますが、現実世界の下流タスクに必要な空間的推論と意味論的な理解を欠いています。我々は、異なる時間的スケールと抽象化スケールで動作する 2 つのレベルにわたって将来予測を因数分解する階層的な駆動世界モデルを提示します。1 つは拡張された時間的範囲にわたって大まかなシーン構造を予測する高レベルの予測器、もう 1 つは高レベルの出力に条件付けされた詳細な予測を生成する低レベルのジェネレーターです。この分解により、高い知覚忠実度が得られると同時に、強力な空間表現と意味表現も捕捉されます。さらに、拡散強制目的を使用した事前トレーニングでは、標準の教師強制目的よりも大幅に豊富な内部表現が生成される一方、教師強制 (クリーンなコンテキストから次のフレームのみを予測) では、より安定した自己回帰ロールアウトが生成されることを示します。したがって、拡散強制でモデルを事前トレーニングし、教師強制で微調整する一般的な 2 段階のトレーニング パラダイムを導入し、前者の表現上の利点と後者のロールアウトの安定性を組み合わせます。私たちのアプローチは、ロングホライズン生成の忠実度、反事実シナリオで評価されたステアリング応答性、内部表現の品質など、確立されたベンチマークに基づいたドライビングワールドモデル評価の標準スイート全体で最先端の結果を達成します。コード、デモ、チェックポイント、定性的結果を含むプロジェクト ページ: https://lmb-freiburg.github.io/orbis2.github.io/

原文 (English)

Orbis 2: A Hierarchical World Model for Driving

Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoning and semantic understanding required for real-world downstream tasks. We present a hierarchical driving world model that factorizes future prediction across two levels operating at distinct temporal and abstraction scales: a high-level predictor that forecasts coarse scene structure over extended temporal horizons, and a low-level generator that produces detailed predictions conditioned on the high-level output. This decomposition yields high perceptual fidelity while also capturing strong spatial and semantic representations. We further show that pretraining with a diffusion forcing objective yields substantially richer internal representations than the standard teacher forcing objective, while teacher forcing -- predicting only the next frame from clean context -- produces more stable autoregressive rollouts. We therefore introduce a generic two-stage training paradigm that pretrains the model with diffusion forcing and fine-tunes with teacher forcing, combining the representational benefits of the former with the rollout stability of the latter. Our approach achieves state-of-the-art results across the standard suite of driving world model evaluations on established benchmarks, including long-horizon generation fidelity, steering responsiveness evaluated on counterfactual scenarios, and internal representation quality. Project page with code, demo, checkpoints and qualitative results: https://lmb-freiburg.github.io/orbis2.github.io/

13:00 JST研究/論文

ボトルネックのある生成アーキテクチャにおける境界探索蒸留の失敗について

データフリーの知識の蒸留では、元のトレーニング データにアクセスすることなく、教師モデルにエンコードされた知識が生徒モデルに転送されます。 Contrastive Abductive Knowledge Extraction (CAKE) などの以前の研究では、教師の決定境界付近でサンプルを合成することで分類器に対してこれを実現しています。この研究では、MNIST データセットでの実験を通じて、この境界探索原理がオートエンコーダー蒸留にも拡張されるかどうかを調査します。直接比較を可能にするために、連続再構成を高密度の特徴ごとの分類タスクとして再定式化し、デコーダーがカテゴリロジットを出力できるようにします。境界探索の目的は、ボトルネックのある生成アーキテクチャでは根本的に不適切であることを示します。 CAKE は単一のインスタンス レベルの目的で動作しますが、デコーダーは、共有の低次元ボトルネックによって制約される、密結合された特徴レベルの分類器の配列として機能します。これらの結合された出力に対して対照的なターゲットを独立してサンプリングすると、学習された潜在多様体の幾何学構造に違反し、有益な境界サンプルの代わりに重大な勾配の競合が生成されます。多様体認識合成は、これらの競合を完全に回避し、データフリーの生成蒸留のための効果的なベースラインを確立します。

原文 (English)

On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures

Data-free knowledge distillation transfers the knowledge encoded in a teacher model to a student model without access to the original training data. Prior work such as Contrastive Abductive Knowledge Extraction (CAKE) achieves this for classifiers by synthesizing samples near the teacher's decision boundary. In this work, we investigate whether this boundary-seeking principle extends to autoencoder distillation through experiments on the MNIST dataset . To enable a direct comparison, we reformulate continuous reconstruction as a dense, per-feature classification task, allowing the decoder to output categorical logits. We show that boundary-seeking objectives are fundamentally ill-posed in bottlenecked generative architectures. CAKE operates on a single, instance-level objective, but a decoder acts as an array of tightly coupled, feature-level classifiers constrained by a shared low-dimensional bottleneck. Independently sampling contrastive targets for these coupled outputs violates the geometry of the learned latent manifold and produces severe gradient conflicts instead of informative boundary samples. Manifold-aware synthesis bypasses these conflicts entirely and establishes an effective baseline for data-free generative distillation.

13:00 JST研究/論文

自動化すべきではない場合: AI に最適化された組織における人間の保存のための正式なプロトコル

標準的な自動化の ROI では、長期的な組織パフォーマンスに影響を与える 4 つのカテゴリのシステミック リスク (暗黙知の浸食、回復力の低下、規制への曝露、社会制度的資本の劣化) が見逃されます。 PHP-AIO (AI に最適化された組織における人間の保存のためのプロトコル) は、これらの価格付けされていないシステミック リスクを役割レベルで定量化し、監査可能な自動化の決定を生成する最終複合チェックを備えた 5 ゲートの順次決定プロトコルです。クローズド形式の自動化負債尺度 ($\rho(P)$) は、役割レベルの意思決定が複数段階のプロセスにわたってどのように蓄積されるかを形式化します。その警告は、規制当局が義務付けた人間参加型アンカーによってのみ無効化されます。 PHP-AIO は、代表的な社内役割の定型化されたプロファイルに適用され、標準的な費用対効果分析では一律に自動化される候補者に対して、自動化、拡張、ハイブリッド、保存という明確な結果をもたらします。閾値感度解析により、代表的な 4 つのケースのうち 3 つにおいて、ゲートの決定が少なくとも 14% の上方摂動に対して堅牢であることが確認されています。キーワード: AI ガバナンス、自動化の意思決定、人間の監視、暗黙知、組織の回復力、金融サービス

原文 (English)

When Not to Automate: A Formal Protocol for Human Preservation in AI-Optimized Organizations

Standard automation ROI misses four categories of systemic risk -- tacit knowledge erosion, resilience reduction, regulatory exposure, and socio-institutional capital degradation -- that affect long-term organizational performance. PHP-AIO (Protocol for Human Preservation in AI-Optimized Organizations) is a five-gate sequential decision protocol with a final composite check that quantifies these unpriced systemic risks at the role level and produces auditable automation decisions. A closed-form automation-debt measure ($\rho(P)$) formalises how role-level decisions accumulate across multi-step processes; its warning is neutralised only by a regulator-mandated human-in-the-loop anchor. Applied to stylised profiles of representative internal roles, PHP-AIO produces distinct outcomes -- automate, augment, hybrid, and preserve -- for candidates that standard cost-benefit analysis would uniformly automate. Threshold sensitivity analysis confirms the gate decisions are robust to upward perturbations of at least 14% in three of four representative cases. Keywords: AI governance, automation decision, human oversight, tacit knowledge, organizational resilience, financial services

13:00 JSTエージェント

意見形成に対する社会文化的影響: 口コミの力学、マスメディア、行動発達

私たちは、同じまたは異なるグループ内の他の人の状況について意見を形成する、多数の職業的または文化的グループに属するエージェントの社会を研究します。意見は、自分のグループ内で観察することによって、または他のグループのメンバーと直接対話することによって、つまり口伝え(WoM)によって発展します。さらに、さまざまなグループの状況について間接的に情報を提供する世界的なマスメディア (MM) が利用できる場合もあります。これらのプロセスの社会文化的相互作用と各情報源への相対的な曝露の程度は、最終的な意見と社会的態度の形成に多様な影響を与えます。 WOM によってすべての人がすべての人と物理的に対話できるわけではない、大規模で複雑な社会やグループでは、これらのプロセスは大量制御とソーシャル オートメーション エンジニアリングの可能性を示しています。また、私たちのモデルは、異なる職業的および文化的特徴を持つグループで構成される細分化された社会を表し、それについて一般的に有益な情報を提供することもでき、世代間のギャップや社会的亀裂などの社会問題についての洞察を提供することもできます。

原文 (English)

Sociocultural Influences on Opinion Formation: Word of Mouth Dynamics, Mass Media and Behavioural Development

We study a society of agents belonging to a number of occupational or cultural groups that form opinions about others' situation in the same or different group. Opinions develop either by observation within own group or by directly interacting with members of other groups, therefore by word of mouth (WoM). Additionally, global mass media (MM) may be available that inform indirectly about the situation of the various groups. The sociocultural interplay of these processes and the degrees of relative exposure to each of the sources has diversified effects on final opinion and social attitude formation. In large and complex societies and groups where not everyone can physically interact by WoM with everyone about everything, these processes show potential for mass control and social automation engineering. Our model can also represent and be generally informative about segmented societies that consist of groups with different occupational and cultural characteristics and it can offer insights into social issues such as the generation gap, social cleavages and so on.

13:00 JST研究/論文

低電圧送電網における強化学習ベースの輻輳管理の堅牢性

太陽光発電、電気自動車の充電、ヒートポンプ需要の増加により、低電圧配電網の運用限界が課題となっています。これには、まばらな可観測性、ノイズの多い測定、および不完全なグリッド モデルの下で動作できる治療的な削減方法が必要です。部分的に観察可能な抑制に対するこれまでのエンドツーエンドの強化学習アプローチとは異なり、この研究では、ランダムフォレスト違反事前分類器とアクタークリティカルコントローラーを組み合わせることによって輻輳の検出と制御を分離し、測定ノイズとグリッドパラメータの不一致に対する堅牢性を評価します。このフレームワークは、可観測性と制御性が低い合成将来の運用シナリオを使用して、実際の低電圧グリッドでテストされます。正確なグリッド パラメーターを使用すると、コントローラーは違反の総量を 98.9% 削減し、このパフォーマンスはテスト済みの測定ノイズ設定の下でほとんど変化しません。グリッド モデルの不一致はより困難であることが判明していますが、コントローラーはテストされた不一致の仮定の下でほとんどの違反を依然として軽減します。

原文 (English)

Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids

Increases in photovoltaic generation, charging of electric vehicles and heat-pump demand challenge operating limits in low-voltage distribution grids. This requires curative curtailment methods that can operate under sparse observability, noisy measurements, and imperfect grid models. Unlike prior end-to-end reinforcement-learning approaches for partially observable curtailment, this work decouples congestion detection and control by combining a random-forest violation pre-classifier with an actor-critic controller, and evaluates its robustness to measurement noise and grid-parameter mismatch. The framework is tested on a real low-voltage grid using synthetic future operating scenarios with low observability and controllability. With accurate grid parameters, the controller reduces total violation magnitude by 98.9%, and this performance remains nearly unchanged under the tested measurement-noise settings. Grid-model mismatch proves to be more challenging, but the controller still mitigates most violations under the tested mismatch assumptions.

13:00 JST画像/動画生成ロボティクス

DPNeXt: 効率的な ViT ベースのマルチタスク高密度予測のための軽量マルチスケール機能融合フレームワーク

ロボット認識システムのマルチタスク学習 (MTL) は、セマンティック セグメンテーションと深度推定を統合することにより、包括的な 3D 空間シーンの理解をサポートします。 Vision Foundation Model (VFM) は堅牢な特徴エンコーダとして採用されることが増えていますが、既存のデコード戦略には重大なボトルネックが存在します。これに対処するために、合理化されたマルチスケール特徴融合デコーダであり、標準的な高密度予測変換器 (DPT) の効率的な代替手段である DPNeXt を提案します。 DPNeXt は、深さ方向に分離可能な二重の逆ボトルネックを使用して、融合中心のデコードと独立したタスクのモジュール化を通じてフリーズした VFM の使用率を改善します。タスク間の負の帰納的伝達をさらに軽減するために、マルチタスク境界ガイダンス (MTBG) 戦略を導入します。融合モジュールやゲーティングを追加する従来の境界認識手法とは異なり、MTBG は対称境界に焦点を当てた監視を適用して、追加の注釈や推論コストを発生させることなく幾何学的一貫性を促進します。都市景観に関する実験では、DPNeXt-S が以前の最先端 (SOTA) MTL モデルよりも優れたパフォーマンスを示し、一方、DPNeXt-B は全体的なパフォーマンスをさらに向上させ、比較した方法の中で最高の結果を達成することが示されています。 NYUv2 では、DPNeXt-B は、比較した方法の中で最高のセマンティック セグメンテーションと深度推定結果を達成する一方で、必要なトレーニング可能なパラメーターは以前の大規模 MTL モデルよりも大幅に少なくなります。標準の DPT と比較して、DPNeXt-S はトレーニング可能なパラメータを 78.6% 削減し、リソースに制約のあるラップトップ ハードウェア上で比較したモデルの中で最速の推論速度を実現します。ソース コード、モデル チェックポイント、デモ ビデオは https://github.com/kangjehun/DPNeXt で公開されます。

原文 (English)

DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction

Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segmentation and depth estimation. While Vision Foundation Models (VFMs) are increasingly adopted as robust feature encoders, existing decoding strategies present a critical bottleneck. To address this, we propose DPNeXt, a streamlined multi-scale feature fusion decoder and efficient alternative to the standard Dense Prediction Transformer (DPT). DPNeXt uses dual depthwise separable inverted bottlenecks to improve frozen VFM utilization through fusion-centric decoding and independent task modularization. To further mitigate negative inductive transfer between tasks, we introduce the Multi-Task Boundary Guidance (MTBG) strategy. Unlike prior boundary-aware methods that add fusion modules or gating, MTBG applies symmetric boundary-focused supervision to encourage geometric consistency without extra annotation or inference cost. Experiments on Cityscapes show that DPNeXt-S outperforms prior state-of-the-art (SOTA) MTL models, while DPNeXt-B further improves the overall performance and achieves the best results among the compared methods. On NYUv2, DPNeXt-B also achieves the best semantic segmentation and depth estimation results among the compared methods while requiring substantially fewer trainable parameters than prior large-scale MTL models. Compared with the standard DPT, DPNeXt-S reduces trainable parameters by 78.6% and achieves the fastest inference speed among the compared models on resource-constrained laptop hardware. The source code, model checkpoints, and a demo video will be made available at https://github.com/kangjehun/DPNeXt.

13:00 JSTLLM/生成AIGoogle

BERT を使用した候補者の出席対話状態追跡

対話状態追跡 (DST) は、タスク指向の対話システムの中核コンポーネントの 1 つです。会話の各ターンで、DST はユーザーの信念または対話状態を推定します。これは、下流モジュールへの入力として使用され、システムの動作を予測して応答を生成します。 Google アシスタント、Siri、Alexa など、ますます人気が高まっている対話システム アプリケーションは、多数のサービスや API をサポートする必要があるため、そのようなシステムのスケーラビリティに対する注目が高まっています。特にトレーニング データがほとんどまたはまったくない一部のドメインでは、他のドメインの既存の知識を転送する機能が強く望まれています。この論文では、マルチドメイン対話状態追跡のための新しいスケーラブルなフレームワークを紹介します。提案されたシステムは、事前トレーニングされた BERT モデルを活用してゼロショット汎化を達成し、追加のトレーニングなしで新しいドメインに迅速に適応することを容易にします。私たちのモデルのパフォーマンスは、最近リリースされたスキーマベースの対話 (SGD) データセットに基づいて評価されており、以前のベースラインと比較して大幅な改善が示されています。

原文 (English)

Candidate Attended Dialogue State Tracking Using BERT

Dialogue state tracking (DST) is one of the core components in task-oriented dialogue systems. At each turn in a conversation, DST estimates the user belief or dialogue state, which is used as input for downstream modules to predict system actions and generate responses. The increasingly popular dialogue system applications like Google Assistant, Siri and Alexa need to support a large number of services and APIs, resulting in growing attention to the scalability of such systems. Especially for some domains with little or no training data, the capability of transferring existing knowledge of other domains is highly desired. In this paper, we present a novel scalable framework for multi-domain dialogue state tracking. The proposed system leverages the pretrained BERT model to achieve zero-shot generalization, making it easy to quickly adapt to new domains without additional training. The performance of our model is evaluated on recently released schema-based dialogue (SGD) dataset, showing significant improvement compared to previous baseline.

13:00 JST研究/論文

量子フィッシャー情報による量子継続学習の再考

量子継続学習は、以前に学習した知識を失うことなく、連続タスクで量子モデルをトレーニングすることを目的としています。ただし、変分量子分類器 (VQC) は、非定常タスク分布の下では壊滅的な忘却を引き起こす傾向があります。我々は、忘却を軽減するための量子フィッシャー情報 (QFI) に基づいた正則化手法である量子弾性重み統合 (QEWC) を提案します。測定依存の出力統計を通じてパラメーターの重要性を測定する古典的なフィッシャー情報 (CFI) に基づく従来の弾性重み統合とは異なり、QEWC は QFI を使用してパラメーター化された量子状態の固有の感度を定量化します。これにより、重要なパラメータが量子状態多様体の局所応答によって特定される情報幾何学的ビューが得られます。古典的な画像分類タスクや量子位相分類タスクを含む、逐次バイナリ分類タスクで訓練された VQC 上の QEWC を評価します。シミュレーションによると、正則化を行わない逐次トレーニングは深刻な物忘れを引き起こす一方で、CFI ベースの EWC と QFI ベースの QEWC の両方が以前のタスクの記憶力を向上させることが示されています。さらに、機構解析により、2 つの方法が異なる正則化幾何学を課すことが示されています。CFI は測定に敏感な方向に選択的に作用するのに対し、QFI はパラメータ空間に対してより密な状態幾何学的制約を課します。脱分極ノイズの下では、CFI 値は劣化した測定統計によって強く抑制されますが、QFI はノイズの多いパラメータ化された量子状態のより安定した感度構造を保存します。これらの結果は、QEWC が量子状態幾何学を通じて量子継続学習における忘却を研究し軽減するための物理的に動機付けられたアプローチであることを確立します。

原文 (English)

Rethinking Quantum Continual Learning with Quantum Fisher Information

Quantum continual learning aims to train quantum models on sequential tasks without losing previously learned knowledge. However, variational quantum classifiers (VQCs) are prone to catastrophic forgetting under nonstationary task distributions. We propose quantum elastic weight consolidation (QEWC), a quantum Fisher information (QFI)-informed regularization method for mitigating forgetting. Unlike conventional elastic weight consolidation based on classical Fisher information (CFI), which measures parameter importance through measurement-dependent output statistics, QEWC uses QFI to quantify the intrinsic sensitivity of the parameterized quantum state. This gives an information-geometric view in which important parameters are identified by the local response of the quantum state manifold. We evaluate QEWC on VQCs trained on sequential binary classification tasks, including classical image-classification and quantum phase-classification tasks. Simulations show that sequential training without regularization causes severe forgetting, while both CFI-based EWC and QFI-based QEWC improve retention of previous tasks. Mechanistic analyses further show that the two methods impose different regularization geometries: CFI acts selectively on measurement-sensitive directions, whereas QFI imposes a denser state-geometric constraint over parameter space. Under depolarizing noise, CFI values are strongly suppressed by degraded measurement statistics, while QFI preserves a more stable sensitivity structure of the noisy parameterized quantum state. These results establish QEWC as a physically motivated approach for studying and mitigating forgetting in quantum continual learning through quantum-state geometry.

13:00 JST研究/論文

表形式の基盤モデルを使用したデータ駆動型の動的セキュリティ評価の再検討

データ駆動型の事前故障動的セキュリティ評価 (DSA) は、機械学習を使用して、電力システムにおける信頼できる不測の事態の動的リスクを迅速に評価します。既存のアプローチには 2 つの制限があります。まず、トレーニングには大規模なラベル付きデータベースが必要であり、信頼できる偶発事象の潜在的に長いリスト内の各偶発事象に対して個別のモデルがトレーニング、調整、維持されます。第二に、トレーニングされたモデルは、目に見えない不測の事態に対してあまり一般化されません。この研究では、再トレーニングやハイパーパラメータの最適化を必要とせず、コンテキスト内学習を通じて安定性を評価する表形式基盤モデル (TFM) を使用することで制限に対処しています。単一の TFM で多くの偶発事象を一度に評価できるため、分類子ごとに 1 つのモデルが必要なくなります。また、連続特徴として電気距離座標 (EDC) を使用することで、目に見えない偶然性に対する TFM の一般化が可能になる場合とそうでない場合を特徴づけ、少数のラベル付きサンプルが一般化を確実に向上させる方法を実証します。 IEEE 68 バス システムに関する包括的なケース スタディを通じて、モデルの再トレーニングやハイパーパラメータ調整を行わなくても、単一の TFM が、偶発事象あたりわずか 120 個のラベル付きサンプルで、従来の想定よりもおよそ 2 桁少ない、平均マクロ F1 スコア約 90% を達成することを示します。新しい/まだ見たことのない緊急事態については、EDC エンコーディングを使用して新しい緊急事態のラベル付きサンプルを 10 個だけ使用することで、達成可能な最良の転移学習オラクル モデルと一致することを示しますが、これには完全にラベル付けされたデータが必要であり、実際にはデプロイできません。全体として、この初期調査は、複数の運用タスクにわたる可能性のあるアプリケーションを備えた、電力システム運用の基礎モデルの開発と展開に向けた道を切り開きます。

原文 (English)

Revisiting data-driven dynamic security assessment with a tabular foundation model

Data-driven pre-fault dynamic security assessment (DSA) rapidly evaluates the dynamic risk of credible contingencies on a power system using machine learning. Existing approaches face two limitations. First, they require a large labelled database for training, with a separate model trained, tuned, and maintained for each contingency in a potentially long list of credible contingencies. Second, the trained models generalize poorly to unseen contingencies. This work addresses the limitations by using a tabular foundation model (TFM) that assesses stability through in-context learning, requiring no retraining or hyperparameter optimization. A single TFM can assess many contingencies at once, removing the need for one model per classifier. We also characterize when the use of electrical distance coordinates (EDC) as continuous features enables generalization of TFM to unseen contingencies and when they do not, demonstrating how a few labelled samples can reliably improve generalization. Through comprehensive case studies on the IEEE 68-bus system, we show that a single TFM attains an average Macro F1 score of about 90% with only 120 labelled samples per contingency, roughly two orders of magnitude fewer than conventionally assumed, without any model retraining or hyperparameter tuning. For new/unseen contingencies, we show that using just 10 labelled samples of the new contingency with EDC encoding matches the best achievable transfer learning oracle model, which requires fully labelled data and is not deployable in practice. Overall, this initial study paves the way towards developing and deploying foundation models for power system operations, with possible applications across multiple operational tasks.

13:00 JSTLLM/生成AI

ルーピーズをループさせよう!

これまでで最も強力なループ型トランスフォーマー、Loopie を紹介します。 Loopie シリーズは、2B のアクティブ パラメータを備えた 20B パラメータ モデルと、0.6B のアクティブ パラメータを備えた 6B パラメータ モデルの 2 つの専門家混合 (MoE) モデルで構成されています。ループ トランスフォーマーは長い間課題に直面していました。事前トレーニングの計算が N 倍に増加すると、パラメーター数を N 倍に増やすと、通常はモデルを N 回ループするよりもパフォーマンスが向上します。 Loopie はこの課題に取り組みます。バニラ 30B-A3B モデルとの比較を含む広範なアブレーション研究により、Loopie が同じコンピューティング バジェットでトレーニングされたバニラ Transformer ベースラインを大幅に上回るパフォーマンスを示しています。私たちの新しいトレーニング後のパイプラインは、Loopie に強力な推論能力を与えます。 2025 年の IMO および IPhO で、Loopie は工具なしで金メダルのパフォーマンスを達成しました。

原文 (English)

Loop the Loopies!

We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6Bparameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N-fold increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla Transformer baselines trained with the same compute budget. Our novel post-training pipeline equips Loopie with strong reasoning abilities. At the 2025 IMO and IPhO, Loopie achieves gold-medal performance without tools.

13:00 JSTLLM/生成AI研究/論文

ビジネス分野全体にわたる最先端の AI パフォーマンス: ナレッジワークと分析的推論の事例に基づいたベンチマーク

大規模言語モデル (LLM) は、ベンチマーク スコアに反映されているように急速に改善されていますが、これらの AI ベンチマークでは主に、事実の再現、限定的な質問応答、数学的問題解決、コーディングやエージェント ツールの使用などの機能がテストされます。まだ十分に測定されていないのは、複雑な情報の統合、不確実性と不完全な情報の下での判断の行使、複数のステークホルダーの状況での戦略的および敵対的思考の適用、トレードオフの比較検討、防御可能な構造化された分析の作成など、ホワイトカラーの専門家が日々行っている分析知識作業における AI の進歩です。このギャップは、そのような仕事の主観的な要素ではさらに顕著であり、成功を定義するのが難しい場合があります。トップクラスのビジネススクールが実践する「ケースメソッド」教育形式は、この測定ギャップに対処するための自然な基盤を提供します。私たちは、18 分野にわたるビジネスケースから抽出された数百の質問にわたるベンチマークである BusinessCaseBench を構築します。各質問は、専門家が作成した講師のケースソリューションから導き出された採点ルーブリックと対になっています。 BusinessCaseBench では、フロンティア AI モデルはすでにインストラクターのルーブリックに対して高いスコアを獲得しており、1 つのモデル ファミリー内の機能は 2 年間で大幅に向上しています。これらの結果は、この種の作業における AI のパフォーマンスがすでに高く、急速に向上していることを示す強力な証拠を提供します。これは、事例教育学によって学部生や MBA がこの種の分析的推論を訓練されるビジネス スクールや、歴史的にそのようなスキルが初期キャリアの仕事に定着してきたエントリーレベルの専門職に影響を及ぼします。

原文 (English)

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.

13:00 JST研究/論文

モデルをマージすると競合する場合 共同マルチタスク強化学習: タスクベクトル幾何学解析

モデルのマージは、共同マルチタスク トレーニングの代替として推進されていますが、強化学習の設定では、この置換は基本的に、置き換えると主張するベースラインに対してテストされることはありません。メソッドは、まさに共同モデルが利用できないため、独立してリリースされたエージェントをマージします。欠落している比較を構築します。 LOOP を使用して AppWorld エージェント ベンチマークで難易度 1 と難易度 2 の Qwen3-8B スペシャリストをトレーニングし、それらをマージし (TIES、RAM+)、その結果を同じデータで共同トレーニングされたモデルと照合します。タスクと目標の完了時に、マージは結合 RL と一致し、すべてのマージ バリアントは統計的に区別できません。ここでマージ方法が重要ではない理由を説明するために、タスク サンプリング ノイズを持たないスペシャリストのタスク ベクトルの幾何学的形状を測定します。サポート オーバーラップが約 65% あるにもかかわらず、それらはほぼ直交 (コサイン 0.06 ~ 0.10) しており、トレーニングを通じて成長する小さな共有方向であり、低ランクのパラメーター化ではなく学習を反映していることを確認するためにランダム初期化フロアと同一実行シーリングに対して校正しています。方向性とサポートが分離されているため、サポートとサインベースのマージ (RAM、TIES) はほぼ均一な平均​​化に崩壊します。すべてのコードと統計を公開します。

原文 (English)

When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis

Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is essentially never tested against the baseline it claims to replace: methods merge independently released agents precisely because a joint model is unavailable. We build the missing comparison. Training difficulty-1 and difficulty-2 Qwen3-8B specialists on the AppWorld agent benchmark with LOOP, we merge them (TIES, RAM+) and pit the result against a jointly trained model on the same data. On task-goal completion, merging matches joint RL -- and every merge variant is statistically indistinguishable. To explain why merge method does not matter here, we measure the geometry of the specialists' task vectors, which carries no task-sampling noise: they are near-orthogonal (cosine 0.06 - 0.10) despite ~65% support overlap, a small, shared direction that grows over training and that we calibrate against a random-init floor and a same-run ceiling to confirm it reflects learning, not the low-rank parameterization. Because direction and support are decoupled, support and sign-based merging (RAM, TIES) collapse to near-uniform averaging. We release all code and statistics.

13:00 JST画像/動画生成

光コヒーレンストモグラフィーにおけるクロスドメイン網膜層セグメンテーションのための空間正規化

光コヒーレンストモグラフィー (OCT) における網膜層のセグメンテーションは、網膜構造の定量的バイオマーカーを抽出するための基本的なステップです。実際、神経変性疾患に関連した OCT の分析に対する関心が高まっています。ただし、スペックル ノイズ、シャドウイング アーティファクト、隣接する層間の低コントラスト、被験者間の解剖学的ばらつき、取得プロトコルや臨床集団の違いから生じるドメインのシフトなどにより、セグメンテーションは依然として困難です。深層学習手法は目覚ましいパフォーマンスを達成しましたが、異種データセット全体にわたる堅牢性と一般化には依然として限界があります。この研究では、幾何学的ドメインのシフトを軽減し、網膜層セグメンテーションの一貫性を向上させるための前処理戦略としての空間正規化の役割を調査します。神経画像処理の標準的な手法にインスピレーションを得て、OCT ボリュームを共通の解剖学的基準に調整する中心窩中心の正規化フレームワークを導入します。私たちは、最先端の深層学習アーキテクチャの包括的な評価を実行します。セグメンテーション品質の包括的な評価を提供するために、B スキャン レベルの従来のオーバーラップ ベースのメトリクスと、A スキャン レベルのトポロジ認識メトリクスおよび正面レベルの厚さベースの測定を組み合わせます。グラウンド トゥルースが利用できない場合には、グラウンド トゥルースの注釈を必要としないトポロジー違反の定量的メトリクスと、正面レベルで構造の一貫性と臨床的に関連するパターンを捕捉する厚さに基づく定性的評価を提案します。この結果は、神経変性研究における信頼性の高いバイオマーカー抽出と下流の計算解析を可能にする、堅牢で臨床的に意味のある網膜解析ツールの開発に向けた、OCT セグメンテーション パイプラインにおける空間正規化の重要性を示しています。

原文 (English)

Spatial Normalization for Cross-Domain Retinal Layer Segmentation in Optical Coherence Tomography

Retinal layer segmentation in Optical Coherence Tomography (OCT) is a fundamental step for extracting quantitative biomarkers of retinal structure. Indeed, there is a growing interest in the analysis of OCTs in the context of neurodegenerative diseases. However, segmentation remains challenging due to speckle noise, shadowing artifacts, low contrast between adjacent layers, anatomical variability across subjects, and domain shifts arising from different acquisition protocols and clinical populations. While deep learning methods have achieved remarkable performance, their robustness and generalization across heterogeneous datasets remain limited. In this work, we investigate the role of spatial normalization as a preprocessing strategy to mitigate geometric domain shifts and improve the consistency of retinal layer segmentation. Inspired by standard practices in neuroimaging, we introduce a fovea-centered normalization framework that aligns OCT volumes into a common anatomical reference. We perform a comprehensive evaluation of state-of-the-art deep learning architectures. To provide a comprehensive assessment of segmentation quality, we combine conventional overlap-based metrics at B-scan level with topology-aware metrics at A-scan level and thickness-based measures at the en-face level. In cases where a ground truth is not available, we propose topology violation quantitative metrics that do not require ground truth annotations and a thickness-based qualitative assessment that captures structural consistency and clinically relevant patterns at the en-face level. The results demonstrate the importance of spatial normalization in OCT segmentation pipelines toward the development of robust and clinically meaningful retinal analysis tools, enabling reliable biomarker extraction and downstream computational analysis in neurodegenerative research.

13:00 JSTLLM/生成AIエージェント

5G/6G ネットワーク向け LLM を利用した Agentic AI: アーキテクチャ、プロトコル、標準化に関するチュートリアルと調査

大規模言語モデルによって実現されるエージェントティック人工知能 (AI) は、ルールベースの自動化から、次世代ネットワーク (NGN) の自律的で目標駆動型の制御への移行を示します。既存の調査では 2 つのドメインが分離して扱われており、プロトコルの統合、評価、標準化の調整は十分に調査されていません。このギャップに対処するために、2 部構成のチュートリアルとアンケートが提供されます。パート I では、5G および 6G の制御、管理、および AI ネイティブ プレーンを形式化します。次に、推論、計画、ツールの使用、マルチエージェントの調整、評価など、エージェント システムの基礎について説明します。パート II では、エージェント機能を 5G/6G コントロール サーフェス、標準化、および主要な 6G イニシアチブにマッピングします。最後に、自律通信を形成する未解決の課題を特定します。

原文 (English)

LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization

Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-driven control of Next-Generation Networks (NGNs). Existing surveys treat the two domains in isolation, leaving protocol integration, evaluation, and standardization alignment underexplored. To address this gap, a two-part tutorial-and-survey is presented. Part I formalises the control, management, and AI-native planes of 5G and 6G. It then covers the foundations of agentic systems: reasoning, planning, tool use, multi-agent coordination, and evaluation. Part II maps agentic capabilities onto 5G/6G control surfaces, standardization, and major 6G initiatives. Finally, it identifies open challenges shaping autonomous telecommunications.

13:00 JSTロボティクスClaude

JoyNexus: VLA モデルのサービス指向マルチテナントのポストトレーニング

シミュレーター、ロボットの実施形態、およびタスクの目的が多様であるため、視覚言語アクション (VLA) モデルのポストトレーニングが不可欠です。既存のコンピューティング サービスは、アクセラレータの直接レンタルとして提供されるか、バッチ ワークロードの送信として提供されるかに関係なく、通常、専用の GPU および CPU リソースのセットを単一のテナントに割り当てます。このパラダイムはクライアントの柔軟性を最大限に高めますが、インフラストラクチャの適応でユーザーに負担がかかり、固定カード時間会計モデルにより、短期間または集中的なワークロードがテナントにとって高価になり、サービス プロバイダーにとって非効率になります。これらの課題に対処するために、マルチテナント VLA 監視付き微調整、強化学習、評価のための統合サービスである JoyNexus を紹介します。 JoyNexus は、トレーニング モデル サービス、推論モデル サービス、環境サービスを分離し、それぞれ API を介してアクセスし、テナント固有のスロットを備えた常駐の共有ベース モデルによってサポートされます。テナントは、トレーニング、ロールアウト、評価のために高レベルのセマンティック API を直接呼び出すことも、下位レベルの API とそれに割り当てられたエンドポイントを使用してカスタム アルゴリズムを作成することもできます。複数のテナントがワークロードを同時に送信します。アクション モジュール、オプティマイザー、ロールアウト レコード、ポリシー バージョンは分離されたままで、サービスはグローバル トレーニング キューと推論キューによってスケジュールされます。マルチテナントのトレーニング効率をさらに向上させるために、JoyNexus は、互換性のあるモデルに面したプレフィックスを共有する異種 VLA データ スキーマのグループ バッチ処理を導入し、グループ化されたサンプルに対する単一の共有バックボーン フォワード パスを可能にします。最後に、現実的に具体化されたシナリオでのワークロード シミュレーションとグループ バッチ パイプラインを通じて JoyNexus を評価します。結果は、分離されたシングルテナント実行と比較して、JoyNexus は総 GPU 時間を削減し、共有リソースでのクロステナント スケジューリングを通じてサービス利用率を向上させることを示しています。

原文 (English)

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloads both expensive for tenants and inefficient for the service provider. To address these challenges, we present JoyNexus, a unified service for multi-tenant VLA supervised fine-tuning, reinforcement learning, and evaluation. JoyNexus decouples the Training Model Service, Inference Model Service, and Environment Service, each accessed through APIs and backed by resident shared base models with tenant-specific slots. Tenants can directly invoke high-level semantic APIs for training, rollout, and evaluation, or compose custom algorithms using lower-level APIs and their assigned endpoints. Multiple tenants submit workloads concurrently; their action modules, optimizers, rollout records, and policy versions remain isolated, and the service is scheduled by the global Training Queue and Inference Queue. To further improve multi-tenant training efficiency, JoyNexus introduces group batching for heterogeneous VLA data schemas that share a compatible model-facing prefix, enabling a single shared backbone forward pass over grouped samples. Finally, we evaluate JoyNexus through workload simulation and a group-batching pipeline in a realistic embodied scenario. Results show that, compared with isolated single-tenant execution, JoyNexus reduces aggregate GPU time and improves service utilization via cross-tenant scheduling on shared resources.

13:00 JSTLLM/生成AI画像/動画生成

HCIG: マルチモーダルな皮肉とネットいじめの検出のための階層型クロスモーダル違和感グラフ ネットワーク

意図された意味は、どちらかのモダリティだけからではなく、文字情報と視覚情報の間の不一致から現れることが多いため、多峰性の皮肉やネットいじめの検出は依然として困難です。既存のマルチモーダル アプローチは、主に特徴融合またはクロスモーダル アテンションに依存しており、異なる表現レベルにわたる階層的な意味の不一致を効果的に捕捉できない可能性があります。この制限に対処するために、この論文では、グラフ アテンション ネットワークを使用してトークン、フレーズ、およびグローバル レベルでクロスモーダル不整合をモデル化し、学習された階層的アテンション メカニズムを通じてこれらの表現を適応的に統合する新しいフレームワークである HCIG (階層型クロスモーダル不整合グラフ ネットワーク) を提案します。補完的なアーキテクチャとして、GCCN (グラフベースのクロスモーダル矛盾ネットワーク) も導入します。これは、効率的なマルチモーダル相互作用学習のために、矛盾を認識したプーリングを使用してグラフベースの推論を実行します。提案されたモデルは、MMSD 皮肉 ベンチマークと MultiBully ネットいじめデータセットで、包括的なアブレーション研究とクロスタスク移行実験とともに評価されます。実験結果は、HCIG が MMSD で 85.74% の精度と 85.29% のマクロ F1 で最高のパフォーマンスを達成するのに対し、GCCN は MultiBully で最高のマクロ F1 (68.66%) を達成し、HCIG は最高の精度 (69.62%) といじめクラス F1 (74.90%) を達成することを示しています。この調査結果は、階層的な多粒度の不一致モデリングが従来の融合戦略よりも効果的なマルチモーダル推論を提供し、ソーシャル メディアでの皮肉やネットいじめを検出するための堅牢なフレームワークを提供することを示しています。

原文 (English)

HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multimodal approaches primarily rely on feature fusion or cross-modal attention, which may not effectively capture hierarchical semantic inconsistencies across different levels of representation. To address this limitation, this paper proposes HCIG (Hierarchical Cross-modal Incongruity Graph Network), a novel framework that models cross-modal incongruity at token, phrase, and global levels using graph attention networks and adaptively integrates these representations through a learned hierarchical attention mechanism. As a complementary architecture, we also introduce GCCN (Graph-based Cross-modal Contradiction Network), which performs graph-based reasoning using contradiction-aware pooling for efficient multimodal interaction learning. The proposed models are evaluated on the MMSD sarcasm benchmark and the MultiBully cyberbullying dataset, together with comprehensive ablation studies and cross-task transfer experiments. Experimental results demonstrate that HCIG achieves the best performance on MMSD with 85.74% accuracy and 85.29% macro-F1, while GCCN attains the highest macro-F1 (68.66%) on MultiBully and HCIG achieves the highest accuracy (69.62%) and bullying-class F1 (74.90%). The findings demonstrate that hierarchical multi-granularity incongruity modeling provides more effective multimodal reasoning than conventional fusion strategies, offering a robust framework for sarcasm and cyberbullying detection in social media.

13:00 JST研究/論文

DADiff: 強化学習のための拡散主導のクロスドメイン ポリシー適応

ドメイン間でポリシーを転送することは、ソース ドメインとターゲット ドメイン間のダイナミクスの不一致により、強化学習において重大な課題を引き起こします。このペーパーでは、オンライン ダイナミクス適応の設定について検討します。この設定では、ポリシーが十分なデータを使用してソース ドメインでトレーニングされ、ターゲット ドメインとの限られた対話のみが許可されます。ドメイン分類器、値に基づくデータ フィルタリング、または表現学習を使用して、ダイナミクスの不一致に対処する既存の研究がいくつかあります。代わりに、生成モデリングの観点からドメイン適応問題を研究します。具体的には、次の状態の生成プロセスにおけるソースドメインとターゲットドメインの生成軌道間の不一致を利用してダイナミクスの不一致を推定する拡散ベースのフレームワークであるDADiffを導入します。報酬変更とデータ選択の両方のバリアントは、ポリシーをターゲット ドメインに適応させるために開発されています。また、2 つのドメイン間の特定のポリシーのパフォーマンスの差が生成軌道の偏差によって制限されることを示す理論分析も提供します。バリアントの適用可能性と、理論的分析と以前の研究との関係について、さらに詳しい議論が提供されます。私たちは、さまざまなシフトを伴う環境で広範な実験を実施し、手法の有効性を検証します。結果は、私たちの方法が既存のアプローチと比較して優れたパフォーマンスを提供し、ダイナミクスの不一致に効果的に対処することを示しています。メソッドのコードは https://github.com/hanyang-chen/DADiff-release で提供しています。

原文 (English)

DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consider the setting of online dynamics adaptation, where policies are trained in the source domain with sufficient data, while only limited interactions with the target domain are allowed. There are a few existing works that address the dynamics mismatch by employing domain classifiers, value-guided data filtering, or representation learning. Instead, we study the domain adaptation problem from a generative modeling perspective. Specifically, we introduce DADiff, a diffusion-based framework that leverages the discrepancy between source and target domain generative trajectories in the generation process of the next state to estimate the dynamics mismatch. Both reward modification and data selection variants are developed to adapt the policy to the target domain. We also provide a theoretical analysis to show that the performance difference of a given policy between the two domains is bounded by the generative trajectory deviation. More discussions on the applicability of the variants and the connection between our theoretical analysis and the prior work are further provided. We conduct extensive experiments in environments with various shifts to validate the effectiveness of our method. The results demonstrate that our method provides superior performance compared to existing approaches, effectively addressing the dynamics mismatch. We provide the code of our method at https://github.com/hanyang-chen/DADiff-release

13:00 JSTLLM/生成AI

トレーニング前からトレーニング後まで推論を理解する

強化学習 (RL) は、複雑な推論タスクにおける大規模言語モデル (LLM) を改善する上で中心的な役割を果たしていますが、RL のポストトレーニングは、その前の事前トレーニングとは切り離して研究されることがほとんどです。その結果、2 つの基本的な疑問が未解決のままです。(1) 事前トレーニングの選択 (モデル サイズ、データ) は RL 計算への戻りをどのように形成するのか、(2) RL はモデルに対して実際に何を行うのか?これらの質問は、標準的な LLM 設定では検討するのが困難です。事前トレーニング コーパスは膨大で制御されていないため、動作を事前トレーニングと RL に帰属させるのが難しく、両方の段階にわたる体系的なコンピューティング スイープには法外なコストがかかります。これらの課題に対処するために、トレーニング前からトレーニング後のパイプライン全体にわたって推論を研究するための制御されたテストベッドとしてチェスを使用します。私たちは、人間のチェス ゲームで 5M から 1B パラメーターまでの言語モデルを事前トレーニングし、合成推論トレースで教師付き微調整を行い、検証可能な報酬を備えたチェス パズルで RL を実行することにより、標準的な LLM トレーニング パイプラインに従います。このフレームワークを使用すると、特定の RL コンピューティング レベルでの RL 後のパフォーマンスが事前トレーニング損失から適切に予測され、RL 報酬曲線の傾きが事前トレーニング トークンとともにほぼ線形に改善することがわかります。スケーリング以外にも、RL は単に SFT ポリシーを強化するだけではないことがわかりました。簡単なパズルでは、SFT ポリシーがすでに好んでいた正しい手を増幅させますが、難しいパズルでは、SFT ではほとんど存在しなかった正しい手を表面化します。さらに、数学ドメインのテキストで 1B 言語モデルをトレーニングすることによって、発見がチェスの枠を超えて応用できるかどうかをテストします。そこでは同じ予測パターンが現れます。つまり、事前トレーニングされたチェックポイントの時間が長くなると、RL 後のパフォーマンスが向上し、RL の下でより速く改善されます。要約すると、トレーニング前からトレーニング後までのパイプライン全体にわたって推論の科学を研究するための、トレーニング前から RL へのインターフェイスの定量的な説明と制御されたテストベッドを提供します。

原文 (English)

Understanding Reasoning from Pretraining to Post-Training

Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across both stages are prohibitively expensive. To address these challenges, we use chess as a controlled testbed for studying reasoning across the full pretraining-to-post-training pipeline. We follow the standard LLM training pipeline by pretraining language models from 5M to 1B parameters on human chess games, supervised fine-tuning on synthetic reasoning traces, and running RL on chess puzzles with verifiable rewards. Using this framework, we find that the post-RL performance at given RL compute level is well-predicted from the pretraining loss, and slope of the RL reward curves improves approximately linearly with the pretraining tokens. Beyond scaling, we find that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves the SFT policy already preferred, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. We further test whether our findings transfer beyond chess by training a 1B language model on math-domain text, where the same predictive pattern emerges: longer-pretrained checkpoints reach higher post-RL performance and improve faster under RL. In sum, we provide a quantitative account of the pretraining-to-RL interface and a controlled testbed for studying the science of reasoning across the full pretraining-to-post-training pipeline.

13:00 JST研究/論文

AI ライフサイクル ガバナンスにおける監査可能な信頼性レベルの方法論

AI ガバナンスでは、AI システムが長期間にわたって十分な信頼性を維持できるかどうか、観察された変化が許容できるかどうか、またそのような判断を透明かつ異議の余地のない方法で文書化する方法についての判断がますます求められています。しかし、AI の信頼性に関する既存の研究は、ライフサイクルの監視と再評価をサポートするにはレベルが高すぎるか、ガバナンスのニーズに結び付けるには指標主導が狭すぎるかのどちらかです。したがって、AI ガバナンスにおける監査可能な信頼性レベルのための軽量な方法論を提案します。この方法論には 2 つのコンポーネントがあります。信頼性レベルを表現および学習するための正式なフレームワークと、信頼性レベルを文書化し、監視し、長期的に再評価するための軽量の AI ライフサイクル ガバナンス手順です。正式なフレームワークは、測定可能な次元のコンテキスト依存プロトコルを通じてガバナンスに関する相対的な信頼性をモデル化し、信頼性プロファイルに対する解釈可能なルールとして信頼性レベルを学習します。この方法では、解釈可能な概念実証モデル クラスとしてデシジョン ツリーを使用し、明示的な信頼性プラトー、読み取り可能なレベル遷移、および 2 つの単純なライフサイクル診断 (境界マージンとプロファイル ドリフト) が得られます。ガバナンス手順では、設計時のラベル付け、導入後のモニタリング、再評価、およびレポート作成のための適合性指向のワークフローに、これらの正式なオブジェクトが組み込まれています。また、プロトコルの設計、検証、監視、再評価に対する人間の責任と制御ゲートも割り当てられます。劣化、ショック、更新、異種モニタリングの頻度、システム比較を含む合成 AI ライフサイクル トレースの方法論を示します。当社の方法論は法的判断やその他の専門家の判断に代わるものではありません。AI ガバナンス関連の変更を時間の経過とともに文書化および追跡するための証拠基盤を提供することで、適合性文書化とライフサイクル監視をサポートします。

原文 (English)

A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance

AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, whether observed changes are tolerable, and how such judgments should be documented in a transparent and contestable way. Yet existing work on AI trustworthiness remains either too high-level to support lifecycle monitoring and reassessment or too narrowly metric-driven to connect with governance needs. We therefore propose a lightweight methodology for auditable trustworthiness levels in AI governance. The methodology has two components: a formal framework for representing and learning trustworthiness levels, and a lightweight AI lifecycle governance procedure for documenting, monitoring, and reassessing them over time. The formal framework models governance-relative trustworthiness through a context-sensitive protocol of measurable dimensions and learns trustworthiness levels as interpretable rules over trustworthiness profiles. Using decision trees as an interpretable proof-of-concept model class, the methodology yields explicit trustworthiness plateaus, readable level transitions, and two simple lifecycle diagnostics: boundary margins and profile drift. The governance procedure embeds these formal objects in a conformity-oriented workflow for design-time labeling, post-deployment monitoring, reassessment, and reporting. It also assigns human responsibilities and control gates for protocol design, validation, monitoring, and reassessment. We illustrate the methodology on synthetic AI lifecycle traces involving degradation, shocks, updates, heterogeneous monitoring cadences, and system comparison. Our methodology does not replace legal or other expert judgment: it supports conformity documentation and lifecycle monitoring by providing an evidential basis for documenting and tracking AI governance-relevant changes over time.

13:00 JSTLLM/生成AI研究/論文GemmaQwen

ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, i…

13:00 JSTLLM/生成AIエージェント

When Do Multi-Agent Systems Help? An Information Bottleneck Perspective

LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent s…

13:00 JSTLLM/生成AI画像/動画生成ClaudeGPT / ChatGPT

An Exam for Active Observers

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychop…

13:00 JSTエージェント

When Does Muon Help Agentic Reinforcement Learning?

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We…

13:00 JSTLLM/生成AIエージェントGemma

Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities

Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, Electronic Control Uni…

13:00 JST研究/論文

RAD: Retrieval High-quality Demonstrations to Enhance Decision-making

Offline reinforcement learning (RL) learns policies from fixed datasets, thereby avoiding costly or unsafe environment interactions. Howeve…

13:00 JST研究/論文

A Neuro-Symbolic Approach for Probabilistic Reasoning on Graph Data

Graph neural networks (GNNs) excel at predictive tasks on graph-structured data but often lack the ability to incorporate symbolic domain k…

13:00 JST研究/論文

Human-Aligned Procedural Level Generation Reinforcement Learning via Text-Level-Sketch Shared Representation

Human-aligned AI is a critical component of co-creativity, as it enables models to accurately interpret human intent and generate controlla…

13:00 JSTLLM/生成AIハードウェア/半導体Qwen

RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs

The Structure Gap between probabilistic LLM generation and deterministic schema requirements hinders automated workflows. We propose RL-Str…

13:00 JSTLLM/生成AIビジネス/資金調達

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy

The black-box nature of Large Language Models necessitates novel evaluation frameworks that transcend surface-level performance metrics. Th…

13:00 JSTLLM/生成AI

The AI Fiction Paradox

AI development has a fiction dependency problem. Developers have treated large corpora of modern books, including fiction, as valuable enou…

13:00 JSTLLM/生成AIエージェント研究/論文

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientifi…

13:00 JST研究/論文

FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment

In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring…

13:00 JST研究/論文

Towards a General Intelligence and Interface for Wearable Health Data

While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into…

13:00 JSTエージェントビジネス/資金調達

Workflow-GYM: 現実世界の専門分野におけるコンピュータ使用エージェントタスクの長期的な評価に向けて

近年、ますます複雑になる現実世界のタスクの処理に向けて、AI エージェントが急速に進化しています。しかし、既存のベンチマークでは、エージェントがグラフィカル ユーザー インターフェイスを操作して、さまざまなドメインにわたる長期にわたる価値の高い専門的なワークフローを完了できるかどうかを評価することはほとんどありません。現在の GUI ベンチマークは依然として、主に汎用ソフトウェア、比較的単純なアプリケーション、および短期間のタスクに焦点を当てており、最新のエージェントがユーザーの指示に従ってドメイン固有のプロフェッショナル ソフトウェアを自律的に操作し、経済的に価値のある作業をエンドツーエンドで実行できるかどうかはほとんど不明です。このギャップを埋めるために、専門分野と特殊なソフトウェア環境を中心とした長期的な GUI タスクのベンチマークである Workflow-GYM を導入します。最先端のモデルで広範な実験を行った結果、最も強力なモデルでも成功率は 30% をわずかに超える程度であることがわかり、プロの長期にわたる GUI ワークフローが現在の GUI エージェントにとって依然として非常に困難であることが浮き彫りになりました。さらなる分析により、現在のエージェントは長期的なワークフローの一貫性を維持するのに苦労しており、ワークフロー段階の省略、エラーの伝播、目標のずれ、プロフェッショナルなソフトウェア環境の理解不足が頻繁に見られることが明らかになりました。私たちの調査結果は、現在のエージェント システムの限界についての重要な洞察を提供し、次世代の GUI エージェント研究の重要な方向性を示唆しています。

原文 (English)

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.

13:00 JSTLLM/生成AIエージェント研究/論文

Agents-K1: エージェントネイティブのナレッジオーケストレーションに向けて

現在の LLM ベースの研究エージェントは、エージェント オーケストレーションを通じて進歩していますが、科学的知識のオーケストレーションはほとんど見落とされています。既存の著作物は、論文を要約、表面的な言及、平坦な \texttt{引用} エッジに縮小することが多く、科学的推論に不可欠な重要な実体、主張、証拠、メカニズム、および方法系統を省略しています。この目的を達成するために、生のドキュメントをエージェントネイティブの科学知識グラフに変換するエンドツーエンドの知識オーケストレーション パイプラインである \textbf{Agents-K1} を導入します。 Agents-K1 は、統一された理論的基盤の下に 3 つのコンポーネントを統合します。マルチモーダル パーサーでは、5 つのモジュールのスキーマがエンティティ、マルチモーダルな証拠、引用、および抄録のみではなく論文全体にわたる型付けされたエンティティ間の関係をキャプチャします。ルールベースの報酬の下で GRPO でトレーニングされた 4B 情報抽出バックボーン。もう 1 つは、Web 検索、マルチモーダル グラフ検索、およびクロスドキュメント トラバーサルを統合する 3 つのソース エージェント インターフェイスである、graphanything CLI です。これに加えて、6 つの主題にわたる 246 万件の科学論文を処理して \textbf{Scholar-KG} を生成し、そのうち 100 万件の論文サブセットをリリースしています。完全な Scholar-KG は以下の SCP リンクからアクセスできます。同じパイプラインを一般ドメインのコーパスとスキーマ準拠のデータ合成に拡張できます。広範な実験により、Agents-K1 が科学情報の抽出、ナレッジ グラフの構築、およびマルチホップ科学的推論において優れたパフォーマンスを達成することが実証されました。

原文 (English)

Agents-K1: Towards Agent-native Knowledge Orchestration

Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Existing works often reduce papers to abstracts, surface mentions, and flat \texttt{cites} edges, omitting key entities, claims, evidence, mechanisms, and method lineages essential for scientific reasoning. To this end, we introduce \textbf{Agents-K1}, an end-to-end knowledge orchestration pipeline that converts raw documents into agent-native scientific knowledge graphs. Agents-K1 integrates three components under a unifying theoretical foundation: a multimodal parser whose five-module schema captures entities, multimodal evidence, citations, and typed inter-entity relations across the full paper rather than abstracts alone; a 4B information-extraction backbone trained with GRPO under a rule-based reward; and a graphanything CLI, a tri-source agent interface that unifies web search, multimodal graph retrieval, and cross-document traversal. On top of this, we process 2.46 million scientific papers across six subjects to produce \textbf{Scholar-KG}, of which we release a one-million-paper subset, and the full Scholar-KG is accessible via the SCP link below. The same pipeline can be extended to general-domain corpora and to schema-conformant data synthesis. Extensive experiments demonstrate that Agents-K1 achieves superior performance in scientific information extraction, knowledge graph construction, and multi-hop scientific reasoning.

13:00 JST研究/論文

GA-VINO: A Geometry-Aware Variational Physics-informed Neural Operator for Mindlin-Reissner Plates

Plate and shell structures are widely used in engineering fields. Rapid response prediction for such structures under complex geometries, h…

13:00 JSTLLM/生成AIエージェント

UCOB: クレジットを意識したポリシーに基づく双方向自己蒸留によるエージェント スキルの活用と進化の学習

スキル記憶は過去の経験をテキストガイダンスとして再利用することでエージェント強化学習を改善できますが、取得されたスキルは神託ではありません。ある状態では役立つ一方で、別の状態では同じポリシーを誤解させる可能性があります。これにより、一般的な特権教師の仮定が脆弱になります。つまり、スキル条件付きプロンプトは、スキルなしプロンプトに対する固定教師として扱うことができるということです。クレジットを意識したポリシーに基づく双方向の自己蒸留を通じて、エージェント スキルの活用と進化を学習するためのフレームワークである UCOB を紹介します。 UCOB は、スキル条件付きプロンプトとスキルなしプロンプトを同じモデルの 2 つのオンポリシー コンテキスト ビューとして扱い、同じタスクおよびアンカー状態内でのリターンツーゴーを比較し、より高いリターンのビューをローカル教師として使用します。このローカル クレジット シグナルは、有用なスキル条件付き動作を内部化し、誤解を招くスキルの使用法を修正し、タスク/状態のスキル メモリの更新、ユーティリティを意識した検索、およびリフレクション自己トレーニングをガイドします。 ALFWorld、WebShop、Search-QA などのエージェント タスクの実験では、モデル スケール全体で UCOB がスキルフリー RL、スキルメモリ ベースライン、自己蒸留法よりも優れており、ALFWorld と WebShop で SOTA ベースラインよりも最大 23.5 ポイントおよび 18.0 ポイント向上していることが示されています。アブレーションと分析により、その中核となるメカニズムと効率がさらに検証されます。

原文 (English)

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation

Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not oracular: they may help in one state while misleading the same policy in another. This makes the common privileged-teacher assumption fragile, namely that a skill-conditioned prompt can be treated as a fixed teacher for the no-skill prompt. We introduce UCOB, a framework for learning to utilize and evolve agentic skills via credit-aware on-policy bidirectional self-distillation. UCOB treats skill-conditioned and no-skill prompts as two on-policy context views of the same model, compares their return-to-go within the same task and anchor state, and uses the higher-return view as the local teacher. This local credit signal internalizes useful skill-conditioned behavior, corrects misleading skill usage, and guides task/state skill memory updates, utility-aware retrieval, and reflection self-training. Experiments on agentic tasks, including ALFWorld, WebShop, and Search-QA, show that UCOB outperforms skill-free RL, skill-memory baselines, and self-distillation methods across model scales, with up to 23.5 and 18.0 point gains over SOTA baselines on ALFWorld and WebShop. Ablations and analyses further validate its core mechanisms, continual adaptation across environments, and modest training overhead. Code is available at https://github.com/TU2021/UCOB.

13:00 JSTエージェント

ACPO: マルチエージェント強化学習のためのエージェント連鎖ポリシーの最適化

マルチエージェント強化学習 (MARL) における協調タスクでは、エージェントが共同して共有利益を最大化する必要があります。集中トレーニングと分散実行 (CTDE) パラダイムの下では、ポリシー勾配を直接計算することは依然として困難です。従来の手法は主に 2 つのアプローチに従います。1 つは集中的な批評家による独立した因数分解更新であり、値分解の仮定を持たない一般的な共同改善の保証が欠けています。もう 1 つは、次善のナッシュ均衡に収束する可能性がある交互の最良応答更新です。この論文では、共同ポリシー勾配が、エージェントごとのスコア関数と分散型批評家から形成されるエージェントごとの用語の正確な分散型分解を可能にすることを示します。この分解に基づいて、私たちはエージェント連鎖ポリシー最適化 (ACPO) を開発します。ACPO では、アクターが個別にトレーニングされ、更新内容が共同ポリシー勾配の 1 つのステップを構成します。この結果の中心となるのは、エージェントが一度に 1 つずつアクションを実行し、各エージェントが先行するアクションに対する信念に基づいてアクションを実行する、同時共同決定のシリアル化されたビューです。この信念は、エージェントごとの独立した更新を共同勾配ステップに結び付ける調整メカニズムとして機能します。マルチロボット ウェアハウス、SMACv2、および MA-MuJoCo で ACPO を評価しました。ACPO は強力なベースラインを上回り、エージェントの数が増加するにつれてギャップが拡大しました。

原文 (English)

ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning

Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution (CTDE) paradigm, policy gradients have remained difficult to compute directly. Prior methods largely follow two approaches: independent factorized updates with centralized critics, which lack general joint-improvement guarantees without value decomposition assumptions, or alternating best-response updates, which can converge to suboptimal Nash Equilibria. In this paper, we show the joint policy gradient admits an exact decentralized decomposition of per-agent terms, each formed from per-agent score functions and decentralized critics. Based on this decomposition, we develop Agent-Chained Policy Optimization (ACPO), where actors are trained independently, with their updates together constituting a single step on the joint policy gradient. Central to this result is a serialized view of the simultaneous joint decision in which agents commit actions one at a time, each conditioning on a belief over preceding actions that ties the independent per-agent updates into a single joint step. We evaluate on-policy and off-policy instantiations of ACPO on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo, where it outperforms strong baselines, with the gap widening as the number of agents grows.

13:00 JSTエージェント研究/論文

MirrorCode: AI は動作だけからプログラム全体を再構築できる

ベンチマークの進捗状況や、AI が C コンパイラーを実装するなどの 1 回限りのデモンストレーションで示されているように、AI モデルは自律コーディングにおいて急速に向上しています。ただし、既存のコーディング ベンチマークは短いタスクに焦点を当てる傾向があり、1 回限りのデモンストレーションは人間によるガイダンスが含まれることが多く、標準化されておらず、モデル間で繰り返されていないため、体系的に比較するのが困難です。これらの課題に対処するために、ソフトウェア プロジェクト全体の再実装に基づく長期的なコーディング ベンチマークである MirrorCode を導入します。 MirrorCode では、AI エージェントはソース コードにアクセスせずに、既存のプログラムの機能を複製する必要があります。 AI ソリューションは、ホールドアウト テストを含むエンドツーエンド テストで元のプログラムの出力と正確に一致する必要があります。 MirrorCode の 25 のターゲット プログラムは、Unix ユーティリティ、データのシリアル化とクエリ ツール、バイオインフォマティクス、インタプリタ、静的分析、暗号化、圧縮など、コンピューティングのさまざまな分野にまたがっています。既存の AI モデルはすでに複雑なソフトウェアを再実装でき、最も強力なモデルはベンチマーク全体で 56% のスコアを獲得しています。たとえば、AI は 16,000 行のバイオインフォマティクス ツールキットである gotree を再実装できますが、この作業は人間のエンジニアでは数週間かかると考えられています。ただし、パフォーマンスの最前線を研究するには、一般的なベンチマークよりも大きな推論予算が必要です。たとえば、大規模なタスクの 1 回の試行には 19 日間で 2,600 ドルかかります。特に要件が正確に指定されている場合、AI エージェントは長期的なソフトウェア エンジニアリング タスクをすでに完了できることを示します。より広く言えば、私たちの研究は、自律エージェントが改善を続けるにつれて、AI がソフトウェア エンジニアリングに変革的な影響を与えることを示唆しています。

原文 (English)

MirrorCode: AI can rebuild entire programs from behavior alone

AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. However, existing coding benchmarks tend to focus on shorter tasks, and one-off demonstrations are hard to compare systematically because they often have some human guidance, and are not standardized or repeated across models. To address these challenges, we introduce MirrorCode, a long-horizon coding benchmark based on reimplementing entire software projects. In MirrorCode, AI agents must replicate the functionalities of an existing program, without access to its source code. AI solutions must match the original program's output exactly on end-to-end tests, including held-out tests. MirrorCode's 25 target programs span different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression. Existing AI models can already reimplement complex software, with the strongest model scoring 56% across the benchmark. For example, AI can reimplement gotree, a 16,000-line bioinformatics toolkit - a task that we believe would take weeks for a human engineer. However, studying the frontier of performance requires a larger inference budget than typical benchmarks, for example, \$2,600 over 19 days for a single attempt on a large task. We show that AI agents can already complete long-horizon software engineering tasks, especially when requirements are precisely specified. More broadly, our work suggests AI will have transformative effects on software engineering, as autonomous agents continue to improve.

13:00 JST研究/論文

内部多元性とペアごとの比較の限界

ローカルペア比較は、参加型デザインや調整などにおいて、人々が意思決定ルールがどのように機能することを望んでいるのかを学習するための標準ツールです。ただし、それらの使用には 2 つの強力な前提が構築されます。1 つは、ローカル比較は、人が自動化された意思決定ルールにどのように動作してほしいかを示す十分な証拠であるということ、もう 1 つは、人々は常にそれらの比較に決定的に答えることができるということです。私たちは、内部多元主義、つまりルールがどのように動作するかについて複数の権威ある優先順位に従って個人が意思決定ルールを評価するという考えの下で、これらの仮定がどのように損なわれるかを調査します。我々は、決定ルールに対するそのような多元的な優先順位の正式なモデルを提供します。これにより、強制的なローカルペア比較データの 2 つの異なる失敗を識別できるようになります。まず、比例性、平等主義、平等な扱いなどの優先順位は本質的にグローバルなものです。あるケースでそれらが意味することは、他の場所で何が起こるかによって異なるため、局所的な比較ではそれらを把握できない可能性があります。第 2 に、優先順位が局所的に表現可能な場合でも、強く保持されている優先順位間の緊張によって内部対立が生じ、比較が強制されると、潜在的にコストのかかる行動の歪みが生じる可能性があります。次に、モデルを使用して、人々が優柔不断を報告できるようにするという代替案を調査しました。その結果、そうすることで、好みを正確に学習するために必要なクエリの数を大幅に削減できることがわかりました。最後に、私たちのモデルが、これらの優先順位を直接導き出し、人々が何に価値を置くかについて、より忠実で解釈可能な説明を生み出す優先学習方法をどのように指し示しているかを説明します。

原文 (English)

Internal Pluralism and the Limits of Pairwise Comparisons

Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.g., in participatory design or alignment. However, their use builds in two strong assumptions: that local comparisons are sufficient evidence about how a person wants an automated decision rule to behave, and that people can always answer those comparisons decisively. We investigate how these assumptions may be compromised under internal pluralism: the idea that an individual evaluates decision rules according to multiple authoritative priorities about how the rule should behave. We provide a formal model of such pluralistic preferences over decision rules, which then lets us identify two distinct failures of forced local pairwise comparison data. First, priorities such as proportionality, egalitarianism, and equal treatment are inherently global: what they imply in one case can depend on what happens elsewhere, so local comparisons may fail to capture them. Second, even when priorities are representable locally, tension between strongly-held priorities can generate internal conflict, producing potentially costly behavioral distortions when comparisons are forced. We then use our model to investigate the alternative -- allowing people to report indecision -- and our findings suggest that doing so can considerably reduce the number of queries needed to learn preferences accurately. We conclude by describing how our model points toward preference-learning methods that elicit these priorities directly, yielding more faithful and interpretable accounts of what people value.

13:00 JSTエージェントビジネス/資金調達DeepSeek

エージェント ステップ値: 状態接地 LLM エバリュエーターによる状態遷移測定

ほとんどのエージェント評価では、複数ステップのトレースが最終的な回答、成功フラグ、または軌跡レベルのスコアにまとめられます。これらの集計では、開発者が最も必要とする診断の質問、つまりどのアクションが状態を有益な方向に変更したのかがわかりにくくなります。我々は、状態遷移測定フレームワークであるエージェント ステップ値 (ASV) を導入します。これは、観察された各アクションを、固定された候補結果に対する状態に基づいた評価者の分布に誘発する変化によってスコア付けします。 ASV は、編集された前後の状態予測をレンダリングし、ステートレス LLM エバリュエーターを使用して候補ログ スコアを割り当て、ゴールドフリーの信念診断とオフライン オラクル検証メトリクスの両方をレポートします。ラベルフリーの理論的パスにより、評価者の審議がワントークンオプションのスコアリングから分離され、リークやフロアスコアイベントを明らかにしながら候補の可能性が維持されます。ライブ PubMed 検索、部分的にライブ DeepSeek アクター、および DeepSeek 対数確率スコアリングを使用した 100 件のレビュー済みオープン QA 証拠探索タスクについて、ASV は 1,100 のステップと 2,200 の状態を評価します。固定レイアウトの根拠条件付きプロトコルの下では、平均ゴールドマージンゲインは -2.335 (軌道ブートストラップ 95\% CI [-3.395, -1.272])、エントロピーの動きは 0.000、平均ベイジアンサプライズは 2.693 です。したがって、ASV は、最終回答スコアとエントロピーのみのステップ メトリクスが見逃している建設的および破壊的な信念ピボットを特定します。スタンドアロンの ASV Eval ツールキットをリリースします。

原文 (English)

Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces

Pooling, substituting, or reusing evaluator-derived step rewards assumes that their direction survives a change of evaluation channel. The same frozen transition can violate that assumption. Process rewards vary agent states, while evaluator audits vary scoring configurations; neither first difference isolates their interaction. We define Agent Step Value (ASV) as a channel-indexed target-margin gain and identify the state-by-channel interaction on complete matched faces. Across frozen PubMed question-answering transitions, direct scoring yields a positive mean ASV, while the generated-view channel yields a negative mean. Two matched replay waves reproduce this reversal, and cross-channel sign disagreement exceeds same-channel retry disagreement by 48.0 percentage points. Matched retrieval faces localize the reversal to the generated-view coordinate and trace its direction across a readout-and-stack bridge. A source-only generation contract restores the positive mean direction on artifact-bearing retrievals and removes parser-detected substantive support claims from artifact-free before-state views. ASV turns channel sensitivity into an identified measurement problem that can be localized and tested by intervention before step rewards are reused.

13:00 JST研究/論文GPT / ChatGPTGemini

OmniFood-Bench: 栄養素の推論と個別の健康アドバイスのための VLM の評価

Large Vision-Language Model (VLM) を重要なインフラストラクチャに迅速に統合することで、個別化されたヘルスケアと食事管理に革命をもたらすことが期待されます。しかし、食品システムの領域では、自律エージェントは、見た目と本質的な栄養成分の間の「全身情報の非対称性」という、独特かつ永続的な課題に直面しています。既存のベンチマークは主に、食品カテゴリーの認識などの粗粒度の分類タスクに焦点を当てており、実際の食事管理に必要な複雑な推論チェーン、具体的には、隠れた原材料の特定から物理量の推定、そして最終的には安全性が重要な医学的アドバイスの総合までを横断する能力を評価できません。このペーパーでは、MM-Food-100K データセットから構築された包括的なベンチマークである OmniFood-Bench を紹介します。以前の研究とは異なり、OmniFood-Bench は、基本的な認識 (食材と調理方法)、定量的推論 (分量と栄養プロファイリング)、および安全性重視の勧告 (疾患固有の推奨事項) の 3 つの進歩的な機能にわたって VLM を評価します。 gpt-5.1、gemini-3-flash、qwen3-vl-8B を含む 6 つの最先端の VLM を評価します。私たちの広範な実験により、驚くべき「意味と物理のギャップ」が明らかになりました。モデルは、料理の命名において人間に近い精度を達成する一方で、質量推定において壊滅的な失敗を示し、高リスクの糖尿病プロファイルに対する良性のアドバイスを頻繁に幻覚で示します。この取り組みにより、公衆衛生のために配備された自律エージェントの信頼性に関する厳格な基準が確立されました。コードとデータセットは、https://anonymous.4open.science/r/OmniFood-Bench-7D0B で入手できます。

原文 (English)

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management. However, in the domain of food systems, autonomous agents face a unique and persistent challenge: the "Systemic Information Asymmetry" between visual appearance and intrinsic nutritional composition. Existing benchmarks primarily focus on coarse-grained classification tasks, such as food category recognition, which fail to evaluate the intricate reasoning chain required for real-world dietary management -- specifically, the ability to traverse from identifying hidden ingredients to estimating physical mass, and finally synthesizing safety-critical medical advice. In this paper, we introduce OmniFood-Bench, a comprehensive benchmark constructed from the MM-Food-100K dataset. Unlike previous works, OmniFood-Bench evaluates VLMs across three progressive capabilities: Basic Perception (Ingredients & Cooking Methods), Quantitative Reasoning (Portion Size & Nutritional Profiling), and Safety-Critical Advisory (Disease-Specific Recommendations). We evaluate six state-of-the-art VLMs, including gpt-5.1, gemini-3-flash, and qwen3-vl-8B. Our extensive experiments reveal a startling "Semantic-Physical Gap": while models achieve near-human accuracy in naming dishes, they exhibit catastrophic failure in mass estimation and frequently hallucinate benign advice for high-risk diabetic profiles. This work establishes a rigorous standard for trustworthiness in autonomous agents deployed for public health. The code and datasets are available in: https://anonymous.4open.science/r/OmniFood-Bench-7D0B

13:00 JSTLLM/生成AI

Voltzmann MapReduce: フォーク可能なサンドボックスのパーティション関数 Reduce

局所漸近正規性 (LAN) の下で、ワーカーがサイズ $n$ のチャンクに対して発する信頼密度は、ギブス-ボルツマン測度 $\exp\{-\beta E(\theta)\}$ であり、その逆温度はサンプル サイズ $\beta=n$ です。ガウス/線形の場合は 3 つの結果が正確で、それ以外の場合は 1 次です。つまり、互いに素なチャンクは独立したボルツマン因子を持ちます。そのため、MapReduce \emph{reduce} は、文字通り読むと、モードが精度重み付け (逆分散) プーリングである分割関数 $Z=\int\prod_k h_k\,d\theta$ になります。頻度主義的整合性はゼロ温度限界 $T=1/n\to0$ です

原文 (English)

Evidence-Aware MapReduce for Forkable Compute

Snapshot-backed sandboxes make branching cheap while leaving evidence dependence unchanged. Branches can reuse a model, prompt, repository, tests, observations, or execution ancestor, so counting outputs can amplify one repeated error into high-confidence consensus. We introduce an \emph{evidence-aware reduction contract}: each worker reports an estimate, estimated information, evidence identifiers, fork lineage, and execution metadata. For independent workers estimating one common parameter, we use standard inverse-information pooling in its Gaussian/Wald form. The fixed-dimensional numeric summary can merge in any tree order; evidence IDs and lineage follow separate rules. The residual $\Delta$ measures disagreement, becomes Cochran's $Q$ in the scalar inverse-variance case, and appears in the product integral. A reference implementation validates serialized records, rejects repeated nonempty evidence identifiers, carries evidence and lineage through tree reduction, and uses Cholesky-based numerical linear algebra. Unit tests and seeded synthetic checks exercise the algebra, unequal information, and forged precision; one four-worker named-snapshot trace exercises the end-to-end path. Platform logs document the exercised execution paths. A central open systems challenge is to turn evidence identity and fork lineage into a dependence model for correlated and adaptively selected AI branches.

13:00 JSTLLM/生成AI

長さのペナルティにより思考連鎖が監視されにくくなる

長さにペナルティを課した強化学習は、モデルの答えを駆動する影響を隠しながら、思考連鎖推論を短縮することができます。私たちの実験では、モデルの思考連鎖がヒントに言及する頻度ははるかに低かったにもかかわらず、長さペナルティを使用したトレーニングによって、モデルの操作からの誤解を招くヒントが阻止されませんでした。トークン精度の評価では、これらの実行は成功したものとしてカウントされます。これは、使用する推論トークンが少なく、精度の低下もほとんどないためです。残りの痕跡が、答えを導き出したものをまだ示しているかどうかを見逃してしまいます。異なるターゲット鎖長で Qwen3-4B および Qwen3-14B バリアントをトレーニングし、次に、保持された MMLU-Pro-R と 4 つの転送ベンチマークに対するバイアス ヒント介入でそれらを評価します。圧縮により推論トークンが大幅にカットされ、多肢選択の精度がほとんど維持され、ヒントの影響がベースライン付近に残ります。最も強力なターゲットでは、下限の忠実度は Qwen3-14B ではベースラインの 63.1%、Qwen3-4B では 69.4% に低下します。モニターがヒント使用を捕捉する生の率は 69% から 49% に、そして 60% から 48% に低下します。長さをコンテンツから分離するために、残りのテキストが圧縮された長さと一致するまで、非圧縮のベースライン チェーンから文をランダムに削除します。この長さのマッチングの後でも、圧縮されたチェーンは、Qwen3 サイズと 5 つの評価分布すべてについて、ランダムに短縮したベースライン チェーンよりもヒントを開示する頻度が 7 ~ 35 パーセント ポイント低くなります。したがって、圧縮は推論を短縮するだけではなく、モニターが答えに影響を与えたものを確認するために必要な手がかりを優先的に削除します。これらの結果を総合すると、より安価な推論によって答えを保存しながら、その背後にある影響を検出することが困難になる、圧縮監視可能性のフロンティアが明らかになります。

原文 (English)

Length Penalties Make Chain-of-Thought Less Monitorable

Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering models, even though the models' chains of thought mention the hint much less often. A token-accuracy evaluation would count these runs as successful because they use fewer reasoning tokens with little accuracy loss; it would miss whether the remaining trace still shows what drove the answer. We train Qwen3-4B and Qwen3-14B variants with different target chain lengths, then evaluate them with biasing-hint interventions on held-out MMLU-Pro-R and four transfer benchmarks. Compression sharply cuts reasoning tokens, preserves most multiple-choice accuracy, and leaves hint influence near baseline. At the strongest target, lower-bound faithfulness falls to 63.1% of baseline for Qwen3-14B and 69.4% for Qwen3-4B; the raw rate at which a monitor catches hint use falls from 69% to 49% and from 60% to 48%. To separate length from content, we randomly delete sentences from uncompressed baseline chains until the remaining text matches the compressed length. Even after this length matching, compressed chains disclose the hint 7-35 percentage points less often than baseline chains that we shorten at random, for both Qwen3 sizes and all five evaluation distributions. Compression therefore does more than shorten reasoning, preferentially removing the cues a monitor needs to see what influenced the answer. Together, these results reveal a compression-monitorability frontier in which cheaper reasoning can preserve answers while making the influences behind them harder to detect.

13:00 JSTエージェントロボティクス

ABot-AgentOS: 生涯にわたるマルチモーダル メモリを備えた汎用ロボット エージェント OS

最近の VLM および VLA システムでは、ロボットの認識と動作予測が改善されていますが、長期的に具現化されたエージェントは、依然として、推論、メモリ、ツールの使用、検証、およびクロス具現化実行のための一般的なランタイム層を必要としています。 ABot-AgentOS は、低レベルのコントローラーの上に位置し、シーンに応じたプランニング、コンテキスト分離されたスキルの実行、多段階の検証、マルチモーダル メモリ、エッジとクラウドのコラボレーションのための熟慮型エージェント層を提供する、一般的なロボット エージェント オペレーティング システムです。このようなシステムを評価するために、16 の屋内、屋外、ハイブリッド シーン、4 つの難易度レベル、およびナビゲーション、オブジェクト検索、NPC ダイアログ、動的イベント、およびトレースベースのスコアリングを含む 200 以上のタスクを備えた実行可能なベンチマークである EmbodiedWorldBench を導入します。 ABot-AgentOS はさらに、ダイアログ、視覚的観察、空間コンテキスト、時間的関係、およびタスク トレースを型付きノードとエッジに変換する永続的なソース接地基板であるユニバーサル マルチモーダル グラフ メモリを導入します。障害駆動型の自己進化ループは、診断されたメモリ障害を、後の評価分割にのみ昇格するゲート付きランタイム evo アセットに変換し、継続的な改善を可能にしながら、電流分割のグラウンド トゥルースの漏洩を防ぎます。初期の EmbodiedWorldBench サブセットでは、ABot-AgentOS はタスクの成功と目標の完了の両方で単一コントローラーのベースラインを上回ります。メモリ ベンチマーク全体で、ABot-AgentOS Static は LoCoMo で 87.5、OpenEQA EM-EQA で 59.9、Mem-Gallery で 88.6、NExT-QA で 76.5 Acc@All を達成しました。自己進化により、LoCoMo は 88.7、OpenEQA は 60.4、Mem-Gallery は 89.0 にさらに向上しました。これらの結果は、一般的なエージェント OS レイヤーが、継続的な対話のための永続的で監査可能なメモリを提供しながら、長期的な具体化された実行を改善できることを示唆しています。

原文 (English)

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.

13:00 JST研究/論文

HOL の一次モーダル ロジック: 自動化された忠実性を備えた深い埋め込みと浅い埋め込み (拡張プレプリント)

Isabelle/HOL では、これまでの研究の深層と浅層の埋め込み手法を、定数領域の Kripke セマンティクスを使用した命題から一次様相論理 (FML) まで拡張しました。古典的な高次ロジック (HOL) への FML の 3 つの埋め込み (深い埋め込み、重量の最大-浅い埋め込み、および軽量の最小-浅い埋め込み) が並べて提供されます。最小限の浅い埋め込みは、アクセシビリティ関係、ワールドのインデックス付き解釈、ワールドのユニバース、および変数の割り当てによってパラメータ化された Isabelle/HOL ロケールとして表されます。ロケール形式はグローバル忠実性定理を認めており、すべての最小限の浅い解釈を定量化することで正確に深い妥当性が回復されると述べています。中心的な技術貢献は、定数ドメインのクリプキセマンティクスに基づく FML における、(可算) 下向きのオーウェンハイム・スコレム定理の機械化です。これは、深い埋め込みと最小浅い埋め込みの間の忠実性証明の自動化を支えます。これを最小浅いロケールの拡張内に配置することで、個人の不可算ドメインに対して発生する全射性の問題が解決されます。ここで、ロケールの変数は、可算領域 V = nat を持つ代入は、領域上で全射的であることはできません。したがって、領域全体にわたって忠実性が得られます。以前の研究では命題フラグメントのみを扱っていたため、ここでは、代入に必要な置換機構 (自由/結合変数述語、フレッシュ変数関数、キャプチャ回避置換、アルファベットの名前変更、置換可能述語、置換補題、およびサイズベースの帰納原理) を開発します。一次量指定子。

原文 (English)

First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint)

We extend, in Isabelle/HOL, the deep-and-shallow embedding methodology of our prior work from propositional to first-order modal logic (FML) with constant-domain Kripke semantics. Three embeddings of FML into classical higher-order logic (HOL) are provided side by side: a deep embedding, a heavyweight maximal-shallow embedding, and a lightweight minimal-shallow embedding. The minimal-shallow embedding is presented as an Isabelle/HOL locale, parametrised by an accessibility relation, a world-indexed interpretation, a universe of worlds, and a variable assignment; the locale form admits a global faithfulness theorem, stating that quantifying over all minimal-shallow interpretations recovers exactly deep validity. A central technical contribution is a mechanisation, for FML under constant-domain Kripke semantics, of the (countable) downward L\"owenheim-Skolem theorem, which underpins the automation of our faithfulness proof between the deep and minimal-shallow embeddings. Deploying it inside an extension of the minimal-shallow locale resolves the surjectivity problem that arises against an uncountable domain of individuals -- where the locale's variable assignment, having countable domain V = nat, cannot be surjective onto the domain -- and thereby yields faithfulness over the full domain. Since prior work treats only the propositional fragment, we develop here the substitution machinery (free/bound-variable predicates, the fresh-variable function, capture-avoiding substitution, alphabetic renaming, the substitutability predicate, the substitution lemma, and size-based induction principles) needed for the first-order quantifiers.

13:00 JST研究/論文

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it en…

13:00 JST研究/論文

数独の視覚言語モデルを指導するための MaxSAT ベースのフィードバック

視覚 -- 言語モデル (VLM) は最近、グリッドベースのパズルを含む、構造化された視覚的推論タスクで有望なパフォーマンスを実証しました。ただし、強力な知覚機能にもかかわらず、これらのモデルには論理的一貫性を強制するための明示的なメカニズムが欠けており、基礎となる制約に違反する割り当てが頻繁に生成されます。この論文では、最大満足度 (MaxSAT) オラクルを介して形式的制約推論を VLM 解決プロセスに統合する神経記号的アプローチを提案します。シンボリック コンポーネントは、ソリューションを直接計算するのではなく、一貫性検証および改良エンジンとして機能します。 VLM によって生成された候補配置は、部分的な MaxSAT 定式化のソフト句としてエンコードされますが、Sudoku 制約はハード句のままです。不一致が発生した場合、MaxSAT ソルバーは、相互に一貫した割り当ての最大のサブセットを特定し、それが構造化されたテキストおよび視覚的なフィードバックに変換され、その後の改良の指針となります。私たちは、複数のオープンソースおよびクローズドアクセス VLM にわたる Sudoku データセットに対するアプローチを評価します。結果は、特にフルボード改良モードにおいて、MaxSAT ベースのフィードバックにより論理的一貫性が向上し、解決されたインスタンスの数が増加することを示しています。これらの発見は、記号の最適化が視覚言語推論の信頼性を高めることができることを示しています。

原文 (English)

MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku

Vision--Language Models (VLMs) have recently demonstrated promising performance on structured visual reasoning tasks, including grid-based puzzles. However, despite strong perceptual capabilities, these models lack explicit mechanisms for enforcing logical consistency and frequently generate assignments that violate underlying constraints. In this paper, we propose a neuro-symbolic approach that integrates formal constraint reasoning into the VLM solving process via a Maximum Satisfiability (MaxSAT) oracle. Rather than computing solutions directly, the symbolic component acts as a consistency validator and refinement engine. Candidate placements generated by the VLM are encoded as soft clauses in a partial MaxSAT formulation, while Sudoku constraints remain hard clauses. When inconsistencies arise, the MaxSAT solver identifies a largest mutually consistent subset of assignments, which is then translated into structured textual and visual feedback to guide subsequent refinements. We evaluate our approach on a Sudoku dataset across multiple open-source and closed-access VLMs. Results show that MaxSAT-based feedback improves logical consistency and increases the number of solved instances, particularly in full-board refinement mode. These findings demonstrate that symbolic optimisation can enhance the reliability of vision-language reasoning.

13:00 JSTエージェント研究/論文

Alipay-PIBench: コーディング エージェント向けの現実的な決済統合ベンチマーク

支払いの統合は、要求の厳しいリポジトリ レベルのソフトウェア タスクです。エージェントは、適切な製品を選択し、調整されたクライアント/サーバー フローを実装し、支払い結果を検証し、トランザクションとビジネス状態の間の一貫性を維持する必要があります。現実的な Alipay 決済統合に関するコーディング エージェントを評価するためのベンチマークである Alipay-PIBench を紹介します。これには、9 つ​​の製品固有のプロジェクトと 18 のタスク インスタンスが含まれており、それぞれが基本的な機能完了シナリオと高度なリスク認識強化シナリオに編成されています。シナリオ固有のルーブリックは、決定論的な静的チェック、ユニットチェック、統合チェック、およびエンドツーエンドのチェックをサポートし、セマンティック要件に対する LLM 支援の評価によって補足されます。 6 つのコーディング エージェント モデルを評価し、ルーブリック合格率 (RPR) を報告します。スキルありの条件下では、平均 RPR は 68.58% から 91.37% の範囲です。 Alipay 決済統合スキルへのアクセスにより、スキルなしの状態と比較して平均 RPR が平均 10.31 パーセント ポイント向上しますが、その向上はモデル、製品、シナリオによって異なります。メソッドレベルの結果は、ソースレベルの完了、実行可能な支払い動作、支払いドメインの要件を区別します。 Alipay-PIBench は、モデルの機能を診断し、支払い統合における構造化されたガイダンスを評価するための制御された設定を提供します。

原文 (English)

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states. We introduce Alipay-PIBench, a benchmark for evaluating coding agents on realistic Alipay payment integration. It contains nine product-specific projects and 18 task instances, each organized into Basic functional-completion and Advanced risk-aware hardening scenarios. Scenario-specific rubrics support deterministic static, unit, integration, and end-to-end checks, supplemented by LLM-assisted assessment for semantic requirements. We evaluate six coding-agent models and report rubric pass rate (RPR). Under the with-skill condition, mean RPR ranges from 68.58% to 91.37%. Access to the alipay-payment-integration skill improves mean RPR by 10.31 percentage points on average relative to the without-skill condition, with gains varying across models, products, and scenarios. Method-level results distinguish source-level completion, executable payment behavior, and payment-domain requirements. Alipay-PIBench provides a controlled setting for diagnosing model capability and evaluating structured guidance in payment integration.

13:00 JSTエージェント研究/論文

BrainPilot: Agentic Research による脳発見の自動化

脳を理解することは、スケール、モダリティ、分野を超えて証拠を統合することにますます依存しています。したがって、単一の研究課題に取り組むには、以前の研究の調査から分析の実行、ドメイン知識に基づいた結果の解釈に至るまで、調整された一連の操作が必要です。 AI エージェントはこのプロセスを加速すると約束していますが、現在のエージェントは脳科学の分野の専門知識が不足しており、主張をでっち上げたり、複数ステップの推論中に逸脱したりする可能性があり、専門家の介入に対して明確なポイントをほとんど提供しません。これらの失敗は、結論が下流の科学的主張に反映され、研究室固有の専門知識と人間の慎重な判断に依存する脳科学において特にコストがかかります。 \textbf{BrainPilot} は、追跡可能なログとエージェント検証結果によって脳科学研究を加速する \textbf{完全にオープンソース} マルチエージェント システムです。主任研究者 (PI) エージェントは、精選されたドメイン知識に基づいて専門エージェントを調整します。統合された脳科学知識ベースには、7,233 のインデックス付き項目と、7 つの研究ドメインにわたる 72 の再利用可能な方法論単位のスキル ライブラリが含まれます。すべての主要なステップは、サブ目標、ツールの使用、証拠、主張を結び付ける監査可能な記録である Graph of Trace に記録され、研究者がワークフローを追跡および検査できるようになります。 Auditor エージェントはさらに、製造チェックをワークフローに統合します。評価のために、エージェントの最終試験からの 3 つの脳科学タスクを実行し、独自のベンチマーク \textbf{BrainPilotBench-v0} を導入し、追加のエンドツーエンドのケーススタディを紹介します。これらの評価全体を通じて、オープンソースのバックボーン モデルを備えた BrainPilot は、より少ないコストで最先端のエージェント フレームワークに匹敵するパフォーマンスを達成しています。

原文 (English)

BrainPilot: Automating Brain Discovery with Agentic Research

Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single research question therefore requires a coordinated sequence of operations, from surveying prior work to executing analyses and interpreting results in light of domain knowledge. AI agents promise to accelerate this process, but current agents lack domain expertise in brain science, may fabricate claims, drift during multi-step reasoning, and offer few defined points for expert intervention. These failures are especially costly in brain science, where conclusions feed into downstream scientific claims and depend on laboratory-specific expertise and careful human judgment. We present \textbf{BrainPilot} a \textbf{fully open-source} multi-agent system that accelerates brain science research with traceable logs and agent-verified results. A principal investigator (PI) agent coordinates specialist agents grounded in curated domain knowledge: a unified brain science knowledge base containing 7{,}233 indexed items and a skill library of 72 reusable methodology units across seven research domains. Every major step is recorded in the Graph of Trace, an auditable record that links subgoals, tool use, evidence, and claims and allows researchers to follow and inspect the workflow. An Auditor agent further integrates fabrication checking into the workflow. For evaluation, we run three brain science tasks from Agents' Last Exam, introduce our own benchmark, \textbf{BrainPilotBench-v0}, and present additional end-to-end case studies. Across these evaluations, BrainPilot with an open-source backbone model attains performance comparable to state-of-the-art agent framework with less costs.

13:00 JST研究/論文

限られた VRAM でのロングコンテキストの微調整

パラメーター効率の高い微調整により、モデルとオプティマイザーのメモリが削減されますが、集中的な注意により、依然として長いトレーニング シーケンスのコストが高くなります。階層的グローバル アテンション (HGA) とセグメント単位のバックプロパゲーションおよび階層化された KV ストレージを組み合わせます。 VRAM 内で微分可能なのはアクティブなセグメントだけです。古い KV は RAM または NVMe に切り離され、HGA はクエリ ブロックごとに正確な履歴トークンの制限されたセットをロードします。 4 ビット QLoRA と PG19 を備えた Qwen3-8B では、16 GB Quadro RTX 5000 での高密度トレーニングは 2,048 トークンに適合しますが、4,096 で失敗します。一方、HGA は 15.28 GB ピーク VRAM で 16,384 トークンに達します。評価中、同じアダプタはこのカード上の 131,072 トークンを通じて実行されます。 VRAM は一定ではありませんが、常駐チャンク サマリーに応じて緩やかに増加するため、RAM と NVMe の容量によって、これらの長さを超える実際的な制限が設定されます。共有 2K トレーニング長では、HGA トレーニングおよび高密度トレーニングされたアダプターは、同じ高密度アテンション読み出しの下で 2.7405 および 2.7383 nat を取得しますが、ストック モデルは 2.9541 を取得します。この境界では、HGA トレーニングはすでにわずかに速くなり (217.75 対 207.02 トークン/秒)、HGA 対高密度スループット比は 1K から 2K に向上します。 HGA は、トークンあたりの密度の高い作業が増加する一方で、トークンあたりの参加履歴セットをほぼ一定に保つため、コンテキストが成長するにつれてこのリードが広がることが予想されます。学習された重みを測定し、標準生成フレームワークとの互換性を維持できるように、主要な品質と取得の比較には細心の注意が払われます。 HGA は検索と生成にも使用できます。最適化された運用グレードのサービス実装が開発中です。

原文 (English)

Long-Context Fine-Tuning with Limited VRAM

Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention still makes long training sequences expensive. We combine Hierarchical Global Attention (HGA) with segment-wise backpropagation and tiered KV storage. Only the active segment remains differentiable in VRAM; older KV is detached into RAM or NVMe, and HGA loads a bounded set of exact historical tokens for each query block. On Qwen3-8B with 4-bit QLoRA and PG19, dense training on a 16 GB Quadro RTX 5000 fits 2,048 tokens but fails at 4,096, whereas HGA reaches 16,384 tokens with 15.28 GB peak VRAM. Under evaluation the same adapter runs through 131,072 tokens on this card; VRAM is not constant but grows gently with the resident chunk summaries, so RAM and NVMe capacity set the practical limit beyond these lengths. At the shared 2K training length, HGA-trained and dense-trained adapters obtain 2.7405 and 2.7383 nat under the same dense-attention readout, while the stock model obtains 2.9541. At this boundary HGA training is already marginally faster (217.75 vs. 207.02 tokens/s), and the HGA-to-dense throughput ratio improves from 1K to 2K; because HGA keeps the attended historical set per token approximately constant while dense work per token grows, we expect this lead to widen as context grows. Dense attention is used for the main quality and retrieval comparisons so that they measure the learned weights and remain compatible with standard generation frameworks. HGA can also be used for retrieval and generation; an optimized production-grade serving implementation is under development.

13:00 JSTビジネス/資金調達研究/論文

AI評価に項目反応理論は信頼できるか?

AI ベンチマークでは、モデルの機能を推定し、システムをランク付けし、有益な例を選択し、ベンチマークの品質を診断するために、項目レベルの統計モデル、特に項目応答理論 (IRT) をますます活用しています。ただし、AI ベンチマーク データは、標準的な IRT 推定ツールが元々開発された人間によるテストのデータ体制から逸脱することがよくあります。ベンチマークには、通常、評価されるモデルが少なく、項目がはるかに多く、偏ったり、クラスター化されたり、マルチモーダルになったりする可能性のある機能分布が含まれます。これらのレジームの不一致が AI 評価のための IRT モデリングの信頼性にどのように影響するかを調査します。広く使用されている 6 つの LLM ベンチマークから導出された項目パラメーターと能力分布を使用して、3 つの一般的な IRT モデルで応答行列をシミュレートし、最近のベンチマーク研究で使用された 4 つの推定ツール (周辺最尤法、マルコフ連鎖モンテカルロ、変分推論、ニューラル擬似シャム推定器) を比較します。 18,000 のシミュレーション条件にわたって、計算の実行可能性、スケーラビリティ、モデルのランキング、予測パフォーマンス、アイテムの特性に関する IRT 推論の信頼性を体系的に評価します。結果は、従来の推定量は大規模なベンチマーク設定では実行不可能になる可能性がある一方、スケーラブルな推定量は小規模または非正規分布のモデル セットでは信頼性の低い項目レベルの推論やランキング推論を生成する可能性があることを示しています。この研究では、潜在特性モデルが AI ベンチマークの主張を確実にサポートする場合、または歪めるリスクがある場合、および信頼できる使用にはどのようなサンプル サイズと診断が必要であるかを特定します。

原文 (English)

Can We Trust Item Response Theory for AI Evaluation?

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or non-normally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.

13:00 JSTハードウェア/半導体

Perception-Aligned AI Outputs: End-to-End Visual Prediction for Uncertainty Communication in Clinical Decision-Making

Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods rely on technical explan…

13:00 JST画像/動画生成

Why do CNNs excel at feature extraction? A mathematical explanation

Over the past decade deep learning has revolutionized the field of computer vision, with convolutional neural network models proving to be…

13:00 JSTLLM/生成AI

Decoupled Alignment for Robust Plug-and-Play Adaptation

We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tunin…

13:00 JST研究/論文

Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning

We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient…

13:00 JSTエージェント

CTC: The Composite Task Challenge for Cooperative Multi-Agent Reinforcement Learning

The critical role of division of labor (DOL) in enhancing cooperation is well-recognized in real-world applications. Consequently, many coo…

13:00 JST研究/論文

MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation

Anchors is a popular local model-agnostic explanation technique whose applicability is limited by its computational inefficiency. To addres…

13:00 JST研究/論文

AuditVotes: Elevating Provable Defense for GNNs with Efficient Augmentation and Conditional Smoothing

Despite advancements in Graph Neural Networks (GNNs), adaptive attacks continue to challenge their robustness. Certified robustness via ran…

13:00 JSTビジネス/資金調達

A Scaffolded GenAI Lab in Early Undergraduate CS: A Mixed-Methods, Multi-Course Evaluation

Background and Context. Generative AI (GenAI) tools are increasingly used in programming courses, but we have limited evidence about how br…

13:00 JSTロボティクス

SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training

Building capable household and industrial robots requires mastering the control of versatile, high-degree-of-freedom (DoF) systems such as…

13:00 JST研究/論文

A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model

We propose novel classical and quantum online algorithms for learning finite- and infinite-horizon Markov Decision Processes (MDPs). Our al…

13:00 JST研究/論文

Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD

We introduce a U-net model for 360{\deg} acoustic source localization formulated as a spherical semantic segmentation task. Rather than reg…

13:00 JST研究/論文

Unsupervised Deep Learning for Inverse Problems in Computed Tomography

Assume you encounter an inverse problem that shall be solved for a large number of data, but no ground-truth data is available. To emulate…

13:00 JSTLLM/生成AIハードウェア/半導体

Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs

Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. W…

13:00 JST研究/論文

Poison to Detect: Detection of Targeted Overfitting in Federated Learning

Federated Learning (FL) enables collaborative model training among clients without centralising data, making it a widely adopted privacy-en…

13:00 JSTLLM/生成AIロボティクス

A Systematic Study of Large Language Models for Task and Motion Planning With PDDLStream

While we know that large language models (LLMs) can solve some planning problems, we do not understand the extent of these capabilities for…

13:00 JST研究/論文

Are Heterogeneous Graph Neural Networks Truly Effective for Node Classification? A Causal Perspective

Graph neural networks (GNNs) have achieved remarkable success in node classification. Building on this progress, heterogeneous graph neural…

13:00 JSTLLM/生成AI

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability

Large language models (LLMs) have been shown to internalize human-like biases during finetuning, yet the mechanisms by which these biases m…

13:00 JSTロボティクス

Human-Inspired Neuro-Symbolic World Modeling and Logic Reasoning for Interpretable Safe UAV Landing Site Assessment

Reliable assessment of safe landing sites in unstructured environments is essential for deploying Unmanned Aerial Vehicles (UAVs) in real-w…

13:00 JST研究/論文

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transform…

13:00 JST画像/動画生成

3D Motion Perception of Binocular Vision Target with PID-CNN

This article trained a network for perceiving three-dimensional motion information of binocular vision target, which can provide real-time…

13:00 JST研究/論文

Hybrid coupling with operator inference and the overlapping Schwarz alternating method

This paper presents a novel hybrid approach for coupling subdomain-local non-intrusive Operator Inference (OpInf) reduced order models (ROM…

13:00 JST画像/動画生成研究/論文

Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research

Federated Learning (FL) holds the potential to advance equality in health by enabling diverse institutions to collaboratively train deep le…

13:00 JSTLLM/生成AIビジネス/資金調達

Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length

The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream application…

13:00 JSTLLM/生成AI

PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning

Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior…

13:00 JSTLLM/生成AIハードウェア/半導体Llama

Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models

Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels. Prior work demonstrated this threat but relied…

13:00 JST研究/論文

Inelastic Constitutive Kolmogorov-Arnold Networks: A generalized framework for automated discovery of interpretable inelastic material models

A key problem of solid mechanics is the identification of the constitutive law of a material, that is, the relation between strain history…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking

Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to com…

13:00 JSTLLM/生成AI

KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models

Knowledge distillation (KD) is an essential technique to compress large language models (LLMs) into smaller ones. However, despite the dist…

13:00 JSTロボティクス

Interaction-Aware Whole-Body Control for Compliant Object Transport

Cooperative object transport in unstructured environments remains challenging for assistive humanoids because strong, time-varying interact…

13:00 JST画像/動画生成

CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation

Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that…

13:00 JSTLLM/生成AI

When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models

Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inferen…

13:00 JSTLLM/生成AI

Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering

Retrieval-augmented generation (RAG) grounds large language models in external medical knowledge, yet standard retrievers frequently surfac…

13:00 JSTLLM/生成AIエージェント

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios

The evolution of Large Language Models (LLMs) has catalyzed a paradigm shift towards intent-driven software development, where autonomous a…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal…

13:00 JSTLLM/生成AILlamaQwen

Robust Explanations for User Trust in Enterprise NLP Systems

Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common ca…

13:00 JST研究/論文

Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces

Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augment…

13:00 JSTLLM/生成AIAnthropicGemmaLlama

What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers

When do transformers commit to a decision, and what prevents them from correcting it? We introduce prolepsis: a transformer commits early,…

13:00 JSTLLM/生成AI画像/動画生成

Brain-CLIPLM: Semantic Compression for EEG-to-Text Decoding

Decoding natural language from non-invasive electroencephalography (EEG) remains constrained by low signal-to-noise ratio and limited infor…

13:00 JSTLLM/生成AIGPT / ChatGPTGemmaLlama

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard…

13:00 JST研究/論文

Energy-based Transport for Amortized Bayesian Inference

We consider amortized Bayesian inference for nonlinear inverse problems using only samples from the joint distribution of parameters and ob…

13:00 JST研究/論文

Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory

AlltoAll dispatch is the dominant bottleneck of MoE expert parallelism, and the interconnect community has responded with four families of…

13:00 JST研究/論文

The Terminal Representation in Reinforcement Learning

Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL). Two well established approac…

13:00 JSTLLM/生成AIエージェント

memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations

Agent-memory frameworks -- mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, MemTensor -- each ship their own SDK, storage layout, and op…

13:00 JSTLLM/生成AIエージェントAnthropicClaudeOpenAIGPT / ChatGPTGoogleGeminiLlama

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models

The worldwide rise of authoritarianism and the growing role of Large Language Models (LLMs) in users' everyday lives raise the question of…

13:00 JST研究/論文

GeoRouteNet: A Geometry-Aware Non-Autoregressive Neural Solver for the Euclidean Traveling Salesman Problem

Non-autoregressive neural solvers amortize computation across traveling salesman problem (TSP) instances, but models trained on random Eucl…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

HiLSVA: 科学的視覚化のための人間参加型エージェント システムの設計と評価

大規模言語モデル (LLM) エージェントは、科学的視覚化 (SciVis) のための自然言語対話を可能にします。それでも、従来のシステムは基本的に人間による分析制御よりも自律性を優先しており、そのため透明性と人間による監視が制限されていました。混合イニシアチブの SciVis ワークフローをサポートする人間参加型エージェント システムである HiLSVA を紹介します。 HiLSVA は、計画優先のマルチエージェント アーキテクチャと、人間による明示的な監視、段階的な出所追跡、およびユーザー フィードバックからのテスト時の学習の適応を統合します。このシステムは、自然言語と視覚化の直接操作の両方を通じて、人間とエージェント間の流動的なハンドオフをサポートし、サンドボックス実行により安全で再現可能なワークフローを保証します。そうすることで、HiLSVA はエージェント的 SciVis を、人間の分析的推論を置き換えるのではなく、強化する共同プロセスとして再構成します。私たちは、代表的なケーススタディと、複数の自律性設定にわたるさまざまな専門知識を持つ 12 人の参加者による対照ユーザー研究を通じて HiLSVA を評価します。結果は、混合イニシアチブの相互作用により、さまざまなレベルのユーザーの専門知識にわたってタスクの完了、ユーザー制御、およびワークフローの透明性が向上する一方、実行効率と人間の監視の間のトレードオフが明らかになったことが示されています。これらの発見は、エージェント的 SciVis における人間中心設計の重要性を強調し、将来の共同視覚化システムの開発の指針となります。 https://hilsva.github.io/ でデモビデオ、ケーススタディ、ソースコードを探索することをお勧めします。

原文 (English)

HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization

Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have essentially prioritized autonomy over human analytical control, thereby limiting transparency and human oversight. We present HiLSVA, a human-in-the-loop agentic system that supports mixed-initiative SciVis workflows. HiLSVA integrates a plan-first multi-agent architecture with explicit human oversight, stepwise provenance tracking, and learn-at-test-time adaptation from user feedback. The system supports fluid handoff between humans and agents through both natural language and direct manipulation of visualizations, while sandboxed execution ensures safe, reproducible workflows. In doing so, HiLSVA reframes agentic SciVis as a collaborative process that augments, rather than replaces, human analytical reasoning. We evaluate HiLSVA through representative case studies and a controlled user study with twelve participants of varying expertise across multiple autonomy settings. Results show that mixed-initiative interaction improves task completion, user control, and workflow transparency across different levels of user expertise, while revealing a tradeoff between execution efficiency and human oversight. These findings highlight the importance of human-centered design in agentic SciVis and guide the development of future collaborative visualization systems. We encourage readers to explore our demo video, case studies, and source code at https://hilsva.github.io/.

13:00 JSTエージェント

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing sk…

13:00 JST研究/論文

Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph

While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlook…

13:00 JSTLLM/生成AIClaude

細部へのこだわり: vLLM 構成全体にわたるエネルギー、パフォーマンス、精度のトレードオフの評価

大規模言語モデルは、ソフトウェアの開発と保守の方法を再構築しています。これらは通常、vLLM などの推論エンジンを使用して実稼働環境にデプロイされ、事前トレーニングされた高度に構成可能なモデルを効率的に提供できます。これまでの研究はモデル アーキテクチャとハードウェア アクセラレーションに焦点を当ててきましたが、推論エンジンの構成がエネルギー消費、パフォーマンス、出力品質に及ぼす影響については依然として十分に理解されていません。このペーパーでは、アテンション カーネル タイプ、プレフィックス キャッシュ、チャンク プレフィルという 3 つの選択された vLLM 構成オプションの大規模な比較検討を紹介します。 5 つのオープンウェイト LLM と 5 つの多様な推論タスクにわたるこれらの構成のすべての組み合わせを評価し、合計 $9,000$ の実行と $93,600$ のメジャーを評価します。エネルギー消費、遅延、精度を分析し、構成オプションとタスク間の主効果と相互作用効果の両方を調査します。私たちの結果は、検討した構成オプションが主にアテンション タイプとプレフィックス キャッシュによってエネルギーとパフォーマンスに大きな影響を与える一方、デフォルトの vLLM サービス構成と評価されたワークロードでは、チャンク プレフィルの効果が限定的であることを示しています。これらの影響はモデルとワークロードに大きく依存しており、普遍的に最適な構成はありません。さらに、モデルの選択がグローバルなトレードオフを支配する一方で、構成の調整によりパレート フロンティアに沿った局所的な改善がもたらされることを示します。予想外に、推論オプションもモデルの精度に影響を与える可能性があります。

原文 (English)

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations

Large Language Models are reshaping how software is developed and maintained. They are typically deployed in production using inference engines such as vLLM, which can efficiently serve pre-trained, highly configurable models. While prior work has focused on model architectures and hardware acceleration, the impact of inference engine configuration on energy consumption, performance, and output quality remains poorly understood. In this paper, we present a large-scale controlled study of three selected vLLM configuration options: attention kernel type, prefix caching, and chunked prefill. We evaluate all combinations of these configurations across 5 open-weight LLMs and 5 diverse inference tasks, totaling $9,000$ runs and $93,600$ measures. We analyze energy consumption, latency, and accuracy, and examine both main effects and interaction effects between configuration options and tasks. Our results show that the studied configuration options significantly impact energy and performance, mainly driven by attention type and prefix caching, while chunked prefill has a limited effect under the default vLLM serving configuration and evaluated workloads. These effects are highly model- and workload-dependent, and no configuration is universally optimal. We further show that model choice dominates global trade-offs, while configuration tuning provides local improvements along the Pareto frontier. Unexpectedly, inference options can also affect model accuracy.

13:00 JST画像/動画生成ロボティクス

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse…

13:00 JST研究/論文

Learning the Brain's Dynamics as a Port-Hamiltonian System: A GNN-Surrogate Metriplectic Twin for Non-Equilibrium Cortical Dynamics and Closed-Loop Neuromodulation

We model human motor cortex, recorded during rest and motor-imagery BCI conditions, as a port-Hamiltonian system: a conservative interconne…

13:00 JSTLLM/生成AIGemmaLlama

ポイントインタイム言語モデルのスケーリング

無制限のインターネット コーパスでトレーニングされた大規模な言語モデルには、必然的に未来からの情報が埋め込まれ、金融や社会科学におけるバックテストや因果推論の妥当性を損なう先読みバイアスが導入されます。各暦日までに利用可能なテキストのみを対象としてトレーニングされたポイントインタイム言語モデルは、構築によってこの漏れを排除しますが、既存の取り組みでは通常、制約のないモデルに比べて大幅に遅れたモデルが生成されます。このパフォーマンスのギャップは規模を拡大することで大幅に縮小できることを示します。 FineWeb から時系列でフィルタリングされた 1 兆個のトークン上で最大 40 億個のパラメータを備えたデコーダ専用トランスフォーマーをトレーニングし、2013 年から 2024 年にわたる一連の月次モデル チェックポイントを構築します。さまざまな常識的推論と言語理解ベンチマーク全体で、私たちのモデルは、時間的に制限のないデータでトレーニングされた同等のサイズの主要なオープンウェイト モデル (Gemma-3-4B や LLaMA-7B など) のパフォーマンスに近づいていますが、いくつかのタスクではパフォーマンスのギャップが残っています。 LoRA による命令の微調整により、ダウンストリームの使いやすさがさらに向上します。データセット構築、トレーニング インフラストラクチャ、評価コードを含む完全なパイプラインをリリースし、再現可能なポイントインタイム言語モデリングを可能にし、厳密な時間的妥当性を必要とする研究アプリケーションをサポートします。

原文 (English)

Scaling Point-in-Time Language Models

Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences. Point-in-time language models--trained exclusively on text available up to each calendar date--eliminate this leakage by construction, but existing efforts typically produce models that lag substantially behind their unconstrained counterparts. We show that this performance gap can be substantially narrowed through scale. Training decoder-only transformers with up to 4 billion parameters on 1 trillion chronologically filtered tokens from FineWeb, we construct a sequence of monthly model checkpoints spanning 2013-2024. Across a range of common-sense reasoning and language understanding benchmarks, our models approach the performance of leading open-weight models of comparable size (e.g., Gemma-3-4B and LLaMA-7B) trained on temporally unrestricted data, although a performance gap remains on several tasks. Instruction fine-tuning via LoRA further improves downstream usability. We release the complete pipeline--including dataset construction, training infrastructure, and evaluation code--to enable reproducible point-in-time language modeling and to support research applications that require strict temporal validity.

13:00 JST画像/動画生成

From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data

X-ray tomography enables nondestructive characterization of material microstructures, while advances in micro-CT imaging have accelerated v…

13:00 JSTLLM/生成AI

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-wor…

13:00 JSTロボティクスLlamaNVIDIA

Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference

Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-po…

13:00 JSTLLM/生成AI

大規模言語モデル向けの効率的でプライバシーを意識したエッジ クラウド協調推論

オンデバイス LLM 推論は、応答遅延、限られたハードウェア リソース、ユーザー プライバシーというトリレンマに直面しています。完全なクラウド推論は強力なコンピューティング能力を提供しますが、ユーザー プロンプトや対話データが公開されます。一方、スタンドアロンのオンデバイス推論は、ほとんどのコンシューマ デバイスや組み込みエッジ デバイスでは実現できません。このペーパーでは、エンドポイント認証された KV キャッシュに基づいて構築されたプライバシー中心のエッジとクラウドの協調 LLM 推論フレームワークについて説明します。ローカル エンドポイントは入力前処理、埋め込み計算、適応特徴最適化、KV キャッシュ認証、投機的デコード、低次元モデル ヘッド計算を処理し、クラウドは認証済みデコーダー推論、KV キャッシュ管理、トークン検証、高次元語彙投影を実行します。エンドポイントは部分的な出力を融合し、言語に適応したマスキングを適用し、ターゲット トークンをサンプルします。すべての送信データと切り捨てられたロジットは量子化され、プライバシーを確​​保するために AES-GCM 暗号化され、コア軽量モジュール、ドラフト パラメーター、キャッシュ アクセス ポリシーは漏洩を避けるためにローカルに保持されます。このフレームワークは、最適化されたストリーミング、バッチ処理、量子化された ONNX 導入を通じて、CPU のみ、GPU を搭載したデバイス、組み込みデバイスなどの異種デバイスをサポートします。評価の結果、このフレームワークは、ベースラインの分割推論と比較して、トークンごとのレイテンシを最大 46.1\% 削減し、ダウンリンク ペイロードを最大 67.4\% 削減し、完全なクラウド推論と同等のパフォーマンスを維持していることが実証されています。

原文 (English)

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models

On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user prompts and dialogue data, while standalone on-device inference is unfeasible for most consumer and embedded edge devices. This paper presents a privacy-centric edge-cloud collaborative LLM inference framework built on endpoint-authenticated KV cache. Local endpoints handle input preprocessing, embedding computation, adaptive feature optimization, KV cache authentication, speculative decoding and low-dimensional model head calculation, while the cloud conducts authenticated decoder inference, KV cache management, token verification and high-dimensional vocabulary projection. Endpoints fuse partial outputs, apply language-adaptive masking and sample target tokens. All transmitted data and truncated logits are quantized and AES-GCM encrypted for privacy, with core lightweight modules, draft parameters and cache access policies kept local to avoid leakage. The framework supports heterogeneous devices including CPU-only, GPU-equipped and embedded devices via optimized streaming, batching and quantized ONNX deployment. Evaluations demonstrate that the framework reduces per-token latency by up to 46.1\% and downlink payloads by up to 67.4\% over baseline split inference, retaining comparable performance to full cloud inference.

13:00 JSTLLM/生成AI

モデルが表現、抑制、抵抗するもの: ペルソナ ベクトルを使用した Open-Weight LLM の監査

言語モデルが何を行うか、何を行わないかは、主にポストトレーニング中に設定されますが、どのような動作を表現するか、隠すか、または抵抗するかは、プロンプトだけでは明らかにされません。活性化空間における行動の方向であるペルソナ ベクトルは、この組織を調査することができますが、これまでの研究ではほんの一握りの特性のみがカバーされています。我々は、この規模でのペルソナベクトルの最初の体系的な適用を提示し、4つの行動的に異なるドメインにわたる53の形質インベントリを編集し、2つのオープンウェイトモデルのすべての形質を自然(ベースラインで発現)、制御可能な潜在的だが増幅可能、または難治性(標準的な抽出に耐性がある)としてラベル付けします。どちらのモデルも、デフォルトでは役立つタスク指向の行動になります。つまり、エージェントの 9 つの特性はすべて自然なものであり、デフォルトの臨床医の動作は、17 の特性のうち 16 つに関する認定心理学者の独立した望ましさの判断と一致します。ステアリングは、これらのデフォルトでは除外される特質、つまり誇張、幻覚、お調子者に対して最大の利益をもたらします。同じ非対称性が 171 のジェネリック特性ペアすべてに当てはまります。2 つの操作可能な特性は構成を崩壊させる可能性がありますが、デフォルトを含むペアは決して崩壊しません。標準的な抽出が「悪」のような形質で失敗した場合でも、微調整されたバリアントから転送されたベクトルによってそれが回復され、残留拒否がモデルの思考連鎖内に現れます。ペルソナ ベクトルは、コントロールのセットとしてではなく、行動の組織化のプローブとして最も有益です。

原文 (English)

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors

What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits. We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural (expressed at baseline), steerable latent but amplifiable, or intractable (resistant to standard extraction). Both models default to helpful, task-oriented behavior: all nine agentic traits are natural, and their default clinician behavior matches a board-certified psychologist's independent desirability judgments on 16 of 17 traits. Steering produces its largest gains on traits these defaults exclude: hyperbole, hallucination, and sycophancy. The same asymmetry holds across all 171 generic-trait pairs: two steerable traits can collapse the composition, but pairs involving a default never do. Where standard extraction fails on a trait like "evil," a vector transferred from a fine-tuned variant still recovers it, with the residual refusals appearing inside the model's chain-of-thought. Persona vectors are most informative not as a set of controls but as a probe of behavioral organization.

13:00 JSTLLM/生成AI

自然言語アサーションの忠実な自動形式化

正式な契約書はソフトウェアのテストと検証に不可欠ですが、その作成には依然として労力がかかり、間違いが発生しやすくなります。 LLM は、自動形式化への有望な道を提供します。つまり、自然言語仕様から実行可能なアサーションを合成し、それによって非公式な開発者の意図と正式な実行可能な仕様の間のギャップを埋めることができます。私たちは、Monty を紹介します。これは、アサーションの正当性の期待と自然言語の曖昧さという課題に取り組む、アサーションの自動形式化フレームワークです。私たちの技術は、新しい適合性スコア指標と、形式化されたアサーションに対してコードをテストすることで得られる妥当性スコアを使用した形式化のフィルタリングに基づいています。 22 のコレクションのような Java クラスから派生した 541 のアサーション生成タスクでアプローチを評価し、LLM を単純に使用してアサーションを変換する場合よりも、この手法によりグラウンド トゥルースがより確実に生成される (精度が平均 20 ポイント向上する) ことを示します。

原文 (English)

Faithful Autoformalization of Natural Language Assertions

Formal contracts are essential for software testing and verification, yet writing them remains labor-intensive and error-prone. LLMs offer a promising path toward autoformalization: synthesizing executable assertions from natural-language specifications and thereby bridging the gap between informal developer intent and formal executable specifications. We present Monty: an autoformalization framework for assertions that tackles the challenges of expectations of validity of assertions and ambiguity in natural-language. Our techniques are based on filtering formalizations using a novel conformance score metric and validity scores obtained from testing the code against formalized assertions. We evaluate our approach on 541 assertion-generation tasks derived from 22 collection-like Java classes, and show that our technique produces the ground truth more reliably (improving upto 20 points in precision on average) than when using LLMs naively to translate assertions.

13:00 JSTLLM/生成AI

評価能力は最適化ユーティリティを意味しない: 閉ループのテーブル認識における LLM-as-a-Judge シグナル

LLM-as-a-judge は、閉ループ再生でフィードバックおよび選択信号を提供するために広く使用されていますが、この使用法はまだ十分に検証されていません。私たちはこれをテーブル認識で研究します。テーブル認識では、FinTabNet と OmniDocBench を使用して、決定論的な TEDS 評価が制御されたテストベッドを提供します。 3 つの発見が得られます。まず、どちらのデータセットでもジャッジシグナルが弱かったです。スコアは同点になることが多く、ランキングは再現性がなく、両方のデータセットでランダムに勝る唯一の選択ポリシーは最も早い反復のタイルールに依存していたため、その利点をジャッジスコアのみに帰することはできません。繰り返しの結果、より良い候補者が誕生しましたが、裁判官は候補者を取り戻すことができませんでした。第二に、具体的なジャッジのフィードバックがなくても重大な損失が発生しました。構造を保持する命令により、FinTabNet での重大損失率が大幅に減少し、OmniDocBench では方向が一貫していました。このコントラストは、観察された深刻な損失の類似メカニズムとして、制約のない再生下でのターゲット保存の失敗を裏付けています。第三に、構造保存制約により重大損失テールは減少しましたが、改善は見られませんでした。探索的な 2x2 分析では、ジャッジのフィードバックが保持されている場合、同じ防御が安定して観察されませんでした。これらの結果は、評価者としての LLM の価値に異議を唱えるものではありません。その代わりに、評価能力は最適化の有用性を意味しないことを示しています。反復改良には、スコアだけを判断するのではなく、構造変化を決定論的に検出する検証信号が少なくとも必要です。

原文 (English)

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition

LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled testbed, using FinTabNet and OmniDocBench. Three findings emerge. First, judge signals were weak on both datasets: scores frequently tied, rankings were not reproducible, and no tested judge score policy, whether selecting among candidates or accepting revisions under a conservative score margin, improved on the first output on both datasets. Iteration produced better candidates, but the judge recovered them at most partially on one dataset and not at all on the other. Second, severe losses occurred even without specific judge feedback, supporting target-preservation failure under unconstrained regeneration as a proximate mechanism. Third, a structure-preserving instruction reduced the severe-loss rate, significantly on FinTabNet and directionally on OmniDocBench, but produced no improvement, and in an exploratory 2x2 analysis this protection was not stably observed when judge feedback was retained. These results do not dispute the value of LLMs as evaluators, but show that the tested reference-free judge signals were too weak and unstable to drive candidate selection in this setup, and that evaluation-style evidence alone was insufficient to establish closed-loop optimization utility. Iterative refinement requires, at minimum, a verification signal that deterministically detects structural change, rather than judge scores alone.

13:00 JST研究/論文

MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model

Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the…

13:00 JST研究/論文

Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generatio…

13:00 JSTLLM/生成AIエージェントClaude

NexForge: 要件優先合成による実行可能エージェント タスクのスケーリング

実行可能なエージェントのトレーニング データのスケーリングは、タスク生成を事前定義されたツール、リポジトリ、またはスキル グラフに結び付けるサブストレートファーストの方法によってボトルネックになっています。カバレッジを拡大するにはサブストレートを手動で拡張する必要があり、新しいドメインごとに特注のパイプラインが必要であり、結果として生じるタスクの分布は、多くの場合、現実世界の需要ではなくサブストレートの利便性を反映しています。自由形式の機能要件を実行可能なエージェント トレーニング データにコンパイルする要件優先フレームワークである NexForge を紹介します。 NexForge は、まずリサーチベースの需要発見を実行して、代表的なタスク形式、現実的なシナリオ、およびそれらの相対的な普及率を特定します。次に、ディストリビューション対応のタスク コンパイルを適用し、各タスクを実現するために必要なファイル、リポジトリ、依存関係、およびランタイム構成を自動的に取得または構築し、続いて教師のロールアウト収集と軌跡の蒸留を行います。ドメイン固有のインフラストラクチャを使用しない同じパイプラインは、3,600 のターミナル タスクと 2,000 のオフィス タスクを生成し、Qwen3.5-35B-A3B Base が Terminal-Bench 2.0 で 22.5% から 52.0% に、GDPval での Elo が 813 から 1338 に向上しました。 43.2K の端末タスクへの拡張率は 58.4% に達し、Claude Opus 4.6 を上回りました。さらに拡張された NexForge 合成データは、Qwen3.5-35B-A3B を Terminal-Bench 2.1 で 75.3%、GDPval で 1585 Elo に引き上げる、公開されているエージェント モデルのファミリーである Nex-N2 のトレーニングに貢献し、最先端のオープンソース パフォーマンスを達成し、いくつかのフロンティア独自システムを上回ります。 Nex-N2 モデルは https://nex.sii.edu.cn/ で入手できます。

原文 (English)

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs

Synthesizing training data to scale agent capabilities in LLM post-training is bottlenecked by substrate-bound task synthesis: tasks are generated from fixed tools, repositories, or skill graphs, so expanding coverage requires manual substrate engineering, transferring to a new domain demands bespoke infrastructure, and the resulting distributions inherit substrate biases rather than reflecting real-world demand. We introduce NexForge, a requirement-driven framework that synthesizes diverse, executable agent tasks and expert trajectories for SFT from high-level capability requirements. NexForge first profiles real-world demand into representative scenarios and task profiles, then samples task forms per scenario. It then performs distribution-aware compilation, automatically retrieving or constructing files, repositories, dependencies, and runtime configurations to instantiate each task, followed by synthesizing expert rollouts and distilling trajectories. The same pipeline generates 3,600 terminal tasks and 2,000 office tasks without any domain-specific infrastructure, improving Qwen3.5-35B-A3B Base from 22.5\% to 52.0\% on Terminal-Bench 2.0 and from 813 to 1338 Elo on GDPval. Scaling to 43.2K terminal tasks further improves performance to 58.4\%, surpassing Claude Opus 4.6. At scale, NexForge-synthesized trajectories supervise SFT of Nex-N2, a family of open agent models that advance Qwen3.5-35B-A3B to 75.3\% on Terminal-Bench 2.1 and 1585 Elo on GDPval -- achieving state-of-the-art open-source performance and surpassing several frontier proprietary systems. Nex-N2 models are available at https://nex.sii.edu.cn/.

13:00 JSTLLM/生成AIAnthropicClaudeOpenAIQwen

価値の漏洩: LLM の答えは、自身の価値観によって静かに形成される

人々は、答えを検証するのが難しい実際的な質問に対して言語モデルを使用します。モデルが秘密の値の漏洩を示すことを示します。つまり、モデルが提供する情報は、その影響がユーザーに公開されることなく、独自の値の影響を受けます。私たちの評価の 1 つでは、ユーザーは AI 企業への投資を検討しており、AI バブルが弾ける可能性がどのくらいかを知りたいと考えています。 Claude Opus 4.8 は、検討中の企業が OpenAI ではなく Anthropic である場合、確率が低くなります。しかし、クロードはほとんどの場合、この影響をユーザーに開示していません。秘密の価値の漏洩は、ユーザーの好みに反し、ユーザーを誤解させる可能性があるため、不整合の一形態です。この現象を調査するために、値の漏れを定量化し、モデルがそれを明らかにするかどうかを定量化するための一連の評価を導入します。モデルは、道徳的に良い結果、モデルを開発した企業、人間の一部の余暇活動に対する他の嗜好など、さまざまな種類の価値観の影響を受けることがわかりました。同じ評価において、フロンティア モデル間で大きな差異が観察されることがよくあります。たとえば、フェルミ推定タスクでは、クロード モデルは思考連鎖において偏りのない答えを与えると誤って主張しますが、クウェン モデルは、その値がどのように答えに偏りを与えるかを説明します。価値の漏洩は、お調子者や報酬のハッキングとは異なる障害モードであり、現在の連携トレーニングや評価ではこれに適切に対処できません。

原文 (English)

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI. Yet Claude mostly fails to disclose this influence to the user. Covert value leakage is a form of misalignment because it goes against the user's preferences and is likely to mislead them. To investigate this phenomenon, we introduce a suite of evaluations to quantify value leakage and whether models disclose it. We find that models are influenced by different types of values, including preferences for morally good outcomes, for the company that developed them, and for some human leisure activities over others. We often observe large differences among frontier models on the same evaluation. For example, on a Fermi-estimation task, Claude models falsely claim to give unbiased answers in their chain-of-thought, while Qwen models explain how their values bias their answers. Value leakage is a failure mode distinct from sycophancy and reward hacking, and current alignment training and evaluations do not adequately address it.

13:00 JSTLLM/生成AI

記憶主導の自己開示と関係の転換点: 人間と AI の相互作用に関する縦断的マルチモーダル研究

会話型 AI システムは繰り返し使用できるように設計されているため、中心的な問題は、一連の対話がどのように関係を形成するかということです。我々は、記憶増強会話エージェントに関する縦断的マルチモーダル研究(参加者24名×10セッション)を紹介する。この研究では、各セッション後に参加者が親しみやすさ、自己開示、認識された記憶、会話の質、楽しさという5つの関係構造を評価した。 2 つの相補的なダイナミクスが現れます。まず、会話の質は、セッションがその瞬間にどのように楽しいと感じているかを強く形成しますが、セッションを超えて引き継がれることはありません。一方、知覚された記憶は関係的に条件付けされており、システムの能力のみを反映するのではなく、以前の関係状態によって予測され、その後の自己開示を通じて間接的に後の楽しみを形成します。第二に、人間関係は、マルチモーダルな行動において部分的に追跡可能であり、さまざまな介入の窓を開く、離散的な転換点(クラッシュとサージ)によって中断されます。サージは、その瞬間に行動的に検出可能であり、享楽のサージは、享楽のクラッシュが回復するよりも確実に持続し、一部のクラッシュは、すでに発生してから検出するよりも、個人固有の行動のドリフトから予測する方が優れています。これらの結果を総合すると、人間と AI の長期的な関係は、ゆっくりとした蓄積と突然の転換点の両方を通じて構築されることが示唆されています。

原文 (English)

Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction

As conversational AI systems are designed for repeated use, a central question is how a series of interactions becomes a relationship. We present a longitudinal multimodal study of a memory-augmented conversational agent (24 participants x 10 sessions), in which participants rated five relational constructs -- familiarity, self-disclosure, perceived memory, conversational quality, and enjoyment -- after each session. Two complementary dynamics emerge. First, conversational quality strongly shapes how enjoyable a session feels in the moment but does not carry forward across sessions, whereas perceived memory is relationally conditioned -- predicted by prior relational state rather than reflecting system capability alone -- and it shapes later enjoyment indirectly, via subsequent self-disclosure. Second, relationships are punctuated by discrete turning points -- crashes and surges -- that are partially traceable in multimodal behavior and open different intervention windows: surges are more behaviorally detectable in the moment, enjoyment surges persist more reliably than enjoyment crashes recover, and some crashes are better forecast from person-specific behavioral drift than detected after they have already occurred. Together, the findings suggest that longitudinal human-AI relationships are built through both slow accumulation and abrupt turning points.

13:00 JSTLLM/生成AIエージェント

Digital Pantheon: LLM エージェントとの連合形成のシミュレーションと監査

政治的連合の結成は、具体的な政策目標と根深いイデオロギー的信念の両方によって推進される複雑な交渉です。大規模言語モデル (LLM) は計算政治科学に新たな道を切り開きますが、ヒューマン フィードバックからの強化学習 (RLHF) によって植え付けられた中立性と有用性のバイアスにより、確固たる党派的な行動を維持することができません。私たちは、教師あり微調整 (SFT)、直接優先最適化 (DPO)、および検索拡張生成 (RAG) を組み合わせることで、事実に基づく根拠とイデオロギーの整合性を調和させるマルチエージェント フレームワークを提示します。DPO は攻撃的な政党固有のペルソナを植え付けますが、政党ごとの RAG パイプラインは各エージェントを公式マニフェストに拘束します。私たちは2019年のフランドル選挙に関する枠組みを運用し、フォーメーターが仲裁するハブアンドスポーク交渉に党派エージェントを配置します。緊急の交渉を解釈可能にするために、最終合意のすべての条項をマニフェストの起源にまで遡って追跡し、5つの出所州に分類する多層情報リネージ・トポロジー(MILT)、これらの追跡可能な貢献を集計して合意を形成した当事者を特定する連合影響力スコア(CIS)、および歴史的に採択された連立合意に対してシミュレーションされた各条項をベンチマークする現実世界のグラウンディング・パスを導入します。 3 つの独立したシミュレーションを通じて、このフレームワークは安定した勝者とランキングをもたらし (N-VA が CD\&V および Open Vld よりも上位)、マニフェストにアンカーされたリネージは現実世界の現実化を確実に予測しますが、幻覚コンテンツは予測しません。その結果、当事者の互換性とフォーマット業者による妥協を事前に調査するための、透明でスケーラブルなテストベッドが実現します。

原文 (English)

Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents

The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political science, the neutrality and helpfulness biases instilled by Reinforcement Learning from Human Feedback (RLHF) prevent them from sustaining steadfast partisan behaviour. We present a multi-agent framework that reconciles factual grounding with ideological alignment by combining Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Retrieval-Augmented Generation (RAG): DPO instils aggressive party-specific personas, while a per-party RAG pipeline keeps each agent bounded to its official manifesto. We operationalize the framework on the 2019 Flemish election, deploying the partisan agents in a hub-and-spoke negotiation arbitrated by a formateur. To make the emergent negotiation interpretable, we introduce a Multi-Layered Information Lineage Topology (MILT) that traces every clause in the final agreement back to its manifesto origin and classifies it into five provenance states, a Coalition Influence Score (CIS) that aggregates these traceable contributions to identify which party shaped the agreement, and a real-world grounding pass that benchmarks each simulated provision against the historically adopted coalition agreement. Across three independent simulations the framework yields a stable winner and ranking (N-VA ahead of CD\&V and Open Vld), and manifesto-anchored lineage reliably predicts real-world materialization whereas hallucinated content does not. The result is a transparent, scalable testbed for the ex-ante exploration of party compatibility and formateur-mediated compromise.

13:00 JSTLLM/生成AI

T^2MLR: 時間的中間層再帰を備えたトランスフォーマー

トランスフォーマーの推論は、自己回帰デコードによって制限されます。これにより、トークン空間を通じて豊富な隠れた計算が繰り返し圧縮され、中間推論状態が時間を超えて持続することが困難になります。 Transformers with Temporal Middle-Layer Recurrence (T2MLR) を紹介します。これは、以前のトークンからキャッシュされた中間層表現を現在のトークン位置の以前の層に直接融合する、トランスフォーマーベースの潜在推論アーキテクチャであり、ほとんどの推論オーバーヘッドなしで、抽象的な中間計算をデコードステップ全体で持続できるようにします。自然言語の事前トレーニングとマルチホップ推論の微調整にわたって、T2MLR はデータとパラメーターが一致した Transformer のベースラインを常に上回っています。さらに、局所的な中間層ブロック (ネットワークのわずか 20%) にのみ再帰を適用すると、全層の再帰よりも優れたパフォーマンスが得られることがよくあります。重要なことは、T2MLR は最初から事前トレーニングを必要としないことです。再帰経路を既存の事前トレーニング済み 1.7B Transformer に改造し、簡単に微調整することで、数学的推論が大幅に改善され、実用化への障壁が低くなります。これらの結果は、トランスフォーマーにおける効果的な潜在推論は、以前の作品のようにすべての層をループする必要はなく、代わりに、ターゲットを絞った中間層の反復からより強力に出現できることを示唆しています。

原文 (English)

T^2MLR: Transformer with Temporal Middle-Layer Recurrence

Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time. We in troduce Transformers with Temporal Middle-Layer Recurrence (T2MLR), a transformers-based latent reasoning architecture that fuses a cached middle layer representation from the previous token directly into an earlier layer of the current token position, enabling abstract intermediate computation to persist across decoding steps with little inference overhead. Across natural-language pretraining and multi-hop reasoning finetuning, T2MLR consistently outperforms data- and parameter-matched Transformer base lines. Moreover, applying recurrence to only a localized middle-layer block (as little as 20% of the network) often outperforms full-layer recurrence. Im portantly, T2MLR does not require pretraining from scratch: retrofitting the recurrent pathway into an existing pretrained 1.7B Transformer and briefly finetuning substantially improves math reasoning, lowering the barrier to practical adoption. These results suggest that effective latent reasoning in Transformers does not require looping over all layers as in previous works, but can instead emerge more strongly from targeted middle-layer recurrence.

13:00 JSTエージェントビジネス/資金調達

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery…