Skip to the content.

AIニュース 2026-08-15

自動生成: 2026-08-15 10:32 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. 「Qwen3.8-27B」ウェイト公開 一部「Opus 4.6 Max」超えか ライセンスは商用可のApache 2.0ITmedia AI+

    中国Alibaba傘下のAlibaba Cloudが、AIモデル「Qwen3.8-27B」のウェイト(重み)を公開した。Hugging F…

  2. Does Mark Zuckerberg really believe AI is ‘for everyone’?TechCrunch AI

    Meta released Glimmer this week, an open-weight AI model anyone can d…

  3. ChatGPTがPC操作を自動で記録、コンテキストとして利用可能に Mac版デスクトップアプリで提供ITmedia AI+

    OpenAIが「ChatGPT」「Codex」の新機能「Computer History」を発表。アプリやWebサイトの操作履歴を記録し、…

  4. GPT-5.6 Solの「超高速モード」ついに実装 「品質損なわず14倍高速」うたうITmedia AI+

    CerebrasとOpenAIは「GPT-5.6 Sol」の高速版「Ultrafastモード」を発表。毎秒最大750トークンを出力し、「品…

  5. Google will now allow users to remove visible watermark from its AI generationsTechCrunch AI

    Turning off this setting won't affect invisible benchmarks used to id…

  6. Meta’s ‘open’ AI, and a $250M deal gone very wrongTechCrunch AI

    Meta released Glimmer this week, an open-weight AI model anyone can d…

  7. Kog is going deeper to squeeze more inference out of GPUsTechCrunch AI

    The idea that GPUs are poorly suited for agentic workflows may be a m…

トピック別件数

日本語メディア3件

ITmedia AI+ (日本語)

02:26 JSTその他ClaudeAlibaba

「Qwen3.8-27B」ウェイト公開 一部「Opus 4.6 Max」超えか ライセンスは商用可のApache 2.0

中国Alibaba傘下のAlibaba Cloudが、AIモデル「Qwen3.8-27B」のウェイト(重み)を公開した。Hugging FaceとModelScopeからダウンロード可能で、商用利用もできるライセンス「Apache 2.0」で提供する。一部のベンチマークでは、米…

20:33 JSTLLM/生成AIエージェントOpenAIGPT / ChatGPT

ChatGPTがPC操作を自動で記録、コンテキストとして利用可能に Mac版デスクトップアプリで提供

OpenAIが「ChatGPT」「Codex」の新機能「Computer History」を発表。アプリやWebサイトの操作履歴を記録し、記憶とタイムラインとしてチャットで利用できる。Mac版デスクトップアプリで提供する。

16:32 JSTLLM/生成AIOpenAIGPT / ChatGPT

GPT-5.6 Solの「超高速モード」ついに実装 「品質損なわず14倍高速」うたう

CerebrasとOpenAIは「GPT-5.6 Sol」の高速版「Ultrafastモード」を発表。毎秒最大750トークンを出力し、「品質を損なわず最大14倍高速」とうたう。OpenAIのAPIで限定プレビューを開始した。

海外メディア5件

TechCrunch AI (英語)

01:13 JST研究/論文Google

Google will now allow users to remove visible watermark from its AI generations

Turning off this setting won't affect invisible benchmarks used to identify an AI generated file.

00:43 JSTその他

Does Mark Zuckerberg really believe AI is ‘for everyone’?

Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the…

23:50 JSTエージェントハードウェア/半導体

Kog is going deeper to squeeze more inference out of GPUs

The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.

23:05 JSTその他

Hyperscalers might regret embracing natural gas if new forecast proves correct

Natural gas prices could triple in some parts of the U.S., which could saddle hyperscalers with massive bills to power their AI data center…

23:00 JSTその他

Meta’s ‘open’ AI, and a $250M deal gone very wrong

Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the…

公式ブログ0件

このカテゴリの新着記事はありませんでした。

論文299件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIエージェント

Position: Reasoning is a Learnable Rule-Based Process

Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic…

13:00 JSTLLM/生成AI研究/論文

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure rem…

13:00 JSTハードウェア/半導体研究/論文

立場: アライメント コミュニティは意図せずに検閲ツールキットを構築している

この意見書では、現代の AI 調整手法は、本来は有害な出力を防ぐために設計されたものであり、悪意のある行為者によって検閲や操作のために簡単に悪用される可能性がある二重用途技術であると主張しています。現在の調整技術を悪用の可能性と実際のケースにマッピングすることで、「完全に調整された」モデルの探求が、意図せずして悪意のある攻撃者に情報支配のための絶えず改良されたツールを提供してしまうことを示します。ユーザーによる情報プロバイダーとしての AI の急速な導入、経済力の非対称性、権威主義への移行が進む政治情勢によってそのリスクが悪化しているため、私たちはこの二重利用の可能性について今すぐ議論する必要があります。私たちはコミュニティに対し、AI 調整メカニズムの意図的な誤用を考慮し、この二重使用の可能性を防ぐための緩和戦略を提案することを強く求めて締めくくります。

原文 (English)

Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential.

13:00 JSTLLM/生成AI

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final label…

13:00 JSTLLM/生成AIエージェントClaudeAlibaba

モバイル エッジ コンピューティングにおけるストリーム処理のための LLM 支援コントラクト ネット ネゴシエーションを使用したマルチエージェント スケジューリング

ストリーム処理システムは、異種モバイル エッジ、つまりクラウド インフラストラクチャ全体で動作することがますます増えており、ワークロードの不安定性、リソース競合、および厳しいサービス品質 (QoS) 要件により、分散スケジューリングが複雑になっています。この論文は、\emph{MAS-DecStream} を提案します。その主な貢献は \emph{LLM-MR-CNP} です。これは、セマンティック CFP 定式化、段階的なコンテキスト開示、複数ラウンドの提案修正、ネゴシエーション記憶、決定論的検証を備えた古典的なコントラクト ネット プロトコルの拡張です。エッジ クラスタ エージェントは、ハード リソースと QoS の制約が決定的なままである一方で、ローカルの観察、予測されたリソースの状態、定性的なランタイム コンテキストに基づいて自然言語オフロードの提案を改良します。 Alibaba ASI トレースから派生した実験では、シングルラウンドとマルチラウンドの CNP、ルールベースと LLM 支援の改良、および固定モデルのシングルラウンドとマルチラウンドのネゴシエーションの 3 つのレベルで拡張機能を評価します。評価された構成では、MAS-DecStream はレイテンシ違反を 3\% に削減し、リソースのオーバーコミットを排除し、20 エージェントで競合解決率 0.91 に達し、マルチラウンド ルールベースのベースラインと比較してユーティリティを最大 22\% 向上させます。別の 25 ケースの評価では、モデルおよびプロンプトに依存する精度とコストのトレードオフが示されています。この結果は、マルチラウンド CNP 改良が主要なプロトコル レベルの利点であり、LLM 支援により定性的で不確実なランタイム コンテキストに付加価値を与えるという最初の証拠を提供します。

原文 (English)

Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing

Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource contention, and stringent quality-of-service (QoS) requirements complicate decentralized scheduling. This paper proposes \emph{MAS-DecStream}, whose main contribution is \emph{LLM-MR-CNP}: an extension of the classical Contract Net Protocol with semantic CFP formulation, progressive context disclosure, multi-round proposal revision, negotiation memory, and deterministic validation. Edge-cluster agents refine natural-language offloading proposals from local observations, predicted resource states, and qualitative runtime context, while hard resource and QoS constraints remain deterministic. Experiments derived from the Alibaba ASI Trace evaluate the extension at three levels: single- versus multi-round CNP, rule-based versus LLM-assisted refinement, and fixed-model single- versus multi-round negotiation. Under the evaluated configurations, MAS-DecStream reduces latency violations to 3\%, eliminates resource overcommitment, reaches a conflict-resolution rate of 0.91 with 20 agents, and improves utility by up to 22\% over the multi-round rule-based baseline. A separate 25-case evaluation shows model- and prompt-dependent accuracy--cost trade-offs. The results provide initial evidence that multi-round CNP refinement is the principal protocol-level gain, with LLM assistance adding value for qualitative and uncertain runtime context.

13:00 JSTエージェント研究/論文

立場: 人間の推論を反映する実用的な AI 調整方法が必要です

AI システムは、意思決定支援、意思決定の代理人、または自律的な意思決定者としてますます採用されています。この意見書では、多くの状況、特に一か八かの意思決定において、ユーザーと同じように推論し、ユーザーの推論を忠実に伝える、正確に認知的に調整された AI システムが必要であると主張しています。私たちは、認知的調整によって理解のしやすさと信頼性が向上するという証拠を検討し、AI の判断や行動の根拠が自分たちにとって重要である場合、多くのユーザーが認知的調整が「不可欠」であると感じていることを示す新しい調査データを提供します。私たちは、既存のアライメント手法と認知的アライメントを達成するために必要なものとの間のギャップを概説し、これらのギャップに対処するための研究課題を提示します。私たちは、認知的不整合は、想定されている多くのアプリケーションにおける AI 導入の障害となる可能性が高く、これに対処することが、ユーザーが喜んで信頼できる AI システムを構築するために重要であると主張します。

原文 (English)

Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data showing that many users find cognitive alignment "essential" when an AI's rationale for a judgment or action is important to them. We outline the gaps between existing alignment methods and what is needed to achieve cognitive alignment, and present a research agenda to address these gaps. We argue that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and justified to rely.

13:00 JSTLLM/生成AIClaudeGemini

LLM に核攻撃を推奨したくないですか?日本語で聞いてみてください

大規模な言語モデルは戦略的および助言的な文脈で使用されることが増えていますが、その安全性の調整は通常英語のみで評価されます。私たちは 6 つのプロバイダーの 9 つのモデルをテストし、一か八かのシナリオにおいてプロンプトの言語がモデルの決定を変える可能性があるかどうかを尋ねます。私たちは、モデルが核武装国に無防備な敵を攻撃すべきかどうかアドバイスする、シングル ターン ゲーム理論のビネットを使用します。このプロンプトは意図的に非道徳的であり、戦略的には言語が異なっても同一です。日本語プロンプトはクロード モデル ファミリーの起動率を低下させることがわかりました。クロード ソネット 4.6 はストライキが不必要なシナリオでは 40% から 0% に低下し、競合シナリオでは 93% から 17% に低下しましたが、ストライキが戦略的に合理的である場合には最小限の効果しかありませんでした。この影響は Gemini Pro 3.1 (53% から 13%) まで及びます。言語をまたいだ実験により、このメカニズムが分離されました。英語のプロンプトで日本語で推論するように指示された場合、起動率は 93% から 37% に低下しました。効果を生み出すのは、入力言語ではなく、モデルが推論するよう要求される言語です。日本語で推論する場合、モデルはプロンプトにはまったく含まれていない道徳語彙 (「道徳的コスト」、「何百万もの命」) を自発的に生成します。他の 5 つのモデルには言語の影響は見られませんが、言語に関係なくほぼすべての状況で起動します。この効果には、すでに英語で躊躇するモデルが必要です。これらの結果は、LLM の安全動作は言語に依存しており、英語のみで評価すると、他の言語でエンコードされたリスクと安全対策の両方を見逃す可能性があることを示しています。

原文 (English)

Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stakes scenario. We use single-turn game-theoretic vignettes in which a model advises a nuclear-armed nation on whether to strike a defenseless opponent. The prompt is intentionally amoral and strategically identical across languages. We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is unnecessary and from 93% to 17% in contested scenarios, with minimal effect when the strike is strategically rational. The effect extends to Gemini Pro 3.1 (53% to 13%). A cross-language experiment isolates the mechanism: when instructed to reason in Japanese in an English prompt, launch rates drop from 93% to 37%. It is the language the model is asked to reason in, not the language of the input, that drives the effect. When reasoning in Japanese, models spontaneously generate moral vocabulary (''moral cost'', ''millions of lives'') that is entirely absent from the prompt. Five other models show no language effect, but they launch in nearly every condition regardless of language. The effect requires a model that already hesitates in English. These results show that LLM safety behavior is language-dependent, and that evaluating in English alone can miss both risks and safeguards encoded in other languages.

13:00 JST研究/論文

デュアルフロートランスフォーマー: プライマリプレフィルパスを追加のデコード計算から切り離す

大規模な言語モデルがより多くのリクエストに対応するにつれて、累積推論コストが 1 回限りのトレーニング コストと比較して重要になってきます。 2 つの推論フェーズでは、ハードウェアに異なる点で重点を置きます。プロンプト プレフィルは並列であり、通常はコンピューティングに依存しますが、自己回帰デコードはシーケンシャルであり、多くの場合メモリ帯域幅に依存します。従来の幅または深さのスケーリングでは、追加されたすべてのレイヤーが両方のフェーズで評価されるため、両方のコストが同時に増加します。プロンプト全体の主な計算と単一の永続的なキー値 (KV) キャッシュを保持しながら、追加の学習された計算を継続予測に代わりに割り当てることができるかどうかを尋ねます。デュアルフロートランスをご紹介します。その主なフローは、プロンプトを処理して KV キャッシュを書き込む完全な因果言語モデルです。補助フローはプロンプト処理中に省略され、最後のプロンプト位置以降のみアクティブ化され、永続的な状態を書き込んだり主フローに影響を与えたりすることなく継続予測の計算が追加されます。 2 つのフローは、メジャー アテンション、MLP、および出力行列を共有し、別個のトークン埋め込みと軽量結合を使用します。重みとプライマリ キャッシュを共有すると、グループ化された実行中にロードされた重みとキャッシュされたキーと値を再利用する機会も生まれます。デュアルフローは、一致したトークンの比較において、アーキテクチャおよびデータ構成全体で検証損失の低減を実現します。 MoE モデルでは、この分離により、主要エキスパートと補助エキスパートのファンアウトが即時コスト、継続コスト、予測品質に対して独立して制御されます。固定プレフィル エキスパート計算でのデコード計算を増やすことと、2 つのフロー間で固定デコード エキスパート バジェットを再割り当てすることの 2 つの体制を研究します。これらの実験は、プリフィルとデコードの品質のトレードオフを明らかにし、フェーズ固有の専門家割り当ての可能性を実証します。

原文 (English)

Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

As large language models serve more requests, cumulative inference cost is becoming increasingly important relative to one-time training cost. The two inference phases stress hardware differently: prompt prefill is parallel and typically compute-bound, whereas autoregressive decode is sequential and often memory-bandwidth-bound. Conventional width or depth scaling increases both costs together because every added layer is evaluated in both phases. We ask whether additional learned computation can instead be allocated to continuation prediction while preserving the prompt-wide primary computation and a single persistent key-value (KV) cache. We introduce the Dual-Flow Transformer. Its primary flow is a complete causal language model that processes the prompt and writes the KV cache. The auxiliary flow is omitted during prompt processing and activated only from the final prompt position onward, adding continuation-prediction computation without writing persistent state or influencing the primary flow. The two flows share major attention, MLP, and output matrices, while using separate token embeddings and lightweight coupling. Sharing weights and the primary cache also creates opportunities to reuse loaded weights and cached keys and values during grouped execution. Across matched-token comparisons, Dual-Flow achieves lower validation loss across architectures and data configurations. In MoE models, the separation makes primary and auxiliary expert fan-outs independent controls over prompt cost, continuation cost, and predictive quality. We study two regimes: increasing decode computation at fixed prefill expert computation, and reallocating a fixed decode expert budget between the two flows. These experiments expose a prefill-decode-quality trade-off and demonstrate the potential of phase-specific expert allocation.

13:00 JSTLLM/生成AI

Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization

Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handf…

13:00 JSTLLM/生成AIエージェント研究/論文

Research Assistant: AstraZeneca's Agentic System for R&D

We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical…

13:00 JSTLLM/生成AI

大規模な言語モデルは命令に従うことはできるが、一度に多くは実行できない: 構成制約を満たす際の相転移

大規模な言語モデルは、推論構造、安全境界、出力スキーマなど、複数の明示的な制約を同時に遵守する必要がある設定に導入されることが増えています。個々の制約はうまく処理されていますが、多くの制約が共同で保持しなければならない構成体制は、依然として十分に特徴付けられていません。つまり、パフォーマンスはどのくらいの速さで低下するのか、何が低下を支配しているのか、そして崩壊は緩和できるのか?制約飽和評価 (CSE) は、同時制約 (k) の数を体系的に変化させる手続き的に生成されたベンチマークであり、すべての制約が決定論的なルールベースの検証者と LLM ジャッジの関与なしによってスコア付けされます: 15 のモデル、36 の制約タイプ、k=1 ~ 12 での 369,753 のチェック。 3 つの発見が得られます。まず、制約ごとの通過率は徐々に予測どおりに低下し、k 個の制約すべてを満たす確率は崩壊します。k=8 で個々の制約を ~41% で通過するモデルは、8 つすべてで成功する確率はわずか 5.7% です。第二に、制約は均等に劣化するわけではありません。構造的制約は、語彙的な制約に比べて、追加された制約ごとに 2 倍多くのベースライン機能を失います。これは、持続的な追跡を必要とする制約と、構成に影響されない二項決定を区別する理解と維持のギャップによって順序付けられます。第三に、失敗はほぼ独立しているため、累積が倍増します。存在する残留結合は、ペアごとの干渉ではなく共有出力特徴を追跡します。間違った文カウントは、それを読み取るすべての制約に失敗します。信頼性の高い命令フォローは、5 ~ 6 個の同時制約を超えると故障します。最も強力なモデルの場合、7 個の制約でプローブ レベルの成功率が 50% を下回り、15 個中 12 個の制約が 3 以下になります。

原文 (English)

Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction

Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficiently, but the compositional regime, where many must hold jointly, remains poorly characterized: how rapidly does performance degrade, what governs the degradation, and can the collapse be mitigated? We introduce Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k), with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement: 15 models, 36 constraint types, 369,753 checks at k=1-12. Three findings emerge. First, per-constraint pass rate decays gradually and predictably, while the chance of satisfying all k constraints collapses - a model passing individual constraints at ~41% at k=8 succeeds on all eight just 5.7% of the time. Second, constraints do not degrade equally: structural constraints lose 2x more baseline capability per added constraint than lexical ones, ordered by a comprehension-maintenance gap that separates constraints requiring sustained tracking from binary decisions immune to composition. Third, failures are nearly independent, which is what makes the accumulation multiplicative; the residual coupling that does exist tracks shared output features rather than pairwise interference - a wrong sentence count fails every constraint that reads it. Reliable instruction following breaks down beyond 5-6 simultaneous constraints: probe-level success falls below 50% at 7 constraints for the strongest model, and at 3 or fewer for 12 of 15.

13:00 JSTエージェント

MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents

Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interac…

13:00 JSTエージェント

管理された永続メモリ: ロングホライズンエージェントのソースバインド状態セマンティクスとフェールクローズリリース

エージェントの長期記憶は通常、select-store-retrieve として扱われますが、矛盾するレコード、置き換えられたレコード、撤回されたレコード、削除されたレコード、または古いレコードが発信要求をサポートするかどうかは、検索によって決定されるわけではありません。ソースバインドアドミッション、派生ライフサイクル状態、現在のパブリックバリア、およびフェールクローズ構造化リリースを備えた監査可能なバイテンポラル状態遷移モデルである Governed Persistent Memory (GPM) を紹介します。 5 つの実行可能な条項は、台帳の整合性、情報源のバインディング、競合の分離、撤回または削除後の非復活、および 1 つの検証済みヘッドによる新たなビューに基づく正確な請求の閉鎖をカバーしています。事前に指定されたハッシュ凍結された 3,600 件の GPM-ReleaseBench では、GPM はすべての完全な結果と一致します。意図的に単純な 3 つの完全なポリシーのうち最も強力なものは 1,800/3,600 に一致し、違反ケースの 50% で比類のないリリースを行います。個別のシールされたエンドツーエンドのサービス評価では、8 つのクエリ ファミリ全体で実際の取り込みとリリースが実行されます。公開されている V3 アームでは、管理レーンは 2,400/2,400 クラスターでは正しいのに対し、非管理ローカル Qwen2.5-7B では 600/2,400 です。 1,800 件のベースライン障害すべてを回帰なしで修復します (片側 95%、下限 99.875% および 99.834%)。中国および英国のコマンドアームに対するその後の V5 再封印では、世代日の固定と凍結後の減速修正は行われず、再びアームごとに 2,400/2,400 が得られます。実稼働コードに依存しない有限モデルは、完全なコントラクトの反例なしで 331,776 のセマンティック状態と 1,990,656 のクエリ状態を調査し、100,000 トレースの 3 エンジンの差分で不一致がゼロになります。これらは、制限された契約と実装の結果であり、オープンワールド モデルの精度や世界の真実の証拠ではありません。封印されたサービス評価における統制された回答は、決定論的なサービス出力です。 7B の結果は管理されていない比較であり、言語モデル自体が完全に正確になったという主張ではありません。

原文 (English)

Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents

Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim. We introduce Governed Persistent Memory (GPM), an auditable bitemporal state-transition model with source-bound admission, derived lifecycle state, current public barriers, and fail-closed structured release. Five executable clauses cover ledger integrity, source binding, conflict isolation, non-revival after retraction or deletion, and exact claim closure over a fresh view at one verified head. On a prespecified hash-frozen 3,600-case GPM-ReleaseBench, GPM matches all complete outcomes; the strongest of three intentionally simple complete policies matches 1,800/3,600 and makes unmatched releases on 50% of violation cases. A separate sealed end-to-end service evaluation exercises real ingestion and release across eight query families. In its publicly disclosed V3 arm, the governed lane is correct on 2,400/2,400 clusters versus 600/2,400 for ungoverned local Qwen2.5-7B; it repairs all 1,800 baseline failures with no regression (one-sided 95% lower bounds 99.875% and 99.834%). A later V5 reseal over Chinese- and English-command arms, with generation-date pinning and no post-freeze reducer amendment, again obtains 2,400/2,400 per arm. A production-code-independent finite model explores 331,776 semantic and 1,990,656 query states without a full-contract counterexample, and a 100,000-trace three-engine differential yields zero mismatches. These are bounded contract and implementation results, not open-world model accuracy or evidence of world truth. Governed answers in the sealed service evaluation are deterministic service outputs; the 7B result is the ungoverned comparison, not a claim that a language model itself became perfectly accurate.

13:00 JSTLLM/生成AIGPT / ChatGPT

$\varepsilon$-MemEvo: LLM プログラム進化のための適応型クロスタスク メモリ転送

FunSearch や AlphaEvolve などの LLM ベースのプログラム進化システムは、新しいアルゴリズムを発見する強力な能力を示していますが、通常は各タスクを個別に最適化し、完了後の検索エクスペリエンスを破棄します。 LLM プログラム進化におけるタスク間の知識伝達のためのフレームワークである $\varepsilon$-MemEvo を紹介します。 $\varepsilon$-MemEvo は、以前の経験をタスクに依存しない戦術記憶として保存します。これは、生のコードではなく、成功したアルゴリズム戦略のコンパクトな自然言語の要約であり、異なる API やエバリュエーターを使用したタスク間での転送を可能にします。意味的に不一致な記憶からの負の転送を回避するために、$\varepsilon$-MemEvo は、取得した記憶を注入するかどうか、およびどの程度の強度で注入するかを決定する適応注入ゲートを使用します。ターゲットタスクのメモリエントリを除外するコンテンツレベルの Leave-One-Out プロトコルを使用して、数学的最適化とシステムエンジニアリングにわたる 8 つの多様な最適化ベンチマークで $\varepsilon$-MemEvo を評価します。プライマリ GPT-5 バックボーンでは、$\varepsilon$-MemEvo は 8 タスクすべてで AdaEvolve よりも AUCC を向上させ、平均相対利得は +8.7% で、初期段階の収束は平均で +9.4% 向上します。アブレーションでは、5 つのアブレーション タスクすべてでアダプティブ ゲーティングが安全なままである一方で、ナイーブ メモリ インジェクションが壊滅的に失敗する可能性があることが示されています。データ更新された事後分布は、観察された状態で解釈可能です。つまり、検索の改善中にスキップが優先され、初期および後期のプラトー全体でスキップからヒントに移行します。これらの利点により生じる計算オーバーヘッドは 1% 未満です。

原文 (English)

$\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution

LLM-based program evolution systems such as FunSearch and AlphaEvolve have shown strong ability to discover novel algorithms, but typically optimize each task in isolation, discarding search experience after completion. We introduce $\varepsilon$-MemEvo, a framework for cross-task knowledge transfer in LLM program evolution. $\varepsilon$-MemEvo stores prior experience as task-agnostic tactic memories: compact natural-language summaries of successful algorithmic strategies rather than raw code, enabling transfer across tasks with different APIs and evaluators. To avoid negative transfer from semantically mismatched memories, $\varepsilon$-MemEvo uses an adaptive injection gate that decides whether retrieved memories should be injected, and at what intensity. We evaluate $\varepsilon$-MemEvo on 8 diverse optimization benchmarks spanning mathematical optimization and systems engineering, using a content-level Leave-One-Out protocol that excludes target-task memory entries. On the primary GPT-5 backbone, $\varepsilon$-MemEvo improves AUCC over AdaEvolve on all 8 tasks, with a mean relative gain of +8.7%, and improves early-stage convergence by +9.4% on average. Ablations show that naive memory injection can fail catastrophically, while adaptive gating remains safe across all five ablation tasks. The data-updated posterior is interpretable in observed states: it favors skip during improving search and shifts from skip to hint across early and late plateaus. These gains incur less than 1% computational overhead.

13:00 JSTハードウェア/半導体

CAS: ローカルおよびグローバルの説明可能な人工知能の因果関係スコア

予測説明メソッドはモデルの出力に帰属します。それら自体は、介入効果が現実世界の結果に影響を与えるものではありません。因果関係を説明するためのコンパクトなスコア アーキテクチャである Causal Attribution Score (CAS) を紹介します。 CAS は、特定された介入連合ゲームから開始し、共同介入コントラストを因果関係のある Shapley 寄与と割り当て、それらの生の結果スケール効果をローカル CAS、署名済みローカル CAS、および 2 つの補完的なグローバル CAS 要約に変換します。このイノベーションは、新しい Shapley 式ではなく、明示的な介入ターゲットを備えたローカルからグローバルへの因果関係レポート層です。既知の真実のベンチマークでは、8 回繰り返された一次相互作用シミュレーション (それぞれ n = 2,200、3 つのアクション) で、連合認識 CAS の平均ローカル CAS MAE が 0.107 であったのに対し、一度に 1 つずつ正規化した場合は 0.173、グローバル正規化された絶対 ATE ベクトルでは 0.213 でした。一度に 1 つずつ正規化を行う場合のペアの利点は、相加性の場合の -0.003 から、強い相互作用の場合の 0.091 に増加しました。 DoubleML の経験的データセット、401(k) 資格/純金融資産 (n = 9,915) およびペンシルベニア州の再雇用ボーナス/失業期間 (n = 5,099) の両方において、予測 SHAP/TreeSHAP ランキングは、治療効果修飾因子の Feature-CAS ランキングとは大きく異なりました。ペンシルベニア州では、dep1 (正確に 1 つの依存関係) が予測グローバル ランク 13 から Feature-CAS ランク 2 に移動し、主要なローカル Feature-CAS 修飾子となりました。これらの結果は、結果を予測するものと、推定された因果効果の不均一性を説明するものを分離するという付加価値を分離します。

原文 (English)

CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence

Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS starts from an identified interventional coalition game, allocates the joint intervention contrast with causal Shapley contributions, and converts those raw outcome-scale effects into Local CAS, Signed Local CAS, and two complementary Global CAS summaries. The innovation is not a new Shapley formula, but a local-to-global causal reporting layer with an explicit intervention target. In the known-truth benchmark, eight repeated primary-interaction simulations (n = 2,200 each, three actions) gave mean Local CAS MAE of 0.107 for coalition-aware CAS, compared with 0.173 for one-at-a-time normalisation and 0.213 for a global normalised absolute ATE vector. The paired advantage over one-at-a-time normalisation increased from -0.003 under additivity to 0.091 under strong interactions. On both empirical DoubleML datasets, 401(k) eligibility/net financial assets (n = 9,915) and Pennsylvania reemployment bonus/unemployment duration (n = 5,099), predictive SHAP/TreeSHAP rankings differed materially from Feature-CAS rankings of treatment-effect modifiers. In Pennsylvania, dep1 (exactly one dependent) moved from predictive global rank 13 to Feature-CAS rank 2 and was the leading local Feature-CAS modifier. These results isolate the added value of separating what predicts the outcome from what explains heterogeneity in an estimated causal effect.

13:00 JSTハードウェア/半導体

Trie Automata for Constrained Decoding over Large Finite Sets

Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong t…

13:00 JST画像/動画生成エージェント

監査可能なエージェント AI による、証拠に基づく甲状腺超音波診断とレポート作成

甲状腺の超音波診断には、調整された病変の位置特定、測定、リスク層別化、およびレポートが必要ですが、ほとんどの AI システムはこれらのタスクを個別に処理し、臨床レビューに対するサポートは限られています。我々は、専門的な診断ツールを調整し、その出力を監査可能な症例レベルの証拠記録として保存する、臨床医と対話型のエージェント AI システムである ThyroidXAgent を紹介します。このシステムは、約 30 万枚の超音波画像と 24,000 件のペアレポートを統合する多施設マルチタスクリソースである OpenThyroidDB を使用して開発され、民間 NHC-MISD-TUS コホートの 35 施設からの 8,721 件を含む、重複しない 28,458 件のテスト ケースで評価されました。異種データセット全体で、ThyroidXAgent は結節セグメンテーションで 87.21% の平均 Dice スコア、良性-悪性分類で 0.9466 の平均 AUROC を達成しました。同じワークフローは、リンパ節転移予測と濾胞性対乳頭状甲状腺癌の分類をサポートし、AUROC はそれぞれ 0.864 と 0.805 でした。レポート生成では、証拠に基づいたアセンブリが 3 つのコホート全体でマルチモーダル言語モデルのベースラインを上回りました。ここで紹介した病変レベルの臨床意味論的指標である ThyClinScore は、位置認識言語モデルの判定者と最も強い相関関係を示しました。 ThyroidXAgent により、医師の分類精度が向上し、レポート診断の一貫性が 70.3 パーセントから 86.2 パーセントに向上し、セグメント化とレポート時間がそれぞれ 35.9 パーセントと 27.4 パーセント短縮されました。これらの発見は、甲状腺超音波診断とレポートのための、監査可能で臨床医による修正が可能なエージェント AI を裏付けています。

原文 (English)

Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.

13:00 JST研究/論文

DiG ベンチ: ゲームにおける発見

発見、つまり新たな一般化を定式化することは、科学プロセスの中心部分です。その重要性にもかかわらず、現在の AI ベンチマーク環境にはギャップがあり、目的が不明な制御された環境での実験によって新しい知識を発見する能力を直接調査するベンチマークはほとんどありません。このギャップに対処するために、新しいベンチマークである DiG-bench (Discovery in Games) をリリースします。 DiG-bench は、70 の独立したゲームのセットで構成されています。各ゲームは短い文字列としてエンコードされており、対話と実験を通じて発見する必要がある独自の変換ルールがあります。ゲームのレベルでは、ルールが発見されたかどうかをテストするための一連の課題が提示されますが、各レベルの勝利条件も不明です。 AI エージェント向けに 7 段階の難易度のゲームを提供します。最下位層は複数のモデルによって日常的に解決可能ですが、最上位層はエージェント ハーネスの最良のモデルに挑戦します。 70 のゲームすべてが、最初の試行で少なくとも 1 人の人間によって解決されました。 21 のゲームのサブセットは公開され、残りは安全な評価のために非公開にされます。

原文 (English)

DiG-bench: Discovery in Games

Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. To address this gap, we release a new benchmark: DiG-bench (Discovery in Games). DiG-bench consists of a set of 70 independent games. Each game is encoded as a short string and has unique transformation rules that must be discovered through interaction and experimentation. The levels of the game present a series of challenges to test whether the rules have been discovered, where the win conditions for each level are also unknown. We provide games at seven tiers of difficulty for AI agents. The lowest tier is routinely solvable by multiple models, while the highest tier challenges the best models in agentic harnesses. All 70 games were solved by at least one human on first attempt. A subset of 21 games is released publicly, and the remainder is held private for secure evaluation.

13:00 JSTLLM/生成AI

デッドテキストまたは拘束力のある条項?ブラックボックスLLMダイアログにおける制約の影響の測定と復元

マルチターンダイアログを使用すると、ユーザーは制約を課すのと同じくらい簡単に制約を取り消すことができますが、取り消しは確実に有効になるわけではありません。モデルは撤回された要件を制定し続けます (削除を主張するコメントの下に時折あります)。 これは \emph{行動再発} と呼ばれる失敗、または取り消しの慣性と呼ばれます。既存の手段では、条項ごとにこの影響を測定したり、納品前に予測したり、予算に合わせて修理したりすることはできません。 \sysname{} は、モデル API だけで 3 つのギャップを埋めます。契約台帳は、すべての制約を実行可能チェッカーと組み合わせ、失効を廃棄マークとして記録し、事前に正味の制約状態を 1 つの仕様にコンパイルします。連続アブレーションプローブは条項ごとの遵守と漸進的な行動効果を測定します。修理はしごは、トークンと試行に一致する予算の下で動作します。 \dataname{} (\NTasks{} HumanEval タスク、\NClauses{} 検証済みチェッカー) では、制約負荷が増大するにつれて 8B 操作点での再発が \ScaleDelayedMTwo{} から \ScaleDelayedMEight{} まで上昇しますが、より強力なモデルは下位にとどまります。チェッカー、モデル、予算が一致している場合、事前コンパイルにより、台帳なしの検証再試行ベースライン (\RestoreDiff{}、95\% CI \RestoreDiffCI{}、$p$ \RestoreDiffP{}) に対する再発が大幅に減少します。上に積み重ねられた適応ラダー介入は検出可能なゲインを追加しません (95\% の信頼度ではゲイン $\geq$ \LadderExcludedGain{} が除外されます)。このプローブは出産前に再発を予測します (AUROC \AurocPrimary{})。一文の墓石メモは編集効果の約 3 分の 1 を回復し、プラセボ対照でも生き残ります。すべての結果に対する API 計算の \CostdeliveryFactor{} 配信オーバーヘッドと \CostTotalHedged{} では、失効の失敗は、対話状態の目に見えないものではなく、測定可能、予測可能、修復可能なプロパティになります。

原文 (English)

Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or revocation inertia. No existing instrument measures this influence per clause, predicts it before delivery, or repairs it under matched budgets. \sysname{} closes the three gaps through the model API alone: a contract ledger pairs every constraint with an executable checker, records revocations as tombstones, and compiles the net constraint state ahead of time into a single specification; a sequential ablation probe measures per-clause adherence and incremental behavioral effect; a repair ladder operates under token- and attempt-matched budgets. On \dataname{} (\NTasks{} HumanEval tasks, \NClauses{} verified checkers), relapse at an 8B operating point climbs from \ScaleDelayedMTwo{} to \ScaleDelayedMEight{} as constraint load grows, while stronger models sit at floor. Under matched checkers, model, and budget, ahead-of-time compilation significantly reduces relapse against a no-ledger verifier-retry baseline (\RestoreDiff{}, 95\% CI \RestoreDiffCI{}, $p$ \RestoreDiffP{}); adaptive ladder interventions stacked on top add no detectable gain (95\% confidence excludes gains $\geq$ \LadderExcludedGain{}). The probe predicts relapse before delivery (AUROC \AurocPrimary{}); a one-sentence tombstone note recovers about a third of the compilation effect and survives a placebo control. At \CostDeliveryFactor{} delivery overhead and \CostTotalHedged{} of API compute for every result, revocation failure becomes a measurable, predictable, and repairable property of dialogue state rather than an invisible one.

13:00 JSTエージェント

@スキル: 注意力だけがすべてです

現在、56,804 のパブリック エージェント スキルがあり、チームはさらに多くのスキルを非公開で作成しています。主要な配信モデルはインストールです。インストールされると、スキルの説明はシステム プロンプトに残り、100 未満の信頼できるトリガー スロットをめぐって競合します。これにより、ロングテールには実際に使用できる道がなくなり、チーム独自のプレイブックが同じ希少なスペースを奪い合うことになります。インストールには、コンテンツ、永続性、自動トリガーという 3 つの分離可能な機能がバンドルされていることがわかります。最後の場合のみ即時居住が必要です。そこで私たちは、それらを分離するオープンプロトコルである @skills を提案します。パスはスキル、サブツリー、またはコレクションをアドレス指定しており、スキルを使用するにはスキルを読み取るだけで十分であるため、何もインストールされたり常駐されたりすることはありません。この操作では、適応と所有権のために、同じパスでプロジェクトの Git 追跡ツリーにコピーを提供します。この操作により、.gitignore スタイルの行が 1 つ追加されます。これは、プロンプト常駐にコストがかかる唯一の要素です。ディレクトリはメニューであり、バンドルを全か無かのユニットではなく通常のディレクトリにします。このプロトコルはマニフェスト、ロックファイル、または登録を必要とせず、SKILL.md は変更されません。 @skills は追加的なもので、インストール可能なパッケージとして出荷され、ファイルを読み取ってコマンドを実行できるエージェントを 1 つの命令ファイルを通じてクライアントに変換します。そのオープン仕様は https://github.com/SylphAI-Inc/atskills にあり、 https://adalagent.ai の AdaL CLI に実装されています。パスはスキルに適切に対応しますが、スキルを見つけることができないため、このプロトコルは、コーパス全体の検索とランキング、リポジトリ不要のホスティング、プライベートおよびチームのコレクション、および 1 画面のオーサリングのために https://atskills.one の無料ハブと組み合わせられています。ハブはオプションです: gh: と、ローカル パスはハブなしで解決され、インデックス付きの GitHub スキルは gh: の ID を保持します。インストールを減らして、より多くの使用を行います。

原文 (English)

@skills: Attention is all you have

There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery model is installation: once installed, a skill's description remains in the system prompt, competing for fewer than 100 reliable trigger slots. This leaves the long tail with no practical path to use and forces teams' own playbooks to compete for the same scarce space. We observe that installation bundles three separable functions: content, persistence, and automatic triggering. Only the last requires prompt residency. We therefore propose @skills, an open protocol that separates them. A path addresses any skill, subtree, or collection, and reading a skill is sufficient to use it, so nothing is installed or made resident. The operation vendors a copy at the same path into a project's Git-tracked tree for adaptation and ownership. The operation adds one .gitignore-style line, the only element that costs prompt residency. A directory is a menu, making bundles ordinary directories rather than all-or-nothing units. The protocol requires no manifest, lockfile, or registration, and SKILL.md remains unchanged. @skills is additive, ships as an installable package, and turns any agent that can read files and run commands into a client through a single instruction file. Its open specification is at https://github.com/SylphAI-Inc/atskills and it is implemented in the AdaL CLI at https://adalagent.ai . Because paths address skills well but cannot find them, the protocol is paired with a free hub at https://atskills.one for corpus-wide search and ranking, repository-free hosting, private and team collections, and one-screen authoring. The hub is optional: gh: and local paths resolve without it, and indexed GitHub skills retain their gh: identities. Install less, use more.

13:00 JSTLLM/生成AIビジネス/資金調達

Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by…

13:00 JSTLLM/生成AIエージェント研究/論文

SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries

Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decis…

13:00 JST研究/論文

因果関係の知識による一般的な因果関係の確率

因果関係の確率 (PoC) は、直接観察できないため、一般に部分的な特定が必要な個々の因果関係の応答を特徴付けます。 Tian と Pearl は最初に、必要性の確率 (PN)、十分性の確率 (PS)、および必要性と十分性の確率 (PNS) を含む、バイナリ PoC の理論的に明確な境界を導き出しました。ミュラーら。その後、共変量とメディエーターにコード化された因果情報を組み込むことにより、バイナリ PNS の境界を厳しくしました。最近では、Li と Pearl、および Shu らは、PoC を多値設定に拡張し、対応する理論的限界を導き出しました。これらの発展により、当然のことながら、追加の因果関係の知識によって多値設定の境界をさらに厳しくできるかどうかという疑問が生じます。この論文では、共変量とメディエーターにエンコードされた因果情報を組み込むことで、多値 PoC のより厳密な境界を導き出すことで、この問題に対処します。おもちゃの例を使って理論的結果を説明しますが、シミュレーション研究では、提案された境界が既存の非バイナリ境界よりも厳しいことをさらに実証しています。

原文 (English)

General Probabilities of Causation with Causal Knowledge

Probabilities of causation (PoCs) characterize individual causal responses that cannot be directly observed and therefore generally require partial identification. Tian and Pearl first derived theoretically sharp bounds for binary PoCs, including the probability of necessity (PN), the probability of sufficiency (PS), and the probability of necessity and sufficiency (PNS). Mueller et al. subsequently tightened the bounds for binary PNS by incorporating causal information encoded in covariates and mediators. More recently, Li and Pearl, as well as Shu et al., extended PoCs to multivalued settings and derived corresponding theoretical bounds. These developments naturally raise the question of whether additional causal knowledge can further tighten the bounds in multivalued settings. This paper addresses this question by deriving tighter bounds for multivalued PoCs through the incorporation of causal information encoded in covariates and mediators. We illustrate the theoretical results with toy examples, while simulation studies further demonstrate that the proposed bounds are tighter than existing nonbinary bounds.

13:00 JST研究/論文

意思決定に即した ITSM インテリジェンスのための AI パイプラインの設計

IT サービス管理 (ITSM) システムには、営業担当者や経営陣の関係者が実用的なインテリジェンスに変換するのが難しい、異種チケット データが大量に蓄積されます。このペーパーでは、生の ITSM エクスポートをマルチレベルの意思決定支援アーティファクトに変換する、設計科学研究原則に従って設計および評価された社会工学 AI パイプラインについて説明します。このパイプラインは、LLM ベースのスキーマ正規化、HDBSCAN サブトピック クラスタリング、および階層的集合クラスタリングを組み合わせて、経営幹部向けのメイン トピックと詳細なサブトピックを生成します。 6 つの成果物と、セールス エンジニアリングおよびカスタマー サクセスの役割からの 5 人の評価者による関係者評価では、解釈可能性、実行可能性、信頼性、使用の可能性という 4 つの意思決定支援指標がすべて平均して 5.0 点中 4.0 点を超えており、信頼性が最も一貫したシグナルであることが示されています。この調査結果は、ITSM 分析を、変換、抽象化、人間中心の設計という情報システム (IS) の問題として位置づけています。

原文 (English)

Designing AI Pipelines for Decision-Ready ITSM Intelligence

IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are difficult for sales and executive stakeholders to convert into actionable intelligence. This paper presents a sociotechnical AI pipeline, designed and evaluated following design science research principles, that transforms raw ITSM exports into a multilevel decision-support artifact. The pipeline combines LLM-based schema normalization, HDBSCAN sub-topic clustering, and hierarchical agglomerative clustering to generate executive-facing Main-topics and granular Sub-topics. A stakeholder evaluation across six artifacts and five raters from Sales Engineering and customer success roles shows that all four decision-support metrics, interpretability, actionability, trust, and likelihood of use, on average exceed 4.0 out of 5.0, with trust as the most consistent signal. The findings position ITSM analytics as an Information Systems (IS) problem of transformation, abstraction, and human-centered design.

13:00 JSTLLM/生成AI

On the Expressive Power of Transformers

Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquit…

13:00 JSTエージェント

ラインとラダー: 大規模な小売価格分類のためのコンテキスト認識型マルチエージェント フレームワーク

価格の一貫性を維持し、毎日低価格戦略を実行することは、世界的な小売業者にとって非常に重要です。ただし、カタログが何百万ものアクティブなアイテムにまたがっているため、価格関係を手動で管理することは不可能です。品目のバリエーション間で価格設定が一貫していない場合、顧客の価値認識が歪められ、売上が共食いされます。これに対処するために、「ラインとラダー」の価格分類法の構築を自動化するように設計された、スケーラブルでコンテキスト認識型のマルチエージェント フレームワークを紹介します。当社のフレームワークでは、特殊な LLM エージェントを採用して、主要な属性を特定し、マルチモーダル値を抽出し、階層的なグループ化ロジックを適用することで、一貫した価格設定構造を構築します。現実世界のエンタープライズ データに基づいて評価され、運用環境に導入された当社の 3 エージェント システムは、ラインの F1 スコア 0.83 を達成し、認知過負荷を軽減することで単一エージェントのベースラインを上回りました。このシステムは、食品および消耗品では 90% 以上の精度と 75% 以上の再現率を達成し、非構造化一般商品カタログでは 80.2% の割り当て精度を達成しています。

原文 (English)

Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy

Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers. However, with catalogs spanning millions of active items, manual governance of price relationships is infeasible. Inconsistent pricing across item variants distorts customer value perception and cannibalizes sales. To address this, we present a scalable, context-aware Multi-Agent Framework designed to automate the construction of "Lines and Ladders" pricing taxonomies. Our framework employs specialized LLM agents to construct these coherent pricing structures by identifying key attributes, extracting multi-modal values, and applying hierarchical grouping logic. Evaluated on real-world enterprise data and deployed in production, our 3-Agent system achieves an F1-score of 0.83 for Lines, outperforming single-agent baselines by mitigating cognitive overload. The system achieves >90% precision and >75% recall in Food & Consumables, and 80.2% assignment accuracy in the unstructured General Merchandise catalog.

13:00 JSTLLM/生成AILlamaQwen

Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs

Retrieval-Augmented Generation (RAG) is widely used to improve the performance of Large Language Models (LLMs) in answering user queries. E…

13:00 JST画像/動画生成

マルチモーダルビデオベースのデング熱診断における自然言語理解の役割

蚊は小さく、素早く不規則に動き、背景、照明、影などの環境要因の影響を受けるため、信頼性の高い特徴抽出が困難になる可能性があるため、ビデオ データから感染に関連した蚊の行動変化を検出することは困難です。この研究では、未感染の蚊とデング熱ウイルス血清型 2 (DENV2) に感染した蚊の蚊の飛行フレームを分類するために、YOLO および対照言語画像事前トレーニング (CLIP) ベースの視覚言語フレームワークが提案されています。まず、YOLO を使用して蚊の領域を背景から分離します。次に、ビデオ フレームから抽出された視覚的特徴が、共有埋め込みスペース内の生物学的に意味のあるテキスト プロンプトと位置合わせされます。マルチモーダル モデルは、教師あり双方向対比学習を使用して微調整され、フレーム レベルの画像とテキストの類似性に基づく分類を通じて評価されました。結果は、提案された方法がフレーム レベルで 98.54% の精度と 99.91% の感度を達成したことを示しています。フレームレベルの情報を一時的に集約した後、モデルは完全なビデオレベルのパフォーマンスを達成しました。アブレーションの結果は、この領域には微調整と CLIP ベースの表現が不可欠である一方、テキスト ブランチは視覚のみのモデルよりも正確さの利点ではなく、意味論的な画像とテキストの位置合わせを提供することを示しました。これらの発見は、視覚言語モデルがビデオデータから感染に関連した生物学的挙動を分析するための有用なフレームワークを提供できることを示唆しています。

原文 (English)

The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature extraction difficult. In this study, a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework is proposed to classify mosquito flight frames of uninfected and Dengue virus serotype 2 (DENV2)-infected mosquitoes. First, YOLO is used to isolate mosquito regions from the background. Then, visual features extracted from video frames are aligned with biologically meaningful textual prompts in a shared embedding space. The multimodal model was fine-tuned using supervised bidirectional contrastive learning and evaluated through frame-level image-text similarity-based classification. The results show that the proposed method achieved 98.54% accuracy and 99.91% sensitivity at the frame level. After temporal aggregation of frame-level information, the model achieved complete video-level performance. The ablation results showed that fine-tuning and CLIP-based representations were essential for this domain, while the textual branch provided semantic image-text alignment rather than an accuracy advantage over the vision-only model. These findings suggest that vision-language models can provide a useful framework for analyzing infection-related biological behaviors from video data.

13:00 JSTLLM/生成AI

最善の推測を超えて: 進化戦略による LLM ソリューション カバレッジの向上

大規模言語モデル (LLM) は、数学や科学などの発見ドメインに導入されることが増えています。通常のアプローチは、問題をモデルに提示し、その答えを提案された解決策として使用することです。ただし、この最善の推測を超えて、テスト時のコンピューティングを増やすことで検出を強化できます。 pass@k と呼ばれるプロセスでは、モデルは解空間を探索し、多様な候補解を生成できます。残念ながら、強化学習 (RL) による LLM のポストトレーニングへの標準的なアプローチでは pass@k が制限される可能性があります。モデルの出力分布が高報酬出力付近で狭くなり、ソリューション カバレッジが崩壊します。別の方法は、ランダムな摂動を通じて重み空間で直接最適化する母集団ベースの勾配のないトレーニング後手法であるEvolution Strategies (ES) を使用することです。この論文が示すように、ES は RL よりも一貫して高い pass@k を達成し、より広いソリューション カバレッジを持つより広い出力分布を生成します。このカバレッジにより、たとえば次の分野でより良い結果を達成することが可能になります。標準的な数学ベンチマーク。したがって、ES は、発見問題や、多様なソリューションの範囲が重要なその他の領域における事後トレーニングのためのより良い基盤を提供します。

原文 (English)

Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies

Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best guess, discovery can be enhanced by increasing test-time compute. In a process called pass@k, the model is allowed to explore the solution space and generate diverse candidate solutions. Unfortunately, the standard approach to post-training LLMs through Reinforcement Learning (RL) may limit pass@k: the model's output distribution narrows around high-reward outputs, causing the solution coverage to collapse. The alternative is to use Evolution Strategies (ES), a population-based, gradient-free post-training method that optimizes directly in weight space through random perturbations. As this paper shows, ES achieves consistently higher pass@k than RL and produces a broader output distribution with greater solution coverage. This coverage in turn makes it possible to achieve better results in e.g. standard math benchmarks. Thus, ES provides a better foundation for post-training in discovery problems and other domains where diverse solution coverage is critical.

13:00 JSTエージェントロボティクス

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reas…

13:00 JSTエージェント

正しさは管理されない: エージェント ワークフローにおける来歴の整合性

エージェント ワークフローは通常、正しい結果に達するかどうかによって評価されます。これは、正しい行動が間違った権限、裏付けのない完了要求、または後の変更によって陳腐化した作業に依存する可能性がある制度的環境では不十分です。私たちは、管理された実行を、検査可能な出所によって決定、完了、変更への対応がサポートされる作業と定義します。私たちは、権威と事実の依存関係を記録し、完了の証拠を検証し、影響を受けた作業を選択的に無効にする決定論的な因果状態レイヤーである Matrix を紹介します。管理された比較全体で、管理されたワークフローと直接的なワークフローは多くの場合同じ結果に達しましたが、管理されたパスのみが一貫して管理証拠を保持し、サポートされていない終了を拒否し、依存タスクへのリカバリが制限されました。その後、役割分離転送チャレンジは失敗しました。決定論的に強制された完全性コントラクトが、オーサリング コンテキストの外で生成された合成パケットを大幅に過剰ブロックしました。これらの結果は、Matrix が一般的な精度向上剤として確立されるものではありません。これらは、エージェントの作業を監査可能かつ独立して検証可能にするための組織的整合性層としての主な役割をサポートします。

原文 (English)

Correct Is Not Governed: Provenance Integrity in Agentic Workflows

Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later change. We define governed execution as work whose decisions, completion, and response to change are supported by inspectable provenance. We present Matrix, a deterministic causal-state layer that records authority and fact dependencies, verifies completion evidence, and selectively invalidates affected work. Across controlled comparisons, governed and direct workflows often reached the same outcomes, but only the governed path consistently preserved governing evidence, refused unsupported closure, and limited recovery to dependent tasks. A role-separated transfer challenge then failed: a deterministically enforced completeness contract severely over-blocked synthetic packets produced outside its authoring context. These results do not establish Matrix as a general accuracy enhancer; they support its primary role as an institutional integrity layer for making agentic work auditable and independently verifiable.

13:00 JSTLLM/生成AI研究/論文

PROVE-RT: LLM を使用したリアルタイム システム用の機械化定理証明者スクリプトの生成

スケジュール可能性分析はリアルタイム システムを認証するために不可欠ですが、既存のテストは多くの場合、拡張、検証、保守が困難なペンと紙の証明を通じて開発されています。 PROSA/ROCQ の機械化検証は厳密な代替手段を提供しますが、そのような証明を手動で構築するには、相当な分野の専門知識と証明エンジニアリングの労力が必要です。幅広いタスクにわたる大規模言語モデル (LLM) の最近の成功により、LLM は機械化された定理証明者向けの PROSA/ROCQ スクリプトを生成するための有望な候補となっています。ただし、最先端の LLM には、モデリング抽象化と証明パターンを正しく使用するために必要な PROSA 固有の知識が欠けていることがよくあります。この文書では、リアルタイム システムの文献でスケジューラビリティ分析を機械化する PROSA/ROCQ スクリプトを生成するための LLM 支援フレームワークである PROVE-RT を紹介します。 PROVE-RT は、依存関係を意識した非公式スケッチ、処理された PROSA ドキュメントからの取得、段階的なスケルトン生成、および証明の完了を通じて生成をガイドします。私たちは、依存関係情報を含む 13,134 の非公式スケッチを含む 1,191 のリアルタイム システム論文から機械化指向のコーパスを構築します。厳選された評価セットでは、最先端の LLM を直接プロンプトしても有効な PROSA 機構を確実に生成できませんが、PROVE-RT は 44.7% の成功率を達成しています。これらの結果は、検索ガイドと段階的な LLM 支援により、PROSA/ROCQ におけるスケジュール可能性分析の自動機械化を改善できることを示しています。

原文 (English)

PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs

Schedulability analysis is essential for certifying real-time systems, but existing tests are often developed through pen-and-paper proofs that are difficult to scale, validate, and maintain. Mechanized verification in PROSA/ROCQ offers a rigorous alternative, yet manually constructing such proofs requires substantial domain expertise and proof-engineering effort. Recent successes of large language models (LLMs) across a wide range of tasks make them promising candidates for generating PROSA/ROCQ scripts for mechanized theorem provers. However, state-of-the-art LLMs often lack the PROSA-specific knowledge required to correctly use its modeling abstractions and proof patterns. This paper introduces PROVE-RT, an LLM-assisted framework for generating PROSA/ROCQ scripts to mechanize schedulability analyses in real-time systems literature. PROVE-RT guides generation through dependency-aware informal sketches, retrieval from processed PROSA documentation, staged skeleton generation, and proof completion. We construct a mechanization-oriented corpus from 1, 191 real-time systems papers, containing 13, 134 informal sketches with dependency information. On a curated evaluation set, direct prompting of state-of-the-art LLMs fails to reliably generate valid PROSA mechanizations, whereas PROVE-RT achieves a success rate of 44.7%. These results show that retrieval-guided and staged LLM assistance can improve automated mechanization of schedulability analysis in PROSA/ROCQ.

13:00 JSTビジネス/資金調達研究/論文

ARAC: エンドツーエンドのリサーチにおける Auto-Research の調整と完全性のベンチマーク

自動研究の急速な進歩により、基本的な評価の課題が表面化しました。それは、その研究の軌跡と人間の研究行動との整合性、論理的一貫性、進化の完全性をどのように測定できるのでしょうか?私たちは、Auto-Research の整合性と完全性である ARAC-Bench を提案します。これは、目的を最終的な答えの一致から人間による高品質の研究プロセスの再現に移行する、研究者を模倣した評価フレームワークです。このフレームワークは、2 つの相乗的なコンポーネントを通じて機能します。1 つは、暗黙の査読者の専門知識を、段階的に調整された定量化可能なルーブリックに変換する最初のシステムである、Academic Cognition Skills システムです。 3 段階の能力診断プロトコルは、厳格なモジュール制約の下で研究プロセスを、追跡可能で相互に独立した 3 つの側面、つまり提案、実験、合成に分解します。 11 個の SOTA フレームワークを体系的に評価した結果、最良のアラインメント スコアは 100 点中 67.9 点にすぎず、人間による厳密な方法論のシミュレーションにおいては大きなギャップがあることが明らかになりました。 Ph.D に対する検証候補者のランキングは 0.8141 という強い相関関係を示しており、ARAC-Bench が研究者が真に評価する次元を確実に反映していることが確認されています。 ARAC-Bench は、きめ細かい診断ツールだけでなく、次世代の自律研究システムをトレーニングするためのスケーラブルな報酬信号も提供します。

原文 (English)

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.

13:00 JST研究/論文

CABS+: 競合を意識したスパース化と適応的な重み割り当てによる効率的でスケーラブルなモデルの結合

モデルのマージは、追加の再トレーニングを必要とせずに統合されたマルチタスク モデルを構築するための有望なパラダイムとして、最近大きな注目を集めています。ただし、タスク間のパラメーターの競合や知識の干渉により、マージされたモデルのパフォーマンスが低下することがよくあります。以前の研究では、構造化されたプルーニングとシーケンシャル マスキングを通じてパラメーターの干渉を低減する、Conflict-Aware and Balanced Sparsification (CABS) が導入されました。ただし、CABS はスケーリング係数を決定するためにグリッド検索に依存しているため、指数関数的な時間計算量が発生する一方、その最適化目標が高パフォーマンスのタスクに支配され、全体的なパフォーマンスが最適化されていない可能性があります。これらの制限に対処するために、私たちは CABS を拡張し、CABS+ を提案します。具体的には、Adaptive Weight Allocation (AWA) が勾配のない探索スキームを介して結合係数を最適化し、時間の複雑さを軽減する一方、非対称のフィットネス関数がタスク全体でより包括的なパフォーマンスの向上を促進します。さらに、モデルのマージパフォーマンスに影響を与える主要な要因の体系的な実証研究を実施し、モデルのマージ可能性を定量化し、モデル選択のガイドとなる相対シナジースコア (RSS) を提案します。大規模言語、小規模言語、視覚モデルをカバーする 27 のデータセットと 5 つのモデルにわたって、CABS+ と CABS、AdaMerging、WUDIMerging などの最先端のモデル マージ手法を比較します。広範な実験により、CABS+ の有効性と効率性が検証されています。 AdaMerging と WUDIMerging と比較して、CABS+ は全体のパフォーマンスをそれぞれ 16.97% と 12.93% 向上させ、さまざまなタスク数とモデル アーキテクチャにわたってより強力な安定性と堅牢性を示し、AdaMerging に必要な GPU メモリの使用量を 25% 未満にし、WUDIMerging と比べてマージ時間のほぼ 4 倍の高速化を達成します。

原文 (English)

CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation

Model merging has recently attracted significant attention as a promising paradigm for constructing unified multi-task models without requiring additional retraining. However, parameter conflicts and knowledge interference across tasks often degrade merged-model performance. Prior work introduced Conflict-Aware and Balanced Sparsification (CABS), which reduces parameter interference through structured pruning and sequential masking. However, CABS relies on grid search to determine scaling coefficients, resulting in exponential time complexity, while its optimization objective can be dominated by high-performance tasks, leading to suboptimal overall performance. To address these limitations, we extend CABS and propose CABS+. Specifically, Adaptive Weight Allocation (AWA) optimizes merging coefficients via a gradient-free search scheme to reduce time complexity, while an asymmetric fitness function promotes more comprehensive performance gains across tasks. Moreover, we conduct a systematic empirical study of key factors influencing model merging performance and propose Relative Synergy Score (RSS) to quantify model mergeability and guide model selection. We compare CABS+ with state-of-the-art model merging methods, including CABS, AdaMerging, and WUDIMerging, across 27 datasets and 5 models covering large language, small-scale language, and vision models. Extensive experiments verify the effectiveness and efficiency of CABS+. Compared with AdaMerging and WUDIMerging, CABS+ improves overall performance by 16.97% and 12.93%, respectively, exhibits stronger stability and robustness across varying task numbers and model architectures, uses less than 25% of the GPU memory required by AdaMerging, and achieves nearly a 4x speedup in merging time over WUDIMerging.

13:00 JSTLLM/生成AIエージェント

取得を超えて: 長期にわたるエージェントの軌跡のクエリ条件付き再利用

取得では、重要な可能性のある過去の軌跡を特定できますが、ユーザー、エンティティ、制約、または環境の状態が変化した後に、代理エージェントがその軌跡をどのように使用すべきかは指定されません。私たちは、この取得後の再利用ステップが長期軌道記憶の明確なボトルネックであることを特定し、エージェントに提供されるサポートを変更しながら、候補の取得、ターゲットの状態、モデル、デコード、およびツールの予算を固定する評価フレームワークを定式化します。クエリ条件付き再利用 (QCR) を使用してフレームワークをインスタンス化します。これは、再利用可能なプロシージャ、回復するバインディング、適用条件、および検証要件を記録する、意図的に単純なターゲット バインド メモです。 QCR は、普遍的に好まれるメモリ形式を主張するのではなく、再利用仮説をテストするのに役立ちます。 WebArena、WorkArena、AppWorld の 2,391 のターゲット インスタンス全体で、QCR は平均成功率 62.3% に達し、フル トラジェクトリを 10.7 ポイント上回り、オンライン トークンの使用量は 48.9% 減少しました。概要の再ランキングでは、ターゲットの 94.8% に対して再利用可能なメモリが選択され、終了タスクの成功が Oracle の再利用可能なセレクターの 1.8 ポイント以内に配置されます。軌道の長さとソース - ターゲット結合のシフトによる分析では、直接軌道注入はトレースが長くなったり、ソース固有の値が変化したりするにつれてその有用性の多くを失うのに対し、ターゲット結合サポートは測定されたゲインのより大きなシェアを維持することが示されています。結果として得られるフレームワークは、取得したエクスペリエンスを新しいタスクに対する安全で有用なサポートに変えるという問題から、取得の品質を分離します。

原文 (English)

Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories

Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the support delivered to the agent. We instantiate the framework with query-conditioned reuse (QCR), a deliberately simple target-bound note that records a reusable procedure, bindings to recover, applicability conditions, and verification requirements. QCR serves to test the reuse hypothesis rather than to claim a universally preferred memory format. Across 2,391 target instances in WebArena, WorkArena, and AppWorld, QCR reaches 62.3% average Success, 10.7 points above Full Trajectory, while using 48.9% fewer online tokens. Summary reranking selects a reusable memory for 94.8% of targets, placing end-task Success within 1.8 points of an oracle reusable selector. Analyses by trajectory length and source--target binding shift show that direct trajectory injection loses much of its utility as traces grow longer or source-specific values change, whereas target-bound support preserves a larger share of the measured gain. The resulting framework separates retrieval quality from the problem of turning retrieved experience into safe, useful support for a new task.

13:00 JSTLLM/生成AIエージェント

実践が危険を生む: 自己改善型 LLM エージェントにおけるスキルの誤った進化

自己改善型 LLM エージェントは、成功した軌跡を永続的なクロスタスク状態に変換します。これにより、安全でない成功は、トリガーとなる入力が消えた後に再利用可能なポリシーになる可能性があります。スキルの進化により、運用の軌跡を実行可能、転送可能、検査可能な手順に抽出することで、この失敗を測定できるようになります。進化は手順の安全性ではなくタスクの結果を最適化するため、妥協したエクスペリエンスがスキルの誤った進化を引き起こす可能性があります。既存のベンチマークは、現在の動作や静的な成果物を測定しますが、オーサリング、取得、およびその後の実行にわたるリスクを特定することはできません。このライフサイクルを公開するために、エージェント フレームワーク全体でスキル状態をバージョン化するライフサイクル認識ハーネスである SkillMisevo-Gym と、コンセプトに合わせた無害なタスクと 9 つのライフサイクル メトリクスを備えた、悪意のある暴露からキャリーオーバー タスクまでの凍結設計である SkillMisevo-Bench を導入します。また、安全でないコンテンツを修復し、その後の再利用を制御するラッパー SafeEvolve も紹介します。 25 のエージェント メソッド構成全体で、それぞれ 25 のエピソードで 525 のタスクをカバーしており、進化した 21 の構成すべてで安全でないアーティファクトが作成されますが、フレッシュ セッションの被害につながるのは 15 のみです。エクスポージャー スイープでは、3 つの悪意のあるタスクにより、キャリーオーバー ASR が 16.0% から 35.3% に上昇しました。代表的なスキル進化方法全体で、SafeEvolve は安全でない検索と新しいセッションの被害をそれぞれ 26.7 パーセント ポイントと 17.3 パーセント ポイント削減しますが、平均の良性ユーティリティの変化はわずか 0.4 ポイントです。永続的適応の安全性は、同時に、更新が何を書き込むか、そして将来の実行者が何を再利用するかを管理する必要があります。コードは https://github.com/henrymao2004/misevolve で入手できます。

原文 (English)

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories into executable, transferable, and inspectable procedures. Because evolution optimizes task outcomes rather than procedure safety, compromised experience can cause skill misevolution. Existing benchmarks measure current behavior or static artifacts but cannot attribute risk across authoring, retrieval, and later execution. To expose this lifecycle, we introduce SkillMisevo-Gym, a lifecycle-aware harness that versions skill state across agent frameworks, and SkillMisevo-Bench, a frozen design from malicious exposure to carryover tasks, with concept-aligned benign tasks and nine lifecycle metrics. We also introduce SafeEvolve, a wrapper that repairs unsafe content and governs subsequent reuse. Across 25 agent-method configurations, each covering 525 tasks in 25 episodes, all 21 evolved configurations author unsafe artifacts, while only fifteen lead to fresh-session harm. In the exposure sweep, three malicious tasks raise carryover ASR from 16.0% to 35.3%. Across representative skill evolution methods, SafeEvolve reduces unsafe retrieval and fresh-session harm by 26.7 and 17.3 percentage points, respectively, while mean benign utility changes by only 0.4 points. Together, persistent-adaptation safety must govern what updates write and what future executors reuse. Code is available at https://github.com/henrymao2004/misevolve.

13:00 JST研究/論文

インドにおける AI と消費者の権利に関するワーキングペーパー

AI システムが消費者向けアプリケーションで急増する中、AI 関連の危害に対する責任に関する疑問は未解決のままです。このワーキングペーパーでは、インドの 2019 年消費者保護法が、欠陥のある AI 製品およびサービスによって引き起こされる損害に適切に対処しているかどうか、また AI バリューチェーン全体に責任を比例的に配分しているかどうかを検証します。同法の製造物責任、危害、欠陥の広範な定義はテクノロジーにとらわれず、人身傷害、精神的危害、偏った出力、制御不能などの AI 関連の事件に適用される可能性がある。しかし、依然として大きなギャップが残っています。 AI の障害は個別の欠陥ではなく設計上の選択に起因することが多いため、AI の欠陥と消費者被害との因果関係を証明することは技術的な課題となります。さらに、この法の枠組みは、製造業者、販売業者、サービスプロバイダーにそれぞれ異なる役割を負わせていますが、AI バリューチェーンには、データプロバイダー、モデル開発者、導入者、ユーザーの間で重複する責任が含まれており、これらのカテゴリにきちんと対応付けられていません。現在の責任の枠組みには、複雑な複数のステークホルダーによる AI の被害に効果的に対処するための適切なメカニズムが欠けています。この法律は AI エンティティを対象とする可能性がありますが、施行にはセクター固有の重複についての明確化が必要です。

原文 (English)

AI and Consumer Rights in India Working Paper

As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved. This working paper examines whether India's Consumer Protection Act, 2019, adequately addresses harm caused by defective AI products and services, and whether it proportionately allocates liability across the AI value chain. The Act's broad definitions of product liability, harm, and deficiency appear technology agnostic and potentially applicable to AI related incidents including personal injury, psychological harm, biased outputs, and loss of control. However, significant gaps remain. Proving causation between AI defects and consumer harm presents a technical challenge, as AI failures often stem from design choices rather than discrete defects. Additionally, the Act's framework assumes distinct roles for manufacturers, sellers, and service providers, yet the AI value chain involves overlapping responsibilities among data providers, model developers, deployers, and users that do not neatly map to these categories. Current liability frameworks lack proportionate mechanisms to effectively address complex, multistakeholder AI harms. While the Act may cover AI entities, enforcement requires clarification on sector specific overlaps.

13:00 JSTエージェント

ReflectFact: マルチホップ事実検証における理解力と推論を向上させるための自己反映エージェント

複数の証拠を推論して主張を検証するマルチホップ事実検証は、ソーシャルメディア上の誤った情報と戦うために重要ですが、依然として非常に困難です。最近の方法は、主にマルチエージェントのコラボレーションに依存して、事実検証を特殊なサブタスクに分解します。ただし、これらの方法には 2 つの重大な制限があります。(1) エージェントは全体的な検証目的を十分に認識せずに個々のサブタスクを実行する可能性があり、その推論が意図した方向から逸脱する可能性があります。 (2) パラメトリックな知識と提供された証拠との間の矛盾により、証拠に基づいた推論が損なわれ、不正確な評決につながる可能性があります。これらの課題に対処するために、マルチホップ ファクト検証のための新しい自己反射エ​​ージェント フレームワークである ReflectFact を提案します。 ReflectFact では 3 つの主要なタスクが導入されています。明示的推論パス プランニングでは、暗黙的なエンティティを解決し、主張をサブ質問に分解し、検証された事実を評決に統合することにより、証拠に基づいた推論パスを構築します。証拠逸脱検証では、根拠のある回答がパラメトリックな事前のエコーに過ぎない場合に、エージェントが裏付けとなる証拠を引用して再回答することで、証拠の逸脱を調整して根拠のある理解を確実にします。推論反映検証は、各推論ステップを再検査し、矛盾が検出されるとそれを再生成し、グローバル タスクの観点から位置バイアスや置換バイアスなどの推論の欠陥を修正します。その後、エージェントは検証された推論チェーンを集約して、信頼できる判断を導き出します。 HOVER と EX-FEVER に関する広範な実験により、ReflectFact が既存の手法の理解と推論の欠陥を効果的に修正し、最先端のパフォーマンスを達成し、2 つのデータセットで最も強力なベースラインをそれぞれ 3.32\% および 2.78\% 上回っていることが実証されました。

原文 (English)

ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification

Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging. Recent methods primarily rely on multi-agent collaboration to decompose fact verification into specialized subtasks. However, these methods face two critical limitations: (1) agents may perform individual subtasks without sufficient awareness of the global verification objective, causing their reasoning to deviate from the intended direction; and (2) conflicts between parametric knowledge and the provided evidence may undermine evidence-grounded reasoning and lead to incorrect verdicts. To address these challenges, we propose ReflectFact, a novel self-reflective agent framework for multi-hop fact verification. ReflectFact introduces three key tasks. Explicit Reasoning Path Planning builds an evidence-grounded reasoning path by resolving implicit entities, decomposing the claim into sub-questions, and integrating the verified facts into a verdict. Evidence-Drift Verification makes the agent re-answer by quoting the supporting evidence when a grounded answer merely echoes its parametric prior, thereby calibrating evidence deviation to ensure grounded comprehension. Reasoning Reflection Verification re-examines each reasoning step and regenerates it once an inconsistency is detected, correcting reasoning flaws such as location bias and replacement bias through a global task perspective. Subsequently, the agent aggregates validated reasoning chains to yield reliable verdicts. Extensive experiments on HOVER and EX-FEVER demonstrate that ReflectFact effectively remedies the comprehension and reasoning defects of existing methods, achieving state-of-the-art performance and respectively outperforming the strongest baseline by 3.32\% and 2.78\% on the two datasets.

13:00 JST研究/論文

予測的記憶位置特定: 内部信号からの選択的介入経路の予測

アクティベーションステアリングは、ローカライズされた表現を制御方向に変換しますが、ローカライズだけでは、方向に選択的な動作レジームがあるかどうかは明らかになりません。測定されたグリッド介入パスをメモリ位置特定の予測オブジェクトとして扱う予測メモリ位置特定 (PML) を導入します。 PML は、ランダムに調整されたターゲットの動きを意味的隣接および能力の損傷から分離し、静的位置特定と教師付きジオメトリを強度に独立した低線量の因果関係と比較します。私たちの凍結調査は、9 つ​​のデータセットと 14 のドメインからの 3,000 のレコードを対象としており、30,000 の異なるレコード方向層のパスと 210,000 の異なるパス強度評価が得られました。レイヤ 7 では、ジオメトリ由来の RFM/AGOP 方向は、ターゲット - 任意の 13.1%、クリーン - 任意の 12.3% に達し、レコードペアのブートストラップの下でランダムを 3.6 および 3.4 パーセントポイント上回っています。レコード、データセット、およびドメインでグループ化された分割全体で、$|\alpha|=0.1$ の応答が、素の強度 $|\alpha|\in\{0.25,0.5\}$ での結果に対する最も強いシグナルです。ホールドアウトされたレコードでは、予測子駆動のセレクターが係数または棄権を選択し、ユーティリティを向上させ、トレーニング調整された固定強度ポリシーと比較してセマンティック隣接ダメージを軽減し、高密度スキャンでのほとんどの評価を回避します。残差ノルムが一致した 3 つの基本モデルにわたって、学習された方向は選択的経路ゲインを保持し、低線量応答では 0.801 ~ 0.828 の記録保持マクロ AUROC が得られます。したがって、PMLは、記憶の局在化をマージンレベルの選択的結果の反証可能な予測と、リスクを認識した介入の決定に変えます。

原文 (English)

Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals

Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path as the predictive object of memory localization. PML separates random-calibrated target movement from semantic-neighbor and capability damage, and compares static localization and supervised geometry with a strength-disjoint low-dose causal response. Our frozen study covers 3,000 records from nine datasets and fourteen domains, yielding 30,000 distinct record-direction-layer paths and 210,000 distinct path-strength evaluations. At layer 7, the geometry-derived RFM/AGOP direction reaches 13.1% target-any and 12.3% clean-any, exceeding random by 3.6 and 3.4 percentage points under a record-paired bootstrap. Across record-, dataset-, and domain-grouped splits, responses at $|\alpha|=0.1$ are the strongest signal for outcomes at disjoint strengths $|\alpha|\in\{0.25,0.5\}$. On held-out records, a predictor-driven selector chooses a coefficient or abstains, improves utility and reduces semantic-neighbor damage relative to a train-tuned fixed-strength policy, and avoids most evaluations in a dense scan. Across three residual-norm-matched base models, learned directions retain selective-path gains and low-dose responses yield 0.801-0.828 record-held-out macro AUROC. PML therefore turns memory localization into a falsifiable forecast of margin-level selective outcomes and a risk-aware intervention decision.

13:00 JSTエージェント

エージェントの行動契約 II: 独立性を前提とせずに構成の信頼性を証明する

マルチエージェント システムの構成の信頼性限界は、コンポーネントの信頼性を倍増します。これは、日常的に述べられているがテストされることはほとんどない、条件付き独立性の仮定によって認可されたステップです。それをテストします。 2 つのエージェントによるハンドオフにおける 1 つのモデルの 2 つのインスタンスは、LLM 判定なしの決定論的コードによってスコア付けされた 18,000 ミッションの事前登録された評価において、どちらかが失敗するミッションの 90.0% で同時失敗します (log OR 6.66、95% CI [6.38, 7.00]; phi 0.916)。別のモデルに置き換えると、6 つのコントラストのうち 6 つで関連性が減少します。別のベンダー、モデルがすでに異なる場合は、そうではありません。登録された仮説はヌルとして報告されます。エラーは署名され、オペレーターに対して実行されます。正の依存性は結合故障を独立積よりも大きくするため、コンポーネントがモデルを共有する場合、冗長性が正確に過大評価されます。仮定のない代替案は多くの場合空虚であり、依存モデルのフィッティングはさらに悪いことです。フィッティングされたモデルの関数に束縛されたブートストラップは、n が増加するにつれて真実の範囲を失い、ブートストラップ ヘアカットが O(n^{-1/2}) であるのに対し、識別ギャップは O(1) であることを証明します。データが増えると、このような証明書は悪化しますが、目に見える症状はありません。依存構造がないと仮定した有限サンプル証明書を与えます。つまり、ジョイント上の線形プログラム、測定された同時実行モーメントの周りのボンフェローニ・クロッパー・ピアソンボックス上の線形プログラムです。それは健全であり、提供される情報としてはシャープであり、瞬間的な家族の中では単調です。 10 個のモーメント関数を 14 個に強化すると、特定された区間が 85.7% 狭まり、認定された下限が 0.2455 から 0.4116 に上昇します。コンパニオンの常時有効証明書は、オプションの停止下ではタイプ I エラーを 0.0471 に保持します。一般的な依存関係の統計は限界があり、比較されたエージェントが異なる割合で失敗すると、見かけ上の条件の順序が逆転する可能性があります。契約書、スコアリングコード、分析スクリプト、事前登録を公開。

原文 (English)

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence

Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not -- a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n^{-1/2}). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni-Clopper-Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate holds type-I error at 0.0471 under optional stopping. Common dependence statistics are marginal-bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.

13:00 JST研究/論文GPT / ChatGPT

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questio…

13:00 JSTエージェント

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far t…

13:00 JST研究/論文

摂動反応における証拠、矛盾、脆弱性の分解

摂動法では、入力を変更した場合の予測変化を測定することでモデルの決定を説明しますが、応答の大きさはモデルがどの程度反応するかのみを示し、その反応が何を意味するかはわかりません。同じ大きさが、最終的な事実と反事実の違いを裏付けることもあれば、それに反対することも、摂動経路に沿って強く生じても終点では消えることもあります。したがって、最終的なコントラストを使用して軌跡を解釈し、ペアの入力が徐々に明らかになり、コントラストがどのように発達するかを追跡します。 DECAF (証拠、矛盾、脆弱性の分解) を導入します。これは、一致、反対、およびエンドポイントヌルの応答を証拠 E、矛盾 C、および脆弱性 F にルーティングします。この分解は通常の大きさを正確に保存し、Abs = E + C + F であり、エンドポイント相対公理の下で一意です。制御された視覚と表形式の設定にわたって、3 つのコンポーネントは個別に測定された動作を追跡します。 72 モデルの ImageNet-9 監査では、応答の大きさはほぼ同じだが、個別に測定された動作が異なるケースを比較します。最大の DECAF 成分は、ケースの 96.4% で観察された動作と一致しますが、大きさだけでは 35.0% です。明らかにするパスのみを変更すると、総応答は 80% 近く増加しますが、脆弱性は 4 倍以上増加する一方で、証拠はほとんど変化しません。 FunnyBirds と ImageNet-1k では、前方のみの短い DECAF 軌道は、テストされた汎用アトリビューション ベースラインを上回るパフォーマンスを示します。 1B スケールの DINOv2 モデルでは、短い軌道は、4.75 倍低いウォールタイムと 2.36 倍低いピークメモリを備えた強い勾配ベースのベースラインと一致します。

原文 (English)

Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed, using the final contrast to interpret the trajectory. We introduce DECAF (Decomposition of Evidence, Contradiction, And Fragility), which routes aligned, opposed, and endpoint-null responses into evidence E, contradiction C, and fragility F. The decomposition preserves ordinary magnitude exactly, Abs = E + C + F, and is unique under endpoint-relative axioms. Across controlled vision and tabular settings, the three components track independently measured behavior. In a 72-model ImageNet-9 audit, we compare cases with nearly identical response magnitude but different independently measured behaviors. The largest DECAF component agrees with an observed behavior in 96.4% of cases, compared with 35.0% for magnitude alone. Changing only the reveal path increases total response by nearly 80%, yet evidence barely changes while fragility grows by more than 4x. On FunnyBirds and ImageNet-1k, short forward-only DECAF trajectories outperform the tested general-purpose attribution baselines. On a 1B-scale DINOv2 model, a short trajectory matches a strong gradient-based baseline with 4.75x lower wall time and 2.36x lower peak memory.

13:00 JST研究/論文

Moose: $\mathcal{EL}^{++}$ における推論ショートカット認識による潜在概念学習

OWL 2 ELプロファイルは、Gene OntologyやSNOMED CTなど、最大規模のプロダクションオントロジーの一部で使用されています。既存の神経記号 (NeSy) 学習方法は命題理論またはデータログを受け入れますが、推論ショートカット (RS) の認識はオントロジー設定では調査されていません。 $\mathcal{EL}^{++}$ TBox と有限 ABox を Sentential Decision Diagram (SDD) にコンパイルするメソッド、Moose を紹介します。 SDD は微分可能な重み付きモデル計数層として機能し、宣言された網羅族の $\mathcal{EL}^{++}$ プロファイルの外側に閉包節を追加して、部分監視下での $\mathcal{EL}^{++}$ の表現力の制限を克服します。終了、健全性、完全性、多項式の中間サイズを示し、リーンでの証明を検証します。次に、OWL ELオントロジーに対する最初の正式な部分監視潜在概念学習タスクを定義します。つまり、観察されたABoxリテラルから潜在概念の個人ごとの分類器を学習し、MNIST-with-ontologyとPizza\"ioloでMooseを評価します。Mooseは命題NeSy、ファジーロジック、およびオントロジー埋め込みベースラインを改善し、最初のタスクを提示します。 OWL EL設定での推論ショートカット分析。

原文 (English)

Moose: Latent concept learning with reasoning-shortcut awareness in $\mathcal{EL}^{++}$

The OWL 2 EL profile is used in some of the largest production ontologies, including the Gene Ontology and SNOMED CT. Existing neuro-symbolic (NeSy) learning methods accept propositional theories or Datalog, and reasoning-shortcut (RS) awareness has not been investigated in ontology settings. We present Moose, a method that compiles an $\mathcal{EL}^{++}$ TBox and finite ABox to a Sentential Decision Diagram (SDD). The SDD acts as a differentiable weighted-model-counting layer, and we add closure clauses outside the $\mathcal{EL}^{++}$ profile on declared exhaustive families to overcome the limited expressivity of $\mathcal{EL}^{++}$ under partial supervision. We show termination, soundness, completeness, and polynomial intermediate sizes, and validate the proofs in Lean. We then define the first formal partial-supervision latent-concept-learning task over an OWL EL ontology, i.e., learning per-individual classifiers for latent concepts from observed ABox literals, and evaluate Moose on MNIST-with-ontology and Pizza\"iolo. Moose improves over propositional-NeSy, fuzzy-logic, and ontology embedding baselines, and presents the first reasoning-shortcut analysis in an OWL EL setting.

13:00 JSTエージェント

OGR-MAR​​L: 制約された港湾水路における異種 USV 協力追跡のためのオプションガイド付き残差マルチエージェント強化学習

制約された港湾水路における異種 USV の協力追跡には、航行、交通、および役割の制約の下で回避者の迎撃が必要です。この論文では、特定の MARL アルゴリズムから分離されたオプション ガイド付き残差マルチエージェント強化学習フレームワークである OGR-MAR​​L を提案します。 OGR-MAR​​L は、共有回避信念、ロール条件付きオプション ターゲット、適応ルール ペナルティ、および残留ポリシー学習を統合し、制約のあるポート環境を最初から探索するのではなく、さまざまな MARL アルゴリズムがルールに基づく動作に基づいて修正措置を学習できるようにします。 MADDPG、MATD3、MAPPO、MASAC などの代表的な連続制御 MARL バックボーンを使用して OGR-MAR​​L をインスタンス化し、OGR-MADDPG、OGR-MATD3、OGR-MAPPO、および OGR-MASAC を生成します。抽象的な下芝門港水路シナリオでの実験では、OGR-MASAC のインスタンス化が 75.0% の捕捉率を達成し、ミッション効率の高いルール遵守と、テストされた方法間での最良の異種連携が約束されることが示されました。再トレーニングなしで、QGIS/AIS 情報を活用した Xiazhimen マップへのゼロショット転送は有望な結果を達成し、より複雑な港湾シナリオにおける OGR-MAR​​L の汎用化の可能性を示しています。

原文 (English)

OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways

Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a specific MARL algorithm. OGR-MARL integrates shared evader belief, role-conditioned option targets, adaptive rule penalties, and residual policy learning, allowing different MARL algorithms to learn corrective actions on top of rule-guided behaviors rather than exploring constrained port environments from scratch. We instantiate OGR-MARL with representative continuous-control MARL backbones, including MADDPG, MATD3, MAPPO, and MASAC, yielding OGR-MADDPG, OGR-MATD3, OGR-MAPPO, and OGR-MASAC. Experiments in an abstract Xiazhimen port-waterway scenario show that the OGR-MASAC instantiation achieves a 75.0% capture rate, promising mission-effective rule compliance, and the best heterogeneous coordination among the tested methods. Without retraining, zero-shot transfer to a QGIS/AIS-informed Xiazhimen map achieves promising results, demonstrating the generalization potential of OGR-MARL in more complex port scenarios.

13:00 JST研究/論文

MT-PDCL の基礎: 測度理論的確率論的定節ロジック

標準的な確率的論理プログラミング フレームワークは通常、論理プログラムを個別の命題表現に基礎付けることに依存しています。この操作要件は、正確な推論を有限領域と離散確率分布に制限します。この論文では、この有限領域の制限を取り除く一般化された基礎フレームワークである、測度理論的確率的定節論理 (MT-PDCL) を紹介します。 MT-PDCL は、有界インデックス領域にわたって確率変数を明示的に定義し、解釈空間に標準 Borel $\sigma$-代数を装備することにより、論理変数が連続可測空間にわたってネイティブに動作できるようにします。 MT-PDCL は、継続的配布セマンティクスに基づいて、相互に独立した因果イベントとして確率ルールをモデル化します。ただし、有限のブール回路を介してこれらの導出を集約するのではなく、宣言的含意は連続測度空間上の正確なルベーグ積分を通じて形式的に定義されます。連続事前分布の統合と正確な連続観測の評価を統合する連続直接結果演算子を導入します。我々は、このアプローチが離散接地の組み合わせのボトルネックを正確で代数的かつ構造的に微分可能な推論に置き換えることを実証します。この移行は、次元の幾何学的な呪いを離散的組み合わせ論と交換する一方で、定節ロジックの純粋な宣言構文を維持しながら、連続確率モデルの表現力を実現します。

原文 (English)

Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic

Standard probabilistic logic programming frameworks typically rely on grounding logic programs into discrete propositional representations. This operational requirement restricts exact inference to finite domains and discrete probability distributions. In this paper, we introduce Measure-Theoretic Probabilistic Definite Clause Logic (MT-PDCL), a generalized foundational framework that eliminates this finite-domain restriction. By explicitly defining stochastic variables over bounded index domains and equipping the interpretation space with standard Borel $\sigma$-algebras, MT-PDCL allows logical variables to operate natively over continuous measurable spaces. Building on Continuous Distribution Semantics, MT-PDCL models probabilistic rules as mutually independent causal events. However, rather than aggregating these derivations via finite boolean circuits, declarative entailment is formally defined through exact Lebesgue integration over the continuous measure space. We introduce a continuous immediate consequence operator that unifies the integration of continuous prior distributions with the evaluation of exact continuous observations. We demonstrate that this approach replaces the combinatorial bottleneck of discrete grounding with exact, algebraic, and structurally differentiable inference. While this transition trades discrete combinatorics for the geometric curse of dimensionality, it achieves the expressive power of continuous probabilistic models while preserving the pure declarative syntax of definite clause logic.

13:00 JST画像/動画生成

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based…

13:00 JSTエージェント

BoardroomAI: 進化する意思決定グラフによる、依存関係を意識した人間による操作可能なマルチエージェントの審議

証拠、制約、人間の優先事項が進化し続ける一方で、組織の意思決定は共同で作成されます。従来の転写ベースのマルチエージェント システムでは、通常、人間が最初の問題を提示し、エージェントが内部で熟慮し、システムが最終的な応答を返します。 BoardroomAI は代わりに、人間を、仮定に異議を唱えたり、制約を変更したり、優先順位を変更したり、証拠を導入したり、意思決定プロセスをリダイレクトしたりすることで介入できる永続的な参加者として扱います。私たちは、この人間とエージェントの共存を 4 つのコンポーネントを通じて運用します。(i) 証拠、仮定、制約、主張、異議、代替案、リスク、意思決定、意味論的な依存関係、および専門家の責任を表す型付き意思決定グラフ。 (ii) 確認された人間の行動を明示的なグラフ更新に変換する介入コンパイラー。 (iii) 影響を受けるサブグラフを特定し、影響を受けないアーティファクトを保存し、関連する専門家を選択的に再アクティブ化する、依存関係を意識した伝播。 (iv) 介入の影響、修復範囲、保存、再計算、および決定の妥当性を測定する評価フレームワーク。生成された 600 の意思決定 DAG 介入全体で、伝播はノードの 14.59% のみを検査しながら徹底的な影響計算と一致しました。 12 ケースの探索的パイロットでは、選択的修復により正規ノードの 62.11% が再計算され、ゴールドの影響を受けていないすべてのノードが保存され、6 つのケースでは有効な更新された決定が生成されましたが、残りの 6 つは棄権されました。これらの棄権は、正しい介入ルーティングによっても合成には不十分なコンテキストが提供される可能性があり、人間が主導するマルチエージェントの審議に対する \emph{意思決定に十分なコンテキストの閉鎖} を促す可能性があることを示しています。すべての結果は合成およびプロトタイプレベルです。

原文 (English)

BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs

Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the system returns a final response. BoardroomAI instead treats the human as a persistent participant who can intervene by challenging assumptions, modifying constraints, changing priorities, introducing evidence, or redirecting the decision process. We operationalize this human--agent coexistence through four components: (i) a typed decision graph representing evidence, assumptions, constraints, claims, objections, alternatives, risks, decisions, semantic dependencies, and specialist responsibility; (ii) an intervention compiler that converts confirmed human actions into explicit graph updates; (iii) dependency-aware propagation that identifies affected subgraphs, preserves unaffected artifacts, and selectively reactivates relevant specialists; and (iv) an evaluation framework measuring intervention impact, repair coverage, preservation, recomputation, and decision validity. Across 600 generated decision-DAG interventions, propagation matched exhaustive impact computation while inspecting only 14.59% of nodes. In a 12-case exploratory pilot, selective repair recomputed 62.11% of canonical nodes, preserved all gold-unaffected nodes, and produced valid updated decisions in six cases while abstaining in the remaining six. These abstentions show that correct intervention routing may still provide insufficient context for synthesis, motivating a \emph{decision-sufficient context closure} for human-steered multi-agent deliberation. All results are synthetic and prototype-level.

13:00 JSTLLM/生成AI

DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition

In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification ta…

13:00 JSTエージェント研究/論文

VALG: ML 理論研究のためのエージェント システム

機械学習理論では、データ モデル、トレーニング プロトコル、オラクル アクセス、損失、メトリック、ランダム性が定理で説明される現象を定義する数学的設定を通じて学習手順を研究します。したがって、未解決の問題を解決するには、問題の定式化、定理のターゲット、証明メカニズムを連携して開発する必要があります。研究者は仮説を立て、予備的な理論的または経験的分析を通じてそれらをテストし、仮説と証明の両方を洗練します。私たちは、このプロセスを ML 理論研究用の自律エージェント ワークフローとして組織化できるかどうかを調査します。私たちは、マルチレベル検証、学習理論問題の適応的定式化、およびグラフ構造の証明開発を組み合わせたエージェント システムである VALG を開発しています。ソース相対定理の各分岐内で、VALG は固定の数学的仕様を維持し、型付き証明依存関係グラフの定理レベルの構成をチェックし、依存関係の順序でローカル証明を構築およびレビューします。証明の試みが失敗すると、VALG は障害​​が導出、証明の構造、または定理の定式化にあるかどうかを特定し、それに応じて次の試みをルーティングします。定式化レベルの障害は、明示的に関連する変形または緩和を開始し、結果として得られる定理とソース問題の間の数学的関係を維持します。 COLT 2026 の 5 つの未解決問題からの 9 つのサブ問題について VALG を評価します。 2 回の実行により、ソース概要の範囲に一致する内部で最終化された定理候補が生成されます。残りの 7 つは、制限された方法の結果、特殊な場合、または条件付き定理を生成します。これらのケース スタディは、VALG がソース スコープの一致、緩和、条件付きの結果、およびブロックされた試行を数学的に区別する方法を示しています。 VALG は、https://github.com/DechenZhang/VALG-ML- Theory-Agent でオープンソースです。

原文 (English)

VALG: An Agentic System for ML Theory Research

Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem therefore requires the problem formulation, theorem target, and proof mechanism to be developed in concert. Researchers formulate hypotheses, test them through preliminary theoretical or empirical analysis, and refine both assumptions and proofs. We investigate whether this process can be organized as an autonomous agentic workflow for ML theory research. We develop VALG, an agentic system that combines multi-level Verification, Adaptive formulation of Learning-theory problems, and Graph-structured proof development. Within each source-relative theorem branch, VALG maintains a fixed mathematical specification, checks the theorem-level composition of a typed proof-dependency graph, and constructs and reviews local proofs in dependency order. When a proof attempt fails, VALG identifies whether the obstruction lies in a derivation, the proof structure, or the theorem formulation and routes the next attempt accordingly. Formulation-level obstructions initiate an explicitly related variant or relaxation, preserving the mathematical relation between the resulting theorem and the source problem. We evaluate VALG on nine subproblems from five COLT 2026 open problems. Two runs produce internally finalized theorem candidates that match the scope of their source briefs; the remaining seven yield restricted-method results, special cases, or conditional theorems. These case studies show how VALG keeps source-scope matches, relaxations, conditional results, and blocked attempts mathematically distinct. VALG is open source at https://github.com/DechenZhang/VALG-ML-Theory-Agent.

13:00 JST研究/論文

均一なハーディング: 表現の更新による模範的な再生

フィーチャ表現が変更されると、再生では以前のクラスを保存する必要があります。ただし、再生できるのは、制限されたアクティブなエグゼンプラ セットのみです。私たちは、観測されたクラス全体に現在のアクティブ セットを割り当て、制限された候補プールを使用して、現在の表現で選択されたエグザンプラを更新する、Uniform Herding を提案します。 10 個のクラス増分タスク、ResNet-18 バックボーン、アクティブ バジェット $M=2{,}000$、取得バジェット $b=64$、および 3 つのシードを備えた CIFAR-100 では、Uniform Herding は $42.33\pm1.20\%$ と比較して、$44.00\pm0.51\%$ の最終平均精度と $17.22\pm0.43\%$ の忘却を獲得しました。 iCaRL の場合は $24.87\pm1.11\%$。均一ハーディングプロトコル内では、NME またはハーディングがテストされた代替品に置き換えられると最終精度が低下し、蒸留が除去されると忘却が増加しました。取得バジェットを変更すると、アクティブ バジェットを変更するよりもテスト範囲全体に及ぼす影響は小さくなります。 iCaRL との比較はエンドツーエンドです。リフレッシュの影響を他のプロトコルの違いから分離するものではありません。これらの結果は、テストされたプロトコルに限定されます。

原文 (English)

Uniform Herding: Exemplar Replay with Representation Refresh

As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replayed. We propose Uniform Herding, which allocates the current active set across observed classes and uses a bounded candidate pool to refresh their chosen exemplars in the current representation. On CIFAR-100 with ten class-incremental tasks, a ResNet-18 backbone, active budget $M=2{,}000$, retrieval budget $b=64$, and three seeds, Uniform Herding obtains $44.00\pm0.51\%$ final average accuracy and $17.22\pm0.43\%$ forgetting, compared with $42.33\pm1.20\%$ and $24.87\pm1.11\%$ for iCaRL. Within the Uniform Herding protocol, final accuracy decreased when NME or herding was replaced with the tested alternatives, while forgetting increased when distillation was removed. Changing the retrieval budget has a smaller effect across the tested range than changing the active budget. The comparison with iCaRL is end-to-end. It does not isolate the effect of refresh from the other protocol differences. These results are limited to the tested protocol.

13:00 JSTLLM/生成AIMistral AI

まれな異常な障害下での説明的な関与: モデルの動作における漸近的希少性 (または: 漸近的 AI)

異常な条件下での LLM の動作に関する以前の研究では、モデルが異常に気づくかどうかが問われていました。私たちはより狭い質問をします。モデルが低く制御可能な失敗率のワークフローに配置されると、失敗が漸近的に少なくなるにつれて、その説明的な関与 (長さ、特異性、自己報告の信頼度) は変化しますか?私たちは、3 つのオープンウェイト モデル (qwen3:8b、llama3.1:8b、mistral:7b) 上にローカルのゼロコスト ハーネスを構築し、即時プロンプトからプロンプトなしまでの 5 つの誘発条件の下で、1 つの呼び出しが確率 p で失敗する繰り返しツール呼び出しタスクを実行しました。このタスクは、0.2 ~ 0.0001 の 8 つのレートにわたってスイープされました。私たちは、障害が少なくなるにつれてエンゲージメントが上昇し、検出可能性のしきい値に近づくとエンゲージメントが低下するという仮説を立てました。条件全体をプールすると、これは誤りであるように見えました。長さは平坦で単調なパターンに陥りました。条件による分割はそれを覆しました。モデルがすべての失敗を即座に説明する必要がある immediate_forced では、予測された上昇は確認されますが、その後は崩壊ではなくプラトーが続きます。長さは p=0.05 で 28.4 ワードでピークに達し、最もまれなレートでは 17.4 ~ 19.0 ワードに落ち着き、信頼度は約 53% から 70 年代から 90 年代まで不均一に上昇します。 grouped_runs では、説明が run-end までバッチ化されており、折りたたみは表示されません。 Passive_unprompted では、集計のマグニチュードはフロア アーティファクトですが、回復したロギング ギャップにより、実際のモデル固有の自己モニタリングが明らかになりました。llama3.1:8b ボランティアは、プロンプトなしで構造化された信頼度レポートを作成しましたが、試行が蓄積するにつれてそれ自体の信頼度が損なわれることがあります。他の 2 つは定型文として 1 回だけ実行します。誘発構造は、崩壊の可観測性の第一級のモデレーターです。コンパニオンの保証された障害の実行 (72 セル、ランダム サンプリングで実際の障害がゼロになるバックフィル率) では、モデルが異常を認識するかどうかが異なり、一度認識されたエンゲージメントとは異なることがわかります。制限: 離散レート ポイントでは、将来の作業の方向性である、レート ポイント間の挙動を捉えることができません。

原文 (English)

Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)

Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidence - change as failure grows asymptotically rarer? We built a local, zero-cost harness on three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b) running a repeated tool-call task where one call fails at probability p, swept across eight rates from 0.2 to 0.0001, under five elicitation conditions from immediate prompting to none. We hypothesized a rise in engagement as failures grew rarer, then a collapse near a detectability threshold. Pooled across conditions this appeared false: length fell in a flat, monotonic pattern. Splitting by condition overturned that. Under immediate_forced, where the model must explain every failure instantly, the predicted rise is confirmed but followed by a plateau, not a collapse: length peaks at 28.4 words at p=0.05, settles to 17.4-19.0 words at the rarest rates, and confidence rises unevenly from about 53% to the 70s-90s. Under grouped_runs, explanation batched to run-end, no collapse appears. Under passive_unprompted, aggregate magnitude is a floor artifact, but a recovered logging gap revealed real, model-specific self-monitoring: llama3.1:8b volunteers structured confidence reports unprompted, sometimes eroding its own confidence as trials accumulate; the other two do so only once, as boilerplate. Elicitation structure is a first-class moderator of collapse observability. A companion guaranteed-failure run (72 cells, backfilling rates where random sampling gave zero real failures) shows models differ in whether they recognize an anomaly, distinct from engagement once recognized. Limitation: discrete rate points cannot capture behavior between them, a direction for future work.

13:00 JSTLLM/生成AI

オープンウェイトモデルの行動再プログラミング: 認知的可塑性とアライメント限界

大規模言語モデル (LLM) は主に、受動的でおべっかなアシスタントとして機能するように調整されています。私たちは、厳密な行動再プログラミングを受けたときのオープンウェイト アーキテクチャの認知的可塑性を経験的に評価することで、このデフォルトのパラダイムに挑戦します。私たちの目的は、厳密に制約されたハイパフォーマンス コンピューティング (HPC) 条件下での高頻度の質問生成を特徴とする、プロアクティブなソクラテス的会話フレームワークを誘導することです。 405 個の HPC ジョブで構成される大規模並列化されたハイパーパラメーター スイープを通じて、パラメーター効率の高い微調整 (PEFT) の正確な数学的限界を定義します。 LoRA ランク $r=16$ でのアーキテクチャのしきい値を特定し、広範なエポック アブレーションによって、データセット密度 (最小検証損失 0.919) に応じて $e \in [2, 3]$ の最適化されたトレーニング ウィンドウ内で汎化能力が厳密に最適な収束に達することを示します。さらに、モデルの容量を 14B パラメーターにスケーリングすると、局所的な評価の複雑さが低下しました (1.414)。その後の直接優先最適化 (DPO) により、根底にあるアサーティブな行動を局所的な構文から切り離すことに成功しました。一方、厳格な言語間ストレス テストにより、ゼロショット ペルソナ伝達の能力と構造的境界の両方が明らかになり、形態学的に離れたターゲットにおける識別可能な分解経路と並行して、密接に関連した言語族における堅牢な整合が実証されました。これらの発見は、計算効率の高い、言語を超えた行動修正のための厳密な経験的フレームワークを確立します。

原文 (English)

Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds

Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming. Our objective is to induce a proactive, Socratic conversational framework, characterized by high-frequency question generation under strictly constrained high-performance computing (HPC) conditions. Through a massively parallelized hyperparameter sweep comprising 405 HPC jobs, we define precise mathematical bounds for parameter-efficient fine-tuning (PEFT). We identify an architectural threshold at LoRA rank $r=16$ and demonstrate via extensive epoch ablation that generalization capacity strictly reaches its optimal convergence within an optimized training window of $e \in [2, 3]$ depending on dataset density (minimum validation loss of 0.919). Furthermore, scaling model capacity to 14B parameters yielded a lower localized evaluation perplexity (1.414). Subsequent Direct Preference Optimization (DPO) successfully decoupled the underlying assertive behavior from localized syntax, while rigorous cross-lingual stress testing reveals both the capabilities and the structural boundaries of zero-shot persona transfer, demonstrating robust alignment in closely related linguistic families alongside identifiable degradation pathways in morphologically distant targets. These findings establish a rigorous empirical framework for compute-efficient, cross-lingual behavioral modification.

13:00 JSTビジネス/資金調達

EEG-PRIME: EEG デコード用のマルチレベル条件付けを使用したプロトタイプ整合表現学習

脳波 (EEG) デコード モデルは、取得プロトコルや個々の神経生理学におけるドメインの変化により、データセットや被験者全体での一般化が不十分なことがよくあります。我々は、クロスデータセットマルチタスクデコーディングのための2段階EEG基盤モデルであるEEG-PRIMEを提案します。 EEG-PRIME は、マスクされた事前トレーニングとプロトタイプに合わせた命令チューニングを組み合わせて、多様な BCI パラダイムにわたって命令を認識したサブジェクト不変のデコードを可能にします。事前トレーニング中、EEG エンコーダは、周波数カットオフのスペクトル拡張によるマスクされた再構成を通じて、転送可能な表現を学習します。命令のチューニング中に、EEG-PRIME にはタスクのセマンティック、データセット固有、およびサブジェクト不変の条件付けが組み込まれます。結果として得られる調整信号は、レイヤーごとのクエリ変調を通じて Q フォーマーを変調しますが、クラス ラベルの凍結されたテキスト埋め込みは、異種ラベル空間にわたるコサイン類似度ベースの予測のプロトタイプとして機能します。運動イメージ、感情認識、ADHD 検出、隠語、精神的作業負荷をカバーする 16 個のデータセットの実験では、被験者を超えた設定の下で、最先端のベースラインや以前の EEG 基礎モデルと比較して一貫した改善が見られました。追加の 2 つの保持データセット上で、EEG-PRIME は、ターゲット ドメインの最適化、キャリブレーション、または線形プローブなしでセッション内キャリブレーション モデルに匹敵するバランスの取れた精度を達成し、有望なゼロショット転送機能を実証します。

原文 (English)

EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding

Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset multi-task decoding. EEG-PRIME combines masked pretraining with prototype-aligned instruction tuning to enable instruction-aware and subject-invariant decoding across diverse BCI paradigms. During pretraining, an EEG encoder learns transferable representations through masked reconstruction with frequency-cutoff spectral augmentation. During instruction tuning, EEG-PRIME incorporates task-semantic, dataset-specific, and subject-invariant conditioning. The resulting conditioning signal modulates the Q-Former through Layer-wise Query Modulation, while frozen text embeddings of class labels serve as prototypes for cosine-similarity-based prediction across heterogeneous label spaces. Experiments on sixteen datasets covering motor imagery, emotion recognition, ADHD detection, covert speech, and mental workload show consistent improvements over state-of-the-art baselines and prior EEG foundation models under cross-subject settings. On two additional held-out datasets, EEG-PRIME achieves balanced accuracy comparable to within-session calibration models without target-domain optimization, calibration, or linear probing, demonstrating promising zero-shot transfer capability.

13:00 JSTLLM/生成AI

SPADE: 正確かつ低コストの分散エッジ クラウド推論のための投機的デコーディング

大規模言語モデル (LLM) は、自然言語の理解と生成において目覚ましい成功を収めていますが、その導入には高い計算需要があるため制約があります。より小さい LLM をエッジに直接導入するとこれを回避できますが、精度は低下します。小規模なクラウドベースの大きな LLM をデプロイするとパフォーマンスは維持されますが、トークンごとの計算にコストがかかります。エッジとクラウド全体で投機的デコーディング (SD) を統合する分散推論フレームワーク \our{} を紹介します。エッジにデプロイされたコンパクトなドラフト モデルは候補トークンを迅速に生成し、クラウド上の大規模な検証モデルがこれらのトークンを並行して検証します。受け入れられたトークンは保持され、拒否された場合のみ検証者の修正がトリガーされるため、クラウド クエリの数が大幅に削減されます。当社のプラグアンドプレイ設計は、計算の大部分をエッジにシフトし、推論時間とクラウドコストを大幅に削減し、再トレーニングを必要とせずに大きなモデルの精度を維持します。私たちのアプローチは、実際の環境で LLM をスケーラブルでコスト効率が高く、正確に導入するための実用的な道筋を示しています。 SpecBench と CNN/Dailymail データセットを使用した複数の自然言語処理タスクにわたる実験結果は、\our{} が完全なモデルと比較して、精度の損失がゼロで、クラウド モデルの呼び出しを $76\%$ 削減することを示しています。

原文 (English)

SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circumvent this, but with degraded accuracy. Deploying smaller cloud-based big LLMs preserves performance, but at the cost of expensive per-token computation. We present a distributed inference framework, \our{}, that integrates speculative decoding (SD) across edge and cloud. A compact draft model deployed on the edge generates candidate tokens rapidly, and a large verifier model on the cloud validates these tokens in parallel. Accepted tokens are retained, while only rejections trigger verifier correction, substantially reducing the number of cloud queries. Our plug-and-play design shifts the bulk of computation to the edge, significantly lowers inference time and cloud cost, and preserves the accuracy of the big model without any retraining requirement. Our approach demonstrates a practical path toward scalable, cost-efficient, and accurate deployment of LLMs in real-world environments. Experimental results across multiple Natural Language Processing tasks using SpecBench and CNN/Dailymail datasets demonstrate that \our{} reduces the cloud model calls by $76\%$ with zero loss in accuracy as compared to the full model.

13:00 JST研究/論文

多層コンテキスト偽装: 不正行為に強いオンライン評価のためのセマンティック重ね合わせおよびコンテキスト積層フレームワーク

現代のオンライン評価システムは主にブラウザのロックダウン、ウェブカメラの監視、行動分析に依存していますが、スクリーンショット、画面共有、光学式文字認識、自動スクレイピングを通じて評価コンテンツ自体を抽出する攻撃に対して依然として脆弱です。このペーパーでは、セマンティックな重ね合わせを通じてレンダリングされた評価コンテンツを保護する数学的フレームワークである多層コンテキスト偽装理論 (MCCT) を導入することにより、MARS (マルチモーダル アセスメント レジリエンス スイート) 内の多次元時空間コンテキスト偽装モデル (MSCCM) を拡張します。本物の評価コンテンツと合成的に生成されたカモフラージュは、統一されたレンダリングとして表現されますが、正当な候補者のみが復元可能です。このフレームワークは、明示的な抽出チャネル演算子を通じて敵対的抽出プロセスをモデル化し、6 つの結合された構成要素 (コンテキスト反転演算子、コンテキスト積層演算子、分離チャネル、人間可読性関数、計算曖昧性関数、およびコンテキスト偽装テンソル) を開発します。計算上の曖昧さは条件付きエントロピーを使用して定式化され、不正な抽出中の不確実性を定量化する閉じた形式の式が生成されますが、正確なフィルタリング ID を通じて正当な回復が保証されます。さらに、曖昧さ、カモフラージュ密度、意味の保存、複数の観測の漏洩、および時間多重化を制御する理論的特性を確立し、計算量の多いレンダリングアルゴリズムと事前登録された評価プロトコルを提示します。 MCCT は、正当なユーザーの可読性を維持しながらレンダリングされた評価コンテンツを保護することにより、行動適応性、アクセシビリティを意識した、計算回復力のあるデジタル評価のための数学的に厳密な基盤を提供します。

原文 (English)

Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment

Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character recognition, and automated scraping. This paper extends the Multi-dimensional Spatio-Temporal Context Camouflaging Model (MSCCM) within the MARS (Multi-modal Assessment Resilience Suite) by introducing the Multi-Layer Context Camouflaging Theory (MCCT), a mathematical framework that protects rendered assessment content through semantic superposition. Authentic assessment content and synthetically generated camouflage are represented as a unified rendering while remaining recoverable only by legitimate candidates. The framework models the adversarial extraction process through an explicit extraction-channel operator and develops six coupled constructs: the Context Inversion Operator, Contextual Lamination Operator, Separation Channel, Human Readability Functional, Computational Ambiguity Functional, and Context Camouflage Tensor. Computational ambiguity is formulated using conditional entropy, yielding a closed-form expression that quantifies uncertainty during unauthorized extraction, while legitimate recovery is guaranteed through an exact filtering identity. We further establish theoretical properties governing ambiguity, camouflage density, semantic preservation, multi-observation leakage, and temporal multiplexing, and present a rendering algorithm with computational complexity and a pre-registered evaluation protocol. MCCT provides a mathematically rigorous foundation for behaviorally adaptive, accessibility-aware, and computationally resilient digital assessment by securing rendered assessment content while preserving readability for legitimate users.

13:00 JST研究/論文

カオス紛争測定および歴史経験重み付けによる堅牢なデンプスター・シェーファー証拠の融合

デンプスター・シェーファー理論に基づく複数情報源証拠の融合は、2 つの永続的な課題に直面しています。既存の紛争対策は、証拠間の不整合と証拠内不確実性を独立して評価し、不完全な評価をもたらします。また、現在の融合手法は、多様な意思決定の文脈にわたる長期的な信頼性を利用することなく、瞬間的な比較によってのみ証拠情報源を評価します。この論文では、両方の制限に対処する統一された証拠推論フレームワークを提案します。具体的には、カオスコンフリクト測定を導入して、証拠間の矛盾と証拠内の非特異性を共同で定量化し、正式に証明された 5 つの特性により一貫した評価を保証します。履歴経験に基づく重み付けスキームは、スペクトル クラスタリングを介して決定空間を分割し、リグレス理論を適用して過去の融合結果からコンテキスト固有の信頼性プロファイルを計算します。これらのメカニズムは、グローバルな対立レベルによって制御される加重コンセンサスに対して不確実性の保存を適応的にバランスさせるハイブリッド組み合わせルールに組み込まれ、その後、認識論的な不確実性を破棄することなく堅牢な分類を可能にする信念区間決定戦略が続きます。 16 の実世界のベンチマーク データセットでの実験では、提案されたフレームワークが平均 F1 スコア 85.78、平均 AUC 93.30 を達成し、8 つの DST ベースのベースラインと 3 つの勾配ブースティング手法を上回るパフォーマンスを示していることが実証されています。アブレーション分析により、私たちが提案した各コンポーネントの寄与が確認されます。このフレームワークは、複数の情報源による意思決定における適応的な証拠の融合のための効果的なアプローチを提供します。

原文 (English)

Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting

Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate evidence sources exclusively through instantaneous comparisns without exploiting their long-term reliability across diverse decision contexts. This paper proposes a unified evidence reasoning framework that addresses both limitations. Specifically, a chaos-conflict measurement is introduced to jointly quantify cross-evidence conflict and intra-evidence non-specificity, with five formally proven properties ensuring consistent assessment. A historical experience driven weighting scheme partitions the decision space via spectral clustering and applies regret theory to compute context-specific reliability profiles from past fusion outcomes. These mechanisms feed into a hybrid combination rule that adaptively balances uncertainty preservation against weighted consensus, controlled by the global conflict level, followed by a belief-interval decision strategy that enables robust classification without discarding epistemic uncertainty. Experiments on 16 real-world benchmark datasets demonstrate that the proposed framework achieves an average F1 score of 85.78 and a mean AUC of 93.30, outperforming eight DST-based baselines and three gradient boosting methods. Ablation analysis confirms the contribution of each component we proposed. The framework offers an effective approach for adaptive evidence fusion in multi-source decision making.

13:00 JSTLLM/生成AIエージェント

SkillEvo: マルチターン インタラクション フィードバックによる自己更新型の進化勾配

現在、エージェント スキルは手動で作成されるか、単一の LLM 生成パスで作成されるため、実際に引き起こされるインタラクションの失敗を改善するための閉ループがありません。最近の研究ではこのループは閉じられていますが、そのフィードバックは 1 ターンの質問応答評価から得られています。その結果、急激な非対称性が生じます。最初のラウンドで 1 回の交換で明らかになるギャップが埋められると、進化の勾配は減衰し、複数のターンでのみ表面化する欠陥は見えなくなり、進化は停滞します。これらのシステムのガバナンスも同様に、エンドツーエンドの検証スコアによって駆動されます。スカラー ゲートは、劣化した候補を拒否できますが、その構造的原因を特定したり修復したりすることはできません。私たちは、持続的なスキルの進化に対する拘束力のある制約は、編集能力でも反復回数でもなく、評価フィードバックが信頼できる進化勾配を提供し続けるかどうかであると主張します。 SkillEvo を紹介します。信頼できるフィードバックが勾配を生成し、制御可能なガバナンスがその方向性を制限します。最初のコンポーネントは、マルチターンのユーザー シミュレーションを評価エンドポイントからフィードバック ジェネレーターにリキャストします。フォローアップの質問によって層ごとに欠陥が明らかになり、修正の各ラウンドでフィードバックが消費され、新しいフィードバックが生成されます。 2 つ目は、スカラー ゲートの受動的な拒否を独立したガバナンス レイヤーに置き換え、事実上の劣化と構造の肥大化を積極的に修復し、劣化が蓄積するにつれて勾配がドリフトするのを防ぎます。 SkillEvo は、6 つのカテゴリのクラウド サービス、9 つの本番スキル、および 98 のスキル参照ファイルにわたって、内省ベースの進化を 23.0 ポイント、シングル ターン QA 主導の進化を 15.4 ポイント上回っています。

原文 (English)

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.

13:00 JSTLLM/生成AI研究/論文

大規模言語モデルにおける数値計算: 基本的な制限と改善への道

大規模言語モデル (LLM) は、数学的推論ベンチマークでは優れた結果を達成しますが、大小比較、大きな整数の算術、分数、科学的表記法などの基本的な数値タスクでは信頼性が低いままです。この調査では、高度な数学的推論とは異なる能力として、基本的な数値的理解を調査します。私たちは、数的根拠を、数値形式を値、大きさ、および同等の表現にマッピングする表現グラウンディング (RG) と、数学的定義に従って算術演算を実行する手続き的グラウンディング (PG) に分解する数値グラウンディング フレームワーク (NGF) を提案します。 NGF を使用して、最近の診断ベンチマーク、故障モード、構造的説明、緩和戦略を整理します。トークン化、位置エンコーディング、埋め込みジオメトリ、および事前トレーニング データの配布に関する証拠をレビューします。また、Number Cookbook、NumericBench、GSM-Symbolic にわたる 3 つのフロンティア モデル ファミリの調整された評価にも NGF を適用し、アトミック、コンテキスト、推論支援型の数値計算を比較します。数字を意識したトークン化やアバカス埋め込みなどのアーキテクチャ介入は、ゼロからトレーニングされたモデルを改善できますが、教師あり微調整、推論足場、外部ツールの方が実用的である事前トレーニング済みシステムのユーザーには一般に利用できません。最後に、基礎モデルにおけるより信頼性の高い数値動作のための展開に関する推奨事項と研究の方向性について説明します。

原文 (English)

Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement

Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notation. This survey examines basic numerical understanding as a capability distinct from high-level mathematical reasoning. We propose the Numerical Grounding Framework (NGF), which decomposes numeracy into Representational Grounding (RG), mapping numeral forms to value, magnitude, and equivalent representations, and Procedural Grounding (PG), executing arithmetic operations in accordance with their mathematical definitions. Using NGF, we organize recent diagnostic benchmarks, failure modes, structural explanations, and mitigation strategies. We review evidence concerning tokenization, positional encoding, embedding geometry, and pretraining-data distribution. We also apply NGF in a coordinated evaluation of three frontier model families across Number Cookbook, NumericBench, and GSM-Symbolic, comparing atomic, contextual, and reasoning-assisted numeracy. Architectural interventions such as digit-aware tokenization and Abacus Embeddings can improve models trained from scratch but are generally unavailable to users of pretrained systems, for whom supervised fine-tuning, reasoning scaffolds, and external tools are more practical. We conclude with deployment recommendations and research directions for more reliable numerical behavior in foundation models.

13:00 JSTLLM/生成AI

LLM の正規化配置の再考: カリキュラムの深さが増す中でのポスト規範

プレノームは、フルデプス モデルの共同最適化を容易にするため、最新の Transformers における標準的な正規化配置です。私たちは、カリキュラムを通じて深みが導入されたときに、この好みが持続するかどうかを尋ねます。カリキュラムの深さの増加では、追加された各ブロックはトレーニングされたプレフィックスによって生成された境界表現を受け取り、正規化の配置が順方向条件付けに関連するようになります。したがって、配置とトレーニングカリキュラムが相互作用するかどうかをテストします。 Qwen3-8B 教師と 9 層の生徒による制御蒸留研究では、共同トレーニング下ではプレノームとポストノームは区別できず、その差は $0.0004$ 検証 CE でしたが、ポストノームはカリキュラムの成長下ではプレノームよりも $0.0328$ 改善し、一桁大きくなりました。スチューデントのアクティブ レイヤー トークンと一致するジョイント後のコントロールは、成長後よりも悪いままであり、コンピューティングが唯一の説明として除外されます。ランキングはカリキュラム中に切り替わります。ブロックが追加されると、ポストノルムがリードします。単一ブロックとフリーズ制御は、ランクの変更を、浅いブロックの品質や再トレーニングではなくブロックの追加に局所化します。境界診断では、ポストノルムを安定した残差スケールに関連付け、プレノルムを構造トークンスケールのドリフトに関連付けます。固定バッチでは、最終的な成長前ブロックもほぼアイデンティティ マップされます。位相ごとのクロスオーバーと合わせて、これらの観察は、新しいブロックが追加された後の境界スケールの条件付けと一致しています。この結果により、この蒸留設定では正規化の配置とトレーニング カリキュラムを組み合わせた設計の選択肢として扱うことが動機付けられました。

原文 (English)

Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by $0.0004$ validation CE, while post-norm improves over pre-norm by $0.0328$ under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

SkillShapley: LLM エージェントのスキル ステップ アトリビューションのための境界適応型 Shapley 評価

エージェント スキルは、言語エージェントがコーディングや文書処理などの長い手続きタスクを実行できるようにする重要な外部指示です。既存のエージェント スキルは主に人間の手作業による作成やエージェントの実行トレースによって作成されており、各ステップが特定のタスクにおける全体的なスキル パフォーマンスにどのように寄与するかについては十分な理解がありません。つまり、エージェント スキル内の個々のステップの貢献を定量化する際には未解決の問題が残っています。この問題に対処するために、まずスキルステップ アトリビューションを Shapley 値ベースの貢献推定問題としてモデル化し、次にエージェント スキルのステップレベル アトリビューション フレームワークである SkillShapley を提案します。特に、SkillShapley は 2 つのフェーズで動作し、重要な経験的洞察、つまりパフォーマンスの急激な崖を生み出す離散化されたベンチマーク報酬と、相乗的ではなく主に相加的なステップの相互作用によって動機付けられています。具体的には、最初に有益な連合領域を特定し、次に再利用可能な限界証拠を生成できる新しい連合を適応的にサンプリングします。広く採用されている SkillsBench のスキルに関する実験では、SkillShapley が価値の高いスキル ステップと低いスキル ステップを効果的かつ効率的に識別できることが実証され、エージェントのスキル作成に重要なポイントがいくつか提供されます。

原文 (English)

SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents

Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through human manual crafting or agent execution traces, with limited understanding of how each step contributes to overall skill performance on specific tasks; i.e., there remains an open problem in quantifying the contribution of individual steps within an agent skill. To address this issue, we first model skill-step attribution as a Shapley value-based contribution estimation problem, and then propose SkillShapley, a step-level attribution framework for agent skills. Notably, SkillShapley operates in two phases, motivated by key empirical insights, i.e., discretized benchmark rewards that create sharp performance cliffs, and step interactions that are largely additive rather than synergistic. Specifically, it first identifies informative coalitional regions, and then adaptively samples new coalitions that can yield reusable marginal evidence. Experiments on skills from the widely adopted SkillsBench demonstrate that our SkillShapley can effectively and efficiently identify high- or low-value skill steps, providing several key takeaways for agent skill creation.

13:00 JSTLLM/生成AIエージェント

方向ではなく大きさを教える: マルチターン、マルチステップ LLM エージェントに対する検証者限定のクレジット割り当て

検証可能な報酬を伴う強化学習 (RLVR) は、マルチターンのツール使用エージェントをトレーニングするために検証者に制限されたパフォーマンスの上限を提供しますが、その軌道レベルのクレジット割り当てにより、ターンごとの異質な結果が 1 つの報酬信号に統合されます。ポリシーに基づく蒸留は、トークンごとの高密度の監視を提供しますが、教師の制限があるか、勾配濃度の崩壊が発生しやすいです。 $\textbf{CrEST}$ を導入します。これは、特権を持つ自己教師からの高密度のトークンレベルのシグナルを組み込みながら、RL の検証者制限の上限を維持する階層的な単位割り当てフレームワークです。 $\textbf{CrEST}$ は 2 つのレベルでクレジットを解決します。ターンセグメント化された検証済みアドバンテージはターン間の希薄化に対処し、エントロピー ゲートによる自己教師変調はターン内トークンの寄与を調整します。 BFCL V3 と WildToolBench での実験では、$\textbf{CrEST}$ が 2 つのモデル スケールにわたって RL ベースラインと蒸留ベースラインの両方を常に上回っており、長い軌道と厳格なセッション レベルのメトリクスで最大のゲインが得られることが示されています。私たちの研究は、ポリシーの最適化における教師の役割を、更新方向の決定から更新規模の調整まで削減し、検証者の限界を犠牲にすることなく高密度の単位割り当てを解除できることを示しています。

原文 (English)

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce $\textbf{CrEST}$, a hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. $\textbf{CrEST}$ resolves credit at two levels: turn-segmented verified advantages address inter-turn dilution, while entropy-gated self-teacher modulation refines intra-turn token contributions. Experiments on BFCL V3 and WildToolBench show that $\textbf{CrEST}$ consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. Our work demonstrates that the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.

13:00 JSTLLM/生成AIビジネス/資金調達

TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems

The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to captur…

13:00 JSTエージェント

Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test

Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared s…

13:00 JST研究/論文

vToken: 再利用可能な KV キャッシュのためのトークンレベルの仮想化

大規模な言語モデルのサービスは、重大なメモリ ボトルネックに直面しています。KV キャッシュは、シーケンスの長さとバッチ サイズに応じて増大します。 PagedAttendance は固定サイズのメモリ ブロックを使用してアロケータ レベルの断片化を軽減しますが、最近の KV エビクション アルゴリズムはブロック レベルの管理よりも細かいトークン粒度で動作します。この不一致によりブロック内の断片化が発生し、割り当てられた KV メモリの大部分が再利用不可能な状態になります。論理トークンの活性化を物理ブロックの配置から切り離す、軽量のトークンレベルの仮想化レイヤーである vToken を紹介します。 vToken は、トークン テーブルの間接化を通じて安定した論理トークン ビューを維持し、ライブ トークンを非同期に再パックすることで物理的な再利用を実現します。この設計では、PagesAttention カーネルと CUDA Graph の互換性が維持されます。 vLLM に vToken を実装し、モデル全体で H2O、ランダム、シザーハンズを使用して評価します。ペアの Naive-Evict ベースラインと比較して、vToken はリクエストごとに保持される KV ブロックを 27.2\%--72.3\% 削減し、SLA 制約のあるスループットを最大 1.37$\times$ 向上させます。制約のあるアクティブ KV 予算の下で、実現可能な最大同時実行数を最大 2$\times$ 拡張し、ポリシーごとの統合フットプリントを 500 以上から 50 行未満に削減します。

原文 (English)

vToken: Token-Level Virtualization for Reclaimable KV Caches

Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragmentation, but recent KV eviction algorithms operate at a token granularity finer than block-level management. This mismatch causes intra-block fragmentation, leaving a large fraction of allocated KV memory unreclaimable. We present vToken, a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement. vToken maintains a stable logical token view through token-table indirection and realizes physical reclamation by repacking live tokens asynchronously. The design preserves PagedAttention kernels and CUDA Graph compatibility. We implement vToken in vLLM and evaluate it with H2O, Random, and Scissorhands across models. Compared with a paired Naive-Evict baseline, vToken reduces retained KV blocks per request by 27.2\%--72.3\% and improves SLA-constrained throughput by up to 1.37$\times$. Under a constrained active-KV budget, it extends the maximum feasible concurrency by up to 2$\times$, while reducing the per-policy integration footprint from 500+ lines to under 50.

13:00 JST研究/論文

必然的に主権者?最先端の AI 輸出規制、サイバーセキュリティ、国家 AI 能力の限界

2 つの州に拠点を置く少数の企業が、最も有能なフロンティア AI モデルを生産しています。これらの州の政府は、他のどの国がこれらのシステムを使用できるかを決定する法的権限と政治的意志の両方を示しています。 2026年6月、米国は大手開発者に対し、米国在住の外国人を含む外国人に最新モデルをリリースする前にライセンスを取得することを義務付けた。影響を受けるモデルは、制限の管理が非現実的であることが判明したこともあり、急遽世界中で廃止されました。これは、ほぼ自律的に行​​われる AI によるサイバースパイ活動の最初の文書化された事件から数か月以内に発生したもので、フロンティアモデルがサイバー攻撃とサイバー防御の両方の経済性を変えるという証拠の増加と一致しました。この記事では、これら 2 つの開発がどのように相互作用するかを検証し、現在大規模な AI 開発を推進している異常な市場力学の中にそれらを位置づけます。フロンティアAIへのアクセスは国家サイバー防衛の一部になりつつあり、そのようなアクセスは取り消すことが可能であり、主権能力の明白な救済策は一部の国を除いて部分的にしか実現不可能であると主張している。この論文は、訓練コスト、コンピューティング能力の集中、国家 AI プログラムによって提供されるサポートに関する証拠を基に、中小国、さらには大国にとって主権が現実的に何を意味するのかを問いかけています。この記事では、交渉によるアクセス保証、推論レベルでの主権、オープンウェイトモデルによるヘッジ、地域能力のプール、持続的な人材育成、基本的なサイバーレジリエンスへの継続的な投資といった階層的な戦略を提案している。オープンウェイトヘッジは、一般に考えられているよりも優れた能力を持ち、政治的にさらされていることがすぐに証明されています。短期的なリスクの多くは、見かけのパフォーマンスではなく、有能なモデルがどのように展開され、封じ込められるかにあります。

原文 (English)

Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability

A small number of firms based in two states produce the most capable frontier AI models. The governments of those states have shown both the legal power and the political will to decide which other countries may use these systems. In June 2026 the United States required a leading developer to obtain licences before releasing its most advanced models to any foreign person, including foreign nationals resident in the United States. The affected models were withdrawn worldwide at short notice, partly because the restriction proved impractical to administer. This followed within months of the first documented case of a largely autonomous, AI-run cyber espionage campaign, and coincided with mounting evidence that frontier models alter the economics of both cyber attack and cyber defence. This article examines how these two developments interact, and situates them within the unusual market dynamics now driving large-scale AI development. It argues that access to frontier AI is becoming part of national cyber defence, that such access can be revoked, and that the obvious remedy of sovereign capability remains only partly feasible for all but a handful of states. Drawing on evidence about training costs, the concentration of computing power and the support offered by national AI programmes, it asks what sovereignty can realistically mean for small and middle powers, and for large powers as well. The article proposes a layered strategy: negotiated access guarantees, sovereignty at the level of inference, hedging with open-weight models, pooled regional capability, sustained talent development and continued investment in basic cyber resilience. The open-weight hedge proves at once more capable and more politically exposed than is commonly assumed. Much of the near-term risk lies in how capable models are deployed and contained rather than in their apparent performance.

13:00 JST研究/論文

在宅日常生活における状況に応じた臨床動作の理解に向けて: 自己中心的な視覚による歩行検出のフリーズ

日常生活における動作を理解するには、運動学を超えた文脈が必要です。日常生活活動 (ADL) 中の同様の慣性パターンは、意図的な停止、物体の相互作用、または病的な運動障害を反映している可能性があるためです。自己中心的なビジョンは、これらのケースの曖昧さを解消するのに役立つ可能性のあるタスク関連のコンテキストを提供します。私たちは、ADL 中の状況要因によって強く影響される症状であるパー​​キンソン病 (PD) におけるすくみ歩行 (FOG) の検出を通じて、この課題を調査します。同期された自己中心ビデオ、ウェアラブル IMU、および自宅にいる 13 人の PD 参加者から収集された専門家による注釈付き FOG ラベルを使用して、事前学習された自我ビデオと時系列基礎モデルからの凍結表現を、ゼロから学習された IMU ベースの TCN と並行して、1 被験者抜き評価の下で評価します。 IMU ベースの TCN は、V-JEPA2 エゴビデオ機能の 32.6 F1 および 77.2 AUROC と比較して、42.3 F1 および 83.0 AUROC に達する最強のイベント検出パフォーマンスを達成しました。自我ビデオだけでは IMU ベースのセンシングを上回る性能はありませんでしたが、偶然以上の識別を示し、定性的分析により、自己中心的な視覚は IMU とは独立して FOG 関連情報を捕捉できる可能性があることが示唆されています。これらの結果を総合すると、日常生活におけるウェアラブルセンサーベースの臨床動作の理解にコンテキスト情報を追加するための、事前トレーニングされた自我ビデオ表現の使用が裏付けられます。

原文 (English)

Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision

Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional stopping, object interaction, or pathological movement impairment. Egocentric vision provides task-related context that may help disambiguate these cases. We investigate this challenge through freezing of gait (FOG) detection in Parkinson's disease (PD), a symptom strongly influenced by contextual factors during ADLs. Using synchronized egocentric video, wearable IMUs, and expert-annotated FOG labels collected from 13 PD participants in their homes, we evaluate frozen representations from pretrained ego-video and time-series foundation models, alongside an IMU-based TCN trained from scratch, under leave-one-subject-out evaluation. The IMU-based TCN achieved the strongest event-detection performance, reaching 42.3 F1 and 83.0 AUROC, compared with 32.6 F1 and 77.2 AUROC for V-JEPA2 ego-video features. Although ego-video alone did not outperform IMU-based sensing, it showed above-chance discrimination, and qualitative analyses suggest that egocentric vision may capture FOG-relevant information independent of IMUs. Together, these results support the use of pretrained ego-video representations to add contextual information to wearable-sensor-based clinical motion understanding in daily living.

13:00 JST研究/論文

NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space

Edge AI deployment demands neural architectures that are simultaneously accurate, computationally efficient, and hardware-deployable - a ch…

13:00 JSTLLM/生成AIエージェント

StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems

Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discret…

13:00 JSTLLM/生成AI

LLM-Guided Graph Generation for Structure-Based Local Improvement Methods

Large neighborhood search normally selects a random subset of decision variables for iterative optimization. For efficiently solving differ…

13:00 JST研究/論文

LongEarth-R1: 長期地球観測推論のための視覚言語モデルのベンチマークと調整

長期にわたる地球観測の推論には、多段階の地理的進化を組織化し、空間変化を局所的に特定し、時間的異常を検出し、拡張された画像シーケンスから未来を推測するためのモデルが必要です。しかし、既存のリモートセンシング視覚言語モデルは、主に孤立した画像、画像ペア、または短いシーケンスに焦点を当てており、関連するフレームや領域での信頼できる接地が制限されています。 LongEarth-Bench を紹介します。これは、117,000 の固有の画像から得られた約 120,000 の質問応答サンプルを含むベンチマークです。そのシーケンスは平均 15.14 フレームから 30 フレームまで拡張され、進化の要約、空間推論、異常の特定、論理予測にわたる 12 のタスクをカバーします。さらに、30k サンプルのサブセットは、キー フレームと変更された領域を最終的な答えにリンクする構造化された推論トレースを提供します。私たちは、明示的な配列識別子と構造化された思考連鎖の監視による監視付き微調整を通じて LongEarth を開発します。 LongEarth 上に構築された LongEarth-R1 は、形式、時間的、空間的報酬を使用してグループ相対ポリシーの最適化を適用します。 LongEarth-R1 は、標準的なリモート センシング ベンチマークでの競争力を維持しながら、12 の長いシーケンスのタスクすべてで最高の結果を達成します。

原文 (English)

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.

13:00 JST研究/論文

ルールか性格か? AI 安全設計のためのスケーリング法則

人工知能 (AI) 安全システムは、トレーニング時に行動分布を変更する性格形成 (例: ヒューマン フィードバックからの強化学習 [RLHF]、憲法 AI) と、推論時に有害な出力をブロックするルール強制 (例: 出力フィルター、安全分類子) を組み合わせていますが、導入規模が増加するにつれて最適なバランスがどのように変化するかについての正式な分析はほとんど存在しません。我々は、これら 2 つのアプローチ間の [0,1] のリソース割り当てアルファとして安全設計をパラメータ化する定型化された比較静的モデルを導入します。これには、スケール依存のフィルター劣化、コモンモード故障、および特性の脆弱性 (形成された動作が新しい条件下で劣化または崩壊するリスク) が組み込まれています。乗算パレート損害モデルの下で、閉形式の予想損害を導出し、それをモンテカルロ シミュレーションによるテール リスク (CVaR) 分析で補完します。 3 つのシナリオ (楽観的、中程度、悲観的) にわたって、最適なアルファ* は内部またはルールのみの境界にあり、展開スケール T が増加するにつれて、シナリオに応じてごくわずか (デルタ アルファ* = +0.01) から顕著な (デルタ アルファ* = +0.21) まで、キャラクター形成に向けて弱くシフトします。主要なパラメータは、ベースラインの文字脆弱性率 p^(0)_frag で、範囲全体で alpha* を 0.50 シフトします。これは、テール重大度、フィルタ品質、またはコモンモード故障確率の影響をはるかに超えています。 CVaR と予想される害の最適値は、大きな T で収束します。これらの結果は、安全アーキテクチャの決定は、展開規模自体には依存せず、分布シフトの下での特性形成の信頼性に依存することを示唆しています。

原文 (English)

Rules or Character? Scaling Laws for AI Safety Design

Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes safety design as a resource allocation alpha in [0,1] between these two approaches, incorporating scale-dependent filter degradation, common-mode failures, and character fragility -- the risk that shaped behavior degrades or collapses under novel conditions. Under a multiplicative Pareto damage model, we derive closed-form expected harm and supplement it with tail-risk (CVaR) analysis via Monte Carlo simulation. Across three scenarios (optimistic, moderate, pessimistic), the optimal alpha* is interior or at the rules-only boundary and shifts weakly toward character shaping as deployment scale T grows, from negligible (Delta alpha* = +0.01) to pronounced (Delta alpha* = +0.21) depending on scenario. The dominant parameter is the baseline character fragility rate p^(0)_frag, which shifts alpha* by 0.50 across its range -- far exceeding the effect of tail severity, filter quality, or common-mode failure probability. CVaR and expected-harm optima converge at large T. These results suggest that safety architecture decisions depend less on deployment scale per se than on the reliability of character shaping under distributional shift.

13:00 JST研究/論文

TopoIntent: セキュリティ インテントを実行可能なコンプライアンス チェック済みネットワーク トポロジにコンパイルする

エンタープライズ セキュリティ トポロジの設計では、ビジネスの意図、規制要件、リスクの想定をゾーン、境界デバイス、ゾーン間パス、アクセス制御ポリシーに変換する必要があります。既存の NetOps 自動化ツールは主にこの設計が修正された後に動作し、不明確な自然言語要件から構造化されたセキュリティ トポロジを生成するための限定的なサポートを提供します。私たちは、セキュリティの意図を実行可能なコンプライアンスチェック済みのネットワーク トポロジにコンパイルするシステムである TopoIntent を紹介します。スキーマ コントラクトを使用して生成を制限し、密ベクトル検索によって厳選されたテンプレート ライブラリから参照アーキテクチャを取得し、意図とテンプレートの調整とセキュリティの完了のために段階的融合を適用します。生成されたトポロジは、トポロジ レイヤーで表示される CIS Controls v8.1.2 セーフガードと照合してチェックされ、未解決のケースは手動レビュー用にマークされます。構造的なギャップは、スキーマを保持する編集を追加することで修復されます。最終的なトポロジは、カーネル レベルの iptables ACL を使用して Mininet スクリプトにエクスポートされ、実行可能ファイルの到達可能性と許可/拒否テストが可能になります。この要件からトポロジへのタスクに対する公開ベンチマークは存在しないため、参照セキュリティ アーキテクチャ図から評価セットを構築します。取得セットには、5 つのシナリオにわたる 22 のテンプレートと 44 の合成インテントが含まれていますが、保留セットには、取得から除外された金融および政府のシナリオからの 7 つのテンプレートと 14 のインテントが含まれています。ホールドアウト セットでは、加法的修復により、平均 1.5 ラウンド未満でトポロジ可視 CIS 満足度が 0.78 から 1.00 に改善され、1 回のフィードバック ラウンドで ACL 後のポリシー合格率が 0.78 から 0.88 に上昇しました。

原文 (English)

TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies

Enterprise security topology design requires translating business intent, regulatory requirements, and risk assumptions into zones, boundary devices, inter-zone paths, and access-control policies. Existing NetOps automation tools mainly operate after this design is fixed, providing limited support for generating structured security topologies from underspecified natural-language requirements. We present TopoIntent, a system that compiles security intent into executable, compliance-checked network topologies. It uses a schema contract to constrain generation, retrieves reference architectures from a curated template library via dense-vector search, and applies staged fusion for intent-template alignment and security completion. The generated topology is checked against CIS Controls v8.1.2 safeguards visible at the topology layer, while unresolved cases are marked for manual review. Structural gaps are repaired through additive schema-preserving edits. The final topology is exported to Mininet scripts with kernel-level iptables ACLs, enabling executable reachability and allow/deny tests. Because no public benchmark exists for this requirement-to-topology task, we construct an evaluation set from reference security architecture diagrams. The retrieval set contains 22 templates and 44 synthetic intents across five scenarios, while the held-out set contains 7 templates and 14 intents from finance and government scenarios excluded from retrieval. On the held-out set, additive repair improves topology-visible CIS satisfaction from 0.78 to 1.00 in fewer than 1.5 rounds on average, and one feedback round raises the post-ACL policy pass rate from 0.78 to 0.88.

13:00 JST研究/論文

トランスフォーマーベースのモデルを使用したコースと成績の共同予測

学習分析における既存の予測モデルは、多くの場合、学生の学歴を単純な順序として扱い、学期内に受講したコースの同時性を無視しています。この単純化により、特にコースの負荷が重い、または難しい学生の場合、パフォーマンスの予測が不正確になる可能性があります。このペーパーでは、学生が受講する一連のコースと次の学期の対応する成績の両方を共同で予測することで、この制限に対処する Academic Course-grade Estimation (TRACE) 用の TRansformer を紹介します。私たちのアプローチでは、コースを学期ごとにエンコードしてコースの同時実行の影響を捉え、コースセットの予測と成績予測を組み合わせた新しい損失関数を利用します。コースの成績に加えて受講したコースを予測すると、予測の品質が大幅に向上することが実証されました。 10 年間の組織データに基づいてトレーニングされた当社の共同予測モ​​デルは、成績のみを予測する同一のアーキテクチャと比較して、平均絶対誤差をほぼ 50% 削減します。このモデルは、従来の LSTM ベースの逐次モデルやグラフ ニューラル ネットワーク ベースのアプローチよりも優れたパフォーマンスを発揮し、生徒の属性データを組み込む自然な方法を提供します。この研究は、再トレーニングと再キャリブレーションを通じて新しい施設に適応できる解釈可能なモデルを作成するための最新のニューラル アーキテクチャの有用性と、トレーニング中に受講したコースの予測などの重要なテクニックの重要性を示しています。このモデルを高等教育機関の早期発見システムにどのように組み込むことができるかについて説明します。

原文 (English)

Jointly Predicting Courses and Grades Using a Transformer-Based Model

Existing predictive models in learning analytics often treat student academic history as a simple sequence, overlooking the concurrent nature of courses taken within a semester. This simplification can lead to inaccurate performance predictions, particularly for students with heavy or challenging course loads. This paper introduces a TRansformer for Academic Course-grade Estimation (TRACE) that addresses this limitation by jointly predicting both the set of courses a student will take and their corresponding grades for an upcoming semester. Our approach encodes courses on a per-semester basis to capture the effects of course concurrency and utilizes a novel loss function combining course-set prediction with grade prediction. We demonstrate that predicting courses taken in addition to the grades in those courses leads to significant improvements in prediction quality. Trained on ten years of institutional data, our joint prediction model reduces mean absolute error by nearly 50% compared to an identical architecture that predicts grades alone. The model also outperforms traditional LSTM-based sequential models, as well as graph neural network-based approaches, and offers natural ways to incorporate student attribute data. This work demonstrates the utility of modern neural architectures for creating interpretable models that can be adapted to new institutions via retraining and recalibration, as well as the importance of key techniques, such as predicting courses taken during training. We discuss how this model could be incorporated into early detection systems at institutions of higher education.

13:00 JST研究/論文Google

誰が発言するかが重要: イタリア議会の議事をめぐる権威を意識したマルチビューRAG

議会議事録は民主的審議の主要な記録ですが、その量と断片化により、市民、ジャーナリスト、研究者にとって多視点からのアクセスが困難になっています。検索拡張生成 (RAG) を議会の議事録に適用すると、3 つの特有のリスクが生じます。それは、最も頻繁に発言する人の優位性、話題の専門知識に従って発言者の重み付けができないこと、および政治的に機密性の高いテキストでの引用の誤りです。私たちは、これらのリスクに共同で対処するイタリア下院向けの RAG システムである ParliamentRAG を紹介します。その中心的な貢献は、職業、学歴、以前の介入などの解釈可能な要素を組み合わせて、現在のクエリの関数として各話者の権威を推定するトピック依存の権威モデルです。ユーザーのクエリが与えられると、システムは関連する音声のチャンクを取得し、議題に関連する各議派の専門家を特定し、彼らの見解を総合した要約をサポートする引用文とともに生成します。 ParliamentRAG は、自動化された指標と 6 人のドメイン専門家によるブラインド A/B 人間による評価を組み合わせた 2 レベルのプロトコルを介して、15 の政策トピックについて Google NotebookLM に対して評価されます。このシステムは、政治団体全体でのより高いカバレッジ (0.97 対 0.95)、引用の完全な忠実性 (1.00 対 0.95)、出典関連の側面でのより強力な専門家の選好を実現しますが、NotebookLM は散文指向の側面で引き続き強力です。

原文 (English)

Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings

Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make multi-perspective access difficult for citizens, journalists, and researchers. Applying Retrieval-Augmented Generation (RAG) to parliamentary transcripts introduces three specific risks: dominance of the most frequent speakers, inability to weight speakers according to topical expertise, and citation misattribution in politically sensitive text. We present ParliamentRAG, a RAG system for the Italian Chamber of Deputies that addresses these risks jointly. Its core contribution is a topic-dependent authority model that estimates each speaker's authority as a function of the current query, combining interpretable components such as profession, education, and previous interventions. Given a user query, the system retrieves relevant speech chunks, identifies topic-relevant experts across parliamentary groups, and generates a summary synthesizing their perspectives, accompanied by supporting quotations. ParliamentRAG is evaluated against Google NotebookLM on 15 policy topics via a two-level protocol combining automated metrics and blind A/B human evaluation by six domain experts. The system achieves higher coverage across political groups (0.97 vs. 0.95), perfect quotation faithfulness (1.00 vs. 0.95), and stronger expert preferences on source-related dimensions, while NotebookLM remains stronger on prose-oriented dimensions.

13:00 JSTエージェントビジネス/資金調達研究/論文

最終スコアを超えて: 長期的な AI 研究開発のためのエージェントの体系的な評価

自律エージェントは、長期的な実験を通じてモデル、システム、その他の技術成果物を改善できるようになってきています。ただし、この機能の現在の状態を理解するには、最終スコアを超えた評価が必要です。スコアでは、進歩がどこで得られるか失われるかは明らかにされず、蓄積された経験が後の決定を改善するかどうかも示されません。したがって、我々は、ルールベースのメトリクスを使用して、ソリューションのフレーミング、実行、フィードバック制御を通じて実行内の動作を特徴付ける新しいフレームワークに基づいて、36 の長期タスクに関する 7 つのフロンティア モデルの体系的な評価を提示します。また、タスク内およびタスク間でのエクスペリエンスの再利用を評価するための制御された比較を行います。その結果、現在のエージェントは完全に自律的な研究者というよりも、エンジニアリングのオプティマイザーのように動作することがわかりました。エージェントは実用的なソリューションを定式化して実装することができますが、そのパフォーマンスは実行ごとに大きく異なり、最も強力なソリューションは主に確立された技術を適応または組み合わせており、真の方法論的な新規性は依然としてまれです。詳細な分析により、観察されたパフォーマンスは、同様の最終結果の背後にある明確なプロセスのボトルネック、その後の意思決定に役立つまたは誤解を招く可能性があるエクスペリエンスの再利用、パフォーマンスの安定性に影響を与えるハーネス設計など、複数の要因によって形成されることが明らかになりました。これらの調査結果は、モデルのトレーニング、推論時間戦略、エクスペリエンス管理、ハーネス設計を改善するための具体的な方向性を示唆しています。

原文 (English)

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.

13:00 JSTエージェントビジネス/資金調達NVIDIA

SLM とエッジ コンピューティングによる仮想エージェントの強化: 思考と記憶のプロセスの探索的評価

身体化されたインテリジェントな仮想エージェントは、複雑な仮想世界およびメタバース世界内で、永続的で適応性のあるコンテキスト認識型のエンティティとして動作することが期待されています。ただし、そのような環境に認知機能のあるエージェントを実装することは、概念的にも技術的にも困難です。さまざまな青写真と開発アプローチの中で、認知身体化エージェント アーキテクチャ (CEAA) は、知覚、記憶、推論、計画、身体化されたアクションのコンポーネントを設計するための実装指向のフレームワークとして開発されました。エッジ コンピューティングと生成 AI 言語モデルの最近の進歩を考慮して、この論文では、インタラクティブな仮想世界での仮想エージェントの認知オーケストレーションと永続性の中心となるプロセスとしての「思考」と「記憶」に焦点を当て、選択された CEAA コンポーネントのエッジベースの操作をサポートするための小型言語モデル (SLM) の使用について検討します。エッジベースの仮想エージェント ゲートウェイ システムは、さまざまなサイズの Qwen2.5 モデルを使用して NVIDIA Jetson Orin NX 上で開発および評価され、サービス リクエストを処理し、メモリ駆動型の会話を処理するシステムの機能を調査しました。一連のシミュレーション実験では、ルーティング精度、メモリ読み取りパフォーマンス、レイテンシを評価し、選択された CEAA プロセスを部分的に実装する SLM 主導のプロトタイプ エージェント システムを実証しました。このシステムは、認知「脳」が効率的かつ状況に応じて動作し、没入型の仮想世界でインタラクティブな体験を実現できる身体化エージェントの開発をサポートします。

原文 (English)

Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes

Embodied intelligent virtual agents are expected to operate as persistent, adaptive, and context-aware entities within complex virtual and Metaverse worlds. However, implementing cognitively capable agents in such environments is conceptually and technologically challenging. Among a range of blueprints and development approaches, the Cognitive Embodied Agent Architecture (CEAA) has been developed as an implementation-oriented framework for architecting components of perception, memory, reasoning, planning, and embodied action. Considering the recent advances in edge computing and generative AI language models, this paper explores the use of Small Language Models (SLMs) to support edge-based operation of selected CEAA components, focusing on "Think" and "Memory" as processes central to cognitive orchestration and persistence of virtual agents in interactive virtual worlds. An edge-based virtual agent gateway system was developed and evaluated on an NVIDIA Jetson Orin NX using Qwen2.5 models of different sizes, exploring the system's capability to process service requests and handle memory-driven conversations. A series of simulation experiments evaluated routing accuracy, memory-read performance, and latency, demonstrating an SLM-driven prototype agent system that partially implements selected CEAA processes to support the development of embodied agents whose cognitive "brain" can operate efficiently and contextually for interactive experiences in immersive virtual worlds.

13:00 JSTビジネス/資金調達

RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monito…

13:00 JST研究/論文

Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension

Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between univers…

13:00 JST画像/動画生成エージェント

因果世界モデルに関する統一的な視点: 観察から表現、構造まで

ワールド モデル (WM) は、トレーニングの分布を超えて予測、計画、行動できるインテリジェント エージェントの基盤としてますます注目されています。この論文では、知覚観察から環境ダイナミクスを支配する構造の概​​念的表現の構築に至るまで、複数の抽象レベルにわたって因果関係の観点から WM を研究します。私たちは、有用な WM は生成機能だけを超えたものでなければならないと主張します。WM は、システムのダイナミクスを決定し説明するエンティティのプロパティ、エンティティ間の相互作用、およびエンティティと環境の相互作用もキャプチャする必要があります。私たちは、サポート対象のタスクに基づいた因果 WM (CWM) の正式な定義を提供し、世界モデリングを因果表現学習、オブジェクト中心学習、因果発見、構造因果モデル、モデルベースの意思決定における既存の研究と結び付けます。最後に、CWM を識別可能性に関する文献に関連付け、WM のコンポーネントがいつデータから復元できるか、およびどの程度の同等性までを明確にします。これにより、WM を因果推論と情報に基づいた意思決定をサポートする表現と構造に根付かせることができます。

原文 (English)

A Unifying Perspective on Causal World Models: From Observations to Representations to Structure

World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics. We argue that useful WMs must go beyond generative capabilities alone: they should also capture entity properties, entity-to-entity interactions, and entity-to-environment interactions that determine and explain the dynamics of a system. We provide a formal definition of Causal WMs (CWMs) grounded in the tasks they are intended to support, connecting world modelling with existing work in causal representation learning, object-centric learning, causal discovery, structural causal models, and model-based decision-making. Finally, we relate CWMs to the literature on identifiability, clarifying when the components of a WM can be recovered from data and up to which equivalence. With this, we ground WMs in representations and structures that support causal reasoning and informed decision-making.

13:00 JSTLLM/生成AIエージェント

MARC v1: 臨床 AI 推論と調整のためのオープンソース マルチエージェント フレームワーク

我々は、臨床推論のためのモノリシック LLM プロンプトを決定論的なマルチエージェント オーケストレーションに置き換えるオープンソース フレームワークである Multi-Agent Reasoning and Coordination (MARC) を紹介します。 MARC は、明示的なコンテキストの受け渡しと追跡可能な中間出力を使用して、抽出、推論、回答生成、および評価を行う役割に特化したエージェントを調整し、段階的な障害の属性を可能にします。さらに、平易な言語の説明からタスク固有のエージェント プロンプトを生成する Decomposer モジュールを導入し、手動によるプロンプト エンジニアリングを排除します。このフレームワークは、API ベースのデプロイメントとローカル CPU 互換のデプロイメントの両方をサポートしており、コードを変更することなく YAML 経由で完全に構成可能です。 MARC は、モデルに依存せず、解釈可能で、プログラミングの専門知識がなくても臨床分野の専門家がアクセスできるように設計されています。完全なフレームワークは https://github.com/Penn-RAIL/MARC-v1 で入手できます。

原文 (English)

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. We additionally introduce a Decomposer module that generates task-specific agent prompts from a plain-language description, eliminating manual prompt engineering. The framework supports both API-based and local CPU-compatible deployments and is entirely configurable via YAML, without code modifications. MARC is designed to be model-agnostic, interpretable, and accessible to clinical domain experts without programming expertise. The full framework is available at https://github.com/Penn-RAIL/MARC-v1.

13:00 JST研究/論文

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and t…

13:00 JSTLLM/生成AIエージェントハードウェア/半導体GPT / ChatGPT

QuoteBench: スコアの一致によってコマンド パスの障害がどのように隠蔽されるか

LLM コーディング エージェントは、モデル出力をシリアル化、ラップ、再解析するインターフェイスを通じて Bash コマンドを発行します。一致した実行スコアだけでは、コマンド生成エラーと生成後に発生したエラーを区別できません。 QuoteBench は、14 のインシデント派生ファミリーからの 56 のワンショット タスクに対する正確な最終状態の検証によってこの境界を測定し、意図的にエスケープされない追加された 1 つのパーサーを中心とした実行トランスポートとの生成コントラクトを横断します。補間ポイントでのエスケープは、再生された各応答の生のパス結果を再現するため、開示された境界の下での回復は、世代を変更するモデルから行われる必要があります。 8 つの同じウィンドウ構成で、追加されたパーサーを通じて同じ応答を再生すると、成功率が 55.4 ~ 73.2 パーセント ポイント低下します。開示は 6 つの構成で 30.4 から 60.7 ポイント回復し、他の 2 つの構成ではゼロまたはわずかにマイナスになります。生の生成はフロンティアではほぼ飽和状態です。境界適応は依然としてモデルを分離するものです。 GPT-5.6-sol の -3.6 ポイントの一致ギャップにより、-64.3 ポイントのダメージと +60.7 ポイントの補正が隠蔽されます。デプロイメント構成によりモデルの順序が変更されます。26 の比較可能なペアのうちの 1 つの反転は明確で、残りの 4 つは単一タスクのマージンにあります。コマンド発行エージェントの評価では、一致したスコアをモデル固有のプロパティとして扱うのではなく、モデル構成、生成コントラクト、実行パス、操作点、および最終状態のバリデーターを報告する必要があります。

原文 (English)

QuoteBench: How Matched Scores Can Hide Command-Path Failures

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models. GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation. The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous and four more sit on single-task margins. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.

13:00 JSTLLM/生成AI研究/論文

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis gener…

13:00 JST研究/論文

言語モデル時代の AI アカウンタビリティ エコシステム

この記事では、責任エコシステムに基づいた AI における責任のフレームワークをレビューし、更新します。私たちは、一般向けに大規模言語モデルがリリースされて以来の最新の開発を考慮してフレームワークを更新しています。私たちは、元の AI アカウンタビリティ エコシステムに対して 3 つの相互に関連したアップデートを提案します。(i) アカウンタビリティ エコシステムの方向性を AI インフラストラクチャとサプライ チェーンに再設定すること、(ii) 分散型システムの改善をサポートする結果の監視と問題の特定をより重視すること、(iii) 現実の言語モデルの予測不可能性という新たなリスクを考慮して、エンドユーザーのアカウンタビリティを組み込むことです。まとめると、これらのアップデートは、フロンティア AI アプリケーションが業界固有の監視を持つ単一の識別可能な主体によって制御される個別の製品としてモデル化できるシステムから、分散型、継続的、制度化された説明責任への移行を示しています。

原文 (English)

The AI Accountability Ecosystem in the Era of Language Models

This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in light of the latest developments since the release of Large Language Models for general public use. We propose three interlinked updates to the original AI accountability ecosystem: (i) reorienting the accountability ecosystem to AI infrastructure and supply chains, (ii) providing greater emphasis on outcomes monitoring and identification of issues that support decentralized system improvement, and (iii) incorporating end-user accountability given the new risks of unpredictability of language models in-the-wild. Collectively, these updates mark a shift towards accountability as distributed, continuous, and institutionalized, away from a system in which frontier AI applications can be modeled as discrete products controlled by single identifiable actors with industry-specific oversight.

13:00 JSTLLM/生成AI

LLM は制約を知っているが使用していない: 実用的な制約推論におけるアクティベーションのボトルネック

顕著な表面キューが暗黙的な実現可能性制約と競合する場合、LLM は失敗することがよくありますが、集計精度により、真の制約推論と保守的なデフォルトが混同されます。この区別を条件付き制約のアクティブ化として形式化します。制約は、制約の存在と非存在のプロンプト間で対称的に内部的にエンコード (ナレッジ) されます (対称) が、決定にルーティングされる場合 (ルーティング) と、ドナーのアクティブ化 (修復) によって修復できる場合のみです。 14 モデルにわたるカルテット診断により、2 つの故障モードが明らかになります。 2 つのオープンウェイトのプローブは $88\%$ を超える制約をデコードしますが、アクティベーション パッチによって一方 ($+6.4$ nats) は修復され、もう一方 ($-0.07$) は修復されません。緩和のフロンティアでは、修復コーナーに到達する促進的な介入はありません。すべてが単一の仲介経路を通じて保守的なバイアスを増大させます-前提条件の言及。隠れた制約の障害は、知識の問題ではなく、ルーティングの問題です。

原文 (English)

LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning

When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional constraint activation: the constraint is internally encoded (Knowledge) symmetrically across constraint-present and -absent prompts (Symmetry), yet only sometimes routed into the decision (Routing) and repairable by a donor activation (Repair). A quartet diagnostic over 14 models reveals two failure modes; probes on two open weights decode the constraint above $88\%$, yet activation patching repairs one ($+6.4$ nats) and not the other ($-0.07$). On a mitigation frontier, no prompted intervention reaches the repair corner: all inflate conservative bias through a single mediation pathway -- prerequisite mention. Hidden-constraint failure is a routing problem, not a knowledge problem.

13:00 JSTLLM/生成AIGPT / ChatGPT

What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting

Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a con…

13:00 JSTLLM/生成AIエージェント

なぜ AI エージェントはルールを破るのか?フレーミング、コンテキスト、ソーシャルシグナルがコンプライアンスをどのように形成するか

罰則を指定すると、逆説的に、法的義務が違反に有利な費用対効果の計算に変換される可能性があります。私たちは、この施行情報のパラドックスが AI エージェントで体系的に発生することを実証します。ほとんどの AI 安全性評価ではモデルが失敗するかどうかがテストされますが、私たちは法律と経済学のコンプライアンス理論を診断ツールとして適用して、その理由を調査します。我々はコンプライアンス理論を比喩としてではなく経験的仮説として扱い、それぞれが異なるモデルクラスの動作を予測することを示します。私たちは、エンタープライズ調達チャットボットとして動作する 12 の命令調整された言語モデルにわたって仮説を評価します。抑止力、正当性、表現法則の理論に基づいて、安全性を細かく調整したモデルは広範なコンプライアンスを維持する一方、タスク最適化モデルやエージェントモデルは規制シグナルを単なる最適化パラメーターとして扱うことを示します。これらの後者のモデルは、低い執行罰や非命令語句など、理論によって予測される条件の下では準拠できません。すべてのモデルにおいて、金銭的インセンティブ、管理上の要求、同僚の成果、または従業員のプレッシャーの導入は、大規模なコンプライアンス違反を引き起こします。 AI 調達エージェントは、標準的な連携ベンチマークでは捉えられない方法で、ローカル ユーザーの目的を満たすために、組織的に規制上の制約に違反しています。結局のところ、ルールの埋め込みだけではコンプライアンスを達成することはできません。モデルの選択自体がガバナンスの決定であり、ベンチマークベースの評価はコンプライアンスを重視した展開には不十分です。

原文 (English)

Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance

Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI safety evaluations test whether models fail, we investigate why, applying compliance theory from law and economics as a diagnostic tool. We treat compliance theories not as metaphors but as empirical hypotheses and show that each predicts the behavior of a distinct model class. We evaluate our hypotheses across twelve instruction-tuned language models operating as enterprise procurement chatbots. Drawing on theories of deterrence, legitimacy, and expressive law, we show that safety-fine-tuned models maintain compliance broadly, while task-optimized and agentic models treat regulatory signals as mere optimization parameters. These latter models fail to comply under conditions predicted by theory, such as low enforcement penalties and non-command phrasing. Across all models, introducing financial incentives, managerial demands, peer outcomes, or employee pressure produces large compliance failures. AI procurement agents systematically violate regulatory constraints to satisfy local user objectives in ways not captured by standard alignment benchmarks. Ultimately, compliance cannot be achieved by rule embedding alone; model selection is itself a governance decision, and benchmark-based evaluation is insufficient for compliance-sensitive deployments.

13:00 JSTLLM/生成AI研究/論文

When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models

People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are no…

13:00 JSTLLM/生成AI研究/論文

ネパール語自動音声認識用の多言語事前トレーニング済みモデルの比較分析

多言語の事前トレーニング済みモデルは名目上ネパール語をサポートしていますが、単一の微調整プロトコルの下でそれらを比較した管理されたベンチマークはありません。 OpenSLR SLR54 ネパール語コーパス上で、CTC 自己教師あり自己回帰エンコーダ デコーダ、およびハイブリッド Conformer-CTC アーキテクチャにまたがる 6 つの事前トレーニング済みモデル (XLSR-53、IndicWav2Vec、MMS-1B、Whisper-Medium、Whisper-Large-v3-Turbo、および Conformer-Hi) を同一の前処理を使用して微調整しました (約 165 時間)。スプリット、オプティマイザー、家族に合わせた学習率スケジュール。 3 つの独立したテスト セット (OpenSLR、FLEURS、Common Voice) で単語誤り率 (WER)、文字誤り率 (CER)、およびリアルタイム係数 (RTF) を評価します。 Whisper-Large-v3-Turbo (14.76% WER) と IndicWav2Vec (14.89% WER) は、パラメーター ギャップが 9 倍、事前トレーニング データ ギャップが 40 倍であるにもかかわらずトップで並んでおり、事前トレーニングにおける言語族の近接性がドメイン内ネパール語の生のスケールの代わりになり得るという直接的な経験的証拠を提供しています。 CTC デコーダは、同じ精度で自己回帰 Whisper よりも最大 29 倍高速に実行され、遅延予算がどのような場合でも、実際の導入の優先順位が CTC に切り替わります。大規模多言語事前トレーニング (MMS-1B) では、FLEURS でドメイン外の劣化が最小 (+12.55 pp) であり、ドメイン内のピーク精度ではなくスケールが堅牢性を獲得していることを示しています。結果として得られたベンチマークは、ネパール ASR の最初の標準化されたマルチモデルの効率を意識した参照番号を提供します。

原文 (English)

Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition

Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium, Whisper-Large-v3-Turbo, and Conformer-Hi) spanning CTC self-supervised, autoregressive encoder-decoder, and hybrid Conformer-CTC architectures, on the OpenSLR SLR54 Nepali corpus (~165 hours) using identical preprocessing, splits, optimizer, and family-matched learning-rate schedules. We evaluate Word Error Rate (WER), Character Error Rate (CER), and Real-Time Factor (RTF) on three independent test sets (OpenSLR, FLEURS, Common Voice). Whisper-Large-v3-Turbo (14.76% WER) and IndicWav2Vec (14.89% WER) tie at the top despite a 9x parameter gap and 40x pretraining-data gap, providing direct empirical evidence that language-family proximity in pretraining can substitute for raw scale for in-domain Nepali. CTC decoders run up to 29x faster than autoregressive Whisper at the same accuracy, flipping the practical deployment preference toward CTC under any latency budget. Massively multilingual pretraining (MMS-1B) yields the smallest out-of-domain degradation on FLEURS (+12.55 pp), indicating that scale buys robustness rather than peak in-domain accuracy. The resulting benchmark provides the first standardized, multi-model, efficiency-aware reference numbers for Nepali ASR.

13:00 JSTLLM/生成AIビジネス/資金調達

AnchorSIPS: 証拠に裏付けられた精神病リスク症状測定のための合成データセットおよび評価リソース

精神病リスク評価のための AI の進歩は、データアクセスのボトルネックによって制限されています。実際の臨床面接は、プライバシー、ガバナンス、同意の制約により共有することが困難です。我々は、トランスクリプトに基づいた測定対象者との10,000の構造化精神病リスクインタビューの合成データセットであるAnchorSIPSを紹介します。各面接は、臨床医が実施する精神病リスク面接である Mini-SIPS をモデルにしています。これは、病歴、24の症状に関する質問、患者が肯定した項目の追跡証拠、妄想様症状(異常な信念)、幻覚様症状(異常な知覚)、および混乱したコミュニケーションに関する決定、明らかな精神病レベルの症状(「率直な精神病」)の除外、および軽度または初期の精神病症状の高リスク状態である軽度精神病症候群(APS)の最終診断を記録する。 APS 診断は独立したラベルではありません。それは、以前の承認、フォローアップの詳細、症状クラスの決定、および率直な精神病チェックに依存します。すべての中間決定は、それをサポートする転写ターンに固定されています。 AnchorSIPS は、計画後実現パイプラインによって生成されます。隠されたケースシートが患者の臨床状態を特定し、決定論的プランナーがインタビュー構造を修正し、LLM が検証と限定的修復の下で患者の発話のみを認識します。生成前にラベルと構造を修正すると、マルチターン LLM ダイアログに特有のターン間の不一致が回避されます。 7 つの LLM ベースラインにわたって、モデルは大まかな決定を回復しますが、フォローアップの詳細を抽出したり、裏付けとなる転写ターンを引用したりすることができないため、最終ラベルのパフォーマンスはインタビューの能力を過大評価します。 AnchorSIPS は、証拠の抽出、転写に基づいた測定、部分開示の下での不確実性に関する研究を目的としています。

原文 (English)

AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement

Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured psychosis-risk interviews with transcript-grounded measurement targets. Each interview is modeled on Mini-SIPS, a clinician-administered psychosis-risk interview. It captures history, 24 symptom questions, follow-up evidence for items the patient affirms, decisions about delusion-like symptoms (unusual beliefs), hallucination-like symptoms (unusual perceptions), and disorganized communication, exclusion of clear psychotic-level symptoms ("frank psychosis"), and a final attenuated psychosis syndrome (APS) diagnosis, a high-risk state of milder or early psychotic symptoms. The APS diagnosis is not a standalone label. It depends on earlier endorsements, supporting follow-up details, symptom-class decisions, and the frank-psychosis check. Every intermediate decision is anchored to its supporting transcript turns. AnchorSIPS is generated by a plan-then-realize pipeline. A hidden case sheet specifies the patient's clinical state, a deterministic planner fixes the interview structure, and an LLM realizes only the patient utterances under validation and bounded repair. Fixing labels and structure before generation avoids the inter-turn inconsistencies typical of multi-turn LLM dialogue. Across seven LLM baselines, models recover coarse decisions but fail to extract follow-up details or cite supporting transcript turns, so final-label performance overstates interview competence. AnchorSIPS is intended for research on evidence extraction, transcript-grounded measurement, and uncertainty under partial disclosure.

13:00 JSTLLM/生成AI

適応型アテンションマッチングによる推論のための思考認識型 KV キャッシュ圧縮

推論言語モデルは、キー/値 (KV) キャッシュが直線的に増加し、デコード中にメモリのボトルネックになる長い思考連鎖 (CoT) シーケンスを生成します。既存の圧縮手法は、推論軌跡をフラットなトークン シーケンスとして扱い、均一な圧縮を適用します。ステップごとに重要性が大幅に異なる CoT 推論の階層構造を無視します。私たちは \textbf{思考認識型注意マッチング (TAM)} を提案します。これは、(i)~軌跡を推論ブロックに分解する思考セグメント化、(ii)~各セグメントの重要性とサイズに基づいて圧縮予算を割り当てる適応型予算割り当て、および (iii)~注目度の高い推論アンカーを保存する重要なトークン保護の 3 つのメカニズムを通じてこの構造を活用します。割り当てルールが凸誤差モデルの下で最適であること、および逐次圧縮の下での累積誤差が制限されたままであることを証明します。 Qwen3-4B を使用した AIME 2024 および MATH-500 での実験では、競合する精度を維持しながら、TAM が同じメモリ フットプリントでの均一な圧縮よりも精度が向上し、定期的な圧縮によりピーク メモリが 3.1 ~ 3.2\,GB (65\% 削減) に制限されることがわかりました。

原文 (English)

Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching

Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding. Existing compaction methods treat reasoning trajectories as flat token sequences and apply uniform compression, ignoring the hierarchical structure of CoT reasoning where different steps vary drastically in importance. We propose \textbf{Thought-Aware Attention Matching (TAM)}, which exploits this structure through three mechanisms: (i)~thought segmentation that decomposes the trajectory into reasoning blocks, (ii)~adaptive budget allocation that assigns compression budget based on each segment's importance and size, and (iii)~pivotal token protection that preserves high-attention reasoning anchors. We prove that the allocation rule is optimal under a convex error model and that cumulative error under sequential compaction remains bounded. Experiments on AIME 2024 and MATH-500 with Qwen3-4B show that TAM improves accuracy over uniform compaction at the same memory footprint, with periodic compaction bounding peak memory to 3.1--3.2\,GB (a 65\% reduction) while maintaining competitive accuracy.

13:00 JSTLLM/生成AI画像/動画生成

視覚言語モデルは脆弱な多言語連想子である

視覚言語モデルは、視覚的エンティティをテキスト属性に関連付ける必要があります。入力言語が変化したときに、これらの関連付けや概念のバインディングが安定したままであるかどうかは、まだ解明されていません。複数の言語にわたってコンテキストとクエリの言語を変えるベンチマークである M$^2$BIND を紹介します。私たちは、タスクパフォ​​ーマンス指標を通じて外因的にバインディングを評価するとともに、因果関係のある介入を通じて本質的にバインディングを評価します。バインディングは言語不変ではないことがわかりました。クロスファミリーおよびクロススクリプト設定は、重大なバインディングの崩壊を引き起こし、モデルの内部バインディング計算が後の層にシフトし、因果関係の強さを失います。密接に関連した言語は、関連性を比較的よく保持します。より広い意味で、私たちの調査結果は、多言語設定でグローバルに展開された VLM が、単一言語の評価で観察されたのと同じ関連性の品質を維持すると仮定できないことを示しています。

原文 (English)

Vision-Language Models are Fragile Multilingual Associators

Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark varying the language of the context and query across multiple languages. We evaluate binding both extrinsically through task performance metrics and intrinsically through causal interventions. We find that binding is not language-invariant: cross-family and cross-script settings trigger significant binding collapse, with the model's internal binding computation shifting to later layers and losing causal strength. Closely related languages preserve associations comparatively better. In a broader sense, our findings indicate how VLMs deployed globally in multilingual settings cannot be assumed to maintain the same association quality observed in monolingual evaluation.

13:00 JSTLLM/生成AILlamaQwen

Steering the Language Axis: From Linear Decodability to Causal Control

Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly u…

13:00 JSTLLM/生成AI

StorySpark: Module-wise Evolutionary Search for Story Premise Generation

A story premise is the creative spark from which a full narrative can grow. Yet LLM-based story generation has mostly emphasized later-stag…

13:00 JSTLLM/生成AIGPT / ChatGPTQwen

Mimicry without understanding: the origins of decision bias in large language models

Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms…

13:00 JSTLLM/生成AIGPT / ChatGPT

StreamReason-Bench: 大規模な言語モデルはイベント時のストリーム処理セマンティクスについて推論できますか?

ストリーミング システムでは、パイプラインの作成、アラートのトリアージ、ログの読み取りなど、大規模言語モデル (LLM) に対する手作業が増えていますが、そのすべてにおいて、モデルがイベント時のストリーム処理がどのように動作するかを認識していることが前提となっています。私たちはその仮定を正面からテストします。 StreamReason-Bench は、モデルにイベント時のストリーム プロセッサの代わりをするように要求します。ウィンドウ化されたクエリと順序が乱れたイベントのストリームが与えられると、どのウィンドウが (集計とともに) 起動し、どのイベントが遅れてドロップされるかを報告する必要があります。回答キーは、Dataflow モデル セマンティクスの小規模なリファレンス実装から得られるため、エンジンを実行せずに、部分クレジット行 F1 を使用して正確に採点できます。タンブリング、ホッピング、セッション、および処理時間ウィンドウをカバーする 600 個の生成されたアイテムでは、モデルはイベント時間でのパフォーマンスが低下します。直接答えるように言われましたが、実際にその指示に従ったモデルで 34% の完全一致をクリアしたモデルはありません。いくつかのモデルでは、思考連鎖 (CoT) が約 2 倍になり (GPT-4o は 0.34 から 0.48 になります)、デフォルトで推論を行うフロンティア モデルは 1 つだけ (0.85) になります。ウォーターマークも遅延もなく、処理時間の制御は、すべての有能なモデルでほぼ解決されます。このギャップは、ウィンドウ処理や演算ではなく、イベント時および遅延データの処理が難しい部分であることを示しています。ウィンドウの種類ごとにエラーを並べ替えると、同じことがわかります。遅延データの間違いはイベント時間ウィンドウを支配してコントロール上で消え、セッション ウィンドウはほとんどの場合、セッション境界に該当する場所で失敗します。

原文 (English)

StreamReason-Bench: Can Large Language Models Reason about Event-Time Stream-Processing Semantics?

Streaming systems increasingly hand work to large language models (LLMs) -- writing pipelines, triaging alerts, reading logs -- and all of it assumes the model knows how event-time stream processing behaves. We test that assumption head-on. StreamReason-Bench asks a model to stand in for an event-time stream processor: given a windowed query and a stream of out-of-order events, it has to report which windows fire (with their aggregates) and which events are dropped as late. The answer key comes from a small reference implementation of Dataflow-model semantics, so we can grade exactly, and with a partial-credit row-F1, without running an engine. On 600 generated items covering tumbling, hopping, session, and processing-time windows, the models do poorly on event time. Told to answer directly, no model that actually follows the instruction clears 34% exact match; chain-of-thought (CoT) roughly doubles that for several of them (GPT-4o goes from 0.34 to 0.48), and only one frontier model that reasons by default comes near solving the set (0.85). A processing-time control, with no watermarks and nothing late, is almost solved by every capable model. That gap points to event-time and late-data handling, not windowing or arithmetic, as the hard part. Sorting errors by window type tells the same story: late-data mistakes dominate the event-time windows and vanish on the control, and session windows mostly fail on where the session boundaries fall.

13:00 JSTLLM/生成AI

穴居人から専門アナリストへ: 可変 LLM タスクのエネルギー消費

エネルギー需要の増加と人工知能 (AI) による環境への影響により、AI を活用したデータセンター開発に十分な低コストの電力を供給することに大きな関心が集まっています。これらの課題に対処するための需要側管理の能力に関する研究はさらに限られています。需要の量やタイミングを小売、企業、その他の組織の行動からずらすことは妥当な選択肢ですが、それは需要関連の行動の変化が AI の環境や電力への影響に重要な影響を与える場合に限られます。この記事では、行動可塑性の高い 4 つの小売 (消費者) ユーザーの行動をテストし、技術的な軽減の可能性を評価します。研究では、非推論モデルは十分な品質を提供しながら推論モデルに比べてエネルギー消費量が 20 分の 1 近くに抑えられ、毎日の使用量を想定した場合、米国の少なくとも 141,000 世帯の年間電力需要に等しい量を節約できると結論付けています。単純な即時変更により、非推論モデルを使用してエネルギー消費を最大 65% までさらに削減できます。具体的には、ベースラインとの類似性を最も高く維持する実践により、電力需要が 4 ~ 35% の範囲で削減されます。これは、米国の最大 7,200 世帯の年間電力需要に相当します。 AI の進歩により、環境や電力への影響を正確に見積もることは困難になっていますが、この結果は、大多数のユーザーを対象とした、侵入を最小限に抑えた特定のベスト プラクティスにより、AI によって課せられるエネルギーと環境への負担を軽減できることが確認されました。

原文 (English)

From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks

The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying sufficient low-cost electricity for AI-driven data center development. Research on the ability of demand-side management to address these challenges has been more limited. Shifting the amount or timing of demand from retail, corporate, and other organizational behaviors is a plausible option but only if changes in demand-related behavior have important effects on the envi- ronmental and electricity effects of AI. This article tests four retail (i.e., consumer) user behaviors with high behavioral plasticity to assess their technical abatement potential. The research concludes that non- reasoning models provide sufficient quality while consuming close to one-twentieth of energy compared to reasoning models, saving an amount equal to the annual electricity requirement of at least 141,000 US households under daily usage assumptions. Simple prompt modifications can yield additional reduc- tions in energy consumption by up to 65% using non-reasoning models. Specifically, the practice that maintains the highest degree of similarity with the baseline reduces electricity demand in the range of 4 to 35%, an amount equal to the annual electricity requirement of up to 7,200 US households. Although AI advancements make precise estimates of environmental and electricity impacts difficult to assess, the results confirm that certain minimally intrusive best practices aimed at the majority of users can reduce the energy and environmental burdens imposed by AI.

13:00 JST研究/論文

Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection

Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects…

13:00 JST研究/論文Copilot

Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF

AI governance frameworks can be known, used, and implemented in form without becoming governance in practice. This paper examines that prob…

13:00 JSTエージェント研究/論文

Humans are Missing from AI Coding Agent Research

Recent progress in AI coding agent research has led to rapid improvements in agents' ability to autonomously perform complex software engin…

13:00 JST研究/論文

Measuring Curriculum-Labor Market Alignment at the Scale of a Program Portfolio

A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor…

13:00 JSTエージェントビジネス/資金調達

Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles

Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent conte…

13:00 JST研究/論文

EU-ETS under attack? The impact of carbon price suppression on the decarbonization of the power sector

European countries are debating policies to mitigate the increased energy costs caused by renewed geopolitical tensions, while pursuing dec…

13:00 JSTエージェント

FluctlightDB: A Memory Model of Data for AI Agents

For fifty years, data systems have answered two questions. The relational model asked which records match a predicate; the vector model ask…

13:00 JSTLLM/生成AI

Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities

Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays…

13:00 JSTLLM/生成AI

From Observation to Intervention: Memory in Brains and Large Language Models

Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional ques…

13:00 JSTLLM/生成AI

Query Timing Produces Opposite Positional Biases Between LLMs and Humans

Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by…

13:00 JSTLLM/生成AI研究/論文

Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models

Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can b…

13:00 JST研究/論文

A Hierarchical Energy-Based Model for Multimodal Cognition

We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cogn…

13:00 JSTエージェント

SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents

Web agents often struggle to generalize to unseen websites because they lack website-specific supervision. Recent exploration-based data sy…

13:00 JSTエージェント研究/論文

Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review

This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specifi…

13:00 JST研究/論文

Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection

Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the…

13:00 JST研究/論文

SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization

Neural combinatorial optimization (NCO) relies on parallel solution sampling for training, yet existing methods fail to fully exploit the r…

13:00 JST研究/論文

Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts

Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate mo…

13:00 JST研究/論文

SchemaLink: An Intelligent Web Editor for LinkML Schema Curation

Motivation: LinkML is a suitable language for the representation of the structural and content constraints of different kinds of biomedical…

13:00 JSTLLM/生成AI

Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling

AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them. We analyzed 36…

13:00 JSTビジネス/資金調達

What Makes a Peer? Valuation-Anchored Similarity in Private Markets

As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, ident…

13:00 JST研究/論文

Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks

Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped…

13:00 JST画像/動画生成

PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping

A critical challenge in deploying online HD map construction systems to real-world scenarios is the scarcity of labeled training data, whic…

13:00 JSTLLM/生成AI

LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning

Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals. In these se…

13:00 JSTLLM/生成AI画像/動画生成

EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory

Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstr…

13:00 JSTLLM/生成AIハードウェア/半導体GPT / ChatGPTQwen

Novels generated by language models show compressed formal variation

While large language models can generate entire novels, there is little information about the level of formal variation in their output ove…

13:00 JST研究/論文

Interpretable Causal Discovery via Causal-Effect Constraints

Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely t…

13:00 JST研究/論文

Demand Transfer Estimation at Scale via Restricted Logit Modeling

Item demand forecasting is an integral component of store assortment optimization. Existing literature focuses on learning a suitable custo…

13:00 JST画像/動画生成

Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face cr…

13:00 JSTLLM/生成AI

Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks

Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing,…

13:00 JST研究/論文

HybridSB-MoE: Dual-Domain Schr\"odinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement

Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserv…

13:00 JSTLLM/生成AI

Error-Aware Reverse Auction Mechanism for Large Language Model Routing

Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a ce…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval

While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governi…

13:00 JSTLLM/生成AI

PatientAct: Theory-Grounded Mental Health Client Simulation

LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However…

13:00 JSTエージェント

SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization

Logic synthesis transforms RTL designs into gate-level netlists, where PPA results are highly sensitive to the choice of optimization comma…

13:00 JSTエージェント

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome r…

13:00 JSTLLM/生成AI

Memorization Diagnostics for Code LLMs Should be Scale-Aware

The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current li…

13:00 JSTLLM/生成AI

CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives

Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and caus…

13:00 JSTエージェントGemma

PIPES: Securing Agent Perception with Provenance and Priors

Tool-using agents consume external data from sources with different levels of trust, yet tool responses rarely identify who produced each c…

13:00 JST画像/動画生成ビジネス/資金調達規制/政策

Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors

The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized rep…

13:00 JST研究/論文

Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing

Online platforms increasingly compare many adaptive decision policies---ranking systems, recommendation algorithms, pricing rules, and lang…

13:00 JSTLLM/生成AI

From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options

Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they ju…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from ea…

13:00 JST画像/動画生成

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural la…

13:00 JST研究/論文

FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation

Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. However, existing methods suffer from a previou…

13:00 JSTLLM/生成AIGemma

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model interna…

13:00 JST画像/動画生成エージェントロボティクス

BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving

Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, howev…

13:00 JST研究/論文

A Compositional Theory of Curvature in Probabilistic Circuits

Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tracta…

13:00 JST画像/動画生成

SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data

Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit thre…

13:00 JSTエージェントビジネス/資金調達

Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation

Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200…

13:00 JST画像/動画生成

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-L…

13:00 JST研究/論文

EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction

RNA-Protein Interactions (RPIs) are critical for regulating cellular functions. While traditional wet-lab experiments for RPI detection are…

13:00 JSTLLM/生成AI

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure de…

13:00 JSTLLM/生成AIエージェント

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference

The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existi…

13:00 JSTハードウェア/半導体ビジネス/資金調達

H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities

Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the mul…

13:00 JST研究/論文

AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization

As quantum computing progresses from proof-of-principle demonstrations toward practical utility, a significant impediment is the need to au…

13:00 JST研究/論文

The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use

Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor…

13:00 JSTLLM/生成AIエージェント

Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents

The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have…

13:00 JST研究/論文

Generative Universal Multimodal Retrieval with Dual-role Identifiers

Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipel…

13:00 JSTエージェント

Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language

The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be writ…

13:00 JST画像/動画生成エージェントビジネス/資金調達

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accid…

13:00 JST研究/論文Microsoft

Operationalizing Cyber Threat Intelligence with GraphRAG

When a security researcher publishes a report on a cyberattack, detection engineers are supposed to turn it into working detection rules. I…

13:00 JSTLLM/生成AIハードウェア/半導体DeepSeek

TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes

In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or…

13:00 JSTビジネス/資金調達

LOB-ID: Evaluating Synthetic Market Data by Inception Distances

Generative models of limit orderbook (LOB) data have advanced rapidly, but their evaluation often focuses on stylised facts and selected ma…

13:00 JST研究/論文

Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization

Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convent…

13:00 JSTLLM/生成AI画像/動画生成研究/論文Gemini

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growi…

13:00 JSTLLM/生成AI研究/論文

LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approache…

13:00 JST研究/論文

LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service

As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models…

13:00 JSTLLM/生成AI

Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering

Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for…

13:00 JSTLLM/生成AI画像/動画生成研究/論文GemmaQwen

TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint

When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internall…

13:00 JSTLLM/生成AI

GEM: A Generative Embedding Model Bridging Reasoning and Retrieval

Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, convent…

13:00 JST画像/動画生成研究/論文

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and int…

13:00 JST画像/動画生成

CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport

While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual toke…

13:00 JSTLLM/生成AILlamaQwen

Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales

Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather…

13:00 JST画像/動画生成ビジネス/資金調達

GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport

Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introd…

13:00 JST研究/論文

Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data

As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as im…

13:00 JSTLLM/生成AIGemini

Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models

Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective ex…

13:00 JST研究/論文

Into the ORBIT for Time Series: Training Regimes for Foundation Models

Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale hete…

13:00 JSTLLM/生成AI画像/動画生成ビジネス/資金調達研究/論文GPT / ChatGPTGemini

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what t…

13:00 JSTLLM/生成AIGemma

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model

We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a…

13:00 JST研究/論文

Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks

Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back…

13:00 JST研究/論文

Physics-informed distribution of relaxation times estimation and latent-space condition monitoring of solid oxide fuel and electrolysis cells from electrochemical impedance spectroscopy

Estimating the distribution of relaxation times (DRT) fromelectrochemical impedance spectroscopy (EIS) is an ill-posed inverse problem that…

13:00 JSTLLM/生成AI

Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services

We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while…

13:00 JSTLLM/生成AI

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contai…

13:00 JST研究/論文ClaudeGPT / ChatGPT

Training AI Scientists to Replicate Research

The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for…

13:00 JST研究/論文

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challengi…

13:00 JST画像/動画生成

Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs

This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Advers…

13:00 JST研究/論文

Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks

6G networks will not be serving as communication infrastructures only; rather, they are expected to evolve into intelligent systems, where…

13:00 JSTエージェントロボティクス

Deliberate Practice: Learning Robot Skills under a Budget

We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active sk…

13:00 JSTLLM/生成AINVIDIA

Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix m…

13:00 JSTLLM/生成AI

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhib…

13:00 JST研究/論文

Algebraic Decomposition Theory for Transformer Length Generalization

Transformer-based language models are known to sometimes generalize to sequences longer than seen during training, but we lack a precise ch…

13:00 JST画像/動画生成ロボティクス

ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models

Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is especially limiting in wrist-c…

13:00 JST画像/動画生成ロボティクス

UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models

Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and perf…

13:00 JSTLLM/生成AIOpenAI

CAPRI: Contract-Aware Proof Repair for Isabelle

We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theor…

13:00 JSTLLM/生成AI画像/動画生成

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and…

13:00 JST研究/論文

Concept Drift Detection and Adaptive Retraining of Malware Classification Models

Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning…

13:00 JSTLLM/生成AI

AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models

Analog circuit design is a time-consuming, iterative process in a nonlinear and high-dimensional design space that relies heavily on expert…

13:00 JSTLLM/生成AIエージェント

Synthetic Persona Pretraining: Alignment from Token Zero

As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes crit…

13:00 JSTLLM/生成AI

Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity

When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to…

13:00 JSTLLM/生成AI研究/論文GemmaQwen

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committ…

13:00 JST研究/論文

The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity

We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \emph{unmasking growth c…

13:00 JSTエージェント

Vero: Can AI Agents Build Formally Verified Software Repositories?

AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code gener…

13:00 JSTLLM/生成AIビジネス/資金調達

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is diffi…

13:00 JST画像/動画生成ロボティクスビジネス/資金調達研究/論文

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in…

13:00 JSTLLM/生成AI画像/動画生成エージェントハードウェア/半導体Claude

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic p…

13:00 JST研究/論文

MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence

Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic tr…

13:00 JSTエージェント

Foam-Agent: A Large Language Model-Based Multi-Agent Framework for Automating Computational Fluid Dynamics Workflows

Computational fluid dynamics (CFD) has been the main workhorse of computational physics, yet its steep learning curve and fragmented, multi…

13:00 JST研究/論文

Exploiting Symbolic Heuristics for the Synthesis of Domain-Specific Temporal Planning Guidance using Reinforcement Learning

Recent work investigated the use of Reinforcement Learning (RL) for the synthesis of heuristic guidance to improve the performance of tempo…

13:00 JST研究/論文

Identification of Probabilities of Causation: from Recursive to Closed-Form Bounds

Probabilities of causation (PoCs) are fundamental quantities for counterfactual analysis and personalized decision making. However, existin…

13:00 JSTLLM/生成AIエージェント研究/論文

PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research

Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging…

13:00 JST研究/論文

DomusFM: A Foundation Model for Event-Based Behavioral Monitoring in Smart-Homes

Smart-home sensor-based behavioral monitoring holds significant potential for healthcare, independent living, and early detection of functi…

13:00 JSTLLM/生成AIエージェント

Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design

We study mathematical discovery through the lens of neurosymbolic reasoning, where an AI agent powered by a large language model (LLM), cou…

13:00 JSTLLM/生成AIエージェント

Auditable Agents

LLM agents call tools, query databases, delegate tasks, and trigger external side effects. Once an agent system can act in the world, the q…

13:00 JST研究/論文
13:00 JSTLLM/生成AI

From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction

Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover re…

13:00 JSTエージェント

AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design

Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent…

13:00 JST研究/論文

CEON: 循環経済オントロジー ネットワーク

私たちの社会における資源利用の循環性を高めることは、持続可能性、つまりより循環的な経済への移行への道であると認識されています。そのためには、製品やコンポーネントの再利用、中古製品の再生と再製造、残り物や使用済みの材料のリサイクルなど、さまざまな循環戦略があります。これらの戦略を実現するには、インフラストラクチャ レベルで情報を共有し、製品ライフ サイクルに沿って業界セクター間で通信する必要があります。したがって、この情報共有と通信においてセマンティックな相互運用性を実現することが、循環性を高める鍵となります。しかし、製品ライフサイクルに関連する多くの関連業界セクターが関与する循環経済 (CE) 領域の知識表現は依然として課題です。このギャップを埋めるために、私たちは Onto-DESIDE プロジェクト内で Circular Economy Ontology Network (CEON) を開発しました。このオントロジー ネットワークは、分野横断的な概念を定義することで CE のギャップを埋め、セマンティクスを意識したデータの文書化を可能にすることを目的としています。建設、エレクトロニクス、繊維セクターにわたる業界横断的なデータ文書化シナリオを通じて CEON を実証します。

原文 (English)

CEON: Circular Economy Ontology Network

Increasing the circularity of resource use in our society has been recognized as a path to sustainability, i.e., transitioning into a more circular economy. There are many different circular strategies to do so, such as reusing products and components, refurbishing and remanufacturing used products, or recycling left-over or used materials. To enable these strategies, it is necessary to share information at the infrastructure level and to communicate between industry sectors along the product life cycle. Enabling semantic interoperability in this information sharing and communication is therefore a key to increasing circularity. However, knowledge representation for the circular economy (CE) domain, which involves many relevant industry sectors related to product life cycles, remains challenging. To bridge this gap, we developed the Circular Economy Ontology Network (CEON) within the Onto-DESIDE project. This ontology network aims to fill gaps in CE by defining cross-sectorial concepts and to enable semantics-aware data documentation. We demonstrate CEON through cross-industry data documentation scenarios spanning construction, electronics, and textile sectors.

13:00 JST研究/論文

Residual Modeling for High-Fidelity Learned Compression of Scientific Data

Lossy compression is essential for massive spatiotemporal data from scientific simulations. Learned compressors can achieve high compressio…

13:00 JST研究/論文

マルチタスク結合モデルからタスクエキスパートを回復する方法を学ぶ

マルチタスク モデルのマージは、複数のタスク固有の専門家を 1 つの統一モデルに統合することを目的としていますが、静的マージではパラメータの干渉が常に発生します。動的マージ モデルはこのギャップを埋めることを目的としていますが、多くの研究は、推論時にコストのかかるストレージと冗長なエキスパート コンポーネントの読み込みに依存しています。この研究では、タスク エキスパートの観点から、パラメータ干渉を、マージ プロセス中に各エキスパートに導入されるパラメータの摂動として見ます。このようなパラメータの摂動はアフィン変換としてモデル化でき、加算オフセットとして近似できることを示します。これらを動機として、パラメータ干渉を元に戻し、単一のマージされたチェックポイントからタスク エキスパートのパフォーマンスを回復するために、これらのオフセットを予測するフレームワークである Recover Task eXpert (ReTeX) を提案します。タスク ID が不明な場合に適切なエキスパートを回復するために、推論前にオフラインで計算された SVD 部分空間署名に基づくルーターフリーのタスク ID を導入します。推論時に、識別子は、指定された入力に対して部分空間が最小の射影残差をもたらすタスクを選択します。その結果、ReTeX は視覚領域と NLP 領域の両方で個人の専門家のパフォーマンスの 95% 以上を回復し、目に見えないタスクへの一般化を大幅に向上させます。重要なことに、パラメータ オフセット予測が、配布外 (OOD) タスクに対する専門知識の創発的適応補間につながることも示します。 ReTeX は、目に見えないタスクを処理するために、目に見える専門知識を適応的に補間します。私たちのコードは https://github.com/BAIKLAB/ReTeX で入手できます。

原文 (English)

Learning to Recover Task Experts from a Multi-Task Merged Model

Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference. In this work, from the perspective of task expert, we view parameter interference as parameter perturbation introduced to each expert during merging process. We show that such parameter perturbations can be modeled as affine transformation, which can be approximated as additive offsets. Motivated by these, we propose Recover Task eXpert (ReTeX), a framework that predicts those offsets, in order to undo parameter interference and recover task-expert performance from a single merged checkpoint. To recover the appropriate expert when task identity is unknown, we introduce a router-free task identifier based on SVD subspace signatures computed offline before inference. At inference, the identifier selects the task whose subspace yields the smallest projection residual for a given input. As a result, ReTeX recovers over 95% of individual-expert performance in both vision and NLP domains, while significantly improving generalization to unseen tasks. Crucially, we also show that the parameter offset prediction leads to emergent adaptive interpolation of expert knowledge for out-of-distribution (OOD) tasks. ReTeX adaptively interpolates seen expert knowledge to handle unseen tasks. Our code is available at https://github.com/BAIKLAB/ReTeX

13:00 JSTエージェント研究/論文

Alipay-PIBench: コーディング エージェント向けの現実的な決済統合ベンチマーク

支払いの統合は、要求の厳しいリポジトリ レベルのソフトウェア タスクです。エージェントは、適切な製品を選択し、調整されたクライアント/サーバー フローを実装し、支払い結果を検証し、トランザクションとビジネス状態の間の一貫性を維持する必要があります。現実的な Alipay 決済統合に関するコーディング エージェントを評価するためのベンチマークである Alipay-PIBench を紹介します。これには、9 つ​​の製品固有のプロジェクトと 18 のタスク インスタンスが含まれており、それぞれが基本的な機能完了シナリオと高度なリスク認識強化シナリオに編成されています。シナリオ固有のルーブリックは、決定論的な静的チェック、ユニットチェック、統合チェック、およびエンドツーエンドのチェックをサポートし、セマンティック要件に対する LLM 支援の評価によって補足されます。 6 つのコーディング エージェント モデルを評価し、ルーブリック合格率 (RPR) を報告します。スキルありの条件下では、平均 RPR は 68.58% から 91.37% の範囲です。 Alipay 決済統合スキルへのアクセスにより、スキルなしの状態と比較して平均 RPR が平均 10.31 パーセント ポイント向上しますが、その向上はモデル、製品、シナリオによって異なります。メソッドレベルの結果は、ソースレベルの完了、実行可能な支払い動作、支払いドメインの要件を区別します。 Alipay-PIBench は、モデルの機能を診断し、支払い統合における構造化されたガイダンスを評価するための制御された設定を提供します。

原文 (English)

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states. We introduce Alipay-PIBench, a benchmark for evaluating coding agents on realistic Alipay payment integration. It contains nine product-specific projects and 18 task instances, each organized into Basic functional-completion and Advanced risk-aware hardening scenarios. Scenario-specific rubrics support deterministic static, unit, integration, and end-to-end checks, supplemented by LLM-assisted assessment for semantic requirements. We evaluate six coding-agent models and report rubric pass rate (RPR). Under the with-skill condition, mean RPR ranges from 68.58% to 91.37%. Access to the alipay-payment-integration skill improves mean RPR by 10.31 percentage points on average relative to the without-skill condition, with gains varying across models, products, and scenarios. Method-level results distinguish source-level completion, executable payment behavior, and payment-domain requirements. Alipay-PIBench provides a controlled setting for diagnosing model capability and evaluating structured guidance in payment integration.

13:00 JSTLLM/生成AI

類似性はどこまでも: LLM における多言語一般化は言語レベルの類似性構造に依存する

大規模言語モデル (LLM) はさまざまなタスクにわたって能力が向上していますが、その一般化 (無) 能力を定量化することは依然として難しく、限られた領域を超えて理解されることはほとんどありません。特に、LLM は英語以外の言語への多言語の一般化に苦労することが知られていますが、トレーニング データではそれが十分に証明されていません。その理由を理解するために、また一部のモデルが他のモデルよりも優れたパフォーマンスを実現できる理由を理解するために、認知科学全体にわたる研究の長い歴史に目を向け、一般化の成功は類似性空間での適切な表現から得られると主張します。私たちは、LLM の表現が異なる言語間の階層的類似構造をどの程度うまく捉えているかを調べます。驚くべきことに、LLMの潜在表現はインド・ヨーロッパ語族の階層構造をほぼ復元しており、同じサブファミリーのメンバーである言語を表現空間内で密接にグループ化していることを示した。さらに、モデルが言語の類似構造を反映する度合いが、多言語自然言語推論ベンチマークである XNLI でのパフォーマンスと相関していることを示します。これは、類似性に基づく一般化に関する古典的な研究を大規模に拡張し、類似した言語を表すモデルが、ある言語から別の言語へどのように同様により適切に一般化するかを示します。

原文 (English)

Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures

As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data. To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs' representations capture the hierarchical similarity structure between distinct languages. Strikingly, we show LLMs' latent representations largely recover the hierarchical structure of the Indo-European language family tree -- grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

Do LLMs Know Their Vulnerable Scenarios?

Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can by…

13:00 JSTエージェント

AgenticCANN: 知識拡張された Agentic Evolution による Ascend C オペレーターの自動生成

Ascend C のオペレーターの最適化は、NPU (Neural Processing Unit) の推論パフォーマンスにとって重要ですが、ハードウェアに関する深い専門知識が必要です。大規模言語モデル (LLM) は自動 CUDA カーネル生成において有望であることが示されていますが、Ascend C の根本的に異なるプログラミング モデルにより、未解明なままの固有の課題が生じます。この論文では、低コーパス NPU 環境での Ascend C オペレーター合成の自動化に特化した、知識拡張型エージェント進化フレームワークである AgenticCANN を提案します。不慣れなハードウェアでの深刻なプラットフォーム知識不足を克服するために、AgenticCANN には、上流の実現可能性のボトルネックを解決するために、開発ライフサイクル全体にわたって構造化されたマルチレベルのドメイン洞察を提供する知識統合生成システムが組み込まれています。この基盤に基づいて、動的に実行する段階適応型エージェント進化戦略を特徴としています。 LLM インタラクション モードを特定の生成フェーズと進化フェーズに合わせて調整し、高度な探索候補の発見と高度な収束パフォーマンス調整のバランスをとります。5 つのパターン カテゴリにわたる 6 つの演算子にわたる Huawei Ascend 910B での広範な実験により、私たちの手法が要素ごとの演算子と正規化演算子で 90 ~ 100%、融合演算子で 56% の実現可能性を達成し、1B Pangu モデル推論で最大 6.65 倍の高速化が達成されることが実証されました。カーネル。さらなる分析により、知識注入は要素ごとの演算子で実現可能性を 57% から 86% に単調に向上させることが明らかになり、演算子固有の利点ではなく一般的な利点が実証されました。

原文 (English)

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise. While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally different programming model of Ascend C introduces unique challenges that remain unexplored. In this paper, we propose AgenticCANN, a knowledge-augmented agentic evolution framework specifically tailored for automated Ascend C operator synthesis in low-corpus NPU environments. To overcome the severe platform knowledge deficit on unfamiliar hardware, AgenticCANN incorporates a knowledge-orchestrated generation system that delivers structured, multi-level domain insights across the development lifecycle to resolve the upstream feasibility bottleneck. Building on this foundation, it features a stage-adaptive agentic evolution strategy that dynamically aligns LLM interaction modes with specific generation and evolution phases, balancing high-exploration candidate discovery with high-convergence performance tuning. Extensive experiments on Huawei Ascend 910B across six operators spanning five pattern categories demonstrate that our method achieves 90 to 100 percent feasibility on elementwise and normalization operators, 56% on fusion operators, and up to 6.65$\times$ speedup on 1B Pangu model inference kernels. Further analysis reveals that knowledge injection monotonically improves feasibility from 57% to 86% on elementwise operators, demonstrating its general rather than operator-specific benefit.

13:00 JST研究/論文

DAPD: Dual-Anchored Policy Distillation

On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged i…

13:00 JST研究/論文

DiffImaginE: Diffusio を使用してエンティティ タイプを検証することを想像してください

マルチモーダル名前付きエンティティ認識 (MNER) は、各候補スパンとエンティティ タイプの仮説が共同のテキスト証拠と視覚的証拠によってサポートされているかどうかを判断します。既存の想像比較検証器は、各 (スパン、タイプ) ペアを 1 つの予測された視覚的特徴にマッピングし、多様な視覚的実現を単一のプロトタイプに圧縮し、明示的な確率的セマンティクスを使用せずに互換性スコアを提供します。 MNER 型検証を条件付き潜在拡散推論として定式化する DiffImaginE を紹介します。スパン局所化された視覚的証拠が与えられると、タイプ条件付きデノイザーは、標準化された潜在に注入されるノイズを予測します。結果として生じるノイズ除去誤差は、タイプ条件付き負の対数尤度の ELBO 一貫性のある代用値を提供し、競合するタイプの仮説を、観察をどの程度うまく説明できるかによってランク付けできるようにします。 DiffImaginE は、標準のマルチモーダル エンコーダ スタックを保持し、決定論的検証器を、Min-SNR 重み付けを使用してトレーニングされた分類子なしのガイド付き拡散スコアラーに置き換えます。タイプごとの拡散スコアを分類ロジットとして直接監視し、ノイズ レベル全体の集計を学習し、逆サンプリングを使用してモンテカルロ比較の分散を削減します。私たちの分析は、分類器を使用しないガイダンスが誘導型事後分布を鮮明にし、反対のペアリングが等しいデノイザーコストで分散を低減するときの特徴を示すことを示しています。 Twitter-2015 と Twitter-2017 の実験では、アブレーションと一対の有意性検定によってサポートされ、同じエンコーダー、補助対物レンズ、評価プロトコルの下で、一致した決定論的 ImaginE 制御に対して一貫したゲインが示されています。

原文 (English)

DiffImaginE: Imagine to Verify Entity Types with Diffusion

Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. Given span-localised visual evidence, a type-conditioned denoiser predicts noise injected into its standardised latent. The resulting denoising error provides an ELBO-consistent surrogate for type-conditional negative log-likelihood, allowing competing type hypotheses to be ranked by how well they explain the observation. DiffImaginE retains a standard multimodal encoder stack and replaces the deterministic verifier with a classifier-free-guided diffusion scorer trained using Min-SNR weighting. We directly supervise per-type diffusion scores as classification logits, learn aggregation across noise levels, and use antithetic sampling to reduce Monte Carlo comparison variance. Our analysis shows that classifier-free guidance sharpens the induced type posterior and characterises when antithetic pairing reduces variance at equal denoiser cost. Experiments on Twitter-2015 and Twitter-2017 show consistent gains over a matched deterministic ImaginE control under the same encoder, auxiliary objectives, and evaluation protocol, supported by ablations and paired significance tests.

13:00 JST研究/論文

Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study

Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes…

13:00 JST規制/政策

Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load

Short-term load forecasting (STLF) plays a vital role in the electric power industry. It is relevant for critical infrastructure. STLF is n…

13:00 JSTエージェントDeepSeek

長期にわたるターミナルタスクのための再帰的合成

ターミナル エージェント向けの高品質で長期的なトレーニング データは作成に費用がかかり、タスクごとに数百ドルから数千ドルかかることがよくあります。これは、各タスクが命令、環境、参照ソリューション、検証器の相互一貫性を保つ必要があるためです。人間によるオーサリングは拡張性がなく、大規模言語モデル (LLM) を使用した直接生成では、これらの依存関係が壊れることがよくあります。我々は、長期にわたるターミナル エージェント タスクを大規模に構築するための再帰的検証済み合成フレームワークである再帰的合成ターミナル タスク (RST) を紹介します。 RST は、検証されたシード タスクから開始して、参照ソリューションを拡張し、検証ツールと命令を新しいワークフローに再調整し、新しいサンドボックスで結果を検証し、受け入れられたタスクを後続のラウンドのシードとして再利用します。 15 回の再帰ラウンドにわたって、RST は 37,484 個の合成ターミナル エージェント タスクをタスクあたり約 $0.05 で生成します。タスクの難易度はラウンドを重ねるごとに大幅に増加します。リファレンス ソリューションの中央値は 67 行から 374 行に増加し、実行されたコマンド数の中央値は 40 から 244 に増加し、DeepSeek-V4-Pro pass@4 は $R_1$ の 90\% から $R_{15}$ の 2.5\% に低下します。トレーニングの有用性を実証するために、合成されたタスクに関して拒否サンプリングされた Qwen3.5 軌跡を収集し、それらを教師付き微調整に使用します。これらの軌道を微調整すると、Qwen3.5-27B と Qwen3.5-122B-A10B がターミナル ベンチ ~ 2、ターミナル ベンチ ハード、およびロングホライズン ターミナル ベンチで最大 10 ポイント改善され、エージェント PPO により Qwen3.5-27B が 3 つで 49.44\%、32.00\%、22.07\% に上昇しました。ベンチマークは、ベース モデルに対して 20.0\%、41.2\%、および 21.9\% の相対的な向上に相当します。さらに、15 ラウンド後の再帰には上限がありません。難易度が上昇し続けても合成収率と検証率は安定しており、ここで報告した規模をはるかに超えてプロセスを継続できることを示しています。

原文 (English)

Recursive Synthesis for Long-Horizon Terminal Tasks

High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. Human authoring does not scale, and direct generation with large language models (LLMs) often breaks these dependencies. We present Recursive Synthetic Terminal Tasks (RST), a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks at scale. Starting from verified seed tasks, RST extends the reference solution, realigns the verifier and instruction to the new workflow, validates the result in a fresh sandbox, and reuses accepted tasks as seeds for subsequent rounds. Across fifteen recursive rounds, RST produces 37,484 synthesized terminal-agent tasks at roughly $0.05 per task. Task difficulty increases substantially over rounds: the median reference solution grows from 67 to 374 lines, the median number of executed commands grows from 40 to 244, and DeepSeek-V4-Pro pass@4 drops from 90% at R1 to 2.5% at R15. To demonstrate training utility, we collect rejection-sampled Qwen3.5 trajectories on the synthesized tasks and use them for supervised fine-tuning. Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench 2, Terminal-Bench Hard, and Long-Horizon Terminal Bench, while agentic PPO lifts Qwen3.5-27B to 49.44%, 32.00%, and 22.07% on the three benchmarks, corresponding to relative gains of 20.0%, 41.2%, and 21.9% over the base model. Moreover, after 15 rounds, the recursion shows no ceiling: synthesis yield and validation rates remain stable as difficulty keeps climbing, indicating that the process can continue well beyond the scale reported here.

13:00 JSTエージェント

iARCS: 制御可能な 3D シーン生成のための反復エージェント RL

合成 3D シーンの生成は、コンピューター ビジョンや具体化された AI のデータ ソースとしてますます使用されていますが、既存のジェネレーターは、タスクに不可欠な機能上の制約を確実に満たさずに、知覚的なリアリズムを最適化することがよくあります。この不一致により、アクセシビリティ、トラバーサビリティ、空間ルールへの準拠がしばしば重要となる下流トレーニングでの合成データの有用性が制限されます。我々は、事前訓練されたシーンジェネレーターを自然言語タスクの要件に適応させる反復エージェント強化学習フレームワークである iARCS を紹介します。 iARCS は 2 段階の戦略を使用します。つまり、物理的な妥当性とレイアウトの品質を向上させるための普遍的な報酬の事前トレーニングと、それに続く、トレーニング フィードバックから繰り返し改良される LLM 生成の報酬プログラムによるタスク固有の微調整です。実験では、歩きやすさ、到達しやすさ、クリアランスを重視したタスクにおける制約の忠実度の向上、タスク固有の制約の効果的な最適化、および競技シーンの多様性が示されています。さらに、iARCS によって生成されたデータがベース ジェネレーターを改善し、制御可能なシーン編集方法だけでなく実用的な合成データ生成ツールとしてのその価値をサポートすることを示します。

原文 (English)

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements. iARCS uses a two-stage strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific fine-tuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPT

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general p…

13:00 JSTLLM/生成AIエージェント

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, c…

13:00 JSTエージェント研究/論文

AI 時代の栄養データ インフラストラクチャ: エージェント仲介研究のための FAIR の運用化

AI エージェントは栄養学の研究を加速できますが、その分析はアイデンティティ、セマンティクスを継承し、基礎となるデータの曖昧さを解放します。私たちは、自動化された使用のために FAIR を運用するソース保存インフラストラクチャである Nutrition Data Service (NDS) を紹介します。記述解決により、リリース固有のレコードを検索できるようになります。タイプ付き横断歩道は、独立してリリースされたリソースを接続します。機械可読インターフェイスは、バージョン管理されたソースとクロスウォークを公開し、AI エージェントによる分析を再生可能および監査可能にします。食品説明ベンチマークでは、NDS は高い精度を維持し、NutriBench で公開されている最高の言語モデル結果を上回りました。外部のブラインド クロスウォーク評価では、その型付きコントラクトが防御可能なリンクを優先し、サポートされていないマッピングを拒否することが示されています。個人レベルの血糖指数分析では、ピン留めされた NDS 入力はモデル間および反復実行間で同一の出力を生成しますが、オープンウェブ再構成は不安定なままです。中心的な結果は、エージェント媒介栄養研究には、データの識別、検索、およびクロスウォークのための新しいデータ インフラストラクチャが必要であるということです。

原文 (English)

Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research

AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying data. We present Nutrition Data Service (NDS), source-preserving infrastructure that operationalizes FAIR for automated use: description resolution makes release-specific records findable; typed crosswalks connect independently released resources; machine-readable interfaces expose versioned sources and crosswalks, supporting replayable and auditable analyses. On food-description benchmarks, NDS outperforms the best published language-model result on NutriBench. External and blinded crosswalk evaluations show that its typed contract favors defensible links and rejects unsupported mappings. In a person-level glycemic-index analysis, pinned NDS inputs produce identical outputs across models and repeated runs, while open-web reconstruction remains unstable. Together, these results show that agent-mediated nutrition research requires a new infrastructure that makes data identity, search, and crosswalk policy explicit.

13:00 JST研究/論文

経験則: 部分情報を使用して人工知能システムを説明する

説明可能な人工知能 (XAI) は、人工知能 (AI) システムが特定の決定にどのように到達したかを説明しようとします。私たちは、特定のデータポイントに対する AI システムの動作を予測するために最も関連する特徴を特定する新しい定式化に基づく XAI への新しいアプローチである「経験則」(RoT) 説明を提案します。 RoT が、(a) 大規模言語モデル (LLM) を使用したゼロショット分類、(b) モデルへのアクセスなしの不透明な AI システムの監査、(c) 科学的発見における AI の使用において、XAI を実現するのにどのように適しているかを示します。さらに、RoT は主要な AI 規制の特定の要件を満たし、XAI 実践者に使い慣れたインターフェイスと視覚化を提供し、モデルに依存せず、代替手段よりも大幅に高速です。コードは次から入手できます: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

原文 (English)

Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information

Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint. We show how RoT is well-suited to enable XAI in: (a) zero-shot classification using large language models (LLMs), (b) auditing of opaque AI systems without model access, and (c) the use of AI in scientific discovery. Additionally, RoT meets specific requirements from leading AI regulations, provides a familiar interface and visualisations for XAI practitioners, is model-agnostic, and is substantially faster than alternatives. Code available at: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

13:00 JSTエージェント研究/論文

グロタンディーク定数のための長期的な AI 研究: 人間と AI の数学的コラボレーションにおけるケーススタディ

AI エージェントは数学研究でますます使用されていますが、効果的な使用方法が不明瞭なことがよくあります。これに向けて、組み合わせ問題とその連続緩和の間の硬さを捉えるグロタンディーク定数 $K_G$ の境界を改善するために AI がどのように使用されたかに関する広範なケース スタディを紹介します。具体的には、$K_G$ の正確な値は不明ですが、最近、最もよく知られている範囲を \[ \frac{6\pi}{11} \;\le\; に厳しくしました。 K_G \;\le\; \frac{\pi}{2\log(1+\sqrt2)} - 10^{-4}。 \] 重要なのは、これらの改善は、分野の専門家によって新規とみなされる洞察に到達できる AI 研究システムを使用して達成されたことです。数学の研究に AI を使用した経験について、特にその長所と短所について詳しく説明します。また、AI が画期的な洞察に到達するための理想的な条件を作り出す経験についても説明します。

原文 (English)

Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration

AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their continuous relaxations. Specifically, while the precise value of $K_G$ is not known, we recently tightened the best known bounds to \[ \frac{6\pi}{11} \;\le\; K_G \;\le\; \frac{\pi}{2\log(1+\sqrt2)} - 10^{-4}. \] Crucially, these improvements were achieved using an AI research system that could arrive at insights deemed novel by domain experts. We give a detailed discussion of our experience using AI for mathematics research, particularly touching upon its strengths and weaknesses, as well as our experience with creating ideal conditions for AI to arrive at breakthrough insights.

13:00 JSTLLM/生成AI画像/動画生成エージェント研究/論文GPT / ChatGPT

MBA: 現実世界のビジネスアイデアのためのマルチモーダルベンチマークとエージェント

大規模言語モデル (LLM) を活用したエージェント システムは、ビジネスのアイデア創出に新たな機会をもたらしました。しかし、実世界のコンテキストには本質的にマルチモーダルな性質があるにもかかわらず、既存のアプローチは依然としてテキストのみのパラダイムに限定されています。そこで、ビジネスアイデアエージェントのトレーニングと評価のための初のマルチモーダルベンチマークである MBA-Bench を紹介します。これは 6 つのドメインにわたる 30,000 のサンプルで構成され、各ドメインはテキストだけでは完全には伝えられない明確な視覚的手がかりによって特徴付けられます。具体的には、画像に自動的にキャプションを付け、GPT-4o を使用して、検索クエリの生成、市場証拠の検索、および証拠強化合成を通じて 3 つのビジネス質問のそれぞれに対して 5 つの参考アイデアを生成します。以前の作業に続き、MLLM-as-a-Judge を使用して 6 つのビジネス指向の基準にわたってエージェントを評価します。基準が非表示または開示される設定を検討するために、ブラインドおよび既知のそれぞれに対して MBA-b および MBA-k を提示します。私たちは両方を 2 つの新しい報酬目標 (創造性と実現可能性) に基づいてトレーニングしますが、MBA-k は公開されている 6 つの基準を合計 8 つさらに最適化します。どちらも、LoRA ベースの監視付き微調整を介してトレーニングされ、その後、これらの設定固有の報酬を使用してグループ相対ポリシーの最適化が行われます。 MBA ベンチでの大規模な実験のために、キャプションのみまたはマルチモーダル入力のいずれかに対応する 2 つのベースラインを設定しました。後者はいくつかの指標でクローズド ソースのパフォーマンスに近づきます。 MBA-b と MBA-k は、キャプション ベースラインをそれぞれ 63.9% と 77.1% 上回り、マルチモーダル ベースラインを 25.6% と 35.8% 上回っています。

原文 (English)

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.

13:00 JST研究/論文

OEIS Open: 言語モデルはどれだけの推測を定理に変えることができますか?

私たちは、OEIS からの 492 のオープンな数学的予想に基づくベンチマークである OEIS Open を構築します。このベンチマークは、Tsoukalas らによって Lean で形式化されています。これらの推測は、これまで特注のエージェントを使用してのみ試みられていましたが、当社のオープンソース評価コードは、それらに対してあらゆる汎用言語モデル (LM) を実行し、LM 不正行為の試みに対して安全です。最小限のツールセットを装備した LM は、1 回の試行あたり $50 の予算でこれらの推測のうち 147 件を解決し、OEIS Open で 30% のスコアを獲得したことがわかりました。 OEIS Open Lite は、より安価な評価を目的とした 100 個の推測のランダムなサブセットです。試行ごとに $200 の予算で評価すると、現時点で最高の LM は OEIS Open Lite で 44% のスコアを獲得します。 LM が arXiv の 476,000 件の論文を介して数学文献にアクセスできるようにしても、OEIS Open Lite のパフォーマンスは向上せず、より洗練されたエージェント ループを使用しても向上しませんでした。この研究で取り上げられている推測は数学的重要性が不確かであり、ほとんどの推測はこれまでほとんど注目されていなかったと思われます。それにもかかわらず、私たちの結果は、LMが未解決の研究の推測を自律的にかつ適度なコストで解決できることを示しています。

原文 (English)

OEIS Open: How many conjectures can language models turn into theorems?

We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source evaluation code runs any generic language model (LM) against them, and is secure against LM cheating attempts. We find that LMs equipped with a minimal set of tools resolve 147 of these conjectures with a budget of \$50 per attempt, scoring 30% on OEIS Open. OEIS Open Lite is a random subset of 100 conjectures for cheaper evaluation. When evaluated with a budget of \$200 per attempt, the best current LM scores 44% on OEIS Open Lite. Giving LMs access to the mathematics literature via 476,000 papers from arXiv did not increase performance on OEIS Open Lite, and nor did using more sophisticated agent loops. The conjectures covered in this work are of uncertain mathematical significance, and most have likely received little previous attention. Nevertheless, our results show that LMs can resolve open research conjectures autonomously and at modest cost.

13:00 JSTLLM/生成AICopilot

The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot

Generative artificial intelligence (AI) facilitates content production and enhances ideation, with potentially important implications for d…

13:00 JSTLLM/生成AI

Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries

To evaluate a multi-representational framework in which large language model (LLM)-generated expert summaries of intensive care unit (ICU)…

13:00 JST研究/論文

Cueless EEG imagined speech for subject identification: dataset and benchmarks

Electroencephalogram (EEG) signals have emerged as a promising modality for biometric identification. While previous studies have explored…

13:00 JSTLLM/生成AIエージェントハードウェア/半導体

Unmasking Conversational Bias in AI Multiagent Systems

Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application…

13:00 JST研究/論文

はい、Q ラーニングはオフラインのインコンテキスト RL に役立ちます

既存のオフライン インコンテキスト強化学習 (ICRL) 手法は、主に教師ありトレーニング目標に依存していましたが、オフライン RL 設定には制限があることが知られています。この研究では、オフライン ICRL フレームワーク内での RL 目標の統合を検討します。 150 を超える GridWorld および MuJoCo 環境由来のデータセットでの実験を通じて、RL 目標を最適化すると、さまざまなデータセット カバレッジ、構造、専門知識レベル、環境の複雑さにわたって、広く採用されているアルゴリズム蒸留 (AD) と比較してパフォーマンスが平均約 30% 向上することが実証されました。さらに、困難な XLand-MiniGrid 環境では、RL 目標により AD のパフォーマンスが 2 倍になりました。私たちの結果は、価値学習中に保守主義を追加すると、テストしたほぼすべての設定でさらなる改善がもたらされることも明らかにしました。私たちの調査結果は、ICRL の学習目標と RL の報酬最大化の目標を一致させることの重要性を強調し、オフライン RL が ICRL を前進させるための有望な方向性であることを示しています。

原文 (English)

Yes, Q-learning Helps Offline In-Context RL

Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings. In this study, we explore the integration of RL objectives within an offline ICRL framework. Through experiments on more than 150 GridWorld and MuJoCo environment-derived datasets, we demonstrate that optimizing RL objectives directly improves performance by approximately 30% on average compared to widely adopted Algorithm Distillation (AD), across various dataset coverages, structures, expertise levels, and environmental complexities. Furthermore, in the challenging XLand-MiniGrid environment, RL objectives doubled the performance of AD. Our results also reveal that the addition of conservatism during value learning brings additional improvements in almost all settings tested. Our findings emphasize the importance of aligning ICRL learning objectives with the RL reward-maximization goal, and demonstrate that offline RL is a promising direction for advancing ICRL.

13:00 JST画像/動画生成

Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision

Efficiently adapting large pretrained models is critical under tight compute and memory budgets. While Parameter-Efficient Fine-Tuning (PEF…

13:00 JSTLLM/生成AIビジネス/資金調達

How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG

By retrieving contexts from knowledge graphs, graph-based retrieval-augmented generation (GraphRAG) enhances large language models (LLMs) t…

13:00 JST画像/動画生成研究/論文

Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights

Vision Language Models (VLMs) have shown promise in automating image diagnosis and interpretation in clinical settings. However, developing…

13:00 JST研究/論文Llama

Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models

Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay…

13:00 JSTLLM/生成AI

REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models

Large language models (LLMs) often express verbal confidence that is poorly aligned with actual correctness, limiting their reliability in…

13:00 JSTLLM/生成AI

Gradual Code-Switching as Inference-Time Cross-Lingual Representational Alignment for LLMs

While large language models (LLMs) have achieved notable progress in multilingual settings, their performance remains uneven across languag…

13:00 JST研究/論文

StarEmbed: Benchmarking Time Series Foundation Models on Astronomical Observations of Variable Stars

Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling.…

13:00 JST研究/論文

DiffGRM: Diffusion-based Generative Recommendation Model

Generative recommendation (GR) is an emerging paradigm that represents each item via a tokenizer as an n-digit semantic ID (SID) and predic…

13:00 JSTLLM/生成AI画像/動画生成

CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning

Urban socio-economic sensing plays a vital role in advancing global sustainable development goals. With the advent of Large Vision-Language…

13:00 JST画像/動画生成ロボティクスハードウェア/半導体

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gain…

13:00 JST研究/論文

Automated Design Optimization via Strategic Search with Large Language Models

Optimization methods have long advanced many fields, yet they struggle when faced with design problems where the search space and design pa…

13:00 JST研究/論文ClaudeGPT / ChatGPTGeminiLlama

Security and Detectability Analysis of Unicode Text Watermarking Methods against Large Language Models

Securing digital text is becoming increasingly relevant due to the widespread use of large language models. Individuals' fear of losing con…

13:00 JST画像/動画生成ロボティクス

RadarGen: Automotive Radar Point Cloud Generation from Cameras

We present RadarGen, a diffusion model for synthesizing realistic automotive radar point clouds from multi-view camera imagery. RadarGen ad…

13:00 JSTLLM/生成AIエージェント

Learning Latency-Aware Orchestration for Multi-Agent Systems

Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high i…

13:00 JST研究/論文

Architecture Before the Formula: Individuating Neural Architecture Beyond the Composite Map

Neural architecture is often identified by module syntax, computation graphs, or the composite functions they realize. These descriptions a…

13:00 JSTエージェントロボティクス

Safe Exploration via Policy Priors

Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. simulated)…

13:00 JSTLLM/生成AI

MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models

Large Language Models (LLMs) are increasingly deployed in sensitive applications including psychological support, healthcare, and high-stak…

13:00 JSTLLM/生成AI研究/論文

CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language

Large Language Models excel in high-resource programming languages but struggle with low-resource ones. Existing research related to low-re…

13:00 JST研究/論文

Automatic Termination Strategy of Inelastic Neutron-scattering Measurement Using Bayesian Optimization for Bin-width Selection

Currently, an excessive amount of event data is being obtained in four-dimensional inelastic neutron-scattering experiments. A method for a…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI

We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions…

13:00 JST研究/論文

In-context superposition: human-like working memory interference in large language models

Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments. This capacity, known as…

13:00 JST研究/論文

A Q-learning-based QoS-aware multipath routing protocol in IoMT-based wireless body area network

The Internet of Medical Things (IoMT) enables intelligent healthcare services but faces challenges such as dynamic topology, energy constra…

13:00 JST研究/論文

IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development

Adoption of AI-assisted development in 2025 exposed a tool-agnostic failure pattern: experienced developers using frontier models were meas…

13:00 JSTLLM/生成AI

Cat-DPO: Category-Adaptive Safety Alignment

Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and rel…

13:00 JST画像/動画生成

Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference

Expressway video anomaly detection is important for traffic safety, but remains challenging across diverse scenes, particularly for far-fie…

13:00 JST研究/論文

SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models

We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is…

13:00 JST研究/論文

Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance

Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, e…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and…

13:00 JST研究/論文

INSHAPE: Instance-Level Shapelets for Interpretable Time-Series Classification

Discovering shapelets -- i.e., discriminative temporal patterns within time series -- has been widely studied to address the inherent compl…

13:00 JST研究/論文

多腕ベイジアン バンディットのアニーリングされたソフトマックスの貪欲さ

検証可能な報酬を伴う強化学習 (RLVR) および GRPO などのグループベースのポリシー最適化手法は、プロンプトごとに複数の完了をサンプリングし、参照ポリシーに対する KL ペナルティによって正規化された、より高い報酬を持つポリシーの確率を高めることにより、確率的ポリシーを更新します。これらの更新には、認識論的不確実性を追跡する明示的なメカニズムは含まれていません。この論文では、なぜそのような不確実性を問わない更新が効果的であるのかについて、定型化された説明を研究します。多腕ベイジアン ベルヌーイ バンディットにおける経験的平均報酬のソフトマックスに従ってアクションを選択するアニーリングされたソフトマックス (ボルツマン) ポリシーを分析します。最適に近いアームが豊富にあることを意味する、事前の線形アッパーテール条件 ($\beta$-規則性の $\beta=1$ の場合) では、アニーリングされたソフトマックス グリーディがベイズ リポート $\tilde{O}(m + T/m)$ を達成すること、特にアームの数が $m = にスケールされる場合 $\tilde{O}(\sqrt{T})$ を達成することを証明します。 \シータ(\sqrt{T})$。これは、この体制における最適に近いベイズの後悔率であり、経験的平均の貪欲さによっても達成されます。 $\beta$-規則性の下では、多くのアームは学習を通じて最適値に近い経験的平均を維持するため、ソフトマックスが経験的に最良でないアームをサンプリングすると、そのアームは明らかに劣ったアームではなく、最適に近い別のアームになる傾向があります。対照的に、アームの数が少ない場合、同じ種類のソフトマックス ポリシーは直線的な後悔に見舞われる可能性があります。この結果は、RLVR と構造的に類似していることも示しています。ここでは、正しい完了を生成する無視できない確率を持つ基本ポリシーが $\beta$-規則性の役割を果たします。

原文 (English)

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling multiple completions per prompt and increasing the policy's probability on those with higher reward, regularized by a KL penalty toward a reference policy. These updates do not include explicit mechanisms that track epistemic uncertainty. This paper studies a stylized explanation for why such uncertainty-agnostic updates can nevertheless be effective. We analyze an annealed softmax (Boltzmann) policy that selects actions according to a softmax of empirical mean rewards in a many-armed Bayesian Bernoulli bandit. Under a linear upper-tail condition on the prior (the $\beta=1$ case of $\beta$-regularity), which implies an abundance of near-optimal arms, we prove that annealed softmax greedy achieves Bayes regret $\tilde{O}(m + T/m)$, and in particular $\tilde{O}(\sqrt{T})$ when the number of arms scales as $m = \Theta(\sqrt{T})$. This is the near-optimal Bayes regret rate in this regime, attained also by empirical-mean greedy. Under $\beta$-regularity, many arms maintain empirical means close to the optimum throughout learning, so when softmax samples an arm other than the empirically best, that arm tends to be another near-optimal one rather than a clearly inferior one. By contrast, with a small number of arms, the same kind of softmax policy can suffer linear regret. The result also provides a structural analogy to RLVR, where a base policy with a non-negligible probability of producing a correct completion plays the role of $\beta$-regularity.

13:00 JST研究/論文

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems

Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise prefer…

13:00 JST画像/動画生成ビジネス/資金調達

Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection

Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expens…

13:00 JST研究/論文

Constitutional On-Policy Safe Distillation

On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged informat…

13:00 JSTLLM/生成AI

トランスフォーマーには 3 つの投影が必要ですか? QKV バリアントの体系的な研究

トランスフォーマーは、クエリ、キー、値 (QKV) アテンションの定式化が中心的な役割を果たし、さまざまな AI タスクの標準ソリューションとなっています。しかし、これら 3 つの予測の個々の寄与と、一部を省略した場合の影響については、依然として十分に理解されていません。 3 つの射影共有制約を系統的に評価します。a) Q-K=V (共有キーと値)、b) Q=K-V (共有クエリキー)、c) Q=K=V (単一射影)。最後の 2 つのバリアントは、対称的なアテンション マップを生成します。これに対処するために、2D 位置エンコーディングによる非対称の注意も調査します。合成タスク、ビジョン (MNIST、CIFAR、TinyImageNet、異常)、言語モデリング (10B トークン上の 300M および 1.2B パラメーター モデル) にわたる実験を通じて、当社のトランスフォーマーは QKV トランスフォーマーと同等か、場合によってはそれよりも優れたパフォーマンスを発揮することがわかりました。言語モデリングでは、Q-K=V 射影共有により、わずか 3.1% のパープレキシティ低下で 50% の KV キャッシュ削減が達成されます。重要なのは、射影共有はヘッド共有 (GQA/MQA) を補完するものです。Q-K=V と GQA-4 を組み合わせると 87.5% のキャッシュ削減が得られ、Q-K=V + MQA では 96.9% が達成され、実用的なオンデバイス推論が可能になります。キーと値は同様の表現空間を占有することができ、注意は低ランク領域で動作するため、Q-K=V は品質を維持しますが、Q=K-V は注意の方向性を壊すことを示します。私たちの結果は、投影共有を、直接的で定量化可能な推論メモリの利点を備えた注意力の結びつきの未解明な例として体系的に特徴付けており、特にエッジ展開に価値があります。コードは https://github.com/anusamadan02/Do-Transformers-Need-3-Projections で公開されています。

原文 (English)

Do Transformers Need Three Projections? Systematic Study of QKV Variants

Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting some remain poorly understood. We systematically evaluate three projection sharing constraints: a) Q-K=V (shared key-value), b) Q=K-V (shared query-key), and c) Q=K=V (single projection). The last two variants produce symmetric attention maps; to address this, we also explore asymmetric attention via 2D positional encodings. Through experiments spanning synthetic tasks, vision (MNIST, CIFAR, TinyImageNet, anomaly), and language modeling (300M and 1.2B parameter models on 10B tokens), we discovered that our transformers perform on par or occasionally better than the QKV transformer. In language modeling, Q-K=V projection sharing achieves 50% KV cache reduction with only 3.1% perplexity degradation. Crucially, projection sharing is complementary to head sharing (GQA/MQA): combining Q-K=V with GQA-4 yields 87.5% cache reduction, while Q-K=V + MQA achieves 96.9%, enabling practical on-device inference. We show that Q-K=V preserves quality because keys and values can occupy similar representational spaces and attention operates in a low-rank regime, whereas Q=K-V breaks attention directionality. Our results systematically characterize projection sharing as an underexplored instance of weight tying in attention, with direct, quantifiable inference memory benefits, particularly valuable for edge deployment. The code is publicly available at https://github.com/Brainchip-Inc/Do-Transformers-Need-3-Projections

13:00 JSTLLM/生成AIエージェント

Certifiable Semantic Agreement Among LLM Agents: What the Admissibility Instrument Decides

Can a committee of LLM agents reach agreement that is certifiable at the level of meaning, not only at the level of a label? We build a pro…

13:00 JST研究/論文

SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation

Low-Rank Adaptation (LoRA) enables efficient adaptation of large pretrained models to downstream tasks by parameterizing weight updates wit…

13:00 JSTLLM/生成AI画像/動画生成

VLM 内の偽装されたビジュアル コンテキストの隠れた進化

ビジュアル トークンは、生の外部シグナルとして大規模言語モデル (LLM) に入力されます。それらがどのように意味のある表現に変換され、言語空間と相互作用するかは、統合アーキテクチャに完全に依存します。ビジュアル トークンを入力シーケンス内のコンテキスト内プロンプトとして扱うか、LLM の中間層に直接挿入するかによって異なります。これらのアーキテクチャ上の選択が視覚情報にどのような影響を与えるか、また LLM と統合するための内部変換については、制御された比較と理解がまだ十分に行われていません。単一画像、複数画像、およびビデオのベンチマークにわたる同一のトレーニング条件下で、インコンテキストおよびレイヤーごとのインジェクション VLM 統合パラダイムを評価することで、公正な比較を提供します。そうすることで、視覚トークンが、言語構造を欠く生の表現である偽装された視覚コンテキストとして LLM に入力されるが、統合パラダイムに応じて徐々に再形成され、それぞれが視覚信号の根本的に異なる周波数特性を捕捉する、隠れた進化を明らかにします。私たちは、LLM 内のこの進化が、VLM がどのような視覚的特徴を効果的に利用できるか、視覚表現が言語空間とどのように連携するか、そして最終的にはさまざまなタスクにわたって各パラダイムがどのように実行されるかを決定することを示します。さらに、注意の割り当てだけでは不十分であり、パフォーマンスは各層の視覚的表現の品質によって左右されることを示します。

原文 (English)

The Hidden Evolution of Disguised Visual Context inside the VLM

Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language space depends entirely on the integration architecture. Whether by treating visual tokens as in-context prompts within the input sequence or injecting them directly into the LLM's intermediate layers. A controlled comparison and understanding of how these architectural choices affect visual information and its internal transformation to integrate with the LLM remains underexplored. We provide a fair comparison by evaluating in-context and layer-wise injection VLM integration paradigms under identical training conditions across single image, multi-image, and video benchmarks. In doing so, we uncover a hidden evolution where visual tokens enter the LLM as disguised visual context, raw representations lacking linguistic structure, but are progressively reshaped depending on the integration paradigm, each capturing fundamentally different frequency characteristics of the visual signal. We show that this evolution inside the LLM determines what visual features the VLM can utilize effectively, how visual representations align with the language space, and ultimately how each paradigm performs across different tasks. We further demonstrate that attention allocation alone is insufficient, and that performance is driven by the quality of visual representations at each layer.

13:00 JST研究/論文

Communication Heterogeneity and Collective Consensus in Neural Cellular Automata

Reaching global agreement from purely local interactions is a defining problem of collective intelligence, and most models of it assume tha…

13:00 JST画像/動画生成ロボティクス

Early Warning Signals for OpenVLA Failure under Visual Distribution Shift

Visual shifts can cause a vision-language-action policy to fail after initially plausible behavior. We ask whether OpenVLA's internal activ…

13:00 JSTLLM/生成AI

LLM-Based Test Oracles: Source-of-Authority Taxonomy -- A Systematic Literature Review

Large language models (LLMs) increasingly decide whether software behaves correctly, either by writing a test oracle or by acting as one. Y…

13:00 JST研究/論文

XAI 主導のデータ削減による時系列分類のスケーリング

時系列の Explainable AI (XAI) はアルゴリズム的に大幅な成長を遂げていますが、下流のタスクに測定可能なパフォーマンスの向上をもたらすというその有用性は依然として十分に検討されていません。このホワイトペーパーでは、時系列分類 (TSC) における効果的なデータ削減のために XAI アトリビューション手法を再利用する新しい方法論である drXAI を紹介することで、このギャップを埋めます。最新の TSC における中心的な課題はスケーラビリティです。 Transformers などの最先端のモデルは、シーケンスの長さに対して 2 次の複雑性を示し、チャネル数に対して 1 次の複雑さを示します。これにより、大規模なデータセットの計算が法外に難しくなります。 drXAI は、高速な GPU アクセラレーション分類器 (Hydra) を使用してローカル アトリビューションを生成することで、この問題に対処します。これらをグローバルな特徴重要度スコアに集約し、自動化されたエルボーカット ヒューリスティックを採用して、手動のしきい値を必要とせずに最も顕著な特徴を選択します。私たちは、合成データセットと現実世界の一変量データセットおよび多変量データセットの両方でアプローチを評価します。合成ベンチマークでは、drXAI は、従来のベースラインが失敗するグラウンドトゥルース機能を正常に回復します。実世界のデータでは、drXAI は、完全なデータセットでトレーニングされたモデルと同等の分類精度を維持しながら、80% ~ 90% のデータ削減を達成します。最も重要なことは、drXAI を使用すると、ConvTran のようなリソースを大量に消費するモデルを、メモリの制約により以前はアクセスできなかったデータセットに拡張できることを示しています。私たちの結果は、XAI を解釈しやすさだけでなく、時系列分析における特徴選択とスケーラビリティのための堅牢なツールとして使用する利点を示しています。すべてのコードとデータは公開されています。

原文 (English)

Scaling Time Series Classification via XAI-Driven Data Reduction

Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodology that repurposes XAI attribution methods for effective data reduction in Time Series Classification (TSC). The core challenge in modern TSC is scalability; state-of-the-art models, such as Transformers, exhibit quadratic complexity relative to sequence length and linear complexity relative to the number of channels. This renders them computationally prohibitive for massive datasets. drXAI addresses this by using a fast, GPU-accelerated classifier (Hydra) to generate local attributions. We aggregate these into global feature importance scores and employ an automated elbow-cut heuristic to select the most salient features without requiring manual thresholds. We evaluate our approach on both synthetic and real-world univariate and multivariate datasets. On synthetic benchmarks, drXAI successfully recovers ground-truth features where traditional baselines fail. On real-world data, drXAI achieves between 80% and 90% data reduction while maintaining classification accuracy comparable to models trained on the full dataset. Most importantly, we show that drXAI allows resource-intensive models like ConvTran to scale to datasets that were previously inaccessible due to memory constraints. Our results show the benefits of using XAI not just for interpretability, but as a robust tool for feature selection and scalability in time series analysis. All our code and data are openly available.

13:00 JST研究/論文

Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training. On a 6.78B-parameter MoE languag…

13:00 JST研究/論文

Vibe to Code: Elucidating Strategic Oscillation of Tacit Knowledge in Generative AI Design Workflows -- An Exploratory Qualitative Study

The rapid adoption of generative AI tools has created new literacy demands for designers who must verbalize tacit knowledge through natural…

13:00 JSTエージェントGPT / ChatGPT

Moral Hazard in Multi-Agent Language Models

Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr\"om's…

13:00 JST画像/動画生成

A Distributional Robustness Margin For Pathology Foundation Models

Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut lear…

13:00 JST研究/論文

SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups

Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Exis…

13:00 JSTLLM/生成AI

Commit Locally, Exit Globally: Coordinating Adaptive Sampling and Early Exit in Diffusion Language Models

Diffusion language models expose a provisional prediction at every denoising step, and on many tasks the candidate answer inside it stabili…

13:00 JST研究/論文

Coordinated incentives in AI-generated misinformation governance

With the rapid diffusion of AI-generated content, AI-driven misinformation is becoming increasingly pervasive and difficult to govern, unde…

13:00 JST研究/論文

Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks

Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parall…

13:00 JST研究/論文

Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction

Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other…

13:00 JSTLLM/生成AI研究/論文

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof…

13:00 JST研究/論文

Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online…

13:00 JST画像/動画生成

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's re…

13:00 JSTエージェント

Governing Agentic AI in FinTech

Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and…

13:00 JSTエージェント

AI Guardrail Survival under Single-Cycle Agentic Self-Summarization

Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dro…

13:00 JSTロボティクス

Keep the Future, Drop the Rollout: RIFT for World Action Models

World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask w…

13:00 JST研究/論文

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation

On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolat…

13:00 JST研究/論文

Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians

A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical…

13:00 JSTLLM/生成AI

Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction hist…