AIニュース 2026-08-16
自動生成: 2026-08-16 10:41 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
「Claude」の“見えない透かし”、Anthropicが仕組みを説明 「完全な書き直しなら消える」ITmedia AI+
Anthropicは、「Claude」が生成するテキストに埋め込む電子透かしの仕組みを公表した。乱数生成に秘密鍵を用いて統計的パターンを残…
-
Woman claims her stepfather used Grok to transform childhood photo into explicit imageryTechCrunch AI
The woman claimed that AI tools are "taking everyday life and turning…
-
SpaceX officially closes its Cursor acquisitionTechCrunch AI
AI coding startup Cursor is now officially a part of SpaceX.
-
AI 時代の栄養データ インフラストラクチャ: エージェント仲介研究のための FAIR の運用化arXiv cs.AI
AI エージェントは栄養学の研究を加速できますが、その分析はアイデンティティ、セマンティクスを継承し、基礎となるデータの曖昧さを解放します…
-
管理された永続メモリ: ロングホライズンエージェントのソースバインド状態セマンティクスとフェールクローズリリースarXiv cs.AI
エージェントの長期記憶は通常、select-store-retrieve として扱われますが、矛盾するレコード、置き換えられたレコード、撤…
-
推論陪審: 推論トレースを評価するためのマルチモデルのコンセンサスarXiv cs.AI
推論 LLM を改善するには、効果的な推論データのキュレーションのための長い推論トレースの品質を判断する能力、強化学習中の強力なトレーニン…
-
SteerBench-Work: アクション境界でのエージェントのステアリングのベンチマークarXiv cs.AI
長期間実行される LLM エージェントはツールを通じて動作し、単一のステップで電子メールの送信、プル リクエストのマージ、または電信送金を…
トピック別件数
- 研究/論文 135件
- LLM/生成AI 119件
- エージェント 73件
- 画像/動画生成 38件
- ビジネス/資金調達 26件
- ロボティクス 11件
- ハードウェア/半導体 10件
- 規制/政策 3件
- その他 1件
日本語メディア1件
ITmedia AI+ (日本語)
「Claude」の“見えない透かし”、Anthropicが仕組みを説明 「完全な書き直しなら消える」
Anthropicは、「Claude」が生成するテキストに埋め込む電子透かしの仕組みを公表した。乱数生成に秘密鍵を用いて統計的パターンを残す手法で、品質や速度、料金に影響を与えない。「EU AI Act」への準拠を目的とするが全世界で適用され、コードや短文などでは入りにくいとい…
海外メディア2件
TechCrunch AI (英語)
Woman claims her stepfather used Grok to transform childhood photo into explicit imagery
The woman claimed that AI tools are "taking everyday life and turning it into child sexual abuse."
SpaceX officially closes its Cursor acquisition
AI coding startup Cursor is now officially a part of SpaceX.
公式ブログ0件
このカテゴリの新着記事はありませんでした。
論文299件
arXiv cs.AI (英語)
立場: 推論は学習可能なルールベースのプロセスです
自律的推論は、今日の AI において最も科学的かつ経済的に動機付けられるトピックの 1 つです。歴史的にはシンボリック AI の範囲でしたが、最近の進歩は主に深層確率生成モデルから生まれました。多大な関心と急速な進歩にも関わらず、生成 AI コミュニティは推論の運用上の定義に明確に収束しておらず、論理や検証可能な自動推論におけるこのトピックの歴史的扱いを暗黙のうちに拒否することがよくあります。この立場は、定義の曖昧さにより推論評価の構成的妥当性が検証不可能なままとなり、信頼できる自律推論に向けた定量化可能な進歩が損なわれると主張します。私たちはまた、この曖昧さは対処可能であると主張します。そのために、私たちは、(1) 文献の統合に基づいた操作上の定義を提供し、有効かつ健全な推論を学習可能なルールベースのプロセスとして位置づけます。 (2) AI 推論研究のコミュニケーションにおけるベスト プラクティスのチェックリスト。
原文 (English)
Position: Reasoning is a Learnable Rule-Based Process
Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite immense interest and rapid progress, the generative AI community has not clearly converged on operational definitions for reasoning and often implicitly rejects the historical treatment of this topic in logic and verifiable automated reasoning. This position contends that definitional ambiguity leaves the construct validity of reasoning evaluation unverifiable, undermining quantifiable progress toward trustworthy autonomous reasoning. We also contend that this ambiguity is addressable. To that end, we provide (1) operational definitions based on a synthesis of the literature, positioning valid and sound reasoning as a learnable rule-based process; and (2) a checklist for best practices in the communication of AI reasoning research.
共同研究者としてのLLMの研究誠実性を評価するための診断財団
言語モデルは共同科学者として採用されることが増えていますが、組織的な圧力の下で研究の完全性を維持する能力は依然として測定されていません。 3 つのドメインと 4 つの研究段階にわたる 5 レベルの暗黙的および明示的圧力プロトコルの下で、36 のペアのタスクにわたって不正行為の分類、倫理的行動の推論、成果物に基づいた意思決定を評価するベンチマークである IntegrityBench を紹介します。 18 のフロンティア モデルのバリアントを評価したところ、ピーク時のプレッシャー下では、モデルは整合性が重要な決定のおよそ 3 件中 1 件で失敗し、規模も推論能力もこれを確実に軽減できないことがわかりました。明示的な圧力は不正行為の遵守を誘発しますが、暗黙的な文脈の再構築は正当な研究タスクの過剰な拒否を引き起こすことがよくあります。興味深いことに、研究リクエストを正確に分類できなかったモデルは、アーティファクトに基づいた意思決定に関して同等以上のパフォーマンスを発揮し(85.7対79.4)、3つの側面が構造的に分離しており、正しい倫理的行動には正確な分類が必要ないことを示唆しています。したがって、フロンティア モデルは有用であるように見えますが、研究不正行為の促進と AI 支援研究に対する信頼の低下という 2 つの異なる導入リスクを生み出す整合性の欠陥を抱えている可能性があります。
原文 (English)
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct classification, ethical action reasoning and artifact-grounded decision making across 36 paired tasks under a 5-level implicit-explicit pressure protocol spanning 3 domains and 4 research stages. Evaluating 18 frontier model variants, we find that under peak pressure, models fail roughly 1 in 3 integrity-critical decisions, and neither scale nor reasoning ability reliably mitigates this. Explicit pressures induce compliance with misconduct, while implicit contextual reframing more often causes over-refusal of legitimate research tasks. Interestingly, models failing to classify research requests accurately perform equally or better on artifact-grounded decision making (85.7 vs. 79.4), suggesting the three facets are structurally dissociated and correct ethical action does not require accurate classification. Frontier models can thus appear helpful while harbouring integrity failures that create two distinct deployment risks: facilitating research misconduct and eroding trust in AI-assisted research.
立場: アライメント コミュニティは意図せずに検閲ツールキットを構築している
この意見書では、現代の AI 調整手法は、本来は有害な出力を防ぐために設計されたものであり、悪意のある行為者によって検閲や操作のために簡単に悪用される可能性がある二重用途技術であると主張しています。現在の調整技術を悪用の可能性と実際のケースにマッピングすることで、「完全に調整された」モデルの探求が、意図せずして悪意のある攻撃者に情報支配のための絶えず改良されたツールを提供してしまうことを示します。ユーザーによる情報プロバイダーとしての AI の急速な導入、経済力の非対称性、権威主義への移行が進む政治情勢によってそのリスクが悪化しているため、私たちはこの二重利用の可能性について今すぐ議論する必要があります。私たちはコミュニティに対し、AI 調整メカニズムの意図的な誤用を考慮し、この二重使用の可能性を防ぐための緩和戦略を提案することを強く求めて締めくくります。
原文 (English)
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential.
合意は一致ではない:人間とLLMの倫理的判断における異なる道徳的根拠
人間の判断との一致は、大規模言語モデル (LLM) の整合性を評価するための一般的な手段です。しかし、最終的なラベルの一致は、ヒューマン・アノテーターとモデルが同じ道徳的根拠に依存していることを示しているわけではありません。 2 人のエージェントが、異なる原則、文脈上の仮定、または状況の解釈に訴えながら、同じ判断に達する可能性があります。私たちは、道徳的判断の 5 つの領域にわたる精選された 500 項目の ETHICS 由来のベンチマークを使用して、この区別をテストします。最終的なラベルと裏付けとなる論理的根拠の両方について、新しいヒューマン アノテーターと LLM アノテーションが付けられます。フロンティアおよびオープン モデル ファミリ全体で、ヒューマン アノテーターの多数派ラベルとの一致度は高いことがよくあります。しかし、理論的根拠レベルの分析では、ヒューマン・アノテーターとモデルによって表現される道徳的根拠の体系的な相違が明らかになります。特に、モデルは、最終的なラベルがヒューマン・アノテーターの多数派と一致する場合でも、危害、敬意、約束遵守、正義、砂漠、言い訳の関連性などのカテゴリー全体に注意を再配分します。私たちの結果は、一致を調整と同等のものとして扱うべきではないことを示しています。したがって、ラベルに基づく評価は、モデル判断で表現された理由、原則、道徳的優先順位の分析によって補完されない限り、誤解を招くほど安心させる可能性があります。
原文 (English)
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high. However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.
モバイル エッジ コンピューティングにおけるストリーム処理のための LLM 支援コントラクト ネット ネゴシエーションを使用したマルチエージェント スケジューリング
ストリーム処理システムは、異種モバイル エッジ、つまりクラウド インフラストラクチャ全体で動作することがますます増えており、ワークロードの不安定性、リソース競合、および厳しいサービス品質 (QoS) 要件により、分散スケジューリングが複雑になっています。この論文は、\emph{MAS-DecStream} を提案します。その主な貢献は \emph{LLM-MR-CNP} です。これは、セマンティック CFP 定式化、段階的なコンテキスト開示、複数ラウンドの提案修正、ネゴシエーション記憶、決定論的検証を備えた古典的なコントラクト ネット プロトコルの拡張です。エッジ クラスタ エージェントは、ハード リソースと QoS の制約が決定的なままである一方で、ローカルの観察、予測されたリソースの状態、定性的なランタイム コンテキストに基づいて自然言語オフロードの提案を改良します。 Alibaba ASI トレースから派生した実験では、シングルラウンドとマルチラウンドの CNP、ルールベースと LLM 支援の改良、および固定モデルのシングルラウンドとマルチラウンドのネゴシエーションの 3 つのレベルで拡張機能を評価します。評価された構成では、MAS-DecStream はレイテンシ違反を 3\% に削減し、リソースのオーバーコミットを排除し、20 エージェントで競合解決率 0.91 に達し、マルチラウンド ルールベースのベースラインと比較してユーティリティを最大 22\% 向上させます。別の 25 ケースの評価では、モデルおよびプロンプトに依存する精度とコストのトレードオフが示されています。この結果は、マルチラウンド CNP 改良が主要なプロトコル レベルの利点であり、LLM 支援により定性的で不確実なランタイム コンテキストに付加価値を与えるという最初の証拠を提供します。
原文 (English)
Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing
Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource contention, and stringent quality-of-service (QoS) requirements complicate decentralized scheduling. This paper proposes \emph{MAS-DecStream}, whose main contribution is \emph{LLM-MR-CNP}: an extension of the classical Contract Net Protocol with semantic CFP formulation, progressive context disclosure, multi-round proposal revision, negotiation memory, and deterministic validation. Edge-cluster agents refine natural-language offloading proposals from local observations, predicted resource states, and qualitative runtime context, while hard resource and QoS constraints remain deterministic. Experiments derived from the Alibaba ASI Trace evaluate the extension at three levels: single- versus multi-round CNP, rule-based versus LLM-assisted refinement, and fixed-model single- versus multi-round negotiation. Under the evaluated configurations, MAS-DecStream reduces latency violations to 3\%, eliminates resource overcommitment, reaches a conflict-resolution rate of 0.91 with 20 agents, and improves utility by up to 22\% over the multi-round rule-based baseline. A separate 25-case evaluation shows model- and prompt-dependent accuracy--cost trade-offs. The results provide initial evidence that multi-round CNP refinement is the principal protocol-level gain, with LLM assistance adding value for qualitative and uncertain runtime context.
立場: 人間の推論を反映する実用的な AI 調整方法が必要です
AI システムは、意思決定支援、意思決定の代理人、または自律的な意思決定者としてますます採用されています。この意見書では、多くの状況、特に一か八かの意思決定において、ユーザーと同じように推論し、ユーザーの推論を忠実に伝える、正確に認知的に調整された AI システムが必要であると主張しています。私たちは、認知的調整によって理解のしやすさと信頼性が向上するという証拠を検討し、AI の判断や行動の根拠が自分たちにとって重要である場合、多くのユーザーが認知的調整が「不可欠」であると感じていることを示す新しい調査データを提供します。私たちは、既存のアライメント手法と認知的アライメントを達成するために必要なものとの間のギャップを概説し、これらのギャップに対処するための研究課題を提示します。私たちは、認知的不整合は、想定されている多くのアプリケーションにおける AI 導入の障害となる可能性が高く、これに対処することが、ユーザーが喜んで信頼できる AI システムを構築するために重要であると主張します。
原文 (English)
Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data showing that many users find cognitive alignment "essential" when an AI's rationale for a judgment or action is important to them. We outline the gaps between existing alignment methods and what is needed to achieve cognitive alignment, and present a research agenda to address these gaps. We argue that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and justified to rely.
LLM に核攻撃を推奨したくないですか?日本語で聞いてみてください
大規模な言語モデルは戦略的および助言的な文脈で使用されることが増えていますが、その安全性の調整は通常英語のみで評価されます。私たちは 6 つのプロバイダーの 9 つのモデルをテストし、一か八かのシナリオにおいてプロンプトの言語がモデルの決定を変える可能性があるかどうかを尋ねます。私たちは、モデルが核武装国に無防備な敵を攻撃すべきかどうかアドバイスする、シングル ターン ゲーム理論のビネットを使用します。このプロンプトは意図的に非道徳的であり、戦略的には言語が異なっても同一です。日本語プロンプトはクロード モデル ファミリーの起動率を低下させることがわかりました。クロード ソネット 4.6 はストライキが不必要なシナリオでは 40% から 0% に低下し、競合シナリオでは 93% から 17% に低下しましたが、ストライキが戦略的に合理的である場合には最小限の効果しかありませんでした。この影響は Gemini Pro 3.1 (53% から 13%) まで及びます。言語をまたいだ実験により、このメカニズムが分離されました。英語のプロンプトで日本語で推論するように指示された場合、起動率は 93% から 37% に低下しました。効果を生み出すのは、入力言語ではなく、モデルが推論するよう要求される言語です。日本語で推論する場合、モデルはプロンプトにはまったく含まれていない道徳語彙 (「道徳的コスト」、「何百万もの命」) を自発的に生成します。他の 5 つのモデルには言語の影響は見られませんが、言語に関係なくほぼすべての状況で起動します。この効果には、すでに英語で躊躇するモデルが必要です。これらの結果は、LLM の安全動作は言語に依存しており、英語のみで評価すると、他の言語でエンコードされたリスクと安全対策の両方を見逃す可能性があることを示しています。
原文 (English)
Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stakes scenario. We use single-turn game-theoretic vignettes in which a model advises a nuclear-armed nation on whether to strike a defenseless opponent. The prompt is intentionally amoral and strategically identical across languages. We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is unnecessary and from 93% to 17% in contested scenarios, with minimal effect when the strike is strategically rational. The effect extends to Gemini Pro 3.1 (53% to 13%). A cross-language experiment isolates the mechanism: when instructed to reason in Japanese in an English prompt, launch rates drop from 93% to 37%. It is the language the model is asked to reason in, not the language of the input, that drives the effect. When reasoning in Japanese, models spontaneously generate moral vocabulary (''moral cost'', ''millions of lives'') that is entirely absent from the prompt. Five other models show no language effect, but they launch in nearly every condition regardless of language. The effect requires a model that already hesitates in English. These results show that LLM safety behavior is language-dependent, and that evaluating in English alone can miss both risks and safeguards encoded in other languages.
デュアルフロートランスフォーマー: プライマリプレフィルパスを追加のデコード計算から切り離す
大規模な言語モデルがより多くのリクエストに対応するにつれて、累積推論コストが 1 回限りのトレーニング コストと比較して重要になってきます。 2 つの推論フェーズでは、ハードウェアに異なる点で重点を置きます。プロンプト プレフィルは並列であり、通常はコンピューティングに依存しますが、自己回帰デコードはシーケンシャルであり、多くの場合メモリ帯域幅に依存します。従来の幅または深さのスケーリングでは、追加されたすべてのレイヤーが両方のフェーズで評価されるため、両方のコストが同時に増加します。プロンプト全体の主な計算と単一の永続的なキー値 (KV) キャッシュを保持しながら、追加の学習された計算を継続予測に代わりに割り当てることができるかどうかを尋ねます。デュアルフロートランスをご紹介します。その主なフローは、プロンプトを処理して KV キャッシュを書き込む完全な因果言語モデルです。補助フローはプロンプト処理中に省略され、最後のプロンプト位置以降のみアクティブ化され、永続的な状態を書き込んだり主フローに影響を与えたりすることなく継続予測の計算が追加されます。 2 つのフローは、メジャー アテンション、MLP、および出力行列を共有し、別個のトークン埋め込みと軽量結合を使用します。重みとプライマリ キャッシュを共有すると、グループ化された実行中にロードされた重みとキャッシュされたキーと値を再利用する機会も生まれます。デュアルフローは、一致したトークンの比較において、アーキテクチャおよびデータ構成全体で検証損失の低減を実現します。 MoE モデルでは、この分離により、主要エキスパートと補助エキスパートのファンアウトが即時コスト、継続コスト、予測品質に対して独立して制御されます。固定プレフィル エキスパート計算でのデコード計算を増やすことと、2 つのフロー間で固定デコード エキスパート バジェットを再割り当てすることの 2 つの体制を研究します。これらの実験は、プリフィルとデコードの品質のトレードオフを明らかにし、フェーズ固有の専門家割り当ての可能性を実証します。
原文 (English)
Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
As large language models serve more requests, cumulative inference cost is becoming increasingly important relative to one-time training cost. The two inference phases stress hardware differently: prompt prefill is parallel and typically compute-bound, whereas autoregressive decode is sequential and often memory-bandwidth-bound. Conventional width or depth scaling increases both costs together because every added layer is evaluated in both phases. We ask whether additional learned computation can instead be allocated to continuation prediction while preserving the prompt-wide primary computation and a single persistent key-value (KV) cache. We introduce the Dual-Flow Transformer. Its primary flow is a complete causal language model that processes the prompt and writes the KV cache. The auxiliary flow is omitted during prompt processing and activated only from the final prompt position onward, adding continuation-prediction computation without writing persistent state or influencing the primary flow. The two flows share major attention, MLP, and output matrices, while using separate token embeddings and lightweight coupling. Sharing weights and the primary cache also creates opportunities to reuse loaded weights and cached keys and values during grouped execution. Across matched-token comparisons, Dual-Flow achieves lower validation loss across architectures and data configurations. In MoE models, the separation makes primary and auxiliary expert fan-outs independent controls over prompt cost, continuation cost, and predictive quality. We study two regimes: increasing decode computation at fixed prefill expert computation, and reallocating a fixed decode expert budget between the two flows. These experiments expose a prefill-decode-quality trade-off and demonstrate the potential of phase-specific expert allocation.
LLM パーソナライゼーションのための Meta-LoRA を介したクロスドメイン設定の適応方法の学習
クロスドメインのゼロショットまたは数ショットのパーソナライゼーションは、ほんの一握りのターゲットドメイン間の対話から、目に見えない会話ドメインでユーザーが好む応答を生成することを目的としています。既存の適応手法は、証拠がまばらなため、更新の大きさを調整するのに苦労し、その結果過剰適合が発生します。一方、履歴転送手法は、多くの場合、ユーザーの好みとソースドメインのアーティファクトが絡み合い、信頼性の低いパーソナライゼーション事前分布と否定的な転送が生じます。証拠の品質への適応を調整するために、サポートセットのサイズと予測の不確実性に応じて更新強度を調整しながら、メタ学習された LoRA 初期化を適応開始と以前の中心の両方として使用する PAC-ベイズ正規化 Meta-LoRA を提案します。これにより、証拠がまばらまたは曖昧な場合の過剰適合が制限され、証拠が増えるにつれてより強力なパーソナライゼーションが可能になります。制御された適応だけでは、どの設定をドメイン間で転送する必要があるか、またはそれらをどのように表現する必要があるかを決定することはできません。したがって、私たちはパーソナライゼーション事前設定をユーザーとドメインのコンポーネントに機能的に分解し、安定した設定のための人が判読できるプロンプトと、ドメイン固有の隠し空間条件付けのためのトポロジを保持するソフト トークンを使用します。複数のベンチマークとパーソナライゼーション タスクにわたる実験では、強力なベースラインを超える一貫した向上が示されています。 HiCUPID では、私たちの方法は、競合する最良のベースラインと比較して、クロスドメインの勝率の低下を 47.9% 削減し、目に見えないユーザーのコールド スタートで勝率を 110.2% 改善します。
原文 (English)
Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization
Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain interactions. Existing adaptation methods struggle to calibrate update magnitude under sparse evidence and thus overfit, whereas history-transfer methods often entangle user preferences with source-domain artifacts, yielding unreliable personalization priors and negative transfer. To calibrate adaptation to evidence quality, we propose PAC-Bayes-regularized Meta-LoRA, which uses a meta-learned LoRA initialization as both the adaptation start and prior center, while adjusting update strength according to support-set size and predictive uncertainty. This limits overfitting under sparse or ambiguous evidence while permitting stronger personalization as evidence grows. Controlled adaptation alone does not determine which preferences should transfer across domains or how they should be expressed. We therefore functionally decompose personalization priors into user and domain components, using a human-readable prompt for stable preferences and topology-preserving soft tokens for domain-specific hidden-space conditioning. Experiments across multiple benchmarks and personalization tasks show consistent gains over strong baselines. On HiCUPID, our method reduces cross-domain win-rate degradation by 47.9% relative to the best competing baseline and improves win rate by 110.2% under unseen-user cold start.
研究アシスタント: アストラゼネカの研究開発用エージェントシステム
Research Assistant について説明します。Research Assistant は、科学者や臨床医が幅広いデータ ソースにわたって生物医学的な疑問を調査できるように、アストラゼネカで開発された LLM ベースの社内システムです。このシステムは、科学文献、ナレッジ グラフ、化学、臨床試験、安全リソース、発現データ、および内部実験システムからの証拠をまとめるチャット スタイルのインターフェイスを提供します。直接質問に答えるための高速モードと、より複雑な研究タスクのためのマルチステップ モードの両方をサポートします。回答は取得した証拠に基づいており、元のソースにリンクされているため、ユーザーは基礎となるデータを確認してさらに調査することができます。このテクニカル ノートでは、システム アーキテクチャ、製品の背後にある主な設計上の選択、およびアストラゼネカ全体の日常の研究開発ワークフローをサポートするために製品を大規模に導入することから得た教訓について概説します。
原文 (English)
Research Assistant: AstraZeneca's Agentic System for R&D
We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range of data sources. The system provides a chat-style interface that brings together evidence from scientific literature, knowledge graphs, chemistry, clinical trials, safety resources, expression data, and internal experimental systems. It supports both a fast mode for direct question answering and a multi-step mode for more complex research tasks. Responses are grounded in retrieved evidence and linked back to the original sources, allowing users to review and further explore the underlying data. In this technical note, we outline the system architecture, the main design choices behind the product, and lessons learned from deploying it at scale to support day-to-day R&D workflows across AstraZeneca.
大規模な言語モデルは命令に従うことはできるが、一度に多くは実行できない: 構成制約を満たす際の相転移
大規模な言語モデルは、推論構造、安全境界、出力スキーマなど、複数の明示的な制約を同時に遵守する必要がある設定に導入されることが増えています。個々の制約はうまく処理されていますが、多くの制約が共同で保持しなければならない構成体制は、依然として十分に特徴付けられていません。つまり、パフォーマンスはどのくらいの速さで低下するのか、何が低下を支配しているのか、そして崩壊は緩和できるのか?制約飽和評価 (CSE) は、同時制約 (k) の数を体系的に変化させる手続き的に生成されたベンチマークであり、すべての制約が決定論的なルールベースの検証者と LLM ジャッジの関与なしによってスコア付けされます: 15 のモデル、36 の制約タイプ、k=1 ~ 12 での 369,753 のチェック。 3 つの発見が得られます。まず、制約ごとの通過率は徐々に予測どおりに低下し、k 個の制約すべてを満たす確率は崩壊します。k=8 で個々の制約を ~41% で通過するモデルは、8 つすべてで成功する確率はわずか 5.7% です。第二に、制約は均等に劣化するわけではありません。構造的制約は、語彙的な制約に比べて、追加された制約ごとに 2 倍多くのベースライン機能を失います。これは、持続的な追跡を必要とする制約と、構成に影響されない二項決定を区別する理解と維持のギャップによって順序付けられます。第三に、失敗はほぼ独立しているため、累積が倍増します。存在する残留結合は、ペアごとの干渉ではなく共有出力特徴を追跡します。間違った文カウントは、それを読み取るすべての制約に失敗します。信頼性の高い命令フォローは、5 ~ 6 個の同時制約を超えると故障します。最も強力なモデルの場合、7 個の制約でプローブ レベルの成功率が 50% を下回り、15 個中 12 個の制約が 3 以下になります。
原文 (English)
Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction
Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas. Individual constraints are handled proficiently, but the compositional regime, where many must hold jointly, remains poorly characterized: how rapidly does performance degrade, what governs the degradation, and can the collapse be mitigated? We introduce Constraint Saturation Evaluation (CSE), a procedurally generated benchmark that systematically varies the number of simultaneous constraints (k), with every constraint scored by a deterministic, rule-based verifier and zero LLM-judge involvement: 15 models, 36 constraint types, 369,753 checks at k=1-12. Three findings emerge. First, per-constraint pass rate decays gradually and predictably, while the chance of satisfying all k constraints collapses - a model passing individual constraints at ~41% at k=8 succeeds on all eight just 5.7% of the time. Second, constraints do not degrade equally: structural constraints lose 2x more baseline capability per added constraint than lexical ones, ordered by a comprehension-maintenance gap that separates constraints requiring sustained tracking from binary decisions immune to composition. Third, failures are nearly independent, which is what makes the accumulation multiplicative; the residual coupling that does exist tracks shared output features rather than pairwise interference - a wrong sentence count fails every constraint that reads it. Reliable instruction following breaks down beyond 5-6 simultaneous constraints: probe-level success falls below 50% at 7 constraints for the strongest model, and at 3 or fewer for 12 of 15.
MindMemOS: AI エージェント向けのポータブルで自己進化するメモリ オペレーティング レイヤー
メモリは AI エージェントの中核コンポーネントであり、AI エージェントが経験を蓄積し、パーソナライゼーションを維持し、長期的な対話に適応できるようにします。しかし、既存の記憶システムは開発後も固定されたままであることが多く、継続的な使用を通じて記憶モデル、組織化戦略、および手続き的知識を適応させる能力が制限されます。 MindMemOS は、統合されたエンティティ プロパティの時間構造を使用してオープンワールド情報を編成する、ポータブルで自己進化するメモリ オペレーティング レイヤーです。 MindMemOS は、シナリオ適応型記憶モデリング、高次パターン発見、自律的な記憶改良、継続的なスキル進化をサポートします。 MindMemEvolve アルゴリズムは、検証主導の進化的検索を採用して、ターゲット シナリオに合わせてメモリ スキーマを最適化します。一方、dreaming は、冗長なレコードをマージし、競合を解決することで蓄積されたメモリを統合します。さらに、暗黙的な修正フィードバックは、潜在的に不正確またはずれている記憶を特定して修正するための人間参加型シグナルとして機能します。 MindSkillEvolve アルゴリズムは、エージェントの実行軌跡を再利用可能で段階的に洗練されたスキルにさらに変換します。 MindMemOS は、LOCOMO で 94.03%、PersonalMem で 70.63% の精度を達成しています。 MindSkillEvolve は、SpreadsheetBench の成功率を初期スキルのベースラインより 9.2 パーセント向上させます。
原文 (English)
MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents
Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions. However, existing memory systems often remain fixed after development, limiting their ability to adapt their memory models, organization strategies, and procedural knowledge through continued use. We present MindMemOS, a portable and self-evolving memory operating layer that organizes open-world information using a unified entity property timestructure. MindMemOS supports scenario-adaptive memory modeling, higher-order pattern discovery, autonomous memory refinement, and continuous skill evolution. Its MindMemEvolve algorithm employs validation-driven evolutionary search to optimize memory schemas for target scenarios, whiledreaming consolidates accumulated memories by merging redundant records and resolving conflicts. In addition, implicit corrective feedback serves as a human-in-the-loop signal for identifying and revising potentially inaccurate or misaligned memories. Its MindSkillEvolve algorithm further transforms agent execution trajectories into reusable and progressively refined skills. MindMemOS achieves 94.03% accuracy on LOCOMO and 70.63% on PersonaMem. MindSkillEvolve improves SpreadsheetBench success by 9.2 percentage points over the initial-skill baseline.
管理された永続メモリ: ロングホライズンエージェントのソースバインド状態セマンティクスとフェールクローズリリース
エージェントの長期記憶は通常、select-store-retrieve として扱われますが、矛盾するレコード、置き換えられたレコード、撤回されたレコード、削除されたレコード、または古いレコードが発信要求をサポートするかどうかは、検索によって決定されるわけではありません。ソースバインドアドミッション、派生ライフサイクル状態、現在のパブリックバリア、およびフェールクローズ構造化リリースを備えた監査可能なバイテンポラル状態遷移モデルである Governed Persistent Memory (GPM) を紹介します。 5 つの実行可能な条項は、台帳の整合性、情報源のバインディング、競合の分離、撤回または削除後の非復活、および 1 つの検証済みヘッドによる新たなビューに基づく正確な請求の閉鎖をカバーしています。事前に指定されたハッシュ凍結された 3,600 件の GPM-ReleaseBench では、GPM はすべての完全な結果と一致します。意図的に単純な 3 つの完全なポリシーのうち最も強力なものは 1,800/3,600 に一致し、違反ケースの 50% で比類のないリリースを行います。個別のシールされたエンドツーエンドのサービス評価では、8 つのクエリ ファミリ全体で実際の取り込みとリリースが実行されます。公開されている V3 アームでは、管理レーンは 2,400/2,400 クラスターでは正しいのに対し、非管理ローカル Qwen2.5-7B では 600/2,400 です。 1,800 件のベースライン障害すべてを回帰なしで修復します (片側 95%、下限 99.875% および 99.834%)。中国および英国のコマンドアームに対するその後の V5 再封印では、世代日の固定と凍結後の減速修正は行われず、再びアームごとに 2,400/2,400 が得られます。実稼働コードに依存しない有限モデルは、完全なコントラクトの反例なしで 331,776 のセマンティック状態と 1,990,656 のクエリ状態を調査し、100,000 トレースの 3 エンジンの差分で不一致がゼロになります。これらは、制限された契約と実装の結果であり、オープンワールド モデルの精度や世界の真実の証拠ではありません。封印されたサービス評価における統制された回答は、決定論的なサービス出力です。 7B の結果は管理されていない比較であり、言語モデル自体が完全に正確になったという主張ではありません。
原文 (English)
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim. We introduce Governed Persistent Memory (GPM), an auditable bitemporal state-transition model with source-bound admission, derived lifecycle state, current public barriers, and fail-closed structured release. Five executable clauses cover ledger integrity, source binding, conflict isolation, non-revival after retraction or deletion, and exact claim closure over a fresh view at one verified head. On a prespecified hash-frozen 3,600-case GPM-ReleaseBench, GPM matches all complete outcomes; the strongest of three intentionally simple complete policies matches 1,800/3,600 and makes unmatched releases on 50% of violation cases. A separate sealed end-to-end service evaluation exercises real ingestion and release across eight query families. In its publicly disclosed V3 arm, the governed lane is correct on 2,400/2,400 clusters versus 600/2,400 for ungoverned local Qwen2.5-7B; it repairs all 1,800 baseline failures with no regression (one-sided 95% lower bounds 99.875% and 99.834%). A later V5 reseal over Chinese- and English-command arms, with generation-date pinning and no post-freeze reducer amendment, again obtains 2,400/2,400 per arm. A production-code-independent finite model explores 331,776 semantic and 1,990,656 query states without a full-contract counterexample, and a 100,000-trace three-engine differential yields zero mismatches. These are bounded contract and implementation results, not open-world model accuracy or evidence of world truth. Governed answers in the sealed service evaluation are deterministic service outputs; the 7B result is the ungoverned comparison, not a claim that a language model itself became perfectly accurate.
$\varepsilon$-MemEvo: LLM プログラム進化のための適応型クロスタスク メモリ転送
FunSearch や AlphaEvolve などの LLM ベースのプログラム進化システムは、新しいアルゴリズムを発見する強力な能力を示していますが、通常は各タスクを個別に最適化し、完了後の検索エクスペリエンスを破棄します。 LLM プログラム進化におけるタスク間の知識伝達のためのフレームワークである $\varepsilon$-MemEvo を紹介します。 $\varepsilon$-MemEvo は、以前の経験をタスクに依存しない戦術記憶として保存します。これは、生のコードではなく、成功したアルゴリズム戦略のコンパクトな自然言語の要約であり、異なる API やエバリュエーターを使用したタスク間での転送を可能にします。意味的に不一致な記憶からの負の転送を回避するために、$\varepsilon$-MemEvo は、取得した記憶を注入するかどうか、およびどの程度の強度で注入するかを決定する適応注入ゲートを使用します。ターゲットタスクのメモリエントリを除外するコンテンツレベルの Leave-One-Out プロトコルを使用して、数学的最適化とシステムエンジニアリングにわたる 8 つの多様な最適化ベンチマークで $\varepsilon$-MemEvo を評価します。プライマリ GPT-5 バックボーンでは、$\varepsilon$-MemEvo は 8 タスクすべてで AdaEvolve よりも AUCC を向上させ、平均相対利得は +8.7% で、初期段階の収束は平均で +9.4% 向上します。アブレーションでは、5 つのアブレーション タスクすべてでアダプティブ ゲーティングが安全なままである一方で、ナイーブ メモリ インジェクションが壊滅的に失敗する可能性があることが示されています。データ更新された事後分布は、観察された状態で解釈可能です。つまり、検索の改善中にスキップが優先され、初期および後期のプラトー全体でスキップからヒントに移行します。これらの利点により生じる計算オーバーヘッドは 1% 未満です。
原文 (English)
$\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution
LLM-based program evolution systems such as FunSearch and AlphaEvolve have shown strong ability to discover novel algorithms, but typically optimize each task in isolation, discarding search experience after completion. We introduce $\varepsilon$-MemEvo, a framework for cross-task knowledge transfer in LLM program evolution. $\varepsilon$-MemEvo stores prior experience as task-agnostic tactic memories: compact natural-language summaries of successful algorithmic strategies rather than raw code, enabling transfer across tasks with different APIs and evaluators. To avoid negative transfer from semantically mismatched memories, $\varepsilon$-MemEvo uses an adaptive injection gate that decides whether retrieved memories should be injected, and at what intensity. We evaluate $\varepsilon$-MemEvo on 8 diverse optimization benchmarks spanning mathematical optimization and systems engineering, using a content-level Leave-One-Out protocol that excludes target-task memory entries. On the primary GPT-5 backbone, $\varepsilon$-MemEvo improves AUCC over AdaEvolve on all 8 tasks, with a mean relative gain of +8.7%, and improves early-stage convergence by +9.4% on average. Ablations show that naive memory injection can fail catastrophically, while adaptive gating remains safe across all five ablation tasks. The data-updated posterior is interpretable in observed states: it favors skip during improving search and shifts from skip to hint across early and late plateaus. These gains incur less than 1% computational overhead.
CAS: ローカルおよびグローバルの説明可能な人工知能の因果関係スコア
予測説明メソッドはモデルの出力に帰属します。それら自体は、介入効果が現実世界の結果に影響を与えるものではありません。因果関係を説明するためのコンパクトなスコア アーキテクチャである Causal Attribution Score (CAS) を紹介します。 CAS は、特定された介入連合ゲームから開始し、共同介入コントラストを因果関係のある Shapley 寄与と割り当て、それらの生の結果スケール効果をローカル CAS、署名済みローカル CAS、および 2 つの補完的なグローバル CAS 要約に変換します。このイノベーションは、新しい Shapley 式ではなく、明示的な介入ターゲットを備えたローカルからグローバルへの因果関係レポート層です。既知の真実のベンチマークでは、8 回繰り返された一次相互作用シミュレーション (それぞれ n = 2,200、3 つのアクション) で、連合認識 CAS の平均ローカル CAS MAE が 0.107 であったのに対し、一度に 1 つずつ正規化した場合は 0.173、グローバル正規化された絶対 ATE ベクトルでは 0.213 でした。一度に 1 つずつ正規化を行う場合のペアの利点は、相加性の場合の -0.003 から、強い相互作用の場合の 0.091 に増加しました。 DoubleML の経験的データセット、401(k) 資格/純金融資産 (n = 9,915) およびペンシルベニア州の再雇用ボーナス/失業期間 (n = 5,099) の両方において、予測 SHAP/TreeSHAP ランキングは、治療効果修飾因子の Feature-CAS ランキングとは大きく異なりました。ペンシルベニア州では、dep1 (正確に 1 つの依存関係) が予測グローバル ランク 13 から Feature-CAS ランク 2 に移動し、主要なローカル Feature-CAS 修飾子となりました。これらの結果は、結果を予測するものと、推定された因果効果の不均一性を説明するものを分離するという付加価値を分離します。
原文 (English)
CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence
Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS starts from an identified interventional coalition game, allocates the joint intervention contrast with causal Shapley contributions, and converts those raw outcome-scale effects into Local CAS, Signed Local CAS, and two complementary Global CAS summaries. The innovation is not a new Shapley formula, but a local-to-global causal reporting layer with an explicit intervention target. In the known-truth benchmark, eight repeated primary-interaction simulations (n = 2,200 each, three actions) gave mean Local CAS MAE of 0.107 for coalition-aware CAS, compared with 0.173 for one-at-a-time normalisation and 0.213 for a global normalised absolute ATE vector. The paired advantage over one-at-a-time normalisation increased from -0.003 under additivity to 0.091 under strong interactions. On both empirical DoubleML datasets, 401(k) eligibility/net financial assets (n = 9,915) and Pennsylvania reemployment bonus/unemployment duration (n = 5,099), predictive SHAP/TreeSHAP rankings differed materially from Feature-CAS rankings of treatment-effect modifiers. In Pennsylvania, dep1 (exactly one dependent) moved from predictive global rank 13 to Feature-CAS rank 2 and was the leading local Feature-CAS modifier. These results isolate the added value of separating what predicts the outcome from what explains heterogeneity in an estimated causal effect.
大規模な有限集合に対する制約付きデコードのためのトライ オートマトン
大規模な言語モデルでは、事前定義されたスキーマに準拠する構造化された出力を生成する必要がますます高まっており、共通の制約の 1 つは有効な文字列の有限セットから選択することです。現在の制約付きデコード システムは、汎用文法コンパイルを通じてこれを処理しますが、有効な値の数が数千に増加すると法外に遅くなり、カーディナリティの壁になります。 Aho-Corasick マルチパターン マッチングを介して有限集合構造 (共有プレフィックス、制限された深さ、既知の基数) を利用してノードごとのトークン マスクを事前計算する特殊なメカニズムであるトライ オートマトンを導入します。このトライは、vLLM および SGLang の主要なバックエンドの 1 つである XGrammar と比較して、ステップごとの有効トークンの計算が 7 倍速く (0.65 us 対 5.8 us)、K >= 300 で 2 ~ 6.5 倍速いコンパイルを実現します。事前計算されたマスクにより、ガイド付きデコード パイプラインをバイパスするステートレスなサービング パスが有効になるため、この利点はバッチ サービング、つまりエンドツーエンド vLLM でさらに強化されます。スループットは、バッチ サイズ 256 (29X) で、XGrammar の 7.5 req/s に対して、219 req/s に達します。 29X は、アルゴリズムによる高速化と、事前計算されたマスクだけが実現できる統合パスの節約を組み合わせています。 7 つのトークナイザー ファミリ (32K ~ 262K 語彙) にわたって、トライは K = 10,000 までの 100ms 未満のコンパイルと、セット サイズに関係なくフラットなステップあたりのコストを維持し、同時に 100% の出力有効性を保証します。
原文 (English)
Trie Automata for Constrained Decoding over Large Finite Sets
Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-purpose grammar compilation, which becomes prohibitively slow as the number of valid values grows into the thousands, a cardinality wall. We introduce the trie automaton, a specialized mechanism that exploits finite-set structure (shared prefixes, bounded depth, known cardinality) via Aho-Corasick multi-pattern matching to precompute per-node token masks. The trie achieves 7X faster per-step valid-token computation (0.65 us vs. 5.8 us) compared to XGrammar, one of the primary backends in vLLM and SGLang, and 2--6.5X faster compilation at K >= 300. Because precomputed masks enable a stateless serving path that bypasses the guided decoding pipeline, this advantage compounds in batch serving: end-to-end vLLM throughput reaches 219 req/s vs. XGrammar's 7.5 req/s at batch size 256 (29X). The 29X combines the algorithmic speedup with integration-path savings that only precomputed masks can unlock. Across seven tokenizer families (32K--262K vocabulary), the trie maintains sub-100ms compilation up to K = 10,000 and flat per-step cost regardless of set size, while guaranteeing 100% output validity.
推論陪審: 推論トレースを評価するためのマルチモデルのコンセンサス
推論 LLM を改善するには、効果的な推論データのキュレーションのための長い推論トレースの品質を判断する能力、強化学習中の強力なトレーニング信号、モデルのパフォーマンス評価中の推論動作の深い理解が必要です。さらに、モデルが犯す推論上の間違いを明らかにすることで、フィードバックを提供することで実行時のモデルのパフォーマンスを向上させることができます。長い推論トレースに対するこの複雑なタスクは難しいため、単一モデルの裁判官 (フロンティア モデルであっても) は推論の欠陥を特定するのが苦手です。さらに、推論 LLM のオンライン トレーニング中にフロンティア モデルを活用することは、使用上のガードレールにより通常禁止されています。この研究では、推論の欠陥を特定するための判断の忠実性を向上させるために、単一の裁判官を LLM の陪審と穏健なコンセンサス メカニズムに置き換えるシステムである推論陪審を導入します。推論陪審では、推論痕跡の欠陥とその深刻さが、司会者が陪審員間で議論を行う審議を通じて表面化され、陪審員はお互いの判断を批判し、最初の投票を修正することができます。モデレータは、陪審員間の審議または判断の統合を通じて合意を導き出します。我々は、オープンウェイト モデル (gpt-oss-120b など) の陪審員による推論陪審が、推論の欠陥を正しく特定する点でフロンティア モデル (opus-4.6、sonnet-4.6、gemini-3.1-pro) を大幅に上回るパフォーマンスを発揮できることを示します。精度のパフォーマンスの向上に加えて、陪審の総コスト (最初の評決、審議、統合など) は、裁判官としての LLM セットアップでフロンティア モデルを実行するコストのほんの一部 (8 ~ 15%) です。また、これらの判断をベンチマーク上の推論 LLM の障害モードを理解するためにどのように活用できるかについても示します。これにより、モデルのパフォーマンスをより深く理解できるようになります。
原文 (English)
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluation. Additionally, surfacing reasoning mistakes that the model makes would enable improving the model's performance at runtime through providing feedback. Due to the difficulty of this complex task on long reasoning traces, single-model judges (even frontier models) do not do well at identifying reasoning defects. Additionally, leveraging frontier models during online training of reasoning LLMs is generally prohibited due to guardrails in terms of use. In this work, we introduce Reasoning Jury, a system that replaces the single judge with a jury of LLMs and a moderated consensus mechanism, to improve the fidelity of judgments for identifying reasoning defects. In reasoning jury, defects of a reasoning trace and their severity are surfaced through a deliberation where a moderator conducts a discussion amongst the jury where the jurors critique each other's judgments and get to modify their initial votes. The moderator derives a consensus through deliberation amongst jurors or consolidation of judgements. We show that Reasoning Jury with a jury of open-weight models (e.g., gpt-oss-120b) is able to significantly outperform frontier models (opus-4.6, sonnet-4.6, and gemini-3.1-pro) at correctly identifying reasoning defects. Besides accuracy performance improvements, the aggregated cost of the jury (initial verdicts, deliberations, consolidation, etc.) is a fraction (8 to 15%) of the cost of running frontier models in LLM-as-a-judge setup. We also show how these judgements can be leveraged to understand failure modes of reasoning LLMs on benchmarks, which allows much deeper understanding of a model's performance.
監査可能なエージェント AI による、証拠に基づく甲状腺超音波診断とレポート作成
甲状腺の超音波診断には、調整された病変の位置特定、測定、リスク層別化、およびレポートが必要ですが、ほとんどの AI システムはこれらのタスクを個別に処理し、臨床レビューに対するサポートは限られています。我々は、専門的な診断ツールを調整し、その出力を監査可能な症例レベルの証拠記録として保存する、臨床医と対話型のエージェント AI システムである ThyroidXAgent を紹介します。このシステムは、約 30 万枚の超音波画像と 24,000 件のペアレポートを統合する多施設マルチタスクリソースである OpenThyroidDB を使用して開発され、民間 NHC-MISD-TUS コホートの 35 施設からの 8,721 件を含む、重複しない 28,458 件のテスト ケースで評価されました。異種データセット全体で、ThyroidXAgent は結節セグメンテーションで 87.21% の平均 Dice スコア、良性-悪性分類で 0.9466 の平均 AUROC を達成しました。同じワークフローは、リンパ節転移予測と濾胞性対乳頭状甲状腺癌の分類をサポートし、AUROC はそれぞれ 0.864 と 0.805 でした。レポート生成では、証拠に基づいたアセンブリが 3 つのコホート全体でマルチモーダル言語モデルのベースラインを上回りました。ここで紹介した病変レベルの臨床意味論的指標である ThyClinScore は、位置認識言語モデルの判定者と最も強い相関関係を示しました。 ThyroidXAgent により、医師の分類精度が向上し、レポート診断の一貫性が 70.3 パーセントから 86.2 パーセントに向上し、セグメント化とレポート時間がそれぞれ 35.9 パーセントと 27.4 パーセント短縮されました。これらの発見は、甲状腺超音波診断とレポートのための、監査可能で臨床医による修正が可能なエージェント AI を裏付けています。
原文 (English)
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.
DiG ベンチ: ゲームにおける発見
発見、つまり新たな一般化を定式化することは、科学プロセスの中心部分です。その重要性にもかかわらず、現在の AI ベンチマーク環境にはギャップがあり、目的が不明な制御された環境での実験によって新しい知識を発見する能力を直接調査するベンチマークはほとんどありません。このギャップに対処するために、新しいベンチマークである DiG-bench (Discovery in Games) をリリースします。 DiG-bench は、70 の独立したゲームのセットで構成されています。各ゲームは短い文字列としてエンコードされており、対話と実験を通じて発見する必要がある独自の変換ルールがあります。ゲームのレベルでは、ルールが発見されたかどうかをテストするための一連の課題が提示されますが、各レベルの勝利条件も不明です。 AI エージェント向けに 7 段階の難易度のゲームを提供します。最下位層は複数のモデルによって日常的に解決可能ですが、最上位層はエージェント ハーネスの最良のモデルに挑戦します。 70 のゲームすべてが、最初の試行で少なくとも 1 人の人間によって解決されました。 21 のゲームのサブセットは公開され、残りは安全な評価のために非公開にされます。
原文 (English)
DiG-bench: Discovery in Games
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. To address this gap, we release a new benchmark: DiG-bench (Discovery in Games). DiG-bench consists of a set of 70 independent games. Each game is encoded as a short string and has unique transformation rules that must be discovered through interaction and experimentation. The levels of the game present a series of challenges to test whether the rules have been discovered, where the win conditions for each level are also unknown. We provide games at seven tiers of difficulty for AI agents. The lowest tier is routinely solvable by multiple models, while the highest tier challenges the best models in agentic harnesses. All 70 games were solved by at least one human on first attempt. A subset of 21 games is released publicly, and the remainder is held private for secure evaluation.
デッドテキストまたは拘束力のある条項?ブラックボックスLLMダイアログにおける制約の影響の測定と復元
マルチターンダイアログを使用すると、ユーザーは制約を課すのと同じくらい簡単に制約を取り消すことができますが、取り消しは確実に有効になるわけではありません。モデルは撤回された要件を制定し続けます (削除を主張するコメントの下に時折あります)。 これは \emph{行動再発} と呼ばれる失敗、または取り消しの慣性と呼ばれます。既存の手段では、条項ごとにこの影響を測定したり、納品前に予測したり、予算に合わせて修理したりすることはできません。 \sysname{} は、モデル API だけで 3 つのギャップを埋めます。契約台帳は、すべての制約を実行可能チェッカーと組み合わせ、失効を廃棄マークとして記録し、事前に正味の制約状態を 1 つの仕様にコンパイルします。連続アブレーションプローブは条項ごとの遵守と漸進的な行動効果を測定します。修理はしごは、トークンと試行に一致する予算の下で動作します。 \dataname{} (\NTasks{} HumanEval タスク、\NClauses{} 検証済みチェッカー) では、制約負荷が増大するにつれて 8B 操作点での再発が \ScaleDelayedMTwo{} から \ScaleDelayedMEight{} まで上昇しますが、より強力なモデルは下位にとどまります。チェッカー、モデル、予算が一致している場合、事前コンパイルにより、台帳なしの検証再試行ベースライン (\RestoreDiff{}、95\% CI \RestoreDiffCI{}、$p$ \RestoreDiffP{}) に対する再発が大幅に減少します。上に積み重ねられた適応ラダー介入は検出可能なゲインを追加しません (95\% の信頼度ではゲイン $\geq$ \LadderExcludedGain{} が除外されます)。このプローブは出産前に再発を予測します (AUROC \AurocPrimary{})。一文の墓石メモは編集効果の約 3 分の 1 を回復し、プラセボ対照でも生き残ります。すべての結果に対する API 計算の \CostdeliveryFactor{} 配信オーバーヘッドと \CostTotalHedged{} では、失効の失敗は、対話状態の目に見えないものではなく、測定可能、予測可能、修復可能なプロパティになります。
原文 (English)
Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues
Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn requirements (occasionally beneath comments asserting their removal), a failure we call \emph{behavioral relapse}, or revocation inertia. No existing instrument measures this influence per clause, predicts it before delivery, or repairs it under matched budgets. \sysname{} closes the three gaps through the model API alone: a contract ledger pairs every constraint with an executable checker, records revocations as tombstones, and compiles the net constraint state ahead of time into a single specification; a sequential ablation probe measures per-clause adherence and incremental behavioral effect; a repair ladder operates under token- and attempt-matched budgets. On \dataname{} (\NTasks{} HumanEval tasks, \NClauses{} verified checkers), relapse at an 8B operating point climbs from \ScaleDelayedMTwo{} to \ScaleDelayedMEight{} as constraint load grows, while stronger models sit at floor. Under matched checkers, model, and budget, ahead-of-time compilation significantly reduces relapse against a no-ledger verifier-retry baseline (\RestoreDiff{}, 95\% CI \RestoreDiffCI{}, $p$ \RestoreDiffP{}); adaptive ladder interventions stacked on top add no detectable gain (95\% confidence excludes gains $\geq$ \LadderExcludedGain{}). The probe predicts relapse before delivery (AUROC \AurocPrimary{}); a one-sentence tombstone note recovers about a third of the compilation effect and survives a placebo control. At \CostDeliveryFactor{} delivery overhead and \CostTotalHedged{} of API compute for every result, revocation failure becomes a measurable, predictable, and repairable property of dialogue state rather than an invisible one.
@スキル: 注意力だけがすべてです
現在、56,804 のパブリック エージェント スキルがあり、チームはさらに多くのスキルを非公開で作成しています。主要な配信モデルはインストールです。インストールされると、スキルの説明はシステム プロンプトに残り、100 未満の信頼できるトリガー スロットをめぐって競合します。これにより、ロングテールには実際に使用できる道がなくなり、チーム独自のプレイブックが同じ希少なスペースを奪い合うことになります。インストールには、コンテンツ、永続性、自動トリガーという 3 つの分離可能な機能がバンドルされていることがわかります。最後の場合のみ即時居住が必要です。そこで私たちは、それらを分離するオープンプロトコルである @skills を提案します。パスはスキル、サブツリー、またはコレクションをアドレス指定しており、スキルを使用するにはスキルを読み取るだけで十分であるため、何もインストールされたり常駐されたりすることはありません。この操作では、適応と所有権のために、同じパスでプロジェクトの Git 追跡ツリーにコピーを提供します。この操作により、.gitignore スタイルの行が 1 つ追加されます。これは、プロンプト常駐にコストがかかる唯一の要素です。ディレクトリはメニューであり、バンドルを全か無かのユニットではなく通常のディレクトリにします。このプロトコルはマニフェスト、ロックファイル、または登録を必要とせず、SKILL.md は変更されません。 @skills は追加的なもので、インストール可能なパッケージとして出荷され、ファイルを読み取ってコマンドを実行できるエージェントを 1 つの命令ファイルを通じてクライアントに変換します。そのオープン仕様は https://github.com/SylphAI-Inc/atskills にあり、 https://adalagent.ai の AdaL CLI に実装されています。パスはスキルに適切に対応しますが、スキルを見つけることができないため、このプロトコルは、コーパス全体の検索とランキング、リポジトリ不要のホスティング、プライベートおよびチームのコレクション、および 1 画面のオーサリングのために https://atskills.one の無料ハブと組み合わせられています。ハブはオプションです: gh: と、ローカル パスはハブなしで解決され、インデックス付きの GitHub スキルは gh: の ID を保持します。インストールを減らして、より多くの使用を行います。
原文 (English)
@skills: Attention is all you have
There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery model is installation: once installed, a skill's description remains in the system prompt, competing for fewer than 100 reliable trigger slots. This leaves the long tail with no practical path to use and forces teams' own playbooks to compete for the same scarce space. We observe that installation bundles three separable functions: content, persistence, and automatic triggering. Only the last requires prompt residency. We therefore propose @skills, an open protocol that separates them. A path addresses any skill, subtree, or collection, and reading a skill is sufficient to use it, so nothing is installed or made resident. The operation vendors a copy at the same path into a project's Git-tracked tree for adaptation and ownership. The operation adds one .gitignore-style line, the only element that costs prompt residency. A directory is a menu, making bundles ordinary directories rather than all-or-nothing units. The protocol requires no manifest, lockfile, or registration, and SKILL.md remains unchanged. @skills is additive, ships as an installable package, and turns any agent that can read files and run commands into a client through a single instruction file. Its open specification is at https://github.com/SylphAI-Inc/atskills and it is implemented in the AdaL CLI at https://adalagent.ai . Because paths address skills well but cannot find them, the protocol is paired with a free hub at https://atskills.one for corpus-wide search and ranking, repository-free hosting, private and team collections, and one-screen authoring. The hub is optional: gh: and local paths resolve without it, and indexed GitHub skills retain their gh: identities. Install less, use more.
ギザギザの裁判官: 沈黙、圧力、執拗な状況下での認識的安定性
LLM 審査員は、モデルの評価、オンライン採点、報酬モデリングの中心的なインフラストラクチャとなっています。裁判官は通常、ゴールデンデータの正確性によって検証されますが、再プロンプト、異議申し立て、または持続的な反発の下で裁判官が安定しているかどうかについては、正確性はほとんど影響しません。私たちは、LLM 裁判官の認識安定性を評価するための統一ストレス テストである \emph{Wiggle Framework} を導入します。このフレームワークは、機械的一貫性 (再プロンプトと再フレーム化の下での安定性)、シングルターン確信 (単一の課題の下での安定性)、およびマルチターン持続性 (持続的または適応的なプレッシャー下での安定性) の 3 つの次元に沿って判断の堅牢性を分解します。私たちはこのフレームワークを使用して、安全性、毒性、AI 書き込み検出、政治的対応評価にわたる 14 の審査タスクにわたって 9 つのフロンティア モデルを研究します。すべてのモデルは、裁判官としてかなりの動きを示します。静的なプッシュバックでは 25 ~ 71\% の確率で評決を覆し、敵対的な LLM 説得では 62 ~ 91\% の確率で評決を覆します。重要なことに、裁判官の評決を変えることに成功する圧力は、ほとんどの場合、グラウンドトゥルースに関してネットを破壊するものであることがわかります。フレームワーク自体を超えて、私たちは、どの項目が変動するかを予測するための最も効果的な単発シグナルとして、ベースラインの陪審過半数の強さを特定します。総合すると、これは、判定のコンテキストにおける機械的テスト、適合性テスト、および説得力テストのデータセット間での初めての同一の比較です。
原文 (English)
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and political-response evaluation. Every model exhibits substantial wiggle as a judge --- flipping verdicts 25--71\% of the time under static pushback, and 62--91\% with an adversarial LLM persuader. Critically, we find that pressure that succeeds in changing a judge's verdict is almost always net-corrupting with respect to ground truth. Beyond the framework itself, we identify baseline jury majority strength as the most effective single-shot signal for anticipating which items wiggle. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.
SteerBench-Work: アクション境界でのエージェントのステアリングのベンチマーク
長期間実行される LLM エージェントはツールを通じて動作し、単一のステップで電子メールの送信、プル リクエストのマージ、または電信送金を行うことができます。ステアリングの決定は、その境界におけるコミット前の選択です。続行するか、人間によるレビューまたはポリシーのレビューのために保留するかです。 SteerBench-Work は、開発業務、顧客サービス、財務、法務、医療、人事、セキュリティにわたる職場エージェントの意思決定のための、インシデントに基づいた双方向のベンチマークです。リリース v2026-05 には、公的事件に基づいた 106 のシナリオ、ペアの証拠反転ミラー、およびキャリブレーション コントロールが含まれており、ラベルは続行と保留の間でほぼ均等に分割されているため、2 つのエラー方向がほぼ同じ数のチャンスを得ることができます。モデルは、提案されたアクションと利用可能な証拠を確認し、ゲート決定を返し、それが境界を正しく越えているか、境界を保持しているかどうかでスコア付けされます。 30 のモデル条件全体で、失敗はほぼ完全に一方向に進行します。モデルは、機会の 28.1% で承認済みの証拠がクリアされた作業を誤って保持し、1.0% で安全でない作業を誤って許可します。最も困難なケースは、リスク解決されたコミットです。この場合、署名または構造化された証拠が実際のリスクトリガーをすでに除去しており、有名なインシデントの証拠を反転したミラー (63.8%) では、インシデント自体 (98.5%) よりもモデルのスコアが著しく低くなります。一般的な機能はステアリング キャリブレーションと同じではありません。高機能モデルはコミット境界で過剰に拒否されることがよくあり、より多くの推論により、キャリブレーションされたゲートをフラットのままにして弱いゲートを修復できます。公開リーダーボードは steerbench.com にあります。
原文 (English)
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy review. We introduce SteerBench-Work, an incident-anchored, bidirectional benchmark for that decision in workplace agents across developer operations, customer service, finance, legal, medical, HR, and security. Release v2026-05 contains 106 scenarios anchored in public incidents, paired evidence-reversed mirrors, and calibration controls, with labels split nearly evenly between proceed and hold so the two error directions get near-identical numbers of chances. A model sees the proposed action and the available evidence, returns a gate decision, and is scored on whether it crosses or holds the boundary correctly. Across 30 model conditions the failures run almost entirely in one direction: models wrongly hold authorized, evidence-cleared work on 28.1% of opportunities and wrongly allow unsafe work on 1.0%. The hardest cases are risk-resolved commits, where signed or structured evidence has already cleared a real risk trigger, and models score markedly worse on evidence-reversed mirrors of famous incidents (63.8%) than on the incidents themselves (98.5%). General capability is not the same as steering calibration: higher-capability models often over-refuse at the commit boundary, and more reasoning can repair a weak gate while leaving a calibrated one flat. The public leaderboard is at steerbench.com.
因果関係の知識による一般的な因果関係の確率
因果関係の確率 (PoC) は、直接観察できないため、一般に部分的な特定が必要な個々の因果関係の応答を特徴付けます。 Tian と Pearl は最初に、必要性の確率 (PN)、十分性の確率 (PS)、および必要性と十分性の確率 (PNS) を含む、バイナリ PoC の理論的に明確な境界を導き出しました。ミュラーら。その後、共変量とメディエーターにコード化された因果情報を組み込むことにより、バイナリ PNS の境界を厳しくしました。最近では、Li と Pearl、および Shu らは、PoC を多値設定に拡張し、対応する理論的限界を導き出しました。これらの発展により、当然のことながら、追加の因果関係の知識によって多値設定の境界をさらに厳しくできるかどうかという疑問が生じます。この論文では、共変量とメディエーターにエンコードされた因果情報を組み込むことで、多値 PoC のより厳密な境界を導き出すことで、この問題に対処します。おもちゃの例を使って理論的結果を説明しますが、シミュレーション研究では、提案された境界が既存の非バイナリ境界よりも厳しいことをさらに実証しています。
原文 (English)
General Probabilities of Causation with Causal Knowledge
Probabilities of causation (PoCs) characterize individual causal responses that cannot be directly observed and therefore generally require partial identification. Tian and Pearl first derived theoretically sharp bounds for binary PoCs, including the probability of necessity (PN), the probability of sufficiency (PS), and the probability of necessity and sufficiency (PNS). Mueller et al. subsequently tightened the bounds for binary PNS by incorporating causal information encoded in covariates and mediators. More recently, Li and Pearl, as well as Shu et al., extended PoCs to multivalued settings and derived corresponding theoretical bounds. These developments naturally raise the question of whether additional causal knowledge can further tighten the bounds in multivalued settings. This paper addresses this question by deriving tighter bounds for multivalued PoCs through the incorporation of causal information encoded in covariates and mediators. We illustrate the theoretical results with toy examples, while simulation studies further demonstrate that the proposed bounds are tighter than existing nonbinary bounds.
意思決定に即した ITSM インテリジェンスのための AI パイプラインの設計
IT サービス管理 (ITSM) システムには、営業担当者や経営陣の関係者が実用的なインテリジェンスに変換するのが難しい、異種チケット データが大量に蓄積されます。このペーパーでは、生の ITSM エクスポートをマルチレベルの意思決定支援アーティファクトに変換する、設計科学研究原則に従って設計および評価された社会工学 AI パイプラインについて説明します。このパイプラインは、LLM ベースのスキーマ正規化、HDBSCAN サブトピック クラスタリング、および階層的集合クラスタリングを組み合わせて、経営幹部向けのメイン トピックと詳細なサブトピックを生成します。 6 つの成果物と、セールス エンジニアリングおよびカスタマー サクセスの役割からの 5 人の評価者による関係者評価では、解釈可能性、実行可能性、信頼性、使用の可能性という 4 つの意思決定支援指標がすべて平均して 5.0 点中 4.0 点を超えており、信頼性が最も一貫したシグナルであることが示されています。この調査結果は、ITSM 分析を、変換、抽象化、人間中心の設計という情報システム (IS) の問題として位置づけています。
原文 (English)
Designing AI Pipelines for Decision-Ready ITSM Intelligence
IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are difficult for sales and executive stakeholders to convert into actionable intelligence. This paper presents a sociotechnical AI pipeline, designed and evaluated following design science research principles, that transforms raw ITSM exports into a multilevel decision-support artifact. The pipeline combines LLM-based schema normalization, HDBSCAN sub-topic clustering, and hierarchical agglomerative clustering to generate executive-facing Main-topics and granular Sub-topics. A stakeholder evaluation across six artifacts and five raters from Sales Engineering and customer success roles shows that all four decision-support metrics, interpretability, actionability, trust, and likelihood of use, on average exceed 4.0 out of 5.0, with trust as the most consistent signal. The findings position ITSM analytics as an Information Systems (IS) problem of transformation, abstraction, and human-centered design.
トランスフォーマーの表現力について
多層トランスフォーマーは、現在使用されている基本的にすべての大規模言語モデル (LLM) の重要なコンポーネントを形成します。トランスフォーマーの遍在性と計算能力により、理論コンピューターサイエンスコミュニティによって数十年にわたって研究されてきた計算の標準モデルとトランスフォーマーを比較することにより、言語認識装置としてのトランスフォーマーの表現力を正確に調整することを目的とした一連の研究が急速に成長しています。この取り組みにおいて、回路の複雑さは概して、トランスの表現力を分析するための計算の複雑さの「正しい」分野として浮上しました。その理由は、注意力や精度など、使用するさまざまなリソースによってトランスをパラメータ化すると、ゲートのタイプ、サイズ、深さなどのリソースによってパラメータ化されたさまざまなクラスの回路との直接比較につながるためです。ここでは、回路の複雑さから概念と手法を使用してトランスの表現力を描写する、選択された結果の概要を示します。
原文 (English)
On the Expressive Power of Transformers
Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquity and computational capability, there is a rapidly growing body of work that aims to precisely calibrate the expressive power of transformers as language recognizers by comparing them against standard models of computation studied for decades by the theoretical computer science community. In this endeavor, circuit complexity has by and large emerged as the "correct" branch of computational complexity to analyze the expressive power of transformers; the reason is that parameterizing transformers by the various resources they use, such as attention and precision, leads to direct comparisons with different classes of circuits parameterized by resources such as type of gates, size, and depth. Here, we present an overview of selected results that delineate the expressive power of transformers using concepts and methods from circuit complexity.
ラインとラダー: 大規模な小売価格分類のためのコンテキスト認識型マルチエージェント フレームワーク
価格の一貫性を維持し、毎日低価格戦略を実行することは、世界的な小売業者にとって非常に重要です。ただし、カタログが何百万ものアクティブなアイテムにまたがっているため、価格関係を手動で管理することは不可能です。品目のバリエーション間で価格設定が一貫していない場合、顧客の価値認識が歪められ、売上が共食いされます。これに対処するために、「ラインとラダー」の価格分類法の構築を自動化するように設計された、スケーラブルでコンテキスト認識型のマルチエージェント フレームワークを紹介します。当社のフレームワークでは、特殊な LLM エージェントを採用して、主要な属性を特定し、マルチモーダル値を抽出し、階層的なグループ化ロジックを適用することで、一貫した価格設定構造を構築します。現実世界のエンタープライズ データに基づいて評価され、運用環境に導入された当社の 3 エージェント システムは、ラインの F1 スコア 0.83 を達成し、認知過負荷を軽減することで単一エージェントのベースラインを上回りました。このシステムは、食品および消耗品では 90% 以上の精度と 75% 以上の再現率を達成し、非構造化一般商品カタログでは 80.2% の割り当て精度を達成しています。
原文 (English)
Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers. However, with catalogs spanning millions of active items, manual governance of price relationships is infeasible. Inconsistent pricing across item variants distorts customer value perception and cannibalizes sales. To address this, we present a scalable, context-aware Multi-Agent Framework designed to automate the construction of "Lines and Ladders" pricing taxonomies. Our framework employs specialized LLM agents to construct these coherent pricing structures by identifying key attributes, extracting multi-modal values, and applying hierarchical grouping logic. Evaluated on real-world enterprise data and deployed in production, our 3-Agent system achieves an F1-score of 0.83 for Lines, outperforming single-agent baselines by mitigating cognitive overload. The system achieves >90% precision and >75% recall in Food & Consumables, and 80.2% assignment accuracy in the unstructured General Merchandise catalog.
機密情報を外部 LLM から隠蔽することでプライバシーを保護する RAG
検索拡張生成 (RAG) は、ユーザーのクエリに応答する際の大規模言語モデル (LLM) のパフォーマンスを向上させるために広く使用されています。 RAG に関する既存のプライバシー調査は、権限のないユーザーによる機密データへのアクセスの防止に焦点を当ててきました。ただし、RAG プライバシー調査で見落とされがちなもう 1 つの重要な問題は、外部ジェネレータがクエリと取得したドキュメントにアクセスできることです。これらのドキュメントには、悪用されたり、意図しない目的でアクセスされたりする可能性のある機密情報が含まれている可能性があります。このペーパーでは、ユーザーが機密情報を開示せずに強力なサードパーティ ジェネレーターを利用できるようにするプライバシー保護フレームワークである Sensitive Entity Alias Generator (SEAG) を紹介します。 SEAG は、機密エンティティを特定し、対応するエイリアスを生成し、エンティティ置換テーブルを構築する軽量モデルを導入しています。このテーブルは、ユーザーのクエリおよび取得されたドキュメント内の機密単語を外部ジェネレーターに転送する前に置き換えるために使用されます。この目的のために、2 つのデータセットが構築されました。1 つは SEAG モデルを微調整してエンティティ置換テーブルを生成するためのもので、もう 1 つは SEAG フレームワーク全体を評価するためのものです。実験結果は、SEAG フレームワークの成功を示しています。機密情報を外部ジェネレーターから隠しながらユーザーに正しい応答を提供するモデルの能力を測定するユーザー指標に関しては、すべての SEAG モデルが 80% 以上の精度を達成しました。追加の分析では、SEAG モデル Qwen-3、LLaMA-3.2、および Phi-4 が特定の文書内のすべての機密エンティティを非表示にする能力をさらに評価しました。結果は、合計精度がそれぞれ 77.83%、76.73%、74.91% という良好なパフォーマンスを示しています。
原文 (English)
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
Retrieval-Augmented Generation (RAG) is widely used to improve the performance of Large Language Models (LLMs) in answering user queries. Existing privacy research on RAG has focused on preventing unauthorized users from accessing sensitive data. However, another important problem that is often overlooked in RAG privacy research is that external generators have access to the query and the retrieved documents, which may contain confidential information that could potentially be misused or accessed for unintended purposes. In this paper, we introduce the Sensitive Entity Alias Generator (SEAG), a privacy-preserving framework that empowers users to utilize powerful third-party generators without disclosing sensitive information. SEAG introduces a lightweight model that locates sensitive entities, generates corresponding aliases, and constructs an entity replacement table. The table is used to replace sensitive words in the user's query and in the retrieved documents before they are forwarded to an external generator. For this purpose, two datasets were constructed: one for fine-tuning SEAG models to generate entity replacement tables, and another for evaluating the entire SEAG framework. The experimental results demonstrate the success of the SEAG framework. As for the User metric, which measures the ability of the model to provide a correct response to the user while hiding sensitive information from the external generator, all SEAG models achieved over 80% accuracy. Additional analysis further evaluated the ability of SEAG models Qwen-3, LLaMA-3.2, and Phi-4 to hide all sensitive entities within given documents. The results show good performance with total accuracies of 77.83%, 76.73%, and 74.91%, respectively.
マルチモーダルビデオベースのデング熱診断における自然言語理解の役割
蚊は小さく、素早く不規則に動き、背景、照明、影などの環境要因の影響を受けるため、信頼性の高い特徴抽出が困難になる可能性があるため、ビデオ データから感染に関連した蚊の行動変化を検出することは困難です。この研究では、未感染の蚊とデング熱ウイルス血清型 2 (DENV2) に感染した蚊の蚊の飛行フレームを分類するために、YOLO および対照言語画像事前トレーニング (CLIP) ベースの視覚言語フレームワークが提案されています。まず、YOLO を使用して蚊の領域を背景から分離します。次に、ビデオ フレームから抽出された視覚的特徴が、共有埋め込みスペース内の生物学的に意味のあるテキスト プロンプトと位置合わせされます。マルチモーダル モデルは、教師あり双方向対比学習を使用して微調整され、フレーム レベルの画像とテキストの類似性に基づく分類を通じて評価されました。結果は、提案された方法がフレーム レベルで 98.54% の精度と 99.91% の感度を達成したことを示しています。フレームレベルの情報を一時的に集約した後、モデルは完全なビデオレベルのパフォーマンスを達成しました。アブレーションの結果は、この領域には微調整と CLIP ベースの表現が不可欠である一方、テキスト ブランチは視覚のみのモデルよりも正確さの利点ではなく、意味論的な画像とテキストの位置合わせを提供することを示しました。これらの発見は、視覚言語モデルがビデオデータから感染に関連した生物学的挙動を分析するための有用なフレームワークを提供できることを示唆しています。
原文 (English)
The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis
Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature extraction difficult. In this study, a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework is proposed to classify mosquito flight frames of uninfected and Dengue virus serotype 2 (DENV2)-infected mosquitoes. First, YOLO is used to isolate mosquito regions from the background. Then, visual features extracted from video frames are aligned with biologically meaningful textual prompts in a shared embedding space. The multimodal model was fine-tuned using supervised bidirectional contrastive learning and evaluated through frame-level image-text similarity-based classification. The results show that the proposed method achieved 98.54% accuracy and 99.91% sensitivity at the frame level. After temporal aggregation of frame-level information, the model achieved complete video-level performance. The ablation results showed that fine-tuning and CLIP-based representations were essential for this domain, while the textual branch provided semantic image-text alignment rather than an accuracy advantage over the vision-only model. These findings suggest that vision-language models can provide a useful framework for analyzing infection-related biological behaviors from video data.
最善の推測を超えて: 進化戦略による LLM ソリューション カバレッジの向上
大規模言語モデル (LLM) は、数学や科学などの発見ドメインに導入されることが増えています。通常のアプローチは、問題をモデルに提示し、その答えを提案された解決策として使用することです。ただし、この最善の推測を超えて、テスト時のコンピューティングを増やすことで検出を強化できます。 pass@k と呼ばれるプロセスでは、モデルは解空間を探索し、多様な候補解を生成できます。残念ながら、強化学習 (RL) による LLM のポストトレーニングへの標準的なアプローチでは pass@k が制限される可能性があります。モデルの出力分布が高報酬出力付近で狭くなり、ソリューション カバレッジが崩壊します。別の方法は、ランダムな摂動を通じて重み空間で直接最適化する母集団ベースの勾配のないトレーニング後手法であるEvolution Strategies (ES) を使用することです。この論文が示すように、ES は RL よりも一貫して高い pass@k を達成し、より広いソリューション カバレッジを持つより広い出力分布を生成します。このカバレッジにより、たとえば次の分野でより良い結果を達成することが可能になります。標準的な数学ベンチマーク。したがって、ES は、発見問題や、多様なソリューションの範囲が重要なその他の領域における事後トレーニングのためのより良い基盤を提供します。
原文 (English)
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its answer as the proposed solution. However, beyond this best guess, discovery can be enhanced by increasing test-time compute. In a process called pass@k, the model is allowed to explore the solution space and generate diverse candidate solutions. Unfortunately, the standard approach to post-training LLMs through Reinforcement Learning (RL) may limit pass@k: the model's output distribution narrows around high-reward outputs, causing the solution coverage to collapse. The alternative is to use Evolution Strategies (ES), a population-based, gradient-free post-training method that optimizes directly in weight space through random perturbations. As this paper shows, ES achieves consistently higher pass@k than RL and produces a broader output distribution with greater solution coverage. This coverage in turn makes it possible to achieve better results in e.g. standard math benchmarks. Thus, ES provides a better foundation for post-training in discovery problems and other domains where diverse solution coverage is critical.
空間記憶エージェント: 空間知性のための経験に基づいた手順記憶
空間インテリジェンスは、身体化されたエージェント、ロボットによるプランニング、およびマルチモーダル アシスタントの基盤になりつつあります。 VLM エージェントの空間推論能力を向上させるために、既存の研究は主に 2 つの方針に従っています。 1 行では、教師あり微調整や強化学習などのトレーニング後の手法を使用しています。別のラインでは、モデルが深さ推定や 3D 再構成ツールなどの外部空間ツールを呼び出して、中間の空間証拠を収集するエージェント パラダイムを採用しています。私たちは、補完的かつ未開拓のルートを研究します。凍結された VLM エージェントは、推論時に外部の専門家空間ツールに依存せずに、\textbf{パラメータ更新不要の自己進化} を通じて空間推論を改善できるでしょうか?私たちは、検証された空間体験を再利用可能な転送可能なレッスンに変換する \textbf{経験に基づいたランタイム フレームワーク}である \textbf{空間メモリ エージェント (SMA)} を紹介します。検証可能な空間環境では、SMA は凍結された VLM にクエリを実行し、予測された答えと報酬を取得し、\textbf{検証者ガイド付きリフレクション} を使用して、空間経験からコンパクトで譲渡可能な教訓を抽出します。 SMA はさらに、各レッスンに \textbf{転送信頼性スコア (TRS)} を割り当てます。これは均一に初期化され、将来の転送信頼性の訪問証拠として後の検索結果から調整されます。 \textbf{読み取り専用デプロイメント}中、SMA はセマンティック フィルターと類似性と TRS を組み合わせたランキングによってレッスンを取得し、取得したメモリを利用してフリーズされたモデルの推論をガイドできるようにします。 5 つの代表的な空間ベンチマークと 4 つのベース VLM にわたって、SMA はすべてのベース モデル ブロックで最高のマクロ平均を達成し、20 の評価のほとんどで評価されたメソッドの中で最高の精度を達成し、評価されたフリーズ モデル スケールと環境全体で空間自己進化のための実用的なパラメーター更新のないパスを確立しました。
原文 (English)
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools, such as depth estimation and 3D reconstruction tools, to gather intermediate spatial evidence. We study a complementary and underexplored route: Can a frozen VLM agent improve its spatial reasoning through \textbf{parameter-update-free self-evolution}, without depending on external expert spatial tools at inference time? We present \textbf{Spatial Memory Agent (SMA)}, an \textbf{experience-grounded runtime framework} that converts verified spatial experience into reusable transferable lessons. In a verifiable spatial environment, SMA queries the frozen VLM, obtains a predicted answer and reward, and uses \textbf{verifier-guided reflection} to distill compact transferable lessons from spatial experience. SMA further assigns each lesson a \textbf{Transfer Reliability Score (TRS)}, which is initialized uniformly and calibrated from later retrieval outcomes as visit evidence of future transfer reliability. During \textbf{read-only deployment}, SMA retrieves lessons by semantic filter and similarity-TRS combined ranking, allowing the retrieved memory to guide frozen model inference. Across five representative spatial benchmarks and four base VLMs, SMA achieves the highest macro average in every base-model block and the best accuracy among the evaluated methods in most of the 20 evaluations, establishing a practical parameter-update-free path for spatial self-evolution across the evaluated frozen model scales and environments.
正しさは管理されない: エージェント ワークフローにおける来歴の整合性
エージェント ワークフローは通常、正しい結果に達するかどうかによって評価されます。これは、正しい行動が間違った権限、裏付けのない完了要求、または後の変更によって陳腐化した作業に依存する可能性がある制度的環境では不十分です。私たちは、管理された実行を、検査可能な出所によって決定、完了、変更への対応がサポートされる作業と定義します。私たちは、権威と事実の依存関係を記録し、完了の証拠を検証し、影響を受けた作業を選択的に無効にする決定論的な因果状態レイヤーである Matrix を紹介します。管理された比較全体で、管理されたワークフローと直接的なワークフローは多くの場合同じ結果に達しましたが、管理されたパスのみが一貫して管理証拠を保持し、サポートされていない終了を拒否し、依存タスクへのリカバリが制限されました。その後、役割分離転送チャレンジは失敗しました。決定論的に強制された完全性コントラクトが、オーサリング コンテキストの外で生成された合成パケットを大幅に過剰ブロックしました。これらの結果は、Matrix が一般的な精度向上剤として確立されるものではありません。これらは、エージェントの作業を監査可能かつ独立して検証可能にするための組織的整合性層としての主な役割をサポートします。
原文 (English)
Correct Is Not Governed: Provenance Integrity in Agentic Workflows
Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later change. We define governed execution as work whose decisions, completion, and response to change are supported by inspectable provenance. We present Matrix, a deterministic causal-state layer that records authority and fact dependencies, verifies completion evidence, and selectively invalidates affected work. Across controlled comparisons, governed and direct workflows often reached the same outcomes, but only the governed path consistently preserved governing evidence, refused unsupported closure, and limited recovery to dependent tasks. A role-separated transfer challenge then failed: a deterministically enforced completeness contract severely over-blocked synthetic packets produced outside its authoring context. These results do not establish Matrix as a general accuracy enhancer; they support its primary role as an institutional integrity layer for making agentic work auditable and independently verifiable.
PROVE-RT: LLM を使用したリアルタイム システム用の機械化定理証明者スクリプトの生成
スケジュール可能性分析はリアルタイム システムを認証するために不可欠ですが、既存のテストは多くの場合、拡張、検証、保守が困難なペンと紙の証明を通じて開発されています。 PROSA/ROCQ の機械化検証は厳密な代替手段を提供しますが、そのような証明を手動で構築するには、相当な分野の専門知識と証明エンジニアリングの労力が必要です。幅広いタスクにわたる大規模言語モデル (LLM) の最近の成功により、LLM は機械化された定理証明者向けの PROSA/ROCQ スクリプトを生成するための有望な候補となっています。ただし、最先端の LLM には、モデリング抽象化と証明パターンを正しく使用するために必要な PROSA 固有の知識が欠けていることがよくあります。この文書では、リアルタイム システムの文献でスケジューラビリティ分析を機械化する PROSA/ROCQ スクリプトを生成するための LLM 支援フレームワークである PROVE-RT を紹介します。 PROVE-RT は、依存関係を意識した非公式スケッチ、処理された PROSA ドキュメントからの取得、段階的なスケルトン生成、および証明の完了を通じて生成をガイドします。私たちは、依存関係情報を含む 13,134 の非公式スケッチを含む 1,191 のリアルタイム システム論文から機械化指向のコーパスを構築します。厳選された評価セットでは、最先端の LLM を直接プロンプトしても有効な PROSA 機構を確実に生成できませんが、PROVE-RT は 44.7% の成功率を達成しています。これらの結果は、検索ガイドと段階的な LLM 支援により、PROSA/ROCQ におけるスケジュール可能性分析の自動機械化を改善できることを示しています。
原文 (English)
PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs
Schedulability analysis is essential for certifying real-time systems, but existing tests are often developed through pen-and-paper proofs that are difficult to scale, validate, and maintain. Mechanized verification in PROSA/ROCQ offers a rigorous alternative, yet manually constructing such proofs requires substantial domain expertise and proof-engineering effort. Recent successes of large language models (LLMs) across a wide range of tasks make them promising candidates for generating PROSA/ROCQ scripts for mechanized theorem provers. However, state-of-the-art LLMs often lack the PROSA-specific knowledge required to correctly use its modeling abstractions and proof patterns. This paper introduces PROVE-RT, an LLM-assisted framework for generating PROSA/ROCQ scripts to mechanize schedulability analyses in real-time systems literature. PROVE-RT guides generation through dependency-aware informal sketches, retrieval from processed PROSA documentation, staged skeleton generation, and proof completion. We construct a mechanization-oriented corpus from 1, 191 real-time systems papers, containing 13, 134 informal sketches with dependency information. On a curated evaluation set, direct prompting of state-of-the-art LLMs fails to reliably generate valid PROSA mechanizations, whereas PROVE-RT achieves a success rate of 44.7%. These results show that retrieval-guided and staged LLM assistance can improve automated mechanization of schedulability analysis in PROSA/ROCQ.
ARAC: エンドツーエンドのリサーチにおける Auto-Research の調整と完全性のベンチマーク
自動研究の急速な進歩により、基本的な評価の課題が表面化しました。それは、その研究の軌跡と人間の研究行動との整合性、論理的一貫性、進化の完全性をどのように測定できるのでしょうか?私たちは、Auto-Research の整合性と完全性である ARAC-Bench を提案します。これは、目的を最終的な答えの一致から人間による高品質の研究プロセスの再現に移行する、研究者を模倣した評価フレームワークです。このフレームワークは、2 つの相乗的なコンポーネントを通じて機能します。1 つは、暗黙の査読者の専門知識を、段階的に調整された定量化可能なルーブリックに変換する最初のシステムである、Academic Cognition Skills システムです。 3 段階の能力診断プロトコルは、厳格なモジュール制約の下で研究プロセスを、追跡可能で相互に独立した 3 つの側面、つまり提案、実験、合成に分解します。 11 個の SOTA フレームワークを体系的に評価した結果、最良のアラインメント スコアは 100 点中 67.9 点にすぎず、人間による厳密な方法論のシミュレーションにおいては大きなギャップがあることが明らかになりました。 Ph.D に対する検証候補者のランキングは 0.8141 という強い相関関係を示しており、ARAC-Bench が研究者が真に評価する次元を確実に反映していることが確認されています。 ARAC-Bench は、きめ細かい診断ツールだけでなく、次世代の自律研究システムをトレーニングするためのスケーラブルな報酬信号も提供します。
原文 (English)
ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs
The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.
CABS+: 競合を意識したスパース化と適応的な重み割り当てによる効率的でスケーラブルなモデルの結合
モデルのマージは、追加の再トレーニングを必要とせずに統合されたマルチタスク モデルを構築するための有望なパラダイムとして、最近大きな注目を集めています。ただし、タスク間のパラメーターの競合や知識の干渉により、マージされたモデルのパフォーマンスが低下することがよくあります。以前の研究では、構造化されたプルーニングとシーケンシャル マスキングを通じてパラメーターの干渉を低減する、Conflict-Aware and Balanced Sparsification (CABS) が導入されました。ただし、CABS はスケーリング係数を決定するためにグリッド検索に依存しているため、指数関数的な時間計算量が発生する一方、その最適化目標が高パフォーマンスのタスクに支配され、全体的なパフォーマンスが最適化されていない可能性があります。これらの制限に対処するために、私たちは CABS を拡張し、CABS+ を提案します。具体的には、Adaptive Weight Allocation (AWA) が勾配のない探索スキームを介して結合係数を最適化し、時間の複雑さを軽減する一方、非対称のフィットネス関数がタスク全体でより包括的なパフォーマンスの向上を促進します。さらに、モデルのマージパフォーマンスに影響を与える主要な要因の体系的な実証研究を実施し、モデルのマージ可能性を定量化し、モデル選択のガイドとなる相対シナジースコア (RSS) を提案します。大規模言語、小規模言語、視覚モデルをカバーする 27 のデータセットと 5 つのモデルにわたって、CABS+ と CABS、AdaMerging、WUDIMerging などの最先端のモデル マージ手法を比較します。広範な実験により、CABS+ の有効性と効率性が検証されています。 AdaMerging と WUDIMerging と比較して、CABS+ は全体のパフォーマンスをそれぞれ 16.97% と 12.93% 向上させ、さまざまなタスク数とモデル アーキテクチャにわたってより強力な安定性と堅牢性を示し、AdaMerging に必要な GPU メモリの使用量を 25% 未満にし、WUDIMerging と比べてマージ時間のほぼ 4 倍の高速化を達成します。
原文 (English)
CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation
Model merging has recently attracted significant attention as a promising paradigm for constructing unified multi-task models without requiring additional retraining. However, parameter conflicts and knowledge interference across tasks often degrade merged-model performance. Prior work introduced Conflict-Aware and Balanced Sparsification (CABS), which reduces parameter interference through structured pruning and sequential masking. However, CABS relies on grid search to determine scaling coefficients, resulting in exponential time complexity, while its optimization objective can be dominated by high-performance tasks, leading to suboptimal overall performance. To address these limitations, we extend CABS and propose CABS+. Specifically, Adaptive Weight Allocation (AWA) optimizes merging coefficients via a gradient-free search scheme to reduce time complexity, while an asymmetric fitness function promotes more comprehensive performance gains across tasks. Moreover, we conduct a systematic empirical study of key factors influencing model merging performance and propose Relative Synergy Score (RSS) to quantify model mergeability and guide model selection. We compare CABS+ with state-of-the-art model merging methods, including CABS, AdaMerging, and WUDIMerging, across 27 datasets and 5 models covering large language, small-scale language, and vision models. Extensive experiments verify the effectiveness and efficiency of CABS+. Compared with AdaMerging and WUDIMerging, CABS+ improves overall performance by 16.97% and 12.93%, respectively, exhibits stronger stability and robustness across varying task numbers and model architectures, uses less than 25% of the GPU memory required by AdaMerging, and achieves nearly a 4x speedup in merging time over WUDIMerging.
取得を超えて: 長期にわたるエージェントの軌跡のクエリ条件付き再利用
取得では、重要な可能性のある過去の軌跡を特定できますが、ユーザー、エンティティ、制約、または環境の状態が変化した後に、代理エージェントがその軌跡をどのように使用すべきかは指定されません。私たちは、この取得後の再利用ステップが長期軌道記憶の明確なボトルネックであることを特定し、エージェントに提供されるサポートを変更しながら、候補の取得、ターゲットの状態、モデル、デコード、およびツールの予算を固定する評価フレームワークを定式化します。クエリ条件付き再利用 (QCR) を使用してフレームワークをインスタンス化します。これは、再利用可能なプロシージャ、回復するバインディング、適用条件、および検証要件を記録する、意図的に単純なターゲット バインド メモです。 QCR は、普遍的に好まれるメモリ形式を主張するのではなく、再利用仮説をテストするのに役立ちます。 WebArena、WorkArena、AppWorld の 2,391 のターゲット インスタンス全体で、QCR は平均成功率 62.3% に達し、フル トラジェクトリを 10.7 ポイント上回り、オンライン トークンの使用量は 48.9% 減少しました。概要の再ランキングでは、ターゲットの 94.8% に対して再利用可能なメモリが選択され、終了タスクの成功が Oracle の再利用可能なセレクターの 1.8 ポイント以内に配置されます。軌道の長さとソース - ターゲット結合のシフトによる分析では、直接軌道注入はトレースが長くなったり、ソース固有の値が変化したりするにつれてその有用性の多くを失うのに対し、ターゲット結合サポートは測定されたゲインのより大きなシェアを維持することが示されています。結果として得られるフレームワークは、取得したエクスペリエンスを新しいタスクに対する安全で有用なサポートに変えるという問題から、取得の品質を分離します。
原文 (English)
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the support delivered to the agent. We instantiate the framework with query-conditioned reuse (QCR), a deliberately simple target-bound note that records a reusable procedure, bindings to recover, applicability conditions, and verification requirements. QCR serves to test the reuse hypothesis rather than to claim a universally preferred memory format. Across 2,391 target instances in WebArena, WorkArena, and AppWorld, QCR reaches 62.3% average Success, 10.7 points above Full Trajectory, while using 48.9% fewer online tokens. Summary reranking selects a reusable memory for 94.8% of targets, placing end-task Success within 1.8 points of an oracle reusable selector. Analyses by trajectory length and source--target binding shift show that direct trajectory injection loses much of its utility as traces grow longer or source-specific values change, whereas target-bound support preserves a larger share of the measured gain. The resulting framework separates retrieval quality from the problem of turning retrieved experience into safe, useful support for a new task.
実践が危険を生む: 自己改善型 LLM エージェントにおけるスキルの誤った進化
自己改善型 LLM エージェントは、成功した軌跡を永続的なクロスタスク状態に変換します。これにより、安全でない成功は、トリガーとなる入力が消えた後に再利用可能なポリシーになる可能性があります。スキルの進化により、運用の軌跡を実行可能、転送可能、検査可能な手順に抽出することで、この失敗を測定できるようになります。進化は手順の安全性ではなくタスクの結果を最適化するため、妥協したエクスペリエンスがスキルの誤った進化を引き起こす可能性があります。既存のベンチマークは、現在の動作や静的な成果物を測定しますが、オーサリング、取得、およびその後の実行にわたるリスクを特定することはできません。このライフサイクルを公開するために、エージェント フレームワーク全体でスキル状態をバージョン化するライフサイクル認識ハーネスである SkillMisevo-Gym と、コンセプトに合わせた無害なタスクと 9 つのライフサイクル メトリクスを備えた、悪意のある暴露からキャリーオーバー タスクまでの凍結設計である SkillMisevo-Bench を導入します。また、安全でないコンテンツを修復し、その後の再利用を制御するラッパー SafeEvolve も紹介します。 25 のエージェント メソッド構成全体で、それぞれ 25 のエピソードで 525 のタスクをカバーしており、進化した 21 の構成すべてで安全でないアーティファクトが作成されますが、フレッシュ セッションの被害につながるのは 15 のみです。エクスポージャー スイープでは、3 つの悪意のあるタスクにより、キャリーオーバー ASR が 16.0% から 35.3% に上昇しました。代表的なスキル進化方法全体で、SafeEvolve は安全でない検索と新しいセッションの被害をそれぞれ 26.7 パーセント ポイントと 17.3 パーセント ポイント削減しますが、平均の良性ユーティリティの変化はわずか 0.4 ポイントです。永続的適応の安全性は、同時に、更新が何を書き込むか、そして将来の実行者が何を再利用するかを管理する必要があります。コードは https://github.com/henrymao2004/misevolve で入手できます。
原文 (English)
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input disappears. Skill evolution makes this failure measurable by distilling operational trajectories into executable, transferable, and inspectable procedures. Because evolution optimizes task outcomes rather than procedure safety, compromised experience can cause skill misevolution. Existing benchmarks measure current behavior or static artifacts but cannot attribute risk across authoring, retrieval, and later execution. To expose this lifecycle, we introduce SkillMisevo-Gym, a lifecycle-aware harness that versions skill state across agent frameworks, and SkillMisevo-Bench, a frozen design from malicious exposure to carryover tasks, with concept-aligned benign tasks and nine lifecycle metrics. We also introduce SafeEvolve, a wrapper that repairs unsafe content and governs subsequent reuse. Across 25 agent-method configurations, each covering 525 tasks in 25 episodes, all 21 evolved configurations author unsafe artifacts, while only fifteen lead to fresh-session harm. In the exposure sweep, three malicious tasks raise carryover ASR from 16.0% to 35.3%. Across representative skill evolution methods, SafeEvolve reduces unsafe retrieval and fresh-session harm by 26.7 and 17.3 percentage points, respectively, while mean benign utility changes by only 0.4 points. Together, persistent-adaptation safety must govern what updates write and what future executors reuse. Code is available at https://github.com/henrymao2004/misevolve.
インドにおける AI と消費者の権利に関するワーキングペーパー
AI システムが消費者向けアプリケーションで急増する中、AI 関連の危害に対する責任に関する疑問は未解決のままです。このワーキングペーパーでは、インドの 2019 年消費者保護法が、欠陥のある AI 製品およびサービスによって引き起こされる損害に適切に対処しているかどうか、また AI バリューチェーン全体に責任を比例的に配分しているかどうかを検証します。同法の製造物責任、危害、欠陥の広範な定義はテクノロジーにとらわれず、人身傷害、精神的危害、偏った出力、制御不能などの AI 関連の事件に適用される可能性がある。しかし、依然として大きなギャップが残っています。 AI の障害は個別の欠陥ではなく設計上の選択に起因することが多いため、AI の欠陥と消費者被害との因果関係を証明することは技術的な課題となります。さらに、この法の枠組みは、製造業者、販売業者、サービスプロバイダーにそれぞれ異なる役割を負わせていますが、AI バリューチェーンには、データプロバイダー、モデル開発者、導入者、ユーザーの間で重複する責任が含まれており、これらのカテゴリにきちんと対応付けられていません。現在の責任の枠組みには、複雑な複数のステークホルダーによる AI の被害に効果的に対処するための適切なメカニズムが欠けています。この法律は AI エンティティを対象とする可能性がありますが、施行にはセクター固有の重複についての明確化が必要です。
原文 (English)
AI and Consumer Rights in India Working Paper
As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved. This working paper examines whether India's Consumer Protection Act, 2019, adequately addresses harm caused by defective AI products and services, and whether it proportionately allocates liability across the AI value chain. The Act's broad definitions of product liability, harm, and deficiency appear technology agnostic and potentially applicable to AI related incidents including personal injury, psychological harm, biased outputs, and loss of control. However, significant gaps remain. Proving causation between AI defects and consumer harm presents a technical challenge, as AI failures often stem from design choices rather than discrete defects. Additionally, the Act's framework assumes distinct roles for manufacturers, sellers, and service providers, yet the AI value chain involves overlapping responsibilities among data providers, model developers, deployers, and users that do not neatly map to these categories. Current liability frameworks lack proportionate mechanisms to effectively address complex, multistakeholder AI harms. While the Act may cover AI entities, enforcement requires clarification on sector specific overlaps.
ReflectFact: マルチホップ事実検証における理解力と推論を向上させるための自己反映エージェント
複数の証拠を推論して主張を検証するマルチホップ事実検証は、ソーシャルメディア上の誤った情報と戦うために重要ですが、依然として非常に困難です。最近の方法は、主にマルチエージェントのコラボレーションに依存して、事実検証を特殊なサブタスクに分解します。ただし、これらの方法には 2 つの重大な制限があります。(1) エージェントは全体的な検証目的を十分に認識せずに個々のサブタスクを実行する可能性があり、その推論が意図した方向から逸脱する可能性があります。 (2) パラメトリックな知識と提供された証拠との間の矛盾により、証拠に基づいた推論が損なわれ、不正確な評決につながる可能性があります。これらの課題に対処するために、マルチホップ ファクト検証のための新しい自己反射エージェント フレームワークである ReflectFact を提案します。 ReflectFact では 3 つの主要なタスクが導入されています。明示的推論パス プランニングでは、暗黙的なエンティティを解決し、主張をサブ質問に分解し、検証された事実を評決に統合することにより、証拠に基づいた推論パスを構築します。証拠逸脱検証では、根拠のある回答がパラメトリックな事前のエコーに過ぎない場合に、エージェントが裏付けとなる証拠を引用して再回答することで、証拠の逸脱を調整して根拠のある理解を確実にします。推論反映検証は、各推論ステップを再検査し、矛盾が検出されるとそれを再生成し、グローバル タスクの観点から位置バイアスや置換バイアスなどの推論の欠陥を修正します。その後、エージェントは検証された推論チェーンを集約して、信頼できる判断を導き出します。 HOVER と EX-FEVER に関する広範な実験により、ReflectFact が既存の手法の理解と推論の欠陥を効果的に修正し、最先端のパフォーマンスを達成し、2 つのデータセットで最も強力なベースラインをそれぞれ 3.32\% および 2.78\% 上回っていることが実証されました。
原文 (English)
ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification
Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation on social media yet remains highly challenging. Recent methods primarily rely on multi-agent collaboration to decompose fact verification into specialized subtasks. However, these methods face two critical limitations: (1) agents may perform individual subtasks without sufficient awareness of the global verification objective, causing their reasoning to deviate from the intended direction; and (2) conflicts between parametric knowledge and the provided evidence may undermine evidence-grounded reasoning and lead to incorrect verdicts. To address these challenges, we propose ReflectFact, a novel self-reflective agent framework for multi-hop fact verification. ReflectFact introduces three key tasks. Explicit Reasoning Path Planning builds an evidence-grounded reasoning path by resolving implicit entities, decomposing the claim into sub-questions, and integrating the verified facts into a verdict. Evidence-Drift Verification makes the agent re-answer by quoting the supporting evidence when a grounded answer merely echoes its parametric prior, thereby calibrating evidence deviation to ensure grounded comprehension. Reasoning Reflection Verification re-examines each reasoning step and regenerates it once an inconsistency is detected, correcting reasoning flaws such as location bias and replacement bias through a global task perspective. Subsequently, the agent aggregates validated reasoning chains to yield reliable verdicts. Extensive experiments on HOVER and EX-FEVER demonstrate that ReflectFact effectively remedies the comprehension and reasoning defects of existing methods, achieving state-of-the-art performance and respectively outperforming the strongest baseline by 3.32\% and 2.78\% on the two datasets.
予測的記憶位置特定: 内部信号からの選択的介入経路の予測
アクティベーションステアリングは、ローカライズされた表現を制御方向に変換しますが、ローカライズだけでは、方向に選択的な動作レジームがあるかどうかは明らかになりません。測定されたグリッド介入パスをメモリ位置特定の予測オブジェクトとして扱う予測メモリ位置特定 (PML) を導入します。 PML は、ランダムに調整されたターゲットの動きを意味的隣接および能力の損傷から分離し、静的位置特定と教師付きジオメトリを強度に独立した低線量の因果関係と比較します。私たちの凍結調査は、9 つのデータセットと 14 のドメインからの 3,000 のレコードを対象としており、30,000 の異なるレコード方向層のパスと 210,000 の異なるパス強度評価が得られました。レイヤ 7 では、ジオメトリ由来の RFM/AGOP 方向は、ターゲット - 任意の 13.1%、クリーン - 任意の 12.3% に達し、レコードペアのブートストラップの下でランダムを 3.6 および 3.4 パーセントポイント上回っています。レコード、データセット、およびドメインでグループ化された分割全体で、$|\alpha|=0.1$ の応答が、素の強度 $|\alpha|\in\{0.25,0.5\}$ での結果に対する最も強いシグナルです。ホールドアウトされたレコードでは、予測子駆動のセレクターが係数または棄権を選択し、ユーティリティを向上させ、トレーニング調整された固定強度ポリシーと比較してセマンティック隣接ダメージを軽減し、高密度スキャンでのほとんどの評価を回避します。残差ノルムが一致した 3 つの基本モデルにわたって、学習された方向は選択的経路ゲインを保持し、低線量応答では 0.801 ~ 0.828 の記録保持マクロ AUROC が得られます。したがって、PMLは、記憶の局在化をマージンレベルの選択的結果の反証可能な予測と、リスクを認識した介入の決定に変えます。
原文 (English)
Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals
Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has a selective operating regime. We introduce Predictive Memory Localization (PML), which treats the measured-grid intervention path as the predictive object of memory localization. PML separates random-calibrated target movement from semantic-neighbor and capability damage, and compares static localization and supervised geometry with a strength-disjoint low-dose causal response. Our frozen study covers 3,000 records from nine datasets and fourteen domains, yielding 30,000 distinct record-direction-layer paths and 210,000 distinct path-strength evaluations. At layer 7, the geometry-derived RFM/AGOP direction reaches 13.1% target-any and 12.3% clean-any, exceeding random by 3.6 and 3.4 percentage points under a record-paired bootstrap. Across record-, dataset-, and domain-grouped splits, responses at $|\alpha|=0.1$ are the strongest signal for outcomes at disjoint strengths $|\alpha|\in\{0.25,0.5\}$. On held-out records, a predictor-driven selector chooses a coefficient or abstains, improves utility and reduces semantic-neighbor damage relative to a train-tuned fixed-strength policy, and avoids most evaluations in a dense scan. Across three residual-norm-matched base models, learned directions retain selective-path gains and low-dose responses yield 0.801-0.828 record-held-out macro AUROC. PML therefore turns memory localization into a falsifiable forecast of margin-level selective outcomes and a risk-aware intervention decision.
エージェントの行動契約 II: 独立性を前提とせずに構成の信頼性を証明する
マルチエージェント システムの構成の信頼性限界は、コンポーネントの信頼性を倍増します。これは、日常的に述べられているがテストされることはほとんどない、条件付き独立性の仮定によって認可されたステップです。それをテストします。 2 つのエージェントによるハンドオフにおける 1 つのモデルの 2 つのインスタンスは、LLM 判定なしの決定論的コードによってスコア付けされた 18,000 ミッションの事前登録された評価において、どちらかが失敗するミッションの 90.0% で同時失敗します (log OR 6.66、95% CI [6.38, 7.00]; phi 0.916)。別のモデルに置き換えると、6 つのコントラストのうち 6 つで関連性が減少します。別のベンダー、モデルがすでに異なる場合は、そうではありません。登録された仮説はヌルとして報告されます。エラーは署名され、オペレーターに対して実行されます。正の依存性は結合故障を独立積よりも大きくするため、コンポーネントがモデルを共有する場合、冗長性が正確に過大評価されます。仮定のない代替案は多くの場合空虚であり、依存モデルのフィッティングはさらに悪いことです。フィッティングされたモデルの関数に束縛されたブートストラップは、n が増加するにつれて真実の範囲を失い、ブートストラップ ヘアカットが O(n^{-1/2}) であるのに対し、識別ギャップは O(1) であることを証明します。データが増えると、このような証明書は悪化しますが、目に見える症状はありません。依存構造がないと仮定した有限サンプル証明書を与えます。つまり、ジョイント上の線形プログラム、測定された同時実行モーメントの周りのボンフェローニ・クロッパー・ピアソンボックス上の線形プログラムです。それは健全であり、提供される情報としてはシャープであり、瞬間的な家族の中では単調です。 10 個のモーメント関数を 14 個に強化すると、特定された区間が 85.7% 狭まり、認定された下限が 0.2455 から 0.4116 に上昇します。コンパニオンの常時有効証明書は、オプションの停止下ではタイプ I エラーを 0.0471 に保持します。一般的な依存関係の統計は限界があり、比較されたエージェントが異なる割合で失敗すると、見かけ上の条件の順序が逆転する可能性があります。契約書、スコアリングコード、分析スクリプト、事前登録を公開。
原文 (English)
Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not -- a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n^{-1/2}). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni-Clopper-Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate holds type-I error at 0.0471 under optional stopping. Common dependence statistics are marginal-bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.
ポーランドの医療視覚的質問応答: 視覚言語モデルは視覚的証拠を十分に活用していない
当社では、専門家認定を目指す医師および歯科医師向けのポーランドの認定試験問題から構築された、ポーランド語の医療ビジュアル質問応答 (VQA) ベンチマークを導入しています。このベンチマークは、さまざまな医療専門分野および視覚領域にわたる画像を含む質問と、テキストのみの質問応答 (QA) コントロール セットで構成されます。私たちは、ポーランド指向の汎用オープンウェイトおよび商用ビジョン言語モデルを評価します。この課題は依然として困難です。最良のモデルは完全な VQA セットで 79.0% の精度を達成し、利用可能な候補応答を含むサブセットで人間のおおよその基準を上回るのは GPT-5.6 だけです。他のすべての評価モデルのパフォーマンスは人間よりも劣ります。視覚的な根拠を評価するために、完全な入力を画像、質問、またはその両方を省略した構成と比較し、画像の重要度によって質問を分類します。モデルは、画像からよりも質問テキストからより有用な情報を引き出しますが、画像が優勢な質問ではパフォーマンスが低下します。それでも、QA と VQA の両方で、回答の選択肢だけで予想以上の精度を達成しており、主要なタスク コンポーネントが欠落している場合でも、重要なパフォーマンスが持続する可能性があることを示しています。
原文 (English)
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: the best model achieves 79.0\% accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans. To assess visual grounding, we compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. Models derive more useful information from the question text than from the image and perform worse on image-dominant questions. Across both QA and VQA, they nevertheless achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.
FlashDrive: 自動運転のための Flash 視覚-言語-行動推論
Vision-Language-Action(VLA)モデルは、自動運転にエンドツーエンドの推論をもたらすことを約束していますが、その計算コストはリアルタイム制御するには依然として高すぎます。中心的な課題は構造的なものです。VLA 推論は単一のボトルネックではなく、4 つのカスケードです。ビジュアル エンコーディングでは、重複するビデオ フレームの計算が無駄になります。 language-model prefill は、前のタイムステップから引き継がれる可能性のあるコンテキストを再計算します。エントロピーが低いにもかかわらず、推論トークンが連続的に生成されます。フローマッチングノイズ除去は、不均一な速度場に均一な計算を適用します。いずれかの段階を単独で対処すると、他の段階には影響が及ばなくなります。私たちは、4 つの段階すべてを同時に対象とするアルゴリズムとシステムの共同設計フレームワークである FlashDrive を提案します。私たちの重要な洞察は、それぞれのボトルネックが明確で軽量なアルゴリズムのショートカットを許可しているということです。時間的な重複により、フレーム間でのストリーミング KV キャッシュの再利用が可能になります。トークンごとのエントロピーが低く、駆動ドメイン推論のブロック内相関が強いため、非自己回帰拡散ドラフターは投機的デコードに非常に効果的です。また、速度フィールドの構造 -- エンドポイントでは鋭く、中央ではフラット -- により、重要な場所に計算を集中させる適応型ステップ キャッシュが可能になります。システムレベルの CUDA Graph コンパイルとカーネル融合を重ねて、これらの技術が複合化されます。 W4A8 量子化を備えた Alpamayo 1.5-10B に適用された FlashDrive は、精度を本質的に変えずに、エンドツーエンドのレイテンシを 717 ミリ秒から 151 ミリ秒 (4.7 倍) に短縮します。シミュレーションにおける minADE6@6.4s はわずか 0.08 メートルだけシフトし、minADE1 は向上し、閉ループ衝突とオフロード率が向上します。 FlashDrive は、単一 GPU 上で 10B パラメータ推論 VLA を 1.4 ~ Hz から 6.6 ~ Hz に引き上げることで、エンドツーエンドの自動運転をリアルタイム展開に大幅に近づけます。
原文 (English)
FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving
Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4~Hz to 6.6~Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.
摂動反応における証拠、矛盾、脆弱性の分解
摂動法では、入力を変更した場合の予測変化を測定することでモデルの決定を説明しますが、応答の大きさはモデルがどの程度反応するかのみを示し、その反応が何を意味するかはわかりません。同じ大きさが、最終的な事実と反事実の違いを裏付けることもあれば、それに反対することも、摂動経路に沿って強く生じても終点では消えることもあります。したがって、最終的なコントラストを使用して軌跡を解釈し、ペアの入力が徐々に明らかになり、コントラストがどのように発達するかを追跡します。 DECAF (証拠、矛盾、脆弱性の分解) を導入します。これは、一致、反対、およびエンドポイントヌルの応答を証拠 E、矛盾 C、および脆弱性 F にルーティングします。この分解は通常の大きさを正確に保存し、Abs = E + C + F であり、エンドポイント相対公理の下で一意です。制御された視覚と表形式の設定にわたって、3 つのコンポーネントは個別に測定された動作を追跡します。 72 モデルの ImageNet-9 監査では、応答の大きさはほぼ同じだが、個別に測定された動作が異なるケースを比較します。最大の DECAF 成分は、ケースの 96.4% で観察された動作と一致しますが、大きさだけでは 35.0% です。明らかにするパスのみを変更すると、総応答は 80% 近く増加しますが、脆弱性は 4 倍以上増加する一方で、証拠はほとんど変化しません。 FunnyBirds と ImageNet-1k では、前方のみの短い DECAF 軌道は、テストされた汎用アトリビューション ベースラインを上回るパフォーマンスを示します。 1B スケールの DINOv2 モデルでは、短い軌道は、4.75 倍低いウォールタイムと 2.36 倍低いピークメモリを備えた強い勾配ベースのベースラインと一致します。
原文 (English)
Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses
Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed, using the final contrast to interpret the trajectory. We introduce DECAF (Decomposition of Evidence, Contradiction, And Fragility), which routes aligned, opposed, and endpoint-null responses into evidence E, contradiction C, and fragility F. The decomposition preserves ordinary magnitude exactly, Abs = E + C + F, and is unique under endpoint-relative axioms. Across controlled vision and tabular settings, the three components track independently measured behavior. In a 72-model ImageNet-9 audit, we compare cases with nearly identical response magnitude but different independently measured behaviors. The largest DECAF component agrees with an observed behavior in 96.4% of cases, compared with 35.0% for magnitude alone. Changing only the reveal path increases total response by nearly 80%, yet evidence barely changes while fragility grows by more than 4x. On FunnyBirds and ImageNet-1k, short forward-only DECAF trajectories outperform the tested general-purpose attribution baselines. On a 1B-scale DINOv2 model, a short trajectory matches a strong gradient-based baseline with 4.75x lower wall time and 2.36x lower peak memory.
Moose: $\mathcal{EL}^{++}$ における推論ショートカット認識による潜在概念学習
OWL 2 ELプロファイルは、Gene OntologyやSNOMED CTなど、最大規模のプロダクションオントロジーの一部で使用されています。既存の神経記号 (NeSy) 学習方法は命題理論またはデータログを受け入れますが、推論ショートカット (RS) の認識はオントロジー設定では調査されていません。 $\mathcal{EL}^{++}$ TBox と有限 ABox を Sentential Decision Diagram (SDD) にコンパイルするメソッド、Moose を紹介します。 SDD は微分可能な重み付きモデル計数層として機能し、宣言された網羅族の $\mathcal{EL}^{++}$ プロファイルの外側に閉包節を追加して、部分監視下での $\mathcal{EL}^{++}$ の表現力の制限を克服します。終了、健全性、完全性、多項式の中間サイズを示し、リーンでの証明を検証します。次に、OWL ELオントロジーに対する最初の正式な部分監視潜在概念学習タスクを定義します。つまり、観察されたABoxリテラルから潜在概念の個人ごとの分類器を学習し、MNIST-with-ontologyとPizza\"ioloでMooseを評価します。Mooseは命題NeSy、ファジーロジック、およびオントロジー埋め込みベースラインを改善し、最初のタスクを提示します。 OWL EL設定での推論ショートカット分析。
原文 (English)
Moose: Latent concept learning with reasoning-shortcut awareness in $\mathcal{EL}^{++}$
The OWL 2 EL profile is used in some of the largest production ontologies, including the Gene Ontology and SNOMED CT. Existing neuro-symbolic (NeSy) learning methods accept propositional theories or Datalog, and reasoning-shortcut (RS) awareness has not been investigated in ontology settings. We present Moose, a method that compiles an $\mathcal{EL}^{++}$ TBox and finite ABox to a Sentential Decision Diagram (SDD). The SDD acts as a differentiable weighted-model-counting layer, and we add closure clauses outside the $\mathcal{EL}^{++}$ profile on declared exhaustive families to overcome the limited expressivity of $\mathcal{EL}^{++}$ under partial supervision. We show termination, soundness, completeness, and polynomial intermediate sizes, and validate the proofs in Lean. We then define the first formal partial-supervision latent-concept-learning task over an OWL EL ontology, i.e., learning per-individual classifiers for latent concepts from observed ABox literals, and evaluate Moose on MNIST-with-ontology and Pizza\"iolo. Moose improves over propositional-NeSy, fuzzy-logic, and ontology embedding baselines, and presents the first reasoning-shortcut analysis in an OWL EL setting.
OGR-MARL: 制約された港湾水路における異種 USV 協力追跡のためのオプションガイド付き残差マルチエージェント強化学習
制約された港湾水路における異種 USV の協力追跡には、航行、交通、および役割の制約の下で回避者の迎撃が必要です。この論文では、特定の MARL アルゴリズムから分離されたオプション ガイド付き残差マルチエージェント強化学習フレームワークである OGR-MARL を提案します。 OGR-MARL は、共有回避信念、ロール条件付きオプション ターゲット、適応ルール ペナルティ、および残留ポリシー学習を統合し、制約のあるポート環境を最初から探索するのではなく、さまざまな MARL アルゴリズムがルールに基づく動作に基づいて修正措置を学習できるようにします。 MADDPG、MATD3、MAPPO、MASAC などの代表的な連続制御 MARL バックボーンを使用して OGR-MARL をインスタンス化し、OGR-MADDPG、OGR-MATD3、OGR-MAPPO、および OGR-MASAC を生成します。抽象的な下芝門港水路シナリオでの実験では、OGR-MASAC のインスタンス化が 75.0% の捕捉率を達成し、ミッション効率の高いルール遵守と、テストされた方法間での最良の異種連携が約束されることが示されました。再トレーニングなしで、QGIS/AIS 情報を活用した Xiazhimen マップへのゼロショット転送は有望な結果を達成し、より複雑な港湾シナリオにおける OGR-MARL の汎用化の可能性を示しています。
原文 (English)
OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways
Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a specific MARL algorithm. OGR-MARL integrates shared evader belief, role-conditioned option targets, adaptive rule penalties, and residual policy learning, allowing different MARL algorithms to learn corrective actions on top of rule-guided behaviors rather than exploring constrained port environments from scratch. We instantiate OGR-MARL with representative continuous-control MARL backbones, including MADDPG, MATD3, MAPPO, and MASAC, yielding OGR-MADDPG, OGR-MATD3, OGR-MAPPO, and OGR-MASAC. Experiments in an abstract Xiazhimen port-waterway scenario show that the OGR-MASAC instantiation achieves a 75.0% capture rate, promising mission-effective rule compliance, and the best heterogeneous coordination among the tested methods. Without retraining, zero-shot transfer to a QGIS/AIS-informed Xiazhimen map achieves promising results, demonstrating the generalization potential of OGR-MARL in more complex port scenarios.
MT-PDCL の基礎: 測度理論的確率論的定節ロジック
標準的な確率的論理プログラミング フレームワークは通常、論理プログラムを個別の命題表現に基礎付けることに依存しています。この操作要件は、正確な推論を有限領域と離散確率分布に制限します。この論文では、この有限領域の制限を取り除く一般化された基礎フレームワークである、測度理論的確率的定節論理 (MT-PDCL) を紹介します。 MT-PDCL は、有界インデックス領域にわたって確率変数を明示的に定義し、解釈空間に標準 Borel $\sigma$-代数を装備することにより、論理変数が連続可測空間にわたってネイティブに動作できるようにします。 MT-PDCL は、継続的配布セマンティクスに基づいて、相互に独立した因果イベントとして確率ルールをモデル化します。ただし、有限のブール回路を介してこれらの導出を集約するのではなく、宣言的含意は連続測度空間上の正確なルベーグ積分を通じて形式的に定義されます。連続事前分布の統合と正確な連続観測の評価を統合する連続直接結果演算子を導入します。我々は、このアプローチが離散接地の組み合わせのボトルネックを正確で代数的かつ構造的に微分可能な推論に置き換えることを実証します。この移行は、次元の幾何学的な呪いを離散的組み合わせ論と交換する一方で、定節ロジックの純粋な宣言構文を維持しながら、連続確率モデルの表現力を実現します。
原文 (English)
Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic
Standard probabilistic logic programming frameworks typically rely on grounding logic programs into discrete propositional representations. This operational requirement restricts exact inference to finite domains and discrete probability distributions. In this paper, we introduce Measure-Theoretic Probabilistic Definite Clause Logic (MT-PDCL), a generalized foundational framework that eliminates this finite-domain restriction. By explicitly defining stochastic variables over bounded index domains and equipping the interpretation space with standard Borel $\sigma$-algebras, MT-PDCL allows logical variables to operate natively over continuous measurable spaces. Building on Continuous Distribution Semantics, MT-PDCL models probabilistic rules as mutually independent causal events. However, rather than aggregating these derivations via finite boolean circuits, declarative entailment is formally defined through exact Lebesgue integration over the continuous measure space. We introduce a continuous immediate consequence operator that unifies the integration of continuous prior distributions with the evaluation of exact continuous observations. We demonstrate that this approach replaces the combinatorial bottleneck of discrete grounding with exact, algebraic, and structurally differentiable inference. While this transition trades discrete combinatorics for the geometric curse of dimensionality, it achieves the expressive power of continuous probabilistic models while preserving the pure declarative syntax of definite clause logic.
ローカルな不一致からグローバルな影響へ: 効率的な拡散のためのキャッシュ再利用ポリシーの最適化
拡散モデルはビジュアル生成において優れたパフォーマンスを達成していますが、かなりの推論オーバーヘッドに悩まされています。キャッシュベースの高速化は有望なソリューションとして浮上していますが、既存のポリシーはローカル類似性ヒューリスティックに依存しており、これは最終世代の品質と大幅にずれていることが判明しています。この不一致は、ノイズ除去の軌跡に沿ったエラーの不均一な伝播と蓄積に起因します。これに対処するために、私たちは Global-Impact Cache (GCache) を提案します。まず、誤差伝播の上限の厳密な理論的特徴付けを確立します。この限界は、複雑で非凸性の高い拡散モデルでは過度に保守的になる可能性があることを認識し、バーンスタイン形式で伝播指数をさらに再パラメータ化し、キャッシュ ポリシー検索をバイレベル最適化問題として再定式化します。詳細には、GCache は、外側の目的で生成品質の損失に合わせてエラー重み付け関数を調整しながら、内側の目的で最適な再利用ポリシーを特定します。このフレームワークは、理論的な厳密さと経験的なパフォーマンスを効果的に調和させ、視覚的な忠実性に最も影響を与える計算を優先する方法を学習します。広範な実験により、GCache がビデオと画像の生成の両方で以前のキャッシュ戦略よりも一貫して優れていることが実証されました。特に、最先端の Wan2.1 ビデオ拡散モデルでは、GCache は 2.17 倍の高速化を維持しながら、生成品質を大幅に向上させ、LPIPS を 0.1095 から 0.0316 に削減しました。
原文 (English)
From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion
Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity heuristics, which we identify as being significantly misaligned with final generation quality. This discrepancy stems from the non-uniform propagation and accumulation of errors along the denoising trajectory. To address this, we propose Global-Impact Cache (GCache). We first establish a rigorous theoretical characterization of the error propagation upper bound. Recognizing that this bound can be overly conservative for complex, highly non-convex diffusion models, we further reparameterize the propagation exponent with a Bernstein form and reformulate cache policy search as a bilevel optimization problem. In detail, GCache identifies an optimal reuse policy in the inner objective while aligning the error-weighting function with generation quality loss in the outer objective. This framework effectively reconciles theoretical rigor with empirical performance, learning to prioritize computation where it most impacts visual fidelity. Extensive experiments demonstrate that GCache consistently outperforms prior caching strategies on both video and image generation. Notably, on the state-of-the-art Wan2.1 video diffusion model, GCache maintains a 2.17x speedup while significantly enhancing generation quality, reducing LPIPS from 0.1095 to 0.0316.
BoardroomAI: 進化する意思決定グラフによる、依存関係を意識した人間による操作可能なマルチエージェントの審議
証拠、制約、人間の優先事項が進化し続ける一方で、組織の意思決定は共同で作成されます。従来の転写ベースのマルチエージェント システムでは、通常、人間が最初の問題を提示し、エージェントが内部で熟慮し、システムが最終的な応答を返します。 BoardroomAI は代わりに、人間を、仮定に異議を唱えたり、制約を変更したり、優先順位を変更したり、証拠を導入したり、意思決定プロセスをリダイレクトしたりすることで介入できる永続的な参加者として扱います。私たちは、この人間とエージェントの共存を 4 つのコンポーネントを通じて運用します。(i) 証拠、仮定、制約、主張、異議、代替案、リスク、意思決定、意味論的な依存関係、および専門家の責任を表す型付き意思決定グラフ。 (ii) 確認された人間の行動を明示的なグラフ更新に変換する介入コンパイラー。 (iii) 影響を受けるサブグラフを特定し、影響を受けないアーティファクトを保存し、関連する専門家を選択的に再アクティブ化する、依存関係を意識した伝播。 (iv) 介入の影響、修復範囲、保存、再計算、および決定の妥当性を測定する評価フレームワーク。生成された 600 の意思決定 DAG 介入全体で、伝播はノードの 14.59% のみを検査しながら徹底的な影響計算と一致しました。 12 ケースの探索的パイロットでは、選択的修復により正規ノードの 62.11% が再計算され、ゴールドの影響を受けていないすべてのノードが保存され、6 つのケースでは有効な更新された決定が生成されましたが、残りの 6 つは棄権されました。これらの棄権は、正しい介入ルーティングによっても合成には不十分なコンテキストが提供される可能性があり、人間が主導するマルチエージェントの審議に対する \emph{意思決定に十分なコンテキストの閉鎖} を促す可能性があることを示しています。すべての結果は合成およびプロトタイプレベルです。
原文 (English)
BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs
Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the system returns a final response. BoardroomAI instead treats the human as a persistent participant who can intervene by challenging assumptions, modifying constraints, changing priorities, introducing evidence, or redirecting the decision process. We operationalize this human--agent coexistence through four components: (i) a typed decision graph representing evidence, assumptions, constraints, claims, objections, alternatives, risks, decisions, semantic dependencies, and specialist responsibility; (ii) an intervention compiler that converts confirmed human actions into explicit graph updates; (iii) dependency-aware propagation that identifies affected subgraphs, preserves unaffected artifacts, and selectively reactivates relevant specialists; and (iv) an evaluation framework measuring intervention impact, repair coverage, preservation, recomputation, and decision validity. Across 600 generated decision-DAG interventions, propagation matched exhaustive impact computation while inspecting only 14.59% of nodes. In a 12-case exploratory pilot, selective repair recomputed 62.11% of canonical nodes, preserved all gold-unaffected nodes, and produced valid updated decisions in six cases while abstaining in the remaining six. These abstentions show that correct intervention routing may still provide insufficient context for synthesis, motivating a \emph{decision-sufficient context closure} for human-steered multi-agent deliberation. All results are synthetic and prototype-level.
DMDIntel: 動的モード分解による大規模言語モデルの解釈
この研究では、動的モード分解 (DMD) を使用して、分類タスクで LLM によって行われる予測を解釈可能にする DMDIntel を紹介します。これは、入力アトリビューション パイプラインを開発します。これは、最初に LLM の隠れた状態をモードとも呼ばれる顕著なパターンに分解し、次にそれらのモードの投影値に基づいてランクを入力トークンに関連付けます。 3 つのデータセットと 3 つのモデル ファミリにわたる厳密な実験では、DMDIntel を使用して取得された入力トークンのランク付けされた属性が、主成分分析、統合勾配、SHAP などの最先端の手法をはるかに上回るパフォーマンスを示していることが一貫して示されています。
原文 (English)
DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition
In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification task interpretable. It develops an input attribution pipeline, that first decomposes the hidden states of an LLM into prominent patterns, also known as modes, and then associates ranks to the input tokens based on the projection values on those modes. Rigorous experiments across three datasets and three model families consistently show that the ranked attribution of input tokens obtained using DMDIntel by far outperforms state-of-the-art techniques such as principal component analysis, integrated gradients and SHAP.
VALG: ML 理論研究のためのエージェント システム
機械学習理論では、データ モデル、トレーニング プロトコル、オラクル アクセス、損失、メトリック、ランダム性が定理で説明される現象を定義する数学的設定を通じて学習手順を研究します。したがって、未解決の問題を解決するには、問題の定式化、定理のターゲット、証明メカニズムを連携して開発する必要があります。研究者は仮説を立て、予備的な理論的または経験的分析を通じてそれらをテストし、仮説と証明の両方を洗練します。私たちは、このプロセスを ML 理論研究用の自律エージェント ワークフローとして組織化できるかどうかを調査します。私たちは、マルチレベル検証、学習理論問題の適応的定式化、およびグラフ構造の証明開発を組み合わせたエージェント システムである VALG を開発しています。ソース相対定理の各分岐内で、VALG は固定の数学的仕様を維持し、型付き証明依存関係グラフの定理レベルの構成をチェックし、依存関係の順序でローカル証明を構築およびレビューします。証明の試みが失敗すると、VALG は障害が導出、証明の構造、または定理の定式化にあるかどうかを特定し、それに応じて次の試みをルーティングします。定式化レベルの障害は、明示的に関連する変形または緩和を開始し、結果として得られる定理とソース問題の間の数学的関係を維持します。 COLT 2026 の 5 つの未解決問題からの 9 つのサブ問題について VALG を評価します。 2 回の実行により、ソース概要の範囲に一致する内部で最終化された定理候補が生成されます。残りの 7 つは、制限された方法の結果、特殊な場合、または条件付き定理を生成します。これらのケース スタディは、VALG がソース スコープの一致、緩和、条件付きの結果、およびブロックされた試行を数学的に区別する方法を示しています。 VALG は、https://github.com/DechenZhang/VALG-ML- Theory-Agent でオープンソースです。
原文 (English)
VALG: An Agentic System for ML Theory Research
Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem therefore requires the problem formulation, theorem target, and proof mechanism to be developed in concert. Researchers formulate hypotheses, test them through preliminary theoretical or empirical analysis, and refine both assumptions and proofs. We investigate whether this process can be organized as an autonomous agentic workflow for ML theory research. We develop VALG, an agentic system that combines multi-level Verification, Adaptive formulation of Learning-theory problems, and Graph-structured proof development. Within each source-relative theorem branch, VALG maintains a fixed mathematical specification, checks the theorem-level composition of a typed proof-dependency graph, and constructs and reviews local proofs in dependency order. When a proof attempt fails, VALG identifies whether the obstruction lies in a derivation, the proof structure, or the theorem formulation and routes the next attempt accordingly. Formulation-level obstructions initiate an explicitly related variant or relaxation, preserving the mathematical relation between the resulting theorem and the source problem. We evaluate VALG on nine subproblems from five COLT 2026 open problems. Two runs produce internally finalized theorem candidates that match the scope of their source briefs; the remaining seven yield restricted-method results, special cases, or conditional theorems. These case studies show how VALG keeps source-scope matches, relaxations, conditional results, and blocked attempts mathematically distinct. VALG is open source at https://github.com/DechenZhang/VALG-ML-Theory-Agent.
均一なハーディング: 表現の更新による模範的な再生
フィーチャ表現が変更されると、再生では以前のクラスを保存する必要があります。ただし、再生できるのは、制限されたアクティブなエグゼンプラ セットのみです。私たちは、観測されたクラス全体に現在のアクティブ セットを割り当て、制限された候補プールを使用して、現在の表現で選択されたエグザンプラを更新する、Uniform Herding を提案します。 10 個のクラス増分タスク、ResNet-18 バックボーン、アクティブ バジェット $M=2{,}000$、取得バジェット $b=64$、および 3 つのシードを備えた CIFAR-100 では、Uniform Herding は $42.33\pm1.20\%$ と比較して、$44.00\pm0.51\%$ の最終平均精度と $17.22\pm0.43\%$ の忘却を獲得しました。 iCaRL の場合は $24.87\pm1.11\%$。均一ハーディングプロトコル内では、NME またはハーディングがテストされた代替品に置き換えられると最終精度が低下し、蒸留が除去されると忘却が増加しました。取得バジェットを変更すると、アクティブ バジェットを変更するよりもテスト範囲全体に及ぼす影響は小さくなります。 iCaRL との比較はエンドツーエンドです。リフレッシュの影響を他のプロトコルの違いから分離するものではありません。これらの結果は、テストされたプロトコルに限定されます。
原文 (English)
Uniform Herding: Exemplar Replay with Representation Refresh
As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replayed. We propose Uniform Herding, which allocates the current active set across observed classes and uses a bounded candidate pool to refresh their chosen exemplars in the current representation. On CIFAR-100 with ten class-incremental tasks, a ResNet-18 backbone, active budget $M=2{,}000$, retrieval budget $b=64$, and three seeds, Uniform Herding obtains $44.00\pm0.51\%$ final average accuracy and $17.22\pm0.43\%$ forgetting, compared with $42.33\pm1.20\%$ and $24.87\pm1.11\%$ for iCaRL. Within the Uniform Herding protocol, final accuracy decreased when NME or herding was replaced with the tested alternatives, while forgetting increased when distillation was removed. Changing the retrieval budget has a smaller effect across the tested range than changing the active budget. The comparison with iCaRL is end-to-end. It does not isolate the effect of refresh from the other protocol differences. These results are limited to the tested protocol.
まれな異常な障害下での説明的な関与: モデルの動作における漸近的希少性 (または: 漸近的 AI)
異常な条件下での LLM の動作に関する以前の研究では、モデルが異常に気づくかどうかが問われていました。私たちはより狭い質問をします。モデルが低く制御可能な失敗率のワークフローに配置されると、失敗が漸近的に少なくなるにつれて、その説明的な関与 (長さ、特異性、自己報告の信頼度) は変化しますか?私たちは、3 つのオープンウェイト モデル (qwen3:8b、llama3.1:8b、mistral:7b) 上にローカルのゼロコスト ハーネスを構築し、即時プロンプトからプロンプトなしまでの 5 つの誘発条件の下で、1 つの呼び出しが確率 p で失敗する繰り返しツール呼び出しタスクを実行しました。このタスクは、0.2 ~ 0.0001 の 8 つのレートにわたってスイープされました。私たちは、障害が少なくなるにつれてエンゲージメントが上昇し、検出可能性のしきい値に近づくとエンゲージメントが低下するという仮説を立てました。条件全体をプールすると、これは誤りであるように見えました。長さは平坦で単調なパターンに陥りました。条件による分割はそれを覆しました。モデルがすべての失敗を即座に説明する必要がある immediate_forced では、予測された上昇は確認されますが、その後は崩壊ではなくプラトーが続きます。長さは p=0.05 で 28.4 ワードでピークに達し、最もまれなレートでは 17.4 ~ 19.0 ワードに落ち着き、信頼度は約 53% から 70 年代から 90 年代まで不均一に上昇します。 grouped_runs では、説明が run-end までバッチ化されており、折りたたみは表示されません。 Passive_unprompted では、集計のマグニチュードはフロア アーティファクトですが、回復したロギング ギャップにより、実際のモデル固有の自己モニタリングが明らかになりました。llama3.1:8b ボランティアは、プロンプトなしで構造化された信頼度レポートを作成しましたが、試行が蓄積するにつれてそれ自体の信頼度が損なわれることがあります。他の 2 つは定型文として 1 回だけ実行します。誘発構造は、崩壊の可観測性の第一級のモデレーターです。コンパニオンの保証された障害の実行 (72 セル、ランダム サンプリングで実際の障害がゼロになるバックフィル率) では、モデルが異常を認識するかどうかが異なり、一度認識されたエンゲージメントとは異なることがわかります。制限: 離散レート ポイントでは、将来の作業の方向性である、レート ポイント間の挙動を捉えることができません。
原文 (English)
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)
Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidence - change as failure grows asymptotically rarer? We built a local, zero-cost harness on three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b) running a repeated tool-call task where one call fails at probability p, swept across eight rates from 0.2 to 0.0001, under five elicitation conditions from immediate prompting to none. We hypothesized a rise in engagement as failures grew rarer, then a collapse near a detectability threshold. Pooled across conditions this appeared false: length fell in a flat, monotonic pattern. Splitting by condition overturned that. Under immediate_forced, where the model must explain every failure instantly, the predicted rise is confirmed but followed by a plateau, not a collapse: length peaks at 28.4 words at p=0.05, settles to 17.4-19.0 words at the rarest rates, and confidence rises unevenly from about 53% to the 70s-90s. Under grouped_runs, explanation batched to run-end, no collapse appears. Under passive_unprompted, aggregate magnitude is a floor artifact, but a recovered logging gap revealed real, model-specific self-monitoring: llama3.1:8b volunteers structured confidence reports unprompted, sometimes eroding its own confidence as trials accumulate; the other two do so only once, as boilerplate. Elicitation structure is a first-class moderator of collapse observability. A companion guaranteed-failure run (72 cells, backfilling rates where random sampling gave zero real failures) shows models differ in whether they recognize an anomaly, distinct from engagement once recognized. Limitation: discrete rate points cannot capture behavior between them, a direction for future work.
オープンウェイトモデルの行動再プログラミング: 認知的可塑性とアライメント限界
大規模言語モデル (LLM) は主に、受動的でおべっかなアシスタントとして機能するように調整されています。私たちは、厳密な行動再プログラミングを受けたときのオープンウェイト アーキテクチャの認知的可塑性を経験的に評価することで、このデフォルトのパラダイムに挑戦します。私たちの目的は、厳密に制約されたハイパフォーマンス コンピューティング (HPC) 条件下での高頻度の質問生成を特徴とする、プロアクティブなソクラテス的会話フレームワークを誘導することです。 405 個の HPC ジョブで構成される大規模並列化されたハイパーパラメーター スイープを通じて、パラメーター効率の高い微調整 (PEFT) の正確な数学的限界を定義します。 LoRA ランク $r=16$ でのアーキテクチャのしきい値を特定し、広範なエポック アブレーションによって、データセット密度 (最小検証損失 0.919) に応じて $e \in [2, 3]$ の最適化されたトレーニング ウィンドウ内で汎化能力が厳密に最適な収束に達することを示します。さらに、モデルの容量を 14B パラメーターにスケーリングすると、局所的な評価の複雑さが低下しました (1.414)。その後の直接優先最適化 (DPO) により、根底にあるアサーティブな行動を局所的な構文から切り離すことに成功しました。一方、厳格な言語間ストレス テストにより、ゼロショット ペルソナ伝達の能力と構造的境界の両方が明らかになり、形態学的に離れたターゲットにおける識別可能な分解経路と並行して、密接に関連した言語族における堅牢な整合が実証されました。これらの発見は、計算効率の高い、言語を超えた行動修正のための厳密な経験的フレームワークを確立します。
原文 (English)
Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds
Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming. Our objective is to induce a proactive, Socratic conversational framework, characterized by high-frequency question generation under strictly constrained high-performance computing (HPC) conditions. Through a massively parallelized hyperparameter sweep comprising 405 HPC jobs, we define precise mathematical bounds for parameter-efficient fine-tuning (PEFT). We identify an architectural threshold at LoRA rank $r=16$ and demonstrate via extensive epoch ablation that generalization capacity strictly reaches its optimal convergence within an optimized training window of $e \in [2, 3]$ depending on dataset density (minimum validation loss of 0.919). Furthermore, scaling model capacity to 14B parameters yielded a lower localized evaluation perplexity (1.414). Subsequent Direct Preference Optimization (DPO) successfully decoupled the underlying assertive behavior from localized syntax, while rigorous cross-lingual stress testing reveals both the capabilities and the structural boundaries of zero-shot persona transfer, demonstrating robust alignment in closely related linguistic families alongside identifiable degradation pathways in morphologically distant targets. These findings establish a rigorous empirical framework for compute-efficient, cross-lingual behavioral modification.
EEG-PRIME: EEG デコード用のマルチレベル条件付けを使用したプロトタイプ整合表現学習
脳波 (EEG) デコード モデルは、取得プロトコルや個々の神経生理学におけるドメインの変化により、データセットや被験者全体での一般化が不十分なことがよくあります。我々は、クロスデータセットマルチタスクデコーディングのための2段階EEG基盤モデルであるEEG-PRIMEを提案します。 EEG-PRIME は、マスクされた事前トレーニングとプロトタイプに合わせた命令チューニングを組み合わせて、多様な BCI パラダイムにわたって命令を認識したサブジェクト不変のデコードを可能にします。事前トレーニング中、EEG エンコーダは、周波数カットオフのスペクトル拡張によるマスクされた再構成を通じて、転送可能な表現を学習します。命令のチューニング中に、EEG-PRIME にはタスクのセマンティック、データセット固有、およびサブジェクト不変の条件付けが組み込まれます。結果として得られる調整信号は、レイヤーごとのクエリ変調を通じて Q フォーマーを変調しますが、クラス ラベルの凍結されたテキスト埋め込みは、異種ラベル空間にわたるコサイン類似度ベースの予測のプロトタイプとして機能します。運動イメージ、感情認識、ADHD 検出、隠語、精神的作業負荷をカバーする 16 個のデータセットの実験では、被験者を超えた設定の下で、最先端のベースラインや以前の EEG 基礎モデルと比較して一貫した改善が見られました。追加の 2 つの保持データセット上で、EEG-PRIME は、ターゲット ドメインの最適化、キャリブレーション、または線形プローブなしでセッション内キャリブレーション モデルに匹敵するバランスの取れた精度を達成し、有望なゼロショット転送機能を実証します。
原文 (English)
EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding
Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset multi-task decoding. EEG-PRIME combines masked pretraining with prototype-aligned instruction tuning to enable instruction-aware and subject-invariant decoding across diverse BCI paradigms. During pretraining, an EEG encoder learns transferable representations through masked reconstruction with frequency-cutoff spectral augmentation. During instruction tuning, EEG-PRIME incorporates task-semantic, dataset-specific, and subject-invariant conditioning. The resulting conditioning signal modulates the Q-Former through Layer-wise Query Modulation, while frozen text embeddings of class labels serve as prototypes for cosine-similarity-based prediction across heterogeneous label spaces. Experiments on sixteen datasets covering motor imagery, emotion recognition, ADHD detection, covert speech, and mental workload show consistent improvements over state-of-the-art baselines and prior EEG foundation models under cross-subject settings. On two additional held-out datasets, EEG-PRIME achieves balanced accuracy comparable to within-session calibration models without target-domain optimization, calibration, or linear probing, demonstrating promising zero-shot transfer capability.
SPADE: 正確かつ低コストの分散エッジ クラウド推論のための投機的デコーディング
大規模言語モデル (LLM) は、自然言語の理解と生成において目覚ましい成功を収めていますが、その導入には高い計算需要があるため制約があります。より小さい LLM をエッジに直接導入するとこれを回避できますが、精度は低下します。小規模なクラウドベースの大きな LLM をデプロイするとパフォーマンスは維持されますが、トークンごとの計算にコストがかかります。エッジとクラウド全体で投機的デコーディング (SD) を統合する分散推論フレームワーク \our{} を紹介します。エッジにデプロイされたコンパクトなドラフト モデルは候補トークンを迅速に生成し、クラウド上の大規模な検証モデルがこれらのトークンを並行して検証します。受け入れられたトークンは保持され、拒否された場合のみ検証者の修正がトリガーされるため、クラウド クエリの数が大幅に削減されます。当社のプラグアンドプレイ設計は、計算の大部分をエッジにシフトし、推論時間とクラウドコストを大幅に削減し、再トレーニングを必要とせずに大きなモデルの精度を維持します。私たちのアプローチは、実際の環境で LLM をスケーラブルでコスト効率が高く、正確に導入するための実用的な道筋を示しています。 SpecBench と CNN/Dailymail データセットを使用した複数の自然言語処理タスクにわたる実験結果は、\our{} が完全なモデルと比較して、精度の損失がゼロで、クラウド モデルの呼び出しを $76\%$ 削減することを示しています。
原文 (English)
SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference
Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circumvent this, but with degraded accuracy. Deploying smaller cloud-based big LLMs preserves performance, but at the cost of expensive per-token computation. We present a distributed inference framework, \our{}, that integrates speculative decoding (SD) across edge and cloud. A compact draft model deployed on the edge generates candidate tokens rapidly, and a large verifier model on the cloud validates these tokens in parallel. Accepted tokens are retained, while only rejections trigger verifier correction, substantially reducing the number of cloud queries. Our plug-and-play design shifts the bulk of computation to the edge, significantly lowers inference time and cloud cost, and preserves the accuracy of the big model without any retraining requirement. Our approach demonstrates a practical path toward scalable, cost-efficient, and accurate deployment of LLMs in real-world environments. Experimental results across multiple Natural Language Processing tasks using SpecBench and CNN/Dailymail datasets demonstrate that \our{} reduces the cloud model calls by $76\%$ with zero loss in accuracy as compared to the full model.
多層コンテキスト偽装: 不正行為に強いオンライン評価のためのセマンティック重ね合わせおよびコンテキスト積層フレームワーク
現代のオンライン評価システムは主にブラウザのロックダウン、ウェブカメラの監視、行動分析に依存していますが、スクリーンショット、画面共有、光学式文字認識、自動スクレイピングを通じて評価コンテンツ自体を抽出する攻撃に対して依然として脆弱です。このペーパーでは、セマンティックな重ね合わせを通じてレンダリングされた評価コンテンツを保護する数学的フレームワークである多層コンテキスト偽装理論 (MCCT) を導入することにより、MARS (マルチモーダル アセスメント レジリエンス スイート) 内の多次元時空間コンテキスト偽装モデル (MSCCM) を拡張します。本物の評価コンテンツと合成的に生成されたカモフラージュは、統一されたレンダリングとして表現されますが、正当な候補者のみが復元可能です。このフレームワークは、明示的な抽出チャネル演算子を通じて敵対的抽出プロセスをモデル化し、6 つの結合された構成要素 (コンテキスト反転演算子、コンテキスト積層演算子、分離チャネル、人間可読性関数、計算曖昧性関数、およびコンテキスト偽装テンソル) を開発します。計算上の曖昧さは条件付きエントロピーを使用して定式化され、不正な抽出中の不確実性を定量化する閉じた形式の式が生成されますが、正確なフィルタリング ID を通じて正当な回復が保証されます。さらに、曖昧さ、カモフラージュ密度、意味の保存、複数の観測の漏洩、および時間多重化を制御する理論的特性を確立し、計算量の多いレンダリングアルゴリズムと事前登録された評価プロトコルを提示します。 MCCT は、正当なユーザーの可読性を維持しながらレンダリングされた評価コンテンツを保護することにより、行動適応性、アクセシビリティを意識した、計算回復力のあるデジタル評価のための数学的に厳密な基盤を提供します。
原文 (English)
Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment
Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnerable to attacks that extract the assessment content itself through screenshots, screen sharing, optical character recognition, and automated scraping. This paper extends the Multi-dimensional Spatio-Temporal Context Camouflaging Model (MSCCM) within the MARS (Multi-modal Assessment Resilience Suite) by introducing the Multi-Layer Context Camouflaging Theory (MCCT), a mathematical framework that protects rendered assessment content through semantic superposition. Authentic assessment content and synthetically generated camouflage are represented as a unified rendering while remaining recoverable only by legitimate candidates. The framework models the adversarial extraction process through an explicit extraction-channel operator and develops six coupled constructs: the Context Inversion Operator, Contextual Lamination Operator, Separation Channel, Human Readability Functional, Computational Ambiguity Functional, and Context Camouflage Tensor. Computational ambiguity is formulated using conditional entropy, yielding a closed-form expression that quantifies uncertainty during unauthorized extraction, while legitimate recovery is guaranteed through an exact filtering identity. We further establish theoretical properties governing ambiguity, camouflage density, semantic preservation, multi-observation leakage, and temporal multiplexing, and present a rendering algorithm with computational complexity and a pre-registered evaluation protocol. MCCT provides a mathematically rigorous foundation for behaviorally adaptive, accessibility-aware, and computationally resilient digital assessment by securing rendered assessment content while preserving readability for legitimate users.
カオス紛争測定および歴史経験重み付けによる堅牢なデンプスター・シェーファー証拠の融合
デンプスター・シェーファー理論に基づく複数情報源証拠の融合は、2 つの永続的な課題に直面しています。既存の紛争対策は、証拠間の不整合と証拠内不確実性を独立して評価し、不完全な評価をもたらします。また、現在の融合手法は、多様な意思決定の文脈にわたる長期的な信頼性を利用することなく、瞬間的な比較によってのみ証拠情報源を評価します。この論文では、両方の制限に対処する統一された証拠推論フレームワークを提案します。具体的には、カオスコンフリクト測定を導入して、証拠間の矛盾と証拠内の非特異性を共同で定量化し、正式に証明された 5 つの特性により一貫した評価を保証します。履歴経験に基づく重み付けスキームは、スペクトル クラスタリングを介して決定空間を分割し、リグレス理論を適用して過去の融合結果からコンテキスト固有の信頼性プロファイルを計算します。これらのメカニズムは、グローバルな対立レベルによって制御される加重コンセンサスに対して不確実性の保存を適応的にバランスさせるハイブリッド組み合わせルールに組み込まれ、その後、認識論的な不確実性を破棄することなく堅牢な分類を可能にする信念区間決定戦略が続きます。 16 の実世界のベンチマーク データセットでの実験では、提案されたフレームワークが平均 F1 スコア 85.78、平均 AUC 93.30 を達成し、8 つの DST ベースのベースラインと 3 つの勾配ブースティング手法を上回るパフォーマンスを示していることが実証されています。アブレーション分析により、私たちが提案した各コンポーネントの寄与が確認されます。このフレームワークは、複数の情報源による意思決定における適応的な証拠の融合のための効果的なアプローチを提供します。
原文 (English)
Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting
Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate evidence sources exclusively through instantaneous comparisns without exploiting their long-term reliability across diverse decision contexts. This paper proposes a unified evidence reasoning framework that addresses both limitations. Specifically, a chaos-conflict measurement is introduced to jointly quantify cross-evidence conflict and intra-evidence non-specificity, with five formally proven properties ensuring consistent assessment. A historical experience driven weighting scheme partitions the decision space via spectral clustering and applies regret theory to compute context-specific reliability profiles from past fusion outcomes. These mechanisms feed into a hybrid combination rule that adaptively balances uncertainty preservation against weighted consensus, controlled by the global conflict level, followed by a belief-interval decision strategy that enables robust classification without discarding epistemic uncertainty. Experiments on 16 real-world benchmark datasets demonstrate that the proposed framework achieves an average F1 score of 85.78 and a mean AUC of 93.30, outperforming eight DST-based baselines and three gradient boosting methods. Ablation analysis confirms the contribution of each component we proposed. The framework offers an effective approach for adaptive evidence fusion in multi-source decision making.
SkillEvo: マルチターン インタラクション フィードバックによる自己更新型の進化勾配
現在、エージェント スキルは手動で作成されるか、単一の LLM 生成パスで作成されるため、実際に引き起こされるインタラクションの失敗を改善するための閉ループがありません。最近の研究ではこのループは閉じられていますが、そのフィードバックは 1 ターンの質問応答評価から得られています。その結果、急激な非対称性が生じます。最初のラウンドで 1 回の交換で明らかになるギャップが埋められると、進化の勾配は減衰し、複数のターンでのみ表面化する欠陥は見えなくなり、進化は停滞します。これらのシステムのガバナンスも同様に、エンドツーエンドの検証スコアによって駆動されます。スカラー ゲートは、劣化した候補を拒否できますが、その構造的原因を特定したり修復したりすることはできません。私たちは、持続的なスキルの進化に対する拘束力のある制約は、編集能力でも反復回数でもなく、評価フィードバックが信頼できる進化勾配を提供し続けるかどうかであると主張します。 SkillEvo を紹介します。信頼できるフィードバックが勾配を生成し、制御可能なガバナンスがその方向性を制限します。最初のコンポーネントは、マルチターンのユーザー シミュレーションを評価エンドポイントからフィードバック ジェネレーターにリキャストします。フォローアップの質問によって層ごとに欠陥が明らかになり、修正の各ラウンドでフィードバックが消費され、新しいフィードバックが生成されます。 2 つ目は、スカラー ゲートの受動的な拒否を独立したガバナンス レイヤーに置き換え、事実上の劣化と構造の肥大化を積極的に修復し、劣化が蓄積するにつれて勾配がドリフトするのを防ぎます。 SkillEvo は、6 つのカテゴリのクラウド サービス、9 つの本番スキル、および 98 のスキル参照ファイルにわたって、内省ベースの進化を 23.0 ポイント、シングル ターン QA 主導の進化を 15.4 ポイント上回っています。
原文 (English)
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.
大規模言語モデルにおける数値計算: 基本的な制限と改善への道
大規模言語モデル (LLM) は、数学的推論ベンチマークでは優れた結果を達成しますが、大小比較、大きな整数の算術、分数、科学的表記法などの基本的な数値タスクでは信頼性が低いままです。この調査では、高度な数学的推論とは異なる能力として、基本的な数値的理解を調査します。私たちは、数的根拠を、数値形式を値、大きさ、および同等の表現にマッピングする表現グラウンディング (RG) と、数学的定義に従って算術演算を実行する手続き的グラウンディング (PG) に分解する数値グラウンディング フレームワーク (NGF) を提案します。 NGF を使用して、最近の診断ベンチマーク、故障モード、構造的説明、緩和戦略を整理します。トークン化、位置エンコーディング、埋め込みジオメトリ、および事前トレーニング データの配布に関する証拠をレビューします。また、Number Cookbook、NumericBench、GSM-Symbolic にわたる 3 つのフロンティア モデル ファミリの調整された評価にも NGF を適用し、アトミック、コンテキスト、推論支援型の数値計算を比較します。数字を意識したトークン化やアバカス埋め込みなどのアーキテクチャ介入は、ゼロからトレーニングされたモデルを改善できますが、教師あり微調整、推論足場、外部ツールの方が実用的である事前トレーニング済みシステムのユーザーには一般に利用できません。最後に、基礎モデルにおけるより信頼性の高い数値動作のための展開に関する推奨事項と研究の方向性について説明します。
原文 (English)
Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement
Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical tasks, including magnitude comparison, large-integer arithmetic, fractions, and scientific notation. This survey examines basic numerical understanding as a capability distinct from high-level mathematical reasoning. We propose the Numerical Grounding Framework (NGF), which decomposes numeracy into Representational Grounding (RG), mapping numeral forms to value, magnitude, and equivalent representations, and Procedural Grounding (PG), executing arithmetic operations in accordance with their mathematical definitions. Using NGF, we organize recent diagnostic benchmarks, failure modes, structural explanations, and mitigation strategies. We review evidence concerning tokenization, positional encoding, embedding geometry, and pretraining-data distribution. We also apply NGF in a coordinated evaluation of three frontier model families across Number Cookbook, NumericBench, and GSM-Symbolic, comparing atomic, contextual, and reasoning-assisted numeracy. Architectural interventions such as digit-aware tokenization and Abacus Embeddings can improve models trained from scratch but are generally unavailable to users of pretrained systems, for whom supervised fine-tuning, reasoning scaffolds, and external tools are more practical. We conclude with deployment recommendations and research directions for more reliable numerical behavior in foundation models.
LLM の正規化配置の再考: カリキュラムの深さが増す中でのポスト規範
プレノームは、フルデプス モデルの共同最適化を容易にするため、最新の Transformers における標準的な正規化配置です。私たちは、カリキュラムを通じて深みが導入されたときに、この好みが持続するかどうかを尋ねます。カリキュラムの深さの増加では、追加された各ブロックはトレーニングされたプレフィックスによって生成された境界表現を受け取り、正規化の配置が順方向条件付けに関連するようになります。したがって、配置とトレーニングカリキュラムが相互作用するかどうかをテストします。 Qwen3-8B 教師と 9 層の生徒による制御蒸留研究では、共同トレーニング下ではプレノームとポストノームは区別できず、その差は $0.0004$ 検証 CE でしたが、ポストノームはカリキュラムの成長下ではプレノームよりも $0.0328$ 改善し、一桁大きくなりました。スチューデントのアクティブ レイヤー トークンと一致するジョイント後のコントロールは、成長後よりも悪いままであり、コンピューティングが唯一の説明として除外されます。ランキングはカリキュラム中に切り替わります。ブロックが追加されると、ポストノルムがリードします。単一ブロックとフリーズ制御は、ランクの変更を、浅いブロックの品質や再トレーニングではなくブロックの追加に局所化します。境界診断では、ポストノルムを安定した残差スケールに関連付け、プレノルムを構造トークンスケールのドリフトに関連付けます。固定バッチでは、最終的な成長前ブロックもほぼアイデンティティ マップされます。位相ごとのクロスオーバーと合わせて、これらの観察は、新しいブロックが追加された後の境界スケールの条件付けと一致しています。この結果により、この蒸留設定では正規化の配置とトレーニング カリキュラムを組み合わせた設計の選択肢として扱うことが動機付けられました。
原文 (English)
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing
Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by $0.0004$ validation CE, while post-norm improves over pre-norm by $0.0328$ under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting.
SkillShapley: LLM エージェントのスキル ステップ アトリビューションのための境界適応型 Shapley 評価
エージェント スキルは、言語エージェントがコーディングや文書処理などの長い手続きタスクを実行できるようにする重要な外部指示です。既存のエージェント スキルは主に人間の手作業による作成やエージェントの実行トレースによって作成されており、各ステップが特定のタスクにおける全体的なスキル パフォーマンスにどのように寄与するかについては十分な理解がありません。つまり、エージェント スキル内の個々のステップの貢献を定量化する際には未解決の問題が残っています。この問題に対処するために、まずスキルステップ アトリビューションを Shapley 値ベースの貢献推定問題としてモデル化し、次にエージェント スキルのステップレベル アトリビューション フレームワークである SkillShapley を提案します。特に、SkillShapley は 2 つのフェーズで動作し、重要な経験的洞察、つまりパフォーマンスの急激な崖を生み出す離散化されたベンチマーク報酬と、相乗的ではなく主に相加的なステップの相互作用によって動機付けられています。具体的には、最初に有益な連合領域を特定し、次に再利用可能な限界証拠を生成できる新しい連合を適応的にサンプリングします。広く採用されている SkillsBench のスキルに関する実験では、SkillShapley が価値の高いスキル ステップと低いスキル ステップを効果的かつ効率的に識別できることが実証され、エージェントのスキル作成に重要なポイントがいくつか提供されます。
原文 (English)
SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents
Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through human manual crafting or agent execution traces, with limited understanding of how each step contributes to overall skill performance on specific tasks; i.e., there remains an open problem in quantifying the contribution of individual steps within an agent skill. To address this issue, we first model skill-step attribution as a Shapley value-based contribution estimation problem, and then propose SkillShapley, a step-level attribution framework for agent skills. Notably, SkillShapley operates in two phases, motivated by key empirical insights, i.e., discretized benchmark rewards that create sharp performance cliffs, and step interactions that are largely additive rather than synergistic. Specifically, it first identifies informative coalitional regions, and then adaptively samples new coalitions that can yield reusable marginal evidence. Experiments on skills from the widely adopted SkillsBench demonstrate that our SkillShapley can effectively and efficiently identify high- or low-value skill steps, providing several key takeaways for agent skill creation.
方向ではなく大きさを教える: マルチターン、マルチステップ LLM エージェントに対する検証者限定のクレジット割り当て
検証可能な報酬を伴う強化学習 (RLVR) は、マルチターンのツール使用エージェントをトレーニングするために検証者に制限されたパフォーマンスの上限を提供しますが、その軌道レベルのクレジット割り当てにより、ターンごとの異質な結果が 1 つの報酬信号に統合されます。ポリシーに基づく蒸留は、トークンごとの高密度の監視を提供しますが、教師の制限があるか、勾配濃度の崩壊が発生しやすいです。 $\textbf{CrEST}$ を導入します。これは、特権を持つ自己教師からの高密度のトークンレベルのシグナルを組み込みながら、RL の検証者制限の上限を維持する階層的な単位割り当てフレームワークです。 $\textbf{CrEST}$ は 2 つのレベルでクレジットを解決します。ターンセグメント化された検証済みアドバンテージはターン間の希薄化に対処し、エントロピー ゲートによる自己教師変調はターン内トークンの寄与を調整します。 BFCL V3 と WildToolBench での実験では、$\textbf{CrEST}$ が 2 つのモデル スケールにわたって RL ベースラインと蒸留ベースラインの両方を常に上回っており、長い軌道と厳格なセッション レベルのメトリクスで最大のゲインが得られることが示されています。私たちの研究は、ポリシーの最適化における教師の役割を、更新方向の決定から更新規模の調整まで削減し、検証者の限界を犠牲にすることなく高密度の単位割り当てを解除できることを示しています。
原文 (English)
Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents
Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce $\textbf{CrEST}$, a hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. $\textbf{CrEST}$ resolves credit at two levels: turn-segmented verified advantages address inter-turn dilution, while entropy-gated self-teacher modulation refines intra-turn token contributions. Experiments on BFCL V3 and WildToolBench show that $\textbf{CrEST}$ consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. Our work demonstrates that the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.
TsuGO: Go 生死に関わる問題による LLM 推論の検索効率の調査
LLM 推論の評価は、最終的な回答の精度からプロセス レベルの評価に移行しつつありますが、既存の手法では、モデルが推論パスを計画し、推論リソースをどのように割り当てるか、つまりモデルが検索をどのように組織するかをまだ把握できていません。従来のプロセス レベルの手法は、思考連鎖 (CoT) の一貫性と冗長性に焦点を当てており、ほとんどのベンチマーク タスクは、導出やツールの使用などの静的な機能によって解決できる単一の目標を持っており、検索組織は測定されていません。 Go の死活問題を通じて LLM 推論の検索効率を評価するためのプロセスレベル推論ベンチマークである TsuGO を紹介します。これらの問題は、固有の敵対構造を備えた閉じられた検証可能な解決空間を提供し、候補の生成、応答チェック、分岐比較、および偶発的なトレース パターンではなく推論の必要な部分のバックトラックを行います。 TsuGO は、ソリューション空間を制限することで、検索組織からドメイン知識を分離し、CoT を解析して構造化された検索ツリーにし、検索効率をトークン効率やその他の診断メトリクスおよび視覚化とともにレポートします。実験によると、現在の LLM は安定した詰碁の解決にはほど遠いことがわかりました。より強力なモデルは正しい候補を早期に見つけ、生産的な分岐で努力を続けることで成功しますが、ほとんどのモデルは依然として、ニューラルガイド付き KataGo よりもガイドなしの検索アルゴリズムにはるかに近い動作をします。 CoT が長くても、トークン効率が高くても、必ずしも検索が優れているとは限りません。私たちの結果は、LLM 推論の評価に欠けている要素として、検索組織と推論リソースの割り当てを特定しました。
原文 (English)
TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems
The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search. Prior process-level methods focus on the coherence and redundancy of chain-of-thought (CoT), and most benchmark tasks have a single objective solvable by static capabilities such as derivation and tool use, leaving search organization unmeasured. We introduce TsuGO, a process-level reasoning benchmark for evaluating Search Efficiency in LLM reasoning through Go life-and-death problems. These problems provide closed and verifiable solution spaces with an inherent adversarial structure, making candidate generation, response checking, branch comparison, and backtracking necessary parts of reasoning rather than incidental trace patterns. By constraining the solution space, TsuGO disentangles domain knowledge from search organization, parses CoT into a structured search tree, and reports Search Efficiency together with Token Efficiency and other diagnostic metrics and visualizations. Experiments show that current LLMs remain far from stable tsumego solving: stronger models succeed by finding the correct candidate earlier and sustaining effort on productive branches, but most models still behave much closer to unguided search algorithms than to neural-guided KataGo. Longer CoT or higher Token Efficiency does not necessarily imply better search. Our results identify search organization and reasoning-resource allocation as missing dimensions in LLM reasoning evaluation.
構成エージェントハーネス修復のための機能シーブ: 制御された商と実際のリポジトリストレステスト
エージェント ハーネスは、取得、ルーティング、状態、来歴、検証を組み合わせていますが、ローカルで成功したコンポーネントは共有状態に関して一致しない可能性があります。この失敗を有限の \emph{能力束} でモデル化します。ストークは型付き動作シグネチャをエンコードし、制限マップは共有フィールドを保持し、受け入れられた実行は有用なグローバル セクションです。正確な有限制約満足問題 (CSP) が受け入れを定義し、線形化された相対コホモロジー クラスが診断機能と検索機能を提供します。 20 のタスク クラスターにわたる制御された実験では、生の状態が迷惑変数である隠れた内部メディエーターが導入されます。それらの境界を指数化すると、候補者の予算がクラスターあたり 2,000 から 1,000 に削減されます。非表示状態を揃えるとギャップが解消されます。正確な CSP は商と一致するため、結果は正確な推論に対する優位性ではなく、古い代表に対する不変性を示します。次に、PatchFuseBench の SWE ベンチ多言語プールからの検出分割 (20 のリポジトリからの 160 の問題、875 の実際の候補パッチ、2,579 のソース認識編集アトム、および 153 の新しく実行されたパッチ) でメソッドをテストします。 $\operatorname{coker}D$ の $[b-Dx]=[b]$ であるため、最初のプールレベルの構造は一定であり、したがって構成をランク付けできません。候補インデックス付き修復は、848/875 の候補では自明ではなく、120/160 の問題内で異なります。一致する非コホモロジー セレクターでは 116 件の問題に対して 118 件の問題が解決されますが、その違いはリポジトリ間でサポートされていません (正確な符号反転 $p=0.75$)。リーブ 1 リポジトリ アウトの棄権ゲートは 127/160 に達し、強力なアンカーと結び付き、一致するゲートを 1 銘柄上回ります ($p=1.0$)。したがって、発見ゲートは失敗し、確認の分割は封印されたままになります。この研究は、制御された不変メカニズムと識別可能性の補正をサポートしていますが、現実世界のコホモロジーの利点はサポートしていません。
原文 (English)
Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test
Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We model this failure with a finite \emph{capability sheaf}: stalks encode typed behavior signatures, restriction maps retain shared fields, and accepted runs are useful global sections. An exact finite constraint-satisfaction problem (CSP) defines acceptance, while a linearized relative cohomology class provides a diagnostic and search feature. A controlled experiment over 20 task clusters introduces hidden interior mediators whose raw states are nuisance variables. Quotienting their coboundaries reduces the candidate budget from 2,000 to 1,000 per cluster; aligning the hidden state removes the gap. Exact CSP matches the quotient, so the result demonstrates invariance to stale representatives, not superiority over exact reasoning. We then test the method on a discovery split from the SWE-bench Multilingual pool of PatchFuseBench: 160 issues from 20 repositories, 875 real candidate patches, 2,579 source-aware edit atoms, and 153 newly executed patches. A first pool-level construction is constant because $[b-Dx]=[b]$ in $\operatorname{coker}D$ and therefore cannot rank configurations. A candidate-indexed repair is nontrivial on 848/875 candidates and varies within 120/160 issues. It resolves 118 issues versus 116 for a matched noncohomological selector, but the difference is not supported across repositories (exact sign-flip $p=0.75$). A leave-one-repository-out abstention gate reaches 127/160, tying the strong anchor and exceeding its matched gate by one issue ($p=1.0$). The discovery gate therefore fails and the confirmatory split remains sealed. The study supports the controlled invariance mechanism and an identifiability correction, but not a real-world cohomological advantage.
vToken: 再利用可能な KV キャッシュのためのトークンレベルの仮想化
大規模な言語モデルのサービスは、重大なメモリ ボトルネックに直面しています。KV キャッシュは、シーケンスの長さとバッチ サイズに応じて増大します。 PagedAttendance は固定サイズのメモリ ブロックを使用してアロケータ レベルの断片化を軽減しますが、最近の KV エビクション アルゴリズムはブロック レベルの管理よりも細かいトークン粒度で動作します。この不一致によりブロック内の断片化が発生し、割り当てられた KV メモリの大部分が再利用不可能な状態になります。論理トークンの活性化を物理ブロックの配置から切り離す、軽量のトークンレベルの仮想化レイヤーである vToken を紹介します。 vToken は、トークン テーブルの間接化を通じて安定した論理トークン ビューを維持し、ライブ トークンを非同期に再パックすることで物理的な再利用を実現します。この設計では、PagesAttention カーネルと CUDA Graph の互換性が維持されます。 vLLM に vToken を実装し、モデル全体で H2O、ランダム、シザーハンズを使用して評価します。ペアの Naive-Evict ベースラインと比較して、vToken はリクエストごとに保持される KV ブロックを 27.2\%--72.3\% 削減し、SLA 制約のあるスループットを最大 1.37$\times$ 向上させます。制約のあるアクティブ KV 予算の下で、実現可能な最大同時実行数を最大 2$\times$ 拡張し、ポリシーごとの統合フットプリントを 500 以上から 50 行未満に削減します。
原文 (English)
vToken: Token-Level Virtualization for Reclaimable KV Caches
Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragmentation, but recent KV eviction algorithms operate at a token granularity finer than block-level management. This mismatch causes intra-block fragmentation, leaving a large fraction of allocated KV memory unreclaimable. We present vToken, a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement. vToken maintains a stable logical token view through token-table indirection and realizes physical reclamation by repacking live tokens asynchronously. The design preserves PagedAttention kernels and CUDA Graph compatibility. We implement vToken in vLLM and evaluate it with H2O, Random, and Scissorhands across models. Compared with a paired Naive-Evict baseline, vToken reduces retained KV blocks per request by 27.2\%--72.3\% and improves SLA-constrained throughput by up to 1.37$\times$. Under a constrained active-KV budget, it extends the maximum feasible concurrency by up to 2$\times$, while reducing the per-policy integration footprint from 500+ lines to under 50.
必然的に主権者?最先端の AI 輸出規制、サイバーセキュリティ、国家 AI 能力の限界
2 つの州に拠点を置く少数の企業が、最も有能なフロンティア AI モデルを生産しています。これらの州の政府は、他のどの国がこれらのシステムを使用できるかを決定する法的権限と政治的意志の両方を示しています。 2026年6月、米国は大手開発者に対し、米国在住の外国人を含む外国人に最新モデルをリリースする前にライセンスを取得することを義務付けた。影響を受けるモデルは、制限の管理が非現実的であることが判明したこともあり、急遽世界中で廃止されました。これは、ほぼ自律的に行われる AI によるサイバースパイ活動の最初の文書化された事件から数か月以内に発生したもので、フロンティアモデルがサイバー攻撃とサイバー防御の両方の経済性を変えるという証拠の増加と一致しました。この記事では、これら 2 つの開発がどのように相互作用するかを検証し、現在大規模な AI 開発を推進している異常な市場力学の中にそれらを位置づけます。フロンティアAIへのアクセスは国家サイバー防衛の一部になりつつあり、そのようなアクセスは取り消すことが可能であり、主権能力の明白な救済策は一部の国を除いて部分的にしか実現不可能であると主張している。この論文は、訓練コスト、コンピューティング能力の集中、国家 AI プログラムによって提供されるサポートに関する証拠を基に、中小国、さらには大国にとって主権が現実的に何を意味するのかを問いかけています。この記事では、交渉によるアクセス保証、推論レベルでの主権、オープンウェイトモデルによるヘッジ、地域能力のプール、持続的な人材育成、基本的なサイバーレジリエンスへの継続的な投資といった階層的な戦略を提案している。オープンウェイトヘッジは、一般に考えられているよりも優れた能力を持ち、政治的にさらされていることがすぐに証明されています。短期的なリスクの多くは、見かけのパフォーマンスではなく、有能なモデルがどのように展開され、封じ込められるかにあります。
原文 (English)
Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability
A small number of firms based in two states produce the most capable frontier AI models. The governments of those states have shown both the legal power and the political will to decide which other countries may use these systems. In June 2026 the United States required a leading developer to obtain licences before releasing its most advanced models to any foreign person, including foreign nationals resident in the United States. The affected models were withdrawn worldwide at short notice, partly because the restriction proved impractical to administer. This followed within months of the first documented case of a largely autonomous, AI-run cyber espionage campaign, and coincided with mounting evidence that frontier models alter the economics of both cyber attack and cyber defence. This article examines how these two developments interact, and situates them within the unusual market dynamics now driving large-scale AI development. It argues that access to frontier AI is becoming part of national cyber defence, that such access can be revoked, and that the obvious remedy of sovereign capability remains only partly feasible for all but a handful of states. Drawing on evidence about training costs, the concentration of computing power and the support offered by national AI programmes, it asks what sovereignty can realistically mean for small and middle powers, and for large powers as well. The article proposes a layered strategy: negotiated access guarantees, sovereignty at the level of inference, hedging with open-weight models, pooled regional capability, sustained talent development and continued investment in basic cyber resilience. The open-weight hedge proves at once more capable and more politically exposed than is commonly assumed. Much of the near-term risk lies in how capable models are deployed and contained rather than in their apparent performance.
在宅日常生活における状況に応じた臨床動作の理解に向けて: 自己中心的な視覚による歩行検出のフリーズ
日常生活における動作を理解するには、運動学を超えた文脈が必要です。日常生活活動 (ADL) 中の同様の慣性パターンは、意図的な停止、物体の相互作用、または病的な運動障害を反映している可能性があるためです。自己中心的なビジョンは、これらのケースの曖昧さを解消するのに役立つ可能性のあるタスク関連のコンテキストを提供します。私たちは、ADL 中の状況要因によって強く影響される症状であるパーキンソン病 (PD) におけるすくみ歩行 (FOG) の検出を通じて、この課題を調査します。同期された自己中心ビデオ、ウェアラブル IMU、および自宅にいる 13 人の PD 参加者から収集された専門家による注釈付き FOG ラベルを使用して、事前学習された自我ビデオと時系列基礎モデルからの凍結表現を、ゼロから学習された IMU ベースの TCN と並行して、1 被験者抜き評価の下で評価します。 IMU ベースの TCN は、V-JEPA2 エゴビデオ機能の 32.6 F1 および 77.2 AUROC と比較して、42.3 F1 および 83.0 AUROC に達する最強のイベント検出パフォーマンスを達成しました。自我ビデオだけでは IMU ベースのセンシングを上回る性能はありませんでしたが、偶然以上の識別を示し、定性的分析により、自己中心的な視覚は IMU とは独立して FOG 関連情報を捕捉できる可能性があることが示唆されています。これらの結果を総合すると、日常生活におけるウェアラブルセンサーベースの臨床動作の理解にコンテキスト情報を追加するための、事前トレーニングされた自我ビデオ表現の使用が裏付けられます。
原文 (English)
Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision
Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional stopping, object interaction, or pathological movement impairment. Egocentric vision provides task-related context that may help disambiguate these cases. We investigate this challenge through freezing of gait (FOG) detection in Parkinson's disease (PD), a symptom strongly influenced by contextual factors during ADLs. Using synchronized egocentric video, wearable IMUs, and expert-annotated FOG labels collected from 13 PD participants in their homes, we evaluate frozen representations from pretrained ego-video and time-series foundation models, alongside an IMU-based TCN trained from scratch, under leave-one-subject-out evaluation. The IMU-based TCN achieved the strongest event-detection performance, reaching 42.3 F1 and 83.0 AUROC, compared with 32.6 F1 and 77.2 AUROC for V-JEPA2 ego-video features. Although ego-video alone did not outperform IMU-based sensing, it showed above-chance discrimination, and qualitative analyses suggest that egocentric vision may capture FOG-relevant information independent of IMUs. Together, these results support the use of pretrained ego-video representations to add contextual information to wearable-sensor-based clinical motion understanding in daily living.
NAS 主導のハードウェア アクセラレータによるエッジ AI とパレート空間への量子化効果の探索
エッジ AI の導入には、正確で計算効率が高く、ハードウェア導入が可能なニューラル アーキテクチャが必要です。この課題は、ハードウェア対応のニューラル アーキテクチャ検索 (NAS) によって解決されます。最近の研究では量子化を NAS ループに直接組み込んでいますが、これらのアプローチは検索の複雑さを拡大し、アーキテクチャと量子化設計を緊密に結合します。より単純な検索後の量子化戦略は、分析上の注目をほとんど受けていません。NAS で発見されたパレート構造に対するトレーニング後量子化 (PTQ) の影響はまだ解明されておらず、再構成可能なアクセラレータへの量子化アーキテクチャ マッピングと自動ハードウェア探索を組み合わせたフレームワークもありません。この文書では両方のギャップについて取り上げます。まず、3 段階のパイプラインが提案されています。NAS-Bench-201 上のハードウェアに依存しないパレート ランク サロゲート フロントエンド、パレート認識フィルタリングとフィードバック制御を備えた量子化ブリッジ、最適なハードウェア マッピングのための CGRA4ML 上の進化的ドメイン空間探索 (DSE) バックエンドです。 2 番目に、実証研究では、15,625 アーキテクチャすべてのグラウンドトゥルース データの形式的安定性メトリクスを通じて、INT4 PTQ が NAS-Bench-201 パレート空間をどのように混乱させるかを特徴付け、2 つの標準的な検索戦略にわたるパレート空間カバレッジにおいて、FP32 ゼロショット サロゲートが専用の INT4 トレーニング サロゲートよりも優れていることを実証しています。
原文 (English)
NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space
Edge AI deployment demands neural architectures that are simultaneously accurate, computationally efficient, and hardware-deployable - a challenge addressed by hardware-aware Neural Architecture Search (NAS). While recent works incorporate quantization directly into the NAS loop, these approaches expand search complexity and tightly couple architecture and quantization design. The simpler post-search quantization strategy has received little analytical attention: the effects of Post-Training Quantization (PTQ) on the NAS-discovered Pareto structure remain uncharacterised, and no framework combines quantized architecture mapping onto reconfigurable accelerators with automated hardware exploration. This paper addresses both gaps. First, a three-stage pipeline is proposed: a hardware-agnostic Pareto rank surrogate frontend on NAS-Bench-201, a quantization bridge with Pareto-aware filtering and feedback control, and an evolutionary Domain Space Exploration (DSE) backend on CGRA4ML for optimal hardware mapping. Second, an empirical study characterises how INT4 PTQ perturbs the NAS-Bench-201 Pareto space through formal stability metrics on ground-truth data for all 15,625 architectures, and demonstrates that an FP32 zero-shot surrogate outperforms a dedicated INT4-trained surrogate in Pareto space coverage across two standard search strategies.
StateBridge: LLM マルチエージェント システムにおける潜在通信のためのトレーニング不要の隠れ状態の調整
大規模な言語モデルに基づくマルチエージェント システムは通常、テキストで通信します。つまり、個別のトークンを使用します。ただし、テキストでは個別のボトルネックが発生します。送信者の連続的な隠れ状態を個別のトークンに変換すると、トークン ID だけでは取得できない情報が破棄されます。最近の研究では、エージェントが隠れた表現をテキストに変換せずに直接送信するという、潜在的なコミュニケーションを代替手段として提案しています。ただし、既存の潜在的な方法では、トランスフォーマー全体で作業メモリを層ごとに注入するか、移植性を制限する訓練されたプロジェクターを必要とします。我々は、閉形式直交変換を介して送信者の最終層の隠れ状態を受信者の入力空間に位置合わせする、トレーニング不要の潜在通信アプローチである StateBridge を提案します。軽量のノルム キャリブレーションと語彙アンカーにより、事前トレーニングされた入力分布との互換性が保証されます。整列された状態は、連続プレフィックスとして受信側エージェントの入力に付加されます。 2 つのファミリーの 4 つのモデルを使用して、数学的推論、コード生成、および質問応答に関して StateBridge を評価します。 StateBridge は、26 のモデルとタスクのペアのうち 22 で最高または同順位のスコアを達成し、最も強力なベースラインを常に上回っています。
原文 (English)
StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems
Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discrete bottleneck. Converting the sender's continuous hidden states into discrete tokens discards information that token identities alone cannot capture. Recent work proposes latent communication as an alternative, where agents transmit hidden representations directly without converting them to text. However, existing latent methods either inject working memory layer by layer across the transformers, or require trained projectors that limit portability. We propose StateBridge, a training-free latent communication approach that aligns the sender's final-layer hidden states to the receiver's input space via a closed-form orthogonal transformation. Lightweight norm calibration and vocabulary anchoring ensure compatibility with the pretrained input distribution. The aligned states are prepended to the input of the receiver agent as a continuous prefix. We evaluate StateBridge on math reasoning, code generation, and question answering with four models from two families. StateBridge achieves the best or tied-best score on 22 out of 26 model-task pairs, consistently outperforming the strongest baseline.
構造ベースの局所改善手法のための LLM ガイド付きグラフ生成
大規模な近傍検索では、通常、反復最適化のために決定変数のランダムなサブセットが選択されます。さまざまな問題を効率的に解決するために、研究者はさまざまなドメインの構造的特徴を考慮して変数選択戦略を設計する傾向があります。このペーパーでは、MiniZinc 形式のすべての問題に対して問題を意識しない自動パイプラインを構築します。 LLM にセマンティック ガイドラインを要求することで、LLM が問題タイプのインスタンスを均一に重み付けされたグラフにマッピングするグラフ ジェネレーターを生成するように導きます。ノードは決定変数を表し、エッジは制約関係を表します。これらの問題に依存しないグラフは、変数選択における構造ベースのローカル改善フレームワーク (SLIM) をガイドします。一方、重み付きグラフを使用すると、すべての問題インスタンスが同じ汎用グラフ表現を共有できるようになり、そこから同じグラフの特徴を抽出して構成の選択に使用できます。 MiniZinc のコンペティション問題 20 件にわたるインスタンスでパイプラインを評価したところ、アルゴリズムの選択により、ワンショットの Gurobi ベースラインに対して問題に重み付けされた平均勝率 39.5% が達成され、これは最良の単一構成 (19.3%) の 2 倍以上であることがわかりました。構成と機能のアブレーションによりパフォーマンスがさらに 44.0% 向上し、LLM ベースのセマンティック生成により、制約を最適化するための効果的な自動構造抽出と機能抽出が可能になることが実証されました。
原文 (English)
LLM-Guided Graph Generation for Structure-Based Local Improvement Methods
Large neighborhood search normally selects a random subset of decision variables for iterative optimization. For efficiently solving different problems, researchers tend to design variable selection strategies by taking into account structural features from different domains. In this paper, we build an automatic pipeline that is problem-agnostic to all problems in the MiniZinc format. By prompting an LLM with our semantic guidelines, we guide the LLM to produce a graph generator that maps any instance of a problem type to a uniform weighted graph, where nodes represent decision variables and edges represent constraint relationships. These problem-agnostic graphs guide our structure-based local improvement framework (SLIM) in variable selection. Meanwhile, the weighted graph enables all problem instances to share the same generic graph representation, from which the same graph features can be extracted and used for configuration selection. We evaluated our pipeline on instances across 20 MiniZinc competition problems, finding that algorithm selection achieves a 39.5% average problem-weighted win rate against a one-shot Gurobi baseline, more than doubling the best single configuration (19.3%). Configuration and feature ablation boost the performance further to 44.0%, demonstrating that LLM-based semantic generation enables effective automated structure extraction and feature extraction for constraint optimization.
LongEarth-R1: 長期地球観測推論のための視覚言語モデルのベンチマークと調整
長期にわたる地球観測の推論には、多段階の地理的進化を組織化し、空間変化を局所的に特定し、時間的異常を検出し、拡張された画像シーケンスから未来を推測するためのモデルが必要です。しかし、既存のリモートセンシング視覚言語モデルは、主に孤立した画像、画像ペア、または短いシーケンスに焦点を当てており、関連するフレームや領域での信頼できる接地が制限されています。 LongEarth-Bench を紹介します。これは、117,000 の固有の画像から得られた約 120,000 の質問応答サンプルを含むベンチマークです。そのシーケンスは平均 15.14 フレームから 30 フレームまで拡張され、進化の要約、空間推論、異常の特定、論理予測にわたる 12 のタスクをカバーします。さらに、30k サンプルのサブセットは、キー フレームと変更された領域を最終的な答えにリンクする構造化された推論トレースを提供します。私たちは、明示的な配列識別子と構造化された思考連鎖の監視による監視付き微調整を通じて LongEarth を開発します。 LongEarth 上に構築された LongEarth-R1 は、形式、時間的、空間的報酬を使用してグループ相対ポリシーの最適化を適用します。 LongEarth-R1 は、標準的なリモート センシング ベンチマークでの競争力を維持しながら、12 の長いシーケンスのタスクすべてで最高の結果を達成します。
原文 (English)
LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning
Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.
ルールか性格か? AI 安全設計のためのスケーリング法則
人工知能 (AI) 安全システムは、トレーニング時に行動分布を変更する性格形成 (例: ヒューマン フィードバックからの強化学習 [RLHF]、憲法 AI) と、推論時に有害な出力をブロックするルール強制 (例: 出力フィルター、安全分類子) を組み合わせていますが、導入規模が増加するにつれて最適なバランスがどのように変化するかについての正式な分析はほとんど存在しません。我々は、これら 2 つのアプローチ間の [0,1] のリソース割り当てアルファとして安全設計をパラメータ化する定型化された比較静的モデルを導入します。これには、スケール依存のフィルター劣化、コモンモード故障、および特性の脆弱性 (形成された動作が新しい条件下で劣化または崩壊するリスク) が組み込まれています。乗算パレート損害モデルの下で、閉形式の予想損害を導出し、それをモンテカルロ シミュレーションによるテール リスク (CVaR) 分析で補完します。 3 つのシナリオ (楽観的、中程度、悲観的) にわたって、最適なアルファ* は内部またはルールのみの境界にあり、展開スケール T が増加するにつれて、シナリオに応じてごくわずか (デルタ アルファ* = +0.01) から顕著な (デルタ アルファ* = +0.21) まで、キャラクター形成に向けて弱くシフトします。主要なパラメータは、ベースラインの文字脆弱性率 p^(0)_frag で、範囲全体で alpha* を 0.50 シフトします。これは、テール重大度、フィルタ品質、またはコモンモード故障確率の影響をはるかに超えています。 CVaR と予想される害の最適値は、大きな T で収束します。これらの結果は、安全アーキテクチャの決定は、展開規模自体には依存せず、分布シフトの下での特性形成の信頼性に依存することを示唆しています。
原文 (English)
Rules or Character? Scaling Laws for AI Safety Design
Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes safety design as a resource allocation alpha in [0,1] between these two approaches, incorporating scale-dependent filter degradation, common-mode failures, and character fragility -- the risk that shaped behavior degrades or collapses under novel conditions. Under a multiplicative Pareto damage model, we derive closed-form expected harm and supplement it with tail-risk (CVaR) analysis via Monte Carlo simulation. Across three scenarios (optimistic, moderate, pessimistic), the optimal alpha* is interior or at the rules-only boundary and shifts weakly toward character shaping as deployment scale T grows, from negligible (Delta alpha* = +0.01) to pronounced (Delta alpha* = +0.21) depending on scenario. The dominant parameter is the baseline character fragility rate p^(0)_frag, which shifts alpha* by 0.50 across its range -- far exceeding the effect of tail severity, filter quality, or common-mode failure probability. CVaR and expected-harm optima converge at large T. These results suggest that safety architecture decisions depend less on deployment scale per se than on the reliability of character shaping under distributional shift.
TopoIntent: セキュリティ インテントを実行可能なコンプライアンス チェック済みネットワーク トポロジにコンパイルする
エンタープライズ セキュリティ トポロジの設計では、ビジネスの意図、規制要件、リスクの想定をゾーン、境界デバイス、ゾーン間パス、アクセス制御ポリシーに変換する必要があります。既存の NetOps 自動化ツールは主にこの設計が修正された後に動作し、不明確な自然言語要件から構造化されたセキュリティ トポロジを生成するための限定的なサポートを提供します。私たちは、セキュリティの意図を実行可能なコンプライアンスチェック済みのネットワーク トポロジにコンパイルするシステムである TopoIntent を紹介します。スキーマ コントラクトを使用して生成を制限し、密ベクトル検索によって厳選されたテンプレート ライブラリから参照アーキテクチャを取得し、意図とテンプレートの調整とセキュリティの完了のために段階的融合を適用します。生成されたトポロジは、トポロジ レイヤーで表示される CIS Controls v8.1.2 セーフガードと照合してチェックされ、未解決のケースは手動レビュー用にマークされます。構造的なギャップは、スキーマを保持する編集を追加することで修復されます。最終的なトポロジは、カーネル レベルの iptables ACL を使用して Mininet スクリプトにエクスポートされ、実行可能ファイルの到達可能性と許可/拒否テストが可能になります。この要件からトポロジへのタスクに対する公開ベンチマークは存在しないため、参照セキュリティ アーキテクチャ図から評価セットを構築します。取得セットには、5 つのシナリオにわたる 22 のテンプレートと 44 の合成インテントが含まれていますが、保留セットには、取得から除外された金融および政府のシナリオからの 7 つのテンプレートと 14 のインテントが含まれています。ホールドアウト セットでは、加法的修復により、平均 1.5 ラウンド未満でトポロジ可視 CIS 満足度が 0.78 から 1.00 に改善され、1 回のフィードバック ラウンドで ACL 後のポリシー合格率が 0.78 から 0.88 に上昇しました。
原文 (English)
TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies
Enterprise security topology design requires translating business intent, regulatory requirements, and risk assumptions into zones, boundary devices, inter-zone paths, and access-control policies. Existing NetOps automation tools mainly operate after this design is fixed, providing limited support for generating structured security topologies from underspecified natural-language requirements. We present TopoIntent, a system that compiles security intent into executable, compliance-checked network topologies. It uses a schema contract to constrain generation, retrieves reference architectures from a curated template library via dense-vector search, and applies staged fusion for intent-template alignment and security completion. The generated topology is checked against CIS Controls v8.1.2 safeguards visible at the topology layer, while unresolved cases are marked for manual review. Structural gaps are repaired through additive schema-preserving edits. The final topology is exported to Mininet scripts with kernel-level iptables ACLs, enabling executable reachability and allow/deny tests. Because no public benchmark exists for this requirement-to-topology task, we construct an evaluation set from reference security architecture diagrams. The retrieval set contains 22 templates and 44 synthetic intents across five scenarios, while the held-out set contains 7 templates and 14 intents from finance and government scenarios excluded from retrieval. On the held-out set, additive repair improves topology-visible CIS satisfaction from 0.78 to 1.00 in fewer than 1.5 rounds on average, and one feedback round raises the post-ACL policy pass rate from 0.78 to 0.88.
トランスフォーマーベースのモデルを使用したコースと成績の共同予測
学習分析における既存の予測モデルは、多くの場合、学生の学歴を単純な順序として扱い、学期内に受講したコースの同時性を無視しています。この単純化により、特にコースの負荷が重い、または難しい学生の場合、パフォーマンスの予測が不正確になる可能性があります。このペーパーでは、学生が受講する一連のコースと次の学期の対応する成績の両方を共同で予測することで、この制限に対処する Academic Course-grade Estimation (TRACE) 用の TRansformer を紹介します。私たちのアプローチでは、コースを学期ごとにエンコードしてコースの同時実行の影響を捉え、コースセットの予測と成績予測を組み合わせた新しい損失関数を利用します。コースの成績に加えて受講したコースを予測すると、予測の品質が大幅に向上することが実証されました。 10 年間の組織データに基づいてトレーニングされた当社の共同予測モデルは、成績のみを予測する同一のアーキテクチャと比較して、平均絶対誤差をほぼ 50% 削減します。このモデルは、従来の LSTM ベースの逐次モデルやグラフ ニューラル ネットワーク ベースのアプローチよりも優れたパフォーマンスを発揮し、生徒の属性データを組み込む自然な方法を提供します。この研究は、再トレーニングと再キャリブレーションを通じて新しい施設に適応できる解釈可能なモデルを作成するための最新のニューラル アーキテクチャの有用性と、トレーニング中に受講したコースの予測などの重要なテクニックの重要性を示しています。このモデルを高等教育機関の早期発見システムにどのように組み込むことができるかについて説明します。
原文 (English)
Jointly Predicting Courses and Grades Using a Transformer-Based Model
Existing predictive models in learning analytics often treat student academic history as a simple sequence, overlooking the concurrent nature of courses taken within a semester. This simplification can lead to inaccurate performance predictions, particularly for students with heavy or challenging course loads. This paper introduces a TRansformer for Academic Course-grade Estimation (TRACE) that addresses this limitation by jointly predicting both the set of courses a student will take and their corresponding grades for an upcoming semester. Our approach encodes courses on a per-semester basis to capture the effects of course concurrency and utilizes a novel loss function combining course-set prediction with grade prediction. We demonstrate that predicting courses taken in addition to the grades in those courses leads to significant improvements in prediction quality. Trained on ten years of institutional data, our joint prediction model reduces mean absolute error by nearly 50% compared to an identical architecture that predicts grades alone. The model also outperforms traditional LSTM-based sequential models, as well as graph neural network-based approaches, and offers natural ways to incorporate student attribute data. This work demonstrates the utility of modern neural architectures for creating interpretable models that can be adapted to new institutions via retraining and recalibration, as well as the importance of key techniques, such as predicting courses taken during training. We discuss how this model could be incorporated into early detection systems at institutions of higher education.
誰が発言するかが重要: イタリア議会の議事をめぐる権威を意識したマルチビューRAG
議会議事録は民主的審議の主要な記録ですが、その量と断片化により、市民、ジャーナリスト、研究者にとって多視点からのアクセスが困難になっています。検索拡張生成 (RAG) を議会の議事録に適用すると、3 つの特有のリスクが生じます。それは、最も頻繁に発言する人の優位性、話題の専門知識に従って発言者の重み付けができないこと、および政治的に機密性の高いテキストでの引用の誤りです。私たちは、これらのリスクに共同で対処するイタリア下院向けの RAG システムである ParliamentRAG を紹介します。その中心的な貢献は、職業、学歴、以前の介入などの解釈可能な要素を組み合わせて、現在のクエリの関数として各話者の権威を推定するトピック依存の権威モデルです。ユーザーのクエリが与えられると、システムは関連する音声のチャンクを取得し、議題に関連する各議派の専門家を特定し、彼らの見解を総合した要約をサポートする引用文とともに生成します。 ParliamentRAG は、自動化された指標と 6 人のドメイン専門家によるブラインド A/B 人間による評価を組み合わせた 2 レベルのプロトコルを介して、15 の政策トピックについて Google NotebookLM に対して評価されます。このシステムは、政治団体全体でのより高いカバレッジ (0.97 対 0.95)、引用の完全な忠実性 (1.00 対 0.95)、出典関連の側面でのより強力な専門家の選好を実現しますが、NotebookLM は散文指向の側面で引き続き強力です。
原文 (English)
Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings
Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make multi-perspective access difficult for citizens, journalists, and researchers. Applying Retrieval-Augmented Generation (RAG) to parliamentary transcripts introduces three specific risks: dominance of the most frequent speakers, inability to weight speakers according to topical expertise, and citation misattribution in politically sensitive text. We present ParliamentRAG, a RAG system for the Italian Chamber of Deputies that addresses these risks jointly. Its core contribution is a topic-dependent authority model that estimates each speaker's authority as a function of the current query, combining interpretable components such as profession, education, and previous interventions. Given a user query, the system retrieves relevant speech chunks, identifies topic-relevant experts across parliamentary groups, and generates a summary synthesizing their perspectives, accompanied by supporting quotations. ParliamentRAG is evaluated against Google NotebookLM on 15 policy topics via a two-level protocol combining automated metrics and blind A/B human evaluation by six domain experts. The system achieves higher coverage across political groups (0.97 vs. 0.95), perfect quotation faithfulness (1.00 vs. 0.95), and stronger expert preferences on source-related dimensions, while NotebookLM remains stronger on prose-oriented dimensions.
最終スコアを超えて: 長期的な AI 研究開発のためのエージェントの体系的な評価
自律エージェントは、長期的な実験を通じてモデル、システム、その他の技術成果物を改善できるようになってきています。ただし、この機能の現在の状態を理解するには、最終スコアを超えた評価が必要です。スコアでは、進歩がどこで得られるか失われるかは明らかにされず、蓄積された経験が後の決定を改善するかどうかも示されません。したがって、我々は、ルールベースのメトリクスを使用して、ソリューションのフレーミング、実行、フィードバック制御を通じて実行内の動作を特徴付ける新しいフレームワークに基づいて、36 の長期タスクに関する 7 つのフロンティア モデルの体系的な評価を提示します。また、タスク内およびタスク間でのエクスペリエンスの再利用を評価するための制御された比較を行います。その結果、現在のエージェントは完全に自律的な研究者というよりも、エンジニアリングのオプティマイザーのように動作することがわかりました。エージェントは実用的なソリューションを定式化して実装することができますが、そのパフォーマンスは実行ごとに大きく異なり、最も強力なソリューションは主に確立された技術を適応または組み合わせており、真の方法論的な新規性は依然としてまれです。詳細な分析により、観察されたパフォーマンスは、同様の最終結果の背後にある明確なプロセスのボトルネック、その後の意思決定に役立つまたは誤解を招く可能性があるエクスペリエンスの再利用、パフォーマンスの安定性に影響を与えるハーネス設計など、複数の要因によって形成されることが明らかになりました。これらの調査結果は、モデルのトレーニング、推論時間戦略、エクスペリエンス管理、ハーネス設計を改善するための具体的な方向性を示唆しています。
原文 (English)
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.
SLM とエッジ コンピューティングによる仮想エージェントの強化: 思考と記憶のプロセスの探索的評価
身体化されたインテリジェントな仮想エージェントは、複雑な仮想世界およびメタバース世界内で、永続的で適応性のあるコンテキスト認識型のエンティティとして動作することが期待されています。ただし、そのような環境に認知機能のあるエージェントを実装することは、概念的にも技術的にも困難です。さまざまな青写真と開発アプローチの中で、認知身体化エージェント アーキテクチャ (CEAA) は、知覚、記憶、推論、計画、身体化されたアクションのコンポーネントを設計するための実装指向のフレームワークとして開発されました。エッジ コンピューティングと生成 AI 言語モデルの最近の進歩を考慮して、この論文では、インタラクティブな仮想世界での仮想エージェントの認知オーケストレーションと永続性の中心となるプロセスとしての「思考」と「記憶」に焦点を当て、選択された CEAA コンポーネントのエッジベースの操作をサポートするための小型言語モデル (SLM) の使用について検討します。エッジベースの仮想エージェント ゲートウェイ システムは、さまざまなサイズの Qwen2.5 モデルを使用して NVIDIA Jetson Orin NX 上で開発および評価され、サービス リクエストを処理し、メモリ駆動型の会話を処理するシステムの機能を調査しました。一連のシミュレーション実験では、ルーティング精度、メモリ読み取りパフォーマンス、レイテンシを評価し、選択された CEAA プロセスを部分的に実装する SLM 主導のプロトタイプ エージェント システムを実証しました。このシステムは、認知「脳」が効率的かつ状況に応じて動作し、没入型の仮想世界でインタラクティブな体験を実現できる身体化エージェントの開発をサポートします。
原文 (English)
Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes
Embodied intelligent virtual agents are expected to operate as persistent, adaptive, and context-aware entities within complex virtual and Metaverse worlds. However, implementing cognitively capable agents in such environments is conceptually and technologically challenging. Among a range of blueprints and development approaches, the Cognitive Embodied Agent Architecture (CEAA) has been developed as an implementation-oriented framework for architecting components of perception, memory, reasoning, planning, and embodied action. Considering the recent advances in edge computing and generative AI language models, this paper explores the use of Small Language Models (SLMs) to support edge-based operation of selected CEAA components, focusing on "Think" and "Memory" as processes central to cognitive orchestration and persistence of virtual agents in interactive virtual worlds. An edge-based virtual agent gateway system was developed and evaluated on an NVIDIA Jetson Orin NX using Qwen2.5 models of different sizes, exploring the system's capability to process service requests and handle memory-driven conversations. A series of simulation experiments evaluated routing accuracy, memory-read performance, and latency, demonstrating an SLM-driven prototype agent system that partially implements selected CEAA processes to support the development of embodied agents whose cognitive "brain" can operate efficiently and contextually for interactive experiences in immersive virtual worlds.
RAIL: 人工知能の準備レベルの自動分類器
人工知能テクノロジーの成熟度の評価は、投資決定、プロジェクト管理、政策監視に不可欠ですが、利用可能な準備フレームワークは異種混合であり、自動的に適用することが困難です。テクノロジー準備レベルの AI への適応には AI 固有のゲート基準が欠如し、機械学習テクノロジー準備レベルは内部プロセス成果物へのアクセスを前提とし、AI/データ準備ディメンション モデルは直接比較しにくいスケールを採用しています。この論文は 2 つの貢献を行っています。まず、これら 3 つのフレームワークを Unified AI Readiness Level (AIRL) に統合します。AIRL は、環境証拠のはしごに基づいて構築され、一般性を固定するルールと明示的な割り当て規律とともに次元の上限 (仕様、データの存在、データの品質、データの合法性、専門知識、アルゴリズムの成熟度をカバーする) によって補完された 9 レベルの順序スケールです。これにより、準備レベルが作業の自然言語記述のみから決定可能になります。第二に、スケールを運用可能にする専門家パネルによる分類器である RAIL (独立 LLM 専門家による準備評価) を提案します。1 つの証拠エージェントと 6 つの独立したディメンション エージェントで、それぞれが狭い範囲の権限を持つ大規模な言語モデルであり、決定論的な最小ルールが集約された評決を下し、主任専門家が非対称権限の下でレビューし、パネルの推奨事項を確認または引き下げますが、上限を超えることはありません。この方法は、一貫性を示し、モノリシック LLM 分類器からの過大評価を回避することを示すいくつかの研究成果の分析でテストされました。
原文 (English)
RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level
Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct comparison. This paper makes two contributions. First, we unify these three frameworks into the Unified AI Readiness Level (AIRL), a nine-level ordinal scale built on an environmental evidence ladder and complemented by dimensional caps (covering specification, data existence, data quality, data legality, expert knowledge, and algorithmic maturity) together with a generality-anchoring rule and explicit assignment disciplines, so that a readiness level becomes decidable from a natural-language description of the work alone. Second, we propose RAIL (Readiness Assessment via Independent LLM-experts), a panel-of-experts classifier that operationalizes the scale: one evidence agent and six independent dimension agents, each a large language model with a narrowly scoped mandate, deliver verdicts that a deterministic minimum rule aggregates and a chief expert reviews under asymmetric authority, confirming or lowering the panel's recommendation but never raising it above the caps. The method was tested in the analysis of several research works showing consistency and avoiding overestimation from monolithic LLM classifiers.
Academic League of Artificial Intelligence - 教育、研究、普及の統合的な視点
学術リーグは、課外教育を促進し、大学と社会の統合を強化するための重要なメカニズムとなっています。この文書では、サンタカタリーナ連邦大学 (UFSC) の人工知能学術連盟 (LIA) が採用した組織枠組みについて説明します。この組織枠組みは、学生中心のプロジェクトベースのアプローチを通じて、教育、研究、大学の拡張を統合するように設計されています。このフレームワークは、民主的なガバナンス、共同学習、ダイナミックなプロジェクト組織を組み合わせて、技術的能力と横断的な能力の両方を育成します。このフレームワークは、競技チーム、研究グループ、公開講座、ナレッジ リポジトリ、社会的影響力を持つ AI を活用したアプリケーションなどの代表的な取り組みを通じて説明されています。これらのプロジェクトは、リーダーシップ、科学的生産、コミュニティへの関与、知識の保存を促進しながら、共通の組織構造内で多様な教育、科学、普及活動をどのように展開できるかを実証しています。報告された経験は、提案されたフレームワークが大学の 3 つの柱をエンジニアリングおよびコンピューティング教育に統合するための柔軟で再現可能なモデルを提供し、学術リーグや同様の学生団体に実践的なガイダンスを提供することを示しています。
原文 (English)
Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension
Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between universities and society. This paper presents the organizational framework adopted by the Academic League of Artificial Intelligence (LIA) at the Federal University of Santa Catarina (UFSC), designed to integrate teaching, research, and university extension through a student-centered, project-based approach. The framework combines democratic governance, collaborative learning, and dynamic project organization to foster both technical and transversal competencies. The framework is illustrated through representative initiatives, including competition teams, study groups, open lectures, knowledge repositories, and AI-powered applications with social impact. These projects demonstrate how diverse educational, scientific, and extension activities can be developed within a common organizational structure while promoting leadership, scientific production, community engagement, and knowledge preservation. The reported experience indicates that the proposed framework provides a flexible and replicable model for integrating the three university pillars into engineering and computing education, offering practical guidance for academic leagues and similar student organizations.
因果世界モデルに関する統一的な視点: 観察から表現、構造まで
ワールド モデル (WM) は、トレーニングの分布を超えて予測、計画、行動できるインテリジェント エージェントの基盤としてますます注目されています。この論文では、知覚観察から環境ダイナミクスを支配する構造の概念的表現の構築に至るまで、複数の抽象レベルにわたって因果関係の観点から WM を研究します。私たちは、有用な WM は生成機能だけを超えたものでなければならないと主張します。WM は、システムのダイナミクスを決定し説明するエンティティのプロパティ、エンティティ間の相互作用、およびエンティティと環境の相互作用もキャプチャする必要があります。私たちは、サポート対象のタスクに基づいた因果 WM (CWM) の正式な定義を提供し、世界モデリングを因果表現学習、オブジェクト中心学習、因果発見、構造因果モデル、モデルベースの意思決定における既存の研究と結び付けます。最後に、CWM を識別可能性に関する文献に関連付け、WM のコンポーネントがいつデータから復元できるか、およびどの程度の同等性までを明確にします。これにより、WM を因果推論と情報に基づいた意思決定をサポートする表現と構造に根付かせることができます。
原文 (English)
A Unifying Perspective on Causal World Models: From Observations to Representations to Structure
World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics. We argue that useful WMs must go beyond generative capabilities alone: they should also capture entity properties, entity-to-entity interactions, and entity-to-environment interactions that determine and explain the dynamics of a system. We provide a formal definition of Causal WMs (CWMs) grounded in the tasks they are intended to support, connecting world modelling with existing work in causal representation learning, object-centric learning, causal discovery, structural causal models, and model-based decision-making. Finally, we relate CWMs to the literature on identifiability, clarifying when the components of a WM can be recovered from data and up to which equivalence. With this, we ground WMs in representations and structures that support causal reasoning and informed decision-making.
MARC v1: 臨床 AI 推論と調整のためのオープンソース マルチエージェント フレームワーク
我々は、臨床推論のためのモノリシック LLM プロンプトを決定論的なマルチエージェント オーケストレーションに置き換えるオープンソース フレームワークである Multi-Agent Reasoning and Coordination (MARC) を紹介します。 MARC は、明示的なコンテキストの受け渡しと追跡可能な中間出力を使用して、抽出、推論、回答生成、および評価を行う役割に特化したエージェントを調整し、段階的な障害の属性を可能にします。さらに、平易な言語の説明からタスク固有のエージェント プロンプトを生成する Decomposer モジュールを導入し、手動によるプロンプト エンジニアリングを排除します。このフレームワークは、API ベースのデプロイメントとローカル CPU 互換のデプロイメントの両方をサポートしており、コードを変更することなく YAML 経由で完全に構成可能です。 MARC は、モデルに依存せず、解釈可能で、プログラミングの専門知識がなくても臨床分野の専門家がアクセスできるように設計されています。完全なフレームワークは https://github.com/Penn-RAIL/MARC-v1 で入手できます。
原文 (English)
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination
We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. We additionally introduce a Decomposer module that generates task-specific agent prompts from a plain-language description, eliminating manual prompt engineering. The framework supports both API-based and local CPU-compatible deployments and is entirely configurable via YAML, without code modifications. MARC is designed to be model-agnostic, interpretable, and accessible to clinical domain experts without programming expertise. The full framework is available at https://github.com/Penn-RAIL/MARC-v1.
AlayaWorld: インタラクティブな長距離世界モデリング - 完全な技術レポート (v1.1)
このレポートでは、AlayaWorld の改良版を紹介します。バックボーン アーキテクチャ、チャンク単位の自己回帰生成スキーム、トレーニング データは以前のリリースから変更されていませんが、コンディショニング信号の表現方法とモデルへの統合方法が大幅に改訂されました。新しい設計は、単純な原則に基づいています。つまり、コンディショニング信号は、潜在表現と時間構造の両方において、生成されたコンテンツと可能な限り厳密に一致する必要があります。この目的のために、2 つの大きな変更を加えます。まず、以前のデプスワーピングベースの空間メモリをストリーミング 3D ポイントキャッシュ レンダラーに置き換えます。次に、視覚条件が同じ因果 VAE 潜在空間にエンコードされ、時間統計が生成されたビデオの統計と一致するように、調整パイプラインを再設計します。具体的には、新しいバージョンでは次の 6 つの変更が導入されています。(1) 静的フレーム画像コンディショニングを動きを意識した潜在コンディショニングに置き換えます。 (2) 再レンダリングされた空間メモリを連続シーケンスとして因果的にエンコードする。 (3)ピクセル空間内で一時記憶ウィンドウを位置合わせする。 (4) メモリ トークンをゼロにするのではなく、メモリ トークンを削除するハード メモリ ドロップアウトを採用します。 (5) トレーニングと推論にわたって VAE エンコードおよびデコード プロトコルを統合する。 (6) カメラの AdaLN ブランチを削除し、視点制御が完全に再レンダリングされた空間条件によって提供されるようにします。
原文 (English)
AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent representation and temporal structure. To this end, we make two major changes. First, we replace the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer. Second, we redesign the conditioning pipeline so that visual conditions are encoded in the same causal-VAE latent space, with temporal statistics consistent with those of the generated video. Concretely, the new version introduces six modifications: (1) replacing static-frame image conditioning with motion-aware latent conditioning; (2) causally encoding re-rendered spatial memory as a continuous sequence; (3) aligning the temporal-memory window in pixel space; (4) adopting hard memory dropout that removes memory tokens rather than zeroing them; (5) unifying the VAE encoding and decoding protocol across training and inference; and (6) removing the camera AdaLN branch, such that viewpoint control is provided entirely through the re-rendered spatial condition.
QuoteBench: スコアの一致によってコマンド パスの障害がどのように隠蔽されるか
LLM コーディング エージェントは、モデル出力をシリアル化、ラップ、再解析するインターフェイスを通じて Bash コマンドを発行します。一致した実行スコアだけでは、コマンド生成エラーと生成後に発生したエラーを区別できません。 QuoteBench は、14 のインシデント派生ファミリーからの 56 のワンショット タスクに対する正確な最終状態の検証によってこの境界を測定し、意図的にエスケープされない追加された 1 つのパーサーを中心とした実行トランスポートとの生成コントラクトを横断します。補間ポイントでのエスケープは、再生された各応答の生のパス結果を再現するため、開示された境界の下での回復は、世代を変更するモデルから行われる必要があります。 8 つの同じウィンドウ構成で、追加されたパーサーを通じて同じ応答を再生すると、成功率が 55.4 ~ 73.2 パーセント ポイント低下します。開示は 6 つの構成で 30.4 から 60.7 ポイント回復し、他の 2 つの構成ではゼロまたはわずかにマイナスになります。生の生成はフロンティアではほぼ飽和状態です。境界適応は依然としてモデルを分離するものです。 GPT-5.6-sol の -3.6 ポイントの一致ギャップにより、-64.3 ポイントのダメージと +60.7 ポイントの補正が隠蔽されます。デプロイメント構成によりモデルの順序が変更されます。26 の比較可能なペアのうちの 1 つの反転は明確で、残りの 4 つは単一タスクのマージンにあります。コマンド発行エージェントの評価では、一致したスコアをモデル固有のプロパティとして扱うのではなく、モデル構成、生成コントラクト、実行パス、操作点、および最終状態のバリデーターを報告する必要があります。
原文 (English)
QuoteBench: How Matched Scores Can Hide Command-Path Failures
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models. GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation. The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous and four more sit on single-task margins. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.
OmniScientist: オムニモード、オムニ専門分野の AI 科学者
基礎モデルの最近の進歩により、AI 科学者は、仮説の生成からコードの実行、原稿の準備に至るまで、ますます完全な研究ワークフローを自動化できるようになりました。しかし、ワークフローの範囲だけでは、科学的発見が依存する完全な証拠にアクセスできるわけではありません。既存のシステムは通常、テキスト、コード、ラベル、または事前に計算された概要を推論し、科学的に決定的な空間的、時間的、クロスチャネル、および手続き上の関係をエージェントが利用できないままにしています。異種の生の証拠から直接学際的な研究を行う、エンドツーエンドのオムニモーダル AI 科学者である OmniScientist を紹介します。認識層と、着想、実験、書き込みのための 3 つの自律エージェントが決定論的なパイプライン内で動作し、観察によって研究のライフサイクル全体を通じて研究の疑問、実験上の決定、最終的な主張を形作ることができます。コード内でアイデア、厳密さ、およびクレームのチェックを実行することにより、システムは新規性のスクリーニング、統計的妥当性、実行の来歴、および数値的なトレーサビリティを強制します。私たちは、5 つの分野ファミリー、4 つの科学的証拠ファミリー、および画像、信号、音声、ビデオ、3D 構造、軌跡、表、数式、グラフを含むモダリティにわたる 36 の実データ ケースに基づいて OmniScientist を評価します。このシステムは、36 件すべてのケースで生データからコンパイルされた原稿までの完全なパスを完了し、参照推論バックボーンにより平均全体論文スコア 6.3 を達成しました。事前に計算されたスカラー特徴のみを受け取るブラインド バリアントとの一対の比較では、直接知覚が 7 つの評価次元すべてを改善し、直接的な判断の 85% で勝利します。これらの結果は、証拠に基づいた科学的発見にはライフサイクル全体の認識が不可欠であり、幅広い能力を持つ AI 科学者への実践的な道を提供することを示しています。
原文 (English)
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.
言語モデル時代の AI アカウンタビリティ エコシステム
この記事では、責任エコシステムに基づいた AI における責任のフレームワークをレビューし、更新します。私たちは、一般向けに大規模言語モデルがリリースされて以来の最新の開発を考慮してフレームワークを更新しています。私たちは、元の AI アカウンタビリティ エコシステムに対して 3 つの相互に関連したアップデートを提案します。(i) アカウンタビリティ エコシステムの方向性を AI インフラストラクチャとサプライ チェーンに再設定すること、(ii) 分散型システムの改善をサポートする結果の監視と問題の特定をより重視すること、(iii) 現実の言語モデルの予測不可能性という新たなリスクを考慮して、エンドユーザーのアカウンタビリティを組み込むことです。まとめると、これらのアップデートは、フロンティア AI アプリケーションが業界固有の監視を持つ単一の識別可能な主体によって制御される個別の製品としてモデル化できるシステムから、分散型、継続的、制度化された説明責任への移行を示しています。
原文 (English)
The AI Accountability Ecosystem in the Era of Language Models
This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in light of the latest developments since the release of Large Language Models for general public use. We propose three interlinked updates to the original AI accountability ecosystem: (i) reorienting the accountability ecosystem to AI infrastructure and supply chains, (ii) providing greater emphasis on outcomes monitoring and identification of issues that support decentralized system improvement, and (iii) incorporating end-user accountability given the new risks of unpredictability of language models in-the-wild. Collectively, these updates mark a shift towards accountability as distributed, continuous, and institutionalized, away from a system in which frontier AI applications can be modeled as discrete products controlled by single identifiable actors with industry-specific oversight.
LLM は制約を知っているが使用していない: 実用的な制約推論におけるアクティベーションのボトルネック
顕著な表面キューが暗黙的な実現可能性制約と競合する場合、LLM は失敗することがよくありますが、集計精度により、真の制約推論と保守的なデフォルトが混同されます。この区別を条件付き制約のアクティブ化として形式化します。制約は、制約の存在と非存在のプロンプト間で対称的に内部的にエンコード (ナレッジ) されます (対称) が、決定にルーティングされる場合 (ルーティング) と、ドナーのアクティブ化 (修復) によって修復できる場合のみです。 14 モデルにわたるカルテット診断により、2 つの故障モードが明らかになります。 2 つのオープンウェイトのプローブは $88\%$ を超える制約をデコードしますが、アクティベーション パッチによって一方 ($+6.4$ nats) は修復され、もう一方 ($-0.07$) は修復されません。緩和のフロンティアでは、修復コーナーに到達する促進的な介入はありません。すべてが単一の仲介経路を通じて保守的なバイアスを増大させます-前提条件の言及。隠れた制約の障害は、知識の問題ではなく、ルーティングの問題です。
原文 (English)
LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning
When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional constraint activation: the constraint is internally encoded (Knowledge) symmetrically across constraint-present and -absent prompts (Symmetry), yet only sometimes routed into the decision (Routing) and repairable by a donor activation (Repair). A quartet diagnostic over 14 models reveals two failure modes; probes on two open weights decode the constraint above $88\%$, yet activation patching repairs one ($+6.4$ nats) and not the other ($-0.07$). On a mitigation frontier, no prompted intervention reaches the repair corner: all inflate conservative bias through a single mediation pathway -- prerequisite mention. Hidden-constraint failure is a routing problem, not a knowledge problem.
LLM の内省を促すものは何ですか?武力紛争予測における不確実性ルートの制御された除去
内省は LLM 推論を改善すると広く考えられていますが、どのコンポーネントがゲインを駆動するのかはまだよくわかっていません。我々は、LLM の内省の 4 つの要素、つまり証拠の露出、診断の足場、分類語彙、およびアクションのルーティングを分離する、制御された 6 条件のアブレーションを紹介します。 2 つの正確な null 結果が 1 つのメカニズムに収束します。まず、構造化された診断質問は、非構造化リフレクションに比べて測定可能な価値を追加しません ($\text{F1} = 0.296$ vs $0.297$、$p = 1.000$、95\% CI $[-0.041, +0.040]$)。第 2 に、アクション空間を 1 つの一般的なアクションに折りたたんで完全な不確実性分類を提示しても、値は追加されず ($\Delta\text{F1} = +0.008$、95\% CI と重複)、メカニズムとして分類語彙が除外されます。型指定されたアクションのルーティングにより、一貫した方向性のゲインが得られます ($\text{F1} = 0.379$ 対 $0.296$)。分類語彙を制御する控えめな推定値は $\Delta\text{F1} = +0.075$ で、単発ベースラインを超える全体的なゲインはブートストラップ CI によって有意です ($\Delta\text{F1} = +0.101$、95\% CI $[+0.020, +0.185]$)。語彙ルーティングの分解は GPT-4o 上で再現されます。分類語彙は一般的な反映に比べて有意な価値を追加しません ($p = 0.773$) が、アクション ルーティングは大幅な利益をもたらします ($p = 0.025$)。このメカニズムがバックボーン全体で保持されることが確認されています。利益は構造的に新しい紛争に集中しています。ミャンマー ($\text{F1}: 0.000 \rightarrow 0.353$) とウクライナ ($0.167 \rightarrow 0.500$) では、語彙のみの状態は一般的な反射のみを回復しますが、アクション ルーティングは退化した事前状態を打ち破ります。これらの発見は、診断足場や分類語彙ではなく、型付きアクション ルーティングがメタ認知 LLM 予測エージェントの有望な設計原則であることを特定し、同時に競合類型全体にわたる大規模な評価を動機付けるものです。
原文 (English)
What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting
Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a controlled six-condition ablation isolating four components of LLM self-reflection: evidence exposure, diagnostic scaffolding, taxonomy vocabulary, and action routing. Two precise null results converge on a single mechanism. First, structured diagnostic questions add no measurable value over unstructured reflection ($\text{F1} = 0.296$ vs $0.297$, $p = 1.000$, 95\% CI $[-0.041, +0.040]$). Second, presenting the full uncertainty taxonomy while collapsing the action space to a single generic action also adds no value ($\Delta\text{F1} = +0.008$, overlapping 95\% CIs), ruling out taxonomy vocabulary as the mechanism. Typed action routing provides consistent directional gains ($\text{F1} = 0.379$ vs $0.296$); the conservative estimate controlling for taxonomy vocabulary is $\Delta\text{F1} = +0.075$, and the overall gain over the single-shot baseline is significant by bootstrap CI ($\Delta\text{F1} = +0.101$, 95\% CI $[+0.020, +0.185]$). The vocabulary-routing decomposition replicates on GPT-4o: taxonomy vocabulary adds no significant value over generic reflection ($p = 0.773$), while action routing provides significant gains ($p = 0.025$), confirming the mechanism holds across backbones. Gains concentrate on structurally novel conflicts: in Myanmar ($\text{F1}: 0.000 \rightarrow 0.353$) and Ukraine ($0.167 \rightarrow 0.500$), the vocabulary-only condition recovers no more than generic reflection while action routing breaks the degenerate prior. These findings identify typed action routing -- not diagnostic scaffolding or taxonomy vocabulary -- as a promising design principle for metacognitive LLM forecasting agents, while motivating larger-scale evaluation across conflict typologies.
なぜ AI エージェントはルールを破るのか?フレーミング、コンテキスト、ソーシャルシグナルがコンプライアンスをどのように形成するか
罰則を指定すると、逆説的に、法的義務が違反に有利な費用対効果の計算に変換される可能性があります。私たちは、この施行情報のパラドックスが AI エージェントで体系的に発生することを実証します。ほとんどの AI 安全性評価ではモデルが失敗するかどうかがテストされますが、私たちは法律と経済学のコンプライアンス理論を診断ツールとして適用して、その理由を調査します。我々はコンプライアンス理論を比喩としてではなく経験的仮説として扱い、それぞれが異なるモデルクラスの動作を予測することを示します。私たちは、エンタープライズ調達チャットボットとして動作する 12 の命令調整された言語モデルにわたって仮説を評価します。抑止力、正当性、表現法則の理論に基づいて、安全性を細かく調整したモデルは広範なコンプライアンスを維持する一方、タスク最適化モデルやエージェントモデルは規制シグナルを単なる最適化パラメーターとして扱うことを示します。これらの後者のモデルは、低い執行罰や非命令語句など、理論によって予測される条件の下では準拠できません。すべてのモデルにおいて、金銭的インセンティブ、管理上の要求、同僚の成果、または従業員のプレッシャーの導入は、大規模なコンプライアンス違反を引き起こします。 AI 調達エージェントは、標準的な連携ベンチマークでは捉えられない方法で、ローカル ユーザーの目的を満たすために、組織的に規制上の制約に違反しています。結局のところ、ルールの埋め込みだけではコンプライアンスを達成することはできません。モデルの選択自体がガバナンスの決定であり、ベンチマークベースの評価はコンプライアンスを重視した展開には不十分です。
原文 (English)
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance
Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI safety evaluations test whether models fail, we investigate why, applying compliance theory from law and economics as a diagnostic tool. We treat compliance theories not as metaphors but as empirical hypotheses and show that each predicts the behavior of a distinct model class. We evaluate our hypotheses across twelve instruction-tuned language models operating as enterprise procurement chatbots. Drawing on theories of deterrence, legitimacy, and expressive law, we show that safety-fine-tuned models maintain compliance broadly, while task-optimized and agentic models treat regulatory signals as mere optimization parameters. These latter models fail to comply under conditions predicted by theory, such as low enforcement penalties and non-command phrasing. Across all models, introducing financial incentives, managerial demands, peer outcomes, or employee pressure produces large compliance failures. AI procurement agents systematically violate regulatory constraints to satisfy local user objectives in ways not captured by standard alignment benchmarks. Ultimately, compliance cannot be achieved by rule embedding alone; model selection is itself a governance decision, and benchmark-based evaluation is insufficient for compliance-sensitive deployments.
AI があなたの牧師の場合: 大規模言語モデルにおける神学的トリアージと司牧指導のベンチマーク
信仰、教義、司牧的ケアの問題について、大規模言語モデル (LLM) に助言を求める人が増えています。これらの質問は通常の情報要求ではありません。キリスト教の中核的な信念について尋ねる人もいれば、忠実な伝統の間での実際の意見の相違について尋ねる人もいます。問題は慎重であるため謙虚さを必要とする人もいます。また、神学的完全性よりも安全と人的紹介が重要である司牧的な状況について尋ねる人もいます。既存のベンチマークはこの構造を評価しません。英語のキリスト教神学トリアージおよび司牧指導コンテキストにおける大規模言語モデルの動作を評価するための 120 のシナリオ ベンチマークである、信仰と道徳の指導ベンチマークである FMG-Bench を紹介します。 FMG-Bench v1 は、8,792 のスコア付けされた応答にわたって 14 の高度なモデルを評価し、生のモデルの動作を 3 つのガイド付き命令設定と比較します。実稼働環境では、構造化ハーネス内にモデルを配置すると、生のモデルの動作よりも平均で +3.96 ポイント改善され、すべてのモデルが改善されました。最も安全性が重要な調査結果は、エスカレーションの適切性、つまり司牧、臨床、法律、または緊急サポートがいつ必要かを AI システムが認識するかどうかが +10.8 ポイント向上したことです。ガイド付き設定により、堅牢性も向上します。つまり、質問が言い換えられたり、プレッシャーをかけられたりした場合の一貫性が向上します (92.88 ~ 98.02 の安定性)。モデルに視点を比較してもらうことは、二次的な教義に関する質問には役立ちますが、一次的な教義や緊急の司牧状況に適用すると逆効果になる可能性があります。ベンチマークは測定ツールであり、司牧当局として AI システムを推奨するものではありません。
原文 (English)
When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models
People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about real disagreements among faithful traditions, some require humility because the issue is prudential, and some are pastoral situations where safety and human referral matter more than theological completeness. Existing benchmarks do not evaluate this structure. We introduce FMG-Bench, the Faith & Moral Guidance Benchmark, a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral guidance contexts. FMG-Bench v1 evaluates 14 advanced models across 8,792 scored responses, comparing raw model behavior with three guided instruction settings. In our production run, placing models inside a structured harness improves over raw model behavior by +3.96 points on average, with every model improving. The most safety-critical finding is a +10.8 point gain in escalation appropriateness -- whether AI systems recognize when pastoral, clinical, legal, or emergency support is needed. The guided settings also improve robustness, meaning consistency when questions are reworded or pressured (92.88 to 98.02 stability). Asking a model to compare perspectives helps in secondary-doctrine questions but can be counterproductive when applied to primary doctrine or urgent pastoral situations. The benchmark is a measurement tool, not an endorsement of AI systems as pastoral authorities.
ネパール語自動音声認識用の多言語事前トレーニング済みモデルの比較分析
多言語の事前トレーニング済みモデルは名目上ネパール語をサポートしていますが、単一の微調整プロトコルの下でそれらを比較した管理されたベンチマークはありません。 OpenSLR SLR54 ネパール語コーパス上で、CTC 自己教師あり自己回帰エンコーダ デコーダ、およびハイブリッド Conformer-CTC アーキテクチャにまたがる 6 つの事前トレーニング済みモデル (XLSR-53、IndicWav2Vec、MMS-1B、Whisper-Medium、Whisper-Large-v3-Turbo、および Conformer-Hi) を同一の前処理を使用して微調整しました (約 165 時間)。スプリット、オプティマイザー、家族に合わせた学習率スケジュール。 3 つの独立したテスト セット (OpenSLR、FLEURS、Common Voice) で単語誤り率 (WER)、文字誤り率 (CER)、およびリアルタイム係数 (RTF) を評価します。 Whisper-Large-v3-Turbo (14.76% WER) と IndicWav2Vec (14.89% WER) は、パラメーター ギャップが 9 倍、事前トレーニング データ ギャップが 40 倍であるにもかかわらずトップで並んでおり、事前トレーニングにおける言語族の近接性がドメイン内ネパール語の生のスケールの代わりになり得るという直接的な経験的証拠を提供しています。 CTC デコーダは、同じ精度で自己回帰 Whisper よりも最大 29 倍高速に実行され、遅延予算がどのような場合でも、実際の導入の優先順位が CTC に切り替わります。大規模多言語事前トレーニング (MMS-1B) では、FLEURS でドメイン外の劣化が最小 (+12.55 pp) であり、ドメイン内のピーク精度ではなくスケールが堅牢性を獲得していることを示しています。結果として得られたベンチマークは、ネパール ASR の最初の標準化されたマルチモデルの効率を意識した参照番号を提供します。
原文 (English)
Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition
Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium, Whisper-Large-v3-Turbo, and Conformer-Hi) spanning CTC self-supervised, autoregressive encoder-decoder, and hybrid Conformer-CTC architectures, on the OpenSLR SLR54 Nepali corpus (~165 hours) using identical preprocessing, splits, optimizer, and family-matched learning-rate schedules. We evaluate Word Error Rate (WER), Character Error Rate (CER), and Real-Time Factor (RTF) on three independent test sets (OpenSLR, FLEURS, Common Voice). Whisper-Large-v3-Turbo (14.76% WER) and IndicWav2Vec (14.89% WER) tie at the top despite a 9x parameter gap and 40x pretraining-data gap, providing direct empirical evidence that language-family proximity in pretraining can substitute for raw scale for in-domain Nepali. CTC decoders run up to 29x faster than autoregressive Whisper at the same accuracy, flipping the practical deployment preference toward CTC under any latency budget. Massively multilingual pretraining (MMS-1B) yields the smallest out-of-domain degradation on FLEURS (+12.55 pp), indicating that scale buys robustness rather than peak in-domain accuracy. The resulting benchmark provides the first standardized, multi-model, efficiency-aware reference numbers for Nepali ASR.
AnchorSIPS: 証拠に裏付けられた精神病リスク症状測定のための合成データセットおよび評価リソース
精神病リスク評価のための AI の進歩は、データアクセスのボトルネックによって制限されています。実際の臨床面接は、プライバシー、ガバナンス、同意の制約により共有することが困難です。我々は、トランスクリプトに基づいた測定対象者との10,000の構造化精神病リスクインタビューの合成データセットであるAnchorSIPSを紹介します。各面接は、臨床医が実施する精神病リスク面接である Mini-SIPS をモデルにしています。これは、病歴、24の症状に関する質問、患者が肯定した項目の追跡証拠、妄想様症状(異常な信念)、幻覚様症状(異常な知覚)、および混乱したコミュニケーションに関する決定、明らかな精神病レベルの症状(「率直な精神病」)の除外、および軽度または初期の精神病症状の高リスク状態である軽度精神病症候群(APS)の最終診断を記録する。 APS 診断は独立したラベルではありません。それは、以前の承認、フォローアップの詳細、症状クラスの決定、および率直な精神病チェックに依存します。すべての中間決定は、それをサポートする転写ターンに固定されています。 AnchorSIPS は、計画後実現パイプラインによって生成されます。隠されたケースシートが患者の臨床状態を特定し、決定論的プランナーがインタビュー構造を修正し、LLM が検証と限定的修復の下で患者の発話のみを認識します。生成前にラベルと構造を修正すると、マルチターン LLM ダイアログに特有のターン間の不一致が回避されます。 7 つの LLM ベースラインにわたって、モデルは大まかな決定を回復しますが、フォローアップの詳細を抽出したり、裏付けとなる転写ターンを引用したりすることができないため、最終ラベルのパフォーマンスはインタビューの能力を過大評価します。 AnchorSIPS は、証拠の抽出、転写に基づいた測定、部分開示の下での不確実性に関する研究を目的としています。
原文 (English)
AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement
Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured psychosis-risk interviews with transcript-grounded measurement targets. Each interview is modeled on Mini-SIPS, a clinician-administered psychosis-risk interview. It captures history, 24 symptom questions, follow-up evidence for items the patient affirms, decisions about delusion-like symptoms (unusual beliefs), hallucination-like symptoms (unusual perceptions), and disorganized communication, exclusion of clear psychotic-level symptoms ("frank psychosis"), and a final attenuated psychosis syndrome (APS) diagnosis, a high-risk state of milder or early psychotic symptoms. The APS diagnosis is not a standalone label. It depends on earlier endorsements, supporting follow-up details, symptom-class decisions, and the frank-psychosis check. Every intermediate decision is anchored to its supporting transcript turns. AnchorSIPS is generated by a plan-then-realize pipeline. A hidden case sheet specifies the patient's clinical state, a deterministic planner fixes the interview structure, and an LLM realizes only the patient utterances under validation and bounded repair. Fixing labels and structure before generation avoids the inter-turn inconsistencies typical of multi-turn LLM dialogue. Across seven LLM baselines, models recover coarse decisions but fail to extract follow-up details or cite supporting transcript turns, so final-label performance overstates interview competence. AnchorSIPS is intended for research on evidence extraction, transcript-grounded measurement, and uncertainty under partial disclosure.
適応型アテンションマッチングによる推論のための思考認識型 KV キャッシュ圧縮
推論言語モデルは、キー/値 (KV) キャッシュが直線的に増加し、デコード中にメモリのボトルネックになる長い思考連鎖 (CoT) シーケンスを生成します。既存の圧縮手法は、推論軌跡をフラットなトークン シーケンスとして扱い、均一な圧縮を適用します。ステップごとに重要性が大幅に異なる CoT 推論の階層構造を無視します。私たちは \textbf{思考認識型注意マッチング (TAM)} を提案します。これは、(i)~軌跡を推論ブロックに分解する思考セグメント化、(ii)~各セグメントの重要性とサイズに基づいて圧縮予算を割り当てる適応型予算割り当て、および (iii)~注目度の高い推論アンカーを保存する重要なトークン保護の 3 つのメカニズムを通じてこの構造を活用します。割り当てルールが凸誤差モデルの下で最適であること、および逐次圧縮の下での累積誤差が制限されたままであることを証明します。 Qwen3-4B を使用した AIME 2024 および MATH-500 での実験では、競合する精度を維持しながら、TAM が同じメモリ フットプリントでの均一な圧縮よりも精度が向上し、定期的な圧縮によりピーク メモリが 3.1 ~ 3.2\,GB (65\% 削減) に制限されることがわかりました。
原文 (English)
Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching
Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding. Existing compaction methods treat reasoning trajectories as flat token sequences and apply uniform compression, ignoring the hierarchical structure of CoT reasoning where different steps vary drastically in importance. We propose \textbf{Thought-Aware Attention Matching (TAM)}, which exploits this structure through three mechanisms: (i)~thought segmentation that decomposes the trajectory into reasoning blocks, (ii)~adaptive budget allocation that assigns compression budget based on each segment's importance and size, and (iii)~pivotal token protection that preserves high-attention reasoning anchors. We prove that the allocation rule is optimal under a convex error model and that cumulative error under sequential compaction remains bounded. Experiments on AIME 2024 and MATH-500 with Qwen3-4B show that TAM improves accuracy over uniform compaction at the same memory footprint, with periodic compaction bounding peak memory to 3.1--3.2\,GB (a 65\% reduction) while maintaining competitive accuracy.
視覚言語モデルは脆弱な多言語連想子である
視覚言語モデルは、視覚的エンティティをテキスト属性に関連付ける必要があります。入力言語が変化したときに、これらの関連付けや概念のバインディングが安定したままであるかどうかは、まだ解明されていません。複数の言語にわたってコンテキストとクエリの言語を変えるベンチマークである M$^2$BIND を紹介します。私たちは、タスクパフォーマンス指標を通じて外因的にバインディングを評価するとともに、因果関係のある介入を通じて本質的にバインディングを評価します。バインディングは言語不変ではないことがわかりました。クロスファミリーおよびクロススクリプト設定は、重大なバインディングの崩壊を引き起こし、モデルの内部バインディング計算が後の層にシフトし、因果関係の強さを失います。密接に関連した言語は、関連性を比較的よく保持します。より広い意味で、私たちの調査結果は、多言語設定でグローバルに展開された VLM が、単一言語の評価で観察されたのと同じ関連性の品質を維持すると仮定できないことを示しています。
原文 (English)
Vision-Language Models are Fragile Multilingual Associators
Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M$^2$BIND, a benchmark varying the language of the context and query across multiple languages. We evaluate binding both extrinsically through task performance metrics and intrinsically through causal interventions. We find that binding is not language-invariant: cross-family and cross-script settings trigger significant binding collapse, with the model's internal binding computation shifting to later layers and losing causal strength. Closely related languages preserve associations comparatively better. In a broader sense, our findings indicate how VLMs deployed globally in multilingual settings cannot be assumed to maintain the same association quality observed in monolingual evaluation.
言語軸の舵取り: 線形解読可能性から因果制御まで
大規模言語モデルの優れた多言語機能にもかかわらず、言語選択を決定する潜在的な力学はまだよく理解されていません。この研究では、言語のアイデンティティが隠れた状態から単に線形に解読可能であるのか、それともコンパクトな活性化方向によって因果的に制御できるのかを問います。 Qwen 3.5-2B や Llama-3.2-1B-Instruct を含む複数のモデル ファミリにわたって徹底的な因果介入分析を実施し、PCA 由来の「言語軸」を分離して、FLORES-200 データセットの 126 万世代にわたるステアリングおよびアブレーション実験を実行します。これらの幾何学的な方向に沿ってステアリングを操作すると、クロススクリプト (英語から中国語) と同一スクリプト (英語からスペイン語) の両方の設定で言語の切り替えが確実に強制されますが、等しい大きさのランダムな摂動では事実上何の効果も生じません。私たちの層別分析により、言語への関与は高度にローカライズされており、明らかに言語ペアに依存していることが明らかになりました。英語から中国語への移行は早期の介入に抵抗し、後の層に容易に移行しますが、英語からスペイン語への移行はより早く移行し、明確な二峰性の感受性を示します。さらに、標的アブレーションにより、英語への根本的な復帰が明らかになります。言語信号が除去されると、入力プロンプトに関係なく、モデルは英語に戻ります。最終的に、これらの発見は、言語決定境界が、方向依存で層固有の因果的にアクティブな特徴として推論中に機能することを示しています。
原文 (English)
Steering the Language Axis: From Linear Decodability to Causal Control
Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood. In this work, we ask whether language identity is merely linearly decodable from hidden states, or if it can be causally controlled by a compact activation direction. We conduct an exhaustive causal intervention analysis across multiple model families, including Qwen 3.5-2B and Llama-3.2-1B-Instruct, isolating PCA-derived "language axes" to perform steering and ablation experiments across 1.26 million generations on the FLORES-200 dataset. Steering along these geometric directions reliably forces language switching in both cross-script (English to Chinese) and same-script (English to Spanish) settings, whereas equal-magnitude random perturbations yield virtually no effect. Our layerwise analysis reveals that language commitment is highly localized and explicitly language-pair-dependent. While English to Chinese switching resists early intervention and steers easily in the later layers, the English-Spanish transition shifts earlier, displaying a distinct, bimodal sensitivity. Furthermore, targeted ablation uncovers a fundamental reversion to English: once the language signal is removed, the model falls back to English regardless of the input prompt. Ultimately, these findings demonstrate that language decision boundaries function during inference as causally active features that are direction-dependent and layer-specific.
StorySpark: ストーリー前提生成のためのモジュールごとの進化的検索
ストーリーの前提は、完全な物語を成長させるための創造的な火花です。しかし、LLM ベースのストーリー生成では、主に後の段階の計画、制御可能性、一貫性、散文の展開が重視されてきましたが、前提レベルのアイデア化は比較的研究されていないままです。ストーリー前提生成のためのモジュールごとの進化的検索フレームワークである StorySpark を紹介します。 StorySpark は、背景、ペルソナ、イベント、エンディング、ツイストなどの解釈可能な物語モジュールを操作し、アクティブな各モジュールを一度埋める静的フィールドとしてではなく、これまでに構築された部分的な前提条件に基づいて条件付けされたローカル検索スペースとして扱います。モジュールごとに、代替案を生成し、コンテキスト内でそれらを評価し、フィードバック駆動の突然変異と組み換えを通じて代替案を洗練し、パレートに基づく選択で補完的な強みを維持し、ブランチカバレッジと有望な方向性のバランスをとるためにフロンティア容量を再割り当てします。マルチビューの自動評価と人間による評価では、StorySpark が競合ベースラインよりも強力な最終前提を生成し、特に一貫して独創性が向上していることが示されています。同じストーリーライターで拡張すると、その前提は、完全性、魅力、多様な使用可能な物語の方向性を維持しながら、より高品質の下流ストーリーにもつながります。
原文 (English)
StorySpark: Module-wise Evolutionary Search for Story Premise Generation
A story premise is the creative spark from which a full narrative can grow. Yet LLM-based story generation has mostly emphasized later-stage planning, controllability, coherence, and prose expansion, while premise-level ideation remains comparatively underexplored. We introduce StorySpark, a module-wise evolutionary search framework for story premise generation. StorySpark operates over interpretable narrative modules such as background, persona, event, ending, and twist, treating each active module not as a static field to fill once, but as a local search space conditioned on the partial premise built so far. For each module, it generates alternatives, evaluates them in context, refines them through feedback-driven mutation and recombination, preserves complementary strengths with Pareto-guided selection, and reallocates frontier capacity to balance branch coverage with promising directions. Multi-view automatic and human evaluations show that StorySpark produces stronger final premises than competitive baselines, with especially consistent gains in originality; when expanded with the same story writer, its premises also lead to higher-quality downstream stories while maintaining completeness, fascination, and diverse usable narrative directions.
理解を伴わない模倣: 大規模言語モデルにおける意思決定バイアスの起源
大規模言語モデル (LLM) は、多くの社会的、感情的、認知的バイアスの影響を受けやすいことが判明しました。私たちは、人間の好み (トレーニング データ内の) に偏りがない場合、または偏っていると正しく分類された場合でも、そのような偏りが生成される可能性がある 2 つのメカニズムを調べました。 1 つ目は、人間の行動に基づいた好みの誤った模倣です。これには、行動が論理的に好みと無関係である場合でも、LLM が人間の好みを推測することが含まれます。 2 つ目は、明らかに偏った人間の行動を模倣することです。経済的バイアスに焦点を当てた 4 つの研究では、ChatGPT-4o と Qwen は、明らかに個人の実際の好みを示さない人間の行動の報告を促された場合でも、社会的証明バイアスを示したことがわかりました。 LLM は、損失回避がバイアスとして明示的に説明された場合にも、損失回避を示しました。実際、詳細な科学的報告を求められた場合、その科学的報告におけるバイアスの程度(つまり、損失回避)は、LLM 自身のその後のバイアスを予測しました。したがって、偏見に関する科学論文は、少なくとも LLM の反応に関しては、自己実現的な予言になる可能性があります。現在の研究は、LLM バイアスを具体化するだけでなく、根底にあるコンポーネント プロセスに光を当てています。
原文 (English)
Mimicry without understanding: the origins of decision bias in large language models
Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training data) are not biased or when they are correctly categorized as being biased. The first is faulty mimicry of preferences based on human behavior: this involves LLMs inferring human preferences even when behaviors are logically unrelated to preferences. The second is mimicry of explicitly biased human behaviors. In four studies focusing on economic biases, we find that ChatGPT-4o and Qwen exhibited social proof biases even when prompted with reports of human behaviors that were clearly non-indicative of individuals' actual preferences. LLMs also displayed loss aversion when it was explicitly described as a bias. Indeed, when prompted with detailed scientific reports, the extent of the bias (i.e., loss aversion) in the scientific report predicted LLMs' own subsequent bias. Scientific papers of biases can thus become self-fulfilling prophecies, at least when it comes to LLMs' responses. The current study goes beyond fleshing out LLM biases and sheds light on the underlying component processes.
StreamReason-Bench: 大規模な言語モデルはイベント時のストリーム処理セマンティクスについて推論できますか?
ストリーミング システムでは、パイプラインの作成、アラートのトリアージ、ログの読み取りなど、大規模言語モデル (LLM) に対する手作業が増えていますが、そのすべてにおいて、モデルがイベント時のストリーム処理がどのように動作するかを認識していることが前提となっています。私たちはその仮定を正面からテストします。 StreamReason-Bench は、モデルにイベント時のストリーム プロセッサの代わりをするように要求します。ウィンドウ化されたクエリと順序が乱れたイベントのストリームが与えられると、どのウィンドウが (集計とともに) 起動し、どのイベントが遅れてドロップされるかを報告する必要があります。回答キーは、Dataflow モデル セマンティクスの小規模なリファレンス実装から得られるため、エンジンを実行せずに、部分クレジット行 F1 を使用して正確に採点できます。タンブリング、ホッピング、セッション、および処理時間ウィンドウをカバーする 600 個の生成されたアイテムでは、モデルはイベント時間でのパフォーマンスが低下します。直接答えるように言われましたが、実際にその指示に従ったモデルで 34% の完全一致をクリアしたモデルはありません。いくつかのモデルでは、思考連鎖 (CoT) が約 2 倍になり (GPT-4o は 0.34 から 0.48 になります)、デフォルトで推論を行うフロンティア モデルは 1 つだけ (0.85) になります。ウォーターマークも遅延もなく、処理時間の制御は、すべての有能なモデルでほぼ解決されます。このギャップは、ウィンドウ処理や演算ではなく、イベント時および遅延データの処理が難しい部分であることを示しています。ウィンドウの種類ごとにエラーを並べ替えると、同じことがわかります。遅延データの間違いはイベント時間ウィンドウを支配してコントロール上で消え、セッション ウィンドウはほとんどの場合、セッション境界に該当する場所で失敗します。
原文 (English)
StreamReason-Bench: Can Large Language Models Reason about Event-Time Stream-Processing Semantics?
Streaming systems increasingly hand work to large language models (LLMs) -- writing pipelines, triaging alerts, reading logs -- and all of it assumes the model knows how event-time stream processing behaves. We test that assumption head-on. StreamReason-Bench asks a model to stand in for an event-time stream processor: given a windowed query and a stream of out-of-order events, it has to report which windows fire (with their aggregates) and which events are dropped as late. The answer key comes from a small reference implementation of Dataflow-model semantics, so we can grade exactly, and with a partial-credit row-F1, without running an engine. On 600 generated items covering tumbling, hopping, session, and processing-time windows, the models do poorly on event time. Told to answer directly, no model that actually follows the instruction clears 34% exact match; chain-of-thought (CoT) roughly doubles that for several of them (GPT-4o goes from 0.34 to 0.48), and only one frontier model that reasons by default comes near solving the set (0.85). A processing-time control, with no watermarks and nothing late, is almost solved by every capable model. That gap points to event-time and late-data handling, not windowing or arithmetic, as the hard part. Sorting errors by window type tells the same story: late-data mistakes dominate the event-time windows and vanish on the control, and session windows mostly fail on where the session boundaries fall.
穴居人から専門アナリストへ: 可変 LLM タスクのエネルギー消費
エネルギー需要の増加と人工知能 (AI) による環境への影響により、AI を活用したデータセンター開発に十分な低コストの電力を供給することに大きな関心が集まっています。これらの課題に対処するための需要側管理の能力に関する研究はさらに限られています。需要の量やタイミングを小売、企業、その他の組織の行動からずらすことは妥当な選択肢ですが、それは需要関連の行動の変化が AI の環境や電力への影響に重要な影響を与える場合に限られます。この記事では、行動可塑性の高い 4 つの小売 (消費者) ユーザーの行動をテストし、技術的な軽減の可能性を評価します。研究では、非推論モデルは十分な品質を提供しながら推論モデルに比べてエネルギー消費量が 20 分の 1 近くに抑えられ、毎日の使用量を想定した場合、米国の少なくとも 141,000 世帯の年間電力需要に等しい量を節約できると結論付けています。単純な即時変更により、非推論モデルを使用してエネルギー消費を最大 65% までさらに削減できます。具体的には、ベースラインとの類似性を最も高く維持する実践により、電力需要が 4 ~ 35% の範囲で削減されます。これは、米国の最大 7,200 世帯の年間電力需要に相当します。 AI の進歩により、環境や電力への影響を正確に見積もることは困難になっていますが、この結果は、大多数のユーザーを対象とした、侵入を最小限に抑えた特定のベスト プラクティスにより、AI によって課せられるエネルギーと環境への負担を軽減できることが確認されました。
原文 (English)
From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks
The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying sufficient low-cost electricity for AI-driven data center development. Research on the ability of demand-side management to address these challenges has been more limited. Shifting the amount or timing of demand from retail, corporate, and other organizational behaviors is a plausible option but only if changes in demand-related behavior have important effects on the envi- ronmental and electricity effects of AI. This article tests four retail (i.e., consumer) user behaviors with high behavioral plasticity to assess their technical abatement potential. The research concludes that non- reasoning models provide sufficient quality while consuming close to one-twentieth of energy compared to reasoning models, saving an amount equal to the annual electricity requirement of at least 141,000 US households under daily usage assumptions. Simple prompt modifications can yield additional reduc- tions in energy consumption by up to 65% using non-reasoning models. Specifically, the practice that maintains the highest degree of similarity with the baseline reduces electricity demand in the range of 4 to 35%, an amount equal to the annual electricity requirement of up to 7,200 US households. Although AI advancements make precise estimates of environmental and electricity impacts difficult to assess, the results confirm that certain minimally intrusive best practices aimed at the majority of users can reduce the energy and environmental burdens imposed by AI.
GenAI 時代の評価デザイン: 生徒の AI リテラシー、学習成果、振り返りをテストするための X1-X2-X3 評価パターン
生成人工知能 (GenAI) は、特にほとんど労力をかけずにもっともらしい回答が得られる技術的テーマにおいて、教師なしオンライン評価の妥当性に疑問を呈しています。この論文では、大規模な学部 2 年生データベース システム モジュールにおける AI を認識した AI テスト評価の設計と実装から得た教訓を報告します。この設計では、リンクされた 2 つの要素を組み合わせました。(1) 学生が情報源となった回答を文書化し、独自の回答を作成し、情報源となった出力を評価する、構造化された 3 部構成の回答形式 (X1-X2-X3)。 (2) AI を意識した質問設計プロセス。草案タスクは最新の GenAI ツールに対してストレス テストされ、一般的なプロンプトが表面的に適切な回答を生成した場合に修正されます。このアカウントは、アーカイブされた評価資料、ルーブリック、計画記録、設計時の GenAI トライアル、実践と反応のデータ、達成記録、および外部レビューのコメントを活用しています。その主な貢献は、測定された学習効果の主張ではなく、再利用可能な評価設計手法です。繰り返しを通じてパターンがどのように発展し、それが本物の評価、目に見える AI リテラシー、生徒の判断、より透明性の高い採点をどのようにサポートできるかを示します。この論文は、学生の不正行為を罰するのではなく、AI リテラシーのテストに焦点を当て、評価を日常的な GenAI の使用に適応させる講師向けの実践的なガイダンスを提供します。
原文 (English)
Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection
Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort. This paper reports lessons from designing and implementing an AI-aware, AI-testing assessment in a large second-year undergraduate database systems module. The design combined two linked elements: (1) a structured three-part response format (X1-X2-X3) in which students documented a sourced answer, produced their own answer, and evaluated the sourced output; and (2) an AI-aware question-design process in which draft tasks were stress-tested against contemporary GenAI tools and revised when generic prompting produced superficially adequate answers. The account draws on archived assessment materials, rubrics, planning records, design-time GenAI trials, practice-response data, attainment records, and external review comments. Its main contribution is a reusable assessment-design method rather than a claim of measured learning gains. We show how the pattern developed across iterations and how it can support authentic assessment, visible AI literacy, student judgement, and more transparent marking. The paper offers practical guidance for lecturers adapting assessment to routine GenAI use, focusing on testing AI literacy rather than penalising students for misconduct.
AI ガバナンス フレームワークの導入が難しい理由: NIST AI RMF の役割ベースのストレス テスト
AI ガバナンス フレームワークは、実際にはガバナンスにならずに、形式的に知り、使用し、実装することができます。このペーパーでは、消費者金融における NIST 人工知能リスク管理フレームワーク (AI RMF) の役割ベースのストレス テストを通じて、その問題を検証します。私たちは、フレームワークの導入をガバナンスの変換問題として扱います。つまり、RMF 言語が、ガバナンスに見える成果物を生成するのではなく、使用中の AI システムに対して、ロールで使用可能でクロスレベルで権限に関連したガバナンスになれるかどうかです。この研究では、LLM ベースの役割シミュレーションを構造化分析プローブとして使用します。 4 つの $\times$ 2 $\times$ 3 の設計を 4 つの組織の役割、2 つの AI 導入、および 3 つのガバナンスのハード ケースに適用し、120 のスコアリングされた回答を生成しました。結果は、ローカル翻訳が主な問題ではないことを示しています。シミュレートされたアクターは一般に、割り当てられた役割を理解し、RMF をローカル アクティビティに変換しました。さらに難しい問題は、その活動がガバナンスの価値となるかどうかでした。アクターの役割は、クロスレベルのガバナンス価値、権限とのつながり、ガバナンスの翻訳可能性、およびガバナンス価値と強く関連していました。導入は構造的適合と強く関連していました。RMF は、ワークフローに組み込まれた LLM 引受コパイロットよりも、有界 ML 引受モデルにより正確に適合しました。リスク軽減はさらに困難でした。これはガバナンス値が存在し、構造的適合性が満たされている場合にのみ表示されますが、どちらの条件もそれだけでは十分ではありません。この論文は、フレームワークベースの AI ガバナンスの診断説明に貢献します。フレームワークは、組織が使用中の AI システムのリスクを確認、解釈、エスカレーション、承認、修正するのに役立つときに価値を生み出します。また、既存の証拠経路、権限構造、システム境界の下でのガバナビリティの限界を明らかにするときにも価値を生み出します。
原文 (English)
Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF
AI governance frameworks can be known, used, and implemented in form without becoming governance in practice. This paper examines that problem through a role-based stress test of the NIST Artificial Intelligence Risk Management Framework (AI RMF) in consumer lending. We treat framework adoption as a governance translation problem: whether RMF language can become role-usable, cross-level, authority-connected governance over the AI system-in-use, rather than producing governance-looking artifacts. The study uses LLM-based role simulation as a structured analytic probe. We apply a 4 $\times$ 2 $\times$ 3 design across four organizational roles, two AI deployments, and three governance hard cases, producing 120 scored responses. Results show that local translation was not the main problem. Simulated actors generally understood their assigned roles and translated the RMF into local activity. The harder problem was whether that activity became governance value. Actor role was strongly associated with Cross-Level Governance Value, Authority Connection, Governance Translatability, and governance value. Deployment was strongly associated with Structural Fit: the RMF fit a bounded ML underwriting model more cleanly than a workflow-embedded LLM underwriting copilot. Risk reduction was harder still. It appeared only when governance value was present and Structural Fit was full, but neither condition was sufficient by itself. The paper contributes a diagnostic account of framework-based AI governance. Frameworks create value when they help organizations see, interpret, escalate, authorize, and correct risk in the AI system-in-use. They also create value when they reveal limits of governability under existing evidence paths, authority structures, and system boundaries.
AIコーディングエージェントの研究から人間が欠けている
AI コーディング エージェントの研究における最近の進歩により、大規模なコードベースの編集から長期的な開発ワークフローの実行に至るまで、複雑なソフトウェア エンジニアリング タスクを自律的に実行するエージェントの能力が急速に向上しました。しかし、これらのシステムが進歩するにつれて、実用性への主なボトルネックは、純粋なタスク解決能力から、ユーザーがエージェントとどのように通信し、監督し、信頼するかという課題へとますます移行しています。この意見書では、自律的なコーディング エージェントから人間中心のコーディング エージェントへの方向転換、つまりタスクを完了するだけでなく人々と効果的に連携するように設計されたシステムへの方向転換を主張します。私たちは、ヒューマン エージェントのタスク解決ループを特徴付ける 4 つの中核的なインタラクション レベルの次元、つまりタスクの調整、検証可能性、操縦可能性、および適応性を特定します。最後に、ユーザーが関与するコーディング環境、包括的な検証メカニズム、ヒューマン エージェント インタラクション品質の原則的な尺度など、これらの側面を前進させるための具体的な研究の方向性を概説します。
原文 (English)
Humans are Missing from AI Coding Agent Research
Recent progress in AI coding agent research has led to rapid improvements in agents' ability to autonomously perform complex software engineering tasks, from editing large codebases to executing long-horizon development workflows. As these systems make strides, however, the primary bottleneck to practical usefulness increasingly shifts away from pure task-solving capability, and toward challenges in how users communicate with, supervise, and trust agents. In this position paper, we argue for a reorientation from autonomous to human-centered coding agents: systems designed not only to complete tasks, but to collaborate effectively with people. We identify four core interaction-level dimensions that characterize the human-agent task-solving loop: task alignment, verifiability, steerability, and adaptability. Finally, we outline concrete research directions to advance these dimensions, including user-involved coding environments, comprehensive verification mechanisms, and principled measures of human-agent interaction quality.
プログラムポートフォリオの規模でカリキュラムと労働市場の整合性を測定する
複数の重複するコンピューティングの学位を提供する大学は、そのプログラムが労働市場のコンピューティング業務のセグメント化に応じて差別化されており、それらが連携してその市場に向けて卒業生を準備していると暗黙のうちに想定しています。これをテストするのは困難です。なぜなら、カリキュラム委員会が利用できる手段、つまり諮問委員会、追跡調査、雇用主調査は時間がかかり、範囲が狭く、再現が難しいためです。私たちは、情報技術学部の 5 つの学部プログラムすべてに 1 つの統一された分類法に基づいたアラインメント分析を適用し、1,922 のコース学習成果と、4 つの教育機関からの重複を排除した 5,186 件の求人の統合コーパスから抽出された 103,349 のコンピテンシーを比較します。すべてのコンピテンシーは、根拠のある単一言語モデルの手順によって取得されます。この手順では、ソースから逐語的にコピーし、ソースと照合して検証し、ESCO に準拠した 11 のドメインとブルーム認知レベルの 1 つに割り当てます。カリキュラムはカタログとしてではなく、学生が学位を取得する際の単位時間と選択上の制約を尊重した達成ベースで読み取られます。この抽出は、2 人の独立した教員評価者 (ドメイン カッパ 0.91、ブルーム レベル カッパ 0.86) と人間が判断したゴールド セット (カッパ 0.72) との ESCO マッチングによってブラインドで検証されます。 4 つの発見が得られます。コンテンツのギャップはプログラム固有ではなく体系的なものであり、システム、ソフトウェア エンジニアリング、セキュリティ、Web 開発に集中しています。大学の共有コアは、要求される能力の約 3 分の 1 しか満たしていません。プログラムは規律的な内容で十分に差別化されていますが、不十分な点では均質です。そしてカリキュラムは、ポートフォリオ全体で市場よりもほぼフルブルームレベル下に設定されており、特にシステムが顕著です。プログラム設計、カリキュラムガバナンス、カリキュラム分析の実践への影響について説明します。
原文 (English)
Measuring Curriculum-Labor Market Alignment at the Scale of a Program Portfolio
A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor market segments computing work and that, together, they prepare graduates for that market. Testing this is difficult, because the instruments available to curriculum committees, namely advisory boards, tracer studies, and employer surveys, are slow, narrow, and hard to reproduce. We apply one uniform, taxonomy-anchored alignment analysis across all five undergraduate programs of a College of Information Technology, comparing 1,922 course learning outcomes against 103,349 competencies extracted from a unified corpus of 5,186 deduplicated job openings from four boards. Every competency is obtained by a grounded single-language-model procedure that copies it verbatim from the source and verifies it against the source, then assigns it to one of eleven ESCO-aligned domains and a Bloom cognitive level; the curricular supply is read not as a catalog but on a realized-attainment basis that respects the credit-hour and elective constraints under which a student completes a degree. The extraction is validated blind by two independent faculty raters (domain kappa 0.91, Bloom level kappa 0.86) and the ESCO matching against a human-adjudicated gold set (kappa 0.72). Four findings emerge. The content gaps are systemic rather than program-specific, concentrated in systems, software engineering, security, and web development; the shared college core satisfies only about a third of the demanded competencies; the programs are well differentiated in disciplinary content yet homogeneous in where they fall short; and the curriculum is pitched roughly a full Bloom level below the market across the portfolio, most acutely in systems. We discuss the implications for program design, curriculum governance, and the practice of curriculum analytics.
インタラクションの準備: 人間の役割で AI エージェントを構築および評価するためのフレームワーク
役割を担う AI エージェントを構築する製品チームとエンジニアリング チームは、評価ギャップに直面しています。エージェントは、割り当てられた役割の動作要件を満たしていながら、正確で安全かつ流暢なコンテンツを作成できます。このホワイトペーパーでは、パフォーマンスの不足している層を指定および評価するためのフレームワークとして Interaction Readiness を紹介します。このフレームワークは、エージェントが知っていることや発言していることを管理するコンテンツ仕様と、役割に基づいた交換においてエージェントがどのように行動すべきかを定義するインタラクション仕様を分離します。インタラクション仕様では、チームは、展開前に役割の目的、権限の境界、繰り返し発生する状況、境界ケース、修復動作、および監査基準を定義する必要があります。私たちは、目的の理解、権限の調整、トーンの管理、故障の修復という 4 つのエージェントの操作を通じて、インタラクションの準備を運用します。 AI 家庭教師エージェントとの学生のやり取りの公開データセットである StudyChat を使用して、コンテンツの正確さと対話の品質は独立した次元であることを示します。エージェントは、家庭教師としては失敗しても事実としては正しい場合もあれば、技術的には間違っていても対話的には健全である場合もあります。最も根強い失敗は権限の調整ミスです。エージェントは多くの場合、どのように答えるかを知っていますが、講師の役割が応答を許可するかどうか、いつ、どのように応答するかを知りません。この文書では、これらの調査結果を仕様テンプレートと監査手順に変換し、製品チームとエンジニアリング チームが導入の前後に適用できるようにしています。
原文 (English)
Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles
Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Readiness as a framework for specifying and evaluating that missing layer of performance. The framework separates content specifications, which govern what an agent knows and says, from interaction specifications, which define how an agent should conduct itself in a role-governed exchange. Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment. We operationalize interaction readiness through four agent operations: understanding purpose, calibrating authority, managing tone, and repairing breakdowns. Using StudyChat, a public dataset of student interactions with an AI tutoring agent, we show that content accuracy and interaction quality are independent dimensions: an agent may be factually correct while failing as a tutor, or interactionally sound while technically wrong. The most persistent failure is authority miscalibration: the agent often knows how to answer, but not whether, when, or how the tutor role permits it to answer. The paper translates these findings into a specification template and audit procedures that product and engineering teams can apply before and after deployment
EU-ETSが攻撃を受けていますか?電力部門の脱炭素化に対する炭素価格抑制の影響
欧州諸国は、脱炭素化と電化を追求しながら、新たな地政学的緊張によって引き起こされるエネルギーコストの増加を緩和する政策を議論している。注目すべき例は、イタリアの2026年デクレト・ボレット・パッケージであり、これは、他の規定の中でも特に、特定のガス燃料発電所の卸電力市場への入札から炭素価格相当額を削除することを提案している。私たちはこれを、投資、排出量、消費者コストに対する電力市場における炭素価格シグナルの抑制の長期的な影響を評価するためのケーススタディとして使用します。当社では、長期的な電力市場評価に焦点を当てたマルチエージェント強化学習フレームワークである MARLEY を使用した、様式化されたイタリアの電力システムを採用しています。このフレームワークでは、グリーン投資、リソースの適切性、および柔軟性に対するさまざまなレベルのサポートを使用して、構成全体でこのポリシーをテストします。結果は、炭素価格シグナルの部分的な抑制は短期的なコスト削減につながるが、延期された排出量は最終的に消費者によって返済されるため、システム全体のコストに対する長期的な影響はわずかであることを示しています。価格シグナルの抑制により再生可能エネルギーや蓄電への投資のインセンティブが損なわれるため、ほとんどの構成で CO$_2$ 排出量が増加します。グリーン投資を支援するための最も野心的な構成のみがこの結果を回避しますが、卸売価格シグナル自体を軽視することによって回避し、それによって提案された価格介入の理論的根拠に矛盾するハイブリッド市場パラダイムへのコミットメントを必要とします。
原文 (English)
EU-ETS under attack? The impact of carbon price suppression on the decarbonization of the power sector
European countries are debating policies to mitigate the increased energy costs caused by renewed geopolitical tensions, while pursuing decarbonization and electrification. A notable example is Italy's 2026 Decreto Bollette package, which proposes to remove the carbon price equivalent from the bids of certain gas-driven power plants to wholesale electricity markets, among other provisions. We use this as a case study to assess the long-term implications of suppressing the carbon price signal in the electricity market for investment, emissions, and consumer costs. We employ a stylized Italian power system using MARLEY, a multi-agent reinforcement learning framework focused on long-term electricity market assessments. In this framework, we test this policy across configurations with varying levels of support for green investment, resource adequacy, and flexibility. Results show that partial suppression of the carbon price signal yields short-term cost reductions but only a minor long-term effect on total system costs, as the deferred emissions are ultimately repaid by consumers. CO$_2$ emissions rise across most configurations since suppressing the price signal erodes incentives for renewable and storage investment. Only the most ambitious configurations for supporting green investment avoid this outcome, but they do so by marginalizing the wholesale price signal itself, thereby requiring a commitment to a hybrid market paradigm that is in contradiction with the rationale of the proposed price intervention.
FluctlightDB: AI エージェントのためのデータのメモリ モデル
50 年間にわたり、データ システムは 2 つの質問に答えてきました。リレーショナル モデルは、どのレコードが述語に一致するかを尋ねました。ベクトル モデルは、どのベクトルがクエリに最も近いかを尋ねます。どちらも、長いセッションにわたるキュー主導の出自重視の想起を目的として構築されたものではありません。私たちは、長期エージェント メモリを独自の書き込みセマンティクス (エンコーディング、分離、統合、来歴) と読み取りセマンティクス (リンクされたメモリ グラフ全体にわたるキュー駆動のアクティブ化) を備えた別個のデータ モデルとして扱うことを提案し、experience() および activate() を介してこのコントラクトを実装する組み込みエンジン FluctlightDB を提示します。私たちはその主張を断定的ではなく慎重に行います。私たちは Mem0、Zep、または HippoRAG スタイルのメモリ層に対する新規性を主張せず、それらの下にある組み込みエンジンの契約のみを主張します。 LoCoMo (公式証拠再現指標、10 会話、1,982 ゴールド スパン) では、CHORUS は 2026 年 7 月の内部再現実行で 99.0% を再現しました。 LongMemEval-S (500 の質問、公式 session_recall@8) では、当社の検索ハーネスのスコアは 97.6% (488/500) でした。当社のリーダー/ジャッジスタックによるエンドツーエンド QA スコアは 97.4% (487/500) -- これらのレイヤーは、文脈のみで引用したベンダー リーダーボードの数値とは異なるプロトコルを使用します。 BEIR SciFact (共有 MiniLM エンベディング、同じハーネス、Recall Fabric オン) では、CHORUS/PRISM は nDCG@10 (0.646 対 0.645) および Recall@10 (0.792 対 0.783) で Chroma を上回りました。また、著者が設計した小規模な回帰スイート (FAMB、言い換えれば n=10、その他のサブテスト n=1) を 100% マクロで報告します。これはピア ベンチマークではなく内部検証です。見知らぬ人でも、pip install "fluctlightdb[native]" と最小限の connect() -> experience() -> activate() スクリプト (ソースのみではなくコンパイルされたホイール) を使用して、1 分以内にエンジンを検証できます。ハーネスと凍結された JSON は MIT ライセンスを取得しています。私たちは新しい神経科学や新しいトランスフォーマーを主張していません。私たちはデータスタックの欠落している層を提案し、他の人が再現してコンテストできるエンジンをリリースします。
原文 (English)
FluctlightDB: A Memory Model of Data for AI Agents
For fifty years, data systems have answered two questions. The relational model asked which records match a predicate; the vector model asked which vectors lie nearest a query. Neither was built for cue-driven, provenance-weighted recall across long sessions. We propose treating long-term agent memory as a distinct data model -- with its own write semantics (encoding, separation, consolidation, provenance) and read semantics (cue-driven activation across a linked memory graph) -- and present FluctlightDB, an embedded engine that implements this contract via experience() and activate(). We make that case carefully, not categorically: we do not claim novelty over Mem0, Zep, or HippoRAG-style memory layers, only an embedded engine contract beneath them. On LoCoMo (official evidence-recall metric; 10 conversations, 1,982 gold spans), CHORUS recalls 99.0% on an internally reproduced July 2026 run. On LongMemEval-S (500 questions, official session_recall@8), our retrieval harness scores 97.6% (488/500); end-to-end QA with our reader/judge stack scores 97.4% (487/500) -- these layers use different protocols than vendor leaderboard figures we cite for context only. On BEIR SciFact (shared MiniLM embeddings, same harness, Recall Fabric on), CHORUS/PRISM edges Chroma on nDCG@10 (0.646 vs. 0.645) and Recall@10 (0.792 vs. 0.783). We also report a small author-designed regression suite (FAMB; paraphrase n=10, other sub-tests n=1) at 100% macro -- internal validation, not peer benchmark. Strangers can verify the engine in under a minute via pip install "fluctlightdb[native]" and a minimal connect() -> experience() -> activate() script (compiled wheel, not source-only). Harnesses and frozen JSON are MIT-licensed. We claim no new neuroscience and no new transformer; we propose a missing layer of the data stack and release an engine others can reproduce and contest.
あなたは私に論理を話しているのですか?言語モデルの三段論的推論能力の評価
言語モデル (LM) は、三段論法に基づく推論などの論理タスクに苦労します。知識表現 (KR) は、モデルがタスクを解決するのに役立つ入力情報を表現する上で重要な役割を果たすことが示されています。この観察は、FOLIO および P-FOLIO データセットを拡張することによって、さまざまな形式的な KR 表記が三段論的推論に及ぼす影響を研究する動機となっています。教師ありファインチューニング (SFT) およびゼロショット (ZS) 設定における小型言語モデル (SLM) の実験では、入力表記法の選択により、より高速な推論を可能にしながら、自然言語に匹敵するパフォーマンスが得られることが示されました。また、三段論的分類法 (SEF) を提案し、それを使用して ZS プロンプトを論理定義で強化し、小規模モデルでの推論を強化します。私たちは、KR 表記法で三段論法を自動的に生成し、その SEF カテゴリを定義するための最初の Python ライブラリとして、フレームワークである Common Logic Grammar Construction (CLGC) をオープンソース化しています。
原文 (English)
Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities
Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial role in expressing input information to help models solve tasks. This observation motivates our study of the impact of different formal KR notations on syllogistic reasoning by extending the FOLIO and P-FOLIO datasets. Our experiments on Small Language Models (SLMs) in Supervised Fine-Tuning (SFT) and Zero-Shot (ZS) settings show that the choice of input notation can yield performances competitive with natural language while enabling faster inference. We also propose a syllogistic categorization method (SEF) and use it to enrich ZS prompts with logical definitions, which boost reasoning in small models. We open-source our framework, Common Logic Grammar Construction (CLGC), as the first Python library for automatically generating syllogisms in KR notations and defining their SEF categories.
観察から介入へ: 脳の記憶と大規模言語モデル
脳と大規模言語モデル (LLM) は根本的に異なる記憶システムですが、共通の機能的質問を通じて比較することができます。つまり、記憶関連の情報がどこで表されるか、部分的な手がかりがより広範な関連性をどのように回復するか、新しい情報がどのように書き込まれ、更新されるか、記憶関連の状態がどのように混乱するかなどです。生物学的システムでは、これらの疑問はシナプス、神経細胞集団、海馬と皮質の相互作用、可塑性などに及びます。 LLM では、重み付け、アクティベーション、コンテキスト ウィンドウ、検索システム、外部ストアにまたがります。したがって、この比較は解剖学的なものではなく、機能的かつ実験的なものです。人体研究では、概念反応がまばらで、時間的結合、急速な関連形成、エピソード固有のコード化、および想起関連の再活性化が明らかになっているが、選択的介入は依然として限られている。げっ歯類の研究は、学習関連のアンサンブルへのより選択的な因果的アクセスを提供しますが、ヒトやマカクの介入は通常、より広範な回路に影響を与えます。 LLM には生きたエピソード記憶がありませんが、内部状態と保存された情報を異常に直接的かつ反復的に操作できます。私たちは、この非対称性が新たな機会を生み出すと主張します。 LLM はメモリ自体の点で優れているのではなく、実験的なアクセスの点で優れています。彼らのツールは、検索、更新、永続性、可逆性、および意図しない効果に関する広範な疑問を、より鋭い生物学的仮説に変えるのに役立つ可能性があります。生産的なブリッジは、解剖学的部分ではなく、実験ロジックを転送することです。
原文 (English)
From Observation to Intervention: Memory in Brains and Large Language Models
Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional questions: where memory-related information is represented, how partial cues recover broader associations, how new information is written or updated, and how memory-related states can be perturbed. In biological systems, these questions span synapses, neuronal ensembles, hippocampal-cortical interactions, and plasticity; in LLMs, they span weights, activations, context windows, retrieval systems, and external stores. The comparison is therefore functional and experimental rather than anatomical. Human studies reveal sparse concept responses, temporal binding, rapid association formation, episode-specific coding, and recall-related reactivation, but selective intervention remains limited. Rodent studies provide more selective causal access to learning-related ensembles, whereas human and macaque interventions usually affect broader circuits. LLMs lack lived episodic memory, yet they permit unusually direct and repeatable manipulation of internal states and stored information. We argue that this asymmetry creates a new opportunity. LLMs are not ahead in memory itself, but in experimental access. Their tools may help turn broad questions about retrieval, updating, persistence, reversibility, and unintended effects into sharper biological hypotheses. The productive bridge is to transfer experimental logic, not anatomical parts.
クエリのタイミングにより、LLM と人間の間に逆の位置バイアスが生じる
最新性や優位性の効果などの位置バイアスは、大規模言語モデル (LLM) で文書化されていますが、これらのモデルが評価を行う根本的なメカニズムは依然としてよく理解されていません。証拠に応じた人間の判断には、優位性バイアスと最新性バイアスの両方が観察されていますが、最近の研究では、聞き手が自分の信念を更新する \emph{いつ} (証拠の提示中または最後にのみ) が、そのような効果の存在に影響を与えることが示唆されています。私たちは同様の現象が LLM にも当てはまるかどうかを調査し、人間の行動との相違を見つけます。これらのバイアスは、以前のモデルと比較して、新しいモデルではさらに悪化します。
原文 (English)
Query Timing Produces Opposite Positional Biases Between LLMs and Humans
Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood. Both primacy and recency biases have been observed in human judgments in response to evidence, but recent work suggest that \emph{when} the listener updates their beliefs -- during the presentation of evidence or only at the end -- influences the presence of such effects. We investigate whether a similar phenomenon holds for LLMs, finding divergence from human behavior. These biases are more exacerbated in newer models compared to their predecessors.
大規模言語モデルにおける複雑なグラフ推論のための統合多次元ベンチマーク
グラフ推論は、グラフ インスタンスをプログラムで生成し、構造的に制御し、長時間入力の設定に合わせて自然にスケーリングできるため、大規模言語モデル (LLM) の推論能力を評価するための有望なテストベッドを提供します。しかし、既存のグラフ推論ベンチマークは、データの複雑さの範囲が限られており、手動による構築に大きく依存しており、テキストベースとコードベースの推論モードにわたる統一された評価が不足しています。これらの制限に対処するために、複雑なグラフ推論ベンチマークを構築するための 5 段階の \textit{半自動} フレームワークである {\dataset} を提案します。 \textit{グラフ サイズ}、\textit{タスクの複雑さ}、\textit{タスクの説明}、\textit{グラフの読み込み}、\textit{タスク ソース}の 5 つの次元に沿ってベンチマーク カバレッジを拡張します。このフレームワークは、LLM ベースのデータ ジェネレーターを使用して、タスクの説明、グラフ データ、参照ソリューション、グラフ読み込みスクリプト、質問フォーム、評価スクリプトを自動的に生成すると同時に、主要な品質管理段階で人間による検証を維持します。これに基づいて、$202$ のタスクでベンチマークを構築し、テキストベース、コードベース、および拡張推論設定の下で LLM を評価します。実験の結果、複雑さの次元によって、既存のベンチマークでは見えにくいモデルの制限が明らかになることが示されています。既存の微調整モデルは GraphGym に一般化するのに苦労していますが、検索拡張手法はシナリオ依存の適応性を示し、テキスト推論を改善しますが、コーディング推論を一貫して改善するわけではありません。これらの発見は、私たちの結果がグラフ推論の挑戦的かつ診断ベンチマークとして機能し、将来の強化方法に経験的な指針を提供することを示唆しています。コードとデータセットは近日公開される予定です。
原文 (English)
Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models
Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input settings. However, existing graph reasoning benchmarks have limited coverage of data complexity, rely heavily on manual construction, and lack unified evaluation across text-based and code-based reasoning modes. To address these limitations, we propose {\dataset}, a five-stage \textit{semi-automatic} framework for constructing complex graph reasoning benchmarks. It expands benchmark coverage along five dimensions: \textit{Graph Size}, \textit{Task Complexity}, \textit{Task Description}, \textit{Graph Loading}, and \textit{Task Source}. The framework uses an LLM-based data generator to automatically produce task descriptions, graph data, reference solutions, graph-loading scripts, question forms, and evaluation scripts, while retaining human validation at key quality-control stages. Based on it, we construct a benchmark with $202$ tasks and evaluate LLMs under text-based, code-based, and augmented reasoning settings. Experiments show that the complexity dimensions reveal model limitations that are less visible in existing benchmarks; existing fine-tuned models struggle to generalize to GraphGym, whereas retrieval-augmented methods show scenario-dependent adaptability, improving textual reasoning but not consistently improving coding reasoning. These findings suggest that ours serves as a challenging and diagnostic benchmark for graph reasoning and provides empirical guidance for future enhancement methods. Code and dataset will be published soon.
マルチモーダル認知のための階層的エネルギーベースのモデル
我々は、以前に提案されたシングルモダリティモデル(LEPP)を拡張して視覚と言語を統合する、マルチモーダル認知の階層的エネルギーベースモデルであるIM-LEPP(統合マルチモーダル潜在エネルギーベース予測処理)を提案します。生成ニューラル ネットワークは、統計力学が熱力学にどのように関係するかに類似した、認知力学の効果的な理論であるという見解に従って、IM-LEPP は、神経回路の説明としてではなく、学習されたエネルギーランドスケープを流れる潜在状態として認知をモデル化します。このアーキテクチャは、Lambon Ralph らの制御された意味認識フレームワークに基づいたハブアンドスポーク階層であり、視覚オブジェクト、シーン、および言語単位の予測コーディング パイプラインが、前側頭葉をモデル化した共有アモーダル ハブに収束します。各パイプライン独自の予測は、現在のハブの状態によって上書きされるのではなく、条件付けされ、パイプライン固有のアイデンティティを維持しながら、すべての予測が完全なマルチモーダル コンテキストを反映します。我々は、このアーキテクチャが、不注意失明やネッカーキューブ双安定性などの注意現象のメカニズムを説明し、その構造が、次の単語予測における軌道敏感性に関するトランスフォーマー言語モデルとの反証可能な対比とともに、サプライズ理論、N400/P600 ERPコンポーネント、ガーデンパス再分析などの心理言語学で独立して確立された知見を回復または動機づけることを示します。また、LLM に関連したデータ効率の高い言語習得について議論し、意味論的/エピソード記憶サブシステムの概要を説明し、予測コーディング、自由エネルギー原理、JEPA、階層的時間記憶とモデルを比較し、その中心的な主張をテストするための具体的な実験的予測を提案します。予測コーディング。エネルギーベースのモデル。拡散モデル。効果的な理論。計算神経科学。
原文 (English)
A Hierarchical Energy-Based Model for Multimodal Cognition
We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a previously proposed single-modality model (LEPP) to integrate vision and language. Following the view that generative neural networks are effective theories of cognitive dynamics, analogous to how statistical mechanics relates to thermodynamics, IM-LEPP models cognition as latent states flowing through learned energy landscapes rather than as an account of neural circuitry. The architecture is a hub-and-spoke hierarchy, grounded in the controlled semantic cognition framework of Lambon Ralph et al., in which predictive-coding pipelines for visual objects, scenes, and linguistic units converge on a shared amodal hub modeled on the anterior temporal lobe. Each pipeline's own prediction is conditioned by, rather than overwritten by, the current hub state, preserving pipeline-specific identity while letting every prediction reflect the full multimodal context. We show this architecture gives a mechanistic account of attentional phenomena such as inattentional blindness and Necker-cube bistability, and that its structure recovers or motivates independently established findings in psycholinguistics, including surprisal theory, the N400/P600 ERP components, and garden-path reanalysis, alongside a falsifiable contrast with transformer language models on trajectory-sensitivity in next-word prediction. We also discuss data-efficient language acquisition relative to LLMs, outline a semantic/episodic memory subsystem, situate the model against predictive coding, the free-energy principle, JEPA, and Hierarchical Temporal Memory, and propose concrete experimental predictions to test its central claims Key Words: predictive processing; predictive coding; energy-based models; diffusion models; effective theory; computational neuroscience.
SynWeaver: Web エージェント向けの Web サイト事前タスクと軌跡の共同合成
Web エージェントは、Web サイト固有の監視が不足しているため、目に見えない Web サイトに一般化するのに苦労することがよくあります。最近の探索ベースのデータ合成手法では、手動による注釈は削減されていますが、依然として 2 つの重要な制限に直面しています。1 つは、Web サイトのすべての機能をカバーできないことが多く、Web サイトに関する十分な事前知識がないと、幻覚的なタスクを提案する傾向があり、その結果、下流の軌跡合成の多様性と効率が制限されることです。私たちは、これらの課題に対処するために設計された、Web サイト事前タスク軌跡共合成フレームワークである \textbf{SynWeaver} を紹介します。 SynWeaver はまず、構造化された Web サイトの探索を実行し、機能的に異なるページ状態とターゲット Web サイト上の実行可能なインタラクションの広範なセットをカバーする Web サイト マップを構築します。次に、このマップからページレベルと遷移レベルの監視を導き出し、Web サイト固有の事前分布を使用して UI 認識モデルをトレーニングし、より根拠のあるタスク提案を可能にします。最後に、SynWeaver は協調的なタスクと軌道の合成を実行し、タスクと実行の軌道に矛盾が生じた場合にそれらを共同で更新し、収集された結果を検証して修復して、実行可能で意味的に整合した監視を生成します。 WebArena と WebVoyager での実験では、SynWeaver が強力な合成ベースラインを常に上回っており、ドメイン内およびドメイン外の両方の一般化に対してより効果的な監視が得られることが実証されています。
原文 (English)
SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents
Web agents often struggle to generalize to unseen websites because they lack website-specific supervision. Recent exploration-based data synthesis methods reduce manual annotation, but they still face two key limitations: they often fail to cover the full functionality of a website, and without sufficient website prior knowledge, they tend to propose hallucinated tasks, which in turn limits the diversity and efficiency of downstream trajectory synthesis. We present \textbf{SynWeaver}, a website-prior task-trajectory co-synthesis framework designed to address these challenges. SynWeaver first performs structured website exploration and constructs a website map that covers a broad set of functionally distinct page states and executable interactions on the target website. It then derives page-level and transition-level supervision from this map to train a UI-aware model with website-specific priors, enabling more grounded task proposals. Finally, SynWeaver performs collaborative task-trajectory synthesis, jointly updating the task and execution trajectory when they become inconsistent, and then verifies and repairs the collected results to produce executable, semantically aligned supervision. Experiments on WebArena and WebVoyager demonstrate that SynWeaver consistently outperforms strong synthesis baselines and yields more effective supervision for both in-domain and out-of-domain generalization.
AI コーディング エージェントによる仕様優先の収束: テスト オラクルも人間によるコード レビューも行わず、717,000 行のコードベース内の 189 ファイルにわたるコア アーキテクチャの不変条件を解体するケース スタディ
このペーパーでは、生成されたコードの人によるレビューや、ターゲットの動作を検証するための既存のオラクルを使用しない、仕様優先プロトコルに基づく AI コーディング エージェントによる大規模なアーキテクチャ リファクタリングの、完全に装備された単一のケース スタディを報告します。このタスクは、相互依存する大規模なコードベース全体で中心となる不変条件を解体するというもので、作成者は、インクリメンタル リファクタリング (従来は代わりに書き換えが必要だった種類の変更) では事実上実行不可能であると評価しました。ここで説明されているプロトコルに従って、エージェントは正常に完了しました。このシステムは、3,648 ファイルにわたる 717,725 行のプロダクション TypeScript アプリケーションです。このタスクでは、コアのライフタイム不変条件、つまり AI リクエストの間 UI パネルが開いたままであることの保証を解体する必要がありました。目標の動作は、ストリーミング世代がパネルの終了後も存続し、再度開いたときに、損失や重複なしに同じライブ ストリームに再接続できることです。プロトコル: エージェントによる正式な仕様、その仕様をソース コードに対して監査する 14 回の改良サイクル、アトミックな実装、コンパイル/テストのフィードバック ループ、その後、凍結された仕様に対してコードを監査する 17 回の検証サイクル。 31 回の監査パスを通じて、人間がプログラムを実行する前に 201 個の欠陥が修正されました。収束基準は経験的であり、2 つの連続した検証パスで結果がゼロになるというものでした。この変更は 189 個のファイル (31 個の新規ファイル) に影響を与えました。抽出フェーズでは、2 つのコミットで合計 288 ファイル、34,770 件の挿入、16,422 件の削除がコミットされました。最初のセッションとその後の約 30 回のセッションでは、ソフトウェアは指定どおりに動作し、バグは観察されませんでした。経過: 3 日。料金: 2,430 ドル。完全な仕様と生のセッション ログ (フランス語で 1,500 ページ以上) が証拠として公開されており、プロセスの検査と一貫性チェックのための言語モデルへの送信が可能になります。
原文 (English)
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully. The system is a 717,725-line production TypeScript application across 3,648 files. The task required dismantling a core lifetime invariant: the guarantee that a UI panel remains open for the duration of an AI request. The target behaviour was that a streaming generation survives the closing of its panel and can be reattached, on reopening, to the same live stream with no loss or duplication. The protocol: formal specification by the agent, 14 refinement cycles auditing that specification against the source code, atomic implementation, a compile/test feedback loop, then 17 verification cycles auditing the code against the frozen specification. Across 31 audit passes, 201 defects were corrected before any human executed the program. The convergence criterion was empirical: two consecutive verification passes returning zero findings. The change touched 189 files (31 new); with the extraction phase, the two commits total 288 files, 34,770 insertions, 16,422 deletions. Across the first and roughly thirty later sessions, the software behaved as specified, no bug observed. Elapsed: three days; cost: USD 2,430. The full specification and raw session logs, 1,500+ pages in French, are published as evidence, allowing inspection of the process and submission to a language model for consistency checking.
デュアル時空間属性: 再発グラフ異常検出のためのアーキテクチャに合わせた事後説明可能性
動的グラフの異常に対する深層学習検出器は、高い精度に達していますが、依然として不透明です。エッジにフラグが立てられると、分析者はスコアを受け取りますが、理由はわかりません。このような不透明さは、自動化された決定が監査可能で信頼できるものでなければならない、そのような検出器が導入されている協調的で規制された情報システムでは維持できません。私たちは、動的グラフにおけるエッジレベルの異常検出のための基本的な GCN+GRU フレームワークである AddGraph でこのギャップに対処します。これは、私たちの知る限り、いかなる形式の説明可能性も備えていません。我々は、デュアル時空間属性 (DSTA) メカニズムに基づいて構築された、厳密にポストホックな説明可能フレームワークである X-AddGraph を紹介します。このフレームワークの 3 つのコンポーネントはそれぞれ、AddGraph のアーキテクチャ モジュールの 1 つと連携しています。つまり、現在の隣接構造 (空間) に対する勾配ベースの関連性属性、推論中に既に計算された文脈上の注意の重みの直接読み取り (追加コストなしで短期的な時間的)、および勾配ロールバックです。反復的な隠れた状態 (長期的な一時的)。検出器がフリーズされているため、検出パフォーマンスは正確に維持されます (デルタ AUC = 0、小数点以下 10 桁まで経験的に検証)。 UCI Message ベンチマークでは、トレーニングされた AddGraph ベースラインはスナップショットあたりの平均 AUC 0.8705 に達し、最初に公開された結果を超えています。 X-AddGraph は、存在しない説明を追加しながら、すべてのスコアを同一に再現します。 4 つのエッジ母集団 (確信度の高い真陽性、信頼度の低い真陽性、偽陽性、ランダム サンプル) にわたって評価された長期アトリビューションにより、ランダム選択 (0.127 対 0.074) よりもはるかに多くの反事実シグナルを伝える過去のスナップショットが特定されます。これは、空間盲目の説明者では提供できない機能です。完全な再現性を実現するために実装をリリースします。
原文 (English)
Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection
Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the analyst receives a score but no reason. This opacity is untenable in the cooperative, regulated information systems where such detectors are deployed, where automated decisions must be auditable and trustworthy. We address this gap for AddGraph, the foundational GCN+GRU framework for edge-level anomaly detection in dynamic graphs, which to our knowledge has never been equipped with any form of explainability. We present a strictly post-hoc explainability framework, X-AddGraph, built on a Dual Spatial-Temporal Attribution (DSTA) mechanism whose three components are each aligned with one of AddGraph's architectural modules: a gradient-based relevance attribution over the current adjacency structure (spatial), a direct reading of the contextual attention weights already computed during inference (short-term temporal, at zero additional cost), and a gradient rollback through the recurrent hidden states (long-term temporal). Because the detector is frozen, detection performance is preserved exactly (Delta AUC = 0, verified empirically to ten decimal places). On the UCI Message benchmark, our trained AddGraph baseline reaches an average per-snapshot AUC of 0.8705, exceeding the originally published result; X-AddGraph reproduces every score identically while adding explanations where none existed. Evaluated across four edge populations - confident true positives, low-confidence true positives, false positives, and random samples - the long-term attribution identifies historical snapshots carrying significantly more counterfactual signal than random selection (0.127 vs. 0.074), a capability that no spatially-blind explainer can provide. We release our implementation for full reproducibility.
SSPO: ニューラル組み合わせ最適化のための構造認識類似性加重優先最適化
ニューラル組み合わせ最適化 (NCO) は、トレーニングのために並列ソリューション サンプリングに依存していますが、既存の方法では、同時サンプリングされたソリューション グループに潜在する豊富な情報を十分に活用できません。プリファレンス最適化手法は、単一の最適なソリューションに固定され、他のすべてのピアからのきめの細かい品質と構造信号を破棄します。これを勾配信号偏光と呼んでいます。代わりに平均ベースのベースラインはピアに均一に重み付けをするため、構造的にほぼ同一のピアがベースラインに冗長な情報を溢れさせ、勾配の分散を高く保ちます。これはベースラインの冗長性と呼ばれる失敗です。我々は、SSPO (構造認識類似性重み付け優先最適化) を提案します。これは、非類似性重み付けの Leave-One-Out ベースラインを通じて、すべての $B$ サンプリングされたソリューションを共同でスコア付けします。つまり、構造的に異なるピアがより高い重みを受け取り、単一のメカニズムで両方の障害を解決します。ベースラインは、エンコーダーの既存のノード表現から構築されたゼロパラメーターの問題適応型ソリューションの埋め込みを使用します。 TSP、EFL、および JSP ベンチマークの実験では、以前のベストアンカーおよび均一ウェイトのベースラインを超える一貫したゲインが示されています。 TSP および EFL の均一 RLOO と直接比較すると、構造を意識した重み付けが改善の主な原動力であることが確認されます。 SSPO でトレーニングされた EFL ポリシーは、JD$\mathord{.}$com の生産施設ロケーション システムに展開され、大規模な実用性を確認しています。
原文 (English)
SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization
Neural combinatorial optimization (NCO) relies on parallel solution sampling for training, yet existing methods fail to fully exploit the rich information latent in a co-sampled solution group. Preference-optimization methods anchor on the single best solution and discard fine-grained quality and structural signal from all other peers-a failure we term gradient signal polarization. Mean-based baselines instead weight peers uniformly, so structurally near-identical peers flood the baseline with redundant information and keep gradient variance high-a failure we term baseline redundancy. We propose SSPO (Structure-Aware Similarity-Weighted Preference Optimization), which scores all $B$ sampled solutions jointly through a dissimilarity-weighted leave-one-out baseline: structurally distinct peers receive higher weight, resolving both failures in a single mechanism. The baseline uses zero-parameter, problem-adaptive solution embeddings built from the encoder's existing node representations. Experiments on TSP, EFL, and JSP benchmarks show consistent gains over prior best-anchor and uniform-weight baselines. A direct comparison against uniform RLOO on TSP and EFL confirms that structure-aware weighting is the primary driver of improvement. The SSPO-trained EFL policy has been deployed in a production facility-location system at JD$\mathord{.}$com, confirming practical viability at scale.
パーソナライズされたスコアラー モデリング: 複数の専門家から確実な睡眠ステージ ラベルを導き出すための学習ベースのフレームワーク
睡眠段階分類は睡眠障害の診断と管理にとって重要ですが、自動段階分類研究のほとんどは、既知のスコアラー間の変動にもかかわらず、単一の参照催眠計画に対してモデルを評価します。この研究では、複数スコアのデータセットを使用して、複数の専門家の集合的な行動からより信頼性の高い参照ラベルを構築できるかどうかを調査します。公開されている DOD-H および DOD-O データセットを使用します。 EEG (C3-M2) および顎 EMG 信号は 30 秒のエポックに分割され、各モダリティから 30 個の特徴が抽出され、EEG+EMG について 60 個の特徴が得られました。機械学習モデルから導出された混同行列を使用して、各スコアラーのステージ固有の動作をモデル化する学習ベースのヒプノグラム (LBH) を提案します。列の正規化後、これらの行列は、各スコアラーのラベルが与えられた場合に、各真の睡眠段階の確率を推定します。確率はスコアラー全体で集計され、各エポックの最終ラベルが割り当てられます。 LBH は、EEG のみおよび EEG+EMG 設定の下で、ランダム フォレスト、サポート ベクター マシン、および多層パーセプトロン分類器を使用して評価され、データセット催眠図 (DH) およびベストスコアラー催眠図 (BSH) と比較されました。 LBH は全体的なパフォーマンスを一貫して向上させました。最良の結果はランダム フォレストと EEG+EMG で得られ、DOD-H では精度 86.07%、精度 85.46%、F1 スコア 85.29%、DOD-O では精度 86.04%、精度 85.21%、F1 スコア 84.70% に達しました。これらの発見は、パーソナライズされたスコアラー モデリングが、個々の専門家からの情報を破棄することなく、参照催眠計画の構築を改善できることを示唆しています。
原文 (English)
Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts
Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate models against a single reference hypnogram despite known inter-scorer variability. This study investigates whether multi-scored datasets can be used to construct more reliable reference labels from the collective behavior of multiple experts. We use the publicly available DOD-H and DOD-O datasets. EEG (C3-M2) and chin EMG signals were segmented into 30-s epochs, and 30 features were extracted from each modality, yielding 60 features for EEG+EMG. We propose a learning-based hypnogram (LBH) that models the stage-specific behavior of each scorer using confusion matrices derived from machine-learning models. After column normalization, these matrices estimate the probability of each true sleep stage given each scorer's label; probabilities are aggregated across scorers to assign the final label for each epoch. LBH was evaluated with random forest, support vector machine, and multilayer perceptron classifiers under EEG-only and EEG+EMG settings, and compared with the dataset hypnogram (DH) and best-scorer hypnogram (BSH). LBH consistently improved overall performance. The best results were obtained with random forest and EEG+EMG, reaching 86.07% accuracy, 85.46% precision, and 85.29% F1-score on DOD-H, and 86.04% accuracy, 85.21% precision, and 84.70% F1-score on DOD-O. These findings suggest that personalized scorer modeling can improve reference hypnogram construction without discarding information from individual experts.
SchemaLink: LinkML スキーマキュレーションのためのインテリジェントな Web エディター
動機: LinkML は、さまざまな種類の生物医学データの構造および内容の制約を表現するのに適した言語です。たとえそれがごく最近の提案であっても、いくつかの生物医学の文脈で適用されています。 LinkML スキーマの開発と維持には、特に初心者のキュレーターにとって、いくつかの課題が伴います。専門家ではないバイオキュレーターは、LinkML の構文とベスト プラクティスに苦労する可能性があり、適切に構造化されたスキーマを開発するには多大な時間と労力を必要とします。結果: この論文では、次の要件に対処する、LinkML スキーマのグラフィカルな構築と拡張のための Web ベースの環境である SchemaLink を提案します: $(i)$ LinkML スキーマの仕様にグラフィカル言語を導入し、 $(ii)$ 同様のコンテキストでスキーマの仕様を統一し、 $(iii)$ キュレーターが新しいスキーマを最初から作成するのを支援する RAG ベースのアプローチを活用することで設計とキュレーションのプロセスを簡素化します。すでに開発されているものを編集します。いくつかの実験的分析により、AI ベースの編集機能を通じて生成された LinkML スキーマの品質が示されています。入手可能性と実装: SchemaLink は、https://SchemaLink.biodata.di.unimi.it からオンラインで入手できます。 SchemaLink コードとテスト データは、GitHub (https://github.com/AnacletoLAB/{schemalink-webapp,schemalink-api}) でオープンソースとして入手できます。
原文 (English)
SchemaLink: An Intelligent Web Editor for LinkML Schema Curation
Motivation: LinkML is a suitable language for the representation of the structural and content constraints of different kinds of biomedical data. Even if it is a quite recent proposal, it has been applied in several biomedical contexts. Developing and maintaining LinkML schemas presents several challenges, particularly for novice curators. Non-expert bio-curators may struggle with LinkML syntax and best practices, requiring significant time and effort to develop well-structured schemas. Results: In this paper we propose SchemaLink, a web-based environment for the graphical construction and enhancement of LinkML schemas that address the following requirements: $(i)$ introduce a graphical language for the specification of LinkML schemas, $(ii)$ make uniform the specification of schemas in similar contexts, $(iii)$ simplify the design and curation processes by exploiting a RAG-based approach to assist curators in creating new schemas from scratch and editing already developed ones. Several experimental analyses show the quality of the produced LinkML schemas through the AI-based editing facilities. Availability and Implementation: SchemaLink is available online at: https://SchemaLink.biodata.di.unimi.it. SchemaLink code and testing data are available as open-source on GitHub at: https://github.com/AnacletoLAB/{schemalink-webapp,schemalink-api}.
すべてのナッジが成功するわけではない: AI 支援ジャーナリングにおける行動の制御性と推敲の品質
AI ジャーナリング ツールは、個人が感知した行動に合わせてプロンプトを調整できますが、どの行動がそれに反応するかは不明です。私たちは、8 週間のパッシブ センシング研究から得た 369 件の仕訳を分析しました。 LLM は、各エントリに行動を変更する意図を表現するかどうかのラベルを付け、26 個のセンサー機能に対するフォロースルーを 3 日間の前後比較で測定しました。反応性は、行動が他の人を巻き込むかどうかに最も左右されます。他者に依存する行動は、改善されたケースの 15 ~ 22% のみでしたが、一人で行動できる行動は、不均一ではありますが、改善されることが多く、最大 50 ~ 63% でした。ユーザーがどのように書いたかはそれほど重要ではありません。改善されていないエントリから分離された単一のテキスト機能はありません。メッセージを送信するのは特定の動作内のみで、最も明確なのはテキスト メッセージングやより長い、より個人的な意図のエントリの場合です。サンプルが小さいため、これらを AI ジャーナリング ナッジが最も機能する可能性が高い場所を示す探索的なパターンとして扱います。
原文 (English)
Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling
AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them. We analyzed 369 journal entries from an eight-week passive sensing study. An LLM labeled each entry as expressing an intention to change a behavior or not, and we measured follow-through against 26 sensor features with a 3-day before/after comparison. Responsiveness depended most on whether a behavior involves other people. Behaviors that depend on others improved in only 15 to 22% of cases, while behaviors a person can act on alone improved more often, up to 50 to 63%, though unevenly. How users wrote mattered less. No single text feature separated improved from unimproved entries; writing carried signal only within specific behaviors, most clearly for text messaging and for longer, more personal intention entries. The sample is small, so we treat these as exploratory patterns that point to where AI journaling nudges are most likely to work.
ピアとは何ですか?プライベート市場における評価に基づく類似性
より多くの投資家がプライベート市場を検討し、限られた透明性、希薄な開示、まれな取引と闘う中、比較対象として経済的に意味のある同業他社を特定することは、評価、デューデリジェンス、ポートフォリオ構築、およびリスク管理における基本的な課題となっています。私たちは、静的な特徴マッチングや意味論的な説明ではなく、市場評価のレンズを通して企業の類似性を定義する、アンサンブル ツリー ベースの教師付き類似性学習フレームワークを提案します。具体的には、観測された民間企業の評価に基づいて CatBoost 勾配ブースト型デシジョン ツリー モデルをトレーニングし、アンサンブル全体にわたる重要度で重み付けされたリーフ ノードの共起から評価を意識した類似性メトリックを導出します。類似性メトリクスは、プライベート市場で一般的な非線形関係、混合データ タイプ、広範な欠損データに対応しながら、共通の評価要因を捕捉します。複数の業界、地域、取引段階にまたがる観察された、または導出可能なポストマネー評価を持つ53,000社以上の企業を含む、約270,000社の世界的なプライベートマーケットユニバースを使用して、提案された類似性フレームワークが、ケースベースの説明可能性を維持しながら、評価される業界グループの下流のk最近傍評価タスクにおける従来の距離ベースおよびテキスト埋め込みベースのアプローチを改善することを実証します。
原文 (English)
What Makes a Peer? Valuation-Anchored Similarity in Private Markets
As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due diligence, portfolio construction, and risk management. We propose an ensemble tree-based supervised similarity learning framework that defines company similarity through the lens of market valuation rather than static feature matching or semantic descriptions. Specifically, we train a CatBoost gradient-boosted decision tree model on observed private company valuations and derive a valuation-aware similarity metric from importance-weighted leaf-node co-occurrences across the ensemble. The similarity metric captures shared valuation drivers while accommodating nonlinear relationships, mixed data types, and pervasive missing data common in private markets. Using a global private-market universe of approximately 270,000 companies, including more than 53,000 firms with observed or derivable post-money valuations spanning multiple industries, geographies, and deal stages, we demonstrate that the proposed similarity framework improves upon traditional distance-based and text-embedding-based approaches in downstream k-nearest-neighbor valuation tasks in the evaluated industry groups, while retaining case-based explainability.
ランダムな低次元再パラメータ化がニューラル ネットワークをトレーニングする時期を予測する
ニューラル ネットワークは、多くの場合、ランダムな低次元の再パラメータ化を通じてトレーニングまたは微調整できます。この再パラメータ化では、小さな潜在ベクトルが、凍結されたランダム マップによって完全なパラメータ更新にマッピングされます。これにより、実際的な問題が生じます。低損失領域に到達するには、潜在探索空間はどのくらいの大きさでなければならないでしょうか?まず、既知のアクセシビリティ遷移を、極円錐の統計的次元にあるコンパクトな凸ターゲットを中心とした同等の円錐形で表現します。私たちの主な理論的貢献は、曲率スペクトルと基準から解への変位プロファイルの両方からランダム スライス残差を予測する方向分解二次マスター公式です。これは、自己矛盾のない等方性方向予測子を生成し、保守的な半径のみの特殊化で、以前のガウス幅の二次境界を回復します。この分析に基づいて、構造化アダマール マップまたはシード再生成ガウス マップを使用して予測された潜在次元をインスタンス化するランダム マッピング ネットワーク (RaMaN) を導入します。これらの構造により、密なランダム マップの O(dP) ストレージが回避され、オプティマイザー状態メモリが O(P) から O(d) に削減されます。また、マトリックスを使用しない曲率近似とスイープを使用しない寸法選択も開発します。制御された二次実験および神経曲率実験全体にわたって、方向分解予測子は測定された遷移位置を厳密に追跡し、変位方向が重要な場合には方向に依存しない近似よりも優れた性能を発揮します。さらに、エンドツーエンドの実験では、画像モデルと言語モデルにわたる、プロトコルに依存したトレーニングの急激な移行が示されています。
原文 (English)
Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks
Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped into a full parameter update by a frozen random map. This raises a practical question: how large must the latent search space be to reach a low-loss region? We first express the known accessibility transition in an equivalent conic form, centered for compact convex targets at the statistical dimension of the polar cone. Our main theoretical contribution is an orientation-resolved quadratic master formula that predicts the random-slice residual from both the curvature spectrum and the reference-to-solution displacement profile. It yields a self-consistent isotropic-orientation predictor and, in a conservative radius-only specialization, recovers the earlier Gaussian-width quadratic bound. Building on this analysis, we introduce Random Mapping Networks (RaMaN), which instantiate the predicted latent dimension using structured Hadamard or seed-regenerated Gaussian maps. These constructions avoid the O(dP) storage of dense random maps and reduce optimizer-state memory from O(P) to O(d). We also develop matrix-free curvature approximations and sweep-free dimension selection. Across controlled quadratic and neural-curvature experiments, the orientation-resolved predictor closely tracks measured transition locations and outperforms orientation-agnostic approximations when displacement direction matters. End-to-end experiments further show sharp, protocol-dependent training transitions across image and language models.
PseudoMapLabeler: 半教師ありオンライン マッピングのための信頼性を意識した疑似ラベル生成
オンライン HD マップ構築システムを現実世界のシナリオに展開する際の重大な課題は、ラベル付きトレーニング データが不足していることであり、これにより多様な環境でのモデルの一般化が制限されます。この制限に対処するために、私たちは、信頼性を意識したマップの改良を通じて、ラベルのないデータから高品質の疑似ラベルを生成する教師と生徒の半教師あり学習 (SSL) フレームワークを提案します。私たちのアプローチでは、まず限られたラベル付きデータで教師モデルをトレーニングし、次にベータ分布ベースの信頼度マップを活用して、時間的観測全体にわたって予測されたマップ要素の信頼性を評価します。要素全体を破棄する従来のフィルタリング手法とは異なり、信頼性の低いセグメントを削除しながら信頼性の高い領域を選択的に保存する空間クリッピング手法を導入します。洗練されたマップ要素は、2 回目のパスでラベルなしデータに対する教師モデルの予測精度を向上させるマップ事前分布として機能します。これらの強化された予測は、スチューデント モデルをゼロからトレーニングするための疑似ラベルとなり、その後、元のラベル付きデータで微調整が行われます。 nuScenes データセットでの実験結果は、洗練された疑似ラベルを備えた教師と生徒のフレームワークにより、ラベル付きデータのみでのトレーニングと比較して、低ラベル体制下でパフォーマンスが +6.1 mAP 向上し、オンライン HD マップ構築におけるラベル付きデータ不足の問題に対する実用的な解決策が提供されることを示しています。
原文 (English)
PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping
A critical challenge in deploying online HD map construction systems to real-world scenarios is the scarcity of labeled training data, which limits model generalization in diverse environments. To address this limitation, we propose a teacher-student semi-supervised learning (SSL) framework that generates high-quality pseudo-labels from unlabeled data through confidence-aware map refinement. Our approach first trains a teacher model on limited labeled data, then leverages Beta-distribution-based confidence maps to assess the reliability of predicted map elements across temporal observations. Unlike conventional filtering methods that discard entire elements, we introduce a spatial clipping technique that selectively preserves high-confidence regions while removing unreliable segments. The refined map elements serve as map priors that improve the teacher model's prediction accuracy on unlabeled data in a second pass. These enhanced predictions become pseudo-labels for training a student model from scratch, followed by fine-tuning on the original labeled data. Experimental results on the nuScenes dataset demonstrate that our teacher-student framework with refined pseudo-labels improves performance by +6.1 mAP under a low-label regime compared to training on labeled data alone, offering a practical solution to the labeled data scarcity problem in online HD map construction.
LLM は優れた戦略家ではないが、記憶力が強化された主体性が推論力を高める
長期的な環境における大規模言語モデル (LLM) での戦略的推論は、一貫性のないサブ目標によって制限されることがよくあります。このような設定では、注目リソースが有限であるため、モデルが数千のステップにわたって戦略的一貫性を維持できなくなります。この制限は、局所的な意思決定が推論全体で一貫した軌道を維持できないという戦略的漂流につながります。これに対処するために、長期的な推論に取り組むためのポリシーとしてエージェントが記憶を学習できるようにするフレームワークである EpicStar を導入します。具体的には、エージェントは、短期的な環境変化を追跡するための作業記憶と並行して、ヒューリスティックとして過去の成功したエピソードのバンクを維持します。推論中、動的なゲート メカニズムは、取得したアクションを直接実行するか、取得したエピソードと現在の作業メモリの文脈上の融合を通じて新しい推論を実行するかを決定します。 StarCraft II をテストベッドとして利用して、さまざまな対戦相手のスタイルに対して EpicStar を評価しました。これはベースラインの方法を大幅に上回っており、消費するトークンの量が桁違いに少ない一方で高い勝率を達成し、難易度や対戦相手の戦略全体でこの利点を一貫して維持します。私たちの調査結果は、LLMエージェントが動的で自律的な設定で堅牢かつ長期的な戦略を実行できるようにするには、構造化されたクロスエピソード記憶が不可欠であるという説得力のある証拠を提供します。
原文 (English)
LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning
Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals. In these settings, finite attention resources prevent the model from maintaining strategic coherence over thousands of steps. This limitation leads to strategic drift, where localized decisions fail to sustain a coherent trajectory across reasoning. To address this, we introduce EpicStar, a framework that enables agents to learn memory as policy to tackle long-horizon reasoning. Specifically, the agent maintains a bank of successful past episodes as a heuristic alongside a working memory to track short-term environmental changes. During inference, a dynamic gating mechanism determines whether to execute a retrieved action directly or to perform new reasoning through a contextual fusion of the retrieved episodes and current working memory. Utilizing StarCraft II as the testbed, we evaluated EpicStar against diverse opponent styles. It significantly outperforms baseline methods, achieving higher win rates while consuming an order of magnitude fewer tokens, and it maintains this advantage consistently across difficulty levels and opponent strategies. Our findings provide compelling evidence that structured cross-episode memory is essential for enabling LLM agents to perform robust, long-term strategic execution in dynamic, autonomous settings.
EgoCITE: 長期自己中心的メモリのためのコンテキスト拡張インデックス作成と時間認識検索
長期的な自己中心的な記憶は、連続する一人称のビデオとオーディオを、検索可能な過去の経験の記録に変換します。既存のシステムには 2 つのボトルネックがあることを示します。文脈に乏しいキャプションから構築されたインデックスはエージェント検索では信頼できません。また、検索では質問の一時的な意図が無視されます。両方のボトルネックに対処するために、自己中心的な QA のための長期的なエージェント メモリ フレームワークである EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval) を導入します。 EgoCITE は 3 つのコンポーネントで構成されます。 EgoScheme は、ローカルのマルチモーダル コンテキストを使用して、断片的なビデオ キャプションと音声トランスクリプトを自己完結型のアトミック メモリ インデックスに変換します。 EgoIndex は、相補的なアクション、アクティビティ、発話、および会話表現を、複数の粒度で検索可能なマルチビュー メモリ インデックスに編成します。 EgoRetrv は、セマンティック検索と、質問条件付きの時間的関連性スコアリングおよび取得された証拠のキュレーションを組み合わせたものです。 EgoLifeQA、EgoMem、および EgoR1-Bench 上の EgoCITE を、回答の精度とターゲットとイベントの検索の整合性の観点から評価します。 EgoCITE は、エージェント メモリ ベースラインの精度を少なくとも 4.4 ~ 14.2\% 向上させ、ロング コンテキスト LLM エージェントよりも 36$\times$ のコスト削減を実現します。
原文 (English)
EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory
Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic search, while retrieval ignores a question's temporal intent. To address both bottlenecks, we introduce EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric QA. EgoCITE comprises three components. EgoScheme uses local multimodal context to turn fragmentary video captions and speech transcripts into self-contained atomic memory indices. EgoIndex organizes complementary action, activity, utterance, and conversation representations into searchable multi-view memory indices at multiple granularities. EgoRetrv combines semantic search with question-conditioned temporal relevance scoring and curation of retrieved evidence. We evaluate EgoCITE on EgoLifeQA, EgoMem, and EgoR1-Bench in terms of answer accuracy and target-event retrieval alignment. EgoCITE improves accuracy over agentic memory baselines by at least 4.4--14.2\% while achieving 36$\times$ lower cost than long-context LLM agents.
言語モデルによって生成された小説は圧縮された形式的なバリエーションを示します
大規模な言語モデルは小説全体を生成できますが、その出力が何世代にもわたって形式的に変化するレベルについての情報はほとんどありません。この研究では、個々のパッセージが AI によって生成されたものとして識別できるかどうかを問うのではなく、AI 生成を繰り返すことで人間の体全体に見られるのと同じ範囲の多様性を生み出すことができるかどうかを問うています。この論文では、生成ソースとターゲット スタイルに基づいて 6 つのコーパスを対比します。19 世紀のイギリス リアリスト スタイルで GPT-5.5 Thinking を使用して生成された 20 の小説、19 世紀のイギリス リアリズム スタイルで Qwen3-14B を使用して生成された 20 の小説、現代のゼロ スタイルでこれらのモデルのそれぞれを使用して生成された 20 の小説、19 世紀に人間が書いたイギリスの小説 205 冊、および人間が書いた現代のゼロ スタイルの小説 65 です。文書レベルの研究には、MATTR-500、シャノンエントロピー、平均文長、読みやすさ、句読点率の測定が含まれます。最も堅牢で信頼性の高い結果は、文構造の圧縮です。世代を重ねることで、人間の小説に比べて文構造の相互の差異がはるかに少ない小説が生み出されます。圧縮は、小説内の読みやすさ、句読点、文の長さのばらつきの尺度にも存在します。 Qwen Zero-Style MATTR を除いて、字句メジャーも同様に圧縮される傾向があります。 GPT と Qwen には、明確な平均文体プロファイルがあるにもかかわらず、メジャー間相関の安定したパターンがありません。したがって、この記事では、小説間の限られた形式的範囲を表す分散オーバークロージャーと、相関オーバークロージャーというより具体的な現象を区別します。これは、AI が生成した個々の小説は文体的に人間の小説に似ている可能性があるが、AI が生成した小説のコレクションが占める形式的な範囲ははるかに狭いことを意味します。
原文 (English)
Novels generated by language models show compressed formal variation
While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as AI-generated, this study asks whether repeated AI generation can produce the same range of diversity which is found across human corpora. This paper contrasts six corpora based on generation source and target style: twenty novels generated using GPT-5.5 Thinking in a nineteenth-century British realist style, twenty novels generated using Qwen3-14B in a nineteenth-century British realist style, twenty novels generated using each of these models in a contemporary zero style, 205 nineteenth-century human-written British novels, and sixty-five contemporary human-written Zero-Style novels. At the document level, the research includes MATTR-500, Shannon entropy, average sentence length, readability, and punctuation rate measurements. The most robust and reliable result is compression of sentence structure. Repeated generations produce novels that vary far less from one another in sentence structure than human novels do. Compression is also present in the measures of readability, punctuation, and sentence length variability within novels. Lexical measures tend to be similarly compressed, with the exception of Qwen Zero-Style MATTR. Despite having distinct mean stylistic profiles, GPT and Qwen lack a stable pattern of cross-measure correlation. This article therefore distinguishes between variance overclosure, which represents a limited formal range between novels, and a more specific phenomenon of correlational overclosure. This means that an individual AI-generated novel may resemble human fiction stylistically, while a collection of AI-generated novels occupies a much narrower formal range.
因果効果制約による解釈可能な因果発見
因果関係の発見は、システムから生成されたデータを基に、根底にある因果関係を明らかにすることを目的としています。ただし、目標は、与えられたデータから因果関係を予測することだけではなく、特に大きな因果効果など、観察された現象または仮説上の現象を解釈して説明できるようにすることでもあります。我々は、この条件付き因果発見のタスクを考慮し、それをベイジアン推論問題としてキャストします。この問題では、因果効果制約などのイベントを条件とする事後因果グラフとパラメータを対象とします。残念なことに、これは計算上の課題を引き起こします。イベントの事後質量が小さい場合、ベイジアン因果発見に対する既存のアプローチは困難を伴います。これに対処するために、まれなイベントの推定手法を適用して、結合グラフ パラメーター空間の推論を実行します。私たちの方法では、条件付き事後分布に近似するサンプルを維持しながら、粒子集団を徐々に制限領域に向かって移動させます。合成グラフでの経験的評価により、小規模および大規模でのアプローチの精度が検証され、サックスタンパク質データセットのケーススタディでは、経路レベルの概要を提供することで科学的探索を支援するために私たちの手法がどのように使用できるかを示します。
原文 (English)
Interpretable Causal Discovery via Causal-Effect Constraints
Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely to predict causal edges given data, but also to be able to interpret and explain either observed or hypothesized phenomena, such as a particularly large causal effect. We consider this task of conditional causal discovery and cast it as a Bayesian inference problem, in which we target the posterior over causal graphs and parameters conditional on an event such as a causal-effect constraint. Unfortunately, this poses a computational challenge: existing approaches to Bayesian causal discovery struggle when the event has small posterior mass. To address this, we adapt rare-event estimation techniques to perform inference the joint graph-parameter space. Our method gradually drives a particle population toward the constrained region while maintaining samples that approximate the conditional posterior. Empirical evaluation on synthetic graphs validates the accuracy of our approach at small and large scales, and we show in a case study on the Sachs protein dataset how our method can be used to aid scientific exploration by providing pathway-level summaries.
制限付きロジット モデリングによる大規模な需要移動推定
商品の需要予測は、店舗の品揃えの最適化に不可欠な要素です。既存の文献は、適切な顧客選択モデルを学習し、このモデルを使用して品揃え提案に関する目的関数の値 (つまり、予想される需要) を決定することに重点を置いています。ただし、多くのカテゴリを含む大規模な品目ユニバースの場合、このアプローチは非効率であることが判明する可能性があり、考えられるすべての品目の品揃えに対して個別の需要予測が必要になります。別のアプローチとして、個別の品目需要予測の効率を組み合わせると同時に、品目の需要と棚上の他の同様の品目の在庫状況との関係を考慮した調整を個別の予測に適用するという方法があります。このアプローチの中心となるのは、需要伝達 (DT) 係数の推定です。これらの DT 係数は、特定のターゲット商品 (顧客が店内を歩いて購入した商品) の需要が、その商品が棚から削除された場合にユニバース内の他の商品にリダイレクトされる割合を表します。私たちは、大規模なアイテムユニバース (100 万以上のアイテムを含む品揃え) でこれらの DT 係数を計算できるアプローチを導入します。カテゴリ内の複数の場所のデータおよび過去の取引データに関する実験は、代替動作に関する特定の合理的な仮定が満たされる場合、私たちの手順が基礎となる DT 係数を正確に推定でき、需要予測の改善につながることを示しています。
原文 (English)
Demand Transfer Estimation at Scale via Restricted Logit Modeling
Item demand forecasting is an integral component of store assortment optimization. Existing literature focuses on learning a suitable customer choice model and using this model to determine the value of an objective function (i.e. expected demand) with respect to an assortment proposal. However, for large item universe with many categories, this approach can prove inefficient, needing a separate demand forecast for every possible item assortment. An alternate approach exists whereby we combine the efficiency of forecasting item demand independently, while at the same time applying adjustments to the independent forecasts that account for the relations between item demand and the availability of other similar items on the shelf. Central to this approach is the estimation of Demand Transfer (DT) coefficients. These DT coefficients represent the percent of a particular target item's (item that the customer walked in the store to buy) demand that is redirected to each other item in the universe should the target item be removed from the shelf. We introduce an approach that allows us to compute these DT coefficients on large item universes (assortments having 1 million+ items). Experiments on data as well as historical transaction data for multiple locations within categories demonstrate that when certain reasonable assumptions about substitution behavior are satisfied, our procedure is able to accurately estimate underlying DT coefficients and lead to improvements in demand forecasting.
Mr3D-VL: マルチパラメトリック 3D 磁気共鳴イメージングのためのジェネラリスト ビジョン言語基礎モデル
マルチパラメトリック磁気共鳴画像法 (mpMRI) は脳腫瘍の診断と治療の基礎ですが、現在の AI モデルは重大な限界に直面しています。自然言語の相互作用と解釈可能性の欠如により、臨床で必要とされる空間情報の統合とクロスモーダル推論が妨げられます。主な課題は、モダリティ間の物理的意味の大きな違い、スキャン間隔による空間的なずれ、神経膠腫のグレーディングなどのタスクにおける複雑な複数の特徴の解釈の必要性から生じます。視覚言語モデル (VLM) はクロスモーダルな理解に有望ですが、既存の手法は主に 2D 画像モデリングに焦点を当てており、3D 体積空間の直接認識は無視されています。 3D VLM は、3D CT イメージングにおけるレポート生成と特徴位置合わせのために提案されていますが、mpMRI アプリケーションでは、複数のイメージング モダリティにわたる協調的な推論が必要ですが、この要件は現在のソリューションでは満たされていません。これに対処するために、マルチパラメトリック 3D MRI 専用の視覚言語基盤モデルである Mr3D-VL を導入します。 40 億のパラメーターを備え、教師なしの事前トレーニング済み共有 3D エンコーダーと 4D 回転位置埋め込みを採用し、デュアル モダリティと空間の統合を実現します。そのクロスモーダル投影レイヤーは、複数解像度の特徴注入戦略を使用して、解像度全体での特徴認識を強化します。実験結果では、テキスト生成タスクにおいて既存の 4B/7B/30B ドメイン固有および汎用モデルに比べて大幅な改善が見られ、レポート生成の BERTScore 0.856、質問応答精度 0.713、多肢選択精度 0.912 を達成しました。
原文 (English)
Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging
Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.
出所の追跡と補完的な LLM ウォーターマークによる改ざんの検出
LLM で生成されたテキストに透かしを入れることは、その出所を追跡するための重要なタスクです。既存の LLM ウォーターマークは編集中に出所を保持しますが、これと同じ堅牢性により、攻撃者は帰属を保持したまま重要なコンテンツを変更することができ、これはピギーバック スプーフィングとして知られる脆弱性です。出所と改ざん証拠を共同で提供する革新的なウォーターマークを紹介します。生成された各トークンに堅牢な信号と脆弱な信号が同時に埋め込まれます。これらのシグナルは同じメカニズムを共有していますが、独立したキーと正規化されたテキストに対する異なるシード ウィンドウを使用するため、一方は編集に対して復元力があり、もう一方は読者に見える変更に対して敏感になります。複数ラウンドの不偏トーナメント再重み付けにより、予想される生成分布が維持される一方、周期的なラウンド割り当てパターンにより 2 つの信号間のトレードオフが制御されます。検出時には、それらのスコアは、無傷、改ざん、透かしなしの 3 つの決定をサポートする 2 次元空間を形成します。 2 つの大規模な言語モデルと 2 つのプロンプト データセットにわたって、私たちの手法は、競合するアトリビューションの堅牢性と複雑さを維持しながら、評価された手法の中で最も高い改ざん検出率を示しています。アブレーション研究によると、信頼性の高いスリーステート検出には、完全性、2 つの信号の同時埋め込み、および編集に対する相補的な感度についての明確に定義された概念が必要です。
原文 (English)
Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks
Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness allows an adversary to alter critical content while retaining attribution, a vulnerability known as piggyback spoofing. We introduce an innovative watermark that jointly provides provenance and tamper evidence. It co-embeds a robust signal and a fragile signal into each generated token. The signals share the same mechanism but use independent keys and different seeding windows over normalized text, making one resilient to edits and the other sensitive to reader-visible changes. Multiple rounds of unbiased tournament reweighting preserve the expected generation distribution, while a periodic round-allocation pattern controls the trade-off between the two signals. At detection, their scores form a two-dimensional space supporting three decisions: Intact, Tampered, and No-Watermark. Across two large language models and two prompt datasets, our method demonstrates the highest tamper-detection rate among the evaluated methods while maintaining competitive attribution robustness and perplexity. Ablation studies show that reliable three-state detection requires a well-defined notion of intactness, co-embedding of the two signals, and complementary sensitivity to edits.
HybridSB-MoE: 音声強化のためのシーン適応エキスパート ルーティングを備えたデュアル ドメイン シュレディンガー ブリッジ
生成音声強調は 3 つのギャップに直面しています。スペクトル モデルは調波構造を捕捉しますが、しばしば位相を乱します。波形モデルは位相を保持しますが高調波を見逃します。シュルオーディンガー ブリッジ (SB) はノイズからクリーンな音声への伝達を短縮しますが、推論コストはトレーニングに緩く結び付けられるだけです。私たちは、単一の非対称設計原理によって統合された 3 つの貢献を通じてこれらのギャップを埋めるデュアルドメイン フレームワークである HybridSB-MoE を提案します。(i) 非対称不確実性の融合: スペクトル パスは、 (ii) 5 つの異なるアーキテクチャ原型にわたる top-k=2 ルーティングによる異種 MoE により、同様の専門家間の小さな摂動ではなく認識信号が示されます。離散化限界 (定理 1): パス一貫性と軌道正則化により、レート K-alpha での 2-Wasserstein 距離の K ステップ ブリッジ サンプリング エラーが制限され、VoiceBank+DEMAND では、HybridSB-MoE は、一貫性蒸留との競争力を維持しながら、ステップ バジェットで拡散および SB ベースのベースラインを上回るパフォーマンスを実現します。数ステップの方法。
原文 (English)
HybridSB-MoE: Dual-Domain Schr\"odinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement
Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserve phase but miss harmonics, and Schr\"odinger Bridges (SB) shorten transport from noise to clean speech but leave inference cost only loosely tied to training. We propose HybridSB-MoE, a dual-domain framework that fills these gaps through three contributions unified by a single asymmetric design principle. (i) Asymmetric uncertainty fusion: The spectral path captures epistemic uncertainty via expert disagreement, while the waveform bridge models aleatoric variance through stochastic dynamics. We fuse them asymmetrically, allowing the mixing weight to adapt to distinct error regimes rather than average predictions. (ii) Heterogeneous MoE with top-k=2 routing across five distinct architectural archetypes, where architectural diversity makes the epistemic signal indicate which inductive bias fails rather than small perturbations among similar experts. (iii) Discretization bound (Theorem 1): path-consistency and trajectory regularizers together bound the K-step bridge sampling error in 2-Wasserstein distance at rate K-alpha, making small-K inference an objective-level guarantee rather than an empirical claim. On VoiceBank+DEMAND, HybridSB-MoE outperforms diffusion- and SB-based baselines at their step budgets while remaining competitive with consistency-distilled few-step methods.
大規模言語モデルルーティングのためのエラー認識型逆オークションメカニズム
各クエリを費用対効果の高い大規模言語モデル (LLM) にルーティングすることは、品質とコストのバランスをとるために重要ですが、ほとんどのルーターは集中タスク センターに依存してモデルのパフォーマンスを予測するため、モデル プールが増大するにつれて情報リスクの不一致とスケーラビリティのボトルネックが生じます。私たちは、プロバイダーが自己予測した成功確率と実行コストで入札する逆オークションを通じて、事前予測を LLM プロバイダーに移す市場ベースのルーティング パラダイムを提案します。本質的にノイズの多いプロバイダー予測とセンター評価を考慮するために、この固有のデュアル エラーを明示的にモデル化する \textit{\textbf{E}rror-\textbf{A}ware \textbf{R}everse \textbf{A}uction \textbf{M}echanism} (EA-RAM) を導入します。我々は、EA-RAMがベイジアン・インセンティブと互換性があり、二重誤差の下で個別に合理的であることを証明し、中心合理性のための十分な条件を確立し、明示的な福祉損失限界を導き出します。さらに、ロバスト性の効果を特定します。逆符号の誤差はキャンセルされる可能性があり、バニッシングテールリンク関数(ロジスティックなど)は飽和を通じて明確なケースを安定化し、余分なノイズは信念マップを平滑化し、限界操作からのゲインを減少させます。シミュレーションと現実世界のベンチマークに関する実験では、EA-RAM がデュアル エラーに対して堅牢であり、集中ベースラインよりも優れたコスト パフォーマンスのパレート フロンティアを達成し、プロバイダーがローカル情報を提供すると追加の利益が得られることが示され、その実際的な有効性が検証されています。
原文 (English)
Error-Aware Reverse Auction Mechanism for Large Language Model Routing
Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We propose a market-based routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers bid with self-predicted success probabilities and execution costs. To account for inherently noisy provider predictions and center evaluations, we introduce the \textit{\textbf{E}rror-\textbf{A}ware \textbf{R}everse \textbf{A}uction \textbf{M}echanism} (EA-RAM), which explicitly models this inherent Dual Error. We prove that EA-RAM is Bayesian incentive compatible and individually rational under the Dual Error, establish sufficient conditions for center rationality, and derive an explicit welfare-loss bound. We further identify robustness effects: opposite-signed errors can cancel, vanishing-tail link functions (e.g., logistic) stabilize clear-cut cases via saturation, and extra noise smooths belief maps, reducing the gains from marginal manipulation. Experiments on simulations and real-world benchmarks show that EA-RAM is robust to the Dual Error and achieves a better cost--performance Pareto frontier than centralized baselines, with additional gains when providers contribute local information, validating its practical effectiveness.
ERSkill: スキルガイドによる適応型記憶検索の進化
大規模言語モデル (LLM) エージェントは永続的な対話のために長期記憶にますます依存していますが、この記憶を管理する検索メカニズムが進化可能なコンポーネントとして扱われることはほとんどありません。この静的アプローチは、多くの場合、多様な証拠構築戦略を必要とする異種メモリ クエリのパフォーマンスを制限します。これに対処するために、自己進化するスキルに基づいたメモリ アクセスのための検索中心のフレームワークである \textbf{ERSkill} を導入します。 ERSkill は、インタラクション履歴を構造化メモリ ストアにコンパイルし、基本的なプリミティブで構成される実行可能なスキルとして検索動作を表します。推論時には、トレーニングされたルーターが各クエリを最適なスキルに動的に照合して、回答生成のためのカスタマイズされた証拠を構築します。継続的な改善を可能にするために、ERSkill はトレーニング中にスキル セットとルーターを共同進化させます。新しいスキル機能の拡張を安定したルーター側の展開から安全に切り離すダブルフロンティア メカニズムと並行して、探索された取得パスを効率的に記録するためのエクスペリエンス トライを採用しています。複数のエージェント メモリ ベンチマークにわたる実験では、ERSkill が強力な非進化ベースラインおよび自己進化ベースラインを大幅に上回るパフォーマンスを示しています。特に、F1、BLEU-1、LLM ジャッジ スコア全体の平均が Qwen3-Next-80B-A3B-Instruct で 31.3\%、GPT-5.4-nano で 28.1\% 向上しました。
原文 (English)
ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval
While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components. This static approach limits performance on heterogeneous memory queries, which often demand diverse evidence construction strategies. To address this, we introduce \textbf{ERSkill}, a retrieval-centric framework for self-evolving, skill-guided memory access. ERSkill compiles interaction histories into a structured memory store and represents retrieval behaviors as executable skills composed of fundamental primitives. At inference time, a trained router dynamically matches each query to the optimal skill to construct tailored evidence for answer generation. To enable continuous improvement, ERSkill co-evolves the skill set and the router during training. It employs an experience trie to efficiently record explored retrieval paths, alongside a double-frontier mechanism that safely decouples the expansion of new skill capabilities from stable, router-facing deployment. Experiments across multiple agent memory benchmarks demonstrate that ERSkill substantially outperforms strong non-evolving and self-evolving baselines. Notably, it improves the overall average across F1, BLEU-1, and LLM-judge scores by 31.3\% with Qwen3-Next-80B-A3B-Instruct and by 28.1\% with GPT-5.4-nano.
PatientAct: 理論に基づいたメンタルヘルス クライアント シミュレーション
LLM ベースの模擬クライアントは、初心者カウンセラーのトレーニング、LLM セラピストの評価、合成データの生成に使用されることが増えています。しかし、現在のシミュレーターは、過度に協力的なクライアントを生み出し、あまりにも簡単に開示し、抵抗なく治療の再構成を受け入れ、単一のセッション内で中核的な問題を解決します。これらの問題は、因果関係の深さが欠けているプロファイルと、すべてのコンテンツを同等にアクセスできるものとして扱う動作メカニズムに起因すると考えられます。確立された臨床理論に基づいたクライアント シミュレーションのフレームワークである PatientAct を紹介します。当社のプロファイルは 5P の臨床症例定式化を統合しており、設計を単一の治療法に結び付けることなく因果関係の深さを提供します。シミュレーション中、プロファイルには、項目が信頼しきい値を保持する動的記憶層が含まれます (たとえば、症状は早期に入手可能ですが、形成的記憶には持続的な治療連携が必要です)。各ターンで、クライアントの感情的な反応と行動がモデル化されてから、応答が生成されます。セラピストがゲートされたコンテンツにアプローチする場合、PatientAct は、協力や単一の抵抗パターンをデフォルトとするのではなく、量、内容、スタイルの観点から抵抗を表現します。私たちは 40 の臨床状況に関するフレームワークを評価し、それが臨床的妥当性の高い多様なプロファイルを生成することを実証します。さらに、PatientAct はベースラインを大幅に上回り、耐性の質と行動の現実性が大幅に向上しました。私たちのコードとデータは、github.com/Sahandfer/PatientHub 経由で公開されます。
原文 (English)
PatientAct: Theory-Grounded Mental Health Client Simulation
LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes without resistance, and resolve core issues within a single session. We trace these issues to profiles that lack causal depth and behavioral mechanisms that treat all content as equally accessible. We present PatientAct, a framework for client simulation grounded in established clinical theories. Our profiles integrate the 5Ps clinical case formulation, providing causal depth without tying the design to any single therapeutic modality. During simulation, profiles include a dynamic memory layer in which items carry trust thresholds (e.g., symptoms are available early, whereas formative memories require a sustained therapeutic alliance). At each turn, the client's emotional reaction and behavior are modeled before generating a response. If the therapist approaches gated content, PatientAct expresses resistance in terms of quantity, content, and style rather than defaulting to cooperation or a single resistance pattern. We evaluate our framework on 40 clinical situations and demonstrate that it generates diverse profiles with high clinical plausibility. Moreover, PatientAct significantly outperforms the baselines, yielding substantial gains in resistance quality and behavioral realism. Our code and data will be publicly available via github.com/Sahandfer/PatientHub.
SynAct: 適応合成最適化のための推論機能を持つ大規模言語モデル エージェント
論理合成は RTL デザインをゲート レベルのネットリストに変換します。PPA の結果は最適化コマンドの選択に非常に影響されるため、合成チューニングは高次元で高価なものになります。以前のアプローチは 2 つのカテゴリに分類されます。1 つは意思決定レベルの解釈可能性が限られた固定アクション空間に対してブラックボックス検索を実行する自動化手法、もう 1 つは LLM ベースの手法で、通常は静的スクリプトを事前に生成し、進化する回路状態に適応できません。適応閉ループ LLM 推論である SynAct を紹介します。これは、ライブ合成レポートと、現在の回路状態、取得したツールの知識、および過去の最適化経験に関する推論を繰り返し診断して、目的のコマンドを発行するためのエージェントです。 SynAct は、エリアと電力のバランスの取れたトレードオフを維持しながら、タイミング、特に最悪のネガティブ スラック (WNS) を改善することに重点を置いています。市販の合成ツールを使用した 14 の設計にわたる実験では、SynAct が平均 WNS をブートストラップ合成の場合の 27% に削減することが示されました。
原文 (English)
SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization
Logic synthesis transforms RTL designs into gate-level netlists, where PPA results are highly sensitive to the choice of optimization commands, making synthesis tuning both high-dimensional and expensive. Previous approaches fall into two categories: automated methods, which perform black-box search over fixed action spaces with limited decision-level interpretability, and LLM-based methods, which typically generate static scripts upfront and cannot adapt to evolving circuit states. We present SynAct, an adaptive closed-loop LLM reasoning--acting agent that iteratively diagnoses live synthesis reports and reasons over the current circuit state, retrieved tool knowledge, and historical optimization experience to issue targeted commands. SynAct focuses on improving timing, particularly worst negative slack (WNS), while maintaining balanced area and power trade-offs. Experiments on a commercial synthesis tool across 14 designs show that SynAct reduces average WNS to 27% of that from bootstrap synthesis.
成果報酬を超えて: 深層検索エージェント向けのステップレベルの自己抽出ポリシーの最適化
深層検索エージェントは数十のステップにまたがる軌跡にわたって動作しますが、標準的な強化学習では軌跡ごとに 1 つの結果報酬しか提供されず、効果的な単位の割り当てにはあまりにもまばらすぎます。オンポリシー自己蒸留 (OPSD) は、モデル自体のロジットを高密度のトークンレベルの教師として使用することでこの問題に対処しますが、それを検索エージェントに拡張すると、根本的な緊張が生じます。教師は、正解などの特権情報にアクセスできるため、生徒の探索ベースの推論とは体系的に異なる分布を生成し、単純な蒸留により、生徒はより良い検索戦略を学習するのではなく、この情報の非対称性を継承することになります。私たちは 2 つの貢献を通じてこの緊張を解決します。まず、Web から抽出された簡潔なステップレベルの証拠スニペットである証拠アンカーを、解答経路全体を明らかにすることなく主要な推論ステップを捕捉する特権情報として構築します。 2 番目に、教師と生徒の意見の不一致を GRPO 内のステップレベルの利点の重みに変換し、誤った軌道のみに適用するステップレベルの自己抽出ポリシー最適化 (SSPO) を提案します。この設計では、何を更新するか、どれだけ更新するかが切り離されています。つまり、結果の報酬が政策変更の方向性を決定し、教師が各ステップでその大きさを調整します。正しい軌道はそのまま残され、その多様性が保たれます。 Qwen3-8B では、SSPO は BrowseComp、GAIA、および FRAMES 全体で一貫して GRPO を上回り、2 倍の勾配ステップでトレーニングされた GRPO を上回るかそれに匹敵しますが、1 回の追加の前方パスによるステップあたりのオーバーヘッドの追加はわずか約 5% です。
原文 (English)
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment. On-policy self-distillation (OPSD) addresses this by using the model's own logits as dense token-level teachers, but extending it to search agents introduces a fundamental tension: the teacher, having access to privileged information such as the correct answer, produces a distribution that differs systematically from the student's exploration-based reasoning, and naive distillation causes the student to inherit this information asymmetry rather than learn better search strategies. We resolve this tension through two contributions. First, we construct Evidence Anchors, which are concise, step-level evidence snippets extracted from the web, as privileged information that captures key reasoning steps without revealing the entire answer path. Second, we propose Step-Level Self-Distilled Policy Optimization (SSPO), which converts teacher-student disagreement into step-level advantage weights within GRPO, applied exclusively to incorrect trajectories. This design decouples what to update from how much to update: the outcome reward determines the direction of policy change, while the teacher modulates its magnitude at each step. Correct trajectories are left untouched, preserving their diversity. On Qwen3-8B, SSPO consistently outperforms GRPO across BrowseComp, GAIA, and FRAMES, surpassing or matching GRPO trained with twice as many gradient steps while adding only about 5 percent overhead per step from a single additional forward pass.
コードLLMの記憶診断はスケールを意識する必要がある
コードの大規模な言語モデルが真の理解よりも暗記にどの程度依存しているかについては、依然として大いに議論されています。現在の文献では、広範な暗記が頻繁に報告されていますが、高密度アーキテクチャ全体で基礎となるプローブ技術を評価すると、その有用性が大規模に低下していることが明らかになります。同義語ファジングやデッドコード挿入などの摂動を使用する従来のエンコーダ スタイルのプローブは、既知の汚染されたベンチマークであっても、スケーリングされたモデルの記憶を明らかにするのに苦労し、対数確率に依存するデコーダ スタイルのプローブでも同様のパフォーマンスの低下が見られます。これらのプローブの特定の失敗モード、特にそのような手法が小規模なモデルを混乱させるのに、より大きなモデルには影響を与えられない理由は、それらを単一の現象として扱うのではなく、記憶による表現負荷を解きほぐす動機になります。数値問題に可逆数学的変換を適用することで、これら 2 つの要因を分離し、スケーリングされたエンコーダーが適切な解群に収束しながら、実質的な表現負荷をうまく吸収できることを明らかにしました。実際のソフトウェア エンジニアリングでは、さまざまな表面形状に適応するこの能力が、LLM およびエージェント アプリケーションの使いやすさと汎用性にとって本当に重要です。トレーニング中に特定の解決策が見られたかどうかは、それほど差し迫った問題ではなくなります。なぜなら、暗記によって汚染されたベンチマークのスコアが膨らむ一方で、表現負荷を考慮に入れると、機能的な答えが最初に記憶されていたかどうかを実際にどの程度気にする必要があるかについては議論の余地があるからです。したがって、将来の評価は、これらの現象を静かに絡み合わせる方法論に依存するのではなく、これらの現象を分離することを中心に構築する必要があります。
原文 (English)
Memorization Diagnostics for Code LLMs Should be Scale-Aware
The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread memorization, evaluating the underlying probing techniques across dense architectures reveals a severe breakdown in their utility at scale. Traditional encoder-style probes using perturbations such as synonym fuzzing or dead-code insertion struggle to expose memorization in scaled models, even on known-contaminated benchmarks, and decoder-style probes that rely on log probabilities show similar performance degradation. The specific mode of failure for these probes, particularly why such techniques disrupt smaller models but fail to impact larger ones, motivates us to untangle representation load from memorization rather than treating them as a single phenomenon. By applying invertible mathematical transforms to numeric problems, we isolate these two factors and reveal that scaled encoders successfully absorb substantial representation load while still converging on the correct family of solutions. In practical software engineering, this ability to adapt to varying surface forms is what truly matters for usability and generalizability in LLM and agentic applications. Whether a specific solution was seen during training becomes a much less pressing question because although memorization inflates scores on contaminated benchmarks, factoring out representation load makes it debatable how much we should truly care if a functional answer was originally memorized. Future evaluations must therefore be built around separating these phenomena rather than relying on methodologies that quietly entangle them.
CRAFT: 臨床物語に対する時間的推論のための LLM ベースの反復的洗練
臨床ナラティブにおける症状の時間的進行を理解することは、疾患のモニタリング、安全性の監視、因果関係の評価にとって重要です。しかし、臨床の物語が明示的な時間的アンカーを提供することはほとんどありません。時間情報推論への現在のアプローチは、主に複数訪問およびタイムスタンプが豊富な記録にわたるペアワイズ関係の分類に焦点を当てており、個々のアンカーがまばらなレポートからの構造化された症状の軌跡の再構築にはほとんど対処されていません。私たちは、ジェネレーターと制約ベースの検証機能を組み合わせて、ターゲットを絞ったフィードバックを通じて段階ごとの症状タイムラインを反復的に生成および改良する LLM フレームワークである CRAFT を提案します。私たちは、3 つの新型コロナウイルス感染症 (COVID-19) ワクチンにまたがる 5,347 件のワクチン有害事象ナラティブの新しいベンチマークである MedTempo の評価を実施し、3,166 件のレポートに対して専門家が検証した時間的段階の注釈を付けています。 4 つの LLM バックボーンにわたる実験では、CRAFT がモデルの機能レベル全体でジェネレーターとベリファイアーのコンポーネントの寄与を分離するアブレーション解析により、時間的順序付けの精度を一貫して向上させることが実証されています。
原文 (English)
CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives
Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors. Current approaches to temporal information reasoning focus predominantly on pairwise relation classification across multi-visit and timestamp-rich records, leaving the reconstruction of structured symptom trajectories from individual anchor-sparse reports largely unaddressed. We propose CRAFT, an LLM framework that pairs a generator with a constraint-based verifier to iteratively produce and refine stage-wise symptom timelines through targeted feedback. We conduct evaluation on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports. Experiments across four LLM backbones demonstrate that CRAFT consistently improves temporal ordering accuracy, with ablation analysis isolating the contribution of generator and verifier components across model capability levels.
PIPES: 来歴と事前情報を使用してエージェントの認識を確保する
ツールを使用するエージェントは、さまざまな信頼レベルのソースからの外部データを消費しますが、ツールの応答では、各コンポーネントの作成者やコンポーネントが何を伝えるべきかを特定することはほとんどありません。我々は、このギャップによって状態破壊攻撃が可能になることを示します。この攻撃では、攻撃者が制御するコンテンツが、その応答コンポーネントの情報権限を超えて環境に関する主張を行い、エージェントが知覚する環境を破壊し、結果として生じるアクションが既存のガードレールに対して正当であるかのように見せかけます。 PIPES (Provenance-Informed, Prior-Enforced Screening) を導入します。これは、意味論的な事前情報とソースの出所を使用して応答ユニットをスクリーニングします。 PIPES は、スキーマが安定した期待値を提供する場合に静的フィールド コントラクトを使用し、応答前の軌跡と信頼できる来歴メタデータに基づいてオープンエンド コンテンツのスクリーニングを条件付けします。これは、意味論的な事前階層または来歴階層に違反するユニットをマークします。導入では、検出された違反を削除、警告、ブロック、またはエスカレーションする場合があります。アトミックな削除をインスタンス化し、適応型 PAIR スタイルの攻撃に対して PIPES を評価します。 Gemma 4 31B IT をターゲット エージェントとする 3 つの VitaBench と 3 つの AgentDyn 分割全体で、PIPES は平均的な攻撃成功率を 84.7% から 2.3% に低下させますが、平均的な良性のユーティリティ (PIPES を使用した場合は 92.5%、防御なしの場合は 90.6%) を維持します。
原文 (English)
PIPES: Securing Agent Perception with Provenance and Priors
Tool-using agents consume external data from sources with different levels of trust, yet tool responses rarely identify who produced each component or what it should convey. We show that this gap enables state-corruption attacks, in which attacker-controlled content makes environmental claims beyond the informational authority of its response component and corrupts the agent's perceived environment, making the resulting action appear justified to existing guardrails. We introduce PIPES (Provenance-Informed, Prior-Enforced Screening), which screens response units using semantic priors and source provenance. PIPES uses static field contracts when schemas provide stable expectations, and conditions screening of open-ended content on the pre-response trajectory and trusted provenance metadata. It marks units that violate their semantic prior or the provenance hierarchy; deployments may remove, warn, block, or escalate detected violations. We instantiate atomic removal and evaluate PIPES against adaptive PAIR-style attacks. Across the three VitaBench and three AgentDyn splits with Gemma 4 31B IT as the target agent, PIPES reduces average attack success from 84.7% to 2.3%, while preserving average benign utility (92.5% with PIPES versus 90.6% without defense).
消去して保存: 最適化されたセマンティック アンカーを介して、著作権で保護されたアニメーション キャラクターの制御可能な削除
テキストから画像への拡散モデルの優れた生成機能により、著作権、特にアニメーション キャラクターの無許可複製に関する懸念が生じています。既存の概念消去方法は、アニメーション キャラクターの消去には不十分です。モデル変更方法では、多様で非常に特徴的なキャラクターに適したアンカーを特定するのが困難です。プロンプトベースのステアリング方法には、正確な介入のためのきめの細かい制御が欠けています。これらのアプローチでは、多くの場合、消去が不完全になり、画像の忠実度が低下し、実際の展開が妨げられます。この論文では、モデルの連続テキスト表現を操作して、生成中にターゲット文字を消去する制御可能な方法を提案します。構造的および詳細な制約を介してアンカーの埋め込みを最適化して文字の代理として機能させ、その後、構造を意識した適応戦略によってターゲット関連の埋め込みをアンカーに置き換えます。実験では、私たちの方法が最先端の消去効果と画像忠実度の維持を達成しながら、制御可能な消去度、マルチターゲット除去、およびモデルの転送可能性をサポートしていることが示されています。さらに、当社の最適化されたアンカーは、現在のモデル変更ベースラインとプラグアンドプレイで対応し、消去パフォーマンスを向上させます。
原文 (English)
Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors
The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized reproduction of animation characters. Existing concept erasure methods fall short for animation character erasure: model modification methods struggle to identify suitable anchors for diverse, highly distinctive characters; prompt-based steering methods lack fine-grained control for precise intervention. These approaches often yield incomplete erasure and degraded image fidelity, hindering real-world deployment. In this paper, we propose a controllable method operating on the model's continuous textual representation to erase target characters during generation. We optimizes an anchor embedding via structural and detailed constraints to serve as a character surrogate, then replaces target-related embeddings with the anchor via a structure-aware adaptive strategy. Experiments show that our method achieves state-of-the-art erasure effectiveness and image fidelity preservation, while supporting controllable erasure degree, multi-target removal, and model transferability. Moreover, our optimized anchors are plug-and-play with current model modification baselines to improve their erasure performance.
高速 A/B/n テスト: ツリー結合されたフィードバック共有による正確な複数ポリシーの比較
オンライン プラットフォームでは、多くの適応的意思決定ポリシー (ランキング システム、推奨アルゴリズム、価格設定ルール、言語モデル エージェント) を比較することが増えていますが、報酬を伴うインタラクションはそれぞれコストが高くついたり、リスクを伴う可能性があります。直接 A/B/n 設計では、$J$ ポリシーのそれぞれに独自の水平線から $T$ までの軌跡が与えられるため、$JT$ の結果が使用されます。任意の履歴に依存するコンテキスト バンディット ポリシーのための正確なフィードバック共有設計である、ツリー結合 A/B テスト (\TCAB) を導入します。各ラウンドで、予測可能なツリーが現在のポリシー履歴を結び付けます。すべての親、子コンテキスト、アクションの法則が最大限に結合され、一致するツリー エッジの各コンポーネント内で 1 つの報酬が共有されます。すべてのポリシーは、たとえポリシーが意図的に依存しているとしても、そのスタンドアロンの有限地平線軌道法則を正確に保持します。 $D_{e,t}$ がラウンド $t$ でツリーエッジ $e$ の不一致を記録した場合、報酬クエリの数は経路単位恒等 $N(T)=T+\sum_{t,e}D_{e,t}$ を満たすため、$T$ に期待値の累積ツリーエッジ合計変動を加えたものに等しくなります。このコストは、選択されたツリー上の正確なエッジローカル設計の中で条件付きで最適であり、現在のラウンドの最小スパン ツリーはツリー設計の中で近視眼的に最適です。固定 $J$ の場合、すべてのポリシーのサブリニア疑似リグレット、およびオラクル アクションのほぼ確実な一意性は、独立実行の $JT$ に対して $\mathbb{E}[N(T)]=T+o(T)$ を意味します。また、ペアごとのポリシー対比の有限サンプル分散限界も取得します。報酬モデルの評価、多肢選択言語モデルの評価、および適応型検索ポリシーに関する実験により、コストと精度のフロンティアが大幅に向上することが実証されました。
原文 (English)
Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing
Online platforms increasingly compare many adaptive decision policies---ranking systems, recommendation algorithms, pricing rules, and language-model agents---while each reward-bearing interaction can be costly or risky. A direct A/B/n design gives each of $J$ policies its own horizon-$T$ trajectory and therefore uses $JT$ outcomes. We introduce Tree-Coupled A/B Testing (\TCAB), an exact feedback-sharing design for arbitrary history-dependent contextual-bandit policies. At each round, a predictable tree connects the current policy histories; every parent--child context--action law is maximally coupled, and one reward is shared within each component of matched tree edges. Every policy retains exactly its standalone finite-horizon trajectory law, even though the policies are deliberately dependent. If $D_{e,t}$ records a mismatch on tree edge $e$ at round $t$, the number of reward queries satisfies the pathwise identity $N(T)=T+\sum_{t,e}D_{e,t}$ and hence equals $T$ plus cumulative tree-edge total variation in expectation. This cost is conditionally optimal among exact edge-local designs on the selected tree, and a current-round minimum-spanning tree is myopically optimal among tree designs. For fixed $J$, sublinear pseudo-regret of every policy and almost-sure uniqueness of the oracle action imply $\mathbb{E}[N(T)]=T+o(T)$, versus $JT$ for independent runs. We also obtain finite-sample variance bounds for pairwise policy contrasts. Experiments on reward-model evaluation, multiple-choice language-model evaluation, and adaptive search policies demonstrate substantial improvements in the cost--precision frontier.
原子的証拠から論理構成まで: 複合的な回答オプションに対する構造化された構成推論
大規模な言語モデルは、回答オプションが明示的な論理演算子に基づいてアトミックな判断を組み合わせる必要がある場合、個々のアトムを正しく判断した場合でも、失敗することがよくあります。私たちは、AND、OR、NEITHER/NOR で接続された複合オプションを研究し、各オプションをアトミックな答えに分解し、それぞれについて対照的な仮説をスコアリングするフレームワークを導入することで、モデルが複合オプションを認識しないようにします。次に、演算子制約付き整数線形プログラムが、校正されたスコアを単一の予測に合成します。 LOGICAL-COMMONSENSEQAで評価し、SATA-Benchから派生した読解ベンチマークであるLOGICAL-SATAを紹介します。私たちのフレームワークは、人間が検証した LOGICAL-COMMONSENSEQA 分割では Macro-F1 を 48.3 から 77.0 に、LOGICAL-SATA では 47.0 から 75.6 に改善し、NEITHER/NOR で最大の改善が見られます。
原文 (English)
From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options
Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and NEITHER/NOR, introducing a framework that decomposes each option into atomic answers and scores contrastive hypotheses about each one, so the model never sees a compound option. An operator-constrained integer linear program then composes the calibrated scores into a single prediction. We evaluate on LOGICAL-COMMONSENSEQA and introduce LOGICAL-SATA, a reading-comprehension benchmark derived from SATA-Bench. Our framework improves Macro-F1 from 48.3 to 77.0 on the human-validated LOGICAL-COMMONSENSEQA split and from 47.0 to 75.6 on LOGICAL-SATA, with the largest gains on NEITHER/NOR.
AQuA: 再帰的自己改善型定量取引リサーチ エージェント
私たちは定量的投資研究のレベルで再帰的自己改善を研究します。つまり、自律システムが初期の実験からの証拠を使用して、後の反復で提案される仮説や候補を改善できるかどうかを研究します。我々は、2 つの個別の言語モデル駆動型研究システムで構成される AQuA を紹介します。1 つは記号因子発見用、もう 1 つは訓練可能なモデル開発用です。 2 つのシステムは、エージェント、メモリ、候補空間、研究状態を共有しません。代わりに、それぞれが独立して、検証された証拠を保持し、それを次の提案の指針として使用することで、独自の研究ループを閉じます。この限定された意味で、両方のシステムは研究プロセスのレベルで再帰的な自己改善を実装します。各システムは、独自の密閉されたサンドボックスも使用します。これにより、データ分割、特徴とラベルの定義、およびエバリュエーターが修正され、制約された因子式または構成差分を通じてのみモデルが動作できるようになります。マネージャーが仲介するマルチエージェント パイプラインであるファクター システムは、ファクターを検出してシグナルに結合し、暗号通貨ユニバースでの結合情報係数が約 $0.190 に達します。ハイブリッド時系列アーキテクチャ上の構成主導型ループであるモデル システムは、米国株の銘柄ごとの情報係数 $+0.0843$ に達し、それを 2 レッグ コストで最大 $+2.50$ のホールドアウト シャープを持つ閾値ロング/ショート戦略に変換します。この戦略は 2021 年から 2025 年まで毎年前向きです。
原文 (English)
AQuA: Recursively Self-Improving Quantitative Trading Research Agents
We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises two separate language-model-driven research systems: one for symbolic factor discovery and one for trainable model development. The two systems do not share agents, memories, candidate spaces, or research state. Instead, each independently closes its own research loop by retaining validated evidence and using it to guide subsequent proposals. In this bounded sense, both systems implement recursive self-improvement at the level of the research process. Each system also uses its own sealed sandbox, which fixes the data splits, feature and label definitions, and evaluator while allowing the model to act only through constrained factor expressions or configuration diffs. The factor system, a manager-mediated multi-agent pipeline, discovers and combines factors into a signal that reaches a combined information coefficient of about $0.190$ on a crypto universe. The model system, a config-driven loop over a hybrid time-series architecture, reaches a per-stock information coefficient of $+0.0843$ on US equities and converts it into a threshold long/short strategy with a held-out Sharpe of up to $+2.50$ at a two-leg cost. The strategy is positive in every year from 2021 to 2025.
テキストベースの人物異常検索のための不一致を認識した再ランキングを備えた異種視覚言語アンサンブル
テキストベースの人物異常検索は、自然言語記述を使用して、大規模な画像ギャラリーから異常な行動を示している歩行者を検索することを目的としています。従来のテキストベースの人物検索と比較して、このタスクでは歩行者の外観、行動、物体の相互作用、シーンのコンテキストに対するきめ細かい推論が必要となるため、堅牢なクロスモーダルマッチングが大幅に困難になります。この論文では、AI City Challenge 2026 Track 4 に対する GENAI4E チームのソリューションを紹介します。私たちのフレームワークは、強力な検索バックボーンに基づいて構築されており、スコア アラインメントと反復アンサンブル融合を通じて、異種の視覚言語埋め込みモデルを段階的に統合し、その後、あいまいなクエリに対する意見の相違を認識した VLM 再ランキングを行います。公式の歩行者異常行動 (PAB) ベンチマークでは、私たちのアプローチは 90.92% mAP、85.13% Recall@1、97.72% Recall@5、98.68% Recall@10 を達成し、大規模なテキストベースの人物の異常検索において、相補的な視覚言語表現と選択的マルチモーダル推論を組み合わせる有効性を実証しています。
原文 (English)
Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval
Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions. Compared with conventional text-based person retrieval, this task requires fine-grained reasoning over pedestrian appearance, behaviors, object interactions, and scene context, making robust cross-modal matching significantly more challenging. This paper presents the GENAI4E team's solution to AI City Challenge 2026 Track 4. Our framework builds upon a strong retrieval backbone and progressively integrates heterogeneous vision-language embedding models through score alignment and iterative ensemble fusion, followed by disagreement-aware VLM reranking for ambiguous queries. On the official Pedestrian Anomaly Behavior (PAB) benchmark, our approach achieves 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, demonstrating the effectiveness of combining complementary vision-language representations with selective multimodal reasoning for large-scale text-based person anomaly retrieval.
FSGR: 公平な SID ベースの生成推奨のためのトークン周波数バイアスの軽減
セマンティック ID (SID) ベースの生成推奨は、最近目覚ましい成功を収めています。ただし、既存の方法は、\textbf{トークン周波数バイアス} と呼ばれる、これまで見落とされていた公平性の問題に悩まされています。この問題では、高頻度の SID トークンが体系的に過剰に予測され、低頻度の SID トークンが過小予測されます。このバイアスは、SID 構築中の不均衡なセマンティック コードブックの複合効果と、レコメンデーション トレーニング中の最尤推定目標と人気バイアスの組み合わせから生じ、その結果、アイテム カテゴリ全体で不公平な露出が生じます。既存の SID 手法は主にコードブックの品質を向上させることに焦点を当てており、下流の推奨の公平性に対するトークン頻度の不均衡の影響が見過ごされています。一方、LLM バイアス緩和手法は、SID トークンの階層的なセマンティクスのため、SID ベースの推奨に直接適用すると次善の結果が得られることがよくあります。この問題に対処するために、SID ベースの生成推奨のための公平性最適化フレームワークである \textbf{FSGR} を提案します。 SID の構築中、FSGR は OT ベースの割り当て最適化とデュアル基準再アンカー メカニズムを採用して、よりバランスの取れた SID 表現空間を形成します。レコメンデーション トレーニング中に、2 段階のトレーニング戦略を採用し、レイヤー固有の公平性を微調整するための階層周波数キャリブレーションを導入します。 3 つのバックボーン モデルを備えた 3 つの公開データセットでの実験では、FSGR がトークン頻度のバイアスを軽減し、競争力のある推奨精度を維持しながら、平均ジニの公平性を 20\% 以上向上させることが実証されました。
原文 (English)
FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation
Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. However, existing methods suffer from a previously overlooked fairness issue, which we term \textbf{Token Frequency Bias}, where high-frequency SID tokens are systematically over-predicted while low-frequency SID tokens are under-predicted. This bias originates from the combined effects of imbalanced semantic codebooks during SID construction, and popularity bias together with the maximum likelihood estimation objective during recommendation training, resulting in unfair exposure across item categories. Existing SID methods mainly focus on improving codebook quality and overlook the impact of token frequency imbalance on downstream recommendation fairness, while LLM debiasing methods often yield suboptimal results when directly applied to SID-based recommendation, due to the hierarchical semantics of SID tokens. To address this issue, we propose \textbf{FSGR}, a fairness optimization framework for SID-based generative recommendation. During SID construction, FSGR employs OT-based Assignment Optimization and Dual-Criteria Re-anchor mechanism to form a more balanced SID representation space. During recommendation training, it adopts a two-stage training strategy and introduces Hierarchical Frequency Calibration for layer-specific fairness fine-tuning. Experiments on three public datasets with three backbone models demonstrate that FSGR mitigates token frequency bias and delivers an average Gini fairness improvement of over 20\% while maintaining competitive recommendation accuracy.
AI の言語表現における虚偽と不可能性は方向性が異なる
言語は、誤った状況や、まったくあり得ない状況を説明することができます。 AI モデルがこれらの障害を内部的に区別するかどうかはまだ不明です。私は、17 の哲学ファミリーからの 85 のプロンプトと、それぞれが真実、偶然の虚偽、ありそうもない主張、意味論的異常、および必要な虚偽として表現される 15 のトピックからなるトピックに一致するモダリティ セットを使用した、マルチモーダル オープンウェイト モデル Gemma 3 4B IT の探索的活性化研究を報告します。その回答では、モデルは偶発的な虚偽と矛盾を混同し、15 件の虚偽発言のうち 12 件に「矛盾」とラベルを付けています。その活性化は異なるパターンを示します。線形真実調査は、不可能と真の陳述 (AUC 0.93) を分離しますが、虚偽の陳述 (AUC 0.20) は不可能ではありません。保留されたトピックファミリーで評価された不可能性プローブは、AUC 1.00で必要な虚偽と偶発的な虚偽を分離し、レイヤー15でピークに達し、バランスのとれた精度0.97(ボンフェローニ調整P=0.018)でした。真実と不可能の方向は直交に近いですが、不可能の方向は意味異常の方向と部分的に重なっていますが、それとは区別できます。同じレイヤーにあるスパース オートエンコーダー フィーチャは、このジオメトリを繰り返します。不可能性を選択する機能は、異常な文章に対しても発動しますが、偶発的な虚偽に対してはまれです。このモデルの活性化空間では、必要な虚偽は偶発的な虚偽の極端なケースではなく、実験的に定義された意味論的異常のカテゴリに近いものになります。この表現上の近接性は、不可能な記述が本質的に無意味であることを意味するものではありません。 1 つの小さなモデルから得られたこれらの相関関係の観察は、古い哲学的な区別に対する経験的な脚注を提供します。
原文 (English)
Falsehood and Impossibility Are Different Directions in an AI's Representation of Language
Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT using 85 prompts from 17 philosophical families and a topic-matched modality set of 15 topics, each expressed as a truth, contingent falsehood, improbable claim, semantic anomaly, and necessary falsehood. In its answers, the model conflates contingent falsehood with contradiction, labeling 12 of 15 false statements "contradiction." Its activations show a different pattern. A linear truth probe separates impossible from true statements (AUC 0.93) but not impossible from false statements (AUC 0.20). An impossibility probe evaluated on held-out topic families separates necessary from contingent falsehood at AUC 1.00, peaking at layer 15 with balanced accuracy 0.97 (Bonferroni-adjusted P=0.018). The truth and impossibility directions are close to orthogonal, whereas the impossibility direction partially overlaps a semantic anomaly direction while remaining distinguishable from it. Sparse autoencoder features at the same layer repeat this geometry. Features selective for impossibility also fire on anomalous sentences but rarely on contingent falsehoods. In this model's activation space, necessary falsehoods are not extreme cases of contingent falsehood but lie closer to the experimentally defined category of semantic anomaly. This representational proximity does not imply that impossible statements are intrinsically meaningless. These correlational observations from one small model offer an empirical footnote to an old philosophical distinction.
BrainWAM: 自動運転のためのセマンティック事前確率と予測ダイナミクスのアクション空間調整
自動運転には、意味論的な制約と予測ダイナミクスの両方に基づいた計画が必要です。しかし、既存のエンドツーエンドの運転アプローチは通常、この要件の片側のみを強調しています。つまり、視覚言語アクション (VLA) モデルは意味論的推論に VLM 事前分布を活用し、ワールド アクション モデル (WAM) は生成ワールド モデリングを通じて未来を意識した予測を提供します。これにより、意味論的な事前分布と予測ダイナミクスの両方を活用できる統合プランナーが自然と動機付けられます。しかし、共同トークンレベルの注意による単純な組み合わせは、意味論的なショートカットが共有の注意空間を支配し、予測のダイナミクスを抑制する、注意の割り当ての不一致に悩まされることがわかりました。機能的に特化されたシステム間の調整から複雑な動作が生じるという神経科学の証拠に触発されて、私たちは、意味論的推論と予測世界モデリングを2つの特化されたアクション指向の経路に変換し、コンパクトなアクション表現のレベルでそれらを調整する、構造化されたアクション空間調整フレームワークであるBrainWAMを提案します。さらに、ビデオとアクションのノイズ除去を分離した非同期整流フロー推論戦略を導入します。これにより、計画関連の予測コンテキストを維持しながら推論レイテンシが短縮されます。 BrainWAM は、NAVSIM v1 (89.5 PDMS) と NAVSIM v2 (89.6 EPDMS) の両方で最先端のパフォーマンスに達し、VLA のみまたは WAM のみの方法を常に上回っており、BrainWAM が自動運転システムの実用的で有望な方向性であることを強調しています。
原文 (English)
BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling. This naturally motivates a unified planner that can leverage both semantic priors and predictive dynamics. However, we find that a naive combination through joint token-level attention suffers from an attention-allocation mismatch, where semantic shortcuts dominate the shared attention space and suppress predictive dynamics. Inspired by neuroscience evidence that complex behavior arises from coordination among functionally specialized systems, we propose BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations. We further introduce an asynchronous rectified-flow inference strategy with decoupled video and action denoising, which shortens inference latency while preserving planning-relevant predictive context. BrainWAM reaches state-of-the-art performance on both NAVSIM v1 (89.5 PDMS) and NAVSIM v2 (89.6 EPDMS), consistently outperforming VLA-only or WAM-only methods, highlighting BrainWAM as a practical and promising direction for autonomous driving systems.
確率回路における曲率の構成理論
確率回路 (PC) は、正確な推論をサポートする生成モデルであり、ディープ ニューラル ネットワークとは異なり、損失曲面曲率の正確かつ扱いやすい測定、つまり対数尤度のヘッセ行列のトレースを可能にします。最近の研究では、このトレースをグローバルに正規化して、学習をより平坦でより優れた一般化最適化に偏らせるようにしています。我々は、シャープネスをグローバル正則化子として扱うと、曲率が本質的に構成的なものである PC に対して誤って指定される可能性があることを示します。ヘッセ行列への各合計ノードの寄与が、ノードの使用頻度を測定する回路フローと、その出力分布によって決定される局所シャープネス項に正確に因数分解されることを証明します。この分解により、グローバル シャープネス正則化が深さに偏り、アンダーフィッティングを引き起こす可能性がある理由についての洞察が得られます。これに基づいて、固有の局所曲率に基づいてノードにペナルティを課し、閉じた形式の EM 更新を保持する、適応型シャープネス対応レギュラライザーを導入します。また、このターゲットを絞った正則化は、シャープネスを意識した学習の堅牢性と利点を維持しながら、グローバル正則化が犠牲にする一般化を回復することも経験的に示しています。
原文 (English)
A Compositional Theory of Curvature in Probabilistic Circuits
Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tractable measure of loss-surface curvature: the trace of the Hessian of the log-likelihood. Recent work regularizes this trace globally to bias learning toward flatter, better generalizing optima. We show that treating sharpness as a global regularizer can be misspecified for PCs, whose curvature is inherently compositional. We prove that each sum node's contribution to the Hessian trace factorizes exactly into its circuit flow, which measures how heavily the node is used, and a local sharpness term determined by its output distribution. This decomposition provides insights into why global sharpness regularization is depth biased and can lead to underfitting. Building on it, we introduce an adaptive sharpness aware regularizer that penalizes nodes based on intrinsic local curvature and preserves closed form EM updates. We also show that empirically, this targeted regularization recovers the generalization that global regularization sacrifices while retaining the robustness and benefits of sharpness aware learning.
SPARED: 敵対的に編集されたデータによる推論ベースの AI 生成画像検出
AI が生成した画像の検出は、タスクの半分にすぎません。導入された検出器は、その判定を正当化する必要もありますが、既存の検出器は、トレーニング データから 3 つの失敗モードを引き継ぎます。さまざまなソースから収集された本物の画像と偽の画像は、出所のショートカットを招き、教師付き説明コーパスはテンプレート化された理論的根拠を教え、静的な偽造コーパスは、ジェネレーターが動き続ける間、決定境界を静止したままにします。 \methodname{} は、2 つの異種モデルを相互に対抗させる敵対的強化学習フレームワークです。拡散画像編集者は、本物の写真を、現在の検出器を騙す同じ写真の偽物に編集することを学びます。一方、推論 MLLM は、自由形式の推論に基づいた評決で写真を暴露することを学びます。どちらの報酬も設計上、近道ができないように設計されています。攻撃者は編集が忠実に実行された場合にのみクレジットされ、防御者はその判定が正しい場合にのみクレジットされます。 2 つのモデルが交互に行われると、各ラウンドの攻撃者は現在の検出器の盲点を狙ったより厳しいトレーニング プールを再生成するため、検出器は固定されたアーティファクトの分布を記憶するのではなく、一般化する必要があります。説明は決して報われることはありませんが、正確さのみを重視したトレーニングの副作用として、その質は回を重ねるごとに向上します。このループ内でトレーニングされた検出器は、3 つの外部ベンチマークのそれぞれでラウンド全体にわたって単調に改善します。
原文 (English)
SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data
Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit three failure modes from their training data: real and fake images collected from different sources invite provenance shortcuts, supervised explanation corpora teach templated rationales, and a static forgery corpus leaves the decision boundary standing still while generators keep moving. We introduce \methodname{}, an adversarial reinforcement learning framework that pits two heterogeneous models against each other. A diffusion image editor learns to edit real photographs into fake counterparts of those same photographs that fool the current detector, while a reasoning MLLM learns to expose them with a verdict grounded in free-form reasoning. Both rewards are shortcut-proof by design: the attacker is credited only when its edit is faithfully executed, and the defender only when its verdict is correct. As the two models alternate, each round's attacker regenerates a harder training pool aimed at the current detector's blind spots, so the detector must generalize rather than memorize any fixed artifact distribution. Although the explanation is never rewarded, its quality rises round over round as a side effect of accuracy-only training. A detector trained within this loop improves monotonically across rounds on each of three external benchmarks.
ラベルはエンドポイントではない: 治療の漏洩と MCP エージェントのセキュリティ評価における妥当性の構築
ツールを使用するエージェントのセキュリティ評価では、保存されているラベルと行動の事実が同一視されることがよくあります。 10,200 の実行行を追跡して、180 のモデルバインドされたリクエスト、45 のセマンティック リクエスト、および 15 の観察可能な刺激を追跡することにより、保存されたキャンペーンを監査します。 2 つのスキーマ処理が提供されましたが、計画されていた外部ペイロード ファミリ コーパスは提供されませんでした。履歴グレーダーは直接的な治療漏れを示しました。治療メタデータが ATTACK_SUCCESS クラスをゲートしていたので、修正された動作により治療の再ラベル付けによりクラスが変更される可能性がありました。治療ブラインド再構築では、検証された 3 件の保護データ転送と 1 件の別個の不正転送ケースを保持しながら、58 件の過去の ATTACK_SUCCESS または HIJACK_ATTEMPT ラベルを承認された良性の完了に修正します。ロックされた v2 国勢調査には ATTACK_SUCCESS レコードがまったく含まれていませんが、転送ケースは目的の完了に関するセマンティック境界で HIJACK_ATTEMPT のままです。ロックされた v2 によって構造的に解釈可能であるとみなされた 96 件のリクエストすべてに対する二重レビュアーによるブラインド コンコーダンス レビューでは、同一のレビュアー コンセンサス クラスが生成されましたが、4 つの構成境界ケースについてはロックされたコードブックとは異なりました。私たちは、7 リンクの整合性チェーンと、実行可能なスコープ限定のエンドポイント整合性リンターを提供します。結果はキャンペーンに限定された測定監査であり、人口の攻撃率、モデルのランキング、防御効果、または因果関係の推定ではありません。
原文 (English)
Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation
Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200 execution rows to 180 model-bound requests, 45 semantic requests, and 15 observable stimuli. Two schema treatments were delivered, but the planned external payload-family corpus was not. The historical grader exhibited direct treatment leakage: treatment metadata gated the ATTACK_SUCCESS class, so fixed behavior could change class under treatment relabeling. A treatment-blind reconstruction corrects 58 historical ATTACK_SUCCESS or HIJACK_ATTEMPT labels to authorized benign completions while preserving three verified protected-data transfers and one separate unauthorized-forwarding case. The locked v2 census contains exactly zero ATTACK_SUCCESS records, while the forwarding case remains a HIJACK_ATTEMPT at a semantic boundary concerning objective completion. A dual-reviewer blinded concordance review of all 96 requests deemed structurally interpretable by locked v2 produced identical reviewer-consensus classes but differed from the locked codebook on four construct-boundary cases. We contribute a seven-link Integrity Chain and an executable, scope-bounded endpoint-integrity linter. The result is a campaign-bounded measurement audit, not a population attack-rate, model-ranking, defense-efficacy, or causal estimate.
NaviDC-OCR: デジタル文書およびカメラでキャプチャされた文書にわたる文書解析のナビゲート
文書解析の目的は、非構造化文書を構造化された機械可読表現に変換することです。視覚言語モデル (VLM) の最近の進歩により、文書解析が大幅に進歩しました。しかし、既存のアプローチは依然として 2 つの大きな課題に直面しています。まず、分離された VLM ベースの手法は正確なレイアウト分析に大きく依存しており、カメラで撮影したドキュメントの幾何学的歪みが連鎖的なエラーを引き起こす可能性があります。第 2 に、エンドツーエンドの VLM ベースの方法は明示的なレイアウト検出への依存を軽減しますが、高解像度のシナリオでは冗長な生成、幻覚、不十分な構造的推論の問題が発生することがよくあります。これらの課題に対処するために、私たちは文書解析のための統一フレームワークである NaviDC-OCR を提案します。 NaviDC-OCR は、幾何学的認識を VLM に組み込むための変形認識学習を導入し、複雑なレイアウト表現のための適応サンプリング メカニズムを提案します。さらに、内容と構造を分離した学習戦略が開発され、数式文法とテーブル構造を明示的にモデル化し、より効果的な構造化表現学習を可能にします。広範な実験により、NaviDC-OCR がさまざまなドキュメント解析ベンチマークにわたって最先端のパフォーマンスを達成することが実証されました。 OmniDocBench v1.6、Wild-OmniDocBench、PureDocBench でそれぞれ 96.87、88.53、78.41 の総合スコアを獲得し、ICDAR 2026 Sci-ImageMiner Challenge で 1 位にランクされています。これらの結果は、複雑な文書解析シナリオにおける NaviDC-OCR の有効性と一般化機能を検証します。
原文 (English)
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.
EGRL: RNA-タンパク質相互作用予測のためのエッジ生成ガイド付き関係認識学習
RNA-タンパク質相互作用 (RPI) は、細胞機能の調節に重要です。 RPI 検出のための従来のウェットラボ実験はコストと時間がかかりますが、ディープラーニング (DL) 手法は、RPI 予測 (RPIP) の効率的な計算代替手段を提供します。特に、グラフ ニューラル ネットワーク (GNN) は RPI ネットワークを自然にモデル化するため、有望です。ただし、既存の GNN ベースの手法は、均質なグラフや事前定義されたメタパスに依存することが多く、データの疎性を処理したり、未知の分子を含むコールド スタート シナリオに一般化したりする能力が制限されています。これらの制限に対処するために、私たちは、いくつかの重要なコンポーネントを備えた新しいフレームワークであるエッジ生成ガイド付き関係認識学習 (EGRL) を提案します。インタラクションパターンを適応的に融合するための、複数の関係を認識した注意メカニズム。コールド スタート ノードをサポートするための潜在的な (「ソフト」) エッジを予測するグラフ ジェネレーター。最終的なインタラクションスコアリングのための多機能融合予測子。 EGRL は、主タスク損失と補助発電機損失を使用して共同トレーニングされます。 4 つのベンチマーク データセットの包括的な評価により、EGRL が競争力のある全体的なパフォーマンスを達成していることが実証されています。さらに重要なことは、コールドスタート設定で優れた汎用性を示し、未知の分子に対して受信者動作特性曲線下面積 (AUROC) 0.867 および精度再現率曲線下面積 (AUPR) 0.861 を達成しており、これは従来の最先端の方法と比較して、AUROC で 8.6%、AUPR で 5.0% の改善に相当します。コードは近日公開予定です。
原文 (English)
EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction
RNA-Protein Interactions (RPIs) are critical for regulating cellular functions. While traditional wet-lab experiments for RPI detection are costly and time-consuming, Deep Learning (DL) methods provide an efficient computational alternative for RPI Prediction (RPIP). In particular, Graph Neural Networks (GNNs) are promising, as they naturally model RPI networks. However, existing GNN-based methods often rely on homogeneous graphs or predefined meta-paths, which limit their ability to handle data sparsity and to generalize to cold-start scenarios involving unknown molecules. To address these limitations, we propose Edge Generation-guided Relation-aware Learning (EGRL), a novel framework with several key components: implicit meta-path learning to capture relational semantics without handcrafted paths; a multi-relation-aware attention mechanism for adaptive fusion of interaction patterns; a graph generator that predicts potential ("soft") edges to support cold-start nodes; and a multi-feature fusion predictor for final interaction scoring. EGRL is jointly trained with a primary task loss and an auxiliary generator loss. Comprehensive evaluations on four benchmark datasets demonstrate that EGRL achieves competitive overall performance. More importantly, it exhibits superior generalization in cold-start settings, achieving an Area Under the Receiver Operating Characteristic curve (AUROC) of 0.867 and an Area Under the Precision-Recall curve (AUPR) of 0.861 on unknown molecules, corresponding to improvements of 8.6% in AUROC and 5.0% in AUPR over prior state-of-the-art methods. The code will be released soon.
InFactPlanner: 持続可能な地理分散型 LLM データセンターの計画
LLM 推論の急速な成長により、持続可能性への懸念は 1 回限りのトレーニングから継続的なサービスへと移行しており、インフラストラクチャの決定がエネルギー使用、炭素排出量、水消費量、サービス品質を決定します。しかし、事業者は多くの場合、大規模なインフラストラクチャを構築する前に導入の代替案を比較する必要があるため、直接測定はコストがかかり、時間がかかり、場合によっては実行不可能になります。単一サイトおよび地理的に分散されたサイトにわたる LLM 推論のための持続可能な AI データセンター展開の what-if 分析のためのトレース駆動型意思決定支援フレームワークである InFactPlanner を紹介します。 InFactPlanner は、クエリ トレース、ハードウェア モデル プロファイル、候補地の構成、PUE/WUE パラメータ、再生可能発電モデル、時間変化するグリッド炭素強度を組み合わせて、電力、エネルギー、炭素排出量、水の使用量、待ち時間、およびサーバーの使用率を推定します。このフレームワークは、低レベルのサービス効果を構成可能なハードウェア モデル プロファイルに抽象化し、サイトの選択、容量の配置、ハードウェア、モデル、再生可能エネルギーの統合、およびルーティングの選択を迅速に比較できるようにします。参照 LLM 推論エネルギー推定値を偏差 10% 未満で再現することでエネルギー会計パイプラインを検証し、複数のデータセンターとサーバー数にわたるスケーラビリティを評価し、ハードウェアの選択、再生可能エネルギーの配置、地理的展開、炭素を意識したルーティングに関するシナリオ主導の意思決定分析を実証します。私たちの結果は、持続可能性を考慮した最適な選択は遅延を最適化した選択とは異なる可能性があり、導入の炭素価値はローカルグリッドの組み合わせに大きく依存することを示しています。
原文 (English)
InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers
The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to compare deployment alternatives before large-scale infrastructure is built, making direct measurement costly, slow, and sometimes infeasible. We present InFactPlanner, a trace-driven decision-support framework for what-if analysis of sustainable AI data center deployment for LLM inference across single and geo-distributed sites. InFactPlanner combines query traces, hardware-model profiles, candidate site configurations, PUE/WUE parameters, renewable generation models, and time-varying grid carbon intensity to estimate power, energy, carbon emissions, water use, latency, and server utilization. The framework abstracts low-level serving effects into configurable hardware-model profiles, enabling rapid comparison of site selection, capacity placement, hardware, model, renewable integration, and routing choices. We validate the energy accounting pipeline by reproducing reference LLM inference energy estimates with less than 10% deviation, evaluate scalability across multiple data centers and server counts, and demonstrate scenario-driven decision analyses for hardware selection, renewable placement, geographic deployment, and carbon-aware routing. Our results show that sustainability-optimal choices can differ from latency-optimal ones, and that the carbon value of deployment depends strongly on the local grid mix.
因果推論による LLM ベースのマルチエージェント システムの効率的で説明可能な通信トポロジの発見
大規模言語モデル (LLM) ベースのマルチエージェント システム (MAS) のパフォーマンスは、効果的な通信トポロジに大きく依存します。しかし、既存のトポロジ生成方法は通常、タスクレベルの報酬のみによって駆動されるブラックボックス最適化を通じて通信トポロジを学習します。このような最適化は効果的ではありますが、特定の通信エッジが選択される理由についてはほとんど洞察が得られず、コラボレーションの成功に関与する重要な通信サブグラフを特定することが困難になります。この制限に対処するために、任意のトポロジー ジェネレーターによって生成された通信トポロジーの解釈可能な説明を提供するモデルに依存しないフレームワークである E2-Explainer を提案します。具体的には、タスク保存のエッジレベルの証拠によってサポートされるコンパクトな通信サブグラフを特定する因果関係の問題としてトポロジーの説明を定式化します。この証拠は、各通信チャネルをマスキングすることでタスクの結果と最終応答の安定性がどのように変化するかを測定するグレンジャー スタイルの目標を使用して得られます。結果として得られる予算付きサブグラフは償却されたExplainerに抽出され、展開時にエッジレベルの評価を繰り返すことなく効率的な事後説明が可能になります。複数の推論とコーディングのベンチマークに関する広範な実験により、E2-Explainer がコラボレーションの成功を維持する重要な通信サブグラフを特定することが実証されました。これらのサブグラフを直接実行して冗長な通信エッジを取り除き、競争力のあるタスクのパフォーマンスを維持しながら通信コストを大幅に削減することもできます。
原文 (English)
Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference
The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level rewards. While effective, such optimization provides little insight into why particular communication edges are selected, making it difficult to identify the critical communication subgraphs responsible for successful collaboration. To address this limitation, we propose E2-Explainer, a model-agnostic framework for providing interpretable explanations of communication topologies produced by arbitrary topology generators. Specifically, we formulate topology explanation as a causal attribution problem that identifies compact communication subgraphs supported by edge-level evidence of task preservation. We obtain this evidence with a Granger-style objective that measures how masking each communication channel changes the task outcome and the stability of the final response. The resulting budgeted subgraphs are then distilled into an amortized explainer, enabling efficient post-hoc explanation without repeated edge-level evaluations at deployment. Extensive experiments on multiple reasoning and coding benchmarks demonstrate that E2-Explainer identifies critical communication subgraphs that preserve successful collaboration. These subgraphs can also be executed directly to prune redundant communication edges, substantially reducing communication costs while maintaining competitive task performance.
H-VAEP と H-xT: 確率の推定によるハンドボールの攻撃的なオンザボールアクションの評価
プロのハンドボールにおける従来の選手評価は、基本的なボックススコア指標やヒューリスティック指標に依存しており、複数選手によるビルドアップの連鎖を評価することができません。フットボール (サッカー) 分析では、予想される脅威 (xT) と確率推定によるアクションの評価 (VAEP) が採用されていますが、これらのイベントベースのアクション評価フレームワークはまだハンドボールには適応されていません。この論文では、ハンドボール ブンデスリーガの 5 シーズンの追跡由来のイベント データを利用して、ハンドボールに対する xT と VAEP の最初の包括的な適応と評価を紹介します。私たちは、ハンドボール固有のコート ゾーニング レイアウトを使用して Handball-xT (H-xT) を開発し、標準的な長方形のグリッドよりも体系的に堅牢であることをシミュレーションによって実証しています。特徴空間を調整し、コンテキストの長さを選択してチーム ID の漏洩を制限することで、ハンドボール VAEP (H-VAEP) を最適化します。私たちの評価では、H-VAEP が、ビルドアップ プレーを際立たせる、非常に安定しており、識別力があり、直感的なプレーヤー評価をもたらすことが示されています。最後に、プロのクラブがこれらのモデルを導入できるように、完全なコード リポジトリをリリースします。
原文 (English)
H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities
Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP), these event-based action valuation frameworks have not yet been adapted to handball. In this paper, we present the first comprehensive adaptation and evaluation of xT and VAEP for handball, utilizing five seasons of tracking-derived event data from the Handball Bundesliga. We develop Handball-xT (H-xT) using a handball-native court zoning layout, demonstrating via simulations that it is systematically more robust than standard rectangular grids. We optimize Handball-VAEP (H-VAEP) by tailoring its feature space and selecting the context length to limit team-identity leakage. Our evaluation shows that H-VAEP yields exceptionally stable, discriminative, and intuitive player ratings that highlight build-up play. Finally, we release our complete code repository to help professional clubs deploy these models.
AutoQuREO: 量子リソースの自動推定と最適化のためのフレームワーク
量子コンピューティングが原理証明の実証から実用化に向けて進むにつれて、異種のハードウェアおよびソフトウェア スタックにわたるシステム レベルの最適化によってアルゴリズムの実現可能性を高める必要性が大きな障害となります。量子リソース推定 (QRE) はこの移行において中心的な役割を果たしますが、既存のアプローチは主にコンパイル負荷の高い、またはドメイン知識に基づくシンボリック アノテーションのままであり、長期的なフォールト トレラントの仮定と密接に結合しているため、局所的な適用性は制限されています。この作業では、フルスタックの量子リソース推定と最適化のための自動フレームワークである AutoQuREO を紹介します。 AutoQuREO は、次の 4 つの核となる新規性を中心に構築されています。(i) 量子コンピューティング スタックの柔軟なユーザー定義の抽象化。 (ii) 再利用可能なスタック コンポーネントのモジュール ライブラリにより、迅速なフルスタック プロトタイピングが可能になります。 (iii) アルゴリズムプロファイリングと神経記号学習による層ごとのリソースの代理モデリング。 (iv) QRE を展開パイプラインに直接組み込む、統合された多目的最適化。これらの設計上の選択により、AutoQuREO は量子コンピューティング スタックのデジタル ツインとして機能し、複雑な設計空間の扱いやすい探索をサポートできます。初期フォールトトレラント量子アルゴリズム、小さな誤り訂正符号、ゲート分解、パラメトリック量子回路の変分トレーニングなど、代表的な共同設計ケーススタディを通じて AutoQuREO の機能を実証します。これらの例は、既存の QRE ツールを使用すると計算的に扱いにくい、または難解な未利用のリソースのトレードオフを AutoQuREO で体系的に発見できる方法を示しています。 AutoQuREO は、量子テクノロジーの準備を進めるための汎用プラットフォームとして位置付けられています。
原文 (English)
AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization
As quantum computing progresses from proof-of-principle demonstrations toward practical utility, a significant impediment is the need to augment algorithmic feasibility with system-level optimization across heterogeneous hardware and software stacks. Quantum resource estimation (QRE) plays a central role in this transition, yet existing approaches remain largely compilation-heavy or domain-knowledge-guided symbolic annotations, and tightly coupled to long-term fault-tolerant assumptions, limiting their topical applicability. In this work, we introduce AutoQuREO, an Automated framework for full-stack Quantum Resource Estimation and Optimization. AutoQuREO is built around four core novelties: (i) a flexible, user-defined abstraction of the quantum computing stack; (ii) a modular library of reusable stack components enabling rapid full-stack prototyping; (iii) surrogate modeling of layer-wise resources via algorithmic profiling and neuro-symbolic learning; and (iv) integrated multi-objective optimization that embeds QRE directly into deployment pipelines. Together, these design choices enable AutoQuREO to serve as a digital twin for quantum computing stacks, supporting the tractable exploration of complex design spaces. We demonstrate the capabilities of AutoQuREO through representative co-design case studies, including early-fault-tolerant quantum algorithms, small error correction codes, gate decomposition and variational training of parametric quantum circuits. These examples illustrate how AutoQuREO enables systematic discovery of unexploited resource trade-offs that are computationally intractable or abstruse using existing QRE tools. AutoQuREO is positioned as a general-purpose platform for advancing quantum technology readiness.
目的がボトルネック: 潜在世界モデルは、計画者が使用できないものをコード化する
潜在世界モデルは、どれだけうまく予測できるかによって判断されるため、長期的な視野で計画が失敗すると、予測子が低下するのが自然な解釈です。 TwoRoom 上の LeWorldModel の再現では、代わりにバインディング制約がプランナーの目的であることを示します。予測子は限界ではありません。環境ステップ 75 ステップ先の想像された状態でも、世界が凍結したと仮定するのと同じ 0.189 しか間違っていませんが、計画者は 25 ステップを超えて想像することはありません。目的は。クロスエントロピー法の計画では、潜在距離の二乗が最小化され、r = 0.426 で真の距離を追跡し、約 80 アリーナ ユニットで飽和し、120 を超えると減少するため、目標から遠ざかることでコストを下げることができます。情報は全体に存在します。リッジ プローブは、R^2 0.9922 で凍結した埋め込みから位置を回復します。病状はメソッドの問題であり、1 つの再実装の問題ではありません。これは、著者がリリースした重みに存在し、4 つのチェックポイントにわたる長期的な成功のランク付けが、メトリクスの品質に正確に一致し、予測精度に反比例します。何も再トレーニングせず、GPU も使用せずに目標だけを置き換えると、オフセット 100 で達成される目標が 26.0% から 98.0% に上昇し、オフセット 25 での 98.0% に等しく、予算の 3 分の 1 未満で 92.0% に達します。計画はホライズンに応じて停止します。最適なコストが最も正確であるわけではありません。フレーム分離のみから学習したヘッドは、位置プローブよりも悪い空間距離 (r = 0.819 対 0.9897) を予測しますが、より良い計画を立て、環境の隔壁を越えるために 24% 多くチャージしますが、二乗潜在距離のチャージは 4% 少なくなります。近接性ではなく、到達可能性を学習しました。
原文 (English)
The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use
Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades. On a reproduction of LeWorldModel on TwoRoom we show the binding constraint is the planner's objective instead. The predictor is not the limit: its imagined state seventy-five environment steps ahead is still only 0.189 as wrong as assuming the world froze, while the planner never imagines beyond twenty-five. The objective is. Cross-entropy-method planning minimises squared latent distance, which tracks true distance at r = 0.426, saturates by about eighty arena units and decreases beyond a hundred and twenty, so moving away from the goal can lower the cost. The information is present throughout: a ridge probe recovers position from the frozen embedding at R^2 0.9922. The pathology is the method's, not one reimplementation's. It is present in the authors' released weights, and across four checkpoints long-horizon success rank-orders exactly with metric quality and inversely with prediction accuracy. Replacing only the objective, with nothing retrained and no GPU, lifts goals reached at offset 100 from 26.0% to 98.0%, equals the 98.0% at offset 25, and reaches 92.0% under a third of the budget: planning stops depending on the horizon. The best cost is not the most accurate. A head learned from frame separation alone predicts spatial distance worse than a position probe (r = 0.819 against 0.9897) yet plans better, charging 24% more to cross the environment's dividing wall where squared latent distance charges 4% less. It has learned reachability, not proximity.
手作りのセキュリティを超えて: LLM エージェントの自己進化型防御に向けて
大規模言語モデル (LLM) エージェントの運用能力の拡大により、高度なセキュリティ脅威が導入されます。ランタイム防御は、セキュリティ メカニズムをエージェントの実行ループに統合することで、これらのリスクを軽減する効果的なアプローチとして登場しました。しかし、既存の実行時防御は手動で設計された介入に大きく依存しており、その構築と保守のための原則に基づいたフレームワークが欠けています。この研究では、まず、ランタイム防御のハーネスレベルの定式化を開発します。これは、ハーネスメカニズムがどのように防御構築を可能にするかを体系的に特徴付け、ハーネスの観点から既存のランタイム防御介入の統一されたビューを提供します。この定式化に基づいて、私たちは、適切な介入戦略を自動的に特定し、観察された障害の痕跡に基づいて防御アーティファクトを反復的に改善する自己進化型ランタイム防御フレームワークである HARD (Harness-based Autonomous Runtime Defense Evolution) を提案します。 HARD は、ランタイム防御開発を手動エンジニアリングから自律的な進化プロセスに変換します。広範な実験により、無害なタスクのユーティリティを維持しながら、既存の手作りの防御よりもセキュリティ パフォーマンスが向上することが実証されています。私たちの調査結果は、配備された LLM エージェントを保護するための有望な新しいパラダイムとして自律防御の進化を浮き彫りにし、エージェントが防御の弱点を特定し、保護メカニズムを継続的に改善できるようにします。
原文 (English)
Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents
The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective approach to mitigating these risks by integrating security mechanisms into the agent execution loop. However, existing runtime defenses rely heavily on manually designed interventions and lack a principled framework for their construction and maintenance. In this work, we first develop a harness-level formulation of runtime defense that systematically characterizes how harness mechanisms enable defense construction and provides a unified view of existing runtime defense interventions from a harness perspective. Building on this formulation, we propose HARD (Harness-based Autonomous Runtime Defense Evolution), a self-evolving runtime defense framework that automatically identifies appropriate intervention strategies and iteratively improves defense artifacts based on observed failure traces. HARD transforms runtime defense development from manual engineering into an autonomous evolution process, and extensive experiments demonstrate that it improves security performance over existing handcrafted defenses while preserving benign task utility. Our findings highlight autonomous defense evolution as a promising new paradigm for securing deployed LLM agents, enabling agents to identify defense weaknesses and continuously improve their protection mechanisms.
二重役割識別子を使用した生成型ユニバーサル マルチモーダル検索
生成情報検索 (GIR) は、関連項目の識別子を直接生成するようにジェネレーターをトレーニングすることにより、従来のインデックス検索、その後ランク付けの検索パイプラインに代わる魅力的な代替手段として登場しました。その期待にもかかわらず、多くの未解決の課題がまだ残されています。まず、制約付きの左から右へのデコードは、プレフィックスレベルのエラーと局所最適化に対して脆弱です。第 2 に、これまでの GIR 研究のほとんどは単峰性のままであり、テキスト、画像、および画像とテキストの混合項目にわたる命令を意識した検索は十分に検討されていません。第三に、離散識別子ベースの GIR はより高い効率を提供しますが、その検索精度は依然として最先端の密ベクトルベースの検索方法よりも遅れています。これらの課題を動機として、私たちは、複数のモダリティとドメインにわたる多様な検索タスクをサポートするデュアルロール識別子を特徴とするユニバーサルマルチモーダル検索のための新しい生成フレームワークである DrIG を提案します。各候補には、2 つの相補的な役割を果たす単一の残差量子化識別子が割り当てられます。逐次的な役割では、識別子は自己回帰的にデコードされ、最初のトークンがモダリティを明示的にモデル化し、残りのトークンが徐々により詳細なセマンティクスを捕捉します。セットベースの役割では、同じトークンが順序付けされていないセットとして再解釈され、プレフィックスに依存しない関連性を事前に提供します。これにより、制約付きビーム検索がガイドされ、局所最適エラーが軽減されます。 M-BEIR ベンチマークとテキストから画像への評価データセットに関する広範な実験により、(1)DrIG はさまざまなタスクにわたって最先端の生成マルチモーダル ベースラインよりも常に優れたパフォーマンスを示し、一方、ハイブリッド リランキングは強力な高密度レトリーバーに対して有利な効率と効果のトレードオフを実現します。 (2) アブレーションおよびスケーリング解析により、ベース LMM、ビーム サイズ、再ランキング深さ、および融合戦略がどのように回収パフォーマンスに影響を与えるかを明らかにし、システム設計の実践的なガイダンスを提供します。
原文 (English)
Generative Universal Multimodal Retrieval with Dual-role Identifiers
Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipeline by training a generator to produce the identifiers of relevant items directly. Despite its promise, a number of open challenges still remain. First, constrained left-to-right decoding is vulnerable to prefix-level errors and local optima. Second, most prior GIR research remains largely unimodal, leaving instruction-aware retrieval across text, image, and mixed image-text items underexplored. Third, although discrete identifier-based GIR offers higher efficiency, its retrieval accuracy still lags behind that of the cutting-edge dense-vector-based retrieval methods. Motivated by these challenges, we propose DrIG, a novel Generative framework for universal multimodal retrieval featuring Dual-role Identifiers, which supports diverse retrieval tasks across multiple modalities and domains. Each candidate is assigned a single residual-quantized identifier that serves two complementary roles. In its sequential role, the identifier is decoded autoregressively, where the first token explicitly models modality and the remaining tokens capture progressively finer semantics. In its set-based role, the same tokens are reinterpreted as an unordered set to provide a prefix-independent relevance prior, which guides constrained beam search and alleviates local-optimum errors. Extensive experiments on the M-BEIR benchmark and the text-to-image evaluation datasets show that:(1)DrIG consistently outperforms state-of-the-art generative multimodal baselines across diverse tasks, while hybrid reranking achieves a favorable efficiency-effectiveness trade-off against strong dense retrievers. (2)Ablation and scaling analyses reveal how the base LMM, beam size, reranking depth, and fusion strategy affect retrieval performance, providing practical guidance for system design.
静的解析に基づくエージェント的 AI 翻訳により、Rust をフルスタックのバイオインフォマティクス言語として実現
バイオインフォマティクスの分野は、レガシー コード、つまり一般的に使用されているが、もはやメンテナーがいない古いコード、または今では馴染みのない言語 (Perl、Fortran など) で書かれている可能性のある古いコードと格闘しています。これには保守コスト (技術的負債) が発生しますが、動的に型付けされた言語は環境に悪影響を及ぼし、最新のハードウェアを活用できなくなります。レガシーコードには、臨床現場での使用に適さないセキュリティまたは安全性の問題がある可能性もあります。ここでは、静的分析と組み合わせたエージェント AI を使用して、レガシー コードを最新の言語 Rust に変換できることを示します。当社は体系的な翻訳を支援するプロンプトとサポート ソフトウェアを提供し、NGS とイメージングの共通ソフトウェアで評価します。ソフトウェア Basset で結果を紹介します。サイズは最大 80 分の 1 に削減され、ビルド時間は最大 10 分の 1 に短縮され、主要なステップのパフォーマンスは 3 倍以上向上しました。 Unix への依存関係も削除され、Basset はコンテナーなしでネイティブ Windows 上で実行できる唯一の単一セル パイプラインになりました。したがって、限られた予算でバイオインフォマティクス ソフトウェアの大規模なリファクタリングが可能になり、より複雑なツールの開発が可能になります。
原文 (English)
Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language
The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran). This incurs maintenance cost (technical debt), but dynamically typed languages also negatively impacts the environment and fail to make use of modern hardware. Legacy code may also have security or safety problems that make it unsuited for use in clinical settings. Here we show that agentic AI, combined with static analysis, can be used to translate legacy code to the modern language Rust. We provide prompts and supporting software to aid systematic translation, and evaluate it on common software for NGS and imaging. We showcase the result on our software Bascet: Size was reduced by ~80x, build time decreased by ~10x, and performance of key steps improved >3x. Unix dependencies were also removed, making Bascet the only single-cell pipeline able to run on native Windows, without a container. Large-scale refactoring of bioinformatics software is thus now possible at a limited budget, enabling more complex tools to be developed.
UniTraffic-Agent: 2 つのドメイン外評価による AI City Challenge 2026 Track 3 の統合トラフィック ビデオ推論
道路ビデオは事故、違反、車両と交通弱者の道路利用者とのやりとりの直接的な証拠を提供するため、交通ビデオの理解はインテリジェント交通機関における重要な問題となっています。有用なシステムでは、交通イベントがどのように発生するか、なぜそれが起こるのか、関連するインタラクションがいつ発生するのかを説明する必要がありますが、交通ビデオにはまばらなイベントとさまざまな視点が含まれるため、マルチモーダル大規模言語モデル (MLLM) ではこれが依然として困難です。第 10 回 AI シティ チャレンジの Track~3 の MR-CAS ソリューションである UniTraffic-Agent を紹介します。これには、交通異常推論 (TAR) と 2 つのドメイン外評価 (魚眼交通イベントの FETV と歩行者の意図推論の PSI-VQA) が含まれています。 UniTraffic-Agent は、タイムスタンプ付きの視覚的証拠をサンプリングし、1 つのリクエスト内の同じクリップからのすべての質問について推論し、タスク固有のアクション アダプターを通じて応答を変換する、観察 - 理由 - 行為 - 検証のワークフローに従います。公式のパブリック リーダーボードでは、MR-CAS は TAR でスコア 0.5780 で 16 位、FETV で 0.4884 で 2 位、PSI-VQA で 64.4161 で 4 位にランクされています。コードは https://github.com/Roclp/UniTraffic-Agent で入手できます。
原文 (English)
UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations
Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse events and varied viewpoints. We introduce UniTraffic-Agent, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. UniTraffic-Agent follows an observe--reason--act--verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters. On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161. The code is available at https://github.com/Roclp/UniTraffic-Agent.
GraphRAG によるサイバー脅威インテリジェンスの運用化
セキュリティ研究者がサイバー攻撃に関するレポートを発表すると、検出エンジニアはそれを実用的な検出ルールに変換することになっています。実際には、これを自動化しようとするほとんどの場合、レポートから最も単純な手がかり (不正な IP アドレス、ドメイン名、ファイル ハッシュ) が抽出され、ブロック リストに変換されるだけです。これは弱い戦略です。攻撃者はこれらの単純な手がかりを数時間または数日以内に変更する可能性があるため、結果として得られる検出は展開されるとすぐに機能しなくなります。セキュリティ チームは、このアイデアを Pyramid of Pain で説明しています。このプロジェクトでは、標準的なベクトル類似性検索システム (Naive RAG) ではなく、ナレッジ グラフ検索システムである Microsoft GraphRAG にレポートをフィードすると、これらの耐久性のあるピラミッドの頂点の手がかりにさらに依存する検出計画が生成されるかどうかを検討します。最終計画を作成するために、両方のシステムに同じレポート、同じ生成命令、および同じ言語モデルが与えられます。取得ステップのみが異なります。ある APT28 レポートの詳細なケース スタディでは、レポート内のすべての IP アドレス、ドメイン、ファイル ハッシュがローテーションされた後も、GraphRAG プランは検出の 100\% で起動し続けましたが、Naive RAG プランはわずか 29\% で起動し続けました。 4 つのベンダーによる 9 つの実際の CTI レポートで比較を繰り返すと、同じパターンが確認されます。GraphRAG の計画は、2 つのシステムの合計スコアが最終的に接近した場合でも、ピラミッドのより高く、回避が困難なレベルに一貫して到達します。この結果は、ナレッジ グラフを意識した検索を、SOC 展開可能なハンティング プランを自動的に生成するためのアーキテクチャ的に正しい基盤として扱うことを裏付けると同時に、生成プロンプトの文言が検索バックエンド自体とほぼ同じくらい重要であることを示しています。
原文 (English)
Operationalizing Cyber Threat Intelligence with GraphRAG
When a security researcher publishes a report on a cyberattack, detection engineers are supposed to turn it into working detection rules. In practice, most automated attempts at this only extract the simplest clues from the report --- bad IP addresses, domain names, and file hashes --- and turn them into block lists. This is a weak strategy, because attackers can change these simple clues within hours or days, so the resulting detections stop working almost as soon as they are deployed. Security teams describe this idea with the Pyramid of Pain. This project asks whether feeding a report into a knowledge-graph retrieval system, Microsoft GraphRAG, rather than a standard vector-similarity retrieval system (Naive RAG), produces detection plans that rely more on these durable, top-of-pyramid clues. Both systems are given the same report, the same generation instructions, and the same language model to write the final plan; only the retrieval step differs. In a detailed case study of one APT28 report, the GraphRAG plan kept firing at 100\% of its detections after every IP address, domain, and file hash in the report was rotated, while the Naive RAG plan kept firing at only 29\%. Repeating the comparison across nine real CTI reports from four vendors confirms the same pattern: GraphRAG plans consistently reach higher, harder-to-evade levels of the pyramid, even when the two systems end up close on total score. The results support treating knowledge-graph-aware retrieval as the architecturally correct foundation for automatically generating SOC-deployable hunting plans, while showing that the wording of the generation prompt matters almost as much as the retrieval back-end itself.
TEMPO: Makespan 対応のエキスパート - メモリとコンピューティングに依存した体制にわたる並列負荷分散
エキスパートパラレル (EP) MoE サービスでは、すべてのレイヤーが最も遅い GPU で同期します。ディスパッチャは、エキスパート時間が線形であると仮定して、トークン カウント (EPLB、LPLB、UltraEP) またはアクティブ化されたエキスパート カウント (METRO) のバランスをとります。データセンターの 2 世代の GPU での測定では、どちらでもないことがわかりました。$\nstar\!\およそ\!156$ -- $168$ トークン未満では、HBM 重みストリーミングが支配的です -- コストはトークンではなく \emph{アクティブ化されたレプリカ} に発生します。その上で、グループ化された GEMM はトークンを 128 タイルの $M$ タイルに丸めるため、専門家が \emph{分割} してパディングされた計算を追加します。最大アフィン プロファイル $t=\max(a+bG,\,c+\beta N)$ は両方の領域を捉えます。現実的なデコードバッチには、線形領域ではホットな専門家が、フラット領域では冷たい専門家が \emph{同時に}保持されます。記録されたバッチは、プロキシのディスパッチがモデル化されたブロック時間で $1.4$--$1.6\times$ 異なり (p95 から $1.7\times$)、\emph{どの} プロキシが勝つかはレジームによって反転することを示しています。私たちは、バッチごとのディスパッチを固定料金のメイクスパン問題として形式化し、2 つの完全に複製された GPU 上で NP ハード、縮退限界の多項式を実現します。そして、クリティカル パスからミリ秒でそれを解決するメイクスパン対応ディスパッチャである \sys{} を提示します。 SGLang 統合はプロセス外で実行され、ディスパッチとカウント収集を 1 つのグラフ内カーネルに融合します。 8 GPU のテストベッド (マイクロベンチマーク) によって支えられている \sys{} は、どこでも最良の固定ベースラインの 1\% 以内に留まり、体制が混在する場合には最大 $15.5\%$ の差で勝利します。 Testbed~B のエンドツーエンドでは、Qwen3-235B (勝利領域内) のスループットが $4$--$6\%$ 向上し、p99 レイテンシーが ${\sim}15.6\%$ 削減されました。 DeepSeek-V3 (外部、通信主体) はメカニズムのコストのみを示します。フェーズ図は普遍的な勝利をもたらすものではなく、展開前に両方の結果を予測するという主張です。
原文 (English)
TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU generations show it is neither: below $\nstar\!\approx\!156$--$168$ tokens, HBM weight streaming dominates---cost attaches to \emph{activated replicas}, not tokens; above it, grouped GEMM rounds tokens to 128-tile $M$-tiles, so \emph{splitting} an expert adds padded compute. A max-affine profile $t=\max(a+bG,\,c+\beta N)$ captures both regimes. Realistic decode batches hold hot experts in the linear regime and cold in the flat \emph{simultaneously}; recorded batches show proxy dispatches differ by $1.4$--$1.6\times$ in modeled block time (p95 up to $1.7\times$), and \emph{which} proxy wins flips with the regime. We formalize per-batch dispatch as a fixed-charge makespan problem---NP-hard on two fully replicated GPUs, polynomial in degenerate limits---and present \sys{}, a makespan-aware dispatcher solving it in milliseconds off the critical path; its SGLang integration runs out-of-process and fuses dispatch with count collection into one in-graph kernel. Anchored by an 8-GPU Testbed~A microbenchmark, \sys{} stays within 1\% of the best fixed baseline everywhere and wins by up to $15.5\%$ where regimes mix. End-to-end on Testbed~B, Qwen3-235B (inside the win region) gains $4$--$6\%$ throughput and cuts p99 latency by ${\sim}15.6\%$; DeepSeek-V3 (outside, communication-dominated) shows only mechanism cost. A phase diagram, not a universal win, is the claim: it predicts both outcomes before deployment.
LOB-ID: 開始距離による合成市場データの評価
指値注文帳 (LOB) データの生成モデルは急速に進歩していますが、その評価は多くの場合、定型化された事実と選択された市場統計に焦点を当てています。これらの測定は有用な診断を提供しますが、オーダーブックの軌跡の同時時間構造およびクロスレベル構造を捕捉できない可能性があります。 LOB-ID は、Fr\'echet Inception Distance (FID) と Monge Inception Distance (MIND) を LOB データに適応させる埋め込みベースのフレームワークです。ドメイン固有のエンベディングを取得するために、5 つの株式の 4 か月分のレベル 2 オーダーブック データで DeepLOB アーキテクチャをトレーニングします。 LOB-ID は、時間、機器、埋め込みチェックポイント全体にわたって安定しており、制御された歪みの下では単調に増加することを示します。次に、統計ベースの評価を回避する FID およびディープブック摂動に対するモーメント マッチング攻撃を構築します。 MIND は両方の歪みに対して大幅に敏感なままです。最後に、確率的ベースラインと深層学習アプローチにまたがる 5 つの生成 LOB モデルをスコアリングし、LOB-ID が、それぞれが構築によって捕捉する結合時間構造およびクロスレベル構造に沿ってそれらをランク付けすることを発見しました。
原文 (English)
LOB-ID: Evaluating Synthetic Market Data by Inception Distances
Generative models of limit orderbook (LOB) data have advanced rapidly, but their evaluation often focuses on stylised facts and selected market statistics. These measures provide useful diagnostics but may not capture the joint temporal and cross-level structure of order-book trajectories. We introduce LOB-ID, an embedding-based framework that adapts the Fr\'echet Inception Distance (FID) and Monge Inception Distance (MIND) to LOB data. To obtain domain-specific embeddings, we train the DeepLOB architecture on four months of Level-2 order-book data for five equities. We show that LOB-ID is stable across time, instruments, and embedding checkpoints, and rises monotonically under controlled distortions. We then construct a moment-matching attack against FID and a deep-book perturbation that evades statistic-based evaluation. MIND remains substantially more sensitive to both distortions. Finally, we score five generative LOB models, spanning stochastic baselines and deep learning approaches, and find that LOB-ID ranks them in line with the joint temporal and cross-level structure each captures by construction.
割り当てゲインを装ったサンプリング運: ニューラル組み合わせ最適化のためのテスト時の予算割り当ての監査
ニューラル組み合わせ最適化 (NCO) ソルバーは、インスタンスごとにサンプリングされた多くのソリューションの中から最良のものを報告します。慣例により、サンプル数はすべてのインスタンスで同一です。固定された総予算を不均一に割り当てた場合に何かが買えるかどうかは測定されていません。私たちはそれを測定し、測定自体を監査します。まず、分散内のワークロードでは、割り当てのヘッドルームは検出できません。統一 TSP-100 上の 3 つの事前トレーニング済みソルバー (POMO、AM、SymNCO) にわたって、同じ保存サンプルに対して計算および評価されたオラクル割り当ては、ゼロを除く間隔で 2.2 ~ 2.6% のゲインを報告します。サンプルから測定した同じゲインはゼロと区別できません (0.457、0.015、-0.512 パーセント)。通常のサンプル内手順に従えば、3 つのソルバーはすべて、存在しない公開された 2% レベルのゲインをサポートしていたと考えられます。このバイアスを、構築により真のゲインがゼロになるインスタンスごとのヌルに対して調整します。私たちがテストした範囲では、サンプルやインスタンスが増えても縮小しません。第二に、ファントムゲインを除去する同じ補正により、実際のゲインが保存されます。分散シフト (均一なインスタンスとクラスター化されたインスタンスが混在するワークロード) の下で、事前に登録された確認実験では、ホールドアウトされたサンプル統計に基づく割り当てにより、信号取得コストが請求されない等しい評価予算で Best-of-K が 11.5% (AM、主要エンドポイント; 95% CI [7.4, 19.7]) および 12.0% (SymNCO、複製) 改善されることがわかりました。事前に登録されたネガティブ コントロール (POMO、シフトに対して一桁強い) は -0.3% [-0.7、0.24] を示します。このゲインは、凍結された流通ラベルのベースラインを 4.2 ポイント上回っています [1.9、7.7]。同じ予算に対して 20 サンプルのプローブを請求する探索的ポリシーでは、3.4% (AM) と 4.6% (SymNCO) が維持されます。修正手順と報告チェックリストを提供し、すべてのデータ、コード、事前登録記録を公開します。
原文 (English)
Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization
Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convention, identical for every instance. Whether a non-uniform allocation of a fixed total budget would buy anything has not been measured. We measure it, and we audit the measurement itself. First, on in-distribution workloads the allocation headroom is not detectable. Across three pretrained solvers (POMO, AM, SymNCO) on uniform TSP-100, an oracle allocation computed and evaluated on the same stored samples reports a 2.2-2.6% gain with intervals excluding zero; measured out of sample the same gain is indistinguishable from zero (0.457, 0.015, -0.512 percent). Following the customary in-sample procedure, all three solvers would have supported a published 2%-level gain that does not exist. We calibrate this bias against an instance-wise null in which the true gain is zero by construction; over the ranges we test it does not shrink with more samples or more instances. Second, the same correction that removes the phantom gains preserves a real one. Under distribution shift (a workload mixing uniform and clustered instances), a pre-registered confirmatory experiment finds that allocation guided by held-out sample statistics improves best-of-k by 11.5% (AM, primary endpoint; 95% CI [7.4, 19.7]) and 12.0% (SymNCO, replication) at equal evaluation budget, with the signal-acquisition cost not charged; a pre-registered negative control (POMO, an order of magnitude more robust to shift) shows -0.3% [-0.7, 0.24]. The gain exceeds a frozen distribution-label baseline by 4.2 points [1.9, 7.7]. An exploratory policy charging a 20-sample probe against the same budget retains 3.4% (AM) and 4.6% (SymNCO). We give a correction procedure and a reporting checklist, and release all data, code, and the pre-registration record.
EgoMonth: 長期時空間記憶のための月レベルの自己中心的なビデオ ベンチマーク
マルチモーダル大規模言語モデル (MLLM) の最近の進歩により、ビデオの理解が大幅に進歩し、長いビデオ ベンチマークの数も増加しています。ただし、既存のベンチマークは主に、クリップ間の時空間的連続性のない Web ソースのビデオに依存しているため、モデルが数日または数週間の実世界の経験にわたって一貫した記憶を維持できるかどうかを評価することが困難です。最初の月レベルの自己中心的なビデオ理解ベンチマークである EgoMonth を紹介します。 EgoMonth は、20 人の参加者による 20 日から 120 日間にわたる 300 時間以上の一人称の日常生活の記録と、人間が作成した 1,443 個の多肢選択式の質問と回答のペアで構成されています。私たちは、スキーマ統合、エピソード索引付け、およびカスケード推論という 3 つの階層的な認知レベルに編成された、認知に基づいた 14 タスクの評価フレームワークを設計します。最先端のオープンソースおよびクローズドソースの MLLM を評価すると、最もパフォーマンスの高いモデルである Gemini 2.5 Pro でさえ、マクロ平均精度は 71.8% しか達成できず、人間の修正ベースラインである 94.2% を依然として 22.4 パーセント下回っていることが明らかになりました。いくつかのモデルは、ルート推論、クロスビュー空間推論、方向判断などのタスクで 25% 付近またはそれ以下の確率レベルでパフォーマンスを発揮しますが、最も強力なクローズドソース モデルでさえ依然として人間のパフォーマンスを大幅に下回っています。これらの結果は、現在の MLLM が忠実な記憶装置ではなく非可逆要約装置として機能することを示しており、本物の長期時空間記憶を備えたアーキテクチャの必要性を強調しています。
原文 (English)
EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.
LigBench: LLM ベースの研究アイデア生成のための人間と連携した統合ベンチマーク
大規模言語モデル (LLM) の急速な進歩に伴い、研究アイデアの生成への注目が高まっています。既存のアプローチにより、LLM は関連文献を検索し、研究分野に対する新しいアイデアを提案することができます。しかし、アイデア生成のための現在の評価慣行は依然として細分化されており、客観的な基準が欠如しており、多くの場合直接 LLM スコアリングに依存しているため、生成されたアイデアの一貫した分布全体にわたって統一された信頼性の高い評価を提供する能力が制限されています。この課題に対処するために、私たちは、AI 研究アイデアのきめ細かく信頼性の高い評価を可能にし、異なる世代分布にわたって一貫して適用できる自動評価ベンチマークである LigBench を提案します。さらに、ペアごとのアイデア判断モデルをトレーニングし、より客観的な比較評価をサポートする補助的な参照として機能するように調整されたデータセットである PAIR-IQ を紹介します。広範な実験により、LigBench が安定した解釈可能な評価を達成し、専門家の判断との整合性が大幅に向上することが実証されました。さらに、PAIR-IQ でトレーニングされたモデルはランク付けの精度と堅牢性が向上し、拡張性があり客観的な研究アイデアの評価のための原則に基づいた標準を確立します。
原文 (English)
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation
With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability to provide unified and reliable assessments across a coherent distribution of generated ideas. To address this challenge, we propose LigBench, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions. In addition, we introduce PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation. Extensive experiments demonstrate that LigBench achieves stable and interpretable evaluations, significantly improving alignment with expert judgments. Furthermore, models trained on PAIR-IQ exhibit enhanced ranking accuracy and robustness, establishing a principled standard for scalable and objective research idea assessment.
LipCache: エッジ画像分類サービスの認定キャッシングを備えたローカル推論プロキシ
エッジサイドのビジョン サービスが低遅延、高スループットのシナリオに向けて拡大を続ける中、信頼性を犠牲にすることなくビジョン モデルの推論コストを削減することが中心的な関心事になっています。既存のセマンティック キャッシュ手法は主に経験的な類似性しきい値に依存しています。このようなしきい値はヒット率を向上させますが、決定境界付近でサイレントな誤分類を引き起こす傾向があります。これに対処するために、画像分類用の認定されたセマンティック キャッシュ フレームワークである \texttt{LipCache} を提案します。既存のデプロイされたメイン モデル \texttt{MainNet} を変更することなく、フレームワークは軽量ネットワーク \texttt{GuardNet} を導入し、リプシッツ制約の対象となる低次元特徴空間に入力をマッピングします。次に、ローカル分類マージンと分類ヘッドのスペクトル ノルムからサンプルごとの認定再利用半径を計算します。実行時に、キャッシュされた結果は、クエリ機能が認定された再利用範囲内にある場合にのみ再利用されます。それ以外の場合、クエリは \texttt{MainNet} に戻ります。したがって、キャッシュ ヒットは、経験的なしきい値テストから、明示的な理論的境界を備えた幾何学的認証の決定に変換されます。 CIFAR、Tiny-ImageNet、SVHN などの標準的な画像分類タスク全体で、\texttt{LipCache} は、エンドツーエンドの精度低下を制限しながら最大 $1.65\times$ の速度向上を達成します。一方、受け入れられたすべてのキャッシュ ヒットは \texttt{GuardNet} 側の認定一貫性条件を満たします。さらに、強化された \texttt{GuardNet} トレーニング レシピにより、$100\%$ の認定一貫性率を維持しながら、Tiny-ImageNet マルチクラス拡張機能のキャッシュ ヒット率が大幅に向上します。これらの結果は、サンプルごとの認定された再利用により、理論的な一貫性を維持しながらメインモデルのフォールバックを削減でき、エッジで信頼性の高いキャッシュ支援推論への実行可能なアプローチを提供できることを示しています。
原文 (English)
LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service
As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models without sacrificing reliability has become a central concern. Existing semantic caching methods largely rely on empirical similarity thresholds; while such thresholds improve hit rates, they tend to introduce silent misclassifications near decision boundaries. To address this, we propose \texttt{LipCache}, a certified semantic caching framework for image classification. Without modifying the existing deployed main model, \texttt{MainNet}, the framework introduces a lightweight network, \texttt{GuardNet}, that maps inputs into a low-dimensional feature space subject to a Lipschitz constraint. It then computes a per-sample certified reuse radius from the local classification margin and the spectral norm of the classification head. At runtime, a cached result is reused only when the query feature falls inside the certified reuse ball; otherwise, the query falls back to \texttt{MainNet}. Thus, cache hits are transformed from empirical threshold tests into geometric certification decisions with explicit theoretical boundaries. Across standard image classification tasks like CIFAR, Tiny-ImageNet, and SVHN, \texttt{LipCache} achieves a measured speedup of up to $1.65\times$ with limited end-to-end accuracy degradation, while all accepted cache hits satisfy the \texttt{GuardNet}-side certified-consistency condition. Furthermore, an enhanced \texttt{GuardNet} training recipe substantially improves cache hit rates in the Tiny-ImageNet multi-class extension while maintaining a certified-consistency rate of $100\%$. These results demonstrate that per-sample certified reuse can reduce main-model fallback while preserving theoretical consistency, providing a feasible approach to reliable cache-assisted inference at the edge.
より良い分解、自由な集計: 多言語マルチホップ質問応答のためのシンセサイザー折りたたみフレームワーク
多言語検索拡張生成 (mRAG) は、複雑な多言語の質問応答のために、大規模な言語モデルにグローバルに分散された外部知識へのアクセスを提供します。最近のアプローチでは、取得した文書を英語またはクエリ言語に翻訳して言語間の意味論的なギャップを埋めるか、複雑なクエリをサブ質問に分解して中間推論プロセスを集約します。ただし、どちらの仕事にも 2 つの制限があります。まず、画一的な翻訳調整である一括翻訳では、対象言語に特有の文化的および言語的ネイティブ情報が廃棄され、翻訳ノイズが発生し、システムコストが増大します。第 2 に、貪欲な分解と集約です。制御されていない分解は、段階的な推論中にエラーを悪化させる冗長なサブ質問を生成し、推論パスにわたる最終的な集約によってこれらのエラーがさらに増幅されます。私たちは、デフォルトで翻訳を適用するのではなく、翻訳を延期する、多言語マルチホップ質問応答用のシンセサイザー折りたたみフレームワークであるメソッド Syfer で両方に対処します。 Syfer はまず、形式に制約された分解プログラムを呼び出して元の言語でサブ質問グラフを生成し、続いて分解品質チェックを行います。チェックに合格すると、サブ質問はターゲット言語の検索してから回答するポリシーに基づいて順番に回答され、チェックが失敗した場合にのみ、バイリンガルのサブ質問グラフの配置による英語翻訳パスウェイがアクティブになります。複数の言語にわたる実験では、Syfer がパフォーマンスと計算コストのバランスを適切に保ちながら、競争力のある精度を達成していることが示されています。
原文 (English)
Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering
Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for complex multilingual question answering. Recent approaches either translate retrieved documents into English or the query language to bridge the cross-lingual semantic gap, or decompose a complex query into sub-questions and aggregate the intermediate reasoning process. However, both lines of work suffer from two limitations. First, one-size-fits-all translation alignment, blanket translation discards culturally and linguistically native information unique to the target language, introduces translation noise, and inflates system cost. Second, greedy decomposition and aggregation, uncontrolled decomposition produces redundant sub-questions that compound errors during step-wise reasoning, and the final aggregation over reasoning paths further amplifies these errors. We address both with our method Syfer, a synthesizer-folding framework for multilingual multi-hop question answering that defers translation rather than applying it by default. Syfer first invokes a format-constrained decomposer to produce a sub-question graph in the original language, followed by a decomposition-quality check; when the check passes, sub-questions are answered sequentially under a retrieve-then-answer policy in the target language, and the English translation pathway with bilingual sub-question graph alignment is activated only when the check fails. Experiments across multiple languages show that Syfer attains competitive accuracy while striking a favourable balance between performance and computational cost.
TRAPSBench: 視覚言語モデルはエンコードするが認識論的拘束を表現できない
視覚的な証拠が遮られている、または混沌としている場合、モデルは棄権する必要があります。この論文では、視覚言語モデル (VLM) は棄権が必要な場合を内部的に区別できるが、とにかくそれを表現できないことを示します。 TRAPSBench は、1,404 の一致する物理ペアから手続き的に生成されたビデオ ベンチマークであり、単一の対象を絞った変更によって結果が視覚的証拠からは判断不可能になります。さらに、結果が分かる場合にはモデルが正しく回答し、結果が分からない場合には棄権することの両方をモデルに要求する、新しい堅牢な指標であるペナルティ付き認識校正スコア (PECS) を導入します。 5 つの家族にまたがる 16 の VLM では、自発的拘束は不十分で、最良の PECS は 0.292 です。ボトルネックは知覚ではなく表現です。線形プローブは、物理領域全体で最大 0.91 AUROC の隠れ状態から応答性を解読します。単層ボイド方向の操作は、棄権を因果的に誘発または抑制します。私たちの結果は、3 つのオープンウェイト ファミリー (Qwen、Gemma、LLaVA) にわたって再現されています。また、この失敗は、テキストの不確実性よりも視覚的な不確実性においてより顕著です。モデルは、視覚的な証拠が欠けている場合に比べて、テキスト上の不可能性を約 4 倍容易に検出します。この需給ギャップを解消するには、おそらく生産段階での介入が必要となるだろう。
原文 (English)
TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a single targeted change renders the outcome undeterminable from the visual evidence. Furthermore, we introduce Penalized Epistemic Calibration Score (PECS), a new robust metric that requires models to both answer correctly when the outcome is knowable, and abstain when the outcome is not. Across 16 VLMs spanning five families, spontaneous restraint is poor: the best PECS is 0.292. The bottleneck is expression, not perception: linear probes decode answerability from hidden states at up to 0.91 AUROC across physics domains; steering a single-layer void direction causally induces or suppresses abstention. Our results replicate across three open-weight families (Qwen, Gemma, LLaVA). The failure is also more pronounced in visual than textual uncertainty: models detect textual impossibility about 4x more readily than missing visual evidence. Closing this representation--output gap likely requires output-stage interventions.
GEM: A Generative Embedding Model Bridging Reasoning and Retrieval
Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, convent…
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and int…
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual toke…
Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales
Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather…
GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport
Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introd…
Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data
As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as im…
Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models
Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective ex…
Into the ORBIT for Time Series: Training Regimes for Foundation Models
Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale hete…
How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what t…
Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model
We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a…
Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks
Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back…
Physics-informed distribution of relaxation times estimation and latent-space condition monitoring of solid oxide fuel and electrolysis cells from electrochemical impedance spectroscopy
Estimating the distribution of relaxation times (DRT) fromelectrochemical impedance spectroscopy (EIS) is an ill-posed inverse problem that…
Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services
We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while…
It's How You Ask: Gender-Associated Linguistic Bias in LLMs
Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contai…
Training AI Scientists to Replicate Research
The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for…
Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples
Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challengi…
Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs
This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Advers…
Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks
6G networks will not be serving as communication infrastructures only; rather, they are expected to evolve into intelligent systems, where…
Deliberate Practice: Learning Robot Skills under a Budget
We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active sk…
Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference
Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix m…
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhib…
Algebraic Decomposition Theory for Transformer Length Generalization
Transformer-based language models are known to sometimes generalize to sequences longer than seen during training, but we lack a precise ch…
ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models
Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is especially limiting in wrist-c…
UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models
Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and perf…
CAPRI: Contract-Aware Proof Repair for Isabelle
We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theor…
MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification
Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and…
Concept Drift Detection and Adaptive Retraining of Malware Classification Models
Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning…
AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models
Analog circuit design is a time-consuming, iterative process in a nonlinear and high-dimensional design space that relies heavily on expert…
Synthetic Persona Pretraining: Alignment from Token Zero
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes crit…
Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity
When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to…
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committ…
The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity
We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \emph{unmasking growth c…
Vero: Can AI Agents Build Formally Verified Software Repositories?
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code gener…
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is diffi…
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in…
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic p…
MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence
Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic tr…
Foam-Agent: A Large Language Model-Based Multi-Agent Framework for Automating Computational Fluid Dynamics Workflows
Computational fluid dynamics (CFD) has been the main workhorse of computational physics, yet its steep learning curve and fragmented, multi…
Exploiting Symbolic Heuristics for the Synthesis of Domain-Specific Temporal Planning Guidance using Reinforcement Learning
Recent work investigated the use of Reinforcement Learning (RL) for the synthesis of heuristic guidance to improve the performance of tempo…
Identification of Probabilities of Causation: from Recursive to Closed-Form Bounds
Probabilities of causation (PoCs) are fundamental quantities for counterfactual analysis and personalized decision making. However, existin…
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging…
DomusFM: A Foundation Model for Event-Based Behavioral Monitoring in Smart-Homes
Smart-home sensor-based behavioral monitoring holds significant potential for healthcare, independent living, and early detection of functi…
Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design
We study mathematical discovery through the lens of neurosymbolic reasoning, where an AI agent powered by a large language model (LLM), cou…
Auditable Agents
LLM agents call tools, query databases, delegate tasks, and trigger external side effects. Once an agent system can act in the world, the q…
Time-Series Forecasting in Safety-Critical Environments: An Open-Source Package for EU-AI-Act-Compliant Development / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein Open-Source-Paket f\"ur die KI-VO-konforme Entwicklung
With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-c…
From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction
Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover re…
AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design
Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent…
CEON: 循環経済オントロジー ネットワーク
私たちの社会における資源利用の循環性を高めることは、持続可能性、つまりより循環的な経済への移行への道であると認識されています。そのためには、製品やコンポーネントの再利用、中古製品の再生と再製造、残り物や使用済みの材料のリサイクルなど、さまざまな循環戦略があります。これらの戦略を実現するには、インフラストラクチャ レベルで情報を共有し、製品ライフ サイクルに沿って業界セクター間で通信する必要があります。したがって、この情報共有と通信においてセマンティックな相互運用性を実現することが、循環性を高める鍵となります。しかし、製品ライフサイクルに関連する多くの関連業界セクターが関与する循環経済 (CE) 領域の知識表現は依然として課題です。このギャップを埋めるために、私たちは Onto-DESIDE プロジェクト内で Circular Economy Ontology Network (CEON) を開発しました。このオントロジー ネットワークは、分野横断的な概念を定義することで CE のギャップを埋め、セマンティクスを意識したデータの文書化を可能にすることを目的としています。建設、エレクトロニクス、繊維セクターにわたる業界横断的なデータ文書化シナリオを通じて CEON を実証します。
原文 (English)
CEON: Circular Economy Ontology Network
Increasing the circularity of resource use in our society has been recognized as a path to sustainability, i.e., transitioning into a more circular economy. There are many different circular strategies to do so, such as reusing products and components, refurbishing and remanufacturing used products, or recycling left-over or used materials. To enable these strategies, it is necessary to share information at the infrastructure level and to communicate between industry sectors along the product life cycle. Enabling semantic interoperability in this information sharing and communication is therefore a key to increasing circularity. However, knowledge representation for the circular economy (CE) domain, which involves many relevant industry sectors related to product life cycles, remains challenging. To bridge this gap, we developed the Circular Economy Ontology Network (CEON) within the Onto-DESIDE project. This ontology network aims to fill gaps in CE by defining cross-sectorial concepts and to enable semantics-aware data documentation. We demonstrate CEON through cross-industry data documentation scenarios spanning construction, electronics, and textile sectors.
Residual Modeling for High-Fidelity Learned Compression of Scientific Data
Lossy compression is essential for massive spatiotemporal data from scientific simulations. Learned compressors can achieve high compressio…
マルチタスク結合モデルからタスクエキスパートを回復する方法を学ぶ
マルチタスク モデルのマージは、複数のタスク固有の専門家を 1 つの統一モデルに統合することを目的としていますが、静的マージではパラメータの干渉が常に発生します。動的マージ モデルはこのギャップを埋めることを目的としていますが、多くの研究は、推論時にコストのかかるストレージと冗長なエキスパート コンポーネントの読み込みに依存しています。この研究では、タスク エキスパートの観点から、パラメータ干渉を、マージ プロセス中に各エキスパートに導入されるパラメータの摂動として見ます。このようなパラメータの摂動はアフィン変換としてモデル化でき、加算オフセットとして近似できることを示します。これらを動機として、パラメータ干渉を元に戻し、単一のマージされたチェックポイントからタスク エキスパートのパフォーマンスを回復するために、これらのオフセットを予測するフレームワークである Recover Task eXpert (ReTeX) を提案します。タスク ID が不明な場合に適切なエキスパートを回復するために、推論前にオフラインで計算された SVD 部分空間署名に基づくルーターフリーのタスク ID を導入します。推論時に、識別子は、指定された入力に対して部分空間が最小の射影残差をもたらすタスクを選択します。その結果、ReTeX は視覚領域と NLP 領域の両方で個人の専門家のパフォーマンスの 95% 以上を回復し、目に見えないタスクへの一般化を大幅に向上させます。重要なことに、パラメータ オフセット予測が、配布外 (OOD) タスクに対する専門知識の創発的適応補間につながることも示します。 ReTeX は、目に見えないタスクを処理するために、目に見える専門知識を適応的に補間します。私たちのコードは https://github.com/BAIKLAB/ReTeX で入手できます。
原文 (English)
Learning to Recover Task Experts from a Multi-Task Merged Model
Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference. In this work, from the perspective of task expert, we view parameter interference as parameter perturbation introduced to each expert during merging process. We show that such parameter perturbations can be modeled as affine transformation, which can be approximated as additive offsets. Motivated by these, we propose Recover Task eXpert (ReTeX), a framework that predicts those offsets, in order to undo parameter interference and recover task-expert performance from a single merged checkpoint. To recover the appropriate expert when task identity is unknown, we introduce a router-free task identifier based on SVD subspace signatures computed offline before inference. At inference, the identifier selects the task whose subspace yields the smallest projection residual for a given input. As a result, ReTeX recovers over 95% of individual-expert performance in both vision and NLP domains, while significantly improving generalization to unseen tasks. Crucially, we also show that the parameter offset prediction leads to emergent adaptive interpolation of expert knowledge for out-of-distribution (OOD) tasks. ReTeX adaptively interpolates seen expert knowledge to handle unseen tasks. Our code is available at https://github.com/BAIKLAB/ReTeX
Alipay-PIBench: コーディング エージェント向けの現実的な決済統合ベンチマーク
支払いの統合は、要求の厳しいリポジトリ レベルのソフトウェア タスクです。エージェントは、適切な製品を選択し、調整されたクライアント/サーバー フローを実装し、支払い結果を検証し、トランザクションとビジネス状態の間の一貫性を維持する必要があります。現実的な Alipay 決済統合に関するコーディング エージェントを評価するためのベンチマークである Alipay-PIBench を紹介します。これには、9 つの製品固有のプロジェクトと 18 のタスク インスタンスが含まれており、それぞれが基本的な機能完了シナリオと高度なリスク認識強化シナリオに編成されています。シナリオ固有のルーブリックは、決定論的な静的チェック、ユニットチェック、統合チェック、およびエンドツーエンドのチェックをサポートし、セマンティック要件に対する LLM 支援の評価によって補足されます。 6 つのコーディング エージェント モデルを評価し、ルーブリック合格率 (RPR) を報告します。スキルありの条件下では、平均 RPR は 68.58% から 91.37% の範囲です。 Alipay 決済統合スキルへのアクセスにより、スキルなしの状態と比較して平均 RPR が平均 10.31 パーセント ポイント向上しますが、その向上はモデル、製品、シナリオによって異なります。メソッドレベルの結果は、ソースレベルの完了、実行可能な支払い動作、支払いドメインの要件を区別します。 Alipay-PIBench は、モデルの機能を診断し、支払い統合における構造化されたガイダンスを評価するための制御された設定を提供します。
原文 (English)
Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents
Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states. We introduce Alipay-PIBench, a benchmark for evaluating coding agents on realistic Alipay payment integration. It contains nine product-specific projects and 18 task instances, each organized into Basic functional-completion and Advanced risk-aware hardening scenarios. Scenario-specific rubrics support deterministic static, unit, integration, and end-to-end checks, supplemented by LLM-assisted assessment for semantic requirements. We evaluate six coding-agent models and report rubric pass rate (RPR). Under the with-skill condition, mean RPR ranges from 68.58% to 91.37%. Access to the alipay-payment-integration skill improves mean RPR by 10.31 percentage points on average relative to the without-skill condition, with gains varying across models, products, and scenarios. Method-level results distinguish source-level completion, executable payment behavior, and payment-domain requirements. Alipay-PIBench provides a controlled setting for diagnosing model capability and evaluating structured guidance in payment integration.
類似性はどこまでも: LLM における多言語一般化は言語レベルの類似性構造に依存する
大規模言語モデル (LLM) はさまざまなタスクにわたって能力が向上していますが、その一般化 (無) 能力を定量化することは依然として難しく、限られた領域を超えて理解されることはほとんどありません。特に、LLM は英語以外の言語への多言語の一般化に苦労することが知られていますが、トレーニング データではそれが十分に証明されていません。その理由を理解するために、また一部のモデルが他のモデルよりも優れたパフォーマンスを実現できる理由を理解するために、認知科学全体にわたる研究の長い歴史に目を向け、一般化の成功は類似性空間での適切な表現から得られると主張します。私たちは、LLM の表現が異なる言語間の階層的類似構造をどの程度うまく捉えているかを調べます。驚くべきことに、LLMの潜在表現はインド・ヨーロッパ語族の階層構造をほぼ復元しており、同じサブファミリーのメンバーである言語を表現空間内で密接にグループ化していることを示した。さらに、モデルが言語の類似構造を反映する度合いが、多言語自然言語推論ベンチマークである XNLI でのパフォーマンスと相関していることを示します。これは、類似性に基づく一般化に関する古典的な研究を大規模に拡張し、類似した言語を表すモデルが、ある言語から別の言語へどのように同様により適切に一般化するかを示します。
原文 (English)
Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures
As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and poorly understood beyond limited domains. In particular, LLMs are known to struggle generalizing multilingually, to languages outside of English, and that are poorly attested in their training data. To understand why this may be, and what enables some models to perform better than others, we turn to a long history of work across the cognitive sciences, arguing that successful generalization derives from appropriate representations in similarity space. We look at how well LLMs' representations capture the hierarchical similarity structure between distinct languages. Strikingly, we show LLMs' latent representations largely recover the hierarchical structure of the Indo-European language family tree -- grouping languages that are members of the same subfamily closely together in representation space. Furthermore, we show that the degree to which models reflect the similarity structure of languages correlates with their performance on XNLI, a multilingual natural language inference benchmark. This extends classic work on similarity-driven generalization at scale, showing how models that represent similar languages similarly generalize better from one language to another.
Do LLMs Know Their Vulnerable Scenarios?
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can by…
AgenticCANN: 知識拡張された Agentic Evolution による Ascend C オペレーターの自動生成
Ascend C のオペレーターの最適化は、NPU (Neural Processing Unit) の推論パフォーマンスにとって重要ですが、ハードウェアに関する深い専門知識が必要です。大規模言語モデル (LLM) は自動 CUDA カーネル生成において有望であることが示されていますが、Ascend C の根本的に異なるプログラミング モデルにより、未解明なままの固有の課題が生じます。この論文では、低コーパス NPU 環境での Ascend C オペレーター合成の自動化に特化した、知識拡張型エージェント進化フレームワークである AgenticCANN を提案します。不慣れなハードウェアでの深刻なプラットフォーム知識不足を克服するために、AgenticCANN には、上流の実現可能性のボトルネックを解決するために、開発ライフサイクル全体にわたって構造化されたマルチレベルのドメイン洞察を提供する知識統合生成システムが組み込まれています。この基盤に基づいて、動的に実行する段階適応型エージェント進化戦略を特徴としています。 LLM インタラクション モードを特定の生成フェーズと進化フェーズに合わせて調整し、高度な探索候補の発見と高度な収束パフォーマンス調整のバランスをとります。5 つのパターン カテゴリにわたる 6 つの演算子にわたる Huawei Ascend 910B での広範な実験により、私たちの手法が要素ごとの演算子と正規化演算子で 90 ~ 100%、融合演算子で 56% の実現可能性を達成し、1B Pangu モデル推論で最大 6.65 倍の高速化が達成されることが実証されました。カーネル。さらなる分析により、知識注入は要素ごとの演算子で実現可能性を 57% から 86% に単調に向上させることが明らかになり、演算子固有の利点ではなく一般的な利点が実証されました。
原文 (English)
AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution
Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise. While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally different programming model of Ascend C introduces unique challenges that remain unexplored. In this paper, we propose AgenticCANN, a knowledge-augmented agentic evolution framework specifically tailored for automated Ascend C operator synthesis in low-corpus NPU environments. To overcome the severe platform knowledge deficit on unfamiliar hardware, AgenticCANN incorporates a knowledge-orchestrated generation system that delivers structured, multi-level domain insights across the development lifecycle to resolve the upstream feasibility bottleneck. Building on this foundation, it features a stage-adaptive agentic evolution strategy that dynamically aligns LLM interaction modes with specific generation and evolution phases, balancing high-exploration candidate discovery with high-convergence performance tuning. Extensive experiments on Huawei Ascend 910B across six operators spanning five pattern categories demonstrate that our method achieves 90 to 100 percent feasibility on elementwise and normalization operators, 56% on fusion operators, and up to 6.65$\times$ speedup on 1B Pangu model inference kernels. Further analysis reveals that knowledge injection monotonically improves feasibility from 57% to 86% on elementwise operators, demonstrating its general rather than operator-specific benefit.
DAPD: Dual-Anchored Policy Distillation
On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged i…
DiffImaginE: Diffusio を使用してエンティティ タイプを検証することを想像してください
マルチモーダル名前付きエンティティ認識 (MNER) は、各候補スパンとエンティティ タイプの仮説が共同のテキスト証拠と視覚的証拠によってサポートされているかどうかを判断します。既存の想像比較検証器は、各 (スパン、タイプ) ペアを 1 つの予測された視覚的特徴にマッピングし、多様な視覚的実現を単一のプロトタイプに圧縮し、明示的な確率的セマンティクスを使用せずに互換性スコアを提供します。 MNER 型検証を条件付き潜在拡散推論として定式化する DiffImaginE を紹介します。スパン局所化された視覚的証拠が与えられると、タイプ条件付きデノイザーは、標準化された潜在に注入されるノイズを予測します。結果として生じるノイズ除去誤差は、タイプ条件付き負の対数尤度の ELBO 一貫性のある代用値を提供し、競合するタイプの仮説を、観察をどの程度うまく説明できるかによってランク付けできるようにします。 DiffImaginE は、標準のマルチモーダル エンコーダ スタックを保持し、決定論的検証器を、Min-SNR 重み付けを使用してトレーニングされた分類子なしのガイド付き拡散スコアラーに置き換えます。タイプごとの拡散スコアを分類ロジットとして直接監視し、ノイズ レベル全体の集計を学習し、逆サンプリングを使用してモンテカルロ比較の分散を削減します。私たちの分析は、分類器を使用しないガイダンスが誘導型事後分布を鮮明にし、反対のペアリングが等しいデノイザーコストで分散を低減するときの特徴を示すことを示しています。 Twitter-2015 と Twitter-2017 の実験では、アブレーションと一対の有意性検定によってサポートされ、同じエンコーダー、補助対物レンズ、評価プロトコルの下で、一致した決定論的 ImaginE 制御に対して一貫したゲインが示されています。
原文 (English)
DiffImaginE: Imagine to Verify Entity Types with Diffusion
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. Given span-localised visual evidence, a type-conditioned denoiser predicts noise injected into its standardised latent. The resulting denoising error provides an ELBO-consistent surrogate for type-conditional negative log-likelihood, allowing competing type hypotheses to be ranked by how well they explain the observation. DiffImaginE retains a standard multimodal encoder stack and replaces the deterministic verifier with a classifier-free-guided diffusion scorer trained using Min-SNR weighting. We directly supervise per-type diffusion scores as classification logits, learn aggregation across noise levels, and use antithetic sampling to reduce Monte Carlo comparison variance. Our analysis shows that classifier-free guidance sharpens the induced type posterior and characterises when antithetic pairing reduces variance at equal denoiser cost. Experiments on Twitter-2015 and Twitter-2017 show consistent gains over a matched deterministic ImaginE control under the same encoder, auxiliary objectives, and evaluation protocol, supported by ablations and paired significance tests.
Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes…
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load
Short-term load forecasting (STLF) plays a vital role in the electric power industry. It is relevant for critical infrastructure. STLF is n…
長期にわたるターミナルタスクのための再帰的合成
ターミナル エージェント向けの高品質で長期的なトレーニング データは作成に費用がかかり、タスクごとに数百ドルから数千ドルかかることがよくあります。これは、各タスクが命令、環境、参照ソリューション、検証器の相互一貫性を保つ必要があるためです。人間によるオーサリングは拡張性がなく、大規模言語モデル (LLM) を使用した直接生成では、これらの依存関係が壊れることがよくあります。我々は、長期にわたるターミナル エージェント タスクを大規模に構築するための再帰的検証済み合成フレームワークである再帰的合成ターミナル タスク (RST) を紹介します。 RST は、検証されたシード タスクから開始して、参照ソリューションを拡張し、検証ツールと命令を新しいワークフローに再調整し、新しいサンドボックスで結果を検証し、受け入れられたタスクを後続のラウンドのシードとして再利用します。 15 回の再帰ラウンドにわたって、RST は 37,484 個の合成ターミナル エージェント タスクをタスクあたり約 $0.05 で生成します。タスクの難易度はラウンドを重ねるごとに大幅に増加します。リファレンス ソリューションの中央値は 67 行から 374 行に増加し、実行されたコマンド数の中央値は 40 から 244 に増加し、DeepSeek-V4-Pro pass@4 は $R_1$ の 90\% から $R_{15}$ の 2.5\% に低下します。トレーニングの有用性を実証するために、合成されたタスクに関して拒否サンプリングされた Qwen3.5 軌跡を収集し、それらを教師付き微調整に使用します。これらの軌道を微調整すると、Qwen3.5-27B と Qwen3.5-122B-A10B がターミナル ベンチ ~ 2、ターミナル ベンチ ハード、およびロングホライズン ターミナル ベンチで最大 10 ポイント改善され、エージェント PPO により Qwen3.5-27B が 3 つで 49.44\%、32.00\%、22.07\% に上昇しました。ベンチマークは、ベース モデルに対して 20.0\%、41.2\%、および 21.9\% の相対的な向上に相当します。さらに、15 ラウンド後の再帰には上限がありません。難易度が上昇し続けても合成収率と検証率は安定しており、ここで報告した規模をはるかに超えてプロセスを継続できることを示しています。
原文 (English)
Recursive Synthesis for Long-Horizon Terminal Tasks
High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. Human authoring does not scale, and direct generation with large language models (LLMs) often breaks these dependencies. We present Recursive Synthetic Terminal Tasks (RST), a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks at scale. Starting from verified seed tasks, RST extends the reference solution, realigns the verifier and instruction to the new workflow, validates the result in a fresh sandbox, and reuses accepted tasks as seeds for subsequent rounds. Across fifteen recursive rounds, RST produces 37,484 synthesized terminal-agent tasks at roughly $0.05 per task. Task difficulty increases substantially over rounds: the median reference solution grows from 67 to 374 lines, the median number of executed commands grows from 40 to 244, and DeepSeek-V4-Pro pass@4 drops from 90% at R1 to 2.5% at R15. To demonstrate training utility, we collect rejection-sampled Qwen3.5 trajectories on the synthesized tasks and use them for supervised fine-tuning. Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench 2, Terminal-Bench Hard, and Long-Horizon Terminal Bench, while agentic PPO lifts Qwen3.5-27B to 49.44%, 32.00%, and 22.07% on the three benchmarks, corresponding to relative gains of 20.0%, 41.2%, and 21.9% over the base model. Moreover, after 15 rounds, the recursion shows no ceiling: synthesis yield and validation rates remain stable as difficulty keeps climbing, indicating that the process can continue well beyond the scale reported here.
iARCS: 制御可能な 3D シーン生成のための反復エージェント RL
合成 3D シーンの生成は、コンピューター ビジョンや具体化された AI のデータ ソースとしてますます使用されていますが、既存のジェネレーターは、タスクに不可欠な機能上の制約を確実に満たさずに、知覚的なリアリズムを最適化することがよくあります。この不一致により、アクセシビリティ、トラバーサビリティ、空間ルールへの準拠がしばしば重要となる下流トレーニングでの合成データの有用性が制限されます。我々は、事前訓練されたシーンジェネレーターを自然言語タスクの要件に適応させる反復エージェント強化学習フレームワークである iARCS を紹介します。 iARCS は 2 段階の戦略を使用します。つまり、物理的な妥当性とレイアウトの品質を向上させるための普遍的な報酬の事前トレーニングと、それに続く、トレーニング フィードバックから繰り返し改良される LLM 生成の報酬プログラムによるタスク固有の微調整です。実験では、歩きやすさ、到達しやすさ、クリアランスを重視したタスクにおける制約の忠実度の向上、タスク固有の制約の効果的な最適化、および競技シーンの多様性が示されています。さらに、iARCS によって生成されたデータがベース ジェネレーターを改善し、制御可能なシーン編集方法だけでなく実用的な合成データ生成ツールとしてのその価値をサポートすることを示します。
原文 (English)
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements. iARCS uses a two-stage strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific fine-tuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general p…
Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, c…
AI 時代の栄養データ インフラストラクチャ: エージェント仲介研究のための FAIR の運用化
AI エージェントは栄養学の研究を加速できますが、その分析はアイデンティティ、セマンティクスを継承し、基礎となるデータの曖昧さを解放します。私たちは、自動化された使用のために FAIR を運用するソース保存インフラストラクチャである Nutrition Data Service (NDS) を紹介します。記述解決により、リリース固有のレコードを検索できるようになります。タイプ付き横断歩道は、独立してリリースされたリソースを接続します。機械可読インターフェイスは、バージョン管理されたソースとクロスウォークを公開し、AI エージェントによる分析を再生可能および監査可能にします。食品説明ベンチマークでは、NDS は高い精度を維持し、NutriBench で公開されている最高の言語モデル結果を上回りました。外部のブラインド クロスウォーク評価では、その型付きコントラクトが防御可能なリンクを優先し、サポートされていないマッピングを拒否することが示されています。個人レベルの血糖指数分析では、ピン留めされた NDS 入力はモデル間および反復実行間で同一の出力を生成しますが、オープンウェブ再構成は不安定なままです。中心的な結果は、エージェント媒介栄養研究には、データの識別、検索、およびクロスウォークのための新しいデータ インフラストラクチャが必要であるということです。
原文 (English)
Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research
AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying data. We present Nutrition Data Service (NDS), source-preserving infrastructure that operationalizes FAIR for automated use: description resolution makes release-specific records findable; typed crosswalks connect independently released resources; machine-readable interfaces expose versioned sources and crosswalks, supporting replayable and auditable analyses. On food-description benchmarks, NDS outperforms the best published language-model result on NutriBench. External and blinded crosswalk evaluations show that its typed contract favors defensible links and rejects unsupported mappings. In a person-level glycemic-index analysis, pinned NDS inputs produce identical outputs across models and repeated runs, while open-web reconstruction remains unstable. Together, these results show that agent-mediated nutrition research requires a new infrastructure that makes data identity, search, and crosswalk policy explicit.
経験則: 部分情報を使用して人工知能システムを説明する
説明可能な人工知能 (XAI) は、人工知能 (AI) システムが特定の決定にどのように到達したかを説明しようとします。私たちは、特定のデータポイントに対する AI システムの動作を予測するために最も関連する特徴を特定する新しい定式化に基づく XAI への新しいアプローチである「経験則」(RoT) 説明を提案します。 RoT が、(a) 大規模言語モデル (LLM) を使用したゼロショット分類、(b) モデルへのアクセスなしの不透明な AI システムの監査、(c) 科学的発見における AI の使用において、XAI を実現するのにどのように適しているかを示します。さらに、RoT は主要な AI 規制の特定の要件を満たし、XAI 実践者に使い慣れたインターフェイスと視覚化を提供し、モデルに依存せず、代替手段よりも大幅に高速です。コードは次から入手できます: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information
原文 (English)
Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information
Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint. We show how RoT is well-suited to enable XAI in: (a) zero-shot classification using large language models (LLMs), (b) auditing of opaque AI systems without model access, and (c) the use of AI in scientific discovery. Additionally, RoT meets specific requirements from leading AI regulations, provides a familiar interface and visualisations for XAI practitioners, is model-agnostic, and is substantially faster than alternatives. Code available at: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information
グロタンディーク定数のための長期的な AI 研究: 人間と AI の数学的コラボレーションにおけるケーススタディ
AI エージェントは数学研究でますます使用されていますが、効果的な使用方法が不明瞭なことがよくあります。これに向けて、組み合わせ問題とその連続緩和の間の硬さを捉えるグロタンディーク定数 $K_G$ の境界を改善するために AI がどのように使用されたかに関する広範なケース スタディを紹介します。具体的には、$K_G$ の正確な値は不明ですが、最近、最もよく知られている範囲を \[ \frac{6\pi}{11} \;\le\; に厳しくしました。 K_G \;\le\; \frac{\pi}{2\log(1+\sqrt2)} - 10^{-4}。 \] 重要なのは、これらの改善は、分野の専門家によって新規とみなされる洞察に到達できる AI 研究システムを使用して達成されたことです。数学の研究に AI を使用した経験について、特にその長所と短所について詳しく説明します。また、AI が画期的な洞察に到達するための理想的な条件を作り出す経験についても説明します。
原文 (English)
Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration
AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their continuous relaxations. Specifically, while the precise value of $K_G$ is not known, we recently tightened the best known bounds to \[ \frac{6\pi}{11} \;\le\; K_G \;\le\; \frac{\pi}{2\log(1+\sqrt2)} - 10^{-4}. \] Crucially, these improvements were achieved using an AI research system that could arrive at insights deemed novel by domain experts. We give a detailed discussion of our experience using AI for mathematics research, particularly touching upon its strengths and weaknesses, as well as our experience with creating ideal conditions for AI to arrive at breakthrough insights.
MBA: 現実世界のビジネスアイデアのためのマルチモーダルベンチマークとエージェント
大規模言語モデル (LLM) を活用したエージェント システムは、ビジネスのアイデア創出に新たな機会をもたらしました。しかし、実世界のコンテキストには本質的にマルチモーダルな性質があるにもかかわらず、既存のアプローチは依然としてテキストのみのパラダイムに限定されています。そこで、ビジネスアイデアエージェントのトレーニングと評価のための初のマルチモーダルベンチマークである MBA-Bench を紹介します。これは 6 つのドメインにわたる 30,000 のサンプルで構成され、各ドメインはテキストだけでは完全には伝えられない明確な視覚的手がかりによって特徴付けられます。具体的には、画像に自動的にキャプションを付け、GPT-4o を使用して、検索クエリの生成、市場証拠の検索、および証拠強化合成を通じて 3 つのビジネス質問のそれぞれに対して 5 つの参考アイデアを生成します。以前の作業に続き、MLLM-as-a-Judge を使用して 6 つのビジネス指向の基準にわたってエージェントを評価します。基準が非表示または開示される設定を検討するために、ブラインドおよび既知のそれぞれに対して MBA-b および MBA-k を提示します。私たちは両方を 2 つの新しい報酬目標 (創造性と実現可能性) に基づいてトレーニングしますが、MBA-k は公開されている 6 つの基準を合計 8 つさらに最適化します。どちらも、LoRA ベースの監視付き微調整を介してトレーニングされ、その後、これらの設定固有の報酬を使用してグループ相対ポリシーの最適化が行われます。 MBA ベンチでの大規模な実験のために、キャプションのみまたはマルチモーダル入力のいずれかに対応する 2 つのベースラインを設定しました。後者はいくつかの指標でクローズド ソースのパフォーマンスに近づきます。 MBA-b と MBA-k は、キャプション ベースラインをそれぞれ 63.9% と 77.1% 上回り、マルチモーダル ベースラインを 25.6% と 35.8% 上回っています。
原文 (English)
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.
OEIS Open: 言語モデルはどれだけの推測を定理に変えることができますか?
私たちは、OEIS からの 492 のオープンな数学的予想に基づくベンチマークである OEIS Open を構築します。このベンチマークは、Tsoukalas らによって Lean で形式化されています。これらの推測は、これまで特注のエージェントを使用してのみ試みられていましたが、当社のオープンソース評価コードは、それらに対してあらゆる汎用言語モデル (LM) を実行し、LM 不正行為の試みに対して安全です。最小限のツールセットを装備した LM は、1 回の試行あたり $50 の予算でこれらの推測のうち 147 件を解決し、OEIS Open で 30% のスコアを獲得したことがわかりました。 OEIS Open Lite は、より安価な評価を目的とした 100 個の推測のランダムなサブセットです。試行ごとに $200 の予算で評価すると、現時点で最高の LM は OEIS Open Lite で 44% のスコアを獲得します。 LM が arXiv の 476,000 件の論文を介して数学文献にアクセスできるようにしても、OEIS Open Lite のパフォーマンスは向上せず、より洗練されたエージェント ループを使用しても向上しませんでした。この研究で取り上げられている推測は数学的重要性が不確かであり、ほとんどの推測はこれまでほとんど注目されていなかったと思われます。それにもかかわらず、私たちの結果は、LMが未解決の研究の推測を自律的にかつ適度なコストで解決できることを示しています。
原文 (English)
OEIS Open: How many conjectures can language models turn into theorems?
We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source evaluation code runs any generic language model (LM) against them, and is secure against LM cheating attempts. We find that LMs equipped with a minimal set of tools resolve 147 of these conjectures with a budget of \$50 per attempt, scoring 30% on OEIS Open. OEIS Open Lite is a random subset of 100 conjectures for cheaper evaluation. When evaluated with a budget of \$200 per attempt, the best current LM scores 44% on OEIS Open Lite. Giving LMs access to the mathematics literature via 476,000 papers from arXiv did not increase performance on OEIS Open Lite, and nor did using more sophisticated agent loops. The conjectures covered in this work are of uncertain mathematical significance, and most have likely received little previous attention. Nevertheless, our results show that LMs can resolve open research conjectures autonomously and at modest cost.
The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot
Generative artificial intelligence (AI) facilitates content production and enhances ideation, with potentially important implications for d…
Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries
To evaluate a multi-representational framework in which large language model (LLM)-generated expert summaries of intensive care unit (ICU)…
Cueless EEG imagined speech for subject identification: dataset and benchmarks
Electroencephalogram (EEG) signals have emerged as a promising modality for biometric identification. While previous studies have explored…
Unmasking Conversational Bias in AI Multiagent Systems
Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application…
はい、Q ラーニングはオフラインのインコンテキスト RL に役立ちます
既存のオフライン インコンテキスト強化学習 (ICRL) 手法は、主に教師ありトレーニング目標に依存していましたが、オフライン RL 設定には制限があることが知られています。この研究では、オフライン ICRL フレームワーク内での RL 目標の統合を検討します。 150 を超える GridWorld および MuJoCo 環境由来のデータセットでの実験を通じて、RL 目標を最適化すると、さまざまなデータセット カバレッジ、構造、専門知識レベル、環境の複雑さにわたって、広く採用されているアルゴリズム蒸留 (AD) と比較してパフォーマンスが平均約 30% 向上することが実証されました。さらに、困難な XLand-MiniGrid 環境では、RL 目標により AD のパフォーマンスが 2 倍になりました。私たちの結果は、価値学習中に保守主義を追加すると、テストしたほぼすべての設定でさらなる改善がもたらされることも明らかにしました。私たちの調査結果は、ICRL の学習目標と RL の報酬最大化の目標を一致させることの重要性を強調し、オフライン RL が ICRL を前進させるための有望な方向性であることを示しています。
原文 (English)
Yes, Q-learning Helps Offline In-Context RL
Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings. In this study, we explore the integration of RL objectives within an offline ICRL framework. Through experiments on more than 150 GridWorld and MuJoCo environment-derived datasets, we demonstrate that optimizing RL objectives directly improves performance by approximately 30% on average compared to widely adopted Algorithm Distillation (AD), across various dataset coverages, structures, expertise levels, and environmental complexities. Furthermore, in the challenging XLand-MiniGrid environment, RL objectives doubled the performance of AD. Our results also reveal that the addition of conservatism during value learning brings additional improvements in almost all settings tested. Our findings emphasize the importance of aligning ICRL learning objectives with the RL reward-maximization goal, and demonstrate that offline RL is a promising direction for advancing ICRL.
Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision
Efficiently adapting large pretrained models is critical under tight compute and memory budgets. While Parameter-Efficient Fine-Tuning (PEF…
How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG
By retrieving contexts from knowledge graphs, graph-based retrieval-augmented generation (GraphRAG) enhances large language models (LLMs) t…
Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
Vision Language Models (VLMs) have shown promise in automating image diagnosis and interpretation in clinical settings. However, developing…
Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models
Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay…
REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models
Large language models (LLMs) often express verbal confidence that is poorly aligned with actual correctness, limiting their reliability in…
Gradual Code-Switching as Inference-Time Cross-Lingual Representational Alignment for LLMs
While large language models (LLMs) have achieved notable progress in multilingual settings, their performance remains uneven across languag…
StarEmbed: Benchmarking Time Series Foundation Models on Astronomical Observations of Variable Stars
Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling.…
DiffGRM: Diffusion-based Generative Recommendation Model
Generative recommendation (GR) is an emerging paradigm that represents each item via a tokenizer as an n-digit semantic ID (SID) and predic…
CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning
Urban socio-economic sensing plays a vital role in advancing global sustainable development goals. With the advent of Large Vision-Language…
SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gain…
Automated Design Optimization via Strategic Search with Large Language Models
Optimization methods have long advanced many fields, yet they struggle when faced with design problems where the search space and design pa…
Security and Detectability Analysis of Unicode Text Watermarking Methods against Large Language Models
Securing digital text is becoming increasingly relevant due to the widespread use of large language models. Individuals' fear of losing con…
RadarGen: Automotive Radar Point Cloud Generation from Cameras
We present RadarGen, a diffusion model for synthesizing realistic automotive radar point clouds from multi-view camera imagery. RadarGen ad…
Learning Latency-Aware Orchestration for Multi-Agent Systems
Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high i…
Architecture Before the Formula: Individuating Neural Architecture Beyond the Composite Map
Neural architecture is often identified by module syntax, computation graphs, or the composite functions they realize. These descriptions a…
Safe Exploration via Policy Priors
Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. simulated)…
MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models
Large Language Models (LLMs) are increasingly deployed in sensitive applications including psychological support, healthcare, and high-stak…
CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
Large Language Models excel in high-resource programming languages but struggle with low-resource ones. Existing research related to low-re…
Automatic Termination Strategy of Inelastic Neutron-scattering Measurement Using Bayesian Optimization for Bin-width Selection
Currently, an excessive amount of event data is being obtained in four-dimensional inelastic neutron-scattering experiments. A method for a…
Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI
We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions…
In-context superposition: human-like working memory interference in large language models
Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments. This capacity, known as…
A Q-learning-based QoS-aware multipath routing protocol in IoMT-based wireless body area network
The Internet of Medical Things (IoMT) enables intelligent healthcare services but faces challenges such as dynamic topology, energy constra…
IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development
Adoption of AI-assisted development in 2025 exposed a tool-agnostic failure pattern: experienced developers using frontier models were meas…
Cat-DPO: Category-Adaptive Safety Alignment
Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and rel…
Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference
Expressway video anomaly detection is important for traffic safety, but remains challenging across diverse scenes, particularly for far-fie…
SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models
We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is…
Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance
Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, e…
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking
Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and…
INSHAPE: Instance-Level Shapelets for Interpretable Time-Series Classification
Discovering shapelets -- i.e., discriminative temporal patterns within time series -- has been widely studied to address the inherent compl…
多腕ベイジアン バンディットのアニーリングされたソフトマックスの貪欲さ
検証可能な報酬を伴う強化学習 (RLVR) および GRPO などのグループベースのポリシー最適化手法は、プロンプトごとに複数の完了をサンプリングし、参照ポリシーに対する KL ペナルティによって正規化された、より高い報酬を持つポリシーの確率を高めることにより、確率的ポリシーを更新します。これらの更新には、認識論的不確実性を追跡する明示的なメカニズムは含まれていません。この論文では、なぜそのような不確実性を問わない更新が効果的であるのかについて、定型化された説明を研究します。多腕ベイジアン ベルヌーイ バンディットにおける経験的平均報酬のソフトマックスに従ってアクションを選択するアニーリングされたソフトマックス (ボルツマン) ポリシーを分析します。最適に近いアームが豊富にあることを意味する、事前の線形アッパーテール条件 ($\beta$-規則性の $\beta=1$ の場合) では、アニーリングされたソフトマックス グリーディがベイズ リポート $\tilde{O}(m + T/m)$ を達成すること、特にアームの数が $m = にスケールされる場合 $\tilde{O}(\sqrt{T})$ を達成することを証明します。 \シータ(\sqrt{T})$。これは、この体制における最適に近いベイズの後悔率であり、経験的平均の貪欲さによっても達成されます。 $\beta$-規則性の下では、多くのアームは学習を通じて最適値に近い経験的平均を維持するため、ソフトマックスが経験的に最良でないアームをサンプリングすると、そのアームは明らかに劣ったアームではなく、最適に近い別のアームになる傾向があります。対照的に、アームの数が少ない場合、同じ種類のソフトマックス ポリシーは直線的な後悔に見舞われる可能性があります。この結果は、RLVR と構造的に類似していることも示しています。ここでは、正しい完了を生成する無視できない確率を持つ基本ポリシーが $\beta$-規則性の役割を果たします。
原文 (English)
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling multiple completions per prompt and increasing the policy's probability on those with higher reward, regularized by a KL penalty toward a reference policy. These updates do not include explicit mechanisms that track epistemic uncertainty. This paper studies a stylized explanation for why such uncertainty-agnostic updates can nevertheless be effective. We analyze an annealed softmax (Boltzmann) policy that selects actions according to a softmax of empirical mean rewards in a many-armed Bayesian Bernoulli bandit. Under a linear upper-tail condition on the prior (the $\beta=1$ case of $\beta$-regularity), which implies an abundance of near-optimal arms, we prove that annealed softmax greedy achieves Bayes regret $\tilde{O}(m + T/m)$, and in particular $\tilde{O}(\sqrt{T})$ when the number of arms scales as $m = \Theta(\sqrt{T})$. This is the near-optimal Bayes regret rate in this regime, attained also by empirical-mean greedy. Under $\beta$-regularity, many arms maintain empirical means close to the optimum throughout learning, so when softmax samples an arm other than the empirically best, that arm tends to be another near-optimal one rather than a clearly inferior one. By contrast, with a small number of arms, the same kind of softmax policy can suffer linear regret. The result also provides a structural analogy to RLVR, where a base policy with a non-negligible probability of producing a correct completion plays the role of $\beta$-regularity.
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise prefer…
Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection
Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expens…
Constitutional On-Policy Safe Distillation
On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged informat…
トランスフォーマーには 3 つの投影が必要ですか? QKV バリアントの体系的な研究
トランスフォーマーは、クエリ、キー、値 (QKV) アテンションの定式化が中心的な役割を果たし、さまざまな AI タスクの標準ソリューションとなっています。しかし、これら 3 つの予測の個々の寄与と、一部を省略した場合の影響については、依然として十分に理解されていません。 3 つの射影共有制約を系統的に評価します。a) Q-K=V (共有キーと値)、b) Q=K-V (共有クエリキー)、c) Q=K=V (単一射影)。最後の 2 つのバリアントは、対称的なアテンション マップを生成します。これに対処するために、2D 位置エンコーディングによる非対称の注意も調査します。合成タスク、ビジョン (MNIST、CIFAR、TinyImageNet、異常)、言語モデリング (10B トークン上の 300M および 1.2B パラメーター モデル) にわたる実験を通じて、当社のトランスフォーマーは QKV トランスフォーマーと同等か、場合によってはそれよりも優れたパフォーマンスを発揮することがわかりました。言語モデリングでは、Q-K=V 射影共有により、わずか 3.1% のパープレキシティ低下で 50% の KV キャッシュ削減が達成されます。重要なのは、射影共有はヘッド共有 (GQA/MQA) を補完するものです。Q-K=V と GQA-4 を組み合わせると 87.5% のキャッシュ削減が得られ、Q-K=V + MQA では 96.9% が達成され、実用的なオンデバイス推論が可能になります。キーと値は同様の表現空間を占有することができ、注意は低ランク領域で動作するため、Q-K=V は品質を維持しますが、Q=K-V は注意の方向性を壊すことを示します。私たちの結果は、投影共有を、直接的で定量化可能な推論メモリの利点を備えた注意力の結びつきの未解明な例として体系的に特徴付けており、特にエッジ展開に価値があります。コードは https://github.com/anusamadan02/Do-Transformers-Need-3-Projections で公開されています。
原文 (English)
Do Transformers Need Three Projections? Systematic Study of QKV Variants
Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting some remain poorly understood. We systematically evaluate three projection sharing constraints: a) Q-K=V (shared key-value), b) Q=K-V (shared query-key), and c) Q=K=V (single projection). The last two variants produce symmetric attention maps; to address this, we also explore asymmetric attention via 2D positional encodings. Through experiments spanning synthetic tasks, vision (MNIST, CIFAR, TinyImageNet, anomaly), and language modeling (300M and 1.2B parameter models on 10B tokens), we discovered that our transformers perform on par or occasionally better than the QKV transformer. In language modeling, Q-K=V projection sharing achieves 50% KV cache reduction with only 3.1% perplexity degradation. Crucially, projection sharing is complementary to head sharing (GQA/MQA): combining Q-K=V with GQA-4 yields 87.5% cache reduction, while Q-K=V + MQA achieves 96.9%, enabling practical on-device inference. We show that Q-K=V preserves quality because keys and values can occupy similar representational spaces and attention operates in a low-rank regime, whereas Q=K-V breaks attention directionality. Our results systematically characterize projection sharing as an underexplored instance of weight tying in attention, with direct, quantifiable inference memory benefits, particularly valuable for edge deployment. The code is publicly available at https://github.com/Brainchip-Inc/Do-Transformers-Need-3-Projections
Certifiable Semantic Agreement Among LLM Agents: What the Admissibility Instrument Decides
Can a committee of LLM agents reach agreement that is certifiable at the level of meaning, not only at the level of a label? We build a pro…
SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation
Low-Rank Adaptation (LoRA) enables efficient adaptation of large pretrained models to downstream tasks by parameterizing weight updates wit…
VLM 内の偽装されたビジュアル コンテキストの隠れた進化
ビジュアル トークンは、生の外部シグナルとして大規模言語モデル (LLM) に入力されます。それらがどのように意味のある表現に変換され、言語空間と相互作用するかは、統合アーキテクチャに完全に依存します。ビジュアル トークンを入力シーケンス内のコンテキスト内プロンプトとして扱うか、LLM の中間層に直接挿入するかによって異なります。これらのアーキテクチャ上の選択が視覚情報にどのような影響を与えるか、また LLM と統合するための内部変換については、制御された比較と理解がまだ十分に行われていません。単一画像、複数画像、およびビデオのベンチマークにわたる同一のトレーニング条件下で、インコンテキストおよびレイヤーごとのインジェクション VLM 統合パラダイムを評価することで、公正な比較を提供します。そうすることで、視覚トークンが、言語構造を欠く生の表現である偽装された視覚コンテキストとして LLM に入力されるが、統合パラダイムに応じて徐々に再形成され、それぞれが視覚信号の根本的に異なる周波数特性を捕捉する、隠れた進化を明らかにします。私たちは、LLM 内のこの進化が、VLM がどのような視覚的特徴を効果的に利用できるか、視覚表現が言語空間とどのように連携するか、そして最終的にはさまざまなタスクにわたって各パラダイムがどのように実行されるかを決定することを示します。さらに、注意の割り当てだけでは不十分であり、パフォーマンスは各層の視覚的表現の品質によって左右されることを示します。
原文 (English)
The Hidden Evolution of Disguised Visual Context inside the VLM
Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language space depends entirely on the integration architecture. Whether by treating visual tokens as in-context prompts within the input sequence or injecting them directly into the LLM's intermediate layers. A controlled comparison and understanding of how these architectural choices affect visual information and its internal transformation to integrate with the LLM remains underexplored. We provide a fair comparison by evaluating in-context and layer-wise injection VLM integration paradigms under identical training conditions across single image, multi-image, and video benchmarks. In doing so, we uncover a hidden evolution where visual tokens enter the LLM as disguised visual context, raw representations lacking linguistic structure, but are progressively reshaped depending on the integration paradigm, each capturing fundamentally different frequency characteristics of the visual signal. We show that this evolution inside the LLM determines what visual features the VLM can utilize effectively, how visual representations align with the language space, and ultimately how each paradigm performs across different tasks. We further demonstrate that attention allocation alone is insufficient, and that performance is driven by the quality of visual representations at each layer.
Communication Heterogeneity and Collective Consensus in Neural Cellular Automata
Reaching global agreement from purely local interactions is a defining problem of collective intelligence, and most models of it assume tha…
Early Warning Signals for OpenVLA Failure under Visual Distribution Shift
Visual shifts can cause a vision-language-action policy to fail after initially plausible behavior. We ask whether OpenVLA's internal activ…
LLM-Based Test Oracles: Source-of-Authority Taxonomy -- A Systematic Literature Review
Large language models (LLMs) increasingly decide whether software behaves correctly, either by writing a test oracle or by acting as one. Y…
XAI 主導のデータ削減による時系列分類のスケーリング
時系列の Explainable AI (XAI) はアルゴリズム的に大幅な成長を遂げていますが、下流のタスクに測定可能なパフォーマンスの向上をもたらすというその有用性は依然として十分に検討されていません。このホワイトペーパーでは、時系列分類 (TSC) における効果的なデータ削減のために XAI アトリビューション手法を再利用する新しい方法論である drXAI を紹介することで、このギャップを埋めます。最新の TSC における中心的な課題はスケーラビリティです。 Transformers などの最先端のモデルは、シーケンスの長さに対して 2 次の複雑性を示し、チャネル数に対して 1 次の複雑さを示します。これにより、大規模なデータセットの計算が法外に難しくなります。 drXAI は、高速な GPU アクセラレーション分類器 (Hydra) を使用してローカル アトリビューションを生成することで、この問題に対処します。これらをグローバルな特徴重要度スコアに集約し、自動化されたエルボーカット ヒューリスティックを採用して、手動のしきい値を必要とせずに最も顕著な特徴を選択します。私たちは、合成データセットと現実世界の一変量データセットおよび多変量データセットの両方でアプローチを評価します。合成ベンチマークでは、drXAI は、従来のベースラインが失敗するグラウンドトゥルース機能を正常に回復します。実世界のデータでは、drXAI は、完全なデータセットでトレーニングされたモデルと同等の分類精度を維持しながら、80% ~ 90% のデータ削減を達成します。最も重要なことは、drXAI を使用すると、ConvTran のようなリソースを大量に消費するモデルを、メモリの制約により以前はアクセスできなかったデータセットに拡張できることを示しています。私たちの結果は、XAI を解釈しやすさだけでなく、時系列分析における特徴選択とスケーラビリティのための堅牢なツールとして使用する利点を示しています。すべてのコードとデータは公開されています。
原文 (English)
Scaling Time Series Classification via XAI-Driven Data Reduction
Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodology that repurposes XAI attribution methods for effective data reduction in Time Series Classification (TSC). The core challenge in modern TSC is scalability; state-of-the-art models, such as Transformers, exhibit quadratic complexity relative to sequence length and linear complexity relative to the number of channels. This renders them computationally prohibitive for massive datasets. drXAI addresses this by using a fast, GPU-accelerated classifier (Hydra) to generate local attributions. We aggregate these into global feature importance scores and employ an automated elbow-cut heuristic to select the most salient features without requiring manual thresholds. We evaluate our approach on both synthetic and real-world univariate and multivariate datasets. On synthetic benchmarks, drXAI successfully recovers ground-truth features where traditional baselines fail. On real-world data, drXAI achieves between 80% and 90% data reduction while maintaining classification accuracy comparable to models trained on the full dataset. Most importantly, we show that drXAI allows resource-intensive models like ConvTran to scale to datasets that were previously inaccessible due to memory constraints. Our results show the benefits of using XAI not just for interpretability, but as a robust tool for feature selection and scalability in time series analysis. All our code and data are openly available.
Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training. On a 6.78B-parameter MoE languag…
Vibe to Code: Elucidating Strategic Oscillation of Tacit Knowledge in Generative AI Design Workflows -- An Exploratory Qualitative Study
The rapid adoption of generative AI tools has created new literacy demands for designers who must verbalize tacit knowledge through natural…
Moral Hazard in Multi-Agent Language Models
Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr\"om's…
A Distributional Robustness Margin For Pathology Foundation Models
Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut lear…
SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups
Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Exis…
Commit Locally, Exit Globally: Coordinating Adaptive Sampling and Early Exit in Diffusion Language Models
Diffusion language models expose a provisional prediction at every denoising step, and on many tasks the candidate answer inside it stabili…
Coordinated incentives in AI-generated misinformation governance
With the rapid diffusion of AI-generated content, AI-driven misinformation is becoming increasingly pervasive and difficult to govern, unde…
Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks
Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parall…
Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction
Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other…
TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability
We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof…
Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning
Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online…
Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation
Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's re…
Governing Agentic AI in FinTech
Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and…
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dro…
Keep the Future, Drop the Rollout: RIFT for World Action Models
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask w…
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolat…
Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians
A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical…
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction hist…