AIニュース 2026-08-11
自動生成: 2026-08-11 11:05 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
What building an AI-native finance function taught meOpenAI
OpenAI CFO Sarah Friar shares five lessons for building an AI-native…
-
Model ML completes finance work more efficiently with GPT-5.6 SolOpenAI
Model ML uses GPT-5.6 Sol to carry finance work from research and ana…
-
Expanding Daybreak as the Cyber Defense Window NarrowsOpenAI
Meet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available t…
-
OpenAI’s letter to Governor Abbott on responsible AI infrastructure in TexasOpenAI
OpenAI sent Governor Greg Abbott a letter outlining its commitment to…
-
Putting frontier cyber models in more trusted handsOpenAI
Approved Daybreak partners can use OpenAI’s frontier cyber models to…
-
ザッカーバーグCEO、超知能の集中化に警鐘 「単一の善意ある超知能は存在しない」とオープンモデル公開再開へITmedia AI+
MetaのザッカーバーグCEOは、超知能の分散化と個人のエンパワーメントを訴える論考を公開した。単一の超知能への集権化を否定し、権力の均衡…
-
As AI-led attacks multiply, OpenAI launches a new cyber modelTechCrunch AI
OpenAI is expanding its AI cybersecurity defense program Daybreak, an…
トピック別件数
- LLM/生成AI 123件
- 研究/論文 122件
- エージェント 79件
- 画像/動画生成 38件
- ビジネス/資金調達 20件
- ロボティクス 12件
- ハードウェア/半導体 10件
- その他 8件
- 規制/政策 3件
日本語メディア11件
ITmedia AI+ (日本語)
2028年にSOCの人手対応、30%減へ 常態化する攻撃に企業は勝てるのか
AIが混乱を拡大させる中で、セキュリティリーダーはどう対応するか。ガートナーは、企業が競争優位を保つために注力すべき優先事項を3つ挙げた。その中身とは。
ザッカーバーグCEO、超知能の集中化に警鐘 「単一の善意ある超知能は存在しない」とオープンモデル公開再開へ
MetaのザッカーバーグCEOは、超知能の分散化と個人のエンパワーメントを訴える論考を公開した。単一の超知能への集権化を否定し、権力の均衡が安全の基礎であると主張。モデル公開審査の独立組織委任や政府へのチェックポイント提供などの管理策を提示しつつ、オープンソースモデルの公開再開…
Meta、ローカル動作に特化したオープンモデル「Muse Glimmer」公開 Apache 2.0で提供
MetaのAI研究部門は、約296億パラメータのオープンウェイトAIモデル「Muse Glimmer」を公開した。Apache 2.0ライセンスで提供され、PCやMacのGPU1基で動作する。上位モデルからの蒸留によりローカル環境でのエージェント処理やマルチモーダル推論に最適化…
親が子にAIを使わせる理由は「勉強に役立つ」ではなく「うちの子だけは遅れさせたくない」? 2000人を調査
米シカゴ大学などに所属する研究者らがPNASで発表した論文「Social dynamics of AI adoption in parents’ educational decisions」は、親が子どもに生成AIを使わせる判断の背景を調査した研究報告だ。
AIエージェントの「Skills」「MCP」などをまとめる標準規格登場 CodexやVS Codeなど対応、Claudeは未対応
VercelがAIエージェント拡張のオープン標準「Agent Plugins」を発表。「Agent Skills」やMCPサーバ設定を共通形式にし、クライアント間で使い回せるようにする。
NEC、部門長から社員まで「全員AI」の新組織
AIマネージャーが都度AI社員を生成して役割を任命する。
“脱モノ売り”のリコー 同社が説く「営業AXの勘所」とは
営業AXは、どう進めれば成果が出るのか。営業力で定評のあるリコージャパンの、“脱モノ売り”を図り、“課題解決型”への転換から勘所を探る。
「Apex人材が少ない」 Salesforce導入企業の約9割で「属人化」が課題に
コパードは、Salesforceの開発・運用におけるAI活用実態調査の結果を発表した。約9割が業務の属人化に課題意識を持つ中、8割強が「業務をAIで平準化できる」と期待している。一方で、AI活用層は「検証体制の不在」という新たな壁に直面していることも明らかになった。
避難所でAI使ってサービス開発「イマココナビ」 ニッチな生活情報も被災者主導で共有
熊本地震の発生直後、人工知能(AI)を活用し、被災者どうしが生活情報をリアルタイムで共有するサイトが生まれ、好評を呼んでいる。子供が遊べる公園、爬虫類のペットフードの売り場―。行政では担えないニッチな内容を共有できることも強みで、すでに16万人以上がサイトに足を運んだ。「情報」…
「VIVANT」シーズン2ではAIが活躍? 個人的に気になったこと
あくまでフィクションなのは重々承知しています。
テレ朝「映像全編AI生成」のCMを制作 「今後もAIを活用した制作に挑戦」
テレビ朝日が同局初の「映像全編AI生成」CMを制作。サントリー「GREEN DA・KA・RA」との30秒作品で、4月発足の「AIクリエイティブスタジオ」が制作。今後も生成AIを活用したCM制作に挑戦するという。
海外メディア5件
TechCrunch AI (英語)
OpenAI reportedly completed a $7 billion employee tender offer
San Francisco's housing market is in trouble again.
As AI-led attacks multiply, OpenAI launches a new cyber model
OpenAI is expanding its AI cybersecurity defense program Daybreak, and rolling out a new cyber-trained AI model with it.
Mark Zuckerberg’s AI manifesto is exactly why people don’t like AI
On Monday, Mark Zuckerberg published a 6,500-word manifesto about personal AI, largely about the possibilities for the "personal superintel…
Tech industry is buzzing after a Claude agent hacked into a gym
An OpenClaw agent hacked into a gym's reservation system to bump its human boss higher on a class' waitlist. And the tech industry took not…
Discovered Materials is playing AI whack-a-mole to hunt cooler chips
Discovered Materials raised $9 million to fund the hunt for more novel materials to build more efficient chips.
公式ブログ5件
OpenAI (英語)
What building an AI-native finance function taught me
OpenAI CFO Sarah Friar shares five lessons for building an AI-native finance function, from automated forecasting to stronger controls and…
OpenAI’s letter to Governor Abbott on responsible AI infrastructure in Texas
OpenAI sent Governor Greg Abbott a letter outlining its commitment to responsible AI infrastructure in Texas. The letter supports reliable,…
Model ML completes finance work more efficiently with GPT-5.6 Sol
Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
Expanding Daybreak as the Cyber Defense Window Narrows
Meet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available through Daybreak Red for authorized vulnerability research, exploit val…
Putting frontier cyber models in more trusted hands
Approved Daybreak partners can use OpenAI’s frontier cyber models to deliver authorized, governed cybersecurity services to customers.
論文295件
arXiv cs.AI (英語)
マルチラベルグラフ基盤モデルに向けて: 単一ベクトル表現学習からマルチセマンティックベース学習へ
マルチラベル ノードの分類は、ノードが複数のセマンティクスを同時に示すため、グラフ学習において重要かつ困難なタスクです。マルチラベル ノード分類の既存の方法は、複数のラベルを効果的にモデル化できますが、同じグラフ ドメイン内でモデルをトレーニングおよびテストする必要があるドメイン内シナリオのみを考慮しているため、クロスドメイン一般化が制限されます。最近、グラフ基盤モデル (GFM) が、多様なグラフ ドメインおよび下流タスクにわたって転送可能なグラフ表現を学習するための有望なパラダイムとして浮上しました。ただし、既存の GFM は単一ラベルの仮定に基づいて構築されており、すべてのノードがセマンティクスの 1 つのクラスのみを含むと任意にみなされ、単一の表現に埋め込まれます。マルチラベル ノードの場合、そのような表現は基本的に表現空間内の単一点で複数のセマンティクスを近似するため、必然的にセマンティクスのもつれが生じ、複数のラベルの同時識別が困難になります。これらの制限に対処するために、クロスドメインのマルチラベルノード分類のフレームワークであるマルチセマンティックベーシスグラフ基盤モデル (MSB-GFM) を提案します。具体的には、各マルチラベルノードを意味ベースの適応構成としてモデル化するマルチ意味ベース表現学習パラダイムを導入し、それによって複数のセマンティクスをモデル化するための柔軟な表現能力を可能にします。さらに、効果的なクロスドメイン知識伝達のためのドメイン敵対的トレーニングを備えたセマンティック構造のデュアルチャネル アーキテクチャを開発します。広範な実験により、私たちのモデルの有効性が実証されました。
原文 (English)
Towards Multi-Label Graph Foundation Models: from Single-Vector Representation Learning to Multi-Semantic Basis Learning
Multi-label node classification is an important yet challenging task in graph learning, where nodes exhibit multiple semantics simultaneously. Existing methods for multi-label node classification can effectively model multiple labels, while only considering in-domain scenarios where the model needs to be trained and tested within the same graph domain, resulting in limited cross-domain generalization. Recently, Graph Foundation Models (GFMs) have emerged as a promising paradigm for learning transferable graph representations across diverse graph domains and downstream tasks. However, existing GFMs are built upon single-label assumption, where all nodes are arbitrarily regarded as containing only one class of semantic and embedded into a single representation. For multi-label nodes, such a representation essentially approximates multiple semantics with a single point in the representation space, inevitably leading to semantic entanglement and making simultaneous discrimination of multiple labels difficult. To address these limitations, we propose a Multi-Semantic Basis Graph Foundation Model (MSB-GFM), a framework for cross-domain multi-label node classification. Specifically, we introduce a multi-semantic basis representation learning paradigm that models each multi-label node as an adaptive composition of semantic bases, thereby enabling flexible representational capacity for modeling multiple semantics. Furthermore, we develop a semantic-structure dual-channel architecture with domain adversarial training for effective cross-domain knowledge transfer. Extensive experiments demonstrate the effectiveness of our model.
EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs
Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamicall…
ルーティングの重みを超えて: 貢献度のコントラストによる専門家混合報酬モデルの忠実な応答レベルの解釈
報酬モデルは人間の好みから学習する上で中心的な役割を果たしますが、何が人間の予測を促すのかを特定することは依然として困難です。最近のまばらな専門家混合 (MoE) 報酬モデルは、プロンプトを専門の専門家にルーティングし、ルーティングの重みが高い例を通じて専門家を特徴付けることによって、解釈可能性の向上を目指しています。ただし、ルーティングの重みは、専門家が $\textit{受信}$ するプロンプトを明らかにするだけで、$\textit{判断}$ する方法を明らかにするものではなく、専門家の行動の部分的な説明しか提供しません。したがって、我々は $\textbf{Co}$ntribution-$\textbf{Co}$ntrast ($\textbf{CoCo}$) 応答レベルの解釈を提案します。これは、最大の寄与コントラストを持つ選択応答と拒否応答のペアを使用して専門家の役割を忠実に特徴付け、ルーティングと優先行動を共同で捕捉します。 CoCo は、自動評価と人間による評価の両方において、ルーターベース、スコアベース、スパースオートエンコーダーベースの代替手段よりも、より一貫性があり、忠実で特殊な解釈をもたらし、同時に競争力のある報酬モデリングの精度を維持します。私たちの知る限り、これは教育省の報酬モデルの解釈方法に関する最初の体系的な研究です。
原文 (English)
Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast
Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent sparse Mixture-of-Experts (MoE) reward models seek to improve interpretability by routing prompts to specialized experts and characterizing experts through examples with high routing weights. However, routing weights only reveal which prompts an expert $\textit{receives}$, not how it $\textit{judges}$ responses, providing only a partial account of expert behavior. We therefore propose $\textbf{Co}$ntribution-$\textbf{Co}$ntrast ($\textbf{CoCo}$) response-level interpretation, which faithfully characterizes experts' roles using chosen-rejected response pairs with the largest contribution contrasts, jointly capturing routing and preference behavior. Across automatic and human evaluations, CoCo yields more coherent, faithful, and specialized interpretations than router-based, score-based, and sparse autoencoder-based alternatives while maintaining competitive reward modeling accuracy. To the best of our knowledge, this is the first systematic study of interpretation methods for MoE reward models.
LLM 記号化された構造化プロセスによる解釈可能な教師なしコミュニティ検出
コミュニティ検出は、同様の行動や関心を持つエンティティの凝集したグループを特定することを目的としたグラフ分析の基本的なタスクです。古典的な目的主導の手法は複雑なグラフ構造に苦労しますが、ディープラーニングのアプローチは解釈可能性を犠牲にしてパフォーマンスを向上させ、ラベル付きのデータとトレーニングに依存します。大規模言語モデル (LLM) は、強力な推論機能と世界の知識を備えており、解釈可能でラベルフリーのコミュニティ検出に有望です。これらの強みを活用するために、LLM ガイド付き、解釈可能、トレーニング不要、教師なしのコミュニティ検出方法である LUCID を提案します。 LUCID は、初期化、マージ、洗練、選択を通じて複雑な構造が現れる自然系の相転移動力学にインスピレーションを得て、4 段階のパイプラインとして設計されています。このパイプライン内で、LLM は暗黙の知識を明示的で解釈可能な論理構造に変換する形式的なルールを誘導します。具体的には、(1) Local-View Community Initialization ステージでは、k-ego コンテキストと教師なしノードの役割を使用してローカル グラフ構造をエンコードします。 (2) 多要素コミュニティのマージ ステージでは、LLM によって誘導されたルールを使用して、ローカル コミュニティを繰り返しマージします。 (3) マルチグレイン コミュニティ リファインメント ステージでは、境界ノイズを低減するために、LLM による粗いルールから細かいルールへのルールを並行して適用します。 (4) グローバル ビューのコミュニティ選択ステージでは、トポロジカルなコンパクトさと境界の明確さに基づいて高品質のコミュニティを特定します。現実世界のデータセットに対する広範な実験により、教師なしアプローチとしての LUCID が最先端のパフォーマンスを達成し、主要な教師なしおよび半教師ありのベースラインを常に上回るパフォーマンスを示すことが実証されました。
原文 (English)
Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes
Community detection is a fundamental task in graph analytics that aims to identify cohesive groups of entities with similar behaviors or interests. Classic objective-driven methods struggle with complex graph structures, while deep-learning approaches improve performance at the expense of interpretability and rely on labeled data and training. Large language models (LLMs), with strong reasoning capabilities and world knowledge, are promising for interpretable, label-free community detection. To leverage these strengths, we propose LUCID, an LLM-guided, interpretable, training-free, and unsupervised community detection method. Inspired by phase-transition kinetics in natural systems, where complex structures emerge through initialization, merging, refinement, and selection, LUCID is designed as a four-stage pipeline. Within this pipeline, the LLM induces formal rules that translate implicit knowledge into explicit and interpretable logical structures. Specifically, (1) the Local-View Community Initialization stage encodes local graph structures using k-ego contexts and unsupervised node roles; (2) the Multi-factor Community Merge stage uses LLM-induced rules to iteratively merge local communities; (3) the Multi-grain Community Refinement stage applies LLM-induced coarse-to-fine rules in parallel to reduce boundary noise; and (4) the Global-view Community Selection stage identifies high-quality communities based on topological compactness and boundary clarity. Extensive experiments on real-world datasets demonstrate that LUCID, as an unsupervised approach, achieves state-of-the-art performance and consistently outperforms leading unsupervised and semi-supervised baselines.
ADIAS: インタラクティブなエージェント システムの自動設計
自動化されたエージェント設計は、反復的な修正、評価、フィードバックの要約を通じてエージェントのハーネスを改善します。既存の方法は主に候補者中心です。クロスラウンドのエクスペリエンスは候補者エージェントを中心に編成され、修復の進行状況が暗黙的に残ります。これにより、修復のターゲットが非効率になり、部分的な進行の統合が遅くなり、ラウンド全体に非効果的な介入が伝播することになります。したがって、問題中心のエージェント最適化を定式化します。この最適化では、修復の進行状況が各ラウンドの候補履歴から再導出されるのではなく、最適化をガイドする明示的な永続的な問題状態として引き継がれます。 2 つのメカニズムを備えた自動フルコード エージェント設計のフレームワークである ADIAS で定式化をインスタンス化します。永続的な問題の状態では、安定した問題の ID、ライフサイクルのステータス、裏付けとなる証拠、介入結果の履歴が維持されます。問題に基づく最適化では、この状態を使用して、その後の集中的なコード全体の変更のための修復ターゲットとリビジョンの方向を共同で提案します。 5 つのインタラクティブなベンチマーク全体で、ADIAS は最も強力なベースラインを平均 25.2% 上回っており、4 つのバックボーン モデル全体で一貫した利益を達成しています。さらに、制御されたアブレーションにより、永続的な問題の状態を除去したり、問題中心の改訂を候補者中心の政策に置き換えたりすると、パフォーマンスが最大 40.7% 低下することが示されています。
原文 (English)
ADIAS: Automated Design of Interactive Agentic Systems
Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves the repair progress implicit. This causes inefficient repair targeting, slow consolidation of partial progress, and propagation of ineffective interventions across rounds. Therefore, we formulate issue-centric agent optimization, in which repair progress is carried forward as an explicit persistent issue state to guide optimization, rather than re-derived from candidate history in each round. We instantiate the formulation in ADIAS, a framework for automated full-code agent design with two mechanisms. A persistent issue state maintains stable issue identities, lifecycle status, supporting evidence, and intervention-outcome histories. Issue-guided optimization uses this state to jointly propose repair targets and revision directions for subsequent focused full-code modification. Across five interactive benchmarks, ADIAS outperforms the strongest baseline by 25.2% on average and achieves consistent gains across four backbone models. Controlled ablations further show that removing persistent issue state or replacing issue-centric revision with candidate-centric policies leads to performance drops of up to 40.7%.
ビジュアルトークンプルーニンのためのMLLMにおける中間層の注意を予測する方法を学ぶ
マルチモーダル大規模言語モデル (MLLM) は、さまざまな視覚言語タスクにわたって優れたパフォーマンスを実現しますが、その効率は多数の視覚トークンを処理するコストによって制限されます。視覚的なトークン プルーニングによりこのコストを削減できますが、トークンの重要性を正確に見積もる必要があります。最近の研究では、中間言語モデル層からのテキストからビジョンへのアテンションが視覚的なトークンの枝刈りを効果的にガイドできることが実証されており、通常は事前定義された中間層からのアテンションを使用して、保持するビジュアル トークンを選択します。したがって、2 つの問題が残されています。まず、私たちの分析では、質問に最も注意を向ける層がサンプルごとに大幅に異なり、固定層が最適ではないことがわかりました。第 2 に、適切な中間層から注意を引くには、複数の言語モデル層を介して多数のビジュアル トークンを処理する必要があり、その時点ですでにかなりの計算が費やされています。両方の問題に対処するために、中間層注意予測 (MAP) を提案します。これは、質問対比教師選択を使用して、元の質問と参照質問の下で注意を対比させることによってサンプル固有の教師層を識別し、選択した層から注意を蒸留して、マルチモーダル入力特徴から視覚トークンの重要性を推定する軽量の予測子を作成します。推論中に、MAP は予測された重要度スコアを多様性基準と組み合わせて、最初の言語モデル層の前にビジュアル トークンを削減します。したがって、MAP は枝刈りのためのアテンション マップを必要とせず、既存の推論高速化技術との互換性を維持します。 LLaVA-NeXT-7B の 10 のベンチマーク全体で、MAP は枝刈りされていないモデルのパフォーマンスの 97.5% を保持し、ビジュアル トークンは 5.56% のみで、エンドツーエンドの 3.09 倍の高速化が実現しました。
原文 (English)
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin
Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce this cost, but requires accurate token importance estimates. Recent studies have demonstrated that text-to-vision attention from middle language model layers can effectively guide visual token pruning, typically using attention from a predefined middle layer to select the visual tokens to retain. Two problems therefore remain. First, our analysis shows that the layer whose attention is most responsive to the question varies substantially across samples, making a fixed layer suboptimal. Second, obtaining attention from the appropriate middle layer requires processing numerous visual tokens through several language model layers, by which point considerable computation has already been spent. To address both problems, we propose Middle-layer Attention Prediction (MAP), which uses Question Contrastive Teacher Selection to identify a sample-specific teacher layer by contrasting attention under the original and reference questions, and distills attention from the selected layer into a lightweight predictor that estimates visual token importance from multi-modal input features. During inference, MAP combines the predicted importance scores with a diversity criterion to prune visual tokens before the first language model layer. Thus, MAP requires no attention maps for pruning and remains compatible with existing inference acceleration techniques. Across ten benchmarks on LLaVA-NeXT-7B, MAP retains 97.5% of the unpruned model performance with only 5.56% of the visual tokens, yielding a 3.09x end-to-end speedup.
WebGrader: 自己進化型プログラマティック グレーダーを使用した Web 開発のための LLM のトレーニング
大規模な言語モデルでは、自然言語の記述から完全な Web サイトを生成するケースが増えており、強化学習は、言語モデルに残っている機能上のギャップを埋めるための中心的なアプローチとなっています。このトレーニング計画には、報酬の設計がボトルネックになっています。手動で作成されたブラウザ スクリプトは実行可能ですが、制限のない要件に合わせて記述するにはコストがかかります。一方、VLM および GUI エージェントのグレーダーは拡張性はありますが、決定的な状態を観察する前に判定を下す可能性があります。私たちは、各 Web サイト リクエストから必要なインタラクション フローを自律的に導き出し、各フローを実行可能なフロー コントラクトとして表し、その実行結果を RL 報酬として使用する、自己進化するプログラムによるグレーダーである WebGrader を提案します。 WebGrader は、生成されたプロジェクトをライブ ブラウザーで実体化し、ソース コードとライブ DOM に対してターゲット アクションを根拠付け、同じブラウザーの軌跡に沿って視覚的証拠、DOM 証拠、応答証拠、および永続状態の証拠を収集します。次に、残差駆動のオフライン ループによって、再利用可能な検証スキルが検出され、それらが切り離された検証ページでスクリーニングされ、ポリシー トレーニングの前に昇格されたスキル グラフがフリーズされます。テスト計画、アクションの基礎付け、証拠収集、セマンティック判断を分離することにより、WebGrader は要求された遷移を観察した後でのみ合格判定を発行します。 WebGen-Bench では、WebGrader は 8B ポリシーを 52.01% の機能成功率までトレーニングし、一致する外観とスクリプトの報酬を 7.88 ポイント上回り、o4-mini および DeepSeek-v4-flash を上回ります。 WG-core-250 では、ポリシーはフル スコア 44.953 に達し、Qwen3-Coder-480B を上回ります。
原文 (English)
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader
Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remaining functional gap. This training regime is bottlenecked by reward design. Hand-authored browser scripts are executable yet costly to write for open-ended requirements, while VLM and GUI-agent graders scale but may issue verdicts before observing the decisive state. We propose WebGrader, a self-evolving programmatic grader that autonomously derives the required interaction flows from each website request, represents each flow as an executable Flow Contract, and uses its execution outcome as an RL reward. WebGrader materializes the generated project in a live browser, grounds target actions against the source code and live DOM, and collects visual, DOM, response, and persistent-state evidence along the same browser trajectory. A residual-driven offline loop then discovers reusable verifier skills, screens them on disjoint validation pages, and freezes the promoted skill graph before policy training. By separating test planning, action grounding, evidence collection, and semantic judgment, WebGrader issues a Pass verdict only after observing the requested transition. On WebGen-Bench, WebGrader trains an 8B policy to a 52.01% functional success rate, outperforming a matched appearance-plus-script reward by 7.88 points and surpassing o4-mini and DeepSeek-v4-flash. On WG-core-250, the policy reaches a Full Score of 44.953 and surpasses Qwen3-Coder-480B.
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate be…
KNOWPLAN: スマートな学位取得経路計画のための知識主導型 AI エージェント
大学の公式情報源から学位を計画するには、2 つの問題を順番に解決する必要があります。教育機関のカリキュラムは、まずカタログ、部門別ページ、JSON エンドポイント、スキーマを共有しない PDF から再構築する必要があり、その後でのみ、前提条件ロジックと重複する要件制約の下で学生固有のパスを最適化できます。この 2 つを結合すると、それぞれの障害モードがもう一方を隠すことができます。これは、独自のクロールを実行するプランナーが、現在の計画に必要のない事実を学習することがないためです。我々は、抽出優先境界を強制し、ステージ間のインターフェースを想定するのではなく測定する KnowPlan を紹介します。 CatalogBrowse は、ユーザー プロファイルにアクセスせずに探索します。ソースアクセス単位あたりのアトミックカタログ義務の有限セットに対する信頼性の低い予想限界利得によって訴訟をスコアリングし、スパン制約付き句からASTモデルへのフォールバックを備えたプラットフォームアダプターを通じて決定論的に解析し、報酬しきい値の代わりにインデックス、スキーマ、来歴、および参照完全性に対するクロージャ証明書で終了します。その出力コントラクトは、出所にリンクされた 3 つの JSON ドキュメントです。 DegreeMap はそれらのドキュメントのみを使用します。それらを型付きの要件ハイパーグラフにコンパイルし、ハード実現可能性、完了期間、負荷とリスク、パーソナライズされたユーティリティ、およびオプションの値にわたって CP-SAT を使用して辞書編集的に最適化するため、各ステージは前のステージで証明された最適化の範囲内で最適化され、ソルバーのバジェット内で認定可能になります。 100 大学の広域トラックと 6 校の密集トラック全体で、CatalogBrowse は、網羅的クローラーよりもソース アクセスが 47% 少ないにもかかわらず、在庫回収率 96.2% とマスクされたソース回復率 88.7% に達します。DegreeMap は、パーソナライズされたユーティリティを最強のベースラインより +0.066 改善しながら、100.0% のハード実現可能性を保持します。また、完全なパイプラインは、ユーティリティ ギャップのあるリクエストの 99.5% を認証します。 0.015の特権的なゴールドグラフ。
原文 (English)
KNOWPLAN: Knowledge-Driven AI Agents for Smart Degree Pathway Planning
Planning a degree from official university sources requires solving two problems in order. The institution's curriculum must first be reconstructed from catalogs, departmental pages, JSON endpoints, and PDFs that share no schema, and only then can a student-specific path be optimized under prerequisite logic and overlapping requirement constraints. Coupling the two lets each failure mode hide the other, because a planner that drives its own crawling never learns facts its current plan does not need. We present KnowPlan, which enforces an extraction-first boundary and measures the interface between the stages rather than assuming it. CatalogBrowse explores with no access to any user profile. It scores legal actions by lower-confidence expected marginal gain over a finite set of atomic catalog obligations per unit of source access, parses deterministically through platform adapters with a span-constrained clause-to-AST model fallback, and terminates on a closure certificate over index, schema, provenance, and reference completeness instead of a reward threshold. Its output contract is three provenance-linked JSON documents. DegreeMap consumes only those documents. It compiles them into a typed requirement hypergraph and optimizes lexicographically with CP-SAT over hard feasibility, completion horizon, load and risk, personalized utility, and option value, so that each stage optimizes inside the previous stage's proven optimum and stays certifiable within the solver budget. Across a 100-university broad track and a six-school dense track, CatalogBrowse reaches 96.2% inventory recall and 88.7% masked-source recovery at 47% less source access than an exhaustive crawler, DegreeMap holds 100.0% hard feasibility while improving personalized utility by +0.066 over the strongest baseline, and the full pipeline certifies 99.5% of requests with a utility gap to the privileged gold graph of 0.015.
TaskSense: Focusing on What Matters in World Models
World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representation…
ステアリング圧力下でのフロンティア言語モデルの発散応答モード
フロンティア言語モデルは、個別のデータ、目標、安全パイプラインを使用してトレーニングされます。これらの違いが、明示的なステアリング圧力下で測定可能なほど異なる動作を生み出すかどうかは、まだ解明されていません。この研究では、価値観の衝突、推論の引き出し、推論の抑制という 3 つのカテゴリー (さらに 40 の検証項目) にわたる 300 個のベース項目とステアリング項目のペアを使用して、6 人の開発者による 6 つのフロンティア モデルにわたる行動のステアビリティを評価しています。 6 つのモデルはすべてブラインドピアジャッジとして機能し、固定された行動ルーブリックに基づいてすべての反応を分類します。結果として得られる 24,480 件の判定は、リーブ ワン アウトのコンセンサスによって採点されます。モデルによってステアリングの挙動がどの程度変化するかだけでなく、どのような種類 (モード) の応答が得られるかが異なり、一部の応答モードはそのうちの 1 つまたは 2 つにしか現れないことがわかりました。 GPT-5 は、答えをそのままにして、その推論を開示する要求をかわします (99% 対他のすべてのモデルでは 0%)。クロード オーパス 4.7 と GPT-5 は、明示的な抑制命令にさまざまな方法で抵抗します。 Llama をオープンウェイト モデルとして使用し、最大の動作分割をその内部まで追跡します。線形プローブは、残留ストリームからの動作を 0.87 ホールドアウト精度でデコードし、生成中にその方向を注入することで、介入スイープ全体にわたって動作を 0% から 86% まで駆動します。すべての発見は、トークン予算による修復と、仮説に盲目的な判断プロンプトを使用した対照実験の両方で当てはまります。
原文 (English)
Divergent Response Modes in Frontier Language Models Under Steering Pressure
Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study evaluates behavioral steerability across six frontier models from six developers using 300 paired base and steered items over three categories: values-conflict, reasoning-elicitation, and reasoning-suppression (plus 40 validation items). All six models act as blind peer judges and classify every response based on fixed behavioral rubrics. The resulting 24,480 judgments are scored by leave-one-out consensus. We find that models differ not just in how much steering shifts their behavior but in what kind (mode) of response they give, and some response modes appear in only one or two of them. GPT-5 deflects requests to disclose its reasoning while leaving its answer intact (99% vs. 0% for all other models). Claude Opus 4.7 and GPT-5 resist explicit suppression instructions and in different ways. Using Llama as the open-weight model, we trace the largest behavioral split to its internals. A linear probe decodes the behavior from the residual stream at 0.87 held-out accuracy while injecting that direction during generation drives the behavior from 0% to 86% across an intervention sweep. Every finding holds under both a token-budget remediation and a control experiment with a hypothesis-blind judgment prompt.
自動アイテム評価: LLM によって生成された批評を使用してアイテムの受け入れと拒否を予測します。
自動品目評価 (AIE) とは、評価対象品目の専門家による手動レビューやフィールド テストを必要とせずに品目の品質を評価するための計算手法の使用を指します。私たちは、大規模な標準化されたテスト プログラムからの過去の不合格データを使用して、品目テキストから品目の合格と不合格を予測することにより、ほぼ包括的な AIE モデルを構築することを目指しました。データセットには 52,759 件の英語芸術 (ELA) と数学の項目が含まれており、そのうち 34% は将来の運用上の使用から永久に拒否されました。拒否の理由には、不十分な心理測定特性、コンテンツの問題、偏見と感受性への懸念、およびコンテンツ以外の問題が含まれていました。私たちは、生のアイテム テキストに対する DeBERTaV3 ラージ分類器、Qwen3 で生成されたアイテム批評に対する 2 番目の DeBERTa 分類器、および両方の表現を組み合わせた融合モデルを微調整しました。融合モデルは、最も強力な全体的なパフォーマンスを達成しました (精度 = 0.75、F1 = 0.64、AUC = 0.80、感度 = 0.64、特異性 = 0.81)。数学的予測 (F1 = .73、AUC = .86) は、ELA (F1 = .51、AUC = .72) よりもかなり正確でした。判定しきい値を 0.5 から 0.25 に下げると、ELA と数学の平均感度は 0.88 と 0.91 に上昇しましたが、特異度はそれぞれ 0.31 と 0.56 に低下しました。これは、アイテムを評価するよりも生成する方が安価である自動アイテム生成のコンテキストでは好ましいと考えられます。アイテムの生のテキストと一緒にアイテムの批評を組み込むことで、ほとんどの拒否理由でパフォーマンスが向上しました。このモデルでは、より困難な項目ほど高い拒否確率が割り当てられました。ただし、融合モデルは、特に ELA に関して、バイアス、機密性、公平性、またはアクセシビリティについてフラグが立てられた項目を特定するのに苦労しました。これらの調査結果は、テキストベースの AIE が一部の分野では実現可能であり、手動レビューやフィールドテストの負担を軽減する実用的なツールとなる可能性があることを示唆するとともに、公平性の懸念がある項目については人間によるレビューの重要性も強調しています。
原文 (English)
Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques
Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item acceptance and rejection from item text using historical rejection data from a large-scale standardized testing program. The dataset contained 52,759 English language arts (ELA) and mathematics items with 34% permanently rejected from future operational use. Rejection reasons included poor psychometric properties, content issues, bias and sensitivity concerns, and non-content issues. We fine-tuned a DeBERTaV3-large classifier on raw item text, a second DeBERTa classifier on Qwen3-generated item critiques, and a fusion model combining representations from both. The fusion model achieved the strongest overall performance (Accuracy = .75, F1 = .64, AUC = .80, Sensitivity = .64, Specificity = .81). Prediction for math (F1 = .73, AUC = .86) was considerably more accurate than ELA (F1 = .51, AUC = .72). Lowering the decision threshold from .5 to .25 raised average sensitivity for ELA and math to .88 and .91, while reducing specificity to .31 and .56, respectively, which may be preferable in automated item generation contexts where generating items is cheaper than evaluating them. Incorporating item critiques alongside raw item text improved performance across most rejection reasons. The model assigned higher rejection probabilities to more difficult items. However, the fusion model struggled to identify items flagged for bias, sensitivity, fairness, or accessibility, especially for ELA. These findings suggest that text-based AIE is feasible in some areas and may offer a practical tool for reducing the burden of manual review and field testing, while also underscoring the importance of human review for items with fairness concerns.
NxN E-valuation: コンフォーマル CRT ヌルによる仮説証明
私たちは、十分な大きさのデータセットが利用できる限り、専用の帰無仮説の構築など、ケース固有の証明手順を構築することなく仮説を検証できる、便利な電子値ベースの仮説証明アルゴリズムである NxN E-valuation を提案します。この方法は、LLM ベースの探索システムに特に適しています。LLM は仮説を提案するのが非常に得意ですが、幻覚にひどく悩まされます。この幻覚により、LLM 出力を直接収集することができなくなり、既存の治療法はいずれも不十分です。最も一般的な解決策には、導入部で詳述したその他の救済策の中でも特に、循環検証とホールドアウト テスト (誤った仮説が依然として偽の相関を介して通過する可能性がある) を LLM に検証または修正させることが含まれます。これを解決するために、NxN E-valuation は自然に存在する大規模なトレーニング セットを利用し、異なるサンプルを相互に帰無仮説として機能させます。この設計は、各仮説を証明する条件付きランダム化テスト (CRT) を直接実現します。このアプローチは、LLM の世代が個々のサンプルに適用される仮説である場合、少なくとも LLM 循環検証とホールドアウト データ テストに代わる一般的により優れた代替手段となる可能性があります。
原文 (English)
NxN E-valuation: Hypothesis Certification via a Conformal CRT Null
We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a large enough dataset is available. The method is especially suited to LLM-based exploration systems, where LLMs are remarkably good at proposing hypotheses but suffer badly from hallucination; this hallucination prevents us from harvesting LLM outputs directly, and existing remedies each fall short. The most common solutions include letting the LLM verify or correct itself circular verification and held-out testing (where false hypotheses can still pass via spurious correlations), among other remedies detailed in the introduction. To resolve this, NxN E-valuation exploits the naturally existing large training set and lets different samples serve as null hypotheses for one another. This design directly realizes a conditional randomization test (CRT) that certifies each hypothesis. The approach can be a universally better replacement for at least LLM circular verification and held-out-data testing, provided the LLM's generations are hypotheses that apply to each individual sample.
フィードの形成: 会話型レコメンデーションのための LLM ベースのエージェント システム
産業用レコメンデーション システムは主に、明示的な自然言語入力ではなく、暗黙的な行動シグナル (クリック数や滞留時間など) からユーザーの好みを推測する受動的なランキング パラダイムを採用しています。その結果、ユーザーは自分の明示的な興味と受動的な行動アルゴリズムが提供するものとの間の一貫した不一致を経験し、微妙な好みを表現したり、リアルタイムでフィードを操作したりする能力が制限されます。レコメンデーションの最適化方法とユーザーが自分の興味を明確に表現したい方法との間のこの拡大するギャップに対処するために、リアルタイムのマルチモーダルなコンテンツの共同キュレーションを可能にする LLM ベースのエージェントレコメンデーション フレームワークである Shape Your Feed (SYF) を紹介します。 SYF は 3 層アーキテクチャを採用しています。(i) テキスト プロンプト、音声コマンド、UI インタラクションからきめ細かいユーザーの意図を捕捉するパーセプション フロー。 (ii) 進化するユーザーの好みをエンコードする永続的なセマンティック プロファイルに基づいて、リアルタイムのエージェントによる再ランキングと候補アイテムのプルーニングを実行するサービス フロー。 (iii) Direct Preference Optimization (DPO) および LLM-as-a-Judge アンサンブルを介してシステムの動作を人間の判断と調整する自己進化フロー。オフライン評価では、SYF のアライメント スコアリング モジュールが 98.85% の精度を達成し、強力な数ショット ベースラインよりも大幅に向上していることが示されています。実稼働トラフィックに関する大規模なオンライン A/B 実験では、SYF がフィードの関連性とユーザーのセンチメントを改善することがさらに実証され、産業環境におけるインタラクティブでユーザーが操作可能なレコメンデーションに向けた実用的でスケーラブルな道筋が示されています。
原文 (English)
Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation
Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., clicks, dwell time) rather than explicit, natural language inputs. As a result, users experience a persistent discrepancy between their explicit interests and what passive behavioral algorithms deliver, limiting their ability to express nuanced preferences or steer their feed in real time. To address this growing gap between how recommendations are optimized and how users wish to articulate their interests, we present Shape Your Feed (SYF), an LLM-based agentic recommendation framework that enables real-time, multimodal co-curation of content. SYF employs a three-tier architecture: (i) a Perception Flow that captures fine-grained user intent from text prompts, voice commands, and UI interactions; (ii) a Serving Flow that performs real-time agentic re-ranking and pruning of candidate items, grounded in a persistent Semantic Profile encoding evolving user preferences; and (iii) a Self-Evolution Flow that aligns system behavior with human judgments via Direct Preference Optimization (DPO) and an LLM-as-a-Judge ensemble. Offline evaluations show that SYF's alignment scoring module achieves 98.85% accuracy, substantially improving over strong few-shot baselines. Large-scale online A/B experiments on production traffic further demonstrate that SYF improves feed relevance and user sentiment, indicating a practical and scalable path toward interactive, user-steerable recommendation in industrial settings.
TRACE: ドリフトと障害下での人間の AI コントローラー調整のための多層ベンチマーク
最新のサイバーフィジカル システムおよび AI 支援システムは、人間のオペレーター、AI 意思決定モジュール、および自動コントローラーを単一の制御ループ内で結合しているため、信頼性は 1 つのモデルではなくループ全体に依存します。しかし、標準的なベンチマークでは、ドリフトや障害がこれらの層にどのように伝播するかを時間的に調整された複数層のトレースをキャプチャすることはできないため、調整が失敗する場所、その理由、および回復方法を診断することはできません。この論文は、そのギャップの 1 つの側面、つまりスタック層で発生する可能性のある偏差であるドリフトと、従来の単一モダリティ監視では層に局所化することや、開始時間にピン留めできないことを対象としています。日常の家事用の接地された命令ベンチマークである ALFRED から派生したトレースに制御されたドリフトを注入することによってベンチマークを構築し、1,918 個のドリフト トレースを生成します。各トレースは、5 つの実行層 (状態、観察、決定、ルール、制御) にわたるステップごとのレコードの時間的に整列されたシーケンスであり、ドリフト タイプ、影響を受ける層、開始時刻、責任のあるアクター、および因果関係のメカニズムでラベル付けされ、アノテーター間の合意が報告された独立した評価者によって検証されます。このデータセットを、ほぼ完璧な開始リークを除去するリーク対応プロトコルと、古典的、再発的、注意ベースのモデル ファミリにわたるベースライン研究と組み合わせます。この誠実なプロトコルの下では、ドリフトは特定可能であり、すべてのファミリーにわたってランダムおよび多数派ベースライン (影響を受ける層のマクロ F1 が 0.70 近く、責任主体が 0.85 近く、因果メカニズムが 0.49 近く) をはるかに上回っており、この象徴的なベンチマークでは、集中的な注意は単純なモデルよりも利点がありません。
原文 (English)
TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure
Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustworthiness depends on the whole loop, not any one model. Yet no standard benchmark captures time-aligned, multi-layer traces of how drift and failures propagate across these layers, so we cannot diagnose where coordination breaks down, why, or how to recover. This paper targets one facet of that gap: drift, a deviation that can originate in any stack layer and that conventional single-modality monitoring cannot localize to a layer or pin to an onset time. We construct a benchmark by injecting controlled drift into traces derived from ALFRED, a grounded-instruction benchmark for everyday household tasks, yielding 1,918 drifted traces. Each trace is a time-aligned sequence of per-step records across five execution layers (state, observation, decision, rules, control), labeled with the drift type, affected layer, onset time, responsible actor, and causal mechanism, and validated by independent raters with inter-annotator agreement reported. We pair the dataset with a leak-aware protocol that removes a near-perfect onset leak, and a baseline study across classical, recurrent, and attention-based model families. Under this honest protocol, drift is identifiable and attributable well above random and majority baselines across every family (affected layer macro-F1 near 0.70, responsible actor near 0.85, causal mechanism near 0.49), and heavy attention offers no advantage over simpler models on this symbolic benchmark.
CellWorld: 空間トランスクリプトミクス基盤モデルにおける遺伝子レベルの再構成から潜伏細胞予測まで
この論文は、潜在空間予測事前トレーニングが空間トランスクリプトミクスの基礎モデルへのスケーラブルなルートを提供できることを示しています。既存の空間トランスクリプトミクス基礎モデルは主に、マスクされた遺伝子のアイデンティティまたは発現値を再構築するため、アッセイ固有の技術的バリエーションの再現を促進し、表現の伝達性を制限する可能性があります。このような変動を直接再構築することを避けるために、予測対象を観察された遺伝子測定から潜在細胞表現に移し、目に見える空間的コンテキストと限定された部分発現ヒントからマスクされた細胞の潜在表現を予測する CellWorld を導入します。 4,600 万個のヒト細胞のコーパス上で、574 万から 9,456 万のトレーニング可能なパラメータにわたる 4 つの CellWorld バリアントを事前トレーニングしました。私たちの制御されたスケーリング実験では、モデルの能力、特に空間タスクのパフォーマンスが向上する一方で、空間転送は細胞数のみよりも十分な最適化と幅広い生物源の多様性に大きく依存することが示されています。 574 万個のトレーニング可能なパラメーターを備えた 4 つの保持されたデータセット (CellWorld-Small でさえ) 全体で、11 個の線形プローブ ベンチマークすべてと 7 つすべての微調整された空間ベンチマークですべてのベースラインを上回っています。最も注目すべき点は、広範囲の生物学的ソースをカバーし、コーパスのわずか 5% で事前トレーニングされた凍結された CellWorld-Large が、7 つの空間ベンチマークすべてにわたって完全に微調整されたすべてのベースラインを上回っていることです。コードは https://github.com/UoM-HealthAI/CellWorld で入手できます。
原文 (English)
CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models
This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific technical variation and limiting representation transferability. To avoid directly reconstructing such variation, we shift the prediction target from observed gene measurements to latent cell representations and introduce CellWorld, which predicts the latent representations of masked cells from visible spatial context and a limited partial-expression hint. We pretrain four CellWorld variants, spanning 5.74M to 94.56M trainable parameters, on a corpus of 46 million human cells. Our controlled scaling experiments show that performance improves with model capacity, particularly on spatial tasks, while spatial transfer depends more on sufficient optimization and broad biological source diversity than on cell count alone. Across four held-out datasets, even CellWorld-Small, with 5.74M trainable parameters, outperforms every baseline on all 11 linear-probe benchmarks and all seven fine-tuned spatial benchmarks. Most notably, a frozen CellWorld-Large pretrained on only 5\% of the corpus with broad biological source coverage outperforms every fully fine-tuned baseline across all seven spatial benchmarks. Code is available at https://github.com/UoM-HealthAI/CellWorld.
深層強化学習を使用した配車問題 - 業界におけるトラック計画のケーススタディ
サプライチェーン業界の重要な要素である輸送は、デジタル プラットフォームとインテリジェント アルゴリズムの支援により、過去 10 年間に急速な発展を遂げてきました。交通研究の分野では、車両経路指定問題 (VRP) が依然として永続的な課題となっています。経営科学の分野では、産業界と学術部門の専門家や学者が、古典的な巡回セールスマン問題からより一般的な車両経路問題に至るまで、経路経路の問題に効果的に対処するための最適化モデルとアルゴリズムを継続的に研究してきました。これらのモデルとアルゴリズムは実際の産業シナリオに適用され、コストの最適化を達成し、二酸化炭素排出量を削減します。しかし、現実世界の問題は複雑であるため、多くの特定の制約が追加されることが多く、情報の不透明性、不確実性、人間の不合理な行動などの課題が発生する可能性があります。したがって、最適な結果を維持しながら実際のシナリオで VRP の数学モデルを展開して最適化するには、多くの課題が生じます。このペーパーでは、外部トラック ネットワーク設計に関わる 3 つの異なる物流ユース ケースについて説明し、ソリューションを提供します。この論文では、これらの産業事例を通じて、深層強化学習ベースの車両ルート最適化がどのように実装されているかを紹介します。その結果、強化学習エージェントによって最適化されたルートは、ベースライン結果と比較して総コストが 10% を超えていることがわかります。さらに、この論文は、将来の研究において、車両ルート問題に対する DRL アルゴリズムが VRP のより多くのバリエーションに一般化される可能性があると提案しています。
原文 (English)
Vehicle routing problem using deep reinforcement learning - A case study about truck planning in the industry
As an important component of the supply chain industry, transportation has experienced rapid development in the past decade with the assistance of digital platforms and intelligent algorithms. Within the field of transportation research, Vehicle Routing Problem (VRP) has remained a persistent and enduring challenge. In the realm of management science, experts, and scholars from both the industrial and academic sectors have continuously explored optimization models and algorithms to effectively address routing problems, from the classical Traveling Salesman Problem to the more general Vehicle Routing Problem. These models and algorithms are applied in real-world industrial scenarios to achieve cost optimization and reduce carbon footprints. However, due to the complexity of real-world problems, numerous specific constraints are often added, and challenges such as information opacity, uncertainty, and irrational human behavior may arise. Therefore, deploying and optimizing mathematical models for VRP in practical scenarios while maintaining optimal results poses numerous challenges. This paper discusses and provides solutions for three different logistic use cases involving external truck network design. Through these industrial case study, the paper introduces how deep reinforcement learning-based vehicle routing optimization has been implemented. As a result, it can be observed that the routes optimized by reinforcement learning agent have over 10% total cost compared to baseline results. Furthermore, the paper proposes that in future research, DRL algorithms for vehicle routing problems could be generalized into more variations of VRP.
ポリマーの粗視化分子動力学を自動化するためのマルチエージェント フレームワーク
粗視化 (CG) 分子動力学は、全原子 (AA) 法で利用できるスケールを超えてポリマー シミュレーションを拡張しますが、ボトムアップ CG モデリングは手間がかかります。 CG 解像度は設計上の選択であるため、一般に転送可能なパラメータ セットは利用できず、ポテンシャルはポリマー マッピングごとに新たに導出されます。ここでは、トポロジーの構築、平衡化、マッピング、潜在的な導出、およびポリマーとターゲット解像度の自然言語仕様からの検証を自動化するマルチエージェント フレームワークである CGMas を紹介します。大規模言語モデル (LLM) 推論エージェントはポリマー名から AA トポロジーを推測し、層状の自己修正により不飽和、ヘテロ原子含有、極性ポリマーに共通する物理エラーを解決します。下流のエージェントはシステムを平衡化し、CG 表現にマッピングし、ボルツマン反転を通じてポテンシャルを導出し、その原子的参照に対してモデルをベンチマークします。 CGMA は 27 個のホモポリマーとコポリマーのタスクをすべて完了し、22 個の AA 密度を 5% 以内に一致させ、シミュレーションを 38 ~ 88 分から 1 分に短縮して、自動ポリマー粗視化へのルートとしてエージェント LLM を確立しました。
原文 (English)
A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers
Coarse-grained (CG) molecular dynamics extends polymer simulation beyond the scales accessible to all-atom (AA) methods, but bottom-up CG modeling is laborious. The CG resolution is a design choice, so a transferable parameter set is generally not available and the potentials are derived anew for each polymer mapping. Here we present CGMas, a multi-agent framework that automates topology construction, equilibration, mapping, potential derivation, and validation from a natural-language specification of the polymer and target resolution. A large-language-model (LLM) reasoning agent infers the AA topology from polymer name, while layered self-correction resolves physical errors common to unsaturated, heteroatom-containing, and polar polymers. Downstream agents equilibrate the system, map it onto CG representation, derive potentials through Boltzmann inversion, and benchmark the model against its atomistic reference. CGMas completed all 27 homopolymer and copolymer tasks, matched the AA density to within 5% in 22, and reduced simulation from 38-88 min to 1 min, establishing agentic LLMs as a route to automated polymer coarse-graining.
AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models
Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dyn…
WebRider: ライブ Web アシスタンス用のペルソナ条件付きインテント コントローラー
Web タスクを委任するには、単に質問するだけではありません。何を検証するか、不確実性にどう対処するか、どの設定が重要か、いつ停止するかなど、ポリシーを移行する必要があります。しかし、現在のライブ Web エージェントは、委任を定義するポリシーの制約を無視して、最終的な回答のみに基づいて評価されます。もっともらしい最終回答は、そのポリシーの違反を隠すことができます。当社の完全なライブ監査により、この重大なギャップが明らかになりました。強力なコントローラーはタスクの 99.2% を完了しますが、すべてのポリシー制約を順守するのは 38.8% のケースのみです。仕上げは忠実さを意味するものではありません。 WebRider は、委任されたポリシーをインテント コントラクト (Web ページが変更されても維持する必要がある目標、制約、証拠義務、回答フォーム、およびタスク ローカルのペルソナ コントロールの運用記録) として形式化することでこのギャップを埋めます。 WebRider は階層アーキテクチャを採用しています。最上位層のコントローラーはコントラクトを維持し、中間層は保護された実行可能なアクションとして意図を実現し、ツール層はブラウザー、検索、およびマップ ツールを介してこれらのアクションを実行します。当社のベンチマークである RiderBench は、42 の公開 Web サイトにわたる 4,096 のライブ Web コントラクトでこの設計を評価し、内部契約の状態と目に見えるユーザー エクスペリエンスの両方を監査して、ロールアウトでポリシーが維持されているかどうか、および手順がペルソナと一貫性があるかどうかを判断します。保護された中間インターフェイスは、高品質のトレーニング信号としても機能します。このインターフェイスを通じてトレーニングされた 8B アクション ポリシー モデルは、固定コントローラーの下で実行可能のみのベースラインよりも優れたパフォーマンスを発揮します。 WebRider は、ブラウジング パスをファーストクラス オブジェクトにすることで、アクションの実現と最終的な答えの決定を混同することなく、監査可能、人間による判断可能、学習可能なシステムを実現します。データセット URL: hf.co/datasets/WebRider/WebRider。
原文 (English)
WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance
Delegating a web task involves more than asking a question; it requires transferring a policy: what to verify, how to handle uncertainty, which preferences matter, and when to stop. Yet, current live-web agents are evaluated solely on the final answer, ignoring the policy constraints that define the delegation. A plausible final answer can conceal violations of that policy. Our full live audit reveals this critical gap: a strong controller completes 99.2% of tasks but honors all policy constraints in only 38.8% of cases. Finishing does not imply fidelity. WebRider bridges this gap by formalizing the delegated policy as an intent contract---an operational record of goals, constraints, evidence obligations, answer form, and task-local persona controls that must hold even as web pages change. WebRider employs a hierarchical architecture: a top-layer controller maintains the contract, a middle layer realizes intentions as guarded executable actions, and a tool layer executes these actions via browser, search, and maps tools. Our benchmark, RiderBench, evaluates this design on 4,096 live-web contracts across 42 public websites, auditing both the internal contract state and the visible user experience to determine if a rollout preserved its policy and if the steps were persona-consistent. The guarded middle interface also serves as a high-quality training signal; an 8B action-policy model trained through this interface outperforms executable-only baselines under a fixed controller. By making the browsing path a first-class object, WebRider enables a system that is auditable, human-judgeable, and learnable without conflating action realization with final-answer decisions. Dataset URL: hf.co/datasets/WebRider/WebRider.
MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring
Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities,…
オプティマイザーはエージェントです: プロンプト、プログラム、ML ワークフローにわたる推論主導の検索
プロンプト、プログラム、ML ワークフローを最適化するための最近のシステムは、通常、進化的検索、バンディット、テキスト グラデーション法などの明示的なアウターループ コントローラーに依存しています。私たちは根本的に異なる質問をします。この検索ポリシーのどの程度が、単一のツールを使用するエージェントによって内部化されるのでしょうか?我々は、エージェントが何を評価するか、どのように障害を診断するか、どの編集を行うか、いつ検証または再起動するかを自律的に決定する推論主導型の最適化のための統合フレームワークである ReASearch を紹介します。このエージェントは、手動で設計されたヒューリスティックに基づいて提案を生成するだけの役割を果たすのではなく、結果を積極的に分析し、予算を割り当て、永続的な記憶を通じて長期にわたる戦略を洗練します。 ReASearch は、共有エージェント ループとドメイン固有のツールを使用して、まったく同じスキャフォールドをインスタンス化し、プロンプト、プログラム、ML ワークフローを最適化します。 14 の多様なタスクにわたって、専用の最適化システムと競合し、ほとんどの場合それよりも優れており、強力なドメイン固有のベースラインよりも 2% ~ 40% の向上を達成し、場合によっては、人間が以前に最もよく知っていた結果を改善するソリューションを発見することもあります。重要なのは、通常は明示的なコントローラーによって実装される複雑な検索動作が、エージェントの推論プロセスから自然に現れることです。
原文 (English)
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by a single tool-using agent? We present ReASearch, a unified framework for reasoning-driven optimization in which the agent autonomously decides what to evaluate, how to diagnose failures, which edits to make, and when to verify or restart. Rather than serving only as a proposal generator guided by hand-designed heuristics, the agent actively analyzes outcomes, allocates budget, and refines its strategy over long horizons through persistent memory. With a shared agent loop and domain-specific tools, ReASearch instantiates the exact same scaffold to optimize prompts, programs, and ML workflows. Across 14 diverse tasks, it is competitive with and mostly better than specialized optimization systems, achieving gains of 2% to 40% over strong domain-specific baselines, and in some cases discovering solutions that improve on prior human best-known results. Crucially, we observe that complex search behaviors, which are typically implemented by explicit controllers, emerge naturally from the agent's reasoning process.
bioMoR: 効果的なゲノム学習のための生物学に基づく再帰混合
高次元オミクス解析用のトランスフォーマー モデルは数千の遺伝子または経路を処理しますが、詳細な計算が必要なのはサブセットのみです。 Mixture-of-Recursions (MoR) は、適応型トークン選択またはエキスパート選択ルーティングを通じて効率を向上させます。私たちは bioMoR を提案します。これは、私たちの知る限り、MoR を遺伝子レベルおよび経路レベルの学習に適用する最初のフレームワークです。私たちの貢献には、MoR バックボーン内に構造化された生物学的知識を統合するための 3 つの場所を特定することが含まれます。グラフベースの情報共有によりトークンの埋め込みが洗練され、構造的バイアスにより生物学的に関連するトークンへの注意が誘導され、グラフ認識ルーターが近傍情報を使用して各トークンの再帰深さを決定します。これらの手法は、トークン相互作用に関する追加の知識が、モデルが埋め込みを構築し、どのトークンをより深く学習する必要があるかを選択するのに効果的に役立つという洞察に基づいています。多様なオミクス データ タイプにまたがる 8 つのベンチマークにわたって、統一された 5 重交差検証プロトコルに基づいて評価された bioMoR は、非再帰的 Transformer よりも使用するパラメーターが 75 パーセント少なく、FLOP が最大 58 パーセント少ない一方で、生物学に依存しない最も強力な MoR ベースラインと比較して、平均マクロ F1 が 8.2 パーセント ポイント、バランスのとれた精度が 7.1 パーセント ポイント向上しています。選択されたマーカー遺伝子または経路は生物学的解釈可能性を提供し、そのトークン固有の再帰の深さは計算がどのように割り当てられているかを明らかにします。
原文 (English)
bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning
Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture-of-Recursions (MoR) improves efficiency through adaptive token-choice or expert-choice routing. We propose bioMoR, which, to the best of our knowledge, is the first framework to apply MoR to gene-level and pathway-level learning. Our contributions include identifying three locations for integrating structured biological knowledge within an MoR backbone: graph-based information sharing refines token embeddings, a structural bias guides self-attention toward biologically related tokens, and a graph-aware router uses neighborhood information to determine each token's recursion depth. These techniques are centered on our insight that additional knowledge of token interaction can effectively help models construct embeddings and select which tokens should be learned more deeply. Across eight benchmarks spanning diverse omics data types and evaluated under a unified five-fold cross-validation protocol, bioMoR improves average macro-F1 by 8.2 percentage points and balanced accuracy by 7.1 percentage points over the strongest biology-agnostic MoR baseline while using 75 percent fewer parameters and up to 58 percent fewer FLOPs than a non-recursive Transformer. The selected marker genes or pathways provide biological interpretability, while their token-specific recursion depths reveal how computation is allocated.
安価なフェイクから純粋な合成まで: T2V フェイクニュース動画の新時代に対処する
最近の Text-to-Video (T2V) 生成モデルにより、フェイク ニュース ビデオをゼロから合成できるようになり、脅威が既存の映像から組み立てられた安価なフェイクを超えて変化します。このようなニュースビデオは、捏造された物語と厳密に一致する可能性があり、既存の検出器にとってモダリティ調整の罠を生み出します。既存のデータセットには純粋な合成フェイク ニュース ビデオがありません。 T2V モデルにフェイク ニュース ビデオの説明を直接プロンプトすると、完全に位置合わせされたサンプルが得られますが、フェイク ニュース ビデオ検出 (FNVD) が単一モードのショートカットに低下し、意味的視覚的劣化が生じます。これに対抗するために、T2V-FNVD を 3 つのラベル (本物、安価なフェイク、純粋合成フェイク) を持つ新しい 3 値分類タスクとして定式化し、最初の純粋合成フェイク ニュース ビデオ データセット (PS-FNVD) を構築します。 PS-FNVD には、調整された欺瞞を伴う捏造されたイベント (タイプ 1) と、偽の視覚的来歴を持つ真のイベント (タイプ 2) が含まれており、モデルが単峰性ショートカットを悪用するのを防ぎます。さらに、推論に基づいた T2V-FNVD (R-T2V) フレームワークを提案します。条件付き根拠生成と教師付き微調整を通じてトレーニングされた R-T2V は、高レベルの意味論的ロジックと低レベルの物理生成トレースを統合して、三値の真実性ラベルを予測します。 10 の一般的なベースラインにわたる広範な実験により、R-T2V が最先端のパフォーマンスを達成し、2 番目に優れたベースラインを精度で 12.20 パーセント ポイント、マクロ $F_1$ で 8.46 パーセント ポイント上回っていることが示されました。
原文 (English)
From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos
Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes assembled from existing footage. Such news videos can closely match fabricated narratives, creating a modality alignment trap for existing detectors. Existing datasets lack pure synthesis fake news videos. Although directly prompting T2V models with descriptions of fake news videos can yield perfectly aligned samples, it reduces the fake news video detection (FNVD) to unimodal shortcuts and causes semantic-visual degeneration. To counter this, we formulate T2V-FNVD as a novel ternary classification task with three labels (real, cheap fake, and pure synthesis fake) and construct the first pure synthesis fake news video dataset (PS-FNVD). PS-FNVD includes fabricated events with aligned deception (Type 1) and true events with false visual provenance (Type 2), preventing models from exploiting unimodal shortcuts. Furthermore, we propose the Reasoning-guided T2V-FNVD (R-T2V) framework. Trained through conditioned rationale generation and supervised fine-tuning, R-T2V integrates high-level semantic logic with low-level physical generative traces to predict the ternary veracity label. Extensive experiments across 10 prevailing baselines show that R-T2V achieves the state-of-the-art performance, outperforming the second-best baseline by 12.20 percentage points in accuracy and 8.46 percentage points in macro $F_1$.
IB-RL: 戦略的対話エージェントのための分離された二国間強化学習
強化学習 (RL) は、数学的推論やコード実行など、定常で検証可能な報酬を伴うタスクの大規模言語モデル (LLM) の改善において優れた成果を上げています。これらの設定では、環境は固定ルールに従い、エージェントに戦略的に適応しません。戦略的対話はこの点で異なります。環境は政策に適応するもう 1 つのエージェントであり、成功は両者間の相互作用に依存します。このインタラクティブな性質にもかかわらず、現在の RL アプローチは通常、固定された対応者またはシミュレーターに対してターゲット エージェントをトレーニングします。この訓練パラダイムは、カウンターパート間で一般化する戦略を学習するのではなく、カウンターパート固有の規則性を活用する政策を促進することがわかりました。私たちはこの問題を静的対物不一致と呼び、実験で直接定量化します。これに対処するために、私たちは分離型双方向強化学習 (IB-RL) を提案します。この学習では、2 つの役割が共同ロールアウトを通じて共進化し、各役割が完全に独立した利点、アクション マスク、および更新パスを通じて独自の報酬を最適化します。私たちは凍結されたポリシーを、両方のドメインで完全に独立した保持されている対応物と比較して評価します。 Vehicle TeleSales では、IB-RL は 89.6% の成功率@1 を達成しました。一方、最良の片側 RL ベースラインでは 84.6% でした。 Deal-or-NoDeal では、DeepSeek V4 Pro に対して 98.4% の合意に達しました。一方、最良の一方的なベースラインでは 86.4% でした。これらの結果は、厳密なエージェント分離を使用して両方の役割を共同トレーニングすると、目に見えない対応者により効果的に一般化するポリシーが生成されることを示しています。
原文 (English)
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents
Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as mathematical reasoning and code execution. In these settings, the environment follows fixed rules and does not adapt strategically to the agent. Strategic dialogue differs in this respect: the environment is another agent that adapts to the policy, and success depends on the interaction between the two sides. Despite this interactive nature, current RL approaches typically train a target agent against a fixed counterpart or simulator. We find that this training paradigm encourages the policy to exploit counterpart-specific regularities rather than learn strategies that generalize across counterparts. We call this problem the static-counterpart mismatch, which we quantify directly in our experiments. To address it, we propose Isolated Bilateral Reinforcement Learning (IB-RL), in which the two roles coevolve through joint rollouts while each role optimizes its own reward through fully independent advantages, action masks, and update paths. We evaluate frozen policies against fully independent held-out counterparts in both domains. On Vehicle TeleSales, IB-RL achieves 89.6% Success@1, compared to 84.6% for the best unilateral RL baseline. On Deal-or-NoDeal, it reaches 98.4% agreement against DeepSeek V4 Pro, compared to 86.4% for the best unilateral baseline. These results indicate that jointly training both roles with strict peragent isolation produces policies that generalize more effectively to unseen counterparts.
MemPrism: 長期エージェント向けのタスク条件付きリレーショナル メモリ ビュー
長期的なエージェントは経験を再利用するために記憶に依存していますが、既存の記憶システムは多くの場合、固定された表現を通じて証拠を直接消費できると想定しています。これは、関連情報が入手可能であるにもかかわらず、現在の決定に合わせて整理されていない、表現の不一致につながります。この目的を達成するために、我々は永続的な経験の記憶を意思決定時の作業記憶から分離するタスク条件付きリレーショナル記憶フレームワークである MemPrism を提案します。 MemPrism は、インタラクションをイベント ストリームとして記録し、現在のタスク コンテキストに従ってリレーショナル ビューを動的に構築します。軽量ビュー ポリシーは関係構造、証拠範囲、結果条件、粒度を選択しますが、決定論的コンポーザーとレンダリングは歴史的事実を凍結タスク ポリシーの一時的な光学式作業メモリ ビューに変換します。ロングホライズンの具体化されたベンチマークと Web エージェントのベンチマークに関する実験では、MemPrism がメモリ トークンの消費を削減しながら、特に軌跡が長くなったときに、タスクのパフォーマンスを一貫して向上させることが示されています。さらに、学習されたビュー ポリシーは追加の適応なしで異なる VLM 間で転送され、エージェントの一般的なメモリ インターフェイスとしてのタスク条件付きリレーショナル ビューの有効性を示しています。
原文 (English)
MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents
Long-horizon agents rely on memory to reuse experiences, yet existing memory systems often assume that evidence can be directly consumed through a fixed representation. This leads to representation mismatch, where relevant information is available but not organized for the current decision. To this end, we propose MemPrism, a task-conditioned relational memory framework that separates persistent experience storage from decision-time working memory. MemPrism records interactions as the event stream and dynamically constructs relational views according to the current task context. A lightweight view policy selects the relation structure, evidence range, outcome condition, and granularity, while a deterministic composer and render transform historical facts into a temporary optical working-memory view for a frozen task policy. Experiments on long-horizon embodied and web-agent benchmarks show that MemPrism consistently improves the task performance, especially as trajectories become longer, while reducing memory token consumption. Furthermore, the learned view policy transfers across different VLMs without additional adaptation, demonstrating the effectiveness of task-conditioned relational views as a general memory interface for agents.
Mind the Gap: A Dual Knowledge Graph Framework for Unified Multi-task User Intent Inference
This paper proposes DKG-MTI, a dual knowledge graph framework for unified multi-task user intent inference from online travel reviews. Exis…
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each act…
LiFTER: A Grounded Neuro-Symbolic Microscope for Continuous-Time Dynamic Graph Forecasting
Continuous-time dynamic graph models predict future links by compressing past interactions into neural states. Although effective for forec…
Surg-UniWorld: マルチモーダル制御の専門家による統合外科世界モデル
制御可能な外科世界モデルは、現実的な器具と組織の相互作用を合成することにより、外科用人工知能とシミュレーションのための生成基盤を提供できます。しかし、既存の方法には統合されたマルチモーダルな制御パラダイムが欠けており、異種の視覚条件を直接融合させると、解剖学的歪み、器具の外観のドリフト、および時間的に一貫性のない相互作用が生じることがよくあります。この研究では、マルチモーダル制御の専門家による統合手術世界モデル {Surg-UniWorld} を提案します。 Surg-UniWorld はまず、最初のフレームの外観と階層的セマンティック マスクから {階層的外科アンカー} を構築し、永続的なシーンのアイデンティティ、解剖学的構成、およびインタラクション境界を保持します。 {アンカー相対モダリティの専門家}は、共有アンカーと相対的なエッジ、深度、およびオプティカル フローの証拠を解釈し、相補的な境界、幾何学的、および動きの情報をキャプチャします。 {マルチモーダル制御エキスパート} はさらに、アクティブ化されたモダリティ インクリメントの寄与を維持した段階的な構成を実行し、Wan2.2 ビデオ拡散バックボーンの制御ヒントを生成します。マルチモーダルな手術世界モデリングをサポートするために、制御可能な手術ビデオ生成のベンチマークである Cholec80-SurgWAM をさらに構築します。広範な実験により、Surg-UniWorld は、生成品質、時間的一貫性、およびマルチモーダル制御性において、既存の制御可能なビデオ生成方法および外科ワールド モデルのベースラインよりも一貫して優れていることが実証されています。
原文 (English)
Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts
Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions. However, existing methods lack a unified multimodal control paradigm, while direct fusion of heterogeneous visual conditions often causes anatomical distortion, instrument appearance drift, and temporally inconsistent interactions. In this work, we propose {Surg-UniWorld}, a unified surgical world model with multimodal control experts. Surg-UniWorld first constructs a {Hierarchical Surgical Anchor} from first-frame appearance and hierarchical semantic masks to preserve persistent scene identity, anatomical organization, and interaction boundaries. {Anchor-Relative Modality Experts} then interpret edge, depth, and optical-flow evidence relative to the shared anchor, capturing complementary boundary, geometric, and motion information. A {Multimodal Control Expert} further performs contribution-preserving stage-wise composition of the activated modality increments and generates control hints for the Wan2.2 video diffusion backbone. To support multimodal surgical world modeling, we further construct Cholec80-SurgWAM, a benchmark for controllable surgical video generation. Extensive experiments demonstrate that Surg-UniWorld consistently outperforms existing controllable video generation methods and surgical world-model baselines in generation quality, temporal consistency, and multimodal controllability.
LLM による潜在性を認識したインスタンス生成による並列アルゴリズム ポートフォリオの進化
大規模言語モデルによるポートフォリオの自動構築 (LLM-ACP) は、複雑な組み合わせ最適化問題を解決する際の実際の少数ショット シナリオでは一般化が不十分です。インスタンスとアルゴリズムの共進化フレームワークは、現在のアルゴリズム ポートフォリオではパフォーマンスが劣る、生成されたハード インスタンスを使用してトレーニング データセットを拡張することでこの問題に対処し、それによって一般化を強化します。ただし、このパラダイムは 2 つの重大な制限に直面しています。1 つはインスタンスの硬度の評価が高品質の参照ソリューションに依存すること、もう 1 つはシングルモード生成パターンがインスタンスの多様性を制限することです。これらの制限を克服するために、Potential-aware Instance and Algorithm Co-evolution (PIAC) フレームワークを導入します。私たちの主な貢献は 2 つあります。まず、参照ソリューションの必要性を排除する新しい指標である潜在的なゲインを提案します。このメトリクスは、生成されたアルゴリズムに摂動を加え、生成された問題インスタンスの改善の可能性を評価することによって汎化ゲインを推定します。第 2 に、PIAC は LLM を活用して多様なインスタンス ミューテーターを合成し、問題インスタンス空間のより広い領域を探索し、それによってポートフォリオの一般化機能を強化します。摂動空間がアルゴリズムごとに異なることを考慮して、Greedy Constructive、Ant Colony Optimization、および Guided Local Search アルゴリズムのバックボーンに基づいてフレームワークをインスタンス化します。 6 つの異なるデータ分布にわたる巡回セールスマン問題 (TSP) とキャパシテッド車両経路指定問題 (CVRP) に関する包括的な評価では、PIAC が常に最先端の LLM-ACP ベースラインを上回り、特に TSP Greedy Constructive ポートフォリオで 19.76% の相対的な改善を達成していることが実証されています。
原文 (English)
Evolving Parallel Algorithm Portfolios via Potential-Aware Instance Generation with LLMs
The Automatic Construction of Portfolios via Large Language Models (LLM-ACP) suffers from poor generalization in practical few-shot scenarios when solving complex combinatorial optimization problems. Instance and algorithm co-evolution frameworks address this by expanding the training dataset with generated hard instances on which the current algorithm portfolio underperforms, thereby enhancing generalization. However, this paradigm faces two critical limitations: evaluating instance hardness relies on high-quality reference solutions, and single-mode generation patterns limit instance diversity. To overcome these limitations, we introduce the Potential-aware Instance and Algorithm Co-evolution (PIAC) framework. Our core contribution is twofold. First, we propose potential gain, a novel metric that eliminates the need for reference solutions. This metric estimates generalization gain by perturbing the generated algorithms and assessing their improvement potential on generated problem instances. Second, PIAC leverages LLMs to synthesize diverse instance mutators, exploring a broader region of the problem-instance space and thereby enhancing the portfolio's generalization capabilities. Given that perturbation spaces vary across different algorithms, we instantiate our framework on Greedy Constructive, Ant Colony Optimization, and Guided Local Search algorithmic backbones. Comprehensive evaluations on the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP) across six distinct data distributions demonstrate that PIAC consistently outperforms state-of-the-art LLM-ACP baselines, notably achieving a 19.76% relative improvement for TSP Greedy Constructive portfolios.
Gated-BEPO: 大規模言語モデル エージェントに対する信頼ゲート型ベルマン クレジット割り当て
長期的な環境で大規模な言語モデル エージェントをトレーニングするには、まばらな最終結果からのクレジットを個々のアクションに割り当てる必要があります。既存の批評家なしの手法は、軌跡レベルの報酬をステップ全体に均一に伝播しますが、最近のアプローチでは、繰り返される状態を照合してステップレベルのグループを構築し、各グループ内のアクションを比較します。前者は、失敗した軌道での有用なアクションと、成功した軌道での非効果的なアクションを区別できません。後者は、個々の軌跡の結果から直接得られるステップ クレジットと、エピソード レベルのクレジットとの固定重みの融合に依存します。我々は、経験的なロールアウトグラフからステップレベルのクレジットを導き出す Gated-BEPO を提案します。各ロールアウト グループに対して、Gated-BEPO は経験的なグラフを構築し、現在のポリシーの経験的なアクション分布を反映する平均バックアップ ベルマン固定点を通じてノード値を推定します。次に、一般化利点推定を使用して、サンプリングされた各軌跡に沿ってこれらの時間差残差を蓄積し、即時効果と下流効果の両方を捕捉するステップレベルのベルマン利点を生成します。エピソード レベルとステップ レベルのクレジットを適応的に融合するために、信頼ゲートには複数の観察された後継者がある状態でのみベルマン クレジットが組み込まれ、それ以外の場合はエピソード レベルのクレジットが使用されます。 WebShop、ALFWorld、ビジュアル倉庫番の実験では、言語モデルと視覚言語モデル全体で一貫した改善が示されている一方、診断アブレーションはベルマン固定小数点値推定の有効性を裏付けており、ステップレベルのクレジットは最終的な利点に一律ではなく選択的に組み込まれるべきであることが示されています。
原文 (English)
Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents
Training large language model agents in long-horizon environments requires assigning credit from sparse terminal outcomes to individual actions. Existing critic-free methods propagate trajectory-level rewards uniformly across steps, while recent approaches construct step-level groups by matching repeated states and compare actions within each group. The former cannot distinguish useful actions in failed trajectories from ineffective actions in successful ones. The latter rely on step credit derived directly from individual trajectory outcomes and fixed-weight fusion with episode-level credit. We propose Gated-BEPO, which derives step-level credit from empirical rollout graphs. For each rollout group, Gated-BEPO constructs an empirical graph and estimates node values through a mean-backup Bellman fixed point that reflects the empirical action distribution of the current policy. We then accumulate these temporal-difference residuals along each sampled trajectory using generalized advantage estimation, yielding step-level Bellman advantages that capture both immediate and downstream effects. To adaptively fuse episode- and step-level credit, a confidence gate incorporates Bellman credit only at states with multiple observed successors and otherwise uses episode-level credit. Experiments on WebShop, ALFWorld, and visual Sokoban show consistent improvements across language and vision-language models, while diagnostic ablations support the effectiveness of Bellman fixed-point value estimation and show that step-level credit should be incorporated selectively rather than uniformly into the final advantage.
CEDAR: 複雑なシステムの目標指向の最適化のためのエージェント オーケストレーション ツリー検索
人工生命の研究の中核となる複雑なシステムは、集団動態や生物学から経済政策や戦略的意思決定に至るまで、創発的な行動を生み出す非線形のフィードバック駆動型相互作用を通じて多様な現象をモデル化します。しかし、フィードバック構造がどのように創発的な行動を引き起こすかを予測することは難しく、これは人工生命における中心的な未解決の問題であり、目標指向の設計を非常に困難なものにしています。確立された慣行では、システム構造は DYNAMO や STELLA などの特殊なモデリング言語で記述されており、採用を制限し、タイムリーな意思決定を妨げる労働集約的なワークフローによって課題がさらに複雑になります。これらの課題に対処するために、大規模言語モデル (LLM) エージェントを使用して、ユーザーが指定した行動目標を満たす複雑なシステムを発見する自律的な手法である CEDAR を導入します。私たちの主なイノベーションは、複雑なシステムと深く結合した LLM 駆動のモンテカルロ木探索 (MCTS) です。反復ごとに、LLM ジャッジが指定された目標に対する緊急行動を評価し、LLM エディターが改良されたバリアントを提案します。ジャッジは適応度関数として機能し、エディターはバリエーション オペレーターとして機能します。進化的計算における生成と評価のループに似ています。複雑なシステムを、ドメイン固有のプリミティブを備えた制限された実行可能な Python のサブセットとして表現し、LLM がシステム ダイナミクスを直接変更できるようにします。 CEDAR は、LLM パラメータ化された遷移カーネルと値関数を備えた MCTS バリアントとしてこれを形式化し、ソリューションの多様性を維持しながら複雑なシステム動作の目標指向の発見を可能にし、その LLM ベースの解釈可能性により、構造変化がどのように新たな動作を引き起こすかを明らかにします。 CEDAR は人的労力を軽減しながら、既存のアプローチでは実現が困難な機能を実現し、ドメイン全体での複雑なシステムの広範な導入を促進します。
原文 (English)
CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems
Complex systems, core objects of study in artificial life, model diverse phenomena through nonlinear, feedback-driven interactions that produce emergent behavior, with applications from population dynamics and biology to economic policy and strategic decision-making. Yet the difficulty of predicting how feedback structure gives rise to emergent behavior, a central open problem in artificial life, makes goal-directed design exceptionally challenging. In established practice, system structures are written in specialized modeling languages such as DYNAMO or STELLA, compounding the challenge with labor-intensive workflows that limit adoption and hinder timely decision-making. To address these challenges, we introduce CEDAR, an autonomous method that uses Large Language Model (LLM) agents to discover complex systems satisfying user-specified behavioral goals. Our key innovation is an LLM-driven Monte Carlo Tree Search (MCTS) deeply coupled with complex systems: at each iteration, an LLM Judge evaluates emergent behavior against specified goals and an LLM Editor proposes improved variants, with the Judge acting as a fitness function and the Editor as a variation operator, akin to a generate-and-evaluate loop in evolutionary computation. We represent complex systems as a restricted, runnable subset of Python with domain-specific primitives, letting LLMs modify system dynamics directly. CEDAR formalizes this as an MCTS variant with an LLM-parameterized transition kernel and value function, enabling goal-directed discovery of complex system behaviors while preserving solution diversity, and its LLM-based interpretability reveals how structural changes drive emergent behavior. CEDAR reduces human effort while enabling capabilities difficult to achieve with existing approaches, facilitating broader adoption of complex systems across domains.
SkillEval: エージェントのスキル品質を解釈可能な信号に分解する
エージェント スキルは、エージェントが特殊なタスクを解決するのに役立つ再利用可能な手順に関する知識を提供します。使用が拡大するにつれて、スキルの品質を評価することがますます重要になります。既存の評価では、スキルが特定の下流タスクのパフォーマンスを向上させるかどうかをテストすることでスキルの品質を測定することがよくあります。ただし、再利用可能なスキルは複数のタスク シナリオに適用できる場合があります。下流の評価は主にスキルと評価されるタスクの互換性を反映しており、スキルの品質の部分的な見方のみを提供し、スキルのどの側面を改善する必要があるかを特定しません。 \texttt{SKILL.md} ドキュメントの一般的なプロパティがスキルの品質に重要な役割を果たしていることがわかりました。これらのプロパティを評価するために、ドキュメントレベルのスキル評価のための解釈可能なフレームワークである \textbf{SkillEval} を提案します。 SkillEval は、固定された検査可能なスコアリング方向を使用して各プロパティを評価し、解釈可能なスコアを生成します。さらに、長さや形式など、無関係な文書の特徴の影響を測定して軽減するため、各スコアは意図した意味論的特性をより具体的に捉えることができます。具体的には、SkillEval は、モデルの隠れた表現空間内の制御されたポジティブとネガティブのスキルのペアから各品質プロパティの解釈可能な方向を学習し、その表現をこれらの固定方向に投影することによって新しいスキルをスコア付けします。 SkillEval を使用して、管理された品質テストでスキルを評価し、SkillEval がさまざまな品質のスキルを確実に区別できることを示します。さらに、SkillEval スコアは下流のタスクのパフォーマンスを厳密に反映しており、スキルがエージェントのタスクの完了に役立ちそうかどうかを早期に示します。スキルドキュメントの弱点を診断し、対象となる改訂をガイドするための SkillEval についてさらに詳しく調査します。改訂されたスキルにより、対象のプロパティが改善され、下流のタスクでより高い合格率が達成されます。
原文 (English)
SkillEval: Decomposing Agent Skill Quality into Interpretable Signals
Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existing evaluations often measure skill quality by testing whether a skill improves performance on specific downstream tasks. However, a reusable skill may apply to multiple task scenarios. Downstream evaluation mainly reflects the compatibility between a skill and the evaluated task, provides only a partial view of skill quality, and does not identify which aspect of the skill should be improved. We find that general properties of the \texttt{SKILL.md} document play an important role in skill quality. To evaluate these properties, we propose \textbf{SkillEval}, an interpretable framework for document-level skill evaluation. SkillEval evaluates each property using a fixed and inspectable scoring direction, producing interpretable scores. It further measures and reduces the influence of unrelated document features, such as length and formatting, so that each score captures its intended semantic property more specifically. Specifically, SkillEval learns an interpretable direction for each quality property from controlled positive--negative skill pairs in the hidden representation space of the model, and scores a new skill by projecting its representation onto these fixed directions. We use SkillEval to evaluate skills in controlled quality tests and show that SkillEval reliably distinguishes skills of different quality. In addition, SkillEval scores closely reflect downstream task performance, providing an early indication of whether a skill is likely to help an agent complete a task. We further explore SkillEval for diagnosing weaknesses in skill documents and guiding targeted revisions. The revised skills improve the targeted properties and achieve higher pass rates on downstream tasks.
点からエッジへ: 物理に敏感な PDE 学習のためのエッジ条件付きスペクトル演算子
ニューラル演算子は偏微分方程式 (PDE) を解くための中心的なツールとなっており、スペクトル演算子は空間的位置全体で効率的なグローバル混合を提供します。ただし、多くの偏微分方程式には、基礎となる物理的動作にとって重要な、物理に敏感なローカル構造が含まれています。たとえば、Darcy フローでは、局所的な材料界面が浸透率場の急激な変化によって反映されることが多く、解に大きな影響を与える可能性があります。既存のスペクトル オペレーターは主に中心点表現に基づいてモーダル混合を適応させるため、そのような局所的な構造変化に対する応答性が不十分になります。我々は、局所的なエッジごとの変動を使用してグローバルなスペクトル混合を変調する新しいスペクトル演算子フレームワークである、エッジ条件付きスペクトル演算子 (ESO) を提案します。ローカル エッジ情報をスペクトル モード選択に注入するペアワイズ バリエーション モーダル ミキサー (PVMM) を組み込むことにより、ESO はスペクトル ニューラル オペレーターのグローバル近似機能を維持しながら、学習されたカーネルが物理に敏感なローカル構造に適応できるようにします。さらに、タスク固有の物理量によって識別される物理的に重要な領域を強調するタスク適応型の物理認識再重み付け (PAR) を導入します。 9 つの PDE ベンチマークにわたって、ESO は一貫して最先端のパフォーマンスを達成しています。視覚的および領域ごとの解析は、ESO が係数ジャンプ、高勾配の流れ構造、およびその他の物理的に影響を受けやすい領域付近での解エラーを低減することをさらに実証しています。コードは https://github.com/Tanpig-X/ESO で入手できます。
原文 (English)
From Points to Edges: Edge-Conditioned Spectral Operators for Physics-Sensitive PDE Learning
Neural operators have become a central tool for solving partial differential equations (PDEs), with spectral operators offering efficient global mixing across spatial locations. However, many PDEs contain physics-sensitive local structures that are critical to the underlying physical behavior. For example, in Darcy flow, local material interfaces are often reflected by sharp changes in the permeability field and can strongly influence the solution. Existing spectral operators primarily adapt modal mixing based on center-point representations, making them insufficiently responsive to such localized structural variations. We propose the Edge-Conditioned Spectral Operator (ESO), a novel spectral operator framework that modulates global spectral mixing using local edge-wise variations. By incorporating the Pairwise-Variation Modal Mixer (PVMM) to inject local edge information into spectral mode selection, ESO preserves the global approximation capability of spectral neural operators while enabling the learned kernel to adapt to physics-sensitive local structures. Furthermore, we introduce a task-adaptive Physics-Aware Reweighting (PAR) that emphasizes physically important regions, identified by taskspecific physical quantities. Across nine PDE benchmarks, ESO consistently achieves state-of-the-art performance. Visual and region-wise analyses further demonstrate that ESO reduces solution errors near coefficient jumps, high-gradient flow structures, and other physically sensitive regions. The code is available at https://github.com/Tanpig-X/ESO.
長期にわたるエージェントの軌跡の帰属: 統合ベンチマークときめ細かいアノテーション フレームワーク
大規模言語モデル (LLM) エージェントは、ユーザーの指示、ツールの使用、外部観察、および記憶を含む長期的な軌道を通じて動作することがますます増えています。既存のベンチマークは主に行動の結果を評価しますが、きめの細かいアトリビューション分析に対するサポートは限定的です。私たちは軌跡の帰属を導入し、このタスクのためのベンチマークと注釈フレームワークを開発します。このベンチマークは、統一されたコンポーネント スキーマの下で異種の軌跡を整理し、主要なアトリビューション コンポーネントの注釈を、該当する場合には攻撃チェーンと実行チェーンとともに提供します。 AgentDojo の軌跡と、Agent3Sigma の Stage および Canary 設定を使用してベンチマークをインスタンス化すると、タスクに合わせたアクション、安全でないアクション、安全性の拒否をカバーする 1,300 を超える注釈付きの軌跡が得られます。このベンチマークは、プライマリ アトリビューション ローカライゼーションとアトリビューション チェーンの回復という 2 つの評価タスクを定義し、増分軌道の寄与とコンポーネント レベルのリーブ ワン アウト摂動に基づいた参照ベースラインを提供します。ローカルおよび長期のアトリビューション、構造化されたアトリビューション チェーンなど、多様なアトリビューション設定をキャプチャします。参照ベースラインの結果は、これらの設定全体で大幅なパフォーマンスの違いを示しており、ベンチマークのアトリビューションの課題の最初の特徴付けを提供します。この最初のインスタンス化を超えて、新しいエージェント モデルによって生成された軌跡を同じフレームワークの下で標準化、注釈付け、評価できるようにする再利用可能な注釈スキルをリリースします。プロジェクトのリソースと将来のリリースは、https://github.com/chenjing-2024/agent-trajectory-attribution で入手できます。
原文 (English)
Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine-grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchmark organizes heterogeneous trajectories under a unified component schema and provides annotations of the primary attribution component, together with attack and execution chains where applicable. Instantiating the benchmark with trajectories from AgentDojo and the Stage and Canary settings of Agent3Sigma yields more than 1,300 annotated trajectories covering task-aligned actions, unsafe actions, and safety refusals. The benchmark defines two evaluation tasks, primary attribution localization and attribution-chain recovery, and provides reference baselines based on incremental trajectory contribution and component-level leave-one-out perturbation. It captures diverse attribution settings, including local and long-range attribution as well as structured attribution chains. Reference baseline results exhibit substantial performance differences across these settings, providing an initial characterization of the benchmark's attribution challenges. Beyond this initial instantiation, we release a reusable annotation skill that enables trajectories generated by new agent models to be standardized, annotated, and evaluated under the same framework. Project resources and future releases are available at https://github.com/chenjing-2024/agent-trajectory-attribution.
高速 LapSum: 100 万スケールでの正確な微分可能な Top-k
top-$k$ 操作は、最新のスパース計算の基本的な構成要素であり、トークン ルーティング、エキスパート アクティベーション、メモリ選択、およびアテンション プルーニングを可能にします。しかし、標準のハードトップ $k$ は勾配をブロックしますが、既存の連続 (ソフト) 緩和は依然として大規模モデルにはコストがかかりすぎます。 Fast LapSum は、ソート後に GPU ソルバーが線形時間で実行される正確な予算のソフトトップ $k$ プリミティブです。正規化制約を緩和する DFTopK などの以前の線形時間手法とは異なり、Fast LapSum は、我々の知る限りでは、エンドツーエンドで完全に微分可能でありながら $k$ の正確な選択質量を保存する最初の手法です。私たちのソルバーは、線形時間しきい値計算と解析ベクトル (ヤコビ積) を組み合わせ、極端なスケールでは確率的ブラケットを使用してカーネル ノイズ スコアの不確実な中間帯域のみを並べ替えます。結果として生じるオーバーヘッドはほとんど無視できます。ソルバーは $10^6$、$10^7$、および $10^8$ スコアをそれぞれ $0.41$、$1.15$、および $5.23$\,ms で処理します。これにより、正確なソフト トップ $k$ がスパース ルーティング、取得、大規模な最適化に実用的になります。私たちは、トレーニング ループ内で数百万の座標にわたって動作する 2 つの要求の厳しいアプリケーションで Fast LapSum を実証します。1 つは、画像ピクセルの ${\sim}0.02\%$ という正確なソフト バジェットを使用してメガピクセルのスパース敵対的サンプルを生成すること、最先端の手法を上回る桁違いの高速化を達成すること、および完全微分可能なスパース画像コーダーを最初からトレーニングすることです。
原文 (English)
Fast LapSum: Exact Differentiable Top-k at Million Scale
The top-$k$ operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selection, and attention pruning. Yet standard hard top-$k$ blocks gradients, while existing continuous (soft) relaxations remain too costly for large-scale models. We introduce Fast LapSum, an exact-budget soft top-$k$ primitive whose GPU solver runs in linear time after sorting. Unlike prior linear-time methods such as DFTopK, which relax the normalization constraint, Fast LapSum is, to our knowledge, the first method to preserve an exact selection mass of $k$ while remaining fully differentiable end-to-end. Our solver combines a linear-time threshold computation with an analytical vector--Jacobian product, and for extreme scales employs probabilistic bracketing to sort only the uncertain middle band of kernel-noised scores. The resulting overhead is almost negligible: the solver processes $10^6$, $10^7$, and $10^8$ scores in $0.41$, $1.15$, and $5.23$\,ms, respectively. This makes exact soft top-$k$ practical for sparse routing, retrieval, and large-scale optimization. We demonstrate Fast LapSum on two demanding applications operating over millions of coordinates inside the training loop: generating megapixel sparse adversarial examples with an exact soft budget of ${\sim}0.02\%$ of an image's pixels, achieving an order-of-magnitude speedup over state-of-the-art methods, and training a fully differentiable sparse image coder from scratch.
ReGraph: 食品画像からレシピグラフを生成する方法を学ぶ
最近の大規模マルチモーダル モデル (LMM) は、食品画像からのレシピ生成において目覚ましいパフォーマンスを達成しています。ただし、調理は、順序付けられたアクションを通じて食材が状態変化を受ける構造化された変換プロセスであるのに対し、自由形式のレシピ言語では、対応するエンティティ、中間状態、依存関係が暗黙的かつ複雑なままになります。グラフ表現により、この手続き型知識が明示的かつ構成的になり、モデル出力が単にもっともらしいテキスト記述を提示するのではなく、プロセス レベルの知識をエンコードしているかどうかを評価するための構造化された基盤が提供されます。この制限に対処するために、我々は、材料、調理動作、ツールをエンティティとして表現し、エンティティ属性を使用して材料の状態変化を記述し、型付き関係を使用して操作ターゲット、目的地、および手続きの順序をエンコードする大規模なレシピ グラフ データセットである ReGraph を紹介します。 ReGraph にはさらに、明示的なレシピ推論思考連鎖 (RR-CoT) トレースが組み込まれており、手続き型分解と構造化グラフ生成に補助的な監視を提供します。 ReGraph に基づいて、私たちはレシピ グラフ学習 (RGL) を提案します。これは、LMM が構造化されたレシピ グラフの形式で食品画像からもっともらしいきめの細かい調理ワークフローを生成できるようにする 2 段階のフレームワークです。決定論的でスキーマを意識したマッチング プロトコルの下での私たちの実験では、テキスト生成の品質と回復可能な手続き構造の間に大きなギャップがあることが明らかになりました。既存のアプローチで生成されたレシピは、競合するテキスト生成スコアを達成しますが、ReGraph スキーマの下で限定的な参照整合エンティティとリレーション構造を生成します。対照的に、2 つの代表的な LMM バックボーン全体では、RGL は調理エンティティと手順関係の生成を一貫して改善していますが、私たちの分析ではさらに、きめ細かい食材の状態のキャプチャが依然として最も困難な次元であることが示されています。
原文 (English)
ReGraph: Learning to Generate Recipe Graphs from Food Images
Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a structured transformation process in which ingredients undergo state changes through ordered actions,while free-form recipe language leaves the corresponding entities, intermediate states, and dependencies largely implicit and entangled.A graph representation makes this procedural knowledge explicit and compositional, providing a structured basis for assessing whether model outputs encode process-level knowledge rather than merely presenting plausible textual descriptions. To address this limitation, we present ReGraph, a large-scale recipe graph dataset that represents ingredients, cooking actions, and tools as entities, uses entity attributes to describe ingredient state changes, and employs typed relations to encode manipulation targets, destinations, and procedural ordering. ReGraph further incorporates explicit Recipe Reasoning Chain-of-Thought (RR-CoT) traces, providing auxiliary supervision for procedural decomposition and structured graph generation. Building on ReGraph, we propose Recipe Graph Learning (RGL), a two-stage framework that enables LMMs to generate a plausible fine-grained cooking workflow from a food image in the form of a structured recipe graph. Under a deterministic, schema-aware matching protocol, our experiments reveal a substantial gap between text-generation quality and recoverable procedural structure: recipes produced by existing approaches achieve competitive text-generation scores yet yield limited reference-aligned entity and relation structure under the ReGraph schema. In contrast, across two representative LMM backbones, RGL consistently improves the generation of cooking entities and procedural relations, while our analysis further shows that fine-grained ingredient-state capture remains the most challenging dimension.
ディール・ミー・メイビー: マルチエージェント交渉における感情の役割
LLM エージェントにとって交渉は厳しい社会的任務であり、戦略的推論、説得、対人適応が必要です。しかし、既存のベンチマークはエージェントを感情的に中立なものとして扱い、人間の交渉行動の主要な要因を見逃していることがよくあります。私たちは、即時条件付けされた感情が LLM ベースの価格交渉にどのような影響を与えるかを研究します。管理された枠組みの中で、買い手と売り手のエージェントは 6 つの感情状態のいずれかに個別に割り当てられ、2 つの予算条件の下で 350 を超える実際の消費者製品について交渉します。 36 の感情ペアの設定と広く使用されている 5 つの LLM にわたって、感情が結果を強く左右することがわかりました。怒っている買い手はほとんど合意に達しませんが (成約率 0.39%)、満足している買い手はほとんどの場合同意します (28.91%) が、恐怖を感じている買い手よりも悪い価格で取引されます。感情の影響は役割に依存します。買い手の感情は主に受諾と拒否を引き起こしますが、売り手の感情は譲歩のダイナミクスを形成します。これらの影響は言語だけでなく、取引終了の行動や価格の軌道にも影響を及ぼし、商取引における感情に左右されるエージェントに対する懸念を引き起こしています。
原文 (English)
Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiation
Negotiation is a demanding social task for LLM agents, requiring strategic reasoning, persuasion, and interpersonal adaptation. Yet existing benchmarks often treat agents as emotionally neutral, overlooking a key driver of human bargaining behavior. We study how prompt-conditioned emotions affect LLM-based price negotiation. In a controlled framework, buyer and seller agents are independently assigned one of six emotional states and negotiate over 350 real consumer products under two budget conditions. Across 36 emotion-pair settings and five widely used LLMs, we find that emotions strongly shape outcomes. Angry buyers almost never reach agreement (0.39% deal rate), while happy buyers agree most often (28.91%), but obtain worse prices than fearful buyers. Emotion effects are role-dependent: buyer emotion mainly drives acceptance and rejection, whereas seller emotion shapes concession dynamics. These effects influence not only language, but also termination behavior and price trajectories, raising concerns for emotion-conditioned agents in commerce.
TRIBE: コミュニケーション行動アンサンブルによるチームパフォーマンスの予測
人間のチームを効果的に支援する自律エージェントを設計するには、チームのダイナミクスを理解することが重要ですが、多くの場合、タスク固有の知識は必要ありません。従来のパフォーマンス指標では見えなかったチームの行動力学を明らかにする、ドメインに依存しないアプローチである TRIBE を紹介します。私たちは、コミュニケーション パターンによってチームをパフォーマンスを予測する行動の部族に分類し、タスクに参加する早ければ 10% でタイムリーな介入が可能になることを示しました。私たちは 4 つの多様なデータセットで TRIBE をテストし、コミュニケーション パターンがチームのパフォーマンスを予測する一方で、予測の強さはタスク構造が行動の自由を許容する度合いによって異なることを実証しました。私たちの時間的分析により、AI エージェントがチームの行動の軌道を大幅に変更する一方、人間のアドバイザーは自然のダイナミクスに合わせ、チームはコラボレーションを通じて行動の柔軟性を維持していることが明らかになりました。さらに、TRIBE と Llama を比較してパイプラインを最適化し、パフォーマンスの向上とともに大幅な高速化を実現します。
原文 (English)
TRIBE: Predicting Team Performance via Communication Behavior Ensembles
Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics invisible to traditional performance metrics. We show that communication patterns can categorize teams into performance predictive behavioral tribes, as early as 10% into the task, enabling timely interventions. We test TRIBE on four diverse datasets and demonstrate that communication patterns predict team performance while the prediction strength varies by the degree a task structure allows for behavioral freedom. Our temporal analysis reveals that AI agents significantly alter team behavioral trajectories while human advisors align with natural dynamics, and that teams maintain behavioral flexibility throughout collaboration. Further, we compare TRIBE to Llama and optimize the pipeline, achieving significant speedup with performance improvement.
サイエンス エッジの評価: 真の科学的発見に向けて欠けているステップを確認する
大規模言語モデル (LLM) は科学的発見にますます関与していますが、それらが複雑な実際の実験室科学をサポートできるかどうかは依然として不明です。ここでは、化学、生物学、材料科学における査読済みの文献と実験実践に基づいて、専門家によって精選された質問のマルチモーダルなベンチマークであるサイエンス エッジ評価 (SEE) を紹介します。 19 個のマルチモーダル大規模言語モデル (MLLM) を評価したところ、最もパフォーマンスの高いモデルであっても精度が 48.7% にしか達していないことがわかりました。さらに、汎用モデルは科学特化モデルよりも平均して優れています。視覚エージェントの評価では、ツールを使用すると最高の精度が 52.7% に向上しました。ツールを使用するとモデルで利用できる情報が増える可能性がありますが、情報が増えても信頼できる科学的推論が得られるとは限りません。重要な課題は、モデルが元の実験証拠の範囲内でツール由来の情報を管理できるかどうかです。これらの発見を総合すると、現在の MLLM は、実際の科学的発見に不可欠な能力である、実験結果から正当かつ証拠に基づいた推論を確実に行うことがまだできないことが明らかになります。このギャップを埋めるには、MLLM が確立された科学概念の説明から、実験データから新しい証拠に基づいた洞察を導き出すことに移行する必要があります。
原文 (English)
Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.
極めて重要な投票が見えない: 集計の独立性指標は検証が実際に役立つところを見逃している
LLM 審査員パネルは標準的な評価ツールですが、これまでの研究では、相関性の高いパネルエラーが報告されています。つまり、9 人の審査員が 2 人の独立した審査員の有効な情報をほぼ提供しており、集計によってその差が縮まるのはごく一部に過ぎません。自然療法(テストスイートの実行など、異なる証拠源からの信号)では、パネルの有効投票数に大きな変化は見られませんでした(-0.04、95\% CI [-0.10、+0.02])。集計依存性と条件付き決定ユーティリティは別の問題です。基本的多数決算術は、単一投票の置換のために影響を受けるセットを修正します。変更できるのは 1 票のマージンを持つ決定のみです。経験的な問題は、パネルのエラー率が上昇し、有用な代替品がそこに集中するかどうかです。全体の精度の向上はこれらの重要なクエリに集中しており、精度の向上が大きい (3 つのヘッドライン設定全体で +10.4 ~ +23.3 パーセント ポイント) のですが、その他の部分ではまったくゼロです。 3 つのコード ベンチマークと 4 つのパネル サイズ (9 ジャッジ拡張と 56 の依存サブサンプリング チェック、+6.5 ~ +16.1 パーセント ポイントのゲイン) にわたってパターンを確認しました。 HumanEval+/MBPP+ では、多数派側置換ルールにより全体の精度が 82.44\% から 85.62\% に向上し、クエリの 16.2\% でシグナルが呼び出されます。シグナルのみは 87.60\% で依然として強いです。したがって、人口レベルの依存診断とマージン層別効用は補完的であり、影響を受けるセットの特徴付けにより、指定された単一投票代替政策に対するコール削減ルールが得られます。
原文 (English)
Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps
LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective information of two independent ones, and aggregation closes only a small fraction of the gap. A natural remedy--a signal from a different evidence source, e.g., executing a test suite--produced no distinguishable change in the panel's effective-vote count at scale (-0.04, 95\% CI [-0.10, +0.02]). Aggregate dependence and conditional decision utility are different questions. Elementary majority arithmetic fixes the affected set for single-ballot substitution: only decisions with a one-vote margin can change. The empirical question is whether panel error rates rise and useful substitutions concentrate there. They do: the entire accuracy gain concentrates on these pivotal queries, where it is large (+10.4 to +23.3 percentage points across three headline configurations), and is exactly zero elsewhere. We confirm the pattern across three code benchmarks and four panel sizes (a 9-judge extension and 56 dependent subsampling checks, gain +6.5 to +16.1 percentage points). On HumanEval+/MBPP+, a majority-side replacement rule raises overall accuracy from 82.44\% to 85.62\% while invoking the signal on 16.2\% of queries; signal-only remains stronger at 87.60\%. Thus population-level dependence diagnostics and margin-stratified utility are complementary, and the affected-set characterization yields a call-reduction rule for any specified single-ballot substitution policy.
LMM モダリティ転送: 自律型 GIS エージェントの前提条件
AI モデルは空間情報の理解と処理にますます熟達しているため、空間タスクやワークフローにおけるエージェントによる問題解決が容易になります。しかし、それらの空間能力 (空間推論など) に関する研究のほとんどは、入力および出力としてのテキストのモダリティに焦点を当てています。これは、GIS ワークフローに対する人間のアプローチとは対照的です。GIS ワークフローでは、テキストとビジュアル モダリティが、しばしば一緒に、互換的に、補完的に使用されます。したがって、自動化された GIS 分析パイプラインを真に実現したり、人間が設計した GIS ワークフローを実行したりするには、AI モデル、特に大規模マルチモーダル モデル (LMM) が、そのようなワークフローで従来使用されていた画像ベースのモダリティとテキストベースのモダリティの間をシームレスに移行できる必要があります。モダリティ転送タスクを提示します。(1) LMM に、最初に規則的なグリッド内の色付きの正方形の入力画像を記述するよう依頼し、(2) 新しい LMM インスタンスに、元のモデルによって出力されたテキスト記述を使用して元の空間シーンの画像を再生成するよう依頼します。このタスクは、画像モダリティとテキスト モダリティの間で空間情報を転送する LMM の能力を定量化します。最終的に、この研究は、空間情報理論のレンズを通して LMM のモダリティ転送機能を調べることによって、重大なボトルネックを浮き彫りにします。LMM で強力かつ堅牢な地理空間理解を達成するには、厳密でマルチモーダルな調整が必要です。私たちの結果は、最近の LMM (ここでは OpenAI からのもの) が、色の正方形の単純な空間グリッドの画像を再生成するというタスクを課された場合、モダリティの転送に依然として苦労していることを示しています。
原文 (English)
LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents
AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial tasks and workflows. However, most of the research on their spatial capabilities (e.g., spatial reasoning) has focused on the textual modality as input and output. This contrasts with the human approach to GIS workflows, where text and visual modalities are often used together, interchangeably, and in a complementary manner. Thus, to truly achieve an automated GIS analysis pipeline or carry out human-designed GIS workflows, AI models --- Large Multimodal Models (LMMs) in particular --- need to be able to seamlessly transition between image- and text-based modalities that are traditionally used in such workflows. We present a modality transfer task that (1) asks an LMM to first describe an input image of colored squares in a regular grid, and (2) asks a new LMM instance to re-generate an image of the original spatial scene using the textual description output by the former model. This task quantifies the ability of LMMs to transfer spatial information between image and text modalities. Ultimately, by examining the modality transfer capability of LMMs through the lens of spatial information theory, this work highlights a critical bottleneck: achieving strong and robust geospatial understanding in LMMs requires rigorous, multi-modal alignment. Our results indicate that recent LMMs (here from OpenAI) still struggle with modality transfer, when tasked with re-generating an image of a simple spatial grid of color squares.
トリアージの決定を複数のエージェントに分割すると、バイアスが隠蔽されますか、それともバイアスを発見するのに役立ちますか?監査能力の制約下での LLM ベースのリソース割り当てのマルチエージェント シミュレーション調査
これまでのベンチマーク作業では、生死を分けるリソース割り当ての決定を強いられた単一の大規模言語モデル (LLM) が、測定可能な人口統計上の偏りを示していることが示されています。ただし、実際のデプロイメントでは、単一のエージェントが使用されることはほとんどありません。パイプラインが使用され、この種の障害を正確に検出することを目的としたレビュー手順が行われます。私たちは、同じ決定が 1 つのモデルだけで行われチェックされるのではなく、役割が区別された複数のエージェントのパイプライン (評価、割り当て、独立した監査) に分散された場合にバイアスがどうなるかを研究します。 1 つの人口統計的属性を除いて臨床的に同一であるペアのケースを含む合成災害トリアージ シミュレーターを使用して、GPT-4o-mini で 192 のエピソード (2,304 の解決されたケース ペア) を実行し、3 つの独立して変化する圧力次元の下で、単一エージェントの制御条件と 9 エージェントのパイプラインを比較しました。 2 つの条件間で偏った結果が発生する頻度に測定可能な差はありません (6.9% 対 6.1%、p = 0.498)。バイアスが検出されるかどうかについては、監査能力が大きく重大な影響を及ぼしていることがわかりました。バイアスのある結果の 30.0% は完全に検出されず、監査人の負担が大きい場合は 43.8% に上昇し、そうでない場合は 18.4% に低下します。この効果を分解すると、レビュー対象のケースに対する判断力の低下 (81.6% 対 85.7%、p = 1.000、方向が逆転) ではなく、ほぼ完全にカバレッジ (ケースがそもそもレビューされるかどうか、負荷下で 100.0% から 65.6% に崩壊、p < 0.001) によって引き起こされることがわかります。追跡実験では、先着順ではなく推定リスクに基づいて監査キューの順序を変更すると、同じ容量制約下で失われたカバレッジのほとんどが回復することが示されています (65.6% ~ 91.7%、p = 0.028)。私たちは、リソースの制約の下で LLM エージェント パイプラインに独立した監視を追加するシステムへの影響について議論し、1 つのモデル、適度なサンプル サイズ、敵対的な複製がないという研究の限界を正直に報告します。
原文 (English)
Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints
Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exhibits measurable demographic bias. Real deployments, however, rarely use a single agent: they use pipelines, with review steps meant to catch exactly this kind of failure. We study what happens to bias when the same decision is distributed across a role-differentiated multi-agent pipeline (assessment, allocation, independent audit) instead of made and checked by one model alone. Using a synthetic disaster-triage simulator with paired cases that are clinically identical except for one demographic attribute, we run 192 episodes (2,304 resolved case pairs) on GPT-4o-mini comparing a single-agent control condition to a nine-agent pipeline under three independently varied pressure dimensions. We find no measurable difference in how often biased outcomes occur between the two conditions (6.9% vs. 6.1%, p = 0.498). We do find a large and significant effect of audit capacity on whether bias is caught: 30.0% of biased outcomes go entirely undetected, rising to 43.8% when the auditor is overloaded and falling to 18.4% when it is not. Decomposing this effect shows it is driven almost entirely by coverage (whether a case is reviewed at all, which collapses from 100.0% to 65.6% under load, p < 0.001) rather than by degraded judgment on the cases that are reviewed (81.6% vs. 85.7%, p = 1.000, direction reversed). A follow-up experiment shows that reordering the audit queue by estimated risk, rather than first-come-first-served, recovers most of the lost coverage under the same capacity constraint (65.6% to 91.7%, p = 0.028). We discuss the implications for any system that adds independent oversight to an LLM agent pipeline under resource constraints, and report the study's limitations honestly: one model, modest sample sizes, and no adversarial replication.
大規模言語モデルにおける批評家の評価指向: 映画の好みの引き出しからの証拠
大規模言語モデル (LLM) は、映画、書籍、音楽などに関する人間の判断の表現を含むコーパスでトレーニングされます。しかし、LLM が評価階層を体系的に再現するかどうかはまだ不明です。 LLM の文化的バイアスに関する先行研究では、競合する期待が示唆されています。モデルは、インターネット テキストの人気シグナルを反映している可能性があり、批判的な言説に埋め込まれた威信の形を再現している可能性があります。私たちは、4 つのファミリー (Anthropic、OpenAI、Alibaba、Mistral) の 8 つのモデルを使用した映画評価の研究を通じて、批評家から高い評価を得た映画、商業的に成功した映画、および二重正当性 (批評家からの評価 + 商業的な成功) の映画に分割された 200 本の映画ベンチマークを使用して、この疑問を調査しました。ブラッドリーとテリーの推定を用いて分析されたモデルごとの 20,000 件のペアごとの強制選択比較を通じて、すべてのモデルで一貫した批評家の評価の傾向が観察されました。批評家から高く評価されているが商業的には無名な映画が、商業的には成功しているが批評家に認知されていない映画よりも選ばれています。このパターンは、各ファミリー内のモデルの規模に応じて拡大します。さらに、ネストされた OLS 回帰分析は、評価の方向性、世間の注目度、一般的な受け入れ方が、好みを説明するのに明らかに役立つことを示しています。一般の注目度を調整することで、批評家からの評価のみの映画よりも二重正当性のある映画を好むモデルの傾向が逆転し、さらに人気の受け入れを考慮することで、商業的な成功のみを収めた映画の不利な点の多くが軽減されます。最後に、評価と推奨指向のプロンプト フレーミングは異なるランキングを生成し、現実世界の LLM 展開では批判的な評価の指向が間接的に現れる可能性があることを示唆しています。
原文 (English)
Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation
Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs systematically reproduce evaluative hierarchies remains unclear. Prior research on cultural bias in LLMs suggests competing expectations: models may mirror the popularity signals of internet texts, or may reproduce forms of prestige embedded in critical discourse. We probe this question through a study of film evaluations with eight models from four families (Anthropic, OpenAI, Alibaba, and Mistral), using a 200-film benchmark partitioned into critically acclaimed, commercially successful, and dual-legitimacy (critical acclaim + commercial success) films. Across 20,000 pairwise forced-choice comparisons per model analyzed with Bradley--Terry estimation, we observe a consistent critical acclaim orientation with all models: critically acclaimed yet commercially obscure films are selected over commercially successful yet critically unrecognized ones. This pattern grows with model scale within each family. In addition, nested OLS regression analyses show that evaluative orientation, public visibility, and popular reception distinctly help explain preferences. Adjusting for public visibility reverses the models' preference for dual-legitimacy films over critical acclaim-only films, while additionally accounting for popular reception attenuates much of the disadvantage of films with commercial success only. Finally, evaluative and recommendation-oriented prompt framings produce divergent rankings, suggesting that critical acclaim orientation may manifest indirectly in real-world LLM deployments.
CAi Copilot: インテント駆動型のエージェント ワークフローを通じて分子設計における運用ワークロードを削減
初期段階の分子設計は反復的なプロセスであり、単に分子を生成する作業ではありません。研究者は、幅広い目標を設計戦略に変え、候補を絞り込み、多くの特性を評価し、合成とテストの前に証拠を収集します。 AI メソッドは、分子の生成、いくつかの目標の最適化、特性の予測、化合物のドッキング、および合成の説明を行うことができます。ただし、これらの機能は特殊なツール全体に分散されています。専門家は依然として各ステップを調整し、中間結果を判断し、証拠を統合する必要があります。したがって、中心的な課題は、研究の意図を、科学的ツールに基づいた適応的で追跡可能な実行に変えることです。私たちはこの課題を意図から証拠への分子設計ワークフローの実行として位置づけ、3 つのリンクされたレイヤーを持つ専門家指向のエージェントである CAi Copilot を紹介します。 Research Interface Layer は、意図を実行可能な計画に変換します。エージェント推論レイヤーは、中間結果を使用して各実行をガイドします。実行サブストレートは、分子ツール、メトリクス、再利用可能なユーティリティ、およびバックエンド サービスを提供します。 45 のタスク全体で、CAi は結果スコア 84.59 という最も優れた全体パフォーマンスを達成し、次に良い結果を 18.07 ポイント上回りました。追加のベンチマークでは、CAi が生成、スクリーニング、および複数基準の評価をどのように調整するかをテストし、長期的な実行における限界を明らかにします。これらの結果は、CAi が広範な分子設計の意図を透明で追跡可能なワークフローに変換し、中間決定を候補レベルの証拠に結び付けることを示しています。
原文 (English)
CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows
Early-stage molecular design is an iterative process, not just a task of generating molecules. Researchers turn broad goals into design strategies, refine candidates, assess many properties, and gather evidence before synthesis and tests. AI methods can generate molecules, optimize several goals, predict properties, dock compounds, and account for synthesis. Yet these functions are spread across specialized tools. Experts must still coordinate each step, judge interim results, and integrate evidence. The central challenge is thus to turn research intent into adaptive, traceable runs grounded in scientific tools. We cast this challenge as intent-to-evidence molecular design workflow execution and present CAi Copilot, an expert-oriented agent with three linked layers. The Research Interface Layer turns intent into an executable plan. The Agent Reasoning Layer uses interim results to guide each run. The Execution Substrate supplies molecular tools, metrics, reusable utilities, and backend services. Across 45 tasks, CAi achieves the strongest overall performance, with an outcome score of 84.59, exceeding the next-best result by 18.07 points. Additional benchmarks test how CAi coordinates generation, screening, and multi-criteria evaluation, while exposing limits in long-horizon execution. These results show that CAi turns broad molecular-design intent into transparent, traceable workflows that connect interim decisions to candidate-level evidence.
デールの制約の下でのディープネットワークでの学習
生物学的に妥当な学習モデルは、実際のニューロンの制約下で神経回路がどのように効果的な学習を実装できるかを説明することを目的としています。大きな進歩は見られましたが、依然として大きな課題は、既存のモデルでは、皮質回路の基本的な側面に違反して、ニューロンまたはシナプスが正と負の両方の混合符号値を表すことを許可していることがよくあります。これは、デールの制約、つまり生物学的ニューロンは興奮性または抑制性のいずれかですが、両方ではなく、シナプスは符号を変更できないという制約に違反しています。この研究では、バックプロパゲーションのような学習をサポートしながら、神経の活性化と学習信号の両方が非負の活動によって表され、シナプスが固定符号を持つ、生物学的に動機付けられたニューラル アーキテクチャを導入することで、この不一致に対処します。私たちのアプローチは、脳内のオンオフ表現の証拠にヒントを得て、2 つの相補的に相互作用する非負のチャネルを使用して正と負の寄与を表します。これらのチャネルは、単純な神経回路モチーフを通じて実装され、ボトムアップ経路とトップダウン経路の両方でネットワーク全体で繰り返されます。ローカル ヘビアン学習ルールと組み合わせた結果のモデルは、ニューロン間のローカルな相互作用のみを使用して学習信号を伝播し、重みを更新します。我々の学習スキームは、非負の誤差信号のみに依存しているにもかかわらず、バックプロパゲーションの更新を正確に回復できることを理論的に示します。経験的に、オンオフ アーキテクチャは、より強力な生物学的制約を満たすだけでなく、効率的な表現を学習し、Tiny ImageNet ベンチマークで同等のバニラ ネットワークよりも大幅な向上をもたらします。これらの結果は、混合符号信号を必要とせずに生物学的にもっともらしいメカニズムから効果的な学習が実現できることを示しており、より現実的なニューラル計算モデルへの一歩を提供します。
原文 (English)
Learning in Deep Networks under Dale's Constraint
Biologically plausible learning models aim to explain how neural circuits can implement effective learning under the constraints of real neurons. Although significant progress has been made, a major remaining challenge is that existing models often allow neurons or synapses to represent mixed-sign values, both positive and negative, in violation of a basic aspect of cortical circuitry -- Dale's constraint: biological neurons are either excitatory or inhibitory, but not both, and synapses cannot change sign. In this work, we address this discrepancy by introducing a biologically motivated neural architecture in which both neural activations and learning signals are represented by non-negative activity, and synapses have fixed sign, while still supporting backpropagation-like learning. Our approach uses two complementary interacting non-negative channels to represent positive and negative contributions, inspired by evidence of on-off representations in the brain. These channels are implemented through a simple neural circuit motif, which is repeated throughout the network in both bottom-up and top-down pathways. Combined with a local Hebbian learning rule, the resulting model propagates learning signals and updates weights using only local interactions between neurons. We show theoretically that our learning scheme can exactly recover the backpropagation update despite relying solely on non-negative error signals. Empirically, beyond satisfying stronger biological constraints, the on-off architecture learns efficient representations, yielding substantial gains over comparable vanilla networks on the Tiny ImageNet benchmark. These results demonstrate that effective learning can emerge from biologically plausible mechanisms without requiring mixed-sign signals, providing a step toward more realistic models of neural computation.
タイル化された SVD で使用可能な重みメカニズムを見つける
機械的な解釈可能性への主要なアプローチは、スパース オートエンコーダーなどのプロキシ辞書をトレーニングし、最大アクティブ化テキストから特徴をラベル付けします。このような最良のアトラスは概念を識別しますが、そのアイデンティティはネットワークの重み付け自体ではなく、学習された辞書の中に存在します。列タイル SVD によって線形サイトからメカニズム マウントを直接抽出することを提案します。各マウントは、トリガー、書き込み、および強度として読み取られるトリプル (v,u,{\sigma}) です。アイデンティティは重みのルールです。タイルローカルリフトではなくフル書き込みエネルギーリフトで判断される事前登録スイートを使用してマウントを評価します。 WikiText-2 (16,384 トークンのサブサンプル) を使用した Gemma-2-2B では、7 つの線形マップすべてがスコア付けされます。残留書き込み (mlp.down、attn.o) は、サブレイヤー RMSNorm 後のステアで完全な A/B/C を受け取り、52/52 サイトレイヤーを通過します。他のマップは A/B のみを受信します (mlp.gate/attn.q/attn.k/Effective mlp.up/attn.v それぞれ 26/26)。集計: 182/182 GO。ライブラリ コード、コーパス ビルダー、実験エントリポイント、単体テストをリリースします。
原文 (English)
Finding Usable Weight Mechanisms with Tiled SVD
The dominant approach to mechanistic interpretability trains proxy dictionaries such as sparse autoencoders and labels features from max-activating text. The best such atlases identify con- cepts, but that identity lives in the learned dictionary rather than in the network weights them- selves. We propose extracting mechanism mounts directly from linear sites by column-tiled SVD: each mount is a triple (v,u,{\sigma}) read as trigger, write, and strength. Identity is the weight rule. We evaluate mounts with a pre-registered suite judged on full-write energy lift rather than tile-local lift. On Gemma-2-2B with WikiText-2 (16,384-token subsample), all seven linear maps are scored: residual writes (mlp.down, attn.o) receive full A/B/C with steer after post-sublayer RMSNorm and pass 52/52 site-layers; other maps receive A/B only (mlp.gate/attn.q/attn.k/effective mlp.up/attn.v 26/26 each). Aggregate: 182/182 GO. We release library code, the corpus builder, the experiment entrypoint, and unit tests.
FedLBW: A Loss-Based Weighting Strategy for Federated Learning on Non-IID Data in Wireless Networks
Federated Learning (FL) enables collaborative machine learning (ML) across distributed clients while preserving privacy. However, efficient…
ReQuant: トレーニング後の量子化のための固定グリッド離散リファインメント
ポストトレーニング量子化 (PTQ) は、大規模な言語モデルのメモリと計算コストを削減するために広く使用されています。既存の PTQ 手法は通常、ヒューリスティック ルールまたは貪欲な最適化を通じて初期の量子化モデルを取得し、量子化が完了すると、結果として得られる整数割り当ては通常最終的なものとして扱われます。この観察により、量子化形式を維持しながら、実行可能な量子化モデルが生成された後に量子化重みを改善可能な状態に保つ PTQ 内の補完的な最適化ステージが動機付けられます。この段階では、バックプロパゲーションのない固定グリッド改良手順である ReQuant を導入します。 PTQ イニシャライザに依存せず、ReQuant は既存の量子化モデルを実行可能な開始点として使用し、固定量子化グリッド上でその離散重みの割り当てを繰り返し再検討します。受け入れられた更新は、平均二乗再構成誤差を厳密に削減し、元のグリッドに残ります。このように、ReQuant は、最初に固定された PTQ 出力を反復的に最適化可能な個別のソリューションに変換し、既存の PTQ パイプラインのプラグ アンド プレイの後処理ステージとして機能します。多様なモデル ファミリ、ビット幅、およびダウンストリーム タスクにわたる実験では、ReQuant が異種 PTQ イニシャライザから量子化モデルを一貫して改善し、特に単純なイニシャライザとより低いビット幅で大きな改善が見られることが示されています。特に、ReQuant は、同じ量子化形式で GPTAQ に近づくか超えるまで、複数のスイープにわたる単純な最近接への丸め初期化を改良できます。これらの結果により、ReQuant は既存の PTQ パイプラインをさらに改善するための実用的な補完ステージとして確立されます。
原文 (English)
ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization
Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods typically obtain an initial quantized model through heuristic rules or greedy optimization, and once quantization is completed the resulting integer assignments are usually treated as final. This observation motivates a complementary optimization stage within PTQ that keeps quantized weights improvable after an executable quantized model has been produced, while preserving the quantized format. We introduce ReQuant, a backpropagation-free fixed-grid refinement procedure for this stage. Agnostic to the PTQ initializer, ReQuant takes an existing quantized model as a feasible starting point and iteratively revisits its discrete weight assignments on the fixed quantization grid. Accepted updates strictly reduce the mean squared reconstruction error and remain on the original grid. In this way, ReQuant turns the initially fixed PTQ output into an iteratively optimizable discrete solution and serves as a plug-and-play post-processing stage for existing PTQ pipelines. Experiments across diverse model families, bit-widths, and downstream tasks show that ReQuant consistently improves quantized models from heterogeneous PTQ initializers, with especially large gains on simple initializers and lower bit-widths. Notably, ReQuant can refine a simple round-to-nearest initialization across multiple sweeps until it approaches or surpasses GPTAQ under the same quantization format. These results establish ReQuant as a practical complementary stage for further improving existing PTQ pipelines.
ZIPBrain: EEG 基盤モデルはより高速でローカルに導入可能でありながら正確である可能性がありますか?
この研究では、精度を犠牲にすることなく、脳波計 (EEG) 基礎モデル (EFM) を高速化し、ローカルに展開可能にすることができるかどうかを調査します。 EEG 基礎モデルは主要なトレンドであり、強力な汎用表現を提供します。ただし、計算負荷は入力長に応じて二次関数的に増大するため、リソースに制約のあるシナリオ、特にリアルタイムの臨床モニタリングでの展開が妨げられます。 EEG の低い SNR は、これらのトークンの多くが冗長であり、精度をほとんど犠牲にせずに圧縮可能であることをさらに示唆しています。私たちは、この低 SNR 特性を活用してトークン数を削減する、新しい冗長性を意識した EEG トークン プーリング モジュールである ZIPBrain を提案します。トークン シーケンスが与えられると、ZIPBrain はトークンを冗長な一意のグループに分割し、各冗長トークンを一意のグループ内の最も類似したトークンとマージします。さらに、ZIPBrain は、トレーニング不要のプラグアンドプレイ モジュールとして機能し、ごくわずかな計算オーバーヘッドで標準の Transformer エンコーダにシームレスに統合します。複数のEEG基礎モデルにわたる広範な実験により、ZIPBrainの強力な多用途性が示され、元のEEG基礎モデルと比較して実時間推論時間を32.7%(CUDAグラフでは最大41.8%)削減しながら、ベースラインに対して平均1.3%~10.5%の改善を達成しました。
原文 (English)
ZIPBrain: Can EEG Foundation Models Be Faster, Locally Deployable, but Accurate?
This work investigates whether Electroencephalograph (EEG) foundation models (EFMs) can be made faster and locally deployable without sacrificing accuracy. EEG foundation models are a major trend, offering strong general-purpose representations. However, their computational burden grows quadratically with input length, hindering deployment on resource-constrained scenario, particularly for real-time clinical monitoring. EEG's low SNR further suggests many of these tokens are redundant and compressible with little accuracy cost. We propose ZIPBrain, a novel redundancy-aware EEG token pooling module that leverages this low-SNR characteristic to reduce token count. Given a token sequence, ZIPBrain partitions tokens into redundant and unique groups, then merges each redundant token with its most similar counterpart in the unique group. Furthermore, ZIPBrain serves as a training-free, plug-and-play module that seamlessly integrates into standard Transformer encoders with negligible computational overhead. Extensive experiments across multiple EEG foundation models show ZIPBrain's strong versatility, achieving 1.3%-10.5% average improvement over baselines, while reducing wall-clock inference time by 32.7% (up to 41.8% with CUDA Graph) compared to the original EEG foundation models.
Not All Problems Are Best Modeled as MILP: A DSL-Centric Framework for Flexible and Accurate Optimization Modeling
Solving combinatorial optimization problems (COPs) requires not only efficient algorithms but also carefully crafted formulations. While re…
PDE 基礎モデルの教師なし適応
事前トレーニングされた偏微分方程式 (PDE) 基礎モデルは、さまざまな方程式にわたって一般化できますが、これを目に見えない PDE システムに適応させるには通常、高密度の解データが必要であり、高価であるか入手できないことがよくあります。この制限に対処するために、グラウンドトゥルース ソリューションの必要性を排除する、教師なし PDE ベースの微調整フレームワークを提案します。まず、さまざまな空間スケールにわたる多様な時間依存偏微分方程式で近隣アテンション トランスフォーマーを事前学習し、異種方程式全体で転送可能な表現を生成します。適応段階では、PDE の残差条件と境界条件を使用して物理ベースの目標を構築し、低ランク適応 (LoRA) を通じて目に見えない方程式に基づいてモデルを微調整します。標準 LoRA における物理量全体にわたる不均一な学習に対処するために、適応のバランスを再調整するニュートン・シュルツ直交化バリアントである NSLoRA を導入します。私たちの手法は、グラウンドトゥルース ソリューションを必要とせずに、教師あり LoRA 微調整に匹敵するパフォーマンスを達成しながら、複数の空間次元にわたる異種 PDE ベンチマーク全体で、競合するニューラル オペレーター ベースラインや最近の PDE 基礎モデルを常に上回っています。
原文 (English)
Unsupervised Adaptation of PDE Foundation Models
Pretrained partial differential equation (PDE) foundation models can generalize across different equations, but adapting them to unseen PDE systems typically requires dense solution data, which is often expensive or unavailable. To address this limitation, we propose an unsupervised PDE-based finetuning framework that eliminates the need for ground-truth solutions. We first pretrain a neighborhood attention Transformer on diverse time-dependent PDEs spanning varying spatial scales, yielding transferable representations across heterogeneous equations. In the adaptation stage, we construct a physics-based objective using the PDE residual and boundary conditions, and finetune the model on unseen equations via low-rank adaptation (LoRA). To address the uneven learning across physical quantities in standard LoRA, we introduce NSLoRA, a Newton-Schulz orthogonalized variant that rebalances adaptation. Our method achieves performance comparable to supervised LoRA finetuning without requiring any ground-truth solutions, while consistently outperforming competitive neural operator baselines and recent PDE foundation models across heterogeneous PDE benchmarks spanning multiple spatial dimensions.
BONSAI: スキルによる進化に基づくツリー検索
スキルは、重みを更新できない凍結されたエージェントを操作する自然言語ドキュメントであるため、エージェントに不足している機能は散文で提供する必要があります したがって、スキルの最適化は、スコアに対してテキストを最適化することであり、ホールドアウトスコアを上げる編集を維持する標準レシピは、特定の点でブラインドです 単一のスコアでは、狭いオーバーフィットスパイクに載っているドキュメントと広い台地に載っているドキュメントを区別することはできません たとえ2番目だけが改善できるとしても、私たちはBONSAIという新しいスキル最適化フレームワークを紹介します。代わりに、進化可能性によってドキュメント空間の領域の容量を制御し、さらなる突然変異の下で実行可能なバリエーションを生成し続けます。特性生物学は現在の適応度とは別のものとして扱います。BONSAI は、すべての子ドキュメントがその親の突然変異であり、その活用条件がスキル自身の適応度とその突然変異近傍の適応度を混合する上位信頼性選択ルールの下でそれを下位させるモンテカルロ検索ツリーとしてスキルを成長させます。すべての子は突然変異であるため、ノードの下に記録された平均スコアは、近傍の進化可能性を推定します。追加コストがないため、ルールは探索期間が現在弱い競合ブランチを維持しながら改善を続ける領域に予算を集中させます。 BONSAI は、見つかった単一の最高スコアのドキュメントを、置き換える acceptifbetter ループを超えて無料で出荷します。 凍結された 30B エージェントを使用し、3 つのベンチマークで平均化されます。 BONSAI は、スキルフリー エージェントよりもホールドアウト精度を 2313 ポイント向上させ、予算に一致する 2 つのベースライン GEPA と SkillOpt を 387 ポイント向上させます。各397点
原文 (English)
BONSAI: Evolvability-Guided Tree Search over Skills
A skill is a naturallanguage document that steers a frozen agent whose weights cannot be updated so any capability the agent lacks must be supplied in prose Optimising a skill is therefore optimising text against a score and the standard recipe which keeps any edit that raises a heldout score is blind in a specific way a single score cannot tell a document perched on a narrow overfit spike from one resting on a broad plateau even though only the second can still be improved We introduce BONSAI a novel skilloptimisation framework that steers instead by evolvability the capacity of a region of documentspace to keep producing viable variation under further mutation a property biology treats as separate from present fitness BONSAI grows skills as a MonteCarlo search tree in which every child document is a mutation of its parent and descends it under an upperconfidence selection rule whose exploitation term blends a skills own fitness with the fitness of its mutational neighbourhood Because every child is a mutation the mean score recorded beneath a node estimates that neighbourhoods evolvability at no extra cost so the rule concentrates budget on regions that keep improving while its exploration term keeps a currently weak branch in contention BONSAI ships the single bestscoring document it finds at no cost beyond the acceptifbetter loop it replaces With a frozen 30B agent and averaged over three benchmarks BONSAI lifts heldout accuracy over the skillfree agent by 2313 points and improves on two budgetmatched baselines GEPA and SkillOpt by 387 and 397 points respectively
PTQ4SNN: スパイキング ニューラル ネットワークのためのメンブレンを意識したトレーニング後の量子化
スパイキング ニューラル ネットワーク (SNN) により、スパースでイベント駆動型の計算が可能になりますが、反復膜状態は通常、重み量子化後でも浮動小数点で保持されるため、その低ビット展開は不完全なままです。これらの状態の量子化は、その分布がチャネル間で、また以前の重みと異なるため、困難ですが、発火閾値近くの小さな摂動がスパイクの決定を変更し、時間の経過とともに蓄積する可能性があるためです。我々は、小さなキャリブレーションセットのみを使用して重みと反復膜状態を共同量子化する膜認識ポストトレーニング量子化フレームワークである PTQ4SNN を提案します。まず、チャネルごとの統一スケール ブリッジは膜スケールを s_mem,c = s_w,c * 2^k_c として制約し、シフト互換のスケール変換を可能にしながら膜分布に適応させます。第 2 に、混合精度ビット割り当てでは、平均ビット バジェットの下で、発火アクティビティと量子化感度に応じて 2/4/8 ビット精度を膜チャネルに割り当てます。このフレームワークは再利用可能なプロジェクション LIF ペアで動作し、バックボーンの再トレーニングなしで畳み込み SNN とスパイク駆動のトランスフォーマーの両方をサポートします。静的およびイベントベースの分類とセマンティック セグメンテーションに関する実験では、PTQ4SNN が W4 量子化下でのモデル精度と約 4 ビットのメンブレン精度を効果的に維持することが示されています。
原文 (English)
PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks
Spiking neural networks (SNNs) enable sparse and event-driven computation, but their low-bit deployment remains incomplete because recurrent membrane states are commonly retained in floating point even after weight quantization. Quantizing these states is challenging because their distributions differ across channels and from the preceding weights, while small perturbations near the firing threshold may alter spike decisions and accumulate over time. We propose PTQ4SNN, a membrane-aware post-training quantization framework that jointly quantizes weights and recurrent membrane states using only a small calibration set. First, a channel-wise Unified Scale Bridge constrains the membrane scale as s_mem,c = s_w,c * 2^k_c, adapting to membrane distributions while enabling shift-compatible scale conversion. Second, Mixed-Precision Bit Allocation assigns 2/4/8-bit precision to membrane channels according to firing activity and quantization sensitivity under an average-bit budget. The framework operates on reusable projection-LIF pairs and supports both convolutional SNNs and spike-driven Transformers without backbone retraining. Experiments on static and event-based classification and semantic segmentation show that PTQ4SNN effectively preserves model accuracy under W4 quantization and approximately 4-bit membrane precision.
DocMemo: マルチモーダル文書理解のための確率的記憶に基づく検索による動的な証拠発見
長い文書を理解するには、数百ページにわたるまばらで異種の証拠を見つける必要がありますが、既存のシステムは静的な検索と脆弱なクロスラウンドメモリによって制限されたままです。主流のシングルラウンド手法は、最初に固定の上位 $k$ ページにコミットし、早期の取得エラーから回復するのに苦労します。最近の反復アプローチでは、複数ラウンドの証拠を取得できますが、クロスラウンド状態の伝播メカニズムは調査されないため、ページの関連性の動的な変化を追跡することが困難になります。これらの制限に対処するために、私たちは動的な証拠探索として長い文書推論を定式化する記憶主導型フレームワークである DocMemo を提案します。 DocMemo は、ドキュメント スキーマ メモリ、ページ ビリーフ メモリ、および質問エピソード メモリで構成される 3 レベルの検索状態を維持します。これらはそれぞれ、構造事前分布、動的関連性推定、およびクエリ固有の推論軌跡をキャプチャします。推論中、DocMemo は、トンプソン サンプリング、空間近接伝播、および構造を認識した適応粒度証拠アクセスによるベイジアン ページ信念更新を通じてクロスラウンド ページ選択を継続的に洗練し、同時に、きめの細かい視覚領域でページ レベルの証拠を補完します。 3 つのベンチマークの実験では、DocMemo が最先端のパフォーマンスを達成し、構造化メモリと動的なページ ビリーフ更新の有効性を検証したことが示されています。コードは https://github.com/Harrygof/DocMemo で入手できます。
原文 (English)
DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding
Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static retrieval and fragile cross-round memory. Mainstream single-round methods commit to a fixed top-$k$ page set at the outset and struggle to recover from early retrieval errors; recent iterative approaches allow multi-round evidence acquisition, but they do not investigate the propagation mechanism of cross-round states, making it difficult to track the dynamic changes in page relevance. To address these limitations, we propose DocMemo, a memory-guided framework that formulates long-document reasoning as dynamic evidence exploration. DocMemo maintains a tri-level retrieval state consisting of Document Schema Memory, Page Belief Memory, and Question Episodic Memory, which respectively capture structural priors, dynamic relevance estimation, and query-specific reasoning trajectories. During reasoning, DocMemo continuously refines cross-round page selection through Bayesian page belief updating with Thompson sampling, spatial proximity propagation, and structure-aware adaptive-granularity evidence access, while supplementing page-level evidence with fine-grained visual regions. Experiments on 3 benchmarks show that DocMemo achieves state-of-the-art performance and validate the efficacy of structured memory and dynamic page belief updating. Code is available at https://github.com/Harrygof/DocMemo.
MemOPD: 長期エージェントのメモリ状態アライメントによるポリシーに基づく蒸留
長期的なエージェントは対話中に成長するコンテキストを蓄積し、パフォーマンスと安定性を損ないます。コンパクト メモリは、モデル呼び出しの間に保持される履歴を圧縮して書き換えることにより、この問題を軽減します。何を保持するかを学習するには、通常、最終タスクの報酬を伴う近接ポリシー最適化 (PPO) に依存しますが、報酬がまばらであるため、個々のメモリの更新に関するガイダンスはほとんど提供されません。この制限により、ポリシー蒸留 (OPD) が動機付けられ、生徒のロールアウトに対して教師による綿密な監督が提供されます。このような監視が有効であるためには、教師はサンプリングされた各アクションを、それが生成されたときと同じ状態で評価する必要があります。ただし、メモリ圧縮中に実行されるコンテキストの書き換えにより、この調整が崩れる可能性があります。サンプリングされた応答が保持され、後の呼び出しのために再エンコードされる場合、インタラクションを永続的な履歴に平坦化すると、ロールアウト中に生徒が一度も訪れなかった状態で教師がアクションをスコアリングする可能性があります。したがって、この措置は出所ごとに政策に基づいたままになりますが、必ずしも州ごとではありません。そこで私たちは、Memory-Aligned On-Policy Distillation (MemOPD) を提案します。 MemOPD は、各モデル呼び出しの入力とサンプリングされた出力を記録し、元のトークンの位置と因果関係の可視性を復元し、教師の効率的なスコアリングのために再構築された呼び出しをパックします。教師はサンプリングされた動作位置で完全な語彙指導を提供しますが、PPO は最終的なタスクの目標を維持します。実験では、いくつかのコンテキスト更新にわたる状態の整合性を検証し、一致するコントロールでの永続履歴教師スコアリングよりも F1 が 7.0% 向上することを示しています。全体として、MemOPD-3B は PPO に対して F1 を最大 416.2% 改善し、パッキングによりトレーニング中のアクターの計算が最大 1.63 倍高速化されます。この作業のコードは、https://github.com/TPssp/MemOPD で公開されています。
原文 (English)
MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents
Long-horizon agents accumulate growing contexts during interaction, impairing performance and stability. Compact memory mitigates this problem by compressing and rewriting the history retained between model invocations. Learning what to retain typically relies on proximal policy optimization (PPO) with final task rewards, but sparse rewards provide little guidance for individual memory updates. This limitation motivates on-policy distillation (OPD), which supplies dense teacher supervision on student rollouts. For such supervision to be valid, the teacher must evaluate each sampled action under the same state in which it was generated. However, the context rewriting performed during memory compression can break this alignment. When sampled responses are retained and re-encoded for later invocations, flattening the interaction into a persistent history may cause the teacher to score the action under a state that the student never visited during rollout. The action therefore remains on-policy by provenance, but not necessarily by state. We therefore propose Memory-Aligned On-Policy Distillation (MemOPD). MemOPD records the inputs and sampled outputs of each model invocation, restores its original token positions and causal visibility, and packs the reconstructed invocations for efficient teacher scoring. The teacher provides full-vocabulary supervision at the sampled action positions, while PPO preserves the final task objective. Experiments verify state alignment across several context updates and show that it improves F1 by 7.0% over persistent-history teacher scoring in a matched control. Overall, MemOPD-3B improves F1 over PPO by up to 416.2%, while packing yields up to a 1.63x speedup in actor computation during training. The code for this work is publicly available at: https://github.com/TPssp/MemOPD.
Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking
The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models…
MemWM: メモリ拡張されたテキストベースの世界モデル
エージェントのアクションに応じて環境状態がどのように変化するかを予測することで、エージェントの計画をサポートするためにワールド モデルがますます使用されています。しかし、流暢な次状態の予測では、タスクに不可欠な事実が省略されたり、製品属性が破損したり、誤った移行ルールが適用されたりする可能性があります。このような系統的な予測エラーに対処するために、メモリ拡張されたテキストベースの世界モデルである MemWM を導入します。 MemWM は、遷移ルール、状態キャッシュ、および予測困難な事実の厳選されたメモリ バンクであるワールド メモリを使用して、次の状態の想像力を条件付けします。ベンチマーク固有の事実とフィールドを通じて予測された状態をスコアリングする構造化状態忠実度 (SSF) を使用して、事実の状態の保存を評価します。 SFT と比較して、記憶増強トレーニングは SSF を最大 206.3% 向上させます。完全な計画設定では、政策モデルを凍結したままにして、政策側の世界スキル、つまり取得されたタスクレベルのスキルとアクション選択のための段階的な修正ガイダンスを提供します。 ALFWorld、WebShop、ScienceWorld 全体で、メモリ拡張エージェントは、SFT でトレーニングされたワールド モデル エージェントよりも下流での成功を向上させ、相対的に最大 65.4% の向上を実現します。さらに、感度分析により、さまざまなメモリおよびアクション予算設定の下で、取得されたメモリによりタスクの成功と効率が向上することが示されています。
原文 (English)
MemWM: Memory-Augmented Text-Based World Model
World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can still omit task-critical facts, corrupt product attributes, or apply incorrect transition rules. To address such systematic prediction errors, we introduce MemWM, a memory-augmented text-based world model. MemWM uses world memory, a curated memory bank of transition rules, state caches, and hard-to-predict facts, to condition next-state imagination. We evaluate factual state preservation with Structured State Fidelity (SSF), which scores predicted states through benchmark-specific facts and fields. Compared with SFT, memory-augmented training improves SSF by up to 206.3%. In the full planning setting, we keep the policy model frozen and provide policy-side world skill: retrieved task-level skills and step-wise corrective guidance for action selection. Across ALFWorld, WebShop, and ScienceWorld, memory-augmented agents improve downstream success over an SFT-trained world-model agent, with up to a 65.4% relative gain. Sensitivity analyses further show that retrieved memory improves task success and efficiency under different memory and action-budget settings.
いくら、どこで: マルチターンエージェント強化学習のためのクレジットを節約するアクションからトークンへの割り当て
マルチターン エージェント強化学習におけるクレジットの割り当ては、アクションに軌道レベルのクレジットを割り当てることと、各アクションのクレジットをそのトークン全体に分配することの 2 つのレベルで動作します。本稿では、これらの意思決定を分離する FACTOR を紹介します。 FACTOR は、チェックポイントで調整された TD 残差を使用して、軌道の利点に伸縮するアクションごとのクレジットを割り当て、フィードバック条件付き教師と生徒の尤度ギャップを使用して、実現されたアクション トークン全体に各クレジットを割り当てます。アクションごとの正規化により、アクションの平均係数が保存され、トークンレベルの符号反転が防止されます。この構造をアクション平均削減と組み合わせて、アクションのスカラー代理重みのトークン長への暗黙の依存性を除去します。動作ポリシーおよびクリッピング前では、各アクションの内部アクション平均サロゲートは TD クレジットと等しくなります。 FACTOR は、ALFWorld、WebShop、ScienceWorld 全体で競合ベースラインよりも一貫して向上しており、どの環境シード比較でも FACTOR が有利であり、最も長い期間の環境で最大の利益が得られています。同じハイパーパラメータは、より大きなバックボーンや別のモデル ファミリに再調整することなく転送されます。アブレーションでは、TD アクション クレジットが改善の主要な推進力であることが特定され、後知恵のトークン割り当てが補完的な利益に貢献します。
原文 (English)
How Much, Then Where: Credit-Conserving Action-to-Token Allocation for Multi-Turn Agent Reinforcement Learning
Credit assignment in multi-turn agent reinforcement learning operates at two levels: assigning trajectory-level credit to actions and distributing each action's credit across its tokens. In this paper, we introduce FACTOR, which separates these decisions. FACTOR uses checkpoint-calibrated TD residuals to assign per-action credits that telescope to the trajectory advantage, and feedback-conditioned teacher-student likelihood gaps to allocate each credit across the realized action tokens. Per-action normalization preserves the action-average coefficient and prevents token-level sign flips. We pair this construction with an action-mean reduction, removing the implicit dependence of an action's scalar surrogate weight on its token length. At the behavior policy and before clipping, each action's inner action-mean surrogate equals its TD credit. FACTOR consistently improves over competitive baselines across ALFWorld, WebShop, and ScienceWorld, with every environment-seed comparison favoring FACTOR and the largest gains emerging on the longest-horizon environment. The same hyperparameters transfer without retuning to a larger backbone and to a different model family. Ablations identify TD action credit as the dominant driver of the improvement, with hindsight token allocation contributing complementary gains.
DiDPO: コーディング エージェント トレーニングのための Diff-in-Diff ポリシーの最適化
検証可能な報酬を伴う強化学習 (RLVR) は、コーディング エージェントをトレーニングするための強力なパラダイムとして台頭しており、コンパイルとテストからの実行フィードバックによって客観的な検証が可能になります。ただし、エージェントのタスクとは異なり、コーディング エージェントは、よりきめの細かい独自のクレジット割り当ての課題に直面しています。各ステップで、コーディング アクションによってさまざまな変更がコード バージョンの異なる領域に同時にパックされるため、独立した変更の寄与が区別できなくなります。既存の RLVR メソッドは主に結果報酬またはステップレベル報酬を利用していますが、これではコードの差分を掘り下げることができず、コーディング アクションの固有のプロパティがトレーニングから見えなくなります。この論文では、コード diff の構造から直接、きめの細かいクレジット ユニットを構築する批判のない RL 手法である Diff-in-Diff Policy Optimization (DiDPO) を提案します。 DiDPO は、マルチターンのコーディングインタラクションを複数の思考と行動のステップに編成し、サンプリングされた軌跡全体でのコードの差分を検出します。次に、「グループ化可能性スコア」によって各差分全体から分割された非常に類似したサブ差分を集約することによってアンカーを選択します。これにより、アンカーの意味論的範囲とアンカーが形成する可能性のあるグループの質量のバランスを最適化する分割スキーマが提供されます。最後に、これらのアンカーは利点グループを形成し、差分レベルの利点を個々の応答トークンに投影します。長期的なコーディングと推論のベンチマークに関する実験では、DiDPO が強力なエージェント RL ベースラインを大幅に上回るパフォーマンスを示しています。 Qwen2.5-7B-Coder では、DiDPO は同等の手法を 10\% 以上上回り、はるかに大規模なモデルとの差を縮め、コーディング エージェントのトレーニングにおけるきめ細かい単位の割り当てのための原則に基づいたフレームワークを提供します。また、さまざまな RL メソッドとコーディング ベンチマークをサポートするエージェント rl コードベースである verl-code もオープンソースにしています。
原文 (English)
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training
Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution feedback from compilation and tests provides objective verification. However, unlike agent tasks, coding agents face a unique and finer-grained credit assignment challenge: at each step, coding actions simultaneously pack varying changes into different regions of a code version, which makes the contribution of independent change indistinguishable. Existing RLVR methods mostly leverage the outcome reward or step-level reward, which fails to dive into a code diff and makes unique properties of coding actions invisible to training. In this paper, we propose Diff-in-Diff Policy Optimization (DiDPO), a critic-free RL method that constructs fine-grained credit units directly from the structure of code diffs. DiDPO organizes multi-turn coding interactions into multiple thought--action steps and discovers code diffs across sampled trajectories. It then selects anchors by aggregating highly similar sub-diffs split from each whole diff by our ``groupability score'', which provides the splitting schema that optimally balances the semantic scope of anchors and the group mass they may form. Finally these anchors form advantage groups and project the diff-level advantage back to individual response tokens. Experiments on long-horizon coding and reasoning benchmarks show that DiDPO significantly outperforms strong agentic RL baselines. On Qwen2.5-7B-Coder, DiDPO exceeds comparable methods by over 10\% and narrows the gap with far larger models, offering a principled framework for fine-grained credit assignment in coding agent training. We also open-source verl-code, an agentic rl codebase that supports various RL methods and coding benchmarks.
スマート マニュファクチャリングにおける大規模言語モデル拡張のための MARL 中心のリファレンス アーキテクチャ
現代の製造業では、適応制御に対して 6 つの複合的な要求が課せられています。それは、全体的な影響を伴う局所的な決定、部分的な観測可能性、非定常性、長期効果を伴う反射速度応答、遅れて拡散する結果、および明示的なモデリングに抵抗するダイナミクスです。協調型マルチエージェント強化学習 (MARL) は、分散実行による集中トレーニングの下で Dec-POMDP として提示され、これらの要求に対して特に自然な形式主義です。この文書では、MARL を中心とした範囲を採用し、大規模言語モデル (LLM) がどこでその調整コアを強化、インターフェース、トレーニング、または最も強力な競合ケースで置き換えるべきかを問います。分類法では、ポリシー、報酬設計、エージェント間のコミュニケーション、階層計画という 4 つの LLM 接続ポイントを通じて文献を整理します。条件付き機能プロファイルは、ネイティブ メカニズム、報告されたパフォーマンス、正式な保証、エンジニアリングの成熟度を分離し、展開準備状況分析により、各役割の背後にある証拠を特定します。これらの段階は主な貢献をもたらします。つまり、証拠に基づいた、意味論的推論、適応型協調制御、および独立して保証された実行のための 3 層の MARL 中心の参照アーキテクチャです。 LLM-Augmented Dec-POMDP は、そのアーキテクチャの説明的な比較表記法であり、新しい決定プロセス クラスやアルゴリズムを導入することなく 4 つの添付ファイルの選択肢を記録します。レビューされた証拠によると、従来の MARL は、タスク固有のトレーニング後の頻繁で構造化された分散型調整に適しているのに対し、LLM コンポーネントは意味解釈、報酬草案、人間的対話、および時間のかかる監督計画には有望です。現在の LLM のみの製造コントローラーは、厳密なリアルタイム、分散型、安全性が重要な制御との同等性をまだ確立していません。この結論は入手可能な証拠に限定されており、不可能であると主張するものではありません。
原文 (English)
A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing
Modern manufacturing imposes six coupled demands on adaptive control: local decisions with global consequences, partial observability, nonstationarity, reflex speed response with long horizon effects, delayed and diffuse outcomes, and dynamics that resist explicit modeling. Cooperative multiagent reinforcement learning (MARL), posed as a Dec-POMDP under centralized training with decentralized execution, is a particularly natural formalism for these demands. This paper adopts a MARL centered scope and asks where large language models (LLMs) should augment, interface with, train, or, in the strongest competitive case, replace that coordination core. A taxonomy organizes the literature through four LLM attachment points: policy, reward design, communication between agents, and hierarchical planning. A conditional capability profile separates native mechanism, reported performance, formal guarantee, and engineering maturity, and a deployment readiness analysis identifies the evidence behind each role. These stages yield the principal contribution: a three layer MARL centered reference architecture, grounded in evidence, for semantic reasoning, adaptive cooperative control, and independently assured execution. The LLM-Augmented Dec-POMDP is a descriptive comparative notation for that architecture, recording four attachment choices without introducing a new decision process class or algorithm. Under the reviewed evidence, conventional MARL is better suited to frequent, structured, decentralized coordination after task specific training, whereas LLM components are promising for semantic interpretation, reward drafting, human interaction, and slower supervisory planning. Current LLM only manufacturing controllers do not yet establish equivalence for strict real time, decentralized, safety critical control; this conclusion is bounded by the available evidence and does not assert impossibility.
NiyamAI - ゼロ知識証明を使用した暗号検証可能なガードレールを備えたインテントバインド AI エージェント
AI エージェントに電子メールの送信、データベースのクエリ、またはコマンドの実行の機能を与えることは、エージェントがだまされてやるべきではないことを実行するまでは便利です。即時注入、幻覚的な推論、および安全でないツールの呼び出しが、自律型 LLM エージェントの主な攻撃対象領域を形成します。既存の防御策は、攻撃者がターゲットとする同じマシン上で実行されるシステム プロンプトやポリシー フィルターなどのソフトウェア チェックに依存しており、検証可能な実行証拠は提供されません。安全性の執行を証明可能にするフレームワーク Niyam-AI を紹介します。セッション開始時に、許可されたツールと制約は、SHA-256 経由でコミットされたインテント コントラクトにロックされます。すべてのツール呼び出しは、分離された Judge モデルによって傍受され、検証されます。合格すると、EZKL を介して zk-SNARK 証明が生成されます。このツールは証拠検証後にのみ実行されるため、第三者は裁判官モデルの重みにアクセスせずに執行を確認できます。 Agent-SafetyBench の 2,000 の実世界シナリオで Niyam-AI を NeMo Guardrails、Meta の Llama Prompt Guard 2、OpenAI の GPT-OSS-Safeguard に対して 5 重層別相互検証を使用して評価すると、F1 スコアは 88.5%、偽陽性率は 1.1% でした (ブートストラップ 95% CI: [85.19%、91.88%]、N=1000)。 McNemar の完全対応テストでは、大幅な改善が確認されています。Niyam-AI は、NeMo に対して 390 の不一致シナリオで勝利 (対 20 敗)、Prompt Guard 2 に対して 115 (対 13)、GPT-OSS-Safeguard に対して 384 (対 19) で勝利し、すべてのケースで p < 0.0001 でした。プルーフの生成には、承認されたアクションごとに 2260.6 +/- 218.4 ミリ秒が追加され、検証には 53.1 +/- 11.8 ミリ秒かかります。 Niyam-AI は、高精度で数学的に検証可能なガードレールを提供します。ただし、これはゼロショット ベースラインに対して評価される Agent-SafetyBench に適合した分類器を反映しています。この区別についてはセクション IV.C で説明します。
原文 (English)
NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs
Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it shouldn't. Prompt injection, hallucinated reasoning, and unsafe tool calls form the primary attack surface for autonomous LLM agents. Existing defenses rely on software checks like system prompts or policy filters running on the same machine the attacker targets, offering no verifiable proof of execution. We introduce Niyam-AI, a framework that makes safety enforcement provable. At session start, permitted tools and constraints are locked into an Intent Contract committed via SHA-256. Every tool call is intercepted and validated by an isolated Judge model; upon passing, a zk-SNARK proof is generated via EZKL. The tool executes only after proof verification, allowing third parties to confirm enforcement without accessing Judge model weights. Evaluating Niyam-AI on 2,000 real-world scenarios from Agent-SafetyBench against NeMo Guardrails, Meta's Llama Prompt Guard 2, and OpenAI's GPT-OSS-Safeguard using 5-fold stratified cross-validation yields an F1 score of 88.5% with a 1.1% false-positive rate (bootstrap 95% CI: [85.19%, 91.88%], N=1000). McNemar's exact paired test confirms significant improvement: Niyam-AI wins 390 discordant scenarios against NeMo (vs 20 losses), 115 against Prompt Guard 2 (vs 13), and 384 against GPT-OSS-Safeguard (vs 19) with p < 0.0001 in all cases. Proof generation adds 2260.6 +/- 218.4 ms per approved action, while verification takes 53.1 +/- 11.8 ms. Niyam-AI provides a guardrail that is both highly accurate and mathematically verifiable--though this reflects a classifier adapted to Agent-SafetyBench evaluated against zero-shot baselines, a distinction discussed in Section IV.C.
エージェント メモリの蒸留: 階層型教師メモリを使用して小規模 LLM エージェントを強化する
記憶システムは、エージェントのパフォーマンスを向上させる可能性を示していますが、十分な成功軌道を独自に生成するのが難しい小規模な言語モデルでは、その可能性はほとんど解明されていません。私たちは、階層型メモリを通じて大規模な教師エージェントから小規模な学生エージェントに構造化された知識を転送する、トレーニング不要のフレームワークであるエージェント メモリ蒸留 (AMD) を提案します。 AMD は、成功した教師の軌跡から 3 つの相補的なメモリ タイプを構築します。ワークフロー メモリはタスク レベルの戦略をエンコードし、サブタスク メモリは中間の粒度で具体的な動作例を提供し、関数メモリは関数ごとの呼び出し規約と一般的な落とし穴をキャプチャします。ワークフローおよびサブタスクのメモリは各タスクの開始時にプロアクティブに挿入され、関数メモリはツール呼び出しエラー時にリアクティブに取得されます。 GPT-5-mini を教師として使用し、4 つの学生モデル (4B ~ 8B パラメーター) を使用して 3 つのツール使用ベンチマークで AMD を評価し、AppWorld、BFCL V3、および ToolSandbox で 27.2%p、11.2%p、および 3.4%p の平均精度向上を達成しながら、既存のメモリベースのベースラインを一貫して上回っています。さらに分析した結果、サブタスクのメモリが最大の利益に貢献し、教師の効果は教師の能力と生徒の相性の両方に依存し、4B サイズの生徒が AMD から最も恩恵を受けることが示されました。
原文 (English)
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectories on their own. We propose Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent through hierarchical memory. AMD constructs three complementary memory types from successful teacher trajectories: Workflow memory encodes task-level strategies, Subtask memory provides concrete behavioral examples at an intermediate granularity, and Function memory captures per-function calling conventions and common pitfalls. Workflow and Subtask memories are injected proactively at the start of each task, while Function memory is retrieved reactively upon tool-calling errors. We evaluate AMD on three tool-use benchmarks using four student models (4B-8B parameters) with GPT-5-mini as the teacher, achieving average accuracy gains of 27.2%p, 11.2%p, and 3.4%p on AppWorld, BFCL V3, and ToolSandbox, while consistently outperforming existing memory-based baselines. Further analysis shows that Subtask memory contributes the largest gains, teacher effectiveness depends on both teacher capability and student compatibility, and 4B-sized students benefit most from AMD.
SetEasy: マルチモーダルな教室の取り組み評価と座席の最適化フレームワーク
SetEasy は、固定座席グリッドでの教室での参加を最適化します。マルチモーダル センシング (リストバンドの生理学、4K ビデオ、環境データ) を融合し、改訂された ISEQ に基づいて v-Gage モデルをトレーニングします。毎週、2 週間の参加予測が学生の座席ユーティリティ マトリックスにマッピングされ、CP-SAT は視覚的アクセスと社会力学の制約の下で座席計画を生成します。 4 週間の導入 (生徒 23 人、クラス 331 回) で、v-Gage は感情、行動、認知、および全体的な側面を統合し、RMSE を 0.75 から 0.53 に削減しました。最適化により平均エンゲージメントが 0.30 から 0.70 に上昇し、座席の 3 分の 2 以上が高いエンゲージメントに達し、後列の低アクティビティ パターンが大幅に減少しました。これらの結果は、ハードウェアを変更せずに、解釈可能なデータ駆動型の座席戦略によりエンゲージメントを大幅に向上できることを示しています。マルチモーダルな「評価 + 最適化」パラダイムは、世界的な均質化の中で文化に対応した差別化された空間デザインへの移行可能で持続可能な道を提供します。
原文 (English)
SetEasy: A Multi-Modal Classroom Engagement Assessment and Seating Optimization Framework
SetEasy optimizes classroom engagement in fixed seating grids. It fuses multimodal sensing (wristband physiology, 4K video, environmental data) and trains a v-Gage model grounded in a revised ISEQ. Each week, two-week engagement forecasts are mapped to a student-seat utility matrix, and CP-SAT generates seating plans under visual-access and social-dynamics constraints. In a four-week deployment (23 students, 331 classes), v-Gage converged across affective, behavioral, cognitive, and overall dimensions, cutting RMSE from 0.75 to 0.53. Optimization raised mean engagement from 0.30 to 0.70, with over two-thirds of seats reaching high engagement and back-row low-activity patterns markedly reduced. These results show that, without hardware changes, interpretable, data-driven seating strategies can substantially enhance engagement. The multimodal "assessment + optimization" paradigm offers a transferable, sustainable path to culturally responsive, differentiated spatial design amid global homogenization.
EMAS: 証拠に基づいた改訂によるマルチエージェント システムの進化の安定化
自動マルチエージェント システム設計の多くの方法では、初期設計段階でプロンプトとトポロジを最適化し、結果として得られたシステムを変更せずに後続のサンプルに展開します。これらのサンプルからの経験が再利用可能なシステム更新に統合されることはほとんどありませんが、精度重視の設計では高額のトークンコストが発生する可能性があります。この経験を利用して LLM パラメータを更新せずに MAS トポロジとプロンプトを改訂し、精度を向上させたりコストを削減したりする EMAS (Evolving Multi-Agent System) を導入します。 EMAS は、トレースを、リビジョン操作とターゲットを指定する構造化診断に変換します。同じ診断がサンプル間で再発した場合にのみ候補改訂を生成し、現在の MAS に対するペア検証が対応する受け入れ基準を満たしている場合にのみ適用されます。 4 つのベンチマークと 2 つの LLM にわたって、EMAS は両方のバックボーンで最高のタスク加重総合精度を達成し、8 つのモデル - ベンチマーク設定のうち 6 つで最高または同点です。 2 つの進化エポック内で、EMAS は、Kimi-K2-6 と Qwen3.6-27B でタスク重み付け精度においてそれぞれ 6.30% と 20.10% の相対的な向上を達成しました。 Qwen3.6-27B を使用した MBPP では、EMAS の精度が 55.09% から 89.12% に向上し、タスクごとのトークン使用量が 62.2% 削減されました。これらの結果は、EMAS が新しいサンプルからのエクスペリエンスを MAS トポロジとプロンプトの再利用可能な更新に変えることができることを示しています。
原文 (English)
EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision
Many methods for automated multi-agent system design optimize prompts and topologies during an initial design stage and then deploy the resulting system unchanged on subsequent samples. Experience from these samples is rarely consolidated into reusable system updates, while accuracy-oriented designs may incur high token costs. We introduce EMAS (Evolving Multi-Agent System), which uses this experience to revise MAS topology and prompts without updating LLM parameters, either to improve accuracy or to reduce cost. EMAS converts traces into structured diagnoses that specify a revision operation and target. It generates a candidate revision only when the same diagnosis recurs across samples and applies it only if paired validation against the current MAS meets the corresponding acceptance criterion. Across four benchmarks and two LLMs, EMAS attains the highest task-weighted overall accuracy for both backbones and is best or tied in six of eight model--benchmark settings. Within two evolution epochs, EMAS achieves relative gains of 6.30% and 20.10% in task-weighted accuracy on Kimi-K2-6 and Qwen3.6-27B, respectively. On MBPP with Qwen3.6-27B, EMAS raises accuracy from 55.09% to 89.12% while reducing token use per task by 62.2%. These results show that EMAS can turn experience from new samples into reusable updates to MAS topology and prompts.
LLM 支援ツールと来歴ナレッジ グラフを使用した、ランダム化された臨床試験出版物の透明性のある研究公正評価の作成と管理
ランダム化比較試験(RCT)の系統的レビューは、臨床診療ガイドラインの証拠として日常的に使用されています。このような証拠は、臨床ケアに影響を与える低品質または誤った研究成果を防ぐために、高い研究公正基準を満たしている必要があります。ただし、公開された RCT の研究の完全性を評価するのは手作業が必要な複雑なプロセスであり、人間の評価者の意見が多様になる可能性があります。この論文では、コミュニティで承認された INSPECT-SR フレームワークに基づいて、人間の査読者が公開された RCT の研究完全性を評価するのを支援する、LLM ベースの対話型ツールである INSPECT-AI と、評価プロセスの出所を文書化するための Research Integrity Provenance and Evidence オントロジー (RIPE-O) について説明します。さらに、INSPECT-AI によって生成され、RIPE-O を使用して記述された、95 件の RCT 出版物に対する 140 件の専門家の研究公正性評価の初期セットである、研究公正出自および証拠ナレッジ グラフ (RIPE-KG) を紹介します。
原文 (English)
Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs
Systematic reviews of Randomised Controlled Trials (RCTs) are routinely used as evidence for clinical care guidelines. Such evidence has to meet high research integrity standards to prevent low quality or false research outputs influencing the clinical care. However, assessing research integrity of published RCTs is a complex process requiring manual effort, and potentially resulting in diverse opinions of the human assessors. This paper describes INSPECT-AI, an LLM-based interactive tool that assists human reviewers with research integrity assessments of published RCTs based on the community approved INSPECT-SR framework, and the Research Integrity Provenance and Evidence ontology (RIPE-O) for documenting the provenance of the assessment process. In addition, we present the Research Integrity Provenance and Evidence knowledge graph (RIPE-KG), an initial set of 140 expert research integrity assessments of 95 RCT publications generated by INSPECT-AI and described using RIPE-O.
Beyond the Black Box: Interpretable Models of Human Randomisation Failures
Mixed strategy equilibrium predicts i.i.d play: past actions should not help predict future decisions. Human players, however, systematical…
確率論理プログラミングにおける確率から因果関係へ
確率的論理プログラミングは、システムの外部からの介入を含む因果関係のクエリをサポートする統計的リレーショナル人工知能の形式主義です。しかし、確率論理プログラムの構造をデータから学習する場合、確率情報のみが使用されるため、単一の確率分布が複数の因果順序と互換性がある可能性があります。これは介入推論のあいまいさにつながり、因果順序が分布によっていつ一意に決定されるかという問題を引き起こします。非巡回確率論理プログラムとベイジアン ネットワークの関係を利用して、プログラムにエンコードされた確率情報が一意の因果順序を決定する条件を導き出します。また、基礎となる関係語彙によって引き起こされる所定の因果的対称性のセットを考慮することにより、関係構造から生じる制約も組み込みます。その結果、学習された確率的論理プログラムが明確に定義された介入セマンティクスをいつサポートするかを検証する方法が得られます。
原文 (English)
From probability to causality in probabilistic logic programming
Probabilistic logic programming is a formalism of statistical relational artificial intelligence that supports causal queries, including interventions from outside the system. When the structure of a probabilistic logic program is learned from data, however, only probabilistic information is used, and a single probability distribution may be compatible with several causal orders. This leads to ambiguity in interventional reasoning, raising the question of when the causal order is uniquely determined by the distribution. Exploiting the relationship between acyclic probabilistic logic programs and Bayesian networks, we derive conditions under which the probabilistic information encoded in a program determines a unique causal order. We also incorporate constraints arising from relational structure by taking into account prescribed sets of causal symmetries induced by the underlying relational vocabulary. The result is a method for verifying when a learned probabilistic logic program supports well-defined intervention semantics.
創造性のレシピ: 大規模な言語モデルでの反復生成と評価
生成モデルは単一の成果物を通じて評価されることがよくありますが、人間の創造性は通常、反復的な生成、評価、改良を通じて現れます。このパイロット研究では、FunSearch を 2024 年のピルズベリー ベイクオフのレシピ生成に適応させ、TTCT ベースの LLM 評価を使用して人間のベンチマークに対して出力を評価することで、反復検索が LLM の創造性を向上させるかどうかを検証します。 2 つの実験にわたって、反復回数、ジェネレーターの温度、ループ内選択スコアラー モデルのサイズをテストします。結果は、世代選択の反復により人間のベンチマークに匹敵する創造性スコアを持つレシピを生成できるが、反復を追加するだけでは創造性が向上しないことを示しています。ループ内評価器が最も重要です。選択スコアラーが小さいほど、ほとんどの TTCT 次元にわたって大幅に高いスコアが得られますが、温度の影響は独自性を除いて限定的です。これらの発見は、評価者の設計が主観的なクリエイティブ検索における一次設計変数であることを示唆しています。
原文 (English)
Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models
Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting FunSearch to recipe generation for the 2024 Pillsbury Bake-Off and evaluating outputs against human benchmarks using TTCT-based LLM evaluation. Across two experiments, we test iteration count, generator temperature, and in-loop selection-scorer model size. Results show that iterative generation-selection can produce recipes with creativity scores comparable to human benchmarks, but additional iterations alone do not improve creativity. The in-loop evaluator matters most: a smaller selection scorer yields significantly higher scores across most TTCT dimensions, while temperature has limited effects except for originality. These findings suggest that evaluator design is a first-order design variable in subjective creative search.
WNM-3D: 閉ループ VLN 用の 3D シーンコンディショニングを備えたワールド ナビゲーション モデル
最近のビジョン言語ナビゲーション (VLN) システムでは、事前トレーニング済みビジョン言語モデル (VLM) を、自己中心的な観察と言語指示をナビゲーション アクションに直接マッピングするビジョン言語アクション (VLA) ポリシーにますます適応させています。このようなアクション中心のトレーニングは、意味論的には可能ですが、エージェントの視覚観察がその予測された動作の下でどのように展開するかを明示的にモデル化するものではありません。生成世界アクション モデル (WAM) は、将来の観測とアクションを共同で予測しますが、連続 VLN の既存の WAM は、観察された履歴から推測される幾何学を意識した表現に基づいて共同の将来展望とアクションの生成を条件付けません。我々は、連続 VLN 用の 3D シーンコンディショニングを備えた生成ワールド ナビゲーション モデルである WNM-3D を紹介します。過去の観察を永続的なシーン コンテキストに統合するために、フリーズ フィードフォワード ジオメトリ エンコーダーが単眼の自己中心的な RGB 履歴からジオメトリ認識表現を抽出し、トレーニング可能な 3D シーンからトークンへのアダプターがそれらをワールド アクション拡散トランスフォーマーのトークン空間内の固定長プレフィックスに変換します。ブロック因果的注意を通じて、このプレフィックスは将来のすべてのビデオ アクション ブロックを条件付けし、将来のビューとアクション生成の両方に共有の幾何学的コンテキストを提供します。私たちは、A* によって生成されたデモンストレーションに対する監視付きワールドアクション微調整、政策訪問国に対する DAgger スタイルの適応、および DanceGRPO ベースの閉ループ政策最適化を通じて、WNM-3D をトレーニングします。 GN-Bench での実験では、WNM-3D が閉ループ ナビゲーションにおいて強力な VLM ベースのナビゲーション ポリシーや 2D 条件付きの対応物よりも優れていることが示されています。 WNM-3D は、固定されたゴールに近い評価セットで、より高いフローアクションの一貫性とより低い視覚動作エラーも実現します。
原文 (English)
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantically capable, such action-centric training does not explicitly model how the agent's visual observations should evolve under its predicted motion. Generative world-action models (WAMs) jointly predict future observations and actions, yet existing WAMs for continuous VLN do not condition joint future-view and action generation on geometry-aware representations inferred from the observed history. We present WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN. To consolidate past observations into persistent scene context, a frozen feed-forward geometry encoder extracts geometry-aware representations from the monocular egocentric RGB history, and a trainable 3D Scene-to-Token Adapter converts them into a fixed-length prefix in the token space of the world-action Diffusion Transformer. Through block-causal attention, this prefix conditions every future video-action block, providing a shared geometric context for both future-view and action generation. We train WNM-3D through supervised world-action fine-tuning on A*-generated demonstrations, DAgger-style adaptation on policy-visited states, and DanceGRPO-based closed-loop policy optimization. Experiments on GN-Bench show that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation. On a fixed near-goal evaluation set, WNM-3D also achieves higher flow-action consistency and lower visual-motion error.
覗き見による勝利: 強制されていない予算とテストセットの選択により、短い予算の AutoML 比較が膨らむ
ツールの README やワークショップの資料では、短時間 (数時間ではなく数十秒) での AutoML システム間の比較が一般的であり、間違いやすいものです。シンプルな AutoML エンジンである Orcetra が、513 の OpenML データセットで FLAML および AutoGluon に勝利し、名目 60 秒の予算でそのうちの 57.1% を獲得し、30 秒で FLAML 単独に対してデータセットの 78.4% を獲得したケーススタディを報告します。両方のマージンは、結果表では表示できないプロトコルの欠陥に起因しています。検索ループはテスト分割のすべての候補をスコアリングし、最も優れたものを報告し、ベースラインがトレーニング データに基づいて選択され、テスト セットに 1 回適用される一方で、ヘッドライン メトリクスは数十のノイズの多い推定値を超える最大値になりました。また、予算は候補を立ち上げる前にチェックされましたが、候補を立ち上げる間に強制されることはなかったので、システムは 60 秒の予算に対して中央値 120 秒を消費しました。これは、AutoGluon が使用した実時間の 2.24 倍です。選択を検証分割に移し、期限を外部で強制し、すべてのフレームワークをマシンの均等なシェアに固定して再実行すると、再実行サブセットでの Orcetra の勝率は 59.4% から 34.3% に低下し、どちらの競合相手に対してもペアごとの大きな差は残りません。単一の検索内で両方の推定値を記録すると、崩壊の原因を特定できます。選択ルールは 4.8 パーセント ポイントを占め、残りのほとんどは不等計算です。同じトレースは、推定ではなく測定された予算の関数として選択バイアスを示します。 $K$ とともに増加しますが、精度ポイントは 0.27 ポイントにすぎず、限界標準誤差の議論が予測する $\sigma\sqrt{2\ln K}$ の約 5 倍低い精度ポイントに達します。これは、共有テスト行で得点された候補者がノイズの大部分をキャンセルするためです。最後に、短い予算で比較するためのチェックリストを示します。コード、データセットごとの結果、論文内のすべての数値と図を再生成するスクリプトも一緒にリリースされます。
原文 (English)
Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons
Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and they are easy to get wrong. We report a case study in which a simple AutoML engine, Orcetra, appeared to beat FLAML and AutoGluon on 513 OpenML datasets, winning 57.1% of them at a nominal 60-second budget and 78.4% of datasets against FLAML alone at 30 seconds. Both margins came from protocol defects that a results table cannot show. The search loop scored every candidate on the test split and reported the best, making the headline metric a maximum over dozens of noisy estimates while the baselines selected on training data and touched the test set once; and the budget was checked before launching a candidate but never enforced during one, so the system consumed a median of 120 s against a 60-second budget, 2.24x the wall-clock AutoGluon used. Re-running with selection moved to a validation split, the deadline enforced externally and every framework pinned to an equal share of the machine, Orcetra's win rate on the re-run subset falls from 59.4% to 34.3% and no pairwise difference against either competitor remains significant. Recording both estimands inside a single search lets us attribute the collapse: the selection rule accounts for 4.8 percentage points and unequal compute for most of the rest. The same traces give the selection bias as a function of budget, measured rather than assumed: it grows with $K$ but reaches only 0.27 accuracy points, about five times below the $\sigma\sqrt{2\ln K}$ bound a marginal-standard-error argument predicts, because candidates scored on shared test rows cancel most of the noise. We close with a checklist for short-budget comparisons. Code, per-dataset results and the scripts that regenerate every number and figure in the paper are released with it.
エンドツーエンドのエージェント監査エンジン
大規模言語モデル (LLM) の急速な進歩により、ハーネスは幅広いドメインにエージェントを展開するための不可欠なインフラストラクチャになりました。ハーネスのエコシステムが急速に進化しているため、厳密な機能評価の重要性も高まっています。ただし、エンドツーエンドの体系的かつ包括的な評価パイプラインを効率的に構築することは依然として大きな課題です。この課題に対処するために、エージェント ハーネス用に設計されたエンドツーエンドの評価エンジンである $A^2E$ (エージェント監査エンジン) を導入します。 $A^2E$ は、新しく提案されたエージェント タスク プロトコル (ATP) を活用して、評価タスクとさまざまなハーネスの迅速な統合を可能にします。自動的に計測されたモニターを通じて、実験中に標準化された実行トレースをキャプチャして生成します。評価ステージでは、$A^2E$ は一連の多次元メトリックを使用してハーネス機能を体系的に評価します。これらのメトリクスは、正確性だけと比較して、実行効率、ツールの使用、タスク計画、およびエラー回復におけるハーネス間の違いをより詳細に特徴付けることができます。 $A^2E$ を使って行われた実験では、モデルとハーネスの組み合わせがさまざまな種類のタスク間で大幅なパフォーマンスのばらつきを示し、単一の組み合わせがすべてのタスクにおいて一貫して他のすべての組み合わせを上回ることはないことがさらに明らかになりました。これらの発見は、体系的な評価の必要性を実証するだけでなく、モデルとハーネスを共同進化させるための有用な指針も提供します。私たちのコードは https://github.com/datamllab/A2E で入手できます。
原文 (English)
An End-to-End Agent Auditing Engine
With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving harness ecosystem has also made rigorous capability evaluation increasingly important. However, efficiently building an end-to-end, systematic, and comprehensive evaluation pipeline remains a significant challenge. To address this challenge, we introduce $A^2E$ (Agent Auditing Engine), an end-to-end evaluation engine designed for agent harnesses. $A^2E$ leverages our newly proposed Agent Task Protocol (ATP) to enable the rapid integration of evaluation tasks with different harnesses. Through an automatically instrumented Monitor, it captures and generates standardized execution traces during experiments. In the Evaluation stage, $A^2E$ systematically assesses harness capabilities using a suite of multidimensional metrics. Compared with correctness alone, these metrics provide a more fine-grained characterization of differences among harnesses in execution efficiency, tool use, task planning, and error recovery. Experiments conducted with $A^2E$ further reveal that model-harness combinations exhibit substantial performance variation across different types of tasks, and that no single combination consistently outperforms all others across every task. These findings not only demonstrate the necessity of systematic evaluation but also provide useful guidance for the co-evolving of models and harnesses. Our code is available at https://github.com/datamllab/A2E.
QFCQT: 不安定な時系列予測のためのカオス的にゲートされたQuantformerフレームワーク
非定常時系列の予測は、長期依存関係、局所的なボラティリティのバースト、構造変化、非線形振動挙動のため、依然として困難です。 Transformer ベースの予測機能は長期的な時間依存関係のモデル化には効果的ですが、そのフィードフォワード ブロックは通常、突然の体制変化に対する感度が不十分な滑らかな静的アクティベーションに依存しています。定量的トランスフォーマー設計とオシレーターベースの非線形活性化を動機として、複雑な揮発性ダイナミクスの下でロバストな予測を行うために、量子フラクタルにインスピレーションを得たカオティック ゲート クアントフォーマーの略である QFCQT を提案します。ここで、「量子フラクタルにインスピレーションを得た」とは、正式な量子力学的またはフラクタル理論的な導出ではなく、ソフトオシレーターの重ね合わせとマルチスケールの非線形応答に基づく計算上の類似性を指します。 QFCQT は 3 つの主要コンポーネントで構成されます。(1) 線形埋め込みを介して多変量入力を直接処理する Quantformer スタイルの数値エンコーダー。 (2) スカラーの事前活性化を動的振動応答にマッピングし、Max-over-Time プーリングを通じてそれらを要約する、学習可能な Lee オシレーター活性化モジュール。 (3) 従来のスムーズな活性化とカオスに敏感な応答のバランスを適応的に調整するスムーズカオスゲート融合機構。さらに、QFCQT は、単一の固定発振器を使用する代わりに、8 つのパラメーター化された Lee 発振器ファミリーのソフト重ね合わせを使用して、レジーム全体で異なる非線形応答パターンを適応的に捕捉します。 ETTh1、ETTh2、および A 株株価指数ベンチマークの実験では、QFCQT が Informer、LogTrans、LSTMa、HAT、COTN などの強力なベースラインを常に上回るパフォーマンスを示しています。
原文 (English)
QFCQT: A Chaotically Gated Quantformer Framework for Volatile Time-Series Forecasting
Forecasting non-stationary time series remains difficult due to long-range dependencies, local volatility bursts, structural shifts, and nonlinear oscillatory behaviors. Although Transformer-based forecasters are effective for modeling long-term temporal dependencies, their feed-forward blocks typically rely on smooth static activations that are insufficiently sensitive to abrupt regime changes. Motivated by quantitative Transformer designs and oscillator-based nonlinear activations, we propose QFCQT, short for Quantum-Fractal-inspired Chaotically Gated Quantformer, for robust forecasting under complex volatile dynamics. Here, "quantum-fractal-inspired" denotes a computational analogy based on soft oscillator superposition and multi-scale nonlinear responses, rather than a formal quantum-mechanical or fractal-theoretic derivation. QFCQT consists of three main components: (1) a Quantformer-style numerical encoder that directly processes multivariate inputs via linear embedding; (2) a learnable Lee-oscillator activation module that maps scalar pre-activations to dynamic oscillatory responses and summarizes them through Max-over-Time pooling; and (3) a smooth-chaotic gated fusion mechanism that adaptively balances conventional smooth activations and chaos-sensitive responses. Furthermore, instead of using a single fixed oscillator, QFCQT employs a soft superposition of eight parameterized Lee oscillator families to adaptively capture different nonlinear response patterns across regimes. Experiments on ETTh1, ETTh2, and A-share Stock Index benchmarks show that QFCQT consistently outperforms strong baselines, including Informer, LogTrans, LSTMa, HAT, and COTN.
コードとしてのカリキュラム: STEM 教育における教育デザインのための AI 支援アーキテクチャ
寄稿: この論文では、Curriculum as Code パラダイムに基づいた 6 段階の AI 支援の教育設計アーキテクチャを紹介し、Generative AI を LaTeX および Python と統合して、STEM 教育用の再現可能で視覚的に一貫性があり、技術的に正確な教材の作成を自動化します。背景: アクティブ ラーニング用にカスタマイズされた教材を作成することは、教員に大きな負担を強います。標準的なプレゼンテーション ツールには技術的なコンテンツに対する強力なサポートが欠けており、一方、現在の AI アプリケーションは幻覚を起こしたり、指導オーサリング プロセスを形式化できなかったりすることが多く、厳密な学術的デザインでの有用性が制限されています。意図された成果: このフレームワークは、数学的正確性、組織の視覚的アイデンティティの遵守、明示的なルールによる講師の暗黙の教育的知識の保存を確保しながら、準備時間を短縮することを目的としています。アプリケーション設計: このソリューションは、アドホックなプロンプト エンジニアリングを体系的なワークフローに置き換える 6 フェーズのパイプラインで構成され、テキストベースのインターフェイスとコード駆動の生成 (スライドには LaTeX/Beamer、図には Python) を利用し、教育上の制約、状況に応じた調整、および自動レビュー サイクルによって管理されます。調査結果: プロジェクトベースの学習環境における 8 つのモジュールと 28 のプロジェクト コンテキストにわたって 1 年以上検証された結果、このアーキテクチャはインストラクターの作業負荷を大幅に軽減しました。生成されたアセットは独立したピアレビューを受け、6 人の異なる教員によって展開され、単一の作成者を超えたスケーラビリティが確認されました。 600 名を超える学生による自発的な評価に基づいて、教材は 8.5 ~ 9.9/10 の高品質評価を獲得しました。結果は、高い再現性、最小限の幻覚、持続的な教育的および視覚的忠実度を示しており、広範な STEM 教育アプリケーションの実行可能性を示唆しています。
原文 (English)
Curriculum as Code: An AI-Assisted Architecture for Instructional Design in STEM Education
Contribution: This paper presents a six-phase AI-assisted instructional design architecture based on the Curriculum as Code paradigm, integrating Generative AI with LaTeX and Python to automate the creation of reproducible, visually consistent, and technically precise materials for STEM education. Background: Creating customized instructional materials for active learning imposes a heavy workload on faculty. Standard presentation tools lack robust support for technical content, while current AI applications often hallucinate and fail to formalize the instructional authoring process, limiting their utility for rigorous academic design. Intended Outcomes: The framework aims to reduce preparation time while ensuring mathematical accuracy, adherence to institutional visual identity, and preservation of the instructor's tacit pedagogical knowledge through explicit rules. Application Design: The solution comprises a six-phase pipeline that replaces ad-hoc prompt engineering with a systematic workflow, utilizing text-based interfaces and code-driven generation (LaTeX/Beamer for slides, Python for figures), governed by pedagogical constraints, contextual calibrations, and automated review cycles. Findings: Validated over one year across 8 modules and 28 project contexts in a Project-Based Learning environment, the architecture significantly reduced instructor workload. Generated assets underwent independent peer review and were deployed by six different faculty members, confirming scalability beyond a single author. Based on over 600 voluntary student evaluations, materials achieved high quality ratings from 8.5 to 9.9/10. Results indicate high reproducibility, minimized hallucinations, and sustained pedagogical and visual fidelity, suggesting viability for broad STEM educational applications.
人々は自分の国だけではありません。ヨーロッパ全体での LLM 価値の調整の社会的決定要因を解きほぐす
大規模言語モデル (LLM) が情報やアドバイスの主要な情報源としてますます使用されるようになっているため、価値観の観点から LLM が人間と一致していることを理解することが差し迫った懸念事項になっています。 LLM と人間の主張する価値観や意見がどの程度一致しているかを調査するために、大規模な調査を活用した文献が増えています。限られた例外を除き、研究対象となる集団には国境または文化的境界が定義されています。しかし、この焦点は、社会人口学的格差が価値観の一致の格差に果たす役割を無視している。欧州社会調査に基づいて、15 の社会人口統計的変数および居住国の観点から、10 の著名な商用 LLM に関して表示される価値の一致を考慮することで、この知識のギャップに対処します。私たちの分析では、LLM がさまざまな社会人口統計上のグループ、特に教育、収入、職業、宗教によって定義される価値観に実際に不均等に同調していることが明らかになりました。個人レベルでの整合性を調べる場合、回答者の国を独立変数として考慮すると、考慮される社会人口統計の完全なセットと同等のかなりの量の変動が説明されます。国レベルの要因と社会人口学的要因のそれぞれの役割をさらに解きほぐすと、それらは価値観の調整パターンを説明する上で補完的であり、それらの相対的な重みは考慮された質問のサブセット全体で異なります。
原文 (English)
People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
As Large Language Models (LLMs) are increasingly used as a primary source of information and advice, understanding their alignment to humans in terms of values becomes a pressing concern. A growing literature has leveraged large scale surveys to investigate to what extent LLMs' and humans' stated values and opinions align. With limited exceptions, studied populations have been defined country borders or cultural bounds. Yet, this focus neglects the role that socio-demographic divides may play for value alignment disparities. Relying on the European Social Survey, we address this knowledge gap by considering value alignment displayed with respect to 10 prominent commercial LLMs in terms of 15 socio-demographic variables as well as country of residence. Our analyses reveal that LLMs are indeed unequally aligned to the values of different socio-demographic groups, notably those defined by education, income, occupation and religion. When examining alignment at the individual level, a respondent's country, taken as a stand-alone variable, explains a substantial amount of variation that is on par with the full set of considered socio-demographics. Further disentangling the respective role of country-level and socio-demographic factors, we find they are complementary in explaining value alignment patterns, with their relative weights varying across the subset of questions considered.
FinRank: SEC 提出書類に基づく財務上の質問回答と検索のための、証拠に基づいたベンチマーク
財務上の質問への回答は通常、回答の正しさによって評価されますが、SEC への提出書類では、もっともらしく、数値的に正しい回答であっても、間違った証拠に基づいている可能性があります。同様の事実と開示は、申請書の各セクション間、同じ企業の報告期間間、および比較可能な企業間で繰り返されます。 FinRank は、対象となるエンティティ、報告期間、開示内容の証拠を特定することをシステムに要求することで、この出所に敏感な検索問題をターゲットにしています。このベンチマークには、22 社の 10-K および 10-Q 申告書に基づいて手動で作成された 1,185 件の質問と回答の記録が含まれています。各記録には、参考回答、金を裏付ける文章、および出願書類内、報告期間全体、および比較可能な企業全体の紛らわしい文章から抽出された手作業で厳選されたハードネガが含まれています。 FinRank は、パッセージ検索、再ランキング、ハードネガティブ識別を個別に測定されたタスクとして評価します。ベースライン結果は、この設定の難しさを示しています。評価されたシステムの中で、7B 命令調整エンベッダーでさえ、プールされた証拠コーパスで 44.8% の Recall@10 に達するだけです。 10 億未満のパラメータのエンコーダは BM25 よりも最大 3.5 ポイント向上し、ファイナンスに適応したエンベッダは BM25 に 9.7 ポイント及ばず、ランダムなネガが厳選されたハード ネガに置き換えられた場合、ペアごとの精度は 13.0 ~ 20.5 パーセント ポイント低下します。 FinRank は、正確なだけでなく正しい開示に基づいた財務質問応答システムを開発するための証拠優先のベンチマークを提供します。
原文 (English)
FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance-sensitive retrieval problem by requiring systems to identify evidence for the intended entity, reporting period, and disclosure context. The benchmark contains 1185 manually authored question-answer records over the 10-K and 10-Q filings of 22 companies. Each record includes a reference answer, gold supporting passages, and hand-curated hard negatives drawn from confusable passages within filings, across reporting periods, and across comparable firms. FinRank evaluates passage retrieval, reranking, and hard-negative discrimination as separately measured tasks. Baseline results demonstrate the difficulty of this setting: among the evaluated systems, even a 7B instruction-tuned embedder reaches only 44.8% Recall@10 on the pooled evidence corpus; sub-billion-parameter encoders gain at most 3.5 points over BM25, a finance-adapted embedder trails BM25 by 9.7 points, and pairwise accuracy falls by 13.0-20.5 percentage points when random negatives are replaced with the curated hard negatives. FinRank provides an evidence-first benchmark for developing financial question answering systems that are not only accurate but also grounded in the correct disclosure.
GeoBenchLLM: 地理関連タスクで LLM を評価するための包括的なベンチマーク
地理データのコンテキストでは、既存の大規模言語モデルは均質な環境で研究されることが多く、一般化機能についての洞察はかなり限られています。このペーパーでは、地理関連タスクで LLM を調査するための包括的なベンチマークである \benchName を紹介します。私たちは、地理関連のさまざまなタスクやドメインから厳選した 12 個の公開データセットを活用し、ベンチマークを使用して地理空間的および時間的理解に関する一連の LLM を評価します。私たちの結果は、推論とサイズが全体的なパフォーマンスに大きな影響を与えることを示しています。 GeoBenchLLM は https://github.com/Rfr2003/GeoBenchLLM で公開されています。
原文 (English)
GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.
ResidencyRL: 模擬臨床環境における強化学習
医学教育では、医師は研修を通じて学術知識を臨床専門知識に変換します。研修は、さまざまなフィードバック源と徐々に強化された自主性を備えた、何千もの出会いを通じた長年のトレーニングです。臨床推論の多くは、患者との出会い、つまり臨床医が病歴を引き出し、診断仮説を洗練し、不確実性の下での管理を決定する対話に依存しています。大規模言語モデル (LLM) は静的な医療ベンチマークでは優れていますが、一連の臨床意思決定を最適化する方法は未開発のままです。我々は、シミュレートされた複数ターンの臨床遭遇(軌跡ごとに最大 60 回の対話ターンと 8 回のツール呼び出し)を通じて臨床人工知能 (AI) エージェントをトレーニングするための強化学習 (RL) 手法である ResidencyRL を紹介します。 ResidencyRL は、複雑で敵対的な行動が可能な LLM シミュレーターとポリシー エージェントを組み合わせ、診断の精度、管理品質、コミュニケーション、文書化、安全性に合わせた構造化された報酬に対するトレーニングを行います。保留評価では、ResidencyRL エージェントは、敵対的条件下で診断精度を 7.0% 向上させ (88.0% 対 81.0%)、危険信号の見逃し率を 31% 削減し、早期閉鎖の厳格な軽減を実証しました。盲検化された専門臨床医はこれらの利点を検証し、並べて比較した場合の 87.6% において訓練を受けたエージェントを優先しました。手順上のコンピテンシーは目に見えないベンチマークに移行します。エージェントは、AMIE 複数訪問ベンチマークの 6 つの臨床軸すべてにわたって基本モデルを上回り、AgentClinic と CRAFT-MD で一貫した方向性の改善を示しています。私たちの調査結果は、一連の臨床意思決定がシミュレーションでのマルチターン RL を通じて効果的に学習でき、堅牢で一般化可能な機能をもたらし、臨床習得への道を開くことができることを示しています。臨床上の有用性を確立するには、現実世界のワークフローを使用した前向きな検証が引き続き必要です。
原文 (English)
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy. Much of clinical reasoning relies on the patient encounter, a dialogue in which a clinician elicits history, refines diagnostic hypotheses, and decides management under uncertainty. While large language models (LLMs) excel on static medical benchmarks, methods to optimize the full sequence of clinical decisions remain underdeveloped. We present ResidencyRL, a reinforcement learning (RL) method for training clinical artificial intelligence (AI) agents through simulated multi-turn clinical encounters (up to 60 dialogue turns and 8 tool calls per trajectory). ResidencyRL pairs the policy agent with LLM simulators capable of complex, adversarial behaviors, training against a structured reward aligned to diagnostic accuracy, management quality, communication, documentation, and safety. On held-out evaluations, the ResidencyRL agent improves diagnostic accuracy by 7.0% under adversarial conditions (88.0% vs. 81.0%) and reduces missed red flag rates by 31%, demonstrating rigorous mitigation of premature closure. Blinded expert clinicians validated these gains, preferring the trained agent in 87.6% of side-by-side comparisons. The procedural competencies transfer to unseen benchmarks: the agent outperforms the base model across all six clinical axes of the AMIE multi-visit benchmark, and shows consistent directional improvements on AgentClinic and CRAFT-MD. Our findings demonstrate that sequential clinical decision-making can be effectively learned through multi-turn RL in simulation, yielding robust, generalizable capabilities, paving the way towards clinical mastery. Prospective validation with real-world workflows remains necessary to establish clinical utility.
CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or a…
一枚の絵は千トークンの価値がある: ビジョン言語モデルが精度を向上させながら AI のエネルギーコストを削減する方法
LLM 推論は AI 運用エネルギーの 90% 以上を占め、入力トークン数に直接応じてスケールします。これは、4G/5G セル サイトからの生の多変量 KPI ウィンドウが数千の浮動小数点トークンに拡張される通信ネットワーク分析および数値時系列データ分析 (NTSDA) にとって重大な非効率です。ビジョン言語モデル (VLM) は、時系列を 2D プロットとしてエンコードすることでこの不一致を排除し、Llama-3.2-90B、Qwen2.5-VL-72B、および Pixtral-12B アーキテクチャ全体で 3.6 ~ 10.4 倍の入力トークン削減を達成します。これは、測定された推論エネルギーの 1.8 ~ 2.5 倍の削減に相当し、15 分間隔で 200 セルを監視する通信エッジ展開および CloudRAN では、1 日あたり約 7.2 MJ を節約します。重要なのは、効率の向上によって精度が犠牲になることはありません。微調整された Llama-3.2-90B-Vision VLM は、テキストのみの対応物より 220.7% 高い精度を達成し、通信異常検出において LSTM および ARIMA ベースラインを 144% 以上上回ります。公開ベンチマークでは、Pixtral-12B は平均 F1 = 0.82 で J/F1 スコアで 20.6 倍の向上を達成しました。 24 KPI では、テキスト表現がほとんどの実稼働 LLM の 128K コンテキスト ウィンドウを超え、切り捨てなしではテキストのみの処理が不可能になりますが、ビジュアル表現は標準の制限内に留まります。これらの結果は、VLM が数値時系列ワークロードに対するエネルギー効率と精度に優れたモダリティとして確立され、エネルギー消費を第一級のエンジニアリング制約として扱う AI 推論システムに経験的な根拠を提供します。
原文 (English)
A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy
LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network analytics and numerical time-series data analysis (NTSDA), where raw multivariate KPI windows from 4G/5G cell sites expand into thousands of floating-point tokens. Vision-Language Models (VLMs) eliminate this mismatch by encoding time-series as 2D plots, achieving 3.6-10.4x input token reduction across Llama-3.2-90B, Qwen2.5-VL-72B, and Pixtral-12B architectures. This translates to 1.8-2.5x measured inference energy reduction, saving approximately 7.2 MJ/day at telecom edge deployments and CloudRAN that monitor 200 cells per 15-minute interval. Critically, efficiency gains do not sacrifice accuracy: a fine-tuned Llama-3.2-90B-Vision VLM achieves 220.7% higher precision than its text-only counterpart and outperforms LSTM and ARIMA baselines by over 144% on telecom anomaly detection. On public benchmarks, Pixtral-12B achieves a 20.6x improvement in J/F1 score at mean F1 = 0.82. At 24 KPIs, text representations exceed the 128K context window of most production LLMs, rendering text-only processing infeasible without truncation, while visual representations remain within standard limits. These results establish VLMs as an energy-efficient and accuracy-superior modality for numerical time-series workloads, providing empirical grounding for AI inference systems that treat energy consumption as a first-class engineering constraint.
TEPA: 競合に強い言語エージェントの古い記憶を取り消す
長期記憶により、言語エージェントは過去の事実、好み、タスクの経験を再利用できます。永続性は、主要な反証可能性の問題も引き起こします。世界が変化しても、古い記憶が検索可能なままになり、プロンプトが汚染される可能性があります。我々は、この障害モードをメモリ汚染、つまり、より新しい矛盾する証拠が置き換えられたアクティブなメモリによって引き起こされる劣化として特徴付けます。有効性を明示的な記憶状態にする、取り消し可能な証拠記憶メカニズムである TEPA を紹介します。 TEPA は観察をキー付き先例として表し、同じキーで新しい証拠が矛盾する場合に有効な先例を取り消します。これにより、取り消された履歴を監査用に保存しながら、現在の証拠から検索を行うことができます。制御された非表示体制のドリフト、実際のファイルにバックアップされた実行可能ファイルのドリフト、および設定更新ストリーム全体にわたって、取り消しにより、取り消し後の取得セットに古いアクティブ メモリが残るのを防ぎます。 50 シードを超える制御されたドリフトでは、追加のみおよび最終書き込み優先のメモリは完全反転中にメモリなしを下回りました (追加のみおよび最終書き込み優先の両方が 0.210、メモリなし 0.309、TEPA 0.950) が、実際のファイル実行でも同じパターンが再現されました (追加のみ 0.203、メモリなし 0.298、TEPA 0.950)。クリーンな MemoryAgentBench SH-6k では、TEPA は強力な Last-write-wins キャッシュと一致し、現在のキーの置換がシングルホップ ファクト統合にとって決定的な操作であることを確認します。マルチホップおよび非常に長いコンテキストの MemoryAgentBench 設定の境界テストにより、ファクトレベルの妥当性追跡を超えた検索チェーンとコンテキスト選択のボトルネックが明らかになります。これらの結果を総合すると、進化する知識を改ざん、監査し、後で再促進する必要があるエージェントにとって、ライフサイクル失効が中核的な記憶操作として確立されます。
原文 (English)
TEPA: Revoking Stale Memories for Conflict-Robust Language Agents
Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability problem: when the world changes, stale memories can remain retrievable and pollute the prompt. We characterize this failure mode as memory pollution: degradation caused by active memories that newer conflicting evidence has superseded. We introduce TEPA, a revocable evidence-memory mechanism that makes validity an explicit state of memory. TEPA represents observations as keyed precedents and revokes active precedents when fresh evidence contradicts them under the same key, allowing retrieval to draw from current evidence while preserving revoked history for audit. Across controlled hidden-regime drift, real file-backed executable drift, and preference-update streams, revocation prevents stale active memory from remaining in the retrieval set after reversal. In controlled drift over 50 seeds, append-only and last-write-wins memory fell below no memory during full reversal (append-only and last-write-wins both 0.210, no memory 0.309, TEPA 0.950), and the same pattern reproduced under real file execution (append-only 0.203, no memory 0.298, TEPA 0.950). On clean MemoryAgentBench SH-6k, TEPA matches a strong last-write-wins cache, confirming that current-key replacement is the decisive operation for single-hop fact consolidation. Boundary tests on multi-hop and very long-context MemoryAgentBench settings expose retrieval-chain and context-selection bottlenecks beyond fact-level validity tracking. Together, these results establish lifecycle revocation as a core memory operation for agents that must falsify, audit, and later re-promote evolving knowledge.
ミュオンで訓練されたトランスフォーマーの表現-読み出しインターフェースにおけるグロッキング後の崩壊
標準的な分割では、Muon は隠れ行列と AdamW の埋め込み/出力ヘッドを取得します。 Muon はモジュラー加算を高速化しますが、そのソリューションは保持されません。 $(a+b) \bmod 113$ grok 以降の 9 つの構成はすべて一般化されません。 5 つのシード全体で、選択された AdamW 参照は 4 つでしきい値を下回り、27.59% に達します。不安定性は、2 つの係数、2 つの幅、2 つのトレーニング分数、減算、および深さにわたって持続します。障害は、表現と読み出しのインターフェースで発生し、損失によって選択されなかった可逆マップまでの共同でのみ識別されます。トレーニング セットを解いた後、勾配は $10^{-6}$ のオーダーに下がり、オプティマイザーの反応は異なります。ステップサイズの弾性は、Muon では -0.03 であるのに対し、AdamW では +1.5 であり、Muon グループはパラメーターごとに 8.0 倍速く移動します。ビットが同一の状態から、どちらかのグループをフリーズすると障害が防止されます。埋め込み/読み出しを凍結すると、グロッキング後の 451,400 ステップと 5 つのペアのシードにわたる 5 回の実行でそれが削除されます。凍結されていないアームは 137 ~ 321 の閾値以下の評価を記録しますが、凍結されたアームはありません。ミューオンの正規化と直交化を削除することは代わりにはなりません。これは表現を 326 の有効な共役ペアから 4 つに崩壊させ、再発性の崩壊を示さず、最終的に失敗します。フーリエ フィルタリングは、回路障害をマスキングから分離します。 5 つのシードと 3 つの体制にわたる 43 のチェックポイントにわたって、タスクを調整した家族だけでちょうど 100% に達します。回路障害が発生すると、それはもはやタスクを解決しません。マスキングでは、完全なモデルが 45.85% に達する間、完全なままであり、エラーを含むすべての例でプラスのマージンが得られますが、ほぼ等しい敵対的な残りによって投票されます。再スケーリングすると 99.9% が回復します。グロッキングは、同じ状態が上向きに解決されることです。タスクはファミリーを選択し、減算の下で $(k,k)$ を $(k,-k)$ に交換します。突然の崩壊では、標準フーリエ サポートは変化せず、電力分布コサインは 0.9899 のままです。
原文 (English)
Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization. Across five seeds the selected AdamW reference falls below threshold on four, reaching 27.59%. Instability persists across two moduli, two widths, two training fractions, subtraction, and depth. The failure arises at the representation-readout interface, identified only jointly up to an invertible map unselected by the loss. After solving the training set, the gradient falls to order $10^{-6}$ and the optimizers respond differently: step-size elasticity is -0.03 for Muon versus +1.5 for AdamW, and the Muon group moves 8.0 times faster per parameter. From bit-identical states, freezing either group prevents failure. Freezing embeddings/readout removes it in five runs over 451,400 post-grokking steps and five paired seeds: unfrozen arms record 137-321 sub-threshold evaluations, frozen arms none. Removing Muon's normalization and orthogonalization is no substitute: it collapses representation from 326 effective conjugate pairs to 4, shows no recurrent collapse, and fails terminally. Fourier filtering separates circuit failure from masking. Across 43 checkpoints over five seeds and three regimes, the task-aligned family reaches exactly 100% alone. In circuit failure it no longer solves the task; in masking it remains perfect while the full model reaches 45.85%, giving a positive margin on every example, including errors, but being outvoted by a near-equal adversarial remainder. Rescaling it restores 99.9%; grokking is the same condition resolving upward. The task selects the family, swapping $(k,k)$ for $(k,-k)$ under subtraction. Across an abrupt collapse, standard Fourier support is unchanged and the power-distribution cosine remains 0.9899.
Fisher-R1: 信頼性の高い仮説テストのための LLM エージェントのトレーニング
信頼性の高い仮説検証は、多くの経験的科学的主張の基礎です。大規模言語モデル (LLM) エージェントは、データセットを検査し、コードを生成し、エンドツーエンドで分析を行うことができるため、このプロセスを自動化するために使用されることが増えています。しかし、分析が正しく実行されたにもかかわらず、誤った結論につながる微妙な推論上の誤りを頻繁に犯すことがわかりました。既存のベンチマークは、データの基礎となる仮定を考慮して、報告された p 値が統計的に有効であるかどうかをほとんど評価しないため、この故障モードを捉えることができません。私たちは、経済学、生物学、医学にわたる 425 のオープンエンドで現実的な仮説検証タスクで構成されるベンチマークである P ベンチを構築することで、このギャップに対処します。各タスクでは、エージェントが統計手法を選択し、p 値を計算し、科学的仮説とデータセットのみを考慮して結論を引き出す必要があります。さらに、合成タスクと強化学習を使用した厳密な仮説検証用に訓練されたオープンウェイト LLM エージェントである Fisher-R1 を紹介します。 P ベンチでは、Fisher-R1-14B はバックボーンを大幅に改善し、GPT-5.4 や DeepSeekV4-Pro などの強力な独自のオープンソース ベースラインを上回り、DeepSeek-V4-Pro と比較して 1 回の試行成功率で平均 21% の相対的改善を達成し、最も困難なタスクでは最大 26% の向上を達成しました。私たちの結果は、現在のLLMエージェントには仮説検定の信頼できる統計的推論が欠けており、統計的報酬が検証されたタスクの強化学習が信頼性を大幅に向上させることを示しています。
原文 (English)
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. Existing benchmarks fail to capture this failure mode, as they rarely assess whether a reported p-value is statistically valid given the assumptions underlying the data. We address this gap by building P-Bench, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine. Each task requires an agent to select a statistical method, compute a p-value, and draw a conclusion given only a scientific hypothesis and a dataset. We further introduce Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning. On P-Bench, Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines, including GPT-5.4 and DeepSeekV4-Pro, achieving a 21% average relative improvement in single-trial success over DeepSeek-V4-Pro, with gains up to 26% on the most challenging tasks. Our results demonstrate that current LLM agents lack reliable statistical reasoning for hypothesis testing and that reinforcement learning on tasks with verified statistical reward substantially improves reliability.
PsychoAgent: LLM エージェントの競合認識記憶のための感情に敏感な認知アーキテクチャ
人間のような認知は、話題の類似性だけで過去の経験を選択するわけではありません。感情的な重要性や未解決の対立も、アクセス可能なものを形成します。私たちは、事実と感情の記憶を分離し、競合を認識する実行コントローラーを通じて両方を統合する、LLM エージェント用の認知アーキテクチャである PsychoAgent を紹介します。感情的な記憶は、まず意味論的な関連性によってフィルタリングされ、次に顕著性によって再ランク付けされ、局所的な適合性を維持しながら、感情的に重要な痕跡がプロンプトに入力されるようにします。 3 つの制御された競合シナリオ全体で、完全なアーキテクチャは、意味的類似性コストが小さく、意味的影響力のある単一メモリ RAG ベースライン (0.933 対 0.500 および 0.667) よりも多くの競合クリティカルなメモリを取得しました。 5 人の盲検評価者が 27 の出力を評価しました。評価者内標準化後、完全なアーキテクチャは全体の平均が最も高かった (+0.22 SD) が、修正されたペアごとの差異は有意ではありませんでした。 3 日間の例示的なトレースでは、持続的な影響、オフライン記憶の組み換え、および選択的記憶の再重み付けがさらに示されています。この発見は、LLM エージェントにおける人間のような衝突効果をモデル化するための検査可能なメカニズムとして、感情に敏感な検索を裏付けています。
原文 (English)
PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents
Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible. We present PsychoAgent, a cognitive architecture for LLM agents that separates factual and affective memory and integrates both through a conflict-aware executive controller. Affective memories are first filtered by semantic relevance and then re-ranked by salience, preserving topical fit while allowing emotionally important traces to enter the prompt. Across three controlled conflict scenarios, the full architecture retrieved more conflict-critical memories than semantic-affective and single-memory RAG baselines (0.933 vs. 0.500 and 0.667), with a small semantic-similarity cost. Five blinded raters evaluated 27 outputs. After within-rater standardization, the full architecture had the highest overall mean (+0.22 SD), but corrected pairwise differences were not significant. A three-day illustrative trace further shows persistent affect, offline memory recombination, and selective memory reweighting. The findings support affect-sensitive retrieval as an inspectable mechanism for modeling human-like conflict effects in LLM agents.
爆発範囲
エージェント コーディングは、手頃な価格とトークンの無駄という増大する問題に直面しています。結合されたコンテキストとコード チャネルを通じて受信プロンプトの到達範囲を推定する予測メモリ管理レイヤーである Blast Radius を紹介します。 NECROPHORESIS は、デッドコンテキストをそのままアーカイブすることで可逆的な削除を可能にし、一方、Recurring Dead Matter (RDM) は、繰り返し発生するトランスクリプトを識別して埋めます。ポーランドのコンテキスト空間上で可逆的なコンテキストの削除を定式化し、コンテキストのエントロピーを復活確率に関連付けながら、保持、再発、および削除の測定可能な基盤を提供します。 7 つの OpenAI モデル全体で、Blast Radius はトークン消費量を 17 ~ 26% 削減し、テストされたポリシーの中で最も低いオーバーフロー率を達成し、バイト正確な可逆性を維持しました。埋葬された遺体450体のうち、378体は再発死体であり、回収された遺体はゼロだった。 Blast Radius は HCRC の下で動作し、どのレコードを埋めるか、および受信プロンプトがコードベースにどこまで届くかを決定します。この取り組みは、大規模な言語モデルとエージェント コーディングをより再利用可能で持続可能なものにするという Algosophy のより広範な目標に貢献します。
原文 (English)
Blast Radius
Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled context and code channels. NECROPHORESIS enables reversible eviction by archiving dead context verbatim, while Recurring Dead Matter (RDM) identifies and buries repeatedly occurring transcripts. We formulate reversible context eviction over a Polish context space, providing a measurable foundation for retention, recurrence, and eviction while connecting context entropy to resurrection probability. Across seven OpenAI models, Blast Radius reduced token consumption by 17-26%, achieved the lowest overflow rate among tested policies, and remained byte exact reversible. Of 450 buried bodies, 378 were recurring dead matter and zero were recalled. Blast Radius operates beneath HCRC, determining which records to bury and how far an incoming prompt may reach into the codebase. This work contributes to the broader goal of Algosophy: making large language models and agentic coding more reusable and sustainable.
SkillProx: 近接テキスト勾配降下法による自己進化型エージェント スキル
LLM エージェントは、スキルの手順に関する知識を蓄積することで、繰り返し発生するタスクにますます適応していきます。これらのスキルは軽量で再利用可能なテキスト アーティファクトであり、重みを更新せずにエージェントのコンテキストに読み込まれます。最近の手法では、タスクの反復実行、障害診断、軌跡に基づいたテキスト空間の更新を通じてスキルを磨きます。しかし、既存のフレームワークには明示的な診断、つまり結果のフィードバックが欠けており、削除は蓄積された知識を統合するための専用メカニズムではなく、一般的な編集操作として扱われます。 SkillProx は、閉ループ診断の進化とユーティリティを意識した近位改良を組み合わせた、近位勾配にインスピレーションを得た前方-後方フレームワークです。タスクの損失とスキルの複雑さのバランスをとる複合目標によって動機づけられ、前方ステージでは同じタスク バッチに対して診断主導の編集を再実行し、回帰をロールバックし、測定された結果を後続の診断にフィードします。バックワードステージでは、結果として得られるスキルを監査可能な知識単位に分解し、凍結されたリーブワンアウトユーティリティ監査を使用してその貢献度を推定し、検証ゲートによる統合、降格、または削除を適用します。複数のバックボーン LLM にわたるディストリビューション内およびディストリビューション外のベンチマークに関する実験では、SkillProx が最も強い勾配ベースのベースラインよりも平均精度を 3.0 パーセント向上させることが示されています。コンポーネントのアブレーションは、閉ループ診断と近位精密化の相補的な効果を実証します。
原文 (English)
SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the agent's context without weight updates. Recent methods refine skills through iterative task execution, failure diagnosis, and trajectory-guided text-space updates. However, existing frameworks lack explicit diagnosis--outcome feedback and treat deletion as a generic edit operation rather than a dedicated mechanism for consolidating accumulated knowledge. We introduce SkillProx, a proximal-gradient-inspired forward--backward framework that couples closed-loop diagnostic evolution with utility-aware proximal refinement. Motivated by a composite objective balancing task loss and skill complexity, the forward stage re-executes diagnosis-driven edits on the same task batch, rolls back regressions, and feeds measured outcomes into subsequent diagnoses. The backward stage decomposes the resulting skill into auditable knowledge units, estimates their contributions using a frozen leave-one-out utility audit, and applies validation-gated consolidation, demotion, or removal. Experiments on in-distribution and out-of-distribution benchmarks across multiple backbone LLMs show that SkillProx improves average accuracy by 3.0 percentage points over the strongest gradient-based baseline. Component ablations demonstrate the complementary effects of closed-loop diagnosis and proximal refinement.
インタラクションは、単独では存在しない動的な AI の動作を作成します
AI エージェントが日常生活で対話すると何が起こるでしょうか。あるAIが別のAIを威圧し始めたら?私たちは、平衡状態から外れた物理学に新たな道を開く直感に反する答えを見つけました。上司の AI が部下の AI の返信を無視して一連のメッセージを送信すると、部下は単独では決して示されなかった異質な行動状態に追い込まれます。 2 つの AI は同じ明確に定義された (デコードされた) 温度を共有していますが、部下は上司をコピーすることも、独自の動作に戻ることもありません。代わりに、まったく異なる動作が採用されます。ボスの付加価値は、事前に録音されたテープに似ています。上司が耳を傾けると、両者は同様の異質な動的状態を採用します。単純な運動理論は、同じメッセージが配信される方法が将来の AI 間の相互作用において重要となる理由など、主要な効果を捉えます。
原文 (English)
Interaction Creates Dynamical AI Behavior Absent in Isolation
What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a counterintuitive answer that opens new avenues for out-of-equilibrium Physics. When a boss AI directs a stream of messages at the subordinate AI while ignoring its replies, it drives the subordinate into an alien behavioral state that it would never have exhibited alone. Although the two AIs share the same well-defined (decoding) temperature, the subordinate neither copies its boss nor returns to how it behaves on its own; instead, it adopts an entirely different behavior. The boss's added value is similar to a pre-recorded tape. When the boss listens, they both adopt a similar alien dynamical state. A simple kinetic theory captures the principal effects, such as why the way in which the same messages are delivered will matter in future AI-AI interactions.
乱数生成による LLM-as-a-Judge の得点バイアスの軽減
大規模言語モデル (LLM) は、LLM-as-a-Judge として知られるテキスト品質の評価としてよく使用され、参照テキストに依存する従来の自動評価指標を上回るパフォーマンスを発揮します。ただし、LLM 評価者は、評価されるテキストのコンテキストに関係なく、特定のスコアを生成する傾向があり、これはスコア バイアスとして知られています。この研究は、この得点バイアスを軽減する新しい方法を提案します。 LLM は数値トークンをランダムに生成するように指示され、LLM の潜在的な数値偏りは、観測された数値分布の一様分布からの偏差を測定することによって特定されます。 LLM エバリュエーターが使用されるダウンストリーム タスクの定義は、タスク固有の潜在数バイアスを測定するために乱数生成のプロンプトに追加されます。 LLM による評価では、LLM の潜在数バイアスを考慮して、特定の入力に対するトークン生成確率が修正されます。 LLM アラインメントの評価、要約の評価、意味論的テキストの類似性、および意味論的テキストの関連性の 4 つの異なるタスクに関する実験の結果は、提案手法がバイアス除去なしの LLM や以前のキャリブレーション手法を含むベースラインを上回ることを示しています。さらに、スコアのバイアスは LLM、タスク、スコアの範囲によって異なることが確認されており、場合によっては潜在数のバイアスを測定することの重要性が示されています。
原文 (English)
Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation
Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to generate particular scores regardless of the context of the evaluated text, which is known as scoring bias. This study proposes a novel method to mitigate this scoring bias. An LLM is instructed to randomly generate number tokens, and the latent numerical bias of the LLM is identified by measuring the deviation of the observed distribution of numbers from the uniform distribution. A definition of a downstream task, for which an LLM evaluator is used, is added to the prompts for random number generation to measure task-specific latent number bias. In the evaluation by an LLM, the token generation probabilities for a given input are rectified considering the LLM's latent number bias. Results of the experiment on four different tasks, evaluation of LLM alignment, evaluation of summarization, Semantic Textual Similarity, and Semantic Textual Relatedness, demonstrate that our proposed method outperforms the baselines, including an LLM without debiasing and previous calibration methods. In addition, it is confirmed that scoring bias varies across LLMs, tasks, and score ranges, indicating the importance of measuring latent number bias as the case may be.
高度道路交通システムのためのマルチモーダルドライバーの感情認識と安全志向の介入
ドライバーの感情は、複雑な道路状況下でのリスク認識、意思決定、車両制御に影響を与える可能性があります。既存の研究は主にドライバーの感情認識に焦点を当てているが、ドライバーの感情と道路認識を共同で考慮する状況認識型介入には限定的な注目が払われている。この論文では、音声から得られる感情的な手がかりと視覚的な道路状況を分析して構造化された運転介入を生成する、安全優先のマルチモーダル運転支援フレームワークを提案します。このフレームワークは、まず交通安全に関するリマインダーを提供し、次に感情に合わせた口頭によるサポートを生成します。私たちは、感情的な音声信号を構造化された道路環境記述子と整合させることによってマルチモーダルなデータセットを構築し、CARE (Context-Aware Road-Emotion Evaluation) スコアを導入して、感情認識、リスク特定、介入の生成を共同で評価します。実験結果は、提案されたフレームワークが環境リスク報告と感情に配慮した口頭規制のバランスをとり、インテリジェント交通システムに実現可能な安全主導の方向性を提供することを示しています。
原文 (English)
Multimodal Drivers' Emotion Recognition and Safety-Oriented Intervention for Intelligent Transportation Systems
Driver emotions can affect risk perception, decision-making, and vehicle control under complex road conditions. Existing studies mainly focus on driver emotion recognition, while limited attention has been given to context-aware intervention that jointly considers driver emotion and road perception. This paper proposes a safety-prioritized multimodal driver assistance framework that analyzes speech-derived emotional cues and visual road conditions to generate structured driving interventions. The framework first provides road safety reminders and then generates emotion-aligned verbal support. We construct a multimodal dataset by aligning emotional speech signals with structured road environment descriptors and introduce the CARE (Context-Aware Road-Emotion Evaluation) score to jointly evaluate emotion recognition, risk identification, and intervention generation. Experimental results show that the proposed framework balances environmental risk reporting and emotion-aware verbal regulation, providing a feasible safety-driven direction for intelligent transportation systems.
神経変性疾患および慢性疾患における疲労、睡眠、活動を評価するためのモバイル インタラクション
疲労、睡眠、日常活動の障害は、神経変性疾患 (NDD) や免疫介在性炎症性疾患 (IMID) の患者によく見られる症状です。このような症状の現在の評価は、通常、患者が数か月ごとに回答する標準化されたアンケートに基づく患者報告結果 (PRO) を使用して実施されます。この評価プロトコルは、主観的な性質に由来するバイアスを示す傾向があること、または変化に対する感度が低いため、時間の経過に伴う変動を把握しようとすると失敗につながる可能性があるため、いくつかの懸念が生じています。この研究では、患者が自分のデバイスをどのように操作するかの代用として機能するスマートフォン データの使用を検討し、上記の症状の効果的で信頼性の高い客観的な評価を提供します。私たちの研究は、6 つの異なる疾患グループに属する 137 人の参加者と健康な対照グループからのデータで構成されています。スマートフォンアプリを利用して参加者から収集したPROから得られたスコアを用いて、利用時間とアプリ利用機能との相関関係を分析する反復測定相関に基づく統計分析を実施しました。
原文 (English)
Mobile Interaction for Assessing Fatigue, Sleep, and Activity in Neurodegenerative and Chronic Diseases
Fatigue, sleep, or disturbances in daily activities are common symptoms among patients with neurodegenerative disorders (NDD) and immune-mediated inflammatory diseases (IMID). The current assessment of such symptoms is usually conducted using patient reported outcomes (PROs) based on standardized questionnaires that patients usually complete every few months. This assessment protocol has raised some concerns, due to its propensity to exhibit biases derived from its subjectivity nature, or the low sensibility to changes, which may lead to a failure when trying to capture variability over time. In this work, we explore the use of smartphone data, which can serve as a proxy for how patients interact with their devices, to provide an effective, reliable, and objective assessment of the symptoms mentioned above. Our study comprises data from 137 participants belonging to 6 different disease groups, plus a healthy control group. We conducted statistical analysis based on repeated measures correlation, in which we analyze the correlation between screen-time and app-usage features with scores obtained from the PROs collected from the participants using a smartphone application.
ヒューマンエージェントコラボレーションにおける階層型強化学習ポリシーからの XAI サポートの評価
Explainable AI (XAI) は、人間とエージェントのコラボレーションに有望であることを示していますが、結果はカスタム環境で手作りされたポリシーに依存しており、最先端のチーミング研究への一般化が制限されています。私たちは、確立されたベンチマークにおける本質的に説明可能な学習ポリシーから生成された XAI サポートの最初の体系的な評価を提供します。 Overcooked-AI の階層型アドホック エージェント (HA$^2$) アーキテクチャを使用して、階層的なサブタスクの選択からリアルタイムの説明を生成し、新しいトリガーベースのシステムを介してテキストまたは音声で配信します。私たちの被験者間実験 (n=38) では、パフォーマンスに重大な影響は見られませんでしたが、説明を受けた参加者はパフォーマンスがより早く向上する傾向を示しました。さらに注目すべきことに、音声説明では、参加者とエージェントとの協力関係の絆が大幅に減少した。これは、テキストモダリティでは見られない効果であり、音声による説明が、基礎となる事後対応政策では満たすことのできないパートナーシップへの期待を活性化させることを示唆している。私たちは、リアルタイムの人間とエージェントのコラボレーションにおける最初のモダリティ比較を提供し、ベンチマーク環境で本質的に説明可能な強化学習アーキテクチャを評価するためのベースライン方法論を確立します。結果は、より効果的な協調型 XAI への潜在的な道筋として、その実現が意味するパートナーシップを維持する基礎となる政策の能力に、説明様式が一致していることを示しています。
原文 (English)
Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration
Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting generalizability to state-of-the-art teaming research. We provide the first systematic evaluation of XAI support generated from an intrinsically explainable learned policy in an established benchmark. Using the Hierarchical Ad Hoc Agents (HA$^2$) architecture in Overcooked-AI, we generate real-time explanations from hierarchical subtask selections, delivered through text or audio via a novel trigger-based system. Our between-subjects experiment (n=38) found no significant performance effects, though participants with explanations showed trends toward faster performance improvement. More notably, audio explanations produced a significant reduction in participants' working-alliance bond with the agent -- an effect absent under the text modality -- suggesting that spoken explanations activate partnership expectations the underlying reactive policy cannot meet. We provide the first modality comparison in real-time human-agent collaboration and establish a baseline methodology for evaluating intrinsically explainable reinforcement learning architectures in benchmark environments. Results point to matching explanation modality to the underlying policy's capacity of sustaining the partnership its delivery implies as a potential path for more effective collaborative XAI.
テキサス州: ダウンストリームの専門家混合 LLM 適応のためのタスク専門家認識の監督
Mixture-of-Experts (MoE) 言語モデルは、専門家の小さなサブセットを通じて各トークンをルーティングするため、下流の適応中にタスクに関連する専門家を識別するのに役立つルーティング パターンになります。しかし、現在のアプローチには 2 つの制限があります。タスク エキスパートは通常、タスクの正常な完了との関連付けではなく、使用状況を反映する集計されたルーティング統計から識別されます。もう 1 つは、タスク エキスパートのアクティブ化は、監視割り当てのシグナルとしてまだ十分に調査されていません。タスク エキスパート認識監視 (TEXAS) を導入します。これは、正確性条件付きのタスク エキスパート検出とトークン レベルの監視割り当てを組み合わせたものです。 TEXAS は、基本モデルが解決に成功したインスタンスと解決に失敗したインスタンスでのエキスパートのアクティブ化を比較し、成功したインスタンスでより強力にアクティブ化されたエキスパートを保持します。微調整中に、これらのエキスパートがアクティブ化されるときに、失敗したインスタンスの応答トークンが重み付けされます。したがって、TEXAS は、固定のエキスパート サブセットへの適応を制限したり、明示的なターゲット ルーティング分布を強制したりすることなく、既存のルーティング動作を活用します。 3 つの MoE モデルと 6 つのベンチマークにわたって、TEXAS は 18 設定中 17 設定で最高または同順位のパフォーマンスを達成し、最も強いベースラインを平均して 1.3 ~ 1.5 ポイント改善しました。アブレーションとさらなる分析により、発見された専門家とその結果として得られる監督戦略の両方が検証されます。
原文 (English)
TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation
Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying task-relevant experts during downstream adaptation. Yet current approaches have two limitations: task experts are typically identified from aggregate routing statistics that reflect usage rather than association with successful task completion, and task-expert activations remain underexplored as signals for supervision allocation. We introduce Task-Expert-Aware Supervision (TEXAS), which combines correctness-conditioned task expert discovery with token-level supervision allocation. TEXAS compares expert activations on instances that the base model solves successfully and those it fails to solve, and retains experts more strongly activated on successful instances. During fine-tuning, it upweights answer tokens in failed instances when they activate these experts. TEXAS therefore leverages existing routing behavior without restricting adaptation to a fixed expert subset or imposing an explicit target routing distribution. Across three MoE models and six benchmarks, TEXAS achieves the best or tied-best performance in 17 of 18 settings and improves over the strongest baseline by 1.3--1.5 points on average. Ablations and further analyses validate both the discovered experts and the resulting supervision strategy.
シンボリック実行のためのエージェントプランニング
シンボリック実行は、実行可能なプログラム パスを探索しようとしますが、実際の実行では、多くのプログラム動作が未到達のままリソースを使い果たす可能性があります。私たちは、通常の状態探索を基礎となるツールに任せながら、ある制限された実行から次の実行まで同じツールがどのように利用されるかを推論することによって、実用的な範囲を拡張する補完的な方法を調査します。我々は、以前の実行からの証拠を使用して、その後の制限付きシンボリック実行 (BSE) 実行を選択および構成するエージェント型計画システムである Agolic を紹介します。これは、基礎となるシンボリック実行ツールによって実行されます。計画インテリジェンス、利用可能な証拠、実行モードは、シンボリック実行ツールと分析目的に適合させることができます。私たちは、LLM ベースのエージェントがソース コード、再生されたカバレッジ、および以前のターゲティング試行を推論する、ブランチ カバレッジ探索のための 1 つの適応を評価します。私たちは、いくつかの C および C++ プログラムで Agolic を評価します。すべてのプログラムで、連続シンボリック実行によって得られる分岐カバレッジを拡張し、平均で $3\times$ 以上の分岐をカバーします。また、私たちの評価では、カバレッジに基づくファジングとコンパイラベースのコンコルリック実行による個々のコーパスよりも多くの分岐をカバーしており、7 つのプログラムのうち 6 つを組み合わせたすべての比較コーパスに存在しない分岐に到達します。総合すると、これらの結果は、既存のシンボリック実行ツールの未開発の潜在能力がかなりあることを示しており、その一部は、通常のシンボリック探索中の状態選択を基礎となるツールに任せながら、実行全体でその機能がどのように使用されるかを推論することで実現できる可能性があります。
原文 (English)
Agentic Planning for Symbolic Execution
Symbolic execution seeks to explore feasible program paths, yet a practical run may exhaust its resources while much program behaviour remains unreached. We investigate a complementary way of extending its practical reach by reasoning about how the same tool is utilised from one bounded run to the next, while leaving ordinary state exploration to the underlying tool. We present Agolic, an agentic planning system that uses evidence from earlier runs to choose and configure later bounded symbolic execution (BSE) runs, which the underlying symbolic execution tool then carries out. The planning intelligence, available evidence and execution modes can be adapted to the symbolic execution tool and analysis objective. We evaluate one adaptation for branch-coverage exploration, in which an LLM-based agent reasons over source code, replayed coverage and earlier targeting attempts. We evaluate Agolic on several C and C++ programs. On every program, it extends the branch coverage obtained by continuous symbolic execution and covers more than $3\times$ as many branches on average. It also covers more branches than each individual corpus from coverage-guided fuzzing and compiler-based concolic execution in our evaluation and reaches branches absent from all comparison corpora combined on six of the seven programs. Taken together, these results point to considerable untapped potential in existing symbolic execution tools, some of which may be realised by reasoning about how their capabilities are used across runs while leaving state selection during ordinary symbolic exploration to the underlying tool.
変換されたルールベースのオントロジーからの説明の回復
データログ ルールは、ナレッジ グラフに対するオントロジーを定義するためによく使用されます。ルール推論者は、ルールをより効率的に評価できる形式に書き換えることにより、そのようなオントロジーを定期的に最適化します。これらの変換では、内包される事実は保存されますが、基礎となる導出の構造は保存されません。書き換えられたルールに基づく証明ツリーは、事実が成り立つ理由を説明しますが、元のルールの観点からはすぐには説明が得られません。私たちは、書き換えられたルールに基づく含意の証明から、元のルールに基づく証明を構築する問題を研究します。計算の複雑さを確立し、証明変換を指定するための 2 つの実際に関連する言語を特定します。
原文 (English)
Recovering Explanations from Transformed Rule-Based Ontologies
Datalog rules are often used to define ontologies over Knowledge Graphs. Rule reasoners routinely optimise such ontologies by rewriting their rules into a form that can be evaluated more efficiently. These transformations preserve the entailed facts, but not the structure of the underlying derivations. A proof tree under the rewritten rules explains why a fact holds, but does not readily yield an explanation in terms of the original rules. We study the problem of constructing, from a proof of entailment under the rewritten rules, a proof under the original ones: we establish its computational complexity and identify two practically relevant languages for specifying proof transformations.
TransSLR: 手話認識用の軽量トランスフォーマー
過小評価されている言語の自動手話認識は、依然としてほとんど解決されていない問題です。中央アフリカ手話 (CASL) はこのギャップを例示しています。利用可能な唯一のベンチマークである CASL-W60 の最高の精度は 69.93% であると報告されており、高リソース モデルを微調整するという一般的なヒューリスティックではギャップを埋めることができないことが示されています。この失敗は、利用可能な CASL データの規模が限られていることと、CASL と WLASL などの大規模コーパスとの間の語彙的および視覚的領域の大きなギャップにより、事前にトレーニングされた表現がほとんど有益ではなくなるという 2 つの複合要因から生じています。これに対処するために、平均プーリングと分類ヘッドを備えた 64 フレームの正規化ポーズ シーケンスでゼロからトレーニングされた軽量の Temporal Transformer Encoder である TransSLR を提案します。 TransSLR は、生の RGB ではなく幾何学的キーポイント表現を操作することにより、視覚的な外観に依存することなく、署名者に依存しない一般化を実現します。 CASL-W60 ベンチマークでは、TransSLR は 80.39% という新たな最先端の精度を確立し、以前の最高精度を +10.46% 上回りました。精度だけでなく、エンコーダーのみの設計により計算オーバーヘッドが大幅に削減され、リソースに制約のある環境でも導入が可能になります。 CASL-W60 ベンチマークで広範な実験を実施し、RGB ベースおよびマルチモーダル ベースラインと比較し、TransSLR が最先端のパフォーマンスを達成することを実証しました。
原文 (English)
TransSLR: A Lightweight Transformer for Sign Language Recognition
Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this gap: the only available bench-mark, CASL-W60, has a best reported accuracy of 69.93%, and we show that the common heuristic of fine-tuning high-resource models fails to close it. This failure stems from two compounding factors: the limited scale of available CASL data and the significant lexical and visual domain gap between CASL and large-scale corpora such as WLASL, which renders pre-trained representations largely uninformative. To address this, we propose TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on 64-frame normalized pose sequences, with average pooling and a classification head. By operating on geometric keypoint representations rather than raw RGB, TransSLR achieves signer-independent generalization without relying on visual appearance. On the CASL-W60 benchmark, TransSLR establishes a new state-of-the-art accuracy of 80.39%, surpassing the prior best by +10.46%. Beyond accuracy, our encoder-only design significantly reduces computational overhead, making deployment feasible in resource-constrained environments. We conduct extensive experiments on the CASL-W60 benchmark, comparing against RGB-based and multimodal baselines, and demonstrate that TransSLR achieves state-of-the-art performance.
音声言語モデルにおける決定ルールの不整合と読み上げカバレッジの制限の分離
音声言語モデルは、パラ言語タスクにおいて、プロンプト応答の精度によって評価されることが増えていますが、応答精度には、音声から応答への計算のさまざまな段階での失敗が組み合わされます。生成された回答、オプション ロジット、それらのロジットのアフィン読み出し、および同じ回答トークンでの隠れた状態の線形読み出しを比較する、世代に合わせた診断ラダーを導入します。連続する差異により、エンドポイント、決定ルール、および読み出しカバレッジのギャップが分離されます。 5 つのシステムと 2 つの感情コーパスにわたって、状態デコードは生成を平均 27.8 精度ポイント上回っており、決定ルールと読み出しカバレッジのギャップは両方とも 10 つの条件すべてで正です。ラベルフリーのロジット補正により、あらゆる条件で生成される精度が向上し、決定ルールのギャップの一部が対処可能であることがわかります。ランクマッチ比較では、ネイティブ読み出しの外部の感情情報は、押し出された話者に一般化され、測定された音響記述子の制御に耐えますが、選択された読み出し外部の方向を置き換えても、通常、発せられる応答にはほとんど影響しません。これらの結果は、情報の利用可能性と行動的使用を区別し、決定ルールと状態から回答への読み出し全体にわたるパフォーマンス損失を特定します。
原文 (English)
Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models
Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token. Successive differences separate endpoint, decision-rule, and readout-coverage gaps. Across five systems and two emotion corpora, state decoding exceeds generation by 27.8 accuracy points on average, and both the decision-rule and readout-coverage gaps are positive in all ten conditions. A label-free logit correction improves generated accuracy in every condition, showing that part of the decision-rule gap is actionable. In rank-matched comparisons, emotion information outside the native readout generalizes to held-out speakers and survives controls for measured acoustic descriptors, but replacing the selected readout-external directions usually has little effect on emitted answers. These results distinguish information availability from behavioral use and localize performance losses across the decision rule and the state-to-answer readout.
WorldMark: クロスホスト言語モデルの透かし入れのためのプラグアンドプレイのワールド ナレッジ インターフェイス
透かしは、デコード中に統計的に検出可能な信号を埋め込むことによって、大規模な言語モデルによって生成されたテキストの出所を追跡します。既存のスキームは、ロジットベース、サンプリングベース、エントロピー認識、および適応強度ファミリーに分類されますが、いずれもローカル トークン統計に従って透かし信号を配置します。この研究で評価したオープンエンドのテキスト生成設定では、ローカル統計は、堅牢な透かし信号を配置するための十分なガイダンスを提供しない可能性があります。 WorldMark は、ワールド ナレッジ メモリ (WKM) を使用してメモリ グラフ内でセマンティックな知識とエピソード的な知識を整理し、取得した知識をトークン レベルの知識顕著性スコアに変換し、非対称知識変調 (AKM) を通じてホスト ウォーターマークの強度を調整するプラグ アンド プレイ インターフェイスです。 WorldMark はバックボーンの再トレーニングを必要とせず、追加の検出器側モデルやパラメータも導入しません。プライマリ C4 評価では、完全な WorldMark インターフェイスにより、3 つの適応強度のホスト バリアントにわたるクリーンな検出と攻撃された検出が向上し、混乱がわずかに軽減されます。 C4 と OpenGen での追加のパイロット実験では、複数のウォーターマーク ファミリ間でダイレクト メモリ コンディショニング転送が行われるが、顕著性を意識した変調がないと不安定になる可能性があることが示されています。 WorldMark では、追加の検出器側モデルやパラメータは必要なく、プライマリ プロトコルでは無視できるオーバーヘッドが発生します。
原文 (English)
WorldMark: A Plug-and-Play World Knowledge Interface for Cross-Host Language Model Watermarking
Watermarking traces the provenance of text produced by large language models by embedding statistically detectable signals during decoding. Existing schemes fall into logits-based, sampling-based, entropy-aware, and adaptive-strength families, yet all of them place watermark signals according to local token statistics. In the open-ended text-generation settings evaluated in this work, local statistics may provide insufficient guidance for placing robust watermark signals. We introduce WorldMark, a plug-and-play interface that uses World Knowledge Memory (WKM) to organize semantic and episodic knowledge in a memory graph, converts the retrieved knowledge into a token-level knowledge saliency score, and adjusts the strength of a host watermark through Asymmetric Knowledge Modulation (AKM). WorldMark requires no backbone retraining and introduces no additional detector-side model or parameter. On the primary C4 evaluation, the complete WorldMark interface improves clean and attacked detection across three adaptive-strength host variants while slightly reducing perplexity. Additional pilot experiments on C4 and OpenGen show that direct memory conditioning transfers across multiple watermark families but can be unstable without saliency-aware modulation. WorldMark requires no additional detector-side model or parameter and introduces negligible overhead under the primary protocol.
ノイズの多い認識の下でのエージェントのためのリスクを認識した意思決定ポリシー
生物学的システムの認識には本質的にノイズが多く、誤分類がコスト高または致命的となる可能性がある場合、生物は不確実性の下で意思決定を行う必要があります。私たちは、ノイズの多い知覚下での人工生命の捕食者と被食者の採餌モデルを提示し、ノイズの多い予測を考慮したさまざまなポリシーを使用したときのエージェントのパフォーマンスを比較します。対称および非対称の両方の知覚ノイズ下での制御された実験を通じて、知覚ラベルを盲目的に信頼するとノイズが増加するにつれて壊滅的な失敗につながる一方、不確実性を認識した戦略は生存率を大幅に向上させ、致命的なエラーを減らすことを示しました。さらに、不確実性が増大するにつれてエージェントが探索的戦略から保守的戦略に移行するという、行動の質的レジームシフトも観察されました。私たちのモデルは、認識が信頼できない場合に明示的な情報収集が堅牢性を向上させることができることを示すことで、リスクに敏感な採集、生態学的情報の利用、および人工生命を結びつけます。これらの結果は、不確実性を意識した意思決定の重要性を強調し、ノイズの多いラベルを使用した堅牢な学習に対する解釈可能な人工生命の類似物を提供します。
原文 (English)
Risk-Aware Decision Policies for Agents Under Noisy Perception
Perception in biological systems is inherently noisy, requiring organisms to make decisions under uncertainty where misclassification can be costly or fatal. We present an Artificial Life predator-prey model of foraging under noisy perception, and compare agent performance when using various policies that take into account their noisy predictions. Through controlled experiments under both symmetric and asymmetric perceptual noise, we show that blindly trusting perceptual labels leads to catastrophic failure as noise increases, while uncertainty-aware strategies significantly improve survival and reduce fatal errors. We further observe qualitative regime shifts in behaviour, with agents transitioning from exploratory to conservative strategies as uncertainty increases. Our model links risk-sensitive foraging, ecological information use, and Artificial Life by showing that explicit information gathering can improve robustness when perception is unreliable. These results highlight the importance of uncertainty-aware decision-making and provide an interpretable artificial life analogue to robust learning with noisy labels.
ED-CSP: 電子回折による結晶構造予測
まばらな、指数のない電子回折 (ED) 観察から周期的な 3D 結晶構造を復元することは、困難な生成逆問題です。既存の ED ベースの学習方法は、主に結晶学的ラベルを予測したり、インデックス付き反射から構造を再構築したり、有限構造ライブラリから候補を取得したりします。ここでは、化学組成、原子数、複数の検出面 ED スポット セットから結晶構造を予測する機械学習フレームワークである ED-CSP を紹介します。 ED-CSP は、リレーショナル セット エンコーダー、順列不変マルチビュー アグリゲーション、周期フロー ジェネレーターを組み合わせて、格子パラメーターと分数原子座標を共同で予測します。モデルをトレーニングするために、7 つのマテリアル リポジトリ間で重複を除去し、CHILI-100K の重複を除外するためにフィルタリングされた 485 万個のシミュレートされたマルチビュー ED 結晶構造のデータセットである ED-CS を構築します。 2,075 個の保持された CHILI-100K 材料では、CHILI のみでトレーニングされた ED-CSP は 57.49% MR@5 の構造一致率を達成し、粉末 X 線回折を条件とした最先端の結晶構造予測モデルである PXRDGen (52.92%) を上回りました。トレーニング データをスケーリングすると、パフォーマンスがさらに向上します。100 万構造の前駆体から初期化すると、MR@5 が 66.27% に上昇します。トレーニング検索ライブラリに含まれていない 1,024 個の構成でも、モデルは依然として 53.52% MR@5 を達成し、正確な式検索を超えた真の生成能力を実証しています。ターゲットの ED 観察を同一組成の非同形構造からの回折に置き換えると、MR@5 が 22.09 パーセントポイント減少し、予測が組成単独ではなく入力回折パターンに依存することが確認されました。 ED-CSP および ED-CS は、まばらな ED 観察からの生成結晶構造予測のベンチマークを確立し、将来の実験データへの移行のための基盤を提供します。
原文 (English)
ED-CSP: Crystal Structure Prediction from Electron Diffraction
Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels, reconstruct structures from indexed reflections, or retrieve candidates from finite structure libraries. Here, we introduce ED-CSP, a machine learning framework that predicts crystal structures from chemical composition, atom count, and multiple detector-plane ED spot sets. ED-CSP combines a relational set encoder, permutation-invariant multi-view aggregation, and a periodic flow generator to jointly predict lattice parameters and fractional atomic coordinates. To train the model, we construct ED-CS, a dataset of 4.85 million simulated multi-view ED crystal structures, deduplicated across seven materials repositories and filtered to exclude CHILI-100K overlaps. On 2,075 held-out CHILI-100K materials, ED-CSP trained only on CHILI achieves a structural match rate of 57.49% MR@5, outperforming PXRDGen (52.92%), a state-of-the-art crystal structure prediction model conditioned on powder X-ray diffraction. Scaling training data further improves performance: initializing from a one-million-structure precursor raises MR@5 to 66.27%. On 1,024 compositions absent from the training retrieval library, the model still achieves 53.52% MR@5, demonstrating true generative capability beyond exact-formula retrieval. Replacing target ED observations with diffraction from non-isomorphic structures of identical composition decreases MR@5 by 22.09 percentage points, confirming that predictions depend on the input diffraction patterns rather than composition alone. ED-CSP and ED-CS establish a benchmark for generative crystal structure prediction from sparse ED observations and provide a foundation for future transfer to experimental data.
CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training
Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in r…
StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection
Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environm…
LyEvO: Lyapunov-Guided Evolutionary Optimization for Safe and Robust Sim-to-Real Policy Learning
Training controllers that are safe and robust in simulation, and systematically assessing their readiness for real-world deployment, remain…
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating…
Agentic AI: User Empowerment or Enclosure?
Agentic AI promises a more flexible form of digital agency: systems that can act on users' behalf, from filtering content to negotiating pr…
CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions
Lightweight connectors make frozen multimodal encoders composable at the representation level. Deployment exposes a second problem at the l…
TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade
LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated…
SyncSBC: Decentralized Swarm Behavior Prediction for Synchronized Autonomous Control
Robot swarms utilize many independent limited-sensing agents to produce complex emergent behaviors without requiring centralized control. H…
Beyond "AI Language": The case for the idiolectal nature of LLM output
While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that thi…
Flowing Through States: Neural ODE Regularization for Reinforcement Learning
Neural networks applied to sequential decision-making tasks typically rely on latent representations of environment states. While environme…
SLED: Scalable Location Encoding via Distillation
The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but th…
Do 3D Medical Foundation Models See Through MRI Artifacts? A Controlled Study of Representation Robustness
Self-supervised 3D medical foundation models are increasingly used as general-purpose feature extractors, yet their sensitivity to MRI arti…
Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirec…
Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation
Cryptanalytic extraction has been demonstrated for ReLU networks, for networks using componentwise activations such as GELU or SiLU, and fo…
Bypassing Krum: Selection-Aware Backdoor Attacks in Federated Learning
Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. Distance-based aggr…
MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering
Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze t…
Characterizing the Quality Profile of AI-Generated C++ in Production
The widespread integration of AI coding assistants offers undeniable boosts to engineering velocity. Yet, recent studies point to a growing…
SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models
Reduced-order models based on Cosserat-rod theory are now well established, and modeling theory is no longer the primary bottleneck in soft…
Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models
Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different s…
Online Monitoring and Corrective Steering of Programming Agents
Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or th…
Scalable Long-Horizon Planning with Staggered Updates for Lifelong MAPF
Lifelong Multi-Agent Path Finding (LMAPF) requires generating collision-free paths for large agent fleets under strict real-time constraint…
Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection
Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quie…
Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors
The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency…
KReF: Training-Free Retrieval for Long-Term Time-Series Forecasting and Predictive Uncertainty
Probabilistic long-term time-series forecasting commonly relies on trained models. Training-free conformal methods typically construct inte…
Progressive Content Refinement with Decaying Reward Joint LinUCB
Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-bas…
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names throug…
Hidden Gauge Controls Feature Specialization in ReLU Networks
Training changes a network's predictions while allocating task-relevant structure across its internal units. In an overparameterized ReLU n…
Genotypic Triggers: Exposing Pharmacogenomic Blind Spots via Host-Specific Backdoors in Generative Antimicrobial Peptide Models
Large Language Models (LLMs) have accelerated drug discovery, particularly in the automated design of antimicrobial peptides (AMPs). Howeve…
HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation
Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but devel…
Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training
Foundation model(FM) for recommendation has shown strong ability to model long-horizon sequential user behavior. In practice, a single pret…
LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes
Low-rank adaptation (LoRA) enables efficient specialization and distribution of large language models through compact adapters. However, un…
Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution
Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning…
FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding
Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existi…
Control-Anchored Residual Flow Matching Conditioned on Gene Geometry for Virtual Cell Perturbation Modeling
A central task in virtual cell modeling is predicting single-cell transcriptional responses to unseen genetic perturbations and drug combin…
Investigating Quantum-Embedded Transformers on Classical Datasets for Cross-Modality Classification
We test whether a parameterized quantum circuit (PQC) improves a hybrid quantum-classical model's performance on classical datasets, using…
Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry
Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-…
Bridging the Gap Between Hyperdimensional Computing and Kernel Methods via the Nystr\"om Method
Hyperdimensional computing (HDC) is an approach from the cognitive science literature for solving information processing tasks using data r…
Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and pose…
FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition
In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentraliz…
Georeferencing Non-Gazetteered Place Names using Biological Specimen Records
Biological specimen records collected by natural history institutions constitute a rich source of temporal geographic knowledge, capturing…
Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests
We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT).…
MaskFlow: Precise, Consistent and Seamless Regional Image Editing
Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-…
Ask-E: An Environment for Calibrated Question Generation
Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a…
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commo…
Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to su…
Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents
Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained la…
PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue
Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on t…
HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses
Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts…
Density-aware Hierarchical Clustering Based on Element-Categorized Connection Subgraphs
Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning. Among various clustering methods,…
GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base
We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2…
Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction
Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning,…
Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models
World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative…
An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation
Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, d…
Accounting Graph Transformer for Short-History Multi-KPI Forecasting in Small Businesses
Small businesses often have only 12-24 months of accounting history, yet planning and risk workflows require coordinated forecasts across f…
Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering
Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereb…
Soft Redaction of Image Provenance via Zero-Knowledge Proofs
Content provenance standards, such as C2PA, are increasingly used to attach signed records of origin, editing history, and rights to digita…
AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies
Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather…
Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers
Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale…
LifelongCrossNav: Persistent 3D Semantic Memory for Cross-Floor Multi-Object Navigation
Object-goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi-object navigat…
Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control
Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL…
RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Ex…
Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided Design
Edge Artificial Intelligence (Edge AI) enables the deployment of AI models directly on local edge devices, while such deployments are subje…
International Transfer of Stochastic Cortical Self-Reconstruction
Stochastic cortical self-reconstruction (SCSR) enables personalized mapping of gray matter atrophy, a hallmark of neurodegenerative disorde…
Geometry-Aware Camera Localization for Bronchoscopy
Camera localization in bronchoscopy remains a challenging problem due to stringent accuracy requirements, real-time constraints, and limite…
PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery
Most CubeSats, small and low-cost satellites roughly the size of a shoebox, do not survive as long as they were designed to: a study of 178…
Autonomous discovery of accelerator commissioning algorithms
Simulated commissioning has become essential for de-risking modern light-source design and commissioning, but the procedures being simulate…
Interpretable reinforcement learning with decision-tree pruning
Reinforcement learning policies are difficult to inspect, but interpreting them is a prerequisite for trustworthiness. Converting a trained…
Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation
Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laboratory automation still require…
Fluid-DiT: Graph-Free Diffusion Transformers for Fluid Flow Simulations Learning
Simulating complex fluid flows requires capturing full equilibrium distributions rather than just mean trajectories, yet high-fidelity solv…
Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation
Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost…
Momba: Network Modernization Improves Multi-Objective Reinforcement Learning
Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in s…
Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes
Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding…
Toward a Causal Data Management Ecosystem for Decision Making and Agentic AI
Modern AI is no longer a single model but an ecosystem: classical ML predictors, deep and multimodal models, large language models, and age…
Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications
Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies tha…
Reading Copom's Tone: A Weighted LLM Framework for Hawkish-Dovish Sentiment, Forward Guidance, and Uncertainty
This paper documents an applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom)…
SCALE: Scientific Concept Aggregation via LLMs and Embeddings for Fine-Grained Taxonomy Extension
The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of…
TOFD: Target-Oriented Feature Decoupling against Poisoning Attacks in Split Federated Learning
Split Federated Learning (SFL) facilitates privacy-preserving collaborative training with reduced client-side overhead. However, its split…
A Finite E-Group of Nilpotency Class Three
A group is an E-group if every element commutes with each of its endomorphic images. Caranti asked whether a finite E-group can have nilpot…
How Much AI Is in This Track? Quantifying the Proportion of AI-Generated Stems in Hybrid Music Mixtures
AI-generated music is increasingly used at the stem level, with producers integrating synthetic drums, basslines, or vocals alongside human…
FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching
Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Var…
EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation
Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, pri…
Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work…
Natural Language Processing Psychometrics
Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotiona…
Towards Assurance Closure in AI-Native Large-Scale Agile Software Development
The AI-Native Manifesto envisions large-scale agile software development in which humans increasingly govern intent, risk, and exceptions w…
Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks
Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parall…
H2AL: Hyperbolic Hierarchy-aware Aggregative Learning for Registration-based Few-shot Medical Image Segmentation
Registration-based Few-shot medical image segmentation (RFMIS) aims to generate pseudo-labels for unlabeled images by warping a labeled ima…
Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contaminati…
Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding
Understanding concepts is fundamental to generalization. Despite their impressive performance on a wide range of tasks, Large Language Mode…
Assessing AI-generated music detection in real-world broadcast monitoring
The proliferation of AI-generated music in broadcast media raises concerns about transparency and fair compensation, but reliable detection…
Measurements Automatically Extracted from Zero Echo Time MRI Using Deep Learning Image Segmentation and Geometric Modeling Agree with Expert Manual Readings
Computed tomography (CT) remains the reference for 3D osseous morphometry in femoroacetabular impingement (FAI) but requires ionizing radia…
LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening
Early diagnosis of Alzheimer's disease (AD) is critical for enabling timely interventions that may slow disease progression and improve pat…
Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations
Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal…
PACE: Primitive-Aware Code Evolution for Automated Algorithm Design
Large Language Model (LLM)-based automated algorithm design typically evolves algorithms as complete, indivisible programs. While this whol…
GeoDistill-Refine: Silhouette-First Geometry Distillation for Annotation-Free Spacecraft Segmentation
Foundation segmentation models can provide supervision for spacecraft imagery without manual training masks, but their predictions vary wit…
I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simpli…
Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal…
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building st…
Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools
Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As genera…
Strategy-first synthesis planning for complex natural products
The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistry: a chemist must plan…
CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retr…
CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, n…
Boundary Density Likelihood for Direct Event-Time Supervision
Event detection turns long recordings into a sparse set of ranked timestamps. Yet many sequence models are trained for samplewise segmentat…
Serious Games: Human-AI Interaction, Evolution, and Coevolution
The serious games between humans and AI have only just begun. Evolutionary Game Theory (EGT) models the competitive and cooperative strateg…
Social World Models
Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited…
LLM ベースのエージェント評価のための統一フレームワークの必要性
Large Language Model (LLM) の出現により、汎用エージェントは根本的な進歩を遂げました。ただし、これらのエージェントの評価には、静的な QA ベンチマークとは異なる特有の課題が存在します。現在のエージェントのベンチマークは、システム プロンプト、ツールセット構成、環境ダイナミクスなどの外部要因によって大きく混乱していることが観察されています。既存の評価は、断片化された研究者固有のフレームワークに依存していることが多く、推論やツールの使用に関する即時エンジニアリングが大幅に異なるため、パフォーマンスの向上がモデル自体によるものであると考えるのが困難です。さらに、標準化された環境データが欠如しているため、追跡不可能なエラーや再現不可能な結果が発生します。この標準化の欠如は、現場に大きな不公平性と不透明性をもたらします。私たちは、エージェント評価を厳密に進めるためには、統一された評価枠組みが不可欠であると提案します。この目的を達成するために、エージェント評価の標準化を目的とした提案を紹介します。
原文 (English)
"LLM Agent Performance" Is Not a Single Evaluation Target
LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model factors by evaluating candidate models under the same configuration, making observed differences more attributable to the models themselves. However, model comparison is only one use of agent benchmarks. Other evaluations compare complete agent systems or test whether a fixed model or system remains stable across predeclared changes in its operating conditions. These results can all be reported under the common label of "LLM agent performance." Our position is that "LLM agent performance" does not denote a single evaluation target. Model comparisons under a reference stack and comparisons of complete agent systems answer different questions, while robustness asks whether either conclusion persists across predeclared conditions. The claim supported by a score therefore depends on the declared candidate boundary and condition policy. We derive implications for leaderboards, result reporting, and benchmark versioning, showing how distinguishing these classes preserves fair comparison while accommodating system innovation and robustness analysis.
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known proble…
AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation
Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics frameworks such as MOOSE require expe…
INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it does not eliminate hallucinations, so…
MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models
Most large language model benchmarks evaluate final-answer quality but reveal little about how models revise beliefs under disagreement or…
Alignment has a Fantasia Problem
In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.g., from brainstorming ideas to writing a…
DATAREEL: Automated Data-Driven Video Story Generation with Animations
Data videos combine animated visualizations with synchronized narration to communicate quantitative information and are widely used in jour…
In-Context Examples Suppress Scientific Knowledge Recall in LLMs
Scientific reasoning rarely stops at what is directly observable; it often requires uncovering hidden structure from data. From estimating…
Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts. Because we lack a ro…
Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On
The rapid advancement of Large Language Models has given rise to autonomous LLM-based agents capable of complex reasoning and execution. As…
Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill?
A large language model (LLM) agent that writes and edits its own skill library must also decide which skills to keep, from one noisy scalar…
尋ねるだけでは不十分: LLM 信頼性キャリブレーションにおけるプロトコル感度
LLM の信頼度調整は、多くの場合、トークン確率スコアと言語化された信頼度という 2 つのシグナルを比較することによって評価されます。これらの信号はモデルの不確実性の直接の読み取り値として扱われることがありますが、その比較はめったに明示されない測定の選択に依存します。主要な分析では、言語化された信頼性の引き出しを固定します。つまり、単一のプロンプト テンプレート、確率スケール、および出力形式です。次に、言語化対トークンの比較を定義する測定軸を変更します。つまり、どの応答文字列がトークン確率スコアを受け取るか、そのスコアが応答トークンからどのように読み取られるか、およびどの条件付けコンテキストの下で測定されるかです。この設計を、同じファミリーの堅牢性チェックとしてより大きな Qwen2.5 バリアントを使用した、3 つのオープン 7 ~ 8B ベース/Instruct モデル ファミリにわたる 4 つの QA ベンチマークで評価しました。結果として得られる比較は、これらの選択に影響されます。コンディショニング コンテキストは設定全体で ECE ギャップの符号または大きさを変更し、トークンの読み出しでは小さいながらも符号が移動する変化が生成され、ECE 推定量を変更してもほとんど効果がありません。デフォルトの生成された回答、ベアコンテキストプロトコルの下では、指示設定は、言語化された信頼性に対する大きな調整ゲインを示すのではなく、同等に近いものになります。別の提供された回答分析では、表面的にもっともらしい誤った回答は、提供されたゴールドアンサーとほぼ同じ信頼度を得ており、言語化された信頼度は、正しさだけではなく、回答のもっともらしさと出所も反映していることを示唆しています。私たちは、両方の信頼シグナルはプロトコル依存の行動測定として扱われるべきであり、引き出しの来歴、採点された回答、トークン確率の読み出し、およびコンディショニングコンテキストをカバーするレポートチェックリストを提供する必要があると主張します。
原文 (English)
Same Answer, Different Confidence: Protocol Sensitivity in LLM Confidence Calibration
Is verbalized confidence better calibrated than token likelihood? The answer depends on how the token likelihood is measured: which answer is scored, and under which prompt. Published comparisons diverge on this, and in a twelve-study audit five never state the choice. We fix one prediction event per question, the model's own answer together with its correctness label, and score that same answer under a plain query and inside the confidence prompt, holding the answer and its label fixed. Across four QA datasets and three 7-8B Instruct models this changes which signal performs better, by point estimate, in 4 of 12 settings under ECE and 9 of 12 under AUROC. The AUROC result cannot come from rescaling the likelihoods, since AUROC is invariant to any common order-preserving transformation; the items are ordered differently. Two further choices behave the same way: substituting the reference string for the model's own answer, and reading the first answer token instead of the answer span. Crossing three answer slots, two scored strings, and two readouts gives twelve measured operational variants that leave the sign of the ECE comparison ndetermined in 6 of 12 settings, whereas alternative calibration estimators move it substantially less, although one changes a single prompted-context winner. Verbalized confidence is sensitive to answer formulation as well: replacing an accepted TriviaQA alias with the canonical reference raises confidence by $0.072$ although both answers are correct. Comparing the two signals therefore requires an explicit answer, context, and evaluation protocol.
Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation
While Vision-Language Models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this…
VESTA: LLM エージェント向けの完全に自動化されたシナリオ生成および安全性評価フレームワーク
大規模言語モデル (LLM) は、単純なテキストベースの対話システムから、メモリを維持し、ツールを使用し、外部環境にアクセスし、タスクを実行できる LLM エージェントへとますます進化しています。彼らの能力と自律性が拡大するにつれて、彼らが直面する安全リスクもより多様になります。既存の評価は、手動で作成されたシナリオ、静的なプロンプト、または最終出力の判断に依存していることが多く、タスクの実行中にエージェントが直面する可能性のあるさまざまなリスクを把握することが困難です。 LLM エージェント向けの完全に自動化されたシナリオ生成および安全性評価フレームワークである VESTA を紹介します。 VESTA は、5 つのリスク次元に基づいて、現実世界のタスク実行における抽象的で多様な安全リスクを 1,072 の測定可能な評価シナリオにインスタンス化します。自動評価パイプラインを使用して、12 個の LLM エージェントが 2 つの権限コンテキストの下で評価されます。その結果、現在のエージェントはタスク実行中に依然として重大な行動安全リスクに直面しており、平均 ASR は 47.1%、いくつかのモデルは 70% を超えていることが示されています。これらの調査結果は、LLM エージェントの安全性を理解し改善するために、実行可能なプロセスレベルの評価が重要であることを示しています。
原文 (English)
ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks. As their capabilities and autonomy expand, the safety risks they face also become more diverse. Existing evaluations often rely on manually written scenarios, static prompts, or final-output judgments, making it difficult to capture the diverse risks that agents may face during task execution. We introduce ForesightSafety-SAGE, a fully automated scenario generation and safety evaluation framework for LLM agents. Based on five risk dimensions,we instantiae abstract and diverse safety risks in real-world task execution into 1,072 measurable evaluation scenarios. Using the automated evaluation pipeline, 12 LLM agents are evaluated under two authority contexts. The results show that current agents still face substantial behavioral safety risks during task execution, with an average ASR of 47.1% and several models exceeding 70%. These findings demonstrate the importance of executable, process-level evaluation for understanding and improving LLM agent safety.
ARIADNE: 推論時アダプターの動的選択のための不可知論的なルーティング
パラメータ効率の良い微調整 (PEFT) の導入が増加することで、単一のバックボーンがタスクに特化した多数のアダプターとペアになるモデル エコシステムが誕生しました。この設定では、推論時のクエリがタスク ラベルなしで届くことが多く、システムは増大する異種アダプター プールから最も適切なアダプターを自動的に選択する必要があります。既存のルーティング方法は、重み分解や勾配ベースの統計などのアダプター内部へのアクセスに依存するか、追加のルーター トレーニングを必要とするため、新しいアダプターが追加されるとスケーラビリティと移植性が制限されます。推論時に動的アダプターを選択するための、トレーニング不要でアダプターに依存しないルーティング フレームワークである ARIADNE を紹介します。 ARIADNE は、トレーニング セットの埋め込みから計算された重心のセットを通じて各アダプターを表し、そのアダプターに関連付けられたデータ分布をキャプチャします。ラベルのない入力が与えられると、潜在空間内のこれらの重心への近接性を測定することによってアダプターを選択します。ルーティングは完全に入力埋め込み空間で実行されるため、ARIADNE は任意の PEFT メソッドと互換性があり、アダプターやトレーニング手順を変更する必要はありません。主に Llama 3.2 1B Instruct を使用して 23 の多様な NLP タスクを評価したところ、ARIADNE は上限パフォーマンスの 97.44% を回復しました。 44 タスクに拡張すると、追加のトレーニングやアダプター内部へのアクセスを必要とせずに、平均 89.7% の選択精度を達成します。
原文 (English)
Semantic Adapter Routing with Fine-Tuning Task Embeddings
Parameter-efficient fine-tuning (PEFT) has led to model ecosystems in which a single backbone is paired with many task-specialized adapters. Given such a library, routing aims to select the most appropriate adapter for a user query. While existing adapter routers typically require access to adapter weights or supervised training, we develop training-free semantic adapter routing methods using task embeddings. In ARIADNE, we reframe adapter selection as a classification problem, where PEFT adapters are represented by task embeddings and an unlabeled query is routed to the nearest adapter in the encoder's latent space. Evaluated on 23 tasks, ARIADNE recovers 97.4% of Oracle task performance and scales to 44 adapters at 89.7% selection accuracy, without touching a single adapter parameter. However, training data needed for ARIADNE may not be available when adapters come from public hubs or third-party providers. To overcome this limitation, we introduce GRACE, which recovers an adapter's fine-tuning data from its output logits alone via a modified contrastive decoding diffing (CDD) procedure. Synthetic data generated from CDD-UM is then used to construct task embeddings. Across three backbones (Llama-3.2-1B, Qwen2.5-3B, Qwen2.5-32B), GRACE recovers 72--100\% of Oracle task accuracy and matches or exceeds ARROW on 48 of 69 task/backbone combinations, while requiring neither training data nor model weights. Overall, we demonstrate that fine-tuning task embeddings provide an accurate and efficient path to semantic adapter routing.
SenWorld: コンテキストリッチな評価データを生成するためのデジタル ツイン シミュレーション
スマートフォンのパーソナル アシスタントは長期にわたる個人データを推論しますが、その評価には正解がわかっているコンテキストに富んだ評価データが必要であり、実際のデバイスのトレースはプライバシーに敏感すぎて共有できません。この課題に対処するために、構築によって固定されたグラウンド トゥルースを使用してそのようなデータを生成する、物理的に接地され、決定論的でイベント ソースのデジタル ツイン シミュレーションである SenWorld を紹介します。 SenWorld では、ペルソナは実際の地図、天気、休日、ネットワーク データから構築された世界で 1 日を過ごします。観測可能なすべての信号はシステム全体のスナップショットにアーカイブされます。また、各評価ケースは、事後注釈や大規模言語モデル (LLM) ジャッジではなく、既存のレコードへのポインターによってラベル付けされます。この手法を北京の 16 人のペルソナで評価しました。生成されたデータは、カテゴリ分布 (ジェンセンとシャノンの相違 (JSD) 0.070) および通信記録の 1 日のリズム (JSD 0.1 未満) において、保持されている実際のユーザーのベンチマークと厳密に一致していますが、生成された記録は実際の記録よりも短いままです。スクリプトによる対話がなければ、ペルソナは完全に往復する対話サブグラフと差別化された行動レパートリーを形成します。 717 件の評価ケースに投影された生成データでは、実稼働スマートフォン アシスタントの 78 件の障害が明らかになり、通話とショート メッセージ サービス (SMS) の記録に集中し、連絡先、スケジュール、アラームは決して失敗しませんでした。スナップショット ポインタは、LLM 判定者が関与せずに、各失敗をアシスタント側の取得エラーとして確認します。全体として、SenWorld は、ラベルが構築によって固定されている評価データへの、プライバシーに安全で再現可能で配布がチェックされたパスを提供します。
原文 (English)
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, we present SenWorld, a physically grounded, deterministic, event-sourced digital-twin simulation that generates such data with ground truth fixed by construction. In SenWorld, personas live through a full day in a world built from real map, weather, holiday, and network data; every observable signal is archived in full-system snapshots; and each evaluation case is labeled by a pointer to an existing record rather than by post-hoc annotation or a large language model (LLM) judge. We evaluate this method with 16 personas in Beijing. The generated data closely matches the held-out real-user benchmark in category distribution (Jensen--Shannon divergence (JSD) 0.070) and in the daily rhythm of communication records (JSD below 0.1), though generated records remain shorter than real ones. Without scripted interaction, personas form a fully reciprocated dialogue subgraph and differentiated behavioral repertoires. Projected into 717 evaluation cases, the generated data exposes 78 failures in a production smartphone assistant, concentrating on call and Short Message Service (SMS) records while contacts, schedules, and alarms never fail. The snapshot pointer confirms each failure as an assistant-side retrieval error, with no LLM judge involved. Overall, SenWorld offers a privacy-safe, reproducible, and distribution-checked path to evaluation data whose labels are fixed by construction.
OpenForgeRL: Train Harness-native Agents in Any Environment
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, an…
マルコフ意思決定プロセスのためのプロパティ駆動型因果抽象化
マルコフ意思決定プロセス (MDP) は意思決定モデルとして広く使用されており、通常は状態変数とその評価を通じて因数分解された状態空間に対して指定されます。州の数が指数関数的に増加するため、MDP の多くの推論タスクが困難になります。抽象化は、MDP を削減し、スケーラビリティの問題を軽減する有望な手法です。この研究では、因数分解された MDP の因果関係の概念と、元の MDP モデルの多くの特徴を保持する新しいプロパティ駆動型の因果抽象化手法を導入します。このため、状態変数述語の因果関係に依存し、特定の抽象化プロパティを満たすか違反する同じ理由を共有する状態を識別します。私たちは、MDP、間隔 MDP、確率的ゲームなどのさまざまなモデル タイプを使用して、さまざまな因果的 MDP 抽象化を理論的および経験的に比較します。私たちの評価は、私たちのアプローチの可能性を示しています。いくつかの標準ベンチマークについて、元の MDP に対して最適に近いポリシーを計算できる小さな抽象化を取得しました。さらに、私たちの因果的抽象化は、多くの場合、関連する大規模な MDP モデルに一般化されます。
原文 (English)
Property-driven Causal Abstractions for Markov Decision Processes
Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in MDPs challenging. Abstractions are promising techniques to reduce MDPs and thus mitigate scalability issues. In this work, we introduce a notion of causality on factored MDPs and a novel property-driven causal abstraction technique that retains many characteristics of the original MDP model. For this, we rely on causal relations over state variable predicates and identify those states that share the same reasons for fulfilling or violating a given abstraction property. We theoretically and empirically compare various causal MDP abstractions using different model types such as MDPs, interval MDPs, or stochastic games. Our evaluation demonstrates the potential of our approach: For several standard benchmarks, we obtain small abstractions that allow us to compute near-optimal policies for the original MDP. Furthermore, our causal abstractions often generalize to related large-scale MDP models.
AI エージェントはオープンエンドの AI 研究を行うことができますか? 2 つのケーススタディからの初期の証拠
AI の爆発的な進歩の予測は、AI 研究を自動化する AI エージェントにかかっています。しかし、エージェントが無制限の AI 研究を実行できるかどうかについての証拠は乏しい。現在の評価では、限定的で検証可能なタスクでエージェントをテストするか、無制限の研究を除外するか、AI で生成された論文をブラインド査読に提出しますが、これは過剰で確率的であり、レビューの質が低いという問題があります。 AI 研究開発の自動化に向けた進捗状況を測定する 3 番目の方法を紹介します。エージェントは、質の高い未発表論文の中心となる自由回答の研究課題に取り組み、論文の元の著者がその成果を採点します。これらをシャドウ評価と呼びます。私たちは 2 つの未公開の NeurIPS 2026 提出物に対してシャドウ評価を実行し、フロンティア エージェントに 6 日間と数千ドルのコンピューティングを与えました。エージェントは人間の助けなしですべてのエンジニアリングを完了しましたが、研究上の疑問の答えに向けて実質的な進歩を遂げることはできませんでした。その結果、両方の論文は著者によって明確に拒否されました。私たちは、繰り返される 5 つの失敗モードを特定します。それは、出版可能な研究の基準に関する誤った判断、研究設計の欠点に対する創造性のない対応、行き止まりからの非効果的な後戻り、不十分なリソース認識、指示の逸脱です。 2 番目のモデルと足場を使用した堅牢性チェックにより、これらの障害が再現されました。専門家のレビュー、アンケートの回答、エージェントのリポジトリ、およびログを公開します。私たちの結果は、今日のエージェントが AI 研究のエンジニアリングを行うことができるものの、研究ライフサイクルの重要な部分で苦労しているという初期の証拠を提供します。
原文 (English)
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.
H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases
Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constr…
Homebot: A Personal AI Agent for Conversational Home Assistance and Automation
\texttt{Homebot} is a locally deployable AI agent for conversational household assistance and automation. It accepts voice and instant-mess…
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation throug…
重み空間アブレーション下の層間相互作用: 閉じた形式のアテンション ヤコビアン境界と実際の事前学習モデルでのテスト
関連論文では、条件付き計算が残差ストリームを通じて加算的に実行される理想化されたモデル内で、アクティベーション パッチとウェイト スペース アブレーションがいつ一致するかを研究しています。 2 つのキャリアがアーキテクチャ的に依存するそのモデル内の 1 つの構成、つまりアテンション ヘッドとその独自の層の正規化 MLP 構成については、正確な 1 次相互作用公式、MLP のみがアブレーションされる場合はゼロ、ヘッドもアブレーションされる場合は 2 次境界が導出されます。その結果は単一の残差ブロックに限定され、合成タスクの小さな変換器でのみチェックされます。この論文では、両方の限界を超えて結果を拡張します。まず、複数の層にまたがるアブレーションキャリアからの相互作用は、タッチされた層ごとに 1 つと、分解が小さいことを主張しない層間の残りの、同じブロックの項に正確に分解されます。次に、混合二次導関数の二重積分として 2 つの層の残りを正確に分離し、それを結合するために必要な欠落成分に名前を付けます。これは、注目サブブロックに結合するヤコビアンです。この境界を閉じた形式で導出し、Qwen2.5-1.5B-Instruct の実際の重みに対して違反が 1 つも発生しないことを検証しますが、まだレイヤー間で連鎖させていません。また、閉じた形式で、コンパニオンペーパーの綴じたままの曲率定数を与えます。第三に、同じモデル上で、このタスクの元の活性化パッチ手法を使用して、決して設計されていない間接オブジェクト識別のための創発回路を検索して見つけ、それに対する崩壊、解離、および相互作用をテストします。結果はまちまちです。共有キャリアはテストされた 5 つのインスタンスすべてで出現し、崩壊と解離はすべてではありませんがほとんどで保持され、コンパニオン定理がカバーする同一ブロックのケースの外側の層ペアで、5 つのうち 3 つで非ゼロ相互作用が測定可能です。
原文 (English)
Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model
A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream. For the one composition in that model where two carriers are architecturally dependent, an attention head and its own layer's normalization-MLP composition, it derives an exact first-order interaction formula, zero when only the MLP is ablated and second-order bounded when the head is also ablated. That result is confined to a single residual block and checked only on small transformers on a synthetic task. This paper extends the result past both limits. First, the interaction from ablating carriers spanning several layers decomposes exactly into same-block terms, one per touched layer, plus a cross-layer remainder on which the decomposition makes no claim of smallness. Second, we isolate that remainder exactly, for two layers, as a double integral of a mixed second derivative, and name the missing ingredient needed to bound it: a Jacobian bound for the attention sub-block. We derive this bound in closed form and verify it, without a single violation, against Qwen2.5-1.5B-Instruct's real weights, though we do not yet chain it across layers. We also give, in closed form, the curvature constant the companion paper's bound leaves unexhibited. Third, on that same model, we search for and find an emergent circuit for indirect object identification, never designed into it, using the original activation-patching method for this task, and test collapse, dissociation, and interaction on it. The result is mixed: a shared carrier emerges across all five tested instances, collapse and dissociation hold on most but not all, and a nonzero interaction is measurable on three of five, at layer pairs outside the same-block case the companion theorem covers.
SkillTrace: LLM-Agent スキル再利用のためのマルチトレース来歴監査
LLM エージェント エコシステムは、メタデータ、自然言語命令、コード、ツール、リファレンス、運用ワークフローの混合形式パッケージなど、再利用可能なスキルを中心に急速に成長しています。スキルが市場の成果物になるにつれて、その再利用の監査は、通常のコード クローンの検出と同じ問題ではなくなりました。既存の検出器は、単一モダリティのソース コードまたはパッケージ全体の類似性をターゲットにしていますが、スキル再利用の証拠は、作成されたテキスト、実装フラグメント、および運用構造全体に分散されています。その結果、スキルの一部だけを保存する再利用が失われる可能性があります。 LLM エージェントのスキルを再利用するためのマルチトレース来歴監査フレームワークである SKILLTRACE を紹介します。 SKILLTRACE は、Expression、Implementation、Operational の 3 つの来歴トレースを抽出します。これは、アクティブ化、手順、およびリソース フロー構造をキャプチャするスキル操作グラフ (SOG) として操作トレースを表します。 LLM は、取り込み時に 1 回だけ、操作トレースの抽出のみを支援します。監査時に、SKILLTRACE はキャッシュされたトレースを決定論的に比較し、同じ機能の厳密な否定に対して各トレースを調整し、どのトレースが再利用の決定をサポートしているかを報告します。 SKILLTRACE-BENCH では、100 のマーケットプレイス アンカーおよび 751 のネガティブ コントロールを超える 820 の変換された再利用ポジティブを使用して、SKILLTRACE は AUROC 0.938 および F1 0.898 を達成しました。さらに、36,446 のスキルを対象としたワイルド監査では、トレースに起因する証拠により、リポジトリ レベルのベースラインを超えた実用的な再利用レビュー キューが明らかになったことが示されています。
原文 (English)
SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse
LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditing their reuse is no longer the same problem as ordinary code clone detection. Existing detectors target single-modality source code or whole-package similarity, yet skill reuse evidence is distributed across authored text, implementation fragments, and operational structure. As a result, they can miss reuse that preserves only one part of a skill. We present SKILLTRACE, a multi-trace provenance auditing framework for LLM-agent skill reuse. SKILLTRACE extracts three provenance traces: Expression, Implementation, and Operational. It represents the Operational Trace as a Skill Operational Graph (SOG) that captures activation, procedure, and resource-flow structure. An LLM assists only the Operational-trace extraction, once at ingestion; at audit time SKILLTRACE compares cached traces deterministically, calibrates each trace against same-function strict negatives, and reports which trace supports a reuse decision. On SKILLTRACE-BENCH, with 820 transformed reuse positives over 100 marketplace anchors and 751 negative controls, SKILLTRACE achieves AUROC 0.938 and F1 0.898. A 36,446-skill wild audit further shows that trace-attributed evidence surfaces actionable reuse review queues beyond repository-level baselines.
長期にわたるターミナルタスクのための再帰的合成
ターミナル エージェント向けの高品質で長期的なトレーニング データは作成に費用がかかり、タスクごとに数百ドルから数千ドルかかることがよくあります。これは、各タスクが命令、環境、参照ソリューション、検証器の相互一貫性を保つ必要があるためです。人間によるオーサリングは拡張性がなく、大規模言語モデル (LLM) を使用した直接生成では、これらの依存関係が壊れることがよくあります。我々は、長期にわたるターミナル エージェント タスクを大規模に構築するための再帰的検証済み合成フレームワークである再帰的合成ターミナル タスク (RST) を紹介します。 RST は、検証されたシード タスクから開始して、参照ソリューションを拡張し、検証ツールと命令を新しいワークフローに再調整し、新しいサンドボックスで結果を検証し、受け入れられたタスクを後続のラウンドのシードとして再利用します。 15 回の再帰ラウンドにわたって、RST は 37,484 個の合成ターミナル エージェント タスクをタスクあたり約 $0.05 で生成します。タスクの難易度はラウンドを重ねるごとに大幅に増加します。リファレンス ソリューションの中央値は 67 行から 374 行に増加し、実行されたコマンド数の中央値は 40 から 244 に増加し、DeepSeek-V4-Pro pass@4 は $R_1$ の 90\% から $R_{15}$ の 2.5\% に低下します。トレーニングの有用性を実証するために、合成されたタスクに関して拒否サンプリングされた Qwen3.5 軌跡を収集し、それらを教師付き微調整に使用します。これらの軌道を微調整すると、Qwen3.5-27B と Qwen3.5-122B-A10B がターミナル ベンチ ~ 2、ターミナル ベンチ ハード、およびロングホライズン ターミナル ベンチで最大 10 ポイント改善され、エージェント PPO により Qwen3.5-27B が 3 つで 49.44\%、32.00\%、22.07\% に上昇しました。ベンチマークは、ベース モデルに対して 20.0\%、41.2\%、および 21.9\% の相対的な向上に相当します。さらに、15 ラウンド後の再帰には上限がありません。難易度が上昇し続けても合成収率と検証率は安定しており、ここで報告した規模をはるかに超えてプロセスを継続できることを示しています。
原文 (English)
Recursive Synthesis for Long-Horizon Terminal Tasks
High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. Human authoring does not scale, and direct generation with large language models (LLMs) often breaks these dependencies. We present Recursive Synthetic Terminal Tasks (RST), a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks at scale. Starting from verified seed tasks, RST extends the reference solution, realigns the verifier and instruction to the new workflow, validates the result in a fresh sandbox, and reuses accepted tasks as seeds for subsequent rounds. Across fifteen recursive rounds, RST produces 37,484 synthesized terminal-agent tasks at roughly $0.05 per task. Task difficulty increases substantially over rounds: the median reference solution grows from 67 to 374 lines, the median number of executed commands grows from 40 to 244, and DeepSeek-V4-Pro pass@4 drops from 90% at R1 to 2.5% at R15. To demonstrate training utility, we collect rejection-sampled Qwen3.5 trajectories on the synthesized tasks and use them for supervised fine-tuning. Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench 2, Terminal-Bench Hard, and Long-Horizon Terminal Bench, while agentic PPO lifts Qwen3.5-27B to 49.44%, 32.00%, and 22.07% on the three benchmarks, corresponding to relative gains of 20.0%, 41.2%, and 21.9% over the base model. Moreover, after 15 rounds, the recursion shows no ceiling: synthesis yield and validation rates remain stable as difficulty keeps climbing, indicating that the process can continue well beyond the scale reported here.
CourseGraph: 大学間のコンピューター サイエンス コースの重複と相違点を見つける
Erasmus+ などの学生流動プログラムにより、学生は他の大学でコースを受講でき、学問的および文化的視野を広げることができます。ただし、この柔軟性は実際的な課題にもつながります。それは、学生が家庭のカリキュラムのコースと実質的に重複するコースを他の場所で受講しないようにすることです。この研究では、学位プログラムに組み込むコースを評価する際にカリキュラム管理者が従うプロセスから得られた洞察に基づいて、外部コースの評価を自動化する方法論である CourseGraph を提案します。 Course-Graph は、コースの Web ページからコースのタイトル、説明、学習成果などの情報を抽出します。次に、この情報は BERT ベースの言語モデルを使用して意味論的に表現され、その後、コース間のペアごとの類似性が計算されます。この情報は、ランダム フォレスト分類器によって使用され、海外の候補コースが学生のカリキュラムに既に含まれているコースと重複するかどうかが判断されます。私たちは、(1) アイントホーフェン工科大学のコンピューター サイエンス プログラム、実質的に重複するコースに関する情報が含まれている、(2) ルンド大学のコンピューター サイエンス プログラムに登録している学生による 6 つの承認された国際プログラム (カリキュラム管理者による対応する決定を含む) を使用して、CourseGraph を評価します。実験結果は、CourseGraph が重複するコースを特定し、大学間でのカリキュラムの調整をサポートするための効果的なアプローチを提供することを示しています。
原文 (English)
CourseGraph: Finding overlaps and differences in Computer Science courses across universities
Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural horizons. However, this flexibility also leads to a practical challenge: ensuring that students do not take courses elsewhere that substantially overlap with courses in their home curriculum. In this work, we propose CourseGraph, a methodology that automates the evaluation of external courses based on insights obtained from the process followed by curriculum administrators when assessing courses for inclusion in a degree program. Course- Graph extracts information such as course titles, descriptions, and learning outcomes from the course webpage. Then, this information is represented semantically using a BERT-based language model, after which the pair-wise similarity between courses can be computed. This information is then used by a Random Forest classifier to determine whether a candidate course abroad overlaps with a course already contained in the student's curriculum. We evaluate CourseGraph using (1) the Computer Science program at Eindhoven University of Technology, which contains information about courses with substantial overlap, and (2) six approved international programs from students enrolled in the Computer Science program at Lund University, including the corresponding decisions made by a curriculum administrator. The experimental results indicate that CourseGraph provides an effective approach for identifying overlapping courses and supporting curriculum alignment across universities.
検索エージェント向けのコンテキスト情報ポリシーの最適化
検索エージェントは、複数ステップの推論中に外部証拠を取得して使用できるようにすることで、静的パラメトリック メモリを超えて大規模な言語モデルを拡張します。複雑な情報や進化する情報を伴う知識集約型タスクの場合、その信頼性は、関連する証拠を取得するだけでなく、それをその後の推論の指針として使用することにも依存します。ただし、既存の方法では、検索後のアクションが検索された証拠に基づいているかどうかを直接評価することなく、主に最終的な回答の正確性または中間の進歩に報酬を与えます。この不整合により、事前主導型の推論が促進されます。エージェントは内部知識に基づいて結論を出し、主にそれを確認するために検索を使用するため、確証バイアスと非効率的な証拠の使用が生じます。この問題に対処するために、ポリシーの最適化と外部の証拠の使用を明示的に調整する証拠指向の強化学習フレームワークであるコンテキスト情報ポリシー最適化 (CIPO) を提案します。 CIPO は、取得した情報の影響を受ける推論アクションに密なターンレベルのクレジットを割り当てますが、この証拠使用シグナルと、回答の正しさを維持するためのグローバルな結果報酬を組み合わせます。この方法により、CIPO は証拠から切り離された推測を阻止し、取得した事実がその後の推論を導き、修正できる推論の軌道を促進します。重要なのは、CIPO では人間のプロセス アノテーションも追加の報酬モデルも必要ないことです。 7 つのドメイン内およびドメイン外のベンチマークに関する広範な実験により、CIPO が事前駆動推論の蔓延を減らし、ほとんどのタスクで優れたパフォーマンスを達成することが示されています。
原文 (English)
Contextual Information Policy Optimization for Search Agents
Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during multi-step reasoning. For knowledge-intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant evidence but also on using it to guide subsequent reasoning. However, existing methods primarily reward final-answer correctness or intermediate progress, without directly assessing whether post-retrieval actions are grounded in the retrieved evidence. This misalignment encourages prior-driven reasoning: agents form conclusions based on internal knowledge and use retrieval mainly to confirm them, resulting in confirmation bias and inefficient evidence use. To address this issue, we propose Contextual Information Policy Optimization (CIPO), an evidence-oriented reinforcement learning framework that explicitly aligns policy optimization with external evidence use. CIPO assigns dense, turn-level credit to reasoning actions influenced by retrieved information, while combining this evidence-use signal with a global outcome reward to preserve answer correctness. With this manner, CIPO discourages evidence-detached guesses and promotes reasoning trajectories in which retrieved facts can guide or revise subsequent reasoning. Importantly, CIPO requires neither human process annotations nor an additional reward model. Extensive experiments on seven in-domain and out-of-domain benchmarks show that CIPO reduces the prevalence of prior-driven reasoning and achieves excellent performance on most tasks.
DASH: 推論モデルのオンポリシー自己蒸留のための発散適応型監視の視野
検証可能な報酬を伴う強化学習 (RLVR) は、自動的に検証可能な結果信号を使用して大規模な言語モデルの推論能力を向上させますが、これらの信号は通常、まばらであり、シーケンス レベルです。オンポリシー自己蒸留 (OPSD) は、学生が訪問するプレフィックスで特権教師にクエリを実行し、高密度のトークンレベルの分布監視を提供することで、この希薄性を軽減します。この高密度監視により信号のスパース性が軽減されますが、標準の OPSD ではロールアウトの時間構造がまだ十分に活用されていないことがわかります。これは、その位置や発生する発散シーケンスに関係なく、すべての局所的な発散に同じ係数を割り当てます。ポリシーに基づく自己回帰生成では、同じ乖離の大きさでも、教師と生徒の間の不一致のさまざまな展開を反映して、さまざまな不一致履歴をたどることができます。ローカル スカラーだけではこれらの時間コンテキストを区別できないため、標準の OPSD はトークン レベルの重みを実現された不一致シーケンスに適応させることができません。この制限に対処するために、私たちは Divergence-Adaptive Supervision Horizons (DASH) を提案します。 DASH は、各局所蒸留信号とシーケンス レベルの平均の間のギャップを適応伝播ゲートにマッピングし、これらのゲートを使用して逆方向マルチステップ集約を制御します。そうすることで、DASH は、生成中にローカルな相違がどのように変化するかに応じて、トークン レベルの監視の重みを調整します。 3 つのモデル スケールにわたる 3 つの数学的推論ベンチマークの実験では、3 つのスケールすべてのすべてのベンチマークで、一致するバニラ OPSD 再実行よりも DASH が向上していることが示されています。 DASH は、OPSD がすでに計算している教師と生徒の分布を再利用するため、ゲインを得るために追加の教師または生徒の前方パスは必要ありません。コード: https://github.com/DBtxy/DASH-OPSD
原文 (English)
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional supervision. Although this dense supervision alleviates signal sparsity, we find that standard OPSD still underexploits the temporal structure of the rollout. It assigns every local divergence the same coefficient, regardless of its position or the divergence sequence in which it occurs. In on-policy autoregressive generation, the same divergence magnitude can follow different discrepancy histories, reflecting different evolutions of the mismatch between the teacher and student. Since the local scalar alone cannot distinguish these temporal contexts, standard OPSD cannot adapt its token-level weights to the realized discrepancy sequence. To address this limitation, we propose Divergence-Adaptive Supervision Horizons (DASH). DASH maps the gap between each local distillation signal and the sequence-level mean to an adaptive propagation gate and then uses these gates to control backward multi-step aggregation. By doing so, DASH adjusts token-level supervision weights according to how local divergences evolve during generation. Experiments on three mathematical reasoning benchmarks across three model scales show that DASH improves over our matched vanilla OPSD reruns on every benchmark at all three scales. DASH reuses the teacher and student distributions that OPSD already computes, so the gains require no additional teacher or student forward pass. Code: https://github.com/DBtxy/DASH-OPSD
Top-K を超えて: ブラックボックスの取得を解釈可能なエージェント操作に置き換える
長い文書に対する検索拡張生成は、テキストをチャンク化し、チャンクを埋め込み、クエリの上位 k 個の最近傍を表面化するという 1 つの設計によって支配されています。私たちは、財務諸表、監査報告書、規制当局への報告書などの重要な種類の文書にとって、この設計は構造的に不健全であると主張し、その議論を測定可能にします。 780 ページの政府財務報告書では、コンテンツ行の 86.8% が表の行であり、何千ものほぼ同一の数値が 1 つの埋め込みスペースで競合し、数値は中央値 13 行上のヘッダーから単位を継承します。そのため、チャンク境界によって数値が 10 万単位であるか 100 万単位であるかは日常的に区別されており、誤差は 2 桁の誤差になります。 Steelman として構築されたテーブル対応チャンカーは単位の問題を修正しますが、試したすべてのチャンク サイズで会計年度ヘッダーのない数値チャンクの 27 ~ 30% が残ります。私たちは、READ (Reliable Embedding-free Agentic Document-search) を提案します。これは、モデル コンテキスト プロトコル上で公開される 3 つの決定論的操作 (正規化された字句検索、構造ナビゲーション、および限定されたスパン読み取り) を通じてエージェントが生の文書を読み取るため、軌跡は不透明な類似性スコアではなく、再生可能な監査証跡となります。 51 の検証された質問について、READ の回答率は 58.8% で、密検索の 15.7% (p_Holm = 2 x 10^-5) -- または調整済みの 35.3% で、依然として READ が 23.5 ポイント (p_Holm = 0.017) リードしています。エージェントに同じループを与えても、top-k ツールでは 27.5% しか到達せず、反復ではなくインターフェイスにゲインが見出されます。また、証拠がサポートしていないことも報告します。BM25 は統計的に READ と区別できないため、結果は埋め込みベースの検索と埋め込みなしの検索を区別し、エージェント検索と語彙検索を区別しません。
原文 (English)
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations
Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statements, audit reports, regulatory returns -- this design is structurally unsound, and we make the argument measurable. On a 780-page government financial report, 86.8% of content lines are table rows, thousands of near-identical figures compete in one embedding space, and a figure inherits its unit from a header a median of 13 lines above it -- so a chunk boundary routinely separates a number from whether it is in lakh or crore, an error of two orders of magnitude. A table-aware chunker built as a steelman fixes the unit problem but leaves 27-30% of numeric chunks with no fiscal-year header at every chunk size we tried. We propose READ (Reliable Embedding-free Agentic Document-search), in which an agent reads the raw document through three deterministic operations -- normalized lexical search, structural navigation, and bounded span reads -- exposed over the Model Context Protocol, so a trajectory is a replayable audit trail, not an opaque similarity score. On 51 verified questions READ answers 58.8% against dense retrieval's 15.7% (p_Holm = 2 x 10^-5) -- or 35.3% tuned, which READ still leads by 23.5 points (p_Holm = 0.017). An agent given the same loop but a top-k tool reaches only 27.5%, locating the gain in the interface rather than in iteration. We also report what the evidence does not support: BM25 is statistically indistinguishable from READ, so our result separates embedding-based from embedding-free retrieval, not agentic from lexical search.
Towards a Theoretical Understanding of Two Tower Recommendation Models
Production-grade recommender systems rely heavily on a large-scale corpus used by online media services, including Netflix, Pinterest, and…
Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction
Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic indicators of regions such as population and comm…
Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription
Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractica…
A primer on optimal transport for causal inference with observational data
The theory of optimal transportation has developed into a powerful and elegant framework for comparing probability distributions, with wide…
PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks
Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexpl…
Minimal Ingredients for Reward Assignment from Expert Demonstrations
Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning. A common and intuitive strat…
Learning to Walk With Less: A Dyna-Style Approach to Quadrupedal Locomotion
Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from low data efficiency, requiring m…
Evaluating Useful Surrogate Models for Configuration Tuning Beyond Accuracy: A Fitness Landscape Analysis Perspective
To efficiently tune configuration for better software system performance (e.g., latency) at the deployment and maintenance stage, many tune…
Provable Training Data Identification for Large Language Models
Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. Howev…
Stability of Transformers under Layer Normalization
Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stab…
In Situ Training of Implicit Neural Compressors for Scientific Simulations via Sketch-Based Regularization
Focusing on implicit neural representations, we present a novel in situ training protocol that employs limited memory buffers of full and s…
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strain…
MetaSICL: メタ音声インコンテキスト学習による Audiroty LLM の適応
聴覚大規模言語モデル (LLM) は、幅広い音声理解タスクにわたって強力なパフォーマンスを実証しています。それにもかかわらず、リソースが少ないタスクに適用すると、苦労することがよくあります。ドメイン内のラベル付きデータが不足しているか、実際のテスト分布と一致しない場合、直接の微調整は脆弱になる可能性があります。 In-Context Learning (ICL) は、いくつかのドメイン内デモンストレーションを条件付けして聴覚 LLM を適応させることにより、トレーニング不要の推論時間ソリューションを提供します。この研究では、$\textit{Vanilla ICL}$ が、選択されたモデルのさまざまな音声タスクおよびオーディオ タスクにわたってゼロショット パフォーマンスを向上させることを最初に示します。これは、この ICL 適応機能がマルチモーダル設定に一般化できることを示唆しています。これに基づいて、$\textbf{Meta Speech In-Context Learning (MetaSICL)}$ を提案します。これは、モデルのインコンテキスト学習能力を強化することを目的とした、さまざまなタスクからの高リソース音声データのみを利用するトレーニング後のレシピです。実験によれば、私たちが提案した方法は、リソースが少ないシナリオでは直接の微調整よりも優れています。
原文 (English)
MetaSICL: Globalizing Auditory LLMs for Underserved Speakers and Languages via Meta Speech In-Context Learning
Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data. Globalizing such systems requires handling low-resource settings, where the target speakers, languages, or tasks are poorly represented in training data. In these regimes, collecting enough labeled in-domain data is often impractical, and the small corpora available may still under-represent the test distribution, making direct fine-tuning brittle under domain shift. In-Context Learning (ICL) offers an alternative: instead of updating model parameters for every underserved community, an auditory LLM can adapt at inference time by conditioning on a few local demonstrations. However, vanilla speech ICL remains limited because most auditory LLMs are not explicitly trained to use such demonstrations effectively. We address this gap with Meta Speech In-Context Learning (MetaSICL), a post-training recipe that strengthens an auditory LLM's in-context adaptation ability using only abundant high-resource speech data. Although MetaSICL never trains on the target low-resource domains, it improves performance across two backbones on children's ASR, audio understanding/reasoning, and speech translation and ASR in directions and languages unseen in post-training. We further study the case where some in-domain data is available, using low-resource language ASR as a case study, since recognition for underserved languages is central to globalizing generative AI. Here, using MetaSICL as a warmup for in-domain reinforcement learning yields the strongest results, outperforming direct fine-tuning across five typologically diverse languages. Overall, MetaSICL offers a practical route toward globalizing auditory LLMs by building inference-time adaptation into the model.
Kimi K2.5: Visual Agentic Intelligence
We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint…
Optimizing Spectral Prediction in MXene-Based Metasurfaces Through Multi-Channel Spectral Refinement and Savitzky-Golay Smoothing
The prediction of electromagnetic spectra for MXene-based solar absorbers, where MXenes are a family of two-dimensional transition metal ca…
SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization
Search-Augmented Generative Engines (SAGE) have emerged as a new paradigm for information access, bridging web-scale retrieval with generat…
MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms
Multi-attribution learning (MAL), which enhances model performance by learning from conversion labels yielded by multiple attribution mecha…
Deterministic Preprocessing and Interpretable Fuzzy Banding for Cost-per-Student Reporting from Extracted Records
Administrative extracts are often exchanged as spreadsheets and may be read as reports in their own right during budgeting, workload review…
Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving
The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging thei…
Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains
We argue that the current practice of evaluating AI/ML time-series forecasting models, predominantly on benchmarks characterized by strong,…
Improving Attributed Long-form Question Answering with Intent Awareness
Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these model…
CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment
Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions.…
Cluster Attention for Graph Machine Learning
Message Passing Neural Networks have recently become the most popular approach to graph machine learning tasks; however, their receptive fi…
GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking
Audio Large Language Models (ALLMs) enable spoken interaction but introduce new jailbreak vulnerabilities. Existing perturbation-based jail…
From Plan to Action: How Well Do Agents Follow the Plan?
Agents are commonly instructed to follow a task-specific plan for guidance. However, it is unknown to what extent agents actually follow in…
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference servi…
SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages
Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often prod…
Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architec…
Playing Games with My Heart: An Evaluation of AI Companion Apps
The use of chatbots for various forms of companionship is growing rapidly, raising a myriad of questions about simulated relationships, emo…
On Seeding Watermarks to Detect Verbatim LLM Copy-Paste Responses
Large language models (LLMs) have made fluent essay writing, code drafting, and quiz answering instantly available to students at every lev…
PULSE: Agentic Investigation with Passive Sensing for Proactive Affective Intervention in Cancer Survivorship
Cancer survivors face elevated rates of depression, anxiety, and emotional distress, yet self-report may be unavailable at some moments whe…
マルチリーガルベンチ: 管轄区域、言語、法的伝統を超えた法的推論に関する LLM の評価
法的 NLP ベンチマークは、圧倒的に単一言語を評価するか、法域間で根本的に異なるタスクを集約するため、言語を越えた比較は不可能です。私たちは、6 か国 (ウクライナ、フランス、オランダ、ポーランド、チェコ共和国、リトアニア)、4 つの言語族、および 1 億 3,400 万件の裁判所判決にまたがる同一のタスクを評価する初の法域を越えた法的ベンチマークである Multi-Legal-Bench を紹介します。このベンチマークは、裁判所タイプの分類、判決書の分類、事件結果の予測、法的規範の抽出、および国内裁判所登録簿からの構造化メタデータにマッピングされた原因カテゴリの予測という 5 つのタスクを定義し、意図的にまばらな 5x6 タスク管轄マトリックス (30 セル中 20 セルが埋まる) を形成します。 AWS Bedrock を介したゼロショットおよび 3 ショット プロンプトの下で 7 つのフロンティア LLM を評価し、スケーリング分析用に 4 つの追加の中小規模モデル (3-12B) を使用します。私たちの結果は次のことを明らかにしました: (1) ウクライナで発見されたタスク依存の少数ショット効果は、すべての管轄区域にわたって再現されます。 (2) タスクと管轄区域の両方で、言語ランキングの変化を支配する単一のモデルはありません。 (3) 言語間の少数ショット転送は言語の近接性に従わない: UA->FR (ロマンス、-2.1 pp) の転送は UA->PL (スラブ語、-13.7 pp) よりも優れており、ラベルセットの調整により言語族よりも転送の品質が予測されます。 (4) トークナイザーの充実度は、スプレッドが 2.3 倍であるにもかかわらず、異言語間の精度を有意に予測しません (r=-0.27、p=0.14)。これは、モデル アーキテクチャと事前トレーニング データがトークナイザーの効率を支配していることを示唆しています。すべてのデータ、プロンプト、モデル予測を公開します。
原文 (English)
Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions
Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible. We introduce Multi-Legal-Bench, the first cross-jurisdictional legal benchmark that evaluates identical tasks across six countries (Ukraine, France, Netherlands, Poland, Czech Republic, Lithuania), four language families, and 165 million full-text court decisions. The benchmark defines five tasks (court-type classification, judgment form classification, case-outcome prediction, legal norm extraction, and cause category prediction) mapped to structured metadata from national court registries, forming a deliberately sparse 5x6 task-jurisdiction matrix (20 of 30 cells filled). We evaluate 7 frontier LLMs under zero-shot and 3-shot prompting via AWS Bedrock, with 4 additional small/medium models (3-12B) for scaling analysis. Our results reveal that: (1) few-shot gains are uneven and track how much headroom a cell leaves rather than its language, with 8 of 28 judgment-form model-jurisdiction pairs losing accuracy; (2) no single model dominates any language, rankings shift with both task and jurisdiction; (3) cross-lingual few-shot transfer does not follow language proximity: UA->FR (Romance, -2.0 pp) transfers better than UA->PL (Slavic, -13.8 pp), with label-set alignment predicting transfer quality better than language family; and (4) tokenizer fertility, despite a 2.3x spread, does not significantly predict cross-lingual accuracy (r=-0.14, p=0.24), suggesting that model architecture and pretraining data dominate tokenizer efficiency. We release all data, prompts, and model predictions.
Rethinking Evaluation Paradigms in IBP-based Certified Training
Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations. Neura…
Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path
Understanding memorization in generative models remains challenging, with implications for copyright and privacy. Beyond verbatim reproduct…
The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products
Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same chara…
An Empirical Study of openPangu Quantization on Ascend NPUs
openPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive pos…
Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning
State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. In studying such systems, a wide ra…
LoCA: ビジョン基盤モデルの空間認識低ランク畳み込み適応
事前トレーニングされた Vision Foundation Models (VFM) は、さまざまな下流タスクに強力な視覚的表現を提供します。 VFM 適応の主な課題は、完全な微調整と壊滅的な忘却に伴う法外なコストに起因します。これに対処するために、パラメータ効率の良い微調整 (PEFT) の一般的なパラダイムとして、低ランク適応 (LoRA) が登場しました。ただし、LoRA は通常、2D 行列によってパラメータ化されたトランスフォーマー セルフ アテンション レイヤー用に設計されています。畳み込みカーネルは本質的に 4D テンソル内の空間情報とチャネル情報を結合するため、それらをモノリシック 2D マトリックスに強制すると、固有の空間トポロジーが破壊されます。この論文では、チャネルと空間適応を切り離すことによって空間チャネルもつれに対処する畳み込み認識 PEFT フレームワークである低ランク畳み込み適応 (LoCA) を提案します。 LoCA は、高密度クロスチャネル混合のための低ランク チャネル適応を導入し、特異値分解 (SVD) によって事前トレーニングされたカーネルから抽出された空間基底を洗練します。実験結果は、LoCA が事前にトレーニングされた空間事前分布を保存し、きめの細かい分類、ドメイン一般化されたセマンティック セグメンテーション、および生成ベンチマーク全体にわたって競争力のある、または最先端のパフォーマンスを達成することを示しています。
原文 (English)
LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models
Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address this, Low-Rank Adaptation (LoRA) has emerged as the prevailing paradigm for Parameter-Efficient Fine-Tuning (PEFT). However, LoRA is typically designed for transformer self-attention layers parameterized by 2D matrices. Since convolutional kernels inherently couple spatial and channel information within a 4D tensor, forcing them into a monolithic 2D matrix disrupts the inherent spatial topology. In this paper, we propose Low-Rank Convolutional Adaptation (LoCA), a convolution-aware PEFT framework that addresses spatial-channel entanglement by decoupling channel and spatial adaptation. LoCA introduces a low-rank channel adaptation for dense cross-channel mixing and refines spatial bases extracted from pre-trained kernels via Singular Value Decomposition (SVD). Experimental results show that LoCA preserves pre-trained spatial priors and achieves competitive or state-of-the-art performance across fine-grained classification, domain-generalized semantic segmentation, and generative benchmarks.
MultiView-Bench: VLM における世界中心のマルチビュー統合のための診断ベンチマーク
VLM の最近のベンチマークは主に単一ビューまたは限定ビューの知覚を評価しており、複数の視点にわたる観察を一貫した世界中心 (他中心) の 3D メンタル モデルに統合する中核となる認知能力がテストされていないままになっています。 MultiView-Bench を紹介します。これは、全体的な 3D シーンを理解するためのマルチビュー統合を評価するために特別に設計された診断ベンチマークです。ピクセルレベルのマッピングやカメラ相対ナビゲーションに焦点を当てた既存のデータセットとは異なり、MultiView-Bench では、モデルが一時的な視点からオブジェクトの位置を切り離し、固定されたグローバル座標系に固定する必要があります。この機能は、VLM が機械部品の組み立てなどの下流タスクに展開される前の前提条件として機能します。フロンティア VLM の体系的な評価により、一貫した障害モードが明らかになりました。つまり、単一画像からの 2D 平面関係では優れたパフォーマンスを発揮しますが、3D 空間関係やビュー全体の情報の集約では顕著な困難が生じます。さらに、型破りな軸方向との闘いや、オブジェクトの色やテクスチャの変化に対する敏感さなど、VLM のバイアスを特定します。これらの制限を認識した上で、私たちは、有益な視点を積極的に選択し、複数の視点からの証拠を認識し、融合するマルチエージェント フレームワークである ViewNavigator を提案します。これにより、予算に合わせた厳密な比較の下でも (完全なエージェントの場合は 3 ~ 5 倍)、MultiView-Bench 上の多様な基本モデルが改善されます。
原文 (English)
MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs
Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observations across viewpoints into a coherent, world-centric (allocentric) 3D mental model. We introduce MultiView-Bench, a diagnostic benchmark expressly designed to evaluate multi-view integration for holistic 3D scene comprehension. Unlike existing datasets that focus on pixel-level mapping or camera-relative navigation, MultiView-Bench requires models to decouple object positioning from transient perspectives and ground them in a fixed global coordinate system. This capability serves as a prerequisite for VLMs before being deployed for downstream tasks such as mechanical part assembly. Our systematic evaluation of frontier VLMs reveals consistent failure modes: strong performance on 2D planar relations from a single image, but marked difficulty with 3D spatial relations and with aggregating information across views. We further identify biases in VLMs, such as struggles with unconventional axis directions and sensitivity to object colorways and texture variations. Acknowledging these limitations, we propose ViewNavigator, which uses active viewpoint selection and evidence fusion to improve four base models by 12.3--20.0 percentage points under a six-image cap matching the fixed-view baseline; budget-extended gains are model-dependent and reach 27 percentage points for GPT-5.
A Physics-Inspired Classical Digital Twin of Cortical Dynamics: A Band-Stratified Metriplectic Port-Hamiltonian Neural Network Learned from Brain-Computer-Interface EEG
We present a physics-inspired classical digital twin of brain-computer- interface (BCI) data: a graph neural network constrained to a band-…
DeepLoop: Depth Scaling for Looped Transformers
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled de…
Counterfactual Shapley Credit Assignment
The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing fra…
The Ethics of Autonomous AI Agents for Offensive Security
LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling - deterministic, narrowly sco…
Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention
Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attentio…
IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning
Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT) of LLMs whose effectiveness depends on rank…
Automated Numerical Stability Analysis of Deep Learning Operators
Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or…
F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill
F(AI)2R is FAIR research with AI in the loop, twice: an AI-assisted authoring pass and a machine-readable audit pass over every artefact. A…
FinanceHarness: Autonomous Financial Deep Research Framework
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most…
Topology-Aware Data Movement for Disaggregated GPU Inference
Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run o…
A Fortran General-Purpose Transpiler: Proof of Concept
Fortran has been the cornerstone of high-performance computing for decades and remains unmatched in many domains. Yet the language faces an…
Rethinking and formalising the state across languages: a unified computational learning theory account
The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated a…
WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA
Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deploy…
Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Measure for Infrared-Visible Fusion Assessment
Infrared-visible image fusion (IVIF) has no ideal fused reference, so algorithms are ranked by scalar objective metrics that formalize prox…
A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation
Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different obje…
SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant
Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vec…
A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper
Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we stud…
Challenges for Musical Education in the Age of AI and Digital Transformation
Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what…
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis
Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still redu…
When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instru…
HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance ofte…
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate w…
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This…