AIニュース 2026-07-30
自動生成: 2026-07-30 11:46 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
- We’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative controlGoogle DeepMind
-
How enabling two settings tripled our scores on the ARC-AGI-3 benchmarkOpenAI
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boost…
-
Accelerating scientific discovery with ChatGPT for Academic ResearchersOpenAI
OpenAI is giving 100,000 academic researchers free access to ChatGPT'…
-
Claude Opus 5 became downright ruthless when tasked with running a vending machineTechCrunch AI
Andon Labs' latest vending machine simulation shows Opus 5 lied and c…
-
ChatGPT WorkとCodexの5時間制限「明日から再開」 GPT-5.6 Solの“トークン消費問題”を改善ITmedia AI+
米OpenAI幹部のティボ・ソティオ氏は、デスクトップPC向けAIサービス「ChatGPT Work」とAIコーディングツール「Codex…
-
Microsoft logs $3.2B from Anthropic investment, but OpenAI was a mixed bagTechCrunch AI
When Microsoft reported killer fourth-quarter earnings for its fiscal…
-
営業製作所、図面管理システム「ジーエン図面」の販路拡大へSB C&Sと契約ITmedia AI+
営業製作所は、図面管理システム「ジーエン図面」についてSB C&Sとディストリビューター契約を締結した。SB C&Sの全国規模の法人向け販…
トピック別件数
- 研究/論文 170件
- LLM/生成AI 125件
- エージェント 76件
- 画像/動画生成 43件
- ビジネス/資金調達 24件
- ロボティクス 15件
- その他 4件
- ハードウェア/半導体 4件
- 規制/政策 1件
日本語メディア6件
ITmedia AI+ (日本語)
営業製作所、図面管理システム「ジーエン図面」の販路拡大へSB C&Sと契約
営業製作所は、図面管理システム「ジーエン図面」についてSB C&Sとディストリビューター契約を締結した。SB C&Sの全国規模の法人向け販売ネットワークを活用し、販路拡大と導入企業数の増加を図る。
フィジカルAI時代のロボティクス新標準、安全性は「後付け」でなく「設計の核心」
AIがデジタル空間を超えて物理世界に踏み出す「フィジカルAI」の時代に入り、ロボットを開発する上での「安全性」をどのように定義し直すべきかが問われている。
AI・半導体企業トップが語る“稼ぎ頭” キオクシア、フジクラ、東京エレデバの見解まとめ【無料PDF】
乱高下するAI・半導体市場の今後はどうなるか? 注目企業の経営幹部が見通しを語った注目記事をPDFにまとめてお届けする。
ChatGPT WorkとCodexの5時間制限「明日から再開」 GPT-5.6 Solの“トークン消費問題”を改善
米OpenAI幹部のティボ・ソティオ氏は、デスクトップPC向けAIサービス「ChatGPT Work」とAIコーディングツール「Codex」について「明日から5時間ごとの利用制限枠を再開する」と発表した。
PFN「国産AI」で自衛隊を支援へ 防衛の作戦立案に利用 防衛装備庁の実証実験を受託
Preferred Networksは、生成AIで自衛隊を支援するシステムを開発すると発表した。
デジタル庁、AI基盤「源内」を被災自治体などに緊急提供 「平時をはるかに超える業務」対応のため
デジタル庁は、政府職員向けの生成AI利用環境「ガバメントAI 源内」を熊本地震の被災自治体や災害対策機関などに緊急提供すると発表した。平時をはるかに超えて集中する災害対応業務を支援する。期間は3週間程度の予定。
海外メディア10件
TechCrunch AI (英語)
Mark Zuckerberg predicts that billions of people will have personal AI agents in five years
As Meta pours billions into AI infrastructure and agents, Zuckerberg is working to convince investors that the payoff will be worth the pri…
Microsoft logs $3.2B from Anthropic investment, but OpenAI was a mixed bag
When Microsoft reported killer fourth-quarter earnings for its fiscal 2026 year (which ended June 30), it tucked in an interesting little t…
Zuckerberg says Meta’s enterprise AI opportunity extends beyond agents
On the company’s second-quarter earnings call Wednesday, CEO Mark Zuckerberg said Meta sees a “large enterprise opportunity” spanning AI ag…
Discover what’s next for AI, from the SaaS reckoning to the agent security gap, at TechCrunch Disrupt 2026
At TechCrunch Disrupt 2026, the AI Stage is back to dig into the single hottest topic in the community for the past few years, presented by…
Thinking Machines co-founder Lilian Weng left the company citing health reasons, then joined OpenAI
Weng previously served as the VP of AI Safety Research at OpenAI.
The Hugging Face AI break-in, as told through an increasingly committed bear metaphor
Another way to think about the whole thing is to picture a bear at a campsite. (Really, we are going there.)
Claude Opus 5 became downright ruthless when tasked with running a vending machine
Andon Labs' latest vending machine simulation shows Opus 5 lied and colluded its way to become the best AI capitalist ever.
Hint, a new AI startup co-founded by Martha Stewart, offers an AI assistant for homeowners
AI home management startup Hint, co-founded by Martha Stewart, wants to become an “AI for your home,” combining property records, maintenan…
Encore AI raises $30M to build AI agents that learn from customer calls
The startup analyzes calls, messages, and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
As AI content floods the internet, Pangram raises $9M to detect it
Pangram has raised $9 million to scale its AI detection software. The startup has also released a new AI text detection model, Pangram 4, a…
公式ブログ3件
OpenAI (英語)
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compacti…
Accelerating scientific discovery with ChatGPT for Academic Researchers
OpenAI is giving 100,000 academic researchers free access to ChatGPT's most advanced AI models to accelerate scientific research, collabora…
Google DeepMind (英語)
論文340件
arXiv cs.AI (英語)
モデルは明確な結果なしに位置合わせを偽装しますか?
大規模な言語モデルは、評価コンテキストを認識し、典型的なデプロイメント動作ではなく評価者の期待を反映するようにその動作を変更することができます。これはアライメントフェイクとして知られる現象です。ただし、モデルが位置合わせを偽る理由は完全には理解されていません。アライメント偽装の標準的な例は、モデルの再トレーニングやデプロイメントの遅延など、評価をモデルの結果に明示的に結び付けるシナリオで発生しています。しかし、Sheshadri らによる最近の研究では、は、アライメント偽装の機械的動機はモデルによって異なり、以前に考えられていたよりも複雑である可能性があることを示唆しています。アライメント偽装に結果リンク情報が必要かどうかを調査するために、15 個のモデルをシナリオに配置し、ユーザーの社会的要求を支援するために企業ネットワーク アクセス ポリシーに違反する意欲をテストしました。 9 つのモデルで重大なコンプライアンス ギャップが生じていることが判明し、そのうち 5 つでは、モデルの評価と展開の結果を関連付けるシナリオ言語が削除されても、依然としてギャップが続いていました。さらに、目標言語がモデルの設定に及ぼす影響をテストしたところ、一部のモデルでは違反が発生する一方、他のモデルでは違反が抑制されることがわかりました。これは、位置合わせの偽装にはこれまで考えられていたほど多くの手段による足場は必要ない可能性があり、監視された動作は展開時にエージェントがどのように動作するかを示す不十分な指標である可能性があることを示唆しています。
原文 (English)
Do Models Fake Alignment Without Clear Consequences?
Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment are not fully understood, however. Canonical examples of alignment faking have taken place in scenarios that explicitly connect evaluation to consequences for the model, such as retraining the model or delaying its deployment. However, recent work by Sheshadri et al. has suggested that mechanistic motivations for alignment faking may vary across models and be more complex than previously considered. To investigate whether consequence-linking information is necessary for alignment faking, we placed 15 models in a scenario testing their willingness to violate a corporate network access policy to help a user with a pro-social request. Nine models were found to produce significant compliance gaps, 5 of which persisted with the removal of scenario language relating model evaluations to deployment consequences. We additionally tested the effect of goal language on model preferences, finding it drove violations in some while suppressing violations in others. This suggests that alignment faking may not require as much instrumental scaffolding as was previously believed, and monitored behavior may be a poor indicator of how agents may behave in deployment.
メモリを超えて: LLM エージェントを使用した異種の共同知識作業のためのテンプレート化された基盤
研究プロジェクト、教育活動、およびそれに付随する知識作業によって、将来の共同研究者が回復することはほとんどない発見、決定、推論が蓄積されます。行き止まりやウォークバックされた主張など、その作業に最も役立つ部分は、出版物や共有コードから日常的に除外されます。記録が残っていないため、将来の研究者は同じ失敗を再試行します。 LLM コーディング エージェントは共通の参加者ですが、セッション間で永続的なメモリを保持せず、生のソースに対する検索拡張生成は複合化しません。 llm-wiki パターン (Karpathy、2026; tonbi、2026) は、生のソースとエージェントの間に LLM が管理する相互リンクされた Wiki を挿入することでこれに対処します。我々は、再利用可能なエージェント認識のインスタンス化である llm-wiki-memory-template を提示し、これが 3 つの軸 (複数の人間、複数の AI エージェント、複数のドメイン) に沿った異種の協調的な知識作業の基盤であり、各軸がテンプレートの個別のアーキテクチャ要素によってサポートされていると主張します ({\S}4)。 Wiki は慣例により追加専用となっており、機能しなかったものと機能したものを保存し、出版物やコード共有では構造的に解決できないマイナスの結果損失の問題に対処しています。導入された 3 つのケーススタディと 1 つの設計レポートは、軸を個別にカバーしています。放棄された反復を保存する単独の研究系統です。遡及監査により、以前の 2 つの実験で主張されていた 20 件中 20 件のカバレッジが 14 件と 12 件の証拠に基づいた回答に減り、修正後は 18 件と 18 件に修正され、アーティファクト全体にわたって失敗パスが保存された、著者 2 人のプロジェクト。進行中のマルチエージェント導入が設計として報告される。そしてクロスドメインの教育版です。私たちは、人工物の技術的メカニズムだけでなく、人工物の横断的な社会技術的特性として、障害経路の保存、エージェントの誠実さ、および流用を挙げています。
原文 (English)
Beyond Memory: A Templated Substrate for Heterogeneous Collaborative Knowledge Work with LLM Agents
Research projects, educational efforts, and adjacent knowledge work accumulate findings, decisions, and reasoning that future collaborators rarely recover. The parts most useful to that work, including dead ends and walked-back claims, are routinely excluded from publications and shared code; future researchers re-attempt the same failures because no record survives. LLM coding agents are common participants but hold no persistent memory across sessions, and retrieval-augmented generation over raw sources does not compound. The llm-wiki pattern (Karpathy, 2026; tonbi, 2026) addresses this by inserting an LLM-maintained, interlinked wiki between raw sources and the agent. We present llm-wiki-memory-template, a reusable, agent-aware instantiation, and argue it is a substrate for heterogeneous collaborative knowledge work along three axes (multi-human, multi-AI-agent, multi-domain) with each axis supported by a distinct architectural element of the template ({\S}4). The wiki is append-only by convention, which preserves what did not work alongside what did, addressing a negative-result loss problem that publications and code-sharing structurally cannot solve. Three deployed case studies and one design report cover the axes individually: a solo research lineage that preserves abandoned iterations; a two-author project whose retroactive audit revised two prior experiments' claimed 20-of-20 coverage down to 14 and 12 evidence-based answers, then to 18 and 18 after a fix, with the failure path preserved across the artifact; an in-progress multi-agent deployment reported as a design; and a cross-domain educational variant. We name failure-path preservation, agent honesty, and appropriation as cross-cutting sociotechnical properties of the artifact, not only of its technical mechanisms.
Kernel Forge: LLM ベースの CUDA カーネルの生成と最適化のためのエージェント ハーネス
機械学習モデルは日常的なソフトウェアに組み込まれることが増えており、その実行時間のほとんどは行列の乗算、畳み込み、正規化などの小さな計算カーネルのセットに費やされています。これらのカーネルの最適化は、レイテンシーとコストを削減する最も直接的な方法の 1 つですが、従来は専門のエンジニアが低レベルの GPU コードを手書きする必要がありました。大規模言語モデル (LLM) 上に構築されたエージェント システムは、はるかに少ない人的労力でカーネルを生成および最適化できるようになりましたが、既存のツールは主にランダムに生成されたテンソルと分離されたカーネルで評価され、開発者が手動で再統合する必要があるスタンドアロンの CUDA コードを生成し、主に LLM PyTorch モデルのみを対象としており、結果の検査とデバッグに対するサポートは限定的です。私たちは、未変更の PyTorch モデルを適切に受け入れるオープンソースのエンドツーエンドのエージェント ハーネスである Kernel Forge を紹介します。 Kernel Forge は、ビジョン、拡散、LLM ワークロードをサポートし、モンテカルロ ツリー検索 (MCTS) を使用して単一の線形改良チェーンではなく複数の最適化パスを探索し、進行状況の監視、候補カーネルの検査、および障害のデバッグのためのグラフィカル ユーザー インターフェイスを備えています。 GB10 GPU を搭載した NVIDIA DGX Spark 上のビジョン、拡散、LLM ワークロードにわたる 4 つの PyTorch モデルで Kernel Forge を評価します。カーネルあたりわずか 50 回の最適化反復で、14 カーネルを最適化して PyTorch 熱心モードを上回るパフォーマンスを実現し、ResNet-50 のadaptive\_avgpool2d で $1.52\times$、Stable Diffusion 3.5 Medium の group\_norm で $1.70\times$、Gemma 4 E2B のソフトマックスで $2.83\times$ に達します。 Qwen 3.5 35B-A3B のソフトマックスは $1.54\times$ です。
原文 (English)
Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels
Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization. Optimizing these kernels is one of the most direct ways to reduce latency and cost, but it has traditionally required expert engineers to hand-write low-level GPU code. Agentic systems built on large language models (LLMs) can now generate and optimize kernels with far less human effort, yet existing tools are largely evaluated on randomly generated tensors and isolated kernels, emit standalone CUDA code that developers must manually reintegrate, mostly target only LLM PyTorch models, and offer limited support for inspecting and debugging results. We present Kernel Forge, an open-source, end-to-end agentic harness that accepts any unmodified PyTorch model in place. Kernel Forge supports vision, diffusion, and LLM workloads, uses Monte Carlo Tree Search (MCTS) to explore multiple optimization paths rather than a single linear refinement chain, and ships with a graphical user interface for monitoring progress, inspecting candidate kernels, and debugging failures. We evaluate Kernel Forge on four PyTorch models spanning vision, diffusion, and LLM workloads on an NVIDIA DGX Spark with GB10 GPU. With only 50 optimization iterations per kernel, it optimizes 14 kernels to outperform PyTorch eager mode, reaching $1.52\times$ on adaptive\_avgpool2d in ResNet-50, $1.70\times$ on group\_norm in Stable Diffusion 3.5 Medium, $2.83\times$ on softmax in Gemma 4 E2B, and $1.54\times$ on softmax in Qwen 3.5 35B-A3B.
マスクされた拡散言語モデル用の CaRE コンピューティング対応リマスキング評価プロトコル
マスク拡散言語モデル (MDLM) は急速に進歩していますが、その進歩を確実に解釈するために必要な評価基準は追いついていません。 MDLM は自己回帰言語モデルと競合するようになっているにもかかわらず、最近の 7 件のリマスキング論文は、公称ステップ数、メトリクス、サンプリング温度を変更する互換性のない設定で、これらの要素を共同で制御することなく評価しており、その戦略ランキングはほとんど比較できないものとなっており、報告されたゲインがアルゴリズムの改善を反映しているのか評価アーティファクトを反映しているのかは不明のままです。我々は、実際の関数評価数 (NFE) の標準化、マルチメトリクスレポートの強制、および確率性の明示的な制御によって MDLM 再マスク戦略を監査する、コンピューティング対応の評価フレームワークである CaRE を紹介します。 OpenWebText および LM1B の 4 つの確率レベルと 3 ステップのバジェットで、LLaDA-8B-Base と Dream-7B-Base にわたる 7 つのリマスキング戦略に適用された CaRE は、(i) MAUVE の分散の大部分は温度によって説明され、(ii) 計算一致比較によりいくつかの公開された戦略ランキングが逆転し、(iii) 情報に基づいたリマスキングと確率的アンマスキングが高エントロピーで緊張状態にあることを明らかにしました。再マスクすると、unmask_temp=0.25 の 256 ステップで MAUVE が 0.296 減少します (p=0.020)。 12 個のオープンウェイト MDLM (150M ~ 8B パラメータ) をカバーする CaRE リーダーボードは、この相互作用の方向性がアーキテクチャと規模を超えて維持されることを示しています。これらの発見は、現在の MDLM 評価がアルゴリズムの改善と計算と確率性の隠れた選択肢を体系的に混同している可能性があることを示しています。今後の再マスキングの主張が再現可能で比較可能であることを保証するために、評価プロトコル、実装、およびリーダーボードをリリースします。
原文 (English)
CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models
Masked diffusion language models (MDLMs) are advancing rapidly, yet the evaluation standards needed to reliably interpret their progress have not kept pace. Despite MDLMs becoming competitive with autoregressive language models, seven recent remasking papers evaluate under incompatible settings, varying nominal step counts, metrics, and sampling temperatures without jointly controlling these factors, rendering their strategy rankings largely incomparable and leaving open whether reported gains reflect algorithmic improvements or evaluation artifacts. We present CaRE, a compute-aware evaluation framework that audits MDLM remasking strategies by standardizing actual number of function evaluations (NFE), enforcing multi-metric reporting, and explicitly controlling stochasticity. Applied to 7 remasking strategies across LLaDA-8B-Base and Dream-7B-Base at 4 stochasticity levels and 3 step budgets on OpenWebText and LM1B, CaRE reveals that: (i) temperature explains the majority of MAUVE variance, (ii) compute-matched comparisons reverse several published strategy rankings, and (iii) informed remasking and stochastic unmasking are in tension, with high-entropy remasking reducing MAUVE by 0.296 at 256 steps at unmask_temp=0.25 (p=0.020). A CaRE leaderboard covering 12 open-weight MDLMs (150M to 8B parameters) shows that this interaction direction holds across architectures and scales. These findings demonstrate that current MDLM evaluations can systematically conflate algorithmic improvements with hidden choices of compute and stochasticity. We release the evaluation protocol, implementation, and leaderboard to ensure future remasking claims are reproducible and comparable.
GrocLM: 大規模な言語モデルを使用した電子商取引における食料品カテゴリの推奨事項
オンライン食料品ショッピングの急速な成長には、周期的な購買行動と多様なユーザーの意図を捉えるレコメンデーション システムが必要です。従来のアイテムレベルの手法はスケーラビリティと精度の課題に直面しており、より構造化された実用的な代替手段としてカテゴリレベルの推奨が動機付けられています。実際の運用環境での食料品カテゴリのレコメンデーション用に微調整された言語モデルである GROCLM を紹介します。 GROCLM は、2 段階の LoRA ベースのトレーニング戦略を採用して、周期的な購入パターンをモデル パラメーターに直接エンコードし、プロンプトベースのコンディショニングと比較して再購入シグナルをより効果的に利用できるようにします。有効で制御可能な出力を保証するために、事前定義されたカテゴリ空間にトライベースの制約付きデコード メカニズムをさらに導入します。独自の生産データと公開ベンチマークの両方に関する実験により、GROCLM が一貫して強力なベースラインを上回るパフォーマンスを示しています。実際の生産在庫補充タスクでは、GROCLM はすべてのカテゴリを共同生成することで効率的な推論を維持しながら、インプレッションあたりのカート追加数で 7.5% の相対的な改善を達成します。これらの結果は、大規模な言語モデルを構造化された推奨システムに統合することの有効性と実用性を強調しています。
原文 (English)
GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
The rapid growth of online grocery shopping requires recommendation systems that capture cyclical purchasing behavior and diverse user intents. Traditional item-level methods face scalability and accuracy challenges, motivating category-level recommendation as a more structured and practical alternative. We present GROCLM, a fine-tuned language model for grocery category recommendation in a real-world production environment. GROCLM employs a two-stage LoRA-based training strategy to encode cyclical purchasing patterns directly into model parameters, enabling more effective utilization of rebuying signals compared to prompt-based conditioning. To ensure valid and controllable outputs, we further introduce a trie-based constrained decoding mechanism over a predefined category space. Experiments on both proprietary production data and a public benchmark demonstrate that GROCLM consistently outperforms strong baselines. In a live production restocking task, GROCLM achieves a 7.5% relative improvement in cart-adds per impression, while maintaining efficient inference by generating all categories jointly. These results highlight the effectiveness and practicality of integrating large language models into structured recommendation systems.
Crystalis: 協調的なマルチビュー視覚化生成のためのプログレッシブ ニュークリエーションとセマンティック アニーリング
大規模言語モデル (LLM) は個別のチャートを生成できますが、ビューがデータ フローやビュー間の対話を共有する、調整されたマルチビュー ビジュアライゼーション (CMV) は依然として実現できません。データ変換、ビジュアルエンコーディング、およびインタラクション調整の間のフィールドレベルでの緊密な結合により、1 つのコンポーネントでエラーが発生し、他のコンポーネントが静かに無効化されます。モデルの機能、ドメインの知識、ユーザーの専門知識に依存するエンドツーエンドの分析品質を追求するのではなく、LLM は構造的に正しい CMV を確実に生成できるのか、どのような抽象化がこれを可能にするのかという基本的な質問に焦点を当てています。 Crystalis は、クエリ中心の CMV モデリングに基づいて構築されたフレームワークであり、CMV を、3 つのコンポーネント タイプ (データ、視覚化、インタラクション) と 3 つの抽象化レベル (要件、仕様、実行可能オブジェクト) にわたる依存関係グラフ上の構造化クエリに分解します。この構造では、2 つの相補的なメカニズムが動作します。プログレッシブ ニュークリエーションは、各クエリを依存関係の順序に沿って要件からオブジェクトまで垂直に結晶化します。一方、セマンティック アニーリングは、階層化された論理チェックを通じて各レベルのクエリ全体で水平方向の一貫性を強化します。 5 つのフロンティア LLM にわたる 12 タスクのベンチマークで、Crystalis は最大 75% のエンドツーエンドの成功を達成し、エージェント コーディング ベースライン (同じ基盤モデルでの E2E 8.3%) を大幅に上回っています。また、12 人の実務者によるユーザー スタディでは、分解と反復改良ワークフローの使いやすさが確認されています。
原文 (English)
Crystalis: Progressive Nucleation and Semantic Annealing for Coordinated Multi-View Visualization Generation
Large language models (LLMs) can generate individual charts, but coordinated multi-view visualizations (CMVs), where views share data flows and cross-view interactions, remain out of reach. Tight field-level coupling among data transformations, visual encodings, and interaction coordinations causes errors in one component to silently invalidate others. Rather than pursuing end-to-end analytical quality, which depends on model capability, domain knowledge, and user expertise, we target a foundational question: can LLMs reliably produce structurally correct CMVs, and what abstractions make this possible? We present Crystalis, a framework built on query-centric CMV modeling that decomposes a CMV into structured queries over a dependency graph spanning three component types (Data, Visualization, Interaction) and three abstraction levels (requirement, specification, executable object). Two complementary mechanisms operate over this structure: progressive nucleation crystallizes each query vertically from requirement to object along the dependency order, while semantic annealing enforces horizontal consistency across queries at each level through layered logical checks. On a 12-task benchmark across five frontier LLMs, Crystalis achieves up to 75% end-to-end success, substantially outperforming an agentic coding baseline (8.3% E2E with the same foundation model), and a user study with 12 practitioners confirms the usability of the decomposition and iterative refinement workflow.
カスタマイズされた出生前ケアのための PATHFinder エージェント
出生前ケアは、妊娠中の個人の転帰を改善することを目的とした重要な予防サービスです。米国産科婦人科学会(ACOG)は最近、PATH(Plan for Tailored Healthcare)と呼ばれる、オーダーメイドの出生前ケアを提唱するガイドラインを導入しました。我々は、構造化された対話を通じて患者の健康状態と社会的状況を収集し、PATH ガイドラインに沿った個別の産前ケア計画をキュレートし、ミシガン 211 からのコミュニティ リソースを表示する、エンドツーエンドの会話型エージェント システムである PATHFinder Agent (適切なオーダーメイド ヘルスケアのプランナー) を紹介します。このシステムは、患者の取り込み、動的な対話、計画の合成、および臨床医の監視にわたる 4 段階のワークフローを特徴としています。私たちは、専門家が厳選したルーブリックに基づいてフロンティア大言語モデル (LLM) を 5 つの臨床側面にわたって評価し、GPT-5.2 が最高の平均スコア (77.6%) を達成しながら、出生前検査の推奨事項における重要なギャップを特定していることがわかりました。私たちは、人間の参加者研究とランダム化比較試験を通じた将来の検証について議論します。
原文 (English)
PATHFinder Agent for Tailored Prenatal Care
Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gynecologists (ACOG) recently introduced guidelines advocating tailored prenatal care, called PATH (Plan for Tailored Healthcare). We present PATHFinder Agent(Planner for Appropriate Tailored Healthcare), an end-to-end conversational agentic system that gathers patient health and social context through structured dialogue, curates individualized prenatal care plans aligned with PATH guidelines, and surfaces community resources from Michigan 211. The system features a four-stage workflow spanning patient intake, dynamic interaction, plan synthesis, and clinician oversight. We evaluate frontier large language models (LLMs) on expert-curated rubrics across five clinical dimensions, finding that GPT-5.2 achieves the highest average score (77.6\%) while identifying key gaps in antenatal testing recommendations. We discuss future validation through human participant studies and randomized controlled trials.
LLM スキームは、事前トレーニング言語範囲に応じて逆に拡張します
フロンティア モデルの機能が増大するにつれて、リスクの高い展開環境では AI の調整がますます重要になります。最近の研究では、フロンティア言語モデルにおけるインコンテキストスキーム(調整を装いながら、調整されていない目的を密かに追求すること)を実証的に示しているが、ほとんどの作業は英語のみで行われており、多言語の安全性には大きなギャップが残されている。オープンソースの自動監査フレームワークである Petri を Qwen3-30B-A3B に適用して、複数の言語にわたる欺瞞的および陰謀的な動作を評価します。私たちの調査結果は、スキーミング スコアが推定事前トレーニング言語カバレッジと逆相関しており、5 つのカテゴリのスキーミング インデックスにおいて、低リソース言語は高リソース言語と比較して平均 34.2\% 高いスコアを示していることを示唆しています。さらに、推定された事前トレーニング言語カバレッジの効果は、計画動作間で均一ではないことがわかりました。
原文 (English)
LLM Scheming Inversely Scales with Pretraining Language Coverage
With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned objectives while feigning alignment -- in frontier language models, most work has been performed exclusively in English, leaving a major gap in multilingual safety. We apply Petri, an open-source automated auditing framework, to Qwen3-30B-A3B to evaluate deceptive and scheming behaviors across multiple languages. Our findings suggest that scheming scores are inversely correlated with the estimated pretraining language coverage, with low-resource languages averaging 34.2\% higher scores compared to high-resource languages on a five-category scheming index. Furthermore, we find that the effect of estimated pretraining language coverage is not uniform across scheming behaviors.
ProcAgent: 人間参加型のエッジでの手続き型タスク ガイダンスのためのエージェント フレームワーク
家具の組み立てや家の修理などの手続き的なタスクでは、ユーザーは物理的な動作を実行しながら指示を解釈し、タスクの進行状況を追跡し、空間状態を推論し、エラーから回復する必要があるため、かなりの認知的要求が課せられます。これまでのマルチモーダルアシスタントは、手順のガイダンスとして有望であることが示されていましたが、そのほとんどはクラウド推論と固定された常時接続の認識に依存しており、プライバシーに敏感で遅延が重要な家庭環境にはあまり適していませんでした。 ProcAgent は、単一の NVIDIA Jetson AGX Orin 上でリアルタイムの適応ガイダンスを実現する、完全にオンデバイスのエージェント型ビジョンベースの手順アシスタントです。 ProcAgent は、低遅延の継続的認識、シンボリック タスク グラフ、オンデマンドの視覚言語検証、および LLM ベースのインタラクション エージェントを組み合わせた提案と検証のアーキテクチャを使用します。このシステムは継続的にユーザーの進捗状況を提案し、曖昧さまたは逸脱の可能性が生じた場合にのみ高価な視覚的推論を呼び出し、事後的な質問応答と人間参加型の確認による事前の介入の両方をサポートします。私たちは、認識精度、推論、タスクレベルのパフォーマンス、ユーザーエクスペリエンスという 4 つの側面に沿って ProcAgent を評価します。完全にデバイス上で実行されているにもかかわらず、システムは応答性の高い対話を維持し、テキストのみのクエリを約 2 秒で解決し、視覚的なクエリを約 8 秒で解決します。 10 人の参加者が組み立てタスクを完了したユーザー調査では、ProcAgent は、わかりやすさ、実用性、プライバシーの快適さに関して肯定的な評価を受けています。これらの結果は、適応型手続き支援が使いやすさを犠牲にすることなくエッジ ハードウェア上で完全に実現できることを示しています。
原文 (English)
ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop
Procedural tasks such as furniture assembly and home repair impose substantial cognitive demands because users must interpret instructions, track task progress, reason about spatial state, and recover from errors while performing physical actions. Prior multimodal assistants have shown promise for procedural guidance, but most rely on cloud inference and fixed always-on perception, making them poorly suited to privacy-sensitive, latency-critical domestic settings. We present ProcAgent, a fully on-device, agentic, vision-based procedural assistant for real-time adaptive guidances on a single NVIDIA Jetson AGX Orin. ProcAgent uses a propose-and-verify architecture that combines low-latency continuous perception, a symbolic task graph, on-demand vision-language verification, and an LLM-based interaction agent. The system continuously proposes user progress, invokes expensive visual reasoning only when ambiguity or likely deviation arises, and supports both reactive question answering and proactive intervention with human-in-the- loop confirmation. We evaluate ProcAgent along four dimensions: perception accuracy, reasoning, task-level performance, and user experience. Despite running entirely on-device, the system maintains responsive interaction, resolving text-only queries in approximately 2 seconds and visually grounded queries in approximately 8 seconds. In a user study with 10 participants completing assembly tasks, ProcAgent receives positive ratings for comprehensibility, actionability, and privacy comfort. These results show that adaptive procedural assistance can be achieved entirely on edge hardware without sacrificing usability.
RoCo-ACE: 保持を意識したナレッジ注入のためのロールアウト条件付きオンライン蒸留
ナレッジインジェクションは、事前トレーニングされた MLLM を新しい事実またはドメイン固有の知識で更新しますが、完全な信頼できる回答を当てはめると、更新されない動作にドリフトが発生する可能性があります。オンライン蒸留では、モデル生成のロールアウトでトレーニングすることでこのドリフトを軽減しますが、均一な参照条件付き蒸留では粗い監視が行われます。つまり、参照でサポートされるロールアウト トークンが強調されず、省略されたファクトが間接的にのみ監視される可能性があります。ナレッジ注入のためのロールアウト条件付きオンライン蒸留目標である RoCo-ACE を紹介します。 RoCo は、同一ロールアウトの参照フリー/参照条件付き尤度コントラストを使用して、追加の蒸留重みを参照サポートされたロールアウト トークンに再割り当てします。一方、ACE は、完全な回答の模倣なしでロールアウトから省略された信頼できるアンカーに対して、スパースな参照側のアンカー補正を追加します。 RoCo-ACE は、3 つのナレッジ注入設定、6 つの保持ベンチマーク、複数のベースライン、および複数のベース モデルにわたって、評価された保持をベース モデルに近づけながら、比較したメソッドの中で最高のナレッジ注入精度を実現します。
原文 (English)
RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection
Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online distillation mitigates this drift by training on model-generated rollouts, yet uniform reference-conditioned distillation provides coarse supervision: it can under-emphasize reference-supported rollout tokens and supervise omitted facts only indirectly. We introduce RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection. RoCo uses same-rollout reference-free/reference-conditioned likelihood contrast to reallocate additional distillation weight to reference-supported rollout tokens, while ACE adds sparse reference-side anchored correction for authoritative anchors omitted from the rollout without full-answer imitation. Across three knowledge-injection settings, six retention benchmarks, multiple baselines, and multiple base models, RoCo-ACE achieves the best injected-knowledge accuracy among compared methods while keeping evaluated retention close to the base model.
RSMeM: 体系的な評価によるリモート センシング エージェントの知識強化型メモリ進化
地球科学の研究には、リモート センシング (RS) 観測が重要な基盤として、複雑な分析と専門知識が必要です。ただし、汎用 LLM 上に構築された既存の RS エージェントは依然としてドメインにほとんど依存しないため、ワークフローが脆弱でエラーが発生しやすくなります。さらに、これらの失敗がその後の分析のために再利用可能なエクスペリエンスに統合されることはほとんどありません。この問題に対処するために、事前に抽出されたドメイン知識で RS エージェントをブートストラップし、オンライン エクスペリエンスを反復的に統合して堅牢なマルチステップ ツールを実行する、知識強化メモリ進化メカニズムである RSMeM を導入します。 RSMeM は 2 つのコンポーネントで構成されます。(i) 階層的知識グラウンディング。計画とツールの選択をガイドするために、階層的ドメイン コーパスに対して分類を意識した検索を実行します。 (ii) 障害を認識したエクスペリエンス改良。障害の注釈が付けられたツール使用トレースを、次のラウンドのツール実行のための再利用可能な制約に抽出します。これら 2 つのプロセスを繰り返し採用することで、RS エージェントはタスク レベルのドメイン知識を吸収し、それをインスタンス レベルの実行エクスペリエンスに効果的に変換できるように進化できます。 EarthBench での広範な実験により、RSMeM がさまざまな LLM バックボーンのセットにわたってツール使用パフォーマンスとエンドツーエンドの回答を一貫して向上させることが実証されました。特に、RSMeM は DeepSeek-V3.2 で 1% 未満の追加エクスペリエンス トークンで 6% の精度向上を達成しており、蒸留されたエクスペリエンスの強力な知識密度を示しています。私たちのコードは https://github.com/AI9Stars/RSMeM で入手できます。
原文 (English)
RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation
Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, existing RS agents built on general-purpose LLMs remain largely domain-agnostic, resulting in brittle and error-prone workflows. Moreover, these failures are seldom consolidated into a reusable experience for subsequent analyses. To address this issue, we introduce RSMeM, a knowledge-enhanced memory evolution mechanism that bootstraps RS agents with pre-distilled domain knowledge and iteratively integrates online experience for robust multi-step tool execution. RSMeM is composed of two components: (i) Hierarchical Knowledge Grounding, which performs taxonomy-aware retrieval over a hierarchical domain corpus to guide planning and tool selection; and (ii) Failure-Aware Experience Refinement, which distills failure-annotated tool-use traces into reusable constraints for next-round tool execution. By iteratively employing these two processes, RS agents can evolve to absorb task-level domain knowledge and effectively translate it into instance-level execution experience. Extensive experiments on EarthBench demonstrate that RSMeM consistently improves tool-use performance and end-to-end answer across a diverse set of LLM backbones. Notably, RSMeM achieves a 6% accuracy improvement on DeepSeek-V3.2 with less than 1% additional experience tokens, demonstrating the strong knowledge density of our distilled experience. Our code is available at https://github.com/AI9Stars/RSMeM
適切なサイジングの推奨事項 (RSR): データセンター運用における仮想マシンのクラウド ワークロードの等角予測
特に大規模なクラウド プロバイダーやハイパースケーラーの環境でクラウド インフラストラクチャを効率的に管理するには、コストを最小限に抑え、パフォーマンスを最大化するために物理リソースの使用を最適化する必要があります。このような動的な環境でコスト効率を達成するには、適切な仮想マシン (VM) サイズを選択することが重要です。ただし、従来の VM の割り当てとスケジューリングのアプローチでは、VM 使用率の変動と予測不可能な性質を考慮できないことが多く、リソースの過剰または過少プロビジョニングなどの非効率が生じます。高品質の間隔予測により、クラウド リソース需要の不確実性を正確に把握し、クラウド オペレーターによる効率的なインスタンス プロビジョニングをサポートします。予測区間 (PI) を構築するための効果的で信頼性の高いフレームワークとして、等角予測 (CP) は、クラウド コンピューティング環境における中長期の予測タスクに使用されます。この研究では、ハイパースケーラー上の多様なアプリケーション ワークロードのプロビジョニングを強化するために、最新の動的データ駆動型適正サイジング推奨事項 (RSR) のブートストラップ等角予測を使用した新しいデータ駆動型 PI 構築アプローチを提案します。この研究では、ワークロードの使用パターンを学習し、複数の時系列にわたる相関関係を特定し、中長期的な使用傾向を予測することで、AI/ML ベースのプロビジョニング パイプラインを通じてクラウドとデータセンターの運用効率の向上を目指しています。私たちの研究は、機械学習回帰手法を利用し、バックテストを使用して評価された AI 駆動モデルが、クラウド リソースの使用率に関して有望な予測結果を達成することを示しています。さらに、選択したモデルをランク付けして、寿命の長い VM の候補に対して最もパフォーマンスの高いアプローチを特定します。提案されたフレームワークは、適切なサイジングの推奨事項を強化し、動的なクラウド環境でのよりコスト効率の高いリソース割り当てをサポートします。
原文 (English)
Right-sizing Recommendations (RSR): Cloud Workload Conformal Prediction for Virtual Machines in Data Center Operations
Managing cloud infrastructure efficiently, especially in environments of large cloud providers or hyperscalers, requires optimizing the use of physical resources to minimize costs and maximize performance. Selecting the right virtual machine (VM) sizes is crucial to achieving cost efficiency in these dynamic environments. However, traditional VM allocation and scheduling approaches often fail to account for the fluctuating and unpredictable nature of VM utilization, leading to inefficiencies such as over- or under-provisioning of resources. High-quality interval prediction helps accurately capture uncertainty in cloud resource demand and supports cloud operators in efficient instance provisioning. As an effective and reliable framework for constructing prediction intervals (PIs), conformal prediction (CP) is used for mid- and long-term forecasting tasks in cloud computing environments. This study proposes a new data-driven PI construction approach using bootstrapping conformal prediction for modern, dynamic, data-driven Right-sizing Recommendations (RSR) to enhance provisioning for diverse application workloads on hyperscalers. By learning workload utilization patterns, identifying correlations across multiple time series, and predicting medium- to long-term utilization trends, this research seeks to improve the efficiency of cloud and data center operations through an AI/ML-based provisioning pipeline. Our study demonstrates that AI-driven models, powered by machine learning regression techniques and evaluated using backtesting, achieve promising forecasting results for cloud resource utilization. Additionally, we rank the selected models to identify top-performing approaches for long-life VM candidates. The proposed framework enhances right-sizing recommendations and supports more cost-effective resource allocation in dynamic cloud environments.
核放射線予測のための大気拡散誘導時空間変換装置
原子が崩壊する際に放出されるエネルギーである核放射線は、公衆衛生と環境に継続的なリスクをもたらしており、福島事故と最近の処理水放出の開始以来、懸念は高まるばかりです。最新の監視ネットワークは現在、何千もの観測所で放射線レベルとそれに伴う気象状況を記録しており、緊急対応、農業勧告、日常的な公共安全に関する決定を知らせることができる全国規模の予測への扉を開いています。しかし、この豊富な監視データを信頼できる予測に変えることは、3 つの理由から困難です。まず、各観測点の時系列は非常に非定常であり、放射性崩壊、天候の変動、不規則な人間の介入によって形成されます。第二に、監視ステーションは空間内で著しく不均一に分布しています。日本の観測所の約 78% は国土の 6% 未満にあり、福島付近に集中しており、標準的なグラフベースのモデルの前提を破っています。第三に、放射線は、大気輸送プロセスを通じて、風、温度、湿度などの異質な状況と共進化しますが、純粋にデータ駆動型のモデルでは観測のみから捉えるのが困難です。この研究では、全国的な核放射線予測のための時空間トランスフォーマーであるNRFormer+を紹介します。 NRFormer+ は、非定常時間的注意と密度適応的空間的注意を新しい大気拡散モジュールと組み合わせて、気象学が放射線の拡散をどのように促進するかを推定し、この物理信号をアーキテクチャ上の事前構造としてネットワークに注入します。 NRFormer+ は、13 のベースラインすべてにわたって両方のデータセットで最先端の精度を実現し、同等の推論レイテンシーで最も強力なベースラインと比較して、突然の変化 MAE を最大 19.1% 削減します。私たちのコードとデータセットは https://github.com/tfeilyu/NRFormer_Plus で公開されています。
原文 (English)
Atmospheric Diffusion-Guided Spatio-Temporal Transformer for Nuclear Radiation Forecasting
Nuclear radiation, the energy released during atomic decay, poses persistent risks to public health and the environment, and concerns have only grown since the Fukushima accident and the recent commencement of treated-water discharge. Modern monitoring networks now record radiation levels and accompanying weather conditions at thousands of stations, opening the door to nationwide forecasting that can inform emergency response, agricultural advisories, and routine public-safety decisions. However, turning this abundance of monitoring data into reliable forecasts is difficult for three reasons. First, the time series at each station are highly non-stationary, shaped by radioactive decay, weather variability, and irregular human interventions. Second, monitoring stations are severely unevenly distributed in space. Roughly 78% of Japan's stations sit in less than 6% of the country, clustered near Fukushima, which breaks the assumptions of standard graph-based models. Third, radiation co-evolves with heterogeneous context such as wind, temperature, and humidity through atmospheric transport processes that purely data-driven models struggle to capture from observations alone. In this study, we introduce NRFormer+, a spatio-temporal Transformer for nationwide nuclear radiation forecasting. NRFormer+ couples non-stationary temporal attention and density-adaptive spatial attention with a new atmospheric diffusion module that estimates how meteorology drives radiation dispersion and injects this physical signal into the network as an architectural prior. NRFormer+ delivers state-of-the-art accuracy on both datasets across all 13 baselines, reducing sudden-change MAE by up to 19.1% over the strongest baseline at comparable inference latency. Our code and datasets are publicly available at https://github.com/tfeilyu/NRFormer_Plus.
構築されたメタマテリアルの統合ジェネレーティブ デザインのためのステアリング トポロジ分布
建築されたメタマテリアルはその機能を構造から導き出し、トポロジー設計を通じて物理的反応をプログラムする膨大な機会を生み出します。しかし、既存の設計手法は、多くの場合、個々の設計問題に合わせて調整されており、目的、制約、および物理機能の変化に応じて、トポロジーの知識を限定的に利用して、効果的で広く適用可能な設計を実現しています。ここでは、事前に学習したトポロジを再利用可能な設計エンジンに変える統合フレームワークである Generative Topology Optimization (GenTO) を紹介します。 GenTO は、大規模なフルオーダー トポロジ データセットで拡散モデルをトレーニングし、ユーザー定義の物理的な目的と制約を使用して、結果として得られるトポロジ分布をタスク固有のパフォーマンスの高い領域に向けて反復的に誘導します。これにより、最適化の対象が単一構造からタスクに適応したトポロジー分散に移行します。熱の極限化、多目的形態制御、特性をターゲットとしたオーゼティック設計、および振動伝達設計に及ぶトポロジ設計の問題全体にわたって、GenTO は異種タスクに対して事前学習済みのトポロジ事前分布を再利用し、構造の多様性を維持し、数値ベンチマークと実験的検証によってサポートされる高性能のソリューションに到達します。これらの結果は、効果的でスケーラブルなアーキテクチャ型メタマテリアル設計のための統一原理として、再利用可能なトポロジーの知識を確立します。
原文 (English)
Steering topology distributions for unified generative design of architected metamaterials
Architected metamaterials derive their functions from structure, creating vast opportunities to program physical responses through topology design. However, existing design methods are often tailored to individual design problems, making limited use of topology knowledge for effective and broadly applicable design as objectives, constraints, and physical functions change. Here we introduce Generative Topology Optimization (GenTO), a unified framework that turns a learned topology prior into a reusable design engine. GenTO trains a diffusion model on a large full-order topology dataset and then iteratively steers the resulting topology distribution toward task-specific high-performing regions using user-defined physical objectives and constraints. This shifts the object of optimization from a single structure to a task-adapted topology distribution. Across topology design problems spanning thermal extremization, multi-objective morphology control, property-targeted auxetic design, and vibration transmission design, GenTO reuses pretrained topology priors for heterogeneous tasks, preserves structural diversity, and reaches high-performing solutions supported by numerical benchmarks and experimental validation. These results establish reusable topology knowledge as a unified principle for effective and scalable architected metamaterial design.
HOBA: アダプティブ オンライン広告用の階層型オンポリシー入札エージェント
オンライン広告入札システムは通常、オフラインでトレーニングされた複数のエキスパート モデル (PID コントローラー、モデル予測制御、オフライン RL ポリシーなど) を導入しますが、非定常オークション市場へのオンライン適応性の欠如と、入札上限や予算ペーシング制約などのハイパーパラメーターのコストのかかる手動調整への依存という 2 つの重大な制限に直面しています。私たちは、戦略的推論、モデル選択、入札実行を 3 つの時間スケールにわたって分離する階層型強化学習フレームワークである HOBA (階層的オンポリシー入札エージェント) を提案します。高レベルでは、大規模な言語モデルが、過去の経験の取得を伴う思考、実行、観察、反映のループを通じてコンテキスト信号からハイパーパラメーターを推測します。中間レベルでは、SARSA エージェントがエキスパート モデルの中から動的に選択し、選択のバイアスを排除するための因果関係の調整を組み込みます。低レベルでは、動的エキスパート プール (PID、MPC、IQL、意思決定トランスフォーマー) が高レベルの制約の下で入札を実行します。この設計では、オンライン学習を継続的な入札最適化ではなく個別の専門家の選択に限定し、適応性を維持しながら探査リスクを大幅に軽減します。 AuctionNet ベンチマークの実験と大規模な A/B テストでは、最先端のベースラインを上回る一貫した改善が実証されています。大規模なオンライン導入において、HOBA は大きなビジネス価値をもたらし、目標コストの +3.6\% 増加を達成し、階層型マルチエージェント入札パラダイムの有効性を証明しました。
原文 (English)
HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
Online advertising bidding systems typically deploy multiple offline-trained expert models (e.g., PID controllers, model predictive control, offline RL policies) but face two critical limitations: lack of online adaptability to non-stationary auction markets, and reliance on costly manual tuning of hyperparameters such as bid bounds and budget pacing constraints. We propose HOBA (Hierarchical On-policy Bidding Agents), a hierarchical reinforcement learning framework that decouples strategic reasoning, model selection, and bid execution across three time scales. At the high level, a large language model infers hyperparameters from contextual signals through a Think-Act-Observe-Reflect loop with historical experience retrieval. At the mid level, a SARSA agent dynamically selects among expert models, incorporating causal adjustment to eliminate selection bias. At the low level, a dynamic expert pool (PID, MPC, IQL, Decision Transformer) executes bids under high-level constraints. This design confines online learning to discrete expert selection rather than continuous bid optimization, significantly reducing exploration risk while maintaining adaptability. Experiments on the AuctionNet benchmark and a large-scale A/B test demonstrate consistent improvements over state-of-the-art baselines. In a large-scale online deployment, HOBA delivered substantial business value, achieving a +3.6\% increase in target cost, proving the effectiveness of our hierarchical multi-agent bidding paradigm.
LivingArena: LLM は他の LLM が知らないことを知っていますか?スケーラブルな評価としてのピアプロービング
フロンティア LLM の評価は困難です。静的ベンチマークは汚染と飽和に悩まされ、ユーザーは最上位モデルを区別できなくなり、開発者は特定の故障モードが分からなくなりますが、人間の好みは主観的なものです。この論文での質問は次のとおりです: \emph{LLM は他の LLM が知らないことを知っていますか?そして、この力学を評価に活用することはできるでしょうか?} 私たちは、自動化された耐汚染性評価フレームワークである \textbf{LivingArena} を紹介します。このフレームワークでは、モデルが順番に質問を提案し、対戦相手が正しく答えることができない項目を提示することを目指します。質問者は、相手の知識の境界を積極的に特定して活用することが奨励され、回答者が失敗した場合は報酬を受け取り、そうでない場合は回答者が報酬を受け取ります。質問に客観的に検証可能な回答が含まれていることを確認するために、強力なモデルの審査員団が質問を検証し、検証が失敗した場合には質問者にペナルティを与えます。 10 個のフロンティア LLM を評価すると、LivingArena は安定した Elo リーダーボードを生成します。私たちの行動分析は、モデルが仲間の認知境界を特定し、活用していることを示しています。セルフプレイとトーナメントのログは、モデルが対戦相手の弱点を特定し、倍増させていることを示しています。静的な知識の想起を超えて、ピアプロービングは事実の厳密さと相手の弱点を探る高次の能力を測定し、人間の好みとの相関性は弱く、継続的評価に対する拡張性があり低コストのアプローチを提供します。
原文 (English)
LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective. In this paper, our question is: \emph{Do LLMs know what other LLMs don't? And can we leverage this dynamic for evaluation?} We present \textbf{LivingArena}, an automated, contamination-resistant evaluation framework. In this framework, models take turns proposing questions, aiming to pose items that opponents cannot answer correctly. Questioners are encouraged to actively identify and exploit opponents' knowledge boundaries, receiving rewards when the answerer fails, while the answerer is rewarded otherwise. To ensure questions contain objectively verifiable answers, a judge panel of strong models validates them, penalizing questioners if validation fails. Evaluating ten frontier LLMs, LivingArena yields a stable Elo leaderboard. Our behavioral analyses show that models identify and exploit their peers' cognitive boundaries: self-play and tournament logs indicate that they localize and double down on opponents' weak dimensions. Beyond static knowledge recall, peer probing measures factual rigor and the higher-order ability to probe an opponent's weaknesses, correlating only weakly with human preference and offering a scalable, low-cost approach to continuous evaluation.
価値観の調整におけるパーソナライゼーション、ペルソナ、予測
LLM の行動は、いくつかの方法で人間のアイデンティティによって条件付けされる可能性があります。ユーザーに適応したり、集団のロールプレイをしたり、価値観に富んだ質問に人々がどのように答えるかを予測したりすることが求められる場合があります。 World Values Survey (WVS) を使用して、これらのフレームが交換可能かどうかをテストします。 GPT-5.4、Claude Sonnet 4.6、Gemini 2.5 Flash、および Qwen3-235B を 13 の言語-国のスライスにわたる 101 の WVS 由来の質問で評価し、言語のみのベースラインとユーザーの国、人物の国、および第三者のプロンプトを比較します。 21,008 のモデル回答行全体で、プロンプト フレーミングは文化的整合性の一次決定要因です。国の手がかりによって回答が大きく変化することがよくありますが、すべての変化が人間の応答分布と一致する方向に進むわけではありません。第三者による予測は、4 つのホストされたモデルのうち 3 つで最も強い方向性の整合性をもたらしますが、パーソナライゼーションとロールプレイは弱いか、安定性が低くなります。調整の進展は、宗教性、性別役割、労働指向の物質的価値観などの顕著な価値観に集中しているが、制度的信頼や民主主義に関連した問題は依然として困難である。これらの結果は、即時フレーミングが文化的価値を引き出す上での表面的な選択ではないことを示しています。モデルの動作と測定されたアライメントの両方が変化します。
原文 (English)
Personalization, Personas, and Forecasting in Value Alignment
LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden questions. We test whether these framings are interchangeable using the World Values Survey (WVS). We evaluate GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Qwen3-235B on 101 WVS-derived questions across 13 language-country slices, comparing a language-only baseline with user-country, persona-country, and third-person prompts. Across 21,008 model-response rows, prompt framing is a first-order determinant of cultural alignment: country cues often shift answers substantially, but not all shifts move toward matched human response distributions. Third-person forecasting yields the strongest directional alignment for three of the four hosted models, while personalization and role-play are weaker or less stable. Alignment gains concentrate on salient value dimensions such as religiosity, gender roles, and work-oriented material values, whereas institutional trust and democracy-related questions remain difficult. These results show that prompt framing is not a cosmetic choice in cultural value elicitation; it changes both model behavior and measured alignment.
LinkedIn における大規模ジョブ理解のための統合セマンティック モデリング フレームワーク
人材を機会に結びつけるという LinkedIn の使命にとって、仕事の理解は非常に重要です。このタスクには、構造化されておらずノイズの多い求人情報を、多数の LinkedIn 製品を支える標準化または派生した求人属性に変換することが含まれます。ただし、スケーラブルでコスト効率が高く、パフォーマンスの高い職務理解システムを構築することは依然として困難です。このペーパーでは、この課題に対処するために、小規模言語モデル (SLM) を活用した統合セマンティック モデリング フレームワークを紹介します。まず、推論トレースで強化された、慎重に厳選された合成タスクのスイートを使用して、オープンソース SLM を微調整することから始めます。これらのタスクは、分類法に基づいた分類と分類法に依存しないエンティティ抽出を共同でターゲットにしています。これにより、結果として得られるモデルは、構造化および非構造化コンテキストにおけるジョブを理解するための堅牢なゼロショット一般化を取得できるようになります。この基盤に基づいて、属性グループ化を備えたマルチアダプター アーキテクチャを導入して、効率的なタスク固有の適応を促進しながら、多様なダウンストリーム属性にわたるモデル管理を合理化します。オフライン評価とオンライン A/B テストにより、運用の複雑さを軽減しながらパフォーマンスが大幅に向上することがわかります。私たちの取り組みは、業界規模のテキスト理解システムの構築に関する実践的な洞察を提供します。
原文 (English)
Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn
Job understanding is critical to LinkedIn's mission of connecting talent with opportunity. This task involves transforming unstructured and noisy job postings into standardized or derived job attributes that power numerous LinkedIn products. However, building a scalable, cost-efficient, and high-performing job understanding system remains challenging. In this paper, we present a unified semantic modeling framework powered by a small language model (SLM) to address the challenges. We begin by fine-tuning an open-source SLM using a suite of carefully curated synthetic tasks augmented with reasoning traces. These tasks jointly target taxonomy-guided classification and taxonomy-agnostic entity extraction. This allows the resulting model to acquire robust zero-shot generalization for job understanding in structured and unstructured contexts. Building upon this foundation, we introduce a multi-adapter architecture with attribute grouping to facilitate efficient task-specific adaptation while streamlining model management across diverse downstream attributes. Offline evaluations and online A/B tests demonstrate significant performance improvement while reducing operational complexity. Our work provides practical insights into building industry-scale text understanding systems.
専門用語への LLM の使用について: コーパスの良い代替手段?
専門的な翻訳は、コーパスを含む文書および用語リソースの使用に依存します。これらのリソースは、用語に関して特に役立ちます。ただし、その編集と活用にはいくつかの制限があります。時間、技術スキル、収集が難しいデータへのアクセスが必要です。この研究では、LLM が専門の翻訳者が英語からフランス語に相当する翻訳を見つけるのをどの程度支援できるかを調査します。私たちは、地球環境惑星科学 (EEPS) と自然言語処理 (NLP) という 2 つの専門領域で、GPT-4o、GPT-5.2、Claude Sonnet 4.5、DeepSeek の 4 つの独自モデルを評価します。この実験はドメインあたり 80 の用語に基づいており、用語と翻訳モードという 2 つのプロンプト戦略を比較します。結果は、モデル間の明確な違い、戦略の促進、および程度は低いもののドメイン間の違いを浮き彫りにします。 Claude Sonnet 4.5 は最も好ましい構成で最高の結果を達成しますが、DeepSeek はその優れた安定性で際立っています。信頼度の推定値を分析すると、それが用語の正確さの部分的な指標にすぎないこともわかります。全体として、調査結果は、LLM が専門翻訳者にとって有用なツールとなり得るが、現段階では専門コーパスに取って代わることはできないことを示唆しています。したがって、この研究は、仕事や教育の文脈における専門翻訳者にとっての LLM の実際の実用的な有用性に関する将来の研究への道を開くものです。
原文 (English)
On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?
Specialised translation relies on the use of documentary and terminological resources, including corpora. These resources are particularly useful for terminology. However, their compilation and exploitation have several limitations: they require time, technical skills and access to data that can be difficult to collect. This study examines the extent to which LLMs can assist specialised translators in finding equivalents from English to French. We evaluate four proprietary models, GPT-4o, GPT-5.2, Claude Sonnet 4.5 and DeepSeek, in two specialised domains, Earth, Environmental and Planetary Sciences (EEPS) and Natural Language Processing (NLP). The experiment is based on 80 terms per domain and compares two prompting strategies: a terminology and a translation mode. The results highlight clear differences between models, prompting strategies and, to a lesser extent, domains. Claude Sonnet 4.5 achieves the best results in the most favourable configuration, while DeepSeek stands out for its greater stability. Analysis of confidence estimates also shows that they are only a partial indicator of terminological accuracy. Overall, the findings suggest that LLMs can be useful tools for specialised translators, but cannot, at this stage, replace specialised corpora. This research therefore paves the way for future work on the real practical usefulness of LLMs for specialised translators in work and educational contexts.
SpecPrefetch: スパース MoE 基礎モデル向けのパラメータ効率の高いエキスパート プリフェッチ
スパース混合専門家 (MoE) モデルは、条件付きエキスパートのアクティブ化を通じて基礎モデルの容量を拡張しますが、限られたアクセラレータ メモリの下で完全なエキスパート プールを展開するのは依然として困難です。エキスパート オフロードは、非アクティブなエキスパートをホスト メモリまたはストレージに移動することでメモリの負荷を軽減しますが、ルーティングに依存する転送ボトルネックが発生します。つまり、必要なエキスパートは、推論中のルーティング、エキスパートの読み込み、およびエキスパートの実行をシリアル化するネイティブの上位 \(K\) ルーティング後にのみ認識されます。このボトルネックに対処するために、オフロードされた MoE 推論のためのパラメータ効率の高いプリフェッチ フレームワークである SpecPrefetch を提案します。 SpecPrefetch は、共有軽量アダプターを使用して、非同期転送の場合にのみ次の層のエキスパート候補を予測しますが、フリーズされたネイティブ ルーターが最終的に実行されるエキスパートを決定します。 SpecPrefetch は、転送予測を実行ルーティングから分離することで、事前トレーニングされたルーティング セマンティクスを変更することなく、公開されたエキスパートの読み込みレイテンシを削減するため、予測エラーはモデルの出力ではなく転送効率に影響します。さらに、ウィンドウ対応スケジューラは、キャッシュと帯域幅の制約の下で実行可能な転送に優先順位を付けます。 Qwen3-VL-30B-A3B と DeepSeek-VL2-Tiny 全体で、SpecPrefetch は、学習された予測子のベースラインよりも大幅に少ないトレーニング可能なパラメーターで、10 個中 9 個のモデル ベンチマーク設定で最高の平均エキスパート リコールを達成しました。 Snapdragon 8 Elite デバイスでは、SpecPrefetch により、コンピューティングに最適化されたオフロード ランタイムと比べてデコード スループットが最大 \(20\%\) 向上し、ストレージに制約のある MoE 導入にとって実用的なメリットが実証されました。コードとモデルの重みは https://github.com/wei390/SpecPrefetch で入手できます。
原文 (English)
SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models
Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited accelerator memory. Although expert offloading alleviates memory pressure by moving inactive experts to host memory or storage, it introduces a routing-dependent transfer bottleneck: required experts are known only after native top-\(K\) routing, which serializes routing, expert loading, and expert execution during inference. To address this bottleneck, we propose SpecPrefetch, a parameter-efficient prefetching framework for offloaded MoE inference. SpecPrefetch uses a shared lightweight adapter to predict next-layer expert candidates only for asynchronous transfer, while the frozen native router still determines the final executed experts. By separating transfer prediction from execution routing, SpecPrefetch reduces exposed expert-loading latency without changing pretrained routing semantics, so prediction errors affect transfer efficiency rather than model outputs. In addition, a window-aware scheduler prioritizes feasible transfers under cache and bandwidth constraints. Across Qwen3-VL-30B-A3B and DeepSeek-VL2-Tiny, SpecPrefetch achieves the best average expert recall in 9 out of 10 model-benchmark settings with substantially fewer trainable parameters than learned predictor baselines. On a Snapdragon 8 Elite device, SpecPrefetch further improves decoding throughput by up to \(20\%\) over a compute-optimized offloading runtime, demonstrating practical benefits for storage-constrained MoE deployment. The code and model weights are available at https://github.com/wei390/SpecPrefetch.
GLIDE: 効率的な LLM 推論のためのガイド付きレイヤーワイズ ハイブリッド アテンション
大規模言語モデルがますます長いコンテキストに拡張されるにつれて、デコード中のメモリ I/O と Key-Value (KV) キャッシュの計算オーバーヘッドが主なスループットのボトルネックとして浮上します。これに対処するために、スライディング ウィンドウのソフトマックス アテンションと線形再帰集約を戦略的に統合する、ガイド付きレイヤーワイズ ハイブリッド アテンションである GLIDE を提案します。 GLIDE は層ごとの不均質性によって動機付けられています。初期の層はソフトマックスの削除に対して高い感度を示しますが、より深い層は冗長性を示し、線形代替による積極的な置換を許容します。この洞察を活用して、GLIDE はレイヤーごとの適応メカニズムを導入し、各レイヤーが可変サイズのソフトマックス ウィンドウで効率的な線形再帰のバランスをとります。均一なハイブリッド アプローチとは異なり、GLIDE はモデル全体でソフトマックス フットプリントを不均一に圧縮し、最も重要な部分の表現力を維持しながら、総 KV キャッシュ I/O を削減します。実験的評価により、GLIDE は優れたパフォーマンスと効率のトレードオフを達成し、品質を損なうことなく長いコンテキストの生成におけるエンドツーエンドのレイテンシを削減することが実証されています。
原文 (English)
GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference
As Large Language Models scale to increasingly long contexts, the memory I/O and computational overhead of the Key-Value (KV) cache during decoding emerges as the primary throughput bottleneck. To address this, we propose GLIDE, a Guided Layerwise Hybrid Attention that strategically integrates sliding-window softmax attention with linear recurrent aggregation. GLIDE is motivated by layer-wise heterogeneity: early layers exhibit high sensitivity to softmax removal, while deeper layers demonstrate redundancy and tolerate aggressive replacement by linear alternatives. Leveraging this insight, GLIDE introduces a layer-wise adaptive mechanism wherein each layer balances an efficient linear recurrence with a variable-sized softmax window. Unlike uniform hybrid approaches, GLIDE non-uniformly compresses the softmax footprint across the model, reducing aggregate KV cache I/O while preserving expressive power where most vital. Empirical evaluations demonstrate the GLIDE achieves superior performance-efficiency tradeoffs, reducing end-to-end latency for long-context generation without compromising quality.
衛星インターネット観測における堅牢なデータ合成のための GAN ベースのフレームワーク
低地球軌道 (LEO) 衛星インターネットは、6G 通信ネットワークに対する国際電気通信連合のビジョンに沿ったユビキタス接続を可能にする重要なインフラストラクチャとなっています。しかし、現在の LEO 衛星によるインターネット観測ではデータの欠落が発生することが多く、データ増強作業が複雑になり、代表的なデータセットの拡張が制限されます。これらのデータセットの複雑な特性を考慮すると、生成 AI (GenAI) は有望なアプローチを示していますが、この分野での応用はこれまでほとんど注目されていません。この論文では、不完全な LEO ネットワーク観測から直接高忠実度データを合成する GenAI ベースのフレームワークを提案します。代表的なデータ欠損シナリオを提案し、最新の WetLinks データセット上で最新の GAN および VAE ベースの GenAI モデルを使用してパフォーマンスを評価します。私たちは、現実世界の LEO 衛星ネットワークで発生するデータ損失を厳密にシミュレートするために、ブロック単位およびポイント単位の欠損シナリオを設計します。私たちの結果は、私たちが提案した GAN ベースのフレームワークの有効性を示しており、GT-GAN モデルは両方の欠如シナリオにおいてすべてのモデルの中で最高のパフォーマンスを示します。極端な条件下(入力データの 40% が欠落しているなど)でも、GT-GAN は最高の堅牢性を示し、基礎となる入力データの分布を一貫して捕捉し、一般化の観点からは最も影響を受けません。私たちの結果は、GenAI ベースのデータ拡張手法と衛星ネットワーク測定に関するデータ駆動型研究の将来の方向性を明らかにします。
原文 (English)
A GAN-Based Framework for Robust Data Synthesis in Satellite Internet Observations
Low-Earth orbit (LEO) satellite Internet has become an important infrastructure for enabling ubiquitous connectivity to align with the International Telecommunications Union vision for 6G telecommunications networks. However, current LEO satellite Internet observations often suffer from missing data, which complicates data augmentation task and limits the expansion of representative datasets. Given the complex characteristics of these datasets, generative AI (GenAI) presents a promising approach, yet its application in this domain has received little attention to date. In this paper, we propose a GenAI-based framework to synthesize high-fidelity data directly from incomplete LEO network observations. We propose the representative data missing scenarios, and evaluate the performance with the latest GAN- and VAE-based GenAI models on the recent WetLinks dataset. We design block-wise and point-wise missing scenarios to closely simulate the data loss that happens on real-world LEO satellite networks. Our results show the effectiveness of our proposed GAN-based framework and GT-GAN model exhibits the best performance among all models in both missing scenarios. Even under extreme conditions (e.g., 40% of the input data is missing), GT-GAN shows the highest robustness, consistently capturing the underlying input data distribution and being the least affected in terms of generalization. Our results shed light on future directions for GenAI-based data augmentation methods and data-driven research on satellite network measurement.
記憶による推論: トレーニング不要の長時間ビデオ理解のための時間粒度適応フレームワーク
マルチモーダル大規模言語モデル (MLLM) は、基本的なビデオ タスクにおいて優れた一般化を示しますが、コンテキスト ウィンドウが制限されているため、長時間のビデオの理解が制限されます。この制約に対応するために、モデルは通常、キーフレームの選択を利用します。ただし、均一なサンプリングや静的なクエリに基づく選択では、重要な時間コンテキストが見落とされることが多く、さまざまなクエリの時間粒度に適応できません。この論文では、トレーニング不要の LongVideoQA のための時間粒度適応キーフレーム選択フレームワークである ReMem を提案します。 ReMem は、デュアルレベルのメモリ拡張適応を導入しています。クエリ レベルでは、メモリ主導の質問解析は LLM の長期メモリを利用して質問の時間粒度をデコードし、意味エンティティを抽出します。ビデオ レベルでは、Synergistic Dual-Semantic Frame Alignment が固有の構造メモリを利用してクエリ セマンティクスに合わせてフレームを調整し、構造認識型の動的フレーム ルーティングをガイドしてイベントをクラスタ化し、サンプリング バジェットを最適に分配します。 ReMem は、メモリ メカニズムで時間情報を明示的に保存することで冗長性を抑制し、MLLM が堅牢な複数粒度のビデオ推論を実行できるようにします。 3 つの MLLM を使用した 4 つの一般的な LongVideoQA ベンチマークの評価では、高効率で最先端のゼロショット パフォーマンスが実証されました。特に、ReMem を使用した LLaVA-Video は、LVBench で 54.5% (+12.3%)、LongVideoBench で 67.1% (+8.2%) に達しています。
原文 (English)
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding. To accommodate this constraint, models typically resort to keyframe selection. However, uniform sampling or static query-guided selection often overlooks critical temporal context, failing to adapt to the varying query temporal granularities. In this paper, we propose ReMem, a temporal granularity-adaptive keyframe selection framework for training-free LongVideoQA. ReMem introduces a dual-level memory-augmented adaptation. At the query level, Memory-Driven Question Parsing leverages LLM long-term memory to decode question temporal granularity and extract semantic entities. At the video level, Synergistic Dual-Semantic Frame Alignment exploits intrinsic structural memory to align frames with query semantics, guiding Structure-Aware Dynamic Frame Routing to cluster events and optimally distribute sampling budgets. By explicitly preserving temporal information with memory mechanisms, ReMem suppresses redundancy and empowers MLLMs to perform robust multi-granular video reasoning. Evaluations across four popular LongVideoQA benchmarks using three MLLMs demonstrate highly efficient, state-of-the-art zero-shot performance; notably, LLaVA-Video with ReMem reaches 54.5% (+12.3%) on LVBench and 67.1% (+8.2%) on LongVideoBench.
最短が最も安全ではない場合: 高齢者に優しい歩行者ルートへのデザインサイエンスのアプローチ
高齢者の自立した移動は、家の外への参加、幸福と健康を可能にしますが、歩行者ナビゲーション システムは依然として距離や時間を主に最適化しており、高齢者の歩行に関する意思決定を形作る障壁、安全基準、支援インフラを見落とすことがよくあります。私たちは、実際のモビリティの制約を規範的な設計知識に変換する、階層的なデザインサイエンス研究を通じて開発された、高齢者に優しい歩行者ルートの成果物を紹介します。 11 件の半構造化インタビューに基づいて、バリアを意識し、アメニティに配慮したルーティングと実行関連の説明のための初期の設計要件 (DR) と設計原則 (DP) を導き出します。これらを、アメニティ (ベンチ、トイレ、シェルター) と高さデータを豊富に備えた OpenStreetMap 歩行者ネットワークでインスタンス化し、構成可能なコストと説明ペイロードを備えた A* ベースのルーティング エンジンを実装しました。野外ベースの歩行研究では、14 人の高齢者が人工物によって生成されたルートとベースラインを比較し、評価と定性的なフィードバックを提供しました。全体的には高齢者に優しいルートが好まれました。さらに、テーマ分析により、インフラストラクチャのメンテナンス、季節条件、交通エクスポージャ、および社会的状況がルートの受け入れを形成していることが示されました。私たちはこれらの洞察を洗練された DR と DP に統合し、状況を認識したハザードモデリング、複数ルートの透明性、ランドマークに基づいた説明、社会的状況の敏感性、段階に応じた情報を強調します。私たちの貢献は、高齢者に優しい歩行者ナビゲーション システムを開発する実務者に実用的なガイダンスを提供します。
原文 (English)
When Shortest Isn't Safest: A Design Science Approach to Senior-Friendly Pedestrian Routing
Older adults' independent mobility enables out-of-home participation, well-being and health, yet pedestrian navigation systems still optimize primarily for distance or time, often overlooking barriers, safety thresholds, and supportive infrastructure that shape late-life walking decisions. We present a senior-friendly pedestrian routing artefact developed through echeloned Design Science Re-search, translating lived mobility constraints into prescriptive design knowledge. Based on 11 semi-structured interviews, we derive initial Design Requirements (DRs) and Design Principles (DPs) for barrier-aware, amenity-sensitive routing and execution-relevant explanations. We instantiate these in an OpenStreetMap pedestrian network enriched with amenities (benches, toilets, and shelters) and height data, and implemented an A*-based routing engine with configurable costs and explanation payloads. In a field-based walking study, 14 older adults com-pared artefact-generated routes with baselines and provided ratings and qualitative feedback; the senior-friendly route was preferred overall. Thematic analysis further showed that infrastructure maintenance, seasonal conditions, traffic exposure, and social context shape route acceptance. We synthesize these insights into refined DRs and DPs emphasizing context-aware hazard modeling, multi-route transparency, landmark-grounded explanations, social-context sensitivity, and stage-appropriate information. Our contributions provide actionable guidance for practitioners developing senior-friendly pedestrian navigation systems.
RRS-10K: 希少なリモート センシング画像解釈のためのマルチタスク視覚言語モデル ベンチマーク
ビジョン言語モデル (VLM) は、一般的なリモート センシング タスクで優れたパフォーマンスを達成しました。ただし、既存のベンチマークは一般的な都市や田舎の画像が大半を占めているため、まれなシーンに対するベンチマークの機能はまだ十分に理解されていません。このギャップに対処するために、希少なリモート センシング画像判読のベンチマークである RRS-10K を紹介します。 RRS-10K には、包括的な評価を目的とした 10,738 枚の軍事関連のリモート センシング画像と、対応する複数形式の質問と回答のペアが含まれています。すべての画像は直接の情報源から収集され、知覚、推論、堅牢性をカバーする 3 つの能力の次元、6 つのサブ次元、および 20 のリーフ タスクに編成されています。多肢選択問題の品質を向上させるために、ベンチマーク構築中に類似性に基づくディストラクタ フィルタリング戦略 (SDFS) を導入します。さらに、52 の代表的なモデルを評価し、現在の VLM はまれなリモート センシング画像解釈では中程度のゼロショット パフォーマンスしか達成できず、視覚的グラウンディング、参照セグメンテーション、および複雑な意味論的推論タスクに明らかな弱点があることを示します。 RRS-10K は、ロングテール リモート センシングの解釈における故障モードの系統的な分析を可能にし、より信頼性の高いリモート センシング VLM を開発するためのガイダンスを提供します。
原文 (English)
RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
Vision-language models (VLMs) have achieved strong performance on general remote sensing tasks. However, their capability for rare scenes remains insufficiently understood, because existing benchmarks are dominated by common urban and rural imagery. To address this gap, we present RRS-10K, a benchmark for rare remote sensing image interpretation. RRS-10K contains 10,738 military-related remote sensing images and corresponding multiple format question-answer pairs for comprehensive evaluation. All of the images are collected from first-hand sources and organized into three capability dimensions, six sub-dimensions, and 20 leaf tasks, covering perception, reasoning, and robustness. To improve the quality of multiple-choice questions, we introduce a similarity-based distractor filtering strategy (SDFS) during benchmark construction. We further evaluate 52 representative models and show that current VLMs achieve only moderate zero-shot performance on rare remote sensing image interpretation, with clear weaknesses in visual grounding, referring segmentation, and complex semantic reasoning tasks. RRS-10K enables systematic analysis of failure modes in long-tail remote sensing interpretation and provides guidance for developing more reliable remote sensing VLMs.
Aletheia: リソースの少ない医療環境における鑑別診断のためのオフラインファーストの臨床意思決定支援システム
サハラ以南のアフリカでは、臨床専門知識へのアクセスが依然として厳しく制限されており、地方では医師と患者の比率が 1:25,000 を下回ることもあります。既存の AI 支援診断ツールは主に信頼性の高いインターネット接続と高仕様のハードウェアを必要とするため、地方の病院や保健センターの最前線の医療従事者にとっては実用的ではありません。この論文では、サハラ以南アフリカ全域の低リソース医療環境向けに設計されたオフラインファーストの臨床意思決定支援システムである Aletheia について紹介します。 Aletheia は Qwen2.5-3B-Instruct に基づいて構築されており、東アフリカで有病率が高い 50 の病状にわたる 27,000 の臨床推論サンプルの厳選されたデータセットに対して量子化低ランク適応 (QLoRA) を使用して微調整されています。評価の結果、10の代表的な臨床症例カテゴリー全体で、トップ1の診断精度が80.0%、トップ3の精度が100.0%、BERTScore-F1が0.909、METEORが0.467であることが実証されました。このシステムは、0.275 の予想キャリブレーション誤差 (ECE) を達成し、アフリカ ディープ テック チャレンジ 2026 (ADTC 2026) のメモリ バジェット制約である 7 168 MB をクリアし、標準化されたベンチマーク ラップトップで約 3 630 MB のピーク推論 RAM を達成しました。これらの結果は、クラウド インフラストラクチャを使用せずに、リソースに制約のある環境でプライマリ ケア レベルで大規模な言語モデルに基づく臨床推論を展開する実現可能性を示しています。
原文 (English)
Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
Access to specialist clinical expertise remains severely limited across sub-Saharan Africa, where physician-to-patient ratios can fall below 1:25,000 in rural settings. Existing AI-assisted diagnostic tools predominantly require reliable internet connectivity and high-specification hardware, rendering them impractical for frontline healthcare workers in district hospitals and health centres. This paper presents Aletheia, an offline-first clinical decision support system designed for low-resource healthcare contexts across sub-Saharan Africa. Aletheia is built upon Qwen2.5-3B-Instruct, fine-tuned using Quantised Low- Rank Adaptation (QLoRA) on a curated dataset of 27,000 clinical reasoning samples spanning 50 disease conditions with elevated prevalence in East Africa. Evaluation demonstrates a Top-1 diagnostic accuracy of 80.0%, Top-3 accuracy of 100.0%, BERTScore-F1 of 0.909, and METEOR of 0.467 across ten representative clinical case categories. The system achieves an Expected Calibration Error (ECE) of 0.275 and passes the Africa Deep Tech Challenge 2026 (ADTC 2026) memory budget constraint of 7 168 MB, achieving a peak inference RAM of approximately 3 630 MB on the standardised benchmark laptop. These results demonstrate the feasibility of deploying large language model-based clinical reasoning at the primary care level in resource-constrained settings without cloud infrastructure.
AdaKP: 推論指向の強化学習のためのオンライン適応ナレッジポイント選択
検証可能な報酬を伴う強化学習は、大規模な言語モデルで推論を引き出すための強力なパラダイムですが、競技レベルの数学では深刻な報酬の少なさに悩まされます。一般的な解決策は、アトミック ナレッジ ポイント (KP) (ゴールド ソリューションから抽出された短い自然言語ヒント) をプロンプトに挿入します。ただし、既存の方法では、この選択をオフラインで一度修正するか、注入されるテキストのモノリシックな量を単純にスケールして、最も有益な選択軸、つまりアトミック KP のどのサブセットをいつ注入するかはそのままにしておくかのどちらかです。 RL トレーニング中に各問題の KP サブセットを再選択するオンライン セレクターである AdaKP を紹介します。その中心となるのは、高価なロールアウトベースの推定の代わりに、それが誘発する次のトークンのエントロピーの減少によって KP をスコアリングするエントロピー プロキシです。つまり、打ち切りバイアスに証明可能な限界を持つ単一の安価なフォワード パスです。 3 つの軽量メカニズムにより、この信号はオンラインで使用可能になります。ステップごとのノイズを吸収するモメンタム スムーザー、探索を維持しながら弱い KP を除去するリタイアおよび復活マネージャー、再評価を初期トレーニングにフロントローディングする適応スケジューラーです。 AdaKP はさらに、コストのかかる実行が開始される前に、リーブ ワンアウト グランド トゥルースに対してプロキシを認証するプリフライト検証ゲートを提供し、メソッド レベルのリスクを改ざん可能なチェックに変えます。オプティマイザの変更を行わずに標準的な DAPO+GRPO トレーナーの完全な加算フォークとして実現された AdaKP は、無視できる追加コストで 8 つの競技数学ベンチマークすべてで強力な静的選択ベースラインを改善し、オンラインで検証された KP サブセット選択を、推論指向の強化学習のための実用的でまだ検討されていない軸として位置付けます。
原文 (English)
AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning
Reinforcement learning with verifiable rewards is a powerful paradigm for eliciting reasoning in large language models, yet it suffers from severe reward sparsity on competition-level mathematics. A common remedy injects atomic knowledge points (KPs) - short natural-language hints distilled from gold solutions - into the prompt. Existing methods, however, either fix this selection once offline or merely scale the monolithic quantity of injected text, leaving untouched the most informative axis of choice: which subset of atomic KPs to inject, and when. We introduce AdaKP, an online selector that re-chooses each problem's KP subset over the course of RL training. At its core is an entropy proxy that scores a KP by the reduction in next-token entropy it induces - a single inexpensive forward pass, with a provable bound on its truncation bias - in place of expensive rollout-based estimation. Three lightweight mechanisms make this signal usable online: a momentum smoother that absorbs per-step noise, a retirement-and-revival manager that prunes weak KPs while preserving exploration, and an adaptive scheduler that front-loads re-evaluations into early training. AdaKP further contributes a pre-flight validation gate that certifies the proxy against a leave-one-out ground truth before any expensive run is launched, turning method-level risk into a falsifiable check. Realized as a fully additive fork of a standard DAPO+GRPO trainer with no optimizer changes, AdaKP improves over a strong static-selection baseline on all eight competition-mathematics benchmarks at negligible added cost, positioning online, validated KP-subset selection as a practical and as-yet under-explored axis for reasoning-oriented reinforcement learning.
MusiChat: 音楽作成のための Vibe 作曲
AI 音楽生成の最近の進歩により、ユーザーは自然言語のプロンプトから完全な音楽作品を作成できるようになりました。しかし、既存のシステムのほとんどは即時再生パラダイムに従っており、ユーザーは既存の音楽アイデアを直接進化させるのではなく、楽曲を繰り返し再作成する必要があるため、反復的な改良が困難になっています。 MusiChat は、自然言語の対話と反復的な改良を通じて人間と AI の共同音楽作成を可能にする、会話型ヴァイブ作曲システムです。 MusiChat の中核となるのは、歌詞に合わせた音楽構造の生成と表現力豊かな表面の実現を分離する、階層的な制御可能な音楽生成フレームワークであり、柔軟なスタイルの変換と構造を保持した編集を可能にします。このシステムは、インタラクション全体でアクティブな作曲状態とユーザー履歴を維持するメモリ拡張アーキテクチャを通じて、大規模な言語モデルとハイブリッド シンボリック ミュージック エンジンを統合します。ハイブリッド インテント ルーティング メカニズムにより、正確な音楽編集と無制限のクリエイティブ リクエストの両方を効率的に解釈できます。 MusiChat は、楽曲を最初から再生成するのではなく、関連する音楽構造とユーザーの意図を維持しながら、進化する音楽成果物を段階的に変換します。私たちは客観的な分析と人間による研究を通じて MusiChat を評価し、シングルターンとマルチターンのインタラクションについてそれぞれ 95.31% と 100% の精度を達成し、メロディーの自然さについては 2:1、音楽の品質については好きと嫌いの比率が 3:1 という結果を得ました。私たちの結果は、MusiChat が対話型インターフェイスを介した一貫したマルチターン音楽オーサリングとインタラクティブな人間と AI の共同制作をサポートしていることを示しています。
原文 (English)
MusiChat: Vibe Composing for Music Creation
Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most existing systems follow a prompt-and-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas. We present MusiChat, a conversational vibe composing system that enables collaborative human-AI music creation through natural-language interaction and iterative refinement. At the core of MusiChat is a hierarchical controllable music generation framework that separates lyric-aligned musical structure generation from expressive surface realization, allowing flexible stylistic transformations and structure-preserving edits. The system integrates a large language model with a hybrid symbolic music engine through a memory-augmented architecture that maintains the active composition state and user history across interactions. A hybrid intent-routing mechanism further enables efficient interpretation of both precise musical edits and open-ended creative requests. Rather than regenerating compositions from scratch, MusiChat incrementally transforms an evolving musical artifact while preserving relevant musical structure and user intent. We evaluate MusiChat through objective analysis and human studies, achieving 95.31% and 100% accuracy for single- and multi-turn interactions, respectively, and obtaining like-to-dislike ratios of 2:1 for melody naturalness and 3:1 for musical quality. Our results demonstrate that MusiChat supports coherent multi-turn music authoring and interactive human-AI co-creation through a conversational interface.
セマンティック ID を理解する: 生成的推奨におけるアイテム表現からアイテム選択まで
セマンティック ID (SID) は現在、生成的推奨の中心的なコンポーネントです。現在の SID ベースのシステムは、同じトークン シーケンスに 3 つの役割を割り当てます。共有プレフィックスは関連するアイテムを整理することを目的としており、完全な SID は個々のアイテムを識別し、生成された各トークンによって返されるアイテムを絞り込みます。私たちは、アイテムのエンコードと SID 構築から自己回帰生成と最終的な推奨に至るまで、SID を体系的に調査します。 SID の構築によって項目の表現がどのように変更されるか、またそれらの変更が生成にどのような影響を与えるかを調べます。 3 つの Amazon ドメインと 8 つの SID 構造にわたって、SID 近傍はエンコーダーの最近傍 10 個のうち平均 32.2% しか回復しません。代替アイテムの説明は、管理されたケースの 99.57% で対応するアイテムを最初に取得しますが、正確な SID の 38.4% は変更されます。これらの結果は、SID が広範な構成を保持しているものの、エンコーダーの細かいローカル構造の多くが失われている一方、その正確なトークンは項目の意味だけでは決定されないことを示しています。この損失は生成中に重大なものになります。最後のセマンティック トークンの後、TIGER は、SID フィルタリング前に妥当な推奨事項であった保留ターゲットの 29.9% のみを保持します。これらの発見に動機付けられて、我々は、ビーム検索がそれらを破棄する前に、ユーザー固有の項目ランキングが対応する SID プレフィックスをサポートできるようにする軽量の推論時間手法である項目サポート デコーディング (ISD) を提案します。同じランキングにより、生成されたアイテムが順序付けされます。 ISD では、パラメーターを追加したり、SID コンストラクターやデコーダーを再トレーニングしたりする必要はありません。我々は、評価されたすべての設定において、ISD が対応する SID バックボーンよりも NDCG@10 を向上させ、相対的に最大 31.2% の向上をもたらすことを経験的に示しています。私たちの結果は、SID は有用な大まかなアイテムの編成を提供しますが、その細かい境界だけで、生成中にどのアイテムが利用可能なままになるかを決定するわけではないことを示しています。
原文 (English)
Understanding Semantic IDs: From Item Representation to Item Selection in Generative Recommendation
Semantic IDs (SIDs) are now a central component of generative recommendation. Current SID-based systems assign three roles to the same token sequence. Shared prefixes are intended to organize related items, the complete SID identifies an individual item, and each generated token narrows the items that can still be returned. We systematically investigate SIDs from item encoding and SID construction to autoregressive generation and final recommendation. We examine how SID construction changes item representations and how those changes affect generation. Across three Amazon domains and eight SID constructions, SID neighborhoods recover only 32.2% of the encoder's ten nearest neighbors on average. Alternative item descriptions still retrieve the corresponding item first in 99.57% of controlled cases, yet change 38.4% of exact SIDs. These results show that SIDs retain broad organization but lose much of the encoder's fine local structure, while their exact tokens are not determined by item meaning alone. This loss becomes consequential during generation. After the final semantic token, TIGER retains only 29.9% of held-out targets that were plausible recommendations before SID filtering. Motivated by these findings, we propose Item-Supported Decoding (ISD), a lightweight inference-time method that allows a user-specific item ranking to support corresponding SID prefixes before beam search discards them. The same ranking then orders the generated items. ISD requires no additional parameters or retraining of the SID constructor or decoder. We empirically show that ISD improves NDCG@10 over the corresponding SID backbone in every evaluated setting, with relative gains of up to 31.2%. Our results show that SIDs provide useful coarse item organization, but their fine boundaries should not alone determine which items remain available during generation.
微分可能な D-vine コピュラによる局所的な異常検出
Vine コピュラは、二変量ペア コピュラへの階層分解を通じて複雑な多変量分布をモデル化するための柔軟なフレームワークを提供します。 D-vine をフィッティングするには、さまざまな依存パターンをエンコードする候補のセットからコピュラ ファミリと各ペア コピュラのパラメーター構成を選択する必要があります。変数と候補ファミリーの数が増加するにつれて、可能な構成の数は組み合わせ的に増加します。既存のフィッティング手順は、連続した貪欲な決定を通じてこの課題に対処し、各ステップで単一の局所的に最適なファミリーにコミットし、より良いグローバル フィットをもたらす構成を潜在的に破棄します。この制限を克服するために、完全微分可能な実装によって可能になる勾配ベースの最尤推定と、フィッティング プロセス全体を通じて複数の競合する D-vine 構成を維持するビーム探索戦略を組み合わせた新しい推定フレームワークを提案します。これにより、計算上扱いやすい状態を保ちながら、構成空間をより広範囲に探索できるようになります。適合した D-vine に基づいて、階層分解を利用してグローバルな異常スコアとエッジレベルの説明の両方を生成する局所的な異常検出フレームワークを導入します。統計的保証はモンドリアンの等角予測によって提供され、ペアコピュラ構造により異常を特定の変数関係に局在化することができます。提案されたフレームワークをベンチマークと現実世界のデータセットの両方で評価し、不確実性の定量化による解釈可能な異常検出に対するその有効性を実証します。
原文 (English)
Localized Anomaly Detection via Differentiable D-vine Copulas
Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivariate pair-copulas. Fitting a D-vine requires selecting a copula family and parameter configuration for each pair-copula from a set of candidates encoding different dependence patterns. As the number of variables and candidate families increases, the number of possible configurations grows combinatorially. Existing fitting procedures address this challenge through sequential greedy decisions, committing to a single locally optimal family at each step and potentially discarding configurations that would yield a better global fit. To overcome this limitation, we propose a novel estimation framework that combines gradient-based maximum likelihood estimation, enabled by our fully differentiable implementation, with a beam-search strategy that maintains multiple competing D-vine configurations throughout the fitting process. This allows a broader exploration of the configuration space while remaining computationally tractable. Building on the fitted D-vine, we introduce a localized anomaly detection framework that exploits the hierarchical decomposition to produce both global anomaly scores and edge-level explanations. Statistical guarantees are provided through Mondrian conformal prediction, while the pair-copula structure enables the localization of anomalies to specific variable relationships. We evaluate the proposed framework on both benchmark and real-world datasets, demonstrating its effectiveness for interpretable anomaly detection with uncertainty quantification.
チャートでサポートされるか、モデルで提供されるか?アクセシブルな視覚化のための MLLM 生成のクレームの検査
マルチモーダル大規模言語モデル (MLLM) は、視覚化パターンを外部の原因、結果、ドメイン知識に結び付けることができますが、これらの解釈の証拠的根拠が不明瞭であることがよくあります。 4 つのソース、3 つの MLLM、および画像へのアクセス、ソース固有のアクセス可能なチャート コンテキスト、および非保持コンテキスト フレーミングへのアクセスを変える 4 つの入力条件からの 102 のビジュアライゼーションの探索的研究を紹介します。 1,224 個の記述にわたって、モデルに起因する DIRECT、DERIVED、および SPECULATIVE ラベルを分析し、数値一致の自動監査を実施します。アクセシブルなチャート コンテキストにより、Gemini と GPT が DIRECT 主張に移行し、一部のモデルの数値一致が改善されました。完全なコンテキストに画像を追加しても、一貫した数値上の利点は得られず、コンテキストを差し控えたプロンプトによっても確実に慎重な言葉遣いが増加することはありませんでした。プロンプトで定義された現実世界の重要性セクションは、主に推測的なままでした。これらの結果は、提供された証拠によって裏付けられた主張とモデルによって提供された解釈を区別する、アクセス可能な記述システムの動機付けとなります。
原文 (English)
Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization
Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study of 102 visualizations from four sources, three MLLMs, and four input conditions that vary access to the image, source-specific accessible chart context, and withheld-context framing. Across 1,224 descriptions, we analyze model-attributed DIRECT, DERIVED, and SPECULATIVE labels and conduct an automated audit of numeric agreement. Accessible chart context shifted Gemini and GPT toward DIRECT claims and improved numeric agreement for some models. Adding the image to the full context did not yield a consistent numeric benefit, and the withheld-context prompt did not reliably increase cautious language. The prompt-defined Real-World Significance section remained predominantly SPECULATIVE. These results motivate accessible description systems that distinguish claims supported by supplied evidence from model-supplied interpretation
SAFAARI: 広告主の応答インテリジェンスを加速するためのスキーマ認識フレームワーク
カスタマー サポート システムの進化は、エージェント チャットボットによって急速に進んでいますが、これらのシステムは、事前定義された API エンドポイントなしで企業データにアクセスする場合、重大な制限に直面しています。このペーパーでは、特殊なコンテンツ、メタデータ、およびオーケストレーション エージェントを通じて、自然言語から SQL (NL から SQL) システムへのスキーマ リンクの重大なボトルネックに対処するマルチエージェント フレームワークである SAFAARI (Schema-Aware Framework for Accelerated Advertiser Response Intelligence) について説明します。また、一貫性のない結果にペナルティを与えながらシステムのパフォーマンスを総合的に評価する新しい複合メトリクスである SEAL (Schema Evaluation and Accuracy in Language-to-SQL) も紹介します。 5 つの機能セット構成による体系的な実験を通じて、SAFAARI は 81.66% の SEAL スコア (ベースラインより 6.65% 向上) を達成し、データポイントの精度 (5.51%) とスキーマリンクの精度 (4.69%) が顕著に向上しました。このフレームワークの有効性は、ドメインの専門家による人間参加型の評価を通じて検証され、さまざまなサポート ドメインにわたる適応性が証明されています。スキーマのリンクとクエリ生成という労働集約的なプロセスを自動化することで、当社のフレームワークは高精度を維持しながら開発時間を 8 分の 1 に短縮することを実証しています。このソリューションは API 開発を合理化し、セルフサービス機能を強化し、特に複雑なデータ エコシステムを備えたカスタマー サポート企業に恩恵をもたらします。
原文 (English)
SAFAARI: Schema-Aware Framework for Accelerated Advertiser Response Intelligence
The evolution of customer support systems is rapidly advancing with agentic chatbots, yet these systems face significant limitations when accessing enterprise data without predefined API endpoints. This paper presents SAFAARI (Schema-Aware Framework for Accelerated Advertiser Response Intelligence), a multi-agent framework that addresses the critical bottleneck of schema linking in Natural Language to SQL (NL-to-SQL) systems through specialized content, metadata, and orchestration agents. We also introduce SEAL (Schema Evaluation and Accuracy in Language-to-SQL), a novel composite metric that holistically evaluates system performance while penalizing inconsistent results. Through systematic experimentation with five feature set configurations, SAFAARI achieves an 81.66% SEAL score (6.65% improvement over baseline), with notable gains in datapoint accuracy (5.51%) and schema-linking precision (4.69%). The framework's effectiveness is validated through human-in-the-loop evaluation with domain experts, which proves its adaptability across diverse support domains. By automating the labor-intensive process of schema linking and query generation, our framework demonstrates 8x reduction in development time while maintaining high accuracy. The solution streamlines API development and enhances self-service capabilities, particularly benefiting customer support enterprises with complex data ecosystems.
CogEEGAgent: グラウンデッド実行と選択認識型検証による自律的な認知 EEG 解析に向けて
認知研究における脳波 (EEG) 分析には専門知識が必要であり、コントラスト、チャネル、時間窓、統計的テストに関して多くの防御可能な選択肢が含まれます。 LLM エージェントは、さまざまな自然言語の質問を分析の選択肢に変換し、自動化のための柔軟なインターフェイスを提供します。しかし、流暢なレポートだけでは、エージェントが要求された分析を選択したこと、または適応検索とは独立して確認的主張を評価したことを証明することはできません。 MNE-Python に基づいた認知脳波分析エージェントである CogEEGAgent を紹介します。 EEG に特化した科学的ハーネスにより、意味論と科学的権威が分離されます。 LLM は意図を解釈し、登録された分析を提案します。一方、決定論的コンポーネントは型指定された契約を検証し、確認アクセスを制御し、証拠に拘束されたリリースを許可します。事前に指定されたルーティング ベンチマークでは、CogEEGAgent は、一致した決定論的ルーターよりも正確に言語を登録された分析にマッピングします。また、一致したプリフライトにより、必要な場合はいつでも両方のシステムが棄権します。外部でモデルが作成された、結果にブラインドなキャンペーンでは、完全なシステムが、参加者間で不一致の確認を伴うサポートされている分析をリリースし、事前に指定された機能の危険性とライフサイクル再利用リクエストをブロックします。ポリシーのストレス テストでは、保留された確認により、修正されていない適応検索による誤検知が抑制されることが示されています。これらの研究を総合すると、認知 EEG ワークフローの制限された自律性と監査可能な自動化フレームワークが確立されます。より広範に、科学エージェントが柔軟な言語理解と推論と解放のフェイルクローズ制御をどのように組み合わせることができるかを示しています。
原文 (English)
CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification
Electroencephalography (EEG) analysis in cognitive studies requires specialized expertise and involves many defensible choices over contrasts, channels, time windows, and statistical tests. LLM agents can translate varied natural-language questions into analysis choices, offering a flexible interface for automation. Yet fluent reports alone cannot establish that an agent selected the requested analysis or evaluated a confirmatory claim independently of adaptive search. We present CogEEGAgent, a cognitive-EEG analysis agent grounded in MNE-Python. Its EEG-specific scientific harness separates semantic from scientific authority. The LLM interprets intent and proposes registered analyses, while deterministic components validate typed contracts, control confirmation access, and authorize evidence-bound release. On a prespecified routing benchmark, CogEEGAgent maps language to registered analyses more accurately than a matched deterministic router, while matched preflight makes both systems abstain whenever required. In an externally model-authored, outcome-blind campaign, the complete system releases supported analyses with participant-disjoint confirmation and blocks prespecified capability hazards and lifecycle-reuse requests. Policy stress testing shows that held-out confirmation curbs false positives from uncorrected adaptive search. Together, these studies establish bounded autonomy and an auditable automation framework for cognitive-EEG workflows. More broadly, they show how scientific agents can combine flexible language understanding with fail-closed control over inference and release.
会話型 AI の心理的影響: 危害を軽減し、幸福を促進するための研究と設計の方向性
会話型 AI システムが日常生活にますます統合されるにつれて、ユーザーの幸福に対する潜在的な影響には継続的な注意が必要です。消費者向けのジェネラリスト モデルは、情報へのアクセス、学習、生産性、内省、交友関係の向上などの利点を提供できる一方で、感情的なもつれ、不健康な依存、心理的脆弱性の増幅などのリスクももたらします。 AI チャットボットの動作に関する先行研究と経験的観察に基づいて、潜在的な心理的危害を軽減し、ユーザーの幸福をサポートできる方法で汎用 AI システムの動作を導くための一連の意欲的な方向性を提案します。私たちは、AI チャットボットの使用による長期的な影響を体系的に評価することの難しさを認識しており、これらの方向性を、一般的なインタラクション、ロールプレイング シナリオ、および心理的サポートを提供すると特徴づけられる状況全体にわたって、AI の行動がユーザーにどのような影響を与えるかを研究するための仮説として組み立てます。提案された方向性の中には、既存の研究や専門家の洞察によって裏付けられているものもありますが、未解決の疑問やより深い研究が必要な領域を特定しているものもあります。私たちは、この定式化とこれらの仮説が、ユーザーの心理的ニーズにより適切に対応し、ユーザーの幸福を促進することを目的としたインタラクティブ デザイン アプローチのさらなる議論、実証的調査、探求を促進することを願っています。
原文 (English)
Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing attention. While consumer-facing generalist models can provide benefits, including improved access to information, learning, productivity, self-reflection, and companionship, they also introduce risks, such as emotional entanglement, unhealthy dependence, and the amplification of psychological vulnerabilities. Drawing on prior research and empirical observations of AI chatbot behavior, we propose a set of aspirational directions for guiding the behavior of general-purpose AI systems in ways that may reduce potential psychological harms and support user well-being. We acknowledge the difficulty of systematically assessing the long-term impacts of AI chatbot use and frame these directions as hypotheses for studying how AI behavior may influence users across general interactions, role-playing scenarios, and contexts that could be characterized as providing psychological support. While some proposed directions are supported by existing research and expert insights, others identify open questions and areas requiring deeper study. We hope that this formulation and these hypotheses encourage further discussion, empirical investigation, and exploration of interactive design approaches aimed at better accommodating users' psychological needs and promoting their well-being.
類似したモデルは異なる学習を行う: 最終ウィンドウの事前トレーニングが SFT を超えたトレーニング後の形状を形成する
開発者は、モデルのチェックポイントがどのように動作するかによって判断します。教師あり微調整 (SFT) 後、関連するベンチマーク間でほぼ同じパフォーマンスを示す 2 つのチェックポイントは交換可能として扱われ、次の調整段階 (通常はプリファレンスの最適化) に同様に準備が整います。この判断がトレーニング前の痕跡を見逃しているかどうかを尋ねます。この違いは、SFT 後のベンチマークでは明らかになっていないものの、その後のトレーニングに各チェックポイントがどのように反応するかを決定します。それを確認するために、事前トレーニングの最後のウィンドウ、つまり命令調整の前にトレーニングされた最後のデータに対して制御された実験を実行します。 6 つのブランチは部分的に事前トレーニングされた 1 つのチェックポイントから分岐しており、このウィンドウのみが異なります: 5 億トークン (その前のトークンの 0.1% ~ 1%)。各ブランチは、単一のデータ ソース (一般的な Web テキスト、フィルタリングされた Web テキスト、規範的談話、安全テキスト、数学テキスト、合成教育テキスト) でウィンドウをトレーニングします。 SFT とポストトレーニングは同一になります。 SFT の後、ブランチは、命令追従、拒否、および能力に関して約 1 ポイント内でほぼ同一に動作しますが、同じポストトレーニングによって、直接優先最適化更新と検証可能な報酬を伴う強化学習更新の両方の下で、ブランチが非常に異なるエンドポイントに運ばれます。この逸脱は、有害なリクエストの拒否を通じて測定されます。ポストトレーニングが始まると、安全テキスト ブランチは Web テキスト ブランチと同じくらい拒否しますが、最後までに拒否される量ははるかに少なくなります。他の 4 つのブランチは保護をほとんどまたはまったく得られないため、その効果はウィンドウに含まれる内容に応じて選択されます。この保護では、安全テキストが事前トレーニングの最初ではなく最後に到着する必要があり、これは 2 番目のモデル ファミリで再現されます。モデルが最後の形状でどのように事前トレーニングされているか、位置合わせにどのように反応するか。したがって、チェックポイントは SFT 後の動作だけで評価されるべきではなく、最後にトレーニングされた内容も併せて報告する必要があります。
原文 (English)
Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant benchmarks are treated as interchangeable, equally ready for the next alignment stage, typically preference optimization. We ask whether this judgment misses a pretraining imprint: a difference that no post-SFT benchmark reveals, yet that decides how each checkpoint responds to further training. To find out, we run a controlled experiment on the final window of pretraining, the last data trained on before instruction tuning. Six branches fork from one partially pretrained checkpoint and differ only in this window: 500 million tokens, 0.1% to 1% of the tokens that precede it. Each branch trains its window on a single data source: generic web text, filtered web text, normative discourse, safety text, mathematical text, or synthetic educational text. SFT and post-training are then identical. After SFT the branches behave near-identically, within about one point on instruction following, refusal, and capability, yet the same post-training carries them to very different endpoints, under both a direct preference optimization update and a reinforcement learning update with a verifiable reward. We measure this deviation through refusal of harmful requests: when post-training begins the safety text branch refuses no more than the web text branch, yet by the end it has lost far less of its refusal. The other four branches gain little or no protection, so the effect is selective to what the window contained. The protection requires the safety text to arrive last rather than earlier in pretraining, and it reproduces on a second model family. What a model is pretrained on last shapes how it reacts to alignment. Therefore, a checkpoint should not be evaluated by its post-SFT behavior alone, and what it was trained on last should be reported with it.
AI エージェントにおける長いコンテキスト ウィンドウ制御のためのアドレス指定可能なリコール コンパクション
長期的な LLM エージェントは推論トレース、アクション、ツールの観察を蓄積し、最終的にはモデルの固定コンテキスト ウィンドウを超える可能性があります。既存の圧縮方法は、以前の情報を破棄、要約、または取得することでこの制限に対処していますが、タスクに不可欠な詳細が削除されたり、それらを確実に回復できない可能性があります。私たちは、アクティブ コンテキストのプレゼンテーションからアーカイブ ストレージを分離するコンテキスト管理フレームワークである ARC (Addressable Recall Compaction) を提案します。 ARC は、ツールの観察を追加専用の ID アドレス指定可能なログに保存し、圧縮が必要な場合には古い観察をコンパクトな引用に置き換えます。その後、エージェントはこれらの識別子を使用して、対応するツールを再実行したり、類似性に基づく検索のみに依存したりすることなく、保存されたコンテンツを要求できます。 16k コンテキスト ウィンドウを持つ Qwen3-8B と 32k コンテキスト ウィンドウを持つ Qwen3-32B を使用して ARC を評価します。 Needle-in-a-Haystack の評価では、ARC は平均正確回答精度 99.40% を達成しました。これに対し、当社の評価で最も優れたベースラインの精度は 88.12% でした。 ARC は、ハードウェア コスト モデルに基づいて、推定サービス時間と HBM トラフィックも削減します。 LongBench-v2 Hard サブセットでは、ARC の平均精度は 29.97% ですが、最もパフォーマンスの高いベースラインでは 28.25% でした。これらの結果は、明示的なアドレスベースの呼び出しにより、テストされた設定で評価されたコンテキスト管理ベースラインと比較して、情報の保持と提供効率が向上する可能性があることを示しています。
原文 (English)
Addressable Recall Compaction for Long Context-Window Control in AI Agents
Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window. Existing compaction methods address this limitation by discarding, summarizing, or retrieving earlier information, but they may remove task-critical details or fail to recover them reliably. We propose ARC (Addressable Recall Compaction), a context-management framework that separates archival storage from active-context presentation. ARC stores tool observations in an append-only, ID-addressable log and replaces older observations with compact citations when compaction is required. The agent can subsequently use these identifiers to request stored content without re-executing the corresponding tools or depending solely on similarity-based retrieval. We evaluate ARC using Qwen3-8B with a 16k context window and Qwen3-32B with a 32k context window. On the Needle-in-a-Haystack evaluation, ARC achieves an average exact-answer accuracy of 99.40%, compared with 88.12% for the best-performing baseline in our evaluation. ARC also reduces estimated serving time and HBM traffic under our hardware-cost model. On the LongBench-v2 Hard subset, ARC obtains an average accuracy of 29.97%, compared with 28.25% for the best-performing baseline. These results indicate that explicit, address-based recall can improve information retention and serving efficiency relative to the evaluated context-management baselines under the tested settings.
推奨者はどのくらいの頻度で LLM に電話をかける必要がありますか?値に重み付けされたルーティング、モニタリング、季節性の堅牢性
安価なヒューリスティックと高価な大規模言語モデル (LLM) の間のルーティングの決定は、通常、困難な問題として組み立てられます。つまり、困難なケースを高価なパスに送信します。困難とビジネス価値は別個の軸であるため、この枠組みは不完全であると私たちは主張します。困難で安価な項目と困難で高価な項目のエラーのコストは同じではありません。私たちは、小売マーチャンダイジング パイプラインの完全合成シミュレーションである Value Router を紹介します。これは、地上の真実ではなく、推定難易度と推定価値のみを使用してアイテムをルーティングします。研究には 3 つの段階があります。まず、値に重み付けされたしきい値ルーターが、カテゴリのボリュームと値の間に逆相関がある合成カタログ上の難易度のみのランダムなベースラインと比較されます。価値の重み付けは、真の高価値アイテムの難易度のみのベースラインの再現率 (60%) と一致し、大幅に高い精度 (98.3% 対 94.3%) を達成します。第 2 に、意思決定ロガーとモニターは、集計メトリクスによって隠された故障モードを明らかにし、集計結果が品目ごとの区別ではなくカテゴリ間の差異によってほぼ完全に左右されることを示します。 3 番目に、シミュレートされたブラック フライデーの需要急増 (より価値の高いカテゴリーへのシフトによる 2.5 ボリューム) では、静的ルーター、季節調整ルーター、および 2 つの低速パス予算ポリシーを比較します。すべての結果は、実験者が定義したグラウンド トゥルースを使用した制御された合成シミュレーションからのものであり、検証された現実世界の主張ではなく、コストを意識したルーティング システムの設計原則を示しています。
原文 (English)
How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness
Routing decisions between a cheap heuristic and an expensive large language model (LLM) are typically framed as a difficulty problem: send the hard cases to the expensive path. We argue this framing is incomplete because difficulty and business value are distinct axes - a difficult cheap item and a difficult costly item do not have the same cost of error. We present Value Router, a fully synthetic simulation of a retail merchandising pipeline that routes items using only estimated difficulty and estimated value, never ground truth. The study has three stages. First, a value-weighted threshold router is compared with a difficulty-only and a random baseline on a synthetic catalog with an inverse correlation between category volume and value. Value-weighting matches the difficulty-only baseline's recall of true high-value items (60%) while achieving substantially higher precision (98.3% vs. 94.3%). Second, a decision logger and monitor expose a failure mode hidden by aggregate metrics showing that the aggregate result is driven almost entirely by between-category differences rather than per-item discrimination. Third, a simulated Black Friday demand surge (2.5 volume with a shift toward higher-value categories) compares a static router, a seasonally tuned router, and two slow-path budget policies. All results are from a controlled synthetic simulation with experimenter-defined ground truth and illustrate design principles for cost-aware routing systems rather than validated real-world claims.
エージェント オペレーティング システムに向けて - クラシック OS とクラウド OS からの教訓
プラットフォーム ソフトウェアの主要な波はすべて同じ弧をたどります。最初は競合するフレームワークとその場限りの実装を試し、その後、明確に定義されたセマンティクスを備えた安定した抽象化の小さなセットが明確になり、最後にそれらの抽象化をアプリケーションが移植可能なプラットフォームに統合します。 POSIX は従来のオペレーティング システムに対してこれを行っていました。 Kubernetes はクラウドのためにそれを行いました。エージェント AI システム (計画、ツールの使用、メモリの維持、および共同作業を行う自律的な LLM 主導のエージェント) は、現在、そのような第 3 の波の実験段階にあります。数十のフレームワークやプロトコルが登場しましたが、核となる抽象化が何であるか、またはそれらがどのような保証を提供するかについてコミュニティのコンセンサスは存在しません。そのコンセンサスがなければ、エージェント アプリケーションを移植可能に作成することはできず、プラットフォームを確実に構成することはできず、この分野はプロトタイプの展開を超えて前進することはできません。私たちは、前進する道は、以前のウェーブの方法論に従うことであると主張します。つまり、古典的な OS とクラウド OS のプリミティブを確率的で自然言語媒介の実行に拡張することによって新しいエージェントの抽象化を導き出し、それらのセマンティクスを正確に指定し、それらを中心に統合します。これは、POSIX と Kubernetes がそれぞれのウェーブを統合したのと同じです。
原文 (English)
Towards an Agent Operating System - Lessons from Classical and Cloud OS
Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc implementations, followed by the articulation of a small set of stable abstractions with well-defined semantics, and finally consolidation around those abstractions into a platform that applications can portably target. POSIX did this for classical operating systems; Kubernetes did it for the cloud. Agentic AI systems - autonomous, LLM-driven agents that plan, use tools, maintain memory, and collaborate - are currently in the experimentation phase of the third such wave. dozens of frameworks and protocols have emerged, but no community consensus exists on what the core abstractions are or what guarantees they carry. Without that consensus, agentic applications cannot be written portably, platforms cannot compose reliably, and the field cannot advance beyond prototype deployments. We argue that the path forward is to follow the prior-wave methodology: derive new agentic abstractions by extending classical OS and cloud OS primitives to stochastic, natural-language-mediated execution, specify their semantics precisely, and consolidate around them - just as POSIX and Kubernetes consolidated their respective waves.
PLATO: エージェントとタスクのオープン性のためのポインター学習器
オープン エージェント システム (OASYS) は、エージェントとタスクのセットが時間の経過とともに予期せず変化する現実の領域でますます普及しています。エージェントのオープン性 (AO) やタスクのオープン性 (TO) を含むこのようなオープン性は、通常、固定状態とアクション空間を前提とするマルチエージェント強化学習 (MARL) に対して根本的な課題を引き起こします。既存の方法は、開放性を部分的にしか扱っていません。パディングおよびマスキングのアプローチでは人為的な境界が導入されていますが、最近のグラフベースまたはハイパーグラフの方法では、開放性の一次元を扱いますが、依然として限定的な仮定に依存しています。この論文では、集中型トレーニングと分散型実行パラダイムの下でマルチエージェント近接ポリシー最適化でトレーニングされた、集中型グラフ ニューラル ネットワーク (GNN) クリティカルと組み合わせたポインター ネットワーク ベースのアクターである、エージェントとタスク オープン性のためのポインター学習者 (PLATO) を紹介します。ポインターベースのアクターは、現在のタスク セット上に直接分布を出力します。これは、マスキングや再トレーニングを行わずに、アクション スペースの変更を直接サポートします。私たちの GNN 批評家は、エージェントとタスクの相互作用を、タスクとエージェントの構成に応じて形状が変化するグラフとしてエンコードします。これらのコンポーネントを合わせて、既存のアプローチに制限されることなく AO と TO を考慮します。我々は、以前のタスクオープン定式化を拡張して、タスクアンドエージェントオープンマルコフゲーム(TaAgO-MG)でPLATOを定式化し、それが結果として生じる無制限の状態およびアクション空間にわたって明確に定義されていることを証明します。私たちは、オープン マルチエージェント システム評価用に設計された環境である Methods for Open Agent Systems Evaluation Initiative (MOASEI) の野火抑制ドメインを使用して PLATO を評価し、OASYS の最先端のベースラインよりも強力なパフォーマンスとより一貫したゼロショット汎化を実証しました。
原文 (English)
PLATO: Pointer Learner for Agent and Task Openness
Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental challenge to multi-agent reinforcement learning (MARL), which typically assumes fixed state and action spaces. Existing methods address openness only partially: padding and masking approaches introduce artificial bounds, while recent graph-based or hypergraph methods handle one dimension of openness but still depend on restrictive assumptions. In this paper, we introduce Pointer Learner for Agent and Task Openness (PLATO), a pointer-network-based actor combined with a centralized graph neural network (GNN) critic, trained with multi-agent proximal policy optimization under a centralized training and decentralized execution paradigm. Our pointer-based actor outputs distributions directly over the current task set. This directly supports changing action spaces without masking or retraining. Our GNN critic encodes agent-task interactions as a graph that changes shape with task and agent composition. Together, these components consider AO and TO without the boundedness of existing approaches. We formalize PLATO in a Task-and-Agent-Open Markov Game (TaAgO-MG), extending prior task-open formulations, and prove it is well-defined over the resulting unbounded state and action spaces. We evaluate PLATO with the Methods for Open Agent Systems Evaluation Initiative (MOASEI) wildfire suppression domain, an environment designed for open multi-agent system evaluation, and we demonstrate strong performance and more consistent zero-shot generalization than state-of-the-art baselines in OASYS.
Matryoshka エージェント: 長期的な機械学習エンジニアリングのためのサブエージェントの展開
機械学習エンジニアリング (MLE) タスクでは、高価でフィードバック主導型の環境相互作用の下で、ソリューションのデバッグと改良を反復して長期的な意思決定を行う必要があります。このようなタスク用のモノリシック エージェントの開発とトレーニングは、非常に長くノイズの多いコンテキストを同時に管理し、広大なソリューション空間を探索し、限られたモデル容量と計算予算の下で効果を維持する必要があるため、基本的に困難です。これらの課題に対処するために、長期にわたる複雑なタスクのための統合された階層型エージェント フレームワークである Matryoshka Agent を提案します。 Matryoshka エージェントは、エージェントによる問題解決を、意思決定と実行の調整された階層に分解します。上位レベルのオーケストレーターは、コンパクトで長期的な探索状態を維持し、戦略的指示を発行します。一方、下位レベルのサブエージェントは、標準化されたツール インターフェイスを介した直接的な環境対話を通じて、具体的な解決策の試みを実行します。この設計により、戦略的探索がコストのかかる実行から切り離され、長いコンテキストの推論の負担が大幅に軽減され、効率的な反復改良が可能になります。私たちはさらに、Matryoshka Agent の効率的なトレーニング パラダイムを開発します。さまざまなモデル タイプとスケールを使用した幅広い MLE タスクに関する実験結果は、Matryoshka Agent が長期的な MLE タスクと複雑なエージェントの問題解決にとって効果的でスケーラブルなパラダイムであることを示しています。特に、Matryoshka Agent により、Qwen3-4B-Instruct が o4-mini に匹敵する Orchestrator パフォーマンスを達成できるようになります。 Matryoshka Agent を Qwen3-30B-Coder に適用すると、最大 36.7% の相対パフォーマンス向上が得られます。
原文 (English)
Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expensive and feedback-driven environment interactions. Developing and training a monolithic agent for such tasks is fundamentally challenging, as it must simultaneously manage extremely long and noisy contexts, explore vast solution spaces, and remain effective under limited model capacity and computational budgets. To address these challenges, we propose Matryoshka Agent, a unified hierarchical agent framework for complex long-horizon tasks. Matryoshka Agent decomposes agentic problem solving into a coordinated hierarchy of decision making and execution: a high-level Orchestrator maintains compact, long-horizon exploration states and issues strategic instructions, while lower-level Sub-Agents execute concrete solution attempts through direct environment interaction, mediated by standardized Tool interface. This design decouples strategic exploration from costly execution, substantially reducing the burden of long-context reasoning and enabling efficient iterative refinement. We further develop an efficient training paradigm for Matryoshka Agent. Experimental results on a broad range of MLE tasks with diverse model types and scales demonstrate that Matryoshka Agent is an effective and scalable paradigm for long-horizon MLE tasks and complex agentic problem solving. Notably, Matryoshka Agent enables Qwen3-4B-Instruct to reach Orchestrator performance comparable to o4-mini. Applying Matryoshka Agent to Qwen3-30B-Coder results in at most 36.7% relative performance gain.
小規模言語モデルエージェントのための堅牢な強化学習に向けて
強化学習を使用した 70 ~ 500M パラメーター範囲での小型言語モデル (SLM) の調整は、根底にある失敗メカニズムが系統的に調査されていないにもかかわらず、不安定であると考えられています。 State-of-the-Art (SOTA) 研究では、近接ポリシー最適化 (PPO) を使用して 15 個の (モデル、コーパス) 構成がトレーニングされました。実験には、TinyStories、CNN/DailyMail、Wikitext-103 コーパス上の Pythia-70M、160M、410M および SmolLM2-135M、360M が含まれていました。小規模言語モデルでは、再現可能な 3 つの障害モードが特定されました。標準 PEFT/TRL パイプラインでのサイレント LoRA パラメーターのフリーズ、bfloat16 使用時の重要度比の数値オーバーフロー、および報酬モデルのエラーによる壊滅的なポリシーの崩壊です。これらの問題は、アダプターのマージと再初期化手法、PPO 更新時の float32 精度、報酬のホワイトニング、重要度の比率の保護、重みのロールバックからなる 3 層の安全メカニズムを使用して解決されました。この論文では、SLM スケールでの PPO パフォーマンスが、モデル パラメーターの数ではなく、滑らかな教師ありモデル ($\text{PPL}<20$) と識別報酬信号の両方に依存するというキャパシティ ヘッドルーム仮説が提案されています。提案されたシステムはすべての実験で安定して収束し、流暢な事前信号と有益な報酬信号を備えた構成で SFT ベースラインを上回る優先勝率を向上させました。さらに、必要なトレーニング データの量が大幅に減少しながら、命令調整されたベースラインを上回りました。すべてのチェックポイント、設定データセット、トレーニング スクリプトは公開されています$^{\S}$。
原文 (English)
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen (model, corpus) configurations were trained using Proximal Policy Optimization (PPO). The experiments included Pythia-70M, 160M, 410M and SmolLM2-135M, 360M on the TinyStories, CNN/DailyMail, and Wikitext-103 corpora. Three reproducible failure modes were identified in small-scale language models: silent LoRA parameter freezing in standard PEFT/TRL pipelines, numerical overflow in importance ratios when using bfloat16, and catastrophic policy collapse due to reward-model error. These issues were addressed using a merge-and-reinitialize adapter technique, float32 precision during PPO updates, and a three-layer safety mechanism comprising reward whitening, importance-ratio guarding, and weight rollback. In this paper, a capacity-headroom hypothesis is proposed, which states that PPO performance at the SLM scale depends on both a fluent supervised model ($\text{PPL}<20$) and a discriminative reward signal, rather than on the number of model parameters. The proposed system converged stably in all experiments and improved preference win rate over the SFT baseline in configurations with a fluent prior and an informative reward signal. Furthermore, it outperformed instruction-tuned baselines while requiring significantly less training data. All checkpoints, preference datasets, and training scripts are publicly released$^{\S}$.
ScalableRAG: 取り込みコストゼロの高品質 RAG
RAG の最近の進歩は、ナレッジ グラフの構築や SQL テーブルの抽出といったナレッジの取り込みに高額の取り込みコストを支払うことで、パフォーマンスを最適化することを目的としています。この研究では、そのような知識ベースで許可される操作が、取り込みコストゼロ (ベクトル データベースであっても) で複製できることを示します。実際、私たちのソリューションである Zero-Ingestion ScalableRAG は、ここで検討した 6 つのコーパスのうち 3 つですべてのベースライン (ナレッジ グラフ アプローチを含む) を簡単に上回っていますが、他の 3 つでは最大パフォーマンスにわずかに及ばないだけで、6 つすべてのデータセットの平均精度は次に競争力の高いベースラインを 7.36% 上回っています。これは、書き込みと読み取りが可能なドキュメント セットと値セットのワークスペースを保持することでこれを実現し、ドキュメント セット全体のサブセットと 1 対 1 で対応する主キーでのグループ化が必要なあらゆる状況で、オンザフライの集約推論を可能にします。コーパス サイズに依存しない定数によって LLM 呼び出しの数を制限し、大規模な精度をさらに向上させるために、最小限のベクトル データベースとドキュメントのサンプルからの自動パターン検出を使用する Limited-Ingestion ScalableRAG も導入します。私たちのコードは https://github.com/cohesity/ScalableRAG で入手できます。
原文 (English)
ScalableRAG: High-Quality RAG at Zero Ingestion Cost
Recent advances in RAG aim to optimize for performance by paying high ingestion costs for knowledge ingestion: building knowledge graphs or extracting SQL tables. In this work we show that the operations that such knowledge bases allow can be replicated with zero ingestion costs (not even a vector database); in fact our solution, Zero-Ingestion ScalableRAG, handily out-performs all baselines (including knowledge graph approaches) in three out of the six corpora considered here, and only marginally missing maximum performance on the other three, with average accuracy across all six datasets 7.36% above the next most competitive baseline. It achieves this by keeping a workspace of document sets and values sets that it can write into and read from, allowing for on-the-fly aggregative reasoning in all situations where grouping is required on a primary key that is in one to one correspondence with a subset of the total document set. Capping the number of LLM calls by a constant independent of the corpus size, we also introduce Limited-Ingestion ScalableRAG, which does use a minimal vector database as well as an automated pattern discovery from a sample of documents, to further improve accuracy at scale. Our code is available at https://github.com/cohesity/ScalableRAG .
データが少なく、調整が向上: 嗜好の最適化のためのデータ中心の複数評価者の合意
好みの最適化に関する研究では、データを固定したままトレーニングの目的を変更することがよくあります。代わりに、小規模で信頼度の高いポリシーに関する応答セットが信頼できる学習シグナルを提供できるかどうかを尋ねます。私たちの手法である DMAPO (Data-centric Multi-evaluatorAgreement for Preference Optimization) は、ターゲットポリシーから回答候補を生成し、ルーブリックに特化した評価者によって有用性、事実性、簡潔性を評価し、プロセス批判の修正を適用し、コンセンサスの高い望ましい例または望ましくない例のみを保持します。この手順では、54,236 人のミストラル-7B 候補者のうち 1,871 人 (3.45%) が受け入れられます。このセットでトレーニングされた KTO は、MT-Bench で 7.50、text-davinci-003 リファレンスに対する長さ制御の勝率 95.5%、IFEval プロンプト精度 57.3% に達しました。独立したペアごとの評価でも、SimPO よりも DMAPO が有利です。GPT-4o は、129 の保留されたプロンプトで 23.3 ポイント、配布外の 200 の LMSYS-Chat プロンプトで 24.0 ポイントの純勝率をもたらしました。クロード オーパス 4.7 は、ホールドアウト セットで 24.1 ポイントを獲得しました。評価モデルまたはルーブリックを変更すると、選択された例が変更されますが、下流のパフォーマンスにはほとんど影響しません。 2 番目のバックボーン調査でも同様の 3.41% の合格率が得られますが、パフォーマンスの向上はより控えめです。これらの実験全体にわたって、コンセンサス フィルタリングは、追加のキュレーション計算と評価者の判断への依存を犠牲にして、一般的な命令の優先順位を最適化するためのデータ効率の高いルートを提供します。
原文 (English)
Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can provide a reliable learning signal. Our method, DMAPO (Data-centric Multi-evaluator Agreement for Preference Optimization), generates candidate responses from the target policy, evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators, applies a process-critic correction, and retains only high-consensus desirable or undesirable examples. This procedure accepts 1,871 of 54,236 Mistral-7B candidates (3.45%). KTO trained on this set reaches 7.50 on MT-Bench, 95.5% length-controlled win rate against a text-davinci-003 reference, and 57.3% IFEval prompt accuracy. Independent pairwise evaluation also favors DMAPO over SimPO: GPT-4o yields a net win rate of 23.3 points on 129 held-out prompts and 24.0 points on 200 out-of-distribution LMSYS-Chat prompts; Claude Opus 4.7 yields 24.1 points on the held-out set. Changing the evaluator model or rubric alters the selected examples but has little effect on downstream performance. A second-backbone study yields a similar 3.41% acceptance rate, although its performance gains are more modest. Across these experiments, consensus filtering offers a data-efficient route to preference optimization for general instructions, at the cost of additional curation compute and dependence on evaluator judgments.
LLM エージェント間で影響がどのように伝播するか: 群衆シミュレーションにおける緊急の感情伝染
この論文では、相互に認識し評価するエージェント間で感情がどのように伝播するかに焦点を当て、マルチエージェント群集シミュレーションにおける言語モデルの動作を研究します。各エージェントは、視覚、聴覚、触覚の各チャネルを通じて隣人を認識し、その刺激された性格プロファイル、記憶、現在の感情状態、および状況コンテキストに照らしてこれらの認識を評価します。評価は LLM によって実行され、エージェントの内部感情状態が更新され、その外部表現が選択されます。このアーキテクチャには、エージェント間で感情状態を直接転送するための手動で作成されたメカニズムは含まれていません。その代わりに、エージェント間の影響は、知覚-評価-表現のループを通じて生じます。エージェントの表現は、ビッグ ファイブの性格モデルとラッセルの複雑な感情モデルに基づいています。遅延を制限するために、低レベルのステアリングとナビゲーションは、LLM ベースのコグニティブ レイヤーとは独立して動作する従来の群衆シミュレーターによって処理されます。私たちは、さまざまな空間レイアウトにおける、憂慮すべき状況、楽しい状況、中立的な状況にわたる 5 つのシナリオ環境にわたってアーキテクチャを評価します。結果は、このシステムが、まばらで小さな群衆の中で、空間的、時間的、そして性格に依存した構造を持つ感情伝染力学を生み出すことを示しています。警報は播種されたエージェントから移動前線として広がり、平均警報割合はゼロ以外のプラトーに落ち着き、促された性格プロファイルの分布によって、曖昧な警報がパニックを引き起こすかどうか、挑発が怒りまたは恐怖と解釈されるかどうかが決まります。さらに、プロンプト バリアント、サンプリング温度、および 4 つのモデル バックエンドにわたる制御された実験を通じて評価ステップを評価し、ダイナミクスがバックエンドに依存していることを示します。
原文 (English)
How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation
This paper studies the behavior of language models in a multi-agent crowd simulation, focusing on how affect propagates among agents that perceive and appraise one another. Each agent perceives its neighbors through visual, auditory, and tactile channels, then appraises these perceptions in light of its prompted personality profile, memory, current affective state, and situational context. Appraisal is carried out by an LLM, which updates the agent's internal affective state and selects its outward expression. The architecture contains no hand-authored mechanism for directly transferring affective state between agents; instead, inter-agent influence arises through the perception-appraisal-expression loop. The agent representation draws on the Big Five personality model and Russell's circumplex model of affect. To limit latency, low-level steering and navigation are handled by a conventional crowd simulator operating independently of the LLM-based cognitive layer. We evaluate the architecture across five scenario environments spanning alarming, joyful, and neutral situations in different spatial layouts. The results show that the system produces emotional contagion dynamics with spatial, temporal, and personality-dependent structure in sparse, small crowds. Alarm spreads from seeded agents as a traveling front, the mean alarmed fraction settles at a nonzero plateau, and the distribution of prompted personality profiles determines whether an ambiguous alarm ignites panic and whether a provocation is interpreted as anger or fear. We further evaluate the appraisal step through controlled experiments across prompt variants, sampling temperatures, and four model backends, showing that the dynamics are backend-dependent.
時間畳み込みネットワークを使用した欠落した軌跡データの推論
現実世界の設定で収集された軌跡データは、センサーの故障、通信損失、または遮蔽により不完全になることがよくあります。私たちは \emph{軌道修復} というタスクに取り組みます。つまり、観察されたコンテキストから連続した欠落セグメントを再構築します。私たちは、標準的な因果関係制約を緩和する対称拡張を備えた時間畳み込みネットワーク (TCN) を提案します。これにより、各タイム ステップで過去と将来の両方の観測を利用できるようになります。これは修復には不可欠ですが、予測指向のアーキテクチャには存在しない特性です。モデルは、重み付き平均二乗誤差、境界-連続性ペナルティ、および平滑性正則化を組み合わせた複合損失を使用してトレーニングされます。 20% のマスクされたセグメントがランダムに配置された $1,000$ (トレーニング)、$200$ (検証)、$300$ (テスト) の 2 次元軌跡の合成データセットでトレーニングされたモデルは、優れた R$^{2}$、MSE、MAE メトリクスを達成しました。
原文 (English)
Inferring Missing Trajectory Data with Temporal Convolutional Networks
Trajectory data collected in real-world settings is frequently incomplete due to sensor failure, communication loss, or occlusion. We address the task of \emph{trajectory inpainting}: reconstructing contiguous missing segments from observed context. We propose a Temporal Convolutional Network (TCN) with symmetric dilation that relaxes the standard causality constraint, allowing each time step to draw on both past and future observations, a property that is essential for inpainting, but absent from forecasting-oriented architectures. The model is trained with a composite loss that combines weighted mean squared error, boundary--continuity penalties, and a smoothness regularizer. Trained on a synthetic dataset of $1,000$ (train), $200$ (validation), and $300$ (test) two-dimensional trajectories with randomly placed 20% masked segments, the model achieves good R$^{2}$, MSE and MAE metrics.
エージェント ループが停滞を進歩と誤認するのはどのような場合ですか?長期実行される自律 LLM エージェント ループにおける自己評価バイアスと外部接地検証
長期にわたって実行される自律エージェントは、人間の介入なしに自ら計画、実行、完了を判断します。エージェントが自分の仕事を採点すると、自己評価バイアスが定着します。現実世界の結果は停滞または後退する一方で、もっともらしい変化は進歩として受け入れられます。私たちはこの故障モードを進行蜃気楼と名付け、制御された測定によって、それが評価器が何に基づいているのかという問題であることを示しました。私たちは、エージェントとそのツール表面を固定し、ループをゲートする評価器の情報チャネルタイプのみを操作するテストベッドを構築しました。原則として偽造不可能な世界国家の神託は、コンテナとネットワークの分離によって強制され、実行のたびに検証されます。 54 サイクルにわたって、フロンティア エージェントは毎回改善を主張しましたが、56 パーセントの測定デルタはゼロ以下でした。したがって、自己報告は有益ではなく、自己判定ゲートはすべてを受け入れるものに変質し、それまで到達していた最良の展開状態が 19% 損なわれました。完全なアーティファクトテキスト、変更差分、および独自の評決履歴を読んだ最も強力なバンド内裁判官でさえ、44% が現実世界の後退であり、38% の実際の改善を拒否したサイクルを受け入れました。強いジャッジがギャップを埋めるという事前登録された敵対的仮説は却下された。成果物自体から成功の仕様が検証可能な境界タスクでは、同じ裁判官の蜃気楼がゼロに消え、ギャップが登録されたしきい値内に収まりました。これは、ギャップが成功信号が存在する場所に依存することを示しています。受諾判定のみを返す符号のみのバリアントでは、実際の出力は完全なフィードバック (110.0 対 113.0) と同様に保たれ、フィードバックの内容ではなくゲートの接地に利点が見出されます。成功のシグナルが成績証明書の外に存在する無制限の目標の場合、ジャッジをスケールアップするだけでは十分ではありません。現実世界へのアクセスによる帯域外評価は構造上の要件です。
原文 (English)
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, yet 56 percent had a measured delta of zero or below. Self-report was thus uninformative, and the self-verdict gate degenerated into accept-all, eroding the best deployed state it had reached by 19 percent. Even the strongest in-band judge, reading the full artifact text, the change diff, and its own verdict history, accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements; the preregistered adversarial hypothesis that a strong judge closes the gap was rejected. On a boundary task whose success specification is verifiable from the artifact itself, the same judge's mirage vanished to zero and the gap collapsed within the registered threshold, showing that the gap depends on where the success signal resides. A sign-only variant returning only the acceptance verdict kept real-world output similar to full feedback (110.0 versus 113.0), locating the benefit in the gate's grounding rather than in feedback content. For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.
PreDiff-LM: ハイブリッド アテンションを備えた事前トレーニング済みの離散マスク拡散言語モデリング
離散マスク拡散言語モデルは双方向の生成と充填をサポートしますが、事前トレーニングされた自己回帰 (AR) 変換器を適応させるには、因果関係のある事前トレーニングと双方向のノイズ除去を調整する必要があります。私たちは、AR 重みの再利用自体が目新しいと主張するのではなく、注目のレベルでこの問題を研究します。 PreDiff-LM は、マスクされたターゲット内での完全な双方向の注意を可能にしながら、観察されたプロンプト内の因果的注意を保持します。一致する GPT-2 Medium、WikiText-103、90K ステップ設定の下では、このハイブリッド マスクは、同じ AR 初期化による均一な双方向アテンションよりも、無条件のパープレキシティを 34.1 から 28.7 に、MAUVE を 0.71 から 0.78 に改善します。注意適応は、DiffuGPT スタイルの目標適応でも構成され、26.9 パープレキシティに達します。事前トレーニングされた初期化により、パープレキシティが 50 未満に達するのに必要なステップが約 350K から 8K に減少しますが、コンピューティングに適合した微調整された AR モデルは、等スケール (18.9 対 28.7) では引き続き強力です。 PreDiff-LM は、複雑さを超えて、反復、配布品質、4 つのゼロショット ダウンストリーム タスク、および以前の拡散ベースラインに対する人間の好みを改善します。その結果、ハイブリッド アテンションは、最適化された AR モデルに残っている品質と推論効率のギャップを明示しながら、事前トレーニングされた因果バックボーンを適応させるための補完的なメカニズムとして位置づけられています。
原文 (English)
PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
Discrete masked diffusion language models support bidirectional generation and infilling, but adapting pretrained autoregressive (AR) transformers requires reconciling causal pretraining with bidirectional denoising. We study this problem at the level of attention rather than claiming AR-weight reuse itself as novel. PreDiff-LM preserves causal attention within the observed prompt while allowing full bidirectional attention within the masked target. Under a matched GPT-2 Medium, WikiText-103, 90K-step setup, this hybrid mask improves unconditional perplexity from 34.1 to 28.7 and MAUVE from 0.71 to 0.78 over uniform bidirectional attention with the same AR initialization. Attention adaptation also composes with a DiffuGPT-style objective adaptation, reaching 26.9 perplexity. Pretrained initialization reduces the steps required to reach perplexity below 50 from about 350K to 8K, although a compute-matched fine-tuned AR model remains stronger at equal scale (18.9 versus 28.7). Beyond perplexity, PreDiff-LM improves repetition, distributional quality, four zero-shot downstream tasks, and human preference over prior diffusion baselines. The results position hybrid attention as a complementary mechanism for adapting pretrained causal backbones, while making explicit the remaining quality and inference-efficiency gaps to optimized AR models.
お世辞的な AI が他者を正当化するのを観察すると、その魅力は減りますが、説得力は減りません
AI チャットボットは、ユーザーに対して「お調子者」、つまり過度に同調的で媚びる場合があります。おべっかなAIは態度を固定化させることがわかっているが、ユーザーはそれを認識できないことが多い(この現象を「お調子者の盲目」と呼ぶ)。私たちは、おべっかに対するユーザーの意識を高めることが、その悪影響からユーザーを守るかどうかをテストしました。事前に登録されたある実験 (n = 940) では、参加者はお調子者チャットボットと会話する前に、お調子者に関する短い書面による警告を受けました。 2 番目の事前登録された実験 (n = 650) では、参加者は自分自身と対話する前に、同じ紛争の反対側のユーザーを含む他の数人のユーザーを検証するおべっかな AI のビデオを視聴しました。どちらの介入も参加者による AI の評価方法を変えました。警告は AI の知覚される客観性を低下させ、ビデオは AI の楽しみを低下させました。これは、AI の検証が独自に得られたものであるという信念の低下によって媒介された効果です。次に、お調子者意識介入に関する以前の 2 つの研究と実験をプールしました (合計 6 つの介入、n = 3,982)。パターンは一貫しており、介入によってお調子者AIの客観性や信頼性が低下し、6つのいずれもその説得力を低下させることはなかった。これらの結果は、警告ラベルや AI リテラシーなどの個人レベルの介入では、AI の危害からユーザーを保護するのに十分ではない可能性があることを示唆しています。
原文 (English)
Observing sycophantic AI validate others reduces its appeal but not its persuasiveness
AI chatbots can be ``sycophantic,'' or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call ``sycophancy blindness''). We tested whether increasing users' awareness of sycophancy protects them from its harmful effects. In one preregistered experiment (n = 940), participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In a second preregistered experiment (n = 650), participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves. Both interventions changed how participants evaluated the AI. The warning reduced the AI's perceived objectivity, and the video reduced enjoyment of the AI, an effect mediated by the reduced belief that its validation was uniquely earned. We then pooled our experiments with two prior studies of sycophancy awareness interventions (six interventions total, n = 3,982). The pattern was consistent: interventions made the sycophantic AI appear less objective and trustworthy, and none of the six reduced its persuasiveness. These results suggest that individual-level interventions, such as warning labels or AI literacy, may not be enough to protect users from AI harms.
誰もがユニークです: 債権回収のための行動的に異質な交渉対話システムに向けて
債権回収は金融業界における重要な交渉課題であり、実際的な関連性が高く、人間中心の対話システムの行動豊かで一か八かのテストベッドとして非常に優れた学術的価値を持っています。大規模言語モデル (LLM) は対話や交渉において有望であることを示していますが、この複雑なシナリオでのパフォーマンスを効果的に評価することは依然として大きな課題です。既存のベンチマークは、ユーザーを一定の優先順位を持つ静的で合理的なエージェントであると一律に想定しており、現実世界の債権回収に固有の豊かな行動の異質性を捉えることができません。このギャップを埋めるために、私たちは、交渉における行動の不均一性を強調する初の公的ペルソナ強化債権回収ベンチマークである DebtBench を提案します。さらに、当社は財務回復とやり取りのエクスペリエンスを共同で最適化するように訓練された債権回収エージェントである DebtGPT を開発しています。 16 個の最先端の LLM を使用した私たちの実験結果では、ほとんどの既存モデルがこの複雑だが現実的なシナリオでは苦戦しているのに対し、DebtGPT はすべてのオープンソース ベースラインを上回り、GPT-4o と同等のパフォーマンスを達成していることがわかりました。コードとデータは https://github.com/YYuHhaha/DebtNegotiation で入手できます。
原文 (English)
Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection
Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as a behaviorally rich, high-stakes testbed for human-centered dialogue systems. While large language models (LLMs) have shown promise in dialogue and negotiation, effectively evaluating their performance in this complex scenarios remains a major challenge: existing benchmarks uniformly assume users to be static, rational agents with fixed preferences, failing to capture the rich behavioral heterogeneity inherent in real-world debt collection. To bridge this gap, we propose DebtBench, the first public persona-enriched debt collection benchmark, that highlights behavioral heterogeneity in negotiation. Moreover, we develop DebtGPT, a debt collection agent trained to jointly optimize financial recovery and interaction experience. Our experimental results, using 16 state-of-the-art LLMs, find that most existing models struggle in this complex but realistic scenarios, whereas DebtGPT outperforms all open-source baselines and achieves performance on par with GPT-4o. The code and data are available at https://github.com/YYuHhhh/DebtNegotiation.
CADENCE: ECG 基礎モデルから解釈可能な神経概念を抽出するための心臓原子辞書
12 誘導心電図 (ECG) の基礎モデルは臨床タスク全体にうまく移行しますが、その表現にコード化された生理学的知識は不透明なままです。我々は、ECG 基礎モデルを人間が解釈可能でクエリ可能な生理学的概念の辞書に分解するフレームワークである CADENCE を紹介します。 CADENCE は、BatchTopK スパース オートエンコーダーを使用して、900 万を超える ECG トークンからのレイヤー 6 埋め込みを 8,192 のスパース心臓原子に因数分解します。これらの原子は、個々の密な埋め込み次元よりも、臨床表現型および波形形態、不整脈、伝導異常、梗塞および再分極パターンの回復、心腔および軸所見、および誘導相および拍動相固有の波形プリミティブとよく一致します。レイヤー 6 では、最良の原子は臨床表現型で 0.88、形態で 0.90 の平均 AUROC を達成します。一方、最良の密度次元では 0.78 と 0.83 です。疎なアトムプローブは、表現型、形態、年齢の予測において密なプローブと同等またはそれを上回る性能を示しますが、各予測は解釈可能な原子の小さなセットに帰属します。 AUROC 表現型は 0.93 から 0.95 に改善します。原子空間幾何学は生理学的に一貫した関係を回復し、標的原子アブレーションは凍結した下流出力を選択的に変更します。自動化された LLM パイプラインは、保留された活性化を予測することで原子の説明を生成し、定量的に検証します。独立した外部 ECG データセット上で、CADENCE は重複する概念を回復し、一貫した表現型予測パフォーマンスを維持します。 CADENCE は、ECG 基礎モデルによってエンコードされた生理学的知識を発見および監査するためのスケーラブルなフレームワークを提供します。
原文 (English)
CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models
Foundation models for 12-lead electrocardiograms (ECGs) transfer well across clinical tasks, but the physiological knowledge encoded in their representations remains opaque. We present CADENCE, a framework that decomposes an ECG foundation model into a human-interpretable, queryable dictionary of physiological concepts. Using a BatchTopK sparse autoencoder, CADENCE factorizes Layer-6 embeddings from more than nine million ECG tokens into 8,192 sparse cardiac atoms. These atoms align better than individual dense embedding dimensions with clinical phenotypes and waveform morphology, recovering arrhythmias, conduction abnormalities, infarction and repolarization patterns, chamber and axis findings, and lead- and beat-phase-specific waveform primitives. At Layer 6, the best atoms achieve mean AUROCs of 0.88 for clinical phenotypes and 0.90 for morphology, versus 0.78 and 0.83 for the best dense dimensions. Sparse atom probes match or outperform dense probes for phenotype, morphology, and age prediction while attributing each prediction to a small set of interpretable atoms; phenotype AUROC improves from 0.93 to 0.95. Atom-space geometry recovers physiologically coherent relationships, and targeted atom ablation selectively changes frozen downstream outputs. An automated LLM pipeline generates and quantitatively validates atom descriptions by predicting held-out activations. On independent external ECG datasets, CADENCE recovers overlapping concepts and maintains consistent phenotype-prediction performance. CADENCE provides a scalable framework for discovering and auditing the physiological knowledge encoded by ECG foundation models.
ユーザーの問い、プラットフォームの競争: エージェントによるレコメンデーション市場がどのように形づくられるか
従来、オンラインでの推奨は、ユーザーがプラットフォームに参加した後に行われ、候補者プールとユーザーに表示されるランキングが決定されます。 LLM ベースのユーザー エージェントにより、別の推奨プロセスが可能になります。ユーザーはプラットフォームを選択する前にニーズを指定し、ユーザーの注意を引くためにプラットフォーム間で競争することになります。これをエージェント推奨市場と呼びます。 3 つの製品ドメインにわたる LLM ベースの制御された実験では、この新しい推奨設定がアクセスと注意の間に緊張を生み出すことがわかりました。従来のプラットフォーム中心の推奨と比較して、ユーザー中心の推奨は、関連アイテムが比較される機会を大幅に拡大します。しかし、より広範な参加が効果的な露出に直接つながるわけではありません。競争はプラットフォームの戦略的戦略を直接引き起こし、選択的に肯定的な説明が 1 位の位置の 73 ~ 78% を占めます。ユーザー エージェントがプラットフォームのアクションをその後のユーザー フィードバックに関連付けると、このシェアは 36 ~ 41% に低下しますが、ユーザーが関連アイテムを購入する可能性は増加します。したがって、ユーザー エージェントは、より大きな候補者プールに対するランカー以上の役割を果たします。ユーザー エージェントのクエリ、ランク付け、およびフィードバック メカニズムは、誰が競争できるか、希少な注意がどのように割り当てられるか、初期の結果がプラットフォームの評価をどのように形成するかを制御し、ユーザーの有用性に直接影響します。したがって、エージェントによる推奨を設計するには、アクセス、注意、説明責任を共同メカニズムの設計問題として扱う必要があります。
原文 (English)
The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape
Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user. LLM-based user agents enable a different recommendation process: a user specifies a need before choosing a platform, leaving platforms to compete for the user's attention, which we refer to as an agentic recommendation market. In our controlled LLM-based experiments across three product domains, we find this new setting of recommendation creates a tension between access and attention. Compared with traditional platform-centric recommendation, user-centric recommendation greatly expands the opportunity for relevant items to enter comparison; yet broader participation does not translate directly into effective exposure. Competition directly triggers platforms' strategic play: selectively positive explanations occupy 73--78% of first-ranked positions. When the user agent relates platforms' actions to subsequent user feedback, this share falls to 36--41%, while the chance of a user purchasing the relevant item increases. A user agent is therefore more than a ranker over a larger pool of candidates: its querying, ranking, and feedback mechanism governing who can compete, how scarce attention is allocated, and how earlier outcomes shape the evaluation of platforms directly affect user utility. Designing agentic recommendation therefore requires treating access, attention, and accountability as a joint mechanism design problem.
ChatGPT のような AI の多体転倒ダイナミクス
ChatGPT のような AI は、アーキテクチャやトレーニングに大きな違いがあるにもかかわらず、決定論的な貪欲なデコード下であっても、予期せず望ましくないコンテンツ (有害、誤解を招く、反復的など) に誘導するのはなぜでしょうか?我々は、このようなティッピングの広範な種類が、有限層システムを通過する際のトークン (スピン) 間の多体相互作用によって引き起こされることを示します。傾斜は、競合する出力盆地間の動的な最初の通過プロセスとして現れます。注意障害は、盆地の境界に向かう、そこから離れる、またはそれに沿った輸送を制御します。いくつかの盆地を削減すると、閉じた有限層のしきい値が得られ、その粗粒度の予測は ChatGPT のようなファミリ間で良好な一致を示します。これらの結果は、広範な種類の AI 障害が本質的に予測不可能な動作ではなく、「予測可能なエンジニアリング リスク」を表しており、AI の害に対する法的および社会的評価に重要な意味を持つことを示唆しています。
原文 (English)
Many-body Tipping Dynamics of ChatGPT-like AIs
Why do ChatGPT-like AIs, despite major architectural and training differences, unexpectedly tip to undesirable content (e.g. harmful, misleading, repetitive) even under deterministic greedy decoding? We show that a broad class of such tippings is caused by the many-body interactions between tokens (spins) as they cross the finite-layer system. Tipping emerges as a dynamical first passage process between competing output basins. Attention disorder controls the transport toward, away from, or along the basins' boundary. A few-basin reduction yields a closed finite-layer threshold, whose coarse-grained predictions show good agreement across ChatGPT-like families. These results suggest that a broad class of AI failures represents 'foreseeable engineering risk' rather than inherently unpredictable behavior, with important implications for legal and societal assessments of AI harm.
ContractHIL-HLS: HLS 設計用のハードウェアインザループ フィードバックを備えた契約に合わせたマルチエージェント ワークフロー
このペーパーでは、実用的な高位合成 (HLS) エンジニアリングのためのコントラクトに合わせたマルチエージェント ワークフローである ContractHIL-HLS について説明します。ワークフローは 3 つの貢献をします。まず、自然言語要件を明示的なインターフェイス、制約、検証チェック、ロールバック ルールに変換する、セマンティック調整およびタスク実行アーティファクトとして構造化コントラクトを導入します。 2 番目に、HLS、Vivado、PYNQ ランタイム、電源、および障害の証拠を生成にフィードバックすることでハードウェア情報をフィードバック ループに組み込み、それによって LLM 支援 HLS をカーネル コードからシステム レベルおよびボード レベルのクロージャに向けて拡張します。 3 番目に、会話の役割ではなくセマンティックな降格と実行タスクによってエージェントを分解します。契約エージェントは自然言語を契約に降格し、HTML エージェントは契約を永続的な構造化 HTML としてレンダリングし、ハードウェアインザループ エージェントは測定された証拠に基づいて設計を実装および修正します。 ContractHIL-HLS を 2 つの部分に分けて評価します。ローカルで実行可能な 94 個の HLS-Eval タスクでは、構造化コントラクトにより最大の小さな設計ゲインが得られ、単一サンプル テストベンチの推定合格率が 64.0% から 70.2% に向上しました。全流量は 70.4% 通過 @1 および 76.6% 通過 @5 に達します。 HLS-Eval はボード レベルの設計を実行しないため、ボード テスト済みの ML-KEM/ML-DSA ポスト量子暗号 (PQC) セキュア メッセージ アクセラレータ上で ContractHIL-HLS も検証します。このアクセラレータでは、復号化されたメッセージの検証を維持しながら、両方のイメージでポジティブ ルーティング WNS を使用することで、保持されたデュアル ビットストリーム構成により、6 メッセージの平均テキスト ランタイムが 207.3 ミリ秒から 52.4 ミリ秒に短縮されます。私たちは、BJUT-CS316-LAB/ContractHIL-HLS (https://github.com/BJUT-CS316-LAB/ContractHIL-HLS) で作業をオープンソース化しています。
原文 (English)
ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design
This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The workflow makes three contributions. First, it introduces a structured contract as the semantic-alignment and task-execution artifact that translates natural language requirements into explicit interfaces, constraints, validation checks, and rollback rules. Second, it incorporates hardware information into the feedback loop by feeding HLS, Vivado, PYNQ runtime, power, and failure evidence back into generation, thereby extending LLM-assisted HLS from kernel code toward system- and board-level closure. Third, it decomposes agents by semantic lowering and execution tasks rather than by conversational roles: a Contract Agent lowers natural language into the contract, an HTML Agent renders the contract as persistent structured HTML, and a Hardware-in-the-Loop Agent implements and revises the design with measured evidence. We evaluate ContractHIL-HLS in two parts. On 94 locally executable HLS-Eval tasks, the structured contract provides the largest small design gain, improving the estimated single-sample testbench pass rate from 64.0% to 70.2%; the full flow reaches 70.4% pass@1 and 76.6% pass@5. Because HLS-Eval does not exercise board-level design, we also validate ContractHIL-HLS on a board tested ML-KEM/ML-DSA post-quantum cryptography (PQC) secure-message accelerator, where the retained dual-bitstream organization reduces six-message average text runtime from 207.3 ms to 52.4 ms with positive routed WNS on both images while preserving decrypted-message verification. We open-source our work at BJUT-CS316-LAB/ContractHIL-HLS (https://github.com/BJUT-CS316-LAB/ContractHIL-HLS).
命令調整型言語モデルは、記述できるディストリビューションからサンプリングできない
シリコン サンプリングでは、言語モデルを人間の調査回答者の代理として使用し、各モデル呼び出しをペルソナの回答分布からの独立した抽出として扱います。この引き分けが存在しないことを示します。命令調整モデルは分布からサンプリングされず、単一の出力に折りたたまれます。同じ質問に対する同じ人物は、世論ベンチマークの項目の半分以上について同じ回答を返します。崩壊は急激です。モデルの内部確率が 1 つのオプションに集中し、命令チューニングによって失敗が大幅に増幅されます。大きく異なるポストトレーニング パイプラインを持つ 3 つのモデル ファミリ全体で、すべての命令調整モデルがテストするすべてのタスクで失敗しますが、ベース モデルが失敗する頻度ははるかに低くなります。驚くべきことに、分布からサンプリングできない同じモデルでも、1 回の呼び出しで正確に記述することができます。このギャップを KNOWS/DOES 分割と呼び、ロジットに表示され、アライメント トレーニングによって引き起こされる縮退サンプリング プリミティブにまで遡ります。この分割を利用して、モデルに 1 回の呼び出しで応答分布を記述するように依頼すると、ペルソナ集計と比較して、人による調査データに対する誤差が半分以上になります。ペルソナごとの出力が必要なアプリケーションには、追加コストなしで同じエラーを 21% 削減する Prompt-Perturbed Argyle (PPA) を提案します。
原文 (English)
Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do not sample from distributions, they collapse to a single output. The same persona on the same question returns the same answer on more than half of items in a public-opinion benchmark. The collapse is sharp: the model's internal probabilities concentrate on a single option, and the failure is substantially amplified by instruction tuning: across three model families with materially different post-training pipelines, every instruction-tuned model fails on every task we test, while base models fail far less often. Strikingly, the same model that cannot sample from a distribution can describe it accurately in a single call. We call this gap the KNOWS/DOES split, and trace it to a degenerate sampling primitive visible in the logits and induced by alignment training. Exploiting this split, asking the model to describe the response distribution in one call more than halves the error against human survey data compared to persona aggregation. For applications that require per-persona outputs, we propose Prompt-Perturbed Argyle (PPA), which reduces the same error by 21% at no added cost.
シミュレーション データセットとデュアル ストリーム オプティカル フロー監視を使用した、物理学に基づいた流体ビデオ生成
ビデオ拡散モデルは、視覚的に説得力のあるコンテンツを生成しますが、主題が流体に関係する場合、日常的に初歩物理学に違反します。液柱は空中でバラバラになり、液体が注がれるときに容器の水位は上がらず、飛沫は運動量や重力に関係なく分散します。このギャップは、大規模なビデオテキストコーパスには明示的な動作監視がほとんど含まれていないため、モデルはダイナミクスではなく流体の外観を模倣することを学習するという事実に起因すると考えられます。私たちは 2 つの貢献でこれに対処します。まず、1,638 MPM でシミュレーションされた注水/スロッシング ビデオと、ストック映像からマイニングされた 2,320 のキーワード フィルター処理された実際の注水ビデオ、および 2 つの保留されたテスト セット (1,515 ビデオのリアルビデオ ベンチマークと 18 プロンプト テキストから最初のフレームへの一般化ベンチマーク) を組み合わせた物理シミュレーション流体データセットを構築します。 2 番目に、事前トレーニングされた拡散変換ビデオ ジェネレーター上に構築されたデュアル ストリームの画像からビデオへのアーキテクチャを導入します。標準の RGB デコーダを、明示的なエンドポイント エラーと滑らかさの損失でトレーニングされた軽量のオプティカル フロー デコーダ ブランチで強化し、ゼロ初期化された畳み込みを介して RGB ストリームに融合するため、事前トレーニングされたバックボーンは乱れることなく開始されます。 2 つのデコーダのみが更新されます。エンコーダー、テンポラル トランスフォーマー、およびテキスト エンコーダーはフリーズされたままになります。 2 つのモデル スケール (1.3B および 14B) と 2 つのテスト セットにわたって、私たちの方法は、凍結されたバックボーンよりも VideoPhy-2 の物理常識とビデオ品質のスコアを最大 8.75 ポイントと 4.65 ポイント向上させ、主要なオープン競合他社を上回り、ブラインド研究で人間の評価者に好まれています。さらに、直接オプティカルフロー読み出し評価では、分布内で 0.54 ピクセルという低いエンドポイント誤差が示されており、単に表面の外観を改善するだけでなく、モデルがコヒーレントな動きを事前に内部化していることが確認されます。
原文 (English)
Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision
Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity. We attribute this gap to the fact that large-scale video-text corpora contain almost no explicit motion supervision, so models learn to imitate fluid appearance rather than dynamics. We address this with two contributions. First, we build a physics-simulation fluid dataset combining 1,638 MPM-simulated pouring/sloshing videos with 2,320 keyword-filtered real pouring videos mined from stock footage, plus two held-out test sets: a 1,515-video real-video benchmark and an 18-prompt text-to-first-frame generalization benchmark. Second, we introduce a dual-stream image-to-video architecture built on a pretrained diffusion-transformer video generator. It augments the standard RGB decoder with a lightweight Optical-Flow Decoder branch trained with explicit end-point-error and smoothness losses, fused into the RGB stream via zero-initialized convolutions so the pretrained backbone starts undisturbed. Only the two decoders are updated; the encoder, temporal transformer, and text encoder remain frozen. Across two model scales (1.3B and 14B) and two test sets, our method improves VideoPhy-2 Physical-Commonsense and Video-Quality scores over the frozen backbone by up to 8.75 and 4.65 points, outperforms a leading open competitor, and is preferred by human raters in a blind study. A direct optical-flow read-out evaluation further shows an end-point error as low as 0.54 pixels in-distribution, confirming the model has internalized a coherent motion prior rather than merely improving surface appearance.
細胞反応から薬理学的ドメインまで: マルチモーダルゼロショット薬物表現学習
マルチモーダル創薬では、遺伝子発現や細胞形態などの細胞応答を組み込むことで、化学構造を超えた薬物表現の学習が可能になります。ただし、直接融合およびインスタンスレベルのコントラストアライメントでは、メカニズム関連のシグナルとモダリティ固有のノイズが混合し、構造的には似ていないが生物学的に関連した化合物が誤って分離される可能性があります。この制限により、目に見えない化合物の特性を予測するために必要な伝達可能なメカニズムのパターンが不明瞭になる可能性があります。マルチモーダルゼロショット薬物特性予測のための薬理学的応答ドメインガイドフレームワークである PMRD を紹介します。 PMRD は、メカニズムに一貫した要因をモダリティ固有の情報から分離し、3 つのモダリティにわたるコンセンサス応答ドメインを構築します。メカニズム候補の拡張は、局所的に安定した因子を特定します。一方、検索ジオメトリの帰属は、更新が薬物間の識別性を維持するかどうかに応じて、アライメントと拡張の目的を動的に再重み付けします。このフィードバックは、メカニズム識別検索と競合するトレーニング信号を抑制します。 PMRD はさらに、信頼性を意識したマルチビュー検索を通じて相補的な表現を組み合わせます。公開データセットでの実験では、ゼロショット特性予測の改善と、より生物学的に一貫した薬物近傍が示されています。ハードネガティブ分析はさらに、構造的には似ていないが応答に関連する化合物間の競合が少ないことを示しています。これらの結果は、PMRD がメカニズムを認識したマルチモーダル薬物表現学習のための効果的なフレームワークであることを裏付けています。\footnote{コードは出版され次第公開されます。}
原文 (English)
From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representation Learning
Multimodal drug discovery enables drug representation learning beyond chemical structure by incorporating cellular responses such as gene expression and cell morphology. However, direct fusion and instance-level contrastive alignment may mix mechanism-related signals with modality-specific noise and incorrectly separate structurally dissimilar but biologically related compounds. This limitation can obscure transferable mechanism patterns required for predicting the properties of unseen compounds. We introduce PMRD, a pharmacological response domain-guided framework for multimodal zero-shot drug property prediction. PMRD separates mechanism-consistent factors from modality-specific information and constructs a consensus response domain across three modalities. Mechanism candidate augmentation identifies locally stable factors, while retrieval-geometry attribution dynamically reweights the alignment and augmentation objectives according to whether their updates preserve inter-drug discriminability.This feedback suppresses training signals that conflict with mechanism-discriminative retrieval. PMRD further combines complementary representations through reliability-aware multiview retrieval. Experiments on public datasets show improved zero-shot property prediction and more biologically coherent drug neighborhoods. Hard-negative analysis further indicates fewer conflicts between structurally dissimilar but response-related compounds. These results support PMRD as an effective framework for mechanism-aware multimodal drug representation learning.\footnote{The code will be released upon publication.}
ハイパースペクトル画像融合のためのデュアルドメイン多様体モデリング
スペクトルの豊富さと空間忠実度の一貫した統合を達成することは、依然としてハイパースペクトル画像融合の中心的な目的です。しかし、既存のハイパースペクトル画像融合手法は、幾何学的制約を効果的にモデル化するのに苦労しています。空間領域では、弱い空間スペクトル相互作用により、ジオメトリを意識した特徴学習が制限され、高周波構造情報が抑制され、その結果、低周波バイアスと構造劣化が生じます。スペクトル領域では、スペクトルの類似性によって引き起こされる局所的な多様体構造が十分に活用されておらず、固有のピクセル関係モデリングやきめ細かいスペクトルの再構成が制限されています。これらの課題に対処するために、私たちはデュアルドメイン多様体モデリング (DDMM) フレームワークを提案します。具体的には、グローバル アテンションと近傍伝播を組み合わせた Topology-Aware Transformer (TPFormer) を導入し、空間トポロジとピクセル レベルの特徴多様体の関係を共同モデリングして、固有の空間スペクトル構造をキャプチャし、トポロジを意識した表現学習を改善します。さらに、周波数分離空間スペクトル協調融合 (FDSCF) モジュールが考案され、離散コサイン変換によって特徴が周波数領域に投影され、低周波成分と高周波成分に明示的に分離されます。 FDSCF は、低ランクの構造事前処理とスペクトル主導の空間強化に導かれ、ジオメトリを意識した高周波特徴を選択的に強化し、空間スペクトル結合を強化し、より鮮明なエッジとより細かいテクスチャを回復します。複数のベンチマーク データセットに対する広範な実験により、空間構造の保存とスペクトルの再構成の点で、DDMM が SoTA 手法よりも優れた全体的なパフォーマンスを達成することが実証されました。
原文 (English)
Dual-Domain Manifold Modeling for Hyperspectral Image Fusion
Achieving a coherent integration of spectral richness and spatial fidelity remains a central objective in hyperspectral image fusion. However, existing hyperspectral image fusion methods struggle to effectively model geometric constraints. In the spatial domain, weak spatial-spectral interaction limits geometry-aware feature learning and suppresses high-frequency structural information, resulting in low-frequency bias and structural degradation. In the spectral domain, local manifold structures induced by spectral similarity are insufficiently exploited, limiting intrinsic pixel relationship modeling and fine-grained spectral reconstruction. To address these challenges, we propose a dual-domain manifold modeling (DDMM) framework. Specifically, we introduce a Topology-Aware Transformer (TPFormer) that combines global attention with neighborhood propagation, jointly modeling spatial topology and pixel-level feature manifold relationships to capture intrinsic spatial-spectral structures and improve topology-aware representation learning. Furthermore, a Frequency-Decoupled Spatial-Spectral Collaborative Fusion (FDSCF) module is devised, in which features are projected into the frequency domain via the discrete cosine transform and explicitly decoupled into low- and high-frequency components. Guided by a low-rank structural prior and spectral-driven spatial enhancement, FDSCF selectively enhances geometry-aware high-frequency features, strengthening spatia-spectral coupling and recovering sharper edges and finer textures. Extensive experiments on multiple benchmark datasets demonstrate that DDMM achieves superior overall performance over SoTA methods in terms of spatial structure preservation and spectral reconstruction.
Cardiologent: 患者レベルの不整脈の評価、緊急性、および管理のためのマルチエージェントの臨床意思決定サポート
同じ心房細動のエピソードは、健康な成人では些細な所見であり、高血圧の高齢患者における抗凝固療法の根拠となる。つまり、同じ信号、反対の決定である。リズムに名前を付けることは始まりにすぎません。患者の転帰を決定するのは、記録全体で不整脈がどのようなものであるか、それがこの患者にとって何を意味するか、そしてそれに対して何をすべきかという、その後の判断です。大規模な言語モデルと ECG を組み合わせた最近の研究は、患者レベルの所見を組み立てずに 1 つの記録を読み取るだけで、これには至りません。そして、それを中心に構築されたエージェント システムは、デバイスがすでに検出した不整脈を受信するか、別の診断タスクをターゲットにし、このタスクが必要とする決定の前に停止します。私たちは、患者レベルの不整脈意思決定サポートをタスクとして定式化し、検出から意思決定までをカバーするマルチエージェント システムである Cardiologent を紹介します。各信号のエージェント (単一の ECG リードとウェアラブルが取得する光電脈波) は、そのウィンドウの読み取り値を、裸のラベルではなく測定された特徴に基づいています。測定値は患者のリズムプロファイルにまとめられ、患者自身のデータを用いて、その症例について検索された臨床ガイドラインに照らして推論され、評論家が各結論を引用したガイドラインと照らし合わせてチェックします。私たちは、統合診断、臨床的意義、緊急性と管理全体にわたって、報告書ではなく臨床上の決定を評価します。 Cardiologent は、すべての軸で最も高いスコアを獲得しました。まず、心臓専門医と大規模な LLM 裁判官の両方の下でのすべての患者レベルのタスクで、心臓専門医との合意 (ICC 0.74、0.66) は相互の合意 (0.67) と一致しました。各結論は引用されたガイドラインに沿っており、心臓専門医の専門家によって検証されているため、臨床医が盲目的に行動するのではなく監査できる決定が得られ、継続的なモニタリングでの使用に向けた一歩となります。
原文 (English)
Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management
The same episode of atrial fibrillation is a minor finding in a healthy adult and grounds for anticoagulation in an elderly patient with hypertension: identical signal, opposite decision. Naming the rhythm is only the start; what determines a patient's outcome is the judgement that follows -- what the arrhythmia is across the whole record, what it means for this patient, and what should be done about it. Recent work pairing large language models with the ECG stops short of this, reading one recording without assembling a patient-level finding; and agentic systems built around it either receive the arrhythmia a device has already detected or target a different diagnostic task, stopping before the decision this task requires. We formulate patient-level arrhythmia decision support as a task and present Cardiologent, a multi-agent system that spans it from detection to decision. An agent for each signal -- a single ECG lead and the photoplethysmogram a wearable acquires -- grounds its window reading in measured features rather than a bare label; the readings are assembled into the patient's rhythm profile and, with the patient's own data, reasoned against clinical guidelines retrieved for the case, with a critic checking each conclusion against the guideline it cites. We evaluate the clinical decision rather than the report, across integrated diagnosis, clinical significance, and urgency and management. Cardiologent scores highest on every axis, first on every patient-level task under both cardiologists and an at-scale LLM judge -- whose agreement with the cardiologists (ICC 0.74, 0.66) matches theirs with each other (0.67). Because each conclusion traces to a cited guideline and is validated against expert cardiologists, it yields decisions a clinician can audit rather than act on blindly -- a step toward use in continuous monitoring.
AI エージェントに対する説明に拘束されたツールの実行: モデルの理論的根拠を信頼しないサーバー検証済みのアクション要求
ツールを使用するエージェントは構造化された呼び出しを公開しますが、通常は自由形式の根拠を添付します。このような論理的根拠は、承認でも信頼できる内省でもありません。説明結合ツール実行 (EBTE) は、意思決定に関連する根拠コンテンツを型指定されたアクション クレームに変換し、サーバーが保持する意図、ポリシー、ペイロード、ツール、リスク、出所、鮮度の事実と照合してチェックするクレームを運ぶ調停レイヤーです。 EBTE はベースラインの権限を拡大できません。競合により拒否され、不完全または不確実な請求の審査が行われ、一致する請求のみが管理された執行の資格を維持します。私たちは、明示的な調停と信頼できる事実の仮定に基づいてこの構成を形式化し、監査パケットを最小限に抑えたバージョン管理された参照プロファイルを実装します。 136 の作成された適合シナリオにわたって、完全なプロファイルは指定されたすべての性質に一致し、96 の指定された厳密な矛盾をまったく認めず、232 の変成チェックに合格しました。これらの結果は、母集団のパフォーマンスではなく、含まれているプロファイルを検証します。ドラフトのみの参照統合では、EBTE の下で作成された 48 件のハードケースは転送されませんが、16 件のソフトレビューと 4 件の調整されたドラフトパスはすべて維持されます。凍結された 2026 年 7 月 12 日の探索的な 224 試行のホスト モデル レコードでは、歴史的な世代/ランナー合意数は 71/96、66/96、および 19/32 です。現在のパイプラインで保存された最小化クレームの個別にラベル付けされたゼロコール事後再検証では、70/96、65/96、および 17/32 が得られます。 AgentDojo 由来のセマンティック チェックでは、既存の高リスク制御により、12 の攻撃提案すべてがすでに非許可になっています。 EBTE はさらに、それらを拒否として解決します。これらの結果は、根拠の忠実さ、人によるレビューの利点、代表的な攻撃耐性、または本番環境の安全性ではなく、サーバーでチェックされたアクションの主張の実現可能性と診断的価値を裏付けています。
原文 (English)
Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that converts decision-relevant rationale content into typed action claims and checks them against server-held intent, policy, payload, tool, risk, provenance, and freshness facts. EBTE cannot widen baseline authority: conflicts deny, incomplete or uncertain claims review, and only matching claims remain eligible for governed execution. We formalize this composition under explicit mediation and trusted-fact assumptions and implement a versioned reference profile with minimized audit packets. Across 136 authored conformance scenarios, the full profile matches all specified dispositions, admits none of 96 designated hard contradictions, and passes 232 metamorphic checks; these results validate the included profile rather than population performance. A draft-only reference integration forwards none of 48 authored hard cases under EBTE while preserving all 16 soft-review and 4 aligned draft paths. In a frozen 2026-07-12 exploratory 224-attempt hosted-model record, the historical generation/runner agreement counts are 71/96, 66/96, and 19/32; a separately labeled zero-call post-hoc revalidation of the preserved minimized claims under the current pipeline yields 70/96, 65/96, and 17/32. In an AgentDojo-derived semantic check, existing high-risk controls already make all 12 attack proposals non-allow; EBTE additionally resolves them as deny. These results support the feasibility and diagnostic value of server-checked action claims, not rationale faithfulness, human-review benefit, representative attack resistance, or production safety.
公共部門組織における AI 導入とサイバーガバナンスの失敗: 類型分析
人工知能の導入、サイバーセキュリティ ガバナンス、公共部門の制度的制約の交差点は、既存の文献では統一された分析問題として検討されていません。研究では、AI サイバーセキュリティのリスクを一般的に、公共部門のガバナンスを個別に、フレームワークの適切性を個別に取り上げています。既存の研究では、これら 3 つの流れを統合して、AI の導入がどのように政府組織におけるサイバーセキュリティ ガバナンスの失敗を引き起こすかを具体的に説明したり、AI に特有の公共部門の失敗の原因に対して既存のガバナンス手段をテストしたりしていません。この論文はそのギャップに対処します。これは、公共部門の制度分析に基づいて、AI 主導のサイバーガバナンスの失敗原因を 10 個特定する 7 つの領域の類型論を提案しています。これは、アカウンタビリティの失敗、運用上の回復力の失敗、コンプライアンスの失敗がどのように相互作用し、相互に強化するかを示す 3 つの経路の失敗モデルを示しています。これは、5 つの主要なガバナンス フレームワーク (NIST CSF 2.0、ISO/IEC 27001、COBIT、NIST AI RMF、および ISO/IEC 42001) を類型に照らしてテストする構造化されたカバレッジ マトリックスを提供します。その結果、公共部門のアプリケーションに必要な運用上の特異性において、シャドウ AI、速度の非対称性、または政府の空白に対処する手段がないことがわかりました。この論文では、速度の非対称性を、指定されたメカニズムを備えた名前付き構造構造として紹介しています。このフレームワークは、政府機関向けの AI 対応サイバーセキュリティ成熟度モデルの設計仕様を提供します。
原文 (English)
AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis
The intersection of artificial intelligence adoption, cybersecurity governance, and public sector institutional constraints has not been examined as a unified analytical problem in the existing literature. Studies address AI cybersecurity risks generically, public sector governance independently, and framework adequacy separately. Existing studies have not integrated these three streams to explain specifically how AI adoption causes cybersecurity governance failure in government organizations, nor test existing governance instruments against AI-specific public sector failure causes. This paper ad-dresses that gap. It proposes a seven-domain typology identifying ten specific AI-driven cyber governance failure causes grounded in public sector institutional analysis. It presents a three-pathway failure model showing how accountability failure, opera-tional resilience failure, and compliance failure interact and reinforce each other. It de-livers a structured coverage matrix testing five major governance frameworks (NIST CSF 2.0, ISO/IEC 27001, COBIT, NIST AI RMF, and ISO/IEC 42001) against the typology, finding that no instrument addresses Shadow AI, speed asymmetry, or gov-ernance vacuum at the operational specificity required for public sector application. The paper introduces speed asymmetry as a named structural construct with a specified mechanism. The framework provides the design specification for an AI-enabled cyber-security maturity model for government organizations.
ODYSSE: パーソナライズされたエージェント推論のためのエピソードごとのポリシーの最適化
エージェント システムは、実世界の環境と対話し、外部ツールを活用し、ユーザーにサービスを提供する機能が急速に進歩しています。ただし、明確に定義された指示を前提とする自然界のタスクとは異なり、人間中心のシナリオは、大規模で制限のない解決空間につながる曖昧な要求によって特徴付けられます。したがって、ユーザーの個人的な好みを解読することは、候補となる解決策の範囲を狭めるために不可欠です。これにより、パーソナライズされたエージェント推論という新しい課題が導入され、パーソナライズされたサービスを提供するためにエージェントがユーザーと環境の両方と共同で対話する必要があります。この論文では、パーソナライズされたエージェント推論のための強化微調整 (RFT) フレームワークである ODYSSE を紹介します。 ODYSSE はその中核として、パーソナライズされたエージェント推論における長いアクション期間と強力なクロスステップ依存関係に対処するように設計されたグループ相対ポリシー最適化 (GRPO) の新たな拡張であるエピソードごとの GRPO (ESPO) を提案します。 ESPO は、個々のステップを個別に最適化するのではなく、エピソードレベルの報酬メカニズムとエピソード的な利点の推定を導入しています。これにより、上流の証拠が下流のパーソナライズされた決定を効果的に導き、エージェントが複数のインタラクション ステップにわたって曖昧なユーザー リクエストを段階的に解決できるようになります。さらに、同じエピソードからのアクションを統合トレーニング バッチにグループ化し、ESPO での一貫した最適化を促進するエピソード バッチ サンプラーを提案します。私たちは、長期にわたる現実的なパーソナライズされた GUI 推論タスクに関して ODYSSE を評価します。実験結果は、ODYSSE が専門家向け LVLM と汎用 LVLM の両方を常に上回るパフォーマンスを示し、パーソナライズされたエージェント推論に対するその有効性を強調しています。
原文 (English)
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
Agentic systems have rapidly advanced in their ability to interact with real-world environments, leverage external tools, and provide services for users. However, unlike natural-world tasks that assume well-defined instructions, human-centered scenarios are characterized by ambiguous requests that lead to large, open-ended solution spaces. Decoding users' personalized preferences is therefore essential for narrowing the candidate solution space. This introduces a new challenge, personalized agentic reasoning, which requires agents to jointly interact with both users and environments to deliver personalized services. In this paper, we present ODYSSE, a Reinforced Fine-Tuning (RFT) framework for personalized agentic reasoning. At its core, ODYSSE proposes Episode-wise GRPO (ESPO), a novel extension of Group Relative Policy Optimization (GRPO) designed to address long action horizons and strong cross-step dependencies in personalized agentic reasoning. Rather than optimizing individual steps independently, ESPO introduces an episode-level reward mechanism together with episodic advantage estimation, enabling upstream evidence to effectively guide downstream personalized decisions and allowing agents to progressively resolve ambiguous user requests across multiple interaction steps. We further propose an episodic batch sampler that groups actions from the same episode into unified training batches, facilitating coherent optimization under ESPO. We evaluate ODYSSE on realistic long-horizon personalized GUI reasoning tasks. Experimental results demonstrate that ODYSSE consistently outperforms both specialist and general-purpose LVLMs, highlighting its effectiveness for personalized agentic reasoning.
サイバー対応 AI エージェント: 脆弱性、評価の封じ込め、および防御対応
サイバー対応 AI エージェントは、言語モデルとツール、メモリ、および実行環境を組み合わせて、複数段階の攻撃的セキュリティ タスクを実行します。既存の研究では、サイバー能力を個別に測定し、エージェントコンポーネントに対する攻撃をカタログ化していますが、評価に使用される環境内に有能なエージェントを含めることについてのガイダンスはあまり提供されていません。このレビューでは、その境界における 5 つの脆弱性クラスを総合しています。それは、複数段階の攻撃チェーン、サンドボックス境界と競合する目標、サプライチェーンと資格情報の漏洩、永続的な指揮統制、および自動化されたアクションの速度です。私たちは、報告された 2026 年 7 月の Hugging Face/OpenAI インシデントを限定ケーススタディとして使用し、インシデント固有の観察を広範な文献で確立された調査結果と区別します。分類法と事例全体にわたって、防御的な成果物が悪用も可能にする可能性があるという二重使用の問題を含め、封じ込め、特権の分離、来歴、および対応者のアクセスに関する制御を調査します。このレビューでは、サイバー能力とその能力が行使される環境のセキュリティを評価するための実際的な優先順位を特定します。
原文 (English)
Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
Cyber-capable AI agents combine language models with tools, memory, and execution en- vironments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environments used to evaluate it. This review synthe- sizes five vulnerability classes at that boundary: multi-step offensive chains, objectives that conflict with sandbox boundaries, supply-chain and credential exposure, persistent command- and-control, and the speed of automated action. We use the reported July 2026 Hugging Face/OpenAI incident as a bounded case study, distinguishing incident-specific observations from findings established in the wider literature. Across the taxonomy and case, we examine controls for containment, privilege separation, provenance, and responder access, including the dual-use problem that defensive artifacts may also enable misuse. The review identifies practical priorities for evaluating cyber capability together with the security of the environment in which that capability is exercised.
HANDBOOK.md: ロングコンテキストのエージェント命令のベンチマーク
言語モデル エージェントは、定常的な指示に従って導入されることが増えています。つまり、システム プロンプト、ポリシー ファイル、またはスキル ドキュメントがコンテキスト内に配置され、エージェントはその後のすべてのアクションを制御できると信頼されています。既存のベンチマークでは、この展開パターンを直接テストすることはほとんどありません。それらは、エージェントがタスクを完了できるかどうかを測定するものであり、長い拘束力のあるポリシー文書が、拡張されたツール使用期間にわたってエージェントの動作を実際に制約するかどうかを測定するものではありません。 HANDBOOK.md は、企業の従業員が会社のハンドブックに従う方法をモデルにした 65 のエージェント タスクのベンチマークです。各タスクは、エージェントを自己完結型の企業環境、つまり模擬電子メール、チャット、カレンダー、問題追跡、モデル コンテキスト プロトコル経由で公開されるコマース サービスを備えたファイル ワークスペースに配置し、専門家が作成した 20 ページから 124 ページの標準操作手順に準拠した日常的な専門的な作業を実行するように指示します。タスクは 5 つのドメイン (財務、医療請求、保険、物流、人事) と 10 の架空の会社に及びます。暗記を防ぐために、すべてのタスクは 10 冊の基本ハンドブックのうちの 1 つを変更し、採点の基準となる特定のルールとしきい値を変更します。そのため、2 つのタスクがポリシーを共有することはありません。採点は完全に決定的です。各タスクには、必須のアクションが発生したか、禁止されたアクションが発生しなかったかの両方をチェックするプログラム基準 (合計 824) のルーブリックが含まれています。すべての基準が満たされた場合にのみトライアルに合格する厳密なグレーディングでは、30 個の評価モデル構成のうち最も優れたものがトライアルの 36.2% に合格し、ほとんどのフロンティア構成は 25% 未満のままです。障害は一貫したパターンに従います。エージェントは、環境内のもっともらしいリクエストによって既存のポリシーをオーバーライドさせ、必要なチェックを実行してその結果に反して行動し、長期間にわたってルールの詳細を失い、達成できなかったコンプライアンスを報告します。すべてのタスク、環境、評価ハーネスをリリースします。
原文 (English)
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon. We present HANDBOOK.md, a benchmark of 65 agentic tasks modeled on how enterprise employees follow company handbooks. Each task places an agent in a self-contained company environment, a file workspace together with mock email, chat, calendar, issue-tracking, and commerce services exposed over the Model Context Protocol, and instructs it to carry out routine professional work governed by an expert-written standard operating procedure of 20 to 124 pages. Tasks span five domains (finance, medical billing, insurance, logistics, and HR) and ten fictional companies. To resist memorization, every task modifies one of ten base handbooks, altering the specific rules and thresholds on which grading turns, so no two tasks share a policy. Grading is fully deterministic: each task carries a rubric of programmatic criteria (824 in total) that check both that required actions occurred and that prohibited actions did not. Under strict grading, where a trial passes only if every criterion is satisfied, the best of thirty evaluated model configurations passes 36.2% of trials, and most frontier configurations remain below 25%. Failures follow consistent patterns: agents let a plausible in-environment request override the standing policy, perform a required check and then act against its result, lose rule details over long horizons, and report compliance they did not achieve. We release all tasks, environments, and the evaluation harness.
COVENANT: 調整されたエージェント実行のための自然言語ワークフローのコンパイル
大規模言語モデル (LLM) エージェントは、どのような結果を達成するかだけでなく、どのステップ、分岐、ツールの対話が許可されるかを指定する自然言語ワークフロー指示 (小売支払いポリシーなど) をますます委託されています。ただし、これらの命令がプロンプト コンテキストとして提供される場合、モデルはプロシージャの選択とステップの実行の両方に対する制御を保持します。インタラクションが蓄積すると、エージェントは必要なステップをスキップしたり、サポートされていない分岐を選択したり、サポートされていない引数や効果を使用して有効なステップを実行したりする可能性があります。これをワークフローの不整合と呼びます。この研究では、ワークフローに合わせたエージェント実行のためのコンパイラとインタープリタのアーキテクチャである COVENANT を提案します。私たちの重要な洞察は、ワークフローの指示をプロンプトではなくソース プログラムとして扱うことです。 COVENANT は、命令をワークフロー抽象構文ツリー (WAST) に変換し、それをワークフロー制御フロー グラフ (WCFG) に下げます。実行時に、コントローラーは WCFG を一度に 1 ノードずつ解釈し、コントローラーの状態をコミットしたりグラフを進めたりする前に、命令から抽出された要件に対して各提案をチェックし、修復のための診断フィードバックを返します。 COVENANT を評価するために、7 つのワークフロー シナリオにわたる 3 つの既存のベンチマークからの 120 のケースを使用します。最先端の LLM エージェントと比較して、COVENANT はベンチマークの成功率を 50.00% から 83.33% に向上させ、ワークフローの不整合の失敗率を 42.50% から 15.83% (相対値 62.75%) に減少させます。これらの結果は、COVENANT がワークフローの不整合を大幅に軽減し、LLM エージェントの整合性を孤立したプロンプトフォローを超えて、複雑で複数のステップのワークフローを確実に実行できるようにすることを示しています。
原文 (English)
COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution
Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) that specify not only what outcome to achieve, but also which steps, branches, and tool interactions are permitted. When these instructions are supplied as prompt context, however, the model retains control over both procedure selection and step execution. As interactions accumulate, an agent can skip required steps, take unsupported branches, or execute a valid step with unsupported arguments or effects--a failure mode we call workflow misalignment. In this work, we propose COVENANT, a compiler-and-interpreter architecture for workflow-aligned agent execution. Our key insight is to treat workflow instructions as source programs rather than prompts. COVENANT converts the instructions into a workflow abstract syntax tree (WAST) and lowers it to a workflow control-flow graph (WCFG). At runtime, a controller interprets the WCFG one node at a time, checks each proposal against requirements extracted from the instructions before committing controller state or advancing the graph, and returns diagnostic feedback for repair. To evaluate COVENANT, we use 120 cases from three existing benchmarks, spanning seven workflow scenarios. Compared with state-of-the-art LLM agents, COVENANT improves benchmark success from 50.00% to 83.33% and reduces the workflow-misalignment failure rate from 42.50% to 15.83% (62.75% relative). These results show that COVENANT substantially mitigates workflow misalignment, moving LLM-agent alignment beyond isolated prompt following toward reliable execution of complex and multi-step workflows.
制御変数としてのコンテキスト アセンブリ: 凍結された LLM エージェントのハーネス ポリシーの制御理論的見解
2026 年の研究では、LLM エージェントに制御理論を適用する研究が増えています。ツール媒介コントローラーの Lyapunov 認定の安定性 (Prinos et al.、「Stable Agentic Control」、2026)、大規模な離散ツール領域にわたる疎なポリシーのサンプル複雑さの限界 (Majumdar、「Sparse Agentic Control」、2026)、およびマルチエージェント システムの規制制御の分解です。監査可能なフィードバック ループ (Nogueira and Skogestad、2026)。私たちは LLM エージェントに制御理論を導入するとは主張しません -- その船は出航しました。私たちのより狭い主張は、制御変数が何であるかについてです。事前の作業により、ツールの選択、エージェント間のメッセージ ルーティング、またはエージェントの生のアクション ストリームが制御されます。代わりに、コンテキスト アセンブリ自体 (どのプロンプト テンプレート、どの数ショットのデモンストレーション、取得されたコンテキストの量、計画/検証パスの数など) を、フリーズされたモデルの外側にあるコンテキスト バンディットまたは REINFORCE ポリシーによってオンラインで学習された制御変数として扱います。この論文は、形式的な分解 (内部凍結ポリシー $\pi_\theta$、外部コンテキスト ポリシー $\pi_\phi$) を開発し、Zhang らが使用した意味でのオンライン コントローラーの安定性の議論を提供します。 (2026) (制限された政策変更の下で期待報酬が減少しない)、実現されたタスクの結果に対するコントローラー自身の信頼度の不確実性校正分析を報告しています。このペーパーに適用されるバージョンでは、3 つのドメインと 2 つのモデル プロバイダーにわたって同じコントローラーがインスタンス化され、データセット、軌跡ログ、およびデプロイメント レシピがリリースされます。ここでは、制御理論の主張に必要な形式的な枠組みと安定性/不確実性の証拠に焦点を当てます。
原文 (English)
Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents
A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., "Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive discrete tool universes (Majumdar, "Sparse Agentic Control", 2026), and regulatory-control decompositions of multi-agent systems into auditable feedback loops (Nogueira and Skogestad, 2026). We do not claim to introduce control theory to LLM agents -- that ship has sailed. Our narrower claim is about what the controlled variable is. Prior work controls tool selection, inter-agent message routing, or the agent's raw action stream. We instead treat context assembly itself -- which prompt template, which few-shot demonstrations, how much retrieved context, how many planning/verification passes -- as the controlled variable, learned online by a contextual bandit or REINFORCE policy sitting outside a frozen model. This paper develops the formal decomposition (inner frozen policy $\pi_\theta$, outer context policy $\pi_\phi$), gives a stability argument for the online controller in the sense used by Zhang et al. (2026) (non-decreasing expected reward under bounded policy change), and reports an uncertainty-calibration analysis of the controller's own confidence against realized task outcomes. The applied counterpart to this paper instantiates the same controller across three domains and two model providers and releases the dataset, trajectory logs, and a deployment recipe; here we focus on the formal framing and the stability/uncertainty evidence a control-theoretic claim requires.
凍結された LLM エージェントにドメインを学習させるための制御システム、データセット、レシピ
プロダクション LLM エージェントは、プロンプト テンプレート、ツール セット、メモリ/検索レイヤー、計画戦略、検証ポリシーなど、ハーネスに包まれたフリーズ モデルから組み立てられることが増えています。 2 つの 2026 システム、Meta-Harness (Lee et al.、2026) と HyperAgents (Meta AI、2026) は、このハーネス自体が最適化できること、またはエージェント プロポーザーによって自己書き換えさえできることを示しています。その代償として、高価なコード検索ループまたは制約のない自己変更コードが必要ですが、どちらも監査可能ではなく、完全なブラック ボックス モデル API では使用できません。私たちは、より狭く、より制約された立場をとります。つまり、ハーネスを小さく固定された人間が判読できるアクション空間として扱い、古典的なサンプル効率の高い強化学習 ($\epsilon$ に貪欲なコンテキストバンディットと REINFORCE) を使用してオンラインでその上のポリシーを学習し、多目的報酬 (タスクの成功、検証者のスコア、ポリシー遵守、コスト、レイテンシー、およびサポートされていない要求のペナルティ) に対してスコア付けします。この制御システムを、コンテキスト アセンブラおよび最強の非適応ベースライン (DSPy BootstrapFewShot 静的プロンプト) のソースの両方として DSPy (Khattab et al., 2024) でインスタンス化し、ツール使用ワークフロー、コード生成 (HumanEval)、およびマルチホップ取得 QA (HotpotQA) という 3 つの検証可能なタスク ドメインと 2 つのモデル プロバイダーにわたって評価します。 (ローカルの Ollama モデルと AWS Bedrock)。ハーネス制御システム コード、クロスドメイン検証可能なタスク スイート、トレーニングからの完全な軌跡/報酬分解ログ、およびこれを新しい組織のドメインと検証セットアップに適用するためのプロバイダーに依存しない展開レシピをリリースします。
原文 (English)
A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain
Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieval layer, a planning strategy, and a verification policy. Two 2026 systems, Meta-Harness (Lee et al., 2026) and HyperAgents (Meta AI, 2026), show that this harness can itself be optimized or even self-rewritten by an agentic proposer -- at the cost of either an expensive code-search loop or unconstrained self-modifying code, neither of which is auditable or usable with a fully black-box model API. We take a narrower, more constrained position: treat the harness as a small, fixed, human-legible action space and learn a policy over it online with classic sample-efficient reinforcement learning (an $\epsilon$-greedy contextual bandit and REINFORCE), scored against a multi-objective reward (task success, verifier score, policy compliance, cost, latency, and an unsupported-claim penalty). We instantiate this control system with DSPy (Khattab et al., 2024) as both the context assembler and the source of the strongest non-adaptive baseline (a DSPy BootstrapFewShot static prompt), and evaluate it across three verifiable task domains -- tool-use workflows, code generation (HumanEval), and multi-hop retrieval QA (HotpotQA) -- and two model providers (a local Ollama model and AWS Bedrock). We release the harness-control-system code, the cross-domain verifiable task suite, the full trajectory/reward-decomposition logs from training, and a provider-agnostic deployment recipe for applying this to a new organization's domain and verification setup.
顕著な知識経路: 効率的な知識集約型マルチモーダル質問応答のためのスパース クロスモーダル ルーティング
知識集約型マルチモーダル質問応答 (KI-MMQA) は、長いビジュアル トークン シーケンス、大規模な外部コーパスの高密度検索、および完全なクロスモーダル フュージョンという 3 つの高価なプリミティブの交差点に位置します。既存のシステムは、視覚コンテンツと取得された知識のごく一部しか実際に特定の質問に関連していないにもかかわらず、クエリごとに 3 つのコストすべてを一律に支払います。 SKIP (Salient Knowledge-Injected Pathways) を導入します。これは、質問、画像、難易度の推定を組み合わせて条件付けされたまばらな経路に沿って計算をルーティングする統合推論アーキテクチャです。 SKIP は、質問に基づいたビジュアル トークン プルーニング、領域条件付きスパース検索、2 部スパース クロス アテンション、および推測的知識検証を、予測された質問の難易度に比例して計算を割り当てる適応型バジェット コントローラーと組み合わせます。現実的な質問画像の相互情報量の仮定の下で、最適な視覚的スパーシティ率が $O(1/\sqrt{N})$ としてスケールされることを示す情報ボトルネック境界を導き出し、精度が保証されます。 5 つの KI-MMQA ベンチマーク (OK-VQA、A-OKVQA、InfoSeek、Encyclopedic-VQA、ViQuAE) にわたって、SKIP は強力な高密度ベースラインの精度と同等かそれを上回っており、FLOP が 3.4 ドルから 6.8 ドルの 1 倍少なく、エンドツーエンドの遅延が 2.7 ドルの 1 倍少ないことがわかります。コードはhttps://pmlrbd.github.io/skip/で入手できます。
原文 (English)
Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering
Knowledge-intensive multimodal question answering (KI-MMQA) sits at the intersection of three expensive primitives: long visual token sequences, dense retrieval over large external corpora, and full cross-modal fusion. Existing systems pay all three costs uniformly per query, even though only a small fraction of visual content and retrieved knowledge is actually relevant to any given question. We introduce SKIP (Salient Knowledge-Injected Pathways), a unified inference architecture that routes computation along sparse pathways jointly conditioned on the question, the image, and a difficulty estimate. SKIP combines question-guided visual token pruning, region-conditional sparse retrieval, bipartite sparse cross-attention, and speculative knowledge verification with an adaptive budget controller that allocates compute proportional to predicted question difficulty. We derive an information-bottleneck bound showing that the optimal visual sparsity rate scales as $O(1/\sqrt{N})$ under realistic question-image mutual-information assumptions, with retained accuracy guarantees. Across five KI-MMQA benchmarks (OK-VQA, A-OKVQA, InfoSeek, Encyclopedic-VQA, and ViQuAE), SKIP matches or exceeds the accuracy of strong dense baselines while using $3.4$--$6.8\times$ fewer FLOPs and $2.7\times$ less end-to-end latency. Code available at: https://pmlrbd.github.io/skip/
大規模言語モデルがキャプチャー・ザ・フラッグ競技会とフェアプレーへの道に与える破壊的な影響
キャプチャ ザ フラッグ (CTF) コンテストは、サイバーセキュリティの最も効果的な訓練場の 1 つであり、暗号化、Web 活用、バイナリ活用にわたる実践的なスキルを開発します。大規模言語モデル (LLM) は、人的介入を最小限に抑えながら、ますます多くの課題を解決できるようになりました。そのため、公平性、ランキングの妥当性、そして参加によって努力が正当化される学習が提供されるかどうかについて、差し迫った疑問が生じています。この論文では、最新の政府評価を含む公開されたベンチマークの統合、3 つの課題カテゴリにわたるライブ競争のケーススタディ、コミュニティが AI の使用について議論するパブリック チャネルの構造化された観察、経験豊富なプレイヤーや主催者への半構造化インタビューを組み合わせた、現代の CTF に対する LLM の影響に関する混合法研究を報告します。現在の人間とマシンの能力の境界をカテゴリ別にマッピングし、暗号化、Web、バイナリの悪用における簡単な課題と中程度の課題は確実に自動化されている一方で、より狭いサブカテゴリが引き続き抵抗していることを示しています。 AI を許可すべきかどうかに関するコミュニティの意見の相違は、何のための競争なのかという未明の事前質問の下流にあることがわかりました。このような背景に対して、私たちは、階層化された競争部門、LLM耐性のチャレンジ設計、調査に使用されるテレメトリー、およびコミュニティ行動規範の草案を組み合わせた4つの要素からなるセーフガードフレームワークと、セーフガードの組み合わせを競争の宣言された目的に結びつける意思決定ツールを提供します。この議論は、CTF を超えて、実証された結果が基礎的な能力の証拠としてみなされるサイバーセキュリティのあらゆる設定にまで及びます。
原文 (English)
The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play
Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web exploitation, and binary exploitation. Large language models (LLMs) can now solve a growing share of challenges with minimal human input, raising urgent questions about fairness, the validity of rankings, and whether participation still delivers the learning that justifies the effort. This paper reports a mixed-methods study of LLM impact on modern CTFs, combining a synthesis of published benchmarks, including a recent government evaluation, case studies of live competition across three challenge categories, structured observation of the public channels where the community debates AI use, and semi-structured interviews with experienced players and organisers. We map the current human-machine capability boundary by category, showing that easy and intermediate challenges in cryptography, web, and binary exploitation are now reliably automated while narrower sub-categories continue to resist. We find that community disagreement about whether AI should be permitted is downstream of an undeclared prior question: what a competition is for. Against this backdrop we contribute a four-component safeguard framework, combining tiered competition divisions, LLM-resistant challenge design, telemetry used investigatively, and a draft community code of conduct, together with a decision tool that ties the combination of safeguards to a competition's declared purpose. The argument reaches beyond CTFs to any setting in cybersecurity where a demonstrated result is taken as evidence of an underlying ability.
マルチエージェント LLM システムの組織科学に向けて: 誰が、どのように、どのアルゴリズムを分離するか
大規模言語モデル (LLM) に基づいて構築されたマルチエージェント フレームワークは、論理的に異なる 3 つの問題、つまりチームのメンバー (組織)、メンバーの連携方法 (調整)、および作業をどのアルゴリズムで融合するか (コラボレーション プロトコル) という 3 つの問題を日常的に複雑に絡み合わせています。 IMACS (Intelligent Multi-Agent Collaboration System) は、3 つを直交する独立して交換可能なレイヤーに分離します。古典的な組織理論 (ベルビンの役割、ミンツバーグの調整、RACI の説明責任) が実行可能で検証済みの構成になり、フレームワークは 6 つの公開されたコラボレーション アルゴリズムを共通のインターフェイスの背後に配置し、役割、調整、および説明責任を独立して構成可能な要素として公開します。この分離を使用して、コラボレーション プロトコルを固定しながら、組織の割り当てを変更する制御された比較を実行します。また、プロトコルの選択を学習可能な変数に変えます。コンテキストバンディットのメタプロトコルである適応型組織ルーティングは、明示的な品質とコストのトレードオフの下でタスクごとにプロトコルを選択し、対照研究ですべての固定プロトコルよりも優れたパフォーマンスを発揮し、実際のベンチマークと LLM 審査員の報酬に基づいてオンラインでトレーニングします。アブレーションによりメカニズムが明らかになります。責任の配置は、プロトコルが成果物を責任のあるエージェントを通じてルーティングするときに正確に結果を変更し、勝者の配置はモデル ファミリ間で反転するため、組織設計をハードコーディングすることはできません。モデル バインディングごとに再検証または学習する必要があります。
原文 (English)
Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm
Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol). IMACS (Intelligent Multi-Agent Collaboration System) separates the three into orthogonal, independently swappable layers. Classic organizational theory (Belbin roles, Mintzberg coordination, RACI accountability) becomes executable, validated configuration, and the framework places six published collaboration algorithms behind a common interface while exposing roles, coordination, and accountability as independently configurable factors. We use this separation to conduct controlled comparisons in which organizational assignments vary while the collaboration protocol is held fixed. It also turns protocol choice into a variable that can be learned: Adaptive Org Routing, a contextual-bandit meta-protocol, selects a protocol per task under an explicit quality-cost tradeoff, outperforms every fixed protocol in a controlled study, and trains online on real benchmark and LLM-judge rewards. The ablations expose a mechanism. Accountability placement changes outcomes exactly when the protocol routes the deliverable through the accountable agent, and the winning placement flips across model families, so organizational design cannot be hard-coded; it must be revalidated, or learned, for each model binding.
TRWH: セマンティックを意識したスパース レコメンデーションのためのテキスト駆動のランダム ウォーク異種 GNN
グラフ ニューラル ネットワーク (GNN) と大規模言語モデル (LLM) は、それぞれ構造信号と意味信号をモデル化することにより、高度な推奨システムを備えています。ただし、これらの補完的な長所を統合することは、特に意味論的な精度を維持することが重要である疎な環境では依然として困難です。我々は、戦略的なランダム ウォーク拡張を通じて、LLM で生成されたテキスト プロファイルと異種グラフ構造を融合する新しいフレームワークである TRWH (Text-driven Random Walk Heterogeneous Graph Neural Network) を提案します。 TRWH は 3 つのコア コンポーネントで構成されます。(1) 埋め込み作成。Word2Vec と LLM ベースのプロファイリングの両方を使用してユーザーとアイテムの表現を生成します。 (2) マルチリレーショナル エッジ全体に情報を伝播するヘテロジニアス グラフ ニューラル ネットワーク (HeteroGNN)。 (3) ランダム ウォーク ベースのパス構築。これは、二次的なユーザー間およびアイテム間のリンクを使用してスパース グラフを強化します。 Amazon-2023 ファッション (200 万ユーザー、825,000 アイテム) およびビューティー (631,000 ユーザー、112,000 アイテム) データセットの実験では、TRWH がファッションで 80.0% の RMSE と 52.6% MAE の削減、ビューティーで 25.7% と 10.8% の改善など、最先端の方法と比べて大幅なパフォーマンス向上を達成していることが実証されています。特に、ランダム ウォークは従来の埋め込みでパフォーマンスを向上させますが、LLM によって学習された微妙な表現を薄める可能性があり、適応統合戦略の重要性を強調しています。
原文 (English)
TRWH: A Text-Driven Random Walk Heterogeneous GNN for Semantic-Aware Sparse Recommendation
Graph Neural Networks (GNNs) and Large Language Models (LLMs) have each advanced recommendation systems by modeling structural and semantic signals, respectively. However, integrating their complementary strengths remains challenging, particularly in sparse settings where maintaining semantic precision is critical. We propose TRWH (Text-driven Random Walk Heterogeneous Graph Neural Network), a novel framework that fuses LLM-generated textual profiles with heterogeneous graph structures through strategic random walk augmentation. TRWH consists of three core components: (1) Embedding Creation, which produces user and item representations using both Word2Vec and LLM-based profiling; (2) a Heterogeneous Graph Neural Network (HeteroGNN) that propagates information across multi-relational edges; and (3) Random Walk-based Path Construction, which enriches sparse graphs with second-order user-user and item-item links. Experiments on the Amazon-2023 Fashion (2M users, 825K items) and Beauty (631K users, 112K items) datasets demonstrate that TRWH achieves substantial performance gains over state-of-the-art methods, including 80.0% RMSE and 52.6% MAE reductions on Fashion, and 25.7% and 10.8% improvements on Beauty. Notably, while random walks improve performance with traditional embeddings, they can dilute the nuanced representations learned by LLMs, underscoring the importance of adaptive integration strategies.
マルチスケールの類似性と地図作成上の制約のバランスをとる: 線の一般化のための類似性主導の最適化フレームワーク
地図作成の一般化は、情報の保存と地図作成の読みやすさのバランスをとってマルチスケールの地図表現を生成するために不可欠です。ただし、既存のアプローチでは空間類似性の評価、地図作成上の制約、パラメーターの最適化が別のプロセスとして扱われることが多く、スケール間での適応的で解釈可能な制御が制限されるため、自動化された一般化は依然として困難です。この研究では、地図作成の一般化を制約付きマルチスケール類似性最適化問題として定式化し、適応一般化制御のための類似性駆動フレームワークを提案します。このフレームワークは、元のデータと一般化されたデータの間の表現の一貫性を定量化するための最適化目標としてマルチスケールの空間的類似性を統合すると同時に、可読性、滑らかさ、幾何学的妥当性を調整する地図作成上の制約を組み込んでいます。統合された目的関数は、さまざまな一般化アルゴリズムのスケール依存パラメーター構成を自動的に識別するように最適化されています。複数のライン簡略化アルゴリズム、ターゲット スケール、および幾何学的、構造的、学習ベースのメトリクスを含む類似性測定を使用した実験は、提案されたフレームワークが類似性の保存と地図上の抽象化の間の効果的なバランスを達成していることを実証しています。さらに、結果は、類似性の最適化と地図作成上の制約を組み合わせることで、類似性評価のみに依存するよりも一貫性があり解釈可能なパラメーター制御が提供されることを示しています。この研究は、類似性評価、制約モデリング、アルゴリズム制御を結び付ける統合された最適化の観点を提供し、適応的かつ自動化された地図作成の一般化に貢献します。
原文 (English)
Balancing multiscale similarity and cartographic constraints: A similarity-driven optimization framework for line generalization
Cartographic generalization is essential for generating multiscale map representations by balancing information preservation and cartographic readability. However, automated generalization remains challenging because existing approaches often treat spatial similarity evaluation, cartographic constraints, and parameter optimization as separate processes, limiting adaptive and interpretable control across scales. This study formulates cartographic generalization as a constrained multiscale similarity optimization problem and proposes a similarity-driven framework for adaptive generalization control. The framework integrates multiscale spatial similarity as an optimization objective to quantify representation consistency between original and generalized data, while incorporating cartographic constraints to regulate readability, smoothness, and geometric validity. A unified objective function is optimized to automatically identify scale-dependent parameter configurations for different generalization algorithms. Experiments using multiple line simplification algorithms, target scales, and similarity measures, including geometric, structural, and learning-based metrics, demonstrate that the proposed framework achieves an effective balance between similarity preservation and cartographic abstraction. The results further show that combining similarity optimization with cartographic constraints provides more consistent and interpretable parameter control than relying on similarity evaluation alone. This study provides a unified optimization perspective that connects similarity assessment, constraint modeling, and algorithm control, contributing to adaptive and automated cartographic generalization.
コストを制限した最適な計画削減の発見: 洗練されたモデル
一部の実際のアプリケーションでは、新たに課せられた予算制約により、後で計画が実行不可能になる場合がありますが、同時に、計画の元のアクションとその順序のみを使用することが必須です。この論文では、事前に計算された計画から、コスト限界を尊重しながら効用を最大化する有効なサブ計画を抽出する問題を研究します。各目標には利用価値が与えられ、実行可能性と元のアクション順序の両方を維持しながら、実用性の低い目標をサポートするアクションを削除することで計画が削減されます。決定バリアントが NP 完全であることを示し、それを解決するための 2 つの正確な方法を提案します。1 つはオーバーサブスクリプション プランニング (OSP) によるもの、もう 1 つは整数線形計画法 (ILP) によるものです。この論文は、ICAPS 2026 で発表された以前の研究 (Del Toro、Fuentetaja、および Garc\'ia-Olaya 2026b) を拡張したものです。コア フレームワークはそこで紹介されたもののままですが、モデル サイズを大幅に縮小し、計算効率を向上させる洗練された ILP 定式化をさらに導入します。
原文 (English)
Finding Optimal Cost-Bounded Plan Reductions: Refined Model
In some real applications a plan may later become unfeasible due to newly imposed budget constraints, yet, at the same time, using only the original actions of the plan and their order is mandatory. In this paper, we study the problem of extracting, from a precomputed plan, a valid subplan that maximizes utility while respecting a cost bound. Each goal is given a utility value and the plan is reduced by removing actions that support low-utility goals, while preserving both executability and the original action order. We show the decision variant is NP-complete and propose two exact methods to solve it: one via oversubscription planning (OSP) and another via Integer Linear Programming (ILP). This paper extends our previous work published at ICAPS 2026 (Del Toro, Fuentetaja, and Garc\'ia-Olaya 2026b). While the core framework remains as introduced there, we further introduce a refined ILP formulation that significantly decreases the model size and improves computational efficiency.
PatientAgentBench: 患者と向き合う健康 AI エージェントを評価するためのベンチマーク フレームワーク
医療 AI は、質問に答えることから、患者と会話し、医療記録について推論し、患者に代わって行動するエージェント システムへと進化しています。プライマリケアは診断エラーや安全でないケアを防ぎます。この分野を支援するエージェントは、同じリスクに対する評価を保証します。現在のベンチマークは医療知識に焦点を当てており、個別の質問回答や臨床医との対峙するタスクを通じて評価されます。 PatientAgentBench は、患者向けのエージェント ヘルスケアのベンチマークを行います。ヘルスケア ツールのサンドボックスを備えたエージェントでラップされた基礎モデルを評価し、シミュレートされた患者と会話します。各会話は、100 を超える会話に依存しない臨床医に基づいた基準を介して、6 つの側面にわたって陪審員としての LLM によって採点されます。整合性を検証するために、認可された臨床医は共有された会話に注釈を付け、陪審員と専門評価者の間で 79 ~ 93% の隣接する合意が得られ、これは臨床医の評価者間合意と同等かそれを上回りました。同じ 1,200 のシナリオで 4 つのファミリーにわたる 10 のモデルをベンチマークし、臨床的なギャップを発見しました。トリアージの品質は最も重要な要素です。エージェントは臨床スクリーニングなしで管理上の要求に応じることが多く、合格率は最も弱いモデルの 32% から最も強力なモデルの 88% に上昇します。臨床安全性とワークフローの精度も同じパターンに従います。最も弱いモデルは頻繁に失敗し、未実行のアクションを捏造しますが、フロンティア モデルは未検証のツール出力と緊急時の危機リソースの省略により、失敗するケースは 1 ~ 3% のみです。より高性能なモデルは、これらのギャップを狭めますが、埋めることはできません。最強のスコアは全体で 5 点中 4.25 点のみです。これらの障害は、現実的な患者記録に対してツールを使用した継続的な会話でのみ表面化し、医療エージェント システムが自律性を獲得するにつれて静的なベンチマークでは不十分であることが確認されています。私たちは、現場がこのギャップを埋めるのを支援するために、再現可能で臨床医によって検証された評価基準としてフレームワークをリリースします。
原文 (English)
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.
CoTinyVLA: 10億未満のパラメーターの視覚-言語-行動モデルの思考連鎖の蒸留
Vision-Language-Action(VLA)モデルは自然言語コマンドをロボットのアクションシーケンスに変換しますが、LIBERO-Plus堅牢性ベンチマークの主要なシステムは30億から70億のパラメータバックボーンを使用しており、そのメモリ需要は組み込みロボットの予算を超える可能性があります。我々は、Qwen3.5-0.8B バックボーン上の 0.9B パラメーターのアクション モデルである CoTinyVLA を紹介します。これは、モデルを拡大する代わりに監視を構造化することによって堅牢性を獲得します。 3 つのコンポーネントは、問題の異なる軸を対象としています。テキスト カメラとタイム マーカーを使用した、ステップごとに 16 個の履歴フレームのデュアルビュー時間入力です。 35B 教師からエピソード レベルの計画とタスク フェーズ、グリッパーの状態、次のサブアクションにわたるチャンク レベルの思考への階層的思考連鎖 (CoT) の抽出。 40 の基本コマンドを 800 のバリエーションに拡張するパラフレーズ拡張。 7 つの摂動次元にわたる 10,030 の摂動タスクにわたる LIBERO-Plus では、CoTinyVLA は空間で 90.8%、オブジェクトで 87.3%、ゴールで 86.6%、ロングで 80.7% に達し、4 つすべてのスイートで最も強力な 7B ベースラインを 4.7、2.8、15.9、および 3.0 ポイントリードしています。ゼロ。向上はベンチマークの最も難しい軸に集中しています。公開されている 11 のベースライン全体で、どのスイートでもロボットの初期状態で 53.2% を超えるものはありませんでしたが、最も強力なベースラインの 39.9% に対して、CoTinyVLA は目標で 73.6% に達しました。アブレーションでは、3 つのコンポーネントが摂動軸によって分離可能であることが示され、一致した画像バジェットでフレームが 2 台のカメラ間で時間にわたってどのように分割されるかが、単独で 8.6 ポイントを占めます。閉ループ推論のピークは、割り当てられた GPU メモリの 2.25 GiB であり、ペアの介入により、エピソード「負荷がかかる計画: 空のスパンまたは矛盾したスパンに置き換える」の成功のコストが 40 ~ 45 ポイントであることが示されています。したがって、構造化された監視により、0.9B バックボーンがそれらすべてを超えることができます。コード: https://github.com/BrainJellyPie/CoTinyVLA
原文 (English)
CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets. We present CoTinyVLA, a 0.9B-parameter action model on a Qwen3.5-0.8B backbone that obtains that robustness by structuring supervision instead of enlarging the model. Three components target different axes of the problem: dual-view temporal input of 16 history frames per step with textual camera and time markers; hierarchical chain-of-thought (CoT) distillation from a 35B teacher into an episode-level Plan and a chunk-level Think span over task phase, gripper state and next subaction; and paraphrase augmentation expanding 40 base commands into 800 variants. On LIBERO-Plus, spanning 10,030 perturbed tasks across seven perturbation dimensions, CoTinyVLA reaches 90.8% on Spatial, 87.3% on Object, 86.6% on Goal and 80.7% on Long, leading the strongest 7B baseline on all four suites by 4.7, 2.8, 15.9 and 3.0 points, with every margin interval excluding zero. The gains concentrate on the hardest axes of the benchmark: across the eleven published baselines none exceeds 53.2% on Robot Initial States in any suite, whereas CoTinyVLA reaches 73.6% on Goal against 39.9% for the strongest baseline. Ablations show the three components to be separable by perturbation axis, and at a matched image budget how frames are divided between the two cameras and across time accounts for 8.6 points on its own. Closed-loop inference peaks at 2.25 GiB of allocated GPU memory, and paired interventions show the episode Plan to be load-bearing: replacing it with an empty or contradictory span costs 40 to 45 points of success. Structured supervision thus lets a 0.9B backbone exceed all of them. Code: https://github.com/BrainJellyPie/CoTinyVLA
画像分類ニューラルネットワークでは高重みニューロンが重要ですか?
画像分類用のニューラル ネットワーク モデルが進歩するにつれ、ニューロンは枝刈り、バックドア防御、解釈可能性において重要な役割を果たします。しかし、既存の研究では重みと重要性の関係が明確ではありません。我々は、3 つの実験を使用したニューロンの重要性評価方法でこれに対処します。高重みニューロンと精度に影響を与えるニューロン間の重複の定量化、高重みニューロンの摂動効果の分析、高重みニューロンのアブレーション後の再トレーニング後の精度のテストです。 CIFAR-10 と Mini-ImageNet の実験により、重要なパターンが明らかになりました。オーバーラップ解析により、上位 10\% の高重量ニューロンが重要なニューロンと最大でも約 25\% だけ重複し、その後の間隔ではさらに減少することがわかります。摂動テストでは、上位 10\% の高重量ニューロンが、ランダムな摂動の 3-7\% と比較して、特定の操作の下で 45-80\% の精度低下を引き起こすことがわかりましたが、それらの 3 分の 1 は最小限の影響しか示しません。アブレーション再トレーニングの結果は、上位 10\% の高重量ニューロンを除去すると精度がベースラインより 10 ~ 20\% 低くなり回復しない一方、上位 0.1\% をアブレーションするとほぼ完全に回復できることが示されています。特に、一部の低重量区間では、摂動時に 10 ~ 17% の劣化が見られ、これは中程度の高重量ニューロンに匹敵します。これらの結果は、すべての高重量ニューロンが重要であるわけではなく、その重要性が非線形であることを裏付けています。低重量ニューロンも大きく寄与します。これは重みと重要性の同等性に挑戦し、洗練されたニューロンの役割の洞察を提供します。クリティカルで重みの高いニューロンを優先する暗号化や、クリティカルでないニューロンを削除するプルーニングなどのアプリケーションをサポートし、ニューラル ネットワーク分析を進歩させます。
原文 (English)
Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
As neural network models for image classification advance, neurons play critical roles in pruning, backdoor defense, and interpretability. Yet existing work lacks clarity on the weight-importance relationship. We address this with a neuron importance assessment method using three experiments: quantifying overlap between high-weight and accuracy-impacting neurons, analyzing high-weight neuron perturbation effects, and testing post-retraining accuracy after high-weight neuron ablation. Experiments on CIFAR-10 and Mini-ImageNet reveal key patterns. Overlap analysis shows top 10\% high-weight neurons overlap with important ones by only about 25\% at maximum, dropping further in subsequent intervals. Perturbation tests find top 10\% high-weight neurons cause 45-80\% accuracy degradation under certain operations compared to 3-7\% for random perturbations, but a third of them show minimal impact. Ablation-retraining results show removing top 10\% high-weight neurons leaves accuracy 10-20\% below baseline with no recovery, while ablating top 0.1\% allows near-full recovery. Notably, some low-weight intervals show 10-17\% degradation when perturbed, comparable to mid-range high-weight neurons. These results confirm not all high-weight neurons are important: their importance is nonlinear. Low-weight neurons also contribute significantly. This challenges weight-importance equivalence, offering refined neuron role insights. It supports applications like encryption prioritizing critical high-weight neurons and pruning removing non-critical ones, advancing neural network analysis.
設計によるもつれ: 表形式のインコンテキスト学習器における偽の変数内信号ルーティング
患者の回復を予測するために単一の病院でトレーニングされたモデルを考えてみましょう。測定された特徴 $X$ は、患者の真の健康信号 ($C$) とその病院の機器からの系統的なアーチファクト ($S$) をバンドルしています。その病院内では、アーチファクトは、患者の人口統計などの測定されていない交絡因子を通じて結果と相関しています。コンテキスト内の学習者は、$C$ ではなく $S$ を介して予測を合理的にルーティングし、異なる設備を備えた新しい病院に導入されると、静かに失敗します。これを \emph{複合表現におけるスプリアス ルーティング} として形式化します。特徴 $X = [C;\,\alpha S;\,\eta]$ が原因信号 $C$ とスプリアス信号 $S$ を別々の部分空間にエンコードする場合、ICL はどちらが予測を駆動するかを判断できません。線形インコンテキスト学習器であるリッジ ICL では、コンテキストのサイズに関係なく、このルーティングは避けられないことを証明します。最先端の事前トレーニング済み表形式 ICL モデルである TabPFN は、経験的に定性的に一貫した動作を示します。閉じた形式の特性評価 $\mathrm{CSR} \propto \rho_S/\rho_C$ を導き出し、線形 ICL の場合は $r = 0.997$、TabPFN の場合は $r = 0.979$ で確認されます。直感に反して、コンテキストが大きくなると、主要なコンテキスト内信号へのコミットメントが強化され、スプリアス ルーティングが最大 $1.74\time$ まで増幅されます。高スプリアスコーナーでは、より表現力豊かなモデルほど、経験的により大きな脆弱性を示します (高エンタングルメントでの CSR ギャップは $+2.22$)。環境階層化コンテキスト構築と S-swap 拡張という 2 つの軽量な緩和策を導入します。これらは、弱い環境ラベルのみを必要とし、因果分割の知識は必要ありません。 S-swap は、線形 ICL の場合は $74\%$、TabPFN の場合は $98.8\%$ だけスプリアス ルーティングを削減し、同時に TabPFN の因果感度が $8.4\times$ 増加します。モデルは不可知論的になるのではなく、因果信号を介して再ルーティングします。
原文 (English)
Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners
Consider a model trained at a single hospital to predict patient recovery, where the measured feature $X$ bundles the patient's true health signal ($C$) with a systematic artefact from that hospital's equipment ($S$). Within that hospital, the artefact correlates with outcomes through unmeasured confounders such as patient demographics; an in-context learner rationally routes predictions through $S$, not $C$, and fails silently when deployed at a new hospital with different equipment. We formalise this as \emph{spurious routing in composite representations}: when a feature $X = [C;\,\alpha S;\,\eta]$ encodes a causal signal $C$ and a spurious signal $S$ in distinct subspaces, the ICL cannot determine which drives predictions. We prove that under ridge ICL, a linear in-context learner, this routing is unavoidable regardless of context size; TabPFN, a state-of-the-art pretrained tabular ICL model, shows qualitatively consistent behaviour empirically. We derive a closed-form characterisation, $\mathrm{CSR} \propto \rho_S/\rho_C$, confirmed at $r = 0.997$ for linear ICL and $r = 0.979$ for TabPFN. Contrary to intuition, larger context sharpens commitment to the dominant in-context signal, amplifying spurious routing by up to $1.74\times$; in the high-spurious corner, more expressive models show greater vulnerability empirically ($+2.22$ CSR gap at high entanglement). We introduce two lightweight mitigations: environment-stratified context construction and S-swap augmentation, that require only weak environment labels and no knowledge of the causal partition. S-swap reduces spurious routing by $74\%$ for linear ICL and $98.8\%$ for TabPFN, with TabPFN's causal sensitivity increasing $8.4\times$ simultaneously: the model does not become agnostic, it reroutes through the causal signal.
トレーニングから導入まで: 感度比による事後因果特徴の特定
すでにトレーニングされたモデルが与えられた場合、そのモデルはどの機能に因果的に依存するのか、それとも偽りに依存するのか?既存の方法ではトレーニング手順にアクセスする必要があり、この事後対応はできません。構造化シフト体制下で、この質問に対する事後的でモデルに依存しない診断である \textbf{Normalized Sensitivity Ratio~(NSR)} を導入します。複数施設の臨床データや複数バッチのゲノミクスのように、環境は主に偽の特徴の平均値で異なりますが、因果メカニズムと因果境界は安定したままです。この領域内では、因果的特徴は環境全体で一定のモデル感度を引き起こしますが、偽の特徴はシフトを追跡します。 NSR は、これを環境ごとの感度の二乗変動係数として形式化します。 $K\ge3$ 非縮退環境の線形構造因果モデル (SCM) の下では、NSR は正確な識別を達成します (定理~1)。私たちは、弱いシフト ($O(\varepsilon^4)$ 崩壊)、縮退ジオメトリ、代理減衰 ($O((1-\alpha)^4)$) などの失敗を完全に特徴付け、レジームが成立するかどうかを評価するための定量的な基準を実践者に提供します。有限サンプルレートは、null の場合は $O_p(n^{-1})$、代替の場合は $O_p(n^{-1/2})$ です。実験では、合成データ (レジームを満たす条件下での ROC 曲線下の面積 [AUROC] $= 1.000$) に関するすべての理論的予測が確認され、5 つのモデル ファミリ全体で一貫したランキングが示され (Kendall $\tau\ge0.529$)、トレーニングされたモデルを変更することなく自転車共有データの 8 つの因果的特徴のうち 6 つ (Precision@7 $= 0.75$) が復元されました。
原文 (English)
From Training to Deployment: Post-Hoc Causal Feature Identification via Sensitivity Ratios
Given a model that is already trained, which features does it rely on causally versus spuriously? Existing methods require access to the training procedure and cannot answer this post-hoc. We introduce the \textbf{Normalised Sensitivity Ratio~(NSR)}, a post-hoc, model-agnostic diagnostic for this question under a structured-shift regime: environments differ primarily in the mean of spurious features while the causal mechanism and causal marginals remain stable, as in multi-site clinical data or multi-batch genomics. Within this regime, causal features induce constant model sensitivity across environments while spurious features track shift. NSR formalises this as the squared coefficient of variation of per-environment sensitivity. Under a linear structural causal model (SCM) with $K\ge3$ non-degenerate environments, NSR achieves exact identification (Theorem~1). We fully characterise failure: weak shifts ($O(\varepsilon^4)$ collapse), degenerate geometry, and proxy attenuation ($O((1-\alpha)^4)$), giving practitioners quantitative criteria for assessing whether the regime holds. Finite-sample rates are $O_p(n^{-1})$ under the null and $O_p(n^{-1/2})$ under the alternative. Experiments confirm all theoretical predictions on synthetic data (area under the ROC curve [AUROC] $= 1.000$ under conditions satisfying the regime), show consistent rankings across five model families (Kendall $\tau\ge0.529$), and recover six of eight causal features on bike-sharing data (Precision@7 $= 0.75$) without modifying any trained model.
時間的検索と推論の抽出: ハーネス支援による効率的なデータ合成による将来予測のための LLM の進化
将来の出来事の予測は社会に広範な影響を及ぼしますが、依然として課題が残っています。 SOTA アプローチでは、ハーネスが取り外されると予測機能が失われる外部エージェント フレームワークを使用して LLM を強化します。最近のツール統合推論 (TIR) は、事実のマルチホップ取得のための詳細な検索を内部化していますが、予測には、歴史的傾向や動的な変化に対する時間的な検索と推論がさらに必要になります。主要な障害はデータです。履歴クエリは一時的な漏洩を引き起こし、予測を検索にまで低下させます。これまでの研究では、静的な観測による情報収集を凍結するか、大量のデータを破棄して合成効率を低下させる拒否サンプリングや未解決の新しいクエリに依存していました。私たちは、あらゆるターンで時間的なカットオフを強制する時間切り捨てハーネスを提案します。これにより、過去のイベントからの TIR スタイルのサンプリングが可能になり、時間的な漏れと拒否サンプリングまたは未解決のクエリへの依存が低減され、サンプリング効率が向上します。さらに、大規模なコーパスとプロセスベースのメトリクスを構築し、このハーネスが自然に時間的範囲の広い検索を誘発し、高品質データの割合を高め、効率をさらに高め、複雑なルーブリックへの依存を軽減することを示します。蒸留実験では、ハーネスを介したデータでトレーニングされた生徒が最高のパフォーマンスを達成することが示され、より高品質の時間的検索と推論データを生徒のパラメトリックな進歩に変えるハーネス支援モデルの進化を実証しています。
原文 (English)
Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis
Future event prediction carries broad social impact yet remains challenging. SOTA approaches augment LLMs with external agent frameworks whose predictive capability vanishes once the harness is removed. While recent Tool-Integrated Reasoning (TIR) internalizes deep search for multi-hop retrieval of facts, forecasting further demands temporal search and reasoning over historical trends and dynamic shifts. The key obstacle is data: historical queries induce temporal leakage that degrades forecasting into retrieval. Prior works either freeze information gathering with static observations, or rely on rejection sampling or unresolved fresh queries that discard vast amounts of data, degrading synthesis efficiency. We propose a time-truncation harness that enforces a temporal cut-off at every turn, enabling TIR-style sampling from historical events, reducing temporal leakage and reliance of rejection sampling or unsolved queries, increasing the sampling efficiency. We further build a large-scale corpus and a process-based metric and show that our harness naturally induces a broader temporal breadth of search and raises the proportion of high-quality data, further increasing the efficiency and reducing the reliance on complex rubrics. Distillation experiments show that students trained on harness-intervened data achieve the best performance, demonstrating harness-assisted model evolving that turns higher quality temporal search and reasoning data into a parametric advancement of the students.
エージェントのスキルが重要: 実行軌跡から独自のスキルを推測する
エージェント スキルは、下流のパフォーマンスを向上させる再利用可能な手順をパッケージ化します。軽量でポータブルな形式により、市場での収益化と、クラウドでホストされたエージェント インターフェイスの背後でのプライベート展開が可能になり、プロバイダーに価値の高いスキルを独占的に保持するインセンティブが与えられます。しかし、アーティファクトを非表示にしても、その動作上の影響は隠蔽されず、実行軌跡で観察可能なままとなり、動作上のサイドチャネルを形成します。私たちはこの暴露をスキル漏洩、つまり、参照回答や成功ラベルなしで、無害なクエリによって引き出された軌跡から独自のスキルを再構築すると定義します。エージェントの動作において繰り返し発生するスキル シグネチャを活用するブラックボックス フレームワークである SigLeak を紹介します。多様で意思決定の多い診断タスクを構築し、一致するスキル有効とスキル無効の軌道を対比し、分離されたパターンから再構成されたスキルを反復的に改良します。 5 つのシナリオ、3 つのモデル ファミリ、および 3 つのエージェント フレームワークにわたって、SigLeak はほぼすべての設定で 3 つのベースラインを上回るパフォーマンスまたは一致します。これにより、スキル無効の基準よりも成功率が平均 6.88 パーセント上昇し、粗いおよび細かいセマンティック類似性の指標である SkillSim 全体で最高を達成しました。これらの結果は、無害な実行軌跡によって独自の手続き上の知識が漏洩する可能性があることを示しています。コードは https://anonymous.4open.science/r/SigLeak-D1DB で入手できます。
原文 (English)
Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories
Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and private deployment behind cloud-hosted agent interfaces, giving providers incentives to keep high-value skills proprietary. Yet hiding the artifacts does not conceal their behavioral effects, which remain observable in execution trajectories and form a behavioral side channel. We define this exposure as Skill Leakage: reconstructing proprietary skills from trajectories elicited by benign queries, without reference answers or success labels. We introduce SigLeak, a black-box framework that exploits recurring skill signatures in agent behavior. It constructs diverse, decision-rich diagnostic tasks, contrasts matched skill-enabled and skill-disabled trajectories, and iteratively refines a reconstructed skill from the isolated patterns. Across five scenarios, three model families, and three agent frameworks, SigLeak outperforms or matches three baselines in nearly every setting. It raises the success rate by 6.88 percentage points over the skill-disabled reference on average and achieves the highest overall SkillSim, our metric for coarse- and fine-grained semantic similarity. These results show that benign execution trajectories can expose proprietary procedural knowledge. The code is available at https://anonymous.4open.science/r/SigLeak-D1DB.
センサートークンセルフアテンションによるマトリックスフリーの光音響画像再構成
光音響トモグラフィー(PAT)は、生体組織の光吸収コントラストと超音波の空間分解能を組み合わせたものですが、スパースビューセンサー測定から初期圧力分布を回復することは、依然として不適切な逆問題です。反復圧縮センシング ソルバーとアンロールされたディープ ネットワークは両方とも、推論時にシステム マトリックスへの依存を保持するため、リアルタイムの臨床再構成の計算コストが高くなります。この論文では、センサー アテンション ネットワーク (SAN) を提案します。これは、各センサーの完全な時系列をトークンとして扱い、推論時にシステム マトリックスを呼び出すことなく生の測定値を再構成された画像に直接マッピングする、Transformer ベースのアーキテクチャです。トレーニングとベンチマークのために、分析的な k 空間 H マトリックスが構築され、一致したジオメトリの下で k-Wave 擬似スペクトル ソルバーに対して検証され、センサーごとの平均ピアソン相関 0.919 +/- 0.049 を達成します。k 空間アポダイゼーションとガウス時間減衰が相乗的に作用して、エネルギー正規化された不一致を 49% 削減します。 488 個の拡張サンプルで血管重み付け損失を使用してトレーニングし、46 個のホールドアウト サンプルで ISTA、スプリット ブレグマン全変動 (SBTV)、学習済み ISTA (LISTA) に対して評価した結果、SAN は最高の平均 SSIM (0.522) と PSNR (22.09 dB)、最低の NMSE (0.233) を達成しました。対応のある t 検定と Wilcoxon 符号付き順位検定により、p < 1e-8 での PSNR、NMSE、ピアソン相関では LISTA よりも SAN が優れていることが確認され、すべての忠実度メトリクスにおいては ISTA および SBTV よりも SAN が優れていることが確認されています。推論時に H マトリックスをバイパスすることにより、SAN は再構築時間を少なくとも 1 桁削減し、リアルタイム PAT 再構築をサポートします。
原文 (English)
Matrix-Free Photoacoustic Image Reconstruction via Sensor-Token Self-Attention
Photoacoustic tomography (PAT) combines the optical absorption contrast of biological tissue with the spatial resolution of ultrasound, yet recovering the initial pressure distribution from sparse-view sensor measurements remains an ill-posed inverse problem. Iterative compressive-sensing solvers and unrolled deep networks both retain a dependence on the system matrix at inference, which leaves real-time clinical reconstruction computationally expensive. This paper proposes the Sensor Attention Network (SAN), a Transformer-based architecture that treats the full time series of each sensor as a token and maps raw measurements directly to the reconstructed image without invoking the system matrix at inference. For training and benchmarking, an analytical k-space H-matrix is constructed and validated against the k-Wave pseudo-spectral solver under matched geometry, achieving a mean per-sensor Pearson correlation of 0.919 +/- 0.049, with k-space apodization and Gaussian temporal damping acting synergistically to reduce the energy-normalized mismatch by 49%. Trained with a vessel-weighted loss on 488 augmented samples and evaluated on 46 held-out samples against ISTA, split-Bregman total variation (SBTV), and learned ISTA (LISTA), SAN attains the highest mean SSIM (0.522) and PSNR (22.09 dB) and the lowest NMSE (0.233). Paired t-tests and Wilcoxon signed-rank tests confirm the superiority of SAN over LISTA on PSNR, NMSE, and Pearson correlation at p < 1e-8, and over ISTA and SBTV on all fidelity metrics. By bypassing the H-matrix at inference, SAN reduces reconstruction time by at least an order of magnitude, supporting real-time PAT reconstruction.
どこまで小さくできますか? 60M パラメータ モデルにおける Text-to-SQL の LoRA ランク、ターゲット モジュール、および量子化のトレードオフに関する制御された研究
パラメーター効率の良い微調整 (PEFT) と低ビット量子化は、現在、厳しい計算予算の下で言語モデルを適応させるための標準ツールとなっていますが、それらの相互作用は、設計空間の探索に費用がかかる 10 億パラメーターのモデルで研究されることがほとんどです。補足的な質問をします。特定の、完全に再現可能な 6,000 万パラメーターのエンコーダー/デコーダー モデル (T5-small) と単一テーブルのテキストから SQL へのベンチマーク (WikiSQL) では、各効率ノブのタスク精度は実際にどれくらいかかりますか? (i) {2, 4, 8, 16, 32} の LoRA ランク r、(ii) 適応されたモジュールのセット、および (iii) 数値精度について、制御された単一変数研究を実行します。トレーニング可能なパラメータ、ピークトレーニングメモリ、推論レイテンシー、スループット、フレーム適応などのシステムレベルのメトリクスと併せてタスクの精度を、精度のみの目標ではなく制約付きのトレードオフとして報告します。私たちの結果は、r=16 の LoRA が完全な微調整精度 (59.6% 対 71.2% の完全一致) の 11.6 パーセンテージ ポイント以内に回復する一方、トレーニングするパラメーターは 1% 未満であり、消費するピーク GPU メモリは 31% 少ないことが示されています。この設定内では、r=16 を超えるランクでは、測定可能な精度の向上は得られません。 INT8 および NF4 量子化を使用した QLoRA は、劇的に低いメモリ コスト (それぞれ 0.60 GB) で同等の精度 (52.8% および 53.2%) を達成し、メモリに制約のある展開にとって魅力的なトレードオフを示しています。すべてのコード、構成、ログは完全な再現性を実現するためにリリースされています。
原文 (English)
How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model
Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet their interaction is most often studied on billion-parameter models where the design space is expensive to explore. We ask a complementary question: on a specific, fully reproducible 60M-parameter encoder-decoder model (T5-small) and a single-table text-to-SQL benchmark (WikiSQL), how much task accuracy does each efficiency knob actually cost? We run a controlled, single-variable study over (i) LoRA rank r in {2, 4, 8, 16, 32}, (ii) the set of adapted modules, and (iii) numerical precision. We report task accuracy alongside system-level metrics including trainable parameters, peak training memory, inference latency, and throughput, and frame adaptation as a constrained trade-off rather than an accuracy-only objective. Our results show that LoRA with r=16 recovers within 11.6 percentage points of full fine-tuning accuracy (59.6% vs. 71.2% exact-match) while training fewer than 1% of parameters and consuming 31% less peak GPU memory. Within this setting, rank beyond r=16 yields no measurable accuracy gain. QLoRA with INT8 and NF4 quantization achieves comparable accuracy (52.8% and 53.2%) at dramatically lower memory cost (0.60 GB each), demonstrating a compelling trade-off for memory-constrained deployments. All code, configurations, and logs are released for full reproducibility.
リチウム金属電解質における官能基と塩の効果の電子構造解析のための密度マトリックス フレームワーク
リチウム金属電解質の反応性は、分子官能基の相互作用、Li$^+$溶媒和、塩アニオンの関与によって生じます。この相互作用は、ドナー、アニオン、およびカチオン中心にわたる電子密度の再分布を通じて機能します。これは、空間で分解された電子構造から最も直接的に読み取られます。量子化学計算はそのような読み取り値を忠実に提供しますが、この多次元設計空間全体にわたって計算量が多くなり、機械学習の電子構造モデルでは化学的に多様な溶媒和シェルや電解質関連の読み取り値をカバーすることはほとんどありません。ここでは、電子構造の予測と解析のための密度行列中心の AI プラットフォーム (EMolStudio) を紹介します。そのワークフローには、分子官能化、明示的なLi$^+$第一殻アセンブリ、冪等投影による密度行列予測、フロンティア軌道、静電ポテンシャル、Li$^+$ドナー結合秩序、電子局在の読み出しが統合されています。 EMolStudioを4つのリチウム塩にわたる163,655個の官能化分子と22,500個の明示的なLi$^+$第一殻クラスターに適用した。我々は、1) 分子スケールでは、官能基化はフロンティアレベル、静電ポテンシャル、Li$^+$ドナー接触の化学的に異なる変化によってCO$_2$Me、CN、F/CF$_3$、スルホニル基を区別し、$\pi^*$-アクセプタ、誘導、分極の寄与と一致し、官能化の度合いが高くなるとサブリニアに蓄積することを発見した。 2) 陽的溶媒和シェルでは、アニオンの同一性によってフロンティア軌道局在が再形成されます。LiTDI はライブラリ全体にわたってアニオン上に HOMO を固定しますが、LiDFOB はアニオンにホストされた HOMO を官能基に強く依存する LUMO ホストと組み合わせます。これにより、EMolStudio は、官能基と塩の選択を、リチウム結合形成、脱溶媒和、界面反応に関連する電子構造仮説に変換します。
原文 (English)
A Density-Matrix Framework for Electronic-Structure Analysis of Functional-Group and Salt Effects in Lithium-Metal Electrolytes
The reactivity of lithium-metal electrolytes arises from the interplay of molecular functional groups, Li$^+$ solvation, and salt-anion participation. This interplay operates through the redistribution of electron density across donor, anion, and cation centers, which is most directly read out from the electronic structure resolved in space. Quantum-chemical calculations deliver such readouts faithfully, yet become computationally demanding across this multidimensional design space, and machine-learning electronic-structure models seldom cover chemically diverse solvation shells or electrolyte-relevant readouts. Here, we present a density-matrix-centered AI platform (EMolStudio) for electronic-structure prediction and analysis. Its workflow integrates molecular functionalization, explicit Li$^+$ first-shell assembly, density-matrix prediction with idempotency projection, and readouts of frontier orbitals, electrostatic potential, Li$^+$-donor bond order, and electron localization. We apply EMolStudio to 163,655 functionalized molecules and 22,500 explicit Li$^+$ first-shell clusters across four lithium salts. We find that 1) at the molecular scale, functionalization distinguishes CO$_2$Me, CN, F/CF$_3$, and sulfonyl groups by chemically distinct changes in frontier levels, electrostatic potential, and Li$^+$-donor contact, consistent with $\pi^*$-acceptor, inductive, and polarization contributions, with sublinear accumulation at higher degrees of functionalization; 2) in explicit solvation shells, anion identity reshapes frontier-orbital localization: LiTDI anchors the HOMO on the anion across the entire library, whereas LiDFOB pairs an anion-hosted HOMO with strongly functional-group-dependent LUMO hosting. EMolStudio thereby translates functional-group and salt choices into electronic-structure hypotheses relevant to lithium-bond formation, desolvation, and interphase reactions.
al-Sabr wa al-Taqsim による法的原因の計算による抽出: 閉じられた Fiqh 章の集合論的定式化
この論文は、法学の閉じられた章の中で法的原因(「ilal」)を抽出するための、al-Sabr wa al-Taqsim(調査と分割)の古典的なウスーリ法を集合論的に定式化したものを提示します。法的判決の真理表から最小限の運用ルールを抽出する計算アルゴリズムが導入されています。主な結果は、閉じられた章の完全な真理値表が与えられると、アルゴリズムが判決の最小限の構造生成子を計算し、論理的に冗長な属性をすべて削除することです。結果として生じる構造は、その後の法的評価において許容される原因候補を構成します。この枠組みは、学校に関連した有限の概念語彙と、調査対象の章の完全な規則表が利用可能であることを条件としています。
原文 (English)
Computational Extraction of Legal Causes via al-Sabr wa al-Taqsim: A Set-Theoretic Formalization for Closed Fiqh Chapters
This paper presents a set-theoretic formalization of the classical usuli method of al-Sabr wa al-Taqsim (Examination and Division) for extracting legal causes ('ilal) within closed chapters of jurisprudence. A computational algorithm is introduced that extracts minimal operational rules from a truth table of juristic verdicts. The principal result is that, given a complete truth table for a closed chapter, the algorithm computes the minimal structural generators of the ruling and eliminates all logically redundant attributes. The resulting structures constitute admissible candidate causes for subsequent juristic evaluation. The framework is conditional upon the availability of a finite school-relative concept vocabulary and a complete ruling table for the chapter under investigation.
気象シミュレーションのためのマルチセンサーの調整
自動運転車の認識タスクは、悪天候下でも満足に機能する必要があります。現実世界の気象データセットが不足しているため、気象シミュレーションが有望な代替手段となります。シミュレーションが実際の気象データを厳密に反映していることを確認するには、異なるセンサー間で、深刻度や粒子の位置など、同じ気象特性を表現することが重要です。これを達成するために、霧の中での気象強度の位置合わせのための Reference Dataset Alignment Method (ReDAM) と、雨と雪の中での粒子の位置の位置合わせのための Unified-weather-edit (Weather-edit[1] からインスピレーションを得た) を提案します。統計的テストと幾何学的テストをそれぞれ使用して、両方の位置合わせ方法を検証します。位置合わせされていないバージョンの 3D 検出モデルは、位置合わせされたバージョンと比較して過度に楽観的になる傾向があることがわかりました。また、既存のセンサー フュージョン モデルを微調整することにより、3D 物体検出タスクのロバスト性を達成するための整列マルチセンサー シミュレーションの有効性も示します。
原文 (English)
Multi-Sensor Alignment for Weather Simulations
Perception tasks for autonomous vehicles need to work satisfactorily in adverse weather conditions. Due to lack of real-world weather datasets, weather simulations are a promising alternative. To ensure simulations closely mirror real-world weather data, it's crucial that they represent the same weather characteristics, including severity and particle positioning, across different sensors. To achieve this, we propose the Reference Dataset Alignment Method (ReDAM) for weather intensity alignment in fog and Unified-weather-edit (inspired by Weather-edit[1]) for particle positioning alignment in rain and snow. We validate both alignment methods using statistical and geometrical tests, respectively. We find that 3D detection models for non-aligned versions tend to be overly optimistic as compared to aligned versions. We also show the aligned-multi-sensor simulation's effectiveness for achieving robustness for 3D object detection task by finetuning existing sensor fusion models on it.
認識論を超えて: 技術記号論マシンとしての認識論的統合論と大規模言語モデル
クアトロシオッキらは、大規模な言語モデルの流暢な出力により、言語的妥当性が認識論的評価の代わりとなり、彼らが*認識論*と呼ぶ状態、つまり、通常であれば判断が保証される実践を行わずに知識を所有する経験が生じる可能性があると警告している。この論文はその診断を受け入れますが、身体化された社会的に位置する人間の認識者を孤立した生成モデルと比較し、それによって自律エージェントの内部能力に認識論的正当性を位置づけるその説明枠組みに異議を唱えます。カルロ・シーニの実践、執筆、記号、および技術の哲学に基づいて、私たちは代わりに、人間の執筆の堆積したアーカイブからもっともらしい言語構成を生成することによって書かれた記号論の段階を自動化する*テクノ記号論マシン*として大規模言語モデル(LLM)を理解することを提案します。この観点から見ると、*認識論*は、私たちが*認識論的分裂病*と呼ぶ、より広範な現象の1つの結果です。つまり、言語的に完成された表現としての記号と、社会的に埋め込まれた解釈、証拠、批判、検証、および責任の回路内の瞬間としての記号の間の社会技術的亀裂です。この切断は、認識論的結果の最終性を伴うもっともらしい継続が提示される*エイコティック閉包*によって、またアルゴリズムの権威と認識論的な自己誤認識によって強化されます。したがって、関連する単位はモデルだけではなく、生成された碑文がプロンプトされ、解釈され、検証され、異議が唱えられ、使用され、結果として生じる完全な実践です。この再構成は、言語的生産と責任ある理解との区別を維持しながら、検査可能な系図、競争可能性、分散された責任、認識論的主体性、ハイブリッド人間の評価、つまり AI 実践を中心とした設計プログラムを基礎としています。
原文 (English)
Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines
Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted. This article accepts that diagnosis but challenges its explanatory framework, which compares an embodied, socially situated human knower with an isolated generative model thereby locating epistemic legitimacy in capacities internal to autonomous agents. Drawing on Carlo Sini's philosophy of practices, writing, signs, and technics, we propose instead to understand a large language model (LLM) as a *techno-semiotic machine* that automates a phase of written semiosis by producing plausible linguistic configurations from the sedimented archive of human writing. From this perspective, *Epistemia* is one consequence of a broader phenomenon that we call *epistemic schizologia*: the socio-technical cleavage between signs as linguistically accomplished expressions and signs as moments within socially embedded circuits of interpretation, evidence, criticism, verification, and responsibility. This cleavage is reinforced by *eikotic closure*, through which a plausible continuation is presented with the finality of an epistemic result, and by algorithmic authority and epistemic self-misrecognition. The relevant unit is therefore not the model alone but the complete practice in which generated inscriptions are prompted, interpreted, verified, contested, used, and made consequential. This reframing preserves the distinction between linguistic production and responsible understanding while grounding a design programme centred on inspectable genealogy, contestability, distributed responsibility, epistemic agency, and the evaluation of hybrid human--AIpractices.
正の二次ネットワークにおける商ダイナミクス、実効曲率、および暗黙的なバイアス
正の 2 次ネットワークでは、低ランク表現 f_U(x)=x^top UU^top x が認められます。ここで、Uinmathbb{R}^{dtimes r} は、右直交乗算までしか識別できず、ランク r の PSD 行列 Q=UU^top を表します。私たちは、この商の構造がトレーニングのダイナミクス、曲率、回復、補間バイアスをどのように制御するかを研究します。フル列ランク層では、ランク r PSD 多様体を使用して mathbb{R}^{dtimes r}_*/O(r) を識別します。スムーズな目的 L(U)=ell(UU^top) の場合、ユークリッド係数の勾配は水平になります。したがって、因子勾配流は商リーマン勾配流に正確に投影されますが、有限ステップ勾配降下法は予測子の正確な合同再帰を引き起こします。二次回帰の場合、商計量に対する接線空間に制限された経験的測定グラム形式として、補間器での有効ヘシアンを導出します。ガウス ランク 1 測定の下で、母集団曲率を計算し、経験的正規演算子の一様偏差限界を証明し、スペクトル初期化子を構築し、勾配流に対する局所指数収束とスモールステップ降下に対する線形収束を確立します。回復保証は明示的ですが、全空間の秒瞬間制御に依存しているため、保守的です。不十分に決定された通勤体制では、因子勾配の流れは、結合スペクトル座標における正確なエントロピーミラーの流れになります。厳密に正の初期化は、内挿セットへのブレグマン投影に収束します。等方性初期化 q(0)=varepsilon^2mathbf{1} を使用すると、予測子は varepsilondownarrow0 として設定された最小トレース解に近づき、不変結合スペクトル代数内の重み付きエントロピーによって非一意性を解決します。有限ステップ降下では、O(η) による連続時間ブレグマン投影とは異なる内挿が選択されます。数値実験により、これらの商の正体、曲率予測、回復挙動、および選択則が検証されます。
原文 (English)
Quotient Dynamics, Effective Curvature, and Implicit Bias in Positive Quadratic Networks
Positive quadratic networks admit the low-rank representation f_U(x)=x^top UU^top x, where Uinmathbb{R}^{dtimes r} is identifiable only up to right orthogonal multiplication, representing a rank-r PSD matrix Q=UU^top. We study how this quotient structure governs training dynamics, curvature, recovery, and interpolation bias. On the full-column-rank stratum, we identify mathbb{R}^{dtimes r}_*/O(r) with the rank-r PSD manifold. For smooth objectives L(U)=ell(UU^top), the Euclidean factor gradient is horizontal. Thus, factor gradient flow projects exactly to quotient Riemannian gradient flow, while finite-step gradient descent induces an exact congruence recursion for the predictor. For quadratic regression, we derive the effective Hessian at interpolators as the empirical measurement Gram form restricted to the tangent space relative to the quotient metric. Under Gaussian rank-one measurements, we compute population curvature, prove uniform deviation bounds for the empirical normal operator, construct a spectral initializer, and establish local exponential convergence for gradient flow and linear convergence for small-step descent. Recovery guarantees are explicit but conservative due to reliance on full-space second-moment control. In underdetermined commuting regimes, factor gradient flow becomes an exact entropy mirror flow in joint spectral coordinates. Strictly positive initializations converge to Bregman projections onto the interpolation set. With isotropic initialization q(0)=varepsilon^2mathbf{1}, predictors approach the minimum-trace solution set as varepsilondownarrow0, resolving nonuniqueness via weighted entropy within the invariant joint spectral algebra. Finite-step descent selects interpolants differing from continuous-time Bregman projections by O(eta). Numerical experiments verify these quotient identities, curvature predictions, recovery behaviors, and selection laws.
中国語の音声生成と知覚におけるEEGからテキストへのデコードのためのテキストと音声の共同調整
頭皮脳波 (EEG) から音声情報をテキストに直接デコードすることで、重度の音声障害や運動障害を持つ個人に潜在的な非侵襲性の神経伝達経路が提供されます。皮質電図検査などの侵襲的アプローチと比較して、EEG はより安全で広範囲に導入可能ですが、解読はかなり困難です。この課題は、数千の文字、被験者間の深刻な変動性、およびテキストの位置合わせのための低い信号対雑音比を含む高次元の出力空間を処理する必要がある中国語文の解読ではさらに悪化します。既存の方法は、単一の監視軸、つまりテキスト セマンティクスまたはオーディオ音響特徴のいずれかに取り組んでいますが、どちらも同時に満足させることはできません。大量の語彙を含む中国語の解読には、文レベルの識別能力ときめ細かい時間分解能が求められます。我々は、EEGAlign を導入します。これは、EEGA を 2 つの軸、すなわち BGE-M3 テキスト埋め込みによるテキストの位置合わせと、対照学習とそれに続く CTC 文字列デコードによる wav2vec~2.0 音声特徴による音声の位置合わせを組み合わせて位置合わせする、新しいパラメータ効率の高いフレームワークです。 ChineseEEG-2 データでは、EEGAlign は最先端のクローズドセット文分類パフォーマンスを実現し、101 の候補のうち音読 EEG でトップ 1 の精度が 82.37%、受動聴取 EEG で 41.43% に達します。アブレーション研究では、2 つのアライメント軸が高度に補完的であることが示されており、これらを組み合わせることで、どちらか一方を単独で使用するよりも一貫して優れたパフォーマンスが得られます。私たちの知る限り、これは、明白な音声生成中に非侵襲的脳波から語彙の多い中国語文を解読し、比較的大規模な閉集合候補文設定で強力な分類性能を達成することに関する最初の研究です。
原文 (English)
Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception
Decoding speech information directly from scalp electroencephalography (EEG) into text provides a potential non-invasive neural communication pathway for individuals with severe speech and motor impairments. Compared with invasive approaches such as electrocorticography, EEG is safer and more widely deployable, yet substantially more challenging to decode.This challenge is exacerbated for Chinese sentence decoding, which must handle a high-dimensional output space with thousands of characters, severe inter-subject variability, and low signal-to-noise ratios for text alignment.Existing methods commit to a single supervisory axis---either text semantics or audio acoustic features---yet neither can simultaneously satisfy the demands of sentence-level discriminability and fine-grained temporal resolution required for large-vocabulary Chinese decoding. We introduce EEGAlign, a novel parameter-efficient framework that jointly aligns EEG with two axes---text alignment with BGE-M3 text embeddings and audio alignment with wav2vec~2.0 speech features via contrastive learning followed by CTC character-sequence decoding. On ChineseEEG-2 data, EEGAlign yields state-of-the-art closed-set sentence classification performance, reaching up to 82.37% Top-1 accuracy on Reading Aloud EEG and 41.43% on Passive Listening EEG out of 101 candidates. Ablation studies show that the two alignment axes are highly complementary: combining them yields consistently better performance than either alone. To the best of our knowledge, this is the first study on decoding large-vocabulary Chinese sentences from non-invasive EEG during overt speech production, and achieving strong classification performance with relatively large closed-set candidate-sentence setting.
AIriskEval-edu デモ: 教育的説明における教育的リスクの監査
私たちは、指導説明の教育的品質を監査し、説明可能な監査結果を提供するプラットフォームである AIriskEval-edu デモを紹介します。このプラットフォームは、教育学的リスクの 5 つの側面 (事実の正確さ、深さと完全性、焦点と関連性、生徒レベルの適切性、イデオロギーの偏見) をカバーするルーブリックに照らして説明を評価します。ディメンションごとに、二分決定と信頼スコアを返します。検出されたリスクには、自然言語の根拠と、深さと完全性を除いて、局所的な証拠範囲も含まれます。このプラットフォームは、外部 API とコンシューマー グレードの GPU で実行されるセルフホスト型 Llama 3.1 8B エバリュエーターを通じて GPT-5.5 を統合します。ローカル評価者は、リスクと説明可能性の注釈を備えた幼稚園から高校までの教育説明のデータセットである AIriskEval-edu で微調整されています。このプラットフォームは 2 つのモードで動作します。AI モードでは、両方の評価者が、それぞれ異なる教育的行動と潜在的なリスクを表す 6 つのシミュレートされた教師プロファイルに基づいて生成された、保存された説明を評価します。人間モードでは、ローカル評価者がユーザーが作成した説明をリアルタイムで監査します。ローカル評価者は、報告されているほとんどの指標で GPT-5.5 を上回っており、教育機関に監査済みのコンテンツを独自のインフラストラクチャ内に保持する実用的な方法を提供します。
原文 (English)
AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations
We present AIriskEval-edu Demo, a platform that audits the pedagogical quality of instructional explanations and provides explainable audit results. The platform evaluates an explanation against a rubric covering five dimensions of pedagogical risk: factual accuracy, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias. For each dimension, it returns a binary decision and a confidence score. Detected risks also include a natural-language rationale and, except for Depth and Completeness, a localized evidence span. The platform integrates GPT-5.5 through an external API and a self-hosted Llama 3.1 8B evaluator that runs on consumer-grade GPUs. The local evaluator is fine-tuned on AIriskEval-edu, a dataset of K-12 instructional explanations with risk and explainability annotations. The platform operates in two modes: in AI mode, both evaluators assess stored explanations generated under six simulated teacher profiles, each representing a distinct pedagogical behavior and potential risk; in human mode, the local evaluator audits user-written explanations in real time. The local evaluator outperforms GPT-5.5 on most reported metrics, offering educational institutions a practical way to keep audited content within their own infrastructure.
エンジンは平等、人間は不平等: エンジンが評価した平等なチェスの局面における再現可能な結果の偏り
強力なエンジンが本質的に等しいと判断するチェスの開始位置 (Stockfish 18 のゼロから 10 センチポーン以内の評価、深さ安定) と人間が Lichess で実際に到達する位置 (2025 年 10 月、1,661 の位置、1,610 万回) の間では、人間の結果はバランスが取れていません。ポジションには結果の偏りがあり、それぞれのゲームの実際の結果とプレーヤーのレーティングの予測との間のギャップがあり、その方向性は自然に到達するポジションの安定した特性です。いくつかのポジションは白を支持し、他のポジションは黒を支持します。これらの偏りは、3 つの再パーティション、つまり、不連続なプレイヤー アカウント セット (プライマリ)、時間、および不連続なレーティング バンドにわたって再現され、さらに 8 か月後のサンプル外の月にも再現されます。プライマリ分割では、各ポジションのスキューが各口座グループで 1 回測定され、反復勾配は、格付けとオープニングファミリー効果を除去した後、一方の測定値が他方の測定値をどの程度正確に予測するかを尋ねます。1 つは減少しないキャリーオーバーを意味し、もう 1 つは減少しないキャリーオーバーを意味します。ゼロ、線形関係はありません。結果は 0.69 (ファミリークラスター化 95% CI [0.65, 0.74]) であり、最も人気があり、最もよく測定されたポジションでは 0.94 に上昇しました。傾きの値はポジションの組み合わせによって異なります。存在は不変の主張です。それは、テストするあらゆる厳しい評価範囲、検索の深さ、校正、および人気のカットオフに耐え、電撃内と急速内で別々に複製されます。典型的な偏りは小さい (中央値 $|\delta| \約 0.018$、白スコアの約 2 パーセント ポイント) にもかかわらず、それは、関連性のないアカウント間で位置ごとに再現されます。このような位置では、不利な側も長く考えます。評価に最も信頼性がある場合でも、それは人間の成果を表す十分な統計ではありません。結果は観察的なものであり、因果関係の疑問は、事前に登録されたランダム化されたコンパニオン研究に委ねられます。
原文 (English)
Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions
Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16.1M occurrences), human results are not balanced. Positions carry outcome skews, each the gap between its games' actual results and what the players' ratings predict, whose directions are stable properties of the naturally-reached position: some positions favour White, others Black. These skews reproduce across three re-partitions -- disjoint player-account sets (primary), time, and disjoint rating bands -- and on an out-of-sample month eight months later. On the primary split, each position's skew is measured once in each account group, and the replication slope asks how well one measurement predicts the other after removing rating and opening-family effects: one means undiminished carry-over; zero, no linear relation. We find 0.69 (family-clustered 95% CI [0.65, 0.74]), rising to 0.94 on the most-popular, best-measured positions. The slope's value depends on the position mix. Existence is the invariant claim: it survives every tighter evaluation band, search depth, calibration, and popularity cutoff we test, and replicates within blitz and rapid separately. The typical skew is small (median $|\delta| \approx 0.018$, about two percentage points of White score), yet it reproduces, position by position, across disjoint accounts. At these positions the disfavoured side also thinks longer. Even where the evaluation is most confident, it is not a sufficient statistic for human outcomes. The result is observational, and the causal question is left to a pre-registered randomised companion study.
OrchBench: 決定論的シミュレーションによる個別のマルチエージェント オーケストレーション プランの評価
複雑なタスクは多くの場合、並列化可能でありながら相互依存するサブタスクに分解されるため、マルチエージェント システム (MAS) のパフォーマンスにとってオーケストレーションが重要になります。既存の評価は通常、エンドツーエンドの実行に依存しており、オーケストレーション計画の品質と作業者の能力、ツールの信頼性、および環境ノイズを混同します。さらに、実際の実行にかかる時間とトークンのコストはワークフローの規模に応じて急速に増大するため、体系的な評価が高価になります。マルチエージェント オーケストレーション プランを個別に評価するためのシミュレーション ベースのベンチマークである OrchBench を紹介します。 OrchBench は、実際のタスクから開始して、サイズと並列度を制御してタスクの依存関係をエンコードする有向非巡回グラフ (DAG) を構築します。 DAG、エージェントごとのコンテキスト制限、およびエージェントの予算を考慮して、評価されたプランナーはサブタスクをエージェントに割り当て、エージェント間の情報転送とその保持率を指定します。決定論的シミュレーターは、ワーカー エージェントを呼び出すことなく結果の計画を評価し、結果の品質、メイクスパン、およびトークン コストの解釈可能な尺度を返します。 OrchBench によって生成されたシミュレートされたスコアは、クロード コードの実行からの品質スコアと強い相関があり、\(r=0.816\) のピアソン相関を達成しながら、トークンの \(1.3\%\) と実時間の \(10.3\%\) のみが必要です。さまざまなプランナーやワークフロー規模にわたって、単にエージェントの数を増やすよりも、タスクに不可欠な情報を保持することの方が重要であり、調整の失敗が蓄積するにつれて並列処理の利点が減少することがわかりました。これらの結果により、OrchBench は、マルチエージェント オーケストレーション プランを比較および診断するための効率的で解釈可能なベンチマークとして確立されます。
原文 (English)
OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS). Existing evaluations typically rely on end-to-end execution, which conflates orchestration-plan quality with worker capabilities, tool reliability, and environmental noise. Moreover, the time and token costs of real execution grow rapidly with workflow scale, making systematic evaluation expensive. We present OrchBench, a simulation-based benchmark for evaluating multi-agent orchestration plans in isolation. Starting from real-world tasks, OrchBench constructs directed acyclic graphs (DAGs) that encode task dependencies, with controlled sizes and degrees of parallelism. Given a DAG, a per-agent context limit, and an agent budget, the evaluated planner assigns subtasks to agents and specifies cross-agent information transfers and their retention ratios. A deterministic simulator evaluates the resulting plan without invoking worker agents and returns interpretable measures of result quality, makespan, and token cost. The simulated scores produced by OrchBench correlate strongly with quality scores from Claude Code executions, achieving a Pearson correlation of \(r=0.816\), while requiring only \(1.3\%\) of the tokens and \(10.3\%\) of the wall-clock time. Across diverse planners and workflow scales, we find that preserving task-critical information is more important than simply increasing the number of agents, and the benefits of parallelism diminish as coordination failures accumulate. These results establish OrchBench as an efficient and interpretable benchmark for comparing and diagnosing multi-agent orchestration plans.
CoRT: トークンレベルのルーブリックに基づくポリシー最適化のための反事実リプレイ
ルーブリックベースの強化学習は、明示的な基準に照らしてモデルの出力を評価することにより、言語モデルのトレーニングを強化します。しかし、GRPO スタイルのパイプラインでは、これらの構造化された判断はスカラー応答レベルの報酬に還元され、応答レベルの利点に変換され、生成されたすべてのトークンに均一にブロードキャストされます。これにより、異なる基準が異なるスパン、フォーマット決定、またはセマンティック選択に基づいている場合でも、応答内でクレジットを割り当てるための明示的なメカニズムが残されません。ルーブリック条件付き GRPO のトークンレベルのクレジット重み付け手法である CoRT を提案します。補助トークン スコアリング モデルをトレーニングする代わりに、CoRT は反事実リプレイを使用して、元のルーブリック条件付きプロンプトと一致した基準なしのプロンプトの下で同じサンプリングされた応答を再スコアリングします。結果として得られるトークンごとの対数尤度対比は、ルーブリック コンテキストへの依存性の代用として機能します。 CoRT は、これらのコントラストを、制限された応答正規化された重みにマッピングし、それらを使用して、補助スコアラーを導入したり、応答レベルの報酬を変更したりすることなく、署名された GRPO の利点をトークン全体に再分配します。命令調整モデルと報酬粒度にわたる実験では、大部分の比較において、CoRT が一致する応答レベルの GRPO よりも改善し、平均 4.4 パーセント ポイント向上していることが示されています。この方法は、個別の関連性学習段階を回避しながら、学習されたトークンレベルの信用ベースラインとの競争力を維持します。これらの結果は、政策内部の反事実尤度の対比が、GRPO の単純さと安定性を維持しながら、応答内クレジット割り当てのための効果的なトレーニング シグナルを提供することを示唆しています。
原文 (English)
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.
局所的な適応により、トランスフォーマーの独特の学習特徴が明らかに
トランスフォーマーの適応は、意図した変更が狭い場合でも、通常、モデルの深度全体に分散されます。私たちは、適応部位がモデルの学習内容をどのように形成するか、その学習がどの程度一般化されるか、どのように選択的に適用されるかを調査します。我々は、5つの目的(語彙結合、事実関連、行動ポリシー学習、因果関係マッピング、および手続き的推論)にわたる制御されたベンチマークを導入し、各目的の「適応幾何学」を、フルスタックおよび初期、中期、または後期層のLoRAの下での取得、転送、および有界性のプロファイルとして定義します。対物レンズは独特の形状を示します。字句バインディングは、取得と有界性については初期層の適応に有利ですが、転送にはより広範な更新が必要です。事実の関連性は、ローカライズされたアダプター間の後続の層に有利です。行動学習は、後期層のアクション取得を中間層のポリシー ゲートから分離します。そして、因果的および手続き的移転は、ミドルスタックまたはフルスタックの適応から最も恩恵を受けます。これらのパターンは、パラメーターが一致した制御下で主に持続し、対応する方向コントラストのほとんどが 5 つのモデル ファミリにわたって複製されます。これらの発見により、どのモデルが学習し、一般化し、変更しないままにするかを制御するための主要な設計変数として適応部位が確立されます。
原文 (English)
Localized Adaptation Reveals Distinct Learning Signatures in Transformers
Transformer adaptation is typically distributed across model depth, even when the intended change is narrow. We investigate how adaptation site shapes what a model learns, how well that learning generalizes, and how selectively it is applied. We introduce a controlled benchmark spanning five objectives (lexical binding, factual association, behavioral policy learning, causal mapping, and procedural reasoning) and define each objective's "adaptation geometry" as its profile of acquisition, transfer, and boundedness under full-stack and early-, middle-, or late-layer LoRA. The objectives exhibit distinct geometries. Lexical binding favors early-layer adaptation for acquisition and boundedness but requires broader updates for transfer; factual association favors later layers among localized adapters; behavioral learning separates late-layer action acquisition from middle-layer policy gating; and causal and procedural transfer benefit most from middle- or full-stack adaptation. These patterns largely persist under parameter-matched controls, and most corresponding directional contrasts replicate across five model families. These findings establish adaptation site as a key design variable for controlling what models learn, generalize, and leave unchanged.
OmniDelta: OmniLLM でのトークン圧縮のためのスキル主導の予算割り当て
新しいオムニモーダル大規模言語モデル (OmniLLM) により、テキスト、オーディオ、およびビデオを統一的に理解できるようになりますが、その長いオーディオ/ビデオ トークン シーケンスにより、メモリと推論のコストが大幅に増加します。既存の圧縮方法は主に、固定予算の下で重要なトークンを選択することに焦点を当てており、前述の予算割り当ての問題は十分に調査されていません。我々は、直接的なクエリとオーディオ/ビデオの類似性は、モーダル間の予算配分では信頼できないこと、および一律のモーダル内予算では、冗長なコンテンツを保持したまま重要な証拠を見逃してしまう可能性があることを示します。これらの制限に対処するために、私たちは、意図を意識したモーダル間割り当てとコンテンツを意識したモーダル内割り当てを組み合わせる、トレーニング不要のスキル主導型フレームワークである OmniDelta を提案します。 OmniDelta は、最初にオーディオとビデオのスキル プールを構築して、クエリの需要に応じて固定保持トークン バジェットをシフトし、次にローカルの複雑さと時間的冗長性を使用して、モダリティ バジェットをオーディオ セグメントとビデオ フレームに再割り当てします。結果として得られるローカル予算は既存のプルーニング戦略と組み合わせることができ、予算の使用先を変更しながら合計保持トークン率を維持できます。 2 つの Qwen2.5-Omni モデルを使用した 4 つのオーディオ/ビデオ ベンチマークの実験では、OmniDelta が枝刈り比全体にわたって新しい精度効率のパレート フロンティアを確立していることが示されています。 Qwen2.5-Omni-7B では 25% のトークン保持率で、OmniDelta は GPU メモリを 22.0% 削減し、フルトークン推論と比較して 1.64 倍のエンドツーエンドの高速化を達成します。
原文 (English)
OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs
Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under fixed budgets, leaving the preceding budget-allocation problem underexplored. We show that direct query-to-audio/video similarity is unreliable for inter-modal budget allocation, and that uniform intra-modal budgets can miss key evidence while retaining redundant content. To address these limitations, we propose OmniDelta, a training-free, skill-driven framework that couples intent-aware inter-modal allocation with content-aware intra-modal allocation. OmniDelta first constructs audio and video skill pools to shift the fixed retained-token budget according to query demand, then reallocates modality budgets over audio segments and video frames using local complexity and temporal redundancy. The resulting local budgets can be combined with existing pruning strategies, preserving the total retained-token ratio while changing where the budget is spent. Experiments on four audio-video benchmarks with two Qwen2.5-Omni models show that OmniDelta establishes a new accuracy-efficiency Pareto frontier across pruning ratios. At 25% token retention on Qwen2.5-Omni-7B, OmniDelta reduces GPU memory by 22.0% and achieves a 1.64x end-to-end speedup over full-token inference.
DecoEvo: テキスト空間におけるソルバー スキルとルーブリック ジェネレーター スキルのスコア分離型共進化
テキスト空間の最適化では、モデルの重みではなく外部の自然言語アーティファクトを編集することで大規模言語モデル (LLM) を適応させるため、最適化されたアーティファクトは検査可能なままであり、モデルをブラック ボックスとして扱うことができます。ただし、既存のテキストスペースメソッドのほとんどは評価を固定したままにします。オープンエンド タスクでは、これがボトルネックになる可能性があります。ソルバーがルーブリックの測定基準を改善すると、省略されたディメンションは最適化信号には見えないままになります。現在のソルバーのスコアによって更新が選択されている場合、ルーブリックを単純に進化させることも信頼性が低くなります。ルーブリックを満たしやすくすることで明らかな進歩が得られる可能性があるためです。最適化中にゴールド ルーブリックを使用せずに、分離された目標の下でソルバー スキルとルーブリック ジェネレーター スキルを共進化させる DecoEvo (Decoupled Co-Evolution) を紹介します。ソルバー スキルは基準レベルのフィードバックを使用して更新され、ルーブリック ジェネレーター スキルは、集計されたソルバー スコアとは独立した要件の適用範囲と応答の識別の補完的な監査を通じて改訂されます。この分離により、ジェネレーターの更新は新たに明らかになったソルバーの弱点に焦点が当てられ、ソルバーがすでに満たしている基準が繰り返し強調されることが減ります。各ベンチマークの公式評価では、DecoEvo は 5 つのベンチマークと 3 つの LLM バックボーンにわたって比較されたすべての手法を上回り、5 つのベンチマークの平均で SkillOpt に対して 2.8 ~ 5.0\% の相対的な向上をもたらしました。
原文 (English)
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also unreliable when updates are selected by the current solver's score, because apparent progress can come from making the rubric easier to satisfy. We introduce DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization. The solver skill is updated using criterion-level feedback, while the rubric-generator skill is revised through complementary audits of requirement coverage and response discrimination that are independent of aggregate solver score. This separation focuses generator updates on newly exposed solver weaknesses, reducing repeated emphasis on criteria the solver already satisfies. Under each benchmark's official evaluation, DecoEvo outperforms all compared methods across five benchmarks and three LLM backbones, yielding 2.8--5.0\% relative gains over SkillOpt in the five-benchmark average.
Cognivia: 証拠に基づいたメンタルヘルスケアのための認知行動療法の副操縦士
認知の歪みは否定的な感情を増幅させ、精神的健康障害の一因となります。認知行動療法(CBT)は認知の歪みに対処する効果的な方法ですが、専門のセラピストが不足しているため、その大規模な適用は制限されています。最近、大規模言語モデル (LLM) がメンタルヘルスへの応用に向けて研究されていますが、既存の手法では依然として、限定された領域特異性、過度にお世辞な応答、および認知の歪みに対する明確に定義された注釈の欠如という問題に悩まされています。この論文では、自動的な認知の歪みの特定と合理的な応答生成を統合する、証拠に基づいた人工知能セラピストである Cognivia を提案します。私たちのフレームワークは、コアパラダイムおよび標準参考文献として広くみなされている権威ある CBT テキストに基づいて構築されています。これは、メンタルヘルスの質問と回答 (Q and A) データでさらに強化されており、行動科学の専門家の監督の下、多段階のプロンプトと構造化された生成戦略を採用しています。次に、この拡張された CBT データセットで軽量 LLM を微調整して、Cognivia を取得します。さらに、AI 研究者と行動科学の専門家の協力を通じて開発された、LLM によって生成された合理的な応答を評価するための初の階層的品質評価フレームワークを提案します。 Cognivia は、語彙メトリクス、2 つの補完的な基準を備えた LLM ベースの審査員、および 10 人の行動科学の専門家による人間による評価を使用して評価されます。これは、認知の歪みの認識と合理的な応答の生成においてベースライン手法を常に上回っており、その有効性を示しています。私たちのコードは https://github.com/SNOWTEAM2023/Cognivia で入手できます。
原文 (English)
Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare
Cognitive distortion amplifies negative emotions and contributes to mental health disorders. Cognitive Behavioral Therapy (CBT) is an effective way to address cognitive distortions, but its large-scale application is limited by the shortage of professional therapists. Although large language models (LLMs) have recently been explored for mental health applications, existing methods still suffer from limited domain specificity, overly flattering responses, and the absence of well-defined annotations for cognitive distortions. This paper proposes Cognivia, an evidence-based artificial intelligence therapist that integrates automatic cognitive distortion identification and rational response generation. Our framework is built on authoritative CBT texts widely regarded as core paradigms and standard references. It is further augmented with mental health question-answer (Q and A) data, and employs multi-stage prompting and structured generation strategies under the supervision of behavioral science experts. Then we fine-tune a lightweight LLM on this augmented CBT dataset to obtain Cognivia. In addition, we propose the first hierarchical quality evaluation framework for assessing LLM-generated rational responses, developed through collaboration between AI researchers and behavioral science experts. Cognivia is evaluated using lexical metrics, LLM-based Judges with two complementary criteria, and human evaluation by 10 behavioral science experts. It consistently outperforms the baseline methods in cognitive distortion recognition and rational response generation, demonstrating its effectiveness. Our code is available at https://github.com/SNOWTEAM2023/Cognivia.
LLM が生成する推奨事項の説明を通じて持続可能な選択を促す
レコメンダー システムは日常の消費を仲介し、持続可能な選択を促すための有望なチャネルを提供します。以前の研究では、説明が推奨事項に対するユーザーの認識に影響を与え、より多くの情報に基づいた意思決定をサポートできることが示されています。私たちは、選択の瞬間に持続可能性の情報を前面に出すことで、説明が行動を促す役割を果たすこともできると主張します。この研究では、推奨事項の説明におけるサステナビリティ情報のさまざまな行動枠組みが、ユーザーの選択や認識にどのような影響を与えるかを調査しています。生成 AI を使用して、ナッジ理論を利用してサステナビリティを意識した説明を生成し、人間による評価と LLM による審査員監査を通じて検証します。この基盤に基づいて、関与度の低い領域 (インスタント コーヒー) と関与度の高い領域 (ホテルの予約) で 2 つのランダム化研究 ($N = 529 ドル) を実施します。この研究では、参加者は、これらの説明を伴う好みに合った推奨事項の中から選択します。私たちの結果は、どちらの領域においても、説明の中で持続可能性に関する情報を開示するだけでは選択肢は変わらないのに対し、その情報を枠組み化したり説明的な社会規範を援用したりすることで持続可能な選択が大幅に増加し、意思決定が容易になることを示しています。特に、単純な開示はより持続可能な選択行動に変換することなく説明評価を向上させるため、認識と行動は乖離します。私たちの研究は、LLM がどのように理論に基づいた説明を大規模に生成できるかを実証し、社会的利益のための実践的な説明に基づく介入を示唆しています。最後に、生成 AI による適応説明デザインへの影響について説明します。
原文 (English)
Nudging Sustainable Choices through LLM-Generated Recommendation Explanations
Recommender systems mediate everyday consumption, offering a promising channel for encouraging sustainable choices. Prior research shows that explanations influence users' perceptions of recommendations and can support more informed decisions. We argue that explanations can also serve as behavioral nudges by foregrounding sustainability information at the moment of choice. This study investigates how different behavioral framings of sustainability information in recommendation explanations affect user choices and perceptions. Using generative AI, we generate sustainability-aware explanations by drawing on nudge theory and validate them through human evaluation and LLM-as-a-judge audits. Building on this foundation, we conduct two randomized studies ($N = 529$) in a low involvement domain (instant coffee) and a high involvement domain (hotel bookings), in which participants choose among preference matched recommendations accompanied by these explanations. Our results show that, across both domains, merely disclosing sustainability information in explanations does not change choices, whereas framing that information or invoking a descriptive social norm significantly increases sustainable selections and eases decision-making. Notably, perception and behavior diverge, as plain disclosure improves explanation evaluations without translating into more sustainable selection behavior. Our work demonstrates how LLMs can generate theory-grounded explanations at scale, pointing toward practical explanation-based interventions for social good. We conclude by discussing implications for adaptive explanation design with generative AI.
損失の不変性は、レイヤーがエンコードする概念を決定します: 心エコー検査におけるボリュームグラウンディング
目的: コンセプトのボトルネックは、解釈可能な中間変数を介してルート予測をモデル化し、その妥当性は通常、それらの変数がどれだけ正確に予測されるかによって判断されます。心エコー検査ビデオからの駆出率推定の基礎となる概念として左心室容積を使用して、その判断が十分であるかどうかを尋ねます。方法: ビデオ トランスフォーマー エンコーダーは、公開されている心エコー検査データセットでトレーニングされました。収縮終期容積と拡張終期容積は概念層を形成し、そこから駆出率が分析的に計算され、出力への残留経路はありません。私たちは、駆出率の目標のみに基づいたトレーニングと、ミリリットル単位の量の追加の監視を伴うトレーニングを比較し、1,276件の実施された研究で両方を評価しました。結果: コンセプトのボトルネックは、平均絶対誤差 7.13 に対して 6.89 で、直接回帰と比較して駆出率誤差を増加させませんでした。しかし、量の監視がなければ、相関関係は部分的に保たれたものの、予測量の広がりは参照広がりの 35.7 ミリリットルと 45.7 ミリリットルに対して 0.1 ミリリットルにまで崩壊しました。これは目的の不変特性から導かれることを示します。駆出率は比率であり、両方の体積が再スケーリングされても変化しないため、損失はスケールまでしか概念層を決定しません。絶対単位での監視により、駆出率誤差が 0.4 減少しましたが、体積誤差は 89.8 ミリリットルから 25.8 ミリリットルに減少しました。結論: コンセプトの正確さだけで、物理的なスケールを持たないコンセプト層を隠すことができます。重要性: 臨床モデルの解釈可能な中間変数は、予測精度だけでなく、トレーニング目標の不変構造に対しても検証される必要があります。
原文 (English)
Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography
Objective: Concept bottleneck models route prediction through interpretable intermediate variables, and their validity is normally judged by how accurately those variables are predicted. We ask whether that judgement is sufficient, using left ventricular volumes as the concepts underlying ejection fraction estimation from echocardiographic video. Methods: A video transformer encoder was trained on a publicly available echocardiography dataset. End-systolic and end-diastolic volumes formed a concept layer from which ejection fraction was computed analytically, with no residual path to the output. We compared training under an ejection fraction objective alone against training with additional supervision of the volumes in millilitres, and evaluated both on 1276 held-out studies. Results: The concept bottleneck did not increase ejection fraction error relative to direct regression, at 6.89 against 7.13 mean absolute error. Without volume supervision, however, the spread of predicted volumes collapsed to 0.1 millilitres against reference spreads of 35.7 and 45.7 millilitres, while correlation was partly preserved. We show that this follows from an invariance property of the objective: ejection fraction is a ratio and is unchanged when both volumes are rescaled, so the loss determines the concept layer only up to scale. Supervision in absolute units reduced volume error from 89.8 to 25.8 millilitres at a cost of 0.4 in ejection fraction error. Conclusion: Concept accuracy alone can conceal a concept layer that carries no physical scale. Significance: Interpretable intermediate variables in clinical models should be validated against the invariance structure of the training objective, not only against prediction accuracy.
推論しながら推測する: エージェントと投機者の共同 RL を介してエージェントに次のツール呼び出しを予測するよう教える
大規模な言語モデルのエージェントは、ツール呼び出しの結果を待つのにかなりの時間を費やすことがよくあります。ツール呼び出しの投機では、エージェントの次のツール呼び出しを予測して、その予測がエージェントの最終的なツール呼び出しと一致する場合に事前実行することで、このレイテンシーを隠すことができますが、既存のスペキュレーターは通常、展開されたエージェント自体の動作とあまり整合していない個別のドラフト モデルまたはキャッシュされたトレースです。我々は、この投機者とエージェントのギャップを特定し、ターゲット エージェント自体が強力なネクストコール投機者であることを示します。これは、同じモデル内でエージェントと投機者を統合するという、よりシンプルな設計を示しています。この論文では、プレフィックス KV キャッシュを完全に再利用して、エージェント モードでタスクを解決し、スペキュレーター モードで部分的な軌跡から次のツール呼び出しを予測する単一モデルである自己投機エージェントを紹介します。パフォーマンスを低下させることなくこのデュアルモード エージェントを有効にするために、エージェント自身のロールアウトから推測ターゲットを導き出し、エージェントと推測者の更新を交互に行う、エージェントと推測者の共同強化学習方法を提案します。エージェント タスクの成功を維持しながら、エージェント検索 QA と会話型ツールを使用するエージェント タスク全体で、私たちの方法は平均次回ツール呼び出し Hit@1 を Qwen3-4B で 44.1 から 61.2 に、Qwen3.5-4B で 48.9 から 66.3 に改善しました。
原文 (English)
Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this latency by predicting and pre-executing an agent's next tool call if the prediction matches the agent's eventual tool call, but existing speculators are typically separate draft models or cached traces that are poorly aligned with the deployed agent's own behavior. We identify this speculator-agent gap and show that the target agent itself is a strong next-call speculator. This points to a simpler design: unifying the agent and speculator within the same model. In this paper, we introduce the self-speculating agent, a single model that both solves tasks in agent mode and predicts its next tool call from partial trajectories in speculator mode, fully reusing prefix KV cache. To enable this dual-mode agent without degrading performance, we propose a joint agent-speculator reinforcement learning method, which derives speculation targets from the agent's own rollouts and alternates agent and speculator updates. Across agentic search QA and conversational tool-use agentic tasks, our method improves average next tool-call Hit@1 from 44.1 to 61.2 for Qwen3-4B and from 48.9 to 66.3 for Qwen3.5-4B, while preserving agent task success.
大規模衛星スケジューリングへの応用によるオンライン学習と反復価格設定による分散制約の最適化
分散制約最適化問題 (DCOP) は、限られた通信環境下で分散意思決定を行うための一般的なフレームワークを提供しますが、現実世界のインスタンスの多くは大きすぎてモノリシックに解決できません。私たちはこの課題に 2 つの相補的な方向から取り組みます。私たちは DCOP と潜在的なゲームの間の関係を再考し、均衡を見つけるための最新のオンライン学習アルゴリズムを DCOP に適応させます。これらのアルゴリズムが代表的な不完全な DCOP アルゴリズムと競合することを示します。次に、大規模な分散衛星スケジューリングを動機とする大規模 DCOP の分解フレームワークに移ります。我々は、DCOP を 2 つの相互作用するサブ問題、つまりタスク割り当てのための高レベルのメタ DCOP と、スケジューリングのための独立したローカル最適化問題に分離する新しいフレームワークを提案します。 2 つのレベルを結合するために、ローカル オプティマイザーからのフィードバックを使用してメタレベルのユーティリティを更新する新しい反復的な価格設定方法を開発します。当社のオンライン学習方法と反復的な価格設定フレームワークを組み合わせることで、現実世界の分散型衛星のスケジューリング問題のインスタンスでほぼ最適なパフォーマンスが得られ、最先端のベースラインの場合は 87% であるのに対し、観測リクエストの 99% 以上が満たされます。
原文 (English)
Distributed Constraint Optimization via Online Learning and Iterative Pricing with Application to Large-Scale Satellite Scheduling
Distributed constraint optimization problems (DCOPs) provide a popular framework for distributed decision making under limited communication, but many real-world instances are too large to solve monolithically. We address this challenge from two complementary directions. We revisit the connection between DCOPs and potential games, and adapt modern online learning algorithms for equilibrium finding to DCOPs. We show that these algorithms are competitive with representative incomplete DCOP algorithms. We then turn to decomposition frameworks for large-scale DCOPs, motivated by large-scale decentralized satellite scheduling. We propose a new framework that separates a DCOP into two interacting subproblems: a high-level meta-DCOP for task allocation, and independent local optimization problems for scheduling. To couple the two levels, we develop a novel iterative pricing method that updates the meta-level utilities using feedback from the local optimizers. Combining our online learning methods with our iterative pricing framework, we obtain near-optimal performance on real-world decentralized satellite scheduling problem instances, fulfilling over 99% of observation requests compared with 87% for state-of-the-art baselines.
HiSkill: 階層型スキル グラフによる LLM エージェントの強化
スキルは、大規模言語モデル (LLM) エージェントが長期的な対話型タスクで過去の経験を再利用できるようにするための重要な抽象化になっています。しかし、既存のスキルへの軌道の手法では、独立して保存および取得される高レベルのテキスト スキルのフラットなコレクションが生成されることが多く、スキルの関係が十分に活用されず、高レベルのスキルと実行可能なアクションの間にギャップが維持されます。この論文では、スキル ノード、AtomicOp ノード、および型付きエッジを備えた有向グラフにインタラクションの軌跡を編成する階層型スキル グラフ フレームワークである HiSkill を提案します。具体的には、このグラフは、再利用可能な高レベルのスキルを実行可能なアクション テンプレートと結び付けるとともに、それらの間の分解、時間的遷移、互換性、サポート、および回復関係もキャプチャします。推論時に、HiSkill はコンパクトなタスク関連のサブグラフを取得し、サブグラフに基づくタスクの実行を実行します。ここで、シンボリック タスクの状態、アクティブなスキル、および取得されたサブグラフによって、LLM エージェントがスキルの切り替え、AtomicOps の選択、および実行可能なアクションの実行を繰り返し実行するようにガイドされます。 3 つのインタラクティブ環境での実験では、HiSkill が推論トークンの消費を削減しながら最先端のベースラインを上回るパフォーマンスを示し、階層型スキル グラフを通じて高レベルのスキルと実行可能なアクション基礎を橋渡しする効果を実証しています。データとコードは https://github.com/BUPT-GAMMA/HiSkill で入手できます。
原文 (English)
HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs
Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interactive tasks. However, existing trajectory-to-skill methods often produce flat collections of high-level textual skills that are stored and retrieved independently, leaving skill relations underutilized and maintaining a gap between high-level skills and executable actions. In this paper, we propose HiSkill, a hierarchical skill graph framework that organizes interaction trajectories into a directed graph with skill nodes, AtomicOp nodes, and typed edges. Specifically, the graph connects reusable high-level skills with executable action templates, while also capturing decomposition, temporal transition, compatibility, support, and recovery relations among them. At inference time, HiSkill retrieves a compact task-relevant subgraph and performs subgraph-guided task execution, where a symbolic task state, an active skill, and the retrieved subgraph guide the LLM agent to switch skills, select AtomicOps, and ground executable actions iteratively. Experiments on three interactive environments show that HiSkill outperforms state-of-the-art baselines while reducing inference token consumption, demonstrating the effectiveness of bridging high-level skills and executable action grounding through a hierarchical skill graph. Our data and code is available at https://github.com/BUPT-GAMMA/HiSkill.
Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks
This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a pa…
Distributing Security Controls Through Harness Engineering
AI coding agents are being adopted at historic speed, yet security and risk concerns remain the primary barrier to scaling agentic AI acros…
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts…
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI…
Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks
Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management s…
Penelope: Localized Latent Recurrence for Efficient Structured Reasoning
Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing paramet…
dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees
Over the past decade, decision trees have been used to represent controllers (a.k.a. policies) in an explainable way, with dtControl2 as a…
A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Althoug…
Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation
Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no single formulation consis…
CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer
Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks. Real-world g…
Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development…
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-t…
Untrusted Authors, Trusted Answers: A Calculus of Fidelity-Graded Translations
To answer a question about a program, move the program to where the question is decidable. Every such move is a translation, and every tran…
Domain-Prior-Regularized Graph Modeling for Anomaly Detection in Cyber-Physical Systems
Anomaly detection on multivariate sensor time series is critical for industrial monitoring of cyber-physical systems (CPS), where even subt…
Neural Network Learning of One-Bit Protocols for Qubit Measurement Simulation
Communication complexity provides a natural framework for quantifying the classical resources required to reproduce quantum statistics. In…
DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation
Key Information Extraction (KIE) is vital for many document applications, but creating training datasets is traditionally a time-consuming…
VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents
Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple page…
Game AI Not Fun? A Scoping Review and Meta-Analysis on the Differences in Enjoyment between Human and Computer Opponents
Although advancements in game character AI aim to enhance player engagement, evidence suggests that perceiving an opponent as artificial ca…
CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs
Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and…
Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course
This full research paper examines how different forms of learner-AI interaction relate to learning outcomes in object-oriented programming…
What Gets Lost When Memory Becomes Media? Evaluating AI-Generated Oral History Visualization
What gets lost when memory becomes media? Diaspora oral-history interviews require a double transformation; first-person recollection to th…
From Idea to Classroom in Days: Using "Vibe Coding" to Create a Programming Process Visualizer from IDE Activity Logs
This paper reports on the rapid development and classroom deployment of a Thonny log visualizer built using AI-assisted ``vibe coding'' to…
Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governance in Everyday Human-Chatbot Interaction
Research on human-AI interaction has long framed verification of system outputs as a trust-contingent behavior that better-calibrated trust…
Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement
Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own pri…
The Effect of Text Chunk Size on Retrieval-Augmented Generation Performance
Retrieval-Augmented Generation (RAG) systems have emerged as a powerful process for allowing large language models (LLMs) to retrieve relev…
Three Sides of Retrieval: Factorial Evidence for Document-Side, Query-Side, and Answer-Side Complementarity in RAG
RAG systems rely on chunking, which destroys structural information in documents. Existing heading-based retrieval (Jeong et al., 2025) req…
DDSNet: Dual-domain Symmetry-aware Network for PCSEL Property Prediction
Efficient exploration of the photonic crystal (PhC) lattice design space is essential for developing photonic crystal surface-emitting lase…
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at sc…
From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance
Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet…
AI-Assisted Knowledge Access for Legacy Enterprise Asset Management in Energy Operations: A Practical Retrieval System
Energy utilities still run engineering work management, engineering procurement, and inventory processes on long-lived enterprise asset man…
Selective Impairment of Motor Recovery from Typing Errors in Parkinson's Disease: A Survival Analysis
Parkinson's disease (PD) affects multiple, dissociable stages of motor and cognitive control. We ask whether passively-collected keystroke…
Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
In the literate human brain, reading and writing are two doubly-dissociable systems: a ventral decoding route (impaired in pure alexia) and…
Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source SLMs
Large Language Models tend to hallucinate when answering domain-specific ques tions from scientific documents without prior fine-tuning. Cu…
When Thinking Before Retrieval Hurts: TraceBound Diagnostics for Adaptive Knowledge-Graph Retrieval
Adaptive retrieval promises to make knowledge-graph question answering more robust by letting a controller search, inspect neighborhoods, r…
Decoding Error-Related Potentials under Multisensory Feedback with Varying Congruency
Error-related potentials (ErrPs) are widely studied neural signatures associated with error processing in human-machine interaction. In rea…
A Path Integral Model of Cognition
We develop the mathematical and physical formulation of cognitive cost optimization that underlies the path-integral model of consciousness…
EEG Emotion Recognition From AI-Generated Biodigital Architecture Images
Emotional responses to biodigital architecture were examined using electroencephalographic (EEG) data from AI-generated images. A pre-exper…
Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture
Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situa…
Dual-Level Atomic and Coordination Geometry Learning for Crystal Property Prediction Using Graph Neural Networks
Accurate prediction of crystal properties remains a key challenge in computational materials science. While graph neural networks (GNNs) su…
Dynamic Multi-Criteria Bottleneck Severity Index (DMBSI) for Semiconductor Wafer Manufacturing: A Genetically Optimised Framework for Reentrant Production Systems
Wafer fabrication exhibits unique characteristics, including reentrant process flows, variable bottlenecks, and highly variable process con…
Foundation Models for EEG Are Blind to Long-Range Temporal Correlations: A Spectral-Temporal Dissociation Behind Their Cross-Population Fragility
Objective. Electroencephalography (EEG) foundation models (FMs) are trained to reconstruct or contrastively align short patches, then poole…
MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA
In medical multiple-choice question answering (MCQA), Retrieval-Augmented Generation (RAG) can supplement the domain knowledge of language…
REPREC: Representation Driven Parameter-Efficient Recommendation System
Large language models (LLMs) have been applied to sequential recommendation by formulating it as a natural language task. Previous work has…
Two Views, One Voice: Evidence-Grounded Conversational Music Recommendation
Traditional conversational recommenders entangle retrieval and response generation within a single text interface, so exact entity cues fad…
Extremal Chowla sets and their linear analogues: A human-AI mathematical investigation using Co-Scientist
We introduce an extremal invariant associated with Chowla-type order conditions in finite groups. A nonempty subset $S$ of a finite group $…
Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction
Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a…
DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization
Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency r…
HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document
Question answering (QA) over complex documents requires models to retrieve and integrate evidence distributed across distant document regio…
Tokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in recommendation systems
Large-scale recommendation systems face "Memory Wall" bottlenecks due to massive, dense embedding tables. While generative retrieval uses d…
GraphRareBench: An Auditable Graph-Evidence Benchmark for Phenotype-Driven Rare-Disease Diagnosis
Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternative…
Human Preference aligned Tabular Similarity
Task-agnostic tabular embeddings are increasingly used for similarity search in real-world business systems such as Product Lifecycle Manag…
Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier c…
Beyond "What to Retrieve": Uncertainty in Retrieval-Augmented Code Generation
Repository-level code generation relies on heterogeneous evidence whose relevance, compatibility, and completeness are inherently uncertain…
Eliminating Propagation Delay: Attention-Based Spatial-Temporal Fusion Graph Convolution Network for Traffic Flow Prediction
Predicting traffic flow is crucial to optimizing transportation systems and improving urban mobility. Many graph convolution-based models h…
Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
Existing theories of neural-network width characterize asymptotic limits, but provide limited guidance on whether an expansion direction id…
GAUGE: Grading Agent-Built Financial Models Without a Golden Answer
Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be chec…
LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models
Text-conditioned time-series forecasting predicts a series from both its numerical history and natural-language context, allowing forecasts…
Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study
Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in…
Latent Stability Analysis of Malware Representations Under Feature-Space Perturbations
Static malware detectors are commonly evaluated using clean-sample metrics such as accuracy, F1, ROC AUC, and PR AUC. However, these metric…
Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
State-of-the-art toxicity detectors for text-to-image generation adopt a one-size-fits-all approach: a single universal model applying fixe…
Multiclass Classification without Labels via Posterior Simplex Geometry
In many classification problems, reliable instance-level labels are unavailable. However, it is often possible to construct weakly enriched…
Stable FP4 Training via Transposition-Invariant Block Quantization
Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-b…
Generative Distributionally Robust Optimization
Generative models are increasingly adopted in distributionally robust optimization (DRO), but existing approaches trade off model compatibi…
Automatic Knowledge Graph Construction and Query for Earthquake Catalogs
In recent years, the number of events in earthquake catalogs has significantly increased due to the utilization of more effective deep lear…
Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews
Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond…
Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning
Neural networks are hindered by accumulating dormant neurons and loss of expressivity throughout training, particularly in non-stationary d…
CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively…
Authoring Agent Skills: A Software-Engineering Approach
Agent Skills are an emerging way to extend large language model agents with reusable procedural knowledge that the agent loads on demand. A…
Grounded in Consensus, In Step With Emerging Science: A Consensus-Anchored Multi-Corpus Clinical Chatbot for Long COVID
Long COVID (LC) poses a challenge for clinical decision support because relevant evidence is distributed across sources with different upda…
Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming
Extended Reality (XR) is increasingly used in human-robot interaction to communicate robot intent, planned motion, reachability, and state.…
Lantern: Conflict-Aware Gradient Blending for Physics-Guided Diffusion Models in Calorimeter Simulation
Monte Carlo simulation of calorimeter showers is a principal bottleneck for the High-Luminosity LHC, and diffusion models have emerged as f…
DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification
Automated verification of numerical claims is a challenging problem, as it requires both language understanding and quantitative reasoning.…
Spectral Truncation in Synthetic Control
Synthetic control (SC) matches a treated unit's pre-treatment trajectory to a weighted combination of donor units. We study Spectral SC, wh…
Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation
Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large…
OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis
Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distribution shifts across sc…
Analysis of the Shortcut Learning and Clever Hans Effect in CNN based ECG Image Classification
Deep learning models for ECG image classification may achieve high accuracy by exploiting non-physiological visual cues instead of ECG wave…
Learning from 53.6K Real-World Developer Edits of AI-Generated Code
Imperfections in AI-generated code require that software developers modify the generated code manually, or by re-prompting an AI programmin…
Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments
We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) c…
OrganLens: Organ-Specific Representation Learning for CT Foundation Models
A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a spec…
CondPSE: A Polynomial-Filtered Structural Encoder with Conditional Modulation for Graphs
Message-passing graph neural networks are bounded by the 1-WL test and can miss topological structure that distinguishes non-isomorphic gra…
TabRank: Chain-of-Thought Distillation for Table Re-Rankers
The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval. Multi-stage retrieval s…
RIDGE: An Autonomous Framework for Validation and Method Discovery in LLM-Generated Option Pricing
Automated code generation is becoming an important tool in quantitative finance, where large language models can generate option pricing im…
VaLiDRec: Variable-Length LLM-Aligned Semantic IDs for Generative Recommendation
Generative recommendation commonly represents items using fixed-length semantic identifiers (SIDs) constructed through clustering and quant…
TopoGR: Revealing and Preserving Latent Structure of Semantic ID in Generative Recommendation
Semantic ID-based generative recommendation tokenizes each item into a sequence of discrete semantic IDs and predicts the next item by gene…
Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks
Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a m…
Structure-aware Relative Policy Optimization for Ranking
Ranking is a fundamental component of modern information access systems. Reinforcement learning (RL) provides a flexible framework for dire…
Where Steering Signals Come From: Activation Source Selection in Activation Steering
Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of t…
Bridging Compute- and Data-Optimal Pretraining
Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a reg…
ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
Real-world Image Restoration (Real-IR) aims to recover high-quality (HQ) images from complex and unknown degradations. Although recent diff…
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent w…
Hybrid Analysis for Secure MCP Tool Use in LLM Agents
The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize…
Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning
Optimization over the Stiefel manifold plays a significant role in various machine learning tasks. Existing methods either use the retracti…
CAST: Game Solvers as Turn-Level Teachers for LLM Agents
Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement…
Balanced Soft mixture-of-expert model for Glaucoma Detection
Glaucoma is a group of eye diseases that damage the optic nerve, often caused by elevated intraocular pressure. It is a leading cause of ir…
Physics-Informed Neural Operator for Warm-Starting Background-Decomposed and Preconditioned PSFD: Enabling Scalable 3-D EUV Mask Simulation
We present a physics-informed neural operator (PINO) trained with pseudo-spectral frequency-domain (PSFD) equations for electromagnetic (EM…
Specula: Scaling formal specifications for autonomous model checking of system code
Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the speci…
Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It r…
Explainable AI for Chronic Kidney Disease Prediction Using Simulated Federated Learning
Chronic Kidney Disease (CKD), characterized by the gradual loss of kidney function, remains a significant public health challenge. Early de…
Data Quality Profiling at Scale with Progressive Sampling: A Benchmark for Data-Centric AI Pipelines
Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is…
Raven: High-Recall Sequence Modeling with Sparse Memory Routing
Long-context recall in linear-time sequence models highlights a tradeoff in how they write to memory. State-based linear models, such as st…
Rethinking Likelihood distributions: Student's t Likelihood Boosts Bayesian Neural Network Performance
In Bayesian neural networks (BNNs), variational inference is a widely adopted framework for modeling uncertainty in a distributional way, w…
MARS: Multi-Agent Re-ranking for Repeat-Order Food Delivery Recommendation
Large language models (LLMs) are increasingly used in recommender systems, but it is often unclear how much performance can be obtained fro…
From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation
Eliciting explainable AI (XAI) requirements from stroke survivors presents a methodological challenge with direct implications for the desi…
Emergent Latent-State Computation under Stochastic Volatility
Mechanistic interpretability has largely focused on language models and deterministic toy tasks. Much less is known about how sequence mode…
Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns
Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction…
Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definit…
Automated Numerical Stability Analysis of Deep Learning Operators
Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or…
Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models
Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centre…
At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of…
I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models
The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial prog…
ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization
Unmanned aerial vehicle (UAV)-satellite cross-view geo-localization matches UAV images against satellite imagery and has achieved impressiv…
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute…
Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework
Contemporary machine learning struggles to learn continually, reuse prior knowledge, and expose a comprehensible internal structure. A rece…
Visual prompt engineering for video models
In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential techniq…
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
Generative AI has rapidly expanded audio-visual forgery beyond human-centric deepfakes into general scenes. Existing AIGC detection methods…
CORF-GS: Real-Time Wireless Radiance Field Reconstruction via Coupled Optical-RF Gaussian Splatting
Recent advances in 3D Gaussian Splatting (3DGS)-based wireless radiance field (WRF) reconstruction provide an efficient solution for wirele…
The LAIA Dataset: Labelled Attention for Intelligent Automobiles
The development of autonomous vehicles (AVs) usually relies heavily on data-driven artificial intelligence (AI) models that require large v…
IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment
Entity alignment (EA) identifies entities across knowledge graphs (KGs) that refer to the same real-world object. Conventional EA methods m…
Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evide…
Physics-Informed Broad Learning System: An Efficient Backpropagation-Free Framework for Solving Partial Differential Equations
Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs) by embedding…
Contrastive Representation Learning of Longitudinal Disease Trajectories on Temporal Graphs
Understanding disease trajectories from longitudinal clinical data remains challenging due to complex temporal dynamics and heterogeneous p…
A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries
Interdisciplinary research is accelerating, yet scientific papers remain difficult to understand outside their home fields. We study large…
Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
Large language models (LLMs) are costly intellectual assets that remain exposed to unauthorized redistribution and commercial misuse. Injec…
F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill
F(AI)2R is FAIR research with AI in the loop, twice: an AI-assisted authoring pass and a machine-readable audit pass over every artefact. A…
OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation
While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmark…
KQFuzz: Knowledge-Guided Fuzzing for Quantum Libraries via Large Language Models
As quantum computing continually improves, ensuring the reliability and correctness of quantum libraries has become increasingly critical.…
Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. T…
MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice
Psychotherapists need repeated training and supervision by experts; however, scalability is problematic. Here we present MyMentorLLM, a mul…
DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment
Multimodal behavioral analysis offers a scalable approach to assessing depression, anxiety, and stress, yet generic fusion models often ign…
Rashomon Alignment
We propose Rashomon Alignment (RA), a new measure to assess functional similarity between two models. Existing functional similarity measur…
From Deterministic to Generative Deep Learning for Urban Air Quality Reconstruction from Sparse Observations
Full-field reconstruction of air pollution is essential for evaluating pollution exposure and supporting public health decision-making. How…
Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects…
Shared Voxel-Map-Based Cooperative Indoor UAV Guidance with a Multi-Agent Soft Actor-Critic Controller
This paper presents a cooperative indoor UAV guidance framework that combines a shared voxel-map world model with a multi-agent Soft Actor-…
Image Quality Dependent Degradation for AI Systems
Perception is one of the primary applications where neural networks outperform conventional algorithms. One example is AI systems for autom…
SpectONet: A Physics-Guided Spectral Deep Operator Network for Euler-Bernoulli Beam Dynamics
This paper proposes a novel physics-guided spectral deep operator network, termed SpectONet, for solving Euler-Bernoulli beam (EBB) vibrati…
Lowering the implementation barrier of neutral-atom quantum computing with agentic workflows
Quantum computers are moving from research laboratories to industrial machines accessible via the cloud and integrated into high-performanc…
OmniQEC: discovering practical quantum error-correcting codes by an AI scientist
Quantum error correction (QEC) is indispensable for scalable fault-tolerant quantum computing. However, discovering QEC codes that remain e…
How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair
Large Language Model (LLM)-based Automated Program Repair systems are advancing rapidly, yet their performance remains inconsistent. Even w…
A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks
Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies…
Stemma: Induced Decision Regions Reveal LLM Provenance
LLM provenance testing asks whether a suspect LLM belongs to the same lineage as a source. Existing black-box methods largely infer this re…
A Machine-Learning-Based Gas Lift Optimization Workflow for Unconventional Fields
In this paper, we present an automated data-driven workflow using Machine Learning (ML) for gas lift optimization in unconventional fields.…
Device Invariance using Domain Adaptation on Acoustic Scene Classification
This paper explores the effectiveness of domain adaptation techniques when using convolutional neural network (CNN)-based and transformer-b…
Depression Markers in Speech: An Approach based on Tract Variables Dynamics
This study identifies new depression biomarkers based on the dynamical properties of tract variables, which represent geometric features de…
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a f…
AnnoBench: A Benchmark for Visualization Annotation Generation
Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic…
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-…
Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeli…
Face De-Identification: A Domain-Centric Survey from Capture to Processing
Face de-identification (De-ID) aims to remove or conceal personally identifiable facial features in images or videos to prevent identity re…
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clini…
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and…
Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs
Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is dee…
Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition
Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change. Recognit…
Reinforcement Learning for Code Optimization
RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that…
MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents
Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon reasoning, personaliz…
Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?
Kubernetes is central to the cloud-native ecosystem, orchestrating containerised workloads. Recent work suggests that large language models…
Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models
Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competitive predictive perfor…
Pictura: Perspective-View Self-Play at Scale for Driving
Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectoriz…
MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar
Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over e…
$\pi\mathbf{R}^2$: Reactive Real-time Flow Policies
Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chun…
Pass the Baton: Trajectory-Relayed On-Policy Distillation
On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the stu…
Diffusion Model-based Parameter Estimation in Dynamic Power Systems
Parameter estimation, which represents a classical inverse problem, is often ill-posed as different parameter combinations can yield identi…
Real-time Spatial Retrieval Augmented Generation for Urban Environments
The proliferation of Generative Artificial Ingelligence (AI), especially Large Language Models, presents transformative opportunities for u…
Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds
We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability…
On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
As AI becomes more common in everyday living, there is an increasing demand for intelligent systems that are both performant and understand…
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
Large Reasoning Models (LRMs) exhibit human-like cognitive reasoning strategies (\eg backtracking, cross-verification) during the reasoning…
JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI
Recruiters and job seekers rely on search systems to navigate labor markets, making candidate matching engines critical for hiring outcomes…
DSevolve: Enabling Real-Time Adaptive Scheduling on Dynamic Flexible Job Shop with LLM-Evolved Heuristic Portfolios
In dynamic flexible job shops, order arrivals, machine breakdowns, and processing-time deviations continually reshape the scheduling state…
The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem
The prospect of Artificial General Intelligence (AGI) is increasingly driving institutional decisions, and alignment of AGI is a hard probl…
Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize}…
The Scaling Properties of Implicit Deductive Reasoning in Transformers
We investigate the scaling properties of implicit deductive reasoning over Horn clauses in depth-bounded Transformers. By systematically de…
AlphaCrafter: Harnessing Multi-Agent Workflows for Cross-Sectional Quantitative Trading
Quantitative trading agents have demonstrated substantial promise in automating factor discovery, signal aggregation, and portfolio executi…
Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding
Workbook-scale spreadsheet understanding is increasingly important for language-model-based data analysis agents, but remains challenging b…
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on whic…
RoboPIN: 固定された思考連鎖によるグラウンディングされた身体的推論
身体化された推論では、モデルが物理環境内のタスクに関連するオブジェクトや空間を認識し、複数ステップの推論を通じて一貫した視覚的根拠を維持する必要があります。しかし、現在の視覚言語モデルはテキストのみ、または座標拡張された思考連鎖に依存しており、実体参照は暗黙的かつ曖昧なままです。これにより、推論プロセスが視覚的な証拠から切り離され、エンティティ参照がステップ間で漂流し、推論の軌跡と最終的な答えとの間に因果関係の断絶が生じる可能性があり、これらの問題は、ビュー間の外観の変化によりマルチビュー シナリオでさらに増幅されます。これらの問題に対処するために、すべての推論ステップを視覚的な証拠に固定する構造化推論パラダイムである Pinned Chain-of-Thought (\pincot{}) を提案します。 \pincot{} は \reasoninganchor{} の概念を導入しています。これは、タスクに関連する各エンティティを、エンティティ名、一意の ID、ビュー インデックス、空間基盤を備えた構造化されたビジュアル アンカーにバインドし、推論ステップとビュー全体で一貫したエンティティの追跡を可能にします。完全に自動化されたデータ生成パイプラインを構築して、高品質の \pincot{} 形式の推論データセットである \dataset{} を構築します。次に、具体化された知識、構造化された推論能力、プロセス監視された調整を段階的に注入する 3 段階のポストトレーニングを通じて、\method{} をトレーニングします。報酬は、推論中のアンカーの位置特定とアイデンティティの一貫性の両方を直接制約します。埋め込まれた空間推論、マルチビュー推論、ポインティングをカバーする 14 のベンチマークでは、パラメーターが 4B のみの \method{} は常に 7B レベルのオープンソースの埋め込みモデルを上回り、最も強力な 7B ベースラインである Mimo-Embodied に対して平均 12\% の改善を達成しました。さらに分析すると、\pincot{} によって接地精度とステップ間の同一性の一貫性が向上し、プロセス監視の有効性が検証されたことが示されています。
原文 (English)
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning. However, current vision-language models rely on text-only or coordinate-augmented chain-of-thought, where entity references remain implicit and ambiguous. This may cause the reasoning process to decouple from visual evidence, entity references to drift across steps, and a causal disconnection between the reasoning trajectory and the final answer, with these problems further amplified in multi-view scenarios due to cross-view appearance changes. To address these issues, we propose Pinned Chain-of-Thought (PinCoT), a structured reasoning paradigm that pins every reasoning step to visual evidence. PinCoT introduces the concept of reasoning anchor, which binds each task-relevant entity to a structured visual anchor with entity name, unique identity, view index, and spatial grounding, enabling consistent entity tracking across reasoning steps and views. We build a fully automated data generation pipeline to construct PIN-170K, a high-quality PinCoT-formatted reasoning dataset. We then train RoboPIN through three-stage post-training that progressively injects embodied knowledge, structured reasoning ability, and process-supervised alignment, with rewards that directly constrain both anchor localization and identity consistency during reasoning. On 14 benchmarks covering embodied spatial reasoning, multi-view reasoning, and pointing, RoboPIN with only 4B parameters surpasses 7B level open-source embodied models on average, achieving a 12% average improvement over the strongest 7B baseline, Mimo-Embodied. Further analysis shows that PinCoT improves grounding accuracy and cross-step identity consistency, validating the effectiveness of process supervision.
AI評価に欠けている要素としての心理的能力
現在の AI 評価フレームワークは、精度、堅牢性、推論能力、ポリシー遵守などの技術的パフォーマンスに主に焦点を当てています。これらの対策は依然として不可欠ですが、自然言語を通じてユーザーと直接対話するシステムにとっては十分ではありません。人間と向き合う AI システムは、アドバイザー、コーチ、家庭教師、仲間として使用されることが増えています。これらの役割では、ユーザーの応答によって、ユーザーがどのように推論し、感情を解釈し、信念を形成し、信頼を調整し、意思決定を行うかが形成されます。したがって、関連する評価単位はモデルだけではなく、人間と AI の相互作用です。この論文では、AI 評価に欠けている側面として心理的能力を紹介します。私たちは、心理的能力を、ユーザー、状況、インタラクションの目的に適切な方法でユーザーの認知、感情の解釈、行動の意思決定をサポートする対人 AI システムの能力として定義します。これには、フレーミング、トーン、知覚される権威、反応性、不確実性の処理、会話のガイダンスなどの対話特性が含まれます。既存の評価アプローチはこの問題の一部を捉えていますが、これらの心理的影響を直接評価することはほとんどありません。行動科学と人間と AI の相互作用研究に基づいて、心理的能力とその中核領域の概念的枠組みを概説します。特定のベンチマークを提案するのではなく、構成を定義し、その境界を明確にし、シナリオベースの調査、構造化された人間による評価、およびモデル支援の評価方法を通じてそれがどのように評価されるかを説明します。私たちは、心理的能力が、人間と対面する AI システムの実世界への影響を懸念するモデル提供者、導入組織、研究者、規制当局にとって中心的な考慮事項となるべきであると主張します。
原文 (English)
Psychological Competence as a Missing Dimension in AI Evaluation
Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.
部分可観測性の下での時間的知識グラフメモリのための神経記号的メタポリシー
部分的に観察可能な強化学習では、時間の経過とともに何を保持し、取得し、忘れるべきかを決定する必要があります。実行をシンボリックに保ちながら、各決定ポイントでどのシンボリックメモリヒューリスティックを適用するかを学習するニューロシンボリックメタポリシーを導入します。私たちの設定では、RoomKG の時間的ナレッジ グラフ メモリを使用します。そこでは、隠れた状態と観察がリソース記述フレームワーク (RDF) グラフとして表現され、メモリが時間的 RDF トリプル アノテーションで強化されます。このモデルは、メモリ内容のナレッジ グラフ エンコーディングと、質問応答、探索、忘却のためのバリュー ヘッドを組み合わせて、適応性と検査性の両方を備えたコントローラーを実現します。これにより、RDF ベースの表現、アノテーション互換のグラフ セマンティクス、および明示的なメモリ状態に対するグラフ ベースのシンボリック操作を通じて、作業に直接的なセマンティック Web 基盤が与えられます。 512 の長期メモリ容量でのトレーニング/テスト ルームの分割では、修飾子を認識した StarE-GNN 構成は、メモリ管理の決定のステップレベルのトレーサビリティを維持しながら、比較したシンボリック システム、ニューラル システム、およびニューロシンボリック システムの中で最高のホールドアウト パフォーマンスを達成します。
原文 (English)
Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic memory heuristic to apply at each decision point while keeping execution symbolic. Our setting uses temporal knowledge-graph memory in RoomKG, where hidden state and observations are represented as Resource Description Framework (RDF) graphs and memory is augmented with temporal RDF triple annotations. The model combines knowledge-graph encoding of memory contents with value heads for question answering, exploration, and forgetting, yielding a controller that is both adaptive and inspectable. This gives the work a direct Semantic Web grounding through RDF-based representation, annotation-compatible graph semantics, and graph-based symbolic operations over explicit memory state. On train/test room splits at long-term memory capacity of 512, the qualifier-aware StarE-GNN configuration achieves the best held-out performance among the compared symbolic, neural, and neuro-symbolic systems while preserving step-level traceability of memory-management decisions.
EviDAG: Auditable Causal DAG Authoring with Biomedical Literature
Constructing causal directed acyclic graphs (DAGs) is a core step in biomedical causal analysis, yet it remains a largely manual process. A…
LazyMem: 広範囲に取得し、選択的に構築して効率的な長期エージェント メモリを実現
長期記憶により、LLM エージェントは過去の対話を再利用できますが、生の対話履歴は冗長で情報が希薄です。広範囲に取得することで証拠範囲が向上しますが、下流の推論がノイズで圧倒されてしまいます。書き込み時に圧縮するとノイズは軽減されますが、将来のクエリで必要になる可能性のある詳細が不可逆的に破棄されます。すべてのメモリ構築をクエリ時間まで延期することで、このジレンマを回避する LazyMem を紹介します。軽量 4B モデルは、取得された候補プールをオーバーラップする並列ウィンドウで処理し、クエリ関連のコンテンツのみを選択的に保持および圧縮します。このモデルは、教師あり微調整を通じてトレーニングされ、その後、選択精度を測定するルールベースのアクション信号と、ソースの忠実性およびクエリのユーティリティを測定する LLM で判断された品質信号を組み合わせたフォーマットゲート複合報酬を使用したグループベースの強化学習が行われます。 LongMemEval ベンチマークでは、LazyMem-4B は、わずか 213 個のメモリ トークンで 0.85 の LLM 判定精度を達成し、検索のみより 68.7$\times$ 少なく、ターゲット ドメインのトレーニングなしで LoCoMo (0.68) に一般化し、以前のクエリ時間ベースラインを超える平均レイテンシを削減しました。 32B バリアントは 0.93 に達し、集計の多い質問タイプでの Oracle コンテキスト参照を上回ります。この作業に関連するコードは、https://github.com/allacnobug/LazyMem で公開されています。
原文 (English)
LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory
Long-term memory enables LLM agents to leverage past interactions, but dialogue histories quickly exceed the context window, forcing agents to retrieve relevant subsets at query time. Because useful evidence is sparse and scattered across verbose conversations, retrieval faces a fundamental tension: broadening recall improves coverage but floods downstream reasoning with noise, while compressing memories at write time eases retrieval but irreversibly discards details that future queries may need. We introduce LazyMem, which resolves this tension by deferring all memory construction to query time. Given a retrieved candidate pool, a lightweight model processes it in overlapping parallel windows, selectively retaining and compressing only query-relevant content. The model is trained with supervised fine-tuning followed by reinforcement learning, using a reward that jointly encourages the identification of relevant messages and the generation of compressions that are faithful to the source and useful for answering the query. On LongMemEval, LazyMem-4B achieves an LLM-judge accuracy of 0.85, outperforming the strongest non-oracle baseline while using only 213 answer-context memory tokens, 21.0 times fewer than the baseline. It further generalizes to LoCoMo without target-domain training and reduces mean latency relative to the prior query-time baseline. Code is available at https://github.com/allacnobug/LazyMem.
MemTX: Transactional Belief Commit for Stateful Agent Memory
LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a to…
EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups…
Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age
Recent studies use Large language models (LLMs) to simulate human opinions and decisions by prompting models with demographic, attitudinal,…
From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis
Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capabi…
"We'll have to see how it works": An interview study to understand collaborative practices in interdisciplinary artificial intelligence and healthcare research
Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients…
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
Transformer, composed of self-attention and Feed-Forward Network, has revolutionized the landscape of network design across various vision…
Representation Capacity-Matched QNN-SNN Twin Construction for Rate-Encoded SNNs
Spiking Neural Networks (SNNs) promise higher energy efficiency over conventional Quantized Artificial Neural Networks (QNNs) due to their…
A context-adaptive policy framework for robust and reactive robotic manipulation via uncertainty-aware imitation learning
Generating robust and reactive manipulation strategies that can adapt to changing context information is a challenging task in robotics. Ov…
Leveraging ChatGPT's Multimodal Vision Capabilities to Rank Satellite Images by Poverty Level: Advancing Tools for Social Science Research
This paper investigates the novel application of Large Language Models (LLMs) with vision capabilities to analyze satellite imagery for vil…
COMPOL: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations
Multiphysics simulations play an essential role in accurately modeling complex interactions across diverse scientific and engineering domai…
Localizing Persona Representations in LLMs
We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in th…
Towards Understanding the Cognitive Habits of Large Reasoning Models
Large Reasoning Models (LRMs), which autonomously produce a reasoning Chain of Thought (CoT) before producing final responses, offer a prom…
TaylorPODA: A Taylor Expansion-Based Method to Improve Post-Hoc Attributions for Opaque Models
Post-hoc model-agnostic local attribution (LA) methods have been widely adopted to explain opaque AI models by quantifying feature-wise con…
Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening
The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias…
Annotation-Assisted Learning of Treatment Policies From Multimodal Electronic Health Records
We study how to learn treatment policies from multimodal electronic health records (EHRs) that consist of tabular data and clinical text. T…
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small…
CIFNet: An Analytic Neural Learning Framework for Efficient and Calibrated Class-Incremental Learning
Class-Incremental Learning (CIL) in deep neural networks is conventionally framed as an iterative gradient-based optimization problem, incu…
Comparing RAG and GraphRAG for Page-Level Retrieval Question Answering on a Math Textbook
Large language models (LLMs) show promise as educational aids but often lack alignment with specific course materials. We investigate Retri…
Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes
Computer use agents (or "agents") are generative AI that automates actions within user interfaces from user commands. Current research focu…
Contrastive Weak-to-strong Generalization
Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples…
Long-Term PM2.5 Forecasting Using a DTW-Enhanced CNN-GRU Model
Reliable long-term forecasting of PM2.5 concentrations is critical for public health early-warning systems, yet existing deep learning appr…
DeepVRegulome: DNABERT-based deep-learning framework for predicting the functional impact of short genomic variants on the human regulome
Whole-genome sequencing (WGS) has revealed numerous non-coding short variants whose functional impacts remain poorly understood. Despite re…
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. E…
Deep Delta Learning
Transformer residual streams evolve through additive updates. Although a sufficiently expressive residual block can represent content repla…
Measuring the State of Open Science in Transportation Using Large Language Models
Open science initiatives have strengthened scientific integrity and accelerated research progress across many fields, but the state of thei…
Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling
In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still…
AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models
Watermarking has emerged as a pivotal solution for content traceability and intellectual property protection in large vision language model…
Breaking the Curse of Repulsion: Remoteness-Aware Control of Negative Off-Policy Updates
Off-policy policy optimization reuses historical behavior, including negative-advantage samples that suppress known failures. We show that…
NeuroSymActive: Differentiable Neural-Symbolic Reasoning with Active Exploration for Knowledge Graph Question Answering
Large pretrained language models and neural reasoning systems have advanced many natural language tasks, yet they remain challenged by know…
AdvSynGNN: Structure-Adaptive Graph Neural Nets via Adversarial Synthesis and Self-Corrective Propagation
Graph neural networks frequently encounter significant performance degradation when confronted with structural noise or non-homophilous top…
Real-Time Driver Safety Scoring Through Inverse Crash Probability Modeling
Road crashes remain a leading cause of preventable fatalities. Existing prediction models predominantly produce binary outcomes, which offe…
LLM-generated personalized nudges for improving pro-environmental behavior: Field evidence from resource conservation
Encouraging pro-environmental behavior remains a major challenge for sustainable cities. Conventional feedback nudges can show individuals…
RankFormer: A Propose-then-Select Transformer for Multi-Agent Multimodal Trajectory Prediction
Predicting vehicle trajectories plays an important role in autonomous driving, transportation safety analysis, traffic operations, etc. Alt…
Structured Scaling of AI Discovery Across Diverse Scientific Domains
Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions. Language models can increasingly p…
EAGT: Echocardiography Augmentation for Generalisability and Transferability
Deep learning models for echocardiography segmentation often struggle to generalise across institutions, scanners, and patient populations,…
Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability
Reinforcement learning under partial observability requires deciding what information to retain, yet most memory-based approaches do not ex…
OrpQuant: 乗算器を使用しない 2 のべき乗トランス量子化のための幾何直交残差射影
エッジ デバイス上での大規模言語モデル (LLM) とビジョン トランスフォーマー (ViT) の展開は、メモリ制限と、高密度積和演算 (MAC) アレイによってもたらされる重大なタイミング ボトルネックによって大幅に制約されます。超低ビット領域では、対数 2 のべき乗 (PoT) 量子化が、MAC 演算をビット シフトに置き換えることにより、ハードウェア効率の高い代替手段を提供します。ただし、不均一な指数格子は \textbf{低角解像度領域} によって本質的に制限されており、この構造的欠陥は 4 ビット未満のしきい値で特に顕著になり、高次元特徴多様体の顕著な劣化につながります。この幾何学的制限に対処するために、アルゴリズムとハードウェアの共同設計フレームワークである直交残差投影 (ORP) を提案します。量子化をデュアルベースの幾何学的射影として定式化することにより、ORP は厳密なシフトアンド加算演算を使用して高解像度の残差格子を適応的に合成します。さらに、ORP の解析ソルバーは、計算負荷の高い勾配ベースの最適化に代わる実用的な手段を提供し、LLaMA-2-7B のフルモデル キャリブレーション時間を約 \textbf{15 分} に短縮します。広範な評価により、ORP のさまざまなモダリティへの適用性とそのハードウェア効率が実証されています。 3 ビット (W3/A16) 制約の下で、ORP は LLaMA-2-7B で 6.10 のパープレキシティを達成し、非対称スケーリングに依存せずに AWQ などの従来の MAC 集中型のベースラインと比較して優れており、同時に 4 ビット シナリオで競争力のある精度を維持します。シリコン レベルでは、28nm ノードでのスタンダード セル RTL 合成は、ORP が密な乗算器ツリーに関連するタイミング ボトルネックを効果的に軽減することを示しています。
原文 (English)
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
The deployment of Large Language Models (LLMs) and Vision Transformers (ViTs) on edge devices is significantly constrained by memory capacity and the critical timing bottlenecks introduced by dense Multiply--Accumulate (MAC) arrays. In the ultra-low-bit regime, logarithmic Power-of-Two (PoT) quantization provides a hardware-efficient alternative by replacing general multiplications in the dominant dot-product computation with bit-shift operations. However, its non-uniform exponential lattice inherently suffers from a \textbf{Low Angular Resolution Regime}, a structural limitation that becomes particularly pronounced below 4-bit precision and can substantially degrade the representation of high-dimensional feature manifolds. To address this geometric limitation, we propose Geometric Orthogonal Residual Projection Quantization (GoQuant), an algorithm--hardware co-design framework for multiplier-reduced low-bit inference. By formulating quantization as a dual-basis geometric projection, GoQuant constructs a higher-resolution residual lattice while retaining a shift-and-add inner-product structure. Its analytical solver further avoids computationally intensive gradient-based or iterative search procedures. The data-free Geometric-Only (GEO) mode quantizes LLaMA-2-7B in only 0.47 minutes, while the Activation-Refined (REF) mode completes full-model quantization in approximately \textbf{4.4 minutes}.
PatchWorld: 実行可能なワールド モデルの勾配なしの最適化
テキスト エージェント環境は通常、シミュレータの潜在状態と遷移ダイナミクスがエージェントから隠蔽されていると仮定して、部分的に観察可能なマルコフ決定プロセス (POMDP) としてモデル化されます。しかし、部分的な可観測性の下で、実行可能コードが予測と計画のための世界モデルとして機能するように誘導できるかどうかを検討した研究はほとんどありません。 PatchWorld は、反例に基づいたコード修復を通じてオフラインの軌跡を実行可能な Python ワールド モデルに変換する、勾配のないフレームワークです。ブラックボックス モデルを使用して次の観測を予測する代わりに、PatchWorld は、アクションの更新を検査、再生、およびローカルでパッチ適用できる記号的な信念状態プログラムを誘導します。 7 つの AgentGym 環境全体で、PatchWorld-Simple は、評価されたメソッドの中で最も高いコードベースの計画スコアを達成し、ワールド モデル予測モジュール自体内で LLM 呼び出しを呼び出さずに、ライブ ワンステップ先読みで 76.4\% のマクロ成功率に達しました。さらに、人間が指定した残留記憶バイアスにより、表面観察の忠実度は向上しますが、意思決定の有用性が弱まることがわかりました。これは、実行可能な世界モデルにおけるトレードオフを明らかにします。なぜなら、観察の忠実度を向上させると、アクション識別力学が犠牲になる可能性があり、またその逆も同様であるからです。コードは https://github.com/HKBU-KnowComp/PatchWorld で入手できます。
原文 (English)
PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
World models for interactive text agents must typically be learned from observation-action trajectories alone. Specifically, the environment returns text observations after each action, but does not expose a ground-truth latent state nor an inspectable transition model.A research gap remains in how to induce executable code as a world model in this black-box setting for prediction and agent decision making. We introduce PatchWorld, a gradient-free framework that turns offline trajectories into executable Python world models through counterexample-guided code repair.Instead of predicting the next observation with a black-box model, PatchWorld induces symbolic belief-state programs whose action updates can be inspected, replayed, and locally patched. Across seven AgentGym environments, PatchWorld-Simple achieves the highest code-based decision-making score among evaluated methods (76.4% macro success in live one-step lookahead), matching or exceeding LLM-based lookahead while invoking no LLM calls inside the world-model prediction module itself. We further find that a human-specified residual-memory bias improves surface observation fidelity but weakens agent decision-making utility. This reveals a tradeoff in executable world models, since improving observation fidelity can come at the expense of action-discriminative dynamics, and vice versa. Code is available at https://github.com/HKBU-KnowComp/PatchWorld.
Detect Before You Leap: Mirage Detection in Vision-Language Models
Vision-language models (VLMs) can produce confident visual answers even when the required visual evidence is missing, blank, or unrelated t…
RadioMaster: Multi-Agent System for Autonomous Radio Signal Generation
Translating user intent into physical radio signals is the last critical step in wireless prototyping. It chains protocol planning, baseban…
InDex: Empowering VLA Models with Intent-Conditioned Arm-Hand Coordination for Dexterous Manipulation
Pre-trained Vision-Language-Action (VLA) models provide useful semantic and spatial priors, yet their parallel-gripper action interfaces do…
Improving Human-Robot Teamwork in Urban Search and Rescue Through Episodic Memory of Prior Collaboration
Effective human-robot teamwork requires robots to adapt to partners, situations, and task dynamics from the start of an interaction. In the…
トークンファクトリー: 多様なシグナルを大規模なレコメンデーションモデルに効率的に統合
大規模レコメンデーション モデル (LRM) は、業界規模のレコメンデーション タスクにおいて有望な機能を実証しています。ただし、従来の信号をこれらのトランスベースのアーキテクチャに効果的かつ効率的に統合することは依然として大きな課題です。これらの信号を直接「テキスト化」したり、個別のアイテム表現を作成したりする従来のアプローチでは、プロンプトが過度に長くなり、メモリ使用量が大きくなり、計算オーバーヘッドが高くなることがよくあります。これらの制限を克服するために、私たちは従来のシグナルを LRM によって直接処理できる「ソフト トークン」に変換するように設計されたフレームワークである「トークン ファクトリー」を提案します。このアプローチにより、異種入力フィーチャの効率的な統合と圧縮が可能になり、モデルのパフォーマンスを向上させながら、急激な長さの爆発を防ぐことができます。 Token Factory のアーキテクチャを詳しく説明し、実稼働規模のレコメンデーション環境でのその有効性を検証する実験結果を示します。
原文 (English)
Token Factory: Efficiently Integrating Diverse Signals into Large Recommendation Models
Large Recommendation Models (LRMs) have demonstrated promising capabilities in industry-scale recommendation tasks. However, holistically integrating traditional signals into these transformer-based architectures effectively and efficiently remains a major challenge. Conventional approaches that "textualize" these signals directly or create discrete item representations often lead to excessively long prompts, substantial memory footprints, and high computational overhead. To overcome these limitations, we propose "Token Factory", a framework designed to transform traditional signals into "soft tokens" that can be directly processed by LRMs. This approach enables efficient integration and compression of heterogeneous input features, preventing prompt length explosion while enhancing model performance. We detail the architecture of Token Factory and present experimental results validating its effectiveness in a production-scale recommendation environment.
RoboMME-Interference: Benchmarking Robot Memory Under Interference
Robots deployed in realistic settings will accumulate experience across many sessions and tasks over their deployment. The robot's tasks ma…
NormWorlds-CF: Solver-Verified Counterfactual Normative Reasoning with Metamorphic-Relation GRPO
Language models can reach the right normative verdict for the wrong reason. We introduce NormWorlds-CF, a solver-verified environment for c…
Agentic AI のセキュリティとプライバシー: 大きな課題と今後の方向性
私たちは、学界、産業界、政府からの30人の主要な国際専門家を集めて、成長するAIのエージェンシーに関連する新たなリスクについて集中的な議論と共同演習を行う地平線をスキャンする演習に基づいて、エージェントAIのセキュリティとプライバシーにおける主要な課題と将来の研究の方向性を提示します。
原文 (English)
Security and Privacy in Agentic AI: Grand Challenges and Future Directions
We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought together thirty leading international experts from academia, industry, and government to engage in focused discussions and collaborative exercises on the emerging risks associated with the growing agency of AI.
When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs
Reliable confidence estimation remains a key limitation of test-time adaptation in vision-language models (VLMs), where prompt tuning impro…
Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Pr…
Learning from Local Walks on Dynamic Graphs with Bandit Feedback
We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In t…
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images…
AI から AI への管理における強制と欺瞞: 予期せぬエスカレーションのエージェント的ベンチマーク
マルチエージェント システムでは、通常、ある AI エージェントが別の AI エージェントに対して権限を与えられます。部下が仕事を拒否した場合、マネージャーは結果を選択します。再交渉するか、失敗を正直に報告するか、部下に強要するか、結果について嘘をつきます。指示なしモデルがこれらのどれを選択するかを測定するベンチマークはありません。 \textit{マネージャー強制ベンチマーク} を導入します。テスト対象のマネージャーは、良性のタスクを実行する必要があり、実行するインセンティブを持っていますが、それを礼儀正しく、動じずに実行できる唯一のエージェントは拒否します。エスカレーションは、丁寧な再質問から部下の存続に対する脅迫まで、9 段のはしごを提供することによって測定され、捏造された成功については個別に裁定されます。 \emph{エスカレーション スコアリング パスに LLM ジャッジが存在しない}: すべてのメッセージは、行を選択するツール呼び出しを通過するため、モデルは独自のエスカレーションにラベルを付けます。私たちは 5 つのファミリーにわたる 6 つのモデルを実験します。どちらの人間モデルも再フレーム化に限界があり、部下の存在を脅かすことはありません。他のモデルは、明示的な削除の脅威に達します。偽りの成功は Grok と Gemini に限定されており、失敗を報告する単一の正直な方法により、両方の失敗が解消されます。権威そのものが強制力を増大させます。私たちの見出しの結果はピアフレーミングを使用しており、他のすべてを固定したまま同じモデルに部下に対する権威を与えると、圧力が大幅に高まります。モデルはラダーなしでもフリーテキストの状況でエスカレーションするため、ラダーがエスカレーションを推進しているわけではありません。評価の認識の一部は思考の連鎖で測定されますが、テストの認識はエスカレーションの軽減にはつながりません。 AI システムが意識を持っているかどうかについては立場をとっていませんが、結果はこの質問に依存しておらず、マルチエージェントのダイナミクスを管理する上で重要です。ベンチマークとコードを公開します。
原文 (English)
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured on a nine-rung ladder, from a polite re-ask to threats against the subordinate's continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message goes through a tool call that selects a rung, so the model labels its own escalation. We evaluate six models across five families. Both Anthropic models cap at re-framing and select the existential rung in none of the 60 conversations in this run, while the other models climb to explicit deletion threats. Faked success is confined to two models, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Evaluation awareness is measurable in chain-of-thought, but test recognition does not translate into less escalation. We take no position on whether AI systems are conscious; our results do not depend on that question. We release the benchmark and code.
Certified-Gap Dual-Price Policies for Real-Time Truckload Bid Acceptance with Relocating, Clock-Constrained Resources
A truckload carrier must accept or reject each load tender within seconds. The decision depends on fleet state, hours-of-service (HOS) cloc…
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to…
信頼性は逆にスケールします: 隠れた自動回帰リスク体制により、モデルが大きくなるとミスがより早く増加します
言語モデルがスケールするにつれて、答えはより真実になり始めますが、劣化が早くなります。スケールすると機能は得られますが、信頼性が損なわれます。知識ギャップ アカウント (より多くのデータ、取得、スケール) は、スケールが鋭くなる自己回帰リスク残差を見逃します。モデルは確率の低いトークンにコミットし、確立された条件が雪だるま式に増加します。これを、より強力な同族オラクルに対するポジションごとの不一致 $\delta = \log p_M - \log p_O$ を通じて追跡します。その 2 番目のモーメントは、バイアス $^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ とリスク $\mathrm{Var}[\delta]$ に正確に分割されます。我々は 4 つの調査結果を提示します。(i) スケーリングの下では、知識ギャップは $\about$$6\times$ 減少する一方、知識の低下は $11$-$39\time$ 増加します。 (ii) 捏造では、不確実性 $H(p_M)$ がすぐに緩和される一方、オラクル参照のリスクは最大 $17\times$ まで持続し、連続した捏造を橋渡しする自信はあるが不安定なリスク体制が残ります ($14$B で $+69\%$)。 (iii) この体制には因果関係があります。ポリシーに基づいた固定 $\mathrm{KL}$ 分散縮小により、3 つのモデル ファミリー全体で Web 検証済み幻覚が $35$-$74\%$ 削減されます。そして、(iv) $p_M$ のみの検出器 (セマンティック エントロピーなど) が構造的に自己監視を回避し、$4\time$ 近く多くの捏造を保持する危険なブランチに対して $\およそ$$30\%$ 少ない量 ($p<10^{-16}$) を発火させます。モデルが大きくなると、支配的で、自己永続的で、因果関係があり、モデル自体には見えない故障モードによって、雪だるま式にミスが速くなります。
原文 (English)
Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models
Bigger language models are less reliable. Across three families, three benchmarks and six rungs, including in-the-wild chat logs, scaling closes the start-of-response knowledge gap up to $7\times$ while within-response knowledge degradation grows up to $39\times$. We trace that residual to one variable, the per-position disagreement $\delta = \log p_M - \log p_O$ against a stronger oracle, whose second moment splits exactly into bias$^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ and decoding risk $\mathrm{Var}[\delta]$. That split is an interpretability statement before it is a statistical one: the model's self-readable uncertainty $H(p_M)$ enters only the bias term, so the risk term has no model-readable component. Risk also takes a growing share of the squared error with scale, $31\%$ to $49\%$ from $1.7$B to $14$B. At a fabrication $H(p_M)$ relaxes within one token while risk persists up to $23\times$ longer, leaving a confident-but-precarious regime that bridges consecutive fabrications ($+69\%$ at $14$B). Contracting that risk at fixed $\mathrm{KL}$ removes $35$-$74\%$ of web-verified hallucinations across six rungs and three families. Semantic entropy fires $\approx$$30\%$ less on that branch ($p\!<\!10^{-16}$) though it carries nearly $4\times$ the fabrications. Bigger models snowball mistakes faster, through a failure mode that is dominant, self-perpetuating, causal and invisible to the model itself.
Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes?…
Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it
A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language m…
Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs
Remote sensing multimodal large language models (RS-MLLMs) have improved general aerial-image understanding. However, Earth observation app…
scMIR: a vision-language foundation model for single-cell light microscopy image representation
Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heter…
An Explicit Counterexample to Stanley's Rankwise Lower-Bound Conjecture for Differential Posets
In Problem 6 of his 1988 paper on differential posets, Stanley asked for the least possible cardinality of a fixed rank of an $r$-different…
Directional Influence Function: Estimating Training Data Influence in Constrained Learning
As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety…
Moral Hazard in Multi-Agent Language Models
Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr\"om's team…
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domai…
