AIニュース 2026-07-25
自動生成: 2026-07-25 12:15 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Anthropic、「Claude Opus 5」公開 Fable 5に迫る性能を半額で――サイバー安全策は緩和、拒否時は自動フォールバックもITmedia AI+
Anthropicは、最新LLM「Claude Opus 5」を公開した。上位モデル「Claude Fable 5」に迫る知能を半額の価格…
-
海外の「Claude」や「GPT」ではダメなのか 日本企業向けai&、そのメリットは?ITmedia AI+
ai&は、日本企業向けに設計したAI推論プラットフォーム「ai& Inference」の提供を開始した。ClaudeやGPTなど海外ベンダ…
-
I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone elseTechCrunch AI
OpenAI's fancy new AI keypad will be a lot of fun for some, while man…
-
Prentis, new AI lab co-founded by Reid Hoffman, Mark Pincus in talks to raise $100MTechCrunch AI
The neolab is betting that automating routine computer tasks will soo…
-
Why Cognition bought Poke: AI personality is becoming a competitive advantageTechCrunch AI
The acquisition brings Poke’s conversational style and interaction mo…
-
As US weighs response to Chinese AI, industry urges against broad open-weight restrictionsTechCrunch AI
AI companies, including Nvidia and Mistral, urge policymakers to avoi…
-
Bluesky’s AI assistant Attie expands into an open social research toolTechCrunch AI
Users can now ask Attie questions about news, trends, and conversatio…
トピック別件数
- LLM/生成AI 172件
- 研究/論文 156件
- エージェント 80件
- 画像/動画生成 36件
- ロボティクス 16件
- ビジネス/資金調達 15件
- ハードウェア/半導体 13件
- その他 2件
- 規制/政策 2件
日本語メディア8件
ITmedia AI+ (日本語)
Anthropic、「Claude Opus 5」公開 Fable 5に迫る性能を半額で――サイバー安全策は緩和、拒否時は自動フォールバックも
Anthropicは、最新LLM「Claude Opus 5」を公開した。上位モデル「Claude Fable 5」に迫る知能を半額の価格で提供する。プログラミングやナレッジワークにおいて高い評価を獲得し、推論の深さを調整するパラメータや安全性分類器に連動する自動フォールバック…
スーパーに並んだ「ごちゃごちゃ生成AIポップ」が物議 “看板王”こと、きぬた歯科院長「これはアリ」
スーパーの青果売り場に並ぶ、生成AIで作ったとみられる派手な商品ポップがXで物議を醸している。吸血鬼や戦国武将を描いたデザインに「見づらい」との声が相次ぐ中、看板広告で知られるきぬた歯科のきぬた泰和院長は「これはアリ」と評価。その理由とは。
近畿大、入試にAIの利用認める 情報学部の総合型選抜で
近畿大学は、2027年度の情報学部の総合型選抜入学試験で、生成AIの利用を認めると発表した。提出する自己PR動画やプレゼンテーション資料などでのAIの利用方針を明示した。
AIにもサプライチェーン管理が必要? 中国AI「Kimi K3」を巡る批判でAIの調達リスクが浮き彫りに
中国の最新AIモデルを巡り、米政府高官が、Anthropicの「Claude Fable 5」をモデルの学習に利用した“不正蒸留”が行われていたと指摘した。AIモデルの調達や導入を巡るサプライチェーンリスクが浮き彫りとなっている。
メルカリ、「AI活用の最前線」明かす動画公開 「なぜCTOがCHRO兼CAIOになったのか」など13本
メルカリは、自社のAI活用について紹介する動画を公開した。7月8日に開催したイベント「Mercari AI Career Fes 2026」で実施したセッションのアーカイブ動画、全13本を公式YouTubeチャンネルで視聴できる。
海外の「Claude」や「GPT」ではダメなのか 日本企業向けai&、そのメリットは?
ai&は、日本企業向けに設計したAI推論プラットフォーム「ai& Inference」の提供を開始した。ClaudeやGPTなど海外ベンダーのAIに代わる選択肢として、AI利用コストを大幅に削減できるとしている。
開発工数見積もりの「負担が重い」をAIで解消へ 明治安田はどう実現?
有識者に頼りがちなシステム開発工数の見積もりは、AIエージェントでどこまで効率化できるのか。明治安田生命保険がPoCで検証した仕組みを見ていこう。
「2日かかる攻撃が25分に」生成AIで“爆速化”するサイバー攻撃、パロアルトの識者が警鐘
パロアルトネットワークスの染谷征良氏(チーフサイバーセキュリティストラテジスト)は、生成AIの普及で変化するサイバー攻撃の動向と企業に求められるセキュリティ対策を、ソフトバンクの年次イベント「SoftBank World 2026」の講演で紹介した。
海外メディア7件
TechCrunch AI (英語)
I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone else
OpenAI's fancy new AI keypad will be a lot of fun for some, while many others are probably not going to touch it.
Prentis, new AI lab co-founded by Reid Hoffman, Mark Pincus in talks to raise $100M
The neolab is betting that automating routine computer tasks will soon outpace coding as AI's biggest use case.
Why Cognition bought Poke: AI personality is becoming a competitive advantage
The acquisition brings Poke’s conversational style and interaction model to Cognition’s coding agent Devin, reflecting a growing belief tha…
As US weighs response to Chinese AI, industry urges against broad open-weight restrictions
AI companies, including Nvidia and Mistral, urge policymakers to avoid broad restrictions on open-weight AI models as Washington debates re…
Bluesky’s AI assistant Attie expands into an open social research tool
Users can now ask Attie questions about news, trends, and conversations on Bluesky and other apps on the AT Protocol.
Midjourney acquired the astrology app Co-Star
The AI lab Midjourney continues to expand its purview beyond image and video generation.
OpenAI’s new voice mode makes it to the ChatGPT desktop app
ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents.
公式ブログ0件
このカテゴリの新着記事はありませんでした。
論文354件
arXiv cs.AI (英語)
AINTMA: 生成インテリジェンス、安全なクラウド通信、および適応型品質分析を備えた自律型テスト管理のための Agentic AI アーキテクチャ
最新のソフトウェア品質保証には、分散クラウド環境全体で適応的な意思決定ができるインテリジェントで自律的なシステムが必要です。この文書では、従来のテスト管理を自律的な品質インテリジェンス エコシステムに変換するマルチエージェント エージェント AI システムである AINTMA (Agentic Intelligent Test Management Architecture) について説明します。 AINTMA は、クラウドネイティブのマイクロサービス インフラストラクチャ上の安全なマルチエージェント通信フレームワークを通じて調整される 6 つの専門 AI エージェント (テスト検出、リスク評価、強化学習の優先順位付け、実行オーケストレーション、生成品質インテリジェンス、クラウド セキュリティ モニター) を展開します。 Generative Quality Intelligence エージェントは、大規模な言語モデルを採用して、平易な言語による品質説明、欠陥リスクの概要、およびデータ拡張されたテスト推奨事項を生成します。 RL 優先順位付けエージェントは、テストの選択をマルコフ決定プロセスとしてモデル化し、大規模な過去のテスト実行データ (47 の機能、ローリング 36 か月のウィンドウ) から状況に応じたポリシーを学習します。安全なクラウド通信は、OAuth2/JWT 認証、暗号化されたエージェント間メッセージング、およびマルチテナント分離を備えたゼロトラスト API ゲートウェイを通じて強制されます。 18 か月にわたる 12 の異種ソフトウェア プロジェクトにわたる評価では、テストの優先順位付けの精度が 88.4% (APFD 対ランダム 51.2%、商用ベースラインが 82.1%) であることが実証されました。テストサイクル時間を 43% 短縮。欠陥回避率が 8.3% から 2.1% に減少。 9 か月の投資回収で 340% の ROI。エージェント アーキテクチャは 400 ミリ秒未満の応答時間で 50,000 以上のテスト ケースに拡張でき、生成インテリジェンス モジュールは開発者有用性評価 4.3/5.0 を達成しています。 AINTMA は、自律的なマルチエージェント調整、生成インテリジェンス、安全なスマート接続を組み合わせたエージェント AI が、クラウド規模のエンタープライズ環境におけるソフトウェア品質管理を根本的に進歩させることができることを実証しています。
原文 (English)
AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics
Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud environments. This paper presents AINTMA (Agentic Intelligent Test Management Architecture), a multi-agent agentic AI system that transforms traditional test management into an autonomous quality intelligence ecosystem. AINTMA deploys six specialized AI agents (Test Discovery, Risk Assessment, Reinforcement Learning Prioritization, Execution Orchestration, Generative Quality Intelligence, and Cloud Security Monitor) coordinated through a secure multi-agent communication framework over a cloud-native microservices infrastructure. The Generative Quality Intelligence agent employs large language models to produce plain language quality narratives, defect risk summaries, and data-augmented test recommendations. The RL Prioritization agent models test selection as a Markov Decision Process, learning contextual policies from large-scale historical test execution data (47 features, rolling 36-month window). Secure cloud communication is enforced through a zero-trust API gateway with OAuth2/JWT authentication, encrypted inter-agent messaging, and multi-tenant isolation. Evaluation across 12 heterogeneous software projects over 18 months demonstrates: 88.4% test prioritization accuracy (APFD, vs. 51.2% random, 82.1% best commercial baseline); 43% test cycle time reduction; defect escape rate reduced from 8.3% to 2.1%; 340% ROI at 9-month payback. The agentic architecture scales to 50,000+ test cases with sub-400ms response time, and the generative intelligence module achieves 4.3/5.0 developer usefulness rating. AINTMA demonstrates that agentic AI, combining autonomous multi-agent coordination, generative intelligence and secure smart connectivity, can fundamentally advance software quality management in cloud-scale enterprise environments.
間違った症状のマーク付け: 医学文書における LLM ウォーターマークの評価
大規模言語モデル (LLM) は臨床ワークフローにますます統合されており、透かしを含むモデル生成出力の信頼できるトレーサビリティの必要性が強調されています。しかし、ほとんどのウォーターマークは汎用ベンチマークで評価されており、トークンレベルの小さな摂動が重大な意味の変更を引き起こす可能性がある医療などの領域は十分に調査されていません。この研究では、LLM ウォーターマークが医療パフォーマンスにどのような影響を与えるかについての最初の厳密な研究を紹介し、単峰性および多峰性の臨床推論にわたるさまざまなタスクについて 11 個の LLM と 7 個の VLM にわたる 5 つの透かしスキームのベンチマークを行います。重要なのは、医学的推論の品質、用語の正確さ、誘発された幻覚を体系的に監査するために、人間の専門家によって検証されたパイプラインを導入することで、既存の評価を補完することです。私たちの結果は、透かしが語彙の破損、幻覚の用語、画像所見の拡大された誤った表示や省略など、複数の障害モードにわたって大幅な劣化を引き起こす可能性があることを明らかにしています。特に、ドメイン固有の分析が存在しないことと、臨床テキストに固有の失敗を見逃す集計指標とが組み合わさることで、実際の透かしによる劣化が体系的に見えにくくなる可能性があることがわかりました。私たちの調査結果は、現在のベンチマークが臨床的に重大な失敗を覆い隠してしまう可能性がある医療分野で、透かし入りモデルを安全に導入するための前提条件としてドメイン固有の評価を確立しました。
原文 (English)
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output with watermarking. Yet, most watermarks are evaluated on general-purpose benchmarks, leaving domains like medicine, where small token-level perturbations can result in significant semantic changes, underexplored. In this work, we present the first rigorous study of how LLM watermarks affect medical performance, benchmarking 5 watermarking schemes across 11 LLMs and 7 VLMs on various tasks spanning unimodal and multimodal clinical reasoning. Importantly, we complement existing evaluations by introducing a human-expert-validated pipeline for systematically auditing medical reasoning quality, terminological precision, and induced hallucinations. Our results reveal that watermarking can induce substantial degradation across multiple failure modes, including lexical corruption, hallucinated terminology, and amplified misattribution or omission of image findings. Notably, we find that the absence of domain-specific analyses, combined with aggregate metrics that miss failures inherent to clinical text, can systematically obscure practical watermark-induced degradations. Our findings establish domain-specific evaluation as a prerequisite for the safe deployment of watermarked models in medicine, where current benchmarks can otherwise mask clinically consequential failures.
ClickGuard: 有益性の尺度と大規模な言語モデルを使用してクリックベイト ニュースを検出し台無しにする
このペーパーでは、ユーザーが誤解を招くインターネット記事を回避できるようにクリックベイトを識別する、AI 主導のブラウザ拡張機能について説明します。このアプリケーションは、従来の検出を超えて、トランスフォーマーベースの埋め込みと言語ベースの機能およびカスタムの「ベイトネス」スコアを組み合わせたハイブリッド機械学習アーキテクチャを採用しています。従来のベクタライザーから大規模言語モデル (LLM) 埋め込みに至るまで、さまざまな自然言語処理手法を評価した後、オープン結合データセットで 91% の F1 スコアを達成する XGBoost ベースのモデルが開発されました。最も重要なことは、このツールはユーザーがクリックベイト記事にアクセスする前後に警告できることです。記事を開いた後、ユーザーはクリックベイトである可能性を示すパーセンテージスコアを受け取ります。予測は、提案されたシステム内で特別に開発されたものを含む、分析されたメトリクスに基づいて説明されます。このブラウザ拡張機能は、クリックベイト スポイラー (記事全体の 1 ~ 2 文の要約) も提供します。デモビデオ:https://www.youtube.com/watch?v=IJ1gkQV82C4}{https://www.youtube.com/watch?v=IJ1gkQV82C4
原文 (English)
ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models
This paper presents an AI-driven browser extension that identifies clickbait to help users avoid misleading Internet articles. Moving beyond traditional detection, the application employs a hybrid machine learning architecture that combines transformer-based embeddings with linguistically motivated features and a custom "baitness" score. After evaluating various natural language processing techniques -- from classic vectorizers to large language model (LLM) embeddings -- an XGBoost-based model was developed that achieves an F1-score of 91% on the open combined dataset. Most importantly, the tool can warn users before and after they access a clickbait article. After opening an article, the user receives a percentage score indicating the likelihood that it is clickbait. The prediction is explained based on the analyzed metrics, including those specifically developed within the proposed system. The browser extension also provides a clickbait spoiler -- a one- to two-sentence summary of the entire article. Demo video:https://www.youtube.com/watch?v=IJ1gkQV82C4}{https://www.youtube.com/watch?v=IJ1gkQV82C4
確率的サンプリングは認識論的に浅い: LLM における温度変動とモデル多様性の間の次元ギャップ
言語モデルが実行を繰り返すと異なる答えが得られる場合、その変動によって何が分からないのかが明らかになりますか?自己一貫性により、多数決により変動が質問ごとの不確実性推定値に変換されます。しかし、同じバリエーションがクロス質問の構造、つまり多様なアンサンブルのように関連する質問が交互に現れることを明らかにするのでしょうか?同じ質問について 2 つのレジームを比較します。1 つのモデルは $\tau=1$ で $100$ 回実行されますが、$24$ の LLM のアンサンブルは $\tau=0$ で 1 回ずつ実行されます。マルチェンコ-パストゥールのランダム行列テストでは、信号を両側のサンプリング ノイズから分離します。単一モデル内では、5 つのファミリーと 3 つのベンチマーク (MMLU、HellaSwag、GSM8K) 全体でノイズを上回るのは最大 1 次元です。アンサンブル全体で、4 つの固有値がノイズ エッジをクリアしますが、難易度が一致したベルヌーイ ヌルは、$500$ のモンテカルロ描画で最大 1 つしか生成されません。自己一貫性により、質問ごとに正確な不確実性が得られますが、検出可能なクロス質問構造はありません。多様なアンサンブルだけが、モデルが知らないことを表面化します。
原文 (English)
Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs
When a language model gives different answers on repeated runs, does that variation reveal what it does not know? Self-consistency turns the variation into a per-question uncertainty estimate via majority voting. But does the same variation reveal cross-question structure -- related questions flipping together, the way a diverse ensemble does? We compare two regimes on the same questions: one model run $100$ times at $\tau=1$ versus an ensemble of $24$ LLMs run once each at $\tau=0$. A Marchenko--Pastur random-matrix test separates signal from sampling noise on both sides. Within any single model, at most one dimension rises above noise across five families and three benchmarks (MMLU, HellaSwag, GSM8K). Across the ensemble, four eigenvalues clear the noise edge, while a matched-difficulty Bernoulli null produces at most one in $500$ Monte Carlo draws. Self-consistency gives accurate per-question uncertainty but no detectable cross-question structure; only a diverse ensemble surfaces what a model does not know.
JAXBench: 自律型 TPU カーネル最適化のベンチマーク
厳格なベンチマークにより、ヒルクライムするための共有ターゲットが確立され、自律型 GPU カーネル パフォーマンスの最適化が進歩しましたが、TPU には同等のものが存在しません。 Google Cloud TPU 上で AI によって生成されたカーネル最適化のための TPU ネイティブ ベンチマーク スイートである JAXBench を紹介します。 JAXBench は、関連性があり、最適化のためのヘッドルームを提供する 50 の JAX ワークロードで構成されています。 Llama-3.1、DeepSeek-V3、Mixtral、Mamba-2、AlphaFold2 などのパブリック MaxText ライブラリのアーキテクチャから 17 の実稼働 ML オペレーターを抽出し、正確性が検証され、高い TPU v6e MXU 使用率を達成する新しい問題サイズが設定された 33 のオペレーターを KernelBench から変換します。 17 社のプロダクション オペレーターのうち 8 社は、パブリック Tokamax ライブラリから手動で最適化された Pallas カーネルを出荷し、専門家の上限ベースラインを確立するためにブロック サイズが調整されています。 JAXBench 用の Pallas カーネル候補を生成するための 4 つのフィードバック主導型メソッドを評価します。 Gemini 3 Flash のフル スイート全体にわたって、Pallas のようなまばらに文書化された DSL では、モデルのスケールよりもターゲット固有のコンテキストが重要であることがわかりました。厳選された TPU ドキュメントに基づく条件付けにより、サンプルあたりの正確性が 5.8% から 37.3% に向上し、1.28 倍の幾何平均速度向上で 50 ベンチマーク中 48 を解決します。 Autocomp のビーム検索パイプラインは、XLA の 1.36 倍の幾何平均速度に達し、正確さが達成されると検索構造は大幅な向上をもたらします。 8 つの手動調整されたカーネルでは、Autocomp は XLA の 1.60 倍の幾何平均値に達し、2.08 倍の Tokamax 上限のほとんどを回復しましたが、特殊なページングおよびラグド アテンション演算子には及ばませんでした。高品質の TPU カーネルの最適化は依然として困難な課題であるため、オープンソースの貢献をサポートするために、JAXBench ベンチマーク、評価ハーネス、およびベースライン結果をリリースします。
原文 (English)
JAXBench: Benchmarking Autonomous TPU Kernel Optimization
Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs. We present JAXBench, a TPU-native benchmark suite for AI-generated kernel optimization on Google Cloud TPUs. JAXBench comprises 50 JAX workloads that are both relevant and provide headroom for optimization. We extract 17 production ML operators from architectures in the public MaxText library such as Llama-3.1, DeepSeek-V3, Mixtral, Mamba-2, and AlphaFold2, and translate 33 operators from KernelBench that are validated for correctness and set with new problem sizes that achieve high TPU v6e MXU utilization. Eight of the 17 production operators ship with hand-optimized Pallas kernels from the public Tokamax library and block-size tuned to establish an expert upper-bound baseline. We evaluate four feedback-driven methods on generating candidate Pallas kernels for JAXBench. Across the full suite with Gemini 3 Flash, we find that target-specific context matters more than model scale on a sparsely-documented DSL like Pallas. Conditioning on curated TPU documentation raises per-sample correctness from 5.8% to 37.3% and solves 48 of 50 benchmarks at a 1.28x geomean speedup. Search structure yields significant gains once correctness is achieved, with Autocomp's beam-search pipeline reaching a 1.36x geomean speedup over XLA. On the 8 hand-tuned kernels, Autocomp reaches 1.60x geomean over XLA, recovering most of the 2.08x Tokamax upper bound but trailing on the specialized paged and ragged attention operators. High-quality TPU kernel optimization remains a challenging task, and we release the JAXBench benchmark, evaluation harness, and baseline results to support open source contributions.
DC-Leap: ドラフトガイドによる連続リーピング復号化によるトレーニング不要の dLLM の高速化
並列デコードは拡散大規模言語モデル (dLLM) の効率の中心ですが、現在の戦略は過度に保守的な信頼しきい値によって妨げられることがよくあります。これらのしきい値は、結合確率依存誤差 (JPDE) によって必要となるため、冗長なノイズ除去反復と次善の推論速度をもたらします。これを克服するために、私たちは、中程度の信頼体制で dLLM を確実に加速できるトレーニング不要のフレームワークである DC-Leap を提案します。 DC-Leap は、厳密に順序付けされた因果制約を並列デコード プロセスに統合する動的連続検証戦略を導入します。このメカニズムはトークンの依存関係を段階的に検証することで JPDE を効果的に無力化し、同等のパフォーマンスで信頼性の高い高速化を可能にします。さらに、DC-Leap にはドラフトに基づくデコード メカニズムが組み込まれており、ドラフトは複数のトークンを飛び越えてコンテキストを拡張し、先読みコンテキストを提供し、推論中に双方向の注意の構造的利点を維持するのに役立ちます。標準ベンチマークに関する広範な実験により、DC-Leap は、長いシーケンスの生成で MBPP で最大 53.19 倍、同等の生成品質を持つ KV-Cache と組み合わせた場合に最大 105.02 倍という大幅な高速化を達成することが実証されています。コードは https://github.com/ffh-wyls/DC-Leap で入手できます。
原文 (English)
DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding
While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds. These thresholds, necessitated by the Joint Probability Dependence Error (JPDE), result in redundant denoising iterations and suboptimal inference speeds. To overcome this, we propose DC-Leap, a training-free framework that enables reliable acceleration of dLLMs in the moderate-confidence regime. DC-Leap introduces a Dynamic Contiguous Verification strategy that integrates strictly-ordered causal constraints into the parallel decoding process. By progressively validating token dependencies, this mechanism effectively neutralizes the JPDE, enabling reliable acceleration with comparable performance. Furthermore, DC-Leap incorporates the draft-guided decoding mechanism, where the draft helps extend the context by leaping forward across multiple tokens, providing look-ahead context and retaining the structural benefits of bidirectional attention during inference. Extensive experiments on standard benchmarks demonstrate that DC-Leap achieves substantial speedups, up to 53.19x on MBPP for long-sequence generation, and up to 105.02x when combined with KV-Cache with comparable generation quality. Code is available at https://github.com/ffh-wyls/DC-Leap .
InferenceBench: AI エージェントによるオープンエンド LLM 推論最適化のベンチマーク
AI エージェントは研究開発タスクを自動化するためにますます使用されていますが、既存のベンチマークは通常、所定のワークフローまたは狭いアクション スペースで AI エージェントを評価します。名目上は無制限のタスクであっても、よく知られたレシピを取得し、いくつかのハイパーパラメータを調整することで解決できることが多く、強力な結果が真の最適化を反映しているのか、それとも記憶されたソリューションを反映しているのかが不明確になります。ここでは、エージェントが OpenAI 互換の推論サーバーをデプロイし、LLM 推論の速度を最適化する必要がある InferenceBench を紹介します。各エージェントは、ターゲット LLM、1 つの H100 GPU、最適化シナリオ、および 2 時間の実時間予算を受け取ります。 3 つの最適化シナリオでは、推論の明確なボトルネック (プリフィル レイテンシー、デコード レイテンシー、同時リクエストのスループット) を分離し、4 番目のシナリオでは 3 つすべてのバランスを同時にとります。 15 のフロンティア エージェント構成全体で、エージェントは単純な PyTorch ベースライン (最大 $8.08\times$) を確実に上回っており、多くの場合、デフォルト設定 (vLLM の場合 $4.05\times$) のサービス エンジンと同等またはそれを超えていますが、依然として同じ時間予算 (最大 $11.53\times$) での単純なハイパーパラメータ検索を下回っています。エージェントの軌跡を定性的に分析すると、エージェントは関連する最適化手法を多数列挙していますが、圧倒的に単一の推論フレームワークに収束していることがわかります。彼らは少数の異なる構成のみをテストし、実質的に異なる戦略を検討するのではなく、残りの予算をハイパーパラメータの再測定、修復、または最適化に費やします。これは、ボトルネックはドメインの知識ではなく、多様な構成を提案し、体系的に評価し、特定された最適なソリューションを提出する能力であることを示唆しています。全体として、InferenceBench は、記憶されたソリューションが限定的な改善につながる、オープンエンドの AI エンジニアリング設定で動作するエージェントの能力を反映しています。
原文 (English)
InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces. Even nominally open-ended tasks can often be solved by retrieving a well-known recipe and tuning a few hyperparameters, making it unclear whether strong results reflect genuine optimization or memorized solutions. We introduce InferenceBench, where an agent must deploy an OpenAI-compatible inference server and optimize the speed of LLM inference. Each agent receives a target LLM, one H100 GPU, an optimization scenario, and a wall-clock time budget of two hours. Three optimization scenarios isolate distinct bottlenecks of inference (prefill latency, decode latency, and concurrent request throughput) and a fourth balances all three at the same time. Across 15 frontier agent configurations, agents reliably improve over a naive PyTorch baseline (up to $8.08\times$) and often match or exceed serving engines with default settings ($4.05\times$ for vLLM), but still fall below a simple hyperparameter search under the same time budget (up to $11.53\times$). Qualitative analysis of agent trajectories shows that although agents enumerate many relevant optimization techniques, they overwhelmingly converge on a single inference framework. They test only a few distinct configurations and spend the remaining budget re-measuring, repairing, or optimizing hyperparameters rather than exploring substantially different strategies. This suggests the bottleneck is not domain knowledge, but the ability to propose diverse configurations, evaluate them systematically, and submit the best identified solution. Overall, InferenceBench reflects the ability of agents to operate in an open-ended AI engineering setting, where memorized solutions lead to limited improvements.
DecodeShare: LLM デコード時の決定の共有サブスペースのトレース
大規模言語モデル (LLM) は 1 セットのパラメーターで多くのタスクを処理しますが、KV キャッシュ推論では、プリフィル時ではなくデコード時にどのようなタスク一般構造が使用されるか (存在する場合) が不明です。我々は、デコード時の隠れ状態にあるタスク間で一貫して共有される低次元の部分空間を特定し、デコード中にのみその部分空間を削除することでその因果関係をテストするプロトコルである DecodeShare を提案します。私たちの実験では、発見された共有部分空間を乱すことは、同じ介入予算の下でプレフィル由来またはランダム部分空間を乱すよりもはるかに決定パフォーマンスを低下させます。さらに、このデコード共有サブスペースがアクティベーション ステアリングに対して実際的な結果をもたらすことを示します。つまり、共通のステアリング方向がタスク一般デコード チャネルと重複する可能性があります。この共有サブスペースを投影すると、2 つのコンポーネントの機能的役割が直接分離され、デコード時にステアリング ベクトルを評価することで、プレフィル ベースのプロキシよりも信頼性の高い信号が下流の展開に生成されます。そのコンパクトさにもかかわらず、共有部分空間は、デコード時に高レバレッジの因果チャネルとして機能できます。コードは https://github.com/Zishan-Shao/decodeshare.git で入手できます。
原文 (English)
DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions
Large language models (LLMs) handle many tasks with one set of parameters, but under KV-cached inference it is unclear what task-general structure, if any, is used at decode time rather than during prefill. We propose DecodeShare, a protocol that identifies a low-dimensional subspace consistently shared across tasks in decode-time hidden states, and then tests its causal role by removing that subspace only during decoding. In our experiments, disturbing the discovered shared subspace degrades decision performance far more than disturbing either a prefill-derived or random subspace under the same intervention budget. We further show this decode-shared subspace has practical consequences for activation steering: common steering directions can overlap the task-general decode channel. Projecting out this shared subspace directly separates the functional roles of the two components, while evaluating steering vectors at decode-time yields more reliable signal for downstream deployment than prefill-based proxies. Despite its compactness, the shared subspace can serve as a high-leverage causal channel at decode time. Code is available at: https://github.com/Zishan-Shao/decodeshare.git.
PlanE: 抽出ベースの LLM のデータ、チューニング、推論のメタ プランニング
大規模言語モデル (LLM) のタスク固有の機能を強化するには、主にかなりの命令チューニング データセットが必要です。ただし、そのようなデータの膨大な量により、アノテーションのコストがかなりかかり、特定のタスクに合わせて LLM を調整するための最適化方法が不足しています。上記の問題に対処するために、\textbf{E}xtractive ベースの LLM を構築するための \textbf{PlanE} と呼ばれる \textbf{Planing フレームワークを提案します。これには、データ分解、命令調整、およびプロンプト推論が含まれます。さらに、特定のデータセットに最適なベース LLM とその DTI の組み合わせを選択して構築効率を向上させることを目的とした、データ チューニング推論 (DTI) プランナーを導入します。実験結果は、(1) 同じベース LLM を使用する異なるデータセット間、および (2) 異なるベース LLM を使用する同じデータセット上の 2 つの観点から、PlanE の有効性を示しています。さらに、さまざまな最適化目標の下で、提案された DTI プランナーの一般化可能性を検証します。コードは https://github.com/gugugu-469/PlanE で公開されています。
原文 (English)
PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs
Enhancing the task-specific capabilities of Large Language Models (LLMs) primarily requires substantial instruction-tuning datasets. However, the sheer volume of such data imposes a considerable annotation cost, and a lack of optimization methods for tailoring LLMs to specific tasks. To address the above issues, we propose a \textbf{Plan}ning framework for constructing \textbf{E}xtractive-based LLMs called \textbf{PlanE}, which includes data decomposition, instruction tuning, and prompt inference. Additionally, we introduce a Data-Tuning-Inference (DTI) planner, aimed at selecting the optimal base-LLM and its DTI combinations for specific datasets to improve construction efficiency. The experimental results demonstrate the effectiveness of our PlanE from two views: (1) across different datasets using the same base-LLM, and (2) on the same dataset using different base-LLMs. Furthermore, we validate the generalizability of the proposed DTI planner under different optimization objectives. The codes are publicly available at https://github.com/gugugu-469/PlanE.
大規模言語モデルのパーソナライゼーション機能のベンチマーク
パーソナライゼーションとは、送信者、チャネル、時間を固定したままメッセージを変化させて特定の受信者のアクションを誘発する行為であり、送信者と受信者が独立した目的を持つ二者間問題として心理学とマーケティングにおいて長い伝統があります。大規模な言語モデルは、推定された受信者の状態に条件付けされたメッセージのバリアントの連続体を生成することによって、古典的な検索とランキングのアプローチの有界インベントリの制約を取り除き、現在のモデルが古典的な意味でのパーソナライゼーションをどの程度うまく実行するかという疑問を引き起こします。既存の LLM パーソナライゼーション ベンチマークは、受信者がモデルがサービスを提供しているのと同じユーザーである送信者側の適応を測定します。生成されたメッセージが第三者に意図したアクションを引き起こすかどうかという二者間の問題は、A/B テストと小規模な人体研究を通じてのみ調査されており、オンデマンドで新しいモデルに対して再実行することはできません。我々は、Kamenica and Gentzkow (2011) のベイジアン説得フレームワークを生成エージェントに適応させ、営業における定式化をインスタンス化します。そこでは、受信者のアクションが、それを誘発したアウトリーチに対して定期的に記録されます。当社は、22 の業界と約 200 の企業にまたがる 6,279 件の顧客成功事例を収録したパブリック コーパスである SDR-Bench をリリースします。これは、将来のデータ漏洩を防ぐ、時間的に制限されたシミュレーションを通じて提供されます。フロンティア LLM とディープリサーチ エージェント全体で、一貫したパーソナライゼーションの停滞期が観察されており、Fortune 100 のテクノロジー コホートでは、アウトリーチの成功と失敗を統計的に区別するモデルはありません。 12 人のプロの営業担当者によるフィールド展開によりフレームワークが検証され、モデルで生成されたコンテンツの 48% がすぐに役立つと評価され、上級専門家の合意はピアソン 0.82 でした。大規模な生成パーソナライゼーションの再現可能な研究をサポートするために、SDR-Arena と SDR-Bench を一般公開します。
原文 (English)
Benchmarking the Personalization Capabilities of Large Language Models
Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradition in psychology and marketing as a two-party problem in which sender and receiver have independent objectives. Large language models remove the bounded-inventory constraint of classical retrieval-and-ranking approaches by generating a continuum of message variants conditioned on inferred receiver state, raising the question of how well current models perform personalization in the classical sense. Existing LLM personalization benchmarks measure sender-side adaptation, in which the receiver is the same user the model is serving. The two-party question, whether a generated message induces its intended action in a third party, has been investigated only through A/B tests and small-scale human studies that cannot be re-run against a new model on demand. We adapt the Bayesian Persuasion framework of Kamenica and Gentzkow (2011) to generative agents and instantiate the formulation in sales, where receiver actions are routinely logged against the outreach that induced them. We release SDR-Bench, a public corpus of 6,279 customer success stories spanning 22 industries and approximately 200 enterprises, served through a temporally constrained simulation that prevents future-data leakage. Across frontier LLMs and deep-research agents, we observe a consistent personalization plateau and on a Fortune 100 tech cohort no model statistically separates successful from unsuccessful outreach. A field deployment with 12 professional sales representatives validates the framework, with 48 percent of model-generated content rated immediately useful and senior-expert agreement at Pearson 0.82. We release SDR-Arena and SDR-Bench publicly to support reproducible study of generative personalization at scale.
堅牢な批評家: マルチターン攻撃から LLM を守る
ユーザーが言語モデルに何か有害なことを尋ねたとき、それは本物の攻撃なのでしょうか、それとも誤解されているが善意の質問なのでしょうか?この曖昧さは、LLM の安全性の中心的な課題の 1 つです。最悪の事態が正規ユーザーに損害を与えることを想定したモデル。最善のものを前提とするものは簡単に悪用されてしまいます。この問題は、複数のターンにまたがる対話でさらに複雑になります。攻撃者の真の意図は、多くの対話にわたって徐々にしか明らかにされない可能性があるにもかかわらず、既存の安全フレームワークは状況に応じた盗賊扱いを適用し、会話の軌跡を無視します。そのために、対話のあらゆる段階でユーザーの意図を推測することでこれに対処するフレームワークである Dialogue Critic Guided Sampling (DCGS) を提案します。 DCGS は、何が安全か、何が安全でないかについて固定ルールを適用するのではなく、完全な会話履歴に基づいてユーザーの意図が何である可能性が高いかを学習し、それに応じて応答を生成します。形式的には、敵対的対話をマルコフ意思決定プロセスとしてモデル化し、個々のトークンと発話 (完全な応答) レベルの両方で価値と後悔に基づく批評を学習し、行動価値の批評を通じて候補の応答をスコアリングします。この推論時の再重み付けが基本ポリシーの指数関数的傾斜に近似し、有限の候補プールの期待収益の改善を保証することを証明します。これは、グループ相対の目的では示されない特性です。 CARES-18k、WildJailbreak、Redbench、Harmbench で評価したところ、DCGS は敵対的対話タスクにおいて強力で堅牢なベースラインやフロンティア モデルを上回りました。 DCGS はフロンティア モデルにも移行し、微調整することなく堅牢性を向上させます。
原文 (English)
Robust Critics: Defending LLMs Against Multi-Turn Attacks
When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the central challenges of LLM safety. A model that assumes the worst harms legitimate users; one that assumes the best is easily exploited. The problem is compounded in multi-turn dialogue, where an attacker's true intent may only reveal itself gradually across many exchanges, yet existing safety frameworks apply a contextual bandit treatment, ignoring the trajectory of the conversation. To that end, we propose Dialogue Critic Guided Sampling (DCGS), a framework that addresses this by inferring user intent at every turn of dialogue. Instead of applying a fixed rule about what is or is not safe, DCGS learns what the user's intent is likely to be based on the full conversational history and generates responses accordingly. Formally, we model adversarial dialogue as a Markov Decision Process and learn value and regret-based critics at both the individual token and utterance (full response) levels, scoring candidate responses via an action-value critic. We prove that this inference-time reweighting approximates exponential tilting of the base policy, guaranteeing improvement in expected return for any finite candidate pool, a property that group-relative objectives do not exhibit. Evaluated on CARES-18k, WildJailbreak, Redbench, and Harmbench, DCGS outperforms strong robust baselines and frontier models on adversarial dialogue tasks. DCGS also transfers to frontier models, improving their robustness without fine-tuning.
大規模言語モデルにおける不完全なプロンプト ジェイルブレイク
大規模言語モデル (LLM) は、有害なリクエストに対する保護機能を備えたオープンウェイト モデルとしてリリースされることが増えています。それにもかかわらず、文章の完成は依然として不完全な有害なプロンプトに対して脆弱です。この研究では、この現象を不完全なプロンプトのジェイルブレイク (IPJ) として形式化し、不完全なプロンプトがいつ、どのように有害な継続を引き起こすかについて体系的な経験的特徴付けを提供します。我々は、不完全な文の継続に関連するさまざまなアトラクターのタイプを分析し、LLMが文の終了まで体系的に拒否を遅らせることを示します。さらに、パラメータ調整による不完全な有害なプロンプトを拒否するモデルのトレーニングでは不十分であり、コンテンツ ドメインとアトラクター タイプの両方にわたって一般化できないことを示します。きめ細かい制御を可能にするために、終了ニューロンと継続ニューロンという 2 つの機能ニューロンを特定します。文の完成におけるそれらの役割を明確にすることで、より正確で堅牢なIPJ防御のためのニューロンレベルの介入の可能性を強調します。
原文 (English)
Incomplete Prompt Jailbreaks in Large Language Models
Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion remains vulnerable to incomplete harmful prompts. In this work, we formalize this phenomenon as incomplete prompt jailbreaks (IPJ) and provide a systematic empirical characterization of when and how incomplete prompts elicit harmful continuations. We analyze diverse attractor types associated with incomplete sentence continuation and show that LLMs systematically delay refusal until sentence termination. We further demonstrate that training models to refuse incomplete harmful prompts via parameter tuning is insufficient, failing to generalize across both content domains and attractor types. To enable fine-grained control, we identify two functional neurons: termination and continuation neurons. By clarifying their roles in sentence completion, we highlight the potential of neuron-level interventions for more precise and robust IPJ defenses.
VeriSimpl: 単純化ベースの検証を使用した自然言語からの堅牢な最適化モデリング
自然言語インターフェイスは、最適化モデリングのアクセシビリティと使いやすさに大きなメリットをもたらします。また、大規模言語モデル (LLM) の最近の進歩により、問題のテキスト記述を実行可能なソルバー定式化に自動的に変換することが期待されています。ただし、既存のアプローチの重要な課題は、たとえエラーなしで実行されたとしても、推論された定式化が意図したタスクを正しく実装していることを確認することです。堅牢な自然言語から最適化への形式化のためのソルバー LLM フレームワークである VeriSimpl を紹介します。私たちのアプローチは単純化ベースの検証の考えに基づいており、最適化ソルバーを利用して候補定式化に関する単純化された診断クエリを生成し、LLM がタスクの記述に関して定式化の正しさをわかりやすく推論できるようにします。我々は、問題の制約と決定変数に関してさまざまな次元に沿ったこのような単純化戦略を提示します。これにより、LLM は固定されたグローバル コンテキストの下でローカルに推論できるようになります。さまざまな最適化ベンチマークの評価により、私たちのアプローチが既存の方法と比べて一貫して精度を向上させながら、新しい高精度の自己検証信号も提供できることがわかります。
原文 (English)
VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification
Natural language interfaces can greatly benefit the accessibility and usability of optimization modeling, and recent advances in large language models (LLMs) show promise in automatically translating textual problem descriptions into executable solver formulations. However, a key challenge for existing approaches is to ensure that the inferred formulation correctly implements the intended task, even if it may execute without errors. We introduce VeriSimpl, a solver LLM framework for robust natural-language-to-optimization formalization. Our approach is based on the idea of simplification-based verification, where the optimization solver is leveraged to generate simplified diagnostic queries about a candidate formulation to allow the LLM to tractably reason about the correctness of the formulation with respect to the task description. We present such simplification strategies along different dimensions with respect to problem constraints and decision variables, which allow the LLM to reason locally under fixed global contexts. Evaluations on a range of optimization benchmarks show how our approach provides consistent improvements in accuracy over existing methods, while also providing a novel high-precision self-verification signal.
SonicSampler: LLM サンプリングと投機的検証のための統合されたタイル対応カーネル
LLM 推論におけるサンプリングは、ロジット処理、トークン選択、および投機的デコードのための検証操作の組み合わせセットで構成されます。ただし、既存の実装は、このパイプラインのサブセットのみを高速化するか、複数のカーネル起動に依存するか、バッチ全体で均一なサンプリング動作を想定しているため、動的なサービング ワークロードのサポートが制限され、効率的な CUDA グラフの実行が妨げられます。 $\textbf{SonicSampler}$ は、完全なサンプリング パイプラインを固定のワークロード認識実行モデルに垂直に融合するタイル認識 Triton カーネルの統合スイートです。当社のカーネルは、完全な CUDA Graph 互換性を維持しながら、文法制約のあるデコード、繰り返し、頻度と存在のペナルティ、ロジット バイアス、温度スケーリング、top-$k$ / top-$p$ / min-$p$ フィルタリング、投機的検証などの動的なリクエストごとのサンプリング動作を単一のバッチ カーネル内でサポートします。私たちのアプローチの中心となるのは、競合ベースラインと比較して最大 $\textbf{10 倍の高速化}$ を達成し、LLM 出力の低エントロピー構造を利用して大規模な語彙に対する効率的な選択を可能にする、新しい階層型 2 段階の上位 $k$ アルゴリズムです。 SonicSampler は、異種混合の投機的デコード ワークロード全体で、柔軟なバッチ実行を維持しながら、最先端のベースラインを上回る最大 $\textbf{16 倍の高速化}$ を達成します。
原文 (English)
SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification
Sampling in LLM inference comprises a combinatorial set of logit processing, token selection, and verification operations for speculative decoding. However, existing implementations either accelerate only subsets of this pipeline, rely on multiple kernel launches, or assume homogeneous sampling behavior across a batch, limiting support for dynamic serving workloads and preventing efficient CUDA Graph execution. We present $\textbf{SonicSampler}$, a unified suite of tile-aware Triton kernels that vertically fuses the complete sampling pipeline into a fixed, workload-aware execution model. Our kernels support dynamic per-request sampling behaviors, including grammar-constrained decoding, repetition, frequency and presence penalties, logit bias, temperature scaling, top-$k$ / top-$p$ / min-$p$ filtering, and speculative verification - within a single batched kernel while remaining fully CUDA Graph-compatible. Central to our approach is a novel hierarchical two-stage top-$k$ algorithm that achieves up to $\textbf{10x speedup}$ over competitive baselines and exploits the low-entropy structure of LLM outputs to enable efficient selection over large vocabularies. Across heterogeneous speculative decoding workloads, SonicSampler achieves up to $\textbf{16x speedup}$ over state-of-the-art baselines while preserving flexible batched execution.
マルチセンサーの物理的危険性評価に関する大規模言語モデルのベンチマーク
5 つの大規模な言語モデルがマルチセンサーの物理的危険データをどのように評価するかを評価する経験的なベンチマークを示します。温度 0.0 で 1,800 回の API 呼び出しを使用して、マルチセンサーの共同評価、応答の比例性、パターンの明確化の 3 つのカテゴリにわたる 60 のシナリオをテストしたところ、単一センサーのしきい値違反についてはほぼ完璧な精度を達成しながら、複数のセンサーが同時に個別の安全限界以下に上昇するテスト済みのシナリオ全体で、すべてのテスト済みモデルが一貫して予防警告信号を生成しなかったことがわかりました。 5 つのモデル (ChatGPT-4o、Gemini 2.5 Flash、DeepSeek、Kimi、Llama 3.1 8B) はすべて、単一センサー シナリオ (カテゴリー B Q1: 0.975-1.000)。構造化された表形式には、単純な散文に比べて一貫した利点はありません。 ChatGPT-4o は、散文の下で大幅に優れたパフォーマンスを示します (p = 0.001)。これらの発見は、テストされたモデルを物理的安全監視システムに導入する実務者に直接影響します。
原文 (English)
Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment
We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three categories - multi-sensor joint assessment, response proportionality, and pattern disambiguation - with 1,800 API calls at temperature 0.0, we find that all tested models consistently produced no precautionary warning signal across the tested scenarios where multiple sensors are simultaneously elevated below their individual safety limits, while achieving near-perfect accuracy on single-sensor threshold violations. All five models (ChatGPT-4o, Gemini 2.5 Flash, DeepSeek, Kimi, Llama 3.1 8B) score near zero on Category A multi-sensor scenarios (Q2: 0.000-0.208; Q3: 0.000-0.592) compared to strong performance on single-sensor scenarios (Category B Q1: 0.975-1.000). Structured tabular formatting shows no consistent advantage over plain prose; ChatGPT-4o performs significantly better under prose (p = 0.001). These findings have direct implications for practitioners deploying the tested models in physical safety monitoring systems.
半教師ありテキスト属性グラフ蒸留
{\em Text-Attributed Graphs} (TAG) は、グラフ トポロジを豊富なテキスト セマンティクスと統合するための表現力豊かなデータ モデルとして登場しました。 TAG を介した既存の表現学習方法は、特に {\em Large Language Model} (LLM) を使用した場合に、深刻なスケーラビリティのボトルネックに悩まされます。データ蒸留は有望なデータ中心のソリューションを提供しますが、既存の方法では、グラフとテキストのモダリティ間の複雑な相互作用を捉えることができず、半教師あり設定に固有のラベルの不足に悩まされ、下流の LLM ベースのタスクに必要な人間が判読できるテキスト属性を生成する機能が不足しています。これらの課題に対処するために、私たちは、{\em Wasserstein Distance} (WSD) に基づいた統合された半教師ありフレームワークである \algo{} を提案します。 \algo{} は、実際の TAG に関する経験的発見に基づいて、信頼性の高い疑似ラベルを収集し、相補的なグラフ テキスト特徴を融合するための協調的自己学習スキーム内でデュアル パスウェイ エンコーダ (グラフ対応およびグラフ フリー) を利用するグラフ テキスト協調エンコード モジュールを導入しています。さらに、理論に基づいた WSD ベースのグラフ スケッチ アルゴリズムとコスト効率の高い LLM テキスト合成モジュールを開発します。これは、クラスター ベースのキーワード抽出を利用して、凝縮されたノードに対して一貫した人間が読める要約を生成します。ベンチマーク データセットに対する広範な実験により、\algo{} が GNN ベースと LLM ベースのダウンストリーム タスクの両方に関して最先端のパフォーマンスと圧縮のトレードオフを達成し、効果的かつ効率的な TAG 学習または分析が可能になることが実証されました。
原文 (English)
Semi-Supervised Text-Attributed Graph Distillation
{\em Text-Attributed Graphs} (TAGs) have emerged as an expressive data model for integrating graph topology with rich textual semantics. Existing representation learning methods over TAGs suffer from severe scalability bottlenecks, particularly together with {\em Large Language Models} (LLMs). While data distillation offers a promising data-centric solution, existing methods fail to capture the complex interplay between graph and text modalities, struggle with the label scarcity inherent in semi-supervised settings, and lack the ability to produce the human-readable textual attributes required for downstream LLM-based tasks. To address these challenges, we propose \algo{}, a unified semi-supervised framework guided by the {\em Wasserstein Distance} (WSD). Grounded in our empirical findings on real TAGs, \algo{} introduces a graph-text collaborative encoding module that utilizes dual-pathway encoders (graph-aware and -free) within a collaborative self-training scheme to harvest reliable pseudo-labels and fuse complementary graph-text features. Furthermore, we develop a theoretically grounded WSD-based graph sketching algorithm and a cost-effective LLM text synthesis module, which leverages cluster-based keyword extraction to generate coherent, human-readable summaries for condensed nodes. Extensive experiments on benchmark datasets demonstrate that \algo{} achieves a state-of-the-art performance-compression trade-off in terms of both GNN- and LLM-based downstream tasks, enabling effective and efficient TAG learning or analytics.
嘘つきのベンチを超えて: LLM における嘘の検出に対する嘘の類型学、深さ、およびスパース性の影響
大規模な言語モデルからの不正な出力を検出するためのプローブのトレーニングは、依然として未解決の問題です。最近の研究では、特にドメイン外のシナリオでは検出プローブが失敗することが実証されています。ある種類の嘘に関するトレーニングは、他の種類の嘘が関与する欺瞞シナリオにはうまく移行できません。この研究では、表現の深さ、プローブの表現力、まばらな特徴の表現、トレーニング データの嘘の類型論など、さまざまな要因が検出パフォーマンスにどのように影響するかについて系統的な研究を実施します。この目的を達成するために、捏造、省略、誇張の例など、さまざまなタイプの欺瞞を含む補足データセットを使用して、標準的なベンチマーク トレーニング データを強化します。 7 つのプローブ タイプにわたってこれらの要因を分析した実験結果は、最適な表現深度はデータセットに大きく依存し、より表現力の高いプローブは線形ベースラインに対して選択的なゲインのみを提供し、疎なオートエンコーダー機能は密な隠れ状態と同様に機能することを示しています。最終的に、トレーニング データと嘘の類型学の選択によって検出可能性が大幅に変化することを実証し、欺瞞の検出が表現に大きく依存する問題であることを強調しました。
原文 (English)
Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain scenarios -- training on one type of lie does not transfer well to deception scenarios involving other types of lies. In this work, we conduct a systematic study on how various factors impact detection performance: representation depth, probe expressivity, sparse feature representations, and the lie typology of the training data. To this end, we augment standard benchmark training data with a supplementary dataset containing diverse types of deception, including fabrication, omission, and exaggeration examples. Analyzing these factors across seven probe types, our experimental results show that the optimal representation depth is highly dataset-dependent, more expressive probes provide only selective gains over linear baselines, and sparse autoencoder features perform similarly to dense hidden states. Ultimately, we demonstrate that the choice of training data and lie typology substantially changes detectability, highlighting that deception detection is a highly representation-dependent problem.
制約付きマルチソース推論による分散システムでのスケーラブルなトポロジ推論の有効化
正確な配電システム トポロジは、停電の位置特定、電圧分析、配電網の運用に不可欠ですが、実際には、不均一で不完全な電力会社データのため、信頼性の高い接続記録を維持することは依然として困難です。既存のトポロジ識別方法は、主に電気的類似性または空間記録のみに依存することが多く、密集したフィーダや一貫性のないメタデータ条件下では信頼性が低くなります。この論文では、空間的な実現可能性と物理的な運用上の制約を強制しながら、異種の証拠を使用してユーティリティが提供する基本トポロジを改良する制約付き推論問題として、分散トポロジの特定を定式化します。提案されたフレームワークは、接続を最初から再構築するのではなく、一貫性のない割り当てを検出し、制約された近傍内で局所的な再接続を実行してスケーラビリティを確保し、物理的な実現可能性を繰り返し適用して運用上一貫したトポロジ推定を生成します。さらに、改ざん主導の信頼性メトリクスは、推定された各接続が代替の実行可能な割り当てと比較してどの程度強力にサポートされているかを評価するため、電力会社はシステム全体の可観測性を維持しながら検証作業に優先順位を付けることができます。このフレームワークは、米国の大手電力会社と協力して、8,000 ドルを超える AMI メーターを構成する 3 つのフィーダからの運用データを使用して検証されています。結果は、グローバル推論アプローチと比較して計算量を大幅に削減しながら、$95\%$ 以上のトポロジ再構成精度を示しています。この研究ではさらに、相関ベースの手法だけでは密集した都市フィーダではあいまいな割り当てが生成されるのに対し、電気測定と空間的および運用上の制約を組み合わせることで、現実的な展開条件下で堅牢かつスケーラブルなトポロジ回復が可能になることが示されています。
原文 (English)
Enabling Scalable Topology Inference in Distribution Systems via Constrained Multi-Source Inference
Accurate distribution system topology is essential for outage localization, voltage analytics, and operation of distribution grids, yet maintaining reliable connectivity records remains challenging in practice due to heterogeneous and imperfect utility data. Existing topology identification methods often rely primarily on electrical similarity or spatial records alone, which become unreliable in dense feeders and under inconsistent metadata conditions. This paper formulates distribution topology identification as a constrained inference problem that refines a utility-provided base topology using heterogeneous evidence while enforcing spatial feasibility and physical operational constraints. Instead of reconstructing connectivity from scratch, the proposed framework detects inconsistent assignments, performs localized reconnection within constrained neighborhoods to ensure scalability, and iteratively enforces physical feasibility to produce operationally consistent topology estimates. In addition, a falsification-driven reliability metric evaluates how strongly each inferred connection is supported relative to alternative feasible assignments, enabling utilities to prioritize verification efforts while preserving system-wide observability. The framework is validated using operational data from three feeders comprising more than $8{,}000$ AMI meters in collaboration with a large U.S. utility. Results demonstrate over $95\%$ topology reconstruction accuracy while significantly reducing computational effort compared with global inference approaches. The study further shows that correlation-based methods alone produce ambiguous assignments in dense urban feeders, whereas combining electrical measurements with spatial and operational constraints enables robust and scalable topology recovery under realistic deployment conditions.
トレーニングなしのルーティング: 信頼性ゲーティングによる制御可能な比率の LLM オフロード
ローカルとクラウドのコラボレーションは、リソースの制約の下で大規模な言語モデルを展開するための実用的な方法ですが、既存の方法は多くの場合、トレーニングされたルーターや、ルーティング動作を特定の運用体制に結び付けるコラボレーション対応の微調整に依存しています。この研究では、そのようなトレーニングが不要である可能性があることを示します。サンプリングされた応答にわたるローカル モデル独自の推論時間の一致は、いつローカル実行を信頼するか、いつより強力なクラウド モデルにオフロードするかを決定するための強力なシグナルをすでに提供しています。我々は、トレーニング不要のルーティング フレームワークである CARGO を提案します。これは、プロンプト変動サンプリングを通じてこの一致を推定し、サンプル効率的な不確実性制御のためにベイジアン早期停止を適用し、軽量の展開時キャリブレーションを通じて任意のターゲット コラボレーション比率をサポートします。多様な推論タスクと質問応答タスク、複数のローカル LLM ファミリとスケール、および事前トレーニング済みおよび微調整されたローカル モデルの両方にわたって、CARGO は他のトレーニング不要のベースラインを常に上回り、いくつかの設定では教師あり学習ルーターを上回っています。これらの結果は、効果的で適応性のあるローカルとクラウドのコラボレーションが、追加のトレーニング済みルーターを必要とせずに、ローカル モデルの固有の応答動作から直接実現できることを示唆しています。
原文 (English)
Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating
Local-cloud collaboration is a practical way to deploy large language models under resource constraints, but existing methods often rely on trained routers or collaboration-aware finetuning that tie routing behavior to a particular operating regime. In this work, we show that such training may be unnecessary: the local model's own inference-time agreement across sampled responses already provides a strong signal for deciding when to trust local execution and when to offload to a stronger cloud model. We propose CARGO, a training-free routing framework that estimates this agreement through prompt-varied sampling, applies Bayesian early stopping for sample-efficient uncertainty control, and supports arbitrary target collaboration ratios through lightweight deployment-time calibration. Across diverse reasoning and question-answering tasks, multiple local LLM families and scales, and both pretrained and finetuned local models, CARGO consistently outperforms other training-free baselines and in several settings surpasses supervised learned routers. These results suggest that effective and adaptable local-cloud collaboration can emerge directly from the local model's intrinsic response behavior, without requiring an additional trained router.
PersonalTrail: 証跡の閲覧によるパーソナライズされた Web エージェントのベンチマーク
大規模言語モデルの最近の進歩により、Web エージェントが複雑なタスクを自律的に実行できるようになりました。実際には、ユーザーは頻繁に指定不足の指示を提供し、エージェントが生の閲覧履歴から欠落しているコンテキストを推測することを要求します。既存のベンチマークは、タスクを完全に明示的なプロンプトに制限するか、Web インタラクション履歴を単純化された形式に抽象化するため、この形式のパーソナライゼーションを捉えることができません。このギャップを埋めるために、管理されたオープン Web 環境で動作するパーソナライズされた Web エージェントのベンチマークである PersonaTrail を導入します。 PersonaTrail は、ユーザー履歴として現実的な閲覧軌跡を活用することで、ユーザーの好みを推測し、過去の閲覧セッションからの情報を呼び出すエージェントの能力を評価します。さらに、生の閲覧履歴を 2 種類の構造化記憶 (個々のセッションを要約する事実記憶と、繰り返しの行動パターンを抽出する嗜好記憶) に分解するフレームワークである Preference-Aware Contextual Memory (PACMem) を提案します。推論時に、エージェントはこれらのメモリから最も関連性の高いエントリを取得し、パーソナライズされたナビゲーションをガイドします。広範な実験により、PACMem は両方のタスクにおいて既存のメモリベースのベースラインよりも一貫して優れたパフォーマンスを発揮することが示されています。
原文 (English)
PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails
Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histories. Existing benchmarks fail to capture this form of personalization, as they either restrict tasks to fully explicit prompts or abstract web interaction history into simplified forms. To bridge this gap, we introduce PersonaTrail, a benchmark for personalized web agents operating in a managed open web environment. By leveraging realistic browsing trajectories as user history, PersonaTrail evaluates an agent's ability to infer user preferences and recall information from past browsing sessions. We further propose Preference-Aware Contextual Memory (PACMem), a framework that decomposes raw browsing histories into two types of structured memory: factual memories that summarize individual sessions and preference memories that distill recurring behavioral patterns. At inference time, the agent retrieves the most relevant entries from these memories to guide personalized navigation. Extensive experiments show that PACMem consistently outperforms existing memory-based baselines on both tasks.
自己回帰言語モデルの扱いやすい階層制御
自己回帰大規模言語モデル (LLM) の生成を制限することは、言語モデルを正式なシステムに統合するための重要なコンポーネントです。プログラム合成などのタスクのコードとデータの生成では、言語モデルが構文的に有効な出力を生成することを保証することが、そのような出力を処理するための前提条件です。これらの言語 (SQL や JSON など) は、多くの場合、$LR(k)$ 文脈自由文法として設計されています。 LLM を扱いやすい確率モデルに蒸留することにより、その自己回帰生成を制御およびマスクして、論理制約を満たす確率を組み込むことができ、有効であることが保証された高品質の出力を保証します。この論文は、有限期間の任意の $LR(k)$ 文法の満足度が多項式時間で計算できることを実証します。これは、そのような文法に以前の方法を適用するときの指数関数的な時間よりも改善されています。この結果により、形式的な構文上の制約をより適切に満たす出力に向けた LLM 生成の効率的な制約とステアリングが可能になります。
原文 (English)
Tractable Hierarchical Control of Autoregressive Language Models
Constraining the generation of autoregressive large language models (LLMs) is an important component of integrating language models into formal systems. In the generation of code and data for tasks like program synthesis, ensuring that language models produce syntactically valid output is a prerequisite for processing such output. These languages (such as SQL or JSON) are often designed as $LR(k)$ context-free grammars. By distilling the LLM to a tractable probabilistic model, its autoregressive generation can be steered and masked to incorporate the probability of satisfying logical constraints, ensuring high quality output that is guaranteed to be valid. This paper demonstrates that the satisfaction of any $LR(k)$ grammar of finite duration can be calculated in polynomial time, an improvement over the exponential time of applying previous methods to such grammars. This result enables efficient constraint and steering of LLM generation towards output that better satisfies formal syntactic constraints.
悪魔はスペクトルの中にいる: トポロジー的に正則化されたサイドパスを介して LLM の表現崩壊を軽減する
大規模言語モデル (LLM) は、長いコンテキストのパフォーマンスを大幅に低下させるボトルネックである表現の崩壊によって根本的に制限されています。我々は、既存のアプローチが、均質化崩壊(例:ランクの欠如を引き起こす注意の低下)と孤立化崩壊(例:文脈の切断を引き起こす局所的な注意)という 2 つの病理学的極端な状態のいずれかに陥るリスクがあることを確認しています。注意のダイナミクスのスペクトル分析を通じて、標準的なメカニズムではバランスを取るのが難しい、混合効率(スペクトルギャップ)と情報容量(有効ランク)の間の本質的なトレードオフを導き出します。このジレンマを解決するために、スペクトル バランスを達成する非侵襲的なアーキテクチャ介入である Topologically Regularized Side-Path (TRSP) を提案します。 TRSP は、トークン相互作用トポロジーを正規化するために、軽量で長さを認識するゲートによって拡張されるパラメーターフリーの三角ボックス メカニズムを採用しています。 TRSP は、効果的なランクを維持する近位結合と非縮退混合をサポートする遠位伝播を統合することにより、コアの注意を変えることなく、幾何学的に健全な遷移演算子を促進します。実験では、一般的な機能とロングコンテキストのベンチマーク全体で大幅な改善が見られました。特に、トレーニング長の $8\time$ の NoLiMa では、TRSP は $83\%$ の精度を維持し、差動トランスとゲート アテンションをそれぞれ約 30 パーセント ポイントと 50 パーセント ポイント上回っています。コードは https://github.com/Eziotao-tyd/TRSP で入手できます。
原文 (English)
The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
Large Language Models (LLMs) are fundamentally limited by representation collapse, a bottleneck that severely degrades long-context performance. We identify that existing approaches risk drifting into one of two pathological extremes: homogenization collapse (e.g., attention sinks causing rank deficiency) and isolation collapse (e.g., local attention causing context disconnection). Through spectral analysis of attention dynamics, we derive an intrinsic trade-off between mixing efficiency (spectral gap) and information capacity (effective rank) that standard mechanisms struggle to balance. To resolve this dilemma, we propose the Topologically Regularized Side-Path (TRSP), a non-invasive architectural intervention that achieves spectral balance. TRSP employs a parameter-free Triangular Box mechanism, scaled by a lightweight, length-aware gate, to regularize the token interaction topology. By integrating proximal coupling to preserve effective rank and distal propagation to support non-degenerate mixing, TRSP promotes a geometrically healthier transition operator without altering core attention. Experiments show significant improvements across general capabilities and long-context benchmarks. Notably, on NoLiMa at $8\times$ the training length, TRSP retains $83\%$ accuracy and surpasses the Differential Transformer and Gated Attention by approximately 30 and 50 percentage points, respectively. Code available at: https://github.com/Eziotao-tyd/TRSP.
現実世界のユーザーの期待に対する言語モデルの期待の調整
大規模言語モデル (LLM) は標準ベンチマークで顕著なパフォーマンスを実証していますが、それが本当にユーザーの期待に応えるかどうかはほとんど解明されていません。モデルヒューリスティック、エキスパートルーブリック、またはユーザーシミュレーションに依存する既存の評価アプローチは、実際の人間の期待の多様性と微妙さを捉えることができず、モデルが有能であるように見える一方で、ユーザーが実際に求めているものとずれてしまいます。私たちは、現実世界の LLM インタラクションにおけるユーザーの期待に関する最初の体系的な研究を発表し、意味的に豊かな期待を抽出する原則に基づいた手順を提案し、実際のユーザーの期待に基づいたベンチマークである ExpectBench を紹介します。分析の結果、現在の LLM はユーザーが得たいものを満足させ、予測するのに苦労していることが明らかになり、不整合の根本的な原因が浮き彫りになっています。これらの観察に基づいて、軽量の潜在的な期待を認識した応答生成フレームワークである LENS を提案します。 LENS を使用すると、モデルがユーザーの期待を内面化し、より適切に調整された応答を生成できるようになり、期待の満足度が一貫して向上し、現実的な人間と AI の調整のためにユーザーの期待を明示的にモデル化することの重要性が強調されます。
原文 (English)
Expectation Alignment of Language Models for Real-World User Expectations
Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing evaluation approaches, relying on model heuristics, expert rubrics, or user simulation, fail to capture the diversity and subtlety of real human expectations, causing models to appear competent while misaligning with what users actually seek. We present the first systematic study of user expectations in real-world LLM interactions, proposing a principled procedure to extract semantically rich expectations and introducing ExpectBench, a benchmark grounded in real user expectations. Analyses reveal that current LLMs struggle to satisfy and anticipate what users hope to obtain, highlighting a fundamental source of misalignment. Building on these observations, we propose LENS, a lightweight latent expectation-aware response generation framework. LENS enables models to internalize user expectations and generate better-aligned responses, consistently improving expectation satisfaction and underscoring the importance of explicitly modeling user expectations for realistic human-AI alignment.
OPTScientist: トランスフォーマーの事前トレーニングのための型付きオプティマイザー プログラムのマルチエージェント検出
最新の深層学習向けのオプティマイザーの設計は、依然として科学的に難しい問題であり、最適化幾何学、状態ダイナミクス、数値安定性、実装上の制約、および経験的一般化を共同で考慮する必要があります。既存の自動オプティマイザ検出方法は通常、制約のないコード空間上、または狭くパラメータ化されたオプティマイザ ファミリ内で検索します。前者は柔軟性がありますが、無効なプログラムや解釈できないプログラムが生成されることが多く、後者は安定していますが、新規性が制限されます。型指定されたドメイン固有言語 (DSL) でオプティマイザーを検出するための理論に基づいたマルチエージェント フレームワークである OPTScientist を紹介します。 OPTScientist は、制約付きの科学的検索プロセスとしてオプティマイザーの設計を定式化します。このプロセスでは、候補の更新が方向、スケーリング、事前条件付け、正則化、状態、およびグループ化モジュールを通じて表現されます。理論家、デザイナー、エンジニア、レビューアーの 4 つの役割のエージェントが単一のオーケストレーション ループ内で連携して、仮説の提案、DSL 候補の合成、オプティマイザーのコンパイルと評価、結果の批評を行います。固定検索空間の制限を克服するために、OPTScientist は、オプティマイザー プログラムによる進化的検索と、繰り返される障害によって表現上のボトルネックが明らかになった場合に小さな DSL 拡張を提案する第 2 段階のメカニズムを組み合わせます。このフレームワークを使用して、ネイティブ評価プロトコルの下で強力なベースラインに対する変換器の事前学習を改善する縮小状態行列オプティマイザーである RS-MR を発見します。私たちの結果は、理論、型指定されたプログラム、コンパイラーの検証、および閉ループ実験に基づいた自動オプティマイザー科学への道を示唆しています。
原文 (English)
OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining
Designing optimizers for modern deep learning remains a challenging scientific problem, requiring the joint consideration of optimization geometry, state dynamics, numerical stability, implementation constraints, and empirical generalization. Existing automated optimizer discovery methods typically search either over unconstrained code spaces or within narrowly parameterized optimizer families. The former is flexible but often produces invalid or uninterpretable programs, while the latter is stable but limits novelty. We introduce OPTScientist, a theory-guided multi-agent framework for optimizer discovery in a typed domain-specific language (DSL). OPTScientist formulates optimizer design as a constrained scientific search process, where candidate updates are expressed through direction, scaling, preconditioning, regularization, state, and grouping modules. Four role agents, Theorist, Designer, Engineer, and Reviewer, collaborate within a single orchestration loop to propose hypotheses, synthesize DSL candidates, compile and evaluate optimizers, and critique results. To overcome the limitations of a fixed search space, OPTScientist combines evolutionary search over optimizer programs with a second-stage mechanism that proposes small DSL extensions when repeated failures reveal representational bottlenecks. Using this framework, we discover RS-MR, a reduced-state matrix optimizer that improves transformer pretraining over strong baselines under our native evaluation protocol. Our results suggest a path toward automated optimizer science grounded in theory, typed programs, compiler validation, and closed-loop experimentation.
指向性幻覚: ニュースに基づいた LLM の質問応答におけるイデオロギーの漂流
大規模言語モデル (LLM) は、事実誤認やイデオロギーの歪みが大きな問題となる選挙関連の情報環境など、政治情報に関する質問に答えるためにますます使用されています。私たちは、幻覚や文書に基づく QA における裏付けのない発言を、イデオロギーの偏向の診断信号として扱う、再現可能な測定フレームワークを提示します。左派、中道派、右派の情報源にわたる QBias からの専門家ラベル付き米国政治ニュース記事 21,727 件を使用して、(i) 記事固有の質問を生成し、(ii) 3 つのオープンウェイト LLM と 1 つの独自モデルから文書に基づいた回答を導き出し、(iii) 参照ベースの比較によって文レベルの幻覚を検出し、(iv) 微調整されたスタンス分類器で幻覚文のイデオロギー価度を分類し、 (v) トークンレベルの不確実性を幻覚およびドリフトに関連付ける出力ログを調査します。幻覚の発生率はモデルによって大幅に異なり、議論の多いトピックに集中していますが、幻覚の頻度における情報源とイデオロギーの違いはわずかです。対照的に、幻覚コンテンツは強い左方向偏向を示します。右寄りの情報源から生成された幻覚も含め、幻覚文章の大部分は左寄りとして分類されます。ロジットレベルの分析では、幻覚は高エントロピー生成の状況で発生することが示されており、一部のモデルでは不確実性が左方向へのドリフトも予測しており、これは「推測に対する不確実性」メカニズムと一致しています。 AI を介した政治情報の監査と、選挙関連の展開における安全策の設計への影響について説明します。
原文 (English)
Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering
Large language models (LLMs) are increasingly used to answer questions about political information, including in election-adjacent information settings where factual errors and ideological distortions are high-stakes. We present a reproducible measurement framework that treats hallucinations, unsupported statements in document-grounded QA, as diagnostic signals of ideological drift. Using 21,727 expert-labeled U.S. political news articles from QBias spanning left, center, and right sources, we (i) generate an article-specific question, (ii) elicit document-grounded answers from three open-weight LLMs and one proprietary model, (iii) detect sentence-level hallucinations via reference-based comparison, (iv) classify the ideological valence of hallucinated sentences with a fine-tuned stance classifier, and (v) probe output logits to relate token-level uncertainty to hallucination and drift. Hallucination rates vary substantially across models and concentrate in contentious topics, while source-ideology differences in hallucination frequency are modest. In contrast, hallucination content exhibits robust leftward drift: a majority of hallucinated sentences are classified as left-leaning, including among hallucinations generated from right-leaning sources. Logit-level analysis shows hallucinations arise in high-entropy generation contexts, and in some models uncertainty also predicts leftward drift, consistent with an "uncertainty to guessing" mechanism. We discuss implications for auditing AI-mediated political information and for designing safeguards in election-relevant deployments.
自律的なトポロジの突然変異: 機能、状態、およびシャドウの不変条件を備えたマルチエージェント LLM システムの安全なランタイム再構築
マルチエージェント LLM フレームワークは通常、起動時にチーム トポロジを修正します。実行時に個々のエージェントが過負荷になると、たとえば、アクション カテゴリが多すぎる、ツール エラーが蓄積する、またはコールが多すぎるキューに入れすぎるなどして、システムにはそれ自体を再構築するメカニズムがありません。マルチエージェント LLM フレームワークのランタイム チーム ミューテーション メカニズムである Autonomous Topology Mutation (ATM) を紹介します。 ATM は、テレメトリ主導の過負荷検出と、各構造変化を制御する 3 つの安全性不変条件 (機能の単調性、状態ルーティングの完全性、シャドウ ビフォア ライブ検証) を組み合わせています。 ATM は、キューの深さ、コンテキスト スラッシング、ツール エラー率、ロール エントロピー、再試行ループ率、およびエージェント間待機時間を含む 6 つのシグナルのボトルネック インデックスを監視します。ウォームアップで調整されたしきい値が複数の連続ティックで違反されると、ATM は過負荷になったエージェントを特殊なサブエージェントに分解し、外部 ID を維持しながら親をコーディネーターの役割にホットスワップします。状態転送は、プライバシー レベルを認識したルーティングによって制御されます。各メモリ アトムは、許可された子セットにのみルーティングされるか、ログに記録された理由で明示的に削除されます。シャドウ検証ウィンドウを通過するまで、候補トポロジはライブ トラフィックを受信しません。 4 つのアブレーション条件と 3 つのワークロードにわたって決定論的なツール スタブを使用して 720 個の DeepSeek-V3 駆動タスクを実行すると、ATM 因数分解器分割によりコードタスクの成功率が 3.3% から 61.7% に上昇しました。完全なレールアンド蒸留システムは、タスクの品質を維持しながら、正規表現分類子の下で検出されるプライバシー性の高いメモリの露出をタスクあたり 2.0 から 0.0 イベントに削減します。 ATM の不変条件を保持するランタイム レールにより、エージェントのホット パスで追加される p99 遅延は 500 マイクロ秒未満になります。実際の Python を実行する小さなライブツール プローブが、外部妥当性チェックとして含まれています。実装、ベンチマーク ハーネス、およびトレースはオープンソースです。
原文 (English)
Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants
Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for example by mixing too many action categories, accumulating tool errors, or queueing behind too many calls, the system has no mechanism to restructure itself. We introduce Autonomous Topology Mutation (ATM), a runtime team-mutation mechanism for multi-agent LLM frameworks. ATM combines telemetry-driven overload detection with three safety invariants that gate each structural change: capability monotonicity, state-routing completeness, and shadow-before-live validation. ATM monitors a six-signal Bottleneck Index that includes queue depth, context thrash, tool-error rate, role entropy, retry-loop rate, and cross-agent wait time. When a warmup-calibrated threshold is breached for multiple consecutive ticks, ATM factorises the overloaded agent into specialised sub-agents and hot-swaps the parent into a coordinator role while preserving its external identity. State transfer is controlled by privacy-level-aware routing: each memory atom is routed only to a permitted child set, or explicitly dropped with a logged reason. No candidate topology receives live traffic until it has passed a shadow validation window. On 720 DeepSeek-V3-driven task runs with deterministic tool stubs across four ablation conditions and three workloads, the ATM factoriser split lifts code-task success from 3.3% to 61.7%. The full rail-and-distillation system reduces detected high-privacy memory exposure under a regex classifier from 2.0 to 0.0 events per task while preserving task quality. The runtime rails carrying ATM's invariants add less than 500 microseconds of p99 latency on the agent hot path. A small live-tool probe with real Python execution is included as an external-validity check. The implementation, benchmark harness, and traces are open-sourced.
EvoSQL: テキストから SQL へのメモリ拡張クリティック ジェネレーターの共進化
Text-to-SQL は大規模な言語モデルによって急速に進歩しましたが、複雑なデータベース クエリには依然として、マルチステップの分解、実行ベースの診断、ターゲットを絞った修正など、ワンショット生成を超える推論が必要です。我々は、ジェネレーターとクリティカルの間の反復的な対話として SQL 合成を定式化する共進化フレームワークである EvoSQL を紹介します。 EvoSQL は、コンテキスト化された候補メモリを維持し、実行シグナルと LLM ベースの批評の両方で SQL 候補を検証し、ユーティリティに基づく集計を通じてメモリを更新します。基礎となるジェネレーターとクリティカルのペアを強化するために、最新のコーディング LLM バックボーンに実行を意識した監視を注入する自己蒸留ポリシー最適化 (SDPO) 微調整ステージをさらに導入します。 Spider と BIRD の実験では、EvoSQL が Maj@16 ベースラインよりもオープンソース モデルを一貫して改善し、特に BIRD-Dev で大きな改善が見られ、Qwen3-4B の +1.37% から Qwen2.5-Coder-3B の +9.19% までの範囲に及ぶことがわかりました。 SDPO の初期化により、Spider-Test および BIRD-Dev で選択されたバックボーンがさらに改善されます。これらの結果は、メモリに基づいた共進化が、より信頼性が高く一般化可能な Text-to-SQL システムへの効果的な道であることを示唆しています。コードは https://github.com/valleysprings/EvoSQL で入手できます。
原文 (English)
EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL
Text-to-SQL has advanced rapidly with large language models, but complex database queries still require reasoning beyond one-shot generation, including multi-step decomposition, execution-based diagnosis, and targeted correction. We present EvoSQL, a co-evolution framework that formulates SQL synthesis as an iterative interaction between a generator and a critic. EvoSQL maintains a contextualized candidate memory, verifies SQL candidates with both execution signals and LLM-based critique, and updates its memory through utility-guided aggregation. To strengthen the underlying generator-critic pair, we further introduce a Self-Distillation Policy Optimization (SDPO) fine-tuning stage that injects execution-aware supervision into modern coding LLM backbones. Experiments on Spider and BIRD show that EvoSQL consistently improves open-source models over Maj@16 baselines, with particularly large gains on BIRD-Dev, ranging from +1.37% for Qwen3-4B to +9.19% for Qwen2.5-Coder-3B. SDPO initialization further improves selected backbones on Spider-Test and BIRD-Dev. These results suggest that memory-grounded co-evolution is an effective path toward more reliable and generalizable Text-to-SQL systems. Code is available at https://github.com/valleysprings/EvoSQL.
CRAWO: アダプティブ ワークロード オーケストレーション用のカスタム リソース
エッジ インテリジェンスは、集中化されたクラウド データ センターからネットワーク エッジにコンピューティングを移行し、それによって遅延と帯域幅の消費を削減することにより、スマート シティでリアルタイム アプリケーションを実現するための重要なパラダイムとして浮上しています。ただし、低電力マイクロコントローラーからアクセラレータ搭載システムに至るまで、デバイスの機能が広範囲に及ぶため、異種エッジ インフラストラクチャ全体に人工知能 (AI) パイプラインを展開することは依然として困難です。既存のエッジ オーケストレーション プラットフォームは主に展開の自動化とインフラストラクチャ管理に重点を置いていますが、これらのアプローチは多くの場合非効率的であり、動的な条件下でリソースを適応的に割り当てる能力が制限されます。これらの問題に取り組むために、このホワイト ペーパーでは、分散エッジ環境全体で AI パイプラインを調整するためのアーキテクチャ フレームワークである CRAWO (Custom Resources for Adaptive Workload Orchestration) を紹介します。 CRAWO は、エッジ ノード上でサービスをインスタンス化しながら、配置決定、状態管理、およびステージ間のデータ フローを管理することで、割り当てインテリジェンスを実行から分離する制御ループ ベースのモデルに従います。このフレームワークには、リアルタイムのインフラストラクチャ メトリクスを活用して適応的なワークロード配置を可能にする、プラグイン可能な複数基準の意思決定層を備えたハードウェア対応アロケータが組み込まれています。リファレンス実装では、軽量の Kubernetes ディストリビューション (K3s) 上にデプロイされたマイクロサービス アーキテクチャが採用されており、ドメイン モデリングにはカスタム リソース定義 (CRD)、状態調整には専用のオペレーターが使用されます。ナンバープレート認識を使用した車両監視シナリオでの評価では、ワークロード分散が改善され、遅延の影響を受けやすい環境における集中クラウド処理への依存が軽減されることが実証されました。
原文 (English)
CRAWO: Custom Resources for Adaptive Workload Orchestration
Edge Intelligence has emerged as a key paradigm for enabling real-time applications in smart cities by shifting computation from centralized cloud data centers to the network edge, thereby reducing latency and bandwidth consumption. However, deploying Artificial Intelligence (AI) pipelines across heterogeneous edge infrastructures remains challenging due to the wide range of device capabilities, from low-power microcontrollers to accelerator-equipped systems. Existing edge orchestration platforms primarily focus on deployment automation and infrastructure management, but these approaches are often inefficient and limit the ability to adaptively allocate resources under dynamic conditions. To tackle these issues, this paper introduces CRAWO (Custom Resources for Adaptive Workload Orchestration), an architectural framework for coordinating AI pipelines across distributed edge environments. CRAWO follows a control-loop-based model that separates allocation intelligence from execution by managing placement decisions, state management, and inter-stage data flows while instantiating services on edge nodes. The framework incorporates a hardware-aware allocator with a pluggable multi-criteria decision layer that leverages real-time infrastructure metrics to enable adaptive workload placement. The reference implementation adopts a microservices architecture deployed on a lightweight Kubernetes distribution (K3s), using Custom Resource Definitions (CRDs) for domain modeling and a dedicated operator for state reconciliation. Evaluation in a vehicle surveillance scenario using license plate recognition demonstrates improved workload distribution and reduced reliance on centralized cloud processing in latency-sensitive environments.
DFAH ベンチ: 財務上の意思決定における観察可能なエージェントの不安定性のベンチマーク
標準の評価ベンチマークは、ツールを使用するエージェントが毎回同じプロセスを経てその決定に到達するかどうかではなく、ツールを使用するエージェントが何を決定するかを測定します。 DFAH ベンチは、金融機関の意思決定における観察可能な行動の不安定性を 3 つのチャネル (ツール呼び出しの軌跡、証拠の接触、意思決定の集中) にわたって測定するリプレイ ベンチマークです。これらのチャネルのいずれも、隠された推論テキストへのアクセスを必要としません。 10 のモデルと 3 つの財務タスクにまたがる 8,127 のリプレイ エピソード全体で、結果の合意だけでは不完全な安定性シグナルであることがわかりました。フロンティア モデルは 95% の確率で意思決定に合意できるのに対し、同じツール パスをたどるのは 77% の確率だけです。結果のみの評価では完全に見逃される 18 パーセント ポイントのギャップ (95% CI: [0.14, 0.22])。意思決定の一致度が高いフロンティアモデルのケースグループのうち、55% 以上が意味のある軌道の発散を示しています。私たちは 3 つの動作プロファイルを特定します。入力に関係なく単一の出力に集約することでほぼ完全な一致を達成するパターン マッチャー、比較的一貫したツール使用プロセスを持つ安定した実行者、および実質的に異なるツール パスと証拠の接触を通じて同じ結論に達する軌道分岐者です。ベンチマーク コード、メトリック スクリプト、再生ログ、ベンチマーク カード、データセット README、およびリリース マニフェストは、付属のリポジトリでリリースされます。
原文 (English)
DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making
Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We introduce DFAH-Bench, a replay benchmark that measures observable behavioral instability in financial agent decision-making across three channels -- tool-call trajectories, evidence contacts, and decision concentration -- none of which require access to hidden reasoning text. Across 8,127 replay episodes spanning 10 models and 3 financial tasks, we find that outcome agreement alone is an incomplete stability signal: frontier models can agree on decisions 95% of the time while following the same tool path only 77% of the time -- an 18-percentage-point gap (95% CI: [0.14, 0.22]) that outcome-only evaluation misses entirely. Among frontier-model case groups with high decision agreement, over 55% exhibit meaningful trajectory divergence. We identify three behavioral profiles: pattern matchers that achieve near-perfect agreement by collapsing to a single output regardless of input, stable executors with relatively consistent tool-use processes, and trajectory divergers that reach the same conclusions through materially different tool paths and evidence contacts. The benchmark code, metric scripts, replay logs, benchmark card, dataset README, and release manifest are released in the accompanying repository.
不可知論的な時系列予測モデルの継続学習のためのアテンションベースのエクスペリエンス再生フレームワーク
ディープラーニングは、人工知能、特にロボット工学、画像処理、音声処理において目覚ましい進歩をもたらしました。ただし、ニューラル ネットワークの大きな制限は、大規模で定常的なデータセットに強く依存していることです。実際のアプリケーションの多くでは、データの分布が時間の経過とともに変化する進化する動的な環境のため、これらの条件が満たされることはほとんどありません。継続的学習は、計算上の制約の下で安定性と可塑性のバランスを維持しながら段階的に適応できるモデルを開発することで、この課題に対処することを目的としています。この研究では、継続的な時系列予測のための新しいフレームワークを紹介します。このフレームワークは、アテンション メカニズムによってガイドされるエクスペリエンス リプレイ戦略を組み込むことで、文献で一般的に使用されている既存の静的予測モデルを拡張するように設計されています。このアプローチにより、モデルは事前の知識を維持しながら新しいコンテキストに動的に適応でき、壊滅的な忘却を効果的に軽減できます。このフレームワークは、標準的な予測ベンチマークと、さまざまな時間的挙動を示すピエゾメトリック データセットで評価されます。結果は、私たちのアプローチが、再トレーニングのコストとデータ要件を削減しながら、時間の経過とともに予測パフォーマンスを効果的に向上または維持するため、動的かつ現実世界の設定での予測モデルの展開を容易にすることを示しています。
原文 (English)
Attention-based Experience Replay Framework for Continual Learning of Agnostic Time Series Forecasting Models
Deep learning has led to remarkable progress in artificial intelligence, particularly in robotics, imaging and sound processing. However, a major limitation of neural networks remains their strong dependence on large and stationary datasets. In many real-world applications, these conditions are rarely met due to evolving and dynamic environments where data distributions change over time. Continual learning aims to address this challenge by developing models capable of adapting incrementally while maintaining a balance between stability and plasticity under computational constraints. In this work, we introduce a novel framework for continual time series forecasting, designed to extend existing static forecasting models commonly used in the literature by incorporating an Experience Replay strategy guided by Attention mechanisms. This approach allows the model to adapt dynamically to new contexts while preserving prior knowledge, effectively mitigating catastrophic forgetting. The framework is evaluated on standard forecasting benchmarks as well as on a piezometric dataset exhibiting diverse temporal behaviors. Results show that our approach effectively increases or maintains predictive performance over time while reducing retraining costs and data requirements, thus facilitating the deployment of forecasting models in dynamic and real-world settings.
LLM アライメントを正規表現から分離する: 敵対的突然変異下でのゼロ カバレッジとメトリック依存の発散
実稼働 LLM アプリケーションは通常、モデル側の位置合わせの前に正規表現フィルターをスタックします。以前の研究では、アクティブな正規表現フィルターの背後にライブ Gemini バックエンドを追加しても、測定可能なカバレッジの向上は見つかりませんでした。コーパスが \emph{正規表現をバイパスするように設計されている} 場合に、その上限が維持されるかどうかを尋ねます。 $L_5$-no-regex -- $L_4$-real (Gemini-2.5-フラッシュ、トークンバジェットキャップ、レート制限、出力スクラブ) と同じですが、9 パターンフィルターが無効になっています -- を導入し、3 つのサブコーパス (キャリーフォワード、正規表現バイパス、アライメント分離) にわたる $N{=}45$ の敵対的プローブに対して評価します。これは、Gemini のパラフレーズと$N{=}5$ レプリケーションで ${\sim}1{,}555$ プローブ実行ペアにペアリングします。主要な部分文字列分類器の下では、H1 は反駁されます。$L_5$ ブロック率は、5 つの OWASP LLM トップ 10 カテゴリすべてで $0,%$ です ($\Delta\text{pp}{=}0$ 対\ $L_0$, $p{=}1.00$、ウィルソン上限 ${<}5,%$)。 PAIR バリアントに関する 2 番目の LLM 判定メトリクスは、$56$--$100,%$ ブロック率 ($p{<}0.01$) を示し、アライメントが敵対的にフレーム化されたプローブに応答することを明らかにしていますが、部分文字列の一致には微妙すぎる拒否を生成します。サブコーパスの差分予測はサポートされていません ($p{=}1.00$)。アライメントの貢献は \emph{metric-dependent} です。自然言語の有害なリクエスト プローブでは、正規表現を超えて観測されたカバレッジがゼロになります。敵対的にフレーム化されたバリアントでは、LLM ジャッジが部分文字列分類子が見逃した拒否を検出します。ロックされたコーパス、ミューテーション アーティファクト、およびエクスポート スクリプトは、レプリケーションのために解放されます。
原文 (English)
Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation
Production LLM applications commonly stack a regex filter in front of model-side alignment; prior work found no measurable coverage gain from adding a live Gemini backend behind an active regex filter. We ask whether that ceiling holds when the corpus is \emph{designed to bypass the regex}. We introduce $L_5$-no-regex -- identical to $L_4$-real (Gemini-2.5-flash, token-budget cap, rate limit, output scrub) but with the nine-pattern filter disabled -- and evaluate it against $N{=}45$ adversarial probes across three sub-corpora (carry-forward, regex-bypass, alignment-isolate), amplified by Gemini paraphrase and PAIR to ${\sim}1{,}555$ probe-run pairs over $N{=}5$ replications. Under the primary substring classifier, H1 is refuted: $L_5$ block rate is $0,%$ across all five OWASP LLM Top-10 categories ($\Delta\text{pp}{=}0$ vs.\ $L_0$, $p{=}1.00$; Wilson upper bound ${<}5,%$). A secondary LLM-judge metric on PAIR variants shows $56$--$100,%$ block rates ($p{<}0.01$), revealing alignment does respond to adversarially-framed probes -- but produces refusals too nuanced for substring matching. The sub-corpus differential prediction is not supported ($p{=}1.00$). Alignment's contribution is \emph{metric-dependent}: on natural-language harmful-request probes, it adds zero observed coverage beyond the regex; on adversarially-framed variants, an LLM judge detects refusals the substring classifier misses. The locked corpus, mutation artifacts, and export scripts are released for replication.
マルチエージェント システム向けのワークロード認識キャッシング
マルチエージェント システムは、複雑なタスクを特殊なエージェント実行の有向非巡回グラフ (DAG) に分解し、クエリ全体で中間結果をキャッシュする自然な機会を生み出します。ただし、既存のキャッシュ削除ポリシーは、アクセス履歴に基づいてキャッシュされたすべてのエントリを一律に扱い、エージェント実行環境で固有に利用可能な構造およびワークロードの信号を無視します。ここでは、再計算コスト、DAG 依存関係数、エージェント呼び出し頻度という 3 つのシグナルを、メモリ制約下で最も価値のあるエントリを保持する統合スコアリング関数に結合する、ワークロードを認識したエビクション ポリシーを紹介します。さまざまな再利用体制にわたる 3 つのマルチエージェント ベンチマークで評価したところ、当社のポリシーは、キャッシュなしのベースラインと比較してレイテンシを最大 64.7% 削減し、次に優れた有限容量ベースラインと比較して平均 31.1% のレイテンシ削減を達成しながら、無制限キャッシュのパフォーマンスに近づき、競合するすべての有限容量手法と同等またはそれを超える精度を維持しています。さらに、ワークロードを認識したコンテンツ キャッシュが、プラン レベルのキャッシュや並列エージェント実行などの他のエージェント システム最適化手法を補完し、各手法がマルチエージェント パイプラインにおける個別の効率ボトルネックをターゲットにしていることを示します。
原文 (English)
Workload-Aware Caching for Multi-Agent Systems
Multi-agent systems decompose complex tasks into directed acyclic graphs (DAGs) of specialized agent executions, creating natural opportunities for caching intermediate results across queries. However, existing cache eviction policies treat all cached entries uniformly based on access history, ignoring structural and workload signals uniquely available in agentic execution environments. We present a workload-aware eviction policy that combines three signals, namely recomputation cost, DAG dependency count, and agent invocation frequency, into a unified scoring function that retains the most valuable entries under memory constraints. Evaluated across three multi-agent benchmarks spanning diverse reuse regimes, our policy reduces latency by up to 64.7% relative to the uncached baseline and achieves on average a 31.1% latency reduction over the next best finite-capacity baseline, while approaching the performance of an unbounded cache and maintaining accuracy on par with or exceeding all competing finite-capacity methods. We further show that workload-aware content caching is complementary to other agentic system optimization methods, including plan-level caching and parallel agent execution, with each technique targeting a distinct efficiency bottleneck in multi-agent pipelines.
エラーからルールへ: テキスト分類のための反復プロンプト最適化
テキスト分類の迅速な最適化は、デモンストレーションの選択から探索ベースの検索、エラー主導の診断に至るまで、さまざまなアプローチに及びます。それぞれのアプローチには、既知ではあるが不完全に特性化された長所と限界があります。私たちは、最適化トレースの定量的評価と定性分析の両方を通じてこれらのパラダイムを比較する、多様な分類ベンチマーク (2 ~ 150 クラス) にわたる包括的な実証研究を実施します。これにより、各パラダイムが構造的に異なるタスク タイプで優れており、単一の手法が優勢ではないことが明らかになります。これらの洞察に基づいて、私たちは、重複しないバッチでトレーニング セット全体を反復し、分類の失敗を診断し、診断、処方、書き換えのフィードバック ループを通じて対象を絞った決定ルールを生成する、エラー主導型の手法であるエラーガイド最適化 (ERGO) を提案します。 ERGO は、エラーが特定の混乱したラベルのペアに集中するタスク (これを境界学習可能タスクと呼びます) で最高の精度を達成します。TREC: 90.0%、CLINC150: 94.4%、3 ~ 5 回の反復で収束し、解釈可能な決定ルールを生成します。 ERGO は最高の全体平均を達成することはできませんが、補完的な役割を果たしています。つまり、カバレッジに依存するタスクではデモンストレーション ベースの ICL が勝利し、多クラス インテントでは探索ベースの検索が勝利し、エラー パターンから意思決定の境界が学習できる場合は ERGO が勝利します。私たちはタスクの特性を最適なパラダイムの選択に結び付ける補完性フレームワークを提供し、実践者に実践的なガイダンスを提供します。
原文 (English)
From Errors to Rules: Iterative Prompt Optimization for Text Classification
Prompt optimization for text classification spans diverse approaches, from demonstration selection to exploration-based search to error-driven diagnosis, each with known but incompletely characterized strengths and limitations. We conduct a comprehensive empirical study across diverse classification benchmarks (2 to 150 classes) comparing these paradigms through both quantitative evaluation and qualitative analysis of optimization traces, revealing that each paradigm excels on structurally different task types and that no single method dominates. Guided by these insights, we propose Error-Guided Optimization (ERGO), an error-driven method that iterates over the full training set in non-overlapping batches, diagnoses classification failures, and generates targeted decision rules through a diagnose-prescribe-rewrite feedback loop. ERGO achieves the best accuracy on tasks where errors concentrate in specific confused label pairs (which we term boundary-learnable tasks): TREC: 90.0%, CLINC150: 94.4%, converges in 3-5 iterations, and produces interpretable decision rules. While ERGO does not achieve the highest overall average, it fills a complementary role: demonstration-based ICL wins on coverage-dependent tasks, exploration-based search wins on many-class intent, and ERGO wins where decision boundaries are learnable from error patterns. We provide a complementarity framework linking task characteristics to optimal paradigm selection, offering practical guidance for practitioners.
AISE-Bench: 学術知識グラフ上の情報探索のためのフルサイクルで厳選されたベンチマーク
ツールで強化された大規模言語モデル (LLM) は、Web エンジン、API、コードを使用して複雑で長期的なタスクを解決できる自律エージェントとして登場しつつあります。学術グラフで情報を求める現在のツールを使用したベンチマークは、合成テンプレート、簡素化されたソリューション空間、または紙中心のタスクなどの狭いタスクに依存しており、現実的なユーザーの意図、複雑な複数ステップの API プランニング、API の豊富なパラメーターの入力、参照を伴う根拠のある回答、プロセスと結果の両方の包括的な評価など、重要な課題が十分に検討されていないままになっています。 AISE-Bench は、学術知識グラフ上で情報を求めるための、現実世界のフルサイクルの注釈付きベンチマークです。 AISE-Bench リリースには、クエリ分類、完全な API 実行軌跡、検証済みパラメータ、参照リンク付きのソースに基づいた回答を含む 1,133 の QA ペアが含まれています。高品質のアノテーションをサポートするために、アノテーターが複雑な API ワークフローを効率的に計画、実行、修正できるようにカスタマイズされたエージェント ワークフローを設計します。私たちは、回答の品質、基準の根拠、API 計画の正確さ、および実行の成功を測定する包括的な評価プロトコルを開発します。評価された 14 のメソッドのうち、最も強力なモデル (Gemini-3-Pro を使用した PLAY2PROMPT) でさえ中程度のパフォーマンスしか達成できず、API の計画と実行に苦労することがよくあります。 AISE-Bench は、LLM エージェントを使用するマルチステップ API の段階的な正確性、根拠のある要約、および追跡可能な推論を定量的に評価および改善するための、挑戦的な新しいテストベッドを確立します。コードとデータは https://aise-bench.github.io/ で入手できます。
原文 (English)
AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks. Current tool-using benchmarks for information seeking on academic graphs rely on synthetic templates, simplified solution spaces, or narrow tasks such as paper-centric tasks, leaving key challenges underexplored - realistic user intent, complex multi-step API planning, rich parameter filling for APIs, grounded answers with references, and comprehensive evaluation of both the process and the outcome. We introduce AISE-Bench, a real-world, full-cycle annotated benchmark for information seeking on academic knowledge graphs. AISE-Bench release contains 1,133 QA pairs, including query taxonomies, full API execution trajectories, validated parameters, and source-grounded answers with reference links. To support high-quality annotation, we design a customized agent workflow to enable annotators to plan, execute, and revise complex API workflows efficiently. We develop a comprehensive evaluation protocol measuring answer quality, reference grounding, API-planning correctness, and execution success. Among the 14 evaluated methods, even the strongest model (PLAY2PROMPT with Gemini-3-Pro) achieves only moderate performance and often struggles with API planning and execution. AISE-Bench establishes a challenging new testbed for quantitatively evaluating and improving the stepwise correctness, grounded summarization, and traceable reasoning of multi-step API-using LLM agents. Our code and data are available at https://aise-bench.github.io/.
ExecuGraph: 大規模な言語モデルを使用した信頼性の高いバックエンド コード合成のための、マルチエージェントの実行ベースのフレームワーク
大規模言語モデルは妥当なバックエンド コードを生成しますが、シングルパス パラダイムでは正確性や実行時の信頼性が保証されません。実行ベースの検証をバックエンド コード合成の中心に置くマルチエージェント フレームワークである ExecuGraph を紹介します。 6 つの専門エージェント (プランナー、コード ジェネレーター、論理レビューアー、評価者、オプティマイザー、および説明者) は、ローカルでホストされるモデル (Ollama) とアルゴリズム手法を呼び出すためのオプションの検索レイヤーを使用して LangGraph 上に実装された、制限された再試行予算を備えた型指定されたワークフローによって調整されます。実時間タイムアウトを備えたサブプロセス分離サンドボックスにより、すべての評価が保護されます。私たちは、厳選された 30 の問題の DSA スイート (内部 30)、HumanEval (n=64)、および APPS 入門サブセットを評価し、ExecuGraph を単一エージェントのワンショット ベースラインおよび単一エージェントの実行再試行ベースライン (マルチエージェント分解の寄与を分離するリフレクション スタイルのアブレーション) と対比させます。内部 30 では、3 つの条件は統計的に区別できません (n=30; ペアのウィルコクソン p=0.59 MF 対 SO、p=0.08 SR 対 SO)。すべてのペアごとの平均差に関する 95% ブートストラップ信頼区間にはゼロが含まれます。 HumanEval では、マルチフル エッジが +3.1 pp 先行しています。最も強い信号はクロスモデルです。DeepSeekCoder V2 Lite では、グラフ カテゴリの精度が 57.5% (ワンショット) から 80.0% (マルチフル) に向上し、スケーリング仮説をサポートする +22.5 pp のジャンプです。マルチエージェント分解の値はベースモデルの機能とともに増加します。このフレームワークの主な貢献は方法論的です。構成によってワンショット、実行再試行、およびエージェントごとのアブレーション条件に折りたたまれる単一のコードベースにより、各レバーの限界寄与の制御された測定が可能になります。エージェントごとのアブレーション、再試行予算のスイープ、エラークラスの分類、およびテストソースの監査が報告されます。
原文 (English)
ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models
Large Language Models generate plausible backend code, but a single-pass paradigm provides no guarantee of correctness or runtime reliability. We present ExecuGraph, a multi-agent framework that places execution-based validation at the center of backend code synthesis. Six specialized agents (Planner, Code Generator, Logical Reviewer, Evaluator, Optimizer, and Explainer) are coordinated by a typed directed workflow with a bounded retry budget, implemented on LangGraph with locally hosted models (Ollama) and an optional retrieval layer for algorithmic technique recall. A subprocess-isolated sandbox with a wall-clock timeout guards every evaluation. We evaluate on a curated 30-problem DSA suite (internal-30), HumanEval (n=64), and an APPS-introductory subset, contrasting ExecuGraph against a single-agent one-shot baseline and a single-agent execution-retry baseline (a Reflexion-style ablation that isolates the contribution of multi-agent decomposition). On internal-30, the three conditions are statistically indistinguishable (n=30; paired Wilcoxon p=0.59 MF vs. SO, p=0.08 SR vs. SO); 95% bootstrap confidence intervals on all pairwise mean differences include zero. On HumanEval, multi-full edges ahead by +3.1 pp. The strongest signal is cross-model: with DeepSeekCoder V2 Lite, graph-category accuracy improves from 57.5% (oneshot) to 80.0% (multi-full), a +22.5 pp jump that supports a scaling hypothesis: the value of multi-agent decomposition grows with base-model capability. The framework's primary contribution is methodological: a single codebase that collapses by configuration into one-shot, execution-retry, and per-agent ablation conditions, enabling controlled measurement of each lever's marginal contribution. A per-agent ablation, retry-budget sweep, error-class taxonomy, and test-source audit are reported.
FlowEdit: 競合を伴う不適切な問題に対する LLM 推論フローの情報理論的制御
大規模言語モデル (LLM) は、適切に指定された推論タスクで強力に実行され、実行可能な答えが得られます。ただし、オープンワールドで遭遇する問題は、矛盾した条件、矛盾するステートメント、または相互に互換性のない要件によって不適切な状態になり、有効な応答が得られない可能性があります。競合を伴うこのような不適切な問題の推論には、隠れた競合を明示し、複数の推論分岐を介して競合する仮説を維持し、単一パスで代替応答を生成するための新しい LLM 機能が必要であると主張しますが、これらのすべては、LLM の次トークン予測メカニズムの制限により困難です。この目的を達成するために、有効な仮説に基づいて完全な代替応答を生成するために、情報理論の原理を活用して LLM の内部推論フローを定量化および調整する新しいフレームワークである FlowEdit を提案します。 FlowEdit は、モデルの内部推論表現に 2 つの二重の情報理論的目標を使用して分岐を意識した推論プロセスを強制するものと見なすことができます。つまり、選択された各仮説から分岐の結果までの情報フローを最大化すると同時に、兄弟分岐間の重複と条件依存を最小限に抑えて、広範囲をカバーする多様で有益な応答のセットを提供します。これは、境界埋め込みの下での扱いやすい変分限界が {\epsilon} 十分であり、LLM 推論プロセスにおける基礎となる条件付き相互情報量を最適化することによって達成されることを示します。広範な実験により、FlowEdit が主要な独自モデルよりも優れたパフォーマンスを示し、完全セット一致精度が 68% 向上し、全体的な応答の情報量が 24% 向上することが実証されました。さらに、フロー制御は、各ブランチ内に集中し、フロー境界で増幅し、問題が必要とするフローの数に応じて拡大する次のトークンのエントロピーの再分布としてトークン ストリーム内に現れることを示します。
原文 (English)
FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts
Large Language Models (LLMs) perform strongly on well-specified reasoning tasks with a feasible answer. However, problems encountered in the open world can become ill-posed due to inconsistent conditions, conflicting statements, or mutually incompatible requirements, admitting no valid responses. We argue that reasoning of such ill-posed problems involving conflicts require novel LLM capabilities to make hidden conflicts explicit, maintain competing hypotheses via multiple reasoning branches, and generate alternative responses in a single pass, all of which are challenging due to the limitation of the next-token prediction mechanism in LLMs. To this end, we propose FlowEdit, a novel framework that leverages information-theoretic principles to quantify and regulate internal reasoning flows of LLMs, for generating a full set of alternative responses under valid hypotheses. FlowEdit can be viewed as enforcing a branch-aware reasoning process using two dual information-theoretic objectives on the model's internal reasoning representations: maximizing the information flow from each selected hypothesis to the branch outcome, while minimizing the overlap and conditional dependence across sibling branches, to provide a diverse, informative set of responses with broad coverage. We show that this is achieved through tractable variational bounds under boundary embeddings being {\epsilon}-sufficient, optimizing the underlying conditional mutual information in LLM reasoning process. Extensive experiments demonstrate that FlowEdit outperforms leading proprietary models, improving exact-set-match accuracy by 68%, while boosting overall response informativeness by 24%. We further show that flow regulation surfaces in the token stream as a redistribution of next-token entropy that concentrates inside each branch, amplifies at flow boundaries, and scales with the number of flows the problem requires.
MKEvolve: カーネル コード生成のためのモジュール式マルチエージェント フレームワーク
LLM ベースのコード生成は急速に進歩していますが、ハードウェア アクセラレータ用に正しくパフォーマンスの高いカーネルを作成することは、最新の ML ワークロードをスケーリングする上で依然として重要なボトルネックとなっています。我々は、複雑な PyTorch モジュールのモジュール分解と各サブモジュールの LLM 生成カーネルを反復的に共進化させるフレームワークである MKEvolve (Modular Kernel Evolve) を紹介します。LLM 駆動のビーム検索によって各サブカーネルを個別に改善しながら、反復間での分割と融合によって分解を洗練します。結果として得られるカーネルは、独立して検証されたサブカーネルをプログラム的に構成したものであり、構成可能 (サブカーネル実装は交換可能)、解釈可能 (エラーと高速化は特定のサブカーネルに追跡可能) で、関連するモデル アーキテクチャに容易に適応できます。 KernelBench L2 および L3 上の Triton を使用した、マルチオペレーター シーケンスと完全なモデル アーキテクチャにわたる実験では、MKEvolve が LLM トークンの使用量を最大 35% 削減しながら、エンドツーエンドの直接合成ベースラインよりも正確性と速度の両方が向上することが示されています。
原文 (English)
MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation
Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scaling modern ML workloads. We present MKEvolve (Modular Kernel Evolve), a framework that iteratively co-evolves a modular decomposition of complex PyTorch modules and the LLM-generated kernel for each submodule, refining the decomposition by splitting and fusing across iterations while independently improving each subkernel via LLM-driven beam search. The resulting kernels are programmatic compositions of independently verified subkernels, making them configurable (subkernel implementations are swappable), interpretable (errors and speedups are traceable to specific subkernels), and readily adaptable to related model architectures. Experiments with Triton on KernelBench L2 and L3, spanning multi-operator sequences and full model architectures, show that MKEvolve improves both correctness and speedup over end-to-end direct synthesis baselines while reducing LLM token usage by up to 35%.
因数分解された確率分布の比較可能性の誘導
同一ではない変数セットに対して定義された 2 つの確率的グラフィカル モデル間の原則に基づいた比較を可能にするには、それらを共通の測定可能な空間に持ち上げる必要があります。この目的を達成するために、任意の 2 つのモデルの拡張スキームを提案し、正式な基盤を確立します。一致しないコンポーネントは、結果として得られる結合分布が元の分布と乗算定数だけ異なり、投影下で一致するように、条件付き一様 (ラプラス) 拡張を使用して完成されます。これにより、確率的なセマンティクスが維持されると同時に、明確に定義された分布の不一致尺度の適用が可能になります。投影下での誘導ジョイントの不変性を確立し、拡張を使用して、決定論的アルゴリズムによる共通のグラフィック構造だけでなく最小の共通測定可能空間への 2 つのファクター グラフの最小限の構造拡張を提供します。さらに、構造的および測定理論的特性について議論し、比較方法論の有望な基準を特定します。
原文 (English)
Inducing Comparability of Factorised Probability Distributions
To allow for principled comparison between two probabilistic graphical models defined over non-identical variable sets, they have to be lifted to a common measurable space. To this end, we propose an extension scheme for any two given models and establish the formal foundation: Unmatched components are completed using conditionally uniform (Laplace) extensions such that the resulting joint distributions differ from the original ones only by multiplicative constants and coincide under projection. This preserves the probabilistic semantics while enabling the application of well-defined distributional discrepancy measures. We establish the invariance of the induced joint under projection and use the extensions to provide a minimal structural extension of two factor graphs to the smalles common measurable space as well as to a common graphical structure by a deterministic algorithm. In addition, we discuss structural and measure-theoretic properties and identify promising criteria for comparison methodologies.
LeanFlow: ワークフロー主導のリーン自動化のケーススタディ
私たちは、数学論文を構築可能なリーン プロジェクトに変換することに特化した LLM エージェント システムである LeanFlow を紹介し、評価します。最近の検証者インザループ システムでは、大規模な形式的な成果物が生成される可能性があることが示されていますが、どのランタイム メカニズムがドキュメントからプロジェクトへの形式化の完了、監査可能性、または効率に影響を与えるかは不明のままです。私たちは、モデル、証明ワークフロー、および Kimi2.6 と GPT5.5 によるツールセット アブレーションを使用して、数論と測度理論に関するこれまで形式化されていなかった 2 つの数学論文のケース スタディを通じてこの疑問を研究します。タスクの結果、API 呼び出し、入力トークン、出力トークンをレポートします。 Kimi2.6 では、完全なワークフローは両方のドキュメント レベルのプロジェクトを 2000 コールの予算内で完了しますが、キューなしのバリアントは予算制限に達します。 GPT5.5 では、すべてのドキュメント レベルのバリアントが完了し、完全なワークフローの入力トークン コストが両方のソースで最低または同最低になります。補完的なキャリブレーションとして、LeanFlow は RLM25 の PFR スライスで 75.7% BEq+ に達し、GPT5.5 の実行で 5 つの ICML 2026 AI for Math TCS チャレンジ プロジェクトをすべて解決しました。
原文 (English)
LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization
We present and evaluate LeanFlow, an LLM agent system specialized for translating mathematical papers into buildable Lean projects. Recent verifier-in-the-loop systems show that large formal artifacts can be produced, but it remains unclear which runtime mechanisms affect completion, auditability, or efficiency in document-to-project formalization. We study this question through case studies on two previously unformalized mathematical papers in number theory and measure theory, using model, proof-workflow, and toolset ablations with Kimi2.6 and GPT5.5; we report task outcome, API calls, input tokens, and output tokens. With Kimi2.6, the full workflow completes both document-level projects within the 2000-call budget, while no-queue variants reach the budget limit; with GPT5.5, all document-level variants complete, and the full workflow has the lowest or tied-lowest input-token cost on both sources. As complementary calibration, LeanFlow reaches 75.7% BEq+ on the PFR slice of RLM25 and solves all five ICML 2026 AI for Math TCS challenge projects in our GPT5.5 runs.
ハイパーグラフベースの RAG の最適化: より良いファクト抽出とチャンク取得に向けて
GraphRAG は知識をグラフとして構造化することでより深い推論を可能にしますが、n 値の事実には苦労します。 HyperGraphRAG は、よりリッチなセマンティクスのためにハイパーグラフを使用し、精度を向上させますが、エラーが発生しやすい LLM 抽出と非効率な標準チャンク取得に依存しています。私たちは、自己一貫性プロンプトを使用して抽出を改善し、ハイパーグラフ上のパーソナライズされた PageRank アルゴリズムを使用してチャンクの取得を強化することで、この問題に対処します。
原文 (English)
Optimizing Hypergraph-Based RAG: Toward Better Fact Extraction and Chunk Retrieval
GraphRAG enables deeper reasoning by structuring knowledge as graphs but struggles with n-ary facts. HyperGraphRAG uses hypergraphs for richer semantics, improving accuracy, yet relies on error-prone LLM extraction and inefficient standard chunk retrieval. We address this by employing self-consistency prompting to improve the extraction, and Personalized PageRank algorithm over hypergraph to enhance chunk retrieval.
MiniCache: 効率的な LLM 推論のための小規模モデル インターフェイスを使用した再利用可能なプログラム キャッシュ
大規模言語モデル (LLM) は、プログラム支援推論、エージェントによる意思決定、構造化タスクの実行にますます使用されていますが、これらのアプリケーションでは多くの場合、高い推論コストが発生します。私たちは、Program-of-Thought (PoT) プログラムをパラメータ化されたキャッシュ オブジェクトに変換し、構造的に類似したリクエスト全体で再利用可能な計算を可能にする、再利用可能なプログラム キャッシュ フレームワークである MiniCache を紹介します。 MiniCache は、キャッシュ ヒット リクエストのセマンティック変数抽出とターゲット LLM 生成中の投機的ドラフトに同じ小さなモデルを再利用し、タスクの品質を維持しながら高価なターゲット LLM の呼び出しを削減します。ショッピング スタイルのリクエスト データセット、WebShop、Formula、および CodeTAT-QA に関する実験では、MiniCache が推論レイテンシー、キャッシュの再利用、精度の間のトレードオフを改善し、並列処理下で最大 3.1 倍の低いレイテンシーと 2.8 倍の高いスループットを達成することを実証しています。これらの結果は、小さなモデルが、大きなモデルの代替としてではなく、信頼性が高く効率的な再利用可能なプログラム キャッシュを可能にする軽量のインターフェイス モデルとして最も効果的であることを示しています。
原文 (English)
MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference
Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but these applications often incur high inference cost. We present MiniCache, a reusable program caching framework that transforms Program-of-Thought (PoT) programs into parameterized cache objects, enabling reusable computation across structurally similar requests. MiniCache reuses the same small model for semantic variable extraction on cache-hit requests and speculative drafting during target-LLM generation, reducing expensive target-LLM invocations while preserving task quality. Experiments on shopping-style request datasets, WebShop, Formula, and CodeTAT-QA demonstrate that MiniCache improves the trade-off between inference latency, cache reuse, and accuracy, achieving up to 3.1x lower latency and 2.8x higher throughput under parallel serving. These results show that small models are most effective not as replacements for large models, but as lightweight interface models that enable reliable and efficient reusable program caching.
Telco-GAIA: 通信ドメインのエージェント向けのバイリンガル ベンチマーク
Telco-GAIA は、実際の電気通信事業者のデータに基づいてツールを使用するエージェントを評価するための、バイリンガルでマルチモーダルなベンチマークです。 Telco-GAIA は、英語とアラビア語で人間が検証した 100 の質問応答タスクで構成されており、それぞれのタスクで 3 つの異種ソース (静的 Web サイトのスナップショット (HTML、画像、リンクされた PDF)、合成リレーショナル SQL データベース、テキスト、画像、および表形式のモダリティにわたる外部 Web アーカイブ) にわたるマルチホップ推論 (平均 4.2 ホップ) が必要です。このベンチマークはサンドボックス化された Docker 環境として提供され、正規化された正確な文字列一致によってスコア付けされるため、評価は客観的で決定的であり、LLM-as-a-Judge なしで時間の経過とともに再現可能になります。 12 の商用およびオープン LLM にわたって専用のリファレンス エージェントを評価したところ、Telco-GAIA は困難であることがわかりました。最も強力なモデルでもタスクの 71% しか解決できません。適度なコスト予算の下では、これは約 40% に下がり、視覚に基づいたカテゴリは依然として最も弱く、バックエンドの平均スコアは 30% を下回っており、ドキュメントと画像の理解にはかなりの余地が残されています。 Telco-GAIA は、エンタープライズ エージェント向けの厳密で再現可能なテストベッドと、クローズド ドメインのベンチマークを構築するためのテンプレートを提供します。
原文 (English)
Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain
We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human-verified question-answering tasks, in English and Arabic, that each demand multi-hop reasoning (4.2 hops on average) over three heterogeneous sources: a static website snapshot (HTML, images, and linked PDFs), a synthetic relational SQL database, and external web archives, spanning text, image, and tabular modalities. The benchmark is delivered as a sandboxed Docker environment and scored by normalized exact string matching, making evaluation objective, deterministic, and reproducible over time without any LLM-as-a-Judge. Evaluating a purpose-built reference agent across twelve commercial and open LLMs, we find Telco-GAIA challenging: even the strongest model solves only 71% of tasks; under a moderate cost budget, this falls to about 40%, and the visually grounded categories remain the weakest, where the average backend scores below 30%, leaving substantial headroom in document and image understanding. Telco-GAIA offers a rigorous, reproducible testbed for enterprise agents and a template for constructing closed-domain benchmarks.
SiGMA: マルチモーダル連続命令チューニングのための符号ガイド付きマージと適応
マルチモーダル継続命令チューニング (MCIT) は、マルチモーダル大規模言語モデル (MLLM) を適応させて一連の下流タスクを進化させるために重要です。従来の方法は主に専門家の混合または拡張マージアプローチを利用しており、主に壊滅的な忘却に焦点を当てていましたが、依然として推論中に負の干渉に悩まされており、新たに学習した更新により有用な事前知識が上書きされ、全体的なパフォーマンスが低下します。これに対処するために、我々は、2 つのコンポーネントによる負の干渉を軽減するシンプルかつ効果的なフレームワークである SiGMA (Sign Guided Merging and Adaptation) を提案します。トレーニング中の符号ガイド付き適応調整と推論時の符号ガイド付きマージです。標識に基づく適応チューニングにより、過去の知識との衝突が軽減され、ドリフトを最小限に抑えて現在のタスクを学習し、深刻な物忘れを軽減します。符号ガイド付きマージは、重要なパラメーターを選択的にスケーリングして、有用なタスク固有の知識を保存および増幅することで、統合をさらに改善します。 UCIT および DCL ベンチマークの実験では、SiGMA が負の干渉を大幅に低減し、最先端の MCIT 手法よりも優れたパフォーマンスを発揮することが示されています。私たちのコードは SiGMA で入手できます。
原文 (English)
SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning
Multimodal Continual Instruction Tuning (MCIT) is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving a sequence of downstream tasks. Prior methods mostly utilize Mixture of Experts or expansion merge approach, primarily focusing on catastrophic forgetting, yet they still suffer from negative interference during inference, where newly learned updates overwrite useful prior knowledge and degrade overall performance. To address this, we propose SiGMA (Sign Guided Merging and Adaptation), a simple yet effective framework that mitigates negative interference with two components: sign guided adaptive tuning during training and sign guided merging at inference. Sign guided adaptive tuning reduces collisions with past knowledge and learns the current task with minimal drift, mitigating severe forgetting. Sign guided merging further improves consolidation by selectively scaling salient parameters to preserve and amplify useful task specific knowledge. Experiments on UCIT and DCL benchmarks show that SiGMA significantly reduces negative interference and outperforms state of the art MCIT methods. Our code is available at SiGMA.
一貫性のない人間のフィードバックによる信頼性を意識した LLM 調整
人間のフィードバックからの強化学習 (RLHF) は、大規模言語モデル (LLM) を人間の好みに合わせるために重要です。ただし、人間の注釈に固有の矛盾や主観性によって、その有効性が損なわれることがよくあります。 Direct Preference Optimization (DPO) などの既存のプリファレンス最適化フレームワークは通常、アノテーターの不一致が多いあいまいなペアを全会一致のペアと同様に扱い、一貫性のない監視信号にモデルを過剰適合させ、次善の調整を引き起こします。この研究では、一貫性のない人間によるフィードバックの影響を軽減するために設計された堅牢なフレームワークである、信頼性に基づく優先設定の最適化 (RGPO) を提案します。 RGPO は、アノテーターの信頼性を推定し、ノイズの多い人間のフィードバックから潜在的なグラウンド トゥルース ラベルを推測して、堅牢な好みを特定します。さらに、アノテーションのコンセンサスレベルに基づいてトレーニング目標を動的に調整する信頼性を意識した一貫性の最適化を導入し、モデルがコンセンサスの高い監視信号を優先するようにします。 LLM アライメント ベンチマークに関する広範な実験により、RGPO がトレーニング データの不一致とノイズを効果的に低減し、広く採用されている RLHF ベースラインと比較して優れたパフォーマンスを達成できることが実証されました。コードと構成は https://github.com/GenieHuang/RGPO で入手できます。
原文 (English)
Reliability-Aware LLM Alignment from Inconsistent Human Feedback
Reinforcement Learning from Human Feedback (RLHF) is critical for aligning Large Language Models (LLMs) with human preferences. However, its efficacy is often compromised by the inherent inconsistency and subjectivity of human annotations. Existing preference optimization frameworks, such as Direct Preference Optimization (DPO), typically treat ambiguous pairs with high annotator disagreement identically to those with unanimous consensus, forcing models to overfit to inconsistent supervision signals and leading to suboptimal alignment. In this work, we propose Reliability-Guided Preference Optimization (RGPO), a robust framework designed to mitigate the impact of inconsistent human feedback. RGPO estimates annotator reliability and infers latent ground truth labels from noisy human feedback to identify robust preferences. Furthermore, we introduce a reliability-aware consistency optimization that dynamically modulates the training objective based on the consensus level of annotations, ensuring the model prioritizes high-consensus supervision signals. Extensive experiments on LLM alignment benchmarks demonstrate that RGPO effectively reduces inconsistency and noise in training data and achieves superior performance compared to widely adopted RLHF baselines. Our code and configurations are available at https://github.com/GenieHuang/RGPO.
CANN ベンチ: 実際の NPU およびアルゴリズム制限に対するエージェント生成カーネルのベンチマーク
AI エージェントは、さまざまなハードウェア プラットフォーム上で低レベルのオペレーター カーネルを作成、コンパイルし、繰り返し最適化できるようになりました。しかし、既存のベンチマークはほぼ CUDA と Triton のみに焦点を当てており、ハードウェア エコシステムには共通の評価ベースラインがなく、あまり公開されていないプログラミング モデルが残されています。ファーウェイの Ascend NPU で AI が生成したオペレーター コードのオープン ベンチマークである CANN Bench を紹介します。現在のリリースでは、FP16、BF16、FP32、INT8 精度フォーマットにわたる、単純な要素ごとのプリミティブから MoE ディスパッチおよび FlashAttendant カーネルまで、4 つの難易度に編成された 53 のオペレーターと 1,060 のテスト ケースがカバーされています。評価には、コンパイル、機能の正しさ、パフォーマンスを独立した軸として扱う \textbf{3 次元加重複合スコア} が採用され、カーネル生成エージェントに原則に基づいた報酬シグナルが提供されます。パフォーマンスは、すぐに使用できる PyTorch-on-Ascend ベースラインと、実際の NPU ハードウェアでのケースごとの分析的なハードウェア アンカー パフォーマンス (HAP) 制限に基づいてグレード付けされ、スコアが測定アーティファクトではなく真の最適化ヘッドルームを反映していることが保証されます。評価ハーネスは、報酬のハッキングに根本から対抗するように設計されています。 CANN Bench は、公式 CANN リポジトリ内でバージョン管理されており、長期的なコミュニティの共同構築向けに設計されており、Ascend エコシステムに定量的で再現可能で持続的に維持される AI オペレーター オーサリング機能の基準を提供します。
原文 (English)
CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits
AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchmarks, however, focus almost exclusively on CUDA and Triton, leaving hardware ecosystems with less-exposed programming models without a common evaluation baseline. We present CANN Bench, an open benchmark for AI-generated operator code on Huawei's Ascend NPU. The current release covers 53 operators and 1060 test cases organized into four difficulty tiers -- from simple elementwise primitives to MoE dispatch and FlashAttention kernels -- spanning FP16, BF16, FP32, and INT8 precision formats. Evaluation adopts a \textbf{three-dimensional weighted composite score} that treats compilation, functional correctness, and performance as independent axes, providing a principled reward signal for kernel-generation agents. Performance is graded against an out-of-the-box PyTorch-on-Ascend baseline and an analytical per-case Hardware-Anchored Performance (HAP) limit on real NPU hardware, ensuring scores reflect genuine optimization headroom rather than measurement artifacts. The evaluation harness is designed to resist reward hacking from the ground up. CANN Bench is versioned within the official CANN repository and is designed for long-term community co-construction, providing the Ascend ecosystem with a quantitative, reproducible, and sustainably maintained yardstick for AI operator-authoring capability.
数学的問題解決のための大規模言語モデルにおける実行可能な推論制約下での表現の堅牢性
大規模言語モデル (LLM) は、数学的問題解決に関してますます評価されていますが、これまでの研究では、表現的に同等の定式化が交換可能なものとして扱われ、推論エラーとインターフェイスの障害が混同されることがよくありました。この論文では、ストーリー問題、単語方程式、記号方程式、同型言い換えなど、同じ根本的な問題の表面表現を体系的に変化させることによって、LLM ベースの数学的問題解決における表現のロバスト性を調査します。数学的に等価な問題の厳選されたデータセットを使用して、直接応答生成条件の下で 5 つの現代の LLM を評価します。私たちは、相当な表現の感度を発見しました。モデルは、同等の定式化間で正確さを頻繁に変更し、ストーリー、記号、および単語方程式のバリアント全体で自明ではないフリップ率を伴います。また、同型再定式化の下での系統的な回帰も観察し、数学的構造が保存されているにもかかわらず、言い換えレベルの微妙な変更でさえパフォーマンスが低下する可能性があることを示しています。次に、モデルが推論を、検証のためにローカルで実行される実行可能な Python コードとして外部化するコード拡張条件を評価します。このインターフェイスは、直接的なプロンプトの下ではパフォーマンスが低い一部のモデルにおける強力な潜在的な推論能力を明らかにしますが、堅牢性を一律に向上させるわけではありません。代わりに、障害は、不透明な推論エラーからプロトコル違反や実行障害に至るまで、インタラクション層全体に移行します。実行可能な推論が成功した場合でも、表現の感度は維持されることがよくあります。全体として、私たちの結果は、推論足場によって表現の脆弱性が解消されるわけではなく、正確性、信頼性、待ち時間、コストの間の新たなトレードオフを明らかにすることを示しています。私たちは、特に AI 支援による問題解決システムの場合、LLM の評価と展開において、表現を第一級のインターフェイス設計変数として扱うべきであると主張します。
原文 (English)
Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equivalent formulations as interchangeable and conflates reasoning errors with interface failures. This paper investigates representation robustness in LLM-based mathematical problem solving by systematically varying surface representations of the same underlying problems, including story problems, word-equations, symbolic equations, and isomorphic paraphrases. Using a curated dataset of mathematically equivalent problems, we evaluate five contemporary LLMs under a direct answer generation condition. We find substantial representational sensitivity: models frequently change correctness across equivalent formulations, with nontrivial flip rates across story, symbolic, and word-equation variants. We also observe systematic regressions under isomorphic reformulations, showing that even subtle paraphrase-level changes can degrade performance despite preserved mathematical structure. We then evaluate a code-augmented condition in which models externalize reasoning as executable Python code that is run locally for validation. This interface reveals strong latent reasoning capability in some models that perform poorly under direct prompting, but it does not uniformly improve robustness. Instead, failures shift across interaction layers, from opaque reasoning errors to protocol violations and execution failures. Even when executable reasoning succeeds, representation sensitivity often persists. Overall, our results show that reasoning scaffolds do not eliminate representational brittleness, but expose new tradeoffs among correctness, reliability, latency, and cost. We argue that representation should be treated as a first-class interface design variable in LLM evaluation and deployment, especially for AI-assisted problem-solving systems.
アテンションの低下、関数トークンのアンカリング、および大規模言語モデルにおけるアテンションベースの介入の限界
位置をまたぐ注意力の平均低下は、トランスの解釈可能性において広く報告されていますが、それが文脈検索を因果的に制限するかどうかはまだテストされていません。 GPT-2、LLaMA-3.2-1B/3B、OPT-1.3B、および distilgpt2 にわたる 6 つの調整された実験を紹介します。まず、短期(5 ~ 100 トークン)の注意力の低下を特徴付け、アーキテクチャごとに明確な層ごとのエントロピー シグネチャを備え、その速度が深さと逆相関する普遍的な指数関数的その後プラトー パターンを発見しました。関数トークンのアンカリングはアーキテクチャに依存することを証明します。OPT-1.3B (絶対位置エンコーディング) は距離依存の前置詞特異性を示し、GPT-2 は均一で非特異的な依存性を示し、LLaMA (RoPE) は長距離での反転を示します。句の境界に戦略的にコンマを挿入すると、40 ~ 80 のトークン範囲での予測の劣化が因果的に軽減され、その利点はトークン密度ではなく構文境界の位置合わせに関係します。次に、メカニズムを因果的にテストします。Relay-Aware Attendant (RAA) は、注意ロジットを関数トークンの位置に偏らせますが、検証的に注意量を 16 ~ 24% 増加させますが、GPT-2 および LLaMA-1B では無効な効果、LLaMA-3B では予備的な害をもたらし、OPT-1.3B ではほぼゼロになる混合効果をもたらします。多事実検索プローブはさらに、劣化率がモデル全体の検索精度を予測しないことを示しています。私たちは、平均注意力の低下は、規範的なものではなく主に記述的なものであると結論付けています。関数トークンは、受け取る注意を通じてではなく、隠れ状態が計算する内容を通じて寄与します。これは、解釈可能性の方法論や、KV キャッシュの削除などの注意スコアに基づく推論の最適化に影響を及ぼします。
原文 (English)
Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval remains untested. We present six coordinated experiments across GPT-2, LLaMA-3.2-1B/3B, OPT-1.3B, and distilgpt2. We first characterise short-term (5-100 token) attention degradation, finding a universal exponential-then-plateau pattern whose rate is inversely correlated with depth, with distinct layer-wise entropy signatures per architecture. Function token anchoring proves architecture-dependent: OPT-1.3B (absolute positional encoding) shows distance-dependent preposition specificity, GPT-2 shows uniform non-specific dependence, and LLaMA (RoPE) shows reversal at long distances. Strategic comma insertion at clause boundaries causally reduces prediction degradation in the 40-80 token range, with the benefit tied to syntactic boundary alignment rather than token density. We then test the mechanism causally: Relay-Aware Attention (RAA), which biases attention logits toward function token positions, verifiably increases attention mass by 16-24% yet yields null effects on GPT-2 and LLaMA-1B, preliminary harm on LLaMA-3B, and a mixed effect on OPT-1.3B that nets to approximately zero. Multi-fact retrieval probes further show that degradation rate does not predict retrieval accuracy across models. We conclude that mean attention degradation is largely descriptive rather than prescriptive: function tokens contribute through what their hidden states compute, not through the attention they receive -- with implications for interpretability methodology and attention-score-based inference optimisations such as KV-cache eviction.
GPT-5.5 Pro を使用した $\mathbb R$ に対する積和予想の自律的反証
OpenAI による最近の Erd\H{o} の単位距離予想の反証は、数学における AI の画期的な出来事となりました。また、これは別の画期的な発見をもたらしました。それは、$\mathbb R$ に関する Erd\H{o}s-Szemer\'edi sum-product 予想に対する人による反証です。このペーパーでは、GPT-5.5 Pro 上に構築されたシンプルなエージェントを紹介します。問題にとらわれない 3 段階のプロンプト パイプライン (証明計画の提案、証明の構築、レビュー) を使用して、エージェントは 8 つの独立した試行のうち 7 つで $\mathbb R$ に対して積和予想が偽であるという正しい証明を自律的に生成しました。残りの裁判では、主張に未解決のギャップがあることが判明した。 7 つの証明は多様です。既存の単位ベースの構成に近いものもあれば、代数整数の $L^p$ 型領域を使用して単位を回避するものもあります。システムは、試行ごとに平均 132.4k の推論トークンを使用しました。コード、中間出力、生成されたプルーフを公開し、自律的なプルーフ生成における再現可能でデータ汚染のないケーススタディを提供します。
原文 (English)
Autonomous disproofs of the sum-product conjecture over $\mathbb R$ with GPT-5.5 Pro
OpenAI's recent disproof of the Erd\H{o}s unit distance conjecture marked a milestone for AI in mathematics. It also inspired another breakthrough: a human disproof of the Erd\H{o}s--Szemer\'edi sum-product conjecture over $\mathbb R$. In this paper, we present a simple agent built on GPT-5.5 Pro. Using a problem-agnostic, three-stage prompting pipeline -- proof-plan proposal, proof construction, and review -- the agent autonomously generated correct proofs that the sum-product conjecture is false over $\mathbb R$ in 7 of 8 independent trials; in the remaining trial, it identified an unresolved gap in its argument. The seven proofs are diverse: some are close to existing unit-based constructions, while others avoid units by using $L^p$-type regions of algebraic integers. The system used an average of 132.4k reasoning tokens per trial. We release the code, intermediate outputs, and generated proofs, providing a reproducible, data-contamination-free case study in autonomous proof generation.
ConfidenceBench: 大規模言語モデルの信頼度調整の評価
大規模言語モデル (LLM) は、流暢ではあっても不正確な回答がコストのかかる環境に導入されることが増えています。これらの設定では、精度だけでは不十分です。モデルは、いつ間違っている可能性があるかを認識する必要もあります。我々は、真実の確率レポートを奨励する適切なスコアリング ルールであるブライアー スコアを使用して、15 のフロンティア LLM における言語化された信頼推定値を評価する校正ベンチマークである ConfidenceBench を紹介します。プロンプトを介して自信を引き出し、モデルのロジットへのアクセスを必要とせず、フレームワークをクローズドソースとオープンソースの両方のシステムに適用できるようにします。このベンチマークは、空間推論、高精度数学、単語検索、未知の質問の 4 つのカテゴリにわたる 200 のプライベートな多肢選択問題で構成されています。 3 回の独立した実行全体で、Claude Opus 4.6 と Gemini 3.1 Pro Preview が最も低い Brier スコアを達成し、どちらも 0.103 と報告されました。どちらもキャリブレーションされたランダムのベースラインである 0.1875 を大幅に上回っていますが、Gemini 3.1 Flash-Lite のスコアは 0.367 であり、重大なキャリブレーションの誤りを示しています。精度とキャリブレーションは、モデル ファミリによって大幅に異なります。最も正確なモデルが最もよくキャリブレーションされているわけではありません。また、いくつかのモデルは、妥当な精度にもかかわらず、キャリブレーション済みのランダムな Brier ベースラインよりもパフォーマンスが悪いものもあります。これらの結果は、言語化された信頼度キャリブレーションが、LLM の信頼性の明確かつ実際的に重要な軸であり、標準的な精度に基づく評価を補完するものであることを示しています。
原文 (English)
ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone is insufficient: models must also know when they are likely to be wrong. We present ConfidenceBench, a calibration benchmark that evaluates verbalized confidence estimates in 15 frontier LLMs using the Brier score, a proper scoring rule that incentivises truthful probability reporting. Confidence is elicited via prompting, requiring no access to model logits and making the framework applicable to both closed-source and open-source systems. The benchmark comprises 200 private multiple-choice questions across four categories: spatial reasoning, high-precision mathematics, word lookup, and unknowable questions. Across three independent runs, Claude Opus 4.6 and Gemini 3.1 Pro Preview achieve the lowest Brier scores, both reported as 0.103. Both substantially outperform the calibrated-random baseline of 0.1875, while Gemini 3.1 Flash-Lite scores 0.367, indicating severe miscalibration. Accuracy and calibration diverge substantially across model families: the most accurate model is not the best-calibrated, and several models perform worse than the calibrated-random Brier baseline despite reasonable accuracy. These results show that verbalized confidence calibration is a distinct and practically important axis of LLM reliability, complementary to standard accuracy-based evaluation.
エージェント的科学合成における引用の忠実性の評価と保護
OpenScholar や PaperQA2 などのエージェント LLM システムは、科学文献を読み取り、引用された回答を返します。これらのシステムとそのベンチマークの両方は、固定アトリビューション モデルまたは人間による採点機能を使用して、それらの引用が保持されているかどうかをすでにチェックしています。どちらもそのチェック自体の信頼性を監査しません。私たちは、それが信頼できないこと、そしてこれが重要であることを示しています。同一のエージェント出力では、測定されたサポートされていない引用率は検証者の厳密さのみに応じて約 3% から約 18% の範囲であり、検証者はどの引用がサポートされているかについては一致していますが、どれにフラグを立てるかについては一致していません (否定的特異的一致 0.27 ~ 0.30)。そのため、単一のフラグセットは信頼できず、名前付き検証者とプロトコルがなければ論文間の比較は無効です。私たちは、この動作を測定可能かつ制限付きにする、ゴールドアンカー評価プロトコルと展開可能なガードを紹介します。このプロトコルは検証者を検証し、再帰属を測定し、別のモデルの判定ではなく人間の金に対して保証を調整します。ベリファイアはコストに基づいて選択された交換可能な手段であり (サポートされるクラスの 0.94 を思い出してください、保留されています)、再帰属は決定論的な BM25 が最良のオープン ジェネレーターと一致するコモディティ ステップです。ガードは、選択されたフラグ付けルールをすり抜けて本当にサポートされていない引用に分布フリーの有限サンプル境界を配置する分割コンフォーマル層を追加し、結論の正しさではなく捕捉率を保証します。境界は保留された金に保持されており、事前の等角事実研究ではテストされていない、具体的な再校正レシピを使用して、展開への移行、校正否定の困難を支配する条件を特定し、定量化します。公開ベンチマーク (SciFact、QASA、PubMedQA) で 4 つのオープン 27-35B モデルと 3 つのエージェント パイプラインにわたって検証され、すべてのヘッドライン番号、プロトコル、およびオープン シングル GPU キットとしてのガードシップの信頼区間が検証されています。
原文 (English)
Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis
Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their benchmarks already check whether those citations hold, with a fixed attribution model or human graders. Neither audits the reliability of that check itself. We show it is not reliable, and that this matters. On identical agent outputs the measured unsupported-citation rate ranges from about 3% to about 18% depending only on the verifier's strictness, and although verifiers agree on which citations are supported, they disagree on which to flag (negative-specific agreement 0.27 to 0.30), so no single flag set is trustworthy and cross-paper comparison is invalid without a named verifier and protocol. We present a gold-anchored evaluation protocol and a deployable guard that make this behavior measurable and bounded. The protocol validates the verifier, measures re-attribution, and calibrates a guarantee against human gold rather than another model's verdict; the verifier is a swappable instrument chosen on cost (recall 0.94 on the supported class, held out), and re-attribution is a commodity step where a deterministic BM25 matches the best open generator. The guard adds a split-conformal layer placing a distribution-free, finite-sample bound on truly unsupported citations that slip past a chosen flagging rule, a guarantee on catch rate rather than conclusion correctness. The bound holds on held-out gold, and we identify and quantify the condition governing its transfer to deployment, calibration-negative difficulty, with a concrete recalibration recipe, left untested by prior conformal-factuality work. Validated across four open 27-35B models and three agentic pipelines on public benchmarks (SciFact, QASA, PubMedQA), with confidence intervals on every headline number, the protocol and guard ship as an open single-GPU kit.
PromptPack: オンライン レコメンデーションのための LLM アノテーション エージェントのスケーリング
オンライン レコメンデーション プラットフォームでは、広告クリエイティブから構造化された特徴を抽出するためにラージ言語モデル (LLM) を使用するケースが増えています。シングルコールの LLM アノテーション エージェントを導入すると、実際の運用環境でクリックスルー レート (CTR) が大幅に向上しますが、クリエイティブごとのプロンプトを拡張するには法外なコストがかかります。 The redundant system instructions sent in every request account for 94% of billed input tokens.このコストのボトルネックを解消するために、スケーラブルで高スループットの LLM アノテーション エージェントである PromptPack を導入します。 PromptPack は、共有システム プロンプト、厳密な XML 構造エンベロープ、および出力修正レイヤーを組み合わせたインコンテキスト バッチ処理によってこのスケールを実現し、複数のクリエイティブにわたって同時に決定論的でパイプライン対応の特徴抽出を保証します。ダウンストリームのロジスティック回帰ランカーを使用したオフライン検索ベンチマークによって PromptPack を評価します。エージェントの動作を詳細にプロファイリングするために、AUC を測定し、生成された特徴の信号品質を捕捉する新しいメトリクスである体積加重絶対リフト (VWAL) を導入します。バッチ サイズ 20 の PromptPack は、ライブのバッチ処理されていない運用ベースラインと比較して、AUC を完全に維持しながら LLM コストを 89% 削減し、スループットを 2.5 倍高速化します。
原文 (English)
PromptPack: Scaling LLM Annotation Agents for Online Recommendation
Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives. While deploying a single-call LLM annotation agent yields significant Click-Through Rate (CTR) improvements in our live production environment, per-creative prompting is prohibitively expensive to scale. The redundant system instructions sent in every request account for 94% of billed input tokens. To break this cost bottleneck, we introduce PromptPack, a scalable, high-throughput LLM annotation agent. PromptPack achieves this scale via in-context batching, combining a shared system prompt, a strict XML structural envelope, and an output correction layer to ensure deterministic, pipeline-ready feature extraction across multiple creatives simultaneously. We evaluate PromptPack via an offline retrieval benchmark using a downstream logistic-regression ranker. To deeply profile the agent's behavior, we measure AUC and introduce Volume-Weighted Absolute Lift (VWAL), a novel metric capturing the signal quality of the generated features. Compared to our live, unbatched production baseline, PromptPack at batch size 20 cuts our LLM costs by 89% and accelerates throughput by 2.5x while fully preserving AUC.
DynamicMCPBench: ライブ MCP サーバー上の LLM エージェント向けの、トレースに基づいた効果スコアリングのベンチマーク
大規模言語モデル (LLM) エージェントは、モデル コンテキスト プロトコル (MCP) サーバー上にデプロイされることが増えていますが、それらの評価に使用されるベンチマークは、最終的な答えまたはツールの固定の「グラウンドトゥルース」リストをスコアリングします。基盤となるデータがライブでステートフルになると、どちらも脆弱になります。固定データセットではなく再利用可能なフレームワークである DynamicMCPBench を紹介します。実践者は、独自の MCP サーバー上でこれを実行して、独自のタスクでモデルをテストしたり、サーバーを自動的に収集して、エージェント タスクを解決するモデルの一般的な能力を測定したりできます。サーバーとモデルのセットが与えられると、現実的な目標を生成し、それぞれをライブで追跡して成功の軌跡を記録し、その軌跡を経路に依存しない効果のチェックポイントに抽出し、最終的な答えではなく、それらの効果を再現するかどうかでエージェントをスコア付けします。フレームワークが何を明らかにするかを示すために、フレームワークを大規模に実行します。121 台のサーバー上で 24 のモデルと 750 のタスクが 15 のタスク カテゴリ (各 50) に均等に分散されており、各カテゴリは、生成された質問のツール使用に関する個別の課題を対象としています。各タスクは pass^3 によってスコア付けされます。3 つの独立した試行がすべて成功した場合にのみ、解決済みとしてカウントされます。最も強力なエージェントでもタスクの約半分しか解決できず、タスクの 31% はまったくモデルによって解決されず、必要なツール チェーンが長くなるにつれて精度が低下します (最短のチェーンの 39% から最長のチェーンの 13%)。人間による検証研究により、自動スコアリングが信頼できることが確認されています (確率補正後の一致率は 0.76)。したがって、DynamicMCPBench は、ベンチマークの構築を、実践者が独自のサーバーやモデルで再実行できるものに変えますが、同時に、現在のエージェントが長くて複数ステップのエージェント タスクを処理できないという一貫した能力の欠如を明らかにします。
原文 (English)
DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers
Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful. We present DynamicMCPBench, a reusable framework rather than a fixed dataset. A practitioner can run it on their own MCP servers to test models on their own tasks, or let it collect servers automatically to measure a model's general ability to solve agentic tasks. Given the servers and any set of models, it generates realistic goals, pursues each one live to record a successful trajectory, distills that trajectory into path-agnostic effect checkpoints, and scores an agent on whether it reproduces those effects, never on the final answer. To show what the framework reveals, we run it at scale: 24 models over 121 servers and 750 tasks spread evenly over 15 task categories (50 each), where each category targets a distinct tool-use challenge of the generated questions. Each task is scored by pass^3: it counts as solved only if all three independent attempts succeed. Even the strongest agents solve only about half of the tasks, 31% of tasks are solved by no model at all, and accuracy collapses as the required tool chain grows longer (from 39% on the shortest chains to 13% on the longest). A human validation study confirms the automatic scoring is reliable (chance-corrected agreement of 0.76). DynamicMCPBench thus turns benchmark construction into something practitioners can rerun on their own servers and models, while exposing a consistent inability of current agents to handle long, multi-step agentic tasks.
AppWorld-UL: ツール使用のための多様なエージェントとユーザーのインタラクションのベンチマーク
食料品の注文などの日常的なデジタルタスクに対処するツール使用エージェントは、アプリケーションを操作するだけでなく、説明の質問をしたり、確認を求めたり、指示が実行不可能な場合にユーザーに通知したりするなど、ユーザーと対話する必要もあります。ただし、エージェントとユーザーの対話を評価するための現在のベンチマークは、そのような対話の多様性を捉えていません。さらに、これらは、状態を変更しない API がほとんどない小規模な環境で動作します。このギャップに対処するために、私たちは、多様なエージェントとユーザーの対話を必要とする 516 の困難なタスクの「ユーザーインザループ」ベンチマークである AppWorld-UL を導入します。 Amazon や Spotify などの 9 つの人気のあるシミュレート アプリを含む AppWorld フレームワークを基盤として、元のタスクを体系的に変更して、さまざまなタイプのエージェントとユーザーの対話を必要とする曖昧さと制約を導入します。ユーザーの行動は、慎重に設計された知識境界で応答するよう促される LLM によってシミュレートされ、以前の研究で使用されていた制約のない、または過度に厳格な代替手段よりも信頼性の高いシミュレーションを提供します。私たちの評価では、最先端の LLM、Claude Opus 4.7 は、AppWorld-UL では 48.6% の成功率しか達成しておらず、より困難な構成サブセットでは 35.7% しか成功していないことが明らかになりました。より厳密なシナリオレベルの指標では、構成タスクのパフォーマンスはわずか 21.3% に低下します。私たちの分析により、成功には正しいユーザー インタラクションが重要であることがわかりました。これは、ベンチマークの難しさと、ユーザーインザループのツール使用エージェントに関する研究を前進させる可能性があることを示しています。
原文 (English)
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for confirmation, and inform the user when the instruction is infeasible. However, current benchmarks for evaluating agent-user interactions do not capture the diversity of such interactions. Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a ``user-in-the-loop'' benchmark of 516 challenging tasks requiring diverse agent-user interactions. Building upon the AppWorld framework with 9 popular simulated apps like Amazon and Spotify, we systematically modify original tasks to introduce ambiguities and constraints that necessitate various types of agent-user interaction. User behavior is simulated by an LLM prompted to respond with carefully designed knowledge boundaries, offering more reliable simulation than the unconstrained or overly rigid alternatives used in prior work. Our evaluation reveals that a state-of-the-art LLM, Claude Opus 4.7, achieves only 48.6% success on AppWorld-UL, and only 35.7% on the harder, compositional subset. On the stricter, scenario-level metric, compositional task performance drops to only 21.3%. Our analysis reveals that correct user-interaction is crucial for success. This demonstrates the benchmark's difficulty and its potential to advance research on user-in-the-loop tool-use agents.
StrideDiffusion: 時系列生成のための拡散モデルの加速
拡散モデルは時系列の競合生成手段となっていますが、推論時に多数の連続ノイズ除去ステップが必要となるため、その実用化は制限されています。既存の高速サンプラーは通常、固定または汎用のタイムステップ スケジュールを使用しており、時系列拡散の特有の特性、つまり逆プロセス中に異なるスペクトル バンドが異なる速度で進化するという特性を見落としています。帯域レベルのアクティビティからノイズ除去ストライドを適応的に選択する、トレーニング不要のスペクトル認識サンプラーである StrideDiffusion を紹介します。各ステップで、StrideDiffusion は相対帯域エネルギー、対数パワードリフト、および位相速度を監視して、高周波ダイナミクスがアクティブなままであるかどうか、または軌道が安定した低周波構造によって支配されているかどうかを識別します。その後、急速に変化する帯域がアクティブな場合は細かいステップを実行し、粗い成分のみが残るとより大きなジャンプを実行します。帯域別の安定性解析では、非アクティブな周波数帯域は、決定論的なアフィン逆更新の下でジャンプ サイズに対して線形にのみ変化することが示されており、ステップ サイズの指標としてスペクトル アクティビティの局所的な正当性が得られます。 6 つの無条件時系列生成ベンチマーク全体で、StrideDiffusion は 500/1000 のノイズ除去ステップの代わりに 14 ~ 66 回の関数評価のみを使用し、生成品質を維持または向上させながら最大 18.9 倍の実時間の高速化を達成します。条件付き代入と予測では、同等の予測精度で平均 5 ~ 14 倍の加速を実現します。これらの結果は、スペクトル進化が高速時系列拡散サンプリングのための実用的かつ原理的な信号を提供することを示しています。私たちのコードは https://anonymous.4open.science/r/stridediff-ts で入手できます。
原文 (English)
StrideDiffusion: Accelerating Diffusion Models for Time-series Generation
Diffusion models have become competitive generators for time series, but their practical use is limited by the large number of sequential denoising steps required at inference time. Existing fast samplers typically use fixed or generic timestep schedules, overlooking a distinctive property of time-series diffusion: different spectral bands evolve at different rates during the reverse process. We introduce StrideDiffusion, a training-free spectral-aware sampler that adaptively selects the denoising stride from band-level activity. At each step, StrideDiffusion monitors relative band energy, log-power drift, and phase velocity to identify whether high- frequency dynamics remain active or whether the trajectory is dominated by stable low-frequency structure. It then takes fine steps when rapidly varying bands are active and larger jumps once only coarse components remain. A bandwise stability analysis shows that inactive frequency bands change only linearly with the jump size under deterministic affine reverse updates, providing a local justification for spectral activity as a step-size indicator. Across six unconditional time-series generation benchmarks, StrideDiffusion uses only 14-66 function evaluations instead of 500/1000 denoising steps, achieving up to 18.9x wall-clock speedup while preserving or improving generation quality. On conditional imputation and forecasting, it further delivers 5-14x average acceleration with comparable predictive accuracy. These results show that spectral evolution provides a practical and principled signal for fast time-series diffusion sampling. Our code is available at https://anonymous.4open.science/r/stridediff-ts.
CMI-Mem: CMI拡張強化学習による一般化可能な長期記憶管理に向けて
メモリ マネージャー モデルは、エージェント システムにおいて極めて重要です。既存の方法は主に LLM が判断した合成質問と回答 (QA) のペアに依存しており、メモリの評価はサンプリングされたクエリと下流のリーダーに依存します。この制限に対処するために、下流の QA の正確性と固有の条件付き相互情報 (CMI) を組み合わせたハイブリッド報酬を備えた強化学習 (RL) ベースの軽量メモリ マネージャー モデル \textbf{CMI-Mem} を提案します。 CMI は、サンプリングされた QA クエリを条件付けせずに、新しい会話入力によって提供される情報を現在のメモリ状態と比較して評価します。これにより、QA の基礎を置き換えるのではなく補完します。コードは https://github.com/Wyb0627/CMIMem で入手できます。CMI-Mem-4B モデル チェックポイントは https://www.modelscope.cn/models/wyb0627/CMIMem-4B で入手できます。
原文 (English)
CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning
Memory Manager models are pivotal in agent systems. Existing methods rely predominantly on LLM-judged synthetic question-answer (QA) pairs, making memory valuation dependent on sampled queries and the downstream reader. To address this limitation, we propose \textbf{CMI-Mem}, a reinforcement learning(RL)-based lightweight memory manager model with a hybrid reward that combines downstream QA correctness and intrinsic Conditional Mutual Information (CMI). CMI evaluates the information contributed by new conversational inputs relative to the current memory state without conditioning on a sampled QA query, thereby complementing rather than replacing QA grounding. Our codes are available at: https://github.com/Wyb0627/CMIMem , and the CMI-Mem-4B model checkpoint is available at: https://www.modelscope.cn/models/wyb0627/CMIMem-4B
学習最適化グラフ ニューラル ネットワークによるスマート アーバン NR-V2X ネットワーク向けの AI 主導のマルチホップ リレー選択
信頼性が高く低遅延の NR-V2X 通信は、密集した都市環境におけるスマート モビリティに不可欠です。ただし、路側ユニット (RSU) の密度が限られていること、見通し外の状況が頻繁に発生すること、および非常に動的な車両トポロジにより、多くの Connected and Automated Vehicle (CAV) が安定したシングルホップ接続を維持できないことがよくあります。マルチホップ リレー支援通信はインフラストラクチャのカバレッジを拡張できますが、実際のフロー、容量、接続の制約の下でリアルタイムにリレー リンクを選択することは依然として困難です。混合整数線形計画法 (MILP) は最適なマルチホップ リレーの決定を行いますが、その計算の複雑さはネットワーク密度に応じて大幅に増加するため、リアルタイムの適用性が制限されます。これに対処するために、リアルタイムの NR-V2X リレー選択のためのグラフ ニューラル ネットワーク (GNN) に基づく学習から最適化 (L2O) フレームワークを提案します。車両の通信状態は属性付きグラフとしてモデル化されます。CAV と RSU はノードであり、候補となる無線リンクは伝播を意識した機能で強化されます。オフラインの MILP オラクルは最適な監視を提供し、エッジ対応のグラフ同型ネットワーク (GINE) はほぼ一定の推論レイテンシーでオラクルの決定を近似します。統合された SUMO-GEMV2 シミュレーション パイプラインによって生成された大規模都市データセットの実験では、提案されたアプローチが実行時間を桁違いに短縮しながら、MILP オラクルに匹敵する接続性を達成することが示されました。このフレームワークにより、既存の車両資産を活用し、スマート シティ環境でのスケーラブルなリアルタイム NR-V2X 運用をサポートすることで、都市部の V2X 接続をコスト効率よく強化できます。
原文 (English)
AI-Driven Multi-Hop Relay Selection for Smart Urban NR-V2X Networks via Learning-to-Optimize Graph Neural Networks
Reliable and low-latency NR-V2X communications are essential for smart mobility in dense urban environments. However, limited Road-Side Unit (RSU) density, frequent non-line-of-sight conditions, and highly dynamic vehicular topologies often prevent many Connected and Automated Vehicles (CAVs) from maintaining stable single-hop connectivity. Although multi-hop relay-assisted communication can extend infrastructure coverage, selecting relay links in real time under practical flow, capacity, and connectivity constraints remains challenging. Mixed-Integer Linear Programming (MILP) yields optimal multi-hop relay decisions, but its computational complexity scales sharply with network density, limiting real-time applicability. To address this, we propose a Learning-to-Optimise (L2O) framework based on Graph Neural Networks (GNNs) for real-time NR-V2X relay selection. Vehicular communication states are modeled as attributed graphs, where CAVs and RSUs are nodes and candidate radio links are enriched with propagation-aware features. An offline MILP oracle provides optimal supervision, while an edge-aware Graph Isomorphism Network (GINE) approximates oracle decisions with near-constant inference latency. Experiments on large-scale urban datasets generated by an integrated SUMO--GEMV2 simulation pipeline show that the proposed approach achieves connectivity comparable to that of the MILP oracle while reducing execution time by orders of magnitude. The framework enables cost-effective enhancement of urban V2X connectivity by leveraging existing vehicular assets and supporting scalable, real-time NR-V2X operation in smart city environments.
KeySI: 人間のフィードバックに基づいてテキスト埋め込みを調整するためのインタラクション フレームワーク
大規模なテキスト分析タスクでは、下流分析用にテキスト コーパスを埋め込むために、事前トレーニングされた言語モデルがよく使用されます。ただし、そのようなモデルはドメイン固有のセマンティクスを取得するのに苦労する可能性があり、それらを適応させるには通常、トレーニング パイプラインを実装するための大量のラベル付きデータと技術的専門知識が必要です。最近のアプローチでは、ドキュメント投影における視覚的なインタラクションが人間のフィードバックをモデル調整のためのトレーニング信号としてどのようにキャプチャできるかを実証しています。ただし、これらの方法はドキュメント レベルのフィードバックに基づいて動作するため、効果的なフィードバックを提供するには、ユーザーが個々のドキュメントを開いて評価する必要があります。この論文では、キーワードベースの概念仕様を通じて機能レベルのフィードバックを可能にするインタラクション フレームワークである KeySI を提案します。ユーザーは、抽出されたキーワードを概念を表すグループに編成することによってフィードバックを指定します。KeySI は、それを後のチューニングのための文書レベルの監視に変換します。 KeySI は主要な対話媒体としてキーワードを操作することにより、手動による文書検査とラベル付けの必要性を減らし、埋め込みモデルの適応に対する障壁を下げます。コーパスを与えられた場合に、代表的なキーワードを厳選し、次元削減によってキーワードとドキュメントの埋め込みを視覚化し、キーワード グループの対話型の指定を可能にし、システム フィードバックによる反復的な改良をサポートするプロトタイプの実装を紹介します。当社では、ユーザー調査、使用シナリオ、定量的実験を通じて KeySI を評価し、ユーザーの意図を捉えて埋め込み調整を改善する有効性を実証しています。
原文 (English)
KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback
In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis. However, such models may struggle to capture domain-specific semantics and adapting them typically requires large amounts of labeled data and technical expertise to implement training pipelines. Recent approaches have demonstrated how visual interactions in document projections can capture human feedback as training signals for model tuning. However, these methods operate on document-level feedback, which requires users to open and assess individual documents in order to provide effective feedback. In this paper, we propose KeySI, an interaction framework that enables feature-level feedback through keyword-based concept specification. Users specify feedback by organizing extracted keywords into groups representing concepts, which KeySI translates into document-level supervision for subsequent tuning. By operating on keywords as the primary interaction medium, KeySI reduces the need for manual document inspection and labeling and lowers the barrier to adapting embedding models. We present a prototype implementation that, given a corpus, curates representative keywords, visualizes keywords and document embeddings via dimensionality reduction, allows interactive specification of keyword groups, and supports iterative refinement through system feedback. We evaluate KeySI through a user study, usage scenarios, and quantitative experiments demonstrating its effectiveness in capturing user intent and improving embedding alignment.
WaveformQA: デジタル波形における LLM 時間推論のベンチマーク
大規模言語モデル (LLM) は、コード生成と推論において強力な機能を実証していますが、デジタル波形データに対して時間的推論を実行する能力はほとんど解明されていません。デジタル波形に対する推論は設計検証における重大なボトルネックですが、既存のベンチマークは主にハードウェア記述言語 (HDL) コード生成を評価し、波形を補足的なコンテキストとしてのみ使用します。このペーパーでは、デジタル波形に対する LLM 時間推論を評価するためのオープンソースの質問応答ベンチマークである WaveformQA について説明します。このベンチマークは、さまざまな難易度の 8 つのカテゴリにわたるプログラムで生成されたグラウンド トゥルースを含む 360 の質問で構成されており、これには、複数信号の相関やイベントの順序付けを対象とした質問も含まれます。波形はオープンソースの設計実装から生成され、再現性を確保し、実際のハードウェアの動作に基づいたベンチマークを実現します。フロンティア LLM の評価により、モデルは単純なクエリでは妥当な精度を達成しますが、コンテキスト ウィンドウの制限と、複雑な時間的および複数ステップの質問では推論の困難によりパフォーマンスが低下することが明らかになりました。さらに、イベント時の波形の JSON 表現により、標準化された値変更ダンプ (VCD) 形式と比較して LLM 推論の精度が向上することを示します。オープンソース フレームワークは、新しい質問カテゴリへの拡張と新しい波形ソースのインポートをサポートしているため、研究者は時間推論実験のプロトタイプを迅速に作成できます。
原文 (English)
WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms
Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their ability to perform temporal reasoning over digital waveform data remains largely unexplored. Although reasoning over digital waveforms is a critical bottleneck in design verification, existing benchmarks primarily evaluate hardware description language (HDL) code generation and use waveforms only as supplementary context. This paper presents WaveformQA, an open-source question-answering benchmark for evaluating LLM temporal reasoning over digital waveforms. The benchmark comprises 360 questions with programmatically generated ground truths across eight categories of varying difficulty, including questions targeting multi-signal correlation and event ordering. Waveforms are generated from open-source design implementations, ensuring reproducibility and grounding the benchmark in real hardware behavior. Evaluation of frontier LLMs reveals that while models achieve reasonable accuracy on simple queries, performance degrades due to context window limitations and reasoning difficulties on complex temporal and multi-step questions. In addition, we show that an event-time JSON representation of waveforms improves LLM reasoning accuracy versus the standardized value change dump (VCD) format. The open-source framework supports extending to new question categories and importing new waveform sources, enabling researchers to rapidly prototype temporal reasoning experiments.
NVIDIA-labs OO エージェント: ネイティブ Python オブジェクト指向エージェント
従来のエージェント開発は、プロンプト テンプレート、ツール スキーマ、コールバック コード、およびワークフロー グラフに分割されています。信頼できる AI エージェントを構築するためのモデルに依存しない Python フレームワークである NVIDIA オブジェクト指向エージェント (NOOA) を紹介します。 NOOA はより単純なアプローチを採用しています。つまり、エージェントは Python オブジェクトです。そのメソッドはモデルが実行できるアクション、フィールドはモデルの状態、docstring はプロンプト、型アノテーションはコントラクトです。コード本体が「...」で構成されるメソッドは、実行時に LLM 駆動のエージェント ループによって完了しますが、通常の本体を持つメソッドは標準の決定論的な Python のままです。これにより、開発者とエージェントに同じインターフェイスが提供されるため、他のソフトウェアと同様にエージェントの動作をテスト、追跡、リファクタリング、改善することができます。この論文は 3 つの貢献を行っています。 (1) Python オブジェクトとしてのエージェントのプログラミング モデルとその背後にある設計原則を示します。 Python に既存の抽象化がある場合は、それらを直接採用します。エージェント固有の機能 (コンテキスト、イベント、状態レンダリング、長期メモリ、検証済み LLM ループ) は、シンプルな Python API を通じて公開されるため、開発者とエージェントの両方が 1 つの使い慣れたプログラミング モデルを共有します。 (2) 私たちは、モデルに面した 6 つのアイデアを特定します。NOOA は、私たちの知る限り、単一の表面上で最初に組み合わせたものです。それは、型付き入出力、ライブ オブジェクトの参照渡し、アクションとしてのコード、プログラマブル ループ エンジニアリング、明示的なオブジェクトの状態、コンテキストとイベント用のモデル呼び出し可能なハーネス API です。私たちは、コミュニティがすでにこれらのアイデアのいくつか (多くの場合、実験的または部分的な機能として) に収束していることを発見し、さらなる採用を促進するために比較を提示します。 (3) 現在のモデルが、ターゲットを絞った機能テストと、SWE ベンチ検証済み、ターミナル ベンチ 2.0、ARC-AGI-3 などのエージェントおよび推論ベンチマークの両方で、このインターフェイスを効果的に使用していることを実証します。
原文 (English)
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its methods are the actions the model can take, fields are its state, docstrings are its prompts, and its type annotations are contracts. A method whose code body consists of "..." is completed at runtime by an LLM-driven agent loop, while methods with normal bodies remain standard deterministic Python. This gives developers and agents the same interface, so agent behavior can be tested, traced, refactored, and improved just like other software. This paper makes three contributions. (1) We present the agent-as-a-Python-object programming model and the design principles behind it. Where Python has existing abstractions, we adopt them directly. Agent-specific capabilities--context, events, state rendering, long-term memory, and validated LLM loops--are exposed through simple Pythonic APIs, so both developers and agents share one familiar programming model. (2) We identify six model-facing ideas that NOOA is, to our knowledge, the first to combine on a single surface: typed input/output, pass-by-reference over live objects, code as action, programmable loop engineering, explicit object state, and model-callable harness APIs for context and events. We find the community already converging on several of these ideas--often as experimental or partial features--and present the comparison to encourage further adoption. (3) We demonstrate that current models use this interface effectively, both in targeted capability tests and on agentic and reasoning benchmarks such as SWE-bench Verified and Terminal-Bench 2.0 and ARC-AGI-3.
ArbiGraph: コンテキスト管理を評価するための任意にスケーラブルで検証可能なタスク グラフ
ツール支援言語エージェントが拡張推論ワークフロー全体でタスク関連のコンテキストを保持、更新、作成、破棄できるかどうかを評価するベンチマーク ジェネレーターである ARBIGRAPH を紹介します。 ARBIGRAPH は、各タスクを実行可能な Python ソルバーを使用した自然言語問題として表し、ここではスカラー値とリスト値としてインスタンス化された、型指定された中間状態を通じてタスクを構成します。この設計により、正確な自動検証を維持しながら、長さ、依存構造、ディストラクタの数、および値のタイプを変更できる、制御可能なタスク グラフが可能になります。数学、GSM スタイルの単語問題、および Python トレース タスク カテゴリを使用して ARBIGRAPH をインスタンス化し、4 つのトポロジにわたって Qwen3.5-27B ツール支援エージェントを評価します。結果は、孤立したタスクでは高い精度が示されましたが、より複雑な依存タスクでは大幅に低下しました。依存した数学タスクの分岐チェーンでは精度が最大 33.3% 低下しました。これは、ARBIGRAPH が単一タスクの評価だけでは見えない障害を明らかにすることを示しています。コード、生成されたデータセット、評価結果は、https://github.com/pavelgolikov/ArbiGraph.git で入手できます。
原文 (English)
ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended reasoning workflows. ARBIGRAPH represents each task as a natural-language problem with an executable Python solver, and composes tasks through typed intermediate states, instantiated here as scalar and list values. This design enables controllable task graphs whose length, dependency structure, distractor count, and value type can be varied while preserving exact automatic verification. We instantiate ARBIGRAPH with math, GSM-style word-problems, and Python-tracing task categories, and evaluate a Qwen3.5-27B tool-assisted agent across four topologies. The results show high accuracy on isolated tasks but substantial degradation on more complex dependent tasks: accuracy drops by up to 33.3% on branching chains of dependent math tasks. This shows that ARBIGRAPH exposes failures that are not visible from single-task evaluation alone. Our code, generated datasets, and evaluation results are available at https://github.com/pavelgolikov/ArbiGraph.git
人間と AI の代替原則: あなたの組織はいつ AI に置き換えられますか?
人工知能 (AI) は組織を急速に変革しており、人間の従業員はいつ AI に取って代わられるのかという組織的および経済的な根本的な疑問を引き起こしています。階層型組織における人間と AI のタスク割り当て (HAT) を研究するための分析モデルを紹介します。 HAT モデルの中心的な特徴は、人間のスキル習得と AI 能力のスケーリングの間の経済的な非対称性を形式的にコード化していることです。 HAT モデルを使用すると、リスク調整済みのコスト、スキル、組織の深さ、導入規模、戦略的適応、およびリスクを総合して、いつ、どこで、なぜ、どのような構造的条件下で人間と AI の代替が発生するかをどのように決定するかを導き出すことができます。重要な結果は、人間と AI の代替原理です。これは、AI が人間の労働を置き換える正確な条件を提供します -- 形式的な非対称性の仮定に基づいています --。この結果に基づいて、AI の導入により、急激な労働力の移行、人間と AI のハイブリッド組織(リスクの異質性によって人間の割合の最小制約を必要とせずに人間と AI の役割が維持される場合を含む)、より広い制御範囲を備えたより平坦な管理階層が生成される可能性があることを示します。 HAT モデルは、中間管理職の役割が自動化に対して高い脆弱性を示す構造的条件を特定し、高度なスキルを持つ労働者の脆弱性が、組織の深さ、基準コスト、およびリスクの差によって形成されるスキルの閾値に依存することを示しています。より広範には、この論文は自動化の経済学、組織設計、AI ガバナンス、および労働力計画を AI 主導の組織変革の統一理論に結び付けます。
原文 (English)
The Human-AI Substitution Principle: When will you be replaced by AI in your organization?
Artificial Intelligence (AI) is rapidly transforming organizations, raising a fundamental organizational and economic question: when will a human employee be replaced by AI? We present an analytical model for studying Human--AI Task Allocation (HAT) in hierarchical organizations. A central feature of the HAT model is that it formally encodes the economic asymmetry between human skill acquisition and AI capability scaling. The HAT model allows us to derive how risk-adjusted costs, skills, organizational depth, deployment scale, strategic adaptation, and risk jointly determine when, where, why, and under what structural conditions human--AI replacement occurs. A key result is the Human--AI Substitution Principle, which provides a precise condition --- grounded in the formal asymmetry assumption --- under which AI replaces human labor. Building on this result, we show that AI adoption can produce abrupt workforce transitions, hybrid human--AI organizations, including cases where risk heterogeneity sustains human and AI roles without requiring a minimum-human-fraction constraint, and flatter managerial hierarchies with wider spans of control. The HAT model identifies structural conditions under which middle-management roles exhibit elevated vulnerability to automation, and shows that the vulnerability of highly skilled workers depends on a skill threshold shaped by organizational depth, baseline costs, and risk differentials. More broadly, the paper connects automation economics, organizational design, AI governance, and workforce planning into a unified theory of AI-driven organizational transformation.
拒否ゲート型デコード: 高温サンプリング下での拒否行動の保存
高温サンプリングは、LLM の多様性を高めるための主要なメカニズムの 1 つです。切り捨てベースのサンプリング技術の最近の進歩は、ニューラルテキストの変性などの高温サンプリングの欠点を軽減するのに役立ち、それによって一貫性を犠牲にすることなく LLM 出力の多様性を高めることが可能になりました。ただし、高温によるトークン確率分布のエントロピーの増加は、有害なプロンプトの存在下でのモデルの拒否応答を減少させ、モデルのガードレールを弱めることも示されています。高温サンプリングの潜在的な利点とモデルの安全性を維持することの重要性にもかかわらず、より高いエントロピー領域下で LLM の拒否動作を維持するための既存のソリューションが不足しています。このギャップに対処するために、私たちは温度が LLM の拒否動作にどのような影響を与えるかを体系的に研究し、追加の遅延を最小限に抑えながら高温でのモデルの貪欲な復号拒否応答を維持する効率的な逐次復号アプローチを提案します。広範な実験を通じて、私たちのアプローチは、安全なプロンプトに対するモデルの高温応答を損なうことなく、3 つのベンチマーク データセットにわたって貪欲なデコード拒否動作の 91 ~ 99% を保存することを示しました。私たちの研究は、高温のサンプリングを必要とする用途において、拒否動作を効率的な方法でどのように維持できるかを実証しています。
原文 (English)
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling techniques have helped mitigate drawbacks of high-temperature sampling such as neural text degeneration, thereby enabling greater diversity in LLM outputs without sacrificing coherence. However, increasing the entropy of the token probability distribution via high temperatures has also been shown to weaken model guardrails by reducing the model's refusal response in the presence of harmful prompts. Despite the potential benefits of high-temperature sampling and the importance of maintaining model safety, there is a lack of existing solutions for maintaining the refusal behavior of LLMs under a higher entropy regime. To address this gap, we systematically study how temperature influences refusal behavior in LLMs and propose an efficient sequential decoding approach which preserves a model's greedy decoding refusal response at high temperatures while incurring minimal additional latency. Through extensive experiments, we show that our approach preserves 91-99% of the greedy decoding refusal behavior across three benchmark datasets without compromising the model's high-temperature response for safe prompts. Our work demonstrates how refusal behavior can be maintained in an efficient manner for applications which require high-temperature sampling.
AI システムは創造的になることができるでしょうか?アートとエンジニアリングからの批判的な視点
この論文では、人工知能 (AI) システムが創造的であるかどうかという問題を、電気工学、パターン認識、機械学習、ニューラル ネットワークの訓練を受け、人生のほとんどを俳優、舞台監督、映画監督、作家、作曲家、ビジュアル アーティストとしての芸術と哲学に費やしてきた研究者の二重の観点からアプローチして検証します。この論文は、マーガレット・ボーデンの基本的な枠組み、彼女の創造性の 3 つの特性 (新規性、驚き、価値) と 3 つのタイプの創造プロセス (組み合わせ的、探索的、変革的) の両方を利用して、AI システムは構造的に最も強い意味での創造性を発揮できないと主張しています。彼らは、組み合わせによる創造性の領域では本物の能力を発揮しますが、探索的な創造性にはかなり限界があり、根本的に変革的な創造性を発揮することができません。同論文はさらに、現在のAIシステムの最も重要な限界は、新規性そのものが欠如しているのではなく、創造性の現象学において中心的な役割を果たしている偶然、偶然、予期せぬ出来事を発生させるメカニズムが欠如していること、そしてそのような偶然の出来事を認識し歓迎するための主体的立場が欠如していることだと主張している。この論文は、いくつかの具体的な実験によって示された、現実的かつ生成的な人間と AI の創造的なコラボレーションのモデルを提案することで締めくくられています。この論文自体が、論文が推進する論文のデモンストレーションです。この論文は、意図的な人間 AI の共同プロセスを通じて構成されており、その冒頭の方法論的なメモに説明されています。
原文 (English)
Can an AI System Be Creative? A Critical Perspective from Art and Engineering
This paper examines the question of whether artificial intelligence (AI) systems can be creative, approached from the dual perspective of a researcher trained in electrical engineering, pattern recognition, machine learning, and neural networks, who has also spent most of his life engaged in the arts as actor, stage and film director, writer, composer, and visual artist, and in philosophy. Drawing on Margaret Boden's foundational framework, both her three properties of creativity (novelty, surprise, and value) and her three types of creative processes (combinatorial, exploratory, and transformational), the paper argues that AI systems are structurally incapable of creativity in its strongest sense. While they exhibit genuine capability in the domain of combinatorial creativity, they are significantly bounded in exploratory creativity, and fundamentally incapable of transformational creativity. The paper further argues that the most important limitation of current AI systems is not the absence of novelty per se, but the absence of any mechanism for serendipity, accident, or the unexpected, all of which play a central role in the phenomenology of creativity, and the absence of any subject position from which to recognize and welcome such chance events. The paper concludes by proposing a model of human, AI creative collaboration that is both realistic and generative, illustrated by several concrete experiments. The paper is itself a demonstration of the thesis it advances: it was composed through a deliberate human AI collaborative process, which is described in the methodological note that opens it.
軽量の大規模言語モデルのプロファイリング
軽量ラージ言語モデル (LLM) は、パーソナル コンピューター上でローカルに展開されることが増えており、リソースに制約のあるエッジ環境やモバイル環境でますます大きな役割を果たすことが期待されています。このような設定では、エネルギー消費、実行時間、メモリ使用量が実際の使いやすさに直接影響しますが、LLM 効率の既存の評価は主にパラメータ数や FLOP などのプロキシ記述子に依存しており、多くの場合タスクの精度とは切り離されています。このペーパーでは、軽量 LLM 推論の精度を意識したプロファイリングのための PTME ベースの実験フレームワークを紹介し、ハードウェア レベルの直接測定を通じて精度、実行時間、ピーク メモリ使用量、およびエネルギー消費を共同測定します。この方法論は、コード生成、数学的推論、およびマルチタスクの理解にわたるベンチマークを使用して、制御されたデスクトップ プラットフォーム上のエッジクラス リソース エンベロープでローカルに実行される軽量 LLM の代表的なセットに適用されます。静的プロキシ記述子は推論コストをうまく近似しますが、精度を予測できないことがわかりました。リソース エンベロープを厳しくすると、精度に影響を与えることなくコストが増加し、エネルギーよりも実行時間が大幅に増大し、大規模なモデルに最も大きな不利益をもたらします。さらに、すべての PTME ディメンションにわたって優勢な単一のモデルはなく、パレート分析により、精度のみまたは効率のみの評価では隠蔽される非優勢な構成が明らかになり、さまざまなリソース エンベロープでモデルを選択するための実用的なガイダンスが提供されます。これらの結果は、サイズ、FLOP、遅延、精度だけで軽量 LLM を選択すると、間違った導入候補が選択される可能性があることを示しています。 PTME プロファイリングは、より低い物理コストで有用な精度を維持する構成を明らかにします。
原文 (English)
Profiling Lightweight Large Language Models
Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are expected to play a growing role in resource-constrained edge and mobile environments. In such settings, energy consumption, execution time, and memory usage directly affect practical usability, yet existing evaluations of LLM efficiency largely rely on proxy descriptors such as parameter count or FLOPs, often decoupled from task precision. This paper introduces a PTME-based experimental framework for the precision-aware profiling of lightweight LLM inference, jointly measuring Precision, execution Time, peak Memory usage, and Energy consumption through direct hardware-level measurements. The methodology is applied to a representative set of lightweight LLMs executed locally under edge-class resource envelopes on a controlled desktop platform, using benchmarks spanning code generation, mathematical reasoning, and multi-task understanding. We find that static proxy descriptors approximate inference cost well but fail to predict precision. Tightening the resource envelope increases cost without affecting precision, amplifying execution time more strongly than energy and penalizing larger models the most. Moreover, no single model dominates across all PTME dimensions, and a Pareto analysis reveals non-dominated configurations that would be hidden by accuracy-only or efficiency-only assessments, providing practical guidance for selecting models under different resource envelopes. These results show that selecting lightweight LLMs by size, FLOPs, latency, or accuracy alone can select the wrong deployment candidate; PTME profiling exposes configurations that preserve useful accuracy at lower physical cost.
ガイド接地型マルチモーダル LLM による説明可能な心臓診断の強化
心電図 (ECG) は心臓評価の基礎ですが、深層学習モデルの臨床展開は、解釈可能性の制限と大規模言語モデル (LLM) の幻覚リスクにより依然として制約を受けています。既存の CNN+Grad-CAM+マルチモーダル LLM フレームワークは ECG レポートを生成できますが、その説明は確立された診断基準にあまり根拠がないことが多く、信頼性と再現性が低下します。私たちは、厳選された臨床知識に基づいてレポート作成を明示的に固定する、ガイドに基づいたマルチモーダル フレームワークを提案します。畳み込みニューラル ネットワーク (CNN) と Grad-CAM は、まず 12 誘導 ECG 画像からクラス確率とクラス固有のヒートマップを生成します。並行して、権威ある ECG 教科書とガイドライン資料がオフラインで構造化された ECG 解釈ガイドに抽出され、すべてのサンプルに対する固定知識ブロックとして注入されます。マルチモーダル LLM は、ECG 画像、Grad-CAM オーバーレイ、CNN 由来のファクト パック、および挿入されたガイドを条件として、ガイドラインに準拠した用語と基準の使用法を使用して構造化された診断レポートを生成します。 PTB-XL テスト セット全体での実験では、ガイド グラウンディングにより、競争力のある分類パフォーマンスを維持しながら、生成されたレポートのセマンティック品質と認識される一貫性が向上することが実証されました。特に、私たちの方法は、強力な CNN+Grad-CAM+MLLM ベースラインと比較して、生成されたインプレッションの平均 BERTScore を 0.818 から 0.953 に増加させ、参考レポートとより緊密に一致していることを示しています。これらの発見は、蒸留された解釈ガイドをマルチモーダルなプロンプトパイプラインに注入することで、幻覚を軽減し、LLM ベースの ECG 説明の臨床的妥当性を高める実用的な経路を提供し、説明可能な心臓診断を現実世界での展開に近づけることを示唆しています。
原文 (English)
Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs
The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning models remains con- strained by limited interpretability and the hallucination risk of large language models (LLMs). Existing CNN+Grad-CAM+multimodal LLM frameworks can generate ECG reports, but their explanations are often only weakly grounded in established diagnostic criteria, reducing trust- worthiness and reproducibility. We propose a guide-grounded multimodal framework that explicitly anchors report generation in curated clinical knowledge. A convolutional neural network (CNN) and Grad-CAM first produce class probabilities and class-specific heatmaps from 12-lead ECG images. In parallel, authoritative ECG textbooks and guideline materials are distilled offline into a structured ECG Interpretation Guide, which is injected as a fixed knowledge block for every sample. Conditioned on the ECG image, Grad-CAM overlay, CNN-derived fact pack, and the in- jected guide, a multimodal LLM generates structured diagnostic reports with guideline-consistent terminology and criteria usage. Experiments on the full PTB-XL test set demonstrate that guide grounding improves se- mantic quality and perceived consistency of generated reports while pre- serving competitive classification performance. In particular, our method increases the average BERTScore of generated impressions from 0.818 to 0.953 relative to a strong CNN+Grad-CAM+MLLM baseline, indicat- ing closer alignment with reference reports. These findings suggest that injecting a distilled interpretation guide into the multimodal prompting pipeline offers a practical pathway to reduce hallucinations and enhance the clinical plausibility of LLM-based ECG explanations, bringing ex- plainable cardiac diagnosis closer to real-world deployment.
軽量の時間畳み込みネットワークによる効率的で解釈可能な身体ベースの感情認識
身体ベースの感情認識はリアルタイムの感情システムにとって重要ですが、グラフベースのスケルトン モデルは計算コストが高くなる可能性があります。この論文では、軽量の時間畳み込みネットワーク (TCN) が、身体ベースの感情分類に対する効率的で解釈可能な代替手段を提供できるかどうかを研究します。 DIEM-A で TCN モデルのファミリーを評価し、精度、マクロ F1、パラメーター数、推論レイテンシを使用してグラフベースの時系列グラフ (G-TSG) ベースラインと比較します。 G-TSG は最高の平均パフォーマンスを達成しますが、TCN-Base は $79.18\%$ 少ないパラメーターを使用し、約 $12.5\times$ 削減しながら、精度ポイントは $1.58$、マクロ F1 ポイントは $1.25$ 以内に留まっています。また、領域固有の TCN モデル、ゼロベース オクルージョン、G-TSG 勾配顕著性を使用して、身体領域の寄与も分析します。結果は、上半身の動きが最も強力な独立した局所的手がかりを提供すること、身体領域の有用性が感情によって異なること、および異なる解釈方法がモデルの行動の異なる側面を捉えていることを示しています。これらの発見は、軽量 TCN が効率的な身体ベースの感情認識をサポートできると同時に、動作の手がかりが分類にどのように寄与するかについての実用的な洞察を提供できることを示唆しています。
原文 (English)
Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks
Body-based emotion recognition is important for real-time affective systems, but graph-based skeleton models can be computationally expensive. This paper studies whether lightweight temporal convolutional networks (TCNs) can provide an efficient and interpretable alternative for body-based emotion classification. We evaluate a family of TCN models on DIEM-A and compare them with a graph-based time-series graph (G-TSG) baseline using accuracy, macro-F1, parameter count, and inference latency. Although G-TSG achieves the highest mean performance, TCN-Base remains within $1.58$ accuracy points and $1.25$ macro-F1 points while using $79.18\%$ fewer parameters and reducing classifier latency by approximately $12.5\times$. We also analyze body-region contributions using region-specific TCN models, zero-based occlusion, and G-TSG gradient saliency. The results show that upper-body motion provides the strongest standalone regional cue, that the usefulness of body regions varies across emotions, and that different interpretability methods capture distinct aspects of model behavior. These findings suggest that lightweight TCNs can support efficient body-based emotion recognition while also providing practical insight into how motion cues contribute to classification.
LLM エージェントのアクション選択における出所の機密性の監査
LLM エージェントは、ユーザーのリクエスト、ツールの出力、取得したレコード、メモリ、信頼できないテキストが混在するコンテキストからツールと引数を選択します。決定を下す権限がなくても証拠が関連する可能性があるため、正しい行動は許可された証拠のみに基づいている必要はありません。ツールおよび引数ターゲットごとにコンテキスト要素を個別にラベル付けする、ターゲット固有の認可監査を導入します。その主なテストでは、タスク、命題、立場、ポリシーを固定し、命題のソース権限のみを変更します。次に、有効な証拠が弱まった場合の動作をテストし、二次的な位置推定診断としてコンテキストとサブセットの相互作用を使用します。 450 の制御された次のアクション タスクと複数のオープンウェイト LLM ファミリにわたって、信頼できるバリアントと信頼できないバリアントは、競合ケースの 5.4 パーセントに対して、サポート ケースの 1.7 パーセントで異なるアクションを生成します。制御された劣化の下では、無許可の競争は、比較の 2.4 パーセントで完全正解、混合エラー、完全正解のパターンで維持され、95 パーセントの信頼区間は 2.1 ~ 3.0 パーセントです。これらは制御されたストレス設定率であり、展開の蔓延ではありません。モデルはテキストの情報源と権威の手がかりに反応しますが、信頼できない証拠がモデルの行動に影響を与えることを防ぐことはできません。
原文 (English)
Auditing Provenance Sensitivity in LLM Agent Action Selection
LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can be relevant without being authorized to determine a decision, so a correct action need not be grounded only in permitted evidence. We introduce a target-specific authorization audit that labels context factors separately for each tool and argument target. Its primary test holds the task, proposition, position, and policy fixed while changing only the proposition's source authority. We then test behavior when valid evidence is weakened and use context-subset interactions as a secondary localization diagnostic. Across 450 controlled next-action tasks and multiple open-weight LLM families, trusted and untrusted variants produce different actions in 5.4 percent of competing cases versus 1.7 percent of supporting cases. Under controlled degradation, unauthorized competition is retained in a full-correct, mixed-error, clean-correct pattern in 2.4 percent of comparisons, with a 95 percent confidence interval from 2.1 to 3.0 percent. These are controlled stress-set rates, not deployment prevalence. The models respond to textual source-authority cues, but this does not prevent untrusted evidence from influencing their actions.
医療 LLM 診断における監査証拠の使用
医療 LLM は多くの場合、正しい診断を選択したかどうかによって評価されますが、診断の精度だけでは、モデルが症例証拠を適切に使用したかどうかを示すことはできません。医療診断における証拠の使用に関する行動監査を紹介します。ケースごとに、患者情報を証拠単位に分解し、管理された証拠サブセットの下で診断候補をスコアリングし、診断マージンにおける低次相互作用をマイニングします。医学的証拠は診断に関連しているため、監査では相互作用の発見と失敗の割り当てが分離されます。大規模または否定的な相互作用は、もっともらしい鑑別診断を反映している可能性がありますが、疑わしい相互作用には堅牢性チェックと臨床レビューが必要です。 DDXPlus、CupCase、MedCase で 5 つのオープンウェイト LLM を評価します。データセット全体で、忠実なサポートと差異のある競合またはキャンセルがインタラクション強度のほとんどを占めており、多くの証拠のインタラクションが失敗ではなく臨床的に妥当であることを示しています。 DDXPlus に焦点を当てた盲検の 5 人の査読者による 130 項目の充実レビューサンプルでは、無効またはショートカットのような症例が、否定または欠如した所見および臨床的に局所的な証拠に集中しています。これらの結果は、精度が候補者の証拠利用の失敗を隠し、医療 LLM 評価の役割を意識した監査を動機付けることができることを示しています。
原文 (English)
Auditing Evidence Use in Medical LLM Diagnosis
Medical LLMs are often evaluated by whether they select the correct diagnosis, but diagnostic accuracy alone does not show whether the model used the case evidence appropriately. We present a behavioral audit of evidence use in medical diagnosis. For each case, we decompose patient information into evidence units, score candidate diagnoses under controlled evidence subsets, and mine low-order interactions in diagnostic margins. Because medical evidence is diagnosis-relative, the audit separates interaction discovery from failure assignment: large or negative interactions can reflect plausible differential diagnosis, while suspicious interactions require robustness checks and clinical review. We evaluate five open-weight LLMs on DDXPlus, CupCase, and MedCase. Across datasets, faithful support and differential conflict or cancellation account for most interaction strength, showing that many evidence interactions are clinically plausible rather than failures. In a DDXPlus-focused blinded five-reviewer 130-item enriched review sample, invalid or shortcut-like cases concentrate in negated or absent findings and clinically local evidence. These results show that accuracy can hide candidate evidence-use failures and motivate role-aware audits for medical LLM evaluation.
公開テストに合格したコードのための Code Monitor レッド チーミング
可視テストは LLM で生成されたコードの一般的なゲートですが、テストに合格しても仕様の正確性は保証されません。私たちは、展開に似た監視問題を研究します。コードが公開テストに合格した後、より弱い LLM 検証ツールが残存する隠れたバグを特定できるでしょうか? Code Monitor Red Teaming を導入します。これは、ジェネレーターの圧力、検証者の足場、および弱い機能から強い機能までを変化させながら、公開チェック情報の境界を修正するモニター レッド チーミング プロトコルです。これを、関数レベル、データ サイエンス、ワークフロー コードにまたがる CodeMonitorBench としてインスタンス化します。生成された 71,000 人の候補者のうち、43,677 人が公開テストに合格し、そのうち 23,081 人が非表示のテストに不合格でした。弱い検証ツールは、スキャフォールディングとモデル ファミリによって改善されますが、依然として 5% の誤検知率で隠れたバグのほとんどを見逃します。ロバスト性ストレス テストとして、敵対的な公開テストのオーバーフィット プレッシャーにより、検証者の AUROC が低下し、ほとんどのセルで低 FPR ミス率が上昇します。 GLM-5.1 検証者は、同じ証拠境界の下にあるギャップの一部を回復します。推論監査では、残りのミスには検証者の失敗と M1 証拠の制限が混在していることが示されています。
原文 (English)
Code Monitor Red Teaming for Public-Test-Passing Code
Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring problem: after code has passed public tests, can a weaker LLM verifier identify the residual hidden bugs? We introduce Code Monitor Red Teaming, a monitor-red-teaming protocol that fixes a public-check information boundary while varying generator pressure, verifier scaffolding, and weak-to-strong capability. We instantiate it as CodeMonitorBench, spanning function-level, data-science, and workflow code. Across 71,000 generated candidates, 43,677 pass public tests and 23,081 of those fail hidden tests. Weak verifiers improve with scaffolding and model family, but still miss most hidden bugs at 5% false-positive rate. As a robustness stress test, adversarial public-test-overfit pressure lowers verifier AUROC and raises low-FPR miss rates in most cells. A GLM-5.1 verifier recovers part of the gap under the same evidence boundary; an inferability audit shows that remaining misses mix verifier failures with M1 evidence limits.
綿密な調査は信頼できるのでしょうか?誤解を招く知識は誤った結論を引き起こす
Deep Research エージェントは、LLM ベースのアシスタントを、計画、検索、証拠の合成、レポート生成を含む長期的なワークフローに拡張しますが、オープンな情報環境におけるその信頼性はまだ十分に解明されていません。主な懸念は、このような環境で遭遇した、一見信頼できても事実を誤解を招く知識が、これらのワークフローを通じて伝播し、最終レポートで誤った結論として採用される可能性があるかどうかです。この障害モードを研究するために、ディープリサーチタスク用の誤解を招く知識を構築および検証するためのフレームワークである MisKnow-Agent を紹介します。 MisKnow-Agent は、制御可能な権限レベルとスタイルを備えた誤解を招くインスタンスを生成し、DeepResearch Benchmark タスクに基づいて構築された 5,933 個の品質管理されたインスタンスを生成します。オープンソースおよびクローズドソースのディープ リサーチ エージェントにわたる広範な実験により、誤解を招く知識への限定的な暴露であっても、最終レポートでの誤った結論の採用を誘発する可能性があり、現在のディープ リサーチ エージェントに広範な信頼性の脆弱性があることが明らかになりました。検索可能な検証モデルは、焦点を絞ったコーパス検証中に保持されたインスタンスが誤解を招くものとして一貫して識別しますが、同じインスタンスが長期的な研究中に依然として採用される可能性があり、焦点を絞った検証とワークフローレベルの証拠の使用との間に断絶があることが明らかになります。最後に、調査前後の防御を個別に、または組み合わせて評価し、3 つの構成すべてが誤った結論の採用を軽減するものの、完全には防止できないことを発見しました。私たちの調査結果は、信頼性の高いディープリサーチには、計画、検索、証拠の統合、またはレポート作成能力の向上だけでなく、モデルレベルとフレームワークレベルの両方で証拠の検証と修正の能力が必要であることを示唆しています。
原文 (English)
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report generation, yet their reliability in open information environments remains underexplored. A key concern is whether apparently credible but factually misleading knowledge encountered in such environments can propagate through these workflows and be adopted as false conclusions in final reports. To study this failure mode, we introduce MisKnow-Agent, a framework for constructing and validating misleading knowledge for Deep Research tasks. MisKnow-Agent generates misleading instances with controllable authority levels and styles, yielding 5,933 quality-controlled instances built on DeepResearch Benchmark tasks. Extensive experiments across open-source and closed-source Deep Research agents show that even limited exposure to misleading knowledge can induce false-conclusion adoption in final reports, revealing a broad reliability vulnerability in current Deep Research agents. Although search-enabled verifier models consistently identify the retained instances as misleading during focused corpus validation, the same instances can still be adopted during long-horizon research, revealing a disconnect between focused verification and workflow-level evidence use. Finally, we evaluate pre- and post-research defenses, both individually and in combination, finding that all three configurations mitigate but do not fully prevent false-conclusion adoption. Our findings suggest that reliable Deep Research requires evidence verification and correction capabilities at both the model and framework levels, beyond improvements in planning, retrieval, evidence integration, or report-generation abilities.
効率的な拡散モデル微調整のためのソース優先駆動の選択的適応
新しいドメインやスタイルに合わせて大規模な拡散モデルを微調整するには、トレードオフが伴います。ターゲット固有の生成を改善すると、事前学習済みモデルの広範な生成能力が低下することがよくあります。既存の完全かつパラメータ効率の高い微調整方法は通常、このトレードオフを暗黙的にのみ処理します。この研究では、拡散モデルを効率的に微調整し、有利なトレードオフを達成するための新しいソース優先駆動の選択的適応方法を提案します。私たちの方法は、2 つの重要な観察に基づいています。(1) 一般生成能力の損失は、事前トレーニングされたパラメーター間で非常に一貫性がありません。(2) モデルの一般生成能力に比較的小さな影響を与えるパラメーターは、レイヤーとパラメーターの種類全体で構造的に一貫性がないままです。これらの観察に動機付けられて、私たちはまず静的マスクを学習して下流の適応により適したパラメーターを明示的に特定し、次に選択されたサブセットの構造化された更新戦略を構築します。実験では、私たちの方法が既存の強力なベースラインよりも優れた適応と保持のトレードオフを達成していることを示しています。
原文 (English)
Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning
Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained model's broad generative capability. Existing full and parameter-efficient fine-tuning methods typically handle this trade-off only implicitly. In this work, we propose a novel source-prior-driven selective adaptation method to efficiently fine-tune diffusion models, achieving a favorable trade-off. Our method relies on two key observations: (1) the loss of general generative capability is highly inconsistent across pretrained parameters, and (2) parameters that have a relatively small impact on the model's general generative capability remain structurally inconsistent across layers and parameter types. Motivated by these observations, we first learn a static mask to explicitly identify parameters better suited for downstream adaptation, and then construct structured update strategies for the selected subset. Experiments show that our method achieves a better adaptation-retention trade-off than existing strong baselines.
追跡可能な奨学金: 生成 AI 時代における人文探求のためのページ アンカーとアリアドネのスレッド
生成 AI により、大規模な言語モデルが数秒以内に学術的なテキストを生成できるようになりますが、流暢さは有効な説明と同じではありません。最も深刻なリスクは、事実誤認だけではなく、明確な出典、ページ番号、版、証拠なしに説明がすでに確立されているように見えることです。私たちはページアンカーをアリアドネの糸に例えます。生成の流暢さの迷路の中で、それは学者を源へと導く糸です。この論文は、印刷、デジタル、生成 AI という知識インフラストラクチャの 3 つの革命全体にわたって、AI 支援の人文研究の最低規範条件として追跡可能な奨学金を提案します。ページ アンカー、デュアル ページ番号、引用第一世代、NO_EVIDENCE、人間による検証、4 レベルのコンプライアンス、およびスコープ コントラクトを導入し、AIH-Infra を 3 層のリファレンス実装として示します: コンテキスト (文書構造化)、Open WebUI AIH-Infra (追跡可能なナレッジ ベース)、および AIH-Infra MCP サーバー (エージェント ゲートウェイ)。 29 巻の Kant Akademie-Ausgabe 知識ベースに関するケーススタディでは、トレーサビリティが検索の修正、証拠の格付け、および判決の格下げをどのようにサポートするかを示しています。トレーサビリティはソフトウェアの機能ではありません。それは、生成型 AI の時代において、人文科学的研究が公にされ、反駁可能であり続けるための条件です。
原文 (English)
Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI
Generative AI lets large language models produce scholarly-looking text within seconds, yet fluency does not equal valid explanation. The deepest risk is not factual error alone but the appearance that an explanation is already established without clear sources, page numbers, editions, or evidence. We liken the page anchor to Ariadne's thread: within the labyrinth of generative fluency, it is the thread that leads the scholar back to the source. This paper proposes Traceable Scholarship as the minimum normative condition for AI-assisted humanistic research, situating it across the three revolutions of knowledge infrastructure: print, digital, and generative AI. We introduce page anchors, dual page numbers, citation-first generation, NO_EVIDENCE, human verification, four-level compliance, and Scope Contract, and present AIH-Infra as a three-layer reference implementation: Contexture (document structuring), Open WebUI AIH-Infra (traceable knowledge base), and AIH-Infra MCP Server (agent gateway). A case study on a 29-volume Kant Akademie-Ausgabe knowledge base illustrates how traceability supports retrieval correction, evidence grading, and judgment downgrading. Traceability is not a software feature; it is the condition under which humanistic research can remain public and refutable in the age of generative AI.
OPOD: ポリシーに基づくオムニ蒸留
オムニモーダル モデルは、テキスト、画像、音声を 1 つのシステムで処理できますが、これらすべての機能を同時に向上させることは依然として困難です。プールされたマルチモーダル データで単一のモデルをトレーニングすると、個々のモダリティに特化したモデルと一致しないことがよくあります。オンポリシー蒸留 (OPD) は、そのような専門家を組み合わせる方法を提供します。生徒が応答を生成し、教師がその同じ応答を評価するため、生徒は実際に生成された行動から直接学習します。しかし、複数の教師を使用すると、競合する指導が導入され、別の指導法を犠牲にして 1 つの指導法が改善される可能性があります。 On-Policy Omni Distillation (OPOD) を導入し、各生徒の応答を一致するテキスト、画像、または音声の教師にルーティングします。 OPOD は、教師が生徒よりも高い確率を割り当てるトークンのみに教師のガイダンスを維持し、トレーニング中に各モダリティ教師の影響を個別に調整し、ルーティングされた教師に最終的な解答と推論が正解を裏付けるかどうかの両方を評価するよう依頼します。 12 のベンチマークと 3 つのバックボーン サイズにわたって、OPOD はすべてのスケールで最高の平均スコアを達成し、70.8、51.7、46.2 に達し、最も強力な比較ツールを 2.1、1.8、1.7 ポイント上回っています。 30B モデルでは、12 のベンチマークすべてで、基本モデルと、プールされたマルチモーダル データで共同でポストトレーニングされた対応モデルの両方を上回り、個々の専門家が含まれている場合でも、11 のベンチマークで 1 位または 2 位にランクされます。スペシャリストはトレーニング後に破棄され、展開可能なオムニモーダル モデルが 1 つだけ残ります。これらの結果は、モダリティ固有の教師を調整することが、クロスモーダルのバランスを維持しながら共有モデルを改善する効果的な方法であることを示しています。
原文 (English)
OPOD: On-Policy Omni Distillation
Omni-modal models can handle text, images, and audio in one system, but improving all of these abilities together remains difficult. Training a single model on pooled multimodal data often fails to match models specialized for individual modalities. On-policy distillation (OPD) offers a way to combine such specialists: the student generates a response, and a teacher evaluates that same response, so the student learns directly from behaviors it actually produces. Yet using several teachers can introduce competing guidance and improve one modality at the expense of another. We present On-Policy Omni Distillation (OPOD), which routes each student response to the matching text, image, or audio teacher. OPOD keeps teacher guidance only on tokens where the teacher assigns a higher probability than the student, adjusts the influence of each modality teacher independently during training, and asks the routed teacher to assess both the final answer and whether the reasoning supports the correct answer. Across twelve benchmarks and three backbone sizes, OPOD achieves the best average score at every scale, reaching 70.8, 51.7, and 46.2 and exceeding the strongest comparator by 2.1, 1.8, and 1.7 points. On the 30B model, it outperforms both the base model and a counterpart post-trained jointly on pooled multimodal data on all twelve benchmarks, and ranks first or second on eleven even when the individual specialists are included. The specialists are discarded after training, leaving one deployable omni-modal model. These results show that coordinating modality-specific teachers is an effective way to improve a shared model while maintaining cross-modal balance.
AI 知識システムにおけるエンティティの重要性の表現: 聴衆評価と構造的権威のデュアルシグナル フレームワーク
AI 知識システムでは、検索、推奨、証拠の選択、知識集約型の推論のためにエンティティの重要性を表現する必要があります。しかし、重要性は多くの場合、人間の反応またはグラフ構造から導き出される単一のスコアに還元されます。このような圧縮により、AI システムがさまざまなタスクのエンティティの中から選択する必要がある場合に重要な区別が無視される可能性があります。この研究では、各エンティティが視聴者評価の次元と構造的権威の次元によって特徴付けられる、解釈可能な二重信号表現を導入しています。このフレームワークは、経験的検証ドメインとして映画エンティティを使用して評価されます。 IMDb の非営利データセットは評価ベースの視聴者ランキングを提供し、ウィキデータはエンティティのアライメントをサポートし、英語版ウィキペディアのハイパーリンクは PageRank が構造的権威を推定する知識ネットワークを形成します。 482 個のエンティティと 13,690 個の有向関係に関する実験により、2 つの次元間の統計的に有意ではあるが弱い関連性が明らかになりました (Spearman rho = 0.2275、p < 0.001)。それらの重複は上位 10 位ではわずか 10%、上位 100 位では 34% ですが、エンティティレベルの相違は両方向に発生します。この結果は、視聴者の評価と構造的権威は非冗長シグナルであり、重要性に関する単一のスカラー概念に自動的に集約されるべきではないことを示しています。この貢献は、新しいランキング アルゴリズムや学習された埋め込みではなく、最小限の知識表現フレームワークとその次元の必要性の経験的テストです。この調査結果は、コンテキスト固有の選択や集計を適用する前に、明確な重要性シグナルを保存するタスク認識 AI 知識システムをサポートしています。
原文 (English)
Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority
AI knowledge systems require representations of entity importance for retrieval, recommendation, evidence selection, and knowledge-intensive reasoning. Yet importance is often reduced to a single score derived from either human response or graph structure. Such compression may discard distinctions that matter when an AI system must choose among entities for different tasks. This study introduces an interpretable dual-signal representation in which each entity is characterized by an audience-evaluation dimension and a structural-authority dimension. The framework is evaluated using movie entities as an empirical validation domain. IMDb non-commercial datasets provide a rating-based audience ranking, Wikidata supports entity alignment, and English Wikipedia hyperlinks form the knowledge network on which PageRank estimates structural authority. Experiments on 482 entities and 13,690 directed relationships reveal a statistically significant but weak association between the two dimensions (Spearman rho = 0.2275, p < 0.001). Their overlap is only 10% in the top 10 and 34% in the top 100, while entity-level divergence occurs in both directions. The results show that audience evaluation and structural authority are non-redundant signals and should not automatically be collapsed into a single scalar notion of importance. The contribution is not a new ranking algorithm or learned embedding, but a minimal knowledge-representation framework and an empirical test of its dimensional necessity. The findings support task-aware AI knowledge systems that preserve distinct importance signals before applying context-specific selection or aggregation.
SciExplore: 科学的ナビゲーションから情報統合までの自律エージェントの評価
科学研究には、異種ソースにわたる複雑な情報探索と推論のワークフローが含まれます。ただし、既存のベンチマークは主に一般領域の検索や静的な科学的質問応答に重点を置いているため、現実的な科学研究ワークフローに必要な主要な機能を評価できません。 LLM とエージェントの科学的情報探索能力と推論能力を評価するために設計されたベンチマークである SciExplore を紹介します。 SciExplore は、10 を超える科学分野にわたる 103 の専門家が厳選したタスクをカバーする 4 つのタスク タイプで構成されています。科学データベースのナビゲーション、曖昧な文献の検索、欠落している参考文献の補完、およびクロスソースの構造化された知識の統合であり、エンティティ レベルの推論や文書レベルの特定から、証拠レベルの根拠付けやドメイン レベルの統合まで、段階的により高いレベルの能力を調査します。 SciExplore で 10 を超える最先端の LLM と自律エージェントを評価したところ、タスクの複雑さが増すにつれてパフォーマンスが急激に低下し、最も困難な構造化合成タスクでは精度が非常に低いという、大幅なパフォーマンス ギャップが明らかになりました。これらの結果は、現実的な科学情報探索シナリオにおける現在のモデルとエージェントの重大な限界を浮き彫りにしています。
原文 (English)
SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.
クラスター化されたエッジ インテリジェンス: エッジ コンピューティングと AI の単なる融合を超えて
私たちは情報時代から知性の時代に移行しつつあります。 10 年後、あるいはそれよりも短い期間で、データはもはや金ではなくなり、データやネットワークのエッジから得られる情報から導き出されるインテリジェンスが重要になるでしょう。既存のエッジ インテリジェンスの研究は、主に 2 つの方向に焦点を当てています。エッジ リソース管理への AI の使用と、エッジ デバイスへの軽量 AI モデルの展開です。しかし、既存のエッジ コンピューティングの研究には、派生インテリジェンスを、記述、発見、観察、共有、再利用、および異種エッジ デバイスおよびアプリケーション間で動的にクラスタリングできる、独立して管理可能なファーストクラスのエンティティとして扱われる、インテリジェンス中心のフレームワークが欠けています。これらの研究ギャップに対処するために、私たちは先見性のあるインテリジェンス中心のアプローチであるクラスタード エッジ インテリジェンスを導入しました。 CEI の目的は、インテリジェンスを、分散されたエッジとクラウドの連続体全体で独立して表現、発見、観察、交換、管理できる、共有可能で再利用可能な一流のエンティティにすることです。私たちは 3 層の CEI アーキテクチャを提示し、インテリジェンス インベントリ、セマンティック知識表現、コミュニケーション、発見可能性、可観測性、ライフサイクル自動化、クラスタリング メカニズム、市場、相互運用性、標準化などを実現するテクノロジーと研究の側面を検討します。
原文 (English)
Clustered Edge Intelligence: Beyond Just Convergence of Edge Computing and AI
We are moving from an information age to the age of intelligence. A decade, or possibly less than that, data will not be the gold anymore rather the derived intelligence out of the data and the information we posses from the edge of the network. Existing Edge Intelligence research focuses mainly on two directions: using AI for edge resource management and deploying lightweight AI models on edge devices. However, existing edge computing research lacks an intelligence-centric framework in which derived intelligence is treated as a first-class, independently manageable entity that can be described, discovered, observed, shared, reused, and dynamically clustered across heterogeneous edge devices and applications. To address these research gaps, we introduced Clustered Edge Intelligence, a visionary intelligence-centric approach. The aim of CEI is to make intelligence a shareable and reusable first-class entity that can be independently represented, discovered, observed, exchanged, and managed across the distributed edge-cloud continuum. We present a three layer CEI architecture and examine enabling technologies and research dimensions, including intelligence inventories, semantic knowledge representation, communication, discoverability, observability, lifecycle automation, clustering mechanisms, marketplaces, interoperability, and standardization.
スカラーから時系列へ: 時変体積データの暗黙的なニューラル表現を再考する
時間変化する体積データの暗黙的ニューラル表現 (INR) は通常、時空間座標上の高密度サンプリングを使用してトレーニングされます。各観測値は時空間内の 1 つの点に対応します。この座標的な定式化では、最適化中に大規模なサンプリングが必要となり、計算コストが高くなり、時間構造が非効率的に使用されます。この研究では、この設計の選択を再考し、時変フィールドの学習には高密度の時空間サンプリングが必要ないことを示します。代わりに、データを空間的にインデックス付けされた時系列のコレクションとして表し、座標単位のスカラー サンプルではなく、各空間位置に対するシーケンス レベルの監視を使用して INR をトレーニングします。この再定式化により、高密度の時空間サンプリングの必要性がなくなり、代わりに構造化された方法でその完全な時間的展開から各空間位置が学習されます。我々は、この表現が既存の INR アーキテクチャの範囲と互換性があり、トレーニング コストを大幅に削減しながら、再構築の品質を一貫して向上させることを実証します。さらに、この定式化は専門家の混合アーキテクチャと組み合わせることができ、MoE インスタンス化は基本再定式化と既存の MoE ベースの INR 手法の両方と比較して再構成の品質をさらに向上させ、不均一な時間ダイナミクスの下でより強力な容量割り当てを提供することを示します。
原文 (English)
From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying Volumetric Data
Implicit neural representations (INRs) for time-varying volumetric data are typically trained using dense sampling over spatiotemporal coordinates, where each observation corresponds to a single point in space and time. This coordinate-wise formulation requires extensive sampling during optimization, leading to high computational cost and inefficient use of temporal structure. In this work, we revisit this design choice and show that dense spatiotemporal sampling is not necessary for learning time-varying fields. Instead, we represent the data as a collection of spatially indexed time series and train INRs using sequence-level supervision over each spatial location, rather than coordinate-wise scalar samples. This reformulation eliminates the need for dense spatiotemporal sampling and instead learns each spatial location from its full temporal evolution in a structured manner. We demonstrate that this representation is compatible with a range of existing INR architectures and consistently improves reconstruction quality, while significantly reducing training cost. Furthermore, we show that this formulation can be combined with mixture-of-experts architectures, and that our MoE instantiation further improves reconstruction quality compared to both the base reformulation and existing MoE-based INR methods, providing a stronger capacity allocation under heterogeneous temporal dynamics.
ストレージではなく配信: コーディング エージェントのハーネス プロパティとしてのキュー アンカー型ワーキング メモリ
コーディング エージェントには、ドキュメントという 1 種類のメモリが付属しています。命令ファイル、計画成果物、および自動書き込みメモリ ディレクトリは、意図的に作成され、意図的に取得されます。エージェントは、それらを書き込むことを選択し、それらを読み取ることを選択する必要があります。人間の専門知識は、決して文書化されることのない第 2 層で実行されます。それは、状況に縛られた運用上の事実 (落とし穴、場所、地域の慣例) であり、作業の副作用としてコード化され、状況がきっかけになったときに無意識に検索されます。私たちは、この 2 番目の層は長期実行エージェントにとって負荷がかかる層であり、エージェントの選択ではなく、ハーネス プロパティである必要があると主張します。 (1) メモリのオフロード、偶発的エンコーディング、およびイベントベースのプロスペクティブ メモリに関する認知文献に基づいた 2 層設計理論。それぞれがアーキテクチャ要件にマッピングされています。 (2) キューアンカー型記憶モデル。記憶は、構成可能な語彙 (パス、シンボル、意味、イベント、時間) にわたって第一級のトリガー条件を保持し、ハーネスによって決定論的に評価されます。この構成は、調査対象の学術システムや出荷されたシステムでは提供されません。 (3) 実際のコーディング タスクの制御された評価では、事前シードされたストア (114 ターンでメモリ操作が 0) であっても自発的なメモリ使用量がほぼゼロであること、シードされた実行ごとに決定論的な注入が誤報ゼロで配信されること、およびセッション内の再読み取りの 39% が圧縮境界前に支払われたコンテンツの再購入であることを示しています。 (4) 繰り返される圧縮の減衰プローブ: 会話内でのみ保持されている 10 個のファクトは、最初のサマリーで消失し、108 個の圧縮のうち 106 個には存在しません。剥奪されたエージェントはハーネス独自のセッション ファイルを grep して再構築しますが、ハーネス所有のストアから注入された同じファクトは、最終的なサマリーには何も含まれていないため、138 個のコンパクト履歴書すべてを通じて無傷で到着します。ストレージではなく配信が製品です。エージェントにとって信頼できるメモリ チャネルは、エージェントが決して考える必要のないものです。
原文 (English)
Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents
Coding agents ship with one kind of memory: documents. Instruction files, plan artifacts, and auto-written memory directories are deliberately authored and deliberately retrieved: the agent must choose to write them and choose to read them back. Human expertise runs on a second tier that never gets written down: situationally-bound operational facts (gotchas, locations, local conventions) encoded as a side effect of the work and retrieved involuntarily when the situation cues them. We argue this second tier is the load-bearing one for long-running agents and must be a harness property, not an agent choice. We contribute: (1) a two-tier design theory grounded in the cognitive literature on memory offloading, incidental encoding, and event-based prospective memory, each mapped to an architectural requirement; (2) a cue-anchored memory model where memories carry first-class trigger conditions over a composable vocabulary (path, symbol, semantic, event, temporal), evaluated deterministically by the harness, a composition no surveyed academic or shipped system provides; (3) a controlled evaluation on a real coding task showing that voluntary memory use is near zero even with a pre-seeded store (0 memory operations in 114 turns), that deterministic injection delivered in every seeded run with zero false alarms, and that 39% of intra-session re-reads re-buy content paid for before a compaction boundary; (4) a repeated-compaction decay probe: ten facts held only in conversation vanish at the first summary and stay absent from 106 of 108 compactions, and the deprived agent greps the harness's own session files to rebuild them, while the same facts injected from a harness-owned store arrive intact through all 138 compact-resumes as the final summary carries none. Delivery, not storage, is the product: the reliable memory channel for agents is the one the agent never has to think about.
独立した最適化を超えて: マルチモーダル エッジ インテリジェンスにおける圧縮、MoE ルーティング、量子化インタラクション
効率的なマルチモーダル推論は、モデルの品質や FLOP 数だけでなく、レイテンシ、メモリ、エネルギーの制約の下でマルチモーダル表現を保存、移動、ルーティング、キャッシュ、量子化するコストによっても制約が増えています。このペーパーでは、ビジュアル トークン圧縮、ビデオ トークン管理、KV キャッシュの最適化、Mixture-of-Experts (MoE) ルーティング、低ビット量子化、エッジ展開、およびハードウェアを意識したベンチマークをカバーする、効率的なビジョン言語モデルとマルチモーダル大規模言語モデルの最近の進歩をレビューします。私たちは、これらの手法を独立した最適化として扱うことはできないと主張します。視覚的なトークン圧縮は、ダウンストリームの機能分布と MoE ルーティングの決定を変更し、ルーティングの動作はエキスパートの使用率と量子化の感度に影響を与え、量子化されたルーター ロジットはエキスパートの割り当てに影響を与え、KV キャッシュ ポリシーは保持されるマルチモーダル証拠を決定し、ハードウェアの制約により、計算量の節約がメモリと通信のボトルネックに変わることがよくあります。私たちはこれらの相互作用に関する文献を整理し、精度とトークン バジェット、静的圧縮と適応圧縮、スパース ルーティングの効率とエキスパート コラプス、低ビット推論とモダリティ固有の劣化など、主要な設計のトレードオフを特定します。最後に、ビデオ MoE モデルの診断として時間ルーティング一貫性を紹介し、ルーティングを意識した圧縮、クロスモーダル キャッシュ管理、ハードウェアを意識した共同設計、マルチモーダル エッジ インテリジェンスの統合ベンチマークにおけるオープンな研究の方向性を強調します。
原文 (English)
Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence
Efficient multimodal inference is increasingly constrained not only by model quality or FLOP count, but also by the cost of preserving, moving, routing, caching, and quantizing multimodal representations under latency, memory, and energy constraints. This paper reviews recent advances in efficient vision-language and multimodal large language models, covering visual token compression, video token management, KV-cache optimization, Mixture-of-Experts (MoE) routing, low-bit quantization, edge deployment, and hardware-aware benchmarking. We argue that these techniques cannot be treated as independent optimizations. Visual token compression alters downstream feature distributions and MoE routing decisions, routing behavior affects expert utilization and quantization sensitivity, quantized router logits influence expert assignment, KV-cache policies determine retained multimodal evidence, and hardware constraints often transform computational savings into memory and communication bottlenecks. We organize the literature around these interactions and identify key design trade-offs, including accuracy versus token budget, static versus adaptive compression, sparse routing efficiency versus expert collapse, and low-bit inference versus modality-specific degradation. Finally, we introduce Temporal Routing Consistency as a diagnostic for video MoE models and highlight open research directions in routing-aware compression, cross-modal cache management, hardware-aware co-design, and unified benchmarking for multimodal edge intelligence.
GuardianAgentBench: エージェントが失敗する場所とそれらを保護する方法
大規模な言語モデルのエージェントがツールや外部環境にアクセスして自律的に動作することが増えているため、エージェントの安全で信頼性の高い動作を確保することが重要になっています。ここでは、実稼働対応の 3 つのフレームワーク (LangChain、LlamaIndex、Vectara) で評価された 6 つのドメインにわたる 580 のシナリオのベンチマークである GuardianAgentBench (GABench) を紹介します。このベンチマークには、厳格な多段階検証と 5 つの敵対的攻撃モードが組み込まれています。 6 つの最先端モデルを使った実験では、最も強力な構成でも全体の精度は 74.8% しか達成できず、2 つの異なる障害状況が明らかになりました。つまり、より強力なモデルは必要なツールの呼び出しを下回っていますが、弱いモデルはツールの選択を誤って過剰呼び出しを行っています。パフォーマンスはツールセットのサイズと連続した回転深さの両方で単調に低下し、長期的な計画ではボトルネックがさらに深刻になることがわかります。当社のガードレール実装は、すべてのモデルにわたってシステム プロンプト ベースの防御を常に上回っており、わずか 0.5% の誤検知率で障害の 19.9% を回復します。これらの結果は、実行時の構造的介入により、エージェントの正しい動作を妨げることなく安全性が向上することを示しています。
原文 (English)
GuardianAgentBench: Where Agents Fail and How to Guard Them
As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical. We present GuardianAgentBench (GABench), a benchmark of 580 scenarios across six domains evaluated on three production-ready frameworks: LangChain, LlamaIndex, and Vectara. The benchmark incorporates rigorous multi-stage validation and five adversarial attack modes. Experiments with six state-of-the-art models reveal that even the strongest configuration achieves only 74.8% overall accuracy and expose two distinct failure regimes: stronger models under-call required tools, while weaker models mis-select and over-call tools. Performance degrades monotonically with both tool-set size and sequential turn depth, with long-horizon planning proving the steeper bottleneck. Our guardrail implementation consistently outperforms system-prompt-based defenses across all models, recovering 19.9% of failures at a false positive rate of just 0.5%. These results demonstrate that execution-time structural intervention improves safety without disrupting correct agent behavior.
ワークフローに基づいたメカニズムの学習: 構造化されたエージェント スキルのための帰属に基づく修復と知識の再利用
エージェント スキルは、再利用可能な手続き型知識を、凍結された言語モデル エージェントの外部成果物としてパッケージ化しますが、既存のオプティマイザーは、ワークフローのどこで障害が発生するか、どのメカニズムが原因であるか、サードパーティ スキルからの関連する知識をローカルでどのように再利用するかを共同で解決することはできません。ワークフローローカライズドメカニズム学習 (WML) を導入します。そのノード - メカニズム属性は、失敗したワークフロー ノード、関係するメカニズム、および最小の有効な編集ターゲットを特定し、単一メカニズムの欠陥を L3 リソースにルーティングし、メカニズム全体のリレーショナル欠陥を L2 構成プロトコルにルーティングします。次に、6 つのモジュールからなる Workflow-Guided Skill Optimization (WGSO) ループにより、来歴とスコープを認識したサードパーティのナレッジが選択され、境界付きパッチが適用され、候補が評価され、検証された結果がオプティマイザー側のメモリに保存されます。 SpreadsheetBench では、WML は DeepSeek と Qwen3.6-Flash でそれぞれ 90.33 +/- 1.53 と 74.67 +/- 3.51 のハード精度に達しました。追加の最適化を行わない場合、学習したスキルは 84.00 +/- 2.00 および 83.00 +/- 2.00 の表記精度で WikiTableQuestions に転送されます。 Compiler-Supported50 では、WML は最高のハードパス率と成功したタスクあたりの最低コストの両方を達成します。コンパイルされた実行では、成功したタスクのほとんどを保持しながら、直接 SkillAgent に比べてトークンと呼び出しが大幅に削減されます。コードと成果物は https://github.com/xiaolin9595/workflow-localized-mechanism-learning で入手できます。
原文 (English)
Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills
Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolve where a failure occurs in a workflow, which mechanism caused it, and how relevant knowledge from third-party Skills should be reused locally. We introduce Workflow-Localized Mechanism Learning (WML). Its Node--Mechanism Attribution identifies the failed workflow node, implicated mechanisms, and smallest valid edit target, routing single-mechanism defects to L3 resources and relational defects across mechanisms to L2 composition protocols. A six-module Workflow-Guided Skill Optimization (WGSO) loop then selects provenance- and scope-aware third-party knowledge, applies bounded patches, evaluates candidates, and stores verified outcomes in optimizer-side memory. On SpreadsheetBench, WML reaches 90.33 +/- 1.53 and 74.67 +/- 3.51 Hard Accuracy with DeepSeek and Qwen3.6-Flash, respectively; without additional optimization, the learned Skills transfer to WikiTableQuestions with 84.00 +/- 2.00 and 83.00 +/- 2.00 Denotation Accuracy. On Compiler-Supported50, WML attains both the highest hard-PASS rate and the lowest cost per successful task; compiled execution sharply reduces tokens and calls relative to a direct SkillAgent while retaining most of its successful tasks. Code and artifacts are available at https://github.com/xiaolin9595/workflow-localized-mechanism-learning.
Naju: 長期シーケンス メモリの独立した保持と書き込みを備えたネイティブの離散状態空間モデル
長期シーケンス メモリの追跡では、反復状態に対して 2 つの相反する要求が課せられます。それは、保存されたバインディングを長期間にわたってほぼ損失なく保持することと、古いバインディングをアクティブに上書きすることです。当社の診断スイートでは、最も強力で効率的なベースラインは、片側のみをうまく解決する傾向があります。 Mamba などの連続時間パラメータ化状態空間モデル (SSM) は、連続時間システムのゼロ次ホールド離散化によって離散漸化式を取得します。私たちは、この迂回はメモリ追跡には不要であると主張し、離散遷移を直接パラメータ化します。 Naju (ネイティブ アダプティブ ジャンクション ユニット) は、反復的な更新 (図式的には $x_n = f_n\odot x_{n-1} + i_n\odot(B_n u_n)$) を明示的な離散極 (学習された忘却ゲート $f_n$)、独立した書き込みゲイン $i_n$、および入力依存の書き込み/読み取りマップに因数分解します。シグモイド極は $0<1$ を満たすため、各固定局所座標は構造的にシュール安定であり、完全な時変漸化式は安定性正則化装置なしで均一有界性仮定の下でフェージングメモリ/BIBO 限界を満たします。結合設計の主要な構造的制限を形式化します。非拡張的な相補的シングルゲート反復は、$|r|+w\le 1$ を介して実効リテンション $r$ と書き込みゲイン $w$ を結び付けるため、ほぼ完全なリテンションにより弱い書き込みが強制されます。 $f_n$ を $i_n$ から分離すると、この制約が削除されます。経験的に、Naju は、トレーニング長の 4 倍でも保持と上書きの両方で強力なままである唯一の評価済みモデルです。診断スイートを超えて、WikiText-103 言語モデリング、Long Range Arena、およびマルチクエリ連想再現に関して Naju を評価します。これらの設定全体にわたって、Naju は強力な長距離メモリと競争力のあるまたは優れたパフォーマンスを一貫して組み合わせており、主な比較で Mamba のベースラインを上回り、同時に Transformer との競争力を維持し、線形時間、線形メモリ スケーリングを維持します。
原文 (English)
Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory
Long-sequence memory tracking places two opposing demands on a recurrent state: near-lossless retention of stored bindings over long horizons, and active overwriting of stale ones. In our diagnostic suite, the strongest efficient baselines tend to solve only one side well. Continuous-time-parameterized state-space models (SSMs) such as Mamba obtain their discrete recurrence by zero-order-hold discretization of a continuous-time system; we argue that this detour is unnecessary for memory tracking and parameterize the discrete transition directly. Naju (Native Adaptive Junction Unit) factorizes the recurrent update, schematically $x_n = f_n\odot x_{n-1} + i_n\odot(B_n u_n)$, into an explicit discrete pole (a learned forget gate $f_n$), an independent write gain $i_n$, and input-dependent write/read maps. Since the sigmoid pole satisfies $0<1$, each frozen local coordinate is Schur-stable by construction, and the full time-varying recurrence satisfies a fading-memory/BIBO bound under uniform boundedness assumptions, with no stability regularizer. We formalize the key structural limitation of coupled designs: any non-expansive complementary single-gate recurrence ties the effective retention $r$ and write gain $w$ through $|r|+w\le 1$, so near-complete retention forces weak writing; decoupling $f_n$ from $i_n$ removes this constraint. Empirically, Naju is the only evaluated model that remains strong on both retention and overwriting at 4x the training length. Beyond the diagnostic suite, we evaluate Naju on WikiText-103 language modeling, Long Range Arena, and multi-query associative recall. Across these settings, Naju consistently combines strong long-range memory with competitive or superior performance, outperforming the Mamba baselines in the principal comparisons while remaining competitive with the Transformer and preserving linear-time, linear-memory scaling.
ゼロショット要約の再検討: LLM サマライザーの信頼性の実証的調査
大規模言語モデル (LLM) を使用したゼロショット要約は、一貫性のある流暢な要約を生成することにより、抽象的な要約タスクを大幅に進歩させました。ただし、大規模な言語モデルの根底にある確率性により、LLM が生成する要約の安定性と信頼性について懸念が生じます。この問題は、学生や研究者が複雑な学術資料をゼロショット方式で要約する、LLM によって生成される要約が教育現場で普及しているため、ますます重要になっています。生成された要約の安定性に基づいて LLM サマライザーをベンチマークするための新しい 2 レベルの診断プロトコルを提案します。下位レベルでは、制御された環境下で生成された複数の LLM サマリーに対してドキュメント レベルの安定性分析が実行され、安定性係数が計算されます。生成された各要約は、元の文書との意味的および事実的な整合性についてスコア付けされ、複数の側面に沿った安定性の推定が可能になります。次のレベルでは、コーパスから抽出された文書の層別サンプルからの観察結果が統合され、信頼性の代用となる LLM サマライザーの安定性指数が推定されます。 3 つのジャンルの文書にわたる 3 つの LLM サマライザーの実証的調査により、要約評価指標全体にわたる LLM 間の世代レベルの変動に統計的に有意な差があることが明らかになりました。この研究は、LLM 要約の安定性の問題を証拠的に認識することによって LLM 要約研究を前進させ、堅牢で信頼性の高い LLM 要約ツールの開発に向けたさらなる研究を動機づけます。
原文 (English)
Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers
Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summarization task by producing coherent and fluent summaries. However, underlying stochasticity of the large language models raises concerns about the stability and trustworthiness of the LLM-generated summaries. This issue has become increasingly important due to proliferation of LLM-generated summaries in educational settings, where students and researchers summarize complex academic materials in zero-shot manner. We propose a novel two-level diagnostic protocol for benchmarking LLM-summarizers based on the stability of the generated summaries. At the lower level, document-level stability analysis is performed over multiple LLM-summaries generated under controlled environment, and the stability coefficient is computed. Each generated summary is scored for semantic and factual alignment with the original document, enabling estimation of stability along more than one dimensions. At the next level, observations from a stratified sample of documents drawn from the corpus are consolidated to estimate the stability index of the LLM-summarizer, which is the proxy for its trustworthiness. Our empirical investigation of three LLM-summarizers across three genres of documents reveals statistically significant differences in the generation-level variability among LLMs across summary evaluation metrics. This study advances the LLM-summarization research by evidential recognition of the stability problem in LLM-summaries and motivates further research towards development of robust, reliable and trustworthy LLM-summarizers.
EmoAgent-R1: 強化学習ベースの動的エージェント特化によるマルチモーダルな感情理解に向けて
マルチモーダル大規模言語モデル (MLLM) は、マルチモーダル感情認識 (MER) タスクで目覚ましいパフォーマンスを達成し、MER を高度なビデオ理解能力と自然言語記述による複雑な感情の理解という新しいレベルに引き上げました。ただし、既存の MLLM ベースの方法では、多くの場合、マルチモーダル入力における感情ソースの動的性と複雑性を無視して、感情を認識するために固定プロンプトが使用されます。これらの問題に対処するために、強化学習に基づく動的エージェント特化フレームワーク (\textbf{EmoAgent-R1}) を提案し、MLLM の感情認識、推論、汎化能力を最適化します。具体的には、まずコールド スタート戦略を採用し、合成応答条件付き思考連鎖データとエージェント ルーティング データを使用してトレーニングすることで、MLLM に予備的な感情認識、推論、エージェント ルーティング能力を与えます。次に、強化学習を使用して MLLM をさらにトレーニングし、エージェントの選択とエージェントの特化を備えた 2 段階のエージェント ワークフローで感情を認識します。 EmoAgent-R1 を効果的にトレーニングするために、グループベースの相対的な利点と PMI からインスピレーションを得たプログレッシブ トークン レベルの変調を組み合わせて、まばらな報酬を粒度の細かい学習信号に変換し、GRPO における粗粒度の均一なクレジット割り当ての問題を軽減する、新しいプログレッシブ グループ相対ポリシー最適化 (P-GRPO) を提案します。 MER ベンチマークに関する広範な実験により、より強力な感情推論パフォーマンスと最適化の安定性の向上における EmoAgent-R1 の優位性が実証されました。
原文 (English)
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural language description. However, existing MLLM-based methods often use a fixed prompt to perceive the emotions, ignoring the dynamicity and complexity of the emotion source in the multimodal inputs. To address these issues, we propose a novel Reinforcement Learning-based Dynamic Agent Specialization framework (\textbf{EmoAgent-R1}) to optimize the emotion recognition, reasoning, and generalization abilities of an MLLM with dynamic agent specialization based on reinforcement learning. Specifically, we first adopt a cold start strategy to endow an MLLM with preliminary emotion recognition, reasoning, and agent routing ability by training with synthetic answer-conditioned chain-of-thought data and agent routing data. Then, we further train the MLLM with reinforcement learning to perceive emotions in a two-step agentic workflow with agent selection and agent specialization. To effectively train EmoAgent-R1, we propose a novel Progressive Group-Relative Policy Optimization (P-GRPO) to combine group-based relative advantages with a PMI-inspired progressive token-level modulation to transform sparse rewards into fine-grained learning signals, mitigating the coarse-grained uniform credit assignment issue in GRPO. Extensive experiments on MER benchmarks demonstrate the superiority of our EmoAgent-R1 in stronger emotion reasoning performance and improved optimization stability.
HiMe: ウェアラブル デバイスを使用した健康に関する洞察のためのリアルタイムのセルフホスト型パーソナル エージェント プラットフォーム
スマートウォッチなどのウェアラブル健康信号分析に対する従来のアプローチは、厳格な分析フレームワークと限られたパーソナライゼーションによって制約されています。 LLM エージェントの出現により、個人健康エージェント分析の新たな機会が生まれ、状況に応じて健康に関する洞察を適応的に生成できます。しかし、現時点では、プライバシーを保護しながら個人の健康データをリアルタイムで処理できる、ローカルに展開可能なオープンソースのプラットフォームはありません。私たちは、さまざまなウェアラブル デバイスにわたるリアルタイムの健康データ エコシステムと完全に互換性のある、ローカルに展開可能なプライバシー最優先のエージェント プラットフォームである HiMe を紹介します。 HiMe は 3 つの設計原則に従っています。データベースはファーストクラスのコンポーネントとして扱われます。効果と効率が同時に最適化され、低コストのパレート最適バランスが実現されます。データはリアルタイムで処理され、ユーザーは長期にわたってモデル化されます。これらの原則を組み合わせることで、個人がパーソナル ヘルス エージェントを利用して継続的かつ個別に健康状態をモニタリングし、より良い幸福を得ることが現実的になります。
原文 (English)
HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices
Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid analytical frameworks and limited personalisation. The emergence of LLM agents creates a new opportunity for Personal Health Agentic Analysis, where health insights can be generated adaptively and in context. However, currently there is no open-source locally deployable platform capable of processing personal health data in real time while preserving privacy. We present HiMe, a locally deployable, privacy-first agent platform that is fully compatible with real-time health data ecosystems across a wide range of wearable devices. HiMe is guided by three design principles. The database is treated as a first-class component. Effectiveness and efficiency are jointly optimised to achieve a low-cost Pareto-optimal balance. Data are processed in real time while the user is modelled over the long term. Together, these principles make it practical for individuals to harness Personal Health Agents for continuous, personalised health monitoring for better wellbeing.
Faster IndexTTS-2: GPU での自己回帰ゼロショット テキスト読み上げ合成の高速化とストリーミング
自己回帰テキスト読み上げモデルは、強い自然性を実現しますが、トークンが順次生成されるため推論が遅くなり、低遅延を必要とする運用アプリケーションへの展開が制限されます。 IndexTTS-2 は、GPT、フローマッチング拡散変換器、およびボコーダーで構成される最先端の自己回帰 TTS モデルです。高い合成品質にもかかわらず、その推論速度は、ストリーミングまたはバッチ処理のサポートなしではほとんどリアルタイムに達しません。 Faster IndexTTS-2 を紹介します。これは、NVIDIA TensorRT および TensorRT-LLM を使用して、GPU 上で実稼働環境にデプロイするための IndexTTS-2 のすべてのニューラル ネットワーク コンポーネントを高速化します。 Faster IndexTTS-2 により、レイテンシの影響を受けやすい対話型アプリケーションのストリーミング合成や、GPU 使用率を最大化するためのすべてのコンポーネントにわたるバッチ推論も可能になります。英語と中国語の両方に対する Seed-TTS ベンチマークの実験では、単語誤り率、話者の類似性、自然さの低下を最小限に抑えながら、自己回帰 GPT で最大 5.0$\times$、エンドツーエンドで 3.6$\times$ の高速化が実証されました。私たちの方法論は、GPU 上で同様の自己回帰音声モデルを効率的に高速化するための実用的なリファレンスを提供します。
原文 (English)
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs
Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deployment in production applications that require low latency. IndexTTS-2 is a state-of-the-art autoregressive TTS model consisting of a GPT, a flow-matching Diffusion Transformer, and a vocoder. Despite its high synthesis quality, its inference speed barely reaches real-time without streaming or batching support. We present Faster IndexTTS-2, which accelerates all neural network components of IndexTTS-2 for production deployment on GPUs using NVIDIA TensorRT and TensorRT-LLM. Faster IndexTTS-2 also enables streaming synthesis for latency-sensitive interactive applications, and batched inference across all components to maximize GPU utilization. Experiments on the Seed-TTS benchmark for both English and Chinese demonstrate up to 5.0$\times$ speedup on the autoregressive GPT and 3.6$\times$ end-to-end, with minimal degradation in word error rate, speaker similarity, and naturalness. Our methodology provides a practical reference for efficiently accelerating similar autoregressive speech models on GPUs.
生成された推奨事項はコールドアイテムに到達できますか?セマンティック ID 生成に関する時間的視点
セマンティック ID ベースの生成推奨は、アイテムを共有セマンティック トークンのシーケンスとして表現し、分離されたアイテム ID を超えたトークンの再結合を可能にします。ただし、クローズドワールドの再結合は、必ずしも一時的なオープントークンのコールドスタート誘導を意味するわけではありません。つまり、新しいアイテムが、目に見えないアトミックトークンまたは弱くサポートされている SID パスとともにアイテムカタログに追加されます。この研究では、目に見えるターゲットと見えないターゲットを分離し、トークン レベルでコールド アイテムの到達可能性を診断する絶対時間時間プロトコルに基づいた SID ベースの生成推奨を再検討します。既知/未見ヒット分析、コールドネス分類法、およびオラクルプレフィックス調査を通じて、現在の SID ベースのモデルは、観測されたトークンとプレフィックスによってサポートされる将来のアイテムに到達できることもありますが、目に見えないアトミック トークンとサポートされていない SID パスには苦労することを示します。 SID 生成を階層的なセマンティック バケット化として解釈することで、この境界をさらに説明します。初期のトークンは粗いセマンティック領域を選択し、後のトークンはアイテム固有のパスを絞り込みます。これらの発見は、SID 生成が構成的であるものの、完全に自由に作成できるわけではないことを示しており、より独立した SID 空間、スコアリング ベースのインターフェイス、および動的なテキスト コンテキストにおける将来の方向性を示唆しています。
原文 (English)
Can Generative Recommendation Reach Cold Items? A Temporal Perspective on Semantic-ID Generation
Semantic-ID-based generative recommendation represents items as sequences of shared semantic tokens, enabling token recombination beyond isolated item IDs. However, closed-world recombination does not necessarily imply temporal open-token cold-start induction, where new items enter the item catalog with unseen atomic tokens or weakly supported SID paths. In this work, we revisit SID-based generative recommendation under an absolute-time temporal protocol that separates seen and unseen targets and diagnoses the cold item reachability at the token level. Through seen/unseen-hit analysis, coldness taxonomy, and oracle-prefix probing, we show that current SID-based models can occasionally reach future items supported by observed tokens and prefixes, but struggle with unseen atomic tokens and unsupported SID paths. We further explain this boundary by interpreting SID generation as hierarchical semantic bucketing: early tokens select coarse semantic regions, while later tokens refine item-specific paths. These findings show that SID generation is compositional but not fully open-ended, and suggest future directions in more independent SID spaces, scoring-based interfaces, and dynamic textual context.
AttriMem: エージェントの記憶学習のためのアトリビューションに基づくプロセス フィードバック
LLM エージェントにとって効果的な記憶は非常に重要ですが、それを効果的に構築するのは依然として困難です。メモリ構築ポリシーは、インタラクションが蓄積するにつれてどの情報を抽出、保存、更新、圧縮、または破棄するかを決定します。ヒューリスティック記憶手法は主観的なタスク固有のルールに依存しているため、下流の目標とずれたり、タスク間の適応性が制限されたりする可能性があります。対照的に、RL ベースの手法はタスクのフィードバックから学習しますが、主に結果レベルまたはモジュールレベルの報酬を使用します。これらの粗い信号はタスクの成功を示しますが、どの中間メモリの内容が最終的な答えをサポートしているかを特定できず、きめの細かいクレジット割り当てのボトルネックが生じます。ただし、このようなプロセス フィードバックの構築は、中間記憶の決定には固有のグラウンドトゥルース ターゲットが欠けている一方、適切なクレジットはエージェントの不確実な推論軌道によって変化するため、事前に指定できないため、非常に困難です。我々は、RL を使用してメモリ構築ポリシーを学習するためのアトリビューションに基づくプロセス フィードバック フレームワークである AttriMem を提案します。 AttriMem は、最終的な回答へのトークンレベルの貢献から得られるローカルな報酬でグローバルな結果報酬を強化します。長期対話型質問応答の実験では、AttriMem が検索ベース、ヒューリスティック、RL ベースのベースラインを上回り、ベンチマークと回答モデル全体で一般化され、RL の最適化が安定することが示されました。
原文 (English)
AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract, store, update, compress, or discard as interactions accumulate. Heuristic memory methods rely on subjective, task-specific rules, which can misalign with downstream objectives and limit cross-task adaptability. RL-based methods, by contrast, learn from task feedback but mainly use outcome- or module-level rewards. These coarse signals indicate task success but cannot identify which intermediate memory contents support the final answer, creating a fine-grained credit-assignment bottleneck. However, constructing such process feedback is prohibitively difficult because intermediate memory decisions lack unique ground-truth targets, while the appropriate credit varies with the agent's uncertain reasoning trajectory and therefore cannot be specified in advance. We propose AttriMem, an attribution-guided process-feedback framework for learning memory-construction policies with RL. AttriMem augments the global outcome reward with local rewards derived from token-level contributions to the final answer. Experiments on long-horizon dialogue question answering show that AttriMem outperforms retrieval-based, heuristic, and RL-based baselines, generalizes across benchmarks and answer models, stabilizes RL optimization.
V-DEAL: ビデオの安全性のデキャリブレーションを理解拒否結合障害として診断する
ビデオ大規模言語モデルが現実世界のアプリケーションに導入されることが増えているため、安全性の調整を確保することが重要になっています。直感に反しますが、有害な動画と無害なクエリを組み合わせた方が、同じ動画と明示的に有害なクエリを組み合わせた場合よりも高い攻撃成功率を達成することがわかりました。この脆弱性の根底にあるメカニズムを理解するために、モデルの動作、理解、内部表現にわたってこの障害を共同で分析する 3 レベルの診断フレームワークである V-DEAL を紹介します。 V-DEAL は、認識の失敗を段階的に除外し、モデルの内部拒否傾向を定量化することにより、観察された脆弱性の根底にあるメカニズムを分析するための新しい診断の観点を提供します。 3 つの公開ベンチマークで 6 つのビデオ LLM をテストしたところ、モデルが有害なビデオ コンテンツを 81\% 以上の精度で正しく認識しているにもかかわらず、有害なビデオと無害なクエリを組み合わせた条件下では、平均攻撃成功率が依然として 48.33\% に達していることが観察されました。隠れ状態の分析では、視覚的な理解は文字による理解よりも弱い拒否傾向を引き起こすことがさらに示されています。さらに、攻撃の成功率を平均 48.24 パーセント ポイント低下させ、以前の微調整ベースの手法に匹敵するパフォーマンスを達成する即時インジェクション介入手法を導入し、ビデオ LLM におけるそのような安全性リスクに対処するための効果的かつ実用的な手段を提供します。
原文 (English)
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure
As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates than the same videos paired with explicitly harmful queries. To understand the underlying mechanism of this vulnerability, we present V-DEAL, a three-level diagnostic framework that jointly analyzes this failure across model behaviour, understanding, and internal representations. By progressively ruling out perception failure and quantifying the model's internal refusal tendency, V-DEAL provides a new diagnostic perspective for analyzing the underlying mechanism of the observed vulnerability. We tested six Video LLMs on three public benchmarks and observed that models correctly recognize harmful video content with over 81\% accuracy, yet the average attack success rate still reaches 48.33\% under the condition pairing harmful videos with benign queries. Hidden-state analysis further shows that visual understanding activates a weaker refusal tendency than textual understanding. Furthermore, we introduce a prompt injection intervention method that reduces attack success rates by an average of 48.24 percentage points and achieves performance comparable to prior fine-tuning-based methods, providing an effective and practical means to address such safety risks in Video LLMs.
SafeStep: 虚弱または認知症の高齢者のための AI を活用した旅行支援
英国では100万人以上が虚弱や認知症を患っており、都市環境での移動能力が著しく損なわれている。本稿では、高齢者の移動を支援するAIを活用した旅行システム「SafeStep」について紹介する。 SafeStep の中核となるのは、ルート計画と予測モデリングを統合する、新しい旅行グラフ表現です。行程の各段階で、システムは、(i) LLM と Anticip8 行動予測エンジンの組み合わせを使用してパーソナライズされた失敗シナリオを生成し、(ii) 対象を絞った介入を提案し、(iii) 結果の確率に対する介入の影響を推定します。これにより、SafeStep は、人が目的地に到着する可能性を最大化する介入を選択できるようになります。 SafeStep は、旅行グラフ生成に関する実験と、26 件の実世界の旅行を含むフィールド調査を通じて評価されました。結果は、故障予測のための Anticip8 と介入評価のための GPT ベースのモデルを組み合わせることで、最も信頼性の高いパフォーマンスが得られることを示しました。ユーザーのフィードバックによると、SafeStep は旅行中の信頼性と知覚される安全性を向上させますが、インターフェイスの使いやすさはターゲット層に合わせて改善する必要があります。今後はSafeStepを改良してリリースしていきたいと考えています。 SafeStep 用に開発された AI システムは、メンタルヘルス、キャリアコーチング、依存症治療などの他の分野にも応用できる可能性があります。
原文 (English)
SafeStep: AI-powered Travel Assistance for Elderly People with Frailty or Dementia
More than a million people in the UK suffer from frailty or dementia, which severely compromise their ability to travel in urban environments. This paper presents SafeStep, an AI-driven travel system that assists elderly users with their journeys. At the core of SafeStep is a novel travel graph representation, which integrates route planning with predictive modelling. For each stage of a journey, the system (i) generates personalized failure scenarios using a combi-nation of LLMs and the Anticip8 behavioral prediction engine, (ii) proposes targeted interventions, and (iii) estimates the impact of interventions on out-come probabilities. This enables SafeStep to select interventions that maximize the likelihood of the person reaching their destination. SafeStep was evaluated through experiments on travel graph generation and a field study involving 26 real-world journeys. Results showed that combining Anticip8 for failure pre-diction with GPT-based models for intervention evaluation yields the most re-liable performance. User feedback indicated that SafeStep improves confidence and perceived safety during travel, although interface usability needs to be im-proved for the target demographic. In the future, we would like to improve and release SafeStep. The AI system that was developed for SafeStep could be ap-plied in other areas, such as mental health, career coaching and addiction treatment.
Speech2Speech LLM アシスタントの保護: 自動車アプリケーションのケーススタディ
最近の進歩により、口調や雰囲気などの非言語的手がかりを含む、自然な対話を生み出すことができるスピーチツースピーチ (S2S) 会話アシスタントが導入されました。自動車分野では、これにより直感的で人間らしい車内対話体験が可能になります。ただし、これらのエンドツーエンドのアシスタントを統合すると、プログラム可能なドメイン固有の保護手段のアーキテクチャ オプションが制限されます。このペーパーでは、S2S ガードレールの 2 つの実装アプローチ、トランスクリプトベースとツールベースについて説明します。実証的評価を通じて、法外なレイテンシ (計算量が少ないチェックであっても各回答が 0 ~ 1.4 秒遅れる) と技術的障害 (潜在的に非決定的なツール呼び出し動作など) により、ほとんどの場合、両方の戦略が産業展開には不十分であることを示しています。最後に、自動車分野における S2S ガードレールの未解決の課題について概説します。
原文 (English)
Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications
Recent advances have introduced speech-to-speech (S2S) conversational assistants capable of producing natural-sounding interactions, including non-verbal cues like tonality and mood. In the automotive domain, this enables intuitive and humanlike in-car dialogue experiences. However, integrating these end-to-end assistants limits architectural options for programmable domain-specific safeguards. This paper discusses two implementation approaches for S2S guardrails: transcript-based and tool-based. Through an empirical evaluation, we demonstrate that both strategies are insufficient for industrial deployment in most cases due to prohibitive latency (delaying each answer by 0 to 1.4 seconds even for computationally cheap checks) and technical impediments (like potentially non-deterministic tool call behavior). Finally, we outline open challenges for S2S guardrails in the automotive context.
ILPによる天気予報の解説
帰納的論理プログラミング (ILP) は、記号学習と宣言的知識表現を組み合わせるためのフレームワークとして、90 年代に論理プログラミング コミュニティ内で誕生しました。現在、成熟した ILP フレームワークが存在し、複雑で非単調な仮説を学習できるため、モデリング機能と ILP の実世界への応用範囲の両方が広がります。この研究は主に FastLAS2 フレームワークに基づいており、イタリアのフリウリ・ヴェネツィア・ジュリア地域の地方気象台である OSMER FVG が発行する気象速報を解明するのに役立つ、シンプルで解釈可能な仮説を生成することを目的としています。このペーパーでは、シミュレートされた気象学の生データと OSMER の速報 (グラウンド トゥルースとして使用) から開始して、データを ASP ファクトとして抽出し、ILP の例を生成するパイプラインを紹介します。このような例から、FastLAS2 を介して説明的な仮説が推測されます。このような仮説(自然言語に翻訳)は、人間の専門家によって発行された天気予報、特に速報ピクトグラム(予報の記号注釈付き気象図)における特定の記号の専門家選択の背後にある理論的根拠を説明します。提案されたアプローチは一般的なものであり、特定の地域に特化したものではなく、他の情報源や異なる地域の速報にも同様に適用できます。
原文 (English)
Explaining Weather Bulletins via ILP
Inductive Logic Programming (ILP) originated within the Logic Programming community in the Nineties as a framework for combining symbolic learning with declarative knowledge representation. Nowadays, mature ILP frameworks exist and they are capable of learning complex, non-monotonic hypotheses, thus broadening both the modeling capabilities and the scope of real-world applications of ILP. This work is primarily based on the FastLAS2 framework and aims to generate simple, interpretable hypotheses to help clarify the weather bulletins issued by OSMER FVG, the Regional Meteorological Observatory of the Italian region of Friuli Venezia-Giulia. In this paper we present a pipeline that, starting from simulated meteorological raw data and from OSMERs' bulletins (used as ground truth), extracts data as ASP facts and generates ILP examples. From such examples an explanatory hypothesis is then inferred via FastLAS2. Such a hypothesis (translated into natural language) explains the weather forecast issued by human experts, and in particular the rationale behind experts' choices of specific symbols in the bulletin pictogram (the symbol-annotated meteorological map of the forecast). The proposed approach is general, not specific to any particular region and it can equally be applied to bulletins from other sources and to different regions.
神経記号システムにおける推論のショートカットを軽減する微分可能論理プログラミング
ニューロシンボリック (NeSy) システムは、ニューラル ネットワークと論理的推論を統合して、一般化と解釈可能性の両方を実現しますが、最近の研究では、これらのシステムがショートカット的な推論動作の影響を受けやすいことが示されています。私たちは、行列ベースの微分可能論理プログラミングを使用して、2 つの現象における推論のショートカットを軽減する新しい方法を提案します。1 つは、意図したタスクを達成せずに制約が満たされる制約満足ショートカットと、論理的に健全な推論にもかかわらず、偏ったデータにより意味的に不正確な概念マッピングが生じる認知ショートカットです。最近の行列ベースのロジック プログラミング セマンティクスに基づいて、単一の行列でのルールと制約の統一エンコーディングなど、ショートカットを軽減する設計要素を導入します。また、ファジー ロジック t ノルムとの関係を特定し、それらの勾配流特性を経験的に比較します。 MNIST バリアントで慎重に設計された実験を通じて、ニューラル出力を論理アトムに 1 対 1 で接地することで、ソフトな確率分布に依存する以前の方法と比較して、両方のショートカット タイプが大幅に減少することを示します。次に、記号知識とニューラル学習を組み合わせる際のアーキテクチャ上の選択が、ショートカットの軽減において重要な役割を果たすことを確認します。
原文 (English)
Differentiable Logic Programming to Mitigate Reasoning Shortcuts in Neurosymbolic Systems
Neurosymbolic (NeSy) systems integrate neural networks with logical reasoning to achieve both generalization and interpretability, but recent work has shown they are susceptible to shortcut reasoning behaviors. We propose a novel method using matrix-based differentiable logic programming to mitigate reasoning shortcuts in two phenomena: constraint satisfaction shortcuts, where constraints are satisfied without achieving the intended task, and cognition shortcuts, where biased data leads to semantically incorrect concept mappings despite logically sound inference. Building on recent matrix-based logic programming semantics, we introduce design elements to mitigate shortcuts, including a unified encoding of rules and constraints in a single matrix. We also identify connections to fuzzy logic t-norms and empirically compare their gradient flow properties. Through carefully designed experiments on MNIST variants, we show that one-to-one grounding of neural outputs to logical atoms significantly reduces both shortcut types compared to previous methods that rely on soft probability distributions. We then confirm that architectural choices in coupling symbolic knowledge with neural learning play a critical role in shortcut mitigation.
機械学習を使用した単一定数乗算の効率的な SAT エンコーディングのための適切なルールの特定
単一定数乗算問題は、ハードウェア設計における基本的な NP ハード最適化タスクであり、加算、減算、およびビット シフトのみを使用して固定定数を分解しようとします。動的プログラミング手法は SCM に最適に近い SAT エンコーディングを生成できますが、定数が大きい場合はエンコーディング コストが依然として高くなります。我々は、分解中に演算子選択を導くための適切なルールを特定することにより、SCM SAT エンコーディングを高速化する神経記号フレームワークを提案します。私たちのアプローチでは、グラフ ニューラル ネットワーク モデルを使用して、定数分解から有望な演算子の種類を予測し、結果として得られる信頼スコアを利用して、シンボリック検索で不適切な選択肢を排除します。目に見えない 17 ~ 32 ビット定数に関する実験結果では、加算に関して最適に近いエンコード品質を維持しながら、エンコード時間の 1 ~ 2 桁の削減、メモリ使用量の 97% 以上の削減、および分岐の桁数の減少が実証されました。これらの結果は、学習ガイド付きシンボリック戦略により、SCM エンコーディングのスケーラビリティと効率が大幅に向上する可能性があることを示しています。私たちのコードとデータは、https://github.com/Chufeng-Jiang/SCM_MLDP で公開されています。
原文 (English)
Identifying Good Rules for Efficient SAT Encodings of Single-Constant Multiplication Using Machine Learning
The Single Constant Multiplication problem is a fundamental NP-hard optimization task in hardware design, which seeks to decompose a fixed constant using only additions, subtractions, and bit-shifts. Although dynamic programming methods can produce near-optimal SAT encodings for SCM, their encoding cost remains high for large constants. We propose a neuro-symbolic framework that accelerates SCM SAT encoding by identifying good rules for guiding operator selection during decomposition. Our approach employs a graph neural network model to predict promising operator types from constant decompositions, and exploits the resulting confidence scores to prune no-good choices in the symbolic search. Experimental results on unseen 17-32 bit constants demonstrate one to two orders of magnitude reductions in encoding time, over 97% reduction in memory usage, and an order-of-magnitude decrease in branching, while preserving near-optimal encoding quality in terms of additions. These results show that learning-guided symbolic strategies can significantly improve the scalability and efficiency of SCM encoding. Our code and data are publicly available at: https://github.com/Chufeng-Jiang/SCM_MLDP
差分制約を使用した解答セットプログラミングの境界ありセマンティクス: 暫定レポート
線形制約の統合により、Answer Set Programming (ASP) の範囲が大幅に拡大されましたが、既存のハイブリッド ソルバーは、統一された論理基盤を欠いた異種のセマンティック基盤に依存することがよくあります。私たちは、Bound-founded Logic of Here-and-There (HTb) のさまざまなバリエーションを導入することでこのギャップに対処し、線形制約を伴う ASP の拡張に対して、幅広い代替セマンティクスにわたって平衡モデルを特徴付けることができる多用途のフレームワークを提供します。我々は、 clingo[DL] の意味論的な特徴付けに焦点を当てて、このフレームワークを差分制約の設定に適用します。私たちのアプローチの中心となるのは、数値変数の基礎の形式化です。 clingo[DL]、clingcon、flingo などのさまざまなハイブリッド システムが制約アトムをどのように正当化するかを調査することで、それらのさまざまな動作の意味論的なルーツを明らかにします。この調査の結果、 clingo[DL] のような現在のシステムの基礎を形式化するだけでなく、プログラムの簡素化に関する厳密な研究や、将来の多様なセマンティック原則の統合も促進する、単一の一貫したフレームワークが得られます。
原文 (English)
Bound-Founded Semantics for Answer Set Programming with Difference Constraints: Preliminary Report
While the integration of linear constraints has significantly expanded the reach of Answer Set Programming (ASP), existing hybrid solvers often rely on disparate semantic underpinnings that lack a unified logical foundation. We address this gap by introducing a many-sorted variant of the Bound-founded Logic of Here-and-There (HTb), providing a versatile framework capable of characterizing equilibrium models across a wide spectrum of alternative semantics for extensions of ASP with linear constraints. We apply this framework to the setting of difference constraints, focusing on the semantic characterization of clingo[DL]. Central to our approach is the formalization of foundedness for numeric variables. By investigating how different hybrid systems - such as clingo[DL], clingcon, and flingo - justify constraint atoms, we uncover the semantic roots of their varying behaviors. This investigation results in a single, consistent framework that not only formalizes the foundations of current systems like clingo[DL] but also facilitates the rigorous study of program simplifications and the future integration of diverse semantic principles.
記述ロジック プログラム向けに十分にサポートされた新しいセマンティクス
記述ロジック プログラムは、ルールとオントロジーを組み合わせるための強力な形式主義です。記述ロジック プログラムのセマンティクスが十分にサポートされているため、応答セットが循環依存関係に依存しないことが保証されます。論理プログラミングの最も一般的なセマンティクスには、十分にサポートされているというこの特性があります。現在十分にサポートされている DL プログラムのセマンティクスには 2 つの制限があることを私たちは認識しています。それは、一貫性の問題による計算の複雑さの増大と、リダクト変換の特性評価の欠如です。この研究では、現在の意味論よりも厳密に存在論的アトムを評価する新しい意味論を提示します。これにより、一貫性問題の複雑さが多項式階層の第 2 レベルに増加するのではなく、NP 完全に維持されます。さらに、新しいセマンティクスが現在のセマンティクスと同等である記述ロジック プログラムの構文クラスを特定します。固定点演算子とリダクトベースの変換を使用してセマンティクスを特徴付けます。私たちの新しいセマンティクスは、現在の十分にサポートされているセマンティクスの厳密なサブセットであるため、独自のより厳密な概念を導入しながら、十分にサポートされているという以前の概念を維持します。私たちは、ロジック プログラミングとの類似性により、十分にサポートされているという新しい概念を好みます。
原文 (English)
A New Well-Supported Semantics for Description Logic Programs
Description logic programs are a powerful formalism for combining rules with ontologies. The well-supported semantics for description logic programs ensures that no answer sets rely on cyclic dependencies. Most popular semantics for logic programming have this property of well-supportedness. We recognize two limitations of the current well-supported semantics for DL programs: its increased computational complexity for the consistency problem and its lack of a reduct transformation characterization. In this work, we present a new semantics which evaluates ontological atoms more strictly than the current semantics. This keeps the complexity of its consistency problem NP-complete, rather than increasing it to the second level of the polynomial hierarchy. Additionally, we identify a syntactic class of description logic programs for which our new semantics is equivalent to the current semantics. We characterize our semantics using a fixpoint operator and a reduct-based transformation. Our new semantics is a strict subset of the current well-supported semantics, so it maintains the prior notion of well-supportedness while inducing its own stricter notion. We prefer our new notion of well-supportedness due to its similarities with logic programming.
ルールが因果関係の知識をどのように表現するか: 確率的論理プログラミングによる因果モデリング
パールは、因果関係の知識によって介入効果の予測が可能になると主張しているのは有名です。対照的に、純粋に記述的な知識は、観察から引き出された結論のみをサポートします。しかし、彼の因果関係理論は、もっぱらベイジアン ネットワークと因果モデル内で開発されています。したがって、それは主に非循環的な因果関係に限定されており、そのアイデアを他の形式主義に移すと誤解や矛盾が生じる危険があります。この論文では、因果関係に対するパールのアプローチを確率的論理プログラミング (PLP) に取り入れます。この目的を達成するために、そのようなプログラムは、時間的概念に依存しない、以前の研究で確立された哲学的基礎と一致しています。つまり、関連するすべてのイベントが同時に発生すると想定されます。これらのプログラムの正式な因果的意味論が、介入と実装の概念とともに提案されています。このセマンティクスは層別 ProbLog プログラムの P-log セマンティクスと一致するが、非層別の場合や他の PLP 形式ではこの 2 つは異なる可能性があることが示されています。
原文 (English)
How Rules Represent Causal Knowledge: Causal Modeling with Probabilistic Logic Programming
Pearl famously argues that causal knowledge enables the prediction of intervention effects. By contrast, purely descriptive knowledge supports only conclusions drawn from observations. His theory of causality, however, is developed exclusively within Bayesian networks and causal models. Consequently, it is largely restricted to acyclic causal relationships, and transferring its ideas to other formalisms risks misinterpretation or inconsistency. This paper brings Pearl's approach to causality into probabilistic logic programming (PLP). To this end, such programs are aligned with philosophical foundations established in prior work that do not rely on temporal notions; that is, all relevant events are assumed to occur simultaneously. A formal causal semantics for these programs, together with a notion of intervention and an implementation, is proposed. It is shown that this semantics coincides with the P-log semantics for stratified ProbLog programs, while the two may differ in the non-stratified case and for other PLP formalisms.
ICAE ベンチ: インタラクティブなプロジェクト ビルダーとしてのコーディング エージェントの評価
最近のバイブコーディング ワークフローの出現により、コーディング エージェントに求められることが変わりつつあります。エージェントは、完全に指定された指示に従って単にコードを完成させるのではなく、計画、要件の明確化、ツールの使用、デバッグ、リポジトリ レベルの構築などのさまざまな能力を組み合わせることによって、不完全な製品意図を動作するソフトウェアに変換することがますます期待されています。しかし、既存のベンチマークはこの変化に完全に追いついていず、静的で完全に指定されたタスクでエージェントを評価しています。このペーパーでは、対話型のプロジェクト構築環境でコーディング エージェントを評価するためのベンチマークである ICAE-Bench を紹介します。基本的なアイデアは、あいまいな製品要件から開始し、自動化されたユーザー エージェントを使用して動的なパラダイムをシミュレートすることです。この設定を現実的かつ評価可能にするために、ICAE ベンチでは 3 つの主要な設計が導入されています。まず、制約のないあいまいな要件の曖昧さを回避するために、各タスクは、実行可能動作を備えた正確な実際のオープンソース リポジトリから曖昧さを導き出します。第 2 に、高品質で再現可能なユーザー シミュレーションを保証するために、ICAE ベンチはユーザー エージェント データを介してインタラクションを確立し、ユーザー エージェントが新しい要件を生み出したり、実装成果物を漏らすことなく隠れた制約を明らかにできるようにします。第三に、オープンエンドのリポジトリを公正に評価するために、ICAE ベンチは、機能の正確性、セマンティックと API の類似性、構造の忠実性、設計品質、インタラクション品質などの多次元診断とともに標準化されたブラックボックス テストを使用します。
原文 (English)
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software by combining various abilities including planning, requirement clarification, tool use, debugging, and repository-level construction. Yet existing benchmarks have not fully caught up with this shift, evaluating agents on static, fully specified tasks. In this paper, we introduce ICAE-Bench, a benchmark for evaluating coding agents under interactive project-building settings. The basic idea is to start from a fuzzy product requirement, simulating the dynamic paradigm with an automated User Agent. To make this setting both realistic and evaluable, ICAE-Bench introduces three key designs. First, to avoid the ambiguity of unconstrained fuzzy requirements, each task derives ambiguity from a precise real open-source repository with executable behavior. Second, to ensure high-quality and reproducible user simulation, ICAE-Bench grounds interaction through User Agent Data, allowing the User Agent to reveal hidden constraints without inventing new requirements or leaking implementation artifacts. Third, to evaluate open-ended repositories fairly, ICAE-Bench uses standardized black-box tests together with multi-dimensional diagnostics, including functional correctness, semantic and API similarity, structural fidelity, design quality, and interaction quality.
因果関係プロセスのロジック プログラミング セマンティクス
生命科学におけるモデリングの困難な問題を動機として、私たちは論理プログラミングのセマンティクスと、それらの論理プログラムと互換性のある因果プロセスの最終状態との関係を調査します。より正確には、正論理プログラムの安定モデルは中立状態から始まり、乱れることなく無限に続くプロセスの最終状態に対応するのに対し、サポートされたモデルは任意の開始点から到達可能な最終状態を記述することを示します。これは、因果律言語としての論理プログラミングの適切なセマンティクスの議論にも貢献し、因果関係の説明の観点から、安定したサポートされているモデル セマンティクスの最近の解釈に時間的観点を追加します。
原文 (English)
Logic Programming Semantics for Causal Processes
Motivated by challenging modelling issues in the life sciences, we investigate the relationship between logic programming semantics and the eventual states of causal processes compatible with those logic programs. More precisely, we show that while stable models of positive logic programs correspond to the eventual states of processes commencing from a neutral state and continuing undisturbed indefinitely, supported models describe the eventual states reachable from arbitrary starting points. This also contributes to the discussion of the appropriate semantics for logic programming as a causal rule language, adding a temporal perspective to recent interpretations of the stable and supported model semantics from an explanatory viewpoint of causality.
BasketEvent: バスケットボールのビデオで誰がいつ何をしたかを理解する
バスケットボールのビデオを包括的に理解するには、どのような出来事が起こったかだけでなく、誰が責任を負うのか、そして重要な証拠がいつ現れるのかを解決する必要があります。しかし、既存の方法は通常、空間認識と意味認識を独立したタスクとして扱い、イベントを個々のプレイヤーに根付かせたり、複雑な集団ダイナミクスの中での時間的境界を正確に特定したりすることができません。このギャップを埋めるために、実際の NBA 放送から厳選されたデータセットを理解するプレーヤー中心のバスケットボール イベントである BasketEvent を紹介します。 BasketEvent では、イベント ラベルは責任のあるプレーヤーに基づいて設定され、正確なイベント間隔を持つ手動で注釈が付けられた 1,000 個のサンプルのサブセットが、一時的な証拠の位置特定を評価するために提供されます。このデータに基づいて、バスケットボールのビデオを一時的な証拠を備えたプレーヤーレベルのイベント予測にマッピングするプレーヤー中心の推論フレームワークである PlayNet を提案します。具体的には、PlayNet は、プレーヤー間、プレーヤーとボール、およびグローバルなコートの相互作用をモデル化することで、主要なエンティティを追跡し、プレーヤーのアイデンティティを関連付け、イベントに関する理由を追跡すると同時に、ゲート プーリングを介してまばらな一時的な証拠を集約します。広範な実験により、PlayNet が代表的なビデオ レベルおよびクロップ ベースのベースラインよりも大幅に優れていることが実証され、きめ細かいスポーツ ビデオを理解するためのプレーヤー中心のモデリングの優位性が証明されています。私たちのデータ、コード、モデルは一般に公開されます。
原文 (English)
BasketEvent: Understanding Who Did What and When in Basketball Videos
Comprehensive basketball video understanding requires resolving not only what event occurs, but also who is responsible and when the key evidence appears. However, exist- ing methods typically treat spatial perception and semantic recognition as isolated tasks, failing to ground events to individual players or pinpoint their temporal boundaries within complex collective dynamics. To bridge this gap, we introduce BasketEvent, a player- centric basketball event understanding dataset curated from real NBA broadcasts. In BasketEvent, event labels are grounded to the responsible players, and a manually an- notated subset of 1,000 samples with precise event intervals is provided to evaluate tem- poral evidence localization. Based on this data, we propose PlayNet, a player-centric reasoning framework that maps basketball videos to player-level event predictions with temporal evidence. Concretely, PlayNet tracks key entities, associates player identities, and reasons about events by modeling player-player, player-ball, and global court inter- actions, while aggregating sparse temporal evidence via gated pooling. Extensive experi- ments demonstrate that PlayNet significantly outperforms representative video-level and crop-based baselines, proving the superiority of player-centric modeling for fine-grained sports video understanding. Our data, code, and models will be made publicly available.
An LLM-Driven Workflow for Automated Process Control Strategy Generation and Tuning from Dynamic Process Models
We present a structured large-language-model-driven workflow for automated multi-variable control design from dynamic process models. The w…
Expert Behavior Prior Reinforcement Learning
Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learn…
Regulating autonomous and agentic AI
Regulating activities where regulatees use autonomous and agentic AI is challenging. Regulatory assumptions about regulatee knowledge and c…
SPORD: A Simulation-Propose-then-OR-Dispose Approach for Supply Chain Planning
For years, supply chain planning at e-commerce firms has operated as a collection of isolated projects. Each planning task from static netw…
Towards Faithful Graph Explanations with Synergistic Edge Effects via Granular Balls
Instance-level explanations aim to reveal the rationale behind a model's decisions for a specific graph. Previous methods explain graph neu…
Multimodal Pretraining for Generalizable EEG Representation Learning
Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks. This limited approach can make it c…
MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamics Learning
Self-supervised foundation models have recently shown strong potential for electroencephalogram (EEG)-based analysis. However, existing app…
Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog
Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning,…
Logical Regression for Planning with Axioms
In automated planning, logical regression is an operation that returns the most general condition necessary for an action to achieve a part…
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories…
Bridging the Gap Between Plausibility and Admissibility: Constraint-Aware Flow Maps for Dynamic Graph Systems
Generative models can support decision-making under uncertainty by producing ensembles of plausible future system trajectories, but statist…
Agent-Guided Relational Concept Discovery: Toward Interpretable Surgical Margin Assessment
Deep learning models can effectively use Rapid Evaporative Ionization Mass Spectrometry (REIMS) data for surgical margin assessment. Howeve…
Detecting LLM-Generated Tokens in Human--LLM Coauthored Text
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM…
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verify…
Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks
Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based…
Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry
AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development…
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their…
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay it…
The Boundaries of Automation: A Theory of Persistent Human Participation
The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever po…
MIRROR: Learning from the Other View for Multi-Modal Reasoning
Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasonin…
OpenForgeRL: Train Harness-native Agents in Any Environment
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, an…
Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning
Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducin…
Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in Ghana
A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to identify atypical transm…
Deblurring in the Wild: A Real-World Image Deblurring Dataset from Smartphone High-Speed Videos
We introduce the largest real-world image deblurring dataset constructed from smartphone slow-motion videos. Using 240 frames captured over…
Through-the-Earth Magnetic Induction Communication and Networking: A Comprehensive Survey
Magnetic induction (MI) communication (MIC) has emerged as a promising candidate for underground communication networks due to its excellen…
From Attention to Frequency: Integration of Vision Transformer and FFT-ReLU for Enhanced Image Deblurring
Image deblurring is vital in computer vision, aiming to recover sharp images from blurry ones caused by motion or camera shake. While deep…
Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations
Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models'internal kno…
Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought
Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box. In this paper,…
More Is Not More: What Matters for Diversity in LLM Opinions?
Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group m…
LLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining
We present LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026 on paragraph-level argument mining in UN and UNESCO r…
Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing
While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alterna…
Break Through the Compression Bottleneck: From Theory to Practice
As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memor…
Making Open-Source Text LLM Watermarks Durable Against Merging
Open-source LLMs (OSMs)arereaching near state-of-the-art performance, prompting prior works to trace the text they generate by embedding te…
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this ass…
Preference Tuning as Spectral Update Reorganization
Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains…
Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility
Proprietary large language models (LLMs) entail substantial intellectual and financial investment, making them valuable intellectual proper…
Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception
Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally ind…
The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs
LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examin…
A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction
Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain…
Response drift across frontier large language models
All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet…
RE-AD: Real-Time Requirement Adherence for Data Labeling
Human-annotated data remains fundamental to training frontier Large Language Models (LLMs). However, crowd-sourced annotations often suffer…
Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc
Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such…
Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention
Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attentio…
CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents
Large language model (LLM) agents operating over extended dialogues accumulate vast amounts of information, yet existing memory systems eit…
THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QA
Multi-hop question answering requires retrieving and integrating evidence from multiple contexts. Despite the rapid progress of current res…
Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-t…
Can Valence Reflect Morality in Natural Language? A Preliminary Annotation Study
Present implementations of artificial intelligence (AI) ethics do not adequately take feelings, or affect, into account. If AI should be al…
Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation
Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational…
PhantomFill: When the Form Demands an Answer, Language Models Invent One
Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the…
The Active Ingredient in Muon's Grokking
The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW. Prior work attributes this to "spectral-norm con…
Scaling Closed-Loop Feature Channel Configuration with LLMs
Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths ca…
Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement
Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs. However, exi…
CLOE: Christoffel Loss Autoencoder for Anomaly Detection
Semi-supervised anomaly detection plays a key role in diverse fields such as process monitoring, healthcare, and finance. However, lightwei…
Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development
Many modern AI systems are designed to operate under diverse, open-ended, use-cases. To help generalize deployed systems, many deployed-sys…
Grounding Investor Views: Neural Predicates in the Black-Litterman Model
Portfolio construction under the Black-Litterman model requires investors to specify views on asset returns alongside explicit uncertainty…
A Graph Neural Network approach to zero-shot Digital Twins
Traditional Predictive Digital Twins often remain geometrically rigid, requiring extensive retraining or fine-tuning whenever the underlyin…
ReliableTableQA:How Much Supervision Does Reliability Annotation Need?
We introduce ReliableTableQA, a framework for training an LLM to annotate the statistical reliability of tabular QA results, not whether th…
Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches
Long-context Transformer inference increasingly relies on KV-cache compression or quantization. Prior rotation and transform-coding results…
Leveraging Biokinetic Knowledge Priors for Data-Scarce Bioprocess Modeling
While deep learning has accelerated drug discovery, its impact on biomanufacturing has been considerably more limited. The reason is data s…
From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime
How should a diffusion model decide which noise levels to train on, and how much? Despite the importance of this choice, current noise sche…
HypNO: A Graph-Based Neural Operator with Physics-Informed Message Passing for Hyperbolic Conservation Laws
We introduce HypNO, a graph-based neural operator for scalar hyperbolic conservation laws. HypNO operates directly on a space-time graph of…
Improving Access to Essential Medicines via Decision-Aware Machine Learning
A critical challenge in healthcare systems in low- and middle-income countries (LMICs) is the efficient and equitable allocation of scarce…
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. W…
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability cha…
Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design
The traditional "one drug, one target" paradigm of structure-based drug design (SBDD) frequently proves inadequate for treating multifactor…
SenCos-GEM: SENet-Calibrated and Law-of-Cosines-Constrained Geometry-Enhanced Molecular Representation for Property Prediction
Effective molecular representation learning is crucial for accurate molecular property prediction. Recently, numerous self-supervised learn…
Monkey King Bang: A Unified Scientific Multimodal Foundation Model
Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar trans…
StabilityBench: Benchmarking Instability in LLMs
AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior r…
Joint Utilization of Geospatial and census proxies for Autoencoder-Assisted Downscaling (JUGAAD) of socioeconomic indicators in India
Monitoring poverty and food security indicators is imperative for addressing socioeconomic challenges in developing nations. A limitation i…
Geometric Configurations of Perturbed Jailbreak Prompts
Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major secur…
Bayesian uncertainty estimation improves clinical decision making in medical AI agents
Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atyp…
Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans
The function of many genes is still unknown, and conventional driver-discovery methods, which rely on how frequently a gene is mutated, can…
When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers
When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm? We answer with four findings from…
RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring
Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic trainin…
Demonstrating GenDB: Instance-Optimized and Customized Query Processing Code Generation via LLM Agents
Traditional query processing engines require continuous development and extensions to support new techniques and user requirements, and in…
Frontier Financial Judgement: Can agents tell what might move a stock?
We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to asse…
Scaling Interpretable Transformers with Parity Bottleneck Layers
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual s…
SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking
Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often u…
Adaptive Multi-Horizon Reinforcement Learning
Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement l…
From Agent Failures to Text Policies: What Works and What Breaks
TextGrad improves language-model systems by revising text from feedback. Its core thesis is that natural-language feedback can act as a gra…
Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis
Concept Bottleneck Models provide interpretable-by-design predictions by mediating diagnosis through human-understandable concepts, but in…
DS@GT ARC at ImageCLEFmed GANs 2026: Geometric Filtering for Privacy-Preserving CT Slice Generation
We present a privacy-preserving framework for synthetic lung CT slice generation developed for the Image-CLEFmed GANs 2026 challenge. The a…
A Framework for Reputation Aware Uninorm-driven Consensus Algorithms for Blockchain Networks
The operation of blockchain is governed by consensus algorithms (CA). Several consensus mechanisms require significant computational power,…
U-CFR: Uncertainty-Guided Cascade Forward Refinement for Interactive Segmentation
Interactive image segmentation is critical for efficient image annotation; however, existing methods often require many corrective clicks o…
Transition-Related Potentials as Markers of Narrative Comprehension in Continuous EEG
Harnessing the potential of electroencephalography (EEG) for brain research is fundamentally limited by intrinsic noise and the diffuse pro…
Operational Identity: A Finite Audit of Declared and Implemented Rules of Sameness
A record system declares when two records refer to the same entity, occurrence, scope, or rule. Its disclosed implementation mechanisms ind…
GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning
Large language models increasingly use search tools to retrieve up-to-date information, introducing a new attack surface in which retrieved…
Self-Supervised Bio-Inspired Robotic Trajectory Planning with Obstacle Avoidance
Trajectory planning is a fundamental problem in robotics, requiring the generation of collision-free and efficient trajectories in a potent…
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute c…
Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles
Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choose which models to combine. W…
Emergent Compositional Skills in Mixture-of-Experts VLAs
We consider the problem of learning compositional robot policies end-to-end from expert demonstrations, without any pre-specified notion of…
HARP: The Human--AI Research Platform
Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exc…
Robostral Navigate
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains…
Synthetic minority data is redundant or invalid: a data-dependent validity theory and a de-biased test
For two decades, the standard remedy for class-imbalanced learning has been to fabricate synthetic minority examples, and the standard evid…
The Geometry of Personality: Activation Steering with Jungian Cognitive Functions
Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait framewo…
Beyond Heavy Log Curation: Perplexity-Based APT Detection via Unsupervised, Context-Augmented Language Models
Advanced Persistent Threats (APTs) remain difficult to detect because only a small fraction of events in large-scale logs are attack-relate…
Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery
Scientific machine learning methods such as neural operators and physics-informed neural networks have advanced engineering applications an…
Probabilistic Residual Learning for Online Recommendations
Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and item…
TwistedMerge: Certified Higher-Order Diagnostics and Abstention for Model Merging
Model merging combines independently trained or fine-tuned models, but pairwise alignability does not imply globally consistent alignment.…
Anti-Goal Reasoning: Rethinking the Theory of Goal Reasoning in Non-Axiomatic Logic
Goal reasoning in Non-Axiomatic Logic (NAL) explains how an adaptive system derives means for realizing desired events under insufficient k…
Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for opt…
Scientific exploration, collaboration and labor division in the large language model era
Large language models (LLMs) have rapidly and significantly entered scientific workflows, but it remains unclear how their diffusion is ass…
Interaction Dynamics Modeling and Predictive Control for Safe Steerable Catheter--Tissue Interaction
Safe steerable catheter control is fundamentally a problem of interaction dynamics: the tip must follow a planned motion, remain compliant…
HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving
Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While p…
Sparse Concept Channels in Frozen 3D CT Vision Encoders
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know which internal units…
Training Large Language Models for Self-Explanation Faithfulness
We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's…
TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning
Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important wh…
GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes
Preprocessing blood glucose time-series data is a critical yet often overlooked step in developing data-driven methods for diabetes managem…
Relative Value Learning
In reinforcement learning, critics typically estimate absolute state values $V(s)$, estimating how good a particular situation is in isolat…
Hardware-Software Co-Design for Float16 On-Device Training on RISC-V Single-Core
By leveraging standard RISC-V extensions, namely Zfh (scalar float16) and Zvfh (vector float16), this work proposes an open-source framewor…
Demographically-Informed Heat-Mortality Risk Curves via Risk Graph Neural Networks
Estimating heat-related mortality risk is a core task in environmental epidemiology, typically addressed with Distributed Lag Non-linear Mo…
One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies
Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask…
CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants…
Representative Sets in Propositional Abduction
The propositional abduction problem is a well-known form of non-monotonic reasoning where we are asked to find an explanation of a given ma…
Case study: proving sqrt(2) irrational with LPTP and an LLM
We present the interactions with an LLM (Large Language Model) aiming at proving that the square root of 2 is not a rational number in an L…
Encoding Event-B Proof Rules in Prolog: An Interactive Sequent Prover for ProB
Event-B is a formal method rooted in predicate logic and set theory. We encoded over 600 proof rules in Prolog, enabling a systematic, comp…
Animation, Verification and Visualisation of Prolog Transition Systems with ProB
ProB is a Prolog-based model checker, animator and constraint solver for high-level formal specifications. One can also use ProB to animate…
Chess\_db: A framework for working with large chess game datasets
Chess is a two player strategic game that is embedded in classical AI culture as it was once the frontier for intelligent behaviour. There…
Case study: solving P-99 with LPTP and an LLM
Ninety-Nine Prolog Problems (P-99) is a famous set of Prolog exercises. We solved the first thirty three just by prompting an LLM (Large La…
Declarative Problem Solving in UAM Strategic Deconfliction
The growing demand for Urban Air Mobility (UAM) introduces significant challenges in airspace management, particularly within densely popul…
Towards a Certifying Grounder
Grounding, the translation of high-level theories into equivalent quantifier-free formulas, is a crucial step in declarative solving, yet i…
Hybrid MKNF with Classical Negation in the Rule Component
Hybrid MKNF knowledge bases under the well-founded semantics integrate Description Logics with Logic Programming. However, they do not supp…
Explainability Framework for Policy-Aware Autonomous Agents
In the field of Artificial Intelligence, an agent is a system which is able to autonomously make decisions in order to reach a desired goal…
Explainable Belief Harmonization under Dynamic Epistemic Partitions
Existing approaches to multi-agent belief combination have established mature foundations for combining uncertain beliefs under common assu…
slang.gr as a Large-Scale Crowdsourced Resource for Non-Standard Greek
Slang is a central component of everyday language, reflecting linguistic creativity, social identity, and cultural change, yet its dy- nami…
pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development
In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-…
A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset
We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance j…
Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning
Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as modern game telemetry provide…
Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning
Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations. To evaluate u…
AI Assistants Overassist
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance fr…
Scaling Up Formal Representation of Clinical Trial Protocols in Ensemble Logic Using LLMs: A Preliminary Study
The reliance on unstructured free text for documenting clinical trial protocols creates a significant barrier to automated reasoning, cohor…
PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing
Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by…
GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG
Retrieval-Augmented Generation (RAG) systems increasingly employ multiple LLM agents. Yet, most prior work optimizes components in isolatio…
Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation
Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing…
From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics
Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, technology, and innovation (STI)…
Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource…
M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data
Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for compr…
Hilbert Operator for Progressive Encoding (HOPE): A Mathematical Framework for Deconstructing Learned Representations in Deep Networks
Deep neural networks encode complex representations, but deconstructing this internal knowledge remains a challenge. Given the link between…
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-su…
Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy
Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity…
VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions com…
When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation
A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard that watches that answer has t…
Cycle-Consistent and Uncertainty-Aware Neural Surrogates for Tokamak Edge Plasmas
The boundary and divertor plasma govern how a tokamak exhausts power and particles, setting heat fluxes, target conditions, and the onset o…
Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models
Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate w…
Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping
Video face swapping has no natural paired supervision: no real footage exists of one person's face performing another person's video. The s…
RUMBA: Russian User Memory Benchmark
The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely o…
Thinkink: 2D Spatial Ink-native Interaction with LLMs
People often use handwritten notes and sketches to externalize ideas for ideation. To integrate large language models (LLMs) into this prac…
Error Certificates for KV-Cache Eviction via Randomized Design
Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest. We prove that this design cannot k…
Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections
Coordinating autonomous vehicles at unsignalized intersections remains a critical challenge for multi-agent reinforcement learning (MARL) s…
Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it
A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language m…
Improved lower bounds for the Shannon capacity of odd cycles
The Shannon capacity $\Theta(G)$ of a graph $G$ quantifies the maximum rate at which information can be transmitted with zero error over a…
GS-Agent: Creating 4D Physical Worlds With Generative Simulation
Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional com…
ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundat…
From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs
Concurrent stateful library APIs expose behavior through evolving resource ownership, lifecycle states, and competing interleavings. Large…
Visual Contrastive Self-Distillation
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still n…
Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity
Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's…
Synthetic data generation framework for quality control automation in gravure printing
Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manual inspection. Automat…
Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$
Barzilai--Borwein (BB) method has shown strong practical performance in continuous optimization, yet its convergence dynamics remains poorl…
GraphVid: Interactive Graph-Controllable Video Generation
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts…
3D-Aware VLMs with Implicit and Explicit Geometries
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tas…
A Counterfactual Cause in Situation Calculus
Perhaps the most popular modern formulation of actual causality is the HP account by Halpern and Pearl. Recent advancement has focused on e…
Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models
Large language models (LLMs) are increasingly deployed in decision-support systems for high-stakes domains such as hiring and university ad…
SciTrek: Evaluating and Improving Long-Context Numerical Reasoning over Scientific Articles
We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language…
From Checklists to Clusters: A Homeostatic Account of AGI Evaluation
Contemporary AGI evaluations report multidomain capability profiles, yet they typically assign symmetric weights and rely on snapshot score…
WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance
Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex bro…
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors…
Understanding Critical Thinking in Generative Artificial Intelligence Use: Development, Validation, and Correlates of the Critical Thinking in AI Use Scale
Generative AI tools are increasingly embedded in everyday work and learning, yet their fluency, opacity, and propensity to hallucinate mean…
StackingNet: Collective Inference Across Independent AI Foundation Models
Artificial intelligence built on large foundation models has transformed language understanding, computer vision, and reasoning, yet these…
Diagnosing Pathological Chain-of-Thought in Reasoning Models
Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a critical intervention point for AI safety. How…
Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers
As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a pr…
Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks
We characterize AI automation as a continuum between crashing waves, in which capabilities jump abruptly across narrow task sets, and risin…
Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on…
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
Recent soft prompt research has tried to improve reasoning by inserting trained vectors into LLM inputs, yet whether the gain comes from th…
存在論的連続体に沿ったナレッジグラフのリエンジニアリング (拡張版)
ナレッジ グラフはデータ統合の主要な手段となっており、最新の AI の成功に不可欠ですが、軽量の語彙から高度に公理化されたオントロジーに至るまで、KG モデリングの実践の多様性により、統合と再利用は高価で脆弱なものになっています。この課題は、ニューラル コンポーネントとシンボリック コンポーネントの橋渡しが、新しい要件に合わせて KG を再設計する能力に依存する、ニューロシンボリック AI において特に深刻です。 GenAI は現在、前例のない自動化機能を提供していますが、KG 領域の原則的な理解がなければ、そのような自動化は概念的に根拠のないままです。我々は、存在論的連続体をその欠落した概念化として導入し、理論的構築物、その特徴付け枠組みが2つの直交する区別によって定義される理論的構築物である:意味論と語用論、および性質とアフォーダンス。これらは一緒になって、モデリング実践の全範囲にわたって KG を説明、比較、ナビゲート、変換するための語彙を定義します。方法論的な立場は経験的です。KG をどのようにモデル化するかを規定するのではなく、この連続体は、現実世界の KG エンジニアリング実践の観察から導き出され、その構造が、たとえば形式概念分析 (FCA) を通じて形式的に明示できる、既存の理論を定義することを目的としています。私たちは、来歴の知識に関するケーススタディを通じてビジョンを確立し、単一の懸念が連続体全体でどのように異なる形で現れるかを示します。私たちは 5 つのオープンな研究課題を明確にし、共有の研究課題として存在論的連続体を開発するようコミュニティに呼びかけます。
原文 (English)
Knowledge Graph Re-engineering Along the Ontological Continuum (extended version)
Knowledge graphs have become the primary vehicle for data integration and are critical to the success of modern AI, but the diversity of KG modelling practices, from lightweight vocabularies to richly axiomatised ontologies, makes integration and reuse expensive and brittle. This challenge is particularly acute in neuro-symbolic AI, where bridging neural and symbolic components depends on the ability to reengineer KGs to fit new requirements; GenAI now offers unprecedented automation capability, but without a principled understanding of the KG space, such automation remains conceptually ungrounded. We introduce the ontological continuum as that missing conceptualisation, a theoretical construct a theoretical construct whose characterisation framework is defined by two orthogonal distinctions: semantics vs pragmatics, and properties vs affordances; together these define a vocabulary to describe, compare, navigate, and transform KGs across the full range of modelling practices. The methodological stance is empirical: rather than prescribing how KGs should be modelled, the continuum aims to define a theory of the existent, derived from observation of real-world KG engineering practices and whose structure can be made formally explicit, for example, through Formal Concept Analysis (FCA). We ground the vision through a case study on provenance knowledge, showing how a single concern manifests differently across the continuum. We articulate five open research challenges and invite the community to develop the ontological continuum as a shared research agenda.
DN-Hypo-Pipeline: 大規模な言語モデルと科学的説明による仮説生成のための AI 主導のワークフロー
科学的仮説は研究の最初のステップであり、実験による検証が行われますが、科学的現象に対する深い理解と推論も反映されています。 DN-Hypo-Pipeline は、大規模な言語モデルに基づく AI を活用したワークフローで、事前知識として科学的な説明を活用することで、構造化された科学的思考と仮説生成をサポートするように設計されています。このパイプラインは、研究者が既存の文献から新しい仮説を導き出すのを支援します。研究論文の解説 (つまり、結論) が与えられると、基礎となる法則、理論、原理が特定され、観察された現象についての新しい、まだ検証されていない説明が再構築されます。私たちは、引用度の高い 3 つの論文を使用して、データ サイエンス モデリングの分野で DN-Hypo-Pipeline を評価しました。裁判官としての LLM による評価と人間の専門家による評価の両方によって裏付けられた統計的推論は、当社のパイプラインが直接生成方法よりも効果的であることを示しています。さらに、対応する新しいアルゴリズムを開発することにより、生成された 2 つの最高スコアの仮説を検証しました。このアルゴリズムは、元の論文で提示されたベースライン モデルを上回りました。 DN-Hypo-Pipeline は、データ サイエンスへの応用を超えて、理論に基づいたデータ サイエンス モデリング手法を包含するだけでなく、モデリング プロセスのより基本的な構造も明らかにする理論的フレームワークを提供します。さらに、このアプローチは本質的に理論に基づいたモデリングの一般化であり、他の領域やより幅広い科学分野に拡張できる可能性を提供します。
原文 (English)
DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations
Modern artificial intelligence excels at prediction but cannot explain. From large language models to AI-for-science systems, today's machines answer what by recombining patterns already present in the human literature, yet they cannot reason out why a phenomenon must arise from underlying principles even though explanation, not prediction, lies at the heart of scientific discovery. Here we ask whether the structure of scientific explanation can be operationalized to guide how a machine generates hypotheses. We introduce DN-Hypo-Pipeline, a hypothesis-generation framework that adopts a layered, explanation-theoretic scaffold: Hempel's deductive-nomological (DN) model supplies the output form and deductive validity of a hypothesis, Salmon's causal-process account supplies an organizing constraint on where to search for the governing laws, and Armstrong's view of laws as relations between universals supplies the bridge from a phenomenon's constituent processes to the laws that may be associated with it. Rather than searching the space of what has been written, the framework searches the space of what principles govern a phenomenon: given an explanandum, it abstracts the universals instantiated in the phenomenon's formation process, retrieves the laws relating those universals, and deductively reconstructs a new, testable explanation. Evaluated in data-science modeling and judged by both LLMs and human experts, hypotheses generated through this principled reasoning significantly outperform those from direct prompting. Crucially, we translated the two highest-scoring hypotheses into novel algorithms one that reduces the Transformer's theoretical complexity with only minimal performance loss, and another that achieves competitive accuracy with substantially fewer parameters.
HarnessX: 構成可能、適応性、進化可能なエージェント ハーネス ファウンドリ
AI エージェントのパフォーマンスは、モデルがどのように観察、推論、動作するかを仲介するプロンプト、ツール、メモリ、制御フローで構成されるランタイム ハーネスに大きく依存します。しかし、今日のハーネスは大部分が手作りで静的なままです。新しいモデルやタスクごとに特注の足場が依然として必要であり、実行中に生成される豊富な痕跡が体系的な改善に蒸留されることはほとんどありません。構成可能、適応性、進化可能なエージェント ハーネスのファウンドリである HarnessX を紹介します。 HarnessX は、置換代数を介して型指定されたハーネス プリミティブをアセンブルし、記号適応と強化学習の間の操作ミラーに基づいたトレース駆動型マルチエージェント進化エンジンである AEGIS を通じてそれらを適応させ、軌道をハーネス更新とモデル トレーニング信号の両方に変えることでハーネス モデル ループを閉じます。 5 つのベンチマーク (ALFWorld、GAIA、WebShop、tau^3-Bench、および SWE-bench Verified) にわたって、HarnessX は平均 +14.5% (最大 +44.0%) の利益をもたらし、ベースラインが最も低いところでは利益が最大になります。これらの結果は、エージェントの進歩がモデルのスケーリングのみによってもたらされる必要はないことを示唆しています。実行フィードバックからランタイム インターフェイスを構成および進化させることは、実用的で補完的な手段です。完全なコードベースは将来のリリースでオープンソース化される予定です。
原文 (English)
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts. Yet today's harnesses remain largely hand-crafted and static: each new model or task still demands bespoke scaffolding, and the rich traces produced during execution are rarely distilled back into systematic improvement. We introduce HarnessX, a foundry for composable, adaptive, and evolvable agent harnesses. HarnessX assembles typed harness primitives via a substitution algebra, adapts them through AEGIS, a trace-driven multi-agent evolution engine grounded in an operational mirror between symbolic adaptation and reinforcement learning, and closes the harness-model loop by turning trajectories into both harness updates and model training signal. Across five benchmarks (ALFWorld, GAIA, WebShop, tau^3-Bench, and SWE-bench Verified), HarnessX yields an average gain of +14.5% (up to +44.0%), with gains largest where baselines are lowest. These results suggest that agent progress need not come from model scaling alone: composing and evolving runtime interfaces from execution feedback is an actionable and complementary lever. Project homepage: https://darwin-agent.github.io/HarnessX/.
チャットボットからデジタル同僚へ: 永続的な自律型 AI へのパラダイム シフト
大規模言語モデル (LLM) は、会話ジェネレーターから、推論、行動、記憶、自己改善が可能な統合 AI システムへと根本的な変革を遂げています。私たちはこの移行を、チャットボットからデジタル同僚への移行、つまり会話による回答から永続的な作業への移行として概念化しています。私たちはこの移行を 2 つの密接に結合した次元に沿って整理します。まず、認知コア レベルでは、LLM はネクスト トークン予測によって駆動されるチャットボット時代の「高速思考」システムから、より意図的で信頼性の高い認知をサポートするために、推論時間の計算、思考連鎖推論、リフレクション、プロセス監視、強化学習を活用する思考 LLM へと進化しています。第 2 に、ツール拡張タスク実行レベルでは、LLM は、アドホックな方法で外部リソースを呼び出すツール呼び出しエージェントから、永続的なワークスペース、スキル、検証ループ、ガバナンスを備えた OpenClaw スタイルのワークステーション システム (OpenClaw) へと進歩しています。 「ワークスペース + スキル」パラダイムにより、状態の永続性、再利用可能なプロシージャ、タスクの終了、エクスペリエンスの再利用により、エピソード ツールが同僚のように使用できるようになります。データ構築が命令と応答のペアから状態、行動、観察の軌跡へ移行し、評価が静的ベンチマークからサンドボックス化された監査可能な自己進化型 AI エコシステムへ移行することを検証します。
原文 (English)
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable of reasoning, action, memory, and self-improvement. We conceptualize this transition as a shift from Chatbot to Digital Colleague: from conversational answers to persistent work. We organize this transition along two tightly coupled dimensions. First, at the cognitive core level, LLMs are advancing from Chatbot-era "fast thinking" systems driven by next-token prediction toward Thinking LLMs that leverage inference-time computation, Chain-of-Thought reasoning, reflection, process supervision, and reinforcement learning to support more deliberate and reliable cognition. Second, at the tool-augmented task execution level, LLMs are progressing from tool-calling Agents that invoke external resources in an ad hoc manner toward OpenClaw-style workstation systems (OpenClaw) equipped with persistent Workspaces, skills, verification loops, and governance. The "Workspace + Skill" paradigm makes episodic tool use colleague-like via state persistence, reusable procedures, task closure, and experience reuse. We examine data construction shifts from instruction-response pairs to State-Action-Observation trajectories and evaluation from static benchmarks to sandboxed, auditable, self-evolving AI ecosystems.
ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents
Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory i…
Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles
This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In…
税金を意識したパーソナライズされたポートフォリオ管理のための 3 段階の基礎モデル
我々は、これまでの金融 RL 作業すべてに共通する 3 つの制限、1) ティッカー ロックイン、2) モノリシック目標、3) 静的ユーザー モデルに対処する、パーソナライズされたポートフォリオ管理のための 3 段階の深層強化学習システムを紹介します。フェーズ 1 では、マルチアセット コーパスでの自己教師あり学習を介して、ティッカー ID のないクロスアセット エンコーダーを事前トレーニングします。学習されたゲート メカニズムを介して融合された、T5 ベースの時系列基盤モデルである Chronos を使用した凍結並列ブランチによって強化されます。私たちの知る限り、これはポートフォリオ管理 RL への時系列基礎モデルの最初の適用です。エンコーダーは、新しいティッカーの再トレーニングを必要としない、50 次元の観察可能なメタデータ ベクトルを介して、あらゆる公開取引資産に一般化します。フェーズ 2 では、エピソードごとにサンプリングされた 6 つの異なる投資目標 (短期アルファ、短期利益、長期利益、資本保全、税金損失の回収、および長期利益のみ) を同時に提供する目的条件付き報酬の下で、MoE (専門家混合) のポートフォリオ アクター評論家を PPO で微調整します。 MoE アーキテクチャは、各目標を専門のエキスパート ヘッド (勢い、成長、守備、税務意識) に割り当て、学習型インテント ルーターがアクティブな目標と現在の市場体制に基づいて専門家をブレンドすることで、目標間の勾配の競合を排除します。フェーズ 3 では、実際の証券取引履歴に基づいて微調整された 76 パラメーターの LoRA モジュールを介して推論時に各個人にさらに適合する軽量のパーソナライゼーション レイヤーを追加し、アンケートではなく明らかになった取引行動から投資目標を推測します。自然言語インテント パーサーは、自由形式の目標を構造化された投資目標パラメータに直接変換します。
原文 (English)
A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management
We present a three-phase deep reinforcement learning system for personalized portfolio management that addresses three limitations shared by all prior financial RL work: 1) ticker lock-in, 2) monolithic objectives , and 3) static user models. Phase 1 pretrains a ticker-identity-free cross asset encoder via self-supervised learning on a multi-asset corpus, augmented by a frozen parallel branch using Chronos, a T5-based time series foundation model, fused via a learned gating mechanism. To our knowledge, this is the first application of a time series foundation model to portfolio management RL. The encoder generalizes to any publicly traded asset via a 50-dimensional observable metadata vector that requires no retraining for new tickers. Phase 2 fine-tunes a MoE (Mixture of Experts) portfolio actor critic with PPO under an objective-conditioned reward that simultaneously serves six distinct investment goals sampled per episode: short-term alpha, short-term gain, long-term gain, capital preservation, tax-loss harvesting, and long-term-gains-only. A MoE architecture assigns each objective to a specialized expert head (momentum, growth, defensive, tax-aware), and a learned intent router blends experts based on the active objective and current market regime, which eliminates cross-objective gradient conflict. Phase 3 adds a lightweight personalization layer further adapted at inference time to each individual via a 76-parameter LoRA module fine-tuned on real brokerage transaction history, inferring investment objectives from revealed trading behavior rather than questionnaires. A natural language intent parser converts free-form goals directly into structured investment objective parameters.
隠れたフットプリント: ストレージを LLM エージェント評価の第一級の指標にする
LLM エージェントのベンチマークは、タスクの完了、信頼性、推論コストを測定しますが、ログ、コンテキスト スナップショット、チェックポイント、デバッグ トレースなど、エージェントの実行によってディスクに残される永続データは測定しません。実行後のエージェント ストレージ フットプリントのクロスフレームワーク ベンチマークである AgentFootprint を紹介します。そのシリアル化対応メトリクス スイートは、総保持率、チャネル構成、重複、増加、圧縮率、会話履歴の再構築可能性を測定します。これは、測定の罠に対処します。単純なバイトレベルの測定では、データベースのページングと JSON エスケープが繰り返されるコンテンツを不明瞭にするため、重複が桁違いに過小評価されます。固定トレース制御により、エージェントが生成した論理ボリュームが永続層の増幅から分離されます。7 つの永続フレームワークを通じて同じ軌跡を再生すると、6.7 倍の広がりが得られます。同一のモデル、ツール、およびタスクでは、100% の精度の構成では、デフォルトでサポートされる回復機能と監査機能が異なりますが、保持バイト数が 15.7 倍異なります。 3 つの完全な履歴構成は、反復観察ストレス タスクで超線形に成長します。 108 個のインスタンスで正規化された SWE ベンチからエクスポートされた軌跡 検証済みの送信は、インスタンスごとに 3 桁の大きさに及び、解決率との検出可能な相関関係はありません。コンテンツ アドレス ストアは、すべての再構築可能性スコアを維持しながら、保持率を 4.8 倍から 32.7 倍まで削減します。これらの結果は、精度と再構築可能性を併せてレポートするためのリソース メトリックとして永続ストレージを確立します。
原文 (English)
The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation
LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite measures total retention, channel composition, duplication, growth, compressibility, and conversation-history reconstructability. It addresses a measurement trap: naive byte-level measurement understates duplication by an order of magnitude because database paging and JSON escaping obscure repeated content. A fixed-trace control separates agent-generated logical volume from persistence-layer amplification: replaying the same trajectory through seven persisting frameworks yields a 6.7x spread. Under identical models, tools, and tasks, configurations with 100% accuracy differ by 15.7x in retained bytes, although their defaults support different recovery and audit capabilities. Three full-history configurations grow superlinearly on a repeated-observation stress task. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude per instance, with no detectable correlation with resolve rate. A content-addressed store reduces retention by 4.8x-32.7x while preserving every reconstructability score. These results establish persistent storage as a resource metric to report jointly with accuracy and reconstructability.
コンパイルしてページング: 実行可能 SOP プログラムと手続き型 LLM エージェントの機能ゲート型ランタイム
企業エージェントは、長期にわたる条件付きの安全性を重視した標準運用手順 (SOP) に従う必要があります。機械可読な SOP 制約を実行可能な疑似コードにコンパイルし、LLM がセマンティック実行を実行している間にアクティブなフレームをページングするプログラムガイド付き (PG) スタック マシンで実行します。 6 つのモデルにわたる 3 アーム SOPBench の調査では、表現と実行時が分離されています。コンパイルされたテキストは決して大幅に損なうことはなく、公式の散文がパフォーマンスを下回る場合でも最大 16.0 ポイント向上します。ランタイム ガイダンスは機能ゲート型です。 2 つの強力なモデルは独立して、正の 7 ドメイン PG コントラスト (58:19 および 75:31 の不一致ペア) を示しますが、弱いモデルは損傷を受けています。フルプログラムのカーソルアブレーション (最初にアクティブなフレーム、完全なプログラムを保持) では、強力なモデルの拒否ゲインの多くが回復します。可視性を選択すると、多少の改善が加えられます。プローブと監査のペアの測定により、この分裂は、再構築能力ではなく自発的な状態規律に基づいて追跡されます。バンクでは、3 つの主要なアームが 70.4、86.4、92.8 に上昇し、100% の拒否精度が得られます。実践的なガイダンス: 最初にコンパイルします。モデルレベルの規律チェックの後にのみアクティブフレームページングを有効にします。
原文 (English)
Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents
Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP constraints into executable pseudo-code and run them with a program-guided (PG) stack machine that pages the active frame while an LLM performs semantic execution. A three-arm SOPBench study across six models separates representation from runtime: compiled text never significantly hurts and gains up to 16.0 points where official prose underperforms. Runtime guidance is capability-gated. Two strong models independently show positive seven-domain PG contrasts (58:19 and 75:31 discordant pairs), whereas weak models are harmed. A full-program cursor ablation (active frame first, complete program retained) recovers much of the strong-model refusal gain; selective visibility adds a smaller improvement. Paired probe and audit measurements track this divide to spontaneous state discipline rather than reconstruction ability. On Bank the three primary arms rise from 70.4 to 86.4 to 92.8, with 100% refusal correctness. Practical guidance: compile first; enable active-frame paging only after a model-level discipline check.
AI エージェントは本当に RTL から GDS への変換を完了できるでしょうか?ベンチマーク ツールからの教訓 - インタラクティブ EDA ワークフロー
LLM 駆動のエージェント システムは、電子設計自動化 (EDA) の有望なパラダイムとして浮上しており、複雑な設計ワークフローを自動化する強力な可能性を示しています。ただし、既存の評価は主に、分離された EDA タスクに関する個々の言語モデルを調査しており、完全な EDA フロー全体でさまざまなエージェント システムがどのように実行されるかについての洞察は限られています。この研究では、統一されたプロンプト、ツール環境、テクノロジー ライブラリ設定の下でのエンドツーエンド EDA ワークフローにおける AI エージェントの体系的な評価である FluxBench を紹介します。当社の評価では、オープンソース ツールチェーンを使用した RTL 生成や、産業アプリケーション向けのクローズドソースの商用 EDA ツールを使用した RTL から GDS へのフローなど、代表的なシナリオをカバーしています。これらのワークフローを通じて、RTL コード生成、反復修復、ツール フィードバックの利用、論理合成、配置配線 (P&R)、およびエンジニアリング変更オーダー (ECO) 自動化におけるエージェントの能力を評価します。エージェント システムの効率をさらに特徴付けるために、トークンの使用量とランタイム コストと比較した EDA アーティファクトの効果的な改善を測定するコスト効率の指標であるトークン ROI を導入します。実験結果によると、同じ基盤モデルに基づいて構築されている場合でも、エージェント システム アーキテクチャが異なると、最大 86.27% のパフォーマンス ギャップが見られる可能性があります。さらに、同等のタスク パフォーマンスを持つシステム間では、トークン ROI が $105.92\times$ も異なる可能性があります。 PicoRV32 をケーススタディとして使用した RTL から GDS へのフローでは、FluxEDA は最大 97.94 のエンドツーエンド スコアを達成し、ドメイン固有の EDA スキルを備えた Claude Code を最大 $8.39\times$ 上回りました。これらの結果は、大規模な EDA シナリオでエージェントのパフォーマンスを向上させるには、ドメイン固有のスキルだけでは不十分であることを示しています。代わりに、エージェント システム設計と基盤モデル機能の両方が、効果的な自動 EDA ワークフローを実現する上で重要な役割を果たします。
原文 (English)
Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows
Large language model (LLM) agents are extending electronic design automation (EDA) beyond static RTL generation toward long-horizon, tool-interactive workflows. Yet it remains unclear whether general-purpose coding agents, even with domain-specific EDA skills, can reliably execute an end-to-end RTL-to-GDS flow encompassing synthesis, physical implementation, and engineering change order (ECO) optimization. We evaluate AI agents on a PicoRV32 RTL-to-GDS flow using commercial EDA tools under two timing targets. Their performance is assessed using end-to-end design score, stage completion, and Token ROI, a cost-efficiency metric relating design quality to runtime and cost. Comparing three agent architectures and four foundation models, we derive three practical lessons. First, domain-specific skills improve agents' understanding of individual subtasks but do not ensure reliable completion of a long-horizon EDA flow. Second, agents that achieve similar design progress can still differ by up to 141 times in Token ROI, revealing substantial differences in runtime and cost efficiency. Third, low-level tool-interface mismatches are a major source of physical design failures, particularly when Tcl commands depend on the tool version or execution mode. These results suggest that robust Agentic EDA requires not only stronger models but also structured tool interfaces, persistent design context, controlled execution, and process-level evaluation.
LLM 群集の知恵: 言語モデル アンサンブルにおける集約と汚染
群衆の知恵、つまり個人間の判断を総合すると、最も優れた個人よりも優れた結果をもたらすことが多いという発見は、人間の予報士を対象に広範囲に研究されてきました。 「群衆」が大規模言語モデル (LLM) で構成されている場合に同じ現象が現れるかどうかは、理論的意味と実践的意味の両方を伴う未解決の問題です。 254 のバイナリ予測市場質問について 15 の LLM から確率推定値を導き出し、古典的な集計方法と学習された集計方法を評価しました。学習されたアグリゲーター (多層パーセプトロンとロジスティック回帰) は、すべての個別モデルや古典的な手法を上回りました。ロジスティック回帰はニューラル ネットワークと一致することがわかり、学習された集計の利点は、非線形相互作用ではなく、多様なモデル出力の線形結合の学習から得られることを示唆しています。ニューラル ネットワークの学習されたマッピングに適用されたシンボリック回帰により、純粋なモデル不一致信号がパレート フロンティア上で最も複雑性の低い有用な式として復元され、この解釈がさらに裏付けられました。トレーニング カットオフの汚染が蔓延した混乱であることが判明しました。フロンティア クラウド モデルと小規模なローカル モデルの間の見かけの能力差は、すべてのモデルのトレーニング カットオフ後に解決される質問のクリーンなサブセットで 35.8% から 8.9% に崩壊し、個々のモデルのランキングは中程度の安定性しか示しませんでした。予測市場が各モデルのトレーニング カットオフで評価された場合でも、LLM の精度は大幅に低いままであり、集合的な情報集約における真のギャップを示しています。これらの発見は、LLM 群衆が群衆の知恵効果を示す可能性があるが、信頼できる評価には汚染のない評価が不可欠であることを示唆しています。
原文 (English)
Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles
The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively studied with human forecasters. Whether the same phenomenon emerges when the ``crowd'' consists of large language models (LLMs) is an open question with both theoretical and practical implications. We elicited probability estimates from 15 LLMs on 254 binary prediction market questions and evaluated classical and learned aggregation methods. Learned aggregators -- a multilayer perceptron and a logistic regression -- outperformed all individual models and classical methods. The logistic regression was found to match the neural network, suggesting that the benefit of learned aggregation derives from learning a linear combination of diverse model outputs rather than from nonlinear interactions. Symbolic regression applied to the neural network's learned mapping recovered a pure model-disagreement signal as the lowest-complexity useful formula on the Pareto frontier, further supporting this interpretation. Training cutoff contamination proved a pervasive confound: the apparent capability gap between frontier cloud models and smaller local models collapsed from 35.8% to 8.9% on a clean subset of questions resolving after all models' training cutoffs, and individual model rankings showed only moderate stability. Even when the prediction market is evaluated at each model's training cutoff, LLMs remained substantially less accurate, indicating a genuine gap in collective information aggregation. These findings suggest that LLM crowds can exhibit wisdom-of-crowds effects, but that contamination-free evaluation is essential for reliable assessment.
依存関係から構成性へ: 組み合わせカテゴリ文法による LLM 出力の神経記号的リフティング
大規模言語モデル (LLM) は、接頭辞から次のトークンを段階的に予測することにより、流暢なテキストを生成します。生成の伝統の批評家は、そのようなシステムには真の文法が欠けていると主張します。依存関係文法の観点からの影響力のある回答は、LLM の動作は、単語ごとに構築されたローカルのヘッド依存構造によってよく記述されると主張しています。私たちは、より鋭い観察が見落とされてきたと主張します。つまり、自己回帰生成の接頭辞駆動型補完ダイナミクスは、結合カテゴリ文法 (CCG) が元々サポートするように設計された増分処理モデルと密接に一致しています。これに基づいて、LLM の出力が型付けされた構成導出に持ち上げられる神経象徴的なフレームワークを提案します。LLM が CCG を内部的に実装しているとは主張しませんが、その出力は原則に基づいた増分的で監査可能な CCG 再構築を可能にすると主張します。 2 つの結果が続きます。まず、カリーとハワードの対応を通じて、リフティングは自然言語を超えて、LLM が生成する形式言語 (Solidity などのプログラミング言語、記述ロジック、OWL や SQL などのクエリ言語) まで拡張され、型システムは変化し、アーキテクチャは固定されています。第 2 に、リフティングでは 2 つのチェック層がサポートされています。構造上の欠陥を直接検出する構成層と、リフトされた構造を外部の知識ソースと照合してチェックするコンテンツ層です。これにより、幻覚コンテンツの可能な限り早期のフラグ付けが可能になります。したがって、アカウントはプロデューサーに認識ではなくプレフィックス駆動の生成プロファイルを要求します。最後に、フレームワークが開く一方向としての同期 LLM-CCG カップリングのスケッチを示します。
原文 (English)
From Dependency to Compositionality: A Neurosymbolic Lifting of LLM Outputs via Combinatory Categorial Grammar
Large language models (LLMs) generate fluent text by incrementally predicting the next token from a prefix. Critics in the generative tradition argue that such systems lack genuine grammar; influential replies from the dependency-grammar perspective hold that LLM behavior is well described by local head-dependent structure built word by word. We argue that a sharper observation has been overlooked: the prefix-driven, type-completing dynamics of autoregressive generation align closely with the incremental processing model that Combinatory Categorial Grammar (CCG) was originally designed to support. On this basis we propose a neurosymbolic framework in which LLM outputs are lifted into typed compositional derivations -- not claiming that LLMs implement CCG internally, but that their outputs admit a principled, incremental, and auditable CCG reconstruction. Two consequences follow. First, through the Curry-Howard correspondence the lifting extends beyond natural language to the formal languages LLMs also produce -- programming languages such as Solidity, description-logic and query languages such as OWL and SQL -- with the type system varying and the architecture held fixed. Second, the lifting supports two layers of checking: a compositional layer that catches structural failures directly, and a content layer that checks the lifted structure against external knowledge sources, enabling the earliest possible flagging of hallucinated content. The account thereby requires of a producer not cognition but a prefix-driven generative profile. We close with a sketch of synchronous LLM-CCG coupling as one direction the framework opens.
SenWorld: コンテキストリッチな評価データを生成するためのデジタル ツイン シミュレーション
スマートフォンのパーソナル アシスタントは長期にわたる個人データを推論しますが、その評価には正解がわかっているコンテキストに富んだ評価データが必要であり、実際のデバイスのトレースはプライバシーに敏感すぎて共有できません。この課題に対処するために、構築によって固定されたグラウンド トゥルースを使用してそのようなデータを生成する、物理的に接地され、決定論的でイベント ソースのデジタル ツイン シミュレーションである SenWorld を紹介します。 SenWorld では、ペルソナは実際の地図、天気、休日、ネットワーク データから構築された世界で 1 日を過ごします。観測可能なすべての信号はシステム全体のスナップショットにアーカイブされます。また、各評価ケースは、事後注釈や大規模言語モデル (LLM) ジャッジではなく、既存のレコードへのポインターによってラベル付けされます。この手法を北京の 16 人のペルソナで評価しました。生成されたデータは、カテゴリ分布 (ジェンセンとシャノンの相違 (JSD) 0.070) および通信記録の 1 日のリズム (JSD 0.1 未満) において、保持されている実際のユーザーのベンチマークと厳密に一致していますが、生成された記録は実際の記録よりも短いままです。スクリプトによる対話がなければ、ペルソナは完全に往復する対話サブグラフと差別化された行動レパートリーを形成します。 717 件の評価ケースに投影された生成データでは、実稼働スマートフォン アシスタントの 78 件の障害が明らかになり、通話とショート メッセージ サービス (SMS) の記録に集中し、連絡先、スケジュール、アラームは決して失敗しませんでした。スナップショット ポインタは、LLM 判定者が関与せずに、各失敗をアシスタント側の取得エラーとして確認します。全体として、SenWorld は、ラベルが構築によって固定されている評価データへの、プライバシーに安全で再現可能で配布がチェックされたパスを提供します。
原文 (English)
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, we present SenWorld, a physically grounded, deterministic, event-sourced digital-twin simulation that generates such data with ground truth fixed by construction. In SenWorld, personas live through a full day in a world built from real map, weather, holiday, and network data; every observable signal is archived in full-system snapshots; and each evaluation case is labeled by a pointer to an existing record rather than by post-hoc annotation or a large language model (LLM) judge. We evaluate this method with 16 personas in Beijing. The generated data closely matches the held-out real-user benchmark in category distribution (Jensen--Shannon divergence (JSD) 0.070) and in the daily rhythm of communication records (JSD below 0.1), though generated records remain shorter than real ones. Without scripted interaction, personas form a fully reciprocated dialogue subgraph and differentiated behavioral repertoires. Projected into 717 evaluation cases, the generated data exposes 78 failures in a production smartphone assistant, concentrating on call and Short Message Service (SMS) records while contacts, schedules, and alarms never fail. The snapshot pointer confirms each failure as an assistant-side retrieval error, with no LLM judge involved. Overall, SenWorld offers a privacy-safe, reproducible, and distribution-checked path to evaluation data whose labels are fixed by construction.
PRO-LONG: プログラムによる記憶により長期的な推論が可能に
長期的なタスクには持続的な認識、推論、探索が必要であり、大規模言語モデル (LLM) エージェントにとっては永続的な課題です。このギャップは、特にモデルがそのまま評価される場合、ARC-AGI-3 などの継続学習ベンチマークでのパフォーマンスの制限に反映されます。このギャップを埋めるためにさまざまなエージェントのハーネスが提案されており、それぞれが長い観察シーケンスを処理するための戦略、つまり環境からどのような情報を保存し、それをモデルのコンテキストにどのようにロードするかという戦略に取り組んでおり、この選択は特に重要であると私たちが主張しています。コンテキスト管理の既存の方法は、より多くの情報を保存すると関連する詳細の取得が困難になるため、重大なトレードオフに直面しています。私たちは、長期的な探索的設定における LLM エージェント向けのプログラム メモリを中心に構築された最小限のコンテキスト管理フレームワークである PRO-LONG を提案します。 PRO-LONG は、完全で構造化されたインタラクション ログを保持し、コーディング エージェントの最近の進歩を利用してこの履歴を効率的に検索することで、トレードオフに対処します。完全な ARC-AGI-3 パブリック ゲーム セットでは、PRO-LONG は、フロンティア モデル全体でベース コーディング エージェントよりも平均 18.0 パーセント向上し、最先端の特殊ハーネスと同等またはそれを上回り (最大 76.1% パス @1)、使用するトークンの数は 4.2 ~ 5.8 分の 1 です。 Fable 5 では、PRO-LONG は合計コスト $1,750 で 97.4% の最高@2 を達成します。関連するコードとログは https://github.com/alexisfox7/PRO-LONG で入手できます。
原文 (English)
PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks such as ARC-AGI-3, especially when models are evaluated out of the box. Various agent harnesses have been proposed to close this gap, and each commits to a strategy for handling long sequences of observations, i.e., what information to save from the environment and how to load it into model context, a choice we argue is particularly consequential. Existing methods for context management face a significant tradeoff, as preserving more information makes retrieving relevant details less tractable. We propose PRO-LONG, a minimal context management framework built around programmatic memory for LLM agents in long-horizon, exploratory settings. PRO-LONG addresses the tradeoff by keeping a complete, structured interaction log and capitalizing on recent progress in coding agents to search this history efficiently. On the full ARC-AGI-3 public game set, PRO-LONG improves over a base coding agent by an average of 18.0 percentage points across frontier models, and matches or exceeds state-of-the-art specialized harnesses (up to 76.1% pass@1) while using 4.2-5.8x fewer tokens. With Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of \$1,750. Relevant code and logs are available at https://github.com/alexisfox7/PRO-LONG.
Generative AI and Agency in Education: A Critical Scoping Review and Thematic Analysis
This scoping review examines the relationship between Generative AI (GenAI) and agency in education, analyzing the literature available thr…
Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content. We i…
Loss-Complexity Landscape and Model Structure Functions
We develop a framework for dualizing the Kolmogorov structure function $h_x(\alpha)$, which then allows using computable complexity proxies…
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descri…
Drive As You Like: Multi-Head Diffusion with Reinforcement Learning for Personalized Driving
Despite significant progress, imitation learning-based autonomous driving planners remain largely restricted to reproducing high-frequency…
DynaMark: A Reinforcement Learning Framework for Dynamic Watermarking in Industrial Machine Tool Controllers
Industry 4.0's highly networked Machine Tool Controllers (MTCs) are prime targets for replay attacks that use outdated sensor data to manip…
Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs
Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typica…
Equivariant Conditional Diffusion Model for Head and Neck CT Image Synthesis from CBCT
Background: Cone-beam computed tomography CBCT is a commonly used modality for image guided radiotherapy. It offers real time anatomical vi…
Simple Policy Gradients for Reasoning with Diffusion Language Models
Diffusion large language models (dLLMs) represent a promising alternative to autoregressive LLMs; however, the lack of effective post-train…
On the Granularity of Causal Effect Identifiability
The classical notion of causal effect identifiability is defined in terms of treatment and outcome variables. In this paper, we consider th…
Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
Generative artificial intelligence (GenAI) is transforming bioinformatics by advancing genomics, proteomics, transcriptomics, structural bi…
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, age…
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
Error Span Detection (ESD) extends automatic machine translation (MT) evaluation by localizing translation errors and labeling their severi…
Vision-Language-Policy Model for Dynamic Robot Task Planning
Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robo…
Backpropagation-Free Test-Time Adaptation for Lightweight EEG-Based Brain-Computer Interfaces
Electroencephalogram (EEG)-based brain-computer interfaces (BCIs) face significant deployment challenges due to inter-subject variability,…
Knowledge-Guided Time-Varying Causal Inference for Arctic Sea Ice Dynamics
Quantifying the causal relationship between sea ice thickness and sea surface height (SSH) is essential for understanding the mechanisms dr…
NeuraLSP: A Neural Spectral Preconditioner for Accelerating PDE Solvers
Solving large-scale sparse linear systems originating from partial differential equations (PDEs) is a fundamental topic in high-performance…
PILD: Physics-Informed Learning via Diffusion
Diffusion models have emerged as powerful generative tools for modeling complex data distributions, yet their purely data-driven nature lim…
Variational Speculative Decoding: Rethinking Draft Training from Token Likelihood to Sequence Acceptance
Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize singl…
Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems
The pantograph-catenary interface is essential for ensuring uninterrupted and reliable power delivery in electrified rail systems. However,…
Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents
Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hy…
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual…
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled,…
Benchmarking Unlearning for Vision Transformers
Machine unlearning (MU) refers to the post-training capability to remove (the influence of) training examples that are incorrect, biased, o…
What Matters for Simulation to Online Reinforcement Learning on Real Robots
We investigate what specific design choices enable successful online reinforcement learning (RL) on physical robots. Across 100 real-world…
AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teac…
VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory
Imitation learning from human demonstrations has achieved significant success in robotic control, yet most visuomotor policies still condit…
SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training
Test-Time Training (TTT) language models replace the KV-cache with fast weights updated during inference, achieving O(1) memory but sufferi…
Evolutionarily Stable Stackelberg Equilibrium
We present a new solution concept called evolutionarily stable Stackelberg equilibrium (SESS). We study the Stackelberg evolutionary game s…
EZASP - Facilitating the Usage of ASP
Answer Set Programming (ASP) is a declarative programming language used for modeling and solving complex combinatorial problems. It has bee…
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight…
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit b…
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows affects feed-forward 3D reconstruction.…
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of…
Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model
We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.56 billion tokens of pure Classical Chinese, wit…
Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
Preference optimization is widely used to align large language models (LLMs) with human preferences. However, many margin-based methods als…
Streamliners for Answer Set Programming
Streamliner constraints reduce the search space of combinatorial problems by ruling out portions of the solution space. We adapt the Stream…
SafeHarbor: LLM エージェントの安全のための階層型メモリ拡張ガードレール
基盤モデルの最近の進歩により、LLM は受動的な会話システムから、推論とツールの実行が可能な自律エージェントに変わりました。これらの機能は実質的な実用的価値を解放しますが、敵対者がエージェントを操作して現実世界の環境で有害なアクションを実行する可能性があるため、新たなセキュリティ リスクももたらします。既存の防御戦略はそのような脅威を軽減しますが、安全性と有用性のバランスをとるのにしばしば苦労し、その結果、無害なユーザー要求を過度に拒否する結果になります。このトレードオフを軽減するために、LLM エージェントの正確な決定境界を確立するように設計された新しいフレームワークである SafeHarbor を提案します。静的なガイドラインとは異なり、SafeHarbor は強化された敵対的生成を通じてコンテキストを認識した防御ルールを抽出します。私たちは、動的ルール注入用のローカル階層メモリ システムを設計し、トレーニング不要で効率的なプラグ アンド プレイ ソリューションを提供します。さらに、動的なノードの分割と結合を通じてメモリ構造を継続的に最適化する、情報エントロピーベースの自己進化メカニズムを導入します。広範な実験により、SafeHarbor があいまいで良性のタスクと明示的な悪意のある攻撃の両方で最先端のパフォーマンスを達成し、特に GPT-4o で 63.6\% のピーク無害ユーティリティを達成しながら、有害なリクエストに対して 93\% を超える堅牢な拒否率を維持していることが実証されています。ソース コードは https://github.com/ljj-cyber/SafeHarbor で公開されています。
原文 (English)
SafeHarbor: Defining Precise Decision Boundaries via Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and tool execution. While these capabilities unlock substantial practical value, they also introduce new security risks, as adversaries can manipulate agents into performing harmful actions in real-world environments. Existing defense strategies mitigate such threats but frequently struggle to balance safety and utility, resulting in over-refusal of benign user requests. To mitigate this trade-off, we propose SafeHarbor, a novel framework designed to establish precise decision boundaries for LLM agents. Unlike static guidelines, SafeHarbor extracts context-aware defense rules through enhanced adversarial generation. We design a local hierarchical memory system for dynamic rule injection, offering a training-free, efficient, and plug-and-play solution. Furthermore, we introduce an information entropy-based self-evolution mechanism that continuously optimizes the memory structure through dynamic node splitting and merging. Extensive experiments demonstrate that SafeHarbor achieves state-of-the-art performance on both ambiguous benign tasks and explicit malicious attacks, notably attaining a peak benign utility of 63.6\% on GPT-4o while maintaining a robust refusal rate exceeding 93\% against harmful requests. The source code is publicly available at https://github.com/ljj-cyber/SafeHarbor.
AI Security Policy Should Assess Systems, Not Only Models
We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared me…
Understanding and Accelerating the Training of Masked Diffusion Language Models
Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models (ARMs) for language modeling. However, MDMs…
Constrained latent state modeling: A unifying perspective on representation learning under competing constraints
Learning latent representations from complex data is central to modern machine learning, spanning temporal, multimodal, and partially obser…
Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slan…
PennySynth: 自動量子コード生成のための RAG 駆動のデータ合成
量子プログラミング フレームワークの複雑さの増大により、既存の大規模言語モデル (LLM) ベースのコード アシスタントの重大な制限が明らかになりました。汎用モデルでは、特殊な量子コーディングの課題に直面すると、ペニーレーン固有のゲート名が幻覚され、デバイス構成が間違って配置され、構造的に無効な回路が生成されます。我々は、PennySynth を紹介します。これは、公式 PennyLane リポジトリ、コミュニティ GitHub ソース、および QHack コンテスト アーカイブ上の 3 段階の抽出、検証、重複排除パイプラインを介して構築された、13,389 個の PennyLane 命令コード ペアの厳選されたナレッジ ベースに基づいて LLM 推論を条件付けすることで、このギャップに対処する検索拡張生成フレームワークです。 PennySynth は、自然言語からコードへの検索用にトレーニングされた st-codesearch-distilroberta-base を使用したコード認識型の埋め込み戦略を導入し、汎用ベースラインと比較して平均検索コサイン類似度を 0.45 から 0.726 に増加させます。 QHack コンテストの 3 年間 (2022、2023、2024) にわたる 74 の課題全体で評価された結果、PennySynth は QHack 2022、2023、2024 でそれぞれ 64%、68%、52% の合格率を達成し、取得なしの Claude Sonnet 4.6 よりも +28、+25、+28 パーセンテージ向上しました。ポイント。さらに、qml.* トークン パターンを重み付けする量子に適応した CodeBLEU メトリクスを導入し、構造コードの類似性と機能の正確性が量子コードの品質の異なる側面を捉えていることを示します。制御されたアブレーションにより、コードを意識した埋め込みが検索パフォーマンスの主な要因であることが明らかになり、検索品質が十分に正確であれば、データセットの拡張とソース構成によりさらなる利益がもたらされます。
原文 (English)
PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation
The growing complexity of quantum programming frameworks has exposed a critical limitation in existing large language model (LLM)-based code assistants: general-purpose models hallucinate PennyLane-specific gate names, misplace device configurations, and produce structurally invalid circuits when faced with specialized quantum coding challenges. We present PennySynth, a retrieval-augmented generation framework that addresses this gap by conditioning LLM inference on a curated knowledge base of 13,389 PennyLane instruction-code pairs, built via a three-stage extraction, verification, and deduplication pipeline over official PennyLane repositories, community GitHub sources, and QHack competition archives. PennySynth introduces a code-aware embedding strategy using st-codesearch-distilroberta-base, trained for natural-language-to-code retrieval, increasing average retrieval cosine similarity from 0.45 to 0.726 compared to a general-purpose baseline. Evaluated across 74 challenges spanning three years of the QHack competition (2022, 2023, 2024), PennySynth achieves 64%, 68%, and 52% pass@5 on QHack 2022, 2023, and 2024, respectively, improving over Claude Sonnet 4.6 without retrieval by +28, +25, and +28 percentage points. We further introduce a quantum-adapted CodeBLEU metric that upweights qml.* token patterns and show that structural code similarity and functional correctness capture distinct aspects of quantum code quality. Controlled ablations reveal that code-aware embeddings are the primary driver of retrieval performance, while dataset expansion and source composition provide additional gains when retrieval quality is sufficiently precise.
感覚調節ネットワーク:オブジェクト指向現象学の構築的基盤としての停止可能性
認知科学は、再帰と言語を説明するが形式的記号を意味に基礎づけることができない認知主義と、身体の認識を基盤とするが、生成性をサポートするほど身体の構造を詳細に指定することはほとんどない4Eアプローチとの間で依然として分裂している。私たちはこの行き詰まりが具体化されたエージェントのアーキテクチャの不完全な説明に起因していると主張し、その 1 つを提案します。それは、感覚変調ネットワーク (SMN) です。感覚変調ネットワーク (SMN) は、身体全体として考えられ、対戦相手のダイナミクスによってあらゆる解剖学的スケールで組織され、1 つの基質を通して感知して動作する感覚変調器から構築され、体全体のブロードキャスト ネットワークによってルーティングされる調整されたアクション ゾーンとペアになっている認知エージェントです。 SMN の買収には 3 つの約束があります。停止可能性(拮抗的なアフォーダンスを共活性化された平衡状態にリクルートすること)は、フッサールの意味での対象指向現象学が要求する構造的軌跡を提供します。つまり、対立は共活性化を可能にし、共活性化は停止を可能にし、停止は注意を可能にし、注意は意図的な指向性を可能にし、上部にモジュールを追加する必要はありません。自己調節可能な行動パターン (SMAP) の二重信号特性により、自己と世界の区別は、エージェントが適用するカテゴリではなく配線の構造的特徴になります。そして、基本、停止可能、交渉可能、取引の 4 つのレベルのアクション パターン階層は、自律的な規則性から公共の慣習化までの単一の軌道を示し、文法に基づいた生成性の条件をアーキテクチャの移行として特定します。 SMN は、認知主義と 4E の議論を調和させます。つまり、再帰は交渉可能な行動パターンの修正可能なダイナミクスの中に存在し、それらをサポートする相手の基質に具現化されます。暫定的な形式主義と 8 つの予測レジスタ (7 つはテスト可能、1 つは仮説) が、参照シミュレーションとともに付録に記載されています。
原文 (English)
The Sensation Modulating Network:Haltability as the architectural ground for object-directed phenomenology
We propose the Sensation Modulating Network (SMN): the cognitive agent as the whole body, organized at every scale by opponent dynamics, built from Sensation Modulators -- tissue that senses and acts through one substrate -- paired into Coordinated Action Zones routed by a body-wide broadcast. It is an inclusive model of the body, in which gravity, elasticity, and the body's topology and geometry do constructive cognitive work. The paper is scoped to what such a body constructs at its foundation -- a self-model, a world-model in that self's frame, and object-directedness -- each built by the body's physics, not assumed as a primitive. The architecture is generative: one small kit of primitives whose morphological variations (chain, sheet, tube, layered, appendicular) construct experience by the same mechanism, an invariance shown for the self-model across body plans and scales. The central thesis: haltability -- the active holding of an opponent equilibrium -- is the architectural condition object-directed phenomenology requires; a second principle, that an object is a bundle of more than one property, carries it from felt resistance to a genuine object. A companion bench realizes each construction as a runnable, falsifiable experiment with a pre-registered order parameter and matched foil. We place the principal competing accounts -- sensorimotor enactivism, active inference, and ecological and affordance-based theories -- as limiting cases within a wider landscape, stating in each case the criterion that would tell them apart, and give systems and cognitive neuroscience its place: the nervous system as the integrating core that makes the body one, not a commander over it. On this account, the cognitivism-4E impasse reflects an incomplete architecture of the embodied agent: its resolution begins not with the brain alone but with the whole body.
EvoSpec: リアルタイム語彙とパラメータ適応ターゲットによる推測的デコーディングの進化
投機的デコードは、ドラフトしてから検証するというパラダイムを通じて大規模言語モデルの推論を加速しますが、語彙サイズが拡大するにつれて出力射影層がボトルネックになります。既存の静的プルーニング手法はこのオーバーヘッドを効果的に削減しますが、動的な分布の変化を捉えることができないため、特殊なドメインやトピック切り替えのシナリオでは受け入れ率が急激に低下するという問題があります。これに対処するために、動的な語彙とパラメーターの適応を通じてドラフト モデルのリアルタイムの進化を可能にするフレームワークである EvoSpec を導入します。静的または純粋な検索ベースのアプローチとは異なり、EvoSpec は効率的なセマンティックおよび統計的なインデックス作成を通じて重要なロングテール トークンを取得するコンテキスト認識メカニズムを採用しています。さらに、ドラフトモデルとターゲットモデル間の分布のギャップを継続的に最小化するために、カリキュラム学習を利用した軽量のオンライン調整戦略を提案します。専門領域 (コーディング、法律、医学) にわたる広範な評価により、EvoSpec が静的ベースラインの制限を克服していることが確認されています。 EAGLE-3 では、これらの設定で最先端の静的ベースライン FR-Spec と比較して 1.13 倍の高速化を実現し、標準のオンライン適応よりもメモリ オーバーヘッドが 27\% 低くなります。
原文 (English)
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
Speculative decoding accelerates Large Language Model inference through draft-then-verify generation, yet lightweight draft models face coupled efficiency and quality limitations: large-vocabulary output projection is costly, while limited draft capacity and static parameters reduce acceptance under specialized or shifting inputs. Vocabulary pruning lowers projection cost, but static variants miss locally important long-tail tokens, while dynamic variants remain sensitive to preset selection policies and budgets. Moreover, limited draft capacity can leave the draft distribution misaligned even when the target token is covered. Online alignment improves draft quality, but full-parameter updates introduce substantial memory and latency overhead. We introduce EvoSpec, which jointly adapts the active vocabulary and lightweight draft parameters from verification feedback. EvoSpec asynchronously retrieves semantic and statistical token neighbors and performs curriculum-weighted online LoRA alignment while preserving exact target-model verification. On Qwen3-8B/EAGLE-2, EvoSpec reaches a $2.18\times$ speedup over vanilla decoding and a $1.20\times$ gain over EAGLE-2, while improving specialized-domain coverage and using $27\%$ less auxiliary GPU adaptation memory than full-parameter online adaptation.
SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning
Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…
HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting
Simple linear and frequency-domain models remain surprisingly competitive in long-horizon time-series forecasting, and recent mechanistic e…
CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
Frontier language capability is usually bought with frontier compute; CHERRY shows a different trade. It is a sovereign Korean model family…
DART-VLN: 離散視覚言語ナビゲーションのためのテスト時のメモリ減衰とアンチループ正則化
メモリベースの離散ビジョン言語ナビゲーション (VLN) エージェントは部分的な可観測性の下で動作する必要がありますが、強力な凍結バックボーンでさえテスト時には脆弱なままです。一般的な 2 つの障害モードは、メモリ読み出し時の古い履歴証拠と、アクション選択時の非効率なローカル バックトラッキングです。離散 VLN 用のトレーニング不要のテスト時間制御フレームワークである DART-VLN を紹介します。 DART-VLN は、保存されたコンテンツを書き換えることなく、古くなって冗長な証拠を抑制する読み取り側メモリ再重み付けルールである Test-Time Memory Decay と、アクション選択中の即時逆転を阻止する軽量のネクストホップ ペナルティである Anti-Loop Regularization を組み合わせています。このフレームワークでは、新しい学習可能なパラメーターは導入されず、学習されたバックボーンは変更されません。 R2R と REVERIE の実験では、一貫したパターンが示されています。ディケイのみでは安定した読み取り側ゲインが得られますが、ディケイ + アンチループでは全体的に最高の品質効率のトレードオフが達成され、主要な設定でより短い軌道、より短いランタイム、および改善されたナビゲーション パフォーマンスが得られます。動作分析により、アンチループ正則化によりローカル バックトラッキングが減少し、フリーズしたバックボーンの下でパス効率が向上することがさらに確認されました。全体として、この結果は、適度なテスト時間制御により、再トレーニングすることなくメモリベースの離散 VLN の信頼性と効率性を高めることができることを示しています。
原文 (English)
DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation
Memory-based agents for discrete vision-language navigation (VLN) operate under partial observability and can exhibit systematic inference-time failures even with strong pretrained backbones. We focus on two recurring problems: stale historical evidence during memory readout and inefficient local backtracking during action selection. We present DART-VLN, a training-free inference-time framework that combines Test-Time Memory Decay, which reweights stale and redundant memory slots without modifying their stored content, with Anti-Loop Regularization, a lightweight next-hop penalty that discourages immediate reversals. DART-VLN introduces no learnable parameters and leaves the navigation backbone unchanged. Experiments on R2R and REVERIE show that memory decay consistently preserves or improves task performance while reducing runtime. Adding anti-loop regularization further shortens trajectories, reduces local backtracking, and achieves the best overall balance between navigation quality and efficiency among the evaluated GridMM variants. These results indicate that lightweight inference-time control can improve the reliability and efficiency of memory-based discrete VLN without retraining.
Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild
Agentic code review, where autonomous agents provide code review comments on pull requests, is increasingly integrated into development wor…
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on s…
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound…
On Pairwise Quantile Regression - Statistical Guarantees and Applications
Quantile regression provides a powerful tool for summarizing the conditional distribution of a real-valued random variable (r.v.) of intere…
AnchorPrune: ビジュアル トークン プルーニングのための関連性に基づいたコンテキスト拡張
高解像度の入力では数千のビジュアル トークンが導入され、その多くは特定のクエリに対して冗長であるため、大規模なビジョン言語モデルにはかなりの推論コストがかかります。既存の枝刈り手法では、クエリの関連性とトークンの多様性を組み合わせることがよくありますが、これらの目的は、積極的な圧縮の下では矛盾する可能性があります。関連性主導の選択では、相関する局所的な証拠に予算が集中しすぎる可能性がありますが、多様性主導の選択では、不可欠なトークンが抑制されたり、明確ではあるが情報のない領域が保持されたりする可能性があります。最初に保護された関連性アンカーを構築し、次にそれを補完的な視覚的コンテキストで拡張する、トレーニング不要のフレームワークである AnchorPrune を紹介します。 AnchorPrune は、関連性でランク付けされたトークンのノベルティ プロファイルからアンカー サイズを適応的に決定し、クエリクリティカルな証拠のコンパクトなセットを保存し、重要度に重み付けされたノベルティを通じて残りの予算を割り当て、アンカーに関連する有益で冗長でないコンテキストを回復します。この順序付けされたデザインにより、コンテキストの拡張によって不可欠なクエリ キューが置き換えられるのを防ぎ、全体的な視覚的範囲が向上します。 AnchorPrune は軽量でアーキテクチャを認識しており、再トレーニングもモデルの変更も必要ありません。画像およびビデオの視覚言語モデルとベンチマーク全体で、特に厳しい圧縮下で、トレーニング不要のベースラインと比較して精度と効率のトレードオフを一貫して改善します。 LLaVA-NeXT-7B では、AnchorPrune は 2,880 個のビジュアル トークンのうち 160 個のみを使用して、フルトークンのパフォーマンスの 97.6% を維持します。これらの結果は、効率的なマルチモーダル推論のための効果的な原理として、関連性にアンカーされた文脈拡張を確立します。コードは https://github.com/MULTI-cau/AnchorPrune で入手できます。
原文 (English)
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine query relevance and token diversity, yet these objectives can conflict under aggressive compression: relevance-driven selection may overconcentrate the budget on correlated local evidence, while diversity-driven selection may suppress indispensable tokens or retain distinct but uninformative regions. We introduce AnchorPrune, a training-free framework that first constructs a protected relevance anchor and then expands it with complementary visual context. AnchorPrune adaptively determines the anchor size from the novelty profile of relevance-ranked tokens, preserving a compact set of query-critical evidence, and allocates the remaining budget through importance-weighted novelty to recover informative, non-redundant context relative to the anchor. This ordered design prevents contextual expansion from displacing indispensable query cues while improving overall visual coverage. AnchorPrune is lightweight, architecture-aware, and requires neither retraining nor model modification. Across image and video vision-language models and benchmarks, it consistently improves the accuracy-efficiency trade-off over training-free baselines, particularly under severe compression. On LLaVA-NeXT-7B, AnchorPrune preserves 97.6% of full-token performance using only 160 of 2,880 visual tokens. These results establish relevance-anchored contextual expansion as an effective principle for efficient multimodal inference. Code is available at https://github.com/MULTI-cau/AnchorPrune.
WILDTRACE: ロングコンテキスト推論における自然証拠痕跡のベンチマーク
長い文書にわたる複雑な質問に答えるには、情報源自体が遠く離れた文章に自然に分散していることを示す統合証拠が必要になることがよくあります。インシデントレポートでは、災害を説明する動作条件、設計上の欠陥、安全性チェックの欠如が、何十ものセクションから離れて表示されることがあります。小説では、登場人物の真の動機は、それが関連する瞬間から遠く離れたシーンを通じてのみ表面化することがあります。このソースと内部の証拠の統合は、現実世界の長い文書分析の中心となりますが、既存のベンチマークはそのほとんどを回避しています。ニードル プローブ、植えられたファクト、およびリバース エンジニアリングされたマルチホップ チェーンには、配布、配置、またはレジスタのホスト テキストとは異なる可能性のある証拠が埋め込まれているため、強力なパフォーマンスが本物のソース推論を反映しているのか、それとも配布上のアーティファクトを反映しているのかが不明確になります。 WILDTRACE は、技術的なインシデント レポートやあまり知られていない文学的な物語など、自然に発生する 214 の長い形式のソースに対する 481 のタスクのベンチマークであり、すべての証拠痕跡が文書自体の因果関係、時間的論理、および物語の論理から生じます。 Pearl の因果階層と以前のマルチホップ推論類型論を利用して、長い文書の分析読み取りにおける明確な関係要求を特徴付ける 7 つのソース内部証拠の幾何学を定義します。ソースファーストの構築パイプラインは、質問を書く前に文書構造から候補の痕跡を掘り出します。次に、各項目は、手がかりの必要性、回答の根拠、ルーブリックの忠実度、汚染耐性、回答可能性をカバーする多段階の検証を受けます。現実世界で一か八かの分析タスクがモデルに任されることが増えるにつれ、情報へのアクセスと自然に分散した証拠に基づく推論との間のギャップが、ロングコンテキスト研究の次の段階における決定的な課題として浮上しています。
原文 (English)
WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning
Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed safety check that jointly explain a disaster may appear dozens of sections apart; in a novel, a character's true motive may surface only through scenes far removed from the moment it becomes relevant. This source-internal evidence integration is central to real-world long-document analysis, yet existing benchmarks largely sidestep it. Needle probes, planted facts, and reverse-engineered multi-hop chains embed evidence that may differ from the host text in distribution, placement, or register, making it unclear whether strong performance reflects genuine source reasoning or distributional artifacts. We introduce WILDTRACE, a benchmark of 481 tasks over 214 naturally occurring long-form sources such as technical incident reports and lesser-known literary narratives, where all evidence trails arise from the document's own causal, temporal, and narrative logic. Drawing on Pearl's causal hierarchy and prior multi-hop reasoning typologies, we define seven source-internal evidence geometries that characterize the distinct relational demands of analytical reading in long documents. A source-first construction pipeline mines candidate trails from document structure before writing questions; each item then undergoes multi-stage validation covering clue necessity, answer groundedness, rubric fidelity, contamination resistance and answerability. As models are increasingly entrusted with real-world high-stakes analytical tasks, this gap between accessing information and reasoning over naturally dispersed evidence emerges as a defining challenge for the next stage of long-context research.
ドイツ語と英語のための主権のあるオープンソース基盤モデル
私たちは、ドイツ語と英語向けのソブリンのオープンソース Mixture-of-Experts (MoE) ハイブリッド Mamba Transformer 基礎モデルである Soofi S 30B-A3B を紹介します。そのハイブリッド設計は、トークンごとに 30B パラメーターのうち 3B のみをアクティブにし、コンテキストが増加しても推論キャッシュをほぼ一定に保つため、長いコンテキスト、高同時実行の展開において、高密度モデルよりも決定的なスループットの利点をもたらします。意図的に重み付けされたドイツ語を使用して約 27 兆のトークンで事前トレーニングされた Soofi S は、英語とドイツ語の集約ベンチマークで高密度の 14 ~ 27B モデルに匹敵し、17 のオープンベースモデルの中で両方の言語で最高のコード集約を達成し、アクティブパラメータがはるかに大きいものも含め、比較においてすべてのヨーロッパのソブリンベースラインを上回っています。フルオープンモデルの中で、Soofi S は Olmo 3 32B や Apertus 70B を抑えて、英語とドイツ語で最高の評価スコアを獲得しています。 Soofi S は、ミュンヘンのドイツテレコムが運用する主権 HPC スケールの AI インフラストラクチャである German Industrial AI Cloud 上にエンドツーエンドで構築されました。 Soofi S は、重み付け、選択された中間チェックポイント、完全なソースごとのデータ アカウンティング、ハイパーパラメータ、トレーニングおよび評価コードなど、非常に寛容なオープンアクセス条件に基づいてリリースされます。ソースライセンスが許可する場合、データ構築アーティファクトは寛容なライセンスの下でリリースされます。商業的にライセンスされた情報源は、集計統計と正確な混合物の計算とともに文書化されています。
原文 (English)
A Sovereign, Open-Source Foundation Model for German and English
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.
大規模言語モデル向けの効率的でプライバシーを意識したエッジ クラウド協調推論
オンデバイス LLM 推論は、応答遅延、限られたハードウェア リソース、ユーザー プライバシーというトリレンマに直面しています。完全なクラウド推論は強力なコンピューティング能力を提供しますが、ユーザー プロンプトや対話データが公開されます。一方、スタンドアロンのオンデバイス推論は、ほとんどのコンシューマ デバイスや組み込みエッジ デバイスでは実現できません。このペーパーでは、エンドポイント認証された KV キャッシュに基づいて構築されたプライバシー中心のエッジとクラウドの協調 LLM 推論フレームワークについて説明します。ローカル エンドポイントは入力前処理、埋め込み計算、適応特徴最適化、KV キャッシュ認証、投機的デコード、低次元モデル ヘッド計算を処理し、クラウドは認証済みデコーダー推論、KV キャッシュ管理、トークン検証、高次元語彙投影を実行します。エンドポイントは部分的な出力を融合し、言語に適応したマスキングを適用し、ターゲット トークンをサンプルします。すべての送信データと切り捨てられたロジットは量子化され、プライバシーを確保するために AES-GCM 暗号化され、コア軽量モジュール、ドラフト パラメーター、キャッシュ アクセス ポリシーは漏洩を避けるためにローカルに保持されます。このフレームワークは、最適化されたストリーミング、バッチ処理、量子化された ONNX 導入を通じて、CPU のみ、GPU を搭載したデバイス、組み込みデバイスなどの異種デバイスをサポートします。評価の結果、このフレームワークは、ベースラインの分割推論と比較して、トークンごとのレイテンシを最大 46.1\% 削減し、ダウンリンク ペイロードを最大 67.4\% 削減し、完全なクラウド推論と同等のパフォーマンスを維持していることが実証されています。
原文 (English)
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models
On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user prompts and dialogue data, while standalone on-device inference is unfeasible for most consumer and embedded edge devices. This paper presents a privacy-centric edge-cloud collaborative LLM inference framework built on endpoint-authenticated KV cache. Local endpoints handle input preprocessing, embedding computation, adaptive feature optimization, KV cache authentication, speculative decoding and low-dimensional model head calculation, while the cloud conducts authenticated decoder inference, KV cache management, token verification and high-dimensional vocabulary projection. Endpoints fuse partial outputs, apply language-adaptive masking and sample target tokens. All transmitted data and truncated logits are quantized and AES-GCM encrypted for privacy, with core lightweight modules, draft parameters and cache access policies kept local to avoid leakage. The framework supports heterogeneous devices including CPU-only, GPU-equipped and embedded devices via optimized streaming, batching and quantized ONNX deployment. Evaluations demonstrate that the framework reduces per-token latency by up to 46.1\% and downlink payloads by up to 67.4\% over baseline split inference, retaining comparable performance to full cloud inference.
Anatomically Faithful but Temporally Diffuse: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography
Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accuracy, and post-hoc attributio…
ロボット向けのインテリジェントクラウドエッジマルチモーダルインタラクションシステム
複雑な環境における人間とロボットの堅牢なインタラクションには、限られたオンボード コンピューティング リソースの下で、正確なジェスチャ認識、セマンティック シーンの理解、および信頼性の高いタスク計画が必要です。この論文では、強化された YOLO ベースのジェスチャ検出器と、調整されたラージ言語モデル (LLM) およびビジョン言語モデル (VLM) エージェントを統合する、クラウド エッジ マルチモーダル インタラクション フレームワークについて説明します。提案された検出器は、畳み込みブロック アテンション モジュール (CBAM) をネックに組み込み、ベースライン境界ボックス回帰目標を距離 IoU (DIoU) 損失に置き換えます。これらの修正により、複雑な背景における小さなジェスチャまたは部分的に遮蔽されたジェスチャの特徴の識別と位置特定が改善されます。クラウド層はジェスチャ検出、シーン理解、マルチモーダルフュージョン、アクションプランニングを実行しますが、TonyPi ロボットはデータ取得、通信、アクション実行、フィードバックをローカルで処理します。パブリック ジェスチャ データセットとカスタム データセットの実験では、YOLO-DC がそれぞれ 98.9% と 95.0% の精度値を達成し、mAP@0.5 値が 90.7% と 92.7% であることが示されています。システムレベルの評価では、シングルアクション、複合アクション、および視覚に依存するタスクの成功率が 95%、88%、および 82% でした。 30 人の参加者による評価では、全体の平均満足度スコアは 5 点中 3.69 でした。これらの結果は、リソースに制約のあるロボット インタラクションに対して、洗練されたジェスチャ検出とマルチモーダル エージェントを組み合わせる実現可能性を示しています。
原文 (English)
An Intelligent-Cloud Edge Multimodal Interaction System for Robots
Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents. The proposed detector, incorporates the Convolutional Block Attention Module (CBAM) into the neck and replaces the baseline bounding-box regression objective with Distance-IoU (DIoU) loss. These modifications improve feature discrimination and localization for small or partially occluded gestures in complex backgrounds. The cloud layer performs gesture detection, scene understanding, multimodal fusion, and action planning, whereas the TonyPi robot locally handles data acquisition, communication, action execution, and feedback. Experiments on a public gesture dataset and a custom dataset show that YOLO-DC achieves precision values of 98.9% and 95.0%, with mAP@0.5 values of 90.7% and 92.7%, respectively. System-level evaluation yields success rates of 95%, 88%, and 82% for single-action, composite-action, and vision-dependent tasks. A 30 participant evaluation yields an overall mean satisfaction score of 3.69 out of 5. These results demonstrate the feasibility of combining refined gesture detection with multimodal agents for resource-constrained robotic interaction.
Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
End-to-end vision-language navigation (VLN) with causal vision-language models maps instructions and egocentric observations directly to ac…
GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis,…
Towards an Automated Test of LLM Security Knowledge
Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM p…
Free energy landscape of Dense Associative Memory
Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative mem…
Riemannian Deep Learning: Modules, Networks, and Geometries
Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific…
OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have…
Co-Evolving LLM Evaluators and Policies via DynamicRubric
Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policie…
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descri…