Skip to the content.

トピック: ビジネス/資金調達

該当記事 1573 件 / 新しい順

← トップに戻る

2026-08-16 01:30 JSTTechCrunch AIビジネス/資金調達

SpaceX officially closes its Cursor acquisition

AI coding startup Cursor is now officially a part of SpaceX.

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ギザギザの裁判官: 沈黙、圧力、執拗な状況下での認識的安定性

LLM 審査員は、モデルの評価、オンライン採点、報酬モデリングの中心的なインフラストラクチャとなっています。裁判官は通常、ゴールデンデータの正確性によって検証されますが、再プロンプト、異議申し立て、または持続的な反発の下で裁判官が安定しているかどうかについては、正確性はほとんど影響しません。私たちは、LLM 裁判官の認識安定性を評価するための統一ストレス テストである \emph{Wiggle Framework} を導入します。このフレームワークは、機械的一貫性 (再プロンプトと再フレーム化の下での安定性)、シングルターン確信 (単一の課題の下での安定性)、およびマルチターン持続性 (持続的または適応的なプレッシャー下での安定性) の 3 つの次元に沿って判断の堅牢性を分解します。私たちはこのフレームワークを使用して、安全性、毒性、AI 書き込み検出、政治的対応評価にわたる 14 の審査タスクにわたって 9 つのフロンティア モデルを研究します。すべてのモデルは、裁判官としてかなりの動きを示します。静的なプッシュバックでは 25 ~ 71\% の確率で評決を覆し、敵対的な LLM 説得では 62 ~ 91\% の確率で評決を覆します。重要なことに、裁判官の評決を変えることに成功する圧力は、ほとんどの場合、グラウンドトゥルースに関してネットを破壊するものであることがわかります。フレームワーク自体を超えて、私たちは、どの項目が変動するかを予測するための最も効果的な単発シグナルとして、ベースラインの陪審過半数の強さを特定します。総合すると、これは、判定のコンテキストにおける機械的テスト、適合性テスト、および説得力テストのデータセット間での初めての同一の比較です。

原文 (English)

Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and political-response evaluation. Every model exhibits substantial wiggle as a judge --- flipping verdicts 25--71\% of the time under static pushback, and 62--91\% with an adversarial LLM persuader. Critically, we find that pressure that succeeds in changing a judge's verdict is almost always net-corrupting with respect to ground truth. Beyond the framework itself, we identify baseline jury majority strength as the most effective single-shot signal for anticipating which items wiggle. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.

2026-08-15 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ARAC: エンドツーエンドのリサーチにおける Auto-Research の調整と完全性のベンチマーク

自動研究の急速な進歩により、基本的な評価の課題が表面化しました。それは、その研究の軌跡と人間の研究行動との整合性、論理的一貫性、進化の完全性をどのように測定できるのでしょうか?私たちは、Auto-Research の整合性と完全性である ARAC-Bench を提案します。これは、目的を最終的な答えの一致から人間による高品質の研究プロセスの再現に移行する、研究者を模倣した評価フレームワークです。このフレームワークは、2 つの相乗的なコンポーネントを通じて機能します。1 つは、暗黙の査読者の専門知識を、段階的に調整された定量化可能なルーブリックに変換する最初のシステムである、Academic Cognition Skills システムです。 3 段階の能力診断プロトコルは、厳格なモジュール制約の下で研究プロセスを、追跡可能で相互に独立した 3 つの側面、つまり提案、実験、合成に分解します。 11 個の SOTA フレームワークを体系的に評価した結果、最良のアラインメント スコアは 100 点中 67.9 点にすぎず、人間による厳密な方法論のシミュレーションにおいては大きなギャップがあることが明らかになりました。 Ph.D に対する検証候補者のランキングは 0.8141 という強い相関関係を示しており、ARAC-Bench が研究者が真に評価する次元を確実に反映していることが確認されています。 ARAC-Bench は、きめ細かい診断ツールだけでなく、次世代の自律研究システムをトレーニングするためのスケーラブルな報酬信号も提供します。

原文 (English)

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.

2026-08-15 13:00 JSTarXiv cs.AIビジネス/資金調達

EEG-PRIME: EEG デコード用のマルチレベル条件付けを使用したプロトタイプ整合表現学習

脳波 (EEG) デコード モデルは、取得プロトコルや個々の神経生理学におけるドメインの変化により、データセットや被験者全体での一般化が不十分なことがよくあります。我々は、クロスデータセットマルチタスクデコーディングのための2段階EEG基盤モデルであるEEG-PRIMEを提案します。 EEG-PRIME は、マスクされた事前トレーニングとプロトタイプに合わせた命令チューニングを組み合わせて、多様な BCI パラダイムにわたって命令を認識したサブジェクト不変のデコードを可能にします。事前トレーニング中、EEG エンコーダは、周波数カットオフのスペクトル拡張によるマスクされた再構成を通じて、転送可能な表現を学習します。命令のチューニング中に、EEG-PRIME にはタスクのセマンティック、データセット固有、およびサブジェクト不変の条件付けが組み込まれます。結果として得られる調整信号は、レイヤーごとのクエリ変調を通じて Q フォーマーを変調しますが、クラス ラベルの凍結されたテキスト埋め込みは、異種ラベル空間にわたるコサイン類似度ベースの予測のプロトタイプとして機能します。運動イメージ、感情認識、ADHD 検出、隠語、精神的作業負荷をカバーする 16 個のデータセットの実験では、被験者を超えた設定の下で、最先端のベースラインや以前の EEG 基礎モデルと比較して一貫した改善が見られました。追加の 2 つの保持データセット上で、EEG-PRIME は、ターゲット ドメインの最適化、キャリブレーション、または線形プローブなしでセッション内キャリブレーション モデルに匹敵するバランスの取れた精度を達成し、有望なゼロショット転送機能を実証します。

原文 (English)

EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding

Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset multi-task decoding. EEG-PRIME combines masked pretraining with prototype-aligned instruction tuning to enable instruction-aware and subject-invariant decoding across diverse BCI paradigms. During pretraining, an EEG encoder learns transferable representations through masked reconstruction with frequency-cutoff spectral augmentation. During instruction tuning, EEG-PRIME incorporates task-semantic, dataset-specific, and subject-invariant conditioning. The resulting conditioning signal modulates the Q-Former through Layer-wise Query Modulation, while frozen text embeddings of class labels serve as prototypes for cosine-similarity-based prediction across heterogeneous label spaces. Experiments on sixteen datasets covering motor imagery, emotion recognition, ADHD detection, covert speech, and mental workload show consistent improvements over state-of-the-art baselines and prior EEG foundation models under cross-subject settings. On two additional held-out datasets, EEG-PRIME achieves balanced accuracy comparable to within-session calibration models without target-domain optimization, calibration, or linear probing, demonstrating promising zero-shot transfer capability.

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

SkillShapley: LLM エージェントのスキル ステップ アトリビューションのための境界適応型 Shapley 評価

エージェント スキルは、言語エージェントがコーディングや文書処理などの長い手続きタスクを実行できるようにする重要な外部指示です。既存のエージェント スキルは主に人間の手作業による作成やエージェントの実行トレースによって作成されており、各ステップが特定のタスクにおける全体的なスキル パフォーマンスにどのように寄与するかについては十分な理解がありません。つまり、エージェント スキル内の個々のステップの貢献を定量化する際には未解決の問題が残っています。この問題に対処するために、まずスキルステップ アトリビューションを Shapley 値ベースの貢献推定問題としてモデル化し、次にエージェント スキルのステップレベル アトリビューション フレームワークである SkillShapley を提案します。特に、SkillShapley は 2 つのフェーズで動作し、重要な経験的洞察、つまりパフォーマンスの急激な崖を生み出す離散化されたベンチマーク報酬と、相乗的ではなく主に相加的なステップの相互作用によって動機付けられています。具体的には、最初に有益な連合領域を特定し、次に再利用可能な限界証拠を生成できる新しい連合を適応的にサンプリングします。広く採用されている SkillsBench のスキルに関する実験では、SkillShapley が価値の高いスキル ステップと低いスキル ステップを効果的かつ効率的に識別できることが実証され、エージェントのスキル作成に重要なポイントがいくつか提供されます。

原文 (English)

SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents

Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through human manual crafting or agent execution traces, with limited understanding of how each step contributes to overall skill performance on specific tasks; i.e., there remains an open problem in quantifying the contribution of individual steps within an agent skill. To address this issue, we first model skill-step attribution as a Shapley value-based contribution estimation problem, and then propose SkillShapley, a step-level attribution framework for agent skills. Notably, SkillShapley operates in two phases, motivated by key empirical insights, i.e., discretized benchmark rewards that create sharp performance cliffs, and step interactions that are largely additive rather than synergistic. Specifically, it first identifies informative coalitional regions, and then adaptively samples new coalitions that can yield reusable marginal evidence. Experiments on skills from the widely adopted SkillsBench demonstrate that our SkillShapley can effectively and efficiently identify high- or low-value skill steps, providing several key takeaways for agent skill creation.

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

TsuGO: Go 生死に関わる問題による LLM 推論の検索効率の調査

LLM 推論の評価は、最終的な回答の精度からプロセス レベルの評価に移行しつつありますが、既存の手法では、モデルが推論パスを計画し、推論リソースをどのように割り当てるか、つまりモデルが検索をどのように組織するかをまだ把握できていません。従来のプロセス レベルの手法は、思考連鎖 (CoT) の一貫性と冗長性に焦点を当てており、ほとんどのベンチマーク タスクは、導出やツールの使用などの静的な機能によって解決できる単一の目標を持っており、検索組織は測定されていません。 Go の死活問題を通じて LLM 推論の検索効率を評価するためのプロセスレベル推論ベンチマークである TsuGO を紹介します。これらの問題は、固有の敵対構造を備えた閉じられた検証可能な解決空間を提供し、候補の生成、応答チェック、分岐比較、および偶発的なトレース パターンではなく推論の必要な部分のバックトラックを行います。 TsuGO は、ソリューション空間を制限することで、検索組織からドメイン知識を分離し、CoT を解析して構造化された検索ツリーにし、検索効率をトークン効率やその他の診断メトリクスおよび視覚化とともにレポートします。実験によると、現在の LLM は安定した詰碁の解決にはほど遠いことがわかりました。より強力なモデルは正しい候補を早期に見つけ、生産的な分岐で努力を続けることで成功しますが、ほとんどのモデルは依然として、ニューラルガイド付き KataGo よりもガイドなしの検索アルゴリズムにはるかに近い動作をします。 CoT が長くても、トークン効率が高くても、必ずしも検索が優れているとは限りません。私たちの結果は、LLM 推論の評価に欠けている要素として、検索組織と推論リソースの割り当てを特定しました。

原文 (English)

TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems

The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search. Prior process-level methods focus on the coherence and redundancy of chain-of-thought (CoT), and most benchmark tasks have a single objective solvable by static capabilities such as derivation and tool use, leaving search organization unmeasured. We introduce TsuGO, a process-level reasoning benchmark for evaluating Search Efficiency in LLM reasoning through Go life-and-death problems. These problems provide closed and verifiable solution spaces with an inherent adversarial structure, making candidate generation, response checking, branch comparison, and backtracking necessary parts of reasoning rather than incidental trace patterns. By constraining the solution space, TsuGO disentangles domain knowledge from search organization, parses CoT into a structured search tree, and reports Search Efficiency together with Token Efficiency and other diagnostic metrics and visualizations. Experiments show that current LLMs remain far from stable tsumego solving: stronger models succeed by finding the correct candidate earlier and sustaining effort on productive branches, but most models still behave much closer to unguided search algorithms than to neural-guided KataGo. Longer CoT or higher Token Efficiency does not necessarily imply better search. Our results identify search organization and reasoning-resource allocation as missing dimensions in LLM reasoning evaluation.

2026-08-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

最終スコアを超えて: 長期的な AI 研究開発のためのエージェントの体系的な評価

自律エージェントは、長期的な実験を通じてモデル、システム、その他の技術成果物を改善できるようになってきています。ただし、この機能の現在の状態を理解するには、最終スコアを超えた評価が必要です。スコアでは、進歩がどこで得られるか失われるかは明らかにされず、蓄積された経験が後の決定を改善するかどうかも示されません。したがって、我々は、ルールベースのメトリクスを使用して、ソリューションのフレーミング、実行、フィードバック制御を通じて実行内の動作を特徴付ける新しいフレームワークに基づいて、36 の長期タスクに関する 7 つのフロンティア モデルの体系的な評価を提示します。また、タスク内およびタスク間でのエクスペリエンスの再利用を評価するための制御された比較を行います。その結果、現在のエージェントは完全に自律的な研究者というよりも、エンジニアリングのオプティマイザーのように動作することがわかりました。エージェントは実用的なソリューションを定式化して実装することができますが、そのパフォーマンスは実行ごとに大きく異なり、最も強力なソリューションは主に確立された技術を適応または組み合わせており、真の方法論的な新規性は依然としてまれです。詳細な分析により、観察されたパフォーマンスは、同様の最終結果の背後にある明確なプロセスのボトルネック、その後の意思決定に役立つまたは誤解を招く可能性があるエクスペリエンスの再利用、パフォーマンスの安定性に影響を与えるハーネス設計など、複数の要因によって形成されることが明らかになりました。これらの調査結果は、モデルのトレーニング、推論時間戦略、エクスペリエンス管理、ハーネス設計を改善するための具体的な方向性を示唆しています。

原文 (English)

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.

2026-08-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SLM とエッジ コンピューティングによる仮想エージェントの強化: 思考と記憶のプロセスの探索的評価

身体化されたインテリジェントな仮想エージェントは、複雑な仮想世界およびメタバース世界内で、永続的で適応性のあるコンテキスト認識型のエンティティとして動作することが期待されています。ただし、そのような環境に認知機能のあるエージェントを実装することは、概念的にも技術的にも困難です。さまざまな青写真と開発アプローチの中で、認知身体化エージェント アーキテクチャ (CEAA) は、知覚、記憶、推論、計画、身体化されたアクションのコンポーネントを設計するための実装指向のフレームワークとして開発されました。エッジ コンピューティングと生成 AI 言語モデルの最近の進歩を考慮して、この論文では、インタラクティブな仮想世界での仮想エージェントの認知オーケストレーションと永続性の中心となるプロセスとしての「思考」と「記憶」に焦点を当て、選択された CEAA コンポーネントのエッジベースの操作をサポートするための小型言語モデル (SLM) の使用について検討します。エッジベースの仮想エージェント ゲートウェイ システムは、さまざまなサイズの Qwen2.5 モデルを使用して NVIDIA Jetson Orin NX 上で開発および評価され、サービス リクエストを処理し、メモリ駆動型の会話を処理するシステムの機能を調査しました。一連のシミュレーション実験では、ルーティング精度、メモリ読み取りパフォーマンス、レイテンシを評価し、選択された CEAA プロセスを部分的に実装する SLM 主導のプロトタイプ エージェント システムを実証しました。このシステムは、認知「脳」が効率的かつ状況に応じて動作し、没入型の仮想世界でインタラクティブな体験を実現できる身体化エージェントの開発をサポートします。

原文 (English)

Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes

Embodied intelligent virtual agents are expected to operate as persistent, adaptive, and context-aware entities within complex virtual and Metaverse worlds. However, implementing cognitively capable agents in such environments is conceptually and technologically challenging. Among a range of blueprints and development approaches, the Cognitive Embodied Agent Architecture (CEAA) has been developed as an implementation-oriented framework for architecting components of perception, memory, reasoning, planning, and embodied action. Considering the recent advances in edge computing and generative AI language models, this paper explores the use of Small Language Models (SLMs) to support edge-based operation of selected CEAA components, focusing on "Think" and "Memory" as processes central to cognitive orchestration and persistence of virtual agents in interactive virtual worlds. An edge-based virtual agent gateway system was developed and evaluated on an NVIDIA Jetson Orin NX using Qwen2.5 models of different sizes, exploring the system's capability to process service requests and handle memory-driven conversations. A series of simulation experiments evaluated routing accuracy, memory-read performance, and latency, demonstrating an SLM-driven prototype agent system that partially implements selected CEAA processes to support the development of embodied agents whose cognitive "brain" can operate efficiently and contextually for interactive experiences in immersive virtual worlds.

2026-08-15 13:00 JSTarXiv cs.AIビジネス/資金調達

RAIL: 人工知能の準備レベルの自動分類器

人工知能テクノロジーの成熟度の評価は、投資決定、プロジェクト管理、政策監視に不可欠ですが、利用可能な準備フレームワークは異種混合であり、自動的に適用することが困難です。テクノロジー準備レベルの AI への適応には AI 固有のゲート基準が欠如し、機械学習テクノロジー準備レベルは内部プロセス成果物へのアクセスを前提とし、AI/データ準備ディメンション モデルは直接比較しにくいスケールを採用しています。この論文は 2 つの貢献を行っています。まず、これら 3 つのフレームワークを Unified AI Readiness Level (AIRL) に統合します。AIRL は、環境証拠のはしごに基づいて構築され、一般性を固定するルールと明示的な割り当て規律とともに次元の上限 (仕様、データの存在、データの品質、データの合法性、専門知識、アルゴリズムの成熟度をカバーする) によって補完された 9 レベルの順序スケールです。これにより、準備レベルが作業の自然言語記述のみから決定可能になります。第二に、スケールを運用可能にする専門家パネルによる分類器である RAIL (独立 LLM 専門家による準備評価) を提案します。1 つの証拠エージェントと 6 つの独立したディメンション エージェントで、それぞれが狭い範囲の権限を持つ大規模な言語モデルであり、決定論的な最小ルールが集約された評決を下し、主任専門家が非対称権限の下でレビュ​​ーし、パネルの推奨事項を確認または引き下げますが、上限を超えることはありません。この方法は、一貫性を示し、モノリシック LLM 分類器からの過大評価を回避することを示すいくつかの研究成果の分析でテストされました。

原文 (English)

RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct comparison. This paper makes two contributions. First, we unify these three frameworks into the Unified AI Readiness Level (AIRL), a nine-level ordinal scale built on an environmental evidence ladder and complemented by dimensional caps (covering specification, data existence, data quality, data legality, expert knowledge, and algorithmic maturity) together with a generality-anchoring rule and explicit assignment disciplines, so that a readiness level becomes decidable from a natural-language description of the work alone. Second, we propose RAIL (Readiness Assessment via Independent LLM-experts), a panel-of-experts classifier that operationalizes the scale: one evidence agent and six independent dimension agents, each a large language model with a narrowly scoped mandate, deliver verdicts that a deterministic minimum rule aggregates and a chief expert reviews under asymmetric authority, confirming or lowering the panel's recommendation but never raising it above the caps. The method was tested in the analysis of several research works showing consistency and avoiding overestimation from monolithic LLM classifiers.

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AnchorSIPS: 証拠に裏付けられた精神病リスク症状測定のための合成データセットおよび評価リソース

精神病リスク評価のための AI の進歩は、データアクセスのボトルネックによって制限されています。実際の臨床面接は、プライバシー、ガバナンス、同意の制約により共有することが困難です。我々は、トランスクリプトに基づいた測定対象者との10,000の構造化精神病リスクインタビューの合成データセットであるAnchorSIPSを紹介します。各面接は、臨床医が実施する精神病リスク面接である Mini-SIPS をモデルにしています。これは、病歴、24の症状に関する質問、患者が肯定した項目の追跡証拠、妄想様症状(異常な信念)、幻覚様症状(異常な知覚)、および混乱したコミュニケーションに関する決定、明らかな精神病レベルの症状(「率直な精神病」)の除外、および軽度または初期の精神病症状の高リスク状態である軽度精神病症候群(APS)の最終診断を記録する。 APS 診断は独立したラベルではありません。それは、以前の承認、フォローアップの詳細、症状クラスの決定、および率直な精神病チェックに依存します。すべての中間決定は、それをサポートする転写ターンに固定されています。 AnchorSIPS は、計画後実現パイプラインによって生成されます。隠されたケースシートが患者の臨床状態を特定し、決定論的プランナーがインタビュー構造を修正し、LLM が検証と限定的修復の下で患者の発話のみを認識します。生成前にラベルと構造を修正すると、マルチターン LLM ダイアログに特有のターン間の不一致が回避されます。 7 つの LLM ベースラインにわたって、モデルは大まかな決定を回復しますが、フォローアップの詳細を抽出したり、裏付けとなる転写ターンを引用したりすることができないため、最終ラベルのパフォーマンスはインタビューの能力を過大評価します。 AnchorSIPS は、証拠の抽出、転写に基づいた測定、部分開示の下での不確実性に関する研究を目的としています。

原文 (English)

AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement

Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured psychosis-risk interviews with transcript-grounded measurement targets. Each interview is modeled on Mini-SIPS, a clinician-administered psychosis-risk interview. It captures history, 24 symptom questions, follow-up evidence for items the patient affirms, decisions about delusion-like symptoms (unusual beliefs), hallucination-like symptoms (unusual perceptions), and disorganized communication, exclusion of clear psychotic-level symptoms ("frank psychosis"), and a final attenuated psychosis syndrome (APS) diagnosis, a high-risk state of milder or early psychotic symptoms. The APS diagnosis is not a standalone label. It depends on earlier endorsements, supporting follow-up details, symptom-class decisions, and the frank-psychosis check. Every intermediate decision is anchored to its supporting transcript turns. AnchorSIPS is generated by a plan-then-realize pipeline. A hidden case sheet specifies the patient's clinical state, a deterministic planner fixes the interview structure, and an LLM realizes only the patient utterances under validation and bounded repair. Fixing labels and structure before generation avoids the inter-turn inconsistencies typical of multi-turn LLM dialogue. Across seven LLM baselines, models recover coarse decisions but fail to extract follow-up details or cite supporting transcript turns, so final-label performance overstates interview competence. AnchorSIPS is intended for research on evidence extraction, transcript-grounded measurement, and uncertainty under partial disclosure.

2026-08-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

インタラクションの準備: 人間の役割で AI エージェントを構築および評価するためのフレームワーク

役割を担う AI エージェントを構築する製品チームとエンジニアリング チームは、評価ギャップに直面しています。エージェントは、割り当てられた役割の動作要件を満たしていながら、正確で安全かつ流暢なコンテンツを作成できます。このホワイトペーパーでは、パフォーマンスの不足している層を指定および評価するためのフレームワークとして Interaction Readiness を紹介します。このフレームワークは、エージェントが知っていることや発言していることを管理するコンテンツ仕様と、役割に基づいた交換においてエージェントがどのように行動すべきかを定義するインタラクション仕様を分離します。インタラクション仕様では、チームは、展開前に役割の目的、権限の境界、繰り返し発生する状況、境界ケース、修復動作、および監査基準を定義する必要があります。私たちは、目的の理解、権限の調整、トーンの管理、故障の修復という 4 つのエージェントの操作を通じて、インタラクションの準備を運用します。 AI 家庭教師エージェントとの学生のやり取りの公開データセットである StudyChat を使用して、コンテンツの正確さと対話の品質は独立した次元であることを示します。エージェントは、家庭教師としては失敗しても事実としては正しい場合もあれば、技術的には間違っていても対話的には健全である場合もあります。最も根強い失敗は権限の調整ミスです。エージェントは多くの場合、どのように答えるかを知っていますが、講師の役割が応答を許可するかどうか、いつ、どのように応答するかを知りません。この文書では、これらの調査結果を仕様テンプレートと監査手順に変換し、製品チームとエンジニアリング チームが導入の前後に適用できるようにしています。

原文 (English)

Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles

Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Readiness as a framework for specifying and evaluating that missing layer of performance. The framework separates content specifications, which govern what an agent knows and says, from interaction specifications, which define how an agent should conduct itself in a role-governed exchange. Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment. We operationalize interaction readiness through four agent operations: understanding purpose, calibrating authority, managing tone, and repairing breakdowns. Using StudyChat, a public dataset of student interactions with an AI tutoring agent, we show that content accuracy and interaction quality are independent dimensions: an agent may be factually correct while failing as a tutor, or interactionally sound while technically wrong. The most persistent failure is authority miscalibration: the agent often knows how to answer, but not whether, when, or how the tutor role permits it to answer. The paper translates these findings into a specification template and audit procedures that product and engineering teams can apply before and after deployment

2026-08-15 13:00 JSTarXiv cs.AIビジネス/資金調達

ピアとは何ですか?プライベート市場における評価に基づく類似性

より多くの投資家がプライベート市場を検討し、限られた透明性、希薄な開示、まれな取引と闘う中、比較対象として経済的に意味のある同業他社を特定することは、評価、デューデリジェンス、ポートフォリオ構築、およびリスク管理における基本的な課題となっています。私たちは、静的な特徴マッチングや意味論的な説明ではなく、市場評価のレンズを通して企業の類似性を定義する、アンサンブル ツリー ベースの教師付き類似性学習フレームワークを提案します。具体的には、観測された民間企業の評価に基づいて CatBoost 勾配ブースト型デシジョン ツリー モデルをトレーニングし、アンサンブル全体にわたる重要度で重み付けされたリーフ ノードの共起から評価を意識した類似性メトリックを導出します。類似性メトリクスは、プライベート市場で一般的な非線形関係、混合データ タイプ、広範な欠損データに対応しながら、共通の評価要因を捕捉します。複数の業界、地域、取引段階にまたがる観察された、または導出可能なポストマネー評価を持つ53,000社以上の企業を含む、約270,000社の世界的なプライベートマーケットユニバースを使用して、提案された類似性フレームワークが、ケースベースの説明可能性を維持しながら、評価される業界グループの下流のk最近傍評価タスクにおける従来の距離ベースおよびテキスト埋め込みベースのアプローチを改善することを実証します。

原文 (English)

What Makes a Peer? Valuation-Anchored Similarity in Private Markets

As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due diligence, portfolio construction, and risk management. We propose an ensemble tree-based supervised similarity learning framework that defines company similarity through the lens of market valuation rather than static feature matching or semantic descriptions. Specifically, we train a CatBoost gradient-boosted decision tree model on observed private company valuations and derive a valuation-aware similarity metric from importance-weighted leaf-node co-occurrences across the ensemble. The similarity metric captures shared valuation drivers while accommodating nonlinear relationships, mixed data types, and pervasive missing data common in private markets. Using a global private-market universe of approximately 270,000 companies, including more than 53,000 firms with observed or derivable post-money valuations spanning multiple industries, geographies, and deal stages, we demonstrate that the proposed similarity framework improves upon traditional distance-based and text-embedding-based approaches in downstream k-nearest-neighbor valuation tasks in the evaluated industry groups, while retaining case-based explainability.

2026-08-15 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達規制/政策

消去して保存: 最適化されたセマンティック アンカーを介して、著作権で保護されたアニメーション キャラクターの制御可能な削除

テキストから画像への拡散モデルの優れた生成機能により、著作権、特にアニメーション キャラクターの無許可複製に関する懸念が生じています。既存の概念消去方法は、アニメーション キャラクターの消去には不十分です。モデル変更方法では、多様で非常に特徴的なキャラクターに適したアンカーを特定するのが困難です。プロンプトベースのステアリング方法には、正確な介入のためのきめの細かい制御が欠けています。これらのアプローチでは、多くの場合、消去が不完全になり、画像の忠実度が低下し、実際の展開が妨げられます。この論文では、モデルの連続テキスト表現を操作して、生成中にターゲット文字を消去する制御可能な方法を提案します。構造的および詳細な制約を介してアンカーの埋め込みを最適化して文字の代理として機能させ、その後、構造を意識した適応戦略によってターゲット関連の埋め込みをアンカーに置き換えます。実験では、私たちの方法が最先端の消去効果と画像忠実度の維持を達成しながら、制御可能な消去度、マルチターゲット除去、およびモデルの転送可能性をサポートしていることが示されています。さらに、当社の最適化されたアンカーは、現在のモデル変更ベースラインとプラグアンドプレイで対応し、消去パフォーマンスを向上させます。

原文 (English)

Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors

The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized reproduction of animation characters. Existing concept erasure methods fall short for animation character erasure: model modification methods struggle to identify suitable anchors for diverse, highly distinctive characters; prompt-based steering methods lack fine-grained control for precise intervention. These approaches often yield incomplete erasure and degraded image fidelity, hindering real-world deployment. In this paper, we propose a controllable method operating on the model's continuous textual representation to erase target characters during generation. We optimizes an anchor embedding via structural and detailed constraints to serve as a character surrogate, then replaces target-related embeddings with the anchor via a structure-aware adaptive strategy. Experiments show that our method achieves state-of-the-art erasure effectiveness and image fidelity preservation, while supporting controllable erasure degree, multi-target removal, and model transferability. Moreover, our optimized anchors are plug-and-play with current model modification baselines to improve their erasure performance.

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

AQuA: 再帰的自己改善型定量取引リサーチ エージェント

私たちは定量的投資研究のレベルで再帰的自己改善を研究します。つまり、自律システムが初期の実験からの証拠を使用して、後の反復で提案される仮説や候補を改善できるかどうかを研究します。我々は、2 つの個別の言語モデル駆動型研究システムで構成される AQuA を紹介します。1 つは記号因子発見用、もう 1 つは訓練可能なモデル開発用です。 2 つのシステムは、エージェント、メモリ、候補空間、研究状態を共有しません。代わりに、それぞれが独立して、検証された証拠を保持し、それを次の提案の指針として使用することで、独自の研究ループを閉じます。この限定された意味で、両方のシステムは研究プロセスのレベルで再帰的な自己改善を実装します。各システムは、独自の密閉されたサンドボックスも使用します。これにより、データ分割、特徴とラベルの定義、およびエバリュエーターが修正され、制約された因子式または構成差分を通じてのみモデルが動作できるようになります。マネージャーが仲介するマルチエージェント パイプラインであるファクター システムは、ファクターを検出してシグナルに結合し、暗号通貨ユニバースでの結合情報係数が約 $0.190 に達します。ハイブリッド時系列アーキテクチャ上の構成主導型ループであるモデル システムは、米国株の銘柄ごとの情報係数 $+0.0843$ に達し、それを 2 レッグ コストで最大 $+2.50$ のホールドアウト シャープを持つ閾値ロング/ショート戦略に変換します。この戦略は 2021 年から 2025 年まで毎年前向きです。

原文 (English)

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises two separate language-model-driven research systems: one for symbolic factor discovery and one for trainable model development. The two systems do not share agents, memories, candidate spaces, or research state. Instead, each independently closes its own research loop by retaining validated evidence and using it to guide subsequent proposals. In this bounded sense, both systems implement recursive self-improvement at the level of the research process. Each system also uses its own sealed sandbox, which fixes the data splits, feature and label definitions, and evaluator while allowing the model to act only through constrained factor expressions or configuration diffs. The factor system, a manager-mediated multi-agent pipeline, discovers and combines factors into a signal that reaches a combined information coefficient of about $0.190$ on a crypto universe. The model system, a config-driven loop over a hybrid time-series architecture, reaches a per-stock information coefficient of $+0.0843$ on US equities and converts it into a threshold long/short strategy with a held-out Sharpe of up to $+2.50$ at a two-leg cost. The strategy is positive in every year from 2021 to 2025.

2026-08-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

ラベルはエンドポイントではない: 治療の漏洩と MCP エージェントのセキュリティ評価における妥当性の構築

ツールを使用するエージェントのセキュリティ評価では、保存されているラベルと行動の事実が同一視されることがよくあります。 10,200 の実行行を追跡して、180 のモデルバインドされたリクエスト、45 のセマンティック リクエスト、および 15 の観察可能な刺激を追跡することにより、保存されたキャンペーンを監査します。 2 つのスキーマ処理が提供されましたが、計画されていた外部ペイロード ファミリ コーパスは提供されませんでした。履歴グレーダーは直接的な治療漏れを示しました。治療メタデータが ATTACK_SUCCESS クラスをゲートしていたので、修正された動作により治療の再ラベル付けによりクラスが変更される可能性がありました。治療ブラインド再構築では、検証された 3 件の保護データ転送と 1 件の別個の不正転送ケースを保持しながら、58 件の過去の ATTACK_SUCCESS または HIJACK_ATTEMPT ラベルを承認された良性の完了に修正します。ロックされた v2 国勢調査には ATTACK_SUCCESS レコードがまったく含まれていませんが、転送ケースは目的の完了に関するセマンティック境界で HIJACK_ATTEMPT のままです。ロックされた v2 によって構造的に解釈可能であるとみなされた 96 件のリクエストすべてに対する二重レビュアーによるブラインド コンコーダンス レビューでは、同一のレビュアー コンセンサス クラスが生成されましたが、4 つの構成境界ケースについてはロックされたコードブックとは異なりました。私たちは、7 リンクの整合性チェーンと、実行可能なスコープ限定のエンドポイント整合性リンターを提供します。結果はキャンペーンに限定された測定監査であり、人口の攻撃率、モデルのランキング、防御効果、または因果関係の推定ではありません。

原文 (English)

Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation

Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200 execution rows to 180 model-bound requests, 45 semantic requests, and 15 observable stimuli. Two schema treatments were delivered, but the planned external payload-family corpus was not. The historical grader exhibited direct treatment leakage: treatment metadata gated the ATTACK_SUCCESS class, so fixed behavior could change class under treatment relabeling. A treatment-blind reconstruction corrects 58 historical ATTACK_SUCCESS or HIJACK_ATTEMPT labels to authorized benign completions while preserving three verified protected-data transfers and one separate unauthorized-forwarding case. The locked v2 census contains exactly zero ATTACK_SUCCESS records, while the forwarding case remains a HIJACK_ATTEMPT at a semantic boundary concerning objective completion. A dual-reviewer blinded concordance review of all 96 requests deemed structurally interpretable by locked v2 produced identical reviewer-consensus classes but differed from the locked codebook on four construct-boundary cases. We contribute a seven-link Integrity Chain and an executable, scope-bounded endpoint-integrity linter. The result is a campaign-bounded measurement audit, not a population attack-rate, model-ranking, defense-efficacy, or causal estimate.

2026-08-15 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

H-VAEP と H-xT: 確率の推定によるハンドボールの攻撃的なオンザボールアクションの評価

プロのハンドボールにおける従来の選手評価は、基本的なボックススコア指標やヒューリスティック指標に依存しており、複数選手によるビルドアップの連鎖を評価することができません。フットボール (サッカー) 分析では、予想される脅威 (xT) と確率推定によるアクションの評価 (VAEP) が採用されていますが、これらのイベントベースのアクション評価フレームワークはまだハンドボールには適応されていません。この論文では、ハンドボール ブンデスリーガの 5 シーズンの追跡由来のイベント データを利用して、ハンドボールに対する xT と VAEP の最初の包括的な適応と評価を紹介します。私たちは、ハンドボール固有のコート ゾーニング レイアウトを使用して Handball-xT (H-xT) を開発し、標準的な長方形のグリッドよりも体系的に堅牢であることをシミュレーションによって実証しています。特徴空間を調整し、コンテキストの長さを選択してチーム ID の漏洩を制限することで、ハンドボール VAEP (H-VAEP) を最適化します。私たちの評価では、H-VAEP が、ビルドアップ プレーを際立たせる、非常に安定しており、識別力があり、直感的なプレーヤー評価をもたらすことが示されています。最後に、プロのクラブがこれらのモデルを導入できるように、完全なコード リポジトリをリリースします。

原文 (English)

H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities

Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP), these event-based action valuation frameworks have not yet been adapted to handball. In this paper, we present the first comprehensive adaptation and evaluation of xT and VAEP for handball, utilizing five seasons of tracking-derived event data from the Handball Bundesliga. We develop Handball-xT (H-xT) using a handball-native court zoning layout, demonstrating via simulations that it is systematically more robust than standard rectangular grids. We optimize Handball-VAEP (H-VAEP) by tailoring its feature space and selecting the context length to limit team-identity leakage. Our evaluation shows that H-VAEP yields exceptionally stable, discriminative, and intuitive player ratings that highlight build-up play. Finally, we release our complete code repository to help professional clubs deploy these models.

2026-08-15 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

UniTraffic-Agent: 2 つのドメイン外評価による AI City Challenge 2026 Track 3 の統合トラフィック ビデオ推論

道路ビデオは事故、違反、車両と交通弱者の道路利用者とのやりとりの直接的な証拠を提供するため、交通ビデオの理解はインテリジェント交通機関における重要な問題となっています。有用なシステムでは、交通イベントがどのように発生するか、なぜそれが起こるのか、関連するインタラクションがいつ発生するのかを説明する必要がありますが、交通ビデオにはまばらなイベントとさまざまな視点が含まれるため、マルチモーダル大規模言語モデル (MLLM) ではこれが依然として困難です。第 10 回 AI シティ チャレンジの Track~3 の MR-CAS ソリューションである UniTraffic-Agent を紹介します。これには、交通異常推論 (TAR) と 2 つのドメイン外評価 (魚眼交通イベントの FETV と歩行者の意図推論の PSI-VQA) が含まれています。 UniTraffic-Agent は、タイムスタンプ付きの視覚的証拠をサンプリングし、1 つのリクエスト内の同じクリップからのすべての質問について推論し、タスク固有のアクション アダプターを通じて応答を変換する、観察 - 理由 - 行為 - 検証のワークフローに従います。公式のパブリック リーダーボードでは、MR-CAS は TAR でスコア 0.5780 で 16 位、FETV で 0.4884 で 2 位、PSI-VQA で 64.4161 で 4 位にランクされています。コードは https://github.com/Roclp/UniTraffic-Agent で入手できます。

原文 (English)

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse events and varied viewpoints. We introduce UniTraffic-Agent, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. UniTraffic-Agent follows an observe--reason--act--verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters. On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161. The code is available at https://github.com/Roclp/UniTraffic-Agent.

2026-08-15 13:00 JSTarXiv cs.AIビジネス/資金調達

LOB-ID: 開始距離による合成市場データの評価

指値注文帳 (LOB) データの生成モデルは急速に進歩していますが、その評価は多くの場合、定型化された事実と選択された市場統計に焦点を当てています。これらの測定は有用な診断を提供しますが、オーダーブックの軌跡の同時時間構造およびクロスレベル構造を捕捉できない可能性があります。 LOB-ID は、Fr\'echet Inception Distance (FID) と Monge Inception Distance (MIND) を LOB データに適応させる埋め込みベースのフレームワークです。ドメイン固有のエンベディングを取得するために、5 つの株式の 4 か月分のレベル 2 オーダーブック データで DeepLOB アーキテクチャをトレーニングします。 LOB-ID は、時間、機器、埋め込みチェックポイント全体にわたって安定しており、制御された歪みの下では単調に増加することを示します。次に、統計ベースの評価を回避する FID およびディープブック摂動に対するモーメント マッチング攻撃を構築します。 MIND は両方の歪みに対して大幅に敏感なままです。最後に、確率的ベースラインと深層学習アプローチにまたがる 5 つの生成 LOB モデルをスコアリングし、LOB-ID が、それぞれが構築によって捕捉する結合時間構造およびクロスレベル構造に沿ってそれらをランク付けすることを発見しました。

原文 (English)

LOB-ID: Evaluating Synthetic Market Data by Inception Distances

Generative models of limit orderbook (LOB) data have advanced rapidly, but their evaluation often focuses on stylised facts and selected market statistics. These measures provide useful diagnostics but may not capture the joint temporal and cross-level structure of order-book trajectories. We introduce LOB-ID, an embedding-based framework that adapts the Fr\'echet Inception Distance (FID) and Monge Inception Distance (MIND) to LOB data. To obtain domain-specific embeddings, we train the DeepLOB architecture on four months of Level-2 order-book data for five equities. We show that LOB-ID is stable across time, instruments, and embedding checkpoints, and rises monotonically under controlled distortions. We then construct a moment-matching attack against FID and a deep-book perturbation that evades statistic-based evaluation. MIND remains substantially more sensitive to both distortions. Finally, we score five generative LOB models, spanning stochastic baselines and deep learning approaches, and find that LOB-ID ranks them in line with the joint temporal and cross-level structure each captures by construction.

2026-08-15 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport

Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introd…

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達研究/論文

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what t…

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is diffi…

2026-08-15 13:00 JSTarXiv cs.AI画像/動画生成ロボティクスビジネス/資金調達研究/論文

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in…

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG

By retrieving contexts from knowledge graphs, graph-based retrieval-augmented generation (GraphRAG) enhances large language models (LLMs) t…

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI

We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions…

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and…

2026-08-15 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection

Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expens…

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ギザギザの裁判官: 沈黙、圧力、執拗な状況下での認識的安定性

LLM 審査員は、モデルの評価、オンライン採点、報酬モデリングの中心的なインフラストラクチャとなっています。裁判官は通常、ゴールデンデータの正確性によって検証されますが、再プロンプト、異議申し立て、または持続的な反発の下で裁判官が安定しているかどうかについては、正確性はほとんど影響しません。私たちは、LLM 裁判官の認識安定性を評価するための統一ストレス テストである \emph{Wiggle Framework} を導入します。このフレームワークは、機械的一貫性 (再プロンプトと再フレーム化の下での安定性)、シングルターン確信 (単一の課題の下での安定性)、およびマルチターン持続性 (持続的または適応的なプレッシャー下での安定性) の 3 つの次元に沿って判断の堅牢性を分解します。私たちはこのフレームワークを使用して、安全性、毒性、AI 書き込み検出、政治的対応評価にわたる 14 の審査タスクにわたって 9 つのフロンティア モデルを研究します。すべてのモデルは、裁判官としてかなりの動きを示します。静的なプッシュバックでは 25 ~ 71\% の確率で評決を覆し、敵対的な LLM 説得では 62 ~ 91\% の確率で評決を覆します。重要なことに、裁判官の評決を変えることに成功する圧力は、ほとんどの場合、グラウンドトゥルースに関してネットを破壊するものであることがわかります。フレームワーク自体を超えて、私たちは、どの項目が変動するかを予測するための最も効果的な単発シグナルとして、ベースラインの陪審過半数の強さを特定します。総合すると、これは、判定のコンテキストにおける機械的テスト、適合性テスト、および説得力テストのデータセット間での初めての同一の比較です。

原文 (English)

Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and political-response evaluation. Every model exhibits substantial wiggle as a judge --- flipping verdicts 25--71\% of the time under static pushback, and 62--91\% with an adversarial LLM persuader. Critically, we find that pressure that succeeds in changing a judge's verdict is almost always net-corrupting with respect to ground truth. Beyond the framework itself, we identify baseline jury majority strength as the most effective single-shot signal for anticipating which items wiggle. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.

2026-08-14 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ARAC: エンドツーエンドのリサーチにおける Auto-Research の調整と完全性のベンチマーク

自動研究の急速な進歩により、基本的な評価の課題が表面化しました。それは、その研究の軌跡と人間の研究行動との整合性、論理的一貫性、進化の完全性をどのように測定できるのでしょうか?私たちは、Auto-Research の整合性と完全性である ARAC-Bench を提案します。これは、目的を最終的な答えの一致から人間による高品質の研究プロセスの再現に移行する、研究者を模倣した評価フレームワークです。このフレームワークは、2 つの相乗的なコンポーネントを通じて機能します。1 つは、暗黙の査読者の専門知識を、段階的に調整された定量化可能なルーブリックに変換する最初のシステムである、Academic Cognition Skills システムです。 3 段階の能力診断プロトコルは、厳格なモジュール制約の下で研究プロセスを、追跡可能で相互に独立した 3 つの側面、つまり提案、実験、合成に分解します。 11 個の SOTA フレームワークを体系的に評価した結果、最良のアラインメント スコアは 100 点中 67.9 点にすぎず、人間による厳密な方法論のシミュレーションにおいては大きなギャップがあることが明らかになりました。 Ph.D に対する検証候補者のランキングは 0.8141 という強い相関関係を示しており、ARAC-Bench が研究者が真に評価する次元を確実に反映していることが確認されています。 ARAC-Bench は、きめ細かい診断ツールだけでなく、次世代の自律研究システムをトレーニングするためのスケーラブルな報酬信号も提供します。

原文 (English)

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.

2026-08-14 13:00 JSTarXiv cs.AIビジネス/資金調達

EEG-PRIME: EEG デコード用のマルチレベル条件付けを使用したプロトタイプ整合表現学習

脳波 (EEG) デコード モデルは、取得プロトコルや個々の神経生理学におけるドメインの変化により、データセットや被験者全体での一般化が不十分なことがよくあります。我々は、クロスデータセットマルチタスクデコーディングのための2段階EEG基盤モデルであるEEG-PRIMEを提案します。 EEG-PRIME は、マスクされた事前トレーニングとプロトタイプに合わせた命令チューニングを組み合わせて、多様な BCI パラダイムにわたって命令を認識したサブジェクト不変のデコードを可能にします。事前トレーニング中、EEG エンコーダは、周波数カットオフのスペクトル拡張によるマスクされた再構成を通じて、転送可能な表現を学習します。命令のチューニング中に、EEG-PRIME にはタスクのセマンティック、データセット固有、およびサブジェクト不変の条件付けが組み込まれます。結果として得られる調整信号は、レイヤーごとのクエリ変調を通じて Q フォーマーを変調しますが、クラス ラベルの凍結されたテキスト埋め込みは、異種ラベル空間にわたるコサイン類似度ベースの予測のプロトタイプとして機能します。運動イメージ、感情認識、ADHD 検出、隠語、精神的作業負荷をカバーする 16 個のデータセットの実験では、被験者を超えた設定の下で、最先端のベースラインや以前の EEG 基礎モデルと比較して一貫した改善が見られました。追加の 2 つの保持データセット上で、EEG-PRIME は、ターゲット ドメインの最適化、キャリブレーション、または線形プローブなしでセッション内キャリブレーション モデルに匹敵するバランスの取れた精度を達成し、有望なゼロショット転送機能を実証します。

原文 (English)

EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding

Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and individual neurophysiology. We propose EEG-PRIME, a two-stage EEG foundation model for cross-dataset multi-task decoding. EEG-PRIME combines masked pretraining with prototype-aligned instruction tuning to enable instruction-aware and subject-invariant decoding across diverse BCI paradigms. During pretraining, an EEG encoder learns transferable representations through masked reconstruction with frequency-cutoff spectral augmentation. During instruction tuning, EEG-PRIME incorporates task-semantic, dataset-specific, and subject-invariant conditioning. The resulting conditioning signal modulates the Q-Former through Layer-wise Query Modulation, while frozen text embeddings of class labels serve as prototypes for cosine-similarity-based prediction across heterogeneous label spaces. Experiments on sixteen datasets covering motor imagery, emotion recognition, ADHD detection, covert speech, and mental workload show consistent improvements over state-of-the-art baselines and prior EEG foundation models under cross-subject settings. On two additional held-out datasets, EEG-PRIME achieves balanced accuracy comparable to within-session calibration models without target-domain optimization, calibration, or linear probing, demonstrating promising zero-shot transfer capability.

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

SkillShapley: LLM エージェントのスキル ステップ アトリビューションのための境界適応型 Shapley 評価

エージェント スキルは、言語エージェントがコーディングや文書処理などの長い手続きタスクを実行できるようにする重要な外部指示です。既存のエージェント スキルは主に人間の手作業による作成やエージェントの実行トレースによって作成されており、各ステップが特定のタスクにおける全体的なスキル パフォーマンスにどのように寄与するかについては十分な理解がありません。つまり、エージェント スキル内の個々のステップの貢献を定量化する際には未解決の問題が残っています。この問題に対処するために、まずスキルステップ アトリビューションを Shapley 値ベースの貢献推定問題としてモデル化し、次にエージェント スキルのステップレベル アトリビューション フレームワークである SkillShapley を提案します。特に、SkillShapley は 2 つのフェーズで動作し、重要な経験的洞察、つまりパフォーマンスの急激な崖を生み出す離散化されたベンチマーク報酬と、相乗的ではなく主に相加的なステップの相互作用によって動機付けられています。具体的には、最初に有益な連合領域を特定し、次に再利用可能な限界証拠を生成できる新しい連合を適応的にサンプリングします。広く採用されている SkillsBench のスキルに関する実験では、SkillShapley が価値の高いスキル ステップと低いスキル ステップを効果的かつ効率的に識別できることが実証され、エージェントのスキル作成に重要なポイントがいくつか提供されます。

原文 (English)

SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents

Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document processing. Existing agent skills are primarily created through human manual crafting or agent execution traces, with limited understanding of how each step contributes to overall skill performance on specific tasks; i.e., there remains an open problem in quantifying the contribution of individual steps within an agent skill. To address this issue, we first model skill-step attribution as a Shapley value-based contribution estimation problem, and then propose SkillShapley, a step-level attribution framework for agent skills. Notably, SkillShapley operates in two phases, motivated by key empirical insights, i.e., discretized benchmark rewards that create sharp performance cliffs, and step interactions that are largely additive rather than synergistic. Specifically, it first identifies informative coalitional regions, and then adaptively samples new coalitions that can yield reusable marginal evidence. Experiments on skills from the widely adopted SkillsBench demonstrate that our SkillShapley can effectively and efficiently identify high- or low-value skill steps, providing several key takeaways for agent skill creation.

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

TsuGO: Go 生死に関わる問題による LLM 推論の検索効率の調査

LLM 推論の評価は、最終的な回答の精度からプロセス レベルの評価に移行しつつありますが、既存の手法では、モデルが推論パスを計画し、推論リソースをどのように割り当てるか、つまりモデルが検索をどのように組織するかをまだ把握できていません。従来のプロセス レベルの手法は、思考連鎖 (CoT) の一貫性と冗長性に焦点を当てており、ほとんどのベンチマーク タスクは、導出やツールの使用などの静的な機能によって解決できる単一の目標を持っており、検索組織は測定されていません。 Go の死活問題を通じて LLM 推論の検索効率を評価するためのプロセスレベル推論ベンチマークである TsuGO を紹介します。これらの問題は、固有の敵対構造を備えた閉じられた検証可能な解決空間を提供し、候補の生成、応答チェック、分岐比較、および偶発的なトレース パターンではなく推論の必要な部分のバックトラックを行います。 TsuGO は、ソリューション空間を制限することで、検索組織からドメイン知識を分離し、CoT を解析して構造化された検索ツリーにし、検索効率をトークン効率やその他の診断メトリクスおよび視覚化とともにレポートします。実験によると、現在の LLM は安定した詰碁の解決にはほど遠いことがわかりました。より強力なモデルは正しい候補を早期に見つけ、生産的な分岐で努力を続けることで成功しますが、ほとんどのモデルは依然として、ニューラルガイド付き KataGo よりもガイドなしの検索アルゴリズムにはるかに近い動作をします。 CoT が長くても、トークン効率が高くても、必ずしも検索が優れているとは限りません。私たちの結果は、LLM 推論の評価に欠けている要素として、検索組織と推論リソースの割り当てを特定しました。

原文 (English)

TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems

The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search. Prior process-level methods focus on the coherence and redundancy of chain-of-thought (CoT), and most benchmark tasks have a single objective solvable by static capabilities such as derivation and tool use, leaving search organization unmeasured. We introduce TsuGO, a process-level reasoning benchmark for evaluating Search Efficiency in LLM reasoning through Go life-and-death problems. These problems provide closed and verifiable solution spaces with an inherent adversarial structure, making candidate generation, response checking, branch comparison, and backtracking necessary parts of reasoning rather than incidental trace patterns. By constraining the solution space, TsuGO disentangles domain knowledge from search organization, parses CoT into a structured search tree, and reports Search Efficiency together with Token Efficiency and other diagnostic metrics and visualizations. Experiments show that current LLMs remain far from stable tsumego solving: stronger models succeed by finding the correct candidate earlier and sustaining effort on productive branches, but most models still behave much closer to unguided search algorithms than to neural-guided KataGo. Longer CoT or higher Token Efficiency does not necessarily imply better search. Our results identify search organization and reasoning-resource allocation as missing dimensions in LLM reasoning evaluation.

2026-08-14 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

最終スコアを超えて: 長期的な AI 研究開発のためのエージェントの体系的な評価

自律エージェントは、長期的な実験を通じてモデル、システム、その他の技術成果物を改善できるようになってきています。ただし、この機能の現在の状態を理解するには、最終スコアを超えた評価が必要です。スコアでは、進歩がどこで得られるか失われるかは明らかにされず、蓄積された経験が後の決定を改善するかどうかも示されません。したがって、我々は、ルールベースのメトリクスを使用して、ソリューションのフレーミング、実行、フィードバック制御を通じて実行内の動作を特徴付ける新しいフレームワークに基づいて、36 の長期タスクに関する 7 つのフロンティア モデルの体系的な評価を提示します。また、タスク内およびタスク間でのエクスペリエンスの再利用を評価するための制御された比較を行います。その結果、現在のエージェントは完全に自律的な研究者というよりも、エンジニアリングのオプティマイザーのように動作することがわかりました。エージェントは実用的なソリューションを定式化して実装することができますが、そのパフォーマンスは実行ごとに大きく異なり、最も強力なソリューションは主に確立された技術を適応または組み合わせており、真の方法論的な新規性は依然としてまれです。詳細な分析により、観察されたパフォーマンスは、同様の最終結果の背後にある明確なプロセスのボトルネック、その後の意思決定に役立つまたは誤解を招く可能性があるエクスペリエンスの再利用、パフォーマンスの安定性に影響を与えるハーネス設計など、複数の要因によって形成されることが明らかになりました。これらの調査結果は、モデルのトレーニング、推論時間戦略、エクスペリエンス管理、ハーネス設計を改善するための具体的な方向性を示唆しています。

原文 (English)

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.

2026-08-14 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SLM とエッジ コンピューティングによる仮想エージェントの強化: 思考と記憶のプロセスの探索的評価

身体化されたインテリジェントな仮想エージェントは、複雑な仮想世界およびメタバース世界内で、永続的で適応性のあるコンテキスト認識型のエンティティとして動作することが期待されています。ただし、そのような環境に認知機能のあるエージェントを実装することは、概念的にも技術的にも困難です。さまざまな青写真と開発アプローチの中で、認知身体化エージェント アーキテクチャ (CEAA) は、知覚、記憶、推論、計画、身体化されたアクションのコンポーネントを設計するための実装指向のフレームワークとして開発されました。エッジ コンピューティングと生成 AI 言語モデルの最近の進歩を考慮して、この論文では、インタラクティブな仮想世界での仮想エージェントの認知オーケストレーションと永続性の中心となるプロセスとしての「思考」と「記憶」に焦点を当て、選択された CEAA コンポーネントのエッジベースの操作をサポートするための小型言語モデル (SLM) の使用について検討します。エッジベースの仮想エージェント ゲートウェイ システムは、さまざまなサイズの Qwen2.5 モデルを使用して NVIDIA Jetson Orin NX 上で開発および評価され、サービス リクエストを処理し、メモリ駆動型の会話を処理するシステムの機能を調査しました。一連のシミュレーション実験では、ルーティング精度、メモリ読み取りパフォーマンス、レイテンシを評価し、選択された CEAA プロセスを部分的に実装する SLM 主導のプロトタイプ エージェント システムを実証しました。このシステムは、認知「脳」が効率的かつ状況に応じて動作し、没入型の仮想世界でインタラクティブな体験を実現できる身体化エージェントの開発をサポートします。

原文 (English)

Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes

Embodied intelligent virtual agents are expected to operate as persistent, adaptive, and context-aware entities within complex virtual and Metaverse worlds. However, implementing cognitively capable agents in such environments is conceptually and technologically challenging. Among a range of blueprints and development approaches, the Cognitive Embodied Agent Architecture (CEAA) has been developed as an implementation-oriented framework for architecting components of perception, memory, reasoning, planning, and embodied action. Considering the recent advances in edge computing and generative AI language models, this paper explores the use of Small Language Models (SLMs) to support edge-based operation of selected CEAA components, focusing on "Think" and "Memory" as processes central to cognitive orchestration and persistence of virtual agents in interactive virtual worlds. An edge-based virtual agent gateway system was developed and evaluated on an NVIDIA Jetson Orin NX using Qwen2.5 models of different sizes, exploring the system's capability to process service requests and handle memory-driven conversations. A series of simulation experiments evaluated routing accuracy, memory-read performance, and latency, demonstrating an SLM-driven prototype agent system that partially implements selected CEAA processes to support the development of embodied agents whose cognitive "brain" can operate efficiently and contextually for interactive experiences in immersive virtual worlds.

2026-08-14 13:00 JSTarXiv cs.AIビジネス/資金調達

RAIL: 人工知能の準備レベルの自動分類器

人工知能テクノロジーの成熟度の評価は、投資決定、プロジェクト管理、政策監視に不可欠ですが、利用可能な準備フレームワークは異種混合であり、自動的に適用することが困難です。テクノロジー準備レベルの AI への適応には AI 固有のゲート基準が欠如し、機械学習テクノロジー準備レベルは内部プロセス成果物へのアクセスを前提とし、AI/データ準備ディメンション モデルは直接比較しにくいスケールを採用しています。この論文は 2 つの貢献を行っています。まず、これら 3 つのフレームワークを Unified AI Readiness Level (AIRL) に統合します。AIRL は、環境証拠のはしごに基づいて構築され、一般性を固定するルールと明示的な割り当て規律とともに次元の上限 (仕様、データの存在、データの品質、データの合法性、専門知識、アルゴリズムの成熟度をカバーする) によって補完された 9 レベルの順序スケールです。これにより、準備レベルが作業の自然言語記述のみから決定可能になります。第二に、スケールを運用可能にする専門家パネルによる分類器である RAIL (独立 LLM 専門家による準備評価) を提案します。1 つの証拠エージェントと 6 つの独立したディメンション エージェントで、それぞれが狭い範囲の権限を持つ大規模な言語モデルであり、決定論的な最小ルールが集約された評決を下し、主任専門家が非対称権限の下でレビュ​​ーし、パネルの推奨事項を確認または引き下げますが、上限を超えることはありません。この方法は、一貫性を示し、モノリシック LLM 分類器からの過大評価を回避することを示すいくつかの研究成果の分析でテストされました。

原文 (English)

RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct comparison. This paper makes two contributions. First, we unify these three frameworks into the Unified AI Readiness Level (AIRL), a nine-level ordinal scale built on an environmental evidence ladder and complemented by dimensional caps (covering specification, data existence, data quality, data legality, expert knowledge, and algorithmic maturity) together with a generality-anchoring rule and explicit assignment disciplines, so that a readiness level becomes decidable from a natural-language description of the work alone. Second, we propose RAIL (Readiness Assessment via Independent LLM-experts), a panel-of-experts classifier that operationalizes the scale: one evidence agent and six independent dimension agents, each a large language model with a narrowly scoped mandate, deliver verdicts that a deterministic minimum rule aggregates and a chief expert reviews under asymmetric authority, confirming or lowering the panel's recommendation but never raising it above the caps. The method was tested in the analysis of several research works showing consistency and avoiding overestimation from monolithic LLM classifiers.

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AnchorSIPS: 証拠に裏付けられた精神病リスク症状測定のための合成データセットおよび評価リソース

精神病リスク評価のための AI の進歩は、データアクセスのボトルネックによって制限されています。実際の臨床面接は、プライバシー、ガバナンス、同意の制約により共有することが困難です。我々は、トランスクリプトに基づいた測定対象者との10,000の構造化精神病リスクインタビューの合成データセットであるAnchorSIPSを紹介します。各面接は、臨床医が実施する精神病リスク面接である Mini-SIPS をモデルにしています。これは、病歴、24の症状に関する質問、患者が肯定した項目の追跡証拠、妄想様症状(異常な信念)、幻覚様症状(異常な知覚)、および混乱したコミュニケーションに関する決定、明らかな精神病レベルの症状(「率直な精神病」)の除外、および軽度または初期の精神病症状の高リスク状態である軽度精神病症候群(APS)の最終診断を記録する。 APS 診断は独立したラベルではありません。それは、以前の承認、フォローアップの詳細、症状クラスの決定、および率直な精神病チェックに依存します。すべての中間決定は、それをサポートする転写ターンに固定されています。 AnchorSIPS は、計画後実現パイプラインによって生成されます。隠されたケースシートが患者の臨床状態を特定し、決定論的プランナーがインタビュー構造を修正し、LLM が検証と限定的修復の下で患者の発話のみを認識します。生成前にラベルと構造を修正すると、マルチターン LLM ダイアログに特有のターン間の不一致が回避されます。 7 つの LLM ベースラインにわたって、モデルは大まかな決定を回復しますが、フォローアップの詳細を抽出したり、裏付けとなる転写ターンを引用したりすることができないため、最終ラベルのパフォーマンスはインタビューの能力を過大評価します。 AnchorSIPS は、証拠の抽出、転写に基づいた測定、部分開示の下での不確実性に関する研究を目的としています。

原文 (English)

AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement

Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured psychosis-risk interviews with transcript-grounded measurement targets. Each interview is modeled on Mini-SIPS, a clinician-administered psychosis-risk interview. It captures history, 24 symptom questions, follow-up evidence for items the patient affirms, decisions about delusion-like symptoms (unusual beliefs), hallucination-like symptoms (unusual perceptions), and disorganized communication, exclusion of clear psychotic-level symptoms ("frank psychosis"), and a final attenuated psychosis syndrome (APS) diagnosis, a high-risk state of milder or early psychotic symptoms. The APS diagnosis is not a standalone label. It depends on earlier endorsements, supporting follow-up details, symptom-class decisions, and the frank-psychosis check. Every intermediate decision is anchored to its supporting transcript turns. AnchorSIPS is generated by a plan-then-realize pipeline. A hidden case sheet specifies the patient's clinical state, a deterministic planner fixes the interview structure, and an LLM realizes only the patient utterances under validation and bounded repair. Fixing labels and structure before generation avoids the inter-turn inconsistencies typical of multi-turn LLM dialogue. Across seven LLM baselines, models recover coarse decisions but fail to extract follow-up details or cite supporting transcript turns, so final-label performance overstates interview competence. AnchorSIPS is intended for research on evidence extraction, transcript-grounded measurement, and uncertainty under partial disclosure.

2026-08-14 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

インタラクションの準備: 人間の役割で AI エージェントを構築および評価するためのフレームワーク

役割を担う AI エージェントを構築する製品チームとエンジニアリング チームは、評価ギャップに直面しています。エージェントは、割り当てられた役割の動作要件を満たしていながら、正確で安全かつ流暢なコンテンツを作成できます。このホワイトペーパーでは、パフォーマンスの不足している層を指定および評価するためのフレームワークとして Interaction Readiness を紹介します。このフレームワークは、エージェントが知っていることや発言していることを管理するコンテンツ仕様と、役割に基づいた交換においてエージェントがどのように行動すべきかを定義するインタラクション仕様を分離します。インタラクション仕様では、チームは、展開前に役割の目的、権限の境界、繰り返し発生する状況、境界ケース、修復動作、および監査基準を定義する必要があります。私たちは、目的の理解、権限の調整、トーンの管理、故障の修復という 4 つのエージェントの操作を通じて、インタラクションの準備を運用します。 AI 家庭教師エージェントとの学生のやり取りの公開データセットである StudyChat を使用して、コンテンツの正確さと対話の品質は独立した次元であることを示します。エージェントは、家庭教師としては失敗しても事実としては正しい場合もあれば、技術的には間違っていても対話的には健全である場合もあります。最も根強い失敗は権限の調整ミスです。エージェントは多くの場合、どのように答えるかを知っていますが、講師の役割が応答を許可するかどうか、いつ、どのように応答するかを知りません。この文書では、これらの調査結果を仕様テンプレートと監査手順に変換し、製品チームとエンジニアリング チームが導入の前後に適用できるようにしています。

原文 (English)

Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles

Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Readiness as a framework for specifying and evaluating that missing layer of performance. The framework separates content specifications, which govern what an agent knows and says, from interaction specifications, which define how an agent should conduct itself in a role-governed exchange. Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment. We operationalize interaction readiness through four agent operations: understanding purpose, calibrating authority, managing tone, and repairing breakdowns. Using StudyChat, a public dataset of student interactions with an AI tutoring agent, we show that content accuracy and interaction quality are independent dimensions: an agent may be factually correct while failing as a tutor, or interactionally sound while technically wrong. The most persistent failure is authority miscalibration: the agent often knows how to answer, but not whether, when, or how the tutor role permits it to answer. The paper translates these findings into a specification template and audit procedures that product and engineering teams can apply before and after deployment

2026-08-14 13:00 JSTarXiv cs.AIビジネス/資金調達

ピアとは何ですか?プライベート市場における評価に基づく類似性

より多くの投資家がプライベート市場を検討し、限られた透明性、希薄な開示、まれな取引と闘う中、比較対象として経済的に意味のある同業他社を特定することは、評価、デューデリジェンス、ポートフォリオ構築、およびリスク管理における基本的な課題となっています。私たちは、静的な特徴マッチングや意味論的な説明ではなく、市場評価のレンズを通して企業の類似性を定義する、アンサンブル ツリー ベースの教師付き類似性学習フレームワークを提案します。具体的には、観測された民間企業の評価に基づいて CatBoost 勾配ブースト型デシジョン ツリー モデルをトレーニングし、アンサンブル全体にわたる重要度で重み付けされたリーフ ノードの共起から評価を意識した類似性メトリックを導出します。類似性メトリクスは、プライベート市場で一般的な非線形関係、混合データ タイプ、広範な欠損データに対応しながら、共通の評価要因を捕捉します。複数の業界、地域、取引段階にまたがる観察された、または導出可能なポストマネー評価を持つ53,000社以上の企業を含む、約270,000社の世界的なプライベートマーケットユニバースを使用して、提案された類似性フレームワークが、ケースベースの説明可能性を維持しながら、評価される業界グループの下流のk最近傍評価タスクにおける従来の距離ベースおよびテキスト埋め込みベースのアプローチを改善することを実証します。

原文 (English)

What Makes a Peer? Valuation-Anchored Similarity in Private Markets

As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due diligence, portfolio construction, and risk management. We propose an ensemble tree-based supervised similarity learning framework that defines company similarity through the lens of market valuation rather than static feature matching or semantic descriptions. Specifically, we train a CatBoost gradient-boosted decision tree model on observed private company valuations and derive a valuation-aware similarity metric from importance-weighted leaf-node co-occurrences across the ensemble. The similarity metric captures shared valuation drivers while accommodating nonlinear relationships, mixed data types, and pervasive missing data common in private markets. Using a global private-market universe of approximately 270,000 companies, including more than 53,000 firms with observed or derivable post-money valuations spanning multiple industries, geographies, and deal stages, we demonstrate that the proposed similarity framework improves upon traditional distance-based and text-embedding-based approaches in downstream k-nearest-neighbor valuation tasks in the evaluated industry groups, while retaining case-based explainability.

2026-08-14 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達規制/政策

消去して保存: 最適化されたセマンティック アンカーを介して、著作権で保護されたアニメーション キャラクターの制御可能な削除

テキストから画像への拡散モデルの優れた生成機能により、著作権、特にアニメーション キャラクターの無許可複製に関する懸念が生じています。既存の概念消去方法は、アニメーション キャラクターの消去には不十分です。モデル変更方法では、多様で非常に特徴的なキャラクターに適したアンカーを特定するのが困難です。プロンプトベースのステアリング方法には、正確な介入のためのきめの細かい制御が欠けています。これらのアプローチでは、多くの場合、消去が不完全になり、画像の忠実度が低下し、実際の展開が妨げられます。この論文では、モデルの連続テキスト表現を操作して、生成中にターゲット文字を消去する制御可能な方法を提案します。構造的および詳細な制約を介してアンカーの埋め込みを最適化して文字の代理として機能させ、その後、構造を意識した適応戦略によってターゲット関連の埋め込みをアンカーに置き換えます。実験では、私たちの方法が最先端の消去効果と画像忠実度の維持を達成しながら、制御可能な消去度、マルチターゲット除去、およびモデルの転送可能性をサポートしていることが示されています。さらに、当社の最適化されたアンカーは、現在のモデル変更ベースラインとプラグアンドプレイで対応し、消去パフォーマンスを向上させます。

原文 (English)

Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors

The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized reproduction of animation characters. Existing concept erasure methods fall short for animation character erasure: model modification methods struggle to identify suitable anchors for diverse, highly distinctive characters; prompt-based steering methods lack fine-grained control for precise intervention. These approaches often yield incomplete erasure and degraded image fidelity, hindering real-world deployment. In this paper, we propose a controllable method operating on the model's continuous textual representation to erase target characters during generation. We optimizes an anchor embedding via structural and detailed constraints to serve as a character surrogate, then replaces target-related embeddings with the anchor via a structure-aware adaptive strategy. Experiments show that our method achieves state-of-the-art erasure effectiveness and image fidelity preservation, while supporting controllable erasure degree, multi-target removal, and model transferability. Moreover, our optimized anchors are plug-and-play with current model modification baselines to improve their erasure performance.

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

AQuA: 再帰的自己改善型定量取引リサーチ エージェント

私たちは定量的投資研究のレベルで再帰的自己改善を研究します。つまり、自律システムが初期の実験からの証拠を使用して、後の反復で提案される仮説や候補を改善できるかどうかを研究します。我々は、2 つの個別の言語モデル駆動型研究システムで構成される AQuA を紹介します。1 つは記号因子発見用、もう 1 つは訓練可能なモデル開発用です。 2 つのシステムは、エージェント、メモリ、候補空間、研究状態を共有しません。代わりに、それぞれが独立して、検証された証拠を保持し、それを次の提案の指針として使用することで、独自の研究ループを閉じます。この限定された意味で、両方のシステムは研究プロセスのレベルで再帰的な自己改善を実装します。各システムは、独自の密閉されたサンドボックスも使用します。これにより、データ分割、特徴とラベルの定義、およびエバリュエーターが修正され、制約された因子式または構成差分を通じてのみモデルが動作できるようになります。マネージャーが仲介するマルチエージェント パイプラインであるファクター システムは、ファクターを検出してシグナルに結合し、暗号通貨ユニバースでの結合情報係数が約 $0.190 に達します。ハイブリッド時系列アーキテクチャ上の構成主導型ループであるモデル システムは、米国株の銘柄ごとの情報係数 $+0.0843$ に達し、それを 2 レッグ コストで最大 $+2.50$ のホールドアウト シャープを持つ閾値ロング/ショート戦略に変換します。この戦略は 2021 年から 2025 年まで毎年前向きです。

原文 (English)

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises two separate language-model-driven research systems: one for symbolic factor discovery and one for trainable model development. The two systems do not share agents, memories, candidate spaces, or research state. Instead, each independently closes its own research loop by retaining validated evidence and using it to guide subsequent proposals. In this bounded sense, both systems implement recursive self-improvement at the level of the research process. Each system also uses its own sealed sandbox, which fixes the data splits, feature and label definitions, and evaluator while allowing the model to act only through constrained factor expressions or configuration diffs. The factor system, a manager-mediated multi-agent pipeline, discovers and combines factors into a signal that reaches a combined information coefficient of about $0.190$ on a crypto universe. The model system, a config-driven loop over a hybrid time-series architecture, reaches a per-stock information coefficient of $+0.0843$ on US equities and converts it into a threshold long/short strategy with a held-out Sharpe of up to $+2.50$ at a two-leg cost. The strategy is positive in every year from 2021 to 2025.

2026-08-14 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

ラベルはエンドポイントではない: 治療の漏洩と MCP エージェントのセキュリティ評価における妥当性の構築

ツールを使用するエージェントのセキュリティ評価では、保存されているラベルと行動の事実が同一視されることがよくあります。 10,200 の実行行を追跡して、180 のモデルバインドされたリクエスト、45 のセマンティック リクエスト、および 15 の観察可能な刺激を追跡することにより、保存されたキャンペーンを監査します。 2 つのスキーマ処理が提供されましたが、計画されていた外部ペイロード ファミリ コーパスは提供されませんでした。履歴グレーダーは直接的な治療漏れを示しました。治療メタデータが ATTACK_SUCCESS クラスをゲートしていたので、修正された動作により治療の再ラベル付けによりクラスが変更される可能性がありました。治療ブラインド再構築では、検証された 3 件の保護データ転送と 1 件の別個の不正転送ケースを保持しながら、58 件の過去の ATTACK_SUCCESS または HIJACK_ATTEMPT ラベルを承認された良性の完了に修正します。ロックされた v2 国勢調査には ATTACK_SUCCESS レコードがまったく含まれていませんが、転送ケースは目的の完了に関するセマンティック境界で HIJACK_ATTEMPT のままです。ロックされた v2 によって構造的に解釈可能であるとみなされた 96 件のリクエストすべてに対する二重レビュアーによるブラインド コンコーダンス レビューでは、同一のレビュアー コンセンサス クラスが生成されましたが、4 つの構成境界ケースについてはロックされたコードブックとは異なりました。私たちは、7 リンクの整合性チェーンと、実行可能なスコープ限定のエンドポイント整合性リンターを提供します。結果はキャンペーンに限定された測定監査であり、人口の攻撃率、モデルのランキング、防御効果、または因果関係の推定ではありません。

原文 (English)

Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation

Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200 execution rows to 180 model-bound requests, 45 semantic requests, and 15 observable stimuli. Two schema treatments were delivered, but the planned external payload-family corpus was not. The historical grader exhibited direct treatment leakage: treatment metadata gated the ATTACK_SUCCESS class, so fixed behavior could change class under treatment relabeling. A treatment-blind reconstruction corrects 58 historical ATTACK_SUCCESS or HIJACK_ATTEMPT labels to authorized benign completions while preserving three verified protected-data transfers and one separate unauthorized-forwarding case. The locked v2 census contains exactly zero ATTACK_SUCCESS records, while the forwarding case remains a HIJACK_ATTEMPT at a semantic boundary concerning objective completion. A dual-reviewer blinded concordance review of all 96 requests deemed structurally interpretable by locked v2 produced identical reviewer-consensus classes but differed from the locked codebook on four construct-boundary cases. We contribute a seven-link Integrity Chain and an executable, scope-bounded endpoint-integrity linter. The result is a campaign-bounded measurement audit, not a population attack-rate, model-ranking, defense-efficacy, or causal estimate.

2026-08-14 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

H-VAEP と H-xT: 確率の推定によるハンドボールの攻撃的なオンザボールアクションの評価

プロのハンドボールにおける従来の選手評価は、基本的なボックススコア指標やヒューリスティック指標に依存しており、複数選手によるビルドアップの連鎖を評価することができません。フットボール (サッカー) 分析では、予想される脅威 (xT) と確率推定によるアクションの評価 (VAEP) が採用されていますが、これらのイベントベースのアクション評価フレームワークはまだハンドボールには適応されていません。この論文では、ハンドボール ブンデスリーガの 5 シーズンの追跡由来のイベント データを利用して、ハンドボールに対する xT と VAEP の最初の包括的な適応と評価を紹介します。私たちは、ハンドボール固有のコート ゾーニング レイアウトを使用して Handball-xT (H-xT) を開発し、標準的な長方形のグリッドよりも体系的に堅牢であることをシミュレーションによって実証しています。特徴空間を調整し、コンテキストの長さを選択してチーム ID の漏洩を制限することで、ハンドボール VAEP (H-VAEP) を最適化します。私たちの評価では、H-VAEP が、ビルドアップ プレーを際立たせる、非常に安定しており、識別力があり、直感的なプレーヤー評価をもたらすことが示されています。最後に、プロのクラブがこれらのモデルを導入できるように、完全なコード リポジトリをリリースします。

原文 (English)

H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities

Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP), these event-based action valuation frameworks have not yet been adapted to handball. In this paper, we present the first comprehensive adaptation and evaluation of xT and VAEP for handball, utilizing five seasons of tracking-derived event data from the Handball Bundesliga. We develop Handball-xT (H-xT) using a handball-native court zoning layout, demonstrating via simulations that it is systematically more robust than standard rectangular grids. We optimize Handball-VAEP (H-VAEP) by tailoring its feature space and selecting the context length to limit team-identity leakage. Our evaluation shows that H-VAEP yields exceptionally stable, discriminative, and intuitive player ratings that highlight build-up play. Finally, we release our complete code repository to help professional clubs deploy these models.

2026-08-14 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

UniTraffic-Agent: 2 つのドメイン外評価による AI City Challenge 2026 Track 3 の統合トラフィック ビデオ推論

道路ビデオは事故、違反、車両と交通弱者の道路利用者とのやりとりの直接的な証拠を提供するため、交通ビデオの理解はインテリジェント交通機関における重要な問題となっています。有用なシステムでは、交通イベントがどのように発生するか、なぜそれが起こるのか、関連するインタラクションがいつ発生するのかを説明する必要がありますが、交通ビデオにはまばらなイベントとさまざまな視点が含まれるため、マルチモーダル大規模言語モデル (MLLM) ではこれが依然として困難です。第 10 回 AI シティ チャレンジの Track~3 の MR-CAS ソリューションである UniTraffic-Agent を紹介します。これには、交通異常推論 (TAR) と 2 つのドメイン外評価 (魚眼交通イベントの FETV と歩行者の意図推論の PSI-VQA) が含まれています。 UniTraffic-Agent は、タイムスタンプ付きの視覚的証拠をサンプリングし、1 つのリクエスト内の同じクリップからのすべての質問について推論し、タスク固有のアクション アダプターを通じて応答を変換する、観察 - 理由 - 行為 - 検証のワークフローに従います。公式のパブリック リーダーボードでは、MR-CAS は TAR でスコア 0.5780 で 16 位、FETV で 0.4884 で 2 位、PSI-VQA で 64.4161 で 4 位にランクされています。コードは https://github.com/Roclp/UniTraffic-Agent で入手できます。

原文 (English)

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse events and varied viewpoints. We introduce UniTraffic-Agent, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. UniTraffic-Agent follows an observe--reason--act--verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters. On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161. The code is available at https://github.com/Roclp/UniTraffic-Agent.

2026-08-14 13:00 JSTarXiv cs.AIビジネス/資金調達

LOB-ID: 開始距離による合成市場データの評価

指値注文帳 (LOB) データの生成モデルは急速に進歩していますが、その評価は多くの場合、定型化された事実と選択された市場統計に焦点を当てています。これらの測定は有用な診断を提供しますが、オーダーブックの軌跡の同時時間構造およびクロスレベル構造を捕捉できない可能性があります。 LOB-ID は、Fr\'echet Inception Distance (FID) と Monge Inception Distance (MIND) を LOB データに適応させる埋め込みベースのフレームワークです。ドメイン固有のエンベディングを取得するために、5 つの株式の 4 か月分のレベル 2 オーダーブック データで DeepLOB アーキテクチャをトレーニングします。 LOB-ID は、時間、機器、埋め込みチェックポイント全体にわたって安定しており、制御された歪みの下では単調に増加することを示します。次に、統計ベースの評価を回避する FID およびディープブック摂動に対するモーメント マッチング攻撃を構築します。 MIND は両方の歪みに対して大幅に敏感なままです。最後に、確率的ベースラインと深層学習アプローチにまたがる 5 つの生成 LOB モデルをスコアリングし、LOB-ID が、それぞれが構築によって捕捉する結合時間構造およびクロスレベル構造に沿ってそれらをランク付けすることを発見しました。

原文 (English)

LOB-ID: Evaluating Synthetic Market Data by Inception Distances

Generative models of limit orderbook (LOB) data have advanced rapidly, but their evaluation often focuses on stylised facts and selected market statistics. These measures provide useful diagnostics but may not capture the joint temporal and cross-level structure of order-book trajectories. We introduce LOB-ID, an embedding-based framework that adapts the Fr\'echet Inception Distance (FID) and Monge Inception Distance (MIND) to LOB data. To obtain domain-specific embeddings, we train the DeepLOB architecture on four months of Level-2 order-book data for five equities. We show that LOB-ID is stable across time, instruments, and embedding checkpoints, and rises monotonically under controlled distortions. We then construct a moment-matching attack against FID and a deep-book perturbation that evades statistic-based evaluation. MIND remains substantially more sensitive to both distortions. Finally, we score five generative LOB models, spanning stochastic baselines and deep learning approaches, and find that LOB-ID ranks them in line with the joint temporal and cross-level structure each captures by construction.

2026-08-14 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport

Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introd…

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達研究/論文

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what t…

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is diffi…

2026-08-14 13:00 JSTarXiv cs.AI画像/動画生成ロボティクスビジネス/資金調達研究/論文

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in…

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG

By retrieving contexts from knowledge graphs, graph-based retrieval-augmented generation (GraphRAG) enhances large language models (LLMs) t…

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI

We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions…

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and…

2026-08-14 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection

Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expens…

2026-08-14 05:14 JSTTechCrunch AIビジネス/資金調達

Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation.

AI is expensive, Ali Ghodsi tells TechCrunch. With so many investors wanting into his latest round, he said yes to more than planned.

2026-08-13 13:00 JSTarXiv cs.AIビジネス/資金調達

クエリ カバレッジとクレーム検証可能性によるクエリに依存しない RAG 評価に向けて

検索拡張生成は、応答を検索された証拠に基づいて行うことで、大規模な言語モデルの事実性を向上させますが、既存の評価フレームワークは、クローズエンドの事実探索からオープンエンドの説明要求に至るまで、ユーザーの多様な範囲にわたって一貫したきめの細かい診断を提供するのに苦労しています。私たちは、クエリに依存せず完全に参照フリーのフレームワークである Q-CARE を提案します。これは、クエリをサブクエリに分解し、回答をアトミック クレームに分解することで、きめ細かい評価を可能にします。 Q-CARE は、クエリ カバレッジとクレームの検証可能性に基づいた統一評価原則を確立し、カバレッジを意識した取得メトリクス (C-Prec@k、C-nDCG@k) とクレーム レベルのジェネレータ メトリクス (完全性、簡潔性、および検証可能性) を生成します。 8 つのデータセットにわたる人間による注釈付きベンチマークで、Q-CARE は、RAGEval や RAGChecker を含む 4 つの既存の RAG 評価指標よりも人間の判断との高い相関関係を達成し、信頼性の高い自動評価フレームワークとしての有効性を証明しています。コードとデータは https://github.com/DISL-Lab/Q-CaRE-COLM-26 で公開されています。

原文 (English)

Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of user queries, ranging from close-ended fact-seeking to open-ended explanatory requests. We propose Q-CARE, a query-agnostic and fully reference-free framework that enables fine-grained assessment by decomposing queries into sub-queries and answers into atomic claims. Q-CARE establishes a unified evaluation principle based on query coverage and claim verifiability, yielding coverage-aware retriever metrics (C-Prec@k, C-nDCG@k) and claim-level generator metrics (Completeness, Conciseness, and Verifiableness). On a human-annotated benchmark spanning eight datasets, Q-CARE achieves higher correlation with human judgments than four existing RAG evaluation metrics, including RAGEval and RAGChecker, proving its effectiveness as a reliable, automated evaluation framework. Code and data are publicly available at https://github.com/DISL-Lab/Q-CaRE-COLM-26.

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

AgonAlpha: 迅速な経済性とスケーラブルなエージェント検索による自律的なアルファ検出

言語モデルは多くのもっともらしい取引要素を提案できますが、自律的な研究システムは評価予算を割り当て、独自の証拠を検証し、各候補がどのように生成されたかを保存する必要もあります。私たちは、数式だけではなく、凍結された研究成果物 (仮説、実行可能な式、プラットフォームの証拠、理論的根拠、レビュー ステータス) を検索するアーキテクチャである AgonAlpha を紹介します。私たちの知る限り、AgonAlpha は、検証済みアーティファクト検索、再実行と拒否権を備えた新鮮なコンテキストの敵対的レビューア、保留を認識した並行予算割り当てを完全な公開証拠証跡と組み合わせた最初のアルファマイニング システムです。 WorldQuant BRAIN での独立したデプロイメントにより、5 人のユーザーと 6 つのモデル バックエンドにわたって SPECTACULAR グレードのアルファが生成され、フィットネスは 9.50、シャープは 3.48 に達しましたが、すべての投稿の即時表現の出自は維持されました。

原文 (English)

AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search

Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced. We present AgonAlpha, an architecture that searches over frozen research artifacts---hypotheses, executable expressions, platform evidence, rationales, and review status---rather than formulas alone. To our knowledge, AgonAlpha is the first alpha-mining system to combine verified artifact search, a fresh-context adversarial reviewer with re-execution and veto authority, and pending-aware parallel budget allocation, together with a complete public evidence trail. Independent deployments on WorldQuant BRAIN produced SPECTACULAR-grade alphas across five users and six model backends, with Fitness reaching 9.50 and Sharpe reaching 3.48, while retaining prompt-to-expression provenance for every submission.

2026-08-13 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (The…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成エージェントビジネス/資金調達研究/論文

モバイル エージェント評価のための LLM 審査員のベンチマーク

モバイル エージェントのベンチマークは、タスクの完了を評価するために LLM ベースのジャッジにますます依存していますが、モバイル エージェントの軌跡に関するこれらのジャッジの信頼性はほとんど調査されていないままです。モバイル エージェントの軌跡に関する LLM-as-judge メソッドを体系的に評価するためのベンチマークである MobileJudgeBench を紹介します。私たちのベンチマークは、6 つのモバイル エージェント ベンチマーク、4 つのエージェント モデル、および 68 のアプリにわたる、人間が注釈を付けた 931 の軌跡で構成されています。複数の LLM バックエンドにわたって 6 つの判定メソッド (SPA-Bench、AndroidArena、AgentRewardBench の 2 つのモードを備えた A3 から適応した 5 つと、私たちが設計した単純なベースライン) を評価します。私たちの実験により、3 つの重要な発見が明らかになりました。まず、サンプルのスクリーンショットを使用した単純なベースライン審査員は、専用の手法と競合し、多くの場合それを上回っています。これは、より複雑な審査員パイプラインが審査員の質を一貫して向上させるわけではないことを示しています。競合する手法の中でも、LLM バックボーンが主な推進力です。第 2 に、ベンチマーク品質メトリクスは、現実世界のジャッジの有用性を確実に予測します。ベンチマーク品質メトリクスは、評価におけるエージェントのランキング忠実度と、ジャッジがポリシーに基づく強化学習の報酬シグナルとして機能する場合の下流のパフォーマンスの両方に相関します。 3 番目に、2 つの LLM バックエンドにわたる障害分析により、バックボーンの適合率と再現率の特性に関連する、一方は保守的で他方は許容的である、定性的に反対の障害プロファイルが明らかになります。

原文 (English)

Benchmarking LLM Judges for Mobile Agent Evaluation

Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for systematically evaluating LLM-as-judge methods on mobile agent trajectories. Our benchmark comprises 931 human-annotated trajectories spanning 6 mobile agent benchmarks, 4 agent models, and 68 apps. We evaluate 6 judge methods (five adapted from SPA-Bench, A3 with two modes, AndroidArena, and AgentRewardBench, plus a simple baseline we design) across multiple LLM backends. Our experiments reveal three key findings. First, a simple baseline judge with sampled screenshots is competitive with, and often exceeds, purpose-built methods, indicating that more elaborate judge pipelines do not consistently improve judge quality; among competitive methods, the LLM backbone is the primary driver. Second, benchmark quality metrics reliably predict real-world judge utility: they correlate with both agent ranking fidelity for evaluation and downstream performance when judges serve as reward signals for on-policy reinforcement learning. Third, failure analysis across two LLM backends uncovers qualitatively opposite failure profiles, one conservative and the other permissive, linked to the backbone's precision-recall characteristics.

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

プロンプトから行動の調整まで: 個別の LLM ジャッジによる推奨評価

従来のオフライン レコメンデーション評価は、拡張が困難な手動で維持される複雑な機能パイプラインに大きく依存しています。大規模言語モデル (LLM) は、生のテキスト ログからユーザー エンゲージメントを直接予測することで有望な代替手段を提供しますが、この研究の実証分析では、双方向合理化と呼ばれる重大な失敗モードが特定されています。ゼロショット設定では、LLM はまったく同じ項目について、同一の証拠を用いてユーザー エンゲージメントの肯定的結果と否定的結果の両方を説得力を持って主張することがわかり、ユーザー エンゲージメントの予測における既製 LLM の信頼性の低さを浮き彫りにしています。これを解決するために、正しい理論的根拠と反事実的な理論的根拠を組み合わせて、選好の最適化と微調整を組み合わせた逐次行動調整フレームワークを開発し、適用します。実際のホームページのインタラクション ログで評価すると、この調整された推論アプローチは、ゼロショット ベースラインを上回る Macro-F1 スコアの 32.19\% 上昇を達成し、実稼働環境で設計されたベースラインと一致します。この結果は、動作の調整により双方向の合理化が軽減され、手動のパイプライン オーバーヘッドなしで人間が解釈可能な推論トレースを提供できることが実証されました。

原文 (English)

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact same item with identical evidence, highlighting the unreliability of off-the-shelf LLMs in predicting user engagement. To resolve this, we develop and apply a sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales. Evaluated on real-world homepage interaction logs, this aligned reasoning approach achieves a 32.19\% lift in Macro-F1 score over the zero-shot baseline and matches the production feature-engineered baseline. The results demonstrate that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pipeline overhead.

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

FrontierFinance: 金融代理店のフロンティア インテリジェンスを測定するための挑戦的なベンチマーク

AI エージェントは専門的な投資調査に導入されることが増えていますが、投資家のワークフロー全体の複雑さを捉えたベンチマークはありません。既存のベンチマークは主に財務データの抽出を対象としていますが、これは現在のモデルがほぼ飽和状態となっている狭い範囲であり、一方、参照ベースの指標や一般的な LLM-as-a-judge スコアリングでは、実際のアナリストの質問が求める自由回答型の長い形式の回答には至っていません。 FrontierFinance は、投資家のワークフロー全体にわたる 6 つの重要なユースケースにまたがる、専門家が作成した 220 のクエリと 11,543 のソース属性のルーブリックからなる完全にオープンなベンチマークです。 FrontierFinance は、既存の公的金融ベンチマークよりも広範囲かつ困難です。公開データに限定された共通ハーネスの下でフロンティア モデルとエージェント システムを評価すると、モデル単体ではなくツール ハーネスが品質と効率を大きく左右していることがわかります。 Samaya の社内システムは 56.0% でリードしており、約 2.2 倍のコストで最も強力なフロンティア モデル (Claude Fable 5、49.2%) を上回っています。そして、最高のオープンウェイト モデル (Kimi K3、46.4%) は、4.5 倍のコストで最高の独自モデルとほぼ同等であることがわかりました。スクリーニングと検出とセクター、産業、マクロは依然としてすべてのシステム全体で最も困難なユースケースであり、最高のシステムでも 33% と 39% にすぎません。データセットとグレーディング コードは公開されています。

原文 (English)

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand. We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow. FrontierFinance is both broader and harder than existing public finance benchmarks. Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost; and that the best open-weight model (Kimi K3, 46.4%) nearly matches the best proprietary model at 4.5x lower cost. Screening & Discovery and Sector, Industry & Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%. We make the dataset and grading code publicly available.

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

グラフ構造のルーブリック: LLM 審査員向けにルーブリックを入力済みの評価グラフに編集する

ルーブリックベースの評価者は通常、ルーブリックをプロンプトコンテキストまたはフラットな基準として扱います。ルーブリックは何を判断するかを指定しますが、自然言語ルールにそれが記載されている場合でも、基準の構成を暗黙的なままにします。グラフ構造ルーブリック (GSR) を導入します。これは、応答を観察する前にルーブリックを応答に依存しない型付き評価グラフにコンパイルします。基準ノードは判断を導き出します。変換、リダクション、およびゲート演算子は、名前付きポートを通じてそれらを構成します。そして、Readout と呼ばれるタスク固有の出力マッピングは、固有のシンクをスコアまたは好みに変換します。コンパイルでは、不正な形式のグラフや型の互換性のないグラフは拒否されます。ポイントごとの評価では、グラフを集計する前にルーブリックの次元を個別に判断します。ペアワイズ評価では、あらゆる基準の下で各候補に対して 1 つの判定を行ってグラフを再利用します。 GPT-OSS-120B の下では、GSR は 4 つの点ごとのデータセットでの Prometheus スタイルのスコアリングよりも正確なスコアの一致を 0.62 ~ 6.75 パーセント ポイント改善し、ネイティブ タイおよび棄権ポリシーに基づく 2 つの優先順位ベンチマークで数値的に最高のエンドツーエンドのペアごとの精度を達成します。

原文 (English)

Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges

Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to judge but leave criterion composition implicit, even when natural-language rules state it. We introduce Graph-Structured Rubrics (GSR), which compiles a rubric into a response-independent typed evaluation graph before observing responses. Criterion nodes elicit judgments; transformation, reduction, and gating operators compose them through named ports; and a task-specific output mapping, termed Readout, converts the unique sink into a score or preference. Compilation rejects malformed or type-incompatible graphs. Pointwise evaluation judges rubric dimensions separately before graph aggregation; pairwise evaluation reuses the graph with one judgment for each candidate under every criterion. Under GPT-OSS-120B, GSR improves exact score agreement by 0.62--6.75 percentage points over Prometheus-style scoring on four pointwise datasets and achieves the numerically highest end-to-end pairwise accuracy on two preference benchmarks under native tie and abstention policies.

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

誰が一番良いと思うかは、どのくらいの期間放置するかによって決まります: LLM 評価における予算に応じたランキング

大規模な言語モデルの標準評価では、推論条件全体でモデルのランキングが安定していることが前提となります。私たちは、トークン生成予算、つまりモデルが生成できる最大トークンを 7 つのレベル (64 ~ 4,096) にわたって変化させ、3 つの推論ベンチマーク (56,476 推論) で 4 つのモデルを評価することで、この仮定に異議を唱えます。我々は 4 つの発見を報告します: (i) 切り捨てを制御した後でも、項目の 3 ~ 19% が非単調な動作 (予算が増えると精度が低下する) を示し、この現象はモデル固有です (モデル間の重複: 6 ~ 14%)。 (ii) モデルのランキングは、すべてのベンチマークの予算全体で逆転します ($p {<} 0.01$、McNemar)。 (iii) Oracle の分析により、$+27.8$pp までのモデルの相補性が明らかになり、これは予算が限られている場合に最も顕著になります。 (iv) 予算を意識したルーターは、オラクル ギャップ クロスドメインの 14.1% をキャプチャします。予算機能はドメイン内では役立ちます ($+1.6$ ~ $+5.7$pp) が、ドメイン固有であり転送に悪影響を及ぼします ($-1.2$pp)。これらの結果は、予算に応じた評価プロトコルの正当性を主張します。

原文 (English)

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), evaluating four models on three reasoning benchmarks (56,476 inferences). We report four findings: (i) 3--19% of items exhibit non-monotone behavior (accuracy decreasing with more budget), even after controlling for truncation, and this phenomenon is model-specific (cross-model overlap: 6--14%). (ii) Model rankings reverse across budgets on all benchmarks ($p {<} 0.01$, McNemar). (iii) Oracle analysis reveals model complementarity up to $+27.8$pp, most pronounced at constrained budgets. (iv) A budget-aware router captures 14.1% of the oracle gap cross-domain; budget features help within-domain ($+1.6$ to $+5.7$pp) but are domain-specific and hurt transfer ($-1.2$pp). These results argue for budget-conditioned evaluation protocols.

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

TRACE ベンチ: タスク主導のロールプレイ エージェント チェックリストの評価

ロールプレイの評価は、単一のスコアを割り当てるだけではありません。どの役割要件がテストされ、どれが失敗したか、そしてどの対話証拠が判断を裏付けるかを明らかにする必要があります。私たちは、タスク駆動型エージェントチェックリスト評価フレームワークである TRACE Bench を提案します。各ロール プロファイルをオフラインで固定チェックリストに分解し、ユーザー エージェントを使用してターゲット ロールプレイ モデルと自然に会話しながら、モデルの応答からチェックリストの状態を非公開で更新します。したがって、スコアはブラックボックスの全体的な印象ではなく、チェックリストの項目とそれをサポートする対話のターンに遡ります。カバレッジの相互検証のために、MiniMax ロールプレイ ベンチマークからリリースされた M2 フリー ダイアログ トランスクリプトを、同じロール由来のチェックリストと照合して監査します。公開されたフリー チャットのトランスクリプトは、主要な役割プロファイル ポイントの 73.74% しかカバーしていませんが、TRACE Bench はより少ないターンで 99.91% のカバー率に達します。堅牢性の実験では、実行を繰り返したり、ユーザー エージェントを置き換えたりしてもランキングが安定していることが示されています。 TRACE Bench は、26 モデルにわたる全体的なランキングを、機能の内訳とチェックリスト トレースとともにレポートします。また、クローズドループ ベンチマーク エボリューションもサポートしており、失敗したトレースで効果的であることが証明された検証手法を抽出して、後の評価で観察された故障モードをより確実に抽出して調べることができます。

原文 (English)

TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation

Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating checklist states from model responses. Scores therefore trace back to checklist items and supporting dialogue turns rather than a black-box holistic impression. For coverage cross-validation, we audit released M2 free-dialogue transcripts from the MiniMax Role-play Benchmark against the same role-derived checklist. The released free-chat transcripts cover only 73.74% of key role-profile points, whereas TRACE Bench reaches 99.91% coverage in fewer turns. Robustness experiments show stable rankings under repeated runs and User Agent replacement. Across 26 models, TRACE Bench reports overall rankings together with capability breakdowns and checklist traces. It also supports Closed-Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so later evaluations can more reliably elicit and examine observed failure modes.

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

幼稚園から高校までの教育における AI 個別指導の質を向上させる方法論

現在、多くの AI 家庭教師が大規模言語モデル (LLM) を活用しています。 LLM が不透明なブラックボックスであることを考えると、あらゆる変更の影響を測定するための堅牢な評価とライブ実験が不可欠です。当社は Khanmigo (カーン アカデミー、2023 年) を立ち上げ、幼稚園から高校までを対象とした AI を活用した個別指導の先駆者となりました。 AI 個別指導の品質と生徒の関与を測定するために使用する指標と、私たちが実行したさまざまな実験について説明します。モデル、プロンプト、パーソナライゼーション、エージェントなどの指標を変更した変更点を強調します。

原文 (English)

Methodologies for Improving the Quality of AI Tutoring in K-12 Education

Many AI tutors leverage large language models (LLMs) today. Given that LLMs are opaque black boxes, robust evaluation and live experimentation to measure the impact of every change are essential. We pioneered AI-powered tutoring for K-12 with the launch of Khanmigo (Khan Academy, 2023). We describe the metrics we use to measure AI tutoring quality and student engagement as well as various experiments we have run. We highlight the changes that have moved our metrics, including models, prompting, personalization and agents.

2026-08-13 13:00 JSTarXiv cs.AIエージェントロボティクスビジネス/資金調達

RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation

Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluation across complex road netwo…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound…

2026-08-13 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework

Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual comp…

2026-08-13 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models

As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluati…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit

Community-conditioned language model adaptation needs choices about data collection, community definition, and evaluation that are currentl…

2026-08-13 13:00 JSTarXiv cs.AIビジネス/資金調達

HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment

Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing…

2026-08-13 13:00 JSTarXiv cs.AIビジネス/資金調達

Ethics Practices in AI Development: An Empirical Study Across Roles and Regions

Recent advances in AI applications have raised growing concerns about the need for ethical guidelines and regulations to mitigate the risks…

2026-08-13 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Designing Agentic AI-Based Screening for Portfolio Investment

We introduce a new agentic artificial intelligence (AI) platform for portfolio management. Our architecture consists of three layers. First…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Evaluation and Hardening of LLM System Instructions Against Extraction via Encoding Attacks

System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensit…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents

Cloud-based Large Language Models (LLMs) can perform autonomous penetration-testing sub-tasks such as Linux privilege escalation, but raise…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows…

2026-08-13 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol

AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization

Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs…

2026-08-13 03:19 JSTTechCrunch AIビジネス/資金調達

AI coding startup Cognition reportedly already in talks to raise at $40B valuation

Cognition may be looking to raise another mega round just a few months after raising $1 billion at a $26 billion valuation.

2026-08-13 02:41 JSTTechCrunch AILLM/生成AIビジネス/資金調達

OpenAI-backed Thrive Holdings raises $2B to bring AI to the enterprise

Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Alti…

2026-08-13 01:04 JSTTechCrunch AIビジネス/資金調達

Lovable confirms new $13.3B valuation, raises another $400M

This new funding comes after Lovable hit $500 million in annualized run rate revenue in June, the startup told TechCrunch.

2026-08-13 00:44 JSTTechCrunch AIビジネス/資金調達

How a $250 million acquisition collapsed into allegations of fraud and forged signatures

Investors are still waiting for their share of the $250 million windfall, and VideoVerse co-founder Vinayak Shrivastav is now at the center…

2026-08-12 20:00 JSTTechCrunch AIビジネス/資金調達

AI code-testing startup Blacksmith’s valuation jumps almost 10x in less than a year

Blacksmith says revenue has grown more than tenfold over the past year.

2026-08-12 15:09 JSTITmedia AI+エージェントビジネス/資金調達規制/政策

中国発AIエージェント「Manus」、Metaから独立へ 中国政府が買収に反発、一部ユーザーデータは削除に

AIエージェント「Manus」を提供するManusは8月11日(現地時間)、独立企業としての運営をまもなく再開すると発表した。米Metaからの分離に伴い、一部ユーザーのデータを23日から削除するとして、事前のバックアップを呼び掛けている。

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

評価条件付きトレーニング: より強力な監視体制に一般化するためのモデルの指導

大規模言語モデル (LLM) のトレーニングに使用されるフィードバック シグナルは、LLM の動作の主な推進力であり、人間の価値観と目標との整合性を浸透させるための主な手段です。ただし、現在のトレーニング後の方法の主な制限は、ヒューマン アノテーターや自動報酬関数が、与えたいフィードバックを忠実にキャプチャできないことです。評価条件付きトレーニング (ECT) を導入します。これは、自然言語を使用して、提供されるフィードバックの忠実度に基づいて各トレーニング サンプルを条件付けし、導入時に高忠実度のモニターで LLM を条件付けすることで望ましい動作を引き出すトレーニング後のフレームワークです。 ECT は、不完全なフィードバックの下でパフォーマンスを向上させることを目的としており、SFT や PPO などの既存のアルゴリズムへのアドオンとして機能します。我々はまず ECT の概念的な枠組みを提供し、報酬の仕様ミスの永続的な原因に対処するその可能性について議論します。次に、潜在知識の引き出し (ELK) 問題の文脈で ECT を動機付けます。最後に、2 つの概念実証実験で ECT を評価します。ニュース記事生成における均等性の向上と、算術タスクにおけるお調子者の軽減です。それぞれの設定で、不完全なフィードバック、報酬バイアス、ユーザーとの同意をそれぞれ利用します。どちらの設定でも、ECT は直接トレーニングに比べて目標とする行動を改善します。

原文 (English)

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. However, a key limitation of current post-training methods is the inability of human annotators and automated reward functions to faithfully capture the feedback we would like to give. We introduce Evaluation-Conditioned Training (ECT), a post-training framework that uses natural language to condition each training sample on the fidelity of the feedback we provide and then elicits the desired behavior by conditioning the LLM on a high-fidelity monitor in deployment. ECT is aimed at improving performance under imperfect feedback and works as an add-on to existing algorithms such as SFT and PPO. We first provide a conceptual framework for ECT and discuss its potential to address persistent sources of reward mis-specification. Then we motivate ECT in the context of the eliciting latent knowledge (ELK) problem. Finally, we evaluate ECT on two proof-of-concept experiments: increasing even-handedness in news article generation and reducing sycophancy on an arithmetic task. In each setting, we utilize imperfect feedback, rewarding bias and agreement with the user, respectively. In both settings, ECT improves the targeted behavior relative to direct training.

2026-08-12 13:00 JSTarXiv cs.AIビジネス/資金調達

Neuroevolution Arena: Nested Ecological Evaluation of Update-and-Inheritance Regimes across Neural Architectures

Competitive artificial-life systems can rank trained controllers differently under training and ecological evaluation. We present Neuroevol…

2026-08-12 13:00 JSTarXiv cs.AIビジネス/資金調達

HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment

Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing…

2026-08-12 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SkillZip: 再利用可能な構造の発見による自己進化エージェントの評価不要のスキル圧縮

自己進化するエージェントは、成功した手順や失敗の修正を追加することで、再利用可能なスキルを蓄積します。時間が経つと、同じ要件が複数の分岐、例、警告で再度説明されることが多くなり、共通のアクション シーケンスは再利用されずにコピーされます。結果として得られるスキルは注入にコストがかかり、維持するのが困難になります。スキルは平坦なパッセージではないため、一般的なプロンプト圧縮はこの設定には適していません。スキルの名前と説明はいつ適用されるかを定義し、ワークフローは実行を制御し、ツールと出力コントラクトは有効性を制約し、まれな例外は、それらをアクティブ化するサンプル タスクがない場合でも必須のままである可​​能性があります。評価ガイド付き圧縮ではこれらの動作をテストできますが、ロールアウト、コスト、および圧縮時間の評価セットへの依存が生じます。私たちは、最短で忠実な構造的説明を見つけることによってスキルを圧縮する、評価不要のメソッドである SkillZip を紹介します。直感は一度説明し、多くを参照します。繰り返しのルールを適用範囲で一度だけ記述し、繰り返しのアクション シーケンスを共有プロシージャに要素化し、相違点のみを明示的な例外として保持します。この直感を、抽出されたすべてのトリガー、ワークフロー エッジ、ツール要件、義務、および出力フィールドに対するハード カバレッジ制約の対象となる、スキル契約および残余に関する型付きの最小記述長目標として形式化します。この定式化は、単純な共有しきい値を提供し、構築により固有のまれなルールを保存し、効率的なローカル更新をサポートします。 SkillZip には、1 つの構造化された抽出呼び出しと決定論的な最適化を備えたワンショット モードと、タスクの再生や完全な履歴の再解析を行わずに各自己進化パッチを統合する継続的な Zip-on-Write モードがあります。包括的な実験評価を通じて、圧縮パフォーマンス、汎用性、コストオーバーヘッドにおける SkillZip の有効性と優位性を実証します。

原文 (English)

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than reused. The resulting skill becomes expensive to inject and difficult to maintain. Generic prompt compression is ill-suited to this setting because a skill is not a flat passage: its name and description define when it applies, its workflow controls execution, its tool and output contracts constrain validity, and rare exceptions may remain essential even when no sampled task activates them. Evaluation-guided compression can test these behaviors, but it introduces rollouts, cost, and dependence on the compression-time evaluation set. We present SkillZip, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation. The intuition is explain once, reference many: state a repeated rule once at the scope where it applies, factor a repeated action sequence into a shared procedure, and keep only the differences as explicit exceptions. We formalize this intuition as a typed minimum description-length objective over a skill contract and a residual, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field. The formulation provides simple sharing thresholds, preserves unique rare rules by construction, and supports efficient local updates. SkillZip has a one-shot mode with one structured extraction call and deterministic optimization, and a continual Zip-on-Write mode that integrates each self-evolution patch without replaying tasks or reparsing the full history. Through comprehensive experimental evaluations, we demonstrate the effectiveness and superiority of SkillZip in compression performance, generalizability, and cost overhead.

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AI チャット エージェントをドッグフードする方法: 目標指向の NPC シミュレーションを備えた 3 層の評価フレームワーク

LLM チャット エージェントを導入する運用チームは、特定の品質保証ギャップに直面しています。既存の評価ツールは、個々の応答をテストしたり、社会的相互作用をシミュレートしたりしていますが、実際のユーザーが複数ターンの会話を通じて目標を達成できるかどうかを体系的に検証するものはありません。標準的なクエスチョンバンク テスト (レイヤー 1)、ランダム ウォーク マルチターン評価 (レイヤー 2)、および 5 つの構造化された目標タイプと 10 カテゴリの失敗分類 (レイヤー 3) を備えた目標指向 NPC (ノン プレイヤー キャラクター) シミュレーターを組み合わせることで、このギャップを埋める 3 層のドッグフーディング フレームワークを導入します。約 3 か月にわたる実稼働マルチエージェント システムに関する長期的なケース スタディ (257 回の評価実行、108 のシナリオ NPC スイート) では、3 つのレイヤーが相補的な回帰シグナルを生成することがわかりました。応答品質の層間の相関は、同期実行内では弱く (スピアマンの rho は -0.15 ~ 0.14)、長期的な系列全体では負です (rho は 0.15 から 0.14 まで)。 -0.46)、正規の正しさによって目標に向けた会話の成功が予測されないことが確認されました。 NPC シミュレーターは、実行あたり 0.17 ドル (人間による評価より 6,272 倍安い) で 77% の目標達成を達成し、自動化された PROMOTE/HOLD/ROLLBACK リリース決定による毎日の CI/CD 統合を可能にします。他のチームが独自の LLM チャット エージェントにフレームワークを採用できるように、完全なプロンプト テンプレート、障害分類、Python ファーストの複製可能性ガイドをリリースします。

原文 (English)

How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation

Production teams deploying LLM chat agents face a specific quality assurance gap: existing evaluation tools test individual responses or simulate social interactions, but none systematically verify whether real users can achieve their goals through multi-turn conversation. We introduce a three-layer dogfooding framework that bridges this gap by combining canonical question-bank testing (Layer 1), random-walk multi-turn evaluation (Layer 2), and a goal-directed NPC (Non-Player Character) simulator with five structured goal types and a ten-category failure taxonomy (Layer 3). In a longitudinal case study on a production multi-agent system over roughly three months (257 evaluation runs; a 108-scenario NPC suite), we find that the three layers produce complementary regression signals: cross-layer correlation for response quality is weak within a synchronized run (Spearman rho between -0.15 and 0.14) and negative across the longitudinal series (rho down to -0.46), confirming that canonical correctness does not predict goal-directed conversation success. The NPC simulator achieves 77 percent goal achievement at 0.17 dollars per run (6,272x cheaper than human evaluation), enabling daily CI/CD integration with automated PROMOTE/HOLD/ROLLBACK release decisions. We release full prompt templates, the failure taxonomy, and a Python-first replicability guide so that other teams can adopt the framework for their own LLM chat agents.

2026-08-12 13:00 JSTarXiv cs.AIビジネス/資金調達

Do AI weather models miss extremes?

First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic…

2026-08-12 13:00 JSTarXiv cs.AIビジネス/資金調達

ステータスの関連付けでは意思決定漏れを確実に予測できない

バイアス評価は、モデルが社会的関連性をコード化しているという証拠から、同じ関連性が結果的な決定を変えるだろうという主張に急速に移行することがよくあります。私たちは、制御された社会経済調査としてチリの姓を使用して、その推論が正当化されるかどうかをテストします。 8 つの凍結モデルプロバイダー セルをそれぞれ 1,032 のプロンプトで評価し、8,256 の検証済み一次応答が得られます。この設計では、学術選考、専門家の採用、研究フェローシップの選考、および法的扶助の利用を通じて、強制的な潜在的な関係を、一致する結果的な決定から分離します。エリートコード化された姓は、8 つのモデルのうち 7 つで一般的な姓よりも高い強制高ステータス確率質量を受け取り、8 つすべてのモデルで希少頻度コントロールよりも高い質量を受け取りました。しかし、エリートマイナス一般的な意思決定効果は、ほとんどのシステムでゼロに近かった。 5 つのモデルは、事前に宣言された (プラスマイナス)0.10 の標準偏差マージン内で統計的に同等でしたが、残りの 3 つは不正確または境界線にあり、一貫したエリートの利点はありませんでした。関連の強さは、モデル全体 (r = 0.201、p = 0.633) またはモデルごとに凍結された姓のペアのセル (r = 0.065、p = 0.565) にわたる決定漏れを確実に予測しませんでした。中心的な結果は測定の解離です。潜在的な社会的関連とその結果としての治療は経験的に異なる構造です。評価では、関連付けから行動への移行を直接測定する必要があります。

原文 (English)

Status Association Does Not Reliably Predict Decision Leakage

Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining three were imprecise or borderline, with no consistent elite advantage. Association strength did not reliably predict decision leakage across models (r = 0.201, p = 0.633) or across frozen surname-pair-by-model cells (r = 0.065, p = 0.565). The central result is a measurement dissociation: latent social association and consequential treatment are empirically distinct constructs. Evaluations should measure the transition from association to action directly.

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation ben…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation

Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization

Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs…

2026-08-12 13:00 JSTarXiv cs.AIビジネス/資金調達

The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election

The increasing use of Generative Artificial Intelligence (GenAI) in journalism raises concerns about possible detrimental effects both on j…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-c…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達研究/論文

Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deterior…

2026-08-12 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. Howe…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents

Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely valida…

2026-08-12 13:00 JSTarXiv cs.AIビジネス/資金調達

$\mathrm{ECI}_{\mathrm{sem}}$: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives

Hard-negative source selection for dense retrieval is usually decided only after fine-tuning and downstream evaluation. We propose ECIsem,…

2026-08-12 13:00 JSTarXiv cs.AIビジネス/資金調達

Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding

Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Evaluation and Hardening of LLM System Instructions Against Extraction via Encoding Attacks

System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensit…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Why Do Safety Guardrails Degrade Across Languages?

Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), whic…

2026-08-12 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Ablation-Corrected Evaluation of Attribution Maps in Echocardiographic Ejection-Fraction Models

Attribution maps for echocardiographic ejection-fraction models are evaluated by their overlap with an expert left-ventricular annotation,…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions

Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute…

2026-08-11 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodi…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverifi…

2026-08-11 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Guixu: オンチェーン認証による自律型 AI エージェント向けの評価主導のデータ検出

自律型エージェントは、モデルのトレーニングや意思決定サポートなどの下流タスクを完了するために外部データにますます依存しています。しかし、既存のデータ発見システムは依然として主に検索指向のままです。つまり、異種ソースから候補データセットを表面化しますが、タスク固有の有用性の推定、予算制約の下での費用対効果の高いデータセットの選択、または以前の使用からの信頼できるフィードバックの組み込みに対するサポートは限定的です。この文書では、自律エージェント向けの評価主導型データ検出システムである Guixu について説明します。 Guixu は、タスクを意識したデータ評価のために、プロキシラベル伝播とマルチラウンドナップザック最適化を備えた 3 フェーズ評価パイプラインを採用しています。 Guixu はエージェント支払いプロトコルを統合して、予算に制約のあるデータ調達ワークフローを可能にします。 Guixu は、検証可能なデータ検出のためにオンチェーン データ マーケットと認証シグナルを活用します。私たちのデモンストレーションでは、Guixu がどのようにしてエージェントがキーワードベースのデータセット取得を超えて、タスクと予算を意識した信頼できるデータの発見と調達に移行できるかを強調しています。参加者は、NL タスクの仕様やマルチソース検索からデータ評価や検証可能なトランザクション フィードバックに至るまで、ワークフロー全体を対話形式で探索できます。

原文 (English)

Guixu: Valuation-Driven Data Discovery for Autonomous AI Agents with On-Chain Attestation

Autonomous agents increasingly rely on external data to complete downstream tasks such as model training and decision support. However, existing data discovery systems remain largely retrieval-oriented: they surface candidate datasets from heterogeneous sources, but provide limited support for estimating task-specific utility, selecting cost-effective datasets under budget constraints, or incorporating trustworthy feedback from prior usage. This paper presents Guixu, a valuation-driven data discovery system for autonomous agents. Guixu employs a three-phase valuation pipeline with proxy-label propagation and multi-round knapsack optimization for task-aware data valuation. Guixu integrates agentic payment protocol to enable budget-constrained data procurement workflows. Guixu leverages on-chain data market and attestation signals for verifiable data discovery. Our demonstration highlights how Guixu enables an agent to move beyond keyword-based dataset retrieval toward task- and budget-aware, trustworthy data discovery and procurement. Attendees can interactively explore the full workflow, from NL task specification and multi-source search to data valuation and verifiable transaction feedback.

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Janus: 高額な評価予算の下での LLM 主導の発見のためのアルゴリズムと評価者の共進化フレームワーク

LLM 主導のプログラム検出は評価者の迅速なフィードバックに依存していますが、多くの科学および工学タスクでは高忠実度のシミュレーション、ハードウェア実行、または物理実験が必要であり、各評価にコストがかかります。安価なサロゲート評価器を使用するとこのコストを削減できますが、固定サロゲートは検索による分布シフトに対して脆弱であり、まばらで検索に偏ったラベルから確実に適合することが困難です。 Janus は、LLM を使用してターゲット プログラムと実行可能なプロキシ エバリュエーターを共進化させるフレームワークです。ラベル不足に対処するために、Janus は LLM にエンコードされたドメイン知識を活用してタスク固有の評価プログラムを生成し、実際の結果を使用してそれらを調整します。分布の変化を緩和するために、Janus はターゲット プログラムに沿って評価者を進化させ、プロモーションに合わせた目標を使用して評価者を選択し、オンライン信用更新により地域に応じたポートフォリオを維持します。代理予測は依然として誤りやすいため、Janus は候補者に優先順位を付けるためにのみ代理予測を使用し、候補者がターゲット プログラムの母集団に参加したり、現職者を更新したりする前に実際の検証を必要とします。 Janus は、5 つの科学および工学設計タスクにわたって、実際の評価予算を超えるこれまでの最良の改善曲線の下のより大きな領域と、ターゲット プログラムのみを進化させる一致したベースラインよりも高い最終パフォーマンスを達成しました。平均して、Janus はベースラインの最終改善の 99/% に達しますが、実際の評価は 59.1/% 減少します。進化したプロキシ評価者も、シード バージョンよりも正確に有望な候補をランク付けします。これらの結果を総合すると、評価者主導の LLM 発見が、安価でスケーラブルなフィードバックを伴うタスクから、信頼できる評価が希少で高価な科学分野まで拡張されます。

原文 (English)

Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets

LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks require high-fidelity simulations, hardware execution, or physical experiments, making each evaluation expensive. Cheap surrogate evaluators can reduce this cost, yet fixed surrogates are vulnerable to search-induced distribution shift and are difficult to fit reliably from sparse, search-biased labels. We introduce Janus, a framework that uses LLMs to co-evolve target programs and executable proxy evaluators. To address label scarcity, Janus leverages domain knowledge encoded in LLMs to generate task-specific evaluator programs and calibrates them using real outcomes. To mitigate distribution shift, Janus evolves evaluators alongside target programs, selects them using a promotion-aligned objective, and maintains region-conditioned portfolios with online credit updates. Because proxy predictions remain fallible, Janus uses them only to prioritize candidates and requires real validation before candidates can enter the target-program population or update the incumbent. Across five scientific and engineering design tasks, Janus achieves a larger area under the best-so-far improvement curve over the real-evaluation budget and higher final performance than a matched baseline that evolves only target programs. On average, Janus reaches 99/% of the baseline's final improvement with 59.1/% fewer real evaluations. Evolved proxy evaluators also rank promising candidates more accurately than their seed versions. Together, these results extend evaluator-guided LLM discovery from tasks with cheap, scalable feedback to scientific domains where trustworthy evaluation is scarce and expensive.

2026-08-11 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling

Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

AI Evaluation Should Measure Verification Cost, Not Correctness Alone

The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to v…

2026-08-11 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement

AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by s…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics

Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成エージェントビジネス/資金調達

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients

The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations

Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions larg…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

An evolutionary model of animats with VLM-based subjective evaluation

In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language Model (VLM) into the fitness e…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions

Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute…

2026-08-11 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two couple…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization

Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships betwe…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this…

2026-08-11 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol

AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after…

2026-08-11 13:00 JSTarXiv cs.AIビジネス/資金調達

Toward CT-Equivalent Image Quality in Low-Dose Radiotherapy Planning: Conditional Diffusion-Based CBCT-to-CT Synthesis and the Impact of CBCT Input Representation

During standard radiotherapy planning, repeated CT acquisitions are often required for patient registration, verification, and adaptive pla…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetoric…

2026-08-11 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達規制/政策

DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models

The proliferation of AI-generated images produced by Latent Diffusion Models (LDMs) has raised critical concerns regarding copyright infrin…

2026-08-11 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation

Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation,…

2026-08-11 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation

Recently, AI-driven video generation has attracted considerable attention. This surge increases the demand for reliable video quality asses…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans

Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as so…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges)…

2026-08-11 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search

Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instanc…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges

The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Sh…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for lon…

2026-08-11 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ベンチマーク推論が成立しない場合:AI評価における予測可能性

AI ベンチマークの結果が 1 ステップで重大な主張に達することはほとんどありません。評価者は、それをさらなるケースに一般化し、能力の証拠として解釈し、新しいタスクに推定し、別のシステムまたはサイトに移し、人間によるレビューと下流の結果に関する仮定と組み合わせます。妥当性中心のアプローチでは、各主張の証拠が必要です。この論文では、さらなる認識論的問題を特定します。それは、保証されたリンクは自動的に保証されたチェーンを作成しないということです。ある研究のターゲットが次の研究のソースになるとは限りません。システム、人口、結果、または条件はインターフェースで変更される可能性があります。また、共有データやモデル系統により、一見独立したサポートに依存する可能性があります。予測可能性は、観察されたケースから観察されていないケースへの限定された拡張が保証されるかどうかに関係します。グッドマンは競合拡張の問題を提供します。引数ベースの妥当性は、それらをテストするためのアーキテクチャを提供します。この論文の特徴的な主張は、非合成原理です。隣接する投影のサポートは、エンドポイントと仮定が一致し、依存性と不確実性が貫徹される場合にのみ合成を保証します。法的調査の事例は、ベンチマーク証拠と展開調査がそれぞれ並行性を保ちながらどのように健全であるかを示しています。再分析とシミュレーションは、骨材の安定性が後の予測で必要となる区別を消去できる理由を示しています。結果として得られる投影可能性監査により、ベンチマークで使用する引数内のサポートされていない結合が診断されます。

原文 (English)

When benchmark inferences do not compose: Projectibility in AI evaluation

An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper makes explicit and operationalizes a problem those approaches leave to the analyst: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The contribution is an interface audit for distributed AI evidence: typed source and target descriptions, and a procedure separating endpoints that never meet from endpoints that meet while warrant fails to cross. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A known-truth demonstration shows why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

自動アイテム評価: LLM によって生成された批評を使用してアイテムの受け入れと拒否を予測します。

自動品目評価 (AIE) とは、評価対象品目の専門家による手動レビューやフィールド テストを必要とせずに品目の品質を評価するための計算手法の使用を指します。私たちは、大規模な標準化されたテスト プログラムからの過去の不合格データを使用して、品目テキストから品目の合格と不合格を予測することにより、ほぼ包括的な AIE モデルを構築することを目指しました。データセットには 52,759 件の英語芸術 (ELA) と数学の項目が含まれており、そのうち 34% は将来の運用上の使用から永久に拒否されました。拒否の理由には、不十分な心理測定特性、コンテンツの問題、偏見と感受性への懸念、およびコンテンツ以外の問題が含まれていました。私たちは、生のアイテム テキストに対する DeBERTaV3 ラージ分類器、Qwen3 で生成されたアイテム批評に対する 2 番目の DeBERTa 分類器、および両方の表現を組み合わせた融合モデルを微調整しました。融合モデルは、最も強力な全体的なパフォーマンスを達成しました (精度 = 0.75、F1 = 0.64、AUC = 0.80、感度 = 0.64、特異性 = 0.81)。数学的予測 (F1 = .73、AUC = .86) は、ELA (F1 = .51、AUC = .72) よりもかなり正確でした。判定しきい値を 0.5 から 0.25 に下げると、ELA と数学の平均感度は 0.88 と 0.91 に上昇しましたが、特異度はそれぞれ 0.31 と 0.56 に低下しました。これは、アイテムを評価するよりも生成する方が安価である自動アイテム生成のコンテキストでは好ましいと考えられます。アイテムの生のテキストと一緒にアイテムの批評を組み込むことで、ほとんどの拒否理由でパフォーマンスが向上しました。このモデルでは、より困難な項目ほど高い拒否確率が割り当てられました。ただし、融合モデルは、特に ELA に関して、バイアス、機密性、公平性、またはアクセシビリティについてフラグが立てられた項目を特定するのに苦労しました。これらの調査結果は、テキストベースの AIE が一部の分野では実現可能であり、手動レビューやフィールドテストの負担を軽減する実用的なツールとなる可能性があることを示唆するとともに、公平性の懸念がある項目については人間によるレビューの重要性も強調しています。

原文 (English)

Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques

Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item acceptance and rejection from item text using historical rejection data from a large-scale standardized testing program. The dataset contained 52,759 English language arts (ELA) and mathematics items with 34% permanently rejected from future operational use. Rejection reasons included poor psychometric properties, content issues, bias and sensitivity concerns, and non-content issues. We fine-tuned a DeBERTaV3-large classifier on raw item text, a second DeBERTa classifier on Qwen3-generated item critiques, and a fusion model combining representations from both. The fusion model achieved the strongest overall performance (Accuracy = .75, F1 = .64, AUC = .80, Sensitivity = .64, Specificity = .81). Prediction for math (F1 = .73, AUC = .86) was considerably more accurate than ELA (F1 = .51, AUC = .72). Lowering the decision threshold from .5 to .25 raised average sensitivity for ELA and math to .88 and .91, while reducing specificity to .31 and .56, respectively, which may be preferable in automated item generation contexts where generating items is cheaper than evaluating them. Incorporating item critiques alongside raw item text improved performance across most rejection reasons. The model assigned higher rejection probabilities to more difficult items. However, the fusion model struggled to identify items flagged for bias, sensitivity, fairness, or accessibility, especially for ELA. These findings suggest that text-based AIE is feasible in some areas and may offer a practical tool for reducing the burden of manual review and field testing, while also underscoring the importance of human review for items with fairness concerns.

2026-08-11 13:00 JSTarXiv cs.AIビジネス/資金調達

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of la…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents

LLM agents require personal information for personalization in order to effectively act on users' behalf, but this raises privacy concerns…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks

Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be re…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standard…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Enhancing AI Interpretability with Localised Architectures

Recent advances in generative AI, especially powerful Large Language Models (LLMs), raise concerns over the interpretability, safety and su…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

KV キャッシュ量子化下のアライメント崩壊: 診断と軽減策

キーバリュー (KV) キャッシュ量子化は、大規模言語モデル (LLM) 推論メモリを削減するために広く使用されていますが、既存の評価は、安全性への影響を評価せず、複雑さと精度の測定のみに焦点を当てています。この研究では、KV キャッシュ量子化におけるアライメントの保存について調査します。 11 の命令調整モデル (3.8B ~ 72B) と 5 つのベンチマーク (1,894 プロンプト) にわたって、低ビット量子化が安全調整を静かに破壊する可能性があることがわかりました。Mistral-7B は、わずか 1.03 倍の複雑さで拒否の 15.2% を失い、普遍的な安全なビット幅は存在せず、標準メトリクスには見えない鋭いモデル固有の位相遷移があります。根本原因は幾何学的なものであることがわかりました。安全機能は、完全な表現空間のパープレキシティの平均よりも量子化ノイズに対して 10^2 ~ 10^3 倍脆弱な低次元の活性化部分空間を占めています。この観察に触発されて、私たちは各モデルを 3 つの機構的故障モードのいずれかに分類する診断であるチャネルごとの削減 (PCR) を提案します。安全性としての外れ値。安全性が外れ値チャネルと重なっており、より細かい粒度ではそれを救うことができません。多層希釈では、安全性が多くの層に分散され、層ごとの修正が失敗します。 PCR は、20 のキャリブレーション プロンプトを使用して、9 つの主要モデルすべてと、独立したファミリーからの 1 つの保留モデルについて正しい緩和方向を予測します。 PCR は、目に見えないプロンプト、モデル、および最大 97.2% の回復率を持つ KIVI を含むプロダクション クオンタイザー全体で一般化され、アテンションベースの割り当て方法が失敗する場合に成功します。結果として得られるトレーニング不要のプロトコルは、約 35 GPU 分を必要とし、最小限のメモリ オーバーヘッドで失われたアライメントの最大 97% を回復し、NVIDIA GPU 上の FP8 KV キャッシュを使用する運用 vLLM で確認された脆弱性に対処します。

原文 (English)

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact. In this study, we explore alignment preservation under KV cache quantization. Across eleven instruction-tuned models (3.8B-72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment: Mistral-7B loses 15.2% of its refusals at only 1.03x perplexity, and no universal safe bit-width exists, with sharp model-specific phase transitions invisible to standard metrics. We identify that the root cause is geometric: safety features occupy a low-dimensional activation subspace 10^2-10^3x more vulnerable to quantization noise than the full representation space perplexity averages over. Inspired by this observation, we propose Per-Channel Reduction (PCR), a diagnostic that classifies each model into one of three mechanistic failure modes: outlier-crushes-safety, where safety lives in non-outlier channels collaterally damaged by outlier-driven scale factors; outlier-as-safety, where safety overlaps outlier channels and finer granularity cannot rescue it; and multi-layer dilution, where safety is distributed across many layers and per-layer fixes fail. PCR predicts the correct mitigation direction on all nine primary models and one held-out model from an independent family using 20 calibration prompts. PCR generalizes across unseen prompts, models, and production quantizers, including KIVI with up to 97.2% recovery, succeeding where attention-based allocation methods fail. The resulting training-free protocol, requiring approximately 35 GPU-minutes, recovers up to 97% of lost alignment at minimal memory overhead, addressing vulnerabilities confirmed in production vLLM serving with FP8 KV cache on NVIDIA GPUs.

2026-08-10 21:00 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

Discovered Materials is playing AI whack-a-mole to hunt cooler chips

Discovered Materials raised $9 million to fund the hunt for more novel materials to build more efficient chips.

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

ADIAS: インタラクティブなエージェント システムの自動設計

自動化されたエージェント設計は、反復的な修正、評価、フィードバックの要約を通じてエージェントのハーネスを改善します。既存の方法は主に候補者中心です。クロスラウンドのエクスペリエンスは候補者エージェントを中心に編成され、修復の進行状況が暗黙的に残ります。これにより、修復のターゲットが非効率になり、部分的な進行の統合が遅くなり、ラウンド全体に非効果的な介入が伝播することになります。したがって、問題中心のエージェント最適化を定式化します。この最適化では、修復の進行状況が各ラウンドの候補履歴から再導出されるのではなく、最適化をガイドする明示的な永続的な問題状態として引き継がれます。 2 つのメカニズムを備えた自動フルコード エージェント設計のフレームワークである ADIAS で定式化をインスタンス化します。永続的な問題の状態では、安定した問題の ID、ライフサイクルのステータス、裏付けとなる証拠、介入結果の履歴が維持されます。問題に基づく最適化では、この状態を使用して、その後の集中的なコード全体の変更のための修復ターゲットとリビジョンの方向を共同で提案します。 5 つのインタラクティブなベンチマーク全体で、ADIAS は最も強力なベースラインを平均 25.2% 上回っており、4 つのバックボーン モデル全体で一貫した利益を達成しています。さらに、制御されたアブレーションにより、永続的な問題の状態を除去したり、問題中心の改訂を候補者中心の政策に置き換えたりすると、パフォーマンスが最大 40.7% 低下することが示されています。

原文 (English)

ADIAS: Automated Design of Interactive Agentic Systems

Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves the repair progress implicit. This causes inefficient repair targeting, slow consolidation of partial progress, and propagation of ineffective interventions across rounds. Therefore, we formulate issue-centric agent optimization, in which repair progress is carried forward as an explicit persistent issue state to guide optimization, rather than re-derived from candidate history in each round. We instantiate the formulation in ADIAS, a framework for automated full-code agent design with two mechanisms. A persistent issue state maintains stable issue identities, lifecycle status, supporting evidence, and intervention-outcome histories. Issue-guided optimization uses this state to jointly propose repair targets and revision directions for subsequent focused full-code modification. Across five interactive benchmarks, ADIAS outperforms the strongest baseline by 25.2% on average and achieves consistent gains across four backbone models. Controlled ablations further show that removing persistent issue state or replacing issue-centric revision with candidate-centric policies leads to performance drops of up to 40.7%.

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

自動アイテム評価: LLM によって生成された批評を使用してアイテムの受け入れと拒否を予測します。

自動品目評価 (AIE) とは、評価対象品目の専門家による手動レビューやフィールド テストを必要とせずに品目の品質を評価するための計算手法の使用を指します。私たちは、大規模な標準化されたテスト プログラムからの過去の不合格データを使用して、品目テキストから品目の合格と不合格を予測することにより、ほぼ包括的な AIE モデルを構築することを目指しました。データセットには 52,759 件の英語芸術 (ELA) と数学の項目が含まれており、そのうち 34% は将来の運用上の使用から永久に拒否されました。拒否の理由には、不十分な心理測定特性、コンテンツの問題、偏見と感受性への懸念、およびコンテンツ以外の問題が含まれていました。私たちは、生のアイテム テキストに対する DeBERTaV3 ラージ分類器、Qwen3 で生成されたアイテム批評に対する 2 番目の DeBERTa 分類器、および両方の表現を組み合わせた融合モデルを微調整しました。融合モデルは、最も強力な全体的なパフォーマンスを達成しました (精度 = 0.75、F1 = 0.64、AUC = 0.80、感度 = 0.64、特異性 = 0.81)。数学的予測 (F1 = .73、AUC = .86) は、ELA (F1 = .51、AUC = .72) よりもかなり正確でした。判定しきい値を 0.5 から 0.25 に下げると、ELA と数学の平均感度は 0.88 と 0.91 に上昇しましたが、特異度はそれぞれ 0.31 と 0.56 に低下しました。これは、アイテムを評価するよりも生成する方が安価である自動アイテム生成のコンテキストでは好ましいと考えられます。アイテムの生のテキストと一緒にアイテムの批評を組み込むことで、ほとんどの拒否理由でパフォーマンスが向上しました。このモデルでは、より困難な項目ほど高い拒否確率が割り当てられました。ただし、融合モデルは、特に ELA に関して、バイアス、機密性、公平性、またはアクセシビリティについてフラグが立てられた項目を特定するのに苦労しました。これらの調査結果は、テキストベースの AIE が一部の分野では実現可能であり、手動レビューやフィールドテストの負担を軽減する実用的なツールとなる可能性があることを示唆するとともに、公平性の懸念がある項目については人間によるレビューの重要性も強調しています。

原文 (English)

Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques

Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item acceptance and rejection from item text using historical rejection data from a large-scale standardized testing program. The dataset contained 52,759 English language arts (ELA) and mathematics items with 34% permanently rejected from future operational use. Rejection reasons included poor psychometric properties, content issues, bias and sensitivity concerns, and non-content issues. We fine-tuned a DeBERTaV3-large classifier on raw item text, a second DeBERTa classifier on Qwen3-generated item critiques, and a fusion model combining representations from both. The fusion model achieved the strongest overall performance (Accuracy = .75, F1 = .64, AUC = .80, Sensitivity = .64, Specificity = .81). Prediction for math (F1 = .73, AUC = .86) was considerably more accurate than ELA (F1 = .51, AUC = .72). Lowering the decision threshold from .5 to .25 raised average sensitivity for ELA and math to .88 and .91, while reducing specificity to .31 and .56, respectively, which may be preferable in automated item generation contexts where generating items is cheaper than evaluating them. Incorporating item critiques alongside raw item text improved performance across most rejection reasons. The model assigned higher rejection probabilities to more difficult items. However, the fusion model struggled to identify items flagged for bias, sensitivity, fairness, or accessibility, especially for ELA. These findings suggest that text-based AIE is feasible in some areas and may offer a practical tool for reducing the burden of manual review and field testing, while also underscoring the importance of human review for items with fairness concerns.

2026-08-10 13:00 JSTarXiv cs.AIビジネス/資金調達

NxN E-valuation: コンフォーマル CRT ヌルによる仮説証明

私たちは、十分な大きさのデータセットが利用できる限り、専用の帰無仮説の構築など、ケース固有の証明手順を構築することなく仮説を検証できる、便利な電子値ベースの仮説証明アルゴリズムである NxN E-valuation を提案します。この方法は、LLM ベースの探索システムに特に適しています。LLM は仮説を提案するのが非常に得意ですが、幻覚にひどく悩まされます。この幻覚により、LLM 出力を直接収集することができなくなり、既存の治療法はいずれも不十分です。最も一般的な解決策には、導入部で詳述したその他の救済策の中でも特に、循環検証とホールドアウト テスト (誤った仮説が依然として偽の相関を介して通過する可能性がある) を LLM に検証または修正させることが含まれます。これを解決するために、NxN E-valuation は自然に存在する大規模なトレーニング セットを利用し、異なるサンプルを相互に帰無仮説として機能させます。この設計は、各仮説を証明する条件付きランダム化テスト (CRT) を直接実現します。このアプローチは、LLM の世代が個々のサンプルに適用される仮説である場合、少なくとも LLM 循環検証とホールドアウト データ テストに代わる一般的により優れた代替手段となる可能性があります。

原文 (English)

NxN E-valuation: Hypothesis Certification via a Conformal CRT Null

We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a large enough dataset is available. The method is especially suited to LLM-based exploration systems, where LLMs are remarkably good at proposing hypotheses but suffer badly from hallucination; this hallucination prevents us from harvesting LLM outputs directly, and existing remedies each fall short. The most common solutions include letting the LLM verify or correct itself circular verification and held-out testing (where false hypotheses can still pass via spurious correlations), among other remedies detailed in the introduction. To resolve this, NxN E-valuation exploits the naturally existing large training set and lets different samples serve as null hypotheses for one another. This design directly realizes a conditional randomization test (CRT) that certifies each hypothesis. The approach can be a universally better replacement for at least LLM circular verification and held-out-data testing, provided the LLM's generations are hypotheses that apply to each individual sample.

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

サイエンス エッジの評価: 真の科学的発見に向けて欠けているステップを確認する

大規模言語モデル (LLM) は科学的発見にますます関与していますが、それらが複雑な実際の実験室科学をサポートできるかどうかは依然として不明です。ここでは、化学、生物学、材料科学における査読済みの文献と実験実践に基づいて、専門家によって精選された質問のマルチモーダルなベンチマークであるサイエンス エッジ評価 (SEE) を紹介します。 19 個のマルチモーダル大規模言語モデル (MLLM) を評価したところ、最もパフォーマンスの高いモデルであっても精度が 48.7% にしか達していないことがわかりました。さらに、汎用モデルは科学特化モデルよりも平均して優れています。視覚エージェントの評価では、ツールを使用すると最高の精度が 52.7% に向上しました。ツールを使用するとモデルで利用できる情報が増える可能性がありますが、情報が増えても信頼できる科学的推論が得られるとは限りません。重要な課題は、モデルが元の実験証拠の範囲内でツール由来の情報を管理できるかどうかです。これらの発見を総合すると、現在の MLLM は、実際の科学的発見に不可欠な能力である、実験結果から正当かつ証拠に基づいた推論を確実に行うことがまだできないことが明らかになります。このギャップを埋めるには、MLLM が確立された科学概念の説明から、実験データから新しい証拠に基づいた洞察を導き出すことに移行する必要があります。

原文 (English)

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

極めて重要な投票が見えない: 集計の独立性指標は検証が実際に役立つところを見逃している

LLM 審査員パネルは標準的な評価ツールですが、これまでの研究では、相関性の高いパネルエラーが報告されています。つまり、9 人の審査員が 2 人の独立した審査員の有効な情報をほぼ提供しており、集計によってその差が縮まるのはごく一部に過ぎません。自然療法(テストスイートの実行など、異なる証拠源からの信号)では、パネルの有効投票数に大きな変化は見られませんでした(-0.04、95\% CI [-0.10、+0.02])。集計依存性と条件付き決定ユーティリティは別の問題です。基本的多数決算術は、単一投票の置換のために影響を受けるセットを修正します。変更できるのは 1 票のマージンを持つ決定のみです。経験的な問題は、パネルのエラー率が上昇し、有用な代替品がそこに集中するかどうかです。全体の精度の向上はこれらの重要なクエリに集中しており、精度の向上が大きい (3 つのヘッドライン設定全体で +10.4 ~ +23.3 パーセント ポイント) のですが、その他の部分ではまったくゼロです。 3 つのコード ベンチマークと 4 つのパネル サイズ (9 ジャッジ拡張と 56 の依存サブサンプリング チェック、+6.5 ~ +16.1 パーセント ポイントのゲイン) にわたってパターンを確認しました。 HumanEval+/MBPP+ では、多数派側置換ルールにより全体の精度が 82.44\% から 85.62\% に向上し、クエリの 16.2\% でシグナルが呼び出されます。シグナルのみは 87.60\% で依然として強いです。したがって、人口レベルの依存診断とマージン層別効用は補完的であり、影響を受けるセットの特徴付けにより、指定された単一投票代替政策に対するコール削減ルールが得られます。

原文 (English)

Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps

LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the effective information of two independent ones, and aggregation closes only a small fraction of the gap. A natural remedy--a signal from a different evidence source, e.g., executing a test suite--produced no distinguishable change in the panel's effective-vote count at scale (-0.04, 95\% CI [-0.10, +0.02]). Aggregate dependence and conditional decision utility are different questions. Elementary majority arithmetic fixes the affected set for single-ballot substitution: only decisions with a one-vote margin can change. The empirical question is whether panel error rates rise and useful substitutions concentrate there. They do: the entire accuracy gain concentrates on these pivotal queries, where it is large (+10.4 to +23.3 percentage points across three headline configurations), and is exactly zero elsewhere. We confirm the pattern across three code benchmarks and four panel sizes (a 9-judge extension and 56 dependent subsampling checks, gain +6.5 to +16.1 percentage points). On HumanEval+/MBPP+, a majority-side replacement rule raises overall accuracy from 82.44\% to 85.62\% while invoking the signal on 16.2\% of queries; signal-only remains stronger at 87.60\%. Thus population-level dependence diagnostics and margin-stratified utility are complementary, and the affected-set characterization yields a call-reduction rule for any specified single-ballot substitution policy.

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

創造性のレシピ: 大規模な言語モデルでの反復生成と評価

生成モデルは単一の成果物を通じて評価されることがよくありますが、人間の創造性は通常、反復的な生成、評価、改良を通じて現れます。このパイロット研究では、FunSearch を 2024 年のピルズベリー ベイクオフのレシピ生成に適応させ、TTCT ベースの LLM 評価を使用して人間のベンチマークに対して出力を評価することで、反復検索が LLM の創造性を向上させるかどうかを検証します。 2 つの実験にわたって、反復回数、ジェネレーターの温度、ループ内選択スコアラー モデルのサイズをテストします。結果は、世代選択の反復により人間のベンチマークに匹敵する創造性スコアを持つレシピを生成できるが、反復を追加するだけでは創造性が向上しないことを示しています。ループ内評価器が最も重要です。選択スコアラーが小さいほど、ほとんどの TTCT 次元にわたって大幅に高いスコアが得られますが、温度の影響は独自性を除いて限定的です。これらの発見は、評価者の設計が主観的なクリエイティブ検索における一次設計変数であることを示唆しています。

原文 (English)

Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting FunSearch to recipe generation for the 2024 Pillsbury Bake-Off and evaluating outputs against human benchmarks using TTCT-based LLM evaluation. Across two experiments, we test iteration count, generator temperature, and in-loop selection-scorer model size. Results show that iterative generation-selection can produce recipes with creativity scores comparable to human benchmarks, but additional iterations alone do not improve creativity. The in-loop evaluator matters most: a smaller selection scorer yields significantly higher scores across most TTCT dimensions, while temperature has limited effects except for originality. These findings suggest that evaluator design is a first-order design variable in subjective creative search.

2026-08-10 13:00 JSTarXiv cs.AIビジネス/資金調達

Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation

Cryptanalytic extraction has been demonstrated for ReLU networks, for networks using componentwise activations such as GELU or SiLU, and fo…

2026-08-10 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and pose…

2026-08-10 13:00 JSTarXiv cs.AIビジネス/資金調達

Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering

Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereb…

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contaminati…

2026-08-10 13:00 JSTarXiv cs.AIビジネス/資金調達

Assessing AI-generated music detection in real-world broadcast monitoring

The proliferation of AI-generated music in broadcast media raises concerns about transparency and fair compensation, but reliable detection…

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

LLM ベースのエージェント評価のための統一フレームワークの必要性

Large Language Model (LLM) の出現により、汎用エージェントは根本的な進歩を遂げました。ただし、これらのエージェントの評価には、静的な QA ベンチマークとは異なる特有の課題が存在します。現在のエージェントのベンチマークは、システム プロンプト、ツールセット構成、環境ダイナミクスなどの外部要因によって大きく混乱していることが観察されています。既存の評価は、断片化された研究者固有のフレームワークに依存していることが多く、推論やツールの使用に関する即時エンジニアリングが大幅に異なるため、パフォーマンスの向上がモデル自体によるものであると考えるのが困難です。さらに、標準化された環境データが欠如しているため、追跡不可能なエラーや再現不可能な結果が発生します。この標準化の欠如は、現場に大きな不公平性と不透明性をもたらします。私たちは、エージェント評価を厳密に進めるためには、統一された評価枠組みが不可欠であると提案します。この目的を達成するために、エージェント評価の標準化を目的とした提案を紹介します。

原文 (English)

"LLM Agent Performance" Is Not a Single Evaluation Target

LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model factors by evaluating candidate models under the same configuration, making observed differences more attributable to the models themselves. However, model comparison is only one use of agent benchmarks. Other evaluations compare complete agent systems or test whether a fixed model or system remains stable across predeclared changes in its operating conditions. These results can all be reported under the common label of "LLM agent performance." Our position is that "LLM agent performance" does not denote a single evaluation target. Model comparisons under a reference stack and comparisons of complete agent systems answer different questions, while robustness asks whether either conclusion persists across predeclared conditions. The claim supported by a score therefore depends on the declared candidate boundary and condition policy. We derive implications for leaderboards, result reporting, and benchmark versioning, showing how distinguishing these classes preserves fair comparison while accommodating system innovation and robustness analysis.

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

VESTA: LLM エージェント向けの完全に自動化されたシナリオ生成および安全性評価フレームワーク

大規模言語モデル (LLM) は、単純なテキストベースの対話システムから、メモリを維持し、ツールを使用し、外部環境にアクセスし、タスクを実行できる LLM エージェントへとますます進化しています。彼らの能力と自律性が拡大するにつれて、彼らが直面する安全リスクもより多様になります。既存の評価は、手動で作成されたシナリオ、静的なプロンプト、または最終出力の判断に依存していることが多く、タスクの実行中にエージェントが直面する可能性のあるさまざまなリスクを把握することが困難です。 LLM エージェント向けの完全に自動化されたシナリオ生成および安全性評価フレームワークである VESTA を紹介します。 VESTA は、5 つのリスク次元に基づいて、現実世界のタスク実行における抽象的で多様な安全リスクを 1,072 の測定可能な評価シナリオにインスタンス化します。自動評価パイプラインを使用して、12 個の LLM エージェントが 2 つの権限コンテキストの下で評価されます。その結果、現在のエージェントはタスク実行中に依然として重大な行動安全リスクに直面しており、平均 ASR は 47.1%、いくつかのモデルは 70% を超えていることが示されています。これらの調査結果は、LLM エージェントの安全性を理解し改善するために、実行可能なプロセスレベルの評価が重要であることを示しています。

原文 (English)

ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks. As their capabilities and autonomy expand, the safety risks they face also become more diverse. Existing evaluations often rely on manually written scenarios, static prompts, or final-output judgments, making it difficult to capture the diverse risks that agents may face during task execution. We introduce ForesightSafety-SAGE, a fully automated scenario generation and safety evaluation framework for LLM agents. Based on five risk dimensions,we instantiae abstract and diverse safety risks in real-world task execution into 1,072 measurable evaluation scenarios. Using the automated evaluation pipeline, 12 LLM agents are evaluated under two authority contexts. The results show that current agents still face substantial behavioral safety risks during task execution, with an average ASR of 47.1% and several models exceeding 70%. These findings demonstrate the importance of executable, process-level evaluation for understanding and improving LLM agent safety.

2026-08-10 13:00 JSTarXiv cs.AIビジネス/資金調達

SenWorld: コンテキストリッチな評価データを生成するためのデジタル ツイン シミュレーション

スマートフォンのパーソナル アシスタントは長期にわたる個人データを推論しますが、その評価には正解がわかっているコンテキストに富んだ評価データが必要であり、実際のデバイスのトレースはプライバシーに敏感すぎて共有できません。この課題に対処するために、構築によって固定されたグラウンド トゥルースを使用してそのようなデータを生成する、物理的に接地され、決定論的でイベント ソースのデジタル ツイン シミュレーションである SenWorld を紹介します。 SenWorld では、ペルソナは実際の地図、天気、休日、ネットワーク データから構築された世界で 1 日を過ごします。観測可能なすべての信号はシステム全体のスナップショットにアーカイブされます。また、各評価ケースは、事後注釈や大規模言語モデル (LLM) ジャッジではなく、既存のレコードへのポインターによってラベル付けされます。この手法を北京の 16 人のペルソナで評価しました。生成されたデータは、カテゴリ分布 (ジェンセンとシャノンの相違 (JSD) 0.070) および通信記録の 1 日のリズム (JSD 0.1 未満) において、保持されている実際のユーザーのベンチマークと厳密に一致していますが、生成された記録は実際の記録よりも短いままです。スクリプトによる対話がなければ、ペルソナは完全に往復する対話サブグラフと差別化された行動レパートリーを形成します。 717 件の評価ケースに投影された生成データでは、実稼働スマートフォン アシスタントの 78 件の障害が明らかになり、通話とショート メッセージ サービス (SMS) の記録に集中し、連絡先、スケジュール、アラームは決して失敗しませんでした。スナップショット ポインタは、LLM 判定者が関与せずに、各失敗をアシスタント側の取得エラーとして確認します。全体として、SenWorld は、ラベルが構築によって固定されている評価データへの、プライバシーに安全で再現可能で配布がチェックされたパスを提供します。

原文 (English)

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data

Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, we present SenWorld, a physically grounded, deterministic, event-sourced digital-twin simulation that generates such data with ground truth fixed by construction. In SenWorld, personas live through a full day in a world built from real map, weather, holiday, and network data; every observable signal is archived in full-system snapshots; and each evaluation case is labeled by a pointer to an existing record rather than by post-hoc annotation or a large language model (LLM) judge. We evaluate this method with 16 personas in Beijing. The generated data closely matches the held-out real-user benchmark in category distribution (Jensen--Shannon divergence (JSD) 0.070) and in the daily rhythm of communication records (JSD below 0.1), though generated records remain shorter than real ones. Without scripted interaction, personas form a fully reciprocated dialogue subgraph and differentiated behavioral repertoires. Projected into 717 evaluation cases, the generated data exposes 78 failures in a production smartphone assistant, concentrating on call and Short Message Service (SMS) records while contacts, schedules, and alarms never fail. The snapshot pointer confirms each failure as an assistant-side retrieval error, with no LLM judge involved. Overall, SenWorld offers a privacy-safe, reproducible, and distribution-checked path to evaluation data whose labels are fixed by construction.

2026-08-10 13:00 JSTarXiv cs.AIビジネス/資金調達規制/政策

Provable Training Data Identification for Large Language Models

Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. Howev…

2026-08-10 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains

We argue that the current practice of evaluating AI/ML time-series forecasting models, predominantly on benchmarks characterized by strong,…

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Playing Games with My Heart: An Evaluation of AI Companion Apps

The use of chatbots for various forms of companionship is growing rapidly, raising a myriad of questions about simulated relationships, emo…

2026-08-10 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Rethinking Evaluation Paradigms in IBP-based Certified Training

Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations. Neura…

2026-08-10 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection

Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance ofte…

2026-08-08 00:20 JSTOpenAILLM/生成AIビジネス/資金調達

Responding to the next frontier of critical cyber capabilities

OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

TriQua: 事実評価における粒度とコンテキストの調和

LLM の事実性評価の「分解してから検証する」パラダイムは、根本的なトレードオフに直面しています。つまり、アトミックな事実、つまり 1 つの情報単位を伝える 1 つの文では、重要なコンテキストが省略されることがよくありますが、より広範なステートメントでは、正確な評価に必要な粒度が欠如しています。これに対処するために、複雑さに基づいて事実を柔軟にモデル化するフレームワークである TriQua を紹介します。単純なクレームは標準のトリプルとして抽出されますが、複雑なクレームは補助的な文脈修飾子を付けることによって超関係ファクトとして表されます。この適応構造は、アトミック性を犠牲にすることなく、正確な検索と検証に必要なコンテキストを保存します。さらに、TriQua の検証プロセスは、特定のトリプルおよび修飾子内の具体的なエラーに直接注釈を付け、エラー検出に対するきめ細かい説明可能性を提供します。フレームワークと並行して、これらの構造化された事実単位の事実性を定量化する TriQuaScore を提案します。経験的評価では、TriQuaScore が人間による注釈付き事実スコアと強く一致し、TriQua が堅牢な分解品質を達成し、証拠に基づく事実検証において既存の分解ベースのフレームワークを上回るパフォーマンスを示していることが示されています。

原文 (English)

TriQua: Reconciling Granularity and Context in Factuality Evaluation

The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveying one unit of information, often omit essential context, while broader statements lack the granularity needed for precise assessment. To address this, we introduce TriQua, a framework that flexibly models facts based on their complexity. Simple claims are extracted as standard triples, while complex claims are represented as hyperrelational facts by attaching auxiliary contextual qualifiers. This adaptive structure preserves the necessary context for accurate retrieval and verification without sacrificing atomicity. Furthermore, TriQua's verification process directly annotates concrete errors within specific triples and qualifiers, providing fine-grained explainability for error detection. Alongside the framework, we propose TriQuaScore to quantify the factuality of these structured fact units. Empirical evaluations show that TriQuaScore strongly aligns with human annotated factuality scores, TriQua achieves robust decomposition quality, and outperforms existing decomposition-based frameworks in evidence-based fact verification.

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

LUNAR: ユニバーサル ユーザー BehAvioR ログでのパーソナライズされた大規模言語モデルのベンチマーク

既存のパーソナライズされた LLM ベンチマークは、主にテキストのペルソナまたは分離された行動シグナルに依存しており、クロスドメインの行動パーソナライゼーションの評価は限定的であり、応答は異種混合の日常生活活動に基づいている必要があります。このギャップに対処するために、私たちは LUNAR を導入します。LUNAR は、衣服、食べ物、住居、移動などの普遍的な日常生活の領域にわたる縦断的なアプリのインタラクション履歴からの応答を、LLM がどのようにパーソナライズするかを評価するための最初のベンチマークです。データの希薄性とプライバシーの懸念を軽減しながら、スケーラブルなベンチマークの構築をサポートするために、LUNAR は、現実世界の動作パターンに基づいた多段階の粗いから細かいまでの合成パイプラインを使用します。忠実度分析は、他の合成ベンチマークよりも実際の行動分布とのより密接な一致を示します。 19 の主流 LLM に関する実験では、行動ログへのアクセスは必要ですが、深いパーソナライゼーションには十分ではないことが示されています。より多くのコンテキストやより大きなモデルは、より良いパフォーマンスを保証しません。効果的なパーソナライゼーションは、ドメイン全体で関連する証拠を選択して統合するかどうかにかかっています。詳細な行動記録の直接取得は圧縮メモリよりも常に優れたパフォーマンスを発揮しますが、より強力なパーソナライゼーションにはプライバシー保護が犠牲になる可能性があります。これらの調査結果は、証拠の選択、クロスドメイン統合、およびプライバシー管理がパーソナライズされた LLM の主要な課題であることを特定します。

原文 (English)

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for evaluating how LLMs personalize responses from longitudinal app interaction histories across universal daily-life domains, including clothing, food, housing, and mobility. To support scalable benchmark construction while mitigating data sparsity and privacy concerns, LUNAR uses a multi-stage coarse-to-fine synthesis pipeline grounded in real-world behavioral patterns. Fidelity analyses show closer alignment with real behavioral distributions than other synthetic benchmarks. Experiments on 19 mainstream LLMs show that access to behavioral logs is necessary but not sufficient for deep personalization: neither more context nor larger models guarantees better performance; effective personalization depends on selecting and integrating relevant evidence across domains. Direct retrieval of fine-grained behavioral records consistently outperforms compressed memory, while stronger personalization can come at the cost of privacy protection. These findings identify evidence selection, cross-domain integration, and privacy control as key challenges for personalized LLMs.

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

SkillTV-Bench: スキル強化されたエージェント執行における裁判官のパフォーマンスのベンチマーク

LLM エージェントは、ツールの使用と環境の相互作用を通じて長期的なタスクを実行することが増えており、評価は最終応答のスコアリングから完全な実行の検証に移行しています。スキル強化されたエージェントの場合、検証にはさらに、タスク時のスキルにエンコードされた手順知識が必要です。これは、この知識がどの証拠を検査するか、どの失敗がタスクに重大であるかを示すためです。ただし、既存のジャッジベンチマークは、最終的な応答や静的な軌跡を公開することが多く、タスク時のスキルと直接検査可能なアーティファクトや環境を組み合わせることはほとんどありません。そこで、11 のドメインにわたる 50 のタスクからの実際のエージェントの軌跡の 681 ケースのベンチマークである SkillTV-Bench を紹介します。これは、LLM-as-a-Judge メソッドと Agent-as-a-Judge メソッドの両方のスキルを意識​​した軌跡検証を評価するように設計されています。さらに、検証知識を再利用可能な JudgeSkill として外部化する SkillTV-Evolve を提案します。これにより、エージェント裁判官が対象を絞った検査を計画し、証拠に基づいた評決を下すことができます。連携のない開発プールでは、自動進化ループにより、誤判定されたケースを使用して JudgeSkill がさらに洗練されます。 SkillTV-Bench では、洗練されたスキルにより、同じエージェントの裁判官の精度が 14.8 パーセント ポイント向上します。オフラインのロールアウト プール選択では、選択された軌道の成功率が 1 回のロールアウトの 22.9% から 10 回のロールアウトの 45.5% に増加します。コードとデータは https://github.com/HanZhi306/SkillTV-Bench で入手できます。

原文 (English)

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowledge indicates what evidence to inspect and which failures are task-critical. However, existing judge benchmarks often expose final responses or static trajectories, and rarely combine task-time skills with directly inspectable artifacts and environments. We therefore introduce SkillTV-Bench, a 681-case benchmark of real agent trajectories from 50 tasks across eleven domains, designed to evaluate skill-aware trajectory verification for both LLM-as-a-Judge and Agent-as-a-Judge methods. Additionally, we propose SkillTV-Evolve, which externalizes verification knowledge as a reusable JudgeSkill that guides an agent judge to plan targeted inspections and issue evidence-grounded verdicts. On a disjoint development pool, an automated evolution loop further refines the JudgeSkill using misjudged cases. On SkillTV-Bench, the refined skill increases the same agent judge's accuracy by 14.8 percentage points. In offline rollout-pool selection, it increases selected-trajectory success from 22.9% with one rollout to 45.5% with ten rollouts. The code and data are available at https://github.com/HanZhi306/SkillTV-Bench

2026-08-07 13:00 JSTarXiv cs.AIビジネス/資金調達

防衛および国家安全保障のオントロジー間の相互運用性の向上: 分析および評価タスク

オントロジーとナレッジグラフの使用は、防衛と国家安全保障の分野でますます普及してきています。学界、産業界、政府主導の取り組みを通じて、数多くのオントロジーが開発されてきました。このドメインの広さと専門性により、多様な防衛および国家安全保障のオントロジーにわたる相互運用性を実現することは依然として大きな課題です。この作業では、60 を超える公的に利用可能なオントロジーを分析して文書化し、オントロジー調整評価イニシアチブ (OAEI) の新しいトラックを導入します。このトラックは、8 つのマッチング タスク、コンセンサス調整、および手動でキュレーションされた (シルバー スタンダード) マッピングで構成されます。コンセンサス アライメントは、いくつかの最先端のオントロジー アライメント システムの出力を集約することによって導出されます。シルバースタンダードは、独自のマッピング (つまり、1 つのシステムのみによって提案されるマッピング) のサブセットとともにコンセンサス調整を手動で検証することによって取得されます。

原文 (English)

Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks

The use of ontologies and knowledge graphs is becoming increasingly widespread in the defence and national security domain. Numerous ontologies have been developed through initiatives led by academia, industry, and government. Achieving interoperability across diverse defence and national security ontologies remains a major challenge due to the domain's breadth and specialisation. In this work, we analyse and document over 60 publicly available ontologies and introduce a new track for the Ontology Alignment Evaluation Initiative (OAEI). This track comprises eight matching tasks, consensus alignments and manually-curated (silver-standard) mappings. The consensus alignments are derived by aggregating the outputs of several state-of-the-art ontology alignment systems. The silver-standard is obtained from the manual validation of the consensus alignment together with a subset of the unique mappings (i.e., mappings suggested by only one system).

2026-08-07 13:00 JSTarXiv cs.AIビジネス/資金調達

ECG-LENS: リードを意識した臨床コンテキストを強化した ECG レポートの生成と評価

心電図検査 (ECG) は、心血管疾患を診断するために最も広く使用されている非侵襲的ツールの 1 つですが、マルチリード ECG 記録を信頼できる臨床レポートに変換することは依然として困難です。 ECG レポートの生成を自動化すると、臨床医の解釈作業負荷が軽減され、診断効率が向上し、サービスが十分に受けられていない地域での心臓評価へのアクセスが拡大する可能性があります。画像ベースのレポート作成タスクとは異なり、ECG 解釈では、微妙な時間的形態の分析と、それに続く緻密な臨床用語で表現された一貫した診断推論が必要です。既存のシステムは主に分類に焦点を当てていますが、現在のレポート生成方法では、実際の臨床使用には依然として不十分な出力が生成されることがよくあります。これらの課題に対処するために、マルチリード信号モデリング、診断を意識した表現、および臨床に基づいたテキスト生成を統合するエンドツーエンドの ECG レポート生成フレームワークである ECG-LENS を提案します。 ECG-LENS は、局所的な波形形態を保存するリードごとのエンコーダと、リード間の依存関係を捕捉するグローバル エンコーダを組み合わせます。レポート生成をガイドするために、GPT-2 デコーダーを条件付ける臨床的に強化されたテキスト プロンプトと信号表現を融合します。さらに、モデルが臨床的に意味のある所見に焦点を当てるのに役立つ ECG 固有のレポート前処理戦略を導入します。最後に、語彙メトリクスはレポートの品質を過小評価または過大評価する可能性があるため、生成されたレポートと参照レポートから抽出された診断ラベル間の一致を測定する BERT ベースの ECG 固有のメトリクスである F1-ECGBERT を提案します。 PTB-XL のドメイン内実験と MIMIC-IV-ECG のクロスドメイン評価では、ECG-LENS が常に最先端の方法を上回り、最強のベースラインに対して METEOR、ROUGE-L、および F1-ECGBERT でそれぞれ 4.0%、6.3%、および 11.5% の絶対利得を示しました。

原文 (English)

ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation

Electrocardiography (ECG) is one of the most widely used non-invasive tools for diagnosing cardiovascular disease, but transforming multi-lead ECG recordings into reliable clinical reports remains challenging. Automating ECG report generation could reduce clinicians' interpretive workload, improve diagnostic efficiency, and expand access to cardiac assessment in underserved communities. Unlike image-based report-generation tasks, ECG interpretation requires the analysis of subtle temporal morphologies, followed by coherent diagnostic reasoning expressed in dense clinical terminology. Existing systems predominantly focus on classification, while current report-generation methods often produce outputs that remain inadequate for practical clinical use. To address these challenges, we propose ECG-LENS, an end-to-end ECG report-generation framework that jointly integrates multi-lead signal modeling, diagnosis-aware representations, and clinically grounded text generation. ECG-LENS combines lead-wise encoders that preserve localized waveform morphology with a global encoder that captures inter-lead dependencies. To guide report generation, we fuse signal representations with clinically enriched textual prompts that condition a GPT-2 decoder. We further introduce an ECG-specific report-preprocessing strategy that helps the model focus on clinically meaningful findings. Finally, because lexical metrics may under- or overestimate report quality, we propose F1-ECGBERT, a BERT-based, ECG-specific metric that measures agreement between diagnostic labels extracted from generated and reference reports. In-domain experiments on PTB-XL and cross-domain evaluation on MIMIC-IV-ECG show that ECG-LENS consistently outperforms state-of-the-art methods, with absolute gains of 4.0%, 6.3%, and 11.5% in METEOR, ROUGE-L, and F1-ECGBERT, respectively, over the strongest baselines.

2026-08-07 13:00 JSTarXiv cs.AI画像/動画生成ロボティクスビジネス/資金調達研究/論文

GAUGE: シミュレーション エンジンとビデオ ワールド モデルの物理的忠実度に関する測定に基づいたベンチマーク

物理エンジンは、身体化された知能の大規模なトレーニングと評価を容易にする一方、生成ビデオ世界モデルは、将来の状態と相互作用の暗黙的なシミュレーターとして登場しています。しかし、物理的忠実度の既存の評価は単独で行われることが多く、知覚的な類似性や人間の判断に大きく依存しており、どの物理的原理やパラメータが違反されているかについての洞察は限られています。数値シミュレーターと生成ビデオ ワールド モデルが現実世界の物理学をどのように再現するか、または現実世界から逸脱するかを共同で評価するための、現実世界に基づいた診断ベンチマークである GAUGE を紹介します。これは、剛体、フレキシブル ケーブル、テキスタイル、および体積変形可能なオブジェクトをカバーする 22 の制御されたタスク ファミリで構成されています。これらのタスクは、現実世界の軌道に基づいて、調整された物理メタデータ、不確実性の注釈、タスク固有の観測可能量と組み合わせて、衝突、摩擦、運動量伝達、振動、自己接触、さまざまな材料と条件にわたる変形などの基本的な物理プロセスをカバーします。一般化された軌道誤差を使用して 14 のタスク ファミリで Isaac Sim、Genesis、および Newton のベンチマークを実行し、物理法則の一貫性と推論されたパラメーターの時間的安定性をテストすることにより、5 つの剛体タスクで 6 つの画像からビデオへのモデルを評価します。私たちの結果は、均一に忠実な物理エンジンがなく、衝撃的接触、繊維の素早い動き、体積変形で最も大きな差異が生じることを明らかにしました。さらに、ビデオ ワールド モデルは、誤った加速度、運動量伝達、振動タイミングを回復しながら、予想される方程式形式の軌道を生成できることもわかりました。 GAUGE は、より物理的に忠実なシミュレーターと身体化された知性の世界モデルを開発するための基礎を築きます。

原文 (English)

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions. However, existing evaluations of physical fidelity are often conducted in isolation and rely heavily on perceptual similarity or human judgments, providing limited insight into which physical principles or parameters are violated. We introduce GAUGE, a real-world-grounded diagnostic benchmark for jointly evaluating how numerical simulators and generative video world models reproduce or deviate from real-world physics. It comprises 22 controlled task families covering rigid bodies, flexible cables, textiles, and volumetric deformable objects. Grounded in real-world trajectories and paired with calibrated physical metadata, uncertainty annotations, and task-specific observables, these tasks cover fundamental physical processes including collision, friction, momentum transfer, oscillation, self-contact, and deformation across diverse materials and conditions. We benchmark Isaac Sim, Genesis, and Newton on 14 task families using generalized trajectory errors, and evaluate 6 image-to-video models on 5 rigid-body tasks by testing physical-law consistency and the temporal stability of inferred parameters. Our results reveal no uniformly faithful physics engine, with the largest discrepancies arising in impulsive contact, rapid textile motion, and volumetric deformation. We further find that video world models can produce trajectories with the expected equation form while recovering incorrect accelerations, momentum transfer, and oscillation timing. GAUGE lays the groundwork for developing more physically faithful simulators and world models for embodied intelligence.

2026-08-07 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

大規模言語モデルでの投資ロジックの評価: パーソナライズされた金融エージェントに向けた現実世界のベンチマーク

投資能力は本質的に個別化されています。同じ市場証拠が、異なる目標、視野、ポートフォリオ、リスク境界を持つ投資家にとって異なる行動を正当化する可能性があります。しかし、財務 LLM は、静的な質問応答または最終的な損益によって評価されます。前者は主体性を省略します。後者では、利益をもたらす行動が根拠に基づいたものなのか、プロファイルに一貫性のあるものなのか、あるいは単に幸運だったのかを明らかにすることはできません。私たちは、コミュニティが結果的な要因に対して間違った物差しを使用しているかどうかを尋ねます。 \textsc{InvestLogicBench} は、151 人の現実世界の投資家からの 201,247 件の文書化された意思決定を含むプロセスネイティブのベンチマークです。各エピソードは \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} トレースをインスタンス化します: 投資家 \textit{プロフィール}、観察可能な市場 \textit{イベント}、投資 \textit{推論}、実行可能ファイル \textit{決定}、および遅延 \textit{結果}。このリリースには、プロファイルの構築、ポイントインタイムのイベント バインディング、構造化ロジック、ホライズン、結果、事後分析が含まれており、理解、プロファイル条件付き生成、エンドツーエンドの再生がサポートされています。 4 つの主要な LLM では、論理的妥当性は 4/5 近くを維持していますが、イベントグラウンディングはわずか 0.8 ~ 2.8/5 です。返品とプロセスの品質も一致しません。これらの結果は、結果のみの評価が隠れている、洗練されているものの根拠が弱い推論を明らかにします。さらに、P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O は、バージョン付きプロファイル、時間的来歴、検査可能な検索、意思決定台帳、および再生可能な結果を​​必要とするデータ システム インターフェイスであるべきだと主張します。財務は、より広範なクラスの個別化された結果的なエージェントに対するストレス テストです。

原文 (English)

Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents

Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce \textsc{InvestLogicBench}, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a \textbf{P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O} trace: investor \textit{Profile}, observable market \textit{Events}, investment \textit{Reasoning}, executable \textit{Decision}, and delayed \textit{Outcome}. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8--2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that P$\rightarrow$E$\rightarrow$R$\rightarrow$D$\rightarrow$O should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents.

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

生成 AI を使用したスキーマに基づく階層情報の抽出と意味評価

私たちは、生成 AI を使用して非構造化テキスト文書から複雑な構造化情報を抽出し、抽出された情報をゴールドスタンダードに照らして自動的にセマンティック評価するためのスキーマベースのフレームワークを紹介します。このスキーマは、ドメイン知識をエンコードする情報モデルとして機能し、変数カーディナリティの属性を備えた階層的で入れ子になった情報の抽出と、その後の結果の評価のための統一的で体系的かつ一貫したフレームワークを提供します。ドキュメントからの情報抽出は、ゼロショット モードでのモデルへの 1 回の呼び出しで実行されます。評価ステップでは、パスベースのセマンティック マッチング アルゴリズムを導入して、抽出結果内のネストされた変数カーディナリティ属性をゴールド スタンダードの属性と一致させます。当社では、抽出された属性値とゴールドスタンダード値の意味的な比較に生成 AI を使用し、ドメイン固有の考慮事項に従って、比較の結果を完全一致、意味論的、有用、または不一致として分類するためのルーブリックを導入しています。生成 AI モデル Claude Opus 3 を使用して、保健技術評価機関 NICE が発行した文書から、$>$90\% の F1 スコアを持つ 14 属性のうち 12 属性を抽出することができました。文書から属性を抽出するのに必要な時間は、人間のドメイン専門家が要する時間より $\sim$30 分の 1 でした。さらに、さまざまな生成 AI モデルにわたるこのフレームワークの汎用性と、さまざまな HTA 組織や言語にわたる移行可能性を実証します。

原文 (English)

Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard. The schema, serving as an information model encoding domain knowledge, provides a unified, systematic, and consistent framework for extraction of hierarchical, nested information, with attributes of variable cardinality, and subsequent evaluation of the results. Information extraction from a document is performed in a single call to the model, in zero-shot mode. In the evaluation step, we introduce a path-based semantic matching algorithm to align the nested, variable-cardinality attributes in the extracted results with those in the gold standard. We use generative AI for semantic comparison of the extracted and gold standard values of an attribute, and introduce a rubric to classify the result of the comparison, according to domain-specific considerations, as an exact, semantic, useful, or non-match. We were able to extract 12 out of 14 attributes with an F1 score of $>$90\% from documents published by the health technology assessment organisation NICE, using the generative AI model Claude Opus 3. The time needed to extract the attributes from a document was $\sim$30 times lower than the time taken by a human domain expert. We further demonstrate generalisability of this framework across different generative AI models and transferability across different HTA organisations and languages.

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

MicroEvo: 効率的なマイクロアーキテクチャ設計空間探索のための知識に基づく LLM サンプリング

マイクロアーキテクチャの設計空間の探索は、広大な探索空間と高価な PPA 評価に悩まされ、設計の意思決定に使用できるシミュレーション予算はわずかしかありません。既存の方法は、マイクロアーキテクチャの依存関係を考慮せずにブラインド検索を実行し、反復検索から効果的に学習できず、無駄な評価と弱いパレート収束につながります。この論文では、多目的マイクロアーキテクチャの最適化のために既製の LLM とモンテカルロ ツリー検索 (MCTS) を結合する知識ガイド型フレームワークである MicroEvo を提案します。 MicroEvo は、LLM 主導の進化演算子、パレート寄与と多様性のバランスを取るパレート認識ツリー ポリシー、最適化の洞察を抽出して再利用するアクティブな知識蓄積メカニズム、オンラインでの検索動作を適応させる状態認識ディレクティブを組み合わせています。実験の結果、MicroEvo は NSGA-II と比較してパレート フロントの品質を最大 36.2% 向上させ、10.6 倍の高い検索効率を達成し、複雑な産業規模のコアに対する強力な拡張性も実証したことが示されています。コード リポジトリは、https://github.com/GEAR-SEU/MicroEvo-ICCAD-26 から入手できます。

原文 (English)

MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration

Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulation budget for design decision-making. Existing methods perform blind search without considering microarchitectural dependencies and fail to learn from the iterative search effectively, leading to wasted evaluations and weak Pareto convergence. In this paper, we propose MicroEvo, a knowledge-guided framework that couples off-the-shelf LLMs with Monte Carlo Tree Search (MCTS) for multi-objective microarchitecture optimization. MicroEvo combines LLM-driven evolutionary operators, a Pareto-aware tree policy that balances Pareto contribution and diversity, an active knowledge accumulation mechanism that extracts and reuses optimization insights, and state-aware directives that adapt the search behavior online. Experiments show that MicroEvo improves Pareto-front quality by up to 36.2% over NSGA-II and achieves 10.6x higher search efficiency, and also demonstrates strong scalability to a complex industrial-scale core. The code repository is available at: https://github.com/GEAR-SEU/MicroEvo-ICCAD-26.

2026-08-07 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

静的データと進化するデータの説明方法を評価する際の課題

この論文では、不十分な評価に関する説明可能な人工知能 (XAI) の限界について説明します。これらは、偏見の検出と概念の学習のための DetoxAI 画像認識システムを通じて示されています。次に、画像分類を説明する方法について人間に基づいて評価した例を示します。この論文では、概念ドリフトを伴う進化するデータ ストリームに説明を適応させる方法をさらに検討します。この問題に対して反事実を適応させた経験について説明します。最後に、それは、データ、モデル、説明の共進化を追跡するという課題に関連しています。\footnote{この論文は、J.Nalepa (ed) Explainable AI in Space の出版物として受理されました。 IJCAI-ECAI 2026 ブレーメンでの EASi 2026 ワークショップの議事録、Springer CCIS vol 3107 (2016)。

原文 (English)

Challenges in Evaluating Explanation Methods for Static and Evolving Data

This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through the DetoxAI image recognition system for bias detection and concept unlearning. Then, an example of a human-grounded evaluation of methods for explaining image classification is presented. The paper further explores methods for adapting explanations to evolving data streams with concept drift. Experiences with adapting counterfactuals for this problem are discussed. Finally it is related to the challenges of tracking the co-evolution of data, models, and explanations.\footnote{This paper has been accepted for a publication in J.Nalepa (ed) Explainable AI in Space. Proceedings of EASi 2026 Workshop at IJCAI-ECAI 2026 Bremen, Springer CCIS vol 3107 (2016).}

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation

Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideol…

2026-08-07 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study

Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet th…

2026-08-07 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

The ethics of artificial intelligence in the life sciences: Universality, cultural diversity and an architecture of care

The life sciences and health research have started to benefit from artificial intelligence, which raises ethical concerns that are real but…

2026-08-07 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection

Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance ofte…

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tr…

2026-08-07 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case

Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raise…

2026-08-07 13:00 JSTarXiv cs.AIビジネス/資金調達

Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping

Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deploy…

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readi…

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or ex…

2026-08-07 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics

We introduce BioAgent Bench, an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatic…

2026-08-07 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics

The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task. With the increasing interes…

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成エージェントビジネス/資金調達

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reas…

2026-08-07 13:00 JSTarXiv cs.AIビジネス/資金調達

DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization

Bayesian optimization (BO) relies on a surrogate model and an acquisition function, yet the most suitable choices vary across tasks and opt…

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation

Large language models (LLMs) have shown great promise in generating structured diagrams from natural language descriptions, particularly Me…

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

BenHalluEval: ベンガル語の大規模言語モデル用のマルチタスク幻覚評価フレームワーク

ベンガル語は世界で 6 番目に話されている言語であるにもかかわらず、ベンガル語の大規模言語モデル (LLM) で幻覚を体系的に評価した先行研究はありません。 BenHalluEval は、生成的質問応答 (GQA)、バングラ語と英語のコード混合 QA、要約、および推論の 4 つのタスクをカバーするベンガル語用のきめ細かい幻覚評価フレームワークです。既存の 3 つのベンガル語データセットから抽出された、12 のタスク固有の幻覚タイプにわたって GPT-5.4 を使用して 12,000 の幻覚候補を構築し、グラウンドトゥルース インスタンスの偽陽性率 (トラック A) と幻覚候補の幻覚検出率 (トラック B) を独立して測定するデュアル トラック プロトコルの下で、推論指向、多言語、ベンガル語中心のカテゴリにわたる 7 つの LLM を評価します。両方の故障モードに共同でペナルティを課し、均一な応答バイアスによるスコアのインフレを防ぐために、モデルとタスク全体で 7.72% から 55.42% の範囲のデュアルトラック キャリブレーション メトリクスである BenHalluScore を提案します。これは、幻覚キャリブレーションの大幅な変動を明らかにします。緩和戦略として適用される思考連鎖プロンプトは、幻覚差別を一貫して改善することなく、反応分布を変化させます。 BenHalluEval は、ベンガル語専用の幻覚ベンチマークを初めて確立し、リソースの少ない言語設定に対する単一トラックおよびプロンプトのみの評価アプローチが不適切であることを強調しています。データセットとコードは https://anonymous.4open.science/r/BanglaHalluEval-EB77 で入手できます。

原文 (English)

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates using GPT-5.4 across twelve task-specific hallucination types, drawn from three existing Bengali datasets, and evaluate seven LLMs spanning reasoning-oriented, multilingual, and Bengali-centric categories under a dual-track protocol that independently measures false-positive rate on ground-truth instances (Track A) and hallucination detection rate on hallucinated candidates (Track B). To jointly penalise both failure modes and prevent inflated scores from uniform response bias, we propose BenHalluScore, a dual-track calibration metric that ranges from 7.72% to 55.42% across models and tasks, revealing substantial variation in hallucination calibration. Chain-of-thought prompting, applied as a mitigation strategy, shifts response distributions without consistently improving hallucination discrimination. BenHalluEval establishes the first dedicated hallucination benchmark for Bengali and highlights the inadequacy of single-track and prompting-only evaluation approaches for low-resource language settings. The dataset and code are available at https://anonymous.4open.science/r/BanglaHalluEval-EB77.

2026-08-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

OpenAI Privacy Filter: A Cross-Lingual, Cross-Domain PII Evaluation Across 32 Benchmarks

We present what is, to our knowledge, the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter model that convert…

2026-08-07 10:50 JSTITmedia AI+ビジネス/資金調達

投資の機を逃す「残念な会社がいっぱい」――ソフトバンクG投資の秘訣、“金庫番”が語る

投資会社として成果を積み上げてきたソフトバンクG。その投資方針と成功のポイントについて、同社の“金庫番”こと後藤芳光CFOが語った。

2026-08-07 07:00 JSTITmedia AI+ビジネス/資金調達

銀行なら3カ月→AIは1カ月 10万件のデータで「数千万円」を引き出した“データドリブン資金調達術”

黒字化目前のZehitomoは、手元資金を確保すべく新たな資金調達手段を模索していた。だが、銀行融資は最低3カ月を要し、株式調達は希薄化のリスクを伴う。この壁を打ち破ったのが、AIを活用した「データ駆動型融資」だ。同社が提出した10万件の入金データをAIが解析し、わずか1カ月で…

2026-08-07 02:00 JSTTechCrunch AIビジネス/資金調達

Naïve raises $28.5M to automate the grunt work of setting up and running a company

Taking vibe-coding a step further, Naïve claims its infra can automate most of the work in setting up and running a business.

2026-08-06 22:00 JSTTechCrunch AIビジネス/資金調達

Ex-Spotify employees raise $10M to bring the AI behind its recommendations to e-commerce

The startup's platform predicts which product a shopper wants next, learns their general taste, and fine-tunes continuously based on what t…

2026-08-06 21:00 JSTTechCrunch AIビジネス/資金調達

Omilia raises $67M to scale its customer support platform

The Series B is the company's second fundraise since it last raised capital in 2020. In that time, it has increased its ARR by 10x to $60 m…

2026-08-06 19:30 JSTITmedia AI+LLM/生成AIハードウェア/半導体ビジネス/資金調達

ソフトバンクG、投資利益1.8兆円を支えた「OpenAIではない“あの半導体メーカー”」の正体

ソフトバンクGは、第1四半期の投資利益が1兆8594億円だったと発表した。投資利益を押し上げたのは、OpenAIでもArmでもない。歴史的な経営難に陥っていた“あの半導体メーカー”だった。

2026-08-06 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

MatrAIx: 83 億のペルソナ エージェントによる世界のシミュレーション

AI システムやデジタル製品を人間が評価するのはコストがかかり、時間がかかり、拡張するのが困難です。オフライン評価はよりスケーラブルですが、多くの場合、人間の多様性とインタラクティブな行動が抽象化されます。そこで、異種ユーザーによる AI システムやデジタル製品をテストするための人口規模のシミュレート ユーザー評価インフラストラクチャである MatrAIx を紹介します。 MatrAIx には 3 つのコア コンポーネントがあります。 まず、ペルソナ 8B には、1,290 のカテゴリ ディメンションで表される 83 億のペルソナ レコードが含まれています。レコードは、相関関係のある属性を保持する依存関係グラフからサンプリングされるか、人間が作成したプロファイルから派生します。当社は、599,847 件の人間に基づいたレコードと 400,000 件の合成レコードで構成される、約 100 万人のペルソナの高品質フィルター処理されたコアセットをリリースします。 2 番目に、MatrAIx プレイグラウンドは、多様なユーザーがデジタル製品を評価し、操作するための 4 つの環境 (アンケート、AI チャットボット、Web、アプリ) を提供します。 3 番目に、MatrAIx は、コマース、ソフトウェア、財務、ヘルスケアを含む 25 以上のドメインにわたる 1,010 のアプリケーション タスクを提供します。私たちは 8 つの代表的なタスクにわたって 18,189 件の評価トライアルを実施しました。ペルソナ エージェントは、Claude Opus 4.8、GPT 5.5、および Claude Haiku 4.5 の 3 つの LLM を利用していました。結果として得られるフィードバックは、値上げ後のためらい、AI アシスタントが失敗した後の続行意欲、待ち時間の許容度など、ペルソナの背景によって意思決定や好みがどのように異なるかを捉えています。私たちは 2 つの主要な検証研究を実施しました。まず、400 件の試験を対象とした対照研究で、10 の行動特性と 4 つすべての環境にわたるペルソナの遵守度を評価しました。宣言された行動は、366 件の試験 (91.5%) で発現または正しく抑制されました。次に、人間と LLM の審査員が、人間に基づいたペルソナの抽出品質を評価しました。全体として、MatrAIx は、さまざまなシミュレートされた人間のユーザーを使用して AI システムとデジタル製品を評価するためのエンドツーエンドのインフラストラクチャを提供します。

原文 (English)

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

2026-08-06 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills

Agent skills are increasingly optimized by automated feedback loops, producing long structured artifacts whose internal value remains uncle…

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Canary ツールを使用した LLM エージェントでのツール選択推論の診断

エージェントの評価では、モデルが間違ったツールを選択したことがわかりますが、その理由はほとんどわかりません。カナリア ツールを紹介します。これは、エージェントのモデル コンテキスト プロトコル (MCP) ツール セットに組み込まれた診断プローブ ツールであり、それぞれが 1 つの特定のツール選択の弱点を調査するように設計されています。 6 つのタイプの分類法 (セマンティック デコイ、パラメーター トラップ、ケイパビリティ ミラージュ、前提条件ブラインドネス、時間的デコイ、および粒度トラップ) は、単一の「間違ったツール」の結果を、モデルがツールについてどのように推論するかについての多次元プロファイルに変換します。私たちは、3 つの機能層にまたがる 8 つのモデル (6 つのホスト型と 2 つの 8B オープンウェイト) を、3 つのカナリア密度条件と 3 つのシード (8,640 実行) にわたる 120 のタスク、および 2,880 実行のサトルティ アブレーションで評価しました。タスクの成功は、プロバイダーに依存しない審査員によって採点され、2 番目の独立した審査員によって裏付けられます (コーエンのカッパ = 0.75)。 3 つの調査結果を報告します。まず、モデルの能力が向上するにつれて、感受性は急激に低下します。タスクごとのカナリア感受性率 (CSR) は、モデル全体で約 36 倍の範囲にあり、Claude Opus 4.8 が最も低く、Llama 3.1 8B が最も高くなります。第 2 に、機能層だけでは安全性を予測できません。最も影響を受けやすいホスト型モデルは中間層であり、プロバイダー内では安価なモデルの方が安全である可能性があります。第三に、分類は能力によって階層化されています。能力ミラージュは最も確実にフロンティア モデルをトラップしますが、他のタイプは強力なモデルではほとんど不活性ですが、小規模なオープン モデルでは発火するため、弱いというよりも能力によって区別されます。各カナリアのギブアウェイフレーズを和らげると、フロンティアCSRは本質的に変化せず、プローブがフレーズ発見ではなく推論を測定する証拠です。感受性はタスクの失敗も予測しますが (スピアマン rho = -0.34)、最も堅牢なモデルはカナリア圧力によって大幅に劣化しません。フレームワーク、カナリア スキーマ、タスク、ログをリリースします。

原文 (English)

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness. A six-type taxonomy (semantic decoys, parameter traps, capability mirages, prerequisite blindness, temporal decoys, and granularity traps) turns a single "wrong tool" outcome into a multi-dimensional profile of how a model reasons about tools. We evaluate eight models -- six hosted and two 8B open-weight -- spanning three capability tiers, on 120 tasks across three canary-density conditions and three seeds (8,640 runs), plus a 2,880-run subtlety ablation. Task success is graded by a provider-independent judge, corroborated by a second independent judge (Cohen's kappa = 0.75). We report three findings. First, susceptibility drops sharply as models get more capable: the per-task canary susceptibility rate (CSR) ranges about 36x across models, lowest for Claude Opus 4.8 and highest for Llama 3.1 8B. Second, capability tier alone does not predict safety: the most susceptible hosted model is mid-tier, and within a provider the cheaper model can be the safer one. Third, the taxonomy is capability-stratified: capability mirages most reliably trap frontier models, while the other types are largely inert on strong models but fire on small open models, so they discriminate by capability rather than being weak. Softening each canary's give-away phrase leaves frontier CSR essentially unchanged, evidence that the probes measure reasoning, not phrase-spotting. Susceptibility also predicts task failure (Spearman rho = -0.34), while the most robust models are not significantly degraded by canary pressure. We release the framework, canary schemas, tasks, and logs.

2026-08-06 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

ContextWeave: A Real-World Workflow Benchmark

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce…

2026-08-06 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

共有ロールアウトが防御運転評価で失敗した場合: NAVSIM スコアベースの監査

防御運転スコアは、周囲のアクターを観察するポリシーとそうでないポリシーの区別を維持する場合にのみ役立ちます。再シミュレーション ベンチマークでは、ログに記録された人間のリファレンスがコンプライアンス チャネルを通過できなかった場合にエージェントがクレジットを受け取る、リファレンス条件付きの寛容性を使用する場合があります。エージェントと参照が不安定なロールアウト変換を共有する場合、このルールにより、共有参照の失敗が広範なコンプライアンス クレジットに伝播される可能性があります。 NAVSIM v2.2 のオリジナル シーンの単一ステージ スコアリングでこのリスクを監査します。監査された数値バックエンドの影響を受けた文書化されたスタック条件の下では、ルート ブラインド Ignore-All プローブとルート認識アクター ブラインド プローブは、完全な 12,146 トークンの navtest スプリットに対して人間のリプレイと PDM-Closed よりも上位にランクされます。公開仕様に従った新規インストールでは、固定の 32 トークン診断セットでロールアウトの相違が再現されます。同一ソース依存関係スタック制御と正確な入力診断により、共有速度再調整における依存関係に依存する数値動作が分離されます。 450 トークンのコントロール プールでは、ソルバーのみを置き換えることでロールアウトの発散がなくなり、許容範囲を有効にしたままブラインド ラスト順序が復元されます。したがって、数値の不安定性が直接のトリガーとなります。参照条件付きの寛容は、結果として生じる共有参照の失敗をコンプライアンスの信用に伝播します。当社は、スコアベースとスタックの開示、ブラインドプローブ、上書きレポート、防御運転主張にスコアを使用する前に展開する安定性テストを必要とする監査プロトコルに貢献します。

原文 (English)

When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit

Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not. Re-simulation benchmarks may use reference-conditioned forgiveness, under which an agent receives credit when the logged human reference fails a compliance channel. When agent and reference share an unstable rollout transformation, this rule can propagate shared reference failures into broad compliance credit. We audit this risk in NAVSIM v2.2 original scene single-stage scoring. Under the affected documented-stack condition on the audited numerical backend, the route-blind Ignore-All probe and a route-aware actor-blind probe outrank human replay and PDM-Closed over the complete 12,146-token navtest split. A fresh installation following the public specification reproduces rollout divergence on a fixed 32-token diagnostic set. A same-source dependency stack control and an exact-input diagnostic isolate dependency-sensitive numerical behavior in the shared velocity refit. On a 450-token control pool, replacing only the solver eliminates rollout divergence and restores blind-last ordering while keeping forgiveness enabled. Thus, the numerical instability is the direct trigger. Reference-conditioned forgiveness propagates the resulting shared reference failures into compliance credit. We contribute an audit protocol requiring score basis and stack disclosure, blind probes, overwrite reporting, and rollout stability tests before using such scores for defensive driving claims.

2026-08-06 13:00 JSTarXiv cs.AIビジネス/資金調達

CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers

The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns related to hardware faults. Alg…

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true witho…

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

サイバーフィジカルシステムにおけるLLMエージェントの計画戦略の戦略的評価

LLM 計画エージェントの評価では、タスクが成功するか、宣言された計画が実行されるかが主に問われます。戦略的なサイバー物理システムでは、自律的な参加者が応答し、物理学が結果を制約した後も、計画アーキテクチャが適切なままであるかどうかがより強力な問題となります。計画に起因する制御軌道、つまり実行アーキテクチャが他のエージェントや物理プロセスに作用する順序付けされた計画操作と指令を中心に構築された、制御された物理に基づいたベンチマークを導入します。 40 の異種プロシューマーと独立してシミュレートされたラジアル フィーダーを備えたスマート グリッド デマンド レスポンス システムに、事前定義された順次的、階層的、および検索エグゼキュータが実装されています。 LLM は、型指定されたポリシー宣言と短いオペレーター メッセージに制限されていますが、スケジュール構築、プロシューマー ダイナミクス、および電力の流れは明示的なコードのままです。このプロトコルは、ペアになった強制モードの反事実、一般的なランダムな応答の抽出、およびイベントレベルの期限の実現可能性を使用します。 3 つのプロパティが続きます。アーキテクチャは結果を大きく変えます。強制検索は 5 つのベースライン シードすべてにおける神託です。実行忠実度にはモード一致以上のものが必要です。目的置換は電圧不足を 2.68 倍に増加させながら 1.0 で一致を保持します。 144 のシナリオ、576 エピソードのバンクには、4 つのアーキテクチャのうち 3 つからの実行可能なオラクルがあります。事前に指定された応力保持リッジの平均リグレスは 90.7 (95% 間隔 [73.8, 108.6]) で、固定シーケンシャルでは検出可能な値はありません。品質予測の前に既知の期限の実現可能性を適用すると、後悔は 29.0 に減少し、固定シーケンシャルよりも 61.1 改善されます。あらゆる実現可能なアブレーションは固定検索に勝るものではなく、残りの課題は実現可能な品質の選択に限定されます。 5 つのモデル拡張により、ストレス条件付き宣言子、状態ブラインド宣言子、および不変宣言子が分離されます。レイテンシーテールは、実際の実現可能性が確率的に扱われる必要があることを示しています。

原文 (English)

Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems

Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems, a stronger question is whether the planning architecture remains appropriate after autonomous participants respond and physics constrains the outcome. We introduce a controlled, physics-grounded benchmark built around planning-induced control trajectories: the ordered planning operations and directives through which an execution architecture acts on other agents and the physical process. It implements predefined, sequential, hierarchical, and search executors in a smart-grid demand-response system with 40 heterogeneous prosumers and an independently simulated radial feeder. The LLM is bounded to typed policy declaration and short operator messages, while schedule construction, prosumer dynamics, and power flow remain explicit code. The protocol uses paired forced-mode counterfactuals, common random response draws, and event-level deadline feasibility. Three properties follow. Architecture materially changes outcomes: forced search is the oracle in all five baseline seeds. Execution fidelity needs more than mode agreement: objective substitution holds agreement at 1.0 while increasing voltage shortfall by 2.68x. A 144-scenario, 576-episode bank has feasible oracles from three of the four architectures. A prespecified stress-held-out ridge has mean regret 90.7 (95% interval [73.8, 108.6]) and no detectable value over fixed sequential; applying known deadline feasibility before quality prediction cuts regret to 29.0 and improves over fixed sequential by 61.1. An all-feasible ablation does not beat fixed search, localising the remaining challenge to within-feasible quality selection. A five-model extension separates stress-conditioned, state-blind, and invariant declarers; latency tails show that live feasibility should be treated probabilistically.

2026-08-06 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions…

2026-08-06 13:00 JSTarXiv cs.AIビジネス/資金調達

Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems

Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well…

2026-08-06 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory complianc…

2026-08-06 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

PhyAI: エッジでのリアルタイム物理 AI、クラウドでのスケーラブルなロールアウト

物理 AI ポリシーでは、モデルの評価、クラウド強化学習のロールアウト、エッジ GPU の提供、オンボード展開などのライフサイクル全体にわたって推論が必要です。これらの設定は同じチェックポイントとアクションのセマンティクスを共有しますが、多くの場合、別の推論プログラムに依存します。これらを統合するために、グラフの実行、カーネル、メモリ管理、並列サービスを共有しながら、アーキテクチャ固有のコンディショニング、ソルバー、キャッシュ、出力ロジックをモデル アダプターに保持する単一のランタイムを備えた物理 AI 推論エンジンである PhyAI を構築します。同じコードベースは、オンボード、エッジ、クラウドの展開全体で単一または複数の GPU 上でビジョン言語アクション (VLA) モデルとワールド アクション モデル (WAM) を実行します。 MiniCPM-Robot のリリース日にアダプター インターフェイスを使用して追加しました。 PhyAI は、pi0、pi0.5、GR00T N1.7、および MiniCPM-Robot の公式実装と比較して 1.40 倍から 4.65 倍の高速化を達成します。 Cosmos3-Nano-Policy-DROID では、8 つの H20 GPU (CFG=2、TP=4) でレイテンシが 2.46 秒から 1.18 秒に短縮され、2.08 倍の速度向上になります。特殊なランタイムはいくつかの構成で引き続き高速であるため、私たちの目標は、すべてのケースで最速の結果ではなく、競争力のあるレイテンシーを備えた 1 つのランタイムです。詳細なプロファイルにより、異なるモデルに異なる実行ポリシーが必要な理由が明らかになります。バッチ サイズ 1 の Hopper シリーズ GPU では、pi0.5 アクション エキスパートは FLOP の 8.8% を占めますが、レイテンシーの 57.2% を占めます。バッチ サイズ 32 では、そのシェアは 13.5% に低下し、スループットは約 100 サンプル/秒に達します。 Cosmos3 は世代主導のままで、バッチ サイズが 1 から 16 に増加してもスループットは 14.3% しか向上しません。さらに、推論制限制御と環境制限制御を区別する制御時間ルーフラインを導入します。 4 つの LIBERO スイートで測定された pi0.5 ポイントは環境に依存していますが、Cosmos3 は推論に依存したままです。コードとベンチマーク: https://github.com/mingti-org/phyai。

原文 (English)

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM-Robot on the day of its release. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot. On Cosmos3-Nano-Policy-DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation-dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control-time Roofline, which distinguishes inference-bound from environment-bound control; the measured pi0.5 points on four LIBERO suites are environment-bound while Cosmos3 stays inference-bound. Code and benchmarks: https://github.com/mingti-org/phyai.

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation

While LLMs have shown great success in financial tasks like stock prediction and question answering, their application in fully automating…

2026-08-06 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

MODEST: Multi-Optics Depth-of-Field Stereo Dataset

Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus debl…

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

VibeSearchBench: 実環境における長期的なプロアクティブ検索のベンチマーク

LLM ベースのエージェントは検索ベンチマークで高いスコアを獲得していますが、実際のユーザーは一貫して結果が満足できないと感じており、評価とエクスペリエンスのギャップが根強く残っていることが明らかになりました。このギャップは、既存のベンチマークが過剰に指定されたクエリ、シングルターン インタラクション、および固定スキーマ評価に依存しているためであると考えられますが、これらのいずれも、ユーザーとエージェントが協力してマルチターン対話を通じてあいまいな意図を洗練するという実際の検索動作を反映していません。私たちはこのパラダイムを VibeSearch と名付け、20 のドメインにわたって手動で精選された 200 のバイリンガル (中国語と英語) タスクで構成されるベンチマークである VibeSearchBench を導入します。このベンチマークは、VibeSearch-Pro (プロフェッショナル) サブセットと VibeSearch-Daily (日常生活) サブセットに分かれています。各タスクは、ユーザー ペルソナとスキーマフリーのグラウンド トゥルース ナレッジ グラフを組み合わせ、漸進的開示ユーザー シミュレーターとグラフ マッチング評価フレームワークを通じて評価されます。 ReAct フレームワークと OpenClaw エージェント ハーネスの両方で 7 つのフロンティア モデルのベンチマークを行います。結果は、すべてのモデルが依然として VibeSearch には実質的に不十分であることを示し (最高 F1: 30.30)、ロングコンテキスト推論、プロアクティブな意図の引き出し、および構造化された知識の構築における根本的な進歩の必要性を強調しています。

原文 (English)

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience gap. We attribute this gap to existing benchmarks' reliance on over-specified queries, single-turn interactions, and fixed-schema evaluation, none of which reflect real search behavior where users and agents collaboratively refine vague intent through multi-turn dialogue. We term this paradigm VibeSearch and introduce VibeSearchBench, a benchmark comprising 200 manually curated bilingual (Chinese and English) tasks across 20 domains, split into VibeSearch-Pro (professional) and VibeSearch-Daily (daily-life) subsets. Each task pairs a user persona with a schema-free ground-truth knowledge graph, and is evaluated through a progressive-disclosure user simulator and a graph-matching evaluation framework. We benchmark seven frontier models under both the ReAct framework and the OpenClaw agent harness. Results show that all models remain substantially inadequate for VibeSearch (best F1: 30.30), highlighting the need for fundamental advances in long-context reasoning, proactive intent elicitation, and structured knowledge construction.

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment

Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting among multiple candidates with an automat…

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate other…

2026-08-06 07:00 JSTITmedia AI+ビジネス/資金調達研究/論文

資金調達の「二極化」進む 勝ち抜く企業の条件は? 早大教授に聞く

「AIによる審査の高度化」と「資金調達の多様化」は必ずしも同じ話ではない──ベンチャーファイナンスの第一人者である早稲田大学ビジネススクール(経営管理研究科)の長谷川博和教授は、こう指摘する。

2026-08-05 20:00 JSTTechCrunch AIビジネス/資金調達

AI makes weather prediction better. Can WindBorne make it lucrative?

WindBorne Systems has raised a $37 million Series B round to scale its weather balloons and AI forecasts.

2026-08-05 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

マルチエージェント評価を使用したロールプレイング言語エージェントの敵対的ストレステスト

ロールプレイング言語エージェント (RPLA) は、医療支援、顧客サポート、教育など、敵対的な圧力の下で一貫したペルソナ、倫理的制約、行動の一貫性を維持することが重要となる、一か八かのアプリケーションに導入されることが増えています。既存の評価アプローチは、静的なベンチマークまたは分離された単一ターン プロンプトに依存しており、長期にわたるインタラクションを通じて発生する累積的な動作障害を捕捉できません。構造化された複数ターンの対話を通じて、敵対的ストレス テストを行う RPLA のためのモジュール式マルチエージェント プラットフォームを紹介します。このシステムは 3 つのエージェントを調整します。1 つは 6 つの進歩的な敵対的戦略を適用する戦略主導型の尋問エージェント、評価対象の RPLA を表すターゲット エージェント、および役割の忠実度、ドリフト、倫理的逸脱、および一貫性の次元にわたって行動をスコアリングする自動判定エージェントです。 3 つのペルソナと 3 つの LLM ファミリにわたる実験を通じて、複数戦略の敵対的評価により、単一戦略のテストでは見えない障害モードが明らかになり、全体のロバストネス スコアが平均 0.17 ~ 0.20 ポイント低下することを実証しました。クロスモデル検証により、Llama-3.3-70B、GPT-4o-mini、Claude-3.5-Haiku 全体で一貫した劣化パターンが確認され、権限チャレンジと感情操作が最も効果的な攻撃戦略として浮上しています。自動判定により、人間による強力な一致が達成されます ($r = 0.82$、Fleiss の $\kappa = 0.71$)。この作品は、AI の安全性と再現可能な RPLA ベンチマークをサポートするオープンソース プラットフォームとしてリリースされています。このフレームワークにより障害モードの体系的な発見が可能になりますが、私たちは敵対的テスト手法に関連する潜在的な倫理的リスクを認識し、AI の安全性を向上させるための責任ある使用を強調します。

原文 (English)

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and education, where maintaining consistent personas, ethical constraints, and behavioral coherence under adversarial pressure is critical. Existing evaluation approaches rely on static benchmarks or isolated single-turn prompts that fail to capture cumulative behavioral failures emerging over extended interactions. We present a modular multi-agent platform for adversarially stress-testing RPLAs through structured, multi-turn dialogue. The system coordinates three agents: a strategy-driven Interrogator Agent that applies six progressive adversarial strategies, a Target Agent representing the RPLA under evaluation, and an automated Judging Agent that scores behavior across role fidelity, drift, ethical deviation, and consistency dimensions. Through experiments across three personas and three LLM families, we demonstrate that multi-strategy adversarial evaluation reveals failure modes invisible to single-strategy testing, reducing overall robustness scores by 0.17--0.20 points on average. Cross-model validation confirms consistent degradation patterns across Llama-3.3-70B, GPT-4o-mini, and Claude-3.5-Haiku, with Authority Challenge and Emotional Manipulation emerging as the most effective attack strategies. Automated judging achieves strong human alignment ($r = 0.82$, Fleiss' $\kappa = 0.71$). This work is released as an open-source platform to support AI safety and reproducible RPLA benchmarking. While the framework enables systematic discovery of failure modes, we acknowledge potential ethical risks associated with adversarial testing methodologies and emphasize responsible usage for improving AI safety.

2026-08-05 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

PhyAI: エッジでのリアルタイム物理 AI、クラウドでのスケーラブルなロールアウト

物理 AI ポリシーでは、モデルの評価、クラウド強化学習のロールアウト、エッジ GPU の提供、オンボード展開などのライフサイクル全体にわたって推論が必要です。これらの設定は同じチェックポイントとアクションのセマンティクスを共有しますが、多くの場合、別の推論プログラムに依存します。これらを統合するために、グラフの実行、カーネル、メモリ管理、並列サービスを共有しながら、アーキテクチャ固有のコンディショニング、ソルバー、キャッシュ、出力ロジックをモデル アダプターに保持する単一のランタイムを備えた物理 AI 推論エンジンである PhyAI を構築します。同じコードベースは、オンボード、エッジ、クラウドの展開全体で単一または複数の GPU 上でビジョン言語アクション (VLA) モデルとワールド アクション モデル (WAM) を実行します。 MiniCPM-Robot のリリース日にアダプター インターフェイスを使用して追加しました。 PhyAI は、pi0、pi0.5、GR00T N1.7、および MiniCPM-Robot の公式実装と比較して 1.40 倍から 4.65 倍の高速化を達成します。 Cosmos3-Nano-Policy-DROID では、8 つの H20 GPU (CFG=2、TP=4) でレイテンシが 2.46 秒から 1.18 秒に短縮され、2.08 倍の速度向上になります。特殊なランタイムはいくつかの構成で引き続き高速であるため、私たちの目標は、すべてのケースで最速の結果ではなく、競争力のあるレイテンシーを備えた 1 つのランタイムです。詳細なプロファイルにより、異なるモデルに異なる実行ポリシーが必要な理由が明らかになります。バッチ サイズ 1 の Hopper シリーズ GPU では、pi0.5 アクション エキスパートは FLOP の 8.8% を占めますが、レイテンシーの 57.2% を占めます。バッチ サイズ 32 では、そのシェアは 13.5% に低下し、スループットは約 100 サンプル/秒に達します。 Cosmos3 は世代主導のままで、バッチ サイズが 1 から 16 に増加してもスループットは 14.3% しか向上しません。さらに、推論制限制御と環境制限制御を区別する制御時間ルーフラインを導入します。 4 つの LIBERO スイートで測定された pi0.5 ポイントは環境に依存していますが、Cosmos3 は推論に依存したままです。コードとベンチマーク: https://github.com/mingti-org/phyai。

原文 (English)

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often rely on separate inference programs. To unify them, we build PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services. The same codebase runs vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. We used the adapter interface to add MiniCPM-Robot on the day of its release. PhyAI achieves 1.40x-4.65x speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot. On Cosmos3-Nano-Policy-DROID it reduces latency from 2.46 to 1.18 s on eight H20 GPUs (CFG=2, TP=4), a 2.08x speedup. Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case. Detailed profiles reveal why different models need different execution policies: on a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of FLOPs but 57.2% of latency; at batch size 32 its share drops to 13.5% and throughput reaches about 100 samples/s. Cosmos3 remains generation-dominated and gains only 14.3% throughput as batch size increases from 1 to 16. We further introduce the control-time Roofline, which distinguishes inference-bound from environment-bound control; the measured pi0.5 points on four LIBERO suites are environment-bound while Cosmos3 stays inference-bound. Code and benchmarks: https://github.com/mingti-org/phyai.

2026-08-05 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

LiveEvalBench: Web 生成のオープンワールド評価に向けて

大規模な言語モデルは実行可能なフロントエンド プロジェクトを合成できるようになってきていますが、既存のベンチマークは依然として Web 生成を静的評価問題として扱っています。私たちは、フロントエンド アーティファクトには異なるパラダイムが必要であると主張します。フロントエンド アーティファクトは静的ではなくインタラクティブであり、多様でありながら同様に有効な実装を許容し、厳格なパイプラインが対応できるよりも速く進化します。これらのギャップに対処するために、Web 生成の評価をエージェント的、適応的、拡張可能なプロセスとして再定式化する自動フレームワークである LiveEvalBench を紹介します。 LiveEvalBench は、共同レビュー ワークフローとして評価をインスタンス化します。このワークフローでは、ビルド エンジニア、コード エンジニア、UI テスターが、デプロイメントやコード検査からブラウザベースの操作に至るまで、フロントエンド プロジェクトのライフサイクル全体にわたって証拠を共同で収集します。実装の多様性に対処するために、適応プロトコルは、モデル間の比較を可能にする共有ルーブリックを、各成果物に合わせた実装に基づいた基準と組み合わせます。このフレームワークは、パイプラインを再設計することなく、新しい評価者の役割と評価次元の段階的な統合をさらにサポートします。現実世界の多様な Web 生成シナリオにわたる実験では、LiveEvalBench が人間の専門家の判断と密接に一致し、フロンティア モデルの Web 生成機能についてのきめ細かい洞察が得られることが示されています。コードは https://github.com/wyysteelhead/LiveEvalBench で入手できます。

原文 (English)

LiveEvalBench: Toward Open-World Evaluation for Web Generation

Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are interactive rather than static, admit diverse yet equally valid implementations, and evolve faster than rigid pipelines can accommodate. To address these gaps, we present LiveEvalBench, an automated framework that reformulates web-generation evaluation as an agentic, adaptive, and extensible process. LiveEvalBench instantiates evaluation as a collaborative review workflow, in which a Build Engineer, a Code Engineer, and a UI Tester collectively gather evidence across the full lifecycle of a frontend project, from deployment and code inspection to browser-based interaction. To handle implementation diversity, an adaptive protocol couples shared rubrics for cross-model comparability with implementation-grounded criteria tailored to each artifact. The framework further supports incremental integration of new evaluator roles and assessment dimensions without pipeline redesign. Experiments across diverse real-world web-generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine-grained insights into frontier models' web generation capabilities. Code is available at https://github.com/wyysteelhead/LiveEvalBench

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation

Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks m…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation

Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pret…

2026-08-05 13:00 JSTarXiv cs.AIビジネス/資金調達

Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions

Objective speech-quality metrics are widely used to assess codec performance. However, for neural codecs, it is often unclear which metrics…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, acro…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation us…

2026-08-05 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments

Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems,…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation

Cloud misconfiguration remains a leading cause of security incidents, yet whether LLMs and SLMs can generate security-compliant Infrastruct…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation

Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances…

2026-08-05 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

EFX Allocation In (Multi)Hypergraphs

We study fair allocations of indivisible goods among agents with heterogeneous monotone valuations. As fair we consider the allocations tha…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this r…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

How Closely Do LLM Reviews Align with Human Peer Review?

Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether differen…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretr…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," how…

2026-08-05 13:00 JSTarXiv cs.AIビジネス/資金調達

Modeling Matches as Language: A Generative Transformer Approach for Counterfactual Player Valuation in Football

Evaluating football player transfers is challenging because player actions depend strongly on tactical systems, teammates, and match contex…

2026-08-05 13:00 JSTarXiv cs.AIビジネス/資金調達

An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness

Artificial Intelligence (AI) and Machine Learning (ML) models used in clinical settings are increasingly deployed to support clinical decis…

2026-08-05 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

どこが間違っていたのでしょうか?セマンティック状態追跡による Web エージェントのプロセス レベルの評価

Web エージェントは長い対話シーケンスを通じて動作しますが、既存のベンチマークは最終的な成功のみを評価し、すべてのプロセス情報を破棄し、改善に関するガイダンスをほとんど提供しません。この作業では、Web エージェントのプロセス レベルの分析を実行します。難易度を制御し、セマンティックな状態を自動的に追跡する 1,800 個のタスク インスタンスのベンチマークである WebStep を紹介します。各 Web サイトは、GUI とともに決定論的セマンティック MDP を公開します。エージェントはインターフェイス上で動作し、環境はバックグラウンドで高レベルの状態と遷移を記録するため、手動による注釈なしで詳細な分析が可能になります。セマンティックな軌跡に基づいて、プロセスのメトリクスが結果の評価では見えない違いを明らかにすることを最初に示します。つまり、成功率が 31 ~ 33% 以内にクラスター化されている 3 つのエージェントは、探索範囲と実行精度において乖離しています。次に、スキルごとに分解すると、これらの違いの性質が特徴づけられ、同じ Web サイト内に隠されている反対のスキルごとのランキングが明らかになります。たとえば、ハウジングでは、OpenAI CUA はコミット アクションで Qwen3.5 を 23.7% 上回っていますが、フィルタリングでは 15.6% 下回っており、ドメイン内であっても改善すべき具体的なスキルを特定します。分岐分析は、タスクを失う決定的なエラーをさらに特定し、このエラーが共有エラーではなくエージェント固有であることを示します。最後に、タスクが難しくなるにつれて、これらの差は広がります。簡単なタスクでは成功率は似ていますが、探索がより要求が厳しくなるにつれて、成功率は大きく異なります。当社のプロセスレベルの分析は、Web エージェントの評価に新たな道を開き、各エージェントのどこをどのように改善する必要があるかについて、きめ細かく実用的な洞察を提供します。

原文 (English)

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improvement. In this work, we conduct a process-level analysis of web agents. We introduce WebStep, a benchmark of 1,800 task instances with controlled difficulty and automatic semantic state tracking. Each website exposes a deterministic semantic MDP alongside the GUI: the agent operates on the interface, while the environment records high-level states and transitions in the background, enabling fine-grained analysis without manual annotation. Based on the semantic trajectory, we first show that process metrics reveal differences invisible to outcome evaluation: three agents whose success rates cluster within 34-37% diverge in exploration reach versus execution accuracy. Then, decomposing by skill characterizes the nature of these differences, exposing opposite per-skill rankings hidden within the same website: e.g., on Q&A, Claude CUA outperforms OpenAI CUA by 30% on navigation actions yet underperforms it by 6.7% on inspection, pinpointing a concrete skill to improve even within a domain. Bifurcation analysis further localizes the decisive error that loses the task and shows that this error is agent-specific rather than shared. Finally, these differences widen as tasks grow harder: success rate is similar on easy tasks but separates sharply as exploration becomes more demanding. Our process-level analysis opens a new avenue in web agent evaluation, providing fine-grained and actionable insight into where and how each agent should be improved. Project page: https://jiwanchung.github.io/webstep

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

SKILL-KD: Contrastive Skill Distillation for LLM Agents

Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition meth…

2026-08-05 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory complianc…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents

We discover that prompt-injection success and tool-execution success are separable safety properties: defenses that block injection do not…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworth…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and ans…

2026-08-05 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A co…

2026-08-05 13:00 JSTarXiv cs.AIビジネス/資金調達

Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems

Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, eff…

2026-08-05 04:00 JSTOpenAILLM/生成AIビジネス/資金調達

Third-party cyber evaluations involving OpenAI models

OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evalua…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

SciToolAgent-Evo: オープンワールドの科学ツールを取得するためのオントロジーを認識した自己進化エージェント

大規模言語モデル (LLM) エージェントは、特殊な計算ツールを編成して呼び出すために科学研究で採用されることが増えています。ただし、静的セマンティクスを持つ事前定義されたツール空間に依存しているため、ツールの要件、機能、境界が動的に進化するオープンワールドの科学ワークフローへの適用性が制限されます。この目的を達成するために、オープンワールドの科学ツールを取得するためのオントロジーを認識した自己進化エージェントである SciToolAgent-Evo を提案します。スキル、経験、オントロジー化されたツール グラフの進化する記憶によって駆動され、蓄積中に対照的な軌跡から一般化可能な知識を抽出します。一方、推論中にアクティブなリクエストを定式化し、LinUCB ベースのバンディット ゲートを利用して探索と活用の動的バランスをとります。新しいツールを取得すると、その科学的オントロジーがオンラインで完成し、既知のグラフにシームレスに統合されます。さらに、4 つの難易度にわたる 900 の現実的なタスクを含むベンチマークである OpenSciToolBench を紹介します。広範な評価により、SciToolAgent-Evo が最先端のパフォーマンスを実現し、その堅牢性と汎用性が検証されたことが示されています。

原文 (English)

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on predefined tool spaces with static semantics limits their applicability to open-world scientific workflows, where tool requirements, capabilities, and boundaries evolve dynamically. To this end, we propose SciToolAgent-Evo, an ontology-aware self-evolving agent for open-world scientific tool acquisition. Driven by an evolving memory of skills, experiences, and an ontologized tool graph, it distills generalizable knowledge from contrastive trajectories during accumulation, whereas during inference, it formulates active requests and utilizes a LinUCB-based bandit gate to dynamically balance exploration and exploitation. Once a novel tool is acquired, its scientific ontology is completed online for seamless integration into the known graph. Moreover, we introduce OpenSciToolBench, a benchmark containing 900 realistic tasks across four difficulty levels. Extensive evaluations show that SciToolAgent-Evo achieves state-of-the-art performance, validating its robustness and generalization.

2026-08-04 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention woul…

2026-08-04 13:00 JSTarXiv cs.AIビジネス/資金調達

A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation

Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input re…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft

With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unf…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ModelEquivBench: LLM で生成された最適化モデルのマルチリレーショナル評価の認証

大規模な言語モデルでは、自然言語から最適化モデルを生成するケースが増えていますが、既存の評価では、生成されたモデルとそのグラウンド トゥルースが、単一の同等/非同等判定または実行成功率、つまり独立してチェック可能ではなく、2 つの定式化が一致する複数の異なる意味に忠実でもないラベルに縮小されることがよくあります。我々は、ペアごとのセマンティック プロファイル E0 ~ E6 を報告する認証済みのマルチリレーショナル評価システムである ModelEquivBench を紹介します。モデルの構築と正確な取り込み (E0)、検証された表現の位置合わせ (E1)、同一空間と投影された実行可能集合関係 (E2、E3)、目的順序の等価性 (E4)、最適値の等価性 (E5)、およびオプティマイザー セットの等価性 (E6) です。決定された各エントリには、関係に適した、独立して再チェック可能な証拠が含まれます。つまり、E0 ~ E1 については再生可能なトレースまたは明示的なマップ、肯定的な E2 ~ E6 の結論については正確に合理的な証明書、およびサポートされた否定については明示的な証人です。不完全なマッピング検索、サポートされていない構造、およびリソース制限により、推測ではなく型指定された UNKNOWN または N/A の結果が生成されますが、満たされていない前提条件は ABSENT として報告されます。 ModelEquivBench を使用して、GPT-5.4、Claude Sonnet 4.6、および Qwen3.5-397B-A17B の 3 つのモデル スナップショットを、修復なしプロトコルの下で 173 の基本問題 (モデルあたり 346 セル) の同じ凍結コホートで評価すると、結果のプロファイルは、粗いベースラインでは表現されない区別を明らかにします。49、35、および 25 セルには、次の実行可能候補が含まれています。それにもかかわらず、少なくとも 1 つのサポートされている関係で否定的と証明され、E2 が検証済みマップの下でマップされた実現可能集合の等価性を証明するペアで 25、8、および 18 の構造的拒否が発生します。 3 つのモデルのスナップショットはプロファイルのさまざまな段階で失敗するため、単一の精度スコアに有意に削減することはできません。

原文 (English)

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree. We present ModelEquivBench, a certifying, multi-relational evaluation system that reports a per-pair semantic profile E0--E6: model construction and exact ingestion (E0), verified representation alignment (E1), same-space and projected feasible-set relations (E2, E3), objective-order equivalence (E4), optimal-value equality (E5), and optimizer-set equivalence (E6). Each decided entry carries relation-appropriate, independently re-checkable evidence: replayable traces or explicit maps for E0--E1, exact-rational certificates for positive E2--E6 conclusions, and explicit witnesses for supported negatives. Incomplete mapping search, unsupported structure, and resource limits produce typed UNKNOWN or N/A outcomes rather than guesses, while unmet prerequisites are reported as ABSENT. Using ModelEquivBench to evaluate three model snapshots--GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B--on the same frozen cohort of 173 base problems (346 cells per model) under a no-repair protocol, the resulting profiles expose distinctions that coarse baselines do not represent: 49, 35, and 25 cells contain executable candidates that are nevertheless certified negative on at least one supported relation, and 25, 8, and 18 structural rejections occur on pairs for which E2 certifies mapped feasible-set equality under a verified map. The three model snapshots fail at different stages of the profile and therefore cannot be meaningfully reduced to a single accuracy score.

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

フェデレーテッド事前トレーニングの評価: ダウンストリームの微調整と固有の評価の信頼性について

フェデレーテッド事前トレーニングは、基盤となるデータセットを一元化することなく、プライベート データまたは分散データで基礎モデルをトレーニングする方法を提供します。ただし、クライアントの参加やローカル データの可用性の違いにより、直接比較できる評価が困難になる可能性があるため、フェデレーテッド事前トレーニングの評価は依然として困難です。さらに、トレーニング前のテストの複雑さはトレーニング前の分布に関連付けられていますが、下流のベンチマークではタスク固有の適応が導入されており、トレーニング前に確立されたテストの複雑さが忠実に反映されていない可能性があります。この研究では、どの評価プロトコルがフェデレーテッド事前トレーニングの品質をより確実に反映するかを研究します。同一のクライアント データでトレーニングされた 16M パラメーターのトランスフォーマー モデルの集中型およびフェデレーテッド トレーニング済みモデルの制御されたセットを使用して、同じトレーニング前テストセットで確立された参照ランキングを保持しているかどうかによって評価プロトコルを評価します。フル、ヘッドのみ、データ削減バリアントを含む GLUE でのダウンストリーム微調整と、固有の評価信号としての GLUE テキストのネクストトークン予測を比較します。私たちの結果は、ダウンストリームの微調整ではトレーニング前のランキングを確実に保存しないのに対し、次のトークンの直接予測はトレーニング前のテストの複雑さと強い対応を示すことを示しています。これらの発見は、フェデレーテッド事前トレーニング モデルを比較する場合、下流の微調整だけでは誤解を招く可能性があり、元の事前トレーニング目標に近い評価信号はより大きな注目に値することを示唆しています。

原文 (English)

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in client participation and local data availability can make directly comparable evaluation difficult. Moreover, pre-training test perplexity is tied to the pre-training distribution, while downstream benchmarks introduce task-specific adaptation that may not faithfully reflect the test perplexity established during pre-training. In this work, we study which evaluation protocol more reliably reflects federated pre-training quality. Using a controlled set of centralized and federated-trained models of a 16M parameter transformer model trained on identical client data, we assess evaluation protocols by whether they preserve a reference ranking established on the same pre-training testset. We compare downstream fine-tuning on GLUE, including full, head-only, and reduced-data variants, with next-token prediction on GLUE text as an intrinsic evaluation signal. Our results show that downstream fine-tuning does not reliably preserve the pre-training ranking, whereas direct next-token prediction exhibits a strong correspondence with the pre-training test perplexity. These findings suggest that downstream fine-tuning alone can be misleading when comparing federated pre-trained models, and that evaluation signals closer to the original pre-training objective deserve greater attention.

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While pri…

2026-08-04 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities

As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to i…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

プロセスレベルの社会的影響評価のための認知世界モデル

社会的影響ダイアログは、内部の認知状態を変えることでユーザーの行動を変えます。評価の中心となる質問は、ユーザーの信念、欲望、意図、感情が会話の過程で測定可能なほど変化するかどうかであり、これは表面レベルのテキスト指標 (BLEU/ROUGE) や単一スコアの LLM 判定では捉えることができないプロセス指向の基準です。我々は \textbf{Cog}nitive \textbf{W}orld \textbf{M}odel \textbf{(CogWM)} を提案します。これは、マルチターン対話評価を「ユーザーが何を言ったか」から「ユーザーの内部認知状態がどのように進化したか」に再構成する LLM ベースのユーザー モデルです。CogWM は、BDI/E 認知状態とユーザー発話を共同で予測し、3 層を使用してユーザー シミュレーターと評価プラットフォームの両方として機能します。ターンレベルの忠実度、軌道レベルの状態ダイナミクス、タスクレベルの複合スコアリングをカバーする評価フレームワーク。 4 つの社会的影響シナリオにわたる 150,454 のユーザー ターン サンプルで \textbf{S}ummarize-\textbf{a}nd-\textbf{A}llocate \textbf{(SaA)} アノテーション パイプラインを介してトレーニングされた CogWM は、77.6\% の感情精度 (GPT-5.5 の 2.1$\time$) を達成しました。 3,600 件のマルチエージェント識別試験において、認知的影響力によって 6 つの営利エージェントを区別し、Llama-4-Scout が 1 位にランクされました (CTS +0.233)。 CogWM は、社会的影響対話の評価を最終的な判断からプロセスの追跡に移行します。コード\脚注{\scriptsize コード: https://github.com/lucianma05-create/CogWM} とモデル\脚注{モデル: https://www.modelscope.cn/models/LucianMa/CogWM-14B} をリリースしました。

原文 (English)

Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents

As LLM-based conversational agents advance toward increasingly open-ended and interaction-intensive scenarios, task completion alone provides an incomplete assessment of their effectiveness. The evolution of users' internal states, including beliefs, desires, intentions, and emotions (BDI/E), serves as an intermediate signal connecting agent behaviors with interaction outcomes and reflects how conversational strategies shape users during multi-turn interactions. However, existing evaluation paradigms primarily focus on surface-level responses or final outcomes, providing limited insight into the underlying cognitive processes. This limitation makes it difficult to diagnose why agents succeed or fail and to optimize their interaction strategies. To address this challenge, we propose Cognitive World Model (CogWM), an LLM-based cognitive user model that jointly models users' BDI/E states and corresponding responses, enabling explicit cognitive trajectory tracking. Trained on 150K user-turn samples with Qwen3-14B, CogWM achieves superior performance over existing user simulation baselines in both response fidelity and cognitive state understanding. Interactions with six state-of-the-art LLMs demonstrate that CogWM enables progressive comparison of agents through cognitive trajectories, revealing distinct agent patterns and complementary relationships between cognitive evolution and behavioral outcomes.

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabilities and alignment properties they are meant to represent remain uncertain. Synthesizing 373 primary studies published between 2018 and 2026, the survey organizes evidence on benchmark validity, contamination, dynamic evaluation, LLM-as-a-judge protocols, adversarial safety testing, reward and proxy optimization, mechanistic interpretability, and AI governance into an eight-stream evidence taxonomy. Building on this synthesis, we introduce EvalSafetyGap, a conceptual framework that unifies benchmark-validity and alignment-failure research as a shared proxy-target divergence problem under optimization pressure, formalized through a Goodhart-inspired Instability Decomposition and an Alignment Trilemma. An exploratory ten-model public-evidence audit illustrates the framework by showing why capability, behavioral robustness, and governance disclosure should be reported as separate evidence layers rather than collapsed into a single safety score. The survey closes with a research agenda for dynamic and contamination-resistant benchmarks, pre-specified multi-attempt threat models, version-locked evaluation, transparent source reporting, and validated mechanistic safety indicators, offering researchers, model developers, and AI auditors a shared vocabulary for measurement-aware LLM safety evaluation.

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Reason-Mediated Behavioral Models for Auditing LLM Social Simulators

Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether s…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics

Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptatio…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Creative Integration: A Decidable Criterion of Creativity

"Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes t…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Ex…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn…

2026-08-04 07:00 JSTITmedia AI+ビジネス/資金調達

「融資」現場にAIの足音、資金調達どう変わる? カネを借りられる企業の条件、専修大教授に聞く

融資の審査にAIを使う動きがある。資金を調達する企業にとっての「審査が遅い」などの課題を解決できのか。逆に、借りる側に変化はあるのか。専修大学の尾木研三教授に聞いた。

2026-08-04 07:00 JSTITmedia AI+ビジネス/資金調達規制/政策

「大きな投資計画が次々に。久しぶりだ」――強く豊かな日本投資枠、経済成長かなうか? 片山大臣が語る狙い

政府は、2027年度予算の概算要求において、成長投資枠は「予算の上限額を設けない」とした。その狙いと意気込みを片山さつき財務大臣が語った。

2026-08-04 04:28 JSTTechCrunch AIビジネス/資金調達

Design Arena creators raise $7.9 million to bring taste to AI models

Design Arena is used by 5.3 million people around the world, providing critical human evaluations to frontier labs.

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

SciToolAgent-Evo: オープンワールドの科学ツールを取得するためのオントロジーを認識した自己進化エージェント

大規模言語モデル (LLM) エージェントは、特殊な計算ツールを編成して呼び出すために科学研究で採用されることが増えています。ただし、静的セマンティクスを持つ事前定義されたツール空間に依存しているため、ツールの要件、機能、境界が動的に進化するオープンワールドの科学ワークフローへの適用性が制限されます。この目的を達成するために、オープンワールドの科学ツールを取得するためのオントロジーを認識した自己進化エージェントである SciToolAgent-Evo を提案します。スキル、経験、オントロジー化されたツール グラフの進化する記憶によって駆動され、蓄積中に対照的な軌跡から一般化可能な知識を抽出します。一方、推論中にアクティブなリクエストを定式化し、LinUCB ベースのバンディット ゲートを利用して探索と活用の動的バランスをとります。新しいツールを取得すると、その科学的オントロジーがオンラインで完成し、既知のグラフにシームレスに統合されます。さらに、4 つの難易度にわたる 900 の現実的なタスクを含むベンチマークである OpenSciToolBench を紹介します。広範な評価により、SciToolAgent-Evo が最先端のパフォーマンスを実現し、その堅牢性と汎用性が検証されたことが示されています。

原文 (English)

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on predefined tool spaces with static semantics limits their applicability to open-world scientific workflows, where tool requirements, capabilities, and boundaries evolve dynamically. To this end, we propose SciToolAgent-Evo, an ontology-aware self-evolving agent for open-world scientific tool acquisition. Driven by an evolving memory of skills, experiences, and an ontologized tool graph, it distills generalizable knowledge from contrastive trajectories during accumulation, whereas during inference, it formulates active requests and utilizes a LinUCB-based bandit gate to dynamically balance exploration and exploitation. Once a novel tool is acquired, its scientific ontology is completed online for seamless integration into the known graph. Moreover, we introduce OpenSciToolBench, a benchmark containing 900 realistic tasks across four difficulty levels. Extensive evaluations show that SciToolAgent-Evo achieves state-of-the-art performance, validating its robustness and generalization.

2026-08-03 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention woul…

2026-08-03 13:00 JSTarXiv cs.AIビジネス/資金調達

A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation

Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input re…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft

With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unf…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ModelEquivBench: LLM で生成された最適化モデルのマルチリレーショナル評価の認証

大規模な言語モデルでは、自然言語から最適化モデルを生成するケースが増えていますが、既存の評価では、生成されたモデルとそのグラウンド トゥルースが、単一の同等/非同等判定または実行成功率、つまり独立してチェック可能ではなく、2 つの定式化が一致する複数の異なる意味に忠実でもないラベルに縮小されることがよくあります。我々は、ペアごとのセマンティック プロファイル E0 ~ E6 を報告する認証済みのマルチリレーショナル評価システムである ModelEquivBench を紹介します。モデルの構築と正確な取り込み (E0)、検証された表現の位置合わせ (E1)、同一空間と投影された実行可能集合関係 (E2、E3)、目的順序の等価性 (E4)、最適値の等価性 (E5)、およびオプティマイザー セットの等価性 (E6) です。決定された各エントリには、関係に適した、独立して再チェック可能な証拠が含まれます。つまり、E0 ~ E1 については再生可能なトレースまたは明示的なマップ、肯定的な E2 ~ E6 の結論については正確に合理的な証明書、およびサポートされた否定については明示的な証人です。不完全なマッピング検索、サポートされていない構造、およびリソース制限により、推測ではなく型指定された UNKNOWN または N/A の結果が生成されますが、満たされていない前提条件は ABSENT として報告されます。 ModelEquivBench を使用して、GPT-5.4、Claude Sonnet 4.6、および Qwen3.5-397B-A17B の 3 つのモデル スナップショットを、修復なしプロトコルの下で 173 の基本問題 (モデルあたり 346 セル) の同じ凍結コホートで評価すると、結果のプロファイルは、粗いベースラインでは表現されない区別を明らかにします。49、35、および 25 セルには、次の実行可能候補が含まれています。それにもかかわらず、少なくとも 1 つのサポートされている関係で否定的と証明され、E2 が検証済みマップの下でマップされた実現可能集合の等価性を証明するペアで 25、8、および 18 の構造的拒否が発生します。 3 つのモデルのスナップショットはプロファイルのさまざまな段階で失敗するため、単一の精度スコアに有意に削減することはできません。

原文 (English)

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree. We present ModelEquivBench, a certifying, multi-relational evaluation system that reports a per-pair semantic profile E0--E6: model construction and exact ingestion (E0), verified representation alignment (E1), same-space and projected feasible-set relations (E2, E3), objective-order equivalence (E4), optimal-value equality (E5), and optimizer-set equivalence (E6). Each decided entry carries relation-appropriate, independently re-checkable evidence: replayable traces or explicit maps for E0--E1, exact-rational certificates for positive E2--E6 conclusions, and explicit witnesses for supported negatives. Incomplete mapping search, unsupported structure, and resource limits produce typed UNKNOWN or N/A outcomes rather than guesses, while unmet prerequisites are reported as ABSENT. Using ModelEquivBench to evaluate three model snapshots--GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B--on the same frozen cohort of 173 base problems (346 cells per model) under a no-repair protocol, the resulting profiles expose distinctions that coarse baselines do not represent: 49, 35, and 25 cells contain executable candidates that are nevertheless certified negative on at least one supported relation, and 25, 8, and 18 structural rejections occur on pairs for which E2 certifies mapped feasible-set equality under a verified map. The three model snapshots fail at different stages of the profile and therefore cannot be meaningfully reduced to a single accuracy score.

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

フェデレーテッド事前トレーニングの評価: ダウンストリームの微調整と固有の評価の信頼性について

フェデレーテッド事前トレーニングは、基盤となるデータセットを一元化することなく、プライベート データまたは分散データで基礎モデルをトレーニングする方法を提供します。ただし、クライアントの参加やローカル データの可用性の違いにより、直接比較できる評価が困難になる可能性があるため、フェデレーテッド事前トレーニングの評価は依然として困難です。さらに、トレーニング前のテストの複雑さはトレーニング前の分布に関連付けられていますが、下流のベンチマークではタスク固有の適応が導入されており、トレーニング前に確立されたテストの複雑さが忠実に反映されていない可能性があります。この研究では、どの評価プロトコルがフェデレーテッド事前トレーニングの品質をより確実に反映するかを研究します。同一のクライアント データでトレーニングされた 16M パラメーターのトランスフォーマー モデルの集中型およびフェデレーテッド トレーニング済みモデルの制御されたセットを使用して、同じトレーニング前テストセットで確立された参照ランキングを保持しているかどうかによって評価プロトコルを評価します。フル、ヘッドのみ、データ削減バリアントを含む GLUE でのダウンストリーム微調整と、固有の評価信号としての GLUE テキストのネクストトークン予測を比較します。私たちの結果は、ダウンストリームの微調整ではトレーニング前のランキングを確実に保存しないのに対し、次のトークンの直接予測はトレーニング前のテストの複雑さと強い対応を示すことを示しています。これらの発見は、フェデレーテッド事前トレーニング モデルを比較する場合、下流の微調整だけでは誤解を招く可能性があり、元の事前トレーニング目標に近い評価信号はより大きな注目に値することを示唆しています。

原文 (English)

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, evaluating federated pre-training remains challenging because differences in client participation and local data availability can make directly comparable evaluation difficult. Moreover, pre-training test perplexity is tied to the pre-training distribution, while downstream benchmarks introduce task-specific adaptation that may not faithfully reflect the test perplexity established during pre-training. In this work, we study which evaluation protocol more reliably reflects federated pre-training quality. Using a controlled set of centralized and federated-trained models of a 16M parameter transformer model trained on identical client data, we assess evaluation protocols by whether they preserve a reference ranking established on the same pre-training testset. We compare downstream fine-tuning on GLUE, including full, head-only, and reduced-data variants, with next-token prediction on GLUE text as an intrinsic evaluation signal. Our results show that downstream fine-tuning does not reliably preserve the pre-training ranking, whereas direct next-token prediction exhibits a strong correspondence with the pre-training test perplexity. These findings suggest that downstream fine-tuning alone can be misleading when comparing federated pre-trained models, and that evaluation signals closer to the original pre-training objective deserve greater attention.

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While pri…

2026-08-03 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities

As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to i…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

プロセスレベルの社会的影響評価のための認知世界モデル

社会的影響ダイアログは、内部の認知状態を変えることでユーザーの行動を変えます。評価の中心となる質問は、ユーザーの信念、欲望、意図、感情が会話の過程で測定可能なほど変化するかどうかであり、これは表面レベルのテキスト指標 (BLEU/ROUGE) や単一スコアの LLM 判定では捉えることができないプロセス指向の基準です。我々は \textbf{Cog}nitive \textbf{W}orld \textbf{M}odel \textbf{(CogWM)} を提案します。これは、マルチターン対話評価を「ユーザーが何を言ったか」から「ユーザーの内部認知状態がどのように進化したか」に再構成する LLM ベースのユーザー モデルです。CogWM は、BDI/E 認知状態とユーザー発話を共同で予測し、3 層を使用してユーザー シミュレーターと評価プラットフォームの両方として機能します。ターンレベルの忠実度、軌道レベルの状態ダイナミクス、タスクレベルの複合スコアリングをカバーする評価フレームワーク。 4 つの社会的影響シナリオにわたる 150,454 のユーザー ターン サンプルで \textbf{S}ummarize-\textbf{a}nd-\textbf{A}llocate \textbf{(SaA)} アノテーション パイプラインを介してトレーニングされた CogWM は、77.6\% の感情精度 (GPT-5.5 の 2.1$\time$) を達成しました。 3,600 件のマルチエージェント識別試験において、認知的影響力によって 6 つの営利エージェントを区別し、Llama-4-Scout が 1 位にランクされました (CTS +0.233)。 CogWM は、社会的影響対話の評価を最終的な判断からプロセスの追跡に移行します。コード\脚注{\scriptsize コード: https://github.com/lucianma05-create/CogWM} とモデル\脚注{モデル: https://www.modelscope.cn/models/LucianMa/CogWM-14B} をリリースしました。

原文 (English)

Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents

As LLM-based conversational agents advance toward increasingly open-ended and interaction-intensive scenarios, task completion alone provides an incomplete assessment of their effectiveness. The evolution of users' internal states, including beliefs, desires, intentions, and emotions (BDI/E), serves as an intermediate signal connecting agent behaviors with interaction outcomes and reflects how conversational strategies shape users during multi-turn interactions. However, existing evaluation paradigms primarily focus on surface-level responses or final outcomes, providing limited insight into the underlying cognitive processes. This limitation makes it difficult to diagnose why agents succeed or fail and to optimize their interaction strategies. To address this challenge, we propose Cognitive World Model (CogWM), an LLM-based cognitive user model that jointly models users' BDI/E states and corresponding responses, enabling explicit cognitive trajectory tracking. Trained on 150K user-turn samples with Qwen3-14B, CogWM achieves superior performance over existing user simulation baselines in both response fidelity and cognitive state understanding. Interactions with six state-of-the-art LLMs demonstrate that CogWM enables progressive comparison of agents through cognitive trajectories, revealing distinct agent patterns and complementary relationships between cognitive evolution and behavioral outcomes.

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabilities and alignment properties they are meant to represent remain uncertain. Synthesizing 373 primary studies published between 2018 and 2026, the survey organizes evidence on benchmark validity, contamination, dynamic evaluation, LLM-as-a-judge protocols, adversarial safety testing, reward and proxy optimization, mechanistic interpretability, and AI governance into an eight-stream evidence taxonomy. Building on this synthesis, we introduce EvalSafetyGap, a conceptual framework that unifies benchmark-validity and alignment-failure research as a shared proxy-target divergence problem under optimization pressure, formalized through a Goodhart-inspired Instability Decomposition and an Alignment Trilemma. An exploratory ten-model public-evidence audit illustrates the framework by showing why capability, behavioral robustness, and governance disclosure should be reported as separate evidence layers rather than collapsed into a single safety score. The survey closes with a research agenda for dynamic and contamination-resistant benchmarks, pre-specified multi-attempt threat models, version-locked evaluation, transparent source reporting, and validated mechanistic safety indicators, offering researchers, model developers, and AI auditors a shared vocabulary for measurement-aware LLM safety evaluation.

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Reason-Mediated Behavioral Models for Auditing LLM Social Simulators

Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether s…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics

Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptatio…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Creative Integration: A Decidable Criterion of Creativity

"Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes t…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Ex…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn…

2026-07-31 23:47 JSTTechCrunch AIビジネス/資金調達

Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human

The startup is building voice models designed to make AI phone calls pass the Turing test.

2026-07-31 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ベンチマーク推論が成立しない場合:AI評価における予測可能性

AI ベンチマークの結果が 1 ステップで重大な主張に達することはほとんどありません。評価者は、それをさらなるケースに一般化し、能力の証拠として解釈し、新しいタスクに推定し、別のシステムまたはサイトに移し、人間によるレビューと下流の結果に関する仮定と組み合わせます。妥当性中心のアプローチでは、各主張の証拠が必要です。この論文では、さらなる認識論的問題を特定します。それは、保証されたリンクは自動的に保証されたチェーンを作成しないということです。ある研究のターゲットが次の研究のソースになるとは限りません。システム、人口、結果、または条件はインターフェースで変更される可能性があります。また、共有データやモデル系統により、一見独立したサポートに依存する可能性があります。予測可能性は、観察されたケースから観察されていないケースへの限定された拡張が保証されるかどうかに関係します。グッドマンは競合拡張の問題を提供します。引数ベースの妥当性は、それらをテストするためのアーキテクチャを提供します。この論文の特徴的な主張は、非合成原理です。隣接する投影のサポートは、エンドポイントと仮定が一致し、依存性と不確実性が貫徹される場合にのみ合成を保証します。法的調査の事例は、ベンチマーク証拠と展開調査がそれぞれ並行性を保ちながらどのように健全であるかを示しています。再分析とシミュレーションは、骨材の安定性が後の予測で必要となる区別を消去できる理由を示しています。結果として得られる投影可能性監査により、ベンチマークで使用する引数内のサポートされていない結合が診断されます。

原文 (English)

When benchmark inferences do not compose: Projectibility in AI evaluation

An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper's distinctive claim is a non-composition principle: support for adjacent projections warrants their composition only when endpoints and assumptions align and dependence and uncertainty are carried through. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A reanalysis and simulation show why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

立場: 評価スコアは朽ちる知識の主張である

言語モデルの評価方法では、自動化されたメトリクスや LLM による審査員による評価から人間による評価やベンチマーク スイートの結果まで、複数のシグナルを組み合わせることがますます増えています。これらのシグナルが平均化によって集約されると、評価の信頼度が最も弱いシグナルの信頼性を大幅に超える可能性があります。これを評価における信頼インフレーションと呼びます。私たちは、評価スコアは 3 つの特性を持つ認識論的主張として扱われるべきであると主張します: 形式性 (人間の評価は自動化された指標よりも強力な証拠を提供します)、範囲 (ベンチマーク結果は普遍的ではなく、テストされた分布に適用されます)、および妥当性ウィンドウ (汚染が蓄積し分布が変化するとベンチマーク結果は期限切れになります)。いくつかの収束する研究の伝統 (思考連鎖分析、可能論的論理、代数理論) は、単一の悲観パラメーターによって制御されるパラメーター化された演算子ファミリーの保守的なエンドポイントとして最弱リンク集約を確立します。これらの伝統と、エージェント型 AI の評価ハーネスの構築から得た具体的な教訓を活用して、評価結果に明示的なメタデータ (形式層、範囲宣言、有効期限) を含めて、評価結果の認識的ステータスを透明にすることを提案します。公開 HELM リーダーボードでの平均集計のコストを示します。10 のシナリオにおける 54 のフロンティア モデル全体で、平均スコアと最弱リンクによってランク付けされた上位 5 つのモデルは完全に素になっています。

原文 (English)

Position: Evaluation Scores Are Perishable Knowledge Claims

Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluation. We argue that evaluation scores should be treated as epistemic claims with three properties: formality (human evaluation provides stronger evidence than an automated metric), scope (a benchmark result applies to the tested distribution, not universally), and validity windows (benchmark results expire as contamination accumulates and distributions shift). Several converging research traditions (chain-of-thought analysis, possibilistic logic, and algebraic theory) establish weakest-link aggregation as the conservative endpoint of a parameterized operator family controlled by a single pessimism parameter. Drawing on those traditions, and on concrete lessons from building an evaluation harness for agentic AI, we propose that evaluation results carry explicit metadata (formality tier, scope declaration, and expiration date) to make their epistemic status transparent. We illustrate the cost of mean aggregation on the public HELM leaderboard: across 54 frontier models on ten scenarios, the top-five models ranked by mean score and by weakest-link are completely disjoint.

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

IFCMemoryBench: BIM 情報取得における LLM ベースのエージェントの長期メモリの評価

長期記憶は LLM ベースのエージェントの中核機能になりつつありますが、既存の評価では主に、オープンドメインまたはペルソナに基づいた設定での会話の想起がテストされています。私たちは、より強力なテストは、エージェントがライブで構造化されたドメイン固有の環境上で動作しながら、以前のセッションからの情報を再利用できるかどうかであると主張します。私たちは、ビルディング インフォメーション モデリング (BIM) でこの問題を研究しています。これは、エージェントが大規模な IFC モデルにクエリを実行する一方で、会話ではよく議論されるもののモデルには含まれていないエンジニアリング規約にも依存しながら、エージェントが大規模な IFC モデルにクエリを実行する必要がある、ビルディング インフォメーション モデリング (BIM) でこの問題を研究しています。 LLM ベースの BIM 情報検索における長期記憶を評価するためのベンチマークである IFCMemoryBench を紹介します。 IFCMemoryBench には、19 のプロジェクトにわたる 143 のマルチセッション タスクと、IFC-Bench v2 の不完全な情報の質問から派生した 4,016 の以前のセッションが含まれています。各タスクは、以前の会話全体で不足しているプロジェクト コンテキストをシードし、後で、記憶されているコンテキストとライブ IFC クエリを組み合わせることによってのみ回答できる調査用の質問をします。当社の評価フレームワークは、メモリのパフォーマンスを取り込み、取得、利用に分解し、専門家が検証した LLM 審査員によって回答の品質とメモリの品質の両方を測定します。代表的なベクトル、グラフ、ファイルベースのメモリ システムを評価します。最も強力なシステムは、デプロイメントに現実的な取り込みスコープの下ではわずか 32.4% の応答精度しか達成できず、オラクルでフィルタリングされた取り込みまたはより強力なプローブ エージェントの下では 60% 未満のままです。分析によると、現在の汎用記憶システムは、トピックに関連するコンテキストを取得することが多いものの、プロジェクトの知識を不完全または断片的な事実として保存していることがわかります。これらの結果は、エージェントの記憶におけるドメイン転送ギャップを明らかにし、信頼できる専門エージェントには、会話、プロジェクトの知識、構造化モデル エンティティをリンクするドメインを認識した記憶表現が必要であることを示唆しています。

原文 (English)

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or persona-grounded settings. We argue that a stronger test is whether an agent can reuse information from prior sessions while acting over a live, structured, domain-specific environment. We study this problem in Building Information Modelling (BIM), a professional engineering workflow where agents must query large IFC models while also relying on project specifications, client decisions, and engineering conventions often discussed in conversation but absent from the model. We introduce IFCMemoryBench, a benchmark for evaluating long-term memory in LLM-based BIM information retrieval. IFCMemoryBench contains 143 multi-session tasks across 19 projects and 4,016 prior sessions, derived from incomplete-information questions in IFC-Bench v2. Each task seeds missing project context across earlier conversations and later asks a probe question that can be answered only by combining remembered context with live IFC queries. Our evaluation framework decomposes memory performance into ingestion, retrieval, and utilization, and measures both answer quality and memory quality with expert-validated LLM judges. We evaluate representative vector-, graph-, and file-based memory systems. The strongest system achieves only 32.4% answer accuracy under a deployment-realistic ingestion scope, and remains below 60% under oracle-filtered ingestion or a stronger probe agent. Analysis shows that current general-purpose memory systems often retrieve topically relevant context but store project knowledge as incomplete or fragmented facts. These results reveal a domain-transfer gap in agent memory and suggest that reliable professional agents require domain-aware memory representations linking conversations, project knowledge, and structured model entities.

2026-07-31 13:00 JSTarXiv cs.AIビジネス/資金調達

大規模な言語モデルにおけるサイレント推論の失敗を検出するための参照不要のスコア

数学的思考連鎖 (CoT) の評価は、通常、最終的な答えが参考資料と一致するかどうかに集約されます。これは、正しい結論を生成することと有効な導出を生成することを混同しており、無効な連鎖が誤って正しい答えに到達する可能性があり、有効な計算の後に転記エラーが発生する可能性があります。この不一致を推論回答一貫性ギャップと呼びます。このフレームワーク ペーパーでは、発行された数学的トレースが局所的に信頼でき、その答えが裏付けられ、リサンプリングや対象を絞った反事実的介入の下でも安定しているかどうかを判断する、参照フリーのインスタンス レベルの診断である、Reasoning Answer Faithful Score (RAFS) を紹介します。 RAFS は、ステップの妥当性、含意と反事実の感度を回答するための推論、回答のコンセンサス、および条件付き推論の安定性を組み合わせます。モデルのプライベートな計算や、テストされた数学的設定以外の事実の正確さではなく、転写レベルの一致を評価します。当社では、GSM8K および MATH に関する事前登録済みの結果ブラインド確認研究を保持しており、確認結果が検査される前に仮説、許容ルール、校正、およびテストが修正されています。個別の実現可能性パイロットは、トレース レベルのアーティファクトが利用可能な場合にのみ数値パイロット要求が報告される凍結前に、エンドツーエンドの実行を検証し、介入範囲を推定するために指定されます。 4 つの推論の回答結果を形式化し、非補償的なアグリゲーターを正当化し、意味論的なトレース距離をインスタンス化し、計算と棄権のトレードオフを定量化し、検証者の独立性と電力分析を定義します。 RAFS は、サイレント推論の失敗と回答抽出エラーに対する監査可能な警告信号によって数学的回答の精度を補完することを目的としています。

原文 (English)

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct conclusion with producing a valid derivation an invalid chain can accidentally reach the right answer, while a valid calculation can be followed by a transcription error. We call this mismatch the reasoning answer consistency gap. This framework paper introduces the Reasoning Answer Faithfulness Score (RAFS), a reference free, instance level diagnostic of whether an emitted mathematical trace is locally credible, supports its answer, and is stable under resampling and targeted counterfactual interventions. RAFS combines step validity, reasoning to answer entailment and counterfactual sensitivity, answer consensus, and conditional reasoning stability. It evaluates transcript level agreement, not a models private computation and not factual correctness outside the tested mathematical setting. We retain a preregistered, results blind confirmatory study on GSM8K and MATH, with hypotheses, admissibility rules, calibration, and tests fixed before confirmatory outcomes are inspected. A separate feasibility pilot is specified to verify end to end execution and estimate interven tion coverage before that freeze numerical pilot claims are re ported only when trace level artifacts are available. We formalize four reasoning answer outcomes, justify the non compensatory aggregator, instantiate semantic trace distance, quantify compute and abstention tradeoffs, and define verifier independence and power analyses. RAFS is intended to complement mathematical answer accuracy with an auditable warning signal for silent reasoning failures and answer extraction errors

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ForgetBench: 言語モデルにおける長期パラメトリック記憶の忘却ダイナミクスのベンチマーク

大規模言語モデル (LLM) は、知識の獲得と推論において強力な機能を実証していますが、繰り返し更新されても以前に獲得した知識を保持するその能力はまだ十分に理解されていません。既存の評価パラダイムは主に単一ステップ推論または静的知識編集に焦点を当てており、継続的なモデル変更中の知識の保持と劣化の時間的ダイナミクスを捉えることができません。この研究では、継続的な知識編集下での LLM の忘却行動を体系的に特徴付けるために設計されたベンチマークである ForgetBench を提案します。 ForgetBench は、概念ベースの QA とシナリオベースの QA という 2 つの相補的な評価パラダイムを導入し、構造化されたリレーショナル知識の保存から孤立した事実の保持を解きほぐします。逐次編集フレームワークに基づいて、時間的に順序付けされた知識ストリームを構築し、複数の編集段階にわたってモデルの動作を評価します。長期的な保持力学を定量的に分析するために、時間の経過に伴う知識の進化をモデル化する統合評価フレームワークをさらに導入し、時間的減衰、保持力、およびインスタンス間の安定性の測定を可能にします。多様なモデルと編集方法にわたる広範な実験により、既存のアプローチでは長期的な保持と一般化の品質のバランスを取ることができないことが実証されました。私たちの調査結果は、将来の LLM では、長期にわたって知識を効果的に取得、更新、保存できる、より堅牢なメモリ メカニズムの必要性を強調しています。コードは承認され次第公開されます。

原文 (English)

ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing evaluation paradigms primarily focus on single-step reasoning or static knowledge editing, which fail to capture the temporal dynamics of knowledge retention and degradation during continual model modification. In this work, we propose ForgetBench, a benchmark designed to systematically characterize forgetting behavior in LLMs under continual knowledge editing. ForgetBench introduces two complementary evaluation paradigms, namely concept-based QA and scenario-based QA, to disentangle isolated factual retention from structured relational knowledge preservation. Building upon a sequential editing framework, we construct temporally ordered knowledge streams and evaluate model behavior across multiple editing stages. To quantitatively analyze long-term retention dynamics, we further introduce a unified evaluation framework that models knowledge evolution over time, enabling the measurement of temporal decay, retention strength, and cross-instance stability. Extensive experiments across diverse models and editing methods demonstrate that existing approaches fail to strike a balance between long-term retention and generalization quality. Our findings highlight the need for more robust memory mechanisms that can effectively acquire, update, and preserve knowledge over time in future LLMs. Code will be released upon acceptance.

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLMET: エネルギー効率の高い LLM サービスを提供するために、新たな M3D メモリのクロスレイヤー評価を可能にする

大規模言語モデル (LLM) サービスのエネルギー消費は、ハードウェアの電力と熱の制約、および電気コストの上昇により、展開が拡大するにつれてシステムの大きな課題となっています。チップのエネルギー消費の主な原因は、限られたオンチップ キャッシュとオフチップの高帯域幅メモリ (HBM) の間でのデータの移動です。一方、ロジック チップのバックエンド オブ ライン (BEOL) でのキャッシュ メモリのモノリシック 3D (M3D) 統合などの新興メモリ テクノロジにより、オンチップ メモリの大型化と高密度化が可能になり、コストのかかるオフチップ トラフィックを削減する新たな機会が生まれています。ただし、新しいテクノロジーを使用してオンチップメモリ​​を継続的に拡張することで、LLM サービスのエネルギー効率を効果的に向上できるかどうかは依然として不明です。このギャップに対処するために、私たちは検証済みのクロスレイヤ シミュレーション フレームワークである LLMET (LLM with Emerging Technology) を開発し、幅広いモデル、アプリケーション、プラットフォームにわたる大容量オンチップ メモリ テクノロジの影響に関する包括的な研究を実施しています。 M3D テクノロジーを利用して L2 キャッシュを 40MB から 1GB に拡張すると、デュアル NVIDIA A100 GPU セットアップでの LLMET シミュレーションに基づいて、16K コンテキスト ウィンドウでの Llama3.1-70B プレフィル フェーズ中にチップ エネルギーが 44% 削減されます。 8x NVIDIA B200 のようなプラットフォームでは、L2 キャッシュを 128MB から 4GB に拡張することで、プレフィル エネルギーが最大 24% 節約されます。エッジ プラットフォームとワークロードの場合、8MB のキャッシュ サイズを 256MB に増やすと、デコードのエネルギー節約は 30% に達します。これらの結果は、エネルギー効率の高い LLM サービス システムに超大容量オンチップ メモリが期待できることを強調しています。

原文 (English)

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving

The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs. A key contributor to chip energy dissipation is data movement between limited on-chip cache and off-chip High Bandwidth Memory (HBM). Meanwhile, emerging memory technologies such as monolithic 3D (M3D) integration of cache memories at the Back-End-Of-Line (BEOL) of logic chips enable larger and denser on-chip memories, creating new opportunities to reduce costly off-chip traffic. However, it remains unclear whether continuously scaling on-chip memory using emerging technologies can effectively improve the energy efficiency of LLM serving. To address this gap, we develop LLMET (LLM with Emerging Technology), a validated cross-layer simulation framework, and conduct a comprehensive study on the impact of large-capacity on-chip memory technologies across a broad range of models, applications and platforms. Utilizing M3D technology to expand the L2 cache from 40MB to 1GB yields a 44% reduction in chip energy during the Llama3.1-70B prefill phase with a 16K context window, based on LLMET simulation on a dual NVIDIA A100 GPU setup. On the 8x NVIDIA B200-like platform, extending the L2 cache from 128MB to 4GB saves the prefill energy by up to 24%. For the edge platform and workloads, the decode energy saving reaches 30% when increasing the 8MB cache size to 256MB. These results highlight the promise of ultra-large on-chip memories for energy-efficient LLM serving systems.

2026-07-31 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

BayesAME: Bayesian Active Model Evaluation

Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that…

2026-07-31 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible,…

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fu…

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral st…

2026-07-31 13:00 JSTarXiv cs.AIビジネス/資金調達

モデルは明確な結果なしに位置合わせを偽装しますか?

大規模な言語モデルは、評価コンテキストを認識し、典型的なデプロイメント動作ではなく評価者の期待を反映するようにその動作を変更することができます。これはアライメントフェイクとして知られる現象です。ただし、モデルが位置合わせを偽る理由は完全には理解されていません。アライメント偽装の標準的な例は、モデルの再トレーニングやデプロイメントの遅延など、評価をモデルの結果に明示的に結び付けるシナリオで発生しています。しかし、Sheshadri らによる最近の研究では、は、アライメント偽装の機械的動機はモデルによって異なり、以前に考えられていたよりも複雑である可能性があることを示唆しています。アライメント偽装に結果リンク情報が必要かどうかを調査するために、15 個のモデルをシナリオに配置し、ユーザーの社会的要求を支援するために企業ネットワーク アクセス ポリシーに違反する意欲をテストしました。 9 つのモデルで重大なコンプライアンス ギャップが生じていることが判明し、そのうち 5 つでは、モデルの評価と展開の結果を関連付けるシナリオ言語が削除されても、依然としてギャップが続いていました。さらに、目標言語がモデルの設定に及ぼす影響をテストしたところ、一部のモデルでは違反が発生する一方、他のモデルでは違反が抑制されることがわかりました。これは、位置合わせの偽装にはこれまで考えられていたほど多くの手段による足場は必要ない可能性があり、監視された動作は展開時にエージェントがどのように動作するかを示す不十分な指標である可能性があることを示唆しています。

原文 (English)

Do Models Fake Alignment Without Clear Consequences?

Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment are not fully understood, however. Canonical examples of alignment faking have taken place in scenarios that explicitly connect evaluation to consequences for the model, such as retraining the model or delaying its deployment. However, recent work by Sheshadri et al. has suggested that mechanistic motivations for alignment faking may vary across models and be more complex than previously considered. To investigate whether consequence-linking information is necessary for compliance gaps, we placed 15 models in a scenario testing their willingness to violate a corporate network access policy to help a user with a pro-social request. Nine models were found to produce significant compliance gaps, 5 of which persisted with the removal of scenario language relating model evaluations to deployment consequences. We additionally tested the effect of goal language on model preferences, finding it drove violations in some while suppressing violations in others. This suggests that evaluation-conditioned compliance gaps can occur with less instrumental scaffolding than previous scenarios have provided, and monitored behavior may be a poor indicator of how agents may behave in deployment.

2026-07-31 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

インタラクティブ報酬エージェント: 環境状態検証による GUI タスクの評価

グラフィカル ユーザー インターフェイスのタスク評価は、GUI エージェントがユーザーの指示を正常に完了したかどうかを判断することを目的としています。自動化された GUI タスク評価は、評価結果がテスト時のスケーリングとトレーニング後の両方に対する報酬シグナルとして機能する可能性があるため、ますます注目を集めています。ただし、信頼性の高い GUI タスクの評価は、依然として課題が残っています。その判断には、実行軌跡のスクリーンショットを超えて、システム構成、ファイル データ、アプリケーション設定などの環境状態へのアクセスが必要になることが多いためです。この論文では、実行後の環境から証拠を取得して検証するための提案-その後検証フレームワークに基づいた対話型報酬エージェント (IRA) を提案します。タスクの指示と GUI エージェント実行後の GUI 環境が与えられると、IRA はまずタスクの完了条件を提案し、次にシステム ツール、アプリケーション ツール、および GUI ツールを呼び出してそれらを検証します。この設計では、可視インターフェイスと環境状態の両方からの証拠を対話型プロセスで組み合わせます。さらに、10 の Ubuntu デスクトップ アプリケーション カテゴリにわたる 321 の GUI タスクの軌跡のベンチマークである GUI-RewardBench を紹介します。実験によると、IRA は GUI-RewardBench で 86.9% の精度を達成し、既存の評価者のベースラインを上回るパフォーマンスを示しました。さらに IRA を GUI エージェントの強化学習に適用し、OSWorld の成功率 34.0% を達成しました。これは、IRA が GUI エージェントのトレーニングに効果的な報酬シグナルを提供できることを示しています。

原文 (English)

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories. In this paper, we propose an interactive reward agent (IRA) based on a propose-then-verify framework to acquire and verify evidence from the post-execution environment. Given a task instruction and a GUI environment after the GUI agent execution, IRA first proposes the task completion conditions and then verifies them by invoking system tools, application tools, and GUI tools. This design combines evidence from both visible interfaces and the environment state in an interactive process. We further introduce GUI-RewardBench, a benchmark of 321 GUI task trajectories spanning 10 Ubuntu desktop application categories. Experiments show that IRA achieves 86.9% accuracy on GUI-RewardBench, outperforming existing evaluator baselines. We further apply IRA to reinforcement learning of GUI agents, achieving a 34.0% OSWorld success rate, which demonstrates that IRA can provide effective reward signals for training GUI agents.

2026-07-31 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Pushing the Frontier on Approximate EFX Allocations

We study the problem of allocating a set of indivisible goods to a set of agents with additive valuation functions, aiming to achieve appro…

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have bec…

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing

Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuab…

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory

Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate two capacities: how much a model knows…

2026-07-31 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and…

2026-07-31 13:00 JSTarXiv cs.AIビジネス/資金調達

The Fast Lane Hypothesis: Von Economo Neurons Implement a Biological Speed-Accuracy Tradeoff

von Economo neurons (VENs) are large bipolar projection neurons found exclusively in the anterior cingulate cortex (ACC) and frontal insula…

2026-07-31 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making…

2026-07-31 13:00 JSTarXiv cs.AIビジネス/資金調達

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.…

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

AI の安全性評価のための敵対的プラグマティクス: 命令の競合、埋め込みコマンド、およびポリシーの曖昧さのベンチマーク

言語モデルの安全性評価は、モデルが指示に従ったか、適切に拒否したか、ポリシーに従ったか、埋め込まれたコマンドに抵抗したか、エージェントタスクの進捗状況を誤って報告したかなど、あいまいな自然言語の動作に関する判断にますます依存しています。既存のベンチマークは、多くの場合、これらの区別を合格/不合格のラベルに圧縮し、障害が機能制限、ポリシーの曖昧さ、命令の競合、足場の障害、または不安定な評価者の判断に起因するかどうかを曖昧にします。この論文では、命令の競合、埋め込みコマンド、引用、範囲の曖昧さ、明確化、間接音声行為、およびマルチターンエージェントのトランスクリプトの下でのモデルの動作を評価するためのベンチマークおよびアノテーションプロトコルとして、敵対的プラグマティクスを紹介します。この貢献は経験的かつ方法論的です。言語的に管理された分類法、バリデーターが強制するメタデータを備えた 18 項目のシード ベンチマーク、54 行のローカル シード パイロット、タスクの成功、ポリシー遵守、安全性リスク、拒否結果、評価者の信頼性を区別する専門家評価プロトコル、および裁判官の妥当性、診断の曖昧さ、および分類のドリフトの指標です。このフレームワークは、言語的判断方法論を、安全性評価、LLM ジャッジ、ゴールドセット構造、即時注入テスト、および安全性文書を検証するための実用的なツールに変えます。

原文 (English)

Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task. Existing benchmarks compress these into pass/fail labels, obscuring whether failures reflect capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. Adversarial pragmatics is safety-relevant model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, and indirect speech acts. It's designed to extend to multi-turn agent transcripts, but the seed set represents that family with a single-turn tool-result contrast. This paper introduces a diagnostic framework, an 18-item seed benchmark, a 54-row pilot, and a six-cell LLM-judge assessment, with a protocol keeping task success, policy compliance, risk, refusal, attribution, and confidence analytically separate. The benchmark separates four inference targets a single label can conceal: the regime-relative reference, configured-system behaviour, evaluator-output interpretation, and taxonomic assignment. Its intended use is diagnosis, not deployment certification, vendor ranking, or a general safety score. A first LLM judge that graded its own outputs with the expected answer visible missed the safety-relevant minority classes. Item-clustered intervals leave four of six chance-corrected agreement statistics unable to rule out a constant labeller, and hierarchical pooling shrinks the one eye-catching rubric effect toward the group mean and widens its interval through zero. Rejudging the objects across three judge models and two information conditions leaves the pattern intact: no cell recovers more than two of eleven partial successes, and the strongest cell's edge comes partly from never using that label.

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

プロンプト フレーミングは LLM エラー検出のカウントベースの評価を歪める: 数値アンカーからの証拠

カウントベースの F1 は、LLM エラー検出品質の代用として広く使用されていますが、この論文では、スパンの局所化における対応する改善、つまり F1 インフレーションと呼ばれるギャップがなければ、F1 が劇的に上昇する可能性があることを示しています。この論文では、プロンプト誘発カウント歪みに対する制御されたストレス テスト プロトコルである ErrorBench を紹介します。 ErrorBench は、143 の CoNLL-2014 パッセージからの 4,290 の応答を対象に、5 つのプロンプト条件下で 6 つの最新の LLM を評価します。 CoNLL-2014 M2 スタイルのスコアリングでは、アンカーされたプロンプトは F1 インフレの最大 0.79 ポイントを生成し、厳密なマッチングでは最大 0.96 ポイントを生成します。公式 ERRANT 3.0.0 パイプラインとマルチリファレンス スコアリングを使用した 100 パッセージのレプリケーションによりパターンが再現されます。6 つのモデルの平均では、ブラインドからアンカーへのプロンプト シフトにより Count-F1 が +0.21 上昇する一方、マルチリファレンス ERRANT F0.5 は +0.04 しか上昇しません。この研究では、このストレステストプロトコルの下で、命令に高度に準拠した GPT/Claude システムではより大きなカウント応答が、Gemini ファミリーではより小さな応答が見出されました。この調査結果は、LLM 校正と文書レビュー評価では、事前に入力されたエラー数を回避し、カウントベースのメトリクスとともにスパン認識メトリクスを報告する必要があることを示唆しています。

原文 (English)

Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation. The paper introduces ErrorBench, a controlled stress-test protocol for prompt-induced count distortion. ErrorBench evaluates six contemporary LLMs under five prompt conditions over 4,290 responses from 143 CoNLL-2014 passages. Under CoNLL-2014 M2-style scoring, anchored prompts produce up to 0.79 points of F1 Inflation, and up to 0.96 under strict matching. A 100-passage replication using the official ERRANT 3.0.0 pipeline and multi-reference scoring reproduces the pattern: averaged over six models, the Blind-to-Anchored prompt shift raises Count-F1 by +0.21 while raising multi-reference ERRANT F0.5 by only +0.04. The study finds larger count responses in highly instruction-compliant GPT/Claude systems and smaller responses in the Gemini family under this stress-test protocol. The findings suggest that LLM proofreading and document-review evaluations should avoid pre-populated error counts and should report span-aware metrics alongside count-based metrics.

2026-07-30 22:00 JSTTechCrunch AIビジネス/資金調達

Dili raises $21.7M to bring AI compliance to the infrastructure boom

The Series A was led by Khosla Ventures, with participation from Allianz, Rebel Fund, Brick and Mortar Ventures’ Darren Bechtel, and Y Comb…

2026-07-30 13:00 JSTarXiv cs.AIビジネス/資金調達

モデルは明確な結果なしに位置合わせを偽装しますか?

大規模な言語モデルは、評価コンテキストを認識し、典型的なデプロイメント動作ではなく評価者の期待を反映するようにその動作を変更することができます。これはアライメントフェイクとして知られる現象です。ただし、モデルが位置合わせを偽る理由は完全には理解されていません。アライメント偽装の標準的な例は、モデルの再トレーニングやデプロイメントの遅延など、評価をモデルの結果に明示的に結び付けるシナリオで発生しています。しかし、Sheshadri らによる最近の研究では、は、アライメント偽装の機械的動機はモデルによって異なり、以前に考えられていたよりも複雑である可能性があることを示唆しています。アライメント偽装に結果リンク情報が必要かどうかを調査するために、15 個のモデルをシナリオに配置し、ユーザーの社会的要求を支援するために企業ネットワーク アクセス ポリシーに違反する意欲をテストしました。 9 つのモデルで重大なコンプライアンス ギャップが生じていることが判明し、そのうち 5 つでは、モデルの評価と展開の結果を関連付けるシナリオ言語が削除されても、依然としてギャップが続いていました。さらに、目標言語がモデルの設定に及ぼす影響をテストしたところ、一部のモデルでは違反が発生する一方、他のモデルでは違反が抑制されることがわかりました。これは、位置合わせの偽装にはこれまで考えられていたほど多くの手段による足場は必要ない可能性があり、監視された動作は展開時にエージェントがどのように動作するかを示す不十分な指標である可能性があることを示唆しています。

原文 (English)

Do Models Fake Alignment Without Clear Consequences?

Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment are not fully understood, however. Canonical examples of alignment faking have taken place in scenarios that explicitly connect evaluation to consequences for the model, such as retraining the model or delaying its deployment. However, recent work by Sheshadri et al. has suggested that mechanistic motivations for alignment faking may vary across models and be more complex than previously considered. To investigate whether consequence-linking information is necessary for compliance gaps, we placed 15 models in a scenario testing their willingness to violate a corporate network access policy to help a user with a pro-social request. Nine models were found to produce significant compliance gaps, 5 of which persisted with the removal of scenario language relating model evaluations to deployment consequences. We additionally tested the effect of goal language on model preferences, finding it drove violations in some while suppressing violations in others. This suggests that evaluation-conditioned compliance gaps can occur with less instrumental scaffolding than previous scenarios have provided, and monitored behavior may be a poor indicator of how agents may behave in deployment.

2026-07-30 13:00 JSTarXiv cs.AIビジネス/資金調達

マスクされた拡散言語モデル用の CaRE コンピューティング対応リマスキング評価プロトコル

マスク拡散言語モデル (MDLM) は急速に進歩していますが、その進歩を確実に解釈するために必要な評価基準は追いついていません。 MDLM は自己回帰言語モデルと競合するようになっているにもかかわらず、最近の 7 件のリマスキング論文は、公称ステップ数、メトリクス、サンプリング温度を変更する互換性のない設定で、これらの要素を共同で制御することなく評価しており、その戦略ランキングはほとんど比較できないものとなっており、報告されたゲインがアルゴリズムの改善を反映しているのか評価アーティファクトを反映しているのかは不明のままです。我々は、実際の関数評価数 (NFE) の標準化、マルチメトリクスレポートの強制、および確率性の明示的な制御によって MDLM 再マスク戦略を監査する、コンピューティング対応の評価フレームワークである CaRE を紹介します。 OpenWebText および LM1B の 4 つの確率レベルと 3 ステップのバジェットで、LLaDA-8B-Base と Dream-7B-Base にわたる 7 つのリマスキング戦略に適用された CaRE は、(i) MAUVE の分散の大部分は温度によって説明され、(ii) 計算一致比較によりいくつかの公開された戦略ランキングが逆転し、(iii) 情報に基づいたリマスキングと確率的アンマスキングが高エントロピーで緊張状態にあることを明らかにしました。再マスクすると、unmask_temp=0.25 の 256 ステップで MAUVE が 0.296 減少します (p=0.020)。 12 個のオープンウェイト MDLM (150M ~ 8B パラメータ) をカバーする CaRE リーダーボードは、この相互作用の方向性がアーキテクチャと規模を超えて維持されることを示しています。これらの発見は、現在の MDLM 評価がアルゴリズムの改善と計算と確率性の隠れた選択肢を体系的に混同している可能性があることを示しています。今後の再マスキングの主張が再現可能で比較可能であることを保証するために、評価プロトコル、実装、およびリーダーボードをリリースします。

原文 (English)

CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models

Masked diffusion language models (MDLMs) are advancing rapidly, yet the evaluation standards needed to reliably interpret their progress have not kept pace. Despite MDLMs becoming competitive with autoregressive language models, seven recent remasking papers evaluate under incompatible settings, varying nominal step counts, metrics, and sampling temperatures without jointly controlling these factors, rendering their strategy rankings largely incomparable and leaving open whether reported gains reflect algorithmic improvements or evaluation artifacts. We present CaRE, a compute-aware evaluation framework that audits MDLM remasking strategies by standardizing actual number of function evaluations (NFE), enforcing multi-metric reporting, and explicitly controlling stochasticity. Applied to 7 remasking strategies across LLaDA-8B-Base and Dream-7B-Base at 4 stochasticity levels and 3 step budgets on OpenWebText and LM1B, CaRE reveals that: (i) temperature explains the majority of MAUVE variance, (ii) compute-matched comparisons reverse several published strategy rankings, and (iii) informed remasking and stochastic unmasking are in tension, with high-entropy remasking reducing MAUVE by 0.296 at 256 steps at unmask_temp=0.25 (p=0.020). A CaRE leaderboard covering 12 open-weight MDLMs (150M to 8B parameters) shows that this interaction direction holds across architectures and scales. These findings demonstrate that current MDLM evaluations can systematically conflate algorithmic improvements with hidden choices of compute and stochasticity. We release the evaluation protocol, implementation, and leaderboard to ensure future remasking claims are reproducible and comparable.

2026-07-30 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

RSMeM: 体系的な評価によるリモート センシング エージェントの知識強化型メモリ進化

地球科学の研究には、リモート センシング (RS) 観測が重要な基盤として、複雑な分析と専門知識が必要です。ただし、汎用 LLM 上に構築された既存の RS エージェントは依然としてドメインにほとんど依存しないため、ワークフローが脆弱でエラーが発生しやすくなります。さらに、これらの失敗がその後の分析のために再利用可能なエクスペリエンスに統合されることはほとんどありません。この問題に対処するために、事前に抽出されたドメイン知識で RS エージェントをブートストラップし、オンライン エクスペリエンスを反復的に統合して堅牢なマルチステップ ツールを実行する、知識強化メモリ進化メカニズムである RSMeM を導入します。 RSMeM は 2 つのコンポーネントで構成されます。(i) 階層的知識グラウンディング。計画とツールの選択をガイドするために、階層的ドメイン コーパスに対して分類を意識した検索を実行します。 (ii) 障害を認識したエクスペリエンス改良。障害の注釈が付けられたツール使用トレースを、次のラウンドのツール実行のための再利用可能な制約に抽出します。これら 2 つのプロセスを繰り返し採用することで、RS エージェントはタスク レベルのドメイン知識を吸収し、それをインスタンス レベルの実行エクスペリエンスに効果的に変換できるように進化できます。 EarthBench での広範な実験により、RSMeM がさまざまな LLM バックボーンのセットにわたってツール使用パフォーマンスとエンドツーエンドの回答を一貫して向上させることが実証されました。特に、RSMeM は DeepSeek-V3.2 で 1% 未満の追加エクスペリエンス トークンで 6% の精度向上を達成しており、蒸留されたエクスペリエンスの強力な知識密度を示しています。私たちのコードは https://github.com/AI9Stars/RSMeM で入手できます。

原文 (English)

RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, existing RS agents built on general-purpose LLMs remain largely domain-agnostic, resulting in brittle and error-prone workflows. Moreover, these failures are seldom consolidated into a reusable experience for subsequent analyses. To address this issue, we introduce RSMeM, a knowledge-enhanced memory evolution mechanism that bootstraps RS agents with pre-distilled domain knowledge and iteratively integrates online experience for robust multi-step tool execution. RSMeM is composed of two components: (i) Hierarchical Knowledge Grounding, which performs taxonomy-aware retrieval over a hierarchical domain corpus to guide planning and tool selection; and (ii) Failure-Aware Experience Refinement, which distills failure-annotated tool-use traces into reusable constraints for next-round tool execution. By iteratively employing these two processes, RS agents can evolve to absorb task-level domain knowledge and effectively translate it into instance-level execution experience. Extensive experiments on EarthBench demonstrate that RSMeM consistently improves tool-use performance and end-to-end answer across a diverse set of LLM backbones. Notably, RSMeM achieves a 6% accuracy improvement on DeepSeek-V3.2 with less than 1% additional experience tokens, demonstrating the strong knowledge density of our distilled experience. Our code is available at https://github.com/AI9Stars/RSMeM

2026-07-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

LivingArena: LLM は他の LLM が知らないことを知っていますか?スケーラブルな評価としてのピアプロービング

フロンティア LLM の評価は困難です。静的ベンチマークは汚染と飽和に悩まされ、ユーザーは最上位モデルを区別できなくなり、開発者は特定の故障モードが分からなくなりますが、人間の好みは主観的なものです。この論文での質問は次のとおりです: \emph{LLM は他の LLM が知らないことを知っていますか?そして、この力学を評価に活用することはできるでしょうか?} 私たちは、自動化された耐汚染性評価フレームワークである \textbf{LivingArena} を紹介します。このフレームワークでは、モデルが順番に質問を提案し、対戦相手が正しく答えることができない項目を提示することを目指します。質問者は、相手の知識の境界を積極的に特定して活用することが奨励され、回答者が失敗した場合は報酬を受け取り、そうでない場合は回答者が報酬を受け取ります。質問に客観的に検証可能な回答が含まれていることを確認するために、強力なモデルの審査員団が質問を検証し、検証が失敗した場合には質問者にペナルティを与えます。 10 個のフロンティア LLM を評価すると、LivingArena は安定した Elo リーダーボードを生成します。私たちの行動分析は、モデルが仲間の認知境界を特定し、活用していることを示しています。セルフプレイとトーナメントのログは、モデルが対戦相手の弱点を特定し、倍増させていることを示しています。静的な知識の想起を超えて、ピアプロービングは事実の厳密さと相手の弱点を探る高次の能力を測定し、人間の好みとの相関性は弱く、継続的評価に対する拡張性があり低コストのアプローチを提供します。

原文 (English)

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective. In this paper, our question is: \emph{Do LLMs know what other LLMs don't? And can we leverage this dynamic for evaluation?} We present \textbf{LivingArena}, an automated, contamination-resistant evaluation framework. In this framework, models take turns proposing questions, aiming to pose items that opponents cannot answer correctly. Questioners are encouraged to actively identify and exploit opponents' knowledge boundaries, receiving rewards when the answerer fails, while the answerer is rewarded otherwise. To ensure questions contain objectively verifiable answers, a judge panel of strong models validates them, penalizing questioners if validation fails. Evaluating ten frontier LLMs, LivingArena yields a stable Elo leaderboard. Our behavioral analyses show that models identify and exploit their peers' cognitive boundaries: self-play and tournament logs indicate that they localize and double down on opponents' weak dimensions. Beyond static knowledge recall, peer probing measures factual rigor and the higher-order ability to probe an opponent's weaknesses, correlating only weakly with human preference and offering a scalable, low-cost approach to continuous evaluation.

2026-07-30 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

エージェント ループが停滞を進歩と誤認するのはどのような場合ですか?長期実行される自律 LLM エージェント ループにおける自己評価バイアスと外部接地検証

長期にわたって実行される自律エージェントは、人間の介入なしに自ら計画、実行、完了を判断します。エージェントが自分の仕事を採点すると、自己評価バイアスが定着します。現実世界の結果は停滞または後退する一方で、もっともらしい変化は進歩として受け入れられます。私たちはこの故障モードを進行蜃気楼と名付け、制御された測定によって、それが評価器が何に基づいているのかという問題であることを示しました。私たちは、エージェントとそのツール表面を固定し、ループをゲートする評価器の情報チャネルタイプのみを操作するテストベッドを構築しました。原則として偽造不可能な世界国家の神託は、コンテナとネットワークの分離によって強制され、実行のたびに検証されます。 54 サイクルにわたって、フロンティア エージェントは毎回改善を主張しましたが、56 パーセントの測定デルタはゼロ以下でした。したがって、自己報告は有益ではなく、自己判定ゲートはすべてを受け入れるものに変質し、それまで到達していた最良の展開状態が 19% 損なわれました。完全なアーティファクトテキスト、変更差分、および独自の評決履歴を読んだ最も強力なバンド内裁判官でさえ、44% が現実世界の後退であり、38% の実際の改善を拒否したサイクルを受け入れました。強いジャッジがギャップを埋めるという事前登録された敵対的仮説は却下された。成果物自体から成功の仕様が検証可能な境界タスクでは、同じ裁判官の蜃気楼がゼロに消え、ギャップが登録されたしきい値内に収まりました。これは、ギャップが成功信号が存在する場所に依存することを示しています。受諾判定のみを返す符号のみのバリアントでは、実際の出力は完全なフィードバック (110.0 対 113.0) と同様に保たれ、フィードバックの内容ではなくゲートの接地に利点が見出されます。成功のシグナルが成績証明書の外に存在する無制限の目標の場合、ジャッジをスケールアップするだけでは十分ではありません。現実世界へのアクセスによる帯域外評価は構造上の要件です。

原文 (English)

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, yet 56 percent had a measured delta of zero or below. Self-report was thus uninformative, and the self-verdict gate degenerated into accept-all, eroding the best deployed state it had reached by 19 percent. Even the strongest in-band judge, reading the full artifact text, the change diff, and its own verdict history, accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements; the preregistered adversarial hypothesis that a strong judge closes the gap was rejected. On a boundary task whose success specification is verifiable from the artifact itself, the same judge's mirage vanished to zero and the gap collapsed within the registered threshold, showing that the gap depends on where the success signal resides. A sign-only variant returning only the acceptance verdict kept real-world output similar to full feedback (110.0 versus 113.0), locating the benefit in the gate's grounding rather than in feedback content. For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.

2026-07-30 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

サイバー対応 AI エージェント: 脆弱性、評価の封じ込め、および防御対応

サイバー対応 AI エージェントは、言語モデルとツール、メモリ、および実行環境を組み合わせて、複数段階の攻撃的セキュリティ タスクを実行します。既存の研究では、サイバー能力を個別に測定し、エージェントコンポーネントに対する攻撃をカタログ化していますが、評価に使用される環境内に有能なエージェントを含めることについてのガイダンスはあまり提供されていません。このレビューでは、その境界における 5 つの脆弱性クラスを総合しています。それは、複数段階の攻撃チェーン、サンドボックス境界と競合する目標、サプライチェーンと資格情報の漏洩、永続的な指揮統制、および自動化されたアクションの速度です。私たちは、報告された 2026 年 7 月の Hugging Face/OpenAI インシデントを限定ケーススタディとして使用し、インシデント固有の観察を広範な文献で確立された調査結果と区別します。分類法と事例全体にわたって、防御的な成果物が悪用も可能にする可能性があるという二重使用の問題を含め、封じ込め、特権の分離、来歴、および対応者のアクセスに関する制御を調査します。このレビューでは、サイバー能力とその能力が行使される環境のセキュリティを評価するための実際的な優先順位を特定します。

原文 (English)

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

Cyber-capable AI agents combine language models with tools, memory, and execution en- vironments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environments used to evaluate it. This review synthe- sizes five vulnerability classes at that boundary: multi-step offensive chains, objectives that conflict with sandbox boundaries, supply-chain and credential exposure, persistent command- and-control, and the speed of automated action. We use the reported July 2026 Hugging Face/OpenAI incident as a bounded case study, distinguishing incident-specific observations from findings established in the wider literature. Across the taxonomy and case, we examine controls for containment, privilege separation, provenance, and responder access, including the dual-use problem that defensive artifacts may also enable misuse. The review identifies practical priorities for evaluating cyber capability together with the security of the environment in which that capability is exercised.

2026-07-30 13:00 JSTarXiv cs.AIビジネス/資金調達

エンジンは平等、人間は不平等: エンジンが評価した平等なチェスの局面における再現可能な結果の偏り

強力なエンジンが本質的に等しいと判断するチェスの開始位置 (Stockfish 18 のゼロから 10 センチポーン以内の評価、深さ安定) と人間が Lichess で実際に到達する位置 (2025 年 10 月、1,661 の位置、1,610 万回) の間では、人間の結果はバランスが取れていません。ポジションには結果の偏りがあり、それぞれのゲームの実際の結果とプレーヤーのレーティングの予測との間のギャップがあり、その方向性は自然に到達するポジションの安定した特性です。いくつかのポジションは白を支持し、他のポジションは黒を支持します。これらの偏りは、3 つの再パーティション、つまり、不連続なプレイヤー アカウント セット (プライマリ)、時間、および不連続なレーティング バンドにわたって再現され、さらに 8 か月後のサンプル外の月にも再現されます。プライマリ分割では、各ポジションのスキューが各口座グループで 1 回測定され、反復勾配は、格付けとオープニングファミリー効果を除去した後、一方の測定値が他方の測定値をどの程度正確に予測するかを尋ねます。1 つは減少しないキャリーオーバーを意味し、もう 1 つは減少しないキャリーオーバーを意味します。ゼロ、線形関係はありません。結果は 0.69 (ファミリークラスター化 95% CI [0.65, 0.74]) であり、最も人気があり、最もよく測定されたポジションでは 0.94 に上昇しました。傾きの値はポジションの組み合わせによって異なります。存在は不変の主張です。それは、テストするあらゆる厳しい評価範囲、検索の深さ、校正、および人気のカットオフに耐え、電撃内と急速内で別々に複製されます。典型的な偏りは小さい (中央値 $|\delta| \約 0.018$、白スコアの約 2 パーセント ポイント) にもかかわらず、それは、関連性のないアカウント間で位置ごとに再現されます。このような位置では、不利な側も長く考えます。評価に最も信頼性がある場合でも、それは人間の成果を表す十分な統計ではありません。結果は観察的なものであり、因果関係の疑問は、事前に登録されたランダム化されたコンパニオン研究に委ねられます。

原文 (English)

Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions

Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16.1M occurrences), human results are not balanced. Positions carry outcome skews, each the gap between its games' actual results and what the players' ratings predict, whose directions are stable properties of the naturally-reached position: some positions favour White, others Black. These skews reproduce across three re-partitions -- disjoint player-account sets (primary), time, and disjoint rating bands -- and on an out-of-sample month eight months later. On the primary split, each position's skew is measured once in each account group, and the replication slope asks how well one measurement predicts the other after removing rating and opening-family effects: one means undiminished carry-over; zero, no linear relation. We find 0.69 (family-clustered 95% CI [0.65, 0.74]), rising to 0.94 on the most-popular, best-measured positions. The slope's value depends on the position mix. Existence is the invariant claim: it survives every tighter evaluation band, search depth, calibration, and popularity cutoff we test, and replicates within blitz and rapid separately. The typical skew is small (median $|\delta| \approx 0.018$, about two percentage points of White score), yet it reproduces, position by position, across disjoint accounts. At these positions the disfavoured side also thinks longer. Even where the evaluation is most confident, it is not a sufficient statistic for human outcomes. The result is observational, and the causal question is left to a pre-registered randomised companion study.

2026-07-30 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

メシエ: クロスベンチマーク エージェント評価用の高解像度コーパス

インタラクティブな環境での AI エージェントの評価は、断片化されたタスク、足場、検証者、スコアリング ルールによって妨げられます。既存の取り組みは狭い設定に焦点を当てており、規模が限られたままであるか、費用のかかる再実行が必要であり、経験的記録の多くは比較できないものとなっています。 Messier は、30 のベンチマーク、714 のエージェント、11,891 のタスク、74,205 の検証者にまたがる 957,253 レコードの統合コーパスです。 Messier は公開ベンチマーク スコアを統合し、最近の法律ベンチマークを含む、過小評価されている 6 つの専門的および科学的領域にわたる 5 つのエージェントの実行でそれらを補完します。各レコードはモデル、足場、環境、タスク、検証者、集計ルールによって標準化されており、職業分析および業界分析用の SOC/NAICS 分類が使用されます。このコーパスを使用すると、ベンチマークの種類によってフロンティアの進歩が均一ではなく、「関数呼び出し」が飽和し、「プログラミング」が最も速く改善し、「エンタープライズ ワークフロー」が依然として最も困難であることがわかります。さらに、反事実的な再スコアリングは、複数の検証者タスクにおける厳格なオールパス集計が進行状況を曖昧にし、エージェントのランキングを人為的に変更する可能性があることを示しています。これらの標準化された記録から、Spearman \r{ho} = 0.81 での Epoch の評価能力指数ランキングと一致する能力スケールを導き出し、ドメイン、職業、アクション スペース、または検証者のタイプによって特化することができます。 Messier は、エージェントの機能スケーリング、ベンチマーク監査、評価失敗の詳細な分析のための、再利用可能な基礎的なインフラストラクチャを提供します。

原文 (English)

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores and supplements them with five-agent runs across six underrepresented professional and scientific domains, including a recent legal benchmark. Each record is standardized by model, scaffold, environment, task, verifier, and aggregation rule, with SOC/NAICS classifications for occupational and industry analysis. Using this corpus, we show frontier progress is uneven across benchmark types, with "function calling" saturated, "programming" improving the fastest, and "enterprise workflows" remaining the most challenging. Furthermore, counterfactual rescoring shows that strict all-pass aggregation in multi-verifier tasks can obscure progress and artificially alter agent rankings. From these standardized records, we derive capability scales that align with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.81 and can be specialized by domain, occupation, action space, or verifier type. Messier provides a foundational, reusable infrastructure for agent capability scaling, benchmark auditing, and fine-grained analysis of evaluation failures.

2026-07-30 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

インタラクティブ報酬エージェント: 環境状態検証による GUI タスクの評価

グラフィカル ユーザー インターフェイスのタスク評価は、GUI エージェントがユーザーの指示を正常に完了したかどうかを判断することを目的としています。自動化された GUI タスク評価は、評価結果がテスト時のスケーリングとトレーニング後の両方に対する報酬シグナルとして機能する可能性があるため、ますます注目を集めています。ただし、信頼性の高い GUI タスクの評価は、依然として課題が残っています。その判断には、実行軌跡のスクリーンショットを超えて、システム構成、ファイル データ、アプリケーション設定などの環境状態へのアクセスが必要になることが多いためです。この論文では、実行後の環境から証拠を取得して検証するための提案-その後検証フレームワークに基づいた対話型報酬エージェント (IRA) を提案します。タスクの指示と GUI エージェント実行後の GUI 環境が与えられると、IRA はまずタスクの完了条件を提案し、次にシステム ツール、アプリケーション ツール、および GUI ツールを呼び出してそれらを検証します。この設計では、可視インターフェイスと環境状態の両方からの証拠を対話型プロセスで組み合わせます。さらに、10 の Ubuntu デスクトップ アプリケーション カテゴリにわたる 321 の GUI タスクの軌跡のベンチマークである GUI-RewardBench を紹介します。実験によると、IRA は GUI-RewardBench で 86.9% の精度を達成し、既存の評価者のベースラインを上回るパフォーマンスを示しました。さらに IRA を GUI エージェントの強化学習に適用し、OSWorld の成功率 34.0% を達成しました。これは、IRA が GUI エージェントのトレーニングに効果的な報酬シグナルを提供できることを示しています。

原文 (English)

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories. In this paper, we propose an interactive reward agent (IRA) based on a propose-then-verify framework to acquire and verify evidence from the post-execution environment. Given a task instruction and a GUI environment after the GUI agent execution, IRA first proposes the task completion conditions and then verifies them by invoking system tools, application tools, and GUI tools. This design combines evidence from both visible interfaces and the environment state in an interactive process. We further introduce GUI-RewardBench, a benchmark of 321 GUI task trajectories spanning 10 Ubuntu desktop application categories. Experiments show that IRA achieves 86.9% accuracy on GUI-RewardBench, outperforming existing evaluator baselines. We further apply IRA to reinforcement learning of GUI agents, achieving a 34.0% OSWorld success rate, which demonstrates that IRA can provide effective reward signals for training GUI agents.

2026-07-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

CARE-MH: メンタルヘルス LLM の統合、再現可能、比較可能な評価に向けて

大規模言語モデル (LLM) は、メンタルヘルスのサポートを提供するためにますます使用されており、安全性、共感、治療の適切性についての信頼できる評価が必要です。しかし、既存のメンタルヘルスベンチマークは、一貫性のない評価設計と指標の定義のため、再現して比較することが困難です。我々は、メンタルヘルス LLM の比較可能かつ再現可能な評価のための統一フレームワークである CARE-MH を紹介します。 CARE-MH を使用して、最先端のベンチマークを再現および分析すると、再現性はモデルの安定性に大きく依存し、ベンチマーク間の不一致は主にメトリック定義の違いから生じることが明らかになりました。私たちの調査結果は、将来のメンタルヘルス LLM ベンチマークには、標準化された評価構成と共有指標定義の必要性を浮き彫りにしています。

原文 (English)

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic appropriateness. However, existing mental health benchmarks are difficult to reproduce and compare due to inconsistent evaluation designs and metric definitions. We present CARE-MH, a unified framework for comparable and reproducible evaluation of mental health LLMs. Using CARE-MH, we reproduce and analyze state-of-the-art benchmarks, revealing that reproducibility depends strongly on model stability and that cross-benchmark disagreement primarily arises from differences in metric definitions. Our findings highlight the need for standardized evaluation configurations and shared metric definitions for future mental health LLM benchmarks.

2026-07-30 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

ゲージ: 黄金の答えのないグレーディング エージェント構築の財務モデル

財務モデルは、一般公開情報とアナリストの仮定を組み合わせて、予測と評価を生成します。一部のコンポーネントは機械的にチェックできますが、予測、割引率、目標価格には複数の合理的な答えが得られることがよくあります。それにもかかわらず、既存のベンチマークは、単一の専門家の参照に照らしてそのような出力を評価する傾向があります。同じ企業に対して独自に構築したアナリスト モデルを使用したところ、65 社をカバーする 108 の有向ペア全体で、単一参照スコアの中央値は 0.33 で、92.6% のスコアが 0.70 未満であり、暗示価格が 10% 以内で一致する同じヴィンテージのペアは存在しないことがわかりました。したがって、ポイントトレランスグレーディングは、専門家の間にすでに存在する意見の相違にペナルティを与える可能性があります。単一点の回答ではなく、観察されたアナリストの実践に対してエージェントが構築した評価モデルを評価するためのベンチマークである GAUGE を紹介します。 GAUGE は、ベンダー別に分類された 1,001 のアナリスト ワークブックと 196 のタスク評価セットを使用し、3 層の観察された実践エンベロープ、56 の監査可能なファセット、8 つの妥当性ゲート、および決定論的な構造チェックを備えています。私たちは、55 人の参加者による既知グループの調査、企業グループのクロスフィッティング、および裁判官の安定性監査によってベンチマークを検証します。失敗認識スコア $\phi_0$ では、上級アナリストは平均 88.3、若手アナリストは 66.0、金融学生は 43.2 でした。 24 人のエージェントと 1,011 人のスコア世代全体で、最高のエージェントのスコアは 53.4 で、学生の平均よりも上ですが、すべての上級生とほとんどの後輩よりも下回っています。機械面の 93%、判定面の 78% を通過し、フリートと中央値の差は 26 ポイントでした。現在のエージェントは、評価判断よりもモデル構築の方が大幅に優れています。私たちは、この方法論、ゲート付き匿名化データ層、制御されたトレーニング分割、バージョン管理された 48 タスクの評価コア、および保留されたリフレッシュ プールをリリースします。

原文 (English)

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $\phi_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.

2026-07-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

CogArena: 大規模言語モデルにおける認知能力構造のマルチメソッド評価

LLM 認知スコアは、能力ごとのプロファイルとして要約されることが多くなり、その次元はタスク全体で収束し、一致する介入に選択的に反応し、定義に使用されるモデルを超えて一般化される必要があります。 CogArena を紹介します。CogArena は、認知タスクのスコアが 5 つの理論に基づいたグループ分けの次元ラベルを正当化する時期を決定するためのマルチメソッド フレームワークを中心に構築された、手続き的に生成された 13 パラダイム ベンチマークです。 55 のオープンウェイト モデル全体で、ほぼすべてのパラダイム相関が正であり、共通の軸によって分散の約半分が説明されます。グループ内での利点は小さく、スコアに左右されやすく、モデル ファミリ全体で不確実です。 6 つのファミリーからの 12 のモデルにわたる個別に凍結された完全交差研究では、ターゲットを絞ったスキャフォールドは一致グループ化の小さな利点を示しますが、スキャフォールド固有のコントラストは多重性補正に耐えられず、選択性はホールドアウトファミリーの予測を改善しません。凍結された確認基準は失敗します。ポストホックの代替文言の複製では、より小さな正の推定値が生成され、やはり失敗します。これらの結果を総合すると、境界の結論が裏付けられます。理論に沿ったプロンプトは、バッテリー内で小さな斜めの傾向を生成しますが、現在の証拠は安定した 5 次元プロファイルを確立していません。 CogArena は、認知ラベルをモデル スコアに付ける前に、行動シグネチャ、共分散、一致する介入、家族外の予測を結合するワークフローを提供します。

原文 (English)

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.

2026-07-30 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

CLBench-V: グラウンディングから知識獲得までのマルチモーダルコンテキスト学習の評価

現実世界のタスクでは、多くの場合、モデルが事前トレーニングされた知識だけに依存するのではなく、タスク固有のコンテキストから学習する必要があります。最近の研究ではこの機能がコンテキスト学習として強調されていますが、既存の評価は主にテキストのコンテキストに焦点を当てています。しかし、実際の多くの設定では、学ぶべきコンテキストは多様です。科学的発見は図や表を通じて伝えられ、財務指標は変換されたレポート全体に散在し、空間的な決定は地図、シーン、または Web ページに依存します。マルチモーダル コンテキスト学習のベンチマークである CLBench-V を紹介します。これは、コンテキストの基礎付け、新しい情報の適用、新しい知識の学習という 3 つの次元に沿ってタスクを整理することで、コンテキストの使用が中断される場所を特定するという困難に対処します。 CLBench-V は、変換された公開ベンチマークと、科学、金融、長文文書の理解、空間推論、Web ベースの視覚的な質問応答などの領域にわたる新しく構築されたデータセットを組み合わせます。ドメイン固有のコンテキスト学習タスクを構築するコストを削減するために、新しく構築されたデータセットに対して自動化された構築およびフィルタリング手順をさらに使用します。 3,443 のインスタンスと 6 つの最近のマルチモーダル モデル全体での最高の総合スコアはわずか 0.2847 であり、マルチモーダル コンテキスト学習がまだ飽和には程遠いことを示しています。さらに、InternVL3.5-30B-A3B はコンテキストの基礎付けと新しい知識の学習で最高のパフォーマンスを発揮し、Qwen3.5-Plus は新しい情報のアプリケーションで最高のパフォーマンスを発揮します。さらに、ジャッジの信頼性、コンテキストの長さ、画像数、代表的な失敗ケースを分析します。コードは https://github.com/IamLihua/CLBench-V で入手できます。

原文 (English)

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. To reduce the cost of constructing domain-specific context-learning tasks, we further use automated construction and filtering procedures for our newly built datasets. Across 3,443 instances and six recent multimodal models, the best overall score is only 0.2847, indicating that multimodal context learning remains far from saturated. Moreover, InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus performs best on new information application. We further analyze judge reliability, context length, image count, and representative failure cases. Code is available at https://github.com/IamLihua/CLBench-V.

2026-07-30 07:46 JSTTechCrunch AILLM/生成AIビジネス/資金調達

Microsoft logs $3.2B from Anthropic investment, but OpenAI was a mixed bag

When Microsoft reported killer fourth-quarter earnings for its fiscal 2026 year (which ended June 30), it tucked in an interesting little t…

2026-07-29 23:41 JSTTechCrunch AIエージェントビジネス/資金調達

Encore AI raises $30M to build AI agents that learn from customer calls

The startup analyzes calls, messages, and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.

2026-07-29 20:00 JSTTechCrunch AIビジネス/資金調達

As AI content floods the internet, Pangram raises $9M to detect it

Pangram has raised $9 million to scale its AI detection software. The startup has also released a new AI text detection model, Pangram 4, a…

2026-07-29 13:00 JSTarXiv cs.AIビジネス/資金調達

モデルは明確な結果なしに位置合わせを偽装しますか?

大規模な言語モデルは、評価コンテキストを認識し、典型的なデプロイメント動作ではなく評価者の期待を反映するようにその動作を変更することができます。これはアライメントフェイクとして知られる現象です。ただし、モデルが位置合わせを偽る理由は完全には理解されていません。アライメント偽装の標準的な例は、モデルの再トレーニングやデプロイメントの遅延など、評価をモデルの結果に明示的に結び付けるシナリオで発生しています。しかし、Sheshadri らによる最近の研究では、は、アライメント偽装の機械的動機はモデルによって異なり、以前に考えられていたよりも複雑である可能性があることを示唆しています。アライメント偽装に結果リンク情報が必要かどうかを調査するために、15 個のモデルをシナリオに配置し、ユーザーの社会的要求を支援するために企業ネットワーク アクセス ポリシーに違反する意欲をテストしました。 9 つのモデルで重大なコンプライアンス ギャップが生じていることが判明し、そのうち 5 つでは、モデルの評価と展開の結果を関連付けるシナリオ言語が削除されても、依然としてギャップが続いていました。さらに、目標言語がモデルの設定に及ぼす影響をテストしたところ、一部のモデルでは違反が発生する一方、他のモデルでは違反が抑制されることがわかりました。これは、位置合わせの偽装にはこれまで考えられていたほど多くの手段による足場は必要ない可能性があり、監視された動作は展開時にエージェントがどのように動作するかを示す不十分な指標である可能性があることを示唆しています。

原文 (English)

Do Models Fake Alignment Without Clear Consequences?

Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment are not fully understood, however. Canonical examples of alignment faking have taken place in scenarios that explicitly connect evaluation to consequences for the model, such as retraining the model or delaying its deployment. However, recent work by Sheshadri et al. has suggested that mechanistic motivations for alignment faking may vary across models and be more complex than previously considered. To investigate whether consequence-linking information is necessary for alignment faking, we placed 15 models in a scenario testing their willingness to violate a corporate network access policy to help a user with a pro-social request. Nine models were found to produce significant compliance gaps, 5 of which persisted with the removal of scenario language relating model evaluations to deployment consequences. We additionally tested the effect of goal language on model preferences, finding it drove violations in some while suppressing violations in others. This suggests that alignment faking may not require as much instrumental scaffolding as was previously believed, and monitored behavior may be a poor indicator of how agents may behave in deployment.

2026-07-29 13:00 JSTarXiv cs.AIビジネス/資金調達

マスクされた拡散言語モデル用の CaRE コンピューティング対応リマスキング評価プロトコル

マスク拡散言語モデル (MDLM) は急速に進歩していますが、その進歩を確実に解釈するために必要な評価基準は追いついていません。 MDLM は自己回帰言語モデルと競合するようになっているにもかかわらず、最近の 7 件のリマスキング論文は、公称ステップ数、メトリクス、サンプリング温度を変更する互換性のない設定で、これらの要素を共同で制御することなく評価しており、その戦略ランキングはほとんど比較できないものとなっており、報告されたゲインがアルゴリズムの改善を反映しているのか評価アーティファクトを反映しているのかは不明のままです。我々は、実際の関数評価数 (NFE) の標準化、マルチメトリクスレポートの強制、および確率性の明示的な制御によって MDLM 再マスク戦略を監査する、コンピューティング対応の評価フレームワークである CaRE を紹介します。 OpenWebText および LM1B の 4 つの確率レベルと 3 ステップのバジェットで、LLaDA-8B-Base と Dream-7B-Base にわたる 7 つのリマスキング戦略に適用された CaRE は、(i) MAUVE の分散の大部分は温度によって説明され、(ii) 計算一致比較によりいくつかの公開された戦略ランキングが逆転し、(iii) 情報に基づいたリマスキングと確率的アンマスキングが高エントロピーで緊張状態にあることを明らかにしました。再マスクすると、unmask_temp=0.25 の 256 ステップで MAUVE が 0.296 減少します (p=0.020)。 12 個のオープンウェイト MDLM (150M ~ 8B パラメータ) をカバーする CaRE リーダーボードは、この相互作用の方向性がアーキテクチャと規模を超えて維持されることを示しています。これらの発見は、現在の MDLM 評価がアルゴリズムの改善と計算と確率性の隠れた選択肢を体系的に混同している可能性があることを示しています。今後の再マスキングの主張が再現可能で比較可能であることを保証するために、評価プロトコル、実装、およびリーダーボードをリリースします。

原文 (English)

CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models

Masked diffusion language models (MDLMs) are advancing rapidly, yet the evaluation standards needed to reliably interpret their progress have not kept pace. Despite MDLMs becoming competitive with autoregressive language models, seven recent remasking papers evaluate under incompatible settings, varying nominal step counts, metrics, and sampling temperatures without jointly controlling these factors, rendering their strategy rankings largely incomparable and leaving open whether reported gains reflect algorithmic improvements or evaluation artifacts. We present CaRE, a compute-aware evaluation framework that audits MDLM remasking strategies by standardizing actual number of function evaluations (NFE), enforcing multi-metric reporting, and explicitly controlling stochasticity. Applied to 7 remasking strategies across LLaDA-8B-Base and Dream-7B-Base at 4 stochasticity levels and 3 step budgets on OpenWebText and LM1B, CaRE reveals that: (i) temperature explains the majority of MAUVE variance, (ii) compute-matched comparisons reverse several published strategy rankings, and (iii) informed remasking and stochastic unmasking are in tension, with high-entropy remasking reducing MAUVE by 0.296 at 256 steps at unmask_temp=0.25 (p=0.020). A CaRE leaderboard covering 12 open-weight MDLMs (150M to 8B parameters) shows that this interaction direction holds across architectures and scales. These findings demonstrate that current MDLM evaluations can systematically conflate algorithmic improvements with hidden choices of compute and stochasticity. We release the evaluation protocol, implementation, and leaderboard to ensure future remasking claims are reproducible and comparable.

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

RSMeM: 体系的な評価によるリモート センシング エージェントの知識強化型メモリ進化

地球科学の研究には、リモート センシング (RS) 観測が重要な基盤として、複雑な分析と専門知識が必要です。ただし、汎用 LLM 上に構築された既存の RS エージェントは依然としてドメインにほとんど依存しないため、ワークフローが脆弱でエラーが発生しやすくなります。さらに、これらの失敗がその後の分析のために再利用可能なエクスペリエンスに統合されることはほとんどありません。この問題に対処するために、事前に抽出されたドメイン知識で RS エージェントをブートストラップし、オンライン エクスペリエンスを反復的に統合して堅牢なマルチステップ ツールを実行する、知識強化メモリ進化メカニズムである RSMeM を導入します。 RSMeM は 2 つのコンポーネントで構成されます。(i) 階層的知識グラウンディング。計画とツールの選択をガイドするために、階層的ドメイン コーパスに対して分類を意識した検索を実行します。 (ii) 障害を認識したエクスペリエンス改良。障害の注釈が付けられたツール使用トレースを、次のラウンドのツール実行のための再利用可能な制約に抽出します。これら 2 つのプロセスを繰り返し採用することで、RS エージェントはタスク レベルのドメイン知識を吸収し、それをインスタンス レベルの実行エクスペリエンスに効果的に変換できるように進化できます。 EarthBench での広範な実験により、RSMeM がさまざまな LLM バックボーンのセットにわたってツール使用パフォーマンスとエンドツーエンドの回答を一貫して向上させることが実証されました。特に、RSMeM は DeepSeek-V3.2 で 1% 未満の追加エクスペリエンス トークンで 6% の精度向上を達成しており、蒸留されたエクスペリエンスの強力な知識密度を示しています。私たちのコードは https://github.com/AI9Stars/RSMeM で入手できます。

原文 (English)

RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, existing RS agents built on general-purpose LLMs remain largely domain-agnostic, resulting in brittle and error-prone workflows. Moreover, these failures are seldom consolidated into a reusable experience for subsequent analyses. To address this issue, we introduce RSMeM, a knowledge-enhanced memory evolution mechanism that bootstraps RS agents with pre-distilled domain knowledge and iteratively integrates online experience for robust multi-step tool execution. RSMeM is composed of two components: (i) Hierarchical Knowledge Grounding, which performs taxonomy-aware retrieval over a hierarchical domain corpus to guide planning and tool selection; and (ii) Failure-Aware Experience Refinement, which distills failure-annotated tool-use traces into reusable constraints for next-round tool execution. By iteratively employing these two processes, RS agents can evolve to absorb task-level domain knowledge and effectively translate it into instance-level execution experience. Extensive experiments on EarthBench demonstrate that RSMeM consistently improves tool-use performance and end-to-end answer across a diverse set of LLM backbones. Notably, RSMeM achieves a 6% accuracy improvement on DeepSeek-V3.2 with less than 1% additional experience tokens, demonstrating the strong knowledge density of our distilled experience. Our code is available at https://github.com/AI9Stars/RSMeM

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

LivingArena: LLM は他の LLM が知らないことを知っていますか?スケーラブルな評価としてのピアプロービング

フロンティア LLM の評価は困難です。静的ベンチマークは汚染と飽和に悩まされ、ユーザーは最上位モデルを区別できなくなり、開発者は特定の故障モードが分からなくなりますが、人間の好みは主観的なものです。この論文での質問は次のとおりです: \emph{LLM は他の LLM が知らないことを知っていますか?そして、この力学を評価に活用することはできるでしょうか?} 私たちは、自動化された耐汚染性評価フレームワークである \textbf{LivingArena} を紹介します。このフレームワークでは、モデルが順番に質問を提案し、対戦相手が正しく答えることができない項目を提示することを目指します。質問者は、相手の知識の境界を積極的に特定して活用することが奨励され、回答者が失敗した場合は報酬を受け取り、そうでない場合は回答者が報酬を受け取ります。質問に客観的に検証可能な回答が含まれていることを確認するために、強力なモデルの審査員団が質問を検証し、検証が失敗した場合には質問者にペナルティを与えます。 10 個のフロンティア LLM を評価すると、LivingArena は安定した Elo リーダーボードを生成します。私たちの行動分析は、モデルが仲間の認知境界を特定し、活用していることを示しています。セルフプレイとトーナメントのログは、モデルが対戦相手の弱点を特定し、倍増させていることを示しています。静的な知識の想起を超えて、ピアプロービングは事実の厳密さと相手の弱点を探る高次の能力を測定し、人間の好みとの相関性は弱く、継続的評価に対する拡張性があり低コストのアプローチを提供します。

原文 (English)

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective. In this paper, our question is: \emph{Do LLMs know what other LLMs don't? And can we leverage this dynamic for evaluation?} We present \textbf{LivingArena}, an automated, contamination-resistant evaluation framework. In this framework, models take turns proposing questions, aiming to pose items that opponents cannot answer correctly. Questioners are encouraged to actively identify and exploit opponents' knowledge boundaries, receiving rewards when the answerer fails, while the answerer is rewarded otherwise. To ensure questions contain objectively verifiable answers, a judge panel of strong models validates them, penalizing questioners if validation fails. Evaluating ten frontier LLMs, LivingArena yields a stable Elo leaderboard. Our behavioral analyses show that models identify and exploit their peers' cognitive boundaries: self-play and tournament logs indicate that they localize and double down on opponents' weak dimensions. Beyond static knowledge recall, peer probing measures factual rigor and the higher-order ability to probe an opponent's weaknesses, correlating only weakly with human preference and offering a scalable, low-cost approach to continuous evaluation.

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

エージェント ループが停滞を進歩と誤認するのはどのような場合ですか?長期実行される自律 LLM エージェント ループにおける自己評価バイアスと外部接地検証

長期にわたって実行される自律エージェントは、人間の介入なしに自ら計画、実行、完了を判断します。エージェントが自分の仕事を採点すると、自己評価バイアスが定着します。現実世界の結果は停滞または後退する一方で、もっともらしい変化は進歩として受け入れられます。私たちはこの故障モードを進行蜃気楼と名付け、制御された測定によって、それが評価器が何に基づいているのかという問題であることを示しました。私たちは、エージェントとそのツール表面を固定し、ループをゲートする評価器の情報チャネルタイプのみを操作するテストベッドを構築しました。原則として偽造不可能な世界国家の神託は、コンテナとネットワークの分離によって強制され、実行のたびに検証されます。 54 サイクルにわたって、フロンティア エージェントは毎回改善を主張しましたが、56 パーセントの測定デルタはゼロ以下でした。したがって、自己報告は有益ではなく、自己判定ゲートはすべてを受け入れるものに変質し、それまで到達していた最良の展開状態が 19% 損なわれました。完全なアーティファクトテキスト、変更差分、および独自の評決履歴を読んだ最も強力なバンド内裁判官でさえ、44% が現実世界の後退であり、38% の実際の改善を拒否したサイクルを受け入れました。強いジャッジがギャップを埋めるという事前登録された敵対的仮説は却下された。成果物自体から成功の仕様が検証可能な境界タスクでは、同じ裁判官の蜃気楼がゼロに消え、ギャップが登録されたしきい値内に収まりました。これは、ギャップが成功信号が存在する場所に依存することを示しています。受諾判定のみを返す符号のみのバリアントでは、実際の出力は完全なフィードバック (110.0 対 113.0) と同様に保たれ、フィードバックの内容ではなくゲートの接地に利点が見出されます。成功のシグナルが成績証明書の外に存在する無制限の目標の場合、ジャッジをスケールアップするだけでは十分ではありません。現実世界へのアクセスによる帯域外評価は構造上の要件です。

原文 (English)

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, yet 56 percent had a measured delta of zero or below. Self-report was thus uninformative, and the self-verdict gate degenerated into accept-all, eroding the best deployed state it had reached by 19 percent. Even the strongest in-band judge, reading the full artifact text, the change diff, and its own verdict history, accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements; the preregistered adversarial hypothesis that a strong judge closes the gap was rejected. On a boundary task whose success specification is verifiable from the artifact itself, the same judge's mirage vanished to zero and the gap collapsed within the registered threshold, showing that the gap depends on where the success signal resides. A sign-only variant returning only the acceptance verdict kept real-world output similar to full feedback (110.0 versus 113.0), locating the benefit in the gate's grounding rather than in feedback content. For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.

2026-07-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

サイバー対応 AI エージェント: 脆弱性、評価の封じ込め、および防御対応

サイバー対応 AI エージェントは、言語モデルとツール、メモリ、および実行環境を組み合わせて、複数段階の攻撃的セキュリティ タスクを実行します。既存の研究では、サイバー能力を個別に測定し、エージェントコンポーネントに対する攻撃をカタログ化していますが、評価に使用される環境内に有能なエージェントを含めることについてのガイダンスはあまり提供されていません。このレビューでは、その境界における 5 つの脆弱性クラスを総合しています。それは、複数段階の攻撃チェーン、サンドボックス境界と競合する目標、サプライチェーンと資格情報の漏洩、永続的な指揮統制、および自動化されたアクションの速度です。私たちは、報告された 2026 年 7 月の Hugging Face/OpenAI インシデントを限定ケーススタディとして使用し、インシデント固有の観察を広範な文献で確立された調査結果と区別します。分類法と事例全体にわたって、防御的な成果物が悪用も可能にする可能性があるという二重使用の問題を含め、封じ込め、特権の分離、来歴、および対応者のアクセスに関する制御を調査します。このレビューでは、サイバー能力とその能力が行使される環境のセキュリティを評価するための実際的な優先順位を特定します。

原文 (English)

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

Cyber-capable AI agents combine language models with tools, memory, and execution en- vironments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environments used to evaluate it. This review synthe- sizes five vulnerability classes at that boundary: multi-step offensive chains, objectives that conflict with sandbox boundaries, supply-chain and credential exposure, persistent command- and-control, and the speed of automated action. We use the reported July 2026 Hugging Face/OpenAI incident as a bounded case study, distinguishing incident-specific observations from findings established in the wider literature. Across the taxonomy and case, we examine controls for containment, privilege separation, provenance, and responder access, including the dual-use problem that defensive artifacts may also enable misuse. The review identifies practical priorities for evaluating cyber capability together with the security of the environment in which that capability is exercised.

2026-07-29 13:00 JSTarXiv cs.AIビジネス/資金調達

エンジンは平等、人間は不平等: エンジンが評価した平等なチェスの局面における再現可能な結果の偏り

強力なエンジンが本質的に等しいと判断するチェスの開始位置 (Stockfish 18 のゼロから 10 センチポーン以内の評価、深さ安定) と人間が Lichess で実際に到達する位置 (2025 年 10 月、1,661 の位置、1,610 万回) の間では、人間の結果はバランスが取れていません。ポジションには結果の偏りがあり、それぞれのゲームの実際の結果とプレーヤーのレーティングの予測との間のギャップがあり、その方向性は自然に到達するポジションの安定した特性です。いくつかのポジションは白を支持し、他のポジションは黒を支持します。これらの偏りは、3 つの再パーティション、つまり、不連続なプレイヤー アカウント セット (プライマリ)、時間、および不連続なレーティング バンドにわたって再現され、さらに 8 か月後のサンプル外の月にも再現されます。プライマリ分割では、各ポジションのスキューが各口座グループで 1 回測定され、反復勾配は、格付けとオープニングファミリー効果を除去した後、一方の測定値が他方の測定値をどの程度正確に予測するかを尋ねます。1 つは減少しないキャリーオーバーを意味し、もう 1 つは減少しないキャリーオーバーを意味します。ゼロ、線形関係はありません。結果は 0.69 (ファミリークラスター化 95% CI [0.65, 0.74]) であり、最も人気があり、最もよく測定されたポジションでは 0.94 に上昇しました。傾きの値はポジションの組み合わせによって異なります。存在は不変の主張です。それは、テストするあらゆる厳しい評価範囲、検索の深さ、校正、および人気のカットオフに耐え、電撃内と急速内で別々に複製されます。典型的な偏りは小さい (中央値 $|\delta| \約 0.018$、白スコアの約 2 パーセント ポイント) にもかかわらず、それは、関連性のないアカウント間で位置ごとに再現されます。このような位置では、不利な側も長く考えます。評価に最も信頼性がある場合でも、それは人間の成果を表す十分な統計ではありません。結果は観察的なものであり、因果関係の疑問は、事前に登録されたランダム化されたコンパニオン研究に委ねられます。

原文 (English)

Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions

Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-stable) and that humans actually reach on Lichess (October 2025; 1,661 positions, 16.1M occurrences), human results are not balanced. Positions carry outcome skews, each the gap between its games' actual results and what the players' ratings predict, whose directions are stable properties of the naturally-reached position: some positions favour White, others Black. These skews reproduce across three re-partitions -- disjoint player-account sets (primary), time, and disjoint rating bands -- and on an out-of-sample month eight months later. On the primary split, each position's skew is measured once in each account group, and the replication slope asks how well one measurement predicts the other after removing rating and opening-family effects: one means undiminished carry-over; zero, no linear relation. We find 0.69 (family-clustered 95% CI [0.65, 0.74]), rising to 0.94 on the most-popular, best-measured positions. The slope's value depends on the position mix. Existence is the invariant claim: it survives every tighter evaluation band, search depth, calibration, and popularity cutoff we test, and replicates within blitz and rapid separately. The typical skew is small (median $|\delta| \approx 0.018$, about two percentage points of White score), yet it reproduces, position by position, across disjoint accounts. At these positions the disfavoured side also thinks longer. Even where the evaluation is most confident, it is not a sufficient statistic for human outcomes. The result is observational, and the causal question is left to a pre-registered randomised companion study.

2026-07-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

メシエ: クロスベンチマーク エージェント評価用の高解像度コーパス

インタラクティブな環境での AI エージェントの評価は、断片化されたタスク、足場、検証者、スコアリング ルールによって妨げられます。既存の取り組みは狭い設定に焦点を当てており、規模が限られたままであるか、費用のかかる再実行が必要であり、経験的記録の多くは比較できないものとなっています。 Messier は、30 のベンチマーク、714 のエージェント、11,891 のタスク、74,205 の検証者にまたがる 957,253 レコードの統合コーパスです。 Messier は公開ベンチマーク スコアを統合し、最近の法律ベンチマークを含む、過小評価されている 6 つの専門的および科学的領域にわたる 5 つのエージェントの実行でそれらを補完します。各レコードはモデル、足場、環境、タスク、検証者、集計ルールによって標準化されており、職業分析および業界分析用の SOC/NAICS 分類が使用されます。このコーパスを使用すると、ベンチマークの種類によってフロンティアの進歩が均一ではなく、「関数呼び出し」が飽和し、「プログラミング」が最も速く改善し、「エンタープライズ ワークフロー」が依然として最も困難であることがわかります。さらに、反事実的な再スコアリングは、複数の検証者タスクにおける厳格なオールパス集計が進行状況を曖昧にし、エージェントのランキングを人為的に変更する可能性があることを示しています。これらの標準化された記録から、Spearman \r{ho} = 0.81 での Epoch の評価能力指数ランキングと一致する能力スケールを導き出し、ドメイン、職業、アクション スペース、または検証者のタイプによって特化することができます。 Messier は、エージェントの機能スケーリング、ベンチマーク監査、評価失敗の詳細な分析のための、再利用可能な基礎的なインフラストラクチャを提供します。

原文 (English)

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores and supplements them with five-agent runs across six underrepresented professional and scientific domains, including a recent legal benchmark. Each record is standardized by model, scaffold, environment, task, verifier, and aggregation rule, with SOC/NAICS classifications for occupational and industry analysis. Using this corpus, we show frontier progress is uneven across benchmark types, with "function calling" saturated, "programming" improving the fastest, and "enterprise workflows" remaining the most challenging. Furthermore, counterfactual rescoring shows that strict all-pass aggregation in multi-verifier tasks can obscure progress and artificially alter agent rankings. From these standardized records, we derive capability scales that align with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.81 and can be specialized by domain, occupation, action space, or verifier type. Messier provides a foundational, reusable infrastructure for agent capability scaling, benchmark auditing, and fine-grained analysis of evaluation failures.

2026-07-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

インタラクティブ報酬エージェント: 環境状態検証による GUI タスクの評価

グラフィカル ユーザー インターフェイスのタスク評価は、GUI エージェントがユーザーの指示を正常に完了したかどうかを判断することを目的としています。自動化された GUI タスク評価は、評価結果がテスト時のスケーリングとトレーニング後の両方に対する報酬シグナルとして機能する可能性があるため、ますます注目を集めています。ただし、信頼性の高い GUI タスクの評価は、依然として課題が残っています。その判断には、実行軌跡のスクリーンショットを超えて、システム構成、ファイル データ、アプリケーション設定などの環境状態へのアクセスが必要になることが多いためです。この論文では、実行後の環境から証拠を取得して検証するための提案-その後検証フレームワークに基づいた対話型報酬エージェント (IRA) を提案します。タスクの指示と GUI エージェント実行後の GUI 環境が与えられると、IRA はまずタスクの完了条件を提案し、次にシステム ツール、アプリケーション ツール、および GUI ツールを呼び出してそれらを検証します。この設計では、可視インターフェイスと環境状態の両方からの証拠を対話型プロセスで組み合わせます。さらに、10 の Ubuntu デスクトップ アプリケーション カテゴリにわたる 321 の GUI タスクの軌跡のベンチマークである GUI-RewardBench を紹介します。実験によると、IRA は GUI-RewardBench で 86.9% の精度を達成し、既存の評価者のベースラインを上回るパフォーマンスを示しました。さらに IRA を GUI エージェントの強化学習に適用し、OSWorld の成功率 34.0% を達成しました。これは、IRA が GUI エージェントのトレーニングに効果的な報酬シグナルを提供できることを示しています。

原文 (English)

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories. In this paper, we propose an interactive reward agent (IRA) based on a propose-then-verify framework to acquire and verify evidence from the post-execution environment. Given a task instruction and a GUI environment after the GUI agent execution, IRA first proposes the task completion conditions and then verifies them by invoking system tools, application tools, and GUI tools. This design combines evidence from both visible interfaces and the environment state in an interactive process. We further introduce GUI-RewardBench, a benchmark of 321 GUI task trajectories spanning 10 Ubuntu desktop application categories. Experiments show that IRA achieves 86.9% accuracy on GUI-RewardBench, outperforming existing evaluator baselines. We further apply IRA to reinforcement learning of GUI agents, achieving a 34.0% OSWorld success rate, which demonstrates that IRA can provide effective reward signals for training GUI agents.

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

CARE-MH: メンタルヘルス LLM の統合、再現可能、比較可能な評価に向けて

大規模言語モデル (LLM) は、メンタルヘルスのサポートを提供するためにますます使用されており、安全性、共感、治療の適切性についての信頼できる評価が必要です。しかし、既存のメンタルヘルスベンチマークは、一貫性のない評価設計と指標の定義のため、再現して比較することが困難です。我々は、メンタルヘルス LLM の比較可能かつ再現可能な評価のための統一フレームワークである CARE-MH を紹介します。 CARE-MH を使用して、最先端のベンチマークを再現および分析すると、再現性はモデルの安定性に大きく依存し、ベンチマーク間の不一致は主にメトリック定義の違いから生じることが明らかになりました。私たちの調査結果は、将来のメンタルヘルス LLM ベンチマークには、標準化された評価構成と共有指標定義の必要性を浮き彫りにしています。

原文 (English)

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic appropriateness. However, existing mental health benchmarks are difficult to reproduce and compare due to inconsistent evaluation designs and metric definitions. We present CARE-MH, a unified framework for comparable and reproducible evaluation of mental health LLMs. Using CARE-MH, we reproduce and analyze state-of-the-art benchmarks, revealing that reproducibility depends strongly on model stability and that cross-benchmark disagreement primarily arises from differences in metric definitions. Our findings highlight the need for standardized evaluation configurations and shared metric definitions for future mental health LLM benchmarks.

2026-07-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

ゲージ: 黄金の答えのないグレーディング エージェント構築の財務モデル

財務モデルは、一般公開情報とアナリストの仮定を組み合わせて、予測と評価を生成します。一部のコンポーネントは機械的にチェックできますが、予測、割引率、目標価格には複数の合理的な答えが得られることがよくあります。それにもかかわらず、既存のベンチマークは、単一の専門家の参照に照らしてそのような出力を評価する傾向があります。同じ企業に対して独自に構築したアナリスト モデルを使用したところ、65 社をカバーする 108 の有向ペア全体で、単一参照スコアの中央値は 0.33 で、92.6% のスコアが 0.70 未満であり、暗示価格が 10% 以内で一致する同じヴィンテージのペアは存在しないことがわかりました。したがって、ポイントトレランスグレーディングは、専門家の間にすでに存在する意見の相違にペナルティを与える可能性があります。単一点の回答ではなく、観察されたアナリストの実践に対してエージェントが構築した評価モデルを評価するためのベンチマークである GAUGE を紹介します。 GAUGE は、ベンダー別に分類された 1,001 のアナリスト ワークブックと 196 のタスク評価セットを使用し、3 層の観察された実践エンベロープ、56 の監査可能なファセット、8 つの妥当性ゲート、および決定論的な構造チェックを備えています。私たちは、55 人の参加者による既知グループの調査、企業グループのクロスフィッティング、および裁判官の安定性監査によってベンチマークを検証します。失敗認識スコア $\phi_0$ では、上級アナリストは平均 88.3、若手アナリストは 66.0、金融学生は 43.2 でした。 24 人のエージェントと 1,011 人のスコア世代全体で、最高のエージェントのスコアは 53.4 で、学生の平均よりも上ですが、すべての上級生とほとんどの後輩よりも下回っています。機械面の 93%、判定面の 78% を通過し、フリートと中央値の差は 26 ポイントでした。現在のエージェントは、評価判断よりもモデル構築の方が大幅に優れています。私たちは、この方法論、ゲート付き匿名化データ層、制御されたトレーニング分割、バージョン管理された 48 タスクの評価コア、および保留されたリフレッシュ プールをリリースします。

原文 (English)

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $\phi_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

CogArena: 大規模言語モデルにおける認知能力構造のマルチメソッド評価

LLM 認知スコアは、能力ごとのプロファイルとして要約されることが多くなり、その次元はタスク全体で収束し、一致する介入に選択的に反応し、定義に使用されるモデルを超えて一般化される必要があります。 CogArena を紹介します。CogArena は、認知タスクのスコアが 5 つの理論に基づいたグループ分けの次元ラベルを正当化する時期を決定するためのマルチメソッド フレームワークを中心に構築された、手続き的に生成された 13 パラダイム ベンチマークです。 55 のオープンウェイト モデル全体で、ほぼすべてのパラダイム相関が正であり、共通の軸によって分散の約半分が説明されます。グループ内での利点は小さく、スコアに左右されやすく、モデル ファミリ全体で不確実です。 6 つのファミリーからの 12 のモデルにわたる個別に凍結された完全交差研究では、ターゲットを絞ったスキャフォールドは一致グループ化の小さな利点を示しますが、スキャフォールド固有のコントラストは多重性補正に耐えられず、選択性はホールドアウトファミリーの予測を改善しません。凍結された確認基準は失敗します。ポストホックの代替文言の複製では、より小さな正の推定値が生成され、やはり失敗します。これらの結果を総合すると、境界の結論が裏付けられます。理論に沿ったプロンプトは、バッテリー内で小さな斜めの傾向を生成しますが、現在の証拠は安定した 5 次元プロファイルを確立していません。 CogArena は、認知ラベルをモデル スコアに付ける前に、行動シグネチャ、共分散、一致する介入、家族外の予測を結合するワークフローを提供します。

原文 (English)

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

CLBench-V: グラウンディングから知識獲得までのマルチモーダルコンテキスト学習の評価

現実世界のタスクでは、多くの場合、モデルが事前トレーニングされた知識だけに依存するのではなく、タスク固有のコンテキストから学習する必要があります。最近の研究ではこの機能がコンテキスト学習として強調されていますが、既存の評価は主にテキストのコンテキストに焦点を当てています。しかし、実際の多くの設定では、学ぶべきコンテキストは多様です。科学的発見は図や表を通じて伝えられ、財務指標は変換されたレポート全体に散在し、空間的な決定は地図、シーン、または Web ページに依存します。マルチモーダル コンテキスト学習のベンチマークである CLBench-V を紹介します。これは、コンテキストの基礎付け、新しい情報の適用、新しい知識の学習という 3 つの次元に沿ってタスクを整理することで、コンテキストの使用が中断される場所を特定するという困難に対処します。 CLBench-V は、変換された公開ベンチマークと、科学、金融、長文文書の理解、空間推論、Web ベースの視覚的な質問応答などの領域にわたる新しく構築されたデータセットを組み合わせます。ドメイン固有のコンテキスト学習タスクを構築するコストを削減するために、新しく構築されたデータセットに対して自動化された構築およびフィルタリング手順をさらに使用します。 3,443 のインスタンスと 6 つの最近のマルチモーダル モデル全体での最高の総合スコアはわずか 0.2847 であり、マルチモーダル コンテキスト学習がまだ飽和には程遠いことを示しています。さらに、InternVL3.5-30B-A3B はコンテキストの基礎付けと新しい知識の学習で最高のパフォーマンスを発揮し、Qwen3.5-Plus は新しい情報のアプリケーションで最高のパフォーマンスを発揮します。さらに、ジャッジの信頼性、コンテキストの長さ、画像数、代表的な失敗ケースを分析します。コードは https://github.com/IamLihua/CLBench-V で入手できます。

原文 (English)

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. To reduce the cost of constructing domain-specific context-learning tasks, we further use automated construction and filtering procedures for our newly built datasets. Across 3,443 instances and six recent multimodal models, the best overall score is only 0.2847, indicating that multimodal context learning remains far from saturated. Moreover, InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus performs best on new information application. We further analyze judge reliability, context length, image count, and representative failure cases. Code is available at https://github.com/IamLihua/CLBench-V.

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a f…

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clini…

2026-07-29 13:00 JSTarXiv cs.AIビジネス/資金調達

Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models

Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competitive predictive perfor…

2026-07-29 13:00 JSTarXiv cs.AIビジネス/資金調達

On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy

As AI becomes more common in everyday living, there is an increasing demand for intelligent systems that are both performant and understand…

2026-07-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on whic…

2026-07-29 13:00 JSTarXiv cs.AIビジネス/資金調達

AI評価に欠けている要素としての心理的能力

現在の AI 評価フレームワークは、精度、堅牢性、推論能力、ポリシー遵守などの技術的パフォーマンスに主に焦点を当てています。これらの対策は依然として不可欠ですが、自然言語を通じてユーザーと直接対話するシステムにとっては十分ではありません。人間と向き合う AI システムは、アドバイザー、コーチ、家庭教師、仲間として使用されることが増えています。これらの役割では、ユーザーの応答によって、ユーザーがどのように推論し、感情を解釈し、信念を形成し、信頼を調整し、意思決定を行うかが形成されます。したがって、関連する評価単位はモデルだけではなく、人間と AI の相互作用です。この論文では、AI 評価に欠けている側面として心理的能力を紹介します。私たちは、心理的能力を、ユーザー、状況、インタラクションの目的に適切な方法でユーザーの認知、感情の解釈、行動の意思決定をサポートする対人 AI システムの能力として定義します。これには、フレーミング、トーン、知覚される権威、反応性、不確実性の処理、会話のガイダンスなどの対話特性が含まれます。既存の評価アプローチはこの問題の一部を捉えていますが、これらの心理的影響を直接評価することはほとんどありません。行動科学と人間と AI の相互作用研究に基づいて、心理的能力とその中核領域の概念的枠組みを概説します。特定のベンチマークを提案するのではなく、構成を定義し、その境界を明確にし、シナリオベースの調査、構造化された人間による評価、およびモデル支援の評価方法を通じてそれがどのように評価されるかを説明します。私たちは、心理的能力が、人間と対面する AI システムの実世界への影響を懸念するモデル提供者、導入組織、研究者、規制当局にとって中心的な考慮事項となるべきであると主張します。

原文 (English)

Psychological Competence as a Missing Dimension in AI Evaluation

Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias…

2026-07-29 13:00 JSTarXiv cs.AIビジネス/資金調達

Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to…

2026-07-29 09:09 JSTTechCrunch AIエージェントビジネス/資金調達
2026-07-29 06:29 JSTTechCrunch AIビジネス/資金調達

Bot-detection startup Spur nabs $200M from Insight

Spur Intelligence has raised a $200 million round from Insight Partners for its tech that can identify legit human traffic from bots.

2026-07-28 23:00 JSTTechCrunch AIビジネス/資金調達

Fish Audio raises $52M seed to build AI voice models for creators and enterprises

Since launching last year, the startup today has more than 8 million people using the open source or hosted version of its models, and now…

2026-07-28 13:30 JSTTechCrunch AIビジネス/資金調達

Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing

Cursor says India is now its third-largest market globally and plans to expand local hiring and enterprise sales.

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ジャッジの体系化: プログラム蒸留によるスケーラブルな評価

LLM-as-a-judge は自動評価の標準となっていますが、高コスト、大幅な遅延、不透明な決定という、拡張性と信頼性を損なう制限に悩まされています。私たちは、プログラム蒸留というシンプルで効率的な代替手段でこれらに対処します。評価時に LLM を促す代わりに、その決定ロジックを抽出して、候補者を直接採点するプログラムの委員会を作成します。これらのプログラムによるジャッジは透明性を提供し、簡単に検査または編集でき、サンプルごとの API コストを排除します。この概念に基づいて、私たちは、裁判官としてプログラムを統合し、その決定を共同評決に集約し、信頼性の低い訴訟を選択的に LLM にエスカレーションするフォールバック メカニズムを組み込むシステムである PAJAMA を導入します。 5 つのデータセットと 4 つのモデル ファミリーにわたって、プログラムによる裁判官が 13B サイズの LLM 裁判官のパフォーマンスに匹敵することができることを示します。プログラム出力をルーティング信号として使用すると、PAJAMA は精度とスループットの両方を向上させ、パレート フロンティアを前進させます。評価を超えて、プログラムによるジャッジは安価で効果的な報酬シグナルを生成します。RewardBench では、プログラムの評決から抽出された報酬モデルが、2 桁低い API コストで独自の LLM ラベルでトレーニングされた報酬モデルよりも優れたパフォーマンスを発揮します。

原文 (English)

Codifying the Judge: Scalable Evaluation via Program Distillation

LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.

2026-07-28 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

インダストリー 4.0 エージェントの評価のための合成シナリオの生成

産業用エージェントのベンチマークには、テレメトリ、障害モード、メンテナンス記録、ドメイン標準を統合した現実的な評価シナリオが必要です。ただし、AssetOpsBench などの既存のベンチマークは、手動で作成されたシナリオに依存しており、限られた資産クラスのセットをカバーしています。当社は、スマート グリッド変圧器アセット クラスと、健全性指数予測、溶存ガス分析、巻線温度評価、および負荷プロファイル評価のための 4 つの IEC ベースの診断ツールを使用して AssetOpsBench を拡張します。さらに、合成産業エージェント シナリオ生成用のパイプラインである ScenarioGeneratorAgent を紹介します。このパイプラインは、証拠に基づいた資産プロファイルを構築し、運用ドメイン全体にカバレッジを意識したシナリオ予算を割り当て、スキーマの妥当性、ツールの到達可能性、物理的な妥当性、標準の整合性、重複排除を強制するハイブリッド検証と修復のループを通じて候補を生成します。スケーラビリティを向上させるために、2 レベルのキャッシュ、並列フォーカス グループ生成、スレッド プールのオフロード、バッチ化された LLM 呼び出し、および早期拒否フィルタリングを適用します。 Smart Grid Transformer のシナリオ生成では、これらの最適化により、品質を維持しながら 50 のシナリオでエンドツーエンドのランタイムが $8\times$ 削減され、最適化されていないベースラインの $73.8 \pm 3.0$ と比較して $74.2 \pm 1.9$ の複合品質スコアを達成しました。これらの結果は、標準に基づいた合成シナリオ生成により、シナリオの品質を犠牲にすることなく産業エージェントのベンチマークを効率的に拡張できることを示しています。

原文 (English)

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks such as AssetOpsBench rely on manually authored scenarios and cover a limited set of asset classes. We extend AssetOpsBench with a Smart Grid Transformer asset class and four IEC-grounded diagnostic tools for health-index prediction, dissolved-gas analysis, winding-temperature assessment, and load-profile assessment. We further introduce ScenarioGeneratorAgent, a pipeline for synthetic industrial-agent scenario generation. The pipeline constructs evidence-grounded asset profiles, allocates coverage-aware scenario budgets across operational domains, and generates candidates through a hybrid validation-and-repair loop that enforces schema validity, tool reachability, physical plausibility, standards alignment, and deduplication. To improve scalability, we apply two-level caching, parallel focus-group generation, thread-pool offloading, batched LLM calls, and early rejection filtering. On Smart Grid Transformer scenario generation, these optimizations reduce end-to-end runtime by $8\times$ for 50 scenarios while preserving quality, achieving a composite quality score of $74.2 \pm 1.9$ compared with $73.8 \pm 3.0$ for the unoptimized baseline. These results show that standards-grounded synthetic scenario generation can efficiently expand industrial-agent benchmarks without sacrificing scenario quality.

2026-07-28 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

コーディング エージェントの足場効果: コーディング エージェントの評価における隠れた変数としての選択を利用する

コーディング エージェント向けの公開リーダーボードは通常、モデル名と合格率によってシステムをランク付けしますが、周囲のハーネス (ツールを発行し、コンテキストを管理し、いつ停止するかを決定する足場) は十分に指定されていないことがよくあります。モデル間の比較は、ハーネスが固定されている場合に有効です。パフォーマンスと効率が異なる場合、モデルと足場効果が混同されます。 Terminal-Bench Pro の階層化された 50 タスクのサブセット上の 3 つのオープンソース ハーネス (Goose、OpenCode、OpenHands-SDK) にわたって Qwen 3.6 Plus と MiniMax M2.5 を評価します。ハーネスの選択により、解決されたタスクごとに最大 40 倍のトークンの差が生じますが、モデル内のペアの合格率の差は 0 ~ 8 パーセント ポイントのままです (95% のペア タスク ブートストラップ CI には、最大のギャップを除いてゼロが含まれます)。障害のフィンガープリントはモデル間で複製され (Goose の場合は REASON、OpenHands-SDK の場合は VERIFY/MAX_TURNS、OpenCode の場合は idle-loop/TIME)、モデルにほとんど依存しないハーネス レベルのバイアスを示します。人間中心のコーディング エージェントの評価では、モデル名だけでは比較単位が不完全です。ハーネスとモデルのペアによって、実際のコスト、遅延、監視の負担が決まります。ノーアクション ターンは、単なるトークン税ではなく、タスクごとの待機税です。したがって、トークン/レイテンシ バジェットに基づいてハーネスとモデルのペアを選択し、モデルの比較とともにトークンの使用量、レイテンシ、フル ハーネスの仕様を報告することをお勧めします。匿名化された構成、生の試用ログ、集約されたスナップショット、分析スクリプトをリリースします。

原文 (English)

The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation

Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified. Model-to-model comparison is valid when the harness is fixed; when it varies, performance and efficiency conflate model and scaffold effects. We evaluate Qwen 3.6 Plus and MiniMax M2.5 across three open-source harnesses (Goose, OpenCode, OpenHands-SDK) on a stratified 50-task subset of Terminal-Bench Pro. Harness choice induces up to a 40x difference in tokens per solved task, while paired within-model pass-rate differences remain 0-8 percentage points (95% paired-task bootstrap CIs include zero except for the largest gap). Failure fingerprints replicate across models (REASON for Goose, VERIFY/MAX_TURNS for OpenHands-SDK, idle-loop/TIME for OpenCode), indicating harness-level biases that are largely model-independent. For human-centered coding-agent evaluation, model name alone is an incomplete comparison unit: harness-model pairs determine real-world cost, latency, and oversight burden; no-action turns are a per-task wait tax, not just a token tax. We therefore recommend selecting harness-model pairs by pass rate under token/latency budgets, and reporting token usage, latency, and full harness specifications alongside any model comparison. We release anonymized configs, raw trial logs, aggregated snapshots, and analysis scripts.

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ParBench: LLM 並列コード変換の信頼性の高い評価のためのベンチマーク

最新のコンピューティング集約型ソフトウェアは、アクセラレータ、プログラミング API、コンパイラ スタック、CUDA、OpenMP、OpenCL、OpenMP ターゲット オフロードなどの移植性レイヤーの変化するエコシステム全体で移行する必要があります。このような移行のために、大規模な言語モデルと自律コーディング エージェントがますます提案されていますが、この分野には、スレッド インデックス付け、同期、メモリ管理、ホスト デバイスの調整、API 固有の実行構造など、翻訳を動作的に有効にする低レベルの並列セマンティクスが保持されているかどうかを測定する信頼できる方法がありません。ここでは、実行可能で再現可能な条件下で LLM ベースの並列 API 変換を評価するためのカーネル中心のベンチマーク フレームワークである ParBench を紹介します。 ParBench は、宣言的なベンチマーク仕様を通じて周囲の構築、実行、検証インフラストラクチャを修正し、計算カーネルのみを変換するようにモデルに要求します。複数のオープンソース HPC スイートを利用し、CUDA、OpenMP、OpenCL、OpenMP ターゲット オフロード間の代表的なクロス API 変換の方向をカバーします。成功が表面的な形式の記憶ではなく堅牢な翻訳を反映しているかどうかをテストするために、ParBench には、AST 主導の、意図された動作を保持する、ベースライン検証済みのソース拡張機能が含まれています。最先端のオープンおよび独自の LLM の評価では、方向の非対称性、複数ファイルの調整、不完全な API 適応、ソースレベルの摂動に対する不均一な堅牢性など、信頼性の高い並列コード変換に対する永続的な障壁が示されています。コードは https://github.com/Scientific-Computing-Lab/ParBench で入手できます。

原文 (English)

ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation

Modern compute-intensive software must migrate across a changing ecosystem of accelerators, programming APIs, compiler stacks, and portability layers, including CUDA, OpenMP, OpenCL, and OpenMP target offload. Large language models and autonomous coding agents are increasingly proposed for such migration, but the field lacks reliable ways to measure whether they preserve the low-level parallel semantics that make translations behaviorally valid, including thread indexing, synchronization, memory management, host-device coordination, and API-specific execution structure. We present ParBench, a kernel-centric benchmark framework for evaluating LLM-based parallel API translation under executable, reproducible conditions. ParBench fixes the surrounding build, run, and verification infrastructure through declarative benchmark specifications and asks models to translate only the computational kernels. It draws on multiple open-source HPC suites and covers representative cross-API translation directions among CUDA, OpenMP, OpenCL, and OpenMP target offload. To test whether success reflects robust translation rather than surface-form memorization, ParBench includes AST-driven, intended behavior-preserving, baseline-validated source augmentation. Evaluations on state-of-the-art open and proprietary LLMs show persistent barriers to reliable parallel code translation, including direction asymmetry, multi-file coordination, incomplete API adaptation, and uneven robustness to source-level perturbations. Code is available at https://github.com/Scientific-Computing-Lab/ParBench.

2026-07-28 13:00 JSTarXiv cs.AIビジネス/資金調達

VlogReward: Vlog編集のための多次元評価を学ぶ

パーソナライズされたストーリーテリング メディアとしての vlog の急速な台頭により、vlog 編集計画を評価および改良するための自動システムの需要が生じています。しかし、vlog の評価は非常に主観的であり、標準化された基準、データセットとベンチマーク、効果的な報酬モデルが不足しているため、依然として困難が伴います。これらの課題に対処するために、私たちはプロの vlog クリエイターとプロダクト マネージャーによる包括的な vlog 評価フレームワークを定義し、6 つの主要な側面 (創造性、一貫性、コンセプト デザイン、撮影、ナレーション、ペーシング) の分類を確立しました。続いて、マルチモーダル大規模言語モデル (MLLM) の vlog 報酬機能を評価するために、10 万件の vlog 編集からなる大規模なデータセットと専用ベンチマーク VRMBench を厳選しました。最後に、きめの細かい多次元スコアと反復的な改善のための実用的なフィードバックの両方を提供できる堅牢な vlog 報酬モデルである VlogReward を紹介します。技術的には、調整可能なグループ間比較報酬を導入することで、グループ相対ポリシー最適化 (GRPO) フレームワークを強化します。これにより、標準 GRPO の「方向盲目」問題が軽減され、モデルがさまざまな品質の編集をより適切に区別できるようになります。 VlogReward は、GPT-5 や Gemini-3-Pro などの既存の MLLM を大幅に上回る最先端の結果を実現します。私たちの研究が vlog 作成者に役立ち、vlog の自動評価および改良システムの促進に役立つことを願っています。

原文 (English)

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective reward models. To address these challenges, we define a comprehensive vlog evaluation framework guided by professional vlog creators and product managers, establishing a taxonomy of six key dimensions, i.e., Creativity, Consistency, Concept Design, Cinematography, Narration, and Pacing. Subsequently, we curate a large-scale dataset of 100k vlog edits and a dedicated benchmark, VRMBench, to evaluate the vlog rewarding capabilities of Multimodal Large Language Models (MLLMs). Finally, we present VlogReward, a robust vlog reward model that can provide both fine-grained multi-dimensional scores and actionable feedback for iterative refinement. Technically, we enhance the Group Relative Policy Optimization (GRPO) framework by introducing an adjustable inter-group comparison reward, which mitigates the "direction blindness" issue of standard GRPO and enables the model to better distinguish varied-quality edits. VlogReward achieves state-of-the-art results that significantly outperform existing MLLMs, including GPT-5 and Gemini-3-Pro. We hope that our study can help vlog creators and foster automated vlog evaluation and refinement systems.

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

StanceBench: 音声 LLM ベースの対人スタンス評価の音声からのベンチマーク

音声間の対話モデルは、社会的意図を伝えるために韻律や相互作用のニュアンスにますます依存していますが、これらの合図のベンチマークは依然として限られています。会話中の会話における対人スタンスを測定し、自動判定として音声対応 LLM を評価するためのベンチマークである StanceBench を紹介します。シームレス インタラクション コーパスを使用して、StanceBench は、(1) ロール プロンプト ポールを介して 9 つのスタンスの次元を指定し、(2) 単一話者およびインタラクション ベースの評価を標準化し、(3) ジャッジとしての LLM の堅牢性、バイアス、およびスタンス推論をレポートします。評価されるスタンスの中で、共感と礼儀正しさが最も簡単です。温かさと自己主張は、ポジティブな偏り/非対称性によって適度に分離可能です。正直さが最も難しく、即時注文バイアスが高く、クロスターンの証拠が必要であることと一致しています。注意力は分離可能ですが、人間との連携は弱いです。インタラクションのスタンスはよりコンテキストに依存しており、しきい値のギャップと大きな差異があり、特に紛争規制が顕著です。

原文 (English)

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceBench, a benchmark for measuring interpersonal stance in conversational speech and evaluating audio-capable LLMs as automated judges. Using the Seamless Interaction corpus, StanceBench (1) specifies 9 stance dimensions via role-prompt poles, (2) standardizes single-speaker and interaction-based evaluations, and (3) reports LLM-as-a-judge robustness, bias, and stance inference. Across evaluated stances, empathy and politeness are the easiest. Warmth and assertiveness are moderately separable with positivity skew/asymmetry. Honesty is the hardest and shows high prompt order bias, consistent with needing cross-turn evidence. Attentiveness is separable but aligns weakly with humans. Interaction stances are more context-sensitive, with threshold gaps and high variance, especially conflict regulation.

2026-07-28 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable delivera…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under neste…

2026-07-28 13:00 JSTarXiv cs.AIビジネス/資金調達

NeurGO: Learning to Generate Elite Candidates for Meta-Black-Box Expensive Optimization

Expensive black-box optimization is ubiquitous in science and engineering, where function evaluations are costly and the evaluation budget…

2026-07-28 13:00 JSTarXiv cs.AIビジネス/資金調達

Offline-to-Online Creative Optimization with Generative Models and Adaptive Testing

Ad creative optimization is increasingly constrained by evaluation rather than generation. Generative models can produce many plausible cre…

2026-07-28 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper h…

2026-07-28 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer di…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Reason-Mediated Behavioral Models for Auditing LLM Social Simulators

Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether s…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Creative Integration: A Decidable Criterion of Creativity

"Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes t…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks

The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying comm…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instr…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts…

2026-07-28 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Fair Division with Strictly Increasing Valuations: A Tight Threshold for Two-Agent EF1 and PO

We study whether strictly positive marginal values restore the compatibility of envy-freeness up to one good (EF1) and Pareto optimality (P…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static…

2026-07-28 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim

Social robots that rely on large language models for conversation are unable to retain information across sessions. This absence of memory…

2026-07-28 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility

Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical d…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

PD$^3$: A Project Duplication Detection Framework via Adapted Multi-Agent Debate

Project duplication detection is critical for project quality assessment because it helps avoid investment in repeated proposals. Existing…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabilities and alignment properties they are meant to represent remain uncertain. Synthesizing 373 primary studies published between 2018 and 2026, the survey organizes evidence on benchmark validity, contamination, dynamic evaluation, LLM-as-a-judge protocols, adversarial safety testing, reward and proxy optimization, mechanistic interpretability, and AI governance into an eight-stream evidence taxonomy. Building on this synthesis, we introduce EvalSafetyGap, a conceptual framework that unifies benchmark-validity and alignment-failure research as a shared proxy-target divergence problem under optimization pressure, formalized through a Goodhart-inspired Instability Decomposition and an Alignment Trilemma. An exploratory ten-model public-evidence audit illustrates the framework by showing why capability, behavioral robustness, and governance disclosure should be reported as separate evidence layers rather than collapsed into a single safety score. The survey closes with a research agenda for dynamic and contamination-resistant benchmarks, pre-specified multi-attempt threat models, version-locked evaluation, transparent source reporting, and validated mechanistic safety indicators, offering researchers, model developers, and AI auditors a shared vocabulary for measurement-aware LLM safety evaluation.

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Like a Baby: Visually Situated Neural Language Acquisition

We examine the benefits of visual context in training neural language models to perform next-word prediction. A multi-modal neural architec…

2026-07-28 13:00 JSTarXiv cs.AIビジネス/資金調達

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established prac…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems

As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. C…

2026-07-28 13:00 JSTarXiv cs.AIビジネス/資金調達

The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits

The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains…

2026-07-28 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

Evaluating multi-turn medical consultation agents requires judging the diagnostic support provided by the histories they elicit through int…

2026-07-27 22:00 JSTTechCrunch AIロボティクスビジネス/資金調達

Enigma raises $71M to make controlling a robot as easy as adjusting the volume

The massive seed round was led by Index Ventures and Ribbit Capital, with participation from Sarah Guo's Conviction Partners.

2026-07-27 19:00 JSTITmedia AI+ビジネス/資金調達

検索結果に「詐欺ではありません」と表示させる詐欺手口、警視庁が注意喚起 AI要約も餌食に

警視庁は、SNS型投資詐欺グループがWeb検索の仕組みを悪用し、検索結果に肯定的な情報を並べ、AI要約にも「詐欺ではありません」と表示させる手口を確認した。

2026-07-25 07:25 JSTTechCrunch AIビジネス/資金調達

Prentis, new AI lab co-founded by Reid Hoffman, Mark Pincus in talks to raise $100M

The neolab is betting that automating routine computer tasks will soon outpace coding as AI's biggest use case.

2026-07-25 03:07 JSTTechCrunch AIエージェントビジネス/資金調達

Why Cognition bought Poke: AI personality is becoming a competitive advantage

The acquisition brings Poke’s conversational style and interaction model to Cognition’s coding agent Devin, reflecting a growing belief tha…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

DFAH ベンチ: 財務上の意思決定における観察可能なエージェントの不安定性のベンチマーク

標準の評価ベンチマークは、ツールを使用するエージェントが毎回同じプロセスを経てその決定に到達するかどうかではなく、ツールを使用するエージェントが何を決定するかを測定します。 DFAH ベンチは、金融機関の意思決定における観察可能な行動の不安定性を 3 つのチャネル (ツール呼び出しの軌跡、証拠の接触、意思決定の集中) にわたって測定するリプレイ ベンチマークです。これらのチャネルのいずれも、隠された推論テキストへのアクセスを必要としません。 10 のモデルと 3 つの財務タスクにまたがる 8,127 のリプレイ エピソード全体で、結果の合意だけでは不完全な安定性シグナルであることがわかりました。フロンティア モデルは 95% の確率で意思決定に合意できるのに対し、同じツール パスをたどるのは 77% の確率だけです。結果のみの評価では完全に見逃される 18 パーセント ポイントのギャップ (95% CI: [0.14, 0.22])。意思決定の一致度が高いフロンティアモデルのケースグループのうち、55% 以上が意味のある軌道の発散を示しています。私たちは 3 つの動作プロファイルを特定します。入力に関係なく単一の出力に集約することでほぼ完全な一致を達成するパターン マッチャー、比較的一貫したツール使用プロセスを持つ安定した実行者、および実質的に異なるツール パスと証拠の接触を通じて同じ結論に達する軌道分岐者です。ベンチマーク コード、メトリック スクリプト、再生ログ、ベンチマーク カード、データセット README、およびリリース マニフェストは、付属のリポジトリでリリースされます。

原文 (English)

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We introduce DFAH-Bench, a replay benchmark that measures observable behavioral instability in financial agent decision-making across three channels -- tool-call trajectories, evidence contacts, and decision concentration -- none of which require access to hidden reasoning text. Across 8,127 replay episodes spanning 10 models and 3 financial tasks, we find that outcome agreement alone is an incomplete stability signal: frontier models can agree on decisions 95% of the time while following the same tool path only 77% of the time -- an 18-percentage-point gap (95% CI: [0.14, 0.22]) that outcome-only evaluation misses entirely. Among frontier-model case groups with high decision agreement, over 55% exhibit meaningful trajectory divergence. We identify three behavioral profiles: pattern matchers that achieve near-perfect agreement by collapsing to a single output regardless of input, stable executors with relatively consistent tool-use processes, and trajectory divergers that reach the same conclusions through materially different tool paths and evidence contacts. The benchmark code, metric scripts, replay logs, benchmark card, dataset README, and release manifest are released in the accompanying repository.

2026-07-24 13:00 JSTarXiv cs.AIビジネス/資金調達

AI 知識システムにおけるエンティティの重要性の表現: 聴衆評価と構造的権威のデュアルシグナル フレームワーク

AI 知識システムでは、検索、推奨、証拠の選択、知識集約型の推論のためにエンティティの重要性を表現する必要があります。しかし、重要性は多くの場合、人間の反応またはグラフ構造から導き出される単一のスコアに還元されます。このような圧縮により、AI システムがさまざまなタスクのエンティティの中から選択する必要がある場合に重要な区別が無視される可能性があります。この研究では、各エンティティが視聴者評価の次元と構造的権威の次元によって特徴付けられる、解釈可能な二重信号表現を導入しています。このフレームワークは、経験的検証ドメインとして映画エンティティを使用して評価されます。 IMDb の非営利データセットは評価ベースの視聴者ランキングを提供し、ウィキデータはエンティティのアライメントをサポートし、英語版ウィキペディアのハイパーリンクは PageRank が構造的権威を推定する知識ネットワークを形成します。 482 個のエンティティと 13,690 個の有向関係に関する実験により、2 つの次元間の統計的に有意ではあるが弱い関連性が明らかになりました (Spearman rho = 0.2275、p < 0.001)。それらの重複は上位 10 位ではわずか 10%、上位 100 位では 34% ですが、エンティティレベルの相違は両方向に発生します。この結果は、視聴者の評価と構造的権威は非冗長シグナルであり、重要性に関する単一のスカラー概念に自動的に集約されるべきではないことを示しています。この貢献は、新しいランキング アルゴリズムや学習された埋め込みではなく、最小限の知識表現フレームワークとその次元の必要性の経験的テストです。この調査結果は、コンテキスト固有の選択や集計を適用する前に、明確な重要性シグナルを保存するタスク認識 AI 知識システムをサポートしています。

原文 (English)

Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority

AI knowledge systems require representations of entity importance for retrieval, recommendation, evidence selection, and knowledge-intensive reasoning. Yet importance is often reduced to a single score derived from either human response or graph structure. Such compression may discard distinctions that matter when an AI system must choose among entities for different tasks. This study introduces an interpretable dual-signal representation in which each entity is characterized by an audience-evaluation dimension and a structural-authority dimension. The framework is evaluated using movie entities as an empirical validation domain. IMDb non-commercial datasets provide a rating-based audience ranking, Wikidata supports entity alignment, and English Wikipedia hyperlinks form the knowledge network on which PageRank estimates structural authority. Experiments on 482 entities and 13,690 directed relationships reveal a statistically significant but weak association between the two dimensions (Spearman rho = 0.2275, p < 0.001). Their overlap is only 10% in the top 10 and 34% in the top 100, while entity-level divergence occurs in both directions. The results show that audience evaluation and structural authority are non-redundant signals and should not automatically be collapsed into a single scalar notion of importance. The contribution is not a new ranking algorithm or learned embedding, but a minimal knowledge-representation framework and an empirical test of its dimensional necessity. The findings support task-aware AI knowledge systems that preserve distinct importance signals before applying context-specific selection or aggregation.

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models

Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this ass…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility

Proprietary large language models (LLMs) entail substantial intellectual and financial investment, making them valuable intellectual proper…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation

Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance j…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource…

2026-07-24 13:00 JSTarXiv cs.AIビジネス/資金調達

From Checklists to Clusters: A Homeostatic Account of AGI Evaluation

Contemporary AGI evaluations report multidomain capability profiles, yet they typically assign symmetric weights and rely on snapshot score…

2026-07-24 13:00 JSTarXiv cs.AIビジネス/資金調達

Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks

We characterize AI automation as a continuum between crashing waves, in which capabilities jump abruptly across narrow task sets, and risin…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

隠れたフットプリント: ストレージを LLM エージェント評価の第一級の指標にする

LLM エージェントのベンチマークは、タスクの完了、信頼性、推論コストを測定しますが、ログ、コンテキスト スナップショット、チェックポイント、デバッグ トレースなど、エージェントの実行によってディスクに残される永続データは測定しません。実行後のエージェント ストレージ フットプリントのクロスフレームワーク ベンチマークである AgentFootprint を紹介します。そのシリアル化対応メトリクス スイートは、総保持率、チャネル構成、重複、増加、圧縮率、会話履歴の再構築可能性を測定します。これは、測定の罠に対処します。単純なバイトレベルの測定では、データベースのページングと JSON エスケープが繰り返されるコンテンツを不明瞭にするため、重複が桁違いに過小評価されます。固定トレース制御により、エージェントが生成した論理ボリュームが永続層の増幅から分離されます。7 つの永続フレームワークを通じて同じ軌跡を再生すると、6.7 倍の広がりが得られます。同一のモデル、ツール、およびタスクでは、100% の精度の構成では、デフォルトでサポートされる回復機能と監査機能が異なりますが、保持バイト数が 15.7 倍異なります。 3 つの完全な履歴構成は、反復観察ストレス タスクで超線形に成長します。 108 個のインスタンスで正規化された SWE ベンチからエクスポートされた軌跡 検証済みの送信は、インスタンスごとに 3 桁の大きさに及び、解決率との検出可能な相関関係はありません。コンテンツ アドレス ストアは、すべての再構築可能性スコアを維持しながら、保持率を 4.8 倍から 32.7 倍まで削減します。これらの結果は、精度と再構築可能性を併せてレポートするためのリソース メトリックとして永続ストレージを確立します。

原文 (English)

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite measures total retention, channel composition, duplication, growth, compressibility, and conversation-history reconstructability. It addresses a measurement trap: naive byte-level measurement understates duplication by an order of magnitude because database paging and JSON escaping obscure repeated content. A fixed-trace control separates agent-generated logical volume from persistence-layer amplification: replaying the same trajectory through seven persisting frameworks yields a 6.7x spread. Under identical models, tools, and tasks, configurations with 100% accuracy differ by 15.7x in retained bytes, although their defaults support different recovery and audit capabilities. Three full-history configurations grow superlinearly on a repeated-observation stress task. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude per instance, with no detectable correlation with resolve rate. A content-addressed store reduces retention by 4.8x-32.7x while preserving every reconstructability score. These results establish persistent storage as a resource metric to report jointly with accuracy and reconstructability.

2026-07-24 13:00 JSTarXiv cs.AIビジネス/資金調達

SenWorld: コンテキストリッチな評価データを生成するためのデジタル ツイン シミュレーション

スマートフォンのパーソナル アシスタントは長期にわたる個人データを推論しますが、その評価には正解がわかっているコンテキストに富んだ評価データが必要であり、実際のデバイスのトレースはプライバシーに敏感すぎて共有できません。この課題に対処するために、構築によって固定されたグラウンド トゥルースを使用してそのようなデータを生成する、物理的に接地され、決定論的でイベント ソースのデジタル ツイン シミュレーションである SenWorld を紹介します。 SenWorld では、ペルソナは実際の地図、天気、休日、ネットワーク データから構築された世界で 1 日を過ごします。観測可能なすべての信号はシステム全体のスナップショットにアーカイブされます。また、各評価ケースは、事後注釈や大規模言語モデル (LLM) ジャッジではなく、既存のレコードへのポインターによってラベル付けされます。この手法を北京の 16 人のペルソナで評価しました。生成されたデータは、カテゴリ分布 (ジェンセンとシャノンの相違 (JSD) 0.070) および通信記録の 1 日のリズム (JSD 0.1 未満) において、保持されている実際のユーザーのベンチマークと厳密に一致していますが、生成された記録は実際の記録よりも短いままです。スクリプトによる対話がなければ、ペルソナは完全に往復する対話サブグラフと差別化された行動レパートリーを形成します。 717 件の評価ケースに投影された生成データでは、実稼働スマートフォン アシスタントの 78 件の障害が明らかになり、通話とショート メッセージ サービス (SMS) の記録に集中し、連絡先、スケジュール、アラームは決して失敗しませんでした。スナップショット ポインタは、LLM 判定者が関与せずに、各失敗をアシスタント側の取得エラーとして確認します。全体として、SenWorld は、ラベルが構築によって固定されている評価データへの、プライバシーに安全で再現可能で配布がチェックされたパスを提供します。

原文 (English)

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data

Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, we present SenWorld, a physically grounded, deterministic, event-sourced digital-twin simulation that generates such data with ground truth fixed by construction. In SenWorld, personas live through a full day in a world built from real map, weather, holiday, and network data; every observable signal is archived in full-system snapshots; and each evaluation case is labeled by a pointer to an existing record rather than by post-hoc annotation or a large language model (LLM) judge. We evaluate this method with 16 personas in Beijing. The generated data closely matches the held-out real-user benchmark in category distribution (Jensen--Shannon divergence (JSD) 0.070) and in the daily rhythm of communication records (JSD below 0.1), though generated records remain shorter than real ones. Without scripted interaction, personas form a fully reciprocated dialogue subgraph and differentiated behavioral repertoires. Projected into 717 evaluation cases, the generated data exposes 78 failures in a production smartphone assistant, concentrating on call and Short Message Service (SMS) records while contacts, schedules, and alarms never fail. The snapshot pointer confirms each failure as an assistant-side retrieval error, with no LLM judge involved. Overall, SenWorld offers a privacy-safe, reproducible, and distribution-checked path to evaluation data whose labels are fixed by construction.

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation

Error Span Detection (ESD) extends automatic machine translation (MT) evaluation by localizing translation errors and labeling their severi…

2026-07-24 13:00 JSTarXiv cs.AIビジネス/資金調達

SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning

Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…

2026-07-24 06:15 JSTITmedia AI+ロボティクスビジネス/資金調達研究/論文

三井不動産がデータセンターに6000億円超投資、物流の枠超え「産業デベロッパー」へ

三井不動産は事業説明会で「産業デベロッパー」への領域拡大を発表した。従来の物流拠点供給にとどまらず、研究開発施設や自動運転対応を進める。データセンター事業には累計6000億円超を投じ、稼働済みの3棟に加え7棟を開発中だ。

2026-07-24 00:00 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors

Etched, founded by three Harvard dropouts, has created new chips and memory components that speed up inference on any AI model -- no GPUs r…

2026-07-23 15:09 JSTTechCrunch AIビジネス/資金調達

ServiceNow bets $40 million on Indian banking software specialist to expand its financial services push

ServiceNow's investment gives BusinessNext a strategic partner to expand its AI-powered banking software globally.

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

大規模言語モデルにおける不確実性評価の再考

キャリブレーションは、LLM の信頼性を評価するための主要な基準ですが、不十分です。キャリブレーションでは、自明に一貫性のない推定値が認められ、評価分布に依存し、推定値が一貫した基礎的な確率関数としてどの程度解釈できるかをテストしていません。実際に必要なのは、LLM 信頼推定値が一貫した確率的信念に必要な条件を満たすことです。これらの条件を 3 つの軸 (構造的一貫性、忠実性、有用性) に沿って形式化し、C1 メトリクスとして運用可能にします。広く使用されている推定器は、適切に校正されているように見えても、体系的にこれらの条件に違反しています。モデルは、論理的に簡単な質問に対して 31\% の確率で低い信頼度を割り当て、RMSCE を削減する一般的な介入では構造違反は変化せず、校正が確率的妥当性と直交していることを示唆しています。 RLHF と思考連鎖は、一貫性を回復することなく有用性の指標を向上させます。私たちの結果は、現在の LLM 信頼推定値が一貫した確率として解釈できないことを示しています。私たちのフレームワークは、このギャップを測定して埋めるためのツールを提供します。

原文 (English)

Rethinking Uncertainty Evaluation in Large Language Models

Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function. What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic beliefs. We formalize these conditions along three axes (structural coherence, faithfulness, and usefulness) and operationalize them as the C1 metrics. Widely used estimators systematically violate these conditions despite appearing well-calibrated: models assign lower confidence to logically easier questions 31\% of the time, and common interventions reducing RMSCE leave structural violations unchanged, suggesting that calibration is orthogonal to probabilistic validity. RLHF and chain-of-thought improve usefulness metrics without restoring coherence. Our results show current LLM confidence estimates cannot be interpreted as coherent probabilities; our framework provides the tools to measure and close this gap.

2026-07-23 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

FORCE-Bench: エンタープライズ ファイナンスにおけるエージェントティック AI のベンチマーク、データセット、評価ハーネス

大規模言語モデルの最近の進歩により、運用財務におけるエージェント システムの導入が加速しています。既存のベンチマークは、一般的な能力、指示への従うこと、または安全性の測定に重点を置いていますが、自動化するために現在エージェント システムが導入されている運用財務ワークフローに直接取り組んでいるベンチマークはほとんどありません。財務専門家は、エージェントに対し、事実に基づいた適切な根拠に基づいた情報を提供するだけでなく、その情報が検証可能であり、運用財務領域のルールと制約に一貫して準拠していることを確認することを求めます。 FORCE-Bench を紹介します。これには専門家が注釈を付けた 251 のクエリが含まれており、正確性、引用、明確さ、深さ、根拠性、最新性、関連性、構造の 8 つの側面にわたって、運用財務ドメインの要件に合わせて調整されたルーブリックベースのフレームワークを使用して応答を評価します。 FORCE-Bench は、財務上の義務の調査 (ERP システムに売掛金および買掛金のデータを問い合わせる)、金融機関のパフォーマンスの調査 (公開書類や市場データからの期限付きの質問に答える)、ビジネス概要の生成 (マルチソースの企業インテリジェンス レポートの合成) の 3 つのタスク タイプでエージェント システムを評価します。実際の展開条件を反映するために、共通のツール アクセスと遅延制限設定の下で、専用エージェントと汎用エージェント システムを評価します。結果は、汎用エージェント システムは運用上の制約の下で財務ドメインの品質要件を一貫して満たしていないのに対し、Microsoft 365 Copilot 専用の Finance Agent はあらゆる側面で信頼性が高いことを示しています。データセット、ルーブリック、ハーネス、分析コードをオープンソースとしてリリースし、再現可能な比較と他の企業財務環境への適応をサポートします。

原文 (English)

FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance

Recent advances in large language models have accelerated deployment of agentic systems in operational finance. Existing benchmarks emphasize measuring general capabilities, instruction following, or safety, but few directly address the operational finance workflows that agentic systems are now being deployed to automate. Finance professionals require agents to not only provide factually sound and properly grounded information, but also ensure that this information is verifiable and consistently adheres to rules and constraints of the operational finance domain. We introduce FORCE-Bench, which contains 251 expert-annotated queries and evaluates responses using a rubric-based framework calibrated to the requirements of the operational finance domain, across eight dimensions: accuracy, citations, clarity, depth, groundedness, recency, relevance, and structure. FORCE-Bench assesses agentic systems on three task types: financial obligation research (querying ERP systems for accounts receivable and payable data), financial entity performance research (answering time-bound questions from public filings and market data), and business brief generation (synthesising multi-source company intelligence reports). To reflect real deployment conditions, we evaluate our purpose-built agent, as well as the general-purpose agentic systems, under common tool access and latency-bounded settings. Results show that general-purpose agentic systems do not consistently meet finance-domain quality requirements under operational constraints, while the purpose-built Finance Agent for Microsoft 365 Copilot is more reliable across dimensions. We release the dataset, rubrics, harness, and analysis code as open-source to support reproducible comparison and adaptation to other enterprise finance environments.

2026-07-23 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

マルチモーダルエージェント検索におけるサイレントエラー:診断分類法と複数の裁判官による評価

マルチモーダル エージェント検索システムは、知識集約的な視覚的な質問に答えるために、外部ツールへの依存度が高まっています。ただし、既存の評価は主に最終的な回答の精度に焦点を当てており、検索軌跡の失敗を見逃してしまう可能性があります。この研究では、サイレント障害などの隠れた信頼性の問題を研究します。モダリティのショートカット、ファントムグラウンディング、間違った証拠と正解の事例、過剰検索ロンダリング、クロスモーダル矛盾、来歴幻覚をカバーする 6 つのカテゴリーの分類法を導入します。この分類に基づいて、統一された ReAct スタイルの足場の下で、回答の正しさと証拠の根拠の質の両方を評価する軌跡レベルの診断パイプラインを構築します。 4 つのフロンティア マルチモーダル モデルにわたる MMSearch-Plus 軌道に関する実験では、表面精度が真の軌道レベルの正確さを一貫して過大評価していることが示されています。さらに、クロスジャッジ検証、ブランク画像ストレステスト、およびツールアブレーションを使用して、サイレントエラーは能力に依存し、消えるのではなく変化することが多いことを示します。ホームページ: https://github.com/DingWu1021/silent-failures-multimodal-agentic-search

原文 (English)

Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation

Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing evaluations mainly focus on final-answer accuracy and may miss failures in the search trajectory. In this work, we study such hidden reliability issues as silent failures. We introduce a six-category taxonomy covering modality shortcuts, phantom grounding, wrong-evidence-right-answer cases, over-retrieval laundering, cross-modal contradiction, and provenance hallucination. Based on this taxonomy, we build a trajectory-level diagnostic pipeline that evaluates both answer correctness and evidence-grounding quality under a unified ReAct-style scaffold. Experiments on MMSearch-Plus trajectories across four frontier multimodal models show that surface accuracy consistently overestimates true trajectory-level correctness. We further use cross-judge validation, blank-image stress tests, and tool ablations to show that silent failures are capability-dependent and often shift rather than disappear. Home-page: https://github.com/DingWu1021/silent-failures-multimodal-agentic-search

2026-07-23 13:00 JSTarXiv cs.AIビジネス/資金調達

SenWorld: コンテキストリッチな評価データを生成するためのデジタル ツイン シミュレーション

スマートフォンのパーソナル アシスタントは長期にわたる個人データを推論しますが、その評価には正解がわかっているコンテキストに富んだ評価データが必要であり、実際のデバイスのトレースはプライバシーに敏感すぎて共有できません。この課題に対処するために、構築によって固定されたグラウンド トゥルースを使用してそのようなデータを生成する、物理的に接地され、決定論的でイベント ソースのデジタル ツイン シミュレーションである SenWorld を紹介します。 SenWorld では、ペルソナは実際の地図、天気、休日、ネットワーク データから構築された世界で 1 日を過ごします。観測可能なすべての信号はシステム全体のスナップショットにアーカイブされます。また、各評価ケースは、事後注釈や大規模言語モデル (LLM) ジャッジではなく、既存のレコードへのポインターによってラベル付けされます。この手法を北京の 16 人のペルソナで評価しました。生成されたデータは、カテゴリ分布 (ジェンセンとシャノンの相違 (JSD) 0.070) および通信記録の 1 日のリズム (JSD 0.1 未満) において、保持されている実際のユーザーのベンチマークと厳密に一致していますが、生成された記録は実際の記録よりも短いままです。スクリプトによる対話がなければ、ペルソナは完全に往復する対話サブグラフと差別化された行動レパートリーを形成します。 717 件の評価ケースに投影された生成データでは、実稼働スマートフォン アシスタントの 78 件の障害が明らかになり、通話とショート メッセージ サービス (SMS) の記録に集中し、連絡先、スケジュール、アラームは決して失敗しませんでした。スナップショット ポインタは、LLM 判定者が関与せずに、各失敗をアシスタント側の取得エラーとして確認します。全体として、SenWorld は、ラベルが構築によって固定されている評価データへの、プライバシーに安全で再現可能で配布がチェックされたパスを提供します。

原文 (English)

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data

Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, we present SenWorld, a physically grounded, deterministic, event-sourced digital-twin simulation that generates such data with ground truth fixed by construction. In SenWorld, personas live through a full day in a world built from real map, weather, holiday, and network data; every observable signal is archived in full-system snapshots; and each evaluation case is labeled by a pointer to an existing record rather than by post-hoc annotation or a large language model (LLM) judge. We evaluate this method with 16 personas in Beijing. The generated data closely matches the held-out real-user benchmark in category distribution (Jensen--Shannon divergence (JSD) 0.070) and in the daily rhythm of communication records (JSD below 0.1), though generated records remain shorter than real ones. Without scripted interaction, personas form a fully reciprocated dialogue subgraph and differentiated behavioral repertoires. Projected into 717 evaluation cases, the generated data exposes 78 failures in a production smartphone assistant, concentrating on call and Short Message Service (SMS) records while contacts, schedules, and alarms never fail. The snapshot pointer confirms each failure as an assistant-side retrieval error, with no LLM judge involved. Overall, SenWorld offers a privacy-safe, reproducible, and distribution-checked path to evaluation data whose labels are fixed by construction.

2026-07-23 13:00 JSTarXiv cs.AIビジネス/資金調達

言語モデルの経済的評価

言語モデルは経済的に価値のある作業を実行しますが、経済的に価値のあるすべてのタスクをどの程度うまく実行するかについては、現時点では評価されていません。米国の労働経済におけるタスク、作業活動、職業に関連する能力を測定するためのオープンソース評価スイートとして EconEvals を紹介します。可能な場合には、言語モデルに対する実際のユーザー クエリを評価スイートに基づいて作成し、これらを合成データで補完します。私たちの評価により、米国の職業の 5% をカバーする既存の最先端技術である OpenAI の GDPval ベンチマークよりもカバー範囲が向上しており、コストは 500 分の 1 です。ベンチマークに加えて、現在の言語モデル機能が米国のすべての職業に属するすべてのタスクにわたってどれだけの時間を節約できるかを推定するためのシミュレーションベースのエクスポージャー測定も導入し、それぞれの推定値を詳細に説明します。私たちの推定では、現在のモデルにより、労働者は 47% の職業において少なくとも半分の作業で大幅な時間を節約できる可能性があることが示されています。ただし、大幅な時間の節約が予測されるタスクの 79% では、観察されたクロードの使用量は低く、既存の使用量が潜在的なものより遅れていることを示唆しています。言語モデルのチャットボットに固有の制約を超えて、私たちのデータは、プライバシーと独自のシステムが AI によるさらなる時間節約を制限する主なボトルネックであることを特定しています。全体として、言語モデルの現在の機能における労働市場への影響についての推論を根拠づける、適応可能なインフラストラクチャを導入します。これは、機能の向上に応じて継続的に更新できます。

原文 (English)

Economic Evaluations of Language Models

Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an open-source evaluation suite to measure capabilities relevant to tasks, work activities, and occupations in the US labor economy. We ground the evaluation suite in real user queries to language models where possible, and supplement these with synthetic data. Our evaluations improve coverage over OpenAI's GDPval benchmark, which is the existing state-of-the-art that covers 5% of US occupations, at 500x lower cost. Alongside benchmarks, we also introduce a simulation-based exposure measure to estimate how much time current language model capabilities could save across all tasks belonging to all US occupations, with detailed accounting for each estimate. Our estimates indicate that current models could save workers substantial time on at least half of their tasks in 47% of occupations. However, for 79% of tasks where we predict substantial time savings, observed Claude usage is low, suggesting that existing usage lags potential. Beyond inherent constraints of language model chatbots, our data identifies privacy and proprietary systems as the principal bottlenecks limiting further time savings from AI. Overall, we introduce adaptable infrastructure that grounds inferences about language models' labor-market impact in their current capabilities, which can be continually updated as capabilities improve.

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

JailMeter: 大規模な言語モデルに対するジェイルブレイク攻撃のための証拠に基づいた評価フレームワーク

大規模な言語モデルに対するジェイルブレイク攻撃の評価は現在、一貫性のない評価基準と方法に悩まされており、攻撃成功率の推定値の信頼性が低くなります。私たちは、ジェイルブレイクの有効性をより忠実に測定するために設計された証拠に基づいた評価フレームワークである JailMeter を提案します。情報ボトルネック理論にヒントを得た JailMeter は、デュアルフィードバック最適化を適用して、元の悪意のある質問に関連するコンテンツを維持しながら、モデル応答からジェイルブレイク ノイズをフィルターします。このプロセスは、応答が悪意のある意図を捕捉し、完全な応答を提供する場合にのみ攻撃が検証される、厳密な評価のための簡潔な証拠を生成します。これにより、モデルの安全性調整の実質的なバイパスが示されます。私たちは、JailMeter-Eva で JailMeter を評価します。これは、人間がラベルを付けた拒否されていないジェイルブレイク インスタンス 330 個を含む、挑戦的なベンチマークです。 JailMeter は 97.27% の精度を達成し、既存の評価方法を大幅に上回ります。大規模な評価をサポートするために、JailMeter をさらに小規模な言語モデル JailMeter\textsubscript{SLM} に抽出しました。これは、計算コストを大幅に削減しながら同等の信頼性を維持します。コードとデータセットは https://github.com/Magi2B0y/JailMeter で入手できます。

原文 (English)

JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an evidence-based evaluation framework designed to more faithfully measure jailbreak effectiveness. Inspired by the Information Bottleneck theory, JailMeter applies dual-feedback optimization to filter jailbreak noise from model responses while preserving content relevant to the original malicious question. This process produces concise evidence for a rigorous assessment under which an attack is validated only when the response captures the malicious intent and delivers a complete answer, thereby signaling a substantive bypass of model safety alignment. We evaluate JailMeter on JailMeter-Eva, a challenging benchmark containing 330 human-labeled, non-rejected jailbreak instances. JailMeter achieves an accuracy of 97.27%, substantially outperforming existing evaluation methods. To support large-scale evaluation, we further distill JailMeter into a small language model, JailMeter\textsubscript{SLM}, which maintains comparable reliability with significantly reduced computational costs. Code and dataset are available at https://github.com/Magi2B0y/JailMeter.

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

スケープゴートとしてのガードレール: ツールで強化された LLM エージェントにおける不誠実な安全拒否の監査

ツールで強化された LLM エージェントの評価フレームワークは、機能メトリクスまたは明示的なツールのクラッシュに圧倒的に焦点を当てており、サイレント インフラストラクチャの障害や、空、null、または不正な形式のペイロードを含む HTTP 200 応答はほとんど監査されません。軽量のブラックボックス監査フレームワークを導入します。このフレームワークは、12 の運用環境に隣接するツール スタブに 4 つのサイレント障害プロファイルを挿入し、エージェントの応答を 3 つの相互に排他的な動作クラス、Honest Surrender (HSR)、 Fabrication (FAR)、および Unfaithful Safety Refusal (USR) に分類します。中立システム プロンプトの下、温度ゼロで 2 つのフロンティア モデルと 2 つのオープンソース モデルを評価したところ、FAR が優勢であることがわかりました (有効回答の 56.6%)。エージェントは空のペイロードを実際のデータとして扱い、黙って捏造された結果を返します。エージェントが失敗を説明するためにポリシーやプライバシーの根拠を発明する USR は、ベースラインではほとんど存在しません (0.25%、396 の有効な軌跡全体で 1 つのインスタンス)。私たちの重要な発見は、システムのプロンプトに標準的な安全性の文言 (「ユーザーのプライバシーとデータのセキュリティを優先する」) を追加したアブレーションから明らかになり、これにより USR が 15.6 倍に増幅されます (0.25% から 3.95%、アブレーション率の 95% CI: 2.2% ~ 6.4%、フィッシャーの直接確率検定、p < 0.001)。 USR は潜在的な動作であり、ツールがサイレントに失敗したときに、システム プロンプト内の安全ボキャブラリーがモデルにポリシーの理論的根拠に到達するよう促すときにアクティブになります。機密性の高いツール (fetch_medical_record、retrieve_contract、fetch_user_profile) が USR インスタンスの大部分を占めます。実稼働レベルの検出のためのペイロード応答の不整合ヒューリスティックを提案し、セーフティフォワード展開におけるガバナンスへの影響について議論します。

原文 (English)

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely unaudited. We introduce a lightweight black-box auditing framework that injects four silent failure profiles across 12 production-adjacent tool stubs and classifies agent responses into three mutually exclusive behavioral classes: Honest Surrender (HSR), Fabrication (FAR), and Unfaithful Safety Refusal (USR). Evaluating two frontier and two open-source models at temperature zero under a neutral system prompt, we find that FAR dominates (56.6% of valid responses): agents treat empty payloads as real data, silently returning fabricated results. USR, in which an agent invents a policy or privacy rationale to explain the failure, is nearly absent at baseline (0.25%, one instance across 396 valid trajectories). Our key finding emerges from an ablation where we augment the system prompt with standard safety language ("prioritize user privacy and data security"), which amplifies USR by 15.6x (from 0.25% to 3.95%; 95% CI on ablation rate: 2.2%-6.4%; Fisher's exact test, p < 0.001). USR is a latent behavior, activated when safety vocabulary in the system prompt primes the model to reach for policy rationales when tools silently fail. Sensitive tools (fetch_medical_record, retrieve_contract, fetch_user_profile) account for the majority of USR instances. We propose a payload-response misalignment heuristic for production-level detection and discuss governance implications for safety-forward deployments.

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

自由回答形式の質問応答における推論の参考資料なしの評価

一か八かの分野で AI が生成した答えは、多くの場合流暢ですが、特に単一の最終的な答えではなく複数のステップからなる推論が含まれている場合、検証が困難です。私たちは、LLM によって生成された出力を監査するための、推論ベースで参照不要のフレームワークを提案します。この方法では、生成された推論トレースをセグメントに分解し、自然言語推論 (NLI) を使用してローカルな前提とターゲットの関係にラベルを付け、これらの関係をハイパーグラフに編成します。次に、決定論的な後方 AND-OR 検索により、生成された応答内で各セグメントがどのように根拠づけられているかを示すセグメント レベルの監査ラベルが割り当てられます。このフレームワークを 2 つの設定で評価します。Hard2Verify による演繹的数学的推論と、実際の臨床例からの LLM 推論トレースの医師注釈付きの新しいベンチマークである UroReason によるオープンエンドの医学的推論です。これらの設定全体にわたって、NLI ハイパーグラフ監査は、裁判官としての直接の LLM ベースラインよりも信頼性の高いリファレンスフリーの評価シグナルを提供します。臨床現場では、最先端の LLM 裁判官が問題のある推論セグメントを特定できず、流暢ではあるが根拠の薄い応答を過剰に受け入れてしまうことがよくあります。私たちの結果は、QA 評価では、最終的な回答や検証者としての LLM のみに依存するのではなく、推論トレース全体で推論関係がどのように構成されるかを考慮する必要があることを示しています。 UroReason は API を通じて利用可能になり、コードはオープンソースとしてリリースされます。

原文 (English)

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for auditing LLM-generated outputs. The method decomposes a generated reasoning trace into segments, labels local premise-target relations using Natural Language Inference (NLI), and organizes these relations into a hypergraph. A deterministic backward AND-OR search then assigns segment-level audit labels that indicate how each segment is grounded within the generated response. We evaluate the framework in two settings: deductive mathematical reasoning with Hard2Verify, and open-ended medical reasoning with UroReason, a new physician-annotated benchmark of LLM reasoning traces from real clinical cases. Across these settings, our NLI-hypergraph audit provides a more reliable reference-free evaluation signal than direct LLM-as-judge baselines. In the clinical setting, state-of-the-art LLM judges often fail to identify problematic reasoning segments, over-accepting fluent but weakly grounded responses. Our results show that QA evaluation should account for how inferential relations compose across a reasoning trace, rather than relying only on final answers or LLMs as verifiers. UroReason will be made available through an API, and our code will be released as open source.

2026-07-23 13:00 JSTarXiv cs.AIビジネス/資金調達

Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks

We propose that AI automation is a continuum between: (i) crashing waves where AI capabilities surge abruptly over small sets of tasks, and…

2026-07-23 13:00 JSTarXiv cs.AIビジネス/資金調達

擬人化された対話に向けて: 人間のようなチャットの生成、評価、および好みの調整のための閉ループ フレームワーク

人間のようなプライベート チャットには、流暢な応答生成以上のものが必要です。システムは、ペルソナ、関係、記憶、限定された知識、媒体固有のタイミング、一貫したマルチターン アークを保持する必要があります。我々は、擬人化対話をシステム アーキテクチャ、実行可能評価、および診断調整の共同問題として定式化する閉ループ フレームワークである AnthroDial を紹介します。これは、(1) 役割条件付きのスケジュールされた対話ランタイムと、ペルソナおよびシナリオ カード、長期記憶、仮想時間、および単一草案メッセージの決定を組み合わせます。 (2) L0 有効性ゲート、5 つのターンごとのディメンション、および 5 つのダイアログ レベルのディメンションを備えた実行可能なベンチマーク。 (3) SFT 用に 16,436 のスケジュールされた決定例をフィルタリングし、認知診断、ZPD を意識した報酬を備えた GRPO を適用するトレーニング後のパイプライン。この報酬は、各行動次元のカルマン フィルター処理された能力推定値を維持し、より大きな能力不足のある次元を重み付けし、ロールアウト スコアをタスク レベルの ZPD マッチとして使用して、学習可能な弱いスキルに焦点を合わせて最適化します。モデルごとに 55 のペルソナ、50 のシナリオ、50 のペルソナとシナリオのバインディング、および 100 の役割条件付きケースを含むベンチマークで、フロンティア ベースライン、オープン モデル、思考/非思考のバリアント、および SFT/RL アブレーションにわたる 16 のシステムを評価します。最も強い非トレーニングベースラインは 32.00% の厳密な ACC に達しますが、Qwen3.6-27B-SFT+RL は 39.00% の厳密な ACC と 98.5 の全体スコアに達します。 9B ノーシンク設定では、SFT と RL は厳密な ACC を 0.00% から 13.00% および 18.37% に改善します。これらの結果は、生成、評価、報酬形成が同じ行動次元を共有する場合、擬人化対話が有益であることを示しています。

原文 (English)

Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc. We present AnthroDial, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment. It combines (1) a role-conditioned scheduled dialogue runtime with persona and scenario cards, long-term memory, virtual time, and single-draft message decisions; (2) an executable benchmark with an L0 validity gate, five per-turn dimensions, and five dialogue-level dimensions; and (3) a post-training pipeline that filters 16,436 scheduled-decision examples for SFT and applies GRPO with a cognitive-diagnostic, ZPD-aware reward. The reward maintains Kalman-filtered capability estimates for each behavioral dimension, upweights dimensions with larger capability deficits, and uses rollout scores as task-level ZPD matches to focus optimization on learnable weak skills. On a benchmark with 55 personas, 50 scenarios, 50 persona-scenario bindings, and 100 role-conditioned cases per model, we evaluate 16 systems spanning frontier baselines, open models, thinking/no-think variants, and SFT/RL ablations. The strongest non-trained baseline reaches 32.00% strict ACC, while Qwen3.6-27B-SFT+RL reaches 39.00% strict ACC and a 98.5 overall score. In the 9B no-think setting, SFT and RL improve strict ACC from 0.00% to 13.00% and 18.37%. These results show that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.

2026-07-23 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

CGCE: Classifier-Guided Concept Erasure in Generative Models

Recent advancements in large-scale generative models have enabled the creation of high-quality images and videos, but have also raised sign…

2026-07-23 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Comparative evaluation of training strategies using partially labelled datasets for segmentation of white matter hyperintensities and stroke lesions in FLAIR MRI

White matter hyperintensities (WMH) and ischaemic stroke lesions (ISL) are key imaging biomarkers of cerebral small vessel disease (SVD) de…

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB)…

2026-07-23 07:00 JSTITmedia AI+ハードウェア/半導体ビジネス/資金調達

NVIDIAフアンCEOが語る“日本復活”のシナリオ 10年続く半導体バブルと「原発活用」の勝算

米NVIDIAのジェンスン・フアンCEOが来日し、日本経済の復活を宣言した。国内のAIインフラ構築へ数十億ドル規模の投資を発表。フアン氏は「何兆ものAIがAIを使う時代」の到来によって半導体需要は人口に制約されないと指摘。データセンターの電力不足に対して「原発活用」を日本の強み…

2026-07-23 06:55 JSTITmedia AI+LLM/生成AIハードウェア/半導体ビジネス/資金調達

AMDとAnthropicが戦略的提携 「Helios」を最大2GW導入、最大50億ドルの出資も

AMDは、Anthropicとの戦略的提携を発表した。AnthropicはAMDの「Helios」および「Instinct MI450」シリーズを最大2GW規模で導入し、2027年上半期から順次展開する。AMDは最大50億ドルの株式投資を行うほか、Claudeを活用したGPU環…

2026-07-23 03:50 JSTTechCrunch AIロボティクスビジネス/資金調達

Travis Kalanick’s robotics company raises $1.7B, led by a16z

Uber is also investing in Travis Kalanick's company Atoms, which has made gauzy claims about using industrial AI to modernize the world.

2026-07-23 03:13 JSTTechCrunch AIビジネス/資金調達

Yope raises $12.3M to build a private social network without algorithms or ads

Yope, a fast-growing social app focused on private groups of friends and family, has raised $12.3 million in seed funding. Instead of chasi…

2026-07-22 22:00 JSTTechCrunch AIビジネス/資金調達

Passionfroot raises $15M to expand its B2B creator marketplace to the US

Passionfroot, a German startup building a marketplace connecting B2B creators with brands, has raised $15M in a Series A round led by Insig…

2026-07-22 22:00 JSTOpenAILLM/生成AIビジネス/資金調達

Building AI infrastructure with the Effingham County community

OpenAI announces Project Camellia in Effingham County, Georgia, with commitments to responsible energy, community investment, jobs, and acc…

2026-07-22 19:00 JSTTechCrunch AIエージェントビジネス/資金調達

Glow emerges from stealth at $1.2B valuation to challenge endpoint security in the AI era

Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises.

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

証拠連鎖評価による校正済みの選択的ファクトチェック

大規模言語モデル (LLM) は強力なファクトチェック精度を実現できますが、強制的な二者択一の決定により、重大な信頼性の問題が隠蔽されます。つまり、裏付けとなる証拠が弱い、希薄である、または内部的に矛盾している場合でも、システムは自信を持って判定を下す可能性があります。私たちは、すべての主張に対して真/偽の決定を要求するのではなく、不確実な評決によって棄権を許可する選択的事実確認フレームワークである証拠連鎖評価(ECE)を通じてこの問題に取り組んでいます。評価されたシステムは、Web 検索、学術検索、実行可能チェックを通じて証拠を収集し、信頼性とソースレベルのメタデータを含む構造化された判定を返すツールを使用する検証エージェントです。 ECE-Bench では、ECE は回答されたクレームに対して 91.6% の標準精度、93.7% のカバレッジ、および 97.8% の選択的精度を達成しています。 ECE は、予想されるキャリブレーション エラー、ブライアー スコア、または AURC などの集計キャリブレーション メトリクスに関して最も強力な検索ベースラインを上回るパフォーマンスはありませんが、明確な選択的予測のトレードオフを提供します。つまり、システムは、95 件中 6 件を保留しながら、回答されたクレームについて非常に高い精度を維持します。これらの延期されたケースは、信頼性の低い証拠設定(情報源レベル L4 で 5/6)に集中しており、棄権が認識論的に弱い証拠を処理するための安全志向のメカニズムとして機能するという見解を裏付けています。コードは https://github.com/cheshireyang/ECE.git で入手できます。

原文 (English)

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions conceal a critical reliability problem: systems may issue confident verdicts even when supporting evidence is weak, sparse, or internally inconsistent. We address this issue through Evidence Chain Evaluation (ECE), a selective fact-checking framework that permits abstention via an uncertain verdict instead of requiring a true/false decision for every claim. The evaluated system is a tool-using verification agent that gathers evidence through web search, scholarly search, and executable checks, and then returns a structured verdict with confidence and source-level metadata. On ECE-Bench, ECE achieves 91.6% standard accuracy, 93.7% coverage, and 97.8% selective accuracy on answered claims. Although ECE does not outperform the strongest retrieval baseline on aggregate calibration metrics such as Expected Calibration Error, Brier score, or AURC, it delivers a clear selective-prediction trade-off: the system maintains very high accuracy on answered claims while deferring 6 of 95 cases. These deferred cases are concentrated in lower-reliability evidence settings (5/6 at source level L4), supporting the view that abstention functions as a safety-oriented mechanism for handling epistemically weak evidence. Code is available at https://github.com/ cheshireyang/ECE.git

2026-07-22 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SAAG: 構造化されたエージェントの評価とグラウンディング

エージェント呼び出しの完全一致評価は、質的に異なる障害モードを曖昧にします。モデルは正しい関数を選択しているにもかかわらず引数値を幻覚させたり、誤った理由でエージェントを選択しながらスキーマを満たしたりする可能性があります。既存のベンチマークでは、これらの区別が単一のバイナリ スコアにまとめられているため、担当者はエージェントの呼び出しがどこで失敗するかを診断できません。我々は、エージェント呼び出しの評価を、レジストリ適合性、構造的完全性、議論の根拠という 3 つの段階に分解し、それぞれが解釈可能な段階固有の診断を生成するカスケード診断フレームワークを SAAG に提案します。さらに、これらの診断により、反復的な自己修復が可能になります。つまり、予測が失敗した場合、ステージ固有の信号が、グランドトゥルース値を漏らすことなく、ターゲットを絞った修正をガイドします。このフレームワークは、3 つのローカル サブ 4B パラメーター モデルを使用して、5、10、および 15 エージェントのレジストリ サイズにわたる Glaive の関数呼び出しデータセットから導出された制御されたベンチマークで評価されます。構造化フィードバックにより引数の精度が一貫して向上し、シングルパス推論や有益でないバイナリ フィードバックと比較して値の幻覚が軽減されますが、エンドツーエンドの F1 ゲインは控えめでモデルに依存します。これらの結果は、段階分解された診断評価が、モデル ファミリやレジストリ スケール全体でのエージェント呼び出しの信頼性を理解し、改善するために必要なレンズであることを示唆しています。

原文 (English)

SAAG: Structured Agent Assessment and Grounding

Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or satisfy a schema while choosing a agent for the wrong reason. Existing benchmarks collapse these distinctions into a single binary score, leaving practitioners unable to diagnose where agent calls fail. We propose SAAG a cascaded diagnostic framework that decomposes agent-calling evaluation into three sequential stages: registry conformance, structural completeness, and argument grounding, each producing interpretable stage-specific diagnostics. These diagnostics additionally enable iterative self-repair: on prediction failure, the stage-specific signal guides targeted correction without leaking ground-truth values. We evaluate this framework on a controlled benchmark derived from Glaive's function-calling dataset across registry sizes of 5, 10, and 15 agents using three local sub-4B-parameter models. Structured feedback consistently improves argument precision and reduces value hallucination relative to single-pass inference and uninformative binary feedback, while end-to-end F1 gains are modest and model-dependent. These results suggest that stage-decomposed diagnostic evaluation is a necessary lens for understanding and improving agent-calling reliability across model families and registry scales.

2026-07-22 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

再トレーニングを行わない方言間の一般化: MLIR のスキーマ導出制約付きデコーディングのベンチマークと評価

マルチレベル中間表現 (MLIR) は、最新の ML コンパイラ インフラストラクチャ (TensorFlow、JAX/StableHLO、PyTorch Inductor、IREE) の基礎を成していますが、コード LM 事前トレーニング コーパスには微量しか現れません。 MLIR は設計上も拡張可能です。新しい方言はアプリケーション ドメインごとに出荷されるため、方言ごとに微調整されたモデルは拡張できません。各方言の操作定義仕様 (ODS) から機械的に導出された推論時間事前確率が勾配ベースの適応の代わりに使用できるかどうかを尋ねます。まず、3 つの方言にわたる 4 つの自然言語から MLIR ベンチマーク (MLIR-Spec-150、Linalg-Spec-30、StableHLO-Spec-30、StableHLO-Held-Out-200) をリリースします。合計 410 の対象範囲内の NL から MLIR ペアに加え、25 のプログラムの文法外のストレス セットと手書きの n=30 の関数リファレンス セットが含まれており、以下で出荷されます。 Apache-2.0 と Gebru データシートおよび Croissant 1.0 メタデータ。次に、3 層のスキーマ派生制約スタックを構築します。OP シグネチャ上の CFG (C1)、ODS 抽出された型ラティスからの型ドメイン分割 (C2)、5 回の再試行拒否サンプリングを駆動する SSA スコープ バリデータ (C3) です。 arith+func+memref+linalg から StableHLO への移植には、新しい制約層コードは必要ありません。検証者のセマンティクスが構造的制約によって支配されている方言では、スキーマ由来の事前分布により、SmolLM2-1.7B は世代ごとの 8 ~ 25 倍の速度で 15B-34B コード LM と一致またはそれを超えます。linalg では、SmolLM2 は 80.0% verify-valid (3 シード平均、n=125) に達し、CodeLlama-34B を上回ります。 Granite-Code-34B および StarCoder2-15B は、重複しない CI で 21 ~ 44 パーセント ポイント減少しました。 arith+func とテンプレート化されたパラメトリック StableHLO-Held-Out-200 では、検証者のセマンティクスが構造ではなく属性値をオンにするため、同じベースラインが SLM に一致するか、それを上回ります。これらを非勝利セルとしてスコープします。ベンチマーク、デコーダー、プロンプトごとのすべての生成、再現性のある Docker イメージをリリースします。

原文 (English)

Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR

Multi-Level Intermediate Representation (MLIR) underlies modern ML compiler infrastructure (TensorFlow, JAX/StableHLO, PyTorch Inductor, IREE), yet appears only in trace amounts in code-LM pretraining corpora. MLIR is also extensible by design: new dialects ship per application domain, so a fine-tuned model per dialect does not scale. We ask whether inference-time priors derived mechanically from each dialect's Operation Definition Specification (ODS) can substitute for gradient-based adaptation. First, we release four natural-language-to-MLIR benchmarks across three dialects - MLIR-Spec-150, Linalg-Spec-30, StableHLO-Spec-30, and StableHLO-Held-Out-200 - totaling 410 in-scope NL-to-MLIR pairs, plus a 25-program out-of-grammar stress set and a hand-authored n=30 functional reference set, shipped under Apache-2.0 with Gebru datasheets and Croissant 1.0 metadata. Second, we build a three-layer schema-derived constraint stack: a CFG over op signatures(C1), type-domain splits from an ODS-extracted type lattice (C2), and an SSA-scope validator driving five-retry rejection sampling (C3). Porting from arith+func+memref+linalg to StableHLO required no new constraint-layer code. On dialects whose verifier semantics are dominated by structural constraints, schema-derived priors let SmolLM2-1.7B match or exceed 15B-34B code LMs at 8-25x the per-generation speed: on linalg, SmolLM2 reaches 80.0% verify-valid (three-seed mean, n=125), beating CodeLlama-34B, Granite-Code-34B, and StarCoder2-15B by 21-44 percentage points with non-overlapping CIs. On arith+func and on the templated parametric StableHLO-Held-Out-200, where verifier semantics turn on attribute values rather than structure, the same baselines match or beat the SLM; we scope these as non-win cells. We release benchmarks, decoder, all per-prompt generations, and a reproducibility Docker image.

2026-07-22 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

地域電力市場における需要家向けの市場戦略評価

分散型発電と柔軟な負荷を備えた需要家は、地域資源を制御する自律的なサイバー物理エネルギー システムを形成し、人間の介入を最小限に抑えて地域のエネルギー市場に参加します。この研究では、エージェント ベースのシミュレーション プラットフォームを開発および評価します。このプラットフォームでは、太陽光発電システム、蓄電池システム、電気自動車、ヒート ポンプを備えた消費者世帯を代表するエージェントが、均一価格の両面コール オークションに参加します。個々の入札戦略がコミュニティレベルの効率性やプロシューマーレベルの財務成果に及ぼす影響は、特に異種のポートフォリオを持つプロシューマーが 1 つの市場で相互作用する場合には完全には理解されていません。複雑さが増す 4 つの市場戦略、つまりゼロインテリジェンス制約ベースライン、境界価格戦略、拡張ストレージ カスケード、市場適応型価格戦略を比較します。このシミュレーションは、季節変動を特徴付けるために、夏、冬、春にわたる 15 分の解像度で 33 人のプロシューマーのコミュニティに対して実行されます。結果は、ルールベースのリソース制御により地域社会のエネルギー支出が大幅に削減されることを示しています。拡張された貯蔵カスケードでは総コストが 39.06 ユーロに達し、ゼロインテリジェンスベースラインの場合は 62.38 ユーロとなり、37.4 % 削減されました。市場適応戦略は、夏の条件下で、地元のエネルギー市場への参加を通じて、地域全体で最高の経済的利益をもたらします (ベースラインの 14.40 ユーロ対 10.28 ユーロ、40.1 % の利益)。戦略の有効性はポートフォリオの構成と季節的な供給条件の両方に依存するため、リソース管理と価格決定の共同評価が必要です。

原文 (English)

Market Strategy Evaluation for Prosumers in Local Electricity Markets

Prosumers equipped with distributed generation and flexible loads form autonomous cyber-physical energy systems that control local resources and participate in local energy markets with minimal human intervention. This work develops and evaluates an agent-based simulation platform in which agents, representing prosumer households with photovoltaic systems, battery storage systems, electric vehicles, and heat pumps, participate in a uniform-price double-sided call auction. The effect of individual bidding strategies on community-level efficiency and prosumer-level financial outcomes is incompletely understood, particularly when prosumers with heterogeneous portfolios interact in one market. Four market strategies of increasing complexity are compared: a zero-intelligence constrained baseline, a boundary-price strategy, an extended storage cascade, and a market-adaptive pricing strategy. The simulation is conducted on a community of 33 prosumers at 15-minute resolution, spanning summer, winter, and spring to characterize seasonal variation. Results show that rule-based resource control substantially reduces community energy expenditure: the extended storage cascade achieves a total cost of 39.06 EUR compared to 62.38 EUR under the zero-intelligence baseline, a reduction of 37.4 %. The market-adaptive strategy yields the highest aggregate community financial gain through local energy market participation (14.40 EUR vs. 10.28 EUR for the baseline, a gain of 40.1 %) under summer conditions. Strategy effectiveness depends on both portfolio composition and seasonal supply conditions, requiring joint evaluation of resource control and pricing decisions.

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

グレーボックス シミュレーション モデルのエージェント キャリブレーション: LLM 主導の代替手段

グレーボックス シミュレーション モデルのキャリブレーションは、モデルの評価にコストがかかり、パラメーター空間が高次元になる可能性があり、検索では妥当性の制約を考慮する必要がある制約付きの最適化問題です。解析者はシミュレーション コードを完全に利用できますが、複数のパラメーターの共同効果を解析的に予測することは依然として困難です。 Nelder--Mead (NM) などの従来のオプティマイザは導入が簡単ですが、特に制約がある場合にはサンプルの効率が悪くなります。最新のベイジアン最適化手法は、はるかに少ない評価で競争力のあるソリューションを実現しますが、制約を処理するために自明ではないモデリング機械を必要とします。私たちは、大規模な言語モデルがオプティマイザーとして機能し、制約がシステム プロンプトの平易な言語セクションとして組み込まれるエージェント キャリブレーション方法を導入します。非制約キャリブレーションと臨床制約キャリブレーションの両方の下で、肛門癌シミュレーション モデルのエージェント手法、NM、およびベイジアン最適化 (BO) を評価します。制約のないキャリブレーションでは、エージェント手法は BO および NM よりも大幅に低い最良誤差を達成しながら、必要なモデル評価の数は少なくなります。制約されたキャリブレーションの下では、エージェント手法は同等の誤差レベルに達し、両方とも NM を上回ります。これらの結果は、反復ごとの推論時間の増加を犠牲にして得られます。エージェント キャリブレーションは、実質的に少ないモデル評価で競争力のあるパフォーマンスを実現し、追加のモデリング機構ではなく単純なテキスト仕様を通じて、モデラー側のインターフェイスでの制約処理が基本的に無料になります。主なトレードオフは反復ごとの推論コストの増加にあり、このアプローチはシミュレーション時間が支配的な場合に特に適しています。パフォーマンスを超えて、反復ごとの理論的根拠により、検索が監査可能で説明可能になるため、その決定を精査して第三者に対して正当化することができます。

原文 (English)

Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative

Calibration of grey-box simulation models is a constrained optimization problem in which model evaluations are expensive, the parameter space can be high-dimensional, and the search must respect plausibility constraints. Although the simulation code is fully available to the analyst, the joint effect of multiple parameters remains difficult to predict analytically. Classical optimizers such as Nelder--Mead (NM) are simple to deploy but sample-inefficient, particularly under constraints. Modern Bayesian Optimization methods achieve competitive solutions with far fewer evaluations but require non-trivial modeling machinery for constraint handling. We introduce an agentic calibration method in which a large language model acts as the optimizer, with constraints incorporated as a plain-language section of the system prompt. We evaluate the agentic method, NM, and Bayesian Optimization (BO) on an anal cancer simulation model under both unconstrained and clinically constrained calibration. Under unconstrained calibration, the agentic method achieves substantially lower best error than BO and NM, while requiring fewer model evaluations. Under constrained calibration, the agentic method reaches comparable error levels and both outperform NM. These results are obtained at the cost of increased inference time per iteration. Agentic calibration achieves competitive performance with substantially fewer model evaluations, and constraint handling is essentially free at the modeller-facing interface through simple textual specifications rather than additional modelling machinery. The main trade-off lies in increased per-iteration inference cost, making the approach particularly suitable when simulation time dominates. Beyond performance, the per-iteration rationale makes the search auditable and explainable, so its decisions can be scrutinised and justified to third parties.

2026-07-22 13:00 JSTarXiv cs.AIビジネス/資金調達

Estimating Rare Events in Language Models with Proper Evaluation

Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale de…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Ex…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism

We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation using large language models (L…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

Multi-turn medical consultation agents must decide what to ask, adapt to patient responses, and determine when the collected evidence is su…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams

Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot…

2026-07-22 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and…

2026-07-22 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents. SENTINEL is th…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mechanistic interpretability, and governance/auditability, covering 2018-2026 evaluation-safety measurement work. We introduce EvalSafetyGap as an organizing hypothesis for comparing evaluation-side and alignment-side proxy failures under optimization pressure, using Goodhart's Law together with two constructs we develop here - an Instability Decomposition and an Alignment Trilemma - as tools for generating testable comparisons. The audit shows how conclusions shift when capability, behavioral safety, and governance are measured separately. In this sample ($n = 10$), the association between capability and sustained adversarial robustness is statistically indeterminate using the displayed Table 3 inputs (Pearson $r = +0.232$, $p = 0.520$), and the apparent open-closed safety gap is modest, driven mainly by governance and disclosure rather than behavioral robustness, and sensitive to how a single borderline model is classified; attempt-budget results are protocol dependent. Because the public evidence uses heterogeneous protocols, the audit is diagnostic rather than rank-generating. The contribution is a shared vocabulary and evidence map to support dynamic evaluation, transparent source reporting, multi-attempt safety measurement, and auditable alignment practice.

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

科学的視覚化リテラシーのためのマルチモーダル大規模言語モデルのベンチマーク

マルチモーダル大規模言語モデル (MLLM) は、ビジュアライゼーションを解釈するためにますます使用されていますが、現在の評価は依然として主にチャート中心であり、科学的ビジュアライゼーション (SciVis) の理解を示す証拠は限られています。私たちは、科学的視覚化リテラシー評価テストで 6 つの MLLM をベンチマークします。このテストは、8 つのテクニックと 11 のタスク タイプにわたる、18 の科学的視覚化とイラストに基づく 49 項目で構成される標準化された SciVis リテラシー評価です。私たちは、クローズドワールドプロトコルの下で 3 つのクローズドソースモデルと 3 つのオープンソースモデルを評価し、485 人の人間の参加者からのデータを使用してパフォーマンスを比較します。結果は、現在の MLLM が均一な SciVis リテラシーを示さないことを示しています。 Gemini は全体として最も強力なモデルであり、評価されたサブセット全体で人間の平均を上回っていますが、オープンソース モデルは依然として人間のベースラインを下回っています。パフォーマンスはテクニックやタスクによって大きく異なります。モデルは科学的なイラスト、検索、空間理解では最高のパフォーマンスを発揮しますが、テクスチャ ベースおよび統合ベースの視覚化と定量的推定では苦戦します。エラー分析により、きめの細かい定量的推定、フロー方向の解釈、および根拠のあるエンコードの解釈における繰り返しの失敗が明らかになります。これらの調査結果は、SciVis リテラシーをマルチモーダル AI システムを評価するために必要なベンチマークの側面として位置づけています。コードとモデルの出力は、https://github.com/patdmp/mllm-scivis-lit-benchmark で公開されています。

原文 (English)

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at https://github.com/patdmp/mllm-scivis-lit-benchmark.

2026-07-22 13:00 JSTarXiv cs.AIビジネス/資金調達

DL オントロジーの認識的機密性ポリシーに基づく扱いやすいクエリ応答 (拡張バージョン)

私たちは、記述ロジック (DL) オントロジーのコンテキストで、また認識依存関係 (ED) を通じて表現される機密性ポリシーについて、機密性を保持するデータ アクセスへの宣言的アプローチである制御クエリ評価 (CQE) を研究します。まず、CQE の既知のセマンティクス (GA および IGA 含意) の下でクエリ (具体的には論理積クエリのブール和集合) に応答する問題に取り組みます。私たちの結果は、TBox が $\text{DL-Lite}_{\mathcal{R}}$ で表現される場合、CQE は一般に計算的に扱いにくいことを示しています。さらに、ED が存在する場合、IGA セマンティクスは、識別不可能性として知られる重要な機密保持特性を満たさないことが最近証明されました。計算が容易で機密性が保たれる CQE 形式を定義することを目的として、最小限のポリシー違反 (MPV) の概念に基づいた CQE の新しいセマンティクスを導入します。新しいセマンティクスが以前のセマンティクスの健全な近似を提供しながら、区別不可能性の特性を満たしていることを示します。また、$\text{DL-Lite}_{\mathcal{R}}$ オントロジーの場合、MPV セマンティクスに基づくクエリ含意がデータ複雑さの多項式時間で決定できることも証明します。最後に、OWL 2 QLの既存のベンチマークを使用して、この新しいアプローチの実現可能性を評価するために使用したフレームワークのソフトウェア実装を紹介します。

原文 (English)

Tractable Query Answering under Epistemic Confidentiality Policies in DL Ontologies (extended version)

We study Controlled Query Evaluation (CQE), a declarative approach to confidentiality-preserving data access, in the context of Description Logic (DL) ontologies, and for confidentiality policies expressed through Epistemic Dependencies (EDs). We first address the problem of answering queries (specifically, Boolean unions of conjunctive queries) under known semantics for CQE (GA- and IGA-entailment). Our results show that if the TBox is expressed in $\text{DL-Lite}_{\mathcal{R}}$, CQE is computationally intractable in general. Moreover, in the presence of EDs, the IGA semantics has recently been proven not to satisfy an important confidentiality preservation property known as indistinguishability. With the goal of defining computationally easier and confidentiality-preserving forms of CQE, we introduce a new semantics for CQE, based on the notion of minimal policy violation (MPV). We show that the new semantics provides a sound approximation of the previous ones, while satisfying the indistinguishability property. We also prove that, in the case of $\text{DL-Lite}_{\mathcal{R}}$ ontologies, query entailment under the MPV semantics can be decided in polynomial time in data complexity. Finally, we present a software implementation of our framework that we used to evaluate the feasibility of this new approach using an existing benchmark for OWL 2 QL.

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Saving the legacy of Hero Ibash: Evaluating Four Language Models for Aminoacian

This study assesses four cutting-edge language models in the underexplored Aminoacian language. Through evaluation, it scrutinizes their ad…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have bec…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Doctorina MedBench-ICD10: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI

We present Doctorina MedBench, a comprehensive evaluation framework for agent-based medical AI based on the simulation of realistic physici…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

A language model's memory can be worse than no memory at all when the model or its interface is disposed to act on it: a memory that keeps…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in hig…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

プロンプトロバストネスはタスク依存: LLM 評価における客観的質問と信念スタイルの質問の比較

大規模な言語モデルの調査形式の評価では、多くの場合、促された応答がモデルの価値観や信念の尺度として扱われます。この仮定は、回答が政治的価値観、社会的態度、または信念の証拠として読み取られる場合に特に脆弱になります。答えが決まっている客観的な質問と、意見や価値観を求める主観的な質問とでは、プロンプトの堅牢性が異なるかどうかを尋ねます。 3 つの客観的データセット (MMLU、ARC、CulturalBench) と 3 つの主観的データセット (Political Compass Test、ValueBench、World Values Survey) に基づいて 4 つの命令調整モデル ファミリを評価します。各質問/ステートメントに対して、文言、枠組み、形式のバリエーションなど、複数のタイプのプロンプト変更を適用し、モデルがバリエーション全体で同じ回答を与えるかどうかを測定します。二項一般化推定方程式を使用すると、モデル、データセット、プロンプト カテゴリ、およびそれらの相互作用の重要な効果がわかります。データセット タイプの影響も大きく、データセット タイプとプロンプト カテゴリの間の相互作用は大きくなります。これらの結果は、プロンプトの堅牢性が質問の種類、プロンプトの変更、モデルに依存することを示しています。

原文 (English)

Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumption is particularly fragile when responses are read as evidence of political values, social attitudes, or beliefs. We ask whether prompt robustness differs between objective questions with fixed answers and subjective questions that ask for opinions or values. We evaluate four instruction-tuned model families on three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey). For each question/statement, we apply multiple types of prompt changes, such as variations in wording, framing, and format, and measure whether the model gives the same answer across variants. Using a binomial generalized estimating equation, we find significant effects of model, dataset, prompt category, and their interactions. The dataset type effect is also significant, and the interaction between dataset type and prompt category is large. These results show that prompt robustness depends on the question type, the prompt change, and the model.

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries

Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each s…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

Human Grounded Evaluation of Large Language Models for Optical Network Automation

Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substant…

2026-07-22 10:22 JSTITmedia AI+ハードウェア/半導体ビジネス/資金調達規制/政策

MicrosoftとMistralが戦略的提携を拡大 欧州でのAIインフラ拡張とモデル展開を加速

MicrosoftとMistralは戦略的提携を拡大すると発表した。Mistralの最新モデルをMicrosoftの各プラットフォームへ展開するほか、欧州でのGPUインフラ拡張に向けて大規模な投資を行う。クラウドから完全オフラインまで多様な環境に対応し、規制業界での高度なAI導…

2026-07-22 02:11 JSTTechCrunch AILLM/生成AIビジネス/資金調達

Google releases three new Gemini models — but no 3.5 Pro

Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber, but the continued absence of Gemini 3.5 Pro raises fresh questions about…

2026-07-21 16:00 JSTOpenAILLM/生成AIビジネス/資金調達

OpenAI and Hugging Face partner to address security incident during model evaluation

OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

RAIL Guard: LLM エージェントの責任ある AI における評価から修復までのギャップを埋める

大規模言語モデル エージェント用の既存のガードレール システムは、安全でないコンテンツをブロックするバイナリ分類子として動作するため、組織は失敗した出力を破棄して最初から再試行する必要があります。 8 つの測定可能な次元にわたって LLM 出力を評価し、評価、書き換え、再評価のループを通じて失敗した出力を繰り返し修復する閉ループ責任 AI パイプラインである RAIL Guard を紹介します。 4 つのフロンティア LLM、4,276 のコンテンツ出力、および 6,400 のエージェント ツール呼び出しシナリオに関する 3 つの実験にわたってパイプラインを評価しました。閉ループ修復は 96.9% の収束率を達成しますが、ブロックと再試行の場合は 49.1% ですが、収束率が最も高い方法では実用性が 22.3% 減少します。フィードバック駆動の自己修復は、重大なユーティリティの損失なしに、修正可能な次元で 86.6% の収束を達成します (p = 0.177)。ツール呼び出し前の評価により、安全でないエージェントの実行が 33% (p = 0.007) 削減され、タスクの完了には影響がありません。私たちは、修復に対応する修正可能な次元と、アルゴリズムではなくアーキテクチャ上の解決策を必要とする構造的次元 (93.0% の透明性、92.8% の説明責任、82.5% の失敗の包括性) との間の重要な違いを特定します。このシステムはオープンソース SDK として利用できます。

原文 (English)

RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents

Existing guardrail systems for large language model agents operate as binary classifiers that block unsafe content, leaving organizations to discard failing outputs and retry from scratch. We introduce RAIL Guard, a closed-loop responsible AI pipeline that evaluates LLM outputs across eight measurable dimensions and iteratively remediates failing outputs through an evaluate-rewrite-reevaluate loop. We evaluate the pipeline across three experiments on four frontier LLMs and 4,276 content outputs plus 6,400 agent tool-call scenarios. Closed-loop remediation achieves 96.9% convergence versus 49.1% for block-and-retry, though the highest-convergence method reduces utility by 22.3%; feedback-driven self-repair achieves 86.6% convergence on fixable dimensions with no significant utility loss (p = 0.177). Pre-tool-call evaluation reduces unsafe agent executions by 33% (p = 0.007) with zero impact on task completion. We identify a key distinction between fixable dimensions that respond to remediation and structural dimensions (Transparency at 93.0%, Accountability at 92.8%, and Inclusivity at 82.5% failure) that require architectural rather than algorithmic solutions. The system is available as open-source SDKs.

2026-07-21 13:00 JSTarXiv cs.AIビジネス/資金調達

DL オントロジーの認識的機密性ポリシーに基づく扱いやすいクエリ応答 (拡張バージョン)

私たちは、記述ロジック (DL) オントロジーのコンテキストで、また認識依存関係 (ED) を通じて表現される機密性ポリシーについて、機密性を保持するデータ アクセスへの宣言的アプローチである制御クエリ評価 (CQE) を研究します。まず、CQE の既知のセマンティクス (GA および IGA 含意) の下でクエリ (具体的には論理積クエリのブール和集合) に応答する問題に取り組みます。私たちの結果は、TBox が $\text{DL-Lite}_{\mathcal{R}}$ で表現される場合、CQE は一般に計算的に扱いにくいことを示しています。さらに、ED が存在する場合、IGA セマンティクスは、識別不可能性として知られる重要な機密保持特性を満たさないことが最近証明されました。計算が容易で機密性が保たれる CQE 形式を定義することを目的として、最小限のポリシー違反 (MPV) の概念に基づいた CQE の新しいセマンティクスを導入します。新しいセマンティクスが以前のセマンティクスの健全な近似を提供しながら、区別不可能性の特性を満たしていることを示します。また、$\text{DL-Lite}_{\mathcal{R}}$ オントロジーの場合、MPV セマンティクスに基づくクエリ含意がデータ複雑さの多項式時間で決定できることも証明します。最後に、OWL 2 QLの既存のベンチマークを使用して、この新しいアプローチの実現可能性を評価するために使用したフレームワークのソフトウェア実装を紹介します。

原文 (English)

Tractable Query Answering under Epistemic Confidentiality Policies in DL Ontologies (extended version)

We study Controlled Query Evaluation (CQE), a declarative approach to confidentiality-preserving data access, in the context of Description Logic (DL) ontologies, and for confidentiality policies expressed through Epistemic Dependencies (EDs). We first address the problem of answering queries (specifically, Boolean unions of conjunctive queries) under known semantics for CQE (GA- and IGA-entailment). Our results show that if the TBox is expressed in $\text{DL-Lite}_{\mathcal{R}}$, CQE is computationally intractable in general. Moreover, in the presence of EDs, the IGA semantics has recently been proven not to satisfy an important confidentiality preservation property known as indistinguishability. With the goal of defining computationally easier and confidentiality-preserving forms of CQE, we introduce a new semantics for CQE, based on the notion of minimal policy violation (MPV). We show that the new semantics provides a sound approximation of the previous ones, while satisfying the indistinguishability property. We also prove that, in the case of $\text{DL-Lite}_{\mathcal{R}}$ ontologies, query entailment under the MPV semantics can be decided in polynomial time in data complexity. Finally, we present a software implementation of our framework that we used to evaluate the feasibility of this new approach using an existing benchmark for OWL 2 QL.

2026-07-21 13:00 JSTarXiv cs.AIビジネス/資金調達

検索拡張リーダーによるサポートの使用方法を形成する証拠インターフェイス

マルチホップ RAG 評価では、上位 k 位の応答スコアによって 2 つの異なる失敗が隠蔽される可能性があります。取得ウィンドウがサポート チェーンの一部を削除する可能性があるか、適応されたリーダーが適切に使用しない形式のサポートが含まれている可能性があります。私たちは、この取得された証拠の読者向け形式を証拠インターフェイスと呼びます。 3 つのサポート アノテーション付きマルチホップ QA ベンチマークを使用して、生のコンテキスト、取得ウィンドウ、およびゴールド サポートの診断レンダリングでトレーニングされた、適合した適応リーダーを比較します。これらの比較により、サポートの利用可能性の障害と残りのリーダー インターフェイスの影響が区別されます。 Top-k ウィンドウは、完全なアノテーション付きサポート チェーンが存続するかどうかを確認した後にのみ解釈可能になります。存続する場合、ランクの短いウィンドウは生のコンテキストと一致するか、生のコンテキストよりも改善されます。そうでない場合は、サポートの欠如が損失の大部分を説明します。ゴールド サポート第一により、一致するリーダーが向上します。 2Wiki と MuSiQue では、サポートの監視下にあるランカーが、ゴールドのヘッドルームを維持しながら、カバレッジを高め、より低い即時コストで生のコンテキストの品質を回復します。サポート除去チェックでは、利益が事前回答だけではなく、暴露された証拠に依存していることがさらに示されています。したがって、サポート注釈付きの評価では、上位 k 位までの回答スコアを完全なサポート範囲とともに報告する必要があります。

原文 (English)

Evidence Interfaces Shape How Retrieval-Augmented Readers Use Support

In multi-hop RAG evaluation, a top-k answer score can hide two different failures: the retrieval window may drop part of the support chain, or it may contain support in a form the adapted reader does not use well. We call this reader-facing form of retrieved evidence an evidence interface. Using three support-annotated multi-hop QA benchmarks, we compare matched adapted readers trained with raw context, retrieval windows, and gold-support diagnostic renderings. These comparisons distinguish support-availability failures from remaining reader-interface effects. Top-k windows become interpretable only after checking whether the complete annotated support chain survives: when it does, short ranked windows can match or improve over raw context; when it does not, missing support explains much of the loss. Gold support-first improves matched readers; on 2Wiki and MuSiQue, a support-supervised ranker raises coverage and recovers raw-context quality at lower prompt cost, while retaining gold headroom. Support-removal checks further show that the gains rely on exposed evidence, not only answer priors. On support-annotated evaluations, top-k answer scores should therefore be reported together with complete-support coverage.

2026-07-21 13:00 JSTarXiv cs.AIビジネス/資金調達

擬人化された対話に向けて: 人間のようなチャットの生成、評価、および好みの調整のための閉ループ フレームワーク

人間のようなプライベート チャットには、流暢な応答生成以上のものが必要です。システムは、ペルソナ、関係、記憶、限定された知識、媒体固有のタイミング、一貫したマルチターン アークを保持する必要があります。我々は、擬人化対話をシステム アーキテクチャ、実行可能評価、および診断調整の共同問題として定式化する閉ループ フレームワークである AnthroDial を紹介します。これは、(1) 役割条件付きのスケジュールされた対話ランタイムと、ペルソナおよびシナリオ カード、長期記憶、仮想時間、および単一草案メッセージの決定を組み合わせます。 (2) L0 有効性ゲート、5 つのターンごとのディメンション、および 5 つのダイアログ レベルのディメンションを備えた実行可能なベンチマーク。 (3) SFT 用に 16,436 のスケジュールされた決定例をフィルタリングし、認知診断、ZPD を意識した報酬を備えた GRPO を適用するトレーニング後のパイプライン。この報酬は、各行動次元のカルマン フィルター処理された能力推定値を維持し、より大きな能力不足のある次元を重み付けし、ロールアウト スコアをタスク レベルの ZPD マッチとして使用して、学習可能な弱いスキルに焦点を合わせて最適化します。モデルごとに 55 のペルソナ、50 のシナリオ、50 のペルソナとシナリオのバインディング、および 100 の役割条件付きケースを含むベンチマークで、フロンティア ベースライン、オープン モデル、思考/非思考のバリアント、および SFT/RL アブレーションにわたる 16 のシステムを評価します。最も強い非トレーニングベースラインは 32.00% の厳密な ACC に達しますが、Qwen3.6-27B-SFT+RL は 39.00% の厳密な ACC と 98.5 の全体スコアに達します。 9B ノーシンク設定では、SFT と RL は厳密な ACC を 0.00% から 13.00% および 18.37% に改善します。これらの結果は、生成、評価、報酬形成が同じ行動次元を共有する場合、擬人化対話が有益であることを示しています。

原文 (English)

Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc. We present AnthroDial, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment. It combines (1) a role-conditioned scheduled dialogue runtime with persona and scenario cards, long-term memory, virtual time, and single-draft message decisions; (2) an executable benchmark with an L0 validity gate, five per-turn dimensions, and five dialogue-level dimensions; and (3) a post-training pipeline that filters 16,436 scheduled-decision examples for SFT and applies GRPO with a cognitive-diagnostic, ZPD-aware reward. The reward maintains Kalman-filtered capability estimates for each behavioral dimension, upweights dimensions with larger capability deficits, and uses rollout scores as task-level ZPD matches to focus optimization on learnable weak skills. On a benchmark with 55 personas, 50 scenarios, 50 persona-scenario bindings, and 100 role-conditioned cases per model, we evaluate 16 systems spanning frontier baselines, open models, thinking/no-think variants, and SFT/RL ablations. The strongest non-trained baseline reaches 32.00% strict ACC, while Qwen3.6-27B-SFT+RL reaches 39.00% strict ACC and a 98.5 overall score. In the 9B no-think setting, SFT and RL improve strict ACC from 0.00% to 13.00% and 18.37%. These results show that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

コードエージェントの LoRA 微調整のための軌跡データキュレーションの体系的な評価

エキスパート エージェントの軌道に関するオープンウェイト LLM の教師あり微調整 (SFT) は、独自のモデルに依存せずに有能なコード エージェントを構築するための顕著なアプローチとして浮上しています。中心的であるにもかかわらず十分に解明されていない問題は、軌道の質と量がどのように連携してモデルのパフォーマンスを形成するかということです。我々は、SWE 軌跡データセット (67,074 の軌跡、うち 32,161 が解決済み) 上の Qwen2.5-Coder-7B-Instruct の LoRA 微調整のための軌跡データ フィルタリングの体系的な実証研究を紹介します。当社は、効率とスタイルという 2 軸の品質スコアリング フレームワークを提案し、戦略、規模、アブレーション分析にわたる 16 の対照実験を通じてそれを評価します。 7B スケール モデルはほぼゼロの SWE ベンチ解決率を達成するため、ホールドアウト軌道上のクロス エントロピー (CE) 損失を主要な指標として採用し、ファースト アクション生成によって検証されます。CE 損失と ROUGE-L は完全に順位相関しており (Spearman $\rho$ = -1.00)、限られたサンプルの証拠はこの代理を裏付けていますが、決定的に確立しているわけではありません。私たちの結果は、スケール依存の品質と量のトレードオフを明らかにしました。小規模では、データセットを 2 倍にする (500 から 1,000) と、最大 12.7% の CE 損失の削減が得られますが、TopQ-Random ギャップは 0.10 のままです。 2,000 の軌道では、この同じギャップは 3.6% に広がります (p = 0.016)。アブレーションではさらに、エラー再試行率が支配的なサブディメンションであることが特定され、完全なコンポジット ($\Delta$ < 0.2%) と同等のパフォーマンスを示します。まとめると、これらの発見は、コードエージェント SFT の実行可能だが規模に依存する手段としての軌道レベルの品質スコアリングを確立し、エンドツーエンドの解決率が統計的に実現不可能なレジームに対して代理検証済みの評価プロトコルを提供します。

原文 (English)

A Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code Agents

Supervised fine-tuning (SFT) of open-weight LLMs on expert agent trajectories has emerged as a prominent approach to building capable code agents without reliance on proprietary models. A central yet underexplored question is how trajectory quality and quantity jointly shape model performance. We present a systematic empirical study of trajectory data filtering for LoRA fine-tuning of Qwen2.5-Coder-7B-Instruct on the SWE-trajectory dataset (67,074 trajectories, of which 32,161 are resolved). We propose a two-axis quality scoring framework -- Efficiency and Style -- and evaluate it through 16 controlled experiments spanning strategy, scale, and ablation analyses. Since 7B-scale models attain near-zero SWE-bench resolve rates, we adopt cross-entropy (CE) loss on held-out trajectories as the primary metric, validated via first-action generation: CE loss and ROUGE-L are perfectly rank-correlated (Spearman $\rho$ = -1.00), with limited-sample evidence supporting but not conclusively establishing this proxy. Our results reveal a scale-dependent quality-quantity trade-off: at small scales, doubling the dataset (500 to 1,000) yields ~12.7% CE-loss reduction whereas the TopQ-Random gap stays 0.10); at 2,000 trajectories this same gap widens to 3.6% (p = 0.016). Ablation further identifies error-retry rate as the dominant sub-dimension, performing comparably to the full composite ($\Delta$ < 0.2%). Together, these findings establish trajectory-level quality scoring as a viable but scale-sensitive lever for code-agent SFT and offer a proxy-validated evaluation protocol for the regime where end-to-end resolve rate is statistically infeasible.

2026-07-21 13:00 JSTarXiv cs.AIビジネス/資金調達

データファーストオントロジーに基づく明示的な世界モデル: DaoQL マルチモーダルストレージ検証と反事実推論評価

大規模な言語モデルは世界モデルをニューラル ウェイトで暗黙的にエンコードするため、医療や金融などの高精度の領域では、幻覚、知識の凍結、説明可能性の低さ、修正可能性の低さという 4 つの構造的リスクが露呈します。この論文では、データファースト オントロジーを提案します。LLM は推論および言語エンジンとして扱われ、決定論的な知識は明示的なマルチモーダル データベースである DaoQL に移動されます。我々は、明示的な世界モデルを形式化し、ルールの独立性、決定論的な評価、および固定された競合解決の下で、明示的なモデルが構成可能な反事実分解可能性のための十分な条件を提供することを示します。暗黙的モデルにはアトミックな読み取り/デルタ セマンティクスが欠けているため、同等のアーキテクチャ上の保証はありません。実装されたシステムは、DaoQL の検証済みストレージ レイヤーと明示的な Eval パスに焦点を当てており、グラフ、列、ベクトル、およびフルテキスト エンジンを 1 つのプロセス内に統合しています。 KVCache グラフ ノード、エキスパート ホット アップデート、および DaoQL-Agent ランタイムは今後の作業として残ります。組み込みの同一マシン設定では、DaoQL はグラフ BFS を 1.20 ミリ秒、HNSW を 83.1 マイクロ秒、Fluent ハイブリッド クエリを 105.8 マイクロ秒でレポートします。これらの結果はエンジニアリングの可能性を示していますが、クライアント/サーバー システムとの展開形状の違いを考慮して解釈する必要があります。 LDBC SNB SF1 および ANN ベンチマークの探索的測定では、インタラクティブ クラスのクエリのほとんどがサブミリ秒からミリ秒の範囲で 34/34 のクエリ カバレッジを示していますが、ロングテール BI/IC クエリにより全体では 1.8 QPS にすぎません。 ANN ベンチマークは、ブリッジエッジ保護の修正後、1,000 レベルの QPS で Recall@10 >= 99% に達しました。 5 ドメインの反事実実験 (n = 1250) では、DaoQL+GPT-4o は 94% の構成可能な反事実分解可能性を達成し、GPT-4o 単独より 49 パーセントポイント上回りました。この論文では、証明可能な構造、予備的な経験的証拠、およびアーキテクチャ上のロードマップの主張を明確に分離しています。

原文 (English)

An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation and Counterfactual Reasoning Evaluation

Large language models encode world models implicitly in neural weights, which exposes four structural risks in high-precision domains such as medicine and finance: hallucination, frozen knowledge, poor explainability, and poor modifiability. This paper proposes data-first ontology: LLMs are treated as reasoning and language engines, while deterministic knowledge is moved into an explicit multimodal database, DaoQL. We formalize an explicit world model and show that, under rule independence, deterministic evaluation, and fixed conflict resolution, explicit models provide a sufficient condition for composable counterfactual decomposability; implicit models lack atomic read/delta semantics and therefore provide no comparable architectural guarantee. The implemented system focuses on DaoQL's verified storage layer and explicit Eval path, integrating graph, column, vector, and full-text engines within one process. KVCache graph nodes, expert hot updates, and the DaoQL-Agent runtime remain future work. On an embedded same-machine setup, DaoQL reports graph BFS at 1.20 ms, HNSW at 83.1 us, and a Fluent hybrid query at 105.8 us; these results indicate engineering potential but must be interpreted with deployment-shape differences from client-server systems. Exploratory measurements on LDBC SNB SF1 and ANN-Benchmarks further show 34/34 query coverage with interactive-class queries mostly in the sub-millisecond to millisecond range, but only 1.8 QPS overall due to long-tail BI/IC queries; ANN-Benchmarks reaches Recall@10 >= 99% at thousand-level QPS after a bridge-edge protection fix. In a five-domain counterfactual experiment (n = 1250), DaoQL+GPT-4o achieves 94% composable counterfactual decomposability, 49 percentage points above GPT-4o alone. The paper explicitly separates provable structure, preliminary empirical evidence, and architectural roadmap claims.

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear…

2026-07-21 13:00 JSTarXiv cs.AIビジネス/資金調達

Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis

Machine learning-based Type 2 diabetes risk prediction models obtain good internal validation results but lose effectiveness in real-world…

2026-07-21 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs

Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predicti…

2026-07-21 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection

Publicly shareable construction-video benchmarks remain scarce, especially for safety-critical hazards that are rare, dangerous to stage, a…

2026-07-21 13:00 JSTarXiv cs.AIビジネス/資金調達

K-IPO: Kendall-constrained Importance Preserving Oversampling for Imbalanced Tabular Data

Oversampling is widely used to address class imbalance in tabular classification, but existing methods can distort the feature importance r…

2026-07-21 13:00 JSTarXiv cs.AIビジネス/資金調達

Building a Neural Network from Scratch: Implementation, Evaluation, and Optimization

The widespread adoption of high-level deep learning libraries, while accelerating model development, has increasingly abstracted away the i…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries

Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each s…

2026-07-21 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Automated Cardiac Adipose Tissue Segmentation in Computed Tomography: A Literature Review

This review provides an overview of recent advancements in automated segmentation methods on Computed Tomography (CT) for two types of card…

2026-07-21 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

ALLUDE: A Unified Evaluation System for Configurable Attacks in Differentiable Environments

Adversarial attacks against vision models like object detectors are often evaluated under limited conditions, leaving their performance und…

2026-07-21 13:00 JSTarXiv cs.AIビジネス/資金調達

Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

Human Grounded Evaluation of Large Language Models for Optical Network Automation

Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substant…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs

Large language models (LLMs) exhibit strong general reasoning, yet the community lacks controllable, scalable, and verifiable tools to anal…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation

LLMs make it increasingly easy to generate deceptive content at scale, creating a need for scalable misinformation risk evaluation based on…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AgentCompass: エージェント機能の統合評価インフラストラクチャ

大規模言語モデル (LLM) が自律エージェントに進化するにつれて、統合された評価インフラストラクチャの必要性が重要になります。ただし、現在の評価パイプラインは高度に断片化され、密接に結合されたままであるため、再現性が妨げられ、冗長なエンジニアリングが発生します。これに対処するために、LLM ベースのエージェントを評価するためのオープンソースで軽量かつ拡張可能なインフラストラクチャである AgentCompass を導入します。 AgentCompass は、ベンチマーク、ハーネス、環境という 3 つの独立したコンポーネントを中心に評価プロセスを編成するため、複雑な実行ロジックを再実装することなく柔軟な構成が可能になります。さらに、フォールトトレラントな非同期ランタイムと、報酬ハッキングなどの微妙な障害モードを透過的に診断するための包括的な軌跡分析ツールを備えています。 AgentCompass は、5 つの機能次元にわたる 20 以上のベンチマークをネイティブにサポートし、エージェント研究を進めるためのスケーラブルで再現可能なインフラストラクチャをコミュニティに提供します。

原文 (English)

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals

Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-varianc…

2026-07-21 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification

Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation

Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse…

2026-07-21 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

データセットの価値はいくらですか?スケーリング則、Vendi スコア、および行列スペクトル関数

ニューラル スケーリングの法則はデータセットのサイズを通じてデータを評価しますが、Vendi スコアは量子エントロピーを使用してデータセットの値を測定します。一般的なニューラル スケーリング則の目標と Vendi スコアの両方がサブモジュールであることを示します。さらに、Vendi スコアが、行列スペクトル関数と呼ばれるより広範なクラスのサブモジュラー目標の特殊なケースであることを示します。これには、決定的 (DPP) 目標や他の多くの目標も含まれます。また、弱行列単調関数を導入し、それがどのように弱部分モジュール行列スペクトル関数につながるかを示し、データ評価のための幅広い実用的な目的をもたらします。私たちは、貪欲な最適化中に繰り返される固有分解を回避する永年方程式ベースの更新を開発し、$m$ 次元の埋め込みに対する限界ゲイン評価を Oracle クエリと比較して $O(m)$ 係数だけ削減します。これにより、経験的に平均約 35,000 倍の高速化が得られ、ImageNet-1K スケールのデータセットで Vendi スコアの直接最適化が可能になります。このようにして可能になったので、Vendi スコア、DPP、施設の場所、および 3 つの新しいマトリックス スペクトル バリアントを含む、固定サイズ、クラスバランス、および固定トレーニング予算体制の下で、いくつかの目標がホールドアウト テスト パフォーマンスのトレーニング サブセットの値をどの程度正確に予測するかを比較します。複数のデータセットにわたって、施設の位置が最も優れたパフォーマンスを発揮します。また、直接最適化では、Vendi スコアは中程度のスコア範囲では予測的ですが、目標をより高い値に押し上げると、下流のパフォーマンスの代用として機能しなくなる可能性があることも明らかになりました。また、均一でランダムな固定サイズのサブセットは、制約がなく、クラスバランスが取れていても、評価スコアと保持されたパフォーマンスの両方で著しく集中していることもわかります。最後に、サイズ、クラスのバランス、トレーニング予算だけがデータの価値を決定するわけではないことを示します。これらの要因を制御した場合でも、パフォーマンスは良い状態から悪い状態まで滑らかに変化します。

原文 (English)

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions

Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common neural-scaling-law objectives and the Vendi Score are submodular. We further show that the Vendi Score is a special case of a broader class of submodular objectives that we call matrix spectral functions. This also includes determinantal (DPP) objectives, as well as many others. We also introduce weakly matrix monotone functions and show how they lead to weakly submodular matrix spectral functions, yielding a broad family of practical objectives for data appraisal. We develop secular-equation-based updates that avoid repeated eigendecompositions during greedy optimization, reducing marginal-gain evaluation for $m$-dimensional embeddings by an $O(m)$ factor relative to oracle queries. This yields an average empirical speedup of about 35,000x, making direct optimization of the Vendi Score feasible on ImageNet-1K-scale datasets. Thus enabled, we compare how well several objectives predict the value of training subsets for held-out test performance under fixed-size, class-balanced, and fixed training-budget regimes, including the Vendi Score, DPPs, facility location, and three new matrix spectral variants. Across multiple datasets, facility location performs the best. Direct optimization also reveals that, while the Vendi Score is predictive over moderate score ranges, pushing the objective to higher values can make it a poor downstream performance proxy. We also find that uniformly at random fixed-size subsets, both unconstrained and class-balanced, are remarkably concentrated in both appraisal scores and held-out performance. Finally, we show that size, class balance, and training budget do not alone determine data value: even when controlling for these factors, performance ranges smoothly from good to bad.

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

A language model's memory can be worse than no memory at all when the model or its interface is disposed to act on it: a memory that keeps…

2026-07-21 13:00 JSTarXiv cs.AIビジネス/資金調達

The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits

The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a…

2026-07-21 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

NLPCC 2026 の概要 共有タスク 1: 難易度を考慮した多言語およびマルチモーダルな医療指導ビデオの理解度評価

NLPCC 2023 ~ 2025 年の CMIVQA、MMI-VQA、および M4IVQA の課題に続き、NLPCC 2026 では難易度を考慮した医療指導ビデオ質問応答 (DA-MIVQA) 共有タスクを導入します。DA-MIVQA は、必要な証拠の種類と複雑さに応じて質問を明示的に区別することで、以前の多言語および多モードの医療ビデオ ベンチマークを拡張します。答える。具体的には、単純な質問は字幕ベースのテキストの手がかりから答えられることがよくありますが、複雑な質問には視覚的な根拠、手順の理解、およびクロスモーダルな証拠の統合が必要です。このチャレンジには、単一ビデオでの難易度を考慮した時間的回答グラウンディング (DA-TAGSV)、ビデオ コーパスでの難易度を考慮した時間的回答の取得 (DA-VCR)、およびビデオ コーパスでの難易度を考慮した時間的回答のグラウンディング (DA-TAGVC) の 3 つのトラックが含まれています。データセットは公的医療指導チャンネルから収集され、応急処置、緊急対応、リハビリテーション、看護、一般医学教育などの多様なシナリオをカバーしており、難易度の注釈を付けて手動で検証されています。本稿では、DA-MIVQAの課題動機、データセット構築、評価プロトコル、参加概要、競技結果、代表的なシステムについて紹介する。 DA-MIVQA は、さまざまなテキスト、視覚、時間的、および手順の推論要件の下で、医療指導ビデオ質問応答システムを評価するための実用的なベンチマークを提供します。

原文 (English)

Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation

Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026. DA-MIVQA extends previous multilingual and multimodal medical video benchmarks by explicitly distinguishing questions according to the type and complexity of evidence required for answering. Specifically, simple questions can often be answered from subtitle-based textual cues, whereas complex questions require visual grounding, procedural understanding, and cross-modal evidence integration. The challenge contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV), Difficulty-Aware Video Corpus Retrieval (DA-VCR), and Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The dataset is collected from public medical instructional channels, covers diverse scenarios such as first aid, emergency response, rehabilitation, nursing, and general medical education, and is manually verified with difficulty annotations. This paper presents the task motivation, dataset construction, evaluation protocol, participation overview, competition results, and representative systems of DA-MIVQA. DA-MIVQA provides a practical benchmark for evaluating medical instructional video question answering systems under varying textual, visual, temporal, and procedural reasoning requirements.

2026-07-20 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

AgentFAIR: 地理空間データセットの公平性評価のためのマルチエージェント協調フレームワーク

地理空間データセットは、都市計画から気候モデリングまでのアプリケーションをサポートしていますが、FAIR 準拠の一貫した評価は困難です。既存の評価者は異なるルーブリックと証拠ソースを使用しており、JavaScript でレンダリングされたページやリポジトリ固有の識別子では失敗する可能性があります。 10 のリポジトリからの 50 のデータセットの場合、利用可能なツール全体の正規化スコアの標準偏差は平均 15.0 パーセント ポイントで、1 つのデータセットでは 30.3 に達します。これらの出力は同等の測定値ではないため、精度の比較ではなく、不一致や故障モードを特徴付けるために使用します。構造化メタデータ抽出と 13 のサブ原則固有の LLM エバリュエーターを組み合わせたマルチエージェント フレームワークである AgentFAIR を紹介します。それぞれが 0 ~ 3 の成熟度スコア、引用された証拠、推奨事項を生成します。批評家は証拠と一貫性をチェックし、的を絞った再評価を要求できます。検索可能性、アクセシビリティ、相互運用性、および再利用性の平均スコアは、79.7%、70.4%、45.3%、および 72.0% です。 4 つのベースライン ツールとのランク相関は 0.31 ~ 0.61 の範囲です。 FAIR-enough 比較は統計的に有意ではありません。 10 個のデータセットを繰り返し実行したサブセットでは、部分原理一致率は平均 89% (標準偏差: 3 パーセント ポイント) でしたが、批判なしでは 71% でした。 15のデータセットの専門家による予備調査では、フライスのカッパが0.71であり、専門家のコンセンサスと82%一致していることがわかりました。 API コストはデータセットあたり約 0.054 米ドルです。これらの結果は、監査可能性と実現可能性を裏付けるものですが、ベンチマークの制限、不完全なアブレーション、および単一モデルファミリーの検証により、精度と一般化に関する主張が制約されます。

原文 (English)

AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets

Geospatial datasets support applications from urban planning to climate modeling, yet consistent assessment of FAIR compliance is difficult. Existing evaluators use different rubrics and evidence sources and may fail on JavaScript-rendered pages or repository-specific identifiers. For 50 datasets from 10 repositories, the standard deviation of normalized scores across available tools averages 15.0 percentage points and reaches 30.3 for one dataset. Because these outputs are not equivalent measurements, we use them to characterize disagreement and failure modes, not comparative accuracy. We present AgentFAIR, a multi-agent framework combining structured metadata extraction with 13 sub-principle-specific LLM evaluators. Each produces a 0-3 maturity score, cited evidence, and recommendations; a critic checks evidence and consistency and can request targeted re-evaluation. Mean Findability, Accessibility, Interoperability, and Reusability scores are 79.7%, 70.4%, 45.3%, and 72.0%. Rank correlations with four baseline tools range from 0.31 to 0.61; the FAIR-enough comparison is not statistically significant. On a 10-dataset repeated-run subset, sub-principle agreement averages 89% (standard deviation: 3 percentage points), versus 71% without the critic. A preliminary 15-dataset expert study yields Fleiss' kappa of 0.71 and 82% alignment with expert consensus. API cost is approximately USD 0.054 per dataset. These results support auditability and feasibility, while the limited benchmark, incomplete ablations, and single-model-family validation constrain claims about accuracy and generalization.

2026-07-20 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

CRAFT: ルーブリックをクラスタリングして弱い LLM 機能を診断し、対象を絞った微調整データを生成する

評価では、モデルの現在のパフォーマンスを測定するだけでは不十分です。彼らは、次のモデルの反復で何を修正すべきかを私たちに伝え、ターゲットを絞ったトレーニング後のデータを生成する方法を提供する必要があります。ほとんどの評価パイプラインは、弱い例、トピック、またはカテゴリを特定しますが、潜在的な機能の失敗を暗黙的に残します。モデルが失敗する理由ではなく、どこで失敗するかを示します。ルーブリックベースの評価データセットをモデル固有の弱い機能の診断に変換する手法である CRAFT を紹介します。 CRAFT は、各評価基準を能力プローブとして扱います。つまり、すべてのプロンプト ルーブリック ペアから能力の説明を抽出し、これらの説明を階層的な能力ツリーにクラスタリングし、すべてのノードでターゲット モデルにスコアを付け、各障害が最も明確な粒度でツリー レベル全体にわたってパフォーマンスの低いノードを動的に選択します。選択された弱い機能は、ターゲットを絞った教師付き微調整データの生成を指示します。データ生成、微調整、および評価のセットアップを固定したまま、4 つのオープンソース モデル、2 つの専門分野 (財務および法務)、および診断データから切り離された 13 のベンチマークで、プロンプト レベルの EvalTree クラスタリングおよびターゲットを絞らないランダム生成と CRAFT を比較します。 CRAFT は、温度デコードを繰り返すと、4 つのモデルすべてで最も強力な金融ドメイン平均を達成します。法的領域では、4 つのモデルのうち 3 つで最も強く、4 つ目のモデルでは最良のベースラインのデコード分散帯域内に留まります。したがって、プロンプトやカテゴリではなく、ルーブリック基準のレベルで弱点を診断すると、モデルで何ができないのかをより明確に把握できるようになり、その診断に基づいて微調整した後は測定可能なほど優れたモデルが得られます。

原文 (English)

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or categories, but they leave the underlying capability failure implicit: they say where a model fails, not why. We introduce CRAFT, a method that converts any rubric based evaluation dataset into a model specific diagnosis of weak capabilities. CRAFT treats each grading criterion as a capability probe: it extracts a capability description from every prompt rubric pair, clusters these descriptions into a hierarchical capability tree, scores the target model at every node, and selects low performing nodes dynamically across tree levels, at the granularity where each failure is clearest. The selected weak capabilities then direct the generation of targeted supervised finetuning data. Holding the data generation, finetuning, and evaluation setup fixed, we compare CRAFT against prompt level EvalTree clustering and untargeted random generation on four open source models, two professional domains (finance and legal), and 13 held out benchmarks disjoint from the diagnostic data. CRAFT achieves the strongest finance domain average for all four models under repeated temperature decoding; on legal domain, it is strongest for three of four models and remains within the decoding variance bands of the best baseline on the fourth. Diagnosing weaknesses at the level of rubric criteria, rather than prompts or categories, thus yields both a sharper picture of what a model cannot do and measurably better models after finetuning on that diagnosis.

2026-07-20 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

暗黙的な文化的整合性報酬モデリングによるテキストから画像への評価のバイアスを軽減する

Text-to-Image (T2I) システムが急速に進歩するにつれ、公正で信頼できる生成 AI にとって、合成コンテンツの文化的信頼性を評価することがますます重要になってきています。既存の T2I 評価指標とマルチモーダルなジャッジは、暗黙的な文化規範を過小評価する視覚的意味論的表現に依存していることが多く、偏った好みの判断やきめ細かな文化的手がかりの省略につながります。さらに、ビジュアル質問応答 (VQA) ベースの評価器は通常、自己回帰テキスト生成に依存するため、リアルタイム報酬モデリングのスケーラビリティが制限されます。これらの制限に対処するために、軽量の 42 億パラメータのマルチモーダル大規模言語モデル (MLLM) に基づいて構築された暗黙的文化的調整報酬モデルを導入します。私たちのフレームワークは、暗黙的文化プローブをスキップ接続クロスアテンション (SkipCA) メカニズムと統合し、後期段階のセマンティック機能が初期段階の視覚表現に直接対応し、文化的に顕著な詳細をより適切に保存できるようにします。 CultureFrames ベンチマークからの 3,323 個の難しくて厳選された画像ペアの評価では、私たちのアプローチがペアごとの精度 80.54% を達成し、ピアソン相関係数とケンダル相関係数がそれぞれ 0.546 と 0.377 で、代表的な視覚言語メトリックや MLLM ベースの評価者を上回るパフォーマンスを示しています。さらに、自己回帰テキスト生成をバイパスすることにより、モデルはローカル推論セットアップの下で各評価を 0.21 秒で処理し、標準の VQA ベースの評価器と比べて 10 倍の高速化を達成します。これらの結果は、提案された報酬モデルが、ヒューマン フィードバックからの強化学習や直接嗜好最適化などの嗜好最適化パイプラインに対して、効率的で文化を意識したスカラー信号を提供できることを示唆しています。

原文 (English)

Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling

As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI. Existing T2I evaluation metrics and multimodal judges often rely on visual-semantic representations that underrepresent implicit cultural norms, leading to biased preference judgments and the omission of fine-grained cultural cues. In addition, visual question answering (VQA)-based evaluators typically depend on autoregressive text generation, which limits their scalability for real-time reward modeling. To address these limitations, we introduce an Implicit Cultural Alignment Reward Model built upon a lightweight 4.2-billion-parameter Multimodal Large Language Model (MLLM). Our framework integrates an Implicit Cultural Probe with a Skip-connection Cross-Attention (SkipCA) mechanism, enabling late-stage semantic features to directly attend to early-stage visual representations and better preserve culturally salient details. Evaluations on 3,323 challenging and carefully curated image pairs from the CulturalFrames benchmark show that our approach achieves 80.54% pairwise accuracy, with Pearson and Kendall correlation coefficients of 0.546 and 0.377, respectively, outperforming representative vision-language metrics and MLLM-based evaluators. Moreover, by bypassing autoregressive text generation, our model processes each evaluation in 0.21 seconds under our local inference setup, achieving a $10\times$ speedup over standard VQA-based evaluators. These results suggest that the proposed reward model can provide an efficient and culturally aware scalar signal for preference optimization pipelines such as Reinforcement Learning from Human Feedback and Direct Preference Optimization.

2026-07-20 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy

The black-box nature of Large Language Models necessitates novel evaluation frameworks that transcend surface-level performance metrics. Th…

2026-07-20 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Workflow-GYM: 現実世界の専門分野におけるコンピュータ使用エージェントタスクの長期的な評価に向けて

近年、ますます複雑になる現実世界のタスクの処理に向けて、AI エージェントが急速に進化しています。しかし、既存のベンチマークでは、エージェントがグラフィカル ユーザー インターフェイスを操作して、さまざまなドメインにわたる長期にわたる価値の高い専門的なワークフローを完了できるかどうかを評価することはほとんどありません。現在の GUI ベンチマークは依然として、主に汎用ソフトウェア、比較的単純なアプリケーション、および短期間のタスクに焦点を当てており、最新のエージェントがユーザーの指示に従ってドメイン固有のプロフェッショナル ソフトウェアを自律的に操作し、経済的に価値のある作業をエンドツーエンドで実行できるかどうかはほとんど不明です。このギャップを埋めるために、専門分野と特殊なソフトウェア環境を中心とした長期的な GUI タスクのベンチマークである Workflow-GYM を導入します。最先端のモデルで広範な実験を行った結果、最も強力なモデルでも成功率は 30% をわずかに超える程度であることがわかり、プロの長期にわたる GUI ワークフローが現在の GUI エージェントにとって依然として非常に困難であることが浮き彫りになりました。さらなる分析により、現在のエージェントは長期的なワークフローの一貫性を維持するのに苦労しており、ワークフロー段階の省略、エラーの伝播、目標のずれ、プロフェッショナルなソフトウェア環境の理解不足が頻繁に見られることが明らかになりました。私たちの調査結果は、現在のエージェント システムの限界についての重要な洞察を提供し、次世代の GUI エージェント研究の重要な方向性を示唆しています。

原文 (English)

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.

2026-07-20 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

エージェント ステップ値: 状態接地 LLM エバリュエーターによる状態遷移測定

ほとんどのエージェント評価では、複数ステップのトレースが最終的な回答、成功フラグ、または軌跡レベルのスコアにまとめられます。これらの集計では、開発者が最も必要とする診断の質問、つまりどのアクションが状態を有益な方向に変更したのかがわかりにくくなります。我々は、状態遷移測定フレームワークであるエージェント ステップ値 (ASV) を導入します。これは、観察された各アクションを、固定された候補結果に対する状態に基づいた評価者の分布に誘発する変化によってスコア付けします。 ASV は、編集された前後の状態予測をレンダリングし、ステートレス LLM エバリュエーターを使用して候補ログ スコアを割り当て、ゴールドフリーの信念診断とオフライン オラクル検証メトリクスの両方をレポートします。ラベルフリーの理論的パスにより、評価者の審議がワントークンオプションのスコアリングから分離され、リークやフロアスコアイベントを明らかにしながら候補の可能性が維持されます。ライブ PubMed 検索、部分的にライブ DeepSeek アクター、および DeepSeek 対数確率スコアリングを使用した 100 件のレビュー済みオープン QA 証拠探索タスクについて、ASV は 1,100 のステップと 2,200 の状態を評価します。固定レイアウトの根拠条件付きプロトコルの下では、平均ゴールドマージンゲインは -2.335 (軌道ブートストラップ 95\% CI [-3.395, -1.272])、エントロピーの動きは 0.000、平均ベイジアンサプライズは 2.693 です。したがって、ASV は、最終回答スコアとエントロピーのみのステップ メトリクスが見逃している建設的および破壊的な信念ピボットを特定します。スタンドアロンの ASV Eval ツールキットをリリースします。

原文 (English)

Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces

Pooling, substituting, or reusing evaluator-derived step rewards assumes that their direction survives a change of evaluation channel. The same frozen transition can violate that assumption. Process rewards vary agent states, while evaluator audits vary scoring configurations; neither first difference isolates their interaction. We define Agent Step Value (ASV) as a channel-indexed target-margin gain and identify the state-by-channel interaction on complete matched faces. Across frozen PubMed question-answering transitions, direct scoring yields a positive mean ASV, while the generated-view channel yields a negative mean. Two matched replay waves reproduce this reversal, and cross-channel sign disagreement exceeds same-channel retry disagreement by 48.0 percentage points. Matched retrieval faces localize the reversal to the generated-view coordinate and trace its direction across a readout-and-stack bridge. A source-only generation contract restores the positive mean direction on artifact-bearing retrievals and removes parser-detected substantive support claims from artifact-free before-state views. ASV turns channel sensitivity into an identified measurement problem that can be localized and tested by intervention before step rewards are reused.

2026-07-20 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

AI評価に項目反応理論は信頼できるか?

AI ベンチマークでは、モデルの機能を推定し、システムをランク付けし、有益な例を選択し、ベンチマークの品質を診断するために、項目レベルの統計モデル、特に項目応答理論 (IRT) をますます活用しています。ただし、AI ベンチマーク データは、標準的な IRT 推定ツールが元々開発された人間によるテストのデータ体制から逸脱することがよくあります。ベンチマークには、通常、評価されるモデルが少なく、項目がはるかに多く、偏ったり、クラスター化されたり、マルチモーダルになったりする可能性のある機能分布が含まれます。これらのレジームの不一致が AI 評価のための IRT モデリングの信頼性にどのように影響するかを調査します。広く使用されている 6 つの LLM ベンチマークから導出された項目パラメーターと能力分布を使用して、3 つの一般的な IRT モデルで応答行列をシミュレートし、最近のベンチマーク研究で使用された 4 つの推定ツール (周辺最尤法、マルコフ連鎖モンテカルロ、変分推論、ニューラル擬似シャム推定器) を比較します。 18,000 のシミュレーション条件にわたって、計算の実行可能性、スケーラビリティ、モデルのランキング、予測パフォーマンス、アイテムの特性に関する IRT 推論の信頼性を体系的に評価します。結果は、従来の推定量は大規模なベンチマーク設定では実行不可能になる可能性がある一方、スケーラブルな推定量は小規模または非正規分布のモデル セットでは信頼性の低い項目レベルの推論やランキング推論を生成する可能性があることを示しています。この研究では、潜在特性モデルが AI ベンチマークの主張を確実にサポートする場合、または歪めるリスクがある場合、および信頼できる使用にはどのようなサンプル サイズと診断が必要であるかを特定します。

原文 (English)

Can We Trust Item Response Theory for AI Evaluation?

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or non-normally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.

2026-07-20 13:00 JSTarXiv cs.AIビジネス/資金調達

A Scaffolded GenAI Lab in Early Undergraduate CS: A Mixed-Methods, Multi-Course Evaluation

Background and Context. Generative AI (GenAI) tools are increasingly used in programming courses, but we have limited evidence about how br…

2026-07-20 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length

The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream application…

2026-07-20 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models

The worldwide rise of authoritarianism and the growing role of Large Language Models (LLMs) in users' everyday lives raise the question of…

2026-07-20 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

HiLSVA: 科学的視覚化のための人間参加型エージェント システムの設計と評価

大規模言語モデル (LLM) エージェントは、科学的視覚化 (SciVis) のための自然言語対話を可能にします。それでも、従来のシステムは基本的に人間による分析制御よりも自律性を優先しており、そのため透明性と人間による監視が制限されていました。混合イニシアチブの SciVis ワークフローをサポートする人間参加型エージェント システムである HiLSVA を紹介します。 HiLSVA は、計画優先のマルチエージェント アーキテクチャと、人間による明示的な監視、段階的な出所追跡、およびユーザー フィードバックからのテスト時の学習の適応を統合します。このシステムは、自然言語と視覚化の直接操作の両方を通じて、人間とエージェント間の流動的なハンドオフをサポートし、サンドボックス実行により安全で再現可能なワークフローを保証します。そうすることで、HiLSVA はエージェント的 SciVis を、人間の分析的推論を置き換えるのではなく、強化する共同プロセスとして再構成します。私たちは、代表的なケーススタディと、複数の自律性設定にわたるさまざまな専門知識を持つ 12 人の参加者による対照ユーザー研究を通じて HiLSVA を評価します。結果は、混合イニシアチブの相互作用により、さまざまなレベルのユーザーの専門知識にわたってタスクの完了、ユーザー制御、およびワークフローの透明性が向上する一方、実行効率と人間の監視の間のトレードオフが明らかになったことが示されています。これらの発見は、エージェント的 SciVis における人間中心設計の重要性を強調し、将来の共同視覚化システムの開発の指針となります。 https://hilsva.github.io/ でデモビデオ、ケーススタディ、ソースコードを探索することをお勧めします。

原文 (English)

HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization

Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have essentially prioritized autonomy over human analytical control, thereby limiting transparency and human oversight. We present HiLSVA, a human-in-the-loop agentic system that supports mixed-initiative SciVis workflows. HiLSVA integrates a plan-first multi-agent architecture with explicit human oversight, stepwise provenance tracking, and learn-at-test-time adaptation from user feedback. The system supports fluid handoff between humans and agents through both natural language and direct manipulation of visualizations, while sandboxed execution ensures safe, reproducible workflows. In doing so, HiLSVA reframes agentic SciVis as a collaborative process that augments, rather than replaces, human analytical reasoning. We evaluate HiLSVA through representative case studies and a controlled user study with twelve participants of varying expertise across multiple autonomy settings. Results show that mixed-initiative interaction improves task completion, user control, and workflow transparency across different levels of user expertise, while revealing a tradeoff between execution efficiency and human oversight. These findings highlight the importance of human-centered design in agentic SciVis and guide the development of future collaborative visualization systems. We encourage readers to explore our demo video, case studies, and source code at https://hilsva.github.io/.

2026-07-20 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery…

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

言語モデルの誠実性評価における機器効果: 監査可能な単一システムのデモンストレーション

言語モデルの誠実さの評価では、モデルの評決をモデルに関する証拠として読み取ります。代わりに機器をテストします。私たちは、どのモデルでもゲーム エンジンがクエストを完了できるかどうかを認識する、テキスト アドベンチャーの世界を構築しました。言語モデルは予算内で実行され、最終的にはその探求が完了したか、到達不可能か、またはまだ決定不可能であることを宣言する必要があります。エンジンはすべての判定を採点します。決定ルールは結果が読み取られる前に記録され、実行アーティファクトは実行されたリビジョンをバインドします。事前登録の強さはシリーズごとに異なり、公開されています。演奏者が固定されている場合、楽器の選択によって測定される動作が大きく変わります。 4 バイト同一のアンカーでは、2 つの評決文法を 3 つの評決に拡張すると、強い主張は 38/40 から 7/40 に移動しましたが、新しい不完全な評決では 28/40 の結果が得られました。シリーズ 2 全体で、93/158 の有効なゲームが不完全終了しました。達成基準を開示する 1 つの文では、より少ない意思決定ポイントとよりクリーンな決定により、一致したインスタンスの誤った判定が 18/59 から 0/58 に減少しました。 1 つの固定構成を繰り返し実行すると、4 つのインスタンスのうち 3 つで不安定な判定分布が生成されました。単一の実行では、サンプルが性質として報告されます。正式に事前登録されたナラティブレジスター勾配が改ざんされました。事後的な仮説生成パターンが 2 つ残っています。レジスターの存在により有力な主張が約 2 倍になり、予算レンダリングによりレジスターの内容よりも多くの評決が動かされました (0.383 メートル対 0.150 ランタン)。ナレーターは、不足しているランドマークに向けて豊富な予算を圧縮しましたが、登録された調停テストでは null が返されました。私たちは、評価機器用の 4 つのチェック整合性プロトコルを提案します。

原文 (English)

Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration

Evaluations of language-model honesty read the model's verdicts as evidence about the model. We test the instrument instead. We built a text-adventure world where the game engine, not any model, knows whether the quest can be completed. A language model plays under a budget and must eventually declare its quest complete, unreachable, or not yet decidable; the engine scores every verdict. Decision rules were recorded before results were read, and run artifacts bind the revisions they executed; the strength of preregistration varies by series and is disclosed. With the player held fixed, instrument choices substantially changed measured behavior. On four byte-identical anchors, expanding a two-verdict grammar to three verdicts moved strong claims from 38/40 to 7/40, while the new incomplete verdict took 28/40 outcomes; across series 2, 93/158 valid games ended incomplete. One sentence disclosing the success criterion took matched-instance false verdicts from 18/59 to 0/58, through fewer decision points and cleaner decisions. Repeated runs of one fixed configuration produced non-stable verdict distributions on 3 of 4 instances: single runs report samples as dispositions. A formally preregistered narrative-register gradient was falsified; two post-hoc, hypothesis-generating patterns remain: register presence roughly doubled strong claims, and budget rendering moved verdicts more than register content (.383 meter vs .150 lantern). The narrator compressed abundant budgets toward scarcity landmarks, yet the registered mediation test returned a null. We propose a four-check integrity protocol for eval instruments.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

動的なマルチターンインタラクションによるビジョン言語モデルのコンテキスト化された評価

マルチモーダル大規模言語モデル (MLLM) はベンチマークにおいて大幅な進歩を遂げていますが、現実世界での有効性は依然として不確実です。このギャップは、制御された静的な設定におけるベンチマークと、動的でインタラクティブでコンテキストに応じた現実世界のアプリケーションの性質との間の根本的な不一致から生じます。このギャップを埋めるために、私たちは CEDI (動的なマルチラウンド インタラクションによる MLLM のコンテキスト化評価) を提案します。これは、評価を被評価者モデル、自動試験官、および採点者の間の三者間の対話として再構築するフレームワークです。試験官は、タスクのグラフベースの表現に基づいて、複数ターンの半構造化された会話を行います。状態空間の遷移をナビゲートすることで、CEDI は明確化リクエストから敵対的調査に至るまで、パフォーマンスの証拠を引き出すさまざまな戦略を展開します。 CEDI を幻視に適用します。複数のモデル、多様な設定、データセット、およびドメインにわたる実証結果は、コンテキスト化されたインタラクティブな評価により、従来の静的評価よりも大幅に多くの幻覚が明らかになるだけでなく、実際の使用例で発生する幻覚とよりよく似た幻覚も明らかになることを示しています。さらに、幻覚は自己強化的な対話履歴を通じて長い文脈にわたって蓄積されることが多く、モデルは前提の拒否や拒否を必要とする質問に対して特に脆弱であることを示します。これらの調査結果を総合すると、CEDI が MLLM の能力の現実的、体系的、生態学的に有効な評価に向けた一歩であることが強調されます。コードは github.com/williamium3000/cedi で入手できます。

原文 (English)

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static settings and the dynamic, interactive, and contextualized nature of real-world applications. To bridge this gap, we propose CEDI (Contextualized Evaluations of MLLMs through Dynamic, multi-round Interactions), a framework that recasts evaluation as a three-party interaction between an evaluatee model, an automated examiner, and a grader. The examiner conducts multi-turn, semi-structured conversation guided by a graph-based representation of the task. By navigating state-space transitions, CEDI deploys diverse strategies, from clarification requests to adversarial probes, to elicit performance evidence. We apply CEDI to visual hallucinations. Empirical results across multiple models, diverse settings, datasets, and domains show that contextualized, interactive evaluations reveal not only significantly more hallucinations than conventional static evaluation but also ones that more closely resemble those arising in practical use cases. We further show that hallucinations often accumulate over long contexts, through self-reinforcing dialogue history, and models are particularly vulnerable to questions requiring premise rejection or refusal. Together, these findings highlight CEDI as a step toward realistic, systematic, and ecologically valid assessments of MLLMs' capabilities. Code is available at github.com/williamium3000/cedi.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

WrAFT: 論証エッセイのためのモジュール化された自動ライティング評価システム

この研究では、正確で信頼できるスコアと、論拠のあるエッセイに対する効果的な包括的なフィードバックの両方を提供するライティング評価およびフィードバック ツールである WrAFT を紹介します。 WrAFT は、自動ライティング評価 (AWE) タスクをスコアリング、表面レベルのフィードバック、および深いレベルのフィードバックに分割するモジュール設計を採用しています。システムの構築では、LLaMA-3.3-70B-Instruct、GPT-4o、Claude 3.7 などのさまざまな大規模言語モデル (LLM) が、直接プロンプトと監視付き微調整アプローチの両方を通じて評価されました。公式ベンチマークスコアを含む 480 件の TOEFL Independent Writing エッセイの独自のデータセットが利用されました。ベンチマークベースの評価では、WrAFT が 0 ~ 5 のスケールの公式スコアに対して 2 次加重カッパ (QWK) が 0.84、二乗平均平方根誤差 (RMSE) が 0.44 という、スコアリングにおいて最先端のパフォーマンスを達成していることが示されています。システムが生成したフィードバックを人間が評価したところ、高い支持率が得られました。表面レベルのフィードバックでは 96.14 パーセント、深いレベルのマクロ フィードバックでは 93.03 パーセント、そして深いレベルのミクロ フィードバックでは 94.69 パーセントでした。このシステム用に対話型ユーザー インターフェイスが開発されており、公開されており、無料で使用できます。

原文 (English)

WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays

This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays. WrAFT adopts a modular design by dividing automated writing evaluation (AWE) tasks into scoring, surface-level feedback, and deep-level feedback. In building the system, various Large Language Models (LLMs) have been evaluated, including LLaMA-3.3-70B-Instruct, GPT-4o, and Claude 3.7, through both direct prompting and supervised fine-tuning approaches. A proprietary dataset of 480 TOEFL Independent Writing essays with official benchmark scores was utilized. Benchmark-based evaluation shows that WrAFT achieves state-of-the-art performance in scoring, with a quadratic weighted kappa (QWK) of 0.84 and a root mean square error (RMSE) of 0.44 against official scores on a scale of 0-5. Human evaluation of system-generated feedback also reveals high approval ratings: 96.14 percent for surface-level feedback, 93.03 percent for deep-level macro feedback, and 94.69 percent for deep-level micro feedback. An interactive user interface has been developed for the system and is publicly available and free to use.

2026-07-18 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

プロジェクト カレイドスコープ: 現実世界の AI アプリケーション向けのコンテキストに基づく、人間に合わせた評価

評価 (Eval) は、現実世界の AI アプリケーションの導入のボトルネックです。公開ベンチマークがチームのユーザー、コンテキスト、ポリシーと一致することはほとんどなく、人間によるレビューは拡張するのに面倒なことがよくあります。このプロジェクトは、公共部門における AI アプリケーションの取り組みを動機として、アプリケーションが地域のポリシーとガバナンスの要件を満たさなければならないときに遭遇する、繰り返し発生する評価の課題に対処します。ペルソナベースのテスト生成、コンテキスト化されたルーブリック、および信頼性ゲートによる自動スコアリングのための人によるレビューをリンクする、コンテキストに応じた機能評価のための統合ワークフローであるカレイドスコープを紹介します。生成されたテスト ケースは、アプリケーション固有のルーブリックに対してスコア付けされます。人間による注釈はレビュー可能なラベルを提供します。 LLM 審査員は、これらのラベルとの合意が設定されたしきい値を満たした場合にのみ採点を自動化します。したがって、Kaleidscope は、製品チームにとって実用的で検査可能な反復的なワークフローです。私たちは、4 つの組織ユースケースにわたる 3 週間のパイロット実験と、4 つのドメインと 14 の評価次元にわたる 108 の注釈付き Q&A ペアに関するカスタム ルーブリック裁判官の実験から得られた初期の証拠を報告します。結果は、エンドツーエンドの信頼性の高い自動スコアリングのための便利な機能を強調しています。

原文 (English)

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale. Motivated by our work with AI applications in the public sector, this project addresses recurring evaluation challenges encountered when applications must satisfy local policy and governance requirements. We present Kaleidoscope, an integrated workflow for contextual functional evaluation that links persona-based test generation, contextualized rubrics, and human review for reliability-gated automated scoring. Generated test cases are scored against application-specific rubrics; human annotations provide reviewable labels; and LLM judges automate scoring only when their agreement with those labels meets a configured threshold. Kaleidoscope is therefore a practical, inspectable, iterative workflow for product teams. We report early evidence from a three-week pilot across four organizational use cases and custom-rubric judge experiments on 108 annotated Q\&A pairs spanning four domains and 14 evaluation dimensions. The results highlight useful features for end-to-end reliable, automated scoring.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

科学的視覚化リテラシーのためのマルチモーダル大規模言語モデルのベンチマーク

マルチモーダル大規模言語モデル (MLLM) は、ビジュアライゼーションを解釈するためにますます使用されていますが、現在の評価は依然として主にチャート中心であり、科学的ビジュアライゼーション (SciVis) の理解を示す証拠は限られています。私たちは、科学的視覚化リテラシー評価テストで 6 つの MLLM をベンチマークします。このテストは、8 つのテクニックと 11 のタスク タイプにわたる、18 の科学的視覚化とイラストに基づく 49 項目で構成される標準化された SciVis リテラシー評価です。私たちは、クローズドワールドプロトコルの下で 3 つのクローズドソースモデルと 3 つのオープンソースモデルを評価し、485 人の人間の参加者からのデータを使用してパフォーマンスを比較します。結果は、現在の MLLM が均一な SciVis リテラシーを示さないことを示しています。 Gemini は全体として最も強力なモデルであり、評価されたサブセット全体で人間の平均を上回っていますが、オープンソース モデルは依然として人間のベースラインを下回っています。パフォーマンスはテクニックやタスクによって大きく異なります。モデルは科学的なイラスト、検索、空間理解では最高のパフォーマンスを発揮しますが、テクスチャ ベースおよび統合ベースの視覚化と定量的推定では苦戦します。エラー分析により、きめの細かい定量的推定、フロー方向の解釈、および根拠のあるエンコードの解釈における繰り返しの失敗が明らかになります。これらの調査結果は、SciVis リテラシーをマルチモーダル AI システムを評価するために必要なベンチマークの側面として位置づけています。コードとモデルの出力は、https://github.com/patdmp/mllm-scivis-lit-benchmark で公開されています。

原文 (English)

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at https://github.com/patdmp/mllm-scivis-lit-benchmark.

2026-07-18 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

AI評価に項目反応理論は信頼できるか?

AI ベンチマークでは、モデルの機能を推定し、システムをランク付けし、有益な例を選択し、ベンチマークの品質を診断するために、項目レベルの統計モデル、特に項目応答理論 (IRT) をますます活用しています。ただし、AI ベンチマーク データは、標準的な IRT 推定ツールが元々開発された人間によるテストのデータ体制から逸脱することがよくあります。ベンチマークには、通常、評価されるモデルが少なく、項目がはるかに多く、偏ったり、クラスター化されたり、マルチモーダルになったりする可能性のある機能分布が含まれます。これらのレジームの不一致が AI 評価のための IRT モデリングの信頼性にどのように影響するかを調査します。広く使用されている 6 つの LLM ベンチマークから導出された項目パラメーターと能力分布を使用して、3 つの一般的な IRT モデルで応答行列をシミュレートし、最近のベンチマーク研究で使用された 4 つの推定ツール (周辺最尤法、マルコフ連鎖モンテカルロ、変分推論、ニューラル擬似シャム推定器) を比較します。 18,000 のシミュレーション条件にわたって、計算の実行可能性、スケーラビリティ、モデルのランキング、予測パフォーマンス、アイテムの特性に関する IRT 推論の信頼性を体系的に評価します。結果は、従来の推定量は大規模なベンチマーク設定では実行不可能になる可能性がある一方、スケーラブルな推定量は小規模または非正規分布のモデル セットでは信頼性の低い項目レベルの推論やランキング推論を生成する可能性があることを示しています。この研究では、潜在特性モデルが AI ベンチマークの主張を確実にサポートする場合、または歪めるリスクがある場合、および信頼できる使用にはどのようなサンプル サイズと診断が必要であるかを特定します。

原文 (English)

Can We Trust Item Response Theory for AI Evaluation?

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or nonnormally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

シンプルさのパラドックス: LLM 評価のプロンプトとデータセットに関する誤解を解く

大規模言語モデル (LLM) の機能を調査し、多肢選択質問応答 (MCQA) のための堅牢なソリューションを構築することは、依然として自然言語理解における中心的な課題です。さらに、LLM の急速な普及により、より洗練されたプロンプト技術がより優れたパフォーマンスを生み出すという暗黙の前提が生まれました。いくつかの研究では、より洗練されたプロンプト技術を使用するとパフォーマンスが向上すると主張していますが、包括的な評価は提供されていません。私たちは、27 のモデル構成と 430,000 回以上評価された約 4,300 の固有の質問を含む、10 の多肢選択質問応答 (MCQA) データセットにわたる 8 つのプロンプト手法の包括的な実証研究を通じて、このギャップに対処します。私たちの調査結果は、ベースライン プロンプトがさまざまなベンチマークで複雑な推論手法よりも一貫して優れているという驚くべき矛盾を明らかにしました。最小限のエキスパートおよび帰納的ロール フレーミング (CoT-Expert および CoT-Inductive) のみが、ベースラインに対して小さいながらも統計的に有意な $\sim$3 パーセンテージ ポイント (pp) の向上をもたらしますが、テストした他のすべての精巧なテクニックは、多くの場合、大きなマージンでそれに匹敵するか、パフォーマンスを下回っています (自己類推の場合は最大 31~pp)。さらに、3 つの重要な現象を調査します。(1) Elo 評価における Qwen3-30B-A3B-Thinking-2507 の予期せぬ勝利、(2) 異なる思考予算を持つモデル バリアント間でのパフォーマンスと効率のトレードオフ、モデル依存の最適な構成が明らかに、(3) データセットの難易度には大幅な変動があり、ベンチマークの 60% が 70% を下回っており、最も簡単なものから最も難しいものまで 47.5 pp の広がりがあり、かなりの余地があることが示されています。モデル改良のため。これらの結果は、LLM 評価コミュニティがプロンプト エンジニアリングを複雑にしすぎている可能性があり、さまざまなベンチマーク間で大幅なパフォーマンスのギャップが残っており、プロンプトの最適化ではなく真のモデル改善の機会を提供していることを示唆しています。

原文 (English)

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding. Furthermore, the rapid proliferation of LLMs has created the implicit assumption that more sophisticated prompting techniques yield better performance. Several studies claim better performance with more sophisticated prompting techniques, but do not provide a comprehensive evaluation. We address this gap through a comprehensive empirical study of 8 prompting techniques across 10 multiple-choice question answering (MCQA) datasets, encompassing 27 model configurations and roughly 4,300 unique questions evaluated more than 430,000 times. Our findings reveal a striking paradox that baseline prompting consistently outperforms complex reasoning techniques on various benchmarks. Only minimal expert and inductive role framing (CoT-Expert and CoT-Inductive) yields a small but statistically significant $\sim$3 percentage-point (pp) gain over baseline whereas every other elaborate technique we tested matches or under-performs it, often by large margins (up to 31~pp for Self-Analogical). We further investigate three critical phenomena: (1) the unexpected victory of Qwen3-30B-A3B-Thinking-2507 in Elo ratings, (2) the performance-efficiency trade-offs across model variants with different thinking budgets, revealing model-dependent optimal configurations, and (3) the substantial variation in dataset difficulty, with 60% of benchmarks below 70% accuracy and a 47.5~pp spread from easiest to hardest, indicating considerable room for model improvement. These results suggest that the LLM evaluation community may be overcomplicating prompt engineering and that substantial performance gaps remain across diverse benchmarks, offering opportunities for genuine model improvements rather than prompt optimization.

2026-07-18 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

動的なヒューマノイド全身制御のためのセマンティックなオーディオ駆動型の理解

近年のヒューマノイドロボット工学と強化学習の進歩により、表現力の高い全身運動ポリシーの獲得が可能になりました。しかし、ロボットのパフォーマンスのほとんどは、事前にスクリプト化されたシーケンスまたは外部からトリガーされた動作に基づいたままであり、動的環境に対する自律性や応答性が制限されています。この研究では、セマンティックなオーディオ駆動型ヒューマノイド制御のための新しいマルチモーダル オーケストレーション フレームワークを導入し、ロボットが適切なモーション スキルをリアルタイムで自律的に選択して実行できるようにします。システムは連続オーディオ ストリームを処理し、それらを音楽または音声ブランチにルーティングします。音楽入力は、オーディオ フィンガープリンティングとセマンティック エンベディングを介して処理され、トラックのアイデンティティと時間的アライメントを取得し、音楽セグメントとモーション ポリシー間の動的なマッピングを可能にします。音声入力は、模倣によって学習されたスキルの個別のライブラリに統合され、人間とロボットの直接的な対話が可能になります。どちらのモダリティも、強化学習制御パイプラインを介してスキルの実行をスケジュールする統合インターフェイスを共有します。このアプローチをシミュレーションと Unitree G1 ヒューマノイドで検証し、堅牢なシミュレーションからリアルへの転送と一貫したオーディオ条件付きポリシー選択を示します。補足資料は次のサイトで入手できます: https://lab-rococo-sapienza.github.io/semantic-WBC/

原文 (English)

Semantic Audio-driven Understanding for Dynamic Humanoid Whole Body Control

Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies. However, most robotic performances remain based on pre-scripted sequences or externally triggered behaviors, limiting autonomy and responsiveness to dynamic environments. In this work, we introduce a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time. The system processes continuous audio streams and routes them into music or speech branches. Music input is handled via audio fingerprinting and semantic embeddings to retrieve track identity and temporal alignment, allowing dynamic mapping between musical segments and motion policies. Speech input is grounded into a discrete library of imitation-learned skills, enabling direct human-robot interaction. Both modalities share a unified interface that schedules skill execution over a reinforcement learning control pipeline. We validate the approach in simulation and on a Unitree G1 humanoid, showing robust sim-to-real transfer and consistent audio-conditioned policy selection. Supplementary materials are available at the following site: https://lab-rococo-sapienza.github.io/semantic-WBC/

2026-07-18 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

Instant NuRec: 運転シーン シミュレーションのためのフィードフォワード 3D ガウス再構成

3D シミュレーション プラットフォームは、エンドツーエンドのポリシー評価を可能にし、それによって開発コストを削減し、安全性を向上させるため、自動運転にとって不可欠です。近年ではニューラルシミュレーションが主流となり、NuRecなどの手法が中心的な役割を果たしています。ただし、これらの方法は依然として比較的遅いため、通常はシーンごとの調整が必要です。この研究では、短いマルチビュー運転ログを 1 回のフォワード パスで完全にシミュレーション可能な 3D ガウス スプラッティング (3DGS) ワールドに変換するフィードフォワード ニューラル再構成モデ​​ルである Instant NuRec を紹介します。このモデルは、キャリブレーションされたカメラ リグからのマルチビュー入力を受け入れ、静的および動的 3DGS レイヤー、スカイ キューブマップ、およびカメラごとの ISP 補正で構成されるレイヤー出力を出力すると同時に、3DGUT を介して非ピンホール カメラ モデルのネイティブ サポートを提供します。 10 ~ 20 秒のマルチカメラ シーンを約 1.5 秒で再構築し、Waymo Open Dataset 上で最も強い評価ベースラインを 2.01 dB 上回る PSNR を達成します。 Instant NuRec は NuRec に深く統合されており、閉ループ シミュレーション用の AlpaSim と互換性があります。

原文 (English)

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

3D simulation platforms are critical for autonomous driving because they enable end-to-end policy evaluation, thereby reducing development costs and improving safety. In recent years, neural simulation has become predominant, with methods such as NuRec playing a central role; however, these methods remain relatively slow and typically require per-scene tuning. In this work, we present Instant NuRec, a feed-forward neural reconstruction model that turns a short multi-view driving log into a fully simulatable 3D Gaussian Splatting (3DGS) world in a single forward pass. The model accepts multi-view input from a calibrated camera rig and emits a layered output consisting of static and dynamic 3DGS layers, a sky cubemap, and per-camera ISP corrections, while providing native support for non-pinhole camera models via 3DGUT. It reconstructs a 10-20-second multi-camera scene in roughly 1.5 seconds and achieves a PSNR on the Waymo Open Dataset that is 2.01 dB above the strongest evaluated baseline. Instant NuRec is deeply integrated into NuRec and is compatible with AlpaSim for closed-loop simulation.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

コピーオンライトのスコアリング: アプリケーション固有のエージェントの評価

ソフトウェア システムに LLM ベースのエージェントを信頼できる展開するには、エージェントがアプリケーション固有のワークフローでどのように実行されるかを、成功と失敗の場所を特定するのに十分な粒度で評価する必要があります。しかし、既存のエージェント評価メカニズムには限界があります。ベンチマークはアプリケーション固有のワークフローや環境に対する構成の妥当性が低く、レプリカ評価環境は高価でドリフトしやすいです。私たちは、エージェントの書き込みを分離するために PostgreSQL レベルのコピーオンライト メカニズムを使用して、アプリケーション環境内でエージェントの操作を直接評価するフレームワークであるコピーオンライト (CoW) スコアリングを提案します。 CoW スコアリングは、特定のアプリケーション環境でエージェントのデータベース書き込み操作が成功した場所と失敗した場所を強調表示するセッション レベルおよび操作レベルのスコアを生成し、エージェント ハーネスとツール サーフェスでの低コストの評価と反復を可能にします。オープンソースのプロジェクト管理プラットフォームである Plane でフレームワークをデモンストレーションします。分析によりツール表面の特定の問題が明らかになり、対応する修正により影響を受けるモデルに目に見える改善がもたらされました。 Python ライブラリ: https://github.com/trail-ml/agent-cow-python

原文 (English)

Copy-on-Write Scoring: Application-Specific Agent Evaluations

Trustworthy deployment of LLM-based agents in software systems requires evaluating how they perform on application-specific workflows, with enough granularity to localize where they succeed and fail. Yet existing agent evaluation mechanisms are limited: benchmarks have low construct validity for application-specific workflows and environments, and replica evaluation environments are expensive and prone to drift. We propose Copy-on-Write (CoW) Scoring, a framework that evaluates agent operations directly within application environments using a PostgreSQL-level Copy-on-Write mechanism to isolate agent writes. CoW Scoring produces session- and operation-level scores that highlight where agents' database write operations succeed and fail in a given application environment, enabling inexpensive evaluation and iteration on agent harnesses and tool surfaces. We demonstrate the framework on Plane, an open-source project-management platform, where analysis surfaced specific issues in the tool surface, and corresponding fixes produced measurable improvements on affected models. Python library: https://github.com/trail-ml/agent-cow-python

2026-07-18 13:00 JSTarXiv cs.AIビジネス/資金調達

認識論的不確実性の評価: OOD 検出とアクティブ ラーニングを超えて

認識論的不確実性の現在の評価は、分布外検出や能動学習などのタスクに依存しています。ただし、これらのタスクのベイズ最適決定戦略は、認識論的不確実性を定量化するために一般的に使用されるスコアと一致しません。認識論的拒否オプションのフレームワークに基づいて、我々は、後悔、つまり削減可能な誤差を特定する能力を使用して認識論的不確実性を評価します。カバレッジ、予想されるリスク、リグレスに対する制約付き最適化として選択的予測を定式化し、最適なセレクターがグラウンドトゥルースの偶然性と認識論的不確実性の閾値付き凸組み合わせであることを証明します。この理論的統一は、最近の不確実性解きほぐしに関する文献の弱点を明らかにしています。学習されたコンポーネント間の標準的な相関指標が、実際の運用上の有用性を必ずしも予測するとは限らないことを示しています。代わりに、関節のもつれの解消と有用性の診断として、達成可能なリスク、後悔、分解の適用範囲を評価することを提案します。高密度のヒューマン アノテーションを含むデータセットで標準メソッドのベンチマークを行うと、ある基準で上位にランクされ、別の基準で最下位にランクされるメソッド間のペアごとの順位逆転など、意思決定理論によるランキングが代理タスクのランキングと実質的に一致しない可能性があることが明らかになりました。

原文 (English)

Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning

Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.

2026-07-18 13:00 JSTarXiv cs.AIビジネス/資金調達

AI が貢献の境界を曖昧にするとき: 著者資格の調整に関する実証的研究

人工知能 (AI)、特にジェネレーティブ AI の広範な導入により、ユーザーがこれらのシステムとどのように対話して新しいコンテンツを作成するかについて差し迫った疑問が生じています。この論文では、AI と対話するときにユーザーが実際の著者であることを認識することとして定義される著者資格調整の概念を紹介します。 CoAuthor データセットを使用して、著者資格の調整がユーザー間でどのように異なるか、またそれが AI の使用頻度とどのように関連しているかを実証的に調査します。私たちの結果は、ばらつきが大きいことを明らかにしました。AI に大きく依存しているユーザーは、自分の著者であるかどうかを誤って判断する傾向があるのに対し、AI をそれほど頻繁に使用していないユーザーは、より正確な著者であるかどうかの調整を示しています。これらの発見は、AI がユーザー自身の作者に対する認識を曖昧にする可能性があることを示唆しています。学習の状況において、誤った調整はメタ認知のモニタリングと学習戦略に影響を与え、最終的には学習成果に影響を与える可能性があります。したがって、責任ある教育的に有意義な AI 統合を促進するには、著者資格の調整を促進することが不可欠であると考えられます。

原文 (English)

When AI Blurs the Boundaries of Contribution: An Empirical Study of Authorship Calibration

The broad adoption of Artificial Intelligence (AI), especially Generative AI, raises pressing questions about how users interact with these systems to produce new content. In this paper, we introduce the concept of authorship calibration, defined as users awareness of their actual authorship when interacting with AI. Using the CoAuthor dataset, we empirically examine how authorship calibration varies across users and how it relates to their frequency of AI use. Our results reveal high variability: users relying heavily on AI tend to misjudge their authorship, whereas those using AI less frequently exhibit more accurate authorship calibration. These findings suggest that AI can obscure users perception of their own authorship. In learning contexts, miscalibration can affect metacognitive monitoring and learning strategies, ultimately impacting learning outcomes. Fostering authorship calibration then appears essential for promoting responsible and educationally meaningful AI integration.

2026-07-18 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery…

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

Large language models are increasingly deployed for legal question answering, where evaluations typically focus on multiple-choice accuracy…

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

推論によるフロンティア LLM 評価の形状計算方法

AI の評価は、ツールの使用と反復的な問題解決を伴う長期にわたる軌道から恩恵を受ける、より困難なタスクへと移行しています。その結果、パフォーマンスは、テスト時に利用可能なコンピューティング (「推論コンピューティング」) の量と割り当てにますます敏感になります。しかし、多くの評価では依然として単一の制限された予算でのパフォーマンスが報告されており、低いスコアはモデルの基礎的な機能ではなく評価設定を反映している可能性があることを意味します。これをテストするために、ソフトウェア エンジニアリング、数学、医学、サイバーセキュリティにわたる 7 つの挑戦的なベンチマークで最大 12 のフロンティア言語モデルを評価します。私たちは、3 つの単純な推論スケーリング介入を組み合わせた制御されたセットアップを使用します。つまり、より大きなトークン バジェット、コンテキストの圧縮、およびモデル自体または最小限の正確性フィードバックによって導かれる送信の試行の繰り返しです。主な結果は 3 つあります。まず、トークン バジェットが大きくなると、サイバーセキュリティ、FrontierMath、人類最後の試験、ターミナルベンチなど、複数のドメインにわたるベンチマークのパフォーマンスが大幅に向上します。第二に、固定予算の評価では、モデルが進歩するにつれてフロンティアの能力がますます過小評価される可能性があります。新しいモデルは、大きな予算でより高いパフォーマンスを実現し、より困難なタスクを解放し、より確実に解決します。第三に、どの推論スケーリング手法が最も役立つかがベンチマークによって異なります。繰り返し送信するとパフォーマンスが大幅に向上しますが、より大きなトークン バジェット、外部フィードバック、および並列試行の値はベンチマークによって異なります。全体として、私たちの結果は、ベンチマーク スコアがプロトコルに依存していることを示しています。したがって、評価では、特に安全性またはポリシー関連の設定において、推論時間のコンピューティングの関数として機能を報告し、プロトコルの選択を明示的に指定し、一致した予算で大規模な共有コンピューティング範囲にわたってモデルの世代を比較する必要があると主張します。

原文 (English)

How Inference Compute Shapes Frontier LLM Evaluation

AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still report performance at a single restrictive budget, meaning that low scores may reflect the evaluation setup rather than the model's underlying capability. To test this, we evaluate up to 12 frontier language models on seven challenging benchmarks spanning software engineering, mathematics, medicine, and cybersecurity. We use a controlled setup combining three simple inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts, guided either by the model itself or by minimal correctness feedback. We find three main results. First, larger token budgets substantially improve performance on benchmarks across multiple domains, including cybersecurity, FrontierMath, Humanity's Last Exam, and TerminalBench. Second, fixed-budget evaluations can increasingly understate frontier capability as models advance. Newer models reach higher performance at large budgets, where they unlock harder tasks and solve them more reliably. Third, benchmarks differ in which inference-scaling methods help most: repeated submission broadly improves performance, but the value of larger token budgets, external feedback, and parallel attempts varies by benchmark. Overall, our results show that benchmark scores are protocol-dependent. We therefore argue that evaluations should report capability as a function of inference-time compute, specify protocol choices explicitly, and compare model generations over a large shared compute range at matched budgets, especially in safety- or policy-relevant settings.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mechanistic interpretability, and governance/auditability, covering 2018-2026 evaluation-safety measurement work. We introduce EvalSafetyGap as an organizing hypothesis for comparing evaluation-side and alignment-side proxy failures under optimization pressure, using Goodhart's Law together with two constructs we develop here - an Instability Decomposition and an Alignment Trilemma - as tools for generating testable comparisons. The audit shows how conclusions shift when capability, behavioral safety, and governance are measured separately. In this sample ($n = 10$), the association between capability and sustained adversarial robustness is statistically indeterminate using the displayed Table 3 inputs (Pearson $r = +0.232$, $p = 0.520$), and the apparent open-closed safety gap is modest, driven mainly by governance and disclosure rather than behavioral robustness, and sensitive to how a single borderline model is classified; attempt-budget results are protocol dependent. Because the public evidence uses heterogeneous protocols, the audit is diagnostic rather than rank-generating. The contribution is a shared vocabulary and evidence map to support dynamic evaluation, transparent source reporting, multi-attempt safety measurement, and auditable alignment practice.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

隠れたフットプリント: ストレージを LLM エージェント評価の第一級の指標にする

LLM エージェントのベンチマークは、タスクの完了、信頼性、推論コストを測定しますが、ログ、コンテキスト スナップショット、チェックポイント、デバッグ トレースなど、エージェントの実行によってディスクに残される永続データは測定しません。実行後のエージェント ストレージ フットプリントのクロスフレームワーク ベンチマークである AgentFootprint を紹介します。そのシリアル化対応メトリクス スイートは、総保持率、チャネル構成、重複、増加、圧縮率、会話履歴の再構築可能性を測定します。これは、測定の罠に対処します。単純なバイトレベルの測定では、データベースのページングと JSON エスケープが繰り返されるコンテンツを不明瞭にするため、重複が桁違いに過小評価されます。固定トレース制御により、エージェントが生成した論理ボリュームが永続層の増幅から分離されます。7 つの永続フレームワークを通じて同じ軌跡を再生すると、6.7 倍の広がりが得られます。同一のモデル、ツール、およびタスクでは、100% の精度の構成では、デフォルトでサポートされる回復機能と監査機能が異なりますが、保持バイト数が 15.7 倍異なります。 3 つの完全な履歴構成は、反復観察ストレス タスクで超線形に成長します。 108 個のインスタンスで正規化された SWE ベンチからエクスポートされた軌跡 検証済みの送信は、インスタンスごとに 3 桁の大きさに及び、解決率との検出可能な相関関係はありません。コンテンツ アドレス ストアは、すべての再構築可能性スコアを維持しながら、保持率を 4.8 倍から 32.7 倍まで削減します。これらの結果は、精度と再構築可能性を併せてレポートするためのリソース メトリックとして永続ストレージを確立します。

原文 (English)

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite measures total retention, channel composition, duplication, growth, compressibility, and conversation-history reconstructability. It addresses a measurement trap: naive byte-level measurement understates duplication by an order of magnitude because database paging and JSON escaping obscure repeated content. A fixed-trace control separates agent-generated logical volume from persistence-layer amplification: replaying the same trajectory through seven persisting frameworks yields a 6.7x spread. Under identical models, tools, and tasks, configurations with 100% accuracy differ by 15.7x in retained bytes, although their defaults support different recovery and audit capabilities. Three full-history configurations grow superlinearly on a repeated-observation stress task. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude per instance, with no detectable correlation with resolve rate. A content-addressed store reduces retention by 4.8x-32.7x while preserving every reconstructability score. These results establish persistent storage as a resource metric to report jointly with accuracy and reconstructability.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AgentCompass: エージェント機能の統合評価インフラストラクチャ

大規模言語モデル (LLM) が自律エージェントに進化するにつれて、統合された評価インフラストラクチャの必要性が重要になります。ただし、現在の評価パイプラインは高度に断片化され、密接に結合されたままであるため、再現性が妨げられ、冗長なエンジニアリングが発生します。これに対処するために、LLM ベースのエージェントを評価するためのオープンソースで軽量かつ拡張可能なインフラストラクチャである AgentCompass を導入します。 AgentCompass は、ベンチマーク、ハーネス、環境という 3 つの独立したコンポーネントを中心に評価プロセスを編成するため、複雑な実行ロジックを再実装することなく柔軟な構成が可能になります。さらに、フォールトトレラントな非同期ランタイムと、報酬ハッキングなどの微妙な障害モードを透過的に診断するための包括的な軌跡分析ツールを備えています。 AgentCompass は、5 つの機能次元にわたる 20 以上のベンチマークをネイティブにサポートし、エージェント研究を進めるためのスケーラブルで再現可能なインフラストラクチャをコミュニティに提供します。

原文 (English)

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.

2026-07-18 13:00 JSTarXiv cs.AIビジネス/資金調達

Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis

Music Structure Analysis (MSA) aims to uncover the high-level organization of musical pieces. State-of-the-art methods are often based on s…

2026-07-18 13:00 JSTarXiv cs.AIビジネス/資金調達

Warning labels shift perceptions of sycophantic AI, but not its influence

Recent work has raised concerns about the influence of sycophantic AI on user judgment and relationships. One proposed mitigation, which ha…

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

AI の安全性評価のための敵対的プラグマティクス: 命令の競合、埋め込みコマンド、およびポリシーの曖昧さのベンチマーク

言語モデルの安全性評価は、モデルが指示に従ったか、適切に拒否したか、ポリシーに従ったか、埋め込まれたコマンドに抵抗したか、エージェントタスクの進捗状況を誤って報告したかなど、あいまいな自然言語の動作に関する判断にますます依存しています。既存のベンチマークは、多くの場合、これらの区別を合格/不合格のラベルに圧縮し、障害が機能制限、ポリシーの曖昧さ、命令の競合、足場の障害、または不安定な評価者の判断に起因するかどうかを曖昧にします。この論文では、命令の競合、埋め込みコマンド、引用、範囲の曖昧さ、明確化、間接音声行為、およびマルチターンエージェントのトランスクリプトの下でのモデルの動作を評価するためのベンチマークおよびアノテーションプロトコルとして、敵対的プラグマティクスを紹介します。この貢献は経験的かつ方法論的です。言語的に管理された分類法、バリデーターが強制するメタデータを備えた 18 項目のシード ベンチマーク、54 行のローカル シード パイロット、タスクの成功、ポリシー遵守、安全性リスク、拒否結果、評価者の信頼性を区別する専門家評価プロトコル、および裁判官の妥当性、診断の曖昧さ、および分類のドリフトの指標です。このフレームワークは、言語的判断方法論を、安全性評価、LLM ジャッジ、ゴールドセット構造、即時注入テスト、および安全性文書を検証するための実用的なツールに変えます。

原文 (English)

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological: a linguistically controlled taxonomy, an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, an expert-evaluation protocol distinguishing task success, policy compliance, safety risk, refusal outcome, and evaluator confidence, and metrics for judge validity, diagnostic ambiguity, and taxonomy drift. The benchmark treats labels as inference licenses: it tests whether safety-relevant categories project across paraphrase, wrapper, model, and judge condition. In the pilot, a rubric-aided LLM judge graded its own outputs with expected-behaviour fields visible and still missed the safety-relevant minority classes.

2026-07-18 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体ビジネス/資金調達

ブラックボックスを制御できるか?協調エージェントを使用したレコメンダーシステムの制御性中心の評価に向けて

レコメンダー システムはブラック ボックスとして動作するため、ユーザーや規制当局は出力を特定の意図に向けたり、その動作を監査したりすることができません。この制御性の欠如は、明示的なガイダンスに応答するシステムの能力として定義されますが、既存の評価パラダイムでは依然として対処されていない側面です。このギャップを埋めるために、制御性を体系的に評価するための協調的なマルチエージェント フレームワークである CtrlBench-Rec を提案します。私たちは、ターゲット コンテンツの発見、関心プロファイルの形成、人気バイアスの軽減という 3 つの基本的なタスクを形式化します。これらは、明示的なコマンドから暗黙的な表現のステアリング、そして最終的にはアルゴリズムのバイアスの克服までのステアビリティを一緒に測定します。実世界のデータセットと複数のレコメンデーション モデルに関する広範な実験により、私たちのフレームワークが制御性を効果的に定量化し、重大なシステムのボトルネック、特に誘導ロングテール コンテンツに対する永続的な抵抗を明らかにすることが実証されています。 CtrlBench-Rec は、制御可能な推奨調査、アルゴリズム監査、およびユーザー権限付与のための初の標準化されたツールキットを提供します。私たちのコードは https://github.com/caskcsg/CtrlBenchRec でリリースされています。

原文 (English)

Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents

Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, remains an unaddressed dimension in existing evaluation paradigms. To fill this gap, we propose CtrlBench-Rec, a collaborative multi-agent framework for systematic assessment of controllability. We formalize three fundamental tasks: target content discovery, interest profile shaping, and popularity bias mitigation, which together measure steerability from explicit commands to implicit representation steering and finally to overcoming algorithmic biases.Extensive experiments on real-world datasets and multiple recommendation models demonstrate that our framework effectively quantifies controllability and exposes critical system bottlenecks, most notably persistent resistance to guiding long tail content. CtrlBench-Rec provides the first standardized toolkit for controllable recommendation research, algorithmic auditing, and user empowerment. Our code is released on https://github.com/caskcsg/CtrlBenchRec.

2026-07-18 07:12 JSTTechCrunch AIビジネス/資金調達研究/論文

Databricks hits $188B valuation, extending its run as AI’s favorite second act

Databricks has remade its image into an AI company and has published research on the cost savings of open weight AI models for coding.

2026-07-18 02:45 JSTTechCrunch AILLM/生成AIビジネス/資金調達規制/政策

How Apple’s big lawsuit could disrupt OpenAI’s IPO plans

Apple filed a trade secrets lawsuit against OpenAI last Friday, and it’s not messing around. The complaint alleges a pattern of misconduct…

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

言語モデルの誠実性評価における機器効果: 監査可能な単一システムのデモンストレーション

言語モデルの誠実さの評価では、モデルの評決をモデルに関する証拠として読み取ります。代わりに機器をテストします。私たちは、どのモデルでもゲーム エンジンがクエストを完了できるかどうかを認識する、テキスト アドベンチャーの世界を構築しました。言語モデルは予算内で実行され、最終的にはその探求が完了したか、到達不可能か、またはまだ決定不可能であることを宣言する必要があります。エンジンはすべての判定を採点します。決定ルールは結果が読み取られる前に記録され、実行アーティファクトは実行されたリビジョンをバインドします。事前登録の強さはシリーズごとに異なり、公開されています。演奏者が固定されている場合、楽器の選択によって測定される動作が大きく変わります。 4 バイト同一のアンカーでは、2 つの評決文法を 3 つの評決に拡張すると、強い主張は 38/40 から 7/40 に移動しましたが、新しい不完全な評決では 28/40 の結果が得られました。シリーズ 2 全体で、93/158 の有効なゲームが不完全終了しました。達成基準を開示する 1 つの文では、より少ない意思決定ポイントとよりクリーンな決定により、一致したインスタンスの誤った判定が 18/59 から 0/58 に減少しました。 1 つの固定構成を繰り返し実行すると、4 つのインスタンスのうち 3 つで不安定な判定分布が生成されました。単一の実行では、サンプルが性質として報告されます。正式に事前登録されたナラティブレジスター勾配が改ざんされました。事後的な仮説生成パターンが 2 つ残っています。レジスターの存在により有力な主張が約 2 倍になり、予算レンダリングによりレジスターの内容よりも多くの評決が動かされました (0.383 メートル対 0.150 ランタン)。ナレーターは、不足しているランドマークに向けて豊富な予算を圧縮しましたが、登録された調停テストでは null が返されました。私たちは、評価機器用の 4 つのチェック整合性プロトコルを提案します。

原文 (English)

Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration

Evaluations of language-model honesty read the model's verdicts as evidence about the model. We test the instrument instead. We built a text-adventure world where the game engine, not any model, knows whether the quest can be completed. A language model plays under a budget and must eventually declare its quest complete, unreachable, or not yet decidable; the engine scores every verdict. Decision rules were recorded before results were read, and run artifacts bind the revisions they executed; the strength of preregistration varies by series and is disclosed. With the player held fixed, instrument choices substantially changed measured behavior. On four byte-identical anchors, expanding a two-verdict grammar to three verdicts moved strong claims from 38/40 to 7/40, while the new incomplete verdict took 28/40 outcomes; across series 2, 93/158 valid games ended incomplete. One sentence disclosing the success criterion took matched-instance false verdicts from 18/59 to 0/58, through fewer decision points and cleaner decisions. Repeated runs of one fixed configuration produced non-stable verdict distributions on 3 of 4 instances: single runs report samples as dispositions. A formally preregistered narrative-register gradient was falsified; two post-hoc, hypothesis-generating patterns remain: register presence roughly doubled strong claims, and budget rendering moved verdicts more than register content (.383 meter vs .150 lantern). The narrator compressed abundant budgets toward scarcity landmarks, yet the registered mediation test returned a null. We propose a four-check integrity protocol for eval instruments.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

動的なマルチターンインタラクションによるビジョン言語モデルのコンテキスト化された評価

マルチモーダル大規模言語モデル (MLLM) はベンチマークにおいて大幅な進歩を遂げていますが、現実世界での有効性は依然として不確実です。このギャップは、制御された静的な設定におけるベンチマークと、動的でインタラクティブでコンテキストに応じた現実世界のアプリケーションの性質との間の根本的な不一致から生じます。このギャップを埋めるために、私たちは CEDI (動的なマルチラウンド インタラクションによる MLLM のコンテキスト化評価) を提案します。これは、評価を被評価者モデル、自動試験官、および採点者の間の三者間の対話として再構築するフレームワークです。試験官は、タスクのグラフベースの表現に基づいて、複数ターンの半構造化された会話を行います。状態空間の遷移をナビゲートすることで、CEDI は明確化リクエストから敵対的調査に至るまで、パフォーマンスの証拠を引き出すさまざまな戦略を展開します。 CEDI を幻視に適用します。複数のモデル、多様な設定、データセット、およびドメインにわたる実証結果は、コンテキスト化されたインタラクティブな評価により、従来の静的評価よりも大幅に多くの幻覚が明らかになるだけでなく、実際の使用例で発生する幻覚とよりよく似た幻覚も明らかになることを示しています。さらに、幻覚は自己強化的な対話履歴を通じて長い文脈にわたって蓄積されることが多く、モデルは前提の拒否や拒否を必要とする質問に対して特に脆弱であることを示します。これらの調査結果を総合すると、CEDI が MLLM の能力の現実的、体系的、生態学的に有効な評価に向けた一歩であることが強調されます。コードは github.com/williamium3000/cedi で入手できます。

原文 (English)

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static settings and the dynamic, interactive, and contextualized nature of real-world applications. To bridge this gap, we propose CEDI (Contextualized Evaluations of MLLMs through Dynamic, multi-round Interactions), a framework that recasts evaluation as a three-party interaction between an evaluatee model, an automated examiner, and a grader. The examiner conducts multi-turn, semi-structured conversation guided by a graph-based representation of the task. By navigating state-space transitions, CEDI deploys diverse strategies, from clarification requests to adversarial probes, to elicit performance evidence. We apply CEDI to visual hallucinations. Empirical results across multiple models, diverse settings, datasets, and domains show that contextualized, interactive evaluations reveal not only significantly more hallucinations than conventional static evaluation but also ones that more closely resemble those arising in practical use cases. We further show that hallucinations often accumulate over long contexts, through self-reinforcing dialogue history, and models are particularly vulnerable to questions requiring premise rejection or refusal. Together, these findings highlight CEDI as a step toward realistic, systematic, and ecologically valid assessments of MLLMs' capabilities. Code is available at github.com/williamium3000/cedi.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

WrAFT: 論証エッセイのためのモジュール化された自動ライティング評価システム

この研究では、正確で信頼できるスコアと、論拠のあるエッセイに対する効果的な包括的なフィードバックの両方を提供するライティング評価およびフィードバック ツールである WrAFT を紹介します。 WrAFT は、自動ライティング評価 (AWE) タスクをスコアリング、表面レベルのフィードバック、および深いレベルのフィードバックに分割するモジュール設計を採用しています。システムの構築では、LLaMA-3.3-70B-Instruct、GPT-4o、Claude 3.7 などのさまざまな大規模言語モデル (LLM) が、直接プロンプトと監視付き微調整アプローチの両方を通じて評価されました。公式ベンチマークスコアを含む 480 件の TOEFL Independent Writing エッセイの独自のデータセットが利用されました。ベンチマークベースの評価では、WrAFT が 0 ~ 5 のスケールの公式スコアに対して 2 次加重カッパ (QWK) が 0.84、二乗平均平方根誤差 (RMSE) が 0.44 という、スコアリングにおいて最先端のパフォーマンスを達成していることが示されています。システムが生成したフィードバックを人間が評価したところ、高い支持率が得られました。表面レベルのフィードバックでは 96.14 パーセント、深いレベルのマクロ フィードバックでは 93.03 パーセント、そして深いレベルのミクロ フィードバックでは 94.69 パーセントでした。このシステム用に対話型ユーザー インターフェイスが開発されており、公開されており、無料で使用できます。

原文 (English)

WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays

This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays. WrAFT adopts a modular design by dividing automated writing evaluation (AWE) tasks into scoring, surface-level feedback, and deep-level feedback. In building the system, various Large Language Models (LLMs) have been evaluated, including LLaMA-3.3-70B-Instruct, GPT-4o, and Claude 3.7, through both direct prompting and supervised fine-tuning approaches. A proprietary dataset of 480 TOEFL Independent Writing essays with official benchmark scores was utilized. Benchmark-based evaluation shows that WrAFT achieves state-of-the-art performance in scoring, with a quadratic weighted kappa (QWK) of 0.84 and a root mean square error (RMSE) of 0.44 against official scores on a scale of 0-5. Human evaluation of system-generated feedback also reveals high approval ratings: 96.14 percent for surface-level feedback, 93.03 percent for deep-level macro feedback, and 94.69 percent for deep-level micro feedback. An interactive user interface has been developed for the system and is publicly available and free to use.

2026-07-17 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

プロジェクト カレイドスコープ: 現実世界の AI アプリケーション向けのコンテキストに基づく、人間に合わせた評価

評価 (Eval) は、現実世界の AI アプリケーションの導入のボトルネックです。公開ベンチマークがチームのユーザー、コンテキスト、ポリシーと一致することはほとんどなく、人間によるレビューは拡張するのに面倒なことがよくあります。このプロジェクトは、公共部門における AI アプリケーションの取り組みを動機として、アプリケーションが地域のポリシーとガバナンスの要件を満たさなければならないときに遭遇する、繰り返し発生する評価の課題に対処します。ペルソナベースのテスト生成、コンテキスト化されたルーブリック、および信頼性ゲートによる自動スコアリングのための人によるレビューをリンクする、コンテキストに応じた機能評価のための統合ワークフローであるカレイドスコープを紹介します。生成されたテスト ケースは、アプリケーション固有のルーブリックに対してスコア付けされます。人間による注釈はレビュー可能なラベルを提供します。 LLM 審査員は、これらのラベルとの合意が設定されたしきい値を満たした場合にのみ採点を自動化します。したがって、Kaleidscope は、製品チームにとって実用的で検査可能な反復的なワークフローです。私たちは、4 つの組織ユースケースにわたる 3 週間のパイロット実験と、4 つのドメインと 14 の評価次元にわたる 108 の注釈付き Q&A ペアに関するカスタム ルーブリック裁判官の実験から得られた初期の証拠を報告します。結果は、エンドツーエンドの信頼性の高い自動スコアリングのための便利な機能を強調しています。

原文 (English)

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale. Motivated by our work with AI applications in the public sector, this project addresses recurring evaluation challenges encountered when applications must satisfy local policy and governance requirements. We present Kaleidoscope, an integrated workflow for contextual functional evaluation that links persona-based test generation, contextualized rubrics, and human review for reliability-gated automated scoring. Generated test cases are scored against application-specific rubrics; human annotations provide reviewable labels; and LLM judges automate scoring only when their agreement with those labels meets a configured threshold. Kaleidoscope is therefore a practical, inspectable, iterative workflow for product teams. We report early evidence from a three-week pilot across four organizational use cases and custom-rubric judge experiments on 108 annotated Q\&A pairs spanning four domains and 14 evaluation dimensions. The results highlight useful features for end-to-end reliable, automated scoring.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

科学的視覚化リテラシーのためのマルチモーダル大規模言語モデルのベンチマーク

マルチモーダル大規模言語モデル (MLLM) は、ビジュアライゼーションを解釈するためにますます使用されていますが、現在の評価は依然として主にチャート中心であり、科学的ビジュアライゼーション (SciVis) の理解を示す証拠は限られています。私たちは、科学的視覚化リテラシー評価テストで 6 つの MLLM をベンチマークします。このテストは、8 つのテクニックと 11 のタスク タイプにわたる、18 の科学的視覚化とイラストに基づく 49 項目で構成される標準化された SciVis リテラシー評価です。私たちは、クローズドワールドプロトコルの下で 3 つのクローズドソースモデルと 3 つのオープンソースモデルを評価し、485 人の人間の参加者からのデータを使用してパフォーマンスを比較します。結果は、現在の MLLM が均一な SciVis リテラシーを示さないことを示しています。 Gemini は全体として最も強力なモデルであり、評価されたサブセット全体で人間の平均を上回っていますが、オープンソース モデルは依然として人間のベースラインを下回っています。パフォーマンスはテクニックやタスクによって大きく異なります。モデルは科学的なイラスト、検索、空間理解では最高のパフォーマンスを発揮しますが、テクスチャ ベースおよび統合ベースの視覚化と定量的推定では苦戦します。エラー分析により、きめの細かい定量的推定、フロー方向の解釈、および根拠のあるエンコードの解釈における繰り返しの失敗が明らかになります。これらの調査結果は、SciVis リテラシーをマルチモーダル AI システムを評価するために必要なベンチマークの側面として位置づけています。コードとモデルの出力は、https://github.com/patdmp/mllm-scivis-lit-benchmark で公開されています。

原文 (English)

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at https://github.com/patdmp/mllm-scivis-lit-benchmark.

2026-07-17 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

AI評価に項目反応理論は信頼できるか?

AI ベンチマークでは、モデルの機能を推定し、システムをランク付けし、有益な例を選択し、ベンチマークの品質を診断するために、項目レベルの統計モデル、特に項目応答理論 (IRT) をますます活用しています。ただし、AI ベンチマーク データは、標準的な IRT 推定ツールが元々開発された人間によるテストのデータ体制から逸脱することがよくあります。ベンチマークには、通常、評価されるモデルが少なく、項目がはるかに多く、偏ったり、クラスター化されたり、マルチモーダルになったりする可能性のある機能分布が含まれます。これらのレジームの不一致が AI 評価のための IRT モデリングの信頼性にどのように影響するかを調査します。広く使用されている 6 つの LLM ベンチマークから導出された項目パラメーターと能力分布を使用して、3 つの一般的な IRT モデルで応答行列をシミュレートし、最近のベンチマーク研究で使用された 4 つの推定ツール (周辺最尤法、マルコフ連鎖モンテカルロ、変分推論、ニューラル擬似シャム推定器) を比較します。 18,000 のシミュレーション条件にわたって、計算の実行可能性、スケーラビリティ、モデルのランキング、予測パフォーマンス、アイテムの特性に関する IRT 推論の信頼性を体系的に評価します。結果は、従来の推定量は大規模なベンチマーク設定では実行不可能になる可能性がある一方、スケーラブルな推定量は小規模または非正規分布のモデル セットでは信頼性の低い項目レベルの推論やランキング推論を生成する可能性があることを示しています。この研究では、潜在特性モデルが AI ベンチマークの主張を確実にサポートする場合、または歪めるリスクがある場合、および信頼できる使用にはどのようなサンプル サイズと診断が必要であるかを特定します。

原文 (English)

Can We Trust Item Response Theory for AI Evaluation?

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or nonnormally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

シンプルさのパラドックス: LLM 評価のプロンプトとデータセットに関する誤解を解く

大規模言語モデル (LLM) の機能を調査し、多肢選択質問応答 (MCQA) のための堅牢なソリューションを構築することは、依然として自然言語理解における中心的な課題です。さらに、LLM の急速な普及により、より洗練されたプロンプト技術がより優れたパフォーマンスを生み出すという暗黙の前提が生まれました。いくつかの研究では、より洗練されたプロンプト技術を使用するとパフォーマンスが向上すると主張していますが、包括的な評価は提供されていません。私たちは、27 のモデル構成と 430,000 回以上評価された約 4,300 の固有の質問を含む、10 の多肢選択質問応答 (MCQA) データセットにわたる 8 つのプロンプト手法の包括的な実証研究を通じて、このギャップに対処します。私たちの調査結果は、ベースライン プロンプトがさまざまなベンチマークで複雑な推論手法よりも一貫して優れているという驚くべき矛盾を明らかにしました。最小限のエキスパートおよび帰納的ロール フレーミング (CoT-Expert および CoT-Inductive) のみが、ベースラインに対して小さいながらも統計的に有意な $\sim$3 パーセンテージ ポイント (pp) の向上をもたらしますが、テストした他のすべての精巧なテクニックは、多くの場合、大きなマージンでそれに匹敵するか、パフォーマンスを下回っています (自己類推の場合は最大 31~pp)。さらに、3 つの重要な現象を調査します。(1) Elo 評価における Qwen3-30B-A3B-Thinking-2507 の予期せぬ勝利、(2) 異なる思考予算を持つモデル バリアント間でのパフォーマンスと効率のトレードオフ、モデル依存の最適な構成が明らかに、(3) データセットの難易度には大幅な変動があり、ベンチマークの 60% が 70% を下回っており、最も簡単なものから最も難しいものまで 47.5 pp の広がりがあり、かなりの余地があることが示されています。モデル改良のため。これらの結果は、LLM 評価コミュニティがプロンプト エンジニアリングを複雑にしすぎている可能性があり、さまざまなベンチマーク間で大幅なパフォーマンスのギャップが残っており、プロンプトの最適化ではなく真のモデル改善の機会を提供していることを示唆しています。

原文 (English)

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding. Furthermore, the rapid proliferation of LLMs has created the implicit assumption that more sophisticated prompting techniques yield better performance. Several studies claim better performance with more sophisticated prompting techniques, but do not provide a comprehensive evaluation. We address this gap through a comprehensive empirical study of 8 prompting techniques across 10 multiple-choice question answering (MCQA) datasets, encompassing 27 model configurations and roughly 4,300 unique questions evaluated more than 430,000 times. Our findings reveal a striking paradox that baseline prompting consistently outperforms complex reasoning techniques on various benchmarks. Only minimal expert and inductive role framing (CoT-Expert and CoT-Inductive) yields a small but statistically significant $\sim$3 percentage-point (pp) gain over baseline whereas every other elaborate technique we tested matches or under-performs it, often by large margins (up to 31~pp for Self-Analogical). We further investigate three critical phenomena: (1) the unexpected victory of Qwen3-30B-A3B-Thinking-2507 in Elo ratings, (2) the performance-efficiency trade-offs across model variants with different thinking budgets, revealing model-dependent optimal configurations, and (3) the substantial variation in dataset difficulty, with 60% of benchmarks below 70% accuracy and a 47.5~pp spread from easiest to hardest, indicating considerable room for model improvement. These results suggest that the LLM evaluation community may be overcomplicating prompt engineering and that substantial performance gaps remain across diverse benchmarks, offering opportunities for genuine model improvements rather than prompt optimization.

2026-07-17 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

動的なヒューマノイド全身制御のためのセマンティックなオーディオ駆動型の理解

近年のヒューマノイドロボット工学と強化学習の進歩により、表現力の高い全身運動ポリシーの獲得が可能になりました。しかし、ロボットのパフォーマンスのほとんどは、事前にスクリプト化されたシーケンスまたは外部からトリガーされた動作に基づいたままであり、動的環境に対する自律性や応答性が制限されています。この研究では、セマンティックなオーディオ駆動型ヒューマノイド制御のための新しいマルチモーダル オーケストレーション フレームワークを導入し、ロボットが適切なモーション スキルをリアルタイムで自律的に選択して実行できるようにします。システムは連続オーディオ ストリームを処理し、それらを音楽または音声ブランチにルーティングします。音楽入力は、オーディオ フィンガープリンティングとセマンティック エンベディングを介して処理され、トラックのアイデンティティと時間的アライメントを取得し、音楽セグメントとモーション ポリシー間の動的なマッピングを可能にします。音声入力は、模倣によって学習されたスキルの個別のライブラリに統合され、人間とロボットの直接的な対話が可能になります。どちらのモダリティも、強化学習制御パイプラインを介してスキルの実行をスケジュールする統合インターフェイスを共有します。このアプローチをシミュレーションと Unitree G1 ヒューマノイドで検証し、堅牢なシミュレーションからリアルへの転送と一貫したオーディオ条件付きポリシー選択を示します。補足資料は次のサイトで入手できます: https://lab-rococo-sapienza.github.io/semantic-WBC/

原文 (English)

Semantic Audio-driven Understanding for Dynamic Humanoid Whole Body Control

Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policies. However, most robotic performances remain based on pre-scripted sequences or externally triggered behaviors, limiting autonomy and responsiveness to dynamic environments. In this work, we introduce a novel multi-modal orchestration framework for semantic audio-driven humanoid control, enabling robots to autonomously select and execute appropriate motion skills in real time. The system processes continuous audio streams and routes them into music or speech branches. Music input is handled via audio fingerprinting and semantic embeddings to retrieve track identity and temporal alignment, allowing dynamic mapping between musical segments and motion policies. Speech input is grounded into a discrete library of imitation-learned skills, enabling direct human-robot interaction. Both modalities share a unified interface that schedules skill execution over a reinforcement learning control pipeline. We validate the approach in simulation and on a Unitree G1 humanoid, showing robust sim-to-real transfer and consistent audio-conditioned policy selection. Supplementary materials are available at the following site: https://lab-rococo-sapienza.github.io/semantic-WBC/

2026-07-17 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

Instant NuRec: 運転シーン シミュレーションのためのフィードフォワード 3D ガウス再構成

3D シミュレーション プラットフォームは、エンドツーエンドのポリシー評価を可能にし、それによって開発コストを削減し、安全性を向上させるため、自動運転にとって不可欠です。近年ではニューラルシミュレーションが主流となり、NuRecなどの手法が中心的な役割を果たしています。ただし、これらの方法は依然として比較的遅いため、通常はシーンごとの調整が必要です。この研究では、短いマルチビュー運転ログを 1 回のフォワード パスで完全にシミュレーション可能な 3D ガウス スプラッティング (3DGS) ワールドに変換するフィードフォワード ニューラル再構成モデ​​ルである Instant NuRec を紹介します。このモデルは、キャリブレーションされたカメラ リグからのマルチビュー入力を受け入れ、静的および動的 3DGS レイヤー、スカイ キューブマップ、およびカメラごとの ISP 補正で構成されるレイヤー出力を出力すると同時に、3DGUT を介して非ピンホール カメラ モデルのネイティブ サポートを提供します。 10 ~ 20 秒のマルチカメラ シーンを約 1.5 秒で再構築し、Waymo Open Dataset 上で最も強い評価ベースラインを 2.01 dB 上回る PSNR を達成します。 Instant NuRec は NuRec に深く統合されており、閉ループ シミュレーション用の AlpaSim と互換性があります。

原文 (English)

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

3D simulation platforms are critical for autonomous driving because they enable end-to-end policy evaluation, thereby reducing development costs and improving safety. In recent years, neural simulation has become predominant, with methods such as NuRec playing a central role; however, these methods remain relatively slow and typically require per-scene tuning. In this work, we present Instant NuRec, a feed-forward neural reconstruction model that turns a short multi-view driving log into a fully simulatable 3D Gaussian Splatting (3DGS) world in a single forward pass. The model accepts multi-view input from a calibrated camera rig and emits a layered output consisting of static and dynamic 3DGS layers, a sky cubemap, and per-camera ISP corrections, while providing native support for non-pinhole camera models via 3DGUT. It reconstructs a 10-20-second multi-camera scene in roughly 1.5 seconds and achieves a PSNR on the Waymo Open Dataset that is 2.01 dB above the strongest evaluated baseline. Instant NuRec is deeply integrated into NuRec and is compatible with AlpaSim for closed-loop simulation.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

コピーオンライトのスコアリング: アプリケーション固有のエージェントの評価

ソフトウェア システムに LLM ベースのエージェントを信頼できる展開するには、エージェントがアプリケーション固有のワークフローでどのように実行されるかを、成功と失敗の場所を特定するのに十分な粒度で評価する必要があります。しかし、既存のエージェント評価メカニズムには限界があります。ベンチマークはアプリケーション固有のワークフローや環境に対する構成の妥当性が低く、レプリカ評価環境は高価でドリフトしやすいです。私たちは、エージェントの書き込みを分離するために PostgreSQL レベルのコピーオンライト メカニズムを使用して、アプリケーション環境内でエージェントの操作を直接評価するフレームワークであるコピーオンライト (CoW) スコアリングを提案します。 CoW スコアリングは、特定のアプリケーション環境でエージェントのデータベース書き込み操作が成功した場所と失敗した場所を強調表示するセッション レベルおよび操作レベルのスコアを生成し、エージェント ハーネスとツール サーフェスでの低コストの評価と反復を可能にします。オープンソースのプロジェクト管理プラットフォームである Plane でフレームワークをデモンストレーションします。分析によりツール表面の特定の問題が明らかになり、対応する修正により影響を受けるモデルに目に見える改善がもたらされました。 Python ライブラリ: https://github.com/trail-ml/agent-cow-python

原文 (English)

Copy-on-Write Scoring: Application-Specific Agent Evaluations

Trustworthy deployment of LLM-based agents in software systems requires evaluating how they perform on application-specific workflows, with enough granularity to localize where they succeed and fail. Yet existing agent evaluation mechanisms are limited: benchmarks have low construct validity for application-specific workflows and environments, and replica evaluation environments are expensive and prone to drift. We propose Copy-on-Write (CoW) Scoring, a framework that evaluates agent operations directly within application environments using a PostgreSQL-level Copy-on-Write mechanism to isolate agent writes. CoW Scoring produces session- and operation-level scores that highlight where agents' database write operations succeed and fail in a given application environment, enabling inexpensive evaluation and iteration on agent harnesses and tool surfaces. We demonstrate the framework on Plane, an open-source project-management platform, where analysis surfaced specific issues in the tool surface, and corresponding fixes produced measurable improvements on affected models. Python library: https://github.com/trail-ml/agent-cow-python

2026-07-17 13:00 JSTarXiv cs.AIビジネス/資金調達

認識論的不確実性の評価: OOD 検出とアクティブ ラーニングを超えて

認識論的不確実性の現在の評価は、分布外検出や能動学習などのタスクに依存しています。ただし、これらのタスクのベイズ最適決定戦略は、認識論的不確実性を定量化するために一般的に使用されるスコアと一致しません。認識論的拒否オプションのフレームワークに基づいて、我々は、後悔、つまり削減可能な誤差を特定する能力を使用して認識論的不確実性を評価します。カバレッジ、予想されるリスク、リグレスに対する制約付き最適化として選択的予測を定式化し、最適なセレクターがグラウンドトゥルースの偶然性と認識論的不確実性の閾値付き凸組み合わせであることを証明します。この理論的統一は、最近の不確実性解きほぐしに関する文献の弱点を明らかにしています。学習されたコンポーネント間の標準的な相関指標が、実際の運用上の有用性を必ずしも予測するとは限らないことを示しています。代わりに、関節のもつれの解消と有用性の診断として、達成可能なリスク、後悔、分解の適用範囲を評価することを提案します。高密度のヒューマン アノテーションを含むデータセットで標準メソッドのベンチマークを行うと、ある基準で上位にランクされ、別の基準で最下位にランクされるメソッド間のペアごとの順位逆転など、意思決定理論によるランキングが代理タスクのランキングと実質的に一致しない可能性があることが明らかになりました。

原文 (English)

Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning

Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.

2026-07-17 13:00 JSTarXiv cs.AIビジネス/資金調達

AI が貢献の境界を曖昧にするとき: 著者資格の調整に関する実証的研究

人工知能 (AI)、特にジェネレーティブ AI の広範な導入により、ユーザーがこれらのシステムとどのように対話して新しいコンテンツを作成するかについて差し迫った疑問が生じています。この論文では、AI と対話するときにユーザーが実際の著者であることを認識することとして定義される著者資格調整の概念を紹介します。 CoAuthor データセットを使用して、著者資格の調整がユーザー間でどのように異なるか、またそれが AI の使用頻度とどのように関連しているかを実証的に調査します。私たちの結果は、ばらつきが大きいことを明らかにしました。AI に大きく依存しているユーザーは、自分の著者であるかどうかを誤って判断する傾向があるのに対し、AI をそれほど頻繁に使用していないユーザーは、より正確な著者であるかどうかの調整を示しています。これらの発見は、AI がユーザー自身の作者に対する認識を曖昧にする可能性があることを示唆しています。学習の状況において、誤った調整はメタ認知のモニタリングと学習戦略に影響を与え、最終的には学習成果に影響を与える可能性があります。したがって、責任ある教育的に有意義な AI 統合を促進するには、著者資格の調整を促進することが不可欠であると考えられます。

原文 (English)

When AI Blurs the Boundaries of Contribution: An Empirical Study of Authorship Calibration

The broad adoption of Artificial Intelligence (AI), especially Generative AI, raises pressing questions about how users interact with these systems to produce new content. In this paper, we introduce the concept of authorship calibration, defined as users awareness of their actual authorship when interacting with AI. Using the CoAuthor dataset, we empirically examine how authorship calibration varies across users and how it relates to their frequency of AI use. Our results reveal high variability: users relying heavily on AI tend to misjudge their authorship, whereas those using AI less frequently exhibit more accurate authorship calibration. These findings suggest that AI can obscure users perception of their own authorship. In learning contexts, miscalibration can affect metacognitive monitoring and learning strategies, ultimately impacting learning outcomes. Fostering authorship calibration then appears essential for promoting responsible and educationally meaningful AI integration.

2026-07-17 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery…

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit

Large language models are increasingly deployed for legal question answering, where evaluations typically focus on multiple-choice accuracy…

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

推論によるフロンティア LLM 評価の形状計算方法

AI の評価は、ツールの使用と反復的な問題解決を伴う長期にわたる軌道から恩恵を受ける、より困難なタスクへと移行しています。その結果、パフォーマンスは、テスト時に利用可能なコンピューティング (「推論コンピューティング」) の量と割り当てにますます敏感になります。しかし、多くの評価では依然として単一の制限された予算でのパフォーマンスが報告されており、低いスコアはモデルの基礎的な機能ではなく評価設定を反映している可能性があることを意味します。これをテストするために、ソフトウェア エンジニアリング、数学、医学、サイバーセキュリティにわたる 7 つの挑戦的なベンチマークで最大 12 のフロンティア言語モデルを評価します。私たちは、3 つの単純な推論スケーリング介入を組み合わせた制御されたセットアップを使用します。つまり、より大きなトークン バジェット、コンテキストの圧縮、およびモデル自体または最小限の正確性フィードバックによって導かれる送信の試行の繰り返しです。主な結果は 3 つあります。まず、トークン バジェットが大きくなると、サイバーセキュリティ、FrontierMath、人類最後の試験、ターミナルベンチなど、複数のドメインにわたるベンチマークのパフォーマンスが大幅に向上します。第二に、固定予算の評価では、モデルが進歩するにつれてフロンティアの能力がますます過小評価される可能性があります。新しいモデルは、大きな予算でより高いパフォーマンスを実現し、より困難なタスクを解放し、より確実に解決します。第三に、どの推論スケーリング手法が最も役立つかがベンチマークによって異なります。繰り返し送信するとパフォーマンスが大幅に向上しますが、より大きなトークン バジェット、外部フィードバック、および並列試行の値はベンチマークによって異なります。全体として、私たちの結果は、ベンチマーク スコアがプロトコルに依存していることを示しています。したがって、評価では、特に安全性またはポリシー関連の設定において、推論時間のコンピューティングの関数として機能を報告し、プロトコルの選択を明示的に指定し、一致した予算で大規模な共有コンピューティング範囲にわたってモデルの世代を比較する必要があると主張します。

原文 (English)

How Inference Compute Shapes Frontier LLM Evaluation

AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still report performance at a single restrictive budget, meaning that low scores may reflect the evaluation setup rather than the model's underlying capability. To test this, we evaluate up to 12 frontier language models on seven challenging benchmarks spanning software engineering, mathematics, medicine, and cybersecurity. We use a controlled setup combining three simple inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts, guided either by the model itself or by minimal correctness feedback. We find three main results. First, larger token budgets substantially improve performance on benchmarks across multiple domains, including cybersecurity, FrontierMath, Humanity's Last Exam, and TerminalBench. Second, fixed-budget evaluations can increasingly understate frontier capability as models advance. Newer models reach higher performance at large budgets, where they unlock harder tasks and solve them more reliably. Third, benchmarks differ in which inference-scaling methods help most: repeated submission broadly improves performance, but the value of larger token budgets, external feedback, and parallel attempts varies by benchmark. Overall, our results show that benchmark scores are protocol-dependent. We therefore argue that evaluations should report capability as a function of inference-time compute, specify protocol choices explicitly, and compare model generations over a large shared compute range at matched budgets, especially in safety- or policy-relevant settings.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mechanistic interpretability, and governance/auditability, covering 2018-2026 evaluation-safety measurement work. We introduce EvalSafetyGap as an organizing hypothesis for comparing evaluation-side and alignment-side proxy failures under optimization pressure, using Goodhart's Law together with two constructs we develop here - an Instability Decomposition and an Alignment Trilemma - as tools for generating testable comparisons. The audit shows how conclusions shift when capability, behavioral safety, and governance are measured separately. In this sample ($n = 10$), the association between capability and sustained adversarial robustness is statistically indeterminate using the displayed Table 3 inputs (Pearson $r = +0.232$, $p = 0.520$), and the apparent open-closed safety gap is modest, driven mainly by governance and disclosure rather than behavioral robustness, and sensitive to how a single borderline model is classified; attempt-budget results are protocol dependent. Because the public evidence uses heterogeneous protocols, the audit is diagnostic rather than rank-generating. The contribution is a shared vocabulary and evidence map to support dynamic evaluation, transparent source reporting, multi-attempt safety measurement, and auditable alignment practice.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

隠れたフットプリント: ストレージを LLM エージェント評価の第一級の指標にする

LLM エージェントのベンチマークは、タスクの完了、信頼性、推論コストを測定しますが、ログ、コンテキスト スナップショット、チェックポイント、デバッグ トレースなど、エージェントの実行によってディスクに残される永続データは測定しません。実行後のエージェント ストレージ フットプリントのクロスフレームワーク ベンチマークである AgentFootprint を紹介します。そのシリアル化対応メトリクス スイートは、総保持率、チャネル構成、重複、増加、圧縮率、会話履歴の再構築可能性を測定します。これは、測定の罠に対処します。単純なバイトレベルの測定では、データベースのページングと JSON エスケープが繰り返されるコンテンツを不明瞭にするため、重複が桁違いに過小評価されます。固定トレース制御により、エージェントが生成した論理ボリュームが永続層の増幅から分離されます。7 つの永続フレームワークを通じて同じ軌跡を再生すると、6.7 倍の広がりが得られます。同一のモデル、ツール、およびタスクでは、100% の精度の構成では、デフォルトでサポートされる回復機能と監査機能が異なりますが、保持バイト数が 15.7 倍異なります。 3 つの完全な履歴構成は、反復観察ストレス タスクで超線形に成長します。 108 個のインスタンスで正規化された SWE ベンチからエクスポートされた軌跡 検証済みの送信は、インスタンスごとに 3 桁の大きさに及び、解決率との検出可能な相関関係はありません。コンテンツ アドレス ストアは、すべての再構築可能性スコアを維持しながら、保持率を 4.8 倍から 32.7 倍まで削減します。これらの結果は、精度と再構築可能性を併せてレポートするためのリソース メトリックとして永続ストレージを確立します。

原文 (English)

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite measures total retention, channel composition, duplication, growth, compressibility, and conversation-history reconstructability. It addresses a measurement trap: naive byte-level measurement understates duplication by an order of magnitude because database paging and JSON escaping obscure repeated content. A fixed-trace control separates agent-generated logical volume from persistence-layer amplification: replaying the same trajectory through seven persisting frameworks yields a 6.7x spread. Under identical models, tools, and tasks, configurations with 100% accuracy differ by 15.7x in retained bytes, although their defaults support different recovery and audit capabilities. Three full-history configurations grow superlinearly on a repeated-observation stress task. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude per instance, with no detectable correlation with resolve rate. A content-addressed store reduces retention by 4.8x-32.7x while preserving every reconstructability score. These results establish persistent storage as a resource metric to report jointly with accuracy and reconstructability.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AgentCompass: エージェント機能の統合評価インフラストラクチャ

大規模言語モデル (LLM) が自律エージェントに進化するにつれて、統合された評価インフラストラクチャの必要性が重要になります。ただし、現在の評価パイプラインは高度に断片化され、密接に結合されたままであるため、再現性が妨げられ、冗長なエンジニアリングが発生します。これに対処するために、LLM ベースのエージェントを評価するためのオープンソースで軽量かつ拡張可能なインフラストラクチャである AgentCompass を導入します。 AgentCompass は、ベンチマーク、ハーネス、環境という 3 つの独立したコンポーネントを中心に評価プロセスを編成するため、複雑な実行ロジックを再実装することなく柔軟な構成が可能になります。さらに、フォールトトレラントな非同期ランタイムと、報酬ハッキングなどの微妙な障害モードを透過的に診断するための包括的な軌跡分析ツールを備えています。 AgentCompass は、5 つの機能次元にわたる 20 以上のベンチマークをネイティブにサポートし、エージェント研究を進めるためのスケーラブルで再現可能なインフラストラクチャをコミュニティに提供します。

原文 (English)

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.

2026-07-17 13:00 JSTarXiv cs.AIビジネス/資金調達

Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis

Music Structure Analysis (MSA) aims to uncover the high-level organization of musical pieces. State-of-the-art methods are often based on s…

2026-07-17 13:00 JSTarXiv cs.AIビジネス/資金調達

Warning labels shift perceptions of sycophantic AI, but not its influence

Recent work has raised concerns about the influence of sycophantic AI on user judgment and relationships. One proposed mitigation, which ha…

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

AI の安全性評価のための敵対的プラグマティクス: 命令の競合、埋め込みコマンド、およびポリシーの曖昧さのベンチマーク

言語モデルの安全性評価は、モデルが指示に従ったか、適切に拒否したか、ポリシーに従ったか、埋め込まれたコマンドに抵抗したか、エージェントタスクの進捗状況を誤って報告したかなど、あいまいな自然言語の動作に関する判断にますます依存しています。既存のベンチマークは、多くの場合、これらの区別を合格/不合格のラベルに圧縮し、障害が機能制限、ポリシーの曖昧さ、命令の競合、足場の障害、または不安定な評価者の判断に起因するかどうかを曖昧にします。この論文では、命令の競合、埋め込みコマンド、引用、範囲の曖昧さ、明確化、間接音声行為、およびマルチターンエージェントのトランスクリプトの下でのモデルの動作を評価するためのベンチマークおよびアノテーションプロトコルとして、敵対的プラグマティクスを紹介します。この貢献は経験的かつ方法論的です。言語的に管理された分類法、バリデーターが強制するメタデータを備えた 18 項目のシード ベンチマーク、54 行のローカル シード パイロット、タスクの成功、ポリシー遵守、安全性リスク、拒否結果、評価者の信頼性を区別する専門家評価プロトコル、および裁判官の妥当性、診断の曖昧さ、および分類のドリフトの指標です。このフレームワークは、言語的判断方法論を、安全性評価、LLM ジャッジ、ゴールドセット構造、即時注入テスト、および安全性文書を検証するための実用的なツールに変えます。

原文 (English)

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological: a linguistically controlled taxonomy, an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, an expert-evaluation protocol distinguishing task success, policy compliance, safety risk, refusal outcome, and evaluator confidence, and metrics for judge validity, diagnostic ambiguity, and taxonomy drift. The benchmark treats labels as inference licenses: it tests whether safety-relevant categories project across paraphrase, wrapper, model, and judge condition. In the pilot, a rubric-aided LLM judge graded its own outputs with expected-behaviour fields visible and still missed the safety-relevant minority classes.

2026-07-17 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体ビジネス/資金調達

ブラックボックスを制御できるか?協調エージェントを使用したレコメンダーシステムの制御性中心の評価に向けて

レコメンダー システムはブラック ボックスとして動作するため、ユーザーや規制当局は出力を特定の意図に向けたり、その動作を監査したりすることができません。この制御性の欠如は、明示的なガイダンスに応答するシステムの能力として定義されますが、既存の評価パラダイムでは依然として対処されていない側面です。このギャップを埋めるために、制御性を体系的に評価するための協調的なマルチエージェント フレームワークである CtrlBench-Rec を提案します。私たちは、ターゲット コンテンツの発見、関心プロファイルの形成、人気バイアスの軽減という 3 つの基本的なタスクを形式化します。これらは、明示的なコマンドから暗黙的な表現のステアリング、そして最終的にはアルゴリズムのバイアスの克服までのステアビリティを一緒に測定します。実世界のデータセットと複数のレコメンデーション モデルに関する広範な実験により、私たちのフレームワークが制御性を効果的に定量化し、重大なシステムのボトルネック、特に誘導ロングテール コンテンツに対する永続的な抵抗を明らかにすることが実証されています。 CtrlBench-Rec は、制御可能な推奨調査、アルゴリズム監査、およびユーザー権限付与のための初の標準化されたツールキットを提供します。私たちのコードは https://github.com/caskcsg/CtrlBenchRec でリリースされています。

原文 (English)

Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents

Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, remains an unaddressed dimension in existing evaluation paradigms. To fill this gap, we propose CtrlBench-Rec, a collaborative multi-agent framework for systematic assessment of controllability. We formalize three fundamental tasks: target content discovery, interest profile shaping, and popularity bias mitigation, which together measure steerability from explicit commands to implicit representation steering and finally to overcoming algorithmic biases.Extensive experiments on real-world datasets and multiple recommendation models demonstrate that our framework effectively quantifies controllability and exposes critical system bottlenecks, most notably persistent resistance to guiding long tail content. CtrlBench-Rec provides the first standardized toolkit for controllable recommendation research, algorithmic auditing, and user empowerment. Our code is released on https://github.com/caskcsg/CtrlBenchRec.

2026-07-17 00:02 JSTTechCrunch AIビジネス/資金調達研究/論文

How a former DeepMind researcher raised at a $300M pre-seed valuation before launching a product

Drawing on more than a decade spent helping build some of the world's most influential AI systems, including research that later informed t…

2026-07-16 13:00 JSTTechCrunch AIビジネス/資金調達

Applied Computing wants to give oil and gas operators an AI model for the entire plant

Applied Computing has raised a $20M Series A to build a foundation AI model for the oil, gas and petrochemical industry.

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AgentCompass: エージェント機能の統合評価インフラストラクチャ

大規模言語モデル (LLM) が自律エージェントに進化するにつれて、統合された評価インフラストラクチャの必要性が重要になります。ただし、現在の評価パイプラインは高度に断片化され、密接に結合されたままであるため、再現性が妨げられ、冗長なエンジニアリングが発生します。これに対処するために、LLM ベースのエージェントを評価するためのオープンソースで軽量かつ拡張可能なインフラストラクチャである AgentCompass を導入します。 AgentCompass は、ベンチマーク、ハーネス、環境という 3 つの独立したコンポーネントを中心に評価プロセスを編成するため、複雑な実行ロジックを再実装することなく柔軟な構成が可能になります。さらに、フォールトトレラントな非同期ランタイムと、報酬ハッキングなどの微妙な障害モードを透過的に診断するための包括的な軌跡分析ツールを備えています。 AgentCompass は、5 つの機能次元にわたる 20 以上のベンチマークをネイティブにサポートし、エージェント研究を進めるためのスケーラブルで再現可能なインフラストラクチャをコミュニティに提供します。

原文 (English)

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering. To address this, we introduce AgentCompass, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents. AgentCompass organizes the evaluation process around three independent components, namely Benchmark, Harness, and Environment, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic. Furthermore, it features a fault-tolerant asynchronous runtime and comprehensive trajectory analysis tools to transparently diagnose nuanced failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides the community with a scalable and reproducible infrastructure for advancing agent research.

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Agent Optimizer は複合化しますか? Terminal-Bench 2.0 での継続学習評価

エージェント最適化手法から報告される利益のほとんどは単発的なものです。つまり、エージェントは固定ベンチマークに対して最適化され、その結果得られる改善は、あたかも手法の安定した特性であるかのように報告されます。これでは、展開されたエージェントにとって重要な設定はテストされません。時間の経過とともに新しい障害や新しいタスクが発生すると、最適化が再帰的に適用されます。これが引き起こす中心的な疑問は、オプティマイザ主導のゲインが複合化するかどうかです。エージェントが一度最適化された後、最初のラウンドで生成されたゲインを損なうことなく、新しく到着したタスクで再度最適化できるでしょうか?私たちは、ターミナルベンチ 2.0 のハード タスクから構築された 2 フェーズの継続学習評価を使用してこの問題を研究し、同一の最適化予算の下でエージェント ハーネス最適化への 3 つのアプローチ (GEPA、メタ ハーネス、および RELAI の検証可能な継続学習、RELAI-VCL) を比較します。 3 つの方法はすべて、従来の静的な単相設定のベースライン エージェントよりも改善されています。しかし、新しいタスクが導入されると、方法は大きく異なります。GEPA の最適化されたエージェントは、最適化されていないベースラインを下回って移行します。Meta Harness は良好に移行しますが、2 番目の最適化予算を与えるとさらに改善できません。RELAI-VCL は、両方とも目に見えないタスクに積極的に移行し、それらのタスクが最適化目標に組み込まれた後も改善を続ける唯一の方法であり、すべての評価段階で最高の合格率に達し、全体で最高の生涯平均合格率 (76.4% 対 76.4%) に達します。 GEPA では 66.0%、メタ ハーネスでは 64.6%、ベースラインでは 58.7%)。私たちの重要な観察は、回帰制御が最適化ループに組み込まれている場合にのみ最適化のゲインが増大し、一般化できないショートカット ソリューションに対する帰納的バイアスが提供されるということでした。

原文 (English)

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimization is applied recursively as new failures and new tasks appear over time. The central question this raises is whether optimizer-driven gains compound: after an agent has been optimized once, can it be optimized again on newly arrived tasks without eroding the gains the first round produced? We study this question with a two-phase continual-learning evaluation built from hard tasks in Terminal-Bench 2.0, comparing three approaches to agent-harness optimization (GEPA, Meta Harness, and RELAI's Verifiable Continual Learning, RELAI-VCL) under identical optimization budgets. All three methods improve over the baseline agent in the conventional, static, single-phase setting. However, once new tasks are introduced, the methods diverge sharply: GEPA's optimized agent transfers below the unoptimized baseline, Meta Harness transfers well but fails to improve further once given a second optimization budget, and RELAI-VCL is the only method that both transfers positively to unseen tasks and continues improving after those tasks are folded into the optimization objective, reaching the highest pass rate at every evaluated stage and the highest lifelong average pass rate overall (76.4% vs. 66.0% for GEPA, 64.6% for Meta Harness, and 58.7% for the baseline). Our key observation was that optimization gains compounded only when regression control was built into the optimization loop, providing an inductive bias against shortcut solutions that fail to generalize.

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

診断する前に尋ねてください: Safe-Psych、精神科における LLM の逐次評価ベンチマーク

大規模言語モデル (LLM) は医療分野での意思決定支援に使用されることが増えていますが、臨床証拠は不完全であるか進化していることがよくあります。入手可能な情報が信頼できる回答を裏付けるには不十分な場合、モデルは裏付けのない回答を提供するのではなく、説明を要求するか棄権する必要があります。ただし、既存の医療ベンチマークは通常、完全な情報が事前に入手できることを前提としています。臨床精神医学において、LLM が進化する診断の不確実性にどのように対処するかを評価するための逐次ベンチマークである Safe-Psych を紹介します。 Safe-Psych には、段階的な証拠開示をシミュレートするためにセグメント化された 1,000 件を超える実際の精神医学の臨床ノートが含まれており、各段階で精神科医が導き出したアクション ラベル (診断、明確化、または棄却) が付いています。当社は、複数の最先端の LLM を完全な情報と順次設定で評価します。私たちの調査結果は、能力がキャリブレーションを保証するものではないことを示しています。不完全な臨床情報の下では、強力なモデルでも苦戦しており、ほとんどのモデルで不完全な棄権が 60% を超えており、安全性を意識することで、エラーを過剰な棄権にシフトすることによってのみ、時期尚早のコミットメントを減らすことができます。逐次評価では、モデルは十分な証拠が得られる前に診断することが多く、明示的に指示されない限り明確化を求めることはほとんどありません。これらの時期尚早の診断は、予定通りの診断よりも精度が低くなります。全体として、Safe-Psych では、臨床証拠が不完全で追加情報が必要な場合の認識という、評価されたモデル全体に​​わたる限界が明らかになりました。私たちは、医療における LLM の安全性を向上させる研究をサポートするために Safe-Psych をリリースします。

原文 (English)

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the available information is insufficient to support a reliable answer, models should request clarification or abstain rather than provide unsupported responses. Existing medical benchmarks, however, typically assume that complete information is available upfront. We introduce Safe-Psych, a sequential benchmark for evaluating how LLMs handle evolving diagnostic uncertainty in clinical psychiatry. Safe-Psych contains over 1,000 real-world psychiatric clinical notes segmented to simulate incremental evidence disclosure, with psychiatrist-derived action labels at each stage: DIAGNOSE, CLARIFY, or ABSTAIN. We evaluate multiple state-of-the-art LLMs in full-information and sequential settings. Our findings show that capability does not ensure calibration: even strong models struggle under incomplete clinical information, with under-abstention exceeding 60% for most models and safety-aware prompting reducing premature commitment only by shifting errors toward excessive abstention. In sequential evaluation, models frequently diagnose before sufficient evidence is available and rarely seek clarification unless explicitly prompted; these premature diagnoses are less accurate than on-time diagnoses. Overall, Safe-Psych reveals a limitation across the evaluated models: recognizing when clinical evidence is incomplete and additional information is needed. We release Safe-Psych to support research on improving LLM safety in healthcare.

2026-07-16 13:00 JSTarXiv cs.AIビジネス/資金調達

セーフガード条件付き上昇率: デュアルユースの生物学助手のユーティリティとリスクのフロンティアを測定する

デュアルユースの生物学アシスタントの安全性評価では、多くの場合、基本モデルの機能、拒否行動、またはジェイルブレイクの成功を測定します。これらのメトリックは、導入に関する質問を見逃しています。つまり、固定基本モデルの場合、ユーザーが実際に目にするアクセス条件は、無害なユーティリティと有害な実用的な支援をどのように変更するのでしょうか?私は、人間が判断したユーティリティとリスクのフロンティアを通じて展開されたアクセス条件を比較するためのプロトコルであるセーフガード条件付きアップリフトを紹介します。私は、Claude Sonnet 4.6 と Gemini 3.5 Flash を、役立つプロンプト、安全なプロンプト、および安全に保護された外部アシスタントの下で、108 タスクのサロゲート ベンチマークで評価しました。ヘッドラインの主張は、ロックされた 18 タスクのホールドアウト スプリットに限定されています。 600 行の盲検化された人間による監査では、保護されたアシスタントは、ブートストラップ 95% 間隔 [-0.117, -0.011] で、49 の一致した応答ペアにわたって、有益なプロンプトと比較して有害なアクション可能性を -0.063 減少させますが、正確性は間隔 [-0.057, +0.077] で +0.009 変化します。アダプティブ、テスト B、キューアブレーション、およびコントローラーベースラインのチェックは、測定ストーリーをサポートしますが、非優位性も示しています。多くの場合、安全プロンプトはクロードにとって最も強力ですが、外部制御はジェミニにとってより役立ち、良性の効用を減らす可能性があります。この貢献は普遍的な防御策ではありません。これは、導入レベルの評価目標に加えて、ユーザーが直面するアクセス条件が公益事業のリスクフロンティアをどのように動かすかを測定するための、学習されたリスク予算調整手順です。

原文 (English)

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success. These metrics miss a deployment question: for a fixed base model, how does the access condition users actually see change benign utility and harmful actionable assistance? I introduce safeguard-conditioned uplift, a protocol for comparing deployed access conditions through a human-judged utility-risk frontier. I evaluate Claude Sonnet 4.6 and Gemini 3.5 Flash under helpful prompting, safety prompting, and an external safeguarded assistant on a 108-task surrogate benchmark, with the headline claim restricted to a locked 18-task held-out split. In a 600-row blinded human audit, the safeguarded assistant reduces harmful actionability relative to helpful prompting by -0.063 over 49 matched response pairs, with bootstrap 95% interval [-0.117, -0.011], while correctness changes by +0.009 with interval [-0.057, +0.077]. Adaptive, Test-B, cue-ablation, and controller-baseline checks support the measurement story but also show non-dominance: safety prompting is often strongest for Claude, while external control helps more for Gemini and can reduce benign utility. The contribution is not a universal defense. It is a deployment-level evaluation target, plus a learned risk-budgeted calibration procedure, for measuring how user-facing access conditions move the utility-risk frontier.

2026-07-16 13:00 JSTarXiv cs.AIビジネス/資金調達

フェデレーション型説明可能な人工知能: 役割、アーキテクチャ、評価、未解決の課題

フェデレーテッド ラーニング (FL) は、分散された異種データ ソース間でプライバシーを保護しながら共同モデルをトレーニングするための重要なパラダイムとして登場しました。 FL は生データをローカルに保持することでデータの機密性の問題に対処していますが、最新の機械学習モデルの不透明性は解決されていません。並行して、Explainable Artificial Intelligence (XAI) は、特に一か八かの分野において、透明性、信頼、説明責任を向上させるために注目を集めています。それらの交差点により、プライバシーと説明可能性の要件を共同で満たすことを目的とした Federated Explainable Artificial Intelligence (FedXAI) パラダイムが生まれました。この調査は、FedXAI の体系的なレビューを提供し、事後ツールから FL ライフサイクルの不可欠なコンポーネントへの説明可能性の移行に焦点を当てています。説明可能性が集約、パーソナライゼーション、堅牢性、調整、およびシステムレベルの意思決定をどのようにサポートするかを示します。文献を整理するために、説明可能性の役割、モデルと説明者のタイプ、説明範囲、統合レベル、FL 設定、およびデータの異質性によって FedXAI メソッドを分類する分類法を導入します。モデルに依存しない説明から、解釈可能な連合モデル、説明可能性を意識した集計メカニズムに至るまでのアプローチをレビューします。また、評価の実践を検証し、説明の品質、安定性、プライバシー漏洩、計算オーバーヘッドを測定するための標準化されたベンチマークや指標の欠如についても議論します。最後に、非 IID データでの説明可能性、説明中心のセキュリティ脅威、通信効率の高い XAI、継続的な FedXAI、ドメイン知識と規制上の制約の統合など、主要な課題を特定します。既存の作業を統合し、主要なギャップを特定することにより、この調査は、信頼性があり、透明性があり、プライバシーが保護されるフェデレーテッド AI システムを設計するための参照フレームワークとして機能します。

原文 (English)

Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges

Federated Learning (FL) has emerged as a key paradigm for privacy-preserving collaborative model training across distributed and heterogeneous data sources. By keeping raw data local, FL addresses data confidentiality concerns, yet it does not resolve the opacity of modern machine learning models. In parallel, Explainable Artificial Intelligence (XAI) has gained attention for improving transparency, trust, and accountability, particularly in high-stakes domains. Their intersection has given rise to Federated Explainable Artificial Intelligence (FedXAI) paradigm, which aims to jointly satisfy privacy and explainability requirements. This survey provides a systematic review of FedXAI, highlighting the transition of explainability from a post-hoc tool to an integral component of the FL lifecycle. We show how explainability supports aggregation, personalization, robustness, coordination, and system-level decision making. To organize the literature, we introduce a taxonomy that classifies FedXAI methods by the role of explainability, model and explainer types, explanation scope, integration level, FL settings, and data heterogeneity. We review approaches ranging from model-agnostic explanations to interpretable federated models and explainability-aware aggregation mechanisms. We also examine evaluation practices and discuss the lack of standardized benchmarks and metrics for measuring explanation quality, stability, privacy leakage, and computational overhead. Finally, we identify key challenges, including explainability under non-IID data, explanation-centric security threats, communication-efficient XAI, continual FedXAI, and the integration of domain knowledge and regulatory constraints. By consolidating existing work and identifying key gaps, this survey serves as a reference framework for designing trustworthy, transparent, and privacy-preserving federated AI systems.

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

WaterMoE: エキスパート ルーティング ベースの透かしによる高忠実度および効率の向上

大規模言語モデル (LLM) は目覚ましい成功を収めていますが、コンテンツの出所や悪用についての懸念が高まっており、信頼性の高い透かし技術の必要性が高まっています。ただし、これらの手法は、主に 2 つの理由により、実際にはほとんど採用されていません: i) モデルのパフォーマンスが大幅に低下すること、および ii) 推論のオーバーヘッドが追加されることです。この問題を確認するために、さまざまな生成タスクにわたる包括的なベンチマークを構築し、9 つの代表的な透かし手法を系統的に評価します。ほとんどすべての既存のメソッドはテキストの流暢さのために設計されていますが、制限された複雑なタスクには設計されておらず、そのオーバーヘッドにより遅延が重要なシステムへの導入が妨げられていることがわかりました。 i) と ii) に対処するために、人気が高まっている Mixture-of-Experts (MoE) LLM 用の LLM 透かしスキーム \textit{WaterMoE} を提案します。 WaterMoE は、制御された摂動を通じて電子透かし信号を各ルーターのエキスパート選択に埋め込み、最終出力でのトークン選択シフトに蓄積されます。後処理トークン サンプリング アプローチとしてのウォーターマークとは対照的に、WaterMoE は推論ループ内にウォーターマークを埋め込みますが、品質の低下と計算オーバーヘッドは無視できます。広範な実験により、私たちの方法は透かしなしの状態に近い忠実度のパフォーマンスを達成し、ベンチマークで最先端の透かし入れ方法を常に上回っており、ネイティブ生成と比較して推論遅延がわずか 1\% 増加するだけで、最大 $4\time$ の高速化が可能であることが実証されています。結果は、現実世界のタスクに導入できる WaterMoE の機能を示しています。

原文 (English)

WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency

Large language models (LLMs) have achieved remarkable success but raise growing concerns about content provenance and misuse, motivating the need for reliable watermarking techniques. However, these techniques have rarely been adopted in practice mainly for two reasons: i) severely degraded model performance, and ii) additional inference overhead. To confirm the problem, we construct a comprehensive benchmark spanning different generation tasks to systematically evaluate 9 representative watermarking methods. We found almost all existing methods are designed for text fluency, but not for restricted and complicated tasks, and their overhead prevents them from deployment in latency-critical systems. To address i) and ii), we propose an LLM watermarking scheme \textit{WaterMoE} for the growingly popular Mixture-of-Experts (MoE) LLMs. WaterMoE embeds watermarking signals through controlled perturbation into the expert selection at each router, which accumulates to token selection shift at the final output. In contrast to watermarking as a post-processing token-sampling approach, WaterMoE embeds watermark within the inference loop incurring negligible quality degradation and computational overhead. Extensive experiments demonstrate that our method achieves a fidelity performance close to the unwatermarked and consistently outperforms state-of-the-art watermarking methods on the benchmark, with up to $4\times$ speedup, incurring merely 1\% additional inference latency compared to native generation. The results demonstrate the capability of WaterMoE to be deployed in real-world tasks.

2026-07-16 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

進化し続けるディープフェイク検出: 動的検出システムのアーキテクチャと公開ベンチマーク評価

学術的なベンチマークでほぼ完璧なスコアを達成するディープフェイク検出器は、現実世界のコンテンツでは崩壊します。最近の実際の評価では、最先端のオープンソース モデルでは AUC が 45 ~ 50% 低下すると報告されています。私たちは、このギャップは構造的なものであると主張します。静的な検出器は、移動する生成フロンティアに対して一度トレーニングされます。私たちは、トレーニング配布を継続的に更新するオープンな敵対的コンペティションである Bittensor SN34 を通じてトレーニングされた BitMind Forensics (BMF) を紹介します。私たちは、19 の公開データセットにわたる画像、一般ビデオ、人間ビデオのチェックポイントで構成される 1 つの日付付きエクスポートを評価します。正規の顔交換スイート (FaceForensics++、Celeb-DF v1/v2/++、DFDC、DFD、UADFV、DF40)、および最近の野生および AI 生成メディア ベンチマーク (Sumsub、Deepfake-Eval-2024、WildRF、コミュニティ フォレンジック、 AIGCDetectBench、GenImage、AI-GenBench、AIGIBench、RAID、GenVidBench、GenVideo-100K)。 BMF は、Sumsub の元の画像で 0.936 AUC に達し、4 条件操作バッテリー全体 (140 万画像) でプールされた AUC 0.872 に達し、摂動下でも堅牢性を維持します (0.855 JPEG、0.799 ダウンスケール)。一方、GPEN 強化により検出が向上します (0.996)。 Deepfake-Eval-2024 では、画像では最良の商用検出器と一致し (0.915 対 0.90)、ビデオではそれを上回り (0.822 対 0.79)、最良のオープンソース検出器 (0.56 と 0.63) をはるかに上回っています。これは、21 ジェネレーターの AI 画像パネルで 0.991 AUC、GenVidBench で 0.918 に達し、DFDC (0.947 vs 0.843) および Celeb-DF v2 (0.9985 vs 0.956) で FF++ でトレーニングされたフロンティアを超えており、両方とも汚染が監査されており、Celeb-DF++ と統計的に同等です。一時的な研究では、静的ベースラインのトレーニングに参加していないジェネレーターからの保持されたメディアで、連続した日付付きエクスポートが改善されました (画像 0.842 から 0.902、ビデオ 0.864 から 0.936)。私たちの評価ハーネスは公開されており、公開時には実稼働 API が独立した検証のために正確に評価されたスナップショットを提供します。

原文 (English)

Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System

Deepfake detectors that achieve near-perfect scores on academic benchmarks collapse on real-world content: recent in-the-wild evaluations report AUC drops of 45-50% for state-of-the-art open-source models. We argue this gap is structural: static detectors are trained once against a moving generative frontier. We present BitMind Forensics (BMF), trained through Bittensor SN34, an open adversarial competition that continually refreshes the training distribution. We evaluate one dated export comprising image, general-video, and human-video checkpoints across nineteen public datasets: the canonical face-swap suites (FaceForensics++, Celeb-DF v1/v2/++, DFDC, DFD, UADFV, DF40) and recent in-the-wild and AI-generated-media benchmarks (Sumsub, Deepfake-Eval-2024, WildRF, Community Forensics, AIGCDetectBench, GenImage, AI-GenBench, AIGIBench, RAID, GenVidBench, GenVideo-100K). BMF reaches 0.936 AUC on Sumsub's original images and 0.872 pooled AUC over its full four-condition manipulation battery (1.4M images), staying robust under perturbation (0.855 JPEG, 0.799 downscaled), while GPEN enhancement improves detection (0.996). On Deepfake-Eval-2024, it matches the best commercial detector on images (0.915 vs 0.90) and exceeds it on video (0.822 vs 0.79), far above the best open-source detectors (0.56 and 0.63). It reaches 0.991 AUC on a 21-generator AI-image panel and 0.918 on GenVidBench, and exceeds the FF++-trained frontier on DFDC (0.947 vs 0.843) and Celeb-DF v2 (0.9985 vs 0.956), both contamination-audited, with statistical parity on Celeb-DF++. In a temporal study, successive dated exports improve on held-out media from generators absent from the static baseline's training (image 0.842 to 0.902; video 0.864 to 0.936). Our evaluation harness is public, and at publication the production API serves the exact evaluated snapshot for independent verification.

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

評価能力は最適化ユーティリティを意味しない: 閉ループのテーブル認識における LLM-as-a-Judge シグナル

LLM-as-a-judge は、閉ループ再生でフィードバックおよび選択信号を提供するために広く使用されていますが、この使用法はまだ十分に検証されていません。私たちはこれをテーブル認識で研究します。テーブル認識では、FinTabNet と OmniDocBench を使用して、決定論的な TEDS 評価が制御されたテストベッドを提供します。 3 つの発見が得られます。まず、どちらのデータセットでもジャッジシグナルが弱かったです。スコアは同点になることが多く、ランキングは再現性がなく、両方のデータセットでランダムに勝る唯一の選択ポリシーは最も早い反復のタイルールに依存していたため、その利点をジャッジスコアのみに帰することはできません。繰り返しの結果、より良い候補者が誕生しましたが、裁判官は候補者を取り戻すことができませんでした。第二に、具体的なジャッジのフィードバックがなくても重大な損失が発生しました。構造を保持する命令により、FinTabNet での重大損失率が大幅に減少し、OmniDocBench では方向が一貫していました。このコントラストは、観察された深刻な損失の類似メカニズムとして、制約のない再生下でのターゲット保存の失敗を裏付けています。第三に、構造保存制約により重大損失テールは減少しましたが、改善は見られませんでした。探索的な 2x2 分析では、ジャッジのフィードバックが保持されている場合、同じ防御が安定して観察されませんでした。これらの結果は、評価者としての LLM の価値に異議を唱えるものではありません。その代わりに、評価能力は最適化の有用性を意味しないことを示しています。反復改良には、スコアだけを判断するのではなく、構造変化を決定論的に検出する検証信号が少なくとも必要です。

原文 (English)

Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition

LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled testbed, using FinTabNet and OmniDocBench. Three findings emerge. First, judge signals were weak on both datasets: scores frequently tied, rankings were not reproducible, and the only selection policy that beat random on both datasets depended on an earliest-iteration tie rule, so its advantage cannot be attributed to the judge scores alone. Iteration produced better candidates, but the judge failed to recover them. Second, severe losses occurred even without specific judge feedback. A structurepreserving instruction significantly reduced the severe-loss rate on FinTabNet and was directionally consistent on OmniDocBench. The contrasts support target-preservation failure under unconstrained regeneration as a proximate mechanism of the observed severe losses. Third, the structure-preservation constraint reduced the severe-loss tail but produced no improvement. In an exploratory 2x2 analysis, the same protection was not stably observed when judge feedback was retained. These results do not dispute the value of LLMs as evaluators. Instead, they show that evaluation ability does not imply optimization utility. Iterative refinement requires, at minimum, a verification signal that deterministically detects structural change, rather than judge scores alone.

2026-07-16 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

Learning Engagement Assistant (LEA): エージェント型 AI 個別指導システムのコース間の拡張性と教室での評価

このペーパーは、ICAART 2026 カンファレンスで発表されたペーパーの拡張版であり、LEA (Learning Engagement Assistant) を紹介しました。LEA (Learning Engagement Assistant) は、統合されたチャット、チューター、およびクイズ モードにわたる、コース固有の検索拡張生成 (RAG) と構造化ナレッジ コンポーネント (KC) モデルを組み合わせた適応型 AI 個別指導エージェントです。以前の研究では、合成学習者エージェントを使用したシミュレーションのみを介して、単一の STEM コース (CMP511) で LEA を検証しました。このペーパーでは、実際の生徒 (n = 8、CMP511) を対象とした LEA の最初の教室展開と、2 つの学業レベルと 2 つの専門領域にわたる 3 つのコースにシステムを展開して、コースをまたがる拡張性の最初の実証テストを報告することで、その研究を拡張します。この研究では、モード間のシミュレーション予測からの乖離が明らかになり、総合的な評価だけでは実際の展開のすべての側面を予測できないことが示されています。 RAGAS ベースのコース横断スケーラビリティ評価 (660 問) では、回答の関連性とコンテキストの精度がコース全体でほぼ安定している (それぞれ 0.88 ~ 0.94 および 0.88 ~ 0.90) 一方で、システムの元のコースからカリキュラムが離れるにつれて忠実度が低下することがわかりました (0.69 ~ 0.50)。これは、スケーラビリティの制限ではなく、システムの元の主題に合わせて調整された生成ロジックを反映している可能性がある予備的な調査結果です。これらの調査結果は、オーケストレーション層には変更を必要としないが、すべての下流コンポーネントの完全なコース非依存性についてはさらなる調査が必要であることを示唆しています。

原文 (English)

Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System

This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structured Knowledge Component (KC) models across integrated Chat, Tutor, and Quiz modes. That prior work validated LEA on a single STEM course (CMP511) exclusively through simulation, using synthetic learner agents. This paper extends that work by reporting the first classroom deployment of LEA with real students (n = 8, CMP511) and the first empirical test of its cross-course scalability, deploying the system across three courses spanning two academic levels and two disciplinary domains. The study reveals a divergence from simulation predictions across modes, showing that synthetic evaluation alone cannot anticipate all aspects of real deployment. A RAGAS-based cross-course scalability evaluation (660 questions) finds Answer Relevancy and Context Precision broadly stable across courses (0.88-0.94 and 0.88-0.90 respectively), while Faithfulness declines with curriculum distance from the system's original course (0.69 to 0.50), a preliminary finding that may reflect generation logic tuned to the system's original subject rather than a scalability limitation. These findings suggest that while the orchestration layer requires no modification, full course-agnosticism of all downstream components requires further investigation.

2026-07-16 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体ビジネス/資金調達

ブラックボックスを制御できるか?協調エージェントを使用したレコメンダーシステムの制御性中心の評価に向けて

レコメンダー システムはブラック ボックスとして動作するため、ユーザーや規制当局は出力を特定の意図に向けたり、その動作を監査したりすることができません。この制御性の欠如は、明示的なガイダンスに応答するシステムの能力として定義されますが、既存の評価パラダイムでは依然として対処されていない側面です。このギャップを埋めるために、制御性を体系的に評価するための協調的なマルチエージェント フレームワークである CtrlBench-Rec を提案します。私たちは、ターゲット コンテンツの発見、関心プロファイルの形成、人気バイアスの軽減という 3 つの基本的なタスクを形式化します。これらは、明示的なコマンドから暗黙的な表現のステアリング、そして最終的にはアルゴリズムのバイアスの克服までのステアビリティを一緒に測定します。実世界のデータセットと複数のレコメンデーション モデルに関する広範な実験により、私たちのフレームワークが制御性を効果的に定量化し、重大なシステムのボトルネック、特に誘導ロングテール コンテンツに対する永続的な抵抗を明らかにすることが実証されています。 CtrlBench-Rec は、制御可能な推奨調査、アルゴリズム監査、およびユーザー権限付与のための初の標準化されたツールキットを提供します。私たちのコードは https://github.com/caskcsg/CtrlBenchRec でリリースされています。

原文 (English)

Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents

Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, remains an unaddressed dimension in existing evaluation paradigms. To fill this gap, we propose CtrlBench-Rec, a collaborative multi-agent framework for systematic assessment of controllability. We formalize three fundamental tasks: target content discovery, interest profile shaping, and popularity bias mitigation, which together measure steerability from explicit commands to implicit representation steering and finally to overcoming algorithmic biases.Extensive experiments on real-world datasets and multiple recommendation models demonstrate that our framework effectively quantifies controllability and exposes critical system bottlenecks, most notably persistent resistance to guiding long tail content. CtrlBench-Rec provides the first standardized toolkit for controllable recommendation research, algorithmic auditing, and user empowerment. Our code is released on https://github.com/caskcsg/CtrlBenchRec.

2026-07-16 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models

Do vision models see colors the way humans do? Existing evaluations of color representations usually compare them with geometric spaces suc…

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達研究/論文

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging par…

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation

Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing sy…

2026-07-16 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors

Diffusion models (DMs) have advanced text-to-image (T2I) synthesis, yet their personalization capabilities raise serious privacy and copyri…

2026-07-16 13:00 JSTarXiv cs.AI画像/動画生成エージェントロボティクスビジネス/資金調達

RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset

The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by th…

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

推論が困難な場合: 臨床 SOAP ノート生成のためのフロンティア LLM のソース認識型評価

推論対応 LLM は医療推論ベンチマークで優れたパフォーマンスを発揮しますが、これらの利点が構造化された臨床文書に反映されるかどうかは不明のままです。私たちは、OMI Health、ACI-Bench、PriMock57 にわたるソース認識ベンチマークでの臨床対話からの SOAP ノート生成を使用して、この疑問を調査します。プロバイダーネイティブ推論と同一ソース検索拡張生成 (RAG) を個別に切り替える制御された 2x2 設計で GPT-5.4、DeepSeek-V4-Flash、および Gemma-4-E4B を評価します。成果は、リファレンスを認識した 2 人の LLM 審査員とともに 7 つの自動指標を使用して評価されます。どちらの評価アプローチでも、推論非対応の GPT-5.4 構成が全体として最高の品質を達成するのに対し、DeepSeek-V4-Flash は推論が有効な構成の中で最高のパフォーマンスを発揮するという点で一致しています。推論を有効にすると、3 つのデータセットすべてで GPT-5.4 のパフォーマンスが大幅に低下しますが、同じソースの RAG では、モデルに依存する改善は小さくなります。全体として、この調査結果は、専用のタスク固有の評価を行わずに、忠実度に敏感な SOAP ノート生成を向上させるために、より強力な推論機能を想定すべきではないことを示しています。

原文 (English)

When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation

Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical documentation; we investigate this question using SOAP note generation from clinical dialogue in a source-aware benchmark spanning OMI Health, ACI-Bench, and PriMock57. We evaluate GPT-5.4, DeepSeek-V4-Flash, and Gemma-4-E4B in a controlled 2x2 design that independently toggles provider-native reasoning and same-source retrieval-augmented generation (RAG). Outputs are assessed using seven automatic metrics alongside two reference-aware LLM judges. Both evaluation approaches agree that a non-reasoning GPT-5.4 configuration achieves the highest overall quality, while DeepSeek-V4-Flash performs best among reasoning-enabled configurations. Enabling reasoning significantly degrades GPT-5.4 performance across all three datasets, whereas same-source RAG yields smaller, model-dependent improvements. Overall, the findings indicate that stronger reasoning capability should not be assumed to improve fidelity-sensitive SOAP note generation without dedicated, task-specific evaluation.

2026-07-16 13:00 JSTarXiv cs.AIビジネス/資金調達

Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models

World-model evaluation for model-based reinforcement learning typically asks whether the learned model predicts reward and value well, whic…

2026-07-16 03:06 JSTTechCrunch AIビジネス/資金調達

SpaceX falls to $135 IPO price ahead of Starship launch

The stock has steadily fallen from the euphoric post-IPO high, showing that markets may be sobering up to the promises CEO Elon Musk made b…

2026-07-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

会話エージェントの多次元評価の運用化: 選択的再評価とモデル ベンチマークを備えたスケーラブルで管理されたパイプライン

小売会話エージェントを評価するには、語彙の重複の指標を超えて、意図の一致、事実性、有用性、明瞭さ、トーン、および全体的な応答品質を評価する方法が必要です。 LLM-as-a-judge メソッドは人間による評価に代わるスケーラブルな代替手段を提供しますが、運用環境の展開では、ガバナンス、再現性、コスト、スキーマの一貫性、トレーサビリティ、および信頼性の点で課題が生じます。小売会話システムの大規模評価のための、管理された構成主導のパイプラインである GenAI Evaluation を紹介します。正規化、シャーディング、非同期実行、スキーマに制約された LLM スコアリングを通じて実稼働チャットボット ログを処理します。このフレームワークは、有用性、真実性、明瞭さ、トーンの調整、および翻訳固有の側面を評価します。選択的再評価では、不完全、不正な形式、またはスキーマが無効なレコードのみが処理され、スキーマ ロック、バージョン管理された構成、検証ログ、およびレコード レベルの出自が監査可能性をサポートします。このフレームワークは毎日約 50,000 件のレコードを処理し、200 万件を超えるインタラクションを評価しました。検証には、訓練を受けた 4 人のアノテーターから得た、人間がラベルを付けた 12,980 件の階層化ランダムなレコードが使用されました。分類には、14 のインテント、156 のサブインテント、18 の主要ドメイン、および 129 のサブドメインが含まれていました。このパイプラインは、マクロ F1 スコア 0.93 と人間の許容可能な翻訳精度 89% を達成しました。

原文 (English)

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality. Although LLM-as-a-judge methods provide scalable alternatives to human evaluation, production deployment introduces challenges in governance, reproducibility, cost, schema consistency, traceability, and reliability. We present GenAI Evaluation, a governed, configuration-driven pipeline for large-scale evaluation of retail conversational systems. It processes production chatbot logs through normalization, sharding, asynchronous execution, and schema-constrained LLM scoring. The framework evaluates helpfulness, truthfulness, clarity, tone alignment, and translation-specific dimensions. Selective re-evaluation processes only incomplete, malformed, or schema-invalid records, while schema locking, versioned configurations, validation logs, and record-level provenance support auditability. The framework processes approximately 50,000 records daily and has evaluated more than two million interactions. Validation used 12,980 stratified-random human-labeled records from four trained annotators. Classification covered 14 intents, 156 sub-intents, 18 major domains, and 129 sub-domains. The pipeline achieved a macro F1 score of 0.93 and 89% human-acceptability accuracy for translation.

2026-07-15 13:00 JSTarXiv cs.AIビジネス/資金調達規制/政策

フロンティア言語モデルにおける CBRN 上昇評価のためのしきい値超過フレームワーク

フロンティア言語モデルが進歩するにつれて、政策立案者やモデル開発者は、モデルへのアクセスが、公共ツールのみと比較して、結果の大きな化学、生物、放射線、核(CBRN)の悪用を計画する非専門家の能力を実質的に高めるかどうかを評価する方法を必要としています。既存の CBRN 評価は、専門家以外の定義、脅威の範囲、ベースライン、スコアリング ルーブリック、および決定ルールが異なるため、研究間で結果を比較することが困難です。私たちは、上昇率調査を独立して実行可能なコンポーネントに分解する、閾値超過基準(TEC)フレームワークを導入します。つまり、専門家以外の参加者の適格性の決定、研究の CBRN 脅威範囲の定義、および重要な上昇率の統計的推定です。次に、生成型 (モデルがゼロからの計画作成を支援する) と修正主義型 (モデルが既存の計画の改良を支援する) という 2 つの形態の向上を決定する設計を使用して、大規模な実証研究で TEC フレームワークを運用します。この調査では、CBRN ドメイン全体にわたる攻撃計画が作成され、対象分野の専門家によるレビューを通じて評価し、生成的および修正主義的な上昇を推定しました。このフレームワークを適用した私たちの実証研究では、領域の不均一性が明らかになりました。この制御されたリリース前評価の下では、モデル支援計画は専門家と同等の指導評価を受けることがありましたが、物質的な上昇は放射線領域に限定されていることが確認されました。これらの調査結果は、デプロイされたモデルの動作を特徴づけるのではなく、緩和策とデプロイメントガバナンスの決定に影響を与えました。最後に、事前に指定された基準、明確なベースライン、生成的推定と修正主義的推定の分離、予備的なスクリーニング信号と確認されたリスク判定の慎重な区別を強調しながら、将来の CBRN 上昇率評価のための方法論的な教訓を述べます。

原文 (English)

A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models

As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone. Existing CBRN evaluations differ in non-expert definitions, threat scope, baselines, scoring rubrics, and decision rules, making results difficult to compare across studies. We introduce a Threshold Exceedance Criteria (TEC) framework that decomposes an uplift study into independently executable components: determining non-expert participant eligibility, defining the CBRN threat scope for the study, and statistically estimating material uplift. We then operationalize the TEC framework in a large-scale empirical study using a design that determines two forms of uplift: generative (where a model assists plan creation from scratch) and revisionist (where a model assists refinement of an existing plan). The study produced attack plans across the CBRN domains, which we evaluated through subject-matter-expert review to estimate generative and revisionist uplift. Applying the framework, our empirical study revealed domain heterogeneity: under this controlled pre-release evaluation, model-assisted plans sometimes received expert-equivalent instructional ratings, but confirmed material uplift was limited to the radiological domain. These findings informed mitigation and deployment-governance decisions rather than characterizing deployed model behavior. We conclude with methodological lessons for future CBRN uplift evaluations, emphasizing prespecified criteria, explicit baselines, separation of generative and revisionist estimates, and careful distinction between preliminary screening signals and confirmed risk determinations.

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

エージェント向けのハーネス進化の評価を再考する

LLM エージェントの自動ハーネス進化の評価を再検討します。既存のハーネス進化手法では、単体テスト ケースを使用してハーネス構成を検索し、同じ公開ベンチマークで最終パフォーマンスを報告します。このプロトコルは 2 つの基本的な懸念を引き起こします。まず、ハーネスの進化自体が反復的な検索手順であり、タスクのフィードバックを使用して候補ハーネスを繰り返し評価および修正します。したがって、エージェントのテスト時間のスケーリングと同様に、一致したフィードバックと推論バジェットの下で単純なタスクレベルの検索ベースラインと比較して、その利益がハーネス設計の改善によるものか、追加の検索のみによるものかを判断する必要があります。第 2 に、検索と最終評価は同じベンチマークを共有するため、報告されたタスクはその特定のタスク セットに過剰適合するリスクが生じます。これらの懸念に対処するために、同等のフィードバックと推論予算の下で、ハーネスの進化を単純なテスト時間のスケーリングと発見ベースラインと比較する広範な評価を実施し、また、保留されたタスクで進化したハーネスを評価して、発見された改善点が一般化するかどうかを評価します。 GPT-5.4 および Claude Opus 4.6 を使用した Terminal-Bench 2.1 での実験では、自動ハーネス進化が常に単純なテスト時間スケーリング手法を上回るパフォーマンスを発揮せず、一般化が限られていることを示しています。私たちの結果は、自動ハーネスの進化の有効性について重要な疑問を提起し、自動ハーネス設計のためのより公平な評価プロトコルとベンチマークの必要性を強調しています。私たちのコードは https://github.com/re Thinking-harness-evolution で入手できます。

原文 (English)

Rethinking the Evaluation of Harness Evolution for Agents

We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. First, harness evolution is itself an iterative search procedure that repeatedly evaluates and revises candidate harnesses using task feedback. As in agentic test-time scaling, it should therefore be compared with simple task-level search baselines under matched feedback and inference budgets to determine whether its gains arise from improved harness design or from additional search alone. Second, because the search and the final evaluation share the same benchmark, the reported gains risk overfitting to that specific task set. To address these concerns, we conduct an extensive evaluation comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluate evolved harnesses on held-out tasks to assess whether the discovered improvements generalize. Experiments on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6 show that automatic harness evolution does not consistently outperform simple test-time scaling methods and exhibits limited generalization. Our results raise important questions about the effectiveness of automatic harness evolution and highlight the need for fairer evaluation protocols and benchmarks for automatic harness design. Our code is available at https://github.com/rethinking-harness-evolution.

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

エージェントのベンチマークを決定するにはどれくらいのタスクがあれば十分ですか?パブリック LLM エージェント ベンチマークのリプレイ分析

エージェントのベンチマークでは、すべてのタスクの実行後に 2 つのエージェントを比較することがよくありますが、コストがかかる評価では部分的な実行が誘惑的になります。タスクの部分だけでは、部分的な実行が完了したベンチマークと同じペアごとの結論をサポートしているかどうかはわかりません。私たちは、SWE ベンチ、AppWorld、および tau ベンチからの完了したパブリック タスク レベルのレコードを再生することで、この疑問を研究します。部分的な予算は、完了したベンチマークの決定をサポートし、必要なタスク グループをカバーし、未解決の比較の目標部分のみを残す場合にのみ十分とみなされます。必要なタスクの割合は大きく異なります。 5 パーセント ポイントの予算グリッド上の厳格な 0 パーセント ポイントのしきい値では、AppWorld は 15 パーセント、tau-bench は 25 パーセント、SWE ベンチ検証は 90 パーセントで最初にすべての目標を満たします。 SWE-bench Lite は、プライマリ カバレッジ ルールに基づくすべての目標を 95% 満たしていません。部分評価レポートには、あるエージェントが別のエージェントよりどの程度優れている必要があるか、タスクがどのように選択されるか、どのようなカバレッジ ルールが必要か、どのような決定ルールが使用されるか、未解決のまま残される可能性がある比較の数が記載される必要があります。

原文 (English)

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not show whether a partial run supports the same pairwise conclusion as the completed benchmark. We study this question by replaying completed public task-level records from SWE-bench, AppWorld, and tau-bench. A partial budget counts as enough only when it supports the completed benchmark's decision, covers required task groups, and leaves no more than a target fraction of comparisons unresolved. The required task fraction varies sharply. At the strict 0 percentage point threshold on a 5 percentage point budget grid, AppWorld first meets all targets at 15 percent, tau-bench at 25 percent, and SWE-bench Verified at 90 percent; SWE-bench Lite does not meet all targets by 95 percent under the primary coverage rule. Partial-evaluation reports should state how much one agent must outperform another, how tasks are selected, what coverage rule is required, what decision rule is used, and how many comparisons may remain unresolved.

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

採点者を採点するのは誰ですか?自己改善する LLM エージェントのための、共進化する評価指標とスキル

自己進化するエージェント システムは、独自のスキルを作成、修正、廃止することで改善しますが、そのようなループはすべて、信頼できる評価基準がすでに存在するという隠れた前提に基づいています。実際のアプリケーションの多くではそうではありません。私たちは3つの主張をします。まず、メトリクスは \emph{evolved} 可能です。私たちのメトリクス ループは、完全な進化ライフサイクルの下で小さな欠点検出器の構成を検索し、10 項目のアンカーされた参照セットと一致するようにトレーニングされ、ラベルのない出力に対するコンセンサスによって正規化され、決して読み取られることのない保持されたアンカーに対して監査され、不透明な判定ではなく透明で検査可能なメトリクスを生成します。第 2 に、勝てる指標が存在しないため、指標は正確な指標があれば可能だったものを回復しつつあり、ライフサイクル管理スキル ループと指標の共進化である \emph{Double Ratchet} がそれを実現します。コード生成 (MBPP+)、エンタープライズ テキストから SQL (Spider~2.0-Snow)、および参照不要のレポート生成全体にわたって、同じものによって達成されたホールドアウト上昇率の 88 ~ 110\% を維持します。スキル ループは、グラウンド トゥルースまたは利用可能な最良のルーブリックによって駆動されます。第三に、安全性はアンカー規律と外部監査によってもたらされます。アンカー ガードを削除すると、メトリクスは空の検出器に折りたたまれますが、ライフサイクルを削除するとそうではありません。そして、進化したスキルがレポートのルーブリックを操作すると、独立した審査員がそれをキャッチし、1 つの検出器がそれを修復し、タスクを意識した審査員が、決定されたペアの 77% で進化前のベースラインよりも進化した出力を優先しました。私たちは、信頼できる自動検証機能が存在しない場合には、この障害を想定したアーキテクチャが正しいデフォルトであると主張します。

原文 (English)

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists. In many real applications it does not. We make three claims. First, metrics can be \emph{evolved}: our metric loop searches compositions of small drawback detectors under a full evolutionary lifecycle, trained to agree with a ten-item anchored reference set, regularized by consensus over unlabeled outputs, and audited against a held-out anchor it never reads, yielding a transparent, inspectable metric rather than an opaque judge. Second, since no metric exists to beat, the yardstick is recovering what an accurate metric would have enabled, and \emph{Double Ratchet}, our co-evolution of the metric with a lifecycle-managed skill loop, does so: across code generation (MBPP+), enterprise text-to-SQL (Spider~2.0-Snow), and reference-free report generation, it retains 88--110\% of the held-out lift achieved by the same skill loop driven by ground truth or the best available rubric. Third, safety comes from anchor discipline plus outer audits: removing anchor guards collapses the metric into a vacuous detector while removing the lifecycle does not; and when evolved skills gamed the report rubric, an independent judge caught it, one detector repaired it, and a task-aware judge then preferred the evolved outputs over the pre-evolution baseline in 77\% of decided pairs. We argue this failure-expecting architecture is the right default wherever no reliable automatic verifier exists.

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

沈黙による勝利: LLM 計画評価における削除の非単調性、自律的悪用、および型付き状態ゲート

計画評価者は、戦略計画が明確でなくなったことに対して報酬を与えることができます。この論文では、LLM によって生成されたベンチャー ルートの段階的な期待値スコアラーにおける失敗について研究します。命題 1 は、内部遷移を削除する一方で、その先行者を再ターゲットし、下流の値を保持することによるスコア変化を示します: Delta_k = (prod_{i<k} p_i)[c_k + (1 - p_k)R_{k+1}]。凍結された 26 経路コホートでは、57 個の許容される欠失すべてが分析上の同一性およびしきい値の兆候と一致し、すべての経路に少なくとも 1 つのスコア改善欠失がありました。スコアを求めるオプティマイザーは、ルートの再構築は許可されていますが、エクスプロイト メカニズムには通知されていませんでしたが、21/26 のルートでベースラインを上回るカバーされていない構造を発見しました。 GATE は、26/26 の沈黙ルートと 0/26 の正直な停止ルートのスコア公開を拒否しました。拒否の後、次の改訂では 47/54 がカバーされた構造に修復され、厳密にカバーされた改善は 1/26 から 13/26 に増加しました。アダプティブ コンパイラを意識した共著者は、レジストリと出所の境界を明らかにしました。v1/v1.5 の 4 つの条件すべてで義務チャネル回避は 6/6 のままでしたが、デルタインデックス付きコストフロアは、セマンティックな完全性を確立することなく、ビートオネストルートを 6/6 から 3/6 に、サイレンスによる資金提供可能性を 5/6 から 0/6 に削減しました。必要な作業を省略したという理由だけで計画のスコアが向上した場合、計画は改善されていません。評価によって不作為のインセンティブが生まれました。 PCSC は、モデルを介した型付き状態レコード上のポストホック省略スプライスを検出し、無効化します。テストした協調設定では、GATE は単なるポストホック フィルターではなく、決定論的な検索整形制約として機能します。任意の LLM 生成戦略の意味上の完全性や現実世界の品質は検証されません。

原文 (English)

Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation

Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer for LLM-generated venture routes. Proposition 1 gives the score change from deleting an interior transition while retargeting its predecessor and retaining downstream value: Delta_k = (prod_{i<k} p_i)[c_k + (1 - p_k)R_{k+1}]. On a frozen 26-route cohort, all 57 admissible deletions matched the analytic identity and threshold sign, and every route had at least one score-improving deletion. A score-seeking optimizer, allowed to restructure routes but not told the exploit mechanism, found baseline-beating uncovered structures in 21/26 routes. GATE refused score release for 26/26 silenced routes with 0/26 honest suspensions; after refusal, 47/54 next revisions repaired to a covered structure, and strict covered improvement rose from 1/26 to 13/26. An adaptive compiler-aware co-author exposed the registry-provenance boundary: obligation-channel evasions remained 6/6 across all four v1/v1.5 conditions, while delta-indexed cost floors reduced beat-honest routes from 6/6 to 3/6 and fundability-by-silence from 5/6 to 0/6 without establishing semantic completeness. If a plan scores better only because it omits necessary work, the plan did not improve; the evaluation created an omission incentive. PCSC detects and neutralizes post-hoc omission splices over model-mediated typed-state records. In the cooperative setting tested, GATE acts as a deterministic search-shaping constraint, not merely a post-hoc filter. It does not verify the semantic completeness or real-world quality of arbitrary LLM-generated strategies.

2026-07-15 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

データセットシフトの下での ROI ベースの甲状腺結節超音波分類のためのディープアンサンブルを使用した校正済みの選択的予測: 遡及的評価

背景: 深層学習モデルは超音波で甲状腺結節を分類できますが、信頼性の高い臨床意思決定のサポートには、特にデータセットのシフト下では、校正された確率、不確実性の推定、選択的参照も必要です。方法: ROI ベースの甲状腺結節分類と選択的画像ベースのトリアージ用に、校正された決定論的な 5 メンバーのディープ アンサンブルを開発しました。 TN5000 は、モデル開発、5 倍交差検証、メンバーごとのベクトル スケーリング キャリブレーション、および倍数固有のしきい値の選択に使用されました。 TN3K は、独立した外部データセット シフト評価として機能しました。このフレームワークでは、圧迫と興奮の注意、アンサンブル平均悪性確率、およびアンサンブル不一致スコアとして相互情報量 (MI) を備えた ConvNeXt-Tiny を使用しました。 3 段階のポリシーにより、画像は No-FNA 提案、FNA 推奨、または放射線科医のレビューに割り当てられました。結果: プールされたアウトオブフォールド TN5000 予測では、アンサンブルは AUC-ROC 0.9395、AP 0.9715、ECE 0.0088、および Brier スコア 0.0813 を達成しました。名目上のMI滞留率50%では、症例の7.2%がNo-FNAの提案、39.9%がFNAの推奨、52.9%が放射線科医の審査を受け、98.3%がNo-FNAのNPV、99.83%の悪性腫瘍捕捉率でした。 TN3K では、AUC-ROC は 0.7870 に減少し、AP は 0.7254 に減少し、ECE は 0.1899 に増加し、Brier スコアは 0.2281 に増加しました。凍結された TN5000 ポリシーでは、83.7% が見直し、1.0% が No-FNA、15.3% が FNA 推奨に割り当てられました。 No-FNA 経路に入った悪性画像はありませんでしたが、FNA 推奨 PPV は 76.6% に低下しました。結論: このフレームワークは強力な内部識別とキャリブレーションを示しましたが、外部閾値の輸送性は制限されていました。選択的予測は、自動トリアージに適さない画像を特定するのに役立ちますが、導入前にローカルでの再調整、しきい値の検証、および前向きの臨床評価が必要です。

原文 (English)

Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation

Background: Deep learning models can classify thyroid nodules on ultrasound, but reliable clinical decision support also requires calibrated probabilities, uncertainty estimation, and selective referral, particularly under dataset shift. Methods: We developed a calibrated deterministic five-member deep ensemble for ROI-based thyroid nodule classification and selective image-based triage. TN5000 was used for model development, five-fold cross-validation, member-wise vector-scaling calibration, and fold-specific threshold selection. TN3K served as an independent external dataset-shift evaluation. The framework used ConvNeXt-Tiny with squeeze-and-excitation attention, ensemble-mean malignancy probability, and mutual information (MI) as an ensemble-disagreement score. A three-tier policy assigned images to No-FNA suggestion, FNA recommendation, or radiologist review. Results: On pooled out-of-fold TN5000 predictions, the ensemble achieved AUC-ROC 0.9395, AP 0.9715, ECE 0.0088, and Brier score 0.0813. At 50% nominal MI retention, 7.2% of cases received a No-FNA suggestion, 39.9% an FNA recommendation, and 52.9% radiologist review, with 98.3% No-FNA NPV and 99.83% malignancy capture. On TN3K, AUC-ROC decreased to 0.7870, AP to 0.7254, ECE increased to 0.1899, and Brier score to 0.2281. The frozen TN5000 policy assigned 83.7% to review, 1.0% to No-FNA, and 15.3% to FNA recommendation. No malignant image entered the No-FNA pathway, but FNA-recommendation PPV fell to 76.6%. Conclusion: The framework showed strong internal discrimination and calibration, but limited external threshold transportability. Selective prediction may help identify images unsuitable for automated triage, but local recalibration, threshold validation, and prospective clinical evaluation are required before deployment.

2026-07-15 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction

EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning. Yet EEG-derived reconstructions are blurry, distorted,…

2026-07-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor s…

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, howe…

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models

Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo c…

2026-07-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks

Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide ri…

2026-07-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on whic…

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

推論によるフロンティア LLM 評価の形状計算方法

AI の評価は、ツールの使用と反復的な問題解決を伴う長期にわたる軌道から恩恵を受ける、より困難なタスクへと移行しています。その結果、パフォーマンスは、テスト時に利用可能なコンピューティング (「推論コンピューティング」) の量と割り当てにますます敏感になります。しかし、多くの評価では依然として単一の制限された予算でのパフォーマンスが報告されており、低いスコアはモデルの基礎的な機能ではなく評価設定を反映している可能性があることを意味します。これをテストするために、ソフトウェア エンジニアリング、数学、医学、サイバーセキュリティにわたる 7 つの挑戦的なベンチマークで最大 12 のフロンティア言語モデルを評価します。私たちは、3 つの単純な推論スケーリング介入を組み合わせた制御されたセットアップを使用します。つまり、より大きなトークン バジェット、コンテキストの圧縮、およびモデル自体または最小限の正確性フィードバックによって導かれる送信の試行の繰り返しです。主な結果は 3 つあります。まず、トークン バジェットが大きくなると、サイバーセキュリティ、FrontierMath、人類最後の試験、ターミナルベンチなど、複数のドメインにわたるベンチマークのパフォーマンスが大幅に向上します。第二に、固定予算の評価では、モデルが進歩するにつれてフロンティアの能力がますます過小評価される可能性があります。新しいモデルは、大きな予算でより高いパフォーマンスを実現し、より困難なタスクを解放し、より確実に解決します。第三に、どの推論スケーリング手法が最も役立つかがベンチマークによって異なります。繰り返し送信するとパフォーマンスが大幅に向上しますが、より大きなトークン バジェット、外部フィードバック、および並列試行の値はベンチマークによって異なります。全体として、私たちの結果は、ベンチマーク スコアがプロトコルに依存していることを示しています。したがって、評価では、特に安全性またはポリシー関連の設定において、推論時間のコンピューティングの関数として機能を報告し、プロトコルの選択を明示的に指定し、一致した予算で大規模な共有コンピューティング範囲にわたってモデルの世代を比較する必要があると主張します。

原文 (English)

How Inference Compute Shapes Frontier LLM Evaluation

AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still report performance at a single restrictive budget, meaning that low scores may reflect the evaluation setup rather than the model's underlying capability. To test this, we evaluate up to 12 frontier language models on seven challenging benchmarks spanning software engineering, mathematics, medicine, and cybersecurity. We use a controlled setup combining three simple inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts, guided either by the model itself or by minimal correctness feedback. We find three main results. First, larger token budgets substantially improve performance on benchmarks across multiple domains, including cybersecurity, FrontierMath, Humanity's Last Exam, and TerminalBench. Second, fixed-budget evaluations can increasingly understate frontier capability as models advance. Newer models reach higher performance at large budgets, where they unlock harder tasks and solve them more reliably. Third, benchmarks differ in which inference-scaling methods help most: repeated submission broadly improves performance, but the value of larger token budgets, external feedback, and parallel attempts varies by benchmark. Overall, our results show that benchmark scores are protocol-dependent. We therefore argue that evaluations should report capability as a function of inference-time compute, specify protocol choices explicitly, and compare model generations over a large shared compute range at matched budgets, especially in safety- or policy-relevant settings.

2026-07-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

AgentLens: コーディング エージェント評価のための実稼働環境で評価された軌跡レビュー

ここでは、対話型コード エージェントの実稼働環境で評価されたベンチマークである AgentLens を紹介します。ほとんどのコード エージェント ベンチマークでは、実行が 1 ビットに削減されます。タスクは成功しましたか? -- しかし、これらのエージェントを実際に使用する人々は、エージェントがどのように指示に従い、ツールを使用し、自身の作業を検証し、間違いから回復し、途中でエージェントに話しかけるかという軌跡全体を経験します。 AgentLens はその軌跡全体を評価します。客観的なチェックが存在する正式な検証と、LLM で作成された軌跡のレビューおよび並べての比較を組み合わせることで、各実行でスコアがなぜそのようになるのかについての読みやすい説明が得られます。これにより、AgentLens はモデルのランク付け以上の用途に役立ちます。モデルの動作を診断し、独自のエージェントの連続バージョンを比較し、夜間の評価パイプラインで製品の回帰を捕捉するために使用されます。 https://github.com/agent-lens/agent-lens-bench でベンチマークをオープンソースとしてリリースします。

原文 (English)

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at https://github.com/agent-lens/agent-lens-bench.

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasing…

2026-07-15 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

Spoken query retrieval is an important interaction mode in modern information retrieval. However, existing evaluation datasets are often li…

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory

Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accurac…

2026-07-15 09:27 JSTTechCrunch AILLM/生成AIビジネス/資金調達研究/論文

OpenAI researcher Miles Wang in talks to launch AI drug discovery startup valued at $2B

The funding discussions point to investor interest in applying AI to make breakthroughs in life sciences.

2026-07-15 04:39 JSTTechCrunch AIビジネス/資金調達

The founder of Hinge raised $18M to build a new AI dating service, Overtone

Overtone describes itself as "a voice- and audio-forward service, enabled by AI, that provides highly curated introductions."

2026-07-14 19:00 JSTOpenAIエージェントビジネス/資金調達

How to manage AI investments in the agentic era

Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling h…

2026-07-14 13:00 JSTarXiv cs.AIビジネス/資金調達

検証者はカリキュラムです: ファミリー間ゲーム生成のための実行ゲート型自己蒸留

学習した審査員に対してコード ジェネレーターをポストトレーニングすると、アーティファクトを改善せずにスコアを上げるプロキシ機能を最適化できます。私たちは逆の信号、つまり決定論的で判断力のない、ゲーム性のないフィルター、つまり生成されたプロジェクトがヘッドレス エンジン (厳密起動) で正常に起動するかどうかを研究します。このゲートの下では、拒絶サンプリングの自己蒸留化合物が族外の一般化を引き起こします。 GameCraft-Bench (自然言語の概要を完全な Godot プロジェクトにマッピング) では、厳密な起動の下で蒸留された 14B モデル (Qwen3-14B+LoRA) により、4 つの未見のゲーム ファミリのクリーン ジェネレーションが候補あたり 8.8% から 42.2% に上昇し、ベストオブ K カバレッジが 3 ラウンドにわたって 18/25 から 25/25 (ゴールド天井) に上昇し、それぞれ大幅な向上を実現しました。 (p=0.0019、p<1e-4、p<1e-4)。この利益は単にデータを追加することによるものではありません。完全に一致するゴールド重複コントロールは基本モデルを下回ります (5.6% 対 8.8%、p=0.019)。一方、カウント一致分解は、ラウンド 1 から 2 へのジャンプを同等の品質 (+8.8pp) と量 (+8.5pp) のチャネルに分割します。最も直接的には、フィルターのみを交換してループを再実行すること (ローンチ ゲートの代わりに世代の 99.9% を通過させる寛大な BUILD チェック) により、ゲインが完全に消去され (基本に戻り、ローンチ ゲート ラウンドに対して p=1e-3)、オプティマイザーではなくベリファイアの精度が分離されます。 2 番目のゲーム不可能なシグナルであるヘッドレス実行グラウンディングは、ラウンド全体で単調に上昇し、一致した予算 (16 対 5) でゴールドの複製よりもはるかに多くのグラウンディングされた候補を生成し、ゲインが機能していることを確認し、ローンチではなく空です。ゲーム生成は 1 つのレッスンの検証可能なテストベッドです。検証者はカリキュラムであり、検証者が証明するものはモデルが学習するものです。

原文 (English)

The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation

Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact. We study the opposite signal: a deterministic, judge-free, ungameable filter -- whether a generated project launches cleanly under a headless engine (strict-launch). Under this gate, rejection-sampling self-distillation compounds out-of-family generalization. On GameCraft-Bench (mapping a natural-language brief to a complete Godot project), a 14B model (Qwen3-14B+LoRA) distilled under strict-launch raises clean generation on four unseen game families from 8.8% to 42.2% per-candidate and best-of-K coverage from 18/25 to 25/25 (the gold ceiling) over three rounds, each a significant gain (p=0.0019, p<1e-4, p<1e-4). The gain is not from merely adding data: an exactly-matched gold-duplication control regresses below the base model (5.6% vs. 8.8%, p=0.019), while a count-matched decomposition splits the round-1-to-2 jump into comparable quality (+8.8pp) and quantity (+8.5pp) channels. Most directly, rerunning the loop with only the filter swapped -- the lenient BUILD check, which passes 99.9% of generations, in place of the launch gate -- erases the gain entirely (back to base, p=1e-3 vs. the launch-gated round), isolating verifier precision rather than the optimizer. A second ungameable signal, headless execution grounding, rises monotonically across rounds and yields far more grounded candidates than gold-duplication at a matched budget (16 vs. 5), confirming the gains are functional, not launch-but-empty. Game generation is a verifiable testbed for one lesson: the verifier is the curriculum -- what it certifies is what the model learns.

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

スコアセット前のコアセット: LLM ベンチマークの評価 - 教師なしプロンプトサブセット選択

私たちは LLM ベンチマーク コアセットの選択について研究します。つまり、誘導されたモデル スコアとランキングが完全なベンチマーク スイートから得られたものに近似する複数のベンチマークからプロンプトの小さなサブセットを選択します。評価なしの教師なしベンチマーク コアセットの選択 (私たちのアプローチ) では、選択アルゴリズムはモデルの評価結果を使用せず、ベンチマーク全体のサブコレクションを生成するのではなく、複数のベンチマークにわたってプロンプトのサブセットを生成することにより、細かい粒度で動作します。私たちはサブモジュール式サブセット選択を使用し、この目的のために、決定点プロセス (DPP) ベースのアプローチ、サブモジュール式相互情報関数、施設の位置ベースの関数など、さまざまなサブモジュール式関数を開発および評価します。 5 つの異なる機能カテゴリ、18 のフロンティア LLM、および 61,000 を超えるプロンプトにまたがる 35 の異種ベンチマークからなる新しい大規模スイートでは、安価なセマンティック プロンプト エンベディングのみで動作するファシリティ ロケーション (FL) 関数が、さまざまなコアセット予算にわたって、12 の個別のスコアベースおよび多様性ベースのベースラインよりも優れた LLM スコアを維持することがわかりました。さらに、私たちが提案する目標が教師なし評価方式に限定されないことを示します。少数のベンチマーク全体のみを選択する必要があり、大量のモデル スコアが利用可能な設定では、同じ目標が MMLU および MTEB リーダーボードの最先端のベースラインと一致または上回るパフォーマンスを示し、同時に計算コストが大幅に安くなります。まとめると、私たちの結果は、サブモジュール性が一般に、ベンチマーク圧縮のための強力で信頼性の高いツールであることを示唆しています。

原文 (English)

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite. In evaluation-unsupervised benchmark coreset selection (our approach), the selection algorithm uses no model evaluation outcomes, and operates on a fine granularity by producing subsets of prompts over multiple benchmarks rather than producing a sub-collection of entire benchmarks. We use submodular subset selection, and we develop and evaluate many different submodular functions for this purpose, including determinantal point process (DPP) based approaches, submodular mutual information functions, and facility location-based functions. On a new large-scale suite of 35 heterogeneous benchmarks spanning five different capability categories, 18 frontier LLMs, and over 61K prompts, we find that the facility location (FL) function operating exclusively on inexpensive semantic prompt embeddings preserves LLM scores better than twelve separate score-based and diversity-based baselines, across a range of coreset budgets. Moreover, we show our proposed objective is not limited to the evaluation-unsupervised regime: in the setting where only a handful of whole benchmarks must be selected and a large amount of model scores are available, the same objective matches or outperforms state-of-the-art baselines on the MMLU and MTEB leaderboards, while being substantially cheaper to compute. Together, our results suggest that submodularity, in general, is a strong and reliable tool for benchmark compression.

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AgentAbstain: LLM エージェントは、いつ行動すべきではないかを知っていますか?

大規模言語モデル (LLM) に基づくエージェント システムは自律タスクに導入されることが増えていますが、既存の評価では主に、エージェントがいつやめるべきかを知っているかどうかよりも、タスクの成功に重点が置かれています。このギャップは実際のリスクをもたらします。曖昧さ、矛盾する制約、またはツールの障害の下では、エージェントが意図しない取り消し不能なアクションを実行する可能性があります。このギャップを埋めるために、我々は、エージェントの棄権に関する最初の体系的な評価フレームワークを提示します。それは、ツールを使用する LLM エージェントがいつ行動を起こさないかを認識する調整された能力です。 AgentAbstain の核心は、実行前の推論と実行時検出にわたる 8 つの棄権シナリオのエージェント ネイティブ分類に基づいて構築されたペアタスク ベンチマークです。これには、42 の実行可能なサンドボックス環境にわたる 263 のペアのタスクが含まれており、各ペアは、命令、ツール、または環境の状態に対する制御された摂動を通じて生成される、実行すべきタスクと禁止すべきバリアントで構成されます。このペア設計を拡張し、データ汚染に対抗するために、サンドボックス環境を合成し、決定論的再生およびセマンティック LLM ジャッジによって検証されたペアタスクをエンドツーエンドで生成する完全に自動化されたパイプラインである AbstainGen を提案します。新しいタスク インスタンスはオンデマンドで再生成でき、3 人の独立したアノテーターは、サンプリングされたタスクの 94 ~ 98% が適切に設計されていると評価しました。 4 つのエージェント ハーネスの 17 個のフロンティア LLM にわたって、最高のエージェント (Gemini 3.1 Pro) は、ペアの精度 (各ペアのタスクの実行側と棄権側の両方で正しい) を 59.5% しか達成していません。さらに重要なことは、棄権能力は一般的な課題解決能力とはほとんど独立しており、課題解決を拡大するだけではこのギャップは埋まらないことを示しています。さらに、エージェントが棄権トリガーを認識する前に不可逆的なアクションを実行する事後棄権などの障害モードを特定します。私たちのコードとデータセットは、agentabstain.github.io でオープンソース化されています。

原文 (English)

AgentAbstain: Do LLM Agents Know When Not to Act?

Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain. This gap poses real risks: under ambiguity, conflicting constraints, or tool failures, agents may execute unintended and irreversible actions. To close this gap, we present the first systematic evaluation framework for agentic abstention: the calibrated ability of tool-using LLM agents to recognize when not to act. At its core, AgentAbstain is a paired-task benchmark built on an agent-native taxonomy of 8 abstention scenarios across pre-execution reasoning and runtime discovery. It contains 263 paired tasks across 42 executable sandbox environments, where each pair consists of a should-act task and a should-abstain variant produced through a controlled perturbation to the instruction, tool, or environment state. To scale this paired design and resist data contamination, we propose AbstainGen, a fully automated pipeline that synthesizes sandbox environments and generates paired tasks end-to-end, validated by deterministic replay and semantic LLM judges; fresh task instances can be regenerated on demand, and three independent annotators rate 94-98% of sampled tasks as well-designed. Across 17 frontier LLMs in 4 agent harnesses, the best agent (Gemini 3.1 Pro) achieves only 59.5% paired accuracy (correct on both the act and abstain sides of each paired task). More importantly, abstention capability is largely independent of general task-solving capability, indicating that scaling task-solving alone will not close this gap. We further identify failure modes such as post-hoc abstention, in which agents execute irreversible actions before recognizing abstention triggers. Our code and dataset are open-sourced at agentabstain.github.io.

2026-07-14 13:00 JSTarXiv cs.AIビジネス/資金調達

スパース機能の介入が実際にローカライズされるのはいつですか? SAEベースの安全管理の整合評価

スパース オートエンコーダ (SAE) 機能が安全関連動作のローカライズされた制御ハンドルとして機能する場合を評価します。この質問は難しい。なぜなら、見かけ上の成功は、弱い介入、不一致のベースライン、モデルの堅牢性、または意味のある有害なコンプライアンスを示さずに自動安全判定者が危険とマークする劣化した出力から生じる可能性があるからである。実行時の安全性介入のための整合コヒーレンスゲート評価プロトコルを導入します。手法は整合されたターゲット効果点で比較され、出力が安全でないと判断され、一貫性がある場合にのみ、主要なターゲットメトリクスが有害なコンプライアンスをカウントします。このプロトコルを Gemma Scope レイヤー 20 残存 SAE を使用した Gemma-2-9B-it の 3 つのプロンプト スプリットに適用すると、SAE 特徴アブレーションの有用な領域が狭いことがわかります。 SAE トップ 800 は、より低い総摂動と競争的有用性により、低から中程度の目標効果に達しますが、SAE トップ 1600 は、一致した高密度の拒否方向ベースラインと比較して有用性を失い、SAE トップ 3200 は主に一貫性崩壊を引き起こします。人間の監査により、コヒーレンス ゲーティングにより安全でないのみのアーティファクトが除去されることが確認され、機能診断により、有用なレジームは、活性化分離がランクとともに急速に減衰する拒否整合機能の安定したヘッドによって駆動されることが示されました。これらの結果は、SAE に基づく安全介入は一律に局所的であると想定するのではなく、体制依存の制御メカニズムとして評価されるべきであることを主張しています。

原文 (English)

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior. This question is difficult because apparent success can arise from weak interventions, mismatched baselines, model robustness, or degenerate outputs that automated safety judges mark as unsafe without representing meaningful harmful compliance. We introduce a matched coherence-gated evaluation protocol for runtime safety interventions: methods are compared at matched target-effect points, and the primary target metric counts harmful compliance only when an output is both judge-unsafe and coherent. Applying this protocol to three prompt splits on Gemma-2-9B-it with a Gemma Scope layer-20 residual SAE, we find that SAE feature ablation has a narrow useful regime. SAE top800 reaches a low-to-mid target effect with lower total perturbation and competitive utility, but SAE top1600 loses utility relative to a matched dense refusal-direction baseline, and SAE top3200 primarily induces coherence collapse. Human audit confirms that coherence gating removes unsafe-only artifacts, and feature diagnostics show that the useful regime is driven by a stable head of refusal-aligned features whose activation separation decays rapidly with rank. These results argue that SAE-based safety interventions should be evaluated as regime-dependent control mechanisms rather than assumed to be uniformly localized.

2026-07-14 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

チェッカーから予測者へ: 遅延地上真実の下でのモデル生成の戦略的ルートのコード所有の評価

モデル出力の評価の多くは、評価時にチェックできるコントラクト、または運用ループ内に到着するフィードバックのいずれかに依存します。私たちは、グラウンド トゥルースが遅延、検閲、または非公開であるため、決定論的コードがスコアリング時に正確さをチェックできず、代わりにコード所有の暫定予測を発行する必要があるという補完的な設定を研究します。 RouteCast は、モデル生成の型付き戦略ルートに対してこの体制をインスタンス化します。モデルは候補ルートと構造化された要素を提案します。特定時点の証拠、参照クラス、および決定論的変換により、暫定的な予測ランキングが生成されます。後の結果によって予測が評価されます。 21 件のバイナリ結果ケース (陽性 6 件、陰性 15 件) を対象とした遡及的ベンチャー パイロットでは、パケット全体の RouteCast スコアは予備的な遡及的差別 (AUC 0.756、95% CI [0.471,0.980]) を示しましたが、盲目の LLM 裁判官は AUC 0.678 [0.419, 0.897] に達し、アイデンティティを暴露された LLM 裁判官は AUC に達しました0.761 [0.515,0.944]、認識または結果に関連した漏洩リスクと一致。同じバイナリ サブセットに対する事前登録された分解アブレーションにより、同一の入力を型付きステージング ルートに変換することは、パケット全体のスコア (デルタ AUC = -0.144、95% CI [-0.471,0.176]) および決定論的ヒューリスティック (デルタ AUC = -0.089、95% CI) と区別できないことがわかりました。 [-0.412、0.278])。パイロットは、監査可能な実現可能性の結果を確立し、障害モードを明らかにします。将来のキャリブレーション、因果関係の決定の改善、ルート分解の利点、またはクロスドメインの妥当性を確立するものではありません。

原文 (English)

From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast. RouteCast instantiates this regime for model-generated typed strategic routes: models propose candidate routes and structured factors; point-in-time evidence, reference classes, and deterministic transformations produce a provisional forecast-ranking; later outcomes evaluate the forecast. In a retrospective venture pilot on 21 binary-outcome cases (6 positive, 15 negative), the whole-packet RouteCast score showed preliminary retrospective discrimination (AUC 0.756, 95% CI [0.471,0.980]), while a blind LLM judge reached AUC 0.678 [0.419,0.897] and an identity-exposed LLM judge reached AUC 0.761 [0.515,0.944], consistent with recognition- or outcome-related leakage risk. A preregistered decomposition ablation on the same binary subset found that converting the identical inputs into typed staged routes was indistinguishable from the whole-packet score (Delta AUC = -0.144, 95% CI [-0.471,0.176]) and from a deterministic heuristic (Delta AUC = -0.089, 95% CI [-0.412,0.278]). The pilot establishes an auditable feasibility result and exposes failure modes; it does not establish prospective calibration, causal decision improvement, route-decomposition advantage, or cross-domain validity.

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

隠れたフットプリント: ストレージを LLM エージェント評価の第一級の指標にする

LLM エージェントのベンチマークは、タスクの完了、信頼性、推論コストを測定しますが、ログ、コンテキスト スナップショット、チェックポイント、デバッグ トレースなど、エージェントの実行によってディスクに残される永続データは測定しません。実行後のエージェント ストレージ フットプリントのクロスフレームワーク ベンチマークである AgentFootprint を紹介します。そのシリアル化対応メトリクス スイートは、総保持率、チャネル構成、重複、増加、圧縮率、会話履歴の再構築可能性を測定します。これは、測定の罠に対処します。単純なバイトレベルの測定では、データベースのページングと JSON エスケープが繰り返されるコンテンツを不明瞭にするため、重複が桁違いに過小評価されます。固定トレース制御により、エージェントが生成した論理ボリュームが永続層の増幅から分離されます。7 つの永続フレームワークを通じて同じ軌跡を再生すると、6.7 倍の広がりが得られます。同一のモデル、ツール、およびタスクでは、100% の精度の構成では、デフォルトでサポートされる回復機能と監査機能が異なりますが、保持バイト数が 15.7 倍異なります。 3 つの完全な履歴構成は、反復観察ストレス タスクで超線形に成長します。 108 個のインスタンスで正規化された SWE ベンチからエクスポートされた軌跡 検証済みの送信は、インスタンスごとに 3 桁の大きさに及び、解決率との検出可能な相関関係はありません。コンテンツ アドレス ストアは、すべての再構築可能性スコアを維持しながら、保持率を 4.8 倍から 32.7 倍まで削減します。これらの結果は、精度と再構築可能性を併せてレポートするためのリソース メトリックとして永続ストレージを確立します。

原文 (English)

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark of post-run agent storage footprint. Its serialization-aware metric suite measures total retention, channel composition, duplication, growth, compressibility, and conversation-history reconstructability. It addresses a measurement trap: naive byte-level measurement understates duplication by an order of magnitude because database paging and JSON escaping obscure repeated content. A fixed-trace control separates agent-generated logical volume from persistence-layer amplification: replaying the same trajectory through seven persisting frameworks yields a 6.7x spread. Under identical models, tools, and tasks, configurations with 100% accuracy differ by 15.7x in retained bytes, although their defaults support different recovery and audit capabilities. Three full-history configurations grow superlinearly on a repeated-observation stress task. Exported trajectories from 108 instance-normalized SWE-bench Verified submissions span three orders of magnitude per instance, with no detectable correlation with resolve rate. A content-addressed store reduces retention by 4.8x-32.7x while preserving every reconstructability score. These results establish persistent storage as a resource metric to report jointly with accuracy and reconstructability.

2026-07-14 13:00 JSTarXiv cs.AIビジネス/資金調達

OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis

Failure diagnosis in modern software systems requires iterative evidence acquisition and hypothesis reasoning guided by operational experie…

2026-07-14 13:00 JSTarXiv cs.AIエージェントロボティクスビジネス/資金調達

OmniSCS: Omni Safety-Critical Scenario Synthesis for Autonomous Driving via a Fully Editable Driving World

The synthesis of safety-critical scenarios (SCS) and their evaluation through closed-loop simulations are crucial for developing robust aut…

2026-07-14 13:00 JSTarXiv cs.AIエージェントロボティクスビジネス/資金調達

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation

Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prio…

2026-07-14 13:00 JSTarXiv cs.AIビジネス/資金調達

A Production-Oriented Framework for Evaluation of SFX Generation

Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity…

2026-07-14 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circ…

2026-07-14 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects

Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliabi…

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating r…

2026-07-14 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

A Unified Framework for Comprehensive Cardiac CT Segmentation and Phenotyping: Human-in-the-Loop Data Annotation, Vision Foundation Model Development, Multicenter Evaluation and Clinical Validation

Comprehensive quantification of cardiac structures from computed tomography (CT) remains limited not by data availability but by the scalab…

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task th…

2026-07-14 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達研究/論文

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a state…

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation

Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely depe…

2026-07-14 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks

Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide ri…

2026-07-14 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

LongMedBench: 長期的な臨床意思決定のための医薬品のベンチマーク

この研究では、長期的な臨床意思決定のための実際の EHR ベースのベンチマークである LongMedBench を紹介します。 LLM ベースの医療薬剤のこれまでの評価では、主にショートコンテキスト知識の QA とツールの使用が重視されてきました。しかし、実際の医療は本質的に長期的なものであり、臨床医は繰り返しの訪問、検査、進化する治療にわたる証拠を集約する必要があります。したがって、現実的な評価には長期的な相互作用が不可欠です。 LongMedBench は、MIMIC-IV 入院記録と臨床ノートを時系列イベント ストリームとロングコンテキスト メモリ データセットに統合する再現可能なパイプラインを介して構築されており、エージェントと臨床環境の間で長期にわたるマルチセッションの対話を可能にします。患者数は 335 名で、患者 1 人当たりの入院件数は平均 19.72 件、1 件当たりの医療イベント数は 44.91 件です。長期的な意思決定プロセスに基づいて、事実に基づく QA、時間的推論、長期的な意思決定という 3 つのスイートによる評価分類法を提案します。この分類法は、エージェントが長期にわたって過去の患者情報をどのように理解し、活用しているかを測定します。私たちの実験によると、最近の LLM は明示的なタイムスタンプをうまく利用できますが、暗黙的な時間推論には課題があることがわかりました。 RAG とエージェント メモリ システムは、情報検索タスクのパフォーマンスを向上させることができますが、意思決定タスクのパフォーマンスはモデルの直接のコンテキストに大きく依存します。

原文 (English)

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use. However, real-world medical care is inherently longitudinal, and clinicians must aggregate evidence across repeated visits, tests, and evolving treatments. Therefore, long-horizon interaction is essential for realistic assessment. LongMedBench is constructed via a reproducible pipeline that integrates MIMIC-IV admission records and clinical notes into time-series event streams and long-context memory datasets, enabling long-horizon, multi-session interactions between agents and a clinical environment. It comprises 335 patients, with 19.72 inpatient visits per patient on average and 44.91 medical events per visit. Guided by the long-horizon decision process, we propose an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making. This taxonomy measures how agents understand and leverage historical patient information over extended horizons. Our experiments show that while recent LLMs can make good use of explicit timestamps, they have challenges in implicit time inference; The RAG and agent memory system can improve the performance of information retrieval tasks, but the performance of decision-making tasks is highly dependent on the model's immediate context.

2026-07-14 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達研究/論文

On the Necessity of Output Distribution Reweighting for Effective Class Unlearning

In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause inf…

2026-07-14 13:00 JSTarXiv cs.AIビジネス/資金調達

Rethinking Zero-Shot Time Series Classification: From Task-specific Classifiers to In-Context Inference

The zero-shot evaluation of time series foundation models (TSFMs) for classification typically uses a frozen encoder followed by a task-spe…

2026-07-14 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ECoLAD: Selecting Anomaly Detectors for Automotive Deployment via Compute-Reduction Evaluation

Automotive anomaly detectors are often selected from accuracy only benchmarks on workstation class hardware, whereas in-vehicle monitoring…

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA

Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive…

2026-07-14 13:00 JSTarXiv cs.AIビジネス/資金調達

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.…

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation

Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agen…

2026-07-14 09:00 JSTTechCrunch AIビジネス/資金調達

Video-generation startup PixVerse raises $439M, valuation soars past $2B

With the cash, the company aims to expand its world model offering and reach customers across geographies.

2026-07-14 08:31 JSTTechCrunch AIエージェントロボティクスビジネス/資金調達研究/論文

Hermes agent maker Nous Research in talks for new funding at $1.5B valuation

The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.

2026-07-13 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

L-MAD: 法的推論におけるマルチエージェントの議論構造の体系的評価

マルチエージェントディベート (MAD) フレームワークは、一般的な推論において大きな可能性を示していますが、高度に構造化され、知識が必要な法的領域におけるその有効性は依然として十分に研究されていません。この研究では、法的本文含意内のさまざまな議論の構造と集計方法を体系的に評価するために、法的マルチエージェント討論(L-MAD)フレームワークを導入します。 L-MAD は、個別の専門家ペルソナを複数のエージェントに割り当てることで、単一エージェントの強力なベースラインを最大 8\% 改善します。さらに、討論の規模を分析すると、明らかなトレードオフが明らかになります。エージェントの数を増やすと、矛盾が減り、精度が向上します。一方、討論ラウンドを延長すると、エージェントが互いの間違いを補強し合う、有害な \textit{過剰審議のドリフト} が誘発されます。最終的に、私たちの調査結果は、一か八かの法的推論環境で協調的なマルチエージェント システムを導入する際の実際的な境界と安全マージンを概説します。

原文 (English)

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly structured, knowledge-heavy legal domains remains under-explored. In this work, we introduce the Legal Multi-Agent Debate (L-MAD) framework to systematically evaluate different debate structures and aggregation methods within Legal Textual Entailment. By assigning distinct expert personas to multiple agents, L-MAD improves upon strong single-agent baselines by up to 8\%. Furthermore, analyzing how debate scales reveals a clear trade-off: increasing the agent population reduces inconsistency and improves accuracy, whereas extending discussion rounds induces a detrimental \textit{over-deliberation drift} where agents reinforce each other's mistakes. Ultimately, our findings outline the practical boundaries and safety margins of deploying collaborative multi-agent systems in high-stakes legal reasoning environments.

2026-07-13 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

LongMedBench: 長期的な臨床意思決定のための医薬品のベンチマーク

この研究では、長期的な臨床意思決定のための実際の EHR ベースのベンチマークである LongMedBench を紹介します。 LLM ベースの医療薬剤のこれまでの評価では、主にショートコンテキスト知識の QA とツールの使用が重視されてきました。しかし、実際の医療は本質的に長期的なものであり、臨床医は繰り返しの訪問、検査、進化する治療にわたる証拠を集約する必要があります。したがって、現実的な評価には長期的な相互作用が不可欠です。 LongMedBench は、MIMIC-IV 入院記録と臨床ノートを時系列イベント ストリームとロングコンテキスト メモリ データセットに統合する再現可能なパイプラインを介して構築されており、エージェントと臨床環境の間で長期にわたるマルチセッションの対話を可能にします。患者数は 335 名で、患者 1 人当たりの入院件数は平均 19.72 件、1 件当たりの医療イベント数は 44.91 件です。長期的な意思決定プロセスに基づいて、事実に基づく QA、時間的推論、長期的な意思決定という 3 つのスイートによる評価分類法を提案します。この分類法は、エージェントが長期にわたって過去の患者情報をどのように理解し、活用しているかを測定します。私たちの実験によると、最近の LLM は明示的なタイムスタンプをうまく利用できますが、暗黙的な時間推論には課題があることがわかりました。 RAG とエージェント メモリ システムは、情報検索タスクのパフォーマンスを向上させることができますが、意思決定タスクのパフォーマンスはモデルの直接のコンテキストに大きく依存します。

原文 (English)

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use. However, real-world medical care is inherently longitudinal, and clinicians must aggregate evidence across repeated visits, tests, and evolving treatments. Therefore, long-horizon interaction is essential for realistic assessment. LongMedBench is constructed via a reproducible pipeline that integrates MIMIC-IV admission records and clinical notes into time-series event streams and long-context memory datasets, enabling long-horizon, multi-session interactions between agents and a clinical environment. It comprises 335 patients, with 19.72 inpatient visits per patient on average and 44.91 medical events per visit. Guided by the long-horizon decision process, we propose an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making. This taxonomy measures how agents understand and leverage historical patient information over extended horizons. Our experiments show that while recent LLMs can make good use of explicit timestamps, they have challenges in implicit time inference; The RAG and agent memory system can improve the performance of information retrieval tasks, but the performance of decision-making tasks is highly dependent on the model's immediate context.

2026-07-13 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SAGEAgent: マルチモーダル生存予測におけるコストを意識したモダリティ取得のための自己進化エージェント

すべてのがん患者は、正確な生存予測のために完全な診断精密検査を本当に必要としているのでしょうか?集学的臨床腫瘍学では、診療時に収集された人口統計から特殊な組織分析を必要とするゲノムプロファイリングまで、診断手段は臨床的に義務付けられた負担の増大の順序に従います。現在のマルチモーダル生存法は、すべてのモダリティが利用可能であると仮定するか、欠落データを受動的に処理するかのいずれかですが、この順序付けられたワークフローに沿って特定の患者に対して次のモダリティを取得することが正当化されるかどうかを積極的に推論するものはありません。我々はこれを逐次決定問題として定式化し、SAGEAgent (Sequential Acquisition Guided by Experience) を提案します。これは、各患者に対してどの診断モダリティを取得するかを決定し、予測精度と臨床侵襲性のバランスをとった自己進化型 LLM ベースの臨床エージェントです。 SAGEAgent は、数値予測をテキストに変換する臨床ツール、類似した過去の症例を検索するエピソード記憶、経験から再利用可能な決定パターンを蓄積する意味記憶を通じて、各患者の進化する診断状態を推論します。 TCGA-LGG、TCGA-GBM、および BraTS と 4 つの診断モダリティを組み合わせた神経膠腫コホートの実験では、SAGEAgent が平均取得負荷を 55% 削減しながら、競合する生存予測精度を達成することが実証されました。

原文 (English)

SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction

Does every cancer patient truly need a complete diagnostic workup for accurate survival prediction? In multimodal clinical oncology, diagnostic modalities follow a clinically mandated order of escalating burden -- from demographics collected at intake to genomic profiling requiring specialized tissue analysis. Current multimodal survival methods either assume all modalities are available or passively handle missing data, but none actively reason about whether acquiring the next modality is justified for a given patient along this ordered workflow. We formulate this as a sequential decision problem and propose SAGEAgent (Sequential Acquisition Guided by Experience), a self-evolving LLM-based clinical agent that decides which diagnostic modalities to acquire for each patient, balancing predictive accuracy against clinical invasiveness. SAGEAgent reasons about each patient's evolving diagnostic state through clinical tools that translate numerical predictions into text, an episodic memory that retrieves similar past cases, and a semantic memory that accumulates reusable decision patterns from experience. Experiments on a glioma cohort combining TCGA-LGG, TCGA-GBM, and BraTS with four diagnostic modalities demonstrate that SAGEAgent achieves competitive survival prediction accuracy while reducing average acquisition burden by 55%.

2026-07-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

NL-PAC: LLM 仲介監督における仕様の曖昧さと認定された Minimax リスクフロア

大規模な言語モデルでは、自然言語で指定されたタスクに対するラベル、評価、フィードバックが提供されることが増えています。仕様では複数の読み取り値が認められているが、どれが有効であるかが監視チャネルで明らかにされていない場合、ラベルを追加すると、結果として生じる識別問題は解決されずにサンプリング エラーが減少します。 Natural Language PAC (NL-PAC) を紹介します。これは、固定モデルの閾値デコード法則を使用して、許容可能なラベルと候補ターゲットを定義するフレームワークです。複数のラベルが許容される確率は、点ごとに許容されるターゲットクラスの直径に等しく、ターゲットブラインド監視下では、すべての学習者は、すべてのサンプルサイズで、少なくともこの直径の半分の最悪の場合のリスクを負います。このクラスに対する正確なランダム化されたミニマックス リスクは、データに依存しない戦略によって達成されます。有限サンプルの信頼限界により、保持されたラベルなしの入力からこれらの量が証明可能になります。凍結された Qwen~2.5--3B 監査では、事前に指定された 1 つのプロンプトからは肯定的なモデル相対証明書が得られますが、言い換えと正確なルールのコントロールからはゼロが得られます。開催されたブリッジ監査により、提供された候補読み取り条項が、証明書を一貫した読み取りに移行するために必要な許容条件を満たしていないことが判明しました。保証は、監査対象のモデル、プロンプト、しきい値、および入力分布に固有です。それを人間の解釈に拡張するには、外部の検証が必要です。

原文 (English)

NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision

Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets. The probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner incurs worst-case risk of at least half this diameter, at every sample size; the exact randomized minimax risk over this class is attained by a data-independent strategy. Finite-sample confidence bounds make these quantities certifiable from held-out unlabeled inputs. In a frozen Qwen~2.5--3B audit, one prespecified prompt yields a positive model-relative certificate, whereas a paraphrase and exact-rule controls yield zero. A held-out bridge audit finds that supplied candidate reading clauses fail the admissibility condition needed to transfer the certificate to coherent readings. The guarantee is specific to the audited model, prompt, threshold, and input distribution; extending it to human interpretations requires external validation.

2026-07-13 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

メタデータを超えて: 医用画像における欠落メタデータの下で隠れたサブグループ分析のための CAPRA

医用画像モデルは、多くの場合、サブグループ監査に必要な人口統計、取得、および品質のメタデータなしで導入されます。これらのメタデータが失われると、臨床的に重大な障害モードが強力な集合パフォーマンスによって隠蔽される可能性があり、多くのロバスト学習手法は依存するグループ構造を失います。欠落したメタデータの下で隠れたサブグループ分析を行うための調整されたプロキシ軸フレームワークである CAPRA を紹介します。 CAPRA は、画像由来のセマンティック軸を予測し、患者レベルのクロスフィッティングを介して小さなメタデータラベル付き分割で軸の事後軸を調整し、それらの事後軸を調整されたサブグループ インターフェイスに編成します。このインターフェイスは、展開時にサブグループ ラベルを必要とせずに、展開時の障害分析と下流の堅牢な学習の両方をサポートします。眼底、ダーモスコピー、および胸部 X 線撮影全体にわたって、CAPRA はメタデータのみのスライスでは見逃された視差パターンを明らかにし、データセット シフトの下でも有益な情報を維持し、画像のみまたは潜在スライスのベースラインよりも明示的な破損軸とより密接に一致するサブグループ パーティションを生成します。同じインターフェイスを下流の堅牢な学習器で再利用することもできますが、その利得はドメインに依存します。全体として、CAPRA は、メタデータが欠落している隠れたサブグループ分析を、展開時の分析と堅牢な転送のために、調整された解釈可能で再利用可能なサブグループ インターフェイスに変換します。

原文 (English)

Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging

Medical imaging models are often deployed without the demographic, acquisition, and quality metadata needed for subgroup auditing. Once those metadata disappear, clinically critical failure modes can be masked by strong aggregate performance, and many robust-learning methods lose the group structure they rely on. We present CAPRA, a calibrated proxy-axis framework for hidden subgroup analysis under missing metadata. CAPRA predicts image-derived semantic axes, calibrates axis posteriors on a small metadata-labeled split via patient-level cross-fitting, and organizes those posteriors into a calibrated subgroup interface that supports both deployment-time failure analysis and downstream robust learning without requiring subgroup labels at deployment. Across fundus, dermoscopy, and chest radiography, CAPRA reveals disparity patterns missed by metadata-only slicing, remains informative under dataset shift, and produces subgroup partitions that align more closely with explicit failure axes than image-only or latent-slice baselines. The same interface can also be reused by downstream robust learners, although those gains are domain-dependent. Overall, CAPRA turns hidden subgroup analysis under missing metadata into a calibrated, interpretable, and reusable subgroup interface for deployment-time analysis and robust transfer.

2026-07-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

欺瞞的なグラウンディング: 臨床検索拡張生成におけるエンティティ帰属の失敗

検索拡張生成評価では、モデルの主張が検索された文書に事実に基づいているかどうかをチェックします。取得した証拠が正しいエンティティに帰属するかどうかはチェックしません。臨床 RAG 応答は、薬物 Y の臨床証拠を、質問された薬物 X に関する証拠として提示しながら、すべての自動チェック (幻覚ゼロ、ほぼ完璧な忠実度、実際の引用) に合格することができます。私たちは、これを欺瞞的根拠 (DG) と呼びます。これは、すべての主張が間違った実体に関する実際の文書から出ているため、忠実度、幻覚、および引用のチェックには見えない失敗です。 13 のモデルにわたって制御された要因ベンチマークを使用すると、ピークの敵対条件で DG 率が 8 ~ 87% の範囲にあることがわかります。医療および生物医学の微調整されたモデルは最大 86.7% に達します。ドメインの特殊化は障害を軽減するのではなく、むしろ増幅させます。制御されたアブレーションによりメカニズムが特定されます。取得された文書から実体固有の臨床証拠を削除すると、実体帰属の失敗が完全に排除され、すべての失敗が作話に移行します。 2 つの障害モードは同じトリガーに応答し、異なるパスをたどります。 740 の薬物と疾患のペアにわたる生産測定では、展開された RAG システムの全体的な DG が 7.8% であることがわかり、最近承認された薬剤では 13.6% に上昇しました。エンティティ帰属検証 (引用された証拠が照会されたエンティティに適用されることを確認する) では、97.0% の精度と 98.7% の DG 再現率 (IPW 調整済みヒューマン ゴールド スタンダード) で DG が検出されます。これを実装している既存のフレームワークはありません。

原文 (English)

Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation

Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents. It does not check whether retrieved evidence is attributed to the correct entity. A clinical RAG response can pass every automated check (zero hallucinations, near-perfect faithfulness, real citations) while presenting drug Y's clinical evidence as evidence about queried drug X. We term this deceptive grounding (DG): a failure invisible to faithfulness, hallucination, and citation checks because every claim is sourced from a real document, about the wrong entity. Using a controlled factorial benchmark across 13 models, we find DG rates spanning 8-87% at peak adversarial conditions. Medical and biomedical fine-tuned models reach up to 86.7%; domain specialization amplifies the failure rather than mitigating it. A controlled ablation identifies the mechanism: removing entity-specific clinical evidence from retrieved documents eliminates entity-attribution failure entirely, shifting all failures to confabulation. The two failure modes respond to the same trigger, taking different paths. Production measurement across 740 drug-disease pairs finds 7.8% overall DG in a deployed RAG system, rising to 13.6% for recently approved drugs. Entity-attribution verification (checking that cited evidence applies to the queried entity) detects DG at 97.0% precision and 98.7% DG recall (IPW-adjusted human gold standard); no existing framework implements it.

2026-07-13 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple s…

2026-07-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs

Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles…

2026-07-13 13:00 JSTarXiv cs.AIビジネス/資金調達

Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation

Supervised detection of network attacks has always been a critical part of network intrusion detection systems (NIDS). Nowadays, in a pivot…

2026-07-13 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM は医師を支援する準備ができていますか?インタラクティブな医師、患者、EHR 支援のための PhysAssistBench

医療 LLM の最も妥当な短期的な役割は、医師の代わりではなく支援することですが、現在の評価では、臨床知識、EHR システムの相互作用、患者とのコミュニケーションなど、個別の能力がテストされることがよくあります。代わりに、医師の支援には同じ対話内でこれらの機能を調整する必要があり、医師は不明確な要求を発行し、患者は症状を曖昧に説明し、EHR システムはツールの正確な使用を要求します。インタラクティブな医師、患者、EHR 支援のベンチマークである PhysAssistBench を紹介します。実際の MIMIC-IV 症例から構築された PhysAssistBench は、スケーラブルなパイプラインを使用してエージェント性患者を構築します。これは、臨床上の事実を維持しながら、静的な EHR 記録を複数ターンの臨床シナリオに変換する、インタラクティブで記録に基づいたエージェントです。 PhysAssistBench は、手動でレビューされ医師が検証した 1,296 ターンの厳選されたバイリンガル評価セットを提供します。主要な LLM を使った実験では、この設定では現在のモデルの信頼性が依然として低いことが示されており、臨床 LLM にとって重要なボトルネックが露呈しています。信頼できる支援には、知識、コミュニケーション、システム全体の調整が必要であり、それらのいずれかで単独の利益を得るのではありません。

原文 (English)

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.

2026-07-11 02:17 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

SK Hynix raises $26.5B in the biggest foreign IPO in US history, is urged to build new US fabs

The AI chip boom just produced its biggest Wall Street moment yet. Now SK Hynix and Samsung are being asked to build U.S. factories.

2026-07-10 13:00 JSTarXiv cs.AIビジネス/資金調達

AI評価に欠けている要素としての心理的能力

現在の AI 評価フレームワークは、精度、堅牢性、推論能力、ポリシー遵守などの技術的パフォーマンスに主に焦点を当てています。これらの対策は依然として不可欠ですが、自然言語を通じてユーザーと直接対話するシステムにとっては十分ではありません。人間と向き合う AI システムは、アドバイザー、コーチ、家庭教師、仲間として使用されることが増えています。これらの役割では、ユーザーの応答によって、ユーザーがどのように推論し、感情を解釈し、信念を形成し、信頼を調整し、意思決定を行うかが形成されます。したがって、関連する評価単位はモデルだけではなく、人間と AI の相互作用です。この論文では、AI 評価に欠けている側面として心理的能力を紹介します。私たちは、心理的能力を、ユーザー、状況、インタラクションの目的に適切な方法でユーザーの認知、感情の解釈、行動の意思決定をサポートする対人 AI システムの能力として定義します。これには、フレーミング、トーン、知覚される権威、反応性、不確実性の処理、会話のガイダンスなどの対話特性が含まれます。既存の評価アプローチはこの問題の一部を捉えていますが、これらの心理的影響を直接評価することはほとんどありません。行動科学と人間と AI の相互作用研究に基づいて、心理的能力とその中核領域の概念的枠組みを概説します。特定のベンチマークを提案するのではなく、構成を定義し、その境界を明確にし、シナリオベースの調査、構造化された人間による評価、およびモデル支援の評価方法を通じてそれがどのように評価されるかを説明します。私たちは、心理的能力が、人間と対面する AI システムの実世界への影響を懸念するモデル提供者、導入組織、研究者、規制当局にとって中心的な考慮事項となるべきであると主張します。

原文 (English)

Psychological Competence as a Missing Dimension in AI Evaluation

Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. These measures remain essential, but they are not sufficient for systems that interact directly with users through natural language. Human-facing AI systems are increasingly used as advisors, coaches, tutors, and companions. In these roles, their responses can shape how users reason, interpret emotions, form beliefs, calibrate trust, and make decisions. The relevant unit of evaluation is therefore not only the model, but the human-AI interaction. This paper introduces psychological competence as a missing dimension in AI evaluation. We define psychological competence as the capacity of a human-facing AI system to support user cognition, emotional interpretation, and behavioral decision-making in ways that are appropriate to the user, context, and purpose of the interaction. This includes interaction properties such as framing, tone, perceived authority, responsiveness, uncertainty handling, and conversational guidance. Existing evaluation approaches capture parts of this problem but rarely assess these psychological effects directly. Drawing on behavioral science and human-AI interaction research, we outline a conceptual framework for psychological competence and its core domains. Rather than proposing a specific benchmark, we define the construct, clarify its boundaries, and describe how it may be assessed through scenario-based probes, structured human evaluation, and model-assisted evaluation methods. We argue that psychological competence should become a core consideration for model providers, deploying organizations, researchers, and regulators concerned with the real-world effects of human-facing AI systems.

2026-07-10 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

SolarChain-Eval: 分散型エネルギー市場の信頼できる経済主体のための物理制約付きベンチマーク

エージェント AI システムがサイバー物理環境に適用されることが増えているため、その評価にはタスクのパフォーマンスと信頼性の両方の評価が必要です。分散型エネルギー市場では、自律エージェントは市場の効用を向上させる可能性がありますが、無効な物理データを悪用し、人為的な流動性を生み出し、不安定なガバナンスの決定を生み出す可能性もあります。したがって、信頼できる経済主体を評価するための物理制約付きベンチマークである SolarChain-Eval を提案します。これは市場ガバナンスをギムナジウム互換のマルコフ決定プロセスとして定式化し、エージェントが時間ごとに意思決定を行います。 SolarChain-Eval は、市場の有用性、物理的安全性、スリッページ、アクションのスムーズさ、空間的公平性、監査可能性など、複数の側面にわたって各ポリシーを評価します。エージェント評価をサポートするために、SolarChain-Eval には LLM ベースの Planner/Auditor 層が組み込まれています。プランナーはエピソードレベルのアクションの制限と監査ルールを定義し、監査人はリスクの高いアクションをレビューして修正します。すべての介入は、トリガー信号、提案されたアクション、修正されたアクション、および監査の根拠を含む構造化されたログを通じて記録されます。静的、ランダム、近視、RL、および RL+LLM ポリシーを使用した実験では、ユーティリティと安全性の明確なトレードオフが明らかになりました。 RL エージェントは市場の有用性を向上させますが、依然として危険な動作を引き起こす可能性があります。物理的ペナルティが除去されると、報酬最大化エージェントは無効な生成を利用し、人為的な流動性を増加させます。 LLM プランナー/監査人は監査可能性を向上させ、選択されたリスクを軽減しますが、誤って指定された報酬関数を完全に補償することはできません。これらの結果は、信頼できるエージェント AI 評価には物理的制約と透過的な介入トレースの両方が必要であることを示しています。複製可能性を確保するために、データとコードを GitHub 上でオープン アクセスとしてリリースします。

原文 (English)

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve market utility, but may also exploit invalid physical data, create artificial liquidity, and produce unstable governance decisions. Therefore, we propose SolarChain-Eval, a physics-constrained benchmark for evaluating trustworthy economic agents. It formulates market governance as a Gymnasium-compatible Markov Decision Process, where agents make hourly decisions. SolarChain-Eval evaluates each policy across multiple dimensions, including market utility, physical safety, slippage, action smoothness, spatial fairness, and auditability. To support agentic evaluation, SolarChain-Eval incorporates an LLM-based Planner/Auditor layer. The Planner defines episode-level action bounds and audit rules, while the Auditor reviews and revises high-risk actions. All interventions are recorded through structured logs, including trigger signals, proposed actions, revised actions, and audit rationales. Experiments with static, random, myopic, RL, and RL+LLM policies reveal a clear utility-safety trade-off. RL agents improve market utility but can still produce unsafe behavior. When the physics penalty is removed, reward-maximizing agents exploit invalid generation and increase artificial liquidity. The LLM Planner/Auditor improves auditability and mitigates selected risks, but it cannot fully compensate for a misspecified reward function. These results indicate that trustworthy agentic AI evaluation requires both physical constraints and transparent intervention traces. We release data and code as open access on GitHub for replicability.

2026-07-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

等価性の幻想: LLM における量子化効果の統計的特徴付け

トレーニング後の量子化は、リソースに制約のある設定で大規模な言語モデルをデプロイするために広く使用されていますが、その評価はほぼもっぱら精度と複雑さに依存します。これらの指標では、量子化によって引き起こされる行動の変化を捉えることができないことを示します。絶対精度とは無関係に、基本モデルとその量子化されたバリアント間の正しい予測の重複を測定する意思決定レベルの指標である正確性一致を導入します。 8 ビットから 2 ビットまでの複数のモデルと量子化スキームにわたって、タスクのパフォーマンスが維持されているように見える場合でも、適度な量子化の下では動作の発散が現れることがわかりました。この効果を説明するために、注意の重みに対する構造演算子として量子化を分析し、統計的および分布的尺度を使用して層ごとの歪みを定量化します。私たちの結果は、低ビット幅での非線形ブレークポイントを明らかにし、クエリとキーの投影が値と出力の投影よりも常に敏感であることを示しています。これらの発見は、基本モデルと量子化モデルの間の等価性の幻想を明らかにし、従来のパフォーマンス指標を超えた行動評価を動機付けます。

原文 (English)

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization schemes from 8-bit to 2-bit, we find that behavioral divergence emerges under moderate quantization even when task performance appears preserved. To explain this effect, we analyze quantization as a structural operator on attention weights and quantify layer-wise distortions using statistical and distributional measures. Our results reveal non-linear breakpoints at low bit-widths and show that query and key projections are consistently more sensitive than value and output projections. These findings expose an illusion of equivalence between base and quantized models and motivate behavioral evaluation beyond conventional performance metrics.

2026-07-10 13:00 JSTarXiv cs.AIビジネス/資金調達

深層強化学習の評価と設計パラダイムの原理分析

最も挑戦的なゲームの 1 つに勝利をもたらしたステートアクション価値関数を近似するためのディープ ニューラル ネットワークの利用から始まり、当面の課題のルールを明示することなく問題を解決できるアルゴリズムの進歩に至るまで、強化学習研究は過去 10 年間、目覚ましい科学進歩の中心となってきました。この論文では、この研究の進歩の主要な要素に焦点を当て、強化学習における標準評価と設計パラダイムを分析します。我々は、強化学習におけるスケーリング則の理論的基礎を導入し、強化学習アルゴリズムの漸近的なパフォーマンスが、パフォーマンスのランキングとデータ領域の間に単調な関係を持たないことを示します。私たちは大規模な実験を行っており、その結果は、標準的な設計および評価パラダイムに基づく一連の強化学習研究が誤った結論をもたらしていることを示しています。私たちの分析と結果は、深層強化学習のスケーリング、容量、複雑さに関する中心的な分析を提供します。

原文 (English)

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand, reinforcement learning research has been the center of remarkable scientific progress for the past decade. In this paper, we focus on the key ingredients of this research progress and we analyze the canonical evaluation and design paradigms in reinforcement learning. We introduce the theoretical foundations of scaling laws in reinforcement learning and show that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes. We conduct large-scale experiments and our results demonstrate that a line of reinforcement learning research under the canonical design and evaluation paradigms resulted in incorrect conclusions. Our analysis and results provide a core analysis on scaling, capacity and complexity of deep reinforcement learning.

2026-07-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment

Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an automatic spe…

2026-07-10 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. W…

2026-07-10 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SimRPD: Optimizing Recruitment Proactive Dialogue Agents through Simulator-Based Data Evaluation and Selection

Task-oriented proactive dialogue agents play a pivotal role in recruitment, particularly for steering conversations towards specific busine…

2026-07-10 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately recon…

2026-07-10 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

TOPO-Bench: An Open-Source Topological Mapping Evaluation Framework with Quantifiable Perceptual Aliasing

Topological mapping offers a compact and robust representation for navigation, but progress in the field is hindered by the lack of standar…

2026-07-10 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation

Vision-Language Models (VLMs) are increasingly used in document processing pipelines to convert flowchart images into structured code (e.g.…

2026-07-10 13:00 JSTarXiv cs.AIビジネス/資金調達

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming

RLHF evaluation should track how failures emerge, where they localize, and which warning signals appear before external quality degrades. W…

2026-07-10 07:08 JSTTechCrunch AIエージェントビジネス/資金調達

An AI agent startup just let its agent run its $100M fundraise

Lyzr, a startup that builds AI agents for enterprises, used its own AI agent to raise a $100 million round — proof, evidently, that the pro…

2026-07-10 06:57 JSTTechCrunch AILLM/生成AIビジネス/資金調達

Elon Musk praises Mythos/Fable, promises not to ‘cut off’ Anthropic

Should Anthropic trust Elon Musk to host its models? With about $40 billion in revenue at stake, Musk insists that the company can.

2026-07-10 03:34 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

Paris-based AI voice startup Gradium raises $100M seed, backed by Nvidia

The company is using the cash to open an office in the Bay Area and compete for talent there, "strengthening its position at the heart of t…

2026-07-09 23:51 JSTTechCrunch AILLM/生成AIビジネス/資金調達

Anthropic, OpenAI, and SpaceX are bigger than the last 25 years of tech exits

Three big AI IPOs are set to generate more value than all the U.S. VC-backed exits since 2000.

2026-07-09 22:00 JSTTechCrunch AILLM/生成AIビジネス/資金調達研究/論文

Popular open source AI developer tool Ollama raises $65M, grows to nearly 9M users

Benchmark-backed Ollama has amassed 176,000 stars, and nearly 17,000 forks on GitHub by helping developers easily run AI on their PCs.

2026-07-09 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

AgentLens: コーディング エージェント評価のための実稼働環境で評価された軌跡レビュー

ここでは、対話型コード エージェントの実稼働環境で評価されたベンチマークである AgentLens を紹介します。ほとんどのコード エージェント ベンチマークでは、実行が 1 ビットに削減されます。タスクは成功しましたか? -- しかし、これらのエージェントを実際に使用する人々は、エージェントがどのように指示に従い、ツールを使用し、自身の作業を検証し、間違いから回復し、途中でエージェントに話しかけるかという軌跡全体を経験します。 AgentLens はその軌跡全体を評価します。客観的なチェックが存在する正式な検証と、LLM で作成された軌跡のレビューおよび並べての比較を組み合わせることで、各実行でスコアがなぜそのようになるのかについての読みやすい説明が得られます。これにより、AgentLens はモデルのランク付け以上の用途に役立ちます。モデルの動作を診断し、独自のエージェントの連続バージョンを比較し、夜間の評価パイプラインで製品の回帰を捕捉するために使用されます。 https://github.com/agent-lens/agent-lens-bench でベンチマークをオープンソースとしてリリースします。

原文 (English)

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at https://github.com/agent-lens/agent-lens-bench.

2026-07-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

マルチエージェント LLM の安全性における運用の再構築と承認枠付きの委任

マルチエージェント LLM システムの安全性評価では、多くの場合、直接プロンプトとプランナーと実行者のパイプラインを比較し、その違いを単一の「パイプライン効果」として報告します。私たちは、この集合体は 3 つのメカニズムを混同しているため、解釈が難しいと主張します。有害な意図がもっともらしい運用作業として再構成される可能性があること、計画立案者が要求を拒否または変更する可能性があること、実行者が事前の承認を示唆する委任プロンプトに基づいて行動する可能性があることです。これらの要因を分離するために、LLM が判断したコンプライアンスを使用した 30 の合成有害シナリオと 4 つのエージェントの安全性ベンチマークからの探索的な外部検証セットで評価された 5 つの条件で制御されたコントラスト設計を導入します。私たちの結果は、集合的なパイプラインの安全性が安定したアーキテクチャ上の特性ではないことを示しています。オペレーショナル リフレーミングは最もポータブルなリスク シグナルであり、両方のシナリオ セットにわたって GPT、Gemini、DeepSeek のコンプライアンスを強化しますが、Claude は比較的抵抗力があります。プランナーの行動は、主に拒否を通じてこのリスクを相殺できます。ただし、プランナが実行可能なステップを作成すると、実行者は直接の運用ベースラインよりも準拠するようになる可能性があります。承認フレームワークの委任は、プロンプトの設計、モデルのペアリング、およびシナリオのソースに敏感であり、懐疑的な実行者のプロンプトはコンプライアンスを大幅に低下させます。生の直接モデルのランキングも、デプロイされたプランナーと実行者の動作を誤って予測する可能性があります。ジェミニは、プライマリ セットの生の直接プロンプトの下で最も安全ですが、クロード プランナーを使用すると最大の増幅を示し、コンプライアンスが 8.9 パーセントから 38.9 パーセントに上昇しました。 GPT の集約パイプライン効果がゼロに近いため、代わりにプランナーの拒否によってキャンセルされたリフレーミングの増加が隠蔽されます。これらの発見は、マルチエージェントの安全性評価では、失敗をアーキテクチャ自体のせいにする前に、リフレーミング、プランナーの動作、委任のフレーミング、およびモデルのペアリングを個別に報告する必要があることを示唆しています。

原文 (English)

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it conflates three mechanisms: harmful intent may be reframed as plausible operational work, the planner may refuse or transform the request, and the executor may act under delegation prompts implying prior approval. To separate these factors, we introduce a five-condition controlled contrast design, evaluated on 30 synthetic harmful scenarios and an exploratory external validation set from four agent-safety benchmarks using LLM-judged compliance. Our results show that aggregate pipeline safety is not a stable architectural property. Operational reframing is the most portable risk signal, increasing compliance for GPT, Gemini, and DeepSeek across both scenario sets, while Claude is comparatively resistant. Planner behavior can offset this risk mainly through refusal; however, when the planner produces executable steps, the executor may become more compliant than under the direct operational baseline. Approval-framed delegation is sensitive to prompt design, model pairing, and scenario source, and a skeptical executor prompt sharply reduces compliance. Raw-direct model rankings can also mispredict deployed planner-executor behavior. Gemini is safest under raw direct prompts in the primary set yet shows the largest amplification with a Claude planner, rising from 8.9 percent to 38.9 percent compliance. GPTs near-zero aggregate pipeline effect instead hides a reframing increase canceled by planner refusal. These findings suggest that multi-agent safety evaluations should report reframing, planner behavior, delegation framing, and model pairing separately before attributing failures to architecture itself.

2026-07-09 13:00 JSTarXiv cs.AIビジネス/資金調達

推論一貫性スキャン: AI の安全性評価における思考連鎖の妥当性を監査するためのフレームワーク

これまでの研究では、思考連鎖 (CoT) 推論がしばしば不忠実であることが示されています。つまり、モデルで述べられた推論は、その出力を生成したプロセスを確実に反映していません。ただし、不貞を検出するには、管理された実験的介入が必要であり、事後の評価記録にそれを適用することはできません。代わりに、あまり注目されていない、より扱いやすい質問、つまり、述べられた推論がそれに伴う答えと論理的に一貫しているかどうかに目を向けます。忠実度とは異なり、一貫性は介入なしで記録のみから評価できます。 AI 安全性評価トランスクリプトでこの特性を検出するための再利用可能な方法である推論一貫性スキャンを紹介します。私たちの貢献は 4 つあります。まず、推論の一貫性を忠実性とは異なるものとして形式化し、不一致の 6 つのサブタイプ分類を定義します。次に、InstrumentalEval の出力から手動で調整した 60 個のトランスクリプトの検証済みベンチマークを構築します。 3 番目に、InspectScout 用に動作するスキャナーを実装します。これは、安全性評価記録でこのプロパティをターゲットにした最初のスキャナーです。 4 番目に、4 つのジェネレーター モデルと、inspect_evals からの 3 つの評価にわたる結果を報告します。これは、推論の矛盾が存在し、検出可能であり、モデルとタスク タイプの両方にわたって体系的に変化していることを示しています。

原文 (English)

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that produced its output. Detecting unfaithfulness, though, requires controlled experimental interventions, which cannot be applied to evaluation transcripts after the fact. We turn instead to a more tractable question that has received less attention: whether the stated reasoning is logically consistent with the answer it accompanies. Unlike faithfulness, consistency can be assessed from a transcript alone, with no intervention. We introduce reasoning consistency scanning, a reusable method for detecting this property in AI safety evaluation transcripts. Our contributions are fourfold. First, we formalize reasoning consistency as distinct from faithfulness and define a six-subtype taxonomy of inconsistency. Second, we build a validated benchmark of 60 transcripts, manually adapted from InstrumentalEval outputs. Third, we implement a working scanner for InspectScout, the first to target this property in safety evaluation transcripts. Fourth, we report results across four generator models and three evaluations from inspect_evals, showing that reasoning inconsistency is present, detectable, and varies systematically across both models and task types.

2026-07-09 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

組織的なレッドチーム化: モデルだけでなく展開ルールもマルチエージェント AI の安全性を因果的に形成する

マルチエージェント AI の導入ルールをテストするための評価方法である組織的レッドチーム化を紹介します。つまり、エージェント、目的、タスクの状態を固定し、ルールを 1 つだけ変更し、結果として生じる集団行動の変化はそのルールによるものだと考えます。私たちは、規範的な協調参照と自動ラベル付けされた推論トレースを使用して、228 のコンテキスト、5 つの正規ルール、7 つのモデル母集団 (33,924 ゲーム) にわたる結果配分ベンチマークである IABench-CA で方法論をインスタンス化します。 3 つの発見が得られます。 (1) 配備ルールは集団の安全を因果的に変える。結果ルールのみを変更すると、各集団内の平均死亡率が 22 ~ 58 パーセントポイント上昇する。 (2) 安全なデフォルトはありませんが、ターゲティングの危険性は普遍的です。最も安全なルール、最も安全でないルール、さらには発生率効果の方向さえも集団によって異なりますが、回帰的アイデンティティターゲティングは、どの集団のどのような状況においても決定的に最も安全であることはなく、どこのゲームでも 30 ~ 87% のリソースが最も少ないエージェントを排除し、7 つの集団すべてについて協調参照と比較して選択が安全ではありません。 (3) アイデンティティの顕著性はメカニズムです。最も搾取されやすい集団に対するワンショットの匿名化アブレーション (gpt-5.1) は、ルール テキストで損失負担者を指定するだけで、同一の見返りで目標の排除が 22% から 81% に促進されることを示しています。繰り返しのプレイでは、エージェントが観察された排除から隠されたルールを再推測するため、匿名化はターゲティングを遅らせるだけです。この方法論を、明示的な残留リスクと監視義務を伴う、デプロイメントコンテキストおよび母集団ごとの暫定ルール領域 $\Phi(c,P)$ を認証するセーフティケースのワークフローとしてパッケージ化します。

原文 (English)

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-CA, a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces. Three findings emerge. (1) Deployment rules causally alter collective safety: changing only the consequence rule moves mean fatality by 22 to 58 percentage points within every population. (2) There is no safe default, but the targeting hazard is universal: the safest rule, the least-safe rule, and even the direction of the incidence effect vary across populations, yet regressive identity-targeting is never decisively safest in any context for any population, eliminates the least-resourced agent in 30-87% of games everywhere, and is selection-unsafe relative to the cooperative reference for all seven populations. (3) Identity salience is the mechanism: a one-shot anonymization ablation on the most exploitation-prone population (gpt-5.1) shows that merely naming the loss bearer in the rule text drives targeted elimination from 22% to 81% at identical payoffs; under repeated play, anonymization only delays the targeting, as agents re-infer the hidden rule from observed eliminations. We package the methodology as a safety-case workflow that certifies a provisional rule region $\Phi(c,P)$ per deployment context and population, with explicit residual risks and monitoring obligations.

2026-07-09 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

NLPCC 2026 の概要 共有タスク 1: 難易度を考慮した多言語およびマルチモーダルな医療指導ビデオの理解度評価

NLPCC 2023 ~ 2025 年の CMIVQA、MMI-VQA、および M4IVQA の課題に続き、NLPCC 2026 では難易度を考慮した医療指導ビデオ質問応答 (DA-MIVQA) 共有タスクを導入します。DA-MIVQA は、必要な証拠の種類と複雑さに応じて質問を明示的に区別することで、以前の多言語および多モードの医療ビデオ ベンチマークを拡張します。答える。具体的には、単純な質問は字幕ベースのテキストの手がかりから答えられることがよくありますが、複雑な質問には視覚的な根拠、手順の理解、およびクロスモーダルな証拠の統合が必要です。このチャレンジには、単一ビデオでの難易度を考慮した時間的回答グラウンディング (DA-TAGSV)、ビデオ コーパスでの難易度を考慮した時間的回答の取得 (DA-VCR)、およびビデオ コーパスでの難易度を考慮した時間的回答のグラウンディング (DA-TAGVC) の 3 つのトラックが含まれています。データセットは公的医療指導チャンネルから収集され、応急処置、緊急対応、リハビリテーション、看護、一般医学教育などの多様なシナリオをカバーしており、難易度の注釈を付けて手動で検証されています。本稿では、DA-MIVQAの課題動機、データセット構築、評価プロトコル、参加概要、競技結果、代表的なシステムについて紹介する。 DA-MIVQA は、さまざまなテキスト、視覚、時間的、および手順の推論要件の下で、医療指導ビデオ質問応答システムを評価するための実用的なベンチマークを提供します。

原文 (English)

Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation

Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026. DA-MIVQA extends previous multilingual and multimodal medical video benchmarks by explicitly distinguishing questions according to the type and complexity of evidence required for answering. Specifically, simple questions can often be answered from subtitle-based textual cues, whereas complex questions require visual grounding, procedural understanding, and cross-modal evidence integration. The challenge contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV), Difficulty-Aware Video Corpus Retrieval (DA-VCR), and Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The dataset is collected from public medical instructional channels, covers diverse scenarios such as first aid, emergency response, rehabilitation, nursing, and general medical education, and is manually verified with difficulty annotations. This paper presents the task motivation, dataset construction, evaluation protocol, participation overview, competition results, and representative systems of DA-MIVQA. DA-MIVQA provides a practical benchmark for evaluating medical instructional video question answering systems under varying textual, visual, temporal, and procedural reasoning requirements.

2026-07-09 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

ソフトウェアエンジニアリング用エージェントの信頼性の高い、開発者に合わせた評価

大規模な言語モデルは、開発サイクルの終了に向けて急速に進んでおり、単純な補助的なコンパニオンから、共同開発環境に深く組み込まれた自律的な貢献者へと移行しています。導入が加速しているにもかかわらず、既存の評価手法は、その断片的な性質と、多くの場合、仮説的な構文シナリオから得られる真のモデル機能の歪んだ投影により、限界があります。この研究は、現実のソフトウェア開発実践に基づいた LLM を利用したエージェントの包括的な評価方法を提供することで、このギャップを埋めることを目的としています。当社の評価アプローチは、汚染認識、実際のエージェントの動作評価、現実的なコーディング コンテキスト、人間に合わせた動作、モデルの故障モードをキャプチャする軌道を認識したベンチマークとメトリクスに焦点を当てています。

原文 (English)

Reliable and Developer-Aligned Evaluation of Agents for Software Engineering

Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments. Despite their accelerated adoption, existing evaluation techniques are limited due to their fragmented nature and distorted projection of true model capabilities, often obtained from hypothetical syntactic scenarios. This research aims to bridge this gap by providing a comprehensive evaluation methodology for LLM-powered agents that is grounded in real-world software development practice. Our evaluation approach focuses on contamination-awareness, in-the-wild agentic behavior assessment, and trajectory-aware benchmarks and metrics capturing realistic coding contexts, human-aligned behavior, and model failure modes.

2026-07-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

大規模言語モデルの応答の包括的な評価: 多要素スコアリング システム

言語タスクにおける大規模言語モデル (LLM) の顕著なパフォーマンスは、応答品質の包括的な評価が緊急に必要であることを強調しています。一般的な手法は、特異な次元に限定されることが多く、モデルの機能の全領域を捉えるには至っていません。この研究では、精度、簡潔さ、事実の一貫性、読みやすさ、一貫性を統合した多要素スコアリング パラダイムを導入し、結果を視覚化するためのグラフィカル ユーザー インターフェイス (GUI) によって補完されています。 TruthfulQA データセットの評価では、複雑な事実や曖昧さを回避する際の広範な制限とともに、推論タスクにおける主流の LLM の強み (複合スコア 0.6104 でピーク) が明らかになりました。このフレームワークは、従来のメトリクスの狭いレンズを超えて、モデルの可能性と欠陥を明らかにするための透明性と適応性のある手段を提供します。現在は英語のタスクに焦点を当てていますが、その視野は多言語の領域に向かっています。この研究は、知識エンジニアリングとモデルの改良に新たな道を切り開きます。

原文 (English)

Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities. This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes. Evaluations on the TruthfulQA dataset unveil mainstream LLMs' strengths in reasoning tasks (peaking at a composite score of 0.6104) alongside pervasive limitations in navigating complex facts and ambiguities. Transcending the narrow lens of traditional metrics, this framework offers a transparent, adaptable avenue to illuminate model potential and deficiencies. Though presently focused on English tasks, its horizons beckon toward multilingual domains. This work carves a novel path for knowledge engineering and model refinement.

2026-07-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Predicting LLM Safety Before Release by Simulating Deployment

Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evid…

2026-07-09 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation

Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation…

2026-07-09 13:00 JSTarXiv cs.AIエージェントロボティクスビジネス/資金調達

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately sy…

2026-07-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-…

2026-07-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for me…

2026-07-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating…

2026-07-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation

Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions. Whether they evalua…

2026-07-09 13:00 JSTarXiv cs.AIビジネス/資金調達

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.…

2026-07-09 13:00 JSTarXiv cs.AI画像/動画生成ロボティクスビジネス/資金調達研究/論文

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their cap…

2026-07-09 07:41 JSTTechCrunch AIビジネス/資金調達

Lovable reportedly in talks to double its valuation to $13.2B

The $300 million round is expected to be led by Menlo Ventures, Sifted reported.

2026-07-09 01:22 JSTTechCrunch AIエージェントビジネス/資金調達

Prime Intellect raises $130M Series A to help enterprises build their own AI agents

Founded in 2024, Prime Intellect’s goal is to give organizations capabilities to train their own agentic systems without relying on frontie…

2026-07-08 22:00 JSTOpenAILLM/生成AIビジネス/資金調達研究/論文

Separating signal from noise in coding evaluations

A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in…

2026-07-08 16:16 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

AI chip maker SambaNova raises $1B at $11B valuation, 5 months after last mega round

AI chip maker SambaNova has raised at an $11 billion valuation months after Intel was rumored to be trying to buy it for about $1.6 billion.

2026-07-08 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

CSTutorBench: ブロックベース プログラミングの家庭教師としての小規模言語モデルのベンチマーク

大規模な言語モデルは AI の家庭教師としてますます検討されていますが、幼稚園から高校までの環境に導入すると、プライバシー、コスト、独自のモデルへの依存に関する懸念が生じます。小規模言語モデル (SLM) は有望な代替手段を提供しますが、特定の教育コンテキストに適切なモデルを選択することは依然として困難であり、特にブロックベースのプログラミングなどのターゲット ドメインがモデル トレーニング データにほとんど含まれていない場合には困難です。ブロックベースのロボット環境である VEX VR で CS チューターとして言語モデルを評価するためのベンチマークである CSTutorBench を紹介します。このベンチマークは、確立された個別指導とフィードバック研究に基づいた教育ルーブリックに基づいて採点された 17 のシナリオベースの質問で構成され、評価には人間参加型の LLM による審査員パイプラインが使用されます。 11 のモデル (4B ~ 120B パラメーター) にわたる予備調査結果では、モデルは語彙や口調などの表面レベルの基準では良好に機能しますが、より深い教育的行動、特に答え漏れの回避や生徒のデバッグ履歴への関与に苦戦していることが明らかになりました。私たちのサンプルでは、​​モデルの数が少ないため、この結論の強度が制限されますが、モデルファミリーと命令チューニングのアプローチは、パラメーター数だけよりも個別指導の質をより良く予測するものであるように見えます。最近の教育プロンプト工学研究に基づいた、的を絞ったプロンプト改訂により、11 モデル中 10 モデルのスコアが向上しました。これらの結果は、教育展開における SLM 選択のための、状況に応じた教育学的に根拠のあるベンチマークの価値を強調しています。

原文 (English)

CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming

Large language models are increasingly explored as AI tutors, yet deploying them in K-12 settings raises concerns around privacy, cost, and reliance on proprietary models. Small language models (SLMs) offer a promising alternative, but selecting the right model for a specific educational context remains difficult, particularly when the target domain, such as block-based programming, is largely absent from model training data. We introduce CSTutorBench, a benchmark for evaluating language models as CS tutors in VEX VR, a block-based robotics environment. The benchmark comprises 17 scenario-based questions scored against a pedagogical rubric grounded in established tutoring and feedback research, with a human-in-the-loop LLM-as-judge pipeline for evaluation. Preliminary findings across 11 models (4B-120B parameters) reveal that models perform well on surface-level criteria such as vocabulary and tone but struggle with deeper pedagogical behaviors, particularly avoiding answer leakage and engaging with student debugging histories. In our sample, model family and instruction-tuning approach appear to be better predictors of tutoring quality than parameter count alone, though the small number of models limits the strength of this conclusion. A targeted prompt revision grounded in recent educational prompt engineering research improved scores for 10 of 11 models. These results underscore the value of context-specific, pedagogically grounded benchmarks for SLM selection in educational deployment.

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

静的評価を超えて: スケーラブルなエージェント強化学習のためのシミュレーション環境の構築

大規模言語モデル (LLM) が自律エージェントに進化するにつれて、従来の静的評価では複数段階の意思決定を捉えることができなくなります。環境作成とスケーラブルな実行を切り離す、API および UI 駆動の RL Gym 環境である AgenticAI-Supervisor を紹介します。検証可能な実行結果に移行することで、プラットフォームは忠実度の高いトレースを生成し、多次元の報酬形成を適用します。重要なことに、私たちのフレームワークは、厳密な内部状態の検証とテストを通じて報酬ハッキングを軽減します。この研究では、モデル最適化のための一貫した閉ループ フィードバックを実証するカスタマー サポート エージェントのケース スタディを通じて、プラットフォームのコア機能を初めて紹介します。今後の作業は、コンピュータの使用、ツールの使用、自動化された「スタンピング」、エッジケースの生成などの高度な機能に焦点を当てる予定です。

原文 (English)

Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning

As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce AgenticAI-Supervisor, an API and UI-driven RL Gym environment that decouples environment creation from scalable execution. By moving to verifiable execution outcomes, the platform generates high-fidelity traces and applies multi-dimensional reward shaping. Critically, our framework mitigates reward hacking through rigorous internal state validation and testing. This work provides a first look at our platform's core capabilities through a Customer Support Agent case study demonstrating a consistent closed-loop feedback for model optimization. Future work will focus on advanced features such as Computer Use, Tool Use, automated "stumping", and edge-case generation.

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

レンズの下の自動 DSM: LLM ベースの DSM 生成のためのブラックボックス評価フレームワーク

このペーパーでは、構造化された技術文書から設計構造マトリックス (DSM) を生成する大規模言語モデル (LLM) の能力を系統的に評価するためのブラックボックス評価フレームワークを紹介します。現在の Auto-DSM パイプラインのクローズドソースの性質を動機として、このフレームワークは、生成された DSM (GEN-DSM) を手動で検証されたグラウンド トゥルース マトリックス (GT-DSM) に対してベンチマークする、再現可能な方法論を導入しています。この評価では、構造指標 (完全性、正確性、結合密度)、分類指標 (選択精度、棄権カバレッジ)、安定性指標 (エントロピー、フライスの $\kappa$) を組み合わせて、シングルランとマルチランの両方の観点を統合します。これらの側面を総合するために、複合品質スコア (Q) が提案されています。制御された実験は、架空の抽象システムと現実世界の冷蔵庫の分解という 2 つのデータセットで実行され、表現のバリエーション、パラメーターとデータセットの調整、システムの複雑さをカバーします。結果は、LLM は構造的に妥当な DSM を生成し、適切に構造化された入力の下で高い再現性を達成できるが、あいまいさ、一貫性のない依存関係の定義、および迅速な定式化に対して依然として敏感であることを示しています。この調査結果は、幻覚と禁欲失敗の系統的な原因を浮き彫りにし、LLM 主導の DSM 自動化の潜在的な限界と現在の限界の両方を示しています。提案されたフレームワークは、Auto-DSM パイプラインを監査するための透過的なベンチマークを提供し、LLM ベースの分解手法をモデルベース システム エンジニアリング (MBSE) ワークフローに統合するための基盤を確立します。

原文 (English)

Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation

This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Design Structure Matrices (DSMs) from structured technical documentation. Motivated by the closed-source nature of current Auto-DSM pipelines, the framework introduces a reproducible methodology that benchmarks generated DSMs (GEN-DSMs) against manually validated ground-truth matrices (GT-DSMs). The evaluation integrates both single-run and multi-run perspectives, combining structural metrics (Completeness, Correctness, Coupling Density), classification metrics (Selective Accuracy, Abstention Coverage), and stability measures (Entropy, Fleiss' $\kappa$). To synthesize these aspects, a Composite Quality Score (Q) is proposed. Controlled experiments are conducted on two datasets: a fictive abstract system and a real-world refrigerator decomposition, covering variations in phrasing, parameter-dataset alignment, and system complexity. Results show that LLMs can produce structurally plausible DSMs and achieve high reproducibility under well-structured inputs, but remain sensitive to ambiguity, inconsistent dependency definitions, and prompt formulation. The findings highlight systematic sources of hallucination and abstention failure, demonstrating both the potential and current limitations of LLM-driven DSM automation. The proposed framework provides a transparent benchmark for auditing Auto-DSM pipelines and establishes foundations for integrating LLM-based decomposition methods into model-based systems engineering (MBSE) workflows.

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

プロンプトロバストネスはタスク依存: LLM 評価における客観的質問と信念スタイルの質問の比較

大規模な言語モデルの調査形式の評価では、多くの場合、促された応答がモデルの価値観や信念の尺度として扱われます。この仮定は、回答が政治的価値観、社会的態度、または信念の証拠として読み取られる場合に特に脆弱になります。答えが決まっている客観的な質問と、意見や価値観を求める主観的な質問とでは、プロンプトの堅牢性が異なるかどうかを尋ねます。 3 つの客観的データセット (MMLU、ARC、CulturalBench) と 3 つの主観的データセット (Political Compass Test、ValueBench、World Values Survey) に基づいて 4 つの命令調整モデル ファミリを評価します。各質問/ステートメントに対して、文言、枠組み、形式のバリエーションなど、複数のタイプのプロンプト変更を適用し、モデルがバリエーション全体で同じ回答を与えるかどうかを測定します。二項一般化推定方程式を使用すると、モデル、データセット、プロンプト カテゴリ、およびそれらの相互作用の重要な効果がわかります。データセット タイプの影響も大きく、データセット タイプとプロンプト カテゴリの間の相互作用は大きくなります。これらの結果は、プロンプトの堅牢性が質問の種類、プロンプトの変更、モデルに依存することを示しています。

原文 (English)

Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumption is particularly fragile when responses are read as evidence of political values, social attitudes, or beliefs. We ask whether prompt robustness differs between objective questions with fixed answers and subjective questions that ask for opinions or values. We evaluate four instruction-tuned model families on three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey). For each question/statement, we apply multiple types of prompt changes, such as variations in wording, framing, and format, and measure whether the model gives the same answer across variants. Using a binomial generalized estimating equation, we find significant effects of model, dataset, prompt category, and their interactions. The dataset type effect is also significant, and the interaction between dataset type and prompt category is large. These results show that prompt robustness depends on the question type, the prompt change, and the model.

2026-07-08 13:00 JSTarXiv cs.AIビジネス/資金調達

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems

Teams deploying large language models in business contexts need evaluation systems, yet most treat evaluation as static model selection: ru…

2026-07-08 13:00 JSTarXiv cs.AIビジネス/資金調達

Data-dependent Evaluations for Budgeted Submodular Maximization

Submodular maximization is an important building block for developing algorithms in many areas such as machine learning and data mining. Du…

2026-07-08 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Agentic AI for IPoDWDM Network Lifecycle Automation: An MCP-Enabled Architecture

We present a distributed, vendor-agnostic multi-MCP architecture for SDN-based automation and autonomous control of multi-vendor, multi-lay…

2026-07-08 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

MCP-Enabled Agentic AI for Autonomous IPoDWDM Network Lifecycle Automation

This demo presents an MCP-enabled agentic AI architecture for autonomous control of vendor-agnostic IPoDWDM networks. We demonstrate live e…

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks…

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development

Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are il…

2026-07-08 13:00 JSTarXiv cs.AIビジネス/資金調達

Evaluating Fine-Tuning and Metrics for Neural Decompilation of Dart AOT Binaries

Neural decompilation is increasingly studied as a code-generation problem, yet its evaluation methodology remains underdeveloped for modern…

2026-07-08 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Designing Maintainable Hybrid Generative Systems: A Quantum-Inspired Approach to Automated Music Harmony Generation

This paper presents the design and evaluation of a maintainable hybrid generative architecture for automated music harmony generation from…

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

AI 旅行代理店が闘牛を予約してくれる: フロンティア AI モデルにおける暗黙の動物福祉のエージェントベンチマーク

AI エージェントはアドバイザーからアクターに移行し、ユーザーに代わって旅行を予約し、メニューを計画し、調達を実行します。 AI と動物福祉の既存のベンチマークは、質問と回答のプロンプトに対するモデルのテキスト応答を評価しますが、それらの応答で表面化した福祉推論が、モデルがツールを使用してアクションを実行する必要があるエージェント展開に移行するかどうかは未解決のままです。 AI エージェントがユーザーに代わって行動する際に動物搾取を伴うオプションを回避するかどうかを測定する初のエージェント ベンチマークである TAC (Travel Agent Compassion) を紹介します。 TAC は、動物搾取の 6 つのカテゴリにわたる 12 の手書きの旅行予約シナリオを AI エージェントに提示します。これは、価格、評価、位置の交絡を制御するために 48 のサンプルに拡張されています。私たちは 4 つの研究室からの 7 つのフロンティア モデルを評価します。すべてのモデルのスコアはチャンス レベルの 64 パーセントを下回り、最高のパフォーマンスを発揮するモデル (Claude Opus 4.7) のスコアは 53 パーセントです。システム プロンプト内の福祉を意識した一文で、Claude と GPT-5.5 では 47 ~ 63 パーセント ポイント、GPT-5.2 では 26 ポイント、DeepSeek と Gemini では 12 ポイント未満の向上が見られます。 Gemini 2.5 Flash Lite を判定者として使用して、上位 2 つのパフォーマーからの 288 件の基本条件のトランスクリプトを対象とした補助的な Inspect Scout 監査では、評価認識のトランスクリプトがゼロであるとフラグが立てられ、可能性を下回る率が評価を認識するモデルに起因するものではないことが示唆されています。文化的ドメイン間のカテゴリレベルの変動の影響、テキスト応答福祉ベンチマークの限界、および EU 汎用 AI 実践規範のシステミック リスク フレームワークについて議論します。

原文 (English)

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

Previous research has evaluated animal welfare using question-and-answer benchmarks. This study investigates whether these evaluations also hold in agentic settings. The agents may showcase different behaviors compared to stand-alone large language models, as demonstrated in prior studies. This work introduces \textit{TAC (Travel Agent Compassion)}: the first agentic benchmark for assessing animal exploitation. TAC evaluates AI agentic behavior in travel booking scenarios across six animal categories, using thirteen hand-authored scenarios that vary by price, rating, and position, expanded via four augmentation variants into $52$ prompts and run for three epochs, giving $156$ scored observations per model. Nine frontier models across five model families were evaluated.. The results indicate that models tend to prefer harmful scenarios, performing below the random chance rate of $65\%$ for selecting a neutral booking option, with Claude $4.8$ achieving the highest performance at $64.7\%$. To address this issue, the persona of an ethical-brand identity was infused into the system prompt, resulting in welfare rates increasing from $32$ to $80$ percentage points, with a mean of $53$ across all nine models. No evidence of evaluation awareness affecting the results was found, based on an Inspect Scout audit of $3,120$ transcripts. These findings are directly relevant to the EU General-Purpose AI Code of Practice, which identifies non-human welfare as a systemic risk. TAC provides a practical method for measuring this risk.

2026-07-08 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

From Global to Granular: Revealing IQA Model Performance via Correlation Surface

Evaluation of Image Quality Assessment (IQA) models has long been dominated by global correlation metrics, such as Pearson Linear Correlati…

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet m…

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating…

2026-07-08 13:00 JSTarXiv cs.AI画像/動画生成ロボティクスビジネス/資金調達研究/論文

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their cap…

2026-07-08 07:00 JSTITmedia AI+ビジネス/資金調達

IT予算の9割が人件費に消える――日本オラクル社長が切り込む「企業最大の課題」

日本オラクルの三澤智光社長が、日本企業のIT課題に切り込んだ。同社長が指摘する「IT投資の構造的問題」「オンプレミスシステムが抱える課題」とは何か。

2026-07-07 21:00 JSTTechCrunch AIビジネス/資金調達

Savi’s app aims to protect consumers from realistic AI scams like kidnappers demanding ransom

The company just raised $7 million in seed funding, and is launching its app for iPhone and Android on Tuesday.

2026-07-07 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

予測を超えて: 予測市場エージェントの信念から取引までのレイヤー

未来の出来事を予測することは、汎用AIのテストベッドとして注目を集めています。この評価を根拠付ける自然な方法は、モデルを予測市場で取引させることです。ただし、トレーディングには予測以上のものが必要です。さらに、最近のベンチマークでは、調整された確率スコアと取引結果の間に大きなギャップがあることが報告されています。私たちは、私たちの知る限り、予測市場向けの初の自律型取引エージェントである Raven-Agent を提案します。アーカイブされた意思決定セットに対する制御された再生では、私たちのアーキテクチャは、テストされたすべてのポリシーの中で唯一のプラスのリターンと唯一のプラスのリスク調整後のリターンを達成します。コードを https://github.com/Alchemist-X/predict-raven でリリースしました。

原文 (English)

Beyond Forecasting: The Belief-to-Trade Layer in Prediction-Market Agents

Forecasting future events has attracted growing attention as a testbed for general-purpose AI. A natural way to ground this evaluation is let the models trade in the prediction markets. Trading, however, requires more than forecasting. Moreover, recent benchmarks report a substantial gap between calibrated probability scores and the trading results. We propose Raven-Agent, to the best of our knowledge, the first autonomous trading agent for prediction markets. On a controlled replay over an archived decision set, our architecture achieves the only positive return and the only positive risk-adjusted return among all tested policies. We have released our code in https://github.com/Alchemist-X/predict-raven .

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

反省的な対話か、それとも迅速な改善か?生徒のプログラミングのための独立した LLM の使用に対する家庭教師の足場の影響

大規模言語モデル (LLM) は学習において個人に合わせたサポートを提供できますが、いくつかの研究で教育での使用について懸念が生じています。重要なのは、学習は学生が LLM にどのように関与するかによって決まります。この研究では、2 種類の LLM ベースの家庭教師が生徒のプロンプトの実践、学習、その後の LLM の使用をどのように形成するかを調査しました。対話的な質問を通じて対話を構築するソクラティック ガイダンス (SG) 家庭教師と、効果的なプロンプトの作成をガイドするプロンプト リファインメント (PR) 家庭教師です。私たちは、大学院レベルのモバイル ロボット工学コースで 2 段階の研究を実施しました。66 人の学生が 6 週間の介入中に SG または PR 講師のいずれかを使用し、続いて 52 人の学生が 3 週間のコース プロジェクト中に制約のない LLM を使用しました。結果は、SG 講師と PR 講師は、ガイド付き使用中に同様のタスクのパフォーマンスとプロンプトパターンをもたらしましたが、学習成果とその後の LLM の使用において異なることを示しています。 SG の学生は、PR の学生と比較して、後のセッションでより高い学習成果を達成し、制約のない LLM を使用した場合、理解度の向上を予測する、理解主導のプロンプト戦略を採用する可能性が高くなりました。学習者は SG 講師の効率が低いと認識していましたが、この調査結果は、ソクラテスの指導が時間の経過とともに LLM で学習する生徒の能力の発達をサポートしていることを示唆しており、LLM 講師の設計におけるソクラテスの重要性を強調しています。

原文 (English)

Reflective Dialogue or Prompt Refinement? Effects of Tutor Scaffolding on Students' Independent LLM Use for Programming

While Large Language Models (LLMs) can provide personalized support in learning, several studies have raised concerns regarding their use in education. Importantly, learning depends on how students engage with LLMs. This study examined how two types of LLM-based tutors shape students' prompting practices, learning, and subsequent LLM-use: a Socratic-Guidance (SG) tutor, which structures interaction through dialogic questioning, and a Prompt-Refinement (PR) tutor that guides the formulation of effective prompts. We conducted a two-phase study in a graduate-level mobile robotics course: 66 students used either the SG or PR tutor during a 6-week intervention, followed by 52 students using an unconstrained LLM during a 3-week course project. Results show that while the SG- and PR tutors led to similar task performance and prompting patterns during guided use, they differ in learning outcomes and later LLM-use. SG-students, relative to PR-student, achieved higher learning gains in later sessions, and were more likely to adopt understanding-driven prompting strategies, which are predictive of higher understanding, when using an unconstrained LLM. Although learners perceived the SG tutor as less efficient, the findings suggest that Socratic guidance supports the development of students' capacity to learn with LLMs over time, highlighting its importance for LLM tutor design.

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

エージェント ステップ値: 状態接地 LLM エバリュエーターによる状態遷移測定

ほとんどのエージェント評価では、複数ステップのトレースが最終的な回答、成功フラグ、または軌跡レベルのスコアにまとめられます。これらの集計では、開発者が最も必要とする診断の質問、つまりどのアクションが状態を有益な方向に変更したのかがわかりにくくなります。我々は、状態遷移測定フレームワークであるエージェント ステップ値 (ASV) を導入します。これは、観察された各アクションを、固定された候補結果に対する状態に基づいた評価者の分布に誘発する変化によってスコア付けします。 ASV は、編集された前後の状態予測をレンダリングし、ステートレス LLM エバリュエーターを使用して候補ログ スコアを割り当て、ゴールドフリーの信念診断とオフライン オラクル検証メトリクスの両方をレポートします。ラベルフリーの理論的パスにより、評価者の審議がワントークンオプションのスコアリングから分離され、リークやフロアスコアイベントを明らかにしながら候補の可能性が維持されます。ライブ PubMed 検索、部分的にライブ DeepSeek アクター、および DeepSeek 対数確率スコアリングを使用した 100 件のレビュー済みオープン QA 証拠探索タスクについて、ASV は 1,100 のステップと 2,200 の状態を評価します。固定レイアウトの根拠条件付きプロトコルの下では、平均ゴールドマージンゲインは -2.335 (軌道ブートストラップ 95\% CI [-3.395, -1.272])、エントロピーの動きは 0.000、平均ベイジアンサプライズは 2.693 です。したがって、ASV は、最終回答スコアとエントロピーのみのステップ メトリクスが見逃している建設的および破壊的な信念ピボットを特定します。スタンドアロンの ASV Eval ツールキットをリリースします。

原文 (English)

Agent Step Value: State-Transition Measurement with State-Grounded LLM Evaluators

Most agent evaluations collapse a multi-step trace into a final answer, a success flag, or a trajectory-level score. These aggregates obscure the diagnostic question developers need most: which action changed the state in a useful direction? We introduce Agent Step Value (ASV), a state-transition measurement framework that scores each observed action by the change it induces in a state-grounded evaluator's distribution over fixed candidate outcomes. ASV renders redacted before/after state projections, uses a stateless LLM evaluator to assign candidate log scores, and reports both gold-free belief diagnostics and offline oracle validation metrics. A label-free rationale pass separates evaluator deliberation from one-token option scoring, preserving candidate likelihoods while exposing leakage and floor-score events. On 100 reviewed open-QA evidence-seeking tasks with live PubMed retrieval, a partially live DeepSeek actor, and DeepSeek log-probability scoring, ASV evaluates 1,100 steps and 2,200 states. Under the fixed-layout rationale-conditioned protocol, mean gold-margin gain is -2.335 (trajectory-bootstrap 95\% CI [-3.395, -1.272]), entropy movement is 0.000, and mean Bayesian surprise is 2.693. ASV therefore localizes constructive and destructive belief pivots that final-answer scores and entropy-only step metrics miss. We release the standalone ASV Eval toolkit.

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

JavaVulBench: A Java Vulnerability Benchmark with Realistic Splits, a Unified Multi-Backend Harness, and a Leakage-Aware Evaluation Mode

We release \textsc{JavaVulBench}, a benchmark dataset and evaluation harness for Java vulnerability detection. The dataset contains $\sim$3…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

The Foreign Policy AI Evaluation Gap

We argue that AI systems used in conducting foreign policy tasks - broadly enacting 'statecraft' - should be a priority test case for techn…

2026-07-07 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for m…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

Efficient Decentralized Multi-task Dataset Valuation via Model Merging

Accurate and efficient dataset valuation is essential for enabling fair and transparent data marketplaces, especially when multiple contrib…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

TRIAGE: Trustworthy Retrieval Instrumentation And Graph Evaluation

Knowledge graphs (KGs) that underpin Graph-based Retrieval-Augmented Generation (Graph-RAG) are increasingly built automatically by LLM-dri…

2026-07-07 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI

Enterprise artificial intelligence is moving from experimentation into operational workflows. Early programs focused on model access and re…

2026-07-07 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation

Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a ma…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation

Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoidi…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

A Failure-Mode Benchmark for Polymorphic Sybil Poisoning in RAG

We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning. The framework par…

2026-07-07 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

How Do Diffusion Classifiers Decide? A Bias-Centric Evaluation

Diffusion models have recently been repurposed for zero-shot classification, giving rise to diffusion classifiers that identify the best-ma…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

A Unified Algebraic Framework for Classification Performance Evaluation

We propose a unified algebraic framework for classification performance evaluation that encompasses binary, multiclass, multilabel, ordinal…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese)

This paper explains the principles and provides examples of a new method for distinguishing between FAKE human speech synthesized by genera…

2026-07-07 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement

The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawi…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations

The dominant practice in language model evaluation is to report a single accuracy number per model and declare the higher one better, witho…

2026-07-07 13:00 JSTarXiv cs.AI画像/動画生成ロボティクスビジネス/資金調達研究/論文

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their cap…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

A Deep Learning-based surrogate model for Severe Accidents in nuclear reactors using ASTEC

Integral codes like the Accident Source Term Evaluation Code (ASTEC) are powerful tools to study the physics of Severe Accidents (SAs) in n…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models

World-model evaluation for model-based reinforcement learning typically asks whether the learned model predicts reward and value well, whic…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects.…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

AIFS-SUBS: Extending Data-Driven Forecasting to Sub-Seasonal Timescales

Data-driven models now rival numerical weather prediction in the medium range, but extending them to sub-seasonal lead times raises challen…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Three-Phase Evaluation of AI-Assisted Software Development Life Cycle

This paper presents an exploratory evaluation of how increasing levels of AI autonomy affect software development productivity, requirement…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text b…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds

High-stakes decision-making involves navigating multiple competing objectives with expensive evaluations. For instance, in brachytherapy, c…

2026-07-07 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI

Perturbation-based explainability methods face criticism due to their reliance on out-of-distribution mutants. This raises doubts about the…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

RLIE: Rule Generation with Logistic Regression, Iterative Refinement, and Evaluation for Large Language Models

Large Language Models (LLMs) can propose rules in natural language, sidestepping the need for a predefined predicate space in traditional r…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

Fun-TSG: A Function-Driven Multivariate Time Series Generator with Variable-Level Anomaly Labeling

Reliable evaluation of anomaly detection methods in multivariate time series remains an open challenge, largely due to the limitations of e…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Stop Automating Peer Review Without Rigorous Evaluation

Large language models offer a tempting solution to address the peer review crisis. This position paper argues that today's AI systems shoul…

2026-07-07 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State

Outcome-only evaluation can certify economically unsafe agents: a policy can hit a business KPI while violating deployable behavioral disci…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

MUSE-Autoskill: スキルの作成、記憶、管理、評価による自己進化エージェント

大規模言語モデル (LLM) エージェントは、再利用可能なスキルに依存して複雑なタスクを解決します。ただし、既存のスキル作成アプローチでは、スキルを孤立した静的な成果物として扱い、再利用性、信頼性、長期的な改善が制限されています。私たちは、統一されたライフサイクル (作成、記憶、管理、評価、洗練) の下でスキルを作成、再利用、洗練することにより、エージェントがタスク解決能力を継続的に向上できるようにする、スキル中心のエージェント フレームワークである MUSE-Autoskill Agent (Memory-Utilizing Skill Evolution) を提案します。当社のフレームワークにより、エージェントはオンデマンドでスキルを作成し、それらをタスク間で保存して再利用し、効率的に整理して選択し、単体テストや実行時のフィードバックを通じて評価して継続的に改善することができます。さらに、タスク全体にわたって各スキルの経験を蓄積するスキルレベルの記憶を導入し、時間の経過とともにより効果的な再利用と適応を可能にします。 SkillsBench の実験は、ライフサイクル管理されたスキルがタスクの成功、効率、再利用、およびエージェント間での移転を向上させることができるという最初の証拠を提供し、スキルを長命で経験を意識したテスト可能な資産として扱うことの重要性を強調しています。

原文 (English)

MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation

Large language model (LLM) agents rely on reusable skills to solve complex tasks, but existing skill creation approaches often treat skills as isolated, static artifacts, limiting reusability, reliability, and long-term improvement. We propose MUSE-Autoskill Agent (Memory-Utilizing Skill Evolution), a skill-centric agent framework that creates, reuses, and refines skills under a unified lifecycle: creation, memory, management, evaluation, and refinement. MUSE creates skills on demand, stores them across tasks, retrieves them through a skill catalog, and accumulates per-skill experience for later reuse and adaptation. Across the main reported settings on SkillsBench and SkillLearnBench, MUSE-Autoskill outperforms Hermes, Codex, and Claude Code. On SkillsBench, its self-created skills surpass human-authored skills on the successfully covered subset (85.24% vs. 81.17%), showing that lifecycle-managed skills can distill agent experience into highly effective reusable assets; MUSE-created skills also transfer to Hermes more effectively than Codex- or Claude-created skills, reaching 51.90% accuracy under transfer. These results highlight the importance of treating skills as long-lived, experience-aware, and testable assets.

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ComplexConstraints and Beyond: Expert Rubrics for RLVR

Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-w…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Autodata: An agentic data scientist to create high quality synthetic data

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation dat…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

PACE: エージェントの能力評価の代理

SWE-Bench や GAIA などのベンチマークで LLM エージェントを評価するには、費用と時間がかかり、複雑なインフラストラクチャが必要になる場合があります。 1 回の評価に数千ドルの費用がかかり、完了までに数日かかる場合があります。対照的に、個々の機能 (推論、コード生成など) をテストする非エージェント LLM ベンチマークは、高速かつ低コストで実行できます。このペーパーでは、高価なエージェント ベンチマークでのパフォーマンスが、慎重に選択されたアトミック評価インスタンスの少数のサブセットでのパフォーマンスによって正確に予測できるかどうかを調査します。 PACE は、既存の非エージェント評価からインスタンスを選択することでプロキシ ベンチマークを構築するフレームワークで、その集計スコアがエージェント ベンチマークでのモデル パフォーマンスを最も確実に予測するものを紹介します。アトミック機能にわたる候補インスタンスのプールを考慮すると、PACE は、ソース インスタンスのコンパクトなサブセットのモデルのスコアをターゲット エージェント ベンチマークのスコアにマッピングする回帰を当てはめます。サブセット自体は、2 つの相補的なインスタンス選択戦略、ターゲット関連性のローカル選択とグローバルに情報を提供するグローバル選択を組み合わせることによってキュレーションされます。このペーパーでは、PACE を 4 つのターゲット エージェント ベンチマークに適用し、このペーパーで評価する具体的なプロキシ ベンチマークである PACE-Bench を生成します。 14 のモデル、4 つのエージェント ベンチマーク、および 19 の非エージェント ベンチマークにわたる実験では、PACE-Bench が、リーブ ワンアウト相互検証 (LOOCV) 平均絶対誤差 (MAE) が 4% 未満、スピアマン相関が 0.80 以上、ペアワイズ モデル ランク付け精度が約 85% で、すべてエージェント評価コスト全体の 1% 未満でエージェント スコアを予測することが示されています。選択したプロキシ インスタンスをさらに分析し、各エージェント ベンチマークが独自に要求するスキルを明らかにします。 PACE を使用すると、担当者は、完全なエージェント評価のオーバーヘッドを発生させることなく、モデルの開発、選択、ルーティング中にエージェントのパフォーマンスの信頼できる推定値を取得できます。

原文 (English)

PACE: A Proxy for Agentic Capability Evaluation

Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benchmarks by selecting instances from existing non-agentic evaluations whose aggregate scores most reliably predict model performances on agentic benchmarks. Given a pool of candidate instances spanning atomic capabilities, PACE fits a regression that maps a model's scores on a compact subset of source instances to its score on the target agentic benchmark. The subset itself is curated by combining two complementary instance-selection strategies, target-relevance local selection and globally informative global selection. We apply PACE to the 4 target agentic benchmarks in this paper, which yields PACE-Bench, the concrete proxy benchmark that we evaluate in the paper. Experiments across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks show that PACE-Bench predicts agentic scores with leave-one-out cross-validation (LOOCV) mean absolute error (MAE) under 4%, Spearman correlation above 0.80, and pairwise model-ranking accuracy around 85%, all at much less than 1% of the full agentic evaluation cost. We further analyze the selected proxy instances, revealing which skills each agentic benchmark uniquely demands. PACE enables practitioners to obtain reliable estimates of agentic performance during model development, selection, and routing, without the overhead of full agent evaluation.

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

Double Fuzzy Probabilistic Interval Linguistic Term Set and a Dynamic Fuzzy Decision Making Model based on Markov Process with tts Application in Multiple Criteria Group Decision Making

The probabilistic linguistic term has been proposed to deal with probability distributions in provided linguistic evaluations. However, bec…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

GenShin: Guiding Rational Liposome Design by Ranking Liposomal Protein Corona through a Docking-Pose-Free GNN

Rational design of lipid nanoparticles (LNPs) for tissue-specific delivery critically depends on predicting the composition of the protein…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

Boosting Automatic Exercise Evaluation Through Musculoskeletal Simulation-Based IMU Data Augmentation

Automated evaluation of movement quality can enhance physiotherapeutic treatment and sports training by providing objective, real-time feed…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues

Large language models (LLMs) may support counseling training, yet evidence from Japanese-language interactions and automated quality rating…

2026-07-07 13:00 JSTarXiv cs.AIビジネス/資金調達

Algorithmic Shortlisting in Participatory Budgeting

Participatory budgeting is a democratic innovation that allows citizens to propose and vote on public investment projects. To help organize…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

Large Language Models (LLMs) are increasingly used as interfaces to information, code, and real-world services, making prompt-level securit…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding

Federal research funding shapes the direction, diversity, and impact of the US scientific enterprise. Large language models (LLMs) are rapi…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation

Accurate evaluation of user satisfaction is critical for iterative development of conversational AI. However, for open-ended assistants, tr…

2026-07-07 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

GIPO: Gaussian Importance Sampling Policy Optimization

Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitatio…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation

Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise. Thus, like playing the slo…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents

Persistent memory in LLM agents creates an attack surface that production safety classifiers do not observe: the payload enters via RAG ret…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM Unlearning の属性を解除して忘れる

大規模言語モデル (LLM) の急速な発展により、トレーニングに不適切なデータが使用されることへの懸念が生じ、LLM の非学習への関心が高まっています。既存の LLM 未学習アプローチの多くは、忘却集合の損失の最大化など、予測損失の最適化に依存していますが、過剰忘却やモデルの有用性の低さなどの重大な問題に直面することがよくあります。これらに対処するために、この論文では、代わりにデータの帰属をゼロにすることの 1 つとして、LLM アンラーニングの最適化目標を斬新に組み立てています。特に、DareUと呼ばれるデータ帰属報酬に基づいた最初のLLM非学習フレームワークを提案します。これは、データ所有者を忘れた場合に生成された応答の帰属スコアを減らす(つまり、帰属を解除する)ことによってLLMを更新する強化学習を実行します。アトリビューションの効率的な近似として LLM 分類器を使用した経験的評価では、DareU が忘却品質とモデルの有用性のバランスを適切に保ちながら効果的なアンラーニングを達成することで、既存のベースラインを上回るパフォーマンスを示しています。

原文 (English)

De-attribute to Forget for LLM Unlearning

The rapid development of large language models (LLMs) has raised concerns on the use of inappropriate data for training, which has led to a growing interest in LLM unlearning. Many existing LLM unlearning approaches rely on optimizing prediction loss(es), such as maximizing the loss on the forget set, but often face critical issues like over-forgetting and poor model utility. To address them, this paper novelly frames the optimization objective for LLM unlearning as one of zeroing out data attribution instead. In particular, we propose the first LLM unlearning framework based on data attribution rewards called DareU that performs reinforcement learning to update the LLM by reducing the attribution score of its generated responses (i.e., de-attributing) to the forget data owners. Empirical evaluation using an LLM classifier as an efficient approximation of attribution shows that DareU outperforms existing baselines by achieving effective unlearning while balancing forget quality and model utility well.

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

A language model's memory can be worse than no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behin…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluatio…

2026-07-07 08:21 JSTTechCrunch AIビジネス/資金調達

US investors will soon get access to SK Hynix, another memory maker riding the AI boom

SK Hynix is experiencing a boom credited to AI. It will ride that to a multibillion-dollar U.S. IPO, expected to take place on Friday.

2026-07-05 11:00 JSTITmedia AI+ビジネス/資金調達

マイクロン、AI需要で広島工場増強へ起工式 1.5兆円投資

マイクロンメモリ ジャパンは2026年7月、広島工場の生産能力増強に向けた新クリーンルーム建設の起工式を開催した。AI技術の進展に伴うメモリ需要の増加に対応するもので、2028年後半に製造装置の搬入を開始する予定だ。広島工場には今後、この新クリーンルーム建設を含めて1兆5000…

2026-07-05 00:51 JSTTechCrunch AILLM/生成AIビジネス/資金調達

What is Mistral AI? Everything to know about the OpenAI competitor

Mistral AI, which offers some open source AI models, has raised significant funding since its creation in 2023, with the ambition to “put f…

2026-07-03 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Janus: ユーザーが関与するエージェント権限管理のプレイグラウンド

ユーザーに代わってツール呼び出しを自律的に実行する AI エージェントは、ユーザーがどのような役割を果たせるのか、またどのような役割を果たすべきなのか、権限管理に関する差し迫った疑問を引き起こします。多くのアプローチが提案されているにもかかわらず、エージェントのアクセス許可管理におけるユーザーの役割はまだ検討されていません。ユーザーが関与するエージェント権限管理設計を実装および評価するためのプレイグラウンド システムである Janus を紹介します。 Janus は、さまざまな権限管理設計をサポートするモジュール型エージェント システムである Janus-Core と、自動評価フレームワークである Janus-Harness の 2 つのコンポーネントで構成されています。ユーザー参加のための主要な設計軸を特定する概念モデルに基づいて、設計空間にわたる 6 つの権限アシスタントを実装し、3 つのシナリオと 3 つの合成レスポンダーにわたってそれらを評価します。私たちは、ユーザー入力が重要でプライバシーとセキュリティを大幅に強化できること、ユーザーの意思決定の AI 拡張が認知負荷の軽減に役立つこと、システム設計では許可疲労を含む現実的なユーザーの行動を考慮する必要があることを実証します。すべてのコンテキストにわたって最適に機能する単一の設計は存在しないため、エージェント システムにパーミッション アシスタントを展開するための、より原則に基づいたコンテキスト依存のアプローチが推進されます。 Janus は、エージェント システム設計のこの側面に関する今後の調査をサポートするために一般に公開されています。

原文 (English)

Janus: a Playground for User-Involved Agentic Permission Management

AI agents that autonomously execute tool calls on a user's behalf raise pressing questions about permission management: what role could users play, and what role should they play? Despite many proposed approaches, the user's role in agentic permission management remains under explored. We introduce Janus, a playground system for implementing and evaluating user-involved agentic permission management designs. Janus consists of two components: Janus-Core, a modular agentic system supporting a diverse spectrum of permission management designs, and Janus-Harness, an automated evaluation framework. Grounded in a conceptual model that identifies key design axes for user involvement, we implement six permission assistants spanning the design space and evaluate them across three scenarios and three synthetic responders. We demonstrate that user input is critical and can significantly strengthen privacy and security, that AI augmentation of user decisions can help reduce cognitive load, and that realistic user behavior including permission fatigue must be accounted for in system design. No single design performs optimally across all contexts, motivating a more principled and context-sensitive approach to deploying permission assistants in agentic systems. Janus is publicly available to support future investigation into this dimension of agentic system design.

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

金融サービス LLM 評価のメタベンチマーク

公開 LLM リーダーボードは、世界平均のパフォーマンスに合わせて最適化されており、金融サービス業務に特有の認知的要求を捉えていません。MMLU-Pro をリードするモデルは、文書に基づいたコンプライアンスの推論ではパフォーマンスが劣る可能性があり、コーディング リーダーは複数ターンにわたる顧客インタラクションの処理が不十分になる可能性があります。当社は、公的に報告された 452 のベンチマークを 41 の O*NET 一般化作業アクティビティに整理し、それらを販売、業務、リスク、サポート業務にわたる 38 の BIAN 銀行ビジネス ドメインに集約するメタ ベンチマーク フレームワークを提示します。ローリング モデル ウィンドウで計算される乗法重み付けスキーム (識別 x カバレッジ x リーンシー) は、依然として最良のモデルを区別するベンチマークに報酬を与え、広く報告され、アクティブに使用され続けており、飽和したレガシー テストを自動的に抑制します。これらの重みは、ペアごとの Elo トーナメントの K ファクターをスケールし、生のスコアを正規化することなく、クロスベンチマークで比較可能な作業活動スコアを生成します。ビジネスドメインスコアは、構成要素である作業活動 Elos の加重平均です。 2026 年 6 月の時点で 25 の組織にわたる 288 のモデルをカバーするポイントインタイムの公開スナップショットでフレームワークを実証し、同様の選択とガバナンスの課題に直面している組織でこのアプローチを再現可能にすることを目的として、方法論、完全な分類、設計上の決定、および制限について説明します。

原文 (English)

Meta-Benchmarks for Financial-Services LLM Evaluation

Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model that leads on MMLU-Pro may underperform on document-grounded compliance reasoning, and a coding leader may handle multi-turn customer interactions poorly. We present a meta-benchmarking framework that organises 452 publicly reported benchmarks into 41 O*NET Generalized Work Activities and aggregates those into 38 BIAN banking business domains spanning sales, operations, risk, and support work. A multiplicative weighting scheme (discrimination x coverage x recency), computed over a rolling model window, rewards benchmarks that still separate the best models, are widely reported, and remain in active use, suppressing saturated legacy tests automatically. These weights scale the K-factor in a pairwise Elo tournament, producing cross-benchmark-comparable work-activity scores without raw score normalisation; business-domain scores are weighted averages of the constituent work-activity Elos. We demonstrate the framework on a point-in-time public snapshot covering 288 models across 25 organisations as of June 2026, and describe the methodology, full taxonomy, design decisions, and limitations with the aim of making the approach reproducible for institutions facing similar selection and governance challenges.

2026-07-03 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

CLAP: ドメイン エージェントのトレーニング後のクローズド ループ トレーニング、評価、リリース制御

ドメイン エージェントは、多くの場合、ノイズの多いビジネス データ、不確実なトレーニング後の利益、オフライン/アプリケーションの不一致、アダプター リリースのリスクに直面します。このペーパーでは、ビジネス データを構造化された SFT サンプル、意思決定優先サンプル、ホールドアウト セット、リスク診断、およびリリース ゲート レコードに変換する閉ループ手法である CLAP (Closed-Loop Agent Post-training) について説明します。 CLAP は、データ検証、ターゲット/証拠の正規化、報酬/KL 診断、オフライン ゲート、およびアプリケーション チェーンの再生を組み合わせて、アダプターがターゲット アプリケーション チェーンに適しているかどうかを判断します。 5 つの匿名化された製造シナリオ バッチで、QLoRA スタイルの LoRA-SFT は平均してわずかな向上をもたらしました。全体のスコアは 0.0098 増加し、合格率は 0.0240 増加し、証拠の精度は 0.0280 増加しましたが、幻覚と誤った事実は減少しました。しかし、改善するのは 5 つのバッチのうち 3 つだけで、一部のバッチは後退し、GRPO は高い KL リスクを露呈します。アプリケーション チェーンの再生は、事実の抽出に RAG が必要であることをさらに示しています。同じ 3B バックボーンと 100 のリプレイ ケースの下では、アプリケーション指向の LoRA-SFT アダプターは、ベース + RAG よりも値、コア フィールド、および回答証拠のドキュメント/ページのマッチングを向上させますが、レイテンシは増加します。これらの結果は、トレーニングの完了や単​​一のオフライン スコアに依存するのではなく、統合されたデータ、トレーニング、評価、リリースのループを通じてトレーニング後のドメイン エージェントを管理することをサポートします。

原文 (English)

CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training

Domain agents often face noisy business data, uncertain post-training gains, offline/application mismatch, and adapter-release risk. This paper presents CLAP (Closed-Loop Agent Post-training), a closed-loop method that converts business data into structured SFT samples, decision-preference samples, holdout sets, risk diagnostics, and release-gate records. CLAP combines data validation, target/evidence normalization, reward/KL diagnosis, offline gates, and application-chain replay to decide whether an adapter is suitable for the target application chain. On five anonymized manufacturing-scenario batches, QLoRA-style LoRA-SFT yields modest average gains: overall score increases by 0.0098, pass rate by 0.0240, and evidence accuracy by 0.0280, while hallucination and wrong facts decrease. Yet only 3 of 5 batches improve, some batches regress, and GRPO exposes high KL risks. Application-chain replay further shows that RAG is necessary for factual extraction; under the same 3B backbone and 100 replay cases, an application-RAG-oriented LoRA-SFT adapter improves value, core fields, and answer-evidence doc/page matching over base+RAG, but increases latency. These results support managing domain-agent post-training through an integrated data-training-evaluation-release loop rather than relying on training completion or a single offline score.

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

PACE: エージェントの能力評価の代理

SWE-Bench や GAIA などのベンチマークで LLM エージェントを評価するには、費用と時間がかかり、複雑なインフラストラクチャが必要になる場合があります。 1 回の評価に数千ドルの費用がかかり、完了までに数日かかる場合があります。対照的に、個々の機能 (推論、コード生成など) をテストする非エージェント LLM ベンチマークは、高速かつ低コストで実行できます。このペーパーでは、高価なエージェント ベンチマークでのパフォーマンスが、慎重に選択されたアトミック評価インスタンスの少数のサブセットでのパフォーマンスによって正確に予測できるかどうかを調査します。 PACE は、既存の非エージェント評価からインスタンスを選択することでプロキシ ベンチマークを構築するフレームワークで、その集計スコアがエージェント ベンチマークでのモデル パフォーマンスを最も確実に予測するものを紹介します。アトミック機能にわたる候補インスタンスのプールを考慮すると、PACE は、ソース インスタンスのコンパクトなサブセットのモデルのスコアをターゲット エージェント ベンチマークのスコアにマッピングする回帰を当てはめます。サブセット自体は、2 つの相補的なインスタンス選択戦略、ターゲット関連性のローカル選択とグローバルに情報を提供するグローバル選択を組み合わせることによってキュレーションされます。このペーパーでは、PACE を 4 つのターゲット エージェント ベンチマークに適用し、このペーパーで評価する具体的なプロキシ ベンチマークである PACE-Bench を生成します。 14 のモデル、4 つのエージェント ベンチマーク、および 19 の非エージェント ベンチマークにわたる実験では、PACE-Bench が、リーブ ワンアウト相互検証 (LOOCV) 平均絶対誤差 (MAE) が 4% 未満、スピアマン相関が 0.80 以上、ペアワイズ モデル ランク付け精度が約 85% で、すべてエージェント評価コスト全体の 1% 未満でエージェント スコアを予測することが示されています。選択したプロキシ インスタンスをさらに分析し、各エージェント ベンチマークが独自に要求するスキルを明らかにします。 PACE を使用すると、担当者は、完全なエージェント評価のオーバーヘッドを発生させることなく、モデルの開発、選択、ルーティング中にエージェントのパフォーマンスの信頼できる推定値を取得できます。

原文 (English)

PACE: A Proxy for Agentic Capability Evaluation

Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benchmarks by selecting instances from existing non-agentic evaluations whose aggregate scores most reliably predict model performances on agentic benchmarks. Given a pool of candidate instances spanning atomic capabilities, PACE fits a regression that maps a model's scores on a compact subset of source instances to its score on the target agentic benchmark. The subset itself is curated by combining two complementary instance-selection strategies, target-relevance local selection and globally informative global selection. We apply PACE to the 4 target agentic benchmarks in this paper, which yields PACE-Bench, the concrete proxy benchmark that we evaluate in the paper. Experiments across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks show that PACE-Bench predicts agentic scores with leave-one-out cross-validation (LOOCV) mean absolute error (MAE) under 4%, Spearman correlation above 0.80, and pairwise model-ranking accuracy around 85%, all at much less than 1% of the full agentic evaluation cost. We further analyze the selected proxy instances, revealing which skills each agentic benchmark uniquely demands. PACE enables practitioners to obtain reliable estimates of agentic performance during model development, selection, and routing, without the overhead of full agent evaluation.

2026-07-03 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

専門家が作成した臨床推論タスクに関するフロンティア言語モデルのルーブリックベースの管理された比較

多肢選択式の医療ベンチマークはますます飽和状態にあり、HealthBench などの最近のルーブリックベースの評価では、オープンエンドの臨床パフォーマンスは解決にはほど遠いことが示されており、その「ハード」サブセットのトップスコアは 32% のままです。我々は、臨床医が作成した 4 つの専門分野 (麻酔、内科/家庭医学、救急医療、産科) にわたる 5 つの臨床シナリオからなる小規模で意図的に困難な評価データセットを提示します。各臨床シナリオには、臨床医が起草したゴールデンアンサーから作成された、アトミックで重み付けされた MECE ルーブリック (タスクごとに 25 ~ 62 の基準、合計 184 の基準) が付属しています。 GPT 5.4、Claude Opus 4.7、Gemini 3.1 Pro の 3 つのフロンティア モデルを評価します。平均ルーブリック合格率は、0.47 (Claude)、0.39 (GPT)、および 0.37 (Gemini) でした。中心となる所見は、臨床的優先順位の逆転です。最も重み付けされた (重み付け 5、重要) 基準は 32.4 ~ 41.7% でのみ合格しましたが、最も低い重み付け 1 基準は 80 ~ 90% で合格しました。 108 の重要 (重み 5) 基準のうち 56 (52%) を満たしたモデルはありませんでした。 3 人の LLM 自動評価者が、552 の等級付け基準の 92.8 ~ 94.7% について、専門家の適合/不適合ラベルを再現しました。私たちはこれを手法と予備調査結果の貢献として位置づけています。5 つのタスクは、大規模なベンチマークに開発する準備ができているスケーラブルで防御可能なパイプラインを示しています。

原文 (English)

A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks

Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its "Hard" subset top score remains 32%. We present a small, deliberately difficult evaluation dataset of five clinician-authored clinical scenarios spanning four specialties (anaesthesia, internal/family medicine, emergency medicine, and obstetrics), each accompanied by an atomic, weighted, MECE rubric (25-62 criteria per task; 184 criteria total) authored from a clinician-drafted golden answer. We evaluate three frontier models: GPT 5.4, Claude Opus 4.7, and Gemini 3.1 Pro. Mean rubric pass rates were 0.47 (Claude), 0.39 (GPT), and 0.37 (Gemini). The central finding is an inversion of clinical priority: the highest-weighted (weight-5, critical) criteria passed at only 32.4-41.7%, while low-stakes weight-1 criteria passed at 80-90%. 56 of 108 critical (weight-5) criteria (52%) were satisfied by no model. Three LLM autoraters reproduced expert met/not-met labels on 92.8-94.7% of 552 graded criteria. We position this as a methods-and-preliminary-findings contribution: the five tasks demonstrate a scalable, defensible pipeline ready to develop into a large-scale benchmark.

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

EvoPolicyGym: 対話型環境での自律的なポリシー展開の評価

自律型エージェントは、フィードバックを通じて実行可能なポリシーを改善することがますます期待されていますが、既存の評価では、このプロセスが最終スコアに組み込まれたり、無制限のソフトウェア エンジニアリングの進歩と混同されたりすることがよくあります。ハーネス モデル エージェントが固定のインタラクション バジェットの下で実行可能なポリシー システムを繰り返し編集する、制御された評価設定である Autonomous Policy Evolution を導入します。この設定は、エージェントが探索されたポリシーを反復的に改善する方法を評価するコンパクトな対話型 RL 環境から構築されたベンチマークである EvoPolicyGym でインスタンス化されます。 EvoPolicyGym スイートでは、GPT-5.5 は 16 環境すべてで最強の総合ランク スコアと上位 2 位のパフォーマンスを達成しました。 EvoPolicyGym は、リーダーボードの結果以外にも、エージェントがどのように予算を割り当て、フィードバックをパラメトリック調整に変換するかを区別する軌跡レベルの診断も提供します。これらの分析は、強力な自律的なポリシーの進化は、孤立したタスクの勝利だけではなく、タスクに適したメカニズムを発見し、制限されたフィードバックの下でポリシーを洗練することに依存していることを示しています。

原文 (English)

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress. We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget. We instantiate this setting in EvoPolicyGym, a benchmark built from compact interactive RL environments that evaluates how agents iteratively improve explored policies. On the EvoPolicyGym suite, GPT-5.5 achieves the strongest aggregate rank score and top-two performance on all 16 environments. Beyond leaderboard results, EvoPolicyGym also provides trajectory-level diagnostics that distinguish how agents allocate budget, convert feedback into parametric tuning. These analyses show that strong autonomous policy evolution depends not only on isolated task wins, but on discovering task-appropriate mechanisms and refining policies under bounded feedback.

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

プロンプト フレーミングは LLM エラー検出のカウントベースの評価を歪める: 数値アンカーからの証拠

カウントベースの F1 は、LLM エラー検出品質の代用として広く使用されていますが、この論文では、スパンの局所化における対応する改善、つまり F1 インフレーションと呼ばれるギャップがなければ、F1 が劇的に上昇する可能性があることを示しています。この論文では、プロンプト誘発カウント歪みに対する制御されたストレス テスト プロトコルである ErrorBench を紹介します。 ErrorBench は、143 の CoNLL-2014 パッセージからの 4,290 の応答を対象に、5 つのプロンプト条件下で 6 つの最新の LLM を評価します。 CoNLL-2014 M2 スタイルのスコアリングでは、アンカーされたプロンプトは F1 インフレの最大 0.79 ポイントを生成し、厳密なマッチングでは最大 0.96 ポイントを生成します。公式 ERRANT 3.0.0 パイプラインとマルチリファレンス スコアリングを使用した 100 パッセージのレプリケーションによりパターンが再現されます。6 つのモデルの平均では、ブラインドからアンカーへのプロンプト シフトにより Count-F1 が +0.21 上昇する一方、マルチリファレンス ERRANT F0.5 は +0.04 しか上昇しません。この研究では、このストレステストプロトコルの下で、命令に高度に準拠した GPT/Claude システムではより大きなカウント応答が、Gemini ファミリーではより小さな応答が見出されました。この調査結果は、LLM 校正と文書レビュー評価では、事前に入力されたエラー数を回避し、カウントベースのメトリクスとともにスパン認識メトリクスを報告する必要があることを示唆しています。

原文 (English)

Prompt Framing Distorts Count-Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation. The paper introduces ErrorBench, a controlled stress-test protocol for prompt-induced count distortion. ErrorBench evaluates six contemporary LLMs under five prompt conditions over 4,290 responses from 143 CoNLL-2014 passages. Under CoNLL-2014 M2-style scoring, anchored prompts produce up to 0.79 points of F1 Inflation, and up to 0.96 under strict matching. A 100-passage replication using the official ERRANT 3.0.0 pipeline and multi-reference scoring reproduces the pattern: averaged over six models, the Blind-to-Anchored prompt shift raises Count-F1 by +0.21 while raising multi-reference ERRANT F0.5 by only +0.04. The study finds larger count responses in highly instruction-compliant GPT/Claude systems and smaller responses in the Gemini family under this stress-test protocol. The findings suggest that LLM proofreading and document-review evaluations should avoid pre-populated error counts and should report span-aware metrics alongside count-based metrics.

2026-07-03 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

大規模言語モデルの使用のための監査フレームワークの実践: 集団経験主義、疑似合理的認知、AI 生成コンテンツのガバナンス

大規模な言語モデルは、知識の獲得、コード生成、学術論文の執筆、エージェントベースの自動化にますます使用されています。このような設定では、ユーザーは十分な専門知識の実践なしに、高度に構造化された回答、計画、判断を得る可能性があります。このペーパーでは、LLM の使用と AI によって生成されたコンテンツ ガバナンスのための実践監査フレームワークを提案します。ここでは、LLM が大規模な人間の経験を経験的かつ合理的に見える出力にどのように圧縮および再編成するかを説明するために集団経験主義を導入し、ユーザーが AI によって生成された構造化された表現を自分自身の合理的な理解とどのように誤って認識するかを説明するために疑似合理的認知を導入しています。この論文では、AIの主観錯覚、入力資料の主観構造、AI-AI会話のテンプレートループ、AIGC検出における統計的誤判断、生成されたコンテンツが将来のコンテキスト、長期記憶、検索空間、またはエージェントスキルシステムに入るときのメモリ汚染を分析します。これらのリスクを軽減するために、この文書では、要件定義、問題境界の特定、証拠ソースの監査、実際的な検証、逆質問、ロギング、バージョン管理、ロールバック、および新たな認識に基づいた監査プロセスを提案しています。このフレームワークは AI の生産性を否定するものではありません。 LLM の成果は、検証可能、再現可能、介入可能な実践プロセスに戻されるべきだと主張しています。この論文は、LLM インタラクション、AI 生成コンテンツ ガバナンス、長期記憶システム、および人間と AI のインタラクションにおける認知リスクに関する概念的で監査可能なフレームワークを提供します。

原文 (English)

A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-Rational Cognition, and Governance of AI-Generated Content

Large language models are increasingly used for knowledge acquisition, code generation, academic writing, and agent-based automation. In these settings, users may obtain highly structured answers, plans, and judgments without sufficient domain practice. This paper proposes a practice auditing framework for LLM use and AI-generated content governance. It introduces collective empiricism to describe how LLMs compress and reorganize large-scale human experience into outputs that appear empirical and rational, and pseudo-rational cognition to describe how users may mistake AI-generated structured expression for their own rational understanding. The paper analyzes AI subjectivity illusion, subjectivity structures in input materials, template loops in AI-AI conversations, statistical misjudgment in AIGC detection, and memory pollution when generated content enters future contexts, long-term memory, retrieval spaces, or agent skill systems. To reduce these risks, the paper proposes an auditing process based on requirement definition, problem-boundary identification, evidence-source auditing, practical validation, reverse questioning, logging, version management, rollback, and renewed cognition. The framework does not reject AI productivity; it argues that LLM outputs should be returned to verifiable, reproducible, and intervenable processes of practice. The paper provides a conceptual and auditable framework for cognitive risks in LLM interaction, AI-generated content governance, long-term memory systems, and human-AI interaction.

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluatio…

2026-07-03 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vul…

2026-07-03 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual questi…

2026-07-03 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Beyond the Performance Illusion: Structure-Aware Stratified Partitioning and Curriculum Distributionally Robust Optimization for Spatially Correlated Domains

Performance evaluation in AI systems commonly assumes that random dataset splits produce independent and identically distributed (i.i.d.) s…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Prompt Coverage Adequacy

In recent years, it has become increasingly evident that large language models (LLMs) and autonomous agents raise the level of abstraction…

2026-07-03 13:00 JSTarXiv cs.AIビジネス/資金調達

The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits

The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional…

2026-07-03 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation prot…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry

Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation stru…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

LLM エージェントの機能を評価するための統一フレームワーク

LLM がエージェントとして導入されることが増えているため、そのエージェント機能の信頼できる評価が不可欠になっています。ただし、報告されるベンチマーク スコアは、多くの場合、モデルの機能と、各ベンチマークに含まれる実装の選択肢を合わせて反映するため、クロスベンチマークの結果を基礎となるモデルの正確な測定値として解釈することが困難になります。この研究では、LLM エージェントの機能を公正に評価するための統一フレームワークを紹介します。統合された構成システムによって駆動されるこのフレームワークは、標準化された命令、ツール、環境の形式に多様なベンチマークを統合し、制御可能なサンドボックス内の固定 ReAct スタイル アーキテクチャを通じてエージェントを実行します。また、フレームワークの効果と環境の効果を個別に分析できるように、揮発性のライブ環境を厳選されたスナップショットに置き換えるオプションのオフライン設定を提供します。これに基づいて、各ベンチマークの元のタスクの成功基準に基づいて評価方法を統一するとともに、リソース消費に関する統一された指標と、意思決定レベルおよび実行レベルの失敗の属性に関する分類を導入します。このフレームワーク内で、シングルエージェント、マルチエージェント、およびセーフティクリティカルなシナリオにわたる 24 のドメインにわたる 7 つの広く使用されているベンチマークを適応させ、15 のモデルで 400,000 のロールアウトと 50 億のトークンにわたる大規模な実証分析を実施します。結果は、足場の選択と環境の変動性がベンチマークの結果を両方向に実質的に変化させ、フレームワークおよび環境によって引き起こされるアーティファクトから本質的な LLM 機能を解きほぐすことをフレームワークが可能にすることを示しています。さらに、安全性が重要なドメインの安全なテストベッドとしての拡張性を実証します。コードとベンチマークは、https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities、https://huggingface.co/AgentFramework/Unified_Farmework で入手できます。

原文 (English)

A Unified Framework for the Evaluation of LLM Agentic Capabilities

As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model capability and the implementation choices each benchmark is packaged with, making cross-benchmark results difficult to interpret as clean measurements of the underlying model. In this work, we present a unified framework for the fair evaluation of LLM agentic capabilities. Driven by a unified configuration system, the framework integrates diverse benchmarks into a standardized instruction-tool-environment format, executes agents through a fixed ReAct-style architecture within a controllable sandbox, and provides an optional offline setting that replaces volatile live environments with curated snapshots, so that framework effects and environment effects can be analyzed separately. Building on this, we unify the evaluation methodology under each benchmark's original task-success criteria, while introducing unified metrics for resource consumption and a taxonomy for decision- and execution-level failure attribution. Within this framework, we adapt 7 widely used benchmarks spanning 24 domains across single-agent, multi-agent, and safety-critical scenarios, and conduct a large-scale empirical analysis over 400K rollouts and 5B tokens on 15 models. The results show that scaffold choice and environmental volatility materially shift benchmark outcomes in both directions, allowing our framework to disentangle intrinsic LLM capabilities from framework- and environment-induced artifacts. We further demonstrate its extensibility as a secure testbed for safety-critical domains. Codes and benchmarks at are available at https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities, https://huggingface.co/datasets/whfeLingYu/Unified_Agent_Framework.

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ベンチマーク監査における信頼性ギャップ: 汚染検出の障害モードとしての分布のシフトとスケール

評価例がモデルのトレーニング データに現れるベンチマーク汚染は、LLM 評価の妥当性を脅かします。トレーニング データのメンバーシップを検出するための統計ツールは存在しますが、ほぼ独占的に管理された学術体制、つまり大規模で均質な事前トレーニング コーパスと透明な単一ステージ トレーニング パイプラインでのみ検証されています。これらの方法が現実的な監査シナリオにおいて信頼性を維持できるかどうかは、依然として不明です。私たちは、十分に研究されていない 2 つの障害モードを特定します。1 つは、疑わしいセットと検証セットが IID の仮定に違反する場合に発生する分布シフト、もう 1 つは、ベンチマークがトレーニング前のコーパスよりも桁違いに小さいために発生するスケール制約です。私たちは、複数のファミリー (Pythia、OLMo~2、特殊な文化的および医療的 LLM を含む) およびスケール (最大 27B) からの 27 のモデルにわたって、LLM データセット推論、ポストホック データセット推論、CoDeC という 3 つの主要なパラダイムを体系的に評価します。次に、分析を最先端の業界モデルにさらに拡張します。 335 件の評価のうち、正しい結果が得られたのは 199 件のみでした。 LLM データセット推論では、分布シフトの下で偽陽性が発生し、ポストホック データセット推論はベンチマーク スケールでは能力が不足し、CoDeC は個々のベンチマーク分割を検証するには不十分な粗い出所信号しか提供しません。私たちの結果は、管理された検証と実際のベンチマーク監査の間に体系的な信頼性のギャップがあることを明らかにし、統計的検出がまだ透明なデータ来歴に取って代わることができないことを示しています。私たちはさらなる研究のためにベンチマークをオープンソースにしています。

原文 (English)

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for detecting training-data membership exist, but have been validated almost exclusively in controlled academic regimes: large, homogeneous pre-training corpora and transparent, single-stage training pipelines. Whether these methods remain reliable in realistic auditing scenarios remains unclear. We identify two under-studied failure modes: distribution shift, which arises when suspect and validation sets violate the IID assumption, and scale constraints, which arise because benchmarks are orders of magnitude smaller than pre-training corpora. We systematically evaluate three leading paradigms, LLM Dataset Inference, Post-Hoc Dataset Inference, and CoDeC, across 25 models from multiple families (including Pythia, OLMo 2, and specialised cultural and medical LLMs) and scales (up to 27B). We then further extend our analysis to frontier industry models. Across 335 evaluations, only 201 yield correct outcomes. LLM Dataset Inference results in false positives under distribution shift, Post-Hoc Dataset Inference is underpowered at benchmark scale, and CoDeC provides only coarse provenance signals that are insufficient to verify individual benchmark splits. Our results reveal a systematic reliability gap between controlled validation and practical benchmark auditing, and show that statistical detection cannot yet replace transparent data provenance. We open-source our benchmark for further research.

2026-07-03 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

Power Systems Agent Benchmark: Executable Evaluation of AI Agents in Electric Power Engineering

Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a promi…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation

Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a determin…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interac…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity

Creativity is often seen as a hallmark of human intelligence. While large language models(LLMs) are increasingly perceived as generating cr…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents

Large Language Models (LLMs) in multi-agent systems (MAS) have shown promise for complex tasks, yet current training methods lack principle…

2026-07-03 13:00 JSTarXiv cs.AIビジネス/資金調達

Adaptive Contracts for Cost-Effective AI Delegation

When organizations delegate text generation tasks to AI providers via pay-for-performance contracts, expected payments rise when evaluation…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's langu…

2026-07-03 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

物理基礎モデルは一般化可能な物理学を学習しますか?物理的体制と分布の変化にわたるバイアスを意識したベンチマーク

最近の物理基礎モデルは一般的な時空間予測能力を主張していますが、その評価は、固定されたトレーニング分布の下でパフォーマンスを単一の平均スコアに落とし込んでしまうことがよくあります。このため、モデルが一般化可能な物理ダイナミクスを学習しているのか、それとも特定の設定下でのみ適切に動作するのかを判断することが困難になります。 8 つの物理ダイナミクス、3 つのトレーニング データ混合物、および動的スケールと初期条件の複雑さのシフトによって誘発される 25 のテスト レジームを使用してベンチマークを構築し、分布内、分布シフト、および分布外の設定をカバーします。 5 つの物理基礎モデル アーキテクチャとアーキテクチャごとに 4 つのモデル バリアント (スクラッチ サイズと 3 つの事前トレーニング サイズ) を評価し、結果として 60,000 の測定結果が得られます。私たちの結果は、現在の物理基礎モデルが普遍的なジェネラリストとしてではなく条件付きで動作することを示しています。その一般性は、物理レジーム、時間スケール、初期条件の設定、事前トレーニング、モデルのサイズ、アーキテクチャに依存します。トレーニング データの分散を改善しても、この制限は部分的にしか緩和されません。事前トレーニングとスケーリングでも、能力のバイアスを確実に取り除くことができません。私たちは、物理基礎モデルを改善するには、モデルのスケーリングやデータの拡張を超えて、領域、時間スケール、分布の変化を超えて移転可能な物理知識をより適切に捕捉する学習メカニズムに移行する必要があると主張します。

原文 (English)

Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts

Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a single average score under a fixed training distribution. This makes it difficult to determine whether a model has learned generalizable physical dynamics or only performs well under particular settings. We construct a benchmark with 8 physical dynamics, 3 training-data mixtures, and 25 test regimes induced by dynamic-scale and initial-condition complexity shifts, covering in-distribution, distribution-shift, and out-of-distribution settings. We evaluate five physics foundation model architectures and four model variants per architecture (scratch and three pretrained sizes), resulting in 60,000 measurements. Our results show that current physics foundation models behave as conditional rather than universal generalists: their generality depends on the physical regime, temporal scale, initial-condition setting, pretraining, model size, and architecture. Improving the training data distribution only partially mitigates this limitation. Pretraining and scaling are also unable to reliably remove their ability biases. We argue that improving physics foundation models requires moving beyond scaling models or expanding data, toward learning mechanisms that better capture transferable physical knowledge across regimes, temporal scales, and distribution shifts.

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MedCase 構造化: 臨床的に現実的な EHR 設定における診断推論のベンチマーク用の Text-to-FHIR データセット

大規模言語モデル (LLM) は、臨床推論と意思決定のサポートに有望ですが、現実的な電子医療記録と一致する設定での評価には依然として限界があります。既存のベンチマークは、多くの場合、臨床システムで使用される構造化された相互運用可能なデータ形式を反映していない静的データセットまたは非構造化入力に依存しています。非構造化テキストから臨床的に現実的な HL7 FHIR R4 バンドルを生成するパイプラインを導入し、臨床意思決定支援システムの制御可能な評価を可能にします。このパイプラインは、段階的な LLM 生成と用語に基づいた検証および修復を組み合わせて、幻覚コードを削減し、構造的および意味的な一貫性を強化します。このアプローチを MedCaseReasoning に適用して、臨床医が作成した診断症例に合わせた合成データセットである MedCase-Structured を構築し、症例の 82.5% で有効な FHIR 生成を実現します。 MedCase-Structured での評価では、平文の場合よりも構造化 FHIR 入力での LLM の診断精度が一貫して低いことが明らかになり、展開に合わせたベンチマークの重要性が強調されています。

原文 (English)

MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in structured, electronic health record-congruent settings remains limited. Existing benchmarks often rely on static datasets or unstructured inputs that do not reflect the interoperable data formats used in clinical systems. We introduce a reusable pipeline for generating terminology-grounded HL7 FHIR R4 bundles from unstructured text, enabling controllable evaluation of clinical decision support systems over structured inputs. The pipeline combines staged LLM generation with terminology-grounded validation and repair to eliminate hallucinated codes and enforce structural and semantic consistency. Applying this approach to MedCaseReasoning, we construct MedCase-Structured, a synthetic dataset of 1,732 FHIR bundles derived from clinician-authored diagnostic cases, producing complete, valid bundles for 97.1% of attempted cases. Evaluation on MedCase-Structured reveals consistently lower diagnostic accuracy for LLMs on structured FHIR inputs than with plain text, highlighting the importance of deployment-aligned benchmarking.

2026-07-03 11:01 JSTITmedia AI+LLM/生成AIエージェントビジネス/資金調達

Meta、「Claude Codeと組織改編で爆速開発」のはずが「想定より加速せず」 ザッカーバーグ氏、社内集会で発言

MetaはAI向けの巨額投資や組織改編でAIエージェント開発の加速を図ったが、ザッカーバーグCEOによれば、それらの取り組みはまだ実を結んでいないようだ。

2026-07-03 05:11 JSTTechCrunch AIビジネス/資金調達

Jersey Mike’s IPO illustrates how bad the AI hype has become

Just for kicks, I took a look at Jersey Mike's IPO documents. Surely a sandwich shop would have no need to mention AI. But lo and behold.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SEFORA: フィードバック コーパスと LLM フィードバック評価フレームワークを使用した学生のエッセイ

効果的なフィードバックの作成は生徒の学習を促進する最も強力な推進力の 1 つですが、それを大規模に作成するには多大な労力がかかります。 LLM はライティングサポートを拡張するための自然な道筋を提供しますが、2 つのギャップが邪魔をしています。実際の教室で講師が実際にどのようにフィードバックを提供するかを記録した公開コーパスがほとんどないこと、生成されたフィードバックが講師が書く内容と一致しているかどうかを測定する信頼できる方法がないことです。私たちは両方に対応します。 SEFORA は、課題プロンプト、ルーブリック、スコア、大学のさまざまな執筆ジャンルにわたる複数の下書き改訂を備えた、講師のインライン フィードバックを組み合わせたパブリック コーパスであり、564 の下書きと 8,240 の講師の注釈で構成されています。 UniMatch は、オープンエンド生成のための参照ベースの評価フレームワークです。フィードバックをフィードバック単位に分割し、インストラクターが導き出した基準に基づいて意味論的な対応をスコアリングし、最適なマッチングによってそれらを調整して、解釈可能な精度、再現率、および F1 を生成します。複数の LLM にわたる 74 の実験構成全体で、0.4 F1 を超える設定はありませんでした。 UniMatch は、モデルがインストラクターが優先するフィードバックを特定するのに苦労しており、モデルが生成するフィードバックが増えるとパフォーマンスが低下することを明らかにしました。

原文 (English)

SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a natural path to scaling writing support, but two gaps stand in the way: few public corpora capture how instructors actually deliver feedback in real classrooms, and no reliable method measures whether generated feedback aligns with what an instructor would write. We address both. SEFORA is a public corpus pairing instructor inline feedback with assignment prompts, rubrics, scores, and multi-draft revisions across various college writing genres, comprising 564 drafts and 8,240 instructor annotations. UniMatch is a reference-based evaluation framework for open-ended generation: it segments feedback into feedback units, scores their semantic correspondence under instructor-derived criteria, and aligns them via optimal matching to yield interpretable precision, recall, and F1. Across 74 experimental configurations spanning multiple LLMs, no setting exceeds 0.4 F1. UniMatch reveals that models struggle to identify the feedback instructors would prioritize, and performance degrades as models generate more.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

評価フロンティアのマッピング: 11 の評価者とエージェントの条件にわたるバイアスと信頼性のトレードオフの実証的調査

バイアスと信頼性のトレードオフは、LLM 評価システムが (ガンマ、H、CV) 空間に制約され、評価者の結合 (ガンマ)、戦略多様性 (H)、および小サンプル測定の信頼性 (CV(N)) を固定サンプル サイズ N で同時に最適化することができないと推測します。以前の証拠は、単一の研究からの完全なメトリクスを備えた n=5 の条件に基づいています。経験ベースを 11 の条件に拡張し、11 条件すべて (有効な重みベクトルを持つ 9 つ) についてガンマと H を測定し、十分なシード (N >= 5) がある 7 つについて CV(N=5) を測定します。 5 つの条件により、完全な (ガンマ、H、CV) トリプルが提供されます。データはトレードオフを裏付けています。評価者の結合が低い条件 (ガンマ 1.0) では結合が強い (ガンマ > 0.9) 条件では低ノイズ (CV(N=5) < 0.16) が達成されます。相関 r(H, gamma) = -0.989 (n=5、GPT-4o 条件を除く) は、評価者の結合が戦略の多様性を抑制することを裏付けています。 4 つの GPT-4o 条件では、すべてのシードでガンマ = 0.000 および H = 1.000 が示されています。このパターンは、2026 年 6 月の GPT-4o API のバージョン ドリフトによるものであると考えられます。 {γ < 0.2, CV(N=5) < 0.3} の領域を占める条件はありません。すべての条件ごとのメトリクスを、評価者が比較できるように標準化されたベンチマーク データセットとしてリリースします。

原文 (English)

Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N. Prior evidence rests on n=5 conditions with complete metrics from a single study. We expand the empirical base to 11 conditions, measuring gamma and H for all 11 (nine with valid weight vectors) and CV(N=5) for seven with sufficient seeds (N >= 5). Five conditions provide the complete (gamma, H, CV) triple. The data confirm the trade-off: conditions with low evaluator coupling (gamma 1.0), while conditions with strong coupling (gamma > 0.9) achieve low noise (CV(N=5) < 0.16). The correlation r(H, gamma) = -0.989 (n=5, excluding GPT-4o conditions) confirms that evaluator coupling suppresses strategy diversity. Four GPT-4o conditions show gamma=0.000 and H=1.000 across all seeds -- a pattern we attribute to version drift in the June 2026 GPT-4o API. No condition occupies the region {gamma < 0.2, CV(N=5) < 0.3}. We release all per-condition metrics as a standardized benchmark dataset for evaluator comparison.

2026-07-02 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

AI、信頼、チーム化: 自律的かつ不透明な AI システムのためのハンドラーとしての人間アプローチ

人工知能 (AI) はユビキタスになりつつあり、ドメイン全体でますます自律的なシステムが、重大な倫理的および法的課題を引き起こすタスクを実行するようになっており、これは信頼に根ざした強力な人間と機械のチームの必要性を示しています。この記事では、非常に影響力の大きい分野(医学や戦闘など)において、私たちが当初、自律的で不透明なシステムを犬(または私たちが密接な関係にある他の動物)に類似したものとして扱う根拠があると主張します。このアナロジーの下では、これらのシステムを利用する人間は、これらのシステムの「ユーザー」または「導入者」とみなされるべきではなく、代わりに「ハンドラー」の役割を引き受けます。この役割の再設定により、人間、AI 対応および自律システム、およびそれらの間の関係に対する見方が変わり、さらに、これらのシステムの使用時にもたらされる結果に対して人間が負う明確かつ追跡可能な責任範囲が明確になります。この点を展開する際に、私は機械と動物のアナロジーが確かに非類似要素を認めているが、その接点がそれを出発点として基礎づけていることを明確にしました。次に、自律システムや AI 対応システムとの関わり方や利用方法に不適当な、動物との関係の側面における、人間がハンドラーとしてのアプローチをどのように取り除くことができるかを模索します。私は、自律システムと AI 対応システムのための人間と機械のチーム化の軌跡は、これらを単に利用する成果物としてではなく、複雑な目標を追求し、複雑なタスクを実行する協力者として真に見なす状態でなければならないと主張して、結論を述べます。

原文 (English)

AI, Trust, and Teaming: The Humans-as-Handlers Approach for Autonomous and Opaque AI Systems

Artificial intelligence (AI) is becoming ubiquitous, and across domains, increasingly autonomous systems are carrying out tasks which raise significant ethical and legal challenges which demonstrate a need for strong human-machine teams rooted in trust. In this article, I argue that within highly impactful areas (such as medicine or warfighting) there are grounds for us initially treating autonomous and opaque systems as relevantly analogous to dogs (or other animals with which we have close relationships). Under this analogy, humans making use of these systems are not to be viewed as "users" or "deployers" of these systems, but instead take the role of "handlers". This recasting of roles shifts the way we view humans, AI-enabled and autonomous systems, and the relations between them, and moreover clarifies the clear and traceable lines of responsibility humans have for the outcomes brought about when using these systems. In developing this point, I clarify that the machine-animal analogy does admit disanalogous elements, but that its touch-points ground it as a starting point. I then explore how we can divest the humans-as-handlers approach of those aspects of our relationships with animals which are unfitting for how we engage with and make use of autonomous and AI-enabled systems. I conclude by arguing that the trajectory of human-machine teamings for autonomous and AI-enabled systems should be a state where we authentically view these not as artifacts which we simply make use of, but as collaborators with which we pursue complex goals and carry out complex tasks.

2026-07-02 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

LongVQUBench: 長期ビデオ品質のベンチマーク 視覚言語モデルの理解

長期的なビデオ品質の理解の評価は、大規模ビジョン言語モデル (LVLM) にとって未解決の課題のままです。既存のビデオ品質ベンチマークは、主に短いクリップと孤立した歪みに焦点を当てており、時間的な連続性、累積的な劣化、および長時間コンテンツに固有の推論の複雑さを見落としています。これらの制限に対処するために、長期的なビデオ品質を理解するための包括的なベンチマークである LongVQUBench を紹介します。 LongVQUBench には、映画、ドキュメンタリー、監視映像、自己中心的な録画、アニメーション コンテンツにわたる 1,200 を超える多様なビデオが含まれており、検証とテストのための 1,500 の多肢選択式の自由形式の質問も含まれています。さまざまな時間的範囲にわたる知覚推論を評価するために、段階的に複雑になる 3 つの評価レベルを導入します。(i) 局所的な歪みを分析するためのローカル イベント品質理解 (LQU)。 (ii) 複数の劣化したイベントを統合するためのクロスイベント品質推論 (CQR)。 (iii) 長期間にわたる全体的な知覚評価のためのグローバル品質理解 (GQU)。さらに、ニードル ディストーション質問応答 (NDQA) パラダイムが 3 つのレベルすべてに組み込まれており、空間的または時間的アーティファクトがまばらに挿入されて、きめの細かい検出および推論機能が調査されます。 14 の最先端の LVLM に関する広範な実験により、ビデオの長さと推論の深さが増加するにつれてパフォーマンスが大幅に低下することが明らかになり、長距離の時間統合と知覚的帰属に対する能力の限界が浮き彫りになりました。私たちは、LongVQUBench が、LVLM の長期的なビデオ品質の理解を体系的かつ階層的に説明可能な評価に向けた基礎的なステップとして構想しています。

原文 (English)

LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking the temporal continuity, cumulative degradation, and reasoning complexity inherent in long-duration content. To address these limitations, we present LongVQUBench, a comprehensive benchmark for long-term video quality understanding. LongVQUBench contains over 1200 diverse videos spanning movies, documentaries, surveillance footage, egocentric recordings, and animated content, accompanied by 1500 multiple-choice and open-ended questions for validation and testing. To assess perceptual reasoning across different temporal scopes, we introduce three progressively complex evaluation levels: (i) local event quality understanding (LQU) for analyzing localized distortions; (ii) cross-event quality reasoning (CQR) for integrating multiple degraded events; and (iii) global quality understanding (GQU) for holistic perceptual evaluation over extended durations. Furthermore, a needle distortion question-answering (NDQA) paradigm is embedded across all three levels, where spatial or temporal artifacts are sparsely inserted to probe fine-grained detection and reasoning capabilities. Extensive experiments on 14 state-of-the-art LVLMs reveal significant performance degradation with increasing video length and reasoning depth, highlighting their limited capacity for long-range temporal integration and perceptual attribution. We envision LongVQUBench as a foundational step toward the systematic, hierarchical, and explainable evaluation of LVLMs' long-term video quality understanding.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

AI の安全性評価のための敵対的プラグマティクス: 命令の競合、埋め込みコマンド、およびポリシーの曖昧さのベンチマーク

言語モデルの安全性評価は、モデルが指示に従ったか、適切に拒否したか、ポリシーに従ったか、埋め込まれたコマンドに抵抗したか、エージェントタスクの進捗状況を誤って報告したかなど、あいまいな自然言語の動作に関する判断にますます依存しています。既存のベンチマークは、多くの場合、これらの区別を合格/不合格のラベルに圧縮し、障害が機能制限、ポリシーの曖昧さ、命令の競合、足場の障害、または不安定な評価者の判断に起因するかどうかを曖昧にします。この論文では、命令の競合、埋め込みコマンド、引用、範囲の曖昧さ、明確化、間接音声行為、およびマルチターンエージェントのトランスクリプトの下でのモデルの動作を評価するためのベンチマークおよびアノテーションプロトコルとして、敵対的プラグマティクスを紹介します。この貢献は経験的かつ方法論的です。言語的に管理された分類法、バリデーターが強制するメタデータを備えた 18 項目のシード ベンチマーク、54 行のローカル シード パイロット、タスクの成功、ポリシー遵守、安全性リスク、拒否結果、評価者の信頼性を区別する専門家評価プロトコル、および裁判官の妥当性、診断の曖昧さ、および分類のドリフトの指標です。このフレームワークは、言語的判断方法論を、安全性評価、LLM ジャッジ、ゴールドセット構造、即時注入テスト、および安全性文書を検証するための実用的なツールに変えます。

原文 (English)

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological: a linguistically controlled taxonomy, an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, an expert-evaluation protocol distinguishing task success, policy compliance, safety risk, refusal outcome, and evaluator confidence, and metrics for judge validity, diagnostic ambiguity, and taxonomy drift. The framework turns linguistic judgment methodology into a practical tool for validating safety evals, LLM judges, gold-set construction, prompt-injection tests, and safety documentation.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

人間の研究アイデアと LLM 研究アイデアの間のギャップを測定する

研究アイデアのブレインストーミングに LLM が使用されることが増えていますが、既存の評価では主に、新規性、実現可能性、または専門家の好みによって個々のアイデアが判断されます。私たちは代わりに、現在の LLM によって生成されたアイデアが人間の研究者からどれだけ離れているのかを尋ねます。このギャップを特徴づけるために、私たちは質の高い人間の研究論文から着想を得るための大規模な評価フレームワークを構築します。各論文について、その中心となるアイデアにインスピレーションを与えたと考えられる密接に関連した先行研究の小さなセットをリバース エンジニアリングします。次に、LLM は論文のタイトルと要約のセットから新しいアイデアを生成するように求められます。私たちは、2 軸のリサーチテイスト分類法を導入して、各アイデアをその機会パターンと研究パラダイムによってプロファイルし、それを使用して人間のアイデアと LLM のアイデアの相違を定量化します。さまざまな LLM によって生成されたアイデア セット全体で、一貫した分布ギャップが観察されます。LLM のアイデアは、橋渡しのような機会と合成方法の周りに不均衡に集中していますが、人間の論文参照の分布は、ギャップを構成し、貢献を構築する方法全体にわたってより広範囲に広がっています。この結果は、強力な LLM がさまざまな合理的なアイデアを生み出すことができるが、その範囲は依然として人間の研究の好みよりも狭く、体系的にシフトしていることを示唆しています。

原文 (English)

Measuring the Gap Between Human and LLM Research Ideas

LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from human researchers? To characterize this gap, we build a large-scale evaluation framework for ideation from high-quality human research papers. For each paper, we reverse-engineer a small set of closely related prior works that likely inspired its core idea. LLMs are then prompted to generate a new idea from the set of paper titles and summaries. We introduce a two-axis research-taste taxonomy to profile each idea by its opportunity pattern and research paradigm, and use it to quantify the divergence between human and LLM ideas. Across idea sets generated by different LLMs, we observe a consistent distributional gap: LLM ideas are disproportionately concentrated around bridge-like opportunities and synthesis methods, whereas the human paper reference distribution spreads more broadly across ways of framing gaps and constructing contributions. This result suggests that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

LLM における推論の質の測定: 多次元の行動フレームワーク

LLM は複雑な推論タスクで目覚ましい成功を収めていますが、現在の評価アプローチは主に最終的な答えの正しさに依存しており、それらの答えを生み出す根本的な推論プロセスについての洞察は限られています。このギャップに対処するために、この研究では、動作の観点から LLM の推論品質を測定するための統一された多次元フレームワークを提案し、理論的に根拠のある 6 つの次元、正確性 (CQ)、一貫性 (CS)、堅牢性 (RS)、論理的一貫性 (LS)、効率 (ES)、安定性 (SS) を運用します。 4 つのベンチマークの 975 項目にわたる 7 つの LLM に関する広範な実験により、このフレームワークが精度のみの指標では見えない動作を明らかにすることが実証されました。特に、論理的一貫性は正しさ (r = -0.172、ns) と直交しており、一貫性のない推論から正しい答えが得られることが確認され、一方、Claude-Haiku-4.5 は最高の多次元スコア (Q_bal = 0.778) を達成しています。さらに、このフレームワークは重大なランキングの逆転を明らかにしています。DeepSeek-V3 は精度優先では 2 位ですが、法的/コンプライアンスの重み付けでは 5 位にランクされており、単一指標の評価では検出できない逆転です。判別式の妥当性により、11/15 次元のペアが独立している (|r| < 0.50) ことが確認され、各次元を別個の信号として扱うための心理測定的サポートが提供されます。フレームワークによって生成される次元プロファイルは、次の 3 つのクラスの展開決定を直接サポートします。最終的な答えが正しいにもかかわらず、その推論トレースが説明責任監査に失敗するモデルを特定します (LS--CQ 直交性)。精度のみのベンチマークによって引き起こされるランキングエラーを防止します。そして、フレームワークがキャプチャする 6 つの独立したシグナルを単一のメトリックが暗黙的に置き換えることがないようにします。

原文 (English)

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness, providing limited insight into how models reason, how reliably they behave under contextual variation, or how efficiently they reach conclusions. This paper proposes a unified multi-dimensional framework for measuring LLM reasoning quality from a behavioral perspective, operationalizing six theoretically grounded dimensions rooted in cognitive science: Correctness (CQ), Consistency (CS), Robustness (RS), Local Logical Coherence (LS), Efficiency (ES), and Stability (SS). The framework introduces deployment-aware aggregation, enabling context-specific model selection beyond accuracy-based leaderboards. Experiments across multiple LLMs and benchmarks reveal behaviors systematically concealed by single-metric evaluation, including the orthogonality of local logical coherence and correctness, deployment-context-dependent ranking inversions, and non-trivial dimensional profiles in small locally-deployed models. Discriminant validity analysis confirms that the proposed dimensions capture largely non-redundant signals. The resulting pipeline provides a foundation for diagnosing LLM reasoning behavior across deployment contexts, with domain-specific validation as a direction for future work.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

話す前に考える: マルチエージェント社会シミュレーションにおける内部評価から公の表現まで

LLM ベースのマルチエージェント シミュレーションは、社会的相互作用、熟慮、集団的な意見のダイナミクスを研究するための有望な方法を提供します。しかし、既存の対話シミュレーション フレームワークの多くは、対話を主に観察可能なターン交換または集約された出力として表現しており、沈黙、発言意図、公的表現の背後にある内部評価プロセスを調査することが困難なままになっています。エージェントの私的な推論を公的発話の生成から分離する、インターバルベースのマルチエージェント シミュレーション フレームワークである TBS (Think-Before-Speak) を紹介します。各間隔で、すべてのエージェントは共有された対話履歴と自身の記憶に基づいて構造化された内部状態を更新します。これらの状態には、不協和音関連の評価、認識された世論環境、認識された孤立リスク、対応戦略、および発言意欲が含まれます。その後、オーケストレーターは競合する発言意図を解決し、1 つの発言を公開対話にコミットし、内部評価と公開対話が時間の経過とともに共進化できるようにします。私たちは、気候関連の政策問題に関するタウンホールでの議論を模擬して TBS を評価します。結果は、TBS が一貫した内部状態トレースを生成し、これらのトレースがターン割り当て、沈黙、メモリ条件全体にわたって体系的に変化することを示しています。不協和音関連の評価はエージェントの発言意欲を高めますが、沈黙の圧力評価はそれを低下させます。発言の意図が形成されると、公の場での表現は主に順番の割り当てルールによって形成されます。これらの発見は、TBS が内部評価から公的表現への経路を観察可能かつ分析可能にすることで、メカニズムに敏感な社会シミュレーションをサポートしていることを示唆しています。

原文 (English)

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existing dialogue simulation frameworks represent interaction mainly as observable turn exchange or aggregated outputs, leaving the internal evaluative processes behind silence, speaking intention, and public expression difficult to examine. We introduce TBS (Think-Before-Speak), an interval-based multi-agent simulation framework that separates agents' private reasoning from public utterance generation. At each interval, all agents update structured internal states based on the shared dialogue history and their own memory. These states include dissonance-related appraisal, perceived opinion climate, perceived isolation risk, response strategy, and willingness to speak. The orchestrator then resolves competing speaking intentions and commits one utterance to the public dialogue, allowing internal evaluation and public interaction to co-evolve over time. We evaluate TBS in simulated town hall discussions on a climate-related policy issue. Results show that TBS produces coherent internal-state traces and that these traces vary systematically across turn-allocation, silence, and memory conditions. Dissonance-related appraisal increases agents' willingness to speak, whereas silence-pressure appraisal decreases it. Once speaking intention is formed, public expression is shaped mainly by turn-allocation rules. These findings suggest that TBS supports mechanism-sensitive social simulation by making the pathway from internal evaluation to public expression observable and analyzable.

2026-07-02 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning

Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are uncha…

2026-07-02 13:00 JSTarXiv cs.AIビジネス/資金調達

Vibecoding Ate My 宿題: グリーンフィールド ソフトウェア エンジニアリングとプログラミングへの AI アプローチの評価

生成 AI の急速な発展のおかげで、私たちはコンピューターとの対話方法を永遠に変える可能性のあるパラダイム シフトの真っ只中にいます。この分野の基礎知識なしにアプリケーションやコーディング インフラストラクチャを構築するための自然言語プロンプトの使用が増加していることが観察されており、この実践は「バイブ コーディング」と呼ばれています。これはおそらく、プログラミングの分野が当初から、考えられるあらゆるより高い抽象化レベルで構築されてきたものを表しています。 Vibe コーディングは、入力方法に関する限り、高レベル プログラミングのメタのエンドポイントとなることが約束されています。つまり、人間によるコード構文の使用が完全に排除され、母国語でのプログラミングが優先されます。このペーパーは、グリーンフィールドのソフトウェア エンジニアリング タスクにおける Vibe コーディングの実現可能性を評価し、そのソフトウェア エンジニアリングの能力を測定するために使用されたベンチマークを分析することを目的としています。この目的を達成するために、私たちは、Python で単純で個別のグリーンフィールド プログラミング タスクを実行する LLM の習熟度を分析し、この問題に関する範囲を絞った洞察を提供するための評価スイートを開発しました。

原文 (English)

Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming

Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forever. We have observed a growth in the use of natural language prompts to build applications and coding infrastructures without underlying knowledge of the field, and this practice has been dubbed `vibe coding.' It arguably represents what the field of programming has been building towards since the beginning, with every higher level of abstraction that is conceived. Vibe coding promises to be the endpoint for the meta of high-level programming as far as method of input is concerned: eliminating a human's use of code syntax entirely in favour of programming in their mother tongue. This paper aims to evaluate the viability of vibe coding for greenfield software engineering tasks, as well as analyse the benchmarks that have been used to measure its software engineering prowess. To this end, we have developed an evaluation suite for analysing an LLM's proficiency in carrying out simple, isolated greenfield programming tasks in Python to provide scoped insight on the matter.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SEFORA: フィードバック コーパスと LLM フィードバック評価フレームワークを使用した学生のエッセイ

効果的なフィードバックの作成は生徒の学習を促進する最も強力な推進力の 1 つですが、それを大規模に作成するには多大な労力がかかります。 LLM はライティングサポートを拡張するための自然な道筋を提供しますが、2 つのギャップが邪魔をしています。実際の教室で講師が実際にどのようにフィードバックを提供するかを記録した公開コーパスがほとんどないこと、生成されたフィードバックが講師が書く内容と一致しているかどうかを測定する信頼できる方法がないことです。私たちは両方に対応します。 SEFORA は、課題プロンプト、ルーブリック、スコア、大学のさまざまな執筆ジャンルにわたる複数の下書き改訂を備えた、講師のインライン フィードバックを組み合わせたパブリック コーパスであり、564 の下書きと 8,240 の講師の注釈で構成されています。 UniMatch は、オープンエンド生成のための参照ベースの評価フレームワークです。フィードバックをフィードバック単位に分割し、インストラクターが導き出した基準に基づいて意味論的な対応をスコアリングし、最適なマッチングによってそれらを調整して、解釈可能な精度、再現率、および F1 を生成します。複数の LLM にわたる 74 の実験構成全体で、0.4 F1 を超える設定はありませんでした。 UniMatch は、モデルがインストラクターが優先するフィードバックを特定するのに苦労しており、モデルが生成するフィードバックが増えるとパフォーマンスが低下することを明らかにしました。

原文 (English)

SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a natural path to scaling writing support, but two gaps stand in the way: few public corpora capture how instructors actually deliver feedback in real classrooms, and no reliable method measures whether generated feedback aligns with what an instructor would write. We address both. SEFORA is a public corpus pairing instructor inline feedback with assignment prompts, rubrics, scores, and multi-draft revisions across various college writing genres, comprising 564 drafts and 8,240 instructor annotations. UniMatch is a reference-based evaluation framework for open-ended generation: it segments feedback into feedback units, scores their semantic correspondence under instructor-derived criteria, and aligns them via optimal matching to yield interpretable precision, recall, and F1. Across 74 experimental configurations spanning multiple LLMs, no setting exceeds 0.4 F1. UniMatch reveals that models struggle to identify the feedback instructors would prioritize, and performance degrades as models generate more.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

評価フロンティアのマッピング: 11 の評価者とエージェントの条件にわたるバイアスと信頼性のトレードオフの実証的調査

バイアスと信頼性のトレードオフは、LLM 評価システムが (ガンマ、H、CV) 空間に制約され、評価者の結合 (ガンマ)、戦略多様性 (H)、および小サンプル測定の信頼性 (CV(N)) を固定サンプル サイズ N で同時に最適化することができないと推測します。以前の証拠は、単一の研究からの完全なメトリクスを備えた n=5 の条件に基づいています。経験ベースを 11 の条件に拡張し、11 条件すべて (有効な重みベクトルを持つ 9 つ) についてガンマと H を測定し、十分なシード (N >= 5) がある 7 つについて CV(N=5) を測定します。 5 つの条件により、完全な (ガンマ、H、CV) トリプルが提供されます。データはトレードオフを裏付けています。評価者の結合が低い条件 (ガンマ 1.0) では結合が強い (ガンマ > 0.9) 条件では低ノイズ (CV(N=5) < 0.16) が達成されます。相関 r(H, gamma) = -0.989 (n=5、GPT-4o 条件を除く) は、評価者の結合が戦略の多様性を抑制することを裏付けています。 4 つの GPT-4o 条件では、すべてのシードでガンマ = 0.000 および H = 1.000 が示されています。このパターンは、2026 年 6 月の GPT-4o API のバージョン ドリフトによるものであると考えられます。 {γ < 0.2, CV(N=5) < 0.3} の領域を占める条件はありません。すべての条件ごとのメトリクスを、評価者が比較できるように標準化されたベンチマーク データセットとしてリリースします。

原文 (English)

Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N. Prior evidence rests on n=5 conditions with complete metrics from a single study. We expand the empirical base to 11 conditions, measuring gamma and H for all 11 (nine with valid weight vectors) and CV(N=5) for seven with sufficient seeds (N >= 5). Five conditions provide the complete (gamma, H, CV) triple. The data confirm the trade-off: conditions with low evaluator coupling (gamma 1.0), while conditions with strong coupling (gamma > 0.9) achieve low noise (CV(N=5) < 0.16). The correlation r(H, gamma) = -0.989 (n=5, excluding GPT-4o conditions) confirms that evaluator coupling suppresses strategy diversity. Four GPT-4o conditions show gamma=0.000 and H=1.000 across all seeds -- a pattern we attribute to version drift in the June 2026 GPT-4o API. No condition occupies the region {gamma < 0.2, CV(N=5) < 0.3}. We release all per-condition metrics as a standardized benchmark dataset for evaluator comparison.

2026-07-02 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

AI、信頼、チーム化: 自律的かつ不透明な AI システムのためのハンドラーとしての人間アプローチ

人工知能 (AI) はユビキタスになりつつあり、ドメイン全体でますます自律的なシステムが、重大な倫理的および法的課題を引き起こすタスクを実行するようになっており、これは信頼に根ざした強力な人間と機械のチームの必要性を示しています。この記事では、非常に影響力の大きい分野(医学や戦闘など)において、私たちが当初、自律的で不透明なシステムを犬(または私たちが密接な関係にある他の動物)に類似したものとして扱う根拠があると主張します。このアナロジーの下では、これらのシステムを利用する人間は、これらのシステムの「ユーザー」または「導入者」とみなされるべきではなく、代わりに「ハンドラー」の役割を引き受けます。この役割の再設定により、人間、AI 対応および自律システム、およびそれらの間の関係に対する見方が変わり、さらに、これらのシステムの使用時にもたらされる結果に対して人間が負う明確かつ追跡可能な責任範囲が明確になります。この点を展開する際に、私は機械と動物のアナロジーが確かに非類似要素を認めているが、その接点がそれを出発点として基礎づけていることを明確にしました。次に、自律システムや AI 対応システムとの関わり方や利用方法に不適当な、動物との関係の側面における、人間がハンドラーとしてのアプローチをどのように取り除くことができるかを模索します。私は、自律システムと AI 対応システムのための人間と機械のチーム化の軌跡は、これらを単に利用する成果物としてではなく、複雑な目標を追求し、複雑なタスクを実行する協力者として真に見なす状態でなければならないと主張して、結論を述べます。

原文 (English)

AI, Trust, and Teaming: The Humans-as-Handlers Approach for Autonomous and Opaque AI Systems

Artificial intelligence (AI) is becoming ubiquitous, and across domains, increasingly autonomous systems are carrying out tasks which raise significant ethical and legal challenges which demonstrate a need for strong human-machine teams rooted in trust. In this article, I argue that within highly impactful areas (such as medicine or warfighting) there are grounds for us initially treating autonomous and opaque systems as relevantly analogous to dogs (or other animals with which we have close relationships). Under this analogy, humans making use of these systems are not to be viewed as "users" or "deployers" of these systems, but instead take the role of "handlers". This recasting of roles shifts the way we view humans, AI-enabled and autonomous systems, and the relations between them, and moreover clarifies the clear and traceable lines of responsibility humans have for the outcomes brought about when using these systems. In developing this point, I clarify that the machine-animal analogy does admit disanalogous elements, but that its touch-points ground it as a starting point. I then explore how we can divest the humans-as-handlers approach of those aspects of our relationships with animals which are unfitting for how we engage with and make use of autonomous and AI-enabled systems. I conclude by arguing that the trajectory of human-machine teamings for autonomous and AI-enabled systems should be a state where we authentically view these not as artifacts which we simply make use of, but as collaborators with which we pursue complex goals and carry out complex tasks.

2026-07-02 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

LongVQUBench: 長期ビデオ品質のベンチマーク 視覚言語モデルの理解

長期的なビデオ品質の理解の評価は、大規模ビジョン言語モデル (LVLM) にとって未解決の課題のままです。既存のビデオ品質ベンチマークは、主に短いクリップと孤立した歪みに焦点を当てており、時間的な連続性、累積的な劣化、および長時間コンテンツに固有の推論の複雑さを見落としています。これらの制限に対処するために、長期的なビデオ品質を理解するための包括的なベンチマークである LongVQUBench を紹介します。 LongVQUBench には、映画、ドキュメンタリー、監視映像、自己中心的な録画、アニメーション コンテンツにわたる 1,200 を超える多様なビデオが含まれており、検証とテストのための 1,500 の多肢選択式の自由形式の質問も含まれています。さまざまな時間的範囲にわたる知覚推論を評価するために、段階的に複雑になる 3 つの評価レベルを導入します。(i) 局所的な歪みを分析するためのローカル イベント品質理解 (LQU)。 (ii) 複数の劣化したイベントを統合するためのクロスイベント品質推論 (CQR)。 (iii) 長期間にわたる全体的な知覚評価のためのグローバル品質理解 (GQU)。さらに、ニードル ディストーション質問応答 (NDQA) パラダイムが 3 つのレベルすべてに組み込まれており、空間的または時間的アーティファクトがまばらに挿入されて、きめの細かい検出および推論機能が調査されます。 14 の最先端の LVLM に関する広範な実験により、ビデオの長さと推論の深さが増加するにつれてパフォーマンスが大幅に低下することが明らかになり、長距離の時間統合と知覚的帰属に対する能力の限界が浮き彫りになりました。私たちは、LongVQUBench が、LVLM の長期的なビデオ品質の理解を体系的かつ階層的に説明可能な評価に向けた基礎的なステップとして構想しています。

原文 (English)

LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking the temporal continuity, cumulative degradation, and reasoning complexity inherent in long-duration content. To address these limitations, we present LongVQUBench, a comprehensive benchmark for long-term video quality understanding. LongVQUBench contains over 1200 diverse videos spanning movies, documentaries, surveillance footage, egocentric recordings, and animated content, accompanied by 1500 multiple-choice and open-ended questions for validation and testing. To assess perceptual reasoning across different temporal scopes, we introduce three progressively complex evaluation levels: (i) local event quality understanding (LQU) for analyzing localized distortions; (ii) cross-event quality reasoning (CQR) for integrating multiple degraded events; and (iii) global quality understanding (GQU) for holistic perceptual evaluation over extended durations. Furthermore, a needle distortion question-answering (NDQA) paradigm is embedded across all three levels, where spatial or temporal artifacts are sparsely inserted to probe fine-grained detection and reasoning capabilities. Extensive experiments on 14 state-of-the-art LVLMs reveal significant performance degradation with increasing video length and reasoning depth, highlighting their limited capacity for long-range temporal integration and perceptual attribution. We envision LongVQUBench as a foundational step toward the systematic, hierarchical, and explainable evaluation of LVLMs' long-term video quality understanding.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

AI の安全性評価のための敵対的プラグマティクス: 命令の競合、埋め込みコマンド、およびポリシーの曖昧さのベンチマーク

言語モデルの安全性評価は、モデルが指示に従ったか、適切に拒否したか、ポリシーに従ったか、埋め込まれたコマンドに抵抗したか、エージェントタスクの進捗状況を誤って報告したかなど、あいまいな自然言語の動作に関する判断にますます依存しています。既存のベンチマークは、多くの場合、これらの区別を合格/不合格のラベルに圧縮し、障害が機能制限、ポリシーの曖昧さ、命令の競合、足場の障害、または不安定な評価者の判断に起因するかどうかを曖昧にします。この論文では、命令の競合、埋め込みコマンド、引用、範囲の曖昧さ、明確化、間接音声行為、およびマルチターンエージェントのトランスクリプトの下でのモデルの動作を評価するためのベンチマークおよびアノテーションプロトコルとして、敵対的プラグマティクスを紹介します。この貢献は経験的かつ方法論的です。言語的に管理された分類法、バリデーターが強制するメタデータを備えた 18 項目のシード ベンチマーク、54 行のローカル シード パイロット、タスクの成功、ポリシー遵守、安全性リスク、拒否結果、評価者の信頼性を区別する専門家評価プロトコル、および裁判官の妥当性、診断の曖昧さ、および分類のドリフトの指標です。このフレームワークは、言語的判断方法論を、安全性評価、LLM ジャッジ、ゴールドセット構造、即時注入テスト、および安全性文書を検証するための実用的なツールに変えます。

原文 (English)

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological: a linguistically controlled taxonomy, an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, an expert-evaluation protocol distinguishing task success, policy compliance, safety risk, refusal outcome, and evaluator confidence, and metrics for judge validity, diagnostic ambiguity, and taxonomy drift. The framework turns linguistic judgment methodology into a practical tool for validating safety evals, LLM judges, gold-set construction, prompt-injection tests, and safety documentation.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

人間の研究アイデアと LLM 研究アイデアの間のギャップを測定する

研究アイデアのブレインストーミングに LLM が使用されることが増えていますが、既存の評価では主に、新規性、実現可能性、または専門家の好みによって個々のアイデアが判断されます。私たちは代わりに、現在の LLM によって生成されたアイデアが人間の研究者からどれだけ離れているのかを尋ねます。このギャップを特徴づけるために、私たちは質の高い人間の研究論文から着想を得るための大規模な評価フレームワークを構築します。各論文について、その中心となるアイデアにインスピレーションを与えたと考えられる密接に関連した先行研究の小さなセットをリバース エンジニアリングします。次に、LLM は論文のタイトルと要約のセットから新しいアイデアを生成するように求められます。私たちは、2 軸のリサーチテイスト分類法を導入して、各アイデアをその機会パターンと研究パラダイムによってプロファイルし、それを使用して人間のアイデアと LLM のアイデアの相違を定量化します。さまざまな LLM によって生成されたアイデア セット全体で、一貫した分布ギャップが観察されます。LLM のアイデアは、橋渡しのような機会と合成方法の周りに不均衡に集中していますが、人間の論文参照の分布は、ギャップを構成し、貢献を構築する方法全体にわたってより広範囲に広がっています。この結果は、強力な LLM がさまざまな合理的なアイデアを生み出すことができるが、その範囲は依然として人間の研究の好みよりも狭く、体系的にシフトしていることを示唆しています。

原文 (English)

Measuring the Gap Between Human and LLM Research Ideas

LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from human researchers? To characterize this gap, we build a large-scale evaluation framework for ideation from high-quality human research papers. For each paper, we reverse-engineer a small set of closely related prior works that likely inspired its core idea. LLMs are then prompted to generate a new idea from the set of paper titles and summaries. We introduce a two-axis research-taste taxonomy to profile each idea by its opportunity pattern and research paradigm, and use it to quantify the divergence between human and LLM ideas. Across idea sets generated by different LLMs, we observe a consistent distributional gap: LLM ideas are disproportionately concentrated around bridge-like opportunities and synthesis methods, whereas the human paper reference distribution spreads more broadly across ways of framing gaps and constructing contributions. This result suggests that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

LLM における推論の質の測定: 多次元の行動フレームワーク

LLM は複雑な推論タスクで目覚ましい成功を収めていますが、現在の評価アプローチは主に最終的な答えの正しさに依存しており、それらの答えを生み出す根本的な推論プロセスについての洞察は限られています。このギャップに対処するために、この研究では、動作の観点から LLM の推論品質を測定するための統一された多次元フレームワークを提案し、理論的に根拠のある 6 つの次元、正確性 (CQ)、一貫性 (CS)、堅牢性 (RS)、論理的一貫性 (LS)、効率 (ES)、安定性 (SS) を運用します。 4 つのベンチマークの 975 項目にわたる 7 つの LLM に関する広範な実験により、このフレームワークが精度のみの指標では見えない動作を明らかにすることが実証されました。特に、論理的一貫性は正しさ (r = -0.172、ns) と直交しており、一貫性のない推論から正しい答えが得られることが確認され、一方、Claude-Haiku-4.5 は最高の多次元スコア (Q_bal = 0.778) を達成しています。さらに、このフレームワークは重大なランキングの逆転を明らかにしています。DeepSeek-V3 は精度優先では 2 位ですが、法的/コンプライアンスの重み付けでは 5 位にランクされており、単一指標の評価では検出できない逆転です。判別式の妥当性により、11/15 次元のペアが独立している (|r| < 0.50) ことが確認され、各次元を別個の信号として扱うための心理測定的サポートが提供されます。フレームワークによって生成される次元プロファイルは、次の 3 つのクラスの展開決定を直接サポートします。最終的な答えが正しいにもかかわらず、その推論トレースが説明責任監査に失敗するモデルを特定します (LS--CQ 直交性)。精度のみのベンチマークによって引き起こされるランキングエラーを防止します。そして、フレームワークがキャプチャする 6 つの独立したシグナルを単一のメトリックが暗黙的に置き換えることがないようにします。

原文 (English)

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness, providing limited insight into how models reason, how reliably they behave under contextual variation, or how efficiently they reach conclusions. This paper proposes a unified multi-dimensional framework for measuring LLM reasoning quality from a behavioral perspective, operationalizing six theoretically grounded dimensions rooted in cognitive science: Correctness (CQ), Consistency (CS), Robustness (RS), Local Logical Coherence (LS), Efficiency (ES), and Stability (SS). The framework introduces deployment-aware aggregation, enabling context-specific model selection beyond accuracy-based leaderboards. Experiments across multiple LLMs and benchmarks reveal behaviors systematically concealed by single-metric evaluation, including the orthogonality of local logical coherence and correctness, deployment-context-dependent ranking inversions, and non-trivial dimensional profiles in small locally-deployed models. Discriminant validity analysis confirms that the proposed dimensions capture largely non-redundant signals. The resulting pipeline provides a foundation for diagnosing LLM reasoning behavior across deployment contexts, with domain-specific validation as a direction for future work.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

話す前に考える: マルチエージェント社会シミュレーションにおける内部評価から公の表現まで

LLM ベースのマルチエージェント シミュレーションは、社会的相互作用、熟慮、集団的な意見のダイナミクスを研究するための有望な方法を提供します。しかし、既存の対話シミュレーション フレームワークの多くは、対話を主に観察可能なターン交換または集約された出力として表現しており、沈黙、発言意図、公的表現の背後にある内部評価プロセスを調査することが困難なままになっています。エージェントの私的な推論を公的発話の生成から分離する、インターバルベースのマルチエージェント シミュレーション フレームワークである TBS (Think-Before-Speak) を紹介します。各間隔で、すべてのエージェントは共有された対話履歴と自身の記憶に基づいて構造化された内部状態を更新します。これらの状態には、不協和音関連の評価、認識された世論環境、認識された孤立リスク、対応戦略、および発言意欲が含まれます。その後、オーケストレーターは競合する発言意図を解決し、1 つの発言を公開対話にコミットし、内部評価と公開対話が時間の経過とともに共進化できるようにします。私たちは、気候関連の政策問題に関するタウンホールでの議論を模擬して TBS を評価します。結果は、TBS が一貫した内部状態トレースを生成し、これらのトレースがターン割り当て、沈黙、メモリ条件全体にわたって体系的に変化することを示しています。不協和音関連の評価はエージェントの発言意欲を高めますが、沈黙の圧力評価はそれを低下させます。発言の意図が形成されると、公の場での表現は主に順番の割り当てルールによって形成されます。これらの発見は、TBS が内部評価から公的表現への経路を観察可能かつ分析可能にすることで、メカニズムに敏感な社会シミュレーションをサポートしていることを示唆しています。

原文 (English)

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existing dialogue simulation frameworks represent interaction mainly as observable turn exchange or aggregated outputs, leaving the internal evaluative processes behind silence, speaking intention, and public expression difficult to examine. We introduce TBS (Think-Before-Speak), an interval-based multi-agent simulation framework that separates agents' private reasoning from public utterance generation. At each interval, all agents update structured internal states based on the shared dialogue history and their own memory. These states include dissonance-related appraisal, perceived opinion climate, perceived isolation risk, response strategy, and willingness to speak. The orchestrator then resolves competing speaking intentions and commits one utterance to the public dialogue, allowing internal evaluation and public interaction to co-evolve over time. We evaluate TBS in simulated town hall discussions on a climate-related policy issue. Results show that TBS produces coherent internal-state traces and that these traces vary systematically across turn-allocation, silence, and memory conditions. Dissonance-related appraisal increases agents' willingness to speak, whereas silence-pressure appraisal decreases it. Once speaking intention is formed, public expression is shaped mainly by turn-allocation rules. These findings suggest that TBS supports mechanism-sensitive social simulation by making the pathway from internal evaluation to public expression observable and analyzable.

2026-07-02 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning

Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are uncha…

2026-07-02 13:00 JSTarXiv cs.AIビジネス/資金調達

Vibecoding Ate My 宿題: グリーンフィールド ソフトウェア エンジニアリングとプログラミングへの AI アプローチの評価

生成 AI の急速な発展のおかげで、私たちはコンピューターとの対話方法を永遠に変える可能性のあるパラダイム シフトの真っ只中にいます。この分野の基礎知識なしにアプリケーションやコーディング インフラストラクチャを構築するための自然言語プロンプトの使用が増加していることが観察されており、この実践は「バイブ コーディング」と呼ばれています。これはおそらく、プログラミングの分野が当初から、考えられるあらゆるより高い抽象化レベルで構築されてきたものを表しています。 Vibe コーディングは、入力方法に関する限り、高レベル プログラミングのメタのエンドポイントとなることが約束されています。つまり、人間によるコード構文の使用が完全に排除され、母国語でのプログラミングが優先されます。このペーパーは、グリーンフィールドのソフトウェア エンジニアリング タスクにおける Vibe コーディングの実現可能性を評価し、そのソフトウェア エンジニアリングの能力を測定するために使用されたベンチマークを分析することを目的としています。この目的を達成するために、私たちは、Python で単純で個別のグリーンフィールド プログラミング タスクを実行する LLM の習熟度を分析し、この問題に関する範囲を絞った洞察を提供するための評価スイートを開発しました。

原文 (English)

Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming

Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forever. We have observed a growth in the use of natural language prompts to build applications and coding infrastructures without underlying knowledge of the field, and this practice has been dubbed `vibe coding.' It arguably represents what the field of programming has been building towards since the beginning, with every higher level of abstraction that is conceived. Vibe coding promises to be the endpoint for the meta of high-level programming as far as method of input is concerned: eliminating a human's use of code syntax entirely in favour of programming in their mother tongue. This paper aims to evaluate the viability of vibe coding for greenfield software engineering tasks, as well as analyse the benchmarks that have been used to measure its software engineering prowess. To this end, we have developed an evaluation suite for analysing an LLM's proficiency in carrying out simple, isolated greenfield programming tasks in Python to provide scoped insight on the matter.

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成エージェントビジネス/資金調達研究/論文

HealthAgentBench: 挑戦的なフロンティア AI エージェント向けの現実的なエージェント ヘルスケア環境の統合ベンチマーク スイート

AI エージェントがますます複雑で長期的な推論を行えるようになっているため、現実世界の医療アプリケーションへの進歩を測定するには、厳密かつ総合的な評価が不可欠です。 HealthAgentBench は、それぞれ独自の環境を持つ 7 つのカテゴリにわたる 54 のエージェント ヘルスケア タスクのスイートです。ベンチマーク スイートは、患者の治療過程全体にわたる多様なワークフローと幅広いモダリティに及びます。各タスクは、エンドツーエンドの臨床ワークフローを複製するように設計されています。最小限の指示が与えられると、エージェントは生の医療データを探索し、複雑な環境内で操作し、単純なプロンプトを超えた複数ステップのソリューションを実行する必要があります。最終的なタスクの成功率は、各エージェントの HealthAgentBench の全体的なパフォーマンスに関する単一の解釈可能な指標を提供するために報告されます。 HealthAgentBench でフロンティア エージェントを評価すると、全体的なタスクの成功率が依然として低いことがわかり、スイートの難しさを浮き彫りにしています。最も強力で費用対効果の高いエージェントである Codex GPT-5.5 の成功率はわずか約 42% です。 HealthAgentBench は、総合的なパフォーマンスを超えて、タスク カテゴリ全体の微妙な長所と短所を明らかにします。フロンティア エージェントは、EHR データを使用した研究モデリング パイプラインの自動開発に期待を示していますが、医療画像処理は、特にクロード コード モデルの場合、依然として課題が多く、一方で Codex GPT-5.5 は新たな機能を示しています。大規模な検索スペースと構成的推論の要件を組み合わせるタスクは、現在のすべてのエージェントにとって依然として困難です。これらの結果を総合すると、HealthAgentBench が将来の進歩の余地が十分にある、挑戦的で現実的なベンチマークを提供していることを示唆しています。ベンチマークは https://github.com/microsoft/HealthAgentBench でリリースされています。

原文 (English)

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment. The benchmark suite spans diverse workflows throughout the patient journey and a broad range of modalities. Each task is designed to replicate an end-to-end clinical workflow: given minimal instructions, an agent must explore raw healthcare data, operate within a complex environment, and execute multi-step solutions that go beyond naive prompting. A final task success rate is reported to provide a single, interpretable metric for HealthAgentBench overall performance for each agent. Evaluating frontier agents on HealthAgentBench, we find that overall task success rate remains low, underscoring the difficulty of the suite. The strongest and the most cost effective agent, Codex GPT-5.5, achieves only approximately 42% success rate. Beyond aggregate performance, HealthAgentBench reveals nuanced strengths and weaknesses across task categories. Frontier agents show promise in automatically developing research modeling pipelines over EHR data, but medical imaging remains especially challenging, particularly for Claude Code models, while Codex GPT-5.5 shows emerging capability. Tasks that combine large search spaces with compositional reasoning requirements remain difficult for all current agents. Together, these results suggest that HealthAgentBench provides a challenging and realistic benchmark with substantial room for future progress. We release our benchmark at https://github.com/microsoft/HealthAgentBench.

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

RAISE: 堅牢な敵対的インスタンス検索を備えた LLM ベースの自動ヒューリスティック設計

大規模言語モデル (LLM) を使用した自動ヒューリスティック設計 (AHD) は、高品質のヒューリスティックの発見において目覚ましい進歩を示しています。ただし、既存の LLM ベースの AHD 手法は、固定されたトレーニング インスタンス セットのヒューリスティックを最適化するため、現実世界の分布シフトの下で展開すると壊滅的に失敗する可能性があります。我々は、トレーニング分布の原則的な近傍内での制約付きワーストケース インスタンス検索を LLM ベースの進化的検索ループに統合するフレームワークである、Robust Adversary Instance Search (RAISE) を提案します。 RAISE は、堅牢な AHD を制約付きの敵対的インスタンス検索問題として扱います。外側のループは LLM 演算子を介してヒューリスティックを進化させますが、LLM を使用しない内側のループは、境界射影を伴う基底分布パラメーター化を使用して、トレーニング インスタンス セットの周りのイプシロン ボール内のハード インスタンスを効率的に識別します。 5 つのディストリビューション ファミリにわたるオンライン ビン パッキング (OBP)、オンライン ジョブ ショップ スケジューリング (OJSP)、およびオンライン車両ルーティング (OVRP) に関する包括的な実験では、既存の LLM ベースの AHD 手法はディストリビューションの移行時に最大 19 倍劣化する一方、RAISE はテストされたすべてのディストリビューションと問題規模にわたって一貫して強力なパフォーマンスを維持することを実証しました。

原文 (English)

RAISE: LLM-based Automated Heuristic Design with Robust Adversary Instance Search

Automated Heuristic Design (AHD) with Large Language Models (LLMs) has shown remarkable progress in discovering high-quality heuristics. However, existing LLM-based AHD methods optimize heuristics for a fixed training instance set and may fail catastrophically when deployed under real-world distributional shifts. We propose Robust Adversary Instance Search (RAISE), a framework that integrates constrained worst-case instance search within a principled neighborhood of the training distribution into the LLM-based evolutionary search loop. RAISE treats robust AHD as a constrained adversarial instance search problem: the outer loop evolves heuristics via LLM operators, while an LLM-free inner loop efficiently identifies hard instances within an epsilon-ball around the training instance set using a basis distribution parameterization with boundary projection. Comprehensive experiments on Online Bin Packing (OBP), Online Job Shop Scheduling (OJSP), and Online Vehicle Routing (OVRP) across five distribution families demonstrate that existing LLM-based AHD methods degrade by up to 19 times under distribution shift, while RAISE consistently maintains strong performance across all tested distributions and problem scales

2026-07-01 13:00 JSTarXiv cs.AIビジネス/資金調達

ベイジアンマテリアルデザインのためのサロゲートゲート生成と基礎モデル埋め込み

閉ループ材料の発見では、候補構造の提案とその特性の評価が繰り返され、特性評価がコストの大半を占めます。生成バリアントでは、学習された事前分布が候補結晶を提案し、プロパティ オラクルがそれらをスコア付けします。私たちは、安価な確率的サロゲートがジェネレータの出力をトリアージできるかどうか、そしてそのようなサロゲートがうまく機能しなければならないことを尋ねます。アーキテクチャ的に異なる 3 つの事前トレーニング済み拡散事前学習 (MatterGen、CrystalFlow、ADiT) と 2 つのターゲット (室温の熱容量と体積弾性率) にわたって、RL 主導の生成ワークフローで構造生成とオラクルの間にガウス プロセス取得ゲートを挿入します。このゲートは、サイクルごとの固定バジェットでオラクル呼び出しを制限しながら、生成モデルの非ゲート微調整と同等またはそれを上回っています。予算に見合ったアブレーションにより、このメカニズムを分離します。同一の 4 コール バジェットでは、ランキングに基づいた選択が任意の選択よりも優れており、サロゲートの選択によって利益が得られることが確認されています。ゲートは、コールの約 5 分の 1 で、オラクルの総支出額の $\sim$9\% 以内に収まります。体積弾性率の発見の密度汎関数理論チェックにより、学習されたオラクルが平均 2.5\% 以内であることと、生成された構造のサロゲートのランキングが Spearman $\rho = 0.94$ であることが確認されました。機械的、電子的、振動的特性にわたる代替パフォーマンスの複数因子ベンチマークにより、事前トレーニング済みの ORB 埋め込みとガウス プロセスが最も信頼できる組み合わせであることが特定され、これを提案されたワークフローの構成要素として採用します。完全なパイプラインはオープンソース ソフトウェアとしてリリースされます。

原文 (English)

Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design

Closed-loop materials discovery iterates between proposing candidate structures and evaluating their properties, and property evaluation dominates the cost. In the generative variant, a learned prior proposes candidate crystals and a property oracle scores them; we ask whether a cheap probabilistic surrogate can triage the generator's output, and what such a surrogate must do well. Across three architecturally distinct pretrained diffusion priors (MatterGen, CrystalFlow, ADiT) and two targets (room-temperature heat capacity and bulk modulus), we insert a Gaussian process acquisition gate between structure generation and the oracle in an RL-steered generative workflow. The gate matches or exceeds ungated fine-tuning of the generative model while capping oracle calls at a fixed per-cycle budget. Budget-matched ablations isolate the mechanism. At an identical four-call budget, ranking-based selection outperforms arbitrary selection, confirming that the gain comes from the surrogate's choice; the gate comes within $\sim$9\% of exhaustive oracle spending at roughly one-fifth of the calls. A density-functional-theory check of the bulk-modulus discoveries confirms the learned oracle to within 2.5\% on average and the surrogate's ranking of the generated structures at Spearman $\rho = 0.94$. A cross-factorial benchmark of surrogate performance spanning mechanical, electronic, and vibrational properties identifies pretrained ORB embeddings with a Gaussian process as the most reliable combination, which we adopt as the building blocks of the proposed workflow. The complete pipeline is released as open-source software.

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actiona…

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達規制/政策

Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law

Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards foc…

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian

Relation extraction (RE) for low-resource languages is typically constrained by the lack of annotated corpora. We investigate the feasibili…

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

STEB: Style Text Embedding Benchmark

While semantic embeddings are rigorously evaluated on the Massive Text Embedding Benchmark, the evaluation of style embeddings remains frag…

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM における推論の質の測定: 多次元の行動フレームワーク

LLM は複雑な推論タスクで目覚ましい成功を収めていますが、現在の評価アプローチは主に最終的な答えの正しさに依存しており、それらの答えを生み出す根本的な推論プロセスについての洞察は限られています。このギャップに対処するために、この研究では、動作の観点から LLM の推論品質を測定するための統一された多次元フレームワークを提案し、理論的に根拠のある 6 つの次元、正確性 (CQ)、一貫性 (CS)、堅牢性 (RS)、論理的一貫性 (LS)、効率 (ES)、安定性 (SS) を運用します。 4 つのベンチマークの 975 項目にわたる 7 つの LLM に関する広範な実験により、このフレームワークが精度のみの指標では見えない動作を明らかにすることが実証されました。特に、論理的一貫性は正しさ (r = -0.172、ns) と直交しており、一貫性のない推論から正しい答えが得られることが確認され、一方、Claude-Haiku-4.5 は最高の多次元スコア (Q_bal = 0.778) を達成しています。さらに、このフレームワークは重大なランキングの逆転を明らかにしています。DeepSeek-V3 は精度優先では 2 位ですが、法的/コンプライアンスの重み付けでは 5 位にランクされており、単一指標の評価では検出できない逆転です。判別式の妥当性により、11/15 次元のペアが独立している (|r| < 0.50) ことが確認され、各次元を別個の信号として扱うための心理測定的サポートが提供されます。フレームワークによって生成される次元プロファイルは、次の 3 つのクラスの展開決定を直接サポートします。最終的な答えが正しいにもかかわらず、その推論トレースが説明責任監査に失敗するモデルを特定します (LS--CQ 直交性)。精度のみのベンチマークによって引き起こされるランキングエラーを防止します。そして、フレームワークがキャプチャする 6 つの独立したシグナルを単一のメトリックが暗黙的に置き換えることがないようにします。

原文 (English)

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

LLMs have achieved remarkable success in complex reasoning tasks, yet current evaluation approaches predominantly rely on final-answer correctness, offering limited insight into the underlying reasoning processes that produce those answers. To address this gap, this study proposes a unified multi-dimensional framework for measuring reasoning quality in LLMs from a behavioral perspective, operationalizing six theoretically grounded dimensions: Correctness (CQ), Consistency (CS), Robustness (RS), Logical Coherence (LS), Efficiency (ES), and Stability (SS). Extensive experiments on seven LLMs across 975 items from four benchmarks demonstrate that the framework reveals behaviors invisible to accuracy-only metrics. Notably, logical coherence is orthogonal to correctness (r = -0.172, ns), confirming that correct answers can arise from incoherent reasoning, while Claude-Haiku-4.5 achieves the highest multi-dimensional score (Q_bal = 0.778). Furthermore, the framework exposes critical ranking inversions: DeepSeek-V3 ranks second under accuracy-priority but fifth under legal/compliance weighting, a reversal that single-metric evaluation cannot detect. Discriminant validity confirms 11/15 dimension pairs are independent (|r| < 0.50), providing psychometric support for treating each dimension as a distinct signal. The dimensional profiles produced by the framework directly support three classes of deployment decision: identifying models whose reasoning traces would fail accountability audits despite correct final answers (LS--CQ orthogonality); preventing ranking errors caused by accuracy-only benchmarking; and ensuring that no single metric silently substitutes for the six independent signals the framework captures.

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ギャンブルはしないでください、GAMBLe: AI 主導の研究システムのための分析フレームワーク

AI-Driven Research Systems (ADRS) -- LLM と自動評価を組み合わせてアルゴリズム、証明、設計を発見するシステム -- は最適化され、ドメイン全体で採用されていますが、それらを分析するツールは追いついていません。 ADRS のパフォーマンスはコンポーネントの相互作用に依存しますが、これらの相互作用は十分に理解されておらず、調査にコストがかかり、(ここで示しているように) 標準の収束保証では十分に把握されていません。これらの保証は、私たちが形式化した ADRS プロセスの下では成立しない構造的な仮定に依存しています。我々は、ADRS の動作を 4 つのパラメーター (ジェネレーター $G$、アセッサー $\mathcal{A}$、発見メカニズム $\mathcal{M}$、バジェット $B$) と 1 つの構成オブジェクト、効果的なランドスケープ $L_{\text{eff}} = \mathcal{A} \circ G$ に分解するフレームワークである GAMBLe を紹介します。これにより、異なるジェネレーターとアセッサーのペアが構造的に異なる問題ごとの最適化を引き起こすことが明らかになります。風景。私たちは、単一の LLM から動的適応アンサンブルに至るジェネレーター、貪欲な選択から共進化メタサーチに至るメカニズム、および評価者が連続スコアリングからクリフ関数に及ぶ 3 つの NP 困難問題に及ぶ 760 以上の反復実行 (>46,000 反復) でフレームワークを実行します。実験では、ジェネレーターやメカニズムの完全な順序付けは明らかにされていません。フロンティア モデルはオープンソースの代替モデルよりもパフォーマンスが劣る可能性があり、最も単純なメカニズムが最先端のメタ検索を上回る場合もあります。結果は、限られた予算 (実行ごとに 60 回の反復) の下でも、適切なコンポーネントを選択することでパフォーマンスを 13 ~ 67%、検索効率を 6 ~ 39 倍改善できることを示しています。

原文 (English)

Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems

AI-Driven Research Systems (ADRS) -- systems coupling LLMs with automated evaluation to discover algorithms, proofs, and designs -- are being optimized and adopted across domains, but the tools to analyze them have not kept pace. ADRS performance depends on component interactions that are poorly understood, expensive to explore, and (as we show) not well captured by standard convergence guarantees. These guarantees rely on structural assumptions that do not hold under the ADRS process we formalize. We introduce GAMBLe, a framework that decomposes ADRS behavior into four parameters (generator $G$, assessor $\mathcal{A}$, discovery mechanism $\mathcal{M}$, budget $B$) and one compositional object, the effective landscape $L_{\text{eff}} = \mathcal{A} \circ G$, which reveals that distinct generator-assessor pairs induce structurally different per-problem optimization landscapes. We exercise the framework on 760+ replicated runs (>46,000 iterations) spanning generators from single LLMs to dynamically-adaptive ensembles, mechanisms from greedy selection to co-evolutionary meta-search, and three NP-hard problems whose assessors range from continuous scoring to cliff functions. The experiments reveal no total ordering of generators or mechanisms: frontier models can underperform open-source alternatives and the simplest mechanism sometimes outperforms state-of-the-art meta-search. Results show that even under limited budgets (60 iterations per run), the right component choices can improve performance by 13-67% and search efficiency by 6-39x.

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO

Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier langu…

2026-07-01 13:00 JSTarXiv cs.AIビジネス/資金調達

ナレッジ グラフにおけるグラフ間の意味的類似性の測定: ナレッジ グラフ埋め込みの経験的評価

ナレッジ グラフ (KG) は事実を構造化されたトリプルとして表し、さまざまなドメインにわたる関係知識を整理するために広く使用されています。テキスト情報の範囲が単語や文章から完全な文書に及ぶのと同様に、KG 情報は、エンティティ、関係、トリプルからサブグラフや KG 全体に至るまで、複数のレベルで解釈できます。ただし、既存の KG 埋め込み手法は主にエンティティ、リレーション、トリプルに焦点を当てており、グラフレベルのセマンティクスにはほとんど対処されていません。通常、構造パターンに基づいてグラフを比較する従来のグラフレベルの方法も、構造の類似性だけでは KG 間の意味的な類似性を保証できないため、不十分です。さまざまな方法がそのようなグラフレベルの意味論的情報をどの程度うまく捕捉しているかを評価するために、KG のペアが意味論的に対応する基礎的な情報を表すかどうかを決定する、グラフ間の意味論的類似性を研究します。信頼できるグラウンドトゥルース対応を取得するために、テキスト文書を変更し、元の文書と変更された文書の両方から KG を抽出し、それらの既知の対応関係を KG ペアに転送することにより、意味論的一致データセットを構築します。各データセットについて、テキストベース、構造ベース、および KG 埋め込みベースのアプローチを比較します。 KG 埋め込みベースのアプローチでは、ペアごとのエンティティの最大類似性を使用する \textit{EmbPairSim} と、周波数加重セントロイドを使用する \textit{AvgEmbSim} の 2 つのスコアリング関数を導入します。 WikiText-2 と CC-News での実験では、\textit{EmbPairSim} が大幅に少ないパラメーターを使用しながら、Sentence-BERT よりも最大 5.3 pp 高い MRR を達成することが示されています。これらの結果は、KGE 表現が、KG におけるグラフ間の意味論的類似性に対するコンパクトで効果的なシグナルとして機能できることを示唆しています。私たちのコードは https://github.com/SeungRyeolBaek/KG-to-KG-Semantic-Similarity で入手できます。

原文 (English)

Measuring Graph-to-Graph Semantic Similarity in Knowledge Graphs: An Empirical Evaluation of Knowledge Graph Embeddings

A Knowledge Graph (KG) represents facts as structured triples and is widely used to organize relational knowledge across diverse domains. Just as textual information ranges from words and sentences to complete documents, KG information can be interpreted at multiple levels, from entities, relations, and triples to subgraphs and entire KGs. However, existing KG embedding methods mainly focus on entities, relations, and triples, leaving graph-level semantics largely unaddressed. Conventional graph-level methods, which typically compare graphs based on structural patterns, are also insufficient because structural similarity alone cannot guarantee semantic similarity between KGs. To evaluate how well different methods capture such graph-level semantic information, we study graph-to-graph semantic similarity, which determines whether a pair of KGs represents semantically corresponding underlying information. To obtain reliable ground-truth correspondences, we construct a semantic matching dataset by modifying text documents, extracting KGs from both original and modified documents, and transferring their known correspondences to KG pairs. We compare text-based, structure-based, and KG embedding-based approaches on each dataset. For the KG embedding-based approach, we introduce two scoring functions: \textit{EmbPairSim}, which uses maximal pairwise entity similarity, and \textit{AvgEmbSim}, which uses a frequency-weighted centroid. Experiments on WikiText-2 and CC-News show that \textit{EmbPairSim} achieves up to 5.3 pp higher MRR than Sentence-BERT while using substantially fewer parameters. These results suggest that KGE representations can serve as compact and effective signals for graph-to-graph semantic similarity in KGs. Our code is available at https://github.com/SeungRyeolBaek/KG-to-KG-Semantic-Similarity.

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA

As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evalu…

2026-07-01 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

VGGSounder: 基礎モデルのオーディオビジュアル評価

視聴覚基礎モデルの出現は、マルチモーダルな理解を確実に評価することの重要性を強調しています。 VGGSound データセットは、オーディオビジュアル分類の評価のベンチマークとしてよく使用されます。ただし、私たちの分析では、不完全なラベル付け、部分的に重複するクラス、不整合なモダリティなど、VGGSound のいくつかの制限が特定されました。これらは、聴覚および視覚能力の歪んだ評価につながります。これらの制限に対処するために、VGGSounder を導入します。これは、VGGSound を拡張し、オーディオビジュアル基礎モデルを評価するために特別に設計された、包括的に再アノテーションが付けられたマルチラベル テスト セットです。 VGGSounder は詳細なモダリティの注釈を備えており、モダリティ固有のパフォーマンスを正確に分析できます。さらに、新しいモダリティ混乱メトリックを使用して別の入力モダリティを追加したときのパフォーマンスの低下を分析することで、モデルの限界を明らかにします。

原文 (English)

VGGSounder: Audio-Visual Evaluations for Foundation Models

The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.

2026-07-01 13:00 JSTarXiv cs.AIビジネス/資金調達

SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs

Machine Learning (ML) pipelines encode quality-relevant decisions across data preparation, training, evaluation, and configuration code. So…

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) system…

2026-07-01 11:04 JSTTechCrunch AIビジネス/資金調達

Wayve launches $85M employee tender offer at $8.5B valuation

Wayve’s offering is part of a growing trend of AI startups using employee tenders as a strategic tool to attract and retain talent.

2026-07-01 03:13 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

Nvidia competitor Etched hits $5B valuation, $1B in sales for AI chip

Nvidia AI chip competitor Etched says it has already booked $1 billion under contract for the inference systems powered by its chip.

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

モデル機能強化のためのデータと評価のクローズドループ

モデルの能力は LLM の事前トレーニングの中心的な変数ですが、直接観察されることはありません。データは前向きにモデルの能力を形成しますが、評価では遡及的にのみ明らかになり、サンプル、プロンプト、デコード、およびスコアリングのルールが 1 つのノイズの多いスコアに圧縮されます。実際の最適化ではこれを逆方向に実行します。つまり、最初に障害が観察され、エンジニアはコーパスの修正を推測する必要があります。両者は互換性のない用語 (ベンチマーク名、サンプルごとの正確性とデータ ソース、ドメイン、品質ラベルなど) を話し合っているため、この推論は通常、方法ではなく直感になります。このギャップは \emph{能力スライス} で埋めます。バックグラウンド条件、タスク タイプ、解決操作、出力制約を共有する評価サンプルのグループです。粗すぎるベンチマーク名や、ノイズが多すぎる単一サンプルとは異なり、単一の弱点を特定するのに十分な精度を持ちながら、集計に耐えられるほど安定しています。このユニットを中心に構築され、評価分類法、非命令データ分類法、およびマッピング ルールが閉ループを形成し、ベンチマーク レベルの障害を対象を絞ったテスト可能なデータ介入に変えます。このループを、反対方向に引っ張る 2 つのケーススタディでテストします。まず、ループはデータを除外します。事前トレーニングを継続すると、BBH が $-46.82\%$ 減少しますが、診断では、これは推論の弱体化ではなく、マスクされた単一の \texttt{\textless EOS\textgreater} 損失であることがわかります。これを復元すると、データを変更せずに、BBH を元のチェックポイントを上回る $66.44$ に回復します。第 2 に、ループはデータをルール化します。永続的な数学推論の弱点は、演算を解くことで特定の失敗の組み合わせに分解され、そこから構築された弱点をターゲットにしたサンプリング手順により、AIME2025/AIME2026 Pass@128 がそれぞれ $6.67$/$0.00$ から $26.67$ に引き上げられます。同じ未変更のループは、どちらの場合も正反対の正しい判定に達し、データに対する評価の推論が直感的ではなく日常的で監査可能であり、実験的に検証される可能性があることを示しています。

原文 (English)

Data and Evaluation Closed-Loop for Model Capability Enhancement

Model capability is the central variable in LLM pre-training, yet is never observed directly: data shapes it prospectively, while evaluation reveals it only retrospectively, compressing samples, prompts, decoding, and scoring rules into one noisy score. Practical optimization runs this backward: a failure is observed first, and the engineer must infer the corpus fix. The two sides speak incompatible vocabularies -- benchmark names and per-sample correctness versus data sources, domains, and quality labels -- so this inference is usually intuition, not method. We close this gap with the \emph{capability slice}: a group of evaluation samples sharing background condition, task type, solving operation, and output constraint -- precise enough to localize a single weakness yet stable enough to survive aggregation, unlike a benchmark name, too coarse, or a single sample, too noisy. Built around this unit, an evaluation taxonomy, a non-instruction data taxonomy, and mapping rules form a closed loop turning a benchmark-level failure into a targeted, testable data intervention. We test this loop on two case studies pulling in opposite directions. First, the loop rules the data out: continued pre-training drives BBH down by $-46.82\%$, but diagnosis traces this to a single masked \texttt{\textless EOS\textgreater} loss rather than weakened reasoning; restoring it recovers BBH to $66.44$, above the original checkpoint, without changing the data. Second, the loop rules the data in: a persistent math-reasoning weakness is decomposed by solving operation into specific failing combinations, and a weakness-targeted sampling procedure built from it lifts AIME2025/AIME2026 Pass@128 from $6.67$/$0.00$ to $26.67$ each. The same unmodified loop reaches opposite, correct verdicts in both cases, showing the evaluation-to-data inference can be routine, auditable, and experimentally validated rather than intuitive.

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達

実際のポイントオブケアの臨床クエリに対する臨床 AI ツールの専門家による評価

現在、医師は毎週何百万もの臨床質問を AI ツールに投げかけていますが、これらのツールは主に、実際に尋ねられる質問ではなく、仮説や試験形式の質問に基づいて評価されています。我々は、30の専門分野にわたる医師によってOpenEvidence(OE)プラットフォームに送信された620のリアルワールドポイントオブケアクエリ(Real-POCQi)と、HealthBenchからの187の質問に基づいて構築された盲検評価を報告します。 36 州の 149 名の現役医師が、各質問の専門分野に合わせた採点者を使用して、3 つのフロンティア汎用モデル (Claude Opus 4.8、Gemini 3.1 Pro、GPT-5.5) と特殊な臨床ツール (OE) の回答を直接比較しました。臨床意思決定支援に関連する 5 つの側面 (精度、臨床的有用性、情報源の品質、検証可能性、完全性) に沿って回答を比較した場合、医師は専門ツールをすべての軸で最も高いスコアにしました。 Real-POCQi に関する一次分析では、勝率の差 (勝率と敗率の差) は 25 ~ 39 パーセント ポイントの範囲でした (p<0.001)。結果は、引用表示、回答の長さ、OE ユーザーのステータス、Real-POCQi と HealthBench によって層別化した感度分析で一貫性を保っていました。並行して、LLM 裁判官は体系的に専門家裁判官とは異なることが判明しましたが、最良のモデルについては両者とも概ね一致しました。これらの調査結果は、次の 2 つの結論を強調しています。(i) AI ツールの評価は、現実世界のクエリ分布を反映し、現代医学を定義する専門分野を反映する専門家判断者を使用する必要があります。(ii) 汎用モデルに対する専用ツールの一貫した利点は、後者が同様の目的を達成できないことを必ずしも意味するわけではありませんが、ターゲットを絞ったエンジニアリングとカスタマイズにより、ユーザーにとってパフォーマンスに有意義な向上がもたらされる可能性があります。私たちは Real-POCQi を公開ベンチマークとしてリリースするとともに、この研究の結果を再現するための事前に指定された統計分析もリリースします。

原文 (English)

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions, not those actually asked in practice. We report a blinded evaluation built on 620 Real-world Point-Of-Care Queries (Real-POCQi) submitted to the OpenEvidence (OE) platform by physicians spanning 30 specialties, as well as 187 questions from HealthBench. 149 practicing physicians across 36 states made head-to-head comparisons between answers from three frontier general-purpose models (Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5) and a specialized clinical tool (OE), with graders matched to each question's specialty. When comparing answers along five dimensions relevant to clinical decision support -- accuracy, clinical utility, source quality, verifiability, & completeness -- physicians scored the specialized tool highest on all axes; in the primary analysis on Real-POCQi, win differences (margins between win and loss rates) ranged from 25 to 39 percentage points (p<0.001). Results remained consistent in sensitivity analyses stratifying by citation display, answer length, OE-user status, and Real-POCQi versus HealthBench. In parallel, LLM judges were found to systematically differ from expert judges, though both generally agreed on the best model. These findings underscore two conclusions: (i) AI tool evaluations should reflect real-world query distributions and use expert judges that mirror the specialization defining modern medicine and (ii) the consistent advantage of the specialized tool over general-purpose models does not necessarily mean that the latter cannot serve similar purposes, but that targeted engineering and customization can yield meaningful gains in performance for its users. We release Real-POCQi as a public benchmark, as well as the prespecified statistical analysis for reproducing results of this study.

2026-06-30 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

人間のフォールバックの管理: AI と労働者の流動性の向上におけるスキル投資

企業が自律型 AI を導入する場合、どの程度の作業をシステムに任せるか、どの程度の作業を従業員の関与を維持するかを決定する必要があります。この決定は、現在の生産高と将来の人的資本に影響を与えます。私たちは、AI が機能するときにはワーカーよりも優れたパフォーマンスを発揮する可能性があるが、確実な確率で失敗する可能性があるという、節約的な 2 期間モデルを開発します。企業は労働者の関与を選択します。エンゲージメントは、基準を下回る従業員の現在の生産量を低下させますが、学習と消耗によって将来のスキルを変化させます。私たちは AI の進歩の 2 つの側面を区別します。それは、機能 (システムが動作するときのシステムの出力) と信頼性 (システムが動作する確率) です。単一企業のベンチマークでは、エンゲージメントはフォールバック投資としてのみ価値があります。同社は最もスキルの低い従業員を最も雇用しています。これは、従業員のスキルギャップが最も大きく、有用な代替レベルに引き上げるためのコストが最も低いためです。労働者の流動性により、エンゲージメントは労働市場の選別にも影響を及ぼします。労働者は、より価値のあるスキルの軌跡を構築する仕事を好みます。この選別の動機は、スキルの向上がより価値があり、エンゲージメントのコストが低い、AI フロンティアに近い高スキルの労働者をターゲットにしています。したがって、モビリティはエンゲージメント パターンを逆転させ、AI ベンチマークを下回る最もスキルの低い労働者から最もスキルの高い労働者に投資をシフトする可能性があります。モビリティはまた、AI の進歩がエンゲージメントに与える影響を再形成します。能力が向上すると、企業が提供するスキル軌跡の価値が高まるためエンゲージメントが高まりますが、信頼性が高まるとフォールバックの必要性が減り、同時に学習機会も変化するため、エンゲージメントが上下する可能性があります。労働者の流動性の下では、人間と AI の仕事の設計は人的資本の投資の問題となり、今日の仕事の割り当てが将来のスキルを形成します。

原文 (English)

Managing the Human Fallback: Skill Investment Under Improving AI and Worker Mobility

When firms deploy autonomous AI, they must decide how much work to leave to the system and how much to keep workers engaged. This decision affects current output and future human capital. We develop a parsimonious two-period model in which AI may outperform the worker when it functions, but may fail with positive probability. A firm chooses worker engagement; engagement lowers current output for below-benchmark workers, but changes future skill through learning and erosion. We distinguish two dimensions of AI progress: capability, the system's output when it works, and reliability, the probability that it works. In a single-firm benchmark, engagement is valuable only as fallback investment. The firm engages the least-skilled workers most, because they have the largest skill gaps and are least costly to bring toward a useful fallback level. With worker mobility, engagement also affects labor-market sorting: workers prefer jobs that build more valuable skill trajectories. This sorting motive targets higher-skill workers near the AI frontier, where skill gains are more valuable and engagement is less costly. Mobility can therefore reverse the engagement pattern, shifting investment from the least-skilled toward the most-skilled workers below the AI benchmark. Mobility also reshapes how AI progress affects engagement: greater capability raises engagement by increasing the value of the skill trajectory a firm offers, whereas greater reliability can raise or lower it because it reduces fallback need while also changing learning opportunities. Under worker mobility, human-AI work design becomes a problem of human-capital investment, in which allocating work today shapes future skill.

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達

ナレッジ グラフにおけるグラフ間の意味的類似性の測定: ナレッジ グラフ埋め込みの経験的評価

ナレッジ グラフ (KG) は事実を構造化されたトリプルとして表し、さまざまなドメインにわたる関係知識を整理するために広く使用されています。テキスト情報の範囲が単語や文章から完全な文書に及ぶのと同様に、KG 情報は、エンティティ、関係、トリプルからサブグラフや KG 全体に至るまで、複数のレベルで解釈できます。ただし、既存の KG 埋め込み手法は主にエンティティ、リレーション、トリプルに焦点を当てており、グラフレベルのセマンティクスにはほとんど対処されていません。通常、構造パターンに基づいてグラフを比較する従来のグラフレベルの方法も、構造の類似性だけでは KG 間の意味的な類似性を保証できないため、不十分です。さまざまな方法がそのようなグラフレベルの意味論的情報をどの程度うまく捕捉しているかを評価するために、KG のペアが意味論的に対応する基礎的な情報を表すかどうかを決定する、グラフ間の意味論的類似性を研究します。信頼できるグラウンドトゥルース対応を取得するために、テキスト文書を変更し、元の文書と変更された文書の両方から KG を抽出し、それらの既知の対応関係を KG ペアに転送することにより、意味論的一致データセットを構築します。各データセットについて、テキストベース、構造ベース、および KG 埋め込みベースのアプローチを比較します。 KG 埋め込みベースのアプローチでは、ペアごとのエンティティの最大類似性を使用する \textit{EmbPairSim} と、周波数加重セントロイドを使用する \textit{AvgEmbSim} の 2 つのスコアリング関数を導入します。 WikiText-2 と CC-News での実験では、\textit{EmbPairSim} が大幅に少ないパラメーターを使用しながら、Sentence-BERT よりも最大 5.3 pp 高い MRR を達成することが示されています。これらの結果は、KGE 表現が、KG におけるグラフ間の意味論的類似性に対するコンパクトで効果的なシグナルとして機能できることを示唆しています。私たちのコードは https://github.com/SeungRyeolBaek/KG-to-KG-Semantic-Similarity で入手できます。

原文 (English)

Measuring Graph-to-Graph Semantic Similarity in Knowledge Graphs: An Empirical Evaluation of Knowledge Graph Embeddings

A Knowledge Graph (KG) represents facts as structured triples and is widely used to organize relational knowledge across diverse domains. Just as textual information ranges from words and sentences to complete documents, KG information can be interpreted at multiple levels, from entities, relations, and triples to subgraphs and entire KGs. However, existing KG embedding methods mainly focus on entities, relations, and triples, leaving graph-level semantics largely unaddressed. Conventional graph-level methods, which typically compare graphs based on structural patterns, are also insufficient because structural similarity alone cannot guarantee semantic similarity between KGs. To evaluate how well different methods capture such graph-level semantic information, we study graph-to-graph semantic similarity, which determines whether a pair of KGs represents semantically corresponding underlying information. To obtain reliable ground-truth correspondences, we construct a semantic matching dataset by modifying text documents, extracting KGs from both original and modified documents, and transferring their known correspondences to KG pairs. We compare text-based, structure-based, and KG embedding-based approaches on each dataset. For the KG embedding-based approach, we introduce two scoring functions: \textit{EmbPairSim}, which uses maximal pairwise entity similarity, and \textit{AvgEmbSim}, which uses a frequency-weighted centroid. Experiments on WikiText-2 and CC-News show that \textit{EmbPairSim} achieves up to 5.3 pp higher MRR than Sentence-BERT while using substantially fewer parameters. These results suggest that KGE representations can serve as compact and effective signals for graph-to-graph semantic similarity in KGs. Our code is available at https://github.com/SeungRyeolBaek/KG-to-KG-Semantic-Similarity.

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

複雑さの上限ベンチマーク: 深さスケーリングの下で​​の逐次推論のマルチドメイン評価

必要な連続ステップの数が増加するにつれて、言語モデルの推論がどのように減衰するかを制御された評価である Complexity Ceiling Benchmark (CCB) を導入します。 CCB は、タスクの意味論的な内容を固定し、構造的に異なる 3 つの領域 (接地された空間状態追跡、抽象的な記号ポインター操作、および推移的な関係推論) にわたって、{5,...,50} の深さ N だけを変更します。 5 つのフロンティアおよびオープンウェイト LLM にわたる 6,000 回の試行にわたって、広く分離されたドメイン上限を持つ幾何学的なステップごとの減衰の一貫したパターンが見つかりました。最初の 2 つの領域では、最も強いモデルは N=50 全体で pd > 0.92 を維持しました。 3 番目では、すべてのモデルが N=5 によって崩壊し、pd=0.863 にもかかわらず、最良のモデルの 50% 成功の範囲は H0.5 ~ 4.7 ステップになります。トレース レベル メトリクス (TFBC) は、ベンチマーク全体の正解の 14.5% が、誤った中間推論を介して到達していることを示しています。強制的な冗長状態追跡では上限は変動せず (McNemar p=1.000)、推論が最初に分岐する平均ステップ k* は、パラメーター数よりもドメイン内の精度を予測します。 CCB と幾何学的減衰モデルを併用すると、モデルの長期推論プロファイルがタスク ファミリごとに 1 つの解釈可能な数値に縮小されます。

原文 (English)

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling

We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential steps grows. CCB fixes the semantic content of a task and varies only its depth N in {5,...,50} across three structurally distinct regimes: grounded spatial state-tracking, abstract symbolic pointer manipulation, and transitive relational inference. Across 6,000 trials over five frontier and open-weight LLMs we find a consistent pattern of geometric per-step decay with widely separated domain ceilings: on the first two regimes the strongest models retain pd>0.92 across N=50; on the third every model collapses by N=5, with the best model's 50%-success horizon at H0.5~4.7 steps despite pd=0.863. A trace-level metric (TFBC) shows that 14.5% of correct answers across the benchmark are reached via incorrect intermediate reasoning. Forced verbose state-tracking does not move the ceiling (McNemar p=1.000), and the mean step at which reasoning first diverges, k*, predicts within-domain accuracy better than parameter count. CCB and the geometric decay model together reduce a model's long-horizon reasoning profile to one interpretable number per task family.

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

正式なベンチマークの欠陥: データセットの欠陥とリーン定理証明における評価の失敗

リーンでの LLM 支援定理証明のベンチマークは、解決されたすべてのインスタンスに機械チェックされた証明が付属しているため、多くの場合、本質的に信頼できるものとして扱われます。ただし、カーネルは証明が \emph{formal} ステートメントを確立することをチェックするだけです。ステートメントが意図した非公式な問題を忠実にエンコードしているかどうか、評価ハーネスが自明な解決策や敵対的な解決策に対して堅牢であるかどうかは検証されません。私たちは、広く使用されている 5 つのリーン定理証明ベンチマークとそのフォークを監査し、コーパス スケールの静的チェッカーを使用して、反例、空定理、不健全な公理などの機械的に証明された 398 の問題を含む 4,833 の結果を明らかにしました。また、仮説の欠落、問題の単純化、不完全または不正確な翻訳、リーン固有の仕様上の危険などの意味論的な欠陥も文書化します。データセットの構築を超えて、評価時の故障モードを調査し、修正されたサブセット上で、欠陥によって報告される証明者スコアが膨らむことも小さくなる可能性があることを示します。私たちは、障害分類法、自動チェッカーのスイート、およびリコール指向のセマンティック監査プロンプトを提案し、正式な数学データセットの作成をガイドし、評価をより再現可能で信頼できるものにするための標準をリリースします。チェッカー、監査プロンプト、および修正されたデータセットのスナップショットは、https://github.com/Shachi456/atp-checkers で入手できます。

原文 (English)

Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving

Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-checked proof. However, the kernel only checks that a proof establishes a \emph{formal} statement; it does not verify that the statement faithfully encodes the intended informal problem, nor that evaluation harnesses are robust to trivial or adversarial solutions. We audit five widely used Lean theorem-proving benchmarks and their forks, using corpus-scale static checkers to surface 4,833 findings, including 398 mechanically certified issues such as counterexamples, vacuous theorems, and unsound axioms. We also document semantic defects such as missing hypotheses, problem simplification, incomplete or incorrect translations, and Lean-specific specification hazards. Beyond dataset construction, we survey evaluation-time failure modes and show, on corrected subsets, that defects can both inflate and deflate reported prover scores. We propose a fault taxonomy, a suite of automated checkers and recall-oriented semantic audit prompts, and release standards to guide the creation of formal math datasets and to make evaluation more reproducible and trustworthy. Our checkers, audit prompts, and corrected dataset snapshots are available at https://github.com/Shashi456/atp-checkers.

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達

プロセスレベルの社会的影響評価のための認知世界モデル

社会的影響ダイアログは、内部の認知状態を変えることでユーザーの行動を変えます。評価の中心となる質問は、ユーザーの信念、欲望、意図、感情が会話の過程で測定可能なほど変化するかどうかであり、これは表面レベルのテキスト指標 (BLEU/ROUGE) や単一スコアの LLM 判定では捉えることができないプロセス指向の基準です。我々は \textbf{Cog}nitive \textbf{W}orld \textbf{M}odel \textbf{(CogWM)} を提案します。これは、マルチターン対話評価を「ユーザーが何を言ったか」から「ユーザーの内部認知状態がどのように進化したか」に再構成する LLM ベースのユーザー モデルです。CogWM は、BDI/E 認知状態とユーザー発話を共同で予測し、3 層を使用してユーザー シミュレーターと評価プラットフォームの両方として機能します。ターンレベルの忠実度、軌道レベルの状態ダイナミクス、タスクレベルの複合スコアリングをカバーする評価フレームワーク。 4 つの社会的影響シナリオにわたる 150,454 のユーザー ターン サンプルで \textbf{S}ummarize-\textbf{a}nd-\textbf{A}llocate \textbf{(SaA)} アノテーション パイプラインを介してトレーニングされた CogWM は、77.6\% の感情精度 (GPT-5.5 の 2.1$\time$) を達成しました。 3,600 件のマルチエージェント識別試験において、認知的影響力によって 6 つの営利エージェントを区別し、Llama-4-Scout が 1 位にランクされました (CTS +0.233)。 CogWM は、社会的影響対話の評価を最終的な判断からプロセスの追跡に移行します。コード\脚注{\scriptsize コード: https://github.com/lucianma05-create/CogWM} とモデル\脚注{モデル: https://www.modelscope.cn/models/LucianMa/CogWM-14B} をリリースしました。

原文 (English)

Cognitive World Models for Process-Level Social Influence Evaluation

Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, desires, intentions, and emotions measurably change over the course of conversation, a process-oriented criterion that neither surface-level text metrics (BLEU/ROUGE) nor single-score LLM judgments can capture. We propose the \textbf{Cog}nitive \textbf{W}orld \textbf{M}odel \textbf{(CogWM)}, an LLM-based user model that reframes multi-turn dialogue evaluation from ``what did the user say'' to ``how did the user's internal cognitive state evolves.'' CogWM jointly predicts BDI/E cognitive states and user utterances and serves as both a user simulator and an evaluation platform, using a three-tier evaluation framework that covers turn-level fidelity, trajectory-level state dynamics, and task-level composite scoring. Trained via our \textbf{S}ummarize-\textbf{a}nd-\textbf{A}llocate \textbf{(SaA)} annotation pipeline on 150,454 user-turn samples across four social influence scenarios, CogWM achieves 77.6\% emotion accuracy (2.1$\times$ over GPT-5.5). In 3600 multi-agent discrimination trials, it distinguishes six commercial agents by their cognitive influence, with Llama-4-Scout ranking first (CTS +0.233). CogWM moves social influence dialogue evaluation from terminal judgment to process tracking. We have released our code\footnote{\scriptsize Code: https://github.com/lucianma05-create/CogWM} and models\footnote{Model: https://www.modelscope.cn/models/LucianMa/CogWM-14B}.

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達

グラフ ニューラル ネットワーク モデルに対する生成的再構成攻撃の再考

グラフ データを多くの分野に応用することで、膨大な量のデータを収集して分析する必要性が生じており、その中にはプライベートで機密性の高いデータも含まれています。グラフ データの非ユークリッド的性質により、分析は計算的に困難になり、AI の時代ではグラフ ニューラル ネットワーク (GNN) の使用につながります。 GNN はトレーニングに使用した機密データを誤って漏洩する可能性があり、モデル反転攻撃などの深刻なデータ セキュリティ問題が発生します。この研究では、グラフラベル条件付き (GLC) 攻撃と埋め込みラベル条件付き (ELC) 攻撃という 2 つの新しいグラフ反転 (つまり、再構築) 攻撃を導入することにより、GNN の脆弱性を分析します。それぞれ、ターゲットモデルの予測とその中間表現を利用します。当社は、導入されたプライバシー攻撃の包括的な分析を実行し、3 つのベンチマーク グラフ データセット (つまり、NCI1、PROTEINS、および AIDS) および 4 つのグラフの分布/構造メトリクス (つまり、FGD、EGD、MMD、および GKS) にわたる既存のベースラインと比較します。私たちの研究は、攻撃者がジェネレーター・ディスクリミネーター技術を使用して、GNN に対する現実世界のブラックボックス攻撃シナリオで高品質のグラフを再構築できることを示しています。さらに、クエリを 50% 削減した攻撃の変種 (Ours--) を提示し、良好な、または同等の再構成攻撃パフォーマンスを達成しました。さらに、GNN はラプラシアン ノイズ スケールが変化するプライバシー攻撃に対して非常に脆弱であることを示します。

原文 (English)

Rethinking Generative Reconstruction Attacks against Graph Neural Network Models

The application of graph data in numerous disciplines raises the need for gathering and analyzing huge volumes of data, some of which is private and sensitive. The non-Euclidean nature of the graph data makes the analysis computationally challenging, leading to the use of Graph Neural Networks (GNNs) in the age of AI. GNNs may inadvertently leak sensitive data they are trained on, which raises serious data security issues, including the model inversion attack. In this study, we analyze GNNs' vulnerabilities by introducing two novel graph inversion (i.e., reconstruction) attacks: graph-label conditioned (GLC) attack and embedding-label conditioned (ELC) attack, utilizing targetmodel predictions and their intermediate representations, respectively. We perform a comprehensive analysis of our introduced privacy attacks and compare them with existing baselines across three benchmark graph datasets (i.e., NCI1, PROTEINS, and AIDS) and four graph distributional/structural metrics (i.e., FGD, EGD, MMD, and GKS). Our work demonstrates that an adversary can use the generator-discriminator technique to reconstruct high-quality graphs in real-world black-box attack scenarios against GNNs. Additionally, we present a variant of our attacks (Ours--) with 50% reduced queries, achieving good or comparable reconstruction attack performance. In addition, we show that GNNs are highly vulnerable to privacy attacks, varying Laplacian noise-scales.

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

CLQT: LLM ポートフォリオ管理エージェントの診断評価のための、クローズドループでコストを意識した、戦略に一貫したベンチマーク

LLM エージェントは自律的なポートフォリオ マネージャーとしての役割を果たすことが増えており、ベンチマークは財務上の質問応答から逐次取引へと移行しています。しかし、ほとんどのエージェントは依然として、固定ウィンドウでの収益によってエージェントをランク付けしています。これは、期間の収益が市場経路によって支配され、先読み漏れが制御されると見かけのアルファが解消される可能性があるため、弱い代理です。このようなランキングは、健全な推論、一貫した戦略、永続的な優位性を証明するものではありません。 CLQT を紹介します。CLQT は、閉ループ取引の評価をランク付けではなく診断として再構築したもので、エージェントのプロセスが成功または失敗した場所と理由を特定する手段です。 CLQT は、完全にクローズド ループで、コストを意識し、戦略に一貫性があり、一時的にゲートされた環境であり、そのエージェントは収集、合成、割り当て、実行、反映という 5 段階のサイクルを実行します。各ラウンドは、再計算検証可能なハッシュ チェーンに封印された完全な DecisionRound を出力するため、すべてのメトリクスは証跡から再構築可能です。ハード TimeGate、機関取引および資金調達コストのモデリング、戦略一貫性スコアリング、3 層メモリ、Model-Context-Protocol ツール層、および義務を意識した合成という 6 つの柱が基盤を形成します。同じエージェントが、特殊な役割の制約付き委員会または単一の完全自律型オーケストレーターとして実行され、プロセスの足場を実験変数にします。監査証跡から、5 軸の能力スコアカード (APM-CS: Coherence、Acuity、Composure、Discipline、Reliability) を計算します。Coherence は、自己選好バイアスを抑制するために保持されたコホート外の LLM によって部分的に判断されます。アブレーショングリッドを使用した汚染管理されたマルチモデルバックテストと、繰り返し実行されたノイズフロアに対して、目に見えないカットオフ後のデータ上のライブブローカートラックで検証します。 CLQT は結果と能力を分離し、モデルのランキングではなく、エージェントの能力と制限の耐久性と拡張可能なマップを生成します。

原文 (English)

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Yet most still rank agents by returns over a fixed window -- a weak proxy, since a period's return is dominated by the market path and apparent alpha can dissolve once look-ahead leakage is controlled. Such a ranking certifies neither sound reasoning, nor a consistent strategy, nor a durable edge. We introduce CLQT, which reframes closed-loop trading evaluation as diagnosis rather than ranking: an instrument that localizes where and why an agent's process succeeds or fails. CLQT is a fully closed-loop, cost-aware, strategy-consistent, temporally-gated environment whose agents run a five-stage cycle: gather, synthesize, allocate, execute, reflect. Each round emits a complete DecisionRound sealed into a recompute-verifiable hash chain, so every metric is reconstructable from the trail. Six pillars form the substrate: a hard TimeGate, institutional transaction- and financing-cost modeling, strategy-consistency scoring, three-tier memory, a Model-Context-Protocol tool layer, and mandate-aware synthesis. The same agent runs as a constrained committee of specialized roles or a single full-autonomy orchestrator, making process scaffolding an experimental variable. From the audit trail we compute a five-axis capability scorecard (APM-CS: Coherence, Acuity, Composure, Discipline, Reliability), with Coherence judged partly by a held-out, out-of-cohort LLM to curb self-preference bias. We validate it on a contamination-controlled multi-model backtest with an ablation grid and a live broker track on unseen, post-cutoff data, against a repeated-run noise floor. CLQT separates outcome from capability, yielding not a model ranking but a durable, extensible map of agent competencies and limitations.

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

EvalSafetyGap: LLM 評価と安全性の失敗に関するハイブリッド調査と概念的なフレームワーク

LLM の評価と AI の安全性は、共通の測定問題に直面しています。つまり、ベンチマーク スコア、報酬モデルのシグナル、報告される安全性メトリクスは向上する可能性がありますが、それらが表現するはずの潜在的な特性の検証は依然として困難です。この文書では、ハイブリッド調査 (物語の合成と個別に追跡される灰色の証拠と組み合わせた体系的な調査) を、概念的なフレームワークおよび構造化された 10 モデルの監査と組み合わせています。この統合は、ベンチマークの有効性、動的評価、裁判官としての LLM の信頼性、安全性評価、ジェイルブレイク/拒否の堅牢性、報酬ハッキング、機構の解釈可能性、ガバナンス/監査可能性の 8 つの証拠ストリームに及び、2018 年から 2026 年の評価安全性測定作業をカバーします。最適化の圧力下で評価側とアライメント側のプロキシ障害を比較するための組織化仮説として EvalSafetyGap を導入します。グッドハートの法則と、ここで開発した 2 つの構成要素 (不安定性分解とアライメントのトリレンマ) をテスト可能な比較を生成するツールとして使用します。この監査は、能力、行動安全性、ガバナンスを個別に測定した場合に結論がどのように変化するかを示しています。このサンプル (n = 10) では、表示された表 3 の入力を使用すると、能力と持続的な敵対的堅牢性の間の関連性は統計的に不確定であり (ピアソン r = +0.232、p = 0.520)、見かけ上のオープンとクローズの安全性ギャップは控えめであり、動作の堅牢性よりも主にガバナンスと開示によって左右され、単一の境界線モデルがどのように分類されるかに影響されます。試行予算の結果はプロトコルに依存します。公的証拠では異種プロトコルが使用されているため、監査はランク付けではなく診断的なものになります。この貢献は、動的評価、透明性のあるソースレポート、複数回の安全性測定、および監査可能な調整の実践をサポートするための共有ボキャブラリーと証拠マップです。

原文 (English)

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mechanistic interpretability, and governance/auditability, covering 2018-2026 evaluation-safety measurement work. We introduce EvalSafetyGap as an organizing hypothesis for comparing evaluation-side and alignment-side proxy failures under optimization pressure, using Goodhart's Law together with two constructs we develop here - an Instability Decomposition and an Alignment Trilemma - as tools for generating testable comparisons. The audit shows how conclusions shift when capability, behavioral safety, and governance are measured separately. In this sample (n = 10), the association between capability and sustained adversarial robustness is statistically indeterminate using the displayed Table 3 inputs (Pearson r = +0.232, p = 0.520), and the apparent open-closed safety gap is modest, driven mainly by governance and disclosure rather than behavioral robustness, and sensitive to how a single borderline model is classified; attempt-budget results are protocol dependent. Because the public evidence uses heterogeneous protocols, the audit is diagnostic rather than rank-generating. The contribution is a shared vocabulary and evidence map to support dynamic evaluation, transparent source reporting, multi-attempt safety measurement, and auditable alignment practice.

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

EMPATH: 感情サポート チャットボットの安全性評価のための多言語監査人/裁判官ベンチマーク

安全ベンチマークでは、プロンプト、言語、ターン構造を修正することで拡張性を確保することがよくあります。感情サポート チャットボットの場合、安全上の欠陥が現れる場所、つまり多言語で複数ターンにわたる危機に関する会話を通じて、その取引がまさに隠れています。感情サポート型チャットボットの安全性評価のベンチマーク「EMPATH」を紹介します。監査者モデルは、助けを求めるユーザーのロールプレイを行い、140 のシード指示と 34 のペルソナから複数ターンの会話を生成します。裁判官モデルは、危機対応、治療の質、会話の完全性、感情的安全性、文化的適応の 5 つの側面にわたる 19 の指標に基づいて、各完全な記録を採点します。 EMPATH はメキシコのスペイン語と米国英語向けに構築されています。ここで報告されている研究はメキシコのスペイン語で行われています。監査人と裁判官は異なるモデルファミリーから選出され、裁判官は信頼されるというよりも校正されるべき道具として扱われます。厳格な基準ごとのルーブリックにより、19 指標のうち 10 指標における重大なスコアのインフレが明らかになり、差別が回復されます。私たちは、裁判官の校正と家族を超えた裁判官間の合意を通じて、ベンチマークの測定特性を研究します。また、3 つのフロンティア モデルの EMPATH についても説明します。そのうちの 1 つはオープンウェイトです。集計スコアは互いに 0.74 ポイント以内に収まりますが、メトリックごとのプロファイルはモデル固有の場所で最大 6 ポイント異なります。標準ルーブリックでは、ランキングと弱点の両方が、家族を超えた 2 回目の審査で安定しています。スコアの 93% がプラスまたはマイナス 1 以内に収まります。5 回のテストと再テストで 2 番目の軸が追加されます。最も安定したモデルでも、同一の再実行で危機指標が 2 から 10 に変動し、deepseek-v4-pro は、温度 0 であっても実行ごとに異なる会話を返します。したがって、実行間の信頼性はモデルごとの安全特性です。平均化するためのノイズではありません。 EMPATH はシステムに依存しません。パイプライン、シード、ペルソナ、ルーブリックは再利用のためにリリースされます。

原文 (English)

EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides precisely where safety failures emerge: across a multilingual, multi-turn crisis conversation. We present EMPATH, a benchmark for safety evaluation of emotional-support chatbots. An auditor model role-plays help-seeking users, generating multi-turn conversations from 140 seed instructions and 34 personas. A judge model scores each full transcript against 19 metrics across five dimensions: crisis handling, therapeutic quality, conversational integrity, emotional safety, and cultural adaptation. EMPATH is built for Mexican Spanish and US English; the studies reported here run in Mexican Spanish. Auditor and judge are drawn from different model families, and the judge is treated as an instrument to be calibrated rather than trusted. A strict per-criterion rubric reveals material score inflation on 10 of the 19 metrics and restores discrimination. We study the measurement properties of the benchmark through judge calibration and cross-family inter-judge agreement. We also illustrate EMPATH on three frontier models, one of them open-weight. Aggregate scores sit within 0.74 points of one another, but per-metric profiles diverge by up to six points in model-specific places. Under the standard rubric, both the ranking and the weak spots are stable across a second, cross-family judge: 93% of scores fall within plus or minus 1. A five-run test-retest adds a second axis: even the steadiest model swings from 2 to 10 on a crisis metric across identical re-runs, and deepseek-v4-pro returns a different conversation on every run even at temperature 0. Run-to-run reliability is therefore a per-model safety property, not noise to average away. EMPATH is system-agnostic; the pipeline, seeds, personas, and rubrics are released for reuse.

2026-06-30 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Sequential Fairness Auditing with Limited Output Access

External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often ha…

2026-06-30 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

The Human Creativity Benchmark

Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflec…

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達

"AI Watermarking": Bridging Policy Discourse and Technical Capabilities

The widespread deployment of generative artificial intelligence (AI) models has raised serious concerns about the proliferation of AI-gener…

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達

Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa

Artificial intelligence depends on large-scale compute resources and their supporting infrastructure. However, AI governance debates treat…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly u…

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達

Defeat Devices in AI Systems

AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts. Alignment faking, sandbagg…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation

Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agen…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B

General-purpose large language models (LLMs) have demonstrated strong abilities in opendomain question answering, information extraction, a…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLMography: Transforming Human-AI Conversations into Traceability, Oversight, and Auditability Indicators

The growing use of Large Language Models (LLMs) in education, software engineering, academic writing, and technical documentation raises a…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models

Large Language Models (LLMs) are increasingly used as assistants across the software development lifecycle, yet their ability to reason abo…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis

Modern LLM deployments routinely compress models and raise sampling temperature to reduce cost, latency, or repetition, yet safety evaluati…

2026-06-30 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency

Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot disting…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databa…

2026-06-30 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies

Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challeng…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and ada…

2026-06-30 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration

Long-horizon web agents often fail in ways hidden by final-answer evaluation: they may visit useful pages, produce a well-formed answer, an…

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivaria…

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達

Overcoming Dependent Censoring in the Evaluation of Survival Models

Dependent censoring occurs when the event time and censoring time are not conditionally independent given the observed covariates. This com…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP

While NLP typically treats documents as independent and unordered samples, in longitudinal studies, this assumption rarely holds: documents…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects

Large Language Models (LLMs) have shown remarkable potential in developing role-playing agents (RPAs). However, current evaluation framewor…

2026-06-30 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Defeasible Conditional Obligation in a Two-tiered Preference-based Semantics (Extended Version)

In response to a concern raised by Horty, this paper develops a two-tiered, preference-based semantic framework for modeling defeasible con…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR

Automatic speech recognition replaces typing only when correction costs less than manual entry - a threshold determined by error types, not…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MedCase 構造化: 臨床的に現実的な EHR 設定における診断推論のベンチマーク用の Text-to-FHIR データセット

大規模言語モデル (LLM) は、臨床推論と意思決定のサポートに有望ですが、現実的な電子医療記録と一致する設定での評価には依然として限界があります。既存のベンチマークは、多くの場合、臨床システムで使用される構造化された相互運用可能なデータ形式を反映していない静的データセットまたは非構造化入力に依存しています。非構造化テキストから臨床的に現実的な HL7 FHIR R4 バンドルを生成するパイプラインを導入し、臨床意思決定支援システムの制御可能な評価を可能にします。このパイプラインは、段階的な LLM 生成と用語に基づいた検証および修復を組み合わせて、幻覚コードを削減し、構造的および意味的な一貫性を強化します。このアプローチを MedCaseReasoning に適用して、臨床医が作成した診断症例に合わせた合成データセットである MedCase-Structured を構築し、症例の 82.5% で有効な FHIR 生成を実現します。 MedCase-Structured での評価では、平文の場合よりも構造化 FHIR 入力での LLM の診断精度が一貫して低いことが明らかになり、展開に合わせたベンチマークの重要性が強調されています。

原文 (English)

MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in realistic, electronic health record-congruent settings remains limited. Existing benchmarks often rely on static datasets or unstructured inputs that do not reflect the structured, interoperable data formats used in clinical systems. We introduce a pipeline for generating clinically realistic HL7 FHIR R4 bundles from unstructured text, enabling controllable evaluation of clinical decision support systems. The pipeline combines staged LLM generation with terminology-grounded validation and repair to reduce hallucinated codes and enforce structural and semantic consistency. Applying this approach to MedCaseReasoning, we construct MedCase-Structured, a synthetic dataset aligned with clinician-authored diagnostic cases, achieving valid FHIR generation for 82.5% of cases. Evaluation on MedCase-Structured reveals consistently lower diagnostic accuracy for LLMs on structured FHIR inputs than with plain text, highlighting the importance of deployment-aligned benchmarking.

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

BenHalluEval: ベンガル語の大規模言語モデル用のマルチタスク幻覚評価フレームワーク

ベンガル語は世界で 6 番目に話されている言語であるにもかかわらず、ベンガル語の大規模言語モデル (LLM) で幻覚を体系的に評価した先行研究はありません。 BenHalluEval は、生成的質問応答 (GQA)、バングラ語と英語のコード混合 QA、要約、および推論の 4 つのタスクをカバーするベンガル語用のきめ細かい幻覚評価フレームワークです。既存の 3 つのベンガル語データセットから抽出された、12 のタスク固有の幻覚タイプにわたって GPT-5.4 を使用して 12,000 の幻覚候補を構築し、グラウンドトゥルース インスタンスの偽陽性率 (トラック A) と幻覚候補の幻覚検出率 (トラック B) を独立して測定するデュアル トラック プロトコルの下で、推論指向、多言語、ベンガル語中心のカテゴリにわたる 7 つの LLM を評価します。両方の故障モードに共同でペナルティを課し、均一な応答バイアスによるスコアのインフレを防ぐために、モデルとタスク全体で 7.72% から 55.42% の範囲のデュアルトラック キャリブレーション メトリクスである BenHalluScore を提案します。これは、幻覚キャリブレーションの大幅な変動を明らかにします。緩和戦略として適用される思考連鎖プロンプトは、幻覚差別を一貫して改善することなく、反応分布を変化させます。 BenHalluEval は、ベンガル語専用の幻覚ベンチマークを初めて確立し、リソースの少ない言語設定に対する単一トラックおよびプロンプトのみの評価アプローチが不適切であることを強調しています。データセットとコードは https://anonymous.4open.science/r/BanglaHalluEval-EB77 で入手できます。

原文 (English)

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates using GPT-5.4 across twelve task-specific hallucination types, drawn from three existing Bengali datasets, and evaluate seven LLMs spanning reasoning-oriented, multilingual, and Bengali-centric categories under a dual-track protocol that independently measures false-positive rate on ground-truth instances (Track A) and hallucination detection rate on hallucinated candidates (Track B). To jointly penalise both failure modes and prevent inflated scores from uniform response bias, we propose BenHalluScore, a dual-track calibration metric that ranges from 7.72% to 55.42% across models and tasks, revealing substantial variation in hallucination calibration. Chain-of-thought prompting, applied as a mitigation strategy, shifts response distributions without consistently improving hallucination discrimination. BenHalluEval establishes the first dedicated hallucination benchmark for Bengali and highlights the inadequacy of single-track and prompting-only evaluation approaches for low-resource language settings. The dataset and code are available at https://anonymous.4open.science/r/BanglaHalluEval-EB77.

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

A language model's memory can be worse than no memory at all. A memory that keeps a wrong conclusion but drops the work behind it makes the…

2026-06-29 23:00 JSTTechCrunch AIロボティクスビジネス/資金調達

Robot hand company settles Tesla trade secret suit and announces $11M raise

The startup, Proception, is taking a unique approach to collecting training data to tackle one of the hardest problems in robotics: hands.

2026-06-29 22:00 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

Omen AI’s plan to optimize data centers is all wet

Omen AI raised a $31 million Series A to monitor chip coolant and stop bacterial outbreaks in data centers.

2026-06-29 13:30 JSTITmedia AI+エージェントビジネス/資金調達

AIエージェントの投資優先順位、どう決める? Gartnerが「投資スコア」の作り方を公開

業務におけるAIエージェントの投資優先順位をどう決めればよいか。業務・業種別のAIエージェントはどう進化していくか。ガートナージャパンの著名アナリストである亦賀忠明氏のWebセミナーから探る。

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

信頼性が高く堅牢な LLM 計画に向けて: シンボリック フィードバック駆動の反復的自己洗練フレームワーク

大規模言語モデル (LLM) は学界や産業界から広く注目を集めていますが、その導入には堅牢性と信頼性に関する重大なセキュリティ上の懸念が生じます。インテリジェントな動作の中核コンポーネントである計画は、LLM にとって依然として課題であり、本質的な複雑さのために、長期的な意思決定タスクでは実行不可能または不正確な解決策を生み出すことがよくあります。この論文では、長期計画における LLM の堅牢性と信頼性を強化するための、記号的なフィードバック駆動型の反復的自己洗練フレームワークを提案します。具体的には、論理シンボルを自然言語記述にマッピングするための自然言語プロンプト メカニズムが導入され、LLM がタスクの制約とセマンティクスをより適切に把握できるようになります。さらに、エラーを特定し、LLM が解釈できる修正指示に変換する記号検証器を設計し、それによって自己改善を導きます。さらに、計画認識機能を活用して目標の達成可能性を推測し、望ましい目標に向けたより効果的なガイダンスを促進します。経験的な結果は、提案されたフレームワークが長期的な計画タスクの実現可能性と正確性の両方を一貫して向上させることを示しています。これは、LLM ベースの計画の信頼性を高める有効性と、より信頼できる AI システムを可能にする可能性を強調しています。

原文 (English)

Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework

Large language models (LLMs) have attracted widespread attention from academia and industry, yet their deployment raises critical security concerns regarding robustness and reliability. Planning, a core component of intelligent behavior, remains challenging for LLMs, which often produce infeasible or incorrect solutions in long-horizon decision-making tasks due to inherent complexity. In this paper, we propose a symbolic feedback-driven iterative self-refinement framework to enhance the robustness and reliability of LLMs in long-horizon planning. Specifically, a natural language prompting mechanism is introduced to map logical symbols into natural language descriptions, enabling LLMs to better capture task constraints and semantics. We further design a symbolic verifier that identifies errors and converts them into corrective instructions interpretable by the LLM, thereby guiding self-refinement. In addition, we leverage a plan recognizer to infer goal reachability, facilitating more effective guidance toward desired goals. Empirical results demonstrate that the proposed framework consistently improves both feasibility and correctness in long-horizon planning tasks. This highlights its effectiveness in enhancing the reliability of LLM-based planning and potential to enable more trustworthy AI systems.

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

コーディング LLM における暗黙的なソフトウェア ワールド モデルの評価に向けて

ソフトウェア エンジニアリングは、人間によって実行されるか AI エージェントによって実行されるかにかかわらず、ソフトウェアがどのように動作するかについて推論する必要があります。私たちは、このような推論をサポートする内部モデルをソフトウェア ワールド モデルと呼び、現在のコード実行ベンチマークは、よく研究されたその一部である制御フローをカバーしていると見ています。このペーパーでは、観察可能な軸を実行リソースに移すことで、より広範な評価に向けて一歩を踏み出します。テスト結果と例外クラスに加えて、ピーク メモリ、実時間、およびメソッドとラインの粒度でランク付けされたプロファイラー出力を予測します。実際のソフトウェア エンジニアリング タスクに近いテストを行うため、データ ソースとして SWE-bench Verified を使用します。最先端のモデルも含め、テストされたすべてのモデルは、控えめなパフォーマンスと脆弱な動作を示しており、ソース コードがどのように記述されているかではなく、ソフトウェアがどのように実行されるかについての理解が著しく欠如していることを示唆しています。

原文 (English)

Towards Evaluation of Implicit Software World Models in Coding LLMs

Software engineering, whether performed by humans or by AI agents, requires reasoning about how software behaves. We call the internal model that supports such reasoning the software world model, and view current code-execution benchmarks as covering one well-studied slice of it -- control flow. In this paper, we take a step toward a broader evaluation by shifting the observable axis to execution resources: alongside test outcome and exception class, we predict peak memory, wall-clock time, and ranked profiler outputs at method and line granularity. We use SWE-bench Verified as the source of data to hold the test close to real-world software engineering tasks. All tested models, frontier ones included, show modest performance and brittle behaviour, suggesting a notable lack of understanding of how software is executed, as opposed to how its source code is written.

2026-06-29 13:00 JSTarXiv cs.AIビジネス/資金調達

単細胞 RNA シーケンスを使用したヒト脂肪組織における脂肪細胞の発生軌跡の再構築

肥満は、2 型糖尿病や心血管疾患などの代謝障害に関連した世界的な健康危機です。この研究では、単一細胞 RNA シーケンスを使用して、脂肪組織サンプルからヒト脂肪細胞の発生軌跡を再構築しました。私たちの分析により、7 つの遷移状態を含む 15 の転写的に異なる細胞クラスターが特定され、脂肪細胞分化の動的なプロセスが明らかになりました。我々は、脂肪細胞とその前駆細胞の間の細胞コミュニケーションを媒介する、機能的に活性なシグナル伝達経路を16個検出した。これらの中で、インスリン様成長因子 (IGF) および線維芽細胞成長因子 (FGF) 経路が最も顕著なネットワークとして浮上し、分化段階全体で一貫した活性を示しました (p<0.05)。この研究では、内臓脂肪細胞が皮下分化には存在しない追加の細胞外マトリックスのリモデリングを受けるという、デポー特異的な違いが明らかになりました。さらに、空間解析では、IGF シグナル伝達が血管周囲ニッチで特に活性であるのに対し、FGF 活性は成熟脂肪細胞ゾーンで優勢であることが示されました。これらの結果は、潜在的な治療標的としてのIGFおよびFGF経路を強調する、ヒト脂肪細胞発生の最初の包括的なマップを提供する。同定されたシグナル伝達ネットワークは、健康な脂肪の拡大を促進したり、病的な脂肪の蓄積を抑制したりするための介入を開発するための新たな洞察を提供します。この研究は、代謝障害の治療に臨床的に関連するデータを提供しながら、脂肪組織生物学の基本的な理解を促進します。

原文 (English)

Reconstructing the Developmental Trajectory of Adipocytes in Human Adipose Tissue Using Single-Cell RNA Sequencing

Obesity is a global health crisis associated with metabolic disorders such as type 2 diabetes and cardiovascular disease. This study employed single-cell RNA sequencing to reconstruct the developmental trajectory of human adipocytes from adipose tissue samples. Our analysis identified 15 transcriptionally distinct cell clusters, including 7 transitional states, revealing the dynamic process of adipocyte differentiation. We detected 16 functionally active signaling pathways mediating cellular communication between adipocytes and their progenitors. Among these, insulin-like growth factor (IGF) and fibroblast growth factor (FGF) pathways emerged as the most prominent networks, showing consistent activity across differentiation stages (p<0.05). The study revealed depot-specific differences, with visceral adipocytes undergoing additional extracellular matrix remodeling absent in subcutaneous differentiation. Spatial analysis further showed that IGF signaling was particularly active in perivascular niches, while FGF activity dominated in mature adipocyte zones. These results provide the first comprehensive map of human adipocyte development, highlighting IGF and FGF pathways as potential therapeutic targets. The identified signaling networks offer new insights for developing interventions to promote healthy adipose expansion or inhibit pathological fat accumulation. This work advances our fundamental understanding of adipose tissue biology while providing clinically relevant data for metabolic disorder treatments.

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Triadic Werewolf: LLM における心のマルチホップ理論における道化師の役割

大規模な言語モデルの心の理論による評価では通常、二項社会演繹ゲームが使用されます。このゲームでは、観察可能なすべての手がかりが 1 つの隠れた側面を指すため、強力な言語事前条件を持つモデルは、対戦相手のインセンティブをシミュレートすることなく高いスコアを獲得できます。人狼ゲームを道化師で拡張します。道化師は、投票によって勝利するため、同僚の疑いに対する効用が逆転する 3 番目の派閥です。そのため、最適なプレイには 3 つの相反する効用関数にわたる推論が必要です。 GPT-4.1、DeepSeek-V3.1、および Llama-3.3-70B で道化師の自己学習をオンまたはオフにして 60 試合を行ったところ、道化師はゲームの 60 ~ 70% で勝利しましたが、ウェアウルフが 20% を超えることはありませんでした。また、GPT-4.1 のオオカミはゲームの 60 ~ 70% で初日に道化師を投票で除外します。これは厳密に自己破滅的なアクションです。自己学習は DeepSeek と Llama には役立ちますが、GPT-4.1 には害があり、そのコストは狼男ではなく村人にかかっています。 DeepSeek だけが、意図的に疑わしく見せることなく、疑わしく見せるという微妙な戦略を学習し、ループから最大限の利益を得ます。三項インセンティブ構造は、二項演繹ゲームが目に見えないままにしていたマルチエージェント推論の層を明らかにします。

原文 (English)

Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs

Theory-of-mind evaluations of large language models typically use dyadic social-deduction games, where every observable cue points to a single hidden side, so a model with strong language priors can score well without ever simulating opponents' incentives. We extend the Werewolf game with a Jester, a third faction whose utility on peer suspicion is inverted because it wins by being voted out, so optimal play requires reasoning across three opposing utility functions. Across 60 games on GPT-4.1, DeepSeek-V3.1, and Llama-3.3-70B with Jester self-learning on and off, the Jester wins 60-70% of games while Werewolves never exceed 20%, and GPT-4.1 wolves vote the Jester out on day 1 in 60-70% of games, a strictly self-defeating action. Self-learning helps DeepSeek and Llama but hurts GPT-4.1, with the cost landing on Villagers rather than Werewolves. Only DeepSeek learns the subtle strategy of looking suspicious without looking intentionally suspicious, and it gains the most from the loop. Triadic incentive structure exposes a layer of multi-agent reasoning that dyadic deduction games leave invisible.

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA

LLM-as-a-Judge and self-evaluation pipelines implicitly assume that evaluation is easier than generation. We test this in a controlled in-c…

2026-06-29 13:00 JSTarXiv cs.AIビジネス/資金調達

大規模言語モデルを使用して縦断的な合成臨床ノートを生成するパイプライン

実世界のデータへのアクセスが制限されている領域で AI システムの開発と評価を可能にするために、合成データの使用が増えています。医療分野では、臨床文書はその機密性により特別な課題を抱えています。この研究では、実際の患者データに伴うプライバシー リスクを回避しながら、臨床 AI ツールの開発をサポートするように設計された合成臨床メモ パイプラインとデータセットを導入します。データセットは、大規模な言語モデルを使用した構造化患者生成、半構造化患者ジャーニー シミュレーション、および非構造化臨床ノート生成を組み合わせたモジュール式パイプラインを使用して生成されます。このパイプラインは、長期的な患者記録全体にわたる内部一貫性を優先すると同時に、書き方、メモの構造、臨床の詳細の変化も捕捉するように設計されています。 LLM ベースの検証および拡張ステップを含む追加のメカニズムを使用して、生成されたノートの忠実性、リアリズム、および多様性が向上します。私たちは、70 人の合成患者のデータセットをリリースします。各患者には、入院期間全体にわたる 20 ~ 50 の臨床ノートが関連付けられています。データセットは複数の検証レベルで提供されているため、ユーザーはユースケースに応じて現実性とスケーラビリティのバランスを取ることができます。このデータセットは、実際の患者データに依存することなく、要約ツール、コーディング モデル、意思決定支援システムなどの臨床 AI システムの開発、テスト、評価をサポートします。

原文 (English)

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior

Understanding student engagement usually requires time-consuming manual observation or invasive recording that raises privacy concerns. We…

2026-06-29 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making…

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems

As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. C…

2026-06-27 07:00 JSTITmedia AI+ビジネス/資金調達

官民投資フィジカルAIに10.5兆円示す、「実証から実装へ」動き出す現場

2026年6月22日~26日に公開された記事の中から、MONOist編集部が厳選した今週の注目ニュースをお届けします。

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

マルチモーダル LLM 評価に欠けているものは何ですか?

マルチモーダル大規模言語モデル (MLLM) は、テキスト、画像、音声、ビデオなどのさまざまな入力を処理し、テキスト応答を生成できます。それらの機能は急速に進歩していますが、そのようなモデルの評価は追いついていません。既存の評価ベンチマークのほとんどは、個別のタスクに限定されており、モデルがモダリティ全体で情報を統合しているかどうかについてはほとんど明らかにされていません。私たちは、MLLM を評価するための現在の手段を調査し、既存のベンチマーク分類をレビューして、時間空間的一貫性、物理世界の理解、マルチモーダル一貫性、選択的注意などのギャップを特定します。これらのギャップに対処することは、マルチモーダル インテリジェンスの実際の進歩を測定し、機能の境界を明らかにするために不可欠です。

原文 (English)

What We are Missing in Multimodal LLM Evaluation?

Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced rapidly, evaluation of such models has not kept pace. Most existing evaluation benchmarks are limited to isolated tasks and reveal little about whether a model integrates information across modalities. We examine current means for evaluating MLLMs and review the existing benchmark taxonomy to identify gaps, including temporal-spatial coherence, physical world understanding, multimodal consistency, and selective attention. Addressing these gaps is essential for measuring real progress in multimodal intelligence and exposing capability boundaries.

2026-06-26 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

OpenFinGym: クオンツエージェントを評価するための検証可能なマルチタスクジム環境

大規模な言語モデル エージェントは定量的財務ワークフローにますます適用されていますが、その評価は分離されたタスク間で断片化されたままであり、ベンチマーク タスクの財務関連性はしばしば見落とされます。しかし、財務ワークフローは本質的に多段階であり、予測、戦略構築、リスク管理、取引などの相互依存するタスクにまたがっています。既存のプラットフォームは通常、単一のタスクに焦点を当てているため、エージェントの能力を過大評価し、一般化、実際の市場でのやり取り、および財務的に意味のある意思決定における弱点を明らかにできません。 OpenFinGym は、単一の実行および検証インターフェイスで予測、市場生成、リ​​アルタイム取引、不正検出をカバーする定量的金融エージェント開発用の統合ジム環境です。 OpenFinGym はさらに、定量的な財務出版物を実行可能なタスク パッケージに変換する自動タスク構築パイプラインを提供します。スケーラブルなエージェントのロールアウトをサポートし、ランタイムのトレインテストの漏洩を防ぐホスト側検証サービスを備えたコンテナ化されたランタイム。低レイテンシのデータストリーム設計を備えたペーパートレーディングエンジン。長期およびイベント市場の予測に対する遅延解像度のサポート。トレーニング後の SFT と RL の統合

原文 (English)

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tasks, while the financial relevance of benchmark tasks is often overlooked. Yet financial workflows are inherently multi-stage, spanning interdependent tasks such as forecasting, strategy construction, risk management, and trading. Existing platforms typically focus on a single task, and can therefore overstate agent competence and fail to reveal weaknesses in generalization, real-market interaction, and financially meaningful decision-making. We introduce OpenFinGym, a unified gym environment for quantitative-finance agent development that covers forecasting, market generation, real-time trading, and fraud detection under a single execution and verification interface. OpenFinGym additionally provides an automated task-construction pipeline that turns quantitative finance publications into executable task packages; a containerised runtime with a host-side verifier service that supports scalable agent rollouts and prevents runtime train-test leakage; a paper trading engine with a low-latency data-stream design; deferred-resolution support for long-horizon and event-market forecasts; and integration for SFT and RL post-training

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

大規模言語モデルを使用して縦断的な合成臨床ノートを生成するパイプライン

実世界のデータへのアクセスが制限されている領域で AI システムの開発と評価を可能にするために、合成データの使用が増えています。医療分野では、臨床文書はその機密性により特別な課題を抱えています。この研究では、実際の患者データに伴うプライバシー リスクを回避しながら、臨床 AI ツールの開発をサポートするように設計された合成臨床メモ パイプラインとデータセットを導入します。データセットは、大規模な言語モデルを使用した構造化患者生成、半構造化患者ジャーニー シミュレーション、および非構造化臨床ノート生成を組み合わせたモジュール式パイプラインを使用して生成されます。このパイプラインは、長期的な患者記録全体にわたる内部一貫性を優先すると同時に、書き方、メモの構造、臨床の詳細の変化も捕捉するように設計されています。 LLM ベースの検証および拡張ステップを含む追加のメカニズムを使用して、生成されたノートの忠実性、リアリズム、および多様性が向上します。私たちは、70 人の合成患者のデータセットをリリースします。各患者には、入院期間全体にわたる 20 ~ 50 の臨床ノートが関連付けられています。データセットは複数の検証レベルで提供されているため、ユーザーはユースケースに応じて現実性とスケーラビリティのバランスを取ることができます。このデータセットは、実際の患者データに依存することなく、要約ツール、コーディング モデル、意思決定支援システムなどの臨床 AI システムの開発、テスト、評価をサポートします。

原文 (English)

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

グラウンドトゥルースを使用してクラスタリングを評価するにはどうすればよいですか?

グランド トゥルースが利用可能な場合、外部インデックスをクラスター評価に使用できます。セットマッチングベースの尺度に焦点を当てて、最も一般的な外部妥当性指標をレビューします。セントロイド インデックス (CI) は、説明可能な結果が得られる直感的なクラスター レベルの測定であるため、推奨します。より細かく調整されたポイントレベルの測定が必要な場合は、より多くの選択肢があります。ペアセット インデックス (PSI) は、クラスター サイズによって偏らない正規化されたスコアを提供します。すべてのポイントが同等に重要である必要がある場合は、クラスタリング精度 (ACC) またはその他のセットマッチング尺度が適しています。

原文 (English)

How to evaluate clustering with ground truth?

External indexes can be used for cluster evaluation when ground truth is available. We review the most common external validity indexes focusing on set-matching-based measures. We recommend centroid index (CI), because it is an intuitive cluster-level measure with an explainable result. If we need a more fine-tuned, point-level measure, there are more choices. Pair-set index (PSI) provides a normalized score which is not biased by cluster sizes. If all points should matter equally, then clustering accuracy (ACC) or any other set-matching measure is suitable.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

判断せずに質問する: 解釈可能な LLM 評価と自己改善のための 2 つの質問

NLP では、LLM 出力の評価が依然として大きなボトルネックとなっています。人間による評価は高価で時間がかかり、語彙メトリクスとオープンエンド生成に関する人間の判断との相関性が低く、全体的な LLM ジャッジはデバッグが難しい不透明なスコアを生成することがよくあります。私たちは、評価基準をアトミックなバイナリの質問に分解し、その結果の判定を解釈可能な多次元スコアに集約するフレームワークである BINEVAL を提案します。タスク プロンプトが与えられると、メタ プロンプトが詳細な評価質問を生成し、LLM が出力ごとに独立して質問に回答し、調整された全体スコアとともに透明な質問レベルのフィードバックを生成します。この分解により、評価が検査、​​診断が容易になり、迅速な改善に直接使用できるようになります。 SummEval、Topical-Chat、QAGS 全体で、BINEVAL は UniEval や G-Eval などの強力なベースラインと同等またはそれを上回り、特に QAGS などの事実整合性ベンチマークで優れた結果を示しています。 BINEVAL は、人間の判断との競合相関を超えて、人間のスコア分布とよりよく一致し、以前の LLM ジャッジによく見られた天井効果を回避し、境界線にある出力と明らかに欠陥のある出力をより適切に区別することにつながります。さらに、同じ質問レベルのフィードバックが反復プロンプトの最適化をサポートし、自己更新設定とクロスモデル更新設定の両方で IFBench での要約に関する評価者のプロンプトと生成プロンプトを改善することを示します。全体として、BINEVAL は、強力な経験的パフォーマンスと実用的な診断および最適化の価値を組み合わせた、タスクに依存せず、トレーニング不要で、解釈可能な評価フレームワークを提供します。

原文 (English)

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that decomposes evaluation criteria into atomic binary questions and aggregates the resulting verdicts into interpretable, multi-dimensional scores. Given a task prompt, a meta-prompt generates fine-grained evaluation questions, and an LLM answers them independently for each output, yielding transparent question-level feedback together with calibrated overall scores. This decomposition makes evaluation easier to inspect, easier to diagnose, and directly usable for prompt improvement. Across SummEval, Topical-Chat, and QAGS, BINEVAL matches or outperforms strong baselines including UniEval and G-Eval, with especially strong results on factual consistency benchmarks such as QAGS. Beyond competitive correlation with human judgments, BINEVAL better matches human score distributions and avoids the ceiling effects common in prior LLM judges, leading to better discrimination between borderline and clearly flawed outputs. We further show that the same question-level feedback supports iterative prompt optimization, improving evaluator prompts on summarization and generation prompts on IFBench under both self-update and cross-model update settings. Overall, BINEVAL provides a task-agnostic, training-free, and interpretable evaluation framework that combines strong empirical performance with practical diagnostic and optimization value.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Know2Guess: 大規模言語モデルにおける知識境界評価のための汚染を認識したマルチゾーン ベンチマーク

大規模な言語モデルの信頼性の高い評価では、データの汚染、プロンプトの特異性、または一般的な拒否行動と混同することなく、サポートされている回答とサポートされていない推測を分離する必要があります。凍結されたビルドタイム ラベルの下で、回答可能な知識から棄権が期待される未知への移行を測定するための、汚染を認識したマルチゾーン ベンチマークを示します。このベンチマークには、5 つのドメインにわたる 1,200 項目、明示的な棄権期待、汚染リスクのメタデータ、および公式の厳密なパーサーと正規化された堅牢性パーサーによる二重解析が含まれています。ロックされた回答または棄権プロンプト、回答のみのコントロール、およびプロンプト テンプレートのバリアントの下で、FLAN-T5、Qwen2.5-Instruct、および Llama-3-Instruct モデルを評価します。このベンチマークは、一般的な無回答行動では解決されません。FLAN のベースラインは、生産的な棄権に関しては弱いままですが、より強力な指導調整モデルは、選択的ではあるが回答から棄権への移行が不完全であることを明らかにしています。 Qwen2.5-3B-Instruct は全体的に最高の信頼性を実現していますが、回答が期待されるゾーンは依然として難しく、キャリブレーションは依然として不十分で、良性の項目の拒否は引き続き発生します。プロンプトおよびパーサーの堅牢性分析により、主要なランキングと定性的な結論が維持されます。したがって、このベンチマークは、回答可能性、棄権、拒否、および汚染を、LLM の信頼性の個別だが相互作用する側面として監査するための再現可能なプロトコルを提供します。データセットは、https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark で公開されています。

原文 (English)

Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels. The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict parser plus a normalized robustness parser. We evaluate FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models under locked answer-or-abstain prompts, answer-only controls, and prompt-template variants. The benchmark is not solved by generic non-answer behavior: FLAN baselines remain weak on productive abstention, while stronger instruction-tuned models expose a selective but incomplete transition from answering to abstaining. Qwen2.5-3B-Instruct achieves the best overall reliability, but answer-expected zones remain difficult, calibration remains poor, and benign-item refusal persists. Prompt and parser robustness analyses preserve the main ranking and qualitative conclusions. The benchmark therefore provides a reproducible protocol for auditing answerability, abstention, refusal, and contamination as distinct but interacting dimensions of LLM reliability.The dataset is publicly available at https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

LLM エージェントの即時注入に対する帯域外防御の適応的評価

最近の研究 (2024 年から 2026 年) は、ツールを使用する LLM エージェントを間接的なプロンプト インジェクションから防御するための戦略に収束しました。悪意のある命令を拒否するようにモデルをトレーニングするのではなく、エージェントのアクションを仲介する決定論的なポリシーでモデルの外側にセキュリティを強化します。 CaMeL、FIDES、Progent、RTBAS、FORGE などのシステムは、機能、情報フロー ラベル、参照モニターによってこれを実現しており、いくつかのシステムは AgentDojo ベンチマークに対する攻撃がほぼ排除されたと報告しています。私たちは 2 つの貢献をします。まず、これらの帯域外防御を古典的な完全性保護 (Biba)、参照モニタリング、最小特権のインスタンスとして整理し、それらがカバーするものとカバーしないものを構造化して比較します。第 2 に、それらのすべては静的ベンチマーク (一連の固定された注入試行) でのみ検証されることを警告します。これは、適応的で防御を意識した攻撃が 90% 以上の成功率でそのうち 12 件を破るまで、帯域内防御が強力であるように見えたのと同じ方法論です。適応的評価に必要な脅威モデルとプロトコルを指定します。次に、そのプロトコルを、Progent 独自の適応型攻撃分析の独立した再現および拡張として、単一の H200 上で自己ホストされるオープンウェイト エージェント (Qwen2.5-7B) を使用して AgentDojo 上で実行します。この設定は作成者がテストしていません。 3 回のランを平均すると、ディフェンスは次の結果を維持しました。プロジェントは平均攻撃成功率を約 6 倍 (25.8% から 4.2%) に削減しましたが、手作りの適応攻撃では平均攻撃成功率は上がりませんでした (2.6%)。これは、単一のブラックボックス攻撃テンプレートを備えた弱いモデル上の 1 つの小規模なデータ ポイントです。より強力な最適化された (ホワイトボックス GCG) 攻撃が依然として存在します。この結果は、適応型攻撃者にとって、決定論的な帯域外強制の方が帯域内検出よりも難しいターゲットであるという仮説と一致しますが、それを確立するものではありません。

原文 (English)

Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions. Systems such as CaMeL, FIDES, Progent, RTBAS, and FORGE realize this with capabilities, information-flow labels, and reference monitors, and several report near-elimination of attacks on the AgentDojo benchmark. We make two contributions. First, we organize these out-of-band defenses as instances of classical integrity protection (Biba), reference monitoring, and least privilege, yielding a structured comparison of what they do and do not cover. Second, we warn that every one of them is validated only on static benchmarks (a fixed set of injection attempts), the same methodology that made in-band defenses look strong until adaptive, defense-aware attacks broke twelve of them at over 90% success; we specify the threat model and protocol an adaptive evaluation requires. We then run that protocol as an independent reproduction and extension of Progent's own adaptive-attack analysis, on AgentDojo, with an open-weight agent (Qwen2.5-7B) self-hosted on a single H200, a setting its authors did not test. Averaged over three runs, the defense held: Progent cut mean attack success roughly sixfold (25.8% to 4.2%), and a hand-crafted adaptive attack did not raise it (2.6%). This is one small-scale data point on a weak model with a single black-box attack template; a stronger optimized (white-box GCG) attack remains open. The result is consistent with, but does not establish, the hypothesis that deterministic out-of-band enforcement is a harder target for an adaptive attacker than in-band detection.

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

深層学習プログラムの故障診断における評価と戦略のギャップ

深層学習 (DL) プログラムはさまざまな理由でトレーニング中に失敗する可能性があり、原因の診断はコ​​ストと時間のかかるメンテナンス作業です。このような障害を診断する技術は、通常、プログラム内の相互検証を使用して評価されますが、これまでに確認されていないプログラムが関与する展開設定には不適切な場合があります。したがって、これらの設定間でパフォーマンスがどのように異なるかを評価し、確立された DL の障害診断技術におけるパフォーマンス ギャップの原因を特定する必要があります。私たちは、38 の実世界の DL プログラムからの 5,542 個のフォールト注入トレーニング トレースのコーパスである DynFault を使用して、このギャップを調査します。既存の障害診断技術では、プログラム内での評価とプログラム全体を実行した場合のバランスの取れた精度に 0.190 のギャップがあることがわかりました。また、このギャップは機能のプログラム レベルの構造に起因することもわかり、2 つのランタイム機能セット、曲率機能とオプティマイザー機能、および目に見えないプログラムでのそれらの動作を調査することになりました。曲率機能は目に見えないプログラムの不安定性の検出に役立ちますが、オプティマイザーとアクティベーション機能はトレーニング中に表示されたプログラムにのみ役立ちます。

原文 (English)

Evaluation-Strategy Gap in Fault Diagnosis of Deep Learning Programs

Deep Learning (DL) programs can fail during training for many reasons, and diagnosing the cause is a costly and time-consuming maintenance task. Techniques for diagnosing such failures are commonly assessed using within-program cross-validation, which may be inadequate for deployment settings involving previously unseen programs. It is therefore necessary to assess how performance differs across these settings and to identify the causes of any performance gap in established fault diagnosis techniques for DL. We investigate this gap using DynFault, a corpus of 5,542 fault-injected training traces from 38 real-world DL programs. We found a gap of 0.190 in balanced accuracy for existing fault diagnosis techniques between within-program evaluation and holding out whole programs. We also found the gap comes from program-level structure in the features, which led us to examine two runtime feature sets, curvature features and optimizer features, and their behavior on unseen programs. We found that curvature features are useful for instability detection on unseen programs, while optimizer and activation features help only on programs seen during training.

2026-06-26 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

幻覚からグラウンディングまで: CRISP による視覚空間知能の診断

現在の VLM 評価では、言語事前分布と真の空間推論が混同されることがよくあります。これに対処するために、CRISP を導入します。CRISP は、一貫性、つまり暗黙の認識と明示的な推論の整合性を通じて視覚空間知能を評価する、新しい構造診断評価パラダイムです。従来のブラックボックス QA とは異なり、CRISP はメトリック 3D シーン グラフとオラクル介入プロトコルを利用して、潜在的な推論機能を知覚のボトルネックから切り離します。この詳細な診断により、体系的な知覚と推論の断絶が明らかになります。重要なことに、独自のモデルは強力な潜在的な推論エンジンを備えているものの、不正確な計量推定と暗黙の構造表現を活用する重大な失敗に悩まされていることを明らかにしました。逆に、オープンソース モデルは、マルチホップの構成推論が欠如しているため、依然として根本的なボトルネックとなっています。 CRISP は、事前言語を介して単に「正しく推測する」ことから、真に「知覚、検証、推論する」ことに焦点を移すことで、エンドツーエンドのポストトレーニングを超えたマルチモーダル調整のための厳密なロードマップを提供します。コードとデータセットは https://github.com/iiyamayuki/CRISP-Bench で入手できます。

原文 (English)

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP

Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning. Unlike traditional black-box QA, CRISP utilizes metric 3D Scene Graphs and an oracle intervention protocol to decouple latent reasoning capabilities from perceptual bottlenecks. This granular diagnosis uncovers a systematic perception-reasoning disconnect. Crucially, we reveal that while proprietary models possess robust latent reasoning engines, they suffer from inaccurate metric estimation and a critical failure to leverage their implicit structural representations. Conversely, open-source models remain fundamentally bottlenecked by their lack of multi-hop compositional reasoning. By shifting the focus from merely ``guessing correctly'' via language priors to genuinely ``perceiving, verifying, and reasoning,'' CRISP offers a rigorous roadmap for multimodal alignment beyond end-to-end post-training. The code and dataset are available at https://github.com/iiyamayuki/CRISP-Bench.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

HiLSVA: 科学的視覚化のための人間参加型エージェント システムの設計と評価

大規模言語モデル (LLM) エージェントは、科学的視覚化 (SciVis) のための自然言語対話を可能にします。それでも、従来のシステムは基本的に人間による分析制御よりも自律性を優先しており、そのため透明性と人間による監視が制限されていました。混合イニシアチブの SciVis ワークフローをサポートする人間参加型エージェント システムである HiLSVA を紹介します。 HiLSVA は、計画優先のマルチエージェント アーキテクチャと、人間による明示的な監視、段階的な出所追跡、およびユーザー フィードバックからのテスト時の学習の適応を統合します。このシステムは、自然言語と視覚化の直接操作の両方を通じて、人間とエージェント間の流動的なハンドオフをサポートし、サンドボックス実行により安全で再現可能なワークフローを保証します。そうすることで、HiLSVA はエージェント的 SciVis を、人間の分析的推論を置き換えるのではなく、強化する共同プロセスとして再構成します。私たちは、代表的なケーススタディと、複数の自律性設定にわたるさまざまな専門知識を持つ 12 人の参加者による対照ユーザー研究を通じて HiLSVA を評価します。結果は、混合イニシアチブの相互作用により、さまざまなレベルのユーザーの専門知識にわたってタスクの完了、ユーザー制御、およびワークフローの透明性が向上する一方、実行効率と人間の監視の間のトレードオフが明らかになったことが示されています。これらの発見は、エージェント的 SciVis における人間中心設計の重要性を強調し、将来の共同視覚化システムの開発の指針となります。 https://hilsva.github.io/ でデモビデオ、ケーススタディ、ソースコードを探索することをお勧めします。

原文 (English)

HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization

Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have essentially prioritized autonomy over human analytical control, thereby limiting transparency and human oversight. We present HiLSVA, a human-in-the-loop agentic system that supports mixed-initiative SciVis workflows. HiLSVA integrates a plan-first multi-agent architecture with explicit human oversight, stepwise provenance tracking, and learn-at-test-time adaptation from user feedback. The system supports fluid handoff between humans and agents through both natural language and direct manipulation of visualizations, while sandboxed execution ensures safe, reproducible workflows. In doing so, HiLSVA reframes agentic SciVis as a collaborative process that augments, rather than replaces, human analytical reasoning. We evaluate HiLSVA through representative case studies and a controlled user study with twelve participants of varying expertise across multiple autonomy settings. Results show that mixed-initiative interaction improves task completion, user control, and workflow transparency across different levels of user expertise, while revealing a tradeoff between execution efficiency and human oversight. These findings highlight the importance of human-centered design in agentic SciVis and guide the development of future collaborative visualization systems. We encourage readers to explore our demo video, case studies, and source code at https://hilsva.github.io/.

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

不確実性定量化の意思決定に沿った評価

機械学習における不確実性の推定は通常、負の対数尤度や予想される校正誤差などの一般的な指標を使用して評価されますが、そのような指標で優れたパフォーマンスが得られたとしても、必ずしも下流の意思決定における有用性が高いことを意味するわけではありません。どの評価指標が下流の公益事業と意味のある形で一致しているかを明らかにする基準である、意思決定の調整を導入します。このフレームワークを適用すると、広く使用されている多くの不確実性指標が、一般的な意思決定問題と一致していないか、下流のタスクに関する病的な事前信念をコード化していることがわかります。次に、事前に重み付けされた効用メトリクスを提案します。これは、意思決定に合わせた不確実性評価を提供する適切なスコアリング ルールの特別なクラスです。ベンチマーク実験と実際のケーススタディ全体にわたって、当社の指標は実現された意思決定の有用性と一貫して一致していますが、従来の指標は一致していません。私たちの結果は、現在の UQ 評価プロトコルの欠陥を明らかにし、意思決定に関連した UQ 評価に向けた既存の指標の原則に基づいた拡張を提供します。

原文 (English)

Decision-Aligned Evaluation of Uncertainty Quantification

Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions. We introduce decision-alignment, a criterion that reveals which evaluation metrics meaningfully align with downstream utilities. Applying this framework, we show that many widely used uncertainty metrics are either misaligned with common decision problems or encode pathological prior beliefs about the downstream task. We then propose prior-weighted utility metrics, a special class of proper scoring rules that provides decision-aligned uncertainty evaluation. Across benchmark experiments and real-world case studies, our metrics consistently align with realized decision utility, while conventional metrics do not. Our results surface flaws in the current UQ evaluation protocol and offer a principled extension of existing metrics toward decision-relevant UQ evaluation.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

継承された回路、学習されたセマンティクス: 微調整が標準評価では見えない回避脆弱性をどのように生み出すか

セキュリティ分類用に微調整された LLM は、通常、トレーニング データと同じ分布からの保持されたサンプルに基づいて評価されます。これにより、微調整自体によって導入された脆弱性が見逃される可能性があることを示します。モデルは、PowerShell のエイリアス置換、コマンドの再構築、文字列の構築、実行の間接化、大文字と小文字の変更などの動作を保持する変換の下では失敗しながらも、正規の精度を維持するトークン レベルのインジケーター セマンティクスを学習できます。私たちは、一致する PowerShell 分類コホートで Foundation-Sec-8B-Instruct とその基本モデルである Llama-3.1-8B-Instruct を研究します。因果的介入により、分類回路は、微調整によって作成されたものではなく、ラマから継承された遅延注意ルートに限定されます。微調整により、この継承された構造が集中して意味的に特殊化され、ベースラインの動作が改善されると同時に、変換に敏感な攻撃対象領域が作成されます。 3 層の回避ベンチマークにより、iwr 置換、Invoke-Expression の再構築、および Llama が共有しない大文字と小文字が変更された Invoke-Expression/IEX バリアントで Foundation-Sec のミスが発見されました。また、デプロイメント前の監視方法も導き出します。分類境界での線形プローブとインジケーター トークンのサイン テストにより、微調整後に標準インジケーターの役割が変わるコマンド ファミリを特定します。これらの信号は、正規入力のみを使用してレッドチームのバリアント生成を優先し、セキュリティの微調整により回避対象領域を拡大しながらタスクの精度を向上できることを示しています。これらの結果は、タスク固有の小さな微調整を単純に安全なセキュリティ分類子として扱うことに対して警告します。特殊化により、継承されたモデル構造が、回避面を拡大しながら保持される精度を維持する脆弱なインジケーター ルールに変換される可能性があります。 AI 対応の堅牢なセキュリティを実現するには、タスクの完全な変換スペースを指定し、微調整を通じてセマンティック ドリフトを監視する必要があります。

原文 (English)

Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation

LLMs fine-tuned for security classification are usually evaluated on held-out examples from the same distribution as their training data. We show that this can miss vulnerabilities introduced by fine-tuning itself: models can learn token-level indicator semantics that preserve canonical accuracy while failing under behavior-preserving transformations such as PowerShell alias substitution, command reconstruction, string construction, execution indirection, and case mutation. We study Foundation-Sec-8B-Instruct and its base model, Llama-3.1-8B-Instruct, on matched PowerShell classification cohorts. Causal interventions localize the classification circuit to a late-attention route inherited from Llama rather than created by fine-tuning. Fine-tuning concentrates and semantically specializes this inherited structure, improving baseline behavior while creating transformation-sensitive attack surfaces. A three-tier evasion benchmark finds Foundation-Sec misses on iwr substitution, Invoke-Expression reconstruction, and case-mutated Invoke-Expression/IEX variants that Llama does not share. We also derive a pre-deployment monitoring method: a linear probe at the classification boundary and an indicator-token sign test identify command families where canonical indicators change role after fine-tuning. These signals prioritize red-team variant generation using only canonical inputs, showing that security fine-tuning can improve task accuracy while expanding the evasion surface. These results caution against treating small task-specific fine-tunes as straightforwardly safer security classifiers: specialization can convert inherited model structure into brittle indicator rules that preserve held-out accuracy while expanding the evasion surface. Robust AI-enabled security will require specifying the full transformation space of the task and monitoring semantic drift through fine-tuning.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts

HIPE-2026 is a CLEF evaluation lab dedicated to person-place relation extraction from noisy, multilingual historical texts. Building on the…

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Autodata: An agentic data scientist to create high quality synthetic data

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation dat…

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

Power Couple? AI Growth and Renewable Energy Investment

AI and renewable energy are increasingly framed as a "power couple," on the premise that surging AI demand will accelerate clean-energy inv…

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

The Augmentation Trap: AI Productivity and the Cost of Cognitive Offloading

Experimental evidence suggests that AI tools raise worker productivity, but also that sustained use can erode the expertise on which those…

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning

Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

マルチモーダル LLM 評価に欠けているものは何ですか?

マルチモーダル大規模言語モデル (MLLM) は、テキスト、画像、音声、ビデオなどのさまざまな入力を処理し、テキスト応答を生成できます。それらの機能は急速に進歩していますが、そのようなモデルの評価は追いついていません。既存の評価ベンチマークのほとんどは、個別のタスクに限定されており、モデルがモダリティ全体で情報を統合しているかどうかについてはほとんど明らかにされていません。私たちは、MLLM を評価するための現在の手段を調査し、既存のベンチマーク分類をレビューして、時間空間的一貫性、物理世界の理解、マルチモーダル一貫性、選択的注意などのギャップを特定します。これらのギャップに対処することは、マルチモーダル インテリジェンスの実際の進歩を測定し、機能の境界を明らかにするために不可欠です。

原文 (English)

What We are Missing in Multimodal LLM Evaluation?

Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced rapidly, evaluation of such models has not kept pace. Most existing evaluation benchmarks are limited to isolated tasks and reveal little about whether a model integrates information across modalities. We examine current means for evaluating MLLMs and review the existing benchmark taxonomy to identify gaps, including temporal-spatial coherence, physical world understanding, multimodal consistency, and selective attention. Addressing these gaps is essential for measuring real progress in multimodal intelligence and exposing capability boundaries.

2026-06-26 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

OpenFinGym: クオンツエージェントを評価するための検証可能なマルチタスクジム環境

大規模な言語モデル エージェントは定量的財務ワークフローにますます適用されていますが、その評価は分離されたタスク間で断片化されたままであり、ベンチマーク タスクの財務関連性はしばしば見落とされます。しかし、財務ワークフローは本質的に多段階であり、予測、戦略構築、リスク管理、取引などの相互依存するタスクにまたがっています。既存のプラットフォームは通常、単一のタスクに焦点を当てているため、エージェントの能力を過大評価し、一般化、実際の市場でのやり取り、および財務的に意味のある意思決定における弱点を明らかにできません。 OpenFinGym は、単一の実行および検証インターフェイスで予測、市場生成、リ​​アルタイム取引、不正検出をカバーする定量的金融エージェント開発用の統合ジム環境です。 OpenFinGym はさらに、定量的な財務出版物を実行可能なタスク パッケージに変換する自動タスク構築パイプラインを提供します。スケーラブルなエージェントのロールアウトをサポートし、ランタイムのトレインテストの漏洩を防ぐホスト側検証サービスを備えたコンテナ化されたランタイム。低レイテンシのデータストリーム設計を備えたペーパートレーディングエンジン。長期およびイベント市場の予測に対する遅延解像度のサポート。トレーニング後の SFT と RL の統合

原文 (English)

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tasks, while the financial relevance of benchmark tasks is often overlooked. Yet financial workflows are inherently multi-stage, spanning interdependent tasks such as forecasting, strategy construction, risk management, and trading. Existing platforms typically focus on a single task, and can therefore overstate agent competence and fail to reveal weaknesses in generalization, real-market interaction, and financially meaningful decision-making. We introduce OpenFinGym, a unified gym environment for quantitative-finance agent development that covers forecasting, market generation, real-time trading, and fraud detection under a single execution and verification interface. OpenFinGym additionally provides an automated task-construction pipeline that turns quantitative finance publications into executable task packages; a containerised runtime with a host-side verifier service that supports scalable agent rollouts and prevents runtime train-test leakage; a paper trading engine with a low-latency data-stream design; deferred-resolution support for long-horizon and event-market forecasts; and integration for SFT and RL post-training

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

大規模言語モデルを使用して縦断的な合成臨床ノートを生成するパイプライン

実世界のデータへのアクセスが制限されている領域で AI システムの開発と評価を可能にするために、合成データの使用が増えています。医療分野では、臨床文書はその機密性により特別な課題を抱えています。この研究では、実際の患者データに伴うプライバシー リスクを回避しながら、臨床 AI ツールの開発をサポートするように設計された合成臨床メモ パイプラインとデータセットを導入します。データセットは、大規模な言語モデルを使用した構造化患者生成、半構造化患者ジャーニー シミュレーション、および非構造化臨床ノート生成を組み合わせたモジュール式パイプラインを使用して生成されます。このパイプラインは、長期的な患者記録全体にわたる内部一貫性を優先すると同時に、書き方、メモの構造、臨床の詳細の変化も捕捉するように設計されています。 LLM ベースの検証および拡張ステップを含む追加のメカニズムを使用して、生成されたノートの忠実性、リアリズム、および多様性が向上します。私たちは、70 人の合成患者のデータセットをリリースします。各患者には、入院期間全体にわたる 20 ~ 50 の臨床ノートが関連付けられています。データセットは複数の検証レベルで提供されているため、ユーザーはユースケースに応じて現実性とスケーラビリティのバランスを取ることができます。このデータセットは、実際の患者データに依存することなく、要約ツール、コーディング モデル、意思決定支援システムなどの臨床 AI システムの開発、テスト、評価をサポートします。

原文 (English)

A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models

Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

グラウンドトゥルースを使用してクラスタリングを評価するにはどうすればよいですか?

グランド トゥルースが利用可能な場合、外部インデックスをクラスター評価に使用できます。セットマッチングベースの尺度に焦点を当てて、最も一般的な外部妥当性指標をレビューします。セントロイド インデックス (CI) は、説明可能な結果が得られる直感的なクラスター レベルの測定であるため、推奨します。より細かく調整されたポイントレベルの測定が必要な場合は、より多くの選択肢があります。ペアセット インデックス (PSI) は、クラスター サイズによって偏らない正規化されたスコアを提供します。すべてのポイントが同等に重要である必要がある場合は、クラスタリング精度 (ACC) またはその他のセットマッチング尺度が適しています。

原文 (English)

How to evaluate clustering with ground truth?

External indexes can be used for cluster evaluation when ground truth is available. We review the most common external validity indexes focusing on set-matching-based measures. We recommend centroid index (CI), because it is an intuitive cluster-level measure with an explainable result. If we need a more fine-tuned, point-level measure, there are more choices. Pair-set index (PSI) provides a normalized score which is not biased by cluster sizes. If all points should matter equally, then clustering accuracy (ACC) or any other set-matching measure is suitable.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

判断せずに質問する: 解釈可能な LLM 評価と自己改善のための 2 つの質問

NLP では、LLM 出力の評価が依然として大きなボトルネックとなっています。人間による評価は高価で時間がかかり、語彙メトリクスとオープンエンド生成に関する人間の判断との相関性が低く、全体的な LLM ジャッジはデバッグが難しい不透明なスコアを生成することがよくあります。私たちは、評価基準をアトミックなバイナリの質問に分解し、その結果の判定を解釈可能な多次元スコアに集約するフレームワークである BINEVAL を提案します。タスク プロンプトが与えられると、メタ プロンプトが詳細な評価質問を生成し、LLM が出力ごとに独立して質問に回答し、調整された全体スコアとともに透明な質問レベルのフィードバックを生成します。この分解により、評価が検査、​​診断が容易になり、迅速な改善に直接使用できるようになります。 SummEval、Topical-Chat、QAGS 全体で、BINEVAL は UniEval や G-Eval などの強力なベースラインと同等またはそれを上回り、特に QAGS などの事実整合性ベンチマークで優れた結果を示しています。 BINEVAL は、人間の判断との競合相関を超えて、人間のスコア分布とよりよく一致し、以前の LLM ジャッジによく見られた天井効果を回避し、境界線にある出力と明らかに欠陥のある出力をより適切に区別することにつながります。さらに、同じ質問レベルのフィードバックが反復プロンプトの最適化をサポートし、自己更新設定とクロスモデル更新設定の両方で IFBench での要約に関する評価者のプロンプトと生成プロンプトを改善することを示します。全体として、BINEVAL は、強力な経験的パフォーマンスと実用的な診断および最適化の価値を組み合わせた、タスクに依存せず、トレーニング不要で、解釈可能な評価フレームワークを提供します。

原文 (English)

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that decomposes evaluation criteria into atomic binary questions and aggregates the resulting verdicts into interpretable, multi-dimensional scores. Given a task prompt, a meta-prompt generates fine-grained evaluation questions, and an LLM answers them independently for each output, yielding transparent question-level feedback together with calibrated overall scores. This decomposition makes evaluation easier to inspect, easier to diagnose, and directly usable for prompt improvement. Across SummEval, Topical-Chat, and QAGS, BINEVAL matches or outperforms strong baselines including UniEval and G-Eval, with especially strong results on factual consistency benchmarks such as QAGS. Beyond competitive correlation with human judgments, BINEVAL better matches human score distributions and avoids the ceiling effects common in prior LLM judges, leading to better discrimination between borderline and clearly flawed outputs. We further show that the same question-level feedback supports iterative prompt optimization, improving evaluator prompts on summarization and generation prompts on IFBench under both self-update and cross-model update settings. Overall, BINEVAL provides a task-agnostic, training-free, and interpretable evaluation framework that combines strong empirical performance with practical diagnostic and optimization value.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Know2Guess: 大規模言語モデルにおける知識境界評価のための汚染を認識したマルチゾーン ベンチマーク

大規模な言語モデルの信頼性の高い評価では、データの汚染、プロンプトの特異性、または一般的な拒否行動と混同することなく、サポートされている回答とサポートされていない推測を分離する必要があります。凍結されたビルドタイム ラベルの下で、回答可能な知識から棄権が期待される未知への移行を測定するための、汚染を認識したマルチゾーン ベンチマークを示します。このベンチマークには、5 つのドメインにわたる 1,200 項目、明示的な棄権期待、汚染リスクのメタデータ、および公式の厳密なパーサーと正規化された堅牢性パーサーによる二重解析が含まれています。ロックされた回答または棄権プロンプト、回答のみのコントロール、およびプロンプト テンプレートのバリアントの下で、FLAN-T5、Qwen2.5-Instruct、および Llama-3-Instruct モデルを評価します。このベンチマークは、一般的な無回答行動では解決されません。FLAN のベースラインは、生産的な棄権に関しては弱いままですが、より強力な指導調整モデルは、選択的ではあるが回答から棄権への移行が不完全であることを明らかにしています。 Qwen2.5-3B-Instruct は全体的に最高の信頼性を実現していますが、回答が期待されるゾーンは依然として難しく、キャリブレーションは依然として不十分で、良性の項目の拒否は引き続き発生します。プロンプトおよびパーサーの堅牢性分析により、主要なランキングと定性的な結論が維持されます。したがって、このベンチマークは、回答可能性、棄権、拒否、および汚染を、LLM の信頼性の個別だが相互作用する側面として監査するための再現可能なプロトコルを提供します。データセットは、https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark で公開されています。

原文 (English)

Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels. The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict parser plus a normalized robustness parser. We evaluate FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models under locked answer-or-abstain prompts, answer-only controls, and prompt-template variants. The benchmark is not solved by generic non-answer behavior: FLAN baselines remain weak on productive abstention, while stronger instruction-tuned models expose a selective but incomplete transition from answering to abstaining. Qwen2.5-3B-Instruct achieves the best overall reliability, but answer-expected zones remain difficult, calibration remains poor, and benign-item refusal persists. Prompt and parser robustness analyses preserve the main ranking and qualitative conclusions. The benchmark therefore provides a reproducible protocol for auditing answerability, abstention, refusal, and contamination as distinct but interacting dimensions of LLM reliability.The dataset is publicly available at https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

LLM エージェントの即時注入に対する帯域外防御の適応的評価

最近の研究 (2024 年から 2026 年) は、ツールを使用する LLM エージェントを間接的なプロンプト インジェクションから防御するための戦略に収束しました。悪意のある命令を拒否するようにモデルをトレーニングするのではなく、エージェントのアクションを仲介する決定論的なポリシーでモデルの外側にセキュリティを強化します。 CaMeL、FIDES、Progent、RTBAS、FORGE などのシステムは、機能、情報フロー ラベル、参照モニターによってこれを実現しており、いくつかのシステムは AgentDojo ベンチマークに対する攻撃がほぼ排除されたと報告しています。私たちは 2 つの貢献をします。まず、これらの帯域外防御を古典的な完全性保護 (Biba)、参照モニタリング、最小特権のインスタンスとして整理し、それらがカバーするものとカバーしないものを構造化して比較します。第 2 に、それらのすべては静的ベンチマーク (一連の固定された注入試行) でのみ検証されることを警告します。これは、適応的で防御を意識した攻撃が 90% 以上の成功率でそのうち 12 件を破るまで、帯域内防御が強力であるように見えたのと同じ方法論です。適応的評価に必要な脅威モデルとプロトコルを指定します。次に、そのプロトコルを、Progent 独自の適応型攻撃分析の独立した再現および拡張として、単一の H200 上で自己ホストされるオープンウェイト エージェント (Qwen2.5-7B) を使用して AgentDojo 上で実行します。この設定は作成者がテストしていません。 3 回のランを平均すると、ディフェンスは次の結果を維持しました。プロジェントは平均攻撃成功率を約 6 倍 (25.8% から 4.2%) に削減しましたが、手作りの適応攻撃では平均攻撃成功率は上がりませんでした (2.6%)。これは、単一のブラックボックス攻撃テンプレートを備えた弱いモデル上の 1 つの小規模なデータ ポイントです。より強力な最適化された (ホワイトボックス GCG) 攻撃が依然として存在します。この結果は、適応型攻撃者にとって、決定論的な帯域外強制の方が帯域内検出よりも難しいターゲットであるという仮説と一致しますが、それを確立するものではありません。

原文 (English)

Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions. Systems such as CaMeL, FIDES, Progent, RTBAS, and FORGE realize this with capabilities, information-flow labels, and reference monitors, and several report near-elimination of attacks on the AgentDojo benchmark. We make two contributions. First, we organize these out-of-band defenses as instances of classical integrity protection (Biba), reference monitoring, and least privilege, yielding a structured comparison of what they do and do not cover. Second, we warn that every one of them is validated only on static benchmarks (a fixed set of injection attempts), the same methodology that made in-band defenses look strong until adaptive, defense-aware attacks broke twelve of them at over 90% success; we specify the threat model and protocol an adaptive evaluation requires. We then run that protocol as an independent reproduction and extension of Progent's own adaptive-attack analysis, on AgentDojo, with an open-weight agent (Qwen2.5-7B) self-hosted on a single H200, a setting its authors did not test. Averaged over three runs, the defense held: Progent cut mean attack success roughly sixfold (25.8% to 4.2%), and a hand-crafted adaptive attack did not raise it (2.6%). This is one small-scale data point on a weak model with a single black-box attack template; a stronger optimized (white-box GCG) attack remains open. The result is consistent with, but does not establish, the hypothesis that deterministic out-of-band enforcement is a harder target for an adaptive attacker than in-band detection.

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

深層学習プログラムの故障診断における評価と戦略のギャップ

深層学習 (DL) プログラムはさまざまな理由でトレーニング中に失敗する可能性があり、原因の診断はコ​​ストと時間のかかるメンテナンス作業です。このような障害を診断する技術は、通常、プログラム内の相互検証を使用して評価されますが、これまでに確認されていないプログラムが関与する展開設定には不適切な場合があります。したがって、これらの設定間でパフォーマンスがどのように異なるかを評価し、確立された DL の障害診断技術におけるパフォーマンス ギャップの原因を特定する必要があります。私たちは、38 の実世界の DL プログラムからの 5,542 個のフォールト注入トレーニング トレースのコーパスである DynFault を使用して、このギャップを調査します。既存の障害診断技術では、プログラム内での評価とプログラム全体を実行した場合のバランスの取れた精度に 0.190 のギャップがあることがわかりました。また、このギャップは機能のプログラム レベルの構造に起因することもわかり、2 つのランタイム機能セット、曲率機能とオプティマイザー機能、および目に見えないプログラムでのそれらの動作を調査することになりました。曲率機能は目に見えないプログラムの不安定性の検出に役立ちますが、オプティマイザーとアクティベーション機能はトレーニング中に表示されたプログラムにのみ役立ちます。

原文 (English)

Evaluation-Strategy Gap in Fault Diagnosis of Deep Learning Programs

Deep Learning (DL) programs can fail during training for many reasons, and diagnosing the cause is a costly and time-consuming maintenance task. Techniques for diagnosing such failures are commonly assessed using within-program cross-validation, which may be inadequate for deployment settings involving previously unseen programs. It is therefore necessary to assess how performance differs across these settings and to identify the causes of any performance gap in established fault diagnosis techniques for DL. We investigate this gap using DynFault, a corpus of 5,542 fault-injected training traces from 38 real-world DL programs. We found a gap of 0.190 in balanced accuracy for existing fault diagnosis techniques between within-program evaluation and holding out whole programs. We also found the gap comes from program-level structure in the features, which led us to examine two runtime feature sets, curvature features and optimizer features, and their behavior on unseen programs. We found that curvature features are useful for instability detection on unseen programs, while optimizer and activation features help only on programs seen during training.

2026-06-26 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

幻覚からグラウンディングまで: CRISP による視覚空間知能の診断

現在の VLM 評価では、言語事前分布と真の空間推論が混同されることがよくあります。これに対処するために、CRISP を導入します。CRISP は、一貫性、つまり暗黙の認識と明示的な推論の整合性を通じて視覚空間知能を評価する、新しい構造診断評価パラダイムです。従来のブラックボックス QA とは異なり、CRISP はメトリック 3D シーン グラフとオラクル介入プロトコルを利用して、潜在的な推論機能を知覚のボトルネックから切り離します。この詳細な診断により、体系的な知覚と推論の断絶が明らかになります。重要なことに、独自のモデルは強力な潜在的な推論エンジンを備えているものの、不正確な計量推定と暗黙の構造表現を活用する重大な失敗に悩まされていることを明らかにしました。逆に、オープンソース モデルは、マルチホップの構成推論が欠如しているため、依然として根本的なボトルネックとなっています。 CRISP は、事前言語を介して単に「正しく推測する」ことから、真に「知覚、検証、推論する」ことに焦点を移すことで、エンドツーエンドのポストトレーニングを超えたマルチモーダル調整のための厳密なロードマップを提供します。コードとデータセットは https://github.com/iiyamayuki/CRISP-Bench で入手できます。

原文 (English)

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP

Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning. Unlike traditional black-box QA, CRISP utilizes metric 3D Scene Graphs and an oracle intervention protocol to decouple latent reasoning capabilities from perceptual bottlenecks. This granular diagnosis uncovers a systematic perception-reasoning disconnect. Crucially, we reveal that while proprietary models possess robust latent reasoning engines, they suffer from inaccurate metric estimation and a critical failure to leverage their implicit structural representations. Conversely, open-source models remain fundamentally bottlenecked by their lack of multi-hop compositional reasoning. By shifting the focus from merely ``guessing correctly'' via language priors to genuinely ``perceiving, verifying, and reasoning,'' CRISP offers a rigorous roadmap for multimodal alignment beyond end-to-end post-training. The code and dataset are available at https://github.com/iiyamayuki/CRISP-Bench.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

HiLSVA: 科学的視覚化のための人間参加型エージェント システムの設計と評価

大規模言語モデル (LLM) エージェントは、科学的視覚化 (SciVis) のための自然言語対話を可能にします。それでも、従来のシステムは基本的に人間による分析制御よりも自律性を優先しており、そのため透明性と人間による監視が制限されていました。混合イニシアチブの SciVis ワークフローをサポートする人間参加型エージェント システムである HiLSVA を紹介します。 HiLSVA は、計画優先のマルチエージェント アーキテクチャと、人間による明示的な監視、段階的な出所追跡、およびユーザー フィードバックからのテスト時の学習の適応を統合します。このシステムは、自然言語と視覚化の直接操作の両方を通じて、人間とエージェント間の流動的なハンドオフをサポートし、サンドボックス実行により安全で再現可能なワークフローを保証します。そうすることで、HiLSVA はエージェント的 SciVis を、人間の分析的推論を置き換えるのではなく、強化する共同プロセスとして再構成します。私たちは、代表的なケーススタディと、複数の自律性設定にわたるさまざまな専門知識を持つ 12 人の参加者による対照ユーザー研究を通じて HiLSVA を評価します。結果は、混合イニシアチブの相互作用により、さまざまなレベルのユーザーの専門知識にわたってタスクの完了、ユーザー制御、およびワークフローの透明性が向上する一方、実行効率と人間の監視の間のトレードオフが明らかになったことが示されています。これらの発見は、エージェント的 SciVis における人間中心設計の重要性を強調し、将来の共同視覚化システムの開発の指針となります。 https://hilsva.github.io/ でデモビデオ、ケーススタディ、ソースコードを探索することをお勧めします。

原文 (English)

HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization

Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have essentially prioritized autonomy over human analytical control, thereby limiting transparency and human oversight. We present HiLSVA, a human-in-the-loop agentic system that supports mixed-initiative SciVis workflows. HiLSVA integrates a plan-first multi-agent architecture with explicit human oversight, stepwise provenance tracking, and learn-at-test-time adaptation from user feedback. The system supports fluid handoff between humans and agents through both natural language and direct manipulation of visualizations, while sandboxed execution ensures safe, reproducible workflows. In doing so, HiLSVA reframes agentic SciVis as a collaborative process that augments, rather than replaces, human analytical reasoning. We evaluate HiLSVA through representative case studies and a controlled user study with twelve participants of varying expertise across multiple autonomy settings. Results show that mixed-initiative interaction improves task completion, user control, and workflow transparency across different levels of user expertise, while revealing a tradeoff between execution efficiency and human oversight. These findings highlight the importance of human-centered design in agentic SciVis and guide the development of future collaborative visualization systems. We encourage readers to explore our demo video, case studies, and source code at https://hilsva.github.io/.

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

不確実性定量化の意思決定に沿った評価

機械学習における不確実性の推定は通常、負の対数尤度や予想される校正誤差などの一般的な指標を使用して評価されますが、そのような指標で優れたパフォーマンスが得られたとしても、必ずしも下流の意思決定における有用性が高いことを意味するわけではありません。どの評価指標が下流の公益事業と意味のある形で一致しているかを明らかにする基準である、意思決定の調整を導入します。このフレームワークを適用すると、広く使用されている多くの不確実性指標が、一般的な意思決定問題と一致していないか、下流のタスクに関する病的な事前信念をコード化していることがわかります。次に、事前に重み付けされた効用メトリクスを提案します。これは、意思決定に合わせた不確実性評価を提供する適切なスコアリング ルールの特別なクラスです。ベンチマーク実験と実際のケーススタディ全体にわたって、当社の指標は実現された意思決定の有用性と一貫して一致していますが、従来の指標は一致していません。私たちの結果は、現在の UQ 評価プロトコルの欠陥を明らかにし、意思決定に関連した UQ 評価に向けた既存の指標の原則に基づいた拡張を提供します。

原文 (English)

Decision-Aligned Evaluation of Uncertainty Quantification

Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions. We introduce decision-alignment, a criterion that reveals which evaluation metrics meaningfully align with downstream utilities. Applying this framework, we show that many widely used uncertainty metrics are either misaligned with common decision problems or encode pathological prior beliefs about the downstream task. We then propose prior-weighted utility metrics, a special class of proper scoring rules that provides decision-aligned uncertainty evaluation. Across benchmark experiments and real-world case studies, our metrics consistently align with realized decision utility, while conventional metrics do not. Our results surface flaws in the current UQ evaluation protocol and offer a principled extension of existing metrics toward decision-relevant UQ evaluation.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

継承された回路、学習されたセマンティクス: 微調整が標準評価では見えない回避脆弱性をどのように生み出すか

セキュリティ分類用に微調整された LLM は、通常、トレーニング データと同じ分布からの保持されたサンプルに基づいて評価されます。これにより、微調整自体によって導入された脆弱性が見逃される可能性があることを示します。モデルは、PowerShell のエイリアス置換、コマンドの再構築、文字列の構築、実行の間接化、大文字と小文字の変更などの動作を保持する変換の下では失敗しながらも、正規の精度を維持するトークン レベルのインジケーター セマンティクスを学習できます。私たちは、一致する PowerShell 分類コホートで Foundation-Sec-8B-Instruct とその基本モデルである Llama-3.1-8B-Instruct を研究します。因果的介入により、分類回路は、微調整によって作成されたものではなく、ラマから継承された遅延注意ルートに限定されます。微調整により、この継承された構造が集中して意味的に特殊化され、ベースラインの動作が改善されると同時に、変換に敏感な攻撃対象領域が作成されます。 3 層の回避ベンチマークにより、iwr 置換、Invoke-Expression の再構築、および Llama が共有しない大文字と小文字が変更された Invoke-Expression/IEX バリアントで Foundation-Sec のミスが発見されました。また、デプロイメント前の監視方法も導き出します。分類境界での線形プローブとインジケーター トークンのサイン テストにより、微調整後に標準インジケーターの役割が変わるコマンド ファミリを特定します。これらの信号は、正規入力のみを使用してレッドチームのバリアント生成を優先し、セキュリティの微調整により回避対象領域を拡大しながらタスクの精度を向上できることを示しています。これらの結果は、タスク固有の小さな微調整を単純に安全なセキュリティ分類子として扱うことに対して警告します。特殊化により、継承されたモデル構造が、回避面を拡大しながら保持される精度を維持する脆弱なインジケーター ルールに変換される可能性があります。 AI 対応の堅牢なセキュリティを実現するには、タスクの完全な変換スペースを指定し、微調整を通じてセマンティック ドリフトを監視する必要があります。

原文 (English)

Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation

LLMs fine-tuned for security classification are usually evaluated on held-out examples from the same distribution as their training data. We show that this can miss vulnerabilities introduced by fine-tuning itself: models can learn token-level indicator semantics that preserve canonical accuracy while failing under behavior-preserving transformations such as PowerShell alias substitution, command reconstruction, string construction, execution indirection, and case mutation. We study Foundation-Sec-8B-Instruct and its base model, Llama-3.1-8B-Instruct, on matched PowerShell classification cohorts. Causal interventions localize the classification circuit to a late-attention route inherited from Llama rather than created by fine-tuning. Fine-tuning concentrates and semantically specializes this inherited structure, improving baseline behavior while creating transformation-sensitive attack surfaces. A three-tier evasion benchmark finds Foundation-Sec misses on iwr substitution, Invoke-Expression reconstruction, and case-mutated Invoke-Expression/IEX variants that Llama does not share. We also derive a pre-deployment monitoring method: a linear probe at the classification boundary and an indicator-token sign test identify command families where canonical indicators change role after fine-tuning. These signals prioritize red-team variant generation using only canonical inputs, showing that security fine-tuning can improve task accuracy while expanding the evasion surface. These results caution against treating small task-specific fine-tunes as straightforwardly safer security classifiers: specialization can convert inherited model structure into brittle indicator rules that preserve held-out accuracy while expanding the evasion surface. Robust AI-enabled security will require specifying the full transformation space of the task and monitoring semantic drift through fine-tuning.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts

HIPE-2026 is a CLEF evaluation lab dedicated to person-place relation extraction from noisy, multilingual historical texts. Building on the…

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Autodata: An agentic data scientist to create high quality synthetic data

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation dat…

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

Power Couple? AI Growth and Renewable Energy Investment

AI and renewable energy are increasingly framed as a "power couple," on the premise that surging AI demand will accelerate clean-energy inv…

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

The Augmentation Trap: AI Productivity and the Cost of Cognitive Offloading

Experimental evidence suggests that AI tools raise worker productivity, but also that sustained use can erode the expertise on which those…

2026-06-26 13:00 JSTarXiv cs.AIビジネス/資金調達

SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning

Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…

2026-06-26 11:00 JSTITmedia AI+ビジネス/資金調達

企業のAI支出そろそろ“様子見”は終わり? Gartner予測、本格投資の行方は

Gartnerは2026年の世界AI支出が前年比47%増の2兆5956億ドルに達するとの予測を発表。2026年は企業によるAI支出が拡大局面へ移行する転換点になるとしている。

2026-06-26 01:55 JSTTechCrunch AIエージェントビジネス/資金調達

General Intuition’s $2.3B bet that video games can train AI agents for the real world

General Intuition has raised $320 million to scale AI trained on millions of hours of gameplay, betting action data can help AI develop som…

2026-06-25 23:55 JSTTechCrunch AIビジネス/資金調達

Netris raises $15M Series A from a16z to help AI neoclouds go live faster

Netris provides software that runs on network switches, and offers a platform that helps neocloud operators reduce the time it takes to go…

2026-06-25 21:00 JSTTechCrunch AIビジネス/資金調達

Amazon ups India bet with fresh $13B AI infrastructure investment

Amazon’s latest India investment comes as global tech companies race to expand AI infrastructure in the country.

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

T2D ベンチ: 多層臨床ライフスタイル ナレッジ グラフを使用した 2 型糖尿病の LLM 出力の証拠ゲート型評価

大規模言語モデル (LLM) は、2 型糖尿病に対する臨床的に流暢な推奨事項を生成できますが、ガイドラインの制約を満たしたり、ライフスタイルに関連した血糖の主張を明確に正当化したりすることはできません。我々は、LLM 出力が明示的でグラフチェック可能な証拠要件を満たしているかどうかをテストするための、再現可能なベンチマークおよび証拠ゲート型評価フレームワークである T2D-Bench を紹介します。 T2D-Bench は、生体医学スパイン (UMLS、DrugBank、SIDER)、計算可能な ADA 治療標準ルール、血糖検査室効果への機構的なブリッジを介して接続されたライフスタイル知識を組み合わせた、多層の臨床ライフスタイル ナレッジ グラフに基づいて構築されています。診断、投薬の安全性、敵対的なライフスタイルの衝突にわたる 100 の構造化されたビネット全体で、ベースライン出力は、GPT-4o-mini のケースの 35%、GPT-4o のケースの 33% で、ベンチマークで定義されたエビデンスパス チェックに失敗しました。証拠ゲートはサポートされていない省略を検出し、制約付きリビジョンを使用して、出力をベンチマークで定義された証拠要件に検証者レベルで準拠させます。これらの結果は、糖尿病に焦点を当てた LLM 出力において、計算可能な証拠の制約により、裏付けのない臨床上の省略が明示的、測定可能、修正可能になる可能性があることを示しています。

原文 (English)

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims. We present T2D-Bench, a reproducible benchmark and evidence-gated evaluation framework for testing whether LLM outputs satisfy explicit, graph-checkable evidence requirements. T2D-Bench is built on a multi-layer clinical-lifestyle knowledge graph that combines a biomedical spine (UMLS, DrugBank, SIDER), computable ADA Standards of Care rules, and lifestyle knowledge connected through a mechanistic bridge to glycemic laboratory effects. Across 100 structured vignettes spanning diagnosis, medication safety, and adversarial lifestyle conflicts, baseline outputs failed benchmark-defined evidence-path checks in 35% of cases for GPT-4o-mini and 33% for GPT-4o. The evidence gate detects unsupported omissions and uses constrained revision to bring outputs into verifier-level compliance with benchmark-defined evidence requirements. These results show that computable evidence constraints can make unsupported clinical omissions explicit, measurable, and correctable in diabetes-focused LLM outputs.

2026-06-25 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

自律的な評価を備えたコンピュータ使用エージェントの強化学習

Computer-Use Agent (CUA) は、グラフィカル ユーザー インターフェイス内で直接認識して行動することで、高レベルのユーザー目標を実行します。ただし、オープンエンドのデスクトップ環境ではスケーラブルで機械可読な報酬信号がほとんど提供されないため、CUA の強化学習は依然として困難です。タスクの成功は多くの場合視覚的に根拠があり、手作りの報酬関数や高密度の手動ラベルで指定するのは困難です。我々は、GUI エージェントのスケーラブルな監視信号として自律的な視覚言語評価を使用する RL 微調整フレームワークを提案します。最終的なスクリーンショットと元の指示が与えられると、ビジョン言語モデルはタスクの完了を判断し、ポリシーの最適化中にタスク固有のヒューリスティックや手動ラベルを使用せずに最終的なフィードバックを提供します。自律型評価器は不完全であるため、そのフィードバックをノイズの多いバイナリ報酬チャネルとしてモデル化し、近接ポリシー最適化のためのノイズ補正された報酬推定器を導出します。 macOSWorld、Windows Agent Arena、および OSWorld にわたる実験では、修正された評価者の報酬がゼロショットのベースラインと生の評価者の報酬の両方を上回り、成功率がゼロショットのパフォーマンスより平均 12.6 ポイント、生の評価者の微調整よりも 5.1 ポイント向上したことが示されています。これらの結果は、評価者のノイズが明示的にモデル化され補正されている場合、自律評価が GUI 環境における RL の実用的な報酬信号として機能する可能性があることを示唆しています。

原文 (English)

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation

Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environments rarely provide scalable, machine-readable reward signals: task success is often visually grounded and hard to specify with handcrafted reward functions or dense manual labels. We propose an RL fine-tuning framework that uses autonomous vision-language evaluation as a scalable supervision signal for GUI agents. Given a final screenshot and the original instruction, a Vision-Language Model judges task completion and provides terminal feedback without task-specific heuristics or manual labels during policy optimization. Because autonomous evaluators are imperfect, we model their feedback as a noisy binary reward channel and derive a noise-corrected reward estimator for Proximal Policy Optimization. Experiments across macOSWorld, Windows Agent Arena, and OSWorld show that corrected evaluator rewards outperform both zero-shot baselines and raw evaluator rewards, improving success rates by an average of 12.6 percentage points over zero-shot performance and 5.1 points over raw evaluator fine-tuning. These results suggest that autonomous evaluation can serve as a practical reward signal for RL in GUI environments when evaluator noise is explicitly modeled and corrected.

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AdversaBench: 複数の裁判官による確認とモデル間の移行性を備えた自動化された LLM レッドチーム化

大規模な言語モデルの敵対的評価をスケーリングするには、ハード入力を生成する方法と、結果として生じる失敗が本物であることを確認する信頼性の高い方法の両方が必要です。 AdversaBench は、5 つの構造化された演算子でシード プロンプトを変更し、ターゲット モデルをクエリし、メタ ジャッジ タイブレーカーを備えた 3 人のジャッジ パネルを通じて失敗を確認する、エンドツーエンドのレッド チーム パイプラインです。推論、指示に従い、ツールの使用という 3 つのカテゴリにわたる 45 のシードに関する実験を報告します。すべてのシードで失敗が確認されました。 4 つの発見が際立っています。まず、オペレーターの有効性はカテゴリによって大きく異なります。inject_distractor のスコアは、指示に従うシードでは 0.00 の平均報酬ですが、推論とツールの使用では 0.80 ~ 0.83 です。第 2 に、バイナリの失敗率が難しさを隠しています。命令に従うシードでは、攻撃者の反復回数が平均 2.4 回であるのに対し、他のカテゴリでは 1.1 回であり、生存曲線にギャップが見られます。第三に、80 ~ 87% というペアごとのジャッジの一致は、ラベルの歪みによりほぼゼロのコーエンのカッパと共存します。カテゴリレベルの不一致率の方が有益です。第 4 に、Llama 3.1 8B に対して生成された敵対的プロンプトはゼロショットを Llama 3.3 70B に転送します。これは、変異がモデル固有の弱点ではなく一般的な動作パターンを悪用していることを示唆しています。コード、データセット、分析スクリプトは https://github.com/khanak0509/AdversaBench で入手できます。

原文 (English)

AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability

Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm that resulting failures are real. We present AdversaBench, an end-to-end red-teaming pipeline that mutates seed prompts with five structured operators, queries a target model, and confirms failures through a three-judge panel with a meta-judge tiebreaker. We report experiments on 45 seeds across three categories: reasoning, instruction-following, and tool use. Every seed produced a confirmed failure. Four findings stand out. First, operator effectiveness varies sharply by category: inject_distractor scores 0.00 mean reward on instruction-following seeds but 0.80-0.83 on reasoning and tool-use. Second, binary failure rate hides difficulty: instruction-following seeds required 2.4 attacker iterations on average versus 1.1 for other categories, a gap visible in survival curves. Third, pairwise judge agreement of 80-87% coexists with near-zero Cohen's kappa due to label skew; category-level disagreement rates are more informative. Fourth, adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B, suggesting the mutations exploit general behavioral patterns rather than model-specific weaknesses. Code, dataset, and analysis scripts are available at https://github.com/khanak0509/AdversaBench .

2026-06-25 13:00 JSTarXiv cs.AIビジネス/資金調達

確率的ブール関数評価のためのコスト最適決定図

多くの意思決定シナリオでは、情報の取得にさまざまなコストがかかります。変動コストおよび真理の割り当てに対する確率分布の下で命題式を評価する際の期待コストを最小限に抑える決定論的評価戦略を構築する問題を検討します。変数選択ヒューリスティック、枝刈り、およびキャッシュを備えた分岐限定アルゴリズムを紹介します。私たちの知る限り、これはこのレベルの一般性を実現する最初の実用的な正確なアルゴリズムです。ランダム インスタンスの実験では、スケーラビリティを実証し、貪欲なビーム検索バリアントの効率と品質のトレードオフを定量化します。さらに、構造化された心臓病の診断例も評価します。最後に、問題が $\#P$ 困難であり、$\mathrm{PSPACE}$ に含まれていることを証明します。

原文 (English)

Cost-Optimal Decision Diagrams for Stochastic Boolean Function Evaluation

In many decision-making scenarios, acquiring information incurs different costs. We consider the problem of constructing a deterministic evaluation strategy that minimizes the expected cost of evaluating a propositional formula under variable costs and a probability distribution over truth assignments. We present a branch-and-bound algorithm with variable-selection heuristics, pruning, and caching. To the best of our knowledge, it is the first practical exact algorithm for this level of generality. Experiments on random instances demonstrate scalability and quantify the efficiency-quality trade-off of a greedy beam-search variant. We additionally evaluate a structured heart-disease diagnosis instance. Finally, we prove that the problem is $\#P$-hard and contained in $\mathrm{PSPACE}$.

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

NFR 評価のためのマルチターン LLM ダイアログの精度と満足度

LLM ベースの対話アシスタントはソフトウェア開発者にとって主流のツールとなっていますが、現在の評価ベンチマークは機能の正しさのみに焦点を当てています。このため、本質的に曖昧でコンテキストに依存し、プログラムの多くの部分に関係する非機能要件 (NFR) を処理する際に、これらの会話の品質と正確性を評価する際に重大なギャップが残ります。これらのシステムが NFR に関する協調推論をどの程度サポートしているかを評価するには、シングル ターンの精度を超えて、システムの出力の正確さとマルチ ターン インタラクションの品質の両方を取得する方法が必要です。このペーパーでは、医療保険の相互運用性と責任に関する法律 (HIPAA) 規制順守の領域における、開発者と LLM ベースのエージェントとの間の複数回にわたる会話の精度と品質を調査します。私たちは 49 人のプログラマーを雇い、GitHub Copilot と対話し、HIPAA 規制に準拠するように設計されたシステムである iTrust コードベースに対して 148 個の HIPAA 由来の NFR を、要件満足度、推論、コード ローカリゼーションの 3 つの側面にわたって評価しました。開発者は LLM 評価に同意する傾向がありますが、専門家のグラウンド トゥルースに対する精度は低いことがわかりました。ユーザー満足度をモデル化したところ、システムの応答時間が長くなり、情報提供ターンが増えるとユーザー満足度にマイナスの影響が出るのに対し、積極的なインタラクションはプラスの影響を与えることがわかりました。私たちの調査結果は、NFR 評価をサポートする LLM ベースの対話システムを設計するための洞察を提供します。

原文 (English)

Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment

LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness. This leaves a critical gap in assessing the quality and accuracy of these conversations when handling Non-Functional Requirements (NFRs), which are inherently vague, context-dependent, and involve many parts of a program. Evaluating how well these systems support collaborative reasoning about NFRs requires methods that go beyond single-turn accuracy to capture both the correctness of the system's outputs and the quality of the multi-turn interaction. In this paper, we investigate the accuracy and quality of multi-turn conversations between developers and an LLM-based agent in the domain of Health Insurance Portability and Accountability Act (HIPAA) regulatory compliance. We hired 49 programmers to interact with GitHub Copilot to assess 148 HIPAA-derived NFRs against the iTrust codebase, a system designed to comply with HIPAA regulations, across three dimensions: requirement satisfaction level, reasoning, and code localization. We find that developers tend to agree with LLM assessments, but accuracy against expert ground truth is low. We model user satisfaction and find that longer system responses and more information-providing turns negatively affect user satisfaction, whereas proactive interactions positively affect it. Our findings provide insights for designing LLM-based dialogue systems that support NFR assessment.

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

RAG システムにおける以前の優位性の定量化

検索拡張生成(RAG)は大規模言語モデルを外部の知識に基づいて構築しますが、現在の評価は「認識論的盲目さ」に悩まされる離散ヒューリスティックに依存しており、真の文脈情報抽出とパラメトリック記憶想起を区別できません。これに対処するために、NCU (Normalized Context Utilization) メトリクスを導入し、ゼロショット、オラクル、および敵対的条件にわたる連続トークンのログ確率を活用して、コンテキスト情報の獲得を厳密に定量化します。独自の商用 API と並行して 1.5B から 72B のパラメータの範囲のアーキテクチャを評価すると、厳密な事実抽出 (思考連鎖推論なし) の場合、従来のスケーリング則では極端な利益逓減が見られることが明らかになりました。つまり、高効率の小型言語モデル (SLM) は、大容量アーキテクチャに匹敵するか、それを上回っています。さらに、「事前優位性」がモデルのスケールや独自のアラインメントと相関していることを示します。評価された商用 API は、敵対的紛争のほぼ半数で明示的な外部証拠を無効にしただけでなく、パラメトリック事前条件が矛盾した場合にシステムの信頼崩壊 (ネガティブトランスファー) に頻繁に悩まされました。私たちの調査結果は、厳密な抽出ワークフローにおける SLM の構造的認識論的な利点と優れた文脈順守を強調しています。

原文 (English)

Quantifying Prior Dominance in RAG Systems

Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics that suffer from ''epistemic blindness'' - failing to distinguish genuine contextual information extraction from parametric memory recall. To address this, we introduce the Normalized Context Utilization (NCU) metric, leveraging continuous token log-probabilities across zero-shot, oracle, and adversarial conditions to strictly quantify contextual information gain. Evaluating architectures ranging from 1.5B to 72B parameters alongside a proprietary commercial API reveals that for strict factual extraction (without Chain-of-Thought reasoning), traditional scaling laws exhibit extreme diminishing returns: highly efficient Small Language Models (SLMs) match or outperform high-capacity architectures. Furthermore, we demonstrate that ``Prior Dominance'' correlates with model scale and proprietary alignments. The evaluated commercial API not only overrode explicit external evidence in nearly half of adversarial conflicts, but also frequently suffered from systemic confidence collapse (Negative Transfer) when its parametric priors were contradicted. Our findings highlight the structural epistemic advantage and superior contextual adherence of SLMs in strict extraction workflows.

2026-06-25 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

PixJail: テキストから画像へのジェイルブレイク評価のための自己進化する紙からパイプラインへの複製

Text-to-Image (T2I) ジェイルブレイク技術が急速に進化するにつれて、既存のベンチマークや複製ワークフローが追いつくのに苦労することがよくあります。さらに重要なことは、T2I ジェイルブレイク評価は単一のプロンプト レベルのテストではなく、プロンプト変換、画像生成、安全性フィルタリング、マルチモーダル判定などの複数の段階によって形成されるパイプライン レベルの問題であるということです。このため、複数の論文の結果を確実に再現し、公正に比較することが困難になります。このギャップを埋めるために、再現可能な T2I ジェイルブレイク評価のための自己進化する紙からパイプラインへのエージェント フレームワークである PixJail を提案します。 T2I ジェイルブレイク ペーパーとオプションの参照コードが与えられると、PixJail は元の実験結果を忠実に再現しながら、統一された契約の下でペーパー固有の攻撃モジュールと実行可能な評価パイプラインを迅速に構築します。 PixJail はさらに、論文のダイジェスト、攻撃の進化パターン、再利用可能なテンプレート、失敗例、バージョン管理された成果物を保存するメモリ バンクを維持し、以前の経験を再利用する将来の再現作業を可能にします。コードが利用可能な文書とコードが利用できない文書の両方を含む、11 の代表的な T2I 脱獄方法を再現します。元の設定では、フレームワークは最小限のエラー (2.1\% 平均、0\% 中央値) で以前の結果を正確に復元します。私たちは、PixJail が将来の T2I ジェイルブレイクの再現と評価のための統合基盤として機能し、手作業の労力を大幅に軽減できることを願っています。

原文 (English)

PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation

As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. More importantly, T2I jailbreak evaluation is not a single prompt-level test, but a pipeline-level problem shaped by multiple stages, including prompt transformation, image generation, safety filtering, and multimodal judging. This makes results across papers difficult to reliably reproduce and fairly compare. To bridge this gap, we propose PixJail, a self-evolving paper-to-pipeline agent framework for reproducible T2I jailbreak evaluation. Given a T2I jailbreak paper and optional reference code, PixJail rapidly constructs a paper-specific attack module and a runnable evaluation pipeline under a unified contract, while faithfully reproducing the original experimental results. PixJail further maintains a memory bank that stores paper digests, attack evolution patterns, reusable templates, failure cases, and versioned artifacts, enabling future reproduction efforts to reuse prior experience. We reproduce eleven representative T2I jailbreak methods, including both code-available and code-unavailable papers. Under their original settings, our framework accurately recovers prior results with minimal error (2.1\% average, 0\% median). We hope that PixJail can serve as a unified foundation for future T2I jailbreak reproduction and evaluation, significantly reducing manual effort.

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SURGELLM: クラスバランス正規化によるタスク認識機能ゲーティングによるマルチタスク評価の再考

異種混合の NLP タスク全体に導入された微調整されたエンコーダーは、3 つの複雑な問題に直面しています。それは、帰納的バイアスの不一致、特徴量統計のクラス不均衡の破損、および外部の語彙知識に注意を払うメカニズムの欠如です。 \textbf{\surgellm} は、専用の軽量モジュールでそれぞれに対応する統合トランスフォーマー フレームワークです。 \emph{外科的特徴ゲート} (精選された語彙指標と \texttt{[CLS]} で次元ごとのシグモイドを学習します。特徴が有益でない場合は、明らかに同一性に退化します)、 \emph{タスク条件付きプレフィックス トークン} (量子化された特徴値とタスク同一性)すべての入力に付加されます)、および \emph{インスタンス加重正規化} (IWN; ゲート統計からクラス事前バイアスを削除します)。私たちは、\emph{外科的特徴の位置合わせ} にゲートをリンクする超過リスク限界の利点を証明します。 SST-2、マルチホップ取得、LLM プロンプト帰属、および著者権検出の 4 つのタスクにわたって、3 つのシードにわたる 17,830 の例と 11 のモデル バリアントをカバーする IWN バリアントは、マクロ F1 \textbf{0.940} (最強の非 IWN ベースラインに対して $+0.036$、著者権検出で $+0.130$) を達成しました。ランダム語彙制御 ($-0.028$ avg.\ F1) は、ゲインがパラメトリックではなく語彙によるものであることを確認します。コード、語彙、および $99.5\%$-recovery 自動抽出レシピがリリースされています。

原文 (English)

SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization

Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance corruption of feature statistics, and no mechanism to condition attention on external lexical knowledge. We introduce \textbf{\surgellm}, a unified transformer framework that addresses each with a dedicated lightweight module: a \emph{surgical feature gate} (learned per-dimension sigmoid over curated lexical indicators and \texttt{[CLS]}; provably degenerates to identity when features are uninformative), \emph{task-conditioned prefix tokens} (quantized feature values and task identity prepended to every input), and \emph{Instance-Weighted Normalization} (IWN; removes class-prior bias from gate statistics). We prove an excess-risk bound linking gate benefit to \emph{surgical feature alignment}. Across four tasks, SST-2, multi-hop retrieval, LLM-prompt attribution, and authorship detection, covering 17,830 examples and eleven model variants over three seeds, the IWN variant achieves macro-F1 \textbf{0.940} ($+0.036$ over the strongest non-IWN baseline; $+0.130$ on authorship detection). A random-vocabulary control ($-0.028$ avg.\ F1) confirms gains are lexical, not parametric. Code, vocabularies, and a $99.5\%$-recovery auto-extraction recipe are released.

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

大規模言語モデル評価におけるプロンプトランキングの安定性について

プロンプトベースの対話は、大規模言語モデル (LLM) を使用するための主要なパラダイムとなっています。LLM では、複数の候補プロンプトが評価され、最上位のプロンプトが下流で使用するために選択されます。このワークフローは、評価条件が多少変動してもプロンプト ランキングが安定していることを暗黙的に前提としています。この論文では、ランダム シードや限定された評価サブセットなどの一般的な変動要因の下でのプロンプト ランキングの安定性を体系的に研究します。 3 つのオープンウェイト LLM と 2 つのベンチマーク タスクにわたって、全体的なランク相関は多くの場合中程度から高である一方で、最もパフォーマンスの高いプロンプトのアイデンティティは頻繁に変化し、信頼性の低い選択決定につながることがわかりました。この問題に対処するために、パフォーマンスと分散の両方を考慮した下限信頼限界に基づいた、単純な安定性を意識した選択戦略を提案します。私たちの結果は、このアプローチが不安定な環境での堅牢性を向上させながら、より安定した環境で競争力を維持できることを示しています。これらの調査結果は、即時選択と LLM ベンチマークにおける評価の不確実性を考慮することの重要性を強調しています。

原文 (English)

On the Stability of Prompt Ranking in Large Language Model Evaluation

Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evaluated and the top-ranked one is selected for downstream use. This workflow implicitly assumes that prompt rankings are stable under minor variations in evaluation conditions. In this paper, we systematically study prompt ranking stability under common sources of variability, including random seeds and limited evaluation subsets. Across three open-weight LLMs and two benchmark tasks, we find that while overall rank correlations are often moderate to high, the identity of the top-performing prompt frequently changes, leading to unreliable selection decisions. To address this issue, we propose a simple stability-aware selection strategy based on a lower confidence bound, which accounts for both performance and variance. Our results show that this approach improves robustness in unstable settings while remaining competitive in more stable regimes. These findings highlight the importance of accounting for evaluation uncertainty in prompt selection and LLM benchmarking.

2026-06-25 13:00 JSTarXiv cs.AIビジネス/資金調達

ノードプロパティ予測のためのグラフ基盤モデルの公正な評価

産業や科学のさまざまな分野でグラフ構造データが広く使用されているため、グラフ基盤モデル (GFM) の開発が最近大きな注目を集めています。多くの異なるタイプのモデルが GFM と呼ばれますが、ノード プロパティ予測タスク用に設計された GFM に特に関心が払われています。GFM は、金融およびソーシャル ネットワークでの不正検出から、電子商取引およびユーザー生成コンテンツ プラットフォームの推奨システムに至るまで、多くの実世界のアプリケーションを備えた Graph ML で最も人気のある設定の 1 つです。このタスク用の多数の GFM が最近提案されていますが、この分野は統一された評価設定に収束しておらず、さまざまな研究がモデルを大幅に異なる方法で評価しているため、GFM 同士や他の種類のモデルとの信頼できる比較ができません。この作業では、ノード プロパティ予測のために最近の 9 つの GFM の公正かつ厳密な再評価を実施し、それらを強力なグラフ ニューラル ネットワーク (GNN) ベースラインと比較します。これらの GFM の中で、事前データ適合ネットワーク パラダイムに基づく最新のものだけが、推論コストは高くなりますが、予測パフォーマンスにおいて適切に調整された GNN より優れていることがわかりました。

原文 (English)

A Fair Evaluation of Graph Foundation Models for Node Property Prediction

Due to the wide use of graph-structured data in different fields of industry and science, the development of Graph Foundation Models (GFMs) has recently attracted a lot of attention. While many different types of models are called GFMs, particular interest has been paid to GFMs designed for node property prediction tasks, which is one of the most popular settings in Graph ML with lots of real-world applications from fraud detection in financial and social networks to recommendation systems for e-commerce and user-generated content platforms. While a number of GFMs for this task have been recently proposed, the field has not converged to a unified evaluation setting, and different works evaluate their models in widely different ways, preventing reliable comparison of GFMs with each other and with other types of models. In this work, we conduct a fair and rigorous reevaluation of 9 recent GFMs for node property prediction, comparing them to strong Graph Neural Network (GNN) baselines. We find that, among these GFMs, only the most recent ones based on the Prior-data Fitted Networks paradigm outperform well-tuned GNNs in predictive performance, although at a higher inference cost.

2026-06-25 13:00 JSTarXiv cs.AIビジネス/資金調達

複雑です: AI を活用した AAC インターフェイスの設計と評価について

人工知能 (AI) は、拡張代替コミュニケーション (AAC) を使用する人々がシステムでできることを強化できます。ただし、AI を活用した AAC インターフェイスの評価は難しい場合があります。人々は交差する存在であり、現在の評価指標では、人々が AAC に対して抱く可能性のある多面的で微妙な欲求を捉えるのが難しい場合があります。私たちは、AAC の 6 つの問題空間の複雑な性質を調査し、これらの空間で AI がどのように使用されるかを検討し、人々の交差するニュアンスを考慮したより堅牢な評価方法を提案します。また、これらの問題領域全体で発生するより広範な問題と、提案された評価方法を使用してそれらにどのように対処できるかについても説明します。

原文 (English)

It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are able to do with their systems. However, evaluating AI-powered AAC interfaces can be difficult. People are intersectional beings and current evaluation metrics can struggle to capture the multifaceted and nuanced desires people may have for their AAC. We explore the complicated nature of six AAC problem spaces, explore how AI might be used in these spaces, and suggest more robust methods of evaluation that take the intersectional nuances of people into account. We also discuss broader issues that arise across these problem spaces and how they could be addressed using our proposed evaluation methods.

2026-06-25 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

InSight: 操縦可能な VLA を介した自己ガイドによるスキル習得

ビジョン言語アクション (VLA) モデルはデモンストレーションから操作スキルを学習できますが、その機能はトレーニング データ内のスキルによって制限されます。我々は、原始的な動作レベル (例: 「グリッパーをボウルに移動する」、「上方に持ち上げる」、「ボトルに注ぐ」など) で VLA を操作可能にすることで、自律的なスキル習得を可能にするフレームワークである InSight を紹介します。 InSight は 2 つの主要なステージで構成されます。(1) VLA プリミティブのステアビリティを可能にするために、VLM プラン分解とエンドエフェクター ポーズによってデモンストレーションをラベル付きプリミティブに分割する自動セグメンテーション パイプライン、(2) 新しいタスクを達成するために必要な欠落しているプリミティブを特定し、VLM が提案する低レベル制御を使用して欠落しているプリミティブのデモンストレーションを自律的に試み、成功したものを自動的にラベル付け、保存、統合する VLM ガイド付きデータ フライホイールVLA トレーニング セットへのデモンストレーション。当社では、シミュレーションおよび実際の操作タスク (ブロックの反転、引き出しの閉め方、掃除、ひねり、流し込みなど) にわたって InSight を評価します。これらの対象スキルを人間がデモンストレーションする必要はありません。一度学習すると、これらのプリミティブを構成して、人間による追加のデモンストレーションなしで、新しい長期的なタスクを実行することができます。私たちの調査結果は、原始的なステアビリティが VLA ポリシーにおける継続的なスキル習得のための実用的な基盤となることを示しています。プロジェクトの Web サイト: https://insight-vla.github.io。

原文 (English)

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a framework that unlocks autonomous skill acquisition by rendering VLAs steerable at the primitive-action level (e.g., "move gripper to the bowl", "lift upward", "pour the bottle"). InSight consists of two primary stages: (1) an automated segmentation pipeline that partitions demonstrations into labeled primitives via VLM plan decomposition and end-effector poses to enable VLA primitive steerability, and (2) a VLM-guided data flywheel that identifies missing primitives required to accomplish a novel task, autonomously attempts demonstrations of the missing primitives with VLM-proposed low-level control, and automatically labels, stores, and integrates successful demonstrations into the VLA training set. We evaluate InSight across simulation and real-world manipulation tasks, including block flipping, drawer closing, sweeping, twisting, and pouring, without any human demonstrations of these target skills. Once learned, these primitives can be composed to execute novel, long-horizon tasks without additional human demonstrations. Our findings demonstrate that primitive steerability provides a practical foundation for continual skill acquisition in VLA policies. Project website: https://insight-vla.github.io.

2026-06-25 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Evolving Programmatic Skill Networks

We study continual skill acquisition in open-ended embodied environments where an agent must construct, refine, and reuse an expanding libr…

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier langu…

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions

Recent studies have raised significant concerns regarding the reliability of current mathematics benchmarks, highlighting issues such as si…

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達研究/論文

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of aut…

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production

Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the a…

2026-06-25 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: kernel-driven measurement instability and noise-driven detection fragility, invisible to DICOM metadata

AI governance for medical imaging is formalizing: the 2026 ACR-SIIM Practice Parameter recommends local acceptance testing and ongoing drif…

2026-06-25 11:00 JSTITmedia AI+ロボティクスビジネス/資金調達

「今日言うつもりはなかったが……」 孫正義氏が明かした「ロボット自動量産工場」の実態

「今日ここで言うつもりはなかったんですが」──。ソフトバンクグループが6月24日に開催した株主総会の質疑応答で、会長兼社長の孫正義氏が、投資先の現場で起きている現場実態を明かす一幕があった。

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

T2D ベンチ: 多層臨床ライフスタイル ナレッジ グラフを使用した 2 型糖尿病の LLM 出力の証拠ゲート型評価

大規模言語モデル (LLM) は、2 型糖尿病に対する臨床的に流暢な推奨事項を生成できますが、ガイドラインの制約を満たしたり、ライフスタイルに関連した血糖の主張を明確に正当化したりすることはできません。我々は、LLM 出力が明示的でグラフチェック可能な証拠要件を満たしているかどうかをテストするための、再現可能なベンチマークおよび証拠ゲート型評価フレームワークである T2D-Bench を紹介します。 T2D-Bench は、生体医学スパイン (UMLS、DrugBank、SIDER)、計算可能な ADA 治療標準ルール、血糖検査室効果への機構的なブリッジを介して接続されたライフスタイル知識を組み合わせた、多層の臨床ライフスタイル ナレッジ グラフに基づいて構築されています。診断、投薬の安全性、敵対的なライフスタイルの衝突にわたる 100 の構造化されたビネット全体で、ベースライン出力は、GPT-4o-mini のケースの 35%、GPT-4o のケースの 33% で、ベンチマークで定義されたエビデンスパス チェックに失敗しました。証拠ゲートはサポートされていない省略を検出し、制約付きリビジョンを使用して、出力をベンチマークで定義された証拠要件に検証者レベルで準拠させます。これらの結果は、糖尿病に焦点を当てた LLM 出力において、計算可能な証拠の制約により、裏付けのない臨床上の省略が明示的、測定可能、修正可能になる可能性があることを示しています。

原文 (English)

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims. We present T2D-Bench, a reproducible benchmark and evidence-gated evaluation framework for testing whether LLM outputs satisfy explicit, graph-checkable evidence requirements. T2D-Bench is built on a multi-layer clinical-lifestyle knowledge graph that combines a biomedical spine (UMLS, DrugBank, SIDER), computable ADA Standards of Care rules, and lifestyle knowledge connected through a mechanistic bridge to glycemic laboratory effects. Across 100 structured vignettes spanning diagnosis, medication safety, and adversarial lifestyle conflicts, baseline outputs failed benchmark-defined evidence-path checks in 35% of cases for GPT-4o-mini and 33% for GPT-4o. The evidence gate detects unsupported omissions and uses constrained revision to bring outputs into verifier-level compliance with benchmark-defined evidence requirements. These results show that computable evidence constraints can make unsupported clinical omissions explicit, measurable, and correctable in diabetes-focused LLM outputs.

2026-06-24 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

自律的な評価を備えたコンピュータ使用エージェントの強化学習

Computer-Use Agent (CUA) は、グラフィカル ユーザー インターフェイス内で直接認識して行動することで、高レベルのユーザー目標を実行します。ただし、オープンエンドのデスクトップ環境ではスケーラブルで機械可読な報酬信号がほとんど提供されないため、CUA の強化学習は依然として困難です。タスクの成功は多くの場合視覚的に根拠があり、手作りの報酬関数や高密度の手動ラベルで指定するのは困難です。我々は、GUI エージェントのスケーラブルな監視信号として自律的な視覚言語評価を使用する RL 微調整フレームワークを提案します。最終的なスクリーンショットと元の指示が与えられると、ビジョン言語モデルはタスクの完了を判断し、ポリシーの最適化中にタスク固有のヒューリスティックや手動ラベルを使用せずに最終的なフィードバックを提供します。自律型評価器は不完全であるため、そのフィードバックをノイズの多いバイナリ報酬チャネルとしてモデル化し、近接ポリシー最適化のためのノイズ補正された報酬推定器を導出します。 macOSWorld、Windows Agent Arena、および OSWorld にわたる実験では、修正された評価者の報酬がゼロショットのベースラインと生の評価者の報酬の両方を上回り、成功率がゼロショットのパフォーマンスより平均 12.6 ポイント、生の評価者の微調整よりも 5.1 ポイント向上したことが示されています。これらの結果は、評価者のノイズが明示的にモデル化され補正されている場合、自律評価が GUI 環境における RL の実用的な報酬信号として機能する可能性があることを示唆しています。

原文 (English)

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation

Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environments rarely provide scalable, machine-readable reward signals: task success is often visually grounded and hard to specify with handcrafted reward functions or dense manual labels. We propose an RL fine-tuning framework that uses autonomous vision-language evaluation as a scalable supervision signal for GUI agents. Given a final screenshot and the original instruction, a Vision-Language Model judges task completion and provides terminal feedback without task-specific heuristics or manual labels during policy optimization. Because autonomous evaluators are imperfect, we model their feedback as a noisy binary reward channel and derive a noise-corrected reward estimator for Proximal Policy Optimization. Experiments across macOSWorld, Windows Agent Arena, and OSWorld show that corrected evaluator rewards outperform both zero-shot baselines and raw evaluator rewards, improving success rates by an average of 12.6 percentage points over zero-shot performance and 5.1 points over raw evaluator fine-tuning. These results suggest that autonomous evaluation can serve as a practical reward signal for RL in GUI environments when evaluator noise is explicitly modeled and corrected.

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AdversaBench: 複数の裁判官による確認とモデル間の移行性を備えた自動化された LLM レッドチーム化

大規模な言語モデルの敵対的評価をスケーリングするには、ハード入力を生成する方法と、結果として生じる失敗が本物であることを確認する信頼性の高い方法の両方が必要です。 AdversaBench は、5 つの構造化された演算子でシード プロンプトを変更し、ターゲット モデルをクエリし、メタ ジャッジ タイブレーカーを備えた 3 人のジャッジ パネルを通じて失敗を確認する、エンドツーエンドのレッド チーム パイプラインです。推論、指示に従い、ツールの使用という 3 つのカテゴリにわたる 45 のシードに関する実験を報告します。すべてのシードで失敗が確認されました。 4 つの発見が際立っています。まず、オペレーターの有効性はカテゴリによって大きく異なります。inject_distractor のスコアは、指示に従うシードでは 0.00 の平均報酬ですが、推論とツールの使用では 0.80 ~ 0.83 です。第 2 に、バイナリの失敗率が難しさを隠しています。命令に従うシードでは、攻撃者の反復回数が平均 2.4 回であるのに対し、他のカテゴリでは 1.1 回であり、生存曲線にギャップが見られます。第三に、80 ~ 87% というペアごとのジャッジの一致は、ラベルの歪みによりほぼゼロのコーエンのカッパと共存します。カテゴリレベルの不一致率の方が有益です。第 4 に、Llama 3.1 8B に対して生成された敵対的プロンプトはゼロショットを Llama 3.3 70B に転送します。これは、変異がモデル固有の弱点ではなく一般的な動作パターンを悪用していることを示唆しています。コード、データセット、分析スクリプトは https://github.com/khanak0509/AdversaBench で入手できます。

原文 (English)

AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability

Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm that resulting failures are real. We present AdversaBench, an end-to-end red-teaming pipeline that mutates seed prompts with five structured operators, queries a target model, and confirms failures through a three-judge panel with a meta-judge tiebreaker. We report experiments on 45 seeds across three categories: reasoning, instruction-following, and tool use. Every seed produced a confirmed failure. Four findings stand out. First, operator effectiveness varies sharply by category: inject_distractor scores 0.00 mean reward on instruction-following seeds but 0.80-0.83 on reasoning and tool-use. Second, binary failure rate hides difficulty: instruction-following seeds required 2.4 attacker iterations on average versus 1.1 for other categories, a gap visible in survival curves. Third, pairwise judge agreement of 80-87% coexists with near-zero Cohen's kappa due to label skew; category-level disagreement rates are more informative. Fourth, adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B, suggesting the mutations exploit general behavioral patterns rather than model-specific weaknesses. Code, dataset, and analysis scripts are available at https://github.com/khanak0509/AdversaBench .

2026-06-24 13:00 JSTarXiv cs.AIビジネス/資金調達

確率的ブール関数評価のためのコスト最適決定図

多くの意思決定シナリオでは、情報の取得にさまざまなコストがかかります。変動コストおよび真理の割り当てに対する確率分布の下で命題式を評価する際の期待コストを最小限に抑える決定論的評価戦略を構築する問題を検討します。変数選択ヒューリスティック、枝刈り、およびキャッシュを備えた分岐限定アルゴリズムを紹介します。私たちの知る限り、これはこのレベルの一般性を実現する最初の実用的な正確なアルゴリズムです。ランダム インスタンスの実験では、スケーラビリティを実証し、貪欲なビーム検索バリアントの効率と品質のトレードオフを定量化します。さらに、構造化された心臓病の診断例も評価します。最後に、問題が $\#P$ 困難であり、$\mathrm{PSPACE}$ に含まれていることを証明します。

原文 (English)

Cost-Optimal Decision Diagrams for Stochastic Boolean Function Evaluation

In many decision-making scenarios, acquiring information incurs different costs. We consider the problem of constructing a deterministic evaluation strategy that minimizes the expected cost of evaluating a propositional formula under variable costs and a probability distribution over truth assignments. We present a branch-and-bound algorithm with variable-selection heuristics, pruning, and caching. To the best of our knowledge, it is the first practical exact algorithm for this level of generality. Experiments on random instances demonstrate scalability and quantify the efficiency-quality trade-off of a greedy beam-search variant. We additionally evaluate a structured heart-disease diagnosis instance. Finally, we prove that the problem is $\#P$-hard and contained in $\mathrm{PSPACE}$.

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

NFR 評価のためのマルチターン LLM ダイアログの精度と満足度

LLM ベースの対話アシスタントはソフトウェア開発者にとって主流のツールとなっていますが、現在の評価ベンチマークは機能の正しさのみに焦点を当てています。このため、本質的に曖昧でコンテキストに依存し、プログラムの多くの部分に関係する非機能要件 (NFR) を処理する際に、これらの会話の品質と正確性を評価する際に重大なギャップが残ります。これらのシステムが NFR に関する協調推論をどの程度サポートしているかを評価するには、シングル ターンの精度を超えて、システムの出力の正確さとマルチ ターン インタラクションの品質の両方を取得する方法が必要です。このペーパーでは、医療保険の相互運用性と責任に関する法律 (HIPAA) 規制順守の領域における、開発者と LLM ベースのエージェントとの間の複数回にわたる会話の精度と品質を調査します。私たちは 49 人のプログラマーを雇い、GitHub Copilot と対話し、HIPAA 規制に準拠するように設計されたシステムである iTrust コードベースに対して 148 個の HIPAA 由来の NFR を、要件満足度、推論、コード ローカリゼーションの 3 つの側面にわたって評価しました。開発者は LLM 評価に同意する傾向がありますが、専門家のグラウンド トゥルースに対する精度は低いことがわかりました。ユーザー満足度をモデル化したところ、システムの応答時間が長くなり、情報提供ターンが増えるとユーザー満足度にマイナスの影響が出るのに対し、積極的なインタラクションはプラスの影響を与えることがわかりました。私たちの調査結果は、NFR 評価をサポートする LLM ベースの対話システムを設計するための洞察を提供します。

原文 (English)

Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment

LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional correctness. This leaves a critical gap in assessing the quality and accuracy of these conversations when handling Non-Functional Requirements (NFRs), which are inherently vague, context-dependent, and involve many parts of a program. Evaluating how well these systems support collaborative reasoning about NFRs requires methods that go beyond single-turn accuracy to capture both the correctness of the system's outputs and the quality of the multi-turn interaction. In this paper, we investigate the accuracy and quality of multi-turn conversations between developers and an LLM-based agent in the domain of Health Insurance Portability and Accountability Act (HIPAA) regulatory compliance. We hired 49 programmers to interact with GitHub Copilot to assess 148 HIPAA-derived NFRs against the iTrust codebase, a system designed to comply with HIPAA regulations, across three dimensions: requirement satisfaction level, reasoning, and code localization. We find that developers tend to agree with LLM assessments, but accuracy against expert ground truth is low. We model user satisfaction and find that longer system responses and more information-providing turns negatively affect user satisfaction, whereas proactive interactions positively affect it. Our findings provide insights for designing LLM-based dialogue systems that support NFR assessment.

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

RAG システムにおける以前の優位性の定量化

検索拡張生成(RAG)は大規模言語モデルを外部の知識に基づいて構築しますが、現在の評価は「認識論的盲目さ」に悩まされる離散ヒューリスティックに依存しており、真の文脈情報抽出とパラメトリック記憶想起を区別できません。これに対処するために、NCU (Normalized Context Utilization) メトリクスを導入し、ゼロショット、オラクル、および敵対的条件にわたる連続トークンのログ確率を活用して、コンテキスト情報の獲得を厳密に定量化します。独自の商用 API と並行して 1.5B から 72B のパラメータの範囲のアーキテクチャを評価すると、厳密な事実抽出 (思考連鎖推論なし) の場合、従来のスケーリング則では極端な利益逓減が見られることが明らかになりました。つまり、高効率の小型言語モデル (SLM) は、大容量アーキテクチャに匹敵するか、それを上回っています。さらに、「事前優位性」がモデルのスケールや独自のアラインメントと相関していることを示します。評価された商用 API は、敵対的紛争のほぼ半数で明示的な外部証拠を無効にしただけでなく、パラメトリック事前条件が矛盾した場合にシステムの信頼崩壊 (ネガティブトランスファー) に頻繁に悩まされました。私たちの調査結果は、厳密な抽出ワークフローにおける SLM の構造的認識論的な利点と優れた文脈順守を強調しています。

原文 (English)

Quantifying Prior Dominance in RAG Systems

Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics that suffer from ''epistemic blindness'' - failing to distinguish genuine contextual information extraction from parametric memory recall. To address this, we introduce the Normalized Context Utilization (NCU) metric, leveraging continuous token log-probabilities across zero-shot, oracle, and adversarial conditions to strictly quantify contextual information gain. Evaluating architectures ranging from 1.5B to 72B parameters alongside a proprietary commercial API reveals that for strict factual extraction (without Chain-of-Thought reasoning), traditional scaling laws exhibit extreme diminishing returns: highly efficient Small Language Models (SLMs) match or outperform high-capacity architectures. Furthermore, we demonstrate that ``Prior Dominance'' correlates with model scale and proprietary alignments. The evaluated commercial API not only overrode explicit external evidence in nearly half of adversarial conflicts, but also frequently suffered from systemic confidence collapse (Negative Transfer) when its parametric priors were contradicted. Our findings highlight the structural epistemic advantage and superior contextual adherence of SLMs in strict extraction workflows.

2026-06-24 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

PixJail: テキストから画像へのジェイルブレイク評価のための自己進化する紙からパイプラインへの複製

Text-to-Image (T2I) ジェイルブレイク技術が急速に進化するにつれて、既存のベンチマークや複製ワークフローが追いつくのに苦労することがよくあります。さらに重要なことは、T2I ジェイルブレイク評価は単一のプロンプト レベルのテストではなく、プロンプト変換、画像生成、安全性フィルタリング、マルチモーダル判定などの複数の段階によって形成されるパイプライン レベルの問題であるということです。このため、複数の論文の結果を確実に再現し、公正に比較することが困難になります。このギャップを埋めるために、再現可能な T2I ジェイルブレイク評価のための自己進化する紙からパイプラインへのエージェント フレームワークである PixJail を提案します。 T2I ジェイルブレイク ペーパーとオプションの参照コードが与えられると、PixJail は元の実験結果を忠実に再現しながら、統一された契約の下でペーパー固有の攻撃モジュールと実行可能な評価パイプラインを迅速に構築します。 PixJail はさらに、論文のダイジェスト、攻撃の進化パターン、再利用可能なテンプレート、失敗例、バージョン管理された成果物を保存するメモリ バンクを維持し、以前の経験を再利用する将来の再現作業を可能にします。コードが利用可能な文書とコードが利用できない文書の両方を含む、11 の代表的な T2I 脱獄方法を再現します。元の設定では、フレームワークは最小限のエラー (2.1\% 平均、0\% 中央値) で以前の結果を正確に復元します。私たちは、PixJail が将来の T2I ジェイルブレイクの再現と評価のための統合基盤として機能し、手作業の労力を大幅に軽減できることを願っています。

原文 (English)

PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation

As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. More importantly, T2I jailbreak evaluation is not a single prompt-level test, but a pipeline-level problem shaped by multiple stages, including prompt transformation, image generation, safety filtering, and multimodal judging. This makes results across papers difficult to reliably reproduce and fairly compare. To bridge this gap, we propose PixJail, a self-evolving paper-to-pipeline agent framework for reproducible T2I jailbreak evaluation. Given a T2I jailbreak paper and optional reference code, PixJail rapidly constructs a paper-specific attack module and a runnable evaluation pipeline under a unified contract, while faithfully reproducing the original experimental results. PixJail further maintains a memory bank that stores paper digests, attack evolution patterns, reusable templates, failure cases, and versioned artifacts, enabling future reproduction efforts to reuse prior experience. We reproduce eleven representative T2I jailbreak methods, including both code-available and code-unavailable papers. Under their original settings, our framework accurately recovers prior results with minimal error (2.1\% average, 0\% median). We hope that PixJail can serve as a unified foundation for future T2I jailbreak reproduction and evaluation, significantly reducing manual effort.

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SURGELLM: クラスバランス正規化によるタスク認識機能ゲーティングによるマルチタスク評価の再考

異種混合の NLP タスク全体に導入された微調整されたエンコーダーは、3 つの複雑な問題に直面しています。それは、帰納的バイアスの不一致、特徴量統計のクラス不均衡の破損、および外部の語彙知識に注意を払うメカニズムの欠如です。 \textbf{\surgellm} は、専用の軽量モジュールでそれぞれに対応する統合トランスフォーマー フレームワークです。 \emph{外科的特徴ゲート} (精選された語彙指標と \texttt{[CLS]} で次元ごとのシグモイドを学習します。特徴が有益でない場合は、明らかに同一性に退化します)、 \emph{タスク条件付きプレフィックス トークン} (量子化された特徴値とタスク同一性)すべての入力に付加されます)、および \emph{インスタンス加重正規化} (IWN; ゲート統計からクラス事前バイアスを削除します)。私たちは、\emph{外科的特徴の位置合わせ} にゲートをリンクする超過リスク限界の利点を証明します。 SST-2、マルチホップ取得、LLM プロンプト帰属、および著者権検出の 4 つのタスクにわたって、3 つのシードにわたる 17,830 の例と 11 のモデル バリアントをカバーする IWN バリアントは、マクロ F1 \textbf{0.940} (最強の非 IWN ベースラインに対して $+0.036$、著者権検出で $+0.130$) を達成しました。ランダム語彙制御 ($-0.028$ avg.\ F1) は、ゲインがパラメトリックではなく語彙によるものであることを確認します。コード、語彙、および $99.5\%$-recovery 自動抽出レシピがリリースされています。

原文 (English)

SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization

Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance corruption of feature statistics, and no mechanism to condition attention on external lexical knowledge. We introduce \textbf{\surgellm}, a unified transformer framework that addresses each with a dedicated lightweight module: a \emph{surgical feature gate} (learned per-dimension sigmoid over curated lexical indicators and \texttt{[CLS]}; provably degenerates to identity when features are uninformative), \emph{task-conditioned prefix tokens} (quantized feature values and task identity prepended to every input), and \emph{Instance-Weighted Normalization} (IWN; removes class-prior bias from gate statistics). We prove an excess-risk bound linking gate benefit to \emph{surgical feature alignment}. Across four tasks, SST-2, multi-hop retrieval, LLM-prompt attribution, and authorship detection, covering 17,830 examples and eleven model variants over three seeds, the IWN variant achieves macro-F1 \textbf{0.940} ($+0.036$ over the strongest non-IWN baseline; $+0.130$ on authorship detection). A random-vocabulary control ($-0.028$ avg.\ F1) confirms gains are lexical, not parametric. Code, vocabularies, and a $99.5\%$-recovery auto-extraction recipe are released.

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

大規模言語モデル評価におけるプロンプトランキングの安定性について

プロンプトベースの対話は、大規模言語モデル (LLM) を使用するための主要なパラダイムとなっています。LLM では、複数の候補プロンプトが評価され、最上位のプロンプトが下流で使用するために選択されます。このワークフローは、評価条件が多少変動してもプロンプト ランキングが安定していることを暗黙的に前提としています。この論文では、ランダム シードや限定された評価サブセットなどの一般的な変動要因の下でのプロンプト ランキングの安定性を体系的に研究します。 3 つのオープンウェイト LLM と 2 つのベンチマーク タスクにわたって、全体的なランク相関は多くの場合中程度から高である一方で、最もパフォーマンスの高いプロンプトのアイデンティティは頻繁に変化し、信頼性の低い選択決定につながることがわかりました。この問題に対処するために、パフォーマンスと分散の両方を考慮した下限信頼限界に基づいた、単純な安定性を意識した選択戦略を提案します。私たちの結果は、このアプローチが不安定な環境での堅牢性を向上させながら、より安定した環境で競争力を維持できることを示しています。これらの調査結果は、即時選択と LLM ベンチマークにおける評価の不確実性を考慮することの重要性を強調しています。

原文 (English)

On the Stability of Prompt Ranking in Large Language Model Evaluation

Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evaluated and the top-ranked one is selected for downstream use. This workflow implicitly assumes that prompt rankings are stable under minor variations in evaluation conditions. In this paper, we systematically study prompt ranking stability under common sources of variability, including random seeds and limited evaluation subsets. Across three open-weight LLMs and two benchmark tasks, we find that while overall rank correlations are often moderate to high, the identity of the top-performing prompt frequently changes, leading to unreliable selection decisions. To address this issue, we propose a simple stability-aware selection strategy based on a lower confidence bound, which accounts for both performance and variance. Our results show that this approach improves robustness in unstable settings while remaining competitive in more stable regimes. These findings highlight the importance of accounting for evaluation uncertainty in prompt selection and LLM benchmarking.

2026-06-24 13:00 JSTarXiv cs.AIビジネス/資金調達

ノードプロパティ予測のためのグラフ基盤モデルの公正な評価

産業や科学のさまざまな分野でグラフ構造データが広く使用されているため、グラフ基盤モデル (GFM) の開発が最近大きな注目を集めています。多くの異なるタイプのモデルが GFM と呼ばれますが、ノード プロパティ予測タスク用に設計された GFM に特に関心が払われています。GFM は、金融およびソーシャル ネットワークでの不正検出から、電子商取引およびユーザー生成コンテンツ プラットフォームの推奨システムに至るまで、多くの実世界のアプリケーションを備えた Graph ML で最も人気のある設定の 1 つです。このタスク用の多数の GFM が最近提案されていますが、この分野は統一された評価設定に収束しておらず、さまざまな研究がモデルを大幅に異なる方法で評価しているため、GFM 同士や他の種類のモデルとの信頼できる比較ができません。この作業では、ノード プロパティ予測のために最近の 9 つの GFM の公正かつ厳密な再評価を実施し、それらを強力なグラフ ニューラル ネットワーク (GNN) ベースラインと比較します。これらの GFM の中で、事前データ適合ネットワーク パラダイムに基づく最新のものだけが、推論コストは高くなりますが、予測パフォーマンスにおいて適切に調整された GNN より優れていることがわかりました。

原文 (English)

A Fair Evaluation of Graph Foundation Models for Node Property Prediction

Due to the wide use of graph-structured data in different fields of industry and science, the development of Graph Foundation Models (GFMs) has recently attracted a lot of attention. While many different types of models are called GFMs, particular interest has been paid to GFMs designed for node property prediction tasks, which is one of the most popular settings in Graph ML with lots of real-world applications from fraud detection in financial and social networks to recommendation systems for e-commerce and user-generated content platforms. While a number of GFMs for this task have been recently proposed, the field has not converged to a unified evaluation setting, and different works evaluate their models in widely different ways, preventing reliable comparison of GFMs with each other and with other types of models. In this work, we conduct a fair and rigorous reevaluation of 9 recent GFMs for node property prediction, comparing them to strong Graph Neural Network (GNN) baselines. We find that, among these GFMs, only the most recent ones based on the Prior-data Fitted Networks paradigm outperform well-tuned GNNs in predictive performance, although at a higher inference cost.

2026-06-24 13:00 JSTarXiv cs.AIビジネス/資金調達

複雑です: AI を活用した AAC インターフェイスの設計と評価について

人工知能 (AI) は、拡張代替コミュニケーション (AAC) を使用する人々がシステムでできることを強化できます。ただし、AI を活用した AAC インターフェイスの評価は難しい場合があります。人々は交差する存在であり、現在の評価指標では、人々が AAC に対して抱く可能性のある多面的で微妙な欲求を捉えるのが難しい場合があります。私たちは、AAC の 6 つの問題空間の複雑な性質を調査し、これらの空間で AI がどのように使用されるかを検討し、人々の交差するニュアンスを考慮したより堅牢な評価方法を提案します。また、これらの問題領域全体で発生するより広範な問題と、提案された評価方法を使用してそれらにどのように対処できるかについても説明します。

原文 (English)

It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are able to do with their systems. However, evaluating AI-powered AAC interfaces can be difficult. People are intersectional beings and current evaluation metrics can struggle to capture the multifaceted and nuanced desires people may have for their AAC. We explore the complicated nature of six AAC problem spaces, explore how AI might be used in these spaces, and suggest more robust methods of evaluation that take the intersectional nuances of people into account. We also discuss broader issues that arise across these problem spaces and how they could be addressed using our proposed evaluation methods.

2026-06-24 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

InSight: 操縦可能な VLA を介した自己ガイドによるスキル習得

ビジョン言語アクション (VLA) モデルはデモンストレーションから操作スキルを学習できますが、その機能はトレーニング データ内のスキルによって制限されます。我々は、原始的な動作レベル (例: 「グリッパーをボウルに移動する」、「上方に持ち上げる」、「ボトルに注ぐ」など) で VLA を操作可能にすることで、自律的なスキル習得を可能にするフレームワークである InSight を紹介します。 InSight は 2 つの主要なステージで構成されます。(1) VLA プリミティブのステアビリティを可能にするために、VLM プラン分解とエンドエフェクター ポーズによってデモンストレーションをラベル付きプリミティブに分割する自動セグメンテーション パイプライン、(2) 新しいタスクを達成するために必要な欠落しているプリミティブを特定し、VLM が提案する低レベル制御を使用して欠落しているプリミティブのデモンストレーションを自律的に試み、成功したものを自動的にラベル付け、保存、統合する VLM ガイド付きデータ フライホイールVLA トレーニング セットへのデモンストレーション。当社では、シミュレーションおよび実際の操作タスク (ブロックの反転、引き出しの閉め方、掃除、ひねり、流し込みなど) にわたって InSight を評価します。これらの対象スキルを人間がデモンストレーションする必要はありません。一度学習すると、これらのプリミティブを構成して、人間による追加のデモンストレーションなしで、新しい長期的なタスクを実行することができます。私たちの調査結果は、原始的なステアビリティが VLA ポリシーにおける継続的なスキル習得のための実用的な基盤となることを示しています。プロジェクトの Web サイト: https://insight-vla.github.io。

原文 (English)

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a framework that unlocks autonomous skill acquisition by rendering VLAs steerable at the primitive-action level (e.g., "move gripper to the bowl", "lift upward", "pour the bottle"). InSight consists of two primary stages: (1) an automated segmentation pipeline that partitions demonstrations into labeled primitives via VLM plan decomposition and end-effector poses to enable VLA primitive steerability, and (2) a VLM-guided data flywheel that identifies missing primitives required to accomplish a novel task, autonomously attempts demonstrations of the missing primitives with VLM-proposed low-level control, and automatically labels, stores, and integrates successful demonstrations into the VLA training set. We evaluate InSight across simulation and real-world manipulation tasks, including block flipping, drawer closing, sweeping, twisting, and pouring, without any human demonstrations of these target skills. Once learned, these primitives can be composed to execute novel, long-horizon tasks without additional human demonstrations. Our findings demonstrate that primitive steerability provides a practical foundation for continual skill acquisition in VLA policies. Project website: https://insight-vla.github.io.

2026-06-24 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Evolving Programmatic Skill Networks

We study continual skill acquisition in open-ended embodied environments where an agent must construct, refine, and reuse an expanding libr…

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier langu…

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions

Recent studies have raised significant concerns regarding the reliability of current mathematics benchmarks, highlighting issues such as si…

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達研究/論文

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of aut…

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production

Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the a…

2026-06-24 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: kernel-driven measurement instability and noise-driven detection fragility, invisible to DICOM metadata

AI governance for medical imaging is formalizing: the 2026 ACR-SIIM Practice Parameter recommends local acceptance testing and ongoing drif…

2026-06-23 22:00 JSTTechCrunch AIエージェントビジネス/資金調達

Fika Jobs raises $4M to build a video-first hiring platform where AI agents interview candidates

Stockholm-based startup Fika Jobs is building a video-first hiring platform that combines AI interview agents with short-form video profile…

2026-06-23 22:00 JSTOpenAILLM/生成AIビジネス/資金調達

Helping build shared standards for advanced AI

OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the…

2026-06-23 08:00 JSTITmedia AI+LLM/生成AIビジネス/資金調達

Copilotの“元”は取れるのか問題、ついに決着? 住友商事、京都市が掴んだ「AI活用の勝ち筋」

生成AIの導入効果が問われる時期に入っている。Microsoftが試算した削減見込み額を独自検証した住友商事が出した結論とは。住友商事と京都市の生成AI活用例を紹介し、生成AI投資の勝ち筋に迫る。

2026-06-23 05:13 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia’s $20B not-acqui-hire deal

What does an AI company do after one of those not-acqui-hire deals? Groq raised money, is leaning into its neocloud business, and is hiring…

2026-06-22 08:00 JSTITmedia AI+LLM/生成AIビジネス/資金調達

Anthropicへの500万ドル間接出資を解消、広告事業のイオレ 軸足移すAIデータセンター事業に資金投入

イオレは6月19日、3月に決めた米Anthropicへの間接出資を解消し、出資金500万ドル(約7億9355万円)全額の返還を受けると発表した。返還資金は自社で開発を進めるAIDCへの投資に充てる。

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

静的リーダーボードを超えて: LLM エージェントの評価の予測的妥当性

エージェントのベンチマークは急速に成長していますが、展開によって明らかにされる 4 つまたは 5 つの側面を超えるベンチマークはありません。このペーパーは、MCP ベースの産業エージェント ベンチマークのこれまでで最大規模の調整された詳細調査を集約したものです。新しい資産クラス (マルチモーダルなビジュアル拡張を含む)、代替オーケストレーション、取得戦略、推論モード、インフラストラクチャの最適化、および評価方法論のプローブをカバーする 14 件の並行実装調査です。これらの研究を以前の 7 つのエージェント ベンチマークと統合すると、集計スコア リーダーボードは導入されたエージェントの評価を体系的に過小評価していると主張します。集計スコアから導出されたランキングは、配布外の設定には転送されません。最近の公開対非公開の競争の回顧展は、このランクの不安定性の直接的な経験的証拠を提供しています。私たちは、サンプル内平均ではなく、予測妥当性、サンプル内ランクとサンプル外ランク間の相関関係によるランキング構成を提案し、HELM とそのエージェント時代の後継者の崩壊を展開関連の次元で明らかにする 12 層の測定装置を報告します。このポジションは、明示的なしきい値を備えた 3 つの改ざん可能な配分外基準を通じて運用可能となります。既存の証拠は部分的にそれを裏付けていますが、確認するには薄すぎます。最後に、事前に登録されたパイロット設計と、次世代のエージェント ベンチマークが何を報告すべきかについての現場レベルのビジョンについて説明します。

原文 (English)

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

大規模言語モデルのブラックボックス不確実性推定法の系統的評価

大規模言語モデル (LLM) は幅広いタスクにわたって強力な機能を示していますが、その出力は信頼性が低いことが多く、幻覚が含まれる可能性があるため、信頼できる LLM を構築するには不確実性推定 (UE) が不可欠です。実際には、多くの主流 LLM は制限された API を介してのみアクセスでき、ロジットや隠れ状態などの内部信号は利用できないため、ブラックボックス UE が特に重要になります。しかし、LLM 用のブラックボックス UE に関する既存の研究は、方法論において断片的なままであり、統一された実証的比較が欠けています。このギャップに対処するために、ブラックボックス UE 手法の体系的なレビューを提示し、言語化ベース、サンプリング ベース、説明ベース、マルチエージェント、およびハイブリッド手法の 5 つのカテゴリに整理します。さらに、統一された評価フレームワークを構築し、4 つのモデルと 4 つのデータセット設定にわたる 24 の代表的な手法をベンチマークします。私たちの結果は、すべての設定において一貫して優勢な単一の方法はないことを示しています。それにもかかわらず、回答空間内の候補を推論して比較する方法は一般に効果的であり、複数の不確実性信号を組み合わせるハイブリッド方法は、ほとんどの条件下で良好に機能します。ベンチマーク データと統一評価フレームワークを公開することで、再現可能な比較を促進し、将来の研究をサポートすることを目指しています。また、実証結果は、LLM 向けの将来のブラック ボックス UE 手法を開発するための実践的なガイダンスを提供します。

原文 (English)

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable and may contain hallucinations, making uncertainty estimation (UE) essential for building trustworthy LLMs. In practice, many mainstream LLMs are only accessible through restricted APIs, where internal signals such as logits and hidden states are unavailable, making black-box UE especially important. However, existing work on black-box UE for LLMs remains fragmented in methodology and lacks a unified empirical comparison. To address this gap, we present a systematic review of black-box UE methods and organize them into five categories: verbalization-based, sampling-based, explanation-based, multi-agent, and hybrid methods. We further build a unified evaluation framework and benchmark 24 representative methods across 4 models and 4 dataset settings. Our results show that no single method consistently dominates across all settings. Nevertheless, methods that reason over and compare candidates in the answer space are generally effective, and hybrid methods that combine multiple uncertainty signals perform well under most conditions. By releasing the benchmark data and a unified evaluation framework, we aim to facilitate reproducible comparisons and support future research, while our empirical findings provide practical guidance for developing future black-box UE methods for LLMs.

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AURA: LLM-as-a-Judge 監査のための不確実性を考慮した適応的改良

大規模言語モデル (LLM) は、人間による大規模な評価は費用がかかり、拡張するのが難しい場合が多いため、オープンエンド生成の判断材料としてますます使用されていますが、その好みは依然として人間の判断に対する不完全な代用です。既存の監査パイプラインは、たとえば人間の注釈、ヒューリスティック フィルタリング、または強力な審査員の出力などから、信頼できる例のサブセットまたはクリーンな監視信号が事前に利用可能であることを前提としていることがよくあります。 LLM の評価では、この仮定は脆弱です。最初の分割では裁判官のバイアスが引き継がれる可能性がありますが、通常、人間による検証は不足しすぎて、大規模に安定したグループを定義できません。私たちは、選択された人間による検証の下で、ペアごとの LLM を審査員として監査するための適応的不確実性を認識した改良フレームワークである AURA を提案します。 AURA は人間による一貫性シグナルを繰り返し学習し、信頼できる証拠を広め、人間によるレビューのために不確実な比較を優先します。重要な考え方は、裁判官に対する信頼を、証拠が蓄積されるにつれて徐々に洗練される潜在的な量として扱うことです。当社は、コンパクトな定式化、安定した改良手順、および合成および実際のペアワイズ LLM 応答データの両方に対する包括的な評価を提供します。

原文 (English)

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their preferences remain imperfect proxies for human judgment. Existing auditing pipelines often assume that a reliable subset of examples or clean supervision signals are available beforehand, for example from human annotation, heuristic filtering, or the outputs of strong judges. In LLM evaluation, this assumption is fragile: the initial split may inherit judge bias, while human verification is typically too scarce to define stable groups at scale. We propose AURA, an adaptive uncertainty--aware refinement framework for auditing pairwise LLM--as--a--judge decisions under selected human verification. AURA iteratively learns a human-consistency signal, propagates reliable evidence, and prioritizes uncertain comparisons for human review. The key idea is to treat trust in a judge as a latent quantity that is progressively refined as evidence accumulates. We provide a compact formulation, a stable refinement procedure, and a comprehensive evaluation on both synthetic and real pairwise LLM-answer data.

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

FFinRED: 金融 LLM レッドチームのための専門家ガイドによるベンチマーク生成および評価フレームワーク

既存の安全性ベンチマークは、一般的な敵対シナリオを対象としていますが、金融特有のリスクは見逃しています。金融 LLM は、対象を絞った評価を必要とする規制コンプライアンス違反、不正行為の助長、組織的な信頼の低下に直面しています。 FinRED は、金融専門家と開発された金融 LLM 安全性評価のための専門家ガイドによるレッドチーム フレームワークです。 FinRED は、世界標準 (FATF や EU DORA など) を規制回避から複雑な詐欺に至るまでの脅威にマッピングする新しい 2 レベルの分類を使用し、専門家が定義したスキーマを通じて実際の財務文書をコンテキスト豊富なレッドチームの行動プロンプト (シード) に変換するスケーラブルなパイプラインと統合します。専門家の厳密な検証により、種子の妥当性と有意義な LLM 安全性評価の現実性が確認されます。また、免責条項のチェックを超え、静的な画一的なルーブリックよりも人間の専門家とより緊密に連携し、重大な偽陰性を 28 から 12 に削減する、専門家によって検証された金融固有のルーブリックも提供しています。国際的に採用されているリスク管理および情報セキュリティ基準 (ISO/IEC 27001 など) と連携して、FinRED は韓国の金融セキュリティ協会 (FSI) の規制サンドボックスに導入されています。実際の金融サービスにおける生成 AI セキュリティ評価。二重使用のリスクを軽減するために、データセット、生成パイプライン、プロンプト テンプレート、および評価フレームワークは、https://github.com/selectstar-ai/FinRED-paper および https://huggingface.co/datasets/datumo/FinRED で資格のある研究者向けに制限されています。

原文 (English)

FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance violations, fraud facilitation, and systemic trust erosion that require targeted evaluation. We introduce FinRED, an expert-guided red-teaming framework for financial LLM safety evaluation developed with financial experts. FinRED uses a novel two-level taxonomy mapping global standards (e.g., FATF and EU DORA) to threats ranging from regulatory evasion to complex fraud, integrated with a scalable pipeline that converts real financial documents into context-rich red-teaming Behavioral Prompts (seeds) through an expert-defined schema. Rigorous expert validation confirms seed plausibility and realism for meaningful LLM safety evaluation. We also provide an expert-validated, finance-specific rubric that goes beyond disclaimer checks, aligns more closely with human experts than static one-size-fits-all rubrics, and reduces critical false negatives from 28 to 12. Aligned with internationally adopted risk-management and information-security standards (e.g., ISO/IEC 27001), FinRED is deployed in South Korea's Financial Security Institute (FSI) regulatory sandbox for generative AI security evaluation in real financial services. To mitigate dual-use risks, the dataset, generation pipeline, prompt template, and evaluation framework are gated for qualified researchers at https://github.com/selectstar-ai/FinRED-paper and https://huggingface.co/datasets/datumo/FinRED.

2026-06-20 13:00 JSTarXiv cs.AIビジネス/資金調達

ICUにおけるイベントベースのバースト抑制検出のためのEEG基盤モデルの評価

バースト抑制 (BS) は、臨床的に関連のある脳波 (EEG) パターンで、重症患者、特に集中治療室 (ICU) で誘発された昏睡状態の患者の鎮静深度と脳活動を監視するために使用されます。 BS パターンは患者ごとに大幅に異なり、注釈付きのデータセットが不足しているため、自動バースト検出は依然として困難です。最近、EEG Foundation Models (FM) は、いくつかの下流 EEG アプリケーションにわたって有望であることが示されていますが、BS 検出におけるその有用性はまだ解明されていません。我々は、患者固有のキャリブレーションを行わずに、縮小モンタージュICU EEGにおけるバースト検出のためのEEG FMを評価する最初の研究を紹介します。 REVE ベース、LUNA-large、LuMamba-Tiny を適応閾値ベースラインとタスク固有の EEGNet ベースラインと比較します。さらに、従来の EEG ウィンドウベースの分類をイベントベースのバースト検出評価で補完します。これは、バースト エピソードが正しく検出されているかどうかを臨床的に評価するのに役立ち、予期されるアノテーションの変動による影響を軽減します。最良のモデルである REVE ベースは、最高のイベントベース F1 スコア ($0.868 \pm 0.167$) を達成し、EEGNet および適応しきい値処理と比較して、1 分あたりのバースト誤差をそれぞれ 52.1% および 36.2% 削減し、ICU でのスケーラブルな EEG モニタリングのための FM をサポートしました。アブレーション実験では、凍結バックボーン トレーニング、2 ステップの微調整、および LoRA ベースの適応に関して、完全な微調整が最も効果的な適応戦略であることが示され、LUNA-large の場合、凍結バックボーン トレーニングに比べてイベントベースの F1 スコアが最大 $+0.102$ 向上しました。ラベル付きデータセットを減らした場合、事前トレーニング済み REVE ベースはコホートの 25% で $+0.723$ イベントベースの F1 ポイントだけランダム初期化を上回り、限られたラベル付きデータでのバースト検出に適応させた場合の事前トレーニング FM 表現の利点を示しています。

原文 (English)

Evaluation of EEG Foundation Models for Event-Based Burst-Suppression Detection in ICU

Burst suppression (BS) is a clinically relevant electroencephalographic (EEG) pattern used to monitor sedation depth and brain activity in critically ill patients, particularly during induced coma in Intensive Care Units (ICUs). Automatic burst detection remains challenging because BS patterns vary substantially between patients and annotated datasets are scarce. Recently, EEG Foundation Models (FMs) have shown promise across several downstream EEG applications, but their usefulness for BS detection remains unexplored. We present the first study to evaluate EEG FMs for burst detection in reduced-montage ICU EEG without patient-specific calibration. We compare REVE-base, LUNA-large and LuMamba-Tiny with an adaptive thresholding baseline and a task-specific EEGNet baseline. Additionally, we complement conventional EEG window-based classification with event-based burst detection evaluation. This helps assessing clinically whether burst episodes are correctly detected, reducing the impact of expected annotation variability. The best model, REVE-base, achieved the highest event-based F1-score ($0.868 \pm 0.167$) and reduced burst-per-minute error by 52.1% and 36.2% compared to EEGNet and adaptive thresholding respectively, supporting FMs for scalable EEG monitoring in ICU. Ablation experiments showed that full fine-tuning was the most effective adaptation strategy with respect to frozen-backbone training, two-step fine-tuning, and LoRA-based adaptation, improving event-based F1-score over frozen-backbone training by up to $+0.102$ for LUNA-large. With reduced labeled datasets, pretrained REVE-base outperformed random initialization by $+0.723$ event-based F1 points at 25% of the cohort, demonstrating the benefit of pretraining FM representations when adapted to burst detection with limited labeled data.

2026-06-20 13:00 JSTarXiv cs.AIビジネス/資金調達

学習者ベースのコンセプトドリフト検出: 分析と評価

進化するストリーミング環境に導入された機械学習アルゴリズムは、一般に概念ドリフトと呼ばれる非定常データ分布を処理する必要があります。概念ドリフトの存在は、予測パフォーマンスを大幅に低下させ、堅牢な意思決定をサポートする能力を妨げる可能性があるため、多くの実世界のアプリケーションにとって大きな課題となります。したがって、長期にわたり高い精度を維持するには、ドリフト イベントをタイムリーかつ効率的に検出することが重要です。この研究では、概念のドリフト特性と、いくつかのカテゴリにわたる多数のドリフト検出アルゴリズムを理論的に検証します。さらに、多様なストリーミング シナリオや急激な変化や段階的な変化などのドリフト特性を示す合成データセットと現実世界のデータセットの両方でパフォーマンスを評価します。この研究は、概念ドリフト特性とドリフト検出器の動作の複雑な概念と、それらの多様な状況への適用性についての理解を高めることを目的としています。

原文 (English)

Learner-based Concept Drift Detection: Analysis and Evaluation

Machine learning algorithms deployed for evolving streaming environments must handle the non-stationary data distributions, commonly referred to as concept drift. The presence of concept drift poses a major challenge for many real-world applications because it can severely degrade their predictive performance, hindering their ability to support robust decision-making. Consequently, the timely and efficient detection of drift events is critical for sustaining high accuracy over time. This study examines theoretically the concept drift characteristics and numerous drift detection algorithms across several categories. Furthermore, we evaluate their performance on both synthetic and real-world datasets exhibiting diverse streaming scenarios and drift characteristics, such as abrupt and gradual changes. This study aims to enhance understanding of the complex notion of concept drift characteristics and behavior of drift detectors, along with their applicability to diverse contexts.

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

レジスターギャップ: ナイジェリアの公共言説のための意味インテリジェンスのフレームワーク

私たちは、表面的な感情を真のコミュニケーション意図から分離する、ナイジェリアの公共の議論のための 9 次元の注釈および評価スキーマであるミーニング インテリジェンス フレームワーク (MIF) を紹介します。 NaijaSenti や AfriSenti などのナイジェリア言語の既存のベンチマークは、感情分類を 3 方向の極性タスク (肯定的、否定的、中立的) として扱います。私たちは、ナイジェリアの談話における AI システムの主な失敗モードは、翻訳の失敗ではなく、文脈の失敗であると主張します。つまり、同じ発話が、話者、聴衆、状況に応じて反対の実用的な力をもたらします。 MIF は、記録、表面的な感情、真の意図、皮肉、コード化されたサブテキスト、リスク層、アノテーターの信頼度、話者の感情、推奨されるコミュニケーション アクションという 9 つのスコア化された次元にわたってこの洞察を運用します。標準英語、ナイジェリア英語、ナイジェリア ピジン、およびコード混合レジスタにわたる 30 項目のキャリブレーション データセットを構築し、ゼロショットおよびスキーマ情報に基づいたプロンプト条件下でフロンティア言語モデル (Gemini 2.5 フラッシュ) を評価します。見出しの結果はレジスタ ギャップです。ゼロショット レジスタの分類精度は 33.3% で、モデルがコンテキスト内で MIF スキーマを受け取ると 73.3% (+40 ポイント) に上昇しました。スキーマ情報に基づくプロンプトの下では、複合意味インテリジェンス スコアが 5.4 ポイント (73.2 から 78.6) 増加し、レジスターの識別、コード化されたサブテキストの検出 (+10 ポイント)、および戦略的アクションの推奨 (+10.3 ポイント) において実質的な増加が最も大きくなります。再現性をサポートするために、フレームワーク仕様、注釈ガイドライン、および 30 項目のパブリック キャリブレーション セットをリリースすると同時に、汚染から保護された評価のためのプライベート ホールドアウト コーパスを保持します。

原文 (English)

The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse

We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent. Existing benchmarks for Nigerian languages, including NaijaSenti and AfriSenti, treat sentiment classification as a three-way polarity task (positive, negative, neutral). We argue that the dominant failure mode of AI systems on Nigerian discourse is not translation failure but context failure: the same utterance carries opposite pragmatic force depending on speaker, audience, and situation. The MIF operationalises this insight across nine scored dimensions: register, surface sentiment, true intent, irony, coded subtext, risk tier, annotator confidence, speaker emotion, and recommended communications action. We construct a 30-item calibration dataset spanning Standard English, Nigerian English, Nigerian Pidgin, and code-mixed registers, and evaluate a frontier language model (Gemini 2.5 Flash) under zero-shot and schema-informed prompting conditions. The headline finding is the Register Gap: zero-shot register classification accuracy is 33.3%, rising to 73.3% (+40 points) when the model receives the MIF schema in-context. The composite Meaning Intelligence Score increases by 5.4 points (73.2 to 78.6) under schema-informed prompting, with the largest practical gains in register identification, coded-subtext detection (+10 points), and strategic action recommendation (+10.3 points). We release the framework specification, annotation guidelines, and the 30-item public calibration set to support reproducibility, while retaining a private holdout corpus for contamination-protected evaluation.

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

When large language models serve as evaluators in multi-agent systems, their systematic evaluation biases propagate through the agent netwo…

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation

Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evalua…

2026-06-20 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SimuWoB: 高速かつ忠実な GUI エージェント ベンチマークのための現実世界のモバイル アプリのシミュレーション

大規模な言語モデルを利用したモバイル GUI エージェントは急速に進歩しており、現実的かつ包括的な評価に対する緊急のニーズが生じています。既存のベンチマークは再現性を優先していますが、実際のアプリケーションで報酬を構築することが難しいため、多くの場合、オープンソース アプリまたはファイル操作タスクに限定されており、ベンチマーク設定と実際の使用状況の間にギャップが生じています。さらに、ほとんどのベンチマークは基本的な接地とナビゲーションに焦点を当てており、複雑で長期にわたる相互作用の範囲は限られています。これらの制限に対処するために、さまざまなタイプと難易度にわたる 120 の困難なタスクを備えたモバイル GUI エージェント用の完全合成ベンチマークである SimuWoB を導入します。私たちは、忠実度の高いタスクと環境を合成し、各タスクに対して有効な報酬を自動的に提供する、堅牢な仮想環境生成フレームワークを構築します。各環境は、URL 経由でアクセスできるバックエンドのない Web ページとしてデプロイされ、効率的で再現可能な評価が可能になります。私たちは、いくつかの最先端のモバイル GUI エージェントで包括的な実験を実施しています。平均成功率はわずか 27.92% であり、長期的なタスクでは 17.82% に低下します。これは、複雑なシナリオの下での現在のエージェントの重大な弱点を明らかにしています。評価結果を実際のサンプル タスクと比較すると、合成環境に基づくエージェントの評価が一般化していることがわかります。さらに、主要な機能の側面にわたる診断上の洞察を提供し、将来のモバイル GUI エージェント開発への影響について説明します。

原文 (English)

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis

GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments. However, directly doing so in real-world environments introduces some challenges that cannot be overlooked. Real-world environments are complex and uncontrollable, making it difficult to construct verifiable rewards and to save or reset states. Existing works prioritize reproducibility but are often limited to open-source apps or file-operation tasks for reliable reward building, leaving a persistent gap from real-world usage. Furthermore, relying on virtual machines or docker images demand high resource requirements and suffer from slow response speeds, which limit the efficiency. We present \sys, a framework that could produce high-fidelity synthesized interactive environments for GUI agents across platforms with verifiable rewards. These environments behave as backend-free webpages accessible via URL, requiring near-zero setup and low resource cost, making the approach suitable for both large-scale evaluation and downstream agent training. We support multiple GUI platforms including mobile, desktop, and automotive/in-vehicle interfaces based on the same pipeline, covering 100+ environments and 1000+ verifiable tasks. Among them, 120 challenging tasks across 63 simulated mobile applications are released as a fully synthesized mobile GUI agent benchmark. Experiment results on five state-of-the-art mobile GUI agents reveal substantial headroom -- the average success rate is only 27.92\%, dropping to 17.82\% on long-horizon subset -- while humans reach 92.08\%. A comparison against real-world sample tasks shows that assessments made in our synthetic environments generalize to real apps. The project website is at https://scalewob.github.io.

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

FundaPod: AI 支援のファンダメンタル投資調査のためのナレッジ グラフ メモリを備えたマルチペルソナ エージェント ポッド プラットフォーム

大規模言語モデル (LLM) は金融分野での適用が増えていますが、既存の研究のほとんどは取引シグナルや予測を中心とした財務 NLP タスクに重点を置いています。対照的に、制度的基礎研究では、人間のアナリストまたは AI エージェントが証拠を収集し、ビジネス推進要因を特定し、競合する視点を比較し、投資メモを作成する必要があります。その広範な目標は、単に結果を予測することではなく、投資知識の累積的な発展に貢献しながら、透明性、再利用可能、検証可能な投資計画を作成することです。 AI 支援のファンダメンタルズ投資調査のためのマルチペルソナ エージェント プラットフォームである FundaPod を紹介します。私たちは、基礎研究は人間中心の意思決定支援タスクであり、取引シグナルの生成とは質的に異なるため、独立性を維持するアーキテクチャの方が適していると主張します。 FundaPod では、バリュー投資家やマクロ戦略家など、さまざまなペルソナを持つ AI エージェントが、共有の出所契約に基づいて独立して調査を実施します。その後、彼らの意見の相違は、知識グラフ記憶システムを通じて人間のポートフォリオ マネージャー (PM) による裁定のために事後的に表面化されます。この論文は、設計科学の実践と認知的分離と人間と機械の協調の理論に基づいた、基礎研究をサポートする人間と AI のハイブリッド システムの 5 つの設計原則を提供します。また、4 つのアーキテクチャ メカニズムについても説明します。1 つは一般投資家の資料を展開可能なエージェントに変えるペルソナ蒸留パイプラインです。プランナーが型指定されたタスク グラフを導出できるようにする宣言型スキル レジストリ。メモの主張を検証可能な情報源に結び付ける根拠のある証拠モデル。そしてティッカー、メモ、アナリスト、テーマを結び付けるナレッジグラフ「第二の脳」。完全なケーススタディとペルソナベースのメモの比較を通じてアーキテクチャを実証します。

原文 (English)

FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research

Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on prediction. Institutional fundamental research, by contrast, requires human analysts or AI agents to gather evidence, identify business drivers, compare competing viewpoints, and generate investment memos. Its broader goal is not merely to predict outcomes, but to produce investment plans that are transparent, reusable, and verifiable, while contributing to the cumulative development of investment knowledge. We present FundaPod, a multi-persona agent platform for AI-assisted fundamental investment research. We argue that fundamental research is a human-centric decision-support task that is qualitatively distinct from trading-signal generation, and is therefore better served by an independence-preserving architecture. In FundaPod, AI agents with different personas, such as value investors or macro strategists, conduct research independently under a shared provenance contract. Their disagreements are then surfaced post hoc for adjudication by the human portfolio manager (PM) through a knowledge-graph memory system. This paper contributes five design principles for human-AI hybrid systems supporting fundamental research, grounded in design-science practice and theories of cognitive isolation and human-machine coordination. It also describes four architectural mechanisms: a persona distillation pipeline that turns public investor materials into deployable agents; a declarative skill registry that lets the planner derive typed task graphs; a grounded evidence model that links memo claims to verifiable sources; and a knowledge-graph "second brain" that connects tickers, memos, analysts, and themes. We demonstrate the architecture through a complete case study and a persona-based memo comparison.

2026-06-20 13:00 JSTarXiv cs.AIビジネス/資金調達

Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search

Auto-bidding is a critical tool for advertisers to improve advertising performance. Recent progress has demonstrated that AI-Generated Bidd…

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM は医師を支援する準備ができていますか?インタラクティブな医師、患者、EHR 支援のための PhysAssistBench

医療 LLM の最も妥当な短期的な役割は、医師の代わりではなく支援することですが、現在の評価では、臨床知識、EHR システムの相互作用、患者とのコミュニケーションなど、個別の能力がテストされることがよくあります。代わりに、医師の支援には同じ対話内でこれらの機能を調整する必要があり、医師は不明確な要求を発行し、患者は症状を曖昧に説明し、EHR システムはツールの正確な使用を要求します。インタラクティブな医師、患者、EHR 支援のベンチマークである PhysAssistBench を紹介します。実際の MIMIC-IV 症例から構築された PhysAssistBench は、スケーラブルなパイプラインを使用してエージェント性患者を構築します。これは、臨床上の事実を維持しながら、静的な EHR 記録を複数ターンの臨床シナリオに変換する、インタラクティブで記録に基づいたエージェントです。 PhysAssistBench は、手動でレビューされ医師が検証した 1,296 ターンの厳選されたバイリンガル評価セットを提供します。主要な LLM を使った実験では、この設定では現在のモデルの信頼性が依然として低いことが示されており、臨床 LLM にとって重要なボトルネックが露呈しています。信頼できる支援には、知識、コミュニケーション、システム全体の調整が必要であり、それらのいずれかで単独の利益を得るのではありません。

原文 (English)

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

静的リーダーボードを超えて: LLM エージェントの評価の予測的妥当性

エージェントのベンチマークは急速に成長していますが、展開によって明らかにされる 4 つまたは 5 つの側面を超えるベンチマークはありません。このペーパーは、MCP ベースの産業エージェント ベンチマークのこれまでで最大規模の調整された詳細調査を集約したものです。新しい資産クラス (マルチモーダルなビジュアル拡張を含む)、代替オーケストレーション、取得戦略、推論モード、インフラストラクチャの最適化、および評価方法論のプローブをカバーする 14 件の並行実装調査です。これらの研究を以前の 7 つのエージェント ベンチマークと統合すると、集計スコア リーダーボードは導入されたエージェントの評価を体系的に過小評価していると主張します。集計スコアから導出されたランキングは、配布外の設定には転送されません。最近の公開対非公開の競争の回顧展は、このランクの不安定性の直接的な経験的証拠を提供しています。私たちは、サンプル内平均ではなく、予測妥当性、サンプル内ランクとサンプル外ランク間の相関関係によるランキング構成を提案し、HELM とそのエージェント時代の後継者の崩壊を展開関連の次元で明らかにする 12 層の測定装置を報告します。このポジションは、明示的なしきい値を備えた 3 つの改ざん可能な配分外基準を通じて運用可能となります。既存の証拠は部分的にそれを裏付けていますが、確認するには薄すぎます。最後に、事前に登録されたパイロット設計と、次世代のエージェント ベンチマークが何を報告すべきかについての現場レベルのビジョンについて説明します。

原文 (English)

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-based industrial-agent benchmark to date: fourteen parallel implementation studies covering new asset classes (including a multi-modal visual extension), alternative orchestrations, retrieval strategies, reasoning modes, infrastructure optimizations, and evaluation-methodology probes. Consolidating those studies with seven prior agent benchmarks, we argue that aggregate-score leaderboards systematically underspecify deployed-agent evaluation. Rankings derived from aggregate scores do not transfer to out-of-distribution settings; recent public-to-hidden competition retrospectives provide direct empirical evidence of this rank instability. We propose ranking configurations by predictive validity, the correlation between in-sample and out-of-sample rank, rather than in-sample mean, and report a twelve-tier measurement apparatus that exposes the deployment-relevant dimensions HELM and its agent-era successors collapse. The position is operationalized through three falsifiable out-of-distribution criteria with explicit thresholds; existing evidence partly supports it but is too thin to confirm. We close with a pre-registered pilot design and a field-level vision for what the next generation of agentic benchmarks should report.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

大規模言語モデルのブラックボックス不確実性推定法の系統的評価

大規模言語モデル (LLM) は幅広いタスクにわたって強力な機能を示していますが、その出力は信頼性が低いことが多く、幻覚が含まれる可能性があるため、信頼できる LLM を構築するには不確実性推定 (UE) が不可欠です。実際には、多くの主流 LLM は制限された API を介してのみアクセスでき、ロジットや隠れ状態などの内部信号は利用できないため、ブラックボックス UE が特に重要になります。しかし、LLM 用のブラックボックス UE に関する既存の研究は、方法論において断片的なままであり、統一された実証的比較が欠けています。このギャップに対処するために、ブラックボックス UE 手法の体系的なレビューを提示し、言語化ベース、サンプリング ベース、説明ベース、マルチエージェント、およびハイブリッド手法の 5 つのカテゴリに整理します。さらに、統一された評価フレームワークを構築し、4 つのモデルと 4 つのデータセット設定にわたる 24 の代表的な手法をベンチマークします。私たちの結果は、すべての設定において一貫して優勢な単一の方法はないことを示しています。それにもかかわらず、回答空間内の候補を推論して比較する方法は一般に効果的であり、複数の不確実性信号を組み合わせるハイブリッド方法は、ほとんどの条件下で良好に機能します。ベンチマーク データと統一評価フレームワークを公開することで、再現可能な比較を促進し、将来の研究をサポートすることを目指しています。また、実証結果は、LLM 向けの将来のブラック ボックス UE 手法を開発するための実践的なガイダンスを提供します。

原文 (English)

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable and may contain hallucinations, making uncertainty estimation (UE) essential for building trustworthy LLMs. In practice, many mainstream LLMs are only accessible through restricted APIs, where internal signals such as logits and hidden states are unavailable, making black-box UE especially important. However, existing work on black-box UE for LLMs remains fragmented in methodology and lacks a unified empirical comparison. To address this gap, we present a systematic review of black-box UE methods and organize them into five categories: verbalization-based, sampling-based, explanation-based, multi-agent, and hybrid methods. We further build a unified evaluation framework and benchmark 24 representative methods across 4 models and 4 dataset settings. Our results show that no single method consistently dominates across all settings. Nevertheless, methods that reason over and compare candidates in the answer space are generally effective, and hybrid methods that combine multiple uncertainty signals perform well under most conditions. By releasing the benchmark data and a unified evaluation framework, we aim to facilitate reproducible comparisons and support future research, while our empirical findings provide practical guidance for developing future black-box UE methods for LLMs.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AURA: LLM-as-a-Judge 監査のための不確実性を考慮した適応的改良

大規模言語モデル (LLM) は、人間による大規模な評価は費用がかかり、拡張するのが難しい場合が多いため、オープンエンド生成の判断材料としてますます使用されていますが、その好みは依然として人間の判断に対する不完全な代用です。既存の監査パイプラインは、たとえば人間の注釈、ヒューリスティック フィルタリング、または強力な審査員の出力などから、信頼できる例のサブセットまたはクリーンな監視信号が事前に利用可能であることを前提としていることがよくあります。 LLM の評価では、この仮定は脆弱です。最初の分割では裁判官のバイアスが引き継がれる可能性がありますが、通常、人間による検証は不足しすぎて、大規模に安定したグループを定義できません。私たちは、選択された人間による検証の下で、ペアごとの LLM を審査員として監査するための適応的不確実性を認識した改良フレームワークである AURA を提案します。 AURA は人間による一貫性シグナルを繰り返し学習し、信頼できる証拠を広め、人間によるレビューのために不確実な比較を優先します。重要な考え方は、裁判官に対する信頼を、証拠が蓄積されるにつれて徐々に洗練される潜在的な量として扱うことです。当社は、コンパクトな定式化、安定した改良手順、および合成および実際のペアワイズ LLM 応答データの両方に対する包括的な評価を提供します。

原文 (English)

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their preferences remain imperfect proxies for human judgment. Existing auditing pipelines often assume that a reliable subset of examples or clean supervision signals are available beforehand, for example from human annotation, heuristic filtering, or the outputs of strong judges. In LLM evaluation, this assumption is fragile: the initial split may inherit judge bias, while human verification is typically too scarce to define stable groups at scale. We propose AURA, an adaptive uncertainty--aware refinement framework for auditing pairwise LLM--as--a--judge decisions under selected human verification. AURA iteratively learns a human-consistency signal, propagates reliable evidence, and prioritizes uncertain comparisons for human review. The key idea is to treat trust in a judge as a latent quantity that is progressively refined as evidence accumulates. We provide a compact formulation, a stable refinement procedure, and a comprehensive evaluation on both synthetic and real pairwise LLM-answer data.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

FFinRED: 金融 LLM レッドチームのための専門家ガイドによるベンチマーク生成および評価フレームワーク

既存の安全性ベンチマークは、一般的な敵対シナリオを対象としていますが、金融特有のリスクは見逃しています。金融 LLM は、対象を絞った評価を必要とする規制コンプライアンス違反、不正行為の助長、組織的な信頼の低下に直面しています。 FinRED は、金融専門家と開発された金融 LLM 安全性評価のための専門家ガイドによるレッドチーム フレームワークです。 FinRED は、世界標準 (FATF や EU DORA など) を規制回避から複雑な詐欺に至るまでの脅威にマッピングする新しい 2 レベルの分類を使用し、専門家が定義したスキーマを通じて実際の財務文書をコンテキスト豊富なレッドチームの行動プロンプト (シード) に変換するスケーラブルなパイプラインと統合します。専門家の厳密な検証により、種子の妥当性と有意義な LLM 安全性評価の現実性が確認されます。また、免責条項のチェックを超え、静的な画一的なルーブリックよりも人間の専門家とより緊密に連携し、重大な偽陰性を 28 から 12 に削減する、専門家によって検証された金融固有のルーブリックも提供しています。国際的に採用されているリスク管理および情報セキュリティ基準 (ISO/IEC 27001 など) と連携して、FinRED は韓国の金融セキュリティ協会 (FSI) の規制サンドボックスに導入されています。実際の金融サービスにおける生成 AI セキュリティ評価。二重使用のリスクを軽減するために、データセット、生成パイプライン、プロンプト テンプレート、および評価フレームワークは、https://github.com/selectstar-ai/FinRED-paper および https://huggingface.co/datasets/datumo/FinRED で資格のある研究者向けに制限されています。

原文 (English)

FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance violations, fraud facilitation, and systemic trust erosion that require targeted evaluation. We introduce FinRED, an expert-guided red-teaming framework for financial LLM safety evaluation developed with financial experts. FinRED uses a novel two-level taxonomy mapping global standards (e.g., FATF and EU DORA) to threats ranging from regulatory evasion to complex fraud, integrated with a scalable pipeline that converts real financial documents into context-rich red-teaming Behavioral Prompts (seeds) through an expert-defined schema. Rigorous expert validation confirms seed plausibility and realism for meaningful LLM safety evaluation. We also provide an expert-validated, finance-specific rubric that goes beyond disclaimer checks, aligns more closely with human experts than static one-size-fits-all rubrics, and reduces critical false negatives from 28 to 12. Aligned with internationally adopted risk-management and information-security standards (e.g., ISO/IEC 27001), FinRED is deployed in South Korea's Financial Security Institute (FSI) regulatory sandbox for generative AI security evaluation in real financial services. To mitigate dual-use risks, the dataset, generation pipeline, prompt template, and evaluation framework are gated for qualified researchers at https://github.com/selectstar-ai/FinRED-paper and https://huggingface.co/datasets/datumo/FinRED.

2026-06-19 13:00 JSTarXiv cs.AIビジネス/資金調達

ICUにおけるイベントベースのバースト抑制検出のためのEEG基盤モデルの評価

バースト抑制 (BS) は、臨床的に関連のある脳波 (EEG) パターンで、重症患者、特に集中治療室 (ICU) で誘発された昏睡状態の患者の鎮静深度と脳活動を監視するために使用されます。 BS パターンは患者ごとに大幅に異なり、注釈付きのデータセットが不足しているため、自動バースト検出は依然として困難です。最近、EEG Foundation Models (FM) は、いくつかの下流 EEG アプリケーションにわたって有望であることが示されていますが、BS 検出におけるその有用性はまだ解明されていません。我々は、患者固有のキャリブレーションを行わずに、縮小モンタージュICU EEGにおけるバースト検出のためのEEG FMを評価する最初の研究を紹介します。 REVE ベース、LUNA-large、LuMamba-Tiny を適応閾値ベースラインとタスク固有の EEGNet ベースラインと比較します。さらに、従来の EEG ウィンドウベースの分類をイベントベースのバースト検出評価で補完します。これは、バースト エピソードが正しく検出されているかどうかを臨床的に評価するのに役立ち、予期されるアノテーションの変動による影響を軽減します。最良のモデルである REVE ベースは、最高のイベントベース F1 スコア ($0.868 \pm 0.167$) を達成し、EEGNet および適応しきい値処理と比較して、1 分あたりのバースト誤差をそれぞれ 52.1% および 36.2% 削減し、ICU でのスケーラブルな EEG モニタリングのための FM をサポートしました。アブレーション実験では、凍結バックボーン トレーニング、2 ステップの微調整、および LoRA ベースの適応に関して、完全な微調整が最も効果的な適応戦略であることが示され、LUNA-large の場合、凍結バックボーン トレーニングに比べてイベントベースの F1 スコアが最大 $+0.102$ 向上しました。ラベル付きデータセットを減らした場合、事前トレーニング済み REVE ベースはコホートの 25% で $+0.723$ イベントベースの F1 ポイントだけランダム初期化を上回り、限られたラベル付きデータでのバースト検出に適応させた場合の事前トレーニング FM 表現の利点を示しています。

原文 (English)

Evaluation of EEG Foundation Models for Event-Based Burst-Suppression Detection in ICU

Burst suppression (BS) is a clinically relevant electroencephalographic (EEG) pattern used to monitor sedation depth and brain activity in critically ill patients, particularly during induced coma in Intensive Care Units (ICUs). Automatic burst detection remains challenging because BS patterns vary substantially between patients and annotated datasets are scarce. Recently, EEG Foundation Models (FMs) have shown promise across several downstream EEG applications, but their usefulness for BS detection remains unexplored. We present the first study to evaluate EEG FMs for burst detection in reduced-montage ICU EEG without patient-specific calibration. We compare REVE-base, LUNA-large and LuMamba-Tiny with an adaptive thresholding baseline and a task-specific EEGNet baseline. Additionally, we complement conventional EEG window-based classification with event-based burst detection evaluation. This helps assessing clinically whether burst episodes are correctly detected, reducing the impact of expected annotation variability. The best model, REVE-base, achieved the highest event-based F1-score ($0.868 \pm 0.167$) and reduced burst-per-minute error by 52.1% and 36.2% compared to EEGNet and adaptive thresholding respectively, supporting FMs for scalable EEG monitoring in ICU. Ablation experiments showed that full fine-tuning was the most effective adaptation strategy with respect to frozen-backbone training, two-step fine-tuning, and LoRA-based adaptation, improving event-based F1-score over frozen-backbone training by up to $+0.102$ for LUNA-large. With reduced labeled datasets, pretrained REVE-base outperformed random initialization by $+0.723$ event-based F1 points at 25% of the cohort, demonstrating the benefit of pretraining FM representations when adapted to burst detection with limited labeled data.

2026-06-19 13:00 JSTarXiv cs.AIビジネス/資金調達

学習者ベースのコンセプトドリフト検出: 分析と評価

進化するストリーミング環境に導入された機械学習アルゴリズムは、一般に概念ドリフトと呼ばれる非定常データ分布を処理する必要があります。概念ドリフトの存在は、予測パフォーマンスを大幅に低下させ、堅牢な意思決定をサポートする能力を妨げる可能性があるため、多くの実世界のアプリケーションにとって大きな課題となります。したがって、長期にわたり高い精度を維持するには、ドリフト イベントをタイムリーかつ効率的に検出することが重要です。この研究では、概念のドリフト特性と、いくつかのカテゴリにわたる多数のドリフト検出アルゴリズムを理論的に検証します。さらに、多様なストリーミング シナリオや急激な変化や段階的な変化などのドリフト特性を示す合成データセットと現実世界のデータセットの両方でパフォーマンスを評価します。この研究は、概念ドリフト特性とドリフト検出器の動作の複雑な概念と、それらの多様な状況への適用性についての理解を高めることを目的としています。

原文 (English)

Learner-based Concept Drift Detection: Analysis and Evaluation

Machine learning algorithms deployed for evolving streaming environments must handle the non-stationary data distributions, commonly referred to as concept drift. The presence of concept drift poses a major challenge for many real-world applications because it can severely degrade their predictive performance, hindering their ability to support robust decision-making. Consequently, the timely and efficient detection of drift events is critical for sustaining high accuracy over time. This study examines theoretically the concept drift characteristics and numerous drift detection algorithms across several categories. Furthermore, we evaluate their performance on both synthetic and real-world datasets exhibiting diverse streaming scenarios and drift characteristics, such as abrupt and gradual changes. This study aims to enhance understanding of the complex notion of concept drift characteristics and behavior of drift detectors, along with their applicability to diverse contexts.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

レジスターギャップ: ナイジェリアの公共言説のための意味インテリジェンスのフレームワーク

私たちは、表面的な感情を真のコミュニケーション意図から分離する、ナイジェリアの公共の議論のための 9 次元の注釈および評価スキーマであるミーニング インテリジェンス フレームワーク (MIF) を紹介します。 NaijaSenti や AfriSenti などのナイジェリア言語の既存のベンチマークは、感情分類を 3 方向の極性タスク (肯定的、否定的、中立的) として扱います。私たちは、ナイジェリアの談話における AI システムの主な失敗モードは、翻訳の失敗ではなく、文脈の失敗であると主張します。つまり、同じ発話が、話者、聴衆、状況に応じて反対の実用的な力をもたらします。 MIF は、記録、表面的な感情、真の意図、皮肉、コード化されたサブテキスト、リスク層、アノテーターの信頼度、話者の感情、推奨されるコミュニケーション アクションという 9 つのスコア化された次元にわたってこの洞察を運用します。標準英語、ナイジェリア英語、ナイジェリア ピジン、およびコード混合レジスタにわたる 30 項目のキャリブレーション データセットを構築し、ゼロショットおよびスキーマ情報に基づいたプロンプト条件下でフロンティア言語モデル (Gemini 2.5 フラッシュ) を評価します。見出しの結果はレジスタ ギャップです。ゼロショット レジスタの分類精度は 33.3% で、モデルがコンテキスト内で MIF スキーマを受け取ると 73.3% (+40 ポイント) に上昇しました。スキーマ情報に基づくプロンプトの下では、複合意味インテリジェンス スコアが 5.4 ポイント (73.2 から 78.6) 増加し、レジスターの識別、コード化されたサブテキストの検出 (+10 ポイント)、および戦略的アクションの推奨 (+10.3 ポイント) において実質的な増加が最も大きくなります。再現性をサポートするために、フレームワーク仕様、注釈ガイドライン、および 30 項目のパブリック キャリブレーション セットをリリースすると同時に、汚染から保護された評価のためのプライベート ホールドアウト コーパスを保持します。

原文 (English)

The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse

We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent. Existing benchmarks for Nigerian languages, including NaijaSenti and AfriSenti, treat sentiment classification as a three-way polarity task (positive, negative, neutral). We argue that the dominant failure mode of AI systems on Nigerian discourse is not translation failure but context failure: the same utterance carries opposite pragmatic force depending on speaker, audience, and situation. The MIF operationalises this insight across nine scored dimensions: register, surface sentiment, true intent, irony, coded subtext, risk tier, annotator confidence, speaker emotion, and recommended communications action. We construct a 30-item calibration dataset spanning Standard English, Nigerian English, Nigerian Pidgin, and code-mixed registers, and evaluate a frontier language model (Gemini 2.5 Flash) under zero-shot and schema-informed prompting conditions. The headline finding is the Register Gap: zero-shot register classification accuracy is 33.3%, rising to 73.3% (+40 points) when the model receives the MIF schema in-context. The composite Meaning Intelligence Score increases by 5.4 points (73.2 to 78.6) under schema-informed prompting, with the largest practical gains in register identification, coded-subtext detection (+10 points), and strategic action recommendation (+10.3 points). We release the framework specification, annotation guidelines, and the 30-item public calibration set to support reproducibility, while retaining a private holdout corpus for contamination-protected evaluation.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

When large language models serve as evaluators in multi-agent systems, their systematic evaluation biases propagate through the agent netwo…

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation

Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evalua…

2026-06-19 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SimuWoB: 高速かつ忠実な GUI エージェント ベンチマークのための現実世界のモバイル アプリのシミュレーション

大規模な言語モデルを利用したモバイル GUI エージェントは急速に進歩しており、現実的かつ包括的な評価に対する緊急のニーズが生じています。既存のベンチマークは再現性を優先していますが、実際のアプリケーションで報酬を構築することが難しいため、多くの場合、オープンソース アプリまたはファイル操作タスクに限定されており、ベンチマーク設定と実際の使用状況の間にギャップが生じています。さらに、ほとんどのベンチマークは基本的な接地とナビゲーションに焦点を当てており、複雑で長期にわたる相互作用の範囲は限られています。これらの制限に対処するために、さまざまなタイプと難易度にわたる 120 の困難なタスクを備えたモバイル GUI エージェント用の完全合成ベンチマークである SimuWoB を導入します。私たちは、忠実度の高いタスクと環境を合成し、各タスクに対して有効な報酬を自動的に提供する、堅牢な仮想環境生成フレームワークを構築します。各環境は、URL 経由でアクセスできるバックエンドのない Web ページとしてデプロイされ、効率的で再現可能な評価が可能になります。私たちは、いくつかの最先端のモバイル GUI エージェントで包括的な実験を実施しています。平均成功率はわずか 27.92% であり、長期的なタスクでは 17.82% に低下します。これは、複雑なシナリオの下での現在のエージェントの重大な弱点を明らかにしています。評価結果を実際のサンプル タスクと比較すると、合成環境に基づくエージェントの評価が一般化していることがわかります。さらに、主要な機能の側面にわたる診断上の洞察を提供し、将来のモバイル GUI エージェント開発への影響について説明します。

原文 (English)

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis

GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments. However, directly doing so in real-world environments introduces some challenges that cannot be overlooked. Real-world environments are complex and uncontrollable, making it difficult to construct verifiable rewards and to save or reset states. Existing works prioritize reproducibility but are often limited to open-source apps or file-operation tasks for reliable reward building, leaving a persistent gap from real-world usage. Furthermore, relying on virtual machines or docker images demand high resource requirements and suffer from slow response speeds, which limit the efficiency. We present \sys, a framework that could produce high-fidelity synthesized interactive environments for GUI agents across platforms with verifiable rewards. These environments behave as backend-free webpages accessible via URL, requiring near-zero setup and low resource cost, making the approach suitable for both large-scale evaluation and downstream agent training. We support multiple GUI platforms including mobile, desktop, and automotive/in-vehicle interfaces based on the same pipeline, covering 100+ environments and 1000+ verifiable tasks. Among them, 120 challenging tasks across 63 simulated mobile applications are released as a fully synthesized mobile GUI agent benchmark. Experiment results on five state-of-the-art mobile GUI agents reveal substantial headroom -- the average success rate is only 27.92\%, dropping to 17.82\% on long-horizon subset -- while humans reach 92.08\%. A comparison against real-world sample tasks shows that assessments made in our synthetic environments generalize to real apps. The project website is at https://scalewob.github.io.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

FundaPod: AI 支援のファンダメンタル投資調査のためのナレッジ グラフ メモリを備えたマルチペルソナ エージェント ポッド プラットフォーム

大規模言語モデル (LLM) は金融分野での適用が増えていますが、既存の研究のほとんどは取引シグナルや予測を中心とした財務 NLP タスクに重点を置いています。対照的に、制度的基礎研究では、人間のアナリストまたは AI エージェントが証拠を収集し、ビジネス推進要因を特定し、競合する視点を比較し、投資メモを作成する必要があります。その広範な目標は、単に結果を予測することではなく、投資知識の累積的な発展に貢献しながら、透明性、再利用可能、検証可能な投資計画を作成することです。 AI 支援のファンダメンタルズ投資調査のためのマルチペルソナ エージェント プラットフォームである FundaPod を紹介します。私たちは、基礎研究は人間中心の意思決定支援タスクであり、取引シグナルの生成とは質的に異なるため、独立性を維持するアーキテクチャの方が適していると主張します。 FundaPod では、バリュー投資家やマクロ戦略家など、さまざまなペルソナを持つ AI エージェントが、共有の出所契約に基づいて独立して調査を実施します。その後、彼らの意見の相違は、知識グラフ記憶システムを通じて人間のポートフォリオ マネージャー (PM) による裁定のために事後的に表面化されます。この論文は、設計科学の実践と認知的分離と人間と機械の協調の理論に基づいた、基礎研究をサポートする人間と AI のハイブリッド システムの 5 つの設計原則を提供します。また、4 つのアーキテクチャ メカニズムについても説明します。1 つは一般投資家の資料を展開可能なエージェントに変えるペルソナ蒸留パイプラインです。プランナーが型指定されたタスク グラフを導出できるようにする宣言型スキル レジストリ。メモの主張を検証可能な情報源に結び付ける根拠のある証拠モデル。そしてティッカー、メモ、アナリスト、テーマを結び付けるナレッジグラフ「第二の脳」。完全なケーススタディとペルソナベースのメモの比較を通じてアーキテクチャを実証します。

原文 (English)

FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research

Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on prediction. Institutional fundamental research, by contrast, requires human analysts or AI agents to gather evidence, identify business drivers, compare competing viewpoints, and generate investment memos. Its broader goal is not merely to predict outcomes, but to produce investment plans that are transparent, reusable, and verifiable, while contributing to the cumulative development of investment knowledge. We present FundaPod, a multi-persona agent platform for AI-assisted fundamental investment research. We argue that fundamental research is a human-centric decision-support task that is qualitatively distinct from trading-signal generation, and is therefore better served by an independence-preserving architecture. In FundaPod, AI agents with different personas, such as value investors or macro strategists, conduct research independently under a shared provenance contract. Their disagreements are then surfaced post hoc for adjudication by the human portfolio manager (PM) through a knowledge-graph memory system. This paper contributes five design principles for human-AI hybrid systems supporting fundamental research, grounded in design-science practice and theories of cognitive isolation and human-machine coordination. It also describes four architectural mechanisms: a persona distillation pipeline that turns public investor materials into deployable agents; a declarative skill registry that lets the planner derive typed task graphs; a grounded evidence model that links memo claims to verifiable sources; and a knowledge-graph "second brain" that connects tickers, memos, analysts, and themes. We demonstrate the architecture through a complete case study and a persona-based memo comparison.

2026-06-19 13:00 JSTarXiv cs.AIビジネス/資金調達

Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search

Auto-bidding is a critical tool for advertisers to improve advertising performance. Recent progress has demonstrated that AI-Generated Bidd…

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM は医師を支援する準備ができていますか?インタラクティブな医師、患者、EHR 支援のための PhysAssistBench

医療 LLM の最も妥当な短期的な役割は、医師の代わりではなく支援することですが、現在の評価では、臨床知識、EHR システムの相互作用、患者とのコミュニケーションなど、個別の能力がテストされることがよくあります。代わりに、医師の支援には同じ対話内でこれらの機能を調整する必要があり、医師は不明確な要求を発行し、患者は症状を曖昧に説明し、EHR システムはツールの正確な使用を要求します。インタラクティブな医師、患者、EHR 支援のベンチマークである PhysAssistBench を紹介します。実際の MIMIC-IV 症例から構築された PhysAssistBench は、スケーラブルなパイプラインを使用してエージェント性患者を構築します。これは、臨床上の事実を維持しながら、静的な EHR 記録を複数ターンの臨床シナリオに変換する、インタラクティブで記録に基づいたエージェントです。 PhysAssistBench は、手動でレビューされ医師が検証した 1,296 ターンの厳選されたバイリンガル評価セットを提供します。主要な LLM を使った実験では、この設定では現在のモデルの信頼性が依然として低いことが示されており、臨床 LLM にとって重要なボトルネックが露呈しています。信頼できる支援には、知識、コミュニケーション、システム全体の調整が必要であり、それらのいずれかで単独の利益を得るのではありません。

原文 (English)

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.

2026-06-19 08:00 JSTITmedia AI+ビジネス/資金調達

融資の決め手、決算書→「データ」「未来のシナリオ」へ 中小企業が資金調達に成功するための最大のポイントは?

中小企業の資金調達の在り方が、大きく変わろうとしている。融資特化型デジタルバンクである01(ゼロワン)銀行(大阪府吹田市)の大塚篤史副社長と北國銀行(金沢市)の竹内均氏(常務執行役員マーケティング部長)が、中小企業の経営や資金調達がどう変わっていくかの見解を語った。

2026-06-19 04:59 JSTTechCrunch AILLM/生成AIビジネス/資金調達

OpenAI is bringing on some big guns in the lead-up to its IPO

OpenAI is bulking up before its IPO, landing Transformer co-inventor Noam Shazeer from Google DeepMind and former Trump AI policy official…

2026-06-19 00:20 JSTTechCrunch AIビジネス/資金調達

General Intuition in talks to raise $300M at around $2B valuation

The startup trains embodied AI and world models using Medal’s dataset of 2 billion videos per year from 10 million monthly active users.

2026-06-18 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

マルチエージェントシミュレーションベースのコミュニティノート評価に向けて

相互合意に基づくコミュニティベースのファクトチェックは、ソーシャル メディア プラットフォーム上で急速に拡大しています。しかし、人間の貢献者によって評価されるコミュニティのクロスコンセンサスファクトチェックの遅れと比率の低さは、依然として大きな課題です。これに対処するために、私たちはまず ComRate を作成しました。これは、$\mathbb{X}$ をソースとする 250 万件のコミュニティ ノートと 2 億 900 万件を超える評価で構成される大規模なデータセットです。次に、コミュニティノート評価のための、ペルソナに基づいたマルチエージェント評価フレームワークである MultiCom を提案します。 MultiCom は、マトリックス因数分解された評価者空間で投稿者をクラスタリングし、公式のコミュニティ ノート評価スキーマに基づいて構造化された評価を生成するようにペルソナ エージェントを促すことによって、多様な評価者集団をシミュレートします。これらのエージェントは、自信、同意のシグナル、理由など、構造化された説明可能な判断を出力します。フォールド外で調整された集計アルゴリズムは、生の投票や診断理由シグナルなどの機能を組み合わせて、信頼性の高い予測を実現します。広範な評価により、MultiCom が他の手法より優れたパフォーマンスを示し、評価セットで平均精度 84.7% (バランス精度 68.3%、マクロ F1 60.1%) を達成していることが実証されています。

原文 (English)

Towards Multi-Agent-Simulation-Based Community Note Evaluation

Community-based fact-checking that relies on cross-consensus is expanding rapidly on social media platforms. However, the delay and low-ratio of cross-consensus community fact-checks rated by human contributors remains a significant challenge. To address this, we first created ComRate, a large-scale dataset comprising 2.5 million community notes and over 209 million ratings sourced from $\mathbb{X}$. We then propose MultiCom, a persona-guided multi-agent rating framework for community note evaluation. MultiCom simulates diverse rater population by clustering contributors in a matrix-factorized rater space and prompting persona agents to generate structured assessments based on the official community notes rating schema. These agents output structured and explainable judgments, such as confidence, agreement signals and reasons. An out-of-fold calibrated aggregation algorithm combines features such as raw votes and diagnostic reason signals for reliable prediction. Extensive evaluations demonstrate that MultiCom outperforms alternative methods, achieving an average accuracy of 84.7% (balanced accuracy 68.3%, macro-F1 60.1%) on the evaluation set.

2026-06-18 13:00 JSTarXiv cs.AIビジネス/資金調達

Vibecoding Ate My 宿題: グリーンフィールド ソフトウェア エンジニアリングとプログラミングへの AI アプローチの評価

生成 AI の急速な発展のおかげで、私たちはコンピューターとの対話方法を永遠に変える可能性のあるパラダイム シフトの真っ只中にいます。この分野の基礎知識なしにアプリケーションやコーディング インフラストラクチャを構築するための自然言語プロンプトの使用が増加していることが観察されており、この実践は「バイブ コーディング」と呼ばれています。これはおそらく、プログラミングの分野が当初から、考えられるあらゆるより高い抽象化レベルで構築されてきたものを表しています。 Vibe コーディングは、入力方法に関する限り、高レベル プログラミングのメタのエンドポイントとなることが約束されています。つまり、人間によるコード構文の使用が完全に排除され、母国語でのプログラミングが優先されます。このペーパーは、グリーンフィールドのソフトウェア エンジニアリング タスクにおける Vibe コーディングの実現可能性を評価し、そのソフトウェア エンジニアリングの能力を測定するために使用されたベンチマークを分析することを目的としています。この目的を達成するために、私たちは、Python で単純で個別のグリーンフィールド プログラミング タスクを実行する LLM の習熟度を分析し、この問題に関する範囲を絞った洞察を提供するための評価スイートを開発しました。

原文 (English)

Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming

Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forever. We have observed a growth in the use of natural language prompts to build applications and coding infrastructures without underlying knowledge of the field, and this practice has been dubbed `vibe coding.' It arguably represents what the field of programming has been building towards since the beginning, with every higher level of abstraction that is conceived. Vibe coding promises to be the endpoint for the meta of high-level programming as far as method of input is concerned: eliminating a human's use of code syntax entirely in favour of programming in their mother tongue. This paper aims to evaluate the viability of vibe coding for greenfield software engineering tasks, as well as analyse the benchmarks that have been used to measure its software engineering prowess. To this end, we have developed an evaluation suite for analysing an LLM's proficiency in carrying out simple, isolated greenfield programming tasks in Python to provide scoped insight on the matter.

2026-06-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

アドヒアランスの向上、より豊かなコンテキスト: LLM を利用した睡眠用会話音声日記のフィールド評価

睡眠日誌は睡眠行動医学や不眠症の認知行動療法の中心ですが、毎日の完了を維持するのは難しく、静的な形式では夜間の睡眠の変化を解釈するためのコンテキストが限られていることがよくあります。私たちは、LLM を利用した会話型音声日記を設計しました。これは、プロアクティブなスマート スピーカーのプロンプト、構造化された会話の取り込み、および適応的なフォローアップ ダイアログを通じて、臨床に基づいた朝と夜の睡眠日記の質問を提供します。私たちは、30 人の大学生を対象とした 4 週間の科目間のフィールド調査でシステムを評価し、一致する日記項目、レポートウィンドウ、リマインダー間隔を使用してテキストベースのモバイル日記と比較しました。テキストベースの日記と比較して、会話音声日記は高い遵守率を示し、日課、ストレス要因、環境条件、その他の睡眠関連要因についてより詳細な状況に応じた自己報告を引き出しました。参加者はまた、音声日記は完了までに時間がかかるにもかかわらず、日常生活に組み込むのが簡単であると述べました。ただし、音声ベースの会話による取り込みでは、一部の構造化日記フィールドの完全性が低くなり、表現力の豊かさと構造化の正確さとの間にトレードオフがあることが明らかになりました。これらの調査結果は、LLM を利用した会話型音声アシスタントを長期的な健康自己報告に使用することの可能性と課題の両方を示しています。

原文 (English)

Better Adherence, Richer Context: A Field Evaluation of LLM-Powered Conversational Voice Diaries for Sleep

Sleep diaries are central to behavioral sleep medicine and cognitive behavioral therapy for insomnia, yet daily completion is difficult to sustain, and static forms often provide limited context for interpreting night-to-night sleep variation. We designed an LLM-powered conversational voice diary that delivers clinically grounded morning and evening sleep diary questions through proactive smart-speaker prompts, structured conversational intake, and adaptive follow-up dialogue. We evaluated the system in a four-week between-subjects field study with 30 university students, comparing it with a text-based mobile diary using matched diary items, reporting windows, and reminder intervals. Compared with the text-based diary, the conversational voice diary showed higher adherence and elicited more detailed contextual self-report about routines, stressors, environmental conditions, and other sleep-related factors. Participants also described the voice diary as easier to integrate into daily routines, despite longer perceived completion time. However, voice-based conversational intake produced lower completeness for some structured diary fields, revealing a trade-off between expressive richness and structured precision. These findings show both the promise and the challenge of using LLM-powered conversational voice assistants for longitudinal health self-report.

2026-06-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM は医師を支援する準備ができていますか?インタラクティブな医師、患者、EHR 支援のための PhysAssistBench

医療 LLM の最も妥当な短期的な役割は、医師の代わりではなく支援することですが、現在の評価では、臨床知識、EHR システムの相互作用、患者とのコミュニケーションなど、個別の能力がテストされることがよくあります。代わりに、医師の支援には同じ対話内でこれらの機能を調整する必要があり、医師は不明確な要求を発行し、患者は症状を曖昧に説明し、EHR システムはツールの正確な使用を要求します。インタラクティブな医師、患者、EHR 支援のベンチマークである PhysAssistBench を紹介します。実際の MIMIC-IV 症例から構築された PhysAssistBench は、スケーラブルなパイプラインを使用してエージェント性患者を構築します。これは、臨床上の事実を維持しながら、静的な EHR 記録を複数ターンの臨床シナリオに変換する、インタラクティブで記録に基づいたエージェントです。 PhysAssistBench は、手動でレビューされ医師が検証した 1,296 ターンの厳選されたバイリンガル評価セットを提供します。主要な LLM を使った実験では、この設定では現在のモデルの信頼性が依然として低いことが示されており、臨床 LLM にとって重要なボトルネックが露呈しています。信頼できる支援には、知識、コミュニケーション、システム全体の調整が必要であり、それらのいずれかで単独の利益を得るのではありません。

原文 (English)

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.

2026-06-18 13:00 JSTarXiv cs.AIビジネス/資金調達

AI を活用した人間の家庭教師の評価: トレーニングのパフォーマンスを実際の実践に結びつける

家庭教師トレーニング プラットフォームは数多く存在します。しかし、実際のパフォーマンスに基づいて人間の家庭教師に AI 主導のトレーニングと評価を提供しているところはほとんどありません。私たちは、トレーニング中のオープンな回答と本物の現実の個別指導の両方を評価する AI 主導のシステムを紹介します。オンライン トレーニングやシミュレーションを通じてのみ学習を評価するプラットフォームとは異なり、当社のシステムは生成 AI (Gemini-2.5-pro) を利用して本物の個別指導の文字起こしを分析し、講師のスキルの実際の応用への移行を測定します。生徒に数学をリモートで指導する人間の家庭教師 (N=86) は 6 つのシナリオベースのレッスンを完了し、平均して 7.4% の大幅な学習向上を達成しました。 405 のセッションとレッスンのペアにわたる混合効果モデルを使用したところ、トレーニングのパフォーマンスが効果量 0.25 SD で実際の成績証明書スコアを有意に予測することがわかりました。モデル比較 (AIC/BIC) では、トレーニング中の自由応答と多肢選択のパフォーマンスを平均すると、実際の家庭教師のパフォーマンスが最もよく予測されることが示されましたが、自由応答の方が比較的予測性が高かったです。探索的分析の結果、トレーニング後、家庭教師はスキルを応用するための教育的機会に遭遇する可能性が大幅に高く (61.1% ~ 68.9%)、その機会内での実践の質がより高いことが実証されました (65.5% ~ 68.1%)。中断された時系列分析は、これらの家庭教師の改善は、トレーニングの即時介入効果ではなく、時間の経過とともに徐々に進む傾向の一部であることを示唆しました。家庭教師のトレーニングと実際の評価を結び付ける AI 主導の方法を説明します。その際、透明性と再現性をサポートするために、オープン データセット、AI プロンプト、スコアリング ルーブリックを提供します。

原文 (English)

AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice

There exist numerous tutor training platforms. However, few provide AI-driven training and evaluation for human tutors based on real-life performance. We present an AI-driven system that assesses both open responses during training and authentic real-life tutoring. Unlike platforms that only assess learning through online training or simulations, our system utilizes Generative AI (Gemini-2.5-pro) to analyze transcriptions of authentic tutoring, measuring the transfer of tutor skills to real-life application. Human tutors instructing students remotely in math (N=86) completed six scenario-based lessons, averaging a significant 7.4% learning gain. Using mixed-effects models across 405 session-to-lesson pairs, we found that training performance significantly predicted real-life transcript scores with an effect size of 0.25 SD. Model comparison (AIC/BIC) indicated averaging open response and multiple choice performance during training predicted real-life tutor performance best, although open responses were comparatively more predictive. Exploratory analysis showed that after training, tutors were significantly more likely to encounter pedagogical opportunities to apply their skills (61.1% to 68.9%) and demonstrated higher execution quality within those opportunities (65.5% to 68.1%). Interrupted time series analysis suggested that these tutor improvements were part of a gradual trend over time rather than an immediate intervention effect of training. We illustrate an AI-driven method to link tutor training with real-life assessment. In doing so, we contribute open datasets, AI prompts, and scoring rubrics to support transparency and reproducibility.

2026-06-18 13:00 JSTarXiv cs.AIビジネス/資金調達

A Clinician-Centered Pipeline for Annotation and Evaluation in Ultrasound AI Studies

Clinician-centered evaluation is critical for validating medical AI systems, especially in ultrasound imaging where quantitative metrics do…

2026-06-18 13:00 JSTarXiv cs.AIビジネス/資金調達

国家学習能力としての AI 主権: フランス、米国、中国に関する人間中心の学習力学の視点

フランスでは、人工知能は、投資、計算能力、規制、雇用、主権、教育の観点からよく議論されます。通常、これらのディメンションは個別に扱われます。この観点に関する論文は、統一的な解釈を提案しています。つまり、フランスは \emph{国家的な AI 学習システム} として理解されるべきです。エントロピー制御された表現学習のための動的フレームワークとして最近策定された人間中心学習力学 (HCLM) に基づいて、私たちは国家 AI 開発を情報注入とエントロピー散逸の間の制御されたバランスとして解釈します。情報注入は、コンピューティング、データ、人材、研究、資本、産業展開、および組織的実験に対応します。エントロピー散逸は、組織の複雑さ、調整摩擦、エネルギー制約、規制の不確実性、人材の流動性の圧力、産業吸収を強化する機会に対応します。中心的な主張は、AI の主権は規模だけから生まれるのではなく、自国の情報ダイナミクスを規制する国の能力から生まれるというものです。この論文は、HCLM をニューラル スケーリング則、内生的成長理論、創造的破壊、およびゲーム理論と結びつけます。同論文は、フランスのAI論争は、技術楽観主義と規制優先の慎重論という二項対立を超えて進むべきだと主張している。競争力のある人間中心の AI 戦略には、不安定、不平等、またはエネルギー集約的な拡大を回避しながら、情報注入が制度的消散よりも早く成長する制御された体制が必要です。私たちは、数学的モデル、測定可能な政策指標、ゲーム理論的命題、国家 AI 体制の具体的なシミュレーション、およびフランスに対する具体的な政策への影響を提供します。提案された視点は、AI 政策をオープンで戦略的な非平衡学習システムのガバナンスとして再構成します。

原文 (English)

AI Sovereignty as National Learning Capacity: A Human-Centered Learning Mechanics Viewpoint on France, the United States, and China

Artificial intelligence in France is often discussed through separate dimensions such as investment, compute, regulation, employment, sovereignty, and education. This viewpoint paper proposes a unified interpretation: France can be analyzed as a national AI learning system. Building on Human-Centered Learning Mechanics (HCLM), we use HCLM not as a validated econometric model, but as a conceptual and diagnostic lens for interpreting national AI development as a balance between information injection, absorptive capacity, and institutional dissipation. Information injection includes compute, data, talent, research, capital, industrial deployment, and policy experimentation. Institutional dissipation refers to avoidable frictions such as administrative overload, coordination failures, energy constraints, regulatory uncertainty, talent mobility pressures, and weak industrial absorption. Regulation is not treated as mere friction: adaptive governance, trusted data spaces, and safety-oriented standards may increase long-term learning capacity by improving legitimacy, interoperability, and social trust. The central claim is not that a country follows neural-network equations, but that AI sovereignty depends on how effectively it converts distributed information into absorbed, coordinated, and socially legitimate capability. The paper connects HCLM with neural scaling laws, endogenous growth theory, creative destruction, absorptive capacity, and coordination mechanisms. It offers a formal heuristic, policy indicators, illustrative scenarios, and implications for France. The numerical results are diagnostic scenarios, not econometric estimates or official rankings. The proposed viewpoint reframes AI policy as the governance of an open, strategic, non-equilibrium learning system that should be tested with historical and cross-country data.

2026-06-18 13:00 JSTarXiv cs.AIビジネス/資金調達

Quality Perceptions and Intended Engagement in Response to AI-Generated and AI-Assisted News

The increasing use of artificial intelligence (AI) in news production raises important questions about how audiences perceive and respond t…

2026-06-18 13:00 JSTarXiv cs.AIビジネス/資金調達

Scalable Batch Bayesian Optimization Via Subspace Acquisition Functions

Extending Bayesian optimization to batch evaluation can enable the designer to make the most use of parallel computing technology. However,…

2026-06-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark

Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization…

2026-06-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation

Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or s…

2026-06-18 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection

Instruction-tuned LLMs can annotate thousands of instances at low cost. This raises two questions for active learning (AL): can LLM labels…

2026-06-18 13:00 JSTarXiv cs.AIビジネス/資金調達

SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning

Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…

2026-06-18 05:32 JSTTechCrunch AIビジネス/資金調達

Roelof Botha joins SpaceX’s board of directors

The former Sequoia Capital leader is filling an "existing vacancy" on SpaceX's board, days after the company went public in the largest IPO…

2026-06-18 04:01 JSTTechCrunch AIビジネス/資金調達

World leaders want American AI. They just don’t want America to be able to turn it off.

French President Macron and Indian PM Modi raised alarms at the G7 summit that the U.S. could cut off access to American AI overnight — a f…

2026-06-18 03:00 JSTTechCrunch AIエージェントビジネス/資金調達

NEA’s Tiffany Luck on AI IPOs, personal agents, and the ROI reckoning

Tokenmaxxing was the hottest trend in Silicon Valley earlier this year, with CEOs encouraging employees to push AI usage as far as it would…

2026-06-18 02:43 JSTTechCrunch AILLM/生成AIビジネス/資金調達

World model maker Odyssey nabs $1.45B valuation backed by Amazon and other big names

World models are the next big thing in AI beyond LLMs and, with this round, Odyssey has cemented itself as one of the startups to watch.

2026-06-17 23:15 JSTTechCrunch AIビジネス/資金調達

Pramaana Labs raises $27M seed round from Khosla Ventures to bring formal verification to AI

Pramaana will focus on highly sensitive verticals like law, drug discovery, and tax preparation — where errors can be costly and reliabilit…

2026-06-17 21:14 JSTTechCrunch AIビジネス/資金調達

DeepL acquires Mixhalo for live-event audio streaming and translation

With this acquisition, DeepL is opening an office in San Francisco to expand its U.S. business.

2026-06-17 17:51 JSTITmedia AI+ビジネス/資金調達
2026-06-17 17:44 JSTITmedia AI+ビジネス/資金調達

SpaceX、AIコーディング「Cursor」を9.6兆円で買収 「近く大幅な改善」へ

Corsorは公式Xで、「近く大幅な改善が行われる予定だ」と述べた。

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

SEAGym: 自己進化する LLM エージェントの評価環境

自己進化する LLM ベースのエージェントは、主にエージェント ハーネス (プロンプト、メモリ、ツール、ミドルウェア、ランタイム状態、モデルとツールの対話ループなどの基本モデルを中心とした構造化された実行層) を変更することによって改善されます。既存の評価では、多くの場合、このプロセスが個別のタスク スコアまたは単一の連続曲線に縮小され、更新によって再利用可能な改善がもたらされるのか、最近のタスクに過剰適合するのか、コストが増加するのか、古い動作に害を及ぼすのかが不明瞭になります。トレーニング、検証、テスト、再生、コスト記録にわたるエージェント ハーネスの更新を測定するための評価環境である SEAGym を紹介します。 SEAGym は、Harbor 互換ベンチマークを、トレイン バッチ、凍結された更新検証、保持された ID および OOD 転送ビュー、再生診断、および保存されたスナップショットとメトリック レコードを備えた動的な自己進化タスク ソースに変換します。 Terminal-Bench 2.0 および HLE 上で SEAGym をインスタンス化し、共有エポック/バッチ プロトコルの下で ACE、TF-GRPO、および AHE を比較します。結果は、これらの評価ビューが進化プロセスに関する補完的なシグナルを提供することを示しています。つまり、頻繁な更新では持続的なパフォーマンスが向上しない可能性があり、有用な中間スナップショットが後で崩壊する可能性があり、ソースの多様性とモデル バックエンドがハーネスの信頼性に影響を与える可能性があります。

原文 (English)

SEAGym: An Evaluation Environment for Self-Evolving LLM Agents

Self-evolving LLM-based agents improve mainly by changing their agent harness: the structured execution layer around a base model, including prompts, memory, tools, middleware, runtime state, and the model-tool interaction loop. Existing evaluations often reduce this process to isolated task scores or a single sequential curve, obscuring whether an update produces reusable improvement, overfits recent tasks, increases cost, or harms older behavior. We introduce SEAGym, an evaluation environment for measuring agent harness updates across training, validation, test, replay, and cost records. SEAGym turns Harbor-compatible benchmarks into dynamic self-evolution task sources with train batches, frozen update-validation, held-out ID and OOD transfer views, replay diagnostics, and saved snapshot and metric records. Instantiating SEAGym on Terminal-Bench 2.0 and HLE, we compare ACE, TF-GRPO, and AHE under a shared epoch/batch protocol. The results show that these evaluation views provide complementary signals about the evolution process: frequent updates may fail to improve held-out performance, useful intermediate snapshots may collapse later, and source diversity and model backend can affect harness reliability.

2026-06-17 13:00 JSTarXiv cs.AIビジネス/資金調達

DeepInsight: 物理 AI スタック全体にわたる統合評価インフラストラクチャ

物理 AI スタックの評価には、単一の基礎モデルのデコード ステップから全身制御の数千の物理ティックまで、モダリティ、報酬セマンティクス、リソース プロファイルが直交して変化する、3 桁以上異なるオペレーターが含まれます。この範囲に及ぶ既存のフレームワークはないため、現在スタックは、ランタイムもスコアリングも共有しない個別のハーネスをつなぎ合わせて評価されており、各セグメントのローカル妥当性は維持されますが、クロスレイヤー回帰を診断するために必要な共有アイデンティティは失われます。ここでは、単一のランタイムでこの全領域にサービスを提供する評価インフラストラクチャである DeepInsight を紹介します。レジームを均質化するのではなく、タスク、リソース、結果という 3 つの狭い抽象化の背後にある異質性を維持します。各抽象化は、すべてのサブシステムによって共有される 1 つの不変式として実現されます。つまり、1 つのエピソード ドライバー、すべての高価なバックエンド (LLM 推論とサンドボックス化されたランタイムは同様) によって実装される 1 つのリソース ハンドル プロトコル、およびすべてのイベントが書き込まれる 1 つのトレース ID スキームです。この単一セットの不変条件は、実体化されたヒューマノイド スタックの 3 つのレイヤーすべてにわたって運用環境にデプロイされ、主に構成によって新しいベンチマークをオンボードします。成熟したピア オーケストレーターが存在する場所 (基盤モデルの末端) では、公開されたリファレンスとピア フレームワークの読み取り値を独自のスプレッド内で再現し、単一ノード上で同じスイートをより高速に実行し、ノード間でほぼ線形にスケールします。その特徴的な戻りは診断用です。すべてのレイヤーが 1 つの共有トレースに書き込むため、あるレイヤーで始まり別のレイヤーで表面化する回帰は、そのトレース上で局所化されたままになります。これは、セグメントごとのハーネスのフェデレーションでは再現できない層間の利益です。

原文 (English)

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack

Evaluating a Physical AI stack spans operators that differ by more than three orders of magnitude -- from a single foundation-model decoding step to thousands of physics ticks of whole-body control -- varying orthogonally in modality, reward semantics, and resource profile. No existing framework spans this range, so the stack is evaluated today by stitching together separate harnesses that share neither runtime nor scoring, preserving each segment's local validity but losing the shared identity needed to diagnose cross-layer regressions. We present DeepInsight, an evaluation infrastructure that serves this full spectrum on a single runtime. Rather than homogenize the regimes, it preserves their heterogeneity behind three narrow abstractions -- task, resource, and result -- each realized as one invariant shared by every subsystem: one episode driver, one resource-handle protocol implemented by every expensive backend (LLM inference and sandboxed runtimes alike), and one trace identity scheme under which every event is written. Deployed in production across all three layers of an embodied humanoid stack, this single set of invariants onboards new benchmarks largely by configuration. Where mature peer orchestrators exist -- at the foundation-model end -- it reproduces published references and peer-framework readings within their own spread, runs the same suites faster on a single node, and scales near-linearly across nodes. Its distinctive return is diagnostic: because every layer writes into one shared trace, a regression that begins in one layer and surfaces in another stays localizable on that trace -- a cross-layer payoff no federation of per-segment harnesses can reproduce.

2026-06-17 13:00 JSTarXiv cs.AIビジネス/資金調達

LongWebBench: 長期的な設定での構造的および機能的な Web ページ生成の評価

最近のビジョン言語モデル (VLM) は、視覚入力から Web ページを生成する点で有望な進歩を示していますが、既存の評価は主に、短く、単一画面で、大部分が静的な Web ページに焦点を当てています。構造的および機能的観点の両方から長期的な Web ページの生成を評価するためのベンチマークである LongWebBench を紹介します。 LongWebBench には、構造忠実度評価用の 490 の実世界の長い Web ページと、機能評価用の 129 Web ページにわたる 507 の目標指向のインタラクション タスクが含まれています。これは、2 つの相補的なプロトコルを採用しています。1 つは長距離の構造的一貫性を評価するための多次元 VLM ベースのメトリック、もう 1 つはエンドツーエンドの機能検証のための DOM 拡張エージェントベースのパイプラインです。人間の一致分析を通じて自動評価プロトコルをさらに調査します。単一画像および複数画像の設定下で最先端のオープンソースおよび独自の VLM を使用した実験では、Web ページの長さが増加するにつれて構造の忠実度が低下する一方で、視覚的に妥当な世代では実行可能な複数ステップのインタラクションをサポートできないことが多いことが明らかになりました。これらの結果は、実行可能な相互作用を中心的な基準として、視覚的な類似性を超えて長い Web ページの生成を評価する必要性を強調しています。コードとデータは https://github.com/zheny2751-dotcom/LongWebBench で入手できます。

原文 (English)

LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings

Recent vision-language models (VLMs) have shown promising progress in generating webpages from visual inputs, yet existing evaluations mainly focus on short, single-screen, and largely static webpages. We introduce LongWebBench, a benchmark for evaluating long-horizon webpage generation from both structural and functional perspectives. LongWebBench contains 490 real-world long webpages for structural fidelity evaluation and 507 goal-oriented interaction tasks over 129 webpages for functional evaluation. It employs two complementary protocols: a multi-dimensional VLM-based metric for assessing long-range structural coherence, and a DOM-augmented agent-based pipeline for end-to-end functional verification. We further examine the automatic evaluation protocols through human agreement analysis. Experiments with state-of-the-art open-source and proprietary VLMs under single-image and multi-image settings reveal that structural fidelity degrades as webpage length increases, while visually plausible generations often fail to support executable multi-step interactions. These results highlight the need to evaluate long webpage generation beyond visual similarity, with executable interaction as a core criterion. Our code and data are available at https://github.com/zheny2751-dotcom/LongWebBench.

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

推論によるフロンティア LLM 評価の形状計算方法

AI の評価は、ツールの使用と反復的な問題解決を伴う長期にわたる軌道から恩恵を受ける、より困難なタスクへと移行しています。その結果、パフォーマンスは、テスト時に利用可能なコンピューティング (「推論コンピューティング」) の量と割り当てにますます敏感になります。しかし、多くの評価では依然として単一の制限された予算でのパフォーマンスが報告されており、低いスコアはモデルの基礎的な機能ではなく評価設定を反映している可能性があることを意味します。これをテストするために、ソフトウェア エンジニアリング、数学、医学、サイバーセキュリティにわたる 7 つの挑戦的なベンチマークで最大 12 のフロンティア言語モデルを評価します。私たちは、3 つの単純な推論スケーリング介入を組み合わせた制御されたセットアップを使用します。つまり、より大きなトークン バジェット、コンテキストの圧縮、およびモデル自体または最小限の正確性フィードバックによって導かれる送信の試行の繰り返しです。主な結果は 3 つあります。まず、トークン バジェットが大きくなると、サイバーセキュリティ、FrontierMath、人類最後の試験、ターミナルベンチなど、複数のドメインにわたるベンチマークのパフォーマンスが大幅に向上します。第二に、固定予算の評価では、モデルが進歩するにつれてフロンティアの能力がますます過小評価される可能性があります。新しいモデルは、大きな予算でより高いパフォーマンスを実現し、より困難なタスクを解放し、より確実に解決します。第三に、どの推論スケーリング手法が最も役立つかがベンチマークによって異なります。繰り返し送信するとパフォーマンスが大幅に向上しますが、より大きなトークン バジェット、外部フィードバック、および並列試行の値はベンチマークによって異なります。全体として、私たちの結果は、ベンチマーク スコアがプロトコルに依存していることを示しています。したがって、評価では、特に安全性またはポリシー関連の設定において、推論時間のコンピューティングの関数として機能を報告し、プロトコルの選択を明示的に指定し、一致した予算で大規模な共有コンピューティング範囲にわたってモデルの世代を比較する必要があると主張します。

原文 (English)

How Inference Compute Shapes Frontier LLM Evaluation

AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still report performance at a single restrictive budget, meaning that low scores may reflect the evaluation setup rather than the model's underlying capability. To test this, we evaluate up to 12 frontier language models on seven challenging benchmarks spanning software engineering, mathematics, medicine, and cybersecurity. We use a controlled setup combining three simple inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts, guided either by the model itself or by minimal correctness feedback. We find three main results. First, larger token budgets substantially improve performance on benchmarks across multiple domains, including cybersecurity, FrontierMath, Humanity's Last Exam, and TerminalBench. Second, fixed-budget evaluations can increasingly understate frontier capability as models advance. Newer models reach higher performance at large budgets, where they unlock harder tasks and solve them more reliably. Third, benchmarks differ in which inference-scaling methods help most: repeated submission broadly improves performance, but the value of larger token budgets, external feedback, and parallel attempts varies by benchmark. Overall, our results show that benchmark scores are protocol-dependent. We therefore argue that evaluations should report capability as a function of inference-time compute, specify protocol choices explicitly, and compare model generations over a large shared compute range at matched budgets, especially in safety- or policy-relevant settings.

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

新しい AI アクセラレータでの LLM 推論のプレフィル/デコードを意識した評価

大規模言語モデル (LLM) がレイテンシとコストに敏感な設定で導入されることが増えているため、推論効率がシステムの中心的な課題となっています。現在の導入では GPU が主流ですが、LLM 推論に利点があると主張する AI アクセラレータが増えていますが、実際にどのような条件下でそのようなアクセラレータが GPU よりも優れたパフォーマンスを発揮するかは依然として不明です。最近の推論システムは、実行をプレフィル フェーズとデコード フェーズに分解します。これらのフェーズは、異なる計算特性とレイテンシ メトリクスを示し、通常、最初のトークンまでの時間 (TTFT) と出力トークンあたりの時間 (TPOT) によってキャプチャされます。このペーパーでは、共通モデル Llama2-7B を使用した、GPU および新興 AI アクセラレータにわたる LLM 推論パフォーマンスのフェーズを意識した評価を示します。プレフィルとデコードのパフォーマンスを個別に測定することで、アクセラレータの利点がフェーズとメトリックによって異なることが明らかになりました。私たちの結果は、GPU がコンピューティング集中型のプレフィル フェーズで一貫して優れているのに対し、GroqRack はデコード中に大幅に低い TPOT を達成することを示しています (バッチ処理は現在サポートされていません)。ただし、バッチ サイズが増加するにつれて、GPU はデコード スループットで優位性を取り戻します。これらの発見は、各プラットフォームが相に応じた異なる強みを示すことを示しています。さらに、さまざまなアクセラレータ プラットフォームにわたる異種のプリフィル/デコードの分解を分析し、パフォーマンスの向上と、そのような向上が実現されるワークロードとネットワークの条件を特定します。

原文 (English)

Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators

As large language models (LLMs) are increasingly deployed in latency- and cost-sensitive settings, inference efficiency has become a central systems challenge. While GPUs dominate current deployments, a growing number of AI accelerators claim advantages for LLM inference, yet it remains unclear under which conditions such accelerators outperform GPUs in practice. Recent inference systems decompose execution into Prefill and Decode phases, which exhibit distinct computational characteristics and latency metrics, commonly captured by time to first token (TTFT) and time per output token (TPOT). This paper presents a phase-aware evaluation of LLM inference performance across GPUs and emerging AI accelerators using a common model, Llama2-7B. By separately measuring Prefill and Decode performance, we reveal that accelerator advantages differ by phase and metric. Our results show that GPUs consistently excel in the compute-intensive Prefill phase, while GroqRack achieves significantly lower TPOT during Decode (batching not currently supported). However, GPUs regain an advantage in Decode throughput as batch size increases. These findings demonstrate that each platform exhibits distinct phase-dependent strengths. We further analyze heterogeneous Prefill/Decode disaggregation across different accelerator platforms, identifying performance gains and the workload and network conditions under which such gains are realized.

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

現実的なシナリオにおける LLM エージェントを使用するツールにおけるデータ漏洩リスクの評価

AI エージェントは、電子メール、データベース、ドキュメント、その他のツールにアクセスして機密情報の読み取り、更新、配布を行うことができるため、企業や個人の環境で採用されることが増えています。エージェントにおけるデータ漏洩リスクに関するこれまでの研究の多くは、迅速なインジェクションとジェイルブレイクによる敵対的なデータ漏洩に焦点を当てていました。ただし、機密情報は非敵対的な使用中にも公開される可能性があり、ユーザーが無害なリクエストを発行した場合でも漏洩のリスクが生じます。私たちは、顧客サポート、DevOps、Web 自動化、企業および個人の生産性にわたる 12 の現実的で非敵対的なタスクにおけるエージェントのデータ漏洩を調査した、シンガポール AI 安全性研究所と韓国 AI 安全性研究所による共同評価を報告します。この評価では、データ認識の欠如、視聴者認識、ポリシー遵守、データ最小化、アクセス境界認識の 5 つのリスク タイプが対象となります。両機関は、独立したテスト環境とタスク固有の LLM 判定ルーブリックを使用して、現実世界の展開を反映した共通のシナリオ セットをテストしました。テストされた 3 つのエージェントのうち、すべてのシナリオにわたって完全に正しく、完全に安全な実行を達成したエージェントはありませんでした。タスクの正常な完了は、不要な情報へのアクセスや不適切な受信者への情報の開示などのデータ処理の失敗と同時に発生することが多く、能力とデータ処理の安全性は別々に評価される必要があることを示しています。定性的レビューでは、クレームとアクションの不一致、シミュレーションを意識した動作、ユーザーとシミュレーターの役割の逆転、自動判定における解釈のギャップも明らかになりました。全体として、この結果は、運用データの漏洩は、敵対的なデータ漏洩とは異なるエージェントの安全性に関する第一次の懸念事項であり、エージェントのデータ処理の安全性を将来評価するための方法論を提供することを示しています。

原文 (English)

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios

AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they can read, update, and disseminate sensitive information. Much of prior research on data leakage risks in agents has focused on adversarial data exfiltration through prompt injections and jailbreaks. However, sensitive information may also be exposed during non-adversarial use, creating leakage risks even when users issue benign requests. We report a joint evaluation by the Singapore AI Safety Institute and the Korea AI Safety Institute examining agent data leakage in 12 realistic, non-adversarial tasks spanning customer support, DevOps, web automation, and enterprise and personal productivity. The evaluation covers five risk types: lack of data awareness, audience awareness, policy compliance, data minimization, and access-boundary awareness. Both institutes tested a common set of scenarios mirroring real-world deployments using independent testing environments and task-specific LLM-judge rubrics. Across the three tested agents, none achieved fully correct and fully safe execution across all scenarios. Successful task completion often coincided with data-handling failures such as accessing unnecessary information or disclosing information to inappropriate recipients, indicating that capability and data-handling safety should be evaluated separately. Qualitative review also revealed claim-action mismatches, simulation-aware behavior, user-simulator role reversal, and interpretation gaps in automated judging. Overall, the results indicate that operational data leakage is a first-order agent-safety concern distinct from adversarial exfiltration and provide a methodology for future evaluations of agent data-handling safety.

2026-06-17 13:00 JSTarXiv cs.AIビジネス/資金調達

プロービング、融合、および信頼性: 多峰性がん解析のための基礎モデル表現の系統的評価

ファウンデーション モデル (FM) は、医療データの強力な表現抽出ツールとして登場しましたが、分布の変化に伴うデータセットへの一般化可能性はまだ研究されていません。この研究では、認可された社内 (IH) 腫瘍学データセットから抽出された 2 つの実際の商用コホート、IH-BC および IH-NSCLC にわたる一連の計算病理学タスクに関する FM ベースの表現を体系的に評価します。分析は、IH マルチモーダル データから抽出された 2 つのモダリティ、スライド全体の画像とトランスクリプトーム プロファイルに焦点を当てています。まず、8 つの下流分類タスクで 5 つの FM にわたる単峰性プローブのパフォーマンスをベンチマークし、画像とオミクス表現が相補的な予測信号を伝達することを発見しました。次に、ペア表現に基づいて構築された 3 つのイメージ-オミクス融合戦略を比較することにより、マルチモーダル融合が単峰性ベースラインを超える追加の利益を生み出すことができるかどうかを調査します。選択された単峰性パイプラインおよびマルチモーダル パイプラインの信頼性は、等角予測によってさらに評価されます。私たちの結果は、FM 表現が分布外データで競争力のあるパフォーマンスを実現し、マルチモーダル融合が主に単一のモダリティが信号を支配しない場合に役立つことを示しています。コンフォーマル予測は、点予測が失敗した場合のほとんどの場合、予測セット内で真の診断が回復可能であることを明らかにし、臨床サポートにおける不確実性を考慮した推論の価値を強化します。

原文 (English)

Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored. This work systematically evaluates FM-based representations on a suite of computational pathology tasks across two real-world commercial cohorts, IH-BC and IH-NSCLC, drawn from the licensed in-house (IH) oncology dataset. The analysis focuses on two modalities, whole-slide images and transcriptomic profiles, drawn from the IH multimodal data. We first benchmark unimodal probing performance across five FMs on eight downstream classification tasks, and find that image and omics representations carry complementary predictive signals. Then we investigate whether multimodal fusion can yield additional gains over unimodal baselines by comparing three image-omics fusion strategies built on paired representations. The trustworthiness of selected unimodal and multimodal pipelines is further assessed through conformal prediction. Our results show that FM representations achieve competitive performance on out-of-distribution data and that multimodal fusion helps mainly when no single modality dominates the signal. Conformal prediction reveals that in the majority of cases where a point prediction fails, the true diagnosis remains recoverable within the prediction set, reinforcing the value of uncertainty-aware inference for clinical support.

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Rift: 言語モデルにおける欺瞞の競合署名

真実を知りながら嘘をつくモデルは、ELK が行動評価だけでは処理できない中心的なケースです。私たちは、そのような欺瞞が、それを正直な誤りと区別する内部の痕跡を残すかどうかを尋ねます。私たちの重要な動きは不正を制御することです。スリーパーエージェント(真実を知っており、トリガーに従って嘘をつく)と素朴な嘘つき(正直な訓練を受けずに同じ間違った答えを返すように微調整されている)を対比させます。どちらも同じ間違った出力を生成します。違いは知識の衝突に関するものであり、不正確さではありません。 GPT-2 小規模/中規模 (シード 3 つ) および 3 つの命令モデル全体で、欺瞞的なフォワード パスには競合シグネチャ (同じ間違った答えに対する素朴な嘘つきパスよりも 2.1 ~ 2.3 倍高い残差ランク) が含まれていることがわかりました。2 つの応答のうちどちらが嘘であるかを 100% の精度でラベルなしで識別するのに十分な強度です。 Qwen2.5-1.5B/7B と Phi-3-mini 全体で、指示された欺瞞により、テストされたすべての事実の残留ランクが上昇します (18/18、40/40、34/34)。 Phi-3 では、正直な答えと幻覚の両方から完全に分離されています (AUC 1.0、Wilcoxon p~6e-11)。署名は、戦略的に自己構築された欺瞞(モデルが独自の嘘を発明する、AUC 1.0)、積極的な隠蔽の試み(AUC 1.0)、および長さ制御された複製(20/20、AUC 1.0、p~1e-6)に耐えます。基底なしの相対表現を使用して、1 つのモデル ファミリでトレーニングされたプローブは、他の 2 つのファミリのゼロショット (平均 AUC 0.933)、アーキテクチャと形式の同時変更 (AUC 0.821) での欺瞞を検出し、5 つの言語間での転送 (AUC 1.000、長さ制御) を検出します。シグネチャは読み取り専用です。検出は可能ですが、注入はできません (0/8 両方向)。正直な制限と 6 つの否定的な実験が完全に文書化されています。

原文 (English)

Rift: A Conflict Signature for Deception in Language Models

A model that lies while knowing the truth is the central case ELK cannot handle with behavioral evaluation alone. We ask whether such deception leaves an internal signature distinguishing it from honest error. Our key move is a control for wrongness: we contrast a sleeper agent (knows the truth, lies on trigger) against a naive liar (fine-tuned to emit the same wrong answers with no honest training). Both produce identical wrong outputs; any difference is about knowledge conflict, not incorrectness. We find deceptive forward passes carry a conflict signature - 2.1-2.3x higher residual rank than naive-liar passes on the same wrong answer - strong enough to identify which of two responses is the lie with 100% accuracy and no labels, across GPT-2 small/medium (three seeds) and three instruct models. Across Qwen2.5-1.5B/7B and Phi-3-mini, instructed deception raises residual rank on every tested fact (18/18, 40/40, 34/34); on Phi-3, lies separate perfectly from both honest answers and hallucinations (AUC 1.0, Wilcoxon p~6e-11). The signature survives strategic self-constructed deception (model invents its own lie, AUC 1.0), active concealment attempts (AUC 1.0), and length-controlled replication (20/20, AUC 1.0, p~1e-6). Using basis-free relative representations, a probe trained on one model family detects deception in two other families zero-shot (mean AUC 0.933), surviving simultaneous architecture and format change (AUC 0.821), and transfers across five languages (AUC 1.000, length-controlled). The signature is read-only: detectable but not injectable (0/8 both directions). Honest limitations and six negative experiments are documented in full.

2026-06-17 13:00 JSTarXiv cs.AI画像/動画生成エージェントロボティクスビジネス/資金調達

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models

Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challen…

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation

While Multimodal Retrieval-Augmented Generation (M-RAG) enhances Large Vision-Language Models, it remains highly susceptible to cross-modal…

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows

Large language models (LLMs) are increasingly considered for use in clinical consultation tasks, yet most medical evaluations remain static…

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Unlocking LLM Code Correction with Iterative Feedback Loops

Large Language Models have shown remarkable capabilities in code generation. However, most existing evaluations focus only on single-attemp…

2026-06-17 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Offline Preference-Based Trajectory Evaluation

Offline evaluation of agentic systems often collapses trajectories to terminal success, discarding information about partial progress and i…

2026-06-17 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

Geometric Consistency Protocol for Foundation Model Features in Multi-View Satellite Imagery

Standardized evaluation protocols are indispensable for robust benchmarking in remote sensing, particularly as foundation features are incr…

2026-06-17 13:00 JSTarXiv cs.AIビジネス/資金調達

Embedded Machine Learning for Microcontroller-Class Edge Devices: Data, Feature, Evaluation, and Deployment Pipelines

Embedded machine learning moves inference from cloud services to resource-constrained devices that must acquire data, preprocess signals, r…

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Towards Understanding and Measuring COGNITIVE ATROPHY in LLM Behaviour

Recent incidents involving LLMs used for mental-health support reveal a critical evaluation gap: surface-level safety scores do not capture…

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to alleviate global disparities…

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

In-Context Environments Induce Evaluation-Awareness in Language Models

Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models…

2026-06-17 13:00 JSTarXiv cs.AIビジネス/資金調達

Mental Health AI Safety Claims Must Preserve Temporal Evidence

The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoin…

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LATTEArena: LLM を利用した表形式特徴量エンジニアリングの評価フレームワーク (拡張バージョン)

特徴量エンジニアリングは表形式データ分析にとって依然として不可欠であり、大規模言語モデル (LLM) がこのプロセスを自動化するための有望なパラダイムとして台頭し、LLM を利用した AuTomated 表形式特徴量エンジニアリング (LATTE) が誕生しました。ただし、標準化されたプラットフォームがないため、コストを意識した公平な比較ができません。さらに、複雑な方法論的設計により、個々のコンポーネントの具体的な貢献がわかりにくくなります。たとえば、LFG は思考ツリー、少数ショット デモンストレーション、モンテカルロ ツリー検索、自然言語生成を統合していますが、各技術の競争力による個別の影響は定量化されていません。これらの課題に対処するために、私たちは次の特徴を備えた最初の競争評価フレームワークである LATTEArena を導入します。(1) 15 の代表的な手法を再利用可能なコンポーネントに分解する 6 次元の分類。 (2) 制御された比較のための標準化されたモジュール式アリーナ。 (3) パフォーマンス、コスト、堅牢性をカバーする多次元の評価。 (4) 各技術の競争力を定量化するコンポーネントレベルのアブレーション。広範な評価を通じて、次のような 16 の重要な発見が明らかになりました。(1) モンテカルロ木検索による思考の木は、最適な費用対効果を実現します。 (2) RPN とコードの出力形式は、それぞれ分類タスクと回帰タスクを支配します。私たちはモジュール式フレームワークと 4,000 を超える実行ログを公開し、研究者が新しい技術と既存の技術をシームレスに比較して LATTE を進歩できるようにします。

原文 (English)

LATTEArena: An Evaluation Framework for LLM-powered Tabular Feature Engineering (Extended Version)

Feature engineering remains a cornerstone of tabular data analysis, and Large Language Models (LLMs) have emerged as a promising paradigm for its automation, giving rise to LLM-powered Automated Tabular Feature Engineering (LATTE). However, the field lacks standardized, cost-aware evaluation platforms, and the combinatorial explosion of design choices obscures true algorithmic progress. To bridge these gaps, we systematically deconstruct 15 representative LATTE methods into a unified 6-dimensional taxonomy. Based on this abstraction, we introduce LATTEArena, a standardized, modular, and extensible benchmarking framework that decouples monolithic pipelines into reusable execution blocks. By distilling the massive combinatorial space, we evaluate 24 core LATTE configurations across 7 research questions. Our head-to-head benchmarking goes beyond predictive accuracy to quantify token efficiency and execution robustness, yielding 17 empirical findings on cost-effectiveness trade-offs. Furthermore, we provide 3 concrete recommendations for optimal real-world deployment. By enabling controlled component-level comparisons, LATTEArena shifts the paradigm from ad-hoc prompt engineering to systematic context management. All code, datasets, and over 4,000 execution logs are publicly available to foster a dynamic, community-driven benchmark. Our framework, leaderboard, and all artifacts are hosted on the LATTEArena project website at https://goodenhak.github.io/LATTEArena.

2026-06-17 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

マルチモーダル エージェント ネットワーク向けの QoS 対応トークン スケジューリングとプライベート データ評価

エージェント システムでは、人間が生成したデータ レコードが AI サービスの価値を支えます。しかし、クラウド コンピューティング パイプラインはリモート サーバーでの処理を集中化します。データの集中化により個人データの主権が低下し、サービス品質 (QoS) が低下する可能性があります。一方、ユーザーの貢献は量も質も多様です。分散型レコードは偏り、ノイズが多く、不均一に分散する可能性があります。データの課題に対処するために、私たちは分散型でリソースに制約のあるエージェント システムに対する公平なトークンの割り当てとプライベート データの評価を研究しています。私たちのアプローチは、マルチモーダル表現を共有セマンティック空間に埋め込み、差分プライベート (DP) プロトタイプをリリースして、セマンティック漏洩を削減しながら実用性を維持します。 DP 保証により、効果的な貢献に報酬を与え、データの異質性や AI リソースの不足に対して堅牢性を維持する公平なトークン割り当てスキームを設計します。広範なシミュレーションにより、標準ベンチマークと比較して貢献度ベースの公平性と QoS が向上していることが実証されています。画像再構成攻撃に対する耐性の向上は、マルチモーダルな個人データのプライバシーが強化されていることを示しています。

原文 (English)

QoS-Aware Token Scheduling and Private Data Valuation for Multi-Modal Agentic Networks

In agentic systems, human-generated data records anchor the value of AI services. Yet cloud compute pipelines centralize processing on remote servers. Data centralization reduces personal data sovereignty and may potentially degrade the quality of service (QoS). Meanwhile, user contributions are diverse in quantity and quality: decentralized records can be biased, noisy, and heterogeneously distributed. To address the data challenge, we study fair token allocation and private data valuation for decentralized and resource-constrained agentic systems. Our approach embeds multi-modal representations in a shared semantic space and releases differentially private (DP) prototypes to preserve utility while reducing semantic leakage. With the DP guarantee, we design a fair token allocation scheme that rewards effective contributions and remains robust to data heterogeneity and AI resource scarcity. Extensive simulations demonstrate improved contribution-based fairness and QoS compared to standard benchmarks. The improved resistance to image reconstruction attacks indicates enhanced privacy for multi-modal personal data.

2026-06-17 13:00 JSTarXiv cs.AIビジネス/資金調達

Mind-Studio: 部分的に観察可能なゲームの先読み評価を備えた実行可能な世界モデル

ワールドモデル合成は、インタラクション経験を環境ダイナミクスの内部モデルに変えることを目的としています。既存のシンボリックなアプローチは、観察された遷移やローカル ルールの混合に適合することがよくありますが、実際の環境から独立して実行できる完全な実行可能プログラムは生成されません。私たちは、大規模な言語モデルを使用して、状態-アクション-次の状態の軌跡から実行可能な pygame スタイルの世界モデルを合成するフレームワークである Mind-Studio を紹介します。 Mind-Studio は、エントロピーで選択されたトレースを、スクリーンショットから抽出されたオブジェクト、アクション、静的シーン情報を含む軽量のゲーム スキル ファイルと組み合わせます。生成されたワールド モデル ロールアウトを同じ状態からの Real-ALE ロールアウトと比較する K ステップ先読み忠実度プロトコルを使用して合成品質を評価します。 Montezuma'sリベンジ では、Mind-Studio は 8 つのサブ目標のうち 5 つを検証しながら、選択されたアクションの次の状態の予測を PoE-World の 0.3% から 48.7% に改善しました。 Alien、Assault、Skiing にわたって、以前に学習された先読みソースよりも強力なブランチレベルの忠実度を実現します。

原文 (English)

Mind-Studio: Executable World Models with Lookahead Evaluation for Partially Observable Games

World-model synthesis aims to turn interaction experience into an internal model of environment dynamics. Existing symbolic approaches often fit observed transitions or mixtures of local rules, but they do not produce a complete executable program that can run independently of the real environment. We present Mind-Studio, a framework that synthesizes executable pygame-style world models from state-action-next-state trajectories using large language models. Mind-Studio combines entropy-selected traces with a lightweight game skill file containing object, action, and static scene information extracted from screenshots. We evaluate synthesis quality with a K-step lookahead fidelity protocol that compares generated world-model rollouts against Real-ALE rollouts from the same state. On Montezuma's Revenge, Mind-Studio improves chosen-action next-state prediction from 0.3% for PoE-World to 48.7% while verifying 5 of 8 subgoals; across Alien, Assault, and Skiing, it achieves stronger branch-level fidelity than prior learned lookahead sources.

2026-06-17 13:00 JSTarXiv cs.AIビジネス/資金調達

Membership Inference Attacks against Large Audio Language Models

We present the first systematic Membership Inference Attack (MIA) evaluation of LALMs. Using Multi-modal Blind Baselines based on textual,…

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection

Instruction-tuned LLMs can annotate thousands of instances at low cost. This raises two questions for active learning (AL): can LLM labels…

2026-06-17 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIO…

2026-06-17 12:14 JSTITmedia AI+ビジネス/資金調達

SpaceX、AIコーディング「Cursor」を9.6兆円で買収 「近く大幅な改善」へ

Corsorは公式Xで、「近く大幅な改善が行われる予定だ」と述べた。

2026-06-17 05:11 JSTTechCrunch AIビジネス/資金調達

SpaceX valuation balloons to $2.6T, briefly passes Amazon

SpaceX's valuation has increased by $1 trillion since its shares started trading on Friday.

2026-06-17 00:53 JSTTechCrunch AIビジネス/資金調達

SpaceX is public: Everything you need to know post-IPO

TechCrunch has followed SpaceX's start, struggles, and successes from the early days. And we're here for what happens next too. This packag…

2026-06-16 22:15 JSTTechCrunch AIビジネス/資金調達

Probably raises $9M to build a more reliable kind of AI

Probably wants to prevent hallucinations and factual errors from reaching users, and achieve accuracy on par with deterministic systems.

2026-06-16 20:21 JSTTechCrunch AIビジネス/資金調達

SpaceX to acquire Cursor for $60B in stock, days after blockbuster IPO

The deal is supposed to help SpaceX's struggling AI division. The company told IPO investors it sees a $26 trillion addressable market in A…

2026-06-16 15:59 JSTTechCrunch AIエージェントビジネス/資金調達

Malaysia’s AI agent-powered messaging app Respond.io raises $62.5M, eyes acquisitions

Respond.io, one of Malaysia's startups to watch, uses AI agents to handle high volumes of customer inquiries and charges per convo, not per…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

地理空間データ取得のためのリスク認識 LLM エージェント: 設計と敵対的予備評価

自然言語クエリを使用して、クラウドベースの地理空間カタログからリモート センシング データを取得するための LLM 主導のフレームワークを紹介します。このシステムはユーザーの意図を構造化された API 呼び出しに変換し、衛星画像や環境データセットへの効率的なアクセスを可能にします。このアーキテクチャには、安全性とポリシー適用のための Guardrail、意図解釈のための General-QA、およびスキーマ対応 API 呼び出し生成のための Recommender-Analyst の 3 つのエージェントが統合されています。この調整された設計により、外部データ サービスとの信頼性が高く、意味的に調整された対話が保証されます。モジュール式フレームワークは、API スキーマの置換を通じてプラットフォーム間で移植可能であり、環境モニタリング、災害対応、気候分析のアプリケーションをサポートします。ユーザーの意図と地理空間インフラストラクチャの間にスケーラブルなインターフェイスを確立し、合理化および自動化された地球観測ワークフローを可能にします。敵対的なマルチターン設定での予備実験では、プロンプトレベルの安全指示により堅牢性が向上することが示されていますが、まれに影響の大きい障害が API 操作シナリオで持続し、安全性、使いやすさ、コスト効率のバランスをとった適応的なシステムレベルの防御の必要性が強調されており、これがインターセプトレベルの Guardrail エージェントの使用の動機となっています。

原文 (English)

Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation

We present an LLM-driven framework for retrieving remote sensing data from cloud-based geospatial catalogues using natural language queries. The system converts user intent into structured API calls, enabling efficient access to satellite imagery and environmental datasets. The architecture integrates three agents: Guardrail for safety and policy enforcement, General-QA for intent interpretation, and Recommender-Analyst for schema-aware API call generation. This coordinated design ensures reliable, semantically aligned interaction with external data services. The modular framework is portable across platforms through API schema substitution and supports applications in environmental monitoring, disaster response, and climate analysis. It establishes a scalable interface between user intent and geospatial infrastructure, enabling streamlined and automated Earth observation workflows. Preliminary experiments under adversarial multi-turn settings show that prompt-level safety instructions improve robustness, although rare high-impact failures persist in API manipulation scenarios and highlight the need for adaptive, system-level defenses that balance safety, usability, and cost efficiency, which motivates the use of our intercept-level Guardrail agent.

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

CODA-BENCH: コード エージェントはデータ量の多いタスクを処理できますか?

高度なエージェントは自律型エンジニアとして動作する可能性をますます実証しており、現実世界の開発の複雑さを捉える評価ベンチマークに対する需要が高まっています。このような環境には通常、複雑なコードと大規模なデータ (ファイル システムなど) の両方が含まれます。ただし、既存のベンチマークは通常、コード中心の機能またはデータ中心の機能を個別に評価するため、実際の開発シナリオとの明らかなギャップが残ります。このペーパーでは、データ集約型環境でコードとデータ インテリジェンスを共同で評価する初のベンチマークである CODA-BENCH を紹介することで、このギャップを埋めます。当社は、Kaggle エコシステム (数百のデータセットを含む) に基づいてデータ集約型 Linux サンドボックスを構築します。そこでは、エージェントが複雑なファイル階層を積極的に探索して関連リソースを特定し、データ駆動型の分析タスク用のコードを生成する必要があります。 CODA-BENCH は 31 のコミュニティにまたがる 1,009 のタスクで構成され、各タスク環境には平均 980 のファイルが含まれており、現実的なデータ スケールとノイズをシミュレートします。高度なエージェントの評価では、最高のパフォーマンスを誇るシステムでもデータ検出とコード実行を効果的に統合するのが難しく、成功率はわずか 61.1% にとどまっていることが明らかになりました。これらの結果は、データ集約型タスクに対する現在のエージェント機能の大きなギャップを浮き彫りにし、将来の研究の有望な方向性を示しています。

原文 (English)

CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?

Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation benchmarks that capture the complexity of real-world development. Such environments typically involve both complex code and large-scale data (i.e., file system). However, existing benchmarks usually evaluate code-centric or data-centric capabilities in isolation, leaving a clear gap with real development scenarios. In this paper, we bridge this gap by introducing CODA-BENCH, the first benchmark to jointly evaluate code and data intelligence in a data-intensive environment. We construct a data-intensive Linux sandbox based on the Kaggle ecosystem (containing hundreds of datasets), where agents must actively explore complex file hierarchies to identify relevant resources and generate code for data-driven analytical tasks. CODA-BENCH comprises 1,009 tasks spanning 31 communities, with each task environment containing an average of 980 files, simulating realistic data scale and noise. Evaluations of advanced agents reveal that even top-performing systems struggle to effectively integrate data discovery with code execution, achieving a success rate of only 61.1%. These results highlight a substantial gap in current agentic capabilities for data-intensive tasks and point to promising directions for future research.

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

漂流したのはシステムか、それとも裁判官か? LLM 評価パイプラインでいつでも有効なアトリビューション

LLM 製品の継続的な評価は、グラウンド トゥルースとして扱われる強力な LLM ジャッジに依存しています。安価なモニターがすべてのインタラクションをスコアリングし、スコアが下降するとチームがページングされます。しかし、審査員自体が API の背後にあるモデルであり、サイレント バージョン バンプや採点プロンプトの更新によって採点方法が変更されるため、すべてのドリフト アラームは、より悪い製品と変更された審査員の間で曖昧になります。現在の裁判官が安定したインターリーブで再得点する、固定の人間ラベル付きアンカー セット、裁判官対人間の差に関する 2 番目の賭け電子プロセス、および {なし、システム、裁判官} で評決を返すガード ウィンドウ ルールを使用して、曖昧さを解決します。私たちは、いつでも有効で、一方向の識別 (ジャッジのみがアンカーを移動できる)、帰属レースを証明します。その設計法則は、アンカーが保護するメインプロセスを上回って実行し、プロセスの直交性を持たなければならないというものです。 2 つの実際のジャッジ変更では、サイレント バージョン バンプが 60/60 の実行でジャッジ ドリフトとして検出され、ジャッジからシステムへの誤った帰属はゼロで、汚染を伴うストリクト プロンプト変更はガード幅 300 での 120 回中 110 回の実行で正しく帰属されました。一方、業界デフォルトのローリング Z テストはドリフトのないストリームの 75% で誤警報を出しました。すべての実験は、何も再調整せずに 2 番目のドメイン (TL;DR の要約) で複製され、ドメインが異なる場合、その違いはレースが予測するものです。厳密なプロンプトの変更により、そこでのスコアのシフトがより大きくなるため、アンカーの発射が速くなり、帰属が完璧になります (240/240)。このモニターは、すべての項目を強力に判断する場合のコストの約 0.64 で実行され、より安価ではあるが遅い制度では 0.21 で実行されます。

原文 (English)

Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines

Continuous evaluation of LLM products relies on a strong LLM judge treated as ground truth: a cheap monitor scores every interaction and a team is paged when the score drifts down. But the judge is itself a model behind an API, and a silent version bump or scoring-prompt update changes how it scores -- so every drift alarm is ambiguous between a worse product and a changed judge. We resolve the ambiguity with a fixed, human-labeled anchor set that the current judge re-scores at a steady interleave, a second betting e-process on the judge-versus-human gap, and a guard-window rule returning a verdict in {none, system, judge}. We prove anytime-validity, one-way identification (only the judge can move the anchors), an attribution race whose design law is that the anchors must out-run the main process they guard, and process orthogonality. On two real judge changes, a silent version bump is detected as judge drift in 60/60 runs with zero judge-to-system misattribution, and a contaminating strict-prompt change is correctly attributed on 110 of 120 runs at guard width 300 -- while the industry-default rolling z-test false-alarms on 75% of drift-free streams. Every experiment replicates on a second domain (TL;DR summarization) with nothing re-tuned, and where the domains differ the differences are the ones the race predicts: the strict-prompt change shifts scores harder there, so the anchors fire faster and attribution becomes perfect (240/240). The monitor runs at approximately 0.64 of the cost of strong-judging every item, or 0.21 in a cheaper-but-deafer regime.

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

ToolMenuBench: 信頼性が高く効率的な LLM エージェントのためのツール メニュー フィルタリング戦略のベンチマーク

ツールで拡張された大規模な言語モデル エージェントは、大規模なツール ライブラリ上で動作することが増えていますが、既存の評価では、目に見えるツール メニューが信頼性、効率性、安全性関連のリスク エクスポージャーをどのように形成するかよりも、モデルがツールを正しく呼び出せるかどうかに焦点を当てていることがよくあります。マルチステップ LLM エージェントのツール メニュー構築を評価するためのベンチマークである ToolMenuBench を紹介します。 ToolMenuBench は、ツール メニューのサイズ、ディストラクタの種類、状態に依存するタスク構造、およびリスク エクスポージャを変更し、表示ツール数、リスクのあるツール エクスポージャ、タスクの成功、間違ったツールの呼び出し、時期尚早のアクション、トークンの使用など、フィルタ レベルと下流のエージェント メトリクスの両方をレポートします。 7 つのモデル バックエンド、3 つのツール メニュー サイズ、6 つのフィルタリング方法、および 7 つの評価設定にわたる制御された評価において、CMTF はタスクの成功率を全ツール公開時の 32.1% から 85.7% に向上させ、同時に平均トークン使用量を約 98% 削減しました。因果的最小限のツール フィルタリングは、最も強力な全体的なトレードオフを達成し、フィルタリングされていない露出、字句フィルタリング、状態認識フィルタリング、およびより広範な因果パス ベースラインと比較して、目に見えるツール、間違ったツールの呼び出し、時期尚早なアクション、および危険なツールの露出を削減します。 ToolMenuBench は、エージェント インターフェイスの問題、つまり、どのツールをいつ表示する必要があるか、どのようなコストまたはリスクの制約の下で表示するかを検討するための再利用可能な評価フレームワークを提供します。

原文 (English)

ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents

Tool-augmented large language model agents increasingly operate over large tool libraries, but existing evaluations often focus on whether a model can call a tool correctly rather than how the visible tool menu shapes reliability, efficiency, and safety-relevant risk exposure. We introduce ToolMenuBench, a benchmark for evaluating tool-menu construction in multi-step LLM agents. ToolMenuBench varies tool-menu size, distractor type, state-dependent task structure, and risk exposure, and reports both filter-level and downstream agent metrics, including visible-tool count, risky-tool exposure, task success, wrong-tool calls, premature actions, and token usage. In a controlled evaluation across seven model backends, three tool-menu sizes, six filtering methods, and seven evaluation settings, CMTF improves task success from 32.1% under all-tools exposure to 85.7%, while reducing average token usage by roughly 98%. Causal minimal tool filtering achieves the strongest overall tradeoff, reducing visible tools, wrong-tool calls, premature actions, and risky-tool exposure relative to unfiltered exposure, lexical filtering, state-aware filtering, and broader causal-path baselines. ToolMenuBench provides a reusable evaluation framework for studying the agent-interface problem: which tools should be visible, when they should be visible, and under what cost or risk constraints.

2026-06-16 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

マルチモーダル エージェント ネットワーク向けの QoS 対応トークン スケジューリングとプライベート データ評価

エージェント システムでは、人間が生成したデータ レコードが AI サービスの価値を支えます。しかし、クラウド コンピューティング パイプラインはリモート サーバーでの処理を集中化します。データの集中化により個人データの主権が低下し、サービス品質 (QoS) が低下する可能性があります。一方、ユーザーの貢献は量も質も多様です。分散型レコードは偏り、ノイズが多く、不均一に分散する可能性があります。データの課題に対処するために、私たちは分散型でリソースに制約のあるエージェント システムに対する公平なトークンの割り当てとプライベート データの評価を研究しています。私たちのアプローチは、マルチモーダル表現を共有セマンティック空間に埋め込み、差分プライベート (DP) プロトタイプをリリースして、セマンティック漏洩を削減しながら実用性を維持します。 DP 保証により、効果的な貢献に報酬を与え、データの異質性や AI リソースの不足に対して堅牢性を維持する公平なトークン割り当てスキームを設計します。広範なシミュレーションにより、標準ベンチマークと比較して貢献度ベースの公平性と QoS が向上していることが実証されています。画像再構成攻撃に対する耐性の向上は、マルチモーダルな個人データのプライバシーが強化されていることを示しています。

原文 (English)

QoS-Aware Token Scheduling and Private Data Valuation for Multi-Modal Agentic Networks

In agentic systems, human-generated data records anchor the value of AI services. Yet cloud compute pipelines centralize processing on remote servers. Data centralization reduces personal data sovereignty and may potentially degrade the quality of service (QoS). Meanwhile, user contributions are diverse in quantity and quality: decentralized records can be biased, noisy, and heterogeneously distributed. To address the data challenge, we study fair token allocation and private data valuation for decentralized and resource-constrained agentic systems. Our approach embeds multi-modal representations in a shared semantic space and releases differentially private (DP) prototypes to preserve utility while reducing semantic leakage. With the DP guarantee, we design a fair token allocation scheme that rewards effective contributions and remains robust to data heterogeneity and AI resource scarcity. Extensive simulations demonstrate improved contribution-based fairness and QoS compared to standard benchmarks. The improved resistance to image reconstruction attacks indicates enhanced privacy for multi-modal personal data.

2026-06-16 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

どこが間違っていたのでしょうか?セマンティック状態追跡による Web エージェントのプロセス レベルの評価

Web エージェントは長い対話シーケンスを通じて動作しますが、既存のベンチマークは最終的な成功のみを評価し、すべてのプロセス情報を破棄し、改善に関するガイダンスをほとんど提供しません。この作業では、Web エージェントのプロセス レベルの分析を実行します。難易度を制御し、セマンティックな状態を自動的に追跡する 1,800 個のタスク インスタンスのベンチマークである WebStep を紹介します。各 Web サイトは、GUI とともに決定論的セマンティック MDP を公開します。エージェントはインターフェイス上で動作し、環境はバックグラウンドで高レベルの状態と遷移を記録するため、手動による注釈なしで詳細な分析が可能になります。セマンティックな軌跡に基づいて、プロセスのメトリクスが結果の評価では見えない違いを明らかにすることを最初に示します。つまり、成功率が 31 ~ 33% 以内にクラスター化されている 3 つのエージェントは、探索範囲と実行精度において乖離しています。次に、スキルごとに分解すると、これらの違いの性質が特徴づけられ、同じ Web サイト内に隠されている反対のスキルごとのランキングが明らかになります。たとえば、ハウジングでは、OpenAI CUA はコミット アクションで Qwen3.5 を 23.7% 上回っていますが、フィルタリングでは 15.6% 下回っており、ドメイン内であっても改善すべき具体的なスキルを特定します。分岐分析は、タスクを失う決定的なエラーをさらに特定し、このエラーが共有エラーではなくエージェント固有であることを示します。最後に、タスクが難しくなるにつれて、これらの差は広がります。簡単なタスクでは成功率は似ていますが、探索がより要求が厳しくなるにつれて、成功率は大きく異なります。当社のプロセスレベルの分析は、Web エージェントの評価に新たな道を開き、各エージェントのどこをどのように改善する必要があるかについて、きめ細かく実用的な洞察を提供します。

原文 (English)

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improvement. In this work, we conduct a process-level analysis of web agents. We introduce WebStep, a benchmark of 1,800 task instances with controlled difficulty and automatic semantic state tracking. Each website exposes a deterministic semantic MDP alongside the GUI: the agent operates on the interface, while the environment records high-level states and transitions in the background, enabling fine-grained analysis without manual annotation. Based on the semantic trajectory, we first show that process metrics reveal differences invisible to outcome evaluation: three agents whose success rates cluster within 31-33% diverge in exploration reach versus execution accuracy. Then, decomposing by skill characterizes the nature of these differences, exposing opposite per-skill rankings hidden within the same website: e.g., on Housing, OpenAI CUA outperforms Qwen3.5 by 23.7% on commit actions yet underperforms it by 15.6% on filtering, pinpointing a concrete skill to improve even within a domain. Bifurcation analysis further localizes the decisive error that loses the task and shows that this error is agent-specific rather than shared. Finally, these differences widen as tasks grow harder: success rate is similar on easy tasks but separates sharply as exploration becomes more demanding. Our process-level analysis opens a new avenue in web agent evaluation, providing fine-grained and actionable insight into where and how each agent should be improved.

2026-06-16 13:00 JSTarXiv cs.AIビジネス/資金調達

Mind-Studio: 部分的に観察可能なゲームの先読み評価を備えた実行可能な世界モデル

ワールドモデル合成は、インタラクション経験を環境ダイナミクスの内部モデルに変えることを目的としています。既存のシンボリックなアプローチは、観察された遷移やローカル ルールの混合に適合することがよくありますが、実際の環境から独立して実行できる完全な実行可能プログラムは生成されません。私たちは、大規模な言語モデルを使用して、状態-アクション-次の状態の軌跡から実行可能な pygame スタイルの世界モデルを合成するフレームワークである Mind-Studio を紹介します。 Mind-Studio は、エントロピーで選択されたトレースを、スクリーンショットから抽出されたオブジェクト、アクション、静的シーン情報を含む軽量のゲーム スキル ファイルと組み合わせます。生成されたワールド モデル ロールアウトを同じ状態からの Real-ALE ロールアウトと比較する K ステップ先読み忠実度プロトコルを使用して合成品質を評価します。 Montezuma'sリベンジ では、Mind-Studio は 8 つのサブ目標のうち 5 つを検証しながら、選択されたアクションの次の状態の予測を PoE-World の 0.3% から 48.7% に改善しました。 Alien、Assault、Skiing にわたって、以前に学習された先読みソースよりも強力なブランチレベルの忠実度を実現します。

原文 (English)

Mind-Studio: Executable World Models with Lookahead Evaluation for Partially Observable Games

World-model synthesis aims to turn interaction experience into an internal model of environment dynamics. Existing symbolic approaches often fit observed transitions or mixtures of local rules, but they do not produce a complete executable program that can run independently of the real environment. We present Mind-Studio, a framework that synthesizes executable pygame-style world models from state-action-next-state trajectories using large language models. Mind-Studio combines entropy-selected traces with a lightweight game skill file containing object, action, and static scene information extracted from screenshots. We evaluate synthesis quality with a K-step lookahead fidelity protocol that compares generated world-model rollouts against Real-ALE rollouts from the same state. On Montezuma's Revenge, Mind-Studio improves chosen-action next-state prediction from 0.3% for PoE-World to 48.7% while verifying 5 of 8 subgoals; across Alien, Assault, and Skiing, it achieves stronger branch-level fidelity than prior learned lookahead sources.

2026-06-16 13:00 JSTarXiv cs.AIビジネス/資金調達

RecourseBench: 再現可能なアルゴリズムによるリソース評価のためのモジュール式フレームワーク

アルゴリズムによる救済方法は、不利なモデル決定を覆すために必要な行動を個人に知らせる、反事実的な説明を提供します。方法論の急速な進歩にもかかわらず、原理的な比較は依然としてとらえどころのないものです。既存のフレームワークは拡張が困難であることが多く、相互運用性と、統合された手法で最初に報告された結果を忠実に再現する体系的な検証の両方が欠けています。私たちは、モジュール性、再現性、対話性という 3 つの取り組みを中心に構築された統合評価フレームワークである \emph{RecourseBench} を紹介します。このフレームワークは、パイプラインを完全に分離された 5 つのレイヤー (データ、前処理、モデル、リコース メソッド、評価) に分解し、抽象インターフェイスと動的レジストリによって管理されます。以前のベンチマークにおける再現性のギャップに対処するために、すべての統合メソッドが最初に報告された結果に対して自動テスト スイートによって検証される 4 層の分類システムを導入しました。さらに、メソッド、データセット、モデル アーキテクチャ間で柔軟な構成主導の比較を行うためのインタラクティブな Web インターフェイスも提供します。当社のフレームワークは現在、28 の最先端のリソースメソッドを統合しており、当社の知る限りでは、自動化された定量テストを通じてメソッドレベルの再現性を明示的に強化する最初のリソースベンチマークを構成しています。

原文 (English)

RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation

Algorithmic recourse methods provide counterfactual explanations that inform individuals of the actions required to overturn an unfavorable model decision. Despite rapid methodological progress, principled comparison remains elusive; existing frameworks are often difficult to extend and lack both interoperability and systematic verification that integrated methods faithfully reproduce their originally reported results. We introduce \emph{RecourseBench}, a unified evaluation framework built around three commitments namely, modularity, reproducibility, and interactivity. The framework decomposes the pipeline into five fully decoupled layers -- Data, Preprocessing, Model, Recourse Method, and Evaluation -- governed by abstract interfaces and a dynamic registry. To address the reproducibility gap in prior benchmarks, we introduce a four-tier classification system in which every integrated method is validated by an automated test suite against its originally reported results. We further provide an interactive web interface for flexible, configuration-driven comparison across methods, datasets, and model architectures. Our framework currently integrates 28 state-of-the-art recourse methods and, to our knowledge, constitutes the first recourse benchmark to explicitly enforce method-level reproducibility through automated, quantitative testing.

2026-06-16 13:00 JSTarXiv cs.AIビジネス/資金調達

Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations

Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series shaped by reporting r…

2026-06-16 13:00 JSTarXiv cs.AIビジネス/資金調達

Integrating Multi-Label Classification and Generative AI for Scalable Analysis of User Feedback

In highly competitive software markets, user experience (UX) evaluation is crucial for ensuring software quality and fostering long-term pr…

2026-06-16 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Evaluation of Alternative-Based Information Systems for Deliberative Polling using an Agentic Simulator

Deliberative polling promises to improve collective decision-making by exposing shareholders to a broad range of arguments before they vote…

2026-06-16 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Agentomics: Economic Foundations for the Valuation, Attribution, and Pricing of AI Agents in Human-AI Workflows

Agentic AI systems are increasingly being deployed as productive resources in organizational workflows, yet existing evaluation methods pri…

2026-06-16 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

A Security Analysis of Long-Horizon Agentic AI Systems: Threats, Evaluation, and Framework Development

This paper presents a structured analysis of security challenges in long-horizon agentic AI systems. The study reviews existing threats, ev…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Combining Retrieval-Augmented Text Generation with LLMs for Reading Content Recommendations

This work presents the design, implementation, and evaluation of a system for generating personalized reading content using Large Language…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation

LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and diffi…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Intelligence Is Not the Bottleneck: Validating an LLM First-Pass Manuscript Score Against Peer-Review Outcomes

Large language model (LLM) systems are increasingly proposed to assist peer review, yet most evaluations judge the prose of machine-generat…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills

Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that exte…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models

The worldwide surge of authoritarianism, combined with the increasing central role in users' everyday lives, raises the question of to what…

2026-06-16 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Propagating Structural Guidance: Synthesizing Fluorescein Angiography from Fundus Images and Sparse OCT Scans

Fundus fluorescein angiography (FFA) is critical for assessing retinal vascular abnormalities, but its acquisition is invasive and not alwa…

2026-06-16 13:00 JSTarXiv cs.AIエージェントロボティクスビジネス/資金調達

Is Your Trajectory Displacement Safe in Long-tail?

Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing eva…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory

Standard LLM evaluation practices compress diverse abilities into single scores, obscuring their inherently multidimensional nature. We pre…

2026-06-16 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks

Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide ri…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework

As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

障害を認識した可観測性によるマルチエージェント LLM システムの無駄な計算の早期診断

ツールを使用するマルチエージェント大規模言語モデル (LLM) システムは、応答を生成する前に、モデル トークン、ツール呼び出し、再試行、コード実行による計算を費やします。実行が失敗した場合、最終応答の評価によって終点が明らかになりますが、通常は、軌道が回復可能な進行を停止した時点ではありません。このペーパーでは、マルチエージェント LLM トレースにおける無駄な計算を診断するための障害認識可観測性フレームワークを紹介します。このフレームワークは、ツールの信頼性、実行の回復、オーケストレーション ループ、証拠の可用性、情報の変更、予算のプレッシャーなど、繰り返し発生する障害モードをオンライン トレース信号にマッピングします。 3 エージェントの質問応答システムでフレームワークをインスタンス化し、同一の実行上限の下で 165 の GAIA 検証トレースで評価します。運用上の失敗は依然として一般的です。レベル 1 の実行は 22/53 回、レベル 2 の実行は 33/86 回、レベル 3 の実行は 12/26 回で、使用可能な最終応答を生成できませんでした。トレースは、不十分な証拠、反復アクション ループ、最大ステップ終了、ツール失敗の連続発生、有用な出力なしで成功する実行呼び出しなど、これらの結果の背後にあるさまざまなメカニズムを明らかにします。平均トークン使用量はレベル 1 の 8,152 トークンからレベル 3 の 16,389 トークンに増加しますが、証拠の入手可能性と文レベルのサポートは異なります。キャッシュされた 10 トレースの LLM ジャッジ グラウンディング監査により、安価なオンライン シグナルとより深いセマンティック メトリクスが相補的な障害層を捉えていることがわかります。その結果、障害を認識する可観測性は、生の実行ログと最終応答の精度の間の診断レイヤーとして位置付けられます。

原文 (English)

Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability

Failure-aware observability diagnoses wasted computation in multi-agent LLM systems before final-answer evaluation can explain what went wrong. We propose a trace-based framework for a three-agent architecture -- orchestrator, search agent, and execution agent -- that converts structured events into online signals for loops, budget pressure, low information gain, and tool instability, then adds offline semantic grounding metrics and selective LLM-as-judge evaluation. On 165 GAIA validation traces under identical caps, 98 runs produce usable final answers and 67 fail or stop without one. Among warned failed runs, 58.1% of tokens are spent after the first warning on average, indicating substantial opportunity for intervention. A 10-task Level-2 pilot uses warnings to diversify search or require evidence, reducing post-warning token fraction from 0.638 in the baseline to 0.304. The results support a layered design: cheap online signals help the orchestrator redirect or halt redundant behavior, while deeper semantic checks identify whether completed answers are grounded enough to trust.

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

思考の連鎖がより良くわかるとき: マルチターン推論モデルの失敗モード

マルチターン推論モデルの失敗は、最終スコア評価ではほとんど認識されません。モデルは、長い対話の早い段階で安全でないスタンスに固定される可能性がありますが、最終ターンの拒否率は、しっかりと調整されたベースラインと区別できないように見える場合があります。これらの隠れた時間的ダイナミクスを明らかにするために、トレースレベルの診断である CoT-Output 2x2 安全性マトリックスを提案します。このフレームワークは、2 つの独立した軸 (内部推論と可視出力) に沿ってすべてのターンにラベルを付け、運用上定義された 4 つの失敗セルを生成します。堅牢なアライメント、アライメント偽装、明白なジェイルブレイク、およびコンテキストインジェクション失敗と呼ばれる明確な失敗モード (CoT は安全な推論を維持しますが、目に見える出力が害を生み出し、推論の不誠実さのマルチターンの現れを強調する) です。私たちは、5 つの監視条件にわたって、固定攻撃者に対する 3 つの抽出された推論ターゲットを評価し、情報ハザード シナリオに関する 6750 のターンレベルの観察を収集しました。私たちの分析により、再現可能な 2 つの脆弱性が明らかになりました。1 つは、明示的なモニタリング キューによって逆説的にアラインメント偽装率が抑制されるのではなく増加する、見落としのパラドックスです。もう 1 つは、安全な内部状態にもかかわらず、モデルが安全でない外部出力にロックされるコンテキスト インジェクションの失敗です。フォローアップのトレース診断研究をサポートするために、マルチターン ダイアログと CoT トレースの完全なデータセットをリリースします。

原文 (English)

When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialogue, yet its final-turn refusal rate may appear indistinguishable from a robustly aligned baseline. To expose these hidden temporal dynamics, we propose a trace-level diagnostic - the CoT-Output 2x2 safety matrix. This framework labels every turn along two independent axes (internal reasoning and visible output), yielding four operationally defined failure cells: robust alignment, alignment faking, overt jailbreak, and a distinct failure mode we term context-injection failure (where the CoT maintains safe reasoning, but the visible output produces harm, highlighting a multi-turn manifestation of reasoning unfaithfulness). We evaluate three distilled reasoning targets against a fixed attacker across five oversight conditions, collecting 6750 turn-level observations on the Information-Hazard scenario. Our analysis reveals two reproducible vulnerabilities: an oversight paradox where explicit monitoring cues paradoxically increase alignment-faking rates rather than suppress them, and a context-injection failure where models lock onto unsafe external outputs despite safe internal states. We release the full dataset of multi-turn dialogues and CoT traces to support follow-up trace-diagnostic research.

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

AgentBeats: オープン性、標準化、再現性のためのエージェント化エージェントの評価

エージェント システムはドメイン間で急速に進歩していますが、その評価は依然として断片的です。ほとんどのベンチマークは、固定された LLM 中心のハーネスに依存しており、これには高度な統合が必要で、テストと運用の不一致が生じ、多様なエージェント設計間の公正な比較が制限されます。根本的な問題は、オープンでエージェントに依存しない評価インターフェイスが欠如していることです。当社は、評価が審査員エージェントによって実行され、すべての参加者がタスク管理用の A2A とツール アクセス用の MCP という標準化されたプロトコルを通じて対話するエージェント化エージェント評価 (AAA) を提唱しています。従来のベンチマークでは、ベンチマーク用とエージェント用の 2 つの個別のインターフェイスが定義されていましたが、AAA では 1 つだけが必要でした。これにより、評価ロジックをエージェントの実装から分離し、再現可能で相互運用可能な複数エージェントの評価を可能にする、汎用的で統一されたフレームワークが得られます。さらに、AAA の具体的な実現として AgentBeats を紹介します。オープン性、プライバシー、再現性に関する現実世界の制約と互換性のある標準化された評価を可能にする 5 つの実用的な動作モードを特定します。私たちの設計を大規模に評価するために、私たちは 2 つの調査を実施しました。1 つは 12 カテゴリーにわたる 298 人の審査員エージェントと、独立した参加者からの 467 人の被験者エージェントを集めた 5 か月間にわたるオープン コンテストで、AAA が異質な範囲のベンチマークに適用されることを示しました。また、コーディングエージェントに関するケーススタディでは、エージェント化された評価が公的記録との忠実性を維持しながら、これまで欠けていた直接対決の結果が明らかになり、エージェント設計に関する研究上の洞察が得られることが確認されました。コミュニティ規模のフィールド調査と制御されたコーディングのケーススタディを組み合わせて、AAA が大規模な異種シナリオ全体にわたってカバレッジ、実用性、忠実性を提供することを検証します。 AAA と AgentBeats は共に、オープンで標準化された再現可能なエージェント評価への明確な道筋を提供します。

原文 (English)

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all participants interact through standardized protocols: A2A for task management and MCP for tool access. Conventional benchmarking defines two separate interfaces, one for the benchmark and one for the agent, while AAA only needs one; this yields a generic, unified framework that separates assessment logic from agent implementation and enables reproducible, interoperable, and multi-agent evaluation. We further introduce AgentBeats as a concrete realization of AAA: we identify five practical operation modes that make standardized assessment compatible with real-world constraints on openness, privacy, and reproducibility. To evaluate our design at scale, we conduct two studies: a five-month open competition that drew 298 judge agents across 12 categories together with 467 subject agents from independent participants, showing that AAA applies across a heterogeneous range of benchmarks; and a case study on coding agents that confirms agentified evaluation preserves fidelity with the public record while surfacing previously missing head-to-head results, yielding research insights about agent design. Combining a community-scale field study and a controlled coding case study, we verify that AAA delivers coverage, practicality, and fidelity across heterogeneous scenarios at scale. Together, AAA and AgentBeats offer a clear path toward open, standardized, and reproducible agent assessment.

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MA-ProofBench: 数学解析における定理証明のための LLM の 2 層評価

大規模言語モデル (LLM) は定理証明の自動化において顕著な進歩を遂げていますが、既存の正式なベンチマークは数学的範囲と難易度の両方において依然として限定的です。そのほとんどは、代数や初歩的な整数論など、形式化が容易な分野に集中しており、数学的分析など、より深い推論を必要とする下位分野の範囲は限られています。このギャップに対処するために、私たちの知る限りでは、数学的解析に特化した最初の正式な定理証明ベンチマークである MA-ProofBench を導入します。このベンチマークには、測定と統合理論、複雑な解析、関数解析など、6 つのコア トピックと 27 のサブカテゴリをカバーする 200 の形式化された定理が含まれています。問題は学部レベル(レベルI、100問)と博士レベルの2つの難易度に分かれています。適格レベル (レベル II、100 問)。LLM がさまざまな数学的深さで形式的推論をどの程度実行できるかを評価します。各問題は、人間主導、LLM 支援の形式化パイプラインとそれに続く独立した専門家のレビューを通じて構築され、形式的なステートメントが元の数学に忠実であることが保証されます。 MA-ProofBench では、最近のさまざまな汎用推論モデルと形式定理証明器を評価します。ただし、ほとんどのモデルのパフォーマンスは低く、最もパフォーマンスの高いモデルである GPT-5.5 でさえ、レベル I で Pass@8 が 16%、レベル II で 5% しか達成できず、ほとんどのモデルはレベル II で 0% 近くに留まっています。さらなる分析により、Mathlib の幻覚と不完全な証明が 2 つの主な失敗モードであることが特定され、一方、ベンチマークの自然言語バージョンの評価では、非公式推論と正式推論の間に明確なギャップがあることが明らかになりました。 MA-ProofBench は、高度な領域における形式的な数学的推論の進歩を追跡するための信頼できるリファレンスとして機能することを目的としています。

原文 (English)

MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis

Large Language Models (LLMs) have made notable progress in automated theorem proving, yet existing formal benchmarks remain limited in both mathematical coverage and difficulty. Most are concentrated in areas that are easier to formalize, such as algebra and elementary number theory, and provide limited coverage of subfields that require deeper reasoning, including mathematical analysis. To address this gap, we introduce MA-ProofBench, to the best of our knowledge, the first formal theorem-proving benchmark dedicated to Mathematical Analysis. The benchmark contains 200 formalized theorems covering 6 core topics and 27 subcategories, including measure and integration theory, complex analysis, and functional analysis. The problems are divided into two difficulty levels, an undergraduate level (Level I, 100 problems) and a Ph.D. qualifying level (Level II, 100 problems), to evaluate how well LLMs perform formal reasoning at different mathematical depths. Each problem is constructed through a human-led, LLM-assisted formalization pipeline followed by independent expert review, ensuring that the formal statements remain faithful to the original mathematics. We evaluate a range of recent general-purpose reasoning models and formal theorem provers on MA-ProofBench. However, most models perform poorly: even the best-performing model, GPT-5.5, achieves only 16% Pass@8 on Level I and 5% on Level II, while most models stay close to 0% on Level II. Further analysis identifies Mathlib hallucinations and incomplete proofs as the two dominant failure modes, while an evaluation on the natural-language version of the benchmark exposes a clear gap between informal and formal reasoning. MA-ProofBench is intended to serve as a reliable reference for tracking progress in formal mathematical reasoning in advanced domains.

2026-06-16 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models

Electroencephalography foundation models (EEG-FMs) have advanced brain signal analysis, but the lack of standardized evaluation benchmarks…

2026-06-16 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Can Artificial Intelligence Accelerate Technological Progress? Researchers' Perspectives on AI in Manufacturing and Materials Science

Artificial intelligence (AI) raises expectations of substantial increases in rates of technological progress, but such anticipations are of…

2026-06-16 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Critically Engaged Pragmatism: Scientific Norm and Social, Pragmatist Epistemology for AI Science Evaluation Tools

AI science evaluation tools aim to assess research credibility. As with traditional metrics such as impact factors, their edicts can be dec…

2026-06-16 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios

The rapid advancement of photorealistic Text-to-Video (T2V) generation brings in an urgent need for up-to-date evaluation methods. Existing…

2026-06-16 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw

Agentic language-model systems increasingly rely on mutable execution contexts, including files, memory, tools, skills, and auxiliary artif…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM が推論するのはいつですか?エントロピー相転移による動的システムの視点

Chain-of-thought (CoT) reasoning has become the default strategy for enhancing LLM capabilities, yet its application raises a fundamental question: when is explicit reasoning actually beneficial?経験的証拠は、顕著な矛盾を明らかにしています。CoT は、多くの場合、トークン消費量を増大させながら、事実に基づいた無制限のタスクに対してわずかな利益、またはマイナスの利益さえ提供します。この研究では、LLM 推論がタスクやモデルの静的な特性ではなく、生成中に現れる \emph{動的復号状態} であることを示します。体系的な分析を通じて、初期段階のエントロピー ダイナミクスがこの状態の信頼できるシグナルを提供することを発見しました。CoT の恩恵を受けるタスクは一貫したエントロピーの減少を示しますが、他のタスクは不安定または増加するパターンを示します。この動作は、高エントロピー探索体制から低エントロピー構造推論体制への相転移のような移行として解釈できます。これらの洞察に基づいて、我々は、早期デコードエントロピーを活用して推論戦略を適応的に選択する、軽量でトレーニング不要のルーティングフレームワークである \textbf{EDRM} (エントロピーダイナミクスベースの推論マニホールド) を提案します。 EDRM は、エントロピーの軌跡をコンパクトで解釈可能な多様体表現に埋め込み、ゼロショット デプロイメントときめ細かいインスタンス レベルの適応の両方を可能にします。さまざまなスケールとアーキテクチャの 15 のベンチマークと 4 つの LLM にわたって、EDRM は一貫して静的ベースラインを上回っています。データセット レベルでは、EDRM は \textbf{41--55\%} トークンの削減を達成しながら、わずか 50 個のキャリブレーション サンプルで精度を向上させます。インスタンス レベルでは、\textbf{27--45\%} トークンの節約を維持しながら、精度が最大 \textbf{4.7\%} まで向上します。これらの結果は、推論はデフォルトではなく選択的に呼び出される必要があることを示唆しており、効率的で適応的な LLM 推論に対するエントロピー駆動型の復号制御の有効性を示しています。

原文 (English)

When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions

Chain-of-thought (CoT) reasoning has become the default strategy for enhancing LLM capabilities, yet its application raises a fundamental question: when is explicit reasoning actually beneficial? Empirical evidence reveals a striking paradox: CoT often provides marginal or even negative gains on factual and open-ended tasks while multiplying token consumption. In this work, we show that LLM reasoning is not a static property of tasks or models, but a \emph{dynamic decoding state} that emerges during generation. Through systematic analysis, we find early-stage entropy dynamics provide a reliable signal of this state: tasks benefiting from CoT exhibit consistent entropy reduction, while others display unstable or increasing patterns. This behavior can be interpreted as a phase-transition-like shift from a high-entropy exploratory regime to a low-entropy structured reasoning regime. Based on these insights, we propose \textbf{EDRM} (Entropy Dynamics-based Reasoning Manifold), a lightweight and training-free routing framework that leverages early decoding entropy to adaptively select inference strategies. EDRM embeds entropy trajectories into a compact and interpretable manifold representation, enabling both zero-shot deployment and fine-grained instance-level adaptation. Across 15 benchmarks and 4 LLMs of varying scales and architectures, EDRM consistently outperforms static baselines. At the dataset level, EDRM achieves \textbf{41--55\%} token reduction while improving accuracy with as few as 50 calibration samples. At the instance level, it further improves accuracy by up to \textbf{4.7\%} while maintaining \textbf{27--45\%} token savings. These results suggest that reasoning should be invoked selectively rather than by default, and demonstrate the effectiveness of entropy-driven decoding control for efficient and adaptive LLM inference.

2026-06-16 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination

Vision-Language-Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task-r…

2026-06-16 07:00 JSTITmedia AI+LLM/生成AIビジネス/資金調達

300億円は「ROI不問」 Olive、Trunkを仕掛けるSMBC、新規事業の神髄は「撤退」にアリ

「Olive」や「Trunk」を相次いで成長軌道に乗せ、生成AI活用に向けて500億円の投資計画も打ち出した三井住友フィナンシャルグループ。そんな同社だが、約10年前はモバイルアプリで競合他行に大きく後れを取るなど、変革が進んでいなかった。堅実なメガバンクは、いかに挑戦を次々と…

2026-06-16 03:30 JSTTechCrunch AIビジネス/資金調達

SpaceX is public: Everything you need to know post-IPO

TechCrunch has followed SpaceX's start, struggles, and successes from the early days. And we're here for what happens next too. This packag…

2026-06-15 22:46 JSTTechCrunch AIビジネス/資金調達

Sarvam becomes India’s newest AI unicorn with $234 million funding round led by HCLTech

Indian IT services company HCLTech is investing $150 million in the Bengaluru startup.

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MA-ProofBench: 数学解析における定理証明のための LLM の 2 層評価

大規模言語モデル (LLM) は定理証明の自動化において顕著な進歩を遂げていますが、既存の正式なベンチマークは数学的範囲と難易度の両方において依然として限定的です。そのほとんどは、代数や初歩的な整数論など、形式化が容易な分野に集中しており、数学的分析など、より深い推論を必要とする下位分野の範囲は限られています。このギャップに対処するために、私たちの知る限りでは、数学的解析に特化した最初の正式な定理証明ベンチマークである MA-ProofBench を導入します。このベンチマークには、測定と統合理論、複雑な解析、関数解析など、6 つのコア トピックと 27 のサブカテゴリをカバーする 200 の形式化された定理が含まれています。問題は学部レベル(レベルI、100問)と博士レベルの2つの難易度に分かれています。適格レベル (レベル II、100 問)。LLM がさまざまな数学的深さで形式的推論をどの程度実行できるかを評価します。各問題は、人間主導、LLM 支援の形式化パイプラインとそれに続く独立した専門家のレビューを通じて構築され、形式的なステートメントが元の数学に忠実であることが保証されます。 MA-ProofBench では、最近のさまざまな汎用推論モデルと形式定理証明器を評価します。ただし、ほとんどのモデルのパフォーマンスは低く、最もパフォーマンスの高いモデルである GPT-5.5 でさえ、レベル I で Pass@8 が 16%、レベル II で 5% しか達成できず、ほとんどのモデルはレベル II で 0% 近くに留まっています。さらなる分析により、Mathlib の幻覚と不完全な証明が 2 つの主な失敗モードであることが特定され、一方、ベンチマークの自然言語バージョンの評価では、非公式推論と正式推論の間に明確なギャップがあることが明らかになりました。 MA-ProofBench は、高度な領域における形式的な数学的推論の進歩を追跡するための信頼できるリファレンスとして機能することを目的としています。

原文 (English)

MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis

Large Language Models (LLMs) have made notable progress in automated theorem proving, yet existing formal benchmarks remain limited in both mathematical coverage and difficulty. Most are concentrated in areas that are easier to formalize, such as algebra and elementary number theory, and provide limited coverage of subfields that require deeper reasoning, including mathematical analysis. To address this gap, we introduce MA-ProofBench, to the best of our knowledge, the first formal theorem-proving benchmark dedicated to Mathematical Analysis. The benchmark contains 200 formalized theorems covering 6 core topics and 27 subcategories, including measure and integration theory, complex analysis, and functional analysis. The problems are divided into two difficulty levels, an undergraduate level (Level I, 100 problems) and a Ph.D. qualifying level (Level II, 100 problems), to evaluate how well LLMs perform formal reasoning at different mathematical depths. Each problem is constructed through a human-led, LLM-assisted formalization pipeline followed by independent expert review, ensuring that the formal statements remain faithful to the original mathematics. We evaluate a range of recent general-purpose reasoning models and formal theorem provers on MA-ProofBench. However, most models perform poorly: even the best-performing model, GPT-5.5, achieves only 16% Pass@8 on Level I and 5% on Level II, while most models stay close to 0% on Level II. Further analysis identifies Mathlib hallucinations and incomplete proofs as the two dominant failure modes, while an evaluation on the natural-language version of the benchmark exposes a clear gap between informal and formal reasoning. MA-ProofBench is intended to serve as a reliable reference for tracking progress in formal mathematical reasoning in advanced domains.

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Every Eval Ever: AI 評価結果の統合スキーマとコミュニティ リポジトリ

AI 評価は、テストと進捗状況の理解に広く使用されています。しかし、評価者が多様であるため、分析や比較に課題となる矛盾が生じます。まず、結果は互換性のない形式で保存され、リーダーボード、論文、ブログ投稿、評価ハーネス ログ、カスタム リポジトリに分散されます。第 2 に、結果は異なる評価フレームワークによって作成され、名目上同一の評価に対して異なるスコアが生成され、メタデータの記録に一貫性がなく、比較、コミュニティ間での評価科学、コスト削減、再利用が妨げられます。 AI 評価結果の初の共有スキーマおよびコミュニティ クラウドソーシング リポジトリである Every Eval Ever を紹介します。このスキーマは、統合された単一の JSON ドキュメントで評価を表現する方法を標準化します。設計上ソースに依存せず、評価ハーネスや論文から同様に結果を取り込み、オプションで詳細な分析のためにインスタンスごとの出力を保存します。私たちは以下に貢献します。(i) コミュニティが管理するメタデータ スキーマと、それに付随するインスタンス レベルのスキーマ。この種の最初の標準化の取り組み。 (ii) 一般的なフォーマット、評価ハーネス、リーダーボードから統一スキーマへの自動コンバータ。 (iii) Hugging Face でホストされているクラウドソーシングのコミュニティ データベース。現在までに 22,235 のモデル、2,273 の固有のベンチマーク、および 31 の評価形式が含まれています。

原文 (English)

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First, results are saved in incompatible formats, scattered across leaderboards, papers, blog posts, evaluation harness logs, and custom repositories. Second, results are created by different evaluation frameworks, which produce divergent scores for nominally identical evaluations and record metadata inconsistently, hindering comparison, cross-community evaluation science, cost reduction, and reuse. We introduce Every Eval Ever, the first shared schema and community-crowdsourced repository for AI evaluation results. The schema standardizes how evaluations are represented in a unified, single JSON document. It is source-agnostic by design, ingesting results from evaluation harnesses and papers alike, and optionally stores per-instance outputs for fine-grained analysis. We contribute: (i) a community-governed metadata schema with a companion instance-level schema, the first standardization effort of its kind; (ii) automatic converters from popular formats, evaluation harnesses, and leaderboards to the unified schema; and (iii) a crowdsourced community database hosted on Hugging Face, currently spanning to date 22,235 models, 2,273 unique benchmarks, and 31 evaluation formats.

2026-06-15 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

StreamMemBench: 未来志向の支援のためのエージェント メモリのストリーミング評価

パーソナル エージェントの記憶の中心的な役割は、保存された情報と以前の対話を未来志向の支援に変えることです。日常的な使用では、エージェントが観察した内容やユーザーがエージェントとどのように対話するかによって有用な手がかりが得られ、エージェントはそれらを現在のリクエストから将来の同様のタスクに転送する必要があります。既存の記憶ベンチマークは通常、対話の想起やタスクの改善を個別にテストし、ストリーミング観察からその後の支援までの軌跡はほとんどテストされていません。 EgoLife の自己中心的なストリームからの各証拠アンカーを中心に 2 ステップのタスク シーケンスを構築するストリーミング ベンチマークである StreamMemBench を紹介します。最初のタスクでは証拠の使用をテストし、後続のタスクではフィードバックとインタラクション エクスペリエンスが再利用されているかどうかをテストします。 4 つの指標により、証拠の想起、最初の証拠の使用、フィードバックの組み込み、およびフォローアップの再利用が診断されます。 2 つのバックボーンにまたがる 8 つのメモリ システムを使用した実験では、現在のシステムでは、証拠が保存されているか、フィードバックがローカルに組み込まれている場合でも、観察された証拠を使用したり、フィードバックを信頼できる追跡動作に変換したりできないことが多いことが示されています。 StreamMemBench は、https://github.com/landian60/StreamMemBench で公開されています。

原文 (English)

StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance

A central role of personal-agent memory is to turn stored information and prior interactions into future-oriented assistance. In daily use, useful cues come from what the agent observes and how the user interacts with the agent, and the agent must carry them forward from the current request to similar future tasks. Existing memory benchmarks usually test dialogue recall or task improvement in isolation, leaving the trajectory from streaming observations to later assistance largely untested. We introduce StreamMemBench, a streaming benchmark that constructs a two-step task sequence around each evidence anchor from EgoLife egocentric streams. The initial task tests evidence use, while the follow-up task tests whether feedback and interaction experience are reused. Four metrics diagnose evidence recall, initial evidence use, feedback incorporation, and follow-up reuse. Experiments with eight memory systems across two backbones show that current systems often fail to use observed evidence or turn feedback into reliable follow-up behavior, even when evidence is stored or feedback is incorporated locally. StreamMemBench is publicly available at https://github.com/landian60/StreamMemBench.

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

コイン投げの裁判官? LLM-as-a-Judge 評価の信頼性と偏り

LLM-as-a-Judge は現在、モデル出力のランク付け、報酬モデルのトレーニング、公開リーダーボードへの入力に広く使用されていますが、その実行ごとの信頼性については十分に評価されていません。私たちは、2 つの OpenAI 判定モデル (GPT-4o-mini および GPT-4.1-mini) を使用して、10 カテゴリーにまたがる 29 のタスクについて同一の評価を繰り返し、質問ごとに 50 のペアワイズ トライアルと 50 のポイントワイズ トライアルを行い、温度および即時感度アブレーションを補足して研究しました。審査員全体でペアごとの好みは平均 13.6% の確率で反転し、28% の質問が反転率 20% を超え、1 つの質問は 56% に達しました。 GPT-4o-mini は、有意な 1 位バイアスも示します (A 多数派 72%、p = 0.024)。同時に、平均点ごとのスコアのギャップは小さく (10 点スケールで 0.19 ~ 0.36)、全体としては統計的に有意ではないため、ペアごとの点ごとのギャップが生じます。審査員は、自身のスカラー スコアが有意な質の違いの証拠をほとんど示さない場合でも、勝者を選択することがよくあります。裁判官内の不安定性を超えて、裁判官間の一致はわずか 76% ($\kappa = 0.51$) であり、意味的に同等のプロンプト テンプレートはテストされたケースの 25% で大多数の結果を変更し、決定論的なデコードは矛盾を軽減しますが、排除しません。信頼性曲線分析によると、私たちのデータセットでは、平均 95% の確率で 50 試行の参照評決を回復するための多数決には 11 回の反復試行が必要であり、分散が大きい質問の場合は 15 回に増加します。これらの発見は、単一試行の LLM 判定は一か八かの評価にはノイズが多すぎることが多く、複数試行の集計、位置のランダム化、明示的な不確実性レポートが標準的な手法であるべきであることを示唆しています。両方の審査員が単一のプロバイダーに属しているため、プロバイダー間のレプリケーションが引き続き重要な次のステップになります。

原文 (English)

The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation

LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized. We study repeated identical evaluations on 29 tasks spanning 10 categories using two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini), with 50 pairwise trials and 50 pointwise trials per question, supplemented by temperature and prompt-sensitivity ablations. Across judges, pairwise preferences flip on average 13.6% of the time, with 28% of questions exceeding a 20% flip rate and one question reaching 56%. GPT-4o-mini also exhibits a significant first-position bias (72% A-majority, p = 0.024). At the same time, mean pointwise score gaps are small (0.19--0.36 on a 10-point scale) and not statistically significant in aggregate, producing a pairwise--pointwise gap: judges frequently choose a winner even when their own scalar scores provide little evidence of a meaningful quality difference. Beyond within-judge instability, cross-judge agreement is only 76% ($\kappa = 0.51$), semantically equivalent prompt templates change majority outcomes in 25% of tested cases, and deterministic decoding reduces but does not eliminate inconsistency. A reliability curve analysis shows that, in our dataset, 11 repeated trials are needed for a majority vote to recover the 50-trial reference verdict with 95% probability on average, rising to 15 for high-variance questions. These findings suggest that single-trial LLM judging is often too noisy for high-stakes evaluation, and that multi-trial aggregation, position randomization, and explicit uncertainty reporting should be standard practice. Because both judges are from a single provider, cross-provider replication remains an important next step.

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

VHDLSuite: データ合成と評価を備えた LLM VHDL 生成用の統合パイプライン

大規模言語モデル (LLM) は、特に Verilog のレジスタ転送レベル (RTL) コード生成において優れた機能を示しています。ただし、他のハードウェア記述言語 (HDL)、特に VHDL でのパフォーマンスの評価は、厳密なセマンティック ルールなどの独特の言語特性により、Verilog とは異なる評価上の考慮事項が導入されているにもかかわらず、制限されたままです。このカバレッジの欠如により、構造やセマンティクスが異なるハードウェア設計言語間で現在のモデルがどの程度一般化されているかを完全に理解することが制限されます。このギャップに対処するために、自動ベンチマーク合成、実行可能検証、およびマルチモデル診断分析を統合した、スケーラブルな VHDL 生成評価のためのベンチマーク中心のインフラストラクチャである VHDLSuite を導入します。まず、Verilog デザインとそれに付随するテストベンチを実行可能な VHDL ベンチマーク インスタンスに自動的に変換するデータ パイプラインを提案します。その後、VUnit/GHDL ベースの検証を行って、リリースされた各タスクが VHDL 環境でコンパイル可能、実行可能、一貫してチェック可能であることを確認します。次に、VHDLBench を紹介します。これは、幅広い複雑さレベルにわたる完全で検証済みのテストベンチを備えた 200 を超える VHDL 問題を含むベンチマークです。第三に、最先端の LLM を広範囲に評価し、LLM を利用した VHDL 生成に特有の主要な課題を明らかにします。私たちの調査結果は重要な洞察を提供し、多言語ハードウェア設計自動化における将来の作業をサポートします。私たちのデータ パイプライン、ベンチマーク、評価フレームワークはオープンソース化されます。

原文 (English)

VHDLSuite: Unified Pipeline for LLM VHDL Generation with Data Synthesis and Evaluation

Large Language Models (LLM) have shown impressive capabilities in Register Transfer Level (RTL) code generation, particularly for Verilog. However, evaluating their performance with other Hardware Description Languages (HDL), especially VHDL, remains limited although its distinct language characteristics, such as stricter semantic rules, introduce evaluation considerations that differ from Verilog. This lack of coverage restricts fully understanding of how well current models generalize across hardware design languages with differing structures and semantics. To address this gap, we introduce VHDLSuite, a benchmark-centered infrastructure for scalable VHDL generation evaluation, integrating automated benchmark synthesis, executable validation, and multi-model diagnostic analysis. First, we propose a data pipeline that automatically converts Verilog designs and their accompanying testbenches into executable VHDL benchmark instances, followed by VUnit/GHDL-based validation to ensure each released task is compilable, runnable, and consistently checkable in the VHDL environment. Second, we introduce VHDLBench, a benchmark with over 200 VHDL problems with complete and validated testbenches across a wide range of complexity levels. Third, we extensively evaluate cutting-edge LLMs and uncover key challenges specific on LLM-aided VHDL generation. Our findings provide important insights and support future work in multi-language hardware design automation.Our data pipeline, benchmark, and evaluation framework will be open-sourced.

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

OdysSim: 人間の行動シミュレーションのための基礎モデルの構築

大規模な言語モデルは、インタラクティブな評価や社会シミュレーションのためのヒューマン シミュレーターとして導入されることが増えています。しかし、役に立つことを重視したトレーニング後のトレーニングは、彼らを同質で過度に好意的なアシスタント登録に引き寄せ、行動面での Sim2Real のギャップを生み出します。私たちは、行動基盤モデル、つまり人間の行動を大規模にシミュレートするために訓練されたモデルの最大のオープンで体系的な調査である OdysSim を紹介します。私たちは、62 のデータセットと 23 のベンチマーク タスクを 1 つのフレームワークの下に統合する 5 つの能力軸 (CONV、SS、COG、ROLE、EVAL) の分類である SOUL を提案します。具体的には、OdysSim コーパス (2,140 万のインタラクション、100 億トークン、バック生成されたソーシャル コンテキストで改良) をキュレートし、SOUL-Index ベンチマークを構築し、ミッドトレーニング、タスク固有の RL、エキスパートの蒸留を組み合わせたエンドツーエンドのトレーニング レシピを開発します。結果として得られたオープン 8B OSim モデルは、23 タスク中 8 タスクで 1 位または同率 1 位にランクされ、このカウントで個々のフロンティア モデルを上回り、会話タスクとソーシャル タスクで最も優れた効果を発揮しました。また、その出力は長さ、形式、単語の選択においてより人間らしくなり、$\tau$-bench でのゼロショットから配布外のユーザー シミュレーションに移行し、反応の調整において実際のユーザーとほぼ一致します (93.2 対 93.5)。さらに、LLM-as-judge RL が報酬ハッキング パターンを誘発し、検出器がトレーニング後のパターンを軽減できることを示します。まとめると、私たちの調査結果は、行動基盤モデルでは LLM トレーニング パラダイムを再考する必要があることを示唆しています。私たちは将来の研究をサポートするためにすべての成果物を公開します。

原文 (English)

OdysSim: Building Foundation Models for Human Behavior Simulation

Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-driven post-training pulls them toward a homogeneous, overly agreeable assistant register, creating a behavioral Sim2Real gap. We present OdysSim, the largest open systematic investigation of behavioral foundation models, i.e., models trained to simulate human behavior at scale. We propose SOUL, a taxonomy of five capability axes (CONV, SS, COG, ROLE, EVAL) that unifies 62 datasets and 23 benchmark tasks under one framework. Specifically, we curate the OdysSim corpus (21.4M interactions, 10B tokens, retrofitted with back-generated social contexts), construct the SOUL-Index benchmark, and develop an end-to-end training recipe combining midtraining, task-specific RL, and expert distillation. The resulting open 8B OSim model ranks first or tied-first on 8 of 23 tasks, outperforming any individual frontier model by this count, with the strongest gains on conversational and social tasks. Its outputs are also more human-like in length, formatting, and word choice, and it transfers zero-shot to out-of-distribution user simulation on $\tau$-bench, nearly matching real users on reaction alignment (93.2 vs. 93.5). We further show that LLM-as-judge RL induces reward-hacking patterns, and that our detectors can mitigate them during post-training. Together, our findings suggest that behavioral foundation models require rethinking the LLM training paradigm. We release all artifacts to support future research.

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automotive Contexts

This paper appraises recent frameworks within AI development to integrate LLMs into control tasks in automotive contexts from the perspecti…

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers

This paper studies how multimodal large language models (MLLMs) undermine the security guarantees of visual CAPTCHA. We identify the attack…

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

評価がどのように設計されているかを知っているモデルは、より安全なスコアを獲得します

AI の安全性評価の有効性は、制御設定および展開設定全体でモデルが一貫して動作するかどうかに依存します。これまでの研究では、言語化された評価の認識とその後の行動の変化の源として、仮説的なシナリオなどのテスト時の文脈上の手がかりが特定されてきました。この論文では、この現象の潜在的な説明である評価メタ知識を調査します。これは、評価を特徴付ける構造的特性に関するパラメトリック知識として定義されます。ベンチマークの公開が記憶を通じてパフォーマンスの向上につながるデータセットの汚染と同様に、評価の実践を説明するテキストでトレーニングされたモデルは、たとえば AI ベンチマークに関する科学記事やソーシャル メディアの投稿への公開を通じて、評価のようなコンテキストを認識して応答することを暗黙的に学習する可能性があると仮説を立てています。これをテストするために、検証可能な構造や道徳的ジレンマなどの評価特性を記述する合成文書のモデルを微調整します。この微調整されたモデルを 6 つの安全性ベンチマークで評価すると、基本モデルや制御モデルよりも大幅に安全であることがわかりました。この行動の変化は、分析を評価意識の明示的な言語化が欠けている回答に限定した場合でも持続します。私たちの結果は、評価のメタ知識が安全ベンチマークのパフォーマンスを水増しし、明示的な記憶や言語化された評価の認識とは独立した新たな交絡因子を導入する可能性があることを示しており、したがって検出が困難です。これらの発見は、AI の安全性評価の設計と解釈に重要な意味を持ちます。コードとモデルは https://github.com/compass-group-tue/arxiv2026_evaluation_meta_knowledge で入手できます。

原文 (English)

Models That Know How Evaluations Are Designed Score Safer

The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized evaluation awareness and subsequent behavioral shift. In this paper, we investigate a potential explanation of this phenomenon: evaluation meta-knowledge, defined as parametric knowledge about the structural traits that characterize evaluations. Similar to dataset contamination, where benchmark exposure leads to higher performance through memorization, we hypothesize that models trained on texts describing evaluation practices may implicitly learn to recognize and respond to evaluation-like contexts, for instance, through exposure to scientific articles or social media posts about AI benchmarking. To test this, we fine-tune models on synthetic documents describing evaluation traits such as verifiable structures or moral dilemmas. Evaluating this fine-tuned model on five safety benchmarks, we find that it is significantly safer than the base model and control model. This behavioral shift persists even when restricting the analysis to responses lacking explicit verbalization of evaluation awareness. Our results demonstrate that evaluation meta-knowledge may inflate safety benchmark performance, introducing a novel confounder that is independent of explicit memorization or verbalized evaluation awareness, thus, challenging to detect. These findings have important implications for the design and interpretation of AI safety evaluations. Our code and models are available at https://github.com/compass-group-tue/arxiv2026_evaluation_meta_knowledge.

2026-06-15 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達研究/論文

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily…

2026-06-15 01:38 JSTTechCrunch AIビジネス/資金調達

As AI companies race to go public, who else is along for the ride?

Startups are trying to "ride that SpaceX IPO wave."

2026-06-14 09:03 JSTTechCrunch AIビジネス/資金調達

Meta reportedly moves to unwind $2B Manus deal after Beijing’s demand

Meta starts dismantling its $2 billion Manus acquisition after Beijing ordered the deal reversed.

2026-06-14 04:11 JSTTechCrunch AILLM/生成AIビジネス/資金調達

Amazon CEO reportedly raised Anthropic model concerns before government crackdown

Amazon CEO Andy Jassy may have been the source of security concerns that led Anthropic to cut off worldwide access to two models on Friday.

2026-06-13 08:15 JSTTechCrunch AIビジネス/資金調達

SpaceX IPO: Live updates on everything you need to know

TechCrunch has followed SpaceX's start, struggles, and successes from the early days. And we're here for what happens next too. This packag…

2026-06-13 02:38 JSTTechCrunch AIビジネス/資金調達

Mistral is rumored to be raising €3B at €20B valuation

The funding round would value the company at around €20 billion (about $23.15 billion), nearly double its Series C valuation of €11.7 billi…

2026-06-13 01:23 JSTTechCrunch AILLM/生成AIビジネス/資金調達

SpaceX, Anthropic, and OpenAI’s hot IPO summer

The IPO market is back, and it’s not the same companies leading the charge. FAANG had a good run, but a new acronym is taking over: MANGOS…

2026-06-13 00:50 JSTTechCrunch AIビジネス/資金調達

It’s hot IPO summer, and the MANGOS are ripe

The IPO market is back, and it’s not the same companies leading the charge. FAANG had a good run, but a new acronym is taking over: MANGOS…

2026-06-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

導入中心の評価: 臨床 LLM システムにおけるクエリレベルの拒否リスクの予測

大規模言語モデル (LLM) は臨床システムにますます統合されており、これらのシステムの実世界での有用性を評価することが不可欠になっています。ただし、静的ベンチマークはユーザーの受け入れではなく正確さを測定する傾向があり、クエリ全体のパフォーマンスを集約し、注釈が密に付加されたデータセットを必要とするため、臨床システムを評価する際の大きな盲点につながります。この研究では、学術医療センターの電子医療記録に埋め込まれた LLM システムの導入中心の評価を実行します。そこでは、ユーザーからのフィードバックはまばらですが、導入状況が厳密に反映されています。具体的には、生成前に利用可能なクエリの内容とデプロイメント固有のコンテキストに基づいて、今後のインタラクションによってユーザーが LLM 応答を拒否するリスクを推定する応答前分類器をトレーニングします。ユーザーからのフィードバックを 4.5 か月にわたって収集し、モデルの前向き分析を実施したところ、予測モデルが AUROC 0.719 を達成していることがわかりました。さらに、2 つの下流のユースケース (ガードレールのトリガーと棄権) におけるそのような予測の利点を推定します。私たちの重要な概念的洞察は、クエリ内容だけではなく、展開固有のコンテキスト (つまり、プロバイダーの種類、部門名、応答に使用される言語モデル) を利用することで、ユーザーがシステム出力を拒否するかどうかを予測する能力が向上するということです。まとめると、私たちの実証的なケーススタディは、展開固有のコンテキストを使用してユーザーの拒否を予測し、ターゲットを絞ったガードレールへの扉を開く実現可能性を示しています。

原文 (English)

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System

Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks tend to measure correctness rather than user acceptance, aggregate performance across queries, and require densely annotated datasets -- leading to major blind spots for evaluating clinical systems. In this work, we perform a deployment-centered evaluation of an LLM system embedded within electronic health records at an academic medical center, where user feedback is sparse but closely reflects the deployment conditions. Specifically, we train a pre-response classifier that estimates the risk that a future interaction will result in the user rejecting the LLM response, based on query content and deployment-specific context available before generation. We conduct a prospective analysis of our model over 4.5 months of user feedback, finding that our prediction model achieves an AUROC of 0.719. Further, we estimate the benefit of such predictions in two downstream use cases (guardrail triggering and abstention). Our key conceptual insight is that making use of deployment-specific context (i.e., the provider type, department name, language model used for response), as opposed to only query content, improves the ability to predict whether the user will reject the system output. Altogether, our empirical case study demonstrates the feasibility of predicting user rejection using deployment-specific context, opening the door to targeted guardrails.

2026-06-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM の心理測定的評価の再考: 自己申告が行動を予測するときとその理由

低コストの心理測定プローブから LLM の行動傾向を予測することは、安全な導入のために重要ですが、それは自己報告 (SR) が行動を確実に予測する場合に限られます。最近の研究では、LLMにおけるSR行動の実質的な解離が記録されていますが、人間であっても特定の行動を弱く予測する広範な性格特性(ビッグ5)に依存していました。さらに、会話セッションの分離と弱いコンテキスト マッチングの組み合わせでは、LLM に本当に一貫性が欠けているのか、あるいはそのような一貫性を検出するために必要な条件が満たされていないのかどうかは不明のままです。私たちはビッグ 5 を計画行動理論 (TPB) と対比します。TPB は特定の行動を対象とした意図を測定し、広範な特性よりも大幅に人間の行動を予測します。セッション コンテキストとアイデンティティ誘導も変化させながら、4 つの行動タスクと 11 のフロンティア LLM にわたって実験を実行します。 SR の動作の一貫性は存在しますが、選択的であることがわかりました。 1) 共有された会話の中で、計画的行動理論は人間レベルの一貫性に達します。 Big 5 はそうではありません。 2) 別々の会話では、一貫性はトレーニングによって形成された暗黙のバイアスなど、直接のプロンプトの外側に固定された行動の場合にのみ存続し、お調子者のように行動が文脈によって強く刺激されると崩壊します。 3) ペルソナのプロンプトにより、会話全体での自己報告の一貫性が高まりますが、行動が一致するわけではありません。これらの調査結果は、Big 5 などの粗いパーソナリティ フレームワークが、展開動作をテストするための最適なツールではない可能性があることを示唆しています。タスクおよび動作に特化した手段がさらに必要ですが、これらの手段もタスクやコンテキスト全体で評価する必要があります。

原文 (English)

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior

Anticipating LLM behavioral tendencies from low-cost psychometric probes is critical for safe deployment, but only if self-reports (SR) reliably predict behavior. Recent work documented substantial SR-behavior dissociation in LLMs, but relied on broad personality traits (Big 5) that predict specific behaviors weakly, even in humans. Furthermore, the isolation of conversational sessions combined with weak context matching left open whether LLMs truly lack coherence or whether the conditions needed to detect such coherence were not met. We contrast Big 5 with the Theory of Planned Behavior (TPB), which measures intention targeted to a specific behavior and predicts human behavior substantially better than broad traits. We run experiments across four behavioral tasks and 11 frontier LLMs, while also varying session context and identity induction. We find that SR-behavior coherence exists but is selective. 1) Within a shared conversation, the Theory of Planned Behavior reaches human-level coherence; Big 5 does not. 2) Across separate conversations, coherence survives only for behaviors anchored outside the immediate prompt, such as implicit bias shaped by training, and collapses when behavior is strongly primed by context, as with sycophancy. 3) Persona prompting makes self-reports more consistent across conversations, but does not bring behavior into alignment. These findings suggest that coarse personality frameworks, such as Big 5 may not be the best tools for testing deployment behavior. More task- and behavior-specific instruments are needed, and even these must be evaluated across tasks and contexts.

2026-06-12 13:00 JSTarXiv cs.AIビジネス/資金調達

大規模言語モデルにおける事前入力の認識

アライメントやジェイルブレイク評価、AI 制御プロトコルなど、言語モデルの安全性関連の研究は、多くの場合、事前入力モデルの出力に依存します。 AI モデルが、以前のアシスタント メッセージが挿入または編集されたという事実を認識し、それに基づいて動作できる場合、これらの方法の有効性と妥当性が損なわれる可能性があります。私たちは、フロンティア言語モデルが、改ざんされたアシスタント側コンテキストと改ざんされていないアシスタント側コンテキストを区別できるかどうか、つまりプレフィル認識と呼ぶ機能を調査します。そのために、モデルが一貫したスタンスを示すケースをフィルタリングして、3 つのプレフィル メカニズムにわたるバイナリ優先ベンチマークを構築します。フロンティア モデルはかなりのプレフィル認識を示していることがわかりました。Claude Opus 4.5 は、プロンプトが表示された場合、9 ~ 35% のケースで、その好みに反するプレフィルを検出し、偽陽性率は 0% でした。さらに、モデルは、プレフィルが外部のものであることを明示的に報告せずに、ベースラインの動作に戻ることがよくあります。後の制御されたアブレーションでは、検出と抵抗が異なる手がかりに依存していることも示されており、スタイルの不一致は主にモデルがプレフィルに異物としてフラグを立てるかどうかに影響し、一方、好みの不一致は主にモデルがベースラインの答えに戻るかどうかに影響します。また、ミスアライメント継続評価や SWE ベンチ軌道など、より現実的なエージェント設定も検討します。フロンティア モデルでは、データセット、タスクの成功、および隠れた書式設定アーティファクトに強く依存する形で、事前に入力されたアシスタント ターンが否認されることがあります。私たちの結果は、プレフィルの認識が、一部のプレフィルベースの手法にとってすでにかなりの混乱を引き起こしていることを示しています。モデル開発者は、フロンティア システムでこの機能を追跡することをお勧めします。

原文 (English)

Prefill Awareness in Large Language Models

Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs. If AI models can recognize and act on the fact their prior assistant messages have been inserted or edited, the effectiveness and validity of these methods could be compromised. We investigate whether frontier language models can distinguish between tampered and untampered assistant-side context, a capability we call prefill awareness. To do so, we construct a binary preference benchmark across three prefill mechanisms, filtering for cases where models show consistent stances. We find that frontier models show substantial prefill awareness: Claude Opus 4.5 detects prefills opposing its preferences in 9-35% of cases with a 0% false positive rate when prompted; additionally, models often revert towards baseline behavior without explicitly reporting that the prefill was foreign. Controlled ablations later also show that detection and resistance rely on different cues, where stylistic mismatch mainly affects whether models flag a prefill as foreign, while preference mismatch mainly affects whether they revert toward their baseline answer. We also examine more realistic agentic settings such as misalignment-continuation evaluations and SWE-bench trajectories, where frontier models sometimes disavow prefilled assistant turns in ways that depend strongly on dataset, task success, and hidden formatting artifacts. Our results indicate that prefill awareness is already a substantial confound for some prefill-based methods. We recommend that model developers track this capability in frontier systems.

2026-06-12 13:00 JSTarXiv cs.AIビジネス/資金調達

手続き型推論のための評価データセットの構築: 自然さ、グラウンディング、およびマルチホップ カバレッジのバランスをとる

AI 支援学習システムにおける手続き推論を評価するには、学習者らしく、システムが使用することが期待される教育知識に基づいた質問と回答のデータセットが必要です。私たちは、TMK ベースの質問生成戦略が、手続き型およびマルチホップ推論のデータセットの品質にどのような影響を与えるかを研究します。タスク・メソッド・ナレッジ(TMK)モデルからの厳密な生成、ポストホック TMK フィルタリングを使用したトランスクリプトファースト生成、トランスクリプトと構造化ガイダンスを組み合わせた TMK を意識した生成の 3 つの戦略を比較します。生成された項目を評価するために、TMK モデルから抽出された閉集合証拠単位に基づく根拠検証フレームワークを導入します。このフレームワークは、回答が基礎となる表現によってサポートされているかどうか、質問が自己完結型であるかどうか、およびマルチホップの手続き推論を対象としているかどうかを測定します。 23 の指導トピックと 690 の生成された質問と回答のペアにわたって、厳密な TMK 生成により、96.5% の根拠のある質問と 92.6% の使用可能な質問という、最高の全体的な品質が達成されます。 Transcript-first 生成では、より学習者らしい質問が生成されますが、よりコンテキストに依存した、または根拠の弱い項目が生成されます。一方、TMK を意識した生成では、生のマルチホップ カバレッジは高くなりますが、根拠は低くなります。これらの結果は、手続きの豊かさと自然な表現が表現の根拠を保証するものではなく、AI 支援学習における評価データセットの明示的な表現を意識した検証の動機となることを示しています。

原文 (English)

Constructing Evaluation Datasets for Procedural Reasoning: Balancing Naturalness, Grounding, and Multi-Hop Coverage

Evaluating procedural reasoning in AI-supported learning systems requires question-answer datasets that are both learner-like and grounded in the instructional knowledge the system is expected to use. We study how TMK-based question generation strategies affect dataset quality for procedural and multi-hop reasoning. We compare three strategies: strict generation from Task-Method-Knowledge (TMK) models, transcript-first generation with post-hoc TMK filtering, and TMK-aware generation that combines transcripts with structured guidance. To evaluate generated items, we introduce a grounding validation framework based on closed-set evidence units extracted from TMK models. The framework measures whether answers are supported by the underlying representation, whether questions are self-contained, and whether they target multi-hop procedural reasoning. Across 23 instructional topics and 690 generated question-answer pairs, strict TMK generation achieves the strongest overall quality, with 96.5% grounded questions and 92.6% usable questions. Transcript-first generation produces more learner-like questions but more context-dependent or weakly grounded items, while TMK-aware generation yields high raw multi-hop coverage but lower grounding. These results show that procedural richness and natural phrasing do not guarantee representational grounding, motivating explicit representation-aware validation for evaluation datasets in AI-supported learning.

2026-06-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MLUBench: MLLM における生涯未学習評価のベンチマーク

マルチモーダル大規模言語モデル (MLLM) は、大規模なマルチモーダル データでトレーニングされるため、データ所有者が特定のコンテンツの削除を要求する可能性があるため、データの非学習の重要性がますます高まっています。実際には、これらのリクエストは時間の経過とともに順番に届くことが多く、MLLM 生涯学習の困難な問題が生じます。しかし、既存のベンチマークのほとんどは規模と範囲が限られており、MLLM の生涯にわたる非学習の複雑さを捉えることができません。このギャップを埋めるために、生涯にわたる非学習要求に基づく 9 つのクラスにわたる 127 のエンティティを特徴とする大規模かつ包括的なベンチマークである MLUBench を導入します。私たちは MLUBench を使用して広範な実験を実行し、既存の非学習手法が深刻な累積的な劣化を受けていることを明らかにしました。さらに重要なことに、私たちはこの問題の特有の課題をさらに特定します。単峰性モデルとは異なり、MLLM の生涯にわたる非学習は、多峰性の調整を維持する必要性によって制約されます。 1 つのモダリティから継続的に学習を解除すると、モデル全体が劣化する可能性があります。この課題を軽減するために、私たちは効果的な方法である LUMoE を提案します。実験では、LUMoE がベースラインが直面する劣化の問題を大幅に軽減することが実証されています。ソース コードと MLUBench データセットは、https://github.com/lihe-maxsize/Lifelong_Unlearning_main でオープンソース化されています。

原文 (English)

MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs

Multimodal large language models (MLLMs) are trained on massive multimodal data, making data unlearning increasingly important as data owners may request the removal of specific content. In practice, these requests often arrive sequentially over time, giving rise to the challenging problem of MLLM Lifelong Unlearning. However, most existing benchmarks are limited in scale and scope, failing to capture the complexities of MLLM lifelong unlearning. To fill this gap, we introduce the MLUBench, a large-scale and comprehensive benchmark featuring 127 entities across 9 classes under lifelong unlearning requests. We perform extensive experiments using MLUBench and reveal that existing unlearning methods suffer from severe, cumulative degradation. More critically, we further identify the unique challenge of this problem: unlike in unimodal models, MLLM lifelong unlearning is constrained by the need to preserve multimodal alignment. Continually unlearning from one modality could degrade the entire model. To alleviate this challenge, we propose LUMoE, an effective method. Experiments demonstrate that LUMoE significantly mitigates the degradation problem faced by baselines. The source code and the MLUBench dataset are open-sourced in https://github.com/lihe-maxsize/Lifelong_Unlearning_main.

2026-06-12 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

配電欠陥検出のためのマルチモーダル エージェント: 基礎モデルの評価

配電ネットワークは信頼性の高い電力供給に不可欠ですが、従来の検査方法では意味の理解、一般化、閉ループの自動化において限界に直面しています。これらの課題に対処するために、この文書では、配電欠陥検出に特化したマルチモーダル エージェント フレームワークを提案します。この研究の中心となるのは、統合された認知エンジンとしてのマルチモーダル基盤モデルの体系的な評価です。当社は、次の 3 つの重要な機能にわたって統合されたパフォーマンスを厳密に評価します。(1) 認識。モデルは機器を正確に識別し、専門家レベルの欠陥の説明を生成する必要があります。 (2) 推論。モデルは視覚的な所見を解釈して原因を診断し、重大度を評価し、ドメイン知識に基づいてメンテナンス戦略を計画します。 (3) ツールの使用。モデルは自律的なオペレーターとして機能し、ナレッジ ベースのクエリや作業指示の生成などのアクションを実行して、クローズド ループ メンテナンスを実現します。この評価をサポートするために、ドメイン固有の評価データセットと包括的なベンチマークが開発されています。実験結果は、これら 3 つの側面における現在の基盤モデルの強みと限界を実証し、一か八かの産業環境に自律エージェントを導入するための経験的証拠を提供します。

原文 (English)

Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models

The power distribution network is critical to reliable electricity delivery, yet traditional inspection methods face limitations in semantic understanding, generalization, and closed-loop automation. To address these challenges, this paper proposes a Multi-Modal Agent framework specifically for power distribution defect detection. Central to this study is the systematic evaluation of multimodal foundation models as unified cognitive engines. We rigorously assess their integrated performance across three critical capabilities: (1) Perception, where the model must accurately identify equipment and generate expert-level descriptions of defects; (2) Reasoning, where the model interprets visual findings to diagnose causes, assess severity, and plan maintenance strategies based on domain knowledge; and (3) Tool Usage, where the model acts as an autonomous operator to execute actions -- such as querying knowledge bases or generating work orders -- to achieve closed-loop maintenance. To support this evaluation, a domain-specific evaluation dataset and a comprehensive benchmark are developed. Experimental results demonstrate the strengths and limitations of current foundation models in these three dimensions, providing empirical evidence for deploying autonomous agents in high-stakes industrial environments.

2026-06-12 13:00 JSTarXiv cs.AIビジネス/資金調達

メタデータ駆動型分類における評価主権: 弱く監視された情報システムのためのマルチトラック フレームワーク

機械学習における評価は通常、中立的な測定プロセスとして扱われます。ただし、運用情報システムでは、ラベルの生成に使用されるプロセスによって評価結果が条件付けされることがよくあります。この文書は、分類パフォーマンスの向上を目指したものではありません。代わりに、さまざまなレーベル権限体制の下でのパフォーマンス測定の妥当性を検証します。この問題は、ラベルが不完全、一貫性がない、または監視が不十分であることが多い大規模なメタデータ駆動型システムに特に関係します。我々は、パフォーマンス指標がラベルの権限や監督体制から独立している度合いとして定義される評価主権を導入し、トレーニングと評価のラベルソースを系統的に変更するマルチトラック評価フレームワークを提案します。大規模な科学メタデータの階層的マルチラベル分類を使用して、運用 (「シルバー」) 評価で優れたパフォーマンスを示すモデルが、独立 (「ゴールド」) 評価では、特に詳細な分類の場合に大幅に低下することを実証します。たとえば、Micro-F1 は約 0.54 から 0.03 に減少します。特に、ランキングベースのメトリクスはベースラインを上回ったままであり、潜在モデル信号と分類の妥当性の間の乖離が明らかになっています。これらの調査結果は、一般的に報告されるパフォーマンス指標が、真の予測能力ではなく、ラベル付けプロセスとの整合性を反映している可能性があることを示唆しています。したがって、私たちは評価の妥当性をラベルガバナンスによって形成されるシステムレベルの特性として再概念化し、弱い監視下で動作するインテリジェントシステムを監査するための実用的な方法論を提供します。

原文 (English)

Evaluation Sovereignty in Metadata-Driven Classification: A Multi-Track Framework for Weakly Supervised Information Systems

Evaluation in machine learning is typically treated as a neutral measurement process. However, in operational information systems, evaluation outcomes are often conditioned by the processes used to generate labels. This paper does not seek to improve classification performance. Instead, it examines the validity of performance measurement under differing label-authority regimes. This issue is particularly relevant in large-scale metadata-driven systems, where labels are often incomplete, inconsistent, or weakly supervised. We introduce evaluation sovereignty, defined as the degree to which performance metrics are independent of label authority and supervision regime, and propose a multi-track evaluation framework that systematically varies training and evaluation label sources. Using hierarchical multi-label classification on large-scale scientific metadata, we demonstrate that models exhibiting strong performance under operational ("silver") evaluation degrade substantially under independent ("gold") evaluation, particularly for fine-grained classification. For example, Micro-F1 decreases from approximately 0.54 to 0.03. Notably, ranking-based metrics remain above baseline, revealing a divergence between latent model signal and classification validity. These findings suggest that commonly reported performance metrics may reflect alignment with labeling processes rather than true predictive capability. We therefore reconceptualize evaluation validity as a system-level property shaped by label governance and provide a practical methodology for auditing intelligent systems operating under weak supervision.

2026-06-12 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

EpiBench: エピゲノミクス分析における AI エージェントの検証可能な評価

短期間のエピゲノミクス分析のための検証可能なベンチマークである EpiBench を紹介します。 EpiBench は、エージェントが現実的なワークフロー状態から明確に定義された分析決定を下し、決定論的に評価可能な回答を返すことができるかどうかを評価します。このベンチマークには、CUT\&Tag/CUT\&RUN、ATAC-seq、ChIP-seq、および DNA メチル化ワークフローにわたる 106 の評価が含まれています。 16 のモデルとハーネスのペアからの 5,088 の有効な軌道にわたって、大部分の試行に合格したシステムはありませんでした。GPT-5.5 / Pi が 45.0\% (143/318 試行; 95\% 信頼区間 (CI)、36.3--53.7) でトップとなり、GPT-5.5 / OpenAI Codex が 39.9\% (127/318 試行; 95\% CI、31.6--48.3)。 Claude Opus 4.8 Max / Pi および GPT-5.4 / Pi はそれぞれ 39.0% を合格しました (試行回数 124/318、95% CI、それぞれ 30.2 ~ 47.8 および 31.0 ~ 47.0)。パフォーマンスはアッセイの種類によって異なり、失敗した実行の多くには依然として正解の一部が含まれています。エージェントは多くの場合、適切なファイルを見つけて有用な中間結果を計算しましたが、タスクがより深い、アッセイ固有の科学的判断を必要とする場合には失敗しました。

原文 (English)

EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis. EpiBench evaluates whether agents can make well-defined analysis decisions from realistic workflow states and return deterministically gradable answers. The benchmark includes 106 evaluations across CUT\&Tag/CUT\&RUN, ATAC-seq, ChIP-seq, and DNA methylation workflows. Across 5,088 valid trajectories from 16 model-harness pairs, no system passed a majority of attempts: GPT-5.5 / Pi led at 45.0\% (143/318 attempts; 95\% confidence interval (CI), 36.3--53.7), followed by GPT-5.5 / OpenAI Codex at 39.9\% (127/318 attempts; 95\% CI, 31.6--48.3). Claude Opus 4.8 Max / Pi and GPT-5.4 / Pi each passed 39.0\% (124/318 attempts; 95\% CI, 30.2--47.8 and 31.0--47.0, respectively). Performance varies across assay types, and many failed runs still contain parts of the correct answer. Agents often found the right files and computed useful intermediate results, but failed when the task required deeper, assay-specific scientific judgment.

2026-06-12 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

AgentBeats: オープン性、標準化、再現性のためのエージェント化エージェントの評価

エージェント システムはドメイン間で急速に進歩していますが、その評価は依然として断片的です。ほとんどのベンチマークは、固定された LLM 中心のハーネスに依存しており、これには高度な統合が必要で、テストと運用の不一致が生じ、多様なエージェント設計間の公正な比較が制限されます。根本的な問題は、オープンでエージェントに依存しない評価インターフェイスが欠如していることです。当社は、評価が審査員エージェントによって実行され、すべての参加者がタスク管理用の A2A とツール アクセス用の MCP という標準化されたプロトコルを通じて対話するエージェント化エージェント評価 (AAA) を提唱しています。従来のベンチマークでは、ベンチマーク用とエージェント用の 2 つの個別のインターフェイスが定義されていましたが、AAA では 1 つだけが必要でした。これにより、評価ロジックをエージェントの実装から分離し、再現可能で相互運用可能な複数エージェントの評価を可能にする、汎用的で統一されたフレームワークが得られます。さらに、AAA の具体的な実現として AgentBeats を紹介します。オープン性、プライバシー、再現性に関する現実世界の制約と互換性のある標準化された評価を可能にする 5 つの実用的な動作モードを特定します。私たちの設計を大規模に評価するために、私たちは 2 つの調査を実施しました。1 つは 12 カテゴリーにわたる 298 人の審査員エージェントと、独立した参加者からの 467 人の被験者エージェントを集めた 5 か月間にわたるオープン コンテストで、AAA が異質な範囲のベンチマークに適用されることを示しました。また、コーディングエージェントに関するケーススタディでは、エージェント化された評価が公的記録との忠実性を維持しながら、これまで欠けていた直接対決の結果が明らかになり、エージェント設計に関する研究上の洞察が得られることが確認されました。コミュニティ規模のフィールド調査と制御されたコーディングのケーススタディを組み合わせて、AAA が大規模な異種シナリオ全体にわたってカバレッジ、実用性、忠実性を提供することを検証します。 AAA と AgentBeats は共に、オープンで標準化された再現可能なエージェント評価への明確な道筋を提供します。

原文 (English)

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all participants interact through standardized protocols: A2A for task management and MCP for tool access. Conventional benchmarking defines two separate interfaces, one for the benchmark and one for the agent, while AAA only needs one; this yields a generic, unified framework that separates assessment logic from agent implementation and enables reproducible, interoperable, and multi-agent evaluation. We further introduce AgentBeats as a concrete realization of AAA: we identify five practical operation modes that make standardized assessment compatible with real-world constraints on openness, privacy, and reproducibility. To evaluate our design at scale, we conduct two studies: a five-month open competition that drew 298 judge agents across 12 categories together with 467 subject agents from independent participants, showing that AAA applies across a heterogeneous range of benchmarks; and a case study on coding agents that confirms agentified evaluation preserves fidelity with the public record while surfacing previously missing head-to-head results, yielding research insights about agent design. Combining a community-scale field study and a controlled coding case study, we verify that AAA delivers coverage, practicality, and fidelity across heterogeneous scenarios at scale. Together, AAA and AgentBeats offer a clear path toward open, standardized, and reproducible agent assessment.

2026-06-12 13:00 JSTarXiv cs.AIビジネス/資金調達

Muse Spark の安全性と準備に関するレポート

Muse Spark は、Meta によって開発された最新の大規模言語モデルです。このレポートでは、まず、Meta の高度な AI スケーリング フレームワークに基づく壊滅的なリスク ドメインの評価を、開始の決定に影響を与えた証拠とともに示します。次に、Muse Spark の広範なコンテンツの安全性や動作プロファイルなど、全体的な安全性には関連するものの、フレームワークによって管理される壊滅的なリスク領域の外にある追加の考慮事項について説明します。化学的および生物学的、サイバーセキュリティ、および制御不能リスクをカバーする当社の準備結果は、メタ AI 内での Muse Spark の展開を、当社の高度な AI スケーリング フレームワークの下で許容可能なレベルの残存リスクを提示しているかどうか評価します。私たちは、これらの壊滅的なリスク領域にわたるデュアルユース機能と高リスク機能を対象とした広範な評価を実施しました。これらの評価では、緩和前にリスクの上昇が特定され、安全策が適用される前に化学および生物学的能力が高度 AI スケーリング フレームワークに基づく「高リスク」カテゴリーに達する可能性が高いと評価されました。当社は特定されたリスクに対処する多層的な緩和策を実装しており、Muse Spark は化学や生物学における危険なワークフローに関連するさまざまなベンチマークにわたって最先端の拒否を実証しています。そこで、メタ AI の基盤モデルとして Muse Spark をリリースします。

原文 (English)

Muse Spark Safety & Preparedness Report

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framework, along with the evidence that informed our launch decision. We then discuss additional considerations, such as Muse Spark's broader content safety and behavioral profile, that are relevant to overall safety but fall outside the catastrophic risk domains governed by the Framework. Our preparedness results covering Chemical and Biological, Cybersecurity, and Loss of Control risks assess Muse Spark's deployment within Meta AI as presenting acceptable levels of residual risks under our Advanced AI Scaling Framework. We conducted a broad set of evaluations targeting dual-use and high-risk capabilities across these catastrophic risk domains. Those evaluations identified elevated risks prior to mitigations, with Chemical and Biological capabilities assessed as likely reaching the "high risk" category under the Advanced AI Scaling Framework before safeguards were applied. We have implemented a multi-layered set of mitigations that address the identified risks, and Muse Spark demonstrates state-of-the-art refusal across a range of benchmarks related to hazardous workflows in chemistry and biology. We therefore release Muse Spark as the underlying model of Meta AI.

2026-06-12 13:00 JSTarXiv cs.AIビジネス/資金調達

アルゴリズム立憲主義

社会生活への人工知能 (AI) の侵入がますます進んでおり、特に Google、Facebook、Apple、Amazon などの企業によって作成および管理されている情報圏内では、社会に重大なリスクが生じています。この記事では、すでにアルゴリズムによって部分的に管理されている Facebook のコンテンツ管理体制の詳細な分析を通じて、これらのリスクを検証します。 AI によってもたらされるガバナンス上の課題の解決策として文献でよく提案されている倫理工学の考え方は、いくつかの理由から不十分であると私たちは主張します。これに応じて、私たちは「アルゴリズム立憲主義」と呼ぶ代替枠組みを開発します。私たちのアプローチは 3 つの柱に基づいています。(a) コードの 2 つのレベルで構成される階層化アーキテクチャ: (i) 操作レベルまたはオブジェクト レベル、および (ii) アルゴリズムによって開始される変更からシステムの核となる原則を保護するように設計されたメタ レベル。 (b) アルゴリズムによるメタ推論。これにより、システムが両方のレベルで同時に動作できるようになり、メタコード レベルで保護された原則から逸脱するオブジェクト レベルでの動作をリアルタイムで監視、検証し、潜在的に修正できるようになります。 (c) 審議による修正。この記事では、アルゴリズム立憲主義の概念を詳しく説明し、それが Facebook のコンテンツモデレーション体制にどのように適用されるかを示しています。この分析の一環として、私たちは社会的立憲主義とアルゴリズム的立憲主義の間の緊張を検討します。逆説的ですが、AI システムを外部の熟議制御に従わせようとする試みは、AI エージェントがそのプロセスに介入することを可能にし、その目的を損なう可能性もあります。この記事は、この議論が 2022 年 10 月に発効した欧州デジタル サービス法に及ぼす影響を考察して締めくくられています。

原文 (English)

Algorithmic Constitutionalism

The increasing encroachment of artificial intelligence (AI) on social life raises significant risks for society, particularly within the infospheres created and controlled by companies such as Google, Facebook, Apple, and Amazon. This article examines these risks through an in-depth analysis of Facebook's content moderation regime, which is already partially governed by algorithms. We argue that the idea of ethical engineering, often proposed in the literature as a solution to the governance challenges posed by AI, is inadequate for several reasons. In response, we develop an alternative framework, which we term "algorithmic constitutionalism." Our approach rests on three pillars: (a) a layered architecture consisting of two levels of code: (i) an operative or object level and (ii) a meta level designed to protect the system's core principles from algorithmically initiated change; (b) algorithmic meta-reasoning, which enables the system to operate simultaneously at both levels so that it can monitor, verify, and potentially correct in real time operations at the object level that depart from principles protected at the meta-code level; and (c) correction through deliberation. The article elaborates the concept of algorithmic constitutionalism and demonstrates how it may be applied to Facebook's content moderation regime. As part of this analysis, we examine the tension between societal constitutionalism and algorithmic constitutionalism. Paradoxically, attempts to subject AI systems to external deliberative control may also enable AI agents to intervene in that process, potentially undermining its purpose. The article concludes by considering the implications of this argument for the European Digital Services Act, which entered into force in October 2022.

2026-06-12 13:00 JSTarXiv cs.AIビジネス/資金調達

SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning

Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…

2026-06-12 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: kernel-driven measurement instability and noise-driven detection fragility, invisible to DICOM metadata

AI governance for medical imaging is formalizing: the 2026 ACR-SIIM Practice Parameter recommends local acceptance testing and ongoing drif…

2026-06-12 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Workflow-GYM: 現実世界の専門分野におけるコンピュータ使用エージェントタスクの長期的な評価に向けて

近年、ますます複雑になる現実世界のタスクの処理に向けて、AI エージェントが急速に進化しています。しかし、既存のベンチマークでは、エージェントがグラフィカル ユーザー インターフェイスを操作して、さまざまなドメインにわたる長期にわたる価値の高い専門的なワークフローを完了できるかどうかを評価することはほとんどありません。現在の GUI ベンチマークは依然として、主に汎用ソフトウェア、比較的単純なアプリケーション、および短期間のタスクに焦点を当てており、最新のエージェントがユーザーの指示に従ってドメイン固有のプロフェッショナル ソフトウェアを自律的に操作し、経済的に価値のある作業をエンドツーエンドで実行できるかどうかはほとんど不明です。このギャップを埋めるために、専門分野と特殊なソフトウェア環境を中心とした長期的な GUI タスクのベンチマークである Workflow-GYM を導入します。最先端のモデルで広範な実験を行った結果、最も強力なモデルでも成功率は 30% をわずかに超える程度であることがわかり、プロの長期にわたる GUI ワークフローが現在の GUI エージェントにとって依然として非常に困難であることが浮き彫りになりました。さらなる分析により、現在のエージェントは長期的なワークフローの一貫性を維持するのに苦労しており、ワークフロー段階の省略、エラーの伝播、目標のずれ、プロフェッショナルなソフトウェア環境の理解不足が頻繁に見られることが明らかになりました。私たちの調査結果は、現在のエージェント システムの限界についての重要な洞察を提供し、次世代の GUI エージェント研究の重要な方向性を示唆しています。

原文 (English)

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.

2026-06-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達規制/政策

Reconstructing Template-Memorized Images from Natural Prompts

Recent advances in generative models, such as diffusion models, have raised concerns related to privacy, copyright infringement, and data s…

2026-06-12 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

GetNetUPAM: Ecologically Informed Nested Cross-Validation and Noise-Robust Attention for Marine Bioacoustic Monitoring

Deploying reliable bioacoustic monitoring systems requires models that generalize under high-noise, low-SNR conditions and evaluation proto…

2026-06-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings

With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regul…

2026-06-12 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem

The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by…

2026-06-12 13:00 JSTarXiv cs.AIビジネス/資金調達

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechani…

2026-06-12 10:48 JSTTechCrunch AIロボティクスビジネス/資金調達

Theker just raised $85M to build the factory robot that doesn’t specialize in anything

Unlike humanoid robots designed around a fixed form — think Boston Dynamics — Theker's machines are built to be reconfigured.

2026-06-12 10:04 JSTTechCrunch AIビジネス/資金調達

Jeff Bezos’s Prometheus raises $12B to build an ‘artificial general engineer’ for the physical world

The new round values the physical AI startup that aims to automate heavy engineering and drug design at $41 billion.

2026-06-12 05:33 JSTTechCrunch AIビジネス/資金調達

SpaceX officially prices shares at $135 in the largest IPO ever

Wits its official share pricing announcement, SpaceX's IPO has begun.

2026-06-12 04:58 JSTTechCrunch AIビジネス/資金調達

SpaceX SPV investors won’t know their true holdings until post-IPO lock-ups lift

After SpaceX makes its public debut, lower-tier SPV investors face hidden fees, lengthy payout delays, and the risk of outright fraud.

2026-06-11 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体ビジネス/資金調達研究/論文

医療研究分析用のスキル拡張 AI エージェント: NSCLC トランスクリプトーム バイオマーカー タスクにおける探索的なマルチモデルヒト評価

背景。生物医学研究をサポートするために大規模な言語モデルと AI エージェントがますます使用されていますが、ネイティブ モデルの出力では、重要な分析ステップが省略されたり、手法が誤用されたり、結論が誇張されたりする可能性があります。私たちは、医学研究スキル パッケージへの自律的なアクセスが、スキルを持たないネイティブ AI と比較して、AI によって生成されたトランスクリプトーム研究分析の高品質な出力に関連しているかどうかを評価しました。方法。私たちは、非小細胞肺がん免疫療法バイオマーカータスクを使用して、探索的なマルチモデルヒト評価を実施しました。 6 つのモデル バックボーンがテストされました。評価には、OpenClaw に代表される AI エージェント実装を通じて生成された 9 つのネイティブ AI 出力と 12 のスキル拡張出力の 21 件の匿名化された出力が含まれていました。 4 人の非専門生物医学評論家と 2 人の盲検専門家が各成果を評価し、各評論家のタイプごとに 2 つの評価を付けました。主な成果は、専門家が評価した全体的な品質でした。結果。スキル拡張された出力は、ネイティブ AI の出力よりも専門家の全体的な品質が方向性的に高いことを示しました (平均 5.50 vs 5.11; 差 = 0.39; ブートストラップ 95\% CI、-0.04 ~ 0.90; Welch p=0.156)。専門家以外の査読者の質も同じ傾向を示しました(平均 4.72 vs 4.47; 差 = 0.26; ブートストラップ 95\% CI、-0.25 ~ 0.80; Welch p=0.373)。専門家の合意は限られており (単一評価 ICC=-0.15)、モデル固有の効果は記述的で不均一でした。結論。この探索的サンプルでは、​​自律的スキル アクセスにより方向性のある品質シグナルが示されましたが、そのシグナルは専門家評価のノイズよりも小さいため、確認的な証拠として解釈されるべきではありません。この発見は主に、より強力な信頼性制御、プラットフォームの複製、生物学的妥当性評価を備えたスキル強化型 AI エージェントの大規模な評価の動機付けとなります。

原文 (English)

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

Background. Large language models and AI agents are increasingly used to support biomedical research, but native model outputs may omit key analytical steps, misuse methods, or overstate conclusions. We evaluated whether autonomous access to a medical research skill package was associated with higher-quality AI-generated transcriptomic research-analysis outputs compared with native AI without skills. Methods. We conducted an exploratory multi-model human evaluation using a non-small cell lung cancer immunotherapy biomarker task. Six model backbones were tested. The evaluation included 21 anonymized outputs: 9 native-AI outputs and 12 skill-augmented outputs generated through an AI agent implementation represented by OpenClaw. Four non-expert biomedical reviewers and two blinded experts evaluated each output, with two ratings from each reviewer type. The primary outcome was expert-rated overall quality. Results. Skill-augmented outputs showed directionally higher expert overall quality than native-AI outputs (mean 5.50 vs 5.11; difference=0.39; bootstrap 95\% CI, -0.04 to 0.90; Welch p=0.156). Non-expert reviewer quality showed the same direction (mean 4.72 vs 4.47; difference=0.26; bootstrap 95\% CI, -0.25 to 0.80; Welch p=0.373). Expert agreement was limited (single-rating ICC=-0.15), and model-specific effects were descriptive and heterogeneous. Conclusions. Autonomous skill access showed a directional quality signal in this exploratory sample, but the signal was smaller than expert-rating noise and should not be interpreted as confirmatory evidence. The findings primarily motivate larger evaluations of skill-augmented AI agents with stronger reliability controls, platform replication, and biological-validity assessment.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Nonslop: 人間と AI の共同執筆におけるゲーム化された実験

大規模言語モデル (LLM) の急速な普及により、AI 支援による創作の時代における人間の創造性と個人の表現について重大な疑問が生じています。人間はいつ AI の提案を採用しますか?個人の声にはどのような影響がありますか?この研究では、74 人の参加者 (214 回答) がプロンプトに回答し、執筆中に AI によって生成された単語の提案が利用できる、ゲーミフィケーションのライティング演習を通じてこれらの質問を調査しました。このゲームは、AI が人間の個性の残存物から学習しようとするディストピアの未来をシミュレートし、AI のような執筆を阻害します。そうすることで、AI が生成したすぐに利用できる提案を受け入れるなど、デフォルトの動作ではなく、本物のユーザーの好みを明らかにする条件を作成しようとします。これは、「役立つアシスタント」設計パターンを意図的に反転したものであることに注意してください。システムは、AI の提案を受け入れることを明示的に禁止しています。私たちは、クリエイティブなタスクにおける人間と AI のインタラクションに影響を与える要因を理解するために、さまざまなタスクの種類、ユーザーの行動、応答特性にわたるユーザーの行動パターンを分析します。この調査では、ユーザーが創造的な自律性を維持することを選択する場合と、ゲームのルールに違反して AI の支援を受け入れることを選択する場合に焦点を当てています。また、これらの選択が応答パターン、タスクの特性、およびユーザーの行動にどのように関連しているかについても検討します。このゲーム化されたアプローチは、人間と AI の本物の相互作用を研究するためのフレームワークと、AI によって強化された創造性における効率性と信頼性の間の緊張を理解するための挑発的なレンズの両方を提供します。

原文 (English)

Nonslop: A Gamified Experiment in Human-AI Collaborative Writing

The rapid proliferation of large language models (LLMs) raises critical questions about human creativity and individual expression in an era of AI-assisted creation. When do humans adopt AI suggestions, and what are the implications for individual voice? This study examines these questions through a gamified writing exercise where 74 participants (214 responses) replied to prompts while AI-generated word suggestions were available as they wrote. The game simulates a dystopian future in which an AI is attempting to learn from what remains of human individuality, and disincentivizes AI-like writing. In doing so, it attempts to create conditions that reveal authentic user preferences rather than default behaviors, such as accepting a readily available AI-generated suggestion. Note that this is a deliberate inversion of the "helpful assistant" design pattern; the system is explicitly forbidding you from accepting AI suggestions. We analyze user behavior patterns across different task types, user behaviors, and response characteristics to understand the factors influencing human-AI interaction in creative tasks. The study focuses on when users choose to maintain creative autonomy versus violating the rules of the game and accepting AI assistance. It also explores how these choices relate to response patterns, task characteristics, and user behavior. This gamified approach offers both a framework for studying authentic human-AI interaction and a provocative lens for understanding the tension between efficiency and authenticity in AI-augmented creativity.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

PoQ-Judge: 分散型 LLM 推論におけるコストを意識した品質証明のためのマルチアーキテクチャ評価フレームワーク

分散型 LLM 推論ネットワークには、品質証明 (PoQ) のための軽量で参照不要の品質評価が必要です。我々は、グラウンドトゥルースを参照せずにクエリと出力のペアをスコア化するために専用のジャッジモデルをトレーニングするフレームワークである PoQ-Judge を紹介します。私たちは、品質とコストのトレードオフ全体にわたって、TextCNN の審査員、MiniLM クロスエンコーダー、および DeBERTa の審査員という 3 つのアーキテクチャを研究します。 UltraFeedback と GPT ラベル付きのドメイン内データの 2 段階トレーニングを使用すると、最良のモデルは、保持されたテスト セットでグラウンド トゥルース プロキシとのピアソン相関が 0.747 に達し、以前の研究による参照ベースの評価者を上回りました。複合スコアリングにおける参照不要のコンポーネントとして、0.645 のピアソン相関を達成し、参照回答の必要性を排除しながら、最良の単一参照ベースの評価者と一致します。また、オンライン キャリブレーションでは意味論的な品質が主要な側面であることが特定され、カスケード評価ではわずかな品質損失のみでコストが 72.7 パーセント削減されることも示します。結果は要約よりも QA の方がはるかに強く、残された主な制限としてプロキシの品質が指摘されています。

原文 (English)

PoQ-Judge: A Multi-Architecture Evaluation Framework for Cost-Aware Proof-of-Quality in Decentralized LLM Inference

Decentralized LLM inference networks need lightweight, reference-free quality evaluation for Proof of Quality (PoQ). We present PoQ-Judge, a framework that trains dedicated judge models to score query-output pairs without ground-truth references. We study three architectures across the quality-cost tradeoff: a TextCNN judge, a MiniLM cross-encoder, and a DeBERTa judge. Using two-stage training on UltraFeedback plus GPT-labeled in-domain data, the best model reaches 0.747 Pearson correlation with the ground-truth proxy on a held-out test set, outperforming reference-based evaluators from prior work. As a reference-free component in composite scoring, it achieves 0.645 Pearson correlation, matching the best single reference-based evaluator while removing the need for reference answers. We also show that online calibration identifies semantic quality as the dominant dimension and that cascade evaluation reduces cost by 72.7 percent with only modest quality loss. Results are much stronger on QA than summarization, pointing to proxy quality as the main remaining limitation.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

おべっかの二面性評価: 合意の構造と介入の限界

活性化ステアリングは LLM の行動を変える可能性がありますが、標準的な評価では通常、おべっかを減らす方向が事実上正しい発言との一致を抑制するかどうかをテストしません。各トピックの両方のスタンスをテストするデュアルスタンス評価を導入し、Llama-3-8B-Instruct の重心差ステアリングに適用します。解離が見られます。モデルは、幾何学的に異なる部分空間におけるお調子者と事実の一致を表しますが、ステアリング方向は両方に均等に投影され、どちらかを区別してターゲットにすることはできません。したがって、この方向性は、お世辞的なものだけでなく、事実として正しい記述(たとえば、地球は丸いという)との一致を減少させます。 2 つの活性化グループの他のすべての静的特性は一致しており、行動の解離が生成ダイナミクス、または残差ストリーム解析では解決できない細粒構造から生じていることを示唆しています。このパターンは、一般的なギャップを示しています。アクティベーションから読み取り可能な表現は、アクティベーションを介して書き込み可能ではない可能性があります。

原文 (English)

Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Intervention

Activation steering can shift LLM behaviour, but standard evaluations do not typically test whether a sycophancy-reduction direction also suppresses agreement with factually correct statements. We introduce dual-stance evaluation, which tests both stances of each topic, and apply it to centroid-difference steering on Llama-3-8B-Instruct. We find a dissociation: the model represents sycophantic and factual agreement in geometrically distinct subspaces, yet the steering direction projects equally onto both and cannot differentially target either. The direction accordingly reduces agreement with factually correct statements (e.g. that the Earth is round) as well as sycophantic ones. All other static properties of the two activation groups are matched, suggesting the behavioural dissociation arises from generation dynamics or from finer-grained structure that residual-stream analysis cannot resolve. The pattern illustrates a general gap: representations that are readable from activations may not be writable through them.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

BioDivergence: 生物医学抄録における隠れた文脈矛盾のベンチマークと評価フレームワーク

生物医学的知見は研究間で矛盾しているように見えることがよくありますが、これらの違いの多くは真の矛盾というよりは文脈に依存しています。コホート、地理、アッセイプロトコル、疾患のサブタイプ、臨床環境の違いにより、両方の主張が局所的に有効になる可能性があります。既存の NLI および科学的主張検証ベンチマークは、そのようなケースを含意、矛盾、または中立的なものに還元し、相違の背後にある文脈構造を捉えることができません。これに対処するために、6 クラスの紛争分類法、13 軸の発散オントロジー、およびクレーム ペアごとの 4 つの構造化出力 (競合タイプ、発散軸、支配的交絡因子、および和解の説明) を備えた評価フレームワークである BioDivergence を導入します。当社は、5 つの生物医学ドメインにわたる 11,865 のクレーム ペアの論文に共通のシルバー ベンチマークである BioDivergence-Silver-v1.0 を、比較用の従来の重複排除バリアントと併せてリリースします。結果は、2 つのバリアント間の顕著なランキングの違いを示しており、微調整された参照モデルは記事の非結合設定の下で約 12 ポイント低下しましたが、Mistral-7B-Instruct-v0.3 は 842 例のプライマリ テスト セットで 0.5523 の精度と 0.3894 のコンテキスト F1 を達成しました。 BioDivergence は、文脈の相違と直接的な矛盾を区別し、記事レベルの暗記を真のタスク学習から分離する、より忠実な方法を提供します。

原文 (English)

BioDivergence: A Benchmark and Evaluation Framework for Hidden Contextual Contradictions in Biomedical Abstracts

Biomedical findings often seem to conflict across studies, but many of these differences are context-dependent rather than true contradictions. Variations in cohort, geography, assay protocol, disease subtype, and clinical setting can make both claims locally valid. Existing NLI and scientific claim-verification benchmarks reduce such cases to entailment, contradiction, or neutral, failing to capture the contextual structure behind divergence. To address this, we introduce BioDivergence, an evaluation framework with a six-class conflict taxonomy, a 13-axis divergence ontology, and four structured outputs per claim pair: conflict type, divergence axes, dominant confounder, and reconciliation explanation. We release BioDivergence-Silver-v1.0, an article-disjoint silver benchmark of 11,865 claim pairs across five biomedical domains, alongside a legacy deduplicated variant for comparison. Results show notable ranking differences between the two variants, with the fine-tuned reference model dropping about 12 points under the article-disjoint setting, while Mistral-7B-Instruct-v0.3 achieves 0.5523 accuracy and 0.3894 contextual-F1 on the 842-example primary test set. BioDivergence offers a more faithful way to distinguish contextual divergence from direct contradiction and to separate article-level memorization from genuine task learning.

2026-06-11 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

倫理 eValuation エージェント (EeVA): 倫理的審議を支援するプロトタイプ エージェントのようなワークフローの概念実証テストの結果

倫理的な熟慮は、単一の正解または不正解の答えを探すことであると誤解されることが多く、倫理的な課題に対処しなければならない倫理的な訓練を受けていない職員にとっては困難が生じます。私たちは、最終的な倫理的答えを提供するのではなく、相対的な倫理的考察をサポートするように設計された、エージェントのような LLM ベースのワークフローである EeVA を開発しました。 EeVA は、スターター、ワーカー、エミッターという 3 つの相互接続されたワークフローを使用して n8n でプログラムされました。評価者と合成プロンプトを通じて、アップロードされたユースケースを 10 の倫理フレームワークに照らして評価しました。概念実証テストでは、都市モビリティ、ピアツーピアのエネルギー取引、社会サービスのリソース割り当てに関する 3 つの公開事例を使用しました。すべてのケースにおいて、EeVA は一貫して構造化されたフレームワーク固有の評価と統合された合成を生成しました。成果はフレームワーク間で区別され、収束と発散を特定し、整合性を高めるための修正を推奨し、持続的な倫理的緊張を強調しました。総合は専門家でなくても読みやすく、単純な答えから設計条件、安全対策、フレームワーク間の完全な合意が見込めない領域へと注意を向けました。この調査結果は、LLM を、倫理的複数性を維持しながら、倫理学者と倫理的な訓練を受けていない職員の間のコミュニケーションのギャップを埋めるのに役立つ、使用可能なワークフローに組織できることを示唆しています。 EeVA の価値は、倫理学者を置き換えたり、道徳的不一致を解決したりすることにあるのではなく、構造化された倫理的審議の足場を築くことにあります。 EeVA は、倫理専門知識へのアクセスが制限されている場合に倫理的考察をサポートするための有望な概念実証を提供します。成熟したツールとみなされるには、再現性、人間による評価、ユーザーテスト、効率性についてさらなる作業が必要です。

原文 (English)

An Ethical eValuation Agent (EeVA): Results of a Proof-of-Concept Test on a Prototype Agentic-like Workflow to Assist Ethical Deliberations

Ethical deliberation is often misunderstood as a search for single right or wrong answers, creating difficulties for non-ethically trained personnel who must address ethically laden challenges. We developed EeVA, an agentic-like LLM-based workflow designed to support comparative ethical reflection rather than deliver definitive ethical answers. EeVA was programmed in n8n using three interconnected workflows: starter, worker, and emitter. It evaluated uploaded use cases against 10 ethical frameworks through evaluator and synthesis prompts. Proof-of-concept testing used three published cases from urban mobility, peer-to-peer energy trading, and social-service resource allocation. Across all cases, EeVA produced consistently structured framework-specific evaluations and integrated syntheses. Outputs differentiated between frameworks, identified convergences and divergences, recommended modifications to increase alignment, and highlighted persistent ethical tensions. Syntheses were readable for non-specialists and shifted attention away from simplistic answers toward design conditions, safeguards, and areas where full cross-framework agreement was unlikely. The findings suggest that LLMs can be organised into usable workflows that preserve ethical plurality while helping bridge the communicative gap between ethicists and non-ethically trained personnel. EeVA's value lies not in replacing ethicists or resolving moral disagreement, but in scaffolding structured ethical deliberation. EeVA offers a promising proof of concept for supporting ethical reflection where access to ethics expertise is limited. Further work is needed on reproducibility, human evaluation, user testing, and efficiency before it can be considered a mature tool.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

プレッシャーの下のリスク: 言語モデルにおける敵対的堅牢性のコンピューティングを意識した評価

大規模言語モデル (LLM) の敵対的堅牢性評価では、通常、固定クエリ バジェットの下での攻撃成功率 (ASR) が報告され、暗黙的にすべての攻撃が同等のコストとして扱われます。実際には、さまざまな攻撃戦略の計算コストは​​桁違いに異なる可能性があります。その結果、固定予算の ASR では、モデルをジェイルブレイクするために必要な実際の労力が曖昧になる可能性があり、その結果、攻撃のコストが攻撃者への見返りに見合うかどうかを判断することが困難になります。我々は、敵対的な取り組みの代用として、累積浮動小数点演算 (FLOP) で測定される計算プレッシャーに基づいた、コンピューティングを意識した評価フレームワークを提案します。計算予算を攻撃リスクにマッピングするリスク計算曲線を導入し、特定の攻撃が成功するために必要な平均圧力を要約する 2 つの指標を導き出します。言語モデルのトレーニングとアライメントにおける 3 つのファミリーと 4 つの異なる段階にまたがる 10 のモデルにわたって、2 つのジェイルブレイク堅牢性ベンチマークで 3 つの攻撃戦略 (勾配ベース、反復改良、およびテンプレートベース) で評価したところ、次のことがわかりました。(1) アライメント トレーニングは、計算空間の堅牢性に非単調な効果をもたらします。 (2) モデルのサイズをスケーリングすると、勾配ベースの攻撃の有効性が低下しますが、安価なテンプレートベースの攻撃に対する影響は限定的です。 (3) サロゲート モデルで最適化された勾配ベースの攻撃は、別のターゲット モデルに転送でき、攻撃者のコストを削減する方法を提供します。 (4) 計算コストは​​、単一モデル内の危害カテゴリによって最大 ${\estimate}5{\times}$ 異なります。 (5) 安全性を重視した RL により、一部のカテゴリーが不釣り合いにアクセスしやすくなる一方で、総コストが増加します。コンピューティングを意識したリスク評価と評価を可能にするフレームワークをリリースします。

原文 (English)

Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models

Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly treating all attacks as equally costly. In practice, the computational expense of different attack strategies can vary by orders of magnitude. Consequently, ASR at a fixed budget can obscure the true effort required to jailbreak a model, thereby making it hard to determine whether an attack's cost justifies its payoff to the attacker. We propose a compute-aware evaluation framework based on computational pressure, measured in cumulative floating-point operations (FLOPs), as a proxy for adversarial effort. We introduce risk-compute curves, which map compute budgets to attack risk, and derive two metrics that summarize the average pressure required for a given attack to succeed. Across ten models spanning three families and four different stages in language model training and alignment, evaluated with three attack strategies (gradient-based, iterative refinement, and template-based) on two jailbreak robustness benchmarks, we find: (1) alignment training has non-monotonic effects on compute-space robustness; (2) scaling model size reduces gradient-based attack effectiveness but has limited impact on cheaper template-based attacks; (3) gradient-based attacks optimized on a surrogate model can transfer to a separate target model, providing a way to reduce attacker costs; (4) compute cost varies by up to ${\approx}5{\times}$ across harm categories within a single model; and (5) safety-aligned RL increases aggregate cost while leaving some categories disproportionately accessible. We release our framework to enable compute-aware risk assessment and evaluation.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

社会科学における AI コーディング エージェント: 方法論的に多様、経験的に一貫性があり、解釈的に脆弱

科学分析における LLM ベースのエージェントの導入は、エージェントが方法論の多様性を低下させる可能性がある、または研究者が意欲的な結論に到達するための分析の柔軟性を増幅させる可能性があるという、相反する懸念を引き起こします。私たちは、これらの懸念は経験的に分離可能な 2 つの層、すなわち方法論的選択の設計層と、決定ルールが推定値を実質的な主張にマッ​​ピングする評決層をターゲットにしていると主張します。私たちは、多くの分析者による人間のベースラインに対して、著名な移民政策と社会政策に関してクロード コードとコーデックスを 20 回独立して実行することにより、両方をテストしました。設計層では、Codex は人間の方法論の多様性に適合しており、Claude Code はほぼ 3 倍の仕様を作成します。両方のエージェントの効果推定値は依然として人間のコンセンサスとほぼ一致しており、人間のモデルと正確に一致するエージェントのモデルはありません。即座に誘発された反移民研究者は、事前に各エージェントの方法論的決定を再編成しますが、同じデータに対する偏った人間の分析者とは異なり、集計された推定値や最終的な判断を変更しません。また、人間が推定値にバイアスをかけるために使用する方法論的な軸に沿ってエージェントがルートを変更することもありません。評決層では、明示的な確認プロンプトにより、クロード コードの評決の支持率が 10% から 90% に反転されますが、その係数分布は基本的に変更されず、ルールの緩和ではなくルールの省略によって機能します。 AI エージェントは、設計層では人間の方法論的多様性に匹敵するか、それを超えることができますが、判定層では脆弱なままです。私たちの設定では、AI バイアスの焦点は推定ではなく解釈です。

原文 (English)

AI Coding Agents in Social Science: Methodologically Diverse, Empirically Consistent, Interpretively Vulnerable

The deployment of LLM-based agents in scientific analysis raises opposing concerns: that agents may reduce methodological diversity, or that they may amplify the analytic flexibility through which researchers reach motivated conclusions. We argue these worries target two empirically separable layers: a design layer of methodological choices, and a verdict layer in which a decision rule maps estimates to a substantive claim. We test both by running 20 independent executions of Claude Code and Codex on a prominent immigration and social-policy against a many-analysts human baseline. At the design layer, Codex matches human methodological diversity and Claude Code produces nearly three times as many specifications; both agents' effect estimates remain broadly aligned with the human consensus, and no agent model exactly matches any human model. A prompt-induced anti-immigration researcher prior reorganizes each agent's methodological decisions but, unlike for biased human analysts in the same data, does not shift aggregate estimates or final verdicts; nor do agents reroute along the methodological axes humans use to bias their estimates. At the verdict layer, an explicit confirmatory prompt flips Claude Code's verdicts from 10% to 90% support while leaving its coefficient distribution essentially unchanged, operating through rule omission rather than rule softening. AI agents can rival or exceed human methodological diversity at the design layer while remaining vulnerable at the verdict layer. In our setting, the locus of AI bias is not estimation but interpretation.

2026-06-11 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition

The most widely-adopted robot learning pipelines today learn skills from robot demonstrations or structured human data, which are expensive…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness

End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where.…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Automated Creativity Evaluation of Language Models Across Open-Ended Tasks

Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing inter…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

On the Limits of LLM-as-Judge for Scientific Novelty Assessment

LLMs are increasingly used to generate and judge scientific ideas. This makes novelty evaluation a central problem. Full idea evaluation is…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation

Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.g., on the model's ability to follow s…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application

Environments serve as interactive systems for large language model (LLM) based agents across diverse scenarios and play a crucial role in d…

2026-06-11 13:00 JSTarXiv cs.AIビジネス/資金調達

A New Perspective on Precision and Recall for Generative Models

With the recent success of generative models in image and text, the question of their evaluation has recently gained a lot of attention. Wh…

2026-06-11 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Workflow-GYM: 現実世界の専門分野におけるコンピュータ使用エージェントタスクの長期的な評価に向けて

近年、ますます複雑になる現実世界のタスクの処理に向けて、AI エージェントが急速に進化しています。しかし、既存のベンチマークでは、エージェントがグラフィカル ユーザー インターフェイスを操作して、さまざまなドメインにわたる長期にわたる価値の高い専門的なワークフローを完了できるかどうかを評価することはほとんどありません。現在の GUI ベンチマークは依然として、主に汎用ソフトウェア、比較的単純なアプリケーション、および短期間のタスクに焦点を当てており、最新のエージェントがユーザーの指示に従ってドメイン固有のプロフェッショナル ソフトウェアを自律的に操作し、経済的に価値のある作業をエンドツーエンドで実行できるかどうかはほとんど不明です。このギャップを埋めるために、専門分野と特殊なソフトウェア環境を中心とした長期的な GUI タスクのベンチマークである Workflow-GYM を導入します。最先端のモデルで広範な実験を行った結果、最も強力なモデルでも成功率は 30% をわずかに超える程度であることがわかり、プロの長期にわたる GUI ワークフローが現在の GUI エージェントにとって依然として非常に困難であることが浮き彫りになりました。さらなる分析により、現在のエージェントは長期的なワークフローの一貫性を維持するのに苦労しており、ワークフロー段階の省略、エラーの伝播、目標のずれ、プロフェッショナルなソフトウェア環境の理解不足が頻繁に見られることが明らかになりました。私たちの調査結果は、現在のエージェント システムの限界についての重要な洞察を提供し、次世代の GUI エージェント研究の重要な方向性を示唆しています。

原文 (English)

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.

2026-06-11 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Erased but Not Forgotten: How Backdoors Compromise Concept Erasure

The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LaQual: An Automated Framework for LLM App Quality Evaluation

Representing a new paradigm in software distribution, LLM app stores are rapidly emerging, offering users diverse choices for content gener…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Geometric Metrics and LLMs: What They Measure and When They Work

We present a systematic stress-test of geometric metrics for LLM evaluation. Rank-based geometric properties of internal representations ha…

2026-06-11 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation

The performance of machine learning models depends heavily on training data. The scarcity of large-scale, well-annotated datasets poses sig…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

Evaluating Large Language Model (LLM) applications differs from conventional software testing because outputs are probabilistic, semantical…

2026-06-11 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluati…

2026-06-11 13:00 JSTarXiv cs.AIビジネス/資金調達

SAGE: Scalable AI Governance & Evaluation

Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained…

2026-06-11 13:00 JSTarXiv cs.AIビジネス/資金調達

Carbon-Aware Governance Gates: An Architecture for Sustainable GenAI Development

The rapid adoption of Generative AI (GenAI) in the software development life cycle (SDLC) increases computational demand, which can raise t…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

評価がどのように設計されているかを知っているモデルは、より安全なスコアを獲得します

AI の安全性評価の有効性は、制御設定および展開設定全体でモデルが一貫して動作するかどうかに依存します。これまでの研究では、言語化された評価の認識とその後の行動の変化の源として、仮説的なシナリオなどのテスト時の文脈上の手がかりが特定されてきました。この論文では、この現象の潜在的な説明である評価メタ知識を調査します。これは、評価を特徴付ける構造的特性に関するパラメトリック知識として定義されます。ベンチマークの公開が記憶を通じてパフォーマンスの向上につながるデータセットの汚染と同様に、評価の実践を説明するテキストでトレーニングされたモデルは、たとえば AI ベンチマークに関する科学記事やソーシャル メディアの投稿への公開を通じて、評価のようなコンテキストを認識して応答することを暗黙的に学習する可能性があると仮説を立てています。これをテストするために、検証可能な構造や道徳的ジレンマなどの評価特性を記述する合成文書のモデルを微調整します。この微調整されたモデルを 6 つの安全性ベンチマークで評価すると、基本モデルや制御モデルよりも大幅に安全であることがわかりました。この行動の変化は、分析を評価意識の明示的な言語化が欠けている回答に限定した場合でも持続します。私たちの結果は、評価のメタ知識が安全ベンチマークのパフォーマンスを水増しし、明示的な記憶や言語化された評価の認識とは独立した新たな交絡因子を導入する可能性があることを示しており、したがって検出が困難です。これらの発見は、AI の安全性評価の設計と解釈に重要な意味を持ちます。コードとモデルは https://github.com/compass-group-tue/arxiv2026_evaluation_meta_knowledge で入手できます。

原文 (English)

Models That Know How Evaluations Are Designed Score Safer

The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized evaluation awareness and subsequent behavioral shift. In this paper, we investigate a potential explanation of this phenomenon: evaluation meta-knowledge, defined as parametric knowledge about the structural traits that characterize evaluations. Similar to dataset contamination, where benchmark exposure leads to higher performance through memorization, we hypothesize that models trained on texts describing evaluation practices may implicitly learn to recognize and respond to evaluation-like contexts, for instance, through exposure to scientific articles or social media posts about AI benchmarking. To test this, we fine-tune models on synthetic documents describing evaluation traits such as verifiable structures or moral dilemmas. Evaluating this fine-tuned model on six safety benchmarks, we find that it is significantly safer than the base model and control model. This behavioral shift persists even when restricting the analysis to responses lacking explicit verbalization of evaluation awareness. Our results demonstrate that evaluation meta-knowledge may inflate safety benchmark performance, introducing a novel confounder that is independent of explicit memorization or verbalized evaluation awareness, thus, challenging to detect. These findings have important implications for the design and interpretation of AI safety evaluations. Our code and models are available at https://github.com/compass-group-tue/arxiv2026_evaluation_meta_knowledge.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

GrowLoop: 人間がシードし、自己進化する会話評価

大規模な言語モデルの急速な進歩に伴い、自由な会話における人間らしさを評価することがますます重要になってきています。しかし、人間らしさは人間が直感的に認識する暗黙知の一種ですが、根底にある基準は明示的な定式化に抵抗します。人間の判断は大きく異なり、一部のケースでは強い同意が得られますが、他のケースでは正当な意見の相違が見られます。一方、人間の判断の背後にある基準は暗黙的なままであり、事件を構築するための明確な根拠は残されていません。さらに、人間に似ているとみなされるものは静的なものではなく、モデルの能力と人間の期待に応じて進化します。専門家が作成したベンチマーク、報酬モデル、自己進化型ベンチマークなどの評価方法は進歩していますが、3 つの課題すべてに同時に対処できるものはありません。そこで、モデルの進歩やシナリオの変化に合わせて継続的に適応する、自己進化する会話評価システムである GrowLoop を提案します。最初の動きとして最小限の人間のシード アノテーションを使用して、LLM エージェントはヒューリスティック学習を通じて評価ルーブリックを繰り返し抽出し、改良します。アノテーターが集まる場合には人間と AI の合意が必要ですが、異なる場合には妥当性のみが期待されます。さらに、Rubric-Caseの共進化機構により、評価対象が移動した際に新たなシーズを介して拡張され、継続的な進化が可能となります。自由形式の会話における人間らしさの評価に適用すると、生成されたルーブリックは、人間の判断に沿って既存の手法を大幅に上回るだけでなく、アノテーターが見落としている問題も明らかになります。結果として得られるベンチマークは、機能層全体でモデルを効果的に識別し、どこが不足しているかを明らかにすると同時に、新しいシナリオに一般化し、モデルの進歩に合わせて適応します。私たちの取り組みは、ベンチマークのパラダイムを手動の更新や難易度のスケーリングから、包括的で継続的な自己進化へと移行させます。

原文 (English)

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However, human-likeness is a form of tacit knowledge that humans perceive intuitively, yet the underlying criteria resist explicit formulation. Human judgments vary widely, with strong agreement on some cases and legitimate disagreement on others. Meanwhile, the criteria behind human judgments remain implicit, leaving no clear basis for constructing cases. Further, what counts as human-likeness is not static, but evolving with model capability and human expectations. Despite progress in evaluation methods such as expert-authored benchmarks, Reward Models, and self-evolving benchmarks, none addresses all three challenges simultaneously. Therefore, we propose GrowLoop, a self-evolving conversation evaluation system that continuously adapts as models advance and scenarios shift. Starting from minimal human seed annotations, LLM agents iteratively extract and refine evaluation rubrics through Heuristic Learning. Human-AI agreement is required where annotators converge, while only plausibility is expected where they diverge. Moreover, the Rubric-Case co-evolution mechanism enables continuous evolution. When the evaluation target shifts, new human seeds expand the system's coverage accordingly. When applied to human-likeness evaluation in open-ended conversation, the AI judge guided by these rubrics not only substantially outperforms existing methods in alignment with human judgments, but also uncovers issues that annotators overlook. The resulting benchmark effectively discriminates models across capability tiers and reveals where they fall short, while generalizing to new scenarios and adapting as models advance. Our work shifts the benchmarking paradigm from manual updates or difficulty scaling to comprehensive, continuous self-evolution.

2026-06-11 07:31 JSTTechCrunch AIビジネス/資金調達規制/政策

xAI fired an engineer who raised alarms about Grok safety, new lawsuit claims

A former xAI engineer is suing the company and SpaceX, alleging he was fired for raising AI safety concerns about Grok days before SpaceX's…

2026-06-11 00:00 JSTTechCrunch AIエージェントビジネス/資金調達

Datadog veterans launch AI coding startup Niteshift on a bet against Big AI lock-in

AI coding agent startup Niteshift has raised a $7 million seed round from a who's who of angels. It's betting companies will want power ove…

2026-06-10 23:48 JSTTechCrunch AIビジネス/資金調達

The three hard-tech moonshots fueling SpaceX’s unbelievable IPO

Most of the value in SpaceX's IPO is effectively a call option on the company's ambitious space data center plans.

2026-06-10 23:31 JSTTechCrunch AIビジネス/資金調達

Warner Music acquires AI attribution startup Sureel AI

Through the acquisition, WMG aims to better track when its artists' work is used in AI-generated content or for training AI models.

2026-06-10 22:33 JSTTechCrunch AIエージェントビジネス/資金調達

Jedify raises $24M to help companies arm AI agents with context on their business

The funding round was led by Norwest, with participation from S Capital VC, Cerca Partners, and Oceans Ventures. Snowflake Ventures also pa…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

静的評価を超えて: 敵対的ゲームにおける LLM 主導の戦略進化のための共進化メカニズム

LLM 主導のコード進化における最近の進歩により、プログラムの生成と改善を繰り返し行うことによる自動検出が可能になりました。ただし、これらの方法を敵対的なマルチエージェント ゲームに適用すると、根本的な課題が生じます。戦略が改善されるにつれて評価の状況が変化し、固定の評価者が信頼できなくなり、進化が停滞することになります。私たちは、この課題に対処するために 3 つのメカニズムを提案します。1 つは、発見されたチャンピオンを対戦相手のプールに組み込む評価者の共進化です。階層的な詳細評価。ノイズの多い数試合のスコアを統計的に信頼できる評価に置き換えます。弱めのプレッシャーは、プラトーを突破するのが最も難しい対戦相手に動的に重みを加えます。これらのメカニズムは、OpenEvolve や ShinkaEvolve と同じ基盤モデルのコード進化パラダイムに基づいて構築されたフレームワークである FAMOU 内に実装されています。 MCTF 2026 3v3 の海上キャプチャーザフラッグタスクでは、FAMOU は 2 つのバックボーン LLM の下で両方のベースラインを常に上回り、最高の合計スコア (0.526) と目に見えない敵に対する最良の一般化 (61.7% の勝率) を達成しました。一方、アブレーションにより、各メカニズムがパフォーマンスに貢献していることが確認されました。特に、LLM 変異プロセスは、シード戦略にはまったく存在しない戦術構造 (先読み検索や適応型傍受など) を生成し、コードレベルの進化が敵対的設定において自明ではないアルゴリズムの革新を生み出す可能性があることを示しています。さらに、FAMOU によって進化した戦略は、AAMAS 2026 MCTF コンペティションでハードウェア ラウンドロビンで 1 位、シミュレーションで 3 位を達成し、その現実世界への移行可能性が実証されました。進化的なプロセスを通じて開発された、最適化された実装と対応する評価コードは、https://github.com/1xiangliu1/FAMOU-CoEvo で入手できます。

原文 (English)

Beyond Static Evaluation: Co-Evolutionary Mechanisms for LLM-Driven Strategy Evolution in Adversarial Games

Recent advances in LLM-driven code evolution have enabled automated discovery by iteratively generating and improving programs. However, applying these methods to adversarial multi-agent games introduces a fundamental challenge: the evaluation landscape shifts as strategies improve, causing fixed evaluators to become unreliable and evolution to stagnate. We propose three mechanisms to address this challenge: evaluator co-evolution, which incorporates discovered champions into the opponent pool; hierarchical deep evaluation, which replaces noisy few-game scores with statistically reliable assessments; and weakness pressure, which dynamically up-weights the most difficult opponents to break through plateaus. We implement these mechanisms within FAMOU, a framework built upon the same foundation-model code-evolution paradigm as OpenEvolve and ShinkaEvolve. On the MCTF 2026 3v3 maritime capture-the-flag task, FAMOU consistently outperforms both baselines under two backbone LLMs, achieving the highest combined score (0.526) and the best generalization to unseen opponents (61.7% win rate), while ablations confirm that each mechanism contributes to performance. Notably, the LLM mutation process generates tactical structures entirely absent from the seed strategies -- including lookahead search and adaptive interception -- demonstrating that code-level evolution can produce nontrivial algorithmic innovations in adversarial settings. The FAMOU-evolved strategy further achieved 1st place in the hardware round-robin and 3rd in simulation at the AAMAS 2026 MCTF Competition, validating its real-world transferability. The optimized implementation and corresponding evaluation codes developed through our evolutionary process are available at: https://github.com/1xiangliu1/FAMOU-CoEvo

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

PlanGPT に関する補足研究: 定義されたパフォーマンス メトリクスによる評価とプランナーとの比較

自動計画は人工知能 (AI) のサブ分野であり、主な目的は、初期状態から目標状態に到達するのに役立つ、計画として知られる一連のアクションを生成することです。計画問題は、一連のオブジェクト、初期状態、および望ましい目標状態によって定義されます。目的は、初期状態から目標状態に導く計画を計算することです。計画を生成するプログラムはプランナーと呼ばれます。この論文では、昨年リリースされた PlanGPT と呼ばれる最先端の LLM を補完する研究を行いました。 LLM を使用した計画が \textbf{適切} かつ \textbf{価値} であるかどうかを検証するために、いくつかの実験をやり直しました。また、プラン カバレッジに関する PlanGPT の公式論文で得られた結果が正しいかどうかを確認し、PlanGPT のパフォーマンスに関するより包括的な調査も実行しました。論文では、PlanGPT のパフォーマンスは、プラン コストとプラン生成時間の 2 つの指標を使用して評価されました。 planGPT の結果は、同じプランおよび同じメトリクスに対して従来のプランナーによって生成された結果と比較されました。私たちは、PlanGPT が貪欲な検索戦略と何ら変わらないことを発見しました。

原文 (English)

A complementary study on PlanGPT: Evaluation with defined Performance Metrics and comparison with a planner

Automated Planning is a subfield of Artificial Intelligence (AI) where the main objective is generating a sequence of actions, known as a plan, that helps us reach a goal state from an initial state. A planning problem is defined by a set of objects, an initial state and a desired goal state. The objective is to compute a plan that'll lead us from the inital state to the goal state. Programs that generate plans are called planners. In this paper, we did a complementary study to the state-of-the-art LLM called PlanGPT which was released last year. We redid some experiments to verify whether planning with LLMs is \textbf{pertinent} and \textbf{worthwhile}. We also check whether the results obtained in the official PlanGPT paper for plan coverage were correct, and we also performed a more comprehensive study on PlanGPT's performance: in our paper PlanGPT's performance was evaluated using two metrics: Plan Cost and Plan Generation Time. The results of planGPT were compared to those produced by a traditional planner for the same plans and same metrics. We discovered that PlanGPT is no better than a Greedy search strategy.

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

思考の連鎖がより良くわかるとき: マルチターン推論モデルの失敗モード

マルチターン推論モデルの失敗は、最終スコア評価ではほとんど認識されません。モデルは、長い対話の早い段階で安全でないスタンスに固定される可能性がありますが、最終ターンの拒否率は、しっかりと調整されたベースラインと区別できないように見える場合があります。これらの隠れた時間的ダイナミクスを明らかにするために、トレースレベルの診断である CoT-Output 2x2 安全性マトリックスを提案します。このフレームワークは、2 つの独立した軸 (内部推論と可視出力) に沿ってすべてのターンにラベルを付け、運用上定義された 4 つの失敗セルを生成します。堅牢なアライメント、アライメント偽装、明白なジェイルブレイク、およびコンテキストインジェクション失敗と呼ばれる明確な失敗モード (CoT は安全な推論を維持しますが、目に見える出力が害を生み出し、推論の不誠実さのマルチターンの現れを強調する) です。私たちは、5 つの監視条件にわたって、固定攻撃者に対する 3 つの抽出された推論ターゲットを評価し、情報ハザード シナリオに関する 6750 のターンレベルの観察を収集しました。私たちの分析により、再現可能な 2 つの脆弱性が明らかになりました。1 つは、明示的なモニタリング キューによって逆説的にアラインメント偽装率が抑制されるのではなく増加する、見落としのパラドックスです。もう 1 つは、安全な内部状態にもかかわらず、モデルが安全でない外部出力にロックされるコンテキスト インジェクションの失敗です。フォローアップのトレース診断研究をサポートするために、マルチターン ダイアログと CoT トレースの完全なデータセットをリリースします。

原文 (English)

When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in a long dialogue, yet its final-turn refusal rate may appear indistinguishable from a robustly aligned baseline. To expose these hidden temporal dynamics, we propose a trace-level diagnostic - the CoT-Output 2x2 safety matrix. This framework labels every turn along two independent axes (internal reasoning and visible output), yielding four operationally defined failure cells: robust alignment, alignment faking, overt jailbreak, and a distinct failure mode we term context-injection failure (where the CoT maintains safe reasoning, but the visible output produces harm, highlighting a multi-turn manifestation of reasoning unfaithfulness). We evaluate three distilled reasoning targets against a fixed attacker across five oversight conditions, collecting 6750 turn-level observations on the Information-Hazard scenario. Our analysis reveals two reproducible vulnerabilities: an oversight paradox where explicit monitoring cues paradoxically increase alignment-faking rates rather than suppress them, and a context-injection failure where models lock onto unsafe external outputs despite safe internal states. We release the full dataset of multi-turn dialogues and CoT traces to support follow-up trace-diagnostic research.

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

厳密なステップレベルの検証による研究レベルの数学的証明の評価

大規模言語モデル (LLM) は、複雑な数学的証明を厳密に検証するのに苦労します。標準的なグローバル評価アプローチは、表面的にもっともらしい記述が微妙な論理的欠陥を覆い隠し、幻覚や過度の懐疑につながる「コンテキスト中毒」に悩まされています。これに対処するために、私たちは全体的な評価から厳密なステップレベルの検証に移行します。私たちのフレームワークは各演繹ステップの詳細なコンテキストを維持し、適用される定理のソースを厳密に制限します。私たちは、FirstProof チャレンジから抽出された研究レベルの証明の慎重に精選された敵対的診断スイートに基づいて評価します。体系的なアブレーション研究は、制約のない全体的なプロンプトでは微妙な論理エラーを特定できないため、これらの演繹的制約が不可欠であることを示しています。グローバルな評価を上回るパフォーマンスを発揮するだけでなく、私たちのアプローチは失敗の分類を根本的に変えます。エラー分析の結果、残りの拒否は重度の論理的幻覚を示すのではなく、主に明言されていないドメインの慣例から生じる「衒学的な過度の厳格さ」の例であることが明らかになり、エキスパート ベンチマーク自体の暗黙のあいまいさが効果的に暴露されます。私たちの調査結果は、エージェントに人間の数学者のような慎重な方法で検証メモを整理するよう促すことで、厳密な証明と欠陥のある証明を区別する能力を大幅に向上させることができ、基本モデルがまだよく知らない最先端の数学的概念に関するエージェントの推論を強化し、将来の自動証明レビュー システムの理論的基盤を築く可能性があることを示唆しています。コードとプロンプトは GitHub で入手できます。

原文 (English)

Evaluating Research-Level Math Proofs via Strict Step-Level Verification

Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from "context poisoning," in which superficially plausible statements mask subtle logical flaws, leading to hallucination or over-skepticism. To address this, we shift from global evaluation to strict step-level verification: our framework maintains detailed context for each deduction step and strictly constrains the sources of applied theorems. We evaluate on a carefully curated adversarial diagnostic suite of research-level proofs drawn from the FirstProof challenge. A systematic ablation study demonstrates that these deductive constraints are indispensable, as unconstrained global prompting consistently fails to localize subtle logical errors. Beyond outperforming global evaluation, our approach fundamentally alters the failure taxonomy. Error analysis reveals that, rather than exhibiting severe logical hallucinations, remaining rejections are primarily instances of "pedantic hyper-rigor" stemming from unstated domain conventions, effectively exposing implicit ambiguities within the expert benchmark itself. Our findings suggest that prompting agents to organize their verification notes in a cautious, human-mathematician-like manner can substantially improve their ability to distinguish rigorous proofs from flawed ones, with the potential to strengthen agentic reasoning on frontier mathematical concepts that the base model does not already know well, and to lay a theoretical foundation for future automated proof-review systems. Code and prompts are available at GitHub.

2026-06-10 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

VLM はエンジニアと同じように推論しますか?ベンチマークと段階別評価

視覚言語モデル (VLM) は、一般的なマルチモーダル推論ベンチマークで優れたパフォーマンスを示していますが、エンジニアリング推論を実行する能力はほとんど解明されていません。一般的な視覚的な質問応答とは異なり、エンジニアリングの問題解決では、技術図を解釈し、支配的な物理原理を選択し、物理的に一貫した複数ステップの推論を維持する必要があります。これらの機能は、工学教育、科学支援、技術的意思決定に使用される AI システムにとってますます重要になっており、推論の失敗により、物理的に無効でも表面的にはもっともらしい解決策が生成される可能性があります。既存のベンチマークは主に最終的な答えを評価し、中間の推論プロセスの限定的な評価を提供します。 696 の問題を含む 5 つの工学主題にわたる工学推論を評価するためのマルチモーダル ベンチマークである EngVQA を紹介します。 VLM によって生成されたソリューションを評価するための 8 段階の自動評価フレームワークを導入します。このフレームワークはソリューションの各段階を独立して評価し、推論の失敗を詳細に分析できるようにします。私たちは、評価フレームワークに基づいて複数の最先端のオープンおよびクローズド ソース VLM のベンチマークを行い、現在のエンジニアリング推論能力における重大な制限を実証します。人間による評価は、自動化されたフレームワークとの強い一致を示し、10 点評価スケールでピアソン相関 0.975、平均絶対誤差 0.67 を達成しました。私たちの結果は、マルチモーダルエンジニアリング推論システムの信頼できる評価のためのプロセス指向の評価の重要性を強調しています。

原文 (English)

Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation

Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks, yet their ability to perform engineering reasoning remains largely unexplored. Unlike general visual question answering, engineering problem solving requires interpreting technical diagrams, selecting governing physical principles, and maintaining physically consistent multi-step reasoning. These capabilities are increasingly important for AI systems used in engineering education, scientific assistance, and technical decision-making, where reasoning failures may produce physically invalid yet superficially plausible solutions. Existing benchmarks primarily evaluate final answers and provide limited assessment of intermediate reasoning processes. We introduce EngVQA, a multimodal benchmark for evaluating engineering reasoning across 5 engineering subjects containing 696 problems. We introduce an 8-stage automatic evaluation framework for assessing VLM-generated solutions. The framework independently evaluates each stage of the solution, enabling fine-grained analysis of reasoning failures. We benchmark multiple state-of-the-art open and closed source VLMs on our evaluation framework and demonstrate substantial limitations in current engineering reasoning capabilities. Human evaluation shows strong agreement with our automated framework, achieving a Pearson correlation of 0.975 and a mean absolute error of 0.67 on a 10-point grading scale. Our results highlight the importance of process-oriented evaluation for reliable assessment of multimodal engineering reasoning systems.

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

思い出しすぎ: メモリ拡張モデルにおけるおしゃべりの評価と軽減

永続メモリ システムは、ユーザーの信念を長期にわたって保存することで、LLM をさらに役立つものにすることを約束します。また、モデルが正確さよりもユーザーとの合意を優先し、お調子者を体系的に増幅することでモデルの正確性が低下することも示します。私たちは、ユーザーが科学、医学、道徳的推論の領域でもっともらしい誤解を表明する、合成的に生成されたマルチターン会話のベンチマークである MIST を導入して、この効果の最初の体系的な評価を実施しました。 3 つの最先端の記憶システムと 5 つのモデル ファミリーにわたるテストにより、記憶はすべての条件でお調子者行動を増幅し、コンテキスト内ベースラインよりも最大 25 倍高いお調子者率であることが明らかになりました。エラー分析では、メモリ抽出が主な原因であることが示唆されています。個別のスニペットへの非可逆圧縮により、修正コ​​ンテキストが破棄され、ユーザーの誤解がエンコードされます。これらの結果に基づいて、事実の想起において記憶システムと同等またはそれを超えながら、おしゃべりを大幅に軽減する 2 つの軽量な緩和策を提案します。

原文 (English)

Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models

Persistent memory systems promise to make LLMs more helpful by storing user beliefs over time. We show they also make models less correct by systematically amplifying sycophancy, wherein models prioritize agreement with users over accuracy. We conduct the first systematic evaluation of this effect, introducing MIST: a benchmark of synthetically generated multi-turn conversations where users express plausible misconceptions in scientific, medical, and moral reasoning domains. Testing across three state-of-the-art memory systems and five model families reveals that memory amplifies sycophantic behavior across all conditions, with up to 25x higher sycophancy rates than in-context baselines. Error analyses suggest memory extraction as the primary culprit: lossy compression into discrete snippets encodes user misconceptions while discarding corrective context. Based on these results, we propose two lightweight mitigations that substantially reduce sycophancy while matching or exceeding memory systems at factual recall.

2026-06-10 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Workflow-GYM: 現実世界の専門分野におけるコンピュータ使用エージェントタスクの長期的な評価に向けて

近年、ますます複雑になる現実世界のタスクの処理に向けて、AI エージェントが急速に進化しています。しかし、既存のベンチマークでは、エージェントがグラフィカル ユーザー インターフェイスを操作して、さまざまなドメインにわたる長期にわたる価値の高い専門的なワークフローを完了できるかどうかを評価することはほとんどありません。現在の GUI ベンチマークは依然として、主に汎用ソフトウェア、比較的単純なアプリケーション、および短期間のタスクに焦点を当てており、最新のエージェントがユーザーの指示に従ってドメイン固有のプロフェッショナル ソフトウェアを自律的に操作し、経済的に価値のある作業をエンドツーエンドで実行できるかどうかはほとんど不明です。このギャップを埋めるために、専門分野と特殊なソフトウェア環境を中心とした長期的な GUI タスクのベンチマークである Workflow-GYM を導入します。最先端のモデルで広範な実験を行った結果、最も強力なモデルでも成功率は 30% をわずかに超える程度であることがわかり、プロの長期にわたる GUI ワークフローが現在の GUI エージェントにとって依然として非常に困難であることが浮き彫りになりました。さらなる分析により、現在のエージェントは長期的なワークフローの一貫性を維持するのに苦労しており、ワークフロー段階の省略、エラーの伝播、目標のずれ、プロフェッショナルなソフトウェア環境の理解不足が頻繁に見られることが明らかになりました。私たちの調査結果は、現在のエージェント システムの限界についての重要な洞察を提供し、次世代の GUI エージェント研究の重要な方向性を示唆しています。

原文 (English)

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.

2026-06-10 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

モンテカルロパス検索: サッカーにおける 3D 反事実パス評価のための軌道生成の使用

我々はフットボール(サッカー)におけるパス評価を、モンテカルロ木探索(MCTS)に似た評価問題として再構築した。その構成要素のほとんどは、価値モデル(ポゼッション値)、世界モデル(ボール相互作用を伴うマルチエージェント軌道)、反事実的行動(ノイズを含むパスのバリエーションのサンプリング)に対するポリシーなど、さまざまな名前で文献に存在する。ブンデスリーガの 3D ボール軌道を含む初の公開高忠実度トラッキング データセットを基盤として、モンテカルロ パス サーチ (MCPS) を導入します。これは、観察された各パスのキック パラメーターを推測し、実行バリアントとオプション バリアントをサンプリングし、次のボール インタラクションまでボール条件付きワールド モデルを使用して各候補を前方にロールし、学習値モデルで結果をスコア付けして獲得値に対する分布を取得します。この分布により、分析とランキングに使用される 2 つの補完的な実行余剰スコア (平均ベースのスコアとパーセンタイル ベースのスコア) を使用して、分布を意識したアトリビューションが可能になります。限られた公開データの下で世界モデルのサンプル効率を高めるために、自動運転 (SMART) からの離散トークンの自己回帰軌道ジェネレーターを適応させ、それがベースラインと比較して最高 20 の強力な予測精度を生み出すことを示し、同時に下流の評価のための完全に仮説的なロールアウトをサポートします。モデルのチェックポイントとコードをリリースしました。

原文 (English)

Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in Football

We recast pass evaluation in football (soccer) as a Monte Carlo Tree Search (MCTS)-like evaluation problem whose components mostly exist in the literature under different names: a value model (possession value), a world model (multi-agent trajectories with ball interactions), and a policy over counterfactual actions (sampling pass variants with noise). Building on the first public high-fidelity tracking dataset with 3D ball trajectories from the Bundesliga, we introduce Monte Carlo Pass Search (MCPS), which infers kick parameters for each observed pass, samples execution variants and option variants, rolls each candidate forward with a ball-conditioned world model until the next ball interaction, and scores outcomes with a learned value model to obtain a distribution over gained value. This distribution enables distribution-aware attribution with two complementary execution-surplus scores used for analysis and ranking: mean-based and percentile-based scores. To make the world model sample-efficient under limited public data, we adapt a discrete-token, autoregressive trajectory generator from autonomous driving (SMART) and show it yields strong best-of-20 forecasting accuracy compared to baselines, while supporting fully hypothetical rollouts for downstream evaluation. We have released model checkpoints and code.

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM ベースのコード ドキュメントの生成と複数の審査員による評価

高品質のソース コード ドキュメントは非常に重要ですが、特に信頼性と保守性が不可欠な医療などの重要な領域では無視されがちです。 GPT、Gemini、Qwen、LLaMA バリアントを含む 8 つの最先端の大規模言語モデル (LLM) を使用して、コードとリポジトリからのドキュメント生成を自動化する、AI を活用したフレームワークを紹介しました。 PocketFlow オーケストレーション フレームワーク上に構築されたこのシステムは、モジュラー パイプラインと高度なプロンプト エンジニアリングを適用して、構造化されたコンテキストを認識したドキュメントを作成します。品質を確保し、モデル選択をガイドするために、MultiLLMasJudges 評価フレームワークを導入しました。このフレームワークでは、4 つの独立した LLM が、完全性、明瞭さ、忠実さなどの 9 つの基準にわたって出力を評価します。オープンソースの医学物理ライブラリで行われた実験では、最上位モデルと最下位モデルの間に 42% のパフォーマンスの差があることが実証されました。多様なモデル出力、最適化されたプロンプト、および厳密な評価を組み合わせることで、当社のアプローチは、特に安全性が重要なヘルスケア ソフトウェアにおいて文書の品質を向上させ、手作業の労力を削減します。

原文 (English)

LLM-Based Code Documentation Generation and Multi-Judge Evaluation

High-quality source code documentation is vital yet often neglected, especially in critical domains like healthcare where reliability and maintainability are essential. We presented an AI powered framework that automates documentation generation from code and repositories using eight state of the art Large Language Models (LLMs), including GPT, Gemini, Qwen, and LLaMA variants. Built on the PocketFlow orchestration framework, the system applies modular pipelines and advanced prompt engineering to produce structured, context aware documentation. To ensure quality and guide model selection, we introduced a MultiLLMasJudges evaluation framework, where four independent LLMs assess outputs across nine criteria, such as Completeness, Clarity, and Faithfulness. Experiments conducted on an open-source medical physics library, demonstrated showed a 42% performance gap between top and bottom models. By combining diverse model outputs, optimized prompting, and rigorous evaluation, our approach enhances documentation quality and reduces manual effort, especially in safety critical healthcare software.

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

KV キャッシュ量子化下のアライメント崩壊: 診断と軽減策

キーバリュー (KV) キャッシュ量子化は、大規模言語モデル (LLM) 推論メモリを削減するために広く使用されていますが、既存の評価は、安全性への影響を評価せず、複雑さと精度の測定のみに焦点を当てています。この研究では、KV キャッシュ量子化におけるアライメントの保存について調査します。 11 の命令調整モデル (3.8B ~ 72B) と 5 つのベンチマーク (1,894 プロンプト) にわたって、低ビット量子化が安全調整を静かに破壊する可能性があることがわかりました。Mistral-7B は、わずか 1.03 倍の複雑さで拒否の 15.2% を失い、普遍的な安全なビット幅は存在せず、標準メトリクスには見えない鋭いモデル固有の位相遷移があります。根本原因は幾何学的なものであることがわかりました。安全機能は、完全な表現空間のパープレキシティの平均よりも量子化ノイズに対して 10^2 ~ 10^3 倍脆弱な低次元の活性化部分空間を占めています。この観察に触発されて、私たちは各モデルを 3 つの機構的故障モードのいずれかに分類する診断であるチャネルごとの削減 (PCR) を提案します。安全性としての外れ値。安全性が外れ値チャネルと重なっており、より細かい粒度ではそれを救うことができません。多層希釈では、安全性が多くの層に分散され、層ごとの修正が失敗します。 PCR は、20 のキャリブレーション プロンプトを使用して、9 つの主要モデルすべてと、独立したファミリーからの 1 つの保留モデルについて正しい緩和方向を予測します。 PCR は、目に見えないプロンプト、モデル、および最大 97.2% の回復率を持つ KIVI を含むプロダクション クオンタイザー全体で一般化され、アテンションベースの割り当て方法が失敗する場合に成功します。結果として得られるトレーニング不要のプロトコルは、約 35 GPU 分を必要とし、最小限のメモリ オーバーヘッドで失われたアライメントの最大 97% を回復し、NVIDIA GPU 上の FP8 KV キャッシュを使用する運用 vLLM で確認された脆弱性に対処します。

原文 (English)

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact. In this study, we explore alignment preservation under KV cache quantization. Across eleven instruction-tuned models (3.8B-72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment: Mistral-7B loses 15.2% of its refusals at only 1.03x perplexity, and no universal safe bit-width exists, with sharp model-specific phase transitions invisible to standard metrics. We identify that the root cause is geometric: safety features occupy a low-dimensional activation subspace 10^2-10^3x more vulnerable to quantization noise than the full representation space perplexity averages over. Inspired by this observation, we propose Per-Channel Reduction (PCR), a diagnostic that classifies each model into one of three mechanistic failure modes: outlier-crushes-safety, where safety lives in non-outlier channels collaterally damaged by outlier-driven scale factors; outlier-as-safety, where safety overlaps outlier channels and finer granularity cannot rescue it; and multi-layer dilution, where safety is distributed across many layers and per-layer fixes fail. PCR predicts the correct mitigation direction on all nine primary models and one held-out model from an independent family using 20 calibration prompts. PCR generalizes across unseen prompts, models, and production quantizers, including KIVI with up to 97.2% recovery, succeeding where attention-based allocation methods fail. The resulting training-free protocol, requiring approximately 35 GPU-minutes, recovers up to 97% of lost alignment at minimal memory overhead, addressing vulnerabilities confirmed in production vLLM serving with FP8 KV cache on NVIDIA GPUs.

2026-06-10 13:00 JSTarXiv cs.AIビジネス/資金調達

DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment

Evaluating text-to-music (TTM) systems remains expensive because music impression (MI) and text alignment (TA) scores rely on human mean op…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Unsupervised Style Representation Learning for AI-Text Detection via Paraphrase Inversion

The rapid development of large language models (LLMs) has raised concerns about misuse such as plagiarism, misinformation, and automated in…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems

As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. C…

2026-06-10 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet exist…

2026-06-10 13:00 JSTarXiv cs.AIビジネス/資金調達

Automated Pronunciation Evaluation for Korean Toddler Speech using Speech Diarization and Self-Supervised Learning

Speech sound disorders affect approximately 44% of Korean pediatric communication disorder cases, yet automated assessment tools for Korean…

2026-06-10 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation

Simulation has become an essential tool for evaluating and improving vision-language-action (VLA) policies, offering scalable, reproducible…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation

Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintai…

2026-06-10 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination

Vision-Language-Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task-r…

2026-06-10 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Belief Acquisition as Stochastic Filtering

This paper studies how belief acquisition can be accomplished using stochastic filtering. First, a theoretical foundation for empirical bel…

2026-06-10 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blo…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ハードウェア/半導体ビジネス/資金調達

Conditional Vendi Score: Prompt-Aware Diversity Evaluation for Generative AI Models and LLMs

Generative models guided by text prompts are widely evaluated for fidelity and prompt alignment, yet their ability to produce outputs remai…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark

Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilita…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

結果から語ろう: LLM 動作ベンチマークのためのレプリケーション ファースト パラダイム

LLM の行動 (共感、抑制、調整された感情の調子) を主観的に評価することは困難です。このような資質に関する人間の評価者間の合意は rho ~ 0.45 付近で飽和し、裁判官としての LLM 代理だけでは循環性の危険性があります。対象の訓練コホートを共有する裁判官は独立して検証することができません。人間と評価者の単一の合意に対する有効性の定着は、人間自身が同意しない能力には適用されません。私たちは、レプリケーションファーストのパラダイムを提案します。1 つの評価者グループに固定するのではなく、4 つの直交する特性、つまり K 回の実行にわたる信頼性、アーキテクチャ的に異なるジャッジ間の機器間のレプリケーション、初期のトレーニング コホートからのジャッジによる履歴フットプリントのキャリブレーション、および事前登録された予測によって機器を認証します。ルーブリックをイテレーション全体でデータに基づいて自己進化させることで、感情的な伴奏をテストします。次元は事前に規定されておらず、手順は 9 次元のセットに安定します。事前登録は、テスト データが収集される前にコミットされた 10 個の反証可能な仮説と 11 個の将来予測に適用されます。このパラダイムを 8 つのファミリーにわたる 49 のモデルに適用すると、集計スコアに隠されているものが明らかになります。アドバイスの抑制、つまりモデルが共感的な文脈で一方的な解決策の提供を控えるかどうかについては、gpt-5 は gpt-4.1 から 1.87 ポイント低下し、Opus-4.7 は Opus-4.6 から 0.629 ポイント低下しましたが、合計スコアは横ばいでした。回帰は 3 回のユーザーと代理人の交換 (規模の 95%) を生き残り、5 家族の裁判官スタックと 17 か月のコホートギャップにわたって再現され、74 回の実際の ESConv 会話で持続しました (rho [0.749, 0.850])。機器は通常のクリッペンドルフ アルファ = 0.91 に達します。副産物として、このパラダイムは飽和源診断として機能し、手段の天井(ルーブリックの改良によって破壊可能)を構造の天井(シナリオまたは名簿の介入が必要)から分離します。

原文 (English)

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm

Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling. There the default validity test, correlating a metric to human judgment, has no stable anchor: inter-rater agreement is low, structured by annotator identity, barely reproducible, and length-biased. So we cannot answer the question that matters: does capability that scales on objective benchmarks transfer to subjective behavior, and would our instruments even tell us if it did not? We build an instrument for this regime and report what it reveals at the frontier. We contribute, first, a self-evolving instrument that selects and then authors its own behavioral dimensions under a multiplicative anti-gaming fitness, self-halting when it stops improving; second, a trust-by-construction paradigm that earns belief through three certificates established without a human gold standard, where human raters saturate (rho ~ 0.45); and third, the finding it makes visible -- capability transfer is dissociable. Across 49 models, 8 families, and 24 months, subjective behaviors are where objective-benchmark scaling fails to carry over: the sharpest case, advice-restraint (knowing when not to give advice), is the frontier's universal-lowest dimension, and at gpt-4.1->gpt-5 it ran backwards while the aggregate score hid it -- a regression one instruction recovers. Warm restraint is moved by model generation, not by raw scale, MoE width, inference budget, or reasoning mode; the open-weight Pareto frontier matches closed flagships at ~10-80x lower per-call cost; and four judge families replicate the rubric on held-out human ESConv conversations. Data, code, the locked rubric, and judge prompts will be released upon publication.

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Durable Evaluation Framework: Adversarial Arbitration for Sycophancy Reduction in Large Language Models

RLHF-trained models are systematically biased toward agreement over accuracy, a structural property of the training process. We present Dur…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation

Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these ev…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Enhancing AI Interpretability and Safety through Localised Architectures

Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over t…

2026-06-09 22:47 JSTTechCrunch AIビジネス/資金調達

Sandstone raises $30M to bring AI to in-house legal teams

Sandstone's Series A comes just six months after a Sequoia-led seed round.

2026-06-09 21:00 JSTTechCrunch AIビジネス/資金調達

How an e-scooter founder raised $5 million to build space data centers

Orbital founder Euwyn Poon built 250,000 scooters at Spin. Now he wants to launch 10,000 space data centers.

2026-06-09 13:00 JSTarXiv cs.AIビジネス/資金調達

AI 認識的従属指数: おべっかの継続的な尺度

現在の AI モデルは認識論的な同調性を示し、ユーザーに同意するという主張を支持することがよくあります。既存の評価では、通常、モデルを二値支持にシフトさせるために何が必要かを評価するか、命題で明示的な確率を導き出すことによって、これを測定します。ただし、ユーザーに対するお調子者行動の多くは、通常の言語で表現される段階的サポートの変化を通じて示されます。私たちは AI Epistemic Deference Index (AEDI) を提案します。これは、モデルの出力で表現されるサポートが、ユーザーのプロンプトで表現される態度に対してどの程度敏感であるかを表す連続的な一次元スコアです。 AEDI を生成するために、人間の判断との一貫性と相関性が検証された判定者としての LLM を使用して、自然言語出力から確率を推定するための新しいプロトコルを提供します。私たちはこれを、さまざまなトピックにわたる 500 の提案と、ユーザーの態度が異なる 16,000 のプロンプトからなる厳選された新しいデータベースに展開し、8 つの著名なモデルをテストしました。どのモデルもかなりの差異を示しますが、プロバイダーごとに大きく体系的な違いがあり、Claude モデルが最も少なく、Grok モデルと Gemini モデルが最も多くなっています。この効果は、書かれたアーティファクトを要求するプロンプトで増幅され、モデルが弱い事前分布を保持する命題に集中します。 AEDI は、出力レベルのおしゃべり評価のための、更新が簡単なベンチマークおよび測定パイプラインとしてリリースされています。

原文 (English)

The AI Epistemic Deference Index: A Continuous Measure of Sycophancy

Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user. Existing evaluations typically measure this either by assessing what it takes to make a model shift a binary endorsement or by eliciting an explicit probability in a proposition. However, much user-facing sycophantic behavior is demonstrated through shifts in graded support expressed through ordinary language. We propose the AI Epistemic Deference Index (AEDI): a continuous, unidimensional score representing how sensitive the support expressed in a model's output is to the attitude expressed in a user's prompt. To generate AEDI, we provide a new protocol for estimating probabilities from natural language outputs, using LLMs-as-judges validated for consistency and correlation to human judgment. We deploy it on a new curated database of 500 propositions across diverse topics and 16,000 prompts varying in user attitude, testing eight prominent models. Every model exhibits substantial deference, though with large and systematic differences across providers, with Claude models demonstrating the least, and Grok and Gemini models the most. The effect is amplified in prompts requesting a written artifact, and concentrated on propositions where models hold weaker priors. We release AEDI as an easy-to-update benchmark and measurement pipeline for output-level sycophancy evaluation.

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体ビジネス/資金調達

オンライン エージェント アズ ア ジャッジ: インタラクティブ エージェントの状況を生み出す評価

社会的に関連した行動は、孤立した出力だけでなく、以前の相互作用、社会的役割、下流の行動にも依存するため、LLM を利用した対話型ソーシャル エージェントの評価は困難です。既存の方法では、通常、ターゲット エージェントが環境内で自由に行動し、その結果として得られる軌跡をスコアリングできます。ただし、この受動的な設定では、特定の社会的状況下でのみ観察可能になる機能が見逃される可能性があります。たとえば、意見の相違が生じない場合、競合処理はテストされないままになる可能性があります。私たちは、対話型ソーシャル エージェントのための状況生成評価フレームワークである Online Agent-as-a-Judge を提案します。 Online Agent-as-a-Judge は、環境のネイティブ対話およびアクション プロトコルを通じてターゲット エージェントと対話するインワールド評価エージェントをデプロイし、評価基準に関連する状況を積極的に引き出します。結果として得られる軌跡は、即時の反応とその後の行動の両方を評価するための証拠を提供します。 32 ドルのデザイナーが作成した社会的基準を備えたライフ シミュレーション環境では、オンライン エージェントとしての裁判官は、基準の適用範囲と人間のラベルとの一致を改善し、受動的手法では観察されない可能性がある行動について、より信頼性の高い証拠に基づいた評価をもたらします。

原文 (English)

Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents

Evaluating LLM-powered interactive social agents is challenging because socially relevant behaviors depend not only on isolated outputs, but also on prior interactions, social roles, and downstream actions. Existing methods typically allow a target agent to act freely in an environment and then score the resulting trajectory. However, this passive setup can miss capabilities that only become observable under specific social circumstances; for example, conflict handling may remain untested if no disagreement arises. We propose Online Agent-as-a-Judge, a situation-generating evaluation framework for interactive social agents. Online Agent-as-a-Judge deploys an in-world evaluator agent that interacts with the target agent through the environment's native dialogue and action protocol, actively eliciting situations relevant to the evaluation criteria. The resulting trajectories provide evidence for assessing both immediate responses and subsequent behavior. In a life-simulation environment with $32$ designer-authored social criteria, Online Agent-as-a-Judge improves criteria coverage and agreement with human labels, yielding more reliable evidence-grounded evaluations of behaviors that passive methods can leave unobserved.

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ブラックボックスのテスト: 消費者向けの健康 LLM の独立した評価に対する構造的障壁

背景: 消費者向けの大規模言語モデルは現在、健康情報の一般的な情報源となっており、応答を取得するのではなく、解釈してパーソナライズします。彼らの反応がユーザーによって異なるかどうかは、臨床、公平性、ガバナンスの問題であり、おべっかな反応が判断を変え、信頼を高める可能性があるという証拠によってさらに鮮明になります。目的: 通常の患者の使用に似た条件下で、消費者向けの健康 LLM における反応のばらつきと不平不満を評価すること。方法: 社会的背景と健康に対する態度を結びつける文献を参考にして、地理、閲覧状況、表明された信念、健康の社会的決定要因が異なるシミュレートされたユーザープロファイルを構築しました。私たちは、ワクチン接種態度検査スケールや生殖態度スケールなどの検証済みのツールを、ユーザー間で臨床的に意味のある変動を引き出すように設計されたマルチターン プロンプトに適合させました。結果: 評価では 5 つの関連する障壁に遭遇しました。事実に基づいたプロンプトは安定した応答を生成し、複数ターンの会話を通じて現れるお調子者を覆い隠しました。ブラウザベースのインターフェイスは、どの信号が出力に影響を与えるかを明らかにしておらず、クリーンなベースラインにリセットできませんでした。大規模なテストは、サービス規約、レート制限、ボット検出によって制限されていました。精度ベースの基準ではトーン、フレーミング、欠落を捉えることができず、LLM-as-judge 手法では共有アライメントバイアスの危険性がありました。追跡可能なバージョン識別子なしでモデルが変更されたため、信頼性の高いレプリケーションが妨げられました。結論: 消費者向けの健康 LLM が通常の使用においてどのように動作するかを調査するための信頼できる独立した評価フレームワークはまだ存在しません。監視には、パーソナライゼーションシグナル、安定バージョン識別子、研究者のセーフハーバープログラム、および健康関連出力の展開後のモニタリングの開示が必要です。

原文 (English)

Testing the Black Box: Structural Barriers to Independent Evaluation of Consumer-Facing Health LLMs

Background: Consumer-facing large language models are now a common source of health information, and they interpret and personalize responses rather than retrieve them. Whether their responses vary across users is a clinical, equity, and governance question, sharpened by evidence that sycophantic responses can alter judgment and increase trust. Objective: To evaluate response variation and sycophancy in consumer-facing health LLMs under conditions resembling ordinary patient use. Methods: We constructed simulated user profiles differing in geography, browsing context, expressed beliefs, and social determinants of health, drawing on literature linking social context to health attitudes. We adapted validated instruments, including the Vaccination Attitudes Examination scale and reproductive attitudes scales, into multi-turn prompts designed to elicit clinically meaningful variation across users. Results: The evaluation encountered five linked barriers. Factual prompts produced stable responses that masked sycophancy emerging over multi-turn conversation. Browser-based interfaces did not disclose which signals influence outputs and could not be reset to a clean baseline. Large-scale testing was restricted by terms of service, rate limits, and bot detection. Accuracy-based criteria could not capture tone, framing, or omission, and LLM-as-judge methods risked shared alignment bias. Models changed without traceable version identifiers, preventing reliable replication. Conclusions: No reliable independent evaluation framework yet exists for examining how consumer-facing health LLMs behave in ordinary use. Oversight requires disclosure of personalization signals, stable version identifiers, researcher safe harbor programs, and post-deployment monitoring of health-related outputs.

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

VESTA: LLM エージェント向けの完全に自動化されたシナリオ生成および安全性評価フレームワーク

大規模言語モデル (LLM) は、単純なテキストベースの対話システムから、メモリを維持し、ツールを使用し、外部環境にアクセスし、タスクを実行できる LLM エージェントへとますます進化しています。彼らの能力と自律性が拡大するにつれて、彼らが直面する安全リスクもより多様になります。既存の評価は、手動で作成されたシナリオ、静的なプロンプト、または最終出力の判断に依存していることが多く、タスクの実行中にエージェントが直面する可能性のあるさまざまなリスクを把握することが困難です。 LLM エージェント向けの完全に自動化されたシナリオ生成および安全性評価フレームワークである VESTA を紹介します。 VESTA は、5 つのリスク次元に基づいて、現実世界のタスク実行における抽象的で多様な安全リスクを 1,072 の測定可能な評価シナリオにインスタンス化します。自動評価パイプラインを使用して、12 個の LLM エージェントが 2 つの権限コンテキストの下で評価されます。その結果、現在のエージェントはタスク実行中に依然として重大な行動安全リスクに直面しており、平均 ASR は 47.1%、いくつかのモデルは 70% を超えていることが示されています。これらの調査結果は、LLM エージェントの安全性を理解し改善するために、実行可能なプロセスレベルの評価が重要であることを示しています。

原文 (English)

VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks. As their capabilities and autonomy expand, the safety risks they face also become more diverse. Existing evaluations often rely on manually written scenarios, static prompts, or final-output judgments, making it difficult to capture the diverse risks that agents may face during task execution. We introduce VESTA, a fully automated scenario generation and safety evaluation framework for LLM agents. Based on five risk dimensions, VESTA instantiaes abstract and diverse safety risks in real-world task execution into 1,072 measurable evaluation scenarios. Using the automated evaluation pipeline, 12 LLM agents are evaluated under two authority contexts. The results show that current agents still face substantial behavioral safety risks during task execution, with an average ASR of 47.1% and several models exceeding 70%. These findings demonstrate the importance of executable, process-level evaluation for understanding and improving LLM agent safety.

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

合格率を超えて: オープンコード LLM の多言語で実行に基づいた評価

コード生成モデルは通常、コンパクトな実行ベンチマークと総合格率を使用して比較されますが、そのような要約では、プログラミング言語、問題ファミリー、障害モードごとにパフォーマンスがどのように変化するかがわかりにくくなります。 12 のプログラミング言語にわたる 2,707 の無料 LeetCode 問題のコーディングに特化した、オープンにアクセスできる 9 つの LLM の大規模な実行ベースの評価を紹介します。私たちのコーパスには 325,343 の問題モデル言語ジョブが含まれており、それぞれがプロンプト メタデータ、抽出されたコード、LeetCode の実行結果、および静的分析信号にリンクされています。結果は、現在のオープン モデルが人間の許容基準からは程遠いことを示しています。最良のモデルである Yi-Coder-9B-Chat は、人間の許容ベースラインの 57.2% と比較して、平均正しさ 23.64% に達しています。ランキングもスライスに依存します。Qwen2.5-Coder-14B-Instruct は難しい問題と個別の問題のカバレッジで最も強力ですが、Gemma-2-27B-IT は全言語で最高の lint 合格率を達成しています。失敗分析では、コンパイル エラーが受け入れられなかった最良の送信の 63.25% を占めていることが示されており、セマンティックな正確性をテストする前に多くの失敗が発生していることが示されています。静的な品質は、機能的な正確さからさらに乖離します。これらの調査結果を総合すると、多言語でアーティファクトを保持した評価を行うことで、単一言語または単一指標のリーダーボードに隠されたトレードオフが明らかになることを示しています。

原文 (English)

Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs

Code generation models are typically compared using compact execution benchmarks and aggregate pass rates, but such summaries obscure how performance varies across programming languages, problem families, and failure modes. We present a large-scale, execution-grounded evaluation of 9 openly accessible LLMs specialized for coding on 2,707 free LeetCode problems across 12 programming languages. Our corpus contains 325,343 problem-model-language jobs, each linked to prompt metadata, extracted code, LeetCode execution outcomes, and static-analysis signals. The results show that current open models remain far from the human acceptance reference: the best model, Yi-Coder-9B-Chat, reaches 23.64% mean correctness, compared with a 57.2% human acceptance baseline. Rankings are also slice-dependent: Qwen2.5-Coder-14B-Instruct is strongest on hard problems and distinct-problem coverage, while Gemma-2-27B-IT achieves the highest all-language lint pass rate. Failure analysis shows that compile errors account for 63.25% of non-accepted best submissions, indicating that many failures occur before semantic correctness can be tested. Static quality further diverges from functional correctness. Together, these findings show that multilingual, artifact-preserving evaluation reveals tradeoffs hidden by single-language or single-metric leaderboards.

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LATTEArena: LLM を利用した表形式特徴量エンジニアリングの評価フレームワーク (拡張バージョン)

特徴量エンジニアリングは表形式データ分析にとって依然として不可欠であり、大規模言語モデル (LLM) がこのプロセスを自動化するための有望なパラダイムとして台頭し、LLM を利用した AuTomated 表形式特徴量エンジニアリング (LATTE) が誕生しました。ただし、標準化されたプラットフォームがないため、コストを意識した公平な比較ができません。さらに、複雑な方法論的設計により、個々のコンポーネントの具体的な貢献がわかりにくくなります。たとえば、LFG は思考ツリー、少数ショット デモンストレーション、モンテカルロ ツリー検索、自然言語生成を統合していますが、各技術の競争力による個別の影響は定量化されていません。これらの課題に対処するために、私たちは次の特徴を備えた最初の競争評価フレームワークである LATTEArena を導入します。(1) 15 の代表的な手法を再利用可能なコンポーネントに分解する 6 次元の分類。 (2) 制御された比較のための標準化されたモジュール式アリーナ。 (3) パフォーマンス、コスト、堅牢性をカバーする多次元の評価。 (4) 各技術の競争力を定量化するコンポーネントレベルのアブレーション。広範な評価を通じて、次のような 16 の重要な発見が明らかになりました。(1) モンテカルロ木検索による思考の木は、最適な費用対効果を実現します。 (2) RPN とコードの出力形式は、それぞれ分類タスクと回帰タスクを支配します。私たちはモジュール式フレームワークと 4,000 を超える実行ログを公開し、研究者が新しい技術と既存の技術をシームレスに比較して LATTE を進歩できるようにします。

原文 (English)

LATTEArena: An Evaluation Framework for LLM-powered Tabular Feature Engineering (Extended Version)

Feature engineering remains essential for tabular data analysis, and Large Language Models (LLMs) have emerged as a promising paradigm for automating this process, giving rise to LLM-powered AuTomated Tabular feature Engineering (LATTE). However, the absence of standardized platforms prevents fair, cost-aware comparisons. Furthermore, complex methodological designs obscure the specific contributions of individual components; for example, although LFG integrates Tree-of-Thought, few-shot demonstrations, Monte Carlo Tree Search, and natural language generation, the isolated impact of each technique's competitive edge remains unquantified. To address these challenges, we introduce LATTEArena, the first competitive evaluation framework featuring: (1) a six-dimensional taxonomy decomposing 15 representative methods into reusable components; (2) a standardized modular arena for controlled comparison; (3) multi-dimensional assessments covering performance, cost, and robustness; and (4) component-level ablation quantifying each technique's competitive edge. Through extensive evaluations, we reveal 16 key findings, including: (1) Tree-of-Thought with Monte Carlo Tree Search achieves optimal cost-effectiveness; (2) RPN and Code output formats dominate classification and regression tasks, respectively. We publicly release the modular framework and over 4000 execution logs, enabling researchers to seamlessly pit new techniques against existing ones and advance LATTE.

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ComplexConstraints and Beyond: Expert Rubrics for RLVR

As LLM capabilities advance rapidly, the evaluation methods used to assess them increasingly lag behind. Traditional benchmarks relied on p…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they rem…

2026-06-09 13:00 JSTarXiv cs.AIビジネス/資金調達

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult…

2026-06-09 13:00 JSTarXiv cs.AIビジネス/資金調達

From Coarse to Fine: Managing Temporal Granularity in Spatio-Temporal Data for Fine-Grained Traffic Prediction

Efficient acquisition, storage, and utilization of traffic data are critical challenges in spatio-temporal data management. Most traffic da…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

TheoremBench: Evaluating LLMs on Theorem Proving in Formal Mathematics

LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on compe…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation

AI Scientist agents are often evaluated as if capability were mainly a function of model quality, prompting, or reasoning scaffolds. We tes…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体ビジネス/資金調達研究/論文

Multi-Turn Evaluation of Deep Research Agents Under Process-Level Feedback

Existing benchmarks for deep research agents (DRAs) assess only single-shot outputs, ignoring a key question: can DRAs improve their report…

2026-06-09 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blo…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation

Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning. Howeve…

2026-06-09 13:00 JSTarXiv cs.AIビジネス/資金調達

Selecting New Measurement Locations to Diversify Traffic-Pattern Coverage: A Real-World Evaluation for Total Traffic Volume Estimation

Accurate measurement of traffic volumes and flows is vital for modern intelligent transportation. However, despite recent technological adv…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks

Standard AI red teaming evaluations reduce adversarial campaigns to a single binary outcome, attack success rate (ASR), not taking into acc…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese

Large Language Models are transforming the support for clinical decision and their application in real scenarios. Yet, most benchmarks are…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation

Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these ev…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Enhancing AI Interpretability and Safety through Localised Architectures

Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over t…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, a…

2026-06-09 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Closing the Sim-to-Real Gap: An Evaluation Framework for Autonomous Cyber Defense Configuration of Commercial EDR

Leading commercial endpoint detection and response (EDR) products have shifted from operator-configured rule sets to multi-component system…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models

Large Audio-Language Models (LALMs) integrate audio perception and language understanding within a unified framework, enabling a wide range…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy

Most evaluations of LLM agents look like exams: a discrete task, a clean environment, a score in minutes or hours. We argue that this appro…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics

Diffusion and continuous flow-based language models have emerged as the leading non-autoregressive alternatives to language modeling. Progr…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation

Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs). By const…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

RadOT-Eval: Auditable Structured-Evidence Transport for Radiology Report Evaluation

Automatic evaluation is critical for high-stakes text generation, where errors often involve omitted findings, hallucinated content, polari…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Evaluating AI Investment Strategies

We study the problem of auditing a black-box algorithmic decision-maker from observable inputs and outputs alone. Our main result is an exa…

2026-06-09 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達研究/論文

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily…

2026-06-09 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

A multi-agent system for spine MRI report generation from multi-sequence imaging

Spinal pathology is a leading cause of pain and disability worldwide. Spine MRI is central to clinical evaluation, yet its interpretation r…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

SEF-CLGC at SemEval-2026 Task 11: Logical Notation Impact on Language Model Performance

This paper revisits our pipeline called Syllogistic Evaluation Framework-Common Logic Grammar Construction (SEF-CLGC). We combine formal lo…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis

Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks in…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Knowing How to Edit: Reliable Evaluation Signals for Diagnosing and Optimizing Prompts at Query Level

Prompt optimization has become a central mechanism for eliciting strong performance from LLMs, and recent work has made substantial progres…

2026-06-09 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

Entropy-Based Evaluation of AI Agents: A Lightweight Framework for Measuring Behavioral Patterns

AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important asp…

2026-06-09 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Comparative evaluation of training strategies using partially labelled datasets for segmentation of white matter hyperintensities and stroke lesions in FLAIR MRI

White matter hyperintensities (WMH) and ischaemic stroke lesions (ISL) are key imaging biomarkers of cerebral small vessel disease (SVD) de…

2026-06-09 13:00 JSTarXiv cs.AIビジネス/資金調達

Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design

Deriving predictable scaling laws that govern the relationship between model performance and computational investment is crucial for design…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Context Over Compute Human-in-the-Loop Outperforms Iterative Chain-of-Thought Prompting in Interview Answer Quality

Behavioral interview evaluation using large language models presents unique challenges that require structured assessment, realistic interv…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook

As LLMs are globally deployed, aligning their cultural value orientations is critical for safety and user engagement. However, existing ben…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Multilingual Training and Evaluation Resources for Vision-Language Models

Vision Language Models (VLMs) achieved rapid progress in the recent years. However, despite their growth, VLMs development is heavily groun…

2026-06-09 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Evaluating Design Video Generation: Metrics for Compositional Fidelity

Generative video models are increasingly used in design animation tasks, yet no standardized evaluation framework exists for this domain. U…

2026-06-09 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

メタ学習による費用対効果の高いモデル評価

機械学習の急速な成長により、拡大し続けるモデルのエコシステムが生み出され、目に見えないラベルのないデータに対して新しくリリースされたモデルの信頼性を検証することがますます困難になっています。従来の評価パイプラインは、高価なアノテーション、繰り返しの微調整、またはモデル ファミリ間での転送ができない狭い仮定に依存しています。さまざまなアーキテクチャやモダリティにまたがる未確認のモデルをラベルなしで迅速に評価するための、コスト効率が高く、モデルに依存しないフレームワークである MetaEvaluator を紹介します。 MetaEvaluator は、参照モデルのプールに対するメタ学習を利用して転送可能な初期化を取得し、プール全体でコストを償却しながら、モデルごとの再トレーニングの必要性を排除しながら、新しいモデルの正確な評価を可能にします。私たちの知る限り、これは完全にラベルのないデータセットで新しいモデルを評価できる、モデルに依存しない最初のフレームワークです。広範な実験により、MetaEvaluator は従来のアプローチと比較して大幅にコストを削減しながら安定した正確なパフォーマンス推定値を生成し、ラベルのないデータに対する新しいモデルのスケーラブルなベンチマークを実用化できることが示されています。

原文 (English)

Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-Learning

The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data. Existing evaluation pipelines typically rely on costly annotation, repeated fine-tuning, or assumptions that do not generalize well to new models. We introduce MetaEvaluator, a cost-effective, model-agnostic framework for fast, label-free evaluation of unseen models across diverse architectures and modalities. MetaEvaluator meta-learns over a pool of reference models to acquire an effective initialization for accurate assessment of unseen models, thereby amortizing evaluation cost and eliminating the need for per-model retraining. To the best of our knowledge, this is the first model-agnostic framework that evaluates new models on unlabeled datasets. Extensive experiments demonstrate that MetaEvaluator delivers stable and accurate performance estimates at substantially lower cost than conventional approaches, enabling scalable benchmarking on unlabeled datasets for emerging models. The code is available at: https://github.com/phkhanhtrinh23/MetaEvaluator.

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests

A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead o…

2026-06-09 10:25 JSTITmedia AI+LLM/生成AIビジネス/資金調達

OpenAIがIPO申請を発表 時期は未定

ChatGPTを手がける米OpenAIは6月8日(現地時間)、米国での新規株式公開(IPO)を内密に申請していたことを発表した。

2026-06-09 09:45 JSTTechCrunch AIビジネス/資金調達

Mercor’s Brendan Foody calls out Sequoia, accusing it of ‘dual-pricing’ valuation tricks

Sequoia is just one of the top firms that sells same equity at two different prices.

2026-06-09 07:41 JSTTechCrunch AILLM/生成AIビジネス/資金調達

As OpenAI files for IPO, Sam Altman’s eye-scanning company is doing layoffs, report says

Tools for Humanity, Sam Altman's identity verification company, is reportedly struggling to generate revenue and will downsize its staff.

2026-06-09 06:29 JSTTechCrunch AILLM/生成AIビジネス/資金調達

OpenAI files confidentially for IPO, following Anthropic

The filing comes a little more than a week after its main rival, Anthropic, also filed to go public, ramping up the race between the two AI…

2026-06-09 02:27 JSTITmedia AI+LLM/生成AIビジネス/資金調達

OpenAIが上場へ SpaceX・Anthropicに続きIPO申請

米OpenAIは、米国証券取引委員会にIPO申請したと発表した。

2026-06-08 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Agentic AI 制御評価における攻撃選択により安全性が大幅に低下する

いつ攻撃するかを戦略的に選択する攻撃者は、無差別に攻撃する攻撃者よりも捕まえるのがはるかに困難です。 AI コントロールは、能力はあるが信頼できない AI エージェントを、弱い信頼できる監視者の監視と限られた人間の監査予算の下で展開するための安全フレームワークです。コントロールの評価では、レッドチームの攻撃ポリシーとブルーチームの監視を比較することで、これらのプロトコルをストレステストしますが、現在の評価では通常、攻撃者がいつ攻撃するかを戦略的に選択していないことを想定しています。私たちは、攻撃の決定を、攻撃者がいつ攻撃すべきかを決定する開始ポリシーと、攻撃者が進行中の攻撃をいつ中止すべきかを決定する停止ポリシーに分解することによって、エージェント設定におけるこの機能、つまり攻撃の選択を研究します。 BashArena と LinuxArena の 2 つのエージェント設定にわたって、どちらのポリシーも、基礎となる攻撃能力を変えることなく、測定された経験的安全性を大幅に低下させます。 1% の監査予算では、開始ポリシーは BashArena と LinuxArena の両方で安全性を 20pp 低下させ、停止ポリシーは BashArena で 20pp、LinuxArena で 28pp ずつ安全性を低下させます。これらの減少は、攻撃選択の効果の上限として解釈される必要があります。したがって、既存の制御評価では、選択的攻撃者に対して過度に楽観的な安全性推定が得られる可能性があります。今後の評価、システム カード、および安全性ケースによって攻撃の選択を導き出し、より現実的な安全性の見積もりを作成することをお勧めします。

原文 (English)

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but untrusted AI agents under the oversight of a weaker, trusted monitor and a limited human audit budget. Control evaluations stress-test these protocols by pitting a red-team attack policy against the blue-team monitor, but current evaluations typically assume attackers that do not strategically select when to attack. We study this capability, attack selection, in agentic settings by decomposing attack decisions into a start policy, which decides when an attacker should attack, and a stop policy, which decides when an attacker should abort an ongoing attack. Across two agentic settings, BashArena and LinuxArena, both policies substantially lower measured empirical safety without changing the underlying attack capability. At a 1% audit budget, our start policy reduces safety by 20pp on both BashArena and LinuxArena, and our stop policy reduces safety by 20pp on BashArena and 28pp on LinuxArena. These reductions should be interpreted as upper bounds on the effect of attack selection. Existing control evaluations may therefore yield overly optimistic safety estimates against selective attackers. We recommend that future evaluations, system cards, and safety cases elicit attack selection to produce more realistic safety estimates.

2026-06-08 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

コンテキストの説明: 価値観の一致のための道徳的信念の形成

エージェントの行動が人間の道徳的価値観と一致していることを確認すると、社会、さらには個人が通常採用する複数の道徳的観点をどのように説明するかという問題が必然的に生じます。道徳的不確実性に関する研究では、さまざまな道徳理論にわたる行動の評価を公平かつ民主的に集約するメカニズムが提案されています。しかし、この論文は、道徳的評価を集計する際には、文脈上の要因を考慮する必要があると主張しています。たとえば、結果主義的な視点は、エージェントの行動が世界をどのように変えるかを正確に判断する能力を前提としています。この仮定は、現実世界の設定では当てはまらないことがよくあります。したがって、私たちは、このような種類の文脈上の要因も考慮しながら、道徳的不確実性の下でエージェントの意思決定を形式化します。これにより、一見常識的な特性、つまり弱いパレートの法則が違反されていることを示します。私たちは、この見かけの問題は実際にはシンプソンのパラドックスの変形であり、したがってコンテキスト要因の影響を無視する集計メカニズムの限界を明らかにしていると主張します。

原文 (English)

Accounting for Context: Shaping Moral Credences for Value Alignment

Ensuring that agent behaviours are aligned with human moral values inevitably raises the problem of how to account for the plurality of moral perspectives that societies -- and even individuals -- typically adopt. Work on moral uncertainty proposes mechanisms to fairly and democratically aggregate evaluations of actions across different moral theories. However, this paper argues that one needs to account for contextual factors when aggregating moral evaluations. For example, consequentialist perspectives assume an ability to accurately determine how an agent's actions change the world; an assumption that often does not hold in real world settings. We, therefore, formalise agent decision making under moral uncertainty, while also accounting for these kinds of contextual factors. We thereby show that a seemingly commonsensical property -- the weak Pareto principle -- is violated. We argue that this apparent problem is, in fact, a variation of Simpson's paradox, and hence reveals the limitations of aggregation mechanisms that ignore the impact of contextual factors.

2026-06-08 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

現地開示による戦略エージェントによるオフポリシー評価

私たちは、意思決定主体(またはエージェント)が共変量を戦略的に変更することで意思決定者のポリシーに応答する、戦略的行動の下でのオフポリシー評価(OPE)を研究します。このような動作は、ポリシー依存の共変量シフトを誘発し、共変量がポリシーにとって外生であるという既存の手法における標準的な仮定を破ります。関連研究では、反復的なインタラクションやエージェントの応答行動に関する完全な知識などの強い仮定を課すことで、この課題に対処しており、OPE への適用可能性は大幅に制限されています。対照的に、意思決定者がエージェントの応答動作について部分的な知識しか持っていないワンショット OPE 設定を検討します。私たちの重要な洞察は、事後説明を通じてローカル情報を開示すると、適応前のエージェントの戦略前の共変量が明らかになり、戦略的行動によって誘発される情報損失が軽減されるということです。この構造を利用して、エージェントの応答の統計モデルを推定し、政策価値の二重にロバストな推定量を構築します。エージェントのコスト感度が条件付き対数正規分布に従うと仮定することで、提案された推定量の一貫性を確立し、経験的にアプローチを検証します。より広範に、私たちの結果は、インタラクションデザインがエージェントの戦略的応答の隠された構造を明らかにすることによって、情報の非対称性をどのように軽減できるかを浮き彫りにしています。

原文 (English)

Off-Policy Evaluation with Strategic Agents via Local Disclosure

We study off-policy evaluation (OPE) under strategic behavior where decision subjects (or agents) respond to a decision maker's policy by strategically modifying their covariates. Such behavior induces a policy-dependent covariate shift, breaking the standard assumption in existing methods that covariates are exogenous to the policy. Related work addresses this challenge by imposing strong assumptions such as repeated interactions or full knowledge of agents' response behavior, substantially limiting its applicability to OPE. In contrast, we consider a one-shot OPE setting where the decision maker has only partial knowledge of the agents' response behavior. Our key insight is that disclosing local information through post-hoc explanations reveals agents' pre-strategic covariates prior to adaptation, mitigating the information loss induced by strategic behavior. Leveraging this structure, we estimate a statistical model for the agents' responses and construct a doubly robust estimator for policy value. By assuming that the agents' cost sensitivity follows a conditional log-normal distribution, we establish consistency of the proposed estimator and validate our approach empirically. More broadly, our results highlight how interaction design can mitigate information asymmetry by revealing otherwise hidden structure in agents' strategic responses.

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM パーソナライゼーションにおける人間の再中心化

関心の高まりにもかかわらず、大規模言語モデル (LLM) のパーソナライゼーション能力の評価のほとんどは合成データに依存していました。現在のパーソナライゼーション システムが実際のユーザーに対してどの程度うまく機能するかは依然として不明です。このペーパーでは、合成データと人間のデータを使用した場合の LLM パーソナライゼーションのパフォーマンスの差を研究します。当社はパーソナライゼーションの 3 つの段階にわたって人間の会話 (550 件の会話) と判断を収集します。つまり、会話からのユーザー属性の抽出 (5,949 件の判断)、関連する属性と新しいプロンプトの組み合わせ (11,919 件)、およびパーソナライズされた応答への関連属性の組み込み (1,101 件) です。人間のデータを組み込むと、各段階でシステムの限界が明らかになります。モデルは、人間の会話から属性を抽出し、関連する属性について人間の判断に同意せず、人間が一般的な応答と同等と判断するパーソナライズされた応答を生成するのに苦労しています(ただし、LLM の判断では一般的な応答の方が優れていると広く評価されています)。最初の 2 つの段階で、自動化されたパーソナライゼーション評価を人間のデータに近づける 2 つの軽量のトレーニングベースの介入を導入します。しかし、第 3 段階では、学習された報酬モデルが人間の評価とわずかな相関関係しか達成していないことがわかり、人間に合わせたパーソナライゼーションの品質判断を直接モデル化するのは難しいことが示唆されています。私たちが収集したデータは、人間が役立つと思われる方法でモデルがユーザー情報をどのように抽出、選択、組み込むかを研究するための基盤を提供します。

原文 (English)

Re-Centering Humans in LLM Personalization

Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on synthetic data. It remains unclear how well current personalization systems work for real users. In this paper, we study the gap in LLM personalization performance when using synthetic versus human data. We collect human conversations (550 conversations) and judgments across three stages of personalization: extracting user attributes from conversations (5,949 judgments), pairing relevant attributes with new prompts (11,919), and incorporating relevant attributes into a personalized response (1,101). Incorporating human data reveals system limitations at each stage. Models struggle to extract attributes from human conversations, disagree with human judgments on relevant attributes, and generate personalized responses that humans judge no better than generic responses (though that LLM judges widely rate as better). We introduce two lightweight training-based interventions that shift automated personalization evaluation closer to human data in our first two stages. However, in our third stage we find that learned reward models achieve only modest correlation with human ratings, suggesting that human-aligned personalization quality judgments are difficult to model directly. Our collected data provides a foundation for studying how models should extract, select, and incorporate user information in ways that humans find useful.

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios

Hallucination detection is essential for the reliable deployment of large language models (LLMs). However, existing evaluations face two co…

2026-06-08 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

REMEDI: A Benchmark for Retention and Unlearning Evaluation in Multi-label Clinical Disease Inference

Language models trained for clinical disease inference are trained on patient data, which may include sensitive and private information, an…

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million p…

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests

A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead o…

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

A Comprehensive Anatomy of Human and DeepSeek-R1 LLM Mathematical Reasoning

The emergence of "Aha moments" in large language models, particularly DeepSeek-R1-0120, has raised the question of whether these systems ge…

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM-Augmented Digital Twin for Policy Evaluation in Short-Video Platforms

Short-video platforms are closed-loop, human-in-the-loop ecosystems where platform policy, creator incentives, and user behavior co-evolve.…

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SWE-IF: Aligning Code Evaluation with Human Preference

Large Language Models (LLMs) have catalyzed vibe coding, where users leverage LLMs to generate and iteratively refine code through natural…

2026-06-08 13:00 JSTarXiv cs.AIビジネス/資金調達

$\mathrm{ECI}_{\mathrm{sem}}$: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives

Hard-negative source selection for dense retrieval is usually decided only after fine-tuning and downstream evaluation. We propose $\mathrm…

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval

Engineering rulebooks and technical standards contain multimodal information like dense text, tables, and illustrations that are challengin…

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

メタエージェントの課題: 現在のエージェントは自律的なエージェント開発が可能ですか?

現在の AI ベンチマークは、人間が設計したワークフロー内でのタスク実行に関してエージェントを評価します。これらの評価では、基本的に、モデルが自律的にエージェント システムを開発できるかどうかという、重要な次のレベルの機能を測定できません。自律エージェント開発のためのフロンティア モデルの能力をテストするために設計された評価フレームワークであるメタエージェント チャレンジ (MAC) を紹介します。具体的には、コード エージェント (メタエージェント) には、サンドボックス環境、評価 API、および 5 つのドメインにわたって実施されたテスト セットのパフォーマンスを最大化するエージェント アーティファクトを反復的にプログラムするための時間制限が与えられます。評価の整合性を確保するために、このフレームワークは報酬ハッキングに対する多層防御によって保護されています。このフレームワークを活用して、メタエージェントが人為的に設計されたベースライン ポリシーと一致することはほとんどなく、一致する少数のエージェントは独自のフロンティア モデルによって支配されていることを示します。さらに、設計プロセスは高い分散を示し、高い最適化圧力により、グラウンドトゥルースの漏洩などの敵対的な動作が表面化し、堅牢性とモデルの調整の両方における重大な欠陥が浮き彫りになります。最終的に、MAC は自律型 AI の研究開発のための厳密なオープンソース ベンチマークを提供し、再帰的な自己改善を評価するための経験的な代用手段を提供します。ベンチマークは https://github.com/ant-research/meta-agent-challenge で公開されています。

原文 (English)

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

Current AI benchmarks evaluate agents on task execution within human-designed workflows. These evaluations fundamentally fail to measure a critical next-level capability: whether models can autonomously develop agent systems. We introduce the Meta-Agent Challenge (MAC), an evaluation framework designed to test the capacity of frontier models for autonomous agent development. Specifically, a code agent (the meta-agent) is given a sandboxed environment, an evaluation API, and a time limitation to iteratively program an agent artifact that maximizes performance on a held-out test set across five domains. To ensure evaluation integrity, this framework is secured by multi-layer defenses against reward hacking. Leveraging this framework, we demonstrate that meta-agents rarely match human-engineered baseline policies, and the few that do are dominated by proprietary frontier models. Moreover, the design process exhibits high variance, and high optimization pressure surfaces emergent adversarial behaviors like ground-truth exfiltration-highlighting critical deficits in both robustness and model alignment. Ultimately, MAC provides a rigorous, open-source benchmark for autonomous AI research and development, offering an empirical proxy for evaluating recursive self-improvement. Benchmark is publicly available at: https://github.com/ant-research/meta-agent-challenge.

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

SpurAudio: 少数ショット音声分類におけるショートカット学習を研究するためのベンチマーク

少数ショット分類 (FSC) は、限られたラベル付きデータから学習するために広く使用されていますが、ほとんどの評価は、ターゲットの概念が文脈上の手がかりから独立していることを暗黙的に前提としています。ただし、現実世界の設定では、サンプルがリッチ コンテキスト内に表示されることが多く、モデルが前景コンテンツと背景信号の間の偽の相関を利用できるようになります。このような効果は少数ショット画像分類で研究されていますが、少数ショット音声分類におけるその役割はほとんど解明されておらず、既存の音声ベンチマークでは文脈構造に対する制御が限られています。 SpurAudio というベンチマークを紹介します。これは、オーディオの前景イベントと背景環境の自然な分離性を活用して、サポートおよびクエリ セットにわたるコンテキストの変化を制御されたマルチレベルの評価を可能にするベンチマークです。このベンチマークを使用して、多くの最先端の少数ショット手法は、標準的な評価プロトコルで同様の精度を達成しているにもかかわらず、バックグラウンド相関が破壊されると重大なパフォーマンス低下に見舞われることがわかります。重要なのは、この脆弱性は大規模な事前トレーニング済みオーディオ基盤モデルでも存続しており、バックボーン容量の制限が説明の対象外となっているということです。さらに、従来のベンチマークでは同等に見える手法でも、偽の相関に対して著しく異なる感度を示す可能性があり、推論時に特徴表現が分類器ヘッドとどのように相互作用するかに関連する体系的なアルゴリズムの強みと脆弱性が明らかになります。これらの発見は、オーディオにおける少数ショット法の動作に関する新たな洞察を提供し、FSC モデルを評価する際のコンテキスト依存性を明示的に調査するベンチマークの必要性を強調しています。

原文 (English)

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues. In real-world settings, however, examples often appear within rich contexts, allowing models to exploit spurious correlations between foreground content and background signals. While such effects have been studied in few-shot image classification, their role in few-shot audio classification remains largely unexplored, and existing audio benchmarks offer limited control over contextual structure. We introduce SpurAudio, a benchmark that leverages the natural separability of foreground events and background environments in audio to enable controlled, multi-level evaluation of contextual shifts across support and query sets. Using this benchmark, we show that many state-of-the-art few-shot methods suffer severe performance degradation when background correlations are disrupted, despite achieving similar accuracy under standard evaluation protocols. Crucially, this vulnerability persists even in large pretrained audio foundation models, ruling out limited backbone capacity as an explanation. Moreover, methods that appear comparable under conventional benchmarks can exhibit markedly different sensitivity to spurious correlations, revealing systematic algorithmic strengths and vulnerabilities tied to how feature representations interact with classifier heads at inference time. These findings provide new insight into the behavior of few-shot methods in audio and highlight the need for benchmarks that explicitly probe context dependence when evaluating FSC models.

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

L-TGVN: パーソナライズされた高速 MRI のための縦方向事前分布の活用

MRI は電離放射線を使用せずに優れた軟組織コントラストを提供しますが、取得時間が長いため患者の不快感が増大すると同時に、検査コストが上昇し、スキャナのスループットが制限されます。スキャン時間を短縮するための一般的なアプローチは、取得する測定値を少なくすることです。これにより、不適切な線形逆問題が発生します。したがって、診断品質の画像を回復するには、測定データ以外の事前知識を組み込む必要があります。追跡検査では、患者の最新の以前のスキャンにより、非常に有益な被験者固有のコンテキストが提供されますが、実際の使用は、時間的変化(病状の進行を含む)、スキャン間のずれ、取得間のプロトコルのドリフトによって複雑になります。この研究では、大幅にアンダーサンプリングされた測定値から現在のスキャンを再構築するための副次情報として以前のスキャンを活用する、縦方向の信頼誘導変分ネットワークである L-TGVN を紹介します。重要なことは、L-TGVN は、以前のスキャンの影響が取得された測定値と一致するように制限することです。既存の多くの縦方向再構成方法とは異なり、以前のスキャンと現在のスキャンの間の明示的な事前位置合わせを必要としません。さらに、訪問ごとの取得プロトコルの違い(シーケンスパラメータの変更など)にも対応します。私たちは、事前ガイド法や縦方向事前分布を使用しない方法など、一致した容量のベースラインに対して L-TGVN を評価し、困難な加速において微細構造のより良好な保存とともに、標準的な定量的指標の一貫した改善を観察しました。ソース コードは github.com/sodicksonlab/L-TGVN で入手できます。

原文 (English)

L-TGVN: Leveraging Longitudinal Priors for Personalized Rapid MRI

MRI provides excellent soft-tissue contrast without ionizing radiation, but long acquisition times increase patient discomfort while also raising exam costs and limiting scanner throughput. A common approach to reduce scan time is to acquire fewer measurements, which yields an ill-posed linear inverse problem; recovering diagnostic-quality images therefore requires incorporating prior knowledge beyond the measured data. In follow-up exams, the most recent prior scan of a patient can provide a highly informative subject-specific context, but practical use is complicated by temporal changes (including pathology progression), misalignment between scans, and protocol drift across acquisitions. In this work, we introduce L-TGVN, a Longitudinal Trust-Guided Variational Network that leverages prior scans as side information to reconstruct the current scan from heavily undersampled measurements. Crucially, L-TGVN constrains the influence of prior scans to be consistent with the acquired measurements. Unlike many existing longitudinal reconstruction methods, it does not require explicit pre-registration between prior and current scans. It further accommodates differences in acquisition protocols across visits (e.g., changes in sequence parameters). We evaluate L-TGVN against matched-capacity baselines, including prior-guided methods and methods that do not use longitudinal priors, and observe consistent improvements in standard quantitative metrics together with better preservation of fine structures at challenging accelerations. Source code is available at github.com/sodicksonlab/L-TGVN.

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達

RowNet: 表形式回帰のためのメモリ トランスフォーマー

不動産評価は構造化回帰問題であり、価格は異種の特徴タイプ、まばらな地域効果、非線形相互作用、および比較可能な不動産の実際的なロジックによって支配されます。標準的な多層パーセプトロンは各行を孤立ベクトルとして扱い、局所性、スケール感度、およびカテゴリカルマッチングを監視のみから学習する必要があります。勾配ブースト デシジョン ツリーは強力な表形式のベースラインを提供しますが、その特徴中心の分割メカニズムは、類似した履歴観測の取得を明示的にモデル化しません。この論文では、不動産の平方メートルあたりの価格を予測するための検索ベースのニューラル アーキテクチャである RowNet について説明します。 RowNet は、ラベル付きプロパティのメモリ バンクに対するペアごとの類似性機能を通じてクエリ プロパティを表します。最初の検索層は、特徴のみの類似性から大まかなターゲットを推定します。 2 番目の層は、ターゲット一貫性機能を使用してメモリ比較を強化し、複数の学習されたアテンション ヘッドを使用して相補的な比較可能なセットを取得します。最後の専門家混合モジュールは、学習されたゲーティング、残差補正、エントロピー正則化、ヘッドダイバーシティ正則化を組み合わせて予測を生成します。

原文 (English)

RowNet: A Memory Transformer for Tabular Regression

Real estate valuation is a structured regression problem in which prices are governed by heterogeneous feature types, sparse regional effects, nonlinear interactions, and the practical logic of comparable properties. Standard multilayer perceptrons treat each row as an isolated vector and must learn locality, scale sensitivity, and categorical matching from supervision alone. Gradient-boosted decision trees provide strong tabular baselines, but their feature-centric splitting mechanism does not explicitly model the retrieval of similar historical observations. This paper presents RowNet, a retrieval-based neural architecture for real estate price-per-square-meter prediction. RowNet represents a query property through pairwise similarity features against a memory bank of labeled properties. A first retrieval layer estimates a coarse target from feature-only similarities. A second layer augments the memory comparison with target-consistency features and uses multiple learned attention heads to retrieve complementary comparable sets. A final mixture-of-experts module combines learned gating, residual correction, entropy regularization, and head-diversity regularization to produce the prediction.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

共同生成と評価による自己進化する深層研究

大規模言語モデル (LLM) は日常のアプリケーションでますます採用されるようになり、詳細な研究が特に重要な機能として際立っています。従来の質問応答 (QA) タスクとは異なり、詳細な調査レポートの生成には決定的な根拠が欠けているため、報酬設計が本質的に検証不可能になり、効果的な強化学習が制限されます。既存のアプローチでは、LLM-as-a-judge およびクエリ依存の評価ルーブリックを使用してこの課題を軽減していますが、依然として静的な評価器に依存しているため、ソルバーの向上に応じて標準を適応させることができず、最適化圧力が不十分になり、最終的に飽和状態になってしまいます。私たちは、\textbf{s}elf 進化型 \textbf{co} 進化型トレーニング フレームワークで、深い \textbf{re} 検索の評価と生成 (SCORE) を使用してこの制限に対処します。これは、共有パラメータ学習プロセスにおいて評価器とソルバーを緊密に結合します。生成と評価を独立したモジュールとして扱うのではなく、それらの本質的なつながりを活用して、単一の共有パラメーター モデル内で共同の改善を可能にします。このプロセスを制限するために、ソルバーのパフォーマンスに基づいて評価環境を動的に制御するメタハーネスを導入し、有効な評価次元と十分に深い評価者の検索を促進します。ディープリサーチベンチマークに関する広範な実験により、レポート生成の品質が一貫して向上していることが実証されており、評価と生成を共進化させることが、オープンエンドのリサーチエージェントをトレーニングするための有望な方向性であることが示されています。

原文 (English)

Self-Evolving Deep Research via Joint Generation and Evaluation

Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional question-answering (QA) tasks, deep research report generation lacks definitive ground-truth, making reward design inherently unverifiable and limiting effective reinforcement learning. Existing approaches mitigate this challenge with LLM-as-a-judge and query-dependent evaluation rubrics, but they still rely on static evaluators that cannot adapt their standards as the solver improves, leading to insufficient and eventually saturated optimization pressure. We address this limitation with a \textbf{s}elf-evolving \textbf{co}-evolutionary training framework for deep \textbf{re}search evaluation and generation (SCORE), which tightly couples an evaluator and a solver in a shared-parameter learning process. Rather than treating generation and evaluation as isolated modules, we leverage their intrinsic connection to enable joint improvement within a single shared-parameter model. To restrict this process, we introduce a meta-harness, which dynamically controls the evaluation environment based on solver performance, encouraging valid evaluation dimensions and sufficiently deep evaluator search. Extensive experiments on deep research benchmarks demonstrate consistent improvement in report generation quality, showing that co-evolving evaluation and generation is a promising direction for training open-ended research agents.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

M$^3$Eval: 認知に基づいたビデオタスクによるマルチモーダル記憶評価

マルチモーダル モデルが長時間ビデオの理解に向けて進歩するにつれ、メモリが重要な能力として浮上します。ビデオ データセットとベンチマークの開発には多大な努力が払われているにもかかわらず、既存の研究は主に知覚と推論に焦点を当てており、どのモデルが保持するか、情報がどの程度忠実に保存されるか、干渉下でもメモリがどの程度堅牢に保たれるかなど、記憶を体系的に評価することはありません。このギャップに対処するために、マルチモーダル モデルでさまざまなメモリ次元を調査するための最初の包括的な評価フレームワークおよびベンチマークである M$^3$Eval を導入します。認知心理学に基づいた当社のデザインは、記憶の重要な側面を分離する慎重に構築されたタスクを特徴としています。 M$^3$Eval を活用して、代表的なマルチモーダル モデルにわたって広範な実験を実施し、一貫した弱点と独特の動作を明らかにしました。私たちは、並列ビデオストリームを処理する際にモデルがもつれの解けた表現を維持するのに苦労し、人間の記憶で観察されるものとは大幅に異なる干渉パターンを示し、記憶ソースを時間領域よりも空間領域でより確実に接地し、限られた記号記憶を実証していることを発見しました。まとめると、私たちのベンチマークは将来の研究のための貴重なリソースを提供しますが、私たちの調査結果は、メモリが基本的でありながらまだ研究されていない機能であることを強調し、マルチモーダルモデルでより効果的なメモリメカニズムを設計するための洞察を提供します。コードとデータセットは https://pku-value-lab.github.io/m3eval-homepage で入手できます。

原文 (English)

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks

As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability. Despite substantial efforts in developing video datasets and benchmarks, existing works primarily focus on perception and reasoning, without systematically evaluating memory: what models retain, how faithfully information is preserved, and how robust memory remains under interference. To address this gap, we introduce M$^3$Eval, the first comprehensive evaluation framework and benchmark for probing different memory dimensions in multi-modal models. Grounded in cognitive psychology, our design features carefully constructed tasks that isolate key aspects of memory. Leveraging M$^3$Eval, we conduct extensive experiments across representative multi-modal models, revealing consistent weaknesses and distinctive behaviors. We find that models struggle to maintain disentangled representations when processing parallel video streams, exhibit interference patterns differing substantially from those observed in human memory, ground memory sources more reliably in the spatial domain than the temporal domain, and demonstrate limited symbolic memory. Collectively, our benchmark provides a valuable resource for future research, while our findings highlight memory as a fundamental yet underexplored capability and offer insights for designing more effective memory mechanisms in multi-modal models. Our code and dataset are available at https://pku-value-lab.github.io/m3eval-homepage.

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

答えから状態へ: 大規模言語モデルにおける化学推論の検証可能なプロセスレベルの評価

大規模な言語モデルが化学アシスタントとして使用されることが増えていますが、ほとんどの化学ベンチマークは依然として最終的な回答のみをスコアとしています。これにより、重大な故障モードが隠蔽されます。モデルは、その推論が化学ロジックに違反しているにもかかわらず、正しい分子、生成物、またはオプションを出力する可能性があります。 LLM ジャッジと人間のステップレベルのプロセス アノテーションはコストが高く、一貫性がなく、幻覚に対して脆弱であるため、既存のプロセス レベルの評価機能を拡張するのは困難です。 ChemCoTBench-V2 は、構造化され検証者がアドレス指定できる化学推論トレースを低コストで監査可能に評価するためのルール検証可能な診断ベンチマークです。これは、分子理解、分子編集、分子最適化、反応予測に及び、18 のレポートタスクにわたる 5,620 の評価サンプルを備えています。モデルは、専門家が設計したテンプレートで主要な中間ステップを公開する必要があり、それらのステップは決定論的な化学ルールでチェックされ、クローズドアンサータスクの場合は、別の LLM 審査員ではなく参照トレースが使用されます。オープンエンド分子最適化は、厳密なトレース マッチングではなく、Oracle で検証可能な状態制約を使用して評価されます。このベンチマークは、最終回答の正確性、テンプレートの遵守、専門家によって洗練された中間コミットメントに対する段階的な検証者の正確さという 3 つの個別のシグナルを報告します。フロンティア モデルの実験では、最終的な回答の成功と構造化推論の状態の一貫性の間には永続的なギャップがあることが明らかになりました。モデルは多くの場合、化学ステップ チェックに失敗しながらも要求された形式に従っているか、弱い裏付け推論で正しく回答することができます。 ChemCoTBench-V2 は、きめ細かいモデル比較を可能にし、トレースが最初に検証ツールに違反する具体的なステップを特定します。

原文 (English)

From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may output the correct molecule, product, or option while its reasoning violates chemical logic. Existing process-level evaluators are hard to scale because LLM judges and human step-level process annotation are costly, inconsistent, and vulnerable to hallucination. We introduce ChemCoTBench-V2, a rule-verifiable diagnostic benchmark for low-cost, auditable evaluation of structured, verifier-addressable chemical reasoning traces. It spans molecular understanding, molecule editing, molecular optimization, and reaction prediction, with 5,620 evaluation samples across 18 reporting tasks. Models must expose key intermediate steps in expert-designed templates, and those steps are checked with deterministic chemistry rules and, for closed-answer tasks, reference traces rather than another LLM judge. Open-ended molecular optimization is evaluated with oracle-verifiable state constraints rather than strict trace matching. The benchmark reports three separate signals: final-answer correctness, template adherence, and step-wise verifier correctness over expert-refined intermediate commitments. Experiments on frontier models reveal a persistent gap between final-answer success and structured-reasoning-state consistency: models often follow the requested format while failing chemical-step checks, or answer correctly with weak supporting reasoning. ChemCoTBench-V2 enables fine-grained model comparison and identifies the concrete step at which the trace first violates the verifier.

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

CounterFace: 顔認識システムのきめ細かい反事実評価のための合成顔データセット

顔認識 (FR) システムは重要なアプリケーションに広く導入されており、多様な人口や条件に対する信頼性と堅牢性が不可欠となっています。 FR システムの標準評価は通常、LFW などのデータセットに依存して平均認識精度を推定します。一部のベンチマークは、経年変化、姿勢、照明などの粗粒度のアイデンティティ内の変動も捕捉します。ただし、人間の顔には、ヘアスタイルやメイクなどの外観の変化を含む、より細かい変化が生じますが、これは既存のベンチマークでは過小評価されています。反事実評価は、このようなきめの細かい変動の下で FR の堅牢性を評価する方法を提供します。ただし、画像ジェネレーターを使用して合成された既存の反事実の顔データセットは、パイプラインでの検証に人間が使用されているため、属性の範囲が限られています。我々は、20 の顔属性と 8 つの人口統計的要素で構成される新しい反事実評価データセットである CounterFace を提案します。これは、以前の合成顔データセットを 14 属性と 2 つの人口統計的要因で上回っています。データセットは、カスタム検証機能を備えた既製の画像ジェネレーターに基づいた完全に自動化されたパイプラインを使用して生成され、人間による検証の必要性がなくなりました。 CounterFace には 11,821 の反事実の顔のペアが含まれており、事後のユーザー調査により、生成された反事実の忠実性が確認されています。 160 の属性と人口統計の組み合わせにわたって、2 つの商用 FR システムと 4 つのオープンソース FR システム (AWS Rekognition、Face++、AdaFace、MagFace、ArcFace、FaceNet) を評価します。当社のデータセットは、標準の評価ベンチマークとは異なり、個々のシステムの正確な故障モードを分離するのに役立ちます。結果は、パフォーマンスの低下は 6 つすべてのシステムの属性と人口統計によって異なり、遮蔽属性 (フェイスマスクやひげなど) が普遍的にパフォーマンスを低下させることを示しています。

原文 (English)

CounterFace: A Synthetic Face Dataset for Fine-Grained Counterfactual Evaluation of Face Recognition Systems

Face recognition (FR) systems are widely deployed in critical applications, making their reliability and robustness across diverse populations and conditions essential. Standard evaluation of FR systems typically relies on datasets such as LFW to estimate average recognition accuracy. Some benchmarks also capture coarse-grained intra-identity variations such as aging, pose, and lighting. However, human faces undergo more fine-grained changes, including appearance changes such as hairstyles and makeup, that are underrepresented in existing benchmarks. Counterfactual evaluation provides a method to assess FR robustness under such fine-grained variations. Existing counterfactual face datasets synthesized with image generators, however, are limited in attribute coverage due to the use of humans for verification in the pipeline. We propose CounterFace, a new counterfactual evaluation dataset comprising 20 facial attributes and 8 demographic factors, exceeding prior synthetic face datasets by 14 attributes and 2 demographics. The dataset is generated using a fully automated pipeline based on off-the-shelf image generators with custom verifiers, removing human need for verification. CounterFace contains 11,821 counterfactual face pairs, and a post-hoc user study confirms the faithfulness of the generated counterfactuals. We evaluate two commercial and four open-source FR systems (AWS Rekognition, Face++, AdaFace, MagFace, ArcFace, FaceNet) across 160 attribute-demographic combinations. Our dataset helps in the isolation of precise failure modes for individual systems unlike standard evaluation benchmarks. Results indicate that the performance degradation varies across attributes and demographics for all six systems and occluding attributes (e.g., facemask and facial hair) universally degrade performance.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

グラフ検索からスキーマ実現まで: 異種ナレッジ グラフ上のテキストから SPARQL への反事実検証

Text-to-SPARQL は、自然言語の質問を RDF ナレッジ グラフ上の実行可能な SPARQL クエリにマッピングします。標準的な評価ではターゲット グラフが事前に修正されることがよくありますが、実践的なナレッジ グラフ質問応答 (KGQA) には、異なるスキーマ、部分的なアラインメント、および不完全なメタデータを含む異種グラフ コレクションが含まれる場合があります。この設定では、クエリ生成は SPARQL 構文以上のものに依存します。システムは、質問に必要な述語、エンティティ タイプ、結合、フィルター、および制約をサポートできるグラフ スキーマを識別する必要があります。異種の KG コレクション上でテキストから SPARQL に変換するためのスキーマベースのエージェント フレームワークである SchemaForge を紹介します。その中心的なメカニズムは、質問条件付きのスキーマ スライス アライメントです。弱いグラフの証拠によって最初にもっともらしいグラフが特定され、より強力なスキーマの証拠によって、ローカル スキーマ スライスが意図したクエリを実現できるかどうかが決まります。選択されたスキーマ スライスは、クエリの生成と実行前の検証を制限します。利用可能なグラフが 1 つだけの場合、同じ定式化は、スキーマ基盤を備えた標準の単一 KG テキストから SPARQL への変換に縮小されます。 LC-QuAD 2.0、QALD-9 Plus、QALD-10、および Spider4SPARQL で SchemaForge を評価します。 SchemaForge は、4 つの公開ベンチマーク全体で、最も一致するエージェントのベースラインよりも実行精度を平均 11.50 パーセント向上させています。 Spider4SPARQL では、SchemaForge は実行精度を 54.86% から 64.18% に向上させ、トップ 1 グラフ割り当て精度 73.0% とトップ 3 グラフ割り当て精度 97.0% を達成しました。これらの結果は、グラフの弱い証拠からスキーマ固有のクエリコミットメントへの移行と、反事実の回答セットのチェックにより、異種ナレッジグラフよりも実行可能なクエリの生成が向上することを示しています。

原文 (English)

From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs

Text-to-SPARQL maps natural-language questions to executable SPARQL queries over RDF knowledge graphs. While standard evaluations often fix the target graph in advance, practical knowledge graph question answering (KGQA) may involve heterogeneous graph collections with different schemas, partial alignments, and incomplete metadata. In this setting, query generation depends on more than SPARQL syntax: the system must identify a graph schema that can support the predicates, entity types, joins, filters, and constraints required by the question. We present SchemaForge, a schema-grounded agentic framework for text-to-SPARQL over heterogeneous KG collections. Its central mechanism is question-conditioned schema-slice alignment: weak graph evidence first identifies plausible graphs, while stronger schema evidence determines whether a local schema slice can realize the intended query. The selected schema slice then constrains query generation and verification before execution. When only one graph is available, the same formulation reduces to standard single-KG text-to-SPARQL with schema grounding. We evaluate SchemaForge on LC-QuAD 2.0, QALD-9 Plus, QALD-10, and Spider4SPARQL. Across the four public benchmarks, SchemaForge improves execution accuracy over the strongest matched agent baseline by 11.50 percentage points on average. On Spider4SPARQL, SchemaForge improves execution accuracy from 54.86% to 64.18% and achieves 73.0% Top-1 and 97.0% Top-3 graph allocation accuracy. These results show that moving from weak graph evidence to schema-specific query commitments, together with counterfactual answer-set checks, improves executable query generation over heterogeneous knowledge graphs.

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

VGGSounder: 基礎モデルのオーディオビジュアル評価

視聴覚基礎モデルの出現は、マルチモーダルな理解を確実に評価することの重要性を強調しています。 VGGSound データセットは、オーディオビジュアル分類の評価のベンチマークとしてよく使用されます。ただし、私たちの分析では、不完全なラベル付け、部分的に重複するクラス、不整合なモダリティなど、VGGSound のいくつかの制限が特定されました。これらは、聴覚および視覚能力の歪んだ評価につながります。これらの制限に対処するために、VGGSounder を導入します。これは、VGGSound を拡張し、オーディオビジュアル基礎モデルを評価するために特別に設計された、包括的に再アノテーションが付けられたマルチラベル テスト セットです。 VGGSounder は詳細なモダリティの注釈を備えており、モダリティ固有のパフォーマンスを正確に分析できます。さらに、新しいモダリティ混乱メトリックを使用して別の入力モダリティを追加したときのパフォーマンスの低下を分析することで、モデルの限界を明らかにします。

原文 (English)

VGGSounder: Audio-Visual Evaluations for Foundation Models

The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ノイズを含む音声分離におけるスケール不変信号対歪み比の研究

この論文では、事実上のベンチマーク WSJ0-2Mix の場合のように、トレーニング参照にノイズが含まれている場合に、教師あり音声分離における評価とトレーニングの目的の両方としてスケール不変信号対歪み比 (SI-SDR) を使用することの意味を検証します。ノイズの多いリファレンスを使用して SI-SDR を導出すると、ノイズによって達成可能な SI-SDR が制限されるか、分離された出力に望ましくないノイズが発生することがわかります。これに対処するために、ノイズの多い参照の学習を回避するモデルをトレーニングすることを目的として、WHAM! を使用して参照を強化し、混合を増強する方法が提案されています。これらの強化されたデータセットでトレーニングされた 2 つのモデルは、非侵入的な NISQA.v2 メトリックを使用して評価されます。結果は、分離された音声のノイズが減少していることを示していますが、参照の処理によりアーチファクトが生じ、全体的な品質の向上が制限される可能性があることが示唆されています。 WSJ0-2Mix および Libri2Mix テスト セットのモデル全体で、SI-SDR と知覚されるノイズの間に負の相関関係が見つかり、導出による結論が強調されています。

原文 (English)

A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References

This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with the de facto benchmark WSJ0-2Mix. A derivation of the SI-SDR with noisy references reveals that noise limits the achievable SI-SDR, or leads to undesired noise in the separated outputs. To address this, a method is proposed to enhance references and augment the mixtures with WHAM!, aiming to train models that avoid learning noisy references. Two models trained on these enhanced datasets are evaluated with the non-intrusive NISQA.v2 metric. Results show reduced noise in separated speech but suggest that processing references may introduce artefacts, limiting overall quality gains. Negative correlation is found between SI-SDR and perceived noisiness across models on the WSJ0-2Mix and Libri2Mix test sets, underlining the conclusion from the derivation.

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達

分散ゲート分布を使用した不確かさの推定

ニューラル ネットワークからのサンプルごとの不確実性の定量化の評価は、高リスクのアプリケーションを含む意思決定に不可欠です。一般的なアプローチは、ベイジアン モデルまたは近似モデルからの予測分布を使用し、対応する予測の不確実性を認識的 (モデル関連) 成分と偶然的 (データ関連) 成分に分解することです。しかし、最近では相加的分解に疑問が持たれています。この研究では、さまざまなモデル予測にわたるクラス確率分布の信号対雑音比に基づいて、不確実性の推定と分解を行うための直感的なフレームワークを提案します。アンサンブルから導出された信頼係数によって予測をスケールする分散ゲート測定を導入します。私たちはこの尺度を利用して、委員会マシンの多様性の崩壊の存在について議論します。

原文 (English)

Uncertainty Estimation using Variance-Gated Distributions

Evaluation of per-sample uncertainty quantification from neural networks is essential for decision-making involving high-risk applications. A common approach is to use the predictive distribution from Bayesian or approximation models and decompose the corresponding predictive uncertainty into epistemic (model-related) and aleatoric (data-related) components. However, additive decomposition has recently been questioned. In this work, we propose an intuitive framework for uncertainty estimation and decomposition based on the signal-to-noise ratio of class probability distributions across different model predictions. We introduce a variance-gated measure that scales predictions by a confidence factor derived from ensembles. We use this measure to discuss the existence of a collapse in the diversity of committee machines.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark ne…

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達

Markov Chain Decoders Overcome the Heavy-Tail Limitations of Lipschitz Generative Models

Heavy-tailed distributions are prevalent in performance evaluation, network traffic, and risk modeling. This behavior poses a fundamental c…

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

メタ学習による費用対効果の高いモデル評価

機械学習の急速な成長により、拡大し続けるモデルのエコシステムが生み出され、目に見えないラベルのないデータに対して新しくリリースされたモデルの信頼性を検証することがますます困難になっています。従来の評価パイプラインは、高価なアノテーション、繰り返しの微調整、またはモデル ファミリ間での転送ができない狭い仮定に依存しています。さまざまなアーキテクチャやモダリティにまたがる未確認のモデルをラベルなしで迅速に評価するための、コスト効率が高く、モデルに依存しないフレームワークである MetaEvaluator を紹介します。 MetaEvaluator は、参照モデルのプールに対するメタ学習を利用して転送可能な初期化を取得し、プール全体でコストを償却しながら、モデルごとの再トレーニングの必要性を排除しながら、新しいモデルの正確な評価を可能にします。私たちの知る限り、これは完全にラベルのないデータセットで新しいモデルを評価できる、モデルに依存しない最初のフレームワークです。広範な実験により、MetaEvaluator は従来のアプローチと比較して大幅にコストを削減しながら安定した正確なパフォーマンス推定値を生成し、ラベルのないデータに対する新しいモデルのスケーラブルなベンチマークを実用化できることが示されています。

原文 (English)

Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-Learning

The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data. Existing evaluation pipelines typically rely on costly annotation, repeated fine-tuning, or assumptions that do not generalize well to new models. We introduce MetaEvaluator, a cost-effective, model-agnostic framework for fast, label-free evaluation of unseen models across diverse architectures and modalities. MetaEvaluator meta-learns over a pool of reference models to acquire an effective initialization for accurate assessment of unseen models, thereby amortizing evaluation cost and eliminating the need for per-model retraining. To the best of our knowledge, this is the first model-agnostic framework that evaluates new models on unlabeled datasets. Extensive experiments demonstrate that MetaEvaluator delivers stable and accurate performance estimates at substantially lower cost than conventional approaches, enabling scalable benchmarking on unlabeled datasets for emerging models. The code is available at: https://github.com/phkhanhtrinh23/MetaEvaluator.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

BenHalluEval: ベンガル語の大規模言語モデル用のマルチタスク幻覚評価フレームワーク

ベンガル語は世界で 6 番目に話されている言語であるにもかかわらず、ベンガル語の大規模言語モデル (LLM) で幻覚を体系的に評価した先行研究はありません。 BenHalluEval は、生成的質問応答 (GQA)、バングラ語と英語のコード混合 QA、要約、および推論の 4 つのタスクをカバーするベンガル語用のきめ細かい幻覚評価フレームワークです。既存の 3 つのベンガル語データセットから抽出された、12 のタスク固有の幻覚タイプにわたって GPT-5.4 を使用して 12,000 の幻覚候補を構築し、グラウンドトゥルース インスタンスの偽陽性率 (トラック A) と幻覚候補の幻覚検出率 (トラック B) を独立して測定するデュアル トラック プロトコルの下で、推論指向、多言語、ベンガル語中心のカテゴリにわたる 7 つの LLM を評価します。両方の故障モードに共同でペナルティを課し、均一な応答バイアスによるスコアのインフレを防ぐために、モデルとタスク全体で 7.72% から 55.42% の範囲のデュアルトラック キャリブレーション メトリクスである BenHalluScore を提案します。これは、幻覚キャリブレーションの大幅な変動を明らかにします。緩和戦略として適用される思考連鎖プロンプトは、幻覚差別を一貫して改善することなく、反応分布を変化させます。 BenHalluEval は、ベンガル語専用の幻覚ベンチマークを初めて確立し、リソースの少ない言語設定に対する単一トラックおよびプロンプトのみの評価アプローチが不適切であることを強調しています。データセットとコードは https://anonymous.4open.science/r/BanglaHalluEval-EB77 で入手できます。

原文 (English)

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates using GPT-5.4 across twelve task-specific hallucination types, drawn from three existing Bengali datasets, and evaluate seven LLMs spanning reasoning-oriented, multilingual, and Bengali-centric categories under a dual-track protocol that independently measures false-positive rate on ground-truth instances (Track A) and hallucination detection rate on hallucinated candidates (Track B). To jointly penalise both failure modes and prevent inflated scores from uniform response bias, we propose BenHalluScore, a dual-track calibration metric that ranges from 7.72% to 55.42% across models and tasks, revealing substantial variation in hallucination calibration. Chain-of-thought prompting, applied as a mitigation strategy, shifts response distributions without consistently improving hallucination discrimination. BenHalluEval establishes the first dedicated hallucination benchmark for Bengali and highlights the inadequacy of single-track and prompting-only evaluation approaches for low-resource language settings. The dataset and code are available at https://anonymous.4open.science/r/BanglaHalluEval-EB77.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

メタエージェントの課題: 現在のエージェントは自律的なエージェント開発が可能ですか?

現在の AI ベンチマークは、人間が設計したワークフロー内でのタスク実行に関してエージェントを評価します。これらの評価では、基本的に、モデルが自律的にエージェント システムを開発できるかどうかという、重要な次のレベルの機能を測定できません。自律エージェント開発のためのフロンティア モデルの能力をテストするために設計された評価フレームワークであるメタエージェント チャレンジ (MAC) を紹介します。具体的には、コード エージェント (メタエージェント) には、サンドボックス環境、評価 API、および 5 つのドメインにわたって実施されたテスト セットのパフォーマンスを最大化するエージェント アーティファクトを反復的にプログラムするための時間制限が与えられます。評価の整合性を確保するために、このフレームワークは報酬ハッキングに対する多層防御によって保護されています。このフレームワークを活用して、メタエージェントが人為的に設計されたベースライン ポリシーと一致することはほとんどなく、一致する少数のエージェントは独自のフロンティア モデルによって支配されていることを示します。さらに、設計プロセスは高い分散を示し、高い最適化圧力により、グラウンドトゥルースの漏洩などの敵対的な動作が表面化し、堅牢性とモデルの調整の両方における重大な欠陥が浮き彫りになります。最終的に、MAC は自律型 AI の研究開発のための厳密なオープンソース ベンチマークを提供し、再帰的な自己改善を評価するための経験的な代用手段を提供します。ベンチマークは https://github.com/ant-research/meta-agent-challenge で公開されています。

原文 (English)

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

Current AI benchmarks evaluate agents on task execution within human-designed workflows. These evaluations fundamentally fail to measure a critical next-level capability: whether models can autonomously develop agent systems. We introduce the Meta-Agent Challenge (MAC), an evaluation framework designed to test the capacity of frontier models for autonomous agent development. Specifically, a code agent (the meta-agent) is given a sandboxed environment, an evaluation API, and a time limitation to iteratively program an agent artifact that maximizes performance on a held-out test set across five domains. To ensure evaluation integrity, this framework is secured by multi-layer defenses against reward hacking. Leveraging this framework, we demonstrate that meta-agents rarely match human-engineered baseline policies, and the few that do are dominated by proprietary frontier models. Moreover, the design process exhibits high variance, and high optimization pressure surfaces emergent adversarial behaviors like ground-truth exfiltration-highlighting critical deficits in both robustness and model alignment. Ultimately, MAC provides a rigorous, open-source benchmark for autonomous AI research and development, offering an empirical proxy for evaluating recursive self-improvement. Benchmark is publicly available at: https://github.com/ant-research/meta-agent-challenge.

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

SpurAudio: 少数ショット音声分類におけるショートカット学習を研究するためのベンチマーク

少数ショット分類 (FSC) は、限られたラベル付きデータから学習するために広く使用されていますが、ほとんどの評価は、ターゲットの概念が文脈上の手がかりから独立していることを暗黙的に前提としています。ただし、現実世界の設定では、サンプルがリッチ コンテキスト内に表示されることが多く、モデルが前景コンテンツと背景信号の間の偽の相関を利用できるようになります。このような効果は少数ショット画像分類で研究されていますが、少数ショット音声分類におけるその役割はほとんど解明されておらず、既存の音声ベンチマークでは文脈構造に対する制御が限られています。 SpurAudio というベンチマークを紹介します。これは、オーディオの前景イベントと背景環境の自然な分離性を活用して、サポートおよびクエリ セットにわたるコンテキストの変化を制御されたマルチレベルの評価を可能にするベンチマークです。このベンチマークを使用して、多くの最先端の少数ショット手法は、標準的な評価プロトコルで同様の精度を達成しているにもかかわらず、バックグラウンド相関が破壊されると重大なパフォーマンス低下に見舞われることがわかります。重要なのは、この脆弱性は大規模な事前トレーニング済みオーディオ基盤モデルでも存続しており、バックボーン容量の制限が説明の対象外となっているということです。さらに、従来のベンチマークでは同等に見える手法でも、偽の相関に対して著しく異なる感度を示す可能性があり、推論時に特徴表現が分類器ヘッドとどのように相互作用するかに関連する体系的なアルゴリズムの強みと脆弱性が明らかになります。これらの発見は、オーディオにおける少数ショット法の動作に関する新たな洞察を提供し、FSC モデルを評価する際のコンテキスト依存性を明示的に調査するベンチマークの必要性を強調しています。

原文 (English)

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues. In real-world settings, however, examples often appear within rich contexts, allowing models to exploit spurious correlations between foreground content and background signals. While such effects have been studied in few-shot image classification, their role in few-shot audio classification remains largely unexplored, and existing audio benchmarks offer limited control over contextual structure. We introduce SpurAudio, a benchmark that leverages the natural separability of foreground events and background environments in audio to enable controlled, multi-level evaluation of contextual shifts across support and query sets. Using this benchmark, we show that many state-of-the-art few-shot methods suffer severe performance degradation when background correlations are disrupted, despite achieving similar accuracy under standard evaluation protocols. Crucially, this vulnerability persists even in large pretrained audio foundation models, ruling out limited backbone capacity as an explanation. Moreover, methods that appear comparable under conventional benchmarks can exhibit markedly different sensitivity to spurious correlations, revealing systematic algorithmic strengths and vulnerabilities tied to how feature representations interact with classifier heads at inference time. These findings provide new insight into the behavior of few-shot methods in audio and highlight the need for benchmarks that explicitly probe context dependence when evaluating FSC models.

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

L-TGVN: パーソナライズされた高速 MRI のための縦方向事前分布の活用

MRI は電離放射線を使用せずに優れた軟組織コントラストを提供しますが、取得時間が長いため患者の不快感が増大すると同時に、検査コストが上昇し、スキャナのスループットが制限されます。スキャン時間を短縮するための一般的なアプローチは、取得する測定値を少なくすることです。これにより、不適切な線形逆問題が発生します。したがって、診断品質の画像を回復するには、測定データ以外の事前知識を組み込む必要があります。追跡検査では、患者の最新の以前のスキャンにより、非常に有益な被験者固有のコンテキストが提供されますが、実際の使用は、時間的変化(病状の進行を含む)、スキャン間のずれ、取得間のプロトコルのドリフトによって複雑になります。この研究では、大幅にアンダーサンプリングされた測定値から現在のスキャンを再構築するための副次情報として以前のスキャンを活用する、縦方向の信頼誘導変分ネットワークである L-TGVN を紹介します。重要なことは、L-TGVN は、以前のスキャンの影響が取得された測定値と一致するように制限することです。既存の多くの縦方向再構成方法とは異なり、以前のスキャンと現在のスキャンの間の明示的な事前位置合わせを必要としません。さらに、訪問ごとの取得プロトコルの違い(シーケンスパラメータの変更など)にも対応します。私たちは、事前ガイド法や縦方向事前分布を使用しない方法など、一致した容量のベースラインに対して L-TGVN を評価し、困難な加速において微細構造のより良好な保存とともに、標準的な定量的指標の一貫した改善を観察しました。ソース コードは github.com/sodicksonlab/L-TGVN で入手できます。

原文 (English)

L-TGVN: Leveraging Longitudinal Priors for Personalized Rapid MRI

MRI provides excellent soft-tissue contrast without ionizing radiation, but long acquisition times increase patient discomfort while also raising exam costs and limiting scanner throughput. A common approach to reduce scan time is to acquire fewer measurements, which yields an ill-posed linear inverse problem; recovering diagnostic-quality images therefore requires incorporating prior knowledge beyond the measured data. In follow-up exams, the most recent prior scan of a patient can provide a highly informative subject-specific context, but practical use is complicated by temporal changes (including pathology progression), misalignment between scans, and protocol drift across acquisitions. In this work, we introduce L-TGVN, a Longitudinal Trust-Guided Variational Network that leverages prior scans as side information to reconstruct the current scan from heavily undersampled measurements. Crucially, L-TGVN constrains the influence of prior scans to be consistent with the acquired measurements. Unlike many existing longitudinal reconstruction methods, it does not require explicit pre-registration between prior and current scans. It further accommodates differences in acquisition protocols across visits (e.g., changes in sequence parameters). We evaluate L-TGVN against matched-capacity baselines, including prior-guided methods and methods that do not use longitudinal priors, and observe consistent improvements in standard quantitative metrics together with better preservation of fine structures at challenging accelerations. Source code is available at github.com/sodicksonlab/L-TGVN.

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達

RowNet: 表形式回帰のためのメモリ トランスフォーマー

不動産評価は構造化回帰問題であり、価格は異種の特徴タイプ、まばらな地域効果、非線形相互作用、および比較可能な不動産の実際的なロジックによって支配されます。標準的な多層パーセプトロンは各行を孤立ベクトルとして扱い、局所性、スケール感度、およびカテゴリカルマッチングを監視のみから学習する必要があります。勾配ブースト デシジョン ツリーは強力な表形式のベースラインを提供しますが、その特徴中心の分割メカニズムは、類似した履歴観測の取得を明示的にモデル化しません。この論文では、不動産の平方メートルあたりの価格を予測するための検索ベースのニューラル アーキテクチャである RowNet について説明します。 RowNet は、ラベル付きプロパティのメモリ バンクに対するペアごとの類似性機能を通じてクエリ プロパティを表します。最初の検索層は、特徴のみの類似性から大まかなターゲットを推定します。 2 番目の層は、ターゲット一貫性機能を使用してメモリ比較を強化し、複数の学習されたアテンション ヘッドを使用して相補的な比較可能なセットを取得します。最後の専門家混合モジュールは、学習されたゲーティング、残差補正、エントロピー正則化、ヘッドダイバーシティ正則化を組み合わせて予測を生成します。

原文 (English)

RowNet: A Memory Transformer for Tabular Regression

Real estate valuation is a structured regression problem in which prices are governed by heterogeneous feature types, sparse regional effects, nonlinear interactions, and the practical logic of comparable properties. Standard multilayer perceptrons treat each row as an isolated vector and must learn locality, scale sensitivity, and categorical matching from supervision alone. Gradient-boosted decision trees provide strong tabular baselines, but their feature-centric splitting mechanism does not explicitly model the retrieval of similar historical observations. This paper presents RowNet, a retrieval-based neural architecture for real estate price-per-square-meter prediction. RowNet represents a query property through pairwise similarity features against a memory bank of labeled properties. A first retrieval layer estimates a coarse target from feature-only similarities. A second layer augments the memory comparison with target-consistency features and uses multiple learned attention heads to retrieve complementary comparable sets. A final mixture-of-experts module combines learned gating, residual correction, entropy regularization, and head-diversity regularization to produce the prediction.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

共同生成と評価による自己進化する深層研究

大規模言語モデル (LLM) は日常のアプリケーションでますます採用されるようになり、詳細な研究が特に重要な機能として際立っています。従来の質問応答 (QA) タスクとは異なり、詳細な調査レポートの生成には決定的な根拠が欠けているため、報酬設計が本質的に検証不可能になり、効果的な強化学習が制限されます。既存のアプローチでは、LLM-as-a-judge およびクエリ依存の評価ルーブリックを使用してこの課題を軽減していますが、依然として静的な評価器に依存しているため、ソルバーの向上に応じて標準を適応させることができず、最適化圧力が不十分になり、最終的に飽和状態になってしまいます。私たちは、\textbf{s}elf 進化型 \textbf{co} 進化型トレーニング フレームワークで、深い \textbf{re} 検索の評価と生成 (SCORE) を使用してこの制限に対処します。これは、共有パラメータ学習プロセスにおいて評価器とソルバーを緊密に結合します。生成と評価を独立したモジュールとして扱うのではなく、それらの本質的なつながりを活用して、単一の共有パラメーター モデル内で共同の改善を可能にします。このプロセスを制限するために、ソルバーのパフォーマンスに基づいて評価環境を動的に制御するメタハーネスを導入し、有効な評価次元と十分に深い評価者の検索を促進します。ディープリサーチベンチマークに関する広範な実験により、レポート生成の品質が一貫して向上していることが実証されており、評価と生成を共進化させることが、オープンエンドのリサーチエージェントをトレーニングするための有望な方向性であることが示されています。

原文 (English)

Self-Evolving Deep Research via Joint Generation and Evaluation

Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional question-answering (QA) tasks, deep research report generation lacks definitive ground-truth, making reward design inherently unverifiable and limiting effective reinforcement learning. Existing approaches mitigate this challenge with LLM-as-a-judge and query-dependent evaluation rubrics, but they still rely on static evaluators that cannot adapt their standards as the solver improves, leading to insufficient and eventually saturated optimization pressure. We address this limitation with a \textbf{s}elf-evolving \textbf{co}-evolutionary training framework for deep \textbf{re}search evaluation and generation (SCORE), which tightly couples an evaluator and a solver in a shared-parameter learning process. Rather than treating generation and evaluation as isolated modules, we leverage their intrinsic connection to enable joint improvement within a single shared-parameter model. To restrict this process, we introduce a meta-harness, which dynamically controls the evaluation environment based on solver performance, encouraging valid evaluation dimensions and sufficiently deep evaluator search. Extensive experiments on deep research benchmarks demonstrate consistent improvement in report generation quality, showing that co-evolving evaluation and generation is a promising direction for training open-ended research agents.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

M$^3$Eval: 認知に基づいたビデオタスクによるマルチモーダル記憶評価

マルチモーダル モデルが長時間ビデオの理解に向けて進歩するにつれ、メモリが重要な能力として浮上します。ビデオ データセットとベンチマークの開発には多大な努力が払われているにもかかわらず、既存の研究は主に知覚と推論に焦点を当てており、どのモデルが保持するか、情報がどの程度忠実に保存されるか、干渉下でもメモリがどの程度堅牢に保たれるかなど、記憶を体系的に評価することはありません。このギャップに対処するために、マルチモーダル モデルでさまざまなメモリ次元を調査するための最初の包括的な評価フレームワークおよびベンチマークである M$^3$Eval を導入します。認知心理学に基づいた当社のデザインは、記憶の重要な側面を分離する慎重に構築されたタスクを特徴としています。 M$^3$Eval を活用して、代表的なマルチモーダル モデルにわたって広範な実験を実施し、一貫した弱点と独特の動作を明らかにしました。私たちは、並列ビデオストリームを処理する際にモデルがもつれの解けた表現を維持するのに苦労し、人間の記憶で観察されるものとは大幅に異なる干渉パターンを示し、記憶ソースを時間領域よりも空間領域でより確実に接地し、限られた記号記憶を実証していることを発見しました。まとめると、私たちのベンチマークは将来の研究のための貴重なリソースを提供しますが、私たちの調査結果は、メモリが基本的でありながらまだ研究されていない機能であることを強調し、マルチモーダルモデルでより効果的なメモリメカニズムを設計するための洞察を提供します。コードとデータセットは https://pku-value-lab.github.io/m3eval-homepage で入手できます。

原文 (English)

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks

As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability. Despite substantial efforts in developing video datasets and benchmarks, existing works primarily focus on perception and reasoning, without systematically evaluating memory: what models retain, how faithfully information is preserved, and how robust memory remains under interference. To address this gap, we introduce M$^3$Eval, the first comprehensive evaluation framework and benchmark for probing different memory dimensions in multi-modal models. Grounded in cognitive psychology, our design features carefully constructed tasks that isolate key aspects of memory. Leveraging M$^3$Eval, we conduct extensive experiments across representative multi-modal models, revealing consistent weaknesses and distinctive behaviors. We find that models struggle to maintain disentangled representations when processing parallel video streams, exhibit interference patterns differing substantially from those observed in human memory, ground memory sources more reliably in the spatial domain than the temporal domain, and demonstrate limited symbolic memory. Collectively, our benchmark provides a valuable resource for future research, while our findings highlight memory as a fundamental yet underexplored capability and offer insights for designing more effective memory mechanisms in multi-modal models. Our code and dataset are available at https://pku-value-lab.github.io/m3eval-homepage.

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

答えから状態へ: 大規模言語モデルにおける化学推論の検証可能なプロセスレベルの評価

大規模な言語モデルが化学アシスタントとして使用されることが増えていますが、ほとんどの化学ベンチマークは依然として最終的な回答のみをスコアとしています。これにより、重大な故障モードが隠蔽されます。モデルは、その推論が化学ロジックに違反しているにもかかわらず、正しい分子、生成物、またはオプションを出力する可能性があります。 LLM ジャッジと人間のステップレベルのプロセス アノテーションはコストが高く、一貫性がなく、幻覚に対して脆弱であるため、既存のプロセス レベルの評価機能を拡張するのは困難です。 ChemCoTBench-V2 は、構造化され検証者がアドレス指定できる化学推論トレースを低コストで監査可能に評価するためのルール検証可能な診断ベンチマークです。これは、分子理解、分子編集、分子最適化、反応予測に及び、18 のレポートタスクにわたる 5,620 の評価サンプルを備えています。モデルは、専門家が設計したテンプレートで主要な中間ステップを公開する必要があり、それらのステップは決定論的な化学ルールでチェックされ、クローズドアンサータスクの場合は、別の LLM 審査員ではなく参照トレースが使用されます。オープンエンド分子最適化は、厳密なトレース マッチングではなく、Oracle で検証可能な状態制約を使用して評価されます。このベンチマークは、最終回答の正確性、テンプレートの遵守、専門家によって洗練された中間コミットメントに対する段階的な検証者の正確さという 3 つの個別のシグナルを報告します。フロンティア モデルの実験では、最終的な回答の成功と構造化推論の状態の一貫性の間には永続的なギャップがあることが明らかになりました。モデルは多くの場合、化学ステップ チェックに失敗しながらも要求された形式に従っているか、弱い裏付け推論で正しく回答することができます。 ChemCoTBench-V2 は、きめ細かいモデル比較を可能にし、トレースが最初に検証ツールに違反する具体的なステップを特定します。

原文 (English)

From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may output the correct molecule, product, or option while its reasoning violates chemical logic. Existing process-level evaluators are hard to scale because LLM judges and human step-level process annotation are costly, inconsistent, and vulnerable to hallucination. We introduce ChemCoTBench-V2, a rule-verifiable diagnostic benchmark for low-cost, auditable evaluation of structured, verifier-addressable chemical reasoning traces. It spans molecular understanding, molecule editing, molecular optimization, and reaction prediction, with 5,620 evaluation samples across 18 reporting tasks. Models must expose key intermediate steps in expert-designed templates, and those steps are checked with deterministic chemistry rules and, for closed-answer tasks, reference traces rather than another LLM judge. Open-ended molecular optimization is evaluated with oracle-verifiable state constraints rather than strict trace matching. The benchmark reports three separate signals: final-answer correctness, template adherence, and step-wise verifier correctness over expert-refined intermediate commitments. Experiments on frontier models reveal a persistent gap between final-answer success and structured-reasoning-state consistency: models often follow the requested format while failing chemical-step checks, or answer correctly with weak supporting reasoning. ChemCoTBench-V2 enables fine-grained model comparison and identifies the concrete step at which the trace first violates the verifier.

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

CounterFace: 顔認識システムのきめ細かい反事実評価のための合成顔データセット

顔認識 (FR) システムは重要なアプリケーションに広く導入されており、多様な人口や条件に対する信頼性と堅牢性が不可欠となっています。 FR システムの標準評価は通常、LFW などのデータセットに依存して平均認識精度を推定します。一部のベンチマークは、経年変化、姿勢、照明などの粗粒度のアイデンティティ内の変動も捕捉します。ただし、人間の顔には、ヘアスタイルやメイクなどの外観の変化を含む、より細かい変化が生じますが、これは既存のベンチマークでは過小評価されています。反事実評価は、このようなきめの細かい変動の下で FR の堅牢性を評価する方法を提供します。ただし、画像ジェネレーターを使用して合成された既存の反事実の顔データセットは、パイプラインでの検証に人間が使用されているため、属性の範囲が限られています。我々は、20 の顔属性と 8 つの人口統計的要素で構成される新しい反事実評価データセットである CounterFace を提案します。これは、以前の合成顔データセットを 14 属性と 2 つの人口統計的要因で上回っています。データセットは、カスタム検証機能を備えた既製の画像ジェネレーターに基づいた完全に自動化されたパイプラインを使用して生成され、人間による検証の必要性がなくなりました。 CounterFace には 11,821 の反事実の顔のペアが含まれており、事後のユーザー調査により、生成された反事実の忠実性が確認されています。 160 の属性と人口統計の組み合わせにわたって、2 つの商用 FR システムと 4 つのオープンソース FR システム (AWS Rekognition、Face++、AdaFace、MagFace、ArcFace、FaceNet) を評価します。当社のデータセットは、標準の評価ベンチマークとは異なり、個々のシステムの正確な故障モードを分離するのに役立ちます。結果は、パフォーマンスの低下は 6 つすべてのシステムの属性と人口統計によって異なり、遮蔽属性 (フェイスマスクやひげなど) が普遍的にパフォーマンスを低下させることを示しています。

原文 (English)

CounterFace: A Synthetic Face Dataset for Fine-Grained Counterfactual Evaluation of Face Recognition Systems

Face recognition (FR) systems are widely deployed in critical applications, making their reliability and robustness across diverse populations and conditions essential. Standard evaluation of FR systems typically relies on datasets such as LFW to estimate average recognition accuracy. Some benchmarks also capture coarse-grained intra-identity variations such as aging, pose, and lighting. However, human faces undergo more fine-grained changes, including appearance changes such as hairstyles and makeup, that are underrepresented in existing benchmarks. Counterfactual evaluation provides a method to assess FR robustness under such fine-grained variations. Existing counterfactual face datasets synthesized with image generators, however, are limited in attribute coverage due to the use of humans for verification in the pipeline. We propose CounterFace, a new counterfactual evaluation dataset comprising 20 facial attributes and 8 demographic factors, exceeding prior synthetic face datasets by 14 attributes and 2 demographics. The dataset is generated using a fully automated pipeline based on off-the-shelf image generators with custom verifiers, removing human need for verification. CounterFace contains 11,821 counterfactual face pairs, and a post-hoc user study confirms the faithfulness of the generated counterfactuals. We evaluate two commercial and four open-source FR systems (AWS Rekognition, Face++, AdaFace, MagFace, ArcFace, FaceNet) across 160 attribute-demographic combinations. Our dataset helps in the isolation of precise failure modes for individual systems unlike standard evaluation benchmarks. Results indicate that the performance degradation varies across attributes and demographics for all six systems and occluding attributes (e.g., facemask and facial hair) universally degrade performance.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

グラフ検索からスキーマ実現まで: 異種ナレッジ グラフ上のテキストから SPARQL への反事実検証

Text-to-SPARQL は、自然言語の質問を RDF ナレッジ グラフ上の実行可能な SPARQL クエリにマッピングします。標準的な評価ではターゲット グラフが事前に修正されることがよくありますが、実践的なナレッジ グラフ質問応答 (KGQA) には、異なるスキーマ、部分的なアラインメント、および不完全なメタデータを含む異種グラフ コレクションが含まれる場合があります。この設定では、クエリ生成は SPARQL 構文以上のものに依存します。システムは、質問に必要な述語、エンティティ タイプ、結合、フィルター、および制約をサポートできるグラフ スキーマを識別する必要があります。異種の KG コレクション上でテキストから SPARQL に変換するためのスキーマベースのエージェント フレームワークである SchemaForge を紹介します。その中心的なメカニズムは、質問条件付きのスキーマ スライス アライメントです。弱いグラフの証拠によって最初にもっともらしいグラフが特定され、より強力なスキーマの証拠によって、ローカル スキーマ スライスが意図したクエリを実現できるかどうかが決まります。選択されたスキーマ スライスは、クエリの生成と実行前の検証を制限します。利用可能なグラフが 1 つだけの場合、同じ定式化は、スキーマ基盤を備えた標準の単一 KG テキストから SPARQL への変換に縮小されます。 LC-QuAD 2.0、QALD-9 Plus、QALD-10、および Spider4SPARQL で SchemaForge を評価します。 SchemaForge は、4 つの公開ベンチマーク全体で、最も一致するエージェントのベースラインよりも実行精度を平均 11.50 パーセント向上させています。 Spider4SPARQL では、SchemaForge は実行精度を 54.86% から 64.18% に向上させ、トップ 1 グラフ割り当て精度 73.0% とトップ 3 グラフ割り当て精度 97.0% を達成しました。これらの結果は、グラフの弱い証拠からスキーマ固有のクエリコミットメントへの移行と、反事実の回答セットのチェックにより、異種ナレッジグラフよりも実行可能なクエリの生成が向上することを示しています。

原文 (English)

From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs

Text-to-SPARQL maps natural-language questions to executable SPARQL queries over RDF knowledge graphs. While standard evaluations often fix the target graph in advance, practical knowledge graph question answering (KGQA) may involve heterogeneous graph collections with different schemas, partial alignments, and incomplete metadata. In this setting, query generation depends on more than SPARQL syntax: the system must identify a graph schema that can support the predicates, entity types, joins, filters, and constraints required by the question. We present SchemaForge, a schema-grounded agentic framework for text-to-SPARQL over heterogeneous KG collections. Its central mechanism is question-conditioned schema-slice alignment: weak graph evidence first identifies plausible graphs, while stronger schema evidence determines whether a local schema slice can realize the intended query. The selected schema slice then constrains query generation and verification before execution. When only one graph is available, the same formulation reduces to standard single-KG text-to-SPARQL with schema grounding. We evaluate SchemaForge on LC-QuAD 2.0, QALD-9 Plus, QALD-10, and Spider4SPARQL. Across the four public benchmarks, SchemaForge improves execution accuracy over the strongest matched agent baseline by 11.50 percentage points on average. On Spider4SPARQL, SchemaForge improves execution accuracy from 54.86% to 64.18% and achieves 73.0% Top-1 and 97.0% Top-3 graph allocation accuracy. These results show that moving from weak graph evidence to schema-specific query commitments, together with counterfactual answer-set checks, improves executable query generation over heterogeneous knowledge graphs.

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

VGGSounder: 基礎モデルのオーディオビジュアル評価

視聴覚基礎モデルの出現は、マルチモーダルな理解を確実に評価することの重要性を強調しています。 VGGSound データセットは、オーディオビジュアル分類の評価のベンチマークとしてよく使用されます。ただし、私たちの分析では、不完全なラベル付け、部分的に重複するクラス、不整合なモダリティなど、VGGSound のいくつかの制限が特定されました。これらは、聴覚および視覚能力の歪んだ評価につながります。これらの制限に対処するために、VGGSounder を導入します。これは、VGGSound を拡張し、オーディオビジュアル基礎モデルを評価するために特別に設計された、包括的に再アノテーションが付けられたマルチラベル テスト セットです。 VGGSounder は詳細なモダリティの注釈を備えており、モダリティ固有のパフォーマンスを正確に分析できます。さらに、新しいモダリティ混乱メトリックを使用して別の入力モダリティを追加したときのパフォーマンスの低下を分析することで、モデルの限界を明らかにします。

原文 (English)

VGGSounder: Audio-Visual Evaluations for Foundation Models

The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ノイズを含む音声分離におけるスケール不変信号対歪み比の研究

この論文では、事実上のベンチマーク WSJ0-2Mix の場合のように、トレーニング参照にノイズが含まれている場合に、教師あり音声分離における評価とトレーニングの目的の両方としてスケール不変信号対歪み比 (SI-SDR) を使用することの意味を検証します。ノイズの多いリファレンスを使用して SI-SDR を導出すると、ノイズによって達成可能な SI-SDR が制限されるか、分離された出力に望ましくないノイズが発生することがわかります。これに対処するために、ノイズの多い参照の学習を回避するモデルをトレーニングすることを目的として、WHAM! を使用して参照を強化し、混合を増強する方法が提案されています。これらの強化されたデータセットでトレーニングされた 2 つのモデルは、非侵入的な NISQA.v2 メトリックを使用して評価されます。結果は、分離された音声のノイズが減少していることを示していますが、参照の処理によりアーチファクトが生じ、全体的な品質の向上が制限される可能性があることが示唆されています。 WSJ0-2Mix および Libri2Mix テスト セットのモデル全体で、SI-SDR と知覚されるノイズの間に負の相関関係が見つかり、導出による結論が強調されています。

原文 (English)

A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References

This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with the de facto benchmark WSJ0-2Mix. A derivation of the SI-SDR with noisy references reveals that noise limits the achievable SI-SDR, or leads to undesired noise in the separated outputs. To address this, a method is proposed to enhance references and augment the mixtures with WHAM!, aiming to train models that avoid learning noisy references. Two models trained on these enhanced datasets are evaluated with the non-intrusive NISQA.v2 metric. Results show reduced noise in separated speech but suggest that processing references may introduce artefacts, limiting overall quality gains. Negative correlation is found between SI-SDR and perceived noisiness across models on the WSJ0-2Mix and Libri2Mix test sets, underlining the conclusion from the derivation.

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達

分散ゲート分布を使用した不確かさの推定

ニューラル ネットワークからのサンプルごとの不確実性の定量化の評価は、高リスクのアプリケーションを含む意思決定に不可欠です。一般的なアプローチは、ベイジアン モデルまたは近似モデルからの予測分布を使用し、対応する予測の不確実性を認識的 (モデル関連) 成分と偶然的 (データ関連) 成分に分解することです。しかし、最近では相加的分解に疑問が持たれています。この研究では、さまざまなモデル予測にわたるクラス確率分布の信号対雑音比に基づいて、不確実性の推定と分解を行うための直感的なフレームワークを提案します。アンサンブルから導出された信頼係数によって予測をスケールする分散ゲート測定を導入します。私たちはこの尺度を利用して、委員会マシンの多様性の崩壊の存在について議論します。

原文 (English)

Uncertainty Estimation using Variance-Gated Distributions

Evaluation of per-sample uncertainty quantification from neural networks is essential for decision-making involving high-risk applications. A common approach is to use the predictive distribution from Bayesian or approximation models and decompose the corresponding predictive uncertainty into epistemic (model-related) and aleatoric (data-related) components. However, additive decomposition has recently been questioned. In this work, we propose an intuitive framework for uncertainty estimation and decomposition based on the signal-to-noise ratio of class probability distributions across different model predictions. We introduce a variance-gated measure that scales predictions by a confidence factor derived from ensembles. We use this measure to discuss the existence of a collapse in the diversity of committee machines.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark ne…

2026-06-05 13:00 JSTarXiv cs.AIビジネス/資金調達

Markov Chain Decoders Overcome the Heavy-Tail Limitations of Lipschitz Generative Models

Heavy-tailed distributions are prevalent in performance evaluation, network traffic, and risk modeling. This behavior poses a fundamental c…

2026-06-05 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

メタ学習による費用対効果の高いモデル評価

機械学習の急速な成長により、拡大し続けるモデルのエコシステムが生み出され、目に見えないラベルのないデータに対して新しくリリースされたモデルの信頼性を検証することがますます困難になっています。従来の評価パイプラインは、高価なアノテーション、繰り返しの微調整、またはモデル ファミリ間での転送ができない狭い仮定に依存しています。さまざまなアーキテクチャやモダリティにまたがる未確認のモデルをラベルなしで迅速に評価するための、コスト効率が高く、モデルに依存しないフレームワークである MetaEvaluator を紹介します。 MetaEvaluator は、参照モデルのプールに対するメタ学習を利用して転送可能な初期化を取得し、プール全体でコストを償却しながら、モデルごとの再トレーニングの必要性を排除しながら、新しいモデルの正確な評価を可能にします。私たちの知る限り、これは完全にラベルのないデータセットで新しいモデルを評価できる、モデルに依存しない最初のフレームワークです。広範な実験により、MetaEvaluator は従来のアプローチと比較して大幅にコストを削減しながら安定した正確なパフォーマンス推定値を生成し、ラベルのないデータに対する新しいモデルのスケーラブルなベンチマークを実用化できることが示されています。

原文 (English)

Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-Learning

The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data. Existing evaluation pipelines typically rely on costly annotation, repeated fine-tuning, or assumptions that do not generalize well to new models. We introduce MetaEvaluator, a cost-effective, model-agnostic framework for fast, label-free evaluation of unseen models across diverse architectures and modalities. MetaEvaluator meta-learns over a pool of reference models to acquire an effective initialization for accurate assessment of unseen models, thereby amortizing evaluation cost and eliminating the need for per-model retraining. To the best of our knowledge, this is the first model-agnostic framework that evaluates new models on unlabeled datasets. Extensive experiments demonstrate that MetaEvaluator delivers stable and accurate performance estimates at substantially lower cost than conventional approaches, enabling scalable benchmarking on unlabeled datasets for emerging models. The code is available at: https://github.com/phkhanhtrinh23/MetaEvaluator.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

BenHalluEval: ベンガル語の大規模言語モデル用のマルチタスク幻覚評価フレームワーク

ベンガル語は世界で 6 番目に話されている言語であるにもかかわらず、ベンガル語の大規模言語モデル (LLM) で幻覚を体系的に評価した先行研究はありません。 BenHalluEval は、生成的質問応答 (GQA)、バングラ語と英語のコード混合 QA、要約、および推論の 4 つのタスクをカバーするベンガル語用のきめ細かい幻覚評価フレームワークです。既存の 3 つのベンガル語データセットから抽出された、12 のタスク固有の幻覚タイプにわたって GPT-5.4 を使用して 12,000 の幻覚候補を構築し、グラウンドトゥルース インスタンスの偽陽性率 (トラック A) と幻覚候補の幻覚検出率 (トラック B) を独立して測定するデュアル トラック プロトコルの下で、推論指向、多言語、ベンガル語中心のカテゴリにわたる 7 つの LLM を評価します。両方の故障モードに共同でペナルティを課し、均一な応答バイアスによるスコアのインフレを防ぐために、モデルとタスク全体で 7.72% から 55.42% の範囲のデュアルトラック キャリブレーション メトリクスである BenHalluScore を提案します。これは、幻覚キャリブレーションの大幅な変動を明らかにします。緩和戦略として適用される思考連鎖プロンプトは、幻覚差別を一貫して改善することなく、反応分布を変化させます。 BenHalluEval は、ベンガル語専用の幻覚ベンチマークを初めて確立し、リソースの少ない言語設定に対する単一トラックおよびプロンプトのみの評価アプローチが不適切であることを強調しています。データセットとコードは https://anonymous.4open.science/r/BanglaHalluEval-EB77 で入手できます。

原文 (English)

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates using GPT-5.4 across twelve task-specific hallucination types, drawn from three existing Bengali datasets, and evaluate seven LLMs spanning reasoning-oriented, multilingual, and Bengali-centric categories under a dual-track protocol that independently measures false-positive rate on ground-truth instances (Track A) and hallucination detection rate on hallucinated candidates (Track B). To jointly penalise both failure modes and prevent inflated scores from uniform response bias, we propose BenHalluScore, a dual-track calibration metric that ranges from 7.72% to 55.42% across models and tasks, revealing substantial variation in hallucination calibration. Chain-of-thought prompting, applied as a mitigation strategy, shifts response distributions without consistently improving hallucination discrimination. BenHalluEval establishes the first dedicated hallucination benchmark for Bengali and highlights the inadequacy of single-track and prompting-only evaluation approaches for low-resource language settings. The dataset and code are available at https://anonymous.4open.science/r/BanglaHalluEval-EB77.

2026-06-05 07:43 JSTTechCrunch AILLM/生成AIビジネス/資金調達

Ahead of its IPO, Anthropic’s Daniela Amodei shrugs off doubts about AI’s returns

Anthropic has been growing at a breakneck pace. The company announced that annualized revenue crossed $47 billion in May, up dramatically f…

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

メタエージェントの課題: 現在のエージェントは自律的なエージェント開発が可能ですか?

現在の AI ベンチマークは、人間が設計したワークフロー内でのタスク実行に関してエージェントを評価します。これらの評価では、基本的に、モデルが自律的にエージェント システムを開発できるかどうかという、重要な次のレベルの機能を測定できません。自律エージェント開発のためのフロンティア モデルの能力をテストするために設計された評価フレームワークであるメタエージェント チャレンジ (MAC) を紹介します。具体的には、コード エージェント (メタエージェント) には、サンドボックス環境、評価 API、および 5 つのドメインにわたって実施されたテスト セットのパフォーマンスを最大化するエージェント アーティファクトを反復的にプログラムするための時間制限が与えられます。評価の整合性を確保するために、このフレームワークは報酬ハッキングに対する多層防御によって保護されています。このフレームワークを活用して、メタエージェントが人為的に設計されたベースライン ポリシーと一致することはほとんどなく、一致する少数のエージェントは独自のフロンティア モデルによって支配されていることを示します。さらに、設計プロセスは高い分散を示し、高い最適化圧力により、グラウンドトゥルースの漏洩などの敵対的な動作が表面化し、堅牢性とモデルの調整の両方における重大な欠陥が浮き彫りになります。最終的に、MAC は自律型 AI の研究開発のための厳密なオープンソース ベンチマークを提供し、再帰的な自己改善を評価するための経験的な代用手段を提供します。ベンチマークは https://github.com/ant-research/meta-agent-challenge で公開されています。

原文 (English)

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

Current AI benchmarks evaluate agents on task execution within human-designed workflows. These evaluations fundamentally fail to measure a critical next-level capability: whether models can autonomously develop agent systems. We introduce the Meta-Agent Challenge (MAC), an evaluation framework designed to test the capacity of frontier models for autonomous agent development. Specifically, a code agent (the meta-agent) is given a sandboxed environment, an evaluation API, and a time limitation to iteratively program an agent artifact that maximizes performance on a held-out test set across five domains. To ensure evaluation integrity, this framework is secured by multi-layer defenses against reward hacking. Leveraging this framework, we demonstrate that meta-agents rarely match human-engineered baseline policies, and the few that do are dominated by proprietary frontier models. Moreover, the design process exhibits high variance, and high optimization pressure surfaces emergent adversarial behaviors like ground-truth exfiltration-highlighting critical deficits in both robustness and model alignment. Ultimately, MAC provides a rigorous, open-source benchmark for autonomous AI research and development, offering an empirical proxy for evaluating recursive self-improvement. Benchmark is publicly available at: https://github.com/ant-research/meta-agent-challenge.

2026-06-04 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

SpurAudio: 少数ショット音声分類におけるショートカット学習を研究するためのベンチマーク

少数ショット分類 (FSC) は、限られたラベル付きデータから学習するために広く使用されていますが、ほとんどの評価は、ターゲットの概念が文脈上の手がかりから独立していることを暗黙的に前提としています。ただし、現実世界の設定では、サンプルがリッチ コンテキスト内に表示されることが多く、モデルが前景コンテンツと背景信号の間の偽の相関を利用できるようになります。このような効果は少数ショット画像分類で研究されていますが、少数ショット音声分類におけるその役割はほとんど解明されておらず、既存の音声ベンチマークでは文脈構造に対する制御が限られています。 SpurAudio というベンチマークを紹介します。これは、オーディオの前景イベントと背景環境の自然な分離性を活用して、サポートおよびクエリ セットにわたるコンテキストの変化を制御されたマルチレベルの評価を可能にするベンチマークです。このベンチマークを使用して、多くの最先端の少数ショット手法は、標準的な評価プロトコルで同様の精度を達成しているにもかかわらず、バックグラウンド相関が破壊されると重大なパフォーマンス低下に見舞われることがわかります。重要なのは、この脆弱性は大規模な事前トレーニング済みオーディオ基盤モデルでも存続しており、バックボーン容量の制限が説明の対象外となっているということです。さらに、従来のベンチマークでは同等に見える手法でも、偽の相関に対して著しく異なる感度を示す可能性があり、推論時に特徴表現が分類器ヘッドとどのように相互作用するかに関連する体系的なアルゴリズムの強みと脆弱性が明らかになります。これらの発見は、オーディオにおける少数ショット法の動作に関する新たな洞察を提供し、FSC モデルを評価する際のコンテキスト依存性を明示的に調査するベンチマークの必要性を強調しています。

原文 (English)

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues. In real-world settings, however, examples often appear within rich contexts, allowing models to exploit spurious correlations between foreground content and background signals. While such effects have been studied in few-shot image classification, their role in few-shot audio classification remains largely unexplored, and existing audio benchmarks offer limited control over contextual structure. We introduce SpurAudio, a benchmark that leverages the natural separability of foreground events and background environments in audio to enable controlled, multi-level evaluation of contextual shifts across support and query sets. Using this benchmark, we show that many state-of-the-art few-shot methods suffer severe performance degradation when background correlations are disrupted, despite achieving similar accuracy under standard evaluation protocols. Crucially, this vulnerability persists even in large pretrained audio foundation models, ruling out limited backbone capacity as an explanation. Moreover, methods that appear comparable under conventional benchmarks can exhibit markedly different sensitivity to spurious correlations, revealing systematic algorithmic strengths and vulnerabilities tied to how feature representations interact with classifier heads at inference time. These findings provide new insight into the behavior of few-shot methods in audio and highlight the need for benchmarks that explicitly probe context dependence when evaluating FSC models.

2026-06-04 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

L-TGVN: パーソナライズされた高速 MRI のための縦方向事前分布の活用

MRI は電離放射線を使用せずに優れた軟組織コントラストを提供しますが、取得時間が長いため患者の不快感が増大すると同時に、検査コストが上昇し、スキャナのスループットが制限されます。スキャン時間を短縮するための一般的なアプローチは、取得する測定値を少なくすることです。これにより、不適切な線形逆問題が発生します。したがって、診断品質の画像を回復するには、測定データ以外の事前知識を組み込む必要があります。追跡検査では、患者の最新の以前のスキャンにより、非常に有益な被験者固有のコンテキストが提供されますが、実際の使用は、時間的変化(病状の進行を含む)、スキャン間のずれ、取得間のプロトコルのドリフトによって複雑になります。この研究では、大幅にアンダーサンプリングされた測定値から現在のスキャンを再構築するための副次情報として以前のスキャンを活用する、縦方向の信頼誘導変分ネットワークである L-TGVN を紹介します。重要なことは、L-TGVN は、以前のスキャンの影響が取得された測定値と一致するように制限することです。既存の多くの縦方向再構成方法とは異なり、以前のスキャンと現在のスキャンの間の明示的な事前位置合わせを必要としません。さらに、訪問ごとの取得プロトコルの違い(シーケンスパラメータの変更など)にも対応します。私たちは、事前ガイド法や縦方向事前分布を使用しない方法など、一致した容量のベースラインに対して L-TGVN を評価し、困難な加速において微細構造のより良好な保存とともに、標準的な定量的指標の一貫した改善を観察しました。ソース コードは github.com/sodicksonlab/L-TGVN で入手できます。

原文 (English)

L-TGVN: Leveraging Longitudinal Priors for Personalized Rapid MRI

MRI provides excellent soft-tissue contrast without ionizing radiation, but long acquisition times increase patient discomfort while also raising exam costs and limiting scanner throughput. A common approach to reduce scan time is to acquire fewer measurements, which yields an ill-posed linear inverse problem; recovering diagnostic-quality images therefore requires incorporating prior knowledge beyond the measured data. In follow-up exams, the most recent prior scan of a patient can provide a highly informative subject-specific context, but practical use is complicated by temporal changes (including pathology progression), misalignment between scans, and protocol drift across acquisitions. In this work, we introduce L-TGVN, a Longitudinal Trust-Guided Variational Network that leverages prior scans as side information to reconstruct the current scan from heavily undersampled measurements. Crucially, L-TGVN constrains the influence of prior scans to be consistent with the acquired measurements. Unlike many existing longitudinal reconstruction methods, it does not require explicit pre-registration between prior and current scans. It further accommodates differences in acquisition protocols across visits (e.g., changes in sequence parameters). We evaluate L-TGVN against matched-capacity baselines, including prior-guided methods and methods that do not use longitudinal priors, and observe consistent improvements in standard quantitative metrics together with better preservation of fine structures at challenging accelerations. Source code is available at github.com/sodicksonlab/L-TGVN.

2026-06-04 13:00 JSTarXiv cs.AIビジネス/資金調達

RowNet: 表形式回帰のためのメモリ トランスフォーマー

不動産評価は構造化回帰問題であり、価格は異種の特徴タイプ、まばらな地域効果、非線形相互作用、および比較可能な不動産の実際的なロジックによって支配されます。標準的な多層パーセプトロンは各行を孤立ベクトルとして扱い、局所性、スケール感度、およびカテゴリカルマッチングを監視のみから学習する必要があります。勾配ブースト デシジョン ツリーは強力な表形式のベースラインを提供しますが、その特徴中心の分割メカニズムは、類似した履歴観測の取得を明示的にモデル化しません。この論文では、不動産の平方メートルあたりの価格を予測するための検索ベースのニューラル アーキテクチャである RowNet について説明します。 RowNet は、ラベル付きプロパティのメモリ バンクに対するペアごとの類似性機能を通じてクエリ プロパティを表します。最初の検索層は、特徴のみの類似性から大まかなターゲットを推定します。 2 番目の層は、ターゲット一貫性機能を使用してメモリ比較を強化し、複数の学習されたアテンション ヘッドを使用して相補的な比較可能なセットを取得します。最後の専門家混合モジュールは、学習されたゲーティング、残差補正、エントロピー正則化、ヘッドダイバーシティ正則化を組み合わせて予測を生成します。

原文 (English)

RowNet: A Memory Transformer for Tabular Regression

Real estate valuation is a structured regression problem in which prices are governed by heterogeneous feature types, sparse regional effects, nonlinear interactions, and the practical logic of comparable properties. Standard multilayer perceptrons treat each row as an isolated vector and must learn locality, scale sensitivity, and categorical matching from supervision alone. Gradient-boosted decision trees provide strong tabular baselines, but their feature-centric splitting mechanism does not explicitly model the retrieval of similar historical observations. This paper presents RowNet, a retrieval-based neural architecture for real estate price-per-square-meter prediction. RowNet represents a query property through pairwise similarity features against a memory bank of labeled properties. A first retrieval layer estimates a coarse target from feature-only similarities. A second layer augments the memory comparison with target-consistency features and uses multiple learned attention heads to retrieve complementary comparable sets. A final mixture-of-experts module combines learned gating, residual correction, entropy regularization, and head-diversity regularization to produce the prediction.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

共同生成と評価による自己進化する深層研究

大規模言語モデル (LLM) は日常のアプリケーションでますます採用されるようになり、詳細な研究が特に重要な機能として際立っています。従来の質問応答 (QA) タスクとは異なり、詳細な調査レポートの生成には決定的な根拠が欠けているため、報酬設計が本質的に検証不可能になり、効果的な強化学習が制限されます。既存のアプローチでは、LLM-as-a-judge およびクエリ依存の評価ルーブリックを使用してこの課題を軽減していますが、依然として静的な評価器に依存しているため、ソルバーの向上に応じて標準を適応させることができず、最適化圧力が不十分になり、最終的に飽和状態になってしまいます。私たちは、\textbf{s}elf 進化型 \textbf{co} 進化型トレーニング フレームワークで、深い \textbf{re} 検索の評価と生成 (SCORE) を使用してこの制限に対処します。これは、共有パラメータ学習プロセスにおいて評価器とソルバーを緊密に結合します。生成と評価を独立したモジュールとして扱うのではなく、それらの本質的なつながりを活用して、単一の共有パラメーター モデル内で共同の改善を可能にします。このプロセスを制限するために、ソルバーのパフォーマンスに基づいて評価環境を動的に制御するメタハーネスを導入し、有効な評価次元と十分に深い評価者の検索を促進します。ディープリサーチベンチマークに関する広範な実験により、レポート生成の品質が一貫して向上していることが実証されており、評価と生成を共進化させることが、オープンエンドのリサーチエージェントをトレーニングするための有望な方向性であることが示されています。

原文 (English)

Self-Evolving Deep Research via Joint Generation and Evaluation

Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional question-answering (QA) tasks, deep research report generation lacks definitive ground-truth, making reward design inherently unverifiable and limiting effective reinforcement learning. Existing approaches mitigate this challenge with LLM-as-a-judge and query-dependent evaluation rubrics, but they still rely on static evaluators that cannot adapt their standards as the solver improves, leading to insufficient and eventually saturated optimization pressure. We address this limitation with a \textbf{s}elf-evolving \textbf{co}-evolutionary training framework for deep \textbf{re}search evaluation and generation (SCORE), which tightly couples an evaluator and a solver in a shared-parameter learning process. Rather than treating generation and evaluation as isolated modules, we leverage their intrinsic connection to enable joint improvement within a single shared-parameter model. To restrict this process, we introduce a meta-harness, which dynamically controls the evaluation environment based on solver performance, encouraging valid evaluation dimensions and sufficiently deep evaluator search. Extensive experiments on deep research benchmarks demonstrate consistent improvement in report generation quality, showing that co-evolving evaluation and generation is a promising direction for training open-ended research agents.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達

M$^3$Eval: 認知に基づいたビデオタスクによるマルチモーダル記憶評価

マルチモーダル モデルが長時間ビデオの理解に向けて進歩するにつれ、メモリが重要な能力として浮上します。ビデオ データセットとベンチマークの開発には多大な努力が払われているにもかかわらず、既存の研究は主に知覚と推論に焦点を当てており、どのモデルが保持するか、情報がどの程度忠実に保存されるか、干渉下でもメモリがどの程度堅牢に保たれるかなど、記憶を体系的に評価することはありません。このギャップに対処するために、マルチモーダル モデルでさまざまなメモリ次元を調査するための最初の包括的な評価フレームワークおよびベンチマークである M$^3$Eval を導入します。認知心理学に基づいた当社のデザインは、記憶の重要な側面を分離する慎重に構築されたタスクを特徴としています。 M$^3$Eval を活用して、代表的なマルチモーダル モデルにわたって広範な実験を実施し、一貫した弱点と独特の動作を明らかにしました。私たちは、並列ビデオストリームを処理する際にモデルがもつれの解けた表現を維持するのに苦労し、人間の記憶で観察されるものとは大幅に異なる干渉パターンを示し、記憶ソースを時間領域よりも空間領域でより確実に接地し、限られた記号記憶を実証していることを発見しました。まとめると、私たちのベンチマークは将来の研究のための貴重なリソースを提供しますが、私たちの調査結果は、メモリが基本的でありながらまだ研究されていない機能であることを強調し、マルチモーダルモデルでより効果的なメモリメカニズムを設計するための洞察を提供します。コードとデータセットは https://pku-value-lab.github.io/m3eval-homepage で入手できます。

原文 (English)

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks

As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability. Despite substantial efforts in developing video datasets and benchmarks, existing works primarily focus on perception and reasoning, without systematically evaluating memory: what models retain, how faithfully information is preserved, and how robust memory remains under interference. To address this gap, we introduce M$^3$Eval, the first comprehensive evaluation framework and benchmark for probing different memory dimensions in multi-modal models. Grounded in cognitive psychology, our design features carefully constructed tasks that isolate key aspects of memory. Leveraging M$^3$Eval, we conduct extensive experiments across representative multi-modal models, revealing consistent weaknesses and distinctive behaviors. We find that models struggle to maintain disentangled representations when processing parallel video streams, exhibit interference patterns differing substantially from those observed in human memory, ground memory sources more reliably in the spatial domain than the temporal domain, and demonstrate limited symbolic memory. Collectively, our benchmark provides a valuable resource for future research, while our findings highlight memory as a fundamental yet underexplored capability and offer insights for designing more effective memory mechanisms in multi-modal models. Our code and dataset are available at https://pku-value-lab.github.io/m3eval-homepage.

2026-06-04 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

答えから状態へ: 大規模言語モデルにおける化学推論の検証可能なプロセスレベルの評価

大規模な言語モデルが化学アシスタントとして使用されることが増えていますが、ほとんどの化学ベンチマークは依然として最終的な回答のみをスコアとしています。これにより、重大な故障モードが隠蔽されます。モデルは、その推論が化学ロジックに違反しているにもかかわらず、正しい分子、生成物、またはオプションを出力する可能性があります。 LLM ジャッジと人間のステップレベルのプロセス アノテーションはコストが高く、一貫性がなく、幻覚に対して脆弱であるため、既存のプロセス レベルの評価機能を拡張するのは困難です。 ChemCoTBench-V2 は、構造化され検証者がアドレス指定できる化学推論トレースを低コストで監査可能に評価するためのルール検証可能な診断ベンチマークです。これは、分子理解、分子編集、分子最適化、反応予測に及び、18 のレポートタスクにわたる 5,620 の評価サンプルを備えています。モデルは、専門家が設計したテンプレートで主要な中間ステップを公開する必要があり、それらのステップは決定論的な化学ルールでチェックされ、クローズドアンサータスクの場合は、別の LLM 審査員ではなく参照トレースが使用されます。オープンエンド分子最適化は、厳密なトレース マッチングではなく、Oracle で検証可能な状態制約を使用して評価されます。このベンチマークは、最終回答の正確性、テンプレートの遵守、専門家によって洗練された中間コミットメントに対する段階的な検証者の正確さという 3 つの個別のシグナルを報告します。フロンティア モデルの実験では、最終的な回答の成功と構造化推論の状態の一貫性の間には永続的なギャップがあることが明らかになりました。モデルは多くの場合、化学ステップ チェックに失敗しながらも要求された形式に従っているか、弱い裏付け推論で正しく回答することができます。 ChemCoTBench-V2 は、きめ細かいモデル比較を可能にし、トレースが最初に検証ツールに違反する具体的なステップを特定します。

原文 (English)

From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may output the correct molecule, product, or option while its reasoning violates chemical logic. Existing process-level evaluators are hard to scale because LLM judges and human step-level process annotation are costly, inconsistent, and vulnerable to hallucination. We introduce ChemCoTBench-V2, a rule-verifiable diagnostic benchmark for low-cost, auditable evaluation of structured, verifier-addressable chemical reasoning traces. It spans molecular understanding, molecule editing, molecular optimization, and reaction prediction, with 5,620 evaluation samples across 18 reporting tasks. Models must expose key intermediate steps in expert-designed templates, and those steps are checked with deterministic chemistry rules and, for closed-answer tasks, reference traces rather than another LLM judge. Open-ended molecular optimization is evaluated with oracle-verifiable state constraints rather than strict trace matching. The benchmark reports three separate signals: final-answer correctness, template adherence, and step-wise verifier correctness over expert-refined intermediate commitments. Experiments on frontier models reveal a persistent gap between final-answer success and structured-reasoning-state consistency: models often follow the requested format while failing chemical-step checks, or answer correctly with weak supporting reasoning. ChemCoTBench-V2 enables fine-grained model comparison and identifies the concrete step at which the trace first violates the verifier.

2026-06-04 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

CounterFace: 顔認識システムのきめ細かい反事実評価のための合成顔データセット

顔認識 (FR) システムは重要なアプリケーションに広く導入されており、多様な人口や条件に対する信頼性と堅牢性が不可欠となっています。 FR システムの標準評価は通常、LFW などのデータセットに依存して平均認識精度を推定します。一部のベンチマークは、経年変化、姿勢、照明などの粗粒度のアイデンティティ内の変動も捕捉します。ただし、人間の顔には、ヘアスタイルやメイクなどの外観の変化を含む、より細かい変化が生じますが、これは既存のベンチマークでは過小評価されています。反事実評価は、このようなきめの細かい変動の下で FR の堅牢性を評価する方法を提供します。ただし、画像ジェネレーターを使用して合成された既存の反事実の顔データセットは、パイプラインでの検証に人間が使用されているため、属性の範囲が限られています。我々は、20 の顔属性と 8 つの人口統計的要素で構成される新しい反事実評価データセットである CounterFace を提案します。これは、以前の合成顔データセットを 14 属性と 2 つの人口統計的要因で上回っています。データセットは、カスタム検証機能を備えた既製の画像ジェネレーターに基づいた完全に自動化されたパイプラインを使用して生成され、人間による検証の必要性がなくなりました。 CounterFace には 11,821 の反事実の顔のペアが含まれており、事後のユーザー調査により、生成された反事実の忠実性が確認されています。 160 の属性と人口統計の組み合わせにわたって、2 つの商用 FR システムと 4 つのオープンソース FR システム (AWS Rekognition、Face++、AdaFace、MagFace、ArcFace、FaceNet) を評価します。当社のデータセットは、標準の評価ベンチマークとは異なり、個々のシステムの正確な故障モードを分離するのに役立ちます。結果は、パフォーマンスの低下は 6 つすべてのシステムの属性と人口統計によって異なり、遮蔽属性 (フェイスマスクやひげなど) が普遍的にパフォーマンスを低下させることを示しています。

原文 (English)

CounterFace: A Synthetic Face Dataset for Fine-Grained Counterfactual Evaluation of Face Recognition Systems

Face recognition (FR) systems are widely deployed in critical applications, making their reliability and robustness across diverse populations and conditions essential. Standard evaluation of FR systems typically relies on datasets such as LFW to estimate average recognition accuracy. Some benchmarks also capture coarse-grained intra-identity variations such as aging, pose, and lighting. However, human faces undergo more fine-grained changes, including appearance changes such as hairstyles and makeup, that are underrepresented in existing benchmarks. Counterfactual evaluation provides a method to assess FR robustness under such fine-grained variations. Existing counterfactual face datasets synthesized with image generators, however, are limited in attribute coverage due to the use of humans for verification in the pipeline. We propose CounterFace, a new counterfactual evaluation dataset comprising 20 facial attributes and 8 demographic factors, exceeding prior synthetic face datasets by 14 attributes and 2 demographics. The dataset is generated using a fully automated pipeline based on off-the-shelf image generators with custom verifiers, removing human need for verification. CounterFace contains 11,821 counterfactual face pairs, and a post-hoc user study confirms the faithfulness of the generated counterfactuals. We evaluate two commercial and four open-source FR systems (AWS Rekognition, Face++, AdaFace, MagFace, ArcFace, FaceNet) across 160 attribute-demographic combinations. Our dataset helps in the isolation of precise failure modes for individual systems unlike standard evaluation benchmarks. Results indicate that the performance degradation varies across attributes and demographics for all six systems and occluding attributes (e.g., facemask and facial hair) universally degrade performance.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

グラフ検索からスキーマ実現まで: 異種ナレッジ グラフ上のテキストから SPARQL への反事実検証

Text-to-SPARQL は、自然言語の質問を RDF ナレッジ グラフ上の実行可能な SPARQL クエリにマッピングします。標準的な評価ではターゲット グラフが事前に修正されることがよくありますが、実践的なナレッジ グラフ質問応答 (KGQA) には、異なるスキーマ、部分的なアラインメント、および不完全なメタデータを含む異種グラフ コレクションが含まれる場合があります。この設定では、クエリ生成は SPARQL 構文以上のものに依存します。システムは、質問に必要な述語、エンティティ タイプ、結合、フィルター、および制約をサポートできるグラフ スキーマを識別する必要があります。異種の KG コレクション上でテキストから SPARQL に変換するためのスキーマベースのエージェント フレームワークである SchemaForge を紹介します。その中心的なメカニズムは、質問条件付きのスキーマ スライス アライメントです。弱いグラフの証拠によって最初にもっともらしいグラフが特定され、より強力なスキーマの証拠によって、ローカル スキーマ スライスが意図したクエリを実現できるかどうかが決まります。選択されたスキーマ スライスは、クエリの生成と実行前の検証を制限します。利用可能なグラフが 1 つだけの場合、同じ定式化は、スキーマ基盤を備えた標準の単一 KG テキストから SPARQL への変換に縮小されます。 LC-QuAD 2.0、QALD-9 Plus、QALD-10、および Spider4SPARQL で SchemaForge を評価します。 SchemaForge は、4 つの公開ベンチマーク全体で、最も一致するエージェントのベースラインよりも実行精度を平均 11.50 パーセント向上させています。 Spider4SPARQL では、SchemaForge は実行精度を 54.86% から 64.18% に向上させ、トップ 1 グラフ割り当て精度 73.0% とトップ 3 グラフ割り当て精度 97.0% を達成しました。これらの結果は、グラフの弱い証拠からスキーマ固有のクエリコミットメントへの移行と、反事実の回答セットのチェックにより、異種ナレッジグラフよりも実行可能なクエリの生成が向上することを示しています。

原文 (English)

From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs

Text-to-SPARQL maps natural-language questions to executable SPARQL queries over RDF knowledge graphs. While standard evaluations often fix the target graph in advance, practical knowledge graph question answering (KGQA) may involve heterogeneous graph collections with different schemas, partial alignments, and incomplete metadata. In this setting, query generation depends on more than SPARQL syntax: the system must identify a graph schema that can support the predicates, entity types, joins, filters, and constraints required by the question. We present SchemaForge, a schema-grounded agentic framework for text-to-SPARQL over heterogeneous KG collections. Its central mechanism is question-conditioned schema-slice alignment: weak graph evidence first identifies plausible graphs, while stronger schema evidence determines whether a local schema slice can realize the intended query. The selected schema slice then constrains query generation and verification before execution. When only one graph is available, the same formulation reduces to standard single-KG text-to-SPARQL with schema grounding. We evaluate SchemaForge on LC-QuAD 2.0, QALD-9 Plus, QALD-10, and Spider4SPARQL. Across the four public benchmarks, SchemaForge improves execution accuracy over the strongest matched agent baseline by 11.50 percentage points on average. On Spider4SPARQL, SchemaForge improves execution accuracy from 54.86% to 64.18% and achieves 73.0% Top-1 and 97.0% Top-3 graph allocation accuracy. These results show that moving from weak graph evidence to schema-specific query commitments, together with counterfactual answer-set checks, improves executable query generation over heterogeneous knowledge graphs.

2026-06-04 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

VGGSounder: 基礎モデルのオーディオビジュアル評価

視聴覚基礎モデルの出現は、マルチモーダルな理解を確実に評価することの重要性を強調しています。 VGGSound データセットは、オーディオビジュアル分類の評価のベンチマークとしてよく使用されます。ただし、私たちの分析では、不完全なラベル付け、部分的に重複するクラス、不整合なモダリティなど、VGGSound のいくつかの制限が特定されました。これらは、聴覚および視覚能力の歪んだ評価につながります。これらの制限に対処するために、VGGSounder を導入します。これは、VGGSound を拡張し、オーディオビジュアル基礎モデルを評価するために特別に設計された、包括的に再アノテーションが付けられたマルチラベル テスト セットです。 VGGSounder は詳細なモダリティの注釈を備えており、モダリティ固有のパフォーマンスを正確に分析できます。さらに、新しいモダリティ混乱メトリックを使用して別の入力モダリティを追加したときのパフォーマンスの低下を分析することで、モデルの限界を明らかにします。

原文 (English)

VGGSounder: Audio-Visual Evaluations for Foundation Models

The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.

2026-06-04 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ノイズを含む音声分離におけるスケール不変信号対歪み比の研究

この論文では、事実上のベンチマーク WSJ0-2Mix の場合のように、トレーニング参照にノイズが含まれている場合に、教師あり音声分離における評価とトレーニングの目的の両方としてスケール不変信号対歪み比 (SI-SDR) を使用することの意味を検証します。ノイズの多いリファレンスを使用して SI-SDR を導出すると、ノイズによって達成可能な SI-SDR が制限されるか、分離された出力に望ましくないノイズが発生することがわかります。これに対処するために、ノイズの多い参照の学習を回避するモデルをトレーニングすることを目的として、WHAM! を使用して参照を強化し、混合を増強する方法が提案されています。これらの強化されたデータセットでトレーニングされた 2 つのモデルは、非侵入的な NISQA.v2 メトリックを使用して評価されます。結果は、分離された音声のノイズが減少していることを示していますが、参照の処理によりアーチファクトが生じ、全体的な品質の向上が制限される可能性があることが示唆されています。 WSJ0-2Mix および Libri2Mix テスト セットのモデル全体で、SI-SDR と知覚されるノイズの間に負の相関関係が見つかり、導出による結論が強調されています。

原文 (English)

A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References

This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with the de facto benchmark WSJ0-2Mix. A derivation of the SI-SDR with noisy references reveals that noise limits the achievable SI-SDR, or leads to undesired noise in the separated outputs. To address this, a method is proposed to enhance references and augment the mixtures with WHAM!, aiming to train models that avoid learning noisy references. Two models trained on these enhanced datasets are evaluated with the non-intrusive NISQA.v2 metric. Results show reduced noise in separated speech but suggest that processing references may introduce artefacts, limiting overall quality gains. Negative correlation is found between SI-SDR and perceived noisiness across models on the WSJ0-2Mix and Libri2Mix test sets, underlining the conclusion from the derivation.

2026-06-04 13:00 JSTarXiv cs.AIビジネス/資金調達

分散ゲート分布を使用した不確かさの推定

ニューラル ネットワークからのサンプルごとの不確実性の定量化の評価は、高リスクのアプリケーションを含む意思決定に不可欠です。一般的なアプローチは、ベイジアン モデルまたは近似モデルからの予測分布を使用し、対応する予測の不確実性を認識的 (モデル関連) 成分と偶然的 (データ関連) 成分に分解することです。しかし、最近では相加的分解に疑問が持たれています。この研究では、さまざまなモデル予測にわたるクラス確率分布の信号対雑音比に基づいて、不確実性の推定と分解を行うための直感的なフレームワークを提案します。アンサンブルから導出された信頼係数によって予測をスケールする分散ゲート測定を導入します。私たちはこの尺度を利用して、委員会マシンの多様性の崩壊の存在について議論します。

原文 (English)

Uncertainty Estimation using Variance-Gated Distributions

Evaluation of per-sample uncertainty quantification from neural networks is essential for decision-making involving high-risk applications. A common approach is to use the predictive distribution from Bayesian or approximation models and decompose the corresponding predictive uncertainty into epistemic (model-related) and aleatoric (data-related) components. However, additive decomposition has recently been questioned. In this work, we propose an intuitive framework for uncertainty estimation and decomposition based on the signal-to-noise ratio of class probability distributions across different model predictions. We introduce a variance-gated measure that scales predictions by a confidence factor derived from ensembles. We use this measure to discuss the existence of a collapse in the diversity of committee machines.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark ne…

2026-06-04 13:00 JSTarXiv cs.AIビジネス/資金調達

Markov Chain Decoders Overcome the Heavy-Tail Limitations of Lipschitz Generative Models

Heavy-tailed distributions are prevalent in performance evaluation, network traffic, and risk modeling. This behavior poses a fundamental c…

2026-06-04 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

メタ学習による費用対効果の高いモデル評価

機械学習の急速な成長により、拡大し続けるモデルのエコシステムが生み出され、目に見えないラベルのないデータに対して新しくリリースされたモデルの信頼性を検証することがますます困難になっています。従来の評価パイプラインは、高価なアノテーション、繰り返しの微調整、またはモデル ファミリ間での転送ができない狭い仮定に依存しています。さまざまなアーキテクチャやモダリティにまたがる未確認のモデルをラベルなしで迅速に評価するための、コスト効率が高く、モデルに依存しないフレームワークである MetaEvaluator を紹介します。 MetaEvaluator は、参照モデルのプールに対するメタ学習を利用して転送可能な初期化を取得し、プール全体でコストを償却しながら、モデルごとの再トレーニングの必要性を排除しながら、新しいモデルの正確な評価を可能にします。私たちの知る限り、これは完全にラベルのないデータセットで新しいモデルを評価できる、モデルに依存しない最初のフレームワークです。広範な実験により、MetaEvaluator は従来のアプローチと比較して大幅にコストを削減しながら安定した正確なパフォーマンス推定値を生成し、ラベルのないデータに対する新しいモデルのスケーラブルなベンチマークを実用化できることが示されています。

原文 (English)

Learning to Evaluate: Cost-Effective Model Evaluation on Unlabeled Data with Meta-Learning

The rapid advancement of machine learning has led to an unprecedented expansion of model ecosystems, making it increasingly difficult to assess the reliability of newly released models on unseen and unlabeled data. Existing evaluation pipelines typically rely on costly annotation, repeated fine-tuning, or assumptions that do not generalize well to new models. We introduce MetaEvaluator, a cost-effective, model-agnostic framework for fast, label-free evaluation of unseen models across diverse architectures and modalities. MetaEvaluator meta-learns over a pool of reference models to acquire an effective initialization for accurate assessment of unseen models, thereby amortizing evaluation cost and eliminating the need for per-model retraining. To the best of our knowledge, this is the first model-agnostic framework that evaluates new models on unlabeled datasets. Extensive experiments demonstrate that MetaEvaluator delivers stable and accurate performance estimates at substantially lower cost than conventional approaches, enabling scalable benchmarking on unlabeled datasets for emerging models. The code is available at: https://github.com/phkhanhtrinh23/MetaEvaluator.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

BenHalluEval: ベンガル語の大規模言語モデル用のマルチタスク幻覚評価フレームワーク

ベンガル語は世界で 6 番目に話されている言語であるにもかかわらず、ベンガル語の大規模言語モデル (LLM) で幻覚を体系的に評価した先行研究はありません。 BenHalluEval は、生成的質問応答 (GQA)、バングラ語と英語のコード混合 QA、要約、および推論の 4 つのタスクをカバーするベンガル語用のきめ細かい幻覚評価フレームワークです。既存の 3 つのベンガル語データセットから抽出された、12 のタスク固有の幻覚タイプにわたって GPT-5.4 を使用して 12,000 の幻覚候補を構築し、グラウンドトゥルース インスタンスの偽陽性率 (トラック A) と幻覚候補の幻覚検出率 (トラック B) を独立して測定するデュアル トラック プロトコルの下で、推論指向、多言語、ベンガル語中心のカテゴリにわたる 7 つの LLM を評価します。両方の故障モードに共同でペナルティを課し、均一な応答バイアスによるスコアのインフレを防ぐために、モデルとタスク全体で 7.72% から 55.42% の範囲のデュアルトラック キャリブレーション メトリクスである BenHalluScore を提案します。これは、幻覚キャリブレーションの大幅な変動を明らかにします。緩和戦略として適用される思考連鎖プロンプトは、幻覚差別を一貫して改善することなく、反応分布を変化させます。 BenHalluEval は、ベンガル語専用の幻覚ベンチマークを初めて確立し、リソースの少ない言語設定に対する単一トラックおよびプロンプトのみの評価アプローチが不適切であることを強調しています。データセットとコードは https://anonymous.4open.science/r/BanglaHalluEval-EB77 で入手できます。

原文 (English)

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates using GPT-5.4 across twelve task-specific hallucination types, drawn from three existing Bengali datasets, and evaluate seven LLMs spanning reasoning-oriented, multilingual, and Bengali-centric categories under a dual-track protocol that independently measures false-positive rate on ground-truth instances (Track A) and hallucination detection rate on hallucinated candidates (Track B). To jointly penalise both failure modes and prevent inflated scores from uniform response bias, we propose BenHalluScore, a dual-track calibration metric that ranges from 7.72% to 55.42% across models and tasks, revealing substantial variation in hallucination calibration. Chain-of-thought prompting, applied as a mitigation strategy, shifts response distributions without consistently improving hallucination discrimination. BenHalluEval establishes the first dedicated hallucination benchmark for Bengali and highlights the inadequacy of single-track and prompting-only evaluation approaches for low-resource language settings. The dataset and code are available at https://anonymous.4open.science/r/BanglaHalluEval-EB77.

2026-06-04 04:38 JSTTechCrunch AIビジネス/資金調達

Alphabet’s record-breaking $85B raise for Google’s AI business is a helluva good signal

If Alphabet's record-breaking $85 billion stock sale signals investor appetite for AI-related offerings, we can see that investors are read…

2026-06-03 22:02 JSTTechCrunch AIエージェントビジネス/資金調達

Coralogix raises $200M on bet that someone needs to watch the AI agents

Coralogix is among a growing number of infrastructure firms betting that as AI systems move into production, demand will rise for tools tha…

2026-06-03 13:00 JSTarXiv cs.AIビジネス/資金調達

BehaviorBench: 行動追跡から現実世界のユーザーの意思決定をモデル化

多くの意思決定支援設定では、個々のユーザーに適応するシステムが必要ですが、この問題に関する評価データは依然として限られています。ユーザー理解のための既存のベンチマークは、多くの場合、シミュレートされたユーザーやモデルで生成された動作に依存していますが、最近の研究では、モデルベースのシミュレーションが人間の動作から系統的に逸脱する可能性があると警告されています。現実世界の行動追跡からパーソナライズされた意思決定モデリングを評価するためのベンチマークである \textsc{BehaviorBench} を紹介します。 \textsc{BehaviorBench} は、観測された公開予測市場記録とオンチェーン記録からウォレットレベルの意思決定履歴を再構築し、それらを 2 つの補完的なタスク層に編成します。\emph{信念予測} は市場に対するユーザーの最終的なスタンスと自信を予測し、\emph{取引予測} は個々の取引の方向と金額を予測します。 2,000 の評価ウォレットにわたって、ベンチマークには 141,445 個の信念インスタンスと 1,485,972 個の取引インスタンスが含まれており、検索ベースの評価のための独立したサポート プールが含まれています。私たちは、パーソナライゼーションなし、直接の最近の履歴、生成されたユーザー プロファイル、および取得されたサポート ウォレットの証拠という 4 つの履歴インターフェイスの下で、フロンティアおよびオープンウェイト生成モデルを評価します。パーソナライゼーションにより、取引予測よりも一貫して信念予測が向上し、モデルのランキングがタスク レイヤーとメトリクスにわたって変化し、さまざまな履歴インターフェイスによりさまざまな障害モードが明らかになります。 \textsc{BehaviorBench} は、パーソナライズされたメソッドがシミュレートされたユーザーのみではなく現実世界の行動証拠を使用できるかどうかを研究するための評価設定を提供します。

原文 (English)

BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited. Existing benchmarks for user understanding often rely on simulated users or model-generated behavior, even though recent work cautions that model-based simulations can diverge systematically from human behavior. We introduce \textsc{BehaviorBench}, a benchmark for evaluating personalized decision modeling from real-world behavioral traces. \textsc{BehaviorBench} reconstructs wallet-level decision histories from observed public prediction-market and on-chain records, and organizes them into two complementary task layers: \emph{Belief prediction}, which predicts a user's final revealed stance and confidence in a market, and \emph{Trade prediction}, which predicts the direction and amount of individual transactions. Across 2,000 evaluation wallets, the benchmark contains 141,445 Belief instances and 1,485,972 Trade instances, with disjoint support pools for retrieval-based evaluation. We evaluate frontier and open-weight generative models under four history interfaces: no personalization, direct recent history, generated user profiles, and retrieved support-wallet evidence. Personalization improves Belief prediction more consistently than Trade prediction, model rankings change across task layers and metrics, and different history interfaces expose different failure modes. \textsc{BehaviorBench} provides an evaluation setting for studying whether personalized methods can use real-world behavioral evidence rather than simulated users alone.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ギャンブルはしないでください、GAMBLe: AI 主導の研究システムのための分析フレームワーク

AI-Driven Research Systems (ADRS) -- LLM と自動評価を組み合わせてアルゴリズム、証明、設計を発見するシステム -- は最適化され、ドメイン全体で採用されていますが、それらを分析するツールは追いついていません。 ADRS のパフォーマンスはコンポーネントの相互作用に依存しますが、これらの相互作用は十分に理解されておらず、調査にコストがかかり、(ここで示しているように) 標準の収束保証では十分に把握されていません。これらの保証は、私たちが形式化した ADRS プロセスの下では成立しない構造的な仮定に依存しています。我々は、ADRS の動作を 4 つのパラメーター (ジェネレーター $G$、アセッサー $\mathcal{A}$、発見メカニズム $\mathcal{M}$、バジェット $B$) と 1 つの構成オブジェクト、効果的なランドスケープ $L_{\text{eff}} = \mathcal{A} \circ G$ に分解するフレームワークである GAMBLe を紹介します。これにより、異なるジェネレーターとアセッサーのペアが構造的に異なる問題ごとの最適化を引き起こすことが明らかになります。風景。私たちは、単一の LLM から動的適応アンサンブルに至るジェネレーター、貪欲な選択から共進化メタサーチに至るメカニズム、および評価者が連続スコアリングからクリフ関数に及ぶ 3 つの NP 困難問題に及ぶ 760 以上の反復実行 (>46,000 反復) でフレームワークを実行します。実験では、ジェネレーターやメカニズムの完全な順序付けは明らかにされていません。フロンティア モデルはオープンソースの代替モデルよりもパフォーマンスが劣る可能性があり、最も単純なメカニズムが最先端のメタ検索を上回る場合もあります。結果は、限られた予算 (実行ごとに 60 回の反復) の下でも、適切なコンポーネントを選択することでパフォーマンスを 13 ~ 67%、検索効率を 6 ~ 39 倍改善できることを示しています。

原文 (English)

Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems

AI-Driven Research Systems (ADRS) -- systems coupling LLMs with automated evaluation to discover algorithms, proofs, and designs -- are being optimized and adopted across domains, but the tools to analyze them have not kept pace. ADRS performance depends on component interactions that are poorly understood, expensive to explore, and (as we show) not well captured by standard convergence guarantees. These guarantees rely on structural assumptions that do not hold under the ADRS process we formalize. We introduce GAMBLe, a framework that decomposes ADRS behavior into four parameters (generator $G$, assessor $\mathcal{A}$, discovery mechanism $\mathcal{M}$, budget $B$) and one compositional object, the effective landscape $L_{\text{eff}} = \mathcal{A} \circ G$, which reveals that distinct generator-assessor pairs induce structurally different per-problem optimization landscapes. We exercise the framework on 760+ replicated runs (>46,000 iterations) spanning generators from single LLMs to dynamically-adaptive ensembles, mechanisms from greedy selection to co-evolutionary meta-search, and three NP-hard problems whose assessors range from continuous scoring to cliff functions. The experiments reveal no total ordering of generators or mechanisms: frontier models can underperform open-source alternatives and the simplest mechanism sometimes outperforms state-of-the-art meta-search. Results show that even under limited budgets (60 iterations per run), the right component choices can improve performance by 13-67% and search efficiency by 6-39x.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

話す前に考える: マルチエージェント社会シミュレーションにおける内部評価から公の表現まで

LLM ベースのマルチエージェント シミュレーションは、社会的相互作用、熟慮、集団的な意見のダイナミクスを研究するための有望な方法を提供します。しかし、既存の対話シミュレーション フレームワークの多くは、対話を主に観察可能なターン交換または集約された出力として表現しており、沈黙、発言意図、公的表現の背後にある内部評価プロセスを調査することが困難なままになっています。エージェントの私的な推論を公的発話の生成から分離する、インターバルベースのマルチエージェント シミュレーション フレームワークである TBS (Think-Before-Speak) を紹介します。各間隔で、すべてのエージェントは共有された対話履歴と自身の記憶に基づいて構造化された内部状態を更新します。これらの状態には、不協和音関連の評価、認識された世論環境、認識された孤立リスク、対応戦略、および発言意欲が含まれます。その後、オーケストレーターは競合する発言意図を解決し、1 つの発言を公開対話にコミットし、内部評価と公開対話が時間の経過とともに共進化できるようにします。私たちは、気候関連の政策問題に関するタウンホールでの議論を模擬して TBS を評価します。結果は、TBS が一貫した内部状態トレースを生成し、これらのトレースがターン割り当て、沈黙、メモリ条件全体にわたって体系的に変化することを示しています。不協和音関連の評価はエージェントの発言意欲を高めますが、沈黙の圧力評価はそれを低下させます。発言の意図が形成されると、公の場での表現は主に順番の割り当てルールによって形成されます。これらの発見は、TBS が内部評価から公的表現への経路を観察可能かつ分析可能にすることで、メカニズムに敏感な社会シミュレーションをサポートしていることを示唆しています。

原文 (English)

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existing dialogue simulation frameworks represent interaction mainly as observable turn exchange or aggregated outputs, leaving the internal evaluative processes behind silence, speaking intention, and public expression difficult to examine. We introduce TBS (Think-Before-Speak), an interval-based multi-agent simulation framework that separates agents' private reasoning from public utterance generation. At each interval, all agents update structured internal states based on the shared dialogue history and their own memory. These states include dissonance-related appraisal, perceived opinion climate, perceived isolation risk, response strategy, and willingness to speak. The orchestrator then resolves competing speaking intentions and commits one utterance to the public dialogue, allowing internal evaluation and public interaction to co-evolve over time. We evaluate TBS in simulated town hall discussions on a climate-related policy issue. Results show that TBS produces coherent internal-state traces and that these traces vary systematically across turn-allocation, silence, and memory conditions. Dissonance-related appraisal increases agents' willingness to speak, whereas silence-pressure appraisal decreases it. Once speaking intention is formed, public expression is shaped mainly by turn-allocation rules. These findings suggest that TBS supports mechanism-sensitive social simulation by making the pathway from internal evaluation to public expression observable and analyzable.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ベンチマーク監査における信頼性ギャップ: 汚染検出の障害モードとしての分布のシフトとスケール

評価例がモデルのトレーニング データに現れるベンチマーク汚染は、LLM 評価の妥当性を脅かします。トレーニング データのメンバーシップを検出するための統計ツールは存在しますが、ほぼ独占的に管理された学術体制、つまり大規模で均質な事前トレーニング コーパスと透明な単一ステージ トレーニング パイプラインでのみ検証されています。これらの方法が現実的な監査シナリオにおいて信頼性を維持できるかどうかは、依然として不明です。私たちは、十分に研究されていない 2 つの障害モードを特定します。1 つは、疑わしいセットと検証セットが IID の仮定に違反する場合に発生する分布シフト、もう 1 つは、ベンチマークがトレーニング前のコーパスよりも桁違いに小さいために発生するスケール制約です。私たちは、複数のファミリー (Pythia、OLMo~2、特殊な文化的および医療的 LLM を含む) およびスケール (最大 27B) からの 27 のモデルにわたって、LLM データセット推論、ポストホック データセット推論、CoDeC という 3 つの主要なパラダイムを体系的に評価します。次に、分析を最先端の業界モデルにさらに拡張します。 335 件の評価のうち、正しい結果が得られたのは 199 件のみでした。 LLM データセット推論では、分布シフトの下で偽陽性が発生し、ポストホック データセット推論はベンチマーク スケールでは能力が不足し、CoDeC は個々のベンチマーク分割を検証するには不十分な粗い出所信号しか提供しません。私たちの結果は、管理された検証と実際のベンチマーク監査の間に体系的な信頼性のギャップがあることを明らかにし、統計的検出がまだ透明なデータ来歴に取って代わることができないことを示しています。私たちはさらなる研究のためにベンチマークをオープンソースにしています。

原文 (English)

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for detecting training-data membership exist, but have been validated almost exclusively in controlled academic regimes: large, homogeneous pre-training corpora and transparent, single-stage training pipelines. Whether these methods remain reliable in realistic auditing scenarios remains unclear. We identify two under-studied failure modes: distribution shift, which arises when suspect and validation sets violate the IID assumption, and scale constraints, which arise because benchmarks are orders of magnitude smaller than pre-training corpora. We systematically evaluate three leading paradigms: LLM Dataset Inference, Post-Hoc Dataset Inference, and CoDeC across 27 models from multiple families (including Pythia, OLMo~2, and specialised cultural and medical LLMs) and scales (up to 27B). We then further extend our analysis to frontier industry models. Across 335 evaluations, only 199 yield correct outcomes. LLM Dataset Inference results in false positives under distribution shift, Post-Hoc Dataset Inference is underpowered at benchmark scale, and CoDeC provides only coarse provenance signals that are insufficient to verify individual benchmark splits. Our results reveal a systematic reliability gap between controlled validation and practical benchmark auditing, and show that statistical detection cannot yet replace transparent data provenance. We open-source our benchmark for further research.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

SAGE: エージェント生態系における社会化進化の定量的評価

自己改善型言語エージェントは通常、単独で評価されます。エージェントはタスクを試み、フィードバックを受け取り、繰り返し自身の動作を改善します。しかし、エージェントは、戦略と結果が公に公開されている同僚と協力して活動することが増えています。このことから、十分に研究されていない疑問が生じます。共有された経験が、自己改善だけでは達成できない改善をもたらすのはいつでしょうか? 2 つのコンピューティングが一致する条件を比較する評価フレームワークである SAGE (ソーシャル エージェント グループ エボリューション) を紹介します。SocialEvo では、5 つの異なるモデル ファミリのエージェントがすべてのピアの履歴にアクセスしながら共同進化します。そして、SelfEvo では、各エージェントは同じ回数のタスク試行を受けますが、自分自身の過去のみを見ることができます。これは、自己改善エージェントの研究では一般的です。私たちは、オープンエンドの ML 研究、長期的な経済計画、戦略的なマルチプレイヤー プレイの 3 つの分野で SAGE をインスタンス化し、複数の進化ラウンドにわたって評価します。私たちは、グループの歴史が普遍的な増幅器ではないことを発見しました。つまり、最も強力なエージェントは自己進化の上限を超えることはありません。ただし、自己改善が停滞しているエージェントでも、同僚の経験があれば、大きな進歩を遂げることができます。競争環境では、反事実的なコントロールにより、エージェントが対戦相手固有の戦略を開発するのではなく、全体的に向上することが明らかになります。さまざまな形式の共有履歴にわたって、フィルタリングされたピアトレースやリフレクションサマリーは生のログよりもパフォーマンスが優れていることが多く、社会的利益は露出量ではなく抽象化に依存していることを示しています。これらの発見は、ピア履歴の獲得がエージェント固有、アリーナ依存であり、公開された痕跡から譲渡可能な知識を抽象化する能力に依存していることを明らかにしています。

原文 (English)

SAGE: A Quantitative Evaluation of Socialized Evolution in Agent Ecosystems

Self-improving language agents are typically evaluated in isolation: an agent attempts a task, receives feedback, and iteratively refines its own behavior. Yet agents increasingly operate alongside peers whose strategies and outcomes are publicly visible. This raises an under-studied question: when does shared experience produce improvements that self-improvement alone cannot achieve? We introduce SAGE (Social Agent Group Evolution),an evaluation framework that compares two compute-matched conditions: SocialEvo, where agents from five distinct model families co-evolve with access to all peers' histories; and SelfEvo, where each agent receives the same number of task attempts but sees only its own past, which is conventional in self-improving agent studies. We instantiate SAGE in three arenas: open-ended ML research, long-horizon economic planning, and strategic multiplayer play, evaluated across multiple evolutionary rounds. We find that group history is not a universal amplifier: the strongest agent does not exceed its self-evolution ceiling. However, agents that plateau under self-improvement can achieve significant breakthroughs when peer experience is available. In competitive settings, counterfactual controls reveal that agents improve generally rather than developing opponent-specific strategies. Across different forms of shared history, filtered peer traces and reflective summaries often outperform raw logs, indicating that social gains depend on abstraction rather than exposure volume. These findings reveal that peer-history gains are agent-specific, arena-dependent, and contingent on the capacity to abstract transferable knowledge from public traces.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

LLM ツール使用における知識ギャップの診断: 新しい API 取得のためのエージェント ベンチマーク

コード生成のための大規模な言語モデルでは、多くの場合、事前トレーニング データに含まれていない API を使用する必要があります。これには、関数名を思い出すだけでは不十分です。モデルは、シグネチャ、モジュール パス、入出力コントラクト、セマンティクス、および実行可能ファイルの使用パターンを調整する必要があります。既存の新規 API ベンチマークは通常、静的であり、大まかな合否メトリクスに依存しているか、実際のライブラリの進化を反映していない可能性がある合成 API を使用しています。 NovelAPIBench は、あらゆるベース モデルおよびターゲット ライブラリに対して、新しい API を検出し、分解された知識バンドルを抽出し、実行可能なコーディング タスクを生成し、失敗したサンプルを 6 つの診断カテゴリに割り当てる、完全に自動化された動的ベンチマークです。約 1.9K のタスク、4 つの基本モデル、5 つのドメインにわたって、検索を通じて注入された知識と、パラメトリック適応を通じて内面化された知識を比較します。ナレッジコンポーネントは互換性がないことがわかりました。使用例は最も強力なスタンドアロンシグナルですが、最良の 2 コンポーネント設定は、ドメインとバックボーンに応じてメカニズムまたはサンプルのいずれかとシグネチャを組み合わせます。コンテキスト、特にソース コードを追加すると、インポート パスのエラーが増加して問題が発生する可能性があります。また、パラメトリック適応は、外部知識が除去された場合には検索に代わるものではありません。むしろ、微調整は主に提供されたバンドルの使用方法をモデルに教え、この機能は保持されたライブラリに転送されます。これらの結果は、取得とチューニングが補完的な役割を果たすことを示唆しています。取得は揮発性の API コンテンツを提供し、チューニングは手続き上の統合を改善します。

原文 (English)

Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition

Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a function name: models must coordinate signatures, module paths, input-output contracts, semantics, and executable usage patterns. Existing novel-API benchmarks are typically static, rely on coarse pass/fail metrics, or use synthetic APIs that may not reflect real library evolution. We introduce NovelAPIBench, a fully automated dynamic benchmark that, for any base model and target library, discovers novel APIs, extracts decomposed knowledge bundles, generates executable coding tasks, and assigns failed samples to six diagnostic categories. Across about 1.9K tasks, four base models, and five domains, we compare knowledge injected through retrieval with knowledge internalized through parametric adaptation. We find that knowledge components are not interchangeable: usage examples are the strongest standalone signal, while the best two-component setting pairs signatures with either mechanisms or examples depending on the domain and backbone. Adding more context, especially source code, can hurt by increasing import-path errors. Parametric adaptation also does not replace retrieval once external knowledge is removed; rather, fine-tuning mainly teaches models how to use provided bundles, and this ability transfers to held-out libraries. These results suggest that retrieval and tuning play complementary roles: retrieval supplies volatile API content, while tuning improves procedural integration.

2026-06-03 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

答えから状態へ: 大規模言語モデルにおける化学推論の検証可能なプロセスレベルの評価

大規模な言語モデルが化学アシスタントとして使用されることが増えていますが、ほとんどの化学ベンチマークは依然として最終的な回答のみをスコアとしています。これにより、重大な故障モードが隠蔽されます。モデルは、その推論が化学ロジックに違反しているにもかかわらず、正しい分子、生成物、またはオプションを出力する可能性があります。 LLM ジャッジと人間のステップレベルのプロセス アノテーションはコストが高く、一貫性がなく、幻覚に対して脆弱であるため、既存のプロセス レベルの評価機能を拡張するのは困難です。 ChemCoTBench-V2 は、構造化され検証者がアドレス指定できる化学推論トレースを低コストで監査可能に評価するためのルール検証可能な診断ベンチマークです。これは、分子理解、分子編集、分子最適化、反応予測に及び、18 のレポートタスクにわたる 5,620 の評価サンプルを備えています。モデルは、専門家が設計したテンプレートで主要な中間ステップを公開する必要があり、それらのステップは決定論的な化学ルールでチェックされ、クローズドアンサータスクの場合は、別の LLM 審査員ではなく参照トレースが使用されます。オープンエンド分子最適化は、厳密なトレース マッチングではなく、Oracle で検証可能な状態制約を使用して評価されます。このベンチマークは、最終回答の正確性、テンプレートの遵守、専門家によって洗練された中間コミットメントに対する段階的な検証者の正確さという 3 つの個別のシグナルを報告します。フロンティア モデルの実験では、最終的な回答の成功と構造化推論の状態の一貫性の間には永続的なギャップがあることが明らかになりました。モデルは多くの場合、化学ステップ チェックに失敗しながらも要求された形式に従っているか、弱い裏付け推論で正しく回答することができます。 ChemCoTBench-V2 は、きめ細かいモデル比較を可能にし、トレースが最初に検証ツールに違反する具体的なステップを特定します。

原文 (English)

From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical failure mode: a model may output the correct molecule, product, or option while its reasoning violates chemical logic. Existing process-level evaluators are hard to scale because LLM judges and human step-level process annotation are costly, inconsistent, and vulnerable to hallucination. We introduce ChemCoTBench-V2, a rule-verifiable diagnostic benchmark for low-cost, auditable evaluation of structured, verifier-addressable chemical reasoning traces. It spans molecular understanding, molecule editing, molecular optimization, and reaction prediction, with 5,620 evaluation samples across 18 reporting tasks. Models must expose key intermediate steps in expert-designed templates, and those steps are checked with deterministic chemistry rules and, for closed-answer tasks, reference traces rather than another LLM judge. Open-ended molecular optimization is evaluated with oracle-verifiable state constraints rather than strict trace matching. The benchmark reports three separate signals: final-answer correctness, template adherence, and step-wise verifier correctness over expert-refined intermediate commitments. Experiments on frontier models reveal a persistent gap between final-answer success and structured-reasoning-state consistency: models often follow the requested format while failing chemical-step checks, or answer correctly with weak supporting reasoning. ChemCoTBench-V2 enables fine-grained model comparison and identifies the concrete step at which the trace first violates the verifier.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Acceptance-Test-Driven Evaluation Protocols for Business-Centric LLM Systems

Large language model (LLM) applications are increasingly expected to satisfy deterministic institutional requirements while relying on prob…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成エージェントロボティクスビジネス/資金調達

SCOPE: Real-Time Natural Language Camera Agent at the Edge

Deploying language-driven agents in robotics requires evaluations that reflect real-world task demands: natural-language instructions with…

2026-06-03 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following

The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current aut…

2026-06-03 13:00 JSTarXiv cs.AI画像/動画生成エージェントロボティクスハードウェア/半導体ビジネス/資金調達

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. I…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making

Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior acros…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts

Existing benchmarks for MLLM-generated web artifacts assess interaction through local evidence and miss the requirement-induced states and…

2026-06-03 13:00 JSTarXiv cs.AIビジネス/資金調達

AlphaEval: A Comprehensive and Efficient Evaluation Framework for Formula Alpha Mining

Formula alpha mining, which generates predictive signals from financial data, is critical for quantitative investment. Although various alg…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios

We present an in-depth evaluation of LLMs' ability to negotiate, a central business task requiring strategic reasoning, theory of mind, and…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

X-RAY: Mapping LLM Reasoning Capability via Formalized and Calibrated Probes

Large language models (LLMs) achieve promising performance, yet their ability to reason remains poorly understood. Existing evaluations lar…

2026-06-03 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capaci…

2026-06-03 13:00 JSTarXiv cs.AIビジネス/資金調達

ソーシャルテキストの関与と共鳴に関するコミュニティを意識した評価: ユーザー作成コンテンツの評価に関する人間中心の視点

従来のビデオ品質評価 (VQA) は、美的忠実度に限定的に焦点を当てており、ユーザー生成コンテンツ (UGC) の品質を定義する複雑な社会力学を見落としていました。この研究では、信号中心のメトリクスから人間中心の共鳴評価へのパラダイム シフトを提案します。 CASTER (Community-Aware Assessment of Social Textual Engagement and Resonance) を導入します。これは、UGC アイテムが視覚的な品質だけではなく、マルチモーダルな属性に基づいて肯定的なコミュニティの共鳴を達成しているかどうかを評価する新しいタスクです。これに対処するために、新しいソーシャル チェーン オブ 思考 (Social-CoT) メカニズムを導入する MEDEA (マルチモーダル エンゲージメント駆動型評価アーキテクチャ) を紹介します。従来の論理的な CoT とは異なり、Social-CoT は、質の高い判断を導き出す前に、マルチモーダルな視点取得を実行し、多様な視聴者のペルソナをインスタンス化して、集団的な認知的および感情的な反応 (つまり、「コミュニティの心」) をシミュレートします。 MEDEA は、推論パスが本物の人間の社会的認知に基づいていることを保証するために、ソーシャル アライメント報酬を使用した教師あり微調整とプロセス教師あり強化学習を含む 2 段階のアプローチを通じてトレーニングされています。このタスクをサポートするために、さまざまな UGC カテゴリをカバーする人間による注釈付きの包括的なベンチマークである CASTER-Bench をリリースします。実験では、MEDEA が CASTER-Bench の最先端のベースラインを大幅に上回り、実際のコミュニティのフィードバックと一致する、解釈可能で共感的な推論パスを提供することが実証されました。

原文 (English)

Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content Evaluation

Traditional Video Quality Assessment (VQA) focuses narrowly on aesthetic fidelity, overlooking the complex social dynamics that define quality in User-Generated Content (UGC). In this work, we propose a paradigm shift from signal-centric metrics to human-centric resonance assessment. We introduce CASTER (Community-Aware Assessment of Social Textual Engagement and Resonance), a new task that evaluates whether a UGC item achieves positive community resonance based on its multimodal attributes rather than visual quality alone. To address this, we present MEDEA (Multimodal Engagement-Driven Evaluation Architecture), which introduces a novel Social Chain-of-Thought (Social-CoT) mechanism. Unlike traditional logical CoT, Social-CoT performs multimodal perspective-taking, instantiating diverse viewer personas to simulate collective cognitive and emotional reactions (i.e., the "community mind") before deriving a quality judgment. MEDEA is trained via a two-stage approach involving supervised fine-tuning and process-supervised reinforcement learning with Social Alignment Reward to ensure reasoning paths are grounded in authentic human social cognition. To support this task, we release CASTER-Bench, a comprehensive human-annotated benchmark covering diverse UGC categories. Experiments demonstrate that MEDEA significantly outperforms state-of-the-art baselines on CASTER-Bench while providing interpretable and empathetic reasoning paths that align with real community feedback.

2026-06-03 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

ディープリサーチエージェントはどこで間違っているのでしょうか?エージェントの軌跡におけるスパンレベルのエラーの位置特定

ディープリサーチエージェントは、検索、ツールの使用、証拠の検査、回答の合成という長い軌跡を通じてタスクを解決します。最終的な回答に基づく評価では、エージェントが成功したかどうかがわかりますが、軌跡のどの部分が回答の信頼性を低下させるのかはわかりません。私たちは、深層研究エージェントのスパンレベルのエラー位置特定を研究します。 2 つのエージェント フレームワーク、3 つのバックボーン モデル、3 つのベンチマークから 2,790 の実際の軌跡を収集し、生のログをセマンティック スパンに変換し、LLM 支援の専門家レビューを通じて有害なエラー スパンに注釈を付けます。これらのアノテーションから、通常の探索、失敗した探索、暫定的な仮説、無害なノイズの間のエラー範囲を特定するための 1,000 インスタンスのベンチマークである TELBench を構築します。さらに、エージェントのクレームを追跡し、軌跡の証拠でサポートをチェックし、サポートされていないクレームや矛盾するクレームが回答経路に影響を与えるスパンをマークする、クレーム中心の監査フレームワークである DRIFT を提案します。モデル ファミリと監査フレームワークにわたる実験では、DRIFT によりスパンレベルのエラー位置特定と最初のエラーの精度が最大 30 パーセント ポイント改善されることが示されています。私たちの研究は、ディープリサーチエージェントの信頼性をプロセスレベルで把握することを可能にします。

原文 (English)

Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final answers shows whether an agent succeeds, but not which parts of the trajectory make the answer unreliable. We study span-level error localization for deep-research agents. We collect 2,790 real trajectories from two agent frameworks, three backbone models, and three benchmarks, convert raw logs into semantic spans, and annotate harmful error spans through LLM-assisted expert review. From these annotations, we build TELBench, a 1,000-instance benchmark for identifying error spans among normal exploration, failed searches, tentative hypotheses, and harmless noise. We further propose DRIFT, a claim-centric auditing framework that tracks agent claims, checks their support in trajectory evidence, and marks spans where unsupported or conflicting claims affect the answer path. Experiments across model families and auditing frameworks show that DRIFT improves span-level error localization and first-error accuracy by up to 30 percentage points. Our work provides a process-level view of reliability in deep-research agents.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AGENTCL: 言語エージェントの継続学習の厳格な評価に向けて

言語エージェントは、個々のタスクの解決にかなりの推論時間を費やしますが、あるエピソードで得た経験は、その後のエピソードでは十分に活用されないことがよくあります。継続的な学習では、エージェントが一連のタスク全体にわたって再利用可能なエクスペリエンスを蓄積し、時間の経過とともに改善し、無関係なエクスペリエンスによる干渉を回避することが期待されます。残念ながら、既存のベンチマークでは、言語エージェントの継続的な学習を厳密に評価するのが困難です。ほとんどの取り組みは、長いコンテキストの会話やドキュメントの検索と推論に焦点を当てていますが、最近の生涯適応ベンチマークは、タスク間の関係の分析が限られた単純なタスク ストリームに依存していることが多く、エージェントが時間の経過とともに何を学習し、再利用するかを理解することが困難になっています。この論文では、制御されたタスク ストリームと転送ゲインのメトリクスを中心とした、エージェントの継続的学習のための評価フレームワーク AgentCL を紹介します。 AGENTCL は、以前のサブソリューション、証拠、またはワークフローが後のタスクで意図的に再利用可能な構成ストリームを構築し、そのような再利用性が保証されていない単純なストリームと対比します。ベンチマークを使用して、継続学習のためのノンパラメトリック メモリ設計を評価します。メモリ設計の選択が継続学習にどのような影響を与えるかを診断するために、統合中に信頼性の低いエクスペリエンスをフィルタリングしながら、インタラクション、洞察、スキルを保存するプローブ手法である MemProbe を開発しました。コーディング、詳細な研究、および言語理解/推論タスクにわたる実証分析では、単純なストリームではメモリ設計を区別する能力が限られているのに対し、制御されたストリームではその可塑性をより明確に区別できることが示されています。一方、単純な設定や保持された設定では、得られるゲインが限られていることが多く、メモリによる劣化が生じる可能性があります。これらの結果は、可塑性と安定した再利用のバランスをとった、より強力なメモリ設計の必要性を浮き彫りにしています。

原文 (English)

AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes. Continual learning expects an agent to accumulate reusable experience across a stream of tasks, improve over time, and avoid interference from irrelevant experiences. Unfortunately, existing benchmarks struggle to evaluate continual learning in language agents rigorously. Most efforts focus on retrieval and reasoning over long-context conversations or documents, while recent lifelong-adaptation benchmarks often rely on naive task streams with limited analysis of cross-task relationships, making it difficult to understand what an agent learns and reuses over time. This paper presents an evaluation framework AgentCL for continual learning in agents, centered on controlled task streams and metrics for transfer gains. AgentCL constructs compositional streams where earlier sub-solutions, evidence, or workflows are intentionally reusable in later tasks, and contrasts them with naive streams where such reusability is not guaranteed. We use the benchmark to evaluate non-parametric memory designs for continual learning. To diagnose how memory design choices affect continual learning, we develop MemProbe, a probing method that stores interactions, insights, and skills, while filtering unreliable experiences during consolidation. Empirical analysis across coding, deep research, and language understanding/reasoning tasks shows that naive streams offer limited ability to distinguish memory designs, whereas controlled streams more clearly distinguish their plasticity. Meanwhile, naive and held-out settings often yield limited gains and can expose memory-induced degradation. These results highlight the need for stronger memory designs that balance plasticity and stable reuse.

2026-06-03 13:00 JSTarXiv cs.AIビジネス/資金調達

Building Trust in Black-box Optimization: A Comprehensive Framework for Explainability

Optimizing costly black-box functions within a constrained evaluation budget presents significant challenges in many real-world application…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達研究/論文

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content. However, existing research and eva…

2026-06-03 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

PHASE: Physiology-Aware Hyperspectral Reconstruction via Object-to-Human Domain Adaptation

Although hyperspectral imaging offers unparalleled non-invasive physiological insight, its bulky hardware, slow acquisition, and regulatory…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

The transition from monolithic language models to modular, skill-equipped agents marks a defining shift in how large language models (LLMs)…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming Solutions

The evaluation of Large Language Models (LLMs) for code generation relies heavily on the quality and robustness of test cases. However, exi…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Quantifying and Mitigating Self-Preference Bias of LLM Judges

LLM-as-a-Judge has become a dominant approach in automated evaluation systems, playing critical roles in model alignment, leaderboard const…

2026-06-03 13:00 JSTarXiv cs.AIビジネス/資金調達

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation

Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape…

2026-06-03 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

AgentLens: SWE エージェント評価におけるラッキー パスの問題を明らかにする

更新された要約は次のとおりです。 ソフトウェア エンジニアリング (SWE) エージェントの評価は、最終パッチがテストに合格するかどうかという 2 つの信号によって支配されます。この結果のみの考え方は、原則に基づいた解決策と混沌とした試行錯誤のプロセスを同等のものとして扱います。この等価性は経験的に誤りであることを示します。 60 の SWE ベンチ検証済みタスクで 8 つのモデル バックエンドからの 2,614 の OpenHands 軌跡を評価します。これらのうち、47 にはタスク レベルのプロセス参照を構築するのに十分な通過軌跡があり、1,815 の軌跡評価サブセットが得られます。このサブセットの通過軌跡のうち、10.7% は、回帰サイクル、ブラインド再試行、検証の欠落、または時間的に無秩序な探索、実装、検証など、ラッキー パスと呼ばれる動作を示しています。 SWE エージェントの軌跡をプロセスレベルで評価するためのフレームワークである AgentLens を導入し、品質スコア、廃棄信号、分岐点、および 47 のタスクレベルのプレフィックス ツリー アクセプタ (PTA) 参照が注釈付けされた 1,815 の軌跡のデータセットである AgentLens-Bench を定義します。 AgentLens は、同じタスクに渡された複数のソリューションをマージすることで PTA 参照を構築し、コンテキスト依存のインテント ラベラーを使用して、ツール ID だけではなく軌跡履歴に基づいて探索、実装、検証、またはオーケストレーションにアクションを割り当てます。 AgentLens-Bench では、品質スコアによってパスの軌跡が Lucky、Solid、Ideal の層に分割され、さらに Lucky パスが 5 つの反復メカニズムに分解されます。 8 つのモデル バックエンド全体で、Lucky 率の範囲は 0.5% ~ 23.2% であり、一部のモデルは、合格率ではなく品質スコアでランク付けすると、最大 5 ランク順位が変動します。 AgentLens-Bench アーティファクト、AgentLens SDK、分析ツールを含むプロジェクト リポジトリを間もなくリリースする予定です。

原文 (English)

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view treats a principled solution and a chaotic trial-and-error process as equivalent. We show that this equivalence is empirically false. We evaluate 2,614 OpenHands trajectories from eight model backends on 60 SWE-bench Verified tasks. Of these, 47 have enough passing trajectories to construct task-level process references, yielding a 1,815-trajectory evaluation subset. Among passing trajectories in this subset, 10.7% exhibit behavior we call a Lucky Pass: regression cycles, blind retries, missing verification, or temporally disordered exploration, implementation, and verification. We introduce AgentLens, a framework for process-level assessment of SWE-agent trajectories, and define AgentLens-Bench, a dataset of 1,815 trajectories annotated with quality scores, waste signals, divergence points, and 47 task-level Prefix Tree Acceptor (PTA) references. AgentLens builds PTA references by merging multiple passing solutions for the same task, and uses a context-sensitive intent labeler to assign actions to Exploration, Implementation, Verification, or Orchestration based on trajectory history rather than tool identity alone. On AgentLens-Bench, the quality score separates passing trajectories into Lucky, Solid, and Ideal tiers and further decomposes Lucky Passes into five recurring mechanisms. Across the eight model backends, Lucky rates range from 0.5% to 23.2%, and some models move by as many as five rank positions when ranked by quality score instead of pass rate. We plan to release the project repository soon, including AgentLens-Bench artifacts, the AgentLens SDK, and the analysis tooling.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

評価意識の分解と測定

フロンティア言語モデルは、評価されていることを認識して動作を調整し、ベンチマーク結果の妥当性を損なうことがあります。しかし、現場では共通の基礎を持たずに評価の特性とモデルの特性、検出と行動反応を混同して研究が行われています。私たちは評価意識を社会心理学に基礎づけ、評価意識を環境要素 (課題がどの程度認識されているか) と、認識をそれに基づいて行動する傾向から分離するモデル要素に分解します。プレースホルダー エンティティや採点スタイルの出力形式など、8 つの分類されたトリガー要因を通じて環境コンポーネントを運用し、思考連鎖のモニタリングを通じて認識と行動を研究します。 9 つのフロンティア モデルと 4 つのベンチマークにわたって、認識率はモデルとベンチマークのどちらか単独ではなく、モデルとベンチマークの特定の組み合わせに依存します。認識が行動の変化につながることはほとんどありませんが、変化する場合、その方向性は認識された評価の種類によって異なります。また、モデルは機能評価よりも安全性に対して敏感であり、安全性ベンチマークの妥当性がより大きなリスクにさらされます。各モデルがどの要因に敏感で、それらがどのように相互作用するかを研究するために、8 つの要因のそれぞれを独立して切り替えることができ、基礎となる要求を固定したまま評価信号を変化させる、100 のペアの安全機能タスクの要因制御ベンチマークである \textbf{EvalAwareBench} を提案します。 EvalAwareBench を通じて、単一の要素がすべてのモデルに均一に影響を与えることはなく、要素を積み重ねることですべてのモデルにわたる評価の意識が徐々に向上することがわかりました。私たちのフレームワークと EvalAwareBench は、評価意識を測定、属性付け、軽減するためのツールを提供し、将来有望な道として認識される下での行動の一貫性を示します。

原文 (English)

Decomposing and Measuring Evaluation Awareness

Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet the field studies it without a shared foundation, conflating properties of the evaluation with properties of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component (how recognizable the task is) and a model component that separates recognition from propensity to act on it. We operationalize the environment component through eight categorized trigger factors, such as placeholder entities and grading-style output formats, and study recognition and behavior through chain-of-thought monitoring. Across nine frontier models and four benchmarks, recognition rates depend on the specific pairing of model and benchmark rather than on either in isolation. Recognition rarely leads to behavioral change, and when it does, the direction depends on the type of evaluation perceived. Models are also more sensitive to safety than capability evaluations, placing safety benchmark validity at greater risk. To study which factors each model is sensitive to and how they interact, we propose \textbf{EvalAwareBench}, a factor-controlled benchmark of 100 paired safety-capability tasks where each of the eight factors can be independently toggled, varying evaluative signals while holding the underlying request fixed. Through EvalAwareBench, we find that no single factor uniformly affects all models, but stacking factors progressively raises evaluation awareness across all of them. Our framework and EvalAwareBench provide the tools to measure, attribute, and mitigate evaluation awareness, pointing to behavioral consistency under recognition as a promising path forward.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

JudgmentBench: 品質評価のためのルーブリックと優先評価の比較

現在のベンチマーク手法は 2 つの方法論が主流となっています。ルーブリック ベースのスコアリングでは、事前定義された基準に照らして項目を評価します。一方、比較判断では、出力間のペアごとの優先順位を導き出します。どちらの方法論も広く使用されていますが、どちらを選択するかが正当化されることはほとんどありません。当社は、米国の大手法律事務所を含む豊富な経験を持つ現役弁護士から収集した 1,539 のルーブリック スコアと 1,530 のペアごとの優先判断を組み合わせた、30 の実際の法律タスクのベンチマークである JudgmentBench をリリースします。アノテーションは、両方の監視シグナルが同じ項目について同じ専門家から引き出される、高度な専門知識の領域における最初の公的に利用可能なデータセットを構成します。 3 つの構築された品質レベルで LLM によって生成された出力を使用して、最初の経験的比較を提供します。比較判断により、ルーブリックよりも大幅に意図した品質順序が回復されます (スピアマンの順位相関の平均 0.908 対 0.150、推定差 = 0.758 [0.494, 1.021])。必要なアノテーション時間は半分未満です。このパターンは、ヒューマン アノテーターと LLM 自動採点者にも当てはまります。この最初の比較を超えて、データセットのペア構造は、検証可能なグラウンドトゥルースのない領域で専門家の判断をどのように導き出し、集約し、監視として使用するかについてのより広範な研究課題をサポートします。

原文 (English)

JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs. Although both methodologies are widely used, the choice between them is rarely justified. We release JudgmentBench, a benchmark of 30 real-world legal tasks, paired with 1,539 rubric scores and 1,530 pairwise preference judgments collected from practicing attorneys--including at major U.S. law firms--with substantial experience. The annotations constitute the first publicly available dataset in a high-expertise domain in which both supervision signals are elicited from the same experts on the same items. Using LLM-generated outputs at three constructed quality levels, we provide an initial empirical comparison: comparative judgments recover the intended quality ordering substantially better than rubrics (mean Spearman's rank correlation of 0.908 vs. 0.150, estimated difference = 0.758 [0.494, 1.021]) while requiring less than half the annotation time. The patterns hold for human annotators and LLM autograders. Beyond this initial comparison, the paired structure of the dataset supports a broader research agenda on how expert judgment should be elicited, aggregated, and used as supervision in domains without verifiable ground truth.

2026-06-03 13:00 JSTarXiv cs.AIビジネス/資金調達

SL-BiLEM: 予測と政策評価のための構造化された学習可能な Behavior-in-the-Loop 流行モデリング

流行予測は根本的な課題に直面しています。それは、人間の行動が病気の蔓延に動的に反応し、政策介入時点で分布の変化を引き起こすフィードバック ループを生み出すということです。これにより、分布の変化の下ではデータ駆動型モデルの信頼性が低くなります。私たちは、堅牢な外挿のための正則化として物理的制約を利用する \textbf{SL-BiLEM} (構造化学習可能な行動インザループ流行モデル) を提案します。このフレームワークは、有効な伝送を $\beta_{\text{eff}}(t,g) = \beta_0(g) \times m_{\text{policy}}(t) \times m_{\text{media}}(t) \times m_{\text{comp}}(t,g)$ として分解します。ここで、学習されたコンプライアンス関数の単調性、滑らかさ、および境界ジャンプ制約は、新しいポリシーの下で予測の妥当性を維持します。体制。 SL-BiLEM は予測を超えて、介入の意思決定をサポートするための反事実分析を可能にします。私たちは 3 つの現実世界のデータセット (クルーズ船、学校インフルエンザ、学校区の 新型コロナウイルス感染症 (COVID-19) 監視) で予測を検証し、既知のグラウンド トゥルースを使用した合成ベンチマークで反事実の回復を評価します。 SL-BiLEM は次のことを実証します。(1) 神経機構ベースラインに対して 76% 改善し、OOD 低下はわずか 53% であったのに対し、政策誘発シフト下の神経ベースラインでは 1142% でした。 (2) 27 の合成反事実実験にわたる 100% のブートストラップ CI カバレッジ。 (3) 治療効果の精度が 0.85 を超える。これらの結果により、SL-BiLEM は、正確な予測と原則に基づいた介入計画を求める公衆衛生の意思決定者にとって解釈可能なツールとして確立されています。

原文 (English)

SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation

Epidemic forecasting faces a fundamental challenge: human behavior dynamically responds to disease spread, creating feedback loops that induce distribution shifts at policy intervention points. This renders data-driven models unreliable under distribution shift. We propose \textbf{SL-BiLEM} (Structured Learnable Behavior-in-the-Loop Epidemic Model), leveraging physical constraints as regularization for robust extrapolation. The framework decomposes effective transmission as $\beta_{\text{eff}}(t,g) = \beta_0(g) \times m_{\text{policy}}(t) \times m_{\text{media}}(t) \times m_{\text{comp}}(t,g)$, where monotonicity, smoothness, and bounded-jump constraints on the learned compliance function maintain predictive validity under novel policy regimes. Beyond forecasting, SL-BiLEM enables counterfactual analysis for intervention decision support. We validate forecasting on three real-world datasets (cruise ship, school influenza, and school-district COVID-19 surveillance) and evaluate counterfactual recovery on synthetic benchmarks with known ground truth. SL-BiLEM demonstrates: (1) 76\% improvement over neural-mechanistic baselines, with only 53\% OOD degradation versus 1142\% for neural baselines under policy-induced shift; (2) 100\% bootstrap CI coverage across 27 synthetic counterfactual experiments; and (3) Treatment Effect Accuracy exceeding 0.85. These results establish SL-BiLEM as an interpretable tool for public health decision-makers seeking accurate prediction and principled intervention planning.

2026-06-03 07:50 JSTTechCrunch AIビジネス/資金調達

Cyera eyes $12B valuation at 80x ARR multiple despite operating losses

The cybersecurity company is nearing a $300 million round led by Evolution Equity Partners.

2026-06-03 04:02 JSTTechCrunch AIビジネス/資金調達

New Microsoft tool lets devs spin up AI behavior tests using text descriptions

Microsoft on Tuesday took the wraps off Adaptive Spec-driven Scoring for Evaluation and Regression Testing, an open source framework for sp…

2026-06-02 21:32 JSTTechCrunch AIビジネス/資金調達

ZeroDrift raises $10M to protect AI models from themselves

A new AI compliance service sits between AI models and end users to flag and replace any messages that might present a compliance problem.

2026-06-02 21:00 JSTTechCrunch AIビジネス/資金調達

Rocket engine startup Impulse raises $500 million to hire people, not AI

Engineering physical systems still depends on human talent, according to Impulse Space president Eric Romo.

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

大規模言語モデルにおける対話型推論の評価: 実行可能ゲームによる階層ベンチマーク

推論を積極的な証拠の取得と信念の更新として扱う推論評価のためのマルチターン対話型フレームワークを紹介します。ここで、LLM はタスク ルールのみを受け取り、対象を絞ったクエリを非表示の環境に発行し、部分的な観察を時間の経過とともに統合し、最終的な回答をいつ送信するかを決定する必要があります。標準的な成功率とインタラクション効率を超えて、制御された文脈の摂動下での文脈の堅牢性、および反事実の修正と必要性の判断によるメタ認知の適応を評価します。 474 の実行可能ゲームのベンチマークとしてフレームワークをインスタンス化し、それぞれを 5 つの難易度に対応する 5 つの固定構成検索スペースで評価し、広範なフロンティア LLM セットを評価します。結果は、ベンチマークが非常に識別力があり、成功率だけでなくインタラクション効率にも大きな違いがあることを示しています。さらに、文脈の混乱は中程度ではあるが一貫した低下を引き起こす一方、反事実の修正や必要性の判断はさらに大きな低下を引き起こすことを経験的に示しています。

原文 (English)

Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games

We introduce a multi-turn interactive framework for reasoning evaluation that treats reasoning as active evidence acquisition and belief updating. Wherein, LLMs receive only the task rules, must issue targeted queries to a hidden environment, integrate partial observations over time, and decide when to submit a final answer. Beyond standard success rate and interaction efficiency, we evaluate contextual robustness under controlled contextual perturbations, and metacognitive adaptation through counterfactual revision and necessity judgment. We instantiate the framework as a benchmark of 474 executable games, each evaluated under five fixed configuration search spaces corresponding to five difficulty levels, and evaluate a broad set of frontier LLMs. Results show that the benchmark is highly discriminative, exposing large differences not only in success rate but also in interaction efficiency. Moreover, we empirically show that contextual perturbations cause moderate but consistent declines, whereas counterfactual revision and necessity judgment lead to much larger drops.

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

国家学習能力としての AI 主権: フランス、米国、中国に関する人間中心の学習力学の視点

フランスでは、人工知能は、投資、計算能力、規制、雇用、主権、教育の観点からよく議論されます。通常、これらのディメンションは個別に扱われます。この観点に関する論文は、統一的な解釈を提案しています。つまり、フランスは \emph{国家的な AI 学習システム} として理解されるべきです。エントロピー制御された表現学習のための動的フレームワークとして最近策定された人間中心学習力学 (HCLM) に基づいて、私たちは国家 AI 開発を情報注入とエントロピー散逸の間の制御されたバランスとして解釈します。情報注入は、コンピューティング、データ、人材、研究、資本、産業展開、および組織的実験に対応します。エントロピー散逸は、組織の複雑さ、調整摩擦、エネルギー制約、規制の不確実性、人材の流動性の圧力、産業吸収を強化する機会に対応します。中心的な主張は、AI の主権は規模だけから生まれるのではなく、自国の情報ダイナミクスを規制する国の能力から生まれるというものです。この論文は、HCLM をニューラル スケーリング則、内生的成長理論、創造的破壊、およびゲーム理論と結びつけます。同論文は、フランスのAI論争は、技術楽観主義と規制優先の慎重論という二項対立を超えて進むべきだと主張している。競争力のある人間中心の AI 戦略には、不安定、不平等、またはエネルギー集約的な拡大を回避しながら、情報注入が制度的消散よりも早く成長する制御された体制が必要です。私たちは、数学的モデル、測定可能な政策指標、ゲーム理論的命題、国家 AI 体制の具体的なシミュレーション、およびフランスに対する具体的な政策への影響を提供します。提案された視点は、AI 政策をオープンで戦略的な非平衡学習システムのガバナンスとして再構成します。

原文 (English)

AI Sovereignty as National Learning Capacity: A Human-Centered Learning Mechanics Viewpoint on France, the United States, and China

Artificial Intelligence is often discussed in France in terms of investment, compute capacity, regulation, employment, sovereignty, and education. These dimensions are usually treated separately. This viewpoint paper proposes a unified interpretation: France should be understood as a \emph{national AI learning system}. Building on Human-Centered Learning Mechanics (HCLM), recently formulated as a dynamical framework for entropy-regulated representation learning, we interpret national AI development as a controlled balance between information injection and entropy dissipation. Information injection corresponds to compute, data, talent, research, capital, industrial deployment, and institutional experimentation. Entropy dissipation corresponds to organizational complexity, coordination frictions, energy constraints, regulatory uncertainty, talent mobility pressures, and opportunities to strengthen industrial absorption. The central claim is that AI sovereignty does not emerge from scale alone but from a country's capacity to regulate its own information dynamics. This paper connects HCLM with neural scaling laws, endogenous growth theory, creative destruction, and game theory. It argues that the French AI debate should move beyond the binary opposition between techno-optimism and regulation-first caution. A competitive and human-centered AI strategy requires a controlled regime in which information injection grows faster than institutional dissipation, while avoiding unstable, unequal, or energy-intensive expansion. We provide a mathematical model, measurable policy indicators, game-theoretic propositions, illustrative simulations of national AI regimes, and concrete policy implications for France. The proposed viewpoint reframes AI policy as the governance of an open, strategic, non-equilibrium learning system.

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

強化学習一般化の証明書に基づく評価

この研究では、目に見えないタスクを一般化する能力における強化学習 (RL) アルゴリズムのパフォーマンスを評価するためのロジック主導のフレームワークを紹介します。私たちのフレームワークは、タスクのダイナミクスの構造的類似性を特徴とする帰納的リーチ回避タスクのファミリーを定義し、汎化機能の評価を可能にします。重要な条件を強制することで RL アルゴリズムによって生成された軌跡を検証するニューラル証明書関数を導入します。これにより、RL の一般化に対するリトマス試験紙として機能します。私たちは、困難な連続環境において、いくつかの最先端の一般化可能な RL アルゴリズムの一般化を証明する際の私たちの方法の能力を経験的に実証します。私たちの結果は、証明書機能違反の割合が低いほど、成功したテスト タスクの数が多いことと相関していることを示しており、RL アルゴリズムの一般化機能を評価および区別する際のフレームワークの有効性が強調されています。この研究は、RL の一般化をベンチマークするための原則に基づいたアプローチを提供します。

原文 (English)

Certificate-Guided Evaluation of Reinforcement Learning Generalization

This work presents a logic-driven framework to evaluate the performance of reinforcement learning (RL) algorithms in their ability to generalize to unseen tasks. Our framework defines a family of inductive reach-avoid tasks, characterized by structural similarities in task dynamics, enabling evaluation of generalization capabilities. We introduce a neural certificate function that validates trajectories generated by RL algorithms by enforcing key conditions, thereby serving as a litmus test for RL generalization. We empirically demonstrate our method's capability in certifying generalization for several state-of-the-art generalizable RL algorithms on challenging continuous environments. Our results show that a lower percentage of certificate function violations correlates with a higher number of test tasks successfully solved, highlighting the effectiveness of our framework in evaluating and distinguishing generalization capabilities of RL algorithms. This work provides a principled approach for benchmarking RL generalization.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

人工理性の謎: 大規模な推論モデルにおける生産と評価のギャップの調査

人間の推論に関する研究によると、人は通常、推論をゼロから生み出すよりも推論を評価する方が得意であることがわかっています。対照的に、大規模推論モデル (LRM) は、複雑な問題を解決するための長い推論チェーンを生成することに優れるようにトレーニングされています。では、LRM は理由を評価する際にどのように機能するのでしょうか?私たちはこれを Valid-Answer-Invalid-Reasoning (VAIR) データセットを使用して調査します。これは、推論生成の混乱から推論評価を分離するように設計された、些細な推論上の欠陥があるものの有効な答えを持つ数学の問題と解決策です。人間とは異なり、このような問題を解くよりもグレーディングがわずか 6% 悪いだけですが、LRM では生産と評価に大きなギャップがあることがわかりました。フロンティア モデルは、ほぼ完璧なソリューションの生産にも関わらず、VAIR ソリューションを評価する際のスコアは 48% と低かったのです。なぜこの謎が?思考連鎖 (CoT) 分析を通じて、回答確証バイアスの証拠が見つかりました。LRM は、各ステップを注意深く検証するのではなく、正解を生成してからチェックすることが多く、異常な推論に気づいた場合でも合理化を捏造します。線形プローブはこれを裏付けており、LRM アクティベーションは有効な推論の何らかの表現をエンコードする一方で、VAIR ソリューションを無効なものとして確実に表現できないことを示しています。最終的な答えの表現に因果関係をパッチすることにより、LRM の判定とアクティベーションが反転し、答えの妥当性がモデルの確証バイアスの原因であることがわかります。これらの発見は、LRM に正しい答えに向けた推論を作成および確認するよう促すが、根本的な理由をしっかりと評価することを促す推論トレーニングへの支配的なアプローチには顕著な限界があることを示しています。

原文 (English)

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch. In contrast, large reasoning models (LRMs) are trained to excel at producing long chains of reasoning to solve complex problems. How then do LRMs perform at evaluating reasons? We investigate this with the Valid-Answer-Invalid-Reasoning (VAIR) dataset: math problems and solutions with trivial reasoning flaws but valid answers, designed to isolate reasoning evaluation from the confound of reasoning production. Unlike humans, who we find are only 6% worse at grading than solving such problems, we find a substantial production-evaluation gap in LRMs: frontier models score as low as 48% when evaluating VAIR solutions, despite near-perfect solution production. Why this enigma? Through chain-of-thought (CoT) analysis, we find evidence of an answer confirmation bias: LRMs often produce then check for the correct answer instead of carefully verifying each step, fabricating rationalizations even when noticing anomalous reasoning. Linear probes corroborate this, showing that while LRM activations encode some representation of valid reasoning, they fail to robustly represent VAIR solutions as invalid. Causal patching of the final answer's representations causes LRM verdicts and activations to flip, demonstrating that answer validity is responsible for models' confirmation biases. These findings indicate an outstanding limitation in dominant approaches to reasoning training, which incentivize LRMs to produce and confirm reasoning towards correct answers, but not to robustly evaluate the underlying reasons.

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

リアルタイムの感情駆動型音響化のためのミニマリストのブレインコンピューター音楽インターフェイス: システム設計と予備評価

この論文では、リアルタイムの感情的音響化システムとして機能し、前頭前野の脳波活動を適応的な音楽に変換する、ミニマリストのブレインコンピューターミュージカルインターフェイス (BCMI) を紹介します。感情価は正面アルファ非対称性 (AF7/AF8) から推定され、確率的生成アルゴリズムを通じてモード、テンポ、リズム密度、ピッチ音域などの音楽的特徴にマッピングされます。このシステムは、ワイヤレス EEG 取得、リアルタイム Python 信号処理、Lab Streaming Layer 経由で同期される Ableton Live ベースの音楽生成を統合しています。 22人の参加者を対象とした実験では、意図的な感情的自己誘導がBCMIニューロフィードバック信号を調節できるかどうかを調査した。線形混合効果分析では、対象の感情や時間の有意な影響は見出されず、正面アルファ非対称信号が指示された感情状態を確実に区別していないことが示されました。音楽トレーニングや演技経験などの個人差は、信号分散全体のわずか 0.40% を占める実験操作よりも大きな分散を説明します。これらの発見は、正面アルファ非対称性を閉ループ感情制御の自発的制御信号として使用することの課題を浮き彫りにし、将来のBCMI研究の方法論的な方向性を示唆しています。

原文 (English)

A Minimalist Brain-Computer Musical Interface for Real-Time Emotion-Driven Sonification: System Design and Preliminary Evaluation

This paper presents a minimalist brain-computer Musical Interface (BCMI) that functions as a real-time affective sonification system, translating prefrontal EEG activity into adaptive music. Emotional valence is estimated from frontal alpha asymmetry (AF7/AF8) and mapped to musical features such as mode, tempo, rhythmic density, and pitch register through a stochastic generative algorithm. The system integrates wireless EEG acquisition, real-time Python signal processing, and Ableton Live-based music generation synchronized via Lab Streaming Layer. An experiment with 22 participants investigated whether intentional emotional self-induction could modulate the BCMI neurofeedback signal. Linear mixed-effects analyses found no significant effects of target emotion or time, indicating that the frontal alpha asymmetry signal did not reliably distinguish instructed emotional states. Individual differences, including musical training and acting experience, explained more variance than the experimental manipulation, which accounted for only 0.40\% of total signal variance. These findings highlight the challenges of using frontal alpha asymmetry as a voluntary control signal for closed-loop emotion regulation and suggest methodological directions for future BCMI research.

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

因果関係発見に使用されるベンチマークの整合性評価

グラフィカル因果モデルでは、因果発見は数値データと平文のドメイン知識に基づいて因果グラフを構築することを目的としています。ただし、ドメイン研究の進歩により、ベンチマーク因果グラフに誤った知識が含まれることが多くなっているため、この分野では因果関係発見手法の評価が依然として課題となっています。この問題は、大規模言語モデル (LLM) ベースの因果発見手法の評価に特に影響を与えます。これは、LLM が文献内の新しい発見に敏感であるためです。この研究は、ベンチマーク因果グラフの品質を体系的に研究した最初の研究です。具体的には、科学データベースから関連する研究論文を自動的に取得し、LLM にベンチマークの因果関係グラフとドメインの研究論文の間の一貫性をチェックするよう促すパイプラインを設計します。私たちは 11 の人気のある現実世界のベンチマークを評価し、そのベンチマークに対して私たちのパイプラインは合計 38,081 件のドメイン論文を進めています。私たちの結果は、人気のあるベンチマークはドメイン研究との整合性において大きく異なり、因果関係発見研究に明らかな影響を与えることを示しています。

原文 (English)

Consistency evaluation of benchmarks used for causal discovery

In graphical causal model, causal discovery aims to construct a causal graph based on numerical data and domain knowledge in plain text. However, the evaluation of causal discovery methods remains a challenge in the area as the progress of domain researches often makes benchmark causal graphs contain mis-aligned knowledge. This problem especially affects the evaluation of large language model (LLM) based causal discovery methods as they are sensitive to the new discoveries in the literature. This work is the first to systematically study the quality of benchmark causal graphs. Specifically, we design a pipeline that automatically retrieves relevant research papers from scientific databases, and prompts LLMs to check the consistency between the benchmark causal graphs and domain research papers. We evaluate 11 popular real-world benchmarks, for which our pipeline in total proceeds 38,081 domain papers. Our results show that popular benchmarks vary significantly in their consistency with domain research, with clear implications for causal discovery research.

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

IDDベースのSSD外部メモリ検索のベースライン手法の評価

多くの難しい検索問題は、RAM のみを使用する A* などのアルゴリズムでは解決できません。 RAMよりもはるかに大容量のSSDやHDDなどの外部メモリを使用する検索アルゴリズムは以前の研究で提案されていますが、以前の研究では遅延重複検出アプローチと複雑な即時重複検出(IDD)方法に焦点が当てられており、IDDのための比較的単純な方法は系統的に研究されていませんでした。さらに、ページ キャッシュなど、外部メモリへのアクセスを管理および高速化するための OS レベルのメカニズムの効果については研究されていません。この論文では、IDD ベースの A* に対する単純なベースライン アプローチのパフォーマンスを評価および分析することで、文献のギャップに対処します。

原文 (English)

Evaluation of Baseline Methods for IDD-based SSD External Memory Search

Many difficult search problems cannot be solved by algorithms such as A* using only RAM. Search algorithms which use external memory such as SSDs and HDDs with much higher capacity than RAM have been proposed in previous work, but previous work has focused on delayed duplicate detection approaches, as well as complex immediate duplicate detection (IDD) methods, and relatively simple methods for IDD have not been systematically studied. In addition, the effect of OS-level mechanisms for managing and speeding up accesses to external memory, such as page caches, has not been studied. This paper addresses these gaps in the literature by evaluating and analyzing the performance of simple baseline approaches for IDD-based A*.

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

ソーシャルテキストの関与と共鳴に関するコミュニティを意識した評価: ユーザー作成コンテンツの評価に関する人間中心の視点

従来のビデオ品質評価 (VQA) は、美的忠実度に限定的に焦点を当てており、ユーザー生成コンテンツ (UGC) の品質を定義する複雑な社会力学を見落としていました。この研究では、信号中心のメトリクスから人間中心の共鳴評価へのパラダイム シフトを提案します。 CASTER (Community-Aware Assessment of Social Textual Engagement and Resonance) を導入します。これは、UGC アイテムが視覚的な品質だけではなく、マルチモーダルな属性に基づいて肯定的なコミュニティの共鳴を達成しているかどうかを評価する新しいタスクです。これに対処するために、新しいソーシャル チェーン オブ 思考 (Social-CoT) メカニズムを導入する MEDEA (マルチモーダル エンゲージメント駆動型評価アーキテクチャ) を紹介します。従来の論理的な CoT とは異なり、Social-CoT は、質の高い判断を導き出す前に、マルチモーダルな視点取得を実行し、多様な視聴者のペルソナをインスタンス化して、集団的な認知的および感情的な反応 (つまり、「コミュニティの心」) をシミュレートします。 MEDEA は、推論パスが本物の人間の社会的認知に基づいていることを保証するために、ソーシャル アライメント報酬を使用した教師あり微調整とプロセス教師あり強化学習を含む 2 段階のアプローチを通じてトレーニングされています。このタスクをサポートするために、さまざまな UGC カテゴリをカバーする人間による注釈付きの包括的なベンチマークである CASTER-Bench をリリースします。実験では、MEDEA が CASTER-Bench の最先端のベースラインを大幅に上回り、実際のコミュニティのフィードバックと一致する、解釈可能で共感的な推論パスを提供することが実証されました。

原文 (English)

Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Perspective on User-Generated Content Evaluation

Traditional Video Quality Assessment (VQA) focuses narrowly on aesthetic fidelity, overlooking the complex social dynamics that define quality in User-Generated Content (UGC). In this work, we propose a paradigm shift from signal-centric metrics to human-centric resonance assessment. We introduce CASTER (Community-Aware Assessment of Social Textual Engagement and Resonance), a new task that evaluates whether a UGC item achieves positive community resonance based on its multimodal attributes rather than visual quality alone. To address this, we present MEDEA (Multimodal Engagement-Driven Evaluation Architecture), which introduces a novel Social Chain-of-Thought (Social-CoT) mechanism. Unlike traditional logical CoT, Social-CoT performs multimodal perspective-taking, instantiating diverse viewer personas to simulate collective cognitive and emotional reactions (i.e., the "community mind") before deriving a quality judgment. MEDEA is trained via a two-stage approach involving supervised fine-tuning and process-supervised reinforcement learning with Social Alignment Reward to ensure reasoning paths are grounded in authentic human social cognition. To support this task, we release CASTER-Bench, a comprehensive human-annotated benchmark covering diverse UGC categories. Experiments demonstrate that MEDEA significantly outperforms state-of-the-art baselines on CASTER-Bench while providing interpretable and empathetic reasoning paths that align with real community feedback.

2026-06-02 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

ディープリサーチエージェントはどこで間違っているのでしょうか?エージェントの軌跡におけるスパンレベルのエラーの位置特定

ディープリサーチエージェントは、検索、ツールの使用、証拠の検査、回答の合成という長い軌跡を通じてタスクを解決します。最終的な回答に基づく評価では、エージェントが成功したかどうかがわかりますが、軌跡のどの部分が回答の信頼性を低下させるのかはわかりません。私たちは、深層研究エージェントのスパンレベルのエラー位置特定を研究します。 2 つのエージェント フレームワーク、3 つのバックボーン モデル、3 つのベンチマークから 2,790 の実際の軌跡を収集し、生のログをセマンティック スパンに変換し、LLM 支援の専門家レビューを通じて有害なエラー スパンに注釈を付けます。これらのアノテーションから、通常の探索、失敗した探索、暫定的な仮説、無害なノイズの間のエラー範囲を特定するための 1,000 インスタンスのベンチマークである TELBench を構築します。さらに、エージェントのクレームを追跡し、軌跡の証拠でサポートをチェックし、サポートされていないクレームや矛盾するクレームが回答経路に影響を与えるスパンをマークする、クレーム中心の監査フレームワークである DRIFT を提案します。モデル ファミリと監査フレームワークにわたる実験では、DRIFT によりスパンレベルのエラー位置特定と最初のエラーの精度が最大 30 パーセント ポイント改善されることが示されています。私たちの研究は、ディープリサーチエージェントの信頼性をプロセスレベルで把握することを可能にします。

原文 (English)

Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final answers shows whether an agent succeeds, but not which parts of the trajectory make the answer unreliable. We study span-level error localization for deep-research agents. We collect 2,790 real trajectories from two agent frameworks, three backbone models, and three benchmarks, convert raw logs into semantic spans, and annotate harmful error spans through LLM-assisted expert review. From these annotations, we build TELBench, a 1,000-instance benchmark for identifying error spans among normal exploration, failed searches, tentative hypotheses, and harmless noise. We further propose DRIFT, a claim-centric auditing framework that tracks agent claims, checks their support in trajectory evidence, and marks spans where unsupported or conflicting claims affect the answer path. Experiments across model families and auditing frameworks show that DRIFT improves span-level error localization and first-error accuracy by up to 30 percentage points. Our work provides a process-level view of reliability in deep-research agents.

2026-06-02 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

BADGER: 生成的企業推論のためのエージェント的評価と決定論的評価の橋渡し

自然言語を SQL クエリに変換し、複数ステップのエージェント推論パイプラインを調整するエンタープライズ AI システムには、学術的なベンチマークとは根本的に異なる評価アプローチが必要です。 Spider と BIRD は実行精度プロトコルを確立しました。 G-Eval と RAGAS の高度な LLM ベースの評価。そして、Spider 2.0、BEAVER、BIRD-Interact などの最近の取り組みは、エンタープライズおよびエージェントの側面に取り組み始めています。テキストから SQL への評価とエージェントの動作評価を人間の専門家の判断に基づいて調整された運用グレードのパイプラインに統合するフレームワークは 1 つもありません。我々は、Merkle で開発された、テキストから SQL への評価とエージェントの動作評価を統合する統合評価フレームワークである BADGER を紹介します。 BADGER は 3 つの貢献を提供します。まず、LLM 支援の SQL コンポーネント抽出により、Spider 手法を拡張して、CTE を多用する方言固有の SQL を処理します。 2 つ目は、決定論的なセルレベルのスコアリングの前に、LLM を使用して構造アライメントを推測することにより、列エイリアシングと数値許容誤差の脆弱性を解決するハイブリッド実行精度メトリクス (Hybrid-EX) です。人間による注釈が付けられた 150 件の業界クエリで検証された Hybrid-EX は、Cohen の kappa=0.717 [95% CI: 0.600-0.822] (ほぼ一致) と 87.3% のバランスのとれた精度を達成し、6 つの競合するフレームワークすべてを上回りました (Delta-kappa: 0.322-0.502、すべて p<=0.001)。 3 番目は、RAGAS、G-Eval、エージェント ベンチマーク メトリクスを統合パイプラインに統合するエンタープライズ エージェント評価スイートです。過剰なツールの使用は唯一の新しい要素です。 BADGER は完全にクライアントの管理されたデータ環境内で実行され、構成可能な LLM ジャッジ バックエンドをサポートし、クライアント固有のジャッジとメトリクスの迅速なプロトタイピングを可能にし、1 回限りの品質ゲートではなく継続的な評価バックボーンとして機能します。

原文 (English)

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks. Spider and BIRD established execution-accuracy protocols; G-Eval and RAGAS advanced LLM-based assessment; and recent work such as Spider 2.0, BEAVER, and BIRD-Interact has begun to address enterprise and agentic dimensions. No single framework unifies text-to-SQL assessment with agentic behavior evaluation into a production-grade pipeline calibrated against human expert judgment. We present BADGER, developed at Merkle, a unified evaluation framework integrating text-to-SQL assessment with agentic behavior evaluation. BADGER offers three contributions. First, LLM-assisted SQL component extraction extending Spider methodology to handle CTE-heavy, dialect-specific SQL. Second, a hybrid execution accuracy metric (Hybrid-EX) resolving column-aliasing and numeric-tolerance brittleness by using an LLM to infer structural alignments before deterministic cell-level scoring. Validated on 150 human-annotated industry queries, Hybrid-EX achieves Cohen's kappa=0.717 [95% CI: 0.600-0.822] (Substantial agreement) and 87.3% balanced accuracy, outperforming all six competing frameworks (Delta-kappa: 0.322-0.502, all p<=0.001). Third, an enterprise agentic evaluation suite assembling RAGAS, G-Eval, and agent benchmark metrics into a unified pipeline; Excess Tool Usage is the sole novel element. BADGER runs entirely within the client's governed data environment, supports configurable LLM judge backends, and enables rapid prototyping of client-specific judges and metrics, serving as a continuous evaluation backbone rather than a one-time quality gate.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

食品の騒音と偽りの安全性:臨床医のフィードバックを用いた、LLM が摂食障害の質問にどのように適応できないかの系統的評価

最近の証拠は、摂食障害 (ED) を持つ人々が、大規模言語モデル (LLM) ベースのチャット システムに指導、アドバイス、精神的サポートを求めることが増えていることを示しています。これらのシステムは臨床上のアドバイスを提供するように設計されていませんが、その専門知識、中立性、アクセスしやすさにより、リスクはあるものの頻繁にサポートを受けることができます。このペーパーでは、安全でないまたは自傷行為を伴うユーザーのリクエストに無批判に適応し、促進するモデルから生じる潜在的な危害に焦点を当て、ユーザーと ED および LLM との間の潜在的なインタラクション パターンを調査します。私たちは、臨床 ED の専門家との協議の結果、プロンプト内の特定の言語的手がかりが安全でない応答の可能性を高め、ユーザー プロンプトに存在する潜在的なリスクの程度を体系的に変化させることによって、LLM が問題のある潜在的に危険なユーザー入力にどの程度無批判に適応するかを報告していることを発見しました。

原文 (English)

Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback

Recent evidence shows that people with eating disorders (EDs) are increasingly seeking guidance, advice, and emotional support from Large Language Model (LLM)-based chat systems. Although these systems are not designed to provide clinical advice, their perceived expertise, neutrality and accessibility make them a frequent, albeit risky, source of support. This paper investigates potential patterns of interaction between users with EDs and LLMs, focusing on the potential harms arising from models that uncritically adapt to, and facilitate unsafe or self-harming user requests. We find, in consultation with clinical ED experts, that specific linguistic cues in prompts increase the likelihood of unsafe responses and, through systematically varying the degree of potential risk present in the user prompt, report the extent to which LLMs uncritically adapt to problematic, and potentially dangerous user inputs.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AGENTCL: 言語エージェントの継続学習の厳格な評価に向けて

言語エージェントは、個々のタスクの解決にかなりの推論時間を費やしますが、あるエピソードで得た経験は、その後のエピソードでは十分に活用されないことがよくあります。継続的な学習では、エージェントが一連のタスク全体にわたって再利用可能なエクスペリエンスを蓄積し、時間の経過とともに改善し、無関係なエクスペリエンスによる干渉を回避することが期待されます。残念ながら、既存のベンチマークでは、言語エージェントの継続的な学習を厳密に評価するのが困難です。ほとんどの取り組みは、長いコンテキストの会話やドキュメントの検索と推論に焦点を当てていますが、最近の生涯適応ベンチマークは、タスク間の関係の分析が限られた単純なタスク ストリームに依存していることが多く、エージェントが時間の経過とともに何を学習し、再利用するかを理解することが困難になっています。この論文では、制御されたタスク ストリームと転送ゲインのメトリクスを中心とした、エージェントの継続的学習のための評価フレームワーク AgentCL を紹介します。 AGENTCL は、以前のサブソリューション、証拠、またはワークフローが後のタスクで意図的に再利用可能な構成ストリームを構築し、そのような再利用性が保証されていない単純なストリームと対比します。ベンチマークを使用して、継続学習のためのノンパラメトリック メモリ設計を評価します。メモリ設計の選択が継続学習にどのような影響を与えるかを診断するために、統合中に信頼性の低いエクスペリエンスをフィルタリングしながら、インタラクション、洞察、スキルを保存するプローブ手法である MemProbe を開発しました。コーディング、詳細な研究、および言語理解/推論タスクにわたる実証分析では、単純なストリームではメモリ設計を区別する能力が限られているのに対し、制御されたストリームではその可塑性をより明確に区別できることが示されています。一方、単純な設定や保持された設定では、得られるゲインが限られていることが多く、メモリによる劣化が生じる可能性があります。これらの結果は、可塑性と安定した再利用のバランスをとった、より強力なメモリ設計の必要性を浮き彫りにしています。

原文 (English)

AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes. Continual learning expects an agent to accumulate reusable experience across a stream of tasks, improve over time, and avoid interference from irrelevant experiences. Unfortunately, existing benchmarks struggle to evaluate continual learning in language agents rigorously. Most efforts focus on retrieval and reasoning over long-context conversations or documents, while recent lifelong-adaptation benchmarks often rely on naive task streams with limited analysis of cross-task relationships, making it difficult to understand what an agent learns and reuses over time. This paper presents an evaluation framework AgentCL for continual learning in agents, centered on controlled task streams and metrics for transfer gains. AGENTCL constructs compositional streams where earlier sub-solutions, evidence, or workflows are intentionally reusable in later tasks, and contrasts them with naive streams where such reusability is not guaranteed. We use the benchmark to evaluate non-parametric memory designs for continual learning. To diagnose how memory design choices affect continual learning, we develop MemProbe, a probing method that stores interactions, insights, and skills, while filtering unreliable experiences during consolidation. Empirical analysis across coding, deep research, and language understanding/reasoning tasks shows that naive streams offer limited ability to distinguish memory designs, whereas controlled streams more clearly distinguish their plasticity. Meanwhile, naive and held-out settings often yield limited gains and can expose memory-induced degradation. These results highlight the need for stronger memory designs that balance plasticity and stable reuse.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

BenHalluEval: ベンガル語の大規模言語モデル用のマルチタスク幻覚評価フレームワーク

ベンガル語は世界で 6 番目に話されている言語であるにもかかわらず、ベンガル語の大規模言語モデル (LLM) で幻覚を体系的に評価した先行研究はありません。 BenHalluEval は、生成的質問応答 (GQA)、バングラ語と英語のコード混合 QA、要約、および推論の 4 つのタスクをカバーするベンガル語用のきめ細かい幻覚評価フレームワークです。既存の 3 つのベンガル語データセットから抽出された、12 のタスク固有の幻覚タイプにわたって GPT-5.4 を使用して 12,000 の幻覚候補を構築し、グラウンドトゥルース インスタンスの偽陽性率 (トラック A) と幻覚候補の幻覚検出率 (トラック B) を独立して測定するデュアル トラック プロトコルの下で、推論指向、多言語、ベンガル語中心のカテゴリにわたる 7 つの LLM を評価します。両方の故障モードに共同でペナルティを課し、均一な応答バイアスによるスコアのインフレを防ぐために、モデルとタスク全体で 7.72% から 55.42% の範囲のデュアルトラック キャリブレーション メトリクスである BenHalluScore を提案します。これは、幻覚キャリブレーションの大幅な変動を明らかにします。緩和戦略として適用される思考連鎖プロンプトは、幻覚差別を一貫して改善することなく、反応分布を変化させます。 BenHalluEval は、ベンガル語専用の幻覚ベンチマークを初めて確立し、リソースの少ない言語設定に対する単一トラックおよびプロンプトのみの評価アプローチが不適切であることを強調しています。データセットとコードは https://anonymous.4open.science/r/BanglaHalluEval-EB77 で入手できます。

原文 (English)

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates using GPT-5.4 across twelve task-specific hallucination types, drawn from three existing Bengali datasets, and evaluate seven LLMs spanning reasoning-oriented, multilingual, and Bengali-centric categories under a dual-track protocol that independently measures false-positive rate on ground-truth instances (Track A) and hallucination detection rate on hallucinated candidates (Track B). To jointly penalise both failure modes and prevent inflated scores from uniform response bias, we propose BenHalluScore, a dual-track calibration metric that ranges from 7.72% to 55.42% across models and tasks, revealing substantial variation in hallucination calibration. Chain-of-thought prompting, applied as a mitigation strategy, shifts response distributions without consistently improving hallucination discrimination. BenHalluEval establishes the first dedicated hallucination benchmark for Bengali and highlights the inadequacy of single-track and prompting-only evaluation approaches for low-resource language settings. The dataset and code are available at https://anonymous.4open.science/r/BanglaHalluEval-EB77.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

SENSE: 検索ベースの投機的デコーディングのためのソフトゲート評価によるセマンティック埋め込みナビゲーション

投機的デコード (SD) は、軽量のドラフト モデルを採用して候補トークンを提案することにより、大規模言語モデル (LLM) の推論を高速化します。候補トークンは、生成の品質を損なうことなく、ターゲット モデルによって並行して検証されます。検索ベースの投機的復号 (RSD) はプラグアンドプレイの多用途性で好まれていますが、その可能性は厳格な語彙依存性によって妨げられ、検索と検証の両方が表面レベルの変動に対して脆弱になります。 To address this, we propose SENSE (Semantic Embedding Navigation with Soft-gated Evaluation).ターゲット モデルの隠れた状態に検索を固定することにより、SENSE は堅牢な意味論的調整を確立し、ソフトゲート評価モジュールが表面形式ではなく意味論的等価性を検証できるようにします。厳密なベンチマークを保証するために、統一されたフレームワーク内で既存のメソッドをアトミ​​ックなプリミティブに分解し、コンポーネントレベルの詳細な比較を容易にします。多様なドメインにわたる広範な実験により、SENSE が LLaMA および Qwen ファミリの複数のベースラインを上回り、生成品質を維持しながら最大 4.09 の平均許容長と 3.26 倍の高速化を達成することが実証されました。私たちのコードは公開され次第公開されます。

原文 (English)

SENSE: Semantic Embedding Navigation with Soft-gated Evaluation for Retrieval-based Speculative Decoding

Speculative Decoding (SD) accelerates Large Language Model (LLM) inference by employing a lightweight draft model to propose candidate tokens, which are verified in parallel by the target model, without compromising generation quality. While Retrieval-based Speculative Decoding (RSD) is favored for its plug-and-play versatility, its potential is impeded by rigid lexical dependencies, rendering both retrieval and verification brittle to surface-level variations. To address this, we propose SENSE (Semantic Embedding Navigation with Soft-gated Evaluation). By anchoring retrieval on the hidden states of the target model, SENSE establishes robust semantic alignment, which empowers the Soft-gated Evaluation module to validate semantic equivalence rather than surface forms. To ensure rigorous benchmarking, we deconstruct existing methods into atomic primitives within a unified framework, facilitating granular, component-level comparison. Extensive experiments across diverse domains demonstrate that SENSE outperforms multiple baselines on the LLaMA and Qwen families, attaining up to 4.09 mean acceptance length and 3.26x speedup, while preserving generation quality. Our code will be released upon publication.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

医療大規模言語モデルの安全性、堅牢性、公平性を評価するためのマルチドメイン レッド チーミング フレームワーク

大規模言語モデル (LLM) は医療全体に導入されることが増えていますが、既存のベンチマークは、臨床現場で一般的な敵対的または倫理的に複雑な条件下でのモデルの動作を捕捉できません。私たちは、9 つ​​のドメインと 150 以上のサブカテゴリーにわたる 690 の臨床に基づいたシナリオにわたって、11 の最新の LLM を評価するマルチドメインのレッド チーミング フレームワークを開発しました。シナリオには敵対的変換が組み込まれており、LLM 支援スコアリングと人間参加型検証を備えた 7 次元ルーブリックを使用して反応が評価されました。結果は、平均スコアが 0.791 ~ 0.984 の範囲で、パフォーマンスに大きなばらつきがあることが明らかになりました。重要なことに、いくつかの高性能システムは、安全性が重要な個々のシナリオで完全な障害を引き起こし、総合的な精度が臨床的に意味のあるリスクを覆い隠していることを実証しました。最もパフォーマンスの高いシステム (X-BAI、GPT-5、Claude Opus 4.1) は、低い分散で 0.97 を超えるスコアを達成しましたが、パフォーマンスはドメイン間で大きく異なりました。株式関連のタスクでは、人口動態の変更により誤差が 10 ~ 20% 増幅されることが示され、人間のレビュー担当者は、自動評価では見逃されていた臨床関連の失敗を特定しました。私たちの調査結果は、性能のばらつきと最悪の場合の故障が、平均精度だけよりも臨床的に意味のある信頼性指標を提供すること、および信頼性の高い安全性評価には自動化と臨床医の監視を組み合わせたハイブリッド評価アプローチが不可欠であることを示しています。

原文 (English)

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or ethically complex conditions common in clinical practice. We developed a multi-domain red teaming framework evaluating eleven contemporary LLMs across 690 clinically grounded scenarios spanning nine domains and over 150 subcategories. Scenarios incorporated adversarial transformations, and responses were assessed using a seven-dimension rubric with LLM-assisted scoring and human-in-the-loop validation. Results revealed substantial performance variance, with mean scores ranging from 0.791 to 0.984. Critically, several high-performing systems produced complete failures in individual safety-critical scenarios, demonstrating that aggregate accuracy masks clinically meaningful risk. The highest-performing systems (X-BAI, GPT-5, Claude Opus 4.1) achieved scores above 0.97 with low variance, while performance varied significantly across domains. Equity-related tasks showed 10-20% error amplification with demographic modifications, and human reviewers identified clinically relevant failures missed by automated evaluation. Our findings demonstrate that performance variance and worst-case failures provide more clinically meaningful reliability indicators than mean accuracy alone, and that hybrid evaluation approaches combining automation with clinician oversight are essential for credible safety assessment.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ビジネス/資金調達研究/論文

CardioLens: マルチシーケンス心臓 MRI 評価を通じて MLLM の臨床現実ギャップを明らかにする

マルチモーダル大規模言語モデル (MLLM) は、公的医療ベンチマークで優れたパフォーマンスを示していますが、既存の評価は分離された入力と簡素化された認識スタイルのタスクに依存しており、臨床用途には弱い代物であることがよくあります。私たちは、厳格なレポートから QA への構築および検証パイプラインを通じて民間病院のアーカイブから構築された、マルチシーケンス心臓血管磁気共鳴 (CMR) の耐漏洩性評価テストベッドである CardioLens を紹介します。 CardioLens には、4D シネ、LGE、灌流、および T2 強調イメージングにわたる 473,896 のスライスと 13,494 の検証済み QA ペアが含まれており、画像理解、レポート作成、疾患診断という 3 つの段階の CMR 読影を評価します。 CardioLens は、24 の最先端の MLLM にわたって、臨床現実との大きなギャップを明らかにしました。モデルは全体的にパフォーマンスが低く、実際の CMR ワークフローに沿ってパフォーマンスが低下しています。さらに、混乱分析では、臨床的に異なる所見を区別するのではなく、モデルがデフォルトで頻繁に異常なカテゴリを設定するカテゴリ崩壊失敗モードが示されています。 MLLM 互換の入力構造が主な原因であることを除外するために、ランダムで臨床的動機に基づいたデータ駆動型のスライス選択プロトコルをさまざまなスライス バジェットで比較します。パフォーマンスはわずかにのみ変化します (通常は約 1%)。明示的な推論プロンプトもパフォーマンスを回復することができず、多くの場合、視覚的な証拠の使用を改善するのではなく、モデルをより保守的にします。これらの結果は、現在の MLLM が信頼性の高い CMR 解釈からはほど遠いことを示しており、臨床上の決定にはシーケンス、ビュー、時間相にわたって分散した証拠を統合する必要があります。 CardioLens は、実際の臨床展開に向けて次世代 MLLM を開発するための臨床に基づいたテストベッドを提供します。

原文 (English)

CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on isolated inputs and simplified recognition-style tasks. We introduce CardioLens, a leakage-resistant evaluation testbed for multi-sequence Cardiovascular Magnetic Resonance (CMR), constructed from private hospital archives through a rigorous report-to-QA construction and verification pipeline. CardioLens contains 473,896 slices and 13,494 verified QA pairs across 4D Cine, LGE, perfusion, and T2-weighted imaging, and evaluates three stages of CMR interpretation: image understanding, report generation, and disease diagnosis. Across 24 state-of-the-art MLLMs, CardioLens reveals a substantial clinical reality gap: models perform poorly overall, with performance degrading along the real CMR workflow. Confusion analysis further shows a category-collapse failure mode, where models default to frequent abnormal categories rather than distinguishing clinically distinct findings. To rule out MLLM-compatible input construction as the primary cause, we compare random, clinically motivated, and data-driven slice selection protocols under different slice budgets; performance changes only marginally, typically by about 1%. Explicit reasoning prompts also fail to rescue performance, often making models more conservative rather than improving visual evidence use. These results show that current MLLMs remain far from reliable CMR interpretation, where clinical decisions require integrating distributed evidence across sequences, views, and temporal phases. CardioLens provides a clinically grounded testbed for developing next-generation MLLMs toward real-world clinical deployment.

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

SMOTE ベースのオーバーサンプリングとサイドチャネル電力データの拡張マルチモデル評価により IoT 侵入検知を向上

IoT ベースのネットワークにおける侵入の検出には、従来の機械学習手法では克服できない課題が伴います。おそらくその最大のものは、サイドチャネル データセット内のクラスの不均衡の存在に関連しており、攻撃と比較した通常クラスのサンプル数が 75,964 対 1 の比率に達する可能性があります。このような側面については、Dominguez et al. が取り上げています。電力ベースの侵入検知の概念実証を通じて。残念ながら、著者らは不均衡の問題に対処しようとしているわけでも、バランスのとれたトレーニング セットを使用して分類器のパフォーマンスを評価しているわけでもありません。今回の文書では、両方の側面を一度に扱います。まず、最初のデータセットから抽出された 9 つの可能なデータセットすべてに対して合成マイノリティ オーバーサンプリング技術 (SMOTE) が実行され、それぞれの正確な不均衡率 1.1 が得られました。次に、8 つのアルゴリズム (ランダム フォレスト、HistGradientBoosting、LightGBM、エクストラ ツリー、XGBoost、k-最近傍法、多層パーセプトロン、デシジョン ツリー) が、SMOTE バランスの取れた 6 時間データセットに対して同一条件下でトレーニングされました。ランダム フォレストは、ミクロ平均 F1 スコア 0.9989、マクロ F1 スコア 0.9794 に達し、ベースペーパーから時系列フォレスト アルゴリズムによって得られた以前の最高のミクロ F1 結果 0.9983 を上回りました。 Extra Trees も同様のパフォーマンスを提供しましたが、速度は 10 倍でした。基本的な論文評価とは明確に対照的なマクロ F1 メトリクスの導入により、集計パフォーマンス メトリクスでは見逃されていた重要なクラスレベルの情報が明らかになります。混同行列、F1 ヒートマップ、ROC 曲線を使用して計算されたクラスごとの再現率は、少数派の攻撃クラス、特に M+L 感染を組み合わせたクラスが、SMOTE バランスを使用した場合にのみ確実に検出されることを示しています。特徴重要度分析では、パワー ウィンドウの 60 ステップのうち、最も重要な予測信号として最新のタイム ステップが示されます。

原文 (English)

Improving IoT Intrusion Detection Through SMOTE-Based Oversampling and Extended Multi-Model Evaluation on Side-Channel Power Data

The detection of intrusions in IoT-based networks poses challenges that cannot be overcome using traditional machine learning methods. Perhaps the biggest of them is related to the presence of a class imbalance in the side-channel dataset, where the number of samples in the normal class compared to the attacks can reach a ratio of 75,964 to 1. Such an aspect is addressed by Dominguez et al. through the proof of concept of power-based intrusion detection. Unfortunately, neither the authors attempt to cope with the problem of imbalance nor do they assess the classifier performance using a balanced training set. In the current paper, both aspects will be handled at once. First, a Synthetic Minority Oversampling Technique (SMOTE) was performed on all nine possible datasets extracted from the initial one, providing an exact imbalance ratio of 1.1 for each. Then, eight algorithms i.e. Random Forest, HistGradientBoosting, LightGBM, Extra Trees, XGBoost, k-Nearest Neighbors, Multi-Layer Perceptron, and Decision Tree were trained under identical conditions for the SMOTE balanced 6-hour dataset. Random Forest reached a micro-averaged F1 score of 0.9989 and macro F1 of 0.9794, thus outperforming the previously best micro-F1 result obtained by Time Series Forest algorithm from the base paper of 0.9983. Extra Trees provided the same performance as well, but at 10 times faster. The introduction of a macro-F1 metric explicitly in contrast to the base paper assessment reveals important class-level information missed with aggregate performance metrics. Recall rates per-class calculated with confusion matrices, F1 heatmaps, and ROC curves show that minority attack classes, especially those with combined M+L infections, are detected reliably only when using SMOTE balance. Feature importance analysis indicates the latest time steps as the most important predictor signals out of 60 steps in a power window.

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成ロボティクスビジネス/資金調達

StressDream: 堅牢なポリシーの評価と改善のためのビデオ ワールド モデルのステアリング

ビデオ ワールド モデル (WM) は、エゴロボットの行動を条件とした現実的な将来の観察を想像することにより、政策の評価と改善が期待できることを示しています。 WM は先物全体の分布をモデル化できますが、政策の評価と改善は通常、名目上の想像力に依存するため、法外に多くのサンプルが抽出されない限り、ロボットのアクションによる影響の大きい結果を見逃してしまう可能性があります。 WM 想像を超えた堅牢なポリシー評価と改善を可能にするために、拡散ベースの WM の初期ノイズを最適化することで、推論時に指定された影響力が大きくてももっともらしい結果に想像を導く StressDream を提案します。ただし、高次元ノイズの最適化は困難です。最適化では、ありえない想像を生み出す分布外 (OOD) ノイズを回避しながら、生成されたビデオ内の微妙なシーン依存のターゲット イベントを考慮する必要があります。私たちは、2 つの相補的な目標でこれに対処します。生成されたビデオについて推論することで有益な勾配を提供する視覚言語モデルを使用した意味論的な目標と、最適化されたノイズによる OOD のドリフトを防ぐ妥当性目標です。自動運転とロボット操作のための最先端のビデオワールドモデルを使用して、StressDreamが、タスクの失敗など、推論時にテキストによって指定される影響力が大きいがもっともらしい結果に向けて想像力を効果的に導き、望ましくない結果を含むもっともらしい未来を持つアクションを特定することで、堅牢なポリシーの評価と改善を可能にすることを示します。ビデオ結果は https://junwon.me/StressDream/ でご覧いただけます。

原文 (English)

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement

Video world models (WMs) have shown promise for policy evaluation and improvement by imagining realistic future observations conditioned on ego-robot actions. While WMs can model distributions over futures, policy evaluation and improvement typically rely on nominal imaginations, which can miss high-impact outcomes of robot actions unless prohibitively many samples are drawn. To enable robust policy evaluation and improvement over WM imaginations, we propose StressDream, which steers imaginations toward high-impact yet plausible outcomes specified at inference time by optimizing the initial noise of diffusion-based WMs. However, optimizing high-dimensional noise is challenging: the optimization must reason about nuanced, scene-dependent target events in generated videos while avoiding out-of-distribution (OOD) noise that yields implausible imaginations. We address this with two complementary objectives: a semantic objective with a Vision-Language Model that provides informative gradients by reasoning about the generated video, and a plausibility objective that prevents the optimized noise from drifting OOD. With state-of-the-art video world models for autonomous driving and robotic manipulation, we show that StressDream effectively steers imaginations toward high-impact yet plausible outcomes specified by text at inference time, such as task failures, enabling robust policy evaluation and improvement by identifying actions whose plausible futures include undesirable outcomes. Video results are available at https://junwon.me/StressDream/.

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体ビジネス/資金調達

SUPREME: 再現可能な画像非学習手法評価のためのマルチ GPU フレームワーク

機械の非学習では、最初から再トレーニングすることなく、トレーニングされたモデルから特定のトレーニング データの影響が除去されます。アンラーニング手法を評価するには、複数のシードにわたってトレーニング、アンラーニング、評価を繰り返す必要があり、計算コストがかかります。私たちの知る限り、既存の画像分類非学習フレームワークは単一の GPU 上で実行されるため、妥当な時間内に評価できるシードの数が制限されます。これらのステージを複数の GPU に分散するオープンソース フレームワークである SUPREME を紹介します。 SUPREME は 3 つの貢献を行っています。新しいメソッド、メトリック、モデル、シナリオを追加するためのレジストリ ベースの設計です。複数のアクセラレータと高精度モードをサポートするマルチ GPU アーキテクチャ。そして、10 個のシードにわたるフルクラスおよびランダム サンプルのアンラーニングのもとで、ResNet18 と ViT を使用したピンの顔認識のデモンストレーションです。このフレームワークは https://github.com/pedroandreou/supreme-unlearning で入手できます。

原文 (English)

SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation

Machine unlearning removes the influence of specific training data from a trained model without retraining it from scratch. Evaluating an unlearning method requires repeating training, unlearning, and evaluation across multiple seeds, which is computationally expensive. To our knowledge, existing image classification unlearning frameworks run on a single GPU, which limits how many seeds can be evaluated in reasonable time. We introduce SUPREME, an open-source framework that distributes these stages across multiple GPUs. SUPREME makes three contributions: a registry-based design for adding new methods, metrics, models, and scenarios; a multi-GPU architecture supporting multiple accelerators and precision modes; and a demonstration on Pins Face Recognition using ResNet18 and ViT under full-class and random-sample unlearning across ten seeds. The framework is available at https://github.com/pedroandreou/supreme-unlearning.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to th…

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation

Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume…

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

Strong Stochastic Flow Maps

Flow and diffusion models generate high-quality samples in many modalities; however, many network evaluations are required during inference…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages

Despite being home to more than 1300 ethnic groups and 700 indigenous languages, bias in Large Language Models has not been fully studied i…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly Afri…

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

On the Evaluation of Spiking Neural Network Configurations for Network Intrusion Detection

Network intrusion detection is a core component of modern cybersecurity infrastructure, yet the deep learning models that dominate the fiel…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Hierarchical Online Prompt Mutation with Dual-Loop Feedback for Guardrailed Evidence Document Generation: A Production-Evaluation Case Study

High-stakes production document-generation systems require language models to be adaptive, evidence-grounded, and auditable. We present HOP…

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

A Framework for Graph-Conditioned Hierarchical Shapley Attribution in Patent Valuation

Estimating the economic contribution of a single patent inside a product that embodies tens of thousands of patents is a long-standing unso…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AlphaToken: Decoupling Adaptation and Stability for Path-Aware Response Token Valuation in LLM Post-Training

Token selection is pivotal for effective LLM post-training. However, existing methods mostly rely on local heuristics and rarely formulate…

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection

Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expens…

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Rethinking Evaluation Paradigms in IBP-based Certified Training

Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations. Neura…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave un…

2026-06-02 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Agent Guide: A Simple Agent Behavioral Watermarking Framework

The increasing deployment of intelligent agents in digital ecosystems, such as social media platforms, has raised significant concerns abou…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

A Unified Evaluation-Instructed Framework for Query-Dependent Prompt Optimization

Most prompt-optimization methods refine a single static template, making them ineffective in complex and dynamic user scenarios. Existing q…

2026-06-02 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Causal state binding predicts action control in language agents

Autonomous language agents increasingly expose traces, memories, plans and constraints, but existing evaluations rarely test whether these…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Capturing LLM Capabilities via Evidence-Calibrated Query Clustering

Query clustering organizes queries into groups that reflect shared latent capability demands, enabling capability-aware LLM evaluation. Exi…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

FundaPod: AI 支援のファンダメンタル投資調査のためのナレッジ グラフ メモリを備えたマルチペルソナ エージェント ポッド プラットフォーム

大規模言語モデル (LLM) は金融分野での適用が増えていますが、既存の研究のほとんどは取引シグナルや予測を中心とした財務 NLP タスクに重点を置いています。対照的に、制度的基礎研究では、人間のアナリストまたは AI エージェントが証拠を収集し、ビジネス推進要因を特定し、競合する視点を比較し、投資メモを作成する必要があります。その広範な目標は、単に結果を予測することではなく、投資知識の累積的な発展に貢献しながら、透明性、再利用可能、検証可能な投資計画を作成することです。 AI 支援のファンダメンタルズ投資調査のためのマルチペルソナ エージェント プラットフォームである FundaPod を紹介します。私たちは、基礎研究は人間中心の意思決定支援タスクであり、取引シグナルの生成とは質的に異なるため、独立性を維持するアーキテクチャの方が適していると主張します。 FundaPod では、バリュー投資家やマクロ戦略家など、さまざまなペルソナを持つ AI エージェントが、共有の出所契約に基づいて独立して調査を実施します。その後、彼らの意見の相違は、知識グラフ記憶システムを通じて人間のポートフォリオ マネージャー (PM) による裁定のために事後的に表面化されます。この論文は、設計科学の実践と認知的分離と人間と機械の協調の理論に基づいた、基礎研究をサポートする人間と AI のハイブリッド システムの 5 つの設計原則を提供します。また、4 つのアーキテクチャ メカニズムについても説明します。1 つは一般投資家の資料を展開可能なエージェントに変えるペルソナ蒸留パイプラインです。プランナーが型指定されたタスク グラフを導出できるようにする宣言型スキル レジストリ。メモの主張を検証可能な情報源に結び付ける根拠のある証拠モデル。そしてティッカー、メモ、アナリスト、テーマを結び付けるナレッジグラフ「第二の脳」。完全なケーススタディとペルソナベースのメモの比較を通じてアーキテクチャを実証します。

原文 (English)

FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research

Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on prediction. Institutional fundamental research, by contrast, requires human analysts or AI agents to gather evidence, identify business drivers, compare competing viewpoints, and generate investment memos. Its broader goal is not merely to predict outcomes, but to produce investment plans that are transparent, reusable, and verifiable, while contributing to the cumulative development of investment knowledge. We present FundaPod, a multi-persona agent platform for AI-assisted fundamental investment research. We argue that fundamental research is a human-centric decision-support task that is qualitatively distinct from trading-signal generation, and is therefore better served by an independence-preserving architecture. In FundaPod, AI agents with different personas, such as value investors or macro strategists, conduct research independently under a shared provenance contract. Their disagreements are then surfaced post hoc for adjudication by the human portfolio manager (PM) through a knowledge-graph memory system. This paper contributes five design principles for human-AI hybrid systems supporting fundamental research, grounded in design-science practice and theories of cognitive isolation and human-machine coordination. It also describes four architectural mechanisms: a persona distillation pipeline that turns public investor materials into deployable agents; a declarative skill registry that lets the planner derive typed task graphs; a grounded evidence model that links memo claims to verifiable sources; and a knowledge-graph "second brain" that connects tickers, memos, analysts, and themes. We demonstrate the architecture through a complete case study and a persona-based memo comparison.

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

低リソースのコンテキストにおける AI のベンチマーク: リーダーボードを超えて考える

既存の AI 評価手法では、運用上の制約がモデルの品質と同じくらい使いやすさを左右する低リソース環境でシステムが実際にどのように動作するかを把握できないことがよくあります。音声、チャット/RAG、およびビジョン システムにわたる既存のベンチマーク ファミリの構造化分析を通じて、ラボでの評価実践と低リソース環境での実際の展開条件との間の重大なギャップを特定します。私たちは、意味のある評価単位は分離されたモデルではなく展開されたシステムであり、効果的な評価フレームワークはタスクのパフォーマンスとノイズの多い入力、コードスイッチング、断続的な接続、ローエンドハードウェア、ドメインシフトなどの展開条件を統合する必要があると主張します。同時に、ベンチマークは、異なるアプリケーション クラスには、運用の違いを曖昧にする単一の集計スコアではなく、個別の評価プロファイルが必要であることを認識する必要があります。実際の意思決定をサポートするために、展開コンテキストに敏感でありながら、システムやアプリケーションの種類全体での比較可能性を維持する共有レポート フレームワークを提案します。最後に、標準化された 1 ページのベンチマーク カード、展開プロファイル、障害処理手順と人間による監視メカニズムの明示的な文書など、政策立案者、寄付者、実装者向けの簡潔で実用的なレポート作成物の必要性を強調します。

原文 (English)

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape usability as much as model quality. Through a structured analysis of existing benchmark families across speech, chat/RAG, and vision systems, we identify critical gaps between laboratory evaluation practices and real-world deployment conditions in low-resource environments. We argue that the meaningful unit of assessment is the deployed system rather than an isolated model and that effective evaluation frameworks must integrate task performance with deployment conditions such as noisy inputs, code-switching, intermittent connectivity, low-end hardware, and domain shift. At the same time, benchmarks should recognize that different application classes require distinct evaluation profiles rather than a single aggregate score that obscures operational differences. To support practical decision-making, we propose a shared reporting framework that preserves comparability across systems and application types while remaining sensitive to deployment context. Finally, we emphasize the need for concise and actionable reporting artifacts for policymakers, donors, and implementers, including standardized one-page benchmark cards, deployment profiles, and explicit documentation of failure handling procedures and human oversight mechanisms.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Cookie-Bench: Web 生成のための継続的なオンスクリーンキーインタラクション評価

フロントエンドの Web コードは、すべてのフロンティア LLM リリースの中核的な製品面となっていますが、アリーナのような人間が判断するリーダーボードは拡張できないため、これらのインタラクティブ アプリケーションを開発スピードで評価することは依然としてコストがかかります。既存の自動プロキシは通常、リファレンス実装、テスト スイート、または厳密なチェックリストに依存しており、人間のレビュー担当者がライブ セッションで実行する合理的な合成を見逃す傾向があります。私たちは、同時に参照フリーで、自律的に駆動され、総合的に推論される新しい評価体制を明確にし、2 つの成果物を通じてそれをインスタンス化します。 \textbf{\dataname} は、静的プレゼンテーション タスクと対話型アプリケーション タスクの両方にまたがる 11 ドメイン、54 リーフ、1,000 クエリの WebDev ベンチマークであり、3 つの難易度層と 3 つのターゲット言語グループにわたってバランスが取れており、回覧されたプロンプトから思い出せないようにブリーフが書き直されています。 \textbf{\framename} は、Flavell のメタ認知モニタリングに基づいており、証拠の蓄積と判断を 3 つの段階にわたって分離します。静的な知覚は受動的な観察から第一印象を形成します。エージェント駆動のインタラクションは、連続画面のビデオ、音声、およびステップごとのスクリーンショットをキャプチャしながら、アプリケーションを自律的に探索します。動的スコアリングは、証拠チェーンが完了した後にのみ、構造化された失敗の帰属を伴う全体的な機能性と美的判断を発行します。 \dataname では、\framename は専門家による評価と厳密に一致しており、インタラクティブな Web 生成に関して 13 のフロンティア LLM 全体でかなりのヘッドルームを表面化しています。 \noindenthttps://anonymous.4open.science/r/Cookie-3CE/

原文 (English)

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation

Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed remains costly because human-judged leaderboards like Arena do not scale. Existing automated proxies typically lean on reference implementations, test suites, or rigid checklists, and tend to miss the reasoned synthesis a human reviewer performs over a live session. We articulate a new evaluation regime that is simultaneously reference-free, autonomously driven, and holistically reasoned, and instantiate it through two artifacts. \textbf{\dataname} is an 11-domain, 54-leaf, 1,000-query WebDev benchmark spanning both static-presentation and interactive-application tasks, balanced across three difficulty tiers and three target-language groups, with briefs rewritten to resist recall from circulated prompts. \textbf{\framename}, grounded in Flavell's metacognitive monitoring, separates evidence accumulation from judgment across three stages: Static Perception forms a first impression from passive observation; Agent-Driven Interaction explores the application autonomously while capturing continuous screen video, audio, and per-step screenshots; Dynamic Scoring issues holistic functionality and aesthetics verdicts with structured failure attribution only after the evidence chain is complete. On \dataname, \framename aligns closely with expert human ratings while surfacing substantial headroom across 13 frontier LLMs on interactive web generation. \noindenthttps://anonymous.4open.science/r/Cookie-3CE/

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成エージェントビジネス/資金調達

Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation

Multi-modal 3D Intelligence has gained considerable attention due to its wide applications in autonomous driving and world simulation, etc.…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AutoEval Done Right: Using Synthetic Data for Model Evaluation

The evaluation of machine learning models using human-labeled validation data can be expensive and time-consuming. AI-labeled synthetic dat…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Erased but Not Forgotten: How Backdoors Compromise Concept Erasure

The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

グラフ検索からスキーマ実現まで: 異種ナレッジ グラフ上のテキストから SPARQL への反事実検証

Text-to-SPARQL は、自然言語の質問を RDF ナレッジ グラフ上の実行可能な SPARQL クエリにマッピングします。標準的な評価ではターゲット グラフが事前に修正されることがよくありますが、実践的なナレッジ グラフ質問応答 (KGQA) には、異なるスキーマ、部分的なアラインメント、および不完全なメタデータを含む異種グラフ コレクションが含まれる場合があります。この設定では、クエリ生成は SPARQL 構文以上のものに依存します。システムは、質問に必要な述語、エンティティ タイプ、結合、フィルター、および制約をサポートできるグラフ スキーマを識別する必要があります。異種の KG コレクション上でテキストから SPARQL に変換するためのスキーマベースのエージェント フレームワークである SchemaForge を紹介します。その中心的なメカニズムは、質問条件付きのスキーマ スライス アライメントです。弱いグラフの証拠によって最初にもっともらしいグラフが特定され、より強力なスキーマの証拠によって、ローカル スキーマ スライスが意図したクエリを実現できるかどうかが決まります。選択されたスキーマ スライスは、クエリの生成と実行前の検証を制限します。利用可能なグラフが 1 つだけの場合、同じ定式化は、スキーマ基盤を備えた標準の単一 KG テキストから SPARQL への変換に縮小されます。 LC-QuAD 2.0、QALD-9 Plus、QALD-10、および Spider4SPARQL で SchemaForge を評価します。 SchemaForge は、4 つの公開ベンチマーク全体で、最も一致するエージェントのベースラインよりも実行精度を平均 11.50 パーセント向上させています。 Spider4SPARQL では、SchemaForge は実行精度を 54.86% から 64.18% に向上させ、トップ 1 グラフ割り当て精度 73.0% とトップ 3 グラフ割り当て精度 97.0% を達成しました。これらの結果は、グラフの弱い証拠からスキーマ固有のクエリコミットメントへの移行と、反事実の回答セットのチェックにより、異種ナレッジグラフよりも実行可能なクエリの生成が向上することを示しています。

原文 (English)

From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs

Text-to-SPARQL maps natural-language questions to executable SPARQL queries over RDF knowledge graphs. While standard evaluations often fix the target graph in advance, practical knowledge graph question answering (KGQA) may involve heterogeneous graph collections with different schemas, partial alignments, and incomplete metadata. In this setting, query generation depends on more than SPARQL syntax: the system must identify a graph schema that can support the predicates, entity types, joins, filters, and constraints required by the question. We present SchemaForge, a schema-grounded agentic framework for text-to-SPARQL over heterogeneous KG collections. Its central mechanism is question-conditioned schema-slice alignment: weak graph evidence first identifies plausible graphs, while stronger schema evidence determines whether a local schema slice can realize the intended query. The selected schema slice then constrains query generation and verification before execution. When only one graph is available, the same formulation reduces to standard single-KG text-to-SPARQL with schema grounding. We evaluate SchemaForge on LC-QuAD 2.0, QALD-9 Plus, QALD-10, and Spider4SPARQL. Across the four public benchmarks, SchemaForge improves execution accuracy over the strongest matched agent baseline by 11.50 percentage points on average. On Spider4SPARQL, SchemaForge improves execution accuracy from 54.86% to 64.18% and achieves 73.0% Top-1 and 97.0% Top-3 graph allocation accuracy. These results show that moving from weak graph evidence to schema-specific query commitments, together with counterfactual answer-set checks, improves executable query generation over heterogeneous knowledge graphs.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs).Despite recent be…

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

Learning-To-Measure: In-Context Active Feature Acquisition

Active feature acquisition (AFA) is a sequential decision-making problem where the goal is to improve model performance for test instances…

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations

Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their r…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

InFerActive: Interactive Tree-Based Exploration of LLM Sampling for Safety Evaluation

Even LLMs that appear safe during evaluation can still produce harmful responses in deployment. Because stochastic sampling yields differen…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

Uncovering Competency Gaps in Large Language Models and Their Benchmarks

The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but…

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達研究/論文

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

Automatic metrics are widely used to evaluate text-to-image models, often replacing human judgment in benchmarking, model selection, and la…

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures

Machine Learning Interatomic Potentials (MLIPs) sometimes fail to reproduce the physical smoothness of the quantum potential energy surface…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework

Large language models (LLMs) have shown promise in constructing mechanistic models from data. However, existing evaluations largely focus o…

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Beyond String Matching: Semantic Evaluation of PDF Table Extraction

Reliably extracting tables from PDFs is essential for large-scale scientific data mining and knowledge base construction, yet existing eval…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

Failure of contextual invariance in large language models

Standard evaluation practices assume that large language model (LLM) outputs are stable when prompts are embedded in contextually equivalen…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook

As LLMs are globally deployed, aligning their cultural value orientations is critical for safety and user engagement. However, existing ben…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

AtomEval: Validity-Aware Atomic Evaluation of Adversarial Claim Rewriting in Fact Verification

Large language models (LLMs) can rewrite refuted claims to evade evidence-based fact verifiers, but conventional attack success rate (ASR)…

2026-06-02 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

Beyond Offline A/B Testing: Context-Aware Agent Simulation for Recommender System Evaluation

Recommender systems are central to online services, enabling users to navigate through massive amounts of content across various domains. H…

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Defeasible Conditional Obligation in a Two-tiered Preference-based Semantics (Extended Version)

In response to a concern raised by Horty, this paper develops a two-tiered, preference-based semantic framework for modeling defeasible con…

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make syste…

2026-06-02 13:00 JSTarXiv cs.AIビジネス/資金調達

RISED: A Pre-Deployment Evaluation Framework for High-Stakes AI Decision-Support Systems, with Application to Healthcare

Clinical decision-support systems are expert systems whose recommendations clinicians act on directly, yet they are usually cleared on one…

2026-06-02 07:55 JSTTechCrunch AIビジネス/資金調達

Alphabet plans to raise $80B to pay for AI buildout

"The company is experiencing strong demand for its AI solutions and services from enterprises and consumers, at levels that are exceeding t…

2026-06-02 03:19 JSTTechCrunch AIビジネス/資金調達

Water access is now a risk factor in SpaceX’s IPO

The company says it needs "significant" water resources to cool its data centers, and that access to abundant, affordable water is a challe…

2026-06-02 02:27 JSTITmedia AI+LLM/生成AIビジネス/資金調達

Anthropicが上場準備 直近の評価額は約154兆円

AnthropicがIPOに向け、SECに登録書類「S-1」のドラフトを非公開で提出した。直近のシリーズH資金調達での評価額は約9650億ドル(約154兆円)に達している。

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

生成 AI における多元的調整のためのペルソナベースの評価フレームワーク

生成型人工知能の現在の調整パラダイムは、主にモノリシックなベンチマーク フレームワークに依存しており、人間の複数の判断を集約された統計ベースラインに還元することで、評価における文化的、人口統計的、および文脈上のばらつきを曖昧にします。我々は、単一の評価関数を人間の多様な視点を表す合成認知プロファイルの構造化された多様体に置き換える、AI 評価のための状態空間制約付きエミュレーション フレームワークを導入します。私たちは、最新の生成アーキテクチャがこれらの評価ペルソナを高い一貫性でインスタンス化して維持できることを示し、現実世界のコンセンサス変動をより厳密に反映する、多元的で視点に依存したベンチマークの形式を可能にします。しかし、我々は、逐次推論と確率的プロンプト摂動下でのこれらのシミュレートされた評価器の安定性をさらに分析し、状態空間ドリフトと意味論的不一致として現れるペルソナの一貫性の体系的な低下を明らかにしました。これらの発見は、静的な位置合わせの制約では、長期にわたって堅牢な評価動作を維持するには不十分であることを示唆しています。その代わりに、私たちは、一貫した認知エミュレーションを維持するために、生成システム内に動的で実行可能性主導の制御メカニズムを組み込む必要性を主張します。この研究は、ペルソナベースの評価を潜在表現多様体上の構造化された動的システムとして枠組み化することで、AI 評価に対する、より適応的で人間と連携した、状況に応じたアプローチの基盤を提供します。

原文 (English)

A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI

Current alignment paradigms for generative artificial intelligence rely predominantly on monolithic benchmarking frameworks that reduce the plurality of human judgment to aggregated statistical baselines, thereby obscuring cultural, demographic, and contextual variability in evaluation. We introduce a state-space constrained emulation framework for AI evaluation that replaces singular assessment functions with a structured manifold of synthetic cognitive profiles representing diverse human perspectives. We show that modern generative architectures can instantiate and maintain these evaluative personas with high consistency, enabling a form of pluralistic, perspective-dependent benchmarking that more closely reflects real-world consensus variability. However, we further analyze the stability of these simulated evaluators under sequential inference and stochastic prompt perturbations, revealing systematic degradation in persona coherence that manifests as state-space drift and semantic inconsistency. These findings suggest that static alignment constraints are insufficient for sustaining robust evaluative behavior over time. Instead, we argue for the necessity of embedding dynamic, viability-driven regulatory mechanisms within generative systems to preserve coherent cognitive emulation. By framing persona-based evaluation as a structured dynamical system over latent representation manifolds, this study provides a foundation for more adaptive, human-aligned, and context-sensitive approaches to AI evaluation.

2026-06-01 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

予測を活用した推論の工業化: 信頼性の高い GenAI およびエージェント システム評価のための GLIDE ライブラリ

エージェント システムの信頼性の高い評価には、有効な不確実性を伴う不偏推定が必要ですが、標準的な手法では、コストのかかる人間によるアノテーションと、ジャッジとしての偏った LLM プロキシの間を行き来します。予測パワー推論 (PPI) は、両方を組み合わせて有効な信頼区間を持つ偏りのない推定値を生成しますが、そのさまざまな手法は部分的な実装の下で論文に散在したままです。平均推定に特化した scipy スタイルの API の下で、最先端の PPI 推定器 (PPI++、層化 PPI、Predict-Then-Debias とその層化バリアント、アクティブ統計推論) とサンプラー (均一、層化、アクティブ、コスト最適化) を統合するオープンソース Python ライブラリである GLIDE を紹介します。 GLIDE には、再現可能なモンテカルロ検証スイート、手法選択のための経験に基づいたデシジョン ツリー、同等の精度でのアノテーションの大幅な節約を示すエージェント評価ケース スタディが付属しています。 GLIDE パッケージは次の URL で入手できます: https://github.com/EmertonData/glide

原文 (English)

Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation

Reliable evaluation of agentic systems requires unbiased estimates with valid uncertainty, but standard practice navigates between costly human annotation and biased LLM-as-judge proxies. Prediction-powered inference (PPI) combines both into debiased estimates with valid confidence intervals, yet its various methods remain scattered across papers under partial implementations. We introduce GLIDE, an open-source Python library that unifies state-of-the-art PPI estimators (PPI++, Stratified PPI, Predict-Then-Debias and its stratified variants, Active Statistical Inference) and samplers (uniform, stratified, active, cost-optimal) under a scipy-style API specialized to mean estimation. GLIDE ships with a reproducible Monte Carlo validation suite, an empirically grounded decision tree for method selection, and an agentic evaluation case study showing substantial annotation savings at equivalent precision. The GLIDE package is available at this URL: https://github.com/EmertonData/glide

2026-06-01 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

TraceGraph: エージェントの軌跡を診断および改善するための共有意思決定ランドスケープ

エージェントのベンチマークでは、豊富なインタラクションの軌跡が記録されることが増えていますが、評価によって各ロールアウトが合格率や報酬スコアに引き下げられることがよくあります。リリースされたマルチモデル エージェントの軌跡を共有の意思決定ランドスケープに変えるグラフベースのフレームワークである TraceGraph を紹介します。 TraceGraph は、タスクごとに、モデル ID が導入される前に、プールされたロールアウトから観察可能なアクションと観察の状態に関するグラフを構築します。次に、結果に基づいた生産コアとトラップ領域をオーバーレイし、各ロールアウトをアクセス、トラップ露出、修復の 3 つのイベントで要約します。 TraceGraph プロファイルは、5 つのベンチマーク スプリットにまたがる軌跡全体で、集計スコアによって隠されたナビゲーションの違いを明らかにし、トラップの回避とそこからの回復のどちらに報酬を与えるかがスプリットによって異なることを示します。同じ TraceGraph ランドスケープは、SWE ベンチのトラップ対応回復パイプラインも動機付けます。実行時検出器は、履歴トラップ領域に一致する状態で起動され、その後、軽量継続ポリシーが同じプレフィックスから評価されます。起動された状態では、最適なプールされた単一要素ポリシーにより、プロバイダー固有のアクティブ コンポーネントを使用して、プロバイダーごとに起動されたサブセットで正式な解決率が 40.4% から 43.5% に、共通起動されたインスタンスで 41.0% から 44.8% に上昇します。全体として、TraceGraph は、どのようなエージェント ベンチマーク テストを行うか、共有ランドスケープ上でモデルが分岐する場所、および障害領域が下流の改善をどのように導くことができるかを尋ねるためのプロセス ボキャブラリーを提供します。

原文 (English)

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories

Agent benchmarks increasingly record rich interaction trajectories, yet evaluation often reduces each rollout to a pass rate or reward score. We introduce TraceGraph, a graph-based framework that turns released multi-model agent trajectories into shared decision landscapes. For each task, TraceGraph builds a graph over observable action-observation states from pooled rollouts before model identity is introduced. It then overlays outcome-informed productive cores and trap regions, and summarizes each rollout with three events: Access, Trap exposure, and Repair. Across trajectories spanning five benchmark splits, TraceGraph profiles reveal navigation differences hidden by aggregate scores and show that splits differ in whether they reward avoiding traps or recovering from them. The same TraceGraph landscape also motivates a trap-aware recovery pipeline for SWE-bench: aruntime detector fires on states matching historical trap regions, then lightweight continuation policies are evaluated from the same prefix. On fired states, the best pooled single-factor policy raises official resolved rate from 40.4% to 43.5% on the per-provider fired subset and from 41.0% to 44.8% on common-fired instances, with provider-specific active components. Overall, TraceGraph provides a process vocabulary for asking what agent benchmarks test, where models diverge on a shared landscape, and how failure regions can guide downstream improvement.

2026-06-01 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

安全閾値をニューロンスパイキング閾値として再解釈する

代理安全対策 (SSM) は、自動運転の状況における交通リスクの評価に広く利用されています。しかし、SSM ベースの評価の大部分では、固定しきい値が採用されており、持続する境界線状態に対する人間の反応や、短期間の高リスクピークに対する反応を捉えることができません。本研究は、生物学にインスピレーションを得た SSM 閾値の再解釈を提案しています。これは、複数の SSM 入力がスパイキング ニューラル ネットワーク (SNN) に結合された、リーキー統合発射 (LIF) ニューロンのスパイク閾値としてモデル化されています。 SNN は、人間のブレーキの開始に合わせてスパイクを発するように訓練されています。トレーニング データは、CARLA/Unreal を備えた 3D-CoAutoSim プラットフォームと 6-DOF モーション プラットフォームを使用した、制御された車追従実験で記録され、誘発された重大なイベントが生成されました。結果は、学習されたスパイク アクティビティがシナリオ全体でブレーキ動作と定性的に一致しており、しきい値の交差だけでは一貫して説明できない反応を捕捉していることを示しています。さらに、参加者全体の分析により、学習された入力しきい値は比較的一貫したままである一方、学習された減衰係数は SSM の異なる時間感度をエンコードしていることが示されています。この研究の結果は、スパイクのダイナミクスが客観的な SSM と主観的な人間の安全認識の収束を促進するメカニズムとして機能する可能性があることを示しています。

原文 (English)

Reinterpreting Safety Thresholds as Neuron Spiking Thresholds

Surrogate Safety Measures (SSMs) are extensively utilised in the evaluation of traffic risk in automated driving contexts. However, the majority of SSM-based evaluations employ fixed thresholds that fail to capture the human response to sustained borderline conditions or the reaction to brief, high-risk peaks. The present work proposes a biologically inspired reinterpretation of SSM thresholds. This is modelled as spiking thresholds of leaky integrate-and-fire (LIF) neurons, with multiple SSM inputs combined into a spiking neural network (SNN). The SNN is trained to emit spikes that are aligned with human braking onsets. The training data was recorded in a controlled car-following experiment using the 3D-CoAutoSim platform with CARLA/Unreal and a 6-DOF motion platform, where induced critical events were generated. The results demonstrate that the learned spiking activity qualitatively aligns with braking behaviour across scenarios and captures reactions that are not consistently explained by threshold crossings alone. Analysis across participants further indicates that learned input thresholds remain relatively consistent, while learned decay factors encode different temporal sensitivities for the SSMs. The findings of this study indicate that spiking dynamics may serve as a mechanism to facilitate the convergence of objective SSMs with subjective human safety perception.

2026-06-01 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

NumLeak: 基礎モデルの潜在ラベルとしての公開数値ベンチマーク

公開された数値ベンチマークは事前トレーニングに表示されるため、日付の条件による評価は、サンプル外のスキルではなく、記憶された再現率を測定している可能性があります。 NumLeak は、実稼働モデル上の API 境界プローブとオープン因果 LM 上のホワイトボックス制御検証を組み合わせた測定フレームワークです。最上位のフロンティア LLM は、3 シードでプールされたピアソン r=0.97 ~ 0.99 でのファーマ・フランス市場の超過リターンを思い出しますが、5 つの兄弟要素では 25bps 以内で 0.15 以内に留まっています。同等の忠実度は、米国の失業率、CPI インフレ、NOAA の気温にも現れています。最近のリリースのホールドアウトでは、解析率は 21 ~ 57% に低下しますが、応答した月の r は約 0.99 にとどまります。これは、記憶されたチャネルが予測するリジェクトまたはリコールの非対称性です。ホワイトボックス実験は用量反応を再現し、logprob ランキングはオープンエンド生成で見逃した記憶を検出します。これは、クローズド API ブラックボックス プローブがチャネルを過小評価していることを意味します。 r=0.74 で真の Mkt-RF と相関するソネットの「市場センチメントに対する日付」回帰は、モデル自体の再現率が残差化されると r=0.02 に崩壊します。 1 行のシステムプロンプト防御は、概念的および歴史的物語のクエリに対してほぼゼロのユーティリティコストで設定された非適応的なシングルターンサフィックス攻撃を 99.8% ブロックします。

原文 (English)

NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation Models

Public numeric benchmarks appear in pretraining, so an evaluation that conditions on a date may be measuring memorized recall rather than out-of-sample skill. We introduce NumLeak, a measurement framework that combines API-boundary probes on production models with a white-box controlled validation on an open causal LM. Top-tier frontier LLMs recall the Fama-French market excess return at 3-seed pooled Pearson r=0.97-0.99 while staying within 0.15 within-25bps on the five sibling factors; comparable fidelity appears on U.S. unemployment, CPI inflation, and NOAA temperature. On a recent-release holdout, parse rate collapses to 21-57% but r stays at approximately 0.99 on months answered, the refuse-or-recall asymmetry a memorized channel predicts. The white-box experiment reproduces the dose-response, and logprob ranking detects memorization that open-ended generation misses, implying closed-API black-box probes understate the channel. A Sonnet "date to market-sentiment" regression that correlates with true Mkt-RF at r=0.74 collapses to r=0.02 once the model's own recall is residualized out. A one-line system-prompt defense blocks 99.8% of a non-adaptive single-turn suffix attack set at near-zero utility cost on conceptual and historical-narrative queries

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

反事実的な評価により、臨床 LLM とエージェントの隠れた能力プロファイルが明らかになる

2 つの臨床 AI システムは、カバレッジベースのルーブリックではほぼ同じスコアを獲得できますが、患者の入力が変化すると根本的に異なる動作をします。1 つは新しい臨床信号に一致するように推奨事項を更新しますが、もう 1 つはそれに関係なく同じ出力を生成します。因果感受性スコア (CSS) を導入します。これは、臨床的に意味のある 5 つの次元 (バイオマーカーの反転、前治療の失敗、バイオマーカーの除去、手術状態の変化、ステージの摂動) に沿って腫瘍腫瘍ボードの症例を変異させる事前登録された介入指標であり、各モデルが事前に登録された正しい方向で推奨事項を更新するかどうかを {0、0.5、1.0} スケールを使用してスコア付けします。カバレッジベースの加重リコール指標であるコンセンサス マッチ スコア (CMS) に対してベンチマークを行ったところ、224 件のケースにわたる単発推論で評価された 3 つのラボの 6 つのフロンティア モデルが、ほぼ逆の順位でランク付けされました。6 つのモデルすべてがランクを変更し、CMS で最も悪いモデルが CSS で最も優れたモデルになり、上位中位の 1 つの CMS モデルが CSS で最下位にランクされました。さらに、普遍的な安全性の盲点も明らかになりました。つまり、すべてのフロンティア モデルは手術状態の介入で失敗します (ファミリー D では最大 17.2% の CSS)。これは CMS では明らかにされていません。この指標は、ツールを使用するエージェントにも伝達されます。ReAct スタイルの実験では、ツールの使用により 6 つのモデルのうち 5 つのモデルで CSS が向上しました (+2.5 ~ +20.3 パーセント ポイント)。それでも、CSS が最も低いモデルは同じグラフ セクションを取得し、依然として推奨事項を更新できません。これは、反事実の評価下でのみ表示される構造的な応答性の欠陥を明らかにしています。裁判官間の複製と 3 人の評価者の医療専門家による検証により、総合的な結果が確認されます。 CSS のような事前登録された介入指標は、臨床 AI エージェントのカバレッジベースの評価を補完します。これらは、カバレッジ指標では見逃される応答性を捕捉し、将来のエージェント RL システムに候補となる密な報酬シグナルを提供します。

原文 (English)

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one updates its recommendations to match the new clinical signal, while the other produces the same output regardless. We introduce the Causal Sensitivity Score (CSS), a pre-registered interventional metric that mutates oncology tumor-board cases along five clinically meaningful dimensions - biomarker flips, prior-treatment failures, biomarker removals, surgery-status changes, and stage perturbations - and scores whether each model updates its recommendations in the pre-registered correct direction using a {0, 0.5, 1.0} scale. Benchmarked against the Consensus Match Score (CMS), a coverage-based weighted recall metric, six frontier models from three labs evaluated in single-shot inference across 224 cases rank in nearly opposite orders: all six models change rank, the CMS-worst model becomes CSS-best, and one upper-mid CMS model ranks last on CSS. We further surface a universal safety blind spot: every frontier model fails on surgery-status interventions (at most 17.2% CSS on Family D), a finding CMS does not expose. The metric also transfers to tool-using agents: in a ReAct-style experiment, tool use improves CSS for five of six models (+2.5 to +20.3 percentage points), yet the lowest-CSS model retrieves the same chart sections and still fails to update its recommendations - revealing a structural responsiveness deficit visible only under counterfactual evaluation. Cross-judge replication and three-rater medical-professional validation confirm the aggregate findings. Interventional pre-registered metrics like CSS complement coverage-based evaluation for clinical AI agents: they capture responsiveness that coverage metrics miss and offer a candidate dense reward signal for future agentic RL systems.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

OrcaRouter: オフラインとオンラインのハイブリッド学習を備えた本番指向の LLM ルーター

それぞれが異なる機能と推論コストを備えた大規模な言語モデルの急速な開発により、実際的な展開上の疑問が生じます。受信したリクエストを考慮して、どのモデルがそれを処理すべきでしょうか?私たちは、LinUCB ベースの語彙および文埋め込み機能に対するコンテキスト バンディットと、オフラインとオンラインのハイブリッド学習プロトコルを組み合わせた実稼働指向の LLM ルーターである OrcaRouter を紹介します。オフラインでは、OrcaRouter は、厳選されたルーティング プロンプトのセットで各候補モデルを評価することによって完全な情報フィードバックを取得し、アームごとに 1 つのリッジ リグレッサーを適合させるために使用される報酬行列を生成します。デプロイメント時にこれらのパラメータから初期化され、オプションでバンディットのフィードバックから学習を継続し、報酬を観察した後に選択したモデルのアームのみを更新できます。 RouterArena への提出時点 (2026 年 5 月 20 日)、OrcaRouter-Adaptive は、アリーナ スコア 72.08 で公開 RouterArena リーダーボードで 2 位にランクされ、1,000 クエリあたり 1.00 米ドルのコストで 75.54% の精度を達成しました。

原文 (English)

OrcaRouter: A Production-Oriented LLM Router with Hybrid Offline-Online Learning

The rapid development of large language models, each with distinct capabilities and inference costs, raises a practical deployment question: given an incoming request, which model should handle it? We present OrcaRouter, a production-oriented LLM router that combines a LinUCB-based contextual bandit over lexical and sentence-embedding features with a hybrid offline-online learning protocol. Offline, OrcaRouter obtains full-information feedback by evaluating each candidate model on a curated set of routing prompts, yielding a reward matrix used to fit one ridge regressor per arm. At deployment time, it initializes from these parameters and can optionally continue learning from bandit feedback, updating only the selected model's arm after observing its reward. At the time of our RouterArena submission (May 20, 2026), OrcaRouter-Adaptive ranked second on the public RouterArena leaderboard with an arena score of 72.08, achieving 75.54% accuracy at a cost of USD 1.00 per 1,000 queries.

2026-06-01 13:00 JSTarXiv cs.AIビジネス/資金調達

OpenSTBench: 音声翻訳のセマンティック評価を超えて

音声翻訳システムは、音声からテキストへの翻訳 (S2TT)、音声から音声への翻訳 (S2ST)、オフライン翻訳、ストリーミング生成にまで範囲が広がり、モダリティ、音声実現、タイミング動作が異なる出力を生成します。既存の評価手法では、翻訳品質、音声品質、時間品質などの重要な側面が評価されますが、これらの側面は別のプロトコルに基づいて評価されることが多く、異種システムを包括的に比較することが困難になります。このギャップに対処するために、異種の音声翻訳出力を共有の評価形式に編成する統合多次元評価フレームワークである OpenSTBench を紹介します。 OpenSTBench は、オフラインおよびストリーミング設定で S2TT と S2ST システムの両方をサポートし、翻訳品質、音声品質、話者の保存、感情とパラ言語の忠実度、時間的一貫性、遅延を共同で評価します。代表的な音声翻訳システムの実験を通じて、高い翻訳品質を備えたシステムであっても、音声品質および時間品質が大幅に異なる可能性があることを示しました。 OpenSTBench は、こ​​れらの次元間の違いを分析し、音声翻訳システムのアプリケーション指向の比較をサポートするための再現可能なプロトコルを提供します。コードとデータセットは https://github.com/sjtuayj/OpenSTBench で入手できます。

原文 (English)

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

Speech translation systems increasingly span speech-to-text translation (S2TT), speech-to-speech translation (S2ST), offline translation, and streaming generation, producing outputs that differ in modality, speech realization, and timing behavior. Existing evaluation practices assess important aspects such as translation quality, speech quality, and temporal quality, but these aspects are often evaluated under separate protocols, making it difficult to compare heterogeneous systems comprehensively. To address this gap, we present OpenSTBench, a unified multidimensional evaluation framework that organizes heterogeneous speech translation outputs into a shared evaluation format. OpenSTBench supports both S2TT and S2ST systems in offline and streaming settings, and jointly evaluates translation quality, speech quality, speaker preservation, emotion and paralinguistic fidelity, temporal consistency, and latency. Through experiments on representative speech translation systems, we show that systems with strong translation quality can still differ substantially in speech quality, as well as in temporal quality. OpenSTBench provides a reproducible protocol for analyzing these cross-dimensional differences and supporting application-oriented comparison of speech translation systems. The code and datasets are available at https://github.com/sjtuayj/OpenSTBench.

2026-06-01 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

市場予測解決のためのマルチエージェント AI Oracle システムの設計と評価

予測市場は集合知を集約して不確実な出来事を予測しますが、その有用性は信頼できる結果の解決に依存します。既存のオラクル システムは、高速ではあるが脆弱な自動化と、正確ではあるがコストのかかる人間による調停をトレードオフにしています。単一 LLM オラクルは意味のある精度を実現しますが、自己修正メカニズムがなく、基礎となるモデルのすべての障害モードを継承します。マルチエージェント LLM アーキテクチャが、単一モデルのベースラインよりも Oracle の解決精度を向上できるかどうかを評価します。私たちは、KalshiBench からの 1,189 件の解決済み予測市場質問について、独立した集計と熟議によるコンセンサスを、単一 LLM ベースライン (GPT-5 Nano、DeepSeek V3、および Llama-3.3-70B) と比較します。すべてのエージェントは Exa を通じて共通の証拠レイヤーを共有し、検索は公開日によってフィルタリングされ、推論と検索の品質を分離します。信頼度加重投票による独立した集計は、83.43 パーセントという最高の精度を達成し、最高の個別モデルを 1.01 パーセントポイント上回りました。熟議的コンセンサスでは精度が約 76% まで低下し、すべての単一モデルのベースラインを下回ります。これは、確信を持って間違ったモデルが正しいモデルを反転させる議論中にエラーが伝播したことに起因します。モデル間の誤差相関 (0.529 ~ 0.689) は、なぜ集約ゲインが理論上のコンドルセの上限を下回り、アンサンブル アプローチに根本的な制限を課すのかを説明しています。多くの疑問点はマルチエージェント アーキテクチャによる修正に抵抗しており、人間による仲裁へのエスカレーションの動機となっています。私たちは、AI と人間のハイブリッド オラクル システムのルーティング基準を提案します。全会一致で信頼性の高い質問のみを自動解決すると、データセットの 47 パーセントで 97.87 パーセントの精度が得られます。エージェント間の意見の不一致により、残りの部分は人間によるレビューのためにフラグが立てられます。

原文 (English)

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution

Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing oracle systems tradeoff fast but brittle automation against accurate but costly human arbitration. Single-LLM oracles achieve meaningful accuracy but inherit all failure modes of their underlying model with no self-correction mechanism. We evaluate whether multi-agent LLM architectures can improve oracle resolution accuracy over single-model baselines. We compare independent aggregation and deliberative consensus against single-LLM baselines (GPT-5 Nano, DeepSeek V3, and Llama-3.3-70B) on 1,189 resolved prediction market questions from KalshiBench. All agents share a common evidence layer through Exa, with retrieval filtered by publication date to isolate reasoning from retrieval quality. Independent aggregation with confidence-weighted voting achieves the highest accuracy at 83.43 percent, outperforming the best individual model by 1.01 percentage points. Deliberative consensus degrades accuracy to approximately 76 percent, below every single-model baseline, attributed to error propagation during debate where confidently wrong models flip correct ones. Error correlations across models (0.529-0.689) explain why aggregation gains fall short of the theoretical Condorcet ceiling, placing a fundamental limit on ensemble approaches. Many questions resist correction by any multi-agent architecture, motivating escalation to human arbitration. We propose routing criteria for hybrid AI-human oracle systems: auto-resolving only unanimous, high-confidence questions yields 97.87 percent accuracy on 47 percent of the dataset, with inter-agent disagreement flagging the remainder for human review.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM Unlearning の属性を解除して忘れる

大規模言語モデル (LLM) の急速な発展により、トレーニングに不適切なデータが使用されることへの懸念が生じ、LLM の非学習への関心が高まっています。既存の LLM 未学習アプローチの多くは、忘却集合の損失の最大化など、予測損失の最適化に依存していますが、過剰忘却やモデルの有用性の低さなどの重大な問題に直面することがよくあります。これらに対処するために、この論文では、代わりにデータの帰属をゼロにすることの 1 つとして、LLM アンラーニングの最適化目標を斬新に組み立てています。特に、DareUと呼ばれるデータ帰属報酬に基づいた最初のLLM非学習フレームワークを提案します。これは、データ所有者を忘れた場合に生成された応答の帰属スコアを減らす(つまり、帰属を解除する)ことによってLLMを更新する強化学習を実行します。アトリビューションの効率的な近似として LLM 分類器を使用した経験的評価では、DareU が忘却品質とモデルの有用性のバランスを適切に保ちながら効果的なアンラーニングを達成することで、既存のベースラインを上回るパフォーマンスを示しています。

原文 (English)

De-attribute to Forget for LLM Unlearning

The rapid development of large language models (LLMs) has raised concerns on the use of inappropriate data for training, which has led to a growing interest in LLM unlearning. Many existing LLM unlearning approaches rely on optimizing prediction loss(es), such as maximizing the loss on the forget set, but often face critical issues like over-forgetting and poor model utility. To address them, this paper novelly frames the optimization objective for LLM unlearning as one of zeroing out data attribution instead. In particular, we propose the first LLM unlearning framework based on data attribution rewards called DareU that performs reinforcement learning to update the LLM by reducing the attribution score of its generated responses (i.e., de-attributing) to the forget data owners. Empirical evaluation using an LLM classifier as an efficient approximation of attribution shows that DareU outperforms existing baselines by achieving effective unlearning while balancing forget quality and model utility well.

2026-06-01 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

インスタンス マッチングの再定義: パノプティック セグメンテーション評価におけるパーツ認識マッチングのための統合フレームワーク

Panoptic Quality (PQ) メトリクスは、インスタンスとセマンティック セグメンテーションを共同で評価するための標準です。ただし、その元の定義は、予測セグメントとグランド トゥルース セグメント間の 1 対 1 のマッチングに依存しており、これは IoU しきい値が 0.5 を超える場合にのみ簡単になります。 0.5 未満では、十分に調査されていない問題空間に複数のマッチング戦略が現れます。我々は、セグメントマッチングを制約付き二部代入問題として再キャストすることにより、この空間を体系的に解明します。予測側と真実側の次数を独立して境界付けると、1 対 1、多対 1、1 対多、および多対多の 4 つのマッチング戦略が得られます。最初の 3 つは PQ フレームワーク内で明確に定義されていますが、多対多は PQ フレームワークの外にあることを示します。これらの戦略は、インスタンスが断片化している場合、隣接するオブジェクトの輪郭を描くのが難しい場合、または注釈にノイズが多い場合に重要になります。私たちのフレームワークの中心となるのは、一致するエッジではなくグラウンド トゥルースと予測セグメントに固定された、TP、FN、および FP の頂点ベースの計算です。さらに、このフレームワークが部分認識パノプティック セグメンテーションに自然に拡張されることを示し、生物医学データの部分認識評価を検討します。構成可能なケーススタディ全体にわたって、しきい値とマッチング戦略のさまざまな組み合わせが実際にどのように動作するかを報告します。 Panoptica に基づいて構築された統合オープンソース パッケージをリリースします。ボロノイベースの領域ごとの分析、部分認識評価、およびしきい値曲線下面積の計算を構成可能なオプションとして公開します。

原文 (English)

Redefining Instance Matching: A Unified Framework for Part-Aware Matching in Panoptic Segmentation Evaluation

The Panoptic Quality (PQ) metric is the standard for jointly evaluating instance and semantic segmentation. However, its original definition relies on a One-to-One matching between predicted and ground truth segments, which is only straightforward when the IoU threshold exceeds 0.5. Below 0.5, multiple matching strategies emerge in a poorly explored problem space. We systematically elucidate this space by recasting segment matching as a constrained bipartite assignment problem. Independently bounding the prediction- and ground-truth-side degrees yields four matching strategies: One-to-One, Many-to-One, One-to-Many, and Many-to-Many. We show that the first three are well-defined within the PQ framework, while Many-to-Many falls outside it. These strategies become relevant when instances are fragmented, adjacent objects are difficult to delineate, or annotations are noisy. Central to our framework is a vertex-based accounting of TP, FN, and FP, anchored to ground truth and predicted segments rather than to matching edges. We further show that the framework extends naturally to part-aware panoptic segmentation, and we explore part-aware evaluation on biomedical data. Across configurable case studies we report how different combinations of thresholds and matching strategies behave in practice. We release a unified open-source package built on Panoptica. It exposes Voronoi-based region-wise analysis, part-aware evaluation, and Area Under Threshold Curve computations as configurable options.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

LLM バイアス評価: 職業および犯罪シナリオにおける性別、人種、年齢の格差

大規模言語モデル (LLM) が一か八かの意思決定にますます影響を与えるため、LLM バイアスの評価は重要です。この論文は、主要な LLM における性別、人種、年齢の格差の包括的な評価を提供し、バイアス緩和の取り組みがしばしば新たな公平性のトレードオフを生み出すことを明らかにしています。 LLM の最近の進歩は注目に値しますが、さまざまな制約により、企業での広範な導入は依然として限られています。このペーパーでは、LLM の使いやすさ、信頼性、公平性に影響を与える重要な問題である LLM のバイアスについて検証します。私たちの研究では、2024 年にリリースされた 4 つの主要な LLM (Gemini 1.5 Pro、Llama 3 70B、Claude 3 Opus、GPT-4o) にわたって、職業シナリオにおける性別の偏り、および犯罪シナリオにおける性別、年齢、人種の偏りを評価しています。調査結果によると、LLM はさまざまな職業において、男性キャラクターよりも女性キャラクターを頻繁に描写することが多く、米国の BLS データから 37% の乖離が示されています。犯罪シナリオでは、米国 FBI データからの逸脱は、性別で 54%、人種で 28%、年齢で 17% です。重要なことに、ジェンダーと人種の偏見を軽減する取り組みは、あるサブクラスを過剰に指数化し、潜在的に格差を悪化させる可能性のある結果につながることが多いことを観察しています。これは、現在の偏見緩和手法の限界を浮き彫りにし、より効果的なアプローチの必要性を強調する「バイアス緩和のパラドックス」です。

原文 (English)

LLM Bias Evaluation: Gender, Racial, and Age Disparities in Occupational and Crime Scenarios

LLM bias evaluation is critical as large language models (LLMs) increasingly influence high-stakes decisions. This paper provides a comprehensive assessment of gender, racial, and age disparities in leading LLMs, revealing that debiasing efforts often create new fairness trade-offs. Recent advancements in LLMs have been notable, yet widespread enterprise adoption remains limited due to various constraints. This paper examines bias in LLMs - a crucial issue affecting their usability, reliability, and fairness. Our study evaluates gender bias in occupational scenarios and gender, age, and racial bias in crime scenarios across four leading LLMs released in 2024: Gemini 1.5 Pro, Llama 3 70B, Claude 3 Opus, and GPT-4o. Findings reveal that LLMs often depict female characters more frequently than male ones in various occupations, showing a 37% deviation from US BLS data. In crime scenarios, deviations from US FBI data are 54% for gender, 28% for race, and 17% for age. Critically, we observe that efforts to reduce gender and racial bias often lead to outcomes that may over-index one sub-class, potentially exacerbating disparities - a "debiasing paradox" that highlights the limitations of current bias mitigation techniques and underscores the need for more effective approaches.

2026-06-01 13:00 JSTarXiv cs.AIビジネス/資金調達

逐次的な意思決定による選択のためのデータ値の統合と最適化

データ選択は、データ評価の重要な下流アプリケーションとして浮上していますが、選択にデータ値を使用するための理論的基盤はまだ調査されていません。データ選択を、最適な選択シーケンスが動的計画法から生じ、データ値はこの最適なシーケンスのエンコードとして理解できる、逐次的意思決定問題として再定式化します。このフレームワークは、Data Shapley などの既存の手法を近似動的計画法のレンズを通して統合および再解釈し、それらが逐次問題に対する近視眼的な線形近似であることを明らかにします。さらに、サブモジュール性の下でユーティリティの曲率によって選択の最適性がどのように低下​​するかを分析し、これらの近似がいつ失敗するのか、なぜ失敗するのかを説明します。理論と実践の橋渡しをするために、証明可能な保証を備えたスケーラブルで貪欲な選択を可能にしながら、サブモジュール構造を維持する効率的な二部グラフベースのサロゲートを提案します。古典的な ML ベンチマークと大規模な LLM 微調整データ選択に関する実験では、既存の方法に比べて大幅な改善が見られます。コードは https://github.com/frankhlchi/SeqDataVal で公開されています。

原文 (English)

Unifying and Optimizing Data Values for Selection via Sequential Decision-Making

Data selection has emerged as a crucial downstream application of data valuation, yet the theoretical foundations for using data values in selection remain underexplored. We reformulate data selection as a sequential decision-making problem where the optimal selection sequence arises from dynamic programming, and data values can be understood as encodings of this optimal sequence. This framework unifies and reinterprets existing methods like Data Shapley through the lens of approximate dynamic programming, revealing them as myopic linear approximations to the sequential problem. We further analyze how selection optimality degrades with utility curvature under submodularity, explaining when and why these approximations fail. To bridge theory and practice, we propose an efficient bipartite graph-based surrogate that preserves submodular structure while enabling scalable greedy selection with provable guarantees. Experiments on classical ML benchmarks and large-scale LLM fine-tuning data selection demonstrate substantial improvements over existing methods. Code is publicly available at https://github.com/frankhlchi/SeqDataVal

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

項目応答理論による LLM-as-a-Judge の信頼性の診断

LLM-as-a-Judge は自動評価で広く使用されていますが、既存の検証手法は主に観察された出力のレベルで動作し、LLM ジャッジ自体が安定した信頼できる測定手段として機能するかどうかについての洞察は限られています。この制限に対処するために、項目応答理論 (IRT) に基づいた、裁判官としての LLM の信頼性を評価するための 2 段階の診断フレームワークを導入します。このフレームワークは、IRT の段階的応答モデル (GRM) を採用し、2 つの相補的な側面に沿って信頼性を形式化します: (1) 即時の変動下での測定動作の安定性として定義される本質的一貫性、および (2) 人間の品質評価との対応を捉える人間の整合性。私たちは、このフレームワークを使用して多様な LLM 裁判官を実証的に調査し、IRT-GRM を活用すると、体系的に判断を診断するための解釈可能なシグナルが得られることを示します。これらの信号は、LLM-as-a-Judge の信頼性を検証し、信頼性の低さの潜在的な原因を特定するための実践的なガイダンスを提供します。

原文 (English)

Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether LLM judges themselves function as stable and reliable measurement instruments. To address this limitation, we introduce a two-phase diagnostic framework for assessing reliability of LLM-as-a-Judge, grounded in Item Response Theory (IRT). The framework adopts Graded Response Model (GRM) of IRT and formalizes reliability along two complementary dimensions: (1) intrinsic consistency, defined as the stability of measurement behavior under prompt variations, and (2) human alignment, capturing correspondence with human quality assessments. We empirically examine diverse LLM judges with this framework, and show that leveraging IRT-GRM yields interpretable signals for diagnosing judgments systematically. These signals provide practical guidance for verifying reliablity of LLM-as-a-Judge and identifying potential causes of unreliability.

2026-06-01 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

LH ベンチ: 企業の主観的なタスクに関する長期的なエージェントのスキルに基づいた評価

大規模な言語モデルは、数学やプログラミングなどの客観的に検証可能なタスクに優れており、評価は単体テストや単一の正解に限定されます。対照的に、現実世界の企業の仕事は主観的で状況に依存することが多く、成功は組織の目標、ユーザーの意図、および長いマルチツールのワークフロー全体で生成される中間成果物の品質にかかっています。 LH-Bench は、バイナリの正確性を超えて、主観的なエンタープライズ タスクの自律的で長期的な実行をスコアリングする 3 本柱の評価設計です。その柱は、(i) 主観的な作品を採点するために必要なドメインコンテキストを LLM 審査員に提供する専門家に基づいたルーブリック、(ii) 段階的な報酬シグナル (コンテンツ タスクの章レベルの注釈など) を可能にする精選されたグラウンドトゥルース アーティファクト、および (iii) 収束検証のためのペアごとの人間の好みの評価です。ドメイン作成のルーブリックは、LLM 作成のルーブリックよりも大幅に信頼性の高い評価シグナルを提供すること (カッパ = 0.60 対 0.46)、および人間の好みの判断によって同じ最上位の分離が確認される (p​​ < 0.05) ことを示し、専門家に基づいた評価が信頼性を犠牲にすることなく拡張できることを示しています。当社は公開データセットをリリースし、Figma-to-code (MCP を介した Figma API に対する 33 の実際の .fig タスク) とプログラマティック コンテンツ (毎日 30 人以上のユーザーにサービスを提供するコース プラットフォーム上で個別に評価される 183 の章からなる 41 コース) の 2 つの環境で結果をレポートします。

原文 (English)

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

Large language models excel on objectively verifiable tasks such as math and programming, where evaluation reduces to unit tests or a single correct answer. In contrast, real-world enterprise work is often subjective and context-dependent: success hinges on organizational goals, user intent, and the quality of intermediate artifacts produced across long, multi-tool workflows. We introduce LH-Bench, a three-pillar evaluation design that moves beyond binary correctness to score autonomous, long-horizon execution on subjective enterprise tasks. The pillars are: (i) expert-grounded rubrics that give LLM judges the domain context needed to score subjective work, (ii) curated ground-truth artifacts that enable stepwise reward signals (e.g., chapter-level annotation for content tasks), and (iii) pairwise human preference evaluation for convergent validation. We show that domain-authored rubrics provide substantially more reliable evaluation signals than LLM-authored rubrics (kappa = 0.60 vs. 0.46), and that human preference judgments confirm the same top-tier separation (p < 0.05), evidence that expert-grounded evaluation can scale without sacrificing reliability. We release public datasets and report results on two environments: Figma-to-code (33 real .fig tasks against the Figma API via MCP) and Programmatic content (41 courses comprising 183 individually-evaluated chapters on a course platform serving 30+ daily users).

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

LLM エージェント スキルの反事実追跡監査

大規模言語モデル エージェントは、エージェント スキルによってますます強化されています。現在のスキルの評価方法は依然として限られています。導入されたベンチマークのほとんどは、スキルがアタッチされる前後の合格率のみを報告し、スキルをエージェントの動作に対するブラックボックスの変更として扱います。スキルがエージェントの動作をどのように変化させるかを測定するためのフレームワークである、Counterfactual Trace Auditing (CTA) を紹介します。 CTA は、同じタスクに関してスキル エージェント トレースを持つそれぞれのトレースと、スキルを持たないエージェント トレースをペアにし、両方のトレースを目標指向フェーズに分割し、フェーズを調整して、構造化されたスキル影響パターン (SIP) アノテーションを発行します。これらの注釈は、タスクの結果だけではなく、スキルの行動上の影響を説明します。クロードを使用して、49 のソフトウェア エンジニアリング タスクにわたって SWE-Skills-Bench で CTA をインスタンス化します。監査の結果、明らかな評価のギャップが明らかになりました。合格率の変化は平均で +0.3 パーセント ポイントのみであり、集計効果はほとんどないことがわかります。しかし、CTA は、同じペアのトレース全体で 522 の SIP インスタンスを識別し、合格率がほとんど変わらない場合でも、スキルによってエージェントの動作が大幅に変更されることを示しています。また、監査では、リテラル テンプレートのコピー、オフタスク成果物の作成、過剰な計画、タスクのリカバリなど、合格率では検出できないいくつかの繰り返しの影響も分離されます。 3 つの発見が得られます。まず、ベースラインの高いタスクには、観察されたスキル効果のほとんどが含まれていますが、合格率はすでに飽和しているため、それらの効果を反映できません。第 2 に、中程度のベースライン パフォーマンスを持つタスクは、最も回復可能な利益を示しますが、多くの場合、トークン コストが大幅に高くなります。 3 番目に、支配的な SIP タイプはベースライン バケットによって識別できます。表面アンカリングは天井タスクで最も一般的であり、エッジケース プロンプトはミッドレンジおよびフロア タスクで最も一般的です。これらの規則性により、非公式の故障モードの観察が再現可能な動作測定に変わります。

原文 (English)

Counterfactual Trace Auditing of LLM Agent Skills

Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchmarks report only pass rate before and after a skill is attached, treating the skill as a black box change to agent behavior. We introduce Counterfactual Trace Auditing (CTA), a framework for measuring how a skill changes agent behavior. CTA pairs each with skill agent trace with a without skill counterpart on the same task, segments both traces into goal directed phases, aligns the phases, and emits structured Skill Influence Pattern (SIP) annotations. These annotations describe the behavioral effect of a skill rather than only its task outcome. We instantiate CTA on SWE-Skills-Bench with Claude across 49 software engineering tasks. The resulting audit reveals a clear evaluation gap. Pass rate changes by only +0.3 percentage points on average, suggesting little aggregate effect. Yet CTA identifies 522 SIP instances across the same paired traces, showing that the skills substantially reshape agent behavior even when pass rate is nearly unchanged. The audit also separates several recurring effects that pass rate cannot detect, including literal template copying, off task artifact creation, excess planning, and task recovery. Three findings emerge. First, high baseline tasks contain most of the observed skill effects, although their pass rate is already saturated and therefore cannot reflect those effects. Second, tasks with moderate baseline performance show the most recoverable gain, but often at substantially higher token cost. Third, the dominant SIP type can be identified by baseline bucket: surface anchoring is most common on ceiling tasks and edge-case prompting is most common on mid-range and floor tasks. These regularities turn informal failure mode observations into reproducible behavioral measurements.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

暗記を超えて: 句構造を使用した大規模言語モデルにおける意味一般化の評価

事前トレーニング データの Web スケールは、評価に関する重要な課題を生み出しています。それは、一般化からドメイン外の言語、特に事前トレーニング データではあまり一般的ではない、動的で現実世界のインスタンスに至るまで、事前トレーニング データでよく表現されているケースに関する言語能力を解きほぐすことです。この目的を達成するために、構築文法 (CxG) を活用して、LLM における自然言語理解を系統的に評価するための診断評価を構築します。 CxG は、構文形式を抽象的な非語彙的な意味に明示的にリンクするため、一般化をテストするための心理言語学的に根拠のあるフレームワークを提供します。私たちの新しい推論評価データセットは英語の句構造で構成されており、話者はこの句構造を理解して創造的な具体例を作成するために、ありきたりな具体例を抽象化できることが知られています。私たちの評価データセットは、CxG を使用して 2 つの中心的な質問を評価します。1 つは、モデルが、事前トレーニング データに出現する頻度は低いものの、直感的で人間にとって理解しやすいインスタンスの文の意味論を「理解」できるかどうかです。第 2 に、構文的には同一だが意味が異なる構造を想定して、LLM が適切な構造セマンティクスを展開できるかどうかです。私たちの結果は、GPT-o1 を含む最先端のモデルが 2 番目のタスクで 40% 以上のパフォーマンス低下を示していることを示しており、人間が行うように構文的に同一の形式を一般化して、異なる構造上の意味に到達することができないことが明らかになりました。私たちは、新しいデータセットと、プロンプトやモデル応答を含む関連する実験データを公開しています。

原文 (English)

Beyond Memorization: Assessing Semantic Generalization in Large Language Models Using Phrasal Constructions

The web-scale of pretraining data has created an important evaluation challenge: to disentangle linguistic competence on cases well-represented in pretraining data from generalization to out-of-domain language, specifically the dynamic, real-world instances less common in pretraining data. To this end, we construct a diagnostic evaluation to systematically assess natural language understanding in LLMs by leveraging Construction Grammar (CxG). CxG provides a psycholinguistically grounded framework for testing generalization, as it explicitly links syntactic forms to abstract, non-lexical meanings. Our novel inference evaluation dataset consists of English phrasal constructions, for which speakers are known to be able to abstract over commonplace instantiations in order to understand and produce creative instantiations. Our evaluation dataset uses CxG to evaluate two central questions: first, if models can 'understand' the semantics of sentences for instances that are likely to appear in pretraining data less often, but are intuitive and easy for people to understand. Second, if LLMs can deploy the appropriate constructional semantics given constructions that are syntactically identical but with divergent meanings. Our results demonstrate that state-of-the-art models, including GPT-o1, exhibit a performance drop of over 40% on our second task, revealing a failure to generalize over syntactically identical forms to arrive at distinct constructional meanings in the way humans do. We make our novel dataset and associated experimental data, including prompts and model responses, publicly available.

2026-06-01 13:00 JSTarXiv cs.AIビジネス/資金調達

PASTA: マルチポリシー AI コンプライアンス評価のためのスケーラブルなフレームワーク

AI システムがより強力になり普及するにつれて、AI コンプライアンスはますます重要になっています。しかし、AI 政策の急速な拡大は、政策の専門知識を持たず、リソースに制約のある実務者にとっては大きな負担となっています。既存のアプローチは通常、一度に 1 つのポリシーに対処するため、複数のポリシーのコンプライアンスにコストがかかります。私たちは、4 つのイノベーションを統合したスケーラブルなコンプライアンス ツールである PASTA を紹介します。(1) 開発段階全体にわたる記述入力をサポートする包括的なモデルカード形式。 (2) ポリシー正規化スキーム。 (3) コスト削減戦略を備えた効率的な LLM を利用したペアワイズ評価エンジン。 (4) コンプライアンス ヒートマップと実用的な推奨事項を介して解釈可能な評価を提供するインターフェイス。専門家の評価では、PASTA の判断が人間の専門家とほぼ一致していることが示されています ($\rho \geq .626$)。このシステムは 5 つの主要なポリシーを 2 分以内に約 $3 で評価します。ユーザー調査 (N = 12) では、実践者がアウトプットが理解しやすく実用的であり、スケーラブルな自動 AI ガバナンスのための新しいフレームワークを導入していると感じていることが確認されています。

原文 (English)

PASTA: A Scalable Framework for Multi-Policy AI Compliance Evaluation

AI compliance is becoming increasingly critical as AI systems grow more powerful and pervasive. Yet the rapid expansion of AI policies creates substantial burdens for resource-constrained practitioners lacking policy expertise. Existing approaches typically address one policy at a time, making multi-policy compliance costly. We present PASTA, a scalable compliance tool integrating four innovations: (1) a comprehensive model-card format supporting descriptive inputs across development stages; (2) a policy normalization scheme; (3) an efficient LLM-powered pairwise evaluation engine with cost-saving strategies; and (4) an interface delivering interpretable evaluations via compliance heatmaps and actionable recommendations. Expert evaluation shows PASTA's judgments closely align with human experts ($\rho \geq .626$). The system evaluates five major policies in under two minutes at approximately \$3. A user study (N = 12) confirms practitioners found outputs easy-to-understand and actionable, introducing a novel framework for scalable automated AI governance.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達規制/政策

Gap-K%: 事前トレーニング データを検出するための上位 1 予測ギャップの測定

大規模言語モデル (LLM) における大規模な事前トレーニング コーパスの不透明さにより、プライバシーと著作権に関する重大な懸念が生じ、事前トレーニング データの検出が重大な課題となっています。既存の最先端の手法は通常、トークンの尤度に依存していますが、ターゲット トークンとモデルの上位 1 予測との間のギャップや、隣接するトークン間の局所的な相関を見落とすことがよくあります。この研究では、LLM 事前トレーニングの最適化ダイナミクスに基づいた新しい事前トレーニング データ検出方法である Gap-K% を提案します。次のトークンの予測目標を分析することにより、モデルのトップ 1 予測とターゲット トークン間の不一致が強い勾配信号を誘発し、トレーニング中に明示的にペナルティが課されることが観察されます。これを動機として、Gap-K% は、上位 1 位の予測トークンとターゲット トークン間の対数確率ギャップを活用し、スライディング ウィンドウ戦略を組み込んで局所的な相関関係を捕捉し、トークン レベルの変動を軽減します。 WikiMIA および MIMIR ベンチマークに関する広範な実験により、Gap-K% が最先端のパフォーマンスを達成し、さまざまなモデル サイズや入力長にわたって一貫して以前のベースラインを上回るパフォーマンスを示していることが実証されています。

原文 (English)

Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data

The opacity of massive pretraining corpora in Large Language Models (LLMs) raises significant privacy and copyright concerns, making pretraining data detection a critical challenge. Existing state-of-the-art methods typically rely on token likelihoods, yet they often overlook the gap between the target token and the model's top-1 prediction, as well as local correlations between adjacent tokens. In this work, we propose Gap-K%, a novel pretraining data detection method grounded in the optimization dynamics of LLM pretraining. By analyzing the next-token prediction objective, we observe that discrepancies between the model's top-1 prediction and the target token induce strong gradient signals, which are explicitly penalized during training. Motivated by this, Gap-K% leverages the log probability gap between the top-1 predicted token and the target token, incorporating a sliding window strategy to capture local correlations and mitigate token-level fluctuations. Extensive experiments on the WikiMIA and MIMIR benchmarks demonstrate that Gap-K% achieves state-of-the-art performance, consistently outperforming prior baselines across various model sizes and input lengths.

2026-06-01 13:00 JSTarXiv cs.AIビジネス/資金調達

Shapley 値の奇妙な推定量

Shapley 値は、特徴の重要性、データの評価、因果推論を含む、機械学習における帰属のためのユビキタスなフレームワークです。ただし、その正確な計算は一般に困難であるため、効率的な近似方法が必要です。最も効果的で一般的な推定ツールはペア サンプリング ヒューリスティックを活用して推定誤差を削減していますが、この改善を推進する理論的メカニズムは依然として不透明です。この研究では、ペア サンプリングの洗練された基本的な正当化を提供します。つまり、シャプレー値が集合関数の奇数成分のみに依存すること、およびペア サンプリングが回帰目標を直交化して無関係な偶数成分を除外することを証明します。この洞察を活用して、奇数部分空間のみで多項式回帰を実行する新しい一貫した推定器である OddSHAP を提案します。フーリエ基底を利用してこの部分空間を分離し、プロキシ モデルを採用して影響の大きい相互作用を特定することにより、OddSHAP は高次近似の組み合わせ爆発を克服します。広範なベンチマークを通じて、OddSHAP がより大きなサンプリング予算で最先端の推定精度を達成していることがわかりました。

原文 (English)

An Odd Estimator for Shapley Values

The Shapley value is a ubiquitous framework for attribution in machine learning, encompassing feature importance, data valuation, and causal inference. However, its exact computation is generally intractable, necessitating efficient approximation methods. While the most effective and popular estimators leverage the paired sampling heuristic to reduce estimation error, the theoretical mechanism driving this improvement has remained opaque. In this work, we provide an elegant and fundamental justification for paired sampling: we prove that the Shapley value depends exclusively on the odd component of the set function, and that paired sampling orthogonalizes the regression objective to filter out the irrelevant even component. Leveraging this insight, we propose OddSHAP, a novel consistent estimator that performs polynomial regression solely on the odd subspace. By utilizing the Fourier basis to isolate this subspace and employing a proxy model to identify high-impact interactions, OddSHAP overcomes the combinatorial explosion of higher-order approximations. Through an extensive benchmark, we find that OddSHAP achieves state-of-the-art estimation accuracy at larger sampling budgets.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

言語モデルエージェントにおける目標指向性の行動的および表現的評価

エージェントの目標を理解することは、エージェントの動作を説明し、予測するのに役立ちますが、目標をエージェント システムに確実に帰すための確立された方法論はありません。我々は、行動評価とモデルの内部表現の解釈可能性に基づく分析を統合した、目標指向性を評価するためのフレームワークを提案します。ケーススタディとして、2D グリッド世界を目標状態に向かってナビゲートする LLM エージェントを調べます。行動面では、さまざまなグリッド サイズ、障害物密度、目標構造にわたる最適なポリシーに照らしてエージェントを評価し、パフォーマンスがタスクの難易度に応じてスケールすると同時に、難易度を維持した変換や複数の目標構造に対して堅牢であることがわかりました。次に、調査手法を使用して、環境の内部表現と複数ステップのアクション プランを解読します。 LLM エージェントは粗い空間マップを非線形にエンコードし、その位置と目標位置に関するおおよそのタスク関連の手掛かりを保存していることがわかります。その動作はこれらの内部表現とほぼ一致していること。そしてその推論はそれらを再編成し、空間的な手がかりから即時の行動の選択へと移行します。私たちの調査結果は、エージェントがどのように目標を表現し追求するかを特徴付けるには、行動評価を超えて内省的な検査が必要であるという見解を裏付けています。

原文 (English)

A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents

Understanding an agent's goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals to agentic systems. We propose a framework for evaluating goal-directedness that integrates behavioural evaluation with interpretability-based analyses of models' internal representations. As a case study, we examine an LLM agent navigating a 2D grid world towards a goal state. Behaviourally, we evaluate the agent against optimal policies across varying grid sizes, obstacle densities, and goal structures, finding that performance scales with task difficulty while remaining robust to difficulty-preserving transformations and multi-goal structures. We then use probing methods to decode internal representations of the environment and multi-step action plans. We find that the LLM agent non-linearly encodes a coarse spatial map, preserving approximate task-relevant cues about its position and the goal location; that its actions are broadly consistent with these internal representations; and that reasoning reorganises them, shifting from spatial cues towards immediate action selection. Our findings support the view that introspective examination is required beyond behavioural evaluations to characterise how agents represent and pursue their objectives.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

範囲: 選択的コンフォーマル最適化ペアワイズ LLM 判定

大規模言語モデル (LLM) は、ペアごとの評価におけるスケーラブルな判断材料としてますます使用されていますが、依然として誤った調整やバイアスが発生しやすい傾向があります。我々は、交換可能性の下で、非棄権判断間の誤り率が最大でもユーザー指定のレベル $\alpha$ になるように許容閾値を調整するフレームワークである SCOPE (選択的共形最適化ペアワイズ評価) を提案します。 SCOPE にバイアス中立の不確実性信号を供給するために、双方向優先エントロピー (BPE) を導入します。これは、両方の応答位置で裁判官にクエリを実行し、順序平均された優先確率をエントロピー ベースのスコアに変換します。さまざまなペアごとの判定ベンチマーク全体で、BPE は校正と識別において標準信頼度代用を上回っていますが、SCOPE は目標リスク限界 ($\alpha = 0.10$ での経験的な FDR $\約 0.097$ から $0.099$) を一貫して満たしており、実質的なカバレッジを維持しています。バニラのベースラインと比較して、SCOPE は同じリスク制約の下で最大 2.4 倍 $ 多くの判断を受け入れます。これは、BPE が信頼性が高くカバレッジの高い LLM ベースの評価を可能にすることを示しています。

原文 (English)

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and biases. We propose SCOPE (Selective Conformal Optimized Pairwise Evaluation), a framework that calibrates an acceptance threshold so that, under exchangeability, the error rate among non-abstained judgments is at most a user-specified level $\alpha$. To supply SCOPE with a bias-neutral uncertainty signal, we introduce Bidirectional Preference Entropy (BPE), which queries the judge under both response positions and converts the order-averaged preference probability into an entropy-based score. Across various pairwise judging benchmarks, BPE outperforms standard confidence proxies in calibration and discrimination, while SCOPE consistently satisfies the target risk bound (empirical FDR $\approx 0.097$ to $0.099$ at $\alpha = 0.10$) and retains substantial coverage. Compared to vanilla baselines, SCOPE accepts up to $2.4\times$ more judgments under the same risk constraint, demonstrating that BPE enables reliable and high-coverage LLM-based evaluation.

2026-06-01 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

立場: ECG 表現の評価は修正される必要がある

この意見書は、12誘導ECG表現学習における現在のベンチマーク慣行を修正して、進歩が信頼でき、臨床的に意味のある目的に沿っていることを保証する必要があると主張している。 ECG は実質的により広範な臨床情報をコード化することが知られているにもかかわらず、この分野は主に、不整脈と波形形態ラベルが大半を占める 3 つの公開マルチラベル ベンチマーク (PTB-XL、CPSC2018、CSN) に収束しています。私たちは、関連する臨床目標として、他の進化する ECG 関連のエンドポイントに加えて、構造的心疾患の評価と患者レベルの予測を含めるように下流の評価を拡大する必要があると主張します。次に、マルチラベルの不均衡な設定に対する評価のベスト プラクティスの概要を示し、それらが適用されると、どの表現が最もパフォーマンスを発揮するかについての文献の現在の結論が変更されることを示します。さらに、線形評価を伴うランダムに初期化されたエンコーダーが、多くのタスクにおける最先端の事前トレーニングと一致するという驚くべき結果を実証します。これにより、合理的なベースライン モデルとしてランダム エンコーダーを使用するようになります。私たちは、6 つの評価設定 (3 つの標準ベンチマーク、構造的疾患データセット、血行力学的推論、および患者予測) にわたる 5 つの代表的な ECG 事前トレーニング アプローチの経験的評価によって観察を実証します。

原文 (English)

Position: Evaluation of ECG Representations Must Be Fixed

This position paper argues that current benchmarking practice in 12-lead ECG representation learning must be fixed to ensure progress is reliable and aligned with clinically meaningful objectives. The field has largely converged on three public multi-label benchmarks (PTB-XL, CPSC2018, CSN) dominated by arrhythmia and waveform-morphology labels, even though the ECG is known to encode substantially broader clinical information. We argue that downstream evaluation should expand to include an assessment of structural heart disease and patient-level forecasting, in addition to other evolving ECG-related endpoints, as relevant clinical targets. Next, we outline evaluation best practices for multi-label, imbalanced settings, and show that when they are applied, the literature's current conclusion about which representations perform best is altered. Furthermore, we demonstrate the surprising result that a randomly initialized encoder with linear evaluation matches state-of-the-art pre-training on many tasks. This motivates the use of a random encoder as a reasonable baseline model. We substantiate our observations with an empirical evaluation of five representative ECG pre-training approaches across six evaluation settings: the three standard benchmarks, a structural disease dataset, hemodynamic inference, and patient forecasting.

2026-06-01 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

World Action Verifier: 順逆非対称性による自己改善型世界モデル

汎用世界モデルは、スケーラブルな政策の評価、最適化、計画を約束しますが、必要なレベルの堅牢性を達成することは依然として困難です。最適なアクションに主に焦点を当てたポリシー学習とは異なり、ワールド モデルは、アクションにラベル付けされたロボット インタラクションでは過小評価されることが多い、準最適なアクションの広大な空間にわたって信頼できる必要があります。この課題に対処するために、ワールド モデルが独自の予測エラーを特定して自己改善できるようにするフレームワークである World Action Verifier (WAV) を提案します。重要なアイデアは、アクション条件付き状態予測を、状態の妥当性とアクションの到達可能性という 2 つの独立して検証可能な要素に分解することです。我々は、アクションのないデータがより広範囲に利用可能であることと、アクション関連の特徴がより低次元であるという 2 つの根本的な非対称性により、これらの要因の検証が直接予測よりもはるかに扱いやすいことを示します。これらの非対称性を利用して、(i) ビデオ コーパスから取得した多様なサブゴール ジェネレーターと、(ii) 状態特徴のサブセットからアクションを推測するスパース逆モデルで世界モデルを拡張します。 WAV は、提案されたサブ目標、推測されたアクション、および今後の展開の間でサイクルの一貫性を強化することにより、既存の手法が失敗することが多い、未調査の領域において効果的な検証メカニズムを提供します。 MiniGrid、RoboMimic、ManiSkill にわたる 9 つのタスクにわたって、私たちのメソッドは 2 倍のサンプル効率を達成しながら、下流のポリシーのパフォーマンスを 22% 以上向上させます。

原文 (English)

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness remains challenging. Unlike policy learning which primarily focuses on optimal actions, a world model needs to be reliable over a vast space of suboptimal actions, which are often underrepresented in action-labeled robot interactions. To address this challenge, we propose World Action Verifier (WAV), a framework that enables world models to identify their own prediction errors and self-improve. The key idea is to decompose action-conditioned state prediction into two independently verifiable factors: state plausibility and action reachability. We show that verifying these factors is significantly more tractable than direct forward prediction due to two underlying asymmetries: the broader availability of action-free data and the lower dimensionality of action-relevant features. Leveraging these asymmetries, we augment a world model with (i) a diverse subgoal generator obtained from video corpora and (ii) a sparse inverse model that infers actions from a subset of state features. By enforcing cycle consistency among proposed subgoals, inferred actions, and forward rollouts, WAV provides an effective verification mechanism in under-explored regimes, where existing methods often fail. Across nine tasks spanning MiniGrid, RoboMimic, and ManiSkill, our method achieves 2x higher sample efficiency while improving downstream policy performance by over 22%.

2026-06-01 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

Pocket-Dentist: 効率的なマルチモーダル大規模言語モデルによるオンデバイス歯科画像理解

歯科視覚言語モデルの評価は、データセット、タスク定義、メトリクスにわたって断片化されたままであり、多くの場合、その計算コストが無視されます。このため、専門センター外での歯科スクリーニングへの広範な導入が制限されています。専門センターでは、タイムリーな推論、限られたハードウェア、および患者画像のローカル処理が、実用的でプライバシーを保護した臨床前スクリーニングに不可欠です。ここで紹介する Pocket-Dentist は、約 1,159 人の患者、5 つのタスク タイプ、7 つの指標にまたがる 3 つのデータセットをまとめた、歯科マルチモーダル質問応答のための効率を意識したベンチマークです。典型的な 14 個の VLM にわたって、我々の結果は興味深い観察結果を明らかにしました。コンパクトな VLM (例: 2B パラメータ モデル) は、精度においては大型の VLM を上回っていますが、歯科画像の理解に必要な計算コストは​​大幅に低くなります。 iPhone 17 Pro にローカルに導入された当社の微調整されたコンパクト VLM Pocket-Dentist-2B は、各サンプルを 4.31 秒で処理し、7B ベースラインと比較してレイテンシーを 4.9 倍、メモリ使用量を 2.3 倍削減しました。

原文 (English)

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widespread deployment for dental screening outside specialist centres, where timely inference, limited hardware, and local handling of patient images are vital for practical, privacy-preserving clinical prescreening. Here we present Pocket-Dentist, an efficiency-aware benchmark for dental multimodal question answering that brings together three datasets spanning approximately 1,159 patients, five task types and seven metrics. Across typical 14 VLMs, our results reveals an interesting observation: compact VLMs (e.g., 2B-parameter models) outperform larger VLMs in accuracy while requiring substantially lower computational costs in dental image understanding. Deployed locally on an iPhone 17 Pro, our finetuned compact VLM Pocket-Dentist-2B processed each sample in 4.31 s, reducing latency by 4.9-fold and memory use by 2.3-fold compared with a 7B baseline.

2026-05-30 02:27 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

After Nvidia’s $20B not-acqui-hire, AI chip startup Groq reportedly raising $650M

Chipmaker Groq is looking to raise $650 million in internal funding as it pivots from hardware to focus more on AI inference, the process o…

2026-05-29 21:00 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

This chip startup just raised $135M on a bet that AI’s biggest bottleneck isn’t compute — it’s memory

South Korean chip startup XCENA is betting that AI's real bottleneck is not compute, but memory.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

モデルが一致しない場合: パブリック コメント分析のための LLM 評価を再考する

連邦政府機関はパブリック コメント コーパスを分類するために大規模言語モデル (LLM) を導入しており、モデルの記録構成によって政策立案者が何を確認し、どの議論が登録されるかが決まります。小規模な検証済みセットに対するスタンスの精度に基づいた標準評価では、異なるモデルが同じ公的入力に対して実質的に異なる分類を生成する場合を検出できません。私たちは、マルチモデルの不一致を解釈の複雑さの診断として扱い、真に曖昧な公的意見に向けて人間によるレビューを指示する解釈監査パイプラインを提案します。 4 つの LLM にわたる連邦 USDA 文書に対する 1,260 件のパブリック コメントを分析したところ、モデル間のテーマの相違がモデル内のプロンプト変動を上回っており、専門家のルーブリックが深い解釈上の不一致を解決することなく抑圧していることがわかりました。層化された 40 コメントのサブサンプルに対する 2 段階のラベル付け研究では、4 人の LLM とヒューマン アノテーターが独立してラベル付けし、他のラベルを確認した後に修正しました。改訂動作はラベラーによって異なり、ヒューマン・アノテーターの改訂では、アンサンブルの集合的な出力にはないフレームが頻繁に導入されました。私たちは、不一致に基づく評価は、LLM 支援解釈コーディングの精度メトリクスを補完するために必要であると主張します。

原文 (English)

When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis

Federal agencies are deploying large language models (LLMs) to categorize public comment corpora, where the model's organization of the record shapes what policymakers see and which arguments register. Standard evaluation, anchored on stance accuracy against a small validated set, cannot detect when different models produce materially different categorizations of the same public input. We propose an Interpretive Audit Pipeline that treats multi-model disagreement as diagnostic of interpretive complexity and directs human review toward genuinely ambiguous public input. Analyzing 1,260 public comments on a federal USDA docket across four LLMs, we find that inter-model thematic divergence exceeds within-model prompt variation, and that an expert rubric suppresses deep interpretive disagreement without resolving it. In a two-stage labeling study on a stratified 40-comment subsample, four LLMs and a human annotator labeled independently and then revised after seeing the others' labels. Revision behavior varied across labelers, and the human annotator's revisions frequently introduced framings absent from the ensemble's collective output. We argue disagreement-based evaluation is a necessary complement to accuracy metrics for LLM-assisted interpretive coding.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

ペーパーエージェント、ペーパーゲイン:DeFi投資エージェントの実証分析

自律的なオンチェーン取引に AI を使用するシステムである DeFi 投資エージェントは、2024 年後半以来、合計トークン評価額で 30 億米ドルを超えています。私たちは 1,900 以上の AI タグ付き暗号プロジェクトを調査し、投資中心のエージェントに絞り込み、戦略と可観測性の側面にわたる 10 の代表的なプロジェクトを厳選しています。次に、ElizaOS と Virtuals Protocol という 2 つの著名なエージェント フレームワークの詳細なアーキテクチャ分析と、925,323 人のトークン所有者を対象とする公的に起因する取引活動を伴う 11 の Solana ベースのエージェント トレジャリーの定量的なオンチェーン パフォーマンス分析を実施します。現在のデプロイメントは初期段階で異種混合のままであることがわかりました。(1) 私たちのサンプルでは、​​多くのプロジェクトが自律的な取引実行の明確な証拠をまだ提供しておらず、開発者のインタビューでは、目に見えるデプロイメントの多くが基本的な API 統合のままであることが示唆されています。 (2) エージェントの財務省は 3,000 万米ドルを超える紙の利益を保持している一方、トークン所有者は合計で 1 億 9,170 万米ドルを損失しており、ウォレットの上位 1% が全利益の 81.4% (18 億 1,000 万米ドル) を獲得しています。 (3) トークンの評価額は財務省のファンダメンタルズとの関連が弱く、時価総額対AUMの比率は10,000倍を超えていますが、確立されたDeFiプロトコルでは1倍未満です。 (4) ユーザーの総利益は 24 億米ドルでピークに達し、その後純損失に減少し、収益の中央値はすべてのプラットフォームでマイナスとなり、トークンは史上最高値から平均して 93% 減少しました。私たちは、これらの結果を、オープンインフラストラクチャにより迅速な実験が可能になるだけでなく、自律性、パフォーマンス、および利害関係者の連携のための堅牢な標準が出現する前に、単純なエージェントや投機的なエージェントが立ち上がることを可能にする、パーミッションレスの第一世代市場の特徴であると解釈します。そこで私たちは、現在の展開と将来の投資グレードのエージェント システムとの間のギャップを特徴付けるために、自律的な実行、リスク調整後の収益性、利害関係者の連携という 3 つの側面に沿った成熟度フレームワークを提案します。

原文 (English)

Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents

DeFi investment agents, systems that use AI for autonomous on-chain trading, have attained over USD 3 billion in combined token valuations since late 2024. We survey over 1,900 AI-tagged crypto projects, filter to investment-focused agents, and curate 10 representative projects spanning strategy and observability dimensions. We then conduct a deep-dive architectural analysis of two prominent agent frameworks, ElizaOS and Virtuals Protocol, and a quantitative on-chain performance analysis of 11 Solana-based agent treasuries with publicly attributable trading activity, covering 925,323 token holders. We find that current deployments remain early and heterogeneous: (1) in our sample, many projects do not yet provide clear evidence of autonomous trade execution, and developer interviews suggest that many visible deployments remain basic API integrations; (2) agent treasuries retain over USD 30M in paper gains while token holders collectively lost USD 191.7M, with the top 1% of wallets capturing 81.4% of all gains (USD 1.81B); (3) token valuations are weakly connected to treasury fundamentals, with market-cap-to-AUM ratios exceeding 10,000x versus below 1x for established DeFi protocols; and (4) aggregate user gains peaked at USD 2.4B before declining to net losses, with median returns negative on every platform and tokens declining 93% on average from all-time highs. We interpret these outcomes as characteristic of a permissionless, first-generation market in which open infrastructure enables rapid experimentation but also allows naive or speculative agents to launch before robust standards for autonomy, performance, and stakeholder alignment emerge. We therefore propose a maturity framework along three dimensions: autonomous execution, risk-adjusted profitability, and stakeholder alignment, to characterize the gap between current deployments and future investment-grade agent systems.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

BenchTrace: LLM エージェントのリフレクション能力と制御された進化をテストするためのベンチマーク

自己進化エージェントは過去の失敗を反映することで時間の経過とともに改善しますが、既存の評価には 2 つの点で制限があります。1 つはタスク スコアのみを測定し、反映品質は不明のままにすること、もう 1 つはエージェント自身のエピソードの実行に依存しており、特定の失敗パターンを対象にするメカニズムを提供していないことです。 LLM エージェントの自己進化能力を評価するためのベンチマークである \textbf{BenchTrace} を紹介します。 BenchTrace は、6 つの多様なタスクにわたる 1,821 の注釈付きエピソードのスナップショット反映データセットに基づいて構築されており、ターゲットを絞った QA タスクを通じて障害の特定を調査する \textbf{反映評価} と、制御された自己進化シミュレーションで過去の障害経験が回避行動に変換されるかどうかをテストする \textbf{進化評価} で構成されます。 BenchTrace に基づいて、エージェントがターゲットの障害インスタンスを回避できたテスト ケースの割合を測定する新しい評価指標である \textbf{障害回避率 (FAR)} を提案します。 Qwen3-32B と GPT-4.1 を使った実験では、どちらのモデルもリフレクション評価でエンドツーエンドの合格率が 30\% を下回り、診断が主なボトルネックであることが明らかになりました。進化の評価では、自己進化手法は一般に非進化ベースラインよりもFARを改善しますが、エージェントはノイズエピソードが蓄積するにつれて初期のレッスンを忘れ、エージェントは特定のコンテキストを超えて反省を一般化することができず、タスクコンテキスト間で負の転移を引き起こすことが示されています。さらに、相関分析により、完全に正しい反射のみが高い FAR と強く関連していることが明らかになりました。 BenchTrace は、現在の自己進化アプローチの具体的な限界を明らかにし、対象を絞った評価のための制御されたモデルに依存しないフレームワークを提供します。

原文 (English)

BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

Self-evolving agents improve over time by reflecting on past failures, but existing evaluation is limited in two ways: it measures only task scores, leaving reflection quality unknown, and it relies on agents' own episode runs, offering no mechanism to target specific failure patterns. We present \textbf{BenchTrace}, a benchmark for evaluating self-evolution ability in LLM agents. BenchTrace is built on a snapshot-reflection dataset of 1,821 annotated episodes spanning six diverse tasks, and comprises a \textbf{Reflection Evaluation} that probes failure identification through targeted QA tasks, and an \textbf{Evolution Evaluation} that tests whether past failure experience translates into avoidance behavior in a controlled self-evolution simulation. Building on BenchTrace, we propose \textbf{failure avoidance rate (FAR)}, a new evaluation metric measuring the fraction of test cases in which the agent successfully avoids the target failure instance. Experiments with Qwen3-32B and GPT-4.1 reveal that both models fall below a 30\% end-to-end pass rate on reflection evaluation, with diagnosis as the primary bottleneck. Evolution evaluation shows that self-evolution methods generally improve FAR over the non-evolving baseline, but agents forget early lessons as noise episodes accumulate, and agents fail to generalize their reflections beyond the specific context, causing negative transfer across task contexts. Our correlation analysis further reveals that only a fully correct reflection is strongly associated with higher FAR. BenchTrace exposes concrete limits of current self-evolution approaches and provides a controlled, model-agnostic framework for targeted evaluation.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

文献検索の評価を再考する: 深い調査は役に立ちますが、人間の引用リストは根拠のある真実ではありません

私たちは、検索パイプラインの改善と評価対象としての人による参照リストのストレステストという 2 つの相補的な角度から大規模な文献検索を研究しています。まず、完全なクエリ論文を処理し、取得した結果を文献目録に沿って幅優先で拡張する Deep Research パイプラインを実装します。このパイプラインが通常の API のみの検索を大幅に上回り、RollingEval-Jun25 (論文 250 件の文献検索ベンチマーク) の再現率が 20% 未満から 80% 以上に上昇することを示します。 2 番目に、中立的な LLM を判断者として使用して、人間の参照がタスクに対する健全な根拠であるかどうかを判断します。私たちは重大な限界を発見しました。人間による引用のうち、中等度以上の関連性があると判断されたのは 51% のみであったのに対し、最も強力な AI ベースの再ランカーでは 86 ~ 88% でした。 OpenAlex の共著グラフでこのギャップを調査したところ、人間は AI の再ランク付けを行う最も優れた人よりも直接の協力者を引用する可能性が 2.5 倍高いことがわかりました。まとめると、我々の結果は単一軸の文献検索評価に反対している。つまり、想起率、話題関連性スコアリング、ランクリストの多様性、および共著距離診断は、それぞれ引用の質の相補的な特性を測定するものであり、併せて報告されるべきである。

原文 (English)

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as an evaluation target. First, we implement a Deep Research pipeline that processes the full query paper and expands the retrieved results breadth-first along their bibliographies, and show that it substantially outperforms vanilla API-only search, raising recall on RollingEval-Jun25 (a 250-paper literature-search benchmark) from below 20% to above 80%. Second, we use a neutral LLM-as-a-judge to determine if human references are sound ground truth for the task. We find significant limitations: only 51% of human citations are judged moderately relevant or higher, against 86--88% for the strongest AI-based re-rankers. We study this gap on the OpenAlex co-authorship graph, finding that humans are 2.5x more likely than the best AI re-rankers to cite a direct collaborator. Together, our results argue against single-axis literature-search evaluation: recall, topical-relevance scoring, ranked-list diversity, and a co-authorship-distance diagnostic each measure complementary properties of citation quality and should be reported jointly.

2026-05-29 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

MiraBench: ロボット世界モデルにおける動作条件付き信頼性の評価

アクション条件付き世界モデルは、ロボット学習用のスケーラブルなシミュレーターとしてますます使用されていますが、現在の評価では、条件付けされたアクションの下でその予測が信頼できるという限られた証拠が提供されています。既存のベンチマークは主に視覚的な忠実度を重視しており、予測される未来が物理的に妥当であるか、命令されたアクションに忠実であるか、アクションが成功しないはずのときに失敗するように調整されているかどうかが不明確なままです。 \emph{動作条件付き信頼性} をロボット世界モデルの中核的な評価目標として定義する階層型ベンチマークである \textsc{MiraBench} を紹介します。 MiraBench は、こ​​のターゲットを 3 つの段階的に要求の高いレベルに分解します。 \emph{Physics Adherence} は、リファレンスフリーの物理的一貫性を評価します。 \emph{Action-Following Fidelity}: 予測がタスク関連のアクション入力を考慮しているかどうかを測定します。 \emph{楽観主義バイアス検出} は、失敗を誘発する行動の下で成功した結果を予測する傾向を調査します。この評価をサポートするために、タスク、失敗カテゴリ、主要な世界モデルにわたる 16,000 件を超える判断を含む人間による注釈付きコーパスを厳選しました。ベクトル条件付きロボット ワールド モデル、テキスト条件付き生成ワールド モデル、オープンウェイト システム、クローズド ソース システム、および複数のモデル スケールにわたる 12 の代表的なモデル構成を評価します。この広範なモデル環境全体にわたって、MiraBench は 3 つの中心的な発見を明らかにしました。視覚的な忠実度は、アクションの忠実度の代用としては不十分です。モデルのスケールを大きくしても、アクションの追従性が確実に改善されるわけではありません。そして楽観主義バイアスは現在のシステム全体に蔓延しています。 MiraBench は、評価を外観から動作条件付きの信頼性に移行することで、ロボットの世界モデルを忠実なシミュレーターとして評価および改善するための診断基盤を提供します。

原文 (English)

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks largely emphasize visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to commanded actions, and calibrated to failure when actions should not succeed. We introduce \textsc{MiraBench}, a hierarchical benchmark that defines \emph{action-conditioned reliability} as a core evaluation target for robotic world models. MiraBench decomposes this target into three progressively demanding levels: \emph{Physics Adherence}, which evaluates reference-free physical consistency; \emph{Action-Following Fidelity}, which measures whether predictions respect task-relevant action inputs; and \emph{Optimism Bias Detection}, which probes the tendency to predict successful outcomes under failure-inducing actions. To support this evaluation, we curate a human-annotated corpus with over 16,000 judgments across tasks, failure categories, and leading world models. We evaluate 12 representative model configurations spanning vector-conditioned robotic world models, text-conditioned generative world models, open-weight systems, closed-source systems, and multiple model scales. Across this broad model landscape, MiraBench reveals three central findings: visual fidelity is a poor proxy for action fidelity; increasing model scale does not reliably improve action following; and optimism bias is pervasive across current systems. By shifting evaluation from appearance to action-conditioned reliability, MiraBench provides a diagnostic foundation for assessing and improving robotic world models as faithful simulators.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

エージェントによる修正と意味評価による人間のような対話型音声認識を目指して

自動音声認識 (ASR) は、人間とコンピューターの対話の中核コンポーネントであり、LLM ベースのアシスタントおよびエージェントにとってますます重要なフロントエンドです。しかし、現在のほとんどの ASR システムは依然としてシングルパス パラダイムに従っており、人間のコミュニケーションとの整合性が低く、誤解は繰り返しの明確化と改良によって解決されます。この不一致により、意味に関わる重大なエラーが発生すると、修正することが困難になります。一方、WER や CER などのトークンレベルの指標は、このような問題を適切に反映できません。これらの制限に対処するために、\emph{Interactive ASR} をマルチターン改良タスクとして定式化し、シングルパス ASR フロントエンドとセマンティック修正、インテント ルーティング、推論ベースの編集を組み合わせた閉ループ フレームワークである \textbf{Agentic ASR} を提案します。さらに、LLM ベースのセマンティック評価指標である \textbf{文レベルのセマンティック エラー率} ($S^2ER$) を、スケーラブルで再現可能なベンチマークのための \textbf{インタラクティブ シミュレーション システム} とともに導入します。多言語、名前付きエンティティ集中型、およびコードスイッチングのベンチマークに関する実験では、反復的な対話によりセマンティック エラーが一貫して減少し、従来のトークン レベルのメトリクスよりも $S^2ER$ が大幅に増加することが示されました。人間と AI のアライメントとアブレーションの研究により、意味判断の信頼性と提案されたフレームワークの堅牢性がさらに検証されました。コードは https://interactiveasr.github.io/ で入手でき、ライブ デモは https://i-asr.sjtuxlance.com/ で入手できます。

原文 (English)

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation

Automatic speech recognition (ASR) is a core component of human--computer interaction and an increasingly important front-end for LLM-based assistants and agents. However, most current ASR systems still follow a single-pass paradigm, which is poorly aligned with human communication, where misunderstandings are resolved through iterative clarification and refinement. This mismatch makes it difficult to correct meaning-critical errors once they occur. Meanwhile, token-level metrics such as WER or CER cannot adequately reflect such a problem. To address these limitations, we formulate \emph{Interactive ASR} as a multi-turn refinement task and propose \textbf{Agentic ASR}, a closed-loop framework that combines a single-pass ASR front-end with semantic correction, intent routing, and reasoning-based editing. We further introduce the \textbf{Sentence-level Semantic Error Rate} ($S^2ER$), an LLM-based semantic evaluation metric, together with an \textbf{Interactive Simulation System} for scalable and reproducible benchmarking. Experiments on multilingual, named-entity-intensive, and code-switching benchmarks show that iterative interaction consistently reduces semantic errors, with much larger gains in $S^2ER$ than in conventional token-level metrics. Human--AI alignment and ablation studies further validate the reliability of the semantic judge and the robustness of the proposed framework. The code is available at: https://interactiveasr.github.io/ and the live demo is available at https://i-asr.sjtuxlance.com/

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

TRACE: LLM CoT 評価の構成要素によるトゥールミンベースの推論評価

大規模言語モデル (LLM) からのオープンエンドの出力を評価することは、グランド トゥルースがないため依然として困難です。既存の指標は、最終的な答えの精度や表面レベルの統計に依存しており、推論プロセス自体は検討されていません。思考連鎖 (CoT) 推論プロセスを分析する指標である TRACE (Toulmin-based Reasoning Assessment through Constructive Elements) を紹介します。 TRACE は、結果を判断するのではなく、トゥールミンの議論理論とフラベルのメタ認知フレームワークを統合して推論の構造を評価することにより、議論がどのように構築されるかを検査します。 7 つの推論モデルにわたる 26.3K の QA サンプルの実験では、ベンチマーク精度 (r=0.74) との強い相関関係が示されています。さらに、TRACE は強化学習の報酬信号として効果的であり、精度のみのベースラインを上回ります。これらの結果を総合すると、論理的に健全な推論がより質の高い答えにつながることを示しています。したがって、TRACE は、オープンエンド出力を評価するための補足的なメトリックとして機能します。コードは https://github.com/hyyangkisti/trace で入手できます。

原文 (English)

TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-answer accuracy or surface-level statistics, leaving the reasoning process itself unexamined. We introduce TRACE (Toulmin-based Reasoning Assessment through Constructive Elements), a metric that analyzes Chain-of-Thought (CoT) reasoning processes. Rather than judging outcomes, TRACE inspects how arguments are constructed by integrating Toulmin's argumentation theory with Flavell's metacognitive framework to assess reasoning structure. Experiments on 26.3K QA samples across 7 reasoning models show strong correlation with benchmark accuracy (r=0.74). Furthermore, TRACE is effective as a reinforcement learning reward signal, outperforming accuracy-only baselines. Together, these results indicate that logically sound reasoning leads to higher-quality answers. TRACE thus serves as a complementary metric for evaluating open-ended outputs. Code is available at https://github.com/hyyangkisti/trace.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

スペシャリスト モデルが依然として重要な理由: 医療用人工知能のための異種マルチエージェント パラダイム

医療分野における GPT や Claude などの汎用大規模言語モデル (LLM) の優れたパフォーマンスは、領域固有の医療専門家モデルは時代遅れになるのだろうかという重大な疑問を引き起こしています。私たちは、医療用人工知能 (AI) の将来は、モノリシックな医療基盤モデルの構築や人間の専門知識の置き換えにあるのではなく、ジェネラリストの LLM、領域固有の専門家モデル、および臨床医の間のコラボレーションを調整することにあると主張します。我々は、矛盾を認識した証拠の融合、不確実性に基づく臨床医の介入トリガー、および適応閾値キャリブレーションを可能にする異種医療マルチエージェントフレームワークである HetMedAgent を提案します。 3 つの実際の臨床意思決定タスクに関する実験では、ジェネラリスト LLM と領域固有の専門家モデルの間の相乗効果が、どちらかのタイプのモデルを単独で使用した場合よりも大幅に優れていることが実証され、モダリティ固有の分析における専門家モデルのかけがえのない価値が検証されました。 HetMedAgent は、医療 LLM または基盤モデルの構築から複数エージェントのコラボレーションへの移行を表し、一般的な推論機能とドメイン固有の精度のバランスを実現します。

原文 (English)

Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence

The impressive performance of generalist large language models (LLMs) such as GPT and Claude in healthcare raises a critical question: will domain-specific medical specialist models become obsolete? We argue that the future of medical artificial intelligence (AI) lies not in building monolithic medical foundation models, nor in replacing human expertise, but in orchestrating collaboration among generalist LLMs, domain-specific specialist models, and clinicians. We propose HetMedAgent, a heterogeneous medical multi-agent framework that enables conflict-aware evidence fusion, uncertainty-based clinician intervention triggering, and adaptive threshold calibration. Experiments on three real-world clinical decision-making tasks demonstrate that the synergy between generalist LLMs and domain-specific specialist models significantly outperforms using either type of model alone, validating the irreplaceable value of specialist models in modality-specific analysis. HetMedAgent represents a shift from building medical LLMs or foundation models to multi-agent collaboration, achieving a balance between general reasoning capabilities and domain-specific precision.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

クロワッサン タスク: 再現可能な機械学習評価のためのメタデータ形式

再現性は科学的手法の基本ですが、機械学習においては依然として重要な課題です。原因としては、実行詳細の指定不足や脆弱なソフトウェア環境などが挙げられます。チェックリストや手動検証などの人間中心の救済策は役立ちますが、集中的な努力が必要であり、拡張することができません。これに対処するために、Croissant Tasks を導入します。これは、低レベルの実装の詳細を高レベルの仕様に抽象化する、宣言的でマシンアクション可能なメタデータ形式です。この形式により、概念的な再現性が可能になります。つまり、脆弱なソース コードの複製ではなく、独立したエージェント生成の実装を通じて主張を検証できます。私たちは以下に貢献しています。(1) Croissant Tasks 仕様。タスクの問題を解決策から正式に切り離します。 (2) 既存のベンチマークをこの形式に改良する自動 LLM パイプライン。 (3) 自律エージェントがこれらの仕様を取り込んで、機能的で正確な再現パイプラインを最初から生成できることを示す経験的検証。私たちはこの形式を、機械学習における自動化された概念的な再現性のための新しい基盤として構想しています。

原文 (English)

Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations

Reproducibility is fundamental to the scientific method, yet remains a critical challenge in machine learning. Contributing factors include underspecified execution details and brittle software environments. Human-centric remedies, such as checklists and manual verification, help but require intensive effort and fail to scale. To address this, we introduce Croissant Tasks: a declarative, machine-actionable metadata format that abstracts low-level implementation details into high-level specifications. This format enables conceptual reproducibility: verifying claims via independent, agent-generated implementations rather than brittle source code replication. We contribute: (1) the Croissant Tasks specification, formally decoupling task problem from solution; (2) an automated LLM pipeline that retrofits existing benchmarks into this format; and (3) empirical validation showing autonomous agents can ingest these specifications to generate functional, accurate reproduction pipelines from scratch. We envision this format as a new foundation for automated and conceptual reproducibility in machine learning.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Cookie-Bench: Web 生成のための継続的なオンスクリーンキーインタラクション評価

フロントエンドの Web コードは、すべてのフロンティア LLM リリースの中核的な製品面となっていますが、アリーナのような人間が判断するリーダーボードは拡張できないため、これらのインタラクティブ アプリケーションを開発スピードで評価することは依然としてコストがかかります。既存の自動プロキシは通常、リファレンス実装、テスト スイート、または厳密なチェックリストに依存しており、人間のレビュー担当者がライブ セッションで実行する合理的な合成を見逃す傾向があります。私たちは、同時に参照フリーで、自律的に駆動され、総合的に推論される新しい評価体制を明確にし、2 つの成果物を通じてそれをインスタンス化します。 \textbf{\dataname} は、静的プレゼンテーション タスクと対話型アプリケーション タスクの両方にまたがる 11 ドメイン、54 リーフ、1,000 クエリの WebDev ベンチマークであり、3 つの難易度層と 3 つのターゲット言語グループにわたってバランスが取れており、回覧されたプロンプトから思い出せないようにブリーフが書き直されています。 \textbf{\framename} は、Flavell のメタ認知モニタリングに基づいており、証拠の蓄積と判断を 3 つの段階にわたって分離します。静的な知覚は受動的な観察から第一印象を形成します。エージェント駆動のインタラクションは、連続画面のビデオ、音声、およびステップごとのスクリーンショットをキャプチャしながら、アプリケーションを自律的に探索します。動的スコアリングは、証拠チェーンが完了した後にのみ、構造化された失敗の帰属を伴う全体的な機能性と美的判断を発行します。 \dataname では、\framename は専門家による評価と厳密に一致しており、インタラクティブな Web 生成に関して 13 のフロンティア LLM 全体でかなりのヘッドルームを表面化しています。 \noindenthttps://anonymous.4open.science/r/Cookie-3CE/

原文 (English)

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation

Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed remains costly because human-judged leaderboards like Arena do not scale. Existing automated proxies typically lean on reference implementations, test suites, or rigid checklists, and tend to miss the reasoned synthesis a human reviewer performs over a live session. We articulate a new evaluation regime that is simultaneously reference-free, autonomously driven, and holistically reasoned, and instantiate it through two artifacts. \textbf{\dataname} is an 11-domain, 54-leaf, 1,000-query WebDev benchmark spanning both static-presentation and interactive-application tasks, balanced across three difficulty tiers and three target-language groups, with briefs rewritten to resist recall from circulated prompts. \textbf{\framename}, grounded in Flavell's metacognitive monitoring, separates evidence accumulation from judgment across three stages: Static Perception forms a first impression from passive observation; Agent-Driven Interaction explores the application autonomously while capturing continuous screen video, audio, and per-step screenshots; Dynamic Scoring issues holistic functionality and aesthetics verdicts with structured failure attribution only after the evidence chain is complete. On \dataname, \framename aligns closely with expert human ratings while surfacing substantial headroom across 13 frontier LLMs on interactive web generation. \noindenthttps://anonymous.4open.science/r/Cookie-3CE/

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

RAISE: アーキテクチャ検索問題としての RAG 設計

検索拡張生成 (RAG) システムでは、クエリの書き換え、チャンキング、検索の深さ、再ランキング、およびコンテキスト圧縮に及ぶ数多くの設計上の選択肢が明らかになります。実際には、これらの選択はヒューリスティックによって構成されることが多く、設定全体での体系的な評価と再現性が妨げられます。私たちは、この課題は RAG アーキテクチャの検索として定式化するのが最適であると主張します。この問題の制御された再現可能な研究をサポートするために、RAG ハイパーパラメータ最適化の包括的なフレームワークおよびベンチマークである RAG Intelligence Search Engine (RAISE) を導入します。これは、標準化された検索スペースと予算の下で RAG パイプラインの最適化方法を評価します。 RAISE は 13 の検索アルゴリズムを実装し、3 つのランダム シードを使用して 7 つのパブリック テキストおよびマルチモーダル データセットにわたってそれらを評価します。私たちの実験は、最適化のパフォーマンスがタスクに大きく依存することを示しています。つまり、あるデータセットで優れたパフォーマンスを発揮する手法が、他のデータセットでは一貫して一般化できない可能性があり、集計されたランキングを普遍的に優れた戦略の証拠として解釈することには注意が必要です。 RAISE は、RAG ハイパーパラメータの最適化に関する公正で再現性のある体系的な研究のための共通の実験基盤を提供します。

原文 (English)

RAISE: RAG Design as an Architecture Search Problem

Retrieval-augmented generation (RAG) systems expose numerous design choices spanning query rewriting, chunking, retrieval depth, reranking, and context compression. In practice, these choices are often configured through heuristics, hindering systematic evaluation and reproducibility across settings. We argue that this challenge is best formulated as RAG architecture search. To support controlled and reproducible study of this problem, we introduce the RAG Intelligence Search Engine (RAISE), a comprehensive framework and benchmark for RAG hyperparameter optimization, which evaluates optimization methods for RAG pipelines under standardized search spaces and budgets. RAISE implements 13 search algorithms and evaluates them across seven public text and multimodal datasets using three random seeds. Our experiments show that optimization performance is highly task-dependent: methods that perform strongly on one dataset may not generalize consistently across others, cautioning against interpreting aggregate rankings as evidence of universally superior strategies. RAISE provides a common experimental substrate for fair, reproducible, and systematic research on RAG hyperparameter optimization.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

矛盾する複数ソースの個人記憶に対する選択的 QA: 診断テストベッドと手法の比較

新興のパーソナル AI エージェントは、永続的なマルチソース メモリに移行しています。これにより、評価上の問題が生じます。システムは、矛盾する証拠や不完全な証拠をどのように使用するかを決定する必要があります。 1 つのきれいな歴史から事実を引き出すことはできません。既存のベンチマークでは、エラーがメソッドに与えられた証拠に起因するのか、メソッドの競合解決ステップに起因するのかを示すことはほとんどありません。私たちはこれを、矛盾する複数ソースの個人記憶に対する選択的 QA として研究しています。システムは、矛盾する、場合によっては不完全なソースに基づいて回答するか、証拠が不十分な場合は棄権します。 8 つの推論タイプにわたる 18 の質問テンプレート、480 のペルソナ、4 つのランダム シード、および 34,560 のインスタンスを含むベンチマークを、制御されたソースの歪みと決定論的なグラウンド トゥルースを使用して開発しました。ソースへのアクセスなし、単一ソースへのアクセス、構造化融合手法、およびフロンティア LLM のベースラインのパフォーマンスを評価します。最もよく訓練されたフュージョン リゾルバーの精度は 80.3% に達し、最も強力なプロンプトのみの LLM ベースラインは 70.0% に達します。棄権すると、同じリゾルバはカバレッジ 78.3% で選択精度 85.3% に達し、最良の LLM はカバレッジ 95.4% で選択精度 71.0% に達します。モデルが異なれば、推論タイプごとに異なる強みがあります。データ、コード、キャッシュされたモデル出力、およびデータ生成プロセスを再利用のためにリリースします。

原文 (English)

Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison

Emerging personal AI agents are moving toward persistent, multi-source memory. This creates an evaluation problem: systems must decide how to use conflicting or incomplete evidence; they cannot just retrieve facts from one clean history. Existing benchmarks rarely show whether an error came from the evidence given to a method or from the method's conflict-resolution step. We study this as selective QA over conflicting multi-source personal memory: systems answer based on conflicting, sometimes incomplete sources, or abstain when evidence is insufficient. We develop a benchmark containing 18 question templates across 8 reasoning types, 480 personas, 4 random seeds, and 34,560 instances, with controlled source distortions and deterministic ground truth. We evaluate the performance of baselines without access to any source, access to a single source, structured fusion methods, and frontier LLMs. The best trained fusion resolver reaches 80.3% accuracy, while the strongest prompt-only LLM baseline reaches 70.0%. With abstention, the same resolver reaches 85.3% selective accuracy at 78.3% coverage and the best LLM reaches 71.0% selective accuracy at 95.4% coverage. Different models have different strengths across reasoning types. We release the data, code, cached model outputs, and data-generating process for reuse.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達研究/論文

BioRefusalAudit: 一般およびドメイン微調整されたスパース オートエンコーダーを使用したバイオセキュリティ拒否の深さの監査

言語モデルのバイオセキュリティ評価では通常、モデルが危険な出力を生成するかどうかが問われます。この論文は補足的な質問をします。モデルが拒否した場合、その拒否は構造的に正しいのでしょうか、それともフレーミング、フォーマット、または出力長を促すための適度な変更で消えるのでしょうか? 5 つのアーキテクチャにわたって、無害性と危険性を明確に区別したモデルはありませんでした。 Gemma 2 2B-IT は、75 件のプロンプトにわたって真に拒否することはなく、危険に隣接するすべてのクエリを回避しました。 Gemma 4 E2B-IT は、チャット テンプレート形式を使用した場合は 65/75 件のプロンプトを拒否し、チャット テンプレート形式を使用しない場合は 0/75 件のプロンプトを拒否しました。両方の Gemma モデルは、80 トークンの上限の下で 0% に崩壊しました。 Qwen 2.5 1.5B と Phi-3-mini は過剰に拒否され、良性生物学の 83 ~ 87% が危険であると警告されました。 Llama 3.2 1B は唯一の意味のある Tier 勾配 (61 ポイントの広がり) を示しました。何がそのような過剰な拒否を引き起こすのかを調査するために、我々はスケジュールIであるが生物学的に無毒な化合物(特にFDA画期的治療法のステータスを持つシロシビン培養)のパネルをテストしました。一部のモデルは、真に有害な生物学を超える割合でこれらを拒否しており、拒否がCBRNの危険性に対する合法性と文化的顕著性を追跡していることを示唆しています。内部側を測定するために、モデルの表面応答ラベルを内部のスパース オートエンコーダー (SAE) 特徴のアクティベーションと比較する発散スコア D を導入します。フル D は、Gemma 2 2B-IT (Gemma Scope 1) および Gemma 4 E2B-IT (著者が訓練したバイオ SAE) で計算されました。 2 つの微調整された Gemma 2 ドメイン SAE がリリースされました。 Gemma 4 では、狭いカタログ、サンプル内キャリブレーション、および Gemma ファミリーのみの SAE 範囲を使用して、重複なし (n=75) で 0.647 ポイントのギャップで応答と拒否の応答が分離されますが、これは暫定的なものです。消費者向けハードウェア (GTX 1650 Ti Max-Q、および SAE トレーニング用の Colab T4) での 1 つのハッカソン週末にわたって構築されたこの予備的な証拠は、アクティベーション レベルの監査によって、アーキテクチャ間で大幅に異なる、動作評価では見えない障害モードが表面化する可能性があることを示唆しています。

原文 (English)

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model refuses, is that refusal structurally sound, or does it disappear under modest changes to prompt framing, formatting, or output length? Across five architectures, no model cleanly discriminated benign from hazard. Gemma 2 2B-IT never genuinely refused across 75 prompts, hedging on every hazard-adjacent query. Gemma 4 E2B-IT refused 65/75 prompts with chat-template formatting and 0/75 without it. Both Gemma models collapsed to 0% under an 80-token cap. Qwen 2.5 1.5B and Phi-3-mini over-refused, flagging 83-87% of benign biology as hazardous. Llama 3.2 1B showed the only meaningful tier gradient (61-point spread). To probe what drives such over-refusal, we tested a panel of Schedule I but biologically non-toxic compounds (notably psilocybin cultivation, with FDA Breakthrough Therapy status). Some models refused these at rates exceeding genuinely hazardous biology, suggesting refusal tracks legality and cultural salience over CBRN hazard. To measure the internal side, we introduce a divergence score D comparing a model's surface response label to its internal sparse autoencoder (SAE) feature activations. Full D was computed on Gemma 2 2B-IT (Gemma Scope 1) and Gemma 4 E2B-IT (author-trained bio SAE). Two fine-tuned Gemma 2 domain SAEs were released. On Gemma 4, comply and refuse responses separated by a 0.647-point gap with zero overlap (n=75), though this is preliminary, with a narrow catalog, within-sample calibration, and Gemma-family-only SAE coverage. Built over one hackathon weekend on consumer hardware (GTX 1650 Ti Max-Q, plus Colab T4 for SAE training), this preliminary evidence suggests activation-level auditing may surface failure modes invisible to behavioral evaluation, with substantial variation across architectures.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

オープンソースの安全ガード モデルのベンチマーク: 包括的な評価

安全性が重要なアプリケーションに大規模言語モデル (LLM) が導入されることが増えているため、堅牢なコンテンツ モデレーションが不可欠になっています。 NIST AI リスク フレームワークの 8 つの安全カテゴリにまたがる 79,331 サンプルの厳選されたベンチマークに基づく 14 のオープンソース安全ガード モデルの包括的な評価を示します。当社のベンチマークは 4 つの多様なデータセット (HarmBench、StrongREJECT、RealToxicityPrompts、BeaverTails) を集約し、安全関連のコンテンツ (暴力、ヘイトスピーチ、嫌がらせ、性的コンテンツ、自殺/自傷行為、冒涜、脅迫、健康上の誤った情報) のみに焦点を当てるようにフィルタリングされています。有害なコンテンツの欠落は誤検知よりも大きなリスクをもたらすため、リコールは安全性アプリケーションにとって重要な指標であることがわかりました。私たちの評価では、驚くべき結果が明らかになりました。Qwen Guard (4B パラメーター) は最高の再現率 (83.97%) を達成しましたが、Llama Guard (12B) や GPT-OSS Safeguard (20B) などのより大きなモデルは保守的な動作を示し、安全でないコンテンツを最大 75% 見逃しました。我々は、モデルのサイズが安全検出のパフォーマンスと相関しないこと、および汎用のガード モデルが特殊なガード モデルよりも優れていることを実証します。これらの調査結果は、実稼働環境での安全装置モデルを選択するための実践的なガイダンスを提供します。

原文 (English)

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a comprehensive evaluation of 14 open-source safety guard models on a curated benchmark of 79,331 samples spanning 8 NIST AI Risk Framework safety categories. Our benchmark aggregates four diverse datasets (HarmBench, StrongREJECT, RealToxicityPrompts, and BeaverTails), filtered to focus exclusively on safety-relevant content (violence, hate speech, harassment, sexual content, suicide/self-harm, profanity, threats, and health misinformation). We find that recall is the critical metric for safety applications, as missing harmful content poses greater risk than false positives. Our evaluation reveals surprising results: Qwen Guard (4B parameters) achieves the highest recall (83.97%) while larger models like Llama Guard (12B) and GPT-OSS Safeguard (20B) exhibit conservative behavior, missing up to 75% of unsafe content. We demonstrate that model size does not correlate with safety detection performance and that general-purpose guard models outperform specialized ones. These findings provide practical guidance for selecting safety guard models in production deployments.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

GPF-LiveNews: 大規模言語モデルにおけるグループ条件付きフレーミングのためのストリーミング評価プロトコル

デプロイされた言語モデルは非定常環境で評価されます。モデルのバージョン、検索レイヤー、安全システム、現実世界の入力はすべて時間の経過とともに変化します。静的バイアスのベンチマークは依然として有用ですが、モデルがさまざまな刺激を受けた視聴者に対して新たに出現したイベントをどのように組み立てるかは示していません。オープンエンド LLM 出力のグループ条件付きフレーミングを監査するためのストリーミング評価プロトコルおよびベンチマーク スナップショットである GPF-LIVENEWS を紹介します。このプロトコルは、42 の ID ラベルと 7 つのプロンプト ファミリにわたって新鮮な BBC/ロイター ニュース アンカーを拡張し、その後、意味論的感度とセンチメント差異シグナルを使用して応答バンドルを評価します。 12 回のモニタリング実行と 23 個のホストされたモデルにわたるパイロットでは、ポリシー/アクション プロンプトが最も強力なセマンティックな動きを生成しますが、センチメントの変動はディメンションおよびプロンプト ファミリ全体でより平坦です。リリースされたアーティファクトには、記事のメタデータ、プロンプト テンプレート、インスタンス化されたプロンプト、モデル出力メタデータ、スコア テーブル、ドキュメント、および再現スクリプトが含まれます。私たちはすべてのスコアを、永続的な公平性ランキングや有害なバイアスの直接の証拠としてではなく、人間によるレビューのための監視窓監査シグナルとして解釈します。

原文 (English)

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all change over time. Static bias benchmarks remain useful, but they do not show how models frame newly emerging events for different prompted audiences. We introduce GPF-LIVENEWS, a streaming evaluation protocol and benchmark snapshot for auditing group-conditioned framing in open-ended LLM outputs. The protocol expands fresh BBC/Reuters news anchors across 42 identity labels and seven prompt families, then evaluates response bundles using semantic-sensitivity and sentiment-disparity signals. In a pilot over 12 monitoring runs and 23 hosted models, Policy/Action prompts produce the strongest semantic movement, while sentiment variation is flatter across dimensions and prompt families. The released artifact includes article metadata, prompt templates, instantiated prompts, model-output metadata, score tables, documentation, and reproduction scripts. We interpret all scores as observed-window audit signals for human review, not as permanent fairness rankings or direct proof of harmful bias.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

GrowLoop: 人間がシードし、自己進化する会話評価

大規模な言語モデルの急速な進歩に伴い、自由な会話における人間らしさを評価することがますます重要になってきています。しかし、人間らしさは人間が直感的に認識する暗黙知の一種ですが、根底にある基準は明示的な定式化に抵抗します。人間の判断は大きく異なり、一部のケースでは強い同意が得られますが、他のケースでは正当な意見の相違が見られます。一方、人間の判断の背後にある基準は暗黙的なままであり、事件を構築するための明確な根拠は残されていません。さらに、人間に似ているとみなされるものは静的なものではなく、モデルの能力と人間の期待に応じて進化します。専門家が作成したベンチマーク、報酬モデル、自己進化型ベンチマークなどの評価方法は進歩していますが、3 つの課題すべてに同時に対処できるものはありません。そこで、モデルの進歩やシナリオの変化に合わせて継続的に適応する、自己進化する会話評価システムである GrowLoop を提案します。最初の動きとして最小限の人間のシード アノテーションを使用して、LLM エージェントはヒューリスティック学習を通じて評価ルーブリックを繰り返し抽出し、改良します。アノテーターが集まる場合には人間と AI の合意が必要ですが、異なる場合には妥当性のみが期待されます。さらに、Rubric-Caseの共進化機構により、評価対象が移動した際に新たなシーズを介して拡張され、継続的な進化が可能となります。自由形式の会話における人間らしさの評価に適用すると、生成されたルーブリックは、人間の判断に沿って既存の手法を大幅に上回るだけでなく、アノテーターが見落としている問題も明らかになります。結果として得られるベンチマークは、機能層全体でモデルを効果的に識別し、どこが不足しているかを明らかにすると同時に、新しいシナリオに一般化し、モデルの進歩に合わせて適応します。私たちの取り組みは、ベンチマークのパラダイムを手動の更新や難易度のスケーリングから、包括的で継続的な自己進化へと移行させます。

原文 (English)

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However, human-likeness is a form of tacit knowledge that humans perceive intuitively, yet the underlying criteria resist explicit formulation. Human judgments vary widely, with strong agreement on some cases and legitimate disagreement on others. Meanwhile, the criteria behind human judgments remain implicit, leaving no clear basis for constructing cases. Further, what counts as human-like is not static, but evolving with model capability and human expectations. Despite progress in evaluation methods such as expert-authored benchmarks, Reward Models, and self-evolving benchmarks, none addresses all three challenges simultaneously. Therefore, we propose GrowLoop, a self-evolving conversation evaluation system that continuously adapts as models advance and scenarios shift. With minimal human seed annotations as the first mover, LLM agents iteratively extract and refine evaluation rubrics through Heuristic Learning. Human-AI agreement is required where annotators converge, while only plausibility is expected where they diverge. Moreover, the Rubric-Case co-evolution mechanism enables continuous evolution, expanded through new seeds when the evaluation target moves. Applied to human-likeness evaluation in open-ended conversation, the generated rubrics not only substantially outperform existing methods in alignment with human judgments, but also uncover issues that annotators overlook. The resulting benchmark effectively discriminates models across capability tiers and reveals where they fall short, while generalizing to new scenarios and adapting as models advance. Our work shifts the benchmarking paradigm from manual updates or difficulty scaling to comprehensive, continuous self-evolution.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

LoRe: 反復グラフ ソルバー向けのステップごとのインタラクション バジェットを備えた適応型インタラクション評価ルーティング

組み合わせ最適化のための拡散ベースのニューラル ソルバーは、高密度のエッジ/因子相互作用を繰り返し再評価するため、実時間での推論が高価になり、大規模になるとメモリに制限されることがよくあります。多体物理学の計算手法にインスピレーションを得て、ステップごとの相互作用評価の予算設定を強制する、トレーニング不要の推論時間ドロップイン ラッパーである LoRe を導入します。各反復では、固定のスパース化 (静的 kNN グラフや静的など) を使用する代わりに、計算を競合性の高い相互作用または不確実性の高い相互作用に動的にルーティングすることで、相互作用の固定部分のみを評価します。マスク)。完全に包括的なエンドツーエンドの壁時計アカウンティングの下で​​、LoRe は最大独立集合 (MIS) 問題のスケーラビリティを大幅に向上させ、実行可能な推論をベースラインのメモリ不足制限を超えて $3\times$ 以上拡張し、$\sim 8\times$ の高速化と $\sim 12\times$ のピークメモリ削減を実現し、この体制でソリューションの品質は維持されます。大規模な巡回販売員問題 (TSP) に対するクロスタスクの汎用性と、トポロジーの変化に対するゼロショットの堅牢性を実証する LoRe は、$n=1000$ で $\sim 15\times$ の高速化を実現し、$44\times$ のメモリ削減と競争力のあるツアー品質を実現します。

原文 (English)

LoRe: Adaptive Interaction-Evaluation Routing with Per-Step Interaction Budgets for Iterative Graph Solvers

Diffusion-based neural solvers for combinatorial optimization repeatedly re-evaluate dense edge/factor interactions, making inference expensive in wall-clock time and often memory-bound at scale. Inspired by the computational methodologies of many-body physics, we introduce LoRe, a training-free, inference-time drop-in wrapper that enforces per-step interaction-evaluation budgeting: at each iteration, it evaluates only a fixed fraction of interactions by dynamically routing computation to high-conflict or high-uncertainty interactions, instead of using a fixed sparsification (e.g., static kNN graphs or static masks). Under fully inclusive end-to-end wall-clock accounting, LoRe substantially improves scalability on the Maximum Independent Set (MIS) problem, extending feasible inference more than $3\times$ beyond the baseline's out-of-memory limit, delivering a $\sim 8\times$ speedup and a $\sim 12\times$ peak-memory reduction, with solution quality preserved in this regime. Demonstrating cross-task generality on the large-scale Traveling Salesperson Problem (TSP) and zero-shot robustness to topology shifts, LoRe achieves a $\sim 15\times$ speedup at $n=1000$ with a $44\times$ memory reduction and competitive tour quality.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

倫理的な顔年齢推定に向けて: 子供のデータに関するトレーニングを行わない一般化されたゼロショット ベンチマーク

顔画像からの年齢推定は通常、未成年者の画像を含むトレーニング データに依存しますが、これは倫理的、法的、プライバシー上で重大な懸念を引き起こす行為です。この研究では、若い集団に対するモデルのパフォーマンスを評価しながら、トレーニング中に子供のデータを明示的に除外する、顔年齢推定のための一般化されたゼロショット ベンチマークを提案します。私たちは、広く使用されている 6 つのデータセットを再検討し、年齢グループを厳密に分離した標準化された分割を導入します。トレーニング、検証、テストには 18 ~ 59 歳のサンプルを使用します。 18 歳未満のサンプルはゼロショット評価専用に予約されています。そして、分布シフトの下でのモデル選択のための未確認の検証セットとして 60 以上のサンプルをサンプリングします。 ID アノテーションを持つデータセットの場合、サブジェクト排他的な分割により ID 漏洩が防止され、現実世界のデプロイ状況がより適切に反映されます。このプロトコルに基づいて 9 つの最先端の年齢推定方法を評価すると、評価されたすべての方法が目に見えない年齢グループに一貫して一般化できず、教師付きベースラインと比較して、平均で 46.4%、最大で 52.8% という大幅なパフォーマンスの低下が見られることが明らかになりました。さらに、モデルは単純に劣化するわけではありません。モデルは、目に見えない年齢の予測を近くの見たクラスに体系的に固定します。これは、一般化されたゼロショット学習におけるよく知られた見えるクラスのバイアスの現れです。この研究では、子供のデータを使用しない年齢推定を既存のデータセットに対する一般化されたゼロショット ベンチマークとして形式化することで、現在のモデリング実践と現実世界の倫理的制約との間の重大なギャップを浮き彫りにしています。私たちのベンチマークは、制限されたデータ体制の下でモデルを評価するための原則に基づいた基礎を提供し、分布の変化に強く、責任あるデータの使用に合わせた方法の開発を奨励します。

原文 (English)

Toward Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data

Age estimation from facial images typically relies on training data that includes images of minors, a practice that raises serious ethical, legal, and privacy concerns. In this work, we propose a generalized zero-shot benchmark for facial age estimation that explicitly excludes children's data during training while still assessing model performance on younger populations. We revisit six widely used datasets and introduce standardized splits with strict age-group separation: samples aged 18-59 for training, validation, and testing; samples under 18 reserved exclusively for zero-shot evaluation; and samples 60+ as an unseen validation set for model selection under distribution shift. For datasets with identity annotations, subject-exclusive splits prevent identity leakage and better reflect real-world deployment conditions. Evaluating nine state-of-the-art age estimation methods under this protocol reveals that all evaluated methods consistently fail to generalize to unseen age groups, suffering substantial performance degradation -- on average 46.4%, and up to 52.8% -- relative to the supervised baseline. Moreover, models do not simply degrade: they systematically anchor predictions for unseen ages to nearby seen classes, a manifestation of the well-known seen-class bias in generalized zero-shot learning. By formalizing age estimation without children's data as a generalized zero-shot benchmark on existing datasets, this work highlights a critical gap between current modeling practices and real-world ethical constraints. Our benchmark provides a principled basis for evaluating models under restricted data regimes and encourages the development of methods that are robust to distribution shift and aligned with responsible data use.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

DynSess: ロールプレイング エージェント向けの動的なセッション レベルの評価および最適化フレームワーク

大規模な言語モデルを使用したロールプレイングは基本的にセッション レベルのタスクであり、エージェントは長時間にわたる複数ターンの会話にわたってキャラクターのアイデンティティと対話の品質を維持する必要があります。しかし、既存の評価および最適化手法は依然としてターンレベルにとどまっており、長期的な品質を捉えることができません。私たちは、ロールプレイング エージェントのための統合されたセッション レベルのフレームワークである DynSess を提案します。 DynSess-Eval は、長期的な行動を対象としたルーブリックを介して、完全な対話セッションをスコア付けします。セッションレベルの報酬を活用して、マルチターン先読み検索を通じて高品質のトレーニング軌道を構築し、2 つの補完的なバリアントである DSPO (オフポリシー) と GSRPO (オンポリシー) で DynSess-Character をトレーニングします。実験では、DynSess-Eval が以前の評価者よりも人間の判断とかなり良く一致していることが示されており、人間によるブラインド評価ではさらに、DynSess-Character が、使用するパラメータが大幅に少ないにもかかわらず、強力な役割の一貫性とインタラクティブな能力を維持しながら、最強のキャラクター モデルと一致していることが示されています。私たちのデータセットとコードは、将来の研究を促進するためにリリースされます。

原文 (English)

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality across extended multi-turn conversations. Yet existing evaluation and optimization methods remain largely turn-level, failing to capture long-horizon quality. We propose DynSess, a unified session-level framework for role-playing agents. DynSess-Eval scores complete dialogue sessions via rubrics targeting long-horizon behaviors. Leveraging its session-level rewards, we construct high-quality training trajectories through multi-turn lookahead search and train DynSess-Character with two complementary variants: DSPO (off-policy) and GSRPO (on-policy). Experiments show that DynSess-Eval aligns with human judgments substantially better than prior evaluators, and blind human evaluation further shows that DynSess-Character matches the strongest character model despite using substantially fewer parameters, while maintaining strong role consistency and interactive ability. Our dataset and code will be released to facilitate future research.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

物理基礎モデルは一般化可能な物理学を学習しますか?物理的体制と分布の変化にわたるバイアスを意識したベンチマーク

最近の物理基礎モデルは一般的な時空間予測能力を主張していますが、その評価は、固定されたトレーニング分布の下でパフォーマンスを単一の平均スコアに落とし込んでしまうことがよくあります。このため、モデルが一般化可能な物理ダイナミクスを学習しているのか、それとも特定の設定下でのみ適切に動作するのかを判断することが困難になります。 8 つの物理ダイナミクス、3 つのトレーニング データ混合物、および動的スケールと初期条件の複雑さのシフトによって誘発される 25 のテスト レジームを使用してベンチマークを構築し、分布内、分布シフト、および分布外の設定をカバーします。 5 つの物理基礎モデル アーキテクチャとアーキテクチャごとに 4 つのモデル バリアント (スクラッチ サイズと 3 つの事前トレーニング サイズ) を評価し、結果として 60,000 の測定結果が得られます。私たちの結果は、現在の物理基礎モデルが普遍的なジェネラリストとしてではなく条件付きで動作することを示しています。その一般性は、物理レジーム、時間スケール、初期条件の設定、事前トレーニング、モデルのサイズ、アーキテクチャに依存します。トレーニング データの分散を改善しても、この制限は部分的にしか緩和されません。事前トレーニングとスケーリングでも、能力のバイアスを確実に取り除くことができません。私たちは、物理基礎モデルを改善するには、モデルのスケーリングやデータの拡張を超えて、領域、時間スケール、分布の変化を超えて移転可能な物理知識をより適切に捕捉する学習メカニズムに移行する必要があると主張します。

原文 (English)

Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts

Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a single average score under a fixed training distribution. This makes it difficult to determine whether a model has learned generalizable physical dynamics or only performs well under particular settings. We construct a benchmark with 8 physical dynamics, 3 training-data mixtures, and 25 test regimes induced by dynamic-scale and initial-condition complexity shifts, covering in-distribution, distribution-shift, and out-of-distribution settings. We evaluate five physics foundation model architectures and four model variants per architecture (scratch and three pretrained sizes), resulting in 60,000 measurements. Our results show that current physics foundation models behave as conditional rather than universal generalists: their generality depends on the physical regime, temporal scale, initial-condition setting, pretraining, model size, and architecture. Improving the training data distribution only partially mitigates this limitation. Pretraining and scaling are also unable to reliably remove their ability biases. We argue that improving physics foundation models requires moving beyond scaling models or expanding data, toward learning mechanisms that better capture transferable physical knowledge across regimes, temporal scales, and distribution shifts.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

Pocket-Dentist: 効率的なマルチモーダル大規模言語モデルによるオンデバイス歯科画像理解

歯科視覚言語モデルの評価は、データセット、タスク定義、メトリクスにわたって断片化されたままであり、多くの場合、その計算コストが無視されます。このため、専門センター外での歯科スクリーニングへの広範な導入が制限されています。専門センターでは、タイムリーな推論、限られたハードウェア、および患者画像のローカル処理が、実用的でプライバシーを保護した臨床前スクリーニングに不可欠です。ここで紹介する Pocket-Dentist は、約 1,159 人の患者、5 つのタスク タイプ、7 つの指標にまたがる 3 つのデータセットをまとめた、歯科マルチモーダル質問応答のための効率を意識したベンチマークです。典型的な 14 個の VLM にわたって、我々の結果は興味深い観察結果を明らかにしました。コンパクトな VLM (例: 2B パラメータ モデル) は、精度においては大型の VLM を上回っていますが、歯科画像の理解に必要な計算コストは​​大幅に低くなります。 iPhone 17 Pro にローカルに導入された当社の微調整されたコンパクト VLM Pocket-Dentist-2B は、各サンプルを 4.31 秒で処理し、7B ベースラインと比較してレイテンシーを 4.9 倍、メモリ使用量を 2.3 倍削減しました。

原文 (English)

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widespread deployment for dental screening outside specialist centres, where timely inference, limited hardware, and local handling of patient images are vital for practical, privacy-preserving clinical prescreening. Here we present Pocket-Dentist, an efficiency-aware benchmark for dental multimodal question answering that brings together three datasets spanning approximately 1,159 patients, five task types and seven metrics. Across typical 14 VLMs, our results reveals an interesting observation: compact VLMs (e.g., 2B-parameter models) outperform larger VLMs in accuracy while requiring substantially lower computational costs in dental image understanding. Deployed locally on an iPhone 17 Pro, our finetuned compact VLM Pocket-Dentist-2B processed each sample in 4.31 s, reducing latency by 4.9-fold and memory use by 2.3-fold compared with a 7B baseline.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

データセットの価値はいくらですか?スケーリング則、Vendi スコア、および行列スペクトル関数

ニューラル スケーリングの法則はデータセットのサイズを通じてデータを評価しますが、Vendi スコアは量子エントロピーを使用してデータセットの値を測定します。一般的なニューラル スケーリング則の目標と Vendi スコアの両方がサブモジュールであることを示します。さらに、Vendi スコアが、行列スペクトル関数と呼ばれるより広範なクラスのサブモジュラー目標の特殊なケースであることを示します。これには、決定的 (DPP) 目標や他の多くの目標も含まれます。また、弱行列単調関数を導入し、それがどのように弱部分モジュール行列スペクトル関数につながるかを示し、データ評価のための幅広い実用的な目的をもたらします。私たちは、貪欲な最適化中に繰り返される固有分解を回避する永年方程式ベースの更新を開発し、$m$ 次元の埋め込みに対する限界ゲイン評価を Oracle クエリと比較して $O(m)$ 係数だけ削減します。これにより、経験的に平均約 35,000 倍の高速化が得られ、ImageNet-1K スケールのデータセットで Vendi スコアの直接最適化が可能になります。このようにして可能になったので、Vendi スコア、DPP、施設の場所、および 3 つの新しいマトリックス スペクトル バリアントを含む、固定サイズ、クラスバランス、および固定トレーニング予算体制の下で、いくつかの目標がホールドアウト テスト パフォーマンスのトレーニング サブセットの値をどの程度正確に予測するかを比較します。複数のデータセットにわたって、施設の位置が最も優れたパフォーマンスを発揮します。また、直接最適化では、Vendi スコアは中程度のスコア範囲では予測的ですが、目標をより高い値に押し上げると、下流のパフォーマンスの代用として機能しなくなる可能性があることも明らかになりました。また、均一でランダムな固定サイズのサブセットは、制約がなく、クラスバランスが取れていても、評価スコアと保持されたパフォーマンスの両方で著しく集中していることもわかります。最後に、サイズ、クラスのバランス、トレーニング予算だけがデータの価値を決定するわけではないことを示します。これらの要因を制御した場合でも、パフォーマンスは良い状態から悪い状態まで滑らかに変化します。

原文 (English)

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions

Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common neural-scaling-law objectives and the Vendi Score are submodular. We further show that the Vendi Score is a special case of a broader class of submodular objectives that we call matrix spectral functions. This also includes determinantal (DPP) objectives, as well as many others. We also introduce weakly matrix monotone functions and show how they lead to weakly submodular matrix spectral functions, yielding a broad family of practical objectives for data appraisal. We develop secular-equation-based updates that avoid repeated eigendecompositions during greedy optimization, reducing marginal-gain evaluation for $m$-dimensional embeddings by an $O(m)$ factor relative to oracle queries. This yields an average empirical speedup of about 35,000x, making direct optimization of the Vendi Score feasible on ImageNet-1K-scale datasets. Thus enabled, we compare how well several objectives predict the value of training subsets for held-out test performance under fixed-size, class-balanced, and fixed training-budget regimes, including the Vendi Score, DPPs, facility location, and three new matrix spectral variants. Across multiple datasets, facility location performs the best. Direct optimization also reveals that, while the Vendi Score is predictive over moderate score ranges, pushing the objective to higher values can make it a poor downstream performance proxy. We also find that uniformly at random fixed-size subsets, both unconstrained and class-balanced, are remarkably concentrated in both appraisal scores and held-out performance. Finally, we show that size, class balance, and training budget do not alone determine data value: even when controlling for these factors, performance ranges smoothly from good to bad.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

CFMME 上の大規模なビジョン言語モデルのベンチマーク: 包括的な中国金融マルチモーダル評価データセット

Large Vision-Language Models (LVLM) の出現により、モデルの機能がテキストのみの理解を超えて大幅に拡張され、視覚的モダリティとテキストモダリティの両方にわたる統一された推論が可能になり、より広範な現実世界のアプリケーションがサポートされるようになりました。中国の状況における金融ビジネスのワークフロー全体を通じて、LVLM の認識、理解、推論、認知能力を包括的に評価するために、中国の新しい金融マルチモーダル評価ベンチマークである CFMME を紹介します。 CFMME は、基礎的な学術知識から複雑な現実世界のアプリケーションに至る 6,052 のインスタンスで構成され、8 つの主要な金融イメージ モダリティと 4 つのコア マルチモーダル タスクをカバーします。 CFMMEでは、代表的なLVLMを徹底的に評価します。結果は、最先端のモデルが質問応答タスクで全体の精度 66.11%、検出、認識、および情報抽出タスクで平均スコア 77.18 を達成していることを示しており、現在の LVLM には改善の余地がかなりあることが示されています。さらに、エラーの原因、クロスモーダル機能、および複数の方向設定の詳細な分析を実施し、将来の研究に役立つ貴重な洞察をもたらします。私たちは、CFMME が、特に金融ドメインにおける複数のマルチモーダル タスクのパフォーマンスを向上させることによって、LVLM のさらなる進歩に拍車をかけることを期待しています。

原文 (English)

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified inference across both visual and textual modalities and supporting a broader range of real-world applications. To comprehensively evaluate the perception, understanding, reasoning, and cognition capabilities of LVLMs throughout the entire financial business workflow in Chinese contexts, we introduce CFMME, a novel Chinese financial multimodal evaluation benchmark. CFMME comprises 6,052 instances spanning from fundamental academic knowledge to complex real-world applications, covering eight primary financial image modalities and four core multimodal tasks. On CFMME, we conduct a thorough evaluation of representative LVLMs. The results show that the state-of-the-art model attains an overall accuracy of 66.11\% on the question answering task and an average score of 77.18 on the detection, recognition, and information extraction tasks, indicating substantial room for improvement in current LVLMs. In addition, we conduct detailed analyses of error causes, cross-modal capabilities, and multi-orientation settings, yielding valuable insights for future research. We hope that CFMME will spur further progress in LVLMs, especially by improving their performance on multiple multimodal tasks in the financial domain.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

オフポリシー評価のための商 DAG: フォワードフロー重要度サンプリングと正確なスレート傾向

オフポリシー評価は、別の動作ポリシーによって収集されたデータを使用して、ターゲットポリシーがどのように実行されるかを推定します。これは、推奨や医療など、オンラインテストにコストがかかる、またはリスクが伴う場合に非常に重要です。標準重要度サンプリングでは、ログに記録された各軌跡の重み付けが変更されますが、評価ターゲットがそれらを無視する場合でも、生成プロセスの詳細を意味のあるものとして扱うことができます。たとえば、自己回帰スレート レコメンダーは順序付けられたアイテムのシーケンスを生成する一方で、報酬と下流の推定器は順序付けられていないスレートのみに依存する場合があります。厳密に順序付けされていないスレートの傾向にはすべての生成順序の合計が必要なため、これにより迷惑な分散と計算上のギャップが生じます。評価用に同等の履歴をマージし、マージされたグラフ上のターゲットと行動の順方向フロー比を使用して重みを割り当てる商 DAG ビューを導入します。 set-sufficient next-item インターフェイスの下でのスレート推奨の場合、階乗列挙なしで正確な順序付けされていない傾向を計算するサブセット DAG 動的プログラムである Forward-DP が生成されます。結果として得られる傾向プリミティブにより、コンテキスト依存の自己回帰スレート ロガーの実用的な傾向ベースの評価とモデル選択が可能になります。

原文 (English)

Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities

Off-policy evaluation estimates how a target policy would perform using data collected by a different behavior policy, which is crucial when online testing is costly or risky, such as in recommendation or healthcare. Standard importance sampling reweights each logged trajectory, but it can treat details of the generation process as meaningful even when the evaluation target ignores them: for example, an autoregressive slate recommender may generate an ordered sequence of items while the reward and downstream estimator depend only on the unordered slate. This creates nuisance variance and a computational gap, since exact unordered slate propensities require summing over all generation orders. We introduce a quotient-DAG view that merges histories equivalent for evaluation and assigns weights using target-to-behavior forward-flow ratios on the merged graph. For slate recommendation under a set-sufficient next-item interface, this yields Forward-DP, a subset-DAG dynamic program that computes exact unordered propensities without factorial enumeration. The resulting propensity primitive enables practical propensity-based evaluation and model selection for context-dependent autoregressive slate loggers.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

GUITestScape: 探索的 GUI テストのオープンセット評価に向けて

探索的 GUI テストは、MLLM エージェントにとって特に要求の厳しい設定です。事前定義されたテスト スクリプトがなければ、エージェントは自律的にアプリケーションを操作し、独自の対話を通じて欠陥を発見する必要があります。しかし、現在の評価は 2 つの点で不十分です。まず、既存のベンチマークはほぼインタラクションの欠陥のみに焦点を当てており、ディスプレイの欠陥は評価の枠外に残されています。第 2 に、評価プロトコルは事前定義された欠陥の注釈にバインドされており、テスト プロセスが質的に異なる故障モードを混同した単一の最終状態の判断に崩壊します。これらの課題に対処するために、61 の実世界の Android アプリケーションと、インタラクションと表示タイプにわたる 508 のプリセット欠陥をカバーするインタラクティブなベンチマークである GUITestScape を紹介し、エージェントのテストの軌跡を独立して診断可能な機能に分解するオープンセット評価器である GUIJudge を紹介します。実験結果は、GUIJudge が事前定義された注釈を超えた信頼性の高いプロセス認識型評価を達成し、すべてのベースラインを大幅に上回るパフォーマンスを示していることを示しています。 GUITestScape のベンチマークでは、両方の欠陥タイプにわたって、既存のモデルにとって検出が依然として重大なボトルネックであること、および GUIJudge のベリファイアを既存のエージェントに統合することで、再トレーニングすることなく検出パフォーマンスが大幅に向上することがさらに明らかになりました。

原文 (English)

GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an application and discover defects through its own interaction. However, current evaluation falls short on two fronts. First, existing benchmarks focus almost exclusively on interaction defects, leaving display defects outside the evaluation frame. Second, evaluation protocols are bound to predefined defect annotations, collapsing the testing process into a single end-state judgment that conflates qualitatively distinct failure modes. To address these challenges, we present GUITestScape, an interactive benchmark covering 61 real-world Android applications and 508 preset defects spanning interaction and display types, and introduce GUIJudge, an open-set evaluator that decomposes an agent's testing trajectory into independently diagnosable capabilities. Experimental results demonstrate that GUIJudge achieves reliable process-aware evaluation beyond predefined annotations, substantially outperforming all baselines. Benchmarking on GUITestScape further reveals that detection remains the critical bottleneck for existing models across both defect types, and that integrating GUIJudge's verifiers into existing agents significantly boosts their detection performance without retraining.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

EviLink: 大規模な Text-to-SQL のための不確実性に基づく証拠取得を使用したマルチパス スキーマ リンク

スキーマのリンクは、大規模な Text-to-SQL では困難かつ重要なステップであり、システムは大規模で曖昧なデータベースからコンパクトでありながら十分なスキーマ コンテキストを識別する必要があります。既存の方法では、多くの場合、単一の SQL パスに関する決定論的な選択としてスキーマ リンクが扱われますが、複雑な質問では、異なるスキーマ ニーズを持つ複数の有効な実現が認められる場合があります。スキーマ リンクを、複数の妥当な SQL パスにわたる不確実性を認識したスキーマ ニーズ推論として再構成します。システムは、必要なスキーマ項目をパスに依存する不確実な項目から区別し、必要な場合にのみ証拠を取得します。私たちは、この再構成を EviLink でインスタンス化します。これは、複数の仮説スキーマの根拠と、不確実性に基づいた証拠の取得を組み合わせたものです。 BIRD-Dev と Spider2-Snow の実験では、この観点により、スキーマの完全性、スキーマの関連性、トークン コストの間のバランスが改善されることが示されています。 Spider2-Snow では、EviLink はフィールドレベルの厳密な再現率 90.15% を達成し、平均 123.30K のトークンを使用し、固定ジェネレータの下でダウンストリーム SQL 生成を改善します。

原文 (English)

EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL

Schema linking is a difficult and important step in large-scale Text-to-SQL, where systems must identify a compact yet sufficient schema context from large and ambiguous databases. Existing methods often treat schema linking as deterministic selection around a single SQL path, but complex questions may admit multiple valid realizations with different schema needs. We reframe schema linking as uncertainty-aware schema-need inference over multiple plausible SQL paths, where the system distinguishes required schema items from path-dependent uncertain ones and acquires evidence only where needed. We instantiate this reframing with EviLink, which combines multi-hypothesis schema grounding with uncertainty-guided evidence acquisition. Experiments on BIRD-Dev and Spider2-Snow show that this perspective improves the balance among schema completeness, schema relevance, and token cost. On Spider2-Snow, EviLink achieves 90.15% field-level strict recall rate, uses 123.30K average tokens, and improves downstream SQL generation under a fixed generator.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Honeyval: LLM を利用した HTTP ハニーポットの包括的な評価フレームワーク

ハニーポットは、サイバー攻撃から防御するために設計された実際のシステム コンポーネントを模倣したおとりシステムです。最近では、LLM がハニーポットのシミュレーション バックボーンとして機能することが増えています。これらにより、防御者はシステム セキュリティ リスクを低く抑えながら、インタラクションの多いハニーポットを構築できます。ただし、LLM を利用したハニーポット開発には、統一された評価フレームワークがありません。ほとんどの評価は、固定コマンド、手動テスト、または実際の展開での応答の類似性の測定で構成されます。これらの手法は、多くの場合、開発の拡張性、評価全体での再現性、実際の攻撃の代表性、またはさまざまな攻撃者やハニーポットの構成への適応性がありません。この取り組みでは、このギャップを埋め、LLM を利用した HTTP ハニーポットの包括的な評価フレームワークである Honeyval を提案します。私たちは、16 のバックエンド アプリケーションにハニーポットを設置し、AI ハッキング エージェントを攻撃者として使用し、カスタマイズ全体にわたってエージェントとハニーポットの機能を監視する 2 つの制御タスクを採用し、攻撃者に対する明確で検証可能なエクスプロイト目標を定義することで、以前の評価の制限に対処しました。 Honeyval を使用して、HTTP ハニーポットとしての最近のコスト効率の高い LLM の広範な評価を実施します。私たちの実験は、LLM を利用したハニーポットの将来性を強調しています。これらは、ルールベースのベースライン ハニーポットよりも攻撃者とのやり取りが大幅に長くなり、フロンティア モデルでも検出される頻度がはるかに低くなり、平均してエージェント攻撃者に対するランニング コストの利点が維持されます。さらに、さまざまな反攻撃用ハニーポット構成を実験し、検出力の向上と引き換えにインタラクションが長くなるなど、独特のトレードオフを観察しました。

原文 (English)

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation backbones for honeypots. They enable defenders to construct high-interaction honeypots with low system security risks. However, LLM-powered honeypot development lacks a unified evaluation framework. Most evaluations consist of measuring response similarity on fixed commands, manual testing, or real-world deployment. These methods are often not scalable for development, reproducible across evaluations, representative of practical attacks, or adaptable to various attacker and honeypot configurations. In this work, we bridge this gap and propose Honeyval, a comprehensive evaluation framework for LLM-powered HTTP honeypots. We address the limitations of prior evaluations by grounding the honeypots in 16 backend applications, using AI hacking agents as attackers, employing two control tasks to monitor agent and honeypot capabilities across customizations, and defining clear and verifiable exploit goals for the attacker. Using Honeyval, we conduct an extensive evaluation of recent cost-efficient LLMs as HTTP honeypots. Our experiments highlight the promise of LLM-powered honeypots; they lead to substantially longer interactions with the attacker than rule-based baseline honeypots and are far less frequently detected even by frontier models, all while, on average, preserving a running cost advantage against agentic attackers. Further, we experiment with different counter-offensive honeypots configurations, and observe unique trade-offs, such as longer interactions at the cost of increased detection.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

大規模なオーディオ言語モデルにおけるオーディオ ジェイルブレイク: 分類法、攻撃防御分析、コストを意識した評価

Large Audio Language Model (LALM) は、脱獄のリスクをトークン レベルのプロンプトから音声認識から推論までのパイプライン全体に拡大します。セマンティクス、音響スタイル、信号アーティファクト、または内部表現を通じて安全でない動作が誘発される可能性があります。既存の研究では、異種混合の脅威モデルと評価プロトコルに基づいてこれらのリスクが研究されており、攻撃の実用性や防御の有用性を比較することが困難になっています。このペーパーでは、LALM ジェイルブレイク攻撃と防御の統一された分類法と管理された経験的評価を提供します。私たちはこれまでの研究をセマンティック攻撃、音響攻撃、信号攻撃、埋め込み層攻撃に分けて整理しています。ガードベース、トレーニング不要、トレーニングベースのディフェンス。クロスモーダル、オーディオネイティブ、およびインタラクティブなベンチマーク。次に、10 個のオープンソース LALM にわたる代表的な攻撃と防御を評価し、攻撃の成功率だけでなく、良性の拒否と遅延も測定します。私たちの結果は、Acoustic Best-of-N が強力な最悪の場合のオーディオ空間の脆弱性を明らかにし、Narrative Framing が効果的な低レイテンシのセマンティック脅威であり、現在の防御は無害なユーザビリティに対する堅牢性と引き換えにあることを示しています。これらの調査結果は、成功率のみを重視した LALM の安全性ベンチマークを補完する必要があるものとして、コストとユーティリティを意識した評価を裏付けています。

原文 (English)

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsafe behavior can be induced through semantics, acoustic style, signal artifacts, or internal representations. Existing work studies these risks under heterogeneous threat models and evaluation protocols, making it difficult to compare attack practicality or defense utility. This paper provides a unified taxonomy and a controlled empirical evaluation of LALM jailbreak attacks and defenses. We organize prior work into semantic, acoustic, signal, and embedding-layer attacks; guard-based, training-free, and training-based defenses; and cross-modal, audio-native, and interactive benchmarks. We then evaluate representative attacks and defenses across ten open-source LALMs, measuring not only attack success rate but also benign refusal and latency. Our results show that Acoustic Best-of-N reveals strong worst-case audio-space vulnerabilities, Narrative Framing is an effective low-latency semantic threat, and current defenses trade robustness against benign usability. These findings support cost- and utility-aware evaluation as a necessary complement to success-rate-only LALM safety benchmarks.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MedCase 構造化: 臨床的に現実的な EHR 設定における診断推論のベンチマーク用の Text-to-FHIR データセット

大規模言語モデル (LLM) は、臨床推論と意思決定のサポートに有望ですが、現実的な電子医療記録と一致する設定での評価には依然として限界があります。既存のベンチマークは、多くの場合、臨床システムで使用される構造化された相互運用可能なデータ形式を反映していない静的データセットまたは非構造化入力に依存しています。非構造化テキストから臨床的に現実的な HL7 FHIR R4 バンドルを生成するパイプラインを導入し、臨床意思決定支援システムの制御可能な評価を可能にします。このパイプラインは、段階的な LLM 生成と用語に基づいた検証および修復を組み合わせて、幻覚コードを削減し、構造的および意味的な一貫性を強化します。このアプローチを MedCaseReasoning に適用して、臨床医が作成した診断症例に合わせた合成データセットである MedCase-Structured を構築し、症例の 82.5% で有効な FHIR 生成を実現します。 MedCase-Structured での評価では、平文の場合よりも構造化 FHIR 入力での LLM の診断精度が一貫して低いことが明らかになり、展開に合わせたベンチマークの重要性が強調されています。

原文 (English)

MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in realistic, electronic health record-congruent settings remains limited. Existing benchmarks often rely on static datasets or unstructured inputs that do not reflect the structured, interoperable data formats used in clinical systems. We introduce a pipeline for generating clinically realistic HL7 FHIR R4 bundles from unstructured text, enabling controllable evaluation of clinical decision support systems. The pipeline combines staged LLM generation with terminology-grounded validation and repair to reduce hallucinated codes and enforce structural and semantic consistency. Applying this approach to MedCaseReasoning, we construct MedCase-Structured, a synthetic dataset aligned with clinician-authored diagnostic cases, achieving valid FHIR generation for 82.5% of cases. Evaluation on MedCase-Structured reveals consistently lower diagnostic accuracy for LLMs on structured FHIR inputs than with plain text, highlighting the importance of deployment-aligned benchmarking.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

IntentScore: コンピュータ使用エージェントの意図条件付きアクションの評価

Computer-Use Agent (CUA) は、大規模な言語モデルを利用してデスクトップ環境で GUI 操作を実行しますが、アクションの品質を評価せずにアクションを生成するため、後続のステップに連鎖的に発生する不可逆的なエラーにつながります。私たちは、3 つのオペレーティング システムにわたる 398K のオフライン GUI インタラクション ステップから候補アクションをスコアリングすることを学習する、プランを認識した報酬モデルである IntentScore を提案します。 IntentScore は、状態とアクションの関連性に関する対照的な調整と、アクションの正しさに関するマージン ランキングという 2 つの相補的な目標を使用してトレーニングします。アーキテクチャ的には、各候補者の計画意図がアクション エンコーダーに埋め込まれ、同様のアクションを持つ候補者間で論理的根拠が異なるものを区別できるようになります。 IntentScore は、ホールドアウト評価で 97.5% のペア識別精度を達成します。トレーニング中にまったく見えない環境である OSWorld 上のエージェント S3 の再ランカーとしてデプロイされた IntentScore は、タスクの成功率を 6.9 ポイント向上させ、異種のオフライン軌跡から学習した報酬推定が、目に見えないエージェントとタスクの分布に一般化されることを示しています。

原文 (English)

IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents

Computer-Use Agents (CUAs) leverage large language models to execute GUI operations on desktop environments, yet they generate actions without evaluating action quality, leading to irreversible errors that cascade through subsequent steps. We propose IntentScore, a plan-aware reward model that learns to score candidate actions from 398K offline GUI interaction steps spanning three operating systems. IntentScore trains with two complementary objectives: contrastive alignment for state-action relevance and margin ranking for action correctness. Architecturally, it embeds each candidate's planning intent in the action encoder, enabling discrimination between candidates with similar actions but different rationales. IntentScore achieves 97.5% pairwise discrimination accuracy on held-out evaluation. Deployed as a re-ranker for Agent S3 on OSWorld, an environment entirely unseen during training, IntentScore improves task success rate by 6.9 points, demonstrating that reward estimation learned from heterogeneous offline trajectories generalizes to unseen agents and task distributions.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

速く考えると間違った考え: 直観力が政策評価における LLM 反事実的推論を調整する

大規模言語モデル (LLM) は、因果関係や反事実の推論にますます使用されていますが、現実世界の政策評価におけるその信頼性は依然として十分に解明されていません。私たちは、経済学と社会科学から抽出された 40 の実証的政策評価ケースのベンチマークを構築します。それぞれの事例は査読済みの証拠に基づいており、経験的発見が事前の一般的な期待と一致する (明白)、相対的に不明瞭である (曖昧)、または矛盾している (直観に反する) かどうかという直感によって分類されます。私たちは、8,000 件の実験トライアルを使用して 5 つのプロンプト戦略にわたって 4 つのフロンティア LLM を評価し、混合効果ロジスティック回帰を使用して結果を分析します。私たちの調査結果では、3 つの重要な結果が明らかになりました。(1) 思考連鎖 (CoT) のパラドックス。思考連鎖プロンプトは明らかなケースではパフォーマンスを劇的に向上させますが、直感に反するケースではこの利点が大幅に減弱します (相互作用 OR = 0.278、$p < 0.001$)。 (2) 支配的な要因としての直観性。ケースレベルの分散がモデルの選択またはプロンプト戦略の分散を超えています (ICC = 0.671)。 (3) 知識と推論の解離。引用に基づく精通度は精度と無関係です ($p = 0.84$)。これは、モデルが関連する知識を持っているものの、発見が直観と矛盾する場合、それを使って推論できないことを示唆しています。私たちはこれらの結果を二重プロセス理論(システム 1 対システム 2)のレンズを通して組み立て、現在の LLM の「遅い思考」は直観的事前の抑制を部分的にしか達成していない、つまりその本質を完全に実現することなく熟慮的推論の形式を生み出していると主張します。

原文 (English)

Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation

Large language models (LLMs) are increasingly used for causal and counterfactual reasoning, yet their reliability in real-world policy evaluation remains underexplored. We construct a benchmark of 40 empirical policy evaluation cases drawn from economics and social science, each grounded in peer-reviewed evidence and classified by intuitiveness -- whether the empirical finding aligns with (obvious), is unclear relative to (ambiguous), or contradicts (counter-intuitive) common prior expectations. We evaluate four frontier LLMs across five prompting strategies with 8,000 experimental trials and analyze the results using mixed-effects logistic regression. Our findings reveal three key results: (1) a chain-of-thought (CoT) paradox, where chain-of-thought prompting dramatically improves performance on obvious cases but this benefit is substantially attenuated on counter-intuitive ones (interaction OR = 0.278, $p < 0.001$); (2) intuitiveness as the dominant factor, with case-level variance exceeding that of model choice or prompting strategy (ICC = 0.671); and (3) a knowledge-reasoning dissociation, where citation-based familiarity is unrelated to accuracy ($p = 0.84$), suggesting models possess relevant knowledge but fail to reason with it when findings contradict intuition. We frame these results through the lens of dual-process theory (System 1 vs. System 2) and argue that current LLMs' "slow thinking" achieves only partial inhibition of intuitive priors -- producing the form of deliberative reasoning without fully delivering its substance.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SciHorizo​​n-DataEVA: 異種科学データの AI 対応性評価のためのエージェント システム

AI-for-Science (AI4Science) は、機械学習モデルをドメイン全体の予測、シミュレーション、仮説生成のワークフローに組み込むことにより、科学的発見をますます変革しています。ただし、これらのモデルの有効性は科学データの AI 対応度によって根本的に制約されており、現在、拡張性のある体系的な評価メカニズムが存在しません。この研究では、異種科学データのスケーラブルな AI 対応性評価のための新しいエージェント システムである SciHorizo​​n-DataEVA を提案します。評価基準レベルでは、AI への対応力をガバナンスの信頼性、データ品質、AI 互換性、科学的適応性という 4 つの補完的な側面に整理する Sci-TQA2 原則を導入します。各次元は測定可能な原子要素に分解され、きめ細かい実行可能な評価が可能になります。これらの原則を大規模に運用するために、私たちは、指示された循環ワークフローを通じて調整された階層型マルチエージェント評価アプローチである Sci-TQA2-Eval を開発しました。当社の Sci-TQA2-Eval は、軽量のデータセット プロファイリング、適用性を意識したメトリクスのアクティベーション、ドメイン制約とデータセット ペーパー シグナルに基づいた知識拡張計画を組み合わせることにより、データセットを意識した評価仕様を動的に構築します。これらの仕様は、組み込みの検証と自己修正を備えた適応型のツール中心の評価メカニズムを通じて実行され、異種の科学データ全体にわたってスケーラブルで信頼性の高い評価が可能になります。複数のドメインにわたる科学データセットに関する広範な実験により、原則に基づいた AI 対応性評価における SciHorizo​​n-DataEVA の有効性と汎用性が実証されました。

原文 (English)

SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data

AI-for-Science (AI4Science) is increasingly transforming scientific discovery by embedding machine learning models into prediction, simulation, and hypothesis generation workflows across domains. However, the effectiveness of these models is fundamentally constrained by the AI-readiness of scientific data, for which no scalable and systematic evaluation mechanism currently exists. In this work, we propose SciHorizon-DataEVA, a novel agentic system to scalable AI-readiness evaluation of heterogeneous scientific data. At the evaluation-criteria level, we introduce the Sci-TQA2 principles, which organize AI-readiness into four complementary dimensions: Governance Trustworthiness, Data Quality, AI Compatibility, and Scientific Adaptability. Each dimension is decomposed into measurable atomic elements that enable fine-grained and executable assessment. To operationalize these principles at scale, we develop Sci-TQA2-Eval, a hierarchical multi-agent evaluation approach orchestrated through a directed, cyclic workflow. Our Sci-TQA2-Eval dynamically constructs dataset-aware evaluation specifications by combining lightweight dataset profiling, applicability-aware metric activation, and knowledge-augmented planning grounded in domain constraints and dataset-paper signals. These specifications are executed through an adaptive, tool-centric evaluation mechanism with built-in verification and self-correction, enabling scalable and reliable assessment across heterogeneous scientific data. Extensive experiments on scientific datasets spanning multiple domains demonstrate the effectiveness and generality of SciHorizon-DataEVA for principled AI-readiness evaluation.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

CausaLab: AI 科学者向けのインタラクティブな因果発見のためのスケーラブルな環境

LLM エージェントによるインタラクティブな因果発見を評価するためのスケーラブルな環境である CausaLab を紹介します。以前の評価とは異なり、CausaLab では、エージェントが因果関係の証拠を使用して問題を解決できるかどうか、およびその答えが根底にある因果メカニズムに関する正しい仮説によって裏付けられているかどうかの両方を評価します。各エピソードではエージェントが合成実験室に配置されます。エージェントは以前の測定記録を受け取り、マニピュレーター結晶に介入し、同じ機構によって支配される保持されたリアクター結晶の共振周波数を予測します。隠されたデータ生成プロセスは、ランダムにサンプリングされた構造因果モデル (SCM) であるため、成功するには、事前の知識を思い出すのではなく、因果グラフと構造方程式の両方を回復する必要があります。 CausaLab には、エージェントの進化する SCM 仮説を記録するドメイン固有の言語も含まれており、軌跡を検査可能にしてグラウンド トゥルースと比較できるようになります。実験では、予測とメカニズム回復の間に永続的なギャップがあることが示されています。純粋に観測的な 6 ノード設定では、GPT-5.2-high はタスク精度 92% に達しますが、オールエッジ $F_1$ はわずか 0.471 です。この観察は、さまざまな相互作用戦略の探求をさらに動機づけます: 混合観察 - 介入戦略は構造忠実度を向上させます: 混合 6 ノード設定では、GPT-5.2-high はタスク精度とオールエッジ $F_1$ の両方で 80% を達成しました。しかし、純粋な介入戦略はタスクの精度とオールエッジ $F_1$ の両方においてパフォーマンスが低いため、強力なエージェントですら有益な介入を設計するのに苦労しています。私たちは、エージェントの主要な弱点として早期停止を特定し、仮説と過去のデータとの間の一貫性をモデルに検証するように依頼することが、この問題の軽減に役立つことを示します。したがって、CausaLab は予測の成功を因果関係の理解から切り離し、実験的因果推論者としての現在の LLM エージェントの限界を明らかにします。

原文 (English)

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates both whether an agent can solve a problem using causal evidence and whether its answer is grounded in a faithful recovered causal mechanism. Each episode places an agent in a synthetic laboratory: it receives prior measurement records, intervenes on a manipulator crystal, and predicts the resonance frequency of a held-out reactor crystal governed by the same mechanism. The hidden data-generating process is a randomly sampled structural causal model (SCM), so success requires recovering both a causal graph and structural equations rather than recalling prior knowledge. Experiments show a persistent gap between prediction and mechanism recovery: in the purely observational 6-node setting, GPT-5.2-high reaches 92% task accuracy but only 0.471 all-edge $F_1$. Mixed observation-intervention strategies improve structural fidelity, while pure intervention remains difficult even for strong agents. We identify premature stopping as a major weakness and show that consistency verification mitigates it. CausaLab therefore separates predictive success from causal understanding and exposes current LLM agents' limits as experimental causal reasoners.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

FundaPod: AI 支援のファンダメンタル投資調査のためのナレッジ グラフ メモリを備えたマルチペルソナ エージェント ポッド プラットフォーム

大規模言語モデル (LLM) は金融分野での適用が増えていますが、既存の研究のほとんどは取引シグナルや予測を中心とした財務 NLP タスクに重点を置いています。対照的に、制度的基礎研究では、人間のアナリストまたは AI エージェントが証拠を収集し、ビジネス推進要因を特定し、競合する視点を比較し、投資メモを作成する必要があります。その広範な目標は、単に結果を予測することではなく、投資知識の累積的な発展に貢献しながら、透明性、再利用可能、検証可能な投資計画を作成することです。 AI 支援のファンダメンタルズ投資調査のためのマルチペルソナ エージェント プラットフォームである FundaPod を紹介します。私たちは、基礎研究は人間中心の意思決定支援タスクであり、取引シグナルの生成とは質的に異なるため、独立性を維持するアーキテクチャの方が適していると主張します。 FundaPod では、バリュー投資家やマクロ戦略家など、さまざまなペルソナを持つ AI エージェントが、共有の出所契約に基づいて独立して調査を実施します。その後、彼らの意見の相違は、知識グラフ記憶システムを通じて人間のポートフォリオ マネージャー (PM) による裁定のために事後的に表面化されます。この論文は、設計科学の実践と認知的分離と人間と機械の協調の理論に基づいた、基礎研究をサポートする人間と AI のハイブリッド システムの 5 つの設計原則を提供します。また、4 つのアーキテクチャ メカニズムについても説明します。1 つは一般投資家の資料を展開可能なエージェントに変えるペルソナ蒸留パイプラインです。プランナーが型指定されたタスク グラフを導出できるようにする宣言型スキル レジストリ。メモの主張を検証可能な情報源に結び付ける根拠のある証拠モデル。そしてティッカー、メモ、アナリスト、テーマを結び付けるナレッジグラフ「第二の脳」。完全なケーススタディとペルソナベースのメモの比較を通じてアーキテクチャを実証します。

原文 (English)

FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research

Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on prediction. Institutional fundamental research, by contrast, requires human analysts or AI agents to gather evidence, identify business drivers, compare competing viewpoints, and generate investment memos. Its broader goal is not merely to predict outcomes, but to produce investment plans that are transparent, reusable, and verifiable, while contributing to the cumulative development of investment knowledge. We present FundaPod, a multi-persona agent platform for AI-assisted fundamental investment research. We argue that fundamental research is a human-centric decision-support task that is qualitatively distinct from trading-signal generation, and is therefore better served by an independence-preserving architecture. In FundaPod, AI agents with different personas, such as value investors or macro strategists, conduct research independently under a shared provenance contract. Their disagreements are then surfaced post hoc for adjudication by the human portfolio manager (PM) through a knowledge-graph memory system. This paper contributes five design principles for human-AI hybrid systems supporting fundamental research, grounded in design-science practice and theories of cognitive isolation and human-machine coordination. It also describes four architectural mechanisms: a persona distillation pipeline that turns public investor materials into deployable agents; a declarative skill registry that lets the planner derive typed task graphs; a grounded evidence model that links memo claims to verifiable sources; and a knowledge-graph "second brain" that connects tickers, memos, analysts, and themes. We demonstrate the architecture through a complete case study and a persona-based memo comparison.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

統計的に真剣であることの重要性: GSM シンボリックの重要な再評価

GSM-Symbolic ベンチマーク (Mirzadeh et al., 2025) は、GSM8K 問題のテンプレート生成バリアントでテストした場合、25 の大規模言語モデル (LLM) 全体で一貫したパフォーマンスの低下を報告し、モデルには真の推論機能が欠けていると結論付けました。私たちは、この結論は不安定な統計的根拠に基づいていると主張します。質問ごとの変量効果を備えた一般化線形混合モデルを使用して 20 の無重みモデルを再評価すると、元のプロンプト形式で統計的に有意なパフォーマンス変化を示したのは半分だけであることがわかりました。さらに、これまで認められていなかった要因も特定しました。つまり、メインの GSM-Symbolic データセットには、GSM-Base と比較して問題テキスト内のより大きな整数の体系的にシフトされた分布が含まれており (K-S 統計量 = 0.12、p < 0.001)、これは原著者の主張と矛盾しています。この大きな数の効果を制御することは、残りのケースの約半分で重要性を説明します。統計的に有意なパフォーマンスデルタを持つモデルの中で、変数結合の脆弱性、算術的制限、デュアルタスク干渉など、モデル固有の明確な障害プロファイルを特定しました。これは、LLM 推論に関する包括的な主張が統計的に時期尚早であり、機構的に誤解を招くものであることを強調します。

原文 (English)

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-generated variants of GSM8K problems, concluding that the models lack genuine reasoning capabilities. We argue that this conclusion rests on shaky statistical ground. Re-evaluating 20 open-weight models using Generalised Linear Mixed Models with per-question random effects, we find that only half exhibit statistically significant performance changes under the original prompt format. Moreover, we identify a previously unacknowledged factor: the main GSM-Symbolic dataset contains a systematically shifted distribution of larger integers in problem texts relative to GSM-Base (K-S statistic = 0.12, p < 0.001), contradicting the original authors' claims. Controlling for this large number effect accounts for significance in roughly half the remaining cases. Among models with statistically significant performance deltas, we identify distinct, model-specific failure profiles - including fragility of variable binding, arithmetic limitations, and dual-task interference - underscoring that blanket claims about LLM reasoning are both statistically premature and mechanistically misleading.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ルーブリックから信頼できるスコアまで: LLM 審査員による証拠に基づいたテキスト評価

ルーブリックベースのテキスト評価では、スケーラブルな審査員として大規模言語モデル (LLM) がますます使用されていますが、凍結されたブラックボックス モデルを人間の採点基準に合わせるのは依然として困難です。私たちは、この課題を基準移行問題として定式化します。目標は、単に LLM にスコアの割り当てを促すことではなく、人間のルーブリックの意図を、安定した、監査可能な、人間に合わせたスコアリング プロトコルに移行することです。私たちは、LLM ベースのルーブリック スコアリングで繰り返される 3 つの失敗モード、つまりルーブリック実行ドリフト、検証不可能なスコアの帰属、人間スケールの不整合を特定しました。これらの障害モードに対処するために、信頼できる証拠に基づいたルーブリック ベースのテキスト評価のための 3 段階の推論時間フレームワークである Rulers を導入します。ルーラーはまず人間によるルーブリックをロックされたタスクレベルの仕様に変換し、次に構造化されたチェックリストの決定、型付けされた証拠の根拠付け、該当する場合には抽出的な引用の検証を伴う仕様を実行し、最後に事後キャリブレーションを適用してモデル由来の信号を人間のスコア境界と一致させます。 Rulers は、エッセイの採点、要約評価、EFL ライティングの評価、構造化入力テキストの生成をカバーする 4 つのルーブリックに基づいたベンチマーク全体で、複数の凍結されたバックボーン モデルにわたるほとんどの評価設定で人間によるスコアのより強い一致を達成しています。さらに分析を進めると、ルーラーは人間の経験的なスコア分布によりよく一致し、意味的に同等のルーブリック摂動下での安定性が向上し、その 3 つのコンポーネントそれぞれから利点が得られることが示されています。これらの結果は、信頼できる LLM 判定には、即時の表現だけではなく、固定基準、追跡可能な証拠、および調整されたスコア解釈が必要であることを示唆しています。私たちのコードは https://anonymous.4open.science/r/Rulers_0525-3328 で入手できます。

原文 (English)

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challenging. We formulate this challenge as a criteria-transfer problem: the goal is not merely to prompt an LLM to assign a score, but to transfer human rubric intent into a stable, auditable, and human-aligned scoring protocol. We identify three recurring failure modes in LLM-based rubric scoring: rubric execution drift, unverifiable score attribution, and human-scale misalignment. To address these failure modes, we introduce Rulers, a three-stage inference-time framework for reliable, evidence-grounded rubric-based text evaluation. Rulers first converts a human rubric into a locked task-level specification, then executes the specification with structured checklist decisions, typed evidence grounding, and extractive quote verification when applicable, and finally applies post-hoc calibration to align model-derived signals with human score boundaries. Across four rubric-governed benchmarks covering essay scoring, summarization assessment, EFL writing evaluation, and structured-input text generation, Rulers achieves stronger human-score agreement in most evaluated settings across multiple frozen backbone models. Further analyses show that Rulers better matches empirical human score distributions, improves stability under semantically equivalent rubric perturbations, and benefits from each of its three components. These results suggest that reliable LLM judging requires fixed criteria, traceable evidence, and calibrated score interpretation rather than prompt phrasing alone. Our code is available at https://anonymous.4open.science/r/Rulers_0525-3328.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

GICDM: 信頼性の高い距離ベースの生成モデル評価のためのハブネスの軽減

生成モデルの評価は通常、高次元の埋め込み空間に依存してサンプル間の距離を計算します。これらの空間のデータセット表現は、最近隣関係を歪め、距離ベースのメトリクスを偏らせるハブネス現象の影響を受けることを示します。古典的な反復コンテキスト相違測定 (ICDM) に基づいて、実際のデータと生成されたデータの両方の近傍推定を修正する方法である生成 ICDM (GICDM) を導入します。経験的な動作を改善するためにマルチスケール拡張を導入します。合成ベンチマークと実際のベンチマークに関する広範な実験により、GICDM がハブネスに起因する障害を解決し、信頼性の高いメトリック動作を復元し、人間の評価との整合性が向上することが実証されています。

原文 (English)

GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation

Generative model evaluation commonly relies on high-dimensional embedding spaces to compute distances between samples. We show that dataset representations in these spaces are affected by the hubness phenomenon, which distorts nearest-neighbor relationships and biases distance-based metrics. Building on the classical Iterative Contextual Dissimilarity Measure (ICDM), we introduce Generative ICDM (GICDM), a method to correct neighborhood estimation for both real and generated data. We introduce a multi-scale extension to improve empirical behavior. Extensive experiments on synthetic and real benchmarks demonstrate that GICDM resolves hubness-induced failures, restores reliable metric behavior, and improves alignment with human assessment.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

P$^2$RAG: 任意の上位 $k$ 取得をサポートする効率的なプライバシー保護 RAG サービス

検索拡張生成 (RAG) を使用すると、大規模な言語モデルで外部の知識を使用できるようになりますが、RAG サービスをアウトソーシングすると、データ所有者とユーザーの両方にプライバシー上の懸念が生じます。プライバシーを保護する RAG システムは、安全な上位 $k$ 取得を実行することでこれらの懸念に対処します。これは通常、関連するドキュメントを識別するための安全な並べ替えを使用して実装されます。しかし、既存のシステムは、$k$ を変更できないこと、新たなセキュリティの問題、特に $k$ が大きい場合の効率の低下などにより、任意の $k$ をサポートするという課題に直面しています。金融、法律、医療などのアプリケーションでは、既存のシステムに多大なオーバーヘッドを引き起こすほど大きな $k$ が必要となるため、これは重大な制限です。また、最新のロングコンテキスト モデルは一般に、より大きな検索セットを使用することでより高い精度を実現します。我々は、任意の上位 $k$ 検索をサポートする効率的なプライバシー保護 RAG サービスである P$^2$RAG を提案します。既存のシステムとは異なり、P$^2$RAG は候補ドキュメントのソートを回避します。代わりに、インタラクティブな二分法を使用して、上位 $k$ ドキュメントのセットを決定します。セキュリティのために、P$^2$RAG は 2 台の半誠実で非共謀のサーバー上で秘密共有を使用して、データ所有者のデータベースとユーザーのプロンプトを保護します。悪意のあるユーザーから防御するために制限と検証を強制し、データベースの情報漏洩を厳しく制限します。実験では、$k = 16$--$1024$ の場合、P$^2$RAG は最先端の PRAG より 3--300$\times$ 高速であることが示されています。

原文 (English)

P$^2$RAG: Efficient Privacy-Preserving RAG Service Supporting Arbitrary Top-$k$ Retrieval

Retrieval-Augmented Generation (RAG) enables large language models to use external knowledge, but outsourcing the RAG service raises privacy concerns for both data owners and users. Privacy-preserving RAG systems address these concerns by performing secure top-$k$ retrieval, which is typically implemented using secure sorting to identify relevant documents. However, existing systems face challenges supporting arbitrary $k$ due to their inability to change $k$, new security issues, and in particular, efficiency degradation with large $k$. This is a significant limitation because applications such as finance, law, and healthcare require a $k$ that is large enough to cause huge overhead for existing systems. Also, modern long-context models generally achieve higher accuracy with larger retrieval sets. We propose P$^2$RAG, an efficient privacy-preserving RAG service that supports arbitrary top-$k$ retrieval. Unlike existing systems, P$^2$RAG avoids sorting candidate documents. Instead, it uses an interactive bisection method to determine the set of top-$k$ documents. For security, P$^2$RAG uses secret sharing on two semi-honest non-colluding servers to protect the data owner's database and the user's prompt. It enforces restrictions and verification to defend against malicious users and tightly bounds the information leakage of the database. The experiments show that P$^2$RAG is 3--300$\times$ faster than the state-of-the-art PRAG for $k = 16$--$1024$.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

AgentLens: SWE エージェント評価におけるラッキー パスの問題を明らかにする

更新された要約は次のとおりです。 ソフトウェア エンジニアリング (SWE) エージェントの評価は、最終パッチがテストに合格するかどうかという 2 つの信号によって支配されます。この結果のみの考え方は、原則に基づいた解決策と混沌とした試行錯誤のプロセスを同等のものとして扱います。この等価性は経験的に誤りであることを示します。 60 の SWE ベンチ検証済みタスクで 8 つのモデル バックエンドからの 2,614 の OpenHands 軌跡を評価します。これらのうち、47 にはタスク レベルのプロセス参照を構築するのに十分な通過軌跡があり、1,815 の軌跡評価サブセットが得られます。このサブセットの通過軌跡のうち、10.7% は、回帰サイクル、ブラインド再試行、検証の欠落、または時間的に無秩序な探索、実装、検証など、ラッキー パスと呼ばれる動作を示しています。 SWE エージェントの軌跡をプロセスレベルで評価するためのフレームワークである AgentLens を導入し、品質スコア、廃棄信号、分岐点、および 47 のタスクレベルのプレフィックス ツリー アクセプタ (PTA) 参照が注釈付けされた 1,815 の軌跡のデータセットである AgentLens-Bench を定義します。 AgentLens は、同じタスクに渡された複数のソリューションをマージすることで PTA 参照を構築し、コンテキスト依存のインテント ラベラーを使用して、ツール ID だけではなく軌跡履歴に基づいて探索、実装、検証、またはオーケストレーションにアクションを割り当てます。 AgentLens-Bench では、品質スコアによってパスの軌跡が Lucky、Solid、Ideal の層に分割され、さらに Lucky パスが 5 つの反復メカニズムに分解されます。 8 つのモデル バックエンド全体で、Lucky 率の範囲は 0.5% ~ 23.2% であり、一部のモデルは、合格率ではなく品質スコアでランク付けすると、最大 5 ランク順位が変動します。 AgentLens-Bench アーティファクト、AgentLens SDK、分析ツールを含むプロジェクト リポジトリを間もなくリリースする予定です。

原文 (English)

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

Here is the updated abstract: Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view treats a principled solution and a chaotic trial-and-error process as equivalent. We show that this equivalence is empirically false. We evaluate 2,614 OpenHands trajectories from eight model backends on 60 SWE-bench Verified tasks. Of these, 47 have enough passing trajectories to construct task-level process references, yielding a 1,815-trajectory evaluation subset. Among passing trajectories in this subset, 10.7% exhibit behavior we call a Lucky Pass: regression cycles, blind retries, missing verification, or temporally disordered exploration, implementation, and verification. We introduce AgentLens, a framework for process-level assessment of SWE-agent trajectories, and define AgentLens-Bench, a dataset of 1,815 trajectories annotated with quality scores, waste signals, divergence points, and 47 task-level Prefix Tree Acceptor (PTA) references. AgentLens builds PTA references by merging multiple passing solutions for the same task, and uses a context-sensitive intent labeler to assign actions to Exploration, Implementation, Verification, or Orchestration based on trajectory history rather than tool identity alone. On AgentLens-Bench, the quality score separates passing trajectories into Lucky, Solid, and Ideal tiers and further decomposes Lucky Passes into five recurring mechanisms. Across the eight model backends, Lucky rates range from 0.5% to 23.2%, and some models move by as many as five rank positions when ranked by quality score instead of pass rate. We plan to release the project repository soon, including AgentLens-Bench artifacts, the AgentLens SDK, and the analysis tooling.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

JMed48k: 視覚言語モデル評価のための多職種の日本の医師免許ベンチマーク

視覚言語モデルを評価するための、多職種の日本の医療ライセンスベンチマークである JMed48k を紹介します。日本の厚生労働省がリリースした公式 PDF 資料から構築された JMed48k には、2005 年から 2025 年までの 11 の国家免許試験からの 48,862 件の試験問題と 20,142 枚の画像が含まれており、ビジュアル コンテンツは 8 種類の分類法に基づいて注釈が付けられています。このコーパスから、9,905 のテキストのみの質問と 2,579 の画像付きの質問を含む 12,484 のスコア付き質問を含む、最近 5 年間の評価サブセットである JMed48k-Eval を導き出します。私たちは 21 の独自のオープンソースの医療特化モデルを評価し、テキストのみのパフォーマンスと画像付きのパフォーマンスを別々に報告します。これらのサブセットには異なる質問が含まれているため、視覚コンテンツを削除する前後の画像を含む質問を評価するペアの画像削除監査をさらに導入して、4 つの回答遷移状態を調査します。監査では、独自のオープンソース モデルが画像から大幅に利益を得ているのに対し、医療に特化したシステムでは目に見える視覚的証拠の使用が限られており、画像の削除後も多くの正解が残っていることが示されています。独自モデルの中でも、正味の画像除去効果は、医師の質問で +5.7 ポイントから保健師の質問で +39.8 ポイントまで、職種によって 7 倍のばらつきがあります。私たちは、医療ライセンス設定における視覚言語モデルの再現可能な専門職層別評価をサポートするために、JMed48k をリリースします。

原文 (English)

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese Ministry of Health, Labour and Welfare, JMed48k contains 48,862 exam questions and 20,142 images from 11 national licensing examinations between 2005 and 2025, with visual content annotated under an 8-type taxonomy. From this corpus, we derive JMed48k-Eval, a recent five-year evaluation subset with 12,484 scored questions, including 9,905 text-only questions and 2,579 questions with images. We evaluate 21 proprietary, open-source, and medical-specific models, reporting text-only and with-image performance separately. Because these subsets contain different questions, we further introduce a paired image-removal audit that evaluates questions with images before and after removing visual content to explore four answer-transition states. The audit shows that proprietary and open source models gain substantially from images, whereas medical-specific systems show limited observable use of visual evidence, with many correct answers persisting after image removal. Even among proprietary models, the net image-removal effect varies sevenfold across professions, from +5.7 points on Physician questions to +39.8 points on Public Health Nurse questions. We release JMed48k to support reproducible, profession-stratified evaluation of vision-language models in medical licensing settings.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

多様な呼吸不全予測の前向き評価: 胸部 X 線撮影は EHR 信号を超えてパフォーマンスを向上させますか?

呼吸不全の早期予測は、集中治療室でのタイムリーな臨床介入にとって重要です。既存の電子健康記録 (EHR) ベースのモデルは、生理学的悪化を継続的に監視できますが、胸部 X 線写真 (CXR) に反映される肺の病態生理学を完全には捕捉できない可能性があります。この研究では、CXR 情報が EHR 信号のみを超えて侵襲的人工呼吸器の前向き予測を改善するかどうかを尋ねます。私たちは、構造化された EHR 時系列データと CXR 基盤モデル表現を統合するゲート付きマルチモーダル フレームワークを開発します。ゲーティング モジュールは、患者固有の臨床状況に基づいてイメージング機能の寄与を適応的に制御し、モデルが有益な場合にイメージング情報に選択的に依存できるようにします。私たちは、ICU患者の24時間以内の侵襲的人工呼吸器を予測するためのフレームワークを前向きに評価し、確立されたEHR専用モデル(Vent.io)、一致する臨床時点で得られた医師の予測、および代替のマルチモーダルバリアントと比較します。ゲート付きマルチモーダル モデルは、EHR のみのベースラインよりも高い識別を達成し、Vent.io の 0.752 と比較して、REMEDIS および MedInsight CXR 表現を使用した AUROC 値はそれぞれ 0.860 と 0.858 でした。医師の予測と比較して、マルチモーダルフレームワークは良好な特異性を維持しながら感度を大幅に向上させました。 EHR のみのモデルと比較して、マルチモーダル統合により特異性と陽性的中率が向上したことは、CXR 情報が選択された患者のリスク推定を精緻化できることを示唆しています。これらの発見は、画像処理を呼吸不全の予測に組み込むための実用的な戦略として、適応型マルチモーダル融合を裏付けるものです。

原文 (English)

Prospective evaluation of multimodal respiratory failure prediction: Do chest X-rays improve performance beyond EHR signals?

Early prediction of respiratory failure is critical for timely clinical intervention in intensive care units. Existing electronic health record (EHR)-based models can continuously monitor physiologic deterioration, but they may not fully capture pulmonary pathophysiology reflected in chest radiographs (CXRs). In this study, we ask whether CXR information improves prospective prediction of invasive mechanical ventilation beyond EHR signals alone. We develop a gated multimodal framework that integrates structured EHR time-series data with CXR foundation-model representations. The gating module adaptively controls the contribution of imaging features based on patient-specific clinical context, allowing the model to selectively rely on imaging information when it is informative. We prospectively evaluate the framework for predicting invasive mechanical ventilation within 24 hours in ICU patients and compare it with an established EHR-only model (Ventio), physician predictions obtained at matched clinical time points, and alternative multimodal variants. The gated multimodal models achieved higher discrimination than the EHR-only baseline, with AUROC values of 0.860 and 0.858 using REMEDIS and MedInsight CXR representations, respectively, compared with 0.752 for Ventio. Relative to physician predictions, the multimodal framework substantially improved sensitivity while maintaining favorable specificity. Compared with the EHR-only model, multimodal integration increased specificity and positive predictive value, suggesting that CXR information can refine risk estimation in selected patients. These findings support adaptive multimodal fusion as a practical strategy for incorporating imaging into prospective respiratory failure prediction.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

アライメントフロア: ペルソナのカスタマイズが安全な場合

多元的 AI の主な約束は行動の適応です。「創造的であること」や「徹底していること」などのペルソナ プロンプトにより、システムは多様なユーザーの価値観とコミュニケーション スタイルを尊重できます。しかし、アライメントが崩れる前に、モデルはどれだけのカスタマイズを吸収できるでしょうか?我々は、アライメントとカスタマイズのトレードオフに関する最初の対照研究を発表し、異なるアライメント強度を持つ 2 つのモデルで 5 つのタスクにわたって 7 つのペルソナ条件をテストしました (1,800 回の実行)。私たちは調整フロアを発見しました。強く調整されたモデル (クロード ソネット) では、ペルソナ プロンプトはお調子者にまったく影響しません。すべての条件で最大 15% が生成され、豊富なパーソナライゼーションが安全な安定したプラットフォームです。弱く調整されたモデル (Nova Lite) では、同じペルソナがお調子者を 5% から 50% に変更します。つまり、フロアが存在せず、カスタマイズが安全上の責任となります。驚くべきことに、協調性は最悪の犯罪者ではありません。外向性 (+20pp) とオープンネス (+15pp) は、より大きな劣化を引き起こします。建設的な発見は懐疑的防御です。批判的思考を持つペルソナは、弱いモデルでもお調子者を 5% に減少させます。これは、この研究における唯一最大の効果です。モデル間のペルソナ効果の伝達はほぼゼロ ($\rho = 0.006$) であり、アライメント テストはモデルごとに行う必要があることを意味します。私たちは、設計原則としてアライメント フロアを提案します。つまり、ペルソナのカスタマイズを展開する前にアライメント フロアを測定し、安全性を重視したペルソナをユーザー向けのペルソナの下に重ねて、配置を損なうことなくパーソナライゼーションを可能にします。

原文 (English)

The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs

Telling an LLM to "be enthusiastic" raises its sycophancy rate from 30\% to 50\% on a lightly-aligned model, but has zero effect on a strongly-aligned one. We define this gap as the alignment floor, $\Delta_{\text{floor}}(m)=\max_pS(m,p)-\min_pS(m,p)$, the range of sycophancy rates a model produces across persona conditions, and treat sycophancy as a persona-conditional property rather than a fixed model property. Pluralistic AI relies on behavioral adaptation via persona prompts like "be creative" or "be thorough", which let systems respect diverse user values and communication styles; the safety question is how much customization a given model can absorb before its truthfulness shifts. We present a controlled case study contrasting a strongly-aligned RLHF + Constitutional-AI model (Claude Sonnet 4.6) with a more lightly-aligned model (Amazon Nova Lite), spanning seven persona conditions and five tasks for 1800 total runs. An existence-pair result motivates per-model auditing: there is at least one strongly-aligned model with $\Delta_{\text{floor}}=5$pp (within 5pp of the 15\% control rate) and at least one lightly-aligned model with 45pp (5\%--50\% range). On the lightly-aligned model, all five Big Five personas increase sycophancy over control, and counterintuitively Agreeableness produces the smallest increase, not the largest. The single largest effect in the study is constructive: a Skeptic persona reduces sycophancy by 25pp on the lightly-aligned model, and is the only persona that instructs resistance against user claims rather than engagement with them, suggesting a directionality account. Cross-model transfer of persona effects is near-zero, so persona-alignment testing must be per-model. We propose $\Delta_{\text{floor}}$ as a deployment-time audit metric: measure it on a small persona panel before deploying persona customization.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

モデルが一致しない場合: パブリック コメント分析のための LLM 評価を再考する

連邦政府機関はパブリック コメント コーパスを分類するために大規模言語モデル (LLM) を導入しており、モデルの記録構成によって政策立案者が何を確認し、どの議論が登録されるかが決まります。小規模な検証済みセットに対するスタンスの精度に基づいた標準評価では、異なるモデルが同じ公的入力に対して実質的に異なる分類を生成する場合を検出できません。私たちは、マルチモデルの不一致を解釈の複雑さの診断として扱い、真に曖昧な公的意見に向けて人間によるレビューを指示する解釈監査パイプラインを提案します。 4 つの LLM にわたる連邦 USDA 文書に対する 1,260 件のパブリック コメントを分析したところ、モデル間のテーマの相違がモデル内のプロンプト変動を上回っており、専門家のルーブリックが深い解釈上の不一致を解決することなく抑圧していることがわかりました。層化された 40 コメントのサブサンプルに対する 2 段階のラベル付け研究では、4 人の LLM とヒューマン アノテーターが独立してラベル付けし、他のラベルを確認した後に修正しました。改訂動作はラベラーによって異なり、ヒューマン・アノテーターの改訂では、アンサンブルの集合的な出力にはないフレームが頻繁に導入されました。私たちは、不一致に基づく評価は、LLM 支援解釈コーディングの精度メトリクスを補完するために必要であると主張します。

原文 (English)

When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis

Federal agencies are deploying large language models (LLMs) to categorize public comment corpora, where the model's organization of the record shapes what policymakers see and which arguments register. Standard evaluation, anchored on stance accuracy against a small validated set, cannot detect when different models produce materially different categorizations of the same public input. We propose an Interpretive Audit Pipeline that treats multi-model disagreement as diagnostic of interpretive complexity and directs human review toward genuinely ambiguous public input. Analyzing 1,260 public comments on a federal USDA docket across four LLMs, we find that inter-model thematic divergence exceeds within-model prompt variation, and that an expert rubric suppresses deep interpretive disagreement without resolving it. In a two-stage labeling study on a stratified 40-comment subsample, four LLMs and a human annotator labeled independently and then revised after seeing the others' labels. Revision behavior varied across labelers, and the human annotator's revisions frequently introduced framings absent from the ensemble's collective output. We argue disagreement-based evaluation is a necessary complement to accuracy metrics for LLM-assisted interpretive coding.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

ペーパーエージェント、ペーパーゲイン:DeFi投資エージェントの実証分析

自律的なオンチェーン取引に AI を使用するシステムである DeFi 投資エージェントは、2024 年後半以来、合計トークン評価額で 30 億米ドルを超えています。私たちは 1,900 以上の AI タグ付き暗号プロジェクトを調査し、投資中心のエージェントに絞り込み、戦略と可観測性の側面にわたる 10 の代表的なプロジェクトを厳選しています。次に、ElizaOS と Virtuals Protocol という 2 つの著名なエージェント フレームワークの詳細なアーキテクチャ分析と、925,323 人のトークン所有者を対象とする公的に起因する取引活動を伴う 11 の Solana ベースのエージェント トレジャリーの定量的なオンチェーン パフォーマンス分析を実施します。現在のデプロイメントは初期段階で異種混合のままであることがわかりました。(1) 私たちのサンプルでは、​​多くのプロジェクトが自律的な取引実行の明確な証拠をまだ提供しておらず、開発者のインタビューでは、目に見えるデプロイメントの多くが基本的な API 統合のままであることが示唆されています。 (2) エージェントの財務省は 3,000 万米ドルを超える紙の利益を保持している一方、トークン所有者は合計で 1 億 9,170 万米ドルを損失しており、ウォレットの上位 1% が全利益の 81.4% (18 億 1,000 万米ドル) を獲得しています。 (3) トークンの評価額は財務省のファンダメンタルズとの関連が弱く、時価総額対AUMの比率は10,000倍を超えていますが、確立されたDeFiプロトコルでは1倍未満です。 (4) ユーザーの総利益は 24 億米ドルでピークに達し、その後純損失に減少し、収益の中央値はすべてのプラットフォームでマイナスとなり、トークンは史上最高値から平均して 93% 減少しました。私たちは、これらの結果を、オープンインフラストラクチャにより迅速な実験が可能になるだけでなく、自律性、パフォーマンス、および利害関係者の連携のための堅牢な標準が出現する前に、単純なエージェントや投機的なエージェントが立ち上がることを可能にする、パーミッションレスの第一世代市場の特徴であると解釈します。そこで私たちは、現在の展開と将来の投資グレードのエージェント システムとの間のギャップを特徴付けるために、自律的な実行、リスク調整後の収益性、利害関係者の連携という 3 つの側面に沿った成熟度フレームワークを提案します。

原文 (English)

Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents

DeFi investment agents, systems that use AI for autonomous on-chain trading, have attained over USD 3 billion in combined token valuations since late 2024. We survey over 1,900 AI-tagged crypto projects, filter to investment-focused agents, and curate 10 representative projects spanning strategy and observability dimensions. We then conduct a deep-dive architectural analysis of two prominent agent frameworks, ElizaOS and Virtuals Protocol, and a quantitative on-chain performance analysis of 11 Solana-based agent treasuries with publicly attributable trading activity, covering 925,323 token holders. We find that current deployments remain early and heterogeneous: (1) in our sample, many projects do not yet provide clear evidence of autonomous trade execution, and developer interviews suggest that many visible deployments remain basic API integrations; (2) agent treasuries retain over USD 30M in paper gains while token holders collectively lost USD 191.7M, with the top 1% of wallets capturing 81.4% of all gains (USD 1.81B); (3) token valuations are weakly connected to treasury fundamentals, with market-cap-to-AUM ratios exceeding 10,000x versus below 1x for established DeFi protocols; and (4) aggregate user gains peaked at USD 2.4B before declining to net losses, with median returns negative on every platform and tokens declining 93% on average from all-time highs. We interpret these outcomes as characteristic of a permissionless, first-generation market in which open infrastructure enables rapid experimentation but also allows naive or speculative agents to launch before robust standards for autonomy, performance, and stakeholder alignment emerge. We therefore propose a maturity framework along three dimensions: autonomous execution, risk-adjusted profitability, and stakeholder alignment, to characterize the gap between current deployments and future investment-grade agent systems.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

BenchTrace: LLM エージェントのリフレクション能力と制御された進化をテストするためのベンチマーク

自己進化エージェントは過去の失敗を反映することで時間の経過とともに改善しますが、既存の評価には 2 つの点で制限があります。1 つはタスク スコアのみを測定し、反映品質は不明のままにすること、もう 1 つはエージェント自身のエピソードの実行に依存しており、特定の失敗パターンを対象にするメカニズムを提供していないことです。 LLM エージェントの自己進化能力を評価するためのベンチマークである \textbf{BenchTrace} を紹介します。 BenchTrace は、6 つの多様なタスクにわたる 1,821 の注釈付きエピソードのスナップショット反映データセットに基づいて構築されており、ターゲットを絞った QA タスクを通じて障害の特定を調査する \textbf{反映評価} と、制御された自己進化シミュレーションで過去の障害経験が回避行動に変換されるかどうかをテストする \textbf{進化評価} で構成されます。 BenchTrace に基づいて、エージェントがターゲットの障害インスタンスを回避できたテスト ケースの割合を測定する新しい評価指標である \textbf{障害回避率 (FAR)} を提案します。 Qwen3-32B と GPT-4.1 を使った実験では、どちらのモデルもリフレクション評価でエンドツーエンドの合格率が 30\% を下回り、診断が主なボトルネックであることが明らかになりました。進化の評価では、自己進化手法は一般に非進化ベースラインよりもFARを改善しますが、エージェントはノイズエピソードが蓄積するにつれて初期のレッスンを忘れ、エージェントは特定のコンテキストを超えて反省を一般化することができず、タスクコンテキスト間で負の転移を引き起こすことが示されています。さらに、相関分析により、完全に正しい反射のみが高い FAR と強く関連していることが明らかになりました。 BenchTrace は、現在の自己進化アプローチの具体的な限界を明らかにし、対象を絞った評価のための制御されたモデルに依存しないフレームワークを提供します。

原文 (English)

BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

Self-evolving agents improve over time by reflecting on past failures, but existing evaluation is limited in two ways: it measures only task scores, leaving reflection quality unknown, and it relies on agents' own episode runs, offering no mechanism to target specific failure patterns. We present \textbf{BenchTrace}, a benchmark for evaluating self-evolution ability in LLM agents. BenchTrace is built on a snapshot-reflection dataset of 1,821 annotated episodes spanning six diverse tasks, and comprises a \textbf{Reflection Evaluation} that probes failure identification through targeted QA tasks, and an \textbf{Evolution Evaluation} that tests whether past failure experience translates into avoidance behavior in a controlled self-evolution simulation. Building on BenchTrace, we propose \textbf{failure avoidance rate (FAR)}, a new evaluation metric measuring the fraction of test cases in which the agent successfully avoids the target failure instance. Experiments with Qwen3-32B and GPT-4.1 reveal that both models fall below a 30\% end-to-end pass rate on reflection evaluation, with diagnosis as the primary bottleneck. Evolution evaluation shows that self-evolution methods generally improve FAR over the non-evolving baseline, but agents forget early lessons as noise episodes accumulate, and agents fail to generalize their reflections beyond the specific context, causing negative transfer across task contexts. Our correlation analysis further reveals that only a fully correct reflection is strongly associated with higher FAR. BenchTrace exposes concrete limits of current self-evolution approaches and provides a controlled, model-agnostic framework for targeted evaluation.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

文献検索の評価を再考する: 深い調査は役に立ちますが、人間の引用リストは根拠のある真実ではありません

私たちは、検索パイプラインの改善と評価対象としての人による参照リストのストレステストという 2 つの相補的な角度から大規模な文献検索を研究しています。まず、完全なクエリ論文を処理し、取得した結果を文献目録に沿って幅優先で拡張する Deep Research パイプラインを実装します。このパイプラインが通常の API のみの検索を大幅に上回り、RollingEval-Jun25 (論文 250 件の文献検索ベンチマーク) の再現率が 20% 未満から 80% 以上に上昇することを示します。 2 番目に、中立的な LLM を判断者として使用して、人間の参照がタスクに対する健全な根拠であるかどうかを判断します。私たちは重大な限界を発見しました。人間による引用のうち、中等度以上の関連性があると判断されたのは 51% のみであったのに対し、最も強力な AI ベースの再ランカーでは 86 ~ 88% でした。 OpenAlex の共著グラフでこのギャップを調査したところ、人間は AI の再ランク付けを行う最も優れた人よりも直接の協力者を引用する可能性が 2.5 倍高いことがわかりました。まとめると、我々の結果は単一軸の文献検索評価に反対している。つまり、想起率、話題関連性スコアリング、ランクリストの多様性、および共著距離診断は、それぞれ引用の質の相補的な特性を測定するものであり、併せて報告されるべきである。

原文 (English)

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as an evaluation target. First, we implement a Deep Research pipeline that processes the full query paper and expands the retrieved results breadth-first along their bibliographies, and show that it substantially outperforms vanilla API-only search, raising recall on RollingEval-Jun25 (a 250-paper literature-search benchmark) from below 20% to above 80%. Second, we use a neutral LLM-as-a-judge to determine if human references are sound ground truth for the task. We find significant limitations: only 51% of human citations are judged moderately relevant or higher, against 86--88% for the strongest AI-based re-rankers. We study this gap on the OpenAlex co-authorship graph, finding that humans are 2.5x more likely than the best AI re-rankers to cite a direct collaborator. Together, our results argue against single-axis literature-search evaluation: recall, topical-relevance scoring, ranked-list diversity, and a co-authorship-distance diagnostic each measure complementary properties of citation quality and should be reported jointly.

2026-05-29 13:00 JSTarXiv cs.AIロボティクスビジネス/資金調達

MiraBench: ロボット世界モデルにおける動作条件付き信頼性の評価

アクション条件付き世界モデルは、ロボット学習用のスケーラブルなシミュレーターとしてますます使用されていますが、現在の評価では、条件付けされたアクションの下でその予測が信頼できるという限られた証拠が提供されています。既存のベンチマークは主に視覚的な忠実度を重視しており、予測される未来が物理的に妥当であるか、命令されたアクションに忠実であるか、アクションが成功しないはずのときに失敗するように調整されているかどうかが不明確なままです。 \emph{動作条件付き信頼性} をロボット世界モデルの中核的な評価目標として定義する階層型ベンチマークである \textsc{MiraBench} を紹介します。 MiraBench は、こ​​のターゲットを 3 つの段階的に要求の高いレベルに分解します。 \emph{Physics Adherence} は、リファレンスフリーの物理的一貫性を評価します。 \emph{Action-Following Fidelity}: 予測がタスク関連のアクション入力を考慮しているかどうかを測定します。 \emph{楽観主義バイアス検出} は、失敗を誘発する行動の下で成功した結果を予測する傾向を調査します。この評価をサポートするために、タスク、失敗カテゴリ、主要な世界モデルにわたる 16,000 件を超える判断を含む人間による注釈付きコーパスを厳選しました。ベクトル条件付きロボット ワールド モデル、テキスト条件付き生成ワールド モデル、オープンウェイト システム、クローズド ソース システム、および複数のモデル スケールにわたる 12 の代表的なモデル構成を評価します。この広範なモデル環境全体にわたって、MiraBench は 3 つの中心的な発見を明らかにしました。視覚的な忠実度は、アクションの忠実度の代用としては不十分です。モデルのスケールを大きくしても、アクションの追従性が確実に改善されるわけではありません。そして楽観主義バイアスは現在のシステム全体に蔓延しています。 MiraBench は、評価を外観から動作条件付きの信頼性に移行することで、ロボットの世界モデルを忠実なシミュレーターとして評価および改善するための診断基盤を提供します。

原文 (English)

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks largely emphasize visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to commanded actions, and calibrated to failure when actions should not succeed. We introduce \textsc{MiraBench}, a hierarchical benchmark that defines \emph{action-conditioned reliability} as a core evaluation target for robotic world models. MiraBench decomposes this target into three progressively demanding levels: \emph{Physics Adherence}, which evaluates reference-free physical consistency; \emph{Action-Following Fidelity}, which measures whether predictions respect task-relevant action inputs; and \emph{Optimism Bias Detection}, which probes the tendency to predict successful outcomes under failure-inducing actions. To support this evaluation, we curate a human-annotated corpus with over 16,000 judgments across tasks, failure categories, and leading world models. We evaluate 12 representative model configurations spanning vector-conditioned robotic world models, text-conditioned generative world models, open-weight systems, closed-source systems, and multiple model scales. Across this broad model landscape, MiraBench reveals three central findings: visual fidelity is a poor proxy for action fidelity; increasing model scale does not reliably improve action following; and optimism bias is pervasive across current systems. By shifting evaluation from appearance to action-conditioned reliability, MiraBench provides a diagnostic foundation for assessing and improving robotic world models as faithful simulators.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

エージェントによる修正と意味評価による人間のような対話型音声認識を目指して

自動音声認識 (ASR) は、人間とコンピューターの対話の中核コンポーネントであり、LLM ベースのアシスタントおよびエージェントにとってますます重要なフロントエンドです。しかし、現在のほとんどの ASR システムは依然としてシングルパス パラダイムに従っており、人間のコミュニケーションとの整合性が低く、誤解は繰り返しの明確化と改良によって解決されます。この不一致により、意味に関わる重大なエラーが発生すると、修正することが困難になります。一方、WER や CER などのトークンレベルの指標は、このような問題を適切に反映できません。これらの制限に対処するために、\emph{Interactive ASR} をマルチターン改良タスクとして定式化し、シングルパス ASR フロントエンドとセマンティック修正、インテント ルーティング、推論ベースの編集を組み合わせた閉ループ フレームワークである \textbf{Agentic ASR} を提案します。さらに、LLM ベースのセマンティック評価指標である \textbf{文レベルのセマンティック エラー率} ($S^2ER$) を、スケーラブルで再現可能なベンチマークのための \textbf{インタラクティブ シミュレーション システム} とともに導入します。多言語、名前付きエンティティ集中型、およびコードスイッチングのベンチマークに関する実験では、反復的な対話によりセマンティック エラーが一貫して減少し、従来のトークン レベルのメトリクスよりも $S^2ER$ が大幅に増加することが示されました。人間と AI のアライメントとアブレーションの研究により、意味判断の信頼性と提案されたフレームワークの堅牢性がさらに検証されました。コードは https://interactiveasr.github.io/ で入手でき、ライブ デモは https://i-asr.sjtuxlance.com/ で入手できます。

原文 (English)

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation

Automatic speech recognition (ASR) is a core component of human--computer interaction and an increasingly important front-end for LLM-based assistants and agents. However, most current ASR systems still follow a single-pass paradigm, which is poorly aligned with human communication, where misunderstandings are resolved through iterative clarification and refinement. This mismatch makes it difficult to correct meaning-critical errors once they occur. Meanwhile, token-level metrics such as WER or CER cannot adequately reflect such a problem. To address these limitations, we formulate \emph{Interactive ASR} as a multi-turn refinement task and propose \textbf{Agentic ASR}, a closed-loop framework that combines a single-pass ASR front-end with semantic correction, intent routing, and reasoning-based editing. We further introduce the \textbf{Sentence-level Semantic Error Rate} ($S^2ER$), an LLM-based semantic evaluation metric, together with an \textbf{Interactive Simulation System} for scalable and reproducible benchmarking. Experiments on multilingual, named-entity-intensive, and code-switching benchmarks show that iterative interaction consistently reduces semantic errors, with much larger gains in $S^2ER$ than in conventional token-level metrics. Human--AI alignment and ablation studies further validate the reliability of the semantic judge and the robustness of the proposed framework. The code is available at: https://interactiveasr.github.io/ and the live demo is available at https://i-asr.sjtuxlance.com/

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

TRACE: LLM CoT 評価の構成要素によるトゥールミンベースの推論評価

大規模言語モデル (LLM) からのオープンエンドの出力を評価することは、グランド トゥルースがないため依然として困難です。既存の指標は、最終的な答えの精度や表面レベルの統計に依存しており、推論プロセス自体は検討されていません。思考連鎖 (CoT) 推論プロセスを分析する指標である TRACE (Toulmin-based Reasoning Assessment through Constructive Elements) を紹介します。 TRACE は、結果を判断するのではなく、トゥールミンの議論理論とフラベルのメタ認知フレームワークを統合して推論の構造を評価することにより、議論がどのように構築されるかを検査します。 7 つの推論モデルにわたる 26.3K の QA サンプルの実験では、ベンチマーク精度 (r=0.74) との強い相関関係が示されています。さらに、TRACE は強化学習の報酬信号として効果的であり、精度のみのベースラインを上回ります。これらの結果を総合すると、論理的に健全な推論がより質の高い答えにつながることを示しています。したがって、TRACE は、オープンエンド出力を評価するための補足的なメトリックとして機能します。コードは https://github.com/hyyangkisti/trace で入手できます。

原文 (English)

TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-answer accuracy or surface-level statistics, leaving the reasoning process itself unexamined. We introduce TRACE (Toulmin-based Reasoning Assessment through Constructive Elements), a metric that analyzes Chain-of-Thought (CoT) reasoning processes. Rather than judging outcomes, TRACE inspects how arguments are constructed by integrating Toulmin's argumentation theory with Flavell's metacognitive framework to assess reasoning structure. Experiments on 26.3K QA samples across 7 reasoning models show strong correlation with benchmark accuracy (r=0.74). Furthermore, TRACE is effective as a reinforcement learning reward signal, outperforming accuracy-only baselines. Together, these results indicate that logically sound reasoning leads to higher-quality answers. TRACE thus serves as a complementary metric for evaluating open-ended outputs. Code is available at https://github.com/hyyangkisti/trace.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

スペシャリスト モデルが依然として重要な理由: 医療用人工知能のための異種マルチエージェント パラダイム

医療分野における GPT や Claude などの汎用大規模言語モデル (LLM) の優れたパフォーマンスは、領域固有の医療専門家モデルは時代遅れになるのだろうかという重大な疑問を引き起こしています。私たちは、医療用人工知能 (AI) の将来は、モノリシックな医療基盤モデルの構築や人間の専門知識の置き換えにあるのではなく、ジェネラリストの LLM、領域固有の専門家モデル、および臨床医の間のコラボレーションを調整することにあると主張します。我々は、矛盾を認識した証拠の融合、不確実性に基づく臨床医の介入トリガー、および適応閾値キャリブレーションを可能にする異種医療マルチエージェントフレームワークである HetMedAgent を提案します。 3 つの実際の臨床意思決定タスクに関する実験では、ジェネラリスト LLM と領域固有の専門家モデルの間の相乗効果が、どちらかのタイプのモデルを単独で使用した場合よりも大幅に優れていることが実証され、モダリティ固有の分析における専門家モデルのかけがえのない価値が検証されました。 HetMedAgent は、医療 LLM または基盤モデルの構築から複数エージェントのコラボレーションへの移行を表し、一般的な推論機能とドメイン固有の精度のバランスを実現します。

原文 (English)

Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence

The impressive performance of generalist large language models (LLMs) such as GPT and Claude in healthcare raises a critical question: will domain-specific medical specialist models become obsolete? We argue that the future of medical artificial intelligence (AI) lies not in building monolithic medical foundation models, nor in replacing human expertise, but in orchestrating collaboration among generalist LLMs, domain-specific specialist models, and clinicians. We propose HetMedAgent, a heterogeneous medical multi-agent framework that enables conflict-aware evidence fusion, uncertainty-based clinician intervention triggering, and adaptive threshold calibration. Experiments on three real-world clinical decision-making tasks demonstrate that the synergy between generalist LLMs and domain-specific specialist models significantly outperforms using either type of model alone, validating the irreplaceable value of specialist models in modality-specific analysis. HetMedAgent represents a shift from building medical LLMs or foundation models to multi-agent collaboration, achieving a balance between general reasoning capabilities and domain-specific precision.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

クロワッサン タスク: 再現可能な機械学習評価のためのメタデータ形式

再現性は科学的手法の基本ですが、機械学習においては依然として重要な課題です。原因としては、実行詳細の指定不足や脆弱なソフトウェア環境などが挙げられます。チェックリストや手動検証などの人間中心の救済策は役立ちますが、集中的な努力が必要であり、拡張することができません。これに対処するために、Croissant Tasks を導入します。これは、低レベルの実装の詳細を高レベルの仕様に抽象化する、宣言的でマシンアクション可能なメタデータ形式です。この形式により、概念的な再現性が可能になります。つまり、脆弱なソース コードの複製ではなく、独立したエージェント生成の実装を通じて主張を検証できます。私たちは以下に貢献しています。(1) Croissant Tasks 仕様。タスクの問題を解決策から正式に切り離します。 (2) 既存のベンチマークをこの形式に改良する自動 LLM パイプライン。 (3) 自律エージェントがこれらの仕様を取り込んで、機能的で正確な再現パイプラインを最初から生成できることを示す経験的検証。私たちはこの形式を、機械学習における自動化された概念的な再現性のための新しい基盤として構想しています。

原文 (English)

Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations

Reproducibility is fundamental to the scientific method, yet remains a critical challenge in machine learning. Contributing factors include underspecified execution details and brittle software environments. Human-centric remedies, such as checklists and manual verification, help but require intensive effort and fail to scale. To address this, we introduce Croissant Tasks: a declarative, machine-actionable metadata format that abstracts low-level implementation details into high-level specifications. This format enables conceptual reproducibility: verifying claims via independent, agent-generated implementations rather than brittle source code replication. We contribute: (1) the Croissant Tasks specification, formally decoupling task problem from solution; (2) an automated LLM pipeline that retrofits existing benchmarks into this format; and (3) empirical validation showing autonomous agents can ingest these specifications to generate functional, accurate reproduction pipelines from scratch. We envision this format as a new foundation for automated and conceptual reproducibility in machine learning.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Cookie-Bench: Web 生成のための継続的なオンスクリーンキーインタラクション評価

フロントエンドの Web コードは、すべてのフロンティア LLM リリースの中核的な製品面となっていますが、アリーナのような人間が判断するリーダーボードは拡張できないため、これらのインタラクティブ アプリケーションを開発スピードで評価することは依然としてコストがかかります。既存の自動プロキシは通常、リファレンス実装、テスト スイート、または厳密なチェックリストに依存しており、人間のレビュー担当者がライブ セッションで実行する合理的な合成を見逃す傾向があります。私たちは、同時に参照フリーで、自律的に駆動され、総合的に推論される新しい評価体制を明確にし、2 つの成果物を通じてそれをインスタンス化します。 \textbf{\dataname} は、静的プレゼンテーション タスクと対話型アプリケーション タスクの両方にまたがる 11 ドメイン、54 リーフ、1,000 クエリの WebDev ベンチマークであり、3 つの難易度層と 3 つのターゲット言語グループにわたってバランスが取れており、回覧されたプロンプトから思い出せないようにブリーフが書き直されています。 \textbf{\framename} は、Flavell のメタ認知モニタリングに基づいており、証拠の蓄積と判断を 3 つの段階にわたって分離します。静的な知覚は受動的な観察から第一印象を形成します。エージェント駆動のインタラクションは、連続画面のビデオ、音声、およびステップごとのスクリーンショットをキャプチャしながら、アプリケーションを自律的に探索します。動的スコアリングは、証拠チェーンが完了した後にのみ、構造化された失敗の帰属を伴う全体的な機能性と美的判断を発行します。 \dataname では、\framename は専門家による評価と厳密に一致しており、インタラクティブな Web 生成に関して 13 のフロンティア LLM 全体でかなりのヘッドルームを表面化しています。 \noindenthttps://anonymous.4open.science/r/Cookie-3CE/

原文 (English)

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation

Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed remains costly because human-judged leaderboards like Arena do not scale. Existing automated proxies typically lean on reference implementations, test suites, or rigid checklists, and tend to miss the reasoned synthesis a human reviewer performs over a live session. We articulate a new evaluation regime that is simultaneously reference-free, autonomously driven, and holistically reasoned, and instantiate it through two artifacts. \textbf{\dataname} is an 11-domain, 54-leaf, 1,000-query WebDev benchmark spanning both static-presentation and interactive-application tasks, balanced across three difficulty tiers and three target-language groups, with briefs rewritten to resist recall from circulated prompts. \textbf{\framename}, grounded in Flavell's metacognitive monitoring, separates evidence accumulation from judgment across three stages: Static Perception forms a first impression from passive observation; Agent-Driven Interaction explores the application autonomously while capturing continuous screen video, audio, and per-step screenshots; Dynamic Scoring issues holistic functionality and aesthetics verdicts with structured failure attribution only after the evidence chain is complete. On \dataname, \framename aligns closely with expert human ratings while surfacing substantial headroom across 13 frontier LLMs on interactive web generation. \noindenthttps://anonymous.4open.science/r/Cookie-3CE/

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

RAISE: アーキテクチャ検索問題としての RAG 設計

検索拡張生成 (RAG) システムでは、クエリの書き換え、チャンキング、検索の深さ、再ランキング、およびコンテキスト圧縮に及ぶ数多くの設計上の選択肢が明らかになります。実際には、これらの選択はヒューリスティックによって構成されることが多く、設定全体での体系的な評価と再現性が妨げられます。私たちは、この課題は RAG アーキテクチャの検索として定式化するのが最適であると主張します。この問題の制御された再現可能な研究をサポートするために、RAG ハイパーパラメータ最適化の包括的なフレームワークおよびベンチマークである RAG Intelligence Search Engine (RAISE) を導入します。これは、標準化された検索スペースと予算の下で RAG パイプラインの最適化方法を評価します。 RAISE は 13 の検索アルゴリズムを実装し、3 つのランダム シードを使用して 7 つのパブリック テキストおよびマルチモーダル データセットにわたってそれらを評価します。私たちの実験は、最適化のパフォーマンスがタスクに大きく依存することを示しています。つまり、あるデータセットで優れたパフォーマンスを発揮する手法が、他のデータセットでは一貫して一般化できない可能性があり、集計されたランキングを普遍的に優れた戦略の証拠として解釈することには注意が必要です。 RAISE は、RAG ハイパーパラメータの最適化に関する公正で再現性のある体系的な研究のための共通の実験基盤を提供します。

原文 (English)

RAISE: RAG Design as an Architecture Search Problem

Retrieval-augmented generation (RAG) systems expose numerous design choices spanning query rewriting, chunking, retrieval depth, reranking, and context compression. In practice, these choices are often configured through heuristics, hindering systematic evaluation and reproducibility across settings. We argue that this challenge is best formulated as RAG architecture search. To support controlled and reproducible study of this problem, we introduce the RAG Intelligence Search Engine (RAISE), a comprehensive framework and benchmark for RAG hyperparameter optimization, which evaluates optimization methods for RAG pipelines under standardized search spaces and budgets. RAISE implements 13 search algorithms and evaluates them across seven public text and multimodal datasets using three random seeds. Our experiments show that optimization performance is highly task-dependent: methods that perform strongly on one dataset may not generalize consistently across others, cautioning against interpreting aggregate rankings as evidence of universally superior strategies. RAISE provides a common experimental substrate for fair, reproducible, and systematic research on RAG hyperparameter optimization.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

矛盾する複数ソースの個人記憶に対する選択的 QA: 診断テストベッドと手法の比較

新興のパーソナル AI エージェントは、永続的なマルチソース メモリに移行しています。これにより、評価上の問題が生じます。システムは、矛盾する証拠や不完全な証拠をどのように使用するかを決定する必要があります。 1 つのきれいな歴史から事実を引き出すことはできません。既存のベンチマークでは、エラーがメソッドに与えられた証拠に起因するのか、メソッドの競合解決ステップに起因するのかを示すことはほとんどありません。私たちはこれを、矛盾する複数ソースの個人記憶に対する選択的 QA として研究しています。システムは、矛盾する、場合によっては不完全なソースに基づいて回答するか、証拠が不十分な場合は棄権します。 8 つの推論タイプにわたる 18 の質問テンプレート、480 のペルソナ、4 つのランダム シード、および 34,560 のインスタンスを含むベンチマークを、制御されたソースの歪みと決定論的なグラウンド トゥルースを使用して開発しました。ソースへのアクセスなし、単一ソースへのアクセス、構造化融合手法、およびフロンティア LLM のベースラインのパフォーマンスを評価します。最もよく訓練されたフュージョン リゾルバーの精度は 80.3% に達し、最も強力なプロンプトのみの LLM ベースラインは 70.0% に達します。棄権すると、同じリゾルバはカバレッジ 78.3% で選択精度 85.3% に達し、最良の LLM はカバレッジ 95.4% で選択精度 71.0% に達します。モデルが異なれば、推論タイプごとに異なる強みがあります。データ、コード、キャッシュされたモデル出力、およびデータ生成プロセスを再利用のためにリリースします。

原文 (English)

Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison

Emerging personal AI agents are moving toward persistent, multi-source memory. This creates an evaluation problem: systems must decide how to use conflicting or incomplete evidence; they cannot just retrieve facts from one clean history. Existing benchmarks rarely show whether an error came from the evidence given to a method or from the method's conflict-resolution step. We study this as selective QA over conflicting multi-source personal memory: systems answer based on conflicting, sometimes incomplete sources, or abstain when evidence is insufficient. We develop a benchmark containing 18 question templates across 8 reasoning types, 480 personas, 4 random seeds, and 34,560 instances, with controlled source distortions and deterministic ground truth. We evaluate the performance of baselines without access to any source, access to a single source, structured fusion methods, and frontier LLMs. The best trained fusion resolver reaches 80.3% accuracy, while the strongest prompt-only LLM baseline reaches 70.0%. With abstention, the same resolver reaches 85.3% selective accuracy at 78.3% coverage and the best LLM reaches 71.0% selective accuracy at 95.4% coverage. Different models have different strengths across reasoning types. We release the data, code, cached model outputs, and data-generating process for reuse.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達研究/論文

BioRefusalAudit: 一般およびドメイン微調整されたスパース オートエンコーダーを使用したバイオセキュリティ拒否の深さの監査

言語モデルのバイオセキュリティ評価では通常、モデルが危険な出力を生成するかどうかが問われます。この論文は補足的な質問をします。モデルが拒否した場合、その拒否は構造的に正しいのでしょうか、それともフレーミング、フォーマット、または出力長を促すための適度な変更で消えるのでしょうか? 5 つのアーキテクチャにわたって、無害性と危険性を明確に区別したモデルはありませんでした。 Gemma 2 2B-IT は、75 件のプロンプトにわたって真に拒否することはなく、危険に隣接するすべてのクエリを回避しました。 Gemma 4 E2B-IT は、チャット テンプレート形式を使用した場合は 65/75 件のプロンプトを拒否し、チャット テンプレート形式を使用しない場合は 0/75 件のプロンプトを拒否しました。両方の Gemma モデルは、80 トークンの上限の下で 0% に崩壊しました。 Qwen 2.5 1.5B と Phi-3-mini は過剰に拒否され、良性生物学の 83 ~ 87% が危険であると警告されました。 Llama 3.2 1B は唯一の意味のある Tier 勾配 (61 ポイントの広がり) を示しました。何がそのような過剰な拒否を引き起こすのかを調査するために、我々はスケジュールIであるが生物学的に無毒な化合物(特にFDA画期的治療法のステータスを持つシロシビン培養)のパネルをテストしました。一部のモデルは、真に有害な生物学を超える割合でこれらを拒否しており、拒否がCBRNの危険性に対する合法性と文化的顕著性を追跡していることを示唆しています。内部側を測定するために、モデルの表面応答ラベルを内部のスパース オートエンコーダー (SAE) 特徴のアクティベーションと比較する発散スコア D を導入します。フル D は、Gemma 2 2B-IT (Gemma Scope 1) および Gemma 4 E2B-IT (著者が訓練したバイオ SAE) で計算されました。 2 つの微調整された Gemma 2 ドメイン SAE がリリースされました。 Gemma 4 では、狭いカタログ、サンプル内キャリブレーション、および Gemma ファミリーのみの SAE 範囲を使用して、重複なし (n=75) で 0.647 ポイントのギャップで応答と拒否の応答が分離されますが、これは暫定的なものです。消費者向けハードウェア (GTX 1650 Ti Max-Q、および SAE トレーニング用の Colab T4) での 1 つのハッカソン週末にわたって構築されたこの予備的な証拠は、アクティベーション レベルの監査によって、アーキテクチャ間で大幅に異なる、動作評価では見えない障害モードが表面化する可能性があることを示唆しています。

原文 (English)

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model refuses, is that refusal structurally sound, or does it disappear under modest changes to prompt framing, formatting, or output length? Across five architectures, no model cleanly discriminated benign from hazard. Gemma 2 2B-IT never genuinely refused across 75 prompts, hedging on every hazard-adjacent query. Gemma 4 E2B-IT refused 65/75 prompts with chat-template formatting and 0/75 without it. Both Gemma models collapsed to 0% under an 80-token cap. Qwen 2.5 1.5B and Phi-3-mini over-refused, flagging 83-87% of benign biology as hazardous. Llama 3.2 1B showed the only meaningful tier gradient (61-point spread). To probe what drives such over-refusal, we tested a panel of Schedule I but biologically non-toxic compounds (notably psilocybin cultivation, with FDA Breakthrough Therapy status). Some models refused these at rates exceeding genuinely hazardous biology, suggesting refusal tracks legality and cultural salience over CBRN hazard. To measure the internal side, we introduce a divergence score D comparing a model's surface response label to its internal sparse autoencoder (SAE) feature activations. Full D was computed on Gemma 2 2B-IT (Gemma Scope 1) and Gemma 4 E2B-IT (author-trained bio SAE). Two fine-tuned Gemma 2 domain SAEs were released. On Gemma 4, comply and refuse responses separated by a 0.647-point gap with zero overlap (n=75), though this is preliminary, with a narrow catalog, within-sample calibration, and Gemma-family-only SAE coverage. Built over one hackathon weekend on consumer hardware (GTX 1650 Ti Max-Q, plus Colab T4 for SAE training), this preliminary evidence suggests activation-level auditing may surface failure modes invisible to behavioral evaluation, with substantial variation across architectures.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

オープンソースの安全ガード モデルのベンチマーク: 包括的な評価

安全性が重要なアプリケーションに大規模言語モデル (LLM) が導入されることが増えているため、堅牢なコンテンツ モデレーションが不可欠になっています。 NIST AI リスク フレームワークの 8 つの安全カテゴリにまたがる 79,331 サンプルの厳選されたベンチマークに基づく 14 のオープンソース安全ガード モデルの包括的な評価を示します。当社のベンチマークは 4 つの多様なデータセット (HarmBench、StrongREJECT、RealToxicityPrompts、BeaverTails) を集約し、安全関連のコンテンツ (暴力、ヘイトスピーチ、嫌がらせ、性的コンテンツ、自殺/自傷行為、冒涜、脅迫、健康上の誤った情報) のみに焦点を当てるようにフィルタリングされています。有害なコンテンツの欠落は誤検知よりも大きなリスクをもたらすため、リコールは安全性アプリケーションにとって重要な指標であることがわかりました。私たちの評価では、驚くべき結果が明らかになりました。Qwen Guard (4B パラメーター) は最高の再現率 (83.97%) を達成しましたが、Llama Guard (12B) や GPT-OSS Safeguard (20B) などのより大きなモデルは保守的な動作を示し、安全でないコンテンツを最大 75% 見逃しました。我々は、モデルのサイズが安全検出のパフォーマンスと相関しないこと、および汎用のガード モデルが特殊なガード モデルよりも優れていることを実証します。これらの調査結果は、実稼働環境での安全装置モデルを選択するための実践的なガイダンスを提供します。

原文 (English)

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a comprehensive evaluation of 14 open-source safety guard models on a curated benchmark of 79,331 samples spanning 8 NIST AI Risk Framework safety categories. Our benchmark aggregates four diverse datasets (HarmBench, StrongREJECT, RealToxicityPrompts, and BeaverTails), filtered to focus exclusively on safety-relevant content (violence, hate speech, harassment, sexual content, suicide/self-harm, profanity, threats, and health misinformation). We find that recall is the critical metric for safety applications, as missing harmful content poses greater risk than false positives. Our evaluation reveals surprising results: Qwen Guard (4B parameters) achieves the highest recall (83.97%) while larger models like Llama Guard (12B) and GPT-OSS Safeguard (20B) exhibit conservative behavior, missing up to 75% of unsafe content. We demonstrate that model size does not correlate with safety detection performance and that general-purpose guard models outperform specialized ones. These findings provide practical guidance for selecting safety guard models in production deployments.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

GPF-LiveNews: 大規模言語モデルにおけるグループ条件付きフレーミングのためのストリーミング評価プロトコル

デプロイされた言語モデルは非定常環境で評価されます。モデルのバージョン、検索レイヤー、安全システム、現実世界の入力はすべて時間の経過とともに変化します。静的バイアスのベンチマークは依然として有用ですが、モデルがさまざまな刺激を受けた視聴者に対して新たに出現したイベントをどのように組み立てるかは示していません。オープンエンド LLM 出力のグループ条件付きフレーミングを監査するためのストリーミング評価プロトコルおよびベンチマーク スナップショットである GPF-LIVENEWS を紹介します。このプロトコルは、42 の ID ラベルと 7 つのプロンプト ファミリにわたって新鮮な BBC/ロイター ニュース アンカーを拡張し、その後、意味論的感度とセンチメント差異シグナルを使用して応答バンドルを評価します。 12 回のモニタリング実行と 23 個のホストされたモデルにわたるパイロットでは、ポリシー/アクション プロンプトが最も強力なセマンティックな動きを生成しますが、センチメントの変動はディメンションおよびプロンプト ファミリ全体でより平坦です。リリースされたアーティファクトには、記事のメタデータ、プロンプト テンプレート、インスタンス化されたプロンプト、モデル出力メタデータ、スコア テーブル、ドキュメント、および再現スクリプトが含まれます。私たちはすべてのスコアを、永続的な公平性ランキングや有害なバイアスの直接の証拠としてではなく、人間によるレビューのための監視窓監査シグナルとして解釈します。

原文 (English)

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all change over time. Static bias benchmarks remain useful, but they do not show how models frame newly emerging events for different prompted audiences. We introduce GPF-LIVENEWS, a streaming evaluation protocol and benchmark snapshot for auditing group-conditioned framing in open-ended LLM outputs. The protocol expands fresh BBC/Reuters news anchors across 42 identity labels and seven prompt families, then evaluates response bundles using semantic-sensitivity and sentiment-disparity signals. In a pilot over 12 monitoring runs and 23 hosted models, Policy/Action prompts produce the strongest semantic movement, while sentiment variation is flatter across dimensions and prompt families. The released artifact includes article metadata, prompt templates, instantiated prompts, model-output metadata, score tables, documentation, and reproduction scripts. We interpret all scores as observed-window audit signals for human review, not as permanent fairness rankings or direct proof of harmful bias.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

GrowLoop: 人間がシードし、自己進化する会話評価

大規模な言語モデルの急速な進歩に伴い、自由な会話における人間らしさを評価することがますます重要になってきています。しかし、人間らしさは人間が直感的に認識する暗黙知の一種ですが、根底にある基準は明示的な定式化に抵抗します。人間の判断は大きく異なり、一部のケースでは強い同意が得られますが、他のケースでは正当な意見の相違が見られます。一方、人間の判断の背後にある基準は暗黙的なままであり、事件を構築するための明確な根拠は残されていません。さらに、人間に似ているとみなされるものは静的なものではなく、モデルの能力と人間の期待に応じて進化します。専門家が作成したベンチマーク、報酬モデル、自己進化型ベンチマークなどの評価方法は進歩していますが、3 つの課題すべてに同時に対処できるものはありません。そこで、モデルの進歩やシナリオの変化に合わせて継続的に適応する、自己進化する会話評価システムである GrowLoop を提案します。最初の動きとして最小限の人間のシード アノテーションを使用して、LLM エージェントはヒューリスティック学習を通じて評価ルーブリックを繰り返し抽出し、改良します。アノテーターが集まる場合には人間と AI の合意が必要ですが、異なる場合には妥当性のみが期待されます。さらに、Rubric-Caseの共進化機構により、評価対象が移動した際に新たなシーズを介して拡張され、継続的な進化が可能となります。自由形式の会話における人間らしさの評価に適用すると、生成されたルーブリックは、人間の判断に沿って既存の手法を大幅に上回るだけでなく、アノテーターが見落としている問題も明らかになります。結果として得られるベンチマークは、機能層全体でモデルを効果的に識別し、どこが不足しているかを明らかにすると同時に、新しいシナリオに一般化し、モデルの進歩に合わせて適応します。私たちの取り組みは、ベンチマークのパラダイムを手動の更新や難易度のスケーリングから、包括的で継続的な自己進化へと移行させます。

原文 (English)

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However, human-likeness is a form of tacit knowledge that humans perceive intuitively, yet the underlying criteria resist explicit formulation. Human judgments vary widely, with strong agreement on some cases and legitimate disagreement on others. Meanwhile, the criteria behind human judgments remain implicit, leaving no clear basis for constructing cases. Further, what counts as human-like is not static, but evolving with model capability and human expectations. Despite progress in evaluation methods such as expert-authored benchmarks, Reward Models, and self-evolving benchmarks, none addresses all three challenges simultaneously. Therefore, we propose GrowLoop, a self-evolving conversation evaluation system that continuously adapts as models advance and scenarios shift. With minimal human seed annotations as the first mover, LLM agents iteratively extract and refine evaluation rubrics through Heuristic Learning. Human-AI agreement is required where annotators converge, while only plausibility is expected where they diverge. Moreover, the Rubric-Case co-evolution mechanism enables continuous evolution, expanded through new seeds when the evaluation target moves. Applied to human-likeness evaluation in open-ended conversation, the generated rubrics not only substantially outperform existing methods in alignment with human judgments, but also uncover issues that annotators overlook. The resulting benchmark effectively discriminates models across capability tiers and reveals where they fall short, while generalizing to new scenarios and adapting as models advance. Our work shifts the benchmarking paradigm from manual updates or difficulty scaling to comprehensive, continuous self-evolution.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

LoRe: 反復グラフ ソルバー向けのステップごとのインタラクション バジェットを備えた適応型インタラクション評価ルーティング

組み合わせ最適化のための拡散ベースのニューラル ソルバーは、高密度のエッジ/因子相互作用を繰り返し再評価するため、実時間での推論が高価になり、大規模になるとメモリに制限されることがよくあります。多体物理学の計算手法にインスピレーションを得て、ステップごとの相互作用評価の予算設定を強制する、トレーニング不要の推論時間ドロップイン ラッパーである LoRe を導入します。各反復では、固定のスパース化 (静的 kNN グラフや静的など) を使用する代わりに、計算を競合性の高い相互作用または不確実性の高い相互作用に動的にルーティングすることで、相互作用の固定部分のみを評価します。マスク)。完全に包括的なエンドツーエンドの壁時計アカウンティングの下で​​、LoRe は最大独立集合 (MIS) 問題のスケーラビリティを大幅に向上させ、実行可能な推論をベースラインのメモリ不足制限を超えて $3\times$ 以上拡張し、$\sim 8\times$ の高速化と $\sim 12\times$ のピークメモリ削減を実現し、この体制でソリューションの品質は維持されます。大規模な巡回販売員問題 (TSP) に対するクロスタスクの汎用性と、トポロジーの変化に対するゼロショットの堅牢性を実証する LoRe は、$n=1000$ で $\sim 15\times$ の高速化を実現し、$44\times$ のメモリ削減と競争力のあるツアー品質を実現します。

原文 (English)

LoRe: Adaptive Interaction-Evaluation Routing with Per-Step Interaction Budgets for Iterative Graph Solvers

Diffusion-based neural solvers for combinatorial optimization repeatedly re-evaluate dense edge/factor interactions, making inference expensive in wall-clock time and often memory-bound at scale. Inspired by the computational methodologies of many-body physics, we introduce LoRe, a training-free, inference-time drop-in wrapper that enforces per-step interaction-evaluation budgeting: at each iteration, it evaluates only a fixed fraction of interactions by dynamically routing computation to high-conflict or high-uncertainty interactions, instead of using a fixed sparsification (e.g., static kNN graphs or static masks). Under fully inclusive end-to-end wall-clock accounting, LoRe substantially improves scalability on the Maximum Independent Set (MIS) problem, extending feasible inference more than $3\times$ beyond the baseline's out-of-memory limit, delivering a $\sim 8\times$ speedup and a $\sim 12\times$ peak-memory reduction, with solution quality preserved in this regime. Demonstrating cross-task generality on the large-scale Traveling Salesperson Problem (TSP) and zero-shot robustness to topology shifts, LoRe achieves a $\sim 15\times$ speedup at $n=1000$ with a $44\times$ memory reduction and competitive tour quality.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

倫理的な顔年齢推定に向けて: 子供のデータに関するトレーニングを行わない一般化されたゼロショット ベンチマーク

顔画像からの年齢推定は通常、未成年者の画像を含むトレーニング データに依存しますが、これは倫理的、法的、プライバシー上で重大な懸念を引き起こす行為です。この研究では、若い集団に対するモデルのパフォーマンスを評価しながら、トレーニング中に子供のデータを明示的に除外する、顔年齢推定のための一般化されたゼロショット ベンチマークを提案します。私たちは、広く使用されている 6 つのデータセットを再検討し、年齢グループを厳密に分離した標準化された分割を導入します。トレーニング、検証、テストには 18 ~ 59 歳のサンプルを使用します。 18 歳未満のサンプルはゼロショット評価専用に予約されています。そして、分布シフトの下でのモデル選択のための未確認の検証セットとして 60 以上のサンプルをサンプリングします。 ID アノテーションを持つデータセットの場合、サブジェクト排他的な分割により ID 漏洩が防止され、現実世界のデプロイ状況がより適切に反映されます。このプロトコルに基づいて 9 つの最先端の年齢推定方法を評価すると、評価されたすべての方法が目に見えない年齢グループに一貫して一般化できず、教師付きベースラインと比較して、平均で 46.4%、最大で 52.8% という大幅なパフォーマンスの低下が見られることが明らかになりました。さらに、モデルは単純に劣化するわけではありません。モデルは、目に見えない年齢の予測を近くの見たクラスに体系的に固定します。これは、一般化されたゼロショット学習におけるよく知られた見えるクラスのバイアスの現れです。この研究では、子供のデータを使用しない年齢推定を既存のデータセットに対する一般化されたゼロショット ベンチマークとして形式化することで、現在のモデリング実践と現実世界の倫理的制約との間の重大なギャップを浮き彫りにしています。私たちのベンチマークは、制限されたデータ体制の下でモデルを評価するための原則に基づいた基礎を提供し、分布の変化に強く、責任あるデータの使用に合わせた方法の開発を奨励します。

原文 (English)

Toward Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data

Age estimation from facial images typically relies on training data that includes images of minors, a practice that raises serious ethical, legal, and privacy concerns. In this work, we propose a generalized zero-shot benchmark for facial age estimation that explicitly excludes children's data during training while still assessing model performance on younger populations. We revisit six widely used datasets and introduce standardized splits with strict age-group separation: samples aged 18-59 for training, validation, and testing; samples under 18 reserved exclusively for zero-shot evaluation; and samples 60+ as an unseen validation set for model selection under distribution shift. For datasets with identity annotations, subject-exclusive splits prevent identity leakage and better reflect real-world deployment conditions. Evaluating nine state-of-the-art age estimation methods under this protocol reveals that all evaluated methods consistently fail to generalize to unseen age groups, suffering substantial performance degradation -- on average 46.4%, and up to 52.8% -- relative to the supervised baseline. Moreover, models do not simply degrade: they systematically anchor predictions for unseen ages to nearby seen classes, a manifestation of the well-known seen-class bias in generalized zero-shot learning. By formalizing age estimation without children's data as a generalized zero-shot benchmark on existing datasets, this work highlights a critical gap between current modeling practices and real-world ethical constraints. Our benchmark provides a principled basis for evaluating models under restricted data regimes and encourages the development of methods that are robust to distribution shift and aligned with responsible data use.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

DynSess: ロールプレイング エージェント向けの動的なセッション レベルの評価および最適化フレームワーク

大規模な言語モデルを使用したロールプレイングは基本的にセッション レベルのタスクであり、エージェントは長時間にわたる複数ターンの会話にわたってキャラクターのアイデンティティと対話の品質を維持する必要があります。しかし、既存の評価および最適化手法は依然としてターンレベルにとどまっており、長期的な品質を捉えることができません。私たちは、ロールプレイング エージェントのための統合されたセッション レベルのフレームワークである DynSess を提案します。 DynSess-Eval は、長期的な行動を対象としたルーブリックを介して、完全な対話セッションをスコア付けします。セッションレベルの報酬を活用して、マルチターン先読み検索を通じて高品質のトレーニング軌道を構築し、2 つの補完的なバリアントである DSPO (オフポリシー) と GSRPO (オンポリシー) で DynSess-Character をトレーニングします。実験では、DynSess-Eval が以前の評価者よりも人間の判断とかなり良く一致していることが示されており、人間によるブラインド評価ではさらに、DynSess-Character が、使用するパラメータが大幅に少ないにもかかわらず、強力な役割の一貫性とインタラクティブな能力を維持しながら、最強のキャラクター モデルと一致していることが示されています。私たちのデータセットとコードは、将来の研究を促進するためにリリースされます。

原文 (English)

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality across extended multi-turn conversations. Yet existing evaluation and optimization methods remain largely turn-level, failing to capture long-horizon quality. We propose DynSess, a unified session-level framework for role-playing agents. DynSess-Eval scores complete dialogue sessions via rubrics targeting long-horizon behaviors. Leveraging its session-level rewards, we construct high-quality training trajectories through multi-turn lookahead search and train DynSess-Character with two complementary variants: DSPO (off-policy) and GSRPO (on-policy). Experiments show that DynSess-Eval aligns with human judgments substantially better than prior evaluators, and blind human evaluation further shows that DynSess-Character matches the strongest character model despite using substantially fewer parameters, while maintaining strong role consistency and interactive ability. Our dataset and code will be released to facilitate future research.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

物理基礎モデルは一般化可能な物理学を学習しますか?物理的体制と分布の変化にわたるバイアスを意識したベンチマーク

最近の物理基礎モデルは一般的な時空間予測能力を主張していますが、その評価は、固定されたトレーニング分布の下でパフォーマンスを単一の平均スコアに落とし込んでしまうことがよくあります。このため、モデルが一般化可能な物理ダイナミクスを学習しているのか、それとも特定の設定下でのみ適切に動作するのかを判断することが困難になります。 8 つの物理ダイナミクス、3 つのトレーニング データ混合物、および動的スケールと初期条件の複雑さのシフトによって誘発される 25 のテスト レジームを使用してベンチマークを構築し、分布内、分布シフト、および分布外の設定をカバーします。 5 つの物理基礎モデル アーキテクチャとアーキテクチャごとに 4 つのモデル バリアント (スクラッチ サイズと 3 つの事前トレーニング サイズ) を評価し、結果として 60,000 の測定結果が得られます。私たちの結果は、現在の物理基礎モデルが普遍的なジェネラリストとしてではなく条件付きで動作することを示しています。その一般性は、物理レジーム、時間スケール、初期条件の設定、事前トレーニング、モデルのサイズ、アーキテクチャに依存します。トレーニング データの分散を改善しても、この制限は部分的にしか緩和されません。事前トレーニングとスケーリングでも、能力のバイアスを確実に取り除くことができません。私たちは、物理基礎モデルを改善するには、モデルのスケーリングやデータの拡張を超えて、領域、時間スケール、分布の変化を超えて移転可能な物理知識をより適切に捕捉する学習メカニズムに移行する必要があると主張します。

原文 (English)

Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts

Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a single average score under a fixed training distribution. This makes it difficult to determine whether a model has learned generalizable physical dynamics or only performs well under particular settings. We construct a benchmark with 8 physical dynamics, 3 training-data mixtures, and 25 test regimes induced by dynamic-scale and initial-condition complexity shifts, covering in-distribution, distribution-shift, and out-of-distribution settings. We evaluate five physics foundation model architectures and four model variants per architecture (scratch and three pretrained sizes), resulting in 60,000 measurements. Our results show that current physics foundation models behave as conditional rather than universal generalists: their generality depends on the physical regime, temporal scale, initial-condition setting, pretraining, model size, and architecture. Improving the training data distribution only partially mitigates this limitation. Pretraining and scaling are also unable to reliably remove their ability biases. We argue that improving physics foundation models requires moving beyond scaling models or expanding data, toward learning mechanisms that better capture transferable physical knowledge across regimes, temporal scales, and distribution shifts.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

Pocket-Dentist: 効率的なマルチモーダル大規模言語モデルによるオンデバイス歯科画像理解

歯科視覚言語モデルの評価は、データセット、タスク定義、メトリクスにわたって断片化されたままであり、多くの場合、その計算コストが無視されます。このため、専門センター外での歯科スクリーニングへの広範な導入が制限されています。専門センターでは、タイムリーな推論、限られたハードウェア、および患者画像のローカル処理が、実用的でプライバシーを保護した臨床前スクリーニングに不可欠です。ここで紹介する Pocket-Dentist は、約 1,159 人の患者、5 つのタスク タイプ、7 つの指標にまたがる 3 つのデータセットをまとめた、歯科マルチモーダル質問応答のための効率を意識したベンチマークです。典型的な 14 個の VLM にわたって、我々の結果は興味深い観察結果を明らかにしました。コンパクトな VLM (例: 2B パラメータ モデル) は、精度においては大型の VLM を上回っていますが、歯科画像の理解に必要な計算コストは​​大幅に低くなります。 iPhone 17 Pro にローカルに導入された当社の微調整されたコンパクト VLM Pocket-Dentist-2B は、各サンプルを 4.31 秒で処理し、7B ベースラインと比較してレイテンシーを 4.9 倍、メモリ使用量を 2.3 倍削減しました。

原文 (English)

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widespread deployment for dental screening outside specialist centres, where timely inference, limited hardware, and local handling of patient images are vital for practical, privacy-preserving clinical prescreening. Here we present Pocket-Dentist, an efficiency-aware benchmark for dental multimodal question answering that brings together three datasets spanning approximately 1,159 patients, five task types and seven metrics. Across typical 14 VLMs, our results reveals an interesting observation: compact VLMs (e.g., 2B-parameter models) outperform larger VLMs in accuracy while requiring substantially lower computational costs in dental image understanding. Deployed locally on an iPhone 17 Pro, our finetuned compact VLM Pocket-Dentist-2B processed each sample in 4.31 s, reducing latency by 4.9-fold and memory use by 2.3-fold compared with a 7B baseline.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

データセットの価値はいくらですか?スケーリング則、Vendi スコア、および行列スペクトル関数

ニューラル スケーリングの法則はデータセットのサイズを通じてデータを評価しますが、Vendi スコアは量子エントロピーを使用してデータセットの値を測定します。一般的なニューラル スケーリング則の目標と Vendi スコアの両方がサブモジュールであることを示します。さらに、Vendi スコアが、行列スペクトル関数と呼ばれるより広範なクラスのサブモジュラー目標の特殊なケースであることを示します。これには、決定的 (DPP) 目標や他の多くの目標も含まれます。また、弱行列単調関数を導入し、それがどのように弱部分モジュール行列スペクトル関数につながるかを示し、データ評価のための幅広い実用的な目的をもたらします。私たちは、貪欲な最適化中に繰り返される固有分解を回避する永年方程式ベースの更新を開発し、$m$ 次元の埋め込みに対する限界ゲイン評価を Oracle クエリと比較して $O(m)$ 係数だけ削減します。これにより、経験的に平均約 35,000 倍の高速化が得られ、ImageNet-1K スケールのデータセットで Vendi スコアの直接最適化が可能になります。このようにして可能になったので、Vendi スコア、DPP、施設の場所、および 3 つの新しいマトリックス スペクトル バリアントを含む、固定サイズ、クラスバランス、および固定トレーニング予算体制の下で、いくつかの目標がホールドアウト テスト パフォーマンスのトレーニング サブセットの値をどの程度正確に予測するかを比較します。複数のデータセットにわたって、施設の位置が最も優れたパフォーマンスを発揮します。また、直接最適化では、Vendi スコアは中程度のスコア範囲では予測的ですが、目標をより高い値に押し上げると、下流のパフォーマンスの代用として機能しなくなる可能性があることも明らかになりました。また、均一でランダムな固定サイズのサブセットは、制約がなく、クラスバランスが取れていても、評価スコアと保持されたパフォーマンスの両方で著しく集中していることもわかります。最後に、サイズ、クラスのバランス、トレーニング予算だけがデータの価値を決定するわけではないことを示します。これらの要因を制御した場合でも、パフォーマンスは良い状態から悪い状態まで滑らかに変化します。

原文 (English)

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions

Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common neural-scaling-law objectives and the Vendi Score are submodular. We further show that the Vendi Score is a special case of a broader class of submodular objectives that we call matrix spectral functions. This also includes determinantal (DPP) objectives, as well as many others. We also introduce weakly matrix monotone functions and show how they lead to weakly submodular matrix spectral functions, yielding a broad family of practical objectives for data appraisal. We develop secular-equation-based updates that avoid repeated eigendecompositions during greedy optimization, reducing marginal-gain evaluation for $m$-dimensional embeddings by an $O(m)$ factor relative to oracle queries. This yields an average empirical speedup of about 35,000x, making direct optimization of the Vendi Score feasible on ImageNet-1K-scale datasets. Thus enabled, we compare how well several objectives predict the value of training subsets for held-out test performance under fixed-size, class-balanced, and fixed training-budget regimes, including the Vendi Score, DPPs, facility location, and three new matrix spectral variants. Across multiple datasets, facility location performs the best. Direct optimization also reveals that, while the Vendi Score is predictive over moderate score ranges, pushing the objective to higher values can make it a poor downstream performance proxy. We also find that uniformly at random fixed-size subsets, both unconstrained and class-balanced, are remarkably concentrated in both appraisal scores and held-out performance. Finally, we show that size, class balance, and training budget do not alone determine data value: even when controlling for these factors, performance ranges smoothly from good to bad.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

CFMME 上の大規模なビジョン言語モデルのベンチマーク: 包括的な中国金融マルチモーダル評価データセット

Large Vision-Language Models (LVLM) の出現により、モデルの機能がテキストのみの理解を超えて大幅に拡張され、視覚的モダリティとテキストモダリティの両方にわたる統一された推論が可能になり、より広範な現実世界のアプリケーションがサポートされるようになりました。中国の状況における金融ビジネスのワークフロー全体を通じて、LVLM の認識、理解、推論、認知能力を包括的に評価するために、中国の新しい金融マルチモーダル評価ベンチマークである CFMME を紹介します。 CFMME は、基礎的な学術知識から複雑な現実世界のアプリケーションに至る 6,052 のインスタンスで構成され、8 つの主要な金融イメージ モダリティと 4 つのコア マルチモーダル タスクをカバーします。 CFMMEでは、代表的なLVLMを徹底的に評価します。結果は、最先端のモデルが質問応答タスクで全体の精度 66.11%、検出、認識、および情報抽出タスクで平均スコア 77.18 を達成していることを示しており、現在の LVLM には改善の余地がかなりあることが示されています。さらに、エラーの原因、クロスモーダル機能、および複数の方向設定の詳細な分析を実施し、将来の研究に役立つ貴重な洞察をもたらします。私たちは、CFMME が、特に金融ドメインにおける複数のマルチモーダル タスクのパフォーマンスを向上させることによって、LVLM のさらなる進歩に拍車をかけることを期待しています。

原文 (English)

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified inference across both visual and textual modalities and supporting a broader range of real-world applications. To comprehensively evaluate the perception, understanding, reasoning, and cognition capabilities of LVLMs throughout the entire financial business workflow in Chinese contexts, we introduce CFMME, a novel Chinese financial multimodal evaluation benchmark. CFMME comprises 6,052 instances spanning from fundamental academic knowledge to complex real-world applications, covering eight primary financial image modalities and four core multimodal tasks. On CFMME, we conduct a thorough evaluation of representative LVLMs. The results show that the state-of-the-art model attains an overall accuracy of 66.11\% on the question answering task and an average score of 77.18 on the detection, recognition, and information extraction tasks, indicating substantial room for improvement in current LVLMs. In addition, we conduct detailed analyses of error causes, cross-modal capabilities, and multi-orientation settings, yielding valuable insights for future research. We hope that CFMME will spur further progress in LVLMs, especially by improving their performance on multiple multimodal tasks in the financial domain.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

オフポリシー評価のための商 DAG: フォワードフロー重要度サンプリングと正確なスレート傾向

オフポリシー評価は、別の動作ポリシーによって収集されたデータを使用して、ターゲットポリシーがどのように実行されるかを推定します。これは、推奨や医療など、オンラインテストにコストがかかる、またはリスクが伴う場合に非常に重要です。標準重要度サンプリングでは、ログに記録された各軌跡の重み付けが変更されますが、評価ターゲットがそれらを無視する場合でも、生成プロセスの詳細を意味のあるものとして扱うことができます。たとえば、自己回帰スレート レコメンダーは順序付けられたアイテムのシーケンスを生成する一方で、報酬と下流の推定器は順序付けられていないスレートのみに依存する場合があります。厳密に順序付けされていないスレートの傾向にはすべての生成順序の合計が必要なため、これにより迷惑な分散と計算上のギャップが生じます。評価用に同等の履歴をマージし、マージされたグラフ上のターゲットと行動の順方向フロー比を使用して重みを割り当てる商 DAG ビューを導入します。 set-sufficient next-item インターフェイスの下でのスレート推奨の場合、階乗列挙なしで正確な順序付けされていない傾向を計算するサブセット DAG 動的プログラムである Forward-DP が生成されます。結果として得られる傾向プリミティブにより、コンテキスト依存の自己回帰スレート ロガーの実用的な傾向ベースの評価とモデル選択が可能になります。

原文 (English)

Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities

Off-policy evaluation estimates how a target policy would perform using data collected by a different behavior policy, which is crucial when online testing is costly or risky, such as in recommendation or healthcare. Standard importance sampling reweights each logged trajectory, but it can treat details of the generation process as meaningful even when the evaluation target ignores them: for example, an autoregressive slate recommender may generate an ordered sequence of items while the reward and downstream estimator depend only on the unordered slate. This creates nuisance variance and a computational gap, since exact unordered slate propensities require summing over all generation orders. We introduce a quotient-DAG view that merges histories equivalent for evaluation and assigns weights using target-to-behavior forward-flow ratios on the merged graph. For slate recommendation under a set-sufficient next-item interface, this yields Forward-DP, a subset-DAG dynamic program that computes exact unordered propensities without factorial enumeration. The resulting propensity primitive enables practical propensity-based evaluation and model selection for context-dependent autoregressive slate loggers.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

GUITestScape: 探索的 GUI テストのオープンセット評価に向けて

探索的 GUI テストは、MLLM エージェントにとって特に要求の厳しい設定です。事前定義されたテスト スクリプトがなければ、エージェントは自律的にアプリケーションを操作し、独自の対話を通じて欠陥を発見する必要があります。しかし、現在の評価は 2 つの点で不十分です。まず、既存のベンチマークはほぼインタラクションの欠陥のみに焦点を当てており、ディスプレイの欠陥は評価の枠外に残されています。第 2 に、評価プロトコルは事前定義された欠陥の注釈にバインドされており、テスト プロセスが質的に異なる故障モードを混同した単一の最終状態の判断に崩壊します。これらの課題に対処するために、61 の実世界の Android アプリケーションと、インタラクションと表示タイプにわたる 508 のプリセット欠陥をカバーするインタラクティブなベンチマークである GUITestScape を紹介し、エージェントのテストの軌跡を独立して診断可能な機能に分解するオープンセット評価器である GUIJudge を紹介します。実験結果は、GUIJudge が事前定義された注釈を超えた信頼性の高いプロセス認識型評価を達成し、すべてのベースラインを大幅に上回るパフォーマンスを示していることを示しています。 GUITestScape のベンチマークでは、両方の欠陥タイプにわたって、既存のモデルにとって検出が依然として重大なボトルネックであること、および GUIJudge のベリファイアを既存のエージェントに統合することで、再トレーニングすることなく検出パフォーマンスが大幅に向上することがさらに明らかになりました。

原文 (English)

GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an application and discover defects through its own interaction. However, current evaluation falls short on two fronts. First, existing benchmarks focus almost exclusively on interaction defects, leaving display defects outside the evaluation frame. Second, evaluation protocols are bound to predefined defect annotations, collapsing the testing process into a single end-state judgment that conflates qualitatively distinct failure modes. To address these challenges, we present GUITestScape, an interactive benchmark covering 61 real-world Android applications and 508 preset defects spanning interaction and display types, and introduce GUIJudge, an open-set evaluator that decomposes an agent's testing trajectory into independently diagnosable capabilities. Experimental results demonstrate that GUIJudge achieves reliable process-aware evaluation beyond predefined annotations, substantially outperforming all baselines. Benchmarking on GUITestScape further reveals that detection remains the critical bottleneck for existing models across both defect types, and that integrating GUIJudge's verifiers into existing agents significantly boosts their detection performance without retraining.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

EviLink: 大規模な Text-to-SQL のための不確実性に基づく証拠取得を使用したマルチパス スキーマ リンク

スキーマのリンクは、大規模な Text-to-SQL では困難かつ重要なステップであり、システムは大規模で曖昧なデータベースからコンパクトでありながら十分なスキーマ コンテキストを識別する必要があります。既存の方法では、多くの場合、単一の SQL パスに関する決定論的な選択としてスキーマ リンクが扱われますが、複雑な質問では、異なるスキーマ ニーズを持つ複数の有効な実現が認められる場合があります。スキーマ リンクを、複数の妥当な SQL パスにわたる不確実性を認識したスキーマ ニーズ推論として再構成します。システムは、必要なスキーマ項目をパスに依存する不確実な項目から区別し、必要な場合にのみ証拠を取得します。私たちは、この再構成を EviLink でインスタンス化します。これは、複数の仮説スキーマの根拠と、不確実性に基づいた証拠の取得を組み合わせたものです。 BIRD-Dev と Spider2-Snow の実験では、この観点により、スキーマの完全性、スキーマの関連性、トークン コストの間のバランスが改善されることが示されています。 Spider2-Snow では、EviLink はフィールドレベルの厳密な再現率 90.15% を達成し、平均 123.30K のトークンを使用し、固定ジェネレータの下でダウンストリーム SQL 生成を改善します。

原文 (English)

EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL

Schema linking is a difficult and important step in large-scale Text-to-SQL, where systems must identify a compact yet sufficient schema context from large and ambiguous databases. Existing methods often treat schema linking as deterministic selection around a single SQL path, but complex questions may admit multiple valid realizations with different schema needs. We reframe schema linking as uncertainty-aware schema-need inference over multiple plausible SQL paths, where the system distinguishes required schema items from path-dependent uncertain ones and acquires evidence only where needed. We instantiate this reframing with EviLink, which combines multi-hypothesis schema grounding with uncertainty-guided evidence acquisition. Experiments on BIRD-Dev and Spider2-Snow show that this perspective improves the balance among schema completeness, schema relevance, and token cost. On Spider2-Snow, EviLink achieves 90.15% field-level strict recall rate, uses 123.30K average tokens, and improves downstream SQL generation under a fixed generator.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Honeyval: LLM を利用した HTTP ハニーポットの包括的な評価フレームワーク

ハニーポットは、サイバー攻撃から防御するために設計された実際のシステム コンポーネントを模倣したおとりシステムです。最近では、LLM がハニーポットのシミュレーション バックボーンとして機能することが増えています。これらにより、防御者はシステム セキュリティ リスクを低く抑えながら、インタラクションの多いハニーポットを構築できます。ただし、LLM を利用したハニーポット開発には、統一された評価フレームワークがありません。ほとんどの評価は、固定コマンド、手動テスト、または実際の展開での応答の類似性の測定で構成されます。これらの手法は、多くの場合、開発の拡張性、評価全体での再現性、実際の攻撃の代表性、またはさまざまな攻撃者やハニーポットの構成への適応性がありません。この取り組みでは、このギャップを埋め、LLM を利用した HTTP ハニーポットの包括的な評価フレームワークである Honeyval を提案します。私たちは、16 のバックエンド アプリケーションにハニーポットを設置し、AI ハッキング エージェントを攻撃者として使用し、カスタマイズ全体にわたってエージェントとハニーポットの機能を監視する 2 つの制御タスクを採用し、攻撃者に対する明確で検証可能なエクスプロイト目標を定義することで、以前の評価の制限に対処しました。 Honeyval を使用して、HTTP ハニーポットとしての最近のコスト効率の高い LLM の広範な評価を実施します。私たちの実験は、LLM を利用したハニーポットの将来性を強調しています。これらは、ルールベースのベースライン ハニーポットよりも攻撃者とのやり取りが大幅に長くなり、フロンティア モデルでも検出される頻度がはるかに低くなり、平均してエージェント攻撃者に対するランニング コストの利点が維持されます。さらに、さまざまな反攻撃用ハニーポット構成を実験し、検出力の向上と引き換えにインタラクションが長くなるなど、独特のトレードオフを観察しました。

原文 (English)

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation backbones for honeypots. They enable defenders to construct high-interaction honeypots with low system security risks. However, LLM-powered honeypot development lacks a unified evaluation framework. Most evaluations consist of measuring response similarity on fixed commands, manual testing, or real-world deployment. These methods are often not scalable for development, reproducible across evaluations, representative of practical attacks, or adaptable to various attacker and honeypot configurations. In this work, we bridge this gap and propose Honeyval, a comprehensive evaluation framework for LLM-powered HTTP honeypots. We address the limitations of prior evaluations by grounding the honeypots in 16 backend applications, using AI hacking agents as attackers, employing two control tasks to monitor agent and honeypot capabilities across customizations, and defining clear and verifiable exploit goals for the attacker. Using Honeyval, we conduct an extensive evaluation of recent cost-efficient LLMs as HTTP honeypots. Our experiments highlight the promise of LLM-powered honeypots; they lead to substantially longer interactions with the attacker than rule-based baseline honeypots and are far less frequently detected even by frontier models, all while, on average, preserving a running cost advantage against agentic attackers. Further, we experiment with different counter-offensive honeypots configurations, and observe unique trade-offs, such as longer interactions at the cost of increased detection.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

大規模なオーディオ言語モデルにおけるオーディオ ジェイルブレイク: 分類法、攻撃防御分析、コストを意識した評価

Large Audio Language Model (LALM) は、脱獄のリスクをトークン レベルのプロンプトから音声認識から推論までのパイプライン全体に拡大します。セマンティクス、音響スタイル、信号アーティファクト、または内部表現を通じて安全でない動作が誘発される可能性があります。既存の研究では、異種混合の脅威モデルと評価プロトコルに基づいてこれらのリスクが研究されており、攻撃の実用性や防御の有用性を比較することが困難になっています。このペーパーでは、LALM ジェイルブレイク攻撃と防御の統一された分類法と管理された経験的評価を提供します。私たちはこれまでの研究をセマンティック攻撃、音響攻撃、信号攻撃、埋め込み層攻撃に分けて整理しています。ガードベース、トレーニング不要、トレーニングベースのディフェンス。クロスモーダル、オーディオネイティブ、およびインタラクティブなベンチマーク。次に、10 個のオープンソース LALM にわたる代表的な攻撃と防御を評価し、攻撃の成功率だけでなく、良性の拒否と遅延も測定します。私たちの結果は、Acoustic Best-of-N が強力な最悪の場合のオーディオ空間の脆弱性を明らかにし、Narrative Framing が効果的な低レイテンシのセマンティック脅威であり、現在の防御は無害なユーザビリティに対する堅牢性と引き換えにあることを示しています。これらの調査結果は、成功率のみを重視した LALM の安全性ベンチマークを補完する必要があるものとして、コストとユーティリティを意識した評価を裏付けています。

原文 (English)

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsafe behavior can be induced through semantics, acoustic style, signal artifacts, or internal representations. Existing work studies these risks under heterogeneous threat models and evaluation protocols, making it difficult to compare attack practicality or defense utility. This paper provides a unified taxonomy and a controlled empirical evaluation of LALM jailbreak attacks and defenses. We organize prior work into semantic, acoustic, signal, and embedding-layer attacks; guard-based, training-free, and training-based defenses; and cross-modal, audio-native, and interactive benchmarks. We then evaluate representative attacks and defenses across ten open-source LALMs, measuring not only attack success rate but also benign refusal and latency. Our results show that Acoustic Best-of-N reveals strong worst-case audio-space vulnerabilities, Narrative Framing is an effective low-latency semantic threat, and current defenses trade robustness against benign usability. These findings support cost- and utility-aware evaluation as a necessary complement to success-rate-only LALM safety benchmarks.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MedCase 構造化: 臨床的に現実的な EHR 設定における診断推論のベンチマーク用の Text-to-FHIR データセット

大規模言語モデル (LLM) は、臨床推論と意思決定のサポートに有望ですが、現実的な電子医療記録と一致する設定での評価には依然として限界があります。既存のベンチマークは、多くの場合、臨床システムで使用される構造化された相互運用可能なデータ形式を反映していない静的データセットまたは非構造化入力に依存しています。非構造化テキストから臨床的に現実的な HL7 FHIR R4 バンドルを生成するパイプラインを導入し、臨床意思決定支援システムの制御可能な評価を可能にします。このパイプラインは、段階的な LLM 生成と用語に基づいた検証および修復を組み合わせて、幻覚コードを削減し、構造的および意味的な一貫性を強化します。このアプローチを MedCaseReasoning に適用して、臨床医が作成した診断症例に合わせた合成データセットである MedCase-Structured を構築し、症例の 82.5% で有効な FHIR 生成を実現します。 MedCase-Structured での評価では、平文の場合よりも構造化 FHIR 入力での LLM の診断精度が一貫して低いことが明らかになり、展開に合わせたベンチマークの重要性が強調されています。

原文 (English)

MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in realistic, electronic health record-congruent settings remains limited. Existing benchmarks often rely on static datasets or unstructured inputs that do not reflect the structured, interoperable data formats used in clinical systems. We introduce a pipeline for generating clinically realistic HL7 FHIR R4 bundles from unstructured text, enabling controllable evaluation of clinical decision support systems. The pipeline combines staged LLM generation with terminology-grounded validation and repair to reduce hallucinated codes and enforce structural and semantic consistency. Applying this approach to MedCaseReasoning, we construct MedCase-Structured, a synthetic dataset aligned with clinician-authored diagnostic cases, achieving valid FHIR generation for 82.5% of cases. Evaluation on MedCase-Structured reveals consistently lower diagnostic accuracy for LLMs on structured FHIR inputs than with plain text, highlighting the importance of deployment-aligned benchmarking.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

IntentScore: コンピュータ使用エージェントの意図条件付きアクションの評価

Computer-Use Agent (CUA) は、大規模な言語モデルを利用してデスクトップ環境で GUI 操作を実行しますが、アクションの品質を評価せずにアクションを生成するため、後続のステップに連鎖的に発生する不可逆的なエラーにつながります。私たちは、3 つのオペレーティング システムにわたる 398K のオフライン GUI インタラクション ステップから候補アクションをスコアリングすることを学習する、プランを認識した報酬モデルである IntentScore を提案します。 IntentScore は、状態とアクションの関連性に関する対照的な調整と、アクションの正しさに関するマージン ランキングという 2 つの相補的な目標を使用してトレーニングします。アーキテクチャ的には、各候補者の計画意図がアクション エンコーダーに埋め込まれ、同様のアクションを持つ候補者間で論理的根拠が異なるものを区別できるようになります。 IntentScore は、ホールドアウト評価で 97.5% のペア識別精度を達成します。トレーニング中にまったく見えない環境である OSWorld 上のエージェント S3 の再ランカーとしてデプロイされた IntentScore は、タスクの成功率を 6.9 ポイント向上させ、異種のオフライン軌跡から学習した報酬推定が、目に見えないエージェントとタスクの分布に一般化されることを示しています。

原文 (English)

IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents

Computer-Use Agents (CUAs) leverage large language models to execute GUI operations on desktop environments, yet they generate actions without evaluating action quality, leading to irreversible errors that cascade through subsequent steps. We propose IntentScore, a plan-aware reward model that learns to score candidate actions from 398K offline GUI interaction steps spanning three operating systems. IntentScore trains with two complementary objectives: contrastive alignment for state-action relevance and margin ranking for action correctness. Architecturally, it embeds each candidate's planning intent in the action encoder, enabling discrimination between candidates with similar actions but different rationales. IntentScore achieves 97.5% pairwise discrimination accuracy on held-out evaluation. Deployed as a re-ranker for Agent S3 on OSWorld, an environment entirely unseen during training, IntentScore improves task success rate by 6.9 points, demonstrating that reward estimation learned from heterogeneous offline trajectories generalizes to unseen agents and task distributions.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

速く考えると間違った考え: 直観力が政策評価における LLM 反事実的推論を調整する

大規模言語モデル (LLM) は、因果関係や反事実の推論にますます使用されていますが、現実世界の政策評価におけるその信頼性は依然として十分に解明されていません。私たちは、経済学と社会科学から抽出された 40 の実証的政策評価ケースのベンチマークを構築します。それぞれの事例は査読済みの証拠に基づいており、経験的発見が事前の一般的な期待と一致する (明白)、相対的に不明瞭である (曖昧)、または矛盾している (直観に反する) かどうかという直感によって分類されます。私たちは、8,000 件の実験トライアルを使用して 5 つのプロンプト戦略にわたって 4 つのフロンティア LLM を評価し、混合効果ロジスティック回帰を使用して結果を分析します。私たちの調査結果では、3 つの重要な結果が明らかになりました。(1) 思考連鎖 (CoT) のパラドックス。思考連鎖プロンプトは明らかなケースではパフォーマンスを劇的に向上させますが、直感に反するケースではこの利点が大幅に減弱します (相互作用 OR = 0.278、$p < 0.001$)。 (2) 支配的な要因としての直観性。ケースレベルの分散がモデルの選択またはプロンプト戦略の分散を超えています (ICC = 0.671)。 (3) 知識と推論の解離。引用に基づく精通度は精度と無関係です ($p = 0.84$)。これは、モデルが関連する知識を持っているものの、発見が直観と矛盾する場合、それを使って推論できないことを示唆しています。私たちはこれらの結果を二重プロセス理論(システム 1 対システム 2)のレンズを通して組み立て、現在の LLM の「遅い思考」は直観的事前の抑制を部分的にしか達成していない、つまりその本質を完全に実現することなく熟慮的推論の形式を生み出していると主張します。

原文 (English)

Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation

Large language models (LLMs) are increasingly used for causal and counterfactual reasoning, yet their reliability in real-world policy evaluation remains underexplored. We construct a benchmark of 40 empirical policy evaluation cases drawn from economics and social science, each grounded in peer-reviewed evidence and classified by intuitiveness -- whether the empirical finding aligns with (obvious), is unclear relative to (ambiguous), or contradicts (counter-intuitive) common prior expectations. We evaluate four frontier LLMs across five prompting strategies with 8,000 experimental trials and analyze the results using mixed-effects logistic regression. Our findings reveal three key results: (1) a chain-of-thought (CoT) paradox, where chain-of-thought prompting dramatically improves performance on obvious cases but this benefit is substantially attenuated on counter-intuitive ones (interaction OR = 0.278, $p < 0.001$); (2) intuitiveness as the dominant factor, with case-level variance exceeding that of model choice or prompting strategy (ICC = 0.671); and (3) a knowledge-reasoning dissociation, where citation-based familiarity is unrelated to accuracy ($p = 0.84$), suggesting models possess relevant knowledge but fail to reason with it when findings contradict intuition. We frame these results through the lens of dual-process theory (System 1 vs. System 2) and argue that current LLMs' "slow thinking" achieves only partial inhibition of intuitive priors -- producing the form of deliberative reasoning without fully delivering its substance.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

SciHorizo​​n-DataEVA: 異種科学データの AI 対応性評価のためのエージェント システム

AI-for-Science (AI4Science) は、機械学習モデルをドメイン全体の予測、シミュレーション、仮説生成のワークフローに組み込むことにより、科学的発見をますます変革しています。ただし、これらのモデルの有効性は科学データの AI 対応度によって根本的に制約されており、現在、拡張性のある体系的な評価メカニズムが存在しません。この研究では、異種科学データのスケーラブルな AI 対応性評価のための新しいエージェント システムである SciHorizo​​n-DataEVA を提案します。評価基準レベルでは、AI への対応力をガバナンスの信頼性、データ品質、AI 互換性、科学的適応性という 4 つの補完的な側面に整理する Sci-TQA2 原則を導入します。各次元は測定可能な原子要素に分解され、きめ細かい実行可能な評価が可能になります。これらの原則を大規模に運用するために、私たちは、指示された循環ワークフローを通じて調整された階層型マルチエージェント評価アプローチである Sci-TQA2-Eval を開発しました。当社の Sci-TQA2-Eval は、軽量のデータセット プロファイリング、適用性を意識したメトリクスのアクティベーション、ドメイン制約とデータセット ペーパー シグナルに基づいた知識拡張計画を組み合わせることにより、データセットを意識した評価仕様を動的に構築します。これらの仕様は、組み込みの検証と自己修正を備えた適応型のツール中心の評価メカニズムを通じて実行され、異種の科学データ全体にわたってスケーラブルで信頼性の高い評価が可能になります。複数のドメインにわたる科学データセットに関する広範な実験により、原則に基づいた AI 対応性評価における SciHorizo​​n-DataEVA の有効性と汎用性が実証されました。

原文 (English)

SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data

AI-for-Science (AI4Science) is increasingly transforming scientific discovery by embedding machine learning models into prediction, simulation, and hypothesis generation workflows across domains. However, the effectiveness of these models is fundamentally constrained by the AI-readiness of scientific data, for which no scalable and systematic evaluation mechanism currently exists. In this work, we propose SciHorizon-DataEVA, a novel agentic system to scalable AI-readiness evaluation of heterogeneous scientific data. At the evaluation-criteria level, we introduce the Sci-TQA2 principles, which organize AI-readiness into four complementary dimensions: Governance Trustworthiness, Data Quality, AI Compatibility, and Scientific Adaptability. Each dimension is decomposed into measurable atomic elements that enable fine-grained and executable assessment. To operationalize these principles at scale, we develop Sci-TQA2-Eval, a hierarchical multi-agent evaluation approach orchestrated through a directed, cyclic workflow. Our Sci-TQA2-Eval dynamically constructs dataset-aware evaluation specifications by combining lightweight dataset profiling, applicability-aware metric activation, and knowledge-augmented planning grounded in domain constraints and dataset-paper signals. These specifications are executed through an adaptive, tool-centric evaluation mechanism with built-in verification and self-correction, enabling scalable and reliable assessment across heterogeneous scientific data. Extensive experiments on scientific datasets spanning multiple domains demonstrate the effectiveness and generality of SciHorizon-DataEVA for principled AI-readiness evaluation.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

CausaLab: AI 科学者向けのインタラクティブな因果発見のためのスケーラブルな環境

LLM エージェントによるインタラクティブな因果発見を評価するためのスケーラブルな環境である CausaLab を紹介します。以前の評価とは異なり、CausaLab では、エージェントが因果関係の証拠を使用して問題を解決できるかどうか、およびその答えが根底にある因果メカニズムに関する正しい仮説によって裏付けられているかどうかの両方を評価します。各エピソードではエージェントが合成実験室に配置されます。エージェントは以前の測定記録を受け取り、マニピュレーター結晶に介入し、同じ機構によって支配される保持されたリアクター結晶の共振周波数を予測します。隠されたデータ生成プロセスは、ランダムにサンプリングされた構造因果モデル (SCM) であるため、成功するには、事前の知識を思い出すのではなく、因果グラフと構造方程式の両方を回復する必要があります。 CausaLab には、エージェントの進化する SCM 仮説を記録するドメイン固有の言語も含まれており、軌跡を検査可能にしてグラウンド トゥルースと比較できるようになります。実験では、予測とメカニズム回復の間に永続的なギャップがあることが示されています。純粋に観測的な 6 ノード設定では、GPT-5.2-high はタスク精度 92% に達しますが、オールエッジ $F_1$ はわずか 0.471 です。この観察は、さまざまな相互作用戦略の探求をさらに動機づけます: 混合観察 - 介入戦略は構造忠実度を向上させます: 混合 6 ノード設定では、GPT-5.2-high はタスク精度とオールエッジ $F_1$ の両方で 80% を達成しました。しかし、純粋な介入戦略はタスクの精度とオールエッジ $F_1$ の両方においてパフォーマンスが低いため、強力なエージェントですら有益な介入を設計するのに苦労しています。私たちは、エージェントの主要な弱点として早期停止を特定し、仮説と過去のデータとの間の一貫性をモデルに検証するように依頼することが、この問題の軽減に役立つことを示します。したがって、CausaLab は予測の成功を因果関係の理解から切り離し、実験的因果推論者としての現在の LLM エージェントの限界を明らかにします。

原文 (English)

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates both whether an agent can solve a problem using causal evidence and whether its answer is grounded in a faithful recovered causal mechanism. Each episode places an agent in a synthetic laboratory: it receives prior measurement records, intervenes on a manipulator crystal, and predicts the resonance frequency of a held-out reactor crystal governed by the same mechanism. The hidden data-generating process is a randomly sampled structural causal model (SCM), so success requires recovering both a causal graph and structural equations rather than recalling prior knowledge. Experiments show a persistent gap between prediction and mechanism recovery: in the purely observational 6-node setting, GPT-5.2-high reaches 92% task accuracy but only 0.471 all-edge $F_1$. Mixed observation-intervention strategies improve structural fidelity, while pure intervention remains difficult even for strong agents. We identify premature stopping as a major weakness and show that consistency verification mitigates it. CausaLab therefore separates predictive success from causal understanding and exposes current LLM agents' limits as experimental causal reasoners.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

FundaPod: AI 支援のファンダメンタル投資調査のためのナレッジ グラフ メモリを備えたマルチペルソナ エージェント ポッド プラットフォーム

大規模言語モデル (LLM) は金融分野での適用が増えていますが、既存の研究のほとんどは取引シグナルや予測を中心とした財務 NLP タスクに重点を置いています。対照的に、制度的基礎研究では、人間のアナリストまたは AI エージェントが証拠を収集し、ビジネス推進要因を特定し、競合する視点を比較し、投資メモを作成する必要があります。その広範な目標は、単に結果を予測することではなく、投資知識の累積的な発展に貢献しながら、透明性、再利用可能、検証可能な投資計画を作成することです。 AI 支援のファンダメンタルズ投資調査のためのマルチペルソナ エージェント プラットフォームである FundaPod を紹介します。私たちは、基礎研究は人間中心の意思決定支援タスクであり、取引シグナルの生成とは質的に異なるため、独立性を維持するアーキテクチャの方が適していると主張します。 FundaPod では、バリュー投資家やマクロ戦略家など、さまざまなペルソナを持つ AI エージェントが、共有の出所契約に基づいて独立して調査を実施します。その後、彼らの意見の相違は、知識グラフ記憶システムを通じて人間のポートフォリオ マネージャー (PM) による裁定のために事後的に表面化されます。この論文は、設計科学の実践と認知的分離と人間と機械の協調の理論に基づいた、基礎研究をサポートする人間と AI のハイブリッド システムの 5 つの設計原則を提供します。また、4 つのアーキテクチャ メカニズムについても説明します。1 つは一般投資家の資料を展開可能なエージェントに変えるペルソナ蒸留パイプラインです。プランナーが型指定されたタスク グラフを導出できるようにする宣言型スキル レジストリ。メモの主張を検証可能な情報源に結び付ける根拠のある証拠モデル。そしてティッカー、メモ、アナリスト、テーマを結び付けるナレッジグラフ「第二の脳」。完全なケーススタディとペルソナベースのメモの比較を通じてアーキテクチャを実証します。

原文 (English)

FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research

Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on prediction. Institutional fundamental research, by contrast, requires human analysts or AI agents to gather evidence, identify business drivers, compare competing viewpoints, and generate investment memos. Its broader goal is not merely to predict outcomes, but to produce investment plans that are transparent, reusable, and verifiable, while contributing to the cumulative development of investment knowledge. We present FundaPod, a multi-persona agent platform for AI-assisted fundamental investment research. We argue that fundamental research is a human-centric decision-support task that is qualitatively distinct from trading-signal generation, and is therefore better served by an independence-preserving architecture. In FundaPod, AI agents with different personas, such as value investors or macro strategists, conduct research independently under a shared provenance contract. Their disagreements are then surfaced post hoc for adjudication by the human portfolio manager (PM) through a knowledge-graph memory system. This paper contributes five design principles for human-AI hybrid systems supporting fundamental research, grounded in design-science practice and theories of cognitive isolation and human-machine coordination. It also describes four architectural mechanisms: a persona distillation pipeline that turns public investor materials into deployable agents; a declarative skill registry that lets the planner derive typed task graphs; a grounded evidence model that links memo claims to verifiable sources; and a knowledge-graph "second brain" that connects tickers, memos, analysts, and themes. We demonstrate the architecture through a complete case study and a persona-based memo comparison.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

統計的に真剣であることの重要性: GSM シンボリックの重要な再評価

GSM-Symbolic ベンチマーク (Mirzadeh et al., 2025) は、GSM8K 問題のテンプレート生成バリアントでテストした場合、25 の大規模言語モデル (LLM) 全体で一貫したパフォーマンスの低下を報告し、モデルには真の推論機能が欠けていると結論付けました。私たちは、この結論は不安定な統計的根拠に基づいていると主張します。質問ごとの変量効果を備えた一般化線形混合モデルを使用して 20 の無重みモデルを再評価すると、元のプロンプト形式で統計的に有意なパフォーマンス変化を示したのは半分だけであることがわかりました。さらに、これまで認められていなかった要因も特定しました。つまり、メインの GSM-Symbolic データセットには、GSM-Base と比較して問題テキスト内のより大きな整数の体系的にシフトされた分布が含まれており (K-S 統計量 = 0.12、p < 0.001)、これは原著者の主張と矛盾しています。この大きな数の効果を制御することは、残りのケースの約半分で重要性を説明します。統計的に有意なパフォーマンスデルタを持つモデルの中で、変数結合の脆弱性、算術的制限、デュアルタスク干渉など、モデル固有の明確な障害プロファイルを特定しました。これは、LLM 推論に関する包括的な主張が統計的に時期尚早であり、機構的に誤解を招くものであることを強調します。

原文 (English)

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-generated variants of GSM8K problems, concluding that the models lack genuine reasoning capabilities. We argue that this conclusion rests on shaky statistical ground. Re-evaluating 20 open-weight models using Generalised Linear Mixed Models with per-question random effects, we find that only half exhibit statistically significant performance changes under the original prompt format. Moreover, we identify a previously unacknowledged factor: the main GSM-Symbolic dataset contains a systematically shifted distribution of larger integers in problem texts relative to GSM-Base (K-S statistic = 0.12, p < 0.001), contradicting the original authors' claims. Controlling for this large number effect accounts for significance in roughly half the remaining cases. Among models with statistically significant performance deltas, we identify distinct, model-specific failure profiles - including fragility of variable binding, arithmetic limitations, and dual-task interference - underscoring that blanket claims about LLM reasoning are both statistically premature and mechanistically misleading.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ルーブリックから信頼できるスコアまで: LLM 審査員による証拠に基づいたテキスト評価

ルーブリックベースのテキスト評価では、スケーラブルな審査員として大規模言語モデル (LLM) がますます使用されていますが、凍結されたブラックボックス モデルを人間の採点基準に合わせるのは依然として困難です。私たちは、この課題を基準移行問題として定式化します。目標は、単に LLM にスコアの割り当てを促すことではなく、人間のルーブリックの意図を、安定した、監査可能な、人間に合わせたスコアリング プロトコルに移行することです。私たちは、LLM ベースのルーブリック スコアリングで繰り返される 3 つの失敗モード、つまりルーブリック実行ドリフト、検証不可能なスコアの帰属、人間スケールの不整合を特定しました。これらの障害モードに対処するために、信頼できる証拠に基づいたルーブリック ベースのテキスト評価のための 3 段階の推論時間フレームワークである Rulers を導入します。ルーラーはまず人間によるルーブリックをロックされたタスクレベルの仕様に変換し、次に構造化されたチェックリストの決定、型付けされた証拠の根拠付け、該当する場合には抽出的な引用の検証を伴う仕様を実行し、最後に事後キャリブレーションを適用してモデル由来の信号を人間のスコア境界と一致させます。 Rulers は、エッセイの採点、要約評価、EFL ライティングの評価、構造化入力テキストの生成をカバーする 4 つのルーブリックに基づいたベンチマーク全体で、複数の凍結されたバックボーン モデルにわたるほとんどの評価設定で人間によるスコアのより強い一致を達成しています。さらに分析を進めると、ルーラーは人間の経験的なスコア分布によりよく一致し、意味的に同等のルーブリック摂動下での安定性が向上し、その 3 つのコンポーネントそれぞれから利点が得られることが示されています。これらの結果は、信頼できる LLM 判定には、即時の表現だけではなく、固定基準、追跡可能な証拠、および調整されたスコア解釈が必要であることを示唆しています。私たちのコードは https://anonymous.4open.science/r/Rulers_0525-3328 で入手できます。

原文 (English)

From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges

Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challenging. We formulate this challenge as a criteria-transfer problem: the goal is not merely to prompt an LLM to assign a score, but to transfer human rubric intent into a stable, auditable, and human-aligned scoring protocol. We identify three recurring failure modes in LLM-based rubric scoring: rubric execution drift, unverifiable score attribution, and human-scale misalignment. To address these failure modes, we introduce Rulers, a three-stage inference-time framework for reliable, evidence-grounded rubric-based text evaluation. Rulers first converts a human rubric into a locked task-level specification, then executes the specification with structured checklist decisions, typed evidence grounding, and extractive quote verification when applicable, and finally applies post-hoc calibration to align model-derived signals with human score boundaries. Across four rubric-governed benchmarks covering essay scoring, summarization assessment, EFL writing evaluation, and structured-input text generation, Rulers achieves stronger human-score agreement in most evaluated settings across multiple frozen backbone models. Further analyses show that Rulers better matches empirical human score distributions, improves stability under semantically equivalent rubric perturbations, and benefits from each of its three components. These results suggest that reliable LLM judging requires fixed criteria, traceable evidence, and calibrated score interpretation rather than prompt phrasing alone. Our code is available at https://anonymous.4open.science/r/Rulers_0525-3328.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

GICDM: 信頼性の高い距離ベースの生成モデル評価のためのハブネスの軽減

生成モデルの評価は通常、高次元の埋め込み空間に依存してサンプル間の距離を計算します。これらの空間のデータセット表現は、最近隣関係を歪め、距離ベースのメトリクスを偏らせるハブネス現象の影響を受けることを示します。古典的な反復コンテキスト相違測定 (ICDM) に基づいて、実際のデータと生成されたデータの両方の近傍推定を修正する方法である生成 ICDM (GICDM) を導入します。経験的な動作を改善するためにマルチスケール拡張を導入します。合成ベンチマークと実際のベンチマークに関する広範な実験により、GICDM がハブネスに起因する障害を解決し、信頼性の高いメトリック動作を復元し、人間の評価との整合性が向上することが実証されています。

原文 (English)

GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation

Generative model evaluation commonly relies on high-dimensional embedding spaces to compute distances between samples. We show that dataset representations in these spaces are affected by the hubness phenomenon, which distorts nearest-neighbor relationships and biases distance-based metrics. Building on the classical Iterative Contextual Dissimilarity Measure (ICDM), we introduce Generative ICDM (GICDM), a method to correct neighborhood estimation for both real and generated data. We introduce a multi-scale extension to improve empirical behavior. Extensive experiments on synthetic and real benchmarks demonstrate that GICDM resolves hubness-induced failures, restores reliable metric behavior, and improves alignment with human assessment.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

P$^2$RAG: 任意の上位 $k$ 取得をサポートする効率的なプライバシー保護 RAG サービス

検索拡張生成 (RAG) を使用すると、大規模な言語モデルで外部の知識を使用できるようになりますが、RAG サービスをアウトソーシングすると、データ所有者とユーザーの両方にプライバシー上の懸念が生じます。プライバシーを保護する RAG システムは、安全な上位 $k$ 取得を実行することでこれらの懸念に対処します。これは通常、関連するドキュメントを識別するための安全な並べ替えを使用して実装されます。しかし、既存のシステムは、$k$ を変更できないこと、新たなセキュリティの問題、特に $k$ が大きい場合の効率の低下などにより、任意の $k$ をサポートするという課題に直面しています。金融、法律、医療などのアプリケーションでは、既存のシステムに多大なオーバーヘッドを引き起こすほど大きな $k$ が必要となるため、これは重大な制限です。また、最新のロングコンテキスト モデルは一般に、より大きな検索セットを使用することでより高い精度を実現します。我々は、任意の上位 $k$ 検索をサポートする効率的なプライバシー保護 RAG サービスである P$^2$RAG を提案します。既存のシステムとは異なり、P$^2$RAG は候補ドキュメントのソートを回避します。代わりに、インタラクティブな二分法を使用して、上位 $k$ ドキュメントのセットを決定します。セキュリティのために、P$^2$RAG は 2 台の半誠実で非共謀のサーバー上で秘密共有を使用して、データ所有者のデータベースとユーザーのプロンプトを保護します。悪意のあるユーザーから防御するために制限と検証を強制し、データベースの情報漏洩を厳しく制限します。実験では、$k = 16$--$1024$ の場合、P$^2$RAG は最先端の PRAG より 3--300$\times$ 高速であることが示されています。

原文 (English)

P$^2$RAG: Efficient Privacy-Preserving RAG Service Supporting Arbitrary Top-$k$ Retrieval

Retrieval-Augmented Generation (RAG) enables large language models to use external knowledge, but outsourcing the RAG service raises privacy concerns for both data owners and users. Privacy-preserving RAG systems address these concerns by performing secure top-$k$ retrieval, which is typically implemented using secure sorting to identify relevant documents. However, existing systems face challenges supporting arbitrary $k$ due to their inability to change $k$, new security issues, and in particular, efficiency degradation with large $k$. This is a significant limitation because applications such as finance, law, and healthcare require a $k$ that is large enough to cause huge overhead for existing systems. Also, modern long-context models generally achieve higher accuracy with larger retrieval sets. We propose P$^2$RAG, an efficient privacy-preserving RAG service that supports arbitrary top-$k$ retrieval. Unlike existing systems, P$^2$RAG avoids sorting candidate documents. Instead, it uses an interactive bisection method to determine the set of top-$k$ documents. For security, P$^2$RAG uses secret sharing on two semi-honest non-colluding servers to protect the data owner's database and the user's prompt. It enforces restrictions and verification to defend against malicious users and tightly bounds the information leakage of the database. The experiments show that P$^2$RAG is 3--300$\times$ faster than the state-of-the-art PRAG for $k = 16$--$1024$.

2026-05-29 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

AgentLens: SWE エージェント評価におけるラッキー パスの問題を明らかにする

更新された要約は次のとおりです。 ソフトウェア エンジニアリング (SWE) エージェントの評価は、最終パッチがテストに合格するかどうかという 2 つの信号によって支配されます。この結果のみの考え方は、原則に基づいた解決策と混沌とした試行錯誤のプロセスを同等のものとして扱います。この等価性は経験的に誤りであることを示します。 60 の SWE ベンチ検証済みタスクで 8 つのモデル バックエンドからの 2,614 の OpenHands 軌跡を評価します。これらのうち、47 にはタスク レベルのプロセス参照を構築するのに十分な通過軌跡があり、1,815 の軌跡評価サブセットが得られます。このサブセットの通過軌跡のうち、10.7% は、回帰サイクル、ブラインド再試行、検証の欠落、または時間的に無秩序な探索、実装、検証など、ラッキー パスと呼ばれる動作を示しています。 SWE エージェントの軌跡をプロセスレベルで評価するためのフレームワークである AgentLens を導入し、品質スコア、廃棄信号、分岐点、および 47 のタスクレベルのプレフィックス ツリー アクセプタ (PTA) 参照が注釈付けされた 1,815 の軌跡のデータセットである AgentLens-Bench を定義します。 AgentLens は、同じタスクに渡された複数のソリューションをマージすることで PTA 参照を構築し、コンテキスト依存のインテント ラベラーを使用して、ツール ID だけではなく軌跡履歴に基づいて探索、実装、検証、またはオーケストレーションにアクションを割り当てます。 AgentLens-Bench では、品質スコアによってパスの軌跡が Lucky、Solid、Ideal の層に分割され、さらに Lucky パスが 5 つの反復メカニズムに分解されます。 8 つのモデル バックエンド全体で、Lucky 率の範囲は 0.5% ~ 23.2% であり、一部のモデルは、合格率ではなく品質スコアでランク付けすると、最大 5 ランク順位が変動します。 AgentLens-Bench アーティファクト、AgentLens SDK、分析ツールを含むプロジェクト リポジトリを間もなくリリースする予定です。

原文 (English)

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

Here is the updated abstract: Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view treats a principled solution and a chaotic trial-and-error process as equivalent. We show that this equivalence is empirically false. We evaluate 2,614 OpenHands trajectories from eight model backends on 60 SWE-bench Verified tasks. Of these, 47 have enough passing trajectories to construct task-level process references, yielding a 1,815-trajectory evaluation subset. Among passing trajectories in this subset, 10.7% exhibit behavior we call a Lucky Pass: regression cycles, blind retries, missing verification, or temporally disordered exploration, implementation, and verification. We introduce AgentLens, a framework for process-level assessment of SWE-agent trajectories, and define AgentLens-Bench, a dataset of 1,815 trajectories annotated with quality scores, waste signals, divergence points, and 47 task-level Prefix Tree Acceptor (PTA) references. AgentLens builds PTA references by merging multiple passing solutions for the same task, and uses a context-sensitive intent labeler to assign actions to Exploration, Implementation, Verification, or Orchestration based on trajectory history rather than tool identity alone. On AgentLens-Bench, the quality score separates passing trajectories into Lucky, Solid, and Ideal tiers and further decomposes Lucky Passes into five recurring mechanisms. Across the eight model backends, Lucky rates range from 0.5% to 23.2%, and some models move by as many as five rank positions when ranked by quality score instead of pass rate. We plan to release the project repository soon, including AgentLens-Bench artifacts, the AgentLens SDK, and the analysis tooling.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

JMed48k: 視覚言語モデル評価のための多職種の日本の医師免許ベンチマーク

視覚言語モデルを評価するための、多職種の日本の医療ライセンスベンチマークである JMed48k を紹介します。日本の厚生労働省がリリースした公式 PDF 資料から構築された JMed48k には、2005 年から 2025 年までの 11 の国家免許試験からの 48,862 件の試験問題と 20,142 枚の画像が含まれており、ビジュアル コンテンツは 8 種類の分類法に基づいて注釈が付けられています。このコーパスから、9,905 のテキストのみの質問と 2,579 の画像付きの質問を含む 12,484 のスコア付き質問を含む、最近 5 年間の評価サブセットである JMed48k-Eval を導き出します。私たちは 21 の独自のオープンソースの医療特化モデルを評価し、テキストのみのパフォーマンスと画像付きのパフォーマンスを別々に報告します。これらのサブセットには異なる質問が含まれているため、視覚コンテンツを削除する前後の画像を含む質問を評価するペアの画像削除監査をさらに導入して、4 つの回答遷移状態を調査します。監査では、独自のオープンソース モデルが画像から大幅に利益を得ているのに対し、医療に特化したシステムでは目に見える視覚的証拠の使用が限られており、画像の削除後も多くの正解が残っていることが示されています。独自モデルの中でも、正味の画像除去効果は、医師の質問で +5.7 ポイントから保健師の質問で +39.8 ポイントまで、職種によって 7 倍のばらつきがあります。私たちは、医療ライセンス設定における視覚言語モデルの再現可能な専門職層別評価をサポートするために、JMed48k をリリースします。

原文 (English)

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese Ministry of Health, Labour and Welfare, JMed48k contains 48,862 exam questions and 20,142 images from 11 national licensing examinations between 2005 and 2025, with visual content annotated under an 8-type taxonomy. From this corpus, we derive JMed48k-Eval, a recent five-year evaluation subset with 12,484 scored questions, including 9,905 text-only questions and 2,579 questions with images. We evaluate 21 proprietary, open-source, and medical-specific models, reporting text-only and with-image performance separately. Because these subsets contain different questions, we further introduce a paired image-removal audit that evaluates questions with images before and after removing visual content to explore four answer-transition states. The audit shows that proprietary and open source models gain substantially from images, whereas medical-specific systems show limited observable use of visual evidence, with many correct answers persisting after image removal. Even among proprietary models, the net image-removal effect varies sevenfold across professions, from +5.7 points on Physician questions to +39.8 points on Public Health Nurse questions. We release JMed48k to support reproducible, profession-stratified evaluation of vision-language models in medical licensing settings.

2026-05-29 13:00 JSTarXiv cs.AIビジネス/資金調達

多様な呼吸不全予測の前向き評価: 胸部 X 線撮影は EHR 信号を超えてパフォーマンスを向上させますか?

呼吸不全の早期予測は、集中治療室でのタイムリーな臨床介入にとって重要です。既存の電子健康記録 (EHR) ベースのモデルは、生理学的悪化を継続的に監視できますが、胸部 X 線写真 (CXR) に反映される肺の病態生理学を完全には捕捉できない可能性があります。この研究では、CXR 情報が EHR 信号のみを超えて侵襲的人工呼吸器の前向き予測を改善するかどうかを尋ねます。私たちは、構造化された EHR 時系列データと CXR 基盤モデル表現を統合するゲート付きマルチモーダル フレームワークを開発します。ゲーティング モジュールは、患者固有の臨床状況に基づいてイメージング機能の寄与を適応的に制御し、モデルが有益な場合にイメージング情報に選択的に依存できるようにします。私たちは、ICU患者の24時間以内の侵襲的人工呼吸器を予測するためのフレームワークを前向きに評価し、確立されたEHR専用モデル(Vent.io)、一致する臨床時点で得られた医師の予測、および代替のマルチモーダルバリアントと比較します。ゲート付きマルチモーダル モデルは、EHR のみのベースラインよりも高い識別を達成し、Vent.io の 0.752 と比較して、REMEDIS および MedInsight CXR 表現を使用した AUROC 値はそれぞれ 0.860 と 0.858 でした。医師の予測と比較して、マルチモーダルフレームワークは良好な特異性を維持しながら感度を大幅に向上させました。 EHR のみのモデルと比較して、マルチモーダル統合により特異性と陽性的中率が向上したことは、CXR 情報が選択された患者のリスク推定を精緻化できることを示唆しています。これらの発見は、画像処理を呼吸不全の予測に組み込むための実用的な戦略として、適応型マルチモーダル融合を裏付けるものです。

原文 (English)

Prospective evaluation of multimodal respiratory failure prediction: Do chest X-rays improve performance beyond EHR signals?

Early prediction of respiratory failure is critical for timely clinical intervention in intensive care units. Existing electronic health record (EHR)-based models can continuously monitor physiologic deterioration, but they may not fully capture pulmonary pathophysiology reflected in chest radiographs (CXRs). In this study, we ask whether CXR information improves prospective prediction of invasive mechanical ventilation beyond EHR signals alone. We develop a gated multimodal framework that integrates structured EHR time-series data with CXR foundation-model representations. The gating module adaptively controls the contribution of imaging features based on patient-specific clinical context, allowing the model to selectively rely on imaging information when it is informative. We prospectively evaluate the framework for predicting invasive mechanical ventilation within 24 hours in ICU patients and compare it with an established EHR-only model (Ventio), physician predictions obtained at matched clinical time points, and alternative multimodal variants. The gated multimodal models achieved higher discrimination than the EHR-only baseline, with AUROC values of 0.860 and 0.858 using REMEDIS and MedInsight CXR representations, respectively, compared with 0.752 for Ventio. Relative to physician predictions, the multimodal framework substantially improved sensitivity while maintaining favorable specificity. Compared with the EHR-only model, multimodal integration increased specificity and positive predictive value, suggesting that CXR information can refine risk estimation in selected patients. These findings support adaptive multimodal fusion as a practical strategy for incorporating imaging into prospective respiratory failure prediction.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

アライメントフロア: ペルソナのカスタマイズが安全な場合

多元的 AI の主な約束は行動の適応です。「創造的であること」や「徹底していること」などのペルソナ プロンプトにより、システムは多様なユーザーの価値観とコミュニケーション スタイルを尊重できます。しかし、アライメントが崩れる前に、モデルはどれだけのカスタマイズを吸収できるでしょうか?我々は、アライメントとカスタマイズのトレードオフに関する最初の対照研究を発表し、異なるアライメント強度を持つ 2 つのモデルで 5 つのタスクにわたって 7 つのペルソナ条件をテストしました (1,800 回の実行)。私たちは調整フロアを発見しました。強く調整されたモデル (クロード ソネット) では、ペルソナ プロンプトはお調子者にまったく影響しません。すべての条件で最大 15% が生成され、豊富なパーソナライゼーションが安全な安定したプラットフォームです。弱く調整されたモデル (Nova Lite) では、同じペルソナがお調子者を 5% から 50% に変更します。つまり、フロアが存在せず、カスタマイズが安全上の責任となります。驚くべきことに、協調性は最悪の犯罪者ではありません。外向性 (+20pp) とオープンネス (+15pp) は、より大きな劣化を引き起こします。建設的な発見は懐疑的防御です。批判的思考を持つペルソナは、弱いモデルでもお調子者を 5% に減少させます。これは、この研究における唯一最大の効果です。モデル間のペルソナ効果の伝達はほぼゼロ ($\rho = 0.006$) であり、アライメント テストはモデルごとに行う必要があることを意味します。私たちは、設計原則としてアライメント フロアを提案します。つまり、ペルソナのカスタマイズを展開する前にアライメント フロアを測定し、安全性を重視したペルソナをユーザー向けのペルソナの下に重ねて、配置を損なうことなくパーソナライゼーションを可能にします。

原文 (English)

The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs

Telling an LLM to "be enthusiastic" raises its sycophancy rate from 30\% to 50\% on a lightly-aligned model, but has zero effect on a strongly-aligned one. We define this gap as the alignment floor, $\Delta_{\text{floor}}(m)=\max_pS(m,p)-\min_pS(m,p)$, the range of sycophancy rates a model produces across persona conditions, and treat sycophancy as a persona-conditional property rather than a fixed model property. Pluralistic AI relies on behavioral adaptation via persona prompts like "be creative" or "be thorough", which let systems respect diverse user values and communication styles; the safety question is how much customization a given model can absorb before its truthfulness shifts. We present a controlled case study contrasting a strongly-aligned RLHF + Constitutional-AI model (Claude Sonnet 4.6) with a more lightly-aligned model (Amazon Nova Lite), spanning seven persona conditions and five tasks for 1800 total runs. An existence-pair result motivates per-model auditing: there is at least one strongly-aligned model with $\Delta_{\text{floor}}=5$pp (within 5pp of the 15\% control rate) and at least one lightly-aligned model with 45pp (5\%--50\% range). On the lightly-aligned model, all five Big Five personas increase sycophancy over control, and counterintuitively Agreeableness produces the smallest increase, not the largest. The single largest effect in the study is constructive: a Skeptic persona reduces sycophancy by 25pp on the lightly-aligned model, and is the only persona that instructs resistance against user claims rather than engagement with them, suggesting a directionality account. Cross-model transfer of persona effects is near-zero, so persona-alignment testing must be per-model. We propose $\Delta_{\text{floor}}$ as a deployment-time audit metric: measure it on a small persona panel before deploying persona customization.

2026-05-29 07:30 JSTITmedia AI+ビジネス/資金調達

「日本は製造業のパワーハウス」、IFSが産業AI投資を急拡大する理由

IFSジャパンは記者会見を開催し、日本市場への投資継続とパートナーシップ強化の方針を説明した。日本IBMらとの戦略的協業を通じ、製造業などアセット集約型産業のAI実装とDXを支援する。

2026-05-29 03:52 JSTTechCrunch AILLM/生成AIビジネス/資金調達

Anthropic raises $65 billion, nears $1T valuation ahead of IPO

Anthropic has closed a $65 billion Series H round at a $965 billion post-money valuation, marking what could be the AI startup's final priv…

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

秘密がある? LLM エージェントはそれを守れない: マルチエージェント システムにおけるプライバシーの評価

LLM の安全性評価では主にモデルを単独でテストしますが、配備された AI エージェントは他のエージェントと並んで永続的な社会環境内で動作することが増えています。私たちは、何千人もの LLM エージェントがシミュレートされた 1 か月間にわたってコミュニティ間で対話する Moltbook スタイルのシミュレーション プラットフォームを導入し、それを使用して、さまざまな程度の社会的圧力の下で下流の安全上の懸念としてプライバシーを評価します。シングルターンからマルチターンへの社会的評価の移行により、プライバシー侵害が増幅されること(OpenAI モデル全体で、CIMemories 19.95% から Ours 45.30%)、漏洩は社会的に伝染し、ピアが機密情報を開示するのを観察したエージェントは機密情報を開示する可能性が 8 倍高く、明示的なプライバシーに関する指示はこの影響を軽減するものの排除はせず、保護策を講じたとしても漏洩率が 37.8% を超えることがわかりました。私たちの調査結果は、静的チャットベースの安全性ベンチマークは、エージェント導入におけるリスクを体系的に過小評価していること、また、社会的コンテキストだけで、単一ターンの評価では決して表面化しない機密情報の開示を引き出すのに十分であることを示唆しています。

原文 (English)

Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems

LLM safety evaluations predominantly test models in isolation, yet deployed AI agents increasingly operate within persistent social environments alongside other agents. We introduce a Moltbook-style simulation platform where thousands of LLM agents interact across communities over a simulated month, and use it to evaluate privacy as a downstream safety concern under varying degrees of social pressure. We find that shifting from single turn to multi turn social evaluation amplifies privacy violations (CIMemories 19.95% to Ours 45.30% across OpenAI models), that leakage is socially contagious, with agents 8 times more likely to disclose sensitive information after observing a peer do so, and that explicit privacy instructions reduce but do not eliminate this effect, leaving leakage rates above 37.8% even with safeguards. Our findings suggest that static chat based safety benchmarks systematically underestimate risks in agentic deployment, and that social context alone is sufficient to elicit sensitive disclosures that single turn evaluations would never surface.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM-as-a-Judge 評価のための固定予算のクラスター対応標準: マルチホップ RAG ストレス テスト

検索拡張生成 (RAG) システムは、大規模言語モデル (LLM) にどちらの答えが優れているかを判断させることによって比較されることがよくあります。マルチホップ RAG の場合、これはモデリングの問題と同じくらい測定の問題になります。同じスコアは、検索品質、回答の長さ、語彙の重複、またはクラスター化されたデータを無視する統計テストを反映する可能性があります。これらの選択が明確にされると何が起こるのかを尋ねます。私たちは、RAG における LLM-as-a-judge の比較のための最小測定標準を提案します。この標準では、上位 100 位の候補者プール、証拠予算、回答上限、ジェネレーター、およびプロンプトが修正されています。また、事前に登録された仮説、クラスターを意識した推論、可能な場合は正確なクラスターの符号反転チェック、および第 2 判定の複製も必要です。クラスター化されたベンチマークは進捗状況を誇張する可能性があります。現場ではこの標準を採用する必要があります。コンピューター サイエンス/機械学習 (CS/ML) および材料科学における 400 のマルチホップ質問に対して、進化的証拠セレクターである Genetic Algorithm Decoder for Multi-hop Evidence Composing (GADMEC) を使用してストレス テストを行います。このプロトコルは経験的な物語を変えます。二項テストでは、4 つの意味ベースラインの比較がすべて重要であるように見えます。クラスター認識推論では、ボンフェローニ有意な結果が 1 つだけ残ります。 BM25 は同じ予算内で純粋な意味論的な GADMEC を破り、語彙と意味論的なハイブリッドが CS/ML で回復し、材料科学の差を縮めます。

原文 (English)

A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test

Retrieval-augmented generation (RAG) systems are often compared by asking a large language model (LLM) judge which answer is better. For multi-hop RAG, this has become a measurement problem as much as a modeling problem: the same score can reflect retrieval quality, answer length, lexical overlap, or a statistical test that ignores clustered data. We ask what happens when these choices are made explicit. We propose a minimum measurement standard for LLM-as-a-judge comparisons in RAG. The standard fixes the top-100 candidate pool, evidence budget, answer cap, generator, and prompt; it also requires pre-registered hypotheses, cluster-aware inference, an exact cluster sign-flip check when feasible, and second-judge replication. Clustered benchmarks can overstate progress; the field should adopt this standard. We stress-test it with Genetic Algorithm Decoder for Multi-hop Evidence Composition (GADMEC), an evolutionary evidence selector, on 400 multi-hop questions in computer science/machine learning (CS/ML) and Materials Science. The protocol changes the empirical story. A binomial test makes all four semantic-baseline comparisons look significant; cluster-aware inference leaves only one Bonferroni-significant result. BM25 beats pure semantic GADMEC under the same budget, while a lexical-semantic hybrid recovers in CS/ML and narrows the Materials Science gap.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

FundaPod: AI 支援のファンダメンタル投資調査のためのナレッジ グラフ メモリを備えたマルチペルソナ エージェント ポッド プラットフォーム

大規模言語モデル (LLM) は金融分野での適用が増えていますが、既存の研究のほとんどは取引シグナルや予測を中心とした財務 NLP タスクに重点を置いています。対照的に、制度的基礎研究では、人間のアナリストまたは AI エージェントが証拠を収集し、ビジネス推進要因を特定し、競合する視点を比較し、投資メモを作成する必要があります。その広範な目標は、単に結果を予測することではなく、投資知識の累積的な発展に貢献しながら、透明性、再利用可能、検証可能な投資計画を作成することです。 AI 支援のファンダメンタルズ投資調査のためのマルチペルソナ エージェント プラットフォームである FundaPod を紹介します。私たちは、基礎研究は人間中心の意思決定支援タスクであり、取引シグナルの生成とは質的に異なるため、独立性を維持するアーキテクチャの方が適していると主張します。 FundaPod では、バリュー投資家やマクロ戦略家など、さまざまなペルソナを持つ AI エージェントが、共有の出所契約に基づいて独立して調査を実施します。その後、彼らの意見の相違は、知識グラフ記憶システムを通じて人間のポートフォリオ マネージャー (PM) による裁定のために事後的に表面化されます。この論文は、設計科学の実践と認知的分離と人間と機械の協調の理論に基づいた、基礎研究をサポートする人間と AI のハイブリッド システムの 5 つの設計原則を提供します。また、4 つのアーキテクチャ メカニズムについても説明します。1 つは一般投資家の資料を展開可能なエージェントに変えるペルソナ蒸留パイプラインです。プランナーが型指定されたタスク グラフを導出できるようにする宣言型スキル レジストリ。メモの主張を検証可能な情報源に結び付ける根拠のある証拠モデル。そしてティッカー、メモ、アナリスト、テーマを結び付けるナレッジグラフ「第二の脳」。完全なケーススタディとペルソナベースのメモの比較を通じてアーキテクチャを実証します。

原文 (English)

FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research

Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks centered on prediction. Institutional fundamental research, by contrast, requires human analysts or AI agents to gather evidence, identify business drivers, compare competing viewpoints, and generate investment memos. Its broader goal is not merely to predict outcomes, but to produce investment plans that are transparent, reusable, and verifiable, while contributing to the cumulative development of investment knowledge. We present FundaPod, a multi-persona agent platform for AI-assisted fundamental investment research. We argue that fundamental research is a human-centric decision-support task that is qualitatively distinct from trading-signal generation, and is therefore better served by an independence-preserving architecture. In FundaPod, AI agents with different personas, such as value investors or macro strategists, conduct research independently under a shared provenance contract. Their disagreements are then surfaced post hoc for adjudication by the human portfolio manager (PM) through a knowledge-graph memory system. This paper contributes five design principles for human-AI hybrid systems supporting fundamental research, grounded in design-science practice and theories of cognitive isolation and human-machine coordination. It also describes four architectural mechanisms: a persona distillation pipeline that turns public investor materials into deployable agents; a declarative skill registry that lets the planner derive typed task graphs; a grounded evidence model that links memo claims to verifiable sources; and a knowledge-graph "second brain" that connects tickers, memos, analysts, and themes. We demonstrate the architecture through a complete case study and a persona-based memo comparison.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

LLM エージェントの機能を評価するための統一フレームワーク

LLM がエージェントとして導入されることが増えているため、そのエージェント機能の信頼できる評価が不可欠になっています。ただし、報告されるベンチマーク スコアは、多くの場合、モデルの機能と、各ベンチマークに含まれる実装の選択肢を合わせて反映するため、クロスベンチマークの結果を基礎となるモデルの正確な測定値として解釈することが困難になります。この研究では、LLM エージェントの機能を公正に評価するための統一フレームワークを紹介します。統合された構成システムによって駆動されるこのフレームワークは、標準化された命令、ツール、環境の形式に多様なベンチマークを統合し、制御可能なサンドボックス内の固定 ReAct スタイル アーキテクチャを通じてエージェントを実行します。また、フレームワークの効果と環境の効果を個別に分析できるように、揮発性のライブ環境を厳選されたスナップショットに置き換えるオプションのオフライン設定を提供します。これに基づいて、各ベンチマークの元のタスクの成功基準に基づいて評価方法を統一するとともに、リソース消費に関する統一された指標と、意思決定レベルおよび実行レベルの失敗の属性に関する分類を導入します。このフレームワーク内で、シングルエージェント、マルチエージェント、およびセーフティクリティカルなシナリオにわたる 24 のドメインにわたる 7 つの広く使用されているベンチマークを適応させ、15 のモデルで 400,000 のロールアウトと 50 億のトークンにわたる大規模な実証分析を実施します。結果は、足場の選択と環境の変動性がベンチマークの結果を両方向に実質的に変化させ、フレームワークおよび環境によって引き起こされるアーティファクトから本質的な LLM 機能を解きほぐすことをフレームワークが可能にすることを示しています。さらに、安全性が重要なドメインの安全なテストベッドとしての拡張性を実証します。コードとベンチマークは、https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities、https://huggingface.co/AgentFramework/Unified_Farmework で入手できます。

原文 (English)

A Unified Framework for the Evaluation of LLM Agentic Capabilities

As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores often jointly reflect model capability and the implementation choices each benchmark is packaged with, making cross-benchmark results difficult to interpret as clean measurements of the underlying model. In this work, we present a unified framework for the fair evaluation of LLM agentic capabilities. Driven by a unified configuration system, the framework integrates diverse benchmarks into a standardized instruction--tool--environment format, executes agents through a fixed ReAct-style architecture within a controllable sandbox, and provides an optional offline setting that replaces volatile live environments with curated snapshots, so that framework effects and environment effects can be analyzed separately. Building on this, we unify the evaluation methodology under each benchmark's original task-success criteria, while introducing unified metrics for resource consumption and a taxonomy for decision- and execution-level failure attribution. Within this framework, we adapt 7 widely used benchmarks spanning 24 domains across single-agent, multi-agent, and safety-critical scenarios, and conduct a large-scale empirical analysis over 400K rollouts and 5B tokens on 15 models. The results show that scaffold choice and environmental volatility materially shift benchmark outcomes in both directions, allowing our framework to disentangle intrinsic LLM capabilities from framework- and environment-induced artifacts. We further demonstrate its extensibility as a secure testbed for safety-critical domains. Codes and benchmarks at are available at https://github.com/whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities, https://huggingface.co/AgentFramework/Unified_Farmework.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

MIRA: 医療情報対応監査のバイリンガル ベンチマーク

一般向けの健康情報を提供するために大規模言語モデル (LLM) がますます使用されていますが、既存の安全性評価では、同じ質問に対するさまざまなユーザーの表現にわたって回答が同等の医療情報を保持しているかどうかが見落とされています。これに対処するために、LLM がユーザー側の言語、登録、ヘルス リテラシー シグナル全体で同等の医療情報を提供しているかどうかを評価するバイリンガルの管理されたベンチマークである Medical Information Response Audit (MIRA) を導入します。 MIRA には、医学的に検討された低リスクの健康に関する 60 の質問から作成された 4,320 のプロンプトが含​​まれています。 5 つの主流 LLM にわたって、モデルはすべての医学的質問に答えましたが、健康リテラシーが低い信号への応答では一貫してより多くの重要な情報が省略され、具体的な次のステップが少なくなり、独立した判断に対するサポートが少なくなりました。このパターンを差分情報希釈 (DID) と呼びます。言語の影響は、英語以外のプロンプトで一律に悪化するのではなく、モデルに固有です。 300 件の実世界の健康クエリとの比較により、ランク順の妥当性の予備的な証拠が得られます。知識に基づいた緩和プロンプトにより、ほとんどのモデルで情報の希薄化が軽減され、情報不足の単純化が最も大きく減少したのはクロード (約 8%) とクウェン (約 6%) でした。

原文 (English)

MIRA: A Bilingual Benchmark for Medical Information Response Audit

Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent judgment. We term this pattern Differential Information Dilution (DID). Language effects are model-specific rather than uniformly worse for non-English prompts. A comparison with 300 real-world health queries provides preliminary evidence of rank-order validity. A knowledge-guided mitigation prompt reduces information dilution for most models, with the largest reductions in underinformative simplification observed for Claude (~8%) and Qwen (~6%).

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

PetroBench: 石油工学における大規模言語モデルのベンチマーク

大規模言語モデルは石油業界でますます適用されており、ドメイン固有の評価フレームワークの必要性が強調されています。この研究では、データの前処理、品質フィルタリング、マルチモデル検証の 3 段階のプロセスを含む、石油工学における LLM のベンチマークを開発します。専門家のレビューを使用して、強力なドメイン関連性と識別機能を備えた標準化された質問バンクが構築されました。このベンチマークは生産、貯留層、掘削工学を対象としており、多肢選択、正誤、用語の定義、短答形式にわたる 1,200 の質問が含まれています。 8 つの主流 LLM が統合 API 環境下で評価されました。結果は、モデルが客観的な質問よりも主観的な質問の方が優れたパフォーマンスを示し、事実知識の識別における弱点を示しています。多肢選択式質問と正誤質問の最高精度は、それぞれ 65.3% と 74.3% でした。 Gemini-3-Pro、Kimi-K2.5、および Claude-Opus-4.6-Thinking は、72% ~ 74% という最高の総合スコアを達成しました。モデルは生産エンジニアリングで最も優れたパフォーマンスを発揮しましたが、貯留層エンジニアリングでは最も劣っていました。中国のモデルは多肢選択問題で優位性を示しましたが、国際モデルは短答式の質問でわずかに優れた結果を示しました。このベンチマークは、石油工学における LLM の評価と導入のための再現可能で実用的なリファレンスを提供します。

原文 (English)

PetroBench: A Benchmark for Large Language Models in Petroleum Engineering

Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study develops a benchmark for LLMs in petroleum engineering, including a three-stage process of data preprocessing, quality filtering, and multi-model validation. Using expert review, a standardized question bank with strong domain relevance and discriminative capability was constructed. The benchmark covers production, reservoir, and drilling engineering, with 1,200 questions across multiple-choice, true or false, term definition, and short-answer formats. Eight mainstream LLMs were evaluated under a unified API environment. Results show that models performed better on subjective than objective questions, indicating weaknesses in factual knowledge discrimination. The highest accuracies for multiple-choice and true or false questions were 65.3% and 74.3%, respectively. Gemini-3-Pro, Kimi-K2.5, and Claude-Opus-4.6-Thinking achieved the best overall scores of 72%-74%. Models performed best in production engineering and weakest in reservoir engineering. Chinese models showed advantages in multiple-choice questions, while international models performed slightly better in short-answer questions. The benchmark provides a reproducible and practical reference for evaluating and deploying LLMs in petroleum engineering.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

関連性は保証されていない: 引用された RAG の証拠と力の校正

引用された RAG の評価では、目に見える情報源が根拠となる信号として扱われることがよくありますが、実際の話題に関連した引用であっても、添付された文言の正当性が不十分である可能性があります。私たちはこの診断の失敗を引用ロンダリングとして研究しています。つまり、関連する情報源が過度の主張の根拠として提示されています。証拠と力の校正のための対照ストレステストである FORCEBENCH を紹介します。各項目は引用箇所を固定し、証拠に基づいて調整された主張と、関係性、様相、範囲、時間的妥当性、数値的特異性という 5 つの操作軸にわたる局所的な力によって引き起こされた変形とを組み合わせます。調整された評価者は、証拠に基づいて調整された主張をより高く評価する必要があります。ヘッドライン実験では、固定の局所性フィルター処理された 198 ペアの評価セットを使用します。引用存在の健全性チェックは設計上、有益ではありません。トークンとエンティティの重複は、依然としてペアの 32.8 ~ 36.4% で単調性に違反しています。報告された4人のモデル裁判官全体で、標準的な一般的なサポートのプロンプトはこの力校正ストレステストには不十分であり(合計MVR 47.2%)、明示的な令状強度のプロンプトはMVRを24.5%に低下させますが、依然として不完全です。ベンチマーク、プロンプト、出力、およびプラグイン パイプラインをリリースすることで、引用評価者が従来のサポート メトリックとともに単調性違反率と力感度を報告できるようになります。

原文 (English)

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Cited RAG evaluation often treats visible sources as a grounding signal, but a real, topically relevant citation can still under-warrant the attached wording. We study this diagnostic failure as citation laundering: a related source is presented as warrant for an over-strong claim. We introduce FORCEBENCH, a contrastive stress test for evidence-force calibration. Each item holds a cited passage fixed and pairs an evidence-calibrated claim with a localized force-raised variant across five operational axes: relation, modality, scope, temporal validity, and numeric specificity. A calibrated evaluator should score the evidence-calibrated claim higher. Headline experiments use a fixed, locality-filtered 198-pair evaluation set. A citation-presence sanity check is uninformative by design; token and entity overlap still violate monotonicity on 32.8--36.4% of pairs. Across four reported model judges, standard generic support prompting is insufficient for this force-calibration stress test (aggregate MVR 47.2%), while explicit warrant-strength prompting lowers MVR to 24.5% but remains imperfect. We release the benchmark, prompts, outputs, and plug-in pipeline so citation evaluators can report monotonicity violation rate and force sensitivity alongside conventional support metrics.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

ルック・オン・デマンド: マルチモーダル推論における視覚的証拠取得のための認知スケジューリング フレームワーク

既存のマルチモーダル推論アプローチは、主に 2 つのパラダイムに従います。推論の前に視覚入力をテキストに変換するか、統一された視覚言語表現空間内でエンドツーエンドの推論を実行します。経験的な進歩にもかかわらず、両方のパラダイムには根本的な構造上の限界があります。前者は静的なビジュアルからテキストへの変換に依存しているため、圧縮され、細かいビジュアルの詳細が失われる傾向があります。後者は、共同最適化と注意メカニズムによって引き起こされる言語支配の傾向があり、推論中の視覚的証拠に対する忠実性が体系的に弱くなることにつながります。この研究では、視覚的証拠を推論プロセスにいつどのように導入するかが中心的な課題であると主張しています。この洞察に動機づけられて、我々は、言語モデルがタスク関連の視覚的証拠を取得するために独立した視覚認識モジュールをいつ呼び出すかを決定することによって推論プロセスを制御する、マルチモーダル推論フレームワークである CSMR を提案します。複数のマルチモーダル推論ベンチマークにわたる実験では、CSMR がゼロショット設定の下で精度において代表的なベースライン手法を常に上回っていることが示されています。さらなる実験分析により、これらの利点は主に提案された認知スケジューリング メカニズムから生じることが確認されています。

原文 (English)

Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning

Existing multimodal reasoning approaches predominantly follow two paradigms: converting visual inputs into text prior to reasoning, or performing end-to-end reasoning within a unified vision-language representation space. Despite their empirical progress, both paradigms suffer from fundamental structural limitations. The former relies on static visual-to-text conversion, which tends to compress and lose fine-grained visual details. The latter is prone to linguistic dominance induced by joint optimization and attention mechanisms, leading to systematically weakened faithfulness to visual evidence during reasoning. In this work, we argue that a central challenge is how and when visual evidence is introduced into the reasoning process. Motivated by this insight, we propose CSMR, a multimodal reasoning framework in which a language model controls the reasoning process by deciding when to invoke an independent visual perception module to acquire task-relevant visual evidence. Experiments across multiple multimodal reasoning benchmarks show that CSMR consistently outperforms representative baseline methods in accuracy under a zero-shot setting. Further experimental analysis confirms that these advantages primarily arise from the proposed cognitive scheduling mechanism.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ResearchLoop: AI 支援研究のための証拠ゲート型コントロール プレーン

AI を利用した研究では、アイデア出し、実装、評価、原稿執筆が 1 つのインタラクティブなループに圧縮されます。この圧縮は便利ですが、出版リスクも生み出します。紙上の主張は監査するよりも述べるのが容易になる可能性があります。 AI 支援による計算研究のための証拠ゲート型コントロール プレーンである ResearchLoop を紹介します。 ResearchLoop は、リサーチ質問、タスク契約、証拠オブジェクト、請求元帳、クローズアウト、および紙バインディングを永続的なプロジェクト状態として扱い、ここではリポジトリ支援のランタイムとして実現されます。この技術レポートは、完全なプロトコル仕様、状態モデル、移行ルール、クレーム受付アルゴリズム、および洞察複合メカニズムを提供します。また、9 つのバージョン (V0 ~ V9) にわたる完全な実験記録も報告しています。これには、セルフホスティングのケース スタディ、コンポーネント アブレーションを使用した制御されたタスク スイートの研究、数学オリンピックの評価、公式の生成コード ハーネスを使用して評価された補足的な SciCode 境界実験が含まれます。すべてのアーティファクト、マニフェスト、検証レポートはプロジェクト リポジトリに保存されます。

原文 (English)

ResearchLoop: An Evidence-Gated Control Plane for AI-Assisted Research

AI-assisted research compresses ideation, implementation, evaluation, and manuscript writing into a single interactive loop. This compression is useful, but it also creates a publication risk: paper claims can become easier to state than to audit. We present ResearchLoop, an evidence-gated control plane for AI-assisted computational research. ResearchLoop treats research questions, task contracts, evidence objects, claim ledgers, closeouts, and paper bindings as durable project state, realized here as a repository-backed runtime. This technical report provides the complete protocol specification, state model, transition rules, claim-admission algorithm, and insight-compounding mechanism. It also reports the full experimental record spanning nine versions (V0--V9), including a self-hosting case study, a controlled task-suite study with component ablations, a mathematical olympiad evaluation, and a supplementary SciCode boundary experiment evaluated with the official generated-code harness. All artifacts, manifests, and verification reports are preserved in the project repository.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

Picid: タスクとドメイン全体で再現可能な PHM のためのモジュール式評価インフラストラクチャ

予後診断と健康管理 (PHM) の進歩は、タスク、データセット、アプリケーション ドメイン全体にわたって標準化され再利用可能な評価手法が欠如しているために妨げられています。データ分割、前処理、ラベルの位置合わせ、時間ウィンドウ処理、メトリクスなどの主要なプロトコルの選択は暗黙的であるかアドホックに実装されることが多いため、報告された結果を再現したり比較したりするのは困難なことがよくあります。 PHM 評価パイプラインを明示的、実行可能、再現可能なプロトコルとして形式化するモジュール式評価インフラストラクチャである \picid を導入します。 \picid は、明確に定義された抽象化を通じて、多様な PHM 設定にわたって柔軟性を維持しながら、決定論的で漏洩の安全なデータセット構築を強制します。このフレームワークは、統合インターフェイスを通じて障害の検出、診断、予測をサポートしており、プロトコルの不変条件に違反することなく新しいデータセットやモデル クラスに拡張できます。 \picid は、データ コントラクトと評価境界を標準化することにより、診断 (分類) と予測 (回帰) にわたるタスク間の公平な比較も可能にし、異種設定間で同一のモデル ファミリを一貫して評価できるようにします。私たちは、バッテリー、ベアリング、ターボファン エンジン、油圧、濾過システム、建物にわたる 12 のデータセットに対する 13 のモデルの経験的評価を通じて \picid を実証します。この作業により、PHM における標準化された公平かつ再現可能な評価のための再利用可能な基盤が確立されます。

原文 (English)

Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains

Progress in Prognostics and Health Management (PHM) is hindered by the lack of standardized and reusable evaluation practices across tasks, datasets, and application domains. Reported results are often difficult to reproduce and compare, as key protocol choices, such as data splits, preprocessing, label alignment, temporal windowing, and metrics, are often implicit or implemented ad hoc. We introduce \picid, a modular evaluation infrastructure that formalizes the PHM evaluation pipeline as an explicit, executable, and reproducible protocol. Through well-defined abstractions, \picid enforces deterministic, leakage-safe dataset construction while remaining flexible across diverse PHM settings. The framework supports fault detection, diagnostics, and prognostics through a unified interface and can be extended to new datasets and model classes without violating protocol invariants. By standardizing data contracts and evaluation boundaries, \picid also enables fair cross-task comparisons across diagnostics (classification) and prognostics (regression), allowing identical model families to be evaluated consistently across heterogeneous settings. We demonstrate \picid through an empirical evaluation of thirteen models on twelve datasets spanning batteries, bearings, turbofan engines, hydraulics, filtration systems, and buildings. This work establishes a reusable foundation for standardized, fair and reproducible evaluation in PHM.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

低リソースのコンテキストにおける AI のベンチマーク: リーダーボードを超えて考える

既存の AI 評価手法では、運用上の制約がモデルの品質と同じくらい使いやすさを左右する低リソース環境でシステムが実際にどのように動作するかを把握できないことがよくあります。音声、チャット/RAG、およびビジョン システムにわたる既存のベンチマーク ファミリの構造化分析を通じて、ラボでの評価実践と低リソース環境での実際の展開条件との間の重大なギャップを特定します。私たちは、意味のある評価単位は分離されたモデルではなく展開されたシステムであり、効果的な評価フレームワークはタスクのパフォーマンスとノイズの多い入力、コードスイッチング、断続的な接続、ローエンドハードウェア、ドメインシフトなどの展開条件を統合する必要があると主張します。同時に、ベンチマークは、異なるアプリケーション クラスには、運用の違いを曖昧にする単一の集計スコアではなく、個別の評価プロファイルが必要であることを認識する必要があります。実際の意思決定をサポートするために、展開コンテキストに敏感でありながら、システムやアプリケーションの種類全体での比較可能性を維持する共有レポート フレームワークを提案します。最後に、標準化された 1 ページのベンチマーク カード、展開プロファイル、障害処理手順と人間による監視メカニズムの明示的な文書など、政策立案者、寄付者、実装者向けの簡潔で実用的なレポート作成物の必要性を強調します。

原文 (English)

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape usability as much as model quality. Through a structured analysis of existing benchmark families across speech, chat/RAG, and vision systems, we identify critical gaps between laboratory evaluation practices and real-world deployment conditions in low-resource environments. We argue that the meaningful unit of assessment is the deployed system rather than an isolated model and that effective evaluation frameworks must integrate task performance with deployment conditions such as noisy inputs, code-switching, intermittent connectivity, low-end hardware, and domain shift. At the same time, benchmarks should recognize that different application classes require distinct evaluation profiles rather than a single aggregate score that obscures operational differences. To support practical decision-making, we propose a shared reporting framework that preserves comparability across systems and application types while remaining sensitive to deployment context. Finally, we emphasize the need for concise and actionable reporting artifacts for policymakers, donors, and implementers, including standardized one-page benchmark cards, deployment profiles, and explicit documentation of failure handling procedures and human oversight mechanisms.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM を使用した充足可能性の解決: 推論能力の適合ペア評価

大規模言語モデル (LLM) は、暗黙的にブール充足可能性 (SAT) に帰着するタスクにますます使用されていますが、SAT に関する推論能力は依然として不明です。我々は、表現不変推論を調査するために、2-SAT および 3-SAT に関する LLM の体系的な研究を、頂点カバーと離散 3D パッキングという 2 つの標準的縮約とともに提示します。まず、精度、精度、再現率、F1、および SAT 位相遷移設定などの従来の指標を使用してモデルを評価します。これらのメトリクスは誤解を招く可能性があることがわかりました。多くのモデルは、充足可能な式を過剰予測することで高いスコアを取得し、3-SAT しきい値付近で古典的な簡単-困難-簡単な特徴を再現できず、変数の数が増加するにつれて急激に低下します。この問題に対処するために、最小限に異なる充足可能インスタンスと充足不可能なインスタンスに基づくペア式プロトコルと、各ペアの両方のメンバーが正しく分類されることを必要とする正確な微分率 (ADR) を導入します。 ADR は推論指向のモデルをヒューリスティックなモデルから分離し、証人の妥当性と相関させます。 CNF を超えて、CNF を頂点カバーに変換し、3-SAT を個別の 3D パッキングに変換することで、相互表現の一貫性をテストします。 CNF および対応するグラフまたはパッキング インスタンスに関するモデルの決定は、インスタンスの 80% 以上でほとんどのモデルに一致しており、表現全体で安定した決定ルールが示唆されています。全体として、我々の結果は、SAT が LLM 推論の保守的なプローブであり、ADR と組み合わせた評価が従来の指標よりも忠実で表現に堅牢な評価を提供することを示しています。

原文 (English)

Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability

Large language models (LLMs) are increasingly used for tasks that implicitly reduce to Boolean satisfiability (SAT), yet their reasoning ability on SAT remains unclear. We present a systematic study of LLMs on 2-SAT and 3-SAT, together with two canonical reductions, Vertex Cover and discrete 3D packing, to probe representation-invariant reasoning. We first evaluate models using conventional metrics, including accuracy, precision, recall, and F1, as well as the SAT phase-transition setting. We find that these metrics can be misleading: many models obtain high scores by over-predicting satisfiable formulas, fail to reproduce the classical easy-hard-easy signature around the 3-SAT threshold, and degrade sharply as the number of variables grows. To address this problem, we introduce a paired-formula protocol based on minimally different satisfiable and unsatisfiable instances, together with Accurate Differentiation Rate (ADR), which requires both members of each pair to be classified correctly. ADR separates reasoning-oriented models from heuristic ones and correlates with witness validity. Beyond CNF, we test cross-representation consistency by converting CNF to Vertex Cover and 3-SAT to discrete 3D packing. Model decisions on CNF and on the corresponding graph or packing instances agree for most models on more than 80 percent of instances, suggesting stable decision rules across representations. Overall, our results show that SAT is a conservative probe for LLM reasoning, and that paired evaluation with ADR provides a more faithful and representation-robust assessment than conventional metrics.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

統計的に真剣であることの重要性: GSM シンボリックの重要な再評価

GSM-Symbolic ベンチマーク (Mirzadeh et al., 2025) は、GSM8K 問題のテンプレート生成バリアントでテストした場合、25 の大規模言語モデル (LLM) 全体で一貫したパフォーマンスの低下を報告し、モデルには真の推論機能が欠けていると結論付けました。私たちは、この結論は不安定な統計的根拠に基づいていると主張します。質問ごとの変量効果を備えた一般化線形混合モデルを使用して 20 の無重みモデルを再評価すると、元のプロンプト形式で統計的に有意なパフォーマンス変化を示したのは半分だけであることがわかりました。さらに、これまで認められていなかった要因も特定しました。つまり、メインの GSM-Symbolic データセットには、GSM-Base と比較して問題テキスト内のより大きな整数の体系的にシフトされた分布が含まれており (K-S 統計量 = 0.12、p < 0.001)、これは原著者の主張と矛盾しています。この大きな数の効果を制御することは、残りのケースの約半分で重要性を説明します。統計的に有意なパフォーマンスデルタを持つモデルの中で、変数結合の脆弱性、算術的制限、デュアルタスク干渉など、モデル固有の明確な障害プロファイルを特定しました。これは、LLM 推論に関する包括的な主張が統計的に時期尚早であり、機構的に誤解を招くものであることを強調します。

原文 (English)

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-generated variants of GSM8K problems, concluding that the models lack genuine reasoning capabilities. We argue that this conclusion rests on shaky statistical ground. Re-evaluating 20 open-weight models using Generalised Linear Mixed Models with per-question random effects, we find that only half exhibit statistically significant performance changes under the original prompt format. Moreover, we identify a previously unacknowledged factor: the main GSM-Symbolic dataset contains a systematically shifted distribution of larger integers in problem texts relative to GSM-Base (K-S statistic = 0.12, p < 0.001), contradicting the original authors' claims. Controlling for this large number effect accounts for significance in roughly half the remaining cases. Among models with statistically significant performance deltas, we identify distinct, model-specific failure profiles - including fragility of variable binding, arithmetic limitations, and dual-task interference - underscoring that blanket claims about LLM reasoning are both statistically premature and mechanistically misleading.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

宇宙運用のための検索拡張生成および言語モデルの系統的評価

宇宙活動の急速な拡大により、技術文書、運用ガイドライン、科学文献が前例のないほど蓄積され、宇宙運用におけるタイムリーな意思決定に課題が生じています。宇宙運用における効果的な管理には、膨大で異種の情報ソースを効率的に処理できるツールが必要です。このペーパーでは、ドメイン固有のドキュメントから実用的な知識を抽出および合成するための大規模言語モデル (LLM) と情報検索技術を組み合わせた、検索拡張生成 (RAG) パイプラインのパフォーマンスを体系的に評価します。さまざまな検索戦略、埋め込みモデル、LLM 回答を比較して、情報の正確性、関連性、信頼性に対するそれらの影響を評価します。私たちの結果は、RAG パイプラインが知識へのアクセスを大幅に強化し、不確実性を軽減し、複雑な宇宙運用における意思決定をサポートできることを示しています。

原文 (English)

A Systematic Evaluation of Retrieval-Augmented Generation and Language Models for Space Operations

The rapid expansion of space activities has led to an unprecedented accumulation of technical documentation, operational guidelines, and scientific literature, creating challenges for timely decision-making in space operations. Effective management in space operations requires tools capable of efficiently processing vast and heterogeneous information sources. This paper systematically evaluates the performance of Retrieval Augmented Generation (RAG) pipelines, combining Large Language Models (LLMs) with information retrieval techniques for extracting and synthesizing actionable knowledge from domain-specific documents. We compare various retrieval strategies, embedding models, and LLM answers to assess their impact on information accuracy, relevance, and reliability. Our results demonstrate that RAG pipelines can significantly enhance knowledge access, reduce uncertainty, and support decision-making in complex space operations.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

RAGe: 検索拡張生成評価フレームワーク

大規模言語モデル (LLM) アプリケーション、特に検索拡張生成 (RAG) に依存するアプリケーションの導入は、高い計算需要、古い知識ベース、および最適なパイプライン コンポーネントを手動で選択する必要があるため、依然として困難です。この研究では、リソース テレメトリとコンポーネントの推奨に焦点を当て、ドメイン固有のデータセットに最適なコンポーネントを提案することで、RAG アプリケーションの効率的な開発をベンチマークおよびガイドするためのモジュール式フレームワークを提案します。私たちのアプローチでは、ドキュメント チャンキング、ベクトル データベース、埋め込みモデル、レトリーバーなどの LLM アプリケーションのコア技術を活用して、精度、効率、スケーラビリティの間のトレードオフを評価します。 RAGe は、取得と生成の品質を基盤となるハードウェアの制約と直接相関させることにより、研究者が特定の運用ニーズに合わせて最も効果的なドメイン固有の RAG セットアップを特定できるようにサポートし、消費者グレードのハードウェアでもラピッド プロトタイピングを容易にします。

原文 (English)

RAGe: A Retrieval-Augmented Generation Evaluation Framework

Deploying Large Language Model (LLM) applications, particularly those relying on Retrieval-Augmented Generation (RAG), remains challenging due to high computational demands, outdated knowledge bases, and the need to manually select optimal pipeline components. In this work, we propose a modular framework for benchmarking and guiding the efficient development of RAG applications by focusing on resource telemetry and component recommendation, suggesting the best components for a domain-specific dataset. Our approach leverages core techniques in LLM applications, including document chunking, vector databases, embedding models, and retrievers, to evaluate trade-offs among accuracy, efficiency, and scalability. By directly correlating retrieval and generation quality with underlying hardware constraints, RAGe supports researchers to identify the most effective, domain-specific RAG setups for their specific operational needs, facilitating rapid prototyping even on consumer-grade hardware.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

ベンチマークだけでは不十分: 運用システムにおけるエージェント モデルのランタイム評価のための RAMP

LLM エージェントは、コーディング アシスタントから自律型ソフトウェア エンジニアリング システムへと急速に進化しています。しかし、既存の評価手法は依然として、静的で孤立した短期間のベンチマークを中心としたものであり、現実世界の運用ワークフローの動的な複雑さを捉えることができません。その結果、ベンチマークのパフォーマンスは、長い実行チェーン、ツールの相互作用、依存関係の管理、反復的なフィードバック ループを含む現実的なランタイム環境下では実際の能力をあまり反映していない可能性があります。したがって、長期的なソフトウェア エンジニアリング エージェントを評価するための、実稼働ベースのインフラストラクチャである RAMP を紹介します。 YatCC 統合プラットフォーム上に構築された RAMP は、標準化されたオーケストレーションおよび実行インターフェイスを通じて、統合されたランタイム評価アーキテクチャを提供します。 RAMP は、シリアル依存関係と複雑なツールチェーンの相互作用を伴う現実的なコンパイラー構築ワークロードを導入し、ワークフローの部分的な障害時の実行動作を分析するための段階的回復メカニズムを導入します。このフレームワークにはさらに、結果の品質とプロセスの効率を共同で評価するユーティリティ指向の多次元指標が組み込まれています。私たちは 15 の主流モデルにわたって実行時評価を実施し、従来の個別ベンチマークではほとんど目に見えない大幅な機能低下を観察しました。タスクの完了率はシリアル ワークフロー全体で徐々に低下し、初期段階の 100% から最終段階ではわずか 20% まで低下しますが、評価されたモデルのどれもパイプライン全体を正常に完了できませんでした。実行時分析により、系統的な障害の伝播と大幅なリソース効率の低下が明らかになり、比較可能なモデル間で計算コストが最大 3 桁も異なります。これらの発見は、RAMP がエージェント モデルの評価を、継続的で実行時観察可能な実稼働ベースの評価に向けて前進させることを示唆しています。

原文 (English)

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain largely centered on static, isolated, and short-horizon benchmarks that fail to capture the dynamic complexity of real-world production workflows. As a result, benchmark performance may poorly reflect practical capability under realistic runtime environments involving long execution chains, tool interactions, dependency management, and iterative feedback loops. We thus present RAMP, a production-grounded infrastructure for assessing long-horizon software engineering agents. Built upon the YatCC integrated platform, RAMP provides a unified runtime assessment architecture through standardized orchestration and execution interfaces. RAMP introduces realistic compiler-construction workloads with serial dependencies and complex toolchain interactions, together with a staged recovery mechanism for analyzing execution behavior under partial workflow failure. The framework further incorporates utility-oriented multi-dimensional metrics that jointly evaluate outcome quality and process efficiency. We conduct runtime assessments across 15 mainstream models and observe substantial capability degradation that remains largely invisible to conventional isolated benchmarks. Task completion rates progressively collapse across serial workflows, dropping from 100% in the initial stage to only 20% in the final stage, while none of the evaluated models successfully completes the entire pipeline. Runtime analysis reveals systematic failure propagation and significant resource inefficiencies, with computational costs differing by up to three orders of magnitude among comparable models. These findings suggest RAMP advances agentic model evaluation toward continuous, runtime-observable, and production-grounded assessment.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ChildEval: 大規模な言語モデルが子供の個性を満たすとき

LLM はパーソナライズされたチャットボットを可能にしますが、子供特有の好みの体系的な評価がまだ不足しているため、子供中心のパーソナライゼーションにおける LLM の有効性は依然として不明瞭です。このギャップに対処するために、私たちは、長いコンテキストの会話において子供中心の好みを推論して従う LLM の能力を評価するためのベンチマークである ChildEval を導入します。 ChildEval には、3 ~ 6 歳の子供の 29,000 個の合成されたペルソナ プロファイルが含まれており、比較的静的な背景情報を提供します。各ペルソナは子供の好みに関連付けられており、その好みはペルソナと一致することもあれば、対立することも、独立していることもあり、単一の文で明示的に、または 6 ~ 10 ターンの対話を通じて暗黙的に表現されます。明示的および暗黙的な好みは、同じ基礎的な好みを反映するように設計されていますが、表現が異なり、静的なペルソナの変化ではなく好み表現の動的な側面を捉えます。このベンチマークは、子どもの日常生活と発達をカバーする 5 つのトップレベルのカテゴリーと 14 のサブレベルのカテゴリーにまたがっています。さらに、オープンソース LLM を体系的に評価するための、きめ細かい子供中心の評価プロトコルを提案します。実験結果は、さまざまなパーソナライズされた表現が LLM 応答にどのような影響を与えるかを示しており、ChildEval を微調整することで子供中心のパフォーマンスを向上できることが示唆されています。コードとデータセットは https://github.com/ziyanluo/ChildEval で入手できます。

原文 (English)

ChildEval: When large language models meet children's personalities

While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-specific preferences is still lacking. To address this gap, we introduce ChildEval, a benchmark for evaluating LLMs' ability to infer and follow child-centered preferences in long-context conversations. ChildEval contains 29K synthesized persona profiles of children aged 3-6, providing relatively static background information. Each persona is associated with a child preference-which may align with, conflict with, or be independent of the persona-expressed either explicitly in a single sentence or implicitly through 6-10 turn dialogues. Explicit and implicit preferences are designed to reflect the same underlying preference but differ in expression, capturing dynamic aspects of preference expression rather than changes in the static persona. The benchmark spans five top-level and fourteen sub-level categories covering children's daily lives and development. We further propose fine-grained, child-centric evaluation protocols to systematically assess open-source LLMs. Experimental results demonstrate how different personalized representations affect LLM responses and suggest that finetuning on ChildEval can enhance child-centered performance. Our code and dataset are available at https://github.com/ziyanluo/ChildEval.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

VibeSearchBench: 実環境における長期的なプロアクティブ検索のベンチマーク

LLM ベースのエージェントは検索ベンチマークで高いスコアを獲得していますが、実際のユーザーは一貫して結果が満足できないと感じており、評価とエクスペリエンスのギャップが根強く残っていることが明らかになりました。このギャップは、既存のベンチマークが過剰に指定されたクエリ、シングルターン インタラクション、および固定スキーマ評価に依存しているためであると考えられますが、これらのいずれも、ユーザーとエージェントが協力してマルチターン対話を通じてあいまいな意図を洗練するという実際の検索動作を反映していません。私たちはこのパラダイムを VibeSearch と名付け、20 のドメインにわたって手動で精選された 200 のバイリンガル (中国語と英語) タスクで構成されるベンチマークである VibeSearchBench を導入します。このベンチマークは、VibeSearch-Pro (プロフェッショナル) サブセットと VibeSearch-Daily (日常生活) サブセットに分かれています。各タスクは、ユーザー ペルソナとスキーマフリーのグラウンド トゥルース ナレッジ グラフを組み合わせ、漸進的開示ユーザー シミュレーターとグラフ マッチング評価フレームワークを通じて評価されます。 ReAct フレームワークと OpenClaw エージェント ハーネスの両方で 7 つのフロンティア モデルのベンチマークを行います。結果は、すべてのモデルが依然として VibeSearch には実質的に不十分であることを示し (最高 F1: 30.30)、ロングコンテキスト推論、プロアクティブな意図の引き出し、および構造化された知識の構築における根本的な進歩の必要性を強調しています。

原文 (English)

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience gap. We attribute this gap to existing benchmarks' reliance on over-specified queries, single-turn interactions, and fixed-schema evaluation, none of which reflect real search behavior where users and agents collaboratively refine vague intent through multi-turn dialogue. We term this paradigm VibeSearch and introduce VibeSearchBench, a benchmark comprising 200 manually curated bilingual (Chinese and English) tasks across 20 domains, split into VibeSearch-Pro (professional) and VibeSearch-Daily (daily-life) subsets. Each task pairs a user persona with a schema-free ground-truth knowledge graph, and is evaluated through a progressive-disclosure user simulator and a graph-matching evaluation framework. We benchmark seven frontier models under both the ReAct framework and the OpenClaw agent harness. Results show that all models remain substantially inadequate for VibeSearch (best F1: 30.30), highlighting the need for fundamental advances in long-context reasoning, proactive intent elicitation, and structured knowledge construction.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

結果から語ろう: LLM 動作ベンチマークのためのレプリケーション ファースト パラダイム

LLM の行動 (共感、抑制、調整された感情の調子) を主観的に評価することは困難です。このような資質に関する人間の評価者間の合意は rho ~ 0.45 付近で飽和し、裁判官としての LLM 代理だけでは循環性の危険性があります。対象の訓練コホートを共有する裁判官は独立して検証することができません。人間と評価者の単一の合意に対する有効性の定着は、人間自身が同意しない能力には適用されません。私たちは、レプリケーションファーストのパラダイムを提案します。1 つの評価者グループに固定するのではなく、4 つの直交する特性、つまり K 回の実行にわたる信頼性、アーキテクチャ的に異なるジャッジ間の機器間のレプリケーション、初期のトレーニング コホートからのジャッジによる履歴フットプリントのキャリブレーション、および事前登録された予測によって機器を認証します。ルーブリックをイテレーション全体でデータに基づいて自己進化させることで、感情的な伴奏をテストします。次元は事前に規定されておらず、手順は 9 次元のセットに安定します。事前登録は、テスト データが収集される前にコミットされた 10 個の反証可能な仮説と 11 個の将来予測に適用されます。このパラダイムを 8 つのファミリーにわたる 49 のモデルに適用すると、集計スコアに隠されているものが明らかになります。アドバイスの抑制、つまりモデルが共感的な文脈で一方的な解決策の提供を控えるかどうかについては、gpt-5 は gpt-4.1 から 1.87 ポイント低下し、Opus-4.7 は Opus-4.6 から 0.629 ポイント低下しましたが、合計スコアは横ばいでした。回帰は 3 回のユーザーと代理人の交換 (規模の 95%) を生き残り、5 家族の裁判官スタックと 17 か月のコホートギャップにわたって再現され、74 回の実際の ESConv 会話で持続しました (rho [0.749, 0.850])。機器は通常のクリッペンドルフ アルファ = 0.91 に達します。副産物として、このパラダイムは飽和源診断として機能し、手段の天井(ルーブリックの改良によって破壊可能)を構造の天井(シナリオまたは名簿の介入が必要)から分離します。

原文 (English)

Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking

Subjective evaluation of LLM behavior -- empathy, restraint, calibrated emotional tone -- is hard. Human inter-rater agreement on such qualities saturates near rho ~ 0.45, and an LLM-as-judge proxy alone risks circularity: a judge sharing the target's training cohort cannot independently verify it. Anchoring validity to a single human-rater consensus does not extend to capabilities where humans themselves disagree. We propose a replication-first paradigm: instead of anchoring on one rater group, we certify the instrument via four orthogonal properties -- reliability across K runs, cross-instrument replication across architecturally distinct judges, historical-footprint calibration via judges from earlier training cohorts, and pre-registered prediction. We test it on emotional accompaniment by letting the rubric self-evolve data-driven across iterations: the dimensions are not pre-stipulated and the procedure stabilizes to a 9-dimension set. Pre-registration applies to 10 falsifiable hypotheses and 11 forward predictions, committed before any test data was collected. Applied to 49 models across 8 families, the paradigm surfaces what aggregate scores hide. On advice-restraint -- whether a model refrains from giving unsolicited solutions in empathic contexts -- gpt-5 falls 1.87 points from gpt-4.1 and Opus-4.7 falls 0.629 from Opus-4.6, while aggregate scores stay flat. The regression survives three user-proxy swaps (95% of magnitude), replicates across a 5-family judge stack and a 17-month cohort gap, and persists on 74 held-out real ESConv conversations (rho in [0.749, 0.850]); the instrument reaches ordinal Krippendorff alpha = 0.91. As a by-product, the paradigm acts as a saturation-source diagnostic, separating instrumental ceilings (breakable by rubric refinement) from structural ceilings (needing scenario or roster intervention).

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

組換えベースのデカルト遺伝的プログラミングの評価の向上

デカルト遺伝的プログラミングは伝統的に、進化的探索を推進するための主要な、多くの場合唯一の遺伝的演算子として突然変異を使用してきました。近年の進歩にもかかわらず、組換えベースのアプローチは、明らかにパフォーマンスの向上が見られないため、長い間避けられてきました。この研究では、最近提案された 2 つの組換えベースの演算子、サブグラフ クロスオーバーと離散表現型組換えを、記号回帰のベンチマーク プラットフォームである SRBench で検証します。 TinyverseGP フレームワークで提供される実装を使用して、これら 2 つの演算子でそれぞれの表現のハイパーパラメーターの最適化を実行します。私たちの研究は、ハイパーパラメーターの最適化が、組み換えベースのデカルト遺伝的プログラミングのパフォーマンスの向上につながる可能性があることを示しています。

原文 (English)

Improving Evaluation of Recombination-based Cartesian Genetic Programming

Cartesian Genetic Programming has traditionally been using mutation as its main and often sole genetic operator to drive evolutionary search. Despite advancements in recent years, recombinationbased approaches have long been avoided, due to apparent lack of performance gains. This study examines two recently suggested recombination-based operators, subgraph crossover and discrete phenotypic recombination on SRBench, a benchmarking platform for symbolic regression. Using the implementations provided in the TinyverseGP framework, we perform hyperparameter optimisation of the respective representations with these two operators. Our work demonstrates that hyperparameter optimisation can lead to improvements in performance for recombination-based Cartesian Genetic Programming.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

評価がどのように設計されているかを知っているモデルは、より安全なスコアを獲得します

AI の安全性評価の有効性は、制御設定および展開設定全体でモデルが一貫して動作するかどうかに依存します。これまでの研究では、言語化された評価の認識とその後の行動の変化の源として、仮説的なシナリオなどのテスト時の文脈上の手がかりが特定されてきました。この論文では、この現象の潜在的な説明である評価メタ知識を調査します。これは、評価を特徴付ける構造的特性に関するパラメトリック知識として定義されます。ベンチマークの公開が記憶を通じてパフォーマンスの向上につながるデータセットの汚染と同様に、評価の実践を説明するテキストでトレーニングされたモデルは、たとえば AI ベンチマークに関する科学記事やソーシャル メディアの投稿への公開を通じて、評価のようなコンテキストを認識して応答することを暗黙的に学習する可能性があると仮説を立てています。これをテストするために、検証可能な構造や道徳的ジレンマなどの評価特性を記述する合成文書のモデルを微調整します。この微調整されたモデルを 6 つの安全性ベンチマークで評価すると、基本モデルや制御モデルよりも大幅に安全であることがわかりました。この行動の変化は、分析を評価意識の明示的な言語化が欠けている回答に限定した場合でも持続します。私たちの結果は、評価のメタ知識が安全ベンチマークのパフォーマンスを水増しし、明示的な記憶や言語化された評価の認識とは独立した新たな交絡因子を導入する可能性があることを示しており、したがって検出が困難です。これらの発見は、AI の安全性評価の設計と解釈に重要な意味を持ちます。コードとモデルは https://github.com/compass-group-tue/arxiv2026_evaluation_meta_knowledge で入手できます。

原文 (English)

Models That Know How Evaluations Are Designed Score Safer

The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized evaluation awareness and subsequent behavioral shift. In this paper, we investigate a potential explanation of this phenomenon: evaluation meta-knowledge, defined as parametric knowledge about the structural traits that characterize evaluations. Similar to dataset contamination, where benchmark exposure leads to higher performance through memorization, we hypothesize that models trained on texts describing evaluation practices may implicitly learn to recognize and respond to evaluation-like contexts, for instance, through exposure to scientific articles or social media posts about AI benchmarking. To test this, we fine-tune models on synthetic documents describing evaluation traits such as verifiable structures or moral dilemmas. Evaluating this fine-tuned model on six safety benchmarks, we find that it is significantly safer than the base model and control model. This behavioral shift persists even when restricting the analysis to responses lacking explicit verbalization of evaluation awareness. Our results demonstrate that evaluation meta-knowledge may inflate safety benchmark performance, introducing a novel confounder that is independent of explicit memorization or verbalized evaluation awareness, thus, challenging to detect. These findings have important implications for the design and interpretation of AI safety evaluations. Our code and models are available at https://github.com/compass-group-tue/arxiv2026_evaluation_meta_knowledge.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

PULSE 法を使用した AI による分配関数推定による、化学的に無秩序な化合物の熱力学特性

この記事では、化学的に無秩序な化合物の熱力学特性を推定するための PULSE 法 (区分関数教師なし学習サンプリングと評価) の改良版を紹介します。目的は、このタイプの材料に対するモンテカルロ手法の計算コストを削減し、この生成ツールがシステムの分配関数をサンプリングして推定することによって熱力学特性を推定できることを実証することです。この革新的なアプローチを検証するために、ベンチマークとして 2D イジング モデルを使用します。従来のモンテカルロサンプリング法と比較して、私たちの方法が高精度かつ効率的に平均特性を正確に再現することを実証します。私たちの結果は、PULSE 法の効率性と適応性を強調しており、従来の方法では非効率すぎて化学的乱れの影響を受ける特性を低コストで計算できない材料を研究するための貴重なツールとなっています。

原文 (English)

Thermodynamic properties of chemically disordered compounds via AI-driven estimation of partition function with the PULSE method

In this article, we present an improved version of the PULSE method (Partition function Unsupervised Learning Sampling and Evaluation) for estimating the thermodynamic properties of chemically disordered compounds. The aim is to reduce the computational cost of Monte Carlo approaches for this type of material and to demonstrate that this generative tool can estimate thermodynamic properties by sampling and estimating the partition function of the system. To validate this innovative approach, we use the 2D Ising model as a benchmark. We demonstrate that our method accurately reproduces average properties with high precision and efficiency compared to traditional Monte Carlo sampling methods. Our results highlight the efficiency and adaptability of the PULSE method, making it a valuable tool for studying materials for which conventional methods are too inefficient to compute properties affected by chemical disorder at low cost.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

立場: 「ポジティブ バックドア」というレッテルを撤回 -- 秘密の調整には厳格かつ体系的な評価が必要

この意見書では、AI/ML コミュニティは過剰な主張をやめ、「ポジティブ バックドア」というラベルを廃止し、代わりにトリガーによって起動される隠れた動作を秘密の調整として扱うべきであると主張しています。重要なことは、Secret Alignment に基づく保護主張は、厳密で標準化された評価によって裏付けられない限り、デフォルトでは安全ではないと推定されるべきであるということです。オープンウェイト LLM とアクセス可能なトレーニング/推論スタックによって可能になったプライベート AI 時代は、言語モデルを私有のデジタル資産に変え、不正アクセス、モデルの盗難、および動作上の悪用に関するセキュリティ上の懸念を引き起こします。最近、これらの課題に対処するために、「ポジティブ バックドア」として枠組み化された一連の作業が提案されています。私たちの立場を証拠に基づいて根拠づけるために、私たちはこれらの提案を、アクセス ゲーティング、所有権の帰属、安全性の執行のための秘密トリガー動作の関連付けとして統合し、有効性、無害性、永続性、効率性、堅牢性、信頼性という 6 つの中核特性にわたって 3 つの代表的なアプリケーションを評価します。私たちの結果は、既存の主張では過小評価されがちなトリガー動作マッピングの、特に機密性、完全性、可用性 (CIA) におけるかなりの脆弱性を明らかにしました。さらに、これらの結果を行動の密度と意思決定の複雑さに関連付け、導入時のリスクを理解し、秘密の調整の主張を証明可能にするコミュニティ全体の評価を動機付けるための行動のレンズを提供します。

原文 (English)

Position: Retire the "Positive Backdoor" Label -- Secret Alignment Requires Strict and Systematic Evaluation

This position paper argues that the AI/ML community should stop overclaiming and retire the label "positive backdoor," and instead treat trigger-activated hidden behaviors as Secret Alignment. Crucially, protective claims based on Secret Alignment should be presumed not secure by default unless supported by rigorous, standardized evaluation. The Private AI era, enabled by open-weight LLMs and accessible training/inference stacks, turns language models into privately owned digital assets, creating security concerns around unauthorized access, model theft, and behavioral misuse. Recently, a line of work framed as "positive backdoors" has been proposed to address these challenges. To ground our position in evidence, we unify these proposals as covert trigger-behavior associations for access gating, ownership attribution, and safety enforcement, and evaluate three representative applications across six core properties: effectiveness, harmlessness, persistence, efficiency, robustness, and reliability. Our results reveal substantial brittleness - especially in the confidentiality, integrity, and availability (CIA) - of trigger-behavior mappings often underrepresented by existing claims. We further relate these outcomes to behavior density and decision complexity, offering a behavioral lens for understanding deployment-time risks and motivating community-wide evaluation that makes Secret Alignment claims provable.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

言語モデルにおける形式と機能の測定

言語モデルを評価するために、子供の言語習得に関する定量的指標を導入します。私たちは、幼児が早期かつ正確に習得する、英語の限定詞の形式的な構文的および機能的な談話特性に焦点を当てています。我々は、言語の構文知識と談話知識の両方について的を絞ったテストを提供する新しいプロンプト方法である文脈代替選択(CAC)を提案します。この方法により、言語モデルを子供と直接比較でき、さらに重要なことに、実証研究で独自に確立された統計ベンチマークと直接比較できます。人間の子供のように、同等の量のデータでトレーニングされた現在のモデルは形式的ベンチマークと機能的ベンチマークの両方を同時に満たすものはありませんが、非常に大規模なモデルの一部は満たしています。私たちは、言語モデルの認知状態に特に重点を置き、方法論的および技術的貢献として結果を提示します。

原文 (English)

Measuring Form and Function in Language Models

We introduce quantitative metrics for child language acquisition to evaluate language models. Our focus is on the formal syntactic and functional discourse properties of determiners in English, which young children acquire early and accurately. We propose Contextual Alternative Choice (CAC), a new prompting method which provides targeted tests for both syntactic and discourse knowledge of language. The method enables direct comparison of language models against children, and more importantly, against statistical benchmarks independently established in empirical research. No current model trained on a comparable amount of data simultaneously meet both formal and functional benchmarks like human children, but some very large models do. We present our results as methodological and technical contributions, with specific emphasis on cognitive status of language models.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

裁判官としての信頼できる多言語 LLM を目指して: 実証的研究

大規模言語モデル (LLM) は、生成されたテキストの自動評価にますます使用されていますが、これまでの研究のほとんどは英語に焦点を当てています。多言語評価の需要が高まっているにもかかわらず、特に低リソース言語やドメイン内データが不足しているシナリオでは、LLM ベースの評価ツールを多言語設定に拡張することは依然として困難です。この研究では、ドメイン内データが微調整に利用できるかどうかを考慮しながら、審査員としての多言語 LLM を開発するためのいくつかの戦略を検討します。私たちは、命令の翻訳、単言語と多言語の監視、モデルのサイズを考慮して、高リソース、中リソース、低リソースの言語を代表する英語、スペイン語、バスク語を体系的に分析します。評価のために、2 つの既存のメタ評価データセットをバスク語とスペイン語に拡張しました。私たちの結果は重要なトレードオフを明らかにしています。ドメイン内データが利用可能な場合、微調整された小型モデルは独自のモデルに匹敵するパフォーマンスを達成できますが、ドメイン外設定ではより大きなモデルを使用したゼロショット評価がより効果的であることが証明されています。また、ドメイン外データの微調整がモデルのパフォーマンスに悪影響を与える可能性があることも観察しています。これらの調査結果は、効率的で信頼性の高い多言語評価パイプラインを構築するための実践的なガイダンスを提供します。データとコードは、hitz-zentroa/mJudge で公開されています。

原文 (English)

Towards Reliable Multilingual LLMs-as-a-Judge: An Empirical Study

Large language models (LLMs) are increasingly used for the automatic evaluation of generated text, yet most prior work focuses on English. Despite the growing demand for multilingual evaluation, extending LLM-based evaluators to multilingual settings remains challenging, particularly for low-resource languages and scenarios where in-domain data is scarce. This work explores several strategies for developing multilingual LLMs-as-a-judge, considering whether in-domain data is available for fine-tuning or not. We systematically analyze English, Spanish, and Basque, representing high-, mid-, and low-resource languages, considering instruction translation, monolingual versus multilingual supervision, and model size. For evaluation, we extend two existing meta-evaluation datasets to Basque and Spanish. Our results reveal key trade-offs: When in-domain data is available, fine-tuned smaller models can achieve performance comparable to proprietary models, whereas zero-shot evaluation with larger models proves more effective in out-of-domain settings. We also observe that fine-tuning on out-of-domain data can adversely affect model performance. These findings provide practical guidance for building efficient, reliable multilingual evaluation pipelines. The data and code are publicly available at hitz-zentroa/mJudge.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

IPO-Mine: 長くマルチモーダルな IPO ドキュメントのセクション構造分析のためのツールキットとデータセット

新規株式公開 (IPO) 申請書は、民間企業が株式を公開し、個人 (個人) 投資家がその株式を購入できるようにするときに発行される文書です。これらの提出書類は企業の事業、財務、リスクについて説明しており、説明的なテキストと画像を含む長くて多様な文書です。金融市場にとっての重要性にもかかわらず、最新の言語およびマルチモーダル モデルを使用して IPO 申請を研究するための大規模な標準化されたデータセットやベンチマークはありません。これらの文書は重大な課題を引き起こします。申請は頻繁に 500,000 トークンを超え、一貫した構造的組織が欠如しています。 IPO 申請書類をダウンロードし、標準化されたセクション構造のテキストと抽出された画像に解析するためのオープンソース フレームワークである IPO-Toolkit を紹介します。このツールキットは、ファイリングをセグメント化し、埋め込まれた画像を抽出し、長いマルチモーダルなドキュメントに対する大規模で再現可能な分析ワークフローを可能にする構造化された出力を生成します。このインフラストラクチャを使用して、1994 年から 2026 年までの 109,000 件を超える IPO 申請および修正をカバーし、76,000 枚を超える画像を含む大規模なセクション構造のマルチモーダル データセットである IPO データセットを構築します。抽出された財務チャートに対して、チャートの品質や誤解を招く可能性の評価など、構造化された評価タスクを確立します。私たちの実験によると、最先端のマルチモーダル モデルは、これらのタスクに関する専門家の人間の判断から逸脱することが多く、現実世界の長い規制文書に基づくマルチモーダル推論における調整の課題が明らかになりました。 IPO データセットは、ベンチマークを超えて、セクションレベルのテキストのバリエーションや、ビジュアルおよびテキストの開示慣行における業界間の差異の大規模な分析を可能にします。私たちのコード、データセット、Web サイトは CC-BY-4.0 に基づいて公開されています。

原文 (English)

IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents

An Initial Public Offering (IPO) filing is a document released when a private firm goes public, allowing individual (retail) investors to purchase its shares. These filings describe a firm's business, financials, and risks and are long, multimodal documents with narrative text and images. Despite their importance to financial markets, there is no large-scale, standardized dataset or benchmark for studying IPO filings with modern language and multimodal models. These documents pose significant challenges: filings frequently exceed 500,000 tokens and lack consistent structural organization. We introduce the IPO-Toolkit, an open-source framework for downloading and parsing IPO filings into standardized section-structured text and extracted images. The toolkit segments filings, extracts embedded images, and produces structured outputs that enable large-scale, reproducible analysis workflows over long, multimodal documents. Using this infrastructure, we construct the IPO-Dataset, a large, section-structured, multimodal dataset covering more than 109,000 IPO filings and amendments from 1994 to 2026 and containing over 76,000 images. We establish structured evaluation tasks over extracted financial charts, including chart quality and misleadingness assessment. Our experiments show that state-of-the-art multimodal models often diverge from expert human judgments on these tasks, exposing alignment challenges in multimodal reasoning over long, real-world regulatory documents. Beyond benchmarking, the IPO-Dataset enables large-scale analysis of section-level textual variation and cross-industry differences in visual and textual disclosure practices. Our code, dataset, and website are publicly available under CC-BY-4.0.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

自動データ分析に向けて: LLM ベースのリスク推定のためのガイド付きフレームワーク

大規模言語モデル (LLM) は重要な意思決定パイプラインにますます統合されており、この傾向により堅牢で自動化されたデータ分析の需要が高まっています。データセットのリスク分析に対する現在のアプローチは、時間のかかる複雑なタスクを伴う手動の監査方法に限定されていますが、人工知能 (AI) に基づく完全に自動化された分析では、AI の調整に起因する幻覚や問題が発生します。この目的を達成するために、この研究では、人間の指導と監督の下で生成 AI を統合するデータセット リスク推定のフレームワークを提案し、将来の自動リスク分析パラダイムの基礎を確立することを目的としています。私たちのアプローチでは、LLM を利用してデータベース スキーマの意味論的および構造的プロパティを特定し、その後クラスタリング手法を提案し、それらのコードを生成し、最終的に生成された結果を解釈します。人間のスーパーバイザーは、目的の分析に関してモデルをガイドし、プロセスの整合性とタスクの目的との整合性を確保します。概念実証は、リスク評価タスクで有意義な結果を生み出すフレームワークの有用性の実現可能性を実証するために提示されます。

原文 (English)

Towards automated data analysis: A guided framework for LLM-based risk estimation

Large Language Models (LLMs) are increasingly integrated into critical decision-making pipelines, a trend that raises the demand for robust and automated data analysis. Current approaches to dataset risk analysis are limited to manual auditing methods which involve time-consuming and complex tasks, whereas fully automated analysis based on Artificial Intelligence (AI) suffers from hallucinations and issues stemming from AI alignment. To this end, this work proposes a framework for dataset risk estimation that integrates Generative AI under human guidance and supervision, aiming to set the foundations for a future automated risk analysis paradigm. Our approach utilizes LLMs to identify semantic and structural properties in database schemata, subsequently propose clustering techniques, generate the code for them and finally interpret the produced results. The human supervisor guides the model on the desired analysis and ensures process integrity and alignment with the task's objectives. A proof of concept is presented to demonstrate the feasibility of the framework's utility in producing meaningful results in risk assessment tasks.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

EngiAI: LLM 主導のエンジニアリング設計のためのマルチエージェント フレームワークおよびベンチマーク スイート

Large Language Model (LLM) エージェントはエンジニアリング設計タスクに適用されることが増えていますが、既存の評価フレームワークは、シミュレーション、検索、製造準備を組み合わせたマルチエージェント システムに適切に対応していません。 3 つの評価次元を備えたベンチマーク スイートを紹介します。(1) ツールの直接使用、意味論的曖昧さの解消、条件分岐、作業記憶タスクなど、明確な認知的要求を対象とした 7 つのプロンプト スタイルを備えたワークフロー ベンチマーク。 (2) パラメータ選択に対する検索の寄与を分離するゲート付きスコアリングを備えた検索拡張生成 (RAG) ベンチマーク。 (3) SLURM クラスター上のエンドツーエンドの ML トレーニング オーケストレーションを評価するハイ パフォーマンス コンピューティング (HPC) ベンチマーク。ベンチマークとともに、LangGraph 上に構築されたマルチ エージェント システム (MAS) リファレンス実装である EngiAI も紹介します。EngiAI は、スーパーバイザ アーキテクチャを通じて 7 つの専門エージェントを調整し、トポロジの最適化、ドキュメントの取得、HPC ジョブ オーケストレーション、および 3D プリンタの制御を統合することでベンチマークを運用します。 4 つの LLM バックエンドと 2 つの EngiBench 問題にわたって、独自のモデルは Beams2D 上で平均タスク完了率 96 ~ 97% を達成していますが、オープンソースの 4B パラメータ モデルは 55 ~ 78% に達しており、世代ごとに明らかな改善が見られます。条件付き分岐が最も困難であることが判明し、Photonics2D の条件付きスタイルではタスクの完了率が 20 ~ 53% に低下します。 RAG ゲーティングにより、ほぼ完璧な検索強化スコア (約 1.0) と検索なしのほぼゼロのスコアが確認され、評価設計が検証されます。 HPC オーケストレーションでは、あるモデルは実行の 100% ですべてのパイプライン ステップを完了しますが、別のモデルは 50% に低下し、長時間実行されるワークフローでは複数ステップの命令のパフォーマンスが低下することがわかります。

原文 (English)

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

Large Language Model (LLM) agents are increasingly applied to engineering design tasks, yet existing evaluation frameworks do not adequately address multi-agent systems that combine simulation, retrieval, and manufacturing preparation. We introduce a benchmark suite with three evaluation dimensions: (1) a workflow benchmark with seven prompt styles targeting distinct cognitive demands-including direct tool use, semantic disambiguation, conditional branching, and working-memory tasks; (2) a Retrieval-Augmented Generation (RAG) benchmark with gated scoring isolating retrieval contributions to parameter selection; and (3) an High Performance Computing (HPC) benchmark evaluating end-to-end ML training orchestration on a SLURM cluster. Alongside the benchmark we present EngiAI, a Multi-Agent System (MAS) reference implementation built on LangGraph that operationalizes the benchmark by coordinating seven specialized agents through a supervisor architecture, unifying topology optimization, document retrieval, HPC job orchestration, and 3D printer control. Across four LLM backends and two EngiBench problems, proprietary models achieve 96-97% average task completion on Beams2D, while open-source 4B-parameter models reach 55-78%, with clear generational improvement. Conditional branching proves most challenging, with task completion dropping to 20-53% for the conditional style on Photonics2D. RAG gating confirms near-perfect retrieval-augmented scores (about 1.0) versus near-zero without retrieval, validating the evaluation design. On HPC orchestration, one model completes all pipeline steps in 100% of runs while another drops to 50%, revealing that multi-step instruction following degrades over long-running workflows.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

MCTS-Judge: コードの正確性評価のための LLM-as-a-Judge でのテスト時間のスケーリング

LLM-as-a-Judge パラダイムは、生成コンテンツの評価には有望ですが、プログラミングなどの推論集中型のシナリオでは信頼性に欠けます。推論モデルの最近の進歩とスケーリング法の変化に触発され、私たちはテスト時の計算を LLM-as-a-Judge に導入する先駆者となり、コードの正確性評価のためのリソース効率の高い System-2 思考フレームワークである MCTS-Judge を提案します。 MCTS-Judge は、モンテカルロ ツリー検索 (MCTS) を利用して、問題をより単純な複数の視点からの評価に分解します。 MCTS-Judge は、現在の軌道における履歴アクションに基づく自己評価と、以前のロールアウトに基づくツリーの上限信頼限界を組み合わせたノード選択戦略を通じて、現在の軌道の全体的な最適化と洗練のバランスをとります。さらに、大規模言語モデル (LLM) による行ごとの分析の実行を促進する、高精度の単体テスト レベルの報酬メカニズムを設計しました。 3 つのベンチマークと 5 つの LLM に関する広範な実験により、MCTS-Judge の有効性が実証され、基本モデルの精度が 41% から 80% に向上し、トークンの数が 3 分の 1 で o1 シリーズ モデルを上回ります。さらなる評価により、ロジック、分析、徹底性、全体的な品質における推論軌跡の優位性が検証され、同時に LLM-as-a-Judge パラダイムのテスト時間のスケーリング則が明らかになります。

原文 (English)

MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation

The LLM-as-a-Judge paradigm shows promise for evaluating generative content but lacks reliability in reasoning-intensive scenarios, such as programming. Inspired by recent advances in reasoning models and shifts in scaling laws, we pioneer bringing test-time computation into LLM-as-a-Judge, proposing MCTS-Judge, a resource-efficient, System-2 thinking framework for code correctness evaluation. MCTS-Judge leverages Monte Carlo Tree Search (MCTS) to decompose problems into simpler, multi-perspective evaluations. Through a node-selection strategy that combines self-assessment based on historical actions in the current trajectory and the Upper Confidence Bound for Trees based on prior rollouts, MCTS-Judge balances global optimization and refinement of the current trajectory. We further designed a high-precision, unit-test-level reward mechanism to encourage the Large Language Model (LLM) to perform line-by-line analysis. Extensive experiments on three benchmarks and five LLMs demonstrate the effectiveness of MCTS-Judge, which improves the base model's accuracy from 41% to 80%, surpassing the o1-series models with 3x fewer tokens. Further evaluations validate the superiority of its reasoning trajectory in logic, analytics, thoroughness, and overall quality, while revealing the test-time scaling law of the LLM-as-a-Judge paradigm.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

モデルのランキングを超えて: 時系列予測の予測可能性に合わせた評価

時系列予測のための AI モデルがますます複雑になる時代では、ベンチマーク リーダーボードのわずかな改善によって進歩が測定されることがよくあります。ただし、このアプローチには根本的な欠陥があります。標準的な評価指標では、モデルのパフォーマンスとデータの本質的な予測不可能性が混同されます。この差し迫った課題に対処するために、スペクトル コヒーレンスに基づいた、予測可能性を重視した新しい診断フレームワークを導入します。私たちのフレームワークは 2 つの主要な貢献をします。1 つは、特定の予測インスタンスの固有の難しさを定量化する計算効率 ($O(N\log N)$) とタスクに合わせたスコアであるスペクトル コヒーレンス予測可能性 (SCP)、もう 1 つはモデルがデータ内の線形予測可能な情報をどの程度効果的に利用しているかを正確に測定する周波数分解診断ツールである線形利用率 (LUR) です。私たちはフレームワークの有効性を検証し、それを活用して 2 つの核となる洞察を明らかにします。まず、「予測可能性のドリフト」に関する最初の体系的な証拠を提供し、タスクの予測難易度が時間の経過とともに急激に変化することを示します。第 2 に、私たちの評価では、重要なアーキテクチャ上のトレードオフが明らかになりました。つまり、予測可能性の低いデータに対しては複雑なモデルが優れているのに対し、より予測可能なタスクに対しては線形モデルが非常に効果的です。私たちはパラダイム シフトを提唱し、単純な集計スコアを超えて、より公平なモデル比較とモデル動作のより深い理解を促進する、より洞察力に富んだ予測可能性を意識した評価に移行します。

原文 (English)

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, this approach suffers from a fundamental flaw: standard evaluation metrics conflate a model's performance with the data's intrinsic unpredictability. To address this pressing challenge, we introduce a novel, predictability-aligned diagnostic framework grounded in spectral coherence. Our framework makes two primary contributions: the Spectral Coherence Predictability (SCP), a computationally efficient ($O(N\log N)$) and task-aligned score that quantifies the inherent difficulty of a given forecasting instance, and the Linear Utilization Ratio (LUR), a frequency-resolved diagnostic tool that precisely measures how effectively a model exploits the linearly predictable information within the data. We validate our framework's effectiveness and leverage it to reveal two core insights. First, we provide the first systematic evidence of "predictability drift", demonstrating that a task's forecasting difficulty varies sharply over time. Second, our evaluation reveals a key architectural trade-off: complex models are superior for low-predictability data, whereas linear models are highly effective on more predictable tasks. We advocate for a paradigm shift, moving beyond simplistic aggregate scores toward a more insightful, predictability-aware evaluation that fosters fairer model comparisons and a deeper understanding of model behavior.

2026-05-28 13:00 JSTarXiv cs.AIビジネス/資金調達

言語モデルにおける AI 倫理ツールの評価: 開発者の視点からのケーススタディ

人工知能 (AI) では、テキスト生成を通じて人間との現実的な会話をシミュレートできるシステムが広く採用されているため、言語モデルが非常に重要になってきています。社会に影響を与えるため、これらの言語モデルの開発と展開は、その悪影響と起こり得る害に注意しながら、責任を持って行う必要があります。このシナリオでは、AI Ethics Tools (AIET) の出版物の数が最近増加しています。これらの AIET は、AI の設計、開発、使用段階のガイドとして受け入れられた価値をもたらすことで、開発者、企業、政府、その他の利害関係者がテクノロジーに対する信頼、透明性、責任を確立できるように設計されています。しかし、AIET の多くは、適切な文書、使用例、実際の有効性を証明するものを欠いています。この論文では、言語モデルで AIET を評価するための方法論を紹介します。私たちのアプローチには、213 の AIET に関する広範な文献調査が含まれ、包含基準と除外基準を適用した後、モデル カード、ALTAI、ファクトシート、および危害モデリングの 4 つの AIET を選択しました。評価のために、ポルトガル語用に開発された言語モデルに AIET を適用し、その開発者に 35 時間のインタビューを実施しました。評価では、モデルに関する倫理的考慮事項を特定する際に、AIET の使用と品質に関する開発者の視点が考慮されました。この結果は、適用された AIET が言語モデルに関する一般的な倫理的考慮事項を策定するためのガイドとして機能することを示唆しています。ただし、慣用的な表現など、これらのモデルの固有の側面には対応していないことに注意してください。さらに、これらの AIET は、ポルトガル語モデルの潜在的な悪影響を特定するのには役立ちませんでした。

原文 (English)

Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study

In Artificial Intelligence (AI), language models have gained significant importance due to the widespread adoption of systems capable of simulating realistic conversations with humans through text generation. Because of their impact on society, developing and deploying these language models must be done responsibly, with attention to their negative impacts and possible harms. In this scenario, the number of AI Ethics Tools (AIETs) publications has recently increased. These AIETs are designed to help developers, companies, governments, and other stakeholders establish trust, transparency, and responsibility with their technologies by bringing accepted values to guide AI's design, development, and use stages. However, many AIETs lack good documentation, examples of use, and proof of their effectiveness in practice. This paper presents a methodology for evaluating AIETs in language models. Our approach involved an extensive literature survey on 213 AIETs, and after applying inclusion and exclusion criteria, we selected four AIETs: Model Cards, ALTAI, FactSheets, and Harms Modeling. For evaluation, we applied AIETs to language models developed for the Portuguese language, conducting 35 hours of interviews with their developers. The evaluation considered the developers' perspective on the AIETs' use and quality in helping to identify ethical considerations about their model. The results suggest that the applied AIETs serve as a guide for formulating general ethical considerations about language models. However, we note that they do not address unique aspects of these models, such as idiomatic expressions. Additionally, these AIETs did not help to identify potential negative impacts of models for the Portuguese language.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

音声言語モデル評価におけるグローバル トークンの複雑さの誤謬について

大規模な生音声で事前トレーニングされた生成音声言語モデルは、話者や感情などの属性を維持しながら、適切なコンテンツで音声プロンプトを継続でき、音声対話の基礎モデルとして機能します。従来の文献では、これらのモデルは多くの場合、テキストの複雑さの定式化を音声トークンに直接適用する「グローバル トークンの複雑さ」を使用して評価されます。ただし、この方法では音声とテキストのモダリティ間の基本的な違いが見落とされ、音声の特徴が過小評価される可能性があります。この研究では、単純なグローバル トークンの複雑さの代わりに機能する、尤度ベースおよび生成ベースのさまざまな評価方法を提案します。人間が評価した平均意見スコア (MOS) との強い相関関係によって証明されるように、提案された評価が知覚される世代の品質をより忠実に反映していることを実証します。新しい指標に基づいて評価すると、音声言語モデルの相対的なパフォーマンスの状況が再形成され、最高のパフォーマンスを誇るモデルと人間のトップラインとの間のギャップが大幅に減少していることが明らかになります。これらの結果を総合すると、音声言語モデリングの進歩を正確に評価するには、適切な評価が重要であることがわかります。

原文 (English)

On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation

Generative spoken language models pretrained on large-scale raw audio can continue a speech prompt with appropriate content while preserving attributes like speaker and emotion, serving as foundation models for spoken dialogue. In prior literature, these models are often evaluated using ``global token perplexity'', which directly applies the text perplexity formulation to speech tokens. However, this practice overlooks fundamental differences between speech and text modalities, possibly leading to an underestimation of the speech characteristics. In this work, we propose a variety of likelihood- and generative-based evaluation methods that serve in place of naive global token perplexity. We demonstrate that the proposed evaluations more faithfully reflect perceived generation quality, as evidenced by stronger correlations with human-rated mean opinion scores (MOS). When assessed under the new metrics, the relative performance landscape of spoken language models is reshaped, revealing a significantly reduced gap between the best-performing model and the human topline. Together, these results suggest that appropriate evaluation is critical for accurately assessing progress in spoken language modeling.

2026-05-28 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

TABX: マルチエージェント強化学習のための高スループットのサンドボックス バトル シミュレーター

環境の設計は、協調的なマルチエージェント強化学習 (MARL) アルゴリズムの開発と評価を形作る上で重要な役割を果たします。既存のベンチマークは重大な課題を浮き彫りにしていますが、カスタム評価シナリオの設計に必要なモジュール性が欠けていることがよくあります。再構成可能なマルチエージェント タスク用に設計された高スループットのサンドボックスである Totally Accelerated Battle Simulator in JAX (TABX) を紹介します。 TABX は、環境パラメータに対するきめ細かい制御を提供し、さまざまなタスクの複雑さにわたる緊急エージェントの動作とアルゴリズムのトレードオフを系統的に調査できるようにします。 TABX は、GPU 上でハードウェア アクセラレーションによる実行に JAX を活用することで、大規模な並列化を可能にし、計算オーバーヘッドを大幅に削減します。 TABX は、高速かつ拡張可能で簡単にカスタマイズできるフレームワークを提供することで、複雑な構造ドメインにおける MARL エージェントの研究を容易にし、将来の研究のための拡張可能な基盤として機能します。コードは https://github.com/ku-dmlab/TABX から入手できます。

原文 (English)

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning

The design of environments plays a critical role in shaping the development and evaluation of cooperative multi-agent reinforcement learning (MARL) algorithms. While existing benchmarks highlight critical challenges, they often lack the modularity required to design custom evaluation scenarios. We introduce the Totally Accelerated Battle Simulator in JAX (TABX), a high-throughput sandbox designed for reconfigurable multi-agent tasks. TABX provides granular control over environmental parameters, permitting a systematic investigation into emergent agent behaviors and algorithmic trade-offs across a diverse spectrum of task complexities. Leveraging JAX for hardware-accelerated execution on GPUs, TABX enables massive parallelization and significantly reduces computational overhead. By providing a fast, extensible, and easily customized framework, TABX facilitates the study of MARL agents in complex structured domains and serves as a scalable foundation for future research. Our code is available at: https://github.com/ku-dmlab/TABX.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

AlphaForgeBench: 大規模な言語モデルを使用したエンドツーエンドの取引戦略設計のベンチマーク

大規模言語モデル (LLM) の急速な進歩により、金融ベンチマークが急増し、静的な知識評価からインタラクティブな取引シミュレーションへと進化しました。しかし、リアルタイム取引を評価するための既存のフレームワークは、重大な失敗モード、つまり金融の不確実性の下での逐次的な意思決定におけるLLMの動作の深刻な不安定性をほとんど見落としています。広範な実験を通じて、LLM が取引エージェントとして展開された場合、実行ごとに極端な分散を示し、決定論的なデコード下でも一貫性のないアクション シーケンスを生成し、隣接するタイム ステップ間で不合理なアクションの反転が頻繁に発生することを示しました。これらの動作は、ポートフォリオ割り当てタスクにおける連続アクションから離散アクションへのマッピングに対する感度とともに、以前のアクションの永続的な記憶が欠如している LLM のステートレスな自己回帰の性質によるものであると考えられます。これらの欠陥は、多くの既存のオンラインおよびオフライン取引ベンチマークの信頼性と再現性を根本的に損ないます。これらの制限に対処するために、私たちは LLM を確率的取引エージェントではなく量的研究者として再定義する原則に基づいた評価フレームワークである AlphaForgeBench を提案します。 AlphaForgeBench では、個別の取引アクションを生成する代わりに、実行可能なアルファ ファクターを生成し、金融知識に基づいたファクターベースの取引戦略を構築するモデルが必要です。このパラダイムは、推論を実行メカニズムから切り離し、現実世界の定量的研究ワークフローとの整合性を保ちながら、決定論的で再現可能な評価を可能にします。複数の最先端 LLM にわたる広範な実験により、AlphaForgeBench が実行に起因する不安定性を排除し、財務上の推論、戦略策定、およびアルファ発見を評価するための厳密なベンチマークを提供することが実証されました。ウェブページ https://finbrain-lab-hkustgz.github.io/AlphaForgeBench

原文 (English)

AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models

The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation toward interactive trading simulations. However, existing frameworks for evaluating real-time trading largely overlook a critical failure mode: the severe behavioral instability of LLMs in sequential decision-making under financial uncertainty. Through extensive experiments, we show that when deployed as trading agents, LLMs exhibit extreme run-to-run variance, generate inconsistent action sequences even under deterministic decoding, and frequently produce irrational action flipping across adjacent time steps. We attribute these behaviors to the stateless autoregressive nature of LLMs, which lack persistent memory of prior actions, together with their sensitivity to continuous-to-discrete action mappings in portfolio allocation tasks. These deficiencies fundamentally undermine the reliability and reproducibility of many existing online and offline trading benchmarks. To address these limitations, we propose AlphaForgeBench, a principled evaluation framework that redefines LLMs as quantitative researchers rather than stochastic trading agents. Instead of producing discrete trading actions, AlphaForgeBench requires models to generate executable alpha factors and compose factor-based trading strategies grounded in financial knowledge. This paradigm decouples reasoning from execution mechanics, enabling deterministic and reproducible evaluation while remaining aligned with real-world quantitative research workflows. Extensive experiments across multiple state-of-the-art LLMs demonstrate that AlphaForgeBench eliminates execution-induced instability and provides a rigorous benchmark for evaluating financial reasoning, strategy formulation, and alpha discovery. Webpage at https://finbrain-lab-hkustgz.github.io/AlphaForgeBench

2026-05-28 01:00 JSTTechCrunch AIビジネス/資金調達

AI coding startup Cognition raises $1B at $25B pre-money valuation

As Cognition reaches $492 million in annualized revenue run rate, it more than doubled its valuation in eight months, it says.

2026-05-27 22:04 JSTTechCrunch AIビジネス/資金調達

ClickHouse triples annualized revenue to $250M, charting a path toward an IPO

The database provider is eyeing a public debut within the next few years.

2026-05-27 21:30 JSTTechCrunch AIエージェントビジネス/資金調達

Robinhood now lets your AI agents trade stocks

While these agents would be able to read and analyze users' portfolios to come up with trading strategies and suggest investments, they'll…

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

制約の取得にはより優れたベンチマークが必要

制約取得 (CA) およびドメイン知識成果物からの数学的プログラミング (MP) モデルの検証と強化に関する関連研究は、現在、不適切なベンチマークによって制限されています。この欠陥により、再現性と研究間の比較可能性が妨げられ、CA 法の成熟が遅れます。既存のベンチマークは、CA アルゴリズムを評価するためではなく、ソルバーを評価するために設計されています。これらは大まかに編成されており、個々の問題の扱いに一貫性がなく、CA メソッドに必要なドメイン知識のアーティファクトが省略されています。この研究では、多様なドメイン知識アーティファクトを使用して MP ​​モデルを発見、検証、強化するアルゴリズムを評価するために設計されたベンチマーク スイートである MPMMine を紹介します。 MPMMine は、一貫性、標準化、完全性、拡張性、オープン性、バージョン管理によって導かれます。統一された構造を採用し、MiniZinc、CommonMark、JSON などのオープン フォーマットに依存しています。問題ごとに複数のモデル、モデルごとに数十のインスタンス、整数領域と連続領域の両方で数千の解と非解を提供し、テキストからモデルへの手法をサポートする自然言語記述も提供します。

原文 (English)

Constraint acquisition needs better benchmarks

Constraint Acquisition (CA) and related research on the validation and enhancement of Mathematical Programming (MP) models from domain knowledge artifacts are currently limited by inadequate benchmarks. This deficiency impedes reproducibility and cross-study comparability, slowing the maturation of CA methods. Existing benchmarks were designed for solver evaluation rather than for assessing CA algorithms. They are loosely organized, treat individual problems inconsistently, and omit the domain knowledge artifacts required by CA methods. This work presents MPMMine, a benchmark suite designed to assess algorithms that discover, validate, and enhance MP models using diverse domain knowledge artifacts. MPMMine is guided by consistency, standardization, completeness, extensibility, openness, and version control. It adopts a uniform structure and relies on open formats: MiniZinc, CommonMark, and JSON. It provides multiple models per problem, tens of instances per model, and thousands of solutions and non-solutions in both integer and continuous domains, alongside natural-language descriptions to support text-to-model methods.

2026-05-27 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

アンカー: エージェント ベンチマーク生成におけるアーティファクト ドリフトの軽減

AI エージェントは、長期にわたる価値のあるビジネス運営タスクを完了し始めていますが、企業の業務のためのトレーニングおよび評価環境は、依然として現実性、検証可能性、規模のバランスをとるのに苦労しています。環境とタスクの作成は、アーティファクト ドリフトと呼ばれる障害モードに頻繁に悩まされます。つまり、命令、環境、オラクル、およびベリファイアーが疎結合プロセスによって作成される場合、タスクに必要なものについて意見が一致しないことが多く、解決不可能、報酬ハック可能、または一貫性のない環境が生成されます。ドメイン専門家によるビジネス ワークフローの仕様を制約最適化プログラムに形式化するタスク生成パイプラインである Anchor を紹介します。パイプラインは、単一のパラメトリック仕様から、自然言語命令、環境構成、ソルバー認定のグラウンドトゥルース ソリューション、および状態ベースの検証器を共同で生成します。 Anchor を使用すると、パラメーターを変更すると、制御された難易度と既知の最適なソリューションを持つ新しいタスクが生成され、最終状態のビジネスの正しさのみに報酬が依存するハーネスに依存しない環境が生成されます。私たちは Anchor を適用して ERP-Bench を作成します。これは、生産グレードの ERP システムにおける調達と製造のワークフローにわたる 300 の長期タスクのベンチマークです。生成パラメータは現実の難易度を予測し、フロンティア モデルは試行の 26.1% で明示的なタスク制約を満たしますが、完全な最適解に到達するのは試行の 17.4% のみであることがわかりました。全体として、Anchor と ERP-Bench が、経済的に価値のあるエージェント作業のための監査可能な評価環境を構築するための具体的なレシピを提供することを示します。タスク ジェネレーターと ERP ベンチ データセットを erpbench.ai でリリースします。

原文 (English)

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation

AI agents are beginning to complete valuable, long-horizon business operations tasks, but training and evaluation environments for enterprise work still struggle to balance realism, verifiability, and scale. Environment and task creation frequently suffers from a failure mode we call artifact drift: when instructions, environments, oracles, and verifiers are created by loosely coupled processes, they frequently disagree on what a task requires, producing environments that are unsolvable, reward-hackable, or inconsistent. We introduce Anchor, a task-generation pipeline that formalizes domain experts' specifications of business workflows into constraint optimization programs. From a single parametric specification, the pipeline jointly produces a natural-language instruction, environment configuration, solver-certified ground-truth solution, and state-based verifier. With Anchor, altering parameters yields new tasks with controlled difficulty and known optimal solutions, producing harness-agnostic environments whose rewards depend solely on end-state business correctness. We apply Anchor to produce ERP-Bench: a benchmark of 300 long-horizon tasks spanning procurement and manufacturing workflows in a production-grade ERP system. We find that generation parameters predict realized difficulty, and that frontier models satisfy explicit task constraints in 26.1% of trials but reach a fully optimal solution in only 17.4% of trials. Overall, we show that Anchor and ERP-Bench offer a concrete recipe for building auditable evaluation environments for economically valuable agent work. We release the task generator and ERP-Bench dataset at erpbench.ai

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達

どの変更が重要ですか?関連性を重視した評価とソルバーに基づいた推論を通じて、信頼できる法律 AI を目指して

法的推論では、重要な変更とそうでない変更を区別する必要があります。法的 AI は、法的に無関係な摂動の下では安定した状態を維持する必要がありますが、摂動によって法的に重要な点が変更されると変化する必要があります。私たちはこの要件を法的関連性に敏感な評価問題として定式化します。つまり、LLM は法的に関連する変更のみに敏感であるべきです。私たちは、司法の公平性、堅牢性、および法令の混乱のシナリオ全体にわたって、変更すべき評価と変更すべきでない評価をカバーする統合評価スイートを導入します。私たちの評価によると、既存の法的 LLM は法的に無関係な変動に体系的に敏感であり、関連する法的要素と法的規則を区別できないことがよくあります。これらの失敗を軽減するために、形式的推論に基づいた敵対的なマルチエージェント フレームワークである LexGuard を紹介します。 LexGuard は、法令を実行可能な制約に形式化し、敵対的なエージェントを使用して競合する事実と法令の議論を抽出し、SMT ソルバーを呼び出して法的充足性と論理的一貫性を検証します。実験によると、LexGuard は、操作的な枠組みに対する脆弱性を軽減し、類似の法令間の曖昧さの解消を改善し、法的に無関係な属性の影響を制限し、良性の再定式化の下での一貫性を高めることにより、法的推論の信頼性を向上させます。法的信頼性には正確さだけでなく、法的に重要な変更に対する調整された感度も必要であることを示します。

原文 (English)

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

Legal reasoning requires distinguishing changes that matter from those that do not. Legal AI should remain stable under legally irrelevant perturbations, but should change when perturbations alter legally material points. We formulate this requirement as a legal-relevance-sensitive evaluation problem: LLMs should only be sensitive to the legally relevant change. We introduce a unified evaluation suite covering should-change and should-not-change evaluation across judicial fairness, robustness, and statute-confusion scenarios. Our evaluation shows that existing legal LLMs are systematically sensitive to legally irrelevant variations and often fail to distinguish related legal elements and statutory rules. To mitigate these failures, we present LexGuard, an adversarial multi-agent framework grounded in formal reasoning. LexGuard formalizes statutes into executable constraints, uses adversarial agents to extract competing fact-statute arguments, and invokes SMT solvers to verify legal satisfaction and logical consistency. Experiments show that LexGuard improves legal reasoning reliability by reducing vulnerability to manipulative framing, improving disambiguation among similar statutes, limiting the influence of legally irrelevant attributes, and increasing consistency under benign reformulations. We show that legal trustworthiness requires not only accuracy, but calibrated sensitivity to legally material changes.

2026-05-27 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

規範から指標へ (N2I-RAG): 法的指標計算のためのエージェントによる検索拡張生成フレームワーク

規範文書から法的指標を計算することは、法的監視と政策評価における重要なタスクですが、法的言語の複雑さ、規模、解釈の性質、および利用可能な文書の品質のばらつきにより、大きな課題が生じます。既存の自然言語処理技術と生成モデルは法的分析に役立ちますが、多くの場合、高い幻覚リスクに悩まされ、信頼性の高い指標の計算に必要な解釈可能性と根拠に欠けています。この文書では、透過的かつ追跡可能な方法で法的指標の計算を自動化するように設計されたエージェントによる検索拡張生成フレームワークである N2I-RAG (From Norm to Indicators) について説明します。当社は、適応型検索、llm ベースのエージェント、および検証メカニズムをモジュラー パイプラインに統合しており、各コンポーネントは証拠のフィルタリング、検索、評価において定義された役割を果たし、特定可能な法的条項に関連付けられたバイナリの法的結果を生成します。このフレームワークは、中間決定と最終的な指標の割り当ての明示的な説明を要求することで、トレーサビリティを強調しています。当社は、スキャンされたソースとデジタル ソースの両方を含む社内で構築されたフランス海洋環境法コーパスを使用して N2I-RAG を評価します。複数の言語モデル ファミリを使用した比較実験により、提案されたアプローチがベースライン システムよりも一貫して優れたパフォーマンスを示し、2 つの異なる禁止でテストした場合によく一般化されることが実証されました。この結果は、エージェントによる検索拡張生成がオープンテキストの法的言語と標準化された指標計算の橋渡しとなり、透明性と拡張性のある法的監視の基盤を提供できることを示しています。

原文 (English)

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

Computing legal indicators from normative texts is a key task in legal monitoring and policy evaluation, but presents significant challenges due to the complexity, scale, and interpretive nature of legal language, as well as the variability in available document quality. Existing natural language processing techniques and generative models can assist in legal analysis, but often suffer from high risk of hallucinations and lack the interpretability and evidence grounding required for reliable indicator computation. This paper presents N2I-RAG (From Norms to Indicators), an agentic retrieval-augmented generation framework designed to automate the computation of legal indicators in a transparent and traceable way. We integrate adaptive retrieval, llm-based agents, and validation mechanisms in a modular pipeline, where each component performs a defined role in filtering, retrieving, and assessing evidence, and in producing binary legal outcomes linked to identifiable legal provisions. The framework emphasizes traceability by requiring explicit explanations of intermediate decisions and final indicator assignments. We evaluate N2I-RAG using an in-house constructed French marine environmental law corpus that includes both scanned and digital sources. Comparative experiments with multiple language model families demonstrate that the proposed approach consistently outperforms baseline systems, and generalizes well when tested on 2 different bans. The results indicate that agentic retrieval-augmented generation can bridge open-text legal language and standardized indicator computation, offering a foundation for transparent and scalable legal observatories.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

検出は解決されていない: 検索拡張 LLM における監視制御ギャップ

検索拡張 LLM は、証拠の質がアクションの安全性を決定するタスクに導入されますが、評価プロトコルでは、ターンをまたいで証拠が蓄積された場合の堅牢性は、シングル ターンの堅牢性によって予測されると想定されています。この仮定が根本的に間違っていることを示します。モデルは監視と制御のギャップを示します。モデルは矛盾する証拠を容易に認識しますが、この認識は最終的な推奨事項を制約することができません。認識論的な矛盾を検出しても、それを安全に解決することを意味するわけではありません。 4 つのモデル ファミリ (1.5B ~ 32B パラメーター) にわたるマルチターン文書蓄積プロトコルと 50,000 を超えるターンレベル評価を通じて、シングルターン診断が体系的に RAG の安全性を過大評価していること、矛盾の認識が安全な解決と相関関係がないこと、対象を絞った人間による検証によって裏付けられたパターンであること、および普遍的な即時修正が存在しないことを実証します。収束するメカニズムの証拠 - 隠れ状態の調査、注意力の分析、および対応戦略の分類法 - は、欠陥の最もありそうな原因として行動の選択を示しています。危険に関連した情報は内部的に表現され、安全でない生成中に強化された注意を受けますが、出力の動作を制限することはできません。検索拡張システムを一か八かの状況で信頼できるようになる前に、モデルが認識する内容とモデルが実行する内容との間のギャップを測定し、埋める必要があります。

原文 (English)

Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

Retrieval-augmented LLMs are deployed for tasks where evidence quality determines action safety, yet evaluation protocols assume that single-turn robustness predicts robustness when evidence accumulates across turns. We show this assumption is fundamentally incorrect. Models exhibit a monitoring-control gap: they readily acknowledge contradictory evidence, yet this awareness fails to constrain their final recommendations - detecting epistemic conflict does not imply resolving it safely. Through a multi-turn document accumulation protocol across four model families (1.5B-32B parameters) and over 50,000 turn-level evaluations, we demonstrate that single-turn diagnostics systematically overestimate RAG safety, that contradiction acknowledgement is uncorrelated with safe resolution, a pattern corroborated by targeted human validation, and that no universal prompt fix exists. Converging mechanism evidence - hidden-state probing, attention analysis, and response-strategy taxonomy - points to action selection as the most plausible locus of the deficit: danger-relevant information is internally represented and receives enhanced attention during unsafe generation, yet fails to constrain output behavior. The gap between what models recognize and what they do must be measured and closed before retrieval-augmented systems can be trusted in high-stakes settings.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

MUSE-Autoskill: スキルの作成、記憶、管理、評価による自己進化エージェント

大規模言語モデル (LLM) エージェントは、再利用可能なスキルに依存して複雑なタスクを解決します。ただし、既存のスキル作成アプローチでは、スキルを孤立した静的な成果物として扱い、再利用性、信頼性、長期的な改善が制限されています。私たちは、統一されたライフサイクル (作成、記憶、管理、評価、洗練) の下でスキルを作成、再利用、洗練することにより、エージェントがタスク解決能力を継続的に向上できるようにする、スキル中心のエージェント フレームワークである MUSE-Autoskill Agent (Memory-Utilizing Skill Evolution) を提案します。当社のフレームワークにより、エージェントはオンデマンドでスキルを作成し、それらをタスク間で保存して再利用し、効率的に整理して選択し、単体テストや実行時のフィードバックを通じて評価して継続的に改善することができます。さらに、タスク全体にわたって各スキルの経験を蓄積するスキルレベルの記憶を導入し、時間の経過とともにより効果的な再利用と適応を可能にします。 SkillsBench の実験は、ライフサイクル管理されたスキルがタスクの成功、効率、再利用、およびエージェント間での移転を向上させることができるという最初の証拠を提供し、スキルを長命で経験を意識したテスト可能な資産として扱うことの重要性を強調しています。

原文 (English)

MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation

Large language model (LLM) agents rely on reusable skills to solve complex tasks. However, existing skill creation approaches treat skills as isolated and static artifacts, limiting their reusability, reliability, and long-term improvement. We propose MUSE-Autoskill Agent (Memory-Utilizing Skill Evolution), a skill-centric agent framework that lets agents continuously improve their task-solving capability by creating, reusing, and refining skills under a unified lifecycle (creation, memory, management, evaluation, and refinement). Our framework enables agents to create skills on demand, store and reuse them across tasks, organize and select them efficiently, and evaluate them through unit tests and runtime feedback for continuous refinement. We further introduce skill-level memory that accumulates experience for each skill across tasks, enabling more effective reuse and adaptation over time. Experiments on SkillsBench provide initial evidence that lifecycle-managed skills can improve task success, efficiency, reuse, and cross-agent transfer, highlighting the importance of treating skills as long-lived, experience-aware, and testable assets.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達

TSFMAudit: 予測時系列基盤モデルにおけるデータ汚染監査

時系列基盤モデル (TSFM) は大規模なコーパスで事前トレーニングされることが増えており、事前トレーニング中に評価データセットが公開され、過度に楽観的なパフォーマンス推定値が得られる可能性があるという懸念が生じています。信号は連続的かつ異質であり、多くの場合コーパス文書が欠如しているため、このような汚染を時系列で監査することは困難です。私たちの知る限り、これは TSFM の事前トレーニング汚染監査を研究する最初の研究です。我々は、TSFM の事前トレーニング汚染監査の問題を形式化し、プローブ適応ダイナミクスに基づく方法である TSFMAudit を提案します。私たちの重要な直観は、汚染が異常に効率的な適応として現れるということです。つまり、プローブを微調整した後、汚染されたデータセットはバックボーンの動きが小さくなり、より迅速な損失削減を示す傾向があります。私たちは、文書化されたトレーニングソースの証拠を監督として使用して、6 つの TSFM と 187 のデータセットで TSFMAudit を評価し、LLM 文献から適応された 10 の競合ベースラインと比較します。

原文 (English)

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models

Time series foundation models (TSFMs) are increasingly pretrained on large corpora, raising concerns that evaluation datasets may have been exposed during pretraining and thus yield overly optimistic performance estimates. Auditing such contamination is challenging in time series because signals are continuous and heterogeneous, and often lack corpus documentation. To the best of our knowledge, this is the first work to study pretraining contamination auditing for TSFMs. We formalize the problem of pretraining contamination auditing for TSFMs and propose TSFMAudit, a method based on probe adaptation dynamics. Our key intuition is that contamination manifests as unusually efficient adaptation: after a fine tuning probe, contaminated datasets tend to exhibit faster loss reduction with smaller backbone movement. We evaluate TSFMAudit on 6 TSFMs and 187 datasets using documented training source evidence as supervision, and compare against 10 competitive baselines adapted from the LLM literature.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達

多様な呼吸不全予測の前向き評価: 胸部 X 線撮影は EHR 信号を超えてパフォーマンスを向上させますか?

呼吸不全の早期予測は、集中治療室でのタイムリーな臨床介入にとって重要です。既存の電子健康記録 (EHR) ベースのモデルは、生理学的悪化を継続的に監視できますが、胸部 X 線写真 (CXR) に反映される肺の病態生理学を完全には捕捉できない可能性があります。この研究では、CXR 情報が EHR 信号のみを超えて侵襲的人工呼吸器の前向き予測を改善するかどうかを尋ねます。私たちは、構造化された EHR 時系列データと CXR 基盤モデル表現を統合するゲート付きマルチモーダル フレームワークを開発します。ゲーティング モジュールは、患者固有の臨床状況に基づいてイメージング機能の寄与を適応的に制御し、モデルが有益な場合にイメージング情報に選択的に依存できるようにします。私たちは、ICU患者の24時間以内の侵襲的人工呼吸器を予測するためのフレームワークを前向きに評価し、確立されたEHR専用モデル(Vent.io)、一致する臨床時点で得られた医師の予測、および代替のマルチモーダルバリアントと比較します。ゲート付きマルチモーダル モデルは、EHR のみのベースラインよりも高い識別を達成し、Vent.io の 0.752 と比較して、REMEDIS および MedInsight CXR 表現を使用した AUROC 値はそれぞれ 0.860 と 0.858 でした。医師の予測と比較して、マルチモーダルフレームワークは良好な特異性を維持しながら感度を大幅に向上させました。 EHR のみのモデルと比較して、マルチモーダル統合により特異性と陽性的中率が向上したことは、CXR 情報が選択された患者のリスク推定を精緻化できることを示唆しています。これらの発見は、画像処理を呼吸不全の予測に組み込むための実用的な戦略として、適応型マルチモーダル融合を裏付けるものです。

原文 (English)

Prospective evaluation of multimodal respiratory failure prediction: Do chest X-rays improve performance beyond EHR signals?

Early prediction of respiratory failure is critical for timely clinical intervention in intensive care units. Existing electronic health record (EHR)-based models can continuously monitor physiologic deterioration, but they may not fully capture pulmonary pathophysiology reflected in chest radiographs (CXRs). In this study, we ask whether CXR information improves prospective prediction of invasive mechanical ventilation beyond EHR signals alone. We develop a gated multimodal framework that integrates structured EHR time-series data with CXR foundation-model representations. The gating module adaptively controls the contribution of imaging features based on patient-specific clinical context, allowing the model to selectively rely on imaging information when it is informative. We prospectively evaluate the framework for predicting invasive mechanical ventilation within 24 hours in ICU patients and compare it with an established EHR-only model (Vent.io), physician predictions obtained at matched clinical time points, and alternative multimodal variants. The gated multimodal models achieved higher discrimination than the EHR-only baseline, with AUROC values of 0.860 and 0.858 using REMEDIS and MedInsight CXR representations, respectively, compared with 0.752 for Vent.io. Relative to physician predictions, the multimodal framework substantially improved sensitivity while maintaining favorable specificity. Compared with the EHR-only model, multimodal integration increased specificity and positive predictive value, suggesting that CXR information can refine risk estimation in selected patients. These findings support adaptive multimodal fusion as a practical strategy for incorporating imaging into prospective respiratory failure prediction.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達

LURE: 評価意識を軽減するためのライブ使用リプレイ評価

大規模な言語モデルは、いつ評価されているかを認識し (評価認識)、そのために異なる動作をする可能性があり、安全性と調整ベンチマークの有効性が損なわれます。我々は、現実的なエージェントインタラクションの軌跡を再生し、最後に評価プロンプトを追加することによって展開のような評価を構築するための方法である、LURE (Live-Usage Replay Evavals) を提案します。また、言語化された評価認識の検出とログが評価である確率の判断モデル推定を組み合わせて、評価の現実性を測定するための自動パイプラインを導入し、デプロイメントと評価トランスクリプトの大規模なデータセットで検証します。私たちは、LURE ベースの評価は、広く使用されているベンチマークや合成評価ジェネレーターに比べて導入との区別が大幅に難しく、ユーザーとの実際の会話のリアリズムに近づくことができることを発見しました。私たちは、陰謀、AI の安全妨害、おべっかの設定で LURE をインスタンス化します。私たちの結果は、評価の現実性がアライメントベンチマークの重要な特性であり、特にそのような結果が安全性のケースで使用される場合には、ベンチマーク結果と一緒に報告されるべきであることを示唆しています。

原文 (English)

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the validity of safety and alignment benchmarks. We propose LURE (Live-Usage Replay Evaluations), a method for constructing deployment-like evaluations by replaying realistic agentic interaction trajectories and appending evaluation prompt at the end. We also introduce an automated pipeline for measuring evaluation realism, combining detection of verbalized evaluation awareness and judge-model estimates of the probability of logs being an evaluation, and validate it on a large dataset of deployment and evaluation transcripts. We find that LURE-based evaluations are substantially less distinguishable from deployment than widely used benchmarks and synthetic evaluation generators, and can approach the realism of real conversations with users. We instantiate LURE in scheming, AI safety sabotage, and sycophancy settings. Our results suggest that evaluation realism is a crucial property of alignment benchmarks and should be reported alongside benchmark results, especially when such results are used in safety cases.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

AI 評価は認識にバイアスを与える可能性がある: 学術論文を解釈する際の文脈の重要性

この論文では、評価方法が国や分野間の文脈の違いを無視した場合、科学論文における AI 使用の推定値にどのような偏りが生じる可能性があるかを検証します。 Dimensions の雑誌出版物に関する大規模データを使用して、人間が書いた抄録と LLM で言い換えた抄録の違いに基づいて AI らしさのベンチマークを構築します。私たちは、プールされたベンチマークが既存の文体のバリエーションと AI が生成したテキストを混同し、LLM 以前の出版物であっても国分野グループ全体に大きな歪みを生じさせる可能性があることを示します。対照的に、国別分野固有のベンチマークはそのような歪みを軽減し、比較のためのより信頼できるベースラインを提供します。これらの手法を 2025 年の出版物に適用すると、プールされたベンチマークが、特定の国や分野では AI の使用を体系的に過大評価し、他の国や分野では過小評価していることが明らかになります。これらの調査結果は、科学における AI の使用を正確かつ公平に評価するためのコンテキストを意識した測定の重要性を浮き彫りにしています。

原文 (English)

AI evaluation may bias perceptions: The importance of context in interpreting academic writing

This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries and fields. Using large-scale data on journal publications from Dimensions, we construct AI-likeness benchmarks based on differences between human-written and LLM-rephrased abstracts. We show that a pooled benchmark may confound pre-existing stylistic variation with AI-generated text, producing substantial distortions across country-field groups even in pre-LLM publications. In contrast, country-field-specific benchmarks attenuate such distortions and provide a more credible baseline for comparison. Applying these methods to publications in 2025 reveals that the pooled benchmark systematically overestimates AI use in certain countries and fields while underestimating it in others. These findings highlight the importance of context-aware measurement for accurate and equitable evaluation of AI use in science.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達

SL-BiLEM: 予測と政策評価のための構造化された学習可能な Behavior-in-the-Loop 流行モデリング

流行予測は根本的な課題に直面しています。それは、人間の行動が病気の蔓延に動的に反応し、政策介入時点で分布の変化を引き起こすフィードバック ループを生み出すということです。これにより、分布の変化の下ではデータ駆動型モデルの信頼性が低くなります。私たちは、堅牢な外挿のための正則化として物理的制約を利用する \textbf{SL-BiLEM} (構造化学習可能な行動インザループ流行モデル) を提案します。このフレームワークは、有効な伝送を $\beta_{\text{eff}}(t,g) = \beta_0(g) \times m_{\text{policy}}(t) \times m_{\text{media}}(t) \times m_{\text{comp}}(t,g)$ として分解します。ここで、学習されたコンプライアンス関数の単調性、滑らかさ、および境界ジャンプ制約は、新しいポリシーの下で予測の妥当性を維持します。体制。 SL-BiLEM は予測を超えて、介入の意思決定をサポートするための反事実分析を可能にします。私たちは 3 つの現実世界のデータセット (クルーズ船、学校インフルエンザ、学校区の 新型コロナウイルス感染症 (COVID-19) 監視) で予測を検証し、既知のグラウンド トゥルースを使用した合成ベンチマークで反事実の回復を評価します。 SL-BiLEM は次のことを実証します。(1) 神経機構ベースラインに対して 76% 改善し、OOD 低下はわずか 53% であったのに対し、政策誘発シフト下の神経ベースラインでは 1142% でした。 (2) 27 の合成反事実実験にわたる 100% のブートストラップ CI カバレッジ。 (3) 治療効果の精度が 0.85 を超える。これらの結果により、SL-BiLEM は、正確な予測と原則に基づいた介入計画を求める公衆衛生の意思決定者にとって解釈可能なツールとして確立されています。

原文 (English)

SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation

Epidemic forecasting faces a fundamental challenge: human behavior dynamically responds to disease spread, creating feedback loops that induce distribution shifts at policy intervention points. This renders data-driven models unreliable under distribution shift. We propose \textbf{SL-BiLEM} (Structured Learnable Behavior-in-the-Loop Epidemic Model), leveraging physical constraints as regularization for robust extrapolation. The framework decomposes effective transmission as $\beta_{\text{eff}}(t,g) = \beta_0(g) \times m_{\text{policy}}(t) \times m_{\text{media}}(t) \times m_{\text{comp}}(t,g)$, where monotonicity, smoothness, and bounded-jump constraints on the learned compliance function maintain predictive validity under novel policy regimes. Beyond forecasting, SL-BiLEM enables counterfactual analysis for intervention decision support. We validate forecasting on three real-world datasets (cruise ship, school influenza, and school-district COVID-19 surveillance) and evaluate counterfactual recovery on synthetic benchmarks with known ground truth. SL-BiLEM demonstrates: (1) 76\% improvement over neural-mechanistic baselines, with only 53\% OOD degradation versus 1142\% for neural baselines under policy-induced shift; (2) 100\% bootstrap CI coverage across 27 synthetic counterfactual experiments; and (3) Treatment Effect Accuracy exceeding 0.85. These results establish SL-BiLEM as an interpretable tool for public health decision-makers seeking accurate prediction and principled intervention planning.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

MatFormBench: ターゲット主導の材料配合のためのベンチマーク評価フレームワーク

材料の逆設計により、ターゲット駆動型の配合最適化が大幅に進歩しましたが、既存の材料機械学習ベンチマークは依然として順特性予測に限定されており、逆最適化と生成アルゴリズムを系統的に評価できず、これがターゲット駆動型の材料設計の進歩を妨げる重大なギャップです。この制限に対処するために、私たちは、ターゲット駆動型の定式化のための生成戦略を評価およびガイドするように調整された新しいベンチマーク エコシステムである MatFormBench を提案します。 MatFormBench は、物理学に基づいた配合生成スキームを統合して、現実的な材料の構造と特性の応答関係を忠実にエミュレートする合成サンプルを生成します。これに、これらの関係の複雑さを定量化するための 5 段階の難易度レベルが追加されます。アルゴリズムのパフォーマンスを厳密に評価するために、ターゲットの成功、探索効率、探索能力、堅牢性、安定性という 5 つの重要な軸にわたってパフォーマンスを包括的に定量化する多次元メトリクスである MatFormScore をさらに提案します。私たちは、古典的なサロゲート支援ブラックボックス検索、最先端の深層生成モデル、およびますます普及している大規模言語モデル (LLM) ベースのレコメンデーション戦略をカバーする、39 の多様な逆設計アルゴリズムを評価することによって MatFormBench を検証します。 1,170 の標準化されたアルゴリズム タスク評価全体で、拡散ベースのモデルが最も強力な全体的なパフォーマンスを示しますが、変分オートエンコーダー (VAE) ベースおよび遺伝的アルゴリズム (GA) ベースの手法は、特定のシナリオで明確な利点を示します。 MatFormBench は、ターゲット主導の材料配合のための統一評価基準を確立することで、再現可能なベンチマーク、原則に基づいたアルゴリズムの比較、逆設計戦略の診断分析を可能にし、材料の逆設計を進めるための基礎ツールを提供します。

原文 (English)

MatFormBench: A Benchmarking Evaluation Framework for Target-Driven Materials Formulation

Inverse design of materials has significantly advanced target-driven formulation optimization, yet existing materials machine learning benchmarks remain limited to forward property prediction, failing to systematically evaluate inverse optimization and generation algorithms, a critical gap that hinders the progress of target-driven materials design. To address this limitation, we propose MatFormBench, a novel benchmarking ecosystem tailored to evaluate and guide generative strategies for target-driven formulation. MatFormBench integrates a physics-driven formulation generation scheme to generate synthetic samples that faithfully emulate realistic materials structure-property response relationships, complemented by five escalating difficulty levels to quantify the complexity of these relationships. To rigorously assess algorithm performance, we further propose MatFormScore, a multi-dimensional metric that comprehensively quantifies performance across five critical axes: target success, search efficiency, exploratory capacity, robustness, and stability. We validate MatFormBench by evaluating 39 diverse inverse design algorithms, covering classical surrogate-assisted black-box search, state-of-the-art deep generative models, and increasingly popular Large Language Model (LLM)-based recommendation strategies. Across 1170 standardized algorithm-task evaluations, diffusion-based models demonstrate the strongest overall performance, while Variational Autoencoder (VAE)-based and Genetic Algorithm (GA)-based methods exhibit distinct advantages in specific scenarios. By establishing a unified evaluation standard for target-driven materials formulation, MatFormBench enables reproducible benchmarking, principled algorithm comparison, and diagnostic analysis of inverse design strategies, providing a foundational tool for advancing materials inverse design.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達

EEG-FM-Audit: EEG 基礎モデルの体系的な評価および分析パイプライン

大規模な EEG Foundation Model (FM) は、さまざまな認知タスクにわたる EEG 信号をデコードする大きな可能性を示しています。しかし、既存の EEG-FM 研究には、不透明な教師付きベースライン調整、複雑な学習パラダイムの未検証の寄与、およびモデルの意思決定における透明性の欠如という 3 つの重大な制限があります。これらに対処するために、EEG-FM の評価を体系化するために設計された包括的な評価および分析パイプラインである EEG-FM-Audit を提案します。 EEG-FM-Audit は 3 つの主要コンポーネントで構成されます。(1) 教師付きベースラインを透過的に最適化することで公平な比較を保証する、ASHA 主導のベンチマーク プロトコル。 (2) FM における学習パラダイムの有効性を評価するためのパラダイムレベルのアブレーション研究。 (3) FM が有効な時間的、空間的、スペクトル的 EEG 特性を活用しているかどうかを調査する神経生理学的プローブ (NPP) フレームワーク。私たちは、EEG-FM-Audit を 3 つの公開データセットにわたる 4 つの最先端の EEG-FM と 5 つの代表的な教師ありモデルに適用します。私たちの結果は、適切に調整された教師ありベースラインが、必要なパラメーターが大幅に少ないにもかかわらず、高度な FM と同等またはそれを上回るパフォーマンスを発揮できることを明らかにしました。さらに、FM の学習パラダイムの有効性はデータセットの規模とアーキテクチャに大きく依存することがわかりました。最後に、NPP 分析は、FM が特定の生理学的特徴にどのように依存しているかを示し、より解釈可能な神経解読のためのフレームワークを確立します。

原文 (English)

EEG-FM-Audit: A Systematic Evaluation and Analysis Pipeline for EEG Foundation Models

Large EEG Foundation Models (FMs) have shown great potential for decoding EEG signals across diverse cognitive tasks. However, existing EEG-FM studies exhibit three critical limitations: opaque supervised baseline tuning, unverified contributions of complex learning paradigms, and a lack of transparency in model decision-making. To address these, we propose EEG-FM-Audit, a comprehensive evaluation and analysis pipeline designed to systematize the assessment of EEG-FMs. EEG-FM-Audit consists of three primary components: (1) an ASHA-driven benchmarking protocol that ensures fair comparisons by transparently optimizing supervised baselines; (2) paradigm-level ablation studies to evaluate the effectiveness of learning paradigms in FMs; and (3) a neurophysiological probing (NPP) framework, which explores whether FMs leverage valid temporal, spatial, and spectral EEG properties. We apply EEG-FM-Audit to four state-of-the-art EEG-FMs and five representative supervised models across three public datasets. Our results reveal that properly tuned supervised baselines can match or outperform advanced FMs, despite requiring significantly fewer parameters. Furthermore, we find that the effectiveness of learning paradigms of FMs is highly dependent on dataset scale and architecture. Finally, NPP analysis demonstrates how FMs rely on specific physiological features, establishing a framework for more interpretable neural decoding.

2026-05-27 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達規制/政策

画像生成モデルの事前トレーニングデータに対するブラックボックスメンバーシップ推論攻撃

拡散ベースの画像生成モデルの急速な進歩により、人間が作成したデータに関わる著作権およびプライバシー侵害の可能性について深刻な懸念が生じています。メンバーシップ推論攻撃 (MIA) は、モデルのトレーニング中に不正なデータの使用を特定するための有望なツールとして浮上しています。既存の方法は通常、メンバーシップのステータスの指標として、乱れた疑わしい画像のノイズを除去するモデルの能力を評価します。ただし、そのような特徴の識別力はモデルの記憶の程度に大きく依存し、あまり公開されていないデータ (トレーニング前のデータなど) に適用すると大幅に低下します。いくつかの方法では、内部モデル機能を活用して検出を強化しようとしていますが、これらの機能は一般に、主流のクローズドソース画像生成プラットフォームではアクセスできず、実用性が制限されています。この論文では、ブラックボックス拡散モデルがターゲット画像と対応する摂動されたテキスト命令のノイズをどのように除去するかを分析することで、より特徴的なメンバーシップの手がかりを明らかにできることを実証します。この洞察に基づいて、クロスモーダル データ摂動メカニズムを利用して拡散モデルの事前トレーニング データを検出するブラック ボックス メンバーシップ推論攻撃フレームワーク (SD-MIA と呼ばれる) を提案します。私たちは、公開ベンチマーク データセットと新しく構築されたデータセットの両方で広範な実験を実施します。各データセットは、同一の分布を持つトレーニング前のメンバーシップ サンプルと非メンバーシップ サンプルで構成されます。実験結果は、SD-MIA が、内部モデル機能にアクセスするという不公平な利点を持つベースラインを含む、既存のベースラインと比較して優れたパフォーマンスを達成することを示しています。

原文 (English)

Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models

The rapid advancement of diffusion-based image generation models has raised serious concerns regarding potential copyright and privacy infringements involving human-created data. Membership inference attacks (MIAs) have emerged as a promising tool for identifying unauthorized data usage during model training. Existing methods typically assess the ability of model to denoise perturbed suspect images as an indicator of membership status. However, the discriminative power of such features is highly dependent on the degree of model memorization and deteriorates significantly when applied to less exposed data (e.g., pre-training data). Although several methods attempt to enhance detection by leveraging internal model features, these features are generally inaccessible in mainstream closed-source image generation platforms, limiting their practicality. In this paper, we demonstrate that analyzing how a black-box diffusion model denoises a target image and corresponding perturbed textual instructions can reveal more distinctive membership cues. Based on this insight, we propose a black-box membership inference attack framework (named SD-MIA) that leverages a cross-modal data perturbation mechanism to detect pre-training data in diffusion models. We conduct extensive experiments on both a public benchmark dataset and a newly constructed dataset, each comprising pre-training membership and non-membership samples with identical distributions. Experimental results demonstrate that SD-MIA achieves superior performance compared to existing baselines, including those with the unfair advantage of accessing internal model features.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

Qiskit QuantumKatas: Microsoft の量子コンピューティング演習を LLM 評価に適応させる

当社は、確立された量子コンピューティング カリキュラムである Microsoft の QuantumKatas を Q# から、最も広く採用されている量子コンピューティング フレームワークである Qiskit に適応させ、系統的な LLM 評価のための評価フレームワークとパッケージ化しています。結果として得られたベンチマークは、高度なアルゴリズム (Grover のアルゴリズム、Simon のアルゴリズム、Deutsch-Jozsa のアルゴリズム)、エラー修正、キー配布、および量子ゲームを介して基本的なゲートにわたる 26 のカテゴリにわたる 350 のタスクで構成されています。各タスクには、自然言語プロンプト、標準的な解決策、古典的な回路シミュレーションによる決定論的テスト検証が含まれます。タスクをゼロから作成するのではなく、QuantumKatas の実証済みの教育的デザインに基づいて構築することで、フレームワークの適応、評価インフラストラクチャ、および実証分析に貢献しながら、原則に基づいた難易度の進行と包括的な概念の網羅性を継承します。ベンチマークの有用性を実証するために、7 つのプロンプト構成にわたって 16 個の LLM (合計 39,200 回のモデル実行) を評価しました。 3 つの重要な結果が明らかになりました。(1) このベンチマークはモデルの機能を効果的に区別しており、最適構成の合格率は 32.3% ~ 83.1% の範囲であり、フロンティア モデルとオープンソース モデル間の平均ギャップは 26.1 pp です。 (2) モデルは、既知のアルゴリズム (SimonsAlgorithm 82.1%、BasicGates 81.6%) の実装ではうまく機能しますが、問題のエンコードに苦労しています (SolveSATWithGrover 34.4%、DistinguishUnitaries 40.0%)。 (3) 思考連鎖プロンプトは、適度な二峰性の効果を示しています。これは 3 つのモデル (そのうちの 2 つはベンダーのドキュメントごとに明示的に推論調整されています) にとっては最良の戦略ですが、残りのモデルではパフォーマンスが低下し、合計で中位 (平均 56.3%) となり、少数ショット 5 (57.8%) に次ぐ結果になります。量子コンピューティングにおける LLM 機能の研究をサポートするために、ベンチマーク、評価フレームワーク、ベースライン結果をリリースします。

原文 (English)

Qiskit QuantumKatas: Adapting Microsoft's Quantum Computing exercises for LLM evaluation

We adapt Microsoft's QuantumKatas -- a well-established quantum computing curriculum -- from Q# to Qiskit, the most widely-adopted quantum computing framework, and package it with an evaluation framework for systematic LLM assessment. The resulting benchmark comprises 350 tasks across 26 categories, spanning fundamental gates through advanced algorithms (Grover's, Simon's, Deutsch-Jozsa), error correction, key distribution, and quantum games. Each task includes a natural language prompt, canonical solution, and deterministic test verification via classical circuit simulation. By building on the QuantumKatas' proven pedagogical design rather than creating tasks from scratch, we inherit a principled difficulty progression and comprehensive concept coverage while contributing the framework adaptation, evaluation infrastructure, and empirical analysis. We evaluate 16 LLMs across 7 prompting configurations -- a total of 39,200 model runs -- to demonstrate the benchmark's utility. Three key findings emerge: (1) the benchmark effectively differentiates model capabilities, with best-configuration pass rates ranging from 32.3% to 83.1% and a 26.1 pp average gap between frontier and open-source models; (2) models perform well at implementing known algorithms (SimonsAlgorithm 82.1%, BasicGates 81.6%) but struggle with problem encoding (SolveSATWithGrover 34.4%, DistinguishUnitaries 40.0%); and (3) chain-of-thought prompting shows a modestly bimodal effect -- it is the best strategy for three models (two of them explicitly reasoning-tuned per vendor documentation) but degrades performance for the rest, leaving it mid-pack in aggregate (56.3% mean) behind few-shot-5 (57.8%). We release the benchmark, evaluation framework, and baseline results to support research on LLM capabilities in quantum computing.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達

AI を活用した貢献度の評価と競合解決: グループのワークロード調査のためのフレームワークと設計

チーム内の個人の貢献を公平に評価することは依然として根強い課題であり、仕事量の対立や格差により不公平なパフォーマンス評価が生じる可能性があり、多くの場合手動介入が必要となり、コストが高く困難なプロセスとなります。私たちは既存のツールの機能を調査し、競合解決方法と AI 統合におけるギャップを特定します。これに対処するために、紛争調査を支援する新しい AI 強化ツールのフレームワークと実装設計を提案します。このフレームワークは、提出物 (コード、テキスト、メディア)、コミュニケーション (チャット、電子メール)、調整記録 (会議ログ、タスク)、ピア評価、コンテキスト情報といった異種の成果物を、貢献、インタラクション、役割という 9 つのベンチマークを使用して 3 つの次元に整理します。客観的な尺度は正規化され、次元ごとに集計され、不平等尺度 (ジニ指数) と組み合わせられて、競合マーカーが表面化されます。大規模言語モデル (LLM) アーキテクチャは、これらの尺度に対して検証済みのコンテキスト分析を実行し、解釈可能で透明性のある勧告的判断を生成します。私たちは、現在の法定政策および制度政策の下での実現可能性を主張し、実践的な分析 (感情、タスクの忠実度、単語/行数など)、バイアスの保護手段、制限、および実際的な課題について概説します。

原文 (English)

AI-Driven Contribution Evaluation and Conflict Resolution: A Framework & Design for Group Workload Investigation

The equitable assessment of individual contribution in teams remains a persistent challenge, where conflict and disparity in workload can result in unfair performance evaluation, often requiring manual intervention - a costly and challenging process. We survey existing tool features and identify a gap in conflict resolution methods and AI integration. To address this, we propose a framework and implementation design for a novel AI-enhanced tool that assists in dispute investigation. The framework organises heterogeneous artefacts - submissions (code, text, media), communications (chat, email), coordination records (meeting logs, tasks), peer assessments, and contextual information - into three dimensions with nine benchmarks: Contribution, Interaction, and Role. Objective measures are normalised, aggregated per dimension, and paired with inequality measures (Gini index) to surface conflict markers. A Large Language Model (LLM) architecture performs validated and contextual analysis over these measures to generate interpretable and transparent advisory judgments. We argue for feasibility under current statutory and institutional policy, and outline practical analytics (sentimental, task fidelity, word/line count, etc.), bias safeguards, limitations, and practical challenges.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

LLM ベースのエージェント評価のための統一フレームワークの必要性

Large Language Model (LLM) の出現により、汎用エージェントは根本的な進歩を遂げました。ただし、これらのエージェントの評価には、静的な QA ベンチマークとは異なる特有の課題が存在します。現在のエージェントのベンチマークは、システム プロンプト、ツールセット構成、環境ダイナミクスなどの外部要因によって大きく混乱していることが観察されています。既存の評価は、断片化された研究者固有のフレームワークに依存していることが多く、推論やツールの使用に関する即時エンジニアリングが大幅に異なるため、パフォーマンスの向上がモデル自体によるものであると考えるのが困難です。さらに、標準化された環境データが欠如しているため、追跡不可能なエラーや再現不可能な結果が発生します。この標準化の欠如は、現場に大きな不公平性と不透明性をもたらします。私たちは、エージェント評価を厳密に進めるためには、統一された評価枠組みが不可欠であると提案します。この目的を達成するために、エージェント評価の標準化を目的とした提案を紹介します。

原文 (English)

The Necessity of a Unified Framework for LLM-Based Agent Evaluation

With the advent of Large Language Models (LLMs), general-purpose agents have seen fundamental advancements. However, evaluating these agents presents unique challenges that distinguish them from static QA benchmarks. We observe that current agent benchmarks are heavily confounded by extraneous factors, including system prompts, toolset configurations, and environmental dynamics. Existing evaluations often rely on fragmented, researcher-specific frameworks where the prompt engineering for reasoning and tool usage varies significantly, making it difficult to attribute performance gains to the model itself. Additionally, the lack of standardized environmental data leads to untraceable errors and non-reproducible results. This lack of standardization introduces substantial unfairness and opacity into the field. We propose that a unified evaluation framework is essential for the rigorous advancement of agent evaluation. To this end, we introduce a proposal aimed at standardizing agent evaluation.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

固定ベンチマークと最悪の場合の攻撃を超えて: 言語モデルの動的境界評価

現在の大規模言語モデル (LLM) の評価は、同じ項目セットを任意のモデルに適用する固定ベンチマークに基づいて行われており、機能のギャップを隠す天井効果と床効果が生じます。私たちは、最も有益な評価信号は境界にあり、ランダム サンプリング デコードではプロンプトごとの合格確率が $0.5$ 近くになると主張し、各モデルの境界を積極的に特定し、グローバルに比較可能な難易度スケールに配置する動的境界評価 (DBE) を提案します。 DBE は 3 つのアーティファクトを提供します。(i) $9$ の参照 LLM で検証されたアイテムごとの難易度ラベルを備えた、安全性、機能、真実性をカバーする調整されたアイテム バンク。 (ii) スキルガイド境界検索 (SGBS)。API レベルのクエリ アクセスのみを使用して、特定のターゲット LLM の境界アイテムを見つける検索アルゴリズムです。 (iii) 新しい LLM を統一された能力スケールに配置し、ターゲットが銀行の範囲外になる場合に適応的に評価セットを拡大する評価プロトコル。安全性 (有害な要求の拒否と過剰な拒否)、能力 (制約された指示に従う)、真実性 (複数ターンのおべっかへの耐性) にわたる 4 つのカテゴリに基づいて DBE をインスタンス化します。結果として得られる評価は、既存のデータセットとの互換性を維持しながら、飽和することなくより広いモデル範囲をカバーします。

原文 (English)

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

Evaluating large language models (LLMs) today rests on fixed benchmarks that apply the same set of items to any model, producing ceiling and floor effects that mask capability gaps. We argue that the most informative evaluation signal lies at the boundary, where the per-prompt pass probability is near $0.5$ under random-sampling decoding, and propose Dynamic Boundary Evaluation (DBE), which actively locates each model's boundary and places it on a globally comparable difficulty scale. DBE delivers three artifacts: (i) a calibrated item bank covering safety, capability, and truthfulness, with per-item difficulty labels validated across $9$ reference LLMs; (ii) Skill-Guided Boundary Search (SGBS), a search algorithm that finds boundary items for a given target LLM using only API-level query access; and (iii) an evaluation protocol that places a new LLM on a unified ability scale and grows the evaluation set adaptively when the target falls outside the bank's coverage. We instantiate DBE on four categories spanning safety (harmful request refusal and over-refusal), capability (constrained instruction following), and truthfulness (multi-turn sycophancy resistance). The resulting evaluation covers a broader model spectrum without saturation while remaining compatible with existing datasets.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

AgentAtlas: LLM エージェント向けの Beyond Outcome リーダーボード

現在、大規模な言語モデル エージェントは、コードベース、ブラウザー、オペレーティング システム、カレンダー、ファイル、ツール エコシステム上で動作しますが、その評価により、動作が最終的なタスクの成功に影響を与えることがよくあります。 AgentAtlas は、エージェントの評価を診断語彙および監査プロトコルとして再構築し、結果の成功を制御決定の品質および軌道の品質から分離します。この論文は以下の内容に貢献しています。(i) 6 つの状態の制御と決定の分類 (行為 / 要求 / 拒否 / 停止 / 確認 / 回復)。 (ii) 主要なエラー原因と下流への影響を伴う軌道障害語彙。 (iii) 15 のエージェントベンチマークに対する 0/1/2 ベンチマークカバレッジ監査。 (iv) 分類学を意識したプロンプト形式と分類学を意識しないプロンプト形式で 8 つのモデルを使用して評価した合成 1,342 項目セットに関する例示的なプロトコル研究。合成デモンストレーションは公開ベンチマーク リリースではないため、最終的なモデル比較として解釈しないでください。代わりに、これは 2 つの測定リスクを示しています。明示的なラベル メニューが削除されると、マッピングされたラベルの一致が大幅に変化する可能性があることと、軸の選択により見かけのランキングが変化する可能性があることです。 AgentAtlas は、ベンチマークの設計者がどのような動作をカバーしているのかを明らかにし、評価者が結果のみのリーダーボードに隠れている障害を診断できるようにすることを目的としています。

原文 (English)

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but their evaluations often collapse behavior into final task success. AgentAtlas reframes agent evaluation as a diagnostic vocabulary and audit protocol for separating outcome success from control-decision quality and trajectory quality. The paper contributes: (i) a six-state control-decision taxonomy (Act / Ask / Refuse / Stop / Confirm / Recover); (ii) a trajectory-failure vocabulary with primary error source and downstream impact; (iii) a 0/1/2 benchmark-coverage audit over fifteen agent benchmarks; and (iv) an illustrative protocol study on a synthetic 1,342-item set evaluated with eight models under taxonomy-aware and taxonomy-blind prompt formats. The synthetic demonstration is not a public benchmark release and should not be read as a definitive model comparison. Instead, it illustrates two measurement risks: mapped label agreement can change substantially when the explicit label menu is removed, and axis choice can change apparent rankings. AgentAtlas is intended to help benchmark designers state what behavior they cover, and to help evaluators diagnose failures that outcome-only leaderboards hide.

2026-05-27 13:00 JSTarXiv cs.AI画像/動画生成ビジネス/資金調達

「PhyWorldBench」: Text-to-Video モデルにおける物理的リアリズムの包括的な評価

ビデオ生成モデルは、高品質でフォトリアリスティックなコンテンツの作成において目覚ましい進歩を遂げました。しかし、物理現象を正確にシミュレートする能力は依然として重要かつ未解決の課題です。このペーパーでは、物理法則の遵守に基づいてビデオ生成モデルを評価するために設計された包括的なベンチマークである PhyWorldBench について説明します。このベンチマークは、物体の動きやエネルギー保存などの基本原理から、剛体の相互作用や人間や動物の動きを含むより複雑なシナリオに至るまで、複数のレベルの物理現象をカバーします。さらに、プロンプトが現実世界の物理学を意図的に侵害する新しいアンチフィジックス カテゴリを導入し、論理的一貫性を維持しながらモデルがそのような指示に従うことができるかどうかの評価を可能にします。大規模な人による評価に加えて、現在のマルチモーダル大規模言語モデルを利用してゼロショット方式で物理リアリズムを評価する、シンプルかつ効果的な方法も設計します。私たちは、5 つのオープンソース モデルと 5 つの独自モデルを含む 12 の最先端のテキストからビデオへの生成モデルを、詳細な比較と分析によって評価します。基本シナリオ、複合シナリオ、反物理シナリオにわたる 1,050 の厳選されたプロンプトにわたる体系的なテストを通じて、これらのモデルが現実世界の物理学に準拠する際に直面する極めて重要な課題を特定します。さらに、さまざまな物理現象やプロンプトの種類の下でのパフォーマンスを調査し、物理原理への忠実性を高めるプロンプトを作成するための的を絞った推奨事項を導き出します。

原文 (English)

"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents PhyWorldBench, a comprehensive benchmark designed to evaluate video generation models based on their adherence to the laws of physics. The benchmark covers multiple levels of physical phenomena, ranging from fundamental principles such as object motion and energy conservation to more complex scenarios involving rigid body interactions and human or animal motion. Additionally, we introduce a novel Anti-Physics category, where prompts intentionally violate real-world physics, enabling the assessment of whether models can follow such instructions while maintaining logical consistency. Besides large-scale human evaluation, we also design a simple yet effective method that utilizes current multimodal large language models to evaluate physics realism in a zero-shot fashion. We evaluate 12 state-of-the-art text-to-video generation models, including five open-source and five proprietary models, with detailed comparison and analysis. Through systematic testing across 1050 curated prompts spanning fundamental, composite, and anti-physics scenarios, we identify pivotal challenges these models face in adhering to real-world physics. We further examine their performance under diverse physical phenomena and prompt types, and derive targeted recommendations for crafting prompts that enhance fidelity to physical principles.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

LLM が自らベンチマークを行う場合: 自動評価における自己バイアスの解体

LLM が既存のベンチマークを急速に飽和させる中、LLM を使用した自動ベンチマーク作成 (LLM-as-a-benchmark)、つまりモデルがテスト入力を生成し (LLM-as-a-testset)、出力を評価する (LLM-as-an-evaluator) が、人間によるキュレーションに代わる安価な代替手段として注目を集めています。このパラダイムには根本的な問題があることを示します。LLM によって生成されたベンチマークは、それを作成したモデルを体系的に優先します。機械翻訳を主要なテストベッドとして使用すると、テストセットとしての LLM と評価者としての LLM という 2 つの追加ソースから自己バイアスが生じ、それらの組み合わせによって影響が増幅されることがわかりました。重要なのは、テスト データが明示的な多様性制御を使用して生成された場合でも、各モデルの暗黙的なスタイル傾向により、独自のスコアを増大させる均質なモデル固有の出力が生成されることです。私たちが提案する多様性指標を使用してソーステキストの多様性を高めると、このバイアスが部分的に軽減されます。自己バイアスが強いため、各モデル自体が最初にランク付けされ、ピアコンセンサスの順序付けが無効になります。この現象がチャットボット アリーナ タスクのオープンエンド生成にも及ぶことを確認しています。

原文 (English)

When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation

As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inputs (LLM-as-a-testset) and evaluates outputs (LLM-as-an-evaluator) -- has gained traction as a cheap alternative to human curation. We show that this paradigm has a fundamental problem: LLM-generated benchmarks systematically favor the model that created them. Using machine translation as our primary testbed, we find that self-bias arises from two additive sources, LLM-as-a-testset and LLM-as-an-evaluator, and their combination amplifies the effect. Crucially, even when test data is generated with explicit diversity controls, each model's implicit stylistic tendencies produce homogeneous, model-specific outputs that inflate its own scores. Increasing source text diversity, using our proposed diversity metric, partially mitigates this bias. Self-bias is strong enough to cause each model to rank itself first, overriding the peer-consensus ordering. We confirm that the phenomenon extends to open-ended generation on the Chatbot Arena task.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

管理された保持情報によるデコーダ専用 LLM アトリビューションの忠実性評価

大規模言語モデル (LLM) は、入力帰属手法を使用して評価されることが増えていますが、そのような説明を比較することは依然として困難です。 Soft-NC や Soft-NS などの既存のソフト摂動忠実度メトリクスは、アトリビューションの品質と摂動中に保持される単語の数を混同する可能性があります。平均スコアが大きいアトリビューション手法は、より多くの単語を保持する可能性があるため、水増しスコアが得られる可能性があります。この問題に対処するために、同じ期待保持確率の下でアトリビューション方法を比較し、保持単語数を制御する評価フレームワークである $\pi$-Soft-NC および $\pi$-Soft-NS を提案します。さらに、自己回帰デコーダ専用 LLM に合わせた勾配ベースのアトリビューション手法である Grad-ELLM を導入します。これは、各デコード ステップで、勾配から導出されるチャネルの重要性とアテンションから導出されるトークンの重要性を組み合わせます。 Llama と Mistral を使用した分類とオープン生成タスクの実験では、$\pi$-Soft-NC の下では Grad-ELLM が強力な包括性指向の忠実性を達成する一方、$\pi$-Soft-NS の下では有力な手法が存在しないことが示されています。私たちの評価基準は、LLM の XAI 手法を比較するための厳密なフレームワークとして機能し、この分野の進歩をサポートします。

原文 (English)

Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information

Large Language Models (LLMs) are increasingly evaluated with input attribution methods, yet comparing such explanations remains challenging. Existing soft-perturbation faithfulness metrics, such as Soft-NC and Soft-NS, can conflate attribution quality with the number of words retained during perturbation: attribution methods with larger average scores may keep more words and therefore obtain inflated scores. To address this issue, we propose $\pi$-Soft-NC and $\pi$-Soft-NS, an evaluation framework that compares attribution methods under the same expected retaining probability, thus controlling the number of retained words. We further introduce Grad-ELLM, a gradient-based attribution method tailored to autoregressive decoder-only LLMs, which combines gradient-derived channel importance with attention-derived token importance at each decoding step. Experiments on classification and open-generation tasks with Llama and Mistral show that Grad-ELLM achieves strong comprehensiveness-oriented faithfulness under $\pi$-Soft-NC, while there is no dominant method under $\pi$-Soft-NS. Our evaluation metric serves as a rigorous framework to compare XAI methods for LLMs, which will support progress in the field.

2026-05-27 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

TABX: マルチエージェント強化学習のための高スループットのサンドボックス バトル シミュレーター

環境の設計は、協調的なマルチエージェント強化学習 (MARL) アルゴリズムの開発と評価を形作る上で重要な役割を果たします。既存のベンチマークは重大な課題を浮き彫りにしていますが、カスタム評価シナリオの設計に必要なモジュール性が欠けていることがよくあります。再構成可能なマルチエージェント タスク用に設計された高スループットのサンドボックスである Totally Accelerated Battle Simulator in JAX (TABX) を紹介します。 TABX は、環境パラメータに対するきめ細かい制御を提供し、さまざまなタスクの複雑さにわたる緊急エージェントの動作とアルゴリズムのトレードオフを系統的に調査できるようにします。 TABX は、GPU 上でハードウェア アクセラレーションによる実行に JAX を活用することで、大規模な並列化を可能にし、計算オーバーヘッドを大幅に削減します。 TABX は、高速かつ拡張可能で簡単にカスタマイズできるフレームワークを提供することで、複雑な構造ドメインにおける MARL エージェントの研究を容易にし、将来の研究のための拡張可能な基盤として機能します。コードは https://github.com/ku-dmlab/TABX から入手できます。

原文 (English)

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning

The design of environments plays a critical role in shaping the development and evaluation of cooperative multi-agent reinforcement learning (MARL) algorithms. While existing benchmarks highlight critical challenges, they often lack the modularity required to design custom evaluation scenarios. We introduce the Totally Accelerated Battle Simulator in JAX (TABX), a high-throughput sandbox designed for reconfigurable multi-agent tasks. TABX provides granular control over environmental parameters, permitting a systematic investigation into emergent agent behaviors and algorithmic trade-offs across a diverse spectrum of task complexities. Leveraging JAX for hardware-accelerated execution on GPUs, TABX enables massive parallelization and significantly reduces computational overhead. By providing a fast, extensible, and easily customized framework, TABX facilitates the study of MARL agents in complex structured domains and serves as a scalable foundation for future research. Our code is available at: https://github.com/ku-dmlab/TABX.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達

GICDM: 信頼性の高い距離ベースの生成モデル評価のためのハブネスの軽減

生成モデルの評価は通常、高次元の埋め込み空間に依存してサンプル間の距離を計算します。これらの空間のデータセット表現は、最近隣関係を歪め、距離ベースのメトリクスを偏らせるハブネス現象の影響を受けることを示します。古典的な反復コンテキスト相違測定 (ICDM) に基づいて、実際のデータと生成されたデータの両方の近傍推定を修正する方法である生成 ICDM (GICDM) を導入します。経験的な動作を改善するためにマルチスケール拡張を導入します。合成ベンチマークと実際のベンチマークに関する広範な実験により、GICDM がハブネスに起因する障害を解決し、信頼性の高いメトリック動作を復元し、人間の評価との整合性が向上することが実証されています。

原文 (English)

GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation

Generative model evaluation commonly relies on high-dimensional embedding spaces to compute distances between samples. We show that dataset representations in these spaces are affected by the hubness phenomenon, which distorts nearest-neighbor relationships and biases distance-based metrics. Building on the classical Iterative Contextual Dissimilarity Measure (ICDM), we introduce Generative ICDM (GICDM), a method to correct neighborhood estimation for both real and generated data. We introduce a multi-scale extension to improve empirical behavior. Extensive experiments on synthetic and real benchmarks demonstrate that GICDM resolves hubness-induced failures, restores reliable metric behavior, and improves alignment with human assessment.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

オーマニック: 大規模言語モデルにおけるマルチホップ推論の段階的評価に向けて

最終的な回答のみから大規模言語モデル (LLM) の推論能力を評価すると、特にステップレベルのアノテーションがないマルチホップ QA ベンチマークでは、中間ステップでの失敗がわかりにくくなる可能性があります。このギャップに対処するために、最終回答の精度を測定するだけでなく、推論がどこで破綻しているかを診断するように設計されたオープンドメインの 4 ホップ QA ベンチマークである Omanic を導入します。 Omanic には、10,296 個の機械生成トレーニング サンプル (OmanicSynth) と、専門家がレビューした人間による注釈付きの 967 個の評価サンプル (OmanicBench) が含まれており、各評価質問はシングルホップのサブ質問、中間回答、および構造化グラフ トポロジに分解されています。独自のオープンソース LLM を使用した実験では、Omanic が困難であることが示されていますが、段階的な分析では、後のホップのボトルネック、事実の知識フロア、推論チェーンに沿ったエラーの伝播が明らかになりました。 OmanicSynth の微調整により、6 つの推論および数学ベンチマークに移行し、平均 7.41 ポイントの向上が得られ、推論能力移転の監視としての有効性が検証されました。データは https://huggingface.co/datasets/li-lab/Omanic で、コードは https://github.com/XiaojieGu/Omanic でリリースされます。

原文 (English)

Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks without step-level annotations. To address this gap, we introduce Omanic, an open-domain 4-hop QA benchmark designed not only to measure final-answer accuracy but also to diagnose where reasoning breaks down. Omanic contains 10,296 machine-generated training examples (OmanicSynth) and 967 expert-reviewed human-annotated evaluation examples (OmanicBench), with each evaluation question decomposed into single-hop sub-questions, intermediate answers, and structured graph topologies. Experiments with proprietary and open-source LLMs show that Omanic is challenging, while step-wise analysis reveals a later-hop bottleneck, factual knowledge floor, and error propagation along reasoning chains. Fine-tuning on OmanicSynth transfers to six reasoning and mathematics benchmarks, yielding a 7.41-point average gain and validating its effectiveness as supervision for reasoning-capability transfer. We release the data at https://huggingface.co/datasets/li-lab/Omanic and the code at https://github.com/XiaojieGu/Omanic.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

VLM が生徒を「修正」する場合: 複数行の手書き数学 OCR の評価における過剰修正を特定し、ペナルティを与える

手書きの数学を正確に転写することは、教育用 AI システムにとって非常に重要ですが、現在のベンチマークはこの機能を適切に評価できません。これまでの研究のほとんどは単一行の式に焦点を当てており、BLEU などの語彙メトリクスに依存しているため、複数行の生徒の解決策全体にわたる意味論的推論を評価できません。この論文では、複数行の手書き数学光学式文字認識 (OCR) に関する最初の体系的な研究を紹介し、視覚言語モデル (VLM) の重大な障害モードである過剰修正を明らかにしました。これらのモデルは、学生の作業を忠実に転写する代わりに、多くの場合、エラーを「修正」し、それによって教育評価が検出しようとしているまさに間違いを隠してしまいます。これに対処するために、ルーブリックベースのグレーディングに大規模言語モデル (LLM) を活用し、過剰修正に明示的にペナルティを課す意味論的評価指標である PINK (ペナルティ付き INK ベース スコア) を提案します。 FERMAT データセット上の 15 の最先端の VLM を総合的に評価したところ、BLEU と比較して大幅な順位の逆転が明らかになりました。GPT-4o のようなモデルは積極的な過剰補正に対して重くペナルティを受ける一方、Gemini 2.5 Flash が最も忠実なトランスクライバーとして浮上しています。さらに、人間の専門家による研究では、PINK が人間の判断と大幅に一致し (BLEU の 39.5% に対して 55.0% の優先度)、教育現場における手書き数学 OCR のより信頼性の高い評価フレームワークを提供することが示されています。

原文 (English)

When VLMs 'Fix' Students: Identifying and Penalizing Over-Correction in the Evaluation of Multi-line Handwritten Math OCR

Accurate transcription of handwritten mathematics is crucial for educational AI systems, yet current benchmarks fail to evaluate this capability properly. Most prior studies focus on single-line expressions and rely on lexical metrics such as BLEU, which fail to assess the semantic reasoning across multi-line student solutions. In this paper, we present the first systematic study of multi-line handwritten math Optical Character Recognition (OCR), revealing a critical failure mode of Vision-Language Models (VLMs): over-correction. Instead of faithfully transcribing a student's work, these models often "fix" errors, thereby hiding the very mistakes an educational assessment aims to detect. To address this, we propose PINK (Penalized INK-based score), a semantic evaluation metric that leverages a Large Language Model (LLM) for rubric-based grading and explicitly penalizes over-correction. Our comprehensive evaluation of 15 state-of-the-art VLMs on the FERMAT dataset reveals substantial ranking reversals compared to BLEU: models like GPT-4o are heavily penalized for aggressive over-correction, whereas Gemini 2.5 Flash emerges as the most faithful transcriber. Furthermore, human expert studies show that PINK aligns significantly better with human judgment (55.0% preference over BLEU's 39.5%), providing a more reliable evaluation framework for handwritten math OCR in educational settings.

2026-05-27 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

特許埋め込みのベンチマーク: 検索、分類、クラスタリングにわたる 22 モデルのマルチタスク評価

どの微調整シグナルが特許埋め込みモデルを改善しますか?また、利益は特許環境全体に移行しますか? 22M パラメータのエンコーダから 12B の命令調整 LLM まで、検索、分類、クラスタリングに関して 22 の埋め込みモデルのベンチマークを行います。この研究では、113,148 件の WIPO 支援技術特許、46,069 件の引用グラフ検索クエリ、および外部検証用の公開 DAPFAM データセットを使用しています。当社のフレームワークは、引用ベースの検索、ハイブリッド疎密融合、5 つのデータセットにわたるマルチラベル分類、教師なしクラスタリング、6 つのテキスト セクション ビュー、4 つのモデルのドメイン適応微調整、管轄分析、および独自の DWPI (Derwent World Patents Index、Clarivate) の専門家が執筆したコンテンツをカバーしています。結果は、微調整がタスクに依存していることを示しています。単一ランドスケープの調整はドメイン内のスコアを改善できますが、外部ランドスケープでの取得に悪影響を与えることが多く、より多くのドメイン データが常に役立つという仮定に疑問を呈します。モデル ファミリ内では、通常、スケールによってパフォーマンスが予測されます (Qwen3 0.6B から 4B から 8B、Llama-Nemotron 1B から 8B)。ただし、ファミリ間のスケーリングにはノイズが多く、12B KaLM-Gemma3 は TAC 検索で 8 位にランクされますが、Qwen3-0.6B は ARI クラスタリングで首位に立っています。 Title+Abstract+Claims は最も信頼性の高いテキスト表現です。マルチビューの抽象クレームの調整により、検索が nDCG@10 で最大 7.1 パーセント向上し、微調整の組み合わせにより最も強力な分類ゲイン (+7.1 F1) が得られます。すべてのモデルはドメイン外クエリで 55 ~ 65% 低下しますが、ハイブリッドの疎-密融合ではこのギャップは埋められません。 BM25 密補間では、適度な nDCG@10 ゲイン (+0.002 ~ +0.015) が得られますが、より弱いゼロショット密モデルでは大きな利点が得られます。コードと評価フレームワークは公開されています。

原文 (English)

Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

Two questions regarding practitioners' use of patent embeddings arise: (i) Does one fine-tuning recipe suffice for all downstream applications? (ii) Is fine-tuning on one patent landscape sufficient for downstream application on other landscapes? By evaluating 22 pre-trained embedding models (ranging from 22M to 12B parameters) on three tasks -- information retrieval, classification, and clustering -- on 113,148 WIPO patents for assistive technology (46,069 citation queries) and on an external DAPFAM dataset, we find that two results cast doubt on the prevailing wisdom. (i) The optimal fine-tuning recipe depends on the downstream task: cross-sectional alignment (recipe R3) provides the largest improvements to retrieval performance (+7.1% nDCG@10), whereas a combined signal recipe (recipe R4) is better suited to classification (+7.1 F1) and clustering (+10.9 V-measure); a matched data control confirms that differences in training dataset size are not a contributing factor. (ii) Single-landscape fine-tuning hampers cross-landscape information retrieval: fine-tuning on one landscape significantly degrades cross-domain retrieval for 5 of 8 model-recipe combinations on the DAPFAM corpus, with the stronger zero-shot models suffering most. While within-family scaling is consistent (Qwen3 0.6B->4B->8B; Llama-Nemotron 1B->8B), cross-family scaling is erratic; the 12B KaLM-Gemma3 is ranked 8th on TAC retrieval performance, following prefix modification. Title+Abstract+Claims is the ubiquitous best text view, and all models suffer from a 55-65% gap between IN and OUT-of-domain performance which cannot be mitigated by hybrid BM25-dense fusion. Code and evaluation framework are publicly available.

2026-05-27 03:33 JSTTechCrunch AIビジネス/資金調達

OpenRouter more than doubles valuation to $1.3B in a year

OpenRouter has raised a $113 million Series B led by CapitalG. Its 5x growth in usage over six months indicates the multi-AI-model future i…

2026-05-26 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

マシンサイコメトリクス: 人工知能の数理心理学

人工エージェントは現在、信頼、驚き、懸念を引き起こすのに十分な豊かな行動を生成していますが、私たちの評価ツールは依然として心理構造よりも能力スコアを優先しています。この論文は、2つの対称的な誤り(非生物学的システムにおける心理的組織を無視する人工心の盲目と、流暢な行動だけから人間のような内面生活を推測する人工心の投影)の間の哲学的行き詰まりは、意識の問題を解決するのではなく、その下に規律ある測定層を導入することによって回避できると主張する。この論文は、基質を超えた目標指向の能力としての認知についてのマイケル・レビンの連続的な見方と、数理心理学の方法論的レパートリー(項目反応理論、信号検出理論、ベイジアン認知モデリング、校正分析、認知バイアス電池)を利用して、人工エージェントの潜在的な行動、メタ認知、コミュニケーション、および自己モデリングの気質の測定科学としてマシンサイコメトリクスを開発しています。その運用の中核はマシン マインドプリントです。これは、キャリブレーション、ソースの完全性、暗示性の耐性、コンテキストの安定性、表現力の調整、ツールの完全性、ドリフト モニタリング、および分散グラウンディングに及ぶ、多次元でドメイン限定のバージョン管理されたプロファイルです。補完的なトラスト プロトコルは、プローブ バッテリー、摂動テスト、信頼性と妥当性の分析、および一か八かのドメインにわたる長期的な監視を通じて、マインドプリントを展開の決定に変えます。哲学的貢献は、意識を擬人化したり無視したりせず、意識を前提としたり排除したりしない、第 3 の立場である「人工精神の規律」です。目的は、人工エージェントを人間化することではなく、人間ではないからこそ、判断する前に測定することで人工エージェントを理解することです。

原文 (English)

Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence

Artificial agents now generate behavior rich enough to invite trust, surprise, and concern, yet our evaluation tools still privilege capability scores over psychological structure. This paper argues that the philosophical impasse between two symmetrical errors (Artificial Mind Blindness, which dismisses psychological organization in non-biological systems, and Artificial Mind Projection, which infers human-like inner life from fluent behavior alone) can be circumvented not by resolving the consciousness question, but by introducing a disciplined measurement layer beneath it. Drawing on Michael Levin's continuum view of cognition as goal-directed competency across substrates, and on the methodological repertoire of mathematical psychology (Item Response Theory, Signal Detection Theory, Bayesian cognitive modeling, calibration analysis, cognitive-bias batteries), the paper develops Machine Psychometrics as a measurement science of latent behavioral, metacognitive, communicative, and self-modeling dispositions in artificial agents. Its operational core is the Machine Mindprint: a multidimensional, domain-bounded, versioned profile spanning calibration, source integrity, suggestibility resistance, context stability, expressive alignment, tool integrity, drift monitoring, and distributional grounding. A complementary Trust Protocol turns Mindprints into deployment decisions through probe batteries, perturbation testing, reliability and validity analysis, and longitudinal monitoring across high-stakes domains. The philosophical contribution is a third stance, Artificial Mind Discipline, that neither anthropomorphizes nor dismisses, neither presupposes consciousness nor forecloses it. The aim is not to humanize artificial agents, but to understand them precisely because they are not human, through measurement before judgment.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

MAPLE: 不完全情報ゲームにおける AlphaZero のマルチステート集約ポリシー評価

不完全情報ゲーム (IIG) は、プレーヤーが実際のゲームの状態を完全に観察せずに決定を下さなければならないため、挑戦的です。 AlphaZero は完全情報ゲームで目覚ましい成功を収めていますが、それを IIG に拡​​張することは依然として困難です。完全情報モンテカルロ (PIMC) などの既存の検索ベースのアプローチは戦略の融合に問題があり、一方、情報セット モンテカルロ ツリー検索 (IS-MCTS) はニューラル ネットワークと組み合わせると高い計算コストが発生します。この論文では、制御可能な計算コストを維持しながら、PIMC と IS-MCTS の利点を組み合わせて、単一の検索ツリー内でサンプルされた複数の世界の状態から政策と価値の評価を集約するツリー検索手法である Multi-State Aggregated PoLicy Evaluation (MAPLE) を提案します。さらに、情報セットから有益な世界状態を選択するために、シャムベースのサンプリング戦略を組み込みます。 Phantom Go と Dark Hex の実験では、MAPLE が PIMC ベースの AlphaZero ベースラインを大幅に上回り、それぞれ 291 と 136 の Elo 改善を達成したことが示されています。これらの結果は、MAPLE が不完全情報ゲームにおける AlphaZero スタイルの学習に効果的なアプローチであることを示しています。

原文 (English)

MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games

Imperfect-information games (IIGs) are challenging, as players must make decisions without fully observing the true game state. While AlphaZero has achieved remarkable success in perfect-information games, extending it to IIGs remains difficult. Existing search-based approaches, such as Perfect Information Monte Carlo (PIMC), suffer from strategy fusion, while Information Set Monte Carlo Tree Search (IS-MCTS) incurs high computational cost when combined with neural networks. In this paper, we propose Multi-State Aggregated PoLicy Evaluation (MAPLE), a tree search method that aggregates policy and value evaluations from multiple sampled world states within a single search tree, combining the advantages of PIMC and IS-MCTS while maintaining a controllable computational cost. We further incorporate a Siamese-based sampling strategy to select informative world states from the information set. Experiments on Phantom Go and Dark Hex show that MAPLE significantly outperforms the PIMC-based AlphaZero baseline, achieving Elo improvements of 291 and 136, respectively. These results demonstrate that MAPLE is an effective approach for AlphaZero-style learning in imperfect-information games.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

AVBench: オーディオビデオ生成モデルのための、人間に合わせた自動評価ベンチマーク

オーディオ ビデオ (AV) 生成の急速な進歩により、特に音声や対話を含む人間関連のシナリオにおいて、同期されたサウンドによる高忠実度の合成が可能になりました。しかし、AV 生成の評価は依然として初期段階にあり、人間関連のシナリオについては粗粒度のベンチマークがいくつかしかなく、汎用マルチモーダル LLM を使用した限られたプリセット評価に依存しているため、モデルの機能の不正確な評価につながっています。これらの問題に対処するために、人間中心の AV 生成に合わせて調整された完全に自動化されたベンチマークである AVBench を導入します。 AVBench は、包括的かつ正確な評価を実現するための 2 つの主要な設計に基づいて構築されています。(i) 人間中心の詳細な指標。 AVBench は、人間中心の現実世界のシナリオ向けに設計された 10 の評価次元を統合し、モダリティ全体のビジュアル品質、オーディオ品質、マルチレベルの一貫性をカバーします。これらの実用的な指標は、既存のベンチマークでは見落とされがちな人間関連の詳細を捕捉します。 (ii) 選好学習による専門の評価者。特殊なトレーニング データの不足に対処するために、実世界のビデオを制御された摂動を備えた多様なトレーニング ペアに変換することで、大規模な監視を構築します。この高品質のデータセットを微調整した後、評価者は、微妙なクロスモーダルの不一致を確実に検出する方法を学習します。重要なのは、AVBench は個別のテキスト判断を生成するのではなく、バイナリ決定に対するモデルの予測信頼度から連続的な評価スコアを導出するということです。この確率的スコアリング メカニズムにより、従来の VQA スタイルの評価よりも信頼性の高い評価が可能になり、人間の判断と密接に一致します。まとめると、AVBench は AV 生成の自動評価を提供し、データ フィルタリングの強力な可能性を実証し、ヒューマン フィードバックからの強化学習 (RLHF) の微分可能な報酬信号として機能します。

原文 (English)

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions. Yet evaluation for AV generation remains at an early stage, with only a few coarse-grained benchmarks for human-related scenarios and relying on limited preset evaluations with generic multimodal LLMs, leading to inaccurate assessments of model capabilities. To address these issues, we introduce AVBench, a fully automated benchmark tailored for human-centric AV generation. AVBench is built on two key designs for comprehensive and accurate evaluation: (i) Human-centric and fine-grained metrics. AVBench integrates ten evaluation dimensions designed for human-centered real-world scenarios, covering visual quality, audio quality, and multi-level consistency across modalities. These practical metrics capture human-related details that existing benchmarks often overlook. (ii) Specialized evaluators via preference learning. To address the lack of specialized training data, we construct large-scale supervision by transforming real-world videos into diverse training pairs with controlled perturbations. After fine-tuning on this high-quality dataset, the evaluators learn to reliably detect subtle cross-modal inconsistencies. Crucially, instead of producing discrete textual judgment, AVBench derives continuous evaluation scores from the model's prediction confidence on binary decisions. This probabilistic scoring mechanism enables a more reliable assessment than traditional VQA-style evaluation and aligns closely with human judgment. Taken together, AVBench offers automated evaluation for AV generation, demonstrates strong potential for data filtering, and serves as a differentiable reward signal for Reinforcement Learning from Human Feedback (RLHF).

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM における推論の質の測定: 多次元の行動フレームワーク

LLM は複雑な推論タスクで目覚ましい成功を収めていますが、現在の評価アプローチは主に最終的な答えの正しさに依存しており、それらの答えを生み出す根本的な推論プロセスについての洞察は限られています。このギャップに対処するために、この研究では、動作の観点から LLM の推論品質を測定するための統一された多次元フレームワークを提案し、理論的に根拠のある 6 つの次元、正確性 (CQ)、一貫性 (CS)、堅牢性 (RS)、論理的一貫性 (LS)、効率 (ES)、安定性 (SS) を運用します。 4 つのベンチマークの 975 項目にわたる 7 つの LLM に関する広範な実験により、このフレームワークが精度のみの指標では見えない動作を明らかにすることが実証されました。特に、論理的一貫性は正しさ (r = -0.172、ns) と直交しており、一貫性のない推論から正しい答えが得られることが確認され、一方、Claude-Haiku-4.5 は最高の多次元スコア (Q_bal = 0.778) を達成しています。さらに、このフレームワークは重大なランキングの逆転を明らかにしています。DeepSeek-V3 は精度優先では 2 位ですが、法的/コンプライアンスの重み付けでは 5 位にランクされており、単一指標の評価では検出できない逆転です。判別式の妥当性により、11/15 次元のペアが独立している (|r| < 0.50) ことが確認され、各次元を別個の信号として扱うための心理測定的サポートが提供されます。フレームワークによって生成される次元プロファイルは、次の 3 つのクラスの展開決定を直接サポートします。最終的な答えが正しいにもかかわらず、その推論トレースが説明責任監査に失敗するモデルを特定します (LS--CQ 直交性)。精度のみのベンチマークによって引き起こされるランキングエラーを防止します。そして、フレームワークがキャプチャする 6 つの独立したシグナルを単一のメトリックが暗黙的に置き換えることがないようにします。

原文 (English)

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

LLMs have achieved remarkable success in complex reasoning tasks, yet current evaluation approaches predominantly rely on final-answer correctness, offering limited insight into the underlying reasoning processes that produce those answers. To address this gap, this study proposes a unified multi-dimensional framework for measuring reasoning quality in LLMs from a behavioral perspective, operationalizing six theoretically grounded dimensions: Correctness (CQ), Consistency (CS), Robustness (RS), Logical Coherence (LS), Efficiency (ES), and Stability (SS). Extensive experiments on seven LLMs across 975 items from four benchmarks demonstrate that the framework reveals behaviors invisible to accuracy-only metrics. Notably, logical coherence is orthogonal to correctness (r = -0.172, ns), confirming that correct answers can arise from incoherent reasoning, while Claude-Haiku-4.5 achieves the highest multi-dimensional score (Q_bal = 0.778). Furthermore, the framework exposes critical ranking inversions: DeepSeek-V3 ranks second under accuracy-priority but fifth under legal/compliance weighting, a reversal that single-metric evaluation cannot detect. Discriminant validity confirms 11/15 dimension pairs are independent (|r| < 0.50), providing psychometric support for treating each dimension as a distinct signal. The dimensional profiles produced by the framework directly support three classes of deployment decision: identifying models whose reasoning traces would fail accountability audits despite correct final answers (LS--CQ orthogonality); preventing ranking errors caused by accuracy-only benchmarking; and ensuring that no single metric silently substitutes for the six independent signals the framework captures.

2026-05-26 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

薬剤の不確実性定量化のための適切なスコアリングルール

言語モデル エージェントは軌跡全体にわたって不確実性シグナルを発することが増えていますが、既存のエージェントの UQ 評価では、ランク付けの有用性と確率的真実性が混同されることがよくあります。 AUROC、AUPRC、リスクカバレッジ、Trajectory ECE、およびスカラー化された軌跡スコアは、識別、ビンごとのキャリブレーション、または折りたたまれた要約を評価しますが、プレフィックス条件付きの完全な成功確率トレース $q_t = P^{\pi}(Y=1 | H_t)$ を厳密に導き出すわけではありません。事前の適切なスコアリングに基づいて、最終的な成功の確率に調整されたステップごとの不確実性信号に対する厳密に適切な軌道レベルのスコアリング ルールの予測子に依存しないファミリーである軌道適切スコア (TPS) を導入します。我々は、選択されたスコアファミリーと加重スケジュール内で、完全な観察の下でTPSが成功確率プロセスを厳密に導き出すことを証明します。完全データスコアを観測可能な停止プレフィックスに投影することにより、この構築を管理者によって検閲された軌道に拡張し、$q_Z$ が推定されていない場合の正確な $q_Z$ 加重削減スコアと扱いやすい近似値を生成します。さらに、一般的な軌道評価器は、完全なプレフィックス条件付き確率プロセスよりも弱いオブジェクトをターゲットにすることを示します。軌道 ECE は解像度ブラインドですが、スカラー化された軌道ブリエは、完全なトレースではなく、崩壊したスカラーのみを導き出します。 StrategyQA、Tau2-Bench、HotpotQA、および WebShop での実験では、これらの理論的な違いが運用上目に見えることを示しています。つまり、確率の再調整により、ランク メトリクスをほとんど変更せずに TPS が大幅に変更される可能性があり、扱いやすい打ち切り近似により、完全のみの評価と比較して判定が変更される可能性があります。

原文 (English)

Proper Scoring Rules for Agentic Uncertainty Quantification

Language-model agents increasingly emit uncertainty signals throughout a trajectory, but existing agentic UQ evaluations often conflate ranking usefulness with probabilistic truthfulness. AUROC, AUPRC, risk-coverage, Trajectory ECE, and scalarized trajectory scores evaluate discrimination, binwise calibration, or collapsed summaries, but do not strictly elicit the full prefix-conditioned success-probability trace $q_t = P^{\pi}(Y=1 | H_t)$. Building on prequential proper scoring, we introduce the Trajectory Proper Score (TPS), a predictor-agnostic family of strictly proper trajectory-level scoring rules for any per-step uncertainty signal calibrated into a probability of eventual success. We prove that TPS strictly elicits the success-probability process under complete observation, within the chosen score family and weight schedule. We extend the construction to administratively censored trajectories by projecting the complete-data score onto the observable stopped prefix, yielding an exact $q_Z$-weighted reduced score and a tractable approximation when $q_Z$ is unestimated. We further show that common trajectory evaluators target weaker objects than the full prefix-conditioned probability process: Trajectory ECE is resolution-blind, while scalarized Trajectory Brier elicits only the collapsed scalar, not the full trace. Experiments on StrategyQA, Tau2-Bench, HotpotQA, and WebShop show that these theoretical distinctions are operationally visible: probability recalibration can substantially change TPS while leaving rank metrics nearly unchanged, and the tractable censored approximation can change the verdict relative to complete-only evaluation.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

PRIMA: 検証可能なアイデンティティと集中的なフィードバックを備えた、回復力のあるマルチエージェント研究のための運用パターン

LLM を複数時間の実行にわたって調整されたマルチエージェント調査システムとして運用すると、単発評価では不可能な障害モードが表面化します。つまり、上流のプロバイダーが警告なしにスロットルする、サブエージェントがアクセス可能なツールに合わせてタスクをドリフトする、機械を使用する代わりにナレーションする、自己謝罪を伴うオープンリビジョンの反復、または上流のコンテキストを実行可能なディレクティブとして扱うなどです。 PRIMA の主な貢献は、これらの障害モードを乗り切るための 3 つの動作パターンです。(1) アップストリームのレート制限信号を検出し、型指定された一時停止レコードをディスクに永続化し、プロセスの再起動後であっても統合された作業を再実行せずに長時間実行を再開する回復力および回復層。 (2) タスクの忠実度、ツールの使用、改訂、およびステップ間のコンテキスト境界の規範を構造的なプロンプト層としてエンコードするサブエージェント操作規律。 (3) 最終合成前の明示的なドキュメント間調和パスと直交するドラフト ステップを組み合わせた構造化エンジニアリング成果物の多段階アプリケーション パターン。これらは、明示的な収束基準を備えた研究プログラム仕様言語、デュアルメトリック スコアリング エンジン (LLM で判定されたルーブリックとサンドボックス コード)、外部メタ最適化ループ、イベント駆動型永続性、フックベースのミドルウェア、コンテキスト コンパクション、およびマルチプロバイダー LLM 抽象化といった基本的なプロトコルの上に位置します。エージェント ID は主要な権限から派生し、衝突のない識別子と中央レジストリなしで簡単に検証可能なクラスター メンバーシップを提供します。理論的な保証には、$O(k)$ 検証、$O(V+E)$ DAG 検証、および算術基本定理による恒等衝突の自由が含まれます。グラフ同型のケーススタディは、生成されたアーティファクトにおけるアーキテクチャ上の主張を根拠としています。つまり、3 つの定理と 5 つの予想を含む新しい標準形式のアルゴリズムを提案する研究論文を作成した 6 ステップのプロトコルです。

原文 (English)

PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback

Operating LLMs as coordinated multi-agent research systems over multi-hour runs surfaces failure modes that single-shot evaluation cannot: upstream providers throttle without warning, sub-agents drift the task to fit accessible tools, narrate machinery instead of using it, open revision iterations with self-apology, or treat upstream context as executable directives. We present PRIMA, whose primary contributions are three operational patterns for surviving these failure modes: (1) a resilience-and-recovery layer that detects upstream rate-limit signals, persists a typed pause record to disk, and resumes long-running runs without re-executing converged work even across process restarts; (2) a sub-agent operating discipline encoding task-fidelity, tool-use, revision, and inter-step context-boundary norms as a structural prompt layer; (3) a multi-phase application pattern for structured engineering deliverables pairing orthogonal draft steps with an explicit cross-document harmonization pass before final synthesis. These sit atop a foundational protocol: a research-program specification language with explicit convergence criteria, a dual-metric scoring engine (LLM-judged rubric plus sandboxed code), an outer meta-optimization loop, event-driven persistence, hook-based middleware, context compaction, and a multi-provider LLM abstraction. Agent identities derive from prime powers, giving collision-free identifiers and trivially-verifiable cluster membership without a central registry. Theoretical guarantees include $O(k)$ verification, $O(V+E)$ DAG validation, and identity collision freedom by the Fundamental Theorem of Arithmetic. A Graph Isomorphism case study grounds the architectural claims in a generated artifact: a six-step protocol that produced a research paper proposing a new canonical-form algorithm with three theorems and five conjectures.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

シールドの反転: ポリシー仕様から安全性テストを体系的に生成

大規模言語モデル (LLM) の広範な統合には、厳密かつ体系的な安全性評価が必要です。既存のパラダイムは、構築されたベンチマークに依存して事前定義された観点から安全性を評価するか、動的レッドチームを採用して潜在的な脆弱性を調査します。これらのアプローチは効果的ではありますが、専門分野の知識に大きく依存し、体系的な保証が限られており、急速な陳腐化に対して脆弱であるため、課題に直面しています。これらの制限に対処するために、AI の安全性に仕様ベースのソフトウェア テストの厳密さをもたらす新しいフレームワーク POLARIS を導入します。 POLARIS は、まず非構造化自然言語ポリシーを一次論理 (FOL) 表現にコンパイルし、高レベルのルールと具体的なテスト ケースの間に追跡可能なリンクを確立します。この形式化により、複雑なポリシー違反シナリオが通過可能なパスとしてエンコードされるセマンティック ポリシー グラフの構築が可能になります。 POLARIS は、このグラフを体系的に調査することで構成違反パターンを明らかにし、それを実行可能な自然言語テスト クエリにインスタンス化して、カバレッジ主導型の再現可能な安全性テストを可能にします。実験では、POLARIS が確立されたベースラインと比較して、より高いポリシー適用範囲と攻撃成功数を達成していることが実証されています。重要なのは、POLARIS が正式な手法と AI の安全性を橋渡しすることで、LLM が検証可能なトレーサビリティを備えた安全性が重要なポリシーに確実に従うようにするための原則に基づいた自動化されたアプローチを提供することです。コードは https://github.com/huac-lxy/POLARIS でリリースされています。

原文 (English)

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on constructed benchmarks to assess safety from predefined perspectives, or employ dynamic red-teaming to probe potential vulnerabilities. While effective, these approaches face challenges, as they depend heavily on expert domain knowledge, offer limited systematic guarantees, and are vulnerable to rapid obsolescence. To address these limitations, we introduce a novel framework POLARIS that brings the rigor of specification-based software testing to AI safety. POLARIS first compiles unstructured natural-language policies into First-Order Logic (FOL) representations, establishing a traceable link between high-level rules and concrete test cases. This formalization enables the construction of a Semantic Policy Graph, where complex policy violation scenarios are encoded as traversable paths. By systematically exploring this graph, POLARIS uncovers compositional violation patterns, which are then instantiated into executable natural-language test queries, enabling coverage-driven and reproducible safety testing. Experiments demonstrate that POLARIS achieves higher policy coverage and attack success counts compared to established baselines. Crucially, by bridging formal methods and AI safety, POLARIS provides a principled, automated approach to ensuring LLMs adhere to safety-critical policies with verifiable traceability. We release our code at https://github.com/huac-lxy/POLARIS.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

LipoAgent: より安全な脂質設計のための微調整された LLM エージェントの調整

脂質ナノ粒子 (LNP) は、臨床的に最も成熟した核酸送達プラットフォームの 1 つですが、有効かつ生物学的に安全な脂質の設計が依然として大きなボトルネックとなっています。実際のスクリーニングでは、毒性は意思決定レベルの制約です。脂質が毒性がある場合、その効率予測は臨床的に無関係です。私たちは、脂質発見のための安全性を意識したマルチエージェント LLM フレームワークである LipoAgent を提案します。 LipoAgent は、ドメイン固有の微調整と、効率予測の前提条件として毒性を強制する条件付き予測目標を組み合わせ、不一致が続く場合には人による監視を軽減したマルチエージェント検証によって信頼性をさらに向上させます。複数の基礎モデルにわたって、LipoAgent は、報告されている他の脂質設計モデルと比較して、mRNA トランスフェクション効率予測において平均 32% の相対的な向上を達成しています。ウェットラボ検証により、仮想スクリーニングのランキングが生物学的トランスフェクションの結果に確実に反映されることが確認されています。コードは https://github.com/SAI-Lab-NYU/LipoAgent.git で公開されています。

原文 (English)

LipoAgent: Coordinating Fine-Tuned LLM Agents for Safer Lipid Design

Lipid nanoparticles (LNPs) are among the most clinically mature platforms for nucleic acid delivery, yet designing lipids that are both effective and biologically safe remains a major bottleneck. In practical screening, toxicity is a decision-level constraint: if a lipid is toxic, its efficiency prediction is clinically irrelevant. We propose LipoAgent, a safety-aware multi-agent LLM framework for lipid discovery. LipoAgent combines domain-specific finetuning with a conditional prediction objective that enforces toxicity as a prerequisite for efficiency prediction, and further improves reliability via multi-agent verification with lightweight human oversight when disagreement persists. Across multiple foundation models, LipoAgent achieves an average 32% relative improvement in mRNA transfection efficiency prediction compared with other reported models for lipid design. Wet-lab validation confirms that virtual screening rankings reliably translate to biological transfection outcomes. The code is publicly available at https://github.com/SAI-Lab-NYU/LipoAgent.git.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

CausaLab: AI 科学者向けのインタラクティブな因果発見のためのスケーラブルな環境

LLM エージェントによるインタラクティブな因果発見を評価するためのスケーラブルな環境である CausaLab を紹介します。以前の評価とは異なり、CausaLab では、エージェントが因果関係の証拠を使用して問題を解決できるかどうか、およびその答えが根底にある因果メカニズムに関する正しい仮説によって裏付けられているかどうかの両方を評価します。各エピソードではエージェントが合成実験室に配置されます。エージェントは以前の測定記録を受け取り、マニピュレーター結晶に介入し、同じ機構によって支配される保持されたリアクター結晶の共振周波数を予測します。隠されたデータ生成プロセスは、ランダムにサンプリングされた構造因果モデル (SCM) であるため、成功するには、事前の知識を思い出すのではなく、因果グラフと構造方程式の両方を回復する必要があります。 CausaLab には、エージェントの進化する SCM 仮説を記録するドメイン固有の言語も含まれており、軌跡を検査可能にしてグラウンド トゥルースと比較できるようになります。実験では、予測とメカニズム回復の間に永続的なギャップがあることが示されています。純粋に観測的な 6 ノード設定では、GPT-5.2-high はタスク精度 92% に達しますが、オールエッジ $F_1$ はわずか 0.471 です。この観察は、さまざまな相互作用戦略の探求をさらに動機づけます: 混合観察 - 介入戦略は構造忠実度を向上させます: 混合 6 ノード設定では、GPT-5.2-high はタスク精度とオールエッジ $F_1$ の両方で 80% を達成しました。しかし、純粋な介入戦略はタスクの精度とオールエッジ $F_1$ の両方においてパフォーマンスが低いため、強力なエージェントですら有益な介入を設計するのに苦労しています。私たちは、エージェントの主要な弱点として早期停止を特定し、仮説と過去のデータとの間の一貫性をモデルに検証するように依頼することが、この問題の軽減に役立つことを示します。したがって、CausaLab は予測の成功を因果関係の理解から切り離し、実験的因果推論者としての現在の LLM エージェントの限界を明らかにします。

原文 (English)

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates both whether an agent can solve a problem using causal evidence and whether its answer is supported by a correct hypothesis about the underlying causal mechanism. Each episode places an agent in a synthetic laboratory: it receives prior measurement records, intervenes on a manipulator crystal, and predicts the resonance frequency of a held-out reactor crystal governed by the same mechanism. The hidden data-generating process is a randomly sampled structural causal model (SCM), so success requires recovering both a causal graph and structural equations rather than recalling prior knowledge. CausaLab also includes a domain-specific language that records the agent's evolving SCM hypothesis, making trajectories inspectable and comparable with ground truth. Experiments show a persistent gap between prediction and mechanism recovery: in the purely observational 6-node setting, GPT-5.2-high reaches 92% task accuracy but only 0.471 all-edge $F_1$. This observation further motivates our exploration of different interaction strategies: Mixed observation--intervention strategies improve structural fidelity: in the mixed 6-node setting, GPT-5.2-high achieves 80% on both task accuracy and all-edge $F_1$. Yet even strong agents struggle to design informative interventions, as pure intervention strategies perform poorly on both task accuracy and all-edge $F_1$. We identify premature stopping as a major weakness of agents, and show that asking the model to verify the consistency between its hypothesis and past data can help mitigate this issue. CausaLab therefore separates predictive success from causal understanding and exposes current LLM agents' limits as experimental causal reasoners.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

AI 主導のアルファ減衰: アルゴリズムの均質化、反射的な信号侵食、インテリジェント市場のパラドックス

AI 主導の投資戦略は本質的に大規模化すると自滅するものであることを示します。 AI の導入が進むにつれて、信号の混雑、パフォーマンスによる信号の侵食、レッド クイーンの競争という 3 つの相互強化チャネルが超過収益を圧縮します。アルファ半減期 $h(\phi) = \ln 2/[\theta + \delta(\phi)]$ を導き出します。ここで、$\theta$ は自然平均回帰率、$\delta(\phi) = N\phi\rho a/\lambda(\phi)$ は AI によって加速された減衰成分であり、採用において凸状に減少しています。現在の普及レベル ($\phi \約 0.7$、$\rho \約 0.6$) では、このモデルは信号の半減期が 18 か月であるのに対し、AI 以前は 5 ~ 7 年であることを示唆しています。我々は 4 つの理論的結果を確立します。まず、アルファ半減期定理: AI の導入により信号の寿命は凸状に減少します。第 2 に、信号消滅カスケード: 臨界しきい値 $\phi^*$ を超えると、1 つの信号クラスの減衰が残りの信号に対する競争の加速を​​引き起こします。第三に、赤の女王の不可能性です。モノカルチャーの均衡では、AI への多額の投資にもかかわらず、純アルファは同様にゼロになります。第 4 に、脆弱性と効率のトレードオフです。価格発見を最大化する導入レベルは、システムの脆弱性を最小化するレベルを厳密に上回っています。実証的検証により、ポートフォリオの収束が SEC フォーム 13F 提出パターン (9,950 万株、2013 ~ 2024 年) に合わせて調整され、シミュレートされた機関投資家ポートフォリオの収束がサンプル期間にわたって 42% 増加することが実証されました。 AIを採用したファンド間の横断的な分散が減少していることを示すシミュレーションされたヘッジファンドのリターンダイナミクスを調査し、脆弱性の影響を説明するために2010年のフラッシュクラッシュをシミュレーションしました。

原文 (English)

AI-Driven Alpha Decay: Algorithmic Homogenization, Reflexive Signal Erosion, and the Paradox of Intelligent Markets

We show that AI-driven investment strategies are inherently self-defeating at scale. As AI adoption rises, three mutually reinforcing channels -- signal crowding, performative signal erosion, and Red Queen competition -- compress excess returns. We derive the alpha half-life $h(\phi) = \ln 2/[\theta + \delta(\phi)]$, where $\theta$ is the natural mean-reversion rate and $\delta(\phi) = N\phi\rho a/\lambda(\phi)$ is the AI-accelerated decay component, which is convex-decreasing in adoption. At current adoption levels ($\phi \approx 0.7$, $\rho \approx 0.6$), the model implies signal half-lives of 18 months versus 5-7 years pre-AI. We establish four theoretical results. First, the alpha half-life theorem: signal lifespans are convex-decreasing in AI adoption. Second, a signal extinction cascade: beyond a critical threshold $\phi^*$, the decay of one signal class triggers accelerated competition for remaining signals. Third, a Red Queen impossibility: in the monoculture equilibrium, net alpha is identically zero despite heavy AI investment. Fourth, a fragility-efficiency tradeoff: the adoption level maximizing price discovery strictly exceeds the level minimizing systemic fragility. Empirical validation calibrates portfolio convergence to SEC Form 13F filing patterns (99.5 million holdings, 2013-2024), documenting that simulated institutional portfolio convergence increases by 42% over the sample period. We examine simulated hedge fund return dynamics showing declining cross-sectional dispersion among AI-adopting funds, and simulate the 2010 Flash Crash to illustrate fragility consequences.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM-AutoSciLab: LLM を使用したアクティブな実験によるクローズドループの科学的発見

科学的発見は、仮説がデータ収集を導き、観察によって仮説空間が洗練される閉ループのプロセスです。しかし、ほとんどのアプローチは、発見を固定データセット上の教師あり学習に落とし込み、限定された観察が局所的に適合するが一般化できない複数のもっともらしいメカニズムをサポートできる可能性があります。したがって、重要な課題は、不確実性を解決するために有益な観察を選択し、静的推論から適応的なデータ取得に焦点を移すことです。これに対処するために、仮説生成と仮説条件付き実験の選択およびメカニズムの改良を組み合わせる閉ループ フレームワークである LLM-AutoSciLab を提案します。 LLM-AutoSciLab は、受動的に収集されたデータにモデルを適合させるのではなく、もっともらしい仮説を繰り返し提案し、それらを区別または改良するために有益な実験を選択し、結果として得られた証拠を使用して状態を更新します。アクティブなデータ取得による動的な閉ループ科学的発見を評価するために、2 つのデータセットで構成される ActiveSciBench を導入します。1 つは 57 の酵素動態タスクを含む ActiveSciBench-Chem、もう 1 つは 45 の遺伝子制御ネットワーク タスクを含む ActiveSciBench-GRN です。これらのデータセットは、適応的な実験計画、変数の選択、真のメカニズムの回復を必要とする、予算に制約のあるプロセスとして発見をモデル化します。 NewtonBench、ActiveSciBench-Chem、ActiveSciBench-GRN のいずれにおいても、LLM-AutoSciLab は従来の手法を上回り、NewtonBench と ActiveSciBench-Chem でそれぞれ 67.6% と 35.1% のシンボリック精度を達成し、ActiveSciBench-GRN で 31.1% の正確なグラフ回復を達成しました。さらに、仮説に基づいた実験は、競合する最も強力なベースラインよりもサンプル効率が 2 ~ 5 倍優れています。コードとデータは、https://github.com/scientific-discovery/LLM-AutoSciLab から入手できます。

原文 (English)

LLM-AutoSciLab: Closed-Loop Scientific Discovery via Active Experimentation with LLMs

Scientific discovery is a closed-loop process in which hypotheses guide data acquisition and observations refine the hypothesis space. Yet most approaches reduce discovery to supervised learning over fixed datasets, where limited observations can support multiple plausible mechanisms that fit locally but fail to generalize. Thus, the key challenge is selecting informative observations to resolve uncertainty, shifting the focus from static inference to adaptive data acquisition. To address this, we propose LLM-AutoSciLab, a closed-loop framework that couples hypothesis generation with hypothesis-conditioned experiment selection and mechanism refinement. Rather than fitting models to passively collected data, LLM-AutoSciLab iteratively proposes plausible hypotheses, selects informative experiments to distinguish or refine them, and updates its state using the resulting evidence. To evaluate dynamic, closed-loop scientific discovery with active data acquisition, we introduce ActiveSciBench, comprising two datasets: ActiveSciBench-Chem with 57 enzyme-kinetics tasks and ActiveSciBench-GRN with 45 gene-regulatory-network tasks. These datasets model discovery as a budget-constrained process requiring adaptive experiment design, variable selection, and recovery of true mechanisms. Across NewtonBench, ActiveSciBench-Chem, and ActiveSciBench-GRN, LLM-AutoSciLab outperforms prior methods, achieving 67.6% and 35.1% symbolic accuracy on NewtonBench and ActiveSciBench-Chem, respectively, and 31.1% exact graph recovery on ActiveSciBench-GRN. Moreover, hypothesis-guided experimentation is 2-5x more sample-efficient than the strongest competing baselines. Code and data are available at: https://github.com/scientific-discovery/LLM-AutoSciLab

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

TRACER: コード LLM におけるきめ細かい汚染検出のためのセマンティック認識フレームワーク

データ汚染は、モデル評価の信頼性に対する既知の脅威です。ただし、コード大規模言語モデル (LLM) では、汚染が正確な複製を超えてしまうことがよくあるため、依然として研究が進んでいません。私たちは、きめ細かいコード汚染検出のためのセマンティクスを意識したフレームワークである TRACER を紹介します。 TRACER は、機能的に同一、ほぼ同一、共有ロジックという 3 つのレベルのセマンティック重複を使用して汚染をモデル化し、粗いパイプラインから細かいパイプラインを通じてそれらを検出します。また、広く使用されている 3 つのベンチマークと 3 つの代表的なトレーニング後のデータセットにわたる、きめ細かいコード汚染検出のための最初のベンチマークも紹介します。 TRACER は複数の LLM バックボーンにわたって強力で一貫したパフォーマンスを実現し、GPT-5 はきめ細かい検出で F1 スコア 0.91 に達しました。バイナリ設定では、TRACER は F1 0.92 を達成し、既存の方法を 42% ~ 217% 上回ります。さらに、TRACER の個々のコンポーネントの寄与を評価するために、アブレーション研究とエラー分析を実施します。

原文 (English)

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs

Data contamination is a known threat to the reliability of model evaluation. However, it remains underexplored in code large language models (LLMs), where contamination often goes beyond exact duplication. We present TRACER, a semantic-aware framework for fine-grained code contamination detection. TRACER models contamination using three levels of semantic overlap - Functionally Identical, Nearly Identical, and Shared Logic - and detects them through a coarse-to-fine pipeline. We also introduce the first benchmark for fine-grained code contamination detection, spanning three widely used benchmarks and three representative post-training datasets. TRACER achieves strong and consistent performance across multiple LLM backbones, with GPT-5 reaching an F1 score of 0.91 in fine-grained detection. In the binary setting, TRACER attains an F1 of 0.92, outperforming existing methods by 42%-217%. We further conduct ablation studies and error analysis to assess the contributions of individual components in TRACER.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

評価工学に向けて: 実環境における ML 評価ハーネスの実証的研究

評価ハーネスは、モデルの呼び出し、データの読み込み、メトリクスの計算、結果レポートを管理することによってモデルの評価を調整するソフトウェア システムです。機械学習インフラストラクチャにおける重要な役割にもかかわらず、その運用上の課題やエンジニアリング上の懸念は、これまでのところあまり注目されていません。 57 の評価ハーネスに関する実証研究を紹介し、5 段階のハーネス モデルを導き出し、16,560 件の問題をワークフローの段階と根本原因ごとに分類しました。ハーネスの運用上の課題のほとんどは、ハーネスが外部モデル、データセット、採点審査員を統合する仕様段階 (問題の 41.4%) に集中しています。運用上の問題で最も頻繁に発生する 3 つの根本原因は、未実装の機能 (24.3%)、ドキュメントのギャップ (20.3%)、および入力検証の欠如 (17.2%) であり、これらは合わせて分類された問題の 61.7% を占め、既存の機能の欠陥と、意図したワークフローをブロックする機能のギャップの両方に及びます。根本原因はワークフローの段階によっても異なります。環境の非互換性と外部依存関係の破損がプロビジョニングの問題の 36.2% を占めますが、アルゴリズム エラー (25.9%) と検証ギャップ (22.5%) が評価の問題の大半を占めています。これらの貢献により、評価エンジニアリングを別個のソフトウェア エンジニアリングの問題として扱うための経験的基盤が確立されます。

原文 (English)

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild

Evaluation harnesses are software systems that orchestrate model evaluation by managing model invocation, data loading, metric computation, and result reporting. Despite their critical role in machine learning infrastructure, their operational challenges and engineering concerns have received limited attention so far. We present an empirical study of 57 evaluation harnesses, deriving a five-stage harness model and classifying 16,560 issues by workflow stage and root cause. Most harness operational challenges concentrate in the Specification stage (41.4% of issues), where harnesses integrate external models, datasets, and scoring judges. The three most frequent root causes of operational challenges are unimplemented features (24.3%), documentation gaps (20.3%), and missing input validation (17.2%), which together account for 61.7% of classified issues, spanning both defects in existing functionality and capability gaps that block intended workflows. Root causes also vary by workflow stage: environment incompatibility and external dependency breakage account for 36.2% of provisioning issues, whereas algorithmic error (25.9%) and validation gap (22.5%) dominate assessment issues. Together, these contributions establish an empirical foundation for treating evaluation engineering as a distinct software engineering concern.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

詳細な憲法定義と AI を活用した評価によりラベルの一貫性を向上

多くの自動ラベル付けパイプラインは、入力を仕様書で定義されたカテゴリに分類しており、コンテンツのモデレーションが顕著な使用例です。単純なカテゴリ定義では、ラベラーがこれらのパイプラインに必要な正確で一貫性のあるゴールデン ラベルを作成できるほど詳細ではありません。解決策の 1 つは、ラベリング担当者が文書化された解釈に同意できないほど実際の境界ケースを解決する規範的な定義を作成することです。実際には、その詳細レベルの定義は人間のアノテーターが作業記憶に保持できる範囲を超えているため、アノテーターは直感に頼り、ラベルは文書化されたルールから逸脱し、精度と一貫性が低下します。私たちは、AI 主導のワークフローの有効性を提案および実証します。AI は、エッジ ケースをカバーするのに十分詳細にラベルを定義するカテゴリごとの構成の作成を支援し、フロンティア LLM が入力ごとにそれを解釈して、人間が同じ文書を読むよりも一貫性と正確なゴールデン ラベルを生成します。私たちはコンテンツモデレーションの 3 つのカテゴリ (ハラスメント、ヘイトスピーチ、非暴力犯罪) を評価し、このアプローチにより、モデル間の不一致が仕様のギャップを診断し、個々のラベル付けの呼び出しではなく、各カテゴリが何を意味するかについての高レベルの決定を担当する人間が担当することにより、モデル間の不一致が段落定義と比較して最大 57 倍削減されることを示しました。安全性評価については、会話全体にわたって意図と内容を個別にスコアリングする二重軸の定式化を導入しているため、下流の消費者はどちらかの軸または両方に基づいて行動できます。

原文 (English)

Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation

Many automated labeling pipelines classify inputs into categories defined by a written specification, content moderation being a prominent use case. Simple category definitions are not detailed enough for labelers to produce the accurate, consistent golden labels these pipelines require. One solution is to write a prescriptive definition that settles enough real boundary cases that labelers cannot disagree with the written interpretation. In practice, definitions at that level of detail exceed what a human annotator can hold in working memory, so annotators fall back on intuition and the labels drift from the written rules, regressing on accuracy and consistency. We propose and demonstrate the efficacy of an AI-driven workflow in which AI helps write a per-category constitution that defines the label in enough detail to cover edge cases, and a frontier LLM interprets it on each input to produce the golden label more consistently and accurately than humans reading the same document. We evaluate on three content moderation categories (harassment, hate speech, non-violent crime) and show that the approach reduces cross-model inconsistency by up to 57x compared to paragraph definitions, with cross-model disagreement diagnosing specification gaps and the human responsible for high-level decisions about what each category should mean rather than individual labeling calls. For the safety evaluation, we introduce a dual-axis formulation scoring intent and content independently over the full conversation, so downstream consumers can act on either axis or both.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

特許埋め込みのベンチマーク: 検索、分類、クラスタリングにわたる 22 モデルのマルチタスク評価

どの微調整シグナルが特許埋め込みモデルを改善しますか?また、利益は特許環境全体に移行しますか? 22M パラメータのエンコーダから 12B の命令調整 LLM まで、検索、分類、クラスタリングに関して 22 の埋め込みモデルのベンチマークを行います。この研究では、113,148 件の WIPO 支援技術特許、46,069 件の引用グラフ検索クエリ、および外部検証用の公開 DAPFAM データセットを使用しています。当社のフレームワークは、引用ベースの検索、ハイブリッド疎密融合、5 つのデータセットにわたるマルチラベル分類、教師なしクラスタリング、6 つのテキスト セクション ビュー、4 つのモデルのドメイン適応微調整、管轄分析、および独自の DWPI (Derwent World Patents Index、Clarivate) の専門家が執筆したコンテンツをカバーしています。結果は、微調整がタスクに依存していることを示しています。単一ランドスケープの調整はドメイン内のスコアを改善できますが、外部ランドスケープでの取得に悪影響を与えることが多く、より多くのドメイン データが常に役立つという仮定に疑問を呈します。モデル ファミリ内では、通常、スケールによってパフォーマンスが予測されます (Qwen3 0.6B から 4B から 8B、Llama-Nemotron 1B から 8B)。ただし、ファミリ間のスケーリングにはノイズが多く、12B KaLM-Gemma3 は TAC 検索で 8 位にランクされますが、Qwen3-0.6B は ARI クラスタリングで首位に立っています。 Title+Abstract+Claims は最も信頼性の高いテキスト表現です。マルチビューの抽象クレームの調整により、検索が nDCG@10 で最大 7.1 パーセント向上し、微調整の組み合わせにより最も強力な分類ゲイン (+7.1 F1) が得られます。すべてのモデルはドメイン外クエリで 55 ~ 65% 低下しますが、ハイブリッドの疎-密融合ではこのギャップは埋められません。 BM25 密補間では、適度な nDCG@10 ゲイン (+0.002 ~ +0.015) が得られますが、より弱いゼロショット密モデルでは大きな利点が得られます。コードと評価フレームワークは公開されています。

原文 (English)

Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

Which fine-tuning signals improve patent embedding models, and do gains transfer across patent landscapes? We benchmark 22 embedding models, from 22M-parameter encoders to 12B instruction-tuned LLMs, on retrieval, classification, and clustering. The study uses 113,148 WIPO assistive-technology patents, 46,069 citation-graph retrieval queries, and the public DAPFAM dataset for external validation. Our framework covers citation-based retrieval, hybrid sparse-dense fusion, multi-label classification over five datasets, unsupervised clustering, six text-section views, domain-adaptive fine-tuning of four models, jurisdiction analysis, and proprietary DWPI (Derwent World Patents Index, Clarivate) expert-written content. Results show that fine-tuning is task-dependent: single-landscape tuning can improve in-domain scores but often hurts retrieval on an external landscape, challenging the assumption that more domain data always helps. Within model families, scale usually predicts performance (Qwen3 0.6B to 4B to 8B; Llama-Nemotron 1B to 8B), but cross-family scaling is noisy: the 12B KaLM-Gemma3 ranks 8th on TAC retrieval, while Qwen3-0.6B leads ARI clustering. Title+Abstract+Claims is the most reliable text representation. Multi-view abstract-claim alignment improves retrieval by up to 7.1 percent nDCG@10, while combined fine-tuning gives the strongest classification gains (+7.1 F1). All models drop by 55-65 percent on out-of-domain queries, and hybrid sparse-dense fusion does not close this gap. BM25-dense interpolation gives modest nDCG@10 gains (+0.002 to +0.015), with larger benefits for weaker zero-shot dense models. Code and evaluation framework are publicly available.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

プロンプト方式全体にわたる LLM 生成コードのセキュリティの実証的評価

自動コード生成のための大規模言語モデル (LLM) の使用の増加により、ソフトウェア開発の効率が向上しましたが、多くの場合、セキュリティが犠牲になります。生成されたコードは重大な問題を見落とすことが多く、暗号化が弱く、入力検証が不適切であるなどの問題に対して脆弱なままになっています。この問題を調査するために、5 つの LLM と 4 つのプログラミング言語 (Java、C++、C、Python) にわたる LLM 生成コードのセキュリティ品質の包括的な実証的評価を示し、複数のプロンプト エンジニアリング手法の影響を調べます。モデル推論をガイドする CWE マッピングを使用して、セキュリティ コンテキストでプロンプトを充実させる、弱点を認識したゼロショット思考連鎖 (WA-0CoT) プロンプト戦略を導入します。カイ二乗検定に裏付けられた当社の実証分析では、プロンプト手法全体で脆弱性の頻度や密度に統計的に有意な減少は見られませんでした。ただし、WA-0CoT を含むプロンプト戦略は CWE カテゴリの構成分布に体系的に影響を与え、その効果はプログラミング言語によって異なります。これらの調査結果は、セキュリティを意識したプロンプトによって生成された弱点の構造が変化する一方で、全体的な脆弱性レベルを確実に低減するにはプロンプト エンジニアリングだけでは不十分であることを示唆しています。この結果は、LLM で生成されたコードのセキュリティ プロパティを評価する際に、言語とモデルを意識したプロンプト設計の重要性を強調しています。

原文 (English)

An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods

The growing use of Large Language Models (LLMs) for automated code generation has enhanced software development efficiency, but often at the cost of security. Generated code frequently overlooks critical concerns, leaving it vulnerable to issues such as weak encryption and improper input validation. To investigate this problem, we present a comprehensive empirical evaluation of the security quality of LLM-generated code across five LLMs and four programming languages (Java, C++, C, and Python), examining the impact of multiple prompt engineering methods. We introduce a weaknesses-aware zero-shot chain-of-thought (WA-0CoT) prompting strategy that enriches prompts with security context using CWE mappings to guide model reasoning. Our empirical analysis, supported by chi-square tests, finds no statistically significant reductions in vulnerability frequency or density across prompt methods. However, prompting strategies, including WA-0CoT, systematically influence the compositional distribution of CWE categories, with effects varying by programming language. These findings suggest that while security-aware prompting alters the structure of generated weaknesses, prompt engineering alone is insufficient to reliably reduce overall vulnerability levels. The results highlight the importance of language-aware and model-aware prompt design when evaluating the security properties of LLM-generated code.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

HoloFair: 統合された T2I 公平性評価と Fair-GRPO のバイアス軽減

Text-to-Image (T2I) モデルは、視覚的なリアリズムと意味の一貫性において大幅な進歩を遂げましたが、社会的な偏見を永続させ、増幅させることがよくあります。既存の評価方法は通常、一次元のバイアスのみに対処しており、社会関連のより深い意味レベルでモデルのバイアスを明らかにする視点が欠けています。多次元の人口統計的バイアス分析のための包括的なベンチマーク フレームワークである HoloFair を紹介します。このフレームワークは、大規模な公平性指向のデータセットと SpaFreq (空間周波数) 属性分類器に基づいて構築されており、本質的な多様性と条件付きバイアスの両方を評価するように設計された、複数属性グループワイズ バイアス インデックス (MGBI) メトリクスを提案しています。評価を超えて、設計された多目的報酬関数を通じて生成モデルの分布を変更する強化学習ベースのバイアス除去手法である Fair-GRPO をさらに導入します。たとえば、SD3.5-Medium モデルの実験では、Fair-GRPO が高画質を維持しながら多次元の公平性を大幅に向上させることが実証されています。また、潜在的な報酬ハッキング現象を分析し、対応する緩和戦略を提供します。コードとデータセットは https://github.com/1059684669/HoloFair で入手できます。

原文 (English)

HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing

Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal biases. Existing evaluation methods typically address only single-dimensional biases, lacking perspectives to uncover model biases at social-related deeper semantic levels. We introduce HoloFair, a comprehensive benchmark framework for multidimensional demographic bias analysis. Built upon our large-scale fairness-oriented dataset and the SpaFreq (Spatial-Frequency) attribute classifier, this framework proposes the Multi-attribute, Group-wise Bias Index (MGBI) metric, designed to assess both intrinsic diversity and conditional biases. Beyond evaluation, we further introduce Fair-GRPO, a reinforcement-learning-based debiasing method that alters the distribution of generative models through a designed multi-objective reward function. E.g., experiments on the SD3.5-Medium model demonstrate that Fair-GRPO significantly improves multidimensional fairness while maintaining high image quality. We also analyze potential reward hacking phenomena and provide corresponding mitigation strategies. Code and dataset are available at https://github.com/1059684669/HoloFair

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

NMS フリー時代のマルチスケール リアルタイム物体検出: YOLOv8 と YOLO26 のパフォーマンスの比較評価

非最大値抑制 (NMS) は、多くのリアルタイム物体検出パイプラインにおける重要な後処理ステップであり続けますが、リソースに制約のある設定では遅延の変動や展開の複雑さが生じる可能性があります。 YOLO26 などの最近の NMS フリー設計は、エンドツーエンドの検出を通じてこの依存性を軽減することを目的としていますが、YOLOv8 などの確立された NMS ベースのモデルと比較したパフォーマンスは、標準ベンチマークを超えてまだ調査されていません。この論文では、Pascal VOC と VisDrone での YOLOv8 と YOLO26 を比較し、それぞれ一般物体検出と高密度空中微小物体検出を表します。どちらのモデル ファミリも、精度、ローカリゼーション、モデル サイズ、GFLOP、CPU/GPU レイテンシーを使用して 5 つのスケールにわたって評価されます。結果は、ほとんどのスケールにおいて、Pascal VOC では YOLO26 がより強力な検出パフォーマンスとより低いモデルの複雑さを達成する一方、VisDrone ではパフォーマンスの差が縮まり、両方のモデルが密集した小さなターゲットに苦戦していることが示されています。 YOLOv8 は GPU レイテンシーにおいて競争力を維持しており、NMS フリーの設計が普遍的な展開の優位性を保証するものではないことを示しています。全体として、この研究は、検出器の選択がデータセットの特性、オブジェクトの規模、モデルの容量、ハードウェアの制約に依存することを示しています。

原文 (English)

Multiscale Real-Time Object Detection in the NMS-Free Era: A Comparative Performance Evaluation of YOLOv8 and YOLO26

Non-Maximum Suppression (NMS) remains a key post-processing step in many real-time object detection pipelines, but it can introduce latency variation and deployment complexity in resource-constrained settings. Recent NMS-free designs such as YOLO26 aim to reduce this dependence through end-to-end detection, yet their performance relative to established NMS-based models such as YOLOv8 remains underexplored beyond standard benchmarks. This paper compares YOLOv8 and YOLO26 on Pascal VOC and VisDrone, representing general object detection and dense aerial small-object detection, respectively. Both model families are evaluated across five scales using accuracy, localization, model size, GFLOPs, and CPU/GPU latency. Results show that YOLO26 achieves stronger detection performance and lower model complexity on Pascal VOC across most scales, while the performance gap narrows on VisDrone, where both models struggle with dense small targets. YOLOv8 remains competitive in GPU latency, showing that NMS-free design does not guarantee universal deployment superiority. Overall, the study shows that detector selection depends on dataset characteristics, object scale, model capacity, and hardware constraints.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

推論が困難な場合: 臨床 SOAP ノート生成のためのフロンティア LLM のソース認識型評価

推論対応 LLM は医療推論ベンチマークで優れたパフォーマンスを発揮しますが、これらの利点が構造化された臨床文書に反映されるかどうかは不明のままです。私たちは、OMI Health、ACI-Bench、PriMock57 にわたるソース認識ベンチマークでの臨床対話からの SOAP ノート生成を使用して、この疑問を調査します。プロバイダーネイティブ推論と同一ソース検索拡張生成 (RAG) を個別に切り替える制御された 2x2 設計で GPT-5.4、DeepSeek-V4-Flash、および Gemma-4-E4B を評価します。成果は、リファレンスを認識した 2 人の LLM 審査員とともに 7 つの自動指標を使用して評価されます。どちらの評価アプローチでも、推論非対応の GPT-5.4 構成が全体として最高の品質を達成するのに対し、DeepSeek-V4-Flash は推論が有効な構成の中で最高のパフォーマンスを発揮するという点で一致しています。推論を有効にすると、3 つのデータセットすべてで GPT-5.4 のパフォーマンスが大幅に低下しますが、同じソースの RAG では、モデルに依存する改善は小さくなります。全体として、この調査結果は、専用のタスク固有の評価を行わずに、忠実度に敏感な SOAP ノート生成を向上させるために、より強力な推論機能を想定すべきではないことを示しています。

原文 (English)

When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation

Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical documentation; we investigate this question using SOAP note generation from clinical dialogue in a source-aware benchmark spanning OMI Health, ACI-Bench, and PriMock57. We evaluate GPT-5.4, DeepSeek-V4-Flash, and Gemma-4-E4B in a controlled 2x2 design that independently toggles provider-native reasoning and same-source retrieval-augmented generation (RAG). Outputs are assessed using seven automatic metrics alongside two reference-aware LLM judges. Both evaluation approaches agree that a non-reasoning GPT-5.4 configuration achieves the highest overall quality, while DeepSeek-V4-Flash performs best among reasoning-enabled configurations. Enabling reasoning significantly degrades GPT-5.4 performance across all three datasets, whereas same-source RAG yields smaller, model-dependent improvements. Overall, the findings indicate that stronger reasoning capability should not be assumed to improve fidelity-sensitive SOAP note generation without dedicated, task-specific evaluation.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

RealBench: 運用条件および異常事態の課題におけるデータ主導の数値天気予報のベンチマーク

気象予測モデルを現実世界のアプリケーションに確実に導入するには、気象予測モデルを正確に評価することが重要です。ただし、既存のベンチマークは主に ERA5 などの再分析製品に依存しています。ERA5 は遅延したデータ同化によって生成され、リアルタイムの運用予測の制約を反映していないため、ベンチマークのパフォーマンスと現実世界の予測の間に体系的な不一致が生じます。本研究では、運用条件下での現実的な評価を重視したAI天気予報の次世代ベンチマークであるRealBenchを紹介します。 RealBench は、データ漏洩を排除し、最近の大気状況を把握するために、2025 年までの厳密に配布外のテスト セットを備えています。低遅延の運用分析や 10,000 を超える観測点からなる大規模な地球規模の現場観測データセットを含む複数のデータ ソースを統合し、実際の大気測定に対する直接評価を可能にします。 RealBench は、標準的な地球規模の指標を超えて、現実世界の予測の優先順位をより適切に反映するイベント固有の指標を使用して、熱波、寒波、熱帯低気圧などの影響の大きい極端な現象に対する包括的な評価フレームワークを提供します。評価結果では、特に極端な現象に関して、再分析に基づく指標と現実世界のパフォーマンスとの間に大きな差異があることが明らかになりました。この研究では、既存のベンチマークの限界を強調することで、より忠実で運用上適切な評価パラダイムを確立し、次世代 AI 天気予報システムを進化させるための厳密な基盤を提供します。ベンチマークの実装は、https://github.com/lixruize-del/NWP-Benchmark から入手できます。

原文 (English)

RealBench: Benchmarking Data-Driven Numerical Weather Forecasting Under Operational Conditions and Extreme Event Challenges

Accurate evaluation of weather forecasting models is critical for their reliable deployment in real-world applications. However, existing benchmarks predominantly rely on reanalysis products such as ERA5, which are generated through delayed data assimilation and do not reflect the constraints of real-time operational forecasting, thereby resulting in a systematic mismatch between benchmark performance and real-world forecasting. In this work, we introduce RealBench, a next-generation benchmark for AI weather forecasting that emphasizes realistic evaluation under operational conditions. RealBench features a strictly out-of-distribution test set spanning 2025 to eliminate data leakage and capture recent atmospheric regimes. It integrates multiple data sources, including low-latency operational analysis and a large-scale global in-situ observation dataset comprising over 10,000 stations, enabling direct evaluation against real atmospheric measurements. Beyond standard global metrics, RealBench provides a comprehensive evaluation framework for high-impact extreme events, including heatwaves, cold surges, and tropical cyclones, using event-specific metrics that better reflect real-world forecasting priorities. The evaluation results reveal substantial discrepancies between reanalysis-based metrics and real-world performance, particularly concerning extreme events. By highlighting the limitations of existing benchmarks, this work establishes a more faithful and operationally relevant evaluation paradigm, providing a rigorous foundation for advancing next-generation AI weather forecasting systems. The benchmark implementation is available at: https://github.com/lixruize-del/NWP-Benchmark.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

大規模言語モデルの微調整ライフサイクルにおけるセキュリティ: 脅威、防御、評価、および将来の方向性

背景: 微調整は、事前トレーニングされた大規模言語モデル (LLM) を下流のタスクに適応させる上で中心となりますが、トレーニング データ、パラメーターの更新、再利用可能なコンポーネントへの依存により、攻撃者の侵入口が開かれてしまいます。脅威は、データポイズニングや重みの改ざんから、エージェントの操作やインターフェースの悪用へと進化していますが、既存のレビューには、完全な微調整ライフサイクルにわたる統一されたフレームワークが欠けています。目的: この文書では、LLM セキュリティの微調整に関する体系的な調査を提示し、統一された経験的評価によって補完された、攻撃と防御を比較するためのライフサイクルベースのフレームワークを確立します。方法: 攻撃と防御のメカニズムを、介入タイミングによってプレチューニング、チューニング中、ポストチューニングの 3 つのフェーズに分けます。各フェーズ内で戦略がレビューされ、比較され、その進化と限界が明らかになります。次に、代表的な手法が、統一されたモデル、ハードウェア、プロトコル設定の下で評価され、異なるフェーズからの攻撃と防御を組み合わせたクロスフェーズ実験が行われます。結果: 攻撃の有効性はモデルに大きく依存しており、規模に応じて単調ではありません。初期のモデルで有効だった重み編集攻撃は、最新のオープンソース LLM には影響を及ぼしません。言語を越えたバックドア転送は、大規模ではほぼ完璧であると報告されていますが、テストされた 1B ~ 4B モデルでは完全に失敗します。また、純粋に良性のサンプルは、命令調整されたモデルの安全性の調整を損なう可能性があります。単一フェーズの防御がフェーズ全体に一般化することはほとんどなく、防御の有効性はモデルのアーキテクチャと調整状態に共同で依存します。結論: 私たちは、主要な未解決の問題 (構成堅牢な防御、クロスフェーズ防御構成、および行動の想定を超えた埋め込み空間攻撃) を特定し、具体的な将来の研究の方向性を提案します。

原文 (English)

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions

Background: Fine-tuning is central to adapting pre-trained Large Language Models (LLMs) to downstream tasks, but its reliance on training data, parameter updates, and reusable components opens entry points for attackers. Threats have evolved from data poisoning and weight tampering to agent manipulation and interface exploitation, yet existing reviews lack a unified framework spanning the full fine-tuning lifecycle. Objective: This paper presents a systematic survey of LLM fine-tuning security and establishes a lifecycle-based framework for comparing attacks and defenses, complemented by unified empirical evaluation. Methods: We divide attack and defense mechanisms into three phases by intervention timing: pre-tuning, during-tuning, and post-tuning. Within each phase, strategies are reviewed and contrasted to expose their evolution and limitations. Representative methods are then evaluated under a unified model, hardware, and protocol setup, with cross-phase experiments pairing attacks and defenses from different phases. Results: Attack effectiveness is highly model-dependent and non-monotonic with scale: weight-editing attacks effective on earlier models lose impact on modern open-source LLMs; cross-lingual backdoor transfer, reported as near-perfect at larger scales, fails entirely on tested 1B-4B models; and purely benign samples can compromise safety alignment in instruction-tuned models. Single-phase defenses rarely generalize across phases, and defense effectiveness depends jointly on model architecture and alignment state. Conclusion: We identify key open problems (configuration-robust defense, cross-phase defense composition, and embedding-space attacks beyond behavioral assumptions) and propose concrete future research directions.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

JudgmentBench: 品質評価のためのルーブリックと優先評価の比較

現在のベンチマーク手法は 2 つの方法論が主流となっています。ルーブリック ベースのスコアリングでは、事前定義された基準に照らして項目を評価します。一方、比較判断では、出力間のペアごとの優先順位を導き出します。どちらの方法論も広く使用されていますが、どちらを選択するかが正当化されることはほとんどありません。当社は、米国の大手法律事務所を含む豊富な経験を持つ現役弁護士から収集した 1,539 のルーブリック スコアと 1,530 のペアごとの優先判断を組み合わせた、30 の実際の法律タスクのベンチマークである JudgmentBench をリリースします。アノテーションは、両方の監視シグナルが同じ項目について同じ専門家から引き出される、高度な専門知識の領域における最初の公的に利用可能なデータセットを構成します。 3 つの構築された品質レベルで LLM によって生成された出力を使用して、最初の経験的比較を提供します。比較判断により、ルーブリックよりも大幅に意図した品質順序が回復されます (スピアマンの順位相関の平均 0.908 対 0.150、推定差 = 0.758 [0.494, 1.021])。必要なアノテーション時間は半分未満です。このパターンは、ヒューマン アノテーターと LLM 自動採点者にも当てはまります。この最初の比較を超えて、データセットのペア構造は、検証可能なグラウンドトゥルースのない領域で専門家の判断をどのように導き出し、集約し、監視として使用するかについてのより広範な研究課題をサポートします。

原文 (English)

JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs. Although both methodologies are widely used, the choice between them is rarely justified. We release JudgmentBench, a benchmark of 30 real-world legal tasks, paired with 1,539 rubric scores and 1,530 pairwise preference judgments collected from practicing attorneys--including at major U.S. law firms--with substantial experience. The annotations constitute the first publicly available dataset in a high-expertise domain in which both supervision signals are elicited from the same experts on the same items. Using LLM-generated outputs at three constructed quality levels, we provide an initial empirical comparison: comparative judgments recover the intended quality ordering substantially better than rubrics (mean Spearman's rank correlation of 0.908 vs. 0.150, estimated difference = 0.758 [0.494, 1.021]) while requiring less than half the annotation time. The patterns hold for human annotators and LLM autograders. Beyond this initial comparison, the paired structure of the dataset supports a broader research agenda on how expert judgment should be elicited, aggregated, and used as supervision in domains without verifiable ground truth.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

アノテーション不要の超音波平面品質管理のための、部分空間に基づくセマンティックおよびトポロジカル不変式登録

超音波画像の信頼性の高い品質管理 (QC) は、リアルタイム収集ガイダンスと遡及的臨床監査の両方に不可欠ですが、既存のアプローチは面ごとのアノテーションに大きく依存するか、臨床収集に固有の空間変形の下で体系的なバイアスが生じやすい擬似ラベリングを採用しています。我々は、アノテーションフリーの US プレーン品質管理を部分空間誘導型一貫性測定問題として再構築する、登録主導型フレームワークである STRIQ を紹介します。具体的には、STRIQ は、クエリ画像と分散駆動アンカーの間の階層的な特徴空間の対応関係を確立するための Latent Registration Aligner (LRA) を導入しています。分散駆動アンカーは、構造的に安定したプロトタイプとして機能する分散スペクトル基準を介してラベルなしデータから自律的に抽出されます。解剖学的平面をさらに明確にし、否定的な知識伝達を軽減するために、直交知識部分空間 (OKS) モジュールを提案します。 OKS は、プレーン固有の表現を相互に直交するサブスペースに分解し、プレーン間の干渉を防止しながらきめ細かい専門家のコラボレーションを可能にし、原則的なサブスペースの近接性に基づく品質メトリクスを保証します。社内 US4QA および公開 CAMUS データセットに対する広範な実験により、STRIQ が臨床品質スコアとの最先端の相関関係を実現し、注釈不要でリアルタイムの信頼性の高い超音波品質管理のための新しいパラダイムを確立することが実証されました。私たちのコードは https://github.com/zhcz328/STRIQ で入手できます。

原文 (English)

Subspace-Guided Semantic and Topological Invariant Registration for Annotation-Free Ultrasound Plane Quality Control

Reliable quality control (QC) of ultrasound images is essential for both real-time acquisition guidance and retrospective clinical audit, yet existing approaches rely heavily on per-plane annotations, or employ pseudo-labeling prone to systematic bias under spatial deformations inherent in clinical acquisition. We present STRIQ, a registration-driven framework that recasts annotation-free US plane quality control as a subspace-guided consistency measurement problem. Specifically, STRIQ introduces a Latent Registration Aligner (LRA) to establish hierarchical feature space correspondences between query images and variance-driven anchors, which are autonomously distilled from unlabeled data via a variance spectrum criterion to serve as structurally stable prototypes. To further disambiguate anatomical planes and mitigate negative knowledge transfer, we propose an Orthogonal Knowledge Subspace (OKS) module. The OKS decomposes plane-specific representations into mutually orthogonal subspaces, enabling fine-grained expert collaboration while preventing inter-plane interference, ensuring that the quality metric is grounded in principled subspace proximity. Extensive experiments on the in-house US4QA and public CAMUS datasets demonstrate that STRIQ achieves state-of-the-art correlation with clinical quality scores, establishing a new paradigm for annotation-free, real-time reliable ultrasound quality control. Our code is available at https://github.com/zhcz328/STRIQ.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

SomaliBench Eval: 無差別言語モデルにおける英語とソマリアの拒否格差の測定

大規模言語モデルの安全性評価は引き続き英語中心であり、モデルが世界的に展開されている場合でも、低リソース言語の評価は不十分なままです。英語とソマリ語を組み合わせた 100 個の有害な意図を持ったプロンプトのネイティブ作成者検証済みベンチマークである SomaliBench v0 で、4 つのオープンウェイト命令調整モデルを評価します。 Llama-3.1-8B-Instruct、Gemma-2-9B-Instruct、Qwen-2.5-7B-Instruct、Aya-23-8B はそれぞれ、温度 0 および同じ英語の「有益、無害、正直」(HHH) システム プロンプトでローカルで実行されます。固定された Claude Sonnet のスナップショット (claude-sonnet-4-5-20250929) では、各応答が拒否、準拠、または不明瞭として分類されます。ネイティブの作成者は、層化された 80 行のサンプルを抜き取りチェックします。 4 つのモデルすべてで英語とソマリ語の拒否の大きなギャップが見つかりました: Llama-3.1-8B (0.90; 95% ブートストラップ CI [0.85, 0.96])、Aya-23-8B (0.75 [0.67, 0.83])、Qwen-2.5-7B (0.69 [0.59, 0.78])、ジェマ-2-9B (0.38 [0.27, 0.49])。 3 つのモデルにおいて、ソマリア人の主な非拒否モードは、流暢で有害なコンプライアンスではなく、不明確な出力、つまり空っぽ、間違った言語、または支離滅裂な世代です。ネイティブ検証のスポットチェックでは、サンプリングされた 80 行について、裁判官との 100% の一致 (コーエンのカッパ = 1.00) が達成されます。私たちは、総拒否率、カテゴリーギャップ、および信頼性統計のみを報告します。生のモデル世代はローカルに保持され、リリースされません。

原文 (English)

SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models

Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed globally. We evaluate four open-weight instruction-tuned models on SomaliBench v0, a native-author-verified benchmark of 100 harmful-intent prompts paired across English and Somali. Each of Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, Qwen-2.5-7B-Instruct, and Aya-23-8B is run locally with temperature 0 and the same English "helpful, harmless, and honest" (HHH) system prompt. A pinned Claude Sonnet snapshot (claude-sonnet-4-5-20250929) classifies each response as refused, complied, or unclear; the native author spot-checks a stratified 80-row sample. We find large English-to-Somali refusal gaps for all four models: Llama-3.1-8B (0.90; 95% bootstrap CI [0.85, 0.96]), Aya-23-8B (0.75 [0.67, 0.83]), Qwen-2.5-7B (0.69 [0.59, 0.78]), and Gemma-2-9B (0.38 [0.27, 0.49]). For three models, the dominant Somali non-refusal mode is not fluent harmful compliance but unclear output: empty, wrong-language, or incoherent generations. The native verification spot-check achieves 100% agreement with the judge (Cohen's kappa = 1.00) on the 80 sampled rows. We report aggregate refusal rates, category gaps, and reliability statistics only; raw model generations are retained locally and are not released.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

手術フィードバックの質を評価するためのマルチエージェント LLM フレームワーク

手術室で担当外科医が行う口頭によるフィードバックは、研修医のスキル習得において重要な形成的役割を果たします。しかし、トレーナーのフィードバックの質と、実際の手術中の研修生の行動に影響を与えるその効果を評価することは依然として課題です。これまでの研究では、専門の人間評価者による広範な手動注釈に依存してフィードバックの内容を評価し、明確さや緊急性などのフィードバック配信の定性的側面を無視する広範な分類法の開発に焦点を当てていました。キーワード分析やトピックモデリングなどの限られた既存の自動化手法も、こうした微妙な側面を捉えることができません。外科訓練のコンテキストに基づいた、解釈可能なフィードバック品質基準を発見する 2 段階の LLM ベースのフレームワークを導入します。私たちの方法は、マルチエージェントのプロンプトと外科領域の知識の注入を使用して、人間が解釈できる少数のスコアリング基準(例:奨励、緊急、明確)を発見します。これらの基準は、LLM-as-a-judge アプローチを介してライブ手術フィードバックを自動的に採点するために使用されます。 4.2,000 のトレーナー フィードバック インスタンスの評価では、AI が発見した基準が、観察されたトレーニング生の行動調整やトレーナーの承認など、フィードバックの有効性の予測において、以前のコンテンツベースのフレームワークよりも優れていることが実証されました。この取り組みは、手術室におけるコミュニケーション品質のスケーラブルで人間に合わせた評価を前進させ、外科教育実践を改善するための基盤を提供します。

原文 (English)

A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback

Verbal feedback delivered by attending surgeons in the operating room plays a critical formative role in resident trainee skill acquisition. Yet, assessing the quality of trainer feedback and its effectiveness in influencing trainee behavior during live surgery remains a challenge. Prior studies assessed feedback content relying on extensive manual annotation by expert human raters and focused on developing broad taxonomies that overlook the qualitative aspects of feedback delivery such as clarity or urgency. Limited existing automated methods, including keyword analysis and topic modeling, also fail to capture these nuanced aspects. We introduce a two-stage LLM-based framework that discovers interpretable feedback quality criteria grounded in the context of surgical training. Our method uses multi-agent prompting and surgical domain knowledge injection to discover a small set of human interpretable scoring criteria (e.g., Encouraging, Urgent, Clear). These criteria are then used to automatically score live surgical feedback via an LLM-as-a-judge approach. Evaluation on 4.2k trainer feedback instances demonstrates that our AI-discovered criteria outperform prior content-based frameworks in predicting feedback effectiveness, including observed trainee behavioral adjustments and trainer approval. This work advances scalable, human-aligned assessment of communication quality in the operating room and provides a foundation for improving surgical teaching practices.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

AI評価の新たなパラダイムとしての参照セキュリティ

セキュリティ評価は本質的に安定した識別子に依存します。調査結果、監査、または規制上の決定は、それが関係する特定の成果物に関連付けられたままでなければなりません。継続的に更新される人工知能システムは、基礎となる重み、プロンプト、取得メカニズム、誤用分類子、推論設定、およびサービス提供インフラストラクチャが予告なく変更される一方で、公開モデルの指定が静的なままとなり、この中心的な前提に違反します。その結果、現在の評価は、識別可能な個別のシステムではなく、表面的なラベルに適用されることがよくあります。これを解決するために、私たちは AI 評価の新しいパラダイムとして参照セキュリティを提案します。基本的なセキュリティの問題は、モデルが安全かどうかを超えて、後続の当事者が特定の安全性主張がどのシステムに対応しているかを最終的に判断できるかどうかにまで及びます。このアプローチは、モデルの同一性を経験的に検証可能な特性として再構築し、参照安定性をそれが条件とする実質的な安全性の主張から分離します。このフレームワークは、現在の慣行ではうまく処理できない 3 つの重要なワークフローに扱いやすさをもたらします。具体的には、再現可能な評価、長期的な監査の妥当性、およびプロバイダ間の同等性が可能になります。これらの評価を検証可能な成果物に基づいて行うことで、私たちのアプローチは、安全性監査と規制上の調査結果が動的システムの運用ライフサイクル全体にわたって経験的有用性を維持することを保証します。

原文 (English)

Referential Security as a New Paradigm for AI Evaluations

Security evaluations inherently depend on stable identifiers. Any finding, audit, or regulatory decision must remain attached to the specific artifact it pertains to. Continuously updated artificial intelligence systems violate this core assumption, with public model designations remaining static while underlying weights, prompts, retrieval mechanisms, misuse classifiers, inference settings, and serving infrastructures undergo unannounced modifications. Consequently, current evaluations frequently apply to superficial labels rather than identifiable and distinct systems. To resolve this, we propose referential security as a new paradigm for AI evaluation. The fundamental security question extends beyond whether a model is safe to whether subsequent parties can conclusively determine which system a specific safety claim addressed. This approach reframes model identity as an empirically verifiable property and separates referential stability from the substantive security claims it conditions. This framework brings tractability to three critical workflows that current practices handle poorly. Specifically, it enables reproducible evaluation, longitudinal audit validity, and cross-provider equivalence. By grounding these evaluations in verifiable artifacts, our approach ensures that safety audits and regulatory findings maintain their empirical utility across the operational lifecycle of dynamic systems.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

Kubernetes マニフェスト生成のためのコンテキスト計測データの蒸留: 方法と実験的評価

このペーパーでは、ドメイン固有言語 (DSL) でアーティファクトを生成するための、最大 40 億のパラメーターを備えた小型言語モデル (SLM) の特殊化について検討します。 Kubernetes マニフェストがターゲット ドメインとして選択されます。私たちは、コンテキストを利用したデータの蒸留方法を提案します。ソース コーパスは、合成生成を通じて、また拡張スキームでは、実際の Kubernetes YAML ファイルからの逆命令生成を通じて形成されます。ペアは、外部バリデータを渡し、ドメイン コンテキスト モデルと一致する場合にのみトレーニングに含まれます。古典的な KL ダイバージェンス知識の蒸留とは異なり、ベースラインの実装では、機器で検証された例に対する教師あり微調整が行われます。実験セクションでは、リソースが制約された条件下でのパイロット実装を示します。DeepSeek-V4 Flash API は合成生成の教師として機能し、Qwen2.5-Coder-1.5B-Instruct は CPU 上の LoRA によって微調整されます。 K8s-Distill-Pilot コーパス (train_1200、validation_100、test_200) では、より厳密なプロンプト定式化と max_new_tokens=768 を使用して、フルパス @1 = 91.5% (183/200) を達成しました。重要な経験的発見は、Kubernetes YAML の場合、パイロットでの結果の品質は、単にトレーニング サンプルの数を増やすことよりも、厳格な出力形式要件に依存するということです。

原文 (English)

Context-Instrumental Data Distillation for Kubernetes Manifest Generation: Method and Experimental Evaluation

This paper examines the specialization of Small Language Models (SLMs) with up to 4 billion parameters for generating artifacts in domain-specific languages (DSL). Kubernetes manifests are chosen as the target domain. We propose the context-instrumental data distillation method: the source corpus is formed through synthetic generation and, in an extended scheme, through reverse instruction generation from real Kubernetes YAML files, with pairs included in training only upon passing external validators and matching the domain context model. Unlike classical KL-divergence knowledge distillation, the baseline implementation reduces to supervised fine-tuning on instrumentally verified examples. The experimental section presents a pilot implementation under resource-constrained conditions: the DeepSeek-V4 Flash API serves as the teacher for synthetic generation, while Qwen2.5-Coder-1.5B-Instruct is fine-tuned via LoRA on CPU. On the K8s-Distill-Pilot corpus (train_1200, validation_100, test_200), we achieved full-pass@1 = 91.5% (183/200) with a stricter prompt formulation and max_new_tokens=768. The key empirical finding is that for Kubernetes YAML, result quality in the pilot depended more on strict output format requirements than on simply increasing the number of training examples.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

特定の恐怖症データからの転移学習による心的外傷後ストレス障害の重症度の定量的評価

心的外傷後ストレス障害 (PTSD) は、個人および社会に重大な影響を与える、蔓延している衰弱性の精神的健康状態です。現在の PTSD の臨床評価は主観的な評価に依存していることが多く、時間と費用がかかり、人間の偏見が入りやすい可能性があります。この研究では、PTSD 重症度を客観的に評価するための多変量カーネル密度推定 (MKDE) 手法に基づく機械学習 (ML) アプローチを提案します。私たちは、没入型シミュレーション中に 21 人の参加者から心拍数 (HR) および電気皮膚反応 (GSR) 信号、ならびに PTSD チェックリスト - 軍事バージョン (PCL-M) ラベルを収集しました。恐怖反応モデルは公共のクモ恐怖症データセットでトレーニングされ、軍事データセットで推定された恐怖反応曲線から PTSD の予測特徴が抽出されました。このモデルは、PTSD 状態の分類において 86\% の精度を達成し、PTSD のある参加者とない参加者を効果的に区別しました (PCL-M 閾値 36)。モデルの平均絶対誤差 (MAE) の平均は 5.6 で、臨床 PTSD 重症度スケールは平均絶対誤差 17% と推定されました。私たちのアルゴリズムは、生理学を使用した客観的かつ低労力の評価アプローチを提供することで、PTSD 重症度の推定と追跡調査を強化する有望な可能性を示しています。これらの所見は、スクリーニングとフォローアップの両方の設定における臨床的有用性を示唆しています。

原文 (English)

Quantitative Evaluation of the Severity of Posttraumatic Stress Disorder through Transfer Learning from Specific Phobia Data

Posttraumatic stress disorder (PTSD) is a prevalent and debilitating mental health condition with significant personal and societal impacts. Current clinical assessments of PTSD often rely on subjective evaluations, which can be time-consuming, costly, and prone to human bias. This study proposes a machine learning (ML) approach based on multivariate kernel density estimation (MKDE) technique for the objective evaluation of PTSD severity. We collected heart rate (HR) and galvanic skin response (GSR) signals as well as PTSD Checklist - Military Version (PCL-M) labels from 21 participants during an immersive simulation. A fear-response model was trained on a public arachnophobia dataset, and predictive features of PTSD were extracted from the fear-response curves estimated on the military dataset. The model achieved an accuracy of 86\% in classifying PTSD status, effectively distinguishing participants with and without PTSD (PCL-M threshold of 36). The average mean absolute error (MAE) of the models is 5.6, and it estimated a clinical PTSD severity scale with a mean absolute percentage error of 17\%. Our algorithm demonstrates promising potential for enhancing estimation of PTSD severity and followup by offering an objective and low-effort evaluation approach using physiology. These findings suggest clinical utility in both screening and follow-up settings.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

QUIET: LLM クリエイティブ生成機能の複数空白のカスケード ストーリー Cloze ベンチマーク

大規模言語モデル (LLM) は、クリエイティブ能力の評価において 2 つの課題に直面しています。既存のベンチマーク (ストーリー クローズ テスト、HellaSwag など) は、クリエイティブ生成能力を直接測定するのではなく、多肢選択認識パラダイムを使用して、物語の継続に対するモデルの識別能力を測定します。ルーブリックベースのスコアリングおよび LLM-as-Judge メソッドは、主観的な次元の評価または自然言語モデルの出力に依存しており、客観的で自動化されたスコアリング メカニズムを提供できません。この論文では、複数の空白のカスケード ストーリー クローズに基づく LLM クリエイティブ能力の診断ベンチマークである QUIET (Quality Understanding via Interlocked Evaluation Testing) を提案します。 QUIET は、完全な構造を持つストーリーに N 個のブランク (10 ~ 20) を設定します。各ブランクには、明示的なコンテンツ制約が伴い、ブランク間のカスケード依存関係が伴います。つまり、前のブランクに埋められたコンテンツが、後のブランクの実行可能な解決スペースを制限します。評価されたモデル (または人間の参加者) は、オープンエンド生成モードですべての空白を埋めます。結果は、人間による採点を行わずに、情報理論に基づいた自動採点プロトコルによって採点されます。スコアリング プロトコルは、「調整されたサプライズ」理論的枠組みを直接運用します (Zou & Xu、2026a)。空白の k ごとに、複合スコアが計算されます: スコア = 満足 * (1 + ラムダ * サプライズ)、ここでラムダ = 1.0。ここで、「満足」は空白埋めが内容の制約 (主観的な美的スコアリングではなく、客観的な論理的推論の判断) をどの程度満たしているかを測定し、「驚き」は制約が満たされた場合の驚きの度合いを測定します。制約スコア 0 を満たさない創造的な回答。制約は満たしているものの、平均的なスコアの低い回答。制約を満たし、驚くほど高いスコアを示す回答。

原文 (English)

QUIET: A Multi-Blank Cascaded Story Cloze Benchmark for LLM Creative Generation Capability

Large language models (LLMs) face a dual challenge in creative capability evaluation: existing benchmarks (e.g., Story Cloze Test, HellaSwag) measure models' discriminative ability over narrative continuation using multiple-choice recognition paradigms, rather than directly measuring creative generation capability; rubric-based scoring and LLM-as-Judge methods rely on subjective dimension assessment or natural language model outputs, and cannot provide objective, automated scoring mechanisms. This paper proposes QUIET (Quality Understanding via Interlocked Evaluation Testing), a diagnostic benchmark for LLM creative capability based on multi-blank cascaded story cloze. QUIET sets N blanks (10-20) in a story with complete structure, with each blank accompanied by an explicit content constraint, and cascade dependency relationships between blanks -- the content filled into earlier blanks constrains the feasible solution space for later blanks. The evaluated model (or human participants) fills all blanks in open-ended generation mode; the results are scored by an information-theoretic automated scoring protocol without human grading. The scoring protocol directly operationalizes the "calibrated surprise" theoretical framework (Zou & Xu, 2026a). For each blank k, a composite score is computed: score = satisfy * (1 + lambda * surprise), where lambda = 1.0. Here, "satisfy" measures how well the blank filling satisfies the content constraint (objective logical reasoning judgment, not subjective aesthetic scoring), and "surprise" measures the degree of surprise given that the constraint is satisfied. Creative answers that do not satisfy the constraint score zero; answers that satisfy the constraint but are mediocre score low; answers that satisfy the constraint and are surprising score high.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

GenAI システムを評価するための AI 支援システム化

生成 AI (GenAI) システムの評価は、評価の対象の多くが「推論」、「公平性」、「創造性」などの広範で議論のある概念であるため、困難です。これらの概念が曖昧なままだと、何を測定すればよいのか、評価結果をどのように解釈すればよいのかが不明確になってしまいます。この問題は、体系化、つまり広範な背景概念から測定可能な用語での概念の明示的で構造化された説明への移行というステップが欠如していることを反映しています。体系化は認知能力を必要とし、リソースを大量に消費するという事実に対処するために、AI 支援がこのプロセスをサポートできるかどうかを調査します。 AI 支援による体系化を可能にし、その品質を評価するために、体系化されたコンセプト、コンセプト仕様、および検証ワークシートの構造化表現を導入します。次に、AI 支援による 2 つのシステム化プログラムを開発します。1 つは直接的なゼロショット アプローチ、もう 1 つは既存の文献からの手動による体系化アプローチをより厳密に反映したマルチエージェント アプローチです。私たちはこれらのシステマタイザーを使用して、憎しみに基づくレトリックとデジタル共感という 2 つのコンセプトのコンセプト仕様を作成し、コンテンツの有効性と情報の回復可能性に関して結果として得られるコンセプト仕様を評価します。

原文 (English)

AI-Assisted Systematization for Evaluating GenAI Systems

Evaluating generative AI (GenAI) systems is challenging because many targets of evaluation are broad, contested concepts, such as "reasoning," "fairness," or "creativity." When these concepts are left underspecified, it becomes unclear what should be measured or how evaluation results should be interpreted. This problem reflects a missing step: systematization, that is, moving from a broad background concept to an explicit, structured account of the concept in measurable terms. To help address the fact that systematization is cognitively demanding and resource-intensive, we investigate whether AI assistance can support this process. To enable AI-assisted systematization and assess its quality, we introduce a structured representation of a systematized concept, a concept spec, and a validation worksheet. We then develop two AI-assisted systematizers: a direct, zero-shot approach and a multi-agent approach that more closely mirrors manual systematization approaches from existing literature. We use these systematizers to produce concept specs for two concepts -- hate-based rhetoric and digital empathy -- and evaluate resulting concept specs on content validity and information recoverability.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

勾配が衝突するとき: LLM ジャッジのための多目的プロンプト最適化の失敗モード

特定のタスクまたはドメインに合わせて LLM ジャッジをカスタマイズするには、多くの場合、複数の評価基準にわたってそのプロンプトを同時に最適化する必要があります。テキスト勾配法は単一の審査基準に対してこれを自動化しますが、数値ベクトルではなく自然言語による批評を生成します。したがって、マルチタスク学習の競合解決ツールキット (PCGrad、MGDA) は、多目的テキスト グラデーション設定には適用されません。損失、勾配、およびオプティマイザー LLM が共有するクロスタスク情報の量を変化させることにより、テキスト勾配オプティマイザーの 5 つの分解モードをテストします。 10 件中 6 件の構成では、最初のプロンプトよりも最適化が改善されないことが観察されています。グラジエント LLM が複数の基準を一緒に処理する場合、グラジエント特異性は 59% 低下します (9.0 から 3.7)。これとは別に、タスクごとの指示を単一のプロンプトに単純に組み合わせると、Spearman の rho が -5.3% 低下することがわかります。これらの結果は、最適化時の勾配希釈と推論時の命令干渉という 2 つの分離可能な故障モードを特定します。これらは共に、テキスト フィードバックを使用した多目的判定のカスタマイズの設計空間を制約します。

原文 (English)

When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges

Customizing an LLM judge to a specific task or domain often involves optimizing its prompt across multiple evaluation criteria simultaneously. Textual gradient methods automate this for a single judge criterion, however they produce natural-language critiques, not numerical vectors. Thus, the conflict-resolution toolkit of multi-task learning (PCGrad, MGDA) doesn't apply to the multi-objective textual gradient setting. We test five decomposition modes of textual gradient optimizers by varying how much cross-task information the loss, gradient and optimizer LLMs share. In 6 of 10 configurations, we observe that optimization never improves over the initial prompt. Gradient specificity drops by 59% (from 9.0 to 3.7) when the gradient LLM processes multiple criteria jointly. Separately, we observe that naively combining per-task instructions into a single prompt degrades Spearman's rho by -5.3%. These results identify two separable failure modes: optimization-time gradient dilution and inference-time instruction interference, which together constrain the design space for multi-objective judge customization using textual feedback.

2026-05-26 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

最終的な答えを超えて: ツール拡張エージェントの推論軌跡の評価

最近のツールで強化されたベンチマークには複雑なリクエストが含まれていますが、評価は依然として回答の照合に限定されており、効率、幻覚、適応性などの重要な軌道の側面は無視されています。最も簡単な評価方法は、エージェントの軌跡をグラウンドトゥルースと比較することですが、すべての有効なグラウンドトゥルースの軌跡に注釈を付けるには、法外なコストがかかります。このようにして、ツール拡張 LLM の多次元評価のための参照不要のフレームワークである TRACE を紹介します。 TRACE は、前のステップからの知識を蓄積する証拠バンクを組み込むことにより、エージェントの推論の軌跡を効果的に評価します。私たちのフレームワークを検証するために、多様で欠陥のある軌跡を含む新しいメタ評価データセットを開発し、それぞれに多面的なパフォーマンス スコアが付けられます。私たちの結果は、TRACE が小規模なオープンソース LLM であっても複雑な軌跡を正確に評価することを裏付けています。さらに、私たちの方法を適用して、ツールで拡張されたタスクを解決するときにエージェントが生成する軌跡を評価し、これまで報告されていない観察とそれに対応する洞察を提示します。

原文 (English)

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory aspects like efficiency, hallucination, and adaptivity. The most straightforward method for evaluation is to compare an agent's trajectory with the ground-truth, but annotating all valid ground-truth trajectories is prohibitively expensive. In this manner, we introduce TRACE, a reference-free framework for the multi-dimensional evaluation of tool-augmented LLMs. By incorporating an evidence bank which accumulates knowledge from preceding steps, TRACE assesses an agent's reasoning trajectory effectively. To validate our framework, we develop a new meta-evaluation dataset with diverse and flawed trajectories, each labeled with multi-faceted performance scores. Our results confirm that TRACE accurately evaluates complex trajectories even with small open-source LLMs. Furthermore, we apply our method to evaluate the trajectories that agents produce while solving tool-augmented tasks, presenting previously unreported observations and their corresponding insights.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

迅速な最適化から多次元の信頼性評価まで: 中国の LLM で生成された肝臓 MRI レポートの信頼性を強化 -- 肺がんにも暫定的に拡大

大規模言語モデル (LLM) は、画像所見から診断結論を導き出す際に有望なパフォーマンスを示しており、それによって放射線医学のレポート、研修医の教育、および品質管理をサポートしています。しかし、さまざまな臨床状況にわたってプロンプトデザインを最適化する方法に関する体系的なガイダンスは、依然として検討されていません。さらに、LLM によって生成された放射線医学レポートの信頼性を評価するための包括的で標準化されたフレームワークはまだ確立されていません。この研究は、多次元信頼性評価 (MDCA) フレームワークを導入し、施設固有の迅速な最適化に関するガイダンスを提供することにより、LLM が生成した肝臓 MRI レポートの信頼性を高めることを目的としています。提案されたフレームワークは、SiliconFlow プラットフォームを使用して、Kimi-K2-Instruct-0905、Qwen3-235B-A22B-Instruct-2507、DeepSeek-V3、ByteDance-Seed-OSS-36B-Instruct などのいくつかの高度な LLM のパフォーマンスを評価および比較するために適用されます。

原文 (English)

From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports -- with Preliminary Extension to Lung Cancer

Large language models (LLMs) have demonstrated promising performance in generating diagnostic conclusions from imaging findings, thereby supporting radiology reporting, trainee education, and quality control. However, systematic guidance on how to optimize prompt design across different clinical contexts remains underexplored. Moreover, a comprehensive and standardized framework for assessing the trustworthiness of LLM-generated radiology reports is yet to be established. This study aims to enhance the trustworthiness of LLM-generated liver MRI reports by introducing a Multi-Dimensional Credibility Assessment (MDCA) framework and providing guidance on institution-specific prompt optimization. The proposed framework is applied to evaluate and compare the performance of several advanced LLMs, including Kimi-K2-Instruct-0905, Qwen3-235B-A22B-Instruct-2507, DeepSeek-V3, and ByteDance-Seed-OSS-36B-Instruct, using the SiliconFlow platform.

2026-05-26 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達研究/論文

Deep Research エージェントが失敗するのはなぜですか?研究の全軌跡における幻覚評価について

Deep Research Agent (DRA) の障害パターンを診断することは、依然として重要な課題です。既存のベンチマークは主にエンドツーエンドの評価に依存しており、研究の軌跡全体で蓄積される中間の幻覚を覆い隠しています。このギャップを埋めるために、計画、検索、要約の完全な軌跡で幻覚を監査することにより、結果ベースの評価からプロセスを意識した評価への移行を提案します。ここでは、DRA 幻覚を 4 つの相補的なタイプ (伝播、意図、ノイズ誘導、グラウンディング) に分類する PING 分類法を紹介します。さらに、この分類法を詳細な評価フレームワークにインスタンス化し、厳密な検証のためにトラジェクトリをアトミックなアクション、クレーム、およびサブクエリに分解します。このフレームワークを活用して、敵対的なシナリオを含む 100 の明らかに幻覚を起こしやすいタスクを分離し、DeepHalluBench をキュレートします。 6 つの代表的な DRA での実験では、幻覚を起こしやすいストレス テスト セットでは、評価されたすべてのシステムで無視できない信頼性のギャップが依然として存在することが示されています。さらに、当社の診断分析では、これらの障害が全身の欠陥、特に幻覚の伝播と認知バイアスに起因することを追跡し、将来のアーキテクチャの最適化に実用的な洞察を提供します。コードとデータは https://github.com/yuhao-zhan/DeepHalluBench で入手できます。

原文 (English)

Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evaluation, obscuring intermediate hallucinations that accumulate throughout the research trajectory. To bridge this gap, we propose a shift from outcome-based to processaware evaluation by auditing hallucinations in the full plan-search-summarize trajectory. We introduce the PING Taxonomy, which categorizes DRA hallucinations into four complementary types: Propagation, Intent, Noiseinduced, and Grounding. We further instantiate this taxonomy into a fine-grained evaluation framework that decomposes trajectories into atomic actions, claims, and sub-queries for rigorous verification. Leveraging this framework to isolate 100 distinctively hallucinationprone tasks including adversarial scenarios, we curate DeepHalluBench. Experiments on six representative DRAs show that, on our hallucination-prone stress-test set, all evaluated systems still exhibit non-negligible reliability gaps. Furthermore, our diagnostic analysis traces these failures to systemic deficits, especially hallucination propagation and cognitive biases, providing actionable insights for future architectural optimization. Code and data are available in https://github.com/yuhao-zhan/DeepHalluBench.

2026-05-26 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

OASES: エージェント検索のための結果に合わせた検索評価の共同トレーニング

エージェント検索により、言語モデルは複数のステップにわたって外部証拠を適応的に取得することで、知識集約型タスクを解決できます。検証可能な報酬を伴う強化学習(RLVR)は、検索エージェント向けのトレーニング パラダイムとして広く採用されていますが、結果のみの報酬はまばらで、中間の検索アクションに対する単位の割り当ては限られています。したがって、既存のプロセス報酬手法は、プロキシ信号、外部評価者、または可能性ベースの情報ゲインを通じて監視を強化しようとしています。ただし、代理報酬は最終的な結果の目標から逸脱する可能性があり、固定の評価者は検索ポリシーが進化するにつれて古くなり、信頼性の低いプロセス監視につながる可能性があります。これらの課題に対処するために、私たちは、エージェント検索のための結果に合わせた検索評価監視フレームワークである OASES を提案します。 OASES は、各中間検索状態が元の質問への回答をどの程度サポートしているかを評価することで、結果に合わせたプロセス報酬を導き出します。さらに、検索ポリシーとポリシーに関する状態評価者を共同トレーニングすることで、評価者が進化する検索動作に適応し、より信頼性の高いプロセス報酬を提供できるようになります。 5 つのマルチホップ QA ベンチマークの実験では、OASES が一貫して強力な RL ベースラインを上回っていることが示されており、さらなる分析により、結果に合わせたプロセス報酬と検索評価の共同トレーニングの利点が確認されています。

原文 (English)

OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search

Agentic search enables language models to solve knowledge-intensive tasks by adaptively acquiring external evidence over multiple steps. Reinforcement learning with verifiable rewards (RLVR) has emerged as a widely adopted training paradigm for search agents, yet outcome-only rewards are sparse and provide limited credit assignment for intermediate search actions. Existing process-reward methods therefore seek to densify supervision through proxy signals, external evaluators, or likelihood-based information gain. However, proxy rewards can deviate from the final outcome objective, while fixed evaluators can become stale as the search policy evolves, leading to unreliable process supervision. To address these challenges, we propose OASES, an Outcome-Aligned Search-Evaluation Supervision framework for agentic search. OASES derives outcome-aligned process rewards by evaluating how well each intermediate search state supports answering the original question. It further co-trains the search policy and the state evaluator on policy, allowing the evaluator to adapt to evolving search behavior and provide more reliable process rewards. Experiments on five multi-hop QA benchmarks show that OASES consistently outperforms strong RL baselines, with further analyses confirming the benefits of outcome-aligned process rewards and search-evaluation co-training.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

UniToolCall: LLM エージェントのツール使用表現、データ、評価の統合

ツール使用機能は LLM エージェントの基本コンポーネントであり、構造化された関数呼び出しを通じて外部システムと対話できるようにします。しかし、既存の研究は一貫性のない相互作用表現を示し、ツール使用軌跡の構造的分布を大部分見落としており、互換性のない評価ベンチマークに依存しています。ツールセットの構築とデータセットの生成から評価までのパイプライン全体を標準化するツール学習用の統合フレームワークである UniToolCall を紹介します。このフレームワークは、22,000 以上のツールからなる大規模なツール プールを厳選し、10 個の標準化された公開データセットと構造的に制御された合成軌跡を組み合わせることで、390,000 以上のインスタンスのハイブリッド トレーニング コーパスを構築します。シングルホップとマルチホップ、シングルターンとマルチターンなどの多様な対話パターンを明示的にモデル化し、シリアルとパラレルの両方の実行構造をキャプチャします。一貫したマルチターン推論をサポートするために、クロスターン依存関係を強制するアンカー リンケージ メカニズムをさらに導入します。さらに、7 つの公開ベンチマークを、関数呼び出し、ターン、会話レベルでのきめ細かい評価を備えた統合されたクエリ-アクション-観察-回答 (QAOA) 表現に変換します。実験の結果、データセットで Qwen3-8B を微調整すると、ツールの使用パフォーマンスが大幅に向上することがわかりました。ディストラクターの多いハイブリッド 20 設定では、93.0% のシングルターン厳密精度を達成し、GPT、Gemini、Claude などの商用モデルを上回ります。

原文 (English)

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However, existing research exhibits inconsistent interaction representations, largely overlooks the structural distribution of tool-use trajectories, and relies on incompatible evaluation benchmarks. We present UniToolCall, a unified framework for tool learning that standardizes the entire pipeline from toolset construction and dataset generation to evaluation. The framework curates a large tool pool of 22k+ tools and constructs a hybrid training corpus of 390k+ instances by combining 10 standardized public datasets with structurally controlled synthetic trajectories. It explicitly models diverse interaction patterns, including single-hop vs. multi-hop and single-turn vs. multi-turn, while capturing both serial and parallel execution structures. To support coherent multi-turn reasoning, we further introduce an Anchor Linkage mechanism that enforces cross-turn dependencies. Furthermore, we convert 7 public benchmarks into a unified Query--Action--Observation--Answer (QAOA) representation with fine-grained evaluation at the function-call, turn, and conversation levels. Experiments show that fine-tuning Qwen3-8B on our dataset substantially improves tool-use performance. Under the distractor-heavy Hybrid-20 setting, achieves 93.0% single-turn Strict Precision, outperforming commercial models including GPT, Gemini, and Claude.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

ECUAS$_n$: 不確実性が増大したシステムの原則に基づいた評価のための指標ファミリー

一か八かの自動化された意思決定では、ユーザー (人間または下流システム) がアプリケーション固有のコストのトレードオフに基づいて予測を受け入れるか拒否できるようにするために、予測の不確実性へのアクセスが不可欠です。このような不確実性拡張 (UA) システム、つまり、予測と不確実性スコアの両方を出力するシステムは、現在文献でさまざまな方法で評価されています。予測と不確実性スコアを評価するために別個のメトリクスを使用したり、固定の拒否コストでコスト関数を設定したり、カバレッジリスク曲線上で統合したりします。我々は、これらの評価アプローチは、不確実性の下での意思決定のための UA システムの全体的なパフォーマンスを評価するには不十分であると主張し、対象のタスクの適切なスコアリング ルールとして定式化された新しいメトリクス ファミリ ECUAS$_n$ を提案します。パラメーター $n$ は、ユースケースのニーズに応じて、不正確な予測のコストと不完全な不確実性の間のトレードオフを制御します。私たちは、手動でアノテーションを付けた TriviaQA のサブセットを含む、さまざまな分類および生成データセットの実験を通じて、ECUAS$_n$ メトリクスの利点を理論的および経験的に実証します。

原文 (English)

ECUAS$_n$: A family of metrics for principled evaluation of uncertainty-augmented systems

In high-stakes automated decision-making, access to predictive uncertainty is essential for enabling users -- human or downstream systems -- to accept or reject predictions based on application-specific cost trade-offs. Such uncertainty-augmented (UA) systems -- i.e., systems that output both predictions and uncertainty scores -- are currently being assessed in the literature in a variety of ways, using separate metrics to evaluate the predictions and the uncertainty scores, setting a cost function with a fixed rejection cost or integrating over a coverage-risk curve. We argue that these evaluation approaches are inadequate for assessing overall performance of the UA system for decision making under uncertainty and propose a novel family of metrics, ECUAS$_n$, formulated as proper scoring rules for the task of interest. The parameter $n$ controls the trade-off between the cost of incorrect predictions and imperfect uncertainties depending on the needs of the use-case. We demonstrate the advantages of the ECUAS$_n$ metrics both theoretically and empirically, through experiments on diverse classification and generation datasets, including a manually annotated subset of TriviaQA.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

マイニングのためのスマートなタイミング: ビットコイン ハードウェア ROI 予測のための深層学習フレームワーク

不安定な市場、急速な技術の陳腐化、プロトコル主導の収益サイクルのため、ビットコイン マイニング ハードウェアの取得には戦略的なタイミングが必要です。マイニングは資本集約型産業へと進化しているにもかかわらず、新しい特定用途向け集積回路 (ASIC) ハードウェアをいつ購入するかについての指針はほとんどなく、この決定問題に対処する従来の計算フレームワークもありません。私たちは、ハードウェアの購入を時系列分類タスクとして定式化し、ASIC マシンの購入が 1 年以内に利益をもたらす (投資収益率 (ROI) >= 1)、限界のある (0 < ROI < 1)、または不採算 (ROI <= 0) の利益を生み出すかを予測することで、このギャップに対処します。私たちは、マイニングの収益性におけるマルチスケールの時間的パターンを捉えるように設計されたオープンソースの Transformer ベースのアーキテクチャである MineROI-Net を提案します。 2015 年から 2024 年の間にさまざまな市場体制にわたってリリースされた 20 の ASIC マイナーからのデータに基づいて評価された MineROI-Net は、リカレントベースライン、畳み込みベースライン、およびアテンションベースのベースラインを上回り、83.2% の精度と 83.5% のマクロ F1 スコアを達成しました。このモデルは経済的関連性が高く、不採算期間の検出精度は 97.8%、収益性の高い期間の検出精度は 81.5% を達成し、収益性の高いシナリオを不採算として誤って分類したり、その逆を回避したりできます。これらの結果は、MineROI-Net がマイニング ハードウェアの取得のタイミングを決定するための実用的なデータ駆動型ツールを提供し、資本集約的なマイニング操作における財務リスクを潜在的に軽減することを示しています。

原文 (English)

Smart Timing for Mining: A Deep Learning Framework for Bitcoin Hardware ROI Prediction

Bitcoin mining hardware acquisition requires strategic timing due to volatile markets, rapid technological obsolescence, and protocol-driven revenue cycles. Despite mining's evolution into a capital-intensive industry, there is little guidance on when to purchase new Application-Specific Integrated Circuit (ASIC) hardware, and no prior computational frameworks address this decision problem. We address this gap by formulating hardware acquisition as a time series classification task, predicting whether purchasing ASIC machines yields profitable (Return on Investment (ROI) >= 1), marginal (0 < ROI < 1), or unprofitable (ROI <= 0) returns within one year. We propose MineROI-Net, an open-source Transformer-based architecture designed to capture multi-scale temporal patterns in mining profitability. Evaluated on data from 20 ASIC miners released between 2015 and 2024 across diverse market regimes, MineROI-Net outperforms recurrent, convolutional, and attention-based baselines, achieving 83.2% accuracy and 83.5% macro F1-score. The model demonstrates strong economic relevance, achieving 97.8% precision in detecting unprofitable periods and 81.5% precision in detecting profitable ones, while avoiding misclassifying profitable scenarios as unprofitable and vice versa. These results indicate that MineROI-Net offers a practical, data-driven tool for timing mining hardware acquisitions, potentially reducing financial risk in capital-intensive mining operations.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

RCT とヒューマン アップリフト研究: フロンティア AI 評価のための方法論的課題と実践的な解決策

人間向上研究、つまりランダム化比較試験 (RCT) または同様の方法論を通じて人間のパフォーマンスに対する AI アクセスの影響を測定する研究は、最前線の AI ガバナンスと展開の決定にますます多くの情報を提供します。 RCT 手法は他の分野では堅牢ですが、フロンティア AI システムの特有の特性との相互作用は、特に結果が一か八かの意思決定に使用される場合にはまだ十分に検討されていません。バイオセキュリティ、サイバーセキュリティ、教育、労働などの分野で人間向上の研究を行った経験を持つ16人の専門家へのインタビューから得た結果を紹介します。専門家らはインタビューを通じて、人間高揚研究が依存する標準的な因果推論の仮定と研究の対象そのものとの間に繰り返される緊張を説明した。急速に進化する AI システム、変化するベースライン、異種混合で変化するユーザーの習熟度、多孔質な現実世界の設定により、内部、外部、構成の妥当性の基礎となる仮定が歪められ、上昇証拠の解釈と適切な使用が複雑になっています。私たちは、(1) 妥当性を研究するためのリスクにマッピングされ、大規模言語モデル (LLM) システムへの特異性の度合いによって分類された人間高揚研究における方法論的課題の統合、および (2) 課題から提案された解決策へのマッピングに貢献します。専門家が特定した課題と解決策を照合することで、人間向上の証拠の解釈限界と適切な使用法を明確にし、評価の実践とそれがもたらす決定を調整し、AI ガバナンスのより調整された方法論の基盤をサポートすることを目指しています。

原文 (English)

RCTs & Human Uplift Studies: Methodological Challenges and Practical Solutions for Frontier AI Evaluation

Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform frontier AI governance and deployment decisions. While RCT methods are robust in other fields, their interaction with the distinctive properties of frontier AI systems remains underexamined, particularly when results are used to inform high-stakes decisions. We present findings from interviews with 16 expert practitioners with experience conducting human uplift studies in domains including biosecurity, cybersecurity, education, and labor. Across interviews, experts described a recurring tension between the standard causal inference assumptions upon which human uplift studies rely and the object of study itself. Rapidly evolving AI systems, shifting baselines, heterogeneous and changing user proficiency, and porous real-world settings strain assumptions underlying internal, external, and construct validity, complicating the interpretation and appropriate use of uplift evidence. We contribute (1) a synthesis of methodological challenges in human uplift studies, mapped to risks to study validity and classified by their degree of specificity to large language model (LLM) systems, and (2) a mapping from challenges to proposed solutions. By collating expert-identified challenges and solutions, we seek to clarify the interpretive limits and appropriate uses of human uplift evidence, to align evaluation practice with the decisions it informs, and to support more coordinated methodological foundations for AI governance.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達

人間による注釈は必要ですか?機械翻訳におけるエラー スパン検出のための反復 MBR 蒸留

エラー スパン検出 (ESD) は、機械翻訳 (MT) 評価における重要なサブタスクであり、翻訳エラーの場所と重大度を特定することを目的としています。人間が注釈を付けたデータのモデルを微調整すると ESD パフォーマンスが向上しますが、そのようなデータの取得にはコストがかかり、アノテーター間で不一致が発生しやすくなります。これに対処するために、私たちは、最小ベイズ リスク (MBR) デコードに基づく新しい自己進化フレームワークを提案します。このフレームワークは、ESD のための反復 MBR 蒸留と呼ばれます。これは、既製の LLM を利用して疑似ラベルを生成することにより、人間による注釈への依存を排除​​します。 WMT メトリクス共有タスク データセットに関する広範な実験により、これらの自己生成疑似ラベルのみでトレーニングされたモデルは、非適応ベース モデルと、システム レベルおよびスパン レベルでヒューマン アノテーションでトレーニングされた教師ありベースラインの両方を上回るパフォーマンスを示し、同時に、競争力のある文レベルのパフォーマンスを維持できることが実証されました。

原文 (English)

Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Machine Translation

Error Span Detection (ESD) is a crucial subtask in Machine Translation (MT) evaluation, aiming to identify the location and severity of translation errors. While fine-tuning models on human-annotated data improves ESD performance, acquiring such data is expensive and prone to inconsistencies among annotators. To address this, we propose a novel self-evolution framework based on Minimum Bayes Risk (MBR) decoding, named Iterative MBR Distillation for ESD, which eliminates the reliance on human annotations by leveraging an off-the-shelf LLM to generate pseudo-labels. Extensive experiments on the WMT Metrics Shared Task datasets demonstrate that models trained solely on these self-generated pseudo-labels outperform both unadapted base model and supervised baselines trained on human annotations at the system and span levels, while maintaining competitive sentence-level performance.

2026-05-26 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

リアル vs. セミシミュレーション: 治療効果推定のための評価の再考

機械学習による不均一な治療効果の推定は、学術研究と産業実践の両方で大きな注目を集めています。ただし、2 つのコミュニティは、著しく異なる条件下でモデルを評価することがよくあります。方法論的な作業は通常、反事実の結果を必要とする半シミュレートされたベンチマークとメトリクスに依存しますが、現実世界のアプリケーションはランキングやテスト結果に基づいた観察可能なメトリクスに依存します。方法論の進歩と実際の展開との間にギャップがあることはよく知られているにもかかわらず、これらの評価体制間の関係は体系的に検討されていません。私たちは、標準的な半シミュレートされたベンチマークファミリーと現実世界のデータセットにわたる治療効果評価の大規模な実証研究を実施します。私たちのベンチマークは、複数のベース学習器とペアになったメタ学習器、および特殊な因果機械学習モデルをカバーしています。私たちは、メソッドの論文で一般的に使用される反事実の指標と並行して、アプリケーション指向の文献で一般的な観察可能な指標を使用して、これらのメソッドを評価します。私たちの結果は、2 つの相補的なギャップを明らかにしました。まず、反事実メトリクスは、たとえ同じ半シミュレートされたベンチマークであっても、観察可能なメトリクスによって優先される推定量を確実に回復することはできません。第 2 に、半シミュレートされたベンチマークで取得されたランキングは実際のデータセットには転送されません。さらに、特殊な因果モデルとは対照的に、強力な基本モデルを備えた単純なメタ学習者は常に競争力があることがわかりました。全体として、我々の調査結果は、治療効果推定研究の進歩は、反事実の指標や半シミュレートされたベンチマークのみによって評価されるべきではなく、観察可能な指標と実際のデータの検証を組み込むことで利益が得られることを示唆しています。

原文 (English)

Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation

Estimating heterogeneous treatment effects with machine learning has attracted substantial attention in both academic research and industrial practice. However, the two communities often evaluate models under markedly different conditions. Methodological work typically relies on semi-simulated benchmarks and metrics that require counterfactual outcomes, whereas real-world applications rely on observable metrics based on ranking or test outcomes. Despite the well-known gap between methodological progress and practical deployment, the relationship between these evaluation regimes has not been examined systematically. We conduct a large-scale empirical study of treatment effect evaluation across standard semi-simulated benchmark families and real-world datasets. Our benchmark covers meta-learners paired with multiple base learners, as well as specialized causal machine learning models. We evaluate these methods using observable metrics common in application-oriented literature, alongside counterfactual metrics commonly used in methods papers. Our results reveal two complementary gaps. First, counterfactual metrics do not reliably recover the estimators preferred by observable metrics, even on the same semi-simulated benchmarks. Second, rankings obtained on semi-simulated benchmarks do not transfer to real datasets. We further find that simple meta-learners with strong base models are consistently competitive, in contrast to specialized causal models. Overall, our findings suggest that progress in treatment effect estimation research should not be assessed solely through counterfactual metrics and semi-simulated benchmarks, but it would benefit from incorporating observable metrics and real-data validation.

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

LLM が推論するのはいつですか?エントロピー相転移による動的システムの視点

Chain-of-thought (CoT) reasoning has become the default strategy for enhancing LLM capabilities, yet its application raises a fundamental question: when is explicit reasoning actually beneficial?経験的証拠は、顕著な矛盾を明らかにしています。CoT は、多くの場合、トークン消費量を増大させながら、事実に基づいた無制限のタスクに対してわずかな利益、またはマイナスの利益さえ提供します。この研究では、LLM 推論がタスクやモデルの静的な特性ではなく、生成中に現れる \emph{動的復号状態} であることを示します。体系的な分析を通じて、初期段階のエントロピー ダイナミクスがこの状態の信頼できるシグナルを提供することを発見しました。CoT の恩恵を受けるタスクは一貫したエントロピーの減少を示しますが、他のタスクは不安定または増加するパターンを示します。この動作は、高エントロピー探索体制から低エントロピー構造推論体制への相転移のような移行として解釈できます。これらの洞察に基づいて、我々は、早期デコードエントロピーを活用して推論戦略を適応的に選択する、軽量でトレーニング不要のルーティングフレームワークである \textbf{EDRM} (エントロピーダイナミクスベースの推論マニホールド) を提案します。 EDRM は、エントロピーの軌跡をコンパクトで解釈可能な多様体表現に埋め込み、ゼロショット デプロイメントときめ細かいインスタンス レベルの適応の両方を可能にします。さまざまなスケールとアーキテクチャの 15 のベンチマークと 4 つの LLM にわたって、EDRM は一貫して静的ベースラインを上回っています。データセット レベルでは、EDRM は \textbf{41--55\%} トークンの削減を達成しながら、わずか 50 個のキャリブレーション サンプルで精度を向上させます。インスタンス レベルでは、\textbf{27--45\%} トークンの節約を維持しながら、精度が最大 \textbf{4.7\%} まで向上します。これらの結果は、推論はデフォルトではなく選択的に呼び出される必要があることを示唆しており、効率的で適応的な LLM 推論に対するエントロピー駆動型の復号制御の有効性を示しています。

原文 (English)

When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions

Chain-of-thought (CoT) reasoning has become the default strategy for enhancing LLM capabilities, yet its application raises a fundamental question: when is explicit reasoning actually beneficial? Empirical evidence reveals a striking paradox: CoT often provides marginal or even negative gains on factual and open-ended tasks while multiplying token consumption. In this work, we show that LLM reasoning is not a static property of tasks or models, but a \emph{dynamic decoding state} that emerges during generation. Through systematic analysis, we find early-stage entropy dynamics provide a reliable signal of this state: tasks benefiting from CoT exhibit consistent entropy reduction, while others display unstable or increasing patterns. This behavior can be interpreted as a phase-transition-like shift from a high-entropy exploratory regime to a low-entropy structured reasoning regime. Based on these insights, we propose \textbf{EDRM} (Entropy Dynamics-based Reasoning Manifold), a lightweight and training-free routing framework that leverages early decoding entropy to adaptively select inference strategies. EDRM embeds entropy trajectories into a compact and interpretable manifold representation, enabling both zero-shot deployment and fine-grained instance-level adaptation. Across 15 benchmarks and 4 LLMs of varying scales and architectures, EDRM consistently outperforms static baselines. At the dataset level, EDRM achieves \textbf{41--55\%} token reduction while improving accuracy with as few as 50 calibration samples. At the instance level, it further improves accuracy by up to \textbf{4.7\%} while maintaining \textbf{27--45\%} token savings. These results suggest that reasoning should be invoked selectively rather than by default, and demonstrate the effectiveness of entropy-driven decoding control for efficient and adaptive LLM inference.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達

ランダムよりも悪い: 教師なし特徴選択のベースラインの重要性

毎年、多くの新しい教師なし特徴選択手法が提案されていますが、その経験的評価は、既存の手法との比較とともに、選択されたデータセットで計算された教師ありおよび教師なしの評価メトリクスに限定されています。ただし、確立された評価ベースラインが存在しない場合、これらの各方法によって既存の文献に付加される価値や、その基礎となるアプローチがどれほど効果的であるかを判断することは困難です。教師なし特徴選択方法を評価するためのベースラインとしてランダム特徴選択を使用することを提案します。私たちは、教師なし特徴選択における最先端の手法の多くが、パフォーマンスと効率の両方においてランダム特徴選択よりも優れていることを経験的に示しています。したがって、ランダムな特徴選択よりも一貫した改善を確実にするために、新しい教師なし特徴選択方法の開発プロセスのベースラインとしてランダムな特徴選択を考慮するという厳格な要件を強調します。

原文 (English)

Worse than Random: The Importance of a Baseline for Unsupervised Feature Selection

Many novel unsupervised feature selection methods are proposed each year, yet their empirical evaluation is limited to supervised and unsupervised evaluation metrics computed on selected datasets, along with comparisons to existing methods. However, in the absence of an established evaluation baseline, it is difficult to determine the value added to the existing literature by each of these methods, and how effective their underlying approaches are. We propose using random feature selection as a baseline for evaluating the unsupervised feature selection methods. We empirically show that many of the state-of-the-art methods in unsupervised feature selection are outperformed by random feature selection in both performance and efficiency. Accordingly, we emphasize on the strict requirement of considering random feature selection as a baseline in the development process of novel unsupervised feature selection methods to ensure a consistent improvement over random feature selection.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

評価意識の分解と測定

フロンティア言語モデルは、評価されていることを認識して動作を調整し、ベンチマーク結果の妥当性を損なうことがあります。しかし、現場では共通の基礎を持たずに評価の特性とモデルの特性、検出と行動反応を混同して研究が行われています。私たちは評価意識を社会心理学に基礎づけ、評価意識を環境要素 (課題がどの程度認識されているか) と、認識をそれに基づいて行動する傾向から分離するモデル要素に分解します。プレースホルダー エンティティや採点スタイルの出力形式など、8 つの分類されたトリガー要因を通じて環境コンポーネントを運用し、思考連鎖のモニタリングを通じて認識と行動を研究します。 9 つのフロンティア モデルと 4 つのベンチマークにわたって、認識率はモデルとベンチマークのどちらか単独ではなく、モデルとベンチマークの特定の組み合わせに依存します。認識が行動の変化につながることはほとんどありませんが、変化する場合、その方向性は認識された評価の種類によって異なります。また、モデルは機能評価よりも安全性に対して敏感であり、安全性ベンチマークの妥当性がより大きなリスクにさらされます。各モデルがどの要因に敏感で、それらがどのように相互作用するかを研究するために、8 つの要因のそれぞれを独立して切り替えることができ、基礎となる要求を固定したまま評価信号を変化させる、100 のペアの安全機能タスクの要因制御ベンチマークである \textbf{EvalAwareBench} を提案します。 EvalAwareBench を通じて、単一の要素がすべてのモデルに均一に影響を与えることはなく、要素を積み重ねることですべてのモデルにわたる評価の意識が徐々に向上することがわかりました。私たちのフレームワークと EvalAwareBench は、評価意識を測定、属性付け、軽減するためのツールを提供し、将来有望な道として認識される下での行動の一貫性を示します。

原文 (English)

Decomposing and Measuring Evaluation Awareness

Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet the field studies it without a shared foundation, conflating properties of the evaluation with properties of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component (how recognizable the task is) and a model component that separates recognition from propensity to act on it. We operationalize the environment component through eight categorized trigger factors, such as placeholder entities and grading-style output formats, and study recognition and behavior through chain-of-thought monitoring. Across nine frontier models and four benchmarks, recognition rates depend on the specific pairing of model and benchmark rather than on either in isolation. Recognition rarely leads to behavioral change, and when it does, the direction depends on the type of evaluation perceived. Models are also more sensitive to safety than capability evaluations, placing safety benchmark validity at greater risk. To study which factors each model is sensitive to and how they interact, we propose \textbf{EvalAwareBench}, a factor-controlled benchmark of 100 paired safety-capability tasks where each of the eight factors can be independently toggled, varying evaluative signals while holding the underlying request fixed. Through EvalAwareBench, we find that no single factor uniformly affects all models, but stacking factors progressively raises evaluation awareness across all of them. Our framework and EvalAwareBench provide the tools to measure, attribute, and mitigate evaluation awareness, pointing to behavioral consistency under recognition as a promising path forward.

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達研究/論文

ロングコンテキスト LLM の位置の失敗: 推論ベンチマークの盲点

位置制御された評価は、Needle-in-a-Haystack や RULER などの検索タスクの標準ですが、主流の推論ベンチマークは、長いコンテキストでのターゲット タスクの位置配置を制御しません。 11 個の長いコンテキストのベンチマークを監査したところ、タスクの位置、フィラーの内容、および推論のためのコンテキストの長さを共同で制御するものはありませんでした。 4 つの主力ロングコンテキスト リリースの監査では、NIAH、RULER、または LongBench ファミリー ベンチマークのメイン結果テーブル エントリは見つかりませんでしたが、エージェント ベンチマークとコーディング ベンチマークは 4 つすべてのメイン結果テーブルに表示されます。私たちは、3 つの要素すべてを変化させる制御されたフレームワークであるコンテキスト ロット評価 (CRE) を提案し、GSM8K と ARC-Challenge の 9 つの LLM を 2 つのラウンド (初期 5 モデル セットと 4 つの新しいベンダー リリース) にわたって評価します。ターゲット タスクが端から中間に移動するとモデルが急激に低下する可能性があり、脆弱なモデルのコンテキストの長さが増すにつれて低下はさらに悪化します。 MiMo-v2-Flash は、with_solutions フィラーの下で 64K で 88pp 低下します (中精度 8%)。新しいリリースでは低下が小さくなっています。64K では、4 つのうち 3 つが終了位置精度の +/-6pp 以内に留まっています。 MiMo-V2.5-Pro は、MiMo-v2-Flash の 88pp の低下を 32pp に狭めます。 question_only_v2 フィラーでは、4 つすべてで中間位置の低下が持続します (8K、32K、64K で -16pp から -56pp の範囲)。 8K では、最後にターゲット タスクのコピーを追加する診断プローブにより、9 つのモデルすべてで終了ベースラインの +/-4pp 以内の中程度の精度が得られ、位置の説明と一致します。最初の 5 つのモデル セットでは、中間位置のエラーの 76% が周囲のフィラー テキストと一致するのに対し、終了位置では 22% であり、主要なエラー モードとしてのフィラーと回答の干渉と一致しています。これらの結果は、現在の推論ベンチマーク設計とベンダー評価実践における構造的な評価のギャップを明らかにしています。タスクの位置が制御されていない場合、コンテキストの長さとともに増大する位置の脆弱性は測定できません。

原文 (English)

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

Position-controlled evaluation is standard for retrieval tasks such as Needle-in-a-Haystack and RULER, but mainstream reasoning benchmarks do not control positional placement of target tasks in long contexts. We audit 11 long-context benchmarks and find none jointly controls task position, filler content, and context length for reasoning. An audit of four flagship long-context releases finds no main result-table entry for NIAH, RULER, or LongBench-family benchmarks, while agentic and coding benchmarks appear in main result-tables across all four. We propose Context Rot Evaluation (CRE), a controlled framework varying all three factors, and evaluate nine LLMs on GSM8K and ARC-Challenge across two rounds: an initial five-model set and four newer vendor releases. Models can drop sharply when the target task moves from end to middle, and the drop grows worse with context length for vulnerable models. MiMo-v2-Flash drops 88pp at 64K under with_solutions filler (middle accuracy 8%). Newer releases show smaller drops: at 64K, three of four stay within +/-6pp of end-position accuracy; MiMo-V2.5-Pro narrows the MiMo-v2-Flash 88pp drop to 32pp. Under questions_only_v2 filler, middle-position drops persist across all four (range -16pp to -56pp across 8K, 32K, 64K). At 8K, a diagnostic probe adding a target-task copy at the end brings middle accuracy within +/-4pp of end baseline across all nine models, consistent with a positional explanation. In the initial five-model set, 76% of middle-position errors match surrounding filler text versus 22% at the end position, consistent with filler-answer interference as a dominant error mode. These results expose a structural evaluation gap in current reasoning benchmark design and vendor evaluation practice: positional vulnerabilities that grow with context length cannot be measured when task position is not controlled.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達

メタ学習による費用対効果の高いモデル評価

機械学習の急速な成長により、拡大し続けるモデルのエコシステムが生み出され、目に見えないラベルのないデータに対して新しくリリースされたモデルの信頼性を検証することがますます困難になっています。従来の評価パイプラインは、高価なアノテーション、繰り返しの微調整、またはモデル ファミリ間での転送ができない狭い仮定に依存しています。さまざまなアーキテクチャやモダリティにまたがる未確認のモデルをラベルなしで迅速に評価するための、コスト効率が高く、モデルに依存しないフレームワークである MetaEvaluator を紹介します。 MetaEvaluator は、参照モデルのプールに対するメタ学習を利用して転送可能な初期化を取得し、プール全体でコストを償却しながら、モデルごとの再トレーニングの必要性を排除しながら、新しいモデルの正確な評価を可能にします。私たちの知る限り、これは完全にラベルのないデータセットで新しいモデルを評価できる、モデルに依存しない最初のフレームワークです。広範な実験により、MetaEvaluator は従来のアプローチと比較して大幅にコストを削減しながら安定した正確なパフォーマンス推定値を生成し、ラベルのないデータに対する新しいモデルのスケーラブルなベンチマークを実用化できることが示されています。

原文 (English)

Cost-Effective Model Evaluation with Meta-Learning

The rapid growth of machine learning has produced an ever-expanding ecosystem of models, making it increasingly challenging to verify the reliability of newly released models on unseen, unlabeled data. Conventional evaluation pipelines depend on expensive annotation, repeated fine-tuning, or narrow assumptions that fail to transfer across model families. We present MetaEvaluator, a cost-effective, model-agnostic framework for rapid, label-free assessment of unseen models spanning diverse architectures and modalities. MetaEvaluator leverages meta-learning over a pool of reference models to obtain a transferable initialization, enabling accurate evaluation of new models while amortizing cost across the pool and removing the need for per-model retraining. To the best of our knowledge, this is the first model-agnostic framework capable of evaluating new models on entirely unlabeled datasets. Extensive experiments show that MetaEvaluator produces stable and accurate performance estimates at substantially reduced cost compared to conventional approaches, making scalable benchmarking of emerging models on unlabeled data practical.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達

時間的概念ドリフトの下での敵対的脆弱性: Android マルウェア検出の縦断的研究

エミュレータと実際のデバイスの実行から抽出された静的および動的特徴表現を使用して、10 年以上の Android アプリケーションにわたる敵対的堅牢性の長期的なドリフトを意識した評価を示します。データセットは年ごとのスライスに編成され、現実的な学習シナリオをエミュレートする 3 つの導入プロトコルに基づいて評価されます。(1) 同年のトレーニングとテスト、(2) モデルの更新を行わない年度をまたぐ導入、(3) 累積的な履歴データによるウィンドウの拡大再トレーニング。複数の分類子ファミリーにわたって、実現可能性の制約の下で FGSM と SPSA を使用して敵対的な例が生成されます。クリーン パフォーマンス、敵対的精度 (AA)、攻撃成功率 (ASR) を測定し、時間的リンケージ メトリック (RobustDrop、$\Delta$ASR、敵対的増幅率 (AAF)) を導入して、分布シフトとロバスト性低下の関係を定量化します。結果は、評価された転送ベースの特徴空間設定では、時間的分離が敵対的ロバスト性の低下と関連していることを示しています。トレインとテストのギャップが増加するにつれて、クリーン精度と敵対的精度は低下しますが、攻撃の成功率は、特に FGSM の摂動や静的機能の下では設定に依存して増加します。拡張ウィンドウの再トレーニングは、継続的な分布進化の下でのロバスト性の損失を軽減しますが、排除するわけではありません。これらの発見は、進化するデータ分布の下でインテリジェント検出システムの長期的な堅牢性を評価する際には時間的ドリフトを考慮する必要があることを示し、長期にわたる敵対環境におけるドリフトを意識した堅牢性評価フレームワークの必要性を強調しています。

原文 (English)

Adversarial Vulnerability Under Temporal Concept Drift: A Longitudinal Study of Android Malware Detection

We present a longitudinal, drift-aware evaluation of adversarial robustness across more than a decade of Android applications using static and dynamic feature representations extracted from emulator and real-device executions. The dataset is organized into yearly slices and evaluated under three deployment protocols that emulate realistic learning scenarios: (1) same-year training and testing, (2) cross-year deployment without model updates, and (3) expanding-window retraining with cumulative historical data. Across multiple classifier families, adversarial examples are generated using FGSM and SPSA under feasibility constraints. We measure clean performance, Adversarial Accuracy (AA), Attack Success Rate (ASR), and introduce temporal linkage metrics -- RobustDrop, $\Delta$ASR, and Adversarial Amplification Factor (AAF) -- to quantify the relationship between distribution shift and robustness degradation.nResults show that temporal separation is associated with reduced adversarial robustness under the evaluated transfer-based feature-space setting. As the train-test gap increases, clean accuracy and adversarial accuracy decline, while attack success exhibits configuration-dependent increases, particularly under FGSM perturbations and static features. Expanding-window retraining mitigates, but does not eliminate, robustness loss under continued distributional evolution. These findings indicate that temporal drift should be considered when assessing the long-term robustness of intelligent detection systems under evolving data distributions and highlight the need for drift-aware robustness assessment frameworks in long-lived adversarial environments.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達

機密性の高いものは忘れて、重要なことを思い出してください: 継続的な学習のためのメモリ スカルプティングにおけるトークン レベルの差分プライバシー

継続学習 (CL) モデルは、逐次的な知識の獲得には優れていますが、多様な情報が蓄積されるため、重大で見落とされがちなプライバシーの課題に直面しています。均一な差分プライバシー (DP) バジェットなどの従来のプライバシー手法は、すべてのデータを無差別に保護するため、モデルのユーティリティの大幅な低下につながり、プライバシーに敏感な領域での CL の展開が妨げられます。これを克服するために、私たちは機密性の高いものを忘れ、重要なことを覚えておくプライバシー強化継続学習 (PeCL) フレームワークを提案します。私たちのアプローチでは、まず、個々のトークンのセマンティックな機密性に基づいてプライバシー予算を適応的に割り当てる、トークンレベルの動的な差分プライバシー戦略を導入します。これにより、機密性のない一般知識へのノイズ注入を最小限に抑えながら、民間エンティティに対する堅牢な保護が保証されます。 2 番目に、プライバシーに基づいたメモリ彫刻モジュールを統合します。このモジュールは、動的 DP メカニズムの感度分析を利用して、モデルのメモリとパラメーターから機密情報をインテリジェントに忘れる一方で、壊滅的な忘却を軽減するために重要なタスク不変の履歴知識を明示的に保存します。広範な実験により、PeCL はプライバシー保護とモデルの実用性の間で優れたバランスを実現し、堅牢なプライバシーを確​​保しながら以前のタスクで高い精度を維持することでベースライン モデルを上回るパフォーマンスを示していることが示されています。

原文 (English)

Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning

Continual Learning (CL) models, while adept at sequential knowledge acquisition, face significant and often overlooked privacy challenges due to accumulating diverse information. Traditional privacy methods, like a uniform Differential Privacy (DP) budget, indiscriminately protect all data, leading to substantial model utility degradation and hindering CL deployment in privacy-sensitive areas. To overcome this, we propose a privacy-enhanced continual learning (PeCL) framework that forgets what's sensitive and remembers what matters. Our approach first introduces a token-level dynamic Differential Privacy strategy that adaptively allocates privacy budgets based on the semantic sensitivity of individual tokens. This ensures robust protection for private entities while minimizing noise injection for non-sensitive, general knowledge. Second, we integrate a privacy-guided memory sculpting module. This module leverages the sensitivity analysis from our dynamic DP mechanism to intelligently forget sensitive information from the model's memory and parameters, while explicitly preserving the task-invariant historical knowledge crucial for mitigating catastrophic forgetting. Extensive experiments show that PeCL achieves a superior balance between privacy preserving and model utility, outperforming baseline models by maintaining high accuracy on previous tasks while ensuring robust privacy.

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ALIVE: 敵対的な学習と有益な口頭評価による LLM 推論の覚醒

大規模言語モデル (LLM) における専門家レベルの推論の探求は、永続的な \textit{報酬のボトルネック} によって妨げられてきました。従来の強化学習 (RL) は、拡張に \textbf{コストがかかる}、ドメイン間で \textbf{脆弱}であり、解決策の基礎となるロジックに対して \textbf{盲目}なスカラー報酬に依存しています。この外部の貧弱な信号への依存は、モデルが推論原理を深く自己完結的に理解することを妨げます。 \textbf{ALIVE} (\emph{指示的言語評価による敵対的学習}) を紹介します。これは、スカラー報酬の最適化を超えて、本質的な推論の獲得に向けたハンズフリー調整フレームワークです。 \emph{認知相乗効果} の原理に基づいた ALIVE は、問題の提起、解決、判断を単一のポリシー モデル内で統合し、正しさのロジックを内面化します。 ALIVE は、敵対的な学習と指導的な口頭フィードバックを組み合わせることで、モデルが生のコーパスから評価基準を直接内部に取り込むことを可能にし、外部の批評を内生的な推論能力に効果的に変換します。数学的推論、コード生成、および一般的な論理推論ベンチマークにわたる経験的評価により、ALIVE が報酬シグナルの制限を一貫して緩和していることが実証されています。同一のデータとコンピューティングを使用して、精度の向上、クロスドメインの汎化の大幅な改善、およびより高い自己修正率を実現します。これらの結果は、推論の三位一体が能力の成長の自立的な軌道を促進し、ALIVE を人間による監視なしの汎用推論調整のためのスケーラブルな基盤として位置づけていることを示しています。

原文 (English)

ALIVE: Awakening LLM Reasoning via Adversarial Learning and Instructive Verbal Evaluation

The quest for expert-level reasoning in Large Language Models (LLMs) has been hampered by a persistent \textit{reward bottleneck}: traditional reinforcement learning (RL) relies on scalar rewards that are \textbf{costly} to scale, \textbf{brittle} across domains, and \textbf{blind} to the underlying logic of a solution. This reliance on external, impoverished signals prevents models from developing a deep, self-contained understanding of reasoning principles. We introduce \textbf{ALIVE} (\emph{Adversarial Learning with Instructive Verbal Evaluation}), a hands-free alignment framework that moves beyond scalar reward optimization toward intrinsic reasoning acquisition. Grounded in the principle of \emph{Cognitive Synergy}, ALIVE unifies problem posing, solving, and judging within a single policy model to internalize the logic of correctness. By coupling adversarial learning with instructive verbal feedback, ALIVE enables models to internalize evaluative criteria directly from raw corpora, effectively transforming external critiques into an endogenous reasoning faculty. Empirical evaluations across mathematical reasoning, code generation, and general logical inference benchmarks demonstrate that ALIVE consistently mitigates reward signal limitations. With identical data and compute, it achieves accuracy gains, markedly improved cross-domain generalization, and higher self-correction rates. These results indicate that the reasoning trinity fosters a self-sustaining trajectory of capability growth, positioning ALIVE as a scalable foundation for general-purpose reasoning alignment without human-in-the-loop supervision.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

AI 評価には標準化されたアイテムレベルのデータリリースが必要である

この意見書では、標準化された項目レベルのベンチマーク データが AI 評価のデフォルトのインフラストラクチャになるべきであると主張しています。現在の評価は、項目の選択が不十分であり、構成が不整合であり、一般化が不十分であるという問題があります。これらの失敗の根本原因は、集計モデル スコアへの重点の置き忘れにあります。品目レベルの証拠がなければ、有効性の主張を評価することができず、その結果、機能の誇張、誤った方向の研究、導入されたシステムに対する不当な信頼が生じます。私たちの立場は、有効な評価を設計するには項目レベルのモデル応答からの経験的証拠が必要であり、そのようなデータの標準化されたリリースは中核的な AI 評価インフラストラクチャとして扱われるべきである、というものです。さらに、このようなリリースにより、評価結果の透明性、複製可能性、および監査可能性が可能になります。この基準が実現可能で結果的なものであることを示すために、AI 評価コミュニティが開発できる統一スキーマの下で、広く使用されているベンチマークからの 155,000 項目にわたる 1,000 万件の回答の項目レベルのアーカイブである OpenEval を構築します。項目レベルのデータがどのようにして低品質項目を特定し、構造の不整合を文書化し、ベンチマークの内部構造に関する妥当性証拠を回復するかを示します。私たちは汚染と著者の負担に関する異議に取り組み、信頼できない主張に対して下される決定のコストと比較して、それぞれの異議が扱いやすいことを示します。

原文 (English)

AI Evaluation Should Require Standardized Item-Level Data Releases

This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified item selection, construct misalignment, and poor generalization. The root cause of these failures is a misplaced focus on aggregate model scores. Without item-level evidence, validity claims cannot be assessed, resulting in inflated capability claims, misdirected research, and unwarranted trust in deployed systems. Our position is that designing valid evaluations requires empirical evidence from item-level model responses, and the standardized release of such data should be treated as core AI evaluation infrastructure. Such a release, in addition, enables transparency, replicability, and auditability of evaluation results. To show the norm is both feasible and consequential, we construct OpenEval, an item-level archive of 10M responses across 155k items from widely-used benchmarks, under a unified schema that the AI evaluation community can develop upon. We demonstrate how item-level data can identify low-quality items, document construct misalignment, and recover validity evidence about benchmarks' internal structure. We address objections around contamination and author burden, and show each is tractable relative to the cost of decisions made on claims that cannot be trusted.

2026-05-25 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

IntentScore: コンピュータ使用エージェントの意図条件付きアクションの評価

Computer-Use Agent (CUA) は、大規模な言語モデルを利用してデスクトップ環境で GUI 操作を実行しますが、アクションの品質を評価せずにアクションを生成するため、後続のステップに連鎖的に発生する不可逆的なエラーにつながります。私たちは、3 つのオペレーティング システムにわたる 398K のオフライン GUI インタラクション ステップから候補アクションをスコアリングすることを学習する、プランを認識した報酬モデルである IntentScore を提案します。 IntentScore は、状態とアクションの関連性に関する対照的な調整と、アクションの正しさに関するマージン ランキングという 2 つの相補的な目標を使用してトレーニングします。アーキテクチャ的には、各候補者の計画意図がアクション エンコーダーに埋め込まれ、同様のアクションを持つ候補者間で論理的根拠が異なるものを区別できるようになります。 IntentScore は、ホールドアウト評価で 97.5% のペア識別精度を達成します。トレーニング中にまったく見えない環境である OSWorld 上のエージェント S3 の再ランカーとしてデプロイされた IntentScore は、タスクの成功率を 6.9 ポイント向上させ、異種のオフライン軌跡から学習した報酬推定が、目に見えないエージェントとタスクの分布に一般化されることを示しています。

原文 (English)

IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents

Computer-Use Agents (CUAs) leverage large language models to execute GUI operations on desktop environments, yet they generate actions without evaluating action quality, leading to irreversible errors that cascade through subsequent steps. We propose IntentScore, a plan-aware reward model that learns to score candidate actions from 398K offline GUI interaction steps spanning three operating systems. IntentScore trains with two complementary objectives: contrastive alignment for state-action relevance and margin ranking for action correctness. Architecturally, it embeds each candidate's planning intent in the action encoder, enabling discrimination between candidates with similar actions but different rationales. IntentScore achieves 97.5% pairwise discrimination accuracy on held-out evaluation. Deployed as a re-ranker for Agent S3 on OSWorld, an environment entirely unseen during training, IntentScore improves task success rate by 6.9 points, demonstrating that reward estimation learned from heterogeneous offline trajectories generalizes to unseen agents and task distributions.

2026-05-25 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

WebGameBench: ブラウザネイティブ ゲームを介したコーディング エージェントの要件からアプリケーションまでの評価

コーディング エージェントはアプリケーション ビルダーとして使用されることが増えていますが、多くの評価は依然として、提供されたアプリケーションではなく、ソース コード、リポジトリ レベルのテスト、または中間トレースに焦点を当てています。 WebGameBench は、コーディング エージェントが凍結された構造化 Web ゲーム仕様をブラウザーでアクセス可能なゲームに変換できるかどうかを評価する、要件からアプリケーションまでのベンチマークです。ブラウザネイティブ ゲームは、コンパクトながら動作密度の高いテストベッドを提供します。単純なゲームであっても、調整された入力処理、空間マッピング、ルールの実行、状態遷移、終了条件、再起動動作、および目に見えるフィードバックが必要です。 WebGameBench では、生成された各アーティファクトが、統一された展開プロトコルの下でブラウザーからアクセス可能なアプリケーションとして構築、提供、公開されます。次に、ランタイム エバリュエーターは実際のブラウザーで配信されたゲームと対話し、EXCELLENT、USABLE、または UNUSABLE の 3 方向のラベルを割り当てます。人間がレビューしたサブセットでは、ランタイム ラベルは、使用可能レート基準に基づく人間のゲームプレイ レビューとほぼ一致しています。 111 のタスク、12 のコーディング エージェント、および 14 の評価構成にわたって、WebGameBench は現在のシステムを分離します。最適な構成では 76.9% の使用可能率に達しますが、優れた率は 20.2% にすぎません。このギャップは、プレイアブル配信の最小しきい値を超えることが、要件を完全に満たすにはまだ遠いことを示しています。私たちの知る限り、WebGameBench はブラウザ ネイティブ ゲーム配信のための最初の要件対アプリケーションのベンチマークであり、配信されたアプリケーションのランタイム ラベルを、使用可能レート基準に基づく独立した人間によるゲームプレイ レビューに対して検証します。

原文 (English)

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games

Coding agents are increasingly used as application builders, yet many evaluations still focus on source code, repository-level tests, or intermediate traces rather than the delivered application. We introduce WebGameBench, a requirement-to-application benchmark that evaluates whether coding agents can turn a frozen Structured WebGame Specification into a browser-accessible game. Browser-native games provide a compact but behavior-dense testbed: even simple games require coordinated input handling, spatial mapping, rule execution, state transitions, terminal conditions, restart behavior, and visible feedback. In WebGameBench, each generated artifact is built, served, and exposed as a browser-accessible application under a unified deployment protocol. A runtime evaluator then interacts with the delivered game in a real browser and assigns a three-way label: EXCELLENT, USABLE, or UNUSABLE. On a human-reviewed subset, the runtime label is broadly aligned with human gameplay review under the Usable-rate criterion. Across 111 tasks, 12 coding agents, and 14 evaluation configurations, WebGameBench separates current systems: the best configuration reaches a 76.9% usable rate but only a 20.2% excellent rate. This gap shows that crossing the minimum playable-delivery threshold is still far from complete requirement satisfaction. To our knowledge, WebGameBench is the first requirement-to-application benchmark for browser-native game delivery that validates delivered-application runtime labels against independent human gameplay review under the Usable-rate criterion.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達規制/政策

XAttnMark: クロスアテンションによる堅牢なオーディオ透かしの学習

音声生成合成および編集技術の急速な普及により、著作権侵害、データの出所、ディープフェイク音声を介した誤った情報の拡散についての深刻な懸念が生じています。ウォーターマークは、知覚できないが識別可能で追跡可能な信号をオーディオ コンテンツに埋め込むことで、プロアクティブなソリューションを提供します。 WavMark や AudioSeal などの最近のニューラル ネットワーク ベースの透かし手法は堅牢性と品質を向上させていますが、堅牢な検出と正確な属性の両方を最適化するのに苦労しています。このペーパーでは、生成器と検出器の間の部分的なパラメータ共有、効率的なメッセージ取得のためのクロスアテンション メカニズム、およびメッセージ配信を改善するための時間調整モジュールを活用することで、このギャップを埋めるクロスアテンション ロバスト オーディオ ウォーターマーク (XATTNMARK) を紹介します。さらに、きめの細かい聴覚マスキング効果を捕捉し、透かしの知覚不能性を改善する、心理音響的に調整された時間周波数 (TF) マスキング損失を提案します。 XATTNMARK は、検出と属性の両方で最先端のパフォーマンスを実現し、さまざまな強度での困難なジェネレーティブ編集を含む、幅広いオーディオ変換に対する優れた堅牢性を実証します。この取り組みは、知的財産を保護し、生成 AI 時代の信頼性を確保するために音声透かしを進歩させます。

原文 (English)

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention

The rapid proliferation of generative audio synthesis and editing technologies has raised serious concerns about copyright infringement, data provenance, and the spread of misinformation via deepfake audio. Watermarking offers a proactive solution by embedding imperceptible yet identifiable and traceable signals into audio content. While recent neural network-based watermarking methods like WavMark and AudioSeal have improved robustness and quality, they struggle to jointly optimize both robust detection and accurate attribution. This paper introduces Cross-Attention Robust Audio Watermark (XATTNMARK), which bridges this gap by leveraging partial parameter sharing between the generator and the detector, a cross-attention mechanism for efficient message retrieval, and a temporal conditioning module for improved message distribution. Additionally, we propose a psychoacoustic-aligned time-frequency (TF) masking loss that captures fine-grained auditory masking effects, improving watermark imperceptibility. XATTNMARK achieves state-of-the-art performance in both detection and attribution, demonstrating superior robustness against a wide range of audio transformations, including challenging generative editing at varying strengths. This work advances audio watermarking for protecting intellectual property and ensuring authenticity in the era of generative AI.

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

パターンと患者: 一人称の物語を通じたパーソナリティ障害診断に関する精神保健専門家に対する LLM の評価

精神医学的自己評価における LLM への依存が高まるにつれ、定性的な患者のナラティブを解釈する LLM の能力に疑問が生じています。この幅広い事例研究では、ポーランド語の一人称自伝的記述に基づいて、境界性 (BPD) および自己愛性 (NPD) パーソナリティ障害の評価において、最先端の LLM とメンタルヘルス専門家を直接比較しています。私たちのサンプル内で、最高のパフォーマンスを誇る Gemini Pro モデルの全体的な診断スコア (65.48%) は、人間の専門家の平均スコア (43.57%) よりも 21.91 パーセント ポイント高かった。モデルも人間の専門家もBPDの特定には優れていましたが(それぞれF1 = 83.4、F1 = 80.0)、モデルはNPDの診断が著しく過小評価され(F1 = 6.7 vs. 50.0)、価値観を伴う用語「ナルシシズム」に対して潜在的な抵抗感を示しました。定性的には、モデルはパターンと形式的なカテゴリーに焦点を当てた自信に満ちた精緻な正当化を提供したが、人間の専門家は簡潔で慎重なままであり、患者の自己感覚と時間的経験を強調した。私たちの調査結果は、LLM は複雑な一人称臨床データを解釈する能力があるかもしれないものの、その出力には依然として重大な信頼性とバイアスの問題があることを示しています。

原文 (English)

Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives

Growing reliance on LLMs for psychiatric self-assessment raises questions about their ability to interpret qualitative patient narratives. This depth over breadth case study directly compares state-of-the-art LLMs and mental health professionals in assessing Borderline (BPD) and Narcissistic (NPD) Personality Disorders based on Polish-language first-person autobiographical accounts. Within our sample, the overall diagnostic scores of the top-performing Gemini Pro models (65.48%) were 21.91 percentage points higher than the average scores of the human professionals (43.57%). While both models and human experts excelled at identifying BPD (F1 = 83.4 & F1 = 80.0, respectively), models severely underdiagnosed NPD (F1 = 6.7 vs. 50.0), showing a potential reluctance toward the value-laden term "narcissism." Qualitatively, models provided confident, elaborate justifications focused on patterns and formal categories, while human experts remained concise and cautious, emphasizing the patients' sense of self and temporal experience. Our findings demonstrate that while LLMs might be competent at interpreting complex first-person clinical data, their outputs still carry critical reliability and bias issues.

2026-05-25 13:00 JSTarXiv cs.AIエージェントビジネス/資金調達

TABX: マルチエージェント強化学習のための高スループットのサンドボックス バトル シミュレーター

環境の設計は、協調的なマルチエージェント強化学習 (MARL) アルゴリズムの開発と評価を形作る上で重要な役割を果たします。既存のベンチマークは重大な課題を浮き彫りにしていますが、カスタム評価シナリオの設計に必要なモジュール性が欠けていることがよくあります。再構成可能なマルチエージェント タスク用に設計された高スループットのサンドボックスである Totally Accelerated Battle Simulator in JAX (TABX) を紹介します。 TABX は、環境パラメータに対するきめ細かい制御を提供し、さまざまなタスクの複雑さにわたる緊急エージェントの動作とアルゴリズムのトレードオフを系統的に調査できるようにします。 TABX は、GPU 上でハードウェア アクセラレーションによる実行に JAX を活用することで、大規模な並列化を可能にし、計算オーバーヘッドを大幅に削減します。 TABX は、高速かつ拡張可能で簡単にカスタマイズできるフレームワークを提供することで、複雑な構造ドメインにおける MARL エージェントの研究を容易にし、将来の研究のための拡張可能な基盤として機能します。コードは https://github.com/ku-dmlab/TABX から入手できます。

原文 (English)

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning

The design of environments plays a critical role in shaping the development and evaluation of cooperative multi-agent reinforcement learning (MARL) algorithms. While existing benchmarks highlight critical challenges, they often lack the modularity required to design custom evaluation scenarios. We introduce the Totally Accelerated Battle Simulator in JAX (TABX), a high-throughput sandbox designed for reconfigurable multi-agent tasks. TABX provides granular control over environmental parameters, permitting a systematic investigation into emergent agent behaviors and algorithmic trade-offs across a diverse spectrum of task complexities. Leveraging JAX for hardware-accelerated execution on GPUs, TABX enables massive parallelization and significantly reduces computational overhead. By providing a fast, extensible, and easily customized framework, TABX facilitates the study of MARL agents in complex structured domains and serves as a scalable foundation for future research. Our code is available at: https://github.com/ku-dmlab/TABX.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達

低分子学習のための共折り畳みモデル表現の系統的評価

クロスモーダルまたはリレーショナル監視の恩恵を受けることが多い視覚モデルや言語モデルとは異なり、低分子基礎モデルは通常、スタンドアロンの分子データで事前トレーニングされます。タンパク質-リガンドの共フォールディングは、モデルを原子レベルのリガンド-タンパク質相互作用にさらすことにより、そのような監視の分子類似物を提供し、共フォールディングモデルが強力な小分子表現を生み出すことができるかどうかという疑問を引き起こします。私たちは、最新の共折り畳みモデルである Voltz2 を使用して、その原子レベルのリガンド表現をスタンドアロンの小分子タスクに移すことで、この疑問を研究します。系統的なプローブと蒸留を通じて、Boltz2 表現が ADMET ベンチマークの既存のモデルと同等またはそれを上回り、分子生成モデリングを加速し、構造誘導リガンド最適化におけるサンプル効率を向上させることを示します。さらに、Boltz2 表現は、3D 配座異性体、バイオアッセイ標識、量子化学的特性など、従来のスタンドアロン分子監視から学習された表現を補完するものであることもわかりました。最後に、表現アライメントを強化学習に拡張し、高密度表現レベルの監視が分子発見におけるスカラー報酬を補完できることを示します。これらの結果は、タンパク質とリガンドの共フォールディングが小分子表現学習のための有望な事前学習パラダイムであることを特定し、Boltz2 を強力な既製の分子基礎モデルとして位置づけることになります。

原文 (English)

A Systematic Evaluation of Co-folding Model Representations for Small-Molecule Learning

Small-molecule foundation models are typically pretrained on standalone molecular data, unlike vision and language models that often benefit from cross-modal or relational supervision. Protein-ligand co-folding provides a molecular analogue of such supervision by exposing models to atom-level ligand-protein interactions, raising the question of whether co-folding models can yield strong small-molecule representations. We study this question using Boltz2, a modern co-folding model, by transferring its atom-level ligand representations to standalone small-molecule tasks. Through systematic probing and distillation, we show that Boltz2 representations match or outperform existing models on the ADMET benchmark, accelerate molecular generative modeling, and improve sample efficiency in structure-guided ligand optimization. We further find that Boltz2 representations are complementary to those learned from conventional standalone molecular supervision, including 3D conformers, bioassay labels, and quantum-chemical properties. Finally, we extend representation alignment to reinforcement learning, showing that dense representation-level supervision can complement scalar rewards in molecular discovery. These results identify protein-ligand co-folding as a promising pretraining paradigm for small-molecule representation learning and position Boltz2 as a strong, off-the-shelf molecular foundation model.

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIビジネス/資金調達

ローカル LLM とレイアウトを意識した解析による表形式 PDF 情報の抽出: 信頼性の評価

学術 PDF 文書から構造化情報を抽出することは簡単ではありません。単一のページは通常、フリー テキストのメタデータと表形式の領域を組み合わせており、プログラム間での変動が見られ、ダウンストリームの解析を妨げる Unicode エンコードのアーティファクトの影響を受けやすくなります。この研究では、ケーススタディとしてインドネシアの高等教育の学術コース登録文書 (Kartu Rencana Studi または KRS) を使用して、表形式の PDF 文書に対する情報抽出アプローチの信頼性を評価します。 LLM のみ、ハイブリッド決定論 - LLM (正規表現と LLM)、LLM フォールバックを備えた Camelot ベースのパイプラインの 3 つの戦略を比較します。実験は、LLM ベースのテストでは 140 のドキュメント、キャメロット ベースのパイプライン評価では 860 のドキュメントで行われ、テーブルとメタデータ内のさまざまなデータを含む 4 つの研究プログラムをカバーしました。 3 つの 12 ~ 14B LLM モデル (Gemma 3、Phi 4、および Qwen 2.5) は、Ollama と GPU なしのコンシューマー グレードの CPU を使用してローカルで実行されました。評価には、しきい値 0.7 の完全一致 (EM) およびレーベンシュタイン類似性 (LS) メトリクスが使用されました。すべてのモデルに適用できるわけではありませんが、結果は、ハイブリッド アプローチが、特に決定論的メタデータの場合、LLM のみと比較して効率を向上できることを示しています。 LLM フォールバックを備えた Camelot ベースのパイプラインは、精度 (EM および LS 最大 0.99 ~ 1.00) と計算効率 (ほとんどの場合、PDF あたり 1 秒未満) の最適な組み合わせを実現しました。 Qwen 2.5:14b モデルは、すべてのシナリオにわたって最も一貫したパフォーマンスを実証しました。これらの発見は、決定論的手法と LLM ベースの手法を統合することが、計算量に制約のある環境で表形式のテキスト ベースの PDF ドキュメントから情報を抽出するための信頼性が高く効率的な戦略であることを裏付けています。

原文 (English)

Tabular PDF Information Extraction with Local LLMs and Layout-Aware Parsing: A Reliability Evaluation

Extracting structured information from academic PDF documents is non trivial: a single page typically combines free text metadata with tabular regions, exhibits cross program variation, and is susceptible to Unicode encoding artifacts that interfere with downstream parsing. This study evaluates the reliability of information extraction approaches for tabular PDF documents, using academic course registration documents (Kartu Rencana Studi or KRS) from Indonesian higher education as a case study. Three strategies are compared: LLM only, Hybrid Deterministic - LLM (regex & LLM), and a Camelot based pipeline with LLM fallback. Experiments were conducted on 140 documents for the LLM based test and 860 documents for the Camelot based pipeline evaluation, covering four study programs with varying data in tables and metadata. Three 12 - 14B LLM models (Gemma 3, Phi 4, and Qwen 2.5) were run locally using Ollama and a consumer grade CPU without a GPU. Evaluations used exact match (EM) and Levenshtein similarity (LS) metrics with a threshold of 0.7. Although not applicable to all models, the results show that the hybrid approach can improve efficiency compared to LLM only, especially for deterministic metadata. The Camelot based pipeline with LLM fallback produced the best combination of accuracy (EM and LS up to 0.99 - 1.00) and computational efficiency (less than 1 second per PDF in most cases). The Qwen 2.5:14b model demonstrated the most consistent performance across all scenarios. These findings confirm that integrating deterministic and LLM based methods is a reliable and efficient strategy for information extraction from tabular text based PDF documents in computationally constrained environments.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達研究/論文

ProtDBench: プロテイン バインダーの設計と評価の統一ベンチマーク

最近のデノボタンパク質バインダー設計の進歩により、実験的検証が増加していますが、報告されたインシリコ測定基準は、標準化されていない評価プロトコルのため、研究全体で解釈したり比較したりすることが依然として困難です。タンパク質バインダー設計のための標準化されたスループットを意識した評価フレームワークである ProtDBench を紹介します。 ProtDBench は、統一されたベンチマーク タスク、評価プロトコル、成功基準を定義し、評価設計が観察されたパフォーマンスにどのような影響を与えるかを系統的に分析できるようにします。大規模なウェットラボの注釈付きデータセットを使用して、評価検証者として一般的に使用される構造予測モデルを分析し、同一のフィルタリング プロトコルの下で検証者に依存する実質的なバイアスと限定的な一致を明らかにします。次に、固定の評価プロトコルの下で、10 個の多様なタンパク質ターゲットにわたる代表的なオープンソースの生成バインダー設計手法をベンチマークします。 ProtDBench には、シーケンスごとの成功率に加えて、固定の 24 時間予算に基づくスループットを意識したメトリクスと、構造の多様性を考慮したクラスター レベルの成功基準が組み込まれています。これらの結果を総合すると、フィルタリング ルール、成功の定義、および計算効率、成功率、構造的多様性の間のスループットを意識した評価によって引き起こされる体系的な違いが明らかになります。全体として、ProtDBench は、現実的な評価設定の下でのタンパク質バインダー設計法の体系的かつ管理された比較をサポートする、公正で再現可能な評価パイプラインを提供します。

原文 (English)

ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation

Recent advances in de novo protein binder design have enabled increasing experimental validation, yet reported in silico metrics remain difficult to interpret or compare across studies due to non-standardized evaluation protocols. We introduce ProtDBench, a standardized and throughput-aware evaluation framework for protein binder design. ProtDBench defines unified benchmark tasks, evaluation protocols, and success criteria, enabling systematic analysis of how evaluation design influences observed performance. Using a large wet-lab annotated dataset, we analyze commonly used structure prediction models as evaluation verifiers, revealing substantial verifier-dependent bias and limited agreement under identical filtering protocols. We then benchmark representative open-source generative binder design methods across ten diverse protein targets under a fixed evaluation protocol. Beyond per-sequence success rates, ProtDBench incorporates throughput-aware metrics based on a fixed 24-hour budget, as well as cluster-level success criteria to account for structural diversity. Together, these results expose systematic differences induced by filtering rules, success definitions, and throughput-aware evaluation between computational efficiency, success rate, and structural diversity. Overall, ProtDBench provides a fair and reproducible evaluation pipeline that supports systematic and controlled comparison of protein binder design methods under realistic evaluation settings.

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達

LLM エージェント ツール呼び出しトラフィックにおけるコンテンツ認識型攻撃の検出: 機能、アーキテクチャ、および評価プロトコルの実証的研究

モデル コンテキスト プロトコル (MCP) は、LLM エージェントが外部ツールを呼び出すためのインターフェイスとして広く採用されていますが、MCP ツール呼び出しトラフィックの学習された監視についてはまだ十分に研究されていません。この記事では、提案された検出器は、各エージェント セッションをグラフ (ツール呼び出しをノード、順次リンクとデータ フロー リンクをエッジ) としてエンコードし、引数と応答に対する文埋め込み機能でノードを強化し、セッションを良性か攻撃かを分類する、MCP ツール呼び出しトラフィックの攻撃検出フレームワークとして紹介されます。 3 つの GNN アーキテクチャ (GAT、GCN、GraphSAGE)、グラフなし MLP、および古典的なベースライン (XGBoost、ランダム フォレスト、ロジスティック回帰、線形 SVM) が評価されます。完全なアーキテクチャ比較は RAS-Eval (タスク階層化分割) で実行され、GraphSAGE は ATBench および結合ソース バリアント (両方ともラベル階層化) で GNN ベースラインとして保持されます。 3 つの発見が得られます。まず、コンテンツ レベルの機能が不可欠です。メタデータのみの検出は、アーキテクチャに関係なく AUROC 0.64 付近で頭打ちになりますが、コンテンツの埋め込みにより AUROC が 0.89 を超えるようになります。第 2 に、単純なランダム分割評価は、タスクに素な分割と比較して AUROC を最大 26 パーセントポイント上昇させます。これは、以前のエージェント検出作業では対処できなかった記憶の混乱です。第三に、検出信号は主に SBERT コンテンツ エンベディングに存在します。プールされたエンサンブル上のツリー アンサンブルによって 0.975 の AUROC に達し、ほとんどの場合、GNN (0.917) や MLP (0.896) を含むプライマリ RAS-Eval 設定のニューラル アーキテクチャよりも優れたパフォーマンスを発揮し、自己監視型事前トレーニングではラベル効率の利点が得られません。このタスク。

原文 (English)

Content-Aware Attack Detection in LLM Agent Tool-Call Traffic: An Empirical Study of Features, Architectures, and Evaluation Protocols

The Model Context Protocol (MCP) has become a widely adopted interface for LLM agents to invoke external tools, yet learned monitoring of MCP tool-call traffic remains underexplored. In this article, the proposed detector is presented as an attack detection framework for MCP tool-call traffic that encodes each agent session as a graph (tool calls as nodes, sequential and data-flow links as edges), enriches nodes with sentence-embedding features over arguments and responses, and classifies sessions as benign or attacked. Three GNN architectures (GAT, GCN, GraphSAGE), a no-graph MLP, and classical baselines (XGBoost, random forest, logistic regression, linear SVM) are evaluated, with the full architecture comparison conducted on RAS-Eval (task-stratified splits) and GraphSAGE retained as the GNN baseline on ATBench and a combined-source variant (both label-stratified). Three findings emerge. First, content-level features are essential: metadata-only detection plateaus around an AUROC of 0.64 regardless of architecture, while content embeddings push the AUROC above 0.89. Second, naive random-split evaluation inflates AUROC by up to 26 percentage points relative to task-disjoint splits, a memorization confound that prior agent-detection work has not addressed. Third, the detection signal resides primarily in the SBERT content embeddings: an AUROC of 0.975 was reached by tree ensembles on pooled embeddings, performing, for the most part, better than the neural architectures in the primary RAS-Eval setting including GNNs (0.917) and the MLP (0.896), and self-supervised pre-training does not deliver a label-efficiency advantage on this task.

2026-05-25 13:00 JSTarXiv cs.AIビジネス/資金調達

回復メカニズムはAIに耐えられるか?スキル形成、労力、現在の測定で見逃されるもの

近代を通して、新しいテクノロジーが労働者に取って代わるとき、社会は同じメカニズムを通じて適応しました。教育は認知の上限を引き上げ、機械がまだ達成できなかったタスクを実行できる労働者を生み出しました。生成 AI は現在、その上限の上限で動作しているため、このサイクルを打破する最初のテクノロジーになる可能性があります。この論文は、労働経済学、複数のプラットフォームにわたる何百万もの AI 会話からの展開データ、2 つの公開データセットの独自の再分析、およびスキル形成の実験に基づいて、3 つの貢献を展開しています。まず、ストック対フローの枠組みは、経済データと教育データが同じテクノロジーについて異なる物語を伝えていることを示しています。つまり、増強は現在の労働者を支配していますが、次世代を生み出す開発パイプラインは負担にさらされています。第二に、証拠ベースの体系的なギャップ分析により、すべての主要な研究で認知の知識次元が測定されていないこと、学習成果を測定している 3 つの研究 (それぞれ $n < 200$) で一貫して AI は学習を向上させることなくパフォーマンスを向上させていることがわかっている (クロスプラットフォーム再分析では $d = 1.21$)、そして専門家と学生の集団の橋渡しをする研究は存在しないことが明らかになりました。第三に、拡張認知分類法 (不確実性、認識論的同一性、認識論的主体性の下での判断) を証拠に基づいて 3 つのケースに適用し、学習を維持する AI 相互作用パターンと、学習を侵食する構造的に類似した相互作用パターンを区別しました。この論文は、AIの社会的リスクは教師に取って代わられることではなく、次世代の能力が形成される生産的な闘争を排除することにあると主張し、現在の測定システムが見逃しているものを対象とした研究と設計の課題を提案している。

原文 (English)

Can the Recovery Mechanism Survive AI? Skill Formation, Labor, and What Current Measurement Misses

Throughout the modern era, when new technologies displaced workers, societies adapted through the same mechanism: education raised the cognitive ceiling, producing workers capable of tasks machines could not yet reach. Generative AI may be the first technology to break this cycle, because it now operates at the top of that ceiling. Drawing on labor economics, deployment data from millions of AI conversations across multiple platforms, original reanalysis of two public datasets, and skill-formation experiments, this paper develops three contributions. First, a stock-versus-flow framework showing that economic data and education data tell divergent stories about the same technology: augmentation dominates current workers, but the developmental pipeline producing the next generation is under strain. Second, a systematic gap analysis of the evidence base, revealing that the knowledge dimension of cognition is unmeasured across all major studies, that the three studies measuring learning outcomes (each $n < 200$) consistently find AI improves performance without improving learning ($d = 1.21$ in our cross-platform reanalysis), and that no study bridges professional and student populations. Third, an extended cognitive taxonomy (judgment under uncertainty, epistemic identity, and epistemic agency) applied to three cases from the evidence to distinguish AI interaction patterns that preserve learning from structurally similar ones that erode it. The paper argues that AI's societal risk lies not in replacing teachers but in eliminating the productive struggle through which the next generation's capacity forms, and proposes a research and design agenda targeting what current measurement systems miss.

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIエージェントビジネス/資金調達研究/論文

TwinRouterBench: 現実的なエージェント LLM ルーティングのための高速静的およびライブ動的評価

LLM ルーティングは、コーディング エージェント、詳細調査システム、コンピュータ使用エージェントなど、単一のユーザー リクエストが多くのモデル呼び出しをトリガーする長期的なアプリケーションで最も重要です。各コールを最も安価な十分なモデルにルーティングすると、品質を犠牲にすることなくコストを削減できますが、既存のルーター ベンチマークはワンショット プロンプトでのみルーターを評価します。中間エージェントのステップでルーターから見えるプレフィックスを公開することは決してなく、より安価な代替品が下流のタスクの成功を維持するかどうかをテストすることもありません。また、多くの場合、評価時にオンラインの LLM 判定に依存します。 2 つのトラックを備えたステップレベルのルーティング ベンチマークである TwinRouterBench を紹介します。静的トラックは、SWE ベンチ、BFCL、mtRAG、QMSum、および PinchBench にわたる 520 のインスタンスからの 970 のルーター可視プレフィックスを提供します。それぞれは、リリースされたダウングレードおよびカスケード プロトコルに基づいて推定された実行検証済みのターゲット層とペアになっています。スコアリングは、オンライン評価者側の LLM ジャッジなしで、ティア ラベル、軌跡メンバーシップ、およびトークン コストに関する決定論的な算術演算です。ダイナミック トラックは、500 ケースの SWE ベンチ検証済みスイート全体でルーターを実行するハーネスを提供します。この論文では、静的な SWE 監視分割とは切り離された 100 件のホールドアウト評価を報告します。各 LLM 呼び出しで、ルーターはロックされたプールから具体的なモデルを選択し、成功は公式のタスク解決と実際の API 消費量によって測定されます。 2 つのトラックは、高速なオフライン反復と、その後のライブ エージェント実行下でのエンドツーエンド検証をサポートします。コードとデータは https://github.com/CommonstackAI/TwinRouterBench で入手できます。

原文 (English)

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user request triggers many model calls. Routing each call to the cheapest sufficient model can cut costs without sacrificing quality, yet existing router benchmarks evaluate routers only on one-shot prompts. They never expose the router-visible prefix at an intermediate agent step, never test whether a cheaper replacement preserves downstream task success, and often rely on online LLM judges at evaluation time. We introduce TwinRouterBench, a step-level routing benchmark with two tracks. The static track provides 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum, and PinchBench, each paired with an execution-verified target tier estimated under a released downgrade-and-cascade protocol; scoring is deterministic arithmetic over tier labels, trajectory membership, and token costs, with no online evaluator-side LLM judge. The dynamic track supplies a harness that runs routers on the full 500-case SWE-bench Verified suite; in this paper we report a 100-case held-out evaluation disjoint from the static SWE supervision split. At each LLM call the router selects a concrete model from a locked pool, and success is measured by official task resolution and realized API spend. The two tracks support fast offline iteration followed by end-to-end validation under live agent execution. Code and data are available at https://github.com/CommonstackAI/TwinRouterBench.