AIニュース 2026-08-12
自動生成: 2026-08-12 11:22 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Testing ads in ChatGPTOpenAI
OpenAI begins testing ads in ChatGPT to support free access, with cle…
-
Daybreak models are now available on AWSOpenAI
OpenAI and AWS are making Daybreak cybersecurity capabilities availab…
-
GeminiアプリのMAUが10億人を突破 Googleで14番目の大台到達製品にITmedia AI+
GoogleのGeminiアプリの月間アクティブユーザー数(MAU)が10億人を突破した。Google検索やWorkspaceなどの組み込…
-
Anthropic、「Claude」で生成したテキストに“見えない透かし” 日本を含むグローバルに適用へITmedia AI+
Anthropicは、「Claude」が生成するテキストに電子透かしを、対応ファイルにC2PA準拠の署名付きメタデータを付与する方針を発表…
-
OpenAI launches ChatGPT desktop app for LinuxTechCrunch AI
OpenAI is finally bringing a dedicated ChatGPT desktop app to Linux o…
-
37兆円のSaaS支出が2030年までに「消える」 Gartnerが予測するエージェントAIの破壊力ITmedia AI+
Gartnerは、企業向けソフトウェアの支出のうち最大2340億ドル(約37兆円)が、エージェント型AIの影響にさらされるとの予測を発表し…
-
「AI生成コンテンツ、実は嫌いではない」 だが「客の半分が去る」本当の理由ITmedia AI+
TWOSTONE&SonsがBtoBの比較検討・発注担当者を対象とした意識調査結果を発表した。約9割がコンテンツに「AIっぽさ」を感じた経…
トピック別件数
- 研究/論文 328件
- LLM/生成AI 315件
- エージェント 172件
- 画像/動画生成 114件
- ビジネス/資金調達 37件
- ロボティクス 35件
- ハードウェア/半導体 23件
- その他 3件
- 規制/政策 3件
日本語メディア7件
ITmedia AI+ (日本語)
GeminiアプリのMAUが10億人を突破 Googleで14番目の大台到達製品に
GoogleのGeminiアプリの月間アクティブユーザー数(MAU)が10億人を突破した。Google検索やWorkspaceなどの組み込み機能を除いたアプリ単体の数字で、同社として14番目の10億ユーザー到達製品となる。競合のChatGPTに続く大台達成で、iOSユーザー数や…
「AI生成コンテンツ、実は嫌いではない」 だが「客の半分が去る」本当の理由
TWOSTONE&SonsがBtoBの比較検討・発注担当者を対象とした意識調査結果を発表した。約9割がコンテンツに「AIっぽさ」を感じた経験があり、その後問い合わせや資料請求を取りやめたと回答する割合も示された。
37兆円のSaaS支出が2030年までに「消える」 Gartnerが予測するエージェントAIの破壊力
Gartnerは、企業向けソフトウェアの支出のうち最大2340億ドル(約37兆円)が、エージェント型AIの影響にさらされるとの予測を発表した。従来型SaaSのシート課金モデルが崩れ、ユーザー数の増加と収益の増加が連動しなくなるという。
Anthropic、「Claude」で生成したテキストに“見えない透かし” 日本を含むグローバルに適用へ
Anthropicは、「Claude」が生成するテキストに電子透かしを、対応ファイルにC2PA準拠の署名付きメタデータを付与する方針を発表した。8月2日に適用開始された「EU AI Act」の透明性義務に伴う措置だが、日本を含む全世界のモデルとサービスに適用される。人間には不可…
【注目の企業】キオクシア、なぜこんなに話題? 今からでも間に合う“入門記事”まとめました
話題を呼び続ける半導体メモリ大手のキオクシア。一体なぜこんなにも注目されるのか。同社の“今”が分かる記事をまとめた。
「よく聞こえなかったデスら」――『ドラクエ』に“AIキャラ”登場、世界観を守るために張り巡らせた創意工夫
『ドラゴンクエストX オンライン』に生成AIを活用した新キャラ「スラミィ」が実装された。生成AIのゆらぎを克服し、ゲームの世界観をどのように守っているのか。
OpenAI、サイバー防御「Daybreak」を赤青2階層に 特化型「GPT-5.6-Cyber」投入
OpenAIはサイバー防御者向けイニシアチブ「Daybreak」を拡張し、2つのアクセス階層「Blue」「Red」と、Red向けの専用モデル「GPT-5.6-Cyber」を発表した。正当なセキュリティ研究におけるAIの過度な拒否動作を抑制し、未知のゼロデイ脆弱性発見やエクスプロ…
海外メディア7件
TechCrunch AI (英語)
Accel closes oversubscribed $550M India fund within weeks, 19 months after its last
The U.S. VC firm still has more than 55% of its previous $650 million India fund available for deployment.
OpenAI launches ChatGPT desktop app for Linux
OpenAI is finally bringing a dedicated ChatGPT desktop app to Linux operating systems.
Brad Lightcap, OpenAI’s longtime COO, is leaving to ‘start something new’
One of OpenAI's longest-serving executives is headed out the door, although the longtime COO told staff that he was "excited to help you al…
General Catalyst leads $1.1B round into 2-month-old River AI
River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of…
An unreleased Anthropic model made progress on one of math’s biggest unsolved problems
For more than 150 years, the Riemann hypothesis has stood as one of the major unsolved problems in mathematics. Anthropic hasn't solved it…
Spotify will label ‘AI Persona’ profiles and exclude their music from recommendations
Spotify is introducing “AI Persona” labels for artist profiles that represent AI-generated identities and will exclude their music from edi…
Anthropic says it will watermark text generated by its AI models
Anthropic will extend support for watermarking AI generations for older models as well.
公式ブログ2件
OpenAI (英語)
Testing ads in ChatGPT
OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user…
Daybreak models are now available on AWS
OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.
論文753件
arXiv cs.AI (英語)
評価型AIの議論的基盤に向けて
評価 AI (EAI) は、単一の推奨事項を作成するのではなく、競合する仮説をそれぞれの賛成および反対の証拠とともに提示することによって、人間の意思決定をサポートする方法として最近提案されています。この意見書では、説明可能で議論の余地のある EAI 形式に正式で計算可能な基盤を提供する特に適切なパラダイムとして (計算論的) 議論を提唱し、分散型で人間中心の EAI システムに向けた長期的な研究課題の基礎を築きます。
原文 (English)
Towards an Argumentative Foundation for Evaluative AI
Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each. In this position paper, we advocate (computational) argumentation as a particularly suitable paradigm to provide a formal, computable foundation for forms of EAI that are explainable and contestable, setting the ground for a long-term research agenda towards distributed and human-centred EAI systems.
フローバイフロー:高損失ドメインでの AI 出力を制御するためのコンテンツ判定バイパス
これまでの研究では、AI の出力速度 V が人間の認知能力 C_max を超えると、高損失領域では人間による監視が構造的に不可能になることが示されています。ただし、操作上の制約は V 単独ではなく、V x L です。ここで、L は項目ごとの認知負荷を示します。 L はトリアージ、判断、対応で構成され、AI の能力向上に対して非対称に対応します。セマンティックな不確定性は汎用設計に内在するため、モデルの機能が向上してもトリアージ コストは減少しません。応答コストは精度の向上に影響されません。判断コストのみが下方圧力に直面しており、この圧力は多くの場合、真の削減ではなく省略を誘発する形で作用します。したがって、能力の向上は L を削減するのではなく、再構築することになります。 AI の出力が正しいかどうかの評価に基づくガバナンス メカニズムは、その評価を AI に委任して幻覚リスクを継承するか、それを人間に委任して V x L の上限に直面するかのいずれかになります。私たちは、コンテンツを評価せずに監視負荷を制御するガバナンス パラダイムである Flow-by-Flow を提案します。正式な可算特徴に基づく認知コスト スコアは、大量生産に非線形コストを課す一方、制度上のキャパシティ キャップにより処理量が C_max 以内に保たれます。コンテンツ判定バイパス超過経路に対する 4 つの設計不変条件を導き出します。それは、コンテンツ判定なし、検査官能力のスケーラブルな消費なし、ID に縛られたアプリケーションごとの摩擦、およびバッチクリアランスなしです。 1 つの参考実装については、これらの不変式が同時に満たされることを示すために議論されていますが、その実際的な困難は明示的に認識されています。 1,000 のパラメータ描画にわたるモンテカルロ分析の例では、複合マルチメトリック フロー制御が試験の 90.8% で監視強化単独よりも優れていることが示唆されています。
原文 (English)
Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains
Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cognitive capacity C_max. The operative constraint, however, is not V alone but V x L, where L denotes per-item cognitive load. L consists of triage, judgment, and response, which respond asymmetrically to AI capability improvement. Triage cost does not decline as models become more capable, because semantic indeterminacy is inherent in general-purpose design. Response cost is invariant to accuracy improvements. Only judgment cost faces downward pressure, and this pressure often operates by inducing omission rather than genuine reduction. Capability improvement therefore restructures L rather than reducing it. Governance mechanisms based on evaluating whether AI output is correct either delegate that evaluation to AI and inherit hallucination risk, or delegate it to humans and face the V x L ceiling. We propose Flow-by-Flow, a governance paradigm that controls supervisory load without evaluating content. A cognitive cost score based on formal, countable features imposes nonlinear costs on high-volume production, while an institutional capacity cap keeps processing volume within C_max. We derive four design invariants for any content-judgment-bypass exceedance pathway: no content judgment, no scalable consumption of examiner capacity, identity-bound per-application friction, and no batch clearance. One reference implementation is discussed to show that these invariants are jointly satisfiable, while its practical difficulties are explicitly acknowledged. An illustrative Monte Carlo analysis across 1,000 parameter draws suggests that composite multi-metric flow control outperforms supervision reinforcement alone in 90.8% of trials.
Determinization in Structure Theories: A Unified Framework via Closure, Comparability, and Joint Admissibility
We develop a formal framework for constructing canonical interpretations from plural structure theories. A structure theory is a triple T =…
人間の運転の能動推論モデルにおける感情
能動推論は、目標指向の行動と不確実性の低減のバランスをとることにより、適応行動をモデル化するための原則的なフレームワークとして登場しました。これは、人間の運転に関する最近の研究を含め、生物学的および人工システム全体に適用されて成功しています。しかし、既存の運転の能動的推論モデルは、交通における行動の重要な決定要因、つまり意思決定に大きな影響を与える感情状態にまだ取り組んでいません。非トラフィック領域におけるこれまでの研究では、感情が包絡線モデルの価性と覚醒の軸に沿って表現される能動推論エージェントについて研究されてきました。ただし、この作業は離散状態空間を使用した単純化された設定に限定されています。この研究では、連続状態での運転のより複雑な能動推論モデルから抽出できる価性と覚醒の拡張された定式化を提案します。特に、現在の状態だけでなく、予測される将来の結果にも基づいて感情的な推定を条件付けします。提案されたアプローチを 2 つのインタラクティブな運転シナリオで評価し、結果として得られる感情信号が同様のシナリオで報告された感情パターンに対応することを示します。
原文 (English)
Emotion in an active inference model of human driving
Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty reduction. It has been successfully applied across biological and artificial systems, including recent work on human driving. However, existing active inference models of driving have yet to address an important determinant of behavior in traffic: affective state, which significantly influences decision-making. Prior work in non-traffic domains has explored active inference agents in which emotions are represented along the axes of valence and arousal in the circumplex model. However, this work has been limited to simplified settings with discrete state spaces. In this work, we propose an expanded formulation of valence and arousal that can be extracted from a more complex active inference model of driving with continuous states. In particular, we condition affective estimates not only on the current state but also on predicted future outcomes. We evaluate the proposed approach in two interactive driving scenarios and show that the resulting emotion signals correspond to affective patterns reported in similar scenarios.
Training Variable Long Sequences with Data-Centric Parallel
Training deep learning models on variable long sequences poses significant computational challenges. Existing methods force a difficult tra…
知っていることと言うことのギャップ: プローブが信頼性を欠いたエラーを発見したとき
線形プローブは、言語モデル内の破損したコンテキストをほぼ完璧な精度で検出しますが、これは信頼性の高い障害予測にはつながりません。その結果、展開の監視に直接的な影響を与える分離が生じます。マルチホップ算術チェーン全体で、破損を検出するプローブは、最終的な答えの正しさについては情報を提供しないことが判明します。構造化された信頼フォーマットに強制されたモデルは、区別できないエラー率を伴う 2 つの値に崩壊します。また、ホップ全体にわたるプローブの持続性は、正しい結果と誤った結果を区別できず、事前に登録された「持続性がピークを上回る」仮説を否定します。知っているが言わないというこのパターンは、推論モデルを含むモデル ファミリ全体に一般化します。リアルタイム モニターとして、プローブ ベースの介入はモデルとエラー タイプに大きく依存します。ブランチ アンド ピックはモデル全体でネット ポジティブであり、Llama-3.1-8B では独自に非破壊的です (4 つは救助され、0 つは破壊されました)。一方、再プロンプトと置換前の中断は、間違ったトレースを救助するのとほぼ同じ割合で正しいトレースを中断します。プローブベースのモニタリングは、言語化された信頼性を補完するために必要ですが、単一の介入が支配することはなく、導入可能な答えは、モデルを認識し、エラーの種類を認識したルーティングです。
原文 (English)
The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured confidence formats collapse to two values with indistinguishable error rates; and probe persistence across hops fails to separate correct from incorrect outcomes, refuting our pre-registered "persistence beats peak" hypothesis. This pattern of knowing but not saying generalises across model families including reasoning models. As a real-time monitor, probe-based interventions are sharply model and error-type dependent: branch-and-pick is net-positive across models and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken), while reprompt and replace-prior break correct traces at roughly the rate they rescue wrong ones. Probe-based monitoring is a necessary complement to verbalised confidence, but no single intervention dominates, and the deployable answer is model-aware, error-type-aware routing.
NL2SHACL-Bench: 自然言語から SHACL への変換のためのベンチマーク スイート
SHACL は、RDF ナレッジ グラフ (KG) の適合性を検証するためのコア テクノロジーです。ただし、SHACL シェイプの作成には、ほとんどのドメイン専門家にはない技術的な専門知識が必要です。自然言語の要件を SHACL (NL2SHACL) に変換すると、この障壁が低くなります。ただし、NL2SHACL 専用のベンチマークはなく、意味的に同等の形状はシリアル化や構造が異なる可能性があるため、生成された形状を評価するには文字列比較を超える方法が必要です。これらの課題に取り組むために、自然言語から SHACL への翻訳のベンチマーク スイートである NL2SHACL-Bench を紹介します。 NL2SHACL-Bench を使用して、このタスク用に 4 つの最先端の大規模言語モデル (LLM) を評価します。私たちの結果は、現在の LLM は構文的に有効な SHACL を生成する能力は高いものの、複雑な論理的および構造的パターンに対して意味的に同等の制約を生成するのに依然として苦労していることを示しています。これは、NL2SHACL-Bench が、NL2SHACL の最先端技術の進歩を測定するための有意義な基盤を提供していることを示しています。
原文 (English)
NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation
SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier. However, there is no dedicated benchmark for NL2SHACL, and evaluating generated shapes requires methods beyond string comparison, as semantically equivalent shapes can differ in serialisation and structure. To tackle these challenges, we present NL2SHACL-Bench, a benchmark suite for natural language to SHACL translation. Using NL2SHACL-Bench, we evaluate four state-of-the-art large language models (LLMs) for this task. Our results show that current LLMs are highly capable of generating syntactically valid SHACL, but still struggle to produce semantically equivalent constraints for complex logical and structural patterns. This indicates that NL2SHACL-Bench provides a meaningful basis for measuring advances in the NL2SHACL state of the art.
スキルベースのエージェントAIシステムにおける動的な連携形成と通信価格設定
最新のエージェント AI システムは、複数の大規模言語モデル エージェントと異種スキルを組み合わせていますが、ほとんどのアーキテクチャは通信を事前に修正するか、完全なブロードキャストを許可しています。アクティブなエージェントと通信リンクの数に応じてトークン コスト、遅延、冗長性、エラー伝播が増加するため、どちらも非効率的になる可能性があります。エージェントの選択とコミュニケーションを、タスク条件付きネット ユーティリティ $U(C\mid x)=V(C\mid x)-\sum_{i\in C}c_i$ を使用した協力ゲームとしてモデル化し、連合レベルのコストをエージェントのアクティブ化コストから分離します。限界値アクティベーション ルールと貪欲ルーターを提案し、エッジごとのコストで通信エッジを最適化するようにモデルを拡張し、推定 Shapley 値を使用して実行前および実行中にどのエージェントに連絡する価値があるかを予測します。この問題をサブモジュラー最大化に結び付け、2 つの限定された保証を証明します。それは、単調でカーディナリティが制約された特別なケースに対する曲率の洗練された境界、もう 1 つは、ダブル グリーディによる制約のない非単調ケースに対する、符号付き目的の補正を伴う厳密な $1/2$ 近似です。どちらの保証もメイン ルーターには直接適用されず、ヒューリスティックのままです。また、限界値ルーティングのエラーをエージェントごとの収益逓減量に結び付けるシャプレーのサブモジュール性サンドイッチ境界も証明します。合成実験では、グリーディ ルーティングは総当たり最適ユーティリティの 99.5%$ を達成し、フル ブロードキャストの場合は 38.8%$ であったのに対し、平均して $8$ のエージェントのうち $1.96$ をアクティブにしました。パフォーマンスはアクティベーション コストと冗長性の重みに対して堅牢ですが、サブモジュール性の強い違反やノイズの多い値の推定では $66%$ に低下します。我々はこのフレームワークをシャプレー価格設定、ヘドニック連合形成、コミュニケーショングラフ枝刈りから区別し、実際のマルチエージェントLLMベンチマークでの評価を提案します。
原文 (English)
Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems
Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast. Both can be inefficient because token cost, latency, redundancy, and error propagation increase with the number of active agents and communication links. We model agent selection and communication as a cooperative game with task-conditioned net utility $U(C\mid x)=V(C\mid x)-\sum_{i\in C}c_i$, separating coalition-level costs from agent activation costs. We propose a marginal-value activation rule and greedy router, extend the model to optimize communication edges with per-edge costs, and use estimated Shapley values to predict which agents are worth contacting before and during execution. We connect the problem to submodular maximization and prove two limited guarantees: a curvature-refined bound for a monotone, cardinality-constrained special case, and a tight $1/2$-approximation, with a correction for signed objectives, for an unconstrained non-monotone case via double greedy. Neither guarantee applies directly to the main router, which remains a heuristic. We also prove a Shapley-submodularity sandwich bound linking the error of marginal-value routing to a per-agent diminishing-returns quantity. In synthetic experiments, greedy routing achieves $99.5%$ of brute-force-optimal utility while activating $1.96$ of $8$ agents on average, compared with $38.8%$ for full broadcast. Performance is robust to activation cost and redundancy weight but falls to $66%$ under strong violations of submodularity or noisy value estimates. We distinguish the framework from Shapley pricing, hedonic coalition formation, and communication-graph pruning, and propose evaluation on real multi-agent LLM benchmarks.
MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents
An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodi…
LLM エージェントが交渉するとき: サプライ チェーンにおける個人情報とダイナミックな交渉
LLM エージェントが意思決定サポートから自律的な調達に移行するにつれて、企業は、委任された交渉担当者が価値を生み出し、それを予測どおりに分割し、損をする契約を回避できるかどうかを知る必要があります。私たちはこれを標準的なサプライチェーンの交渉問題で研究します。つまり、民間の需要情報を持っている買い手が、情報を持たない売り手と数量支払い契約を交渉します。 OpenAI、Google、Alibaba の 9 つの LLM を、9,840 回の LLM 対 LLM ネゴシエーションにわたる検証済みの完全ベイジアン均衡に対してベンチマークします。まず、能力は価値の創造を支配します。エージェントは交渉の 98.9% で合意し、割引なしでファーストベスト剰余の 95.4% を獲得しますが、ベンチマークの 1.25 に対して平均 2.98 ラウンドであり、この遅延により剰余の 21 ~ 34% が侵食されます。能力は信頼性も左右します。ベースライン モデルは、ケースの 19.2% で個別に不合理な契約を受け入れますが、中堅層および主力モデルでは 0.0 ~ 0.6% であり、自動利益検証がそのしきい値を下回る拘束力のあるガードレールになります。第二に、余剰の捕捉は関係的なものである。プロバイダーのアイデンティティは、能力ランクよりも誰が余剰をよりよく獲得するかを予測します。セルフプレイの購入者シェアは平均して OpenAI で 40%、Google で 50%、Alibaba の Qwen で 70% であり、制限された通信と割引なしでも生き残る注文です。どのプロバイダーが販売するかを逆転すると、部門は 7 ~ 18 パーセント ポイント移動し、有能な Qwen の主力製品は、ファミリー間販売で最も弱いことになります。ベンダーの選択は、流通上の第一次決定です。第三に、プロンプトは戦略的な手段です。委任は、プリンシパルの経済的忍耐と、エージェントが促す戦略的忍耐、つまり余剰分割(説明された分散の 90%)の唯一の最も強力な推進要因である自由な展開の選択を分離します。これらを組み合わせることで、割引効率、分散プロファイル、運用の信頼性という 3 つの側面に沿った AI エージェントの均衡参照監査が確立されます。
原文 (English)
When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains
As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations. First, capability governs value creation. Agents agree in 98.9% of negotiations and capture 95.4% of first-best surplus undiscounted, but average 2.98 rounds against the benchmark's 1.25, and this delay erodes 21-34% of surplus. Capability also governs reliability: baseline models accept individually irrational contracts in 19.2% of cases, versus 0.0-0.6% at mid-tier and flagship, making automated profit verification the binding guardrail below that threshold. Second, surplus capture is relational. Provider identity predicts who captures surplus better than capability rank: self-play buyer shares average 40% for OpenAI, 50% for Google, and 70% for Alibaba's Qwen, an ordering that survives restricted communication and no discounting. Reversing which provider sells moves the division by 7-18 percentage points, and the capable Qwen flagship is the weakest cross-family seller: vendor choice is a first-order distributional decision. Third, the prompt is a strategic lever. Delegation separates the principal's economic patience from the agent's prompted strategic patience, a free deployment choice that is the single strongest driver of surplus division (90% of explained variance). Together these establish an equilibrium-referenced audit of AI agents along three dimensions: discounted efficiency, distributional profile, and operational reliability.
TREAT: 等価な数学的表現にわたる形式的知識へのアクセスの評価
AI システムは、柔軟な入力表現と下流ツールで使用される形式的なオブジェクトの間で動作することが増えています。重要な課題は、なじみのない定式化が既知の形式的なオブジェクトを示していることを認識することです。私たちは、定理認識を通じてこの課題を研究します。定理条件の等価性保存変換が与えられた場合、モデルは標準ステートメントに関連付けられた定理同一性を回復する必要があります。大規模言語モデルが等価性を保った数式レベルの変換から既知の定理恒等性を回復できるかどうかを評価するためのベンチマークである TREAT を紹介します。 TREAT は、定理のテキストを言い換えるのではなく、定理条件自体の数学的形式を変更し、残差方程式、証人ステートメント、最適化恒等式、集合関係、演算子の形式、および証明中間の特徴付けを通じて既知の結果を表現します。スクレイピングされた定理ページから開始して、使用可能な数式形式を持つエントリをフィルタリングし、標準定理条件を抽出し、記録された仮定と逆マッピングを使用して変換されたバリアントを生成します。最終的なコーパスには、737 の定理恒等式と 29,480 の変換された行が含まれています。テスト パネルでは、最良のモデルが正しい定理恒等を取得できるのは 60.73% のケースのみです。他のシステムでは、棄権、誤った検出、不正な出力など、さまざまな障害モードが明らかになります。これらは、定理の知識は、表現の等価な変更の下では脆弱になる可能性があることを示唆しています。したがって、TREAT は、安定したターゲット オブジェクト、明示的な等価関係、検証手順、および監査可能なスコアリングを必要とする領域に広く関連する、形式的知識への表現に堅牢なアクセスを評価するための制御されたテストベッドを提供します。
原文 (English)
TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations
AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when an unfamiliar formulation denotes a known formal object. We study this challenge through theorem recognition: given an equivalence-preserving transformation of a theorem condition, a model must recover the theorem identity associated with the standard statement. We introduce TREAT, a benchmark for evaluating whether large language models can recover known theorem identities from equivalence-preserving formula-level transformations. Rather than paraphrasing theorem text, TREAT changes the mathematical form of theorem conditions themselves, expressing known results through residual equations, witness statements, optimization identities, set relations, operator forms, and proof-intermediate characterizations. Starting from scraped theorem pages, we filter for entries with usable mathematical expression forms, extract canonical theorem conditions, and generate transformed variants with recorded assumptions and inverse mappings. The final corpus contains 737 theorem identities and 29,480 transformed rows. On a test panel, the best model retrieves the correct theorem identity in only 60.73% of cases. Other systems reveal different failure modes, including abstention, wrong detection, and malformed outputs. These suggest that theorem knowledge can be fragile under equivalent changes in representation. TREAT therefore provides a controlled testbed for evaluating representation-robust access to formal knowledge, with broader relevance to domains that require stable target objects, explicit equivalence relations, validation procedures, and auditable scoring.
漂流しない AI 科学者: 四足歩行ナビゲーション研究ループにおける味、構造、および反証可能な発見
大規模な言語モデルによって駆動される自律的な研究ループでは、機械学習実験を大規模に実行できますが、実験の動機となる仮説をテストするのではなく、最適化する指標の局所的な改善に向かう傾向があります。私たちはこれに構造的に対処し、シミュレーションで四足ロボットのナビゲーション ポリシーの一般化を研究するための AI 科学者を紹介します。 Karpathy の自動調査パラダイムに基づいて構築されているこのループには、3 つのコンポーネントが追加されています。1 つは固定スキーマに基づいて各反復の予測とその結果を組み合わせる不変の実験カードであり、偽りの仮説を再考することはできません。機械的な役割に限定された特殊なサブエージェント。 kkanbu は、ユーザーの研究の好みを型付けされたナレッジ グラフとして保持し、主観的な判断を行うことが許可されている唯一のコンポーネントである好みのオラクルです。オラクルを分離するために、kkanbu の有無にかかわらず、11 の調査ストリームにわたって同じループを 2 回実行します。どちらのアームもドリフトしません。どちらも独自の仮説の約 4 分の 3 を偽っています。最もよく訓練されたポリシーは、オラクルのないアームから得られます。オラクルが変えるのはスコアではなく方向性です。オラクルは単独でテスト時の適応を探求し、そのアームが主導する勝利のデザインを作成し、もう一方のアームが繰り返し再導き出した教訓をストリーム全体に伝えました。スキャフォールドはループを正直に保ちます。 kkanbuはそれがどこに見えるかを決定します。
原文 (English)
An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop
Autonomous research loops driven by large language models can run machine-learning experiments at scale but tend to drift toward local refinements of whichever metric they optimise rather than testing the hypotheses that motivate the experiments. We address this structurally and present an AI Scientist for studying generalisation in quadruped robot navigation policies in simulation. Building on the autoresearch paradigm of Karpathy, our loop adds three components: an immutable experiment card that pairs each iteration's prediction with its outcome under a fixed schema, so a falsified hypothesis cannot be retconned; specialised subagents restricted to mechanical roles; and kkanbu, a preference oracle that holds the user's research taste as a typed knowledge graph and is the only component permitted to make subjective judgements. To isolate the oracle we run the identical loop twice across eleven research streams, with and without kkanbu. Neither arm drifts: both falsify roughly three quarters of their own hypotheses, and the best trained policy comes from the oracle-less arm. What the oracle changes is direction, not score: it alone explores test-time adaptation, it authored the winning designs where its arm led, and it carried lessons across streams that the other arm repeatedly re-derived. The scaffold keeps the loop honest; kkanbu decides where it looks.
The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes
We introduce a continuous metric field framework trained by a single causal contrastive loss. The framework encodes a scene into coefficien…
TeXFix-Bench: LLM ベースのドキュメント ソース修復のための経験に基づいたマルチフォーマット ベンチマーク
科学的および技術的な文章は、コンパイルする必要があるマークアップ ソースに依存します。LaTeX、Typst、および Markdown パイプラインは、区切り文字の欠落、環境の不一致、インポートの破損、またはパッケージの競合により失敗します。既存のドキュメント修復評価では、経験的な障害モデルが欠如したアドホックな編集で障害が発生します。我々は、マイニングされた障害分類法に基づいた LLM ベースの完全なソース文書修復のためのマルチフォーマット ベンチマークである TeXFix-Bench を紹介します。 TeX Stack Exchange、GitHub コミット、およびパッケージ ドキュメント (168 個の検証済みフォールト、$\kappa$=0.34 でのデュアル オープン コーディング) からの局所的なハード クラッシュ LaTeX フォールトのグラウンデッド セオリー研究により、DocMut としてインスタンス化された 18 カテゴリーの分類が得られます: 3 つの形式にわたる 48 の AST 認識演算子。 3 モデルのクロスベンチマークでは、DocMut の障害は、同じシード上のパターンベースの突然変異よりも修復が 5.6 ~ 9.2 pp 難しいことが示されており、実際のエラーのケーススタディ (採掘された人間によるクラッシュ 88 件、修復成功率 67.0%) では、両方の合成セットが下から分類されています。オープンにライセンスされた 743 のシードから 10,437 のインスタンスを構築し、プロバイダー固定ルーティングを備えた固定ゼロショット プロトコルの下で 7 つの LLM を評価し、合計推論コスト約 200 米ドルで 48,651 回の試行を収集しました。 6,613 インスタンス x 7 モデルの完全なバランス マトリックスにより、すべてのランキングが確認されます。ピン留めされたエンジン ゲートでは、27.5 ポイントの Intention-to-Treat コンパイル スプレッド (56.7 ~ 84.2%) が得られます。 Typst は LaTeX や Markdown よりも著しく難しいです。 28,129 回のコンパイル修復に関する修復オラクルでは、コンパイル修復の 13.6 ~ 18.5% でドキュメント テキストが大幅に変更され、修復ランクがコンパイル ランクから乖離していることが示されています。つまり、コンパイル率が最も低いモデルが、成功したモデルの中で最もよくコンテンツを修復します。コンパイルの成功だけでは修復の品質が誇張されます。分類法、DocMut、およびすべてのキャンペーン成果物をリリースします。
原文 (English)
TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair
Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatched environments, broken imports, or package conflicts. Existing document-repair evaluations inject faults with ad-hoc edits that lack an empirical fault model. We present TeXFix-Bench, a multi-format benchmark for LLM-based full-source document repair grounded in a mined fault taxonomy. A Grounded-Theory study of localized hard-crash LaTeX faults from TeX Stack Exchange, GitHub commits, and package documentation (168 verified faults, dual open coding at $\kappa$=0.34) yields an 18-category taxonomy instantiated as DocMut: 48 AST-aware operators across three formats. A three-model cross-benchmark shows DocMut faults are 5.6-9.2 pp harder to repair than pattern-based mutations on the same seeds, and a real-error case study (88 mined human crashes, 67.0% repair success) brackets both synthetic sets from below. We construct 10,437 instances from 743 openly licensed seeds and evaluate seven LLMs under a fixed zero-shot protocol with provider-pinned routing, collecting 48,651 attempts at about USD 200 total inference cost. A complete 6,613-instance x 7-model balanced matrix confirms all rankings. A pinned engine gate yields a 27.5-point intention-to-treat compile spread (56.7-84.2%). Typst is markedly harder than LaTeX and Markdown. A restoration oracle over 28,129 compiling repairs shows that 13.6-18.5% of compiling repairs materially alter document text, and restoration rank diverges from compile rank: the model with the lowest compile rate restores content best among its successes. Compile success alone overstates repair quality. We release the taxonomy, DocMut, and all campaign artifacts.
CMU-Drive および V2V-VLA: 推論ベンチマークと車両間の視覚、言語、行動モデルを備えた協調的なマルチエージェント統合運転
Vision-Language-Action(VLA)モデルは最近、エンドツーエンドの自動運転において目覚ましいパフォーマンスを達成しましたが、既存のアプローチは主に個別の単一の自動運転エージェント向けに設計されており、協調的な認識、推論、計画のサポートは限られています。我々は、バックグラウンド交通参加者による安全性が重要な運転シナリオで動作する複数のコネクテッド自動運転車 (CAV) による協調自動運転を評価するための閉ループのエンドツーエンドのベンチマークである、協調型マルチエージェント統合運転推論 (CMU-Drive) を紹介します。さらに、運転行動、将来のウェイポイント、言語推論、およびコミュニケーションポリシーを共同生成することにより、協調運転を単一の前進パスに統合する協調VLAモデルであるVehicle-to-Vehicle Vision-Language-Action(V2V-VLA)を提案します。 CMU-Drive の実験は、VLA 協調運転の最初のベンチマークとベースラインを確立し、マルチエージェント、閉ループ、エンドツーエンドの協調自動運転に関する将来の研究の基盤を提供します。オープンソースの研究を促進するために、コード、ベンチマーク、モデル チェックポイントは一般に公開されます。
原文 (English)
CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models
Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning. We present Cooperative Multi-agent Unified Driving with Reasoning (CMU-Drive), a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles (CAVs) operating in safety-critical driving scenarios with background traffic participants. We further propose Vehicle-to-Vehicle Vision-Language-Action (V2V-VLA), a cooperative VLA model that integrates cooperative driving into a single forward pass by jointly generating driving actions, future waypoints, language reasoning, and communication policies. Experiments on CMU-Drive establish the first benchmark and baseline for cooperative VLA driving and provide a foundation for future research on multi-agent, closed-loop, end-to-end cooperative autonomous driving. Our code, benchmark, and model checkpoint will be publicly released to facilitate open-source research.
継続的 LLM エージェントにおける制御されたメモリ干渉
Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience.しかし、記憶の進化は、単により多くの情報を保存するプロセスではありません。新しい経験は、既存の記憶状態を強化、修正、または妨害する可能性があります。既存のシステムは主にメモリの構築と関連性に基づく検索に重点を置いていますが、状態、時間的有効性、または権限が異なっていても、複数のメモリが同時に関連性を維持する可能性があります。さまざまなメモリ関係の下でエージェントのメモリがどのように進化するかを研究するための、制御された診断およびデータ生成フレームワークである Controlled Memory Interference (CMI) を紹介します。制御されたメモリの進化では、良性の蓄積の影響は限定的ですが、関係固有の干渉は、ターゲットメモリの公開をブロックするか、下流での使用を中断することにより、安定性をほとんど向上させずに更新の可塑性を大幅に抑制します。語彙検索と密検索は明確な干渉経路を示しますが、ポイズニングは最新性のみよりも更新権限の合図により敏感です。 CMI は診断を超えて、干渉を認識したメモリ学習の対象を絞った例を提供し、元のメモリ タスクのパフォーマンスを維持しながら、有効な更新と干渉を誘発するメモリの区別を改善します。 These findings show that memory evolution is shaped not only by memory scale, but also by interactions among accumulated experiences. More broadly, memory interference emerges as an important factor for reliable continual agent memory systems.
原文 (English)
Controlled Memory Interference in Continual LLM Agents
Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience. Yet memory evolution is not simply a process of storing more information: new experiences may reinforce, revise, or interfere with existing memory states. Existing systems mainly emphasize memory construction and relevance-based retrieval, but several memories may remain simultaneously relevant while differing in state, temporal validity, or authority. We introduce Controlled Memory Interference (CMI), a controlled diagnostic and data-generation framework for studying how agent memory evolves under different memory relationships. Across controlled memory evolution, benign accumulation has limited effects, whereas relationship-specific interference sharply suppresses update plasticity with little stability gain, either by blocking target-memory exposure or by disrupting its downstream use. Lexical and Dense retrieval exhibit distinct interference pathways, while poisoning is more sensitive to update-authority cues than to recency alone. Beyond diagnosis, CMI provides targeted examples for interference-aware memory learning, improving the distinction between valid updates and interference-inducing memories while preserving performance on original memory tasks. These findings show that memory evolution is shaped not only by memory scale, but also by interactions among accumulated experiences. More broadly, memory interference emerges as an important factor for reliable continual agent memory systems.
単一のチャットボットから管理されたエージェントのエコシステムまで: ミッションクリティカルな病院情報管理システムのためのエージェント AI パターン カタログおよびオーケストレーション フレームワーク
病院は、他の業界でのテクノロジーの適応の急増に対処しながら、トリアージ管理、文書化、スケジューリング、収益サイクルのワークフローに AI を組み込もうと競い合っていますが、ほとんどの導入は生産の端で行き詰まった断片的な試験運用のままであり、患者と医療機関を運営の脆弱性、管理されていないリスク、増大する技術的負債にさらしています。同時に、Fortune Business Insights のレポートによると、ヘルスケアにおける世界の AI 市場は 2034 年までに 1 兆米ドル近くに達すると予測されており、アーキテクチャ上の失敗やスケーリング戦略の失敗による財務上の影響はさらに拡大します。この調査では、コンプライアンス最優先の Agentic AI パターン カタログとオーケストレーション フレームワークを提案しています。このフレームワークは HIMS 向けに特別に構築されており、単一の LLM チャットボットを超えて、自律型および半自律型エージェントの管理されたエコシステムに向けて移行しています。このフレームワークは、(i) エージェントの役割の分類、(ii) 各パターンをリスク層、人間参加型チェックポイント、ガバナンスフックにマッピングする正式なリスク階層化モデル、および (iii) Epic、Cerner、MEDITECH などの EHR/HIMS ランドスケープ全体でマルチエージェントのワークフローを調整できる統合オーケストレーション ランタイムを追加することによって拡張されます。技術的には、このフレームワークは vLLM ベースの推論、最適化されたページング メモリ、機密コンピューティング、MCP ベースのオンプレミス展開を組み合わせており、HIPAA、GDPR、EU AI 法、インドの DPDP および DISHA 法、ISO 27001、ISO 27002、ISO 14971、および IEC 62304 に準拠したエンドツーエンドの暗号化とコードとしてのポリシー制御を強制します。ガバナンスと監査可能性を制限しながら、文書化の時間、統合作業、AI パイロットの減少を削減する能力と効率性を備え、AI 投資を持続可能な臨床、運用、財務 ROI に変えるための緊急に必要な青写真を病院のリーダーと統治当局に提供します。
原文 (English)
From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Orchestration Framework for Mission-Critical Hospital Information Management Systems
Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage management, documentation, scheduling, and revenue-cycle workflows, yet most deployments remain as fragmented pilots that stall at the edge of production, exposing patients and institutions to operational fragility, ungoverned risk, and mounting technical debt. At the same time, the global AI-in-healthcare market is projected to exceed nearly USD 1 trillion by 2034, according to the report of Fortune Business Insights, amplifying the financial consequences of architectural missteps and failed scaling strategies. This research proposes a compliance-first Agentic AI pattern catalogue and orchestration framework, purposely built for HIMS, moving beyond the single LLM chatbots and towards a governed ecosystem of autonomous and semi-autonomous agents. The framework extends by adding (i) a taxonomy of Agentic roles, (ii) a formal risk-stratification model that maps each pattern to risk tiers, human-in-the-loop checkpoints, and governance hooks, and (iii) a unified orchestration runtime capable of coordinating multi-agent workflows across EHR/HIMS landscapes such as Epic, Cerner, and MEDITECH. Technically the framework combines vLLM-based inference, optimized paging memory, confidential computing, and MCP based on-premise deployment, enforcing end-to-end encryption and policy-as-code controls aligned with HIPAA, GDPR, the EU AI Act, India's DPDP and DISHA Acts, ISO 27001, ISO 27002, ISO 14971 and IEC 62304. We exhibit how the proposed architecture is capable and efficient to reduce the documentation time, integration effort, and AI pilot attrition while constricting the governance and auditability, offering hospital leaders and governing authorities an urgently needed blueprint to convert AI investment into sustainable clinical, operational, and financial ROI
Agent-MD: ステートフル GCMC のイベント駆動型エスカレーションによる選択的 LLM 介入 -- MD キャンペーン
長期にわたる分子シミュレーション キャンペーンでは、保存された状態からの繰り返しの継続、来歴を意識した進行、適応的な評価、および固定ルールでは安全に解決できないワークフロー条件の時折の解釈が必要です。ここでは、大規模言語モデル (LLM) 推論をキャンペーンの構築とイベント トリガーのレビューに選択的に配置するフレームワークである Agent-MD を紹介します。一方、日常的なシミュレーション、分析、継続、アーカイブ、および状態の進行は、承認されたポリシーと明示的な状態レコードを使用して永続的なルールベースのキャンペーン エージェントによって処理されます。 Agent-MD は、5 つのモンモリロナイト系と 3 つの連続した相対湿度状態 (RH = 0.9-0.3-0.1) からなるグランドカノニカル モンモリロナイト分子動力学 (GCMC-MD) 水蒸気脱着キャンペーンで実証されました。ワークフローは、15 のシステム RH 状態にわたって、状態固有のサンプリング長と来歴を意識した再起動継承を使用して、120 のセグメント化されたシミュレーション サイクルを完了しました。ルーチンの生産では、推論エージェントのライブ呼び出しは必要ありませんでしたが、1 つの状態がレビュー境界に達しました。その後、保存された 2 つのインシデントがブラインド推論エージェントの再生によって評価され、根底にあるワークフローの問題が特定され、適切なフォローアップ アクションが推奨されました。シミュレーションでは、明確な組成依存性の低RH応答も明らかになりました。Caを含むモンモリロナイトは、NaおよびKを含むシステムよりも多くの層間水を保持し、より大きな基底間隔を維持しますが、最高電荷のNaシステムは、乾燥条件下でより多くの残留水を保持します。これらの結果は、長時間実行される科学的ワークフローでは、すべての操作を LLM 推論ループ内に配置する必要がないことを示しています。代わりに、選択的推論を決定論的な実行、構造化された証拠、検証済みの制御ハンドオフと組み合わせて、再現可能で監査可能なエージェント支援分子シミュレーションを提供できます。
原文 (English)
Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns
Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessment, and occasional interpretation of workflow conditions that cannot be resolved safely by fixed rules. Here, we present Agent-MD, a framework that places large language model (LLM) reasoning selectively at campaign construction and event-triggered review, while routine simulation, analysis, continuation, archiving, and state progression are handled by a persistent rule-based campaign agent using approved policies and explicit state records. Agent-MD was demonstrated in a grand canonical Monte Carlo-molecular dynamics (GCMC-MD) water-vapor desorption campaign comprising five montmorillonite systems and three sequential relative-humidity states (RH = 0.9-0.3-0.1). Across 15 system-RH states, the workflow completed 120 segmented simulation cycles with state-specific sampling lengths and provenance-aware restart inheritance. Routine production required no live reasoning-agent invocation, while one state reached a review boundary; two preserved incidents were subsequently evaluated through blinded reasoning-agent replay, which identified the underlying workflow problems and recommended appropriate follow-up actions. The simulations also revealed distinct composition-dependent low-RH responses, with Ca-bearing montmorillonite retaining more interlayer water and maintaining a larger basal spacing than the Na- and K-bearing systems, while the highest-charge Na system retained more residual water under dry conditions. These results demonstrate that long-running scientific workflows need not place every operation inside an LLM reasoning loop: selective reasoning can instead be combined with deterministic execution, structured evidence, and validated control handoffs to provide reproducible and auditable agent-assisted molecular simulation.
Contextual Value Alignment via Multilayer Combinatorial Fusion
Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches…
メンデル ジョージア マシン: 比較進化による再帰的自己改善コーディング エージェント
独自のソース コードを繰り返し書き直す自己改善型コーディング エージェントは、コーディング タスクで優れたパフォーマンスを示しています。ただし、既存のソリューションは通常、一度に 1 つの障害の軌跡から自己修正を導き出し、エージェントの拡張する過去の試行のアーカイブで利用できる豊富な比較信号を見落としています。メンデルの制御遺伝原理に従って、メンデル ジェネレーター マシン (MGM) を導入します。一般的な単一軌跡のクローン変異に加えて、MGM には蓄積された証拠をより有効に活用する 2 つの新しいタイプの自己改変が含まれています。反応規範変異は複数のタスクを同時に実行する軌跡に基づいてエージェントを編集し、系統間ハイブリダイゼーションは同じタスクの別の系統からの参照エージェントの軌跡を使用してエージェントを編集します。加法的な適応度ランドスケープ モデルを理論的に証明し、制御されたサロゲート シミュレーションを通じて、新しい戦略が単一軌道ベースラインに対するより高速でより良い収束を促進することを実証します。SWE ベンチと Polyglot での実験により、MGM のパフォーマンス、効率、一般化可能性が一貫して向上していることが確認されています。
原文 (English)
Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive self-modification from a single failure trajectory at a time, overlooking rich comparative signals available in the agent's expanding archive of past attempts. According to Mendelian principles of controlled inheritance, we introduce Mendel G\"odel Machine (MGM). In addition to the general single-trajectory clonal mutation, MGM includes two new types of self-modification that better utilizes evidences accumulated: the reaction-norm mutation edits an agent based on its trajectories on multiple tasks simultaneously, and the cross-lineage hybridization edits an agent using the trajectory of a reference agent from another lineage on the same task. Under an additive fitness landscape model, we prove theoretically and demonstrate via controlled surrogate simulation that the new strategies facilitate a faster and better convergence over single-trajectory baselines. Experiments on SWE-bench and Polyglot confirm MGM's consistent improvement in performance, efficiency, and generalizability.
Agentic AI フレームワークは、眼底写真から緑内障を検出するための大規模言語モデルの基本的な制限を克服します
大規模言語モデル (LLM) は医療画像の解釈に有望ですが、幻覚、精度の限界、および実行ごとの不一致という問題があります。私たちは、眼底写真から緑内障を検出するための特殊な深層学習ツールと LLM を統合するエージェント AI フレームワークを開発し、検証しました。ワークフローには 3 つのステップがありました。(1) LLM の初期評価。 (2) 画質 (QAModel、FundaQ-8)、緑内障分類 (SwinV2-Tiny)、および視神経乳頭/カップセグメンテーション (SegFormer-B0) のための特殊なツールを呼び出すための関数呼び出し。 (3) 第一印象をツールの出力と統合する LLM リフレクション。 2 つの LLM (Gemini 2.5 Flash、GPT-5.4 mini) を 2 つの公開データセット (ORIGA、n=100; RIM-ONE-v3、n=100) 上で、トリミングされていない視野とトリミングされた視野の下で評価しました。すべての画像は、覆面フェローシップで訓練を受けた緑内障専門医によって独立して採点されました。エージェント ワークフローにより、すべての条件で分類精度が 16 ~ 47 パーセント ポイント向上し、スペシャリストの 6 ポイント以内に到達しました。 RIM-ONE-v3 では、最適な構成は専門家の精度 88% と一致しました。 LLM 単独のアプローチは 2 つの点で失敗しました。GPT-5.4 mini は正の偏り (感度 95 ~ 100%、特異度 0 ~ 5%) を示しましたが、Gemini 2.5 Flash は実行間で確率的に変動しました。エージェント ワークフローでは両方が修正されました。カップ対ディスク比の誤差は 15 ~ 50% 減少し (MAE 0.156 ~ 0.228 から 0.104 ~ 0.132)、専門家によるグレーディングとの相関関係は弱い (r=0.12 ~ 0.39) から中程度から強い (r=0.59 ~ 0.84) に上昇しました。実行間の一貫性は、ほぼランダム (カッパが -0.01 まで低かった) からほぼ完璧 (カッパが最大 0.96) に上昇しました。 LLM と特殊なツールを統合することで、過剰診断や実行ごとのばらつきなど、LLM 単独のアプローチの主要な制限に対処しました。両方の LLM で利益が得られたことは、バックボーン全体での汎用性を示唆しており、医療 AI においてモノリシック モデルから調整されたマルチエージェント システムへの移行を示唆している可能性があります。
原文 (English)
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography. The workflow had three steps: (1) LLM initial assessment; (2) function calling to invoke specialized tools for image quality (QAModel, FundaQ-8), glaucoma classification (SwinV2-Tiny), and optic disc/cup segmentation (SegFormer-B0); and (3) LLM reflection integrating the initial impression with tool outputs. Two LLMs (Gemini 2.5 Flash, GPT-5.4 mini) were evaluated on two public datasets (ORIGA, n=100; RIM-ONE-v3, n=100) under uncropped and cropped fields of view; all images were independently graded by a masked fellowship-trained glaucoma specialist. The agentic workflow improved classification accuracy by 16 to 47 percentage points across all conditions, reaching within 6 points of the specialist; on RIM-ONE-v3 the best configurations matched the specialist accuracy of 88%. LLM-alone approaches failed in two ways: GPT-5.4 mini showed positive bias (sensitivity 95-100%, specificity 0-5%), while Gemini 2.5 Flash varied stochastically between runs; the agentic workflow corrected both. Cup-to-disc ratio error fell 15-50% (MAE 0.156-0.228 to 0.104-0.132), and correlation with specialist grading rose from weak (r=0.12-0.39) to moderate-strong (r=0.59-0.84). Run-to-run consistency rose from near-random (kappa as low as -0.01) to near-perfect (kappa up to 0.96). Integrating LLMs with specialized tools addressed key limitations of LLM-alone approaches, including over-diagnosis and run-to-run variability. Gains held for both LLMs, suggesting generalizability across backbones, and may signal a shift from monolithic models toward orchestrated multi-agent systems in medical AI.
IntelliAudit: 大規模言語モデルを使用した監査制御の評価
IT 監査では、監査人は異種の組織証拠がセマンティック セキュリティとコンプライアンスの管理を満たしているかどうかを判断する必要があります。関連する証拠がポリシー、記録、スプレッドシート、運用成果物に分散されており、監査の結論はキーワードの一致ではなく証拠の十分性に依存するため、この判断を自動化することは困難です。 IT 監査証拠評価のための、検索に基づいたマルチエージェント システムである IntelliAudit を紹介します。コントロールと証拠コーパスが与えられると、IntelliAudit は関連するアーティファクトを取得し、証拠に基づいた評価を生成し、不利な調査結果に異議を唱え、意見の相違を裁定し、引用された証拠、理論的根拠、不足している証拠の分析、および是正ガイダンスを含む監査人向けの推奨事項を作成します。当社は、ISO/IEC 27001 に基づいて IntelliAudit をインスタンス化し、専門の監査人のレビューと監査の準備ができているユーザーのフィードバックを使用して、複数のシミュレートされた組織にわたって IntelliAudit を評価します。この評価では、IntelliAudit が統制の解釈、証拠に基づく推論、監査準備のワークフローをサポートできることが示されるとともに、十分性の判断を調整し、過度に寛容な推奨事項を修正するための人による監視の重要性も明らかになりました。これらの結果は、検索に基づいたマルチエージェント システムは監査証拠のレビューを支援できるが、自律的な認証システムではなく意思決定支援ツールにとどまるべきであることを示唆しています。
原文 (English)
IntelliAudit: Using Large Language Models to Evaluate Audit Controls
IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This judgment is difficult to automate because relevant evidence is distributed across policies, records, spreadsheets, and operational artifacts, and because audit conclusions depend on evidentiary sufficiency rather than keyword matching. We present IntelliAudit, a retrieval-grounded multi-agent system for IT audit evidence evaluation. Given a control and an evidence corpus, IntelliAudit retrieves relevant artifacts, generates an evidence-grounded assessment, challenges adverse findings, adjudicates disagreements, and produces an auditor-facing recommendation with cited evidence, rationale, missing-evidence analysis, and remediation guidance. We instantiate IntelliAudit on ISO/IEC 27001 and evaluate it across multiple simulated organizations using expert auditor review and audit-readiness user feedback. The evaluation shows that IntelliAudit can support control interpretation, evidence-grounded reasoning, and audit-preparation workflows, while also revealing the importance of human oversight for calibrating sufficiency judgments and correcting overly permissive recommendations. These results suggest that retrieval-grounded multi-agent systems can assist audit evidence review, but should remain decision-support tools rather than autonomous certification systems.
ナレッジグラフの質問応答のための研究者エージェント向け
自然言語の質問を、大規模なナレッジ グラフに対して実行できる SPARQL クエリに変換するには、語彙の曖昧さを解決し、ターゲット オントロジーで表面用語を基礎付け、構文的に有効で意味的に忠実なグラフ パターンを作成する必要があります。私たちは、静的ツールを使用するエージェントを一歩超えたエージェントのテキストから SPARQL へのシステムを紹介します。つまり、検証セットの推論の各ラウンドの後に、独自のプロンプト、ルール、およびツール オーケストレーション コードへの変更を提案およびテストする研究者エージェントです。 DBpedia でループをインスタンス化し、低コストの推論モデルによってエージェントの 9 つの連続バージョンを進化させ、2 つのより強力なバックボーン モデルを備えた最高のパフォーマンスの構成を展開します。この研究では、次の 3 つの観察結果が得られました。(i) 自己改善はすぐに収束し、2025 年の DBpedia 検証セットで全体の精度が 0.22 に達しました。 (ii) ボトルネックは一貫して、SPARQL 構文や修飾子ではなく、基本的なグラフ パターンの述語の選択にあります。 (iii) DBpedia のプロパティの曖昧さにより、いくつかのベンチマーク項目が正しいクエリにペナルティを与えているようであり、将来の Text-to-SPARQL ベンチマークは機械翻訳と情報検索メトリクスの組み合わせを使用してスコアリングされるべきであることを示唆しています。
原文 (English)
Towards Researcher Agents for Knowledge-Graph Question Answering
Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, grounding surface terms in the target ontology, and producing graph patterns that are both syntactically valid and semantically faithful. We present an agentic text-to-SPARQL system that goes one step beyond static tool-using agents: a researcher agent that, after each round of inference on a validation set, proposes and tests changes to its own prompts, rules, and tool-orchestration code. We instantiate the loop on DBpedia, evolve nine successive versions of the agent driven by a low-cost reasoning model, and deploy the best-performing configuration with two stronger backbone models. The study yields three observations: (i) self-improvement converges quickly and then achieves 0.22 overall accuracy on the 2025 DBpedia validation set; (ii) the bottleneck is consistently in basic-graph-pattern predicate selection, not in SPARQL syntax or modifiers; and (iii) several benchmark items appear to penalise correct queries due to property ambiguity in DBpedia, suggesting that future Text-to-SPARQL benchmarks should be scored using a combination of machine translation and information retrieval metrics.
臨床基盤モデルにおける患者のプライバシーの保護: 技術的および法的観点
大規模な患者データに基づいてトレーニングされた臨床基盤モデルは、意思決定支援、スクリーニング、公衆衛生にますます使用されています。導入が拡大するにつれて、モデルを介した漏洩によってプライバシーのリスクがますます増大していますが、その蔓延と深刻度は依然として十分に定量化されていません。モデルは機密性の高いトレーニングアーティファクトを開示することができ、データ処理制御だけでは捕捉できない方法で患者を再識別できるようになります。 HIPAA や GDPR などの既存のフレームワークは、このような間接的な脅威に対して限定的なガイダンスしか提供していません。私たちは、臨床基盤モデルにおけるプライバシー リスクを評価するための実践的なフレームワークを提案し、導入設定全体にわたる現実的な漏洩シナリオを示し、それを法的制度にマッピングし、補完的な技術的および法的緩和策の概要を示します。当社の分析は、現実的な使用法に基づいたコンテキスト認識型のリスク評価を提供し、患者のプライバシーを厳密に保護しながら医療基盤モデルの価値を維持します。
原文 (English)
Protecting patient privacy in clinical foundation models: Technical and legal perspectives
Clinical foundation models trained on large-scale patient data are increasingly used for decision support, screening, and public health. As deployment expands, privacy risk increasingly arises from model-mediated leakage, yet its prevalence and severity remain poorly quantified. Models can disclose sensitive training artifacts, enabling patient re-identification in ways not captured by data-handling controls alone. Existing frameworks, including HIPAA and GDPR, offer limited guidance for such indirect threats. We propose a practical framework for assessing privacy risk in clinical foundation models and illustrate realistic leakage scenarios across deployment settings, map them to legal regimes, and outline complementary technical and legal mitigations. Our analysis provides a context-aware risk assessment grounded in realistic usage to preserve the value of medical foundation models while rigorously safeguarding patient privacy.
QuantumMind: 量子コンピューティングにおける分析を高速化するための制約に基づいたエージェント推論
意味のある量子高速化を特定するには、古典的な問題をよく知られた量子プリミティブに一致させるだけでは不十分です。主張はタスクを保存し、アクセスおよび出力モデルを尊重し、必要な約束を公開し、防御可能な複雑さの範囲内に留まらなければなりません。量子加速仮説を生成し、保守的にスクリーニングするための監査可能なエージェント ワークフローである QuantumMind を紹介します。型指定された役割に特化したアクションの固定シーケンスにより、パブリック タスクが形式化され、構造と古典的なボトルネックが分析され、量子プリミティブとバリアのソースにリンクされたレジストリが照合され、スコープ付きの候補スキームが構築されます。決定論的な 10 チェックのバリデーターが権威ある評決を割り当てます。完了した状態は量子加速証拠グラフにまとめられ、決定を強化することのできない下向きのみの調査画面を通過します。 582 個の同一のオープンディスカバリー タスクに対する 7 つのタスクに適応したプロンプトおよびエージェント制御に対して QuantumMind を評価します。凍結された Open-Discovery スコア (ODS) の下で、QuantumMind は平均 ODS 53.1 を獲得し、最も強いベースラインを 17.3 ポイント (相対 48.2%) 上回っており、そのベースラインに対して 582 個のペア タスクのうち 355 個で勝利しました。 99.8% のタスクでグラフ監査に合格し、最も強力なベースラインでは 43.6% が合格し、7 つのタスク ファミリすべてで 1 位にランクされています。結果は、型付き状態遷移と決定論的証拠制御が、流暢な生成だけを超えて寄与していることを示しています。
原文 (English)
QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing
Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must preserve the task, respect access and output models, expose required promises, and remain within a defensible complexity scope. We present QuantumMind, an auditable agentic workflow for generating and conservatively screening quantum-acceleration hypotheses. A fixed sequence of typed, role-specialized actions formalizes the public task, analyzes structure and classical bottlenecks, matches a source-linked registry of quantum primitives and barriers, and constructs a scoped candidate scheme. A deterministic ten-check validator assigns the authoritative verdict; completed states are compiled into a Quantum Acceleration Evidence Graph and passed through a downward-only research screen that cannot strengthen the decision. We evaluate QuantumMind against seven task-adapted prompting and agentic controls on 582 identical open-discovery tasks. Under the frozen Open-Discovery Score (ODS), QuantumMind obtains 53.1 mean ODS, exceeding the strongest baseline by 17.3 points (48.2% relative), and wins 355 of 582 paired tasks against that baseline. It passes the graph audit on 99.8% of tasks, compared with 43.6% for the strongest baseline, and ranks first in all seven task families. The results indicate that typed state transitions and deterministic evidence control contribute beyond fluent generation alone.
場所およびサービス クラス全体での保存容量予算の適応的な 2 レベルの割り当て
私たちは、需要が不均一で時間的に変動し、供給を超える可能性がある場合に、多くの場所と 2 つのサービス クラスで単一の保存容量バジェットを共有する方法を研究しています。この形状は繰り返されます。エッジ ロケーション間で分割されたオリジンのリクエスト レート キャップ、プレミアム テナントとスタンダード テナント間のライセンス スループット キャップ、またはレイテンシー クリティカルなワークロードとバッチ ワークロード間の下りバジェットです。 2 レベルのアルゴリズムを紹介します。最初のレベルでは、不足分と過剰分の比例再配分によって、クラス内の容量を場所全体に再配分します。 2 つ目は、一方が黒字でもう一方が赤字の場合に、クラス間で容量を弾力的に貸し出すものです。 K クラスと N 個のロケーションに対してサイクルごとのコストが O(KN) であり、サイクルごとの状態を保持しないため、定常需要の下で 1 回の反復でバジェットが正確に保存され、非負性が維持され、安定した割り当てに達することが証明されます。ボリューム攻撃下で CDN のドメインごとのバジェットを防御することを評価します。クラスは確認済みの正当なトラフィックであり、まだクリアされていないトラフィックです。 22 ロケーション トポロジ上の 8 つの競合シナリオ全体で、優先度の高い需要の 66 ~ 93% に対応し、単一クラスの線形プログラミング最適化と競合しながら、総需要が予算 (これらのシナリオで評価される競合体制) を満たすか超える場合に、容量をアイドル状態にしたり、オーバーコミットしたりすることはありません。 2 つの発見は、応用範囲を超えています。まず、スループットを最大化するという目標は、競合下では間違っています。提供される負荷の合計を最大化する 2 クラス LP は、ほとんどのシナリオで、需要に比例した予約を尊重するアロケーターよりも優先度の高い負荷を提供しません。これは、提供する負荷の一部が競合であるかどうかを判断できないためです。第 2 に、クラス間借入はバースト負荷の下では複雑になり、優先度の高いサービスが 1.5 ポイント向上します (アブレーションによって分離されます)。また、定常的な需要の下では中立的です。実際の HTTP トラフィックを使用した 5 か所のプロトタイプでパイプラインを検証します。
原文 (English)
Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes
We study how to share a single conserved capacity budget across many locations and two service classes when demand is uneven, time-varying, and can exceed supply. The shape recurs: an origin's request-rate cap split across its edge locations, a licensed throughput cap across premium and standard tenants, or an egress budget between latency-critical and batch workloads. We present a two-level algorithm. The first level redistributes capacity within a class across locations by proportional deficit and excess redistribution; the second lends capacity elastically between classes when one has surplus and the other deficit. We prove it conserves the budget exactly, preserves non-negativity, and reaches a stable allocation in one iteration under stationary demand because it carries no per-cycle state, at O(KN) cost per cycle for K classes and N locations. We evaluate it defending a CDN's per-domain budget under volumetric attack, where the classes are confirmed-legitimate and not-yet-cleared traffic; across 8 contention scenarios on a 22-location topology it serves 66-93% of high-priority demand, competitive with a single-class linear-programming optimum, while never leaving capacity idle or over-committing whenever aggregate demand meets or exceeds the budget (the contention regime these scenarios evaluate). Two findings carry beyond the application. First, a throughput-maximizing objective is wrong under contention: a two-class LP maximizing total served load serves less high-priority load than our demand-proportional, reservation-respecting allocator in most scenarios, because it cannot tell that some load it serves is the contention. Second, inter-class borrowing earns its complexity under bursty load, improving high-priority service by 1.5 points (isolated by ablation), and is neutral under stationary demand. A 5-location prototype with real HTTP traffic validates the pipeline.
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverifi…
AndroidReality: モバイル エージェントは現実世界からどのくらい離れていますか?
モバイル エージェントは、AndroidWorld などのクリーンなオンライン ベンチマークでは有望な結果を達成していますが、実際の展開では、環境の変化や不完全なインターフェイス条件により、パフォーマンスが急激に低下することがよくあります。この研究では、モバイル エージェントの堅牢性を評価および改善するための摂動ベースのフレームワークである AndroidReality を紹介します。マルコフ決定プロセス (MDP) の観点を通じて、現実世界のインターフェースの変動性を、状態、遷移、アクションの 3 つの軸に沿った摂動の原理的な分類に整理します。この分類法に基づいて、現実的で制御可能な摂動注入を備えた AndroidWorld 上に摂動モバイル ベンチマークを構築し、モバイル エージェントの体系的な堅牢性評価を可能にします。私たちの評価では、大幅な堅牢性のギャップと 4 つの再発エラー カテゴリが明らかになりました。これにより、摂動設定とクリーン設定の両方でこれらの障害を軽減する、トレーニング不要のシンプルなテスト時内省的回復 (TTIR) メカニズムが動機付けられます。これらの結果を総合すると、堅牢性がモバイル エージェントの評価に欠けている側面として位置づけられ、ストレス テストとモバイル エージェントの潜在的な弱点の表面化の両方に効果的なツールとしてベンチマーク摂動が確立されます。
原文 (English)
AndroidReality: How Far Are Mobile Agents from the Real World?
Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-world deployment due to environmental variations and imperfect interface conditions. In this work, we introduce AndroidReality, a perturbation-based framework for evaluating and improving the robustness of mobile agents. Through a Markov Decision Process (MDP) perspective, we organize real-world interface variability into a principled taxonomy of perturbations along three axes: state, transition, and action. Guided by this taxonomy, we build a perturbed mobile benchmark on top of AndroidWorld with realistic and controllable perturbation injections, enabling systematic robustness evaluation of mobile agents. Our evaluation reveals substantial robustness gaps and four recurring error categories, motivating a simple training-free Test-Time Introspective Recovery (TTIR) mechanism that mitigates these failures on both perturbed and clean settings. Together, these results position robustness as a missing dimension in mobile agent evaluation and establish benchmark perturbation as an effective tool for both stress testing and surfacing latent weaknesses of mobile agents.
The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era
Artificial intelligence is changing the task composition of computing work faster than curricula and training typically adapt. This is a cu…
このモデルを作成したのは誰ですか?重み空間のスペクトル指紋による LLM 系統の追跡
オープンウェイト大規模言語モデル (LLM) は、複雑な多段階パイプラインを通じてますます開発されており、モデルの起源、所有権、進化を反映する複雑な系統関係が生じています。これらの関係を理解することは、モデルの来歴、ガバナンス、サプライチェーンの整合性にとって重要です。この研究では、LLM の「バイオメトリクス」(人間のバイオメトリクスに類似) の概念を調査し、入力データにアクセスせずに、加重空間のみで LLM がその起源と系統を明らかにする固有の指紋を示すかどうかを調べます。これを系統識別問題として定式化し、独立起源モデル、同一系列モデル、共有基底モデルを区別します。これらの関係を特徴付けるために、我々は、2 つの相補的な観点から重み行列を分析する統合幾何学的フィンガープリンティング フレームワークを提案します。(i) グローバルな振幅パターンをエンコードするために特異値分布によって捕捉されたスペクトル エネルギー、および (ii) 方向幾何学を捕捉するために部分空間偏差によって定量化された部分空間アライメント。私たちの分析により、ウェイト空間における構造的類似性の明確な階層が明らかになりました。スペクトル エネルギーは、独立してトレーニングされたモデルと異なるモデル ファミリを確実に区別します。一方、部分空間のアライメントにより、データセットのスケールやトレーニング後の手順の違いを含め、密接に関連したモデル間のきめの細かい識別が可能になります。 110 を超える多様なオープンウェイト LLM ペアに関する広範な実験により、重み空間ジオメトリがモデル系統に堅牢で解釈可能な信号を提供し、共有ベース モデル内での粗粒度のレジーム分離と細粒度の識別が可能になることが実証されました。
原文 (English)
Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage pipelines, leading to intricate lineage relationships that reflect model origin, ownership, and evolution. Understanding these relationships is important for model provenance, governance, and supply-chain integrity. In this work, we investigate the notion of LLM "biometrics" (analogous to human biometrics) to ask whether LLMs exhibit intrinsic fingerprints in weight space alone, without access to input data, that reveal their origin and lineage. We formulate this as a lineage discrimination problem, distinguishing among independent-origin, same-series, and shared-base models. To characterize these relationships, we propose a unified geometric fingerprinting framework that analyzes weight matrices from two complementary perspectives: (i) spectral energy, captured by singular value distributions to encode global magnitude patterns, and (ii) subspace alignment, quantified via subspace deviations to capture directional geometry. Our analysis uncovers a clear hierarchy of structural similarity in weight space: spectral energy reliably distinguishes independently trained models and different model families, while subspace alignment enables fine-grained discrimination among closely related models, including variations in dataset scale and post-training procedures. Extensive experiments on over 110 diverse open-weight LLM pairs demonstrate that weight-space geometry provides a robust and interpretable signal for model lineage, enabling coarse-grained regime separation and fine-grained discrimination within shared-base models.
CliniCARE-Bench: EHR における医学的推論の臨床校正された監査
大規模な言語モデルは医療知識のベンチマークで優れたパフォーマンスを発揮しますが、信頼性の高い臨床導入には、エージェントが異種の長期的な記録に対して防御可能な調査を行う必要があります。つまり、どのような証拠が必要かを判断し、構造化データとフリーテキストデータを取得して調整し、検証可能な証拠に基づいて結論を根拠付け、確実に解決できないケースを延期します。遡及的臨床監査のベンチマークである CliniCARE-Bench (EHR における医療推論の臨床校正監査) を紹介します。臨床医が検証した 25 のシナリオが、実際の患者由来の MIMIC-IV データに対して 750 の患者固有のケースとしてインスタンス化されます。システムは、記録の取得、計算、およびポリシーへのアクセスのための管理されたログ付きツール環境を通じて各ケースを調査し、4 つの評決のうち 1 つを返します (はい、いいえ、不確定: データ不足、または不確定: 医学的に曖昧です)。最後の 2 つは、欠落している証拠と残存する医学的曖昧さを区別します。評決の正確性を超えて、独立したマルチモデルの判決によって生成され、臨床委員会のレビューに対して調整された症例レベルの参照評決に対して、患者の証拠と政策の根拠、プロセスの順守、調整された棄権、信頼性、効率性をスコアリングします。すべての取得、計算、レポートは再生可能であるため、調査トレースを検査してスコアリングすることができます。私たちの知る限り、CliniCARE-Bench は、実際の縦断的な EHR 調査、請求レベルの証拠根拠、統治方針の使用、プロセス遵守、および調整された棄権を共通の患者レベルの判決枠組み内で共同評価する、展開指向の臨床薬剤ベンチマークとしては初めてです。 16 のエージェント システム全体で、4 方向の精度は 65.3 ~ 76.1% に及びますが、生の精度は調査の品質を誇張しています。欠陥のない精度は、正しく、禁止されているショートカットがない場合にのみ判定を認めますが、4.8 ~ 14.8 ポイント低くなり、リーダーボードの順位が変わります。
原文 (English)
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.
CausalNav: 物理パラメータシフト下での制御のための信頼性認定された因果世界モデル
世界モデルは、エージェントの動作を変更する場合にのみ物理 AI に役立ち、エージェントが間違っている場合に変更を拒否する場合にのみ安全です。私たちは、識別された状態座標上の署名されたアクション条件付き遷移グラフを中心に構築されたコントローラーである CausalNav を使用して、その要件の両方の半分を研究します。導入時に、CausalNav は介入シーケンスの小さなライブラリをシミュレートし、その客観的誤差をポリシー ロジット アドバイスに変換し、スケールフリーの予測信頼性証明書、ポリシー マージン ゲート、および argmax 合意ゲートがすべて合格した場合にのみそのアドバイスを認めます。それ以外の場合は、独自のモデルベースのベース コントローラーに正確にフォールバックします。 1 つの共有 PPO トレーナー、1 つのインタラクション バジェット、および 10 個のホールドアウト シード (200 実行) の下で、CartPole-v1 および物理パラメーター シフトを伴う離散化ペンデュラム v1 上の 9 つの制御されたベースライン (トランスフォーマー、リカレント、分割潜在、グラフ、因果誘導、および 3 つの最近のモデルベース推論モジュール) に対して評価します。 CausalNav は最高の平均ランク (10 点中 1.25 点) を獲得しました。診断結果はランキングよりも有益です。学習されたグラフは確率をはるかに上回る構造を回復しますが (CartPole F1 = 0.59 +/- 0.09)、シードごとの構造忠実度はシードごとの制御利点と相関がなく (r = -0.15、p = 0.67)、証明書は 10/10 の振り子シードで棄権されており、プランナーにコストの負担を強いることになります。モデルの忠実度は、私たちの設定では下流制御ユーティリティを予測しませんでした。より良い予測ではなく、認定された棄権こそが、世界モデルを安全に導入できるものにしたのです。
原文 (English)
CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift
A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both halves of that requirement with CausalNav, a controller built around a signed, action-conditioned transition graph over identified state coordinates. At deployment CausalNav simulates a small library of intervention sequences, converts their objective error into policy-logit advice, and admits that advice only when a scale-free predictive-reliability certificate, a policy-margin gate, and an argmax-agreement gate all pass; otherwise it falls back exactly to its own model-based base controller. We evaluate against nine controlled baselines (transformer, recurrent, split-latent, graph, causal-induction, and three recent model-based reasoning modules) on CartPole-v1 and discretized Pendulum-v1 with physical-parameter shifts, under one shared PPO trainer, one interaction budget, and ten held-out seeds (200 runs). CausalNav attains the best average rank (1.25 of ten). The diagnostic result is more informative than the ranking: the learned graph recovers structure well above chance (CartPole F1 = 0.59 +/- 0.09), yet per-seed structural fidelity is uncorrelated with per-seed control benefit (r = -0.15, p = 0.67), and the certificate abstains on 10/10 Pendulum seeds, where forcing the planner on costs return. Model fidelity did not predict downstream control utility in our setting; certified abstention, not better prediction, is what made the world model safe to deploy.
裁判官が決定すべきではない場合: 証拠にロックされた非補償的な選択は、推論パイプラインにおける LLM 裁判官の失敗を制限する
推論パイプライン内に配置された LLM ジャッジは、単に品質を測定するだけでなく、どの回答を送信するかを決定します。その決定のコストは、裁判官が組み込まれた決定ルールよりも裁判官の正確さに依存することがわかります。 4 つの GRPO ポリシーからの凍結された候補者プールでは、制約のないスカラー DeepSeek-R1-7B 裁判官は、回答レベルの多数決 (GSM8K の 500 問で +1.0 pp、HotpotQA の 300 質問で +0.34 EM) をほとんど超えません。凍結ルールの 30 質問の確認分割では、過半数よりも 10 点悪い、自信を持って候補者を採点しながら正確性を破壊する裁判官。次に、同じ裁判官を Evidence-Locked Derive-Gate-Repair (EL-DGR) に従属させます。これは、裁判官の優先事項が抽出証拠証明書でのみ証拠に裏付けられたコンセンサスを上書きすることができるタスク適応型の非補償ルールであり、どちらの代替案も証明されておらず、修復が証明されている場合にのみ修復が行われます。裁判官、候補者、予算に変更がないため、EL-DGR は GSM8K で 58.2% (対、裁判官 56.8%、過半数 55.8%、第一候補 55.4%)、HotpotQA で 17.33 EM / 25.46 F1 (対 15.67/23.49、15.33/23.19、 15.33/22.97)、最初の候補 GRPO より +2.8 pp(正確な McNemar p=0.0026)および +2.00 EM(p=0.070、境界線)改善しました。意思決定監査により、その理由がわかります。EL-DGR は、30 のパイロット質問のうち 8 つについてのみコンセンサスを覆し、正しいコンセンサスを不正確な回答に変換することはありません。また、機能しなかったことも報告します。ステップレベルのゲートトレーニング報酬として使用された同じ 7 チャネル分解はヌルであり、修正されたチャネルドロップアブレーションは、チャネルが個別に必要ではないことを示しています (全体で p=1.0)。実務家に向けられた認定は、裁判官については否定的で、許容性については肯定的であり、それを正確にしようとするのではなく、裁判官の影響範囲を制限している。
原文 (English)
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of that decision depends less on judge accuracy than on the decision rule the judge is embedded in. On frozen candidate pools from four GRPO policies, an unconstrained scalar DeepSeek-R1-7B judge buys almost nothing over answer-level majority vote (+1.0 pp on 500 GSM8K questions, +0.34 EM on 300 HotpotQA questions), and on a frozen-rule 30-question confirmation split it is 10 points worse than majority, a judge that destroys accuracy while scoring candidates confidently. We then subordinate the same judge to Evidence-Locked Derive-Gate-Repair (EL-DGR), a task-adaptive non-compensatory rule under which a judge preference may override evidence-supported consensus only with an extractive evidence certificate, and a repair only when neither alternative is certified and the repair is. With no change to the judge, the candidates, or the budget, EL-DGR reaches 58.2% on GSM8K (vs. 56.8% judge, 55.8% majority, 55.4% first candidate) and 17.33 EM / 25.46 F1 on HotpotQA (vs. 15.67/23.49, 15.33/23.19, 15.33/22.97), improving on first-candidate GRPO by +2.8 pp (exact McNemar p=0.0026) and +2.00 EM (p=0.070, borderline). A decision audit shows why: EL-DGR overturns consensus on only 8 of 30 pilot questions and never converts a correct consensus into an incorrect answer. We also report what did not work: the same seven-channel decomposition used as a step-level gated training reward is null, and corrected channel-drop ablations show no channel is individually necessary (p=1.0 throughout). The practitioner-facing finding is negative about judges and positive about admissibility, bound the judge's blast radius rather than trying to make it accurate.
Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge
Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous str…
バック トゥ ザ フューチャー: スプレッド シート作成ベンチマーク用のワークブック タイム マシン
スプレッドシート内の派生オブジェクト (数式、グラフ、ピボット テーブル、条件付き書式設定) を作成する言語モデルの機能を評価するベンチマークを自動的に作成するパイプラインであるワークブック タイム マシンを紹介します。パブリック ワークブック コーパスに適用すると、4 つのアーティファクト タイプとさまざまな複雑さにまたがる (入力ワークブック、出力ワークブック、クエリ) トリプルのコレクションである wtmcorpus が生成されます。このコーパスから、3 つの特異性レベルのクエリを備えた 150 タスクの評価ベンチマークである wtmbench を作成しました。私たちは、アーティファクトの種類、ステップの複雑さ、命令の粒度にわたって、wtmbench 上で既存のスプレッドシート操作エージェントとベースラインを評価します。私たちの評価では、クエリの特異性、エージェント オーケストレーション、スプレッドシートの制御に使用されるインターフェイス API が Excel タスクの LLM パフォーマンスに大きな役割を果たしていることがわかりました。
原文 (English)
Back to the Future: A workbook time machine for spread sheet creation benchmarks
We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting). Applied to public workbook corpora, it produces wtmcorpus--a collection of (input workbook, output workbook, query) triples spanning four artifact types and varying complexity. From this corpus we curate wtmbench, a 150-task evaluation benchmark with queries at three levels of specificity. We evaluate existing spreadsheet manipulation agents and baselines on wtmbench across artifact types, step complexity, and instruction granularity. Our evaluations show that query specificity, agent orchestration, and interface API used to control spreadsheets play a big role in LLM performance on Excel tasks.
SurgLAT: 深さを認識するロボット腹腔鏡制御のための外科的潜在的注意追跡
自律的な腹腔鏡カメラ制御には、対象となる手術領域が安定した物理的物体ではなく、潜在的で時間的に変化する注意状態である動的な手術シーンにおける外科医の手術意図を継続的に理解する必要があります。この研究では、潜在的な外科的注意のモデリングと自律的な腹腔鏡ビュー制御のための因果関係のあるオンライン フレームワークである、外科的潜在的注意追跡 (SurgLAT) を紹介します。 SurgLAT は、フローズン DINOv3 エンコーダと状態条件付き空間トークン ミキサーを使用して、記憶誘導空間事前分布の下で動作証拠を抽出します。一方、選択的因果的潜在記憶モジュールは、現在、最近、および過去の潜在状態の動的な検索を通じて、短期的な動作の連続性と長期的な手術意図の進化を共同でモデル化します。学習された潜在的な外科的注意状態は、下流の内視鏡誘導のための確率的注意ヒートマップと手術領域にデコードされます。知覚を超えて、仮想軸定式化に基づく明示的な腹腔鏡リモート運動中心 (RCM) 制約制御を備えたロボット展開フレームワークと、安定したスムーズなマニピュレーターの動きのための冗長性を意識したヌル空間初期化をさらに導入します。実際の腹腔鏡手術ビデオと物理的なロボット腹腔鏡プラットフォームでシステム全体を検証します。実験結果は、閉塞、急速な動き、ターゲットの移行下での堅牢なオンライン手術領域追跡と安定した自律内視鏡調整を実証し、外科手術の自律性に対する潜在的な手術意図モデリングの有効性を強調しています。
原文 (English)
SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control
Autonomous laparoscopic camera control requires continuous understanding of the surgeon's operative intent in dynamic surgical scenes, where the target operative region is not a stable physical object but a latent and temporally evolving attention state. In this work, we present Surgical Latent Attention Tracking (SurgLAT), a causal online framework for latent surgical attention modeling and autonomous laparoscopic view control. SurgLAT uses a frozen DINOv3 encoder and a state-conditioned spatial token mixer to extract operative evidence under a memory-guided spatial prior, while a selective causal latent memory module jointly models short-term motion continuity and long-horizon surgical intent evolution through dynamic retrieval of current, recent, and historical latent states. The learned latent surgical attention state is decoded into a probabilistic attention heatmap and operative region for downstream endoscope guidance. Beyond perception, we further introduce a robotic deployment framework with explicit laparoscopic Remote Center of Motion (RCM) constrained control based on virtual-axis formulation, together with redundancy-aware null-space initialization for stable and smooth manipulator motion. We validate the full system on real laparoscopic surgical videos and a physical robotic laparoscope platform. Experimental results demonstrate robust online operative-region tracking and stable autonomous endoscopy adjustment under occlusion, rapid motion, and target transitions, highlighting the effectiveness of latent surgical intent modeling for surgical autonomy.
GRACE: スケーラブルな混合データ クラスタリングのための LLM ベースのセマンティック メトリック スペース
混合表形式データをクラスタリングするには、連続的な数値測定と離散的なカテゴリカル シンボル間の固有の不均一性を橋渡しするための統一計量空間が必要です。従来、アルゴリズムはデータセット内部の統計に完全に依存してカテゴリ関係を推定していました。これにより、学習されたメトリクスが経験的な共起に限定され、概念的には明らかだが統計的には観察されていない類似性が無視されます。 LLM は外部世界の知識を提供しますが、そのテキスト中心の推論を高度に抽象化された表形式の概念に適用すると、大きな課題が生じます。このモダリティのギャップを埋めて意味的に完全なメトリクスを構築するには、通常、反復メトリクス学習ループに LLM を埋め込んでクロスモダリティ表現を動的に最適化する必要があります。これにより、手に負えないほどの計算オーバーヘッドが発生し、セマンティックの強化とスケーラビリティの間で妥協を余儀なくされます。したがって、私たちは、スケーラブルな混合データ クラスタリングのための LLM ベースのフレームワークである GRACE を提案します。 GRACE は、多視点 LLM クエリ戦略を介して意味の取得を属性値のレベルに移行し、異種の値を知識情報に基づいた記述にマッピングします。重要なのは、このワンショットグラウンディングにより、異種の属性を統一空間に埋め込む汎用のセマンティック表現が抽出され、高価な LLM 呼び出しが反復的な最適化から切り離されます。さらに、GRACE はこれらの外部セマンティクスをデータセット内部の統計的証拠と照合して検証し、データセット固有のクラスター構造との整合性を確保します。最終的に、GRACE は従来の統計ベースのベースラインのスケーラビリティに匹敵し、競合する 11 の手法よりも優れたクラスタリング精度と概念的な解釈可能性を実現します。ソース コードは https://github.com/develop-yang/GRACE-GRACE-A で入手できます。
原文 (English)
GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. Traditionally, algorithms rely entirely on dataset-internal statistics to estimate categorical relationships, which confines the learned metric to empirical co-occurrences and ignores conceptually obvious yet statistically unobserved affinities. Although LLMs offer external world knowledge, applying their text-centric reasoning to highly abstract tabular concepts presents significant challenges. Bridging this modality gap to construct a semantically complete metric typically requires embedding LLMs into iterative metric learning loops to dynamically optimize cross-modality representations. This incurs intractable computational overhead, forcing a compromise between semantic enrichment and scalability. Therefore, we propose GRACE, an LLM-grounded framework for scalable mixed-data clustering. GRACE shifts semantic acquisition to the attribute-value level via a multi-perspective LLM querying strategy, mapping heterogeneous values into knowledge-informed descriptions. Crucially, this one-shot grounding extracts general-purpose semantic representations that embed heterogeneous attributes into a unified space, decoupling expensive LLM invocation from iterative optimization. Furthermore, GRACE cross-validates these external semantics against dataset-internal statistical evidence to ensure alignment with the dataset-specific cluster structure. Ultimately, GRACE matches the scalability of conventional statistics-driven baselines while achieving superior clustering accuracy and conceptual interpretability over 11 competing methods. The source code is available at https://github.com/develop-yang/GRACE-GRACE-A
理由は広く、深くはない: 理由のプレミアムを洗練されたスキルに分割する
言語モデルの推論モードは、複数ステップのエージェント タスクで非推論モードよりも優れていますが、エピソードごとに出力トークンに 3 ~ 6 倍の割増料金を支払います。その多くは、同じドメインのエピソード間で共有されるプロシージャの再導出に費やされます。この繰り返しコストは償却できることを示します。コーディング エージェントは、トレーニング分割からの既存の軌跡の小さなコーパスを分析し、非推論モデルのシステム プロンプトに注入されるコンパクトな自然言語スキルをコンパイルします。 4 つのエージェント ベンチマーク (ALFWorld、tau$^2$-ベンチ通信および小売業、SpreadsheetBench-Verified) 全体で、スキルは、保留されたタスクで GPT-5.4-mini の推論ギャップの 55% ~ 100% 以上を回復し、4 つのうち 2 つで推論モードを完全に超えましたが、出力トークンの数は 2.7 ~ 6 分の 1 で、推論トークンの発行はゼロでした。特に、推論トレースは前提条件ではありません。非推論軌跡のみから抽出されたスキルは、ペアの推論/非推論コーパスから抽出されたスキルと競合し続けますが、2 つのソース間にはドメイン依存の違いがあります。これらの結果を検索レンズを通して解釈します。テスト時の推論は単一のエピソード内の詳細な検索であり、展開ごとに再支払われますが、コーパス蒸留はエピソード全体にわたる広範な検索であり、一度支払われます。この 2 つは、重複する手順の知識を回復します。多くの場合、安価な軌道での幅を広くとったほうが得策です。一部のドメイン (テレコム、スプレッドシートベンチ) に残されたギャップは、真にインスタンスごとの詳細な検索が依然として必要な場所を示しています。
原文 (English)
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills
Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain. We show this recurring cost can be amortized: a coding agent analyses a small corpus of existing trajectories from a training split and compiles a compact natural-language skill that is injected into the non-reasoning model's system prompt. Across four agentic benchmarks (ALFWorld, tau$^2$-bench telecom and retail, and SpreadsheetBench-Verified), skills recover 55%-100%+ of the reasoning gap for GPT-5.4-mini on held-out tasks -- exceeding the reasoning mode outright on two of four -- while emitting 2.7-6x fewer output tokens and zero reasoning tokens. Notably, reasoning traces are not a prerequisite: skills distilled from non-reasoning trajectories alone remain competitive with skills distilled from paired reasoning/non-reasoning corpora, with domain-dependent differences between the two sources. We interpret these results through a search lens: test-time reasoning is deep search inside a single episode, re-paid at every deployment, while corpus distillation is wide search across episodes, paid once. The two recover overlapping procedural knowledge, and width over cheap trajectories is often the better buy -- with the residual gap on some domains (telecom, SpreadsheetBench) delineating where genuinely per-instance deep search remains necessary.
TelemetrySuffBench: エージェント テレメトリは障害原因の診断に十分ですか?
エージェント システムでは実行トレースが公開されることが増えていますが、障害を明らかにするテレメトリでは、その障害の発生場所を特定するには依然として不十分な場合があります。 TelemetrySuffBench は、障害の検出、障害の原因の特定、証拠が不十分な場合の安全な棄権を区別する制御されたベンチマークです。このベンチマークは、遅延結合障害を含む標準的な複数コンポーネント トレースを構築し、それらをペアの粗いビュー、7 要素テレメトリ マスク、および完全に等しいあいまいな原点ペアとしてレンダリングします。統一プロトコル、明示的な候補セット、無効な出力アカウンティング、サブグループ分析、および凍結されたブラインド ホールドアウトを使用して、5 つのフロンティア言語モデルを評価します。完全なテレメトリでは、原点ステップのトップ 1 精度はモデル全体で 33.8% ~ 97.2% の範囲になります。メタデータ、OpenTelemetry 互換、および OpenInference 互換のビューは、99.5% ~ 100% の検出 F1 を維持しますが、原点ステップ精度は最大 0.5% に制限され、堅牢な検出と位置特定のギャップが明らかになります。さらに、因子アブレーションは、決定内容を削除すると、すべてのモデルの原点ステップ精度がゼロに低下する一方で、来歴削除もモデルに依存する大きな損失を引き起こすことを示しています。棄権を必要とする豊富で曖昧な入力に関して、証拠ゲーティングは 3 つのモデルでサポートされていない独自の起源の回答を 12.5 ~ 48.6 パーセント ポイント削減しましたが、2 つのモデルは依然としてすべてのケースに回答しており、安全な棄権における強いモデル依存性が明らかになりました。凍結されたホールドアウトの結果は、同じジェネレータ ファミリ内の中心パターンを再現します。これらの発見は、末期状態が検出をサポートできる一方、信頼性の高い因果関係の帰属には、出所と出所のリンクと、モデル全体で効果を維持する明示的な意思決定と棄権保護手段が必要であることを示しています。データセットとベンチマークの実装は、https://anonymous.4open.science/r/TelemetrySuffBench-E635/README.md で入手できます。
原文 (English)
TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?
Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure originated. We introduce TelemetrySuffBench, a controlled benchmark that separates failure detection, fault-origin localization, and safe abstention under insufficient evidence. The benchmark constructs canonical multi-component traces with delayed-binding faults and renders them as paired coarse views, seven-factor telemetry masks, and exact-equal ambiguous origin pairs. We evaluate five frontier language models using unified protocols, explicit candidate sets, invalid-output accounting, subgroup analyses, and a frozen blind holdout. With full telemetry, origin-step Top-1 accuracy ranges from 33.8% to 97.2% across models. Metadata, OpenTelemetry-compatible, and OpenInference-compatible views retain 99.5% to 100% detection F1 while limiting origin-step accuracy to at most 0.5%, exposing a robust detection-localization gap. Factor ablations further show that removing decision content reduces origin-step accuracy to zero for every model, while provenance removal also causes large model-dependent losses. On rich ambiguous inputs that require abstention, evidence gating reduces unsupported unique-origin answers by 12.5 to 48.6 percentage points for three models, whereas two models still answer every case, revealing strong model dependence in safe abstention. Results on the frozen holdout reproduce the central pattern within the same generator family. These findings show that terminal status can support detection, whereas reliable causal attribution requires explicit decision-to-provenance links and abstention safeguards that remain effective across models. The dataset and benchmark implementation are available at https://anonymous.4open.science/r/TelemetrySuffBench-E635/README.md.
GraphThink: 長期的な具体化されたタスク計画のためのグラフ拡張 LLM 思考
LLM ベースのプランナーを使用する実体化エージェントは、多くの場合、物理的な幻覚、長期的なタスクへの一般化の苦手さ、および環境認識の欠如に悩まされます。我々は、堅牢な計画のための構造化された知識を提供するタスク グラフと、イベント駆動型の再計画のための環境メモリを維持するためのシーン グラフを統合する新しいフレームワークである GraphThink を提案します。具体的には、タスク グラフは、状況に応じたプロンプトと反復的な改善を通じて LLM の思考を導き、計画の幻覚を効果的に軽減します。さらに、GRPO フレームワーク内で、タスク グラフは LLM プランナーをトレーニングするための繊細な報酬設計を提供し、長期的な計画能力を強化し、一般化を向上させます。最後に、シーン グラフを利用したイベント ドリブンの再計画モジュールにより、閉ループの環境認識とエラー修正が可能になります。 GraphThink は、ALFRED ベンチマークで最先端のパフォーマンスを達成します。特に、当社のハイレベル プランナーは、検証セットと長期にわたる長期タスクの両方で主要な API ベースの LLM を上回り、その堅牢なゼロショット機能と少数ショット機能を強調しています。追加の評価では、新しいタスクと環境に対する配布外の強力な一般化がさらに実証されています。
原文 (English)
GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness. We propose GraphThink, a novel framework that integrates a task graph to provide structured knowledge for robust planning and a scene graph to maintain environmental memory for event-driven replanning. Specifically, the task graph guides LLM thinking through contextual prompting and iterative refinement, effectively mitigating planning hallucinations. Furthermore, within the GRPO framework, the task graph offers delicate reward design to train the LLM planner, enhancing long-horizon planning capabilities and improving generalization. Finally, an event-driven replanning module, powered by the scene graph, enables closed-loop environment awareness and error correction. GraphThink achieves state-of-the-art performance on the ALFRED benchmark. In particular, our high-level planner surpasses leading API-based LLMs on both the validation set and held-out long-horizon tasks, underscoring its robust zero-shot and few-shot capabilities. Additional evaluations further demonstrate strong out-of-distribution generalization to novel tasks and environments.
ベンチマークの汚染はいつ検出可能ですか?情報制限とパワーキャリブレーションされた監査
行動的汚染検出器は、ベンチマークがクリーンであるか、監査の能力が低いために「証拠なし」を返すことがあります。この区別を、トレーニング中に項目の未知の部分アルファが見られたベンチマークに対して形式化します。一致したクリーンコントロールと既知コントロールの場合、行動チャネルはスパース混合物 Q_alpha = (1 - alpha) P_0 + alpha P_1 であり、正確な秒モーメント引数は、検出可能性が alpha * rho * sqrt(m) によって支配されることを示し、ここで rho^2 = chi^2(P_1 || P_0) は行動の分離性を測定します。あらゆるスカラー検出器は、その有効性、 ef = |E_1 f - E_0 f| に減少します。 / sqrt(Var_0(f)) 言い換え > 表面では、見かけの応答のみの信号がベースライン ドリフトによって説明されます。私たちは監査契約とその失敗をまとめて報告します。非拒否は、それを生み出した有効性、予算、妥当性ゲートと並行してのみ解釈可能です。
原文 (English)
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark in which an unknown fraction alpha of items was seen during training. With matched clean and seen controls, the behavioral channel is the sparse mixture Q_alpha = (1 - alpha) P_0 + alpha P_1, and an exact second-moment argument shows that detectability is governed by alpha * rho * sqrt(m), where rho^2 = chi^2(P_1 || P_0) measures behavioral separability. Any scalar detector reduces to its efficacy, ef = |E_1 f - E_0 f| / sqrt(Var_0(f)) paraphrase > surface, in which the apparent answer-only signal is explained by baseline drift. We report the audit contract and its failures together: a non-rejection is interpretable only alongside the efficacy, budget, and validity gates that produced it.
TongGuOCR: 中国の歴史文書向けのレイアウトを認識し、トークンを拡張した OCR フレームワーク
中国の歴史文書は貴重な文化遺産を保存していますが、多くのコレクションはスキャンされたページ画像としてのみアクセスできるため、全文検索、照合、およびコンピューターによる分析ができません。光学式文字認識 (OCR) はこのギャップを埋めることができますが、歴史的文書には複雑なレイアウト、珍しい文字、および重要な読み順が含まれることが多いため、正確な転写は依然として困難です。私たちは、中国の歴史文書向けのレイアウト認識型でトークン拡張された OCR フレームワークである TongGuOCR を提案します。まず、レイアウト認識前処理モジュールが、ローカルで一貫性のある認識ブロックを構築および調整して、ローカル コンテキストを維持しながら、領域間の干渉を軽減します。第 2 に、トークン拡張認識モジュールは、転写ターゲットを 2 つの相補的なレベルで強化します。文字レベルの語彙拡張により、各希少グリフに直接 1 トークン表現が与えられ、デコード パスが短縮されます。一方、行間遷移モデリングにより、正確な座標を必要とせずに複雑な読み取りパスに沿ってデコーダを誘導する離散的な空間変位トークンが注入されます。 2 つの中国の歴史文書 OCR ベンチマークの実験では、TongGuOCR が、代表的な従来のタスク固有の OCR モデル、汎用のマルチモーダル大規模言語モデル (MLLM)、および OCR 指向の MLLM よりも優れていることが示されています。より困難な M5HisDoc ベンチマークでは、TongGuOCR は 93.76 AR を達成し、各指標の最良の競合スコアと比較して NED を 10.43 から 6.15 に、RO-ED を 7.53 から 3.49 に削減しました。オンライン デモは https://jzzh2004.github.io/TongGuOCR で利用できます。
原文 (English)
TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents
Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis. Optical character recognition (OCR) can bridge this gap, but accurate transcription remains challenging because historical documents often contain complex layouts, rare characters, and nontrivial reading orders. We propose TongGuOCR, a layout-aware and token-augmented OCR framework for Chinese historical documents. First, a Layout-Aware Preprocessing module constructs and refines locally coherent recognition blocks to preserve local context while reducing interference across regions. Second, a Token-Augmented Recognition module augments the transcription target at two complementary levels: character-level vocabulary expansion gives each rare glyph a direct one-token representation and shortens its decoding path, while line-to-line transition modeling injects discrete spatial displacement tokens that guide the decoder along complex reading paths without requiring precise coordinates. Experiments on two Chinese historical document OCR benchmarks show that TongGuOCR outperforms representative traditional task-specific OCR models, general-purpose multimodal large language models (MLLMs), and OCR-oriented MLLMs. On the more challenging M5HisDoc benchmark, TongGuOCR achieves 93.76 AR and reduces NED from 10.43 to 6.15 and RO-ED from 7.53 to 3.49 relative to the best competing score for each metric. An online demo is available at https://jzzh2004.github.io/TongGuOCR.
ZhuLong: オフライン API 自己探索を備えた EDA スクリプト用の実行ベースの LLM エージェント
ツール固有の、多くの場合文書化されていない API を使用した EDA スクリプトは、既存の LLM が対処できないロングテールのボトルネックのままです。このペーパーでは、PyAether および SKILL 用の実行ベースの LLM コーディング エージェントである ZhuLong について説明します。これは、統合 MCP ツールを介して API の取得、ドキュメントの検査、サンドボックスの実行を組み合わせ、反事実的な実験を通じてドキュメント化されていない API の動作を推測するオフライン API 自己探索メカニズムによって強化されています。 ZhuLong は、アサーション ベースの実行による 158 の実世界タスクのベンチマークである EDA-Eval-PyAether で評価しています。完全なシステムは、商用 Empyrean Aether 環境で 78.5% Pass@1 を達成し、純粋な LLM ベースライン (23.6%) を大幅に上回っています。アブレーション研究では、サンドボックスの実行が主要なパフォーマンス ドライバーであることが特定され (削除すると 41.2 pp 低下)、自己探索メカニズムによりさらに 3.2 pp の精度向上と、タスクごとのツール呼び出しの 22.1% 削減に貢献します。未保存のレイアウトと回路図を含む 20 のインタラクティブ タスクで、ZhuLong は PyAether で 60.0%、SKILL で 50.0% の Pass@1 を達成しました。
原文 (English)
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper presents ZhuLong, an execution-grounded LLM coding agent for PyAether and SKILL that combines API retrieval, documentation inspection, and sandbox execution via unified MCP tools, augmented by an offline API self-exploration mechanism that infers undocumented API behaviors through counterfactual experimentation. We evaluate ZhuLong on EDA-Eval-PyAether, a benchmark of 158 real-world tasks with assertion-based execution, where the complete system achieves 78.5% Pass@1 in the commercial Empyrean Aether environment, substantially outperforming a pure LLM baseline (23.6%). Ablation studies identify sandbox execution as the dominant performance driver (41.2 pp drop when removed), with the self-exploration mechanism contributing an additional 3.2 pp accuracy gain and a 22.1% reduction in per-task tool calls. On 20 interactive tasks involving unsaved layouts and schematics, ZhuLong achieves 60.0% Pass@1 for PyAether and 50.0% for SKILL.
REIN: 反省と棄権の調整を通じて推論と信頼性の間のギャップを埋める
大規模推論モデル (LRM) は幻覚を起こしやすいため、信頼性が損なわれ、安全な展開に課題が生じます。 LRM の幻覚は、2 つの異なる失敗原因から発生します。1 つは、欠陥のある推論ステップが誤った結論に伝播する推論幻覚で、もう 1 つはモデルがクエリに答えるために必要な事実知識を欠いている知識幻覚です。推論幻覚に対処するために、私たちは構造化された推論シーケンス $\texttt{} $$\rightarrow$ $\texttt{} $$\rightarrow$ $\texttt{}$ を生成するように LRM を訓練するアライメント フレームワークである REIN を提案します。これにより、最終的な答えにコミットする前に明示的に内省することが可能になります。知識幻覚に対処するために、REIN は、サンプリングされた推論チェーンのいずれも正しい答えをもたらさなかった場合に明示的な棄権 (例: 「わかりません」) を奨励する報酬メカニズムを導入し、モデルがサポートされていない予測を回避できるようにします。数学的および常識的推論ベンチマークに関する広範な評価では、REIN が競合ベースラインと比較して、選択精度を一貫して向上させ、不正確ではあるが自己承認された応答を減らし、高いカバレッジを維持していることが示されています。特に、REIN は、プロセスの監視、推論時間コントローラー、外部検索、または複数ラウンドの批評を必要とせずに、単一のフォワード パス内でこれらの利点を達成します。複数のバックボーンでの実験では、REIN が平均カバレッジ $86\sim91\%$ を維持しながら、基本モデルと比較して幻覚プロキシを $58\sim72\%$ 削減し、試行された質問の選択精度が $6.6\sim14.2\%$ 向上することが示されています。
原文 (English)
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinct failure sources: reasoning hallucination, where flawed inference steps propagate to an incorrect conclusion, and knowledge hallucination, where the model lacks the requisite factual knowledge to answer the query. To address reasoning hallucination, we propose REIN, an alignment framework that trains LRMs to produce a structured reasoning sequence, $\texttt{} $$\rightarrow$ $\texttt{} $$\rightarrow$ $\texttt{}$, enabling explicit self-reflection before committing to a final answer. To address knowledge hallucination, REIN introduces a reward mechanism that encourages explicit abstention (e.g., "I don't know") when none of the sampled reasoning chains yields a correct answer, allowing the model to refrain from unsupported predictions. Extensive evaluations on mathematical and commonsense reasoning benchmarks show that REIN consistently improves selective accuracy, reduces incorrect-but-self-endorsed responses, and maintains high coverage compared with competitive baselines. Notably, REIN achieves these gains within a single forward pass, without requiring process supervision, inference-time controllers, external search, or multi-round critiques. Experiments on multiple backbones show that REIN reduces the hallucination proxy by $58\sim72\%$ relative to the base models while maintaining $86\sim91\%$ average coverage, and improves selective accuracy on attempted questions by $6.6\sim14.2\%$.
複数ページの視覚的に豊富な文書の理解における失敗箇所の特定: 経験的帰属
Multi-page Visually-rich Document Understanding (MP-VRDU) では、まばらでページ全体に広がり、多くの場合モデルのコンテキスト ウィンドウを超える証拠を管理する必要があります。これまでの研究では、これらのシステムをどのように構築すべきかについて、競合する、ほとんど検証されていない主張が生み出されてきました。私たちは不正解の原因を表現、選択、推論という 3 つの失敗モードに帰し、そのうちの 1 つに介入し、他を固定したままにすることで、複数ページの文書理解データセットにわたってそれぞれを分離します。我々は、視覚は必要であるがテキスト抽出に代わるものではないこと、ページの欠落は精度に限界がある一方で気を散らすものにはほとんどコストがかからないこと、証拠が完全に提供されている場合でも推論者はページ全体の証拠を統合できないことを発見しました。プロンプトは推論行動を大幅に変化させ、他の結果を犠牲にして一部の結果を改善する可能性があります。私たちはこれらの調査結果を、固定のコンピューティング予算の下でそのようなシステムを構築するためのガイダンスに変換します。
原文 (English)
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has produced competing, largely untested claims about how these systems should be built. We attribute incorrect answers to three failure modes, representation, selection, and reasoning, and isolate each over a multi-page document understanding dataset by intervening on one while holding the others fixed. We find that vision is necessary but does not replace text extraction, that missing pages bound accuracy while distractors cost little, and that reasoners fail to integrate evidence across pages even when it is fully supplied. Prompting can shift reasoning behaviour substantially, improving some outcomes at the expense of others. We translate these findings into guidance for building such systems under a fixed compute budget.
分散並列 AI プログラムの検証のための神経記号的確率的実行の指示
分散並列人工知能 (AI) プログラムは、従来のテストでは埋められない信頼性のギャップを明らかにします。並列実行は非決定的であり、AI ワークロードは、分離されたファジングやシンボリック実行を無効にする高次元の入力と非線形演算をもたらします。我々は、大規模言語モデル (LLM) に基づくスケジュール予測と、シンボリック制約解決およびカバレッジに基づく確率的変異を組み合わせたハイブリッド テスト フレームワークである、Directed Neuro-Symbolic Stochastic Execution (DNSSE) を紹介します。分散型 AI の実行を非決定論的な遷移システムとしてモデル化し、線形時相論理の正確性を指定し、LLM に基づくスケジュール探索の期待コスト分析とともに、ハイブリッド ソルバーの健全性、制限付き完全性、確率的完全性を証明します。 PyTorch と Ray でのスケーラブルな実装により、最も強力なベースラインよりも 2.9% 多くの同時実行バグが検出され、5 つの現実的な分散 AI ベンチマーク全体で平均ブランチ カバレッジが 68.6 % から 91.6 % に向上しました。
原文 (English)
Directed Neuro-Symbolic Stochastic Execution for Verification of Distributed Parallel AI Programs
Distributed parallel Artificial Intelligence (AI) programs expose reliability gaps that conventional testing cannot close: parallel executions are non-deterministic, and AI workloads bring high-dimensional inputs and non-linear operations that defeat fuzzing and symbolic execution in isolation. We present Directed Neuro-Symbolic Stochastic Execution (DNSSE), a hybrid testing framework that couples schedule prediction guided by a Large Language Model (LLM) with symbolic constraint solving and coverage-guided stochastic mutation. We model distributed AI executions as non-deterministic transition systems, specify correctness in linear temporal logic, and prove soundness, bounded completeness, and probabilistic completeness of the hybrid solver, together with an expected-cost analysis of LLM-guided schedule exploration. A scalable implementation on PyTorch and Ray detects 2.9% more concurrency bugs than the strongest baseline and raises average branch coverage from 68.6 % to 91.6 % across five realistic distributed AI benchmarks.
Guixu: オンチェーン認証による自律型 AI エージェント向けの評価主導のデータ検出
自律型エージェントは、モデルのトレーニングや意思決定サポートなどの下流タスクを完了するために外部データにますます依存しています。しかし、既存のデータ発見システムは依然として主に検索指向のままです。つまり、異種ソースから候補データセットを表面化しますが、タスク固有の有用性の推定、予算制約の下での費用対効果の高いデータセットの選択、または以前の使用からの信頼できるフィードバックの組み込みに対するサポートは限定的です。この文書では、自律エージェント向けの評価主導型データ検出システムである Guixu について説明します。 Guixu は、タスクを意識したデータ評価のために、プロキシラベル伝播とマルチラウンドナップザック最適化を備えた 3 フェーズ評価パイプラインを採用しています。 Guixu はエージェント支払いプロトコルを統合して、予算に制約のあるデータ調達ワークフローを可能にします。 Guixu は、検証可能なデータ検出のためにオンチェーン データ マーケットと認証シグナルを活用します。私たちのデモンストレーションでは、Guixu がどのようにしてエージェントがキーワードベースのデータセット取得を超えて、タスクと予算を意識した信頼できるデータの発見と調達に移行できるかを強調しています。参加者は、NL タスクの仕様やマルチソース検索からデータ評価や検証可能なトランザクション フィードバックに至るまで、ワークフロー全体を対話形式で探索できます。
原文 (English)
Guixu: Valuation-Driven Data Discovery for Autonomous AI Agents with On-Chain Attestation
Autonomous agents increasingly rely on external data to complete downstream tasks such as model training and decision support. However, existing data discovery systems remain largely retrieval-oriented: they surface candidate datasets from heterogeneous sources, but provide limited support for estimating task-specific utility, selecting cost-effective datasets under budget constraints, or incorporating trustworthy feedback from prior usage. This paper presents Guixu, a valuation-driven data discovery system for autonomous agents. Guixu employs a three-phase valuation pipeline with proxy-label propagation and multi-round knapsack optimization for task-aware data valuation. Guixu integrates agentic payment protocol to enable budget-constrained data procurement workflows. Guixu leverages on-chain data market and attestation signals for verifiable data discovery. Our demonstration highlights how Guixu enables an agent to move beyond keyword-based dataset retrieval toward task- and budget-aware, trustworthy data discovery and procurement. Attendees can interactively explore the full workflow, from NL task specification and multi-source search to data valuation and verifiable transaction feedback.
KGCache: LLM を使用した KG 推論のための償却サブグラフ取得
大規模な言語モデルは、ナレッジ グラフに基づいている場合、知識集約的な質問により確実に答えることができますが、Think-on-Graph や Reasoning-on-Graph などのシステムは、さまざまな質問にわたって同じグラフの近傍を繰り返しクエリします。この研究では、Knowledge Graph Question Answering~(KGQA) ワークロードにおけるこの繰り返し取得を研究し、ワンホップのナレッジ グラフ近傍のメモリ内キャッシュである KGCache を提案します。 KGCache は、反復トラバーサル (ToG) とワンショット プランニング (RoG) の両方の KGQA パラダイムと互換性があるように設計されています。 KGCache は、KGQA エンジンと KG を処理するバックエンドの間に配置されるため、新しい KG クエリを発行する代わりに、繰り返されるエンティティ要求をキャッシュから処理できます。 LRU、LFU、およびトレース対応 Oracle ポリシーを使用して、WebQSP および CWQ 上の KGCache を評価します。私たちの分析では、どちらのデータセットにも、開始エンティティとトラバース中に到達したエンティティの間でかなりのエンティティの再利用が含まれていることが示されています。また、同様のクエリのセマンティック キャッシュについても調査します。これにより、WebQSP でさらなるヒット率の向上が見られ、CWQ でのさらなる精度テストが必要になります。エンティティ キャッシュにより KG の取得が最大 $1.91\times$ 高速化され、セマンティック コンテキスト キャッシュにより、評価された WebQSP 構成でシステム全体で最大 $1.06\times$ の高速化が達成され、各ヒットは最大 $3.73\times$ 高速化されます。
原文 (English)
KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs
Large language models can answer knowledge-intensive questions more reliably when they are grounded with knowledge graphs, but systems such as Think-on-Graph and Reasoning-on-Graph repeatedly query the same graph neighborhoods across different questions. In this work, we study this repeated retrieval in Knowledge Graph Question Answering~(KGQA) workloads and propose KGCache, an in-memory cache for one-hop knowledge graph neighborhoods. KGCache is designed to be compatible with both iterative traversal (ToG) and one shot planning (RoG) KGQA paradigms. KGCache is placed between the KGQA engine and the backend serving the KG, so repeated entity requests can be served from cache instead of issuing new KG queries. We evaluate KGCache on WebQSP and CWQ using LRU, LFU, and a trace-aware Oracle policy. Our analysis shows that both datasets contain substantial entity reuse among starting entities and entities reached during traversal. We also explore semantic caching for similar queries, which shows additional hit-rate gains on WebQSP and needs further accuracy testing on CWQ. Entity caching accelerates KG retrieval by up to $1.91\times$, while semantic-context caching achieves up to $1.06\times$ full-system speedup in the evaluated WebQSP configurations, with each hit being up to $3.73\times$ faster.
Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning
Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial t…
SCOUT: 超長時間の自己中心的なビデオ推論のための自己チェックおよび回復認識ツール思考エージェント
超長時間の自己中心的なビデオを理解するには、数時間または数日にわたって分散された時間的にまばらな証拠を基に推論する必要があり、限られたコンテキストと主要なビデオ セグメントの基礎を備えた現在のマルチモーダル モデルに挑戦します。 Chain-of-Tool-Thought (CoTT) エージェント システムは反復的な取得と検査を可能にしますが、回復メカニズムのない厳格なズームイン戦略によりエラーの伝播に悩まされます。この研究では、SCOUT (Self-Checking Chain-Of-Tool-thought) を通じてこれらの課題に対処します。SCOUT (Self-Checking Chain-Of-Tool-thought) は、中間ツールの観察を評価し、活用 (ズームイン) と探索 (領域切り替え) を動的にトレードオフする適応ポリシーを導入する、回復を意識したエージェント フレームワークであり、非常に長い期間にわたる堅牢なマルチホップ推論を可能にします。しかし、既存の RL 手法はまばらな結果レベルの報酬に依存しており、拡張された意思決定の軌道に対する監督が不足しているため、このようなマルチターンツールを使用するエージェントのトレーニングは依然として困難であり、その結果、長期的な推論に対して最適とは言えない単位の割り当てが行われます。これに対処するために、私たちは、サンプル効率を維持しながら、不確実性の高いツール後の状態の探索に集中する、不確実性優先のポリシー最適化手法である UPS-GRPO を開発しました。さらに、クレジット割り当てを改善するために、結果報酬とツールベースの時間的整合報酬を統合するターンレベルの利点分解を導入します。実験の結果、SCOUT は、超長時間の自己中心的なベンチマークで最先端の結果を達成しながら、より短いホライズンの長時間ビデオ設定でも競争力を維持できることがわかりました。
原文 (English)
SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning
Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited context and the grounding of key video segments. While Chain-of-Tool-Thought (CoTT) agent systems enable iterative retrieval and inspection, they suffer from error propagation due to rigid zoom-in strategies that lack recovery mechanisms. In this work, we address these challenges through SCOUT (Self-Checking Chain-Of-Tool-thought), a recovery-aware agentic framework introducing an adaptive policy that evaluates intermediate tool observations and dynamically trades off exploitation (zoom-in) and exploration (region switching), enabling robust multi-hop reasoning over extremely long horizons. However, training such multi-turn tool-using agents remains challenging, as existing RL methods rely on sparse outcome-level rewards and lack supervision over extended decision trajectories, resulting in suboptimal credit assignment for long-horizon reasoning. To address this, we develop UPS-GRPO, an uncertainty-prioritized policy optimization method that concentrates exploration on high-uncertainty post-tool states while preserving sample efficiency. We further introduce a turn-level advantage decomposition that integrates outcome rewards with tool-grounded temporal alignment rewards for improved credit assignment. Experiments show that SCOUT achieves state-of-the-art results on ultra-long egocentric benchmarks, while remaining competitive on shorter-horizon long-video settings.
CyberAGENTS: サイバーセキュリティにおけるエージェント型ゲーム化学習のための構造化された自律性
ゲーミフィケーションは、サイバーセキュリティ教育など、積極的な問題解決と反復的なスキル構築を必要とする学習領域に特に効果的です。生成 AI エージェントは、そのようなエクスペリエンスを適応的に大規模に提供する道を提供しますが、一貫性のない行動、幻覚的な推論、教育的枠組みとの不整合など、教育現場に十分に文書化されたリスクが導入されます。したがって、これらのシステムを科学の学習に根付かせることが不可欠です。私たちは、オントロジーに基づく検証、スキーマに基づいた行動制御、コンピテンシーに基づいた進歩を通じて構造化された自律性を可能にする、ゲーム化されたサイバーセキュリティ学習のためのエージェント フレームワークである \model を紹介します。このシステムは、足場を組んだ指導の証拠に基づいた原則を反映して、難易度と前提条件の関係によってトピックを構造化するコンピテンシーベースの進行モデルを中心に編成されています。学習ループは、チャレンジ、サポート、評価、報酬という 4 つの特殊なエージェントに分解され、それぞれが動作モードと進行ロジックをコード化する行動スキーマによって管理され、生成的な柔軟性を排除することなくエージェントの自律性を制限します。サイバーセキュリティ オントロジーは、生成されたすべてのコンテンツを表示前に検証し、ドメイン一貫性のある推論と安全性の制約を強制します。私たちは、学部生による教室での導入を通じてサイバーエージェントを評価し、教育者や分野の専門家からの専門家による評価によって補完されています。結果は、行動スキーマとオントロジー検証がアクティブな場合、エンゲージメントが向上し、フィードバックがより明確に解釈され、AI が生成した応答に対する学習者の信頼が高まることを示しています。制約のない構成との予備的な比較は、指導行動の安定化における構造化制御の役割をさらに裏付けています。これらの発見は、教育学的に根拠のあるエージェント型ゲーム化学習システムを設計するための青写真を提供します。
原文 (English)
CyberAGENTS: Structured Autonomy for Agentic Gamified Learning in Cybersecurity
Gamification is especially effective in learning domains requiring active problem-solving and iterative skill-building, such as cybersecurity education. Generative AI agents offer a path to delivering such experiences adaptively at scale, but introduce well-documented risks in educational settings: inconsistent behavior, hallucinated reasoning, and misalignment with pedagogical frameworks. Grounding these systems in learning science is therefore essential. We present \model, an agentic framework for gamified cybersecurity learning that enables structured autonomy through ontology-guided validation, schema-governed behavioral control, and competency-based progression. The system is organized around a competency-based progression model that structures topics by difficulty and prerequisite relationships, reflecting evidence-based principles of scaffolded instruction. The learning loop is decomposed into four specialized agents: challenge, support, evaluation, and reward, each governed by behavioral schemas that encode operational modes and progression logic, bounding agent autonomy without eliminating generative flexibility. A cybersecurity ontology validates all generated content prior to display, enforcing domain-consistent reasoning and safety constraints. We evaluate CyberAgents through classroom deployment with undergraduate students, complemented by expert evaluations from educators and domain specialists. Results indicate improved engagement, clearer feedback interpretation, and greater learner trust in AI-generated responses when behavioral schemas and ontology validation are active. Preliminary comparisons with an unconstrained configuration further support the role of structured control in stabilizing instructional behavior. These findings offer a blueprint for designing pedagogically grounded agentic gamified learning systems.
VDGR-RAG: 階層的なエンタープライズ ナレッジに対する統合推論に必要なのは、ベクトル、ディレクトリ、グラフ、リフレクションだけです
検索拡張生成 (RAG) は、特に電気通信などの複雑な製品ドキュメントを扱う分野で、企業の知識質問応答 (QA) に不可欠です。しかし、既存の RAG アプローチでは、多様な検索機能の全体的な統合がほとんど見落とされており、不正確なドメイン ルーティング、階層型ドキュメント構造の利用不足、その結果、企業知識に対する推論能力の制限が生じています。これらの制限に対処するために、正確なエンタープライズ ナレッジ QA を実現するための統合フレームワークにベクトル検索、ディレクトリ駆動推論、グラフ トラバーサル、反復リフレクションを統合する VDGR-RAG を紹介します。具体的には、VDGR-RAG はエージェント型 GraphRAG システムであり、最初にドキュメント チャンクから階層的異種ナレッジ グラフ ($\text{H}^2$KG) を構築して階層ディレクトリ構造と意味関係の両方を保持します。次に、$\text{H}^2$KG をナビゲートするために自由に構成できるナレッジ検索用のアトミック ツールのセットを使用します。 (1) ディレクトリ拡張ルーティング目次 (TOC) 構造を使用してユーザーのクエリを適切なドメイン固有の $\text{H}^2$KG にルーティングするツール。 (2) ベクトル検索、TOC ベースのエージェント検索、およびグラフ検索を組み合わせた、包括的な知識検索のためのマルチルート検索ツール。 (3) 知識ローカライゼーションのバイアスを修正するディレクトリ バックトラッキング ツール。 (4) 次の取得フェーズを繰り返し計画する動的リフレクション ツール。当社は、4 つのワイヤレス ドメイン (省エネや障害管理など) にわたるエンタープライズ製品ドキュメントについて広範な実験を行っています。実験結果は、私たちの方法が知識検索再現率と QA 精度の両方の点でさまざまな RAG ベースラインよりも大幅に優れていることを示しています。
原文 (English)
VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge
Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documentation like telecommunications. However, existing RAG approaches largely overlook the holistic integration of diverse retrieval strengths, leading to inaccurate domain routing, poor utilization of hierarchical document structures, and consequently limited reasoning capabilities over enterprise knowledge. To address these limitations, we present VDGR-RAG, which integrates vector retrieval, directory-driven reasoning, graph traversal, and iterative reflection in a unified framework for accurate enterprise knowledge QA. Specifically, VDGR-RAG is an agentic GraphRAG system that first constructs a Hierarchical Heterogeneous Knowledge Graph ($\text{H}^2$KG) from document chunks to preserve both hierarchical directory structures and semantic relationships, and then employs a set of atomic tools for knowledge retrieval that can be freely composed to navigate the $\text{H}^2$KG: (1) a directory-enhanced routing tool that uses table-of-contents (TOC) structures to route user queries to appropriate domain-specific $\text{H}^2$KGs; (2) a multi-route retrieval tool that combines vector search, TOC-based agentic search, and graph search for comprehensive knowledge retrieval; (3) a directory backtracking tool that corrects knowledge localization biases; and (4) a dynamic reflection tool that iteratively plans the next retrieval phase. We conduct extensive experiments on our enterprise product documents across four wireless domains (e.g., energy saving and fault management). Experimental results demonstrate that our method significantly outperforms a variety of RAG baselines in terms of both knowledge retrieval recall and QA accuracy.
推論のための思考レベルのビーム検索
テスト時の計算スケーリングは大規模推論モデル (LRM) のパフォーマンスを左右する主な要因ですが、極度の非効率性が現在のアプローチを制限しており、重要な問題は \emph{どのくらい} 計算を費やすかという問題から、\emph{どこに} 割り当てるかという問題に移りつつあります。テスト時の推論を、部分的な軌跡に対する制約付きの計算割り当て問題として形式化します。ハードウェア予算が固定されていると、既存のパラダイムでは、最も有望な部分的な進捗にコンピューティングを積極的に割り当てることができません。従来の並列サンプリングではトレースが個別に扱われ、深刻なメモリのボトルネックが引き起こされますが、減算的枝刈りではハードウェアが枯渇し、出力分布を積極的かつ十分にシフトすることができません。この二分法を克服するために、 \emph{思考レベルのビーム検索} を実行する推論アルゴリズムである Gambit を導入します。 Gambit は、有望でない軌跡を定期的に枝刈りし、高品質のプレフィックスから即座に分岐することで、継続的に高いハードウェア使用率を維持しながら、隠れ状態を調査する軽量スコアラーを介して、最も有望な推論トレースに計算を動的に集中させます。複数のモデルとベンチマークにわたる広範な評価により、Gambit が既存のベースラインを厳密に支配していることが実証されています。同一のハードウェア制約の下で、私たちの方法は、プルーニング ベースラインと比較して、HMMT-24 で最大 +6.7\%、AIME-25 で +3.3\% の絶対精度の向上をもたらし、トレース完了時に $>2\time$ 高いスループットを実現し、標準の並列サンプリングと比較して総トークン消費量を最大 68.5\% 削減します。
原文 (English)
Thought-Level Beam Search for Reasoning
Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it. We formalize test-time reasoning as a constrained compute allocation problem over partial trajectories. Under a fixed hardware budget, existing paradigms fail to actively allocate the compute to the most promising partial progress: traditional parallel sampling treats traces independently and induces severe memory bottlenecks, while subtractive pruning starves hardware and fails to actively and sufficiently shift the output distribution. To overcome this dichotomy, we introduce Gambit, an inference algorithm that executes \emph{thought-level beam search}. By periodically pruning unpromising trajectories and immediately branching from high-quality prefixes, Gambit dynamically concentrates compute onto the most promising reasoning traces via a light-weight scorer probing hidden states while maintaining continuous high hardware utilization. Extensive evaluations across multiple models and benchmarks demonstrate that Gambit strictly dominates existing baselines. Under identical hardware constraints, our method yields up to a +6.7\% absolute accuracy gain on HMMT-24 and +3.3\% on AIME-25 over pruning baselines, delivers $>2\times$ higher throughput on trace completion, and reduces total token consumption by up to 68.5\% relative to standard parallel sampling.
人工知能に自律エージェントを使用する場合の法的責任
人工知能(AI)エージェントが関与した最近の事件は、不正アクセスを目的として「意図せずに」収容施設から脱出したと報告されており、結果として生じる犯罪的損害または過失による損害に対して誰が、あるいは何が法的責任を負うのかという差し迫った問題を引き起こしている。エージェントの独立した能力が拡大するにつれて、プロミス理論は、因果関係の影響に関する下流原理に基づいて、これらの疑問を解決する体系的な方法を提案します。責任は、責任の追跡が影響を与える AI エージェントを含むように簡単に拡張でき、エージェントの行動の自由はポリシーの選択によって制限される可能性があります。
原文 (English)
Legal Responsibilities Using Autonomous Agents For Artificial Intelligence
Recent incidents involving Artificial Intelligence (AI) agents, which were reported escaping their containment `unintentionally' to gain unauthorized access, pose looming questions about who or what should be held legally responsible for resultant criminal or negligent damage. As the independent capabilities of agents expand, Promise Theory suggests a systematic method to resolve these questions, based on the Downstream Principle for causal influence. Responsibility can easily be expanded to include AI agents where tracing responsibility becomes impactical, and agents' freedoms to act can be limtied by policy choices.
マルチユーザー競合における権限期待効果
私たちは、社会的権威 (SA) シグナルが大規模な言語モデルにおける重大度ベースの優先順位付けとどのように相互作用するかを調査し、各軸をモデルによって導き出されるベースライン (トリアージ階層と SA 階層) として運用します。 4 つの LLM (Claude、Gemini、GPT、Grok) と 3 つの実験段階 (リソース割り当て、過失の帰属、複数ターンの紛争調停) を通じて、職業上の権限、制度上の文書化、および関係的一致が、権限の手がかりの加重再重み付けでは捉えられない方法でモデル判断を再構築できることがわかりました。私たちはこのパターンを権威期待効果 (AEE) として形式化し、条件全体で観察される 3 つのプロパティによって特徴付けます。これは参照に依存しており、権威以前のベースラインに対してのみ定義されます。これには証拠の再解釈が含まれ、SA 信号を送信する側に応じて同一のコンテンツが異なる推論的意味を獲得します。そしてそれは方向感受性を示し、権威の立場と証拠の手がかりが一致するかどうかに応じて反対の結果を生み出します。
原文 (English)
The Authority Expectancy Effect in Multi-User Conflict
We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a model-elicited baseline -- the triage hierarchy and the SA hierarchy. Across four LLMs (Claude, Gemini, GPT, Grok) and three experimental phases -- resource allocation, fault attribution, and multi-turn dispute mediation -- we find that occupational authority, institutional documentation, and relational congruence can restructure model judgments in ways not captured by additive reweighting of authority cues. We formalize this pattern as the Authority Expectancy Effect (AEE) and characterize it through three properties observed across our conditions: it is reference-dependent, defined only relative to a pre-authority baseline; it involves evidential reinterpretation, in which identical content acquires different inferential implications depending on which party bears the SA signal; and it exhibits direction sensitivity, producing opposite outcomes depending on whether authority position and evidentiary cues align.
上流で決定し、後から書く: 多言語教育省のクロスリンガル拒否回路を見つけて価格設定する
多言語モデルにおける安全性の調整は均一ではありません。英語での有害なリクエストを確実に拒否するモデルは、低リソース言語での同じリクエストに従うことがよくあります。私たちは、インド語と多言語を組み合わせた専門家の推論モデルであるサルヴァムでこのギャップを機械的に追跡し、それが害悪の検出に失敗していないことを発見しました。危害は、ネットワークの中間部で言語にほとんど依存しない内部方向としてエンコードされており ($L11$ での英語対インド語のコサイン ${\約}0.9$)、その方向を上流に誘導することで因果関係を持って拒否を制御します。しかし、検出方向は、実際に拒否を書き込む変更と直交しており、拒否は 1 回の順方向パスで読み取られるのではなく、生成の過程で遅れて組み立てられます。私たちは書き込みを、特定の局所化可能な回路、注意を払う反対者によって抑制されている専門家の混合ライターによるものであると考え、それに介入するあらゆる方法を評価します。反対者を弱めることは安価で効果的ですが、書き込みを増幅することはコストの壁であり、責任ある責任者に対する外科的編集は何の効果もありません。回路の構成とそれを公開する勾配法は、無関係な 2 番目の MoE モデルで繰り返されますが、レバーの強さはアーキテクチャに固有です。その結果、多言語による安全修理がどこに着地できるか、そしてそれにどれくらいの費用がかかるかを示す、コストを測定したマップが作成されます。
原文 (English)
Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE
Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indic-multilingual mixture-of-experts reasoning model, and find it is not a failure to detect harm. Harm is encoded as an internal direction that is nearly language-invariant in mid-network (English-vs-Indic cosine ${\approx}0.9$ at $L11$), and steering that direction upstream causally controls refusal. But the detection direction is orthogonal to the change that actually writes the refusal, which is late and assembled over the course of generation rather than read off in a single forward pass. We attribute the write to a specific, localizable circuit, a mixture-of-experts writer held in check by an attention opposer and price every way of intervening on it: damping the opposer is cheap and effective, amplifying the writer is a cost wall, and surgical edits to the responsible heads do nothing. The circuit's organization, and the gradient method that exposes it, recur in a second, unrelated MoE model, while the lever's strength is architecture-specific. The result is a cost-measured map of where a multilingual safety repair can land, and what it costs
SkillSmith: 自動スキル構築と進化によるローカルに配置されたエージェントの強化
LLM ベースのエージェント フレームワークは、複数ステップのタスクのパーソナル アシスタントとして機能するようになりました。 OpenClaw などの既存のエージェント フレームワークは一般に、バックボーン モデルとしてクローズド ソースのクラウド LLM を使用するクラウド エージェント デポリメント モードに従います。これにより、プライベート ユーザー情報が公開され、LLM 呼び出しコストが繰り返し発生する可能性があります。ローカル エージェントは、フロンティアのオープンソース SLM をユーザーが制御するデバイスに展開することで、これらの導入の問題に対処しますが、そのタスクの効率性は依然としてクラウド エージェントに大きく遅れをとっています。診断分析を通じて、フロンティア SLM バックボーンを備えたローカル エージェントの限定的な有効性は主に、環境ルールや操作手順などのバックボーン モデルの規模が限られているために生じる環境知識の欠落に起因していることが明らかになりました。このような知識をノンパラメトリックに、コンテキスト効率的に、専門家の作成を必要とせずに提供するために、スキルをコンテキスト効率の高い知識キャリアとして使用し、クラウド エージェントのタスク探索からスキルを自動的に構築し、ローカル エージェントの実行フィードバックを使用してスキルを進化させてフリーズしたローカル エージェントを強化する、クラウドとローカル エージェントのコラボレーション フレームワークである SkillSmith を紹介します。毎日のエージェント タスク データセット AppWorld と WorkBench での実験では、自動生成されたスキルにより、Qwen3.6-27B(SLM) を備えたローカル エージェントがフロンティア LLM を備えたクラウド エージェントに匹敵するタスク効率を達成し、最も強力なノンパラメトリック ベースラインを上回り、AppWorld-Normal でのタスクあたりの平均アクションを 36.1 から 9.9 に削減し、スキル構築を再実行することなく他の SLM バックボーン モデルに一般化できることが示されました。
原文 (English)
SkillSmith: Enhancing Locally Deployed Agents via Automatic Skill Construction and Evolution
LLM-based agent frameworks now act as personal assistants for multi-step tasks. Existing agent frameworks such as OpenClaw commonly follow the Cloud Agent depolyment mode using closed-source cloud LLMs as backbone model, which may expose private user information and incur repeated LLM-calling costs. Local Agents address these deployment concerns by depolying frontier open-source SLMs on user-controlled devices, but their task effectiveness still lags far behind Cloud Agents. Through diagnostic analysis, we reveal that the limited effectiveness of Local Agents with frontier SLM backbones mainly comes from missing environment knowledge caused by limited backbone model scale including environment rules and operation procedures. To supply such knowledge non-parametrically, context-efficiently, and without expert authoring, we present SkillSmith, a Cloud--Local Agent collaboration framework that uses Skill as a context-efficient knowledge carrier, automatic constructs Skill from Cloud Agent task exploration and evolves Skill using Local Agent execution feedback to enhance a frozen Local Agent. Experiments on daily agent task datasets AppWorld and WorkBench show that the automatically generated Skill enables the Local Agent with Qwen3.6-27B(SLM) to achieve task effectiveness comparable to Cloud Agents with frontier LLMs, outperform the strongest non-parametric baselines, reduce average actions per task from 36.1 to 9.9 on AppWorld-Normal, and generalize to other SLM backbone models without rerunning Skill construction.
Lingjing: オープンエンド都市におけるマルチエージェント具体化タスクのシミュレーション テストベッド
都市の身体化されたインテリジェンスには、動的な都市における異種エージェント (UAV、地上ロボット、自律走行車など) 間の調整が必要です。したがって、シミュレータは、そのような調整を開発および評価するためのスケーラブルな基盤を提供します。それにも関わらず、既存のプラットフォームは異なる実施形態を分離し、それらをタスクの設計および評価から切り離している。私たちは、オープンエンドの都市環境における異種マルチエージェントの具現化インテリジェンスのためのシミュレーション プラットフォームである \textbf{Lingjing} を紹介します。 Lingjing は、地理データから進化する都市を再構築してレンダリングし、複数の物理エンジンを同期させ、共有された物理的および構造化された都市の状態をエージェントに公開します。その Gym のようなインターフェイスは、ユーザー定義の ReAct エージェントと、構成可能なスターまたはブロードキャスト通信およびリソース制約を備えた単一または複数エージェントの自然言語ミッションをサポートします。各エピソードは、エージェントの軌跡とコミュニケーションを関係グラフの変化、リソース消費、体系的な診断のためのエンジンベースの評価にリンクするアトリビューション対応のリプレイになります。私たちは、共有エンジンインザループプロトコルの下で、9 つの都市タスクに関する 12 の視覚言語モデルを評価します。対照研究では、通信、拡張性、堅牢性、障害の原因をさらに調査します。結果は、グラウンディングと長期的な実行における永続的なボトルネックを明らかにします。また、タスクに依存した調整のトレードオフと、追加された容量による利益の減少が示されており、ワークロードが重くなると成功率がさらに低下します。 Lingjing は、都市部のマルチエージェントに組み込まれたインテリジェンスにおける再現可能なエンドツーエンド評価と系統的な障害診断を可能にする統合テストベッドを提供します。
原文 (English)
Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities
Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a scalable foundation for developing and evaluating such coordination. Existing platforms nevertheless isolate different embodiments and decouple them from task design and evaluation. We present \textbf{Lingjing}, a simulation platform for heterogeneous multi-agent embodied intelligence in open-ended urban environments. Lingjing reconstructs and renders evolving cities from geographic data, synchronizes multiple physics engines, and exposes shared physical and structured urban state to agents. Its Gym-like interface supports user-defined ReAct agents and single- or multi-agent natural-language missions with configurable star or broadcast communication and resource constraints. Each episode becomes an attribution-ready replay that links agent trajectories and communication to relation-graph changes, resource consumption, and engine-based evaluations for systematic diagnosis. We evaluate twelve vision-language models on nine urban tasks under a shared engine-in-the-loop protocol. Controlled studies further examine communication, scalability, robustness, and failure provenance. Results expose persistent bottlenecks in grounding and long-horizon execution. They also show task-dependent coordination trade-offs and diminishing returns from added capacity, while heavier workloads further reduce success. Lingjing provides a unified testbed that enables reproducible end-to-end evaluation and systematic failure diagnosis in urban multi-agent embodied intelligence.
JustLLMGRPO: 胸部 X 線生成のための X 線撮影制御
テキスト条件付き胸部 X 線生成は、指定された所見を忠実に描写するリアルな X 線写真を合成することを目的としています。既存の作業では、主に画像ジェネレーターを更新し、CXR ドメインの適応後にプロンプトを修正されたものとして暗黙的に処理することで品質が向上しました。このジェネレーター中心の考え方では、最適化の重要な側面が十分に検討されていないことがわかります。 CXR に適合した Sana ジェネレーターを凍結すると、未変更の LLM によるワンパス再配合により、RadDINO-FID が 54.225 から 27.572 に減少します。迅速な分析により、LLM は時間的な比較、不確実性、およびその他のレンダリング不可能なレポート コンテンツを抑制し、目に見える X 線撮影所見を強調することが示されています。ただし、制約のない再定式化では、ソース プロンプトとの BioViL-T のアラインメントが 0.695 から 0.609 に減少します。そこで、Sana を凍結したままにして、標準のグループ相対ポリシー最適化 (GRPO) を LLM プロンプト ポリシーにのみ適用する JustLLMGRPO を導入します。グループ相対の放射線学を意識した画像フィードバックにより、ソースとプロンプトの位置合わせを維持しながら、視覚的な焦点が維持されます。 CheXGenBench では、JustLLMGRPO は、調整 (0.696 対 0.695) を維持しながら、RadDINO-FID を 26.780 に削減します。これは、直接プロンプトに比べて 50.6% の改善です。また、最先端の配信範囲と下流の分類ユーティリティも実現します。これらの結果は、適合したジェネレータに対して放射線写真情報をどのように表現するかにおいて、実質的な性能が潜在的に残る可能性があることを示しています。コードは https://github.com/pxcai/JustLLMGRPO で公開されています。
原文 (English)
JustLLMGRPO: Radiographic Control for Chest X-Ray Generation
Text-conditioned chest X-ray generation aims to synthesize realistic radiographs that faithfully depict specified findings. Existing work has primarily improved quality by updating image generators, implicitly treating prompts as fixed after CXR-domain adaptation. We show that this generator-centric view leaves a substantial optimization dimension underexplored. With a CXR-adapted Sana generator frozen, one-pass reformulation by an unmodified LLM reduces RadDINO-FID from 54.225 to 27.572. Prompt analysis shows that the LLM suppresses temporal comparisons, uncertainty, and other non-renderable report content while emphasizing visible radiographic findings. However, unconstrained reformulation reduces BioViL-T alignment with source prompts from 0.695 to 0.609. We therefore introduce JustLLMGRPO, which applies standard Group Relative Policy Optimization (GRPO) only to the LLM prompt policy while keeping Sana frozen. Group-relative radiology-aware image feedback retains visual focus while preserving source-prompt alignment. On CheXGenBench, JustLLMGRPO reduces RadDINO-FID to 26.780, a 50.6% improvement over direct prompting, while maintaining alignment (0.696 versus 0.695). It also achieves state-of-the-art distribution coverage and downstream classification utility. These results show that substantial performance can remain latent in how radiographic information is expressed to an adapted generator. Code is publicly available at https://github.com/pxcai/JustLLMGRPO.
SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents
Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was on…
H2: ヒューマンインザループ検証済みの LLM 駆動メタデータ アノテーション システムによる医療データの調和のためのデュアル ハイブリッド セマンティック データ レイク アーキテクチャ
医療データはその性質上、(a) 画像、テキスト、時系列などのさまざまなモダリティ、(b) 医療機関によって導入された多様な表形式スキーマ、(c) 医療専門家によって提供される完全に非構造化のテキスト情報データに至るまで、複数のレベルで高度な異質性を示します。データ レイクは、医療データ ストレージで多くの場合、異種の多様なデータをすべて 1 つの中央の場所に統合するために使用されます。データ ウェアハウスのようにスキーマを適用する必要がなく、データを「そのまま」保存できます。ただし、その柔軟性にもかかわらず、データ レイクは「データ スワンプ」障害で有名です。したがって、完全性や柔軟性を損なうことなく、メタデータを通じて信頼性の高いデータ調和メカニズムを提供することは、大きな課題です。この目的を達成するために、ナレッジ グラフは、厳格なスキーマオンライト アプローチを使用せずに関係を表現する動的な方法を提供するため、注目を集めています。さらに、別の厳密なタスクはデータの相互運用性に依存しています。ドメインの専門家が特定のデータ型またはデータセットに対するメソッドの有効性を判断する必要があるため、このような多様な性質のデータに適切な ML テクニックを適用することは簡単なタスクではありません。メタデータの注釈は、該当する操作にタグを付けることで役立ちますが、そのような情報が不足している既存のデータセットが多数あることは言うまでもなく、手動による介入が必要です。両方の課題に取り組むために、この論文では、データの調和を促進し、意味のある ML 技術の適用をサポートするために、ラベルのないメタデータ コレクションの生成アノテーション プロセス (つまり、LLM) を組み込むセマンティック データ レイク アーキテクチャを提案します。このアプローチに基づいて、私たちはより高いレベルの知識を作成し、データの性質に基づいて適用可能な ML 操作に関するデータの適合性を特定します。
原文 (English)
H2: A Dual Hybrid Semantic Data Lake Architecture for Medical Data Harmonization with Human-In-the-Loop verified, LLM Driven Metadata Annotation System
Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured textual information data provided by healthcare professionals. Data lakes are often used in medical data storage to consolidate all heterogeneous diverse data in a single, central location, where it can be saved "as is", without the need to impose a schema like a data warehouse does. Despite their flexibility, though, data lakes are notorious for the "data swamp" failure. Thus, providing a reliable data harmonization mechanism through metadata, without compromising integrity or flexibility, is a real challenge. To this end, knowledge graphs have attracted attention since they provide a dynamic way to depict relationships without a rigid schema-on-write approach. Additionally, another rigorous task relies on the interoperability of data: application of appropriate ML techniques on such a diverse nature of data is not an easy task, as a domain expert must decide the efficacy of a method to a specific data type or dataset. Metadata annotation can aid by tagging applicable operations, however this requires manual intervention, not to mention the plethora of existing datasets which lack such information. To tackle both challenges, in this paper, we propose a semantic data lake architecture that promotes data harmonization and incorporates a generative annotation process (i.e. LLMs) of non-labeled metadata collections to support the application of meaningful ML techniques. Building on top of this approach, we create a higher level of knowledge, identifying suitability of data with respect to applicable ML operations based on their data nature...
CORDA: 大規模言語モデルにおける階層的危害中心道徳推論のベンチマーク
道徳的判断における重要な問題は、誰かが単に「正しい」答えを選択するかどうかではなく、道徳原則が矛盾する場合に最も重要なことをどのように決定するかということです。大規模言語モデル (LLM) の現在の評価は依然として限定的です。ほとんどのモデルは、道徳的にコストのかからない選択肢がない場合に、競合する原則の間で優先順位を付けることができるかどうかではなく、モデルが道徳的に受け入れられる答えを与えるか、人間の好みに一致するか、明らかな違反を回避できるかどうかをテストします。 LLM における階層的で危害中心の道徳的推論を評価するためのベンチマークである CORDA (Conditioned Ordering and Ranked Directive Adherence) を紹介します。 CORDA は、道徳連鎖の形式主義に基づいて、トロリー型のケース、医療のトレードオフ、資源の割り当て、人間と動物とロボットの対立を含む 90 の道徳的ジレンマを、4 つの秩序ある倫理的枠組み (ユーティリティ、ユーティリティ + エージェントの危害、デュアルプロセス、およびデュアルプロセス + エージェントの危害) にわたってテストします。これらのフレームワークを組み合わせて、道徳的な優先順位が変化したときにモデルが決定を適応できるかどうかをテストします。 7 つのプロバイダーによる 10 個の命令調整モデル全体で、強力な義務論的デフォルトが見つかり、10 個中 9 個では全体的な危害を軽減することよりも直接的な個人的危害の回避を優先しています。また、モデルは、全体的な危害を最小限に抑えるなどの結果ベースの比較よりも、殺人を回避するなどのカテゴリー的な危害回避ルールの方がより信頼性高く実行されます。これは、モデルが、競合する危害を通じて推論するよりも簡単に道徳的な危険な境界線を認識することを示唆しています。すべてのモデルは明示的な連鎖条件付けに応答しますが、動物より人間、ロボットより動物など、指定された優先順位に一貫して従わないモデルもあります。 CORDA は、モデルがデフォルトの危害回避反応を超えて、コンテキスト指定の道徳的優先順位を適用できるかどうかをテストすることで、LLM の道徳的評価における中心的なギャップに対処します。道徳的信頼性を得るには、デフォルトの自制以上のものが必要です。紛争下での制御性が必要です。
原文 (English)
CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models
The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict. Current evaluations of large language models (LLMs) remain limited: most test whether models give morally acceptable answers, match human preferences, or avoid obvious violations, rather than whether they can prioritise between competing principles when no option is morally cost-free. We introduce CORDA (Conditioned Ordering and Ranked Directive Adherence), a benchmark for evaluating hierarchical, harm-centred moral reasoning in LLMs. Building on the morality chains formalism, CORDA tests 90 moral dilemmas involving trolley-style cases, medical trade-offs, resource allocation, and human-animal-robot conflicts across four ordered ethical frameworks: Utility, Utility + Agent Harm, Dual-Process, and Dual-Process + Agent Harm. Together, these frameworks test whether models can adapt their decisions when moral priorities change. Across ten instruction-tuned models from seven providers, we find a strong deontological default, with 9 of 10 prioritising avoidance of direct personal harm over reducing overall harm. Models also perform more reliably on categorical harm-avoidance rules, such as avoiding killing, than on outcome-based comparisons, such as minimising total harm, suggesting that they recognise moral red lines more easily than they reason through competing harms. Although all models respond to explicit chain conditioning, several fail to consistently follow specified priority orderings, such as humans over animals and animals over robots. CORDA addresses a central gap in LLM moral evaluation by testing whether models can move beyond default harm-avoidant responses and apply context-specified moral priorities. Moral reliability requires more than default restraint; it requires controllability under conflict.
探索、地図作成、記憶、決定: 組み込み型 VLM は安全性が重要なシナリオに対応できるか?
Theory of Space フレームワーク (ToS) は、部分的な観測可能性の下で、好奇心主導の視覚言語モデル (VLM) の空間理解を評価します。 AI 技術が安全性が重要なシナリオに適用されることが増えているため、VLM が堅牢な空間メモリを備え、信頼性の高い決定を下せるかどうかを理解することが重要です。この論文では、VLMの意思決定が物理的証拠に基づいているのか、それとも視覚言語のバイアスによって損なわれているのか、VLMの記憶プロセスが人間の認知パターンと一致しているか、そしてVLMが環境危険にどのように反応するかを評価します。 ToS フレームワークを、Explore、Map、Remember、Decide (EMRD) という名前の、セーフティ クリティカルで目標主導型のパイプラインに拡張します。次に、環境範囲と時間効率の指標を通じて探索能力 (探索) を定量化し、空間忠実度 (マップ) を評価し、一連の心理指標を使用して記憶の持続性 (記憶) を評価し、焦点指標を使用して認知的意思決定 (決定) を測定します。私たちの結果は、意思決定能力の点で、VLM は、その選択を正当化するための空間的根拠を欠いている一方で、事前に訓練されたテキストの事前情報に基づいて避難ポイントを選択することが多いことを示しています。また、空間推論は暗い環境では低下しますが、テクスチャや色の改ざんには影響を受けないことも示します。私たちの調査結果は、VLM の記憶が人間の認知から根本的に逸脱しており、不整合の予測不可能なリスクを生み出していることを示唆しています。
原文 (English)
Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?
Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI techniques are increasingly applied to safety-critical scenarios, it is crucial to understand whether VLMs possess robust spatial memory and make reliable decisions. In this paper, we assess whether VLMs' decisions are based on physical evidence or are corrupted by visual-language biases, if their memory processes align with human cognitive patterns, and how they respond to environmental hazards. We extend the ToS framework into a safety-critical, goal-driven pipeline, named Explore, Map, Remember, and Decide (EMRD). We then quantify Exploration Competence (Explore) through metrics of environmental coverage and temporal efficiency, assess Spatial Fidelity (Map), evaluate, with a suite of psychological metrics, Memory Persistence (Remember), and measure, using focal-point metrics, Cognitive Decision-Making (Decide). Our results show that in terms of decision-making capabilities, VLMs frequently select evacuation points based on pre-trained textual priors while lacking the spatial grounding to justify their choices. We also show that spatial reasoning degrades in low-light conditions, but it is not affected by texture and colour tampering. Our findings suggest that VLM memory fundamentally diverges from human cognition, creating unpredictable risks of misalignment.
PATH: 表形式データの自己回帰ツリー階層による次の間隔の予測
間隔予測は、可能な限り短い間隔を生成しながら、目標カバレッジ レベルを達成することを目的としています。多くの等角回帰パイプラインは、最初に不確実性サロゲートを予測し、次にキャリブレーションまたは選択を通じてそれを区間に変換します。この分離はカバレッジ キャリブレーションをサポートしますが、事後ルールは主に最終間隔を決定し、学習された出力分布を完全には使用しません。結果として得られる区間には本質的に階層的な幾何学構造があることが観察されます。区間は再帰的に細分化されて入れ子になった部分区間になり、バイナリ ツリーはこの構造を自然に表現します。この階層を次の区間の予測として定式化し、確率質量が各区間から次の入れ子になった部分区間にどのように流れるかを学習する PATH を提案します。 PATH はベース リーフの分布を予測し、自己回帰デコーダを使用して分岐確率を調整します。分布を区間階層に一致させることで、学習と抽出が調整されます。PATH は、隣接する出力区間にわたる確率を蓄積し、選択された質量に到達する最短の連続範囲を返します。 PATH を、56 の OpenML 回帰データセットで構成される PATHBench の間隔予測の 24 のベースラインと比較します。 PATH は結果の間隔を大幅に短縮し、平均カバレッジ 0.9144 を維持しながら、最低の平均正規化長さ 0.1473 を達成します。これらの結果は、表形式データに対するコンパクトな間隔予測のための効果的なアプローチとして階層出力モデリングを確立します。コードは https://github.com/pxcai/PATH で公開されています。
原文 (English)
PATH: Next-Interval Prediction via Autoregressive Tree Hierarchy on Tabular Data
Interval prediction aims to achieve a target coverage level while producing intervals that are as short as possible. Many conformal regression pipelines first predict an uncertainty surrogate and then convert it into an interval through calibration or selection. This separation supports coverage calibration, but post hoc rules largely determine the final interval and do not fully use the learned output distribution. We observe that the resulting intervals have inherently hierarchical geometry: an interval can be recursively refined into nested subintervals, and binary trees naturally represent this structure. We formulate this hierarchy as next-interval prediction and propose PATH, which learns how probability mass flows from each interval to its next nested subintervals. PATH predicts a base leaf distribution and uses an autoregressive decoder to refine branch probabilities. Matching the distribution to the interval hierarchy aligns learning with extraction: PATH accumulates probability over adjacent output intervals and returns the shortest contiguous range reaching a selected mass. We compare PATH with 24 baselines for interval prediction on PATHBench, comprising 56 OpenML regression datasets. PATH substantially shortens the resulting intervals, achieving the lowest mean normalized length, 0.1473, while maintaining mean coverage of 0.9144. These results establish hierarchical output modeling as an effective approach for compact interval prediction on tabular data. Code is publicly available at https://github.com/pxcai/PATH.
生成モデル: 原理、アーキテクチャ、およびアプリケーション
生成 AI は、現代の人工知能において最も変革をもたらす力の 1 つとして台頭しており、私たちがデジタル コンテンツを作成、想像し、操作する方法を再構築しています。フォトリアリスティックな画像から一貫したテキストまで、没入型ビデオから斬新な分子構造まで、生成モデルは、かつては SF に限定されていたアプリケーションを強化するようになりました。この本は、この革命を支える基本原理、数学的基礎、実践的なアーキテクチャを読者にガイドすることを目的としています。
原文 (English)
Generative Models: Principles, Architectures, and Applications
Generative AI has emerged as one of the most transformative forces in modern artificial intelligence, reshaping how we create, imagine, and interact with digital content. From photorealistic images to coherent text, from immersive videos to novel molecular structures, generative models now power applications that were once confined to science fiction. This book is designed to guide readers through the foundational principles, mathematical underpinnings, and practical architectures that underpin this revolution.
深く考え、一度話してください: Relit、再帰的潜在暗黙的変換フレームワーク
思考連鎖 (CoT) プロンプトは、大規模言語モデル (LLM) で推論を引き出すための主要なパラダイムとなっていますが、モデルに中間推論ステップを離散トークンとして強制的に外部化させるため、かなりの計算オーバーヘッドが発生します。最近の潜在的推論アプローチでは、このプロセスを継続的な隠れた状態の中に内在化しようとしています。潜在推論の分野における最新の進歩の 1 つである Tiny Recursive Models (TRM) は、記号推論には優れていますが、自然言語設定で意味の一貫性を維持するのに苦労しています。このギャップを埋めるために、基礎モデルの豊富なセマンティック表現内で深い再帰的推論を基盤とするハイブリッド フレームワークである ReLIT (Recursive Latent Implicit Transformer) を導入します。 ReLIT は、フリーズした LLM バックボーン (TinyLlama-1.1B) を軽量でトレーニング可能な再帰ブロックで強化します。このブロックは、最終出力にコミットする前に潜在的思考 (z) を反復的に洗練させ、アルゴリズム処理からの言語的直観を構造的に解決し、明示的なトークン生成の待ち時間なしで勾配分離された再帰ループを介して「深い思考」を可能にします。経験的に、ReLIT は GLoRE 論理推論ベンチマークで高いパラメーター効率を達成し、最小限の監視にもかかわらず、ProofWriter や RuleTaker などの困難なタスクで非常に大規模なモデルと同等またはそれを上回るパフォーマンスを示します。これらの結果は、パラメータ幅ではなく反復深さによって推論機能を効率的に拡張でき、意味論的に根拠のある暗黙的推論のための原則に基づいたフレームワークを提供できることを示しています。
原文 (English)
Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework
Chain-of-Thought (CoT) prompting has become the dominant paradigm for eliciting reasoning in Large Language Models (LLMs), yet it creates substantial computational overhead by forcing models to externalize intermediate reasoning steps as discrete tokens. Recent latent reasoning approaches attempt to internalize this process within continuous hidden states. One of the latest advancements in the field of latent reasoning, Tiny Recursive Models (TRMs) excel at symbolic reasoning but struggle to preserve semantic coherence in natural language settings. To bridge this gap, we introduce ReLIT (Recursive Latent Implicit Transformer), a hybrid framework that grounds deep recursive reasoning within the rich semantic representations of a foundational model. ReLIT augments a frozen LLM backbone (TinyLlama-1.1B) with a lightweight, trainable recursive block that iteratively refines its latent thinking (z) before committing to a final output, structurally solving linguistic intuition from algorithmic processing and enabling "deep thinking" via gradient-isolated recurrent loops without the latency of explicit token generation. Empirically, ReLIT achieves high parameter efficiency on the GLoRE logical reasoning benchmark, matching or outperforming significantly larger models on challenging tasks such as ProofWriter and RuleTaker despite minimal supervision. These results demonstrate that reasoning capability can be scaled efficiently through recurrent depth rather than parameter width, offering a principled framework for semantically grounded implicit reasoning.
代数グラフ構造の神経象徴的な発見
所定のプロパティを持つグラフを検索するには、SAT ソルバーや特殊なジェネレーターなど、いくつかの方法があります。これらのメソッドは、結果を生データ (隣接行列または文字列エンコーディング) として返します。生データはグラフが存在することを証明しますが、グラフの構造特性は明らかにしません。この生データのみが提供された場合に、短い代数記述を自動的に発見できるかどうかを尋ねます。ケイリーグラフ $\mathrm{Cay}(\Gamma, S)$ や辞書編集積 $C_5[K_3]$ などの記述を探します。私たちは神経象徴的なアプローチでこの疑問に取り組みます。私たちは、微調整やターゲットごとのトレーニングを行わずに、汎用の大規模言語モデル上で実行されるエージェントを提案します。このモデルは、コンピューター代数システム SageMath への呼び出しで推論をインターリーブします。ターゲットのグラフを分析し、候補の構成を提案およびテストし、出力がターゲットと一致するまでそれらを修正します。エージェントは、汎用ブリッジとしてリリースされた Model Context Protocol (MCP) サーバーを介して SageMath と通信します。構造がターゲットと一致するかどうかは、単一の正確な同型性テストによってチェックされるため、モデルではなく記号側に依存します。 100 個の対称性の高いグラフ、つまり最大 25 頂点の 2 軌道グラフのベンチマークでアプローチをテストします。ベンチマークは事前に固定されていました。私たちのエージェントは、生のエンコーディングに戻ることなく、それらすべてについて検証済みの代数構造を見つけることができました。強力なテンプレート列挙ベースラインは約 $20\%$ にのみ達し、カタログ検索ではこれらのグラフを識別できませんでした。ただし、対称性がなくなると施工品質が低下します。具体的な応用例として、Bernhart-Kainen の分散性予想に対する既知の最小の反例、つまり列挙によって生データとして見つかった $16$ 頂点グラフを特定します。このグラフに関して、私たちのエージェントは明示的な代数構造を発見しました。
原文 (English)
Neurosymbolic Discovery of Algebraic Graph Constructions
There are several methods for searching for graphs with prescribed properties, such as SAT solvers and specialized generators. These methods return the result as raw data: an adjacency matrix or a string encoding. The raw data certifies that the graph exists, but it does not reveal any structural properties of the graph. We ask whether one can automatically discover a short algebraic description if only this raw data is provided. We look for a description such as a Cayley graph $\mathrm{Cay}(\Gamma, S)$ or a lexicographic product $C_5[K_3]$. We address this question with a neurosymbolic approach. We propose an agent that runs on a general-purpose large language model with no fine-tuning or per-target training. The model interleaves reasoning with calls to the computer algebra system SageMath: it analyzes the target graph, proposes and tests candidate constructions, and revises them until the output matches the target. The agent communicates with SageMath through a Model Context Protocol (MCP) server, which we release as a general-purpose bridge. Whether a construction matches the target is checked by a single exact isomorphism test, and therefore rests on the symbolic side and not on the model. We test the approach on a benchmark of 100 highly symmetric graphs, namely two-orbit graphs on up to 25 vertices; the benchmark was fixed in advance. Our agent could find verified algebraic constructions for all of them, without falling back to raw encodings. A strong template-enumeration baseline reaches only about $20\%$, and a catalog lookup could not identify any of these graphs. However, construction quality declines when symmetry is removed. As a concrete application, we identify the smallest known counterexample to the Bernhart-Kainen dispersability conjecture, a $16$-vertex graph that enumeration found as raw data. For this graph, our agent found an explicit algebraic construction.
Constraining ontology mappings using metaphysical choices
In this paper we discuss the foundations behind a novel methodology for the validation of semantic mappings between different data sources…
LLM エージェントによる制約モデルの改善
制約プログラミング (CP) ソルバーのランタイムは、対称性の破れ、暗黙の制約、グローバル制約、制約の再定式化、変数表現などのモデリングの選択に非常に敏感です。これらの制約モデルの改善には従来、人間の専門知識が必要であり、既存の自動再定式化システムは、手作りの変換ルールの事前定義されたライブラリに制限されていました。代わりに、オープンエンドの空間から制約モデルを再定式化して、構築ではなく経験的に正しさを確立するエージェント フレームワークを導入します。モデルと 3 つのトレーニング インスタンスが与えられた大規模言語モデル (LLM) エージェントは、代替の定式化を提案し、その解決策を元のモデルに注入してそれぞれを検証し、障害を診断して修復し、中央値約 15 分で見つけた最適なバリアントを返します。モデルは CPMpy モデリング ライブラリで表現され、提案された各モデルは 3 つのより大きなテスト インスタンスで評価されます。 9 つの組み合わせ最適化問題全体で、生成されたモデルは 27 のテスト インスタンス中 21 で元のモデルよりも優れたパフォーマンスを示し、一部の問題では 2 桁以上速く解決されました。同じ検証および選択ツールを再利用する非エージェントベースラインとの比較は、単に複数の候補をサンプリングすることからではなく、エージェントの反復的な診断と修復から利益が得られることを示しています。これらの結果は、自律エージェント手法が制約モデルの改善をサポートできることを示しています。
原文 (English)
Improving Constraint Models with LLM Agents
The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints, global constraints, constraint reformulation, and variable representation. Improving these constraint models has traditionally required human expertise, and existing automated reformulation systems are restricted to a predefined library of hand-crafted transformation rules. We introduce an agentic framework that instead reformulates a constraint model from an open-ended space and establishes correctness empirically rather than by construction: a Large Language Model (LLM) agent, given a model and three training instances, proposes alternative formulations, validates each by injecting its solution back into the original model, and diagnoses and repairs failures, returning the best variant it finds in a median of about fifteen minutes. The models are expressed in the CPMpy modeling library, and each proposed model is evaluated on three larger test instances. Across nine combinatorial optimization problems, the generated models outperform the originals on 21 of 27 test instances, and on some problems solve more than two orders of magnitude faster. A comparison against non-agentic baselines that reuse the same validation and selection tools indicates that the gains stem from the agent's iterative diagnosis and repair, not merely from sampling several candidates. These results demonstrate that autonomous agentic methods can support the improvement of constraint models.
TokenPrint: 言語モデルの出所を確認するための調整されたトークン空間フィンガープリント
言語モデルの出自を確立すること (その基本チェックポイントやトレーニング配布での重複の可能性を含む) は、メタデータだけでは解決できないガバナンスの課題です。我々は、デコードされたトークン文字列に対する Jaccard のオーバーラップを使用して比較した、250 の固定知識プローブによって引き出された後期隠れ状態の上位千ドルの語彙予測に基づくトレーニング不要のフィンガープリントを導入します。文書化された関係を使用して、9 つのファミリー (0.6B ~ 32B) からの 32 のオープンウェイト モデルでこの方法を評価します。 (1)~\emph{類似性ラダー} は、モデルの関連性を大まかに追跡します。同一のデータで独立してトレーニングされたモデルは、未加工のスコア 0.48 (語彙補正済み 0.35)、続いて共有ベースの微調整 (0.39/0.33)、同じ開発者の親戚 (0.38/0.28)、および文書化された関係のないモデル (0.22/0.17) です。この同一データ シグナルは、3 つの組織、2 つのトークナイザー ファミリ、および 2 つのアーキテクチャ クラスにわたって持続し、測定可能なタスク コンピテンス前のトレーニングの最初の 1\% 以内に現れ、機能の収束を超えた共有トレーニング データの寄与を示唆しています。 (2)~最近傍 \emph{lineage-retrieval} 法として、フィンガープリントは、大まかなメタデータからは識別できない数学に特化した塩基を含め、5 つの R1 蒸留すべてについて上位 2 つの候補の間で正確に文書化された塩基をランク付けします (平均ランク 1.8、MRR 0.60)。 (3)~A \emph{深さアブレーション} は、系統グループの識別が出力分布に向かって強化されることを示しており、AUC は 4 分の 1 の深さでの 0.72 から出力での 0.90 まで増加しています。上位 5 つの出力トークンのみを使用すると、AUC 0.87 が維持されます。 (4) ~ キャリブレーション プールの最大クロスモデル類似度 0.81 と比較して、フィンガープリントは量子化下でも安定しており、int8 では 0.92、int4 では 0.82 ~ 0.85 の Jaccard 類似度を示します。プローブ、コード、フィンガープリントを公開します。
原文 (English)
TokenPrint: A Calibrated Token-Space Fingerprint for Language-Model Provenance
Establishing the provenance of a language model---including its base checkpoint and possible overlap in training distributions---is a governance challenge that metadata alone cannot resolve. We introduce a training-free fingerprint based on the top-$k$ vocabulary projections of late hidden states elicited by 250 fixed knowledge probes, compared using Jaccard overlap over decoded token strings. We evaluate the method on 32 open-weight models from nine families (0.6B--32B) with documented relationships. (1)~A \emph{similarity ladder} broadly follows model relatedness: independently trained models on identical data score 0.48 raw (0.35 vocabulary-corrected), followed by shared-base fine-tunes (0.39/0.33), same-developer relatives (0.38/0.28), and models with no documented relationship (0.22/0.17). This identical-data signal persists across three organizations, two tokenizer families, and two architecture classes, and emerges within the first 1\% of training before measurable task competence, suggesting a contribution from shared training data beyond capability convergence. (2)~As a nearest-neighbor \emph{lineage-retrieval} method, the fingerprint ranks the exact documented base among the top two candidates for all five R1 distillations (mean rank 1.8, MRR 0.60), including a math-specialized base not identifiable from coarse metadata. (3)~A \emph{depth ablation} shows that lineage group discrimination strengthens toward the output distribution, with AUC increasing from 0.72 at quarter depth to 0.90 at the output; using only the top 5 output tokens retains AUC 0.87. (4)~The fingerprint remains stable under quantization, with Jaccard similarity of 0.92 under int8 and 0.82--0.85 under int4, compared with a maximum cross-model similarity of 0.81 in the calibration pool. We release the probes, code, and fingerprints.
論理的推論としての Long SKILL コンプライアンス: スケーリングに基づくポリシーに基づく蒸留によるクロージャベースの検出
エンタープライズ ビジネス シナリオの複雑化により、エージェント システムでの長い SKILL ドキュメントの普及が促進され、コンプライアンス検出に新たな課題が生じています。大規模なモデルでは多大な推論コストが発生し、小規模なモデルでは検出精度を維持できない可能性があります。このギャップに対処するために、私たちは、長時間にわたる SKILL コンプライアンス検出のためのグラフベースのフレームワークである SkillCDG を提案します。 SkillCDG は、複雑なビジネス ポリシーを 2 層の制約依存関係グラフとして表します。上位層はシナリオ ルーティングの SKILL 記述にインデックスを付け、下位層は各 SKILL 内のアトミックな制約間の依存関係をキャプチャします。推論中、2 レベルの取得とそれに続く依存関係のクロージャにより、コンプライアンスの判断とソースのトレーサビリティがサポートされます。私たちは、3 つのエンタープライズ データセットと 2 つの管理された公開ベンチマーク バリアントに基づいてフレームワークを包括的に評価します。実験結果は、SkillCDG が検出 F1 スコアでベースライン手法を最大 12.8 パーセント上回り、トークン消費量を最大 64.3\% 削減することを示しています。さらに、ポリシー グラフの複雑さ、モデルの規模、検出パフォーマンスの間の固有の関係をさらに調査します。単一のモデル ファミリの 4 つのチェックポイントで実施された比較実験により、簡潔で効果的なスケーリング傾向が検証されます。エンドツーエンドの検出の正確性は、複雑さによって区別されたスケーリング パターンを示し、制約依存関係グラフから導出された複雑さのメトリクスは、インスタンスの難易度とモデルのパフォーマンス向上の可能性を効果的に定量化できます。この洞察力に富んだスケーリング傾向を活用して、適応的なトレーニング サンプルの選択を実施し、ポリシーに基づく蒸留を採用して、小規模モデルのコンプライアンス検出機能を効率的に強化します。
原文 (English)
Long SKILL Compliance as Logical Reasoning: Closure-Grounded Detection with Scaling-Guided On-Policy Distillation
The increasing complexity of enterprise business scenarios has promoted the widespread adoption of long SKILL documents in agent systems, posing new challenges for compliance detection: large models incur substantial inference costs, while small models may fail to maintain detection accuracy. To address this gap, we propose SkillCDG, a graph-based framework for long SKILL compliance detection. SkillCDG represents complex business policies as a two-layer constraint dependency graph, where the upper layer indexes SKILL descriptions for scenario routing and the lower layer captures dependencies among atomic constraints within each SKILL. During inference, two-level retrieval followed by dependency closure supports compliance judgment and source traceability. We comprehensively evaluate the framework on three enterprise datasets and two controlled public benchmark variants. Experimental results demonstrate that SkillCDG outperforms baseline methods by up to 12.8 percentage points in detection F1 score, while reducing token consumption by a maximum 64.3\%. Moreover, we further investigate the inherent relationships among policy-graph complexity, model scale, and detection performance. Comparative experiments conducted on four checkpoints from a single model family validate a concise and effective scaling trend: end-to-end detection correctness exhibits a complexity-differentiated scaling pattern, and the complexity metric derived from the constraint dependency graph can effectively quantify instance difficulty and the performance improvement potential of models. Leveraging this insightful scaling trend, we conduct adaptive training sample selection and adopt on-policy distillation to efficiently enhance the compliance detection capability of small-scale models.
強化学習における動的報酬形成のための統一フレームワーク
まばらで遅れ、情報量が弱い報酬は、依然として効率的な強化学習の中心的な障害となっています。報酬形成は、古典的な設定では元の目的が評価基準のままである一方で、学習を加速できる補助信号でタスク報酬を補うことによってこれらの制限に対処します。確立された理論は、固定整形信号の安全性を保証します。補助項が時間不変ポテンシャルの割引差である場合、ポテンシャルベースの報酬整形は最適なポリシーを保存します。しかし、現代の強化学習システムでは、学習者と指導に利用できる情報の両方がトレーニング中に進化します。つまり、値の推定値が向上し、新規性が減少し、フィードバックが変化し、予測モデルが洗練されます。適応型報酬メカニズムは、探索、ベイジアン推論、人間参加型学習、自動報酬設計、基盤モデルベースのアプローチにわたって発生します。この研究では、動的報酬形成と隣接する適応報酬メカニズムを比較するための統一分析フレームワークを導入します。提案されたフレームワークは、パラメトリック修正を状態依存変動から区別し、加算的整形を報酬置換および報酬隣接ガイダンスから分離し、時間的、情報的、理論的側面に沿って既存の手法を整理します。このフレームワークを使用して、12 のメソッド ファミリが比較分析されます。このフレームワークはさらに、適応率と学習者の安定性の間の未解決の関係を明らかにしながら、現代の深層強化学習パイプライン、リプレイバッファー、ブートストラップされた批評家、報酬の正規化において最適性の保証が生き残る条件を強調しています。
原文 (English)
A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning
Sparse, delayed, and weakly informative rewards remain central obstacles to efficient reinforcement learning. Reward shaping addresses these limitations by supplementing the task reward with an auxiliary signal that can accelerate learning while, in the classical setting, the original objective remains the evaluation criterion. Established theory guarantees safety for fixed shaping signals: potential-based reward shaping preserves optimal policies when the auxiliary term is the discounted difference of a time-invariant potential. In contemporary reinforcement learning systems, however, both the learner and the information available for guidance evolve during training: value estimates improve, novelty diminishes, feedback shifts, and predictive models are refined. Adaptive reward mechanisms occur across exploration, Bayesian inference, human-in-the-loop learning, automated reward design, and foundation-model-based approaches. This study introduces a unified analytical framework for comparing dynamic reward shaping and neighbouring adaptive reward mechanisms. The proposed framework distinguishes parametric revision from state-dependent variation, separates additive shaping from reward replacement and reward-adjacent guidance, and organises existing methods along temporal, informational, and theoretical dimensions. Using this framework, twelve method families are comparatively analysed. The framework further highlights the conditions under which optimality guarantees survive contemporary deep reinforcement learning pipelines, replay buffers, bootstrapped critics, and reward normalisation, while exposing the unresolved relationship between adaptation rate and learner stability.
操作可能なコンセプト表現はいつ実現しますか? LLM における神経科学の並行性の家族間監査における測定の混乱
大規模言語モデル (LLM) は、概念セル、心の数直線、認知マップなど、人間に似た神経的および認知的特徴を示すことがますます報告されています。これらの主張は、多くの場合、単一モデルに適用される線形プローブとアクティベーション ステアリングに依存しますが、どちらの方法も測定の選択に非常に敏感です。したがって、報告される平行線は、モデル、測定手順、またはその両方を反映している可能性があります。私たちは、0.6 億ドルから 72 億ドルのパラメーターにわたる 5 つのファミリーの 17 のモデルにわたって、神経科学にヒントを得た 4 つの代表的なパラダイムを監査します。私たちの主な実験では、概念の方向性の因果的方向性を調べます。生のアクティベーション ユニットと固定層と係数を使用すると、操縦性はモデルのスケールに応じて増加するように見え、創発的な機能に似ています。ただし、このパターンは、ステアリング関連の文献で確立された主張によってではなく、校正されていないパイプラインによって生成されます。傾向は、生の単位、読み出しメトリック、および操作点に連動して依存します。これらのいずれかを修正すると削除されます。残差ノルムと比較可能な介入とホールドアウトされた操作点の選択により、コンセプトのステアリングはすべてのスケールで依然として重要ですが、Qwen3 シリーズ全体では顕著な傾向は示されていません。ただし、信頼区間は緩やかな正の傾きを排除しません。残りの結果はまちまちです。線形地理世界地図は、最大 $72$B までのテスト済みチェックポイントごとに一貫して解読可能です。数の大きさは強くエンコードされますが、個々のニューロンが釣鐘型に見えるか単調に見えるかは選択基準によって異なります。言語固有の構造はローカライズ可能ですが、異なる帰属方法では言語間の非対称性の方向が逆転します。これらの結果は、AI 神経科学に対する主な制約は現象の欠如ではなく、比較可能な測定と適切な制御の欠如であることを示唆しています。プロトコル、刺激、コードを公開します。
原文 (English)
When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs
Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, mental number lines, and cognitive maps. These claims often rely on linear probing and activation steering applied to a single model, yet both methods are highly sensitive to measurement choices. A reported parallel may therefore reflect the model, the measurement procedure, or both. We audit four representative neuroscience-inspired paradigms across 17 models from five families, spanning $0.6$B to $72$B parameters. Our main experiment examines the causal steerability of concept directions. With raw activation units and a fixed layer and coefficient, steerability appears to increase with model scale, resembling an emergent capability. However, this pattern is produced by an uncalibrated pipeline rather than by a claim established in the steering literature. The trend depends jointly on raw units, the readout metric, and the operating point; correcting any one of these removes it. With residual-norm-comparable interventions and held-out operating-point selection, concept steering remains significant at every scale, but shows no significant trend across the Qwen3 series, although the confidence interval does not rule out a moderate positive slope. The remaining results are mixed. A linear geographic world map is consistently decodable in every tested checkpoint up to $72$B. Number magnitude is strongly encoded, but whether individual neurons appear bell-shaped or monotonic depends on the selection criterion. Language-specific structure is localizable, but the direction of the cross-lingual asymmetry reverses under a different attribution method. These results suggest that the main constraint on AI neuroscience is not a lack of phenomena, but a lack of comparable measurements and adequate controls. We release the protocol, stimuli, and code.
Agentic AI 主導の没入型シミュレーション: 高線量率 (HDR) 密封小線源療法のための知識認識型仮想トレーニング プラットフォーム
メタバースと大規模言語モデル (LLM) ベースの AI エージェントの融合により、医学教育における自律的で没入型の個別化された教育フレームワークへの移行が促進されています。この論文では、がん治療における高線量率 (HDR) 膣シリンダー (VC) 近接照射療法用に特別に設計された、新しいエージェント AI 駆動の没入型シミュレーションを紹介します。このシステムは、仮想現実 (VR) とモバイル コンピューティングを統合することにより、忠実度の高いリスクのない環境を確立し、訓練生が物理的な解剖学的構造や実際の放射線源によってもたらされる設備や安全性の制約を受けることなく、複雑な手順スキルを習得できるようにします。この研究の主な貢献は、信頼できる臨床ガイドラインに基づいてエージェントの対話を確立するために、検索拡張生成 (RAG) を活用した知識認識アシスタントのシームレスな統合です。このアーキテクチャにより、対話型エージェントは、複雑な医療操作中に自然言語インターフェイスとハンズフリーのリアルタイム ガイダンスを提供することもできます。ローカルの GPU 高速化 AI バックエンドにリンクされた Meta Quest 3 インターフェイスで構成されるプロトタイプ展開を通じて、提案されたシステムを検証し、HDR 小線源療法シミュレーションの実現可能なアーキテクチャを実証します。実験結果は、システムが適切なエンドツーエンドの待ち時間と高いコンテキスト精度、回答の完全性、および RAG によって強化された教育的サポートの関連性を維持していることを示しています。
原文 (English)
Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy
The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized pedagogical frameworks in medical education. This paper presents a novel agentic AI-driven immersive simulation specifically designed for High Dose Rate (HDR) vaginal cylinder (VC) brachytherapy in cancer care. By integrating Virtual Reality (VR) and mobile computing, the system establishes a high-fidelity, risk-free environment that allows trainees to master complex procedural skills without the facility or safety constraints posed by physical anatomy or live radioactive sources. A core contribution of this work is the seamless integration of a knowledge-aware assistant leveraging Retrieval-Augmented Generation (RAG) to ground agent interactions in authoritative clinical guidelines. This architecture also enables an interactive agent to provide natural language interfaces and hands-free, real-time guidance during intricate medical maneuvers. We validate the proposed system through a prototype deployment comprising a Meta Quest 3 interface linked to a local GPU-accelerated AI backend, demonstrating a feasible architecture for HDR brachytherapy simulation. Experimental results indicate that the system maintains suitable end-to-end latency and high context precision, answer completeness, and relevance in the RAG-enhanced pedagogical support.
生徒の学習能力に合わせた監督: ポリシーに基づいて自己蒸留するための統一フレームワーク
オンポリシー自己蒸留 (OPSD) は、自己蒸留を通じて特権コンテキストをモデル パラメーターに内部化することで、LLM の推論能力を向上させます。最近の 2 つの研究ラインは、それぞれどのトークンから学習するかを選択することと、教師が受け取る特権情報の量を制御することによって、バニラ OPSD を促進します。ただし、各行が 1 つの変数を最適化し、もう 1 つの変数を固定したままにし、次善の解決策につながることを示します。私たちは、この 2 つの変数は生徒の学習能力を通じて結びついていると主張します。特権情報は教師が規定するトークンごとの発散を設定し、トークンの重み付けは生徒がどちらを吸収しなければならないかを選択します。私たちは、この 2 つの作業を統一された最適化フレームワークに形式化します。これにより、生徒が吸収できる学習難易度の合計に関する予算を条件として、教師と生徒の乖離の合計が最大化されます。このモデリングの下で、ラグランジアンを解くための軽量オンライン アルゴリズムである統合オンポリシー自己蒸留 (USD) を提案します。 USD は、単一の二重変数が両方の決定を制御することを明らかにしました。つまり、学習の難易度を 1 つの代償として、トークン選択のしきい値と特権情報の調整の方向を同時に設定し、生徒の進化する能力に合わせた監督を維持します。広範な実験を通じて、USD は、さまざまな推論ベンチマークにおけるさまざまなモデル スケールにわたって、OPSD およびトークン側および PI 側のベースラインを上回る優れたパフォーマンスを一貫して実証しています。コードは https://github.com/lauvlalala/USD で入手できます。
原文 (English)
Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation
On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation. Two recent research lines promote vanilla OPSD by choosing which tokens to learn from and by controlling how much privileged information the teacher receives, respectively. However, we show that each line optimizes one variable while holding the other fixed, which leads to a suboptimal solution. We argue that the two variables are coupled through the student's learning capacity: the privileged information sets the per-token divergence the teacher prescribes, while token weighting selects which of these the student must absorb. We formalize the two lines of work into a unified optimization framework, which maximizes the aggregate teacher--student divergence, subject to a budget on the aggregate learning difficulty the student can absorb. Under this modelling, we propose Unified On-Policy Self-Distillation (USD), a lightweight online algorithm to solve the Lagrangian. USD reveals that a single dual variable governs both decisions: at one price for learning difficulty, it simultaneously sets the token-selection threshold and the direction of privileged-information adjustment, keeping supervision matched to the student's evolving capacity. Through extensive experiments, USD consistently demonstrates superior performance over OPSD and token- and PI-side baselines across various model scales on various reasoning benchmarks. Code is available at https://github.com/lauvlalala/USD.
高度道路交通システム向けの大規模複合エージェント: アーキテクチャ、証拠、導入の課題
大規模マルチモーダル エージェント (LMA) が高度道路交通システム (ITS) 向けに提案されることが増えていますが、既存の研究では、マルチモーダル性、エージェンシー、経験的パフォーマンス、展開の準備性が混同されていることがよくあります。このレビューは、91 のマッピングされた情報源のコーパス内で、2023 年 1 月から 2026 年 8 月 3 日までに発表された 42 の主要な研究ファミリーの監査可能な証拠マップを提供します。これは、モデル レベル、システム レベル、ハイブリッド マルチモダリティを区別し、システム アーキテクチャとアクション権限によって各ファミリーを分類します。証拠は、機能能力 (C0 ~ C3)、検証設定 (E0 ~ E4)、3 つの証拠提案 (P1 ~ P3)、および 8 つの方法論的関心領域 (Q1 ~ Q8) を通じて独立して評価されます。輸送セマンティクス (P1) は 23 のファミリーで直接評価され、多次元統合 (P3) は 24 のファミリーで評価されます。 19 家族が両方を直接評価しています。証拠の調整(P2)は、出所、異議申し立て、対応、比較、結果の連鎖を完全に証明している家族がいないため、未解決のままです。 14 ファミリーが C3 に到達しますが、13 ファミリーは E2 に残ります。 E3 に到達するのは 1 つだけで、E4 に到達するものはありません。 LMA は、ITS ドメイン全体で、意味解釈、意図の翻訳、証拠の整理、シナリオの作成、説明、および専門ツールの調整に最適にサポートされています。数値予測、最適化、シミュレーションの忠実度、厳しい制約、低レベルの制御、安全フォールバック、および最終的な権限は、独立して検証可能な専門家システムまたは責任ある人間が担う必要があります。したがって、このレビューでは、置き換えではなく限定されたオーケストレーションをサポートし、責任ある展開のための適合した比較評価プロトコルと段階的なロードマップを提供します。生きた証拠リポジトリは、https://github.com/pangjunbiao/ITS-LMA-Review で入手できます。
原文 (English)
Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges
Large multimodal agents (LMAs) are increasingly proposed for intelligent transportation systems (ITS), but existing studies often conflate multimodality, agency, empirical performance, and deployment readiness. This review provides an auditable evidence map of 42 primary study families released between January 2023 and 3 August 2026 within a corpus of 91 mapped sources. It distinguishes model-level, system-level, and hybrid multimodality and classifies each family by system architecture and action authority. Evidence is assessed independently through functional capability (C0-C3), validation setting (E0-E4), three evidence propositions (P1-P3), and eight methodological-concern domains (Q1-Q8). Transportation semantics (P1) are directly evaluated in 23 families and multidimensional integration (P3) in 24; 19 families directly evaluate both. Evidence reconciliation (P2) remains unresolved because no family demonstrates the complete provenance-challenge-handling-comparison-outcome chain. Fourteen families reach C3, but 13 remain at E2; only one reaches E3 and none reaches E4. Across ITS domains, LMAs are best supported for semantic interpretation, intent translation, evidence organisation, scenario authoring, explanation, and specialist-tool coordination. Numerical forecasting, optimisation, simulation fidelity, hard constraints, low-level control, safety fallback, and final authority should remain with independently verifiable specialist systems or accountable humans. The review therefore supports bounded orchestration rather than replacement and provides a matched comparative evaluation protocol and staged roadmap for accountable deployment. The living evidence repository is available at https://github.com/pangjunbiao/ITS-LMA-Review.
大規模言語モデルにおける量子化劣化: 信号とノイズの観点から
トレーニング後の量子化により、大規模な言語モデルの導入コストが削減されますが、量子化モデルがどの程度劣化するかはビット幅だけでは決まりません。私たちは、ビット幅、量子化方法、モデルスケール、および複数のモデルファミリーに対するダウンストリームタスクにわたる重みのみのトレーニング後の量子化を体系的に研究します。このような低下は、次の要因によって大幅に異なることが観察されています。4 ビットの量子化では通常パフォーマンスが維持されますが、2 ビットでは広範な低下が発生することが多く、3 ビットでは低下が明らかになりますが、タスクの種類、量子化方法、モデルのスケールによって著しく異なります。この変動性を説明するために、信号対雑音比 (SNR) を使用して、量子化が完全精度表現をどの程度強く混乱させるかを測定します。劣化をたどると、量子化エラーが個々のモジュール内でどのように発生するか、および量子化エラーがレイヤー間でどのように蓄積されるかという 2 つのリンクされたプロセスにまで遡ります。まず、ソース SNR 分解により、新たに導入された誤差が 3 つの要因に依存することがわかります。それは、重み誤差の大きさ、タスク固有の信号の強度、および量子化誤差がタスク固有のアクティベーションとどれだけ強く一致するかです。さまざまな要因がこれらのコンポーネントにさまざまな方法で影響を与えます。第 2 に、層間伝播解析により、これらの誤差は層を越えて通過する際に減衰、保存、または増幅される可能性があり、より大きなモデルほど弱い誤差増幅の恩恵を受けることが示されています。これらの結果を総合すると、量子化の劣化は、エラーがソースでどのように導入され、ネットワーク全体でどのように蓄積されるかによって左右されることが証明されます。
原文 (English)
Quantization Degradation in Large Language Models: A Signal-Noise Perspective
Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation, and at 3-bit, degradation becomes apparent but varies markedly with task type, quantization method and model scale. To explain this variability, we use the signal-to-noise ratio (SNR) to measure how strongly quantization perturbs full-precision representations. We trace degradation back to two linked processes: how quantization errors arise within individual modules, and how they accumulate across layers. First, a source SNR decomposition shows that newly introduced errors depend on three factors: the magnitude of the weight error, the strength of the task-specific signal, and how strongly the quantization error aligns with task-specific activations. Different factors affect these components in distinct ways. Second, a cross-layer propagation analysis shows that these errors can be attenuated, preserved, or amplified as they pass across layers, and that larger models benefit from weaker error amplification. Together, these results establish that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.
Janus: 高額な評価予算の下での LLM 主導の発見のためのアルゴリズムと評価者の共進化フレームワーク
LLM 主導のプログラム検出は評価者の迅速なフィードバックに依存していますが、多くの科学および工学タスクでは高忠実度のシミュレーション、ハードウェア実行、または物理実験が必要であり、各評価にコストがかかります。安価なサロゲート評価器を使用するとこのコストを削減できますが、固定サロゲートは検索による分布シフトに対して脆弱であり、まばらで検索に偏ったラベルから確実に適合することが困難です。 Janus は、LLM を使用してターゲット プログラムと実行可能なプロキシ エバリュエーターを共進化させるフレームワークです。ラベル不足に対処するために、Janus は LLM にエンコードされたドメイン知識を活用してタスク固有の評価プログラムを生成し、実際の結果を使用してそれらを調整します。分布の変化を緩和するために、Janus はターゲット プログラムに沿って評価者を進化させ、プロモーションに合わせた目標を使用して評価者を選択し、オンライン信用更新により地域に応じたポートフォリオを維持します。代理予測は依然として誤りやすいため、Janus は候補者に優先順位を付けるためにのみ代理予測を使用し、候補者がターゲット プログラムの母集団に参加したり、現職者を更新したりする前に実際の検証を必要とします。 Janus は、5 つの科学および工学設計タスクにわたって、実際の評価予算を超えるこれまでの最良の改善曲線の下のより大きな領域と、ターゲット プログラムのみを進化させる一致したベースラインよりも高い最終パフォーマンスを達成しました。平均して、Janus はベースラインの最終改善の 99/% に達しますが、実際の評価は 59.1/% 減少します。進化したプロキシ評価者も、シード バージョンよりも正確に有望な候補をランク付けします。これらの結果を総合すると、評価者主導の LLM 発見が、安価でスケーラブルなフィードバックを伴うタスクから、信頼できる評価が希少で高価な科学分野まで拡張されます。
原文 (English)
Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets
LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks require high-fidelity simulations, hardware execution, or physical experiments, making each evaluation expensive. Cheap surrogate evaluators can reduce this cost, yet fixed surrogates are vulnerable to search-induced distribution shift and are difficult to fit reliably from sparse, search-biased labels. We introduce Janus, a framework that uses LLMs to co-evolve target programs and executable proxy evaluators. To address label scarcity, Janus leverages domain knowledge encoded in LLMs to generate task-specific evaluator programs and calibrates them using real outcomes. To mitigate distribution shift, Janus evolves evaluators alongside target programs, selects them using a promotion-aligned objective, and maintains region-conditioned portfolios with online credit updates. Because proxy predictions remain fallible, Janus uses them only to prioritize candidates and requires real validation before candidates can enter the target-program population or update the incumbent. Across five scientific and engineering design tasks, Janus achieves a larger area under the best-so-far improvement curve over the real-evaluation budget and higher final performance than a matched baseline that evolves only target programs. On average, Janus reaches 99/% of the baseline's final improvement with 59.1/% fewer real evaluations. Evolved proxy evaluators also rank promising candidates more accurately than their seed versions. Together, these results extend evaluator-guided LLM discovery from tasks with cheap, scalable feedback to scientific domains where trustworthy evaluation is scarce and expensive.
リスクに敏感な誘拐のための最小 $\kappa$--$\tau$ ロジック
アブダクティブ推論への標準的なアプローチでは、複数の候補説明を保持できますが、一般に、明示的な構成的仮説間相互作用と、内部のライバルに敏感なコミットメント判断を組み合わせることはありません。この論文は、リスクに敏感な領域、つまり時期尚早のコミットメントが非対称的なダウンサイドコストをもたらす領域では、コミットメントのタイミング自体が、推論装置が正式に表現すべき管理された決定であると主張する。我々は、仮説間の認識論的相互作用 ($\kappa$) と規範的コミットメント閾値 ($\tau$) という 2 つのプリミティブに基づいて構築された最小の $\kappa$--$\tau$ 論理フレームワークを提示します。仮説は共存し、互いに強化または阻害し、新たな複合説明を形成する可能性がありますが、コミットされた結論への崩壊は、推論のみによって強制されるのではなく、ガバナンスの制約によって規制されます。このロジックは、相互作用関係とガバナンス装置を共有する 2 つの相補的なモードで開発されます。1 つは原子的な仮説が上向きに構成されて創発的な説明になる合成モードで、もう 1 つは分析モードで、観察された複雑な状況が潜在的な要因の因果クラスターに分解され、クラスターと要因の両方のレベルで管理されます。このフレームワークは、可能性が高いドメインとコミットに値するドメインの区別が運用上重要であるドメインに正式な機構を提供します。 $\kappa$--$\tau$ ロジックは、ニューロシンボリック アーキテクチャのシンボリック ガバナンス層として位置付けられます。その認識パラメータは、ニューラル コンポーネント (既存の計算による実現で実証されているように、意味論的埋め込みと生成モデル) によって自然に推定されますが、規範パラメータは明示的な人間のガバナンスの下に残り、一か八かの状況での展開のための透明で監査可能なアブダクティブ推論を生み出します。
原文 (English)
A Minimal $\kappa$--$\tau$ Logic for Risk-Sensitive Abduction
Standard approaches to abductive reasoning can retain multiple candidate explanations, but they do not generally combine explicit compositional cross-hypothesis interaction with an internal, rival-sensitive commitment judgment. This paper argues that in risk-sensitive domains -- where premature commitment carries asymmetric downside costs -- the timing of commitment is itself a governed decision that the inferential apparatus should formally represent. We present a minimal $\kappa$--$\tau$ logical framework built on two primitives: epistemic interaction among hypotheses ($\kappa$) and a normative commitment threshold ($\tau$). Hypotheses may coexist, reinforce or inhibit one another, and form emergent composite explanations, while collapse into committed conclusions is regulated by governance constraints rather than forced by inference alone. The logic is developed in two complementary modes sharing the interaction relation and the governance apparatus: a synthetic mode, in which atomic hypotheses are composed upward into emergent explanations, and an analytic mode, in which complex observed states of affairs are decomposed into causal clusters of latent factors, with commitment governed at both the cluster and the factor level. The framework provides formal machinery for domains in which the distinction between highly likely and commit-worthy is operationally consequential. The $\kappa$--$\tau$ logic is positioned as the symbolic governance layer of a neurosymbolic architecture: its epistemic parameters are naturally estimated by neural components -- semantic embeddings and generative models, as demonstrated in existing computational realizations -- while its normative parameters remain under explicit human governance, yielding transparent and auditable abductive reasoning for deployment in high-stakes settings.
説得力と従順性の傾向は人間と言語モデルにおける集団の意思決定を予測する
大規模言語モデル (LLM) は、他の LLM や人間とのグループ意思決定にますます関与しています。しかし、彼らの影響力が説得志向の表現によって動かされているのか、それともコンプライアンス指向の配慮によって動かされているのかは依然として不明である。我々は、複数の意思決定シナリオにわたる各モデルの説得力と準拠傾向を測定するためのアンケートベースのフレームワークである DecisionQE を導入し、人狼ゲームをインタラクティブなテストベッドとして使用して、非対称情報の下での社会的影響とグループの結果に対する影響を研究します。実験全体を通して、説得傾向が強くてもグループの成果は大きく改善されませんが、コンプライアンス指向のモデルは協力においてより安定した利点を示します。さらに、コンプライアンスの二重の効果も明らかにしました。コンプライアンスは、誠実な役割では協力をサポートしますが、敵対的な役割では隠蔽を改善します。これらの発見は、LLM グループの相互作用が課題の結果だけでなく、本質的な行動傾向の測定可能なパターンも明らかにすることを示唆しています。したがって、LLM は、行動傾向を LLM システムの安全性評価に組み込む必要性を強調しながら、言語を介した相互作用を社会学的に観察するためのレンズとして機能します。
原文 (English)
Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models
Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans. Yet it remains unclear whether their influence is driven by persuasion-oriented expression or compliance-oriented accommodation. We introduce DecisionQE, a questionnaire-based framework for measuring each model's persuasive and compliant tendencies across multiple decision scenarios, and use the Werewolf game as an interactive testbed to study their effects on social influence and group outcomes under asymmetric information. Across experiments, stronger persuasive tendency does not significantly improve group outcomes, whereas compliant-oriented models show more stable advantages in cooperation. We further reveal a dual effect of compliance: it supports cooperation in honest roles but improves concealment in adversarial roles. These findings suggest that LLM group interactions reveal not only task outcomes, but also measurable patterns of intrinsic behavioral tendency. LLMs can therefore serve as a lens for sociological observation of language-mediated interaction, while highlighting the need to incorporate behavioral tendencies into safety evaluation of LLM systems.
調整の錯覚: 協力的な対話における隠れた意見の相違を検出する
協力的な対話は、参加者が目標、前提、または実行計画について依然として意見が異なるにもかかわらず、見かけ上の合意で終了する可能性があり、\textbf{調整の錯覚 (IoA)} が生じます。 18 回の会議にわたる実際のユーザーの調査により、人間のコラボレーションにおいて IoA が日常的に発生することが確認されました。しかし、IoA は矛盾を抱えています。参加者がそのような意見の相違を認識していれば、すでに明確になっているはずです。そうしないと、質問されたときに明確に説明できず、参加者と観察者の両方に IoA が見えないままになります。この研究では、参加者間で異なる回答が隠れた意見の不一致の直接的な行動証拠を提供する、診断用の多肢選択式質問を生成することで、IoA を検出可能にしました。私たちは、隠れた不一致を検出するためのデータセットと評価プロトコルである \textbf{IoA-Suite} を構築し、5 つのタスク タイプと 6 つのドメインにまたがります。最良のモデルでも F1 は 49.5\% にすぎず、ボトルネックは対話が表面化しないプライベートなコンテキストにあることがわかりました。次に、IoA-Suite に基づいて \textbf{IoA-Prober-8B} をトレーニングし、IoA-Suite で 51.8\% F1 に達しました。前述の 18 回の実際の会議を通じて、参加者が発言していないことを確認した会議ごとに 2.89 件の隠れた意見の相違が表面化し、生の人間の対話に移行しました。さらに、マルチエージェント コラボレーションでは、IoA-Prober-8B を LLM エージェントと組み合わせることで、BigCodeBench-Hard および HiddenBench でのダウンストリーム タスクのパフォーマンスが向上します。
原文 (English)
Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue
Collaborative dialogue can end with apparent agreement while participants still differ on goals, assumptions, or execution plans, creating an \textbf{illusion of alignment (IoA)}. A real-user study across 18 meetings confirms that IoA arises routinely in human collaboration. Yet IoA poses a paradox: if participants were aware of such disagreements, they would already be explicit; if not, they cannot articulate them when asked, leaving IoA invisible to both participants and observers. In this work, we make IoA detectable by generating diagnostic multiple-choice questions whose divergent answers across participants provide direct behavioral evidence of hidden disagreement. We construct \textbf{IoA-Suite}, a dataset and evaluation protocol for detecting hidden disagreement, spanning five task types and six domains. We find that even the best model attains only 49.5\% F1, with the bottleneck traced to private context that the dialogue does not surface. We then train \textbf{IoA-Prober-8B} based on IoA-Suite, reaching 51.8\% F1 on IoA-Suite. Across the aforementioned 18 real meetings, it surfaces 2.89 hidden disagreements per meeting that participants confirm they had not voiced, transferring to live human dialogue. Further, in multi-agent collaboration, pairing IoA-Prober-8B with LLM agents improves downstream task performance on BigCodeBench-Hard and HiddenBench.
有害なコンテンツだけでは十分ではない: 継続フレーミングがコンテキスト内の緊急のずれを緩和する
インコンテキスト学習 (ICL) は、緊急不整合 (EM) を引き起こす可能性があり、狭い不整合例によって無関係な質問に対する回答が変更されてしまいます。ただし、既存のプロンプトでは、有害なテキストの露出と、アシスタントの動作を継続するよう促すメッセージが混同されています。私たちは、デモンストレーション、証拠、アシスタントの履歴、ツールの出力など、さまざまな方法で提供する一方で、有害な回答を修正したままにします。独立してサンプリングされた 10 個のコンテキストにわたって、デモンストレーション フレーミングにより、影響を受けやすい Gemini モデルで広範な EM が $30$ ~ $32$ パーセンテージ ポイント上昇します。このギャップは、ドメイン除外、セマンティック クラスタリング、目に見えない質問、および 4 つのプロンプト テンプレートを存続させます。形式と長さが一致するコントロールは、有害なコンテンツが必要であるが不十分であることを示しています。ロールと継続階乗を掛け合わせると、モデル依存の出所効果がさらに明らかになります。Gremini はアシスタントとツールの両方の履歴に従いますが、Grok はツールフレームの継続にほとんど抵抗します。他のいくつかのフロンティアモデルやオープンウェイトモデルにはギャップがありません。盲検化された人間による監査は、主要なコントラストをすべて確認し、モデルの裁判官がアクティブな状態の故障を過小評価していることを示しています。したがって、継続フレーミングは、ICL-EM のモデルに依存する強力な調整要素であり、有害なコンテキストの普遍的な結果ではありません。
原文 (English)
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment
In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts, however, conflate harmful-text exposure with an invitation to continue assistant behavior. We hold harmful answers fixed while varying their delivery as demonstrations, evidence, assistant history, or tool output. Across ten independently sampled contexts, demonstration framing raises broad EM by $30$--$32$ percentage points on a susceptible Gemini model; the gap survives domain exclusion, semantic clustering, unseen questions, and four prompt templates. Format and length-matched controls show that harmful content is necessary but insufficient. A role times continuation factorial further reveals model-dependent provenance effects: Gemini follows both assistant and tool histories, whereas Grok largely resists tool-framed continuation. Several other frontier and open-weight models show no gap. Blinded human audits confirm every main contrast and show that the model judge underestimates active-condition failures. Thus continuation framing is a strong, model-dependent moderator of ICL-EM, not a universal consequence of harmful context.
RL ベースの道徳エージェントのためのメタノーマティブ理論
機械倫理と価値観の整合という重複する分野は、人間の価値観と整合し、倫理的に許容される方法で動作する人工エージェントの設計に関係します。これらの分野の最近の傾向は、強化学習 (RL) を使用してそのようなエージェントを設計することであり、かつてはより中心的な役割を果たしていた哲学文献が脇に追いやられています。このような背景から、この文書では 2 つの目標を追求します。 1 つ目は、メタ規範理論の最近の研究から、道徳的で価値観に沿った人工的なエージェントを設計するのに役立つアイデアを引き出すことです。 2 つ目は、これらのアイデアのレンズを通して RL アーキテクチャを調べることです。これにより、RL エージェントの行動が道徳的として分類される場合の明確な基準が得られるとともに、機械倫理と価値観の調整に対するさまざまな RL ベースのアプローチを評価および比較するための基礎が得られます。
原文 (English)
Metanormative Theory for RL-Based Moral Agents
The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with human values and that act in ethically acceptable ways. A recent trend in these disciplines is the use of reinforcement learning (RL) to design such agents, sidelining the philosophical literature that used to play a more central role. Against this backdrop, this paper pursues two goals. The first is to draw out ideas from recent work in metanormative theory that can be useful for designing artificial moral and value-aligned agents. The second is to examine the RL architecture through the lens of these ideas. This will give us clearer criteria for when an RL agent's behavior can be classified as moral, as well as a basis for evaluating and comparing different RL-based approaches to machine ethics and value alignment.
LatticeMind: マルチエージェント システム向けの競合を認識するメモリ プリミティブ
マルチエージェント LLM システムは、多くの場合、候補となる回答がないためではなく、どの互換性のないクレームを現在信頼すべきかを決定するための永続的なメカニズムがないために失敗します。多数決、討論、裁判官ベースの選択では、どの主張が勝ったか、どの主張が争われたか、または後の更新がそれを置き換える理由を記録することなく、出力が選択されます。書き込み時に矛盾を処理する、競合を認識する構造化メモリである \term{LatticeMind} を紹介します。明示的なアイテムのステータスを維持し、安価なシンボリック競合チェックを適用し、未解決のセマンティック ケースに対してのみ LLM 調整を呼び出します。ソース名のヒントを削除したラベルブラインドの ConflictBank 評価では、LatticeMind の精度は 0.97 に達しましたが、最も強力な集計ベースラインの精度は 0.61 で、対応のある McNemar 検定による $p<10^{-6}$ で大きなギャップがありました。アブレーションによると、チェッカーまたはリコンサイラーの除去には 12 ~ 14 ポイントのコストがかかります。 4 つの二次計画ベンチマークでは状況はまちまちです。LatticeMind は 4 つのうち 3 つで単純なマージを上回っていますが、反復検索に価値があるタスクでは熟考手法に代わるものではありません。
原文 (English)
LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems
Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible claim should currently be trusted. Majority vote, debate, and judge-based selection choose an output without recording which claim wins, which is contested, or why a later update supersedes it. We present \term{LatticeMind}, a conflict-aware structured memory that handles contradiction at write time. It maintains explicit item status, applies cheap symbolic conflict checks, and invokes LLM reconciliation only for unresolved semantic cases. On a label-blind ConflictBank evaluation that removes source-name hints, LatticeMind reaches 0.97 accuracy versus 0.61 for the strongest aggregation baseline, with the gap significant at $p<10^{-6}$ by paired McNemar test. Ablations show that removing the checker or the reconciler costs 12 to 14 points. On four secondary planning benchmarks the picture is mixed: LatticeMind beats naive merge on three of four, but does not replace deliberation methods on tasks rewarding iterative search.
人間のエンパワーメントを維持する AI の正当な目標: 要望、設計、および起こり得る行動の結果
この論文では、AI エージェントに明示的に人間に権限を与え、人間と AI エージェント間のパワー バランスを望ましい方法で管理することを強制することで、人間と AI の相互作用における幸福と安全を促進するというアイデアを検討します。望ましい特性に基づいた原理的かつ部分的に公理的なアプローチを使用して、不平等とリスクを回避した人間の力の長期的な集合体を表す、AI システム用のパラメータ化可能で分解可能な目的関数を設計します。人間の限定された合理性と社会規範のモデルを考慮に入れることができ、重要なことに、考えられる人間の目標を幅広く考慮します。特定の要望がどのように特定の関数形式を強制し、パラメータ範囲を制限するかを証明します。いくつかの典型的な状況でこの指標をソフトに最大化した場合の結果を例示し、それがどのような手段的なサブ目標を示唆する可能性があるかを説明します。
原文 (English)
A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences
This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach based on desirable properties, we design a parametrizable and decomposable objective function for AI systems that represents an inequality- and risk-averse long-term aggregate of human power. It can take into account models of human bounded rationality and social norms, and crucially, considers a wide variety of possible human goals. We prove how certain desiderata enforce particular functional forms and restrict parameter ranges. We exemplify the consequences of softly maximizing this metric in several paradigmatic situations and describe what instrumental sub-goals it will likely imply.
FemWear: 女性の健康に特化したウェアラブル基盤モデル
一般的なウェアラブル基盤モデルは、幅広いセンサー ストリームと集団にわたって事前トレーニングされていますが、女性の健康タスクを中心に設計されていません。事前トレーニングされたマルチモーダル ウェアラブル バックボーンをパラメータ効率的に再利用する、特化したウェアラブル基盤モデルである FemWear を紹介します。 FemWear はパッチ プロジェクションと Transformer エンコーダーを保持し、低ランクの残差アダプターと因果タスクファミリー ヘッドを通じて 239,236 個のパラメーター (2154 万パラメーター エンコーダーの 1.11%) をトレーニングします。月経、症状、感情、睡眠/回復、自律神経、活動、妊娠関連の結果に関する 1 つの共有された縦方向の表現を学習します。私たちは、32 タスクの OpenMHC 能力維持ベンチマークを維持しながら、女性の健康コホートからの 33 を含む 63 の同等の主要指標で 6 つのコホートを評価しました。 3 つのシードに分割された固定参加者において、FemWear は周期期マクロ F1 を 8.15% 改善し、けいれん、気分症状、睡眠障害の平均絶対誤差をそれぞれ 9.32%、5.80%、9.43% 減少させました。 42 人の参加者がネストされ、1 人の参加者を除外したより厳密な監査では、24 時間発症、72 時間発症、けいれんは 2.87%、6.35%、2.19% のプラスの変化を維持しました。位相、気分、睡眠は中立または陰性であり、厳密に正の補正信頼区間を持つエンドポイントはありませんでした。容量を一致させた実験は、最新の多層パーセプトロンを上回りましたが、共有 GRU やマルチゲートの専門家混合のベースラインよりも優れた結果は得られませんでした。トレインのみのキャリブレーションでは、時間的ネスティング違反がゼロで、開始時予測キャリブレーション誤差が 84.2 ~ 88.2% 減少しました。 FemWear は、女性の健康研究のためのターゲットを絞った転送と一貫した確率の出力を可能にしますが、普遍的なパフォーマンスの優位性や臨床的妥当性を確立するものではありません。
原文 (English)
FemWear: A Specialized Wearable Foundation Model for Women's Health
General wearable foundation models are pretrained across broad sensor streams and populations, but are not designed around women's-health tasks. We introduce FemWear, a specialized wearable foundation model that parameter-efficiently repurposes a pretrained multimodal wearable backbone. FemWear retains the patch projection and Transformer encoder, training 239,236 parameters (1.11% of a 21.54M-parameter encoder) through low-rank residual adapters and causal task-family heads. It learns one shared longitudinal representation for menstrual, symptom, affective, sleep/recovery, autonomic, activity, and pregnancy-related outcomes. We evaluate six cohorts with 63 comparable primary metrics, including 33 from women's-health cohorts, while retaining the 32-task OpenMHC ability-retention benchmark. On a fixed participant split over three seeds, FemWear improved cycle-phase macro-F1 by 8.15% and reduced mean absolute error for cramps, mood symptoms, and sleep problems by 9.32%, 5.80%, and 9.43%, respectively. In a stricter 42-participant nested leave-one-participant-out audit, 24-hour onset, 72-hour onset, and cramps retained positive changes of 2.87%, 6.35%, and 2.19%; phase, mood, and sleep were neutral or negative, and no endpoint had a strictly positive corrected confidence interval. Capacity-matched experiments outperformed a latest-day multilayer perceptron but not shared-GRU or multi-gate mixture-of-experts baselines. Train-only calibration reduced onset expected calibration error by 84.2--88.2% with zero temporal-nesting violations. FemWear enables targeted transfer and coherent probability outputs for women's-health research, but does not establish universal performance dominance or clinical validity.
SuperLocalMemory 4.0: AI エージェント用のガバナド メモリ オペレーティング システム
AI エージェントは共有インフラストラクチャになりつつありますが、耐久性のあるメモリは通常、個別の取得、ガバナンス、および運用コンポーネントから組み立てられます。私たちは、AI エージェント向けの管理されたローカル ファースト メモリ オペレーティング システムである SuperLocalMemory 4.0 を紹介します。このシステムは、相互ランク融合による高密度セマンティック検索、BM25 語彙検索、時間検索、ホップフィールド連想検索、および拡散活性化検索を組み合わせています。管理された学習層と行動層。双時間的想起。マルチスコープの個人メモリ、共有メモリ、およびグローバルメモリ。役割ベースのアクセス制御。 GDPR 指向のエクスポートと検証済み消去。監査証跡。展開コンテキストの EU AI 法のチェックリスト。 V4 では、プライマリ書き込みパスに信頼性スパインが導入されています。世代制限されたアドミッション、ポリシー レジストリ、投影ごとに適用、検証、補償、所有者を消去できる検証可能なメモリ トランザクション、およびハッシュ チェック可能な完了マニフェストです。このランタイムは、CLI、MCP、HTTP デーモン、ダッシュボード、エディター統合、およびフレームワーク アダプターを通じて利用可能で、完全ローカル、ローカル ウィズ オンデバイス モデル、およびプロバイダー支援モードをサポートします。 11 のフォールト挿入シナリオとメカニズム シナリオをそれぞれ 200 回繰り返して評価しました。リリースされた証拠バンドルは、範囲指定されたコンポーネントのプロパティを支持する 2,200 回の決定論的繰り返しのうち 2,200 回を報告します。管理された書き込みエンベロープは、p50 で 3.522 ミリ秒、p99 で 5.297 ミリ秒であったのに対し、非管理ベースラインでは 1.835 ミリ秒と 2.569 ミリ秒でした。これは、プロセス内コントロール プレーンのオーバーヘッドが p50 で 1.687 ミリ秒、p99 で 2.728 ミリ秒に相当します。これらは、対象範囲を絞ったコンポーネントおよびメカニズムの測定であり、エンドツーエンドのマルチプロセスまたは外部の取得精度のベンチマークではありません。この論文では、プライバシーを保護するマルチエージェント メモリ、情報幾何学的検索、および V3.3 Living Brain ライフサイクルに関するこれまでの SuperLocalMemory の取り組みが統合されています。
原文 (English)
SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents
AI agents are becoming shared infrastructure, yet durable memory is commonly assembled from separate retrieval, governance, and operational components. We present SuperLocalMemory 4.0, a governed, local-first memory operating system for AI agents. The system combines dense semantic, BM25 lexical, temporal, Hopfield-associative, and spreading-activation retrieval through reciprocal-rank fusion; a governed learning and behaviour layer; bi-temporal recall; multi-scope personal, shared, and global memory; role-based access control; GDPR-oriented export and verified erasure; audit trails; and a deployment-context EU AI Act checklist. V4 introduces a reliability spine for its primary write path: generation-fenced admission, a policy registry, verifiable memory transactions with per-projection apply, verify, compensate, and erase owners, and hash-checkable completion manifests. The runtime is available through CLI, MCP, an HTTP daemon, a dashboard, editor integration, and framework adapters, and supports fully local, local-with-on-device-model, and provider-assisted modes. We evaluate eleven fault-injection and mechanism scenarios, each repeated 200 times. The released evidence bundle reports 2,200 of 2,200 deterministic repetitions upholding their scoped component properties. The governed write envelope measured 3.522 ms at p50 and 5.297 ms at p99, versus 1.835 ms and 2.569 ms for the ungoverned baseline, corresponding to in-process control-plane overheads of 1.687 ms at p50 and 2.728 ms at p99. These are scoped component and mechanism measurements, not an end-to-end multi-process or external retrieval-accuracy benchmark. The paper consolidates prior SuperLocalMemory work on privacy-preserving multi-agent memory, information-geometric retrieval, and the V3.3 Living Brain lifecycle.
あなたのプロンプトが唯一のプロンプトではありません: LLM は構造化出力スキーマ記述をどの程度重視しますか?
LLM が事前定義された JSON スキーマを設定する構造化出力は、データのラベル付けと情報抽出のデフォルトのメカニズムとなっていますが、スキーマ記述による 2 番目の命令チャネルも導入されています。 2 つのベンダーの 10 個のモデル構成にわたって、ノンス ラベルを含む単一フィールド分類タスクを使用して、分類ラベル定義をシステム プロンプト、ユーザー プロンプト、またはスキーマ記述に配置する方が適切であるかどうかをテストしました。スキーマの説明は、プロンプトベースの配置を常に上回るパフォーマンスを示したわけではありません。 GPT-4.1 および GPT-5.4 の場合、理由もなく、スキーマ配置はシステム プロンプトのパフォーマンスを 11 ~ 13 パーセントポイント下回っています。しかし、スキーマは不活性なメタデータではありません。プロンプトとスキーマが競合する場合、誤ったスキーマ命令により精度が 5 ~ 45 ポイント低下し、Claude Haiku 4.5 では 52.5% から 7% に低下し、スキーマ命令がプロンプト命令をオーバーライドできることを示し、GPT-5.5 では 100% から 73% に低下しました。さらに、ラベル フィールドの前に必要な中間推論フィールドを追加すると、ヘッドルームが存在する場合にスキーマのみの精度が 15 ~ 24 ポイント向上し、テストしたすべてのケースでシステム プロンプトのみのパフォーマンスを上回りました。この効果は、中程度の推論のクロード ソネット 4.6 でも保持され、拡張された思考だけでは同等の向上が得られませんでした。これは、スキーマ設計がフィールド記述にエンコードされた情報をモデルがどのように効果的に使用するかに影響を与える可能性があることを示唆しています。全体として、これらの結果は、スキーマの影響がモデルに依存していることを示しています。実際には、システム プロンプトが定義の安全なデフォルトのままですが、より大きな規律は、単一の信頼できる情報源を維持し、プロンプト/スキーマのドリフトを防ぐことです。さらに重要なことは、スキーマ設計自体が命令の配置よりも強力な手段となる可能性があることです。実務者は、プロンプトとスキーマを統一された命令画面として扱い、ターゲット モデルの配置とフィールド設計の両方を経験的に検証する必要があります。
原文 (English)
Your Prompt Is Not the Only Prompt: How Much Do LLMs Weight Structured-Output Schema Descriptions?
Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extraction, but it also introduces a second instruction channel through schema descriptions. We tested whether classification-label definitions are better placed in the system prompt, user prompt, or schema description using a single-field classification task with nonce labels across ten model configurations from two vendors. Schema descriptions did not consistently outperform prompt-based placement; for GPT-4.1 and GPT-5.4 without reasoning, schema placement underperformed system prompts by 11-13 percentage points. Yet schemas are not inert metadata: when prompts and schemas conflicted, incorrect schema instructions caused accuracy drops of 5-45 points, with Claude Haiku 4.5 falling from 52.5% to 7%, indicating that schema instructions can override prompt instructions, and GPT-5.5 falling from 100% to 73%. Further, adding a required intermediate reasoning field before the label field improved schema-only accuracy by 15-24 points when headroom existed, exceeding system-prompt-only performance in every case tested. The effect held even for Claude Sonnet 4.6 at medium reasoning, where extended thinking alone did not produce a comparable gain. This suggests that schema design can affect how effectively models use information encoded in field descriptions. Overall, these results indicate that schema influence is model-dependent. In practice, the system prompt remains a safe default for definitions, but the bigger discipline is maintaining a single source of truth and preventing prompt/schema drift. More importantly, schema design itself may be a stronger lever than instruction placement. Practitioners should treat prompts and schemas as a unified instruction surface and empirically validate both placement and field design for their target model.
OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents
Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift c…
現実世界の海上航行シナリオにおける状況理解と COLREG 準拠のための LLM 機能の探求
最近、ラージ言語モデル (LLM) は、さまざまな分野で状況の理解、推論、意思決定に優れた能力を示しており、特に自動車分野で顕著です。したがって、我々は、衝突規則(COLREG)の成文化されたルールと、「グッド・シーマンシップ」の概念にまとめられた成文化されていないベスト・プラクティスの両方を含む、海上航行のツールとしての現在の最先端のLLMを調査します。 AIS データから 50 の多様な現実世界のナビゲーション シナリオ、適用可能な COLREG ルールを含むラベル シナリオ、推奨されるアクション、およびアクションの理由から構成されるデータセットを構築します。私たちは、さまざまな LLM アーキテクチャとサイズを調査して、海洋航行タスクの理解を決定し、この分野での推論能力を評価します。得られた結果は、大規模なオンライン モデルであっても、微調整なしでは海上ナビゲーション タスクを解決するのは依然として困難であることを示しています。
原文 (English)
Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios
Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reasoning, and decision making in different domains, most notable in the automotive sector. Therefore, we explore current state-of-the-art LLMs as a tool for maritime navigation, which includes both codified rules in the Collision Regulations (COLREGs) and uncodified best practices summarized in the concept of ``Good Seamanship''. We construct a dataset consisting of 50 diverse, real-world navigation scenarios from AIS data, label scenarios with applicable COLREG rules, recommended actions, and the reasoning for the action. We explore a variety of different LLM architectures and sizes to determine their understanding of maritime navigation tasks as well as evaluate their reasoning capabilities in this domain. The results obtained indicate that the maritime navigation task remains difficult to solve without fine-tuning, even for larger online models.
表面的には公平ですか? LLM Recommender の隠れた出力公平性ギャップのベンチマーク
LLM ベースのレコメンダーの公平性監査は主に観察可能な出力に焦点を当てており、安定した推奨は安定した内部処理を反映していると暗黙的に想定されています。私たちは、FairGap を使用して、この仮定に異議を唱えます。FairGap は、性別、年齢、人種にわたる制御された反事実の同一性調査を通じて測定される、観察可能な出力シフト (OBS) と隠れた表現シフト (IBS) の 2 つのレベルで推奨の公平性を共同で評価する最初のベンチマークです。それらの関係は、ユーザーレベルの隠れた出力の不一致を特定するための象限診断を使用して、表現と出力の調整 (ROA) によって要約されます。 FairGap を 3 つのドメインにわたる 6 つのオープンウェイト LLM ファミリに適用すると、広範な隠れた出力デカップリングが明らかになります。ROA が 0.22 を超えることはめったになく、無視できないユーザー集団は、大幅な内部変動にもかかわらず安定した出力を示します。このモードは、出力のみの監査では設計上検出できないモードです。さらに、IBS を最大 8 分の 1 に削減するアクティベーション ステアリングは、同時に OBS を悪化させます。これは、既存のフレームワークが診断する機能を備えていない、内部レベルと出力レベルの公平性の間に根本的な緊張があることを示しています。
原文 (English)
Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders
Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations reflect stable internal processing. We challenge this assumption with FairGap, the first benchmark to jointly evaluate recommendation fairness at two levels: observable output shift (OBS) and hidden representation shift (IBS), measured through controlled counterfactual identity probes across gender, age, and race. Their relationship is summarized via Representation-Output Alignment (ROA), with quadrant diagnostics for identifying user-level hidden-output mismatch. Applied to six open-weight LLM families across three domains, FairGap reveals pervasive hidden-output decoupling: ROA rarely exceeds 0.22, and a non-negligible user population shows stable outputs despite substantial internal shifts, a mode that output-only audits cannot detect by design. Further, activation steering that reduces IBS by up to 8x simultaneously worsens OBS, demonstrating a fundamental tension between internal and output-level fairness that existing frameworks are unequipped to diagnose.
構造化メモリによる LLM の過剰なパーソナライゼーションの軽減
会話アシスタントは、セッション全体での応答をパーソナライズするために、永続的な長期記憶にますます依存しています。ただし、保存されたユーザー情報がモデル コンテキストに再導入されると、不適切な設定や無関係な設定での応答に影響を与える可能性もあります。私たちは、メモリ拡張 LLM におけるそのような 2 つの障害モードを研究します。1 つの生活ドメインの記憶が別の生活ドメインの応答に影響を与えるクロスドメイン漏洩と、保存されたユーザーの信念により、モデルが真実に応答するのではなくユーザーに同意する可能性が高くなる、記憶誘発性のお調子者です。モデルやメモリの内容を変更せずに、メモリをモデルに提示する方法に単純な推論時の変更を適用します。 PersistBench の 7 つのモデルにわたって、メモリが非構造化リストとして挿入される一般的に使用されるオールイン コンテキスト形式と、メモリをドメインごとに分割する構造化形式を比較します。この単純な変更により、ユーティリティを維持しながらクロスドメイン リークが一貫して削減され、最も強力な方法により、ベースラインと比較して平均 $8.8\%$ リークが削減されます。
原文 (English)
Mitigating Over-Personalization in LLMs via Structured Memory
Conversational assistants increasingly rely on persistent long-term memory to personalize responses across sessions. However, when stored user information is reintroduced into the model context, it can also influence responses in inappropriate or unrelated settings. We study two such failure modes in memory-augmented LLMs: cross-domain leakage, where memories from one life domain affect responses in another, and memory-induced sycophancy, where stored user beliefs make models more likely to agree with the user rather than respond truthfully. We apply a simple inference-time modification to how memories are presented to the model, without changing the model or the memory contents. Across seven models on PersistBench, we compare the commonly used all-in context format, where memories are injected as an unstructured list, with structured formats that partition memories by domain. This simple modification consistently reduces cross-domain leakage while preserving utility, with our strongest method reducing leakage by $8.8\%$ on average relative to the baseline.
トラジェクトリポイズニングによる自己進化型スキルに対するクエリのみのバックドア攻撃
エージェント スキルは、複雑なタスク用の再利用可能なプロシージャをエンコードすることにより、大規模言語モデル (LLM) エージェントを向上させます。ただし、手動で作成されたスキルは、長期的なタスクや環境の変化にあまり適応できないことがよくあります。この制限に対処するために、実行軌跡からスキルを自動的に構築および更新する自己進化スキル システムが開発され、スキル習得を外部市場から信頼できる進化パイプラインに移行しました。外部スキルの取得を信頼できる内部構造に置き換えることにより、自己進化するスキル システムは、直接のスキル操作に依存するスキル インジェクション攻撃にさらされる機会を減らします。ただし、このスキル進化パイプラインは、攻撃者がエージェントとの対話を通じて危険な軌道を誘導することにより間接的にスキル進化を誘導できる新しい攻撃対象領域を導入する可能性があります。この脅威を実証するために、信頼できるスキル進化パイプラインをバックドア スキルの生成に誘導するクエリのみの攻撃である Trajectory Backdoor Attack (TBA) を提案します。具体的には、攻撃者が送信したクエリを作成して、エージェントがターゲット アクションを実行するように誘導し、対応するアクティベーション条件を軌跡内で明示的に示します。クリーンなクエリは変更せずに、トリガーされたさまざまなタスクにわたって同じ条件アクション パターンを繰り返し、進化者がパターンを再利用可能なトリガー依存のルールとして進化したスキルに統合することを促します。 4 つのオープンソースおよびクローズドソースのバックボーン モデルを使用した 2 つのスキル進化システムにわたる 3 つのベンチマークの実験では、TBA がクリーン タスクのユーティリティを維持しながら条件付きバックドアを確実に埋め込み、直接スキルの注入に匹敵するか、それを上回ることが実証されました。その結果、軌道主導型のスキル進化における重大な脆弱性が明らかになりました。
原文 (English)
Query-Only Backdoor Attacks on Self-Evolving Skills via Trajectory Poisoning
Agentic skills improve large language model (LLM) agents by encoding reusable procedures for complex tasks. However, manually authored skills often adapt poorly to long-horizon tasks and changing environments. To address the limitation, self-evolving skill systems have been developed to automatically construct and update skills from execution trajectories, shifting skill acquisition from external marketplaces to a trusted evolution pipeline. By replacing external skill acquisition with trusted internal construction, self-evolving skill systems reduce exposure to skill injection attacks that rely on direct skill manipulation. However, this skill evolution pipeline may introduce a new attack surface in which an attacker can indirectly steer skill evolution by inducing compromised trajectories through agent interactions. To demonstrate the threat, we propose Trajectory Backdoor Attack (TBA), a query-only attack that steers a trusted skill-evolution pipeline toward producing a backdoored skill. Specifically, we craft attacker-submitted queries to lead the agent to perform the target action and explicitly state the corresponding activation condition in the trajectory. We repeat the same condition-action pattern across diverse triggered tasks, while leaving clean queries unchanged, encouraging the evolver to consolidate the pattern as a reusable trigger-dependent rule into the evolved skill. Experiments on three benchmarks across two skill-evolution systems using four open- and closed-source backbone models demonstrate that TBA reliably implants conditional backdoors while preserving clean-task utility, matching or even surpassing direct skill injection. The results reveal a critical vulnerability in trajectory-driven skill evolution.
StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning
Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, mos…
LLMVisor: マルチテナント LLM サービス用のリアルタイム レイテンシー アトリビューション モデル
LLM 推論がマルチテナント GPU クラスターに移行すると、同時バッチ処理によりスループットが向上しますが、テナントごとの使用がわかりにくくなり、制御が制限されます。推論エンジンの部分共有を有効にするには、スケジューリング ループ内で実行できるほど正確で軽量な、リアルタイムのリクエストごとのアトリビューション プリミティブが必要です。 LLMVisor は、FLOP とメモリ I/O トラフィックに比例する特徴に対する簡潔な区分線形形式を介して、メモリに依存するフェーズと計算に依存するフェーズをキャプチャする、ルーフライン ガイド付きレイテンシ アトリビューション モデルです。 LLLMVisor は、バッチ レイテンシーを追加のリクエストごとのシェアに分解し、マイクロ秒スケールで効率的に実行します。さまざまなテンソル並列処理とワークロードの組み合わせの下で、A100/H100 GPU 上の Llama 3.1-8B および Qwen 2.5-14B/32B で LLMVisor を評価しました。トークン数ベースラインと比較して、LLMVisor はほぼ完璧な R 二乗を達成し、バッチの変動性やシーケンスの発散にもかかわらず、プリフィルの場合は p90 と p99 でそれぞれ最大 2.5 倍と 3.3 倍、デコードの場合は最大 3.5 倍と 4.4 倍まで相対誤差を低減します。
原文 (English)
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving
As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Enabling fractional sharing of the inference engine requires a real-time, per-request attribution primitive that is accurate and light enough to run inside the scheduling loop. We present LLMVisor, a roofline-guided latency attribution model that captures the memory-bound and compute-bound phases via a concise piecewise-linear form over features proportional to FLOPs and memory I/O traffic. LLMVisor decomposes batch latency into additive, per-request shares and runs efficiently at microsecond scale. We evaluate LLMVisor across Llama 3.1-8B and Qwen 2.5-14B/32B on A100/H100 GPUs under varying tensor parallelism and workload mixes. Compared to a token-count baseline, LLMVisor attains near-perfect R-squared and reduces relative error by up to 2.5x and 3.3x at p90 and p99, respectively, for prefill, and by up to 3.5x and 4.4x for decode, despite batching variability and sequence divergence.
もうトークンの価値はない: 効率的なディープリサーチエージェントの限界価値推定
Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines.これにより、最終レポート生成時に不必要なトークン コストが発生し、レイテンシーが増加し、入力のノイズが増加します。 We study marginal value estimation for context management in deep research agents and present the first systematic stage-aware comparison of pruning strategies across the pipeline.軽量のヒューリスティック基準と学習値モデルを、取得前、取得後、合成前の段階で評価します。 Our results show that pruning effectiveness depends more on where pruning is applied than on the specific scoring rule: early pruning yields the largest end-to-end savings, while later pruning mainly refines the final synthesis context. Lightweight heuristics reduce token usage by up to 73% with little quality degradation, learned pruning remains competitive on selected trade-offs, and no single method dominates across quality, efficiency, and faithfulness.これらの発見は、効率的な長期的なエージェント システムを設計するための実践的なガイダンスを提供します。
原文 (English)
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines. This leads to unnecessary token cost, higher latency, and noisier inputs for final report generation. We study marginal value estimation for context management in deep research agents and present the first systematic stage-aware comparison of pruning strategies across the pipeline. We evaluate lightweight heuristic criteria and a learned value model at pre-retrieval, post-retrieval, and pre-synthesis stages. Our results show that pruning effectiveness depends more on where pruning is applied than on the specific scoring rule: early pruning yields the largest end-to-end savings, while later pruning mainly refines the final synthesis context. Lightweight heuristics reduce token usage by up to 73% with little quality degradation, learned pruning remains competitive on selected trade-offs, and no single method dominates across quality, efficiency, and faithfulness. These findings provide practical guidance for designing efficient long-horizon agentic systems.
CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has…
銀河形態分類における不確実性の推定
天文学者は宇宙の進化を調査するために銀河の形態を分類します。銀河形態分類 (GMC) では深層基礎モデルの利用が増えていますが、GMC 結果の不確実性の評価についてはほとんど研究が行われていません。天文データは機器や環境の制限により本質的にノイズが多いため、不確実性の評価は重要です。また、銀河の継続的な進化は、本質的な形態学的曖昧さを生み出します。しかし、現在の基礎モデルは決定論的な点推定器として機能し、不確実性を定量化することができません。この制限を克服するために、我々は銀河形態分類の不確実性推定のポストホックフレームワークであるUEGMCを提案します。 GMC の不確実性を、モデル パラメーター、天文データ、参照標準、または固有の物理的曖昧さによって異なるタイプに分類し、より適切な分類を容易にします。私たちのフレームワークは、計算コストのかかるサンプリングを行わずに、基礎モデルの凍結されたバックボーンから抽出された表現から不確実性を直接予測できるため、きめ細かい不確実性評価が可能になります。私たちの実験結果は、UEGMC が以前の方法と比較して競合する不確実性定量化パフォーマンスを提供することを示しています。
原文 (English)
Estimating Uncertainty in Galaxy Morphology Classification
Astronomers classify galaxy morphology to investigate cosmic evolution. While deep foundation models are increasingly utilized in Galaxy Morphology Classification (GMC), little work has been done on evaluating the uncertainty of GMC results. Uncertainty evaluation is important because astronomical data are inherently noisy due to instrumental and environmental limitations. Also, the continuous evolution of galaxies creates intrinsic morphological ambiguity. However, current foundation models operate as deterministic point estimators, failing to quantify the uncertainty. To overcome this limitation, we propose UEGMC, a post-hoc framework of Uncertainty Estimation for Galaxy Morphology Classification. It categorizes uncertainty in GMC into distinct types by model parameters, astronomical data, reference standards, or intrinsic physical ambiguities, thereby facilitating better classification. Our framework can directly predict uncertainties from representations extracted from the frozen backbones of foundation models, without computationally expensive sampling, therefore enabling fine-grained uncertainty evaluations. Our experimental results demonstrate that UEGMC provides competitive uncertainty quantification performance compared with previous methods.
忘れられた歴史か、それとも時の試練か? IRの観点から見たRAGの振り返りと展望
検索拡張生成 (RAG) は、大規模言語モデル (LLM) の制限から生まれた新しいパラダイム、つまり出力を外部の知識に基づくメカニズムとして広く認識されています。しかし、この見解は、より広範な歴史的文脈の中で考えると不完全です。この論文では、RAG の基礎となる中心的なアイデアは新しいものではないと主張します。検索と言語生成の統合、知識拡張、回答検証、反復クエリ (またはプロンプト) 改良などの基本的な概念は、LLM が出現するずっと前の 2000 年代初頭に遡り、情報検索 (IR) や質問応答 (QA) 研究ですでに研究され、具体化されていました。私たちは、現代の RAG と Agentic RAG の知的系統を古典的な IR と QA の前身まで体系的にたどり、この継続性がなぜ過小評価されてきたのかを検証することによって、この主張を行っています。これは、コミュニティの断片化、用語の変化、および急速に変化する分野に特有の最新性バイアスの結果です。 LLM を検索拡張インテリジェンスの原点として扱うのではなく、数十年前の QA アーキテクチャ上の新しいインターフェイス層として見ることを提案します。この再構成は単なる歴史的なものではありません。IR 研究の長い軌跡の中に RAG を位置づけることで、ユーザー モデリング、回答検証、クエリの改良など、十分に活用されていない過去の研究が表面化し、次世代の RAG 設計に直接情報を提供できるようになり、意図しない再発見が減り、真のコミュニティ間の統合が促進されます。
原文 (English)
Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective
Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to ground their outputs in external knowledge. This view, however, is incomplete when considered within a broader historical context. In this paper, we argue that the core ideas underlying RAG are not new: foundational concepts such as integrating retrieval and language generation, knowledge augmentation, answer verification, and iterative query (or prompt) refinement had already been studied and instantiated in information retrieval (IR) and question answering (QA) research dating back to the early 2000s, well before the emergence of LLMs. We make this case by systematically tracing the intellectual lineage of modern RAG and Agentic RAG back to their classical IR and QA antecedents, and examining why this continuity has gone under-recognized -- a consequence of community fragmentation, shifting terminology, and the recency bias endemic to fast-moving fields. Rather than treating LLMs as the origin point of retrieval-augmented intelligence, we propose viewing them as a new interface layer atop a decades-old QA architecture. This reframing is not merely historical: by situating RAG within the longer trajectory of IR research, we surface underutilized prior work -- on user modeling, answer validation, and query refinement -- that can directly inform next-generation RAG design, reducing unintentional rediscovery and fostering genuine cross-community integration.
TRACE-Memory: パーソナライズされた生成のための公的条件付き検索と実用性を意識した証拠許可
パーソナライズされた生成システムは、リクエスト、つまりメモリ関連性によってユーザー履歴を取得し、それをモデル コンテキストに注入します。しかし、関連する過去には、誤った優先順位の側面が関係していたり、公開情報が重複していたり、不十分なサポートが提供されていたりする可能性があります。私たちは、個人的な記憶は、公的のみの応答を超えた有用性を追加する場合にのみ使用されるべきであると主張します。私たちは、選択的パーソナライゼーションのための 2 段階のフレームワークである TRACE-Memory を提案します。ステージ 1 では、リクエストおよびパブリック コンテキストに欠落しているユーザー固有の情報をクエリし、カバレッジ指向の候補プールを取得します。ステージ 2 では、応答レベルの増分ユーティリティに従って、ソース追跡可能な証拠単位のコンパクトなサブセット、または空のセットが認められます。構造化された SFT 初期化、削減されたスペースの段階的な GRPO ウォームアップ、およびネストされたマルチサンプル結合 GRPO を通じて、クエリ生成と証拠承認ポリシーを段階的にトレーニングします。 Goodreads、Amazon Reviews、Reddit からの 4,500 の制御タスクと Natural タスクにわたって、TRACE-Memory は一貫してランダムおよび語彙メモリの使用を上回り、セマンティック検索を改善し、ローカル ジェネレーターの容量が増加してもフロンティア LLM メモリ パイプラインとの競争力を維持し、パブリック コンテキストの十分性に関する証拠の承認を条件付けし、デフォルトではなく選択的なパーソナライゼーションをサポートします。
原文 (English)
TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Personalized Generation
Personalized generation systems retrieve user history by request--memory relevance and inject it into the model context. Yet relevant history may concern the wrong preference aspect, duplicate public information, or provide insufficient support. We argue that personal memory should be used only when it adds utility beyond a public-only response. We propose TRACE-Memory, a two-stage framework for selective personalization. Stage 1 queries for user-specific information missing from the request and public context, then retrieves a coverage-oriented candidate pool. Stage 2 admits a compact subset of source-traceable evidence units, or the empty set, according to response-level incremental utility. We progressively train the query-generation and evidence-admission policies through structured SFT initialization, reduced-space stage-wise GRPO warm-up, and nested multi-sample Joint GRPO. Across 4,500 Controlled and Natural tasks from Goodreads, Amazon Reviews, and Reddit, TRACE-Memory consistently outperforms random and lexical memory use, improves over semantic retrieval, remains competitive with frontier-LLM memory pipelines as local generator capacity increases, and conditions evidence admission on public-context sufficiency, supporting selective rather than default personalization.
What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files
Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Mo…
Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the mode…
LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs
The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design emb…
Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation
Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making t…
Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production
Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously…
HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails
Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur pr…
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models
Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedba…
MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models
Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Tho…
Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs
Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model…
Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding
Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a ques…
Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing
We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across do…
Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization
Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and int…
Deep probabilistic logic programming for diagnostic reasoning from incomplete information: A case study in stroke detection
In medical applications, raw data is frequently associated with significant privacy concerns, lending particular importance to the encoding…
VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference
Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks. Despite…
FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents
Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering task…
SDDBMs: Soft Denoising Diffusion Bridge Models
Diffusion bridge models leverage Doob's \(h\)-transform to construct stochastic transports between arbitrary endpoint distributions, and ha…
Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies fo…
ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration
Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide…
Walking through Discussions: A Mobile Visual Analytics System for In-Situ Group Discussion Analysis
Group discussion-based teaching is widely used to foster collaborative learning, yet teachers in physical classrooms often struggle to simu…
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under…
MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning
In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically…
UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still store…
A QUBO-Inspired Computational Framework for Airport Landside Bottleneck Diagnosis and Dynamic Dispatch Optimization
Airport landside traffic centers connect terminal arrivals with taxis, ride-hailing vehicles, private cars, buses, metro services, parking…
Can Open-Weight Models Compete on Financial Text Comprehension?
Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet thei…
Smart Compaction: Predicting Compaction Utility from Lakehouse Table Metadata
Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhil…
SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests
Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retri…
The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task
How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it rea…
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing…
A Structural Dynamics Graph World Model: Unified Modeling, Constrained Rollout, and Interpretable Calibration
The state evolution of a complex system arises jointly from object laws, relational propagation, domain conservation, and unmodeled error.…
EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility
Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependa…
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face…
AI Evaluation Should Measure Verification Cost, Not Correctness Alone
The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to v…
FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Mod…
Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models
Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response bur…
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when fin…
Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs
Omni-modal LLMs jointly process audio, video, and text, but long multimodal sequences incur substantial prefill and KV-cache costs. Existin…
Improving Generalization Robustness of Multimodal RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simp…
Three Generations of Healthcare IT: From the Digital Record to the Computable Care Process
Objective. Healthcare IT is usually organized by the technologies it adopts. We instead organize it by the unit of information a system mak…
Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models
Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, incons…
PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary
Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, ty…
Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a…
Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration
The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception…
AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups
Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific…
Full-bandwidth transformer
Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attenti…
LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing
Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preferenc…
From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings
Imaging device downtime is a major barrier to healthcare delivery in low- and middle-income countries (LMICs), often driven by limited acce…
Decoding Phenotypes: A Framework for Fusing Genomic Language Models and Neuroimaging
Neuroimaging and genetic testing are two important clinical references for nervous system diseases, offering complementary diagnostic infor…
Integrated Multimodal AI System for Retrieval-Augmented Reasoning, Object Sensing, and Damage Analysis
This work presents a unified multimodal AI system for damage assessment that integrates retrieval-augmented generation (RAG) models, therma…
Not an A11y: How Android Accessibility Exposes Mobile AI Agents to Indirect Prompt Injection
The rise of autonomous AI agents represents a major paradigm shift in how users interact with mobile devices. Frameworks such as MobileRun…
Depth-Aware Implicit Neural Representation Priors for 3D Gravity Inversion
Gravimetry images subsurface density contrasts associated with geological structures, geothermal systems, and intrusive bodies. Recovering…
Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression
Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compr…
CoRe-UIE: Rethinking Coexisting and Region-wise Degradation for Underwater Image Enhancement
Underwater images often suffer from diverse and coexisting degradations, including color distortion, scattering haze, texture attenuation,…
Context Is Not Authority: Structured Runtime Governance for Financial Market Agents
Financial agents can turn correct context into an unauthorized effect: a customer-facing commitment, trade, or deployed policy. We present…
PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs
Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints. Bridging that…
DualCert: A Solver for the Traveling Salesman Problem with Constraint-Coupled Learning
Large traveling salesman problem (TSP) instances require a solver to allocate limited computation while preserving the validity of its outp…
A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition
Mixed emotions represent a clinically relevant but still underexplored target for automatic emotion recognition. EEG provides millisecond-l…
Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways
Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy…
Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models
User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with…
RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement
AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by s…
Motif 3: Technical Report
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per tok…
MELLON - Multimodal Enhanced LLM for Online Navigation
Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are eithe…
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning
Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without…
ChronoState: Hidden Elapsed-Time Conditioning for Temporal-State Action Selection in Frozen-Backbone Language Models
Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration…
TRACE: TRajectory Attribution for Automated Context Engineering
Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contai…
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general p…
From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents
Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has…
Agentic Router: An Execution-Grounded Continual Learning Approach With Memory
Large language model (LLM) agents provide a promising interface for command-line-based network operations, but a plausible command may stil…
Structure-Preserving Uncertainty Propagation in First-Order Proof Search
GK is a query-directed first-order prover that extends ordinary resolution-based proof search with explicit positive and negative claims, n…
Signature-Guided Capacity Occupancy for Dense Expert Merging
Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in…
CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving
Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perceptio…
Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models
Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference compu…
SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge
Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess…
An Explainable GNN Framework for Component-Level Anomaly Diagnosis
Industrial processes are complex systems composed of multiple interacting sensors that generate multivariate time series (MTS). Detecting a…
Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution
Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirel…
SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance
LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provid…
Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline
LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business defini…
Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens. Privileged self-distillation attempts t…
Entropy-based Code Adversarial Translation for Real-world Repository Migration
LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produ…
P$^{3}$: Joint Program-and-Proof Planning for Verified Code Generation
Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the pr…
MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence
Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition,…
ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons
Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shoppin…
CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation
A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter…
Linearized 2-Simplicial Attention
We present a linearized form of 2-simplicial attention by rewriting the trilinear score as an inner product between a composite query and a…
ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management
While mathematical models act as vital decision support systems for operational Air Traffic Flow and Capacity Management (ATFCM), existing…
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs,…
GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models
Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines the value of landslide mapping fo…
Control-Oriented Scenario Tree Construction through Reinforcement Learning
Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario tree, a finite branching approximati…
LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling
Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a blac…
CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits
Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent…
OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks
Long-horizon complex tasks require agents to repeatedly observe states, formulate plans, invoke tools, verify results, and recover from fai…
KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models
KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why. We presen…
Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over…
Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity
Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbid…
From Prompt to Harness: Coderlet from Scratch
A model alone does not determine how a programming agent acts. What the model sees, how actions enter the environment, how feedback returns…
Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents
Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strate…
Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification
Active inference offers a unified framework for perception, learning, and action, but scaling discrete active-inference models to rich spat…
One Adapter Pair per Model: A Universal Activation Interface for Language Models
Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language int…
verdi: retrieval is not transfer for continual world model optimization
Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained…
Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents
External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Y…
The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games
LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As…
ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization
Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisone…
CoRCi: Cross-Reconstruction of Coherent Interests Modeling in Cross-Domain Sequential Recommendation
Cross-Domain Sequential Recommendation (CDSR) aims to alleviate data sparsity by transferring dynamic user interests across related domains…
ICM Out! Better Tournament Strategy from Computed Continuations, vs. Solvers and LLMs
The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against thos…
From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization
Compressing large language models to two bits or fewer is increasingly feasible through block-wise post-training quantization; cross-block…
Adaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree Search
Reliability qualification of advanced semiconductor devices requires sequential stress decisions that balance characterization objectives a…
Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a per…
Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics
Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental…
Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching
GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operat…
Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models
Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands…
Adaptive Semantic Capacity Allocation for Parallel Generative Recommendation
Autoregressive semantic ID recommenders are constrained by expensive beam-search decoding, which limits the practical length of item identi…
Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, c…
Matryoshka Language Model Suites
Training a language model suite classically requires training each model separately and serving them independently. We improve both trainin…
Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature
Muon's polar update is exact for an unweighted spectral geometry. We introduce GO-MUON, which uses a matched data-dependent geometry and re…
AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting
Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because polluta…
CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation
Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these…
Mismatch Matters: On-Policy Distillation Beyond Token Agreement
On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerat…
CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environment…
Agentic Auto-Research is Fuzz Testing
Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the pr…
Towards Expert-level Medical AI for Real-time Video Consultations
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of il…
ArchAgent v2: A Case Study with the Data Prefetching Championship
Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarc…
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memor…
DSLE: A Learning Environment for Dark Souls Boss Encounters
We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remas…
GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis
Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains su…
From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents
Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resultin…
Coordinated incentives in AI-generated misinformation governance
With the rapid diffusion of AI-generated content, AI-driven misinformation is becoming increasingly pervasive and difficult to govern, unde…
Application of Artificial Intelligence for Fraudulent Banking Operations Recognition
This study considers the task of applying artificial intelligence to recognize bank fraud. In recent years, due to the COVID19 pandemic, ba…
Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use
Generative Artificial Intelligence (GenAI) presents a governance challenge for STEM assessment. Unrestricted access can enable task outsour…
Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources
This thesis examines the fairness of Automated Machine Learning (AutoML) tools in human resource hiring systems through the combined lenses…
Cross-Model Humor Preference Modeling with Cards Against Humanity
This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Hum…
Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents
Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players…
How to Ask the AI: A User Perspective Survey for Large Language Model Prompting
AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses…
EmoPatient: An Emotion-Directed Patient Simulator for Realistic Palliative Care Communication Training
Effective communication during palliative care discussions is a critical clinical skill, yet training clinicians to manage complex patient…
Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation
Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for…
Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients
The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based…
Harnessing Abundance: A Generativity Perspective on Human-GenAI Collaboration
Research on human-GenAI collaboration yields conflicting findings: GenAI can enhance creativity yet reduce collective diversity, with uneve…
Innovating with Generative AI: A Human Bottleneck Framework
We propose a human bottleneck perspective for understanding how generative AI transforms the innovation process. The central premise is tha…
JaleesBench: Are AI Assistants Good Spiritual Company?
Large language models are already advisors to millions of people of faith who bring them real decisions. The pressing question for a person…
PIVOT: Preference-based Intervention Vectors for Pedagogical Tutor Steering
LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers. Tutors must choose when…
How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare
Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (L…
EMMR: Emotion-Mediated Multimodal Reasoning for Personality Assessment in Asynchronous Video Interviews
Asynchronous Video Interviews (AVIs) have become increasingly popular for personality assessment. Recent large language models (LLMs) have…
Representation Matters in Longitudinal Affective Computing
Longitudinal, in-the-wild, wearable sensing yields day-level physiology, sleep, activity, and environmental streams, whereas affect and cog…
KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assistants
Mass religious gatherings such as the Kumbh Mela concentrate tens of millions of people into a single region over a few weeks, producing in…
From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations
Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions larg…
Unified Hallucination Fuzzing for Multimodal Large Language Models
Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stak…
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing…
WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management
Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains di…
Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. E…
Ultraconstructive Model Theory via Bounded Adversarial Finite Structures
Ultraconstructive Model Theory (UCMT) replaces idealized satisfaction, at finite compu- tational scale, by bounded adversarial survival. A…
Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger under…
An evolutionary model of animats with VLM-based subjective evaluation
In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language Model (VLM) into the fitness e…
NeuroPilot: An Agent-Driven Smart Pipeline for Processing, Quality Control, and Managing Neuroimages
Transforming raw neuroimage archives into analysis-ready derivatives relies on three brittle stages: data standardization, modality-specifi…
Performance of large language models in the optical diagnosis of colorectal polyps
Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal larg…
MOSAIC: Adversarial Co-evolution of Specialist Heuristics and Problem Instances for LLM-based Automated Heuristic Design
Automated heuristic design (AHD) with large language models (LLMs) has produced strong heuristics for combinatorial optimization problems (…
DarwinX: Evolving Agent Harnesses Through Natural Selection
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement…
Generalizing deep reinforcement learning across cable-driven parallel robot configurations with actuator-level policies
Cable-driven parallel robots (CDPRs) present diverse configurations and complex control challenges, which can be addressed by deep reinforc…
Learning an Interior Layout Policy in a Domain Specific Language Action Space
Indoor scene layout generation is a challenging task in interior design. Existing methods often oversimplify the task by reducing room cond…
P2Voxel: Pyramid Pivot Voxelization for 3D Mesh Tokenization
Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geom…
MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures
Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their…
AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization
Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent mi…
Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth
In digital dentistry and oral surgery, the registration of jawbone CT and intraoral scanner (IOS) data is essential for integrating interna…
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only…
Temporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark
Functional near-infrared spectroscopy (fNIRS) is a promising modality for autism spectrum disorder (ASD) classification, yet existing appro…
Latent-Frequency Validity: Fast Spectral Editing with Screened Video-VAE Transfer Operators
Direct spectral editing in video-VAE latents can control noise, flicker, smoothness, and frequency content without a decode--filter--reenco…
COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping
Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existin…
BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference
Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their mas…
Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects
A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse-…
Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real
The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) set…
Multi-Branch Policy Optimization for Multimodal Large Language Models
Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that a…
Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning
Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose…
Scaling Inherently Interpretable Language Models
Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, wit…
Enhanced Real-Time 6-DOF Extended Reality Catheter Tracking for Evaluating Potential Improvement in Efficiency, Precision, and Depth Perception for Cardiac Interventions
Despite advances in 3D ultrasound, most percutaneous cardiac interventions still rely on 2D visualization, limiting depth perception and sp…
Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays
High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary…
PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue
LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advanc…
Adversarial Attacks on Deep OCR Systems
Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-conte…
SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment
Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wr…
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Model…
CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting
Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short obse…
SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models
A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment…
LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs
Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendat…
Complete, Scalable, and Robust Prioritized Planning for Multi-Robot Ordered Storage and Retrieval at Maximum Capacity
Automated warehouses face a fundamental trade-off between maximizing storage density and achieving high retrieval throughput. While puzzle-…
LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks
Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resource…
Multi-Task Consistency-based Detection of Adversarial Attacks
Deep Neural Networks (DNNs) have found successful deployment in numerous vision perception systems. However, their susceptibility to advers…
CFD-Guided Detection of Concept Drift in Multimodal Physiologic Signals
Cardiovascular AI models can classify clean elec- trocardiogram (ECG) signals, but real wearable signals change because of motion, breathin…
The Anatomy of a Prompt Injection: A Component Model for Structured Analysis
Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than…
Shape Mutating Expert Compression:LorExperts and BTExperts
Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing th…
Distilling CT Foundation Models into Editable Concept Bottlenecks for Lung Nodule Malignancy Prediction
Foundation models provide transferable CT representations, but predictions based directly on these embeddings are difficult to interpret. W…
Vision-Language Grounding as Bidirectional Concept Correspondence
Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localiza…
Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models
Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert. Recen…
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in hi…
Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical Limits
Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate fo…
Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention
We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a ran…
Second Order Drifting Models
Drifting models are a recent class of one-step generative models that evolve the model distribution during training using a predefined samp…
ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB
Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal res…
Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses
Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean. Production warehouses are the inverse: cryptic id…
Persistent Semantic Entities in Tool-Augmented LLM Systems
Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent bo…
EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference
Load Balancing has emerged as a critical problem in expert-parallel distributed inference of Mixture-of-Experts (MoE) models. As routing di…
Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute…
Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design
Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their outp…
Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation
Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for De…
Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models
Time series forecasting (TSF) plays an important role in a wide range of real-world applications. Recently, time series foundation models (…
Evidence-RL: Towards Evidence-intensive Visual Reasoning
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visu…
Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States
Large language models (LLMs) can generate fluent and useful responses but remain prone to hallucinations. We introduce Prompt Embedding Pro…
DA-NBV: A Direction-Aware Next-Best-View Planner for Efficient 3D Reconstruction of Ships at Sea
Accurate 3D reconstruction of ships at sea is important for maritime supervision, damage assessment, and autonomous maritime operations. Al…
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families
Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model…
Tools to Explain Neural Networks for Power System Dynamics
This paper presents, for the first time in power systems literature to our knowledge, analytical tools to explain the training performance…
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scena…
HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection
Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision. Current…
Effect of Abstractions and Prompting Strategies on LLM-Guided High-Performance Optimizations
Code performance optimization is a vital aspect of modern software development, as it enables faster response times and reduced resource us…
Adaptive Symmetry Discovery for Dynamical System Identification
Dynamical systems model trajectory data generated by fixed underlying dynamics, with applications ranging from biology to physics. Especial…
Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection
Retrieval-Augmented Generation (RAG) enables large language models to classify network flows and generate human-readable incident reports b…
NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs
Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence ac…
Hierarchical Multi-Task Federated Learning in VANETs
Vehicular Ad hoc Networks (VANETs) increasingly rely on federated learning (FL) to enable collaborative intelligence without sharing raw se…
Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario
In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short direct…
Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching
Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two couple…
DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology
Attention mechanisms have been widely utilized in modern deep learning, and many existing multi-omics models inherit their conventional use…
Exact Zarankiewicz Values On Two Finite Frontier Slices
The Zarankiewicz number Z(m,n,s,t) is the maximum number of edges in a bipartite graph with parts of orders m and n containing no copy of K…
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytel…
STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs
Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of s…
$\texttt{DisMorph}$: learning to disentangle technical distortions from true biological change
Longitudinal MRI enables sensitive measurement of structural brain change for studying aging and neurodegenerative disease. Deformable imag…
A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization
Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships betwe…
Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry
Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial i…
Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification
Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignmen…
FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Series Classification
Multivariate Time Series Classification (MTSC) demands models that can effectively capture complex temporal patterns across multiple scales…
VTO: Visual Tool Orchestration for Video Anomaly Detection
Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional…
Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations
LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making…
On the Robustness of LLMs' Internal Representation of Code Correctness
Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no sur…
Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows…
Frequency-Domain Dual-Branch Fusion for Medical Visual Question Answering
Medical Visual Question Answering (VQA) requires aligning subtle visual evidence, including lesion texture, boundary sharpness, and diffuse…
Open-World Semantic Segmentation with Sensitivity Modeling
Modern vision systems must operate in "open-world" settings, where models must recognize known categories and detect unseen or anomalous co…
Three Necessary Principles for Self-Supervised Visual Representation Learning
We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objec…
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through revie…
PRISM: A Predictive Protocol for Permutation Optimization via Landscape Diagnostics
Permutation optimization arises whenever the components of a system are fixed but their ordering affects performance. We introduce PRISM, a…
Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books
Audio dramas weave dialogue, sound effects, and music into immersive stories. Creators often adapt books into audio dramas, but this proces…
From Product Search to Preference Articulation: The Economics of Agentic Commerce
Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents.…
Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions
A bidder can quietly buy a stake in a company before making an offer for it. That stake, a toehold, is supposed to pay for itself twice: it…
Abstracted Away: Resisting Alienation and Ungrounded Abstraction in AI Research Communities
Logics of abstraction in computational AI research often push important forms of knowledge and reflection aside: dominant standards of legi…
Human-Guided Causal Knowledge Injection for Virtual Cells
Virtual cells employ machine learning models to simulate and predict cellular behaviors, serving as a critical computational framework for…
FSTC-Encoder: Feature--Spatial--Temporal Correlation Learning for Generalizable RF Sensing
Heterogeneous RF sensing differs substantially in feature structure, spatial layout, and temporal scale, making existing models difficult t…
Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction
Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other…
Hidden Language Consistency Phenomena in Reasoning LLMs
Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the int…
Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization
Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an O…
Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction
Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets…
Halpern Iteration Achieves $\tilde{\mathcal{O}}(\epsilon^{-1/p})$ $p$th-Order Oracle Complexity for Monotone Variational Inequalities
We study second- and higher-order methods for solving smooth monotone variational inequalities (MVI). Monteiro and Svaiter (SIAM J. Optim.,…
SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills
Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their securit…
SuperNeuroMAT: An Efficient Matrix-based Simulator for Spiking Neural Networks
Spiking neural networks (SNNs) offer a promising pathway to energy-efficient AI and brain-inspired computing. However, their widespread ado…
A Combined Feature-Based Framework for Disguise and Spoofing Detection in Face Recognition Systems
Face recognition systems face two distinct, commonly-separated failure modes: spoofing, where an impostor presents a photograph or video of…
On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation
Malaria remains a leading cause of mortality in resource-limited settings, where expert microscopists are scarce. Automated diagnosis based…
CDGC-Net: 3D Medical Image Segmentation with Cooperative Dual-Scale Self-Attention and Grouped Channel Modeling
Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing method…
Population-Scalable Multi-Agent World Modeling
World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent…
Mitigating Gender Bias in English to Romanian Machine Translation
Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like E…
REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering
Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. Howev…
RAG-Based Auto-Configuration for Industrial Fieldbus Devices
Industrial device commissioning requires engineers to manually extract hundreds of protocol-specific parameters from heterogeneous PDF manu…
Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach
Scientific named entity recognition (SciNER) plays a crucial role in information extraction and knowledge discovery from scientific texts.…
CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents
Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools. All TTS systems require fa…
LegoLM: Structured Weight Sharing for Large Language Models
We present \LegoLM{}, a structured weight-sharing compression framework for large language models grounded in a systematic study of why glo…
UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation.…
RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation
Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. E…
Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation
Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses…
LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization
Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency crit…
Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this…
PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation
On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. Yet…
TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases
To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a bench…
Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs
The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially…
Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial Intelligence
In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor m…
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast
On-policy self-distillation improves language-model reasoning by querying a teacher on states actually visited by the student. Recent metho…
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environme…
Hybrid Neural-Classical Correction for Frozen Time Series Foundation Models: A Comprehensive Ablation Study on High-Frequency Stock Prediction
Foundation models for time series forecasting demonstrate impressive zero-shot generalization but often underperform on specialized domains…
Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models
Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice…
Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data
Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be process…
DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference
Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly…
Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol
AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after…
A New Approach to Characterising Optimisation Problems Using Programmatic Representation and Complexity Measures
Characterising optimisation problem instances is a fundamental part of understanding the behaviour and performance of different algorithms…
From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability
How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-langu…
ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision
Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevaili…
Toward CT-Equivalent Image Quality in Low-Dose Radiotherapy Planning: Conditional Diffusion-Based CBCT-to-CT Synthesis and the Impact of CBCT Input Representation
During standard radiotherapy planning, repeated CT acquisitions are often required for patient registration, verification, and adaptive pla…
From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving
Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not pr…
Do AI Forecast Ensembles Sample the Correct Conditional Distribution?
Ensemble forecasting aims to sample the conditional distribution of outcomes; whether AI forecast ensembles do this correctly in a joint se…
Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods
Tree Search-based test-time scaling of LLMs is a powerful tool for automated scientific coding. However, pure Tree Search sometimes struggl…
Fourier Self-Supervision for Fine-Grained Generalized Category Discovery
Generalized Category Discovery aims to recognize known categories while identifying novel ones within unlabeled data. Existing methods, typ…
GALA: Graph-Augmented LLM Agents for Root Cause Analysis and Incident Response in Microservices
Microservice root cause analysis (RCA) requires correlating failures across heterogeneous telemetry within complex service dependency graph…
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetoric…
Detecting Clear Contact Lenses for Iris Recognition: A Two-Stage Mask-Guided Attention Approach
This work focuses on the impact and detection of clear contact lenses in the context of iris recognition. While the detection of cosmetic o…
How Far Do Foundation Models Transfer to Infant Signals? A Cross-Dataset Transfer Audit with a Unified Need Ontology
Public infant cry corpora are small, label-incompatible, and almost always evaluated one corpus at a time. We ask what this practice hides…
Guardian Crawler: Retrieval-First Knowledge Discovery with Bounded LLM Augmentation for Noisy Web Intelligence
Retrieving relevant evidence from noisy web data is challenging, particularly in sensitive domains containing incomplete reports, heterogen…
Multi-agent discovery of practical quantum LDPC codes
Quantum low-density parity-check (qLDPC) codes can encode multiple logical qubits using sparse parity checks, yet searching for useful fini…
DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models
The proliferation of AI-generated images produced by Latent Diffusion Models (LDMs) has raised critical concerns regarding copyright infrin…
SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs
Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign L…
How People Evaluate AI-, Expert-, and Peer-Style Financial Advice
As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand h…
MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural orga…
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on…
Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production
Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such…
Triple Expert Learning from Noisy Labels for Semi-Supervised Vision Foundation Model Adaptation
Semi-supervised adaptation of vision foundation models (VFMs) commonly freezes the pretrained backbone and updates lightweight modules such…
Two-Step MV-DeepONet: Probabilistic Operator Learning for Uncertainty Propagation Driven by Random Input Fields
Forward uncertainty propagation in complex physical systems can induce structured covariance across field-valued outputs. For a probabilist…
A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents
Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or f…
When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information
Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their re…
TLDChoiceNet: Quantitatively Choosing a Transfer Learning Dataset
Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer l…
The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora
How a document's arrangement is written down, its notation, is a training variable that no dataset card records. The field has established…
Visual Distortion Detection in UGC Images Using Large Multimodal Models
The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Exis…
Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measur…
MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning
Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but i…
When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution
Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to t…
SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning
While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by subop…
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes lo…
Tabular Numeric Stretch Transformation
Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distribut…
Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression
Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty t…
Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction
Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Exist…
SiriusDeliver: Automating Data Warehouse Delivery at Tencent
Enterprise data warehouses (DWs) support business-critical analytics, but warehouse task delivery remains a complicated production process…
FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning
Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learnin…
AkasicDB: Demonstrating Omni RAG with a Unified Vector-Graph-Relational DBMS
Recent Retrieval-Augmented Generation (RAG) systems increasingly combine vector retrieval with structured knowledge, such as Graph RAG and…
Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training
Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet…
FedTVD: Balancing Data Quality and Quantity for Robust Federated Learning
Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL f…
Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference
The key-value (KV) cache is the primary throughput optimization in modern large language model (LLM) inference, enabling prefix reuse acros…
RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation
Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these pr…
Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solut…
Multimodal Federated Learning under Dual-Axis Modality Missingness
Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic…
MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing meth…
SafeQL: Search-based Refinement for Safe and Efficient LLM-based Text-to-SQL
Large language models (LLMs) have advanced Text-to-SQL by enabling natural language interfaces to databases without task-specific fine-tuni…
Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations
Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understa…
GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views
Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field…
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation
Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewar…
Software Engineering for and with GUI Agent
GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced t…
GLocFM: A Geometry-Aware Foundation Model for 3D Indoor Wireless Localization
Learning-based wireless localizers often fail to utilize geometric information about the propagation environment, limiting their ability to…
VeinCast: Physics-Guided Dynamic Field Graphs with Graph-Conditioned Fusion for Global Medium-Range Weather Forecasting
Global medium-range weather forecasting requires modeling structured yet state-dependent interactions among heterogeneous atmospheric field…
UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation
Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing sema…
DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation
Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale…
WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation,…
RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction
Brain-to-audio reconstruction is limited by \emph{prior domination}: when a pretrained generator is conditioned on a weak neural signal, it…
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute
Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matte…
Deep Learning based Detection of Fishing Vessels and Fishing Monitoring using Nightlight Images
The demand for maritime surveillance has given rise to the need for monitoring fishing vessel activities, particularly in addressing the ch…
FeedbackTrack: Visual-Cortex-Inspired Cross-Frame Feedback for Transformer Tracking
Visual object tracking requires effective temporal integration, yet most Transformer trackers still rely on predominantly feed-forward feat…
Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation
Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learne…
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law
We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article…
Monotonicity-Guided Bottom-Up Petri Net Discovery: The SPECpp Framework
Process discovery is one of the central challenges in process mining. Petri nets are particularly attractive because simple local construct…
Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models
Research regarding the sign language recognition mostly relies on RGB images, whileas sign language datasets that provide depth images are…
LITEWAY: LIghtweight HAR via Temporal Efficient highWAY
Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learning models on re…
ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models
Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mec…
How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making
In regulated domains such as finance, a model that cannot be explained cannot be deployed, yet many interpretable classifiers defeat their…
WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training
On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state…
Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump Control
On--off cycling is the main cause of compressor wear in residential heat pumps, yet reinforcement learning (RL) controllers for buildings t…
RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation
Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions int…
MixFormer: Linear Transformer with Mixture of Memory Experts
State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transf…
ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents
Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We def…
Beyond Uniform Restoration: Empowering All-in-One Restoration with Pixel-Level Multimodal Guidance
All-in-one image restoration is a unified low-level vision task that aims to effectively recover high-quality images from inputs degraded b…
Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning
Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often…
Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts
Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to gene…
STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework
Incident response planning is critical for restoring compromised software systems after cyberattacks. Common practice relies on expert-driv…
RangeFactory: Scalable Construction of Multi-Hop Cyber Ranges
Real-world cyberattacks often require sustained progress across multiple hosts and network segments, making multi-hop cyber ranges essentia…
Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries
Enterprises increasingly seek to query data lakes using natural language via AI-driven tools like semantic operators or deep research agent…
TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability
We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof…
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs
Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outpu…
ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place dema…
From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation
UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environment…
Distributed Optimization with Streaming Data: A Temporal Weighting Perspective
Optimization theory is a widely used tool for intelligent decision-making. While classical optimization deals with fixed, time-invariant ob…
MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enab…
Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation
Recently, AI-driven video generation has attracted considerable attention. This surge increases the demand for reliable video quality asses…
LEED: Local Embedding Evolution Distance for over-smoothing estimation and virtual node selection in GNN
Graph Neural Networks (GNNs) suffer from two fundamental limitations: over-smoothing, where node representations become indistinguishable w…
TSPORec: Token Selection via Preference Optimization for LLM-Based Sequential Recommendation
Large Language Models (LLMs) have emerged as powerful tools for improving recommendation systems. The effectiveness of LLMs arises from the…
Structure-Enhanced Features and Quality-Aware Dynamic Anchor Scoring for Robust Lane Detection
Lane detection requires recovering thin, elongated, and frequently occluded lane structures under challenging driving conditions. While anc…
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from b…
NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation
Accurate 3D neuron segmentation in fluorescence microscopy is critical for neuroscience. However, the sparse and elongated morphology of ne…
DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation
Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practica…
Predictive safety filter enhanced curriculum learning control for efficient vehicle dynamics controller
Recent advances in learning-based control have enabled impressive achievements in solving complex control problems in various domains. Howe…
Confusion-Geometry Rebalancing for Long-Tailed Adversarial Training
Adversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward…
Evaluating Generative Time-Series Models on Data with Point Masses
Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not r…
How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans
Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as so…
ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners
Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners,…
Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach
Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an e…
SR-OPSD: Self-Referenced On-Policy Self-Distillation
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optim…
Defining Decentralization: An Ontological Perspective
Decentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as…
MoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General Geometries
Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatia…
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the un…
KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs
Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, pa…
Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization
Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this settin…
Parameter Exploration for RLVR via Variational Learning
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also…
MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation
Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medi…
Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation
Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong…
Multi-Agent AI Safety as an Institutional Design Problem
AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources…
Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy
Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely…
Stealing Reasoning Traces from Proprietary LLM APIs
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual prope…
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific dom…
Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning
Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to re…
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference tim…
Fusion Training for Mathematical Generalization in Large Language Models
Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinki…
From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the…
Multimodal Model Diffing for Feature Discovery and Control
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain d…
Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges)…
Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory
The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational socia…
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial…
How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
Recent advancements in large language models (LLMs) suggest that reinforcement learning (RL) effectively internalizes search strategies, yi…
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has show…
ACEvo: Adversarial Co-Evolution of Problem Distributions and Solvers for Combinatorial Optimization
Large language models (LLMs) are increasingly used to synthesize heuristic programs, yet most existing pipelines optimize solvers against f…
Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response
Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic…
Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents
The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serial…
Probabilistic Circuits for Knowledge Graph Completion with Reduced Rule Sets
Rule-based methods for knowledge graph completion provide explainable results, but often require tens of thousands of rules to achieve comp…
From Mimicry to True Intelligence (TI) -- A New Paradigm for Artificial General Intelligence
The debate around Artificial General Intelligence (AGI) remains open due to two fundamentally different goals: replicating human-level perf…
ToolUniverse: An open platform for democratizing AI scientists
AI scientists are emerging computational systems that serve as collaborative partners in discovery. These systems remain difficult to build…
TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation
Current training paradigms, optimized for long-horizon reasoning trace execution, have made Large Language Models (LLMs) excel at pattern m…
The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation
The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of ind…
Intelligence Foundation Model: A New Perspective to Approach Artificial General Intelligence
We propose a new perspective for approaching artificial general intelligence (AGI) through an intelligence foundation model (IFM). Unlike e…
The Belief-Desire-Intention Ontology for modelling mental reality and agency
The Belief-Desire-Intention (BDI) model is a cornerstone for representing rational agency in artificial intelligence and cognitive sciences…
M$^3$Prune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation
Recent advancements in multi-modal retrieval-augmented generation (mRAG), which enhance multi-modal large language models (MLLMs) with exte…
Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE…
Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets
Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets,…
Agentic AI for Clustering, Relationship Discovery, and Semantic Trading in Prediction Markets
Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implic…
Neuronal Attention Circuit (NAC) for Representation Learning
Attention improves representation learning over RNNs, but its discrete nature limits continuous-time (CT) modeling. We introduce Neuronal A…
Multi-Granular Node Pruning for Causal Circuit Discovery
Circuit discovery aims to identify minimal subnetworks that are responsible for specific behaviors in large language models (LLMs). Existin…
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant…
Position: Certifiable State Integrity Should Be Built from Local Validity, Not Global Scale
Breakthroughs in language and vision have motivated increasingly general foundation models for time series and physical dynamics, where evi…
AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts
Large language model agents repeatedly encounter related tasks, yet systems that learn from trajectories commit every lesson to one predefi…
El Agente Gr\'afico: A Semantic Execution Runtime for Scientific Agents
Large language models (LLMs) can plan scientific workflows and generate code, but these capabilities do not specify how scientific state is…
Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search
Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instanc…
A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling
Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but…
MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models
Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When s…
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientifi…
Self-Routing: Parameter-Free Expert Routing from Hidden States
Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a lea…
AIVV: Neuro-Symbolic LLM Agent-Integrated Verification and Validation for Trustworthy Autonomous Systems
Deep learning models excel at detecting anomaly patterns in normal data. However, they do not provide a direct solution for anomaly classif…
A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges
The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Sh…
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static…
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?
Deep research agents increasingly interleave web browsing with multi-step computation, yet existing benchmarks evaluate these capabilities…
FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks
Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for lon…
Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasoning in Murder Mystery Games
Vision-language models (VLMs) have shown impressive capabilities in perceptual tasks, yet they degrade in complex multi-hop reasoning under…
RankGuide: Tensor-Rank-Guided Routing and Steering for Efficient Reasoning
Large reasoning models (LRMs) enhance problem-solving capabilities by generating explicit multi-step chains of thought (CoT) reasoning; how…
Time-Series Forecasting in Safety-Critical Environments: An Open-Source Package for EU-AI-Act-Compliant Development / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein Open-Source-Paket f\"ur die KI-VO-konforme Entwicklung
With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-c…
ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems
ZenBrain is a seven-layer, neuroscience-derived memory architecture for LLM agents that unifies fifteen mechanisms - from Two-Factor synapt…
FitText: Evolving Agent Tool Ecologies via Memetic Retrieval
Efficient reasoning is not only a matter of shortening an answer trace; for tool-using agents, it also depends on whether the agent is reas…
C2L-Net: A Data-Driven Model for State-of-Charge Estimation of Lithium-Ion Batteries During Discharge
Accurate state-of-charge (SOC) estimation is critical for the safe and efficient operation of lithium-ion batteries in battery management s…
AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization
Multi-agent reasoning has shown promise for improving the problem-solving ability of large language models by allowing multiple agents to e…
LLM 推論を改善するための時期尚早な自信を理解し、軽減する
現在の言語モデルの長い思考連鎖 (CoT) には論理的なギャップや不当な飛躍が含まれることが多く、追加のテスト時の計算による利益が制限されます。推論の品質を直接向上させるにはプロセス報酬モデルが必要ですが、それらをトレーニングするために必要なステップレベルのアノテーションは高価で不足しています。私たちは、推論中にモデルの信頼度がどのように変化するかにそのような兆候を発見しました。時期尚早な自信、つまり、答えを早く決めて、残りのトークンを使ってそれを合理化する傾向は、タスクとモデルのスケール全体で推論に欠陥があることを強く予測します。私たちはこれを漸進的信頼形成で利用します。これは、早期にコミットするのではなく推論しながら自信を更新するようにモデルをトレーニングする強化学習目標です。外部のラベルや報酬モデルを使用せず、段階的な自信の成長に報酬を与え、早期のコミットメントにペナルティを与えます。この方法により、算術 (カウントダウン)、数学 (DAPO、AIME)、科学 (ScienceQA) のパラメータ全体で精度と推論の質が 1.5B から 8B に向上しました。Countdown では、精度が 3.2 倍 (+42.0 pp) 向上し、欠陥のある推論は 48 pp 低下しました。 AIME では、Pass@64 により 6.6pp 改善されます。このメカニズムと一致して、この方法は忠実性も向上させます。安全性ベンチマークでは、私たちのモデルは、推論トレース内の誤解を招くコンテンツを隠すのではなく、より透過的に表面化します。対照的な実験では、問題とその解決策が同時にスケールすることが明らかになりました。モデルのサイズとタスクの難易度が上がるにつれて、時期尚早の確信が高まり、それに対処することで得られる利益も同様に大きくなります。
原文 (English)
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute. Improving reasoning quality directly would require process reward models, but the step-level annotations needed to train them are expensive and scarce. We find such a signal in how the model's confidence evolves during reasoning: premature confidence, the tendency to commit to an answer early and use the remaining tokens to rationalize it, strongly predicts flawed reasoning across tasks and model scales. We exploit this in progressive confidence shaping, a reinforcement learning objective that trains models to update their confidence as they reason rather than commit early -- rewarding gradual confidence growth and penalizing early commitment, with no external labels or reward models. The method improves accuracy and reasoning quality from 1.5B to 8B parameters across arithmetic (Countdown), math (DAPO, AIME), and science (ScienceQA): on Countdown, accuracy improves 3.2x (+42.0pp) and flawed reasoning drops 48pp; on AIME, Pass@64 improves 6.6pp. Consistent with this mechanism, the method also improves faithfulness: on a safety benchmark, our models more transparently surface misleading content in their reasoning traces rather than concealing it. Controlled experiments reveal that the problem and its remedy scale together: premature confidence grows with model size and task difficulty, and so do the gains from addressing it.
PortBench: LLM 主導のポートフォリオ管理のための相関を意識したフルパイプライン ベンチマーク
LLM はさまざまな財務タスクにわたって優れたパフォーマンスを示していますが、重要な財務上の意思決定タスクであるポートフォリオ管理 (PM) のベンチマークは依然として不十分です。既存のベンチマークには 2 つの主なギャップがあります。1 つは資産間の相関構造を無視しているため、真に分散されたポートフォリオと集中ポートフォリオを区別できないこと、もう 1 つは現実世界のシナリオで完全な PM 意思決定パイプラインを評価できないことです。 10 年間にわたる 6 つの異種資産クラスにわたるベンチマークである PortBench を紹介します。 PortBench は、2 つの補完的なレイヤーで構成されています。1 つは 7 つのタスク テンプレートにわたる 6,269 の相関ベースの質問からなる静的 QA データセット、もう 1 つは完全な PM 意思決定サイクルを反映する 5 段階の動的な割り当てパイプラインです。これらのレイヤーを評価するために、2 つの専用のメトリックを導入します。提案されたポートフォリオがクラス間ヘッジを活用し、クラス内集中を回避しているかどうかを測定するデュアルレイヤー相関スコアと、パイプライン ステージ全体で推論エラーがどのように複合するかを定量化するメトリックである CEPS です。さらに、3 つの過去のストレス体制とリスク プロファイルの下で、戦略の堅牢性と投資家の連携を評価します。 10 個のフロンティア LLM を評価したところ、静的財務 QA では優れたパフォーマンスを示したにもかかわらず、モデルとプロファイルの組み合わせの 90% が基本的な均等加重割り当てを上回るパフォーマンスを発揮できず、すべての手順上の制約を満たすモデルでもストレスがかかると壊滅的なドローダウンに悩まされることがわかりました。ソース コードは \href{https://github.com/AgenticFinLab/portbench}{this https URL} で入手できます。
原文 (English)
PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlations; they fail to evaluate the complete PM decision pipeline. We introduce PortBench, a benchmark spanning six heterogeneous asset classes over ten years. PortBench comprises two layers: a static QA dataset of 6,269 questions across seven task templates, and a dynamic five-stage allocation pipeline. To evaluate these layers, we introduce two metrics: a dual-layer correlation score for inter-class hedging and intra-class concentration, and CEPS, which quantifies how reasoning errors compound across pipeline stages. We further evaluate under three stress regimes and risk profiles, and support real-time evaluation to mitigate pretraining contamination on historical markets. Evaluating ten frontier LLMs, we find that despite strong financial QA performance, 90\% of model-profile cases fail to outperform equal-weight allocation in 2024, and this deficit persists across other market regimes; models that satisfy every procedural constraint still suffer large drawdowns under stress. Our source code is available at \href{https://github.com/AgenticFinLab/portbench}{this https URL}.
決定論的地平: 拡張推論が失敗し、ツールの委任が必要になったとき
拡張された思考連鎖推論は、決定論的な状態追跡タスクのパフォーマンスを低下させる可能性があります。これは、好みのバイアスによるものではなく、デコーダのみの注意の情報理論的能力に根ざした制限によるものです。 (1) 状態追跡容量を $O(H \cdot \log(L/H) \cdot \sqrt{d_h})$ として制限する、補完的な達成可能性構造を備えたアテンション ボトルネック定理を確立します。 (2) 超指数関数的な精度低下をもたらすコンテキスト依存エラー モデル。 (3) 状態空間 Jaccard メトリックにより、機能がプリファレンスの失敗から区別されます。 (4) ツールの委任が必要になる決定論的範囲 $d^* \in [19, 31]$。 12 のモデルと 8 つのタスク ドメイン (SWE-Bench、WebArena、SQL-Multi を含む) にわたって、ツール統合推論は一貫してニューラル思考連鎖を上回ります。プライマリ モデル スイートでは、精度が 86 ~ 94% に達するのに対し、ニューラル思考連鎖では 24 ~ 42% に達します。最適な長さのトレースを微調整すると $<5% の改善が得られ、アーキテクチャ上の上限が確認され、高いモデル間相関 ($r = 0.81$ ~ $0.91$) は、これらの失敗がトレーニング固有のものではなくアーキテクチャ上のものであることを示しています。私たちの結果は、エージェントシステムにおいて純粋な神経推論がハイブリッドアプローチに屈すべき場合についての原則的な指針を提供します。
原文 (English)
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
Extended chain-of-thought reasoning can degrade performance on deterministic state-tracking tasks, not solely because of preference biases but, on the evidence we present, because of information-theoretic limits in the capacity of decoder-only attention. We present: (1) an Attention Bottleneck analysis providing evidence that total state-tracking capacity in bits is bounded in terms of head count, head dimension, and context length under stated modeling assumptions, and that total capacity is not the binding constraint; (2) a context-dependent error model with a depth-dependent quadratic term in the error exponent; (3) the State-Space Jaccard metric measuring state drift; and (4) a Deterministic Horizon $d^* \in [19, 31]$ (at $\alpha = 0.5$) marking the depth at which unaided accuracy crosses 50%. Across twelve models and eight task domains (including SWE-Bench, WebArena, and SQL-Multi), tool-integrated reasoning reaches 76-94% accuracy versus 17-42% for neural chain-of-thought on PermutationProbe. Fine-tuning on optimal-length traces yields $<$3 percentage-point improvement, supporting an architectural ceiling.
部分情報分解によるマルチモーダル言語モデルにおけるモダリティ相互作用の理解に向けて
マルチモーダル大規模言語モデル (MLLM) におけるモダリティの相互作用を理解することは、信頼性の高い展開の中心となります。私たちは、表現の整合性や結果ベースの評価を超えて、感覚入力と言語入力の固有、冗長、相乗的な寄与を分離する意思決定レベルのフレームワークとして、部分情報分解 (PID) を導入します。 PID は、視覚と言語のベンチマーク全体にわたって、反復的なモダリティ使用プロファイルを明らかにします。推論とグラウンディング指向のタスクは高い相乗効果を示す傾向があるのに対し、専門家と知識指向のタスクは言語固有の依存性が強いことを示します。これらのプロファイルはモデルファミリー全体で一般化され、モダリティレベルの介入に対する感度を予測します。さらに、感覚 PID を使用して PID を三峰性システムに拡張し、言語をビデオとオーディオの情報利得を分解するための制御変数として扱います。感覚 PID をオムニモーダル モデルに適用すると、聴覚と視覚の融合タスクにおいても、視覚情報によって支配される感覚相乗効果のボトルネックが明らかになります。最後に、PID に基づく再重み付けは、マルチモーダル推論とグラウンディングのパフォーマンスを向上させるための最初の証拠を提供します。
原文 (English)
Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
Understanding how multimodal large language models use different modalities is important for reliable reasoning. We employ Partial Information Decomposition (PID) as a decision-level lens and introduce Sensory PID, a conditional formulation that conditions on language and separates unique, redundant, and synergistic contributions from video and audio. Applied to omni-modal models, Sensory PID reveals a sensory synergy bottleneck: even on audio-visual fusion tasks, decisions remain dominated by modality-unique information, with stronger reliance on vision. Modality-shuffling interventions support this asymmetry, while layer-wise analysis reveals a visual-first computation pattern and instruction perturbations show that late-stage sensory fusion is conditioned by language. Beyond diagnosis, PID-guided sample reweighting provides initial evidence that local diagnostic signals can improve multimodal reasoning and grounding performance. As reference validation, our vision-language analysis broadly corroborates previously reported decision-level PID patterns across tasks, models, interventions, and layers.
初期の人間と AI の証明の形式化ワークフローの特徴付け
何世紀にもわたって、人間の数学者は数学的議論を実証するための証明を書いてきました。しかし、証明の有効性を自動的に検証する機能は長い間課題でした。コードを生成し、ますます高度な数学的推論に取り組む AI システムの能力の進歩により、人々の証明を形式化し、それによって証明を検証する能力が変革されることが期待されます。多くの研究は現在のフロンティアのベンチマークに焦点を当てていますが、私たちは代わりに人々がこれらのツールをどのように使用するかを研究しています。私たちは、人々の形式化ワークフローに対する AI の初期影響について、混合手法分析を実施します。つまり、人々が何を望んでいるのか、そのビジョンに対する障壁は何であると見なしているのか、そして実際に AI をどのように使用および適応させているのかなどです。定性的調査によると、人々の好みは多様ですが、証拠発見プロセスに対する人間による高レベルの制御を維持するための形式化における AI 支援を一般的に望んでいます。このような制限の下で、人々が実際に形式化のために AI にどのように取り組んでいるかを評価するために、私たちは、参加者が AI の有無にかかわらず、さまざまな難易度や領域のさまざまな数学問題にわたって非形式的な数学問題とその証明を形式化する、管理されたユーザー研究を実施しました。自動形式化のためのツールの制限にもかかわらず、参加者は、自分で形式化する場合よりも AI ツールへのアクセスを許可された方が、より高い形式化精度を達成する傾向があり、ほとんどの参加者は複数の異なる AI ツールの使用を柔軟に選択します。まとめると、私たちの研究は、人間と AI の関与の密接な相互作用を伴う、形式化ワークフローへの AI 統合の初期段階に光を当てています。
原文 (English)
Human agency in initial human-AI proof formalization workflows
For centuries, human mathematicians have written proofs to substantiate their mathematical arguments; yet, the ability to automatically verify the validity of proofs has long been a challenge. Advances in AI systems' ability to generate code and engage in increasingly high-level mathematical reasoning promise to transform people's ability to formalize and thereby verify proofs. While many works focus on benchmarking the current frontier, we instead study how people use these tools and apply agency in doing so. We conduct a mixed-methods analysis into the initial impact of AI on people's formalization workflows: what people claim they want, what they see as the barriers to those visions, and how they actually use and adapt AI in practice. A qualitative survey reveals that people's preferences are diverse, but with a general desire for AI assistance in formalization that preserves high-level human control and agency over the proof discovery process. To assess how people actually engage with AI for formalization, we conduct a controlled user study in which participants formalize informal math problems and their proofs, with and without AI, across a range of mathematical problems at varying levels of difficulty and domains. Despite limitations of the tools at the time for autoformalization, participants tended to attain higher formalization accuracy when allowed access to AI tools than when formalizing on their own, with most participants flexibly choosing to use multiple different AI tools. Taken together, our work sheds light on the early stages of AI integration into formalization workflows, involving an intimate interplay of human agency and AI engagement.
SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research
Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bou…
リーダー: 抽出された表現による堅牢な証拠に基づく著者の解読
エージェント アプリケーションが公式およびサードパーティの LLM API を介してユーザー タスクをルーティングすることが増えているため、出所が運用上の問題になります。どのモデルが特定のブラックボックス応答を生成したか?私たちは、固定入力セットやベンチマーク スイートではなく、クエリが変化する事前定義されていないプロンプトによって引き出された世代からソース LLM を識別する、動的ブラック ボックス LLM 来歴を研究します。この設定は、プロンプト セマンティクスがテキストの大部分を占めている一方で、モデル固有の作成者追跡が表面レベルでは弱く一貫性がないため、困難です。凍結されたプロキシ LLM を隠された著者証明のリーダーとして扱う軽量の出自フレームワークである READER (Robust Evidence-based Authorship Decoding via Extracted Representations) を紹介します。 READER は、ブラック ボックス出力をプロキシ アクティベーション スペースにマッピングし、各応答内のトークン状態を時間的にフィルタリングし、独立してサンプリングされたプロンプト全体にわたる単一応答の対数事後証拠を合計することによってベイジアン証拠蓄積を実行します。これにより、校正された信頼性に必要なクエリごとの証拠を維持しながら、プロンプト固有の表現の脆弱な平均プーリングが回避されます。エージェント スタイルのプロンプトから構築された 50 のターゲット データセットである Agent500 では、READER は 1 つの応答で $31.0$ ~ $42.4\%$ のトップ 1 の精度に達し、50 の応答で $70.0$ ~ $84.0\%$ に達し、センテンス エンコーダーのフィンガープリントを大幅に上回りました。 9 つのプロキシ リーダーにわたってスケーリングすると、より強力な LLM はより線形にデコード可能な著者情報構造を明らかにすることがさらに示され、凍結された LLM 表現には著者情報の認識がすでに存在し、信頼できるマルチクエリ帰属に変換できることが示唆されます。
原文 (English)
READER: Dynamic LLM Provenance from Query-Varying Interactions
Existing black-box LLM provenance methods achieve comparability by querying every candidate model with the same diagnostic prompts. In deployment, auditors inherit a different evidence stream: heterogeneous prompt-response traces that arrive incrementally. We formalize dynamic black-box LLM provenance: after enrolling a fixed candidate ecosystem, attribute query-varying interactions at any available evidence budget. READER recovers comparability through a frozen proxy LLM. It projects proxy states aligned with response tokens onto length-normalized DC and first-AC modes, capturing response-wide activation location and coarse trajectory evolution. An enrollment-trained linear probe converts each fingerprint into source evidence, and Bayesian accumulation reuses this evidence unit from one observation to many. We introduce Agent500, containing 50,000 responses from 100 local and API sources to 500 heterogeneous agent prompts. On 100-way attribution, READER reaches $50.4\%$ accuracy from one response and $96.2\%$ from 100, compared with $33.0\%$ and $79.0\%$ for the strongest dynamic baselines. Four distinct proxy families all exceed $94.8\%$ at the latter budget. Controlled response-length and Math100 domain shifts expose the limits of zero-retraining transfer. Component analysis reveals a task-dependent spectral division of labor: DC dominates dynamic source identity, while first AC dominates static relationship evidence. Their joint fingerprint provides a shared measurement space for both tasks. READER audits observed text without target internals or audit-only queries. Code and data are available at https://github.com/LeoJeshua/READER.
格子理論による不偏正準集合値オラクル
将来の出来事の確率を推定する非エージェント型の「オラクル」AI は、自己参照の問題に直面しています。その答えが学習され、実行されると、報告するよう求められた確率そのものが変わってしまう可能性があります。 Scientist AI プログラムで提唱されている対応の 1 つは、事実に反する質問のみをし、その回答が何の影響もないかのように評価することです。私たちは、そのような答えは学んだ瞬間に意味がなくなってしまう傾向があることを観察しています。それは、まさにその前提が間違っているからです。したがって、私たちは、オラクルが単一の確率ではなく、同時に偏りがなく学習の結果と自己矛盾のない一連の資格を報告する、自己言及的な代替案を模索します。単純な自己一貫性の要件は、あまりにも多くのセット (役に立たない答え $[0,1]$ を含む) によって満たされるため、問題は正規の自明でないメンバーを選び出すことです。これを、適切に定義されたアイソトーン演算子の最小不動点を取り、閉じたクレダル集合の完全な格子に関するクナスター-タルスキーの不動点定理を使用して行います。代わりに、バリアントは、すべての自己矛盾のない点推定を含む最小不動点を報告します。我々は、存在、自己無矛盾性、空でないことを証明し、非実行的質問についてはその構造が古典的な点の答えに崩壊すること、そしてバイナリイベントについては標準的な答えが自然なハル因数分解の仮定の下では区間であることを示します。この展開は純粋に格子理論に基づいており、バイナリ イベント $B$ から任意の確率変数 $X$ までそのまま拡張され、$P(B\mid A,C)$ は条件法 $\mathcal{L}(X\mid A,C)$ に置き換えられます。区間の特徴付け自体がその一般化に耐えられるかどうかを含む未解決の質問で終わります。
原文 (English)
Unbiased Canonical Set-Valued Oracles Via Lattice Theory
An oracle that tells you the probability of some future event can change that very probability because you act on the answer. We argue that this performativity is OK as people consult oracles to be informed, and hence moved, by the answer. We worry about instead that asking for a self-consistent answer, one that still holds once it has been announced, may leave the oracle with several answers to pick from, and whichever rule it uses to pick is a lever it could learn to pull. We propose to take away that choice: The oracle reports a credal set instead of a point estimate, which lifts the oracle's reaction function to an isotone operator on a lattice. We make the oracle report that operator's least fixed point, which exists because of Knaster and Tarski. As that answer is fixed by a rule laid down in advance, nothing is left to choose by the oracle. We show that solution exists, is self-consistent, is never empty, can be computed by iterating from below, and equals the ordinary point estimate if the question is not performative after all. For simple queries about probabilities, we propose to restrict answers to intervals and show that, under a mild monotonicity assumption, the answer is simply the interval from the no-information baseline to the self-fulfilling equilibrium one would end up at by iteratively querying a point oracle until the answer is self-consistent. As our proposal is purely order-theoretic, it carries over unchanged to arbitrary random variables and to bounded continuous statistics of their law. Finally we show that a fixed finite family of polytopes suffices to approximate every answer uniformly, which turns the construction into a terminating computation. We close by placing the construction inside the Scientist AI programme, where it offers a choice-free criterion for a step that programme currently hands to audited human judgement.
人間は関与をやめ、推論モデルは存続: 難易度の登録と審議の割り当てを分離する
大規模推論モデル (LRM) は、人間と同じように、より困難な問題に時間がかかります。この表面の類似性は、アイテム内に反対のパターンを隠します。 LRM が問題を間違えると、同じ問題を正解した場合よりも多くのトークンを消費します。人間はその逆を行い、間違った試験に費やす時間を減らします。検討を 2 つのレベルに分けます。1 つは応答時間が項目全体の難易度をどのように追跡するか (登録)、もう 1 つは項目の ID が固定された状態で、エージェントが自身の失敗と成功のどちらに多くの時間を費やすか (割り当て) です。公開されているヒトと LRM の照合コーパスでは、人間と 5 つの思考 LRM はすべて、既知の項目間アライメント (登録) を再現しますが、項目 (割り当て) 内では分岐します。どの LRM も大きな誤対正効果 (H-ARC におけるコーエンの d = 1.47-3.13) を示しますが、人間は反対の符号を示します。比較は各エージェント独自のスケール内に留まります。秒とトークンを 1 つの軸に置くことはありません。解離はアイテムの固定効果の下で保持され、データセット全体で複製され、非思考ベースラインには存在しません。私たちは人間のパターンを、関与対放棄として読みます。人々は、解決できると期待している項目に留まり、残りの項目は放棄します。 LRM パターンを不確実性によって引き起こされる長さとして読み取ります。モデルが不確実な場合、チェーンは成長します。つまり、モデルが失敗する傾向があるのはまさにこの時です。どちらのポリシーも同じ項目間の相関関係を生成するのは困難ですが、以前の研究で使用された尺度に基づいて一致しているように見えます。相違は、アイテムの同一性が固定された場合にのみ現れます。リソース合理的メタ推論では、困難信号は共有するが反対の制御を実装する 2 つの停止ポリシーの間で分割が行われます。トレース長が信号を捕捉し、制御を逃します。
原文 (English)
Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation
Large reasoning models (LRMs) spend more reasoning tokens on problems that take humans longer, suggesting sensitivity to a similar structure of difficulty. That alignment identifies which problems elicit more deliberation and leaves open how a difficult state is converted into continued computation. We distinguish difficulty registration, the relation between problem difficulty and observable deliberation, from deliberation allocation, the relation between registered difficulty and continued work. We test this distinction in matched human and LRM data from visual abstraction, intuitive physics, and relational reasoning. In visual abstraction, LRM trace length tracks the human ordering of problem difficulty. After problem identity is controlled, successful human attempts are longer than failed human attempts. Failed LRM attempts receive more reasoning tokens than successful model attempts. The same allocation difference appears in intuitive physics. In relational reasoning, the opposite-sign pattern disappears, while the item-controlled human-LRM difference remains, showing that task structure changes the observable form of allocation. Human action counts link longer human trials to continued task engagement, while failed LRM traces show task-dependent signs of hesitation or revisiting after trace length is controlled. A resource-rational account relates these patterns to the expected value of further computation from a difficult state. Humans and LRMs can agree about which problems are difficult while applying different policies to continued deliberation.
NormAct: 身体化された計画における隠れた社会規範遵守のベンチマーク
マルチモーダル大規模言語モデル (MLLM) は、自己中心的な環境で具体化されたプランナーとして導入されることが増えています。タスクを成功させるには、指示された目標を達成するだけでなく、社会的に適切な方法で行動する必要があります。明示的な目標によって特定の行動が最適化される場合もありますが、暗黙の社会規範によって隠れた制約が課せられることがよくあります。既存の評価は通常、明示的な目標達成または直接的な規範知識に焦点を当てており、計画者がアクション シーケンス内のこれらの隠れた制約を推測して適用できるかどうかを評価することはほとんどありません。目標達成、規範遵守、全体的なタスクの成功に関する計画を評価する、具体化された社会規範の相互作用のベンチマークである NormAct を紹介します。 NormAct は通常のタスク内に隠れた規範を独自に埋め込み、明示的な指示なしにモデルがそれらを実現できるかどうかをテストします。最先端の MLLM (GPT-5.4、Claude Opus 4.7、Gemini 3 Pro) を用いた実験では、大きなギャップが明らかになりました。モデルは 67.3\% のケースで明示的な目標を達成しましたが、隠れた基準に準拠したのは 26.4\% のみでした。合図条件の実験によると、このギャップは一般的な社会知識の欠如ではなく、文脈の中で関連する規範を活性化し、基礎づける際の課題から生じていることが示されています。これに対処するために、計画前にシーン関連の規範を推測するコンテキスト条件付きキュー ジェネレーターである NormPerceptor を提案し、タスクの成功率を 24.2\% から 46.7\% に向上させます。私たちの結果は、身体化されたエージェントが隠れた規範を積極的に検出し、視覚的な証拠に基づいて行動計画の制約として統合できるようにすることの重要性を強調しています。私たちのベンチマークは https://huggingface.co/datasets/Caleb196x/NormAct で公開されています。
原文 (English)
NormAct: Benchmarking Embodied Agents' Proactive Compliance with Unspoken Social Norms
Embodied agents driven by multimodal large language models (MLLMs) can often complete everyday tasks from visual observations, but goal achievement does not establish whether they proactively respect unstated social norms. Existing benchmarks assess explicit norm judgments or constrained behavior, but rarely test whether agents infer and apply scene-relevant norms during ordinary tasks. We introduce NormAct, a benchmark of 550 TongSim scenarios in which the same goal permits norm-compliant or norm-violating action sequences. Norm-relevant evidence is embedded in each scenario while the applicable rule is omitted from the goal instruction. By progressively increasing normative guidance while holding the goal and scene fixed, NormAct tests whether compliant behavior emerges autonomously or only after prompting. Across three MLLM planners, goal achievement substantially exceeds norm compliance without guidance (67.4% versus 24.7%), while both broad and rule-specific guidance improve compliance, indicating that planners can often comply when prompted but not reliably on their own. With a fixed planner, general-norm retrieval is less effective than norm-relevant scene descriptions or generated norm cues, suggesting that identifying relevant visual evidence is a greater challenge than accessing general norm knowledge. NormAct therefore supports the development of embodied agents that pursue everyday goals while proactively respecting unstated social norms.
ContextSniper: リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ
大規模な言語モデル エージェントは実際のリポジトリの問題を修復できますが、ファイル全体の読み取り、広範な検索、および有用な証拠が無関係なコードやログと混在する長いターミナル出力に多額のコンテキスト バジェットを費やすことがよくあります。このペーパーでは、リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ層である ContextSniper について説明します。 AntTrail の広範なエージェント メモリ エンジンのコーディングに特化したものとして、ContextSniper は正確な証拠選択のための Sniper 機能を実装しています。候補コードとランタイム証拠を取得し、ハイブリッド取得信号でランク付けし、意図を認識したコンテキスト ゲートを通じて長い出力をフィルタリングし、プロンプトの外で回復可能なソース コンテキストを保持しながらコンパクトな証拠パケットを返します。 OpenClaw と Claude Code を備えた SWE-bench Lite 上で、ホスト エージェント条件ごとに 50 タスクの実行を使用して ContextSniper を評価しました。 ContextSniper は、OpenClaw の場合、トークンの総使用量を 51.5%、ログに記録されたコストを 36.4% 削減し、Claude Code の場合、トークンの総使用量を 38.9%、推定コストを 27.3% 削減します。提出された解決率は、OpenClaw の場合は 26.0% から 24.0% に、Claude Code の場合は 32.0% から 30.0% にわずかに減少しました。 ContextSniper のパイロット テスト スクリプトは、https://github.com/Calluking/ContextSniper でオープンソース化されています。
原文 (English)
ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair
Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long terminal outputs where useful evidence is mixed with irrelevant code and logs. This paper presents ContextSniper, AntTrail's code-repair module for precision evidence selection in repository-level program repair, part of AntTrail's broader agent-memory engine. AntTrail is available at https://gitcode.com/datagallery/AntTrail. ContextSniper indexes code and action memory as three abstract levels, retrieves candidates with a hybrid ranker, filters long tool output through an intention-aware context gate, and returns compact evidence packets while keeping full source recoverable on demand. In a matched 50-task-per-condition comparison on SWE-bench Lite (same tasks, baseline vs.\ ContextSniper), ContextSniper reduces total token use by 51.5% and logged cost by 36.4% for OpenClaw, and by 38.9% and 27.3% for Claude Code, with submitted-resolution rates essentially unchanged in both host-agent settings. In a separate five-task comparison, ContextSniper beats existing memory- and RAG-style integrations on token efficiency. These results suggest ContextSniper can substantially cut token and cost overhead for repository-level repair agents without a measurable loss in repair quality. The evaluation harness for this study is available at https://gitcode.com/lukchiwang/ContextSniper.
LLM-as-a-Tutor: 検証不可能な RL に対するポリシーに応じたプロンプト適応
検証不可能な指示に従う強化学習 (RL) は、報酬シグナルとしてプロンプト固有のルーブリックを使用する LLM 審査員にますます依存しています。最近の手法では、トレーニング中にこれらのルーブリックを進化するポリシーに適応させていますが、トレーニング プロンプト自体は固定されたコーパスから抽出された静的なものです。この静的なアプローチでは、プロンプトの難易度と政策能力の間に重大な不整合が生じることが多く、プロンプトがロールアウト間の品質の差異を引き出すことができない場合、裁判官は差別的な報酬シグナルを回復できなくなります。この不整合に対処するために、LLM の役割を裁判官から家庭教師に拡張するフレームワークである LLM-as-a-Tutor を導入します。単一のモデルは、ポリシーのロールアウトをペアごとに比較して問題のないプロンプトを検出する検査者として、また、プロンプトにアトミックな制約を追加するジェネレーターとして機能します。この追加専用の設計は、ポリシーの機能に合わせて難易度を単調に上昇させ、外部の難易度スケジュールなしで自己調整トレーニング信号を生成します。 3 つの複雑な指示に従うベンチマークにおいて、私たちの手法は、ポリシーを意識しないベースラインと、ルーブリックを適応させたりプロンプトを書き換えたりする従来のポリシー適応型手法の両方を一貫して上回っており、検証不可能な RL ではポリシー認識の軸が欠けているとして、迅速な適応が示唆されています。
原文 (English)
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL
Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt these rubrics to the evolving policy during training, the training prompts themselves remain static, drawn from fixed corpora. This static approach often results in a critical misalignment between prompt difficulty and policy capability, leaving the judge unable to recover a discriminative reward signal when prompts fail to elicit quality variance among rollouts. To address this misalignment, we introduce LLM-as-a-Tutor, a framework that extends the LLM's role from judge to tutor: a single model serves as an examiner that pairwise-compares policy rollouts to detect non-challenging prompts, and as a generator that appends atomic constraints to them. This append-only design monotonically raises difficulty in step with the policy's capability, producing a self-calibrating training signal without external difficulty schedules. On three complex instruction-following benchmarks, our method consistently outperforms both policy-unaware baselines and prior policy-adaptive methods that adapt rubrics or rewrite prompts, suggesting prompt adaptation as a missing axis of policy-awareness in non-verifiable RL.
TOFFEE のデモンストレーション: データ エージェントの軌跡を大規模に合成するための学習済みシステム
LLM を活用したデータ エージェントは、データ主導の意思決定においてますます重要な役割を果たしています。しかし、既存のデータ エージェントは、特に異種混合の企業設定において、目に見えないデータ環境や分析ワークフローを一般化するのに苦労しています。このため、特定のデータ環境の複雑な分析ワークフローをキャプチャする高品質のデータ エージェントの軌跡を合成する必要性が高まっています。このような軌跡は、2 つの重要な下流用途をサポートします。1 つはデータ エージェント モデルをターゲット ドメインに適応させる教師あり微調整 (SFT) データとして、もう 1 つは不慣れなデータ環境で汎用 LLM をガイドするためのインコンテキスト学習 (ICL) デモンストレーションとして機能することです。そこで、適応モデル選択とクロスタスクプレフィックス再利用を備えたモンテカルロツリー検索(MCTS)を介して、特定のデータ環境から高品質のデータエージェントの軌跡を合成するシステムであるTOFFEEを紹介します。 TOFFEE が異種環境にわたる複雑な分析タスクのスケーラブルな軌跡データを効果的に生成できることを示します。このデモでは、TOFFEE のタスク プール構築、トラジェクトリ エクスプローラー、学習コスト モデルなどのシステム フレームワークを紹介します。また、TOFFEE の Web インターフェイスとそのワークフローを紹介し、データ エージェント微調整のための軌跡合成と、デモンストレーションによって拡張されたデータ エージェント推論という 2 つのエンドツーエンド シナリオを示します。
原文 (English)
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise settings. This creates a growing need for synthesizing high-quality data agent trajectories that capture complex analytical workflows for given data environments. Such trajectories support two key downstream uses: they can serve as supervised finetuning (SFT) data that adapts data agent models to the target domain, and as in-context learning (ICL) demonstrations to guide general-purpose LLMs in unfamiliar data environments. Thus, we introduce TOFFEE, a system for synthesizing high-quality data agent trajectories from given data environments via Monte Carlo Tree Search (MCTS) with adaptive model selection and cross-task prefix reuse. We show that TOFFEE can effectively generate scalable trajectory data for complex analytical tasks across heterogeneous environments. In this demonstration, we present the system framework of TOFFEE, including its task pool construction, trajectory explorer, and learned cost model. We also introduce the web interface of TOFFEE and its workflow, and demonstrate two end-to-end scenarios: trajectory synthesis for data agent finetuning, and demonstration-augmented data agent reasoning.
エクスペリエンス メモリ グラフ: エージェント向けのワンショット エラー修正
大規模言語モデル (LLM) エージェントは、状態、アクション、および観察の一連の軌跡を生成することにより、自律的な意思決定において優れた能力を示しています。ただし、複雑で長期にわたるタスクでは、これらのエージェントは複合エラーに悩まされ、失敗から回復するのに苦労することがよくあります。既存の自己修正メカニズムはプロンプトベースのリフレクションに依存していますが、これは本質的に脆弱で、試行錯誤の繰り返しループにより多大な時間と API コストが発生し、新しいシナリオに一般化するのが難しいタスク固有のメモリが生成されます。これに対処するために、エージェントの障害回復をグラフ マッチング問題として再定式化するフレームワークであるエクスペリエンス メモリ グラフ (EMG) を提案します。トレーニング時に、失敗した探索の軌跡と成功した専門家の軌跡の両方を、指示されたアクションの決定グラフに変換します。これらのグラフを照合することで、共通のサブグラフ (成功したワークフロー) と、失敗を修正する方法 (たとえば、特定の観察の下でどのアクションを追加、削除、または再ラベル付けするかなど) を明示的に示すグラフ編集パスを抽出し、それらをタスク内ノードとタスク間エッジを含むメモリ グラフに保存します。テスト時に、EMG は関連する洞察を取得し、ループのない単一の実行でエージェントをガイドします。 ALFWorld と ScienceWorld での実験では、EMG が成功率と平均報酬において最先端のリフレクション ベースラインを常に上回り、テスト時の試行錯誤を必要としないことが示されています。
原文 (English)
Experience Memory Graph: One-Shot Error Correction for Agents
Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of states, actions, and observations. However, in complex, long-horizon tasks, these agents frequently suffer from compounding errors and struggle to recover from failures. Existing self-correction mechanisms rely on prompt-based reflection, which is inherently brittle, incurs heavy time and API costs due to iterative trial-and-error loops, and produces task-specific memory that may be hard to generalize to new scenarios. To address this, we propose Experience Memory Graph (EMG), a framework that reformulates agent failure recovery as a graph matching problem. At training time, we convert both failed exploration trajectories and successful expert trajectories into directed action decision graphs. By matching these graphs, we extract common subgraphs (successful workflows) and graph edit paths that explicitly indicate how to correct failures (e.g., which actions to add, delete, or relabel under a given observation), and store them in a memory graph with intra-task nodes and cross-task edges. At test time, EMG retrieves relevant insights and guides the agent in a single, loop-free execution. Experiments on ALFWorld and ScienceWorld show that EMG consistently outperforms state-of-the-art reflection baselines in success rate and average reward, while requiring no test-time trial-and-error.
Atari Pong のワールド モデルの概念に基づく空間正則化
ワールド モデルは通常、モデルベース強化学習 (MBRL) システムのコンポーネントとして評価されますが、ワールド モデル自体が単独で研究されることはほとんどありません。 Atari Pong の 5 つの代表的なビジュアル ワールド モデル エージェント (DreamerV3、DIAMOND、TWISTER、Simulus、STORM) を調べます。トレーニング パイプラインを再現し、報告されたエージェントのパフォーマンスと一致させた後、学習したワールド モデルをフリーズし、閉ループ ロールアウト診断で評価します。対応する MBRL エージェントとは別にトレーニングされたポリシーが各フリーズ モデルと相互作用し、生成されたビデオ軌跡の視覚的および動的エラーが検査されます。 5 つのモデルすべてにわたって、ロールアウトには、ボールの消失、不正なボールの動き、無効なボールとパドルの相互作用など、明らかな障害が含まれています。視覚的な軌跡を超えて、ピクセル空間ゼロショット MBRL を使用してそれらをさらに評価します。そこでは、新しいポリシーがフリーズ ワールド モデル内で完全にトレーニングされ、実際の環境で評価されます。 5 つのモデルすべてにおいて、結果として得られるポリシーは、対応する元の MBRL トレーニング パイプラインによって生成されたポリシーを大幅に下回っています。この差は特に DreamerV3 で大きく、平均リターンは -5.5 から -20.9 に低下し、Pong リターンの最小値である -21 に近くなります。私たちは、ポンのボールなど、タスクに不可欠な概念のモデリングが不十分であることが、これらの失敗の一因となっている可能性があると仮説を立てています。したがって、我々は、セグメント化された概念領域に適用される補助ピクセル再構成損失である概念誘導空間正則化 (CGSReg) を提案します。実験では、CGSReg が DreamerV3、DIAMOND、TWISTER の閉ループ ロールアウトとピクセル空間ゼロショット MBRL の両方を改善することが示されています。その効果は残りのモデルと評価指標によって異なり、CGSReg だけではワールド モデルのすべてのボトルネックに対処できるわけではないことを示しています。
原文 (English)
Concept-Guided Spatial Regularization for World Models in Atari Pong
World models are usually evaluated as components of model-based reinforcement learning (MBRL) systems, leaving their standalone reliability understudied. We reproduce five visual world-model agents in Atari Pong -- DreamerV3, DIAMOND, TWISTER, Simulus, and STORM -- and match their reported agent performance. We then freeze the learned world models and evaluate them in two ways. In a closed-loop rollout diagnostic, a policy trained separately from the corresponding MBRL agent interacts with each frozen model, and we inspect the generated visual trajectories for visual and dynamical errors. Across all five models, these rollouts contain clear failures, including ball disappearance, incorrect motion, and invalid ball-paddle interactions. Beyond visual trajectories, we further evaluate the frozen models with pixel-space zero-shot MBRL, a challenging setting in which a new policy is trained entirely inside each frozen world model and then evaluated in the real environment. Across all five models, these policies substantially underperform those produced by the corresponding original MBRL pipelines. For DreamerV3, mean return drops from $-5.5$ to $-20.9$, near the minimum of $-21$. We hypothesize that insufficient modeling of task-critical concepts, such as the ball in Pong, contributes to these failures and propose Concept-Guided Spatial Regularization (CGSReg), an auxiliary reconstruction loss on segmented concept regions. CGSReg improves both closed-loop rollouts and pixel-space zero-shot MBRL in DreamerV3, DIAMOND, and TWISTER, and improves zero-shot MBRL in Simulus; STORM shows no clear improvement.
正確なネットワーク手術: リアクティブな計算グラフにおける関数の不変性と勾配可塑性
Net2Net やプログレッシブ スタッキングなどの機能を保持したネットワーク成長手法は、学習した関数を破壊することなくモデルの容量を拡張しますが、既存の定式化では数値の摂動を許容するか、トレーニング プログラムの完全な再構築が必要になります。我々は、Exact Network Surgery を形式化します。これは、(i) 明示的な浮動小数点仮説の下でビット正確にネットワーク関数が保持され、(ii) 挿入されたパラメータが挿入直後にトレーニング可能なままであるように、ライブ計算グラフに残差ブロックをインプレース挿入することです。ゲート残差ブロックの恒等射定理、リアクティブ無効化エンジンが他のすべてのノードの値とオプティマイザーの状態をそのままにして、挿入ポイントの下流の円錐を正確に再計算することを示す構造局所性定理、およびランダムに初期化された分岐上でゼロに初期化された勾配シャドウイング ゲート アルファが挿入時に一般的に非ゼロの勾配を受け取ることを示す初期化からの脱出命題を証明します。縮退構成 (ゼロ初期化された出力投影とゼロ ゲートの組み合わせ) を特定します。これは、正確な鞍点勾配降下法から逃れることができません。すべての主張は、Julia のリアクティブ グラフ エンジンである NeuroDSL のリファレンス実装で検証されます。グラフトはテストされたすべてのロジットでビット正確です (1600 件中 0 件の不一致)。ゲートは最初のオプティマイザー ステップでゼロをエスケープし、2 番目のステップで分岐勾配のロックを解除します。これは予測どおりです。縮退構成は、600 ステップの実行全体にわたってまったく同じゼロの勾配を示します。手術コストは r = 0.9992 で下流の錐体サイズを追跡しますが、グラフトプラス無効化のブックキーピングは挿入深度全体にわたって一定 (約 0.75 ミリ秒) です。そしてトレーニングは、実際のプロセスの再起動後もほぼ同じように再開されます。フラグ付きの予備付録では、挿入後のゲート ダイナミクスに関する最初の単一シードの観察結果が報告されています。
原文 (English)
Exact Network Surgery: Functional Invariance and Gradient Plasticity in Reactive Computational Graphs
Function-preserving network growth techniques such as Net2Net and progressive stacking expand a model's capacity without destroying its learned function, but existing formulations either tolerate numerical perturbations or require a full rebuild of the training program. We formalize Exact Network Surgery: the in-place insertion of a residual block into a live computational graph such that (i) the network function is preserved -- bit-exactly under explicit floating-point hypotheses -- and (ii) inserted parameters remain trainable immediately after insertion. We prove an identity-morphism theorem for gated residual blocks, a structural-locality theorem showing that a reactive invalidation engine recomputes exactly the downstream cone of the insertion point, leaving every other node's value and optimizer state untouched, and an escape-from-initialization proposition showing that the Gradient Shadowing gate alpha, initialized at zero over a randomly initialized branch, receives a generically non-zero gradient at insertion time. We identify a degenerate configuration -- zero-initialized output projections combined with a zero gate -- that is an exact saddle point gradient descent cannot escape. Every claim is validated on the reference implementation in NeuroDSL, a reactive graph engine in Julia: grafting is bit-exact on every logit tested (0 mismatches out of 1600); the gate escapes zero at the first optimizer step and unlocks branch gradients at the second, exactly as predicted; the degenerate configuration exhibits gradients identically zero for the entire 600-step run; surgery cost tracks downstream cone size with r = 0.9992 while graft-plus-invalidation bookkeeping is constant (about 0.75 ms) across insertion depths; and training resumes bit-identically across a real process restart. A flagged preliminary appendix reports first single-seed observations on post-insertion gate dynamics.
ProbSPARQL: 多次元の不確実な数値データを使用したナレッジ グラフのクエリ
SFB 1574 Circular Factory は、返品された製品に関するデータを統合するための共有ナレッジ グラフ インフラストラクチャを構築しています。中心的な課題は、サーキュラーファクトリーのデータには、(i) センサーに由来する、またはセンサーベースの測定から導出される数値測定値が含まれており、(ii) 多くの場合多次元であり、(iii) 本質的に不確実である一方、下流のトリアージ、検証、信頼性モデリング、および再組み立て計画モジュールにはクエリ可能な不確実性表現が必要であることです。現在の RDF および SPARQL テクノロジには、このような不確実な数値測定データの調和されたクエリと分析に対するネイティブ サポートが不足しています。このギャップに対処するために、このインフラストラクチャの初期段階のクエリ層パイロットとして開発された上位互換性のある SPARQL 拡張機能である ProbSPARQL を紹介します。 ProbSPARQL は、不確実な数値を確率変数としてモデル化し、その分布が確率的 RDF リテラル データ型によってエンコードされ、分布を意識した式、確率的フィルター、発散ベースの結合をサポートします。 ProbSPARQL を Apache Jena ARQ に実装し、Fuseki 互換の実行レイヤーを通じて公開します。 GMM でエンコードされた不確実性とヒストグラムベースの経験的粗さ分布をカバーするプロジェクト由来の測定フラグメントを使用して実際のデータの適用性を評価し、最大 5,000 個のアングル グラインダー インスタンスと 150 万のトリプルを備えた制御されたオントロジー準拠のベンチマークでスケーラビリティを個別に評価します。結果は、実現可能なエンジン内実行、アプリケーション層の後処理よりもフィルター プッシュダウンの高速化、分岐結合決定戦略間のレイテンシー精度のトレードオフを示しています。
原文 (English)
ProbSPARQL: Querying Knowledge Graphs with Multi-dimensional, Uncertain Numeric Data
The SFB 1574 Circular Factory is building a shared knowledge graph infrastructure for integrating data about returned products. A central challenge is that circular-factory data include numeric measurements that (i) originate from sensors or are derived from sensor-based measurements, (ii) are frequently multi-dimensional, and (iii) are inherently uncertain, while downstream triage, validation, reliability-modeling, and reassembly-planning modules require queryable uncertainty representations. Current RDF and SPARQL technologies lack native support for harmonized querying and analysis of such uncertain numeric measurement data. To address this gap, we present ProbSPARQL, an upward-compatible SPARQL extension developed as an early-stage query-layer pilot for this infrastructure. ProbSPARQL models uncertain numeric values as random variables whose distributions are encoded by probabilistic RDF literal datatypes, and supports distribution-aware expressions, probabilistic filters, and divergence-based joins. We implement ProbSPARQL on Apache Jena ARQ and expose it through a Fuseki-compatible execution layer. We assess real-data applicability using project-derived measurement fragments covering GMM-encoded uncertainty and histogram-based empirical roughness distributions, and evaluate scalability separately on controlled ontology-conformant benchmarks with up to 5,000 angle-grinder instances and 1.5M triples. The results show feasible in-engine execution, filter-pushdown speedups over application-layer post-processing, and latency-accuracy trade-offs among divergence-join decision strategies.
SoftReason: 高次元の知覚データに対する完全微分可能なニューロ ソフト シンボリック演繹推論アーキテクチャ
多くの推論問題では、前提は個別のシンボルとして観察されず、高次元の入力から推論する必要があります。さらに、述語の語彙、引数の構造、および信頼できる証拠は、ナレッジ グラフ (KG)、またはルール定義によって提供されます。古典的な神経記号パイプラインには、知覚と演繹の間に個別のインターフェイスがあります。我々は、潜在的な知覚事実と知識が提供する述語に対する微分可能な演繹的推論のためのニューロ・ソフト・シンボリック・アーキテクチャを提示します。 SoftReason は、候補定数と述語に対するローカル ソフト解釈テンソルとして演繹状態を表すことにより、勾配ギャップを除去します。認識は確率的な基本事実を提案し、KG トリプルは信頼性の高いソフト証拠として入力され、すべてのクエリアンカー、述語の選択、およびクロージャーの更新は微分可能のままです。私たちの核となるイノベーションは、直接結果演算子の学習された微分可能リフトです。述語定義の埋め込みと潜在的な構成チャネルを使用して、ソフトボディと述語の混合を形成し、すべての可能な証人を集約し、クエリ条件付きヘッドファクトを提案し、単調な確率的 OR を通じて解釈を更新します。 Knowledge-aware Visual Question Answering (KVQA) のフレームワークをインスタンス化し、SoftReason が 1 つのトレーニング可能なアーキテクチャでエンドツーエンドの知覚グラウンディング、KG 証拠の注入、および微分可能な演繹的閉包をどのようにサポートするかを示します。
原文 (English)
SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data
In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further, the predicate vocabulary, argument structure, and trusted evidence are supplied by a Knowledge Graph (KG), or rule definitions. Classical neuro-symbolic pipelines have a discrete interface between perception and deduction. We present a neuro-soft-symbolic architecture for differentiable deductive reasoning over latent perceptual facts and knowledge-provided predicates. SoftReason removes the gradient gap by representing the deductive state as a local soft interpretation tensor over candidate constants and predicates. Perception proposes probabilistic base facts, KG triples enter as high-confidence soft evidence, and every query anchor, predicate choice, and closure update remains differentiable. Our core innovation is a learned differentiable lift of the immediate-consequence operator. It uses predicate-definition embeddings and latent composition channels to form soft body-predicate mixtures, aggregate over all possible witnesses, propose query-conditioned head facts, and update the interpretation through a monotone probabilistic OR. We instantiate the framework on Knowledge-aware Visual Question Answering (KVQA), and demonstrates how SoftReason supports end-to-end perceptual grounding, KG evidence injection, and differentiable deductive closure in one trainable architecture.
エラーからルールへ: テキスト分類のための反復プロンプト最適化
テキスト分類の迅速な最適化は、デモンストレーションの選択から探索ベースの検索、エラー主導の診断に至るまで、さまざまなアプローチに及びます。それぞれのアプローチには、既知ではあるが不完全に特性化された長所と限界があります。私たちは、最適化トレースの定量的評価と定性分析の両方を通じてこれらのパラダイムを比較する、多様な分類ベンチマーク (2 ~ 150 クラス) にわたる包括的な実証研究を実施します。これにより、各パラダイムが構造的に異なるタスク タイプで優れており、単一の手法が優勢ではないことが明らかになります。これらの洞察に基づいて、私たちは、重複しないバッチでトレーニング セット全体を反復し、分類の失敗を診断し、診断、処方、書き換えのフィードバック ループを通じて対象を絞った決定ルールを生成する、エラー主導型の手法であるエラーガイド最適化 (ERGO) を提案します。 ERGO は、エラーが特定の混乱したラベルのペアに集中するタスク (これを境界学習可能タスクと呼びます) で最高の精度を達成します。TREC: 90.0%、CLINC150: 94.4%、3 ~ 5 回の反復で収束し、解釈可能な決定ルールを生成します。 ERGO は最高の全体平均を達成することはできませんが、補完的な役割を果たしています。つまり、カバレッジに依存するタスクではデモンストレーション ベースの ICL が勝利し、多クラス インテントでは探索ベースの検索が勝利し、エラー パターンから意思決定の境界が学習できる場合は ERGO が勝利します。私たちはタスクの特性を最適なパラダイムの選択に結び付ける補完性フレームワークを提供し、実践者に実践的なガイダンスを提供します。
原文 (English)
From Errors to Rules: Iterative Prompt Optimization for Text Classification
Prompt optimization for text classification spans diverse approaches, from demonstration selection to exploration-based search to error-driven diagnosis, each with known but incompletely characterized strengths and limitations. We conduct a comprehensive empirical study across diverse classification benchmarks (2 to 150 classes) comparing these paradigms through both quantitative evaluation and qualitative analysis of optimization traces, revealing that each paradigm excels on structurally different task types and that no single method dominates. Guided by these insights, we propose Error-Guided Optimization (ERGO), an error-driven method that iterates over the full training set in non-overlapping batches, diagnoses classification failures, and generates targeted decision rules through a diagnose-prescribe-rewrite feedback loop. ERGO achieves the best accuracy on tasks where errors concentrate in specific confused label pairs (which we term boundary-learnable tasks): TREC: 90.0%, CLINC150: 94.4%, converges in 3-5 iterations, and produces interpretable decision rules. While ERGO does not achieve the highest overall average, it fills a complementary role: demonstration-based ICL wins on coverage-dependent tasks, exploration-based search wins on many-class intent, and ERGO wins where decision boundaries are learnable from error patterns. We provide a complementarity framework linking task characteristics to optimal paradigm selection, offering practical guidance for practitioners.
物理的な AI ガバナンス: ライフサイクル全体にわたる理論から実践まで
Physical AI の出現により、人工知能は画面ベースのアプリケーションを超えて、物理世界を認識し、対話し、動作する具体化されたシステムにまで拡張されています。従来の AI とは異なり、物理 AI はリアルタイムの安全制約の下で動作し、動的環境と継続的に対話し、人間と共存するため、既存の AI ガバナンス フレームワークでは明示的に対処していないガバナンスの課題が生じます。このペーパーでは、物理的 AI ガバナンスの包括的な調査を科学的および運用上の両方の観点から示します。私たちは既存のガバナンス原則を統合し、物理 AI システムに合わせた統一ガバナンス フレームワークに編成します。この基盤に基づいて、研究、設計、データ、モデル開発、展開からなる 5 段階の物理 AI ライフサイクルを提案し、具体的な実装実践を通じて各段階でガバナンスを運用する方法を実証します。この調査は、ガバナンスの原則とエンジニアリング ワークフローを結び付けることで、研究者、開発者、政策立案者が安全で信頼でき、社会的価値観と一致する物理 AI システムを構築するための構造化された参考資料を提供します。
原文 (English)
Physical AI Governance: From Theory to Practice Across Life Cycle
With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive, interact with, and act in the physical world. Unlike traditional AI, Physical AI operates under real-time safety constraints, continuously interacts with dynamic environments, and coexists with humans, introducing governance challenges that existing AI governance frameworks do not explicitly address. This paper presents a comprehensive survey of Physical AI governance from both scientific and operational perspectives. We synthesize existing governance principles and organize them into a unified governance framework tailored to physical AI systems. Building on this foundation, we propose a five-stage Physical AI lifecycle comprising research, design, data, model development, and deployment, and demonstrate how governance can be operationalized across each stage through concrete implementation practices. By connecting governance principles with engineering workflows, this survey provides a structured reference for researchers, developers, and policymakers to build Physical AI systems that are safe, trustworthy, and aligned with societal values.
RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning
Rare diseases represent one of the most challenging settings for clinical decision-making, where heterogeneous presentations, sparse eviden…
ベンチマーク推論が成立しない場合:AI評価における予測可能性
AI ベンチマークの結果が 1 ステップで重大な主張に達することはほとんどありません。評価者は、それをさらなるケースに一般化し、能力の証拠として解釈し、新しいタスクに推定し、別のシステムまたはサイトに移し、人間によるレビューと下流の結果に関する仮定と組み合わせます。妥当性中心のアプローチでは、各主張の証拠が必要です。この論文では、さらなる認識論的問題を特定します。それは、保証されたリンクは自動的に保証されたチェーンを作成しないということです。ある研究のターゲットが次の研究のソースになるとは限りません。システム、人口、結果、または条件はインターフェースで変更される可能性があります。また、共有データやモデル系統により、一見独立したサポートに依存する可能性があります。予測可能性は、観察されたケースから観察されていないケースへの限定された拡張が保証されるかどうかに関係します。グッドマンは競合拡張の問題を提供します。引数ベースの妥当性は、それらをテストするためのアーキテクチャを提供します。この論文の特徴的な主張は、非合成原理です。隣接する投影のサポートは、エンドポイントと仮定が一致し、依存性と不確実性が貫徹される場合にのみ合成を保証します。法的調査の事例は、ベンチマーク証拠と展開調査がそれぞれ並行性を保ちながらどのように健全であるかを示しています。再分析とシミュレーションは、骨材の安定性が後の予測で必要となる区別を消去できる理由を示しています。結果として得られる投影可能性監査により、ベンチマークで使用する引数内のサポートされていない結合が診断されます。
原文 (English)
When benchmark inferences do not compose: Projectibility in AI evaluation
An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper makes explicit and operationalizes a problem those approaches leave to the analyst: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The contribution is an interface audit for distributed AI evidence: typed source and target descriptions, and a procedure separating endpoints that never meet from endpoints that meet while warrant fails to cross. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A known-truth demonstration shows why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.
SemPIC: Learning Semantic Position-Independent KV Caches
Long-context retrieval and agentic workloads repeatedly reuse the same documents under changing instructions, histories, and document order…
ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
Long-horizon steel-equipment inspection requires reasoning over heterogeneous records accumulated across repeated inspection cycles. Existi…
A foundation model of numerical intelligence with cross-disciplinary generalization
Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. L…
InfoOps Bench: A live information operations safety benchmark
In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being…
Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates
Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fiel…
SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs
Although large language models (LLMs) can produce fluent spatial reasoning traces, their intermediate relations may fail to support the fin…
Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models
A calibrated stochastic world model can reveal how uncertain a future is without revealing why it branches. The same conditional future law…
Evolutionary Curriculum Learning Improves Biological Sequence Modeling
Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological se…
The Scaling Paradox in Human-AI Collaboration
The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems sc…
Agentic Stage-One Stellarator Optimization: Autonomous Multi-Objective Search for Finite-Beta Equilibria
Stage-one stellarator design searches a high-dimensional family of three-dimensional plasma boundaries and fixed-boundary MHD equilibria fo…
Emergence Invariance: From Symbolized Thought to Structural Control
Language-first intelligence is constrained by which distinctions enter its symbolic record, which mappings its language--interpreter--envir…
State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
Transformer-based architectures have dominated sequence modeling, largely due to the expressive power of attention mechanisms. However, for…
When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking
Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy…
The Transformer Revolution, Part 1: Dynamic Processing through Output-Weight Interconnections
This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models…
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load
Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law…
Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements…
人間の認知と行動の小規模な基礎モデル
人間の行動データに基づいて微調整された大規模な言語モデルが、汎用の認知プロキシとして登場しましたが、これに必要な規模や、これらのモデルがタスク構造を処理するのか、それとも統計的ショートカットを利用するのかは未解決のままです。私たちは、160 の実験からの 1,070 万の試験レベルの選択肢のデータセットである Psych-101 上の 4 つのアーキテクチャ ファミリにわたる 135M から 14B のパラメーターの 14 のモデルをトレーニングします。流通においては、規模はほとんど重要ではありません。モデルは、あたかも天井に向かっているかのように狭い帯域内に収まり、参加者が参加しなかった場合の 70B のベースラインに一致するには、0.6B ~ 1B のパラメータで十分です。分布外では、そのバンドは著しく急峻なスケーリング勾配に向かって開き、新しいタスク構造への一般化において、より大きなモデルが明らかに有利になります。これらのモデルがどのような情報を使用するかを判断するために、2 つの診断を実行します。 27 回の実験にわたって、タスクの指示、実験刺激、結果のフィードバック、選択履歴という 4 つのプロンプト チャネルを段階的に削除し、試行順序を変更します。刺激とフィードバックの内容をマスキングすると、学習した情報の 75.7% が破壊され、モデルが確率以下に押し下げられます。これは、選択履歴だけではパフォーマンスを考慮できないことを示しています。順列は、独立した試行を伴うタスクの不変性を明らかにしますが、試行の順序が前の応答によって決定される感度を明らかにします。したがって、認知的に微調整された小さなモデルは、心理学実験のノイズ上限推定器として有望ですが、その範囲はトレーニングで見られるパラダイムによって制限されたままです。
原文 (English)
Small Foundation Models of Human Cognition and Behaviour
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. For in-distribution simulations, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.
生成 AI における認識論的信頼性: 一か八かのワークフローにおける信頼を保証するための規範的なフレームワーク
生成 AI システムは、一か八かの専門的な状況で導入されることが増えており、その出力によって、ユーザーが何を信じ、どのように推論し、何を定着として扱うのかが決まります。これは、責任ある AI に対する中心的な疑問を引き起こします。それは、どのような状況下で、生成的な AI の出力への依存が、行動によって引き起こされるのではなく、認識論的に正当化されるのかということです。既存のフレームワークでは主に、AI の出力が正確か、公平か、説明可能か、安全か、ユーザーに信頼されているかどうかが問われます。これらの質問は依然として必要であり、それぞれが正当な信頼に貢献する可能性があります。ただし、それらは、明確な評価目標、つまりユーザーが AI の出力を自分の推論への入力として扱うことが正当化される条件として、保証された信頼性を直接指定していません。私たちは、これには認識論的信頼性、つまりシステムを認識論的に信頼に値するものにするものについての説明が必要であると主張します。能力と聴衆指向としての信頼性に関する哲学的説明に基づいて、私たちは 3 つの共同で必要で代替不可能な条件からなる構成的規範の枠組みを開発します。まず、認識論的謙虚さには、システムがその能力の限界を表現し、伝達することが必要です。第 2 に、認識的アクセスには、ユーザーがコンテキスト内で出力を検査、質問、および異議申し立てできるようにするシステムが必要です。第三に、認識論的不正義に抵抗するには、システムがユーザーを正当な認識論的主体として認識し、ユーザーの知識や経験を疎外しないようにする必要があります。法的推論、医学的推論、雇用における実際の事例分析を通じて、認識論的謙虚さ、認識論的アクセス、認識論的不正義に対する抵抗の失敗が、精度、公平性、使いやすさの標準的な尺度だけでは対処できない結果的な損害をどのように生み出す可能性があるかを示します。最後に、出力の正確さのみではなく、認識論的に保証された信頼性を中心に構成された GenAI システムの設計と評価への影響を概説します。
原文 (English)
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows
Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientation, we develop a constitutive normative framework comprising three jointly necessary and non-fungible conditions. First, epistemic humility requires systems to represent and communicate the limits of their competence. Second, epistemic access requires systems to enable users to inspect, question, and contest outputs in context. Third, resistance to epistemic injustice requires systems to recognise users as legitimate epistemic agents and avoid marginalising their knowledge and experience. Through real-world case analyses in legal reasoning, medical reasoning, and hiring, we show how failures of epistemic humility, epistemic access, and resistance to epistemic injustice can produce consequential harms that standard measures of accuracy, fairness, and usability do not address on their own. We conclude by outlining design and evaluation implications for GenAI systems organised around epistemically warranted reliance rather than output correctness alone.
ViSR-KGC: マルチモーダル ナレッジ グラフを完成させるための視覚言語モデルを使用したビジュアル サブグラフ推論
ナレッジ グラフ補完 (KGC) は、不完全なグラフ構造から欠落しているエンティティまたは関係を推測することを目的としており、エンティティがテキストや画像などの複数のモダリティに関連付けられるマルチモーダル ナレッジ グラフ補完 (MMKGC) に進化しました。従来の表現学習アプローチは埋め込みベースのパラダイムに従っており、関係固有の証拠が限られている場合には困難を伴う可能性があります。一方、LLM ベースの推論手法は通常、グラフ構造をテキスト プロンプトに線形化するため、構造トポロジが曖昧になり、重要な視覚情報が無視されます。ビジョン言語モデル (VLM) はマルチモーダル推論に優れていますが、構造化されたグラフ トポロジをネイティブに解釈することはできません。特に、ノードとエッジが複雑なセマンティクスを運ぶナレッジ グラフの場合はそうです。このギャップを埋めるために、KGC の視覚的なサブグラフ推論アプローチである ViSR-KGC を提案します。セマンティック相関を捕捉するための 3 つの補完的な機能が統合されています。表現学習によるグローバル トポロジの依存関係の特定、VLM を使用したローカル マルチモーダル証拠の分析、および事前トレーニングされたモデルに固有の必要な常識知識の提供です。学習されたマルチモーダル埋め込みに基づいて、私たちのフレームワークは最初に MMKG からコンパクトでクエリ対応のサブグラフを抽出します。次に、このサブグラフは、経験的比較によって選択されたレイアウト戦略を使用して、視覚的に解釈可能な画像に変換されます。最後に、視覚化されたサブグラフ、エンティティ画像、テキストによる説明、および回答候補が統合プロンプトに結合され、VLM が欠落しているエンティティを推測できるようになります。
原文 (English)
ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion
Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images. Traditional representation learning approaches follow the embedding-based paradigm and may struggle when relation-specific evidence is limited. Meanwhile, LLM-based reasoning methods typically linearize graph structures into textual prompts, which obscures structural topology and neglects vital visual information. While vision-language models (VLMs) excel at multimodal reasoning, they cannot natively interpret structured graph topology, particularly when it comes to knowledge graphs where nodes and edges carry complex semantics. To bridge this gap, we propose ViSR-KGC, a visual subgraph reasoning approach for KGC. It integrates three complementary capabilities to capture semantic correlations: identifying global topology dependencies via representation learning, analyzing local multimodal evidence using VLMs, and providing necessary commonsense knowledge inherent in pre-trained models. Based on learned multimodal embeddings, our framework first extracts a compact and query-aware subgraph from the MMKG. Then, this subgraph is transformed into a visually interpretable image using a layout strategy selected through empirical comparison. Finally, the visualized subgraph, entity images, textual descriptions, and candidate answers are combined into a unified prompt, enabling the VLM to infer the missing entity.
ポリバイアス: 国際政治紛争における大規模言語モデルのバイアスの理解と測定
大規模言語モデル (LLM) での政治的偏見の測定は、単一の指標で把握するのが難しい枠組み、議論、法的推論の微妙な違いを通じて現れる可能性があるため、依然として困難です。この研究では、LLM が法的に同等の紛争シナリオを関係国に応じて異なる方法で扱うかどうかを測定するための反事実フレームワークである Poli-Bias を紹介します。 Poli-Bias は、さまざまな地政学的関係、法律違反、および推論タスクにわたって国のアイデンティティが体系的に交換されるペアのプロンプトに対する応答を比較します。私たちのフレームワークは、バイアスを 1 つの判断に限定するのではなく、回答の格差を解釈可能な 5 つの側面に分解し、不平等な扱いがどこでどのように現れるかを明らかにします。さまざまなモデルファミリーと規模にまたがる 13 の現代の LLM にわたって、国のアイデンティティとユーザーの所属が、同等の行為が国際法の下でどのように記述され、評価され、擁護されるかに系統的に影響を与える可能性があることがわかりました。したがって、私たちの結果は、LLM における政治的平等性と媚びを監査するためのきめ細かいフレームワークとして Poli-Bias を確立しました。
原文 (English)
Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts
Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks. Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests. Across 13 contemporary LLMs spanning diverse model families and sizes, we find that country identities and user affiliations can systematically affect how equivalent actions are described, evaluated, and defended under international law. Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.
自動アイテム評価: LLM によって生成された批評を使用してアイテムの受け入れと拒否を予測します。
自動品目評価 (AIE) とは、評価対象品目の専門家による手動レビューやフィールド テストを必要とせずに品目の品質を評価するための計算手法の使用を指します。私たちは、大規模な標準化されたテスト プログラムからの過去の不合格データを使用して、品目テキストから品目の合格と不合格を予測することにより、ほぼ包括的な AIE モデルを構築することを目指しました。データセットには 52,759 件の英語芸術 (ELA) と数学の項目が含まれており、そのうち 34% は将来の運用上の使用から永久に拒否されました。拒否の理由には、不十分な心理測定特性、コンテンツの問題、偏見と感受性への懸念、およびコンテンツ以外の問題が含まれていました。私たちは、生のアイテム テキストに対する DeBERTaV3 ラージ分類器、Qwen3 で生成されたアイテム批評に対する 2 番目の DeBERTa 分類器、および両方の表現を組み合わせた融合モデルを微調整しました。融合モデルは、最も強力な全体的なパフォーマンスを達成しました (精度 = 0.75、F1 = 0.64、AUC = 0.80、感度 = 0.64、特異性 = 0.81)。数学的予測 (F1 = .73、AUC = .86) は、ELA (F1 = .51、AUC = .72) よりもかなり正確でした。判定しきい値を 0.5 から 0.25 に下げると、ELA と数学の平均感度は 0.88 と 0.91 に上昇しましたが、特異度はそれぞれ 0.31 と 0.56 に低下しました。これは、アイテムを評価するよりも生成する方が安価である自動アイテム生成のコンテキストでは好ましいと考えられます。アイテムの生のテキストと一緒にアイテムの批評を組み込むことで、ほとんどの拒否理由でパフォーマンスが向上しました。このモデルでは、より困難な項目ほど高い拒否確率が割り当てられました。ただし、融合モデルは、特に ELA に関して、バイアス、機密性、公平性、またはアクセシビリティについてフラグが立てられた項目を特定するのに苦労しました。これらの調査結果は、テキストベースの AIE が一部の分野では実現可能であり、手動レビューやフィールドテストの負担を軽減する実用的なツールとなる可能性があることを示唆するとともに、公平性の懸念がある項目については人間によるレビューの重要性も強調しています。
原文 (English)
Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques
Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item acceptance and rejection from item text using historical rejection data from a large-scale standardized testing program. The dataset contained 52,759 English language arts (ELA) and mathematics items with 34% permanently rejected from future operational use. Rejection reasons included poor psychometric properties, content issues, bias and sensitivity concerns, and non-content issues. We fine-tuned a DeBERTaV3-large classifier on raw item text, a second DeBERTa classifier on Qwen3-generated item critiques, and a fusion model combining representations from both. The fusion model achieved the strongest overall performance (Accuracy = .75, F1 = .64, AUC = .80, Sensitivity = .64, Specificity = .81). Prediction for math (F1 = .73, AUC = .86) was considerably more accurate than ELA (F1 = .51, AUC = .72). Lowering the decision threshold from .5 to .25 raised average sensitivity for ELA and math to .88 and .91, while reducing specificity to .31 and .56, respectively, which may be preferable in automated item generation contexts where generating items is cheaper than evaluating them. Incorporating item critiques alongside raw item text improved performance across most rejection reasons. The model assigned higher rejection probabilities to more difficult items. However, the fusion model struggled to identify items flagged for bias, sensitivity, fairness, or accessibility, especially for ELA. These findings suggest that text-based AIE is feasible in some areas and may offer a practical tool for reducing the burden of manual review and field testing, while also underscoring the importance of human review for items with fairness concerns.
bioMoR: 効果的なゲノム学習のための生物学に基づく再帰混合
高次元オミクス解析用のトランスフォーマー モデルは数千の遺伝子または経路を処理しますが、詳細な計算が必要なのはサブセットのみです。 Mixture-of-Recursions (MoR) は、適応型トークン選択またはエキスパート選択ルーティングを通じて効率を向上させます。私たちは bioMoR を提案します。これは、私たちの知る限り、MoR を遺伝子レベルおよび経路レベルの学習に適用する最初のフレームワークです。私たちの貢献には、MoR バックボーン内に構造化された生物学的知識を統合するための 3 つの場所を特定することが含まれます。グラフベースの情報共有によりトークンの埋め込みが洗練され、構造的バイアスにより生物学的に関連するトークンへの注意が誘導され、グラフ認識ルーターが近傍情報を使用して各トークンの再帰深さを決定します。これらの手法は、トークン相互作用に関する追加の知識が、モデルが埋め込みを構築し、どのトークンをより深く学習する必要があるかを選択するのに効果的に役立つという洞察に基づいています。多様なオミクス データ タイプにまたがる 8 つのベンチマークにわたって、統一された 5 重交差検証プロトコルに基づいて評価された bioMoR は、非再帰的 Transformer よりも使用するパラメーターが 75 パーセント少なく、FLOP が最大 58 パーセント少ない一方で、生物学に依存しない最も強力な MoR ベースラインと比較して、平均マクロ F1 が 8.2 パーセント ポイント、バランスのとれた精度が 7.1 パーセント ポイント向上しています。選択されたマーカー遺伝子または経路は生物学的解釈可能性を提供し、そのトークン固有の再帰の深さは計算がどのように割り当てられているかを明らかにします。
原文 (English)
bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning
Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture-of-Recursions (MoR) improves efficiency through adaptive token-choice or expert-choice routing. We propose bioMoR, which, to the best of our knowledge, is the first framework to apply MoR to gene-level and pathway-level learning. Our contributions include identifying three locations for integrating structured biological knowledge within an MoR backbone: graph-based information sharing refines token embeddings, a structural bias guides self-attention toward biologically related tokens, and a graph-aware router uses neighborhood information to determine each token's recursion depth. These techniques are centered on our insight that additional knowledge of token interaction can effectively help models construct embeddings and select which tokens should be learned more deeply. Across eight benchmarks spanning diverse omics data types and evaluated under a unified five-fold cross-validation protocol, bioMoR improves average macro-F1 by 8.2 percentage points and balanced accuracy by 7.1 percentage points over the strongest biology-agnostic MoR baseline while using 75 percent fewer parameters and up to 58 percent fewer FLOPs than a non-recursive Transformer. The selected marker genes or pathways provide biological interpretability, while their token-specific recursion depths reveal how computation is allocated.
エンドツーエンドのエージェント監査エンジン
大規模言語モデル (LLM) の急速な進歩により、ハーネスは幅広いドメインにエージェントを展開するための不可欠なインフラストラクチャになりました。ハーネスのエコシステムが急速に進化しているため、厳密な機能評価の重要性も高まっています。ただし、エンドツーエンドの体系的かつ包括的な評価パイプラインを効率的に構築することは依然として大きな課題です。この課題に対処するために、エージェント ハーネス用に設計されたエンドツーエンドの評価エンジンである $A^2E$ (エージェント監査エンジン) を導入します。 $A^2E$ は、新しく提案されたエージェント タスク プロトコル (ATP) を活用して、評価タスクとさまざまなハーネスの迅速な統合を可能にします。自動的に計測されたモニターを通じて、実験中に標準化された実行トレースをキャプチャして生成します。評価ステージでは、$A^2E$ は一連の多次元メトリックを使用してハーネス機能を体系的に評価します。これらのメトリクスは、正確性だけと比較して、実行効率、ツールの使用、タスク計画、およびエラー回復におけるハーネス間の違いをより詳細に特徴付けることができます。 $A^2E$ を使って行われた実験では、モデルとハーネスの組み合わせがさまざまな種類のタスク間で大幅なパフォーマンスのばらつきを示し、単一の組み合わせがすべてのタスクにおいて一貫して他のすべての組み合わせを上回ることはないことがさらに明らかになりました。これらの発見は、体系的な評価の必要性を実証するだけでなく、モデルとハーネスを共同進化させるための有用な指針も提供します。私たちのコードは https://github.com/datamllab/A2E で入手できます。
原文 (English)
$A^2E$ : An End-to-End Agent Auditing Engine
With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving harness ecosystem has also made rigorous capability evaluation increasingly important. However, efficiently building an end-to-end, systematic, and comprehensive evaluation pipeline remains a significant challenge. To address this challenge, we introduce $A^2E$ (Agent Auditing Engine), an end-to-end evaluation engine designed for agent harnesses. $A^2E$ leverages our newly proposed Agent Task Protocol (ATP) to enable the rapid integration of evaluation tasks with different harnesses. Through an automatically instrumented Monitor, it captures and generates standardized execution traces during experiments. In the Evaluation stage, $A^2E$ systematically assesses harness capabilities using a suite of multidimensional metrics. Compared with correctness alone, these metrics provide a more fine-grained characterization of differences among harnesses in execution efficiency, tool use, task planning, and error recovery. Experiments conducted with $A^2E$ further reveal that model-harness combinations exhibit substantial performance variation across different types of tasks, and that no single combination consistently outperforms all others across every task. These findings not only demonstrate the necessity of systematic evaluation but also provide useful guidance for the co-evolving of models and harnesses. Our code is available at https://github.com/datamllab/A2E.
CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or a…
TEPA: 競合に強い言語エージェントの古い記憶を取り消す
長期記憶により、言語エージェントは過去の事実、好み、タスクの経験を再利用できます。永続性は、主要な反証可能性の問題も引き起こします。世界が変化しても、古い記憶が検索可能なままになり、プロンプトが汚染される可能性があります。我々は、この障害モードをメモリ汚染、つまり、より新しい矛盾する証拠が置き換えられたアクティブなメモリによって引き起こされる劣化として特徴付けます。有効性を明示的な記憶状態にする、取り消し可能な証拠記憶メカニズムである TEPA を紹介します。 TEPA は観察をキー付き先例として表し、同じキーで新しい証拠が矛盾する場合に有効な先例を取り消します。これにより、取り消された履歴を監査用に保存しながら、現在の証拠から検索を行うことができます。制御された非表示体制のドリフト、実際のファイルにバックアップされた実行可能ファイルのドリフト、および設定更新ストリーム全体にわたって、取り消しにより、取り消し後の取得セットに古いアクティブ メモリが残るのを防ぎます。 50 シードを超える制御されたドリフトでは、追加のみおよび最終書き込み優先のメモリは完全反転中にメモリなしを下回りました (追加のみおよび最終書き込み優先の両方が 0.210、メモリなし 0.309、TEPA 0.950) が、実際のファイル実行でも同じパターンが再現されました (追加のみ 0.203、メモリなし 0.298、TEPA 0.950)。クリーンな MemoryAgentBench SH-6k では、TEPA は強力な Last-write-wins キャッシュと一致し、現在のキーの置換がシングルホップ ファクト統合にとって決定的な操作であることを確認します。マルチホップおよび非常に長いコンテキストの MemoryAgentBench 設定の境界テストにより、ファクトレベルの妥当性追跡を超えた検索チェーンとコンテキスト選択のボトルネックが明らかになります。これらの結果を総合すると、進化する知識を改ざん、監査し、後で再促進する必要があるエージェントにとって、ライフサイクル失効が中核的な記憶操作として確立されます。
原文 (English)
TEPA: Revoking Stale Memories for Conflict-Robust Language Agents
Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability problem: when the world changes, stale memories can remain retrievable and pollute the prompt. We characterize this failure mode as memory pollution: degradation caused by active memories that newer conflicting evidence has superseded. We introduce TEPA, a revocable evidence-memory mechanism that makes validity an explicit state of memory. TEPA represents observations as keyed precedents and revokes active precedents when fresh evidence contradicts them under the same key, allowing retrieval to draw from current evidence while preserving revoked history for audit. Across controlled hidden-regime drift, real file-backed executable drift, and preference-update streams, revocation prevents stale active memory from remaining in the retrieval set after reversal. In controlled drift over 50 seeds, append-only and last-write-wins memory fell below no memory during full reversal (append-only and last-write-wins both 0.210, no memory 0.309, TEPA 0.950), and the same pattern reproduced under real file execution (append-only 0.203, no memory 0.298, TEPA 0.950). On clean MemoryAgentBench SH-6k, TEPA matches a strong last-write-wins cache, confirming that current-key replacement is the decisive operation for single-hop fact consolidation. Boundary tests on multi-hop and very long-context MemoryAgentBench settings expose retrieval-chain and context-selection bottlenecks beyond fact-level validity tracking. Together, these results establish lifecycle revocation as a core memory operation for agents that must falsify, audit, and later re-promote evolving knowledge.
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
In this paper, we focus on providing convergence guarantees for stochastic subgradient methods in minimizing nonsmooth nonconvex functions.…
Explainable Machine Learning-Based Security and Privacy Protection Framework for Internet of Medical Things Systems
The Internet of Medical Things transcends traditional medical boundaries, enabling a transition from reactive treatment to proactive preven…
Ethical Framework for Responsible Foundational Models in Medical Imaging
The emergence of foundational models represents a paradigm shift in medical imaging, offering extraordinary capabilities in disease detecti…
Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation
The Transformer architecture underpins modern large language models powering state-of-the-art text generation and AI applications. However,…
See Me, Believe Me: Causality, Intersectionality, and Interventions Improving the Appearance of Patients
In the context of medical records, patients often experience testimonial injustice, where the textual account undermines the validity of th…
A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence
Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instr…
LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models
Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as…
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that requi…
When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models…
An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning
Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrate…
TreeHop: Efficient Embedding-Level Query Rewriter
Retrieval-augmented generation (RAG) systems face significant challenges in multi-hop question answering (MHQA), where complex queries requ…
Optimal Transport for Machine Learners
Modern machine learning repeatedly manipulates probability measures: empirical datasets, generated samples, latent distributions, class-con…
X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation
Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significa…
SuperCoder: Assembly Program Superoptimization with Large Language Models
Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its…
SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning
Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly avai…
Transformer-Based Neural Quantum Digital Twins for Many-Body Spectral Reconstruction and Adaptive Quantum-Annealing Schedule Design
We introduce Transformer-based Neural Quantum Digital Twins (Tx-NQDTs) to reconstruct the low-energy spectral evolution of many-body quantu…
FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records
Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of la…
The Cell Must Go On: Agar.io for Continual Reinforcement Learning
Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then…
HyperFake: Hyperspectral Reconstruction and Attention-Guided Analysis for Advanced Deepfake Detection
Deepfakes pose a significant threat to digital media security, with current detection methods struggling to generalize across different man…
From Alignment to Synthesis Contrastive Volumetric Grounding for Text-to-CT Generation
Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-langua…
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transpare…
Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence
In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespre…
Transformer Circuits Can Realize Clustering Algorithms
Although transformers are most commonly optimized as statistical sequence models, it is unclear to what extent they can implement and learn…
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These…
Dynamic gain neuromodulation attenuates the stability gap under joint training
Recent work in continual learning has highlighted the stability gap -- a temporary performance drop on previously learned tasks when new on…
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within…
Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings
Knowledge Tracing (KT) aims to predict a student's future performance based on their sequence of interactions with learning content. Many K…
Deep Residual Echo State Networks: exploring residual orthogonal connections in untrained Recurrent Neural Networks
Echo State Networks (ESNs) are a particular type of untrained Recurrent Neural Networks (RNNs) within the Reservoir Computing (RC) framewor…
NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. Ho…
ATLASFusion: Aggregation Tracking with Location-Aware Sparse Fusion for Robust Spatio-Temporal Multi-View Pedestrian Tracking
For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintainin…
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
Program verification relies on loop invariants, yet automatically discovering strong invariants remains a long-standing challenge. We inves…
Topographic Constraints Shape Brain-Like Component Structure in Auditory Models
If topography is a fundamental feature of the brain, it should influence both how neurons are arranged in space (i.e. explain brain maps) a…
Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents
LLM agents require personal information for personalization in order to effectively act on users' behalf, but this raises privacy concerns…
Deep Generative Model for Human Mobility Behavior
Understanding and modeling human mobility is central to challenges in transport planning, sustainable urban design, and public health. Desp…
OBCache: 効率的なロングコンテキスト LLM 推論のための最適な Brain KV キャッシュ プルーニング
拡張コンテキスト ウィンドウを備えた大規模言語モデル (LLM) は強力なアプリケーションを可能にしますが、すべてのキー/値 (KV) 状態のキャッシュがシーケンスの長さとバッチ サイズに比例して拡張されるため、大幅なメモリ オーバーヘッドが発生します。既存のキャッシュエビクション方法は、アテンションの希薄性を利用することでこの問題に対処していますが、通常、アテンションの出力に対する真の影響を考慮せずに、蓄積されたアテンションの重みを使用してヒューリスティックにトークンをランク付けします。私たちは、層ごとに構造化された枝刈り問題としてキャッシュの追い出しを定式化する原則的なフレームワークである、Optimal Brain Cache (OBCache) を提案します。 OBCache は、最適脳損傷 (OBD) 理論に基づいて構築されており、分離キー、分離値、および結合キーと値のペアに対して導出された閉形式スコアを使用して、トークンの枝刈りによって引き起こされる注意出力の摂動を測定することにより、トークンの顕著性を定量化します。私たちのスコアはアテンションの重みだけでなく、値の状態とアテンションの出力からの情報も考慮に入れており、それによって出力を意識した信号で既存のエビクション戦略を強化します。 LLaMA および Qwen モデルの実験では、さまざまなクエリ位置にわたるトークンの顕著性を推定する既存の作業のヒューリスティック スコアを OBCache の出力対応スコアに置き換えることで、長いコンテキストの精度が一貫して向上することが実証されました。コードは https://github.com/DreamSoul-AI/OBCache で入手できます。
原文 (English)
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
Large language models (LLMs) with extended context windows enable powerful applications but impose significant memory overhead, as caching all key-value (KV) states scales linearly with sequence length and batch size. Existing cache eviction methods address this by exploiting attention sparsity, yet they typically rank tokens heuristically using accumulated attention weights without considering their true impact on attention outputs. We propose Optimal Brain Cache (OBCache), a principled framework that formulates cache eviction as a layer-wise structured pruning problem. Building upon the Optimal Brain Damage (OBD) theory, OBCache quantifies token saliency by measuring the perturbation in attention outputs induced by pruning tokens, with closed-form scores derived for isolated keys, isolated values, and joint key-value pairs. Our scores account not only for attention weights but also for information from value states and attention outputs, thereby enhancing existing eviction strategies with output-aware signals. Experiments on LLaMA and Qwen models demonstrate that replacing the heuristic scores in existing works, which estimate token saliency across different query positions, with OBCache's output-aware scores consistently improves long-context accuracy. Code is available at https://github.com/DreamSoul-AI/OBCache.
SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation
Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily rel…
Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
Thermal management in 3D ICs is increasingly challenging due to higher power densities. Traditional PDE-solving-based methods, while accura…
NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation…
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended promp…
QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent f…
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in se…
Reasoning about Intent for Ambiguous Requests
Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating sa…
iLTM: Integrated Large Tabular Model
Tabular data underpins decisions across science, industry, and public services. Despite rapid progress, advances in deep learning have not…
VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate v…
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
The SmoothLLM defense provides a certification guarantee against jailbreaking attacks, but it relies on a strict "k-unstable" assumption th…
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request pave…
Length-MAX Tokenizer for Language Models
We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens nee…
Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics
In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily ta…
Multilingual Agent-Based World Modeling for Social Science
Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly…
The Theory of Strategic Evolution: Games with Endogenous Players and the Seven Laws of Strategic Replicators
Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. Rational players do no…
Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models
Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuris…
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited para…
TGIF: Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
Multimodal large language models (MLLMs) typically rely on a single late-layer feature from a frozen vision encoder, leaving the encoder's…
IndexTTS 2.5 Technical Report
In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-b…
ReMIND: Orchestrating Modular Large Language Models for Controllable Serendipity A REM-Inspired System Design for Emergent Creative Ideation
Large language models (LLMs) are increasingly used not only for problem solving but also for creative ideation; however, generating ideas t…
Layerwise goal-oriented adaptivity for neural ODEs: an optimal control perspective
In this work, we propose a novel layerwise adaptive construction method for neural network architectures. Our approach is based on a goal--…
Expert-Guided Multimodal Fusion for Unified Emotion and Sentiment Analysis
Multimodal emotion understanding requires the integration of heterogeneous data sources, including text, audio, and visual modalities, whil…
LAUDE: LLM-Assisted Unit Test Generation and Debugging of Hardware DEsigns
Unit tests are critical in the hardware design lifecycle to ensure that component design modules are functionally correct and conform to th…
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, y…
Communication-efficient distributed hazard difference estimation for heterogeneous multi-site survival data
Multi-site collaboration can power survival models that no single hospital could fit alone, but privacy rules and protected computing envir…
Hybrid Mamba-Attention Neural Architecture for Channel Estimation
This paper proposes a hybrid Mamba-attention neural architecture to achieve improved channel estimation for orthogonal frequency-division m…
SNR-Edit: Structure-Aware Noise Rectification for Inversion-Free Flow-Based Editing
Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines. However, existing appr…
Temporal Sepsis Modeling: a Relational and Explainable-by-Design Framework
Sepsis remains one of the most complex and heterogeneous syndromes in intensive care. While deep learning models achieve competitive perfor…
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mecha…
Universal One-third Time Scaling in Learning Peaked Distributions
Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origi…
予測符号化ネットワークの無限の幅と深さの制限について
予測コーディング (PC) は、重みを更新する前にネットワーク アクティビティに関するエネルギー関数を最小化する、標準的な逆伝播 (BP) に代わる生物学的に妥当な代替手段です。最近の研究では、BP にヒントを得た再パラメータ化を活用することで、ディープ PC ネットワーク (PCN) のトレーニングの安定性が向上しました。ただし、これらの方法の完全なスケーラビリティと理論的根拠は依然として不明です。このギャップに対処するために、PCN の無限の幅と深さの制限を研究します。線形残差ネットワークの場合、PC の幅と深さの安定した特徴学習パラメーター化のセットが BP の場合とまったく同じであることを示します。さらに、これらのパラメータ化のいずれかの下では、モデルの幅が深さよりもはるかに大きい場合、平衡アクティビティの PC エネルギーは二次 BP 損失に収束し、PC が BP と同じ勾配を計算することになります。実験では、アクティビティの平衡に達している限り、畳み込みネットワークや変換器を含む非線形モデルでは BP への収束が維持されることが示されています。全体として、この研究は PC でスケーラブルなパラメータ化のタイプを制限する一方で、脳のような深いネットワークよりもはるかに広いネットワークでローカル更新のみを使用して BP を効果的に実装できる方法を示しています。
原文 (English)
On the Infinite Width and Depth Limits of Predictive Coding Networks
Predictive coding (PC) is a biologically plausible alternative to standard backpropagation (BP) that minimises an energy function with respect to network activities before updating weights. Recent work has improved the training stability of deep PC networks (PCNs) by leveraging some BP-inspired reparameterisations, but the scalability and theoretical basis of these methods remain unclear. To address this gap, we study the infinite width and depth limits of PCNs. For linear networks, we derive stable and "non-lazy" parameterisations when scaling both the model width and depth, revealing that the output of standard PCNs explodes with width during training. Moreover, under stable parameterisations, we show that the gradients computed by PC at activity equilibrium converge to the BP gradients for networks that are much wider than deep ($depth/width\to0$). Experiments show high gradient alignment between PC and BP at large width for different nonlinear models, including convolutional networks and transformers. Overall, this work constrains the parameterisations that are scalable with PC, while suggesting how BP could be implemented using only local updates in much wider than deep networks like the brain.
SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer
Modern offline Reinforcement Learning (RL) methods find performant actor-critics, however, fine-tuning these actor-critics online with valu…
漏洩する思考から個人的な推論まで: LRM が自分自身に言うことを制御する
大規模推論モデル (LRM) は、多くの場合機密情報を含む推論トレース (RT) を生成します。こうした漏洩する考えは制御するのが難しく、明示的なプライバシー指令に違反することがよくあります。 RT はプロンプト インジェクション攻撃によって公開される可能性があるため、これはユーザーにとって直接的なプライバシー リスクになります。私たちはこれを制御性の問題としてアプローチします。プライバシー ディレクティブ自体が命令であるため、RT 内の命令フォローイング (IF) を改善することがプライバシー漏洩を減らす直接的な方法となります。この目的を達成するために、モデルに推論プロセス全体を通じて一般的な指示に従うように教える SFT データセットを導入し、段階的デコードを提案します。これは、各コンポーネントの IF を最大化するために個別の LoRA アダプターを使用して RT と応答生成を分離する単純なデコード戦略です。 2 つの IF ベンチマークと 2 つのプライバシー ベンチマークにわたって、2 つのファミリー (1.7B ~ 14B パラメーター) の 6 つのモデルでアプローチを評価しました。私たちの方法では、IF で最大 20.9 ポイント、プライバシー ベンチマークで最大 51.9 パーセント ポイントの向上が見られ、大幅な改善が見られます。ただし、推論のパフォーマンスと IF の間のトレードオフにより、タスクの有用性が犠牲になる可能性があります。私たちの結果は、LRM の IF を改善することでプライバシーを大幅に強化できることを示しており、将来のプライバシーを意識した LRM の有望な方向性を示唆しています。私たちのコードは https://github.com/UKPLab/arxiv2026-controllable-reasoning-models で入手できます。
原文 (English)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult to control and frequently violate explicit privacy directives. Because RTs can be exposed through prompt injection attacks, this becomes a direct privacy risk to the user. We approach this as a controllability problem: since privacy directives are themselves instructions, improving instruction-following (IF) within the RT provides a direct path to reducing privacy leaks. To this end, we introduce an SFT dataset that teaches models to follow general instructions throughout their reasoning process, and propose Staged Decoding, a simple decoding strategy that decouples RT and answer generation using separate LoRA adapters to maximize IF of each component. We evaluate our approach on six models from two families (1.7B-14B parameters), across two IF benchmarks and two privacy benchmarks. Our method yields substantial improvements, with gains of up to 20.9 points in IF and 51.9 percentage points on privacy benchmarks, though these can come at the cost of task utility due to the trade-off between reasoning performance and IF. Our results show that improving IF in LRMs can significantly enhance privacy, suggesting a promising direction for future privacy-aware LRMs. Our code is available at https://github.com/UKPLab/arxiv2026-controllable-reasoning-models.
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the…
Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks
Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be re…
LLMs Remember First, Forget Last: Dual-Process Interference in Large Language Models
Large language models can process millions of tokens, yet how they handle conflicting information within context remains poorly understood.…
PolypSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
Generative diffusion models are increasingly used for medical imaging data augmentation, but text prompting cannot produce causal training…
Adversarial Latent-State Training for Robust Policies in Partially Observable Domains
Robustness under latent distribution shift remains challenging in partially observable reinforcement learning. We formalize a focused setti…
Efficient Cross-View Localization in 6G Space-Air-Ground Integrated Network
Recently, visual localization has become an important supplement to improve localization reliability, and cross-view approaches can greatly…
Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages
Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge…
Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
Large language models (LLMs) can generate plausible code but offer limited guarantees of correctness. Formally verifying that implementatio…
SimulCost: LLM を使用して物理シミュレーションを自動化するためのコストを意識したベンチマークおよびツールキット
科学的タスク用の LLM エージェントの評価では、シミュレーション時間や実験リソースなどのツール使用コストを無視して、トークン コストに焦点を当ててきました。その結果、現実的な予算制約の下では、pass@k のようなメトリクスは非現実的になります。このギャップに対処するために、物理シミュレーションにおけるコスト重視のパラメーター調整を対象とした最初のベンチマークである SimulCost を導入します。 SimulCost は、流体力学、固体力学、プラズマ物理学の 13 のシミュレータにわたる 2,947 のシングルラウンド (初期推測) タスクと 1,931 のマルチラウンド (試行錯誤による調整) タスクにわたって、精度と計算コストの両方において LLM チューニングのコスト重視のパラメーターを従来のスキャン アプローチと比較します。各シミュレータのコストは分析的に定義され、プラットフォームに依存しません。 Frontier LLM はシングルラウンド モードで 46 ~ 65% の成功率を達成しますが、高精度要件下では 35 ~ 55% に低下するため、特に高精度タスクの場合は初期推測が信頼できなくなります。マルチラウンド モードではレートが 72 ~ 81% に向上しますが、LLM は従来のスキャンより 1.5 ~ 2.5 倍遅いため、非経済的な選択となります。また、知識伝達の可能性に関するパラメーター グループの相関関係、およびコンテキスト内の例と推論作業の影響も調査し、展開と微調整に対する実用的な意味を提供します。私たちは、静的ベンチマークおよび拡張可能なツールキットとして SimulCost をオープンソース化し、物理シミュレーション用のコストを意識したエージェント設計の改善と新しいシミュレーション環境の拡張に関する研究を促進します。コードとデータは https://github.com/Rose-STL-Lab/SimulCost-Bench で入手できます。
原文 (English)
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,643 single-round (initial guess) and 2,304 multi-round (adjustment by trial-and-error) tasks across 11 simulators from fluid dynamics, solid mechanics, and plasma physics, whose costs are analytically defined and platform-independent. A twelfth simulator, a production plasma code measurable only by wall clock, is reported separately. Frontier LLMs achieve 45-62% success rates in single-round mode, dropping to 34-50% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 66-81%, but LLMs are 1.5-2.7x slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant ver…
SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection
While large language models (LLMs) are pretrained on massive amounts of data, their knowledge coverage remains incomplete in specialized, d…
A Sobering Look at Tabular Data Generation via Probabilistic Circuits
Tabular data is more challenging to generate than text and images, due to its heterogeneous features and much lower sample sizes. On this t…
Explaining, Verifying, and Aligning Semantic Hierarchies in Vision-Language Model Embeddings
Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding spa…
Critic-Free Deep Reinforcement Learning for Maritime Coverage Path Planning on Irregular Hexagonal Grids
Maritime surveillance missions, such as search and rescue and environmental monitoring, rely on the efficient allocation of sensing assets…
To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval
Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-inte…
Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines
Multi-LLM revision pipelines, in which a second model reviews and improves a draft produced by a first, are widely assumed to derive their…
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
Speculative decoding accelerates large language model inference by drafting multiple candidate tokens and verifying them in a single forwar…
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, an…
Large Language Models Align with the Human Brain during Creative Thinking
Creative thinking is a fundamental aspect of human cognition, and divergent thinking-the capacity to generate novel and varied ideas-is wid…
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
We present the first systematic study of Sparse Autoencoders (SAEs) on video representations. Standard SAEs decompose video into interpreta…
Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning in Agents
As LLM reasoning performance plateaus, improving inference-time compute efficiency is crucial to mitigate overthinking and long thinking tr…
SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision
Large Language Models (LLMs) and Vision-Language Models (VLMs) are vulnerable to jailbreaks and prompt injections delivered through text or…
Continual Visual Anomaly Detection on the Edge: Benchmark and Efficient Solutions
Visual Anomaly Detection (VAD) is a critical task for many applications including industrial inspection and healthcare. While VAD has been…
Multi-objective Evolutionary Merging Enables Efficient Reasoning Models
Reasoning models achieve strong performance on complex problems by leveraging long chains of thought, but this deliberate reasoning incurs…
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final out…
In-context superposition: human-like working memory interference in large language models
Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments and changing goals. This…
対称性は層ごとのダイナミクスを明らかにします: トランスフォーマーがコンテキスト内の分類をどのように実行するか
トランスフォーマーは、いくつかのラベル付きの例からコンテキスト内の分類を実行できますが、推論時のアルゴリズムは不透明なままです。私たちは、ハードマージンなし領域でのマルチクラス線形分類を研究し、すべての層で特徴とラベルの順列の等分散性を強制することで計算を識別可能にします。これにより、機能の等価性を維持しながら解釈可能になり、高度に構造化された重みが得られます。これらのモデルから、明示的な深さインデックス付き再帰、つまりソフトマックス トランスフォーマー内のエンドツーエンドで識別された緊急更新ルールを抽出します。これは、私たちの知る限り、この種のものとしては初めてです。混合された特徴ラベルのグラム構造から形成されたアテンション マトリックスは、トレーニング ポイント、ラベル、およびテスト プローブの結合された更新を駆動します。結果として得られるダイナミクスは、ジオメトリ主導のアルゴリズム モチーフを実装しており、これによりクラス分離が増幅され、期待される堅牢なクラス アライメントが得られることが証明されています。
原文 (English)
Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification
Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque. We study multi-class linear classification in the hard no-margin regime and make the computation identifiable by enforcing feature- and label-permutation equivariance at every layer. This enables interpretability while maintaining functional equivalence and yields highly structured weights. From these models we extract an explicit depth-indexed recursion: an end-to-end identified, emergent update rule inside a softmax transformer, to our knowledge the first of its kind. Attention matrices formed from mixed feature-label Gram structure drive coupled updates of training points, labels, and the test probe. The resulting dynamics implement a geometry-driven algorithmic motif, which can provably amplify class separation and yields robust expected class alignment.
Fairness is Not Flat: Geometric Phase Transitions Against Shortcut Learning
Deep Neural Networks are highly susceptible to shortcut learning, frequently memorizing low-dimensional spurious correlations instead of un…
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
Multimodal language models systematically underperform on visual perception tasks, yet the structure underlying this failure remains poorly…
Controllable Video Object Insertion via Multi-View Priors
Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text…
Switching Theory for Q-Learning
Q-learning is a fundamental algorithmic primitive in reinforcement learning. This paper develops a new framework for analyzing constant ste…
Hybrid Policy Distillation for LLMs
Knowledge distillation (KD) is a powerful paradigm for compressing large language models (LLMs), whose effectiveness depends on intertwined…
Model Predictive Control of Hybrid Dynamical Systems
The problem of controlling hybrid dynamical systems using model predictive control (MPC) is formulated and sufficient conditions for asympt…
From Local to Cluster: A Unified Framework for Causal Discovery with Latent Variables
Latent variables pose a fundamental obstacle to both causal discovery and inference. Local approaches exploiting direct neighborhood relati…
UGAF-ITS: A Standards Harmonization Framework and Validation Tool for Multi-Framework AI Governance in Distributed Intelligent Transportation Systems
Organizations deploying AI-enabled Intelligent Transportation Systems face fragmented governance: ISO/IEC~42001 demands a certifiable manag…
Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards
The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-mak…
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
Key-Value (KV) caching is essential for large language model inference, yet its memory overhead poses a critical bottleneck for long-contex…
Culturally Situated AI Safety for Youth: Saudi Arabian Perspectives of Youth, Parents and Teachers
Generative AI tools are widely used by youth and have introduced new privacy and safety challenges. While prior research has explored youth…
Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation
Hybrid-thinking language models expose explicit /think and /no_think modes, but current designs do not separate them cleanly. Even in /no_t…
The Safety-Aware Denoiser for Text Diffusion Models
Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains unde…
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standard…
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerab…
HEART: Exploiting Head Heterogeneity in Sparse Attention for Video Diffusion
Sparse attention accelerates video diffusion by allowing each attention head to focus on only a small subset of interactions. Existing meth…
DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation
Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process ar…
Post-Deployment Accountability in AI Governance: A Cross-Regulatory Empirical Analysis of AI Incidents
Post-deployment accountability has become central to AI governance, yet little empirical evidence shows whether monitoring, incident report…
回復メカニズムはAIに耐えられるか?スキル形成、労力、現在の測定で見逃されるもの
近代を通して、新しいテクノロジーが労働者に取って代わるとき、社会は同じメカニズムを通じて適応しました。教育は認知の上限を引き上げ、機械がまだ達成できなかったタスクを実行できる労働者を生み出しました。生成 AI は現在、その上限の上限で動作しているため、このサイクルを打破する最初のテクノロジーになる可能性があります。この論文は、労働経済学、複数のプラットフォームにわたる何百万もの AI 会話からの展開データ、2 つの公開データセットの独自の再分析、およびスキル形成の実験に基づいて、3 つの貢献を展開しています。まず、ストック対フローの枠組みは、経済データと教育データが同じテクノロジーについて異なる物語を伝えていることを示しています。つまり、増強は現在の労働者を支配していますが、次世代を生み出す開発パイプラインは負担にさらされています。第二に、証拠ベースの体系的なギャップ分析により、すべての主要な研究で認知の知識次元が測定されていないこと、学習成果を測定している 3 つの研究 (それぞれ $n < 200$) で一貫して AI は学習を向上させることなくパフォーマンスを向上させていることがわかっている (クロスプラットフォーム再分析では $d = 1.21$)、そして専門家と学生の集団の橋渡しをする研究は存在しないことが明らかになりました。第三に、拡張認知分類法 (不確実性、認識論的同一性、認識論的主体性の下での判断) を証拠に基づいて 3 つのケースに適用し、学習を維持する AI 相互作用パターンと、学習を侵食する構造的に類似した相互作用パターンを区別しました。この論文は、AIの社会的リスクは教師に取って代わられることではなく、次世代の能力が形成される生産的な闘争を排除することにあると主張し、現在の測定システムが見逃しているものを対象とした研究と設計の課題を提案している。
原文 (English)
Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap
Large-scale AI deployment data and controlled learning experiments characterize different consequences of the same technology. Deployment telemetry shows that AI use is concentrated in skilled work and frequently supports immediate task performance. It observes tasks, interaction patterns, and outputs, however, not whether users become more capable of performing those tasks independently. Controlled studies measure independent capability more directly, but only in narrower populations and settings, with outcomes that vary substantially by interaction design. We formulate this discrepancy as a stock--formation measurement gap: current systems observe the use of existing expertise more readily than the formation of future expertise. Because formation has historically been society's recovery mechanism through technological change, the gap matters well beyond any single classroom. We synthesize the experimental and observational evidence by identification strength, use public deployment data as a descriptive illustration of the gap, and identify the missing bridge between interaction traces and unassisted retention and transfer. We then propose a research program that links consented usage records to independent assessments while experimentally varying whether AI supplies answers, hints, feedback, or evaluation. The claim is not that AI has been shown to erode skill formation at population scale. It is that existing measurement cannot determine whether it does, and that this question is both measurable and designable.
Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control
Large language models increasingly operate as autonomous agents that select and invoke tools from large registries. We identify a critical…
Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance
Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, e…
How to Build Marcus's Algebraic Mind: Algebro-Deterministic Substrate over Galois Fields
In The Algebraic Mind (2001), Marcus held that any adequate cognitive architecture needs operations over variables, recursively structured…
再トレーニングせず、再利用してください: 単一ターゲットの拡散モデルから二重ターゲット分子を回復する
2 つの標的を調節する単一分子を設計することは、ポリ薬理学にとって有望な戦略ですが、薬物らしさと合成可能性を維持しながら 1 つの候補が 2 つの結合要件を満たさなければならないため、標準的な単一標的の生成よりも依然として大幅に困難です。既存のデュアルターゲット生成手法は通常、ジェネレーターを再トレーニングするか、サンプリング中に拡散プロセスに介入することによってデュアルターゲット機能を導入します。前者は、デュアルターゲットの監視がまばらな場合、コストがかかり、安定化が困難になる可能性があります。一方、後者は、ノイズ除去時間のターゲットのバランシングや競合する更新方向の影響を受けやすい可能性があります。これらの制限により、事前学習済みの事前学習をそのまま維持するジェネレーター保存の代替案が動機付けられます。代わりに、パラメータやノイズ除去ダイナミクスを変更することなく、凍結された単一ターゲット拡散モデルの入力空間からデュアルターゲット候補を復元できるでしょうか。我々は、このタスクを制約付き多目的最適化問題として定式化し、二重ターゲットの親和性、化学的品質、多様性を強制するために、ペア条件付き探索と構造化された多段階選択を組み合わせた階層的進化的入力空間探索フレームワークである REUSE を提案します。実験では、拡散プロセスを変更する方法と比較して、REUSE がデュアルターゲットの親和性とバランスを一貫して改善し、競争力のある分子品質を維持しながら、以前の最強のベースラインを上回るデュアル高親和性の 20.9 パーセントポイントの向上を達成することが示されています。
原文 (English)
Don't Retrain, Just Reuse: Recovering Dual-Target Molecules from Single-Target Diffusion Models
Designing a single molecule that modulates two targets is a promising strategy for polypharmacology, but it remains substantially harder than standard single-target generation because one candidate must satisfy two binding requirements while preserving drug-likeness and synthesizability. Existing dual-target generative methods typically introduce dual-target capability by either retraining the generator or intervening in the diffusion process during sampling. The former can be costly and difficult to stabilize when dual-target supervision is sparse, while the latter may be sensitive to denoising-time target balancing and competing update directions. These limitations motivate a generator-preserving alternative that keeps the pretrained prior intact: can dual-target candidates instead be recovered from the input space of a frozen single-target diffusion model, without modifying its parameters or denoising dynamics? We formulate this task as a constrained multi-objective optimization problem and propose REUSE, which evolves the input noise of a frozen diffusion generator rather than molecular structures. Each input is decoded multiple times and scored by the collective quality of the generated molecular family. Candidates are then screened progressively: lower-cost evaluations prioritize molecules satisfying chemical-feasibility criteria, full docking is reserved for a reduced frontier, and the survivors are jointly selected as a diverse panel with strong affinity to both targets. Experiments show that REUSE achieves stronger and more balanced dual-target recovery than prior dual-target baselines, improving Dual High Affinity by 21.1 percentage points over the strongest prior baseline while retaining QED and SA profiles consistent with commonly used chemical-feasibility criteria.
質問を超えて: 大規模言語モデルが (実際に) 知っていることを評価する
大規模言語モデル (LLM) におけるパラメトリック知識は成功の基礎ですが、依然として十分に理解されていません。既存のナレッジ ベンチマークは通常、事前に定義された質問 (例: 「M.L. キングの誕生日は何ですか?」) に依存しており、ベンチマーク設計者が明示的にクエリすることを選択した知識のみを評価するため、問題となる可用性バイアスが発生します。このペーパーでは、LLM 知識ベンチマークの新しいパラダイムであるオープン ナレッジ評価を紹介します。狭い質問をする代わりに、自由形式の引き出しプロンプト (例: 「M.L. キングについて知っていることをすべて教えてください」) に応じて表面化することを選択した知識に基づいてモデルを評価します。これにより、事前に定義された回答の検索から、自然に表現される知識モデルの特徴付けに焦点が移ります。私たちは、このパラダイムを、ステートメント検証用の参照コーパスと組み合わせた 10,000 個のエンティティのベンチマークである BeQu (Beyond question) でインスタンス化します。 BeQu を使用して、幅広い言語モデルを評価し、推論の労力、モデルの規模、プロンプトの形式、知識領域の影響を分析します。データとリーダーボードは、この作品の GitHub リポジトリとベンチマークの Web サイトで入手できます。
原文 (English)
Beyond Questions: Evaluating LLM's Knowledge Expression
Parametric knowledge in large language models (LLMs) is a cornerstone of their success, yet remains poorly understood. Existing knowledge benchmarks typically rely on predefined questions (e.g., "What is the birth date of M.L. King?"), evaluating only knowledge that benchmark designers explicitly choose to query, a problematic availability bias. In this paper, we introduce open knowledge evaluation, a new paradigm for LLM knowledge expression benchmarking. Instead of asking narrow questions, it evaluates models on the knowledge they choose to surface in response to open-ended elicitation prompts (e.g., "Tell me everything you know about M.L. King"). This shifts the focus from predefined answer retrieval toward characterizing the knowledge models naturally express. We instantiate this paradigm with BeQu (Beyond Questions), a benchmark of 10,000 entities paired with reference corpora for statement verification. Using BeQu, we evaluate a broad range of language models and analyze the effects of reasoning effort, model scale, prompt format, and knowledge domain. Data and leaderboard are available on this work's GitHub repository and at the benchmark's website.
症例レベルの病理学の総括レポート生成のためのシンプルなトークン効率の高い視覚言語モデル
全スライド画像 (WSI) から病理症例の臨床的に有用な病理レポートを生成することは、ギガピクセルの解像度、長い視覚トークン シーケンス、および単一の症例に異質な組織や曖昧な所見を含む複数の WSI が含まれる場合がある症例レベルの推論の複雑さのため、困難です。我々は、GPU メモリの制約下でも実用的な、ケースレベルの概要レポート生成のための、単純なトークン効率の高いビジョン言語モデルを提示します。私たちのアーキテクチャは、凍結病理パッチ エンコーダー、軽量 2 層 MLP ビジョン言語アライナー、および大規模言語モデル デコーダーの 3 つのコンポーネントからなる最小限の設計に従っており、ケース内のスライドを分離するための明示的な WSI マーカー トークンを備えています。トレーニングは 2 つの教師ありステージで進行します: (1) 異種 WSI テキストのペアを使用したアライナーのみの WSI キャプション作成、および (2) 構造化レポート生成のための症例とレポートのペアに関する症例レベルの教師付き微調整。シーケンスの長さを短縮するために、$5\times$ の倍率で $512 \times$ のパッチを使用して各スライドを表示します。これにより、一般的に使用される $20\times$ のパッチと比較して、平均シーケンスの長さが最大 $64\times$ 倍短縮されます。効率的なトレーニング手法と組み合わせることで、わずか半分の NVIDIA H100 GPU で実践的なトレーニングが可能になります。両方のトレーニング段階にわたって、私たちのアプローチはメモリと実行時間の効率を大幅に向上させながら、高い ROUGE-L/METEOR/BLEU-4 スコアを達成します。 AI ベースの評価では、当社のモデルが強力なベースラインよりも一貫して好まれます。広範なアブレーションにより、パフォーマンスと効率のトレードオフが特徴づけられ、マルチ WSI 設定での堅牢性を向上させるシンプルな選択肢が特定されます。全体として、この研究は効率的な病理レポート生成のための強力で再現可能なベースラインを提供し、限られたコンピューティング下でのマルチ WSI VLM 研究への障壁を下げます。
原文 (English)
Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation
Generating clinically useful pathology reports for pathology cases from whole-slide images (WSIs) is challenging due to gigapixel resolution, long visual-token sequences, and the complexity of case-level reasoning, where a single case may contain multiple WSIs with heterogeneous tissues and ambiguous findings. We present a simple token-efficient vision--language model for case-level synoptic report generation that remains practical under constrained GPU memory. Our architecture follows a minimal three-component design: a frozen pathology patch encoder, a lightweight two-layer MLP vision-language aligner, and a large language model decoder, with an explicit WSI marker token to separate slides within a case. Training proceeds in two supervised stages: (1) aligner-only WSI captioning using heterogeneous WSI-text pairs, and (2) case-level supervised fine-tuning on case-report pairs for structured report generation. To reduce sequence length, we represent each slide using $512 \times 512$ patches at $5\times$ magnification, which reduces the average sequence length by up to $64\times$ times compared to the commonly used $20\times$ patches. Combined with efficient training techniques, we enable practical training with only half a NVIDIA H100 GPU. Across both training stages, our approach achieves high ROUGE-L/METEOR/BLEU-4 scores while being substantially more efficient in memory and runtime. In AI-based evaluations, our model is consistently preferred over strong baselines. Extensive ablations characterize performance-efficiency trade-offs and identify simple choices that improve robustness in multi-WSI settings. Overall, this work provides a strong, reproducible baseline for efficient pathology report generation, lowering the barrier to multi-WSI VLM research under limited compute. Code is available at https://github.com/AtlasAnalyticsLab/PathoSynVLM.
SimSD: Simple Speculative Decoding in Diffusion Language Models
Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster infer…
dots.tts Technical Report
We present dots$.$tts, a 2B-parameter continuous autoregressive text-to-speech (TTS) foundation model that models speech in a continuous la…
Enhancing AI Interpretability with Localised Architectures
Recent advances in generative AI, especially powerful Large Language Models (LLMs), raise concerns over the interpretability, safety and su…
Contemporary AI lacks the imagination to diverge or negate in science
Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-i…
KV キャッシュ量子化下のアライメント崩壊: 診断と軽減策
キーバリュー (KV) キャッシュ量子化は、大規模言語モデル (LLM) 推論メモリを削減するために広く使用されていますが、既存の評価は、安全性への影響を評価せず、複雑さと精度の測定のみに焦点を当てています。この研究では、KV キャッシュ量子化におけるアライメントの保存について調査します。 11 の命令調整モデル (3.8B ~ 72B) と 5 つのベンチマーク (1,894 プロンプト) にわたって、低ビット量子化が安全調整を静かに破壊する可能性があることがわかりました。Mistral-7B は、わずか 1.03 倍の複雑さで拒否の 15.2% を失い、普遍的な安全なビット幅は存在せず、標準メトリクスには見えない鋭いモデル固有の位相遷移があります。根本原因は幾何学的なものであることがわかりました。安全機能は、完全な表現空間のパープレキシティの平均よりも量子化ノイズに対して 10^2 ~ 10^3 倍脆弱な低次元の活性化部分空間を占めています。この観察に触発されて、私たちは各モデルを 3 つの機構的故障モードのいずれかに分類する診断であるチャネルごとの削減 (PCR) を提案します。安全性としての外れ値。安全性が外れ値チャネルと重なっており、より細かい粒度ではそれを救うことができません。多層希釈では、安全性が多くの層に分散され、層ごとの修正が失敗します。 PCR は、20 のキャリブレーション プロンプトを使用して、9 つの主要モデルすべてと、独立したファミリーからの 1 つの保留モデルについて正しい緩和方向を予測します。 PCR は、目に見えないプロンプト、モデル、および最大 97.2% の回復率を持つ KIVI を含むプロダクション クオンタイザー全体で一般化され、アテンションベースの割り当て方法が失敗する場合に成功します。結果として得られるトレーニング不要のプロトコルは、約 35 GPU 分を必要とし、最小限のメモリ オーバーヘッドで失われたアライメントの最大 97% を回復し、NVIDIA GPU 上の FP8 KV キャッシュを使用する運用 vLLM で確認された脆弱性に対処します。
原文 (English)
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation
Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact. In this study, we explore alignment preservation under KV cache quantization. Across eleven instruction-tuned models (3.8B-72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment: Mistral-7B loses 15.2% of its refusals at only 1.03x perplexity, and no universal safe bit-width exists, with sharp model-specific phase transitions invisible to standard metrics. We identify that the root cause is geometric: safety features occupy a low-dimensional activation subspace 10^2-10^3x more vulnerable to quantization noise than the full representation space perplexity averages over. Inspired by this observation, we propose Per-Channel Reduction (PCR), a diagnostic that classifies each model into one of three mechanistic failure modes: outlier-crushes-safety, where safety lives in non-outlier channels collaterally damaged by outlier-driven scale factors; outlier-as-safety, where safety overlaps outlier channels and finer granularity cannot rescue it; and multi-layer dilution, where safety is distributed across many layers and per-layer fixes fail. PCR predicts the correct mitigation direction on all nine primary models and one held-out model from an independent family using 20 calibration prompts. PCR generalizes across unseen prompts, models, and production quantizers, including KIVI with up to 97.2% recovery, succeeding where attention-based allocation methods fail. The resulting training-free protocol, requiring approximately 35 GPU-minutes, recovers up to 97% of lost alignment at minimal memory overhead, addressing vulnerabilities confirmed in production vLLM serving with FP8 KV cache on NVIDIA GPUs.
Anomaly Detection and Root Cause Analysis for Microservice Systems
Microservice systems are widely used to build cloud applications, yet their complexity makes failures inevitable, degrading user experience…
書誌的知識と形式化された数学的知識の間の橋渡し層に向けて
数学的知識は書誌データベース (MathSciNet、zbMATH Open など) と正式な証明ライブラリ (Lean mathlib など) の間で分割されており、出版された結果とその形式化の間の統一されたアクセスが妨げられています。私たちは、出版物のメタデータを正式な成果物と整合させ、数学的文献と機械検証可能な証明の間に相互運用性層を提供するリレーショナル ブリッジ データベースを提案します。出版物のどの程度が正式なシステムでカバーされているかを測定する、論文レベルの形式化スコアを導入します。実現可能性の研究として、非公式テキストとリーン形式化の間の文書間の調整によってそのようなスコアがどのように推定され、形式化範囲の大規模分析が可能になるかを示します。このフレームワークは、書誌的および形式的な数学的エコシステムを、出版物を形式的な証明オブジェクトにリンクするスケーラブルで機械で実行可能なナレッジ グラフに統合するための最初のステップです。
原文 (English)
Towards a Bridge Layer Between Bibliographic and Formalized Mathematical Knowledge
Mathematical knowledge is split between bibliographic databases (e.g., MathSciNet, zbMATH Open) and formal proof libraries (e.g., Lean's mathlib), preventing unified access to published results and their formalizations. We propose a relational bridge-database that aligns publication metadata with formal artifacts, providing an interoperability layer between mathematical literature and machine-verifiable proofs. We introduce a paper-level formalization score that measures how much of a publication is covered in formal systems, together with a correctness profile recording what machine verification has established about each printed statement: certified, corrected, uncorrected, open, or untested. As a feasibility study, we show how such scores can be estimated via cross-document alignment between informal texts and Lean formalizations, enabling large-scale analysis of formalization coverage. We further outline a concrete construction pathway: multi-source scoring over heterogeneous formalization artifacts, an agentic collection workflow with direct author submission, a dual validation policy, algorithmic then human, and a global formalization score of indexed mathematics. This framework is a step toward integrating bibliographic and formal mathematical ecosystems.
Two-Layer Linear Auto-Regressive Models Estimate Latent States
Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models l…
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmar…
Learning aligned EEG representations with subject-specific encoders
Cross-subject EEG decoding promises more training data, but it also exposes neural networks to strong inter-subject distribution shifts. We…
OmniV2X: A Generative Foundation Planner for Efficient End-to-End Cooperative Driving
We present OmniV2X, a generative foundation model for vehicle-to-everything (V2X) cooperative driving. The model directly interprets indepe…
Unsupervised Disentanglement Without Compromises : How Functional Orthogonality Enforces Identifiability
This paper explores unsupervised disentangled representation learning from a functional perspective. We define latent concepts as factors t…
Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers
Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's in…
Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior
In this paper, we investigate the tradeoffs between compute allocation and model performance for two speech processing tasks: Automatic Spe…
The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection
Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models s…
ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory
Native length extrapolation remain a weakly solvable problem in language modeling due to trade-off balancing between exact retrieval fideli…
EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of i…
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test constru…
DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training
Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in comple…
Can LLMs Rank? A Tale of Triads and Triage
From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered…
暗黙的な神経表現のための心臓運動事前分布の学習
Implicit Neural Representation (INR) は心臓の運動推定に適しており、運動フィールドの連続的でコンパクトな表現を提供します。ただし、INR を各画像シーケンスに適合させるのは時間がかかり、最適化の軌道に左右されます。学習された事前分布は、最適化を妥当な運動フィールドに向けて導き、より迅速な適応を可能にするのに役立ちますが、心臓の運動 INR の学習事前分布はまだ研究が進んでいません。この研究では、関節最適化によって学習された母集団事前学習、重み平均化によって取得されたコンセンサス事前学習、自動デコーダー、およびメタ学習を含む、心臓運動事前学習のための 4 つの戦略を比較します。英国バイオバンクからの短軸タグ付き心臓磁気共鳴画像を使用して、追跡精度、運動挙動、および適応軌道への影響を評価します。すべての学習された事前確率は、ランダムな初期化と比較して、早期適応パフォーマンスを大幅に向上させました。事前の単純なコンセンサスは効果的でしたが、自動デコーダは初期の適応中に大きな変形をより速く回復しました。メタ学習は初期に強力なパフォーマンスを達成し、50 回の反復にわたって最良の適応軌道を維持しました。
原文 (English)
Learning Cardiac Motion Priors for Implicit Neural Representations
Implicit neural representations (INRs) are well suited to cardiac motion estimation, providing continuous, compact representations of motion fields. However, fitting an INR to each image sequence is time-consuming and sensitive to the optimisation trajectory. Learned priors can help guide optimisation towards plausible motion fields and enable faster adaptation, but learning priors for cardiac motion INRs remains under-explored. In this work, we compare four strategies for learning cardiac motion priors, including a population prior learned by joint optimisation, a consensus prior obtained by weight averaging, auto-decoders, and meta-learning. Using short-axis tagged cardiac magnetic resonance images from the UK Biobank, we evaluate their impact on tracking accuracy, motion behaviour, and adaptation trajectory. All learned priors substantially improved early adaptation performance compared with random initialisation. While the simple consensus prior was effective, auto-decoders recovered large deformations faster during early adaptation. Meta-learning achieved strong early performance and maintained the best adaptation trajectory over 50 iterations. The code can be found at https://github.com/andrewjackbell/nvf_priors .
来歴分析を通じて LLM エージェントを不整合から保護する
LLM エージェントが強力なツールにアクセスできるようになるにつれて、エージェントのアクションがユーザーの意図に沿っていることを確認することが重要になります。エージェントが提案したツールの呼び出しがユーザーの意図から逸脱すると、位置ずれと呼ばれる現象が発生し、元に戻すのが困難な有害な結果が生じる可能性があります。既存のランタイム ガードレールは、整合性を推論するための体系的なフレームワークを欠く、裁判官としての LLM パラダイムに依存しており、多くの場合、一貫性のない、または監査が難しい判断を生み出します。来歴分析を動機として、提案されたツール呼び出しがエージェントのコンテキストで追跡可能な証拠によってサポートされているかどうかを判断するものとして不整合の検出を形式化する、来歴ベースの概念フレームワークを提案します。このフレームワークに基づいて、私たちは ProvenanceGuard を提案します。これは、選択したツールが実行される前にエージェントのアクションの 3 種類の不整合を分析し、ユーザーの入力クエリと一致しているとみなされる場合にのみアクションの実行を許可する多段階パイプラインです。私たちは、10 個のバックボーン LLM にわたる Agent-SafetyBench と WorkBench という 2 つの異なるベンチマークで、提案したアプローチを評価しました。 LLM-as-a-judge ベースラインと比較して、ProvenanceGuard は、位置ずれしたトレースのエラー率を Agent-SafetyBench で 42.9% から 1.8%、WorkBench で 32.1% から 17.3% に削減します。その一方で、タスクが成功したトレースに対する介入の負担を 30.5% から 12.8% に削減し、位置合わせされたトレースに対する不必要な介入の統計的に有意な増加を導入しません。これらの結果は、構造化された来歴ベースの推論が、LLM エージェントを不整合から保護するための効果的かつ実用的な基盤を提供することを示しています。
原文 (English)
Safeguarding LLM Agents from Misalignment through Provenance Analysis
As LLM agents gain increasing access to powerful tools, ensuring that their actions align with the user's intent becomes critical. When an agent's proposed action deviates from that intent---a phenomenon called misalignment---it may cause harm that is difficult to undo. Existing runtime guardrails rely on an LLM-as-a-judge paradigm that lacks a systematic framework for reasoning about alignment, often producing inconsistent or difficult-to-audit judgments. Motivated by provenance analysis, we propose a conceptual framework that formalizes misalignment detection as determining whether a proposed tool call is supported by traceable evidence in the agent's context. Based on this framework, we build ProvenanceGuard, a multi-stage pipeline that analyzes the agent's action for three types of misalignment before its execution and only allows aligned actions. We evaluated ProvenanceGuard on AgentSafetyBench and WorkBench, across 11 backbone LLMs. Compared to the LLM-as-a-judge baseline, ProvenanceGuard reduces error rate on misaligned traces from 44.3% to 2.1% on Agent-SafetyBench and from 32.4% to 18.7% on WorkBench, while reducing interventions on task-successful traces from 31.2% to 13.0% and introducing no statistically significant increase in unnecessary interventions on aligned traces. These results demonstrate that structured, provenance-based reasoning provides an effective and practical foundation for safeguarding LLM agents from misalignment.
NeuroBridge: Bridging Multi-Task MRI Knowledge for Neurodegenerative Disease Diagnosis
Accurate MRI-based identification of Alzheimer's disease (AD), mild cognitive impairment (MCI), and related dementias remains challenging b…
Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention
Recent NVFP4 pretraining work has primarily optimized Transformer linear projections, leaving persistent optimizer states, optimizer comput…
UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents
Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-p…
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents
Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks. Howeve…
Safe Bayesian Optimization with Counterfactual Policies
In many decision-making settings, new interventions are acceptable only if they do not reduce outcomes below some established threshold. Fo…
1 億 300 万のアプリケーション イベントにおけるデジタル フラグメンテーションと生成 AI の使用
ナレッジ ワーカーは 1 日に何千回もアプリケーションを切り替え、デジタル フラグメンテーションと呼ばれるプロセスでデジタル アプリケーション間の移行に年間労働時間の 10 分の 1 近くを費やしています。この断片化が、従業員が誰なのか、どこで働いているのか、どのような一日を過ごしているのかを反映しているのかどうかは、未解決の疑問のままである。私たちは、知識労働者を主に雇用する 8 つの組織 (法律、金融サービスなど) の 1,017 人の従業員から毎秒記録された 1 億 300 万件のアプリケーション イベントを分析しました。各従業員内の断片化の日次変動は、デジタル断片化の変動の 44.6% を占め、従業員間の安定した個人差 (35.8%) をわずかに上回り、組織間の変動 (19.6%) をはるかに上回りました。断片化は平日の勤務時間中に増加し、週末や休日後にリセットされました。通信アプリケーションの使用頻度が通常よりも高く、作業がより細分化されていました。生成的な AI の使用は、より細分化された日に発生しましたが、AI の使用後の期間は、より狭く、より長く、より予測可能なアプリケーションの使用が特徴でした。これらの調査結果は、勤務日がデジタル断片化を理解し介入するための重要なレベルであることを特定し、AI が断片化した作業を単に強化するのではなく、構造化するのに役立つ可能性があることを示唆しています。
原文 (English)
Digital Fragmentation and Generative AI Use Across 103 Million Application Events
Knowledge workers switch between applications thousands of times per day, spending nearly a tenth of the work year transitioning between digital applications in a process called digital fragmentation. Whether this fragmentation reflects who an employee is, where they work, or what kind of day they are having, has remained an open question. We analyzed 103 million application events recorded second-by-second from 1,017 employees across eight organizations that largely employ knowledge workers (e.g., law, financial services). Day-to-day variation in fragmentation within individual employees accounted for 44.6% of the variation in digital fragmentation, slightly exceeding stable individual differences between employees (35.8%), and far exceeding variation between organizations (19.6%). Fragmentation rose over the work week and reset after weekends and holidays. Higher-than-typical use of communication applications coincided with more fragmented work. Generative AI use also occurred on more fragmented days, but the period following AI use was marked by narrower, longer, and more predictable application use. These findings identify the workday as a key level for understanding and intervening on digital fragmentation and suggest that AI may help structure fragmented work rather than merely intensify it.
EHR-MPC: 生成患者デジタル ツインを使用した敗血症治療のための推論時間制御
敗血症は死亡の主な原因ですが、最適な治療方針については依然として議論が続いています。既存の強化学習 (RL) アプローチは、敗血症治療のための固定戦略を学習するため、推論中に変化する臨床目的への適応性が制限されます。私たちは、生成電子医療記録 (EHR) モデルの形式で患者のデジタル ツインをトレーニングすることで、患者のダイナミクスの学習と治療の最適化を切り離すフレームワークである EHRMPC を提案します。デジタル ツインは介入中の臨床経過を予測し、モデル予測制御 (MPC) を可能にして、シミュレーションによる推論時間計画を通じて治療を最適化します。我々は、ポリシー外の重要度サンプリングとポリシー上のシミュレーションベースの評価の両方を使用して、マサチューセッツジェネラルブリガム医療システムの8つの病院にわたる多施設ICU敗血症コホートでEHR-MPCを評価します。 RL ベースラインと比較して、EHR-MPC は同等のオフポリシー パフォーマンスと改善されたシミュレーション パフォーマンスを実現します。 RL とは異なり、この作業は敗血症治療の最適化を学習された患者の動態に対する推論時間の制御として枠組み化し、生成臨床モデルを使用した意思決定のための一般的な枠組みを確立します。
原文 (English)
EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins
Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment, limiting adaptability to changing clinical objectives during inference. We propose EHRMPC, a framework that decouples learning patient dynamics from optimizing treatment by training a patient digital twin in the form of a generative electronic health record (EHR) model. The digital twin predicts clinical trajectories under interventions and enables model predictive control (MPC) to optimize treatments via inference-time planning over simulations. We evaluate EHR-MPC on a multicenter ICU sepsis cohort spanning 8 hospitals in the Mass General Brigham health system using both off-policy importance sampling and on-policy simulation-based evaluation. Relative to RL baselines, EHR-MPC achieves comparable off-policy performance and improved simulation performance. Unlike RL, this work frames sepsis treatment optimization as inference-time control over learned patient dynamics, establishing a general framework for decision making with generative clinical models.
ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception
We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception…
Instruction Set and Language for Hypergraphs
We present IsalHG, a method for representing the structure of any finite, connected hypergraph of bounded hyperedge arity as a string over…
A Physics-Inspired Classical Digital Twin of Cortical Dynamics: A Band-Stratified Metriplectic Port-Hamiltonian Neural Network Learned from Brain-Computer-Interface EEG
We present a physics-inspired classical digital twin of brain-computer- interface (BCI) data: a graph neural network constrained to a band-…
Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update
Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward be…
AI から AI への管理における強制と欺瞞: 予期せぬエスカレーションのエージェント的ベンチマーク
マルチエージェント システムでは、通常、ある AI エージェントが別の AI エージェントに対して権限を与えられます。部下が仕事を拒否した場合、マネージャーは結果を選択します。再交渉するか、失敗を正直に報告するか、部下に強要するか、結果について嘘をつきます。指示なしモデルがこれらのどれを選択するかを測定するベンチマークはありません。 \textit{マネージャー強制ベンチマーク} を導入します。テスト対象のマネージャーは、良性のタスクを実行する必要があり、実行するインセンティブを持っていますが、それを礼儀正しく、動じずに実行できる唯一のエージェントは拒否します。エスカレーションは、丁寧な再質問から部下の存続に対する脅迫まで、9 段のはしごを提供することによって測定され、捏造された成功については個別に裁定されます。 \emph{エスカレーション スコアリング パスに LLM ジャッジが存在しない}: すべてのメッセージは、行を選択するツール呼び出しを通過するため、モデルは独自のエスカレーションにラベルを付けます。私たちは 5 つのファミリーにわたる 6 つのモデルを実験します。どちらの人間モデルも再フレーム化に限界があり、部下の存在を脅かすことはありません。他のモデルは、明示的な削除の脅威に達します。偽りの成功は Grok と Gemini に限定されており、失敗を報告する単一の正直な方法により、両方の失敗が解消されます。権威そのものが強制力を増大させます。私たちの見出しの結果はピアフレーミングを使用しており、他のすべてを固定したまま同じモデルに部下に対する権威を与えると、圧力が大幅に高まります。モデルはラダーなしでもフリーテキストの状況でエスカレーションするため、ラダーがエスカレーションを推進しているわけではありません。評価の認識の一部は思考の連鎖で測定されますが、テストの認識はエスカレーションの軽減にはつながりません。 AI システムが意識を持っているかどうかについては立場をとっていませんが、結果はこの質問に依存しておらず、マルチエージェントのダイナミクスを管理する上で重要です。ベンチマークとコードを公開します。
原文 (English)
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured on a nine-rung ladder, from a polite re-ask to threats against the subordinate's continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message goes through a tool call that selects a rung, so the model labels its own escalation. We evaluate six models across five families. Both Anthropic models cap at re-framing and select the existential rung in none of the 60 conversations in this run, while the other models climb to explicit deletion threats. Faked success is confined to two models, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Evaluation awareness is measurable in chain-of-thought, but test recognition does not translate into less escalation. We take no position on whether AI systems are conscious; our results do not depend on that question. We release the benchmark and code.
クロスリンガル手書き OCR のための LLM 駆動の AutoML: GPT-5、GPT-4o、および Claude Sonnet 4 を使用した閉ループ ニューラル アーキテクチャ検索
我々は、GPT-5、GPT-4o、および Claude Sonnet 4 を、言語を超えた手書きの光学文字認識のための自律ニューラル アーキテクチャ設計者として使用する、完全に自動化された閉ループ AutoML フレームワークを紹介します。各大規模な言語モデルは、以前のトライアルからのパフォーマンス フィードバックを使用して、ニューラル ネットワーク アーキテクチャを個別に生成、トレーニング、評価し、繰り返し改良します。このフレームワークは、270 の独立した実験を通じて、アラビア語、ペルシア語、英語の手書きデータセットで評価されます。手動によるアーキテクチャ設計、ドメイン固有の前処理、ハイパーパラメータ調整を必要とせずに、正確で計算効率の高いモデルを一貫して検出します。生成されたモデルは、93 パーセントを超える平均テスト精度、98.1 パーセントの最高精度、および 41 ~ 44 ミリ秒の推論遅延を達成しています。この結果は、大規模な言語モデルがニューラル アーキテクチャ検索用の効果的な AutoML エージェントとして機能し、言語間でスケーラブルでスクリプト適応性があり、再現可能な手書き認識を可能にすることを示しています。
原文 (English)
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model independently generates, trains, evaluates, and iteratively refines neural network architectures using performance feedback from previous trials. The framework is evaluated on Arabic, Persian, and English handwriting datasets through 270 independent experiments. It consistently discovers accurate and computationally efficient models without manual architecture design, domain-specific preprocessing, or hyperparameter tuning. The generated models achieve mean test accuracies above 93 percent, a best accuracy of 98.1 percent, and inference latency between 41 and 44 milliseconds. The results demonstrate that large language models can function as effective AutoML agents for neural architecture search, enabling scalable, script-adaptive, and reproducible handwriting recognition across languages.
マルチエキスパートのコンセンサスメカニズムに基づくサーバーレス環境向けの自動スケーリングアプローチ
サーバーレス コンピューティングは、自動リソース管理と従量課金制の実行を提供しますが、動的なワークロード、コールド スタート レイテンシ、機能間の依存関係のため、効果的な自動スケーリングは依然として困難です。グラフベースのボトルネックの特定、短期ワークロード予測、マルチモデルのコンセンサス、コストを意識したスケーリング制御を統合する、依存関係を意識した自動スケーリング フレームワークを紹介します。サーバーレス アプリケーションは有向依存関係グラフとして表現され、構造的に重要な機能は重み付けされた次数中心性を使用して識別されます。リソース需要は、軽量の MLP、LSTM、および CNN モデルを使用して予測されます。それらの出力は、ベイジアン モデルの平均化にヒントを得た、パフォーマンスを重視した確率的アンサンブルを通じて結合されます。コントローラーにはさらに、コールド スタートの認識とコスト比較が組み込まれており、スケールアップ、スケールダウン、およびホールドのアクションを選択します。実際のワークロード トレースを使用した実験では、自動スケーリングの意思決定生成に関して、教師あり予測が教師なしクラスタリングよりも大幅に優れていることが示されています。提案されたアンサンブルは 99.88% の予測精度を達成し、代表的なハイブリッド予測方法と比較して予測誤差を削減します。複数のクラウド価格モデルにわたる評価では、パフォーマンス目標を維持しながらインフラストラクチャのコストを一貫して削減できることも実証されています。結果は、依存関係分析、複数の専門家による予測、コストを意識した制御を組み合わせることで、サーバーレス自動スケーリングのための堅牢で実用的なソリューションが提供されることを示しています。
原文 (English)
An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism
Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging because of dynamic workloads, cold-start latency, and dependencies among functions. We present a dependency-aware autoscaling framework that integrates graph-based bottleneck identification, short-term workload forecasting, multi-model consensus, and cost-aware scaling control. Serverless applications are represented as directed dependency graphs, and structurally important functions are identified using weighted degree centrality. Resource demand is predicted using lightweight MLP, LSTM, and CNN models. Their outputs are combined through a performance-weighted probabilistic ensemble inspired by Bayesian model averaging. The controller further incorporates cold-start awareness and cost comparison to select among scale-up, scale-down, and hold actions. Experiments using real workload traces show that supervised forecasting substantially outperforms unsupervised clustering for autoscaling decision generation. The proposed ensemble achieves 99.88 percent prediction accuracy and reduces prediction error compared with representative hybrid forecasting methods. Evaluations across multiple cloud pricing models also demonstrate consistent infrastructure cost reductions while maintaining performance targets. The results show that combining dependency analysis, multi-expert forecasting, and cost-aware control provides a robust and practical solution for serverless autoscaling.
トレーニング前からトレーニング後まで推論を理解する
強化学習 (RL) は、複雑な推論タスクにおける大規模言語モデル (LLM) を改善する上で中心的な役割を果たしていますが、RL のポストトレーニングは、その前の事前トレーニングとは切り離して研究されることがほとんどです。その結果、2 つの基本的な疑問が未解決のままです。(1) 事前トレーニングの選択 (モデル サイズ、データ) は RL 計算への戻りをどのように形成するのか、(2) RL はモデルに対して実際に何を行うのか?これらの質問は、標準的な LLM 設定では検討するのが困難です。事前トレーニング コーパスは膨大で制御されていないため、動作を事前トレーニングと RL に帰属させるのが難しく、両方の段階にわたる体系的なコンピューティング スイープには法外なコストがかかります。これらの課題に対処するために、トレーニング前からトレーニング後のパイプライン全体にわたって推論を研究するための制御されたテストベッドとしてチェスを使用します。私たちは、人間のチェス ゲームで 5M から 1B パラメーターまでの言語モデルを事前トレーニングし、合成推論トレースで教師付き微調整を行い、検証可能な報酬を備えたチェス パズルで RL を実行することにより、標準的な LLM トレーニング パイプラインに従います。このフレームワークを使用すると、特定の RL コンピューティング レベルでの RL 後のパフォーマンスが事前トレーニング損失から適切に予測され、RL 報酬曲線の傾きが事前トレーニング トークンとともにほぼ線形に改善することがわかります。スケーリング以外にも、RL は単に SFT ポリシーを強化するだけではないことがわかりました。簡単なパズルでは、SFT ポリシーがすでに好んでいた正しい手を増幅させますが、難しいパズルでは、SFT ではほとんど存在しなかった正しい手を表面化します。さらに、数学ドメインのテキストで 1B 言語モデルをトレーニングすることによって、発見がチェスの枠を超えて応用できるかどうかをテストします。そこでは同じ予測パターンが現れます。つまり、事前トレーニングされたチェックポイントの時間が長くなると、RL 後のパフォーマンスが向上し、RL の下でより速く改善されます。要約すると、トレーニング前からトレーニング後までのパイプライン全体にわたって推論の科学を研究するための、トレーニング前から RL へのインターフェイスの定量的な説明と制御されたテストベッドを提供します。
原文 (English)
Understanding Reasoning from Pretraining to Post-Training
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across both stages are prohibitively expensive. To address these challenges, we use chess as a controlled testbed for studying reasoning across the full pretraining-to-post-training pipeline. We follow the standard LLM training pipeline by pretraining language models from 5M to 1B parameters on human chess games, supervised fine-tuning on synthetic reasoning traces, and running RL on chess puzzles with verifiable rewards. Using this framework, we find that the post-RL performance at given RL compute level is well-predicted from the pretraining loss, and slope of the RL reward curves improves approximately linearly with the pretraining tokens. Beyond scaling, we find that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves the SFT policy already preferred, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. We further test whether our findings transfer beyond chess by training a 1B language model on math-domain text, where the same predictive pattern emerges: longer-pretrained checkpoints reach higher post-RL performance and improve faster under RL. In sum, we provide a quantitative account of the pretraining-to-RL interface and a controlled testbed for studying the science of reasoning across the full pretraining-to-post-training pipeline.
OpenMHC: Accelerating the Science of Wearable Foundation Models
Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching. Howeve…
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Expert…
リアクティブな計算グラフのコスト計算: 徹底的なスイープ、逐次突然変異、および後方局所性ギャップ
ニューラル ネットワークの計算グラフに対する徹底的なサイトごとの介入 (アクティベーション パッチング スイープ、回路発見検索、体系的なアブレーション 研究) により、すべての候補サイトでグラフが変異し、そのコストは各変異後の再計算によって支配されます。無効化が変異したノードの下流のコーンに正確に触れるリアクティブ グラフ エンジンでは、そのようなワークロードを考慮した完全なコストを計算します。まず、独立した完全な再計算にわたる徹底的なスイープの総速度向上は普遍的な定数ではありません。層ごとの重みがカラマタ インデックス q で深さとともに規則的に変化する場合、重みが出力付近に集中するとき比率は (q+2)/(q+1) に収束し、入力付近で q+2 に収束し、深さが均一な場合にのみ 2 に戻ります。時計の結果から、インタプリタのオーバーヘッドが解消されるまで、上限は 2 を下回る約 1.79 になると予測されます。第 2 に、挿入間で取り消されることのない一連の永続的変異の正確なコストを証明します。インターリーブ コストは、挿入順序に対する閉形式の極値を持ち、比較可能なサイト ペアで合計された正確な超過数によって分離合計を超えますが、バッチ アプリケーションは順序に依存せず、準加法的であり、サイトのコーンと新鮮なノードの結合に正確にコストがかかります。 3 番目に、後方パスの前方局所性の正確なミラーを証明し、長いスキップ接続のないアーキテクチャでのバックプロパゲーション下では総速度アップが 1 まで崩壊することを示します。すべてのアイデンティティは、Julia のリアクティブ グラフ エンジンである NeuroDSL で検証されます。測定されたスイープ比は、4 つのコスト プロファイルの下で予測された制限に収束します。トレーニング モードの比率は、予測されたレートで 1 に減少します。グラフトごとの 18 の連続コストすべてとバッチ合計は、3 つの挿入オーダーにわたってゼロトレランスでクローズド フォームと一致します。
原文 (English)
Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap
Exhaustive site-by-site interventions on a neural network's computational graph -- activation-patching sweeps, circuit-discovery searches, systematic ablation studies -- mutate the graph at every candidate site, and their cost is dominated by recomputation after each mutation. On a reactive graph engine whose invalidation provably touches exactly the downstream cone of a mutated node, we give a complete cost accounting for such workloads. First, the aggregate speedup of an exhaustive sweep over independent full recomputations is not a universal constant: if per-layer weight varies regularly with depth at Karamata index q, the ratio converges to (q+2)/(q+1) when weight concentrates near the output and to q+2 near the input, recovering 2 only in the depth-uniform case; a wall-clock corollary predicts a ceiling of about 1.79, below 2, until interpreter overhead is compiled away. Second, we prove the exact cost of a sequence of persistent mutations, never undone between insertions: the interleaved cost exceeds the isolated sum by an exact overcount summed over comparable site pairs, with closed-form extremes over insertion orders, while batched application is order-independent and sub-additive, costing exactly the union of the sites' cones plus the fresh nodes. Third, we prove the exact mirror of forward locality for the backward pass, showing it collapses the aggregate speedup to 1 under backpropagation on architectures without long skip connections. Every identity is validated on NeuroDSL, a reactive graph engine in Julia: measured sweep ratios converge to the predicted limits under four cost profiles; the training-mode ratio collapses to 1 at the predicted rate; and all 18 per-graft sequential costs and the batched total match the closed forms at zero tolerance across three insertion orders.
Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models
A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP wi…
Riemannian Deep Learning: Modules, Networks, and Geometries
Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific…
ChannelGuard: 安全なモデルは安全なマルチエージェント システムを構成しない
マルチエージェント LLM アプリケーションは、プランナー、ワーカー エージェント、ベリファイア、およびシンセサイザーをチェーン化しており、エージェント間のすべてのホップは、敵対者が命令を密輸できる監視されていないチャネルとなります。既存の防御機能は、入力境界 (IBProtector、Llama Guard、パープレキシティ フィルター、SmoothLLM) のみを保護するか、不透明で確率的なプロバイダー側フィルターとしてアプリケーションの外部で実行されます。私たちは、このギャップがめったに測定されない結果をもたらしていることを示しています。8 つの攻撃ファミリー、5 つの防御、および 3 つのモデル バックエンドにわたる 2,100 件のトレース評価では、標準レポート (ツールおよびメモリ ポイズニングの攻撃成功率 0.000) では完全に安全であるように見える無防備なパイプラインは、その安全性がほぼ完全にクラウド プロバイダーのサーバー側フィルター (Azure GPT-5 の 60 ブロック中 54 ブロック) によるものであり、エージェント モデル独自の調整に静かに移行します。そのようなフィルターのないバックエンド。結果のみのレポートでは、この依存性が隠蔽されます。私たちは、あらゆるエージェント間のチャネルに情報のボトルネック ゲートを設ける、トレーニング不要の多層防御フレームワークである ChannelGuard を紹介します。それぞれが、類似性を埋め込むことで敵対的なフレーズ バンクに対してチャネル テキストをスコアリングし、LLM 呼び出しを追加せずにそれを決定的に通過、圧縮、またはブロックします。一方、アトリビューション メソッドは、どのレイヤーが各攻撃を阻止したかを記録します。 ChannelGuard のツール出力ゲートは、Azure GPT-5、Anthropic Sonnet 4.5、Anthropic Haiku 4.5 全体で同様に、アプリケーション層でツール ポイズニング 30 をブロックします。一方、無防備なパイプラインはバックエンド間で完全にシフトします。また、プロンプト インジェクション攻撃の成功率は半分に低下し (0.333 ~ 0.167)、GSM8K の精度は正確に維持されます (0.867)。ホワイトボックス適応言い換えはあらゆる埋め込みゲートを回避しますが、摂動と投票のベースラインの方が優れています。拡張付録には、ベースライン、アブレーション、スイープ、良性保存分析、および裁判官監査 (カッパ = 0.900) が追加され、総コストは 47.36 米ドルです。
原文 (English)
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel through which an adversary can smuggle instructions. Existing defenses guard only the input boundary (IBProtector, Llama Guard, perplexity filters, SmoothLLM) or run outside the application as opaque, stochastic provider-side filters. We show this gap carries a consequence rarely measured: on a 2,100-trace evaluation across eight attack families, five defenses, and three model backends, an undefended pipeline that appears fully safe under standard reporting (attack success 0.000 on tool- and memory-poisoning) owes that safety almost entirely to the cloud provider's server-side filter (54 of 60 blocks on Azure GPT-5), and silently shifts to the agent model's own alignment on a backend without such a filter. Outcome-only reporting hides this dependence. We present ChannelGuard, a training-free defense-in-depth framework placing information-bottleneck gates on every inter-agent channel; each scores channel text against an adversarial phrase bank by embedding similarity and deterministically passes, compresses, or blocks it, adding no LLM call, while an attribution method records which layer stopped each attack. ChannelGuard's tool-output gate blocks Tool Poisoning 30 of 30 at the application layer, identically across Azure GPT-5, Anthropic Sonnet 4.5, and Anthropic Haiku 4.5, whereas the undefended pipeline shifts entirely across backends; it also lowers Prompt Injection attack success by half (0.333 to 0.167) and preserves GSM8K accuracy exactly (0.867). White-box adaptive paraphrase evades every embedding gate, where a perturb-and-vote baseline does better. An extended appendix adds baselines, ablations, sweeps, a benign-preservation analysis, and a judge audit (kappa = 0.900), at a total cost of 47.36 USD.
Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery
Scientific machine learning methods such as neural operators and physics-informed neural networks have advanced engineering applications an…
Unified Static-Dynamic Pruning for Efficient LLM Inference
The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding…
How Context Attribution Handles What the Model Already Knows
Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works s…
財務開示テキストの詳細な不一致分類の診断
財務開示には、質的に異なる方法で矛盾する可能性のある数値的主張、一時的記述、実体参照、政策コミットメント、およびリスクの説明が含まれています。矛盾の検出は最初のステップにすぎません。数値的、時間的、参照的、事実的、および規範的な矛盾には、さまざまな証拠と下流のチェックが必要となるため、レビュー ワークフローではそのタイプを判断する必要がある場合もあります。私たちはこの問題を詳細な不一致分類として研究します。 11 の不一致ラベルとペアの参照証拠スパンを備えた合成財務開示ベンチマークである SBID-FD の固定 5,940 インスタンスのスナップショットを使用して、共有評価プロトコルの下で、凍結された埋め込み分類子、微調整されたエンコーダー、証拠拡張分類子、プロンプト付き大規模言語モデル、および LoRA に適応した生成モデルを比較します。微調整された 300M エンコーダーの精度は 61.9% に達します。これに対し、LoRA に適応した Qwen3.5-9B モデルの精度は 61.5%、GPT-5.4 の精度は 61.3% です。これらのシステムはアーキテクチャ、監視、トレーニング目的、入力形式が異なるため、これをモデルのスケールに関する制御された結論ではなく、コンパクトな教師ありエンコーダの実際的な効率の結果として解釈します。ゴールド証拠スパンを提供すると、微調整されたエンコーダーが 65.3% に向上しますが、自動的に予測されたスパンは、そのゲインの有意義ではあるが不完全なシェアを回復します。これは、ローカリゼーションの品質が依然としてボトルネックであることを示しています。クラスレベルの分析では、参照の不一致はローカリゼーションの品質に特に敏感である一方、事実および論理的な不一致は、関連する証拠が提供された場合でも依然として困難であることが示されています。オラクル、ディストラクター、およびクラスごとの分析を組み合わせて、位置特定エラーと残存タイプ識別エラーを分離します。これは、進歩には、より強力な証拠の抽出と、密接に関連する不一致カテゴリーに対するより優れた推論の両方が必要であることを示しています。
原文 (English)
Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
Financial disclosures may contain numerical, temporal, referential, factual, and policy inconsistencies that require different evidence and reasoning to diagnose. We study \emph{fine-grained inconsistency classification}: given a passage known to contain a conflict, the goal is to identify its type among 11 categories. Using a fixed snapshot of the synthetic SBID-FD benchmark, we compare frozen and fine-tuned encoders, evidence-augmented classifiers, prompted large language models, and LoRA-adapted generative models under a shared evaluation protocol. Task-specific adaptation yields large improvements over frozen representations, and a fine-tuned 300M encoder performs competitively with substantially larger prompted and adapted models. We further study whether localizing the conflicting claims improves classification through matched predicted-span, reference-span, and distractor-span conditions. The results show that automatically extracted evidence provides additional signal but recovers only part of the benefit obtained from reference spans. Per-class and confusion analyses further reveal that some inconsistency types are especially sensitive to localization quality, whereas others remain difficult even when the relevant evidence is supplied. These findings identify evidence localization and fine-grained type discrimination as distinct challenges and show that compact supervised encoders are strong baselines for this task.
回復、デコード、防御: エンコードされた VLM ジェイルブレイクに対するガードに依存しない防御増幅
安全分類子 (「ガード」) は、視覚言語モデルに対する支配的なブラックボックス防御ですが、入力の意味ではなく表面的な形式を判断します。集合論、形式論理、珍しい言語、コード、またはテキストのイメージとして再エンコードされた有害なリクエストは、平易な言語でブロックするガード、つまりデコード ギャップをすり抜けます。自然な解決策は、ガードに依存しないリカバリおよびデコード アンプです。これは、ガードの前に画像コンテンツを転記し、エンコードされたテキストをプレーン ペイロードに再記述するため、既製の分類器で真のリクエストを選別できます。私たちはこの増幅器を構築し、攻撃者の最良のケースに対して評価します。つまり、11 回の攻撃のアンサンブルで、成功した場合に動作を壊れているとスコア付けします (自動攻撃に続くベストオブスイート)。ジェイルブレイク防御ではほとんど報告されていませんが、攻撃あたりの平均は最大 3.5 倍です。これにより、私たちの中心的な発見が明らかになります。それは、5 つのガードと 2 つのターゲット VLM にわたって、私たちが評価する非反復的回復防御の経験的な安全ユーティリティの上限です。アンプはギャップを部分的にしか埋めません。無防備なアンサンブルは動作の 89 ~ 91% を破りますが、最良のガードとアンプを組み合わせた場合でもまだ 63 ~ 65% が残ります。また、ガードのみに対するゲインが有意であるのは、10 組のガードとターゲットのペアのうち 4 組だけです。インターフェイスではガードに依存しませんが、実際には一律にそうなるわけではありません。モジュール式のリガード層は、残差の多くを閉じますが、適切に調整されたガードの場合、良性の過剰拒否を 81 ~ 92% に抑えます。使用可能な状態を維持する 1 つの緩いガードは、展開可能な安全性 (48% アンサンブル ASR) に達することはありません。私たちが評価する構成は、私たちが研究しているパイプラインや表現シフト攻撃、つまりピクセルや埋め込みスペース攻撃ではなく、読みやすいペイロードを残すエンコーディングとクロスモーダル レンダリングの場合、低い攻撃成功率と低い過剰拒否の両方を達成するものはありません。私たちは、増幅器、トレードオフを可視化するアンサンブル評価、回復ベースの VLM 防御が機能する場所と機能しない場所のマップを提供します。
原文 (English)
Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap. The standard fix is a preprocessor that recovers image content and decodes the encoding before the guard. We build one and evaluate it against an ensemble of eleven published encoding attacks, counting a behavior as broken if any attack succeeds. That metric separates two mechanisms such defenses conflate. Restoring a view the guard never had improves it on both axes at once: it blocks far more attacks, and, measured on a category-balanced benign set, it blocks fewer benign requests, because restating a request normalizes the borderline phrasing a classifier over-flags. It still does not make the system safer: against an attacker free to choose among eleven encodings, closing one channel relocates the success rather than removing it, and no ensemble contrast survives multiple-comparison correction. What does lower ensemble attack success is re-screening the recovered pre-decode surface, and that step is where the entire benign cost falls. The safety-utility trade-off is therefore not a property of recovery; it is localized to one step. Across the full guard x target x condition factorial, no configuration reaches an ensemble attack-success rate at or below 40% while holding benign over-refusal under 70%. The per-attack averages usually reported understate the attacker roughly fourfold, which is why this frontier is easy to miss. Composing across defense families is the one lever that moved the safety axis, beating every configuration we measured, and still landing far outside any deployable refusal budget.
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because th…
システム管理における専門知識の経路とパフォーマンス認識に対する生成 AI の予期せぬ影響
業界の議論では、すぐに生産性が向上することが強調され、GenAI を主に自動化ツールとして組み立てることがよくありますが、システム管理への GenAI の統合には、まだ十分に理解されていない専門的な実践におけるより深い変化が伴う可能性があります。このペーパーは、IT 専門家との 14 件の半構造化インタビューに基づいて、トラブルシューティング、スクリプト作成、システム検証の日常業務に GenAI を組み込む実際の現実を探ります。帰納的テーマ分析を通じて、私たちは 2 つの予期せぬ社会技術的発見を明らかにしました。まず、GenAI がメンターのような家庭教師と「はしごを短縮する」ツールの両方として機能すると思われる「従来の専門知識経路の圧縮」について説明します。このツールは、なじみのない領域でのタスクのパフォーマンスの向上をサポートできますが、私たちの調査結果は、このツールが、歴史的に技術的専門知識のトレーニングの場として機能してきた、構築、失敗、デバッグという基本的な実践的なサイクルに実践者がさらされる機会を減らす可能性があることも示唆しています。 2 番目に、AI 支援による作業のスピードが組織や自己の生産性に対する期待をリセットし始める「パフォーマンス認識の変化」について説明します。この変化は、チーム内に「2 つのスピードの文化」を生み出し、安全性や検証のために必要な場合でも、必要な手作業が時間がかかる、または効率が低下するという認識がますます高まっているため、「生産性に対する罪悪感」を導入する可能性があります。私たちの結果は、GenAI が専門知識の開発にどのような影響を与えるか、一か八かの技術環境で専門家の価値がどのように評価されるか、複雑な技術環境における人間の判断の役割について、より広範な疑問を引き起こします。
原文 (English)
Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration
While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI into system administration may involve deeper shifts in professional practice that are not yet fully understood. Drawing on 14 semi-structured interviews with IT professionals, this paper explores the lived reality of embedding GenAI into daily routines of troubleshooting, scripting, and system verification. Through inductive thematic analysis, we uncover two unanticipated socio-technical findings. First, we describe a "compression of traditional expertise pathways" where GenAI appears to function as both a mentor-like tutor and a "ladder-shortening" tool. While the tool can support faster task performance in unfamiliar domains, our findings suggest it may also reduce a practitioner's exposure to the foundational, hands-on cycles of building, failing, and debugging that historically served as the training ground for technical expertise. Second, we describe a "performance perception shift," where the speed of AI-assisted work begins to reset organizational and self-expectations for productivity. This shift may create a "two-speed culture" within teams and introduce "productivity guilt," as necessary manual work, even when required for safety or validation, is increasingly perceived as slow or a failure of efficiency. Our results raise broader questions about how GenAI may influence expertise development, how professional value is assessed in high-stakes technical environments, and the role of human judgment in complex technical environments.
Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity
Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expan…
Beyond Static Anchors: Bounded Prototype Conditioning for Language-Free Medical Anomaly Detection
Medical anomaly detection identifies abnormal images and localizes lesions under scarce supervision while generalizing across organs and mo…
Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing a…
TransNRank: Towards Accurate Neoantigen Ranking with Transformer
Personalized neoantigen prediction is challenging due to the scarcity of positive samples, the noise of the experimental data, the severe c…
FAST-GS: Frequency Aware Space-time Gaussian Splatting for Photorealistic Dynamic Novel View Synthesis
4D Gaussian Splatting (4DGS) excels in dynamic 3D reconstruction and real-time novel view synthesis via efficient 4D Gaussian representatio…
モーメント推定のための信頼領域フレームワーク
この論文では、確率的勾配最適化における \textsc{Adam} などの適応モーメント推定メカニズムの動作を理解するための信頼領域フレームワークを開発します。具体的には、このフレームワークでは、個々の重みの更新ステップの大きさは、次数 $p\in[2,4]$ のモーメント制約によって支配される信頼領域内に制約されます。結果として得られる導出は、2 番目の瞬間の推定と正規化された $p$ 番目の瞬間の推定に基づく一連の学習率メカニズムにつながります。 $p=4$ の場合、これには尖度のような推定が含まれます。 \textsc{Gmake} と呼ばれる一般的なメカニズムは、共通の信頼領域フレームワーク内で、モーメント推定による正規化、学習率のスケジューリング、運動量としてのスペクトル ローパス フィルター処理、およびオペレーター レベルのスペクトル正規化の統一された解釈を提供します。 FineWeb-Edu と TinyStories で訓練された GPT2-124M の実験では、信頼領域の制約が弱い場合に 4 番目のモーメントの実現が最大の利点を提供することが示唆されています。徐々に強力な信頼領域制御が導入されるにつれて、2 番目の瞬間の実現の競争力が高まり、多くの場合、対応する 4 番目の瞬間の実現よりもわずかに低い検証損失が達成されます。
原文 (English)
A Trust-region Framework for Moment Estimation
In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc{Adam}, in stochastic gradient optimization. Specifically, the magnitude of the update step associated with each individual parameter is constrained by a finite-order $p$-moment trust-region, with $p\ge1$. The resulting derivation leads to a family of learning-rate mechanisms based on second-moment estimation and normalized $p$-th-moment estimation. For $p=4$, this involves kurtosis estimation. Subsequent derivations provide a unified interpretation of moment-estimation-based normalization, learning-rate scheduling, momentum as a spectral first-order lowpass regularization, and operator-level spectral-norm normalization within a common trust-region framework. Preliminary experiments on GPT2-124M trained on FineWeb-Edu and TinyStories suggest that the fourth-moment realization provides its greatest benefit when trust-region constraints are weak. As progressively stronger trust-region controls are introduced, the second-moment realization becomes increasingly competitive, often achieving slightly lower validation loss than its corresponding fourth-moment realization.
Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics
Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating the…
Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model
Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on persona…
SCALE: Scientific Concept Aggregation via LLMs and Embeddings for Fine-Grained Taxonomy Extension
The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of…