AIニュース 2026-07-16
自動生成: 2026-07-16 11:58 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
GPT-Red: Unlocking Self-Improvement for RobustnessOpenAI
Explore GPT-Red, OpenAI’s automated red teaming system that uses self…
-
The US is advancing AI safety through state and federal actionOpenAI
OpenAI outlines a “reverse federalism” approach to AI governance, whe…
-
Thinking Machines、初のAIモデル「Inkling」公開──オープンウェイトで「自分のものにできる」基盤モデルITmedia AI+
ミラ・ムラティ氏率いるThinking Machines Labが、同社初のAIモデル「Inkling」を発表した。テキスト・画像・音声対…
-
NVIDIAが「Jetson Thor」に新モジュール追加、高騰するメモリの使用量削減技術もITmedia AI+
NVIDIAは、組み込みAIボード「Jetsonシリーズ」の最新製品である「NVIDIA Jetson AGX Thor」の新たな量産モジ…
-
Claude、NotebookLM、Genspark……238社が選んだ「現場が支持する」AIサービスITmedia AI+
IT製品・サービスの選定から導入まで、企業はどのような視点で判断しているのか。238件の読者アンケートの回答から、製品選びの基準や導入時の…
-
SpaceX falls to $135 IPO price ahead of Starship launchTechCrunch AI
The stock has steadily fallen from the euphoric post-IPO high, showin…
-
Microsoft patches record number of security vulnerabilities, citing its use of AITechCrunch AI
Microsoft's monthly release of security fixes, dubbed Patch Tuesday,…
トピック別件数
- 研究/論文 114件
- LLM/生成AI 106件
- エージェント 66件
- 画像/動画生成 36件
- ビジネス/資金調達 19件
- ロボティクス 16件
- その他 9件
- ハードウェア/半導体 7件
- 規制/政策 3件
日本語メディア9件
ITmedia AI+ (日本語)
Thinking Machines、初のAIモデル「Inkling」公開──オープンウェイトで「自分のものにできる」基盤モデル
ミラ・ムラティ氏率いるThinking Machines Labが、同社初のAIモデル「Inkling」を発表した。テキスト・画像・音声対応のMoE型オープンウェイトモデルで、Hugging Faceで全ウェイトを公開。「最強のモデルではない」とした上で、同社のファインチューニ…
NVIDIAが「Jetson Thor」に新モジュール追加、高騰するメモリの使用量削減技術も
NVIDIAは、組み込みAIボード「Jetsonシリーズ」の最新製品である「NVIDIA Jetson AGX Thor」の新たな量産モジュールとして、消費電力や搭載メモリ容量などを抑えた「Jetson T3000」と「Jetson T2000」を追加すると発表した。
OpenAI、初のハードウェア「Codex Micro」を230ドルで発売 Apple提訴の渦中にある端末とは別物
OpenAIは、コーディング支援AI「Codex」向けの専用キーパッド「Codex Micro」を発売した。キーボードメーカーのWork Louderと共同開発した同社初のハードウェア製品で、価格は230ドル。なお、Appleによる営業秘密不正取得訴訟の渦中にある、開発中のAI…
Claude、NotebookLM、Genspark……238社が選んだ「現場が支持する」AIサービス
IT製品・サービスの選定から導入まで、企業はどのような視点で判断しているのか。238件の読者アンケートの回答から、製品選びの基準や導入時の課題、注目を集めるサービスの傾向を探る。
Anthropicと組んだNEC それでも森田社長が「4つの主権」にこだわる真意
米Anthropicとわずか3週間で電撃提携し、日本企業初のパートナーとなったNEC。最新AI「Mythos」が誇るバグ発見能力を巡り、サイバー悪用のリスクへ各国の懸念が高まる中、いかにして「実装」の壁を越え、自らを守るのか。森田隆之社長が語る「4つの主権(ソブリン)」の真意と…
ゼロから分かる「Claude」の教科書 ChatGPTと比べて分かった強みとは?
「ChatGPT」の陰に隠れがちでありながら、高い性能を持ち、AIユーザーから高い支持を集める「Claude」。最近では、その高い性能から話題になっている最新モデル「Claude Mythos」の登場でも大きな話題を呼んでいる。本ブックレットでは、このClaudeの実力を「Ch…
【一時非公開のお知らせ】「バズるほど赤字だった」──野田クリスタルのAIペットカードゲーム、公開停止からの復活劇を本人に聞いた
野田クリスタルさんがGeminiで開発した「ペットカードジェネレーター」は初日に130万アクセスを集めるも、AI利用料の高騰で一時停止に。「うちのこカード伝説」として復活するまでの舞台裏を、野田さんと電通の開発者に聞いた。
「GPT-5.6」の好きな部分をXに投稿→100ドル分のクレジットがもらえるキャンペーン 先着1万人限定
米OpenAIが、最新のAIモデル「GPT-5.6」の好きな部分や同モデルで作成したものをXに投稿すると、ChatGPTで100ドル分のクレジットを提供するキャンペーンを開催している。
孫正義氏が予想する「2040年」 “AI中心の社会”を生き抜く方法とは
2040年、AIはどれほど進歩し、企業を取り巻く状況はどう変化しているのか。ソフトバンクグループの孫正義代表取締役会長が語った。
海外メディア12件
TechCrunch AI (英語)
Microsoft is reportedly training salespeople to talk down OpenAI and Anthropic
Microsoft is looking to sell its in-house AI models as more efficient and cost-effective than its competitors' models.
SpaceX falls to $135 IPO price ahead of Starship launch
The stock has steadily fallen from the euphoric post-IPO high, showing that markets may be sobering up to the promises CEO Elon Musk made b…
Hack suggests AI music generator Suno scraped YouTube for training data
The hacker used an employee's credentials to access source code, which revealed how Suno scraped decades of audio.
Whatnot acquires Shaped to power real-time live shopping recommendations
Livestream shopping platform Whatnot has acquired AI startup Shaped, a machine learning company focused on real-time recommendations and se…
Microsoft patches record number of security vulnerabilities, citing its use of AI
Microsoft's monthly release of security fixes, dubbed Patch Tuesday, resolved a record 570 security vulnerabilities across the company's pr…
Apple Intelligence approved for launch in China with Alibaba’s Qwen AI
The deal, which was rumored to be in the works last year, marks an important step for Apple's AI ambitions in a key market.
Inside Ode with Anthropic, the startup betting AI services are the future of enterprise
Can a handful of engineers really do the work of an army of consultants? That’s the bet behind Ode with Anthropic — the joint venture dedic…
Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just models
Anthropic-backed Ode launches as AI labs bet that embedding forward-deployed engineers inside enterprises is the key to accelerating enterp…
Reelful’s AI turns your camera roll into short-form videos for social media
The app is designed for people who want to create social content, but find traditional video editing tools too complex or time-consuming.
Rime picks up $24M Series A to help enterprises field customer calls
Rime is handling over 100 million calls each month across multiple companies.
Indian AI coding startup Emergent becomes a unicorn with $130M Series C
The startup has reached a $120 million annualized revenue run rate and more than 200,000 paying customers.
Vint Cerf is working on a plan to unleash AI agents on the open internet
The guy behind TCP/IP is working on a standard for identifying AI agents in the wild.
公式ブログ2件
OpenAI (英語)
The US is advancing AI safety through state and federal action
OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.
GPT-Red: Unlocking Self-Improvement for Robustness
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
論文256件
arXiv cs.AI (英語)
最適で適応的なマーケットメイキング:永久先物市場における高利回り流動性供給のための理論的枠組み
私たちは、メーカー手数料ゼロの永久先物市場で最適なマーケットメイクを実現するための厳密な理論的フレームワークを開発します。マーケットメーカーの問題を、フィルタリングされた確率空間上の確率的最適制御問題としてモデル化します。ここでの制御は、2 つの取引所にわたる適応的な買値と売値のスプレッドと在庫ヘッジの決定です。当社の貢献には次のものが含まれます。(i) 収益をスプレッド収益、逆選択損失、在庫維持コスト、ヘッジ摩擦、および資金調達率エクスポージャーに分離する損益分解定理。 (ii) 検証定理を備えた CARA ユーティリティの下での共同スプレッド在庫ヘッジ制御問題のハミルトン・ヤコビ・ベルマン方程式。 (iii) マスター APY 式で頂点に達する、5 つの無次元パラメータを通じて収益性の高い地域を特徴付ける高 APY レジーム定理。 (iv) 最適な参入・退出閾値を備えた分散型永久取引所におけるゼロ手数料経済学の分析。 (v) 資金調達レートのダイナミクスとヘッジ制度の三分法による最適なクロス為替ヘッジ政策。 (vi) パラメータの不確実性許容度を定量化するロバストネスマージン。 (vii) 指数関数的ドローダウン確率限界と普遍的な APY-VaR アイデンティティ。 (viii) ベイジアン適応推定による最適制御下のエルゴーディックな在庫分布。 (ix) 遺跡境界を伴うケリー最適レバレッジ。 (x) 分散飽和の結果を伴うマルチペアのポートフォリオ配分。 23 個の数値を使用した数値分析により、収益性の高い体制と不採算な体制の間の相転移が明らかになります。私たちのフレームワークは、現代の分散型会場マイクロ構造向けに、アベジャネダ-ストイコフ、グアント-レハレ-フェルナンデス-タピア、グロステン-ミルグロムのパラダイムを統合し、拡張します。
原文 (English)
Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets
We develop a rigorous theoretical framework for optimal market making in perpetual futures markets with zero maker fees. We model the market maker's problem as a stochastic optimal control problem on a filtered probability space, where the controls are adaptive bid-ask spreads and inventory hedging decisions across two exchanges. Our contributions include: (i) a PnL decomposition theorem separating revenue into spread income, adverse selection loss, inventory carrying cost, hedging friction, and funding rate exposure; (ii) the Hamilton-Jacobi-Bellman equation for the joint spread-inventory-hedging control problem under CARA utility with a verification theorem; (iii) High-APY Regime Theorems characterizing profitable regions via five dimensionless parameters, culminating in a Master APY Formula; (iv) analysis of zero-fee economics on decentralized perpetual exchanges with optimal entry-exit thresholds; (v) optimal cross-exchange hedging policies with funding rate dynamics and a hedge regime trichotomy; (vi) a robustness margin quantifying parameter uncertainty tolerance; (vii) exponential drawdown probability bounds and a universal APY-VaR identity; (viii) ergodic inventory distribution under optimal control with Bayesian adaptive estimation; (ix) Kelly-optimal leverage with ruin boundaries; and (x) multi-pair portfolio allocation with diversification saturation results. Numerical analysis with twenty-three figures reveals phase transitions between profitable and unprofitable regimes. Our framework unifies and extends the Avellaneda-Stoikov, Gueant-Lehalle-Fernandez-Tapia, and Glosten-Milgrom paradigms for modern decentralized venue microstructure.
非定常性下でのコンテキスト内強化学習: 調査
意思決定事前トレーニング済みトランスフォーマー、アルゴリズム蒸留、ロングコンテキスト メタ RL、および検索拡張エージェントの開発により、インコンテキスト強化学習 (ICRL) に対する新たな関心が高まっています。ICRL とは、テスト時のパラメーターを更新することなく、インタラクション コンテキストから潜在的なタスク ルールを推論し、将来の動作を改善する、事前トレーニング済みまたは微調整された意思決定モデルの機能です。この一連の作業では、試行錯誤の証拠、報酬、遷移、デモンストレーション、フィードバック、または取得されたエクスペリエンスによって、コンテキスト ウィンドウ内で学習のような計算がいつ行われるかを問います。ただし、ICRL の既存の調査では、主に事前トレーニングの目的、アーキテクチャ、コンテキスト形式、評価プロトコル、理論的メカニズムを中心に分野が整理されている一方で、非定常設定については比較的十分に調査されていないままです。変化する環境では、蓄積されたコンテキストは固定タスクに関する単なる証拠ではありません。報酬の仕様、遷移カーネル、観察チャネル、アクション インターフェイス、制約モデル、またはデモンストレーションとメモリの分布が現在の体制と一致しなくなる可能性があります。したがって、以前は便利だったコンテキストが、古い体制に戻ると、古くなったり、誤解を招くものになったり、再び役に立ったりする可能性があります。展開されたポリシーパラメータが固定されたままである一方で、コンテキストを通じて適応するという問題として、非定常ICRLを調査します。ポリシーは、現在の決定ルールと、そのルールをサポートしている蓄積された証拠のどの部分の両方を推論する必要があります。私たちは非定常 ICRL を定義し、それをメタ RL、意思決定シーケンス モデリング、検索拡張 RL、価値とモデルを認識した ICRL、報酬フィードバック エージェントに関連付け、何が変化するか、変化がどのように展開するか、変化がエージェントにとってどの程度観察可能であるかという 3 つの質問に沿って文献を整理します。
原文 (English)
In-Context Reinforcement Learning under Non-Stationarity: A Survey
The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renewed interest in in-context reinforcement learning (ICRL): the ability of a pretrained or fine-tuned decision model to infer latent task rules and improve future behavior from interaction context, without test-time parameter updates. This line of work asks when trial-and-error evidence, rewards, transitions, demonstrations, feedback, or retrieved experience can make learning-like computation happen inside the context window. However, existing surveys of ICRL mainly organize the field around pretraining objectives, architectures, context formats, evaluation protocols, and theoretical mechanisms, while the non-stationary setting remains comparatively underexamined. In changing environments, accumulated context is not merely more evidence about a fixed task: the reward specification, transition kernel, observation channel, action interface, constraint model, or demonstration and memory distribution can fall out of alignment with the current regime. Previously useful context can therefore become stale, misleading, or useful again when an old regime returns. We survey non-stationary ICRL as the problem of adapting through context while deployed policy parameters remain fixed: the policy must infer both the current decision rule and which parts of its accumulated evidence still support that rule. We define non-stationary ICRL, relate it to meta-RL, decision sequence modeling, retrieval-augmented RL, value- and model-aware ICRL, and reward-feedback agents, and organize the literature along three questions: what changes, how the change unfolds, and how observable the change is to the agent.
ソブリンエンタープライズ言語モデルのオントロジー増幅蒸留とコンテキスト性監査: メカニズム証明と否定結果手法を組み合わせた研究
データ常駐ルールに基づいて運営されている規制対象の金融機関には、金融機関の境界内で実行できるテナント所有の言語モデルが必要です。この論文は、2 つの関連する FAOS 研究を 1 つのメカニズムと制御の記事にまとめたものです。まず、オントロジー増幅蒸留の低電力機構実証研究を報告します。Qwen3.6-27B の学生は、フロンティア教師の軌道に関する教師付き微調整と、オントロジーに基づいた直接選好最適化 (DPO) を通じて財団 AgenticOS オントロジーに適応し、47 の合成英語言語クロスドメイン選好ペアから単一の Apple M5 Max でローカルにトレーニングされます。 40 のベトナム金融ドメインの課題について、抽出された学生は 40 のタスクのうち 36 を根拠づけました (根拠率 0.90; 0.50 を下限としたメトリクスでの平均オントロジー用語カバレッジ r_onto = 0.95)。これは GPT-5 フロンティア ベースラインと同等であり、これも 40 のうち 36 を根拠としています。結果は等価性を確立するには力不足です: 対差分 95%信頼区間は +/-4 のタスクにまたがっており、この実行では、生徒がフロンティアを超えるはずであるという事前に登録された増幅予測をテストしたり示したりすることはありません。第 2 に、この論文では、エンタープライズ エージェント ルーティングのためのコンテキスト性監査手法が統合されています。別の否定的な結果のパイロットでは、ローカル Qwen 実行と明示的にラベル付けされた Gemma レプリケーション チェックの両方で、すべてのフェーズ 1.3 グループの修正された標準的なデフォルトによるコンテキスト性の度合いがゼロです。有用なシグナルは直接的な影響と構築結合であり、残存する文脈性ではありません。これらの研究では、オントロジーに基づいたモデル構築メカニズムと、明らかな不一致がいつ迅速な標準化、マルチエージェント統合、または人間によるレビューを引き起こすかを決定するためのガバナンス診断を組み合わせています。証拠は、展開可能性、安全性、優位性、統計的同等性、またはコンテキスト性を考慮したルーティング ルールをサポートしていません。
原文 (English)
Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study
Regulated financial institutions operating under data-residency rules need tenant-owned language models that can run inside the institution's perimeter. This paper combines two related FAOS studies into one mechanism-and-control article. First, it reports a reduced-power proof-of-mechanism study of ontology-amplified distillation: a Qwen3.6-27B student is adapted to the Foundation AgenticOS ontology through supervised fine-tuning on frontier-teacher trajectories and ontology-grounded direct preference optimization (DPO), trained locally on a single Apple M5 Max from 47 synthetic, English-language, cross-domain preference pairs. On 40 held-out Vietnamese financial-domain tasks, the distilled student grounds 36 of 40 tasks (grounded rate 0.90; mean ontology term-coverage r_onto = 0.95 on a metric floored at 0.50), equal to the GPT-5 frontier baseline, which also grounds 36 of 40. The outcome is underpowered to establish equivalence: the paired-difference 95% confidence interval spans +/-4 tasks, and the run does not test or show the pre-registered amplification prediction that the student should exceed the frontier. Second, the paper consolidates a contextuality-audit method for enterprise-agent routing. In a separate negative-results pilot, the corrected canonical Contextuality-by-Default degree is zero for all Phase 1.3 groups in both the local-Qwen run and an explicitly labeled Gemma replication check; the useful signal is direct influence and construct coupling, not surviving residual contextuality. Together, the studies pair an ontology-grounded model-building mechanism with a governance diagnostic for deciding when apparent disagreement should trigger prompt standardization, multi-agent synthesis, or human review. The evidence supports neither deployability, safety, superiority, statistical equivalence, nor a contextuality-positive routing rule.
GRID: エンタープライズ SQL 生成のための文法レール デコーディング
大規模な言語モデルでは SQL を記述することができますが、エンタープライズ展開では、もっともらしいテキスト以上のものが必要です。出力は構文的に有効である必要があり、ロールごとおよびスキーマごとのポリシーを尊重する必要があり、証明可能な (ベストエフォート型ではない) 保証を保持する必要があり、世代が成長するにつれて速度が低下してはならず、すべての決定についてコンプライアンス グレードの記録を残さなければなりません。我々は、トークンシーケンスではなくパーサー構成(レクサースキャン状態×LALR(1)スタック)で正確なネクストトークンマスクをキー設定する文法制約型デコードエンジンであるGRID(Grammar-Railed Decoding)を紹介し、段階的に高度なLALR(1)パーサー自体を実行可能なプレフィックスオラクルとして使用します。 LLM トークンは、コンテキスト独立/コンテキスト依存の分割を伴うバイトレベルのトライ ウォークによって文法端末にブリッジされ、キャッシュ キーの健全性が構築によって保持されます。役割ベースのアクセス制御は言語にコンパイルされます。役割投影は文法の生成のサブセットであり、スキーマ辞書は識別子の終端を制限するため、禁止された動詞と識別子はマスク レベルでは到達できません。 4 つの保証 (健全性、完全性、終了、トークンごとのほぼ一定のコスト) が明示的な前提条件とともに記載されており、それぞれがテストまたはベンチマークと対になっています。 Rust カーネルは、トークンごとのマスクを 3.6 ~ 6.7 us の中央値に設定し、誤拒否ゼロの 2 つのトークナイザーの p50 および p90 でのガイダンスを上回ります。トークンごとのガードコストは、n=16,000 でポジションフラットです。 Spider では、制約付きデコードは 0.5B で +13 実行精度ポイントの価値があり、マスク強制不可能であることが証明されている残留物 (列レベルのポリシー) に対するチェッカー ガイドに基づく修復パス 1 回により、7B モデルの実行率が 94.5% に引き上げられます。ハッシュチェーンされたトークンごとの監査証跡は、100% の改ざん検出でビット同一に再生されます。マスクで何ができないのか (配布の忠実性、列レベルの RBAC、非 LALR(1) 言語)、および測定されたコストがどこに残るのかを明確に述べます。
原文 (English)
GRID: Grammar-Railed Decoding for Enterprise SQL Generation
Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must respect per-role and per-schema policy, must carry provable (not best-effort) guarantees, must not slow down as generations grow, and must leave a compliance-grade record of every decision. We present GRID (Grammar-Railed Decoding), a grammar-constrained decoding engine that keys exact next-token masks on parser configurations (lexer scan state x LALR(1) stack) rather than on token sequences, and uses the incrementally advanced LALR(1) parser itself as a viable-prefix oracle. LLM tokens are bridged to grammar terminals by a byte-level trie walk with a context-independent/context-dependent split that makes cache-key soundness hold by construction. Role-based access control is compiled into the language: role projections subset the grammar's productions and schema lexicons restrict identifier terminals, so forbidden verbs and identifiers are unreachable at mask level. Four guarantees (soundness, completeness, termination, and near-constant per-token cost) are stated with explicit preconditions and each paired with a test or benchmark. Rust kernels bring the per-token mask to a 3.6-6.7 us median, ahead of llguidance at p50 and p90 on two tokenizers with zero false rejects; per-token guard cost is position-flat at n=16,000. On Spider, constrained decoding is worth +13 execution-accuracy points at 0.5B, and one checker-guided repair pass over the provably mask-unenforceable residue (column-level policy) lifts a 7B model to 94.5% executable. A hash-chained per-token audit trail replays bit-identically with 100% tamper detection. We state plainly what the mask cannot do (distribution faithfulness, column-level RBAC, non-LALR(1) languages) and where measured cost remains.
スマート温室における強化学習制御のためのキャリブレーションファースト報酬コンポーネント監査
温室強化学習は、作物の実験だけでは達成するのが難しい速度と規模で気候制御のアイデアをテストできます。しかし、スマート温室制御の場合、単一のシミュレータのリターンだけでは十分ではありません。栽培者や制御エンジニアは、政策が加熱、CO2 濃縮、通気、湿度管理、スクリーンの展開、またはランプの使用をいつ行うかを知る必要もあります。私たちは、シミュレータ トレーニング、施設に適応したロールアウト、ログに記録された自律型温室チャレンジ記録、およびアクチュエーター ルールの蒸留。 GreenLight-Gym では、フレームワークはスカラー報酬を条件付きの温度、CO2、湿度、蒸気圧不足、スクリーン、および作動代理項に分解します。 GreenLight を第 2 回自律温室チャレンジで記録された気候追跡に適応させます。そして、ログに記録された温室データの同じコンポーネントにスコアを付けます。
原文 (English)
Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses
Greenhouse reinforcement learning can test climate-control ideas at a speed and scale that is difficult to achieve with crop experiments alone. For smart-greenhouse control, however, a single simulator return is not enough: a grower or control engineer also needs to know when the policy heats, enriches CO2, vents, manages humidity, deploys screens, or uses lamps.We propose a reproducible calibration-first reward audit framework that keeps named greenhouse-control reward components comparable across simulator training, facility-adapted rollouts, logged Autonomous Greenhouse Challenge records, and actuator-rule distillation. In GreenLight-Gym, the framework decomposes the scalar reward into conditional temperature, CO2, humidity and vapor-pressure-deficit, screen, and actuation-proxy terms; adapts GreenLight to the second Autonomous Greenhouse Challenge logged climate traces; and scores the same components on logged greenhouse data.
必要なのは最適化だけではない
2019 年、OpenAI は、機械生成テキストの検出を支援するために、非文法的で半分壊れた 200 万個の GPT-2 出力をリリースしました。より流暢な後継者を生み出した連携は、通常、エンジニアリングの成果とみなされます。私たちはそれを、最適化文化の最新の表現として解釈します。つまり、テクノロジーよりも古い、事前に定義された軸に沿った測定可能な改善によって価値の問題が解決されるという信念です。その確信をスタック (事前トレーニング、デコード、プリファレンス調整、ベンチマーク、インターフェース) を通してたどり、監査協会の系譜をたどると、限界に到達します。最適化手順では、生成されたテキストの一部がどの程度ありそうもないかを測定できます。その可能性が誤りなのか発明なのかはわかりません。それにも関わらず、その区別ができない手順が、5 年以内に、正当な言語のプロトコルを設定する権限を引き継いだのです。何世紀にもわたってアカデミーや学校、文法学者や試験官によって保持されてきたこの権限は、損失関数、報酬モデル、ベンチマーク、およびシステムプロンプト、つまり判断能力のない判断官庁を実行する装置に譲渡されました。
原文 (English)
Optimization Is Not All You Need
In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering achievement; we read it instead as the newest expression of optimization culture: the conviction, older than the technology, that measurable improvement along predefined axes exhausts the question of value. Tracing that conviction through the stack-pretraining, decoding, preference tuning, benchmarking, interface-and back through its genealogy in the audit society, we arrive at the limit: an optimization procedure can measure how improbable a piece of generated text is; it cannot tell whether that unlikelihood is error or invention. A procedure that cannot make that distinction has nonetheless, within half a decade, assumed the authority to set the protocols of legitimate language. Held for centuries by academies and schoolrooms, grammars and examiners, this authority has been given over to loss functions, reward models, benchmarks, and system prompts: an apparatus that executes the office of judgment with no capacity for judging.
LP2Graph を使用した LP マイニング: 鉄道のスケジュール変更の使用例
多くの最適化主導の分野と同様に、鉄道の再スケジュールは混合整数線形計画法 (MILP) に依存していますが、この分野のモデリングの知識は互換性のない表記法で何百もの論文に散在しており、ナラティブ調査はそれを主観的に整理しています。つまり、構造ではなく語彙によってモデルを分類しており、どちらも再現していません。 LP2Graph を使用した LP マイニングを紹介します。これは、公開された LP および MILP 定式化の構造をマイニングして、再現可能なデータセットと誘導分類法を作成する方法です。そのコアである LP2Graph は、標準文法によって認められる各定式化を、型付き変数 (単一の標準モデルから派生した方程式グラフ) として表します。ソースがそのモデルに抽出されると、下流のすべてが決定的になります。各ソースはこのモデルに解析され、相同化され、ボトムアップで (変数、次に制約と目的、次にモデル全体の構造)、アプリケーション ドメインとソリューションのアプローチによって個別にクラスター化されます。結果として得られるグループは、ルールシードされた自己更新分類器によってラベル付けされます。私たちは表現を仮定するのではなく検証します。クラスターごとの代表は独立した LaTeX として再生成され、ソース論文で報告されている最適値に対して CBC、HiGHS、および Gurobi 全体で再解決されます。その成果は、変数、制約、モデル タイプの客観的で再現可能な分類法です。これは、自動化された鉄道スケジュール変更モデル開発の raiLPminer ラインが構築される原則的な基盤です。
原文 (English)
LP Mining with LP2Graph: A Use Case for Railway Rescheduling
Like many optimization-driven domains, railway rescheduling relies on Mixed-Integer Linear Programming (MILP), yet the field's modeling knowledge is scattered across hundreds of papers in incompatible notations, and narrative surveys organize it subjectively: they classify models by vocabulary rather than by structure, and reproduce neither. We present LP Mining with LP2Graph, a method that mines the structure of published LP and MILP formulations into a reproducible dataset and an induced taxonomy. Its core, LP2Graph, represents each formulation admitted by its canonical grammar as a typed variable--equation graph derived from a single canonical model; once a source is extracted into that model, everything downstream is deterministic. Each source is parsed into this model, homologized, and clustered bottom-up (over variables, then constraints and the objective, then whole-model structure) and, separately, by application domain and solution approach; the resulting groups are labeled by a rule-seeded, self-updating classifier. We validate the representation rather than assume it: per-cluster representatives are regenerated as independent LaTeX and re-solved across CBC, HiGHS and Gurobi against the optimum reported in the source paper. The outcome is an objective, repeatable taxonomy of variables, constraints and model types: the principled foundation on which our raiLPminer line of automated railway-rescheduling model development builds.
AI Web エージェント向けのエージェント対応 Web サイトの設計: 機械可読性、実行可能性、および意思決定の信頼性のためのフレームワーク
オンライン ショッピングは、AI エージェントが独自に商品を検索し、オプションを比較し、制約を評価し、ユーザーの購入プロセスの一部を実行するモデルにますます移行しています。 Web サイトのデザインは、人間とエージェントを介した対話の両方をサポートする必要があります。このペーパーでは、AI エージェント向けの電子商取引プラットフォームの可読性、解釈可能性、検証可能性、および実行可能性を強化するための設計フレームワークである、エージェント対応 Web サイトについて紹介します。既存の Web デザイン、SEO、生成エンジン最適化 (GEO) の指標では、エージェントを介したインタラクションに対する Web サイトの能力を完全には評価できません。提案されたフレームワークは、エージェントの解釈可能性、エージェントの実行可能性、および機械可読性、意味論的な明瞭さ、エージェントの動作可能性、およびコンテキスト上の決定の信頼性信号などの機能によってサポートされるエージェントの決定の信頼性の 3 つの次元を中心に構造化されています。このフレームワークは、人間指向のベースラインと、同一のカタログ、価格設定、在庫、ショッピングのワークフローを備えた同一の Web サイト プロトタイプのエージェント対応バージョンとを比較する対照実験を通じて評価されます。評価には 5 つのタスク、3 つのブラウザー エージェント モデル (GPT-4.1、Gemini-2.5 Flash、および Grok-4 Fast)、および 300 回の実行が含まれ、PASS、PARTIAL、FAIL の結果、厳密および機能的な成功率、エラー パターン、ステップ数、およびトークン消費量を測定しました。エージェント対応 Web サイトでは、ベースラインの 150 件中 74 件に対して 150 件中 74 件の PASS 実行を達成し (厳密な成功率は 89.3% 対 49.3%)、製品詳細の抽出、比較、および複数制約の選択で最大の向上が見られました。また、部分的な結果が 43 から 3 に減少し、平均歩数が 9.31 から 6.49 に減少しました。これらの結果は、構造の明瞭さ、アクションの合図、証拠シグナル、および時間的妥当性インジケーターを強化することで、AI ブラウザー エージェントの信頼性と効率を大幅に向上できるという予備的な証拠を提供します。
原文 (English)
Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate constraints, and carry out parts of the purchasing process for users. Website design must now support both human and agent-mediated interaction. This paper introduces the agent-ready website, a design framework for enhancing the readability, interpretability, verifiability, and actionability of e-commerce platforms for AI agents. Existing web design, SEO, and generative engine optimization (GEO) metrics do not fully assess a website's capacity for agent-mediated interaction. The proposed framework is structured around three dimensions agent interpretability, agent executability, and agent decision reliability supported by features such as machine readability, semantic clarity, agent actionability, and contextual decision-reliability signals. The framework is evaluated through a controlled experiment comparing a human-oriented baseline and an agent-ready version of an identical website prototype, with identical catalogs, pricing, stock, and shopping workflows. The evaluation involved five tasks, three browser-agent models (GPT-4.1, Gemini-2.5 Flash, and Grok-4 Fast), and 300 runs, measuring PASS,PARTIAL,FAIL outcomes, strict and functional success rates, error patterns, step counts, and token consumption. The agent-ready website achieved 134 PASS runs out of 150 versus 74 out of 150 for the baseline (strict success rates of 89.3% vs. 49.3%), with the largest gains in product detail extraction, comparison, and multi-constraint selection. It also reduced PARTIAL outcomes from 43 to 3 and lowered the average step count from 9.31 to 6.49. These results provide preliminary evidence that enhanced structural clarity, action cues, evidence signals, and temporal validity indicators can substantially improve the reliability and efficiency of AI browser agents.
グラフフィードバックは、無重み言語モデル集団における合意形成と派閥形成を制御する
マルチエージェント言語モデル システムでは、ローカル インタラクションのルーティングがますます増えていますが、ランタイム インタラクション グラフは実装の詳細として扱われることがよくあります。私たちは、命名ゲームプロトコルを使用して、1.1B〜32Bパラメーターにわたる無重力LM集団における規則形成を研究します。トークナイザーセーフラベルに対する制限された最初のトークンスコアにより、プロンプト条件付きスコア状態分布を測定し、状態類似性グラフを構築し、サンプリングされたラベルの一致を潜在的な状態空間の一致から分離することができます。制御された介入全体にわたって、主要なオープンウェイト修復グリッドでは、パートナーラベルの証拠を保持することが必要ですが、十分ではありません。同種閾値類似性ルーティングは、クロスベースエクスポージャを削除し、断片化を増幅しますが、ブリッジシーキングルーティングは、メモリが利用可能な場合に断片化を修復することがよくあります。 3 シード混合 4 モデル グリッドでは、189 回の設定シード実行では、しきい値の類似性によって最終的な動作または状態のコンセンサスが生成されませんが、状態コンポーネントとラベルの不一致ブリッジは、14/18 回の保持メモリ実行で最終的な動作のコンセンサスを回復します。同種のモデル集団全体にわたって、保存された履歴は一般に、断片化されたダイナミクスをコンセンサスに向けてシフトさせます。最も明確なケースは Qwen2.5-32B で、保持された履歴がよく混合された 18 の設定すべてで安定した動作と最終状態のコンセンサスに達しましたが、189 の設定ではしきい値の類似性がどちらの形式のコンセンサスにも達しませんでした。状態のしきい値、母集団サイズ、語彙サイズに対する堅牢性により定性的な順序が維持され、初期ウィンドウのグラフ エネルギー機能により有用なグリッド内診断が提供されます。
原文 (English)
Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations
Multi-agent language-model systems increasingly route local interactions, yet the runtime interaction graph is often treated as an implementation detail. We study convention formation in open-weight LM populations spanning 1.1B-32B parameters with a naming-game protocol. Restricted first-token scores over tokenizer-safe labels let us measure prompt-conditioned score-state distributions, construct state-similarity graphs, and separate sampled-label agreement from latent state-space consensus. Across controlled interventions, in the main open-weight repair grids, retained partner-label evidence is necessary but not sufficient: homophilous threshold-similarity routing deletes cross-basin exposure and amplifies fragmentation, while bridge-seeking routing often repairs fragmentation when memory is available. In a three-seed mixed four-model grid, threshold-similarity produces no final behavioral or state consensus in 189 setting-seed runs, whereas state-component and label-disagreement bridges recover final behavioral consensus in 14/18 retained-memory runs. Across homogeneous model populations, retained history generally shifts fragmented dynamics toward consensus; the clearest case is Qwen2.5-32B, which reaches stable behavioral and final state consensus in all 18 retained-history well-mixed settings, while threshold-similarity reaches neither form of consensus in 189 settings. Robustness over state thresholds, population size, and vocabulary size preserves the qualitative ordering, and early-window graph-energy features provide useful within-grid diagnostics.
会話エージェントの多次元評価の運用化: 選択的再評価とモデル ベンチマークを備えたスケーラブルで管理されたパイプライン
小売会話エージェントを評価するには、語彙の重複の指標を超えて、意図の一致、事実性、有用性、明瞭さ、トーン、および全体的な応答品質を評価する方法が必要です。 LLM-as-a-judge メソッドは人間による評価に代わるスケーラブルな代替手段を提供しますが、運用環境の展開では、ガバナンス、再現性、コスト、スキーマの一貫性、トレーサビリティ、および信頼性の点で課題が生じます。小売会話システムの大規模評価のための、管理された構成主導のパイプラインである GenAI Evaluation を紹介します。正規化、シャーディング、非同期実行、スキーマに制約された LLM スコアリングを通じて実稼働チャットボット ログを処理します。このフレームワークは、有用性、真実性、明瞭さ、トーンの調整、および翻訳固有の側面を評価します。選択的再評価では、不完全、不正な形式、またはスキーマが無効なレコードのみが処理され、スキーマ ロック、バージョン管理された構成、検証ログ、およびレコード レベルの出自が監査可能性をサポートします。このフレームワークは毎日約 50,000 件のレコードを処理し、200 万件を超えるインタラクションを評価しました。検証には、訓練を受けた 4 人のアノテーターから得た、人間がラベルを付けた 12,980 件の階層化ランダムなレコードが使用されました。分類には、14 のインテント、156 のサブインテント、18 の主要ドメイン、および 129 のサブドメインが含まれていました。このパイプラインは、マクロ F1 スコア 0.93 と人間の許容可能な翻訳精度 89% を達成しました。
原文 (English)
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality. Although LLM-as-a-judge methods provide scalable alternatives to human evaluation, production deployment introduces challenges in governance, reproducibility, cost, schema consistency, traceability, and reliability. We present GenAI Evaluation, a governed, configuration-driven pipeline for large-scale evaluation of retail conversational systems. It processes production chatbot logs through normalization, sharding, asynchronous execution, and schema-constrained LLM scoring. The framework evaluates helpfulness, truthfulness, clarity, tone alignment, and translation-specific dimensions. Selective re-evaluation processes only incomplete, malformed, or schema-invalid records, while schema locking, versioned configurations, validation logs, and record-level provenance support auditability. The framework processes approximately 50,000 records daily and has evaluated more than two million interactions. Validation used 12,980 stratified-random human-labeled records from four trained annotators. Classification covered 14 intents, 156 sub-intents, 18 major domains, and 129 sub-domains. The pipeline achieved a macro F1 score of 0.93 and 89% human-acceptability accuracy for translation.
Playtrace 再構築パーティショニングによる時間の経過に伴うレベルの表現と生成
ビデオ ゲームは、時間の経過とともに体験できるダイナミックなメディアです。ビデオ ゲーム レベルを生成するための手続き型コンテンツ生成 (PCG) アプローチは数多くありますが、それらは多くの場合、この動的な性質を抽象化した表現を使用します。この論文では、動的な情報を暗黙的にエンコードする、時間の経過に伴うゲーム レベルの新しい、ドメインに依存しない「ケーキ」表現を紹介します。このケーキ表現のために特別に開発された、新しいレベル生成アプローチ Playtrace Reconstructive Partitioning (PRP) を紹介します。 \textit{倉庫番} のゲーム ドメインにおける 6 つの最先端の PCG アプローチと比較したところ、私たちのアプローチはソリューションの多様性を犠牲にすることなく有効なレベルを生成できることがわかりました。私たちのケーキ表現は、既存の表現と比較してゲームの暗黙的な動的性質をより適切にエンコードしていると考えており、これによりドメインに依存しないレベル生成アルゴリズム PRP が可能になります。
原文 (English)
Representing and Generating Levels Over Time through Playtrace Reconstructive Partitioning
Video games are a dynamic medium experienced over time. While there are many Procedural Content Generation (PCG) approaches for generating video game levels, they often use representations that abstract away this dynamic nature. In this paper, we introduce a novel, domain-independent ``cake'' representation for game levels over time which implicitly encodes dynamic information. We present a novel level generation approach Playtrace Reconstructive Partitioning (PRP) specifically developed for this cake representation. We compare against six state-of-the-art PCG approaches in the game domain of \textit{Sokoban}, and find that our approach can generate valid levels without sacrificing solution diversity. We believe our cake representation more neatly encodes the implicit dynamic nature of games compared to existing representations, which allows for our domain-agnostic level generation algorithm PRP.
建設によるつながり: 巡回セールスマンの問題に対する扱いやすいツアー近くの限界点を学ぶ
巡回セールスマン問題 (TSP) の学習ベースの手法は、解読または検索後に作成されるツアーを通じて評価されることがよくありますが、学習されたオブジェクト自体は、ヒートマップ、割り当て、構築ポリシー、または検索指導スコアなどの代理空間に存在することがよくあります。これには、デコード前に実際にどのようなハミルトニアン構造が学習されたのかという根本的な疑問が隠されています。この研究では、ハミルトニアン構造の大部分を最終解読段階に残すのではなく、構造的に意味のある潜在オブジェクトを通じて TSP を学習することで、この質問に直接答えます。 Connected-by-Construction ルート $1$-tree Gibbs ファミリーに基づいて、\emph{C2TSP} と呼ばれるエンドツーエンドの教師なし学習パイプラインを提案します。パイプラインは、暗黙的な微分を通じて不偏 TSP コストから残留エッジ摂動を学習します。構造修正の場合、平滑化された Held-Karp 層は期待される次数のバランスを復元し、証明書に基づくシャープ化は接続された分布をよりツアーのような構造へとさらに推し進めます。実験では、C2TSP が解釈可能な構造情報を保持しながら、強力なデコード性能を発揮することが示されています。アブレーションでは、エッジ摂動と証明書に基づく研磨が共同してツアーコストとツアーのような構造の両方を改善することをさらに検証します。
原文 (English)
Connected by Construction: Learning Tractable Near-Tour Marginals for Traveling Salesman Problems
Learning-based methods for the traveling salesman problem (TSP) are often evaluated through the tours produced after decoding or search, but the learned object itself frequently lives in a surrogate space such as heatmaps, assignments, construction policies, or search-guidance scores. This hides the fundamental question: what Hamiltonian structure has actually been learned before decoding? In this study, we directly answer this question by learning TSP through a structurally meaningful latent object, rather than leaving most of the Hamiltonian structure to the final decoding stage. Based on a connected-by-construction rooted $1$-tree Gibbs family, we propose an end-to-end unsupervised learning pipeline called \emph{C2TSP}. The pipeline learns residual edge perturbations from unbiased TSP cost through implicit differentiation. For structural correction, a smoothed Held--Karp layer restores expected degree balance, while certificate-guided sharpening further pushes the connected distribution toward more tour-like structures. Experiments show that C2TSP yields strong decoding performance while preserving interpretable structural information. Ablations further verify that edge perturbation and certificate-guided sharpening jointly improve both tour cost and tour-like structure.
地理空間基盤モデルの新たなパラダイム: 事前トレーニングからエージェント推論まで
衛星画像と航空画像の解析は、基礎モデルの出現により新しい時代に入りました。このペーパーでは、さまざまな方法論を通じて大規模な地理空間データセットで事前トレーニングされた人工知能/機械学習 (AI/ML) モデルである地理空間基盤モデル (GeoFM) の概念について説明します。まず、GeoFM によって可能になる核となるパラダイム シフトについて明確に説明します。職務の分離です。大規模なモデル プロバイダーが計算集約的な事前トレーニングを実行し、ドメインの専門家がこれらのモデルを迅速に微調整したり、特定のミッション クリティカルなタスクを実行できるようにします。このアプローチは、下流タスクのセキュリティと機密性を維持しながら、最先端の AI/ML へのアクセスを民主化します。次に、マスクされた自動エンコーディングなどの自己教師あり手法によって生成される微調整可能な視覚モデルと、オープン語彙画像分析などのゼロショット タスクを可能にする対照学習によって生成される視覚言語モデルを区別しながら、さまざまなタイプの GeoFM によって解き放たれる新しい機能を探索します。次に、パフォーマンスとコストの分析からより広範な MLOps エコシステムに至るまで、GeoFM を運用するための実際的な考慮事項について説明します。そのために、モデル適応戦略の分類を導入し、ドメイン専門家が特定のミッションセットに対して最もコスト効率の高い適応アプローチを選択できるフレームワークを提案します。最後に、Agentic Geospatial Reasoning の将来を見据えたビジョンを紹介します。そこでは、大規模言語モデルがインテリジェント オーケストレーターとして機能し、GeoFM をツールとして利用して、自然言語で高レベルのユーザー クエリに答え、複雑な分析ワークフローを自動化し、分野を知覚から認知に移行させます。
原文 (English)
The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning
The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models. This paper describes the concept of Geospatial Foundation Models (GeoFMs), which are artificial intelligence/machine learning (AI/ML) models pre-trained on massive geospatial datasets through varied methodologies. We first articulate the core paradigm shift that GeoFMs enable: a separation of duties, where large-scale model providers perform the computationally intensive pretraining, allowing domain experts to rapidly fine-tune or prompt these models for specific, mission-critical tasks. This approach democratizes access to state-of-the-art AI/ML while maintaining the security and confidentiality of the downstream task. We then explore the novel capabilities unlocked by different types of GeoFMs, distinguishing between the finetunable vision models produced by self-supervised techniques like masked auto-encoding, and the vision-language models produced by contrastive learning which enable zero-shot tasks like open-vocabulary image analysis. Next, we discuss the practical considerations for operationalizing GeoFMs, from performance-cost analysis to the broader MLOps ecosystem. To that end, we introduce a taxonomy of model adaptation strategies and propose a framework for domain experts to select the most cost-effective adaptation approach for their particular mission set. Finally, we present a forward-looking vision of Agentic Geospatial Reasoning, where Large Language Models act as intelligent orchestrators, leveraging GeoFMs as tools to answer high-level user queries in natural language and automate complex analytical workflows, moving the field from perception to cognition.
コストガバナンド RAG: マルチテナント LLM システムでの取得と生成にわたるテナントごとの統合コスト帰属
エンタープライズ検索拡張生成 (RAG) の導入は、重大なガバナンス ギャップに直面しています。LLM 生成コストはトークンごとに計測されますが、取得レイヤー (ベクトル メモリ、類似性計算、埋め込み API 呼び出し) は帰属されない共有コストのままであり、テナント間の目に見えない相互補助が可能になります。我々は、コードブックを無視したベクトル インデックス (TurboVec) とマルチテナント LLM ガバナンス ゲートウェイを統合するアーキテクチャであるコスト ガバナンス RAG を紹介します。これにより、埋め込み、取得、生成のコストがテナントごとに共同して発生する統合された可観測性スタックが作成されます。このアーキテクチャは、TurboVec の決定論的で閉じた形式のメモリ式を利用して、ほぼ正確なテナントごとの取得コスト計算を可能にします。これは、非線形メモリ オーバーヘッドを持つグラフベースのインデックスでは利用できない特性です。クラウド データ プラットフォームのガバナンス境界内の Snowpark Container Services にデプロイされたこのシステムは、クエリ遅延の 0.04% 未満のテレメトリ オーバーヘッドで、シミュレートされた 100 のテナント (1,000 万ベクトル、対数正規サイズ分布) 全体で 99.96% のエンドツーエンドのコスト帰属精度を達成します。このアーキテクチャにより、セクション IV で詳述する価格設定の仮定の下で、マネージド ベクトル データベース サービスと比較して、検索インフラストラクチャのコストが 3.1 ~ 9.0 倍削減されます。我々は 3 層のコスト モデルを形式化し、コードブックを意識しない量子化により、テナントごとの決定的なコスト帰属が可能になると同時に、トレーニングされた量子化器に存在する共有コードブックの漏洩曲面も除去できることを実証します。後者の観察は探索的なものであり、セクション VII で説明する制限の影響を受けます。
原文 (English)
Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems
Enterprise Retrieval-Augmented Generation (RAG) deployments face a critical governance gap: while LLM generation cost is metered per token, the retrieval layer - vector memory, similarity compute, and embedding API calls - remains an unattributed shared cost, enabling invisible cross-subsidization among tenants. We present Cost-Governed RAG, an architecture that integrates a codebook-oblivious vector index (TurboVec) with a multi-tenant LLM governance gateway, creating a unified observability stack where embedding, retrieval, and generation costs are jointly attributable per tenant. The architecture exploits TurboVec's deterministic, closed-form memory formula to enable near-exact per-tenant retrieval cost calculation - a property unavailable in graph-based indexes with non-linear memory overhead. Deployed on Snowpark Container Services within a cloud data platform's governance boundary, the system achieves 99.96% end-to-end cost attribution accuracy across 100 simulated tenants (10M vectors, log-normal size distribution) with telemetry overhead below 0.04% of query latency. The architecture reduces retrieval infrastructure cost by 3.1-9.0x compared to managed vector database services under the pricing assumptions detailed in Section IV. We formalize a three-layer cost model and demonstrate that codebook-oblivious quantization enables deterministic per-tenant cost attribution while also removing the shared-codebook leakage surface present in trained quantizers - the latter observation being exploratory and subject to the limitations described in Section VII.
フロンティア言語モデルにおける CBRN 上昇評価のためのしきい値超過フレームワーク
フロンティア言語モデルが進歩するにつれて、政策立案者やモデル開発者は、モデルへのアクセスが、公共ツールのみと比較して、結果の大きな化学、生物、放射線、核(CBRN)の悪用を計画する非専門家の能力を実質的に高めるかどうかを評価する方法を必要としています。既存の CBRN 評価は、専門家以外の定義、脅威の範囲、ベースライン、スコアリング ルーブリック、および決定ルールが異なるため、研究間で結果を比較することが困難です。私たちは、上昇率調査を独立して実行可能なコンポーネントに分解する、閾値超過基準(TEC)フレームワークを導入します。つまり、専門家以外の参加者の適格性の決定、研究の CBRN 脅威範囲の定義、および重要な上昇率の統計的推定です。次に、生成型 (モデルがゼロからの計画作成を支援する) と修正主義型 (モデルが既存の計画の改良を支援する) という 2 つの形態の向上を決定する設計を使用して、大規模な実証研究で TEC フレームワークを運用します。この調査では、CBRN ドメイン全体にわたる攻撃計画が作成され、対象分野の専門家によるレビューを通じて評価し、生成的および修正主義的な上昇を推定しました。このフレームワークを適用した私たちの実証研究では、領域の不均一性が明らかになりました。この制御されたリリース前評価の下では、モデル支援計画は専門家と同等の指導評価を受けることがありましたが、物質的な上昇は放射線領域に限定されていることが確認されました。これらの調査結果は、デプロイされたモデルの動作を特徴づけるのではなく、緩和策とデプロイメントガバナンスの決定に影響を与えました。最後に、事前に指定された基準、明確なベースライン、生成的推定と修正主義的推定の分離、予備的なスクリーニング信号と確認されたリスク判定の慎重な区別を強調しながら、将来の CBRN 上昇率評価のための方法論的な教訓を述べます。
原文 (English)
A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models
As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases a non-expert actor's ability to plan high-consequence Chemical, Biological, Radiological, or Nuclear (CBRN) misuse relative to public tools alone. Existing CBRN evaluations differ in non-expert definitions, threat scope, baselines, scoring rubrics, and decision rules, making results difficult to compare across studies. We introduce a Threshold Exceedance Criteria (TEC) framework that decomposes an uplift study into independently executable components: determining non-expert participant eligibility, defining the CBRN threat scope for the study, and statistically estimating material uplift. We then operationalize the TEC framework in a large-scale empirical study using a design that determines two forms of uplift: generative (where a model assists plan creation from scratch) and revisionist (where a model assists refinement of an existing plan). The study produced attack plans across the CBRN domains, which we evaluated through subject-matter-expert review to estimate generative and revisionist uplift. Applying the framework, our empirical study revealed domain heterogeneity: under this controlled pre-release evaluation, model-assisted plans sometimes received expert-equivalent instructional ratings, but confirmed material uplift was limited to the radiological domain. These findings informed mitigation and deployment-governance decisions rather than characterizing deployed model behavior. We conclude with methodological lessons for future CBRN uplift evaluations, emphasizing prespecified criteria, explicit baselines, separation of generative and revisionist estimates, and careful distinction between preliminary screening signals and confirmed risk determinations.
優れたベンチマーク
良いタスクとは、正確で、解決可能で、検証可能で、詳細に指定されており、興味深い理由から難しいものです。最良のタスクは、経験豊富な実践者が認識する実際の問題を、アプローチではなく結果を検証するテストを使用して、実践者が使用する言語で説明します。
原文 (English)
Good Benchmarks
Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons. The best tasks describe a real problem an experienced practitioner would recognize, in language a practitioner would use, with tests that verify the outcome rather than the approach.
エージェント向けのハーネス進化の評価を再考する
LLM エージェントの自動ハーネス進化の評価を再検討します。既存のハーネス進化手法では、単体テスト ケースを使用してハーネス構成を検索し、同じ公開ベンチマークで最終パフォーマンスを報告します。このプロトコルは 2 つの基本的な懸念を引き起こします。まず、ハーネスの進化自体が反復的な検索手順であり、タスクのフィードバックを使用して候補ハーネスを繰り返し評価および修正します。したがって、エージェントのテスト時間のスケーリングと同様に、一致したフィードバックと推論バジェットの下で単純なタスクレベルの検索ベースラインと比較して、その利益がハーネス設計の改善によるものか、追加の検索のみによるものかを判断する必要があります。第 2 に、検索と最終評価は同じベンチマークを共有するため、報告されたタスクはその特定のタスク セットに過剰適合するリスクが生じます。これらの懸念に対処するために、同等のフィードバックと推論予算の下で、ハーネスの進化を単純なテスト時間のスケーリングと発見ベースラインと比較する広範な評価を実施し、また、保留されたタスクで進化したハーネスを評価して、発見された改善点が一般化するかどうかを評価します。 GPT-5.4 および Claude Opus 4.6 を使用した Terminal-Bench 2.1 での実験では、自動ハーネス進化が常に単純なテスト時間スケーリング手法を上回るパフォーマンスを発揮せず、一般化が限られていることを示しています。私たちの結果は、自動ハーネスの進化の有効性について重要な疑問を提起し、自動ハーネス設計のためのより公平な評価プロトコルとベンチマークの必要性を強調しています。私たちのコードは https://github.com/re Thinking-harness-evolution で入手できます。
原文 (English)
Rethinking the Evaluation of Harness Evolution for Agents
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. First, harness evolution is itself an iterative search procedure that repeatedly evaluates and revises candidate harnesses using task feedback. As in agentic test-time scaling, it should therefore be compared with simple task-level search baselines under matched feedback and inference budgets to determine whether its gains arise from improved harness design or from additional search alone. Second, because the search and the final evaluation share the same benchmark, the reported gains risk overfitting to that specific task set. To address these concerns, we conduct an extensive evaluation comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluate evolved harnesses on held-out tasks to assess whether the discovered improvements generalize. Experiments on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6 show that automatic harness evolution does not consistently outperform simple test-time scaling methods and exhibits limited generalization. Our results raise important questions about the effectiveness of automatic harness evolution and highlight the need for fairer evaluation protocols and benchmarks for automatic harness design. Our code is available at https://github.com/rethinking-harness-evolution.
4B でのオンデバイスの詳細な調査: 忠実性の露出限界、カバレッジ限界の取得取得
デバイス上の調査エージェントは、個人のラップトップでコーパスを検索し、情報源を読み、引用された要約を作成します。引用が忠実であるかどうか、またそのコストは、展開可能な小規模モデルでは測定できません。この研究では、24 GB ラップトップ上の 1 つの 4B ジェネレーターを修正し、その引用が忠実である理由を尋ねます。通常は 1 つの数値として報告される 2 つの量を分離します。引用された主張の忠実性は、引用された情報源がその主張を支持しているかどうかを尋ねます。信頼できる報道では、エージェントが適切な情報源も引用しているかどうかが問われます。この研究では、ジェネレーターが各ソースをどれだけ認識するか (400 文字と 1500 文字)、提供されたソースの品質、ゴールド ペーパーと取得されたペーパーを比較します。 2 つのレバーが外れて、異なる結果に作用します。露出によって忠実さが決まります。各ソースを増やすと、忠実度が取得ソースでは 0.45 から 0.58 に、金ソースでは 0.37 から 0.58 に上昇し、2 つの設定は収束するため、忠実度はソースが正しいかどうかではなく、露出によって決まります。暴露リフトは、第 2 の独立した裁判官に対して堅牢です。正確な収束は、第一の裁判官の下では厳密ですが、第二の裁判官の下では近似に過ぎません。取得によりカバレッジが設定されます。再現率は 0.40 付近に保たれているため、露出によってどの情報源が引用されているかを特定することはできません。追加のエクスポージャには約 235 出力トークンがかかります。実際的なレシピは、まずソースごとの露出を低コストで増やし、次に残りの唯一の手段として検索リコールを扱うことです。
原文 (English)
On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage
On-device research agents search a corpus, read sources, and write a cited brief on a personal laptop. Whether their citations are faithful, and at what cost, is unmeasured for a deployable small model. This study fixes one 4B generator on a 24 GB laptop and asks what makes its citations faithful. It separates two quantities usually reported as one number. Cited claim faithfulness asks whether the cited source supports the claim. Trustworthy coverage asks whether the agent also cites the right sources. The study crosses how much of each source the generator sees, 400 against 1500 characters, with the quality of the sources supplied, gold papers against retrieved papers. Two levers fall out, and they act on different outcomes. Exposure sets faithfulness. More of each source lifts faithfulness from 0.45 to 0.58 on retrieved sources and from 0.37 to 0.58 on gold sources, and the two settings converge, so faithfulness is bound by exposure, not by whether the source is correct. The exposure lift is robust to a second, independent judge; the exact convergence is tight under the primary judge and only approximate under the second. Retrieval sets coverage. Trustworthy coverage stays near 0.22 on retrieved sources at any exposure, because recall is held near 0.40, so exposure cannot fix which sources are cited. The extra exposure costs about 235 output tokens. The practical recipe is to raise per source exposure first, cheaply, and then treat retrieval recall as the only remaining lever.
エージェントのベンチマークを決定するにはどれくらいのタスクがあれば十分ですか?パブリック LLM エージェント ベンチマークのリプレイ分析
エージェントのベンチマークでは、すべてのタスクの実行後に 2 つのエージェントを比較することがよくありますが、コストがかかる評価では部分的な実行が誘惑的になります。タスクの部分だけでは、部分的な実行が完了したベンチマークと同じペアごとの結論をサポートしているかどうかはわかりません。私たちは、SWE ベンチ、AppWorld、および tau ベンチからの完了したパブリック タスク レベルのレコードを再生することで、この疑問を研究します。部分的な予算は、完了したベンチマークの決定をサポートし、必要なタスク グループをカバーし、未解決の比較の目標部分のみを残す場合にのみ十分とみなされます。必要なタスクの割合は大きく異なります。 5 パーセント ポイントの予算グリッド上の厳格な 0 パーセント ポイントのしきい値では、AppWorld は 15 パーセント、tau-bench は 25 パーセント、SWE ベンチ検証は 90 パーセントで最初にすべての目標を満たします。 SWE-bench Lite は、プライマリ カバレッジ ルールに基づくすべての目標を 95% 満たしていません。部分評価レポートには、あるエージェントが別のエージェントよりどの程度優れている必要があるか、タスクがどのように選択されるか、どのようなカバレッジ ルールが必要か、どのような決定ルールが使用されるか、未解決のまま残される可能性がある比較の数が記載される必要があります。
原文 (English)
How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not show whether a partial run supports the same pairwise conclusion as the completed benchmark. We study this question by replaying completed public task-level records from SWE-bench, AppWorld, and tau-bench. A partial budget counts as enough only when it supports the completed benchmark's decision, covers required task groups, and leaves no more than a target fraction of comparisons unresolved. The required task fraction varies sharply. At the strict 0 percentage point threshold on a 5 percentage point budget grid, AppWorld first meets all targets at 15 percent, tau-bench at 25 percent, and SWE-bench Verified at 90 percent; SWE-bench Lite does not meet all targets by 95 percent under the primary coverage rule. Partial-evaluation reports should state how much one agent must outperform another, how tasks are selected, what coverage rule is required, what decision rule is used, and how many comparisons may remain unresolved.
PM-Bench: LLM エージェントの将来の記憶の評価
エージェント AI における重要な課題は、将来の記憶、つまり他のアクティビティが進行中に、将来の特定の合図または状態で意図を実行する能力です。最新の LLM エージェントで予想されるメモリ能力を測定するためのテキストベースのベンチマークである PM-Bench を紹介します。認知科学の Virtual Week パラダイムに触発された PM-Bench は、LLM エージェントがユーザーの意図をどの程度維持し、遅延した意図を実行し、潜在的な環境変化を監視するかを評価します。シミュレートされた週 7 日間にわたって、エージェントは、延期されたタスクの期限が来るかどうかを判断しながら、進行中のアクティビティを継続する必要があります。 PM-Bench 上の 8 つの最先端の LLM を 8 つの異なるエージェント構成で比較します。 PM-Bench はすべての設定において困難であることが判明しました。最良の方法である GPT-5.4 エージェントは、私たちの評価では 65.1\% の F1 スコアしか達成できませんでした。さらに、予測記憶を向上させるための単一の戦略がモデル全体を支配することはありません。私たちは、これらの障害を診断し、信頼できる将来の動作をサポートするトレーニングや推論時の介入を開発するための制御されたテストベッドとして PM-Bench をリリースします。
原文 (English)
PM-Bench: Evaluating Prospective Memory in LLM Agents
A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing. We introduce PM-Bench, a text-based benchmark for measuring prospective memory capabilities in modern LLM agents. Inspired by the Virtual Week paradigm from cognitive science, PM-Bench evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes. Over the course of a simulated seven-day week, agents must continue an ongoing activity while deciding whether any deferred task is due. We compare eight state-of-the-art LLMs on PM-Bench under eight different agent configurations. PM-Bench proves challenging across all settings: the best method, a GPT-5.4 agent, reaches only 65.1\% F1 score under our evaluation. Furthermore, no single strategy for improving prospective memory dominates across models. We release PM-Bench as a controlled testbed for diagnosing these failures and developing training or inference-time interventions that support reliable prospective behavior.
Critic Experience Bank: LLM エージェント向けの自己進化するステップレベルの信頼推定
LLM エージェントは、各アクションが後の意思決定の条件となる状態を変更し、単一の間違ったステップがインタラクション バジェットを浪費したり、最終的な失敗が観察されるずっと前に取り返しのつかない副作用を引き起こしたりする外部環境で動作します。したがって、信頼性の高い展開には \emph{ステップレベルの信頼推定}、つまり、提案された各アクションが生産的であるという調整された確率が必要であり、アクションが実行される前に \emph{前}に入手可能です。既存の LLM 信頼度推定器は、指定されたプロンプトからの応答をスコアリングするように設計されていますが、エージェントの信頼度は、実行の結果、つまり、環境が応答した後、同様の状況での同様のアクションが実際にタスクを進めたかどうかにも依存します。 \method (\methodshort) を導入します。これは、LLM 批評家が自身の過去の判断とその観察された結果から証拠を蓄積する、自己進化する批評家フレームワークです。各軌跡の後、完全な実行フィードバックを確認する後知恵 LLM が、各ステップが生産的であったかどうかを投票します。結果として得られる疑似ラベルはメモリ バンクに格納され、同様のステップが繰り返されるたびに、そこから関連する生産的な経験と非生産的な経験が批評家のプロンプトに取得されます。 \methodshort はトレーニングを必要とせず、グラウンド トゥルース ステップ ラベルを使用しません。 3 つのエージェント ベンチマークと 3 つの批評家バックボーンにわたって、\methodshort はすべてのデータセットと批評家の組み合わせで最高のキャリブレーション (ECE および Brier) とランキング (AUC) を達成し、最も強力なトレーニングなしのベースラインと比較して ECE を最大 $54\%$ 削減します。
原文 (English)
Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents
LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final failure is observed. Reliable deployment therefore requires \emph{step-level confidence estimation}: a calibrated probability that each proposed action is productive, available \emph{before} the action is executed. Existing LLM confidence estimators are designed to score a response from the given prompt, but agent confidence also depends on execution consequences: whether similar actions in similar situations actually advanced the task after the environment responded. We introduce the \method (\methodshort), a self-evolving critic framework in which an LLM critic accumulates evidence from its own past judgments and their observed consequences. After each trajectory, a hindsight LLM that sees the full execution feedback votes on whether each step was productive. The resulting pseudo-labels populate a memory bank from which related productive and unproductive experiences are retrieved into the critic's prompt whenever a similar step recurs. \methodshort requires no training and uses no ground truth step labels. Across three agent benchmarks and three critic backbones, \methodshort attains the best calibration (ECE and Brier) and ranking (AUC) in every dataset--critic combination, reducing ECE by up to $54\%$ relative to the strongest training-free baseline.
LLM エージェント システムの安全性のための第一級原則としての分離: 概念、分類、課題、および将来の方向性
システムの「頭脳」として機能する LLM エージェントの機能により、スタンドアロン モデルを超えて分析の範囲が根本的に拡張されます。その結果、安全性はもはや入力と出力のコンテンツの整合性だけを考慮するものではなくなりました。これは、システムの動作や現実世界の実行結果にも関係します。ただし、現在の文献は、攻撃の種類、アプリケーション、ベンチマークごとに断片化されています。このため、プロンプト インジェクション、ツールの誤用、メモリ ポイズニングなどの障害が同じ構造的原因を共有することが多い理由と、それらの障害がエージェント ワークフローを通じてどのように広がるかを説明することが困難になります。この調査では、LLM エージェント システムの安全性のための第一級の原則として隔離を扱います。分離とは、ユーザー入力、ツール アクセス、実行チャネル、エージェント間通信、環境由来のコンテキストの分離を指します。私たちは、ユーザー - エージェント、エージェント - ツール、エージェント - 実行、エージェント - エージェント、およびシステム - 環境という 5 つの境界の境界中心の分類法を使用して文献を整理しています。このビューは、分離の喪失が最初に発生する場所、侵害が境界を越えてどのように伝播するか、各インターフェイスでどの防御が最も適切であるかを特定するのに役立ちます。また、境界を越えた障害経路を要約し、未解決の課題について議論し、将来のエージェント システムにおける構築による隔離に関する研究課題の概要を示します。
原文 (English)
Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions
The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequently, safety is no longer only about input--output content alignment. It also concerns system behavior and real-world execution outcomes. However, the current literature is fragmented across attack types, applications, and benchmarks. This makes it hard to explain why failures such as prompt injection, tool misuse, and memory poisoning often share the same structural cause, and how they spread through an agent workflow. In this survey, we treat isolation as a first-class principle for LLM-agent system safety. By isolation, we refer to the separation of user inputs, tool access, execution channels, inter-agent communication, and environment-originated context. We organize the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment. This view helps identify where the loss of isolation first occurs, how compromise propagates across boundaries, and which defenses are most relevant at each interface. We also summarize cross-boundary failure paths, discuss open challenges, and outline a research agenda for isolation-by-construction in future agent systems.
受け入れられるプレフィックスだけが必要なわけではありません: PEFT ベースのブロック拡散ドラフティングでは否定的な結果が得られます
投機的デコードでは、安価なドラフターを使用して複数の将来のトークンとそれらを検証するターゲット モデルを提案することで、自己回帰言語モデルの推論を加速します。したがって、共通の設計目標は、補助パラメータとシステムのオーバーヘッドを削減しながら、ドラフトの品質を向上させることです。我々は、LoRA のようなアダプターが自己回帰検証器のブロック拡散ドラフターとして機能する同一バックボーンの投機的復号法である PEFT-BD を通じて、この方向に対する否定的な結果を研究しました。 PEFT-BD は、いくつかの魅力的な特性によって動機付けられています。トークナイザーの不一致を回避し、別のドラフト モデルのロードを回避し、トレーニング可能なパラメーターを少数のみ追加し、BD3LM スタイルのノイズ除去目標を使用してトークンのブロックを並列に提案します。これらの利点にもかかわらず、PEFT-BD は Qwen3-0.6B の実験では実用的な速度向上をもたらしませんでした。このメソッドは重要な受け入れられたプレフィックスを取得しますが、プロファイリングでは、各投機ステップでアダプターが有効になっているフルバックボーンのドラフト パスと、それに続くアダプターが無効になっているフルバックボーンの検証パスが必要であることが示されています。したがって、ドラフターはパラメータ効率は高くなりますが、計算効率は高くありません。私たちの結果は、投機的復号を成功させるための単純だが重要な条件を特定します。それは、起草者は検証者よりも大幅に実行コストが低くなければなりません。ドラフト計算が検証者規模のままである場合、より長く受け入れられたプレフィックスだけでは補償できません。
原文 (English)
Accepted Prefixes Are Not All You Need: A Negative Result on PEFT-Based Block-Diffusion Drafting
Speculative decoding accelerates autoregressive language model inference by using a cheap drafter to propose multiple future tokens and a target model to verify them. A common design goal is therefore to improve draft quality while reducing auxiliary parameters and systems overhead. We study a negative result for this direction through PEFT-BD, a same-backbone speculative decoding method in which a LoRA-like adapter acts as a block-diffusion drafter for an autoregressive verifier. PEFT-BD is motivated by several attractive properties: it avoids tokenizer mismatch, avoids loading a separate draft model, adds only a small number of trainable parameters, and uses a BD3LM-style denoising objective to propose a block of tokens in parallel. Despite these advantages, PEFT-BD does not yield a practical speedup in our Qwen3-0.6B experiments. Although the method obtains nontrivial accepted prefixes, profiling shows that each speculative step requires an adapter-enabled full-backbone draft pass followed by an adapter-disabled full-backbone verification pass. Thus, the drafter is parameter-efficient but not compute-efficient. Our results isolate a simple but important condition for successful speculative decoding: the drafter must be substantially cheaper to execute than the verifier. Longer accepted prefixes alone cannot compensate when draft computation remains verifier-scale.
EVOQUANT: 堅牢な定量取引のための自己進化する検証者主導の戦略最適化
定量的な戦略の最適化は依然として大部分が手動で行われており、分野の専門家が弱いシグナルを特定し、リスク管理ルールを調整し、反復的な修正を繰り返し検証する必要があります。大規模な言語モデルはこのプロセスを加速できますが、トレーディング戦略を書き直すために言語モデルに直接依存すると、幻覚的な編集、戦略のドリフト、バックテストの過剰適合が発生することがよくあります。私たちは、定量的取引における戦略最適化のための自己進化型検証者ガイド付きフレームワークである EVOQUANT を提案します。私たちの手法では、LLM を利用してパフォーマンスのボトルネックを深く診断し、意味的に制御された編集候補を生成し、多段階の検証パイプラインを通じて最適な戦略を選択し、最適化の経験を再利用可能な知識に蒸留して継続的な自己改善を実現します。私たちは 7 つの代表的な戦略を使用してメソッドを評価します。そのうち 4 つは A 株市場から、3 つは仮想通貨市場からのものです。実験結果は、私たちの方法がテストされたすべての戦略でシャープ比を大幅に改善することを示しています。テストの平均シャープは -0.298 から 0.538 に増加し、最もパフォーマンスの高い戦略では 199% の相対的な改善が達成されます。より厳格な条件下でのアブレーション研究とストレステストにより、フレームワークの有効性と堅牢性がさらに検証されます。全体として、この研究は定量的戦略の最適化を、コストのかかる手動の試行錯誤から自動化された検証可能な反復パラダイムに変換し、大規模な言語モデルを財務戦略研究に適用するための新しい道を提供します。
原文 (English)
EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading
Quantitative strategy optimization remains largely manual, requiring domain experts to identify weak signals, tune risk-control rules, and repeatedly validate iterative revisions. Large language models can accelerate this process, but directly relying on them to rewrite trading strategies often introduces hallucinated edits, strategy drift, and backtest overfitting. We propose EVOQUANT, a self-Evolving Verifier-guided framework for strategy Optimization in Quantitative trading. Our method utilizes LLMs to deeply diagnose performance bottlenecks, generates semantically controlled candidate edits, selects the best strategy through a multi-stage verification pipeline, and distills optimization experience into reusable knowledge for continual self-improvement. We evaluate our method using seven representative strategies: four from the A-share market and three from the Crypto market. Experimental results show that our method significantly improves the Sharpe ratio across all tested strategies: the average test Sharpe increases from -0.298 to 0.538, and the best-performing strategy achieves a 199% relative improvement. Ablation studies and stress tests under stricter conditions further validate the effectiveness and robustness of the framework. Overall, this work transforms quantitative strategy optimization from costly manual trial and error into an automated and verifiable iterative paradigm, offering a new path for applying large language models to financial strategy research.
交通予測におけるグローバル空間情報抽出には本当に変圧器が必要なのでしょうか?
既存の交通予測モデルは一般に、空間依存関係、特に個々のノードと交通ネットワーク全体のすべてのノードの間の相互作用を通じて得られる表現を特徴付けるグローバル空間情報を抽出することに重点を置いています。しかし、そのようなグローバル情報がモデル化および抽出される根本的なメカニズムは、依然として十分に研究されていません。グローバル情報を自由度の高い適応的注意によって抽出する必要があるのか、それとも単純なグローバル集約演算子によって取得できるのかは不明のままです。この目的のために、空間混合モジュールのみを置き換えて注意に基づくグローバルな相互作用をテストする制御アブレーション フレームワークを設計します。 6 つのトラフィック ベンチマーク全体で、均一なフルレンジ ミキシングと標準空間アテンションは、それぞれ 3 つのデータセットで低い MAE を実現し、平均 MAE の差はわずか 0.14% であり、前者はノード スケールの空間ミキシングの複雑さを O(N2) から O(N) に軽減します。メカニズム分析は、空間的注意を行均一のグローバル背景と不均一な残差にさらに分解します。残差はデータセットに依存する限界値を示しており、行均一のグローバル バックグラウンドを超えた安定したゲインによって空間的注意が正当化されるべきであることを示唆しています。対応するソース コードは https://github.com/uuesti/U-Trans で公開されています。
原文 (English)
Do We Really Need Transformers for Global Spatial Information Extraction in Traffic Forecasting?
Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which characterizes the representations obtained through interactions between each individual node and all nodes across the traffic network. However, the underlying mechanism by which such global information is modeled and extracted remains insufficiently investigated. Whether global information must be extracted by high-degree-of-freedom adaptive attention or can be captured by a simple global aggregation operator remains unclear. For this purpose, we design a controlled ablation framework that replaces only the spatial mixing module to test attention-based global interaction. Across six traffic benchmarks, uniform full-range mixing and standard spatial attention each achieve lower MAE on three datasets, with only a 0.14% difference in mean MAE, while the former reduces node-scale spatial mixing complexity from O(N2) to O(N). Mechanism analysis further decomposes spatial attention into a row-uniform global background and a non-uniform residual. The residual shows dataset-dependent marginal value, suggesting that spatial attention should be justified by stable gains beyond a row-uniform global background. The corresponding source code is publicly available at: https://github.com/uuesti/U-Trans
コーディング エージェント基盤モデルの中間トレーニングとしての機能を意識した中間補充
コーディング エージェントは、外部ツールのリターンを継続的な推論に統合する必要があります。これは、コードに対する標準の左から右への事前トレーニングが順方向でのみ公開する機能です。コーディング エージェントのアクション - 観察 - 継続ループは構造的に関数呼び出しサイトと同形であることがわかります。呼び出し元は引数をバインドし、呼び出し先は別の場所で計算された値を返し、ダウンストリーム コードはその値を消費します。この条件付け構造は、通常のコード内にインターネット規模で存在します。私たちはこれを、関数を意識した中間補充 (FIM) 中間トレーニングを通じて活用します。これは、プログラムの依存関係グラフ分析と複雑さの推論の二重基準によって選択された関数をマスクする自己監視型の目標です。 968 の GitHub リポジトリから抽出された 2.6B トークンの汚染除去されたコーパス上で Qwen2.5-Coder-Instruct (7B/14B) と Qwen3-8B を中間トレーニングし、既存のエージェントのポストトレーニング パイプラインを適用します。中間トレーニングでは、SWE-Bench-Verified が 7B/14B で +2.8/+3.0、Qwen3-8B で +3.2 向上します。 SWE-Bench-Lite のゲインは、同じモデルで +3.7/+4.0/+5.4 です。この改善は、2 つのポストトレーニング パイプライン (R2E-Gym、SWE-Smith) および非 Qwen2.5 ベース (SWE-Lego を使用した Qwen3-8B) に当てはまります。ドメイン内のゲインだけでなく、トレーニング中は、エージェントのポストトレーニングが非エージェントコーディング (LiveCodeBench など) や非コーディングツール使用ベンチマーク (tau-bench、BFCL) に与える能力の低下も軽減します。トレーニング途中のコーパスには Python コードのみが含まれていますが、関数呼び出しの帰納的バイアスはトレーニング後も存続し、一貫したゲインが得られます。
原文 (English)
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent is structurally isomorphic to a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This conditioning structure exists at internet scale in ordinary code. We exploit it through function-aware fill-in-the-middle (FIM) mid-training: a self-supervised objective that masks functions selected via program dependency graph analysis and a complexity-inferability double criterion. We mid-train Qwen2.5-Coder-Instruct (7B/14B) and Qwen3-8B on a 2.6B-token decontaminated corpus drawn from 968 GitHub repositories, then apply existing agentic post-training pipelines. Mid-training improves SWE-Bench-Verified by +2.8/+3.0 at 7B/14B and by +3.2 on Qwen3-8B; SWE-Bench-Lite gains are +3.7/+4.0/+5.4 on the same models. The improvement holds across two post-training pipelines (R2E-Gym, SWE-Smith) and on a non-Qwen2.5 base (Qwen3-8B with SWE-Lego). Beyond in-domain gains, mid-training also mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding (e.g., LiveCodeBench) and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.
観察から洞察へ: 機械論的世界モデルと自律的発見の探求
基礎モデルの最近の進歩により、科学向け AI が変革され、タンパク質のフォールディングから天気予報に至るまでの領域にわたって、驚くほど正確な予測パフォーマンスが可能になりました。しかし、予測だけでは科学的発見にはなりません。科学的理解は、観察を生成する再利用可能な説明メカニズムを明らかにすることにかかっていますが、現代の機械学習は、説明構造ではなく、予測マッピングを中心に基本的に編成されたままです。この論文では、科学的発見は基本的に知識の組織化の問題であると主張します。この目的を達成するために、再利用可能なメカニズムを表現、計算、学習の中心に置く新しい設計パラダイムである Mechanistic World Models を導入します。科学哲学からの洞察を利用して、発見に必要な計算能力を導き出し、説明的な知識の出現を促す設計原理と帰納的圧力を特定し、メカニズム中心の世界モデルの構造を形式化します。最後に、機構的解釈可能性、因果表現学習、方程式発見、モジュラーアーキテクチャなどの多様な研究方向が、統一されたフレームワークを欠きながら、このパラダイムの補完的な要素をどのように捉えているかを示します。私たちは、AI を予測予測を超えて自律的な科学的発見に向けて前進させるための概念基盤および計算青写真として、機械論的世界モデルを提案します。
原文 (English)
From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
Recent advances in foundation models have transformed AI for Science, enabling remarkably accurate predictive performance across domains ranging from protein folding to weather forecasting. Yet prediction alone does not constitute scientific discovery. Scientific understanding depends on uncovering the reusable explanatory mechanisms that generate observations, whereas contemporary machine learning remains fundamentally organised around predictive mappings rather than explanatory structure. In this paper, we argue that scientific discovery is fundamentally a problem of knowledge organisation. To this end, we introduce Mechanistic World Models, a new design paradigm that places reusable mechanisms at the centre of representation, computation and learning. Drawing on insights from the philosophy of science, we derive the computational capabilities required for discovery, identify the design principles and inductive pressures that encourage explanatory knowledge to emerge, and formalise the anatomy of a mechanism-centric world model. Finally, we show how diverse research directions including mechanistic interpretability, causal representation learning, equation discovery and modular architectures capture complementary ingredients of this paradigm while lacking a unified framework. We propose Mechanistic World Models as a conceptual foundation and computational blueprint for moving AI beyond predictive forecasting towards autonomous scientific discovery.
TRACE: 監査可能なエージェントのコミットメントのための運用推論スキーマ
この文書では、TRACE (Typed Reasoning And Commitment Evidence) を定義します。これは、推論トレースを記録するための型付き、バージョン付きのスキーマ、それに対してレコードを書き込むための参照手順、および 1 つの運用規律、つまりレコードなしでは永続的な状態変更はありません。この論文は、推論が言語モデルに含まれていないことを 3 つの層で主張しています。自己回帰メカニズムはネイティブに関連性を計算します。思考連鎖と強化学習はその限界を受け継いでいます。そして、ソクラテスの手続きからパールのはしごに至るまで、推論理論の形式的な構造は機械としては存在しません。スキーマは、フィールドとテストで欠如に答えます。TraceRecord とその因果関係の特殊化、8 段階のリファレンス ライター、ゲートファースト測定レジーム、TRACE-Bench プロトコル、およびコンシューマ、メモリ アドミッション、プラン ゲーティング、一時的後悔、および評決の再利用 (より監査可能な決定が記録の尺度となる) です。レコードと消費者の契約では、レコードが何を保証するのか、また消費者はその見返りとして何を尊重しなければならないのかを規定し、スキーマを受動的な文書ではなく操作可能なインターフェースにします。本文には 2 つの実際に実行された例が掲載されています。1 つは文からタイプされた評決まで追跡され、関連付け、介入、処方箋を分離した音楽レッスンの議論です。そして、洪水捜索救助のビネットでは、予測世界モデルが自信を持って計画の成功を報告しますが、それ自体のサポートと分布外のスコアは矛盾しているため、記録はコミットメントを延期し、限定された観察を要求し、追加のみを改訂し、別のブランチをクリアします。このビネットは例示的なものであり、経験的なものではありません。クローズドループの評価は今後の作業に委ねられるため、貢献するのはスキーマとその契約であり、パフォーマンスの要求ではありません。付録には、完全なスキーマ、ライター アルゴリズムとコスト モデル、臨床およびポリシーの図、ベンチマーク プロトコル、収束メトリック、および使用シナリオが含まれています。
原文 (English)
TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments
This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference procedure for writing records against it, and one operating discipline, no durable state change without a record. The paper argues in three layers that reasoning is not in the language model: the autoregressive mechanism natively computes association; chain-of-thought and reinforcement learning inherit its limits; and the formal constructs of reasoning theory, from Socratic procedure to Pearl's ladder, are absent as machinery. The schema answers the absence with fields and tests: the TraceRecord and its causal specialization, an eight-stage reference writer, a gate-first measurement regime, the TRACE-Bench protocol, and the consumers, memory admission, plan gating, temporal regret, and verdict reuse, whose more auditable decisions are the measure of the record. A record-consumer contract states what a record guarantees and what a consumer must honor in return, making the schema an operational interface rather than a passive document. Two worked examples run in the main text: a music-lessons argument traced from sentence to typed verdict, separating association, intervention, and prescription; and a flood search-and-rescue vignette in which a predictive world model reports confident plan success that its own support and out-of-distribution scores contradict, so the record defers the commitment, requests a bounded observation, revises append-only, and clears a different branch. The vignette is illustrative, not empirical; closed-loop evaluation is left to future work, so the contribution is the schema and its contract, not a performance claim. Appendices carry the full schema, writer algorithms and cost model, clinical and policy illustrations, the benchmark protocol, convergence metrics, and usage scenarios.
モデルはあなたではなくあなたのプロジェクトを知っています: NameRank を使用した LLM の認識の測定
フロンティアモデルが、検索ステップの前に、それ自身の重みから人や道具について思い出す内容が、人間が最初に目にする記述を形作ることが多く、そのパラメトリックコーパスの存在が測定上の問題になります。モデルが研究者を認識できるかどうかの約 3 分の 1 は、引用によって説明されます。残差をターゲットにして、NameRank ([0,1]) 認識スコアを構築します。54 コホートの 4,685 エンティティのそれぞれが、36 のモデルにわたる 1 つの自由回答形式の質問で調査され、独立した裁判官が、厳選された金に対して 2 値評決を返します。モデルは、この正確なエンティティについて、特定の推測不可能な事実を述べていますか? -- つまり、幻覚、文脈エコー、推測は何も得られません。合成ヌル エンティティは下限をゼロ付近に保ち、判定はモデルではなくエンティティを追跡します。ある論文は調査結果を整理しており、認定は資格情報や肩書きではなく、名前付きのインデックス可能な成果物に支払われます。メダルには名前の付いた成果物が付属していないため、オリンピック形式のすべての資格は現役研究者の基準を下回っているが、ノーベル賞、チューリング賞、フィールズ賞受賞者がパネルを埋め尽くす主要層ではランキングが逆転する。独立したクリエイターにとって、このツールはその作成者よりも上位にあり、伝播される資格は名前付きのメソッドまたは受賞した論文です。対照的に、有名なアーティファクトに名を連ねる多数の寄稿者の 1 人であることは、ほとんど何も得られません。主力モデルのレポートやシステム カードに記載されている著者は、認定フロアの近くに座っています。なぜなら、認知はそのアーティファクトの背後にある名簿ではなく、そのアーティファクト自体の固有の名前に関連付けられるからです。認識を適切に予測する書誌情報はありません。最高密度の機関は、一致する引用数で同業機関を上回っています。 258 件のニュース イベントの認識は、永続性ではなく、ピークの顕著性に基づいて行われます。自己報告プローブは、内省がコーパス自体の知識ではなく、事前にコーパスを読み取ることを示しています。
原文 (English)
The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank
What a frontier model recalls about a person or tool from its own weights -- before any retrieval step -- often shapes the first description a human sees, making that parametric corpus presence a measurement problem. Citations explain about a third of whether a model recognizes a researcher; we target the residual and build NameRank, a [0,1] recognition score: each of 4,685 entities in 54 cohorts is probed with one open-ended question across 36 models, and an independent judge returns a binary verdict against a curated gold -- did the model state a specific, non-guessable fact about this exact entity? -- so hallucination, context echo, and guesses earn nothing. Synthetic-null entities hold the floor near zero, and verdicts track the entity, not the model. One thesis organizes the findings: recognition is paid to named, indexable artifacts, not to credentials or titles. Every Olympic-style credential sits below a working-researcher baseline, because no named artifact ships with the medal, yet the ranking inverts at the marquee tier, where Nobel, Turing, and Fields laureates saturate the panel. For independent creators the tool out-ranks its maker, and the credential that does propagate is a named method or awarded paper. Being one of many named contributors to a celebrated artifact, by contrast, earns almost nothing -- the authors listed on a flagship model report or system card sit near the recognition floor -- because recognition attaches to the artifact's own distinctive name, not to the roster behind it. No bibliometric predicts recognition well; top-density institutions out-recognize peers at matched citations; and on 258 news events recognition loads on peak salience, not persistence. A self-report probe shows introspection reads a corpus prior, not its own knowledge.
筋骨格ケアのための証拠に基づいた AI
筋骨格系疾患は世界中で障害の主な原因の一つであり、世界的にリハビリテーションに対する最大のニーズを生み出しています。回復、リモデリング、変性は数カ月から数年かけて進行することが多いため、筋骨格ケアには、進化する患者の証拠、外部の医学知識、段階別の機能目標を繰り返し統合する長期的な管理が必要です。日常診療では、この証拠は訪問、部門、病院システム全体で断片化されており、個別化された証拠に基づいたケアが制限されています。ここでは、継続的な筋骨格管理のために病院のデータ ストリームと信頼できる外部の知識を統合する、大規模な言語モデルを搭載した臨床人工知能システムである OrthoPilot について報告します。 OrthoPilot は、リアルタイムの画像データ、検査データ、病理学データ、注文データを自律的に取得し、入院診断からリハビリテーション計画に至るまで、進化する患者の状態を証拠に基づいた決定に変換します。私たちは、1,000 の疾患コードにわたる実際の電子医療記録から専門家によって検証されたベンチマークを確立しました。コンプリートケア経路全体にわたる読者調査で、OrthoPilot は 81 人の整形外科医と比較され、診断推論、臨床意思決定、管理計画において 25 年の経験を持つ専門家を上回りました。また、60 の外部臨床センターで評価されたすべてのインテリジェント システムを上回りました。 1,870 件の複雑なケースを対象とした前向き研究で、OrthoPilot はフルチェーン管理の成功率を 10.6% 向上させました。 8,240 人の入院患者を対象とした 8 か月間にわたる無作為化導入により、ベッドあたりの累積感染者数が 9.7% 増加し、患者報告による健康情報へのアクセスが改善されました。これらの結果により、臨床 AI は、孤立したイベントの予測から、完全な筋骨格ケア経路にわたる長期的な管理の実行へと移行します。
原文 (English)
Evidence-Grounded AI for Musculoskeletal Care
Musculoskeletal diseases are among the leading causes of disability worldwide and create the greatest global need for rehabilitation. Because recovery, remodelling and degeneration often unfold over months to years, musculoskeletal care requires longitudinal management that repeatedly integrates evolving patient evidence, external medical knowledge and stage-specific functional goals. In routine practice, this evidence is fragmented across visits, departments and hospital systems, limiting individualized, evidence-based care. Here we report OrthoPilot, a clinical artificial intelligence system powered by a large language model that integrates hospital data streams with authoritative external knowledge for continuous musculoskeletal management. OrthoPilot autonomously retrieves real-time imaging, laboratory, pathology and order data and converts evolving patient states into evidence-based decisions from admission diagnosis to rehabilitation planning. We established a specialist-validated benchmark from real-world electronic health records spanning 1,000 disease codes. In a reader study across the complete care pathway, OrthoPilot was compared with 81 orthopaedic physicians and surpassed experts with 25 years of experience in diagnostic reasoning, clinical decision-making and management planning. It also outperformed all evaluated intelligent systems across 60 external clinical centres. In a prospective study of 1,870 complex cases, OrthoPilot increased full-chain management success by 10.6%. During an 8-month randomised deployment involving 8,240 inpatients, it increased cumulative cases per bed by 9.7% and improved patient-reported access to health information. These results move clinical AI from predicting isolated events toward executing longitudinal management across complete musculoskeletal care pathways.
EU AI 法に基づく高リスク AI システムの垂直標準化: アルゴリズム採用のためのドメイン固有のフレームワーク
最近の欧州の法律によると、欧州人工知能 (AI) 法で概説されているリスク管理、データの品質とガバナンス、ロギングとトレーサビリティ、技術文書、透明性、人間の監視、精度などの特定の分野に関連する要件に準拠するには、高リスク AI システムを適応させる必要があります。 AI の標準化プロセスは引き続き反復的であると予想されており、これまでのところ、アルゴリズム採用の課題を完全にカバーする AI に関する欧州標準は存在しないため、欧州委員会が指定する関連 AI 分野に関連した標準化指向の具体的な推奨事項を提案します。これらの各分野について、AI 法に基づく要件に沿って、高リスク領域、特に採用分野の AI システムが満たすべき要件と、その適切な使用と望ましいパフォーマンスを確保するために実行する必要がある活動を説明することでコンテキストを設定します。 AI ガバナンスと標準化に対する既存の水平的なアプローチとは異なり、この論文は、ライフサイクル差別リスク、公平性を意識したデータ ガバナンス、説明可能性、人間による監視、採用システムにおける導入後のモニタリングに焦点を当て、AI 法の要件を具体的な標準化推奨事項にマッピングすることにより、アルゴリズム採用、特にランキングベースの採用システムのための垂直的で領域固有のフレームワークに貢献します。私たちの推奨事項は欧州プロジェクト FINDHR の結果に基づいていますが、プロジェクトの技術的成果物とは結びついておらず、代替の方法、ツール、またはガバナンス メカニズムを使用して実装することができます。
原文 (English)
Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring
According to the recent European legislation, high-risk AI systems will have to adapt in order to comply with requirements related to specific areas, like risk management, data quality and governance, logging and traceability, technical documentation, transparency, human oversight, and accuracy, as outlined in the European Artificial Intelligence (AI) Act. As the standardisation process for AI is expected to remain iterative and, so far, there are no European standards on AI fully covering the challenges of algorithmic hiring, we propose specific standardisation-oriented recommendations related to the relevant AI areas specified by the European Commission. For each of these areas, we set the context by describing the requirements that AI systems in high-risk domains, and especially in recruitment, should fulfil, as well as the activities that should be carried out to ensure their appropriate use and desired performance, in line with the requirements deriving from the AI Act. Unlike existing horizontal approaches to AI governance and standardisation, this paper contributes a vertical, domain-specific framework for algorithmic hiring, and especially ranking-based recruitment systems, by mapping the requirements of the AI Act to concrete standardisation recommendations, focusing on lifecycle discrimination risks, fairness-aware data governance, explainability, human oversight, and post-deployment monitoring in recruitment systems. Even though our recommendations were informed by the outcomes of the European project FINDHR, they are not tied to the project's technical artefacts and could be implemented using alternative methods, tools, or governance mechanisms.
エージェントティック サービス指向コンピューティング: サービス指向コンピューティングの次のフロンティアへのマニフェスト
LLM を利用した自律型および半自律型エージェントの急速な出現により、ソフトウェア システムは、静的な要求と応答のコンポーネントから、目標指向型の適応型でツールを使用する計算アクターへと再構築されています。これらのエージェントが孤立したコグニティブ プロトタイプから複雑な分散ワークフローに移行するにつれて、サービス指向コンピューティング コミュニティが 20 年以上にわたって研究してきた課題、つまり構成、相互運用性、サービス品質、ライフサイクル管理、ガバナンス、セキュリティ、信頼に直面します。しかし、今日のエージェント AI エコシステムの多くは、信頼できる企業や社会への展開に必要な厳密なエンジニアリングを行わずに、これらの基盤をアドホックに開発しています。このペーパーでは、エージェントをサービスとしてエンジニアリングし、自律および半自律エージェントを介してサービスを調整し、信頼、サイバーセキュリティ、コンプライアンス、パフォーマンス、説明責任の制約の下でエージェントとサービスのエコシステムを管理することに関する新しい研究および実践領域として、エージェントティック サービス指向コンピューティング (ASOC) を紹介します。私たちは、ASOC の 6 つの基本原則 (ハーネス能力、構成可能性、ライフサイクル エンジニアリング、設計による信頼性、目標主導のオーケストレーション、観察可能性/説明責任) を明確にし、以下に及ぶ 5 次元の研究課題を組織します。(i) エージェント サービスの基盤とライフサイクル エンジニアリング。 (ii) 構成、オーケストレーション、および相互運用性。 (iii) ガバナンス、観察可能性、説明責任。 (iv) セキュリティ、信頼、およびリスク管理。 (v) 評価、認証、およびエージェント QoS。私たちは、サービス コンピューティング コミュニティは、エージェント AI を断片的なデモンストレーションから人間と組織の信頼に値する信頼できるサービスベースのシステムに変換し、この新興分野に概念的およびエンジニアリングの骨組みを提供するのに特に有利な立場にあると主張します。
原文 (English)
Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing
The rapid emergence of LLM-powered autonomous and semi-autonomous agents is reshaping software systems from static, request-response components into goal-directed, adaptive, and tool-using computational actors. As these agents move from isolated cognitive prototypes into complex distributed workflows, they confront challenges that the Service-Oriented Computing community has studied for more than two decades: composition, interoperability, quality of service, lifecycle management, governance, security, and trust. Yet much of today's agentic AI ecosystem is developing these foundations ad hoc, without the engineering rigour required for dependable enterprise and societal deployment. This paper introduces Agentic Service-Oriented Computing (ASOC) as a new research and practice area concerned with engineering agents as services, orchestrating services through autonomous and semi-autonomous agents, and governing ecosystems of agents and services under constraints of trust, cybersecurity, compliance, performance, and accountability. We articulate six foundational principles of ASOC (harness-ability, composability, lifecycle engineering, trustworthiness by design, goal-driven orchestration, and observability/accountability) and organise a five-dimensional research agenda spanning: (i) agentic services foundations and lifecycle engineering; (ii) composition, orchestration, and interoperability; (iii) governance, observability, and accountability; (iv) security, trust, and risk management; and (v) evaluation, certification, and Agentic QoS. We argue that the Services Computing community is especially well positioned to provide the conceptual and engineering spine for this emerging field, transforming agentic AI from fragmented demonstrations into dependable, service-based systems worthy of human and organisational trust.
X の原子単位: インテリジェンスの圧縮層
この論文は、インテリジェンスを原子圧縮と構成再利用のプロセスとして理解するための理論的枠組みを提案します。私たちは、認知システム、生物学システム、計算システム、および組織システムは、複雑な現象を高次構造に再結合できる再利用可能な原子単位に分解することによって、スケーラブルなインテリジェンスを実現すると主張します。この論文は、認知科学、情報理論、進化生物学、ソフトウェア工学、医学、法的推論、教育、音楽、人工知能からの証拠を利用して、効率、伝達、解釈可能性、進化可能性をサポートする基本的な圧縮層としての原子単位の概念を開発しています。中心的な貢献は Compression Calculus です。これは、表面レベルの表現を原子表現と比較し、抽象化レイヤー全体で圧縮ゲインがどのように合成されるかを説明するための正式なフレームワークです。我々は、複合カスケード理論を導入します。これによれば、抽象化の各層が追加されると、単に増分節約が追加されるのではなく、表現効率が乗算的に向上します。この論文はさらに、現代の AI システムは、安定した概念レベルの原子構造ではなく、トークン レベルの処理やドキュメント レベルの検索に依存し、次善の表現レベルで動作することが多いと主張しています。この観点では、大規模な言語モデルは、完全な知識アーキテクチャとしてではなく、原子単位のナビゲーション、順序付け、および再結合が可能な動的融合エンジンとして最もよく理解されます。このフレームワークは、時間の経過とともに新しいプリミティブを発見、洗練、構成できる自己進化する知識システムを設計するための基盤を提供します。この論文は、構成的抽象化による圧縮としてインテリジェンスを再構成することにより、専門知識、知識表現、説明可能な AI、および適応型インテリジェント システムの将来のアーキテクチャに関する統一的な視点を提供します。
原文 (English)
Atomic Units of X: The Compression Layer of Intelligence
This paper proposes a theoretical framework for understanding intelligence as a process of atomic compression and compositional reuse. We argue that cognitive, biological, computational, and organizational systems achieve scalable intelligence by decomposing complex phenomena into reusable atomic units that can be recombined into higher-order structures. Drawing on evidence from cognitive science, information theory, evolutionary biology, software engineering, medicine, legal reasoning, education, music, and artificial intelligence, the paper develops the concept of atomic units as fundamental compression layers that support efficiency, transfer, interpretability, and evolvability. The central contribution is the Compression Calculus, a formal framework for comparing surface-level representations with atomic representations and for describing how compression gains compound across abstraction layers. We introduce the Compounding Cascade thesis, according to which each additional layer of abstraction multiplicatively increases representational efficiency rather than merely adding incremental savings. The paper further argues that contemporary AI systems often operate at suboptimal levels of representation, relying on token-level processing or document-level retrieval rather than stable, concept-level atomic structures. In this view, large language models are best understood not as complete knowledge architectures, but as dynamic fusion engines capable of navigating, sequencing, and recombining atomic units. The framework provides a foundation for designing self-evolving knowledge systems that can discover, refine, and compose new primitives over time. By reframing intelligence as compression through compositional abstraction, the paper offers a unifying perspective on expertise, knowledge representation, explainable AI, and the future architecture of adaptive intelligent systems.
小規模言語および視覚言語モデル Web エージェントにおける学習率ゲート型 GRPO の失敗: 制御されたヌルとそのメカニズム
検証可能な報酬を伴う強化学習、特にグループ相対ポリシー最適化 (GRPO) は、より強力なエージェントを生成することを目的として、教師付きチェックポイントで定期的に実行されるようになりました。 4B から 8B スケールの小さな言語およびビジョン言語モデルの Web エージェントにスキルを追加するのか、それとも教師ありモデルがすでに備えている動作を主に再構築するのかを尋ねます。学習率、KL 重み、シード、初期化、およびクリッピングを変化させる 18 回の実行の制御グリッド全体にわたって、エージェントがほぼ習得したタスクに関する強力な教師ありベースラインの成功率を確実に向上させる構成はありません。テキストトラックでは、学習率が中程度から高いと確実に悪化します。ヌルは、ペアテスト、25 の評価シード、6 つのトレーニングシード、レシピの変更、テキストとセットオブマークのスクリーンショットの両方の観察、およびバックボーンの 8B へのスケーリングの下で保持されます。信頼できる被害はテキストトラックの発見であり、セット・オブ・マークの下では名目上のものにすぎません。 null が壊れたパイプラインではなく設定を反映していることを示すために、サンプリングによって報酬に到達できるタスクで同一のハーネス、報酬、レシピを実行すると、ゼロを除外したペアの間隔で成功率が 22 ポイント上昇しました。したがって、GRPO は、上昇する余地がある場合にのみ役立ちます。つまり、サンプリングされたポリシーがすでに貪欲なポリシーよりも成功する頻度が高いことを意味します。次に、失敗について説明します。中程度の学習率ではエージェントが劣化し、高い学習率ではエージェントが崩壊し、2 つのレジームは二重の解離を形成します。グラフティングにより劣化レジームは注意と MLP ブロックに局所化されますが、崩壊レジームは単一グループに追跡できず、重みの動きを支配する埋め込み変化は因果的に不活性です。 4B では、後期層の実効ランクは両方向の能力を追跡します。 8Bで両者は離れます。この結合はより小さいモデルに固有であるため、スケールに依存するものとして報告します。
原文 (English)
A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism
Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to a small language and vision-language model web agent at the 4B to 8B scale, or whether it mostly reshapes behavior the supervised model already has. Across a control grid of 18 runs that varies learning rate, KL weight, seed, initialization, and clipping, no configuration credibly improves the success rate of a strong supervised baseline on tasks the agent has largely mastered. On the text track, moderate to high learning rates make it credibly worse. The null holds under paired testing, 25 evaluation seeds, 6 training seeds, changes to the recipe, both text and Set-of-Marks screenshot observations, and scaling the backbone to 8B; the credible harm is a text-track finding and is only nominal under Set-of-Marks. To show that the null reflects the setting and not a broken pipeline, we run the identical harness, reward, and recipe on tasks whose reward is reachable by sampling, and there the success rate rises by 22 points with a paired interval that excludes zero. GRPO therefore helps only when there is headroom to climb, meaning the sampled policy already succeeds more often than the greedy one. We then explain the failure. A middle learning rate degrades the agent and a high one collapses it, and the two regimes form a double dissociation: grafting localizes the degrade regime to the attention and MLP blocks, while the collapse regime cannot be traced to any single group, and the embedding change that dominates the weight movement is causally inert. At 4B, effective rank in the late layers tracks capability in both directions; at 8B the two come apart. This coupling is specific to the smaller model, so we report it as scale-dependent.
エージェントのインターネット: クローズドループ IoT オーケストレーションのためのネットワーク化された AI エージェント
この論文では、エージェントティック AI、IoT、サイバー物理システム、物理 AI、エッジ コンピューティング、デジタル ツインを統合された閉ループ オーケストレーション フレームワークに統合するアーキテクチャ フレームワークである、エージェントティック シングスのインターネット (IoAT) について紹介します。提案されたアーキテクチャは、分散したサイバー物理環境全体で認識、推論、調整、作動する自律型 AI エージェントを介して接続されたクラウド、エッジ/フォグ、物理 IoT レイヤーで構成されています。この論文では、エージェントの計画と物理的な実行をリンクする形状動的プログラミング フレームワークを使用して、入れ子になった戦略的および戦術的な意思決定を伴うワークフロー制御の結合問題として IoAT を形式化しています。スマート ビルディング オーケストレーションが代表的なユース ケースとして紹介され、安全性、セキュリティ、ガバナンス、復元力、信頼性の高い導入に関連する主要な研究課題について説明します。
原文 (English)
Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration
The paper introduces the Internet of Agentic Things (IoAT), an architectural framework that integrates agentic AI, IoT, cyber-physical systems, Physical AI, edge computing, and digital twins into a unified closed-loop orchestration framework. The proposed architecture consists of cloud, edge/fog, and physical IoT layers connected through autonomous AI agents that perceive, reason, coordinate, and actuate across distributed cyber-physical environments. The paper formalizes IoAT as a coupled workflow-control problem with nested strategic and tactical decision making using a hylomorphic dynamic programming framework that links agentic planning with physical execution. Smart-building orchestration is presented as a representative use case, and key research challenges related to safety, security, governance, resilience, and trustworthy deployment are discussed.
数独の視覚言語モデルを指導するための MaxSAT ベースのフィードバック
視覚 -- 言語モデル (VLM) は最近、グリッドベースのパズルを含む、構造化された視覚的推論タスクで有望なパフォーマンスを実証しました。ただし、強力な知覚機能にもかかわらず、これらのモデルには論理的一貫性を強制するための明示的なメカニズムが欠けており、基礎となる制約に違反する割り当てが頻繁に生成されます。この論文では、最大満足度 (MaxSAT) オラクルを介して形式的制約推論を VLM 解決プロセスに統合する神経記号的アプローチを提案します。シンボリック コンポーネントは、ソリューションを直接計算するのではなく、一貫性検証および改良エンジンとして機能します。 VLM によって生成された候補配置は、部分的な MaxSAT 定式化のソフト句としてエンコードされますが、Sudoku 制約はハード句のままです。不一致が発生した場合、MaxSAT ソルバーは、相互に一貫した割り当ての最大のサブセットを特定し、それが構造化されたテキストおよび視覚的なフィードバックに変換され、その後の改良の指針となります。私たちは、複数のオープンソースおよびクローズドアクセス VLM にわたる Sudoku データセットに対するアプローチを評価します。結果は、特にフルボード改良モードにおいて、MaxSAT ベースのフィードバックにより論理的一貫性が向上し、解決されたインスタンスの数が増加することを示しています。これらの発見は、記号の最適化が視覚言語推論の信頼性を高めることができることを示しています。
原文 (English)
MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku
Vision--Language Models (VLMs) have recently demonstrated promising performance on structured visual reasoning tasks, including grid-based puzzles. However, despite strong perceptual capabilities, these models lack explicit mechanisms for enforcing logical consistency and frequently generate assignments that violate underlying constraints. In this paper, we propose a neuro-symbolic approach that integrates formal constraint reasoning into the VLM solving process via a Maximum Satisfiability (MaxSAT) oracle. Rather than computing solutions directly, the symbolic component acts as a consistency validator and refinement engine. Candidate placements generated by the VLM are encoded as soft clauses in a partial MaxSAT formulation, while Sudoku constraints remain hard clauses. When inconsistencies arise, the MaxSAT solver identifies a largest mutually consistent subset of assignments, which is then translated into structured textual and visual feedback to guide subsequent refinements. We evaluate our approach on a Sudoku dataset across multiple open-source and closed-access VLMs. Results show that MaxSAT-based feedback improves logical consistency and increases the number of solved instances, particularly in full-board refinement mode. These findings demonstrate that symbolic optimisation can enhance the reliability of vision-language reasoning.
LLM は煙は見えるが火は見えない: エレンコスによるアブダクティブ推論の評価
大規模言語モデル (LLM) はパターン認識とテキスト生成に優れていますが、そのアブダクティブ推論 (観察された動作を説明する潜在的な仮説を推論する) の能力はまだ十分に理解されていません。ここでは、アブダクティブ推論を構造的逆問題として測定する生成的評価フレームワークである Elenchos (ソクラテスの反対尋問法にちなんで名付けられました) を紹介します。ラムダ計算などの参照形式システムと、変異する可能性のある対応物が与えられると、エージェントは変異が発生したかどうかを判断し、結果として生じる動作の違いの原因となるルールの変更を推測する必要があります。フロンティアおよび中間層の LLM を評価すると、一貫した検出と属性の解離が明らかになります。モデルは多くの場合、システムが変更されたことは認識しますが、観察された不一致の原因となっている潜在的な変異を特定するのに苦労します。モデルが基礎となる突然変異のサブセットのみを回復することが多い場合、相互作用する突然変異の下ではパフォーマンスが大幅に低下します。予備的な証拠はまた、推論時間の増加による推論の利益が減少することを示唆しており、より大きな推論予算の下ではわずかな改善しか見られませんが、この発見にはさらなる検証が必要です。
原文 (English)
LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos
Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce Elenchos (named after the Socratic method of cross-examination), a generative evaluation framework that measures abductive reasoning as a structural inverse problem. Given a reference formal system, such as the lambda-calculus, and a potentially mutated counterpart, agents must determine whether a mutation has occurred and infer the rule modifications responsible for the resulting behavioral differences. Evaluating frontier and mid-tier LLMs reveals a consistent detection-attribution dissociation: models often recognize that a system has been altered but struggle to identify the latent mutations causing the observed discrepancies. Performance degrades substantially under interacting mutations, where models frequently recover only a subset of the underlying mutations. Preliminary evidence also suggests diminishing returns from increased inference-time reasoning, with only modest improvements under larger reasoning budgets, though this finding requires further validation.
成功の流れからエージェントの失敗を追跡する
LLM ベースのエージェント システムの障害の原因特定、つまり、障害の軌跡のどのステップがタスクの失敗の原因となったかを特定することは、これらのシステムのデバッグと改善にとって重要です。既存のアプローチは、計算コストがかかるプロンプトベースのパイプラインに依存するか、ステップレベルのエラー注釈を使用した障害軌跡に関する事後トレーニングが必要ですが、収集コストが高く、拡張が困難です。私たちは、実用的な障害属性モデルは軽量であり、障害データに対するステップレベルの監視なしでトレーニング可能であるべきだと主張します。この目的を達成するために、教師なしの失敗の帰属、つまり、成功した軌跡のみをトレーニングし、失敗した軌跡が与えられた推論時にエラー ステップを特定することに取り組みます。我々は、この問題を神経制御微分方程式による 1 クラス学習として捉え、潜在空間における成功軌道の動的パターンをモデル化する OAT を提案します。推論時に、失敗した軌跡の各ステップには、成功した軌跡で学習されたダイナミクスからの偏差に基づいて異常スコアが割り当てられ、その後、それを使用して一連のエラー ステップを形成します。わずか 100 件の成功した軌跡でのトレーニングによる実験では、OAT はプロンプトベースのベースラインより 200 ~ 5000 $\times$ 高速であり、同時にドメイン内および分布外のデータセットの両方で一貫してそれらを上回り、それぞれ +20% および +7% の F1 スコアを示し、OAT がエージェント システムの障害を診断するための有望かつ効率的な方向であることを示しています。
原文 (English)
Tracing Agentic Failure from the Flow of Success
Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approaches either rely on prompting-based pipelines, which are computationally expensive, or require post-training on failure trajectories with step-level error annotations, which are costly to collect and difficult to scale. We argue that a practical failure attribution model should be lightweight and trainable without step-level supervision on failure data. To this end, we address unsupervised failure attribution, i.e., training exclusively on successful trajectories and identifying error steps at inference time given a failure trajectory. We propose OAT, which casts this problem as one-class learning with neural controlled differential equations, modeling the dynamical pattern of successful trajectories in latent space. At inference time, each step in a failure trajectory is assigned an anomaly score based on its deviation from the dynamics learned on successful trajectories, which is then used to form a set of error steps. With training on only 100 successful trajectories, experiments show that OAT is 200--5000 $\times$ faster than prompting-based baselines, and, at the same time, consistently outperforms them in both in-domain and out-of-distribution datasets with +20% and +7% F1 scores, respectively, demonstrating that OAT is a promising and efficient direction for diagnosing agentic system failures.
長さバイアス下での精度と正規化精度: 分析、ガイドライン、およびベイジアン代替法
条件付きの対数確率によって候補の完了をランク付けする多肢選択ベンチマークは、長さのバイアスに悩まされます。対数確率はトークンの合計であるため、実際には、長い回答は短い回答に比べてペナルティを受ける傾向があります。一般的な緩和策は、完了の長さによってスコアを正規化することですが、このヒューリスティックは頻繁に過剰修正を行い、代わりに長い回答へのバイアスを導入することが経験的に示されています。まず、これらのスコアリング ルールを分析し、標準精度と長さ正規化精度がいつ適切であるか、およびその長さのバイアスが完了長の分布にどのように依存するかを特徴付けます。この分析を動機として、解答の長さに対する明示的な事前事前条件の下で各候補の事後確率を計算するスコアリング ルールである \emph{ベイジアン精度} を導入し、それによって線形長の影響を除去します。ベイジアン精度は、尤度ベースの多肢選択評価のドロップイン代替品であり、追加のフォワードパスを必要とせず、ベンチマークおよび少数ショット設定全体で、標準精度および長さ正規化精度の両方よりも経験に基づく長さバイアスが一貫して低くなります。
原文 (English)
Accuracy and Normalized Accuracy under Length Bias: Analysis, Guidelines, and a Bayesian Alternative
Multiple-choice benchmarks that rank candidate completions by conditional log-probability suffer from a length bias: because log-probabilities sum over tokens, longer answers tend to be penalized relative to shorter ones in practice. A common mitigation is to normalize scores by completion length, but we show empirically that this heuristic frequently over-corrects, introducing a bias toward longer answers instead. We first analyze these scoring rules, characterizing when standard and length-normalized accuracy are appropriate and how their length biases depend on the distribution of completion lengths. Motivated by this analysis, we introduce \emph{Bayesian accuracy}, a scoring rule that computes the posterior probability of each candidate under an explicit prior over answer length, thereby removing linear length effects. Bayesian accuracy is a drop-in replacement for likelihood-based multiple-choice evaluation, requires no additional forward passes, and consistently exhibits lower empirical length bias than both standard and length-normalized accuracy across benchmarks and few-shot settings.
1B パラメータを超えるマルチモーダル感情言語モデルは本当に必要ですか?
マルチモーダル大規模言語モデル (MLLM) の最近の進歩により、マルチモーダル感情認識 (MER) のパフォーマンスが大幅に向上し、ビデオ、オーディオ、言語などを共同モデリングすることで解釈可能な記述の生成が可能になりました。ただし、これらのパフォーマンスの向上には、多くの場合、モデル パラメーター サイズの増加 (例: 少なくとも 7B) が伴い、同時に高い計算コストが発生し、推論効率が低下するため、ロボットやモバイルなどのリソースに制約のあるプラットフォームでのリアルタイム展開が妨げられます。デバイス。これにより、基本的な疑問が生じます。高品質の MER には、1B パラメーターを超えるマルチモーダル MER モデルが本当に必要なのでしょうか。この論文では、より大きなモデルが本質的に必要であるという仮定に異議を唱え、知識の蒸留を通じてより優れた、より迅速なマルチモーダル感情の理解と認識を実現する軽量 MER フレームワーク (Light-MER と呼ばれる) を提案します。これは、強力で大規模な教師モデルから軽量のサブビリオンパラメータの生徒モデルに知識を転送することができ、展開効率を大幅に向上させながら、豊かなマルチモーダルな感情推論と認識を維持することを目指しています。具体的には、知識伝達を強化するための 2 つの新しい最適化戦略を導入します。(1) スライス ワッサーシュタイン距離と隠れ状態アライメントを組み合わせた新しい最適輸送損失、(2) 学生モデルの学習能力をさらに強化することを目的とした、MER のパフォーマンスと効率のバランスをとる GRPO に基づく新しい複数報酬最適化戦略。 9 つのベンチマーク データセットに対する広範な実験により、Light-MER が推論効率を大幅に向上させながら最先端のパフォーマンスを達成することが実証されました。これは、将来の研究において、小規模でマルチモーダルな感情言語モデルの強力な可能性を強調しています。コードは https://github.com/GAIR-Lab/Light-MER で入手できます。
原文 (English)
Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. However, these performance improvements are often accompanied by an increase in model parameter size (e.g, at least 7B), which simultaneously incurs high computational costs and reduces inference efficiency, thereby hindering real-time deployment on resource-constrained platforms such as robots and mobile devices. This raises a fundamental question: do we really need the multimodal MER model larger than 1B parameters for high-quality MER? In this paper, we challenge the assumption that larger models are inherently necessary and proposes a lightweight MER framework (called Light-MER), which achieves better and faster multimodal sentiment understanding and recognition through knowledge distillation. It can transfer knowledge from a strong, large-scale teacher model to a lightweight sub-billion-parameter student model, aiming to preserve rich multimodal emotion reasoning and recognition while substantially improving deployment efficiency. Specifically, we introduce two new optimization strategies to enhance knowledge transfer: (1) a new optimal transport loss that combines Sliced Wasserstein Distance with hidden-state alignment, and (2) a new multi-reward optimization strategy based on GRPO that balances MER performance and efficiency, aimed at further enhancing the learning capabilities of student models. Extensive experiments on nine benchmark datasets demonstrate that Light-MER achieves state-of-the-art performance while significantly improving inference efficiency. This highlights the strong potential of small multimodal emotion language models for future research. Code is available at https://github.com/GAIR-Lab/Light-MER.
採点者を採点するのは誰ですか?自己改善する LLM エージェントのための、共進化する評価指標とスキル
自己進化するエージェント システムは、独自のスキルを作成、修正、廃止することで改善しますが、そのようなループはすべて、信頼できる評価基準がすでに存在するという隠れた前提に基づいています。実際のアプリケーションの多くではそうではありません。私たちは3つの主張をします。まず、メトリクスは \emph{evolved} 可能です。私たちのメトリクス ループは、完全な進化ライフサイクルの下で小さな欠点検出器の構成を検索し、10 項目のアンカーされた参照セットと一致するようにトレーニングされ、ラベルのない出力に対するコンセンサスによって正規化され、決して読み取られることのない保持されたアンカーに対して監査され、不透明な判定ではなく透明で検査可能なメトリクスを生成します。第 2 に、勝てる指標が存在しないため、指標は正確な指標があれば可能だったものを回復しつつあり、ライフサイクル管理スキル ループと指標の共進化である \emph{Double Ratchet} がそれを実現します。コード生成 (MBPP+)、エンタープライズ テキストから SQL (Spider~2.0-Snow)、および参照不要のレポート生成全体にわたって、同じものによって達成されたホールドアウト上昇率の 88 ~ 110\% を維持します。スキル ループは、グラウンド トゥルースまたは利用可能な最良のルーブリックによって駆動されます。第三に、安全性はアンカー規律と外部監査によってもたらされます。アンカー ガードを削除すると、メトリクスは空の検出器に折りたたまれますが、ライフサイクルを削除するとそうではありません。そして、進化したスキルがレポートのルーブリックを操作すると、独立した審査員がそれをキャッチし、1 つの検出器がそれを修復し、タスクを意識した審査員が、決定されたペアの 77% で進化前のベースラインよりも進化した出力を優先しました。私たちは、信頼できる自動検証機能が存在しない場合には、この障害を想定したアーキテクチャが正しいデフォルトであると主張します。
原文 (English)
Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists. In many real applications it does not. We make three claims. First, metrics can be \emph{evolved}: our metric loop searches compositions of small drawback detectors under a full evolutionary lifecycle, trained to agree with a ten-item anchored reference set, regularized by consensus over unlabeled outputs, and audited against a held-out anchor it never reads, yielding a transparent, inspectable metric rather than an opaque judge. Second, since no metric exists to beat, the yardstick is recovering what an accurate metric would have enabled, and \emph{Double Ratchet}, our co-evolution of the metric with a lifecycle-managed skill loop, does so: across code generation (MBPP+), enterprise text-to-SQL (Spider~2.0-Snow), and reference-free report generation, it retains 88--110\% of the held-out lift achieved by the same skill loop driven by ground truth or the best available rubric. Third, safety comes from anchor discipline plus outer audits: removing anchor guards collapses the metric into a vacuous detector while removing the lifecycle does not; and when evolved skills gamed the report rubric, an independent judge caught it, one detector repaired it, and a task-aware judge then preferred the evolved outputs over the pre-evolution baseline in 77\% of decided pairs. We argue this failure-expecting architecture is the right default wherever no reliable automatic verifier exists.
視覚言語モデル推論における視覚アクセス境界
思考連鎖 (CoT) プロンプトは、ビジョン言語モデル (VLM) のテスト時間スケーリング戦略として広く使用されていますが、VLM がより長い推論トレースを生成するときに何が拡張されるのかは依然として不明です。 CoT は画像トークンへの継続的なアクセスを必要とするのか、それとも、主にフォワード パスの早い段階で既に利用可能になった視覚情報を基に動作するのかを尋ねます。レイヤーの深さと生成時間に沿って、生成されたトークン クエリからイメージ トークン キーへの注意をマスクする因果的介入であるビジュアル アクセス スイープを導入し、タスクの精度を維持する最小アクセス領域としてビジュアル アクセス境界 (VAB) を定義します。 Qwen2.5-VL および InternVL3 の 6 つのモデル構成にわたって、CoT なしの直接応答と CoT プロンプトの両方が有限の VAB を示します。 14B および 38B スケールの Qwen2.5-VL-32B および InternVL3 では、CoT が非 CoT フルアクセス ターゲットに対して評価される場合、実質的に長い世代にも関わらず、その VAB 層は最大 2 層だけ非 CoT 境界と異なります。これは、CoT が推論トレース全体で直接イメージ トークン アクセスを延長することによって主にパフォーマンスを向上させるのではなく、イメージ由来の隠れ状態情報に対する言語側の計算を拡張することによってパフォーマンスを向上させることを示唆しています。さらに、CoT ゲインが知覚的読み出しによって制限されることを示します。 CoT は、クエリされた視覚属性がモデルによって確実に読み取れる場合には役立ちますが、その読み出しが信頼できない場合には役に立ちません。シンボリック属性のオラクルは、グラウンドトゥルース属性がテキストとして提供されると CoT によってカウントが向上することを示し、一方、単一オブジェクトのプローブ対デコードのチェックは、ハード属性が隠れた状態から線形に回復可能であるものの、モデル自体が出力するのは難しいことを示しています。これらの分析を組み合わせると、カウントではなく読み出しにボトルネックが生じます。
原文 (English)
Visual Access Boundaries in Vision-Language Model Reasoning
Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass. We introduce Visual Access Sweep, a causal intervention that masks attention from generated-token queries to image-token keys along layer depth and generation time, and define the Visual Access Boundary (VAB) as the minimal access region that preserves task accuracy. Across six model configurations from Qwen2.5-VL and InternVL3, both no-CoT direct answering and CoT prompting exhibit finite VABs. In Qwen2.5-VL-32B and InternVL3 at 14B and 38B scales, when CoT is evaluated against the no-CoT full-access target, its VAB layer differs from the no-CoT boundary by at most two layers, despite substantially longer generations. This suggests that CoT does not primarily improve performance by prolonging direct image-token access throughout the reasoning trace, but by extending language-side computation over image-derived hidden-state information. We further show that CoT gains are constrained by perceptual readout. CoT helps when the queried visual attribute can be reliably read out by the model, but not when that readout is unreliable. A symbolic-attribute oracle shows that CoT can improve counting once ground-truth attributes are supplied as text, while a single-object probe-vs-decode check shows that hard attributes can be linearly recoverable from hidden states yet difficult for the model itself to output. Together, these analyses place the bottleneck at readout rather than counting.
神経可塑性トレーニング環境としての人間と AI エージェントの相互作用
AI エージェントとの対話は、日常のデジタル ライフで最も頻繁に行われるアクティビティの 1 つになっています。アシスタントとの会話、コーディング副操縦士との作業、画像の生成のいずれであっても、対話は共通の反復ループに従います。つまり、リクエストが発行され、結果が返され、評価され、リクエストが修正されます。私たちは、このループが接触イベントの高頻度のストリームであること、つまり結果が人と出会い、意図的な評価の前に条件反射が発動する瞬間であることを観察し、日常のエージェントの相互作用を認識されない神経可塑性の訓練環境にしています。結果が期待外れになると、焦り、完璧主義、フラストレーション、自己批判といった反応的なパターンが繰り返し引き起こされ、活動依存性のシナプス可塑性のもとで、中断されない各サイクルが長期的な増強を通じて根底にある経路を深めます。したがって、通常のエージェントの使用は、それが引き起こすまさにパターンを静かに強化する可能性があります。私たちは、同じトレーニング環境が逆効果になる可能性があることを提案します。条件付けされた反応パターンを、短い調節ギャップを開く認知前の感情トーンによって活性化される物理的なニューロン経路として扱うことにより、そのギャップで、反応性の再プロンプトの代わりに人が舞台裏で観察を行うフレームワークを開発します。つまり、カスケードが完了せず、長期的な抑制が経路を強化するのではなく、神経プロセスが動作するのを観察します。我々は、この実践を 3 つの観察層と 2 つの適用モードを通じて特徴付けます。1 つは既存のツールを変更する必要のないユーザーガイド付きモード、もう 1 つはギャップでの観察をサポートするように通常のエージェントが軽く構成されているエージェント支援モードです。私たちは、生成イメージプロンプトを通じてフレームワークを説明し、単一のイライラするセッションが、観察されるかどうかに関係なく、行動的にはほぼ同じであるにもかかわらず、神経学的には反対であることを示します。
原文 (English)
Human-AI Agent Interaction as a Neuroplastic Training Environment
Interaction with AI agents has become one of the most frequent activities of everyday digital life. Whether conversing with an assistant, working with a coding copilot, or generating images, the interaction follows a common iterative loop: a request is issued, a result returned, appraised, and the request revised. We observe that this loop is a high-frequency stream of contact events -- moments at which a result meets a person and a conditioned response may fire before deliberate appraisal -- making everyday agent interaction an unrecognised neuroplastic training environment. When a result disappoints, reactive patterns of impatience, perfectionism, frustration, and self-criticism are repeatedly evoked, and under activity-dependent synaptic plasticity each uninterrupted cycle deepens the underlying pathway through long-term potentiation. Ordinary agent use may thus quietly strengthen the very patterns it provokes. We propose that the same training environment can be engaged to the opposite effect. Treating conditioned reactive patterns as physical neurone paths -- activated through a pre-cognitive feeling tone that opens a brief regulatory gap -- we develop a framework in which, at that gap, in place of the reactive re-prompt, a person performs behind-the-scenes observation: watching the neural process operate so the cascade does not complete and long-term depression weakens the path rather than potentiation strengthening it. We characterise this practice through three layers of observation and two modes of application: a user-guided mode requiring no change to existing tools, and an agent-assisted mode in which an ordinary agent is lightly configured to support observation at the gap. We illustrate the framework through generative image prompting, showing how a single frustrating session is behaviourally nearly identical whether or not it is observed, yet neurologically opposite.
Hempel の統計的曖昧性問題の解決と Causal AI
この論文は、統計法則から矛盾する予測が導出される、帰納統計推論における統計的曖昧さというカール ヘンペルの長年の問題を扱います。このような予測を回避するために、カール ヘンペルは、推論に使用される統計法則に対する最大特異性 (RMS) の要件を提案しました。 Wesley Salmon、Alberto Coffa、James Fetzer によって行われた RMS の改良の分析により、次のような最大限に具体的な統計法則の定義が導かれました。「適切な説明の法則に似た前提は、その存在の有無がその説明現象の発生に違いをもたらす性質のすべてと唯一を特定しなければならない。」しかし、この定義に基づく統計的曖昧さの問題の解決策の証明はありませんでした。背景コンテキスト全体で確率を高めるナンシー・カートライトの原因の定義を使用し、因果規則の概念を導入します。次に、統計的に関連するすべての情報を組み込むことによって、これらの因果関係ルールを段階的に改良する特別な意味論的確率的推論手順を定義します。この手順により最大特異的因果関係 (MSCR) が得られ、そこから導出された予測が一貫していることが証明されます (定理 1)。これにより、統計的な曖昧さの問題が解決されます。意味論的確率的推論手順は、因果 AI や因果機械学習などの新しい分野で使用できる確率的因果学習システムを提供します。彼らは、複雑なシステム内の因果関係を理解するためのツールとして、因果推論を基本的に研究しています。 RMS に類似したプロパティについては、まだ議論中です。 RMS に関連するいくつかの概念 (不変特徴学習、不変因果予測、および擬似関連) が考慮されます。
原文 (English)
Solution of the Hempel's statistical ambiguity problem and Causal AI
This paper addresses Carl Hempel's longstanding problem of statistical ambiguity in inductive-statistical inference, in which contradictory predictions are derived from statistical laws. To avoid such predictions, Carl Hempel proposed the Requirement of Maximal Specificity (RMS) for the statistical laws used in the inference. An analysis of the RMS refinements made by Wesley Salmon, Alberto Coffa, and James Fetzer led to the following definition of maximally specific statistical laws: "the lawlike premises of an adequate explanation must specify all and only those properties whose presence or absence made a difference to the occurrence of its explanandum-phenomenon." However, there was no proof of a solution to the statistical ambiguity problem based on this definition. We use Nancy Cartwright's definition of causes that raise probabilities across background contexts, and then introduce the concept of Causal Rules. Then we define a special semantic probabilistic inference procedure that incrementally refines these causal rules by incorporating all statistically relevant information. This procedure yields Maximally Specific Causal Relationships (MSCRs), for which we prove (Theorem 1) that predictions derived from them are consistent. This resolves the statistical ambiguity problem. The semantic probabilistic inference procedure provides a probabilistic causal learning system, which may be used in such new areas as Causal AI and Causal Machine Learning. They fundamentally explore causal inference as a tool for understanding cause-and-effect relationships within complex systems. Properties similar to RMS remain under discussion. Several notions related to RMS are considered: invariant feature learning, invariant causal prediction, and spurious association.
自律的で微調整不要の臨床症状検出のためのマルチエージェント システム: 開発および検証研究
臨床ノートには、患者を治療に導く兆候や症状の多くが含まれていますが、この情報が構造化された領域に到達することはほとんどありません。既存の抽出アプローチは、誤検知を生成するコンテキストに依存しないルールか、大幅な微調整を必要とする教師ありモデルに依存しています。手動によるプロンプト エンジニアリングや微調整を行わずに、臨床コンセプトの抽出プロンプトを自律的に作成および最適化するマルチエージェント システムである Pythia を紹介します。ローカルでホストされたオープンウェイト モデル上で実行される Pythia は、ローカル インフラストラクチャ上に臨床メモを保持し、開発セットの感度と特異性を使用してプロンプトを選択します。私たちは、387 人の患者を表す 400 の臨床ノートからの 72 の兆候と症状にわたって厳選された辞書と Pythia を比較しました。開発セット (n=300) と検証セット (n=100) は、コンセプトごとに独立して分割されました。 Pythia は、語彙集の平均感度 0.82 および 0.76 と比較して、平均感度 0.76 および特異度 0.95 を達成し、直接比較可能な 62 の概念のうち 20 について両方の指標で語彙集と同等またはそれを上回りました。辞書がすべてのノートをポジティブとラベル付けした 14 の概念について、Pythia は、用語のテキストでの言及ではなく、現在時制の患者に起因する所見を要求することで、平均特異度 0.97 を回復しました。特異性は開発から検証まで、有病率全体にわたって最小限の低下で移行しましたが、感度の移行は有病率 5% 以下で弱まり、有病率 2% 以下では平均ギャップ 0.25 に達しました。同じ開発セットで概念ごとに微調整された BERT 分類器は、平均感度 0.23 を達成しましたが、普及率が約 5% 未満の概念では感度がゼロに崩壊しました。これらの調査結果は、自律的で微調整不要のプロンプト最適化により、ローカル インフラストラクチャへの展開可能性を維持しながら、開発から検証まで効果的に一般化する症状抽出プロンプトを生成できることを示唆しています。
原文 (English)
A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study
Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reaches structured fields. Existing extraction approaches either rely on context-insensitive rules that generate false positives or on supervised models that require substantial fine-tuning. We present Pythia, a multi-agent system that autonomously writes and optimizes extraction prompts for clinical concepts without manual prompt engineering or fine-tuning. Running on a locally hosted open-weights model, Pythia keeps clinical notes on local infrastructure and selects prompts using development-set sensitivity and specificity. We compared Pythia with a curated lexicon across 72 signs and symptoms from 400 clinical notes representing 387 patients. Development (n=300) and validation (n=100) sets were partitioned independently for each concept. Pythia achieved mean sensitivity of 0.76 and specificity of 0.95, compared with 0.82 and 0.76 for the lexicon, and matched or exceeded the lexicon on both metrics for 20 of 62 directly comparable concepts. For 14 concepts where the lexicon labeled every note positive, Pythia recovered mean specificity of 0.97 by requiring a present-tense, patient-attributed finding rather than any textual mention of a term. Specificity transferred from development to validation with minimal degradation across prevalences, whereas sensitivity transfer weakened below 5% prevalence, reaching a mean gap of 0.25 below 2% prevalence. A BERT classifier fine-tuned per concept on the same development set achieved mean sensitivity of 0.23 and collapsed to zero sensitivity for concepts below roughly 5% prevalence. These findings suggest that autonomous, fine-tuning-free prompt optimization can produce symptom extraction prompts that generalize effectively from development to validation while remaining deployable on local infrastructure.
MemOps: 長期的な会話におけるライフサイクル メモリ操作のベンチマーク
長期記憶は、長期にわたるマルチセッションの対話にわたってユーザーに同行する LLM ベースのエージェントの基礎的な機能となっています。しかし、既存のベンチマークは、そのようなメモリをほぼ独占的に下流の質問応答を通じて評価し、最終的な回答の正しさだけをスコアリングします。このブラックボックスの定式化では、関連する事実の導入の欠落、操作を間違ったターゲットにバインドする、修正後の古い値に依存するなど、メモリ障害のさまざまな原因が混同されます。その結果、一貫性のない、または安全でない記憶状態に依存しているにもかかわらず、正解を認定することができます。この論文では、長期にわたる動的インタラクションにおいて、記憶は事実の静的な集合ではなく、記憶すること、忘れること、更新すること、反映すること、およびそれらの構成を含む明示的な操作のライフサイクルであると主張します。 MemOps は、会話型メモリを一連のライフサイクル操作として再定式化し、トリガー、ターゲット、スコープ、状態遷移、および裏付けとなる証拠を指定する構造化トレースで各メモリ イベントを表すベンチマークです。制御可能な生成パイプラインは、これらの操作をタスク指向の長い会話に埋め込み、隣接証拠と長いコンテキスト設定の両方で評価される 6 つのカテゴリの操作レベルのプローブとともにゴールド操作トレースを生成します。 MemOps は、ロングコンテキスト、検索ベース、パラメトリックおよびマネージド メモリ システム全体にわたって、最終応答の精度だけでは隠蔽されている障害モードを解きほぐし、現在のシステムが依然として均一な信頼性から程遠いことを明らかにします。たとえば、セッションレベルの取得はターンレベルの取得よりも優れていますが、ロングコンテキストモデルは、順序付けられたメモリ状態の軌跡を再構築するのが依然として著しく弱いままです。これらの結果は、長期記憶の評価を、最終的な回答スコアリングから、解釈可能な操作レベルの診断へと移行させます。
原文 (English)
MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions. Existing benchmarks, however, evaluate such memory almost exclusively through downstream question answering, scoring only the correctness of a final answer. This black-box formulation conflates the heterogeneous causes of memory failure, such as missing the introduction of a relevant fact, binding an operation to the wrong target, or relying on stale values after a correction. As a result, it can credit correct answers despite their reliance on inconsistent or unsafe memory states. In this paper, we argue that, in dynamic long-horizon interactions, memory is not a static collection of facts but a lifecycle of explicit operations, including remembering, forgetting, updating, reflecting, and their compositions. We introduce MemOps, a benchmark that reformulates conversational memory as a sequence of lifecycle operations and represents each memory event with a structured trace specifying its trigger, target, scope, state transition, and supporting evidence. A controllable generation pipeline embeds these operations into long, task-oriented conversations and produces gold operation traces together with six categories of operation-level probes, evaluated under both adjacent-evidence and long-context settings. Across long-context, retrieval-based, parametric and managed-memory systems, MemOps disentangles failure modes that final-answer accuracy alone conceals, revealing that current systems remain far from uniformly reliable. For instance, session-level retrieval outperforms turn-level retrieval, and long-context models remain notably weak at reconstructing ordered memory-state trajectories. These results move long-term memory evaluation from final-answer scoring toward interpretable, operation-level diagnosis.
パラメータ化された行動マルコフ決定プロセスのための知識と勾配に基づく強化学習
この論文では、パラメータ化されたアクションのマルコフ決定プロセス (PAMDP) における強化学習を研究します。このプロセスでは、各決定は記号アクションと数値パラメーターで構成されます。このような設定では、強化学習アルゴリズムは通常、ワンショット推定器を使用してパラメーターを決定するため、トレーニング サンプルが非効率になります。ほとんどの PAMDP 環境では、明示的ではあるが不完全な知識 (ルール、安全制約、エキスパートヒューリスティックなど) が利用可能ですが、それが強化学習エージェントのトレーニングのサンプル効率を高めるために直接使用されることはほとんどありません。私たちはこのギャップに踏み込み、新しい神経記号知識および勾配誘導強化学習 (KGRL) アルゴリズムを提案します。 KGRL は、Datalog 知識ベースのドメイン知識を使用して、特定の状態に適用可能なアクションと実行可能なパラメーターのセットを導き出します。これにより、適用できないアクションを決定空間から取り除き、残りのアクションのパラメータ空間を制約することができます。次に、勾配ベースのパラメータ調整ループを使用して、エージェントのトレーニングおよび展開中に最適なパラメータを推定します。 KGRL は、アクティブ化されたルールを軌跡に沿って記録することにより、アクションの刈り込みとパラメータの制約に関するローカルな手順の説明をさらに提供します。全体として、KGRL は、トレーニング中のサンプル効率を高めながら、エージェントの探索と展開を実行可能かつ制約を意識した決定に向けて導きます。 KGRL は、サンプル効率とエピソードリターンの両方において、PAMDP の最先端の RL ベースラインを上回ります。
原文 (English)
Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes
In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes their training sample inefficient. Though in most PAMDP environments explicit but incomplete knowledge (e.g., rules, safety constraints, or expert heuristics) is available, it is rarely directly used to increase the sample-efficiency of training Reinforcement Learning agents. We step into this gap and propose our novel Neuro-Symbolic Knowledge- and Gradient-Guided Reinforcement Learning (KGRL) algorithm. KGRL uses domain knowledge in a Datalog knowledge base to derive the set of applicable actions and feasible parameters for a given state. This allows it to prune non-applicable actions from the decision-space and constrain the parameter spaces of the remaining actions. We then use a gradient-based parameter refinement loop to estimate the optimal parameters during training and deployment of the agent. By recording activated rules along the trajectory, KGRL additionally provides local procedural explanations on the pruning of actions and constraining of parameters. Overall, KGRL guides the agent's exploration and deployment toward feasible and constraint-aware decisions, while increasing sample efficiency during training. KGRL outperforms state-of-the-art RL baselines for PAMDPs in both, sample efficiency and episodic return.
FormalAnalyticGeo: マルチモーダル解析幾何問題生成のためのニューラルシンボリックベースのフレームワーク
数学的推論は、マルチモーダル大規模言語モデル (MLLM) の急速な進歩により大幅な進歩を遂げていますが、解析幾何学は、主に注釈付きのサンプルが不足しているため、ほとんど研究されていません。既存のダイアグラム生成アプローチは、解析ジオメトリに苦労しています。テンプレート メソッドは制約駆動のレイアウトを処理できず、生成モデルには注釈付きの円錐曲線を正しくレンダリングするための幾何学的精度が不足しています。私たちは、マルチモーダルな解析幾何学問題を完全に自動生成するためのスケーラブルなフレームワークである FormalAnalyticGeo を紹介します。形式言語の厳密性を活用して、CDL (条件記述言語) を中心としたフレームワークを設計します。これは、自由形式の問題テキストと、符号付き距離フィールド (SDF) エンジンを介した正確な図のレンダリングを橋渡しする形式的な中間表現です。このフレームワークは、4 つの特殊な LLM コンポーネントを順番に使用します。さまざまな解析幾何学問題を生成するジェネレーター、SDF ベースのレンダリング用に各問題を CDL に変換するフォーマライザー、レンダリングされたダイアグラムのビジョンベースの測定を通じてグランドトゥルースの答えを抽出する測定器、および 3 つの段階で出力をチェックする品質検証器です。 Quality Verifier からの構造化されたフィードバックにより自動再試行が行われ、人間による注釈の必要性を排除する閉ループが形成されます。 FormalAnalyticGeo を大規模に適用すると、7K を超える検証済みのマルチモーダル問題のデータセットである AnalyticGeo7K が生成され、それぞれに位置合わせされたテキスト、図、正式な注釈、グラウンド トゥルースが含まれます。実験によると、生成された問題は、グラウンド トゥルース相対誤差の中央値 0.70\% に達し、回答の 82.3\% が正確なシンボリック解の 5\% 以内に収まります。私たちのフレームワークとデータセットは一般に公開されます。
原文 (English)
FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation
Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remains largely underexplored, primarily due to the scarcity of annotated samples. Existing diagram generation approaches struggle with analytic geometry: template methods cannot handle constraint-driven layouts, and generative models lack the geometric precision to render annotated conic curves correctly. We present FormalAnalyticGeo, a scalable framework for fully automatic generation of multimodal analytic geometry problems. Leveraging the rigor of formal languages, we design the framework around CDL (Condition Description Language), a formal intermediate representation that bridges free-form problem text with precise diagram rendering via a Signed Distance Field (SDF) engine. The framework employs four specialized LLM components in sequence: a Generator that produces diverse analytic geometry problems, a Formalizer that converts each problem into CDL for SDF-based rendering, a Measurer that extracts ground-truth answers through vision-based measurement on the rendered diagrams, and a Quality Verifier that checks outputs at three stages. Structured feedback from the Quality Verifier drives automatic retry, forming a closed loop that eliminates any need for human annotation. Applying FormalAnalyticGeo at scale yields AnalyticGeo7K, a dataset of over 7K verified multimodal problems, each with aligned text, diagram, formal annotation, and ground truth.Experiments show that the generated problems achieve a median ground-truth relative error of 0.70\%, with 82.3\% of answers falling within 5\% of the exact symbolic solution. Our framework and dataset will be publicly released.
抵抗して更新: インセンティブ対応 LLM の反事実報告の調整
調整された言語モデルは、証拠のないインセンティブ圧力の下で日常的に誤った報告をします。つまり、自信のあるユーザーに同意したり、ユーザーの内部信念が変わっていない場合でも確信度を誇張したりします。我々は、これを内部インセンティブ互換性(IC)の失敗として位置づけ、モデルのレポートを因果関係の契約に保持する反事実レポートメディエーターを学習および認定する方法を提示します。つまり、禁止された影響(圧力、威信、スタイル変更)に対して不変であり、ライセンスされた影響(本物の証拠)に応答します。抵抗と更新という 2 つの要求は、反対方向に引っ張られます。私たちはそれらを、既知の事後分布を使用したベイジアン・ウィットネス・ベンチマークで研究します。このベンチマークでは、同じユーザーの意見の相違が、純粋に述べられたソースの信頼性によって認可された証拠または禁止された圧力となります。我々は、(i) プローブの精度ではなく交換介入によって、ほぼ直交で独立して制御可能な、回答、信頼度、警告に関する低ランクのレポート座標を因果的に特定し、(ii) 反事実的にインセンティブが中立化されたコンテキストの下でモデル自身のレポートを参照する、トレーニング不要の反事実レポート座標 (CRC) クランプを導入します。ウィットネス ベンチマークでは、2 パス クランプは耐性と 1.00 の更新を合わせて達成 (Wilson 95% CI [0.99,1.00])、展開されたソリューションではなく、構築可能な参照の下での因果関係の証明書です。グローバルなデコードとステアリングでは、単一パラメータのトレードオフが示されます。出力レベルの微調整は、両方が列挙されている場合にのみ、両方の目的と一致します。抵抗のみのトレーニングは証拠への反応性を失います。デプロイ可能なシングルパス コンパイルには非可逆性があります (0.73/0.97)。メカニズムとクランプは 3 つのモデル ファミリにわたって再現され、自然なおしゃべりベンチマーク (SycophancyEval) に移行されます。私たちの貢献はインターフェイスと認証方法です。内部 IC の構造プリミティブとしてのアクティベーション レベルの反事実的インセンティブ不変性です。
原文 (English)
Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when their internal belief is unchanged. We cast this as a failure of internal incentive-compatibility (IC) and present a method for learning and certifying counterfactual report mediators that hold a model's reports to a causal contract: invariant to forbidden influences (pressure, prestige, restyling) and responsive to licensed ones (genuine evidence). These two demands, resist and update, pull in opposite directions. We study them on a Bayesian-witness benchmark with known posteriors, in which the same user disagreement is licensed evidence or forbidden pressure purely by stated source reliability. We (i) causally identify, by interchange interventions rather than probe accuracy, low-rank report coordinates for answer, confidence, and caveat that are near-orthogonal and independently controllable, and (ii) introduce a training-free counterfactual report-coordinate (CRC) clamp that references the model's own report under a counterfactually incentive-neutralized context. On the witness benchmark the two-pass clamp attains resist and update of 1.00 jointly (Wilson 95% CI [0.99,1.00]), a causal certificate under a constructible reference, not a deployed solution. Global decoding and steering show a single-parameter tradeoff; output-level fine-tuning matches both objectives only when both are enumerated; resist-only training loses evidence-responsiveness. The deployable single-pass compilation is lossy (0.73/0.97). The mechanism and clamp reproduce across three model families and transfer to a natural sycophancy benchmark (SycophancyEval). Our contribution is the interface and certification method: activation-level counterfactual incentive-invariance as a structural primitive for internal IC.
沈黙による勝利: LLM 計画評価における削除の非単調性、自律的悪用、および型付き状態ゲート
計画評価者は、戦略計画が明確でなくなったことに対して報酬を与えることができます。この論文では、LLM によって生成されたベンチャー ルートの段階的な期待値スコアラーにおける失敗について研究します。命題 1 は、内部遷移を削除する一方で、その先行者を再ターゲットし、下流の値を保持することによるスコア変化を示します: Delta_k = (prod_{i<k} p_i)[c_k + (1 - p_k)R_{k+1}]。凍結された 26 経路コホートでは、57 個の許容される欠失すべてが分析上の同一性およびしきい値の兆候と一致し、すべての経路に少なくとも 1 つのスコア改善欠失がありました。スコアを求めるオプティマイザーは、ルートの再構築は許可されていますが、エクスプロイト メカニズムには通知されていませんでしたが、21/26 のルートでベースラインを上回るカバーされていない構造を発見しました。 GATE は、26/26 の沈黙ルートと 0/26 の正直な停止ルートのスコア公開を拒否しました。拒否の後、次の改訂では 47/54 がカバーされた構造に修復され、厳密にカバーされた改善は 1/26 から 13/26 に増加しました。アダプティブ コンパイラを意識した共著者は、レジストリと出所の境界を明らかにしました。v1/v1.5 の 4 つの条件すべてで義務チャネル回避は 6/6 のままでしたが、デルタインデックス付きコストフロアは、セマンティックな完全性を確立することなく、ビートオネストルートを 6/6 から 3/6 に、サイレンスによる資金提供可能性を 5/6 から 0/6 に削減しました。必要な作業を省略したという理由だけで計画のスコアが向上した場合、計画は改善されていません。評価によって不作為のインセンティブが生まれました。 PCSC は、モデルを介した型付き状態レコード上のポストホック省略スプライスを検出し、無効化します。テストした協調設定では、GATE は単なるポストホック フィルターではなく、決定論的な検索整形制約として機能します。任意の LLM 生成戦略の意味上の完全性や現実世界の品質は検証されません。
原文 (English)
Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation
Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer for LLM-generated venture routes. Proposition 1 gives the score change from deleting an interior transition while retargeting its predecessor and retaining downstream value: Delta_k = (prod_{i<k} p_i)[c_k + (1 - p_k)R_{k+1}]. On a frozen 26-route cohort, all 57 admissible deletions matched the analytic identity and threshold sign, and every route had at least one score-improving deletion. A score-seeking optimizer, allowed to restructure routes but not told the exploit mechanism, found baseline-beating uncovered structures in 21/26 routes. GATE refused score release for 26/26 silenced routes with 0/26 honest suspensions; after refusal, 47/54 next revisions repaired to a covered structure, and strict covered improvement rose from 1/26 to 13/26. An adaptive compiler-aware co-author exposed the registry-provenance boundary: obligation-channel evasions remained 6/6 across all four v1/v1.5 conditions, while delta-indexed cost floors reduced beat-honest routes from 6/6 to 3/6 and fundability-by-silence from 5/6 to 0/6 without establishing semantic completeness. If a plan scores better only because it omits necessary work, the plan did not improve; the evaluation created an omission incentive. PCSC detects and neutralizes post-hoc omission splices over model-mediated typed-state records. In the cooperative setting tested, GATE acts as a deterministic search-shaping constraint, not merely a post-hoc filter. It does not verify the semantic completeness or real-world quality of arbitrary LLM-generated strategies.
アンサンブル決定 MCTS のための動的リソース割り当て
シミュレーション ベースのアルゴリズムは、ランダム性や隠された情報が重要な要素を持つ敵対的なボード ゲームなど、不確実性の高い環境に特に適しています。特に、そのようなドメインでは、いくつかのモンテカルロ ツリー検索 (MCTS) バリアントが一般的に使用されます。この論文では、動的なリソース割り当てに 2 つの軸を導入する、アンサンブル決定 MCTS の一連の拡張機能を提案します。まず、Dynamic Number of Determinizations は、これまでの検索の動作に応じて、現在使用されている決定ツリーの数を増減します。 2 番目の動的シミュレーション割り当てでは、シミュレーション間の決定を使用して、シミュレーション予算を決定ツリー間で不均一に分割し、潜在的に最良の知識獲得が得られるツリーを選択します。ベンチマーク ドメインとして、Jaipur、Lost Cities、Splendor の 3 つの人気のあるテーブルトップ ゲームを使用しました。提案した拡張機能を反復ベースおよび時間ベースの設定でテストしたところ、特定の構成ではアルゴリズムの強度が統計的に有意に向上することがわかりました。
原文 (English)
Dynamic Resource Allocation for Ensemble Determinization MCTS
Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games with significant elements of randomness and hidden information. In particular, several Monte Carlo Tree Search (MCTS) variants are commonly used in such domains. In this paper, we propose a series of enhancements for Ensemble Determinization MCTS, introducing two axes for dynamic resource allocation. First, Dynamic Number of Determinizations, increases or decreases the number of currently used determinization trees depending on the behavior of so-far search. Second, Dynamic Simulation Allocation, splits the simulation budget nonuniformly across the determinization trees, using simulation-to-simulation decisions to choose the tree with potentially the best knowledge gain. As benchmark domains, we used three popular tabletop games: Jaipur, Lost Cities, and Splendor. Testing our proposed enhancements in iteration- and time-based settings showed that particular configurations yield a statistically significant increase in the algorithm's strength.
凍結離散拡散言語モデルを使用したオーディオネイティブ音声認識
自動音声認識は、一度に 1 つのトークンを発行する自己回帰デコーダーによって支配されています。代わりに、離散拡散言語モデルが音声を書き起こし、少数のノイズ除去ステップで並行して書き起こし全体を洗練できるかどうかを尋ねます。最近の拡散言語モデルに一般的な吸収マスク スキームではなく、均一なランダム トークンの離散拡散によってテキストを生成する 26B の専門家混合モデルである DiffusionGemma のオーディオ ネイティブ インターフェイスをトレーニングします。フリーズした Whisper エンコーダは音響機能を提供し、軽量のプロジェクターはそれらをモデル埋め込み空間にマッピングし、低ランクのアダプタはフリーズしたバックボーンを新しいモダリティに対応させます。約 4,200 万のパラメータがトレーニングされます。これはバックボーンの 0.16 パーセントに相当します。音声の勾配は、すでに音声を無視した注意を介してのみプロジェクターに到達するため、自然なトレーニングの目的は音声を接地することができないことがわかります。凍結された出力ヘッドを通じて適用されるコネクショニストの時間分類損失は、この行き詰まりを打開します。結果として得られたモデルは、LibriSpeech テストクリーンで単語誤り率 6.6% に達し、発話の長さに関係なく、およそ 8 つの並列ステップで文字起こしし、6 つの言語でトレーニングされた単一のアダプターを使用します。ここでは英語、ヒンディー語、中国語について評価します。
原文 (English)
Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 26B mixture-of-experts model that generates text by uniform, random-token discrete diffusion rather than the absorbing-mask scheme common to recent diffusion language models. A frozen Whisper encoder supplies acoustic features, a lightweight projector maps them into the model embedding space, and low-rank adapters let the frozen backbone attend to the new modality. About 42M parameters are trained, which is 0.16 percent of the backbone. We find that the natural training objectives fail to ground the audio because their gradient reaches the projector only through attention that has already dismissed it. A connectionist temporal classification loss applied through the frozen output head breaks this deadlock. The resulting model reaches 6.6 percent word error rate on LibriSpeech test-clean, transcribes in roughly eight parallel steps regardless of utterance length, and uses a single adapter trained on six languages, which we evaluate here on English, Hindi, and Mandarin.
AI エージェントはタスクが単純であることを認識していますか?複雑さを意識した推論と実行に向けて
大規模言語モデル (LLM) エージェントは、複数ステップのエンジニアリングおよび情報学のワークフローをますます自動化していますが、タスクに実際にどれくらいの労力が必要かについて尋ねることはほとんどありません。彼らは多くの場合、最大コンテキスト優先戦略、つまりすでに確認したファイルと依存関係を再読み込みし、1 行の編集を小さなコードベースの監査に変えます。私たちは、欠けている機能はタスクを認識した実行範囲の推定、つまりタスクの難易度、本当に必要な情報、予算をコミットする前に信頼できる最短経路を判断することであると主張します。私たちは、最小十分実行とエージェント認知冗長率 (ACRR) を形式化し、E3 (推定、実行、拡張) を提案します。つまり、エージェントは初期動作点を推定し、実行可能な最小パスを実行し、検証が失敗した場合にのみ範囲を拡張します。 MSE ベンチ (機能制御シミュレーターでの 121 回の編集の決定論的ベンチマーク) では、E3 は最も強力なベースラインの 100% の成功に匹敵し、コストを 85%、トークンを 91%、検査されたファイルを 92% 削減し、さらに強力な適応型検索ベースラインを 16% 上回りました。利益は、保留された命令の文言と基本的にすべてのコストの重み付けに耐えます。コンパニオンのリアル モデル ハーネス (LLM-Case) は、実際のオープンソース ライブラリを編集するライブ gpt-4o エージェントへの影響を裏付けます。すべての候補パッチは、測定されたオラクルに対してプロジェクトの実際の pytest スイートを実際に実行することによってグレーディングされます。オーバーリードは穏やかですが現実であり、E3 は同等のタスクの成功において最も無駄がなく最速のポリシーです。その 1 つの欠点はプロバイダーのレート制限であり、間違った編集ではありません。私たちはこれを、展開されたエージェントの測定ではなく、実行の冗長性の制御された調査として組み立て、タスク認識型の実行を、エンジニアリングに基づいた AI (EGAI)、つまりタスクのエンジニアリング上の現実に取り組みを固定するエージェントに向けたステップとして位置づけています。フレームワークとベンチマークを公開します。
原文 (English)
Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution
Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires. They often follow a maximum-context-first strategy--re-reading files and dependencies they have already seen--turning a one-line edit into a small code-base audit. We argue the missing capability is task-aware execution-scope estimation: judging a task's difficulty, the information it truly needs, and the shortest reliable path before committing budget. We formalize minimum-sufficient execution and the Agent Cognitive Redundancy Ratio (ACRR), and propose E3 (Estimate, Execute, Expand): the agent estimates an initial operating point, executes a minimum viable path, and expands scope only when verification fails. On MSE-Bench--a deterministic benchmark of 121 edits in a capability-controlled simulator--E3 matches the strongest baseline's 100% success while cutting cost by 85%, tokens by 91%, and inspected files by 92%, and further beats a strong adaptive retrieval baseline by 16%; the gains survive held-out instruction wording and essentially every cost weighting. A companion real-model harness (LLM-Case) corroborates the effect on a live gpt-4o agent editing a real open-source library, with every candidate patch graded by actually running the project's real pytest suite against a measured oracle: the over-reading is milder but real, and E3 is the leanest and fastest policy at comparable task success--its one shortfall a provider rate-limit, not a wrong edit. We frame this as a controlled probe of execution redundancy, not a measurement of any deployed agent, and position task-aware execution as a step toward engineering-grounded AI (EGAI)--agents whose effort is anchored in the engineering reality of the task. We release the framework and benchmark.
参照せずに答える: AI 検索がウェブの経済的取引をどのように書き換えるか
検索エンジンは長い間、ユーザーをクエリから Web サイトに誘導することで Web に注目を集めてきました。 AI 検索では、情報ニーズを仲介者内で解決できるため、この仕組みが変わります。 URL レベルの Comscore US デスクトップ クリックストリームを使用して、ChatGPT と Google の情報検索機会を比較し、ChatGPT 検索アクセスの拡張を利用して従来の検索の置き換えを推定します。 ChatGPT が生成するアウトバウンド クリックは会話セッションのわずか 5.2% であり、Google の紹介率をはるかに下回っています。残りのクリックは、縮小された Google ストリームではなく、広告でサポートされているサイトから離れて、専門的な目的地に偏っています。アクセスの拡大により検索の使用が 9.4% 削減され、検索参照の損失は情報カテゴリで最大となっています。私たちの調査結果は、デジタル仲介における中心的な経済変化を特定しています。AI 検索は、オープン ウェブ上で検索、トラフィック、コンテンツ制作を結び付けてきた紹介取引を弱める一方で、仲介業者内部の情報ニーズを満たしている可能性があります。
原文 (English)
Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain
Search engines have long allocated attention on the web by routing users from queries to websites. AI search changes this arrangement because information needs can be resolved inside the intermediary. Using URL-level Comscore U.S. desktop clickstream, we compare ChatGPT and Google information-seeking occasions and exploit ChatGPT Search access expansions to estimate traditional search displacement. ChatGPT produces outbound clicks in only 5.2% of conversation sessions, far below Google's referral ratio. The remaining clicks are not a scaled-down Google stream: they skew toward specialized destinations and away from ad-supported sites. Wider access cuts search use by 9.4%, with search-referral losses largest for informational categories. Our findings identify a central economic shift in digital intermediation: AI search might satisfy information needs inside the intermediary while weakening the referral bargain that has linked search, traffic, and content production on the open web.
FAIR GraphRAG: セマンティック データ分析のための検索拡張生成アプローチ
検索拡張生成 (RAG) は、ドメイン固有の質問に回答する際の大規模言語モデル (LLM) の制限に対処します。 GraphRAG などのグラフベースの RAG アプローチは、ナレッジ グラフ (KG) 内の意味論的な関係をキャプチャすることで検索を強化します。 FAIR 原則 (検索可能性、アクセシビリティ、相互運用性、再利用性) は科学データ管理、特に医学などの複雑な領域で普及しつつありますが、既存の RAG アプローチには基礎となる知識リソースの構造化された FAIR 化が欠けています。この欠如により、これらのドメインでの FAIR 情報取得の可能性が制限されます。このギャップに対処するために、グラフベースの検索システムの基本ユニットとして FAIR デジタル オブジェクト (FDO) を統合する新しいフレームワークである FAIR GraphRAG を紹介します。各グラフ ノードは、コア データ、メタデータ、永続的な識別子、およびセマンティック リンクを組み込んだ FDO を表します。 LLM を活用して、スキーマの構築と、データ ソースからのコンテンツとメタデータの自動抽出をサポートします。このフレームワークは、技術的および臨床的関連性を確保するために、医師とコンピューター科学者によって共同設計されました。私たちは FAIR GraphRAG を消化器病学の生物医学データセットに適用し、RNA シーケンス データへの適用性を実証します。 FAIR GraphRAG は、FAIR 原則への準拠を保証するだけでなく、特にメタデータとオントロジー リンクを含む複雑なクエリについて、質問応答の精度、カバレッジ、説明可能性を大幅に向上させます。この研究は、FAIR データの実践とグラフベースの検索技術を組み合わせる実現可能性を示しています。私たちは、教育やビジネスなどの他の専門分野にも私たちのアプローチを適用できる可能性があると考えています。
原文 (English)
FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis
Retrieval-Augmented Generation (RAG) addresses the limitations of Large Language Models (LLMs) when providing responses to domain-specific questions. Graph-based RAG approaches, such as GraphRAG, enhance retrieval by capturing semantic relationships within knowledge graphs (KGs). While the FAIR principles (Findability, Accessibility, Interoperability, and Reusability) are becoming prevalent for scientific data management, especially in complex domains such as medicine, existing RAG approaches lack a structured FAIRification of the underlying knowledge resources. This lack limits their potential for FAIR information retrieval in these domains. To address this gap, we introduce FAIR GraphRAG, a novel framework that integrates FAIR Digital Objects (FDOs) as the fundamental units of a graph-based retrieval system. Each graph node represents an FDO that incorporates core data, metadata, persistent identifiers, and semantic links. We leverage LLMs to support schema construction and automated extraction of content and metadata from data sources. The framework was co-designed by physicians and computer scientists to ensure technical and clinical relevance. We apply FAIR GraphRAG to a biomedical dataset in gastroenterology, demonstrating its applicability to RNA-sequencing data. Beyond ensuring adherence to the FAIR principles, FAIR GraphRAG significantly improves question answering accuracy, coverage, and explainability, particularly for complex queries involving metadata and ontology links. This work shows the feasibility of combining FAIR data practices with graph-based retrieval techniques. We see potential for applying our approach to other specialized fields such as education and business.
ポイントインタイム言語モデルのスケーリング
無制限のインターネット コーパスでトレーニングされた大規模な言語モデルには、必然的に未来からの情報が埋め込まれ、金融や社会科学におけるバックテストや因果推論の妥当性を損なう先読みバイアスが導入されます。各暦日までに利用可能なテキストのみを対象としてトレーニングされたポイントインタイム言語モデルは、構築によってこの漏れを排除しますが、既存の取り組みでは通常、制約のないモデルに比べて大幅に遅れたモデルが生成されます。このパフォーマンスのギャップは規模を拡大することで大幅に縮小できることを示します。 FineWeb から時系列でフィルタリングされた 1 兆個のトークン上で最大 40 億個のパラメータを備えたデコーダ専用トランスフォーマーをトレーニングし、2013 年から 2024 年にわたる一連の月次モデル チェックポイントを構築します。さまざまな常識的推論と言語理解ベンチマーク全体で、私たちのモデルは、時間的に制限のないデータでトレーニングされた同等のサイズの主要なオープンウェイト モデル (Gemma-3-4B や LLaMA-7B など) のパフォーマンスに近づいていますが、いくつかのタスクではパフォーマンスのギャップが残っています。 LoRA による命令の微調整により、ダウンストリームの使いやすさがさらに向上します。データセット構築、トレーニング インフラストラクチャ、評価コードを含む完全なパイプラインをリリースし、再現可能なポイントインタイム言語モデリングを可能にし、厳密な時間的妥当性を必要とする研究アプリケーションをサポートします。
原文 (English)
Scaling Point-in-Time Language Models
Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences. Point-in-time language models--trained exclusively on text available up to each calendar date--eliminate this leakage by construction, but existing efforts typically produce models that lag substantially behind their unconstrained counterparts. We show that this performance gap can be substantially narrowed through scale. Training decoder-only transformers with up to 4 billion parameters on 1 trillion chronologically filtered tokens from FineWeb, we construct a sequence of monthly model checkpoints spanning 2013-2024. Across a range of common-sense reasoning and language understanding benchmarks, our models approach the performance of leading open-weight models of comparable size (e.g., Gemma-3-4B and LLaMA-7B) trained on temporally unrestricted data, although a performance gap remains on several tasks. Instruction fine-tuning via LoRA further improves downstream usability. We release the complete pipeline--including dataset construction, training infrastructure, and evaluation code--to enable reproducible point-in-time language modeling and to support research applications that require strict temporal validity.
非常に多くの意見、非常に多くの LLM: オープンエンド調査分析のための大規模言語モデルと従来の機械学習の比較
自由回答型調査は貴重な洞察を提供しますが、大規模な分析が難しいことで知られています。テキストを分類するために従来の機械学習を採用した以前の研究 (「So Many Responses, So Little Time: A Machine-Learning Approach to Analyzing Open-Ended Survey Data」) [1] に基づいて、この研究は、さまざまな大規模言語モデル (LLM) が NSSE 自由回答式アンケートの回答をどのように理解して分析するかを調査します。私たちは、OpenAI の GPT シリーズ、Twitter の roBERTa ベース モデル、Meta の LLaMA など、いくつかの最先端の LLMS に焦点を当て、感情分析やテーマ分類などのタスクにおいて、それらのパフォーマンスを以前の機械学習モデルと比較します。私たちの調査分析では、モデルの一致性、分類の精度、推論の解釈可能性を評価します。この調査結果により、現在の LLM は、分類精度、特に学生の回答における複雑な気分やテーマのパターンの理解において、古典的な機械学習モデルを日常的に上回っていることが明らかになりました。 LLM は優れた精度を持っていますが、予測をどのように明示的かつ一貫して正当化し、カテゴリ境界を適用するかという点で大きく異なります。これらの違いは、定性分析に LLM を使用する場合の重大なトレードオフを浮き彫りにします。つまり、予測強度の向上には、一貫性と説明可能性の問題が伴います。私たちの調査結果は、大規模な定性研究にさまざまな LLM を利用する利点と欠点を示しており、自動化と解釈の厳密さのバランスをとろうとしている研究者に実践的なアドバイスを提供しています。
原文 (English)
So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis
Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditional machine learning to classify text ("So Many Responses, So Little Time: A Machine-Learning Approach to Analyzing Open-Ended Survey Data") [1], this study investigates how different large language models (LLMs) understand and analyze NSSE open-ended survey responses. We focus on several cutting-edge LLMSs-OpenAI's GPT series, Twitter-roBERTa-base model, and Meta's LLaMA-and compare their performance to the previous machine learning models in tasks like sentiment analysis and thematic classification. Our research analysis assesses model agreement, classification accuracy, and interpretability of reasoning. The findings reveal that current LLMs routinely beat classic machine learning models in classification accuracy, particularly in understanding complex mood and theme patterns in student replies. While LLMs have superior accuracy, they differ greatly in how explicitly and consistently they justify their predictions and apply category boundaries. These distinctions highlight crucial trade-offs when using LLMs for qualitative analysis: increased predictive strength comes with issues in consistency and explainability. Our findings illustrate the benefits and drawbacks of utilizing various LLMs for large-scale qualitative research, and we provide practical advice for researchers looking to balance automation and interpretive rigor.
CANDI: ニッチ領域のコンテキストの調整 質問応答
医療診断や財務アドバイスなどの特殊な領域に大規模言語モデル (LLM) を導入するには、一般知識を超えた評価機能が必要です。従来の質問応答ベンチマークでは、これらの分野に必要な微妙な文脈の基礎、ユーザーの認識、ドメインの理解を捉えることができないことがよくあります。これに対処するために、CANDI-QA (Contextual Alignment for Niche Domains Question Answering) を導入します。これは、特殊な設定で正確でコンテキストに応じた、ユーザーに合わせた回答を提供することに関して LLM を評価する新しいデータセットです。 CANDI-QA は、専門家が厳選した質問と回答のペアを 2 つのカテゴリに構造化しています。(1) 正確な抽出を必要とする直接的な事実クエリである情報支援質問、(2) 実用的な洞察を生成するために状況推論を必要とするマルチホップ推論タスクである応用推論質問。私たちは、コンパクトなオープンソースから最先端の独自システムに至るまで、10 を超える多様な言語モデルを評価します。堅牢なベースラインとして、ニューラル検索とルールベースの推論を組み合わせた軽量の神経記号フレームワークである MTSS-Net を紹介します。私たちの調査結果は、ニッチな領域でコンテキストの整合性を達成するという深刻な課題を浮き彫りにし、コンテキストまたはシンボリック統合を強化しないと現在の LLM の限界を明らかにしています。最終的に、CANDI-QA は、コンテキスト認識言語モデルの研究を進めるための重要なベンチマークとして機能し、一か八かの分野向けの堅牢で信頼できる AI の開発を促進します。
原文 (English)
CANDI: Contextual Alignment for Niche Domains Question Answering
The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge. Traditional question-answering benchmarks often fail to capture the nuanced contextual grounding, user awareness, and domain understanding these fields require. To address this, we introduce CANDI-QA (Contextual Alignment for Niche Domains Question Answering), a novel dataset evaluating LLMs on delivering accurate, context-sensitive, and user-aligned answers in specialized settings. CANDI-QA features expert-curated question-answer pairs structured into two categories: (1) Information Assistance Questions, which are direct, factual queries requiring precise extraction, and (2) Applied Inference Questions, which are multi-hop reasoning tasks needing situational inference to generate actionable insights. We evaluate over ten diverse language models, from compact open-source to state-of-the-art proprietary systems. As a robust baseline, we present MTSS-Net, a lightweight neuro-symbolic framework combining neural retrieval with rule-based reasoning. Our findings highlight the profound challenges of achieving contextual alignment in niche domains, revealing the limitations of current LLMs without enhanced contextual or symbolic integration. Ultimately, CANDI-QA serves as a critical benchmark for advancing research in context-aware language models, stimulating the development of robust, trustworthy AI for high-stakes domains.
G-SHARE: ヒューマンファクターイベント診断のためのガイドラインベースの構造化推論フレームワーク
人的要因によるイベント診断は、原子力発電所の運転イベントから学ぶために不可欠ですが、その品質は、ナラティブレポートの専門家の解釈とガイドラインに基づいた推論に大きく依存します。既存のデータ駆動型またはワンショットの大規模言語モデルのアプローチには、構造化された推論が欠けていることが多く、正式な診断ガイドラインとの整合性が限られており、論理的に矛盾した結論が生成される可能性があります。この問題に対処するために、この研究では、CNNP の 9 ステップの人的要因イベント診断ガイドラインを多段階の診断パイプラインに運用化する、ガイドラインベースの構造化推論フレームワークである G-SHARE を提案します。このフレームワークは、証拠の抽出、段階的な診断推論、事後整合性修復で構成され、レポート証拠の明示的な使用、中間根拠の生成、診断出力の論理検証を可能にします。実際の人的要因事象レポートのデータセットは中国の原子力産業の情報源から構築され、各分野の専門家によって注釈が付けられたゴールドスタンダードのサブセットが評価に使用されました。結果は、G-SHARE がワンショット プロンプトや従来の機械学習ベースラインを大幅に上回っており、最も強力なバージョンが最高の全体的な精度とマクロ F1 を達成していることを示しています。アブレーションの結果はさらに、構造化された推論と一貫性の強制が、特に弱い刺激条件下での堅牢な診断にとって重要であることを示しています。この調査結果は、専門家の診断ガイドラインを監査可能な推論ワークフローに変換することの価値を実証し、安全性が重要な産業におけるインテリジェントな人的要因分析のための実用的な道筋を提供します。
原文 (English)
G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis
Human-factor event diagnosis is essential for learning from operational events in nuclear power plants, yet its quality depends strongly on expert interpretation of narrative reports and guideline-based reasoning.Existing data-driven or one-shot large language model approaches often lack structured reasoning, have limited alignment with formal diagnostic guidelines, and may generate logically inconsistent conclusions. To address this issue, this study proposes G-SHARE, a guideline-based structured reasoning framework that operationalizes the CNNP nine-step human-factor event diagnosis guideline into a multi-stage diagnostic pipeline.The framework consists of evidence extraction, stepwise diagnostic reasoning, and post-hoc consistency repair, enabling explicit use of report evidence, intermediate rationale generation, and logical validation of diagnostic outputs. A dataset of real human-factor event reports was constructed from Chinese nuclear industry sources, and a gold-standard subset annotated by domain experts was used for evaluation. Results show that G-SHARE substantially outperforms one-shot prompting and traditional machine learning baselines, with the strongest version achieving the best overall accuracy and macro-F1. Ablation results further indicate that structured reasoning and consistency enforcement are critical to robust diagnosis, especially under weak prompting conditions. The findings demonstrate the value of transforming expert diagnostic guidelines into auditable reasoning workflows, providing a practical pathway for intelligent human-factor analysis in safety-critical industries.
申し訳ありませんが、点字についてはどうすることもできません: 最先端の LLM におけるアクセシビリティの障害の解明
大規模言語モデル (LLM) は多くの言語タスクで強力に機能しますが、点字などの構造的に制約があり、アクセシビリティが重要なモダリティにおける LLM の機能はまだ不明です。私たちは、人間による注釈付きデータセットを使用して、双方向の韓国語点字翻訳に関する最先端の LLM を評価します。多言語で命令に調整されたモデルは、テキスト表現を介して点字に一般化できるという期待にもかかわらず、一貫して貧弱で不安定な出力と、人間の判断との実質的な不一致が見つかりました。これらの結果は、点字を意識したトークン化が欠落していることと、韓国語と点字のパターン間の整合性が弱いことを示しています。対照的に、同じデータに対する小規模モデル (T5-small) の教師あり微調整では、標準メトリクス (SacreBLEU、ChrF++、CER、BLEU、ROUGE-L、METEOR、CIDEr) 全体でゼロショットおよびプロンプト LLM ベースラインに対して大きく安定したゲインが得られます。私たちの調査結果は、現在のLLMの体系的な限界を明らかにし、タスク固有の適度な監視の有効性を実証しています。
原文 (English)
I'm Sorry, but I Can't Help with Braille: Revealing Accessibility Failures in State-of-the-Art LLMs
Large Language Models (LLMs) perform strongly on many language tasks, but their capability in structurally constrained, accessibility-critical modalities such as Braille remains unclear. We evaluate state-of-the-art LLMs on bidirectional Korean-Braille translation using a human-annotated dataset. Despite expectations that multilingual, instruction-tuned models can generalize to Braille via text representations, we find consistently poor, unstable outputs and substantial disagreement with human judgments. These results point to missing Braille-aware tokenization and weak alignment between Korean and Braille patterns. In contrast, supervised fine-tuning of a small model (T5-small) on the same data yields large and stable gains over zero-shot and prompted LLM baselines across standard metrics (SacreBLEU, ChrF++, CER, BLEU, ROUGE-L, METEOR, CIDEr). Our findings reveal a systematic limitation of current LLMs and demonstrate the effectiveness of modest task-specific supervision.
ロシアとウクライナの電報チャネル間の偽情報ナラティブ拡散のグラフベースの検出
オンライン コンテンツの拡大規模、急速な進化、言語の多様性により、ソーシャル メディア上の偽情報の物語を検出することは困難です。私たちは、弱い監視と伝播グラフ分析を組み合わせることにより、テレグラムエコシステム内の偽情報ナラティブを特定および分析するためのグラフベースのフレームワークを提案します。このアプローチは、意味的に関連する主張を物語レベルのクラスターに集約し、相互接続されたチャネル全体での拡散をモデル化します。これにより、ポストレベル分析だけでは捕捉するのが難しい、調整された物語の増幅を検出できるようになります。私たちの結果は、テキスト信号をネットワーク構造と統合することで、偽情報ナラティブを検出するためのスケーラブルな方法が提供され、大規模なメッセージング環境内で偽情報ナラティブがどのように伝播するかについての洞察が得られることを示しています。
原文 (English)
Graph-Based Detection of Disinformation Narrative Diffusion between Russian and Ukrainian Telegram Channels
Detecting disinformation narratives on social media is challenging due to the scale of amplification, rapid evolution, and linguistic variability of online content. We propose a graph-based framework for identifying and analyzing disinformation narratives in Telegram ecosystems by combining weak supervision with propagation graph analysis. The approach aggregates semantically related claims into narrative-level clusters and models their diffusion across interconnected channels. This enables the detection of coordinated narrative amplification that is difficult to capture through post-level analysis alone. Our results demonstrate that integrating textual signals with network structure provides a scalable method for detecting disinformation narratives and offers insights into how they propagate within large-scale messaging environments.
OmniPMNet: オムニクエリ ニューラル プロセスを介して離散型およびグリッド型の PM10 予測を橋渡し
粒子状物質 (PM10) の予測には、特に激しい砂嵐の際に、ステーション規模の精度と連続した空間フィールドの両方が必要です。化学輸送モデル (CTM) は格子状の予測を提供しますが、局所的な偏りを保持します。一方、グラフ ニューラル ネットワーク (GNN) は短いリードタイムで監視サイトを適切に追跡しますが、格子状の出力は生成しません。ここでは、共有空間表現内でこれら 2 つの予測タイプを調整する畳み込み条件付きニューラル プロセス (ConvCNP) ベースの融合モデルである OmniPM-Net を紹介します。地形を認識したガウス セットの畳み込みは、不規則な GNN 観測点の予測を規則的なグリッドに引き上げます。そこでは、マルチスケールの空間ソース アテンション (SSA) モジュールがそれらをコペルニクス大気監視サービス (CAMS) の予測とブレンドします。次に、共有されたオムニクエリの読み出しにより、この表現がデコードされ、観測点またはグリッド セルのいずれかで 108 時間にわたる一貫した PM10 予測が生成されます。 2024 年通年にわたって中国全土の 1,618 か所の大気質監視ステーションで評価された OmniPM-Net は、より強力な GNN ベースラインのステーション レベルの精度 (平均絶対誤差 21.14 対 22.00 ug/m3) と一致し、CAMS 平均絶対誤差を 30% 削減すると同時に、離散 GNN では不可能なグリッド フィールドを提供します。最も明らかな増加は、高濃度尾部であり、90 パーセンタイル MAE が GNN に対して 9%、CAMS に対して 25% 低下します。また、粉塵エピソード中には、進化する空間軌跡を追跡しながらカテゴリカル検出スキルが向上します。
原文 (English)
OmniPMNet: Bridging discrete and gridded PM10 forecasts via omni-query neural processes
Forecasting particulate matter (PM10) requires both station-scale accuracy and continuous spatial fields, especially during severe dust storms. Chemical transport models (CTMs) provide gridded forecasts but retain local biases, whereas graph neural networks (GNNs) track monitoring sites well at short lead times but do not produce gridded outputs. Here we present OmniPM-Net, a Convolutional Conditional Neural Process (ConvCNP)-based fusion model that reconciles these two forecast types within a shared spatial representation. A terrain-aware Gaussian set convolution lifts irregular GNN station forecasts onto a regular grid, where a multi-scale Spatial Source Attention (SSA) module blends them with Copernicus Atmosphere Monitoring Service (CAMS) forecasts; a shared omni-query readout then decodes this representation into consistent PM10 predictions at either stations or grid cells over a 108 h horizon. Evaluated across 1,618 air-quality monitoring stations throughout China over the full year of 2024, OmniPM-Net matches the station-level accuracy of the stronger GNN baseline (mean absolute error 21.14 versus 22.00 ug/m3) and reduces the CAMS mean absolute error by 30%, while simultaneously delivering the gridded fields that the discrete GNN cannot. Its clearest gains are in the high-concentration tail, where the 90th-percentile MAE falls by 9% relative to the GNN and 25% relative to CAMS, and during dust episodes, where it improves categorical detection skill while tracking the evolving spatial trajectory.
SeqGPT: マルチパネル複合構造の逆設計のための制約付き変換エージェント
連続ターゲット (積層や座屈パラメータなど) と個別の製造制約を一致させるために複合材料の積層シーケンスを最適化することは、特に数値最適化アプローチ (バイステップ、バイレベル構成) が使用される場合に複合材料設計で定期的に発生する、困難な組み合わせ逆問題を表します。マルチパネル構成では、この複雑さは、異なるパネルのスタッキング間のグローバルな互換性/連続性要件であるブレンディングによってさらに強化されます。この研究では、計算コストのかかる反復手法を置き換えるために開発された条件付き Transformer エージェントである SeqGPT を紹介します。世界的な継続性と建設による製造の実現可能性の両方を確保するために、私たちはハイブリッド神経記号解読戦略を実装しました。 SeqGPT は、ブレンディング ルールに違反するブランチが厳密にプルーニングされる、制約付きビーム検索をガイドする条件付き分布を予測します。 18 パネルの馬蹄ベンチマークの数値実験では、SeqGPT が進化的手法に匹敵する座屈性能でほぼ瞬時に解を生成し、最先端技術と比較して大幅な速度向上を実現することが実証されています。
原文 (English)
SeqGPT: A Constrained Transformer Agent for the Inverse Designof Multi-Panel Composite Structures
Optimizing composite stacking sequences to match continuous targets (e.g., Lamination or Buckling Parameters) with discrete manufacturing constraints represents a challenging combinatorial inverse problem that regularly occurs in composite design especially when numerical optimization approaches are used (bi-step, bi-level configurations). In multipanel configurations, this complexity is further intensified by blending, a global compatibility/continuity requirement between the different panel stackings. This study presents SeqGPT, a conditional Transformer agent developed to replace computationally expensive iterative methods. To ensure both global continuity and manufacturing feasibility by construction, we implemented a hybrid neurosymbolic decoding strategy. SeqGPT predicts a conditional distribution that guides a Constrained Beam Search, where any branch violating blending rules is strictly pruned. Numerical experiments on the 18-panel horseshoe benchmark demonstrate that SeqGPT generates solutions near-instantaneously with buckling performance comparable to evolutionary methods, offering a significant speed-up compared to the state of the art.
自己進化するエージェントに向けて: 遺伝子ネットワーク プログラミングのための人間からインスピレーションを得た適応型探索/活用フレームワーク
エージェント AI の最近の進歩は、説明可能で人間中心の非線形推論ワークフローの需要に後押しされて、グラフベースの手法にますます移行しています。顕著な例は、遺伝的ネットワーク プログラミング (GNP) です。これは、有向グラフを利用してエージェントが解釈可能な意思決定構造を進化させる自己進化型アルゴリズムです。ほとんどの進化的アルゴリズムと同様に、探査と活用の効果的なバランスをとることが GNP の重要な側面です。しかし、このトレードオフは、GNP の文献ではあまり注目されていません。このギャップに対処するために、私たちは人間の発達パターンからインスピレーションを得ています。子供たちは熟考よりも幅広い実験や行動を優先し、この傾向は年齢とともに逆転します。 GNP の判断ノードから熟慮へ、処理ノードから行動への遷移をマッピングすることにより、進化のプロセス全体を通じて探索と搾取のバランスを動的に制御する新しい適応フレームワークである Human-Inspired GNP (HGNP) を提案します。この方法は、新しい適応クロスオーバーおよび突然変異オペレーター、およびサイクル除去メカニズムで構成されます。 HGNP は進化プロセスを改善するだけでなく、ターゲット環境とその探索空間の特性に基づいて探索と活用のバランスを調整するためのフレームワークも提供します。このアプローチは、標準的な GNP の交叉確率や突然変異確率による調整よりも効果的です。この変更は一般的なもので、ほぼすべての GNP バリアントに適用できます。標準 GNP および最近導入された 2 つの GNP バリアントと統合し、Tileworld ベンチマークで評価した場合、HGNP はエージェントの戦略におけるパフォーマンスの大幅な向上を示しました。 HGNP と状況ベースの GNP (HGNP-SBGNP) を組み合わせると、全体的に最高の結果が得られました。
原文 (English)
Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming
Recent advancements in agentic AI have increasingly moved toward graph-based methods, driven by the demand for explainable, human-centered, and non-linear reasoning workflows. A prominent example is Genetic Network Programming (GNP), a self-evolving algorithm that utilizes directed graphs to evolve interpretable decision structures for agents. As in most evolutionary algorithms, effectively balancing exploration and exploitation is a key aspect of GNP. However, this trade-off has received limited attention in the GNP literature. To address this gap, we draw inspiration from human developmental patterns, where children prioritize broad experimentation and action over deliberation, with this tendency reversing with age. By mapping transitions between GNP's judgment nodes to deliberation and processing nodes to action, we propose Human-Inspired GNP (HGNP), a novel adaptive framework that dynamically regulates the exploration-exploitation balance throughout the evolutionary process. The method consists of novel adaptive crossover and mutation operators, and a cycle elimination mechanism. HGNP not only improves the evolutionary process but also provides a framework for adjusting the exploration-exploitation balance based on the characteristics of the target environment and its search space. This approach is more effective than tuning via crossover and mutation probabilities in standard GNP. The modifications are general and can be applied to almost all GNP variants. When integrated with standard GNP and two recently introduced GNP variants and evaluated on the Tileworld benchmark, HGNP demonstrated significant performance improvement in agents' strategy. The combination of HGNP with Situation-based GNP (HGNP-SBGNP) achieved the best overall results.
バーストスパイキングニューラルネットワーク
現在のスパイキング ニューラル ネットワーク (SNN) 研究の中心的な目標は、人工ニューラル ネットワーク (ANN) の低電力代替に向けて精度を向上させることです。この研究はさらに、この目標を実現するには精度だけでなく、入力の摂動下で正しい予測を維持する能力として定義される堅牢性も向上させる必要があると主張しています。私たちは、堅牢性を損なう既存の SNN 手法の 2 つの重要な問題を特定します。まず、バイナリ スパイクの活性化は、小さな摂動の下で大きな活性化状態の変化を引き起こす可能性があります。第 2 に、有効な重み制約が欠如しているため、ネットワーク出力が入力変動の影響をより受けやすくなります。この目的を達成するために、バースト強化スパイキング ニューロン (BSN) と動的重み制約 (DWC) メカニズムに基づいて構築されたバースト スパイキング ニューラル ネットワーク (BuSNN) を提案します。 BSN にはバースト ファイアリングが組み込まれており、段階的なスパイク パターンを提供します。このスパイク機構は、摂動によって引き起こされる活性化状態の遷移を緩和し、それによってロバスト性を強化します。 DWC は、アクティブ化状態に基づいて接続の重みにペナルティを課し、重みの大きさを効果的に削減し、精度を維持しながら堅牢性を向上させます。これらの堅牢性の効果をサポートする理論的分析を提供します。実験結果はさらに、CIFAR-10 などの小規模ベンチマークでは、BuSNN が精度と堅牢性の点で SNN と ANN の両方の対応物よりも優れていることを示しています。大規模な ImageNet では、MS ResNet-34 バックボーンを備えた BuSNN により、対応する SNN ベースラインと比べてトップ 1 の精度と破損耐性がそれぞれ 3.18% と 2.66% 向上します。スパイクベースのアクティベーションを使用しているにもかかわらず、BuSNN は 4 ビットのアクティベーション量子化 ANN ベースラインを上回り、ImageNet 上の 8 ビット ANN ベースラインに近づきます。また、SNN の低電力の利点も維持されます。この研究では、SNN の精度と堅牢性の問題を研究し、堅牢でエネルギー効率の高いアプリケーションでの実用性を高めます。
原文 (English)
Burst Spiking Neural Networks
A central goal of current Spiking Neural Network (SNN) research is to improve their accuracy toward becoming low-power alternatives to Artificial Neural Networks (ANNs). This work further argues that realizing this ambition requires improving not only accuracy but also robustness, defined as the ability to maintain correct predictions under input perturbations. We identify two key issues in existing SNN methods that undermine robustness. First, binary spiking activations can produce large activation-state changes under small perturbations. Second, the lack of effective weight constraints makes network outputs more sensitive to input variations. To this end, we propose Burst Spiking Neural Networks (BuSNNs), built upon Burst-enhanced Spiking Neurons (BSNs) and a Dynamic Weight Constraint (DWC) mechanism. BSNs incorporate burst firing to provide a graded spiking pattern. This spiking mechanism mitigates perturbation-induced transitions in activation states and thereby enhances robustness. DWC penalizes connection weights based on activation states, effectively reducing weight magnitudes and improving robustness while preserving accuracy. We provide theoretical analyses to support these robustness effects. Experimental results further show that, on smaller-scale benchmarks such as CIFAR-10, BuSNNs outperform both SNN and ANN counterparts in accuracy and robustness. On large-scale ImageNet, BuSNN with the MS ResNet-34 backbone further improves top-1 accuracy and corruption robustness over the corresponding SNN baseline by 3.18% and 2.66%, respectively. Despite using spike-based activations, BuSNNs surpass 4-bit activation-quantized ANN baselines and approach 8-bit ANN baselines on ImageNet. They also preserve SNNs' low-power advantage. This work studies the accuracy-robustness problem in SNNs, advancing their practical viability in robust and energy-efficient applications.
QDEvo: 自動ヒューリスティック設計のための多目的品質多様性フレームワーク
大規模言語モデル (LLM) と進化的計算の統合は、組み合わせ最適化における自動ヒューリスティック設計の強力なパラダイムとして浮上しています。しかし、既存のアプローチはモード崩壊に悩まされており、意味論的な多様性を欠き、完全なアルゴリズム空間を探索できない均質な集団に収束してしまいます。私たちは、Quality-Diversity Evolution (QDEvo) を提案します。これは、Quality-Diversity の最適化と LLM 駆動のヒューリスティック検索を統合する多目的フレームワークで、事前にトレーニングされたコード埋め込みを使用して意味的に多様なアルゴリズムの無制限のアーカイブを維持し、進化のプロセスを導くための階層的な自己反映を組み込みます。標準ベンチマークと実際の産業アプリケーションにわたる広範な実験により、QDEvo がハイパーボリュームと反転世代距離メトリクスの両方で最先端の手法を大幅に上回ることが実証されました。私たちのフレームワークは、高性能で計算効率が高く、意味的に多様であるヒューリスティックの発見を可能にし、複雑な最適化問題に対するソリューションの豊富なポートフォリオを実務者に提供します。
原文 (English)
QDEvo: A Multi-Objective Quality-Diversity Framework for Automated Heuristic Design
The integration of Large Language Models (LLMs) with evolutionary computation has emerged as a powerful paradigm for automated heuristic design in combinatorial optimization. However, existing approaches suffer from mode collapse, converging to homogeneous populations that lack semantic diversity and fail to explore the full algorithmic space. We propose Quality-Diversity Evolution (QDEvo), a multi-objective framework that integrates Quality-Diversity optimization with LLM-driven heuristic search, maintaining an unbounded archive of semantically diverse algorithms using pre-trained code embeddings and incorporating hierarchical self-reflection to guide the evolutionary process. Extensive experiments across standard benchmarks and real-world industrial applications demonstrate that QDEvo significantly outperforms state-of-the-art methods in both Hypervolume and Inverted Generational Distance metrics. Our framework enables the discovery of heuristics that are simultaneously high-performing, computationally efficient, and semantically diverse, providing practitioners with a rich portfolio of solutions for complex optimization problems.
AAAI-26 デュアルサブミッション: 新たな挑戦
同一または実質的に類似した論文が、相互引用や開示なしに 1 つ以上のアーカイブ会場に同時に提出される二重投稿は、AAAI 会議およびその他の科学出版会場にとってますます問題となっています。これらの投稿は査読システムの負担を増大させ、科学的記録を汚染します。 AAAI-26 審査プロセスの一環として、私たち (カンファレンス主催者) は、AAAI のメイントラック提出物を、審査期間が重複する他の 9 つのアーカイブ会場と比較しました。また、AAAI-26 メイン トラック内の二重投稿も検索しました。タイトルと要約の類似性評価を使用して、LLM ベースの重複評価ツールによる後続のトリアージのために類似性の高い論文のペアを優先し、その後、最も重大度の高いペアを手動でレビューしました。このようなペアを手動でレビューした結果、141 件の AAAI-26 メイントラック提出物が机上で却下されました。私たちは、二重投稿の大幅な増加について、将来の主催者やより広範な人工知能研究コミュニティに警告したいと考えています。完全に重複した投稿の発生率は検出が容易ですが、同じ投稿を説明するために異なる単語を使用している論文の数に比べて、検出に非常に時間がかかります。この現象の成長は、生成 AI ツールへのアクセスが増えることで促進される可能性があります。この課題に対処するための推奨事項として、(1) AAAI 複数投稿ポリシーを更新し、許容される慣行についてコミュニティに教育すること、(2) 提出が終了する前に二重投稿チェックツールを導入すること、(3) 二重投稿の発生率を減らすために会場全体で協力して一貫したポリシーと罰則を統一すること、(4) 堅牢な検出ツールの開発を加速するためにコミュニティ主導の敵対的チャレンジを作成することなどが挙げられます。
原文 (English)
AAAI-26 Dual Submissions: Novel Challenges
Dual submissions, in which identical or substantially similar papers are simultaneously submitted to one or more archival venues, without cross-citation or disclosure, are a growing problem for the AAAI Conference and other scientific publication venues. These submissions increase the burden on the peer-review system and pollute the scientific record. As part of the AAAI-26 review process, we (conference organizers) compared AAAI main-track submissions to nine other archival venues with overlapping review periods. We also searched for dual submissions within the AAAI-26 main track. We employed title+abstract similarity assessment to prioritize highly similar paper pairs for subsequent triage by an LLM-based overlap assessment tool, followed by manual review of the highest severity pairs. Manual review of such pairs led to the desk-rejection of 141 AAAI-26 main-track submissions. We seek to alert future organizers, and the broader artificial intelligence research community, to the enormous growth in dual submissions. The incidence of exact duplicate submissions, which are easy to detect, has been eclipsed by the number of papers that use different words to describe the same contribution, which are extremely time-consuming to detect. The growth in this phenomenon is likely facilitated by increasing access to generative AI tools. We include several recommendations for addressing this challenge, including (1) updating the AAAI Multiple Submission Policy and educating the community about acceptable practice, (2) having dual-submission checking tools in place before submissions close, (3) working across venues to converge on consistent policies and penalties to aid in reducing the incidence of dual submission, and (4) creating a community-driven adversarial challenge to accelerate the development of robust detection tools.
覚えていますか?メモリ中心のマルチモーダル AI に向けて
人間の記憶は再構築されるものであり、忠実な記録ではありません。現在のマルチモーダル LLM (MLLM) にはこの機能がありません。つまり、フリーズされたビジュアル エンコーダを通じて画像を処理し、ワンショットのテキスト出力を生成し、内部表現を破棄します。 MLLM に再構築メモリを導入する 3 段階のアーキテクチャである DoYouRemember を紹介します。(1) VQ-VAE が画像を個別のビジュアル トークンに圧縮し、(2) LoRA で微調整された LLM がビジュアル トークンとテキスト トークンを共同で処理し、(3) 拡散デコーダが LLM の隠れた状態から画像を再構築します。 1,000 個の 3D 顔の皮膚テクスチャ マップと 99,000 個のラベルなしの顔画像では、LLM の隠れ状態には回復可能な視覚情報がほぼゼロであることがわかりました。VQ-VAE トークン (LLM 前) から明確な再構成を生成する同じデコーダーは、LLM の隠れ状態 (LLM 後) から純粋なノイズを生成し、LLM が画像を理解しているが記憶していないことを示しています。バックプロパゲーション下での共有メモリ行列 M のトレーニングは、勾配のキャンセル (O(1/sqrt(N)) の減衰) により体系的に失敗します。 3 つの根本原因を特定し、ローカル EMA 更新により 3 つすべてが解決されることを示します。つまり、各イメージは 64 スロットのうち上位 8 スロットのみを更新し、スロット間の多様性を維持します。結果として得られる M (229K パラメータ、16 倍圧縮) は、未確認のテスト画像の VQ 上限に近づきます。 M の連続表現により VQ 量子化エラーが回避されるため、1,024 スロットへのスケーリングはそれを上回ります (LPIPS 0.056 対 0.071)。私たちはこれらの発見を情報理論の枠組みの下で統合します。つまり、記憶は非可逆圧縮であり、想起は解凍であり、幻覚は欠陥ではなく非可逆解凍に固有の特性です。
原文 (English)
Do You Remember? Toward Memory-Centric Multimodal AI
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability: they process images through a frozen visual encoder, produce a one-shot text output, and discard internal representations. We present DoYouRemember, a three-stage architecture introducing reconstructive memory into MLLMs: (1) a VQ-VAE compresses images into discrete visual tokens, (2) a LoRA-fine-tuned LLM jointly attends to visual and text tokens, and (3) a Diffusion Decoder reconstructs images from the LLM's hidden states. On 1,000 3D facial skin texture maps and 99,000 unlabeled facial images, we find that LLM hidden states contain approximately zero recoverable visual information -- the same Decoder producing clear reconstructions from VQ-VAE tokens (pre-LLM) produces pure noise from LLM hidden states (post-LLM), demonstrating that the LLM understands images but does not remember them. Training a shared memory matrix M under backpropagation systematically fails due to gradient cancellation (O(1/sqrt(N)) attenuation). We identify three root causes and show that local EMA updating resolves all three: each image updates only its top-8 slots out of 64, preserving inter-slot diversity. The resulting M (229K parameters, 16x compressed) approaches the VQ upper bound on unseen test images. Scaling to 1,024 slots surpasses it (LPIPS 0.056 vs. 0.071), as M's continuous representation avoids VQ quantization error. We unify these findings under an information-theoretic framework: memory is lossy compression, recall is decompression, and hallucination is an inherent property of lossy decompression rather than a defect.
主観的な期待効用の最大化に対する感受性: LLM 意思決定への応用例を含む方法論的研究
ラベル付けされた結果が不足していたり、コストがかかっていたり、運と混同されたりする場合、不確実性の下で行われた決定を評価することは困難です。私たちは、主観的期待効用 (SEU) の最大化を規定の基準として扱い、エージェントの適合性の段階的な尺度 (SEU 感度) を定義します。この車両は、SEU 値の代替車両の感度パラメーター $\alpha$ を備えたソフトマックス選択モデルです。寄与分は、$\alpha$ と信念パラメータおよび効用パラメータ $(\beta, \delta)$ の一連の識別可能性結果であり、有限サンプルの警告はそのままで、事前の予測チェック、パラメータ回復、およびシミュレーションベースのキャリブレーション (SBC) によって Stan で検証されています。不確実性選択のみのモデル $m_0$ では、期待効用ベクトル $\eta$ が与えられると $\alpha$ は特定可能であり、急激に回復しますが、$(\beta, \delta)$ についてはほとんど情報が得られません。事後関数はかろうじて収縮し、$\beta$-$\delta$ のトレードオフに集中します。拡張モデル $m_1$ では、$\delta$ は $\beta$ のない危険なブロックを介して原則として識別可能になりますが、現実的なサンプルサイズでの実際の回復利得は無視でき(一致した数の CI 幅の減少は 1% 未満)、そのブロックは一致した選択肢の数で $\alpha$ 精度の利得を検出しません。これらは 2 つの異なる現象です。$\delta$ の場合、識別可能性は現実的な $n$ での正確な推定可能性を意味しません。 $\alpha$ の場合、識別可能性は有限 $n$ 精度を支配するものについては沈黙しています。境界 SBC は、関節後方の情報が弱い場合でも両方のモデルで合格します。この境界線は私たちが正確に設定しています。 2×2 アプリケーション (GPT-4o と Claude 3.5 Sonnet、それぞれ保険金請求のトリアージとエルズバーグ型の骨壷を使用し、サンプリング温度をレバーとして使用) は、実際の LLM 選択データをエンドツーエンドで実行し、4 つのセルのうち 2 つで構造化された比較 $\alpha$ 効果を検出します。
原文 (English)
Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making
Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck. We treat subjective expected utility (SEU) maximization as a stated standard and define a graded measure -- SEU sensitivity -- of an agent's conformity to it. The vehicle is a softmax choice model with a sensitivity parameter $\alpha$ on SEU-valued alternatives; the contribution is a sequence of identifiability results for $\alpha$ and for belief and utility parameters $(\beta, \delta)$, validated in Stan via prior predictive checks, parameter recovery, and simulation-based calibration (SBC), with finite-sample caveats intact. In the uncertain-choice-only model $m_0$, $\alpha$ is identifiable given the expected-utility vector $\eta$ and sharply recovered, while $(\beta, \delta)$ are only weakly informed: the posterior barely contracts and concentrates on a $\beta$-$\delta$ trade-off. In the extended model $m_1$, $\delta$ becomes identifiable in principle via a $\beta$-free risky block, but its practical recovery gain at realistic sample sizes is negligible (matched-count CI-width reduction under 1%), and that block yields no detected $\alpha$-precision gain at matched choice count. These are two distinct phenomena: for $\delta$, identifiability does not imply precise estimability at realistic $n$; for $\alpha$, identifiability is silent about what governs finite-$n$ precision. Marginal SBC passes for both models even where the joint posterior is weakly informed -- a demarcation we make precise. A two-by-two application (GPT-4o and Claude 3.5 Sonnet, each on insurance-claims triage and Ellsberg-style urns, with sampling temperature as the lever) runs end-to-end on real LLM choice data, detecting a structured comparative $\alpha$ effect in two of four cells.
データサイエンスの数学
この本はデータサイエンスの数学的基礎について書かれています。 1. はじめに 2. 高次元の呪い、祝福、そして驚き 3. 特異値分解と主成分分析 4. 線形回帰と正則化 5. グラフ、ネットワーク、クラスタリング 6. 非線形次元削減と拡散マップ 7. ランダム射影による線形次元削減 8. データ サイエンスの最適化 9. 分類 10. Aディープラーニングの数学的入門 11. グラフラプラシアンの大サンプル限界 12. コミュニティ 13. メジャーの集中とガウス解析 14. 行列の集中不平等 15. 圧縮センシングとスパース性 16. 低ランク行列の回復
原文 (English)
Mathematics of Data Science
This book is about the mathematical foundations of data science. 1. Introduction 2. Curses, Blessings, and Surprises in High Dimensions 3. Singular Value Decomposition and Principal Component Analysis 4. Linear Regression and Regularization 5. Graphs, Networks, and Clustering 6. Nonlinear Dimension Reduction and Diffusion Maps 7. Linear Dimension Reduction via Random Projections 8. Optimization for Data Science 9. Classification 10. A Mathematical Introduction to Deep Learning 11. Large Sample Limit of Graph Laplacians 12. Community 13. Concentration of Measure and Gaussian Analysis 14. Matrix Concentration Inequalities 15. Compressive Sensing and Sparsity 16. Low-Rank Matrix Recovery
CARE-LoRA: メモリ効率の高い LoRA のための圧縮アクティベーション再構築
大規模な事前トレーニング済みモデルの規模が拡大し続けるにつれて、限られたメモリ予算の下でモデルを微調整することがますます困難になってきています。現在最も広く採用されているパラメータ効率の良い微調整 (PEFT) 手法の 1 つである低ランク適応 (LoRA) は、低ランク適応行列のみを最適化することでこの課題を軽減し、それによってトレーニング可能なパラメータの数を大幅に削減します。パラメータのオーバーヘッドが大幅に削減されたため、バックプロパゲーションのために保持されるアクティベーションが、LoRA の微調整中に残る主要なメモリ ボトルネックとして浮上しました。これに対処するために、私たちはデータを認識した Compressed Activation REconstruction フレームワークである CARE-LoRA を提案します。 LoRA の固有の投影構造を利用することにより、CARE-LoRA は、完全な入力アクティベーションを、LoRA ブランチによって自然に生成される低ランクの圧縮アクティベーションに置き換えます。さらに、フォワードパス中に追加の計算コストを無視して軽量の再構成行列を計算します。これはバックプロパゲーション中に勾配信号を再構築するために使用され、それによって LoRA 行列を完全にトレーニング可能に保ちます。多様なモデルとダウンストリーム タスクにわたる広範な実験により、CARE-LoRA が全体のメモリ フットプリントを大幅に削減しながら、標準 LoRA および代表的な LoRA バリアントと比較して競争力、またはさらに優れたパフォーマンスを実現することが実証されました。私たちのコードは https://github.com/fishandyu/CARE-LoRA で公開されています。
原文 (English)
CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA
As the scale of large pre-trained models continues to grow, fine-tuning them under limited memory budgets has become increasingly challenging. Low-Rank Adaptation (LoRA), currently one of the most widely adopted parameter-efficient fine-tuning (PEFT) methods, mitigates this challenge by optimizing only low-rank adaptation matrices, thereby greatly reducing the number of trainable parameters. With the parameter overhead substantially reduced, the activations retained for backpropagation have emerged as the primary remaining memory bottleneck during LoRA fine-tuning. To address this, we propose CARE-LoRA, a data-aware Compressed Activation REconstruction framework. By exploiting the inherent projection structure of LoRA, CARE-LoRA replaces the full input activation with the low-rank compressed activation naturally produced by the LoRA branch. It further computes a lightweight reconstruction matrix during the forward pass with negligible additional computation cost, which is used during backpropagation to reconstruct the gradient signal, thereby keeping LoRA matrices fully trainable. Extensive experiments across diverse models and downstream tasks demonstrate that, while substantially reducing the overall memory footprint, CARE-LoRA achieves competitive or even superior performance compared with standard LoRA and representative LoRA variants. Our code is publicly available at https://github.com/fishandyu/CARE-LoRA .
クエリの可視性が KV キャッシュの圧縮ランキングをどのように変えるか: 予算に応じた監査
KV キャッシュ圧縮方法は主に、圧縮前にコンテキストに追加されたクエリ (クエリ認識プロトコル) を使用して評価されます。しかし、圧縮された KV キャッシュの経済的なケースは再利用です。ドキュメントを一度圧縮すれば、それに対する将来の多くの質問に答えることができます。その展開では、質問が表示される前に、クエリに依存しない圧縮が行われる必要があります。我々は、3 つのオープン 7-9B モデル上の 3 つの自明なベースラインに対する 6 つの公開された圧縮方法の予算に応じた監査を提示します (RULER-8192 で 144,300 件のペア評価、LongBench で 40,800 件、全体で 50,000 件のリサンプル ペア ブートストラップ)。スコアリング ルールを除き、モデル、圧縮率、インスタンス、デコードなど、すべてが固定されます。 3つの発見。 (1) クエリの可視性によってランキングが変わります。不可知論的プロトコルの下では、共通のアテンション バックエンドを共有する 5 つの監査対象メソッドのうち、KeyDiff だけがベストオブ 3 の自明なベースラインを一貫して上回っており (36 セル中 31 セル)、最も広く導入されているメソッドである SnapKV は平均で「開始と最近のウィンドウを維持する」という点で負けています (-0.066)。 (2) 2 つのプロトコル間のメソッドごとのドロップは、各メソッドのスコアリング信号に対する質問の可視度に応じて順序付けされており、ソース コードで判読できます。SnapKV の Delta=+0.198 (質問は 64 トークンの観察ウィンドウ内にあります) から KeyDiff の Delta=+0.011 (スコアにはクエリ用語がまったく含まれていません) までです。
原文 (English)
How Query Visibility Changes KV-Cache Compression Rankings: A Matched-Budget Audit
KV-cache compression methods are predominantly evaluated with the query appended to the context before compression -- a query-aware protocol. Yet the economic case for a compressed KV cache is reuse: compress a document once, answer many future questions against it. In that deployment, compression must happen query-agnostic -- before any question is seen. We present a matched-budget audit of six published compression methods against three trivial baselines on three open 7-9B models (144,300 paired evaluations on RULER-8192; 40,800 on LongBench; 50,000-resample paired bootstrap throughout). Everything is held fixed -- model, compression ratio, instances, decoding -- except the scoring rule. Three findings. (1) Query visibility changes the rankings: under the agnostic protocol, of the five audited methods that share a common attention backend, only KeyDiff beats a best-of-3 trivial baseline consistently (31 of 36 cells), and the most widely deployed method, SnapKV, loses to "keep the start and the recent window" on average (-0.066). (2) The per-method drop between the two protocols is ordered consistently with how visible the question is to each method's scoring signal, legible in its source code: from Delta=+0.198 for SnapKV (the question sits inside its 64-token observation window) down to Delta=+0.011 for KeyDiff (its score contains no query term at all).
BattVAE-GP: 不確実性の定量化による長期的なバッテリー劣化の生成モデリング
長期にわたる物理ベースのバッテリー劣化シミュレーションは、機構に関する洞察を提供しますが、依然として計算コストが高くつくため、延長されたサイクル寿命にわたる動作条件の緻密な調査への使用は制限されます。ここでは、目に見えない充電率でのリチウムイオン電池の劣化軌跡の代理モデリングのための、ハイブリッド物理学と確率論的学習フレームワークを提案します。 PyBaMM の DFN/P2D 電気化学モデルで生成されたサイクル分解劣化データは、まず容量に合わせた電圧および微分特徴に変換され、変分オートエンコーダー (VAE) を使用してエンコードされます。結果として得られる 2 次元の潜在空間は、サイクルの進行と充電プロトコルの両方に従って劣化の軌跡を組織します。次に、サイクル数と C レートを入力変数として使用して、この潜在空間でスパース マルチタスク ガウス プロセス (GP) をトレーニングし、事後不確実性推定とともに潜在劣化ダイナミクスの連続補間を提供します。プロトコルレベルのホールドアウト評価の下で、潜在空間 GP は目に見えない C レート軌道を正確に回復し、トレーニング データのサポートと一致する不確実性挙動を示します。目に見えない内部 C レートでクエリが実行されると、モデルは、隣接するシミュレートされたプロトコル間に一貫して配置された潜在軌道を生成します。フリーズ VAE デコーダを介して GP で予測された潜在状態をデコードすると、スムーズな電圧容量の進化が得られます。一方、補助的な潜在状態対健康状態 (SOH) 予測器を介した GP 潜在の事後モンテカルロ伝播により、不確実性を考慮した SOH 推定値が得られます。したがって、提案されたBattVAE-GPフレームワークは、計算効率が高く不確実性を認識した長期劣化モデリングの代替手段を提供し、より豊富な動作条件や将来のシミュレーションと実験の融合に向けてバッテリーの状態予測を拡張するための構造化された基盤を提供します。
原文 (English)
BattVAE-GP: Generative Modeling of Long-Horizon Battery Degradation with Uncertainty Quantification
Long-horizon physics-based simulations of battery degradation provide mechanistic insight but remain computationally expensive, limiting their use for dense exploration of operating conditions over extended cycle life. Here, we propose a hybrid physics-probabilistic learning framework for surrogate modeling of lithium-ion battery degradation trajectories at unseen charging rates. Cycle-resolved degradation data generated with a DFN/P2D electrochemical model in PyBaMM are first transformed into capacity-aligned voltage and derivative features and encoded using a Variational Autoencoder (VAE). The resulting two-dimensional latent space organizes degradation trajectories according to both cycle progression and charging protocol. A sparse multitask Gaussian process (GP) is then trained in this latent space using cycle number and C-rate as input variables, providing continuous interpolation of latent degradation dynamics together with posterior uncertainty estimates. Under protocol-level holdout evaluation, the latent-space GP accurately recovers unseen C-rate trajectories and exhibits uncertainty behavior consistent with the support of the training data. When queried at unseen interior C-rates, the model generates latent trajectories that remain coherently positioned between neighboring simulated protocols. Decoding the GP-predicted latent states through the frozen VAE decoder yields smooth voltage-capacity evolution, while Monte Carlo propagation of the GP latent posterior through an auxiliary latent to State of Health (SOH) predictor provides uncertainty-aware SOH estimates. The proposed BattVAE-GP framework therefore offers a computationally efficient and uncertainty-aware surrogate for long-horizon degradation modeling, providing a structured basis for extending battery health prediction toward richer operating conditions and future simulation-experiment fusion.
リスクリライトを使用した一般化された分散フリーの半教師あり学習
一般的な半教師あり学習 (SSL) 手法は分布の仮定に依存しており、これに違反するとパフォーマンスが低下します。リスク書き換え手法である PNU 学習は、分布を使用しない代替手段を提供しますが、バイナリ分類に限定されており、その分散の最適性は不明のままです。この論文では、構成要素リスクの線形結合を使用し、PNU 学習を包含し、マルチクラス分類に拡張して不偏リスク推定量を構築する一般化されたフレームワークを提案します。達成可能な最小分散を導出し、非対称損失シナリオにおいて推定器が PNU よりも低い分散を達成できることを示しています。さらに、この分散の減少を学習パフォーマンスの向上に直接結び付ける一般化限界を確立します。これらの理論的な洞察に基づいて、バイナリおよびマルチクラスのベンチマークで経験的に既存のアプローチと同等またはそれを上回る 2 つの実用的な SSL メソッドを紹介します。
原文 (English)
Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite
Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated. While PNU learning, a risk rewriting method, offers a distribution-free alternative, it is restricted to binary classification and its variance optimality remains unclear. In this paper, we propose a generalized framework that constructs unbiased risk estimators using linear combinations of component risks, subsuming PNU learning and extending to multiclass classification. We derive the minimum achievable variance, demonstrating our estimator can attain lower variance than PNU in asymmetric loss scenarios. Furthermore, we establish a generalization bound directly linking this variance reduction to improved learning performance. Based on these theoretical insights, we introduce two practical SSL methods that empirically match or outperform existing approaches on binary and multiclass benchmarks.
BAT-RM: 臨床的に導入された子宮頸がん放射線療法の自動輪郭形成のための、領域認識型多方向マンバを備えた境界認識型トランスフォーマー
我々は、Sobelゲート境界アテンション、長距離コンテキスト用の線形時間多方向Mambaモジュール、および境界スケルトン誘導融合ゲートを統合するハイブリッドアーキテクチャであるRegion-Aware Mambaを備えたBoundary-Aware Transformer(BAT-RM)を核とした、子宮頸がんの放射線治療計画用に臨床的に導入されたエンドツーエンドの自動輪郭システムを紹介します。この設計は、長距離コンテキスト モデリングの線形時間の複雑さを実現し、完全な空間的自己注意による二次コストを回避します。完全なパイプラインは、複数の施設にわたるデータ収集、厳格な評価者間の品質保証、独立したコホートにおける外部検証、および Varian、RayStation、および Monaco とネイティブ互換性のある Web ベースの臨床インターフェイスに及びます。 4 つのベースラインに対して、BAT-RM は 7 つの解剖学的クラスにわたって優れたパフォーマンスを達成し、GTV や CTV などの標的体積、および直腸や膀胱などのリスクにさらされている臓器において統計的に有意な改善をもたらします。 13 人の放射線腫瘍医が参加した前向き多施設読者研究では、AI 支援により若手腫瘍科医の IoU が 0.899 から 0.965 に向上し、上級レベルの精度に近づき、輪郭形成時間を 80% 以上短縮できることが実証されました。このシステムはまた、専門家による相談率を削減し、読者間の一貫性を向上させ、効率と品質保証の両方の向上を反映しました。提携病院での臨床導入後、このシステムはスタッフを追加配置することなく患者の待ち時間を数日から数時間に短縮し、日常的な症例では同日または翌日に治療を開始できるようになりました。 BAT-RM は、放射線治療の需要が専門家の能力をはるかに超えているリソースに制約のある環境において、データのキュレーションから臨床展開に至る厳密な研究パイプラインが、測定可能な患者利益に直接つながる可能性があることを実証しています。
原文 (English)
BAT-RM: A Boundary-Aware Transformer with Region-Aware Multi-Directional Mamba for Clinically Deployed Cervical Cancer Radiotherapy Auto-Contouring
We present a clinically deployed end-to-end auto-contouring system for cervical cancer radiotherapy planning, anchored by the Boundary-Aware Transformer with Region-Aware Mamba (BAT-RM), a hybrid architecture that integrates Sobel-gated boundary attention, a linear-time, multi-directional Mamba module for long-range context, and a boundary-skeleton-guided fusion gate. This design achieves linear-time complexity for long-range context modeling, avoiding the quadratic cost of full spatial self-attention. The full pipeline spans multi-institutional data collection, rigorous inter-rater quality assurance, external validation in an independent cohort, and a web-based clinical interface natively compatible with Varian, RayStation, and Monaco. Against four baselines, BAT-RM achieves superior performance across seven anatomical classes, with statistically significant improvements in target volumes, including GTV and CTV, and in organs at risk such as the rectum and bladder. A prospective multi-center reader study involving 13 radiation oncologists demonstrated that AI assistance elevates junior oncologists' IoU from 0.899 to 0.965, approaching senior-level accuracy, while reducing contouring time by more than 80%. The system also reduced expert consultation rates and improved inter-reader consistency, reflecting gains in both efficiency and quality assurance. Following clinical deployment at a partner hospital, the system reduced patient wait times from days to hours without additional staffing, enabling same-day or next-day initiation of treatment for routine cases. BAT-RM demonstrates that a rigorous research pipeline, from data curation to clinical deployment, can translate directly into measurable patient benefit in resource-constrained settings where the demand for radiotherapy far exceeds specialist capacity.
希少な神経データに対するスケールを意識した注意: Sleep-EDF EEG 上の RG-Flow Transformer
脳野電位はスケールフリーです。そのパワースペクトルは $1/f^{\beta}$ の法則に従い、その非周期指数 $\beta$ が皮質の状態を追跡し、特に睡眠の深さは $\beta$ の変化です。明示的な繰り込み群 (RG) 誘導バイアスを備えた変換器 (学習可能な異常次元 $\gamma$、ブロックスピンの粗視化、およびエントロピーゲート同期ブリッジを備えたスケール認識ストリームに通常の自己注意を結合する RG-Flow 変換器) が、\emph{real, rarce} 上でパラメータが一致したバニラ変換器よりも優れているかどうかを尋ねます。脳波。厳密なリークフリーの被験者別ホールドアウトを備えた PhysioNet Sleep-EDF コーパスを使用して、(i) パラメータが一致したバニラトランスフォーマーおよび 5 クラス AASM 睡眠ステージング上の階層のみのアブレーションに対して RG-Flow をベンチマークし、(ii) データが不足しているときに予測される誘導バイアス クロスオーバーを探すために被験者ごとのデータ バジェットを調べ、(iii) かどうかをテストします。 RG-Flow が学習した $\gamma$ は、サンプル外で測定されたスペクトル指数 $\beta$ を追跡します。これはバニラ モデルには存在しない量です。 1 被験者を除外した相互検証の下で、$5$ の被験者と $5$ のシード間で、RG-Flow とバニラ トランスフォーマーは、5 クラスのステージングでは統計的に区別できません (77.3\% 対 77.0\% の精度、ペア $p=0.294$)、予測された希少データのクロスオーバーは現れません。データが限られたすべての予算でバニラが数値的に優れています。モデルを分けるのは解釈可能性です -- RG-Flow はサンプル外の連続スペクトル指数 ($\beta$-recovery $R^2 = 0.416$) を回復します。これはバニラ アーキテクチャには類似した機能がありません。
原文 (English)
Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG
Brain field potentials are scale-free: their power spectra follow a $1/f^{\beta}$ law whose aperiodic exponent $\beta$ tracks cortical state, and sleep depth in particular is a shift in $\beta$. We ask whether a transformer endowed with an explicit renormalization-group (RG) inductive bias -- the RG-Flow Transformer, which couples ordinary self-attention to a scale-aware stream with a learnable anomalous dimension $\gamma$, block-spin coarse-graining, and an entropy-gated synchronization bridge -- has an advantage over a parameter-matched vanilla transformer on \emph{real, scarce} EEG. Using the PhysioNet Sleep-EDF corpus with a strict leakage-free by-subject hold-out, we (i) benchmark RG-Flow against a param-matched vanilla transformer and a hierarchy-only ablation on 5-class AASM sleep staging, (ii) sweep the per-subject data budget to look for the inductive-bias crossover predicted when data are scarce, and (iii) test whether RG-Flow's learned $\gamma$ tracks the measured spectral exponent $\beta$ out-of-sample -- a quantity the vanilla model does not possess. Across $5$ subjects and $5$ seeds under leave-one-subject-out cross-validation, RG-Flow and the vanilla transformer are statistically indistinguishable on 5-class staging (77.3\% vs 77.0\% accuracy; paired $p=0.294$), and the predicted scarce-data crossover does not appear: vanilla is numerically ahead at every data-limited budget. What does separate the models is interpretability -- RG-Flow recovers the continuous spectral exponent out-of-sample ($\beta$-recovery $R^2 = 0.416$), a capability the vanilla architecture has no analogue for.
極端な臨床コード予測のためのグラフ制約ポリシー学習
臨床コード予測は、構造化されていない退院サマリーを、大規模でまばらで深い階層ラベル空間の ICD-10-CM リーフ コードにマッピングします。ほとんどのシステムは、タスクをフラットなマルチラベル分類として扱い、コードを個別にスコアリングし、まれなラベルに対して限定的なトレーニング信号を提供します。我々は、枝刈りされたコード階層にわたる有限水平決定プロセスとして ICD 予測を定式化する、グラフ制約付きトラバーサル ポリシーを提案します。単一の言語モデルはグラフをレベルごとに下降し、請求可能なリーフ コードに達するまで有効な子ノードを選択します。これにより、構造的に有効な出力が保証されながら、極端なマルチラベル予測が階層を意識したスパースなサブセット決定に変換されます。 MIMIC-IV 退院サマリーでは、当社の最良の監視ポリシーである SFT-1+ は、精選された 50 コードのサブセットで 0.709 マイクロ F1、15,761 コード空間全体で 0.527 マイクロ F1 を達成し、CAML、LAAT、PLM-ICD を含むフラット ベースラインを上回っています。完全な設定では、SFT-1+ は最も強いフラット ベースラインよりも 0.044 マイクロ F1 および 0.157 マクロ F1 改善されており、グラフ制約付き分解によりレア コードのボトルネックが軽減されることが示唆されます。制御された要因研究では、アーキテクチャ、トレーニング アルゴリズム、データ予算が評価されます。どちらのスケールでも、1 つの共有ポリシーが 3 人の専門家のカスケードに一致し、フルスペースのテスト ノートの 28 ~ 32% でのコンテキスト ウィンドウのオーバーフローを回避します。教師付き軌跡データを増やすことは、一貫してパフォーマンスを向上させる唯一の介入ですが、GRPO 強化学習には、一致したデータによる教師付き継続よりも利点はありません。これらの結果は、単純なグラフ制約ポリシー学習が、極端な臨床コード予測において、より複雑なフラット、カスケード、強化学習の代替手段よりも優れたパフォーマンスを発揮できることを示しています。
原文 (English)
Graph-Constrained Policy Learning for Extreme Clinical Code Prediction
Clinical code prediction maps unstructured discharge summaries to ICD-10-CM leaf codes in a large, sparse, and deeply hierarchical label space. Most systems treat the task as flat multi-label classification, scoring codes independently and providing limited training signal for rare labels. We propose a graph-constrained traversal policy that formulates ICD prediction as a finite-horizon decision process over a pruned code hierarchy. A single language model descends the graph level by level, selecting valid child nodes until billable leaf codes are reached. This converts extreme multi-label prediction into sparse, hierarchy-aware subset decisions while guaranteeing structurally valid outputs. On MIMIC-IV discharge summaries, our best supervised policy, SFT-1+, achieves 0.709 micro-F1 on a curated 50-code subset and 0.527 micro-F1 on the full 15,761-code space, outperforming flat baselines including CAML, LAAT, and PLM-ICD. In the full setting, SFT-1+ improves over the strongest flat baseline by 0.044 micro-F1 and 0.157 macro-F1, suggesting that graph-constrained decomposition mitigates the rare-code bottleneck. A controlled factorial study evaluates architecture, training algorithm, and data budget. Across both scales, one shared policy matches a three-specialist cascade while avoiding its context-window overflow on 28-32% of full-space test notes. Increasing supervised trajectory data is the only intervention that consistently improves performance, while GRPO reinforcement learning provides no benefit over supervised continuation with matched data. These results show that simple graph-constrained policy learning can outperform more complex flat, cascaded, and reinforcement-learning alternatives for extreme clinical code prediction.
加重 k 最近傍回帰およびソフトラベル予測のための正確な認証済みデータ Shapley
Data Shapley は、どのトレーニング ポイントにどのような価値があるのかについての標準的な原則に基づいた回答であり、その k 近傍 (KNN) 特化は実際に展開されるバージョンであり、pyDVL や OpenDataVal などのツールキットに同梱されている正確な推定器です。正確なアルゴリズムは、非重み付き KNN と重み付き KNN 分類で知られていますが、重み付き KNN 回帰とソフトラベル予測は抵抗がありました。唯一の正確な方法は、近傍サイズ K で指数関数的な O(N^K) 総当たり法です。障害は、重み付き回帰予測は 2 つの連立に依存する和の比であり、その正規化分母が以前の多項式アルゴリズムが依存していた加法、しきい値、および重複の構造を壊します。私たちはこのギャップを埋めます。 (i) 加重 KNN 回帰 Data Shapley 用の最初の擬似多項式時間正確アルゴリズム (固定格子精度での N と K の多項式)、結合整数状態 (w の合計、w*y の合計) にわたる計数動的プログラムであり、12,716 個の敵対的インスタンスで不一致がゼロの徹底的な列挙に対して検証されます。 (ii) 86,400 回のチェックにわたって違反が一度もなかった、機械チェック可能な値ごとのエラー証明書を備えた、継続的な分銅と目標に対する認定された FPTAS。 (iii) 無条件の Omega(D_w) 出力サイズの下限とアクセス モデルの硬度の結果を含む複雑さの状況。 (iv) 重み付けされたソフトラベルのマルチクラス拡張。私たちは、オープンソースの CPU 専用ライブラリと、最初の正確な重み付き回帰 Data Shapley のグラウンド トゥルースをリリースします。ダウンストリームの誤ったラベルの検出では、正確な値は、事前に登録された結果であるモンテカルロ データ シャプレー (データセット レベル TOST、n=8、p<10^-4) と統計的に同等です。正確さの価値は代わりに、決定論、認定誤差限界、および監査推定量の正確な参照です。モンテカルロでは、最大 3,000 の順列 (~1.28e6 のユーティリティ評価) まで、テストされたどの予算でも正確な上位 10% のランキングが再現されませんでした。
原文 (English)
Exact and Certified Data Shapley for Weighted k-Nearest-Neighbor Regression and Soft-Label Prediction
Data Shapley is the standard principled answer to which training points are worth what, and its k-nearest-neighbor (KNN) specialization is the version deployed in practice: the exact estimator shipped by toolkits such as pyDVL and OpenDataVal. Exact algorithms are known for unweighted KNN and for weighted KNN classification, but weighted KNN regression and soft-label prediction have resisted: the only exact method is an O(N^K) brute force, exponential in neighborhood size K. The obstruction: the weighted regression prediction is a ratio of two coalition-dependent sums, whose normalization denominator breaks the additive, threshold, and duplication structures the prior polynomial algorithms rely on. We close this gap. We give (i) the first pseudo-polynomial-time exact algorithm (polynomial in N and K at fixed lattice precision) for weighted KNN-regression Data Shapley, a counting dynamic program over the joint integer state (sum of w, sum of w*y), verified against exhaustive enumeration with zero mismatch on 12,716 adversarial instances; (ii) a certified FPTAS for continuous weights and targets, with a machine-checkable per-value error certificate never violated across 86,400 checks; (iii) a complexity landscape, including an unconditional Omega(D_w) output-size lower bound and access-model hardness results; and (iv) a weighted soft-label multi-class extension. We release an open-source, CPU-only library and the first exact weighted-regression Data Shapley ground truth. On downstream mislabel detection our exact values are statistically equivalent to Monte-Carlo Data Shapley (dataset-level TOST, n=8, p<10^-4), the pre-registered outcome; the value of exactness is instead determinism, a certified error bound, and an exact reference for auditing estimators: Monte-Carlo did not reproduce the exact top-10% ranking at any budget tested, up to 3,000 permutations (~1.28e6 utility evaluations).
慢性腎臓病の早期予測のための機械学習モデルの信頼性の評価: データ漏洩と予測子の安定性の系統的レビュー
機械学習を使用した慢性腎臓病の早期検出は、医療関連のコンピューター サイエンスで大きな関心を集めています。この分野の急速な進歩にもかかわらず、報告された研究の多くは依然として一貫性がなく、誤解を招く可能性があります。重大な欠点は、方法論上の懸念に関する組織的な評価が欠如していることです。主な問題には、データ漏洩、一時的な患者記録へのアクセスの制限、報告される臨床指標の不一致などが含まれます。この研究は、解釈可能な機械学習技術を使用した既存のCKD予測研究の系統的な文献レビューを提供しており、主要な学術データベース全体の体系的な検索を通じて19の関連研究が選択されています。方法論的な信頼性を評価するために、この研究では、情報漏洩の構造化された分類法と、CKD 予測研究全体にわたる信頼性を系統的に評価するための定量的漏洩スコアリング フレームワークを導入しています。分析により、漏れとパフォーマンスの膨張の間に強い関係があることが明らかになりました。ここで、漏れの多い研究では平均精度が 95.48% であるのに対し、漏れのない研究では 80.2% と報告されており、これは約 15.28% の増加を反映しています。さらに、クロススタディの特徴安定性分析では、一貫して再現できる予測変数はごく一部であり、80% 以上が信頼性に欠けていることが示されています。全体として、この調査結果は、報告されているパフォーマンスの向上の多くは、真の予測能力ではなく、方法論的な制限に起因していることを示唆しています。
原文 (English)
Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability
The early detection of Chronic Kidney Disease using machine learning has attracted significant interest in healthcare-related computer science. Despite rapid advancements in this field, many reported studies remain inconsistent and potentially misleading. A significant drawback is the lack of organized evaluation regarding methodological concerns. Key issues include data leakage, limited access to temporal patient records and inconsistency in reported clinical indicators. This research offers a systematic literature review of existing CKD prediction studies using interpretable machine learning techniques, where nineteen relevant studies were selected via systematic searches across major academic databases. To assess methodological reliability, this study introduces a structured taxonomy of information leakage and a quantitative leakage scoring framework to systematically evaluate reliability across CKD prediction studies. The analysis reveals a strong relationship between leakage and inflated performance. Here, High leakage-studies report an average accuracy of 95.48%, compared to 80.2% for leakage-free studies, reflecting an increase of approximately 15.28%. Furthermore, a cross-study feature stability analysis shows that only a small subset of predictors is consistently reproducible, with over 80% lacking reliability. Overall, the findings suggest that many reported performance improvements stem from methodological limitations rather than true predictive capability.
座標ゲージを超えて: 神経崩壊後のドナー固有の機能的指紋を検出するための監査済みプロトコル
独立してトレーニングされたニューラル ネットワークには共有のニューロン インデックス参照フレームがないため、それらを比較するには座標の自由度を考慮する必要があります。 Neural Collapse はこの問題をさらに明確にします。ネットワークは共有された低次元ジオメトリに向かって収束し、収束後も軌道固有の機能変動が区別可能なままであるかどうかという問題が生じます。我々は、検出可能性、移植可能性、因果関係の持続性という 3 つの主張を区別し、最初の主張に取り組みます。 MNIST 上で Neural Collapse を再構築する 5 つの独立してトレーニングされたネットワークを使用して、ドナー頭部をレシピエント座標にマッピングする検証済みのアフィン補正アライメントを適用します。ドナー固有の機能的指紋は、レシピエントレベルのベースライン補正後も区別可能なままです。20 個の順序付けされたドナーとレシピエントのペアすべてが正確に識別され、正確な順列 p=0.0083 で、漏洩監査に対して堅牢です。これらの所見は、ここで使用された検査での検出可能性を確立しますが、移植可能性や因果関係の持続性は確立しません。この研究では、調整、曖昧さ診断、漏れ制御をどのように組み合わせて、制御された設定でネットワーク間の変動をテストするかを示しています。これがそれを超えて一般化するかどうかは不明です。
原文 (English)
Beyond Coordinate Gauge: An Audited Protocol for Detecting Donor-Specific Functional Fingerprints after Neural Collapse
Independently trained neural networks have no shared neuron-index reference frame, so comparing them requires accounting for coordinate freedom. Neural Collapse sharpens this problem: networks converge toward a shared, low-dimensional geometry, raising the question of whether trajectory-specific functional variation remains distinguishable after convergence. We distinguish three claims - detectability, transplantability, and causal persistence - and address the first. Using five independently trained networks reconstructing Neural Collapse on MNIST, we apply a verified affine-correct alignment mapping donor heads into recipient coordinates. Donor-specific functional fingerprints remain distinguishable after recipient-level baseline correction: all 20 ordered donor-recipient pairs are correctly identified, with an exact permutation p=0.0083, robust to a leakage audit. These findings establish detectability under the test used here, but not transplantability or causal persistence. The study shows how alignment, ambiguity diagnostics, and leakage control combine to test cross-network variation in a controlled setting; whether this generalizes beyond it is open.
MU-MISO システムにおけるパイロットからビームフォーマーへの直接設計のための自己進化型インコンテキスト学習
マルチユーザー多入力単出力 (MU-MISO) システムにおけるパイロットベースのビームフォーミングのパフォーマンスを向上させるための、強化されたインコンテキスト学習 (ICL) フレームワークを開発します。提案された方式は、ICL-Transformer バックボーンをパイロット エンコーダ/デコーダ ネットワーク (EDN) およびビームフォーマ EDN と統合します。私たちの ICL ネットワークの重要な特徴は、モデル固有のコンテキスト データセットの構築によって可能になり、再トレーニングなしで複数のチャネル モデルを処理できることです。収束性とロバスト性を向上させるために、3 つの重要なイノベーションを導入します。(a) 教師あり LMMSE ラベル付き模倣から教師なし合計レート最大化にスムーズに移行するカリキュラム学習 (CL) 戦略、(b) CL ベースのトレーニング中にすべてのチャネル モデルのコンテキスト データセットを動的に拡張および洗練する自己進化メカニズム、(c) 一般的な ICL フレームワークにいくつかの不一致を組み込み、明示的な不一致をバイパスする不一致認識拡張チャンネルの校正。アブレーション研究は、コンテキスト内アーキテクチャと強化されたトレーニング戦略の有効性を検証します。多様な通信環境におけるシミュレーション結果は、提案された方式が、勾配ベースのパラメータ更新なしで、目に見えるチャネルモデルと目に見えないチャネルモデルの両方に迅速に適応でき、インテリジェントなコンテキスト構築を通じて不一致の問題を軽減できることを示しています。さらに、私たちの方式は、WMMSE ベンチマークや最近の Transformer ベースの方法など、パイロットベースの設定の下で既存のビームフォーミング方式よりも一貫して優れています。
原文 (English)
Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems
We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-input single-output (MU-MISO) systems. The proposed scheme integrates the ICL-Transformer backbone with the pilot encoder-decoder network (EDN) and the beamformer EDN. A crucial feature of our ICL network is that it can handle multiple channel models without retraining, enabled by the construction of model-specific context datasets. To improve convergence and robustness, we introduce three key innovations: (a) a curriculum learning (CL) strategy that smoothly transitions from supervised LMMSE-labeled imitation to unsupervised sum-rate maximization, (b) a self-evolving mechanism that dynamically expands and refines the context datasets for all channel models during CL-based training, and (c) a mismatch-aware extension that incorporates several mismatches into the general ICL framework and bypasses explicit channel calibrations. Ablation studies validate the effectiveness of the in-context architecture and enhanced training strategies. Simulation results over diverse communication environments show that the proposed scheme is able to rapidly adapt to both seen and unseen channel models without gradient-based parameter updates, and can mitigate the mismatch issues via intelligent context constructions. Furthermore, our scheme consistently outperforms the existing beamforming schemes under pilot-based settings, including the WMMSE benchmark and the recent Transformer-based methods.
離散化の学習: スペクトル ガイダンスを備えた拡散ベースの適応メッシュ
ほとんどの神経偏微分方程式 (PDE) サロゲートは、グリッドが選択された後にフィールドがどのように展開するかを学習します。ただし、演算子が適用される前に、グリッドは空間、解像度、スペクトル帯域幅にわたってモデリング容量がどのように割り当てられるかをすでに決定しています。私たちは、この隠れた設計の選択自体が学習可能であるべきだと主張し、標準的なオペレーターの学習とは異なる質問につながります。つまり、サロゲートは場の進化を予測する前に解像度が存在するべき場所を学習できるでしょうか?適応的離散化を、有効なメッシュ変位に対する物理的制約付きの条件付き生成問題として定式化します。 PDE 場予測における拡散モデルの成功は、同様の構造化制約の下で適応的離散化を学習できる可能性を示唆しています。これにより、2 段階の拡散フレームワークが得られます。段階 1 では、観察されたダイナミクスに基づいて条件付けされた r 適応変位メッシュを学習し、一方、段階 2 では、メッシュ情報に基づいた表現から解の展開を予測します。メッシュ ジェネレーターは、適応が物理的に解釈可能で数値的に正当なままであるように、物理を意識したプロキシ チャネル、幾何学的妥当性制約、およびローカル スペクトル集中によって正規化されます。結果は、5 つの PDE レジームにわたって、拡散ベースの学習済み離散化が適応メッシュおよび低次数ベースラインと競合し、特に固定または手動による割り当てが不十分なレジームで大きな効果が得られることを示しています。主な結論は、普遍的な最適なメッシュ規則が存在するということではなく、離散化はレジームに依存した方法で学習されるべきであるということです。つまり、異なる空間構造とスペクトル構造は異なる割り当て動作を優先します。これにより、ニューラル PDE ソルバーの適応メッシングがソルバー固有のヒューリスティックから生成表現学習問題に再構成されます。
原文 (English)
Learning to Discretize: Diffusion-Based Adaptive Mesh with Spectral Guidance
Most neural partial differential equation (PDE) surrogates learn how fields evolve after a grid has already been chosen. However, before any operator is applied, the grid has already determined how modeling capacity is allocated across space, resolution, and spectral bandwidth. We argue that this hidden design choice should itself be learnable, leading to a question different from standard operator learning: can a surrogate learn where resolution should exist before predicting field evolution? We formulate adaptive discretization as a physics-constrained conditional generation problem over valid mesh displacements. The success of diffusion models in PDE field prediction suggests their potential for learning adaptive discretizations under similar structured constraints. This leads to a two-stage diffusion framework: Stage 1 learns an r-adaptive displacement mesh conditioned on the observed dynamics, while Stage 2 predicts the solution evolution from the mesh-informed representation. The mesh generator is regularized by physics-aware proxy channels, geometric validity constraints, and local spectral concentration so that adaptation remains physically interpretable and numerically legal. Across five PDE regimes, the results show that diffusion-based learned discretization is competitive with adaptive-mesh and reduced-order baselines, with particularly strong gains in regimes where fixed or handcrafted allocation is insufficient. The main conclusion is not that there exists a universal optimal mesh rule, but that discretization should be learned in a regime-dependent manner: different spatial and spectral structures favor different allocation behaviors. This reframes adaptive meshing for neural PDE solvers from a solver-specific heuristic into a generative representation-learning problem.
機械の非学習のための信号に基づく最適化
現在の機械の非学習手法は、主にグローバルで粗粒度の介入戦略に依存しています。それらは、非学習プロセスをガイドするための正確なパイロット信号を欠いており、さまざまな非学習タスク間で差別化可能なガイダンスを提供できません。元のトレーニング中のサンプルの記憶強度が異なるため、このような一律の戦略は 2 つの問題を引き起こします。1 つは、一部のサンプルが過剰に学習され、モデルの有用性に悪影響を与えることです。一方、十分に学習されていない情報が残っており、プライバシー攻撃に悪用される可能性があります。この論文では、GSUO を提案します。GSUO は、タスク固有のきめ細かいガイダンス信号を設計して、アンラーニング プロセスを制御し、ランダム サブセット タスクとクラスごとの忘却タスクの両方に適用できる、ガイダンス信号を認識した非学習最適化フレームワークです。広範な実験により、GSUO は、高効率と大幅な高速化を達成しながら、非学習の有効性と一般化の両方の点で 14 のベースラインを上回っていることが実証され、信頼性の高い機械の非学習に対するその有効性が検証されています。
原文 (English)
Signal-Guided Optimization for Machine Unlearning
Current machine unlearning methods predominantly rely on global, coarse-grained intervention strategies. They lack precise pilot signals to guide the unlearning process and fail to provide differentiable guidance across different unlearning tasks. Due to the varying memorization strengths of samples during original training, such a uniform strategy leads to two problems: some samples are over-unlearned, which harms model utility; while others are under-unlearned, leaving residual information that can be exploited by privacy attacks. In this paper, we propose GSUO, a guidance-signal-aware unlearning optimization framework that designs task-specific fine-grained guidance signals to steer the unlearning process and is applicable to both random-subset and class-wise forgetting tasks. Extensive experiments demonstrate that GSUO outperforms 14 baselines in terms of both unlearning effectiveness and generalization, while achieving high efficiency and significant speedups, validating its effectiveness for reliable machine unlearning.
遺伝子発現に基づいた共同制御による精密な分子設計のための生成モデリング
精密分子設計は、生物学的関連性や分子設計戦略などの複数の条件を共同制御することで、個別化された医薬品候補を発見することを目的としています。生物学的関連性は疾患または摂動条件下での細胞の機能状態を反映し、分子設計戦略は構造的意図と特性の最適化に関して補完的な指針を提供します。本研究では、遺伝子発現プロファイルによってコード化された生物学的状態とテキストで表現された分子構造情報、および数値によって定量化された化学的特性を統一モデリングフレームワーク内で統合する、共同制御された高精度分子生成モデルであるJoPMolを提案します。この定式化により、結合状態の制御下での候補分子の調整された生成と最適化が可能になります。実験結果は、JoPMol が複数の評価基準にわたって最先端の方法よりも優れていることを示しています。さらに、JoPMol は転送タスクと生物学的に根拠のあるシミュレーション シナリオの両方で強力な一般化能力を実証し、精密な分子設計に対するその有効性を検証します。ソース コードは https://github.com/hala-yh/JoPMol で公開されています。
原文 (English)
Gene Expression-Informed Jointly Controlled Generative Modeling for Precision Molecular Design
Precision molecular design aims to discover personalized drug candidates through joint control of multiple conditions, such as biological relevance and molecular design strategies. Biological relevance reflects cellular functional states under disease or perturbation conditions, while molecular design strategies provide complementary guidance in terms of structural intentions and property optimization. In this study, we propose JoPMol, a jointly controlled precision molecular generative model that integrates biological states encoded by gene expression profiles with molecular structure information expressed in text, and chemical properties quantified by numerical values within a unified modeling framework. This formulation enables coordinated generation and optimization of candidate molecules under joint condition control. Experimental results show that JoPMol outperforms state-of-the-art methods across multiple evaluation metrics. Moreover, JoPMol demonstrates strong generalization ability in both transfer tasks and biologically grounded simulation scenarios, validating its effectiveness for precision molecular design. The source code is publicly available at https://github.com/hala-yh/JoPMol.
応答条件全体にわたる不均一な信頼性の評価: 自動エッセイ採点で示される条件付き一般化可能性フレームワーク
集合的な信頼性推定は、応答条件全体にわたる測定設計負荷の不均一性を曖昧にする可能性があるため、単一の G スタディまたは D スタディでは、特定の階層に対する設計の適切性を誤って評価する可能性があります。この研究では、3 つのコンポーネントからなる条件付き一般化可能性フレームワークを導入します。まず、自動スコアリング構成 (固定パイプライン内で許容されるエンコーダー アーキテクチャとスコアリング ヘッド ファミリ) は、付随的なモデリングの選択肢ではなく、許容される測定条件の世界として扱われます。次に、分析的な D スタディの予測が、有限のスコアリング プールにわたる経験的な構成スイープと比較され、設計の妥当性の 2 つの推定値が得られ、その一致または発散によって実現された構成の世界が診断されます。第三に、証拠はエントロピーで定義された応答層に条件付けされており、エントロピーを執筆の品質に関する構成的主張ではなく、操作層化変数として扱います。最近の一般化可能性理論の拡張機能は応答側で AI によって生成された項目のバリアントに対処しますが、このフレームワークは同様のスコアリング側の問題、つまり AI を介したスコアリング構成に対処します。時間制限された L2 ライティングの自動エッセイ スコアリングで実証され、実現されたデザインは全体として信頼できるものでした (ファイ約 0.76)。エントロピー層内で再推定したところ、信頼性は高いままでしたが、適度かつ頑強に低下しました (ファイ = 0.88、0.87、0.84)。この勾配は、さまざまな意思決定研究要件を意味し、最も高いエントロピー層は最も多くの交差条件を必要とします。このフレームワークは、不均一な信頼性を評価するための移植可能なワークフローを提供します。
原文 (English)
Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring
Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-study may mischaracterize a design's adequacy for particular strata. This study introduces a conditional generalizability framework with three components. First, automated scoring configurations -- the encoder architectures and scoring-head families admissible within a fixed pipeline -- are treated as a universe of admissible measurement conditions rather than incidental modeling choices. Second, analytical D-study projections are compared with empirical configuration sweeps over a finite scoring pool, yielding two estimands of design adequacy whose agreement or divergence diagnoses the realized configuration universe. Third, evidence is conditioned on entropy-defined response strata, treating entropy as an operational stratification variable, not a construct claim about writing quality. Whereas recent generalizability-theory extensions address AI-generated item variants on the response side, this framework addresses the analogous scoring-side problem: AI-mediated scoring configurations. Demonstrated with automated essay scoring of timed L2 writing, the realized design was dependable in aggregate (Phi approx 0.76). Re-estimated within entropy strata, dependability stayed high but declined modestly and robustly (Phi = 0.88, 0.87, 0.84) -- a gradient implying different decision-study requirements, the highest-entropy stratum requiring the most crossed conditions. The framework offers a portable workflow for evaluating nonuniform dependability.
除去可能な欠陥: 経済学と意図的な欠陥の限界
スペシャリストはゼネラリストが許容しない盲点を許容します。通常、これは最小限に抑えるべきコストとして扱われます。私たちはそれを設計変数として扱います。不足が維持されるのは、それが致命的になるようなまれな状況では、要求に応じて補償チャネルにルーティングすることで支払いおよび削除されるためです。 3 つの結果を示します。第一に、不足を維持することが計算可能な経済的地位となる有利な条件。構造的には、タウンゼントの高価な状態検証技術としての検出器を使用して、能力ギャップに適用されるエールリッヒ・ベッカーの市場対自己保険のマージンです。第二に、除去可能性の両面の特徴付けです。結合補題は、欠陥が認識の粗大化である場合、どのスイッチも利益と害を分離できず、逆の結果 (交絡した検出器はプレミアムを獲得せず、プラスのプレミアムを主張する欠陥内の政策は、乗算力学の下でマイナスの長期成長に駆動される) と達成可能性の結果 (欠陥の外側の検出器はプラスのプレミアムを獲得する) を生み出すことを示しています。同時に、重大度上限またはミス率 O(1/L) を持つ構造化された不確実性クラス: 検出器関連の区別が制限を乗り越え、有利な条件が維持される場合、欠陥は有益に除去されます。プレミアムは、経済価格ベクトルで設定されたクラスの ROC のサポート関数です。第三に、観測上の欠陥と容量上の欠陥は、展開配布へのアクセスによってそれらが救われるかどうかという点でまったく異なります。ギャップはクロスリークとクロージャの不足として分解され、タスクごとのランダム化によって後者が買い戻されますが、前者は決して買い戻されません。検出器は、損失重大度が線形 (対数係数まで) のトレーニング料金で宣言された致命的カテゴリから学習できます。結果は、チョウの拒否オプション、破滅下でのケリーの成長、および選択的予測を総合したものです。
原文 (English)
Removable Defects: The Economics and Limits of Deliberate Deficiency
A specialist tolerates blind spots that a generalist does not. Usually this is treated as a cost to be minimized. We treat it as a design variable: a deficiency can be kept because it pays and removed on demand in the rare situation where it would be fatal, by routing to a compensation channel. We give three results. First, an advantage condition under which keeping the deficiency is a computable economic position; structurally it is the Ehrlich-Becker market-vs-self-insurance margin applied to a competence gap, with the detector as a Townsend costly-state-verification technology. Second, a two-sided characterization of removability. A coupling lemma shows that when the deficiency is a coarsening of perception, no switch can separate benefit from harm, yielding a converse (a confounded detector earns zero premium, and any within-defect policy insisting on positive premium is driven, under multiplicative dynamics, to negative long-run growth) and an achievability result (a detector outside the deficiency earns a positive premium). Together, over structured uncertainty classes with severity capped or miss rate O(1/L): a defect is profitably removable iff the detector-relevant distinction survives the restriction and the advantage condition holds; the premium is the support function of the class's ROC set at an economic price vector. Third, observation defects and capacity defects differ exactly on whether access to the deployment distribution rescues them; the gap decomposes as cross-leak plus a closure deficit, and per-task randomization buys back the latter, never the former. The detector can be learned from declared fatal categories at a training bill linear in loss severity (up to a log factor). The results synthesize Chow's reject option, Kelly growth under ruin, and selective prediction.
トランスフォーマー FFN ニューロンの疎な層間依存関係
フィードフォワード ネットワーク (FFN) ブロックは、Transformer アーキテクチャのパラメーターと計算の大部分を占めますが、残差ストリームによって引き起こされる加算的な重ね合わせのため、内部構造の解釈は依然として困難です。 FFN ニューロンの活性化が、先行するニューロンの活性化と注意出力のまばらなセットによって説明できるかどうかを調べます。ターゲットニューロンの活性化に対する上流ニューロンと注意出力の相対的な影響を推定する、トレーニング不要のアトリビューション手法を導入します。経験的に、モデルとレイヤー全体で、残りのすべての入力が平均値でマスクされている場合、先行する活性化と注意出力の小さなサブセットがニューロンの活性化を高い忠実度で保存するのに十分であることがわかります。上流層の固有の活性化の疎性を考慮すると、実効的な疎性はさらに大きくなります。さらに、ニューロン固有のマスクをすべての層に同時に適用すると、誘発された偏差がネットワークを通じて伝播するため、中程度のスパース性レベルではモデルの複雑さがほとんど変化しません。これらの結果は、高密度のパラメータ化にもかかわらず、FFN がニューロン レベルで疎で構造化された層間依存関係を示すことを示しています。私たちの方法は、回路レベルの解釈可能性のための実用的でスケーラブルなツールを提供し、効率的な推論に潜在的な影響を与える候補の疎な経路を特定します。
原文 (English)
Sparse Inter-Layer Dependencies of Transformer FFN Neurons
Feedforward network (FFN) blocks account for a large fraction of the parameters and computation in Transformer architectures, yet their internal structure remains difficult to interpret due to the additive superposition induced by the residual stream. We examine whether the activation of an FFN neuron can be explained by a sparse set of preceding neuron activations and attention outputs. We introduce a training-free attribution method that estimates the relative influence of upstream neurons and attention outputs on a target neuron's activation. Empirically, across models and layers, we find that small subsets of preceding activations and attention outputs suffice to preserve neuron activations with high fidelity when all remaining inputs are masked with their average values. Effective sparsity is even greater when accounting for the inherent activation sparsity of upstream layers. Moreover, applying the neuron-specific masks in all layers simultaneously, such that the induced deviations propagate through the network, leaves model perplexity largely unchanged at moderate sparsity levels. These results demonstrate that, despite dense parameterization, FFNs exhibit sparse and structured inter-layer dependencies at the neuron level. Our method provides a practical, scalable tool for circuit-level interpretability and identifies candidate sparse pathways with potential implications for efficient inference.
階層的で信頼できる構造を持つデータにおけるクラスの不均衡の影響を軽減する
Common Weakness Enumeration (CWE) 分類法を使用してサイバーセキュリティの脆弱性を分類することは、極端なクラスの不均衡と弱点カテゴリ間の強い階層的依存関係により困難です。合成マイノリティ オーバーサンプリング技術 (SMOTE) や適応合成サンプリング (ADASYN) などのオーバーサンプリング技術は、クラスの不均衡を緩和するために広く採用されていますが、階層的な CWE テキスト分類に対するその有効性はほとんど解明されていません。この論文では、学習可能な親クラスの埋め込みを通じて CWE 構造情報を明示的に組み込み、分類学的一貫性を維持する、階層を意識した RoBERTa フレームワークを提案します。私たちの実験は、高次元の埋め込み空間での合成補間が CWE 階層の固有の親子制約に違反し、古典的な ML モデルにわずかな利点しか提供せず、深層学習アーキテクチャを一貫して低下させることを示しています。 CWE Research Concept データセットで評価された提案モデルは、データ拡張なしで 0.76 の加重 F1 スコアを達成し、BERT ベースラインと比較して F1 スコアが 0.40 から 0.60 に改善したクラス カテゴリを含む、少数派クラスで顕著な向上を示し、すべてのベースラインを上回りました。私たちの結果は、階層を意識した表現学習が、構造化された脆弱性分類のためのオーバーサンプリングに代わるより原理的な代替手段であることを示唆しています。
原文 (English)
Mitigating The Effect of Class Imbalance in Data with Hierarchical and Dependable Structure
Classifying cybersecurity vulnerabilities using the Common Weakness Enumeration (CWE) taxonomy is challenging due to extreme class imbalance and strong hierarchical dependencies among weakness categories. Although oversampling techniques such as Synthetic Minority Oversampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN) are widely adopted to mitigate class imbalance, their effectiveness for hierarchical CWE text classification remains largely unexplored. This paper proposes a Hierarchy-Aware RoBERTa framework that explicitly incorporates CWE structural information through learnable parent-class embeddings, preserving taxonomic consistency. Our experiments demonstrate that synthetic interpolation in high-dimensional embedding spaces violates the inherent parent-child constraints of the CWE hierarchy, offering only marginal benefits for classical ML models while consistently degrading deep learning architectures. Evaluated on a CWE Research Concept dataset, the proposed model achieves a weighted F1-score of 0.76 without data augmentation, outperforming all baselines with notable gains on minority classes, including the Class category whose F1-score improved from 0.40 to 0.60 over the BERT baseline. Our results suggest that hierarchy-aware representation learning is a more principled alternative to oversampling for structured vulnerability classification.
適切なモデルをマージしていますか? LLM のモデル結合に対するエキスパート トレーニング期間の影響
マルチタスク モデルのマージでは、個別にトレーニングされたエキスパート モデルを、共同トレーニングなしですべてのタスクを処理する単一のモデルに結合します。標準的な手法では、最適な検証損失で専門家を結合します。私たちは、ドメイン専門家のトレーニング期間がマージされたモデルの品質にどのような影響を与えるかを体系的に研究することで、この慣例に挑戦します。 3 つのモデル サイズ (Qwen 3.5 0.8B、2B、および 4B) にわたる 5 つのドメイン (数学、コード、命令追従、多言語、安全性) について専門家を微調整し、最適なトレーニング ステップの 25\% ~ 500\% のチェックポイントを節約し、各期間で 5 つのマージ方法を評価します。私たちの調査結果は、手法に依存する顕著なパターンを明らかにしました。単純な平均化は過学習により急激に低下しますが、スパース化ベースの手法は検証の最適値をはるかに超えて最高のパフォーマンスを達成します。これをバイアス分散分解分析を通じて形式化し、分散の高い個々の学習者から平均化のメリットが得られるランダム フォレストとの類似点を描きます。これらの結果は、トレーニング期間と結合方法は独立して選択するのではなく、組み合わせて選択する必要があることを示唆しています。
原文 (English)
Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs
Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal validation loss. We challenge this convention by systematically studying how training duration of domain experts affects the quality of the merged model. We fine-tune experts on five domains (Math, Code, Instruction Following, Multilingual, and Safety) across three model sizes (Qwen 3.5 0.8B, 2B, and 4B), saving checkpoints from 25\% to 500\% of the optimal training steps and evaluating five merging methods at each duration. Our findings reveal a striking method-dependent pattern: simple averaging degrades sharply with overfitting, while sparsification-based methods achieve their best performance well past the validation optimum. We formalize this through bias-variance decomposition analysis, drawing a parallel to random forests where averaging benefits from high-variance individual learners. These results suggest that training duration and merging method should be chosen jointly rather than independently.
V-JEPA と深層時空間学習を使用した HPC 対応ビデオベースの海岸波パラメータ推定
高い導入コスト、狭い空間範囲、嵐の影響を受けやすいことはすべて、従来の現場での手法が直面する課題です。この論文では、単眼海岸ビデオからの 5 つの海岸波パラメータ、すなわち有義波高 (Hs)、最大波高 (Hmax)、ピーク周期 (Tp)、ゼロアップクロス周期 (Tz) および波の方向 (θ) をセンサフリーで共同推定するための、ビデオベースのハイ パフォーマンス コンピューティング (HPC) 対応ディープ ラーニング フレームワークを紹介します。提案されたアーキテクチャは、視覚的に困難なシナリオで堅牢な時空間特徴抽出を実現する V-JEPA (自己監視型) ViT Small バックボーン、流体力学的に砕ける領域とうねり領域の両方で波動を広帯域で表現するためのデュアルストリーム SlowFast 時間エンコーダ、流体力学的にアクティブな波の波長帯域に重点を置いて構造に顕著性情報を追加するファーンバック オプティカル フロー アルゴリズムに基づくオプティカル フロー ストリーム、およびマルチタスク回帰層で構成されます。分散制約付き (エアリー波分散 lambda_p = 0.1)。このモデルは NVIDIA DGX A100 クラスターでトレーニングされ、エポック 31 で早期に停止され、Hs、Hmax、Tp、Tz および波の方向についてそれぞれ 0.451、0.578、0.643、0.680、および 0.832 のピアソン相関係数を達成し、地理的に多様なテスト データ サイトへの一般化機能を備えました。データが限られたレジーム (6 つの注釈付きトレーニング シーン) で動作している間、フレームワークは統計的に有意な時間相関 (PCC 0.451 ~ 0.832) を示し、概念実証の実現可能性を確認します。 R2 値 (最大 0.246) は、アノテーション付きデータセットが大きくなると分散キャプチャが向上することを示しています。
原文 (English)
HPC-Enabled Video-based Coastal Wave Parameter Estimation Using V-JEPA and Deep Spatiotemporal Learning
High deployment cost, poor spatial coverage and susceptibility to storm conditions are all challenges faced by traditional in-situ methods. This paper presents a video-based and high performance computing (HPC) enabled deep learning framework for joint sensor free estimation of five coastal wave parameters, namely significant wave height (Hs), maximum wave height (Hmax), peak period (Tp), zero upcrossing period (Tz) and wave direction (theta) from monocular coastal video. The proposed architecture comprises of a V-JEPA (self supervised) ViT Small backbone for robust spatiotemporal feature extraction in visually challenging scenarios, a dual-stream SlowFast temporal encoder for broad bandwidth representation of wave motion in both hydrodynamic breaking and swell regimes, an optical flow stream based on Farneback optical flow algorithm for adding saliency information to the structure with emphasis on hydrodynamically active wavelength bands of waves, and a multi-task regression layer with dispersion constraints (Airy wave dispersion lambda_p = 0.1). The model was trained on an NVIDIA DGX A100 cluster and was early stopped at epoch 31 and achieved Pearson correlation coefficients of 0.451, 0.578, 0.643, 0.680 and 0.832 for Hs, Hmax, Tp, Tz and wave direction respectively, with generalization ability to geographically diverse held out test data sites. While operating in a data-limited regime (6 annotated training scenes), the framework demonstrates statistically significant temporal correlations (PCC of 0.451 to 0.832), confirming proof of concept feasibility; R2 values (max 0.246) indicate that variance capture will improve with larger annotated datasets.
異種医療視覚的質問応答のための継続学習の実証分析
医療視覚的質問応答 (MedVQA) システムを実際の臨床現場に導入するには、以前に取得した知識を忘れることなく、新しい臨床タスクに適応するモデルが必要です。継続学習 (CL) は、この設定のための実用的なフレームワークを提供します。医療視覚言語モデルの急速な進歩にも関わらず、異種の MedVQA タスクにわたってこれらのモデルをトレーニングする際の CL メソッドの動作はまだ解明されていません。この研究では、分類、マルチラベル分類、検出、細胞計数、レポート作成など、さまざまな臨床目的にわたる MedVQA の CL の体系的な評価が示されています。具体的には、(1) 既存の CL 手法が壊滅的な物忘れを軽減する能力を調査します。 (2) タスクの順序に対する敏感さ。さまざまなタスクの順序がパフォーマンスの保持と忘れにどのように影響するかを分析します。 (3) 新しいタスクが学習されるにつれて低ランクの適応パラメータが進化し、さまざまな CL メソッドの下での重みドリフトのパターンが明らかになります。我々の調査結果は、既存の CL 手法では、異なる目的と監視形式を持つタスクが交互に存在する場合、安定性と可塑性のバランスを維持するのに苦労していることを示唆しています。コードと完全な実験セットアップは一般に公開されます。
原文 (English)
An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering
Deploying medical visual question answering (MedVQA) systems in real-world clinical settings requires models that adapt to new clinical tasks without forgetting previously acquired knowledge. Continual learning (CL) provides a practical framework for this setting. Despite rapid progress in medical vision-language models, the behavior of CL methods when training these models across heterogeneous MedVQA tasks remains underexplored. This work presents a systematic evaluation of CL for MedVQA across diverse clinical objectives, including classification, multi-label classification, detection, cell counting, and report generation. Specifically, we explore (1) the ability of existing CL methods to mitigate catastrophic forgetting; (2) their sensitivity to task ordering, analyzing how different task sequences influence performance retention and forgetting; and (3) the evolution of low-rank adaptation parameters as new tasks are learned, revealing patterns of weight drift under different CL methods. Our findings suggest that existing CL methods struggle to maintain stability-plasticity balance when tasks with different objectives and supervision formats are interleaved. Code and full experimental setup will be publicly available.
トレーニング不要の合成画像帰属における表現と参照の選択
合成画像の帰属は、特定の AI 生成画像の原因となっている生成元を特定することを目的としています。タスク固有の分類子を再トレーニングするのではなく、ソース固有の参照を追加することで新しく登場したジェネレーターを組み込むことができるため、トレーニング不要の参照ベースのアトリビューション手法は拡張が容易です。それらのパフォーマンスは、比較に使用される表現空間とソース固有の参照の構築方法という 2 つの要因が組み合わされて決まります。しかし、これら 2 つの要因間の相互作用はほとんど解明されていないままです。この論文では、参照と既製の事前トレーニング済み表現を使用して、この相互作用の制御された分析を提供します。私たちは、CLIP と DINOv2 のさまざまな層から抽出された表現と、さまざまな意味論的制約を持つ 3 つの参照選択方法 (任意参照、意味論的に調整された参照、および再合成ベースの参照) を研究します。私たちの結果は、属性の精度が中間表現レベルで一貫してピークに達することを示しており、これは、強力な意味論的抽象化が優勢になる前に、ソースを識別する手がかりがよりアクセスしやすいことを示しています。さらに、中間表現は完全に意味的に中立ではなく、参照の選択が重要になることを示します。意味的に制約された参照は、特に限られた参照予算の下で、クエリと参照の不一致を減らし、帰属を改善します。再合成は参照が少ない場合に最も役立ちますが、中程度のサイズの参照プールが利用可能な場合は、意味的に調整された参照により精度とコストのトレードオフが向上します。私たちの調査結果は、トレーニング不要の参照ベースのアトリビューションは、画像が比較される場所、参照セットの構築方法、および利用可能な参照の数の間の相互作用として理解されるべきであることを示しています。
原文 (English)
Representation and Reference Selection in Training-Free Synthetic Image Attribution
Synthetic image attribution aims at identifying the generator responsible for a given AI-generated image. Training-free reference-based attribution methods are easily scalable, since newly emerging generators can be incorporated by adding source-specific references rather than retraining a task-specific classifier. Their performance depends on two coupled factors: the representation space used for comparison and the way source-specific references are constructed. However, the interaction between these two factors remains largely unexplored. In this paper, we provide a controlled analysis of this interaction using references and off-the-shelf pretrained representations. We study representations extracted from different layers of CLIP and DINOv2, along with three reference selection methods with varying semantic constraints: arbitrary, semantically aligned, and resynthesis-based references. Our results show that attribution accuracy consistently peaks at intermediate representation levels, indicating that source-discriminative cues are more accessible before strong semantic abstraction dominates. We further show that intermediate representations are not completely semantically neutral, making reference selection critical: semantically constrained references reduce query-reference mismatch and improve attribution, especially under limited reference budgets. Resynthesis is most useful in low-reference regimes, while semantically aligned references provide a better accuracy-cost trade-off when a moderate-sized reference pool is available. Our findings show that training-free reference-based attribution should be understood as the interaction between where images are compared, how the reference set is constructed, and how many references are available.
AutoTrace: エージェントのプロシージャ間の探索によるパッチからトリガーまで
脆弱性を修正するコミットが与えられると、トリガーのローカリゼーションは、どの特定のステートメントが脆弱なプログラムの状態を具体的な安全でない操作に変えるかを尋ねます。この質問は、バイナリ脆弱性の検出よりも難しいです。なぜなら、その答えには手続き間の因果推論が必要だからです。現実世界の CVE のかなりの部分では、トリガーとなるステートメントは、パッチ適用された関数の外側のいくつかの呼び出し層にあり、静的ルール セットやパターン マッチング言語モデルの範囲を超えています。ここでは、コード プロパティ グラフを層ごとに探索することで脆弱性トリガーを特定するエージェント パイプラインである AutoTrace を紹介します。LLM エージェントは次にどこを探すかを決定し、決定論的な許容ゲートはトリガーが報告される前にどのような証拠が必要かを決定します。エージェントは決して自らの権限でトリガーを受け入れることはありません。報告されたすべてのトリガーはグラフから引き出された明確な証拠によって裏付けられているため、パイプラインは根拠のないモデルの判断に依存することなく、プロシージャ内およびプロシージャ間の両方の脆弱性をカバーします。 InterPVD ベンチマーク全体では、AutoTrace は VulnHit 75.0%、FuncHit 80.8% に達し、同じコーパスにおける従来の最先端技術を上回っています。同じ機械を基盤として、SinkTrace-Bench を構築します。このデータセットは、攻撃者が制御する入力から伝播を経て危険な操作に至るソースからシンク (S2S) の因果関係の連鎖として、一致する脆弱性とパッチ適用されたプログラムの状態から抽出された、各脆弱性を明らかにします。これは、検証者によって確認された 1,542 個の、脆弱性と安全性の完全にバランスがとれたサンプルで構成されており、そのラベルの忠実性は専門家の注釈に対して監査されます。フロンティア LLM をベンチマークすると、最も強力なペアでも一致したペアを分離するのに苦労し、ローカリゼーション ターゲットを引き起こす因果推論のギャップが明らかになりました。アーティファクトは https://github.com/Erroristotle/AutoTrace で入手できます。
原文 (English)
AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration
Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe operation. This question is harder than binary vulnerability detection because the answer demands interprocedural, causal reasoning: in a substantial fraction of real-world CVEs the triggering statement lies several call layers outside the patched function, beyond the reach of static rule sets and pattern-matching language models alike. We present AutoTrace, an agentic pipeline that localizes vulnerability triggers by exploring a code property graph layer by layer, with LLM agents deciding where to look next and deterministic admissibility gates deciding what evidence is required before a trigger can be reported. Agents never accept a trigger on their own authority; every reported trigger is backed by explicit evidence drawn from the graph, so the pipeline covers both intra- and interprocedural vulnerabilities without relying on ungrounded model judgment. On the full InterPVD benchmark, AutoTrace reaches 75.0% VulnHit and 80.8% FuncHit, surpassing the prior state of the art on the same corpus. Building on the same machinery, we construct SinkTrace-Bench, a dataset that exposes each vulnerability as a source-to-sink (S2S) causal chain from attacker-controlled input through propagation to the dangerous operation, drawn from matched vulnerable and patched program states. It comprises 1,542 verifier-confirmed, perfectly balanced vulnerable/safe samples whose label fidelity we audit against expert annotations. Benchmarking frontier LLMs on it, we find that even the strongest struggle to separate the matched pairs, exposing the causal-reasoning gap that trigger localization targets. Artifact available at https://github.com/Erroristotle/AutoTrace.
24 時間農業ロボットの実現: 夜間の視覚ナビゲーションのための監視なしの昼夜クロスモーダル画像変換
視覚ナビゲーションは農業ロボット工学において広く研究されてきましたが、既存のシステムのほとんどは日中の条件を前提としています。実際、自律型ロボットを夜間に導入すると、24 時間の作物や土壌の監視、果物の収穫、夜間の害虫の検出など、大きな利点が得られます。しかし、最新のビジョンベースのシステムは、注釈が付けられた大規模な画像データセットに大きく依存しており、夜間の運用シナリオでデータを取得するのは依然として困難です。これに対処するために、ピクセル間の監視を必要とせずに、日中の植物の列の RGB 画像を近赤外線 (NIR) の夜間の対応する画像に変換する、教師なし画像変換フレームワークを提案します。これにより、夜間の知覚モデルをトレーニングするために昼間のセマンティック ラベルを直接再利用できるようになります。特に、事前にトレーニングされた対照言語画像事前トレーニング (CLIP) モデルを組み込むことにより、提案されたフレームワークは、昼夜を問わず翻訳中に意味の一貫性を維持するように設計されています。さらに、夜間シーンでの NIR 照明の有効範囲が限られていることを考慮して、可視性マスクが導入されています。私たちは、最先端の画像翻訳ベースラインとの比較評価を実施し、夜間の視覚ナビゲーションのための下流のセマンティック セグメンテーションのパフォーマンスの向上によってサポートされる、より高い画像品質を実証します。評価には、農業圃場で暗視機能を備えた移動ロボットを使用して収集され、ピクセル単位のセマンティックラベルが手動で注釈付けされた428枚の昼間画像と549枚の夜間画像で構成される新しいデータセットであるAgriNightを利用し、夜間農業視覚ナビゲーションの最初のベンチマークとして導入します。また、夜間に稼働する物理ロボットによるリアルタイム自律航行実験も行っています。データとコードは https://github.com/mamorobel/AgriNight から入手できます。
原文 (English)
Enabling 24-hour Agricultural Robotics: Unsupervised Day-to-Night Cross-Modal Image Translation for Nighttime Visual Navigation
While visual navigation has been extensively studied in agricultural robotics, most existing systems assume daytime conditions. In fact, deploying autonomous robots at night offers significant advantages, including 24-hour crop and soil monitoring, fruit harvesting, and nocturnal pest detection. Modern vision-based systems, however, rely heavily on large-scale well-annotated image datasets, which remains challenging to obtain for nighttime operation scenarios. To address this, we propose an unsupervised image translation framework that converts daytime plant-row RGB images into near-infrared (NIR) nighttime counterparts without requiring pixel-to-pixel supervision. This enables the direct reuse of daytime semantic labels for training nighttime perception models. In particular, by incorporating a pre-trained Contrastive Language-Image Pre-training (CLIP) model, the proposed framework is designed to preserve semantic consistency during day-to-night translation. Additionally, a visibility mask is introduced to account for the limited effective range of NIR illumination in nighttime scenes. We conduct comparative evaluations with state-of-the-art image translation baselines and demonstrate higher image qualities, as supported by improved performance in downstream semantic segmentation for nighttime visual navigation. For evaluation, we utilize AgriNight--a novel dataset comprising 428 daytime and 549 nighttime images collected using night-vision-equipped mobile robots in agricultural fields and manually annotated with pixel-wise semantic labels--and introduce it as the first benchmark for nighttime agricultural visual navigation. We also perform real-time autonomous navigation experiments with a physical robot operating at night. The data and code are available at: https://github.com/mamorobel/AgriNight.
データセットシフトの下での ROI ベースの甲状腺結節超音波分類のためのディープアンサンブルを使用した校正済みの選択的予測: 遡及的評価
背景: 深層学習モデルは超音波で甲状腺結節を分類できますが、信頼性の高い臨床意思決定のサポートには、特にデータセットのシフト下では、校正された確率、不確実性の推定、選択的参照も必要です。方法: ROI ベースの甲状腺結節分類と選択的画像ベースのトリアージ用に、校正された決定論的な 5 メンバーのディープ アンサンブルを開発しました。 TN5000 は、モデル開発、5 倍交差検証、メンバーごとのベクトル スケーリング キャリブレーション、および倍数固有のしきい値の選択に使用されました。 TN3K は、独立した外部データセット シフト評価として機能しました。このフレームワークでは、圧迫と興奮の注意、アンサンブル平均悪性確率、およびアンサンブル不一致スコアとして相互情報量 (MI) を備えた ConvNeXt-Tiny を使用しました。 3 段階のポリシーにより、画像は No-FNA 提案、FNA 推奨、または放射線科医のレビューに割り当てられました。結果: プールされたアウトオブフォールド TN5000 予測では、アンサンブルは AUC-ROC 0.9395、AP 0.9715、ECE 0.0088、および Brier スコア 0.0813 を達成しました。名目上のMI滞留率50%では、症例の7.2%がNo-FNAの提案、39.9%がFNAの推奨、52.9%が放射線科医の審査を受け、98.3%がNo-FNAのNPV、99.83%の悪性腫瘍捕捉率でした。 TN3K では、AUC-ROC は 0.7870 に減少し、AP は 0.7254 に減少し、ECE は 0.1899 に増加し、Brier スコアは 0.2281 に増加しました。凍結された TN5000 ポリシーでは、83.7% が見直し、1.0% が No-FNA、15.3% が FNA 推奨に割り当てられました。 No-FNA 経路に入った悪性画像はありませんでしたが、FNA 推奨 PPV は 76.6% に低下しました。結論: このフレームワークは強力な内部識別とキャリブレーションを示しましたが、外部閾値の輸送性は制限されていました。選択的予測は、自動トリアージに適さない画像を特定するのに役立ちますが、導入前にローカルでの再調整、しきい値の検証、および前向きの臨床評価が必要です。
原文 (English)
Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation
Background: Deep learning models can classify thyroid nodules on ultrasound, but reliable clinical decision support also requires calibrated probabilities, uncertainty estimation, and selective referral, particularly under dataset shift. Methods: We developed a calibrated deterministic five-member deep ensemble for ROI-based thyroid nodule classification and selective image-based triage. TN5000 was used for model development, five-fold cross-validation, member-wise vector-scaling calibration, and fold-specific threshold selection. TN3K served as an independent external dataset-shift evaluation. The framework used ConvNeXt-Tiny with squeeze-and-excitation attention, ensemble-mean malignancy probability, and mutual information (MI) as an ensemble-disagreement score. A three-tier policy assigned images to No-FNA suggestion, FNA recommendation, or radiologist review. Results: On pooled out-of-fold TN5000 predictions, the ensemble achieved AUC-ROC 0.9395, AP 0.9715, ECE 0.0088, and Brier score 0.0813. At 50% nominal MI retention, 7.2% of cases received a No-FNA suggestion, 39.9% an FNA recommendation, and 52.9% radiologist review, with 98.3% No-FNA NPV and 99.83% malignancy capture. On TN3K, AUC-ROC decreased to 0.7870, AP to 0.7254, ECE increased to 0.1899, and Brier score to 0.2281. The frozen TN5000 policy assigned 83.7% to review, 1.0% to No-FNA, and 15.3% to FNA recommendation. No malignant image entered the No-FNA pathway, but FNA-recommendation PPV fell to 76.6%. Conclusion: The framework showed strong internal discrimination and calibration, but limited external threshold transportability. Selective prediction may help identify images unsuitable for automated triage, but local recalibration, threshold validation, and prospective clinical evaluation are required before deployment.
解釈可能な分布外検出のためのスパース オートエンコーダ
機械学習モデルを安全に導入するには、配布外 (OOD) サンプルを確実に検出することが重要です。ニューラル ネットワークは、トレーニング データから逸脱した入力に対して自信過剰な予測を生成することが多く、パフォーマンスの大幅な低下につながります。多くの OOD 検出方法は最終出力層に焦点を当てていますが、中間ネットワーク層に存在する豊富な階層情報は無視されています。この論文では、スパース オートエンコーダ (SAE) を活用して、これらの中間アクティベーションから解釈可能な特徴を学習する新しいアプローチを紹介します。インディストリビューション (ID) および OOD データが、これらのまばらな特徴の個別のセットを活性化することがわかりました。我々は、テストサンプルのまばらな特徴アクティベーションとIDクラスの平均アクティベーションの間のコサイン類似度から導出される新しいOODスコアを提案します。当社のポストホック検出手法は、標準的な OOD 検出ベンチマークで最先端のパフォーマンスを達成するだけでなく、分布のシフトが学習された表現にどのような影響を与えるかについて解釈可能な洞察をもたらします。
原文 (English)
Sparse Autoencoders for Interpretable Out-of-Distribution Detection
Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models. Neural networks often produce overconfident predictions for inputs that deviate from their training data, leading to significant degradation in performance. While many OOD detection methods focus on the final output layer, they neglect the rich hierarchical information present in intermediate network layers. This paper introduces a novel approach that leverages sparse autoencoders (SAEs) to learn interpretable features from these intermediate activations. We find that in-distribution (ID) and OOD data activate distinct sets of these sparse features. We propose a new OOD score derived from the cosine similarity between the sparse feature activations of a test sample and the mean activations of ID classes. Our post-hoc detection method not only achieves state-of-the-art performance on standard OOD detection benchmarks, but yields interpretable insights into how distribution shift affects learned representations.
PFAdapter: パーソナライズされたフェデレーション MLLM のための階層型 LoRA 分解
エージェント AI システムは、ネットワーク エッジでデータのプライバシーを維持しながら、共同学習が可能な自律型インテリジェント エージェントを展開することで、通信とネットワーキングを再構築しています。分散ネットワーク環境内では、マルチモーダル大規模言語モデル (MLLM) がエッジ デバイスのコグニティブ エンジンとして機能しますが、フェデレーテッド微調整は、異種ネットワーク条件下でグローバルな知識の集約とローカルな適応のバランスをとるという大きな課題に直面しています。従来のフェデレーション プロトコルは通常、均一なパラメーターの集約に依存しており、ドメインに不変の機能とクライアント固有のニュアンスが混同されるため、最適ではないパーソナライゼーションと過剰な通信オーバーヘッドが生じます。これらの課題に対処するために、私たちは PFAdapter を提案します。これは、アダプター パラメーターをグローバル共有コンポーネントとローカルプライベート コンポーネントに明示的に分離するための階層型 LoRA 分解を導入した通信効率の高いフレームワークです。クエリとキーの投影は、ネットワーク全体で普遍的なマルチモーダル セマンティクスをキャプチャするためにグローバル同期に割り当てられますが、値と出力の投影はエッジ固有の適応のためにローカライズされたままになります。さらに、フロベニウス ノルムに基づく直交性正則化により、これらのコンポーネント間の厳密な分離が強制され、冗長な特徴学習が防止されます。選択的集約プロトコルは、フェデレーション ネットワーク全体でグローバルに共有されるコンポーネントのみを同期するため、ローカルの専門知識が維持され、通信コストが 50% 近く削減されます。 VQA-RAD、SLAKE、Hateful Memes、CrisisMMD データセットに関する広範な実験では、PFAdapter が常に最先端のベースラインを上回り、さまざまなエッジ インテリジェンス タスク全体で 2.4% から 4.8% の範囲の精度向上を達成していることが実証されています。その結果、私たちのフレームワークは、リソースに制約のある通信ネットワークにおけるエージェント AI 導入のための効率的なソリューションを確立します。
原文 (English)
PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs
Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges. Within distributed network environments, Multimodal Large Language Models (MLLMs) serve as cognitive engines for edge devices, yet federated fine-tuning faces substantial challenges in balancing global knowledge aggregation with local adaptation under heterogeneous network conditions. Conventional federated protocols typically rely on uniform parameter aggregation, which conflates domain-invariant features with client-specific nuances, thereby resulting in suboptimal personalization and excessive communication overhead. To address these challenges, we propose PFAdapter, a communication-efficient framework introducing hierarchical LoRA decomposition to explicitly separate adapter parameters into global-shared and local-private components. Query and key projections are assigned to global synchronization for capturing universal multimodal semantics across the network, while value and output projections remain localized for edge-specific adaptation. Additionally, orthogonality regularization based on the Frobenius norm enforces strict separation between these components, preventing redundant feature learning. Selective aggregation protocols synchronize only global-shared components across the federated network, preserving local expertise and reducing communication costs by nearly 50%. Extensive experiments on VQA-RAD, SLAKE, Hateful Memes, and CrisisMMD datasets demonstrate that PFAdapter consistently outperforms state-of-the-art baselines, achieving accuracy improvements ranging from 2.4% to 4.8% across diverse edge intelligence tasks. Consequently, our framework establishes an efficient solution for agentic AI deployment in resource-constrained communication networks.
フェデレーテッド MLLM 微調整のための弾性正則化と合成再生による継続的学習
分散ネットワーク全体でマルチモーダル大規模言語モデル (MLLM) を連携して微調整することで、進化するデータ ストリームへのプライバシーに配慮した適応が可能になりますが、根本的な障害が動的環境での堅牢な展開を妨げています。それは、壊滅的な忘却です。連続したタスクの更新により、視覚的、言語的、クロスモーダルな表現にわたって以前に取得した知識が消去されます。この課題に対処することは、事前の知識を確実に保持することがシステムの完全性を支えるコンテンツ管理など、安全性が重視されるドメインで動作する自律型ネットワーク AI にとって特に重要です。これを克服するために、私たちは Federated Continual Multimodal Learning (FedCMM) を提案します。これは、3 つの相補的なレベルでフェデレーテッド最適化ループに継続学習の保護手段を組み込むフレームワークです。パラメータ レベルでは、モダリティを意識した弾性重み統合により、ビジョン エンコーダ、言語バックボーン、クロスモーダル プロジェクターの個別のフィッシャー情報行列が計算され、モダリティ固有の忘却に対するきめ細かな非対称性を意識した保護が提供されます。データ レベルでは、各クライアントは軽量のローカル生成再生モジュールをトレーニングして、生データを共有せずに生データのない埋め込みレベルのマルチモーダル再生タプルを合成します。集約レベルでは、タスクの類似性を認識した勾配集約により、勾配コサイン類似度によってクライアント更新が自動的にフィルタリングおよび再重み付けされ、矛盾する方向が抑制され、グローバルな学習軌道が安定します。 2 つのベンチマークに関する広範な実験により、FedCMM が精度と逆方向転送に関して最近のベースラインを常に上回っていることが実証され、全体的でモダリティを意識した最適化により、異種ネットワーク化された AI 導入全体にわたって堅牢な進化的適応が可能になることが確認されました。
原文 (English)
Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning
Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential task updates erase previously acquired knowledge across visual, linguistic, and cross-modal representations. Addressing this challenge is especially critical for autonomous networked AI operating in safety-sensitive domains, such as content moderation, where reliable retention of prior knowledge underpins system integrity. To overcome this, we propose Federated Continual Multimodal Learning (FedCMM), a framework that embeds continual-learning safeguards into the federated optimization loop at three complementary levels. At the parameter level, modality-aware elastic weight consolidation computes separate Fisher information matrices for the vision encoder, language backbone, and cross-modal projector, providing granular, asymmetry-aware protection against modality-specific forgetting. At the data level, each client trains a lightweight local generative replay module to synthesize raw-data-free embedding-level multimodal replay tuples without any raw data sharing. At the aggregation level, Task-similarity-aware gradient aggregation autonomously filters and reweights client updates by gradient cosine similarity, suppressing conflicting directions and stabilizing the global learning trajectory. Extensive experiments on two benchmarks demonstrate that FedCMM consistently outperforms recent baselines on accuracy and backward transfer, confirming that holistic, modality-aware optimization enables robust evolutive adaptation across heterogeneous networked AI deployments.
信頼できる自律科学に向けて: 2 年間のコミュニティ ロードマップ
1 年前の AISLE ロードマップでは、自律的な研究所は孤立した島のように運営され、5 つの重要な側面を中心に組織された草の根ネットワークが提案されました。その後、この分野は予想よりも早く進んでいます。マルチエージェント システムは実験的に検証された仮説を生み出し、自動運転研究所はより相互運用性と統合性を高め、推論訓練されたドメイン基盤モデルは能力の上限を引き上げ、ジェネシス ミッションは自律実験を米国連邦科学戦略の中心に据え、産業界が主要な主体として浮上しました。進歩は、訂正された主力発見結果、クローズドエンド型の質問で専門家に匹敵するエージェントが依然としてオープンエンド型研究のほんの一部しか完了していないことを示すベンチマーク、主要な会場で表面化した捏造された引用など、厳粛な逆流に遭遇している。私たちはこれをこの分野の決定的な緊張と解釈します。発見の候補を生み出すことはもはや難しいことではありませんが、それを検証することは難しいことであり、この非対称性により、生のモデルの機能以上に自律科学が制限されるようになりました。私たちはロードマップを 7 つの側面に沿って更新し、元の 5 つの側面を再考し、以前の 2 つの分野横断的な懸念事項である信頼、検証、再現性、および安全性、セキュリティ、ガバナンスを第一級のステータスに引き上げます。私たちは、元のマイルストーン (M1 ~ M14) を達成、部分的に達成、再構築、またはオープンとして評価し、4 つの新しいマイルストーン (M15 ~ M18) を追加し、今後 2 年の期間に向けての方向性を検討します。 1 年目はインターフェイス、プロトコルの採用、検証の足場に重点を置き、2 年目はフェデレーション、ゼロトラスト調整、ガバナンスを対象とします。全体を通して、私たちは草の根ネットワークを、国家プログラム、国際的な取り組み、商用プラットフォームをサイロ化するのではなく接続できる相互運用性の構造として位置づけています。
原文 (English)
Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap
One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions. The field has since moved faster than anticipated. Multi-agent systems have produced experimentally validated hypotheses, self-driving laboratories have grown more interoperable and orchestrated, reasoning-trained and domain foundation models have raised the capability ceiling, and the Genesis Mission has placed autonomous experimentation at the center of U.S. federal science strategy, with industry emerging as a primary actor. Progress has met a sobering counter-current, including a corrected flagship discovery result, benchmarks showing that agents which rival experts on closed-ended questions still complete only a fraction of open-ended research, and fabricated citations surfacing at leading venues. We read this as the defining tension of the field. Producing a candidate discovery is no longer the hard part, but verifying it is, and this asymmetry now limits autonomous science more than raw model capability. We update the roadmap around seven dimensions, revisiting the original five and elevating two former cross-cutting concerns, trust, verification, and reproducibility, and safety, security, and governance, to first-class status. We assess the original milestones (M1 through M14) as achieved, partially achieved, reframed, or open, add four new milestones (M15 through M18), and scope the path forward to a two-year horizon. The first year concentrates on interfaces, protocol adoption, and the scaffolding of verification, and the second targets federation, zero-trust coordination, and governance. Throughout, we position the grassroots network as the interoperability fabric that lets national programs, international initiatives, and commercial platforms connect rather than re-silo.
GaitSpan: 歩行から走行までのヒューマノイドの移動動作の成長
歩くことができるヒューマノイドは、ジョギングやランニングのために移動運動を最初から学び直す必要はありません。しかし、現在のアプローチは、歩行スケジュールを規定したり、モーションクリップを模倣したり、専門家を訓練してスキルを切り替えたり、スキルを 1 つのポリシーに抽出したりすることによって、歩行の多様性を獲得することがよくあります。これらの戦略は印象的な動作を生み出すことができますが、連続的な速度コマンド、地形、および形態にわたる柔軟性には限界があります。私たちは、事前トレーニングされた基本的な歩行ポリシーをより速い移動に拡張するフレームワークである GaitSpan を使用してスキルの成長を研究します。歩行をシードスキルとして扱います。これは、バランス、サポート、体の調整、接触の移行のための再利用可能な運動構造であり、新しいリズムで再生し、より長い/より高い歩幅に拡張し、残留適応によって修正することができます。この拡張には 3 つの側面があります。1) リズム生成。複数の内部クロックを使用してフリーズ ウォーキング ポリシーを調整し、結果として得られる正規アクションのコマンド条件付きの組み合わせを学習します。 2) ストライドシェイピング。バネ仕掛けの倒立振子のダイナミクスにヒントを得た、物理的に接地された対物レンズを使用して、より高い指令速度に適した動的な移動パターンを与えます。 3) 残差適応。リズムの生成やストライドの形成では考慮されない動作の詳細を捕捉します。 GaitSpan は、連続的な速度範囲をカバーするウォーキング、ジョギング、ランニングのようなレジームにまたがる単一のコマンド条件付きヒューマノイド ポリシーを初めて提供し、形態を超えて移動し、目に見えないシミュレーション間および現実世界の地形でゼロショットを展開します。複数の専門家によって訓練されたベースライン、または人間の模倣によって訓練されたベースラインと比較して、より速く学習し、より強力な歩行パフォーマンスを実現します。
原文 (English)
GaitSpan: Growing Humanoid Locomotion from Walking to Running
A humanoid that can walk should not relearn locomotion from scratch to jog or run. Yet current approaches often obtain gait diversity by prescribing gait schedules, imitating motion clips, training experts to switch between or distilling skills into one policy. These strategies can produce impressive behaviors, but offer limited flexibility across continuous speed commands, terrains, and morphologies. We study skill growth with GaitSpan, a framework that expands a pretrained, basic walking policy into faster locomotion. It treats walking as a seed skill: reusable motor structure for balance, support, body coordination, and contact transition that can be regenerated at new rhythms, extended into longer/higher strides, and corrected by residual adaptation. This expansion has three aspects: 1) rhythm generation, which modulates the frozen walking policy with multiple internal clocks and learns command-conditioned combinations of the resulting canonical actions; 2) stride shaping, which rewards dynamic locomotion patterns appropriate for higher commanded speeds using a physically grounded objective inspired by spring-loaded inverted pendulum dynamics; and 3) residual adaptation, which captures motion details not accounted for by rhythm generation or stride shaping. GaitSpan is the first to deliver a single command-conditioned humanoid policy that spans walking, jogging, and running-like regimes covering a continuous speed range, transfers across morphologies, and deploys zero-shot on unseen sim-to-sim, and real-world terrains. Compared with baselines either trained with multi-experts or imitation from humans, it learns faster and achieves stronger gait performance.
Self-Consistent Flow: Unifying Velocity and Endpoint Prediction for Rectified Flow Models
In rectified-flow-based generative models, the neural network can be trained to predict two different targets, such as the instantaneous ve…
From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data
X-ray tomography enables nondestructive characterization of material microstructures, while advances in micro-CT imaging have accelerated v…
TRAIL: A Platform for Configurable Human--AI Teaming Experiments
An AI teammate's design properties (personality, communication style, when it speaks) can shape a team's trust, coordination, and decisions…
Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing
Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans…
The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests
We show that the Benjamini--Hochberg procedure can fail to control the false discovery rate (FDR) at its nominal level for correlated two-s…
RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls
Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and…
Partial Identification with Multiple Nonlinear Measurements of a Latent Regressor
We study linear regression when the regressor is latent and observed only through multiple noisy measurements, each a smooth but possibly n…
Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals
Large language model (LLM) trading agents show promising performance in equity markets, yet remain narrowly focused on US equities with lit…
Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents
Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy…
Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs
As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-wor…
The Sound of Absence: Audio-Language Embedding Models Struggle with Negation
Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this af…
A Longitudinal Analysis of Public Discourse on AI Ethics in Education Using Twitter Data
The rapid integration of artificial intelligence (AI) and generative AI (GenAI) into education presents significant opportunities to enhanc…
A Comparative Analysis of Institutional and Course Generative AI Policies within Higher Education: Implications for Instruction in Computing Education
With the increased use of generative AI (GenAI) applications such as ChatGPT, higher education institutions (HEIs) have released a range of…
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Ans…
Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)
Artificial Intelligence (AI) technologies, while serving as a foundational enabler for modern social media and digital health services, exe…
Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction
EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning. Yet EEG-derived reconstructions are blurry, distorted,…
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent ap…
Demonstration of the common dual-channel feature decoupling characteristic of front-door mediation causal inference methods in whole-slice image classification
Causal inference using front door intervention and multi-instance learning (MIL) has advanced the analysis of Whole Slide Images (WSI) in d…
ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning
Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that dept…
The Computational Basis of Confidence in Large Language Models
Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language mode…
An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge
We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Lar…
Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric
Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor s…
OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning
Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow…
Mind the Gap: Promises and Pitfalls of Hierarchical Planning in LeWorldModel
We investigate whether temporal hierarchy can improve LeWorldModel on long-horizon goal-conditioned control. We introduce Hi-LeWM, an exten…
Traceback Translators Against Forgetting in Continual Fake Speech Detection
Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem,…
Deep Learning-based Surrogate Modelling of the LOD Method for Multiscale Problems
Multiscale problems are notoriously difficult to tackle using traditional numerical methods, as accurately resolving fine-scale features of…
Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction
The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics…
Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs
Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are ofte…
Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?
As Large Language Models (LLMs) are increasingly deployed as autonomous agents in high-stakes domains, understanding contextual factors tha…
Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs
Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted…
Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference
Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-po…
Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings
Computer-Aided Design (CAD) floor plan drawings contain both graphical primitives and textual annotations, which provide complementary geom…
From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation
LLMs can perform language-based quantitative prediction from unstructured inputs, but remain susceptible to hallucinations and overconfiden…
Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts
Sparse Mixture-of-Experts (MoE) models have become an important approach for scaling Large Language Models (LLMs), but their inference effi…
Line-Anchored Feedback Cuts Token Costs and Improves Correctness in AI Code Editing
Generated tokens are a direct driver of the cost, latency, and energy of generative AI (GAI) code editing. We show the format of feedback i…
Bulkhead: Automated Semantic Detection and Remediation of Container Escape Vulnerabilities
Filesystem isolation in container ecosystems is often weakened by cross-boundary path misresolution, causing path traversal (PaTra) vulnera…
Learning-based Probabilistic Load Forecasting with Post-hoc and In-model Uncertainty
Smart-building load forecasters are often trained offline on dense, multivariate, high-frequency data, but deployment may provide only hour…
Weakly Supervised Spatio-Temporal Candidate Discovery of Dairy Farm Sites from Seasonal Satellite Imagery
Farm site discovery from satellite imagery is a spatiotemporal candidate ranking problem because farm evidence is distributed across pastur…
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision wi…
Practical Judgment, Virtue, and Intuition in the Use of Opaque AI-Enabled Systems
AI-enabled systems are seeing increasing deployment across numerous domains, with many being "black boxes" with respect to core functions a…
Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination
Federated Reinforcement Learning (FedRL) enables coordination of distributed energy resources without sharing raw local data, but standard…
HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition
This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition. For multi-task learning with sim…
When Close Enough Is Not Enough: Autoregressive Drift in Quantum Circuit Synthesis
Quantum circuit optimization for fault-tolerant computing requires exact functional equivalence while minimizing expensive non-Clifford res…
Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels
Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, howe…
The One-Word Census: Answer-Choice Conformity Across 44 Language Models
When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same…
Autonomous Tracking and Terminal Guidance of Moving Targets for Fixed-Wing UAVs
This study introduces a unified control framework for fixed-wing unmanned aerial vehicles (UAVs) fitted with a pan-tilt (PT) camera, intend…
PixelLoop: Shortcut Topological Navigation with Pixel-Level Loops
Although topological mapping and navigation have been studied extensively, the specific role and downstream effect of loop closures in pure…
Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques
Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However,…
Reproducible Reservoir Computing with Thermally Driven Superparamagnets: Controlling Temperature Sensitivity
Unconventional computing systems must demonstrate robust performance under real-world environmental conditions to enable practical deployme…
ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation
A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty. Refere…
Unveiling Complex Collective Behaviors from Simple Rewards
Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates s…
UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies
Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, guide policy learning,…
Real-time fall detection based on vision for low-power edge platforms
Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame i…
ViHoRec: A Quality-Controlled Vietnamese Hotel Recommendation Dataset and Cold-Start Benchmark
Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interaction resource. Building…
Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models
Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo c…
PalmClaw: A Native On-Device Agent Framework for Mobile Phones
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the resu…
TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale
Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to…
DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents
Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary…
Rethinking Reward Models for Multi-Domain Test-Time Scaling
The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward m…
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
While multimodal large language models can describe visual content, their ability to generate executable procedures remains underexplored.…
RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories
Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often…
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide ri…
Calculating Mutual Information between a Reward Maximizer and its Environment
An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world. In t…
Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation
Simulating large-scale human mobility is fundamental to understanding population movement patterns and supporting real-world geospatial app…
Learning When to Trust in Contextual Social Bandits
Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global bu…
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
Remote sensing underpins crucial applications such as disaster relief and ecological field surveys, where systems must understand complex s…
ReLope: KL-Regularized LoRA Probes for Multimodal LLM Routing
Routing has emerged as a promising strategy for balancing performance and cost in large language model (LLM) systems that combine lightweig…
Quantification of Credal Uncertainty: A Distance-Based Approach
Credal sets, i.e., closed convex sets of probability measures, provide a natural framework to represent aleatoric and epistemic uncertainty…
Mistake gating leads to energy and memory efficient continual learning
Synaptic plasticity is metabolically expensive, yet animals continuously update their internal models without exhausting energy reserves. H…
Action-Aware Generative Sequence Modeling for Short Video Recommendation
With the rapid development of the Internet, users have increasingly higher expectations for the recommendation accuracy of online content c…
Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR
As LLMs become credible readers of earnings calls, investor-relations Q\&A, guidance, and disclosure language, supervised financial NLP ben…
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
Language models draw on two knowledge sources: facts baked into weights (parametric memory, PM) and information in context (working memory,…
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on whic…
Brain Vascular Age Prediction Using Cerebral Blood Flow Velocity and Machine Learning Algorithms
Defining vascular age in terms of physiological function has become one focal point of the extensive studies to categorize and track chrono…
推論によるフロンティア LLM 評価の形状計算方法
AI の評価は、ツールの使用と反復的な問題解決を伴う長期にわたる軌道から恩恵を受ける、より困難なタスクへと移行しています。その結果、パフォーマンスは、テスト時に利用可能なコンピューティング (「推論コンピューティング」) の量と割り当てにますます敏感になります。しかし、多くの評価では依然として単一の制限された予算でのパフォーマンスが報告されており、低いスコアはモデルの基礎的な機能ではなく評価設定を反映している可能性があることを意味します。これをテストするために、ソフトウェア エンジニアリング、数学、医学、サイバーセキュリティにわたる 7 つの挑戦的なベンチマークで最大 12 のフロンティア言語モデルを評価します。私たちは、3 つの単純な推論スケーリング介入を組み合わせた制御されたセットアップを使用します。つまり、より大きなトークン バジェット、コンテキストの圧縮、およびモデル自体または最小限の正確性フィードバックによって導かれる送信の試行の繰り返しです。主な結果は 3 つあります。まず、トークン バジェットが大きくなると、サイバーセキュリティ、FrontierMath、人類最後の試験、ターミナルベンチなど、複数のドメインにわたるベンチマークのパフォーマンスが大幅に向上します。第二に、固定予算の評価では、モデルが進歩するにつれてフロンティアの能力がますます過小評価される可能性があります。新しいモデルは、大きな予算でより高いパフォーマンスを実現し、より困難なタスクを解放し、より確実に解決します。第三に、どの推論スケーリング手法が最も役立つかがベンチマークによって異なります。繰り返し送信するとパフォーマンスが大幅に向上しますが、より大きなトークン バジェット、外部フィードバック、および並列試行の値はベンチマークによって異なります。全体として、私たちの結果は、ベンチマーク スコアがプロトコルに依存していることを示しています。したがって、評価では、特に安全性またはポリシー関連の設定において、推論時間のコンピューティングの関数として機能を報告し、プロトコルの選択を明示的に指定し、一致した予算で大規模な共有コンピューティング範囲にわたってモデルの世代を比較する必要があると主張します。
原文 (English)
How Inference Compute Shapes Frontier LLM Evaluation
AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still report performance at a single restrictive budget, meaning that low scores may reflect the evaluation setup rather than the model's underlying capability. To test this, we evaluate up to 12 frontier language models on seven challenging benchmarks spanning software engineering, mathematics, medicine, and cybersecurity. We use a controlled setup combining three simple inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts, guided either by the model itself or by minimal correctness feedback. We find three main results. First, larger token budgets substantially improve performance on benchmarks across multiple domains, including cybersecurity, FrontierMath, Humanity's Last Exam, and TerminalBench. Second, fixed-budget evaluations can increasingly understate frontier capability as models advance. Newer models reach higher performance at large budgets, where they unlock harder tasks and solve them more reliably. Third, benchmarks differ in which inference-scaling methods help most: repeated submission broadly improves performance, but the value of larger token budgets, external feedback, and parallel attempts varies by benchmark. Overall, our results show that benchmark scores are protocol-dependent. We therefore argue that evaluations should report capability as a function of inference-time compute, specify protocol choices explicitly, and compare model generations over a large shared compute range at matched budgets, especially in safety- or policy-relevant settings.
Matilda: Engine-Agnostic Search with Human Policy Guidance
Chess engines have evolved from search-based systems optimized solely for strength to neural policies capable of modeling human decisions a…
マルチエージェント LLM チームにとってパーソナリティ構成が重要になるのはどのような場合ですか?
パーソナリティプロンプトは、大規模な言語モデルがどのようにコミュニケーションするかを形成しますが、これらの行動の変化が客観的なタスクの結果に影響を与えるかどうかは、まだ十分に調査されていません。これまでの研究では、低い同調度で促されたエージェントは敵対的な言葉を発し、高い同調度で促されたエージェントは協力的になることが示されているが、コミュニケーションスタイルとタスクパフォーマンスの関係は複数の領域にわたって体系的に調査されていない。この研究では、構造化コーディング、無制限の研究協力、競争的交渉という 3 つのタスク ドメインでフロンティア LLM 全体の性格特性を操作することにより、性格構成がマルチエージェント チームのパフォーマンスに重要であるかどうかを調査します。性格への影響はタスクの構造に大きく依存することがわかりました。コーディングタスクでは、協調性が低いとコミュニケーションに大きな変化が生じ、マイルストーンの完了にはほとんど影響しません。無制限のコラボレーションや交渉では、同じ操作がパフォーマンスを大幅に低下させます。マルチエージェントシステム設計への影響と人格操作の限界について説明します。
原文 (English)
When Does Personality Composition Matter for Multi-Agent LLM Teams?
Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-explored. Prior work shows that agents prompted with low agreeableness produce adversarial language, while those prompted with high agreeableness become cooperative, but the relationship between communication style and task performance has not been systematically examined across multiple domains. In this work, we investigate whether personality composition matters for multi-agent team performance by manipulating personality traits across frontier LLMs on three task domains: structured coding, open-ended research collaboration, and competitive bargaining. We find that personality effects depend critically on task structure. In coding tasks, low agreeableness leads to large communication shifts that have little effect on milestone completion. In open-ended collaboration and bargaining, the same manipulation substantially degrades performance. We discuss implications for multi-agent system design and the limits of personality manipulation.
OSWorld2.0: 長期にわたる現実世界のタスクにおけるコンピュータ使用エージェントのベンチマーク
既存のコンピュータ使用ベンチマークは、現実世界のコンピュータ使用の現実性、複雑さ、長期的な要求を捉えることができず、フロンティア エージェントの限界を明らかにする能力が制限されています。 OSWorld 2.0 は、複雑で困難な現実世界の現象を捉えるように設計された、日常業務から専門業務にわたる 108 の長期的なコンピューター使用ワークフローのベンチマークです。各タスクは現実的なエンドツーエンドのワークフローを表しており、人間のユーザーが完了するまでに中央値で約 1.6 時間かかり、Claude Opus 4.7 では最大限の思考を使用して平均 318 回のツール呼び出しが必要ですが、OSWorld 1.0 では約 30 回です。 OSWorld 2.0 は、ストリーミング インタラクションや動的環境などのインタラクション設計の課題だけでなく、クロスソース推論、暗黙的状態推論、視覚空間精度などのエージェント パターンの課題にも及ぶ、実際のワークフローでは一般的であるものの、以前のベンチマークでは過小評価されていた課題現象をターゲットにしています。タスクは本物の入力アーティファクトに基づいており、現実的なステートフル ユーザー プロファイル データと相互参照され、安全性を重視した実行を監査する個別の安全性レポートが含まれています。 500 ステップの主要なバイナリ完了基準では、最大限の思考とバッチ化されたツールを備えた Claude Opus 4.8 が最高のスコアを示していますが、それでも部分スコア 54.8% でタスクの 20.6% しか完了していません。 GPT-5.5 はトークン効率がはるかに優れていますが、13% 付近で頭打ちになっています。これらの結果は、現在のエージェントがまだプロレベルのコンピュータ使用には程遠いことを示しています。基本的な GUI 制御やコーディングでつまずくのではなく、制約を見失い、タスクの途中で届く情報を見逃し、ユーザーに尋ねるのではなく推測し、検証をスキップし、タスクが回復する必要がある隠れた状態に左右されるときに最も苦労します。
原文 (English)
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmark of 108 long-horizon computer-use workflows across everyday and professional tasks, designed to capture complex and challenging real-world phenomena. Each task represents a realistic end-to-end workflow that takes human users a median of about 1.6 hours to complete and requires an average of 318 tool calls with Claude Opus 4.7 using maximum thinking, compared with about 30 in OSWorld 1.0. OSWorld 2.0 targets challenge phenomena that are common in real workflows yet underrepresented in prior benchmarks, spanning interaction-design challenges such as streaming interaction and dynamic environments, as well as agent-pattern challenges such as cross-source reasoning, implicit-state inference, and visual-spatial precision. Tasks are grounded in authentic input artifacts and cross-referenced against realistic stateful user profile data, and include separate safety reports auditing safety-sensitive execution. Under our primary binary-completion metric at 500 steps, Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score; GPT-5.5 is far more token-efficient yet plateaus near 13%. These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden state they must recover.
DRIFTLENS: パーソナライズされた言語モデルにおける記憶誘発推論ドリフトの測定
パーソナライゼーションにより、モデルがユーザーに伝える内容が変わります。反応を正当化するために使用される推論の軌道も変更できることを示します。最新の LLM は、ユーザーの属性、設定、以前のコンテキストを保存し、この情報を将来のプロンプトに挿入することによって、インタラクションをパーソナライズします。私たちは、そのような記憶が、単一の真実の答えが存在しない自由形式の質問に対する推論を再構成するかどうかを研究します。この効果を定量化するために、表現された各推論ステップを値カテゴリにマッピングし、質問の記憶のない軌道と注入されたユーザー属性記憶の下での軌道との間の乖離を測定する、グラウンドトゥルースフリーのフレームワークである DRIFTLENS を導入します。まず、DRIFTLENS が内容のない実用的なノイズと実質的な推論の変更を区別することを検証します。年齢、職業、障害を含む 4 つの LLM と 10 のユーザー属性カテゴリにわたって、ユーザー属性の記憶は、最終的な回答が流暢で、主題に合致しており、もっともらしいままである場合でも、各モデルの実用的なノイズ フロアを超える中規模から大規模な推論のドリフトを引き起こします。次に、ドリフトを軽減するための GRPO および DPO ベースのポストトレーニング方法を評価します。どちらもドリフトを軽減しますが、どちらも均一に支配的ではありません。下流の能力、有用性、指示への従うことへの影響は、モデルと報酬に依存します。これらの結果は、記憶に起因する推論ドリフトは測定可能であり、部分的にのみ緩和されるパーソナライズされた言語モデルの障害モードであることを示唆しています。
原文 (English)
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future prompts. We study whether such memory reshapes reasoning on open-ended questions where no single ground-truth answer exists. To quantify this effect, we introduce DRIFTLENS, a ground-truth-free framework that maps each expressed reasoning step to a value category and measures divergence between a question's no-memory trajectory and its trajectory under injected user-attribute memory. We first validate that DRIFTLENS distinguishes content-free pragmatic noise from substantive reasoning changes. Across four LLMs and 10 user-attribute categories, including age, occupation, and disability, user-attribute memory induces medium-to-large reasoning drift above each model's pragmatic-noise floor, even when final answers remain fluent, on-topic, and plausible. We then evaluate GRPO- and DPO-based post-training methods for reducing drift. Both reduce drift, but neither uniformly dominates; effects on downstream capability, helpfulness, and instruction following are model-and reward-dependent. These results suggest that memory-induced reasoning drift is a measurable and only partly mitigated failure mode of personalized language models.
TOFFEE のデモンストレーション: データ エージェントの軌跡を大規模に合成するための学習済みシステム
LLM を活用したデータ エージェントは、データ主導の意思決定においてますます重要な役割を果たしています。しかし、既存のデータ エージェントは、特に異種混合の企業設定において、目に見えないデータ環境や分析ワークフローを一般化するのに苦労しています。このため、特定のデータ環境の複雑な分析ワークフローをキャプチャする高品質のデータ エージェントの軌跡を合成する必要性が高まっています。このような軌跡は、2 つの重要な下流用途をサポートします。1 つはデータ エージェント モデルをターゲット ドメインに適応させる教師あり微調整 (SFT) データとして、もう 1 つは不慣れなデータ環境で汎用 LLM をガイドするためのインコンテキスト学習 (ICL) デモンストレーションとして機能することです。そこで、適応モデル選択とクロスタスクプレフィックス再利用を備えたモンテカルロツリー検索(MCTS)を介して、特定のデータ環境から高品質のデータエージェントの軌跡を合成するシステムであるTOFFEEを紹介します。 TOFFEE が異種環境にわたる複雑な分析タスクのスケーラブルな軌跡データを効果的に生成できることを示します。このデモでは、TOFFEE のタスク プール構築、トラジェクトリ エクスプローラー、学習コスト モデルなどのシステム フレームワークを紹介します。また、TOFFEE の Web インターフェイスとそのワークフローを紹介し、データ エージェント微調整のための軌跡合成と、デモンストレーションによって拡張されたデータ エージェント推論という 2 つのエンドツーエンド シナリオを示します。
原文 (English)
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise settings. This creates a growing need for synthesizing high-quality data agent trajectories that capture complex analytical workflows for given data environments. Such trajectories support two key downstream uses: they can serve as supervised finetuning (SFT) data that adapts data agent models to the target domain, and as in-context learning (ICL) demonstrations to guide general-purpose LLMs in unfamiliar data environments. Thus, we introduce TOFFEE, a system for synthesizing high-quality data agent trajectories from given data environments via Monte Carlo Tree Search (MCTS) with adaptive model selection and cross-task prefix reuse. We show that TOFFEE can effectively generate scalable trajectory data for complex analytical tasks across heterogeneous environments. In this demonstration, we present the system framework of TOFFEE, including its task pool construction, trajectory explorer, and learned cost model. We also introduce the web interface of TOFFEE and its workflow, and demonstrate two end-to-end scenarios: trajectory synthesis for data agent finetuning, and demonstration-augmented data agent reasoning.
AgentLens: コーディング エージェント評価のための実稼働環境で評価された軌跡レビュー
ここでは、対話型コード エージェントの実稼働環境で評価されたベンチマークである AgentLens を紹介します。ほとんどのコード エージェント ベンチマークでは、実行が 1 ビットに削減されます。タスクは成功しましたか? -- しかし、これらのエージェントを実際に使用する人々は、エージェントがどのように指示に従い、ツールを使用し、自身の作業を検証し、間違いから回復し、途中でエージェントに話しかけるかという軌跡全体を経験します。 AgentLens はその軌跡全体を評価します。客観的なチェックが存在する正式な検証と、LLM で作成された軌跡のレビューおよび並べての比較を組み合わせることで、各実行でスコアがなぜそのようになるのかについての読みやすい説明が得られます。これにより、AgentLens はモデルのランク付け以上の用途に役立ちます。モデルの動作を診断し、独自のエージェントの連続バージョンを比較し、夜間の評価パイプラインで製品の回帰を捕捉するために使用されます。 https://github.com/agent-lens/agent-lens-bench でベンチマークをオープンソースとしてリリースします。
原文 (English)
AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at https://github.com/agent-lens/agent-lens-bench.
Infinity-Parser2 テクニカルレポート
我々は、エンドツーエンドの文書解析のための制御可能なデータ合成パイプラインとマルチタスク強化学習を組み合わせた大規模なマルチモーダル モデルである Infinity-Parser2 を紹介し、忠実に注釈が付けられた解析コーパスの持続的な不足に対処します。私たちの貢献は 3 つあります。まず、制御可能なレンダリング フレームワークと反復改良ループを組み合わせたスケーラブルな合成エンジンを構築し、それを使用して Infinity-Doc2-5M を構築し、オープンソース化します。これは、さまざまな文書タイプにまたがる 500 万サンプルのバイリンガル (中国語/英語) コーパスであり、要素の境界ボックス、正規のコンテンツ フォーム (Markdown、HTML、LaTeX、SMILES、構造化チャート)、およびフルページで注釈が付けられています。読む順番。次に、検証可能なマルチタスク報酬システムを導入します。これにより、8 つの共同トレーニング目標 (文書解析、レイアウト分析、表解析、数式解析、チャート解析、化学式解析、文書 VQA、および一般的なマルチモーダル理解) にわたって共同強化学習を可能にし、単一の最適化信号で認識、構造、および推論を統合します。 3 番目に、共有アーキテクチャの下で 2 つのバリアントをリリースします。Infinity-Parser2-Flash は、Infinity-Parser-7B と比較して $3.68\times$ のスループット向上を実現し、低遅延推論用に最適化されています。もう 1 つは、精度が重要な設定向けに設計された Infinity-Parser2-Pro です。 Infinity-Parser2-Pro は、olmOCR-Bench で 87.6%、ParseBench で 74.3% に達し、DeepSeek-OCR-2、PaddleOCR-VL-1.5、MinerU2.5 を上回り、チャート、化学式、ドキュメント VQA に対する強力な汎用性を備えています。
原文 (English)
Infinity-Parser2 Technical Report
We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of faithfully annotated parsing corpora. Our contributions are threefold. First, we build a scalable synthesis engine, pairing a controllable rendering framework with an iterative refinement loop, and use it to construct and open-source Infinity-Doc2-5M: a 5-million-sample bilingual (Chinese/English) corpus spanning diverse document types, annotated with element bounding boxes, canonical content forms (Markdown, HTML, LaTeX, SMILES, structured charts), and full-page reading order. Second, we introduce a verifiable, multi-task reward system that enables Joint Reinforcement Learning across eight co-trained objectives (document parsing, layout analysis, table parsing, math formula parsing, chart parsing, chemical formula parsing, document VQA, and general multimodal understanding), unifying perception, structure, and reasoning in a single optimization signal. Third, we release two variants under a shared architecture: Infinity-Parser2-Flash, optimized for low-latency inference with a 3.68x throughput gain over Infinity-Parser-7B, and Infinity-Parser2-Pro, engineered for precision-critical settings. Infinity-Parser2-Pro reaches state-of-the-art 87.6% on olmOCR-Bench and 74.3% on ParseBench, surpassing DeepSeek-OCR-2, PaddleOCR-VL-1.5, and MinerU2.5, with strong generalization to charts, chemical formulas, and document VQA.
MAG: マルチモーダル アクションとガイド生成のための Web エージェント ベンチマークとハーネス
デジタル アダプション プラットフォーム (DAP) は、Web システムで広く使用されている埋め込みオーバーレイで、ページ内の操作をユーザーにガイドし、不慣れなインターフェイスをすぐに使い始めるのに役立ちます。ただし、実際のタスクを完了するということは、1 つのページ上でいくつかのボタンをクリックすることを意味することはほとんどありません。ページの状態が変化するたびに展開される一連のアクションが必要です。また、以前の研究では、自動化された Web エージェントのアクションとガイド テキストの生成を 2 つの別個の問題として扱っており、そのほとんどは人間が実際に操作するレンダリングされた画面ではなく、DOM やアクセシビリティ ツリーなどのテキスト ページ表現をモデルにフィードします。この作業では、タスクの実行とガイドの書き込みを 1 つのマルチモーダル アクションとガイド タスクに統合する最初のベンチマークである MAG を紹介します。このベンチマークには、スクリーンショット上の 2 つの基礎スキーム (セット オブ マーク要素の選択と生のピクセル座標) が含まれます。さらに、LLM 支援によるアノテーション、人間による検証、トレーニング、ライブ環境での評価、およびアクションとガイドの共同メトリクスをカバーする、この複合タスクのための完全なハーネスを構築します。このハーネスを使用して、フロンティア API モデルとオープン マルチモーダル モデルを評価し、詳細な分析をレポートします。最後に、専門家の軌跡を追加した GRPO トレーニング方法を設計します。これにより、監視された 9B エージェントの成功率がほぼ 2 倍 (6.9% から 13.2%) になり、同時にガイドの品質が向上します。最も強力なモデルでも完了するタスクは 40% 未満であり、将来の研究の余地は十分にあります。
原文 (English)
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly. Completing a real task, however, rarely means clicking a few buttons on a single page: it takes a sequence of actions that unfolds across changing page states. Prior studies have also treated automated web agent actions and guide text generation as two separate problems, and most of them feed models textual page representations such as the DOM or accessibility trees rather than the rendered screens that humans actually operate on. In this work we introduce MAG, the first benchmark that unifies task execution and guide writing into a single Multimodal Action and Guide task, with two grounding schemes over screenshots: Set-of-Mark element selection and raw pixel coordinates. We further build a complete harness for this compound task, covering annotation with LLM assistance and human verification, training, evaluation in live environments, and joint metrics for actions and guides. With this harness we evaluate frontier API models and open multimodal models, and report detailed analyses. Finally, we design a GRPO training method augmented with expert trajectories, which nearly doubles the success rate of a supervised 9B agent (from 6.9% to 13.2%) and improves guide quality at the same time. Even the strongest model completes fewer than 40% of the tasks, leaving ample room for future research.
IdeaTrail: 科学的アイデアのためのフルプロセス エージェントの軌跡
科学研究は、テキスト生成という単一の行為ではなく、複雑な多段階のワークフローです。通常、アイデアのプロセスは、文献検索、論文の読解、ツールの使用、クレームの確認、論文間の統合、ブレーンストーミング、弱い指示の拒否、および反復的な執筆を通じて現れます。既存のリソースはこのプロセスの個々のコンポーネントをキャプチャしますが、ツールの使用、証拠の取得、中間成果物の進化、アイデアまたは提案レベルのエンドポイントを共同で記録するデータセットは依然として限られています。このレポートでは、科学的アイデアと提案生成のためのマルチターン プロセス軌跡データセットである \method を紹介します。各インスタンスは、証拠の収集からアイデアの選択または提案の作成までの調査プロセスを記録します。 \method は軌道を自由に作成するのではなく、人間が選択した高品質の研究論文と提案成果物から開始し、ジェネレーターとアドバイザーの合成ループを使用します。ジェネレーターはアクション、観察、アーティファクトの編集を通じて目に見える軌道を生成しますが、アドバイザーは完全な生成コンテキストにアクセスして、グラウンディング、因果関係の順序、自然性、および隠れたターゲットからの漏れをチェックします。この逆から順の手順により、実際の科学成果との整合性を保ちながら、研究実践の不確実性、証拠の使用、段階的な収束を近似した複数ターンの研究データが生成されます。 \method は、科学研究エージェント向けにプロセス監視データを合成するためのデータセットと一般的なレシピの両方を提供します。
原文 (English)
IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation
Scientific ideation unfolds over multiple stages, including literature search, paper reading, tool use, claim checking, cross-paper synthesis, brainstorming, rejection of weak directions, and iterative writing. Yet most existing resources capture isolated components or final artifacts rather than the process connecting them. We introduce IdeaTrail, a dataset of 1,170 multi-turn trajectories for scientific ideation and proposal generation. Each trajectory follows a research process from evidence gathering to either idea selection or proposal construction, jointly recording tool use, acquired evidence, intermediate artifacts, and reasoning. IdeaTrail is synthesized from human-selected research papers and proposal artifacts through a Generator--Advisor loop. The Generator produces the visible sequence of actions, observations, and artifact edits, while the Advisor uses the full generation context to check grounding, causal order, naturalness, and leakage from hidden targets. This reverse-to-forward design keeps trajectories aligned with real scientific artifacts while retaining the uncertainty, evidence use, and staged convergence characteristic of research practice. IdeaTrail provides both reusable process supervision and a general recipe for constructing scientific-research-agent data.
エージェントはただ同意するだけではなく、覚えている: ステートフル個人エージェントにおける永続的な媚びのベンチマーク
ステートフル パーソナル エージェントは、長期的なユーザー プロファイル、エピソード記憶、再利用可能なスキルを維持することがますます増えています。この持続性により、会話のお調子者が状態記述の失敗に変わります。受け入れられたユーザー中心の主張は、永続的な設定、背景事実、またはワークフローとしてコミットされ、元の会話が終わった後に再利用される可能性があります。私たちはこれを永続的なお調子者と呼び、パーソナル エージェント お調子者ベンチマーク (PASB) を導入します。これは、会話による要求が受け入れられ、永続的なエージェント状態に書き込まれ、後の中立的なクエリで再利用されるかどうかを追跡する 1,600 タスクのベンチマークです。事前に書き込まれたメモリを提供する以前のベンチマークとは異なり、PASB は何を保存するかを決定する実際のエージェント (Hermes-Agent および OpenClaw) を評価します。 4 つのシナリオ フレームと 4 つの時間配信パターンを組み合わせ、5 ターンの永続ステージをクリアされた 3 ターンのクエリ ステージから分離することで書き込みプロセスを分離し、ダウンストリームの影響が永続状態からのみ発生するようにします。 12 のモデル全体で、コミット境界が重要な変曲点です。ダウンストリームの障害は、セッションのみのエピソードの 45.0% からコミット後は 71.9% に増加し、一貫して 27.0 パーセント増加しています。コミットされたクレームには、ステータスの昇格、帰属の削除、範囲の拡大という 3 つの書き込み時間パターンが見られます。これらのパターンは、記憶に似たフレーミングや手順的なフレーミング、繰り返しの強化、さらにはドメインの境界を越えた場合でもより強力になります。これらの結果は、エージェントのおべっかが根本的に国家の統治問題であることを示している。ユーザー コンテンツが永続メモリに保存されると、エージェントの発言だけでなく、エージェントが書き込む内容も安全性によって管理されなければなりません。 PASB は、応答レベルの軽減策を超えて保存されたコンテンツのソース、役割、および範囲を維持しながら、危険なコミットをゲートするために必要な書き込み時の制御を特定します。
原文 (English)
Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents
Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversational sycophancy into a state-writing failure: accepted user-centric claims can be committed as lasting preferences, background facts, or workflows and later reused after the original conversation is gone. We call this persistent sycophancy and introduce the Personal Agent Sycophancy Benchmark (PASB), a 1,600-task benchmark that traces whether a conversational claim is accepted, written into durable agent state, and reused in a later neutral query. Unlike prior benchmarks that provide pre-written memories, PASB evaluates real agents (Hermes-Agent and OpenClaw) that decide what to store. It isolates the write process by combining four scenario framings with four temporal delivery patterns and separating a five-turn persist stage from a cleared three-turn query stage, ensuring downstream effects arise only from durable state. Across twelve models, the commit boundary is the key inflection point: downstream failure increases from 45.0% in session-only episodes to 71.9% after commitment, a consistent increase of 27.0 percentage points. Committed claims exhibit three write-time patterns: status promotion, attribution removal, and scope broadening. These patterns become stronger under memory-like or procedural framing, repeated reinforcement, and even across domain boundaries. These results show that agent sycophancy is fundamentally a state-writing governance problem. Once user content is committed to durable memory, safety must govern what agents write, not only what they say. PASB identifies the write-time controls needed to gate risky commits while preserving the source, role, and scope of stored content beyond response-level mitigations.
QwenPaw-Data: 自律型エンタープライズ データ分析の橋渡しとなる事実、方法論、および実行
エンタープライズ データ分析は、自律エージェントの明確なフロンティアとして浮上しています。汎用のインタラクションやソフトウェア エンジニアリングと比較すると、オープンで曖昧な、継続的に進化する環境で動作します。これらの特性には、セマンティクス、方法論、実行、進化をシステムの第一級の関心事として扱うデータ エージェント アーキテクチャが必要です。この目的を達成するために、企業のインテリジェントなデータ分析のために設計されたエージェント データ システムである QwenPaw-Data を導入します。 QwenPaw-Data は、ウェアハウス、ダッシュボード、ドキュメント、インタラクション ログ、および履歴タスクからの異種資産を、再利用可能で管理可能で進化可能な分析資産に統合し、自然言語リクエストを、データの理解、取得、分析、レポート生成、意思決定支援に及ぶエンドツーエンドの分析ワークフローに変換します。そのアーキテクチャは、問題を 3 つの協調サブシステムに分解します。DataBridge は、相互接続されたメタデータ、ナレッジ、およびトレース グラフを通じて信頼できるセマンティック基盤を提供します。 Skill-Hub は、専門家の分析手法を再利用可能で検証可能なスキルに体系化します。そしてホストは、これらの証拠とメソッド資産を制御可能なアーティファクト中心のランタイム実行に具体化します。これらのサブシステム全体で、セマンティクス、メソッド、トレース、およびフィードバックが継続的にシステムに戻され、自己進化する資産フライホイールを形成します。公開ベンチマークと実際の産業用 BI ワークロードに関する実験では、QwenPaw-Data が検証可能なデータ アクセス機能とより高度な分析品質の両方を向上させ、信頼性があり、追跡可能で、継続的に改善されるエンタープライズ データ エージェントのための実用的な基盤を提供することが示されています。
原文 (English)
QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics
Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics call for a data-agent architecture that treats semantics, methodology, execution, and evolution as first-class system concerns. To this end, we introduce QwenPaw-Data, an agentic data system designed for enterprise intelligent data analysis. QwenPaw-Data consolidates heterogeneous assets from warehouses, dashboards, documents, interaction logs, and historical tasks into reusable, governable, and evolvable analysis assets, then turns natural-language requests into end-to-end analytical workflows spanning data understanding, retrieval, analysis, report generation, and decision support. Its architecture decomposes the problem into three collaborative subsystems: DataBridge provides trustworthy semantic grounding through interconnected metadata, knowledge, and trace graphs; Skill-Hub codifies expert analytical methodology into reusable and verifiable skills; and Host materializes these evidence and method assets into controllable, artifact-centric runtime execution. Across these subsystems, semantics, methods, traces, and feedback are continuously deposited back into the system, forming a self-evolving asset flywheel. Experiments on public benchmarks and real-world industrial BI workloads show that QwenPaw-Data improves both verifiable data access capability and higher-level analytical quality, offering a practical foundation for reliable, traceable, and continuously improving enterprise data agents.
流体知能研究に規則帰納法を戻す?ヒトにおけるARC-AGIベンチマークの初期検証
流体インテリジェンス (gf) 測定に関する 2 つの競合する視点は、パフォーマンスが主に作業記憶容量または新しい関係を誘発する能力のいずれかによって制約されることを提案しています。限られた繰り返しルールの使用から明らかなように、現在、最初の観点が測定において支配的ですが、2 番目の観点は多くの定義に反映されていますが、測定にはほとんど存在しません。 ARC-AGI ベンチマークは主にルール帰納を必要とし、人間と人工システムの両方の gf の尺度として提案されました。ただし、その心理測定特性は人間のサンプルではまだ調査されていません。そこで、我々は 100 人の参加者を対象とした最初の研究で、ARC-AGI の心理測定特性と規範論的ネットワークを調査しました。 ARC-AGI 項目の編集では良好な心理測定特性が示され、図形推論テストで測定された図形の流動性知能と実質的に相関していました (\r{ho} = 0.63)。図形の独創性との関連性は弱かった。これらの発見は、人間の流体知能の尺度としての ARC-AGI の妥当性に対する最初の裏付けを提供します。将来の研究には、追加の多変量共変量だけでなく、より多くのルール帰納タスクが含まれる必要があります。この研究は、当初は機械用に設計されたタスクを人間で研究するという珍しいものです。より体系的な評価と学際的な協力を可能にするために、AI ベンチマークを人間の認知能力の規範論的ネットワークに体系的に埋め込むことを提案します。
原文 (English)
Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans
Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in measurement, as evident from the use of a limited set of recurring rules, whereas the second perspective is reflected in many definitions but rarely present in measurement. The ARC-AGI benchmark predominantly requires rule induction and was proposed as a measure of gf for both humans and artificial systems. However, its psychometric properties have not yet been examined in human samples. We therefore investigated the psychometric characteristics and nomological network of ARC-AGI in a first study with 100 participants. A compilation of ARC-AGI items showed good psychometric properties and correlated substantially with figural fluid intelligence as measured by a figural reasoning test (\rho= .63). Associations with figural originality were weak. These findings provide initial support for the validity of ARC-AGI as a measure of human fluid intelligence. Future research should include more rule induction tasks as well as additional multivariate covariates. This study is unusual by studying a task in humans that was initially designed for machines. We suggest systematically embedding AI benchmarks into the nomological network of human cognitive abilities to enable more systematic evaluation and interdisciplinary cooperation.
Propheticus: Machine Learning Framework for the Development of Predictive Models for Reliable and Secure Software
The growing complexity of software calls for innovative solutions that support the deployment of reliable and secure software. Machine Lear…
Diversity-Enriched Option-Critic
Temporal abstraction allows reinforcement learning agents to represent knowledge and develop strategies over different temporal scales. The…
Seeing Through Uncertainty: Free-Energy-Inspired Real-Time Adaptation for Robust Visual Navigation
Navigation in the natural world is a feat of adaptive inference, where biological organisms maintain goal-directed behaviour despite noisy…
Enabling Energy-Efficient Simultaneous Multi-Task Reinforcement Learning through Spiking Neural Networks with Active Dendrites for Bio-inspired Generalist Agents
Reinforcement learning (RL) has demonstrated remarkable capabilities in training agents to solve complex tasks autonomously, such as mobile…
Modeling Story Expectations: A Generative Framework using LLMs
Consumers' engagement with stories is shaped by their expectations about what will happen next, yet modeling these forward-looking beliefs…
Toward Metaphor-Fluid Conversation Design for Voice User Interfaces
Metaphors play a critical role in shaping user experiences with Voice User Interfaces (VUIs), yet existing designs often rely on static, hu…
Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI
Background: Federated learning (FL) enables collaborative training of clinical AI models without centralizing patient data, but adoption is…
SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation
We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural langua…
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically re…
Real-Time Model Checking for Closed-Loop Robot Reactive Planning
Reactive obstacle avoidance methods often cause agents to become trapped in local minima, because they can often only reason one step ahead…
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models
Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains di…
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) ha…
Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task
We investigate how embedding dimension affects the emergence of an internal "world model" in a transformer trained with reinforcement learn…
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasing…
A Neurosymbolic Approach to Natural Language Formalization and Verification
Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits…
Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning
Algorithmic reasoning -- the ability to perform step-by-step logical inference -- is a synthetic benchmark for evaluating multi-step reason…
First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safe…
Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entro…
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
Relational Foundation Models (RFMs) facilitate data-driven decision-making by learning from complex multi-table databases. However, the div…
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter
Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks. To address the challenges of e…
Self-Regulated Reading with AI Support: An Eight-Week Study with Students
College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions sha…
Declarative by Design, Assistable Only by Convention: Benchmarking Multi-Agent Frameworks for AI-Assistability
Multi-agent frameworks (MAFs) promise to simplify LLM-driven software development, yet no principled metric captures how well AI coding ass…
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise
Spoken query retrieval is an important interaction mode in modern information retrieval. However, existing evaluation datasets are often li…
Egocentric Bias in Vision-Language Models
Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce Flip…
Xray-Visual Models: Scaling Vision models on Industry Scale Data
We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social…
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to i…
Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast
Demographic attributes can be predicted from medical images, raising concerns about bias in clinical AI systems. In X-ray imaging, acquisit…
TADPO: Reinforcement Learning Goes Off-road
Off-road autonomous driving poses significant challenges such as navigating unmapped, variable terrain with uncertain and diverse dynamics.…
Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts
Large language models are increasingly used in scholarly work, yet it remains unclear whether their productivity gains are accompanied by c…
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accurac…
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual struct…
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone…
Neuro-Symbolic ODE Discovery with Latent Grammar Flow
Understanding natural and engineered systems often relies on symbolic formulations, such as differential equations, which provide interpret…
タイムマシン: 効率的な知覚のための動きの力について
ビデオ表現学習は近年、目覚ましい進歩を遂げています。これは、トレーニングの規模や、言語と対照的にトレーニングされた視覚モデルの成功など、多くの要因によって推進されています。これらの要因は、ビデオ モデルができることの限界を押し広げていますが、同時に独自の制限も導入しています。まず、ビデオ モデルをスケーリングすると法外なコストに達する可能性があり、第 2 に、言語から学習すると、キャプション内の概念を学習できる範囲が制限されます。その結果、ビデオモデルは依然として時間的な理解に苦労しています。この論文では、ビデオ表現の中心的なモダリティとして動きを使用する新しいアプローチを提案します。特に、ポイント トラックの形式でビデオ内の動きが与えられると、マスクされたオートエンコーダを使用してトラックの一部をマスクし、失われたトラックを再構築するようにオートエンコーダをトレーニングします。これにより、自己教師ありの方法で表現を学習できるようになります。私たちは、モーションを使用してビデオを表現することで、ビデオ テクノロジーの核となる制限の両方に実際に対処できることを示します。まず、モーションは本質的に外観に依存しないため、適切に一般化するために必要なサンプルが少なくなるため、トレーニング データの規模を大幅に削減できます。第二に、動作により、言語に依存したトレーニング パラダイムを回避して、より詳細な概念を学習できるようになります。その結果は、TIME (Temporally Informed Motion Embedding) と呼ばれる埋め込みであり、合成モーション データのみでトレーニングされた表現です。この埋め込みを幅広いタスクでゼロショット方式でテストします。付加機能がなければ、パフォーマンスは最大 4 桁少ないトレーニング データを使用する最先端のモデルと同等であることがわかります。これは、より時間的な認識とよりスケーラブルなビデオ モデルの新しいパラダイムへの足がかりです。
原文 (English)
The TIME Machine: On The Power of Motion for Efficient Perception
Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have pushed the boundaries of what video models can do, they also introduce their own set of limitations: first, scaling video models can reach prohibitive costs and second, learning from language restricts the range of concepts that can be learned to those in captions. As a result, video models still struggle with temporal understanding. In this paper we propose a novel approach that uses motion as the central modality for video representation. In particular, given the motion in a video in the form of point-tracks, we use a masked-autoencoder to mask some of the tracks and train the autoencoder to reconstruct the missing tracks. This allows us to learn a representation in a self-supervised manner. We show that using motion to represent videos actually addresses both of the core limitations of video technology. First, it allows us to massively reduce the scale of training data, as motion is inherently appearance-independent and hence needs fewer examples to generalize well. Second, motion allows us to bypass the language-dependent training paradigm, learning better fine-grained concepts. The result is an embedding that we call TIME (Temporally Informed Motion Embedding), a representation trained exclusively on synthetic motion data. We test this embedding on a wide set of tasks in a zero-shot manner. We observe that without bells and whistles, performance is on par with state-of-the-art models using up to 4 orders of magnitude less training data. This is a stepping stone towards a new paradigm of video models that are both more temporally aware as well as more scalable.
注意力の散漫によって引き起こされる視覚的なぼやけを修正して幻覚を軽減する: アルゴリズムと理論
マルチモーダル大規模言語モデル (MLLM) は、物体の幻覚に悩まされることがよくありますが、この失敗の根底にある視覚知覚メカニズムはまだ十分に理解されていません。この研究では、幻覚が人間のような注意散漫現象と強く関連していることを明らかにしました。この現象では、分割焦点下にある人間は視覚の明瞭度が低下し、不正確な説明を生成しますが、モデルでは同じメカニズムが、複数頭の注意における空間的な不一致と、デコード中の画像トークンへの注意の一時的な薄れとして現れます。さらに、注意の分散によってモデルの複雑さが増大し、分類の一般化が低下するという理論的な洞察も提供します。これらの発見に動機づけられて、我々は、画像認識を改善するための注意集中アプローチ(AFIP)を提案します。これは、クロスヘッド注意の強化を通じて注意の散漫を修正し、動的な歴史的注意の強化を通じて視覚の基礎を強化します。複数のベンチマークとモデルに関する広範な実験により、追加のトレーニングなしで AFIP の有効性が検証されます。
原文 (English)
Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory
Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this failure remains poorly understood. In this work, we reveal that hallucinations are strongly associated with a human-like attention distraction phenomenon, where humans under divided focus experience degraded visual clarity and produce inaccurate descriptions, while in models the same mechanism manifests as spatial inconsistency in multi-head attention and temporal fading of attention to image tokens during decoding. We further provide theoretical insights that attention dispersion increases model complexity and degrades classification generalization. Motivated by these findings, we propose an Attention-Focused Approach for Improved Image Perception (AFIP), which corrects attention distraction via cross-head attention enrichment and reinforces visual grounding through dynamic historical attention enhancement. Extensive experiments on multiple benchmarks and models validate the effectiveness of AFIP without additional training. Code is available at: https://github.com/MIKUZ12/AFIP.
dMX: 低精度浮動小数点フォーマットの微分可能な混合精度代入
大規模言語モデル (LLM) を低精度の浮動小数点表現に量子化することは、効率的な展開の中心となりますが、単一のビット幅をすべてのレイヤーに均一に適用することは、パフォーマンスと精度の両方の点で最適とは言えません。この研究では、学習可能な浮動小数点ビット幅割り当てのための微分可能な混合精度量子化フレームワークである dMX を紹介します。私たちは、オープン コンピューティング プロジェクト (OCP) 標準によって定義されたデータ型のマイクロスケーリング浮動小数点 (MXFP) ファミリへの応用を研究します。レイヤごとのビット幅の割り当ては、各レイヤの浮動小数点形式がスカラー パラメータによってパラメータ化され、多変量設計空間を単一の学習可能なオフセットに折りたたむ連続最適化問題として定式化されます。トレーニング中、このオフセットは連続値をとり、離散量子化形式間の突然の振動を回避します。温度ベースのアニーリング スケジュールにより、学習されたオフセットが段階的に離散化され、トレーニング動作と推論動作の間で突然移行することなく、最終的な構成がハードウェア互換の MXFP 形式にマッピングされることが保証されます。ターゲットを意識した正則化用語は、平均ビット幅をユーザー指定の予算に向けて導き、推論コストの大まかな代理として機能し、モデルの品質と展開効率のバランスをとります。私たちは Llama、Qwen3、SmolLM2 などのさまざまな LLM ファミリで実験を実行し、WikiText-2 での複雑性と 4 つのゼロショット推論ベンチマークでの精度を評価しました。これらの設定全体にわたって、dMX は一貫してパレート支配モデルを生成し、カルバック ライブラー (KL) 発散ベースのレイヤー選択ヒューリスティックを改善し、モデルの品質と平均ビット幅の間のトレードオフを効率的にナビゲートします。
原文 (English)
dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats
Quantizing large language models (LLMs) to low-precision floating-point representations is central to efficient deployment, yet applying a single bit-width uniformly across all layers is sub-optimal in terms of both performance and accuracy. This work introduces dMX, a differentiable mixed-precision quantization framework for learnable floating-point bit-width assignment. We study its application for the microscaling floating-point (MXFP) family of data types defined by the Open Compute Project (OCP) standard. The per-layer bit-width assignment is formulated as a continuous optimization problem in which each layer's floating-point format format is parameterized by a scalar parameter, folding the multi-variate design space into a single learnable offset. During training this offset takes continuous values, avoiding sudden oscillations between discrete quantization formats. A temperature-based annealing schedule progressively discretizes the learned offsets, ensuring that the final configuration maps to hardware-compatible MXFP formats without abrupt transitions between training and inference behavior. A target-aware regularization term steers the average bit-width toward a user-specified budget, serving as a coarse-grained proxy for inference cost and balancing model quality against deployment efficiency. We performed experiments on different families of LLM, such as Llama, Qwen3, and SmolLM2, evaluating perplexity on WikiText-2 and accuracy on four zero-shot reasoning benchmarks. Across these settings, dMX consistently yields Pareto-dominating models and improves over Kullback-Leibler (KL) divergence-based layer-selection heuristics, efficiently navigating trade-offs between model quality and average bit-width.
The Score Hamiltonian: Mapping Diffusion Models to Adiabatic Transport
We exhibit an exact correspondence between sampling with score-based diffusion models and adiabatic transport of ground states for a family…
話題の感情がイデオロギーの認識を引き起こすのか?政治ニュース記事における人間による注釈と LLM の注釈の比較
私たちは、トピックの感情が政治的イデオロギーの認識に因果関係を持っているかどうか、そしてその答えは誰がイデオロギーのラベルを割り当てるかに依存するかどうかを尋ねます。 AllSides の記事と Llama-3.3-70b-versatile の共有センチメント アノテーションを組み合わせて、人間の専門アノテーター、GPT-4o-mini (ベースラインおよび微調整)、および Llama-3.3-70B からのイデオロギー ラベルを比較します。私たちは、Double Machine Learning (DML) とコミュニティレベルのメディエーション分析を 4 つのアノテーション パラダイムすべてに適用します。人間による注釈は、コミュニティレベルでは重大な因果関係をもたらしません。微調整された GPT-4o-mini は最高の分類精度 (F1=72.48) を達成し、コミュニティレベルでの顕著な治療効果と調停における重大な自然直接効果 (NDE) を生み出す唯一のアノテーター パラダイムです。私たちはこれを近道学習の証拠として解釈します。イデオロギーでラベル付けされたデータを微調整すると、モデルは偽の感情、つまりこのタスクでは人間の判断では機能しないイデオロギーの結合を内部に取り込んでしまいます。この結合は構造的に F1 ベースの評価には見えず、下流の因果関係分析におけるシルバー ラベルや人間の判断の代理として LLM アノテーションを使用することに影響を及ぼします。
原文 (English)
Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles
We ask whether topic sentiment has a causal effect on perceived political ideology, and whether the answer depends on who assigns the ideology label. Using articles from AllSides, paired with shared sentiment annotations from Llama-3.3-70b-versatile, we compare ideology labels from expert human annotators, GPT-4o-mini (baseline and finetuned), and Llama-3.3-70B. We apply Double Machine Learning (DML) and mediation analysis across all four annotation paradigms. Zero-shot LLMs regularly inflate effect sizes relative to human annotations, while fine-tuning often attenuates them back toward the human scale. Our results have implications for the use of LLM annotations as silver labels and as proxies for human judgment in downstream causal analyses: they may be reliable for recovering the presence and direction of effects on the partisan topics, but not their magnitude, leading to over- or under-prediction of some ideology given particular topics.
Explaining Data Mixing Scaling Laws
Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures. However, a theoretical u…
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is f…
ISE: マルチターン OS エージェントの軌跡のための実行ベースのレシピ
有能な OS エージェントをトレーニングするには、構造化されたユーザーの意図、複数ターンのタスク委任、および根拠のあるツールの実行を同時にキャプチャするデータが必要ですが、これらのプロパティは既存のデータセットには存在しません。我々は、これらのギャップに共同で対処する 3 段階の合成パラダイムである ISE (Intent -> Simulate -> Execute) を提案します。ステージ 1 では、4D フレームワーク (ペルソナ x ドメイン x タスク x 複雑さ) を介して約 50,000 の構造化インテントを構築します。重複排除後、プールには 43956 個の一意のインテントが含まれ、mpnet-base-v2 埋め込み (コサイン カーネル、q=1) のプール全体で 61.57 の Vendi スコアを達成しました。ステージ 2 では、ロールロックされたユーザー シミュレータを介してマルチターンのユーザー エージェント インタラクションを推進し、各ユーザー ターンを実際の実行結果に基づいて実行し、平均 8.12 ユーザー ターンと合計 68.24 のダイアログ ターンに相当する 23132 の完全な軌跡を生成します。ステージ 3 では、ライブの分離された OS ワークスペース内ですべてのツール呼び出しが実行され、シミュレートされた応答ではなく、本物の障害回復ダイナミクスが生成されます。 ISETrace の微調整により、標準プロトコルのエージェント ツール使用タスクで Qwen3-8B を使用し、ClawEval pass@1 が 19.3 から 37.7 に改善されました。この結果は、ゼロショット GPT-4o や 4 倍大きい Qwen3-32B ベース モデルよりも優れています。ステージ 2 のアブレーションは、マルチターン シミュレーションがパフォーマンス向上の大部分をもたらすことを証明しています。すべてのソース コードとデータセットは https://github.com/Valiere01/ISE-Trace でリリースされます。
原文 (English)
ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories
Training capable OS agents requires data that simultaneously captures structured user intents, multi-turn task delegation, and grounded tool execution--properties absent from existing datasets. We propose ISE (Intent -> Simulate -> Execute), a three-stage synthesis paradigm that addresses these gaps jointly. Stage 1 constructs roughly 50000 structured intents via a 4D framework (Persona x Domain x Task x Complexity); after deduplication the pool contains 43956 unique intents and attains a Vendi Score of 61.57 over the entire pool on mpnet-base-v2 embeddings (cosine kernel, q=1). Stage 2 drives multi-turn user-agent interaction through a role-locked user simulator that grounds each user turn in actual execution outcomes, producing 23132 complete trajectories averaging 8.12 user turns and 68.24 total dialogue turns. Stage 3 runs every tool call inside a live, isolated OS workspace, generating authentic failure-recovery dynamics instead of simulated responses. Fine-tuning on ISETrace improves ClawEval pass@1 from 19.3 to 37.7 using Qwen3-8B on agent tool-use tasks with a standard protocol. This result outperforms zero-shot GPT-4o and the larger Qwen3-32B base model which is four times bigger. An ablation on Stage 2 proves multi-turn simulation brings a large portion of the performance gain. We release all source code and dataset at https://github.com/Valiere01/ISE-Trace.
Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy
Hematoxylin and eosin (H&E) staining is the cornerstone of histopathology, yet scalable, quantitative analysis of H&E whole-slide images (W…
Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents
Long-horizon tool-use reinforcement learning learns from outcome verification, but trajectory-level advantages are broadcast over reasoning…
Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens
Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and qu…
LLM 主導の段階的改良による解釈可能かつ検証可能なハードウェア生成
大規模言語モデル (LLM) は、ソフトウェア開発において目覚ましい成功を収めています。ただし、幻覚の影響を受けやすいため、微妙な意味的および論理的エラーが発生する可能性があります。チップの設計と製造には大きなリスクが伴うため、ハードウェア エンジニアは依然としてレジスタ転送レベル (RTL) の生成に LLM に依存することに消極的です。この論文では、LLM の創造性と幅広い知識を、形式的手法の説明可能性と数学的厳密性と組み合わせたハードウェア生成フレームワークを提案します。具体的には、さまざまな設計上の決定とハードウェア機能をカバーする一連の変換ルールを考案します。これらのルールを繰り返し適用することで、LLM エージェントは、正確性が保証された設計仕様を RTL プログラムに変換できます。実験結果は、フレームワークの有効性と効率性を示しています。
原文 (English)
Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement
Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, meaning that they can introduce subtle semantic and logical errors. Due to the high stakes in chip design and manufacturing, hardware engineers are still reluctant to rely on LLMs for register-transfer level (RTL) generation. In this paper, we propose a hardware generation framework that combines the creativity and broad knowledge of LLMs with the explainability and mathematical rigor of formal methods. Specifically, we devise a set of transformation rules that cover various design decisions and hardware features. By iteratively applying these rules, an LLM agent can convert a design specification into an RTL program with guaranteed correctness. Experimental results demonstrate the effectiveness and efficiency of the framework.
整流フローによる命令ガイド付きオーディオ編集のためのハイブリッド拡散トランス
オーディオ編集の目的は、残りの音響コンテンツを維持しながら、自然言語命令に従って既存のオーディオ クリップ内の特定のコンテンツを変更することです。拡散モデルの目覚ましい進歩にも関わらず、既存のトレーニングベースの編集手法は主に、畳み込み U-Net バックボーンにおける局所的な帰納的バイアスとクロスアテンション相互作用に依存しており、長距離の意味論的整合や命令の正確な理解と位置特定を妨げることがよくあります。対照的に、拡散トランスフォーマーは、より強力なグローバル モデリングとマルチモーダル フュージョンを提供しますが、既存の編集アーキテクチャは通常、MMDiT ブロックと DiT ブロックの単純なスタックを採用しています。すべてのブロック内の連結されたオーディオ トークンとテキスト トークンに共同注意を適用すると、トークンの長さに関して 2 次の複雑さが生じます。編集パフォーマンスと効率のバランスをとるために、整流されたフローマッチングに基づいた命令ガイド付きオーディオ編集用のハイブリッド 2 ステージ拡散トランス アーキテクチャを提案します。音声トークンとテキスト トークンに対して共同アテンションを実行して、低解像度段階で大まかなセマンティック アライメントを確立し、その後、交互の共同アテンション ブロックとクロス アテンション ブロックに切り替えて、高解像度段階で編集の詳細を調整します。この粗いものから細かいものまでの戦略により、効率的かつ正確な指示に基づくオーディオ編集が可能になります。実験の結果、提案されたフレームワークは、コンパクトなモデルで編集効率を大幅に向上させながら、重複するオーディオ イベントや複雑な命令を含む困難な編集タスクで顕著なパフォーマンスの向上を達成することが示されています。
原文 (English)
RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers
Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models, existing training-based editing methods mainly rely on the local inductive biases and cross-attention interaction in convolutional U-Net backbones, which often hinder long-range semantic alignment and precise understanding and localization of instructions. In contrast, diffusion transformers provide stronger global modeling and multimodal fusion, but existing editing architectures usually adopt a simple stack of diffusion transformer blocks. Applying joint attention over concatenated audio and text tokens in all blocks results in quadratic complexity with respect to token length. To balance editing performance and efficiency, we propose a novel instruction-guided audio editing framework based on rectified flow matching (RFM), named RFM-Editing 2, built on a hybrid two-stage diffusion transformer. The proposed model performs joint attention over audio and text tokens to establish coarse semantic alignment at the low-resolution stage, then switches to alternating joint-attention and cross-attention blocks to refine editing details at the high-resolution stage. This coarse-to-fine strategy enables efficient and accurate instruction-guided audio editing. Experiments show that the proposed framework achieves notable performance gains on challenging editing tasks involving overlapping audio events and complex instructions, while substantially improving editing efficiency.
Prime Fourier Embeddings: A Principled Basis for Modular Arithmetic
Numbers have algebraic structure that standard neural embeddings often fail to expose. We introduce Prime Fourier Embeddings (PFE), which e…
Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking
The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion es…
鳩穴: 不適切なプロンプトはモデルを傷つけ、モデルが崩れたり、間違いを犯したりする
一般に、コンテキスト内学習は大規模言語モデル (LLM) で効果的であることが示されていますが、不適切なコンテキストはパフォーマンスの低下やモードの崩壊を引き起こす可能性があり、これを「ピジョンホール」と呼んでいます。 **意図せず悪い** コンテキストは、悪意のあるジェイルブレイクの意図がなくても発生する可能性があります。たとえば、ユーザーがモデルに間違った数学定理を正当化するように要求したり、モデルのバグのあるコードの修正に失敗したりします。具体的には、(1) ユーザーが解決策を提案したとき、および (2) 会話のコンテキストにアシスタントの以前の (不正確な) 応答が含まれているときの 2 つのシナリオで「ピジョンホール」を調査します。10 個の異なるモデルを使用した 10 個の検証可能なオープンエンドのタスクにわたる実験では、ピジョンホールがいくつかの方法で現れることがわかりました: (1) コンテキストから不正解を繰り返す (38 ~ 40% のパフォーマンス低下につながる)、(2) 狭いセットに収束する(3) ユーザーまたはアシスタントの以前の主張に合わせて、議論の的となっているトピックに対するスタンスを反転する グループ分けは、会話のターン数に応じてほぼ単調に悪化することがわかりました (繰り返される間違いが 1 から 5 に増加するにつれて、パフォーマンスはさらに 14% 以上低下します)。また、提供された例が正しい場合でも、グループ分けに起因するモード崩壊が発生する可能性があります。緩和へのステップとして、モデルを改善する合成エラーを含む RLVR を提案します。バニラ RLVR ベースラインと比較して、不正なコンテキストでは 43 ~ 60%。
原文 (English)
Pigeonholing: how bad prompts hurt models, causing collapse and mistakes
While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing." **Unintentionally bad** contexts can happen without malicious jailbreaking intents: For example, a user asks the model to justify an incorrect math theorem or fails to correct the model's buggy code. Specifically, we investigate ``pigeonholing" in two scenarios: (1) when the user suggests a solution, and (2) when the conversation context includes the assistant's previous (incorrect) responses. Our experiments across 10 verifiable and open-ended tasks with 10 different models show that pigeonholing manifests in several ways: (1) repeating the incorrect answers from context (leading to 38-40% performance drop), (2) converging on a narrow set of answers in coding and text generation without exploring alternatives, and (3) flipping stance on controversial topics to align with the user or the assistant's previous claims. We find that pigeonholing worsens almost monotonically with the number of conversation turns (performance drops by additional 14+% as repeated mistakes increase from 1 to 5), and pigeonholing-induced mode collapse can happen even when the provided example is correct. As a step toward mitigation, we propose RLVR with synthetic errors which improves models by 43-60% under bad contexts compared to vanilla RLVR baselines.
ハイブリッド プライバシーを意識したセマンティック検索: SVD で切り詰められたドキュメント ジオメトリと、制限された脅威モデルに基づく CKKS 暗号化クエリの再ランキング
高密度埋め込みはセマンティック検索と検索拡張生成を強化しますが、埋め込み反転攻撃はベクトルからソース テキストを再構築する可能性があります。ベクトル データベースが漏洩すると、その背後にある文書も漏洩します。教科書的な防御策は極端です。検索全体を準同型的に暗号化するのは健全ですが、100 万文書規模では遅すぎます。その一方で、保護するずっと前にプライバシー ノイズによってランキングが低下します。静的コレクションと動的クエリの間の非対称性を利用した中間パスを研究します。コレクションは幾何学的に保護されています。各ベクトルは低次元の SVD 部分空間上で切り詰められ、所有者のみが知っている秘密の直交変換によって回転されます。クエリは暗号的に保護されています。クエリは CKKS 準同型暗号化の下で再ランク付けされるため、正直だが好奇心旺盛なサーバーはクエリやスコアを見ることはありません。 CKKS パラメータは、小規模なオフライン ベンチマークから取得されます。私たちは、保護された部分空間に限定された攻撃者の再構成エラーの厳しい下限を証明します。 100 万のドキュメントと 5 つのエンコーダでは、このスキームは 1 秒未満のレイテンシでランキングの品質を維持し (線形デノイザーとして強力なエンコーダでわずかに向上します)、保護されたスペースに対する既製の反転攻撃はノイズ フロアまで崩壊します。次に、より強力な敵対者をテストします。既知の平文攻撃者は、保持された次元とほぼ同じ数の漏洩ペアから直交プロクラステスによる回転を回復します。公開されている積量子化コードは、最近傍構造を保存します。ランダム投影、校正済みノイズ、および BEIR ベースラインは、切り捨てが無料のデノイザーではなく、エンコーダーに依存する精度コストであることを示しています。私たちは限界を述べています。クエリの機密性は暗号化されていますが、ドキュメントの保護は経験的な難読化レイヤー (SVD の切り捨てと秘密のローテーション) であり、暗号化のプリミティブではありません。また、各主張の脅威モデルを区切ります。
原文 (English)
Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model
Semantic search creates an asymmetric disclosure problem: query embeddings may reveal user intent, while returning exact provider vectors distributes reusable representations. We evaluate a deliberately restricted hybrid design. A public corpus-fitted SVD basis and PQ index support client-side candidate selection; candidate IDs are disclosed, while the projected query is encrypted with CKKS. A block-SIMD kernel scores 100 plaintext provider vectors and returns one ciphertext. At 672 dimensions, median server time falls from 1689.1 to 224.5 ms and response size by a factor of 99.5; provider-only saturation reaches 25.99 requests/s with 16 workers. In a frozen post-exploratory revision-analysis subset, disjoint from validation and containing 3,235 canonical BEIR queries, five of six collections satisfy a +/-0.002 nDCG@10 equivalence rule between actual CKKS and plaintext reranking of the same shortlist, while ArguAna is inconclusive. Projection controls favor SVD over random and coordinate truncation. Leakage audits show that disclosed candidate sets are highly linkable and reproduce part of the exact neighbourhood, while public PQ reveals approximate corpus geometry. The design therefore conditionally hides numerical query slots, but does not provide semantic-query, document, unlinkability, circuit, or access-pattern privacy.
Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders
Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their p…
SHARD: アライメント耐性のあるプライベート密検索のためのセルキー残差分割
高密度の埋め込みはセマンティック検索と RAG を支えていますが、漏洩したベクトル ストアにより、基礎となるテキストの多くがそれを保持している人の手に渡されます。これを可能にする攻撃 (少数ショットのアライメント、ゼロショットの反転、教師なしクロススペース変換) には 1 つの弱点があります。それは、保護されたストアが、既知のジオメトリにアライメントできる単一のグローバル ジオメトリであるということです。通常の軽量防御である秘密のグローバル回転も例外ではありません。攻撃者が既知のペアでほぼ亜空間次元を取得すると、直交プロクラステスはそれを回復します。この弱い軸を取り除く検索保存埋め込み変換である Shard を紹介します。中央に配置された埋め込みは、短いパブリック プレフィックス (ステージ 1 取得用) と、個別の秘密キーの下で C セルにシャーディングされたプライベート残差に分割されます。残差は CKKS の下で再ランク付けされ、キーはキャンセルされて内積が正確のままになります。単一のパラメーター C は、置き換えられるグローバル線形ベースライン (C=1) からドキュメントごとのマイクロキー (C=N) までデザインを実行します。リランクはフル次元であるため、Shard は、半 SVD 切り捨てが放棄された生の空間 nDCG@10 を返します。また、残差はセルローカルにキー付けされるため、拡散既知平文リークのもとで残差を共通フレームにマッピングし直すには、暗号化されたクエリ数が少ない場合、およそ C 倍のアンカー (C=256 で中央値 200 ~ 102,400) のコストがかかります。短いパブリック プレフィックスにより、近隣構造の漏洩がはるかに少なくなり、マイクロキー制限により、リンク不可能で更新可能なテンプレートで残差グラフがゼロになります。この障壁は、学習済み、非線形、および教師なしのアライナに対して保持されており、整合ユーティリティ ノイズ防御がほぼすべてのプローブを匿名化解除するのに対して、シャードは匿名化を解除しません。私たちはその制限について明確にしています。セル内ではキーがキャンセルされ、標的型攻撃者が必要とするのは d_priv アンカー程度だけであり、重複する参照コーパスは依然としてプレフィックスを介して漏洩します。シャードは攻撃を認識する幾何学的防御であり、暗号化を保証するものではありません。
原文 (English)
SHARD: cell-keyed residual splitting for alignment-resistant private dense retrieval
Dense retrieval systems expose document geometry when vector stores are compromised, and a global protective transform can often be aligned from known pairs. We study SHARD, which splits PCA coordinates into a short routing prefix and a residual protected by independent cell-local orthogonal keys. It supports CKKS ciphertext--plaintext reranking but is evaluated as a leakage trade-off, not a cryptographic document-privacy guarantee. Corrected scoring uses centered document coordinates and an uncentered scoring query, preserving raw ranking up to a query-dependent constant. Across ten BEIR/MIRACL configurations it reproduces raw nDCG@10 and recall, whereas centering both sides loses up to 0.080 nDCG. Cell keys spread diffuse known-pair evidence across compartments, but minimum-norm alignment recovers useful signal far below full key rank, so there is no hard de-anonymization threshold. Real CKKS has maximum score error 2.29e-6 and no top-1 flips; block packing cuts query upload by 74--87% but raises in-process p50 latency by 14--26%. In a strengthened GTR case, an unknown key lowers token-F1 from 0.665 to 0.242; a wide prefix and eight pairs restore much. Under 25--90% release overlap, the unchanged prefix and clean residual norm link persistent rows with R@1 at least 0.9996, although cell-Gram linkage degrades under churn. A formally calibrated Gaussian release gives nDCG@10 at most 0.011 at epsilon=1; its only three strict utility matches occur at epsilon=32768 with linkage R@1 at least 0.995. SHARD preserves retrieval and compartmentalizes alignment evidence, but does not provide DP, unlinkability, or cancellable templates.
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing sk…
FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images
We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from on…
SkillSelect-Serve: 小規模 LLM エージェント向けの予算管理可能で QoS を意識したスキル サービスの推奨と構成
再利用可能なスキル ライブラリは、大規模言語モデル (LLM) エージェントにとって重要なインフラストラクチャになりつつありますが、既存の選択方法では、スキルを取得可能なドキュメントとして扱い、固定の上位 K リストを返すことがよくあります。この文書では、エージェントのスキル選択をスキル サービスの推奨および構成として定式化する、予算管理可能で QoS を意識したフレームワークである SkillSelect-Serve について説明します。 SkillSelect-Serve は、機能の説明、依存関係、コンテキスト コスト、リスク、QoS 関連の属性を備えた構造化されたスキル サービスとして生のスキルを表します。ローカルの Micro-Agent Requirement Planner が自然言語タスクを構造化されたサービス要件に変換し、共有ディスカバリー バックボーンが大規模なレジストリから候補サービスを取得します。次に、このフレームワークは、スキルレベルの限界適合性推定と、カバレッジ、冗長性、コスト、およびリスクのトレードオフに関するバンドルレベルの調整を行う二重粒度ユーティリティモデリングを実行します。 35,353 のスキルと 586 のタスク クエリに関する実験では、SkillSelect-Serve が固定の上位 K 取得ベースラインと比較して、同一予算バンドルの再現率と平均ユーティリティを一貫して向上させることが示されています。
原文 (English)
SkillSelect-Serve: QoS-Aware Budgeted Skill Service Recommendation for LLM Agents
Reusable agent skills are emerging as a service-oriented capability layer for Large Language Model (LLM) agents. Unlike plain retrieval items, a skill exposes functional capabilities, input-output assumptions, tool dependencies, context cost, and risk metadata. Selecting skills is particularly challenging for small LLM agents, which can load only a few capability units under restricted context, tool availability, and risk tolerance. Existing fixed Top-k methods rank skills by textual relevance and overlook requirement satisfaction, deliverability, and operational constraints. We present SkillSelect-Serve, a QoS-aware, budget-constrained Skill Service recommendation framework. Raw skills are profiled as structured Skill Services, the task is converted into a structured requirement object, and candidates discovered from a large-scale registry are ranked by a calibrated task-conditioned suitability estimator and packed by a constrained projection enforcing token-budget, aggregated-risk, and tool-availability constraints, using only deployment-observable features. On a registry of 35,353 skills with pooled multi-positive relevance judgments verified by two independent assessors, the unconstrained top-5 recommendation fits a realistic 4,000-token context for only 9.1% of tasks; the constrained projection restores 100% deliverability at a cost of only 1.14 points of hit rate, outperforming retrieve-and-rerank, budget truncation, and diversity-based selection under identical budgets. The same mechanism halves delivered risk exposure and eliminates the 44-81% tool-violation rates of tool-agnostic recommendation. At an identical three-service budget, hit rate improves from 0.8864 to 0.9091 over fixed Top-3 retrieval. The results support managing reusable agent skills as discoverable, comparable, and constraint-aware service units instead of plain retrievable documents.
Cross-Receiver Open-Set Radio Frequency Fingerprinting via Structure-First Adaptation
Radio frequency fingerprint identification (RFFI) provides a critical physical-layer security mechanism for dynamic Internet of Things (IoT…
PLGSA-Transformer: Periocular Landmark-Guided Attention with Occlusion-Adaptive Cosine Thresholding for Cross-Modal Masked and Unmasked Face Recognition
The widespread adoption of facial masks, accelerated by COVID-19 and mandated in security-sensitive settings, has exposed limitations of co…
Git-Assistant: Git リポジトリ更新のための計画ベースのサポート
バージョン管理システムは共同ソフトウェア開発に不可欠ですが、git のようなツールは多くの実務者にとって依然として困難です。大規模言語モデル (LLM) の最近の進歩により、開発者の意図を解釈するための有望な機能が提供されていますが、リポジトリ管理タスクにおける LLM の有効性は、形式的な推論の必要性によって制限されています。この作業では、LLM と自動計画を組み合わせて、開発者による重要な Git 操作の実行をサポートする AI ベースのアシスタントである Git-Assistant を紹介します。アシスタントはリポジトリのコンテキストを分析し、自然言語リクエストを実行可能なコマンド シーケンスに変換し、正確さと安全性を確保するための計画テクニックを組み込みます。合成およびランダム化された Git 環境を使用した体系的な評価方法論を提示し、LLM のみのバリアントと計画拡張バリアントのパフォーマンスを複数のメトリクスにわたって比較します。実験結果は、形式的推論を LLM と統合することで信頼性が向上し、リポジトリ管理におけるエラーが減少することを示しており、インテリジェントな開発者支援のためのハイブリッド AI アプローチの可能性を強調しています。
原文 (English)
Git-Assistant: Planning-Based Support for Updating Git Repositories
Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners. Recent advances in Large Language Models (LLMs) offer promising capabilities for interpreting developer intent, but their effectiveness in repository management tasks is limited by the need for formal reasoning. This work introduces Git-Assistant, an AI-based assistant that combines LLMs with automated planning to support developers in executing non-trivial git operations. The assistant analyzes repository context, translates natural language requests into actionable command sequences, and incorporates planning techniques to ensure correctness and safety. We present a systematic evaluation methodology using synthetic and randomized git environments, comparing the performance of LLM-only and planning-augmented variants across multiple metrics. Experimental results demonstrate that integrating formal reasoning with LLMs improves reliability and reduces errors in repository management, highlighting the potential of hybrid AI approaches for intelligent developer assistance.
ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-st…
Physics-Informed Structure Anchoring With Capture-Aware Prototype Calibration for Cross-Environment RF Fingerprinting
Radio frequency fingerprint identification (RFFI) exploits transmitter-specific hardware imperfections as physicallayer identity cues for I…
Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization
Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but…
Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory…
Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
Adaptive-compute world models -- early-exit or mixture-of-depths predictors that spend variable depth per step -- assume depth buys better…
ABot-N1: Toward a General Visual Language Navigation Foundation Model
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse…
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP
Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its descripti…
RepTran: Search-Based Repair of Transformer Models
To ensure the overall quality of AI-enabled software, not only traditional software components but also AI components need to be tested and…
An Empirical Study for Android-to-OpenHarmony GUI Test Migration
To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existin…
PRISM Edit: One Vector for All Temporal Answers
Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing lo…
Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous…