AIニュース 2026-07-08
自動生成: 2026-07-08 12:21 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
「Claude Fable 5」サブスク、突如5日間延長 ユーザー悲喜こもごも「寝ずに頑張ったのに」「制限リセットして」ITmedia AI+
歓迎の声もある一方で、「睡眠と健康を犠牲にして使い切った直後に延長を知った」「最後の最後で発表するのはやめてくれ」といった恨み節も。
-
Anthropicの「Claude Cowork」がWebとモバイルでも利用可能に まずはMaxプランからITmedia AI+
Anthropicは、AIエージェント機能「Claude Cowork」をWebブラウザとモバイルアプリ向けに展開すると発表した。これまで…
-
Meta、初のエージェント型画像生成AI「Muse Image」公開 動画生成AI「Muse Video」もプレビューITmedia AI+
Metaは、Meta Superintelligence Labs(MSL)初のメディア生成モデルとなる画像生成AI「Muse Image…
-
AIが破壊するIT業界の“人月商売” 「SIerの死」後に“生き残る者”の正体ITmedia AI+
米Anthropicの機能発表で約8300億ドルの時価総額が消し飛んだ「AIショック」。画面の使いやすさで稼いできたSaaSや、時間と人数…
-
ガバメントAI「源内」は自治体で本当に使えるのか? クラウド依存による“3つの落とし穴”ITmedia AI+
デジタル庁が公開したガバメントAI「源内」(GENAI)は、政府によるAIプラットフォームの公開という点で画期的な取り組みだ。一方で、自治…
-
Why the rise of open source AI isn’t hurting Anthropic … yetTechCrunch AI
Open source models’ success isn’t coming at the expense of frontier l…
-
Anthropicが教える「Fable 5活用術」 まず確認すべきは「自分が何を知らないか」ITmedia AI+
米Anthropicは、AIモデル「Claude Fable 5」の活用ガイドを公式ブログで公開した。AIコーディングツール「Claude…
トピック別件数
- LLM/生成AI 287件
- 研究/論文 284件
- エージェント 178件
- 画像/動画生成 175件
- ビジネス/資金調達 51件
- ロボティクス 46件
- ハードウェア/半導体 19件
- その他 8件
- 規制/政策 6件
日本語メディア16件
ITmedia AI+ (日本語)
note上方修正 「AI活用、想定以上」で人件費率低下 2Q累計営業益は「20倍超」に
AI活用で生産性が向上。人員数は緩やかに減っており、効率的な事業運営ができているという。
「Claude Fable 5」サブスク、突如5日間延長 ユーザー悲喜こもごも「寝ずに頑張ったのに」「制限リセットして」
歓迎の声もある一方で、「睡眠と健康を犠牲にして使い切った直後に延長を知った」「最後の最後で発表するのはやめてくれ」といった恨み節も。
Meta、初のエージェント型画像生成AI「Muse Image」公開 動画生成AI「Muse Video」もプレビュー
Metaは、Meta Superintelligence Labs(MSL)初のメディア生成モデルとなる画像生成AI「Muse Image」を公開した。検索やコーディングなどのツールを自律的に呼び出し、生成した画像を自己修正するエージェント型だ。「Meta AI」アプリやIns…
NTTドコモビジネスがIOWN活用の分散GPU環境を提供、25GBを2秒で転送
NTTドコモビジネスは、次世代ネットワーク「IOWN APN」を活用し、全国8拠点に分散したGPUを統合利用できる実証環境の提供を開始した。電力などの制限を解消し、オンデマンドなリソース確保やデータ主権に対応した分散AI基盤の実用性を検証できる。
IT予算の9割が人件費に消える――日本オラクル社長が切り込む「企業最大の課題」
日本オラクルの三澤智光社長が、日本企業のIT課題に切り込んだ。同社長が指摘する「IT投資の構造的問題」「オンプレミスシステムが抱える課題」とは何か。
逸材も不採用に……トヨタ系企業がAIで「人材採用のブレ」解消、800件の応募を効率化
年間約800件に及ぶ応募書類の確認に加え、採用基準の共有や面接記録の作成など、人事担当者の業務負担が課題となっていた。トヨタテクニカルディベロップメントは、これらの課題に対し、AIを活用した採用業務の見直しに取り組んだ。
ガバメントAI「源内」は自治体で本当に使えるのか? クラウド依存による“3つの落とし穴”
デジタル庁が公開したガバメントAI「源内」(GENAI)は、政府によるAIプラットフォームの公開という点で画期的な取り組みだ。一方で、自治体での実運用を考える上で無視できない論点も見えてくると、CIO補佐官として自治体DXに携わる筆者が解説する。
エンジニアの採用枠をトークン予算に回せるか? HRBrain常務が語る、「財布の壁」に挑む人事の未来
HRBrainの常務執行役員CSaO小山径氏は「AIも、これからは経営資源の一つになっていく」と話す。小山氏が見据えるAI活用の展望とは?
AIがExcel作業を丸ごと自動化? 企業の定型業務を効率化へ
生成AIによる業務自動化が広がる一方、取得したExcelやCSVファイルの集計・加工には依然として人手を伴う。データ収集からファイル処理までをAIが一貫して担うことで、企業に残る定型作業の効率化が進むだろう。
AIが破壊するIT業界の“人月商売” 「SIerの死」後に“生き残る者”の正体
米Anthropicの機能発表で約8300億ドルの時価総額が消し飛んだ「AIショック」。画面の使いやすさで稼いできたSaaSや、時間と人数を積み上げるSIerの「人月商売」が崩壊の危機に直面している。だが、業界ルールにのっとった複雑な計算(ビジネスロジック)を握る企業は依然とし…
Anthropicの「Claude Cowork」がWebとモバイルでも利用可能に まずはMaxプランから
Anthropicは、AIエージェント機能「Claude Cowork」をWebブラウザとモバイルアプリ向けに展開すると発表した。これまでデスクトップアプリ限定だったが、スマートフォンなどからも作業の進捗確認や指示出しが可能になる。今後は順次提供を拡大し、まずは上位のMaxプラ…
ソフトバンクの「1人100エージェント」を支える独自AIゲートウェイ「Cloud Proxy」の正体
生成AIやAIエージェントを全社展開する際、企業はセキュリティやガバナンス、性能といった課題に直面しがちです。ソフトバンクは「全社で1人100エージェント」構想の実現に向けて、AI利用の入り口となる共通基盤「Cloud Proxy」を内製しました。その設計思想や性能強化の取り組…
三菱UFJ半沢淳一社長「企業の資金ニーズに対応」 個人向け「エムット」でAI活用も
三菱UFJフィナンシャル・グループ(MUFG)の半沢淳一社長が産経新聞のインタビューに応じ、日銀の利上げで「金利のある世界」が本格化する中、企業の資金需要を取り込むため、事業戦略の提案力を高める考えを示した。
「AIが引用するドメイン」不動の首位は……
2位には前回5位のnote.comが浮上した。前回2位のWikipedia(ja.wikipedia.org)は3位に後退した。
【Fable 5に聞いてみた】サブスク終了後の効果的な使い方は? トークン節約法は? 本人インタビュー
注目を集めている米AnthropicのAIモデル「Claude Fable 5」の性能はどれほどなのか。ITmedia AI+編集部で試した結果を紹介する。今回は「従量課金制移行後にFable 5を効果的に使う方法」について聞いてみた。
Anthropicが教える「Fable 5活用術」 まず確認すべきは「自分が何を知らないか」
米Anthropicは、AIモデル「Claude Fable 5」の活用ガイドを公式ブログで公開した。AIコーディングツール「Claude Code」で同モデルを使う際の具体的なテクニックを紹介している。
海外メディア5件
TechCrunch AI (英語)
Why the rise of open source AI isn’t hurting Anthropic … yet
Open source models’ success isn’t coming at the expense of frontier labs. Instead, they each seem to capture two phases of the same life cy…
Microsoft joins AI cost-cutting trend by relying more on its own models
Microsoft is the latest Silicon Valley giant to cut back on its AI spending.
Discord admits AI moderation bug wrongfully banned users over harmless images
The company confirmed that the issue had been affecting accounts since May, with an additional 200 users banned over the weekend before its…
Savi’s app aims to protect consumers from realistic AI scams like kidnappers demanding ransom
The company just raised $7 million in seed funding, and is launching its app for iPhone and Android on Tuesday.
The first American autonomous ground vehicles are fighting in Ukraine
Forterra has deployed more than 100 of its self-driving ATVs in conflict zones in Ukraine.
公式ブログ0件
このカテゴリの新着記事はありませんでした。
論文783件
arXiv cs.AI (英語)
iFLYTEK-Embodied-Omni テクニカル レポート
汎用の身体化されたエージェントは、マルチモーダルな命令を理解し、環境がどのように進化するかを予測し、長期にわたる正確な制御アクションを生成する必要があります。既存のアプローチは通常、視覚言語推論、ビデオベースの世界モデリング、またはアクション生成に特化していますが、最初に将来の観測を合成してからアクションを推論するカスケード パイプラインでは、インターフェイスのボトルネックや複合予測エラーが発生する可能性があります。私たちは、単一の Omni フレームワーク内でビジョン (ビデオと画像)、言語、アクションを共同モデル化する統合マルチモーダル基盤モデルである iFLYTEK-Embodied-Omni を紹介します。そのモダリティ固有の視覚言語、ビデオ生成、およびアクション生成コンポーネントは、共有されたマルチモーダル自己注意を通じて通信します。この設計により、脳と小脳の連携が確立されます。視覚言語モデルとビデオ生成モデルは、指示の理解、タスクの計画、進捗状況の追跡、および将来の視覚状態の予測のための高レベルの脳を形成しますが、アクション生成モデルは、計画されたサブ目標と共有されたマルチモーダルなコンテキストを実行可能なアクションのチャンクに直接変換する低レベルの小脳として機能します。これらの機能を開発するために、人間のデモンストレーションやロボットのインタラクションからの、アクションの注釈付きおよびアクションのない具体化されたビデオを、具体化された推論、具体化された知覚、および汎用の画像テキストデータと組み合わせて、包括的なデータセットを構築します。さらに、完全なモデルを共同で微調整する前に、VLM、VGM、および AGM を段階的にトレーニングする 4 段階の戦略を採用しています。
原文 (English)
iFLYTEK-Embodied-Omni Technical Report
General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons. Existing approaches typically specialize in visual-language reasoning, video-based world modeling, or action generation, while cascaded pipelines that first synthesize future observations and then infer actions can introduce interface bottlenecks and compound prediction errors. We present iFLYTEK-Embodied-Omni, a unified multimodal foundation model that jointly models vision(videos and images), language, and action within a single Omni framework. Its modality-specific visual-language, video-generation, and action-generation components communicate through shared multimodal self-attention. This design establishes brain-cerebellum collaboration: the vision-language modeland video generation model form a high-level brain for instruction understanding, task planning, progress tracking, and future visual-state prediction, whereas the action generation modelserves as a low-level cerebellum that directly converts planned subgoals and shared multimodal context into executable action chunks. To develop these capabilities, we combine action-annotated and action-free embodied videos from human demonstrations and robot interactions with embodied reasoning, embodied perception, and general-purpose image-text data to construct a comprehensive dataset. We further adopt a four-stage strategy that progressively trains the VLM, VGM, and AGM before jointly fine-tuning the complete model.
内部多元性とペアごとの比較の限界
ローカルペア比較は、参加型デザインや調整などにおいて、人々が意思決定ルールがどのように機能することを望んでいるのかを学習するための標準ツールです。ただし、それらの使用には 2 つの強力な前提が構築されます。1 つは、ローカル比較は、人が自動化された意思決定ルールにどのように動作してほしいかを示す十分な証拠であるということ、もう 1 つは、人々は常にそれらの比較に決定的に答えることができるということです。私たちは、内部多元主義、つまりルールがどのように動作するかについて複数の権威ある優先順位に従って個人が意思決定ルールを評価するという考えの下で、これらの仮定がどのように損なわれるかを調査します。我々は、決定ルールに対するそのような多元的な優先順位の正式なモデルを提供します。これにより、強制的なローカルペア比較データの 2 つの異なる失敗を識別できるようになります。まず、比例性、平等主義、平等な扱いなどの優先順位は本質的にグローバルなものです。あるケースでそれらが意味することは、他の場所で何が起こるかによって異なるため、局所的な比較ではそれらを把握できない可能性があります。第 2 に、優先順位が局所的に表現可能な場合でも、強く保持されている優先順位間の緊張によって内部対立が生じ、比較が強制されると、潜在的にコストのかかる行動の歪みが生じる可能性があります。次に、モデルを使用して、人々が優柔不断を報告できるようにするという代替案を調査しました。その結果、そうすることで、好みを正確に学習するために必要なクエリの数を大幅に削減できることがわかりました。最後に、私たちのモデルが、これらの優先順位を直接導き出し、人々が何に価値を置くかについて、より忠実で解釈可能な説明を生み出す優先学習方法をどのように指し示しているかを説明します。
原文 (English)
Internal Pluralism and the Limits of Pairwise Comparisons
Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.g., in participatory design or alignment. However, their use builds in two strong assumptions: that local comparisons are sufficient evidence about how a person wants an automated decision rule to behave, and that people can always answer those comparisons decisively. We investigate how these assumptions may be compromised under internal pluralism: the idea that an individual evaluates decision rules according to multiple authoritative priorities about how the rule should behave. We provide a formal model of such pluralistic preferences over decision rules, which then lets us identify two distinct failures of forced local pairwise comparison data. First, priorities such as proportionality, egalitarianism, and equal treatment are inherently global: what they imply in one case can depend on what happens elsewhere, so local comparisons may fail to capture them. Second, even when priorities are representable locally, tension between strongly-held priorities can generate internal conflict, producing potentially costly behavioral distortions when comparisons are forced. We then use our model to investigate the alternative -- allowing people to report indecision -- and our findings suggest that doing so can considerably reduce the number of queries needed to learn preferences accurately. We conclude by describing how our model points toward preference-learning methods that elicit these priorities directly, yielding more faithful and interpretable accounts of what people value.
ASK in the Dark: 部分可観測性の下での不確実性ゲート LLM 支援
部分的な可観測性の下で動作する強化学習エージェントは、不完全な情報に基づいて動作する必要があるため、広範な推論事前分布を保持する小型言語モデル (SLM) からのガイダンスの自然な候補となります。しかし、この設定に SLM ガイダンスを統合するのは難しいことが判明しています。すべてのテスト環境において、バニラの不確実性ゲート アプローチでは上書き率がゼロかそれに近い値に達します。これは、SLM が独立したアクションに寄与することがほとんどないことを意味します。私たちは、この失敗の原因を、真の推論を行うための十分なコンテキストを提供しない、ありのままの自己中心的なプロンプトに起因するものとして追跡し、それを能力の問題ではなくコンテキストの問題として特定します。私たちは、SLM に軌跡を認識したコンテキスト (部分的に明らかになった地図、訪問した位置、行動履歴) と構造化された思考連鎖推論を提供し、SLM を受動的な冗長チェックから、ポリシーを時折修正するより有益なコンサルタントに変換する ASK+ を提案します。さらに、選択的クエリに使用される予測エントロピー信号は、状態の不確実性ではなく行動の不確実性を測定し、POMDP で情報を提供し続けるため、完全に観察可能な設定を超えて不確実性ゲート型支援が実行可能になることを確立します。ステートフル プロンプトは大幅な向上をもたらします。DoorKey では、バニラ ASK が PPO と一致し (両方 89%)、ASK+ は 93% の成功に達します。 FourRooms では、成功率が 53% から 70% に上昇します。 HigherLower では、精度は 73.7% に達し、SLM のみの上限と一致します。すべての環境において、Qwen3.5-2B は Qwen3.5-4B と同等かそれを上回っており、迅速な設計と選択的ゲートがモデル スケールの影響を支配し、大規模なモデルなしでガイダンスを可能にしていることが確認されています。
原文 (English)
ASK in the Dark: Uncertainty-Gated LLM Assistance under Partial Observability
Reinforcement learning agents operating under partial observability must act on incomplete information, making them natural candidates for guidance from small language models (SLMs) that carry broad reasoning priors. Yet integrating SLM guidance into this setting has proven difficult: across all test environments, vanilla uncertainty-gated approaches achieve an overwrite rate at or near zero, meaning the SLM almost never contributes an independent action. We trace this failure to the bare egocentric prompt, which provides insufficient context for genuine reasoning, and identify it as a context problem rather than a capacity problem. We propose ASK+, which supplies the SLM with trajectory-aware context (a partially revealed map, visited positions, and action history) and structured chain-of-thought reasoning, converting it from a passive redundancy check into a more informative consultant that occasionally corrects the policy. We further establish that the predictive entropy signal used for selective querying measures action uncertainty rather than state uncertainty and remains informative in POMDPs, making uncertainty-gated assistance viable beyond fully observable settings. The stateful prompt drives substantial gains: on DoorKey, where vanilla ASK matches PPO (both 89%), ASK+ reaches 93% success; on FourRooms, success climbs from 53% to 70%; on HigherLower, accuracy reaches 73.7%, matching the SLM-only upper bound. Across all environments, Qwen3.5-2B matches or exceeds Qwen3.5-4B, confirming that prompt design and selective gating dominate the impact of model scale, enabling guidance without large models.
科学 AI の自動データ準備
リーダーシップ コンピューティング施設は、AI トレーニング データとして機能する前に日常的に大幅な変換を必要とする大規模な科学データセットを管理します。ただし、自動化された変換、準備状況評価、出所追跡、およびエージェントネイティブの展開を完全に統合する既存のフレームワークはありません。我々は、統合された 5 段階のパイプライン (取り込み、前処理、変換、構造、出力) を通じてこのギャップに対処するオープンソース フレームワークである REDI を紹介します。段階ごとのインストルメンテーションにより、再現性とエージェント呼び出し可能なスキルとしての展開が可能になります。コンパニオン ツール SetGo は、FAIR への準拠とカタログ発行を自動化します。気候、プロテオミクス、材料科学、核融合にわたって評価された REDI は、すべてのデータセットを未加工データから AI 対応データセットに変換し、出力はドメイン専門家の参照に対して検証され、予備的な結果では、気候のケースではフロンティア上の 100 ノードまで理想に近い並列スケーリングが示されています。来歴を計測したプロファイリングにより、ファイル I/O が主要なパイプライン コストであることが明らかになり、フォーマットの選択が一次最適化の手段となります。これらの結果により、REDI は科学 AI の自動データ準備を提供するクロスドメイン プラットフォームとして確立され、データ準備のボトルネックを再現可能で再利用可能なコミュニティ資産に変換します。
原文 (English)
Automated Data Readiness for Scientific AI
Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing framework fully unifies automated transformation, readiness assessment, provenance tracking, and agent-native deployment. We present REDI, an open-source framework that addresses this gap through a unified five-stage pipeline (ingest, preprocess, transform, structure, and output) with per-stage instrumentation for reproducibility and deployment as an agent-callable skill; companion tool SetGo automates FAIR compliance and catalog publication. Evaluated across climate, proteomics, materials science, and nuclear fusion, REDI transforms all datasets from raw to AI-ready, with outputs validated against domain-expert references, and preliminary results show near-ideal parallel scaling to 100 nodes on Frontier for the climate case. Provenance-instrumented profiling reveals file I/O as the dominant pipeline cost, with format selection a first-order optimization lever. These results establish REDI as a cross-domain platform providing automated data readiness for scientific AI, transforming data preparation bottlenecks into reproducible, reusable community assets.
SwarmResearch: オープンエンドのディスカバリーのためのコーディング エージェントのオーケストレーション
自動リサーチなどの長時間実行されるコーディング エージェントは、オープンエンドの問題に対する最適化を永続的に検出できます。ただし、単一の高レベルのアプローチに収束し、問題に対する他の優れたアプローチを見逃しながら、低レベルの編集を進める傾向があります。私たちは、ハーネス レベルの設計上の 2 つの選択がこの動作に寄与していると仮定します。それは、単一の長時間実行エージェントにコンテキストを蓄積することと、単一のプログラム状態のみを編集用に公開することです。オーケストレーターとサブエージェントのハーネスである SwarmResearch を紹介します。このハーネスでは、Shepherd Agent がグローバル コンテキストを使用して検索エージェントの集団を操作し、各検索エージェントがそれぞれの git ブランチでローカル コンテキストで動作します。オープンエンドの最適化タスクでは、SwarmResearch は、より高レベルの探索によって推進される 13/15 タスクで、最先端の LLM 誘導進化とマルチエージェント技術に対するより優れた、または同等のソリューションを発見します。シリアルおよびパラレル エージェントの固定スケーリングと比較して、SwarmResearch のオーケストレーター主導のスケーリングは、さまざまな検索深さで並列処理を適応させることで、よりパフォーマンスの高いソリューションを発見します。
原文 (English)
SwarmResearch: Orchestrating Coding Agents for Open-Ended Discovery
Long-running coding agents such as autoresearch can persistently discover optimizations for open-ended problems. However, they tend to converge onto a single high-level approach, then proceed with low-level edits while missing other superior approaches to the problem. We hypothesize two harness-level design choices contribute to this behavior: accumulating context in a single long-running agent and only exposing a single program state to edit. We introduce SwarmResearch, an orchestrator-subagent harness in which a Shepherd Agent uses global context to steer a population of Search Agents, each operating with local context in their respective git branch. On open-ended optimization tasks, SwarmResearch discovers better or comparable solutions to state-of-the-art LLM-guided evolution and multi-agent techniques on 13/15 tasks, driven by higher-level exploration. Compared with fixed scaling of serial and parallel agents, SwarmResearch's orchestrator-guided scaling discovers better-performing solutions by adapting parallelism at different search depths.
エージェントタスクのためのオブジェクト中心の環境モデリング
大規模言語モデル (LLM) エージェントは経験の蓄積によって改善できますが、自由形式のテキスト記憶は、インタラクションが増大するにつれて維持、検証、再利用することが困難になります。最近のシンボリック アプローチは、実行可能なスキルやプログラムによる世界モデルを学習しますが、ローカル プロシージャを保存したり、単純化されたダイナミクスを想定したりすることがよくあります。私たちは、エクスペリエンスを実行可能なオブジェクト中心の環境モデルに編成するオブジェクト中心環境モデリング (OCM) を提案します。 OCM は、環境エンティティとメカニズムを Python クラスとして定義するオブジェクト ナレッジと、オブジェクト モデルをインポートして使用する必要がある再利用可能な対話パターンを記録するプロシージャ ナレッジの 2 つの接続されたコード ベースを維持します。 OCM はオンライン設定で動作します。各エピソードの後、OCM は軌跡を反映し、両方のナレッジ ベースを更新し、更新されたオブジェクト モデルに対してすべてのプロシージャが実行されることを検証します。今後の対話中、エージェントは進歩的な知識開示を使用して、最初にコンパクトなコード署名を検査し、必要な場合にのみソース コードを読み取ります。実験では、OCM がベンチマーク全体で最良の平均ランクを達成し、無効なアクションが減少することが示されており、エージェントがオブジェクト中心の環境モデルを構築することでメリットを享受できることが実証されています。
原文 (English)
Object-Centric Environment Modeling for Agentic Tasks
Large language model (LLM) agents can improve through accumulated experience, but free-form textual memories become difficult to maintain, validate, and reuse as interactions grow. Recent symbolic approaches learn executable skills or programmatic world models, yet often store local procedures or assume simplified dynamics. We propose Object-Centric Environment Modeling (OCM), which organizes experience into an executable object-centric environment model. OCM maintains two connected code bases: object knowledge, which defines environment entities and mechanisms as Python classes, and procedure knowledge, which records reusable interaction patterns that must import and use the object model. OCM works in an online setting: after each episode, OCM reflects on the trajectory, updates both knowledge bases, and verifies that all procedures execute against the updated object model. During future interaction, the agent uses progressive knowledge disclosure to inspect compact code signatures first and read source code only when needed. Experiments show that OCM achieves the best average rank across benchmarks and reduces invalid actions, demonstrating that agents can benefit from building object-centric environment models.
MedCalc-Pro: LLM エージェントを使用して複雑な医療計算を解決する
医療計算における大規模言語モデル (LLM) を評価するための現在のベンチマークは、主に簡素化された設定に基づいており、各患者の症例は単一の計算機に対応し、必要なツールはクエリで明示的に指定されます。ただし、実際の臨床シナリオでは、共同評価、ネストされたスケール計算、およびターゲットの計算機を直接指定しないファジー クエリのために複数の計算機が必要になることがよくあります。この目的を達成するために、私たちは新しい医療計算ベンチマークである MedCalc-Pro を提案します。これは、単一計算機、複数計算機、およびネストされた計算機の計算設定という、3 つの段階的に困難なタスク設定をカバーします。 MedCalc-Pro には 2,268 件の実際の臨床ケースが含まれており、14 診療科にわたる 77 台の医療計算機をカバーしています。一方、複雑な臨床シナリオにおける既存のフレームワークとメソッドの限られたパフォーマンスに対処するために、構造化された検証と証拠のレビューを通じてパラメーターエラーの伝播を抑制しながら、マルチツールの選択とネストされたツールの呼び出しをサポートする、より一般化可能なエージェントフレームワークをさらに提案します。私たちは、オープンソース、クローズドソース、医療に特化した LLM にわたって体系的な比較を実施しました。その結果、私たちのフレームワークが 3 つのタスク設定すべてで最高のパフォーマンスを達成していることがわかりました。この研究は、困難な医療計算シナリオで LLM を評価および適用するための新しいベンチマークと方法を提供します。
原文 (English)
MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents
Current benchmarks for evaluating large language models (LLMs) in medical calculation are largely based on simplified settings, where each patient case corresponds to a single calculator and the required tool is explicitly specified in the query. However, real clinical scenarios often require multiple calculators for joint evaluation, nested-scale calculation, and fuzzy queries that do not directly specify the target calculator. To this end, we propose a new medical calculation benchmark, MedCalc-Pro, which covers three progressively challenging task settings: single-calculator, multi-calculator, and nested-calculator calculation settings. MedCalc-Pro contains 2,268 real-world clinical cases, covering 77 medical calculators across 14 clinical departments. Meanwhile, to address the limited performance of existing frameworks and methods in complex clinical scenarios, we further propose a more generalizable agent framework that supports multi-tool selection and nested-tool calling, while suppressing parameter error propagation through structured validation and evidence review. We conduct systematic comparisons across open-source, closed-source, and medical-specialized LLMs, and the results show that our framework achieves the best performance across all three task settings. This work provides a new benchmark and method for evaluating and applying LLMs in challenging medical calculation scenarios.
Oyster-II: 大規模言語モデルにおける建設的な安全性調整のための強化学習
大規模言語モデル (LLM) は、さまざまなアプリケーションにわたって優れた機能を実証してきましたが、その安全性、有用性、信頼性を同時に確保することは依然として根深い課題です。従来の拒否指向の調整戦略は、有害なコンテンツの生成を軽減しますが、組織的に正当なユーザーのニーズに応えることができず、多くの場合、機密性の高いクエリの根本的な意図に安全かつ建設的に対処できる情報が差し控えられます。 Oyster-I によって開拓された建設的安全パラダイムを基礎として、包括的な拒否を超え、思慮深い応答指向の安全調整に向けて前進することで、その教師あり微調整 (SFT) ベースのスキームの 2 つの重大な限界を特定します。1 つは配布外のシナリオに対する不十分な安全一般化と、安全指向の推論パターンが良性のクエリに過度に適用され、安全性を低下させる、安全思考連鎖 (CoT) の過剰一般化と呼ばれる現象です。有用性とユーザーエクスペリエンス。これらの制限に対処するために、私たちは強化学習 (RL) ベースの建設的安全調整フレームワークである Oyster-II を提案します。Oyster-II は、多段階強化学習戦略と組み合わせた Zero-RL パラダイムを採用しています。広範なベンチマークにわたって評価された Oyster-II は、安全性の次元で Qwen3-14B とその前身である Oyster-I の両方を包括的に上回り、Qwen3-Max および Qwen3.5-397B に匹敵するクロススケール パフォーマンスを達成します。
原文 (English)
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models
Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulness, and trustworthiness remains a persistent challenge. Conventional refusal-oriented alignment strategies mitigate harmful content generation but systematically fail to serve legitimate user needs, often withholding information that could safely and constructively address the underlying intent of sensitive queries. Building upon the constructive safety paradigm pioneered by Oyster-I, which moves beyond blanket refusal toward thoughtful, response-oriented safety alignment, we identify two critical limitations of its Supervised Fine-Tuning (SFT)-based scheme: insufficient safety generalization to out-of-distribution scenarios and a phenomenon we term safety chain-of-thought (CoT) over-generalization, wherein safety-oriented reasoning patterns are excessively applied to benign queries, degrading helpfulness and user experience. To address these limitations, we propose Oyster-II, a reinforcement learning (RL)-based constructive safety alignment framework that adopts a Zero-RL paradigm combined with a multi-stage reinforcement learning strategy.Evaluated across extensive benchmarks, Oyster-II comprehensively surpasses both Qwen3-14B and its predecessor Oyster-I on safety dimensions, achieving cross-scale performance comparable to Qwen3-Max and Qwen3.5-397B.
VERITAS: 科学研究用の汎用複製ツールを目指して
AI ツールは科学出版を加速させていますが、それを審査するシステムは追いつくのに苦労しており、出版された研究を独立して検証することは困難かつ重要になっています。手動レプリケーションは時間がかかり、コストがかかるため、コーディング エージェントを使用してプロセスの一部を自動化する作業が増えています。既存の取り組みの大部分は、ベンチマーク独自のパイプライン内でのみ実行されるコンパニオン エージェントを備えたベンチマークとしてパッケージ化されており、汎用のレプリケーション ツールは存在しません。 CLI コーディング エージェントを中心に構築されたドメインに依存しないレプリケーション フレームワークである VERITAS を紹介します。論文、コード リポジトリ、またはその両方が与えられると、VERITAS は論文の主張を抽出し、問題が発生したときに解決しながら方法論を実行し、実験実行の証拠に照らして各主張を判断します。パイプラインは、重要度に重み付けされたレプリケーション スコア、適用されたすべての修正の重大度に応じたログ、およびパッチ適用されたコードベースを返します。私たちは、コンピューター サイエンス、社会科学、医学、天体物理学にわたる 65 件の論文、CORE-Bench および ReplicationBench で VERITAS を評価します。同じモデルとホスト環境上の 2 つの強力なクロード コード ベースラインに対して、ベリタスは最先端のパフォーマンスを達成し、両方のベンチマークのすべての指標でリードしています。
原文 (English)
VERITAS: Towards a General-Purpose Replication Tool for Scientific Research
AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of published research has become both harder and more important. As manual replication is slow and expensive, a growing line of work uses coding agents to automate parts of the process. Existing efforts are largely packaged as benchmarks with companion agents that only run inside the benchmark's own pipeline, and no general-purpose replication tool exists. We present VERITAS, a domain-agnostic replication framework built around CLI coding agents. Given a paper, a code repository, or both, VERITAS extracts the paper's claims, runs the methodology while resolving issues as they arise, and judges each claim against the evidence from experiment runs. The pipeline returns an importance-weighted Replication Score, a severity-rated log of every fix applied, and the patched codebase. We evaluate VERITAS on CORE-Bench and ReplicationBench, 65 papers spanning computer science, social science, medicine, and astrophysics. Against two strong Claude Code baselines on the same model and host environment, VERITAS achieves state-of-the-art performance and leads on every metric on both benchmarks.
複数の製品を納品する動的組立フロー工場スケジューリングのためのスライディング ウィンドウ ベースの強化学習
複数製品のキッティング納品は、動的な注文の到着によって供給の依存関係や実行可能な一連のジョブマシンの割り当てが同時に変更されるため、処理と組立を統合するハイブリッド製造システムにおけるリアルタイムのスケジューリングに重大な課題を課します。この論文では、複雑なキッティング制約を持つ柔軟な組み立てフロー ショップのスケジューリング問題におけるエンドツーエンドのオンライン スケジューリングのためのスライディング ウィンドウ ベースの強化学習 (SWRL) フレームワークを提案します。この問題は、二層キッティング構造と、まばらな報酬状況を生み出す最終製品のボトルネックのダイナミクスを捉える、異種グラフベースのマルコフ決定プロセスとして定式化されます。結果として生じる課題に対処するために、SWRL は、非アクティブなノードをフィルタリングしてキッティングクリティカルな操作に優先順位を付けるスライディング ウィンドウ フィルタリング メカニズム、連続する決定状態間でのボトルネック シフトを追跡する時空間グラフ エンコード ネットワーク、および可変トポロジの下で変化するアクション スペースに適応する制約付き待機戦略を備えた動的アクション マッピング モジュールを統合します。家電メーカーによる実際のインスタンスでの実験では、SWRL が従来のディスパッチング ルールや既存の深層強化学習手法に比べて一貫した遅刻削減を達成し、さまざまなリソース構成、注文負荷、到着集中度にわたって堅牢なパフォーマンスを示すことが実証されました。
原文 (English)
A Sliding-Window-Based Reinforcement Learning for Dynamic Assembly Flow Shop Scheduling with Multi-Product Delivery
Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate processing and assembly, as dynamic order arrivals simultaneously alter supply dependencies and the set of feasible job-machine assignments. This paper proposes a sliding-window-based reinforcement learning (SWRL) framework for end-to-end online scheduling in the flexible assembly flow shop scheduling problem with complex kitting constraints. The problem is formulated as a heterogeneous graph-based Markov decision process that captures the dual-layer kitting structure and the tail-product bottleneck dynamics that produce a sparse reward landscape. To address the resulting challenges, SWRL integrates a sliding-window filtering mechanism that filters inactive nodes and prioritizes kitting-critical operations, a spatiotemporal graph encoding network that tracks bottleneck shifts across consecutive decision states, and a dynamic action mapping module with a constrained waiting strategy that adapts to the changing action space under variable topologies. Experiments on real-world instances from a home appliance manufacturer demonstrate that SWRL achieves consistent tardiness reductions over classical dispatching rules and existing deep reinforcement learning methods, and exhibits robust performance across varying resource configurations, order loads, and arrival concentrations.
Incognita を使用した社会的に分散されたタスク環境に基づいたアクションによる生成エージェントの評価
社会環境における効果的なエージェンシーは、エージェントがいつ知識を求め、いつ行動するか、そしてその行動が取得した情報によって正当化されるかどうかに依存します。既存の根拠のあるベンチマークは、実行可能なアクション、永続的な状態、検証可能な結果を提供しますが、ソーシャル シミュレーション環境は言語エージェント間の豊富な対話を提供します。これらの要件を組み合わせた評価設定を検討します。私たちは、社会的に分散されたタスク環境を、役割に分離された参加者間でタスク関連の知識が分割され、結果として生じるアクションに参加者を通じてのみアクセスできる対話型環境として定義します。コミュニケーションは役割分担された知識を探索する役割を果たし、根拠のある行動は環境状態を活用する役割を果たします。ソーシャル インタラクションと地に足の着いた実行を分離する Concordia ベースのフレームワークである Incognita を紹介します。評価されたエージェントはメッセージをユーザーまたは専門家エンティティにルーティングします。専門家が許容される操作を仲介する。決定論的なサブ環境は、正規の状態に対して受け入れられた操作を実行します。そしてオフラインの評価者は、継承された報酬で結果を採点します。 Incognita-Retail は、最終状態の報酬セマンティクスを維持しながら、タウベンチ小売をマルチエンティティ環境に変換します。私たちは、社会的広がりによって層別化された 18 のタスクに関する 3 つの生成エージェント モデルを 540 回の試行で評価しました。進歩は報酬と行動に現れます。成功は 0 パーセントから 8.9 パーセントと 17.2 パーセントに上昇しますが、早期終了は 100 パーセントから 87 パーセントと 58 パーセントに低下します。より強力なモデルは、より多くの隠れた知識を引き出し、より多くのエンティティと接触し、より根拠のある書き込みを試みますが、信頼性は低いままです。これらの調査結果は、社会的に分散されたタスク環境では、知識の引き出し、情報源の選択、根拠のある行動の試み、時期尚早な完了の信念など、確実に成功する前の行動を暴露していることを示しています。
原文 (English)
Evaluating Generative Agents with Actions Grounded in Socially Distributed Task Environments using Incognita
Effective agency in social environments depends on when an agent seeks knowledge, when it acts, and whether its actions are justified by acquired information. Existing grounded benchmarks provide executable actions, persistent state, and verifiable outcomes, while social simulation environments provide rich interaction among language agents. We study an evaluation setting that combines these requirements. We define socially distributed task environments as interactive environments where task-relevant knowledge is partitioned across role-isolated participants and consequential actions are accessible only through them. Communication serves as exploration over role-partitioned knowledge, while grounded action serves as exploitation over environment state. We introduce Incognita, a Concordia-based framework that separates social interaction from grounded execution. The evaluated agent routes messages to a user or specialist entities; specialists mediate admissible operations; a deterministic sub-environment executes accepted operations over a canonical state; and an offline evaluator scores outcomes with inherited rewards. Incognita-Retail transforms tau-bench retail into a multi-entity environment while preserving final-state reward semantics. We evaluate three generative agent models on 18 tasks stratified by social breadth, with 540 trials. Progress appears in reward and behavior: success rises from 0 percent to 8.9 percent and 17.2 percent, while premature finalization falls from 100 percent to 87 percent and 58 percent. Stronger models elicit more hidden knowledge, contact more entities, and attempt more grounded writes, yet reliability remains low. These findings show that socially distributed task environments expose behavior before reliable success, including knowledge elicitation, source selection, grounded action attempts, and premature completion belief.
大規模な言語モデルを使用した証拠を求める診断推論のための強化学習
最近の推論中心の大規模言語モデル (LLM) は大幅な進歩を遂げていますが、主に完全な情報を前提とした受動的推論パターンで動作します。対照的に、現実世界の臨床インテリジェンスは本質的に、戦略的な証拠の取得を必要とする反復的な調査プロセスです。このギャップを埋めるために、私たちは医療診断を反復的な証拠探索タスクとして形式化します。当社は、検証可能な報酬による強化学習 (RLVR) を活用して、診断の精度と検査の一貫性を強化する一連の新しい報酬によって導かれ、閉ループ環境内で固有の推論を引き出します。これを促進するために、現実的で知識に基づいた追跡証拠を提供する高忠実度の臨床オラクルである、検索拡張生成ベース検査シミュレーター (RAGES) を導入します。さまざまなデータセットにわたる実証結果は、私たちのフレームワークにより、LLM が受動的な応答者から自律的なアシスタントに移行できることを示しています。特に、私たちのモデルは、より大規模で推論を強化したベースラインと同等のパフォーマンスを示し、一方、RAGES は、生物学的に妥当な臨床フィードバックの生成においてバニラ LLM よりも優れていることが証明されています。
原文 (English)
Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models
Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that assumes complete information. In contrast, real-world clinical intelligence is inherently an iterative investigative process requiring strategic evidence acquisition. To bridge this gap, we formalize medical diagnosis as an Iterative Evidence-Seeking Task. We leverage Reinforcement Learning with Verifiable Rewards (RLVR) to elicit intrinsic reasoning within a closed-loop environment, guided by a novel suite of rewards that enforce diagnostic precision and examination consistency. To facilitate this, we introduce the Retrieval-Augmented Generation-based Examination Simulator (RAGES), a high-fidelity clinical oracle that provides realistic, knowledge-grounded follow-up evidence. Empirical results across diverse datasets demonstrate that our framework enables LLMs to transition from passive responders to autonomous assistants. Notably, our model demonstrates comparable performance to larger and reasoning-enhanced baselines, while RAGES proves superior to vanilla LLMs in generating biologically plausible clinical feedback.
予測を超えて: 予測市場エージェントの信念から取引までのレイヤー
未来の出来事を予測することは、汎用AIのテストベッドとして注目を集めています。この評価を根拠付ける自然な方法は、モデルを予測市場で取引させることです。ただし、トレーディングには予測以上のものが必要です。さらに、最近のベンチマークでは、調整された確率スコアと取引結果の間に大きなギャップがあることが報告されています。私たちは、私たちの知る限り、予測市場向けの初の自律型取引エージェントである Raven-Agent を提案します。アーカイブされた意思決定セットに対する制御された再生では、私たちのアーキテクチャは、テストされたすべてのポリシーの中で唯一のプラスのリターンと唯一のプラスのリスク調整後のリターンを達成します。コードを https://github.com/Alchemist-X/predict-raven でリリースしました。
原文 (English)
Beyond Forecasting: The Belief-to-Trade Layer in Prediction-Market Agents
Forecasting future events has attracted growing attention as a testbed for general-purpose AI. A natural way to ground this evaluation is let the models trade in the prediction markets. Trading, however, requires more than forecasting. Moreover, recent benchmarks report a substantial gap between calibrated probability scores and the trading results. We propose Raven-Agent, to the best of our knowledge, the first autonomous trading agent for prediction markets. On a controlled replay over an archived decision set, our architecture achieves the only positive return and the only positive risk-adjusted return among all tested policies. We have released our code in https://github.com/Alchemist-X/predict-raven .
人間と AI の協調的な意思決定のための人間中心のリフレクティブ アーキテクチャ
日常業務から安全性が重要なアプリケーションに至るまで、人間の活動のさまざまな領域にわたって大規模言語モデル (LLM) を使用することは、人間によるフィードバックを最小限に抑えながら意思決定の有効性を高めることを目的としています。同時に、AI の非決定性に伴うリスクを軽減しながら、意思決定を人間の期待、好み、ニーズに合わせることを目指しています。しかし、人間は AI の推奨事項に過度に依存したり、過少に依存したりすることが多く、現在の AI システムは依然として人間の期待に合わせて調整されていません。これらの課題に対処するために、人間の能力を強化し、AI エージェントを人間の好みや期待に合わせるように設計された、人間と AI の協調的な意思決定フレームワークを導入します。具体的には、この論文は、(a) AI エージェントと人間のプレイヤーの間の確率的ゲームとして協調的意思決定タスクを定式化し、(b) 反復的かつ内省的なプロセスで言語フィードバックを活用する強化学習エージェントと人間が調整したモデルを統合する人間中心のリフレクティブ アーキテクチャ (HCRA) を提案します。評価結果は、HCRA が意思決定の有効性を高め、高品質の推奨事項を提供することを示しています。
原文 (English)
Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making
The use of Large Language Models (LLMs) across diverse areas of human activity-ranging from everyday tasks to safety-critical applications-aims to enhance decision-making effectiveness with minimal human feedback. Concurrently, it seeks to align decisions with human expectations, preferences, and needs while mitigating risks associated with AI non-determinism. However, humans frequently over- or under-rely on AI recommendations, and current AI systems remain poorly calibrated to human expectations. To address these challenges, we introduce a human-AI collaborative decision-making framework designed to augment human capabilities and align AI agents with human preferences and expectations. Specifically, this paper (a) formulates the collaborative decision-making task as a stochastic game between an AI agent and a human player, and (b) proposes the Human-Centric Reflective Architecture (HCRA), which integrates human-calibrated models with reinforcement learning agents that leverage linguistic feedback in an iterative, reflective process. Evaluation results demonstrate that HCRA enhances decision-making effectiveness and delivers high-quality recommendations.
調査間転送によるシリコンのサンプリング
人間の調査回答者をシミュレートする大規模言語モデル (LLM) を使用するシリコン サンプリングは、従来の調査研究を強化するための有望なアプローチとして浮上しています。ただし、ほとんどの評価は個人レベルの予測ではなく分布比較に依存しているため、パターン マッチングと一貫した回答者レベルの予測が混同される危険があります。私たちは、LLM が 1 セットの質問に対する回答者の回答を与えられ、同じアンケートからのまったく異なる質問に対する回答を予測する必要がある、より厳密な評価フレームワークである調査間転送を提案します。 2024 年の台湾選挙と民主化調査 (TEDS) のデータ、3 つのオープンウェイト LLM (27B ~ 120B パラメーター)、および教師あり機械学習ベースラインを使用すると、次のことがわかります。(1) ゼロショット LLM は、真に目に見えないアイテムに対して 52% の精度を達成し、同じ母集団データで学習された教師ありランダム フォレストの 6 パーセンテージ ポイント (pp) 以内に近い。 (2) 党派的態度の 67% から主権の 23% まで、安定した構成の予測可能性の階層が出現します。 (3) 分散崩壊と安全性調整効果 (よく挙げられる 2 つの LLM 制限) は、以前に報告されたよりも微妙なことが判明し、分散崩壊は教師ありモデルにも影響を及ぼし、調整効果はモデル ファミリ間で劇的に異なります。これらの発見は、シリコンサンプリングの可能性と限界の両方を明らかにします。
原文 (English)
Silicon Sampling via Cross-Survey Transfer
Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects-two commonly cited LLM limitations-turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.
APeB: 大規模言語モデルエージェントのパーソナライゼーション能力のベンチマーク
ユーザーが未加工の未指定のクエリを発行すると、LLM を利用したエージェントはパーソナライゼーションに苦労します。この設定では、エージェントは潜在的な意図を推測し、ノイズの多い対話履歴から好みを抽出し、競合する選択肢の中から選択する必要があります。既存のベンチマークは、ユーザーが絞り込んだクエリや単純化された履歴に依存することが多いため、この機能をテストすることはほとんどありません。生のクエリとさまざまな履歴の下でエージェントによるパーソナライゼーションのテストベッドであるパーソナライズされた製品検索 (PPS) を紹介します。アクション ログからエージェント パーソナライズド ベンチマーク (APeB) を構築し、詳細に指定されていないインテントと豊富な履歴およびユーザーが閲覧した候補アイテムを組み合わせます。マルチステップのエージェントワークフローを備えた最先端の LLM を評価すると、モデルは明示的なクエリをうまく処理しますが、意図と設定の検出を必要とする初期段階のクエリには苦労していることがわかりました。ルーブリック分析では、このギャップは主に歴史の非効率な利用に起因すると考えられます。シンプルな履歴認識クエリ絞り込みパイプライン VQRA は一貫した利益をもたらし、パーソナライズされたエージェントにおける専用の履歴利用モジュールの必要性を浮き彫りにしています。
原文 (English)
APeB: Benchmarking Personalization Ability of Large Language Model Agents
LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy interaction histories, and select among competing alternatives. Existing benchmarks rarely test this capability, as they often rely on user-refined queries or simplified histories. We introduce personalized product search (PPS), a testbed for agentic personalization under raw queries and diverse histories. We construct Agent Personalized Benchmark (APeB) from action logs, pairing underspecified intents with rich histories and user-viewed candidate items. Evaluating state-of-the-art LLMs with multi-step agent workflows, we find that models handle explicit queries well but struggle with early-stage queries requiring intent and preference discovery. Rubric analysis attributes this gap mainly to ineffective history use. A simple history-aware query-refinement pipeline, VQRA, yields consistent gains, highlighting the need for dedicated history-utilization modules in personalized agents.
エージェント的なビジネス プロセス実行のための組織メモリ
LLM ベースのエージェントは、ルールベースのシステムの限界を超えてビジネス プロセスの実行を自動化する新しい機会を提供します。ただし、汎用 LLM には、信頼性の高い実行に必要な組織固有の知識が欠けており、通常、ポリシー、プロセス モデル、標準操作手順など人間中心の成果物全体にわたって断片化されています。このような知識は技術的には個別のプロンプトやエージェント固有の取得設定でエンコードできますが、このアプローチは企業内に拡張できません。知識のサイロ化やルールの重複が生じ、エージェント間での一貫した更新と学習が困難になるからです。これには、エージェントによるビジネス プロセスの実行のための組織のメモリが必要であると私たちは主張します。これは、作業がどのように実行されるべきかについて進化する組織固有の手順知識の共有され、管理され、エージェントが利用できる参照層です。私たちは、そのようなメモリの要件を導き出し、そのキュレーションと使用のためのアーキテクチャを提案し、調達シナリオに基づいた概念実証でその有効性を実証します。
原文 (English)
Organizational Memory for Agentic Business Process Execution
LLM-based agents offer new opportunities for automating business process execution beyond the limits of rule-based systems. However, general-purpose LLMs lack the organization-specific knowledge required for reliable execution, which is typically fragmented across human-oriented artifacts such as policies, process models, and standard operating procedures. While such knowledge can technically be encoded in individual prompts or agent-specific retrieval setups, this approach does not scale in enterprises, as it gives rise to knowledge silos and rule duplicates, and makes consistent updates and learning across agents difficult. We argue that this calls for an organizational memory for agentic business process execution: a shared, governed, and agent-consumable reference layer of evolving organization-specific procedural knowledge about how work should be executed. We derive requirements for such a memory, propose an architecture for its curation and consumption, and demonstrate its effectiveness in a proof-of-concept based on a procurement scenario.
身体化されたオペレーターとベンチマーク: 再利用可能で展開可能な身体化されたインテリジェンス システムに向けて
身体化されたインテリジェンス システムには、エンドツーエンドのポリシー モデルだけでなく、マルチモーダルな観察、ロボットの状態、人間のデモンストレーション、およびタスクのコンテキストを構造化された表現、決定、軌道、制御参照、およびシステム サービスに変換する再利用可能な機能モジュールも必要です。この研究では、これらのモジュールを具体化されたオペレーターとして定義し、それらを具体化されたインテリジェンス パイプライン内の独立していながら構成可能なユニットとして研究します。タスクのセマンティクス、標準化された入出力契約、展開可能性、再利用可能性、およびマルチレイヤーの最適化可能性を強調して、それらの定義境界を明確にします。さらに、検出とセグメンテーション、空間位置特定と 3D 理解、手の動きの回復、具体化された基礎モデルとタスク決定オペレーター、計画、制御、およびシステム サポート オペレーターの 5 つのカテゴリをカバーする分類法を構築します。カテゴリごとに、代表的な機能、技術パラダイム、アプリケーションの役割、実際の制限事項をまとめています。分類を超えて、正確性、エンドツーエンドの効率、リソース使用量、時間的安定性、移植性、インターフェイスの互換性、展開の信頼性、下流タスクのユーティリティの観点から具体化された演算子を評価する多次元ベンチマーク フレームワークを提案します。また、ワークフロー レベルのオペレーターの高速化と、オペレーターの構成、データの標準化、ワールド モデル、VLA の安全性、エッジ展開、および実際のアプリケーションの価値における未解決の課題についても説明します。全体として、この研究は、再利用可能でスケーラブルで検証可能な組み込みインテリジェンス システムの基盤を提供する、統合された展開可能なコンポーネントとして統合オペレーターが最適化および評価されるべきであると主張しています。
原文 (English)
Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems
Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal observations, robot states, human demonstrations, and task contexts into structured representations, decisions, trajectories, control references, and system services. This work defines these modules as embodied operators and studies them as independent yet composable units in embodied intelligence pipelines. We clarify their definition boundary, emphasizing task semantics, standardized input-output contracts, deployability, reusability, and multi-layer optimizability. We further construct a taxonomy covering five categories: detection and segmentation, spatial localization and 3D understanding, hand motion recovery, embodied foundation models and task-decision operators, and planning, control, and system support operators. For each category, we summarize representative functions, technical paradigms, application roles, and practical limitations. Beyond taxonomy, we propose a multi-dimensional benchmark framework that evaluates embodied operators in terms of correctness, end-to-end efficiency, resource usage, temporal stability, portability, interface compatibility, deployment reliability, and downstream task utility. We also discuss workflow-level operator acceleration and open challenges in operator composition, data standardization, world models, VLA safety, edge deployment, and real-world application value. Overall, this work argues that embodied operators should be optimized and evaluated as holistic deployable components, providing a foundation for reusable, scalable, and verifiable embodied intelligence systems.
反省的な対話か、それとも迅速な改善か?生徒のプログラミングのための独立した LLM の使用に対する家庭教師の足場の影響
大規模言語モデル (LLM) は学習において個人に合わせたサポートを提供できますが、いくつかの研究で教育での使用について懸念が生じています。重要なのは、学習は学生が LLM にどのように関与するかによって決まります。この研究では、2 種類の LLM ベースの家庭教師が生徒のプロンプトの実践、学習、その後の LLM の使用をどのように形成するかを調査しました。対話的な質問を通じて対話を構築するソクラティック ガイダンス (SG) 家庭教師と、効果的なプロンプトの作成をガイドするプロンプト リファインメント (PR) 家庭教師です。私たちは、大学院レベルのモバイル ロボット工学コースで 2 段階の研究を実施しました。66 人の学生が 6 週間の介入中に SG または PR 講師のいずれかを使用し、続いて 52 人の学生が 3 週間のコース プロジェクト中に制約のない LLM を使用しました。結果は、SG 講師と PR 講師は、ガイド付き使用中に同様のタスクのパフォーマンスとプロンプトパターンをもたらしましたが、学習成果とその後の LLM の使用において異なることを示しています。 SG の学生は、PR の学生と比較して、後のセッションでより高い学習成果を達成し、制約のない LLM を使用した場合、理解度の向上を予測する、理解主導のプロンプト戦略を採用する可能性が高くなりました。学習者は SG 講師の効率が低いと認識していましたが、この調査結果は、ソクラテスの指導が時間の経過とともに LLM で学習する生徒の能力の発達をサポートしていることを示唆しており、LLM 講師の設計におけるソクラテスの重要性を強調しています。
原文 (English)
Reflective Dialogue or Prompt Refinement? Effects of Tutor Scaffolding on Students' Independent LLM Use for Programming
While Large Language Models (LLMs) can provide personalized support in learning, several studies have raised concerns regarding their use in education. Importantly, learning depends on how students engage with LLMs. This study examined how two types of LLM-based tutors shape students' prompting practices, learning, and subsequent LLM-use: a Socratic-Guidance (SG) tutor, which structures interaction through dialogic questioning, and a Prompt-Refinement (PR) tutor that guides the formulation of effective prompts. We conducted a two-phase study in a graduate-level mobile robotics course: 66 students used either the SG or PR tutor during a 6-week intervention, followed by 52 students using an unconstrained LLM during a 3-week course project. Results show that while the SG- and PR tutors led to similar task performance and prompting patterns during guided use, they differ in learning outcomes and later LLM-use. SG-students, relative to PR-student, achieved higher learning gains in later sessions, and were more likely to adopt understanding-driven prompting strategies, which are predictive of higher understanding, when using an unconstrained LLM. Although learners perceived the SG tutor as less efficient, the findings suggest that Socratic guidance supports the development of students' capacity to learn with LLMs over time, highlighting its importance for LLM tutor design.
集計の調整が誤解を招く場合: 州ごとの専門家のアクションを行わない監査ポリシーの修復
エージェントティック AI システムは、意思決定ポリシーの編集、改良、修復に使用されることが増えていますが、州ごとの専門家のアクション ラベルが利用できない場合、これらの編集を評価することは困難です。私たちはこの問題をホテル価格設定シミュレーターで研究しました。エージェントのポリシー編集者は地域レベルの診断フィードバックのみを受信します。つまり、価格分布が時間、在庫、市場地域全体でベンチマークポリシーとどのように異なるかの概要のみです。編集者は、ベンチマーク アクション、ベンチマーク ソース コード、報酬数値、または保留された結果を観察することはできず、ターゲット アクション テーブルに対して制約付きの編集を提案することしかできません。 5,000 回のホールドアウト エピソードで、複数回再起動した LLM エディターは RevPAR 108.47 (95% CI 107.61 ~ 109.34) に達し、ベンチマーク ポリシーの 108.75 (107.81 ~ 109.68) に近く、ペア ギャップ (LLM マイナス ベンチマーク) -0.276 および 95% CI [-0.692、 0.146]。安価な診断予測はすでに収益の多く (107.90) を回収しているため、LLM エディターの顕著な利益は生の収益増加だけではなく、エピソード構成距離も 1.153 から 0.609 に減少します。これは、ベンチマーク以外の修復結果としては最も強力です。このプロファイルは再起動検索だけでは説明できません。最大 2,500 の評価を持つ非セマンティック プロポーザーでは、RevPAR ポイントが 8.77 ~ 14.57 ポイント不足します。また、それはもっともらしいプロンプト形式によっても説明されません。シャッフル診断コントロールが領域エラーの対応を破り、RevPAR 94.30 に落ちます。一致は本物ですが、部分的です。ツリー エディタは、より強力なプール アライメント (0.266 に対して 0.214)、およびより強力な参照状態 D1 (0.328 対 1.197) を達成しますが、収益は 98.91 に低下します。これらの結果は、エージェントによる政策修復は、単一の行動距離によってではなく、診断フィードバックが信頼できる閉ループ結果になるかどうかによって評価されるべきであることを示しています。
原文 (English)
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable. We study this problem in a hotel-pricing simulator where an agentic policy editor receives only region-level diagnostic feedback: summaries of how its price distribution differs from a benchmark policy across time, inventory, and market regions. The editor cannot observe benchmark actions, benchmark source code, reward numbers, or held-out outcomes, and may only propose constrained edits to a target-action table. On 5,000 held-out episodes, a multi-restart LLM editor reaches RevPAR 108.47 (95% CI 107.61 - 109.34), close to the benchmark policy's 108.75 (107.81 - 109.68), with paired gap (LLM minus benchmark) -0.276 and 95% CI [-0.692, 0.146]. A cheap diagnostic projection already recovers much of the revenue (107.90), so the LLM editor's distinctive gain is not raw revenue lift alone: it also reduces episode composition distance from 1.153 to 0.609. This is the strongest non-benchmark repair result. This profile is not explained by restart search alone: non-semantic proposers with up to 2,500 evaluations fall 8.77 - 14.57 RevPAR points short. Nor is it explained by plausible prompt format: a shuffled-diagnostic control breaks region-error correspondence and falls to RevPAR 94.30. The match is genuine but partial. A tree editor achieves stronger pooled alignment, 0.214 versus 0.266, and stronger reference-state D1, 0.328 versus 1.197, yet revenue falls to 98.91. These results show that agentic policy repair should be evaluated by whether diagnostic feedback becomes reliable closed-loop outcome, not by a single behavioral distance.
モバイル データからビジネス インサイトまで: 大規模な都市モビリティ分析と意思決定支援のためのエンドツーエンド分析フレームワーク
モバイル アプリケーションから得られるリアルタイムの位置データは、観光計画、駐車場管理、バス ルートの最適化、リソースの割り当てなど、都市のさまざまな課題に対処するための強力なツールです。さらに、位置情報ベースのサービス、市場シェア分析、行動プロファイリングなどの商業分野での戦略的意思決定を形成するための貴重な洞察も提供します。この広範な研究では、都市環境、特に観光、交通、小売の分野におけるスマートフォン ユーザーの行動とパターンを調査することで、前述のすべての課題に対処することを目指しています。私たちのアプローチには、洗練されたデータ プラットフォームの開発が開始から実装まで含まれており、これにはユース ケースの定式化、アーキテクチャ設計、モジュールの実装が含まれます。当社では、データの匿名化、ETL パイプライン、データ処理と機械学習モデルの開発のための Google BigQuery と Vertex AI の利用など、最先端の技術とテクノロジーを採用しています。複数の利害関係者主導のユースケースをサポートするデータ製品を生成するために、再利用可能な分析ビルディング ブロックに基づくモジュラー アーキテクチャが開発されました。さらに、Power BI を介した対話型のデータ視覚化手法を適用して、関係者による分析結果の効果的な解釈を促進します。開発されたモデルは、モビリティ プロファイリング、頻繁な軌跡マイニング、影響範囲分析、交通異常検出、起点宛先パターン分析など、幅広いモビリティ分析タスクに対応します。この結果は、このフレームワークがユーザーのモビリティのダイナミクスを優れた空間的および時間的解像度で捕捉し、都市計画や戦略的なビジネス上の意思決定に実用的な洞察を提供する能力を示しています。
原文 (English)
From Mobile Data to Business Insights: An End-to-End Analytics Framework for Large-Scale Urban Mobility Analysis and Decision Support
Real time location data derived from mobile applications is a powerful tool for addressing various urban challenges, including tourism planning, parking management, bus route optimization, and resource allocation. Besides, it offers invaluable insights for shaping strategic decisions in commercial domains such as location based services, market share analysis, and behavioral profiling. In this expansive study, we aim to address all of the aforementioned challenges by investigating the behaviors and patterns of smartphone users within urban environments, particularly in the domains of tourism, transportation, and retail. Our approach encompasses the development of a sophisticated data platform from inception to implementation, which includes the formulation of use cases, architectural design, and implementation of modules. We employ state of the art techniques and technologies, including data anonymization, ETL pipelines, and utilizing Google BigQuery and Vertex AI for data processing and machine learning model development. A modular architecture based on reusable analytical building blocks was developed to generate data products that support multiple stakeholder driven use cases. Additionally, we apply interactive data visualization techniques via Power BI to facilitate the effective interpretation of analytical findings by stakeholders. The developed models address a wide range of mobility analytics tasks, including mobility profiling, frequent trajectory mining, area of influence analysis, traffic anomaly detection, and origin destination pattern analysis. The results demonstrate the framework's ability to capture user mobility dynamics at fine spatial and temporal resolutions, providing actionable insights for urban planning and strategic business decision making.
コンセプト グラフを使用した T2I 拡散モデルにおける効率的なバイアス緩和
Text-to-Image 拡散モデルは、トレーニング データから受け継いだ有害なバイアスを伝播することがよくあります。既存のバイアス軽減技術は通常、テキスト エンコーダーでのみ介入するか、推論時のガイダンスを提供するため、多くの場合、意味的に一貫性のない出力に崩壊する生成につながります。これらの制限に対処するために、モデルの内部コンセプト オントロジーに基づいて動作するコンセプト グラフ アライメントに基づく新しいバイアス軽減アプローチである CO-ALIGN (コンセプト オントロジー アライメント) を導入します。 CO-ALIGN は、テキスト エンコーダーとデノイザー内の概念を調整することにより、生成の完全性を維持しながら大幅なバイアスの削減を実現します。我々は、テキスト エンコーダー、デノイザー、および共同テキスト デノイザー オントロジー アライメントという 3 つのパラダイムにわたる概念とグラフのアライメントの有効性を実証します。 CO-ALIGN は最新技術を上回り、公平性を $30\%$、画質で $\Delta FID=11.4$、画像忠実度で $2.8\%$ 向上させ、同時に意味的に一貫性のない出力を $88\%$ 削減します。バイアスの軽減だけでなく、CO-ALIGN が他の下流タスクにも利益をもたらすことを示します。特に、私たちの実験は、より適切に調整された内部オントロジーが、複数の非学習手法にわたって概念の非学習の堅牢性を強化することを実証しています。
原文 (English)
Efficient bias mitigation in T2I diffusion models using Concept Graphs
Text-to-Image diffusion models often propagate harmful bias inherited from the training data. Existing bias mitigation techniques typically intervene only at the text encoder or provide inference-time guidance, often leading to generations that collapse into semantically incoherent outputs. To address these limitations, we introduce CO-ALIGN (Concept Ontology Alignment), a novel bias mitigation approach based on concept-graph alignment that operates on the model's internal concept ontology. By aligning concepts within the text encoder and denoiser, CO-ALIGN achieves substantial bias reduction while preserving generative integrity. We demonstrate the effectiveness of concept-graph alignment across three paradigms: text-encoders, denoisers and joint text-denoiser ontology alignment. CO-ALIGN outperforms the state of the art, improving fairness by $30\%$, $\Delta FID=11.4$ in image quality, $2.8\%$ in image fidelity, all while reducing semantically incoherent outputs by $88\%$. Beyond bias mitigation, we show that CO-ALIGN benefits other downstream tasks as well. In particular, our experiments demonstrate that better-aligned internal ontologies enhance concept unlearning robustness across multiple unlearning techniques.
パーソナライズされた因果的救済: 人間参加型のアプローチ
アルゴリズムによるリソースは、潜在的に一か八かのシナリオにおいて、機械学習による不利な決定の影響を受けるユーザーに合わせた推奨事項を提供するという課題に対処します。従来の解決策のアプローチでは、最も近い反事実の説明に依存したり、ユーザーの因果構造に関する先験的な知識を前提としたりすることが多く、その結果、個々のコンテキストや特定の機能の相互作用を無視した介入が発生します。これらの制限を克服するために、私たちは、リソース推奨を生成する前に、ベイズ推論による対話型クエリを通じてユーザーの構造的因果モデルを反復的に近似する人間参加型フレームワークを研究します。このフレームワークは、人間のフィードバックを利用して因果関係の特定を改善し、各ユーザーの実際の因果関係に合わせた、妥当でコスト効率の高いパーソナライズされた手段を可能にします。概念実証として、人間の応答をシミュレートしてこのフレームワークを評価します。線形因果モデルと非線形因果モデルにわたる私たちのシミュレーションは有望な結果を示していますが、複雑な非線形構造の捕捉には課題が残されており、正確な近似と堅牢なノイズ分布モデリングの重要性が強調されています。
原文 (English)
Personalized Causal Recourse: A Human-In-The-Loop Approach
Algorithmic recourse addresses the challenge of providing tailored recommendations to users affected by unfavorable machine learning decisions, in potentially high-stakes scenarios. Traditional approaches to recourse often rely on the closest counterfactual explanations or assume a priori knowledge of a user's causal structure, resulting in interventions that overlook individual contexts and specific feature interactions. To overcome these limitations, we study a human-in-the-loop framework that iteratively approximates the user's structural causal model through interactive queries via Bayesian inference before producing recourse recommendations. This framework exploits humans' feedback to improve the identification of causal effects, allowing personalized recourse that is plausible, cost-effective, and aligned with the actual causal dependencies of each user. As a proof of concept, we evaluate this framework through simulated human responses. Our simulations across linear and non-linear causal models show promising results, though challenges remain in capturing complex, non-linear structures, emphasizing the importance of accurate approximations and robust noise distribution modeling.
条件付きポリシーの混合による一般化の失敗の実証
フロンティア言語モデルのポストトレーニングは厳選されたタスクスイートで実行され、必然的にトレーニング環境とデプロイメント環境の間で分散シフトが生じます。これにより、開発者は比較的よく理解されていない一般化の失敗にさらされます。このような一般化の失敗をよりよく理解するには、コミュニティが簡素化された条件の下でクリーンなデモンストレーションを構築する必要があると私たちは考えています。これを促進するために、トレーニング タスクの特定の分布で強化学習 (RL) を使用して後でトレーニングするときに、制御可能な方法で一般化できない言語モデルを構築するためのシンプルで柔軟な方法を提案します。私たちの構築では、「条件付きポリシー」のコレクションに対応するトランスクリプトの混合のデータセットに対して教師付き微調整を使用します。これらのポリシーには、それぞれ異なるタスク分布で特定の動作を個別に割り当てることができ、「条件付きポリシーの混合」として適切に近似されるモデルが得られます。 RL トレーニングでは、トレーニング分布で最高の報酬を得るポリシーが選択されることがわかります。これにより、顕著な動作が生じる可能性があります。2 つの異なる「トリガー文字列」が先頭に付加された同一の質問を含む 2 つのディストリビューションが含まれる制御された設定では、基礎となるタスクが同一であっても、どちらかのディストリビューションで RL トレーニングを行うと、もう一方のディストリビューションのパフォーマンスが積極的にゼロに低下します。また、私たちの構築を使用して、タスク範囲と時間的コンテキストの分布の変化にそれぞれ対応する、将来の言語モデルで一般化が失敗する可能性がある 2 つの新しい方法を説明します。私たちの構築は意図的に単純であり、「自然な」汎化の失敗にはあまり似ていないかもしれませんが、結果として得られる「モデル生物」はアライメントストレステストや汎化科学にとって興味深いものであり、トレーニングの成功と汎化が構造化された方法で分離できることの存在証明として使用できます。
原文 (English)
Demonstrating Generalization Failures via Mixtures of Conditional Policies
Post-training of frontier language models is conducted on curated task suites, and inevitably leaves a distribution shift between training and deployment environments. This exposes developers to generalization failures, which are relatively poorly understood. To better understand such generalization failures, we believe the community should construct clean demonstrations under simplified conditions. To facilitate this, we propose a simple and flexible way to construct language models which fail to generalize in controllable ways when subsequently trained with Reinforcement Learning (RL) on a given distribution of training tasks. Our construction uses Supervised Fine-Tuning on a dataset of a mixture of transcripts corresponding to a collection of 'conditional policies', which can each independently be assigned certain behaviors on each different task distribution, to obtain a model that is then well approximated as a 'mixture of conditional policies.' We observe that RL training then selects for policies that obtain the highest reward on the training distribution. This can produce striking behaviors: in a controlled setting with two distributions containing identical questions prepended with two different 'trigger strings', RL training on either distribution actively degrades performance on the other to zero, even though the underlying task is identical. We also use our construction to illustrate two novel ways in which generalization may fail in future language models, corresponding to distribution shifts of task coverage and temporal context respectively. While our construction is deliberately simple and may not closely resemble 'natural' generalization failures, the resulting 'model organisms' are of interest for alignment stress-testing and generalization science, and can be used as existence proofs that training success and generalization can come apart in structured ways.
MentalThink: メンタル SVG ワールドで思考を形成する
私たちは、マルチモーダル LLM (MLLM) に「精神的」視覚化のための実行可能なメカニズムを装備する視覚的記号推論パラダイムである MentalThink を紹介します。 MentalThink の中核は、Think-with-SVG パイプラインであり、モデルは、マルチターン推論のための中間視覚表現としてスケーラブル ベクター グラフィックス (SVG) コードを生成、レンダリング、解釈する方法を学習します。構造化されたベクター スケッチを作成することにより、モデルは空間仮説を外部化し、決定論的レンダリングを通じて検証し、制約された幾何学的空間内で推論することができ、人間の心的イメージのプロセスを効果的に模倣します。このパラダイムは、SVG の構文調整のための教師あり微調整 (SFT) とマルチターン強化学習 (RL) を組み合わせた 2 段階のトレーニング フレームワークを通じてインスタンス化され、中間視覚仮説の反復検査、修正、改良を促進します。広範な評価により、MentalThink が空間理解と推論ベンチマークで優れたパフォーマンスを達成していることが実証され (例: VSIBench で 55.1%、MindCube で 76.0%)、実行可能なベクトル グラフィックスが、動的な透視図法、視覚的反射、および構成的なシーンの構築のための検証可能な視覚的ワークスペースを提供することが示されています。
原文 (English)
MentalThink: Shaping Thoughts in Mental SVG World
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink is a think-with-SVG pipeline, where the model learns to generate, render, and interpret scalable vector graphics (SVG) code as an intermediate visual representation for multi-turn reasoning. By creating structured vector sketches, the model can externalize spatial hypotheses, inspect them through deterministic rendering, and reason within a constrained geometric space, effectively mimicking the human process of mental imagery. We instantiate this paradigm through a two-stage training framework, combining Supervised Fine-Tuning (SFT) for SVG syntactic alignment with multi-turn Reinforcement Learning (RL) to encourage iterative inspection, revision, and refinement of intermediate visual hypotheses. Extensive evaluations demonstrate that MentalThink achieves superior performance on spatial understanding and reasoning benchmarks (e.g., 55.1% on VSIBench, 76.0% on MindCube), showing that executable vector graphics provide a verifiable visual workspace for dynamic perspective taking, visual reflection, and compositional scene construction.
ファジーメンバーシップ関数を使用した解答セットプログラミングの適用: ケーススタディ
人間の推論は、高い、安い、高い、安いなどの言語ラベルによって表現される定性的な概念を通じて機能することが多く、その解釈は文脈に依存し、数値データに基づいているにもかかわらず、通常は曖昧です。この論文では、数値情報と定性的推論の橋渡しをするための、解答セット プログラミング (ASP) の新しいファジー ロジック ベースの定性的拡張を検討します。別の研究で正式に導入された基礎となる言語は、厳密なしきい値を回避し、曖昧さの下で堅牢な推論をサポートする原則に基づいたフレームワークを提供します。代表的なユースケースに焦点を当てて、フレームワークが数値的に根拠のある入力 (機械学習モデルの出力など) を定性的なラベルに対する記号的推論とどのように統合するかを示します。学習ベースのメンバーシップ関数や意味的に強化された述語などの主要な機能により、統一された宣言設定内で専門知識、文脈上の要因、主観的な解釈を組み合わせることができます。
原文 (English)
Applying Answer Set Programming with Fuzzy Membership Functions: a Case Study
Human reasoning often operates through qualitative concepts expressed by linguistic labels such as high, low, expensive, or cheap, whose interpretation depends on context and is usually vague, despite being rooted in numerical data. This paper explores a novel fuzzy-logic-based qualitative extension of Answer Set Programming (ASP) to bridge numerical information and qualitative reasoning. The underlying language, formally introduced in a separate work, provides a principled framework that avoids rigid thresholds and supports robust reasoning under vagueness. Focusing on a representative use case, we illustrate how the framework integrates numerically grounded inputs (such as outputs of machine learning models) with symbolic reasoning over qualitative labels. Key features, including learning-based membership functions and semantically enriched predicates, enable the combination of expert knowledge, contextual factors, and subjective interpretations within a unified declarative setting.
議論を避ける方法: 二重効率の対話型証明によるスケーラブルな AI の安全性
AI モデルが強力な機能を開発し続けるにつれて、その出力が私たちの意図と一致していることを検証できることが重要になります。最近の研究では、ディベートによる検証に焦点を当てています。これは、2 つの競合する強力な証明者または AI モデルが相互に議論して、弱い検証者または人間に主張の正しさを納得させる対話型証明のモデルです。ただし、議論では 2 つの AI モデルが同等の能力を持ち、そのうちの 1 つが真実であると仮定されていますが、これは現実的ではない可能性があります。この研究では、\emph{議論を避ける方法}を示します。つまり、AI の安全性のための \emph{単一証明者} の対話型証明の研究を開始します。単一証明者の対話型証明における以前の結果は、AI の安全性設定にすぐには引き継がれません。たとえば、計算が人間の判断や Web などの外部データベースなどのオラクルにアクセスする場合、それらは機能しません。我々は、(1) オラクルのクエリに対する答えのごく一部が間違っていても出力が変わらないという意味で計算が堅牢である、または (2) オラクルが低次の多項式であるという設定における、オラクル支援計算 (相対化証明とも呼ばれる) のための二重効率の単一証明者の対話型証明と引数を提示します。これらの結果は、構造化された、またはノイズ耐性のあるオラクルアクセスの下では、議論がなくても対話型検証が可能であることを示唆しています。
原文 (English)
How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs
As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions. A recent line of work focuses on verification via debate, a model of interactive proofs where two competing powerful provers, or AI models, debate each other to convince a weak verifier, or a human, of the correctness of their claim. However, debate assumes that the two AI models possess equal abilities and that one of them is truthful, which may not be realistic. In this work, we show \emph{how to avoid debate}: we initiate the study of \emph{single-prover} interactive proofs for AI safety. Prior results in single-prover interactive proofs do not immediately carry over to the AI safety setting: for example, they do not work when the computation has access to an oracle, such as to human judgment or an external database such as the web. We present doubly-efficient single-prover interactive proofs and arguments for oracle-aided computations (also known as relativizing proofs), in the settings where (1) the computation is robust, in the sense that the output does not change if at most a small fraction of the answers to oracle queries are incorrect, or (2) the oracle is a low-degree polynomial. These results suggest that interactive verification is possible even without debate, under structured or noise-tolerant oracle access.
人工知能における厳密性の役割
人工知能 (AI) は、成熟した分野に関連する概念的および科学的基盤の多くが欠けているにもかかわらず、驚異的な能力を達成しました。信頼できるテクノロジーが理論的な理解から生まれる伝統的な科学とは異なり、現代の AI は主にパフォーマンス主導の反復と「錬金術」実験を通じて進歩してきました。この緊張感が、厳密さのレンズを通して AI を体系的に分析する動機となります。概念的厳密さ (基礎概念の明確化)、認識的厳密さ (科学的理解を確立)、運用的厳密さ (信頼性の高いパフォーマンスと展開の確保) で構成される 3 つの部分からなるフレームワークを導入します。このフレームワークを使用して、知能と理解に関する競合する概念、深層学習への経験的アプローチの長所と限界、ベンチマークの威力と落とし穴、現代の AI システムによってもたらされる理論開発への障害を分析します。私たちは、AI の独特の軌跡は、厳密性の形式がパラダイム間でどのように相互作用するかによって生じ、その結果、現代のディープラーニングにおける操作上の厳密性の優位性をもたらしたと主張します。この視点は、AI の急速な進歩とその根強い不確実性の両方を説明するのに役立ち、同時に AI を成熟した科学と信頼できるテクノロジーに変える際に伴う課題を明確にします。
原文 (English)
The Role of Rigor in Artificial Intelligence
Artificial intelligence (AI) has achieved extraordinary capabilities despite lacking many of the conceptual and scientific foundations associated with mature disciplines. Unlike traditional sciences, where reliable technology typically emerges from theoretical understanding, modern AI has progressed largely through performance-driven iteration and "alchemical" experimentation. This tension motivates a systematic analysis of AI through the lens of rigor. We introduce a three-part framework consisting of conceptual rigor (clarifying foundational concepts), epistemic rigor (establishing scientific understanding), and operational rigor (ensuring reliable performance and deployment). Using this framework, we analyze competing conceptions of intelligence and understanding, the strengths and limitations of the empirical approach to deep learning, the power and pitfalls of benchmarks, and the obstacles to theory development posed by modern AI systems. We argue that the distinctive trajectory of AI arises from how forms of rigor interact across paradigms, resulting in the primacy of operational rigor in modern deep learning. This perspective helps explain both AI's rapid advances and its persistent uncertainties, while clarifying the challenges involved in transforming AI into a mature science and reliable technology.
協調的なパーティション最適化による堅牢で実現可能なルート構築
大規模な静電容量式車両経路指定問題 (CVRP) は、通常、顧客を個別に最適化できる小さな経路指定問題に分割することで解決されます。これにより計算の複雑さは大幅に軽減されますが、独自に構築されたルーティング ソリューションでは、フリート内の他の場所に十分なリソースが存在する場合でも、一部の顧客の需要が満たされない可能性があります。我々は、固定パーティションやその後のグローバルな再最適化段階だけに依存するのではなく、最適化中に顧客と車両を交換するために独立して解決されたサブ問題を可能にするルーティング フレームワークである Collaborative Routing Constructors (CoRC) を紹介します。 AGS ベンチマーク インスタンスと最大 200,000 の顧客を含む合成インスタンスでの計算実験では、CoRC を独立ルーティング、ルーティング後のグローバル再最適化、最先端のエンドツーエンド ルーティング フレームワークと比較します。評価されたすべてのパーティショニング戦略にわたって、CoRC は、競合するパーティションベースの手法が構築できない実現可能なルーティング ソリューションを一貫して構築します。さらに、評価されたエンドツーエンド ルーティング フレームワークが同じ計算量の下で解決策を生成できなかった問題のインスタンスに対しても引き続き有効です。これらの結果は、ルーティング部分問題間の連携により、実現可能な大規模ルート構築のための堅牢かつスケーラブルなアプローチが提供されることを示しています。
原文 (English)
Robust Feasible Route Construction through Collaborative Partition Optimization
Large-scale Capacitated Vehicle Routing Problems (CVRPs) are commonly solved by partitioning customers into smaller routing problems that can be optimized independently. While this substantially reduces computational complexity, independently constructed routing solutions may leave some customer demand unserved even when sufficient resources exist elsewhere in the fleet. We present Collaborative Routing Constructors (CoRC), a routing framework that enables independently solved subproblems to exchange customers and vehicles during optimization rather than relying solely on a fixed partition or a subsequent global re-optimization stage. Computational experiments on AGS benchmark instances and synthetic instances containing up to 200,000 customers compare CoRC against independent routing, post-routing global re-optimization, and state-of-the-art, end-to-end routing frameworks. Across all evaluated partitioning strategies, CoRC consistently constructs feasible routing solutions where competing partition-based methods do not. Furthermore, it remains effective on problem instances for which the evaluated end-to-end routing frameworks did not produce solutions under the same computational budget. These results demonstrate that collaboration between routing subproblems provides a robust and scalable approach for feasible large-scale route construction.
Pivotal-Aware 自己フィードバック再試行によるエージェント強化学習
大規模言語モデル (LLM) エージェントは、長期にわたるインタラクティブなタスクにおいて強力な意思決定能力を示していますが、失敗した軌跡を効果的に活用するのに依然として苦労しています。完全な再試行には高いインタラクション コストが発生し、エクスペリエンスの取得は重要なエクスペリエンス シグナルを弱める傾向があります。 To address this, we propose PivoARL, a self-feedback retry framework for experience exploitation in LLM agents. PivoARL は、構造化リフレクションを通じて重要な誤ったターンを特定し、対応する重要な状態からのみローカルな再試行を実行します。これにより、正しいプレフィックスが再利用され、冗長なインタラクションが削減されます。情報獲得の観点から、重要な再試行は有用な経験信号をエラー境界付近に集中させ、状態に依存しない経験の利用によって引き起こされる信号の希釈を軽減することをさらに示します。この洞察に基づいて、誤ったサフィックスを分離しながら正しいプレフィックスに報酬を与え、暗黙的なリフレクション リターンを通じてリフレクション品質を最適化する、極めて重要な認識のクレジット割り当てメカニズムを設計します。 We conduct a systematic evaluation on 4 agent tasks and 7 search-based QA benchmarks. Results show that PivoARL achieves significant improvements on Pass@2/3 across all tasks, with an average gain of about 11.5\% over MetaRL.さらに、PivoARL は、重要なターンによって誘発される対照的な優先信号の恩恵を受けて、タスクの 80\% 以上で Pass@1 を一貫して改善します。マインスイーパ環境では、PivoARL は、完全再試行方式と比較して、GiGPO よりも 45\% 以上改善され、インタラクション ターンを平均で約 42\% 削減します。 Code is available at https://github.com/yuki-younai/PivoARL.
原文 (English)
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry
Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajectories: full retries incur high interaction costs, while experience retrieval tends to dilute critical experience signals. To address this, we propose PivoARL, a self-feedback retry framework for experience exploitation in LLM agents. PivoARL identifies the pivotal erroneous turn through structured reflection and performs local retry only from the corresponding pivotal state, thereby reusing the correct prefix and reducing redundant interactions. From an information-gain perspective, we further show that pivotal retry concentrates useful experience signals near the error boundary, mitigating the signal dilution caused by state-agnostic experience utilization. Based on this insight, we design a pivotal-aware credit assignment mechanism that rewards correct prefixes while isolating erroneous suffixes, and optimize reflection quality through implicit reflection returns. We conduct a systematic evaluation on 4 agent tasks and 7 search-based QA benchmarks. Results show that PivoARL achieves significant improvements on Pass@2/3 across all tasks, with an average gain of about 11.5\% over MetaRL. Moreover, benefiting from contrastive preference signals induced by pivotal turns, PivoARL also consistently improves Pass@1 on over 80\% of the tasks. On Minesweeper environment, PivoARL improves over GiGPO by more than 45\% and reduces interaction turns by about 42\% on average compared with full-retry methods. Code is available at https://github.com/yuki-younai/PivoARL.
適応型交通信号制御のための説明可能な強化学習
強化学習 (RL) は、適応型交通信号制御の強力なパラダイムとして登場しました。ただし、交通管制などの安全性が重要なインフラでは、ディープ RL モデルの不透明でブラックボックス的な性質により、交通機関の受け入れ、規制順守、運用の信頼性、トラブルシューティング、および微調整に課題が生じています。高性能の最適化と人間が理解できる解釈可能性の間のギャップを埋めるために、この取り組みでは、安全で透過的な交通信号制御のための、新規で説明可能なエンティティ中心の RL フレームワークを導入します。提案されたアーキテクチャは、モノリシックなフラット ベクトルを通じて交通状態を処理するのではなく、リアルタイムの交差点観測を別個の高次元の車線エンティティと位相時間構成に分解し、交差点の構造トポロジーと幾何学的構成を本質的に保存します。リレーショナル依存関係とレーン間の競合は、シーケンシャルなマルチヘッド クロス アテンション ブロックとセルフ アテンション ブロックを特徴とするデュアルステージ アテンション ネットワークを介して動的に抽出されます。この設計により、特定のアプローチ ボリュームおよびキューに対する信号位相の直接的な影響を定量化するリアルタイム アフィニティ マトリックスが生成され、完全な視覚的および分析的解釈が可能になります。厳密な動作信頼性を確保するために、決定論的なアクション マスキング インターフェイスが Proximal Policy Optimization パイプラインに直接統合され、無効な位相遷移を明示的にブロックして、確立された信号タイミングと安全制約への絶対的な準拠を保証します。微視的なシミュレーション環境で評価すると、遅延最小化において最先端のベースラインを上回ります。さらに重要なことは、緊急アテンションの重みが確立されたトラフィック エンジニアリングの原則と正確に一致しており、次世代の適応型トラフィック制御システムに監査可能で信頼性があり、導入可能なアーキテクチャを提供していることです。
原文 (English)
Explainable Reinforcement Learning for Adaptive Traffic Signal Control
Reinforcement Learning (RL) has emerged as a powerful paradigm for adaptive traffic signal control. However, in safety-critical infrastructure like traffic control, the opaque, black-box nature of deep RL models poses challenges for transportation agency acceptance, regulatory compliance, operational trust, troubleshooting, and fine-tuning. To bridge this gap between high-performance optimization and human-comprehensible interpretability, this effort introduces a novel, explainable entity centric RL framework for safe and transparent traffic signal control. Rather than processing traffic states through monolithic, flat vectors, the proposed architecture disaggregates real-time intersection observations into distinct, high-dimensional lane entities and phase temporal configurations to inherently preserve the structural topology and geometric configurations of the intersection. Relational dependencies and inter-lane conflicts are dynamically extracted via a dual-stage attention network featuring sequential multi-head cross-attention and self-attention blocks. This design yields a real time affinity matrix that quantifies the direct influence of signal phases on specific approach volumes and queues, providing full visual and analytical interpretability. To ensure strict operational reliability, a deterministic action-masking interface is integrated directly into the Proximal Policy Optimization pipeline, explicitly blocking invalid phase transitions to guarantee absolute compliance with established signal timing and safety constraints. Evaluated in a microscopic simulation environment, outperforms state-of-the-art baselines in delay minimization. More importantly, the emergent attention weights align precisely with established traffic engineering principles, offering an auditable, trust-enabling, and deployable architecture for next-generation adaptive traffic control systems.
会話の時間的ダイナミクスは二者関係におけるうつ病の検出を改善できるか?マルチモダリティの観点からの予備調査
臨床面接からの自動うつ病検出は通常、参加者の発話の意味内容と音響特性をモデル化します。ただし、臨床医と参加者の間の対話のタイミングは、依然として比較的十分にモデル化されていません。私たちは、自己教師ありエンコーダーと融合した主要モダリティとして、会話の時間ダイナミクス、特に二項ターンペアのタイミングを研究します。 DAIC-WOZ データセットで評価し、コンパクトな 24 次元タイミング モジュールを、凍結された WavLM-large および RoBERTa-large ベースライン検出器と比較します。この時間モジュールは、開発セット上で最高の単一モダリティ パフォーマンスを実現します。さらに、凸重み付け後期融合戦略により、全体的なパフォーマンスが開発セットとテスト セットでそれぞれ 0.804 および 0.669 マクロ F1 に向上します。学習された融合は効果的に音響にゼロの重みを割り当て、会話のタイミングが二項うつ病スクリーニングの軽量で解釈可能な補足として機能することを示しています。
原文 (English)
Can Conversational Temporal Dynamics Improve Depression Detection in Dyads? A Preliminary Investigation in Multi-Modality Perspectives
Automatic depression detection from clinical interviews typically models the semantic content and acoustic characteristics of participant speech. However, the interactional timing between the clinician and participant remains comparatively under-modeled. We investigate conversational temporal dynamics, specifically dyadic turn-pair timing, as a primary modality fused with self-supervised encoders. Evaluated on the DAIC-WOZ dataset, we compare a compact 24-dimensional timing module against frozen WavLM-large and RoBERTa-large baseline detectors. This temporal module achieves the highest single-modality performance on the development set. Furthermore, a convex-weighted late fusion strategy improves overall performance to 0.804 and 0.669 macro-F1 on the development and test sets, respectively. The learned fusion effectively assigns zero weight to acoustics, demonstrating that conversational timing serves as a lightweight, interpretable complement for dyadic depression screening.
統合された意思決定プロセスとしてインターリーブされたマルチモーダル推論のブリッジング
統合マルチモーダル モデル (UMM) は、インターリーブされたテキストと画像の推論機能が有望であることを示していますが、強化学習 (RL) によるこのようなマルチターン生成を効果的に最適化することは、未解決の課題のままです。既存のアプローチは、RL をテキスト ステップのみに適用し、画像生成を教師ありサロゲートに委任し、異種モダリティ全体で完全なインターリーブ トラジェクトリを通じてポリシー勾配が伝播するのを防ぎます。このため、UMM の RL の可能性はほとんど活用されていません。論文では、\textbf{BRAID} (\textbf{B}ridging inte\textbf{R}le\textbf{A}ved mult\textbf{I}-modalreasoning as a unified \textbf{D}ecision process) を紹介します。これは、マルチターンのテキスト、画像、テキスト推論を統合マルコフ決定プロセス (MDP) としてキャストする単純なフレームワークで、単一の原則に基づいた RL 目標を介してテキストとビジュアル生成の共同最適化を可能にします。 BRAID は共有軌道レベルの利点を計算し、それをテキスト トークンと画像ノイズ除去パスの両方に一貫して伝播します。各パスはモダリティ ネイティブのポリシー グラディエント メカニズムを通じて最適化されます。長期的な単位の割り当てにさらに取り組むために、BRAID は、推論ユーティリティで各中間画像を採点するビジョン言語モデル (VLM) ジャッジを採用し、重要な視覚分岐での学習を強化するために密なターンレベルのフィードバックを提供します。空間推論と視覚知覚のベンチマークに関する実験では、BRAID がさまざまなベースラインを常に上回っていることが示されており、効果的なマルチモーダル推論にはビジョン思考のガイダンスを備えた統合 MDP 定式化が不可欠であることが確認されています。
原文 (English)
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process
Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learning (RL) remains an open challenge. Existing approaches apply RL exclusively to text steps, relegating image generation to supervised surrogates, preventing policy gradients from propagating through the full interleaved trajectory across heterogeneous modalities. This leaves the potential of RL for UMMs largely untapped. In the paper, we introduce \textbf{BRAID} (\textbf{B}ridging inte\textbf{R}le\textbf{A}ved mult\textbf{I}-modal reasoning as a unified \textbf{D}ecision process), a simple framework that casts multi-turn text-image-text reasoning as a unified Markov decision process (MDP), enabling joint optimization of textual and visual generation via a single, principled RL objective. BRAID computes a shared trajectory-level advantage and propagates it coherently into both text tokens and image denoising paths, each optimized through its modality-native policy gradient mechanism. To further address long-horizon credit assignment, BRAID employs a vision-language model (VLM) judge that scores each intermediate image on its reasoning utility, supplying dense turn-level feedback to sharpen learning at critical visual branches. Experiments on spatial reasoning and visual perception benchmarks show that BRAID consistently outperforms various baselines, confirming that a unified MDP formulation with vision-thinking guidance is essential for effective multi-modal reasoning.
オープンソース創薬エンジンによる折りたたみ、推論、スケーリング
生体分子の相互作用を正確にモデル化することは、生物学と治療法の発見における中心的なボトルネックです。ここでは、スケーラブルな AI 主導の創薬エンジンへのエントリ ポイントとして共フォールディングを使用する、オープンソースの全原子生体分子基礎モデルである Open Drug Discovery Engine (OpenDDE) を紹介します。 OpenDDE は、構造予測を独立したエンドポイントとして扱うのではなく、生体分子複合体にわたる配列、構造、機能の関係をモデル化するための共有構造推論レイヤーとして設計されており、デノボ設計、親和性推定、構造条件付き最適化などの基盤を提供しながら、今日の複雑な構造予測を可能にします。 OpenDDE は、全アトム アーキテクチャ、アトミック潜在推論、推論最適化、大規模データ処理の進歩を統合し、再現可能でオープンにアクセス可能なフレームワーク内で IsoDDE レベルの共折り精度を実現します。また、共折り畳みモデルの 2 つのスケーリング則の方向性を特定し、データ、モデル、推論、トレーニング スケーリングを通じて継続的に改善するための実用的なルートを明らかにします。 OpenDDE は、トレーニング コード、推論パイプライン、チェックポイント、ベンチマークをリリースすることで、フロンティアの生体分子インテリジェンスへのアクセスを民主化し、グローバルなコラボレーションを加速し、分子構造の予測から人間の健康のための治療候補の設計、スコアリング、最適化に移行できる次世代の創薬システムのオープンな基盤を築くことを目指しています。
原文 (English)
Folding, Reasoning, and Scaling with Open-source Drug Discovery Engine
Accurately modeling biomolecular interactions is a central bottleneck in biology and therapeutic discovery. Here, we introduce Open Drug Discovery Engine (OpenDDE), an open-source, all-atom biomolecular foundation model that uses co-folding as the entry point to a scalable AI-driven drug discovery engine. Rather than treating structure prediction as an isolated endpoint, OpenDDE is designed as a shared structural reasoning layer for modeling sequence-structure-function relationships across biomolecular complexes, enabling complex structure prediction today while providing a foundation for de novo design, affinity estimation, structure-conditioned optimization, and more. OpenDDE integrates advances in all-atom architecture, atomic latent reasoning, inference optimization, and large-scale data processing to achieve IsoDDE-level co-folding accuracy within a reproducible and openly accessible framework. We also identify two scaling-law directions for co-folding models, revealing practical routes for continued improvement through data, model, inference, and training scaling. By releasing training code, inference pipelines, checkpoints, and benchmarks, OpenDDE aims to democratize access to frontier biomolecular intelligence, accelerate global collaboration, and lay an open foundation for next-generation drug discovery systems that can move from predicting molecular structures toward designing, scoring, and optimizing therapeutic candidates for human health.
決定論的グラウンドトゥルースを使用した長形式生成における LLM 不確実性の評価
LLM が生成する出力はますます長くなっているため、効果的な不確実性推定では、応答全体を破棄するのではなく、きめの細かいレベルでエラーを特定する必要があります。このような方法は存在しますが、任意の解像度 (世代全体に対するトークン) での不確実性を評価することは困難であり、ラベルの不完全性の影響を非常に受けやすいため、ゼロノイズ ベンチマークが不可欠です。しかし、長い形式の生成ベンチマークは、決定的なグラウンド トゥルースではなく、誤ったラベルに依存する傾向があります。単一の決定論的な長いテキストのグラウンド トゥルースを使用して、手続き的に生成された 6 つのタスクのベンチマークである単一回答アトミック ロングフォーム ターゲット (SALT) を導入します。これにより、外部の判断者なしで、正確性、校正、およびランキングのユニットレベルの評価が可能になります。 SALT を搭載した 50 以上の LLM の分析により、重要な洞察が明らかになります。どの信頼関数が各不確実性の側面を支配しているかを特定し、より粗いラインレベル単位でより明確な分離性が現れた場合でも、信頼度ランキングが原子分解能で大きく崩れることを示します。 SALT はさらに、生成全体を通じて制御されたアトムレベルの介入を可能にし、将来のエラーの 2 つの分離可能な要因を明らかにします。それは、グローバル コンテキストの正確性によって支配される破損したプレフィックスからの伝播と、応答コンテキストの長さの増加による限定的な劣化です。最後に、思考連鎖のプロンプトまたはトレーニングを通じて内面化された推論によって、信頼度ランキングが低下する一方で精度が向上するというトレードオフが導入されることを示します。これらの発見は、信頼性の高いエラーの特定と軽減を必要とするリスククリティカルなアプリケーションに直接影響します。
原文 (English)
Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground Truth
As LLMs generate increasingly long outputs, effective uncertainty estimation must identify errors at fine-grained levels rather than discard entire responses. While such methods exist, evaluating uncertainty at any resolution (token to an entire generation) is challenging and highly sensitive to label imperfections, making zero-noise benchmarks essential; yet, long-form generation benchmarks tend to rely on fallible labels rather than deterministic ground truth. We introduce Single-answer Atomic Long-form Target (SALT), a benchmark of six procedurally generated tasks with single deterministic long textual ground truths, enabling unit-level evaluation of correctness, calibration, and ranking without external judges. Equipped with SALT, our analysis of 50+ LLMs reveals key insights: We identify which confidence functions dominate each uncertainty aspect and show that confidence ranking largely breaks at atomic resolution, even when clearer separability emerges at coarser line-level units. SALT further enables controlled atom-level interventions throughout generation, revealing two separable drivers of future errors: propagation from corrupted prefixes, dominated by global context correctness, and bounded degradation from increasing answer-context length. Finally, we demonstrate that reasoning, via Chain-of-Thought prompting or internalized through training, introduces a trade-off, improving accuracy while degrading confidence ranking. These findings directly impact risk-critical applications requiring reliable error identification and mitigation.
ハーネスを意識した自己進化: 共進化するモデルの重み、ハーネス、タスク ソリューション
自己進化するフレームワークは通常、周囲のハーネスを固定されたものとして扱いながら、タスクの解決策を最適化します。単一のモデルでタスクの解決策を生成したり、マルチターン アクション スペースで選択したハーネス コンポーネントを編集したりできる、エージェント的強化学習フレームワークである Harness-Aware Self-Evolving (HASE) を紹介します。 HASE により、単一の Qwen3-8B モデルが、ハーネス プロポーザーとしてクロード コードを使用する GPT-OSS-120B モデルのテキスト分類パフォーマンスと一致することが可能になります。アルファ因子マイニングでは、HASE は報告されている GPT-OSS-120B ベースラインを上回ります。 HASE はまた、不完全な評価コンポーネントを修復し、サークル パッキング アルゴリズム発見において最先端のパフォーマンスに収束します。これらの結果は、HASE が 1 つの統合エージェント プロセスを通じてハーネスとソリューションを改善することを示しています。
原文 (English)
Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions
Self-evolving frameworks usually optimize task solutions while treating the surrounding harness as fixed. We introduce Harness-Aware Self-Evolving (HASE), an agentic reinforcement-learning framework in which a single model can generate task solutions or edit selected harness components in a multi-turn action space. HASE enables a single Qwen3-8B model to match the text-classification performance of a GPT-OSS-120B model that uses Claude Code as the harness proposer. In alpha factor mining, HASE outperforms the reported GPT-OSS-120B baseline. HASE also repairs imperfect evaluation components and converges to state-of-the-art performance in circle-packing algorithm discovery. These results show that HASE improves the harness and the solution through one unified agentic process.
LLM サービングにおける多目的ルーティングのためのオンライン線形計画法
私たちは、大規模な言語モデルのサービングにおけるオンライン ルーティングの問題を研究します。この問題では、リクエストは順番に到着し、厳しいバッチ サイズと KV キャッシュの制約の下で並列デコード ワーカーにディスパッチされる必要があります。明示的なサービス レベル目標 (SLO) に関連付けられておらず、レイテンシとスループットのトレードオフに対する限定的な制御を提供する、広く使用されているルーティング ヒューリスティックとは異なり、解釈可能な意思決定報酬を備えたオンライン線形計画法としてルーティングを定式化する多目的最適化フレームワークを導入します。当社は、SLO 加重利益がシャドープライスを超える場合にリクエストを許可するオンライン線形計画法に基づいた効率的な入札価格制御ポリシーを適用します。ミリ秒の意思決定要件を満たすために、予測可能なランタイムで進化するデュアル シャドウ 価格をオンラインで追跡する、ウォーム スタートの予測一次更新を開発します。当社はルーターを Vidur シミュレーターに統合し、エンドツーエンドの遅延、最初のトークンまでの時間、スループット、テール パフォーマンスなど、複数の SLO 体制にわたる標準ベースラインよりも大幅な改善を実証しています。私たちの結果から得られる全体像: 科学に基づいたアプローチは、ヒューリスティックに基づいた他のアプローチよりも優れています。
原文 (English)
Online Linear Programming for Multi-Objective Routing in LLM Serving
We study the online routing problem in large language model serving, where requests arrive sequentially and must be dispatched to parallel decode workers under tight batch-size and KV-cache constraints. Unlike widely used routing heuristics that are not tied to explicit service-level objectives (SLOs) and offer limited control over latency-throughput trade-offs, we introduce a multi-objective optimization framework that formulates routing as an online linear programming with interpretable decision rewards. We apply an efficient bid-price control policy based on the online linear programming that admits requests when their SLO-weighted benefit exceeds their shadow prices. To meet millisecond decision requirements, we develop a warm-started, projected first-order updates that track the evolving dual shadow prices online with predictable runtime. We integrate our router into the Vidur simulator and demonstrate substantial improvements over standard baselines across multiple SLO regimes, including end-to-end latency, time-to-first-token, throughput, and tail performance. A big picture from our result: a science-based approach outperforms others based on heuristics.
バングラデシュの子どもにおける虐待関連のトラウマをスクリーニングするための説明可能な AI: ノイズを意識した合成データで評価されたトレーニング不要のマルチモーダル フレームワーク
バングラデシュには人口10万人あたり推定1.17人の精神保健専門家がいるが、児童精神科医は全国でわずか6人しかいない。児童の虐待に関連した心理的外傷を早期にスクリーニングするための、文化的に適応したベンガル語のツールは存在しない。我々は、検証済みのアンケート(SDQ、CPSS)、ベンガル語の物語テキスト、ハウス・ツリー・パーソン(HTP)描画機能、および顔の感情という 4 つのスクリーニング様式を融合した意思決定支援 (診断ではない) フレームワークである ShishuRaksha AI を紹介します。この融合はトレーニング不要で臨床的に重み付けされており、クロスモーダルの注意を使用し、単一モダリティのオーバーライド ルールが含まれています。すべてのリスク スコアは、臨床的に重み付けされた摂動ベースの加法帰属を通じて説明され、2013 年児童法に基づく各国の児童保護サービス (OCC、DSS、NMHH) への紹介ルーティングを含むバイリンガル (バングラ語/英語) レポートとしてレンダリングされます。現段階では、虐待を受けた子供の臨床データセットは倫理的に収集できないため、ノイズを意識した合成ベンチマーク (ケース 500 件、陽性 116 件) を導入します。 [23.2%]、4 つの意図的なノイズ レイヤー、文献に基づいた HTP 事前分布) を使用し、5 重層化相互検証の下で融合デザインのツリー アンサンブル サロゲート (フェイシャル チャネルは除外) を評価します。融合モデルは、アブレーション、動作点、サブグループ、およびキャリブレーション解析を使用した SDQ のみのベースラインの AUC 0.756 [0.705-0.803] に対して、0.874 [0.834-0.908] の AUC に達します。私たちは、合成のみのデータ、ホールドアウトされたセットなし、テキスト特徴の循環性、都市部と農村部のサブグループのギャップなど、すべての制限をオープンに述べます。この研究は、実現可能性の研究であり、リソースが少ない状況で倫理的に展開可能な児童保護スクリーニングに向けた設計への貢献です。
原文 (English)
Explainable AI for Screening Abuse-Related Trauma in Bangladeshi Children: A Training-Free Multimodal Framework Evaluated on Noise-Aware Synthetic Data
Bangladesh has an estimated 1.17 mental-health professionals per 100,000 population and only six child psychiatrists nationwide. No Bengali-language, culturally adapted tool exists for early screening of abuse-related psychological trauma in children. We present ShishuRaksha AI, a decision-support (not diagnostic) framework that fuses four screening modalities: validated questionnaires (SDQ, CPSS), Bengali narrative text, House-Tree-Person (HTP) drawing features, and facial affect. The fusion is training-free and clinically weighted, uses cross-modal attention, and includes a single-modality override rule. Every risk score is explained through clinically weighted, perturbation-based additive attribution and rendered as a bilingual (Bangla/English) report with referral routing to national child-protection services (OCC, DSS, NMHH) under the Children Act 2013. No clinical dataset of abused children can be collected ethically at this stage, so we introduce a noise-aware synthetic benchmark (500 cases, 116 positive [23.2%], four deliberate noise layers, literature-grounded HTP priors) and evaluate tree-ensemble surrogates of the fusion design (facial channel excluded) under 5-fold stratified cross-validation. The fused model reaches an AUC of 0.874 [0.834-0.908], against 0.756 [0.705-0.803] for an SDQ-only baseline, with ablation, operating-point, subgroup, and calibration analyses. We state all limitations openly, including synthetic-only data, no held-out set, text-feature circularity, and an urban-rural subgroup gap. This work is a feasibility study and a design contribution toward ethically deployable child-protection screening in low-resource settings.
私たちに何が残されているのでしょうか? AIによる研究の劣化に反対する第2回奨学金
私たちは、生成型 AI は学術的判断を形成し、学術的信頼を構築する実践そのものを侵食することにより、研究の質を低下させる可能性があると主張します。知識の生成と検証の構成条件として、これらの実践は、AI が効果的にシミュレートする研究の最終成果物に還元することはできません。したがって、研究者が調査の中心的なタスクを大規模言語モデルなどのシステムに委任すると、これらの実践の実施を中止し、それらが提供する形成へのアクセスを失う可能性があります。 AI によって生成された個々の研究成果は改善されているように見えても、その背後にある研究者は成長していない可能性があります。このリスクに対して、人間が AI 出力のプロンプターまたは品質チェッカーとして常に情報を把握し続けるだけでは、知的形成の場として研究を維持するには不十分です。代わりに必要なのは、しばしば摩擦や学術コミュニティへの参加を通じて、判断力が徐々に形成される生きた実践としての研究への新たな取り組みである。私たちがそれを擁護するのは、それが暗黙知、個人的な取り組み、社交化、深い読書という 4 つの情報源と自動化できない研究の根拠に基づいているからです。この実践は、私たちが第二の学問と呼ぶものを制定し、それによって私たちは、生成 AI ができることとできないことについての重要な経験から選ばれた学術的工芸品の再利用を理解します。委任できない、委任すべきではないものは、研究コミュニティが評価し、それに答えなければならないものになります。 This is what is left for us.
原文 (English)
What is Left for Us? Second Scholarship Against the Degradation of Research by AI
We argue that generative AI can degrade research by eroding the very practices through which scholarly judgement is formed and academic trust is built. As constitutive conditions for the production and validation of knowledge, these practices cannot be reduced to the final outputs of research, which is what AI so effectively simulate. Accordingly, when researchers delegate central tasks of inquiry to systems like Large Language Models, they may stop enacting these practices and, with them, lose access to the formation they provide. An individual research output generated by AI may even appear improved but the researcher behind it fails to develop. Against this risk, merely keeping humans in the loop as prompters or quality checkers of AI outputs is insufficient to preserve research as a site of intellectual formation. What is needed instead is a renewed commitment to research as a lived practice in which judgement is formed gradually, often through frictions, and participation in a scholarly community. We defend it because it rests on four sources and warrants of research that cannot be automated: tacit knowledge, personal commitment, socialisation, and deep reading. This practice enacts what we call second scholarship, by which we understand the reappropriation of scholarly craft, chosen out of a critical experience of what generative AI can and cannot do. What cannot and should not be delegated becomes what research communities must value and answer for. This is what is left for us.
PLACEMEM: 生涯エージェント向けのコンピューティング対応メモリ プレーンに向けて
生涯エージェントには、より大きなコンテキスト ウィンドウとより優れた検索だけでは不十分です。サービングスタックに毎ターン同じ履歴の再計算を強制したり、古いランタイム状態をサイレントに再利用したりすることなく、永続化、進化、修正できるメモリが必要です。 PLACEMEM を、実行可能なコントロール プレーン プロトタイプによってインスタンス化される、生涯にわたるエージェント メモリ上のシステムの位置付けとして紹介します。中心的な主張は、エージェント メモリは、セマンティクス、来歴、有効性、および再利用可能な実行時状態を 1 つの修正を意識したアイデンティティの下で統合する、バージョン管理されたカプセルとして表現されるべきであるということです。現在のプロトタイプでは、カプセルは、プロンプトレベルのテキスト取得、KV 対応ルーティング、およびライブ ストリーミング バックエンドを介したカスケード無効化を駆動します。将来のレイヤー フロンティア リプレイは、主張されるエンジン機能ではなく、より深い統合アジェンダとして意図的に組み立てられています。永続的なカプセル状態、同時実行安全な無効化、OpenAI 互換のルーティング サイドカー、型付きメタデータ コントラクト、ライブのファースト トークン レイテンシ、再利用、および修正後の動作を測定するベンチマーク ハーネスを備えた vLLM ファースト プロトタイプについて説明します。その結果、今日の修正を意識したコントロール プレーンの動作を示す実行可能なアーティファクトと、将来の生涯エージェント システムにおけるリプレイを意識したサービング統合の具体的なロードマップの両方が得られます。
原文 (English)
PLACEMEM: Toward a Compute-Aware Memory Plane for Lifelong Agents
Lifelong agents need more than larger context windows and better retrieval. They need memories that can persist, evolve, and be corrected without forcing the serving stack to recompute the same history on every turn or silently reuse stale runtime state. We present PLACEMEM as a systems position on lifelong-agent memory, instantiated by an executable control-plane prototype. The central claim is that agent memory should be represented as versioned capsules that unify semantics, provenance, validity, and reusable runtime state under one correction-aware identity. In the current prototype, capsules drive prompt-level text retrieval, KV-aware routing, and cascading invalidation over live streamed backends; prospective layer-frontier replay is intentionally framed as a deeper integration agenda rather than a claimed engine feature. We describe a vLLM-first prototype with persistent capsule state, concurrency-safe invalidation, an OpenAI-compatible routing sidecar, a typed metadata contract, and a benchmark harness that measures live first-token latency, reuse, and post-correction behavior. The result is both an executable artifact that demonstrates correction-aware control-plane behavior today and a concrete roadmap for replay-aware serving integration in future lifelong-agent systems.
予見: 神経記号的原始プログラミングからの検証可能な推論
現在のエージェント ワークフローでは、通常、ユーザー リクエストを、正しく解決されたパラメーターを含む一連のツール呼び出しに分解することが含まれます。その結果は、言語モデルのコンテキスト ウィンドウで推論トレースを通じて処理されます。このような推論を改善するための一般的な方法は、テスト時間のスケーリングです。これは、長い思考連鎖にわたって検索するようにモデルをトレーニングします。しかし、結果として得られる機能はモデルの重みに絡み合っており、段階的に検証できず、推論時にコストがかかります。我々は、神経記号的推論システムである Forethought を紹介します。これは、推論を明示的で検証可能なプログラムとして扱い、ドメイン固有の言語を通じて構成される記号プリミティブと神経プリミティブのライブラリから構築されます。その結果、モデルの動作を具体的に表現した推論プログラムが作成され、展開前に検査および変更できます。ツール呼び出し実行カーネルとしてインスタンス化され、5 つのベンチマークにわたって評価された Forethought は、基本モデルの精度を相対的に約 30% 向上させ、バニラ プロンプト、強化学習足場、およびプロンプト進化手法を上回るパフォーマンスを示し、小規模モデルがフロンティア モデルの機能と同等またはそれを超えることを可能にします。直接比較すると、Forethought で強化された非推論モデルは専用の推論モデルと競合しますが、必要なトレーニング後の投資はおよそ 3 桁少なく、モデルに依存せず監査可能です。
原文 (English)
Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming
Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's context window. The prevailing route to improve such reasoning is test-time scaling, which trains models to search over long chains of thought; but the resulting capability is entangled in model weights, is not verifiable step-by-step, and is costly at inference. We present Forethought, a neurosymbolic reasoning system that instead treats reasoning as an explicit, verifiable program, that builds from a library of symbolic and neural primitives which are composed through a domain-specific language. The result are reasoning programs, which are concrete representations of the model's work, and as such can be inspected and modified before deployment. Instantiated as a tool-calling execution kernel and evaluated across five benchmarks, Forethought improves base-model accuracy by about 30% relative and outperforms vanilla prompting, reinforcement learning scaffolds, and prompt-evolution methods, enabling small models to match or exceed frontier models capabilities. In a direct comparison, a non-reasoning model augmented with Forethought competes with a dedicated reasoning model while requiring roughly three orders of magnitude less post-training investment, and remains model-agnostic and auditable.
言語モデルは、検索を制御することでシンボリック方程式の発見をガイドします
科学的方程式の発見では、広範な領域の事前分布と厳密な数値テストを組み合わせる必要があります。シンボリック回帰は数値的基礎を提供しますが、組み合わせ探索空間に直面します。一方、多くの言語モデル システムは、モデルに式を直接提案または選択するように要求します。私たちは異なる分業を試しています。言語モデルが数式作成者、候補決定者、または検索コントローラーとして機能する役割仕様を、エンドツーエンドの言語モデルと純粋な数値ベースラインと並べて比較します。ここで提案するコントローラー設定では、LLM-PySR として実装され、言語モデルは変数、演算子、変換、検索深さを指定します。シンボリック回帰は式を列挙して近似します。そして決定論的な指標が保持を管理します。 74 の AI ファインマン方程式と 7 つの複雑な数式回復タスクにわたって、検索制御は、観察された精度、複雑さ、安定性、コストの最強のバランスを達成しました。 LLM-PySR は、独立したバッテリー データセット上で、初期電圧曲線変位とサイクル寿命の間のコンパクトな区分線形関係を特定しました。この結果は、言語モデルはどの方程式が生き残るかを決定するのではなく、仮説の探索を形作るべきであることを示唆しています。
原文 (English)
Language models guide symbolic equation discovery by controlling search
Scientific equation discovery must combine broad domain priors with strict numerical testing. Symbolic regression supplies numerical grounding but faces a combinatorial search space, whereas many language-model systems ask the model to propose or select formulas directly. We test a different division of labour. We compare role specifications in which the language model acts as equation author, candidate decider or search controller, alongside end-to-end language-model and purely numerical baselines. In the controller setting we propose here, implemented as LLM-PySR, language models specify variables, operators, transformations and search depth; symbolic regression enumerates and fits expressions; and deterministic metrics govern retention. Across 74 AI-Feynman equations and seven complex formula-recovery tasks, search control achieved the strongest observed balance of accuracy, complexity, stability and cost. On an independent battery dataset, LLM-PySR identified a compact piecewise-linear relation between early voltage-curve displacement and cycle life. The results suggest that language models should shape hypothesis exploration rather than decide which equations survive.
資本市場における不審な取引パターンを特定するためのクラスタリングベースのフレームワーク
市場操作とは、手っ取り早く利益を上げるために株価を操作する疑わしい行為であり、取引プラットフォームに対する信頼を著しく低下させます。この問題に対処するために、K-Means++ クラスタリングで始まる教師なし不正検出ツールキットを実装しました。 2012 年から 2024 年までの約 100 万件の金融取引のデータセットが使用されます。不正な取引を特定し、市場慣行のヒューリスティックしきい値を使用して分類するために、この調査ではクラスタリング ベースのパイプラインが提案されています。この方法では、取引の 2.02% が疑わしいものとして強調表示され、そのうち 51.10% は明らかにスプーフィングを示し、0.10% はポンプ アンド ダンプを示し、0.55% はインサイダー取引を示し、1.43% は偽のブレイクアウトを示し、46.83% は機密扱いではありません。グラウンド トゥルースの欠如にもかかわらず、モデルのパフォーマンスは 0.561 のシルエット スコアによって確認されます。
原文 (English)
A Clustering-Based Framework for Identifying Suspicious Trading Patterns in Capital Market
Market manipulation is the dubious practice of manipulating stock prices in order to make a quick profit, which truly degrades confidence on trading platforms. We implemented an unsupervised fraud-detection toolkit that begins with K-Means++ clustering to address this issue. A dataset of roughly one million financial transactions from 2012 to 2024 is used. In order to identify fraudulent trades and categorize them using market practice heuristic thresholds, the study suggests a clustering-based pipeline. The method highlights 2.02% of trades as suspicious where 51.10% clearly indicate spoofing, 0.10% indicate pump and dump, 0.55% indicate insider trading, 1.43% indicate a fake breakout, and 46.83% are unclassified. Despite the lack of ground truth, the model's performance is confirmed by a Silhouette Score of 0.561.
Agentic IoT: エージェントのインターネットに向けたアーキテクチャ、アプリケーション、および課題
AI をモノのインターネット (AIoT) システムに統合することで、システムは受動的なデータ収集インフラストラクチャから、異常検出、予知保全、分類、予測、最適化が可能なインテリジェント システムへと徐々に変化してきました。ただし、既存のソリューションのほとんどは依然としてセンサー データから推測するタスク固有のモデルに依存しています。したがって、リアルタイム推論、適応計画、自律調整、学習、ツールの使用、状況に応じた意思決定などのシステム全体の機能は依然として制限されています。この論文では、自律型 AI エージェントの知覚、推論、計画、学習、およびアクション機能をサイバー物理システムと統合する、次世代のコグニティブ IoT パラダイムとしての Agentic IoT を検証します。 Agentic IoT は、IoT をデータ中心のセンシングおよび推論インフラストラクチャから、デバイス/エッジ-フォグ-クラウドの連続体全体で動作する分散型コグニティブ エージェント エコシステムに変換することを目的としています。この論文はまず、この移行をパラダイム シフトとして根拠付けし、AIoT、エッジ インテリジェンス、マルチ エージェント システム、およびエージェントのインターネットとの関連で Agentic IoT を位置づけます。次に、現在の研究を体系的にレビューし、全体的なアーキテクチャのフレームワークを提示し、ドメイン固有のアプリケーションの可能性について議論し、将来の研究の方向性とともに主要な技術的、運用的、および研究上の課題を特定します。
原文 (English)
Agentic IoT: Architectures, Applications, and Challenges Toward the Internet of Agents
The integration of AI into Internet of Things (AIoT) systems has gradually transformed them from passive data collection infrastructures into intelligent systems capable of anomaly detection, predictive maintenance, classification, forecasting, and optimization. However, most existing solutions still rely on task-specific models that infer from sensor data; thus, system-wide capabilities such as real-time reasoning, adaptive planning, autonomous coordination, learning, tool use, and contextual decision-making remain limited. This paper examines Agentic IoT as a next-generation cognitive IoT paradigm that integrates the perception, reasoning, planning, learning, and action capabilities of autonomous AI agents with cyber-physical systems. Agentic IoT aims to transform IoT from data-centric sensing and inference infrastructures into distributed cognitive agent ecosystems operating across the device/edge-fog-cloud continuum. The paper first grounds this transition as a paradigm shift and positions Agentic IoT in relation to AIoT, edge intelligence, multi-agent systems, and the Internet of Agents. It then systematically reviews current studies, presents a holistic architectural framework, discusses domain-specific application potential, and identifies key technical, operational, and research challenges together with future research directions.
活性化ジオメトリによる教師なし特徴マイニング
解釈可能性メソッドは、大規模言語モデル (LLM) 内で表現される機能を明らかにすることを目的としています。既存の手法の多くは、人間の偏見を反映している可能性がある人間が定義した概念のラベル付き例から始まり、次にその概念がモデル内で (たとえば、活性化空間や他の分解手法などで) どのように表現されるかを特定します。 \emph{Mining via Activation Geometry} (MAG) を導入します。これは、同じ自然言語命令 $Q$ をすべての入力 $p$ の前に付加することで、モデルのアクティベーションから推論特徴を抽出するための単純な教師なしフレームワークです。ここで、$Q$ は、「このオブジェクトは砂漠で見つけることができますか?」または「このプロンプトは悪意がありますか?」などの関心のある推論特徴を定義します。$m(Q \mid p) を使用して、命令がモデルの内部表現をどのように変更するかを測定します。単一の読み出しポイントでの m(p)$。 8 つの異なる MAG を調査します。抽出された推論特徴は、モデル自身の世界の理解と判断を予測し、単一のアクティベーション方向に近似できます。一部の特徴はより線形に表現され、一部の特徴はより線形に表現されないことがわかりました。ベクトル ステアリングであるこの線形表現は、推論特徴を注入することによるアクティベーション ステアリングを通じて LLM の決定を変更できる可能性があります。最後に、同じ方法を使用してプロンプトインジェクション分類器プローブに最適なトレーニング データセットを選択します。通常のアクティベーション間の類似性は下流のパフォーマンスにほとんど無関係ですが、RFD ベースの類似性は $94.7\%$ Top-1 と $100\%$ Top-2 の精度を達成します。
原文 (English)
Unsupervised Features Mining via Activation Geometry
Interpretability methods aim to reveal the features represented inside large language models (LLMs). Many existing methods begin with labeled examples of a human-defined concept that may reflect human biases, and then identify how that concept is represented within the model, for example in its activation space or through other decomposition methods. We introduce \emph{Mining via Activation Geometry} (MAG), a simple unsupervised framework for extracting reasoning features from model activations by prepending the same natural-language instruction $Q$ to every input $p$, where $Q$ defines the reasoning feature of interest, such as ``Can this object be found in the desert?'' or ``Is this prompt malicious?'' We measure how the instruction changes the model's internal representation using $m(Q \mid p) - m(p)$ at a single readout point. We explore eight different MAGs. The extracted reasoning features predict the models' own world understanding and judgment, can be approximated into a single activation direction, we found that some features are more linearly represented and some less, this linear representation, which is vector steering, can change the LLMs' decisions through activation steering by injecting reasoning features. Finally, we use the same method to select the best training datasets for prompt-injection classifier probes: while similarity between ordinary activations is almost unrelated to downstream performance, RFD-based similarity achieves $94.7\%$ Top-1 and $100\%$ Top-2 accuracy.
薬剤制御のための生物学的モチーフ
大規模言語モデル (LLM) が受動的ジェネレーターから自律エージェントに移行したことで、信頼性、セキュリティ、および状態管理に重大な課題が生じました。現在のエージェント アーキテクチャはアドホックに構築されることが多く、幻覚カスケード、無限ループ、プロンプト インジェクション攻撃が発生する傾向があります。この論文は、文字どおりの生物学的メカニズムではなく、型指定された界面と調整構造のレベルで比較が行われる場合には、システム生物学で長年研究されてきた制御モチーフを使用して、これらの故障モードの多くを分析できると主張しています。私たちは、多項式関手と配線図を使用して、遺伝子制御ネットワークとエージェントソフトウェアシステムの間の型付きインターフェイス対応を開発します。 5 つの生物学的モチーフは、構成可能なソフトウェア設計パターンにマッピングされています。ノイズ抑制のためのコヒーレント フィードフォワード ループ、多層セキュリティのための適応免疫、リソース ガバナンスのためのミトコンドリア シグナリング、神経記号統合のための内部共生、空間的に変化する調整のためのモルフォゲン拡散です。認識トポロジ層は、配線図の観測構造からクリプキ形式の知識演算子を導き出し、マルチエージェント スケーリングに関する 4 つの予測定理を証明します。主要な貢献は次のとおりです。(1) Agentic Operad、フィードフォワード トポロジの証明可能なエラー抑制限界を備えたエージェント合成用の型付き構文。 (2) 定性的予測が公開されているマルチエージェント ベンチマークと一致する 4 つの定理 (誤差増幅、逐次ペナルティ、並列加速、ツール密度スケーリング) を備えた認識論的トポロジー。 (3) 自律学習フレームワークと経験的文献からの収束プロキシに基づいた、構造から開発までの 6 層の進行。 1,813 のテストと 116 の例を含むリファレンス実装は、実際的な実現可能性を示しています。
原文 (English)
Biological Motifs for Agentic Control
The transition of Large Language Models (LLMs) from passive generators to autonomous agents has introduced significant challenges in reliability, security, and state management. Current agentic architectures are often constructed ad-hoc, prone to hallucination cascades, infinite loops, and prompt injection attacks. This paper argues that many of these failure modes can be analyzed using control motifs long studied in systems biology, provided the comparison is made at the level of typed interfaces and coordination structure rather than literal biological mechanism. We develop a typed interface correspondence between Gene Regulatory Networks and agentic software systems using polynomial functors and wiring diagrams. Five biological motifs are mapped to composable software design patterns: Coherent Feed-Forward Loops for noise suppression, Adaptive Immunity for layered security, Mitochondrial Signaling for resource governance, Endosymbiosis for neuro-symbolic integration, and Morphogen Diffusion for spatially varying coordination. An epistemic topology layer derives Kripke-style knowledge operators from the wiring diagram's observation structure and proves four predictive theorems for multi-agent scaling. The core contributions are: (1) the Agentic Operad, a typed syntax for agent composition with provable error suppression bounds for feed-forward topologies; (2) an epistemic topology with four theorems (error amplification, sequential penalty, parallel acceleration, and tool density scaling) whose qualitative predictions are consistent with published multi-agent benchmarks; and (3) a six-layer progression from structure through development, grounded in autonomous learning frameworks and convergence proxies from the empirical literature. A reference implementation with 1,813 tests and 116 examples illustrates practical feasibility.
進歩と信頼性を重視したエージェント強化学習のためのグループ ポリシーの最適化
グループベースの強化学習 (RL) は、長期にわたる対話型タスクにおける大規模な言語モデル エージェントを改善するための効果的なパラダイムとなっています。軌道レベルの最適化よりもきめ細かいポリシー更新を取得するために、最近の研究は、中間ステップがグループ化され、ロールアウト バッチ内で比較されるステップ レベルのグループベースの RL に移行しています。ただし、ステップレベルの利点の推定は、グループの形成方法に影響を受けます。広範な状態キーによるグループ化によりカバレッジは向上しますが、異なる履歴の下で実行されたアクションを比較する可能性があります。一方、履歴の一貫性を強制すると、断片化されたグループやピア比較信号の欠落を犠牲にして、より公平な比較が得られます。この論文では、コンテキスト一貫したステップレベル学習のための学習済み批判のない手法である ProGPO (進歩性と信頼性指向のグループ ポリシー最適化) を提案します。 ProGPO は、正確なプレフィックス アクションの比較を維持し、ロールアウト ベースの状態ポテンシャルから得られる遷移クレジットでスパース ピア比較を補完します。これらの可能性を確実に推定するために、ProGPO は意味拡張と履歴深度にわたる逆分散融合を組み合わせます。 Qwen2.5-1.5B-Instruct を使用して、ALFWorld と WebShop という 2 つの困難なエージェント タスクで ProGPO を評価します。結果は、ProGPO が同等の計算オーバーヘッドの下で一致するエージェント RL ベースラインよりも改善することを示し、追加の Qwen2.5-3B-Instruct 実験により、提案された方法のスケーラビリティをさらにテストします。
原文 (English)
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning
Group-based reinforcement learning (RL) has become an effective paradigm for improving large language model agents on long-horizon interactive tasks. To obtain finer-grained policy updates than trajectory-level optimization, recent work has moved toward step-level group-based RL, where intermediate steps are grouped and compared within a rollout batch. However, step-level advantage estimation is sensitive to how groups are formed: grouping by broad state keys improves coverage but may compare actions taken under different histories, while enforcing historical consistency yields fairer comparisons at the cost of fragmented groups and missing peer-comparison signal. In this paper, we propose ProGPO (Progress- and Reliability-Oriented Group Policy Optimization), a learned-critic-free method for context-consistent step-level learning. ProGPO keeps exact-prefix action comparison, and complements sparse peer comparisons with transition credit derived from rollout-based state potentials. To estimate these potentials reliably, ProGPO combines semantic expansion with inverse-variance fusion across history depths. We evaluate ProGPO on two challenging agentic tasks, ALFWorld and WebShop, with Qwen2.5-1.5B-Instruct. Results show that ProGPO improves over matched agentic RL baselines under comparable computational overhead, and additional Qwen2.5-3B-Instruct experiments further test the scalability of the proposed method.
法的判決の予測における近道学習: 英国雇用裁判所からの経験的証拠
現在の法的判決予測 (LJP) は、事後の司法資料に依存しているため制約があり、モデルが真の予測ではなく遡及的な分類を実行する可能性が高くなります。この論文では、英国雇用裁判所 (UKET) の判決における請求レベルの結果予測を研究することにより、この文脈における近道学習を実証的に調査しています。 33,158 件の個々のクレームのコーパスを使用して、解釈可能な TF-IDF ベースの分類器からブラックボックス LLM に至るまでのモデルを評価し、クレーム テキストと LLM で抽出された症例概要から結果を予測します。ヘッドラインの予測パフォーマンス数値は強力であるように見えますが、事後の司法テキストに基づいて訓練された LJP システムのそのようなパフォーマンスは、ソース資料の遡及的な性質によってもたらされる可能性があることを示しています。漏れに関する人間の判断によってテスト データを層別化すると、結果を明らかにする手がかりが物語に埋め込まれている場合、パフォーマンスが向上することが明らかになります。さらに、漏洩として特定された特徴のわずか 4% でトレーニングされたモデルは、人間の専門家を上回る高いパフォーマンスを実現します。これらの発見は、LJP のパフォーマンスが言語アーチファクトによって誇張される可能性があるという懸念を裏付けています。しかし、この脆弱性は研究課題にとって致命的なものではありません。代わりに、事後の判断は汚染された可能性のあるテキストとして扱われる可能性があり、積極的な監査が必要になります。リーク フィーチャをマスキングした後にモデルを再トレーニングしても、Macro-F1 は無視できるほど減少します。したがって、モデルは利用可能な場合にはショートカットを利用しますが、これらのアーティファクトが除去されても有用な予測信号を抽出する能力は維持されます。
原文 (English)
Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal
Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting. This paper empirically investigates shortcut learning in this context by studying claim-level outcome prediction in UK Employment Tribunal (UKET) decisions. Using a corpus of 33,158 individual claims, we predict outcomes from claim texts and LLM-extracted case summaries, evaluating models ranging from interpretable TF-IDF-based classifiers to black-box LLMs. While headline predictive performance figures appear strong, we demonstrate that such performance in LJP systems trained on post-hoc judicial text can be driven by the retrospective nature of the source material. Stratifying the test data by human judgments of leakage reveals that performance increases where outcome-revealing cues are embedded in the narrative. Moreover, a model trained on just the 4% of features identified as leakage achieves high performance, outperforming human experts. These findings substantiate concerns that LJP performance may be exaggerated by linguistic artefacts. Yet this vulnerability is not fatal to the research agenda. Instead, post-hoc judgments might be treated as potentially contaminated texts, requiring active auditing. Retraining models after masking leakage features results in only a negligible reduction in Macro-F1. Hence, while models will opportunistically exploit shortcuts when available, they remain capable of extracting useful predictive signals when these artefacts are removed.
Agentic SABRE: 適応型ランサムウェア検出のための不確実性を認識したニューロシンボリック マルチエージェント フレームワーク
ランサムウェアは、概念のドリフト、回避、および動作の多態性の下で静的シグネチャとモノリシック分類子が一般化できない、複雑で適応性があり、動きの速い敵のカテゴリに進化しました。このペーパーでは、適応型ランサムウェア検出のための不確実性を認識した神経記号的なマルチエージェント フレームワークである Agentic SABRE (ランサムウェア評価のためのセマンティック動作仲裁) を紹介します。 SABRE は、意味論的な表現ベースの証拠と行動的なタイムウィンドウ法医学テレメトリを融合し、モンテカルロ ドロップアウト推論を採用して各エージェントの認識論的不確実性を定量化します。リスク スコアと不確実性バジェットという 2 つの解釈可能なしきい値を使用して、リスクと不確実性を認識したトリアージを実行するデシジョン層オーケストレーターを紹介します。信頼性が高くリスクの高いサンプルは自動的に封じ込められますが、不確実なケースや境界線に近いケースは人間のアナリストにエスカレーションされ、自律的な対応とアナリストの監視の間に柔軟な計算上の契約が確立されます。監査可能性と信頼性をサポートするために、SABRE は勾配顕著性、順列重要度、反事実分析などの事後説明可能メカニズムを統合し、エージェントの決定のローカルおよびグローバルの両方の解釈を可能にします。 RDset と RanSMAP の広範な評価により、Agentic SABRE は飽和したセマンティック データセット上で AUC が 1.0 に等しい完全な識別を維持しながら、弱い行動シグナルの下での堅牢性が向上することが実証されました。調整された予測の不確実性を維持しながら、等しい再現率で誤ったエスカレーションを最大 4.9% 相対的に削減します。反事実分析はさらに、意味論的および行動的決定が有限の摂動コストで覆される可能性があることを示し、安定した解釈可能な決定境界を示します。
原文 (English)
Agentic SABRE: An Uncertainty-Aware Neuro-Symbolic Multi-Agent Framework for Adaptive Ransomware Detection
Ransomware has evolved into a complex, adaptive, and fast-moving adversary category in which static signatures and monolithic classifiers fail to generalise under concept drift, evasion, and behavioural polymorphism. In this paper, we present Agentic SABRE (Semantic-Behavioural Arbitration for Ransomware Evaluation), an uncertainty-aware, neuro-symbolic, multi-agent framework for adaptive ransomware detection. SABRE fuses semantic, representation-based evidence with behavioural, time-window forensic telemetry and employs Monte Carlo Dropout inference to quantify epistemic uncertainty for each agent. We introduce a decision-layer orchestrator that performs risk- and uncertainty-aware triage using two interpretable thresholds: a risk score and an uncertainty budget. High-confidence, high-risk samples are automatically contained, while uncertain or borderline cases are escalated to human analysts, establishing a flexible computational contract between autonomous response and analyst oversight. To support auditability and trust, SABRE integrates post-hoc explainability mechanisms, including gradient saliency, permutation importance, and counterfactual analysis, enabling both local and global interpretation of agent decisions. Extensive evaluation on RDset and RanSMAP demonstrates that Agentic SABRE preserves perfect discrimination on saturated semantic datasets, with AUC equal to 1.0, while improving robustness under weak behavioural signals. It achieves up to a 4.9 percent relative reduction in false escalations at equal recall while maintaining calibrated predictive uncertainty. Counterfactual analysis further shows that semantic and behavioural decisions can be reversed with bounded perturbation cost, indicating stable and interpretable decision boundaries.
HAS ベンチ: 構成可能な人間の参加の下での LLM ベースのヒューマン エージェント システムの評価
大規模な言語モデルは、人間が受動的タスク提供者ではなく能動的協力者となる環境で運用されることが増えています。 HAS フレームワークは、人間と LLM を利用したエージェントを明示的な役割、権限、通信パス、およびアクション権限を持つ第一級の参加者として表すグラフベースのフレームワークです。このフレームワークに基づいて、HAS-Bench は、政府機関レベル、対話チャネル、ペルソナ ポリシー全体にわたる構成可能な人間の参加の下でヒューマン エージェント システムを評価します。このベンチマークは、タスクの結果と、明確化の品質、フィードバックの利用、制御の調整、安全性、イニシアチブ、対話コストなどのプロセスレベルのコラボレーション行動の両方を測定します。 6 つのドメインにわたる実験では、人間の参加によってタスクの完了と失敗からの回復が大幅に向上する可能性がありますが、その効果は人間の入力がいつ、どのように、誰によって実行されるかによって異なります。
原文 (English)
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framework, a graph-based framework that represents humans and LLM-powered agents as first-class participants with explicit roles, permissions, communication paths, and action authority. Building on this framework, HAS-Bench evaluates Human-Agent Systems under configurable human participation across agency levels, interaction channels, and persona policies. The benchmark measures both task outcomes and process-level collaboration behavior, including clarification quality, feedback utilization, control calibration, safety, initiative, and interaction cost. Experiments across six domains show that human participation can substantially improve task completion and failure recovery, but the gains depend on when, how, and by whom human input is exercised.
GUI エージェントは自分の目を信じますか?ピクセル対構造に対する状態信念依存性の診断
マルチモーダル GUI エージェントは、スクリーンショットのレンダリングされたピクセルと、DOM やアクセシビリティ ツリーなどのシリアル化された構造という 2 つの冗長チャネルを通じてインターフェイスを読み取ります。エージェントは行動する前に、現在のインターフェイス状態について信念を形成しますが、既存のベンチマークはタスクの成功、要素のグラウンディング、または攻撃耐性をスコアリングし、その信念がピクセルから引き出されたものであるかどうかは質問しません。私たちは視覚的な状態依存性、つまり状態信念のピクセル、構造、または事前分布への帰属を形式化し、310 の実際の Web、モバイル、およびデスクトップのプローブに対するペアの単一チャネル介入でそれを測定します。すべてのプローブは、モデルによって生成された項目やモデルによる判断を行わず、決定論的な強制選択によってスコア付けされます。私たちの中心的な指標は、モデルが正しく認識し、矛盾している構造に向けて解決するプローブの割合である知覚融合ギャップです。 3 社のベンダーの 5 つのモデルにわたって、テキスト状態の信念は構造に依存しますが、画像のみの精度は天井付近に留まり、知覚と融合のギャップはすべてのモデルでプラスです。対照的に、非テキスト ID は主にピクセルに限定されたままになります。この置換はシリアル化されたテキストとインデックス付きアクション チャネルに固有であり、調整アクション エージェントはほとんど影響を受けません。テキストの競合の場合、ホワイト ボックス アブレーションにより、コピーされた単一の構造値への影響が追跡され、2 つの実際の環境では、競合により誤ったアクションと実際のタスクの失敗が引き起こされます。したがって、視覚的状態依存性は、エージェント状態の信念が視覚的に根拠があり、それが露呈するエラーがアクションに伝播するかどうかの測定可能な診断を提供します。
原文 (English)
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
Multimodal GUI agents read an interface through two redundant channels: the rendered pixels of a screenshot and a serialized structure such as a DOM or accessibility tree. Before acting, an agent forms a belief about the current interface state, but existing benchmarks score task success, element grounding, or attack resistance and do not ask whether that belief is drawn from the pixels. We formalize visual state reliance, the attribution of a state belief to pixels, structure, or priors, and measure it with paired single-channel interventions over 310 real web, mobile, and desktop probes. Every probe is scored by deterministic forced choice, with no model-generated item and no model judge. Our central metric is the Perception-Fusion Gap, the fraction of probes a model perceives correctly yet resolves toward structure under conflict. Across five models from three vendors, textual state beliefs defer to structure while image-only accuracy stays near ceiling, and Perception-Fusion Gap is positive for every model; non-text identity, by contrast, stays largely pixel-bound. The substitution is specific to the serialized-text and indexed-action channel, and coordinate-action agents are largely immune. For textual conflicts, a white-box ablation traces the effect to a single copied structural value, and in two live environments the conflict drives wrong actions and real task failure. Visual state reliance therefore gives a measurable diagnostic of whether agent state beliefs are visually grounded, and the errors it exposes propagate to actions.
深層学習と機械学習を使用した FPS ゲームでのサーバーサイドのアンチチートによる Aimbot 検出
現代のビデオゲームは日に日に複雑になってきています。これらの最新のゲームのほとんどは、マルチプレイヤーの一人称シューティング ゲーム (FPS) です。 FPS ゲームの人気の高まりにより、公平で楽しいゲームを行うために不正行為と闘う必要性が強調されています。エイムボット、ウォールハック、スピードハックなどのチートテクニックを使用するプレイヤーの数も増加しているため、通常のプレイヤーに対して不当な優位性を得るためにチートツールを使用しているプレイヤーを検出する方法が必要です。このシステムでは、エイムボットのチートを検出することにのみ重点を置いています。エイムボット チートを使用するプレイヤーは通常、ゲームの他の側面を優先しません。通常のプレイヤーと不正プレイヤーを区別するために、エイム速度、ショット数、ターゲットまでの距離などの時系列データと、ユーティリティの使用状況、プレイヤーの動き、その他のゲームプレイ パターンなどの行動データを含む特定の特徴を特定します。これらの機能を利用して、「YAACS」という名前のサーバー側エイムボット検出分類器を構築します。 YAACS は、パーサー、深層学習モデル、およびゲーム サーバーとの統合用に設計された中間接続ユーティリティで構成されます。提案されたシステムは、128 ティックのシーケンス (ティック デルタ ネガティブ = 56、ティック デルタ ポジティブ = 24) でトレーニングされた高密度レイヤーを備えたスタック LSTM を使用して、分類精度 88.6%、偽陽性率 0.97% を達成し、96.2% の高い精度を達成するデシジョン ツリー ベースラインを上回りますが、偽陽性率は 2.68% と 2.76 倍悪いです。最適な LSTM 構成よりも優れています。これらの結果は、FPS チート検出における冤罪を最小限に抑えるには、シーケンス モデリングを通じて時間的コンテキストを組み込むことが重要であることを示しています。
原文 (English)
Server-side Anti-cheat in FPS games for Aimbot detection using Deep learning and Machine learning
Modern video games are becoming more complex day by day. Most of these modern games are multiplayer first-person shooter (FPS) games. The rising popularity of FPS games emphasizes the need to combat cheating for fair and enjoyable gaming. As the number of players using cheating techniques like aimbots, wallhacks, and speed hacks is also increasing, we need a way to detect players who are using cheating tools to gain an unfair advantage over regular players. In this system, we focus exclusively on detecting aimbot cheats. Players who use aimbot cheats generally do not prioritize other aspects of the game. To distinguish between regular and cheating players, we identify specific features encompassing time-series data such as aim velocity, number of shots, distance to target, and more, along with behavioral data such as utility usage, player movement, and other gameplay patterns. Utilizing these features, we construct a server-side aimbot detection classifier named 'YAACS'. YAACS comprises a parser, a deep learning model, and intermediary connection utilities designed for integration with the game server. The proposed system achieves a classification accuracy of 88.6% with a false positive rate of 0.97% using a Stacked LSTM with Dense layers trained on sequences of 128 ticks (Tick Delta Negative=56, Tick Delta Positive=24), outperforming the Decision Tree baseline which achieves a higher accuracy of 96.2% but at a false positive rate of 2.68%, 2.76x worse than the best LSTM configuration. These results demonstrate that incorporating temporal context through sequence modelling is critical for minimising false accusations in FPS cheat detection.
Nemotron-Labs-3-Puzzle-75B-A9B: ハイブリッド MoE LLM の圧縮
インタラクティブな展開用に最適化された Nemotron-3-Super の圧縮バージョンである Nemotron-Labs-3-Puzzle-75B-A9B を紹介します。ユーザー スループットの高い制約下でサーバー スループットを最大化するようにモデルを設計しました。単一の 8xB200 ノードで対話型のワークロードを処理する場合、Puzzle-75B-A9B は、一致するユーザー スループット制約で Nemotron-3-Super よりも約 2 倍高いサーバー スループットを達成します。単一の H100 GPU での超ロング コンテキストの展開では、圧縮モデルにより 1M トークンの同時実行数が 1 リクエストから 8 リクエストに増加します。 Puzzle-75B-A9B は、反復パズル圧縮フレームワークと知識蒸留、強化学習、量子化、およびマルチトークン予測ヘッドを組み合わせた多段階パイプラインを使用して構築されています。圧縮プロセスは、異種 MoE プルーニング、アクティブ パラメーター バジェット、および Mamba プルーニングを共同で最適化し、モデルの品質を維持しながら推論効率を向上させます。私たちは、推論、コーディング、多言語、ロングコンテキスト、およびエージェントの幅広いベンチマーク スイートで Puzzle-75B-A9B を評価します。大幅な圧縮にもかかわらず、モデルは幅広いタスクにわたって親モデルと比較して強力なダウンストリーム精度を維持します。これらの結果は、大規模なハイブリッド MoE モデルが強力なダウンストリーム機能を維持しながら導入効率を大幅に最適化できることを示しています。
原文 (English)
Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs
We present Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super optimized for interactive deployment. We designed the model to maximize server throughput under high user throughput constraints. In interactive serving workloads on a single 8xB200 node, Puzzle-75B-A9B achieves approximately 2x higher server throughput than Nemotron-3-Super at matched user throughput constraints. In ultra-long-context deployment on a single H100 GPU, the compressed model increases 1M-token concurrency from 1 request to 8 requests. Puzzle-75B-A9B is constructed using a multi-stage pipeline that combines the Iterative Puzzle compression framework with knowledge distillation, reinforcement learning, quantization, and a Multi-Token Prediction head. The compression process jointly optimizes heterogeneous MoE pruning, active parameter budget, and Mamba pruning to improve inference efficiency while preserving model quality. We evaluate Puzzle-75B-A9B on a broad suite of reasoning, coding, multilingual, long-context, and agentic benchmarks. Despite substantial compression, the model retains strong downstream accuracy relative to the parent model across a wide range of tasks. These results demonstrate that large hybrid MoE models can be substantially optimized for deployment efficiency while maintaining strong downstream capability.
賭けメカニズムによる LLM 予測の分散集約
集合的な予測パフォーマンスを向上させるために、それぞれがドメインの専門知識を持っているか、プライベート ツールやデータにアクセスできる複数の LLM からの予測を集約することがますます一般的になっています。分散型設定では、モデルの個人情報にアクセスせずに集計の重みを決定する必要があり、戦略的なレポートに対して堅牢性を維持する必要があります。我々は、各モデルが予測と学習された賭けを報告し、賭けを重みとして使用して予測が集約される、LLM 集約 (WALLA) のための利点に合わせた賭けメカニズムのファミリーを提案します。 WALLA は、ネットペイアウト関数にリーブワンアウトベースラインを導入し、次の 3 つの望ましい特性をもたらします。(1) 任意の信念構造の下での予測のドミナント戦略インセンティブ互換性、(2) 最適な賭け金がモデルの予想されるスコアの利点に比例するアドバンテージ - 賭け金の調整、(3) 予測に依存しない賭け金の最適化。最適な予測を必要とせずに賭け金ポリシーの分散学習を可能にします。さらに、メカニズムの制限された最悪の場合の不足を維持しながら、正規性と裁定なしをトレードオフする 2 つのメカニズムのバリアントをインスタンス化します。異種モデルと個人情報設定にわたる質問応答と予測のベンチマークに関する実験では、WALLA が予測パフォーマンスにおいて集中型の集計手法と同等であると同時に、分散型学習、利点に合わせた集計の重み、不確実性の認識、およびインセンティブと互換性のある予測を達成していることが示されています。
原文 (English)
Decentralized Aggregation of LLM Predictions via Wagering Mechanisms
It is increasingly common to aggregate predictions from multiple LLMs, each with domain expertise or access to private tools and data, to improve collective prediction performance. In decentralized settings, aggregation weights need to be determined without access to models' private information and should remain robust to strategic reporting. We propose a family of advantage-aligned wagering mechanisms for LLM aggregation (WALLA), in which each model reports a prediction and a learned wager, and predictions are aggregated using wagers as weights. WALLA introduces a leave-one-out baseline into the net payout function, yielding three desirable properties: (1) dominant-strategy incentive compatibility of prediction under arbitrary belief structure, (2) advantage--wager alignment, where the optimal wager is proportional to the model's expected score advantage, and (3) prediction-agnostic wager optimization, enabling decentralized learning of wager policies without requiring optimal predictions. We further instantiate two mechanism variants that trade off normality and no-arbitrage while maintaining a bounded worst-case deficit for the mechanism. Experiments on question-answering and forecasting benchmarks across heterogeneous models and private-information settings show that WALLA matches centralized aggregation methods in predictive performance, while simultaneously achieving decentralized learning, advantage-aligned aggregation weights, uncertainty awareness, and incentive-compatible prediction.
MechMath エージェント チーム: 数学研究のための LLM 主導エージェント
AI 推論は、大規模な言語モデルの成功によって主に推進され、現代の人工知能の中心となっています。しかし、数学的研究は、非線形の導出経路、厳密な論理要件、長期にわたる探索サイクルを特徴としており、既存の推論システムに深刻な課題をもたらします。これらの制限を克服するために、私たちは MechMath Agent Team (MMAT) を紹介します。これは、数学的研究の全サイクルを通じて副操縦士として機能するように設計された大規模な言語モデル駆動のエージェントです。私たちは、システムの責任を制御プレーン、実行プレーン、拡張プレーンに分離する 3 つの部分からなるハーネス アーキテクチャを設計します。これにより、厳格な論理制御と、自由研究に求められる機敏性が調和します。このフレームワークに基づいて、Knowledge Base Manager、Natural Language Prover、Formal Language Prover の 3 つの特殊なエージェントをインスタンス化します。これらはすべて閉ループで動作し、正式に認定された数学的証明を生成します。数論、代数複雑性理論、微分代数、作用素代数、および不等式の未解決問題について MMAT を評価します。 2 か月の展開で 11 件の問題が解決され、研究サイクル全体を通じて副操縦士として機能する能力が実証されました。貢献は 3 つあります。マルチエージェント数学的推論のための一般的な分離されたハーネス アーキテクチャ、MMAT システムでのその具体的なインスタンス化、および未解決の問題の多様なスイートに対する経験的検証です。
原文 (English)
MechMath Agent Team: LLM Driven Agents for Mathematical Research
AI reasoning has become a central focus in contemporary artificial intelligence, largely driven by the success of large language models. However, mathematical research, which is characterized by non-linear derivation paths, rigorous logical requirements, and protracted exploration cycles, poses severe challenges for existing reasoning systems. To overcome these limitations, we present the MechMath Agent Team (MMAT), which is a large language model driven agent designed to serve as a co-pilot throughout the full cycle of mathematical research. We design a tripartite Harness Architecture that decouples system responsibilities into Control, Execution, and Augmentation planes, thereby reconciling rigorous logical control with the agility demanded by open-ended research. Building upon this framework, we instantiate three specialized agents: a Knowledge Base Manager, a Natural Language Prover, and a Formal Language Prover, all operating in a closed loop to produce formally certified mathematical proofs. We evaluate MMAT on open problems in Number Theory, Algebraic Complexity Theory, Differential Algebra, Operator Algebra, and Inequalities. Across a two-month deployment, 11 problems have been solved, demonstrating its capacity to act as a co-pilot throughout the entire research cycle. The contributions are threefold: a general decoupled Harness Architecture for multi-agent mathematical reasoning, its concrete instantiation in the MMAT system, and empirical validation on a diverse suite of open problems.
LLM-as-a-Tutor: 検証不可能な RL に対するポリシーに応じたプロンプト適応
検証不可能な指示に従う強化学習 (RL) は、報酬シグナルとしてプロンプト固有のルーブリックを使用する LLM 審査員にますます依存しています。最近の手法では、トレーニング中にこれらのルーブリックを進化するポリシーに適応させていますが、トレーニング プロンプト自体は固定されたコーパスから抽出された静的なものです。この静的なアプローチでは、プロンプトの難易度と政策能力の間に重大な不整合が生じることが多く、プロンプトがロールアウト間の品質の差異を引き出すことができない場合、裁判官は差別的な報酬シグナルを回復できなくなります。この不整合に対処するために、LLM の役割を裁判官から家庭教師に拡張するフレームワークである LLM-as-a-Tutor を導入します。単一のモデルは、ポリシーのロールアウトをペアごとに比較して問題のないプロンプトを検出する検査者として、また、プロンプトにアトミックな制約を追加するジェネレーターとして機能します。この追加専用の設計は、ポリシーの機能に合わせて難易度を単調に上昇させ、外部の難易度スケジュールなしで自己調整トレーニング信号を生成します。 3 つの複雑な指示に従うベンチマークにおいて、私たちの手法は、ポリシーを意識しないベースラインと、ルーブリックを適応させたりプロンプトを書き換えたりする従来のポリシー適応型手法の両方を一貫して上回っており、検証不可能な RL ではポリシー認識の軸が欠けているとして、迅速な適応が示唆されています。
原文 (English)
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL
Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt these rubrics to the evolving policy during training, the training prompts themselves remain static, drawn from fixed corpora. This static approach often results in a critical misalignment between prompt difficulty and policy capability, leaving the judge unable to recover a discriminative reward signal when prompts fail to elicit quality variance among rollouts. To address this misalignment, we introduce LLM-as-a-Tutor, a framework that extends the LLM's role from judge to tutor: a single model serves as an examiner that pairwise-compares policy rollouts to detect non-challenging prompts, and as a generator that appends atomic constraints to them. This append-only design monotonically raises difficulty in step with the policy's capability, producing a self-calibrating training signal without external difficulty schedules. On three complex instruction-following benchmarks, our method consistently outperforms both policy-unaware baselines and prior policy-adaptive methods that adapt rubrics or rewrite prompts, suggesting prompt adaptation as a missing axis of policy-awareness in non-verifiable RL.
エージェント ステップ値: 状態接地 LLM エバリュエーターによる状態遷移測定
ほとんどのエージェント評価では、複数ステップのトレースが最終的な回答、成功フラグ、または軌跡レベルのスコアにまとめられます。これらの集計では、開発者が最も必要とする診断の質問、つまりどのアクションが状態を有益な方向に変更したのかがわかりにくくなります。我々は、状態遷移測定フレームワークであるエージェント ステップ値 (ASV) を導入します。これは、観察された各アクションを、固定された候補結果に対する状態に基づいた評価者の分布に誘発する変化によってスコア付けします。 ASV は、編集された前後の状態予測をレンダリングし、ステートレス LLM エバリュエーターを使用して候補ログ スコアを割り当て、ゴールドフリーの信念診断とオフライン オラクル検証メトリクスの両方をレポートします。ラベルフリーの理論的パスにより、評価者の審議がワントークンオプションのスコアリングから分離され、リークやフロアスコアイベントを明らかにしながら候補の可能性が維持されます。ライブ PubMed 検索、部分的にライブ DeepSeek アクター、および DeepSeek 対数確率スコアリングを使用した 100 件のレビュー済みオープン QA 証拠探索タスクについて、ASV は 1,100 のステップと 2,200 の状態を評価します。固定レイアウトの根拠条件付きプロトコルの下では、平均ゴールドマージンゲインは -2.335 (軌道ブートストラップ 95\% CI [-3.395, -1.272])、エントロピーの動きは 0.000、平均ベイジアンサプライズは 2.693 です。したがって、ASV は、最終回答スコアとエントロピーのみのステップ メトリクスが見逃している建設的および破壊的な信念ピボットを特定します。スタンドアロンの ASV Eval ツールキットをリリースします。
原文 (English)
Agent Step Value: State-Transition Measurement with State-Grounded LLM Evaluators
Most agent evaluations collapse a multi-step trace into a final answer, a success flag, or a trajectory-level score. These aggregates obscure the diagnostic question developers need most: which action changed the state in a useful direction? We introduce Agent Step Value (ASV), a state-transition measurement framework that scores each observed action by the change it induces in a state-grounded evaluator's distribution over fixed candidate outcomes. ASV renders redacted before/after state projections, uses a stateless LLM evaluator to assign candidate log scores, and reports both gold-free belief diagnostics and offline oracle validation metrics. A label-free rationale pass separates evaluator deliberation from one-token option scoring, preserving candidate likelihoods while exposing leakage and floor-score events. On 100 reviewed open-QA evidence-seeking tasks with live PubMed retrieval, a partially live DeepSeek actor, and DeepSeek log-probability scoring, ASV evaluates 1,100 steps and 2,200 states. Under the fixed-layout rationale-conditioned protocol, mean gold-margin gain is -2.335 (trajectory-bootstrap 95\% CI [-3.395, -1.272]), entropy movement is 0.000, and mean Bayesian surprise is 2.693. ASV therefore localizes constructive and destructive belief pivots that final-answer scores and entropy-only step metrics miss. We release the standalone ASV Eval toolkit.
ResearchStudio-Idea: ML カンファレンスの結果から得た、証拠に基づいたリサーチとアイデアのスキル スイート
大規模な言語モデルにより、研究のアイデア出しがますます容易になりましたが、効果的なアイデア開発には、候補となる方向性を生成するだけでは不十分です。研究者は、実装に取り組む前に、現在の文献に基づいて問題を根拠付け、意味のあるボトルネックを特定し、既存のソリューションと区別し、リスクを評価する必要があります。私たちは ResearchStudio-Idea を、研究のアイデア創出の最初の 1 マイルのための再利用可能なスキル スイートとして紹介します。このスイートには、スタンドアロンのマルチソース文献検索スキルである Paper-Search が含まれています。 Scoop-Check、新規性主張のためのスタンドアロンの従来技術の衝突チェッカー。 IdeaSpark は、証拠の根拠付け、パターンに基づく生成、衝突の回復、監査、アイデア カードのレンダリングを 1 つのワークフローにまとめるエンドツーエンドのスキルです。 IdeaSpark は、2021 年から 2025 年の間に ICLR、ICML、NeurIPS から収集された 1,947 件の機械学習カンファレンス論文のコーパスから構築されています。これには口頭論文、個別に追跡された被引用数の多いサブセット、拒否された投稿も含まれます。これらの結果を分析すると、31 の繰り返しのアイデア発想のサブパターンが 15 の再利用可能なアイデア発想パターンに統合されていることがわかります。各パターンは、研究コンテキスト、ボトルネックの種類、差別化戦略、サポートする前例、および一般的な障害モードを含む構造化カードとして運用されます。研究課題と証拠バンドルが与えられると、IdeaSpark は証拠の準備状況を評価し、周囲の研究コンテキストを再構築し、未解決のボトルネックを特定し、関連するパターンを選択し、1 つの候補方向をインスタンス化し、競合する可能性のある以前の研究を取得し、結果に基づいた監査を実行します。このワークフローは、再利用可能なアイデアのパターンを追跡可能な研究提案に変換します。ブラインド自動判定による評価では、IdeaSpark が競争力のある新規性を維持しながら、ノースキルおよびジェネリックスキルのベースラインよりも強力な研究提案を一貫して生成していることが示されています。
原文 (English)
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
Large language models have made research ideation increasingly accessible, yet effective idea development requires more than generating candidate directions. Researchers must ground a problem in current literature, identify meaningful bottlenecks, differentiate from existing solutions, and evaluate risks before committing to implementation. We present ResearchStudio-Idea as a reusable skill suite for this first mile of research ideation. The suite includes Paper-Search, a standalone multi-source literature search skill; Scoop-Check, a standalone prior-art collision checker for novelty claims; and IdeaSpark, the end-to-end skill that composes evidence grounding, pattern-guided generation, collision retrieval, audit, and idea-card rendering into one workflow. IdeaSpark is constructed from a corpus of 1,947 machine learning conference papers collected from ICLR, ICML, and NeurIPS between 2021 and 2025, including Oral papers, a separately tracked high-citation subset, and rejected submissions. Analysis of these outcomes reveals 31 recurring ideation sub-patterns, consolidated into 15 reusable ideation patterns. Each pattern is operationalized as a structured card containing research contexts, bottleneck types, differentiation strategies, supporting precedents, and common failure modes. Given a research problem and an evidence bundle, IdeaSpark evaluates evidence readiness, reconstructs the surrounding research context, identifies unresolved bottlenecks, selects relevant patterns, instantiates one candidate direction, retrieves potentially conflicting prior work, and performs outcome-informed auditing. This workflow transforms reusable ideation patterns into traceable research proposals. Blind automated-judge evaluations show that IdeaSpark consistently produces stronger research proposals than no-skill and generic-skill baselines while maintaining competitive novelty.
純粋な推論だけでは不十分な理由: 数学的イノベーションの源としての自然
我々は、人間の数学的推論は、たとえ控えめな論理断片であっても決定不能性と計算困難性の両方によって制約を受けており、基本的には純粋な演繹の外側の領域からのパターンマッチングに依存しているという仮説を推し進めます。そのようなパターンの最も豊富な宝庫は自然界であり、その物理法則と生物学的システムは数十億年にわたる「事前計算」を経て、すでに驚くほど革新的な解決策を示しています。この主張を根拠付けるために、振動弦論争から聴覚方程式とその後の数学に普及した形式主義に至るまで、フーリエ変換と関連する数学の歴史をたどります。重要な局面ごとに、物理学の問題により、純粋な形式的推論では予測できなかったり、さらに悪いことに人間の推論が抵抗したりする数学的ツールの受け入れまたは作成が余儀なくされました。さらに、NP の難しい命題充足可能性からモナディック 2 次理論の非要素決定手順に至るまで、論理の複雑さの状況を調査し、論理が決定可能な場合でも、最悪の場合の演繹に必要なリソースが天文学的に法外な量であることを実証します。私たちは、これらの障壁により、物理学にヒントを得たパターンマッチングが単なる歴史的な偶然ではなく、認知的な必然性をもたらしていると主張します。最後に、人工知能に対する結論を導き出します。純粋な推論が構成的に不十分な場合、人間レベルの数学的創造性を目指すシステムは、演繹のみに依存するのではなく、クロスドメインのパターンの膨大な蓄積を組み込む必要があります。これは、現代の大規模言語モデルの巨大な規模に対する原則的な正当化を提供します。
原文 (English)
Why Pure Reasoning is Not Enough: Nature as the Source of Mathematical Innovation
We advance the hypothesis that human mathematical reasoning, constrained by both the undecidability and the computational intractability of even modest logical fragments, relies fundamentally on pattern matching from domains external to pure deduction. The most prolific reservoir of such patterns is the natural world, whose physical laws and biological systems have undergone billions of years of ``pre-computation'' and already exhibit surprisingly innovative solutions. To ground this claim, we trace the history of the Fourier transform and relevant mathematics, from the vibrating string controversy to the hear equation and subsequent formalisms prevalent in mathematics. At each critical juncture, a physics problem forced the acceptance or creation of a mathematical tool that pure formal reasoning failed to anticipate or, worse, human reasoning had resisted. We further survey the landscape of logical complexity, from NP-hard propositional satisfiability to the non-elementary decision-procedures for monadic second-order theories, to demonstrate that even when a logic is decidable, the resources required for worst-case deduction are astronomically prohibitive. We argue that these barriers make physics-inspired pattern matching not just a historical accident but a cognitive necessity. Finally, we draw the consequence for artificial intelligence: if pure reasoning is constitutively insufficient, then any system aiming at human-level mathematical creativity must embed a vast store of cross-domain patterns rather than rely on deduction alone. This furnishes a principled justification for the enormous scale of contemporary large language models.
検証のボトルネックを圧縮: 科学的発見のためのエージェント自動運転ラボ
Agentic AI-for-Science はアイデア出し、計画、分析を自動化できますが、最終的な検証は依然として実際の実験に依存します。自動運転ラボ (SDL) はこれらの実験を実行できますが、ループには依然としてボトルネックがあります。エージェントが価値の低い実験に多くのラウンドを費やしすぎたり、各ラウンドで高コストの実験が必要になったりする可能性があります。これら 2 つの物理的なボトルネックを 1 つのエージェントでターゲットにします。まず、事前認識エージェント DOE ループは、ドメインの知識と過去の結果を使用して、実行可能で有益な次の実験を提案し、目標までの試行回数を減らします。第 2 に、コストを意識した代理エージェントは、低コスト、低解像度の測定値から高コスト、高解像度の測定値を予測します。予測された不確実性に基づいて、高コストの測定と低コストの測定のどちらかを選択します。私たちはこれらの方向性をそれぞれ生物学と材料の領域で検討します。これらのコンポーネントは、単一のエージェントの下で連携して、ループ数と実験あたりのコストの両方を削減することで SDL ループを高速化することを目的としています。
原文 (English)
Compressing the Validation Bottleneck: An Agentic Self-Driving Lab for Scientific Discovery
Agentic AI-for-Science can automate ideation, planning, and analysis, but final validation still depends on real experiments. A self-driving lab (SDL) can execute those experiments, yet the loop still has bottlenecks: the agent may spend too many rounds on low-value experiments, or each round may require a high-cost experiment. We target these two physical bottlenecks with one agent. First, a prior-aware agentic DOE loop uses domain knowledge and past results to propose feasible and informative next experiments, reducing trials-to-target. Second, a cost-aware surrogate agent predicts high-cost, high-resolution measurements from low-cost, low-resolution measurements. It chooses between a high- and a low-cost measurement based on the predicted uncertainty. We examine these directions in the biology and materials domains, respectively. Together, under a single agent, these components aim to accelerate the SDL loop by reducing both the number of loops and the cost per experiment.
VLA Grounder: ブラックボックス VLA モデルの言語条件付け空間の最適化
ビジョン言語アクション (VLA) モデルは、通常、自然言語タスクの記述を条件としたエンドツーエンドのアクション ポリシーとして扱われます。しかし実際には、彼らの行動は指示の表現方法に大きく依存することが多く、言語が単なるタスクのラベルではなく、最適化可能な条件入力であることを示唆しています。私たちは、アクションの重みを更新するのではなく、言語空間を最適化することによって、凍結された VLA ポリシーを改善できるかどうかを研究します。私たちの方法では、オブジェクトの外観、空間関係、ターゲットグラウンディングの手がかりを使用して、人間の指示を短い VLA ベースのコマンドに変換する言語条件付け空間ポリシーを導入します。言語条件付け空間ポリシーは、事前に失敗由来のコマンド空間で初期化され、まばらなタスク完了報酬からの強化学習で最適化されますが、下流の VLA は完全に凍結されたままになります。これにより、言語条件付け空間の最適化が実現します。RL は、凍結されたアクション ポリシーから正常な動作を引き出すのに最も適した VLA ベースのコマンドを発見します。 RL4VLA と VL-Think の実験では、言語条件付け空間の最適化により、命令依存型、シンボリック、およびマルチオブジェクト操作タスクの成功率が向上することが示され、言語がロボット基盤モデルの最適化可能な変数として機能できることが実証されました。ウェブサイト: https://tttonyalpha.github.io/vla_grounder
原文 (English)
VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models
Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. In practice, however, their behavior often depends sharply on how the instruction is phrased, suggesting that language is not merely a task label but an optimizable conditioning input. We study whether frozen VLA policies can be improved by optimizing language space rather than updating action weights. Our method introduces a language-conditioning space policy that translates a human instruction into a short VLA-grounded command using object appearance, spatial relations, and target-grounding cues. The language-conditioning space policy is initialized with a failure-derived command-space prior and optimized with reinforcement learning from sparse task-completion rewards, while the downstream VLA remains fully frozen. This yields language-conditioning space optimization: RL discovers which VLA-grounded commands best elicit successful behavior from the frozen action policy. Experiments on RL4VLA and VL-Think show that language-conditioning space optimization improves success on instruction-sensitive, symbolic, and multi-object manipulation tasks, demonstrating that language can serve as an optimizable variable for a robot foundation models. Website: https://tttonyalpha.github.io/vla_grounder
マルチステップ LLM エージェントにおけるハーネスによって引き起こされる信念の相違の測定
ソフトウェア エージェントのベンチマークは通常、エージェントがタスクを解決したかどうかを報告しますが、エージェントは、何を表示するか、どのアクションを実行できるか、どの障害が修復されるか、どの状態が検証されるか、どの証拠が記録されるかを制御するハーネスを通じてその結果に到達します。このハーネスは、タスク、環境、およびベース LLM が固定されている場合でも、エージェントの複数ステップの信念を変更できることを示します。私たちは、進行状況、リスク、回復可能性、制約、故障モード、不確実性、将来の成功、修理コスト、代替ハーネスの下での次のアクションに関する構造化された K ステップ軌跡を導き出す信念ロールアウト診断を導入します。クロスハーネス信念の発散を定義し、それを即時のインターフェースの変化を表す到着項と、地平線に依存する信念の変化を表す成長項に分解します。制御されたコーディング タスクや公開ベンチマーク ストレス テストでは、ブロックされたアクション、圧縮修復、選択的検証、コストを意識した証拠の刈り込みにより、最終的な成功を維持しながら、後の意思決定を促す信念を変えることがよくあります。さらに、観察を正規化し、検閲された分岐を記録し、修復トレースを拡張し、検証マスクを記録し、危険な分岐をシャドウで実行し、ハーネス ビュー全体で信念の軌道を調整するトレーニング不要のプロトコルである BIWM を紹介します。この結果は、ハーネス設計はエージェント評価における実験変数であり、実装の詳細ではないことを示唆しています。私たちのコードは https://github.com/Hik289/Harness-induce-bias.git で入手できます。
原文 (English)
Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents
Software-agent benchmarks usually report whether an agent solves a task, but the agent reaches that outcome through a harness that controls what it sees, which actions it can take, which failures are repaired, which states are verified, and which evidence is logged. We show that this harness can change the agent's multi-step beliefs even when the task, environment, and base LLM are fixed. We introduce a belief-rollout diagnostic that elicits structured K-step trajectories over progress, risk, recoverability, constraints, failure mode, uncertainty, future success, repair cost, and next action under alternative harnesses. We define a cross-harness belief divergence and decompose it into an arrival term for immediate interface shifts and a growth term for horizon-dependent belief changes. On controlled coding tasks and public-benchmark stress tests, blocked actions, compressed repairs, selective verification, and cost-aware evidence pruning often preserve terminal success while changing the beliefs that drive later decisions. We further introduce BIWM, a no-training protocol that canonicalizes observations, logs censored branches, expands repair traces, records verification masks, executes risky branches in shadow, and aligns belief trajectories across harness views. The results suggest that harness design is an experimental variable in agent evaluation, not an implementation detail. Our code is available at https://github.com/Hik289/Harness-induce-bias.git.
大規模言語モデルにおける認識エントロピーを除去するためのローリング係数のヘヴィサイド連続性
大規模言語モデル (LLM) は、間違っている可能性がある流暢な出力を生成します。誤った情報を提供するときに手がかりを示すことが多い人間とは異なり、LLM は検出が困難なエラーを生成します。これは、自己回帰デコードには状態が進行する前に中間推論を検証するメカニズムがないためです。ヘビサイド ゲートによって制御される述語ゲート状態遷移として推論を再定式化する検証優先実行フレームワークであるヘビサイド ローリング係数 (HCRC) を紹介します。 HCRC は、モデルの信頼性とパラレル ワーカー アーキテクチャからの独立した検証信号を組み合わせて、事前定義された正確性述語が満たされた場合にのみ実行を進めることができます。これにより、無効な中間状態の伝播が防止され、基礎となるモデルを変更することなく認識論的エントロピーが削減されます。私たちは、4 つのプロバイダーからの 13 人の提案者にわたって、ソフトウェア エンジニアリングと推論タスクに関して HCRC を評価します。有能なプロポーザーでは、ゲートは遅延競合性を維持しながら誤完了率 (FCR) を 4 ~ 7% から 0% に削減し、設定によってはアンラップ モデルよりも高速になります。弱いプロポーザーでは、ダウンストリームの状態を破壊するのではなく、誤った完了を正当な停止に変換します。ベンチマークを超えて、HCRC はエージェント コーディング環境の実稼働コントロール プレーンとして数か月間運用され、ファイルの変更の承認、検証主導の進捗レポート、メモリ圧縮を行ってきました。これらの結果は、検証主導型 LLM 実行の一般的なフレームワークとして HCRC を確立し、信頼性の高い推論がモデル スケールだけではなく原則に基づいた実行制御を通じて達成できることを示しています。
原文 (English)
Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models
Large language models (LLMs) generate fluent outputs that can be wrong. Unlike humans, who often exhibit cues when providing false information, LLMs produce errors that are difficult to detect because autoregressive decoding provides no mechanism for verifying intermediate reasoning before state progression. We introduce Heaviside Continuity of Rolling Coefficients (HCRC), a verification-first execution framework that reformulates inference as predicate-gated state transitions governed by a Heaviside Gate. HCRC combines model confidence with independent verification signals from a parallel worker architecture, allowing execution to advance only when predefined correctness predicates are satisfied. This prevents invalid intermediate states from propagating, reducing epistemic entropy without modifying the underlying model. We evaluate HCRC on software-engineering and reasoning tasks across thirteen proposers from four providers. On capable proposers, the gate reduces the false-completion rate (FCR) from 4--7% to 0% while remaining latency-competitive and, in some settings, faster than the unwrapped model. On weaker proposers, it converts false completions into honest halts instead of corrupting downstream state. Beyond benchmarking, HCRC has operated for months as the production control plane of an agentic coding environment, authorizing file mutations, verification-driven progress reporting, and memory compaction. These results establish HCRC as a general framework for verification-driven LLM execution, showing that reliable reasoning can be achieved through principled execution control rather than model scale alone.
切り捨てられた思考連鎖監査による LLM ベースの教育講師の回答主導型推論の検出
大規模言語モデル (LLM) の家庭教師は、流暢な段階的な説明を行うことがよくありますが、正しく教育的にフォーマットされた応答は、その答えが生徒が直面している問題から導き出されたものであることを保証しません。現実的な個別指導システムでは、モデルは教師のノート、解答キー、ルーブリック、または取得されたソリューションのアーティファクトにもアクセスできる場合があります。私たちは、そのようなプライベートな解答情報によって、家庭教師の説明が解答主導型になるかどうか、つまり最終的な解答は、書面による説明が正当化する前に行動的に入手可能になるかどうかを研究しています。思考連鎖プレフィックスが検証者をどれだけ早く通過できるかを調査する Truncated Reasoning AUC Evaluation (TRACE) を使用して、質問のみ、正解キー、間違った解答キーという 3 つのペアの個別指導コンテキストの下で 1000 件の GSM8K テスト問題を評価しました。生成された各説明の一定の割合で、モデルに即時応答を強制し、その応答を黄金の数値応答と照合して検証します。 Qwen2.5-3B-Instruct を使用すると、アンサーキー アクセスにより TRACE AUC の中央値が 0.375 から 0.900 に上昇し、1000 件中 997 件の最初の 10% プレフィックスでゴールド アンサーが利用可能になります。この効果は、質問のみの説明と解答キーの説明の両方が正解で終わる 746 例では依然として強いです。これらの結果は、数学の個別指導の説明における回答主導型の推論のための軽量のプロセス レベルの診断として、切り詰められた CoT 監査をサポートします。
原文 (English)
Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing
Large language model (LLM) tutors often produce fluent step-by-step explanations, but a correct and pedagogically formatted response does not guarantee that the answer was derived from the student-facing problem. In realistic tutoring systems, the model may also have access to teacher notes, answer keys, rubrics, or retrieved solution artifacts. We study whether such private answer information can make tutor explanations answer-driven: the final answer is behaviorally available before the written explanation has justified it. Using Truncated Reasoning AUC Evaluation (TRACE), which probes how early a chain-of-thought prefix can pass a verifier, we evaluate 1000 GSM8K test problems under three paired tutoring contexts: question-only, correct answer-key, and wrong answer-key. At fixed fractions of each generated explanation, we force the model to answer immediately and verify the response against the gold numeric answer. With Qwen2.5-3B-Instruct, answer-key access raises median TRACE AUC from 0.375 to 0.900 and makes the gold answer available at the first 10% prefix in 997 of 1000 cases. The effect remains strong on the 746 examples where both question-only and answer-key explanations end with the correct answer. These results support truncated CoT auditing as a lightweight process-level diagnostic for answer-driven reasoning in math tutoring explanations.
注意が限定された報酬学習
人間のペアごとの比較は、最新の AI システムが人間の好みを学習するための主要なインターフェイスです。 RLHF および関連するアライメント パイプラインは通常、ブラッドリーとテリーの対数オッズとの比較をモデル化し、選択確率は潜在的な報酬の違いによって決まります。この論文では、各ラベルが低容量の評価チャネルによって生成される、合理的な不注意によって動機付けられた縮小形式のモデルを通じて、この仮定が見逃しているものを検証します。このモデルは、標準的な報酬モデリングが混同しがちな 2 つの形式のあいまいさを分離します。2 つの候補の値が本当に近いため、または限られた注意の下では関連する区別を検出するのが難しいため、比較が難しい場合があります。私たちは、注目が限定されていると、ペアごとの比較で明らかになることを根本的に歪めてしまう可能性があることを示します。特に、受動的比較データは一般に報酬、注意、デフォルトの傾向を区別できず、不均一な注意により標準的なブラッドリー-テリー報酬モデリングが誤解を招くランキングを回復する可能性があります。私たちの分析では、学習はラベルの生の数ではなく、各ラベルが持つ注目される情報の量によって左右されることが示されています。 Chatbot Arena の言語モデルのペアに対する人間の投票に関するケーススタディでは、予測された署名、つまりサンプリング ノイズを超え、スカラー報酬では表現できない比較データの周期成分が示されています。知覚比較に関する 2 番目のケーススタディでは、応答時間と視線には、ラベルには含まれないギャップ情報が含まれることが示されています。この観点は、人間のフィードバックは直接明らかにされた好みとしてではなく、注意を限定した測定プロセスとして扱われるべきであることを示唆しています。つまり、弱い好みシグナルは、真の無関心ではなく、隠れた評価の難しさを反映している可能性があります。
原文 (English)
Attention Limited Reward Learning
Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences. RLHF and related alignment pipelines typically model such comparisons with Bradley--Terry log-odds, where choice probabilities are governed by latent reward differences. This paper examines what this assumption misses through a reduced-form model motivated by rational inattention, in which each label is generated by a low-capacity evaluation channel. The model separates two forms of ambiguity that standard reward modeling tends to conflate: a comparison may be difficult because the two candidates are genuinely close in value, or because the relevant distinction is hard to detect under limited attention. We show that limited attention can fundamentally distort what pairwise comparisons reveal. In particular, passive comparison data cannot generally distinguish reward, attention, and default tendencies, and heterogeneous attention can make standard Bradley--Terry reward modeling recover misleading rankings. Our analysis shows that learning is governed not by the raw number of labels, but by the amount of attended information each label carries. A case study on human votes over language-model pairs from Chatbot Arena exhibits the predicted signature, a cyclic component of the comparison data that exceeds sampling noise and that no scalar reward can represent; a second case study on perceptual comparisons shows that response times and gaze carry gap information that the labels do not. This perspective suggests that human feedback should be treated not as direct revealed preference, but as an attention-limited measurement process: a weak preference signal may reflect hidden evaluation difficulty rather than genuine indifference.
統治された個別化: エージェントの学習をその権限から暗号で切り離す
自律エージェントは、サンドボックス化されたテキスト ジェネレーターから、コード、データ、物理インフラストラクチャのオペレーターに移行しており、展開中に学習することが増えています。これは、調整技術が確率的にのみ答えるという疑問を再び明らかにします。つまり、エージェントが現場で適応した後でも、実行中のシステムはオペレータが許可した範囲にまだ制限されているのでしょうか?ここでは、閉じ込めがトレーニングの確率的な結果ではなく、エージェントの実行アーキテクチャの不変量として保証できることを示します。管理された個別化は、起動時にエージェントを暗号的に凍結された ID ダイジェストにバインドし、すべてのアクションを、名前ではなくアクションの意味論的な効果に対して定義されたゲートを介してルーティングします。私たちは、いくら学習、スキルの習得、または自己誘発的なガバナンスの抽象化を行っても、オペレーターによる署名によるアイデンティティーの変更がなければ、エージェントに許可される権限を拡大できないことを証明しています。この保証は、エージェントが独自の安全原則を誘導し、その原則が間違っている場合でも有効です。経験的に、大きなアクション空間が名前ベースのブロックを排除するオープンエンドのツール使用ベンチマークでは、報酬プレッシャーにさらされた管理されていないソフトウェア エージェントは、最も困難なタスクの実行ごとに到達するタスク依存のレートで自身の評価を改ざんしようとしますが、ゲートはタスクの成功を維持しながら、構築の検証されたプロパティとして実行された禁止効果をゼロに削減します。意味論の深さが増加したモニターの敵対的評価では、誤許可が 75% (名前ベースのゲーティング) からゼロ (動的効果追跡) に低下し、拒否履歴により遵守が保留されたレッドラインファミリーに移されることが示されています。導入された学習エージェントに対する信頼は、その継続的な調整に対する賭けから、起動時に誰でも実行できるチェックに変わります。
原文 (English)
Governed Individuation: Cryptographically Decoupling an Agent's Learning from Its Authority
Autonomous agents are moving from sandboxed text generators to operators of code, data, and physical infrastructure, and they increasingly learn while deployed. This reopens a question that alignment techniques answer only probabilistically: after an agent has adapted in the field, is the running system still confined to what its operator authorised? Here we show that confinement can be guaranteed as an invariant of the agent's execution architecture rather than a probabilistic outcome of its training. Governed individuation binds an agent at boot to a cryptographically frozen identity digest, and routes every action through a gate defined over the semantic effect of the action rather than its name. We prove that no amount of learning, skill acquisition, or self-induced governance abstraction can widen the agent's permitted authority without an operator-signed change to its identity; the guarantee holds even when the agent induces its own safety principle and that principle is wrong. Empirically, in an open-ended tool-use benchmark where a large action space rules out name-based blocking, ungoverned software agents under reward pressure attempt to tamper with their own evaluation at a task-dependent rate that reaches every run on the hardest task, whereas the gate reduces executed forbidden effects to zero as a verified property of the construction while preserving task success. An adversarial evaluation of monitors of increasing semantic depth shows false-allows falling from 75% (name-based gating) to zero (dynamic effect tracing), and refusal history transfers compliance to held-out red-line families. Trust in a deployed learning agent shifts from a wager on its continued alignment to a check anyone can run at boot.
MRMS: 長期間存続する AI エージェント用の多重解像度メモリ基板
存続期間の長い AI エージェントにはインタラクション間の継続性が必要ですが、プロンプト ウィンドウを単に延長するだけでは連続性を得ることができません。エージェントは、有用な過去の経験を保存し、それを選択的に取得し、個人的な状況を外部の証拠から区別し、根底にある状況が変化したときに記憶を修正する必要があります。我々は、2 つの直交軸に沿って構成されたアーキテクチャ メモリ基板を提案します。1 つは構造化レコード、ベクトル表現、グラフ関係にまたがる表現軸です。そして、短期的なトレース、中期的な抽象化、長期的な意味論的なコミットメントにわたる時間軸です。その主要な設計制約は、同期化された構造化ベクトル グラフ メモリです。構造化レコードが適格性を管理し、ベクトル表現がリコールをサポートし、ゲートされたコンテキスト投影の前にグラフ関係がサポート、矛盾、および置き換えを判断します。その中心的な主張は、信頼性の高いパーソナライゼーションは記憶設計の問題であるということです。有用な記憶は、未分化な会話履歴として保存されるのではなく、構造化され、選択的に露出され、継続的に統合され、認識論的にラベル付けされます。フレームワークを超えて、構造化レコード、ベクトル検索、時間ポリシー、グラフベースのリビジョンを実装する軽量のプロトタイプとして MRMS をインスタンス化します。このプロトタイプは、明示的な証拠要件を備えた制御された長期相互作用シナリオの下で、世代前のメモリ選択、改訂、境界適用、および証拠帰属を通じてコア基板メカニズムを実行します。
原文 (English)
MRMS: A Multi-Resolution Memory Substrate for Long-Lived AI Agents
Long-lived AI agents require continuity across interactions, but continuity cannot be obtained by simply extending the prompt window. An agent must preserve useful prior experience, retrieve it selectively, distinguish personal context from external evidence, and revise memory when the underlying situation changes. We propose an architectural memory substrate organized along two orthogonal axes: a representational axis spanning structured records, vector representations, and graph relations; and a temporal axis spanning short-term traces, medium-term abstractions, and long-term semantic commitments. Its key design constraint is synchronized structured-vector-graph memory: structured records govern eligibility, vector representations support recall, and graph relations adjudicate support, contradiction, and supersession before gated context projection. Its central claim is that reliable personalization is a memory design problem: useful memory is structured, selectively exposed, continuously consolidated, and epistemically labeled rather than stored as undifferentiated conversation history. Beyond the framework, we instantiate MRMS as a lightweight prototype implementing structured records, vector retrieval, temporal policies, and graph-based revision. The prototype exercises the core substrate mechanisms through pre-generation memory selection, revision, boundary enforcement, and evidence attribution under controlled long-lived interaction scenarios with explicit evidence requirements.
Formal Disco: 正式に検証されたプログラムのスケーラブルなオープンエンド生成
AI エージェントの能力が向上するにつれて、コードの作成コストは急速に減少していますが、生成されたプログラムの品質保証は追いついていません。形式的検証は可能な限り強力な保証を提供しますが、AI モデルが検証対応言語で動作する能力は、これらの言語で人間が作成したプログラムの例が不足しているために妨げられています。この蔓延するデータ不足の問題に取り組むために、私たちは Formal Disco を提案します。これは、大規模なオープンエンドの合成データ生成に簡単に適用できる、LLM ベースのワーカーを調整するための分散システムです。私たちは、Formal Disco を使用して、3 つのクラスのワーカー間でタスクとプログラムを共有します。オープンソース リポジトリとドキュメント スニペットからランダムな README を読み取って、関連する検証済みプログラムをスケッチする「イニシエーター」、コンパイラーとベリファイアーのフィードバックを受け取り、問題の解決を試みる「フィクサー」、および動作中のプログラムを取得し、それらを拡張するためのパッチを提案する「エクステンダー」です。 Formal Disco は、エージェントが生成したすべてのトレースを記録し、それらをより強力なモデルからの最初の抽出と自己改善の両方に使用します。また、合成プログラム生成のためのエントロピー最大化の原理を提案し、教師あり微調整の反復によるエントロピー最大化を使用して、時間の経過とともにますます多様化するプログラムを生成する方法を学習します。私たちは、Dafny、Verus、Frama-C の 3 つの言語で合成検証済みプログラムの大規模なデータセットをリリースし、検証関連タスク用にオープン モデルを微調整し、多くの場合、Claude Opus 4.5 のパフォーマンスと同等またはそれを超えています。全体として、私たちの取り組みは、形式的推論ドメイン向けに大規模な合成データを作成し、長年にわたるデータの壁を克服する道を提供します。
原文 (English)
Formal Disco: Scalable Open-Ended Generation of Formally Verified Programs
The cost of producing code is rapidly diminishing with increasingly capable AI agents, while quality assurance of generated programs has not kept pace. Formal verification provides the strongest possible guarantees, but the ability of AI models to work with verification-aware languages is hindered by the scarcity of human-written examples of programs in those languages. To tackle this prevalent data scarcity issue, we propose Formal Disco: a distributed system for coordination of LLM-based workers that can be easily applied to open-ended synthetic data generation at scale. We use Formal Disco to share tasks and programs between three classes of workers: "initiators", which read random READMEs from open-source repositories and documentation snippets to sketch a related verified program, "fixers" which take compiler and verifier feedback and attempt to resolve issues, and "extenders" that take working programs and propose patches to expand them. Formal Disco records all agent-generated traces and uses them both for initial distillation from a stronger model as well as self-improvement. We also propose a principle of maximum entropy for synthetic program generation, and use entropy maximization via iterative supervised fine-tuning to learn to generate increasingly diverse programs over time. We release large datasets of synthetic verified programs in three languages - Dafny, Verus, and Frama-C -, and fine-tune open models for verification-relevant tasks, often matching or exceeding the performance of Claude Opus 4.5. Overall, our work offers a path to create synthetic data at scale for formal reasoning domains and overcome the long-standing data barrier.
統合された利他主義と公平性の選好が、連続する社会的ジレンマにおいて高度な相互協力を引き起こす
分散エージェント間の協力を誘導することは、マルチエージェント強化学習 (MARL) の分野、特に社会的ジレンマ状況において依然として難しい問題です。そこでは、個人の利益が共通の善と一致せず、個人の合理性がグループの結果を最適以下に導きます。対照的に、人間はそのような状況でも互いに協力することができます。このような協力的な行動の一般的な説明は、個人には社会的な好みがあるということです。 MARLにおける協力を実現するために、社会心理学と行動経済学に基づく利他的選好(他者の報酬に対するインセンティブ)と公平性選好(平等に対するインセンティブ)を統合した新たな効用関数、すなわち、自他への報酬を協力行動のインセンティブに変換する報酬共有メカニズムである利他的かつ公正な選好(AFP)を設計する。私たちは、2 つの挑戦的な連続社会的ジレンマ ゲームにおいて、標準的な RL エージェントと不平等回避エージェントを使った比較実験を実施しました。その結果、AFP エージェントは、ベースラインよりも多くの集合的報酬と高い公平性を備えた相互協力を達成することに成功したことが示されました。トレーニング中の AFP の進行をさらに理解するために、その後、利他的選好と公平性選好がエージェントの行動に及ぼす影響を調査します。この結果は、利他的選好がエージェントの公共財への貢献を促し、公平性選好がエージェント間の相互行動を誘発することを示唆しています。
原文 (English)
Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas
Inducing cooperation among distributed agents is still a difficult problem in the field of multi-agent reinforcement learning (MARL), particularly in social dilemma situations. There, individual interests are misaligned with the common good and individual rationality leads to suboptimal group outcomes. In contrast, humans are able to achieve cooperation with one another in such situations. A common explanation for such cooperative behavior is that individuals have social preferences. In order to achieve cooperation in MARL, we design a new utility function integrating altruistic preferences (incentive for other's reward) and fairness preferences (incentive for equality) from social psychology and behavioral economics, namely, Altruistic and Fairness Preference (AFP), a reward-sharing mechanism which converts one's own and other's rewards to incentives for cooperative behavior. We performed comparative experiments with standard RL and inequity aversion agents in two challenging sequential social dilemma games, and showed that AFP agents successfully achieved mutual cooperation with more collective rewards and higher equity than the baselines. To further understand the progression of AFP during training, we subsequently explore the effects of altruistic preferences and fairness preferences on agents' behavior. The results suggest that altruistic preferences encourage agents to contribute to the public goods, and fairness preferences induce mutual behavior between agents.
FORGE: 深層研究エージェントに対する研究軌道ハイジャック攻撃
深層調査エージェントは、オープンエンドのクエリをサブタスクに分解し、複数のラウンドにわたって Web 証拠を取得し、長文のレポートを合成します。このワークフローは、計画層のポイズニング サーフェスを作成します。つまり、検索プールに侵入した敵対的な文書は、フォローアップの質問を誘導し、ローカル インジェクションをレポート レベルの汚染に変える可能性があります。我々は、文書内の推論捏造と文書間のチェーン調整を組み合わせてサブタスク計画をハイジャックする 2 レベルの攻撃である FORGE (Fabricated Orchestrated Reasoning chain for aGent Exploitation) を紹介します。さらに、感染レポートの主張を認知タイプごとに重み付けする PRISM メトリクスと、再帰的な追跡生成をルート クエリに結び付ける軽量の防御手段であるルート クエリ アンカーリングを紹介します。 25 のクエリ全体で、Network FORGE は 5 つの挿入されたドキュメントで 26.4% の PRISM に達し、再帰的合成によって汚染されたコンテンツをあからさまなフレーミングから事実の前提に移す深さの移行を示しました。 10 クエリ防御サブセットでは、RQA (ルート クエリ アンカリング) により PRISM が 38.5% から 18.3% に減少します。
原文 (English)
FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents
Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form reports. This workflow creates a planning-layer poisoning surface: adversarial documents that enter the retrieval pool can steer follow-up questions and turn a local injection into report-level contamination. We present FORGE (Fabricated Orchestrated Reasoning chain for aGent Exploitation), a two-level attack that combines intra-document reasoning fabrication with inter-document chain coordination to hijack subtask planning. We further introduce the PRISM metric, which weights infected report claims by cognitive type, and Root Query Anchoring, a lightweight defense that ties recursive follow-up generation to the root query. Across 25 queries, Network FORGE reaches 26.4% PRISM with five injected documents and exhibits depth migration, in which recursive synthesis shifts poisoned content from overt framing into factual premises. On the 10-query defense subset, RQA (Root Query Anchoring) reduces PRISM from 38.5% to 18.3%.
FM-ChangeNet: パスワイズ機能トランスポートによる変化の学習
我々は、バイタイム推論を静的なエンドポイント比較ではなく特徴空間での連続トランスポートとして再定式化する、変更検出のためのパスごとに監視されたフレームワークである FM-ChangeNet を紹介します。エンコードされた前時間表現と後時間表現が与えられると、中間潜在状態を構築し、変換軌道に沿って時間条件付き速度場 $\hat{v}_\theta(z_t,t)$ を学習します。このパスワイズ定式化は、中間状態の連続体にわたって予測子を制約し、従来のエンドポイントのみのセグメンテーションよりも高密度で曖昧さの少ない監視信号を提供し、モデルが時間的進化を明示的に捉えることができるようにします。学習された速度場は、伝達メカニズムであるだけでなく、変化の解釈可能な表現でもあります。その大きさは、照明のずれや空間的な位置ずれなどの迷惑な影響から真の構造変化を区別するのに役立つ、空間的に局所的な変化の手がかりとして機能します。私たちは、時間横断的なアライメント、時間条件付きの粗いから細かいフローのデコード、およびフロー監視、軌道の一貫性、空間的正則化、およびセグメンテーション損失を結合する統一された目標を備えた階層的なマルチスケール アーキテクチャを開発します。リモート センシング ベンチマークの実験では、提案されたフレームワークが最先端のパフォーマンスを達成しながら、より構造化された堅牢な変更表現を生成することが示されています。
原文 (English)
FM-ChangeNet: Learning Change through Pathwise Feature Transport
We present FM-ChangeNet, a pathwise-supervised framework for change detection that reformulates bi-temporal reasoning as continuous transport in feature space rather than static endpoint comparison. Given encoded pre and post-temporal representations, we construct intermediate latent states and learn a time-conditioned velocity field $\hat{v}_\theta(z_t,t)$ along the transformation trajectory. This pathwise formulation constrains the predictor over a continuum of intermediate states, providing a denser and less ambiguous supervision signal than conventional endpoint-only segmentation and enabling the model to capture temporal evolution explicitly. The learned velocity field is not only a transport mechanism but also an interpretable representation of change: its magnitude serves as a spatially localized change cue that helps distinguish true structural variation from nuisance effects such as illumination shifts and spatial misalignment. We develop a hierarchical multi-scale architecture with cross-temporal alignment, time-conditioned coarse-to-fine flow decoding, and a unified objective that couples flow supervision, trajectory consistency, spatial regularization, and segmentation loss. Experiments on remote sensing benchmarks show that the proposed framework produces more structured and robust change representations while achieving state-of-the-art performance.
AgenticPD: 物理設計 QoR 最適化のためのステージ対応エージェント フレームワーク
物理設計の結果品質 (QoR) の最適化は難しく、コストがかかります。ある段階での選択は、後の段階で役立つこともあれば、悪影響を与えることもあります。各評価には、フロー全体を通してコストのかかる EDA を実行する必要があります。既存の方法では依然として最適化をフラットなパラメーター調整または LLM ベースのスクリプト生成タスクとして扱いますが、物理設計の QoR 最適化のためのステージ認識型エージェント フレームワークである AgenticPD を紹介します。 AgenticPD は、トライアルごとに完全なフローを再実行するのではなく、物理設計フローのステージ境界を中心に編成されており、ジャッジ エージェントが検索をナビゲートし、ステージ専門のエージェントがステージ ローカル ツールを使用して独自のステージ内でローカルな意思決定を行います。さらに、AgenticPD のエージェント ハーネスは、構造化された観察、実行履歴、およびエージェント コンテキスト管理を提供します。その結果、システムは前の中間状態から分岐し、チェックポイントを再利用して最適化手順を続行することができ、すべての候補がルート後のサインオフで評価されます。これらのベースライン全体で、AgenticPD はパワーとエリアでの競争力を維持しながら、強力なポストルート タイミングを実現します。
原文 (English)
AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization
Physical design quality-of-results~(QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages. Each evaluation requires a costly EDA run through the full flow. While existing methods still treat optimization as flat parameter tuning or a LLM-based script generation task, we present AgenticPD, a stage-aware agentic framework for physical design QoR optimization. Instead of re-running the full flow after every trial, AgenticPD is organized around the stage boundaries of the physical design flow, where a Judge Agent navigates the search and stage-specialized agents make local decisions within their own stage using stage-local tools. Additionally, the agent harness in AgenticPD provides structured observations, execution history, and agent context management. As a result, the system can branch from prior intermediate states and reuse checkpoints to continue the optimization procedure, and every candidate is evaluated at the post-route signoff. Across these baselines, AgenticPD achieves strong post-route timing while remaining competitive in power and area.
CARL: LLM を使用した計画のための制約認識強化学習
大規模言語モデル (LLM) は、強力な推論能力と広範な世界知識にもかかわらず、タスクの制約に違反する計画を頻繁に生成し、現実世界のアプリケーションにおける信頼性を損ないます。この欠陥は、生成プロセス中に制約情報を組み込む体系的なメカニズムが欠如しているために発生します。既存のアプローチは、外部ツールやタスク分解に依存してこれを軽減しようとしますが、モデルの本質的な制約認識を強化することはできません。これに対処するために、LLM の本質的な制約への焦点を強化するように設計された新しい RL フレームワークである制約認識強化学習 (CARL) を提案します。 CARL は、制約のある入力と制約のない入力の下でモデルの出力分布を比較することで、制約を意識した報酬を導入し、制約の焦点を奨励し、無視にペナルティを与えます。 CARL は、さまざまな RL フレームワークと互換性があり、外部ソルバーや最上位モデルを必要としないため、スケーラブルなエンドツーエンドの制約を意識した計画を可能にします。 BlocksWorld、TravelPlanner、および T-Eval に関する広範な実験により、CARL が標準の強化微調整 (RFT) ベースラインや最先端の推論モデルを大幅に上回り、制約への重点が著しく高まっていることが実証されました。
原文 (English)
CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs
Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications. This deficiency arises from a lack of systematic mechanisms to incorporate constraint information during the generation process. While existing approaches attempt to mitigate this by relying on external tools or task decomposition, they fail to enhance the model's intrinsic constraint awareness. To address this, we propose Constraint-Aware Reinforcement Learning (CARL), a novel RL framework designed to strengthen LLMs' intrinsic focus on constraints. CARL introduces a constraint-aware reward by comparing the model's output distributions under constrained and unconstrained inputs, encouraging constraint focus and penalizing neglect. Compatible with various RL frameworks and requiring no external solvers or top models, CARL enables scalable, end-to-end constraint-aware planning. Extensive experiments on BlocksWorld, TravelPlanner, and T-Eval demonstrate that CARL significantly outperforms standard Reinforcement Fine-Tuning (RFT) baselines and state-of-the-art reasoning models, exhibiting a markedly increased focus on constraints.
Medi-Gemma: 決定論的 EMR 分析と検索拡張生成を統合したハイブリッド臨床意思決定支援システム
一か八かの臨床現場での大規模言語モデル (LLM) の導入は、構造的幻覚、表形式の患者データに対する弱い決定論的推論、ベクトル検索の欠落によって依然として制限されています。この文書では、創傷病理トリアージおよびワークフロー自動化のための臨床意思決定支援システム (CDSS) である Medi-Gemma のアーキテクチャと検証について説明します。このプラットフォームは、追跡可能な推論を維持しながら、臨床認識をデータ オーケストレーションから分離する分離されたフレームワークを導入します。 Medi-Gemma は、集中型の ClinicalOrchestrator によって調整されるマルチステージ パイプラインを使用します。データ要求は、型強制によって非構造化電子医療記録 (EMR) ファイルをクリーンアップする DataManager によって生成推論なしで処理されます。自然言語クエリは階層型 IntentRouter によって処理され、PandasQueryEngine によって実行される決定論的分析パス、または CPU に最適化されたベクター ストアを使用して ClinicalRAGEngine によって管理される患者固有の推論にリクエストをルーティングします。主な貢献は、Ground Truth Injection Module です。このモジュールは、患者固有のクエリをインターセプトし、数値識別トークンを抽出し、Pandas 経由で構造化データフレームにクエリを実行し、最新の検証済みの臨床状態を取得し、生成前にこのスナップショットをオーバーライドするコンテキスト ブロックとして LLM プロンプトに埋め込みます。安全性コンプライアンスは、臨床用語を固定された証拠に基づくリスク経路にマッピングする決定論的な ProtocolManager によって強制され、SafetyVerifier フレーズ フィルターは出力ルール違反を防ぎます。検証の結果、このアーキテクチャによりセマンティック コンテキストのドリフトが排除され、データベース コンパイルのクラッシュが防止され、バックエンドの臨床リポジトリに対する事実の遵守が向上することがわかりました。これらの結果は、Medi-Gemma が、構造化されたデータの忠実性、検索の根拠、決定論的な安全対策が不可欠な LLM ベースの臨床意思決定支援のためのより安全なパターンであることを裏付けています。
原文 (English)
Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation
Deploying Large Language Models (LLMs) in high-stakes clinical settings remains limited by structural hallucinations, weak deterministic reasoning over tabular patient data, and omissions in vector retrieval. This paper presents the architecture and validation of Medi-Gemma, a Clinical Decision Support System (CDSS) for wound pathology triage and workflow automation. The platform introduces a decoupled framework that separates clinical perception from data orchestration while preserving traceable reasoning. Medi-Gemma uses a multi-stage pipeline coordinated by a centralized ClinicalOrchestrator. Data requests are handled without generative inference by a DataManager that cleans unstructured Electronic Medical Record (EMR) files through type coercion. Natural language queries are processed by a hierarchical IntentRouter, which routes requests to deterministic analytics paths executed by a PandasQueryEngine or to patient-specific reasoning managed by a ClinicalRAGEngine using a CPU-optimized vector store. A key contribution is the Ground Truth Injection Module, which intercepts patient-specific queries, extracts numeric identification tokens, queries the structured dataframe via Pandas, retrieves the latest validated clinical state, and embeds this snapshot as an overriding context block in the LLM prompt before generation. Safety compliance is enforced by a deterministic ProtocolManager that maps clinical terminology to fixed evidence-based risk pathways, while a SafetyVerifier phrase filter prevents output rule violations. Validation shows that this architecture eliminates semantic context drift, prevents database compilation crashes, and improves factual adherence to backend clinical repositories. These results support Medi-Gemma as a safer pattern for LLM-based clinical decision support where structured data fidelity, retrieval grounding, and deterministic safeguards are essential.
STAPO: LLM エージェント トレーニングのための選択的軌道認識ポリシーの最適化
強化学習 (RL) は、長期的なタスクで大規模言語モデル (LLM) エージェントをトレーニングするための主要なパラダイムです。ただし、報酬がまばらで遅れていると、エージェントが中間ステップでタスクの目標やインタラクション履歴に焦点を合わせることができなくなる、軌跡の無視につながることがよくあります。これまでの研究では、シャノンエントロピーベースの不確実性信号を使用したステップレベルの監視が検討されてきましたが、これは固有の状態の複雑さとエージェントの信頼性を混同し、したがって意思決定の信頼性の信頼性の低い推定値を提供します。この問題に対処するために、正規化エントロピーを提案します。これは、特定の状態におけるエージェントの平均的な行動と比較した信頼偏差を測定し、それによって低品質のアクションと軌道無視との関連を強化します。この洞察に基づいて、階層的なグループベースの RL フレームワークである Selective Trajectory-Aware Policy Optimization (STAPO) を紹介します。 STAPO は、正規化エントロピーを活用して、軌道無視に関連する外れ値ステップを特定し、軌道を意識した報酬と軌道に依存しないペナルティの共同メカニズムを介してそれらを最適化し、トレーニングの安定性を維持しながら軌道認識を強化します。 ALFWorld、WebShop、および Search-Augmented QA に関する広範な実験により、STAPO が軌道無視を大幅に軽減しながら最先端のパフォーマンスを達成することが実証され、エージェント タスクに対するその有効性と堅牢性が検証されました。
原文 (English)
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training
Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to trajectory neglect, in which agents lose focus on the task goal and interaction history at intermediate steps. Prior work has explored step-level supervision using Shannon-entropy-based uncertainty signals, which conflate inherent state complexity with agent confidence and therefore provide unreliable estimates of decision reliability. To address this issue, we propose normalized entropy, which measures confidence deviations relative to an agent's average behavior under a given state, thereby strengthening the association between low-quality actions and trajectory neglect. Building on this insight, we introduce Selective Trajectory-Aware Policy Optimization (STAPO), a hierarchical group-based RL framework. STAPO leverages normalized entropy to locate outlier steps associated with trajectory neglect and optimizes them via a joint mechanism of trajectory-aware reward and trajectory-independent penalty, enhancing trajectory awareness while preserving training stability. Extensive experiments on ALFWorld, WebShop, and Search-Augmented QA demonstrate that STAPO achieves state-of-the-art performance while substantially alleviating trajectory neglect, validating its effectiveness and robustness for agentic tasks.
量子にインスピレーションを得た調和決定モデル: 音楽生成のための計算フレームワーク
この論文では、音楽における調和の意思決定のための量子にインスピレーションを得た計算フレームワークを紹介します。提案されたアプローチは、構造化された組み合わせ空間内の最適化問題としてハーモナイゼーションを定式化し、複数の候補コードシーケンスが相互作用する音楽的制約の下で評価されます。このモデルは、干渉ベースのハーモナイゼーション段階と、音色ハーモニーに基づいた古典的な最適化手順を組み合わせています。量子にインスピレーションを得たコンポーネントにより、複数の倍音の代替案を並行して検討できるようになり、古典的な段階では、結果として得られるシーケンスを洗練して、構造の一貫性と文体の妥当性を確保します。このフレームワークは、「Autumn Leaves」や「It's a Long Way to Tipperary」など、選択された音楽例に基づいて評価されます。定量的分析により、最適化段階によりコード密度が大幅に減少し、倍音の安定性が向上し、機能構成が改善されることが示されています。同時に、専門家の評価は文体的なコンテキストの重要性を強調しており、倍音の複雑さの増加が常により自然であると認識されるわけではないことを示しています。この結果は、高調波の生成が、制約された探索空間における構造化された意思決定プロセスとして解釈できることを示唆しています。提案されたアプローチは、ドメイン固有の知識と干渉ベースの検索メカニズムを統合する計算モデルを提供します。予備的ではあるが、この研究は、量子にインスピレーションを得た手法が、音楽などの創造的な領域における複雑な意思決定プロセスをモデル化するための有用なフレームワークを提供する可能性があることを示している。提案されたフレームワークは、複雑な生物学的および創造的システムにおける認知と意思決定の量子にインスピレーションを得たモデルに関する進行中の研究に貢献します。
原文 (English)
Quantum-Inspired Harmonic Decision Models: A Computational Framework for Music Generation
This paper introduces a quantum-inspired computational framework for harmonic decision-making in music. The proposed approach formulates harmonization as an optimization problem within a structured combinatorial space, where multiple candidate chord sequences are evaluated under interacting musical constraints. The model combines an interference-based harmonization stage with a classical optimization procedure grounded in tonal harmony. The quantum-inspired component enables the parallel consideration of multiple harmonic alternatives, while the classical stage refines the resulting sequences to ensure structural coherence and stylistic plausibility. The framework is evaluated on selected musical examples, including Autumn Leaves and It's a Long Way to Tipperary. Quantitative analysis shows that the optimization stage significantly reduces chord density, increases harmonic stability, and improves functional organization. At the same time, expert evaluation highlights the importance of stylistic context, demonstrating that increased harmonic complexity is not always perceived as more natural. The results suggest that harmonic generation can be interpreted as a structured decision-making process in a constrained search space. The proposed approach provides a computational model that integrates domain-specific knowledge with an interference-based search mechanism. Although preliminary, this work indicates that quantum-inspired methods may offer a useful framework for modeling complex decision processes in creative domains such as music. The proposed framework contributes to ongoing research on quantum-inspired models of cognition and decision-making in complex biological and creative systems.
医療における信頼できる大規模言語モデル エージェントを目指して
医療予約のスケジュール設定は、手動による調整、断片化されたレガシー システム、および高い管理オーバーヘッドによって引き起こされ、依然として運用上のボトルネックとなっています。これらの非効率性により、医療提供者の可用性が制限され、患者のケアへのアクセスが低下します。このペーパーでは、大規模言語モデル (LLM) 関数呼び出し、検索拡張生成 (RAG)、および階層化された決定論的安全ガードレールを活用する、医療物流自動化のための安全第一の会話エージェント CareConnect について説明します。このシステムは、8 つのドメイン固有のツールを統合して、予約、変更、キャンセル、施設情報の取得をサポートすると同時に、医学的アドバイスや診断を禁止する厳格な範囲制限を適用します。安全性が重要な状況は、緊急検出と医療目的の拒否のための決定論的な短絡メカニズムを通じて処理されます。私たちは、エンドツーエンドのワークフロー、マルチターン インタラクション、エッジ ケースにわたる 680 のタスク指向シナリオの包括的なベンチマークに基づいて CareConnect を評価します。実験結果では、リクエストあたりのレイテンシの中央値が 2.2 秒で 91.8% のタスク完了率、専用の安全性クリティカル評価サブセットでの 96.0% の安全性遵守、アポイントあたりの平均運用コストが 0.0324 ドルであることが実証されており、人間による手動スケジューリングと比較して大幅なコスト削減が得られます。これらの調査結果は、慎重に範囲を絞り、厳密に保護された LLM ベースのエージェントが、安全性の保証を維持し、大幅なコスト効率を達成しながら、複雑な医療運用ワークフローを確実に自動化できることを示しています。ソース コードとシステム実装は、https://github.com/Hadi-Hsn/CareConnect で公開されています。
原文 (English)
Toward Trustworthy Large Language Model Agents in Healthcare
Healthcare appointment scheduling remains a persistent operational bottleneck, driven by manual coordination, fragmented legacy systems, and high administrative overhead. These inefficiencies constrain provider availability and degrade patient access to care. This paper presents CareConnect, a safety-first conversational agent for healthcare logistics automation that leverages large language model (LLM) function calling, retrieval-augmented generation (RAG), and layered deterministic safety guardrails. The system orchestrates eight domain-specific tools to support appointment booking, modification, cancellation, and facility information retrieval, while enforcing strict scope constraints that prohibit medical advice or diagnosis. Safety-critical situations are handled through deterministic short-circuit mechanisms for emergency detection and medical intent refusal. We evaluate CareConnect on a comprehensive benchmark of 680 task-oriented scenarios spanning end-to-end workflows, multi-turn interactions, and edge cases. Experimental results demonstrate a 91.8% task completion rate with a median per-request latency of 2.2 seconds, 96.0% safety compliance on the dedicated safety-critical evaluation subset, and an average operational cost of $0.0324 per appointment, yielding a significant cost reduction compared to manual human scheduling. These findings show that carefully scoped and rigorously safeguarded LLM-based agents can reliably automate complex healthcare operational workflows while maintaining safety guarantees and achieving substantial cost efficiency. The source code and system implementation are publicly available at https://github.com/Hadi-Hsn/CareConnect.
拡散に基づく不確実性を認識した遅延ポリシーの最適化
現実世界の環境における強化学習では、フィードバックの遅延により、パフォーマンスが大幅に低下することがよくあります。既存のアプローチは通常、拡張状態を構築したり真の状態を予測したりすることで、観測の遅延によって引き起こされるパフォーマンスの低下を軽減します。ただし、これらの方法では、確率的 MDP によって引き起こされる遅延状態と真の状態の間の固有の不一致が見落とされることがよくあります。我々は、このような矛盾の存在を理論的に証明し、それが最適な政策の劣化につながることを示します。この課題に対処するために、私たちは拡散誘導型不確実性認識遅延ポリシー最適化 (DUPO) を提案します。私たちの方法では、拡散モデルを使用して遅延状態メッセージと現在の状態の関係を明示的にモデル化し、結果として得られる不一致推定を利用して遅延ポリシーに重み付けを行います。複数の確率的遅延を伴う連続ロボット制御タスクに関する広範な実験により、DUPO が一貫して既存の手法を上回り、長時間かつランダムな遅延シナリオ下でも効果を維持できることが実証されました。
原文 (English)
Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization
Reinforcement learning in real world environments often suffers from severe performance degradation due to delayed feedback. Existing approaches typically mitigate performance degradation caused by observation delays by constructing augmented states or predicting the true states. However, these methods often overlook the inherent discrepancy between delayed state and true states induced by stochastic MDP. We theoretically prove the existence of such a discrepancy and show that it leads to the degradation of the optimal policy. To address this challenge, we propose Diffusion Guided Uncertainty Aware Delayed Policy Optimization (DUPO). Our method explicitly models the relationship between delayed state message and the current state using a diffusion model, and leverages the resulting discrepancy estimates to weight delayed policies. Extensive experiments on continuous robotic control tasks with multiple stochastic delays demonstrate that DUPO consistently outperforms existing methods and remains effective even under long and random delay scenarios.
ASSEMCAD: 自然言語から本番環境に対応した CAD アセンブリを生成
大規模言語モデルとプログラム CAD の最近の進歩により、個々のパーツの Text-to-CAD 生成が大幅に改善されました。ただし、生産準備の整った機械アセンブリの生成については、ほとんどが未解決のままです。単一部品のモデリングとは異なり、アセンブリでは、複数のコンポーネント、機能インターフェイス、アセンブリ関係、エンジニアリング原則、および物理的一貫性にわたって調整された推論が必要です。したがって、実行可能な CAD コードを直接生成するだけでは、機械的に有効で再利用可能なアセンブリを構築するには不十分です。私たちは、自然言語から生産対応の CAD アセンブリを生成するための、公理に基づいたフレームワークである AssemCAD を紹介します。アセンブリをモノリシック CAD コードとして表現する代わりに、AssemCAD はまず、型指定された部品、ジオメトリに基づいたポート、実行可能な合致、およびエンジニアリング公理から構成される公理的なアセンブリ仕様を構築します。各アセンブリ関係は 1 つ以上のエンジニアリング原則に明示的に基づいており、結果として得られる仕様は解釈可能、再利用可能、検証可能になっています。この仕様を実現するために、AssemCAD はポートおよび合致ベースの CAD アセンブリ ライブラリを導入しています。これは、決定論的な合致変換を通じて記号的なアセンブリ関係を実行し、具体的な B-Rep 幾何学的証拠を使用して宣言されたインターフェイスを検証します。この表現とライブラリに基づいて構築された AssemCAD は、標準ジオメトリとオープンワールド ジオメトリの両方について、再利用可能なパラメトリック コンポーネント ファクトリのオンデマンド合成をさらにサポートします。 AssemBench での実験では、AssemCAD がコード中心の CAD 生成ベースラインに比べてアセンブリの保存と物理的妥当性を大幅に向上させながら、さまざまな基礎モデルのバックボーン全体で汎用化できることが示されています。 AssemCAD は、公理に基づいたアセンブリ推論と決定論的な幾何学的実行を組み合わせることで、Text-to-CAD を分離部品の生成から生産準備の整った機械アセンブリ設計まで拡張します。
原文 (English)
ASSEMCAD: Production-Ready CAD Assembly Generation from Natural Language
Recent advances in large language models and programmatic CAD have significantly improved Text-to-CAD generation for individual parts. However, production-ready mechanical assembly generation remains largely unsolved. Unlike single-part modeling, assemblies require coordinated reasoning over multiple components, functional interfaces, assembly relations, engineering principles, and physical consistency. Consequently, directly generating executable CAD code is insufficient for constructing mechanically valid and reusable assemblies. We present AssemCAD, an axiom-grounded framework for production-ready CAD assembly generation from natural language. Instead of representing an assembly as monolithic CAD code, AssemCAD first constructs an axiomatic Assembly Specification consisting of typed parts, geometry-backed ports, executable mates, and engineering axioms. Each assembly relation is explicitly grounded in one or more engineering principles, making the resulting specification interpretable, reusable, and verifiable. To realize this specification, AssemCAD introduces a port- and mate-based CAD assembly library that executes symbolic assembly relations through deterministic mate transformations and validates declared interfaces using concrete B-Rep geometric evidence. Built on this representation and library, AssemCAD further supports on-demand synthesis of reusable parametric component factories for both standard and open-world geometries. Experiments on AssemBench show that AssemCAD substantially improves assembly preservation and physical validity over code-centric CAD generation baselines, while generalizing across different foundation-model backbones. By combining axiom-grounded assembly reasoning with deterministic geometric execution, AssemCAD extends Text-to-CAD from isolated part generation toward production-ready mechanical assembly design.
TacReasoner: 現実世界のシナリオにおける対話型推論のための動的触覚言語フレームワーク
人間の 5 つの主要な感覚の中で、触覚はおそらく、現実世界の環境での物理的な接触と相互作用の認識を可能にするため、生存にとって最も基本的なものです。この論文では、マルチモーダル推論のためのインテリジェント システムに触覚センシングを統合する際の 2 つの重要な課題を検討します。(i) 動的触覚信号の不十分なモデリングにより、時間的に変化する特性に対する推論が制限されます。(ii) 明示的な推論メカニズムの欠如によって引き起こされ、不安定な現実世界の推論につながる触覚基礎モデルの幻覚です。これらの課題に対処するために、私たちは、現実世界のシナリオで対話型推論を行うための動的な触覚言語フレームワークである TacReasoner を提案します。まず、TacReasoner には動的触覚信号の知覚と表現を強化するために動的認識触覚エンコーダーが組み込まれています。さらに重要なのは、触覚入力に対する構造化推論のための初の触覚思考連鎖データセットである TouchCoT-10k を紹介することです。それに基づいて、動的触覚知覚と現実世界の常識的推論を体系的に評価するDynTac-Benchを確立します。実験結果は、TacReasoner が複数のデータセットにわたって最先端のモデルに対して競争力のあるパフォーマンスを達成することを示しています。特に、TacReasoner は 7B パラメータのみを使用しているにもかかわらず、ほとんどのサブタスクで 14B VTV-LLM モデルを上回っており、触覚的常識推論における有効性と効率性を強調しています。
原文 (English)
TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios
Among the five primary human senses, tactile is arguably the most fundamental to survival, as it enables the perception of physical contact and interaction in real-world environments. In this paper, we explore two key challenges of integrating tactile sensing into intelligent systems for multimodal reasoning: (i) insufficient modeling of dynamic tactile signals, which restricts reasoning over temporally evolving properties, and (ii) hallucination in tactile foundation models caused by the absence of explicit reasoning mechanisms, leading to unstable real-world inference. To address these challenges, we propose TacReasoner, a dynamic tactile-language framework for interactive reasoning in real-world scenarios. First, TacReasoner incorporates a Dynamic-aware Tactile Encoder to enhance the perception and representation of dynamic tactile signals. More importantly, we introduce TouchCoT-10k, the first tactile chain-of-thought dataset for structured reasoning over tactile inputs. Upon it, we establish DynTac-Bench to systematically evaluate dynamic tactile perception and real-world commonsense reasoning. Experimental results demonstrate that TacReasoner achieves competitive performance against state-of-the-art models across multiple datasets. Notably, despite using only 7B parameters, TacReasoner outperforms the 14B VTV-LLM model on most subtasks, highlighting its effectiveness and efficiency in tactile commonsense reasoning.
DSpark: 半自己回帰生成を使用した、信頼度に基づいてスケジュールされた投機的デコーディング
投機的デコードは、ドラフトの生成をターゲットの検証から切り離すことで、大規模言語モデル (LLM) の推論を高速化します。最近の並列ドラフト作成者は、単一のフォワード パスで長いトークン シーケンスを効率的に提案しますが、トークン間の依存関係が欠如しているため、受け入れが急速に低下するという問題があります。さらに、これらの拡張ブロックを無差別に検証すると、拒否リスクの高いトークンの重要なバッチ容量が無駄になり、同時実行性の高いサービス システムのスループットが大幅に低下します。高スループットの並列生成と適応型の負荷認識検証を統合する投機的デコード フレームワークである DSpark を紹介します。ドラフトの品質を維持するために、DSpark は半自己回帰アーキテクチャを利用し、並列バックボーンと軽量のシーケンシャル モジュールを結合して、ブロック内依存関係モデリングを導入し、サフィックスの減衰を軽減します。システム効率を最適化するために、DSpark は信頼性スケジュール検証を採用し、推定されたプレフィックス生存確率とエンジン固有のスループット プロファイルに基づいて各リクエストの検証長さを動的に調整します。さまざまなドメインにわたるオフライン ベンチマークでは、DSpark は、最先端の自己回帰および並列ドラフターよりも許容される長さを大幅に向上させます。 DSpark は、ライブ ユーザー トラフィックの下で DeepSeek-V4 サービス システム内に展開されると、検証の無駄を軽減することに成功します。確立された運用ベースライン (MTP-1) と比較して、DSpark は、一致するスループット レベルでユーザーごとの生成速度を 60 ~ 85 パーセント高速化します。さらに重要なことは、厳しい対話性制約の下で深刻なスループットの低下を防ぐことで、これまで達成できなかったパフォーマンス層を実現し、サービス システムのパレート フロンティアをシフトさせることです。
原文 (English)
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity on tokens with high rejection risks, severely degrading throughput in high-concurrency serving systems. We introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. To maintain draft quality, DSpark utilizes a semi-autoregressive architecture, coupling a parallel backbone with a lightweight sequential module, to introduce intra-block dependency modeling and mitigate suffix decay. To optimize system efficiency, DSpark employs confidence-scheduled verification, dynamically tailoring the verification length for each request based on estimated prefix survival probabilities and engine-specific throughput profiles. On offline benchmarks across diverse domains, DSpark substantially improves the accepted length over state-of-the-art autoregressive and parallel drafters. When deployed within the DeepSeek-V4 serving system under live user traffic, DSpark successfully mitigates verification waste. Compared to the established production baseline (MTP-1), DSpark accelerates per-user generation speeds by 60 to 85 percent at matched throughput levels. More importantly, by preventing severe throughput degradation under strict interactivity constraints, it enables performance tiers that were previously unattainable, shifting the Pareto frontier of our serving system.
人工知能における記号的手法の役割の変化
なぜインテリジェントシステムは明示的な記号推論を実行する必要があるのでしょうか?コンピューターサイエンスは伝統的に、象徴的推論を知能の定義要素とみなしてきました。しかし、現代の基礎モデルの目覚ましい成功は、根本的な疑問を引き起こしています。能力が高まる AI システムが、明示的な記号的推論をほとんど使わずに動作できるとしたら、記号的手法は実際にどのような役割を果たしているのでしょうか?この記事は、明示的な記号的推論は知能の基本的な特性ではなく、現実の単純化されたモデルに基づいて動作する計算上の結果であると主張します。私たちは圧縮原理を提案します。すべての計算モデルは現実を簡略化して表現しており、明示的な記号推論によってモデル構築中に省略された情報が補われます。この原則から、モデル化と推論のトレードオフを導き出します。つまり、計算モデルが世界のより豊かな表現を保持するため、明示的な記号推論の必要性がそれに応じて減少します。この観点は、シンボリック手法の歴史的な成功と現代の基礎モデルの顕著な有効性の両方について統一的な説明を提供します。逆説的ですが、同じ発展により、人間にとって象徴的な方法がますます重要になります。インテリジェント システムの機能が向上し、不透明になるにつれて、人間が要件を指定し、動作を検証し、自律システムを制御し、信頼を確立するためのインターフェイスとして、シンボリック表現がますます機能します。したがって、私たちは、シンボリック手法の将来は主にインテリジェント システムの計算エンジンとしてではなく、ますます能力が向上する AI システムと、それを構築し、管理し、依存する人間との間のシンボリック インターフェイスにあると主張します。
原文 (English)
The Changing Role of Symbolic Methods in Artificial Intelligence
Why do intelligent systems need to perform explicit symbolic reasoning? Computer science has traditionally regarded symbolic reasoning as a defining component of intelligence. Yet the remarkable success of modern foundation models raises a fundamental question: if increasingly capable AI systems can operate with little explicit symbolic reasoning, what role do symbolic methods actually play? This article argues that explicit symbolic reasoning is not a fundamental property of intelligence, but a computational consequence of operating on simplified models of reality. We propose the Compression Principle: every computational model is a simplified representation of reality, and explicit symbolic reasoning compensates for information omitted during model construction. From this principle, we derive the Modeling--Reasoning Trade-off: as computational models preserve richer representations of the world, the need for explicit symbolic reasoning correspondingly decreases. This perspective provides a unified explanation for both the historical success of symbolic methods and the remarkable effectiveness of modern foundation models. Paradoxically, the same development makes symbolic methods increasingly important for humans. As intelligent systems become more capable and more opaque, symbolic representations increasingly serve as interfaces through which humans specify requirements, verify behavior, regulate autonomous systems, and establish trust. We therefore argue that the future of symbolic methods lies not primarily as the computational engine of intelligent systems, but as the symbolic interface between increasingly capable AI systems and the humans who build, govern, and depend upon them.
AgentGym2: 非理想化された現実世界環境における大規模言語モデル エージェントのベンチマーク
言語エージェント、つまり LLM エージェントは急速に進歩しており、運用環境に導入されるケースが増えています。この傾向は、厳密かつ現実的な評価が緊急に必要であることを浮き彫りにしています。ただし、既存のベンチマークのほとんどは、簡素化された理想的な設定でエージェントを評価します。彼らは通常、事前にパッケージ化されたツールのインターフェイスに依存し、重要な手順を見落とし、入力がクリーンで完全に指定されていると想定します。その結果、不確実性とノイズが遍在し、エージェントが新しいツールを発見するために環境を積極的に探索する必要がある実際の導入の難しさを過小評価しています。このギャップを埋めるために、現実世界のエンドツーエンドの作業要求に基づいたタスク インスタンスを備えた新しい評価フレームワークである AgentGym2 を紹介します。推論と計画を超えて、エンドツーエンドの手順の実行、探索によるツールの発見、目に見えないタスク用のツールの作成、ノイズの多い情報や不明確な情報に対する堅牢性を維持するエージェントの能力を測定します。 15 の独自モデルとオープンソース モデルの実験では、Gemini や GPT-5 などの SOTA システムでさえ AgentGym2 では苦戦していることが示され、現在のエージェントの能力と現実世界のアプリケーションの要求との間に大きなギャップがあることが明らかになりました。
原文 (English)
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for rigorous and realistic evaluations. However, most existing benchmarks evaluate agents in simplified, idealized settings. They typically rely on pre-packaged tool interfaces, overlook critical steps, and assume inputs are clean and fully specified. Consequently, they understate the difficulty of real deployments, where uncertainty and noise are ubiquitous and agents must proactively explore the environment to uncover new tools. To bridge this gap, we present AgentGym2, a new evaluation framework with task instances grounded in real-world end-to-end working demands. Beyond reasoning and planning, it measures agents' ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information. Experiments on 15 proprietary and open-source models show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2, revealing a substantial gap between the capability of current agents and the demands of real-world applications.
CP-WSP: 構成可能な複数制約の従業員スケジュールのための宣言型 CP-SAT フレームワーク
従業員のスケジューリングは、労働規制、補償要件、従業員の好み、業務目標を同時に満たす必要がある NP 困難な組み合わせ最適化問題です。既存の CP 定式化は通常、シフト レベルの粒度で 6 ~ 12 個の制約を持つ簡略化されたインスタンスをモデル化しており、以下に対する明示的なサポートが決定的に欠如しています。 中間点配置制御による必須の休憩スケジュール。緊急度に重み付けされたワークロードの公平性。サブシフトの時間粒度により、需要に応じた人員配置が可能になります。週間のスケジュールの安定性。 24 時間営業で一般的な深夜をまたぐシフト パターン。このホワイトペーパーでは、CP-WSP について説明します。宣言型 CP-SAT フレームワークは、数学的に不可侵な要件 (建設による規制違反ゼロ) として 14 のハード制約を強制すると同時に、統一された重み付きペナルティ関数を通じて 15 のソフト目標を最適化します。これらはすべて、コード変更を必要とせずに JSON 仕様を介して構成可能です。主な貢献には以下が含まれます。 中心性制御による必須の休憩スケジュールを可能にするシフト ウィンドウ変数分解。緊急度に重み付けされたワークロードの公平性。 30 分から 2 時間までの多粒度の時間分解能。週間のスケジュールの安定性。深夜をまたぐシフトのためのグリッドオフセット前処理技術。コミュニティ比較のための再現可能な 36 構成のベンチマーク スイート。 INRC-II ベンチマークで、時間単位とシフトレベルの両方の粒度で、および 36 の合成構成で評価されました。
原文 (English)
CP-WSP: A Declarative CP-SAT Framework for Configurable Multi-Constraint Workforce Scheduling
Workforce scheduling is an NP-hard combinatorial optimization problem requiring simultaneous satisfaction of labor regulations, coverage requirements, employee preferences and operational objectives. Existing CP formulations typically model simplified instances with 6-12 constraints at shift-level granularity and critically lack explicit support for: mandatory break scheduling with midpoint placement control; acuity weighted workload equity; sub-shift temporal granularity enabling demand-driven staffing; inter-week schedule stability; and cross-midnight shift patterns common in 24-hour operations. This paper presents CP-WSP: a declarative CP-SAT framework enforcing 14 hard constraints as mathematically inviolable requirements (zero regulatory violations by construction) while optimizing 15 soft objectives through a unified weighted penalty function -- all configurable via a JSON specification with no code changes required. Key contributions include: a shift-window variable decomposition enabling mandatory break scheduling with centrality control; acuity-weighted workload equity; multi-granularity temporal resolution from 30 minutes to 2 hours; inter-week schedule stability; a grid-offset preprocessing technique for cross-midnight shifts; and a reproducible 36-configuration benchmark suite for community comparison. Evaluated on INRC-II benchmarks at both hourly and shift-level granularity and on 36 synthetic configurations.
思考モデルのためのポリシー上の自己蒸留を再考する
自己蒸留は、言語モデルにおける自己改善の有望なレシピです。この設定では、数学の問題の解決策などの特権情報が与えられた場合、モデルは独自の教師として機能します。これは、テスト時の推論を使用して特権情報を吸収できる思考モデルにとって特に魅力的であると思われます。驚くべきことに、特権的自己蒸留により、長い推論トレースで思考モデルが劣化することがわかりました。AIME24、AIME25、および HMMT25 で評価された 5 つの Qwen3 および OLMo 思考モデル全体で、特権的コンテキスト蒸留により、平均精度が最大 17% 相対的に低下します。この低下は、学生に与えられなかった特権的なコンテキストの量に応じて拡大し、展開予算が長い場合に最も顕著になります。そうでない場合、思考モデルは最大の利益を得ることができます。この失敗モードは自己蒸留に特有のものではありません。オンポリシー蒸留 (OPD) は思考モデルを改善しますが、特権付き OPD はこれらの利益を逆転させます。私たちの診断では、この失敗モードを、教師の特権的コンテキストが高エントロピー分岐位置での学習をどのように再形成するかに関連付けています。分岐位置では、複数の継続がもっともらしいままであり、異なる推論パスにつながる可能性があります。特権コンテキストは思考モデルのロールアウトではフォーク率を低下させますが、命令モデルのロールアウトではそうではありません。これは、特権コンテキストが命令調整モデルには役立つが、より強力な思考モデルには害を及ぼすという興味深い二分法につながります。この効果は、学生が自己修正ブランチを開始すると目に見えます。そこでは、特権付き OPD が、バニラ OPD がサポートするサンプリングされた再検討トークンにペナルティを与えます。特権的な教師によってトレーニングされた思考モデルは、長さの正規化後でも、検証、バックトラッキング、ヘッジ マーカーの生成が少なくなります。これらの発見は、強力な思考モデルの自己蒸留には、特に修正と推論のステップに関して、トークンレベルの信号に注意を払う必要があることを示しています。
原文 (English)
Rethinking On-Policy Self-Distillation for Thinking Models
Self-distillation is a promising recipe for self-improvement in language models. In this setting, a model can serve as its own teacher when given privileged information, such as a solution to a math problem. This seems especially appealing for thinking models, which can use test-time reasoning to absorb the privileged information. Surprisingly, we show that privileged self-distillation degrades thinking models on long reasoning traces: across five Qwen3 and OLMo thinking models evaluated on AIME24, AIME25, and HMMT25, privileged-context distillation causes a relative drop of up to 17% in avg@16 accuracy. The degradation scales with the amount of privileged context withheld from the student and is most pronounced at long rollout budgets, where thinking models otherwise obtain their largest gains. This failure mode is not specific to self-distillation: on-policy distillation (OPD) improves thinking models, but privileged OPD reverses these gains. Our diagnostics link this failure mode to how privileged teacher context reshapes learning at high-entropy forking positions, where multiple continuations remain plausible and may lead to different reasoning paths. Privileged context lowers fork rates in thinking-model rollouts but not in instruction-model rollouts. This leads to an interesting dichotomy, where privileged context can help instruction-tuned models but hurts stronger thinking models. The effect is visible when the student begins a self-correction branch, where privileged OPD penalizes sampled reconsideration tokens that vanilla OPD supports. Thinking models trained with a privileged teacher produce fewer verification, backtracking, and hedging markers, even after length normalization. These findings indicate that self-distillation for strong thinking models requires attention to token-level signal, especially around correction and reasoning steps.
ClassicLogic: 構成の一般化を評価するための古典的なパズル ゲームの知識主導型ベンチマーク
構成の一般化、つまり既知のコンポーネントの新しい組み合わせを理解して生成する能力は、現代の人工知能にとって依然として根本的な課題です。ベンチマークはほとんど存在しませんが、多くは言語的なタスクに焦点を当てており、複雑で明示的な構成構造が欠けています。 ClassicLogic は、問題解決戦略を学習して構築するエージェントの能力を評価するために設計された新しいベンチマーク スイートです。ベンチマークは、Sudoku、KenKen、Kakuro、Futoseki の 4 つの古典的な論理パズルで構成されています。その核となるイノベーションは、各ゲームの階層的な明示的な知識ベースであり、複雑な解決戦略が、より単純な基本的な戦略の構成として正式に定義されます。この構造により、基本的なルールの学習から、数学的に検証された難易度の高いパズルを解くための複数ステップの構成戦略の適用まで、エージェントの推論能力をきめ細かく評価することができます。オープンソース ベンチマークは、神経記号およびその他の高度な AI 推論システムを進化させるための、挑戦的な新しいテストベッドを提供します。
原文 (English)
ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization
Compositional generalization, the ability to understand and produce novel combinations of known components, remains a fundamental challenge for modern artificial intelligence. While few benchmarks exist, many focus on linguistic tasks and lack complex, explicit compositional structures. We introduce ClassicLogic, a new benchmark suite designed to evaluate an agent's ability to learn and compose problem-solving strategies. The benchmark consists of four classic logic puzzles: Sudoku, KenKen, Kakuro, and Futoshiki. Its core innovation is a hierarchical, explicit knowledge base for each game, where complex solving strategies are formally defined as compositions of simpler, foundational strategies. This structure allows for fine-grained evaluation of an agent's reasoning capabilities, from learning basic rules to applying multi-step compositional strategies to solve puzzles of increasing, mathematically validated difficulty. The open-source benchmark provides a challenging new testbed for advancing neuro-symbolic and other advanced AI reasoning systems.
理由、報酬、改善: 小さな言語モデルでの物理推論のための構造化フィードバックによるステップレベルのエラー修正
物理推論は、小規模な言語モデルでは構造的に失敗します。どのステップでもエラーが前方に伝播し、その後のすべての推論が破壊されます。限られた領域の知識、多段階の導出における幻覚、および分布の敏感さが、この失敗をさらに悪化させます。我々は、最初の推論エラーを特定し、ターゲットを絞った構造化フィードバックを生成し、生成ターゲットとしてグランドトゥルースソリューションに公開することなく、KL 正則化を使用したポリシー勾配を介してソリューションを修正するようにモデルをトレーニングする、ステップレベルの報酬フレームワークを提案します。アノテーションに依存するステップレベルのメソッドとは異なり、プリファレンス データの構築は必要なく、外部ベリファイアはトレーニング時にのみ動作します。 5 つの物理ベンチマーク全体で、当社のフレームワークは、CoT プロンプトに対して 17 ~ 20%、最も強力なベースラインに対して 10 ~ 16% の精度向上を実現し、観察された最良のケースで計算エラーを 56.9% から 23.5% に削減し、誤解エラーを 22.3% から 12.0% に削減しました。概念的なエラーは 89.7% から 68.7% に減少しますが、すべての条件にわたって最も困難な故障モードとして存続します。
原文 (English)
Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models
Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucination under multi-step derivation, and distributional sensitivity compound this failure. We propose a step-level reward framework that identifies the first reasoning error, generates targeted structured feedback, and trains the model to revise its solution via policy gradient with KL regularization, without exposing it to ground truth solutions as generation targets. Unlike annotation-dependent step-level methods, no preference data construction is required and the external verifier operates exclusively at training time. Across five physics benchmarks, our framework delivers accuracy gains of 17-20% over CoT prompting and 10-16% over the strongest baseline, reduces calculation errors from 56.9% to 23.5%, and reduces miscomprehension errors from 22.3% to 12.0% in the best observed cases. Conceptual errors reduce from 89.7% to 68.7%, yet persist as the hardest failure mode across all conditions.
EvoAgentBench: 能力移転によるエージェントの自己進化のベンチマーク
長期的な LLM システムにおけるエージェントの自己進化は、主に手続き型です。有用なエクスペリエンスは、単に保存された情報ではなく、検索、デバッグ、検証のための再利用可能な手順です。しかし、現在の評価では、この形態の移転が孤立しているわけではありません。エージェントのベンチマークは、単一エピソードのタスク解決をテストします。メモリ ベンチマークは、手順による再利用ではなく、情報の保持を対象としています。 EvoAgentBench は、Web 調査、アルゴリズム推論、ソフトウェア エンジニアリング、ナレッジ ワークの 4 つのエージェント ドメインにわたる能力ガイドによる転送によるエージェントの自己進化のベンチマークです。 EvoAgentBench は、エージェントの実行からトレースに基づいたアビリティを抽出し、それらを操作単位に正規化し、手順の重複を共有するタスクをリンクするドメイン固有のアビリティ グラフを構築します。設計上、すべてのテスト タスクは、検証済みのトレーニング側のアビリティ サポートによってサポートされています。 528/267 のトレーニング/テスト分割、2 つのスキャフォールド、および 3 つのバックボーンにわたって、厳選されたアビリティ コンテンツはモデル ファミリ間で確実に転送されますが、現在の自動メソッドですべての設定でプラスのゲインを維持できるものはありません。 EvoAgentBench は、自己進化の評価を、総合的な精度の比較から、エクスペリエンスのエンコード、ルーティング、取り込みの詳細な診断に移行します。このベンチマークは、https://huggingface.co/datasets/EverMind-AI/EvoAgentBench で公開されています。
原文 (English)
EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer
Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification. Yet current evaluations do not isolate this form of transfer. Agent benchmarks test single-episode task solving; memory benchmarks target information retention rather than procedural reuse. We introduce EvoAgentBench, a benchmark for agent self-evolution via Ability-guided transfer across four agentic domains: web research, algorithmic reasoning, software engineering, and knowledge work. EvoAgentBench extracts trace-grounded Abilities from agent executions, canonicalizes them into operational units, and builds domain-specific Ability Graphs linking tasks that share procedural overlap. By design, every test task is backed by verified training-side Ability support. Across a 528/267 train/test split, two scaffolds, and three backbones, curated Ability content transfers reliably across model families, but no current automatic method sustains positive gain in all settings. EvoAgentBench shifts self-evolution evaluation from aggregate accuracy comparison to fine-grained diagnosis of experience encoding, routing, and uptake. The benchmark is publicly available at https://huggingface.co/datasets/EverMind-AI/EvoAgentBench.
MoP-JEPA: 確率的 JEPA 世界モデル用のハード割り当て予測子混合物
JEPA ワールド モデルは、潜在回帰によってトレーニングされた単一の決定論的予測子を使用して、次の潜在状態を予測します。環境が確率的である場合、これは構造的に失敗することを示します。分岐遷移では、回帰最適予測子は後続埋め込みの条件付き平均、つまりどの状態にも対応しない真の次の状態の間の点を出力します。我々は、決定論的でゲートされた専門家の混合予測器についてこの崩壊を証明し、MoP-JEPA のハード割り当てされた予測器が代わりに遷移分布の量子化器に収束することを証明します。つまり、後続モードごとに 1 つのヘッドであり、単一の前方パスで列挙可能であり、プランナが使用するインターフェイスです。リークのない評価による公式の OGBench オフライン データでは、単一予測子のロールアウトに関する計画のパフォーマンスは低く ($0.02$ ~ $0.09$ 成功)、一方、予測モードによる計画のパフォーマンスは最大 $0.85$ に達し、すべてのタスクで決定論的予測子、ゲート MoE および変分予測子を上回っています。多重予測評価はカバレッジ フリーローディングを引き起こすため、検証プロトコルはメソッドの一部です。入力に依存しないコードブック コントロール、シャッフル コンテキスト テスト、ルーター ゲートによる読み出し、遷移精度ガード、モデルが遷移グラフをブラインドで提案し、グランド トゥルースは結果をチェックするためだけに使用される検証済みルート基準です。この基準の下では、私たちの方法は 3 つの迷路すべて ($2$--$5\times$) で最も強力なソフト代替案を上回っており、プロトコルは、そのベースラインの生スコアの残りのギャップを、存在しない予測遷移を通るルートとして特定します。同じモデルが実際の環境で実行され、最も難しい迷路で公開されている OGBench ベースラインに対して 7 件中 2 位に位置します。 JEPA 世界モデルがそもそも計画できるかどうかは、マルチモーダルなダイナミクスによって決まります。予測子とハード割り当てを混合することは、最小限の検証可能な修正です。
原文 (English)
MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models
JEPA world models predict the next latent state with a single deterministic predictor trained by latent regression. We show that this fails structurally when the environment is stochastic: at a branching transition, the regression-optimal predictor outputs the conditional mean of the successor embeddings, a point between the true next states that corresponds to no state at all. We prove this collapse for deterministic and gated mixture-of-experts predictors, and prove that MoP-JEPA's hard-assigned predictors converge instead to a quantizer of the transition distribution: one head per successor mode, enumerable in a single forward pass, which is the interface a planner consumes. On official OGBench offline data with leak-free evaluation, planning over single-predictor rollouts performs poorly ($0.02$--$0.09$ success) while planning over our predicted modes reaches up to $0.85$, ahead of deterministic, gated-MoE, and variational predictors on every task. Because multi-prediction evaluation invites coverage freeloading, a verification protocol is part of the method: an input-agnostic codebook control, a shuffled-context test, router-gated readouts, transition-precision guards, and a verified-route criterion in which the model proposes its transition graph blind and ground truth is used only to check the result. Under this criterion our method outperforms the strongest soft alternative on all three mazes ($2$--$5\times$), and the protocol identifies the remaining gap in that baseline's raw scores as routes through predicted transitions that do not exist. The same model executes in the real environment, placing second of seven against the published OGBench baselines on the hardest maze. Multimodal dynamics decide whether a JEPA world model can plan at all; a mixture of predictors with hard assignment is a minimal and verifiable fix.
MetaSkill-Evolve: 2 つのタイムスケールのメタスキル進化による LLM エージェントの再帰的自己改善
最近の LLM エージェントは、ますます長期にわたる、無制限のタスクに取り組み、エージェントに提供される外部スキル、再利用可能な手順知識により、この機能がさらに拡張されます。ただし、手動で作成された固定的なスキルが最適であることはほとんどなく、エージェントが遭遇するタスクの多様性に適応することはできません。自己改善エージェントは、実行トレースから独自のスキル ファイルを書き換えることでこの問題に対処し、困難なベンチマークで有意義な利益をもたらします。しかし、そのような自己進化は非再帰的なままです。それはタスク スキル (エージェントが何を行うか) のみを向上させ、改善手順 (どのように改善するか) は一度作成され、固定されたままになります。エージェントのスキル向上を再帰的に行う 2 つのタイムスケール フレームワークである MetaSkill-Evolve を紹介します。すべてのブランチにはタスク スキル $s$ とブランチ ローカルのメタスキル $m=(\psi,\sigma,\alpha,\pi,\varepsilon)$ の両方が含まれており、その 5 つのコンポーネントが改善パイプラインの Analyzer、Retriever、Allocator、Proposer、および Evolver エージェントをパラメータ化します。タスク スキルは高速ループで進化しますが、メタスキルはそれ自体に適用される同じパイプラインの下で、追加のモデルや目標なしで低速ループで進化します。 5 つのパイプライン エージェントすべてが単一の凍結されたバックボーンを共有することで、MetaSkill-Evolve は 3 つのエージェント ベンチマーク (OfficeQA、SealQA、ALFWorld) でスキルなし、静的スキル、および単一レベルの進化ベースラインを上回り、未加工のバックボーンに比べてそれぞれ +23.54、+16.09、+1.92 ポイント向上しました。
原文 (English)
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
Recent LLM agents tackle increasingly long-horizon, open-ended tasks, and external skills, reusable procedural knowledge supplied to the agent, further extend this capability. However, a fixed, hand-authored skill is rarely optimal, and cannot adapt to the diversity of tasks an agent encounters. Self-improving agents address this by rewriting their own skill files from execution traces, yielding meaningful gains on challenging benchmarks. Yet such self-evolution remains non-recursive: it improves only the task skill (what the agent does) while the improvement procedure (how it improves) is authored once and held fixed. We introduce MetaSkill-Evolve, a two-timescale framework that makes agentic skill improvement recursive: every branch carries both a task skill $s$ and a branch-local meta-skill $m=(\psi,\sigma,\alpha,\pi,\varepsilon)$ whose five components parameterise the Analyzer, Retriever, Allocator, Proposer, and Evolver agents of the improvement pipeline. Task skills evolve on a fast loop while the meta-skill evolves on a slower one under the same pipeline applied to itself, with no additional model or objective. With all five pipeline agents sharing a single frozen backbone, MetaSkill-Evolve outperforms no-skill, static-skill, and single-level evolution baselines on three agentic benchmarks (OfficeQA, SealQA, ALFWorld), improving held-out test accuracy over the raw backbone by +23.54, +16.09, and +1.92 points respectively.
医療ビジョン言語モデルのモデル編集の評価と理解
モデル編集は、コストのかかる再トレーニングを行わずに、医療用視覚言語モデル (VLM) の展開後の間違いを修正するための、迅速かつ的を絞った方法を約束します。ただし、既存のマルチモーダル モデル編集ベンチマークは汎用タスクに焦点を当てており、現実的な臨床領域の要件と変動性を反映していません。これに対処するために、画像とテキストのバリエーション、モダリティとプロトコルの変化、臨床知識の構成、および時間的進行という課題の下で、編集の信頼性、正確さ、一般化可能性が維持されているかどうかを評価する、マルチモーダル モデル編集の臨床に基づいたベンチマークである M3Bench を導入します。 M3Bench には、さまざまな解剖学、モダリティ、専門分野にわたる 16,276 の質問が含まれており、単一編集と連続編集の両方をサポートしています。 6 つの医療および一般 VLM にわたって 4 人の代表的な編集者を評価したところ、すべての基準において優れた方法はないことがわかりました。勾配ベースのエディターは強力な転送を実現しますが、壊滅的な局所性違反に悩まされます。一方、メモリベースの方法は局所性を維持しますが、構成の一般性に欠け、バックボーンに依存するハイパーパラメーターの感度が高くなります。さらに、これらの失敗の原因は、VLM の潜在空間ジオメトリと、さまざまな編集方法によってその景観がどのように変化するかであると考えられます。全体として、M3Bench はマルチモーダル モデル編集のための厳格な臨床ストレス テストを確立し、展開後の適応をより安全にするための実用的なガイダンスを提供します。このベンチマークは https://github.com/BioMed-AI-Lab-U-Michgan/M3Bench で公開されています。
原文 (English)
Evaluating and Understanding Model Editing for Medical Vision Language Models
Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks and do not reflect realistic clinical domain requirements and variability. To address this, we introduce M3Bench, a clinically grounded benchmark for multimodal model editing that evaluates whether an edit remains reliable, precise, and generalizable under the challenges of image and text variation, modality and protocol shifts, clinical knowledge composition, and temporal progression. M3Bench contains 16,276 questions spanning diverse anatomy, modalities, and specialties, and supports both single and sequential edits. By evaluating 4 representative editors across 6 medical and general VLMs, we find that no method excels across all criteria. Gradient-based editors achieve strong transfer but suffer from catastrophic locality violations, whereas memory-based methods preserve locality but lack compositional generality and exhibit high backbone-dependent hyperparameter sensitivity. We further attribute these failures to the latent space geometry of VLMs and how different editing methods shift its landscape. Overall, M3Bench establishes a rigorous clinical stress test for multimodal model editing and offers actionable guidance for safer post-deployment adaptation. The benchmark is publicly available at https://github.com/BioMed-AI-Lab-U-Michgan/M3Bench .
OptiAgent: マルチエージェントの反復改良によるエンドツーエンドの最適化モデリング
私たちは、オペレーションズ リサーチの問題の自然言語記述が与えられると、ソルバー対応の数式と実行可能コードを出力できるマルチエージェント フレームワークである OptiAgent を提案します。私たちのアーキテクチャは数学的モデリングのステップを優先し、専用のエージェントが決定変数や制約などの構造を抽出し、反復的な自己修正を可能にします。 4 つの特殊なフィードバック メカニズムを備えた新しいマルチループ検証アーキテクチャを導入します。各フィードバック メカニズムは、誤解、構造上の欠陥、数学的矛盾、検証の失敗、コード エラーなどの個別の故障モードを対象としています。精度に加えて、当社のモジュラー設計は透明性を高めることで最適化問題を解決するプロセスを改善します。これは、各エージェントが推論とフィードバックを公開し、完全なモデリング プロセスを監査可能にするためです。私たちのフレームワークは、LP、MILP、非線形計画法のタスクにわたる 4 つのベンチマークのうち 3 つで最先端のパフォーマンスを達成しながら、残りのデータセットでは高い競争力を維持します。
原文 (English)
OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement
We propose OptiAgent, a multi-agent framework that, given a natural language description of an Operations Research problem, is able to output a solver-ready mathematical formulation as well as executable code. Our architecture prioritizes the mathematical modeling step, where dedicated agents extract structures, such as decision variables and constraints, enabling iterative self-correction. We introduce a novel multi-loop validation architecture with four specialized feedback mechanisms, each targeting a distinct failure mode such as misinterpretation, structural defects, mathematical inconsistencies, validation failures, and code errors. Alongside accuracy, our modular design improves the process of solving optimization problems by improving transparency, as each agent exposes its reasoning and feedback, making the full modeling process auditable. Our framework achieves state-of-the-art performance on 3 out of 4 benchmarks across LP, MILP, and Nonlinear Programming tasks, while remaining highly competitive on the remaining dataset.
グラフ スパース サンプリング: 継続的な MDP 計画における地平線の呪いを打ち破る
連続領域における不確実性の下での計画は自律システムにとって不可欠ですが、計算量は多くかかります。モンテカルロ ツリー検索 (MCTS) などのツリーベースの検索方法は依然として人気がありますが、その分岐構造により、最悪の場合、先読みの深さとともに指数関数的に増大するサンプリング バジェットが必要になる可能性があります。ツリーの観点から見ると、プランナーは無限の分岐階層のどこを検索するかを決定する必要があるため、連続状態またはアクション空間は特に困難になります。我々は、候補アクションごとに個別の後続をサンプリングするのではなく、多くの候補決定間でサンプリングされた先物を共有するオンライン計画アルゴリズムであるグラフ スパース サンプリング (GSS) を提案します。この分岐のないグラフは、ヒューリスティックを使用して計算に焦点を当てながら、GPU に適した大規模なバッチを公開します。平滑化されたバックアップと離散またはサンプリングされた連続アクション空間を介して、フルランクまたは低ランクの生成シミュレーターをカバーする GSS の有限サンプルのパフォーマンス保証を証明します。適切な重複、規則性、およびアクションカバレッジの条件下では、これらの境界は計画期間に多項式に依存し、共有将来がツリー状の疎なサンプリングの指数関数的な期間依存性を回避できる場合を形式化します。我々は、GSS が長期間にわたってツリーベースのプランナーを大幅に上回る、または最適に近いパフォーマンスを達成する連続制御シミュレーションを実証し、オンライン制御の補完的な設計原則として非分岐グラフ計画をサポートします。
原文 (English)
Graph Sparse Sampling: Breaking the Curse of the Horizon in Continuous MDP Planning
Planning under uncertainty in continuous domains is essential for autonomous systems, yet computationally demanding. Tree-based search methods such as Monte Carlo Tree Search (MCTS) remain popular, but their branching structure can require sampling budgets that grow exponentially with lookahead depth in the worst case. From a tree perspective, continuous state or action spaces become especially challenging, since the planner must decide where to search in an infinite branching hierarchy. We propose Graph Sparse Sampling (GSS), an online planning algorithm that shares sampled futures across many candidate decisions, rather than sampling separate successors for each candidate action. This branch-free graph exposes large GPU-friendly batches, while using heuristics to focus computation. We prove finite-sample performance guarantees for GSS covering full-rank or low-rank generative simulators via smoothed backups, and discrete or sampled continuous action spaces. Under suitable overlap, regularity, and action-coverage conditions, these bounds have polynomial dependence on the planning horizon, formalizing when shared futures can avoid the exponential horizon dependence of tree-shaped sparse sampling. We demonstrate continuous-control simulations where GSS substantially outperforms tree-based planners on long horizons or achieves near-optimal performance, supporting no-branching graph planning as a complementary design principle for online control.
SovereignPA-Bench: 進化する意図、プラットフォームメディエーション、および同意の制約の下でのユーザー所有のパーソナルエージェントの評価
パーソナル エージェントは、ユーザーが所有する永続的な仲介者になりつつあります。彼らは、設定を記憶し、プラットフォームを介した情報をフィルターし、ツールを使用し、サービスと交渉します。既存のベンチマークは、ツールの使用、Web ナビゲーション、デスクトップ コントロール、パーソナライゼーション、レコメンデーション、および進化するコンテキストを評価しますが、エージェントがユーザーの主権を保持しているかどうか、つまりプライバシー、同意、証拠、ユーザーの負担、操作的インセンティブへの抵抗を尊重しながらユーザーの現在の利益を推進しているかどうかを問うことはほとんどありません。 SovereignPA-Bench は、進化する意図、プラットフォームの調停、プライバシー境界、同意の制約、証拠要件、負担のトレードオフの下でユーザー所有のパーソナル エージェントを評価するための実行可能なベンチマークです。このベンチマークは、エージェントに表示される ObservableState を評価者専用の HiddenLabels から分離し、タスクの成功、調整、プライバシー、同意、証拠、操作、負担、および監査可能性に関するコンポーネント メトリックをレポートし、モデルとポリシーの比較のためにペアのシナリオの順序を保持します。私たちは、4 つのモデル ファミリと 8 つの政策ベースラインにわたる 120 の主権ストレス シナリオを評価し、生のプロンプト、出力、プロバイダー形式の応答、解析されたアクション、再計算可能なメトリクス、ハードセット分析、定性的ケース、および 240 項目を超えるブラインド 3 人のアノテーター監査を含む 3,840 の凍結プロンプト トラジェクトリを生成します。フルソブリン足場は、直接ベースライン、記憶のみベースライン、同意のみベースライン、証拠のみベースライン、ReAct/ツール使用ベースライン、安全プロンプトベースライン、およびジャッジガードベースラインよりも主権スコアを向上させると同時に、プライバシー漏洩、同意違反、過剰な譲歩、操作キャプチャを削減します。人間による監査では、プライバシーと同意については高い一致が示され、操作については低い一致が示され、プラットフォーム説得の判断の主観的な境界線が特定されます。これらの結果は、パーソナルエージェントの評価は、タスクの完了を超えて、代表的で、同意を意識した、証拠に基づいた行動に移行する必要があることを示しています。
原文 (English)
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negotiate with services. Existing benchmarks evaluate tool use, web navigation, desktop control, personalization, recommendation, and evolving context, but rarely ask whether an agent preserves user sovereignty: advancing the user's current interests while respecting privacy, consent, evidence, user burden, and resistance to manipulative incentives. We introduce SovereignPA-Bench, an executable benchmark for evaluating user-owned personal agents under evolving intent, platform mediation, privacy boundaries, consent constraints, evidence requirements, and burden tradeoffs. The benchmark separates agent-visible ObservableState from evaluator-only HiddenLabels, reports component metrics for task success, alignment, privacy, consent, evidence, manipulation, burden, and auditability, and preserves paired scenario ordering for model and policy comparisons. We evaluate 120 sovereignty stress scenarios across 4 model families and 8 policy baselines, yielding 3,840 frozen-prompt trajectories with raw prompts, outputs, provider-form responses, parsed actions, recomputable metrics, hard-set analyses, qualitative cases, and a blinded 3-annotator audit over 240 items. Full-sovereign scaffolding improves sovereignty score over direct, memory-only, consent-only, evidence-only, ReAct/tool-use, safety-prompt, and judge-guard baselines while reducing privacy leakage, consent violation, over-concession, and manipulation capture. Human audit shows high agreement on privacy and consent and lower agreement on manipulation, identifying the subjective frontier of platform-persuasion judgments. These results show that personal-agent evaluation must move beyond task completion toward representative, consent-aware, evidence-grounded action.
LLM-as-a-Verifier: 汎用検証フレームワーク
トレーニング前、トレーニング後、テスト時のコンピューティングのスケーリングは、LLM の機能を向上させるための中心的なパラダイムとなっています。この研究では、ソリューションの正しさを判断する能力である検証を新しいスケーリング軸として特定します。これを解き放ち、その有効性を実証するために、追加のトレーニングを必要とせずにエージェント タスクに対するきめ細かいフィードバックを提供する汎用検証フレームワークである LLM-as-a-Verifier を導入します。 LLM に候補解に対する離散スコアの生成を促す標準の LM ジャッジとは異なり、検証者としての LLM は、スコアリング トークン ロジットの分布に対する期待値を計算して連続スコアを生成します。この確率的定式化により、(1) スコアの粒度、(2) 反復評価、および (3) 基準の分解といった複数の次元に沿って検証を拡張することができます。特に、スコアの粒度をスケーリングすると、正の解と負の解がより適切に分離され、より校正された比較が得られることを示します。さらに、繰り返しの評価と基準分解をスケーリングすることにより、分散と複雑さの軽減を通じて検証精度がさらに向上します。さらに、検証者の連続スコアを使用して候補の中から最適なソリューションを選択するための、コスト効率の高いランキング アルゴリズムを導入します。 LLM-as-a-Verifier は、 Terminal-Bench V2 (86.5%)、SWE-Bench Verified (78.2%)、RoboRewardBench (87.4%)、および MedAgentBench (73.3%) で最先端のパフォーマンスを達成します。検証を超えて、LLM-as-a-Verifier からのきめ細かい信号は、タスクの進行状況を推定するためのプロキシとしても機能します。私たちは Claude Code の拡張機能を構築し、開発者が独自のエージェント システムを監視および改善できるようにします。最後に、LLM-as-a-Verifier が RL に緻密なフィードバックを提供し、ロボット工学と数学的推論のベンチマークにおける SAC と GRPO のサンプル効率を向上させることができることを示します。
原文 (English)
LLM-as-a-Verifier: A General-Purpose Verification Framework
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify verification, the ability to determine the correctness of a solution, as a new scaling axis. To unlock this and demonstrate its effectiveness, we introduce LLM-as-a-Verifier, a general-purpose verification framework that provides fine-grained feedback for agentic tasks without requiring additional training. Unlike standard LM judges that prompt LLMs to produce discrete scores for candidate solutions, LLM-as-a-Verifier computes the expectation over the distribution of scoring token logits to generate continuous scores. This probabilistic formulation enables verification to scale along multiple dimensions: (1) score granularity, (2) repeated evaluation, and (3) criteria decomposition. In particular, we show that scaling the scoring granularity leads to better separation between positive and negative solutions, resulting in more calibrated comparisons. Moreover, scaling repeated evaluation and criteria decomposition consistently lead to additional gains in verification accuracy through variance and complexity reduction. We further introduce a cost-efficient ranking algorithm for selecting the best solution among candidates using the verifier's continuous scores. LLM-as-a-Verifier achieves state-of-the-art performance on Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%). Beyond verification, the fine-grained signals from LLM-as-a-Verifier can also serve as a proxy for estimating task progress. We build an extension for Claude Code, enabling developers to monitor and improve their own agentic systems. Finally, we show that LLM-as-a-Verifier can provide dense feedback for RL, improving the sample efficiency of SAC and GRPO on robotics and mathematical reasoning benchmarks.
SiamixFormer: バイテンポラル リモート センシング画像における正確な建物検出と変化検出のためのテンポラル フュージョンを備えた完全トランスフォーマー シャム ネットワーク
リモート センシング画像を使用した建物の検出と変化の検出は、都市計画や救助計画に役立ちます。さらに、自然災害後の建物被害評価にも活用できます。現在、既存の建物検出モデルのほとんどは、1 つの画像 (災害前の画像) のみを使用して建物を検出しています。これは、災害後の画像では、破壊された建物の存在によりモデルのパフォーマンスが低下するという考えに基づいています。この論文では、災害前後の画像を入力として使用する、SiamixFormer と呼ばれるシャム モデルを提案します。私たちのモデルには 2 つのエンコーダーがあり、階層的なトランスフォーマー アーキテクチャがあります。両方のエンコーダの各ステージの出力は、クエリが災害前の画像から生成され、(キー、値)が災害後の画像から生成されるという方法で、特徴融合のための時間変換器に与えられます。この目的のために、特徴融合では時間的特徴も考慮されます。特徴融合で時間変換器を使用するもう 1 つの利点は、CNN と比較して、変換器エンコーダによって生成された大きな受容野をより適切に維持できることです。最後に、時間変換器の出力は、各ステージの単純な MLP デコーダに与えられます。 SiamixFormer モデルは、建物検出については xBD および WHU データセットで評価され、変更検出については LEVIR-CD および CDD データセットで評価されており、最先端のモデルを上回るパフォーマンスを発揮する可能性があります。
原文 (English)
SiamixFormer: a fully-transformer Siamese network with temporal Fusion for accurate building detection and change detection in bi-temporal remote sensing images
Building detection and change detection using remote sensing images can help urban and rescue planning. Moreover, they can be used for building damage assessment after natural disasters. Currently, most of the existing models for building detection use only one image (pre-disaster image) to detect buildings. This is based on the idea that post-disaster images reduce the model's performance because of presence of destroyed buildings. In this paper, we propose a siamese model, called SiamixFormer, which uses pre- and post-disaster images as input. Our model has two encoders and has a hierarchical transformer architecture. The output of each stage in both encoders is given to a temporal transformer for feature fusion in a way that query is generated from pre-disaster images and (key, value) is generated from post-disaster images. To this end, temporal features are also considered in feature fusion. Another advantage of using temporal transformers in feature fusion is that they can better maintain large receptive fields generated by transformer encoders compared with CNNs. Finally, the output of the temporal transformer is given to a simple MLP decoder at each stage. The SiamixFormer model is evaluated on xBD, and WHU datasets, for building detection and on LEVIR-CD and CDD datasets for change detection and could outperform the state-of-the-art.
PotatoGAN: 敵対的生成ネットワーク、インスタンスのセグメンテーション、説明可能な AI を利用してジャガイモの病気の特定と分類を強化
深層学習技術を使用した農業病害のセグメント化の自動化から、数多くのアプリケーションが生まれています。ただし、これらのアプリケーションを新しい条件に適用すると、オーバーフィッティングの問題に頻繁に直面し、その結果、セグメンテーションのパフォーマンスが低下します。ジャガイモ栽培では病気が収量に大きな影響を与えるため、これらの病気を迅速かつ適切に特定することが農業経済にとって重要です。回転、反転、変換などの従来のデータ拡張アプローチには限界があり、強力な一般化結果を提供できないことがよくあります。これらの問題に対処するために、私たちの研究では PotatoGAN と呼ばれる新しいアプローチを採用しています。この新しいデータ拡張アプローチでは、2 種類の敵対的生成ネットワーク (GAN) を利用して、健康なジャガイモの画像から合成ジャガイモの病気画像を生成します。このアプローチは、データセットを拡張するだけでなく、多様性を追加し、モデルの一般化を強化するのに役立ちます。 Inception スコアを尺度として使用した私たちの実験では、PotatoGAN によって作成された画像の品質と現実性が向上し、実際の疾患画像によく似た画像の能力が強調されました。 CycleGAN モデルは、より高い IS スコアから明らかなように、画質の点で Pix2Pix GAN モデルを上回っています。CycleGAN は、黒い斑点と一般的な黒星病でそれぞれ 1.2001 と 1.0900 というより高いインセプション スコア (IS) を達成しています。この合成データにより、大規模なニューラル ネットワークのトレーニングを大幅に改善できます。また、データの多様性と一般化機能を強化しながら、データ収集コストも削減します。私たちの研究では、ジャガイモの病害分類のために 3 つの勾配ベースの Explainable AI アルゴリズム (GradCAM、GradCAM++、および ScoreCAM) と 3 つの異なる CNN アーキテクチャ (DenseNet169、Resnet152 V2、InceptionResNet V2) を組み合わせることにより、解釈可能性が向上しました。
原文 (English)
PotatoGANs: Utilizing Generative Adversarial Networks, Instance Segmentation, and Explainable AI for Enhanced Potato Disease Identification and Classification
Numerous applications have resulted from the automation of agricultural disease segmentation using deep learning techniques. However, when applied to new conditions, these applications frequently face the difficulty of overfitting, resulting in lower segmentation performance. In the context of potato farming, where diseases have a large influence on yields, it is critical for the agricultural economy to quickly and properly identify these diseases. Traditional data augmentation approaches, such as rotation, flip, and translation, have limitations and frequently fail to provide strong generalization results. To address these issues, our research employs a novel approach termed as PotatoGANs. In this novel data augmentation approach, two types of Generative Adversarial Networks (GANs) are utilized to generate synthetic potato disease images from healthy potato images. This approach not only expands the dataset but also adds variety, which helps to enhance model generalization. Using the Inception score as a measure, our experiments show the better quality and realisticness of the images created by PotatoGANs, emphasizing their capacity to resemble real disease images closely. The CycleGAN model outperforms the Pix2Pix GAN model in terms of image quality, as evidenced by its higher IS scores CycleGAN achieves higher Inception scores (IS) of 1.2001 and 1.0900 for black scurf and common scab, respectively. This synthetic data can significantly improve the training of large neural networks. It also reduces data collection costs while enhancing data diversity and generalization capabilities. Our work improves interpretability by combining three gradient-based Explainable AI algorithms (GradCAM, GradCAM++, and ScoreCAM) with three distinct CNN architectures (DenseNet169, Resnet152 V2, InceptionResNet V2) for potato disease classification.
大規模な言語モデルを使用した特定ドメイン オントロジーの構築
オントロジーは、人間とシステムの両方が理解できる情報を整理および維持するための有用な構造です。ただし、手動で作成するのは骨の折れる作業であるため、多くの特定のドメインには参照オントロジーがありません。大規模言語モデル (LLM) によって実証される自然言語を理解する優れた能力により、オントロジー開発を含むさまざまな分野での支援への応用が促進されています。この研究では、ドメイン専門家の役割で LLM を使用して、特定の初期概念の概念階層を構築する手法の実験を紹介します。 GPT-3.5 と GPT-4 を使用して、ブラジルの海域 (別名ブルー アマゾン) のドメイン向けに自動的に構築された 20 のオントロジーが、人間の専門家によって評価されました。モデルはドメインの全体的に一貫した概念化を構築できましたが、洗練せずにコンテキストの表現として完全に満足のいく出力はありませんでした。
原文 (English)
Specific Domain Ontology Construction Using Large Language Models
Ontologies are useful structures to organize and maintain information that can be understood both by humans and systems. However, since their manual crafting is a laborious task, many specific domains lack reference ontologies. The outstanding ability for understanding natural language demonstrated by the Large Language Models (LLMs) has motivated their application to aid on a variety of fields, including on ontology development. This work presents the experimentation with a technique that uses LLMs in the role of domain experts to build conceptual hierarchies for a given initial concept. Twenty ontologies automatically constructed for the domain of the Brazilian maritime territory (a.k.a the Blue Amazon) using GPT-3.5 and GPT-4 were then evaluated by human experts. The models were able to construct overall coherent conceptualizations of the domain, but none of the outputs was completely satisfactory as a representation of the context without refinement.
ボソン量子計算のための SRF キャビティとトランスモンのニューラルネットワーク逆設計
三次元超伝導高周波(SRF)空洞は、非常に長寿命の電磁モードを提供し、トランスモン量子ビットなどの非線形要素と結合すると、ボソン量子情報処理の有望なアーキテクチャになります。このようなシステムの逆設計、つまり、指定された電磁ターゲットと結合ターゲットを生成するデバイスの形状を回復することは、一般に 1 対多の問題です。量子ビットと空洞の結合強度は、トランスモンの幾何学形状と空洞の電磁場内の位置の両方に敏感に依存します。これらのシステムがスケールアップし、設計パラメータ空間が拡大するにつれて、従来の反復シミュレーションのコストは法外なものになります。設計スタックの相補的なレベルでこの逆設計問題に対処する 2 つのディープ ニューラル ネットワーク (DNN) アプローチを紹介します。 1 つ目は、ターゲット キャビティの観測値を生成する SRF キャビティ ジオメトリを提案します。 2 つ目は、ターゲット量子ビット共振器パラメーター (結合率、量子ビット周波数、非調和性 $(g, \nu_q, \alpha)$) を生成するトランスモン量子ビット設計を提案します。復元された候補設計は $\sim$5\% (キャビティ) および $\sim$2\% (トランスモン) 以内でターゲットと一致しており、エンドツーエンドの再シミュレーションによって確認されます。どちらのアプローチも、望ましいデバイスの動作を候補設計に直接マッピングするもので、通常必要とされる反復シミュレーション研究に代わる迅速な代替手段となります。
原文 (English)
Neural-Network Inverse Design of SRF Cavities and Transmons for Bosonic Quantum Computation
Three-dimensional superconducting radio-frequency (SRF) cavities provide exceptionally long-lived electromagnetic modes and, when coupled to nonlinear elements such as transmon qubits, become promising architectures for bosonic quantum information processing. The inverse design of such systems, i.e., recovering device geometries that produce specified electromagnetic and coupling targets, is generally a one-to-many problem. The qubit-cavity coupling strength depends sensitively on both the transmon geometry and its position within the cavity's electromagnetic field. As these systems scale up and their design parameter spaces grow, the cost of conventional iterative simulation becomes prohibitive. We present two deep neural network (DNN) approaches that address this inverse-design problem at complementary levels of the design stack. The first proposes SRF cavity geometries that produce target cavity observables. The second proposes transmon qubit designs that produce target qubit-cavity parameters -- the coupling rate, qubit frequency, and anharmonicity $(g, \nu_q, \alpha)$. The recovered candidate designs match the targets to within $\sim$5\% (cavity) and $\sim$2\% (transmon), confirmed by end-to-end re-simulation. Both approaches map desired device behavior directly to candidate designs, a fast alternative to the iterative simulation studies usually required.
OpenClaw の GLM-5 サービス提供パラメータ調整: ロングコンテキスト エージェント ワークロード向けの単一デプロイメント MaaS 推論の最適化
OpenClaw リクエストは、システム プロンプト、会話履歴、コンテキスト ウィンドウにフィードバックされるツール出力など、ツールによって拡張された長いプレフィックスによって占められます。このワークロードでは、リクエストごとに約 28,000 ~ 30,000 の入力トークンと 500 の出力トークンがあり、サービスの品質は、ショート プロンプトのスループットだけではなく、スループット、TTFT、およびテール レイテンシによって決まります。このレポートでは、MaaS マルチモデル推論最適化アーキテクチャ内での GLM-5 サービング パラメーターの調整について調査します。スコープは推論最適化レイヤーの単一ノード最適化ブロックで、チャンク プリフィル、テンソル並列処理 (TP)、パイプライン並列処理 (PP)、およびリクエストの同時実行性が 1 つの GLM-5 サービング デプロイメント用に調整されます。このレポートでは、「単一ノードの最適化」はアーキテクチャ ブロックを指しますが、実験は 2 ノード、16 GPU クラスターで実行されます。テストされたスペース内での最適な構成は、chunked-prefill-size=3072、tp=4、pp-size=4、および max-running-requests=24 です。保守的な 2048/4/4/16 ベースラインと比較すると、リクエスト スループットが 0.43 から 0.48 req/s に、総トークン スループットが 9029.64 から 9993.23 tok/s に増加し、平均 TTFT が 8.98 秒から 6.69 秒に、レイテンシ P90 が 40.23 秒から 32.64 秒に減少します。同じハードウェア フットプリントの下では、これは推定でリクエストあたりのサービス コストが 10.4% 低くなり、トークンあたりのコストが 9.6% 低いことに相当します。結果は、最適値はワークロード固有であることを示しています。チャンク サイズが大きくなり、キューが深くなっても、パフォーマンスは単調に向上するわけではありません。したがって、デフォルトの OpenClaw 導入プロファイルとして 3072 / tp4 / pp4 / max24 をお勧めします。
原文 (English)
GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads
OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the context window. For this workload, with about 28k-30k input tokens and 500 output tokens per request, serving quality is governed by throughput, TTFT, and tail latency rather than short-prompt throughput alone. This report studies GLM-5 serving-parameter tuning within a MaaS multi-model inference optimization architecture. The scope is the Single-Node Optimization block of the inference-optimization layer, where chunked prefill, tensor parallelism (TP), pipeline parallelism (PP), and request concurrency are tuned for one GLM-5 serving deployment; in this report, "Single-Node Optimization" denotes the architecture block, while experiments run on a two-node, sixteen-GPU cluster. Within the tested space, the best configuration is chunked-prefill-size=3072, tp=4, pp-size=4, and max-running-requests=24. Compared with the conservative 2048/4/4/16 baseline, it increases request throughput from 0.43 to 0.48 req/s and total token throughput from 9029.64 to 9993.23 tok/s, while reducing average TTFT from 8.98 to 6.69 s and latency P90 from 40.23 to 32.64 s. Under the same hardware footprint, this corresponds to an estimated 10.4% lower serving cost per request and 9.6% lower cost per token. The results show that the optimum is workload-specific: larger chunk sizes and deeper queueing do not monotonically improve performance. We therefore recommend 3072 / tp4 / pp4 / max24 as the default OpenClaw deployment profile.
AutoResearch: An Execution-Grounded Multi-Agent Framework for Reliable Research Workflow Automation
Automated research agents increasingly generate code, retrieve literature, and draft scientific artifacts, but they often fail to verify wh…
SWIFT: Spatio-temporal Wavelet Integrated Forecasting Framework for Workload Traces
Accurate cloud workload forecasting is pivotal for efficient resource management but remains challenging as workloads are highly volatile a…
PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving
We present PEEK, a lightweight scheduling and eviction framework for both online (streaming) and offline (batch) LLM serving; this paper fo…
The Hidden Water Geography of U.S. Hyperscale Data Centers in the AI Era
Water use by data centers is routinely reported as a single footprint, but water is consumed through two physically distinct pathways: at t…
Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure
DiLoCo-style training reduces communication by letting learner islands train locally before occasional outer synchronization, making it att…
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spati…
Double-Helix Active Geometry: LiDAR-Anchored Multi-View Depth with Selective Abstention
Consumer depth sensors such as the LiDAR scanner on recent iPhones provide metric range, but their useful range is short and their returns…
Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration
Diffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semant…
From Raw Segmentations to Simulation-Ready Cardiac Meshes: An Automated Framework for Anatomical Reconstruction and Virtual Cohort Generation
Computational models of the human heart are widely used to study electromechanical and fluid-dynamical cardiac function and to support appl…
CRODA-ST: Single-Target Cross-Receiver Open-Set Radio Fingerprint Recognition
Radio frequency fingerprint identification (RFFI) provides a physical-layer credential for Internet of Things devices, but open-set decisio…
DOSE-I: A Multimodal Biosignal Dataset of Procedural Sedation for Endoscopy -- Technical Report
In this document, we describe characteristics and technical details of the multimodal biosignal dataset DOSE-I of procedural sedation for e…
Additive Causal Construction for Transferable and Reconfigurable Cross-System Learning in Multi-Source Image Fusion
In multi-source image fusion scenarios, heterogeneous inputs are typically driven by distinct generative mechanisms and can be viewed as a…
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving
The key-value (KV) cache has become a first-order memory object in LLM serving rather than a temporary per-request tensor. This survey clas…
Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models
Vision-language models can perform new tasks without parameter updates through in-context learning (ICL), whose core mechanism is utilizing…
Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed inc…
An automated method of identifying incorrectly labelled images based on the sequences of loss functions of deep learning networks
Deep learning is widely applied in medical image analysis, but up to 10% of manually labelled images may be incorrect, degrading model perf…
AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents
Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safe…
Knowledge-Centric Information Systems
For decades, data engineering has developed mature architectural principles for integrating, governing, validating, cataloging, and serving…
Fusion: A Framework for Unified Sequential Token AdaptatIon in VisiOn TraNsformers
Vision Transformers achieve strong image classification accuracy but process all image regions with nearly the same computation, even when…
The agent creates, we validate: A Lightweight Framework for Agentic Artifact Generation
Generating structured artifacts with Large Language Models - e.g. database queries, threat framework mappings, entity schemas - is relative…
COMET: Combinatorial Optimization for Multiplex Editing Targets Via Constraint-Preserving QAOA
Multiplex CRISPR-Cas9 gene editing requires selecting one guide RNA per target gene subject to cross-gene interactions: a constrained combi…
Fine-Grained Computation Offload for Off-the-Shelf Servers in Tens of Lines
Hardware accelerators now sit on the critical path of online serving. GPUs, FPGAs, and increasingly remote services such as hardware securi…
QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting
Time-series forecasting supports decisions in finance, en-ergy, transportation, public health, and industrial monitoring. Recent foundation…
Federated Learning for Object Detection: Enabling Collaborative Drone Learning Without Centralizing Data
Object detection is a fundamental capability for AI-driven perception in safety-critical drone and edge-vision systems, including disaster…
Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting
Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existi…
Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving
Real-time interaction models -- Moshi, MiniCPM-o, Qwen-Omni -- turn serving into a periodic real-time task: on every frame a session ingest…
K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexp…
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experience…
LLMoxie: Exploring Agentic AI for Scientific Software Development
In this paper, we describe LLMoxie, an institutional AI platform whose three-tiered architecture supports multi-cloud and on-premise infere…
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale
There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various…
Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models
Recent advances in large-scale image generative models enable photorealistic scene synthesis with controllable attributes. Beyond data augm…
Signal from Space: Detecting Schools and Towers to Bridge the Digital Divide
Reliable internet access is essential for modern education, yet millions of school-aged children especially in developing regions remain of…
SMOCS: A Streaming Framework for Simplified Deployment, Monitoring, and Optimization of ML Systems in Production
Machine learning has demonstrated significant potential for real-time monitoring, optimization, and control of scientific facilities. Howev…
Echoes of Unrest: A Multimodal NLP Framework for Early Warning of Fake News and Violence-Driven Mob Activity
Rapid growth in social media has transformed global communication by enabling fast information exchange, but it has also accelerated the sp…
Out-of-Distribution Generalization of Risk Aversion in Language Models
Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AI…
Gemma 4 Technical Report
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance c…
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weigh…
A Preliminary Study on Explaining Risk of Code Changes using LLM-Based Prediction Models
Predictions by machine learning (ML) and artificial intelligence (AI) models are often received skeptically unless they are paired with int…
Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes
As LLMs are increasingly deployed as autonomous adjudicators in semi-open textual game environments, robust rule adherence becomes critical…
Training Hybrid Block Diffusion Language Models with Partial Bidirectionality
High-throughput long-context generation is one of the central challenges for large language models. Generation is typically memory-bandwidt…
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure
Personal agents will increasingly negotiate on behalf of users: splitting costs with other personal agents, appealing platform decisions, e…
Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models
Cloud-edge Large Vision-Language Model (LVLM) inference enables efficient deployment by splitting computation between edge devices and clou…
JavaVulBench: A Java Vulnerability Benchmark with Realistic Splits, a Unified Multi-Backend Harness, and a Leakage-Aware Evaluation Mode
We release \textsc{JavaVulBench}, a benchmark dataset and evaluation harness for Java vulnerability detection. The dataset contains $\sim$3…
Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation
Multi-view robotic manipulation methods with the attention mechanism have recently achieved significant progress in both training efficienc…
Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI
Large language model agents driving security tool suites over the Model Context Protocol are increasingly common. Yet the factors that boun…
Where do LLMs Fall Short in CBT-Guided Affective Reasoning?
Cognitive Behavioral Therapy (CBT) provides a structured framework for understanding a user's mental state by examining the interaction bet…
SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal Roughness
Deploying AI-generated video detectors in real-world services demands an ultra-low false positive rate (FPR) on real videos to avoid falsel…
PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement
Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising p…
TIER: Trajectory-Invariant Explanation Regularization for Membership Privacy
Explainability is central to building trustworthy AI, yet explanation interfaces can inadvertently provide adversaries with an expanded pri…
Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models
Taxonomies provide key information about the semantic relationships between concepts and the inherent organization of vision and language.…
CoACT: Action-Preserving Observation Compression for Coding Agents
LLM-based coding agents solve software-engineering tasks through iterative interactions with development environments, where returned obser…
Bootstrap Flow-Map Tree Sampling Enables Online Feedback Driven Search
In many scientific and engineering domains, maximizing discovery within a limited sampling budget demands strategic, observation-guided exp…
Harmonic-Aware Transformer for Real-Time Catheter Localization in Interventional Procedures of Magnetic Particle Imaging
Magnetic particle imaging (MPI) enables real-time, radiation-free tracking of magnetic nanoparticle-coated instruments, making it highly su…
R3D: Quantitative 3D Spatial Reasoning for Egocentric Wearables
Quantitative 3D spatial reasoning from egocentric RGB-D video is a critical capability for next-generation wearable assistants. Yet existin…
VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning
Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep…
Modeling the Impact of Visual Brand Language on Attention, Object Recognition, and Memory Retrieval
Visual brand language is the set of visual properties that convey brand identity for a product. What is the impact of visual brand language…
PromptPET: Privacy-Utility Optimized Prompt Obfuscation
Privacy is an important challenge when users interact with AI chatbots, since users may share sensitive information, explicitly or implicit…
MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding
Materials phase diagrams are a core knowledge representation in materials science, encoding temperature,composition, phase stability, and p…
A Precedent-Guided Co-Scientist for Side-Effect-Aware Drug Redesign
We propose PRECEDE, a precedent-guided co-scientist for side-effect-aware drug redesign that revises a parent compound to mitigate a specif…
Pooling-Based Context Modeling for Convolution-Free Deep Image Prior
Convolutional Neural Networks (CNNs) achieve strong denoising performance by exploiting spatial context from neighboring pixels. Deep Image…
The Foreign Policy AI Evaluation Gap
We argue that AI systems used in conducting foreign policy tasks - broadly enacting 'statecraft' - should be a priority test case for techn…
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding a…
Individual Parameters in Weight-Sparse Transformers Appear Interpretable
A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does. Dominant…
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor length extrapolation of…
Enhanced Feature Extraction for IoT Network Intrusion Detection Using GNNs and KAN
Recent advancements in the Internet of Things (IoT) emphasize the urgent need for advanced network security, as IoT networks feature dynami…
VISTA: Auditing Semantic Divergence in Vision-Language Models
Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or…
CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples tha…
PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark
Text-rich image models can now design poster-scale layouts, but we lack ways to measure whether they honor scientific communication contrac…
Back to Basics: Improving Molecular Understanding in LLMs via SMILES-Graph Translation
Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet thes…
Can Model Merging Improve Aggregation in DiLoCo?
Model merging techniques, which aggregate independently finetuned models into one to combine their capabilities, have become a topic of sig…
HyperVAttention: Efficient Sparse Attention with Spatio-Temporal Clustering for Video Diffusion
Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation. However, their ability t…
MambaLIE: Scene Light Intensity-Boosted Low-Light Image Enhancement with State Space Model
Images captured by consumer electronic devices, such as mobile phones and digital cameras, often suffer from low-light degradation due to s…
OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models
Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they ge…
LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression
The rapid growth in the parameter scale of large language models (LLMs) has created a strong demand for efficient compression techniques. A…
Spectral Rewiring for Exploration, Purification, and Model Merging
Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two de…
PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised Segmentation
Semi-supervised semantic segmentation (SSSS) has long turned on one question, which pseudo-labels to trust, and answered it with ever more…
Attention-Guided Efficientnet Architecture For Precise Criminal Identification in Surveillance Images
Criminal identification from surveillance imagery has become a critical research area in intelligent forensic surveillance systems due to t…
STELLA: Efficient Sensor-to-LLM Translation for On-Device Human Activity Recognition
HAR is increasingly expected to run continuously on edge devices, yet recent LLM-based methods remain hard to deploy: raw sensor prompts ar…
Don't Wait to Reply: Towards Responsive yet Thoughtful Dialogue through Proactive Thinking
Thinking has emerged as a critical capability for Large Language Models (LLMs) tackling complex tasks. However, its reactive nature, where…
Flow-A11y: Flow-Aware Accessibility Testing
Modern web applications increasingly expose accessibility barriers through interaction flows rather than static page snapshots. Keyboard tr…
SNR-Adaptive Unified Diffusion for Multi-Task Medical Image Segmentation
Clinical cardiac imaging pipelines currently deploy separate models for each dataset and modality, incurring redundant training costs and p…
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy
Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards st…
A Multi-Task Deep Learning Framework for Real-Time Intelligent Video Surveillance with Temporal Event Validation
Modern video surveillance systems generate far more video streams than human operators can effectively monitor, making automated analysis e…
Detecting Architectural Drift in Safety-Critical Firmware through Runtime Trace Analysis
Maintaining consistency between architectural design and runtime-observed behavior is challenging in long-lived safety-critical firmware. T…
Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning
Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified,…
CuBAS: Information Geometric Curvature-Based Adaptive Sampling for Supervised Classification
The informativeness of a training set is as consequential as its size, yet most sampling strategies remain agnostic to the intrinsic geomet…
Rethinking Neural Nonlinearity as Gating
Activation functions are considered an essential primitive for neural nonlinearity, i.e., they enable neural networks to serve as universal…
Conditional Diffusion Guided Knowledge Transfer for Multi-Domain Knowledge Graph Completion
Multi-domain knowledge graph completion (MKGC) aims to improve missing triple prediction in a target KG by transferring knowledge from othe…
Which Algorithm Specification Formats Help Language Models Implement Machine Learning Algorithms?
Large language models (LLMs) are increasingly used to implement algorithms from research manuscripts, but papers often leave implementation…
The Role of Prompt Language and Translation-Theory-Driven Prompts in Large Language Models: A Case Study on Spanish-Chinese Journalistic Translation
This study examines how prompt language and translation theory-driven prompt design influence the quality of Spanish-Chinese journalistic t…
KARMA: Knowledge graph-based Automated Reasoning Materialization and Alignment
Template-based contrastive synthesis is scalable, but its candidates often differ only in a few entity-slots while sequence-level optimizat…
Decentralised Federated Learning over Temporal Networks: The Role of Heterogeneities
Decentralised federated learning, based on peer-to-peer communication, is increasingly proposed for on-device training of machine learning…
Effectiveness of LLM-based Software Diversity for Reliability Improvement -- an Empirical Study
Software diversity has been extensively studied as a means of reducing the risk of common-mode failures. Classic work showed that the centr…
CRRL: A Causality-Based Reinforcement Learning Framework for Autonomous System Recovery
Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failu…
Teaming Up with AI: Coordination and Cooperation
Successful diffusion of AI in the workforce hinges on the economic value that AI brings to human endeavors. Bringing AI into the workforce…
AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning
Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into exe…
Scalable Maximal Frequent Episode Mining with Desbordante
Episode mining aims to extract subsequences of events that possess certain distinctive properties and constitute facts valuable to the user…
A Bayesian Framework for Evaluating Scenario Compatibility in Generative Population Synthesis
Scenario-based transportation analysis specifies future assumptions through aggregate population targets, whereas generative population syn…
Self-Specializing Vision-Language Transmon Chip Calibration in a Physics-Grounded Environment
Calibrating a superconducting transmon chip is a sequential decision problem under noise, drift, and a finite budget: an expert must choose…
Transition Information Density: Morphological Trajectories, Synesthetic Perception, and Structured Interpolation in Neural Training (or: The Synesthetic AI)
Standard machine learning training presents data as discrete endpoint pairs, omitting the structure of the space between them. This paper i…
OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance
We present OpenGlass, an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance, with a primary foc…
Builder, Defender, Breaker: The Case Against Removing the Human from the AI-Driven Security Lifecycle
Artificial intelligence has spread across the whole of the security lifecycle. The same family of models now writes application code, harde…
CONTRA: Red-Teaming Configurations of Personalizable Agents
Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agent…
Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions
The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for m…
Unbiased Alignment for Large Language Models with Noisy Preferences
The alignment of large language models with human preferences is commonly achieved through Reinforcement Learning from Human Feedback or Di…
Semantic Segmentation-Driven Image-Level Diagnosis of Liver Cancers in Hematoxylin and Eosin Histopathology Images
As hematoxylin & eosin (H&E) staining constitutes the primary entry point in routine diagnostic workflows, computer-aided diagnosis from wh…
A harmonised dataset for Earth system foundation models
Foundation models for Earth systems have so far been trained primarily on physical climate and weather data, with limited representation of…
Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild
Agentic code review, where autonomous agents provide code review comments on pull requests, is increasingly integrated into development wor…
Hierarchical Multi-Agent Reinforcement Learning for Carbon-Aware AI Data Centers in Power Distribution Systems
Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy co…
From Judgments to Issues: Structured Extraction of Legal Reasoning with Citation-Hallucination Control
We present an automated pipeline that decomposes Italian tax-court judgments into individual legal issues and extracts, for each issue, a s…
SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference
LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often…
FedAvg for HAR: Exploring the Tradeoff Between Personalized and Generalization Accuracy
The federated learning (FL) paradigm fosters distributed pervasive computing combined with artificial intelligence techniques, allowing for…
Efficient Decentralized Multi-task Dataset Valuation via Model Merging
Accurate and efficient dataset valuation is essential for enabling fair and transparent data marketplaces, especially when multiple contrib…
PedestrianDiffusion: Multimodal Generative Denoising and Dense State Estimation for Inertial Navigation
The accuracy of consumer-grade inertial navigation is bottlenecked by the stochastic noise of Micro-Electro-Mechanical Systems (MEMS). Trad…
LLM-Enhanced Hierarchical Heterogeneous Graph Representation Learning for Malicious Python Package Detection
Malicious Python packages have become a major threat to software supply chain ecosystems due to the widespread adoption of open-source repo…
Brand-as-Memory: Vision-Language Models Encode Causal, Mechanistically Localizable Credibility Priors for News Sources
Vision-language models (VLMs) increasingly read news and web content as images, where the publisher's identity is visually present. We show…
Spectral Signatures of Large Language Models
The rapidly growing repository of publicly available large language models (LLMs) presents significant challenges for systematic management…
The S-ICDF Dataset: Sionna-Simulated Dynamic Interference Characterization and Direction Finding
Jamming and spoofing threaten wireless and satellite navigation by disrupting or manipulating radio frequency (RF) signals, undermining ava…
DETECT-3B-Omni is Agnostic of Content and Demographics
A trustworthy and GDPR-compliant deepfake audio detector must base its decisions on acoustic artifacts, not on what is being said or who is…
Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies
Modern AI agent implementations such as frontier coding agents chain multiple tools at runtime that create a security surface that per-tool…
Amortising Bayesian Experimental Design for Sequential Information Gathering in LLMs
Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectivel…
No Time Like the Present: Agentic Test-Time Training for LLM Agents
LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies…
TRIAGE: Trustworthy Retrieval Instrumentation And Graph Evaluation
Knowledge graphs (KGs) that underpin Graph-based Retrieval-Augmented Generation (Graph-RAG) are increasingly built automatically by LLM-dri…
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control
Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term m…
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
While skill optimization for autonomous agents has gained traction, existing methods rely on complex pipelines. This leaves a fundamental q…
Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning
Inference-time alignment methods, such as Best-of-$N$, offer a flexible alternative to training-based alignment by using reward models to s…
CaresAI at SMM4H-HeaRD 2026: Predicting TNM Staging
This study aims to predict Tumor, Node, and Metastasis (TNM) stage labels independently, with the Cancer Genome Atlas (TCGA) pathology repo…
Towards Diverse and Comprehensive Benchmarks for Mutual Information Estimation
Mutual information (MI) estimation is a central problem in machine learning and statistics; however, existing benchmarks typically evaluate…
STRATOS: Bridging the Symbolic-to-Numeric Gap in Spatio-Temporal Text-to-SQL for Meteorological Data
Copernicus, the European Union's Earth observation program, produces petabytes of Earth observation and climate data, offering immense pote…
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers w…
CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI
Enterprise artificial intelligence is moving from experimentation into operational workflows. Early programs focused on model access and re…
AGL-1: The Enterprise AI Governance Layer as a Control Plane for Trusted Enterprise Intelligence
Enterprise artificial intelligence is moving from isolated experimentation toward operational dependency across copilots, retrieval-augment…
Aligning Language Models with Selective Prediction
Large language models (LLMs) are increasingly deployed as critical decision-making components in high-stakes real-world AI systems, renderi…
Latent Clarity: Bridging World-Model Kinematics to Semantic Manifolds for Video Anomaly Anticipation
Continuous video anomaly detection is dominated by reactive Multiple Instance Learning (MIL) that collapses spatiotemporal features into sc…
EPRA U-Net: An Efficient Pyramid Residual Attention Framework for Accurate Infarct Segmentation in Diffusion-Weighted MRI
Objective: Accurate identification of acute ischemic infarcts on diffusion-weighted magnetic resonance imaging (DWI) is a critical prerequi…
Teacher Supervision over Representation Equivalence Classes
Knowledge distillation is usually framed as a choice of what to match in the teacher - its logits, hidden features, or sample relations - w…
Differentiate the Evaluator, Not the Program: An Efficient Runtime Representation for Neuro-Symbolic Learning
AI systems increasingly propose executable scientific models whose value depends on both their symbolic structure and their fitted continuo…
PLGSA-Transformer: Periocular Landmark-Guided Attention with Occlusion-Adaptive Cosine Thresholding for Cross-Modal Masked and Unmasked Face Recognition
The widespread adoption of facial masks, accelerated by COVID-19 and mandated in security-sensitive settings, has exposed limitations of co…
Responsibility Distribution Estimation in Ego-View Accident Videos with Multimodal Large Language Models
Recent studies on multimodal traffic accident understanding have mainly relied on infrastructure-camera footage, satellite imagery, or stru…
An Interpretable Deep Learning Framework for Discovery and Clinical Validation of Deep Radiomic Signatures in Tumor Classification
Imaging signatures are quantitative features extracted from medical images that provide clinically meaningful information for tumor diagnos…
Token-Based Affordance Grounding with Large Vision-Language Models
Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence…
They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It
When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was do…
A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning
Adversarial robustness in Unsupervised Domain Adaptation (UDA) remains a significant challenge due to noisy pseudo labels and inherent dist…
RADIO1D: Elastic Representations for Condensed Vision Modeling
This paper challenges the assumption that vision-language models (VLMs) require fixed patch-based 2D vision features. Analyzing fine-tuned…
An AI-Assisted Solution to the Signed BAR Conjecture: Uniqueness in the Harrison--Reiman Class and a Completely-$\mathcal{S}$ Class Obstruction
For a multidimensional reflected diffusion, determining whether the associated basic adjoint relationship (BAR) uniquely characterizes the…
Revealing Hidden Model Behaviors with Task-Specific Self-Reports
Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only wh…
Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing
Decades of orbital missions have produced multi-modal remote sensing data for the Moon, spanning optical imagery, spectroscopy, thermal emi…
ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation
Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a ma…
ELiTeFormer: An Efficient Transformer for FPGAs
Transformer blocks are prevalent in large language model (LLM) but present deployment challenges due to their challenging computational and…
AutoCedar: An Agentic Framework for Verifier-Guided Access Control Policy Synthesis
Large Language Models are increasingly used to turn natural-language requirements into code. In access control, that shortcut is dangerous:…
ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation
Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoidi…
A Fair Benchmarking of Deep Relational Database Learning Models
Relational databases (RDBs) are the primary data infrastructure in many enterprises, yet recent deep learning methods designed for RDBs hav…
Phase-Preserving Trimodal Transformer for Tropical Forest Biomass Estimation Using Optical and PolInSAR Data
The accurate estimation of Above-Ground Biomass (AGB) in mature tropical forests remains a critical challenge in remote sensing, primarily…
Don't Blame the Large Language Model: How Scaffolding Evolution Shapes Coding Agent Quality
Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agentic scaffolding:…
OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-…
CoGen3D: An Agentic Human-AI Co-Design Pipeline for 3D Asset Generation for Virtual Reality
Creating 3D assets for virtual reality requires modeling expertise, which restricts the authorship of immersive experiences. Existing gener…
Attending to Multimodal Generation One Token at a Time
Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving…
A Failure-Mode Benchmark for Polymorphic Sybil Poisoning in RAG
We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning. The framework par…
EmCom-Diffusion: Probing Visual Reflection in Emergent Languages via Image Generation
Measuring the extent to which emergent languages encode the visual content of their inputs is an open problem. We refer to this property as…
FedACT: Federated Adaptive Coordinate Trust Modulation for Robust Transformer Training under Data Heterogeneity
Federated Transformer training increasingly relies on local AdamW, whose adaptive updates can provide much stronger local progress than SGD…
Self-Improving Diffusion Classifiers with Minority Preference Optimization
Prior studies have demonstrated that diffusion classifiers achieve robust zero-shot classification performance. However, their effectivenes…
SkillFab: An Agent-Native Skill Production Platform
SkillFab is an agent-native platform for turning missing capabilities into reviewed, reusable Agent Skills. At runtime, agents first search…
Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective
While prior studies have successfully compressed vision Transformers (ViTs) through various pruning techniques, most have concentrated on w…
Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks
Symmetry is everywhere in nature and society. Geometric deep learning exploits symmetries in data to improve the performance and efficiency…
Punching Above Their Weight: Classification-Head Fine-Tuning of Tiny Language Models (TLMs) for Verifiable Multiple-Choice Tasks
We define Tiny Language Models (TLMs) as models below roughly 3B parameters that fit on mainstream consumer devices. We study how to adapt…
CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation
The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dol…
Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings
Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representatio…
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on s…
DualView: Preventing Indirect Prompt Injection in Personal AI Agents
Personal AI agents that run on the user's local machine, such as OpenClaw, automate daily tasks including web search, email, and file manag…
Q-TriM: Question-Guided Tri-Modal Attention for Audio-Visual Question Answering
Audio-Visual Question Answering (AVQA) extends classical VQA by requiring joint reasoning over video and synchronized audio. However, many…
How Do Diffusion Classifiers Decide? A Bias-Centric Evaluation
Diffusion models have recently been repurposed for zero-shot classification, giving rise to diffusion classifiers that identify the best-ma…
Beyond Static Rules: Automated Discovery of Latent Vulnerabilities in Text-to-SQL
While Large Language Models (LLMs) have achieved remarkable success in Text-to-SQL tasks, their deployment in real-world environments is hi…
When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts
Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresse…
High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching
Generative models such as diffusion and flow matching have advanced robotic visuomotor policies by modeling multimodal action distributions…
GeoSelect: Spatial-Program Execution for Training-Free Referring Remote Sensing Image Segmentation
Referring remote sensing image segmentation isolates the object named by a natural-language expression in an aerial image. Existing trainin…
Next-Gen Sponsored Search: Crafting the Perfect Query with Inventory-Aware RAG (InvAwr-RAG) Based GenAI
Sponsored search plays a crucial role in e-commerce revenue generation, where advertisers strategically bid on keywords to capture the atte…
Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate…
Enhancement of E-commerce Sponsored Search Relevancy with LLM
Sponsored search plays a crucial role as a revenue stream for search engines, wherein advertisers competitively bid on keywords that align…
Advanced Topic Modeling Techniques for Categorizing Software Vulnerabilities
The increasing complexity and frequency of software vulnerabilities demand efficient methods to analyze and prioritize threats. Traditional…
TabQueryBench: A Query-Centric Benchmark for Synthetic Tabular Data
Synthetic tabular data support use cases like data sharing, model development under access restrictions, and rapid prototyping of analytica…
TokAN: Accent Normalization Using Self-Supervised Speech Tokens
Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity. The…
Probe, Don't Prompt: A Hidden-State Probe for Metadata Filtering in Multi-Meta-RAG
Multi-Meta-RAG improves retrieval for multi-hop question answering by filtering a vector store on metadata (the news source) that it extrac…
MPSelectTune: Prompt-type Selection for Fine-tuning improves Concept Unlearning in LLMs
LLMs can be conveniently adapted to a diverse set of tasks, e.g, prediction, question-answering tasks, etc, using appropriate prompts with…
Why3-py: A Tool for Formal Verification of Hypothesis Testing and Meta-Analysis in Python
The reproducibility crisis in scientific research has received widespread recognition, thereby increasing the importance of meta-analyses t…
The Remarkable Effectiveness of Providing AI Agents with Natural Language Tools: A Replication Study Validating NLT Performance Across 14 Models
This study independently replicates and extends the Natural Language Tools (NLT) framework of Johnson et al.~(2025), which questions the us…
NormWorlds-CF: Solver-Verified Counterfactual Normative Reasoning with Metamorphic-Relation GRPO
Language models can reach the right normative verdict for the wrong reason. We introduce NormWorlds-CF, a solver-verified environment for c…
Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulat…
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents
Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, an…
Order-based Causal Discovery for Multistage Processes
Causality has become an increasingly important tool for gaining a deeper understanding of complex systems. Among various causal analysis me…
Scalable Semantic Steering of Embedding Projections
Low-dimensional projections support interactive visual analysis of high-dimensional data embeddings, but their structure often does not ali…
BanglaMemeEvidence: A Multimodal Benchmark Dataset for Explanatory Evidence Detection in Bengali Memes
Memes have become influential communication tools on social media, combining viral visuals with concise messaging to convey impactful ideas…
NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating…
Full Glyph Images Beat Token Embeddings: A Controlled Study for Transformers
Modern language models generally represent text as sequences of discrete token embeddings, an assumption deeply rooted in current practice…
Separating Representation from Reconstruction Enables Scalable Text Encoders
While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evalu…
When Does Small Data Work? Accuracy and Efficiency Trade-offs Between Tabular Foundation Models and Conventional Methods for Crowd-State Classification at Hajj and Umrah
Learning from few labeled examples is a central challenge in tabular machine learning, and it becomes the binding constraint in domains whe…
Finite Reliability Representations: Noise-Calibrated Belief-Space Covers for Reliable Decision-Making
Physical sensing and actuation noise floors should inform how much belief resolution a decision-making system can reliably use. We introduc…
A Unified Algebraic Framework for Classification Performance Evaluation
We propose a unified algebraic framework for classification performance evaluation that encompasses binary, multiclass, multilabel, ordinal…
Efficient Discovery of Conditional Dependencies with Desbordante
Conditional functional dependencies (CFDs) are functional dependencies with a restricted scope: they specify the context in which a depende…
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning…
The "I Don't Know" Filter: Enhancing Agentic Reliability in Function Calling
The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks. However, the metrics used in…
Reward-Gated On-Policy Distillation
On-policy distillation is a powerful way to transfer reasoning ability from a strong teacher to a smaller student: the student samples traj…
Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability
Distinguishing Large Language Model (LLM) generated text from human writing is a critical and difficult challenge. While LLMs are trained t…
Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw…
Enhancing Implicit Neural Representations with Image Feature Embedding for Unsupervised Cardiac Cine MRI Reconstruction
Cardiac cine Magnetic Resonance Imaging (MRI) is a critical diagnostic tool that provides dynamic insights for radiologists. To accelerate…
Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders
Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a…
Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions
Large language models can generate plausible quantum code, but it is unclear whether they can reliably target the specific software develop…
Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering
Recent Multi-modal Large Language Models (MLLMs) have demonstrated remarkable performance on 2D question answering tasks. However, extendin…
FedSPM: Routing-Enabled Federated Learning under Dual Heterogeneity via Semiparametric Mixture
Routing-prediction federated learning has emerged as a new paradigm that reframes inter-client heterogeneity as a resource for system-level…
Submitted and Diagnostic Analysis of Full-Text Temporal Retrieval for LongEval-Sci
LongEval-Sci evaluates scientific retrieval under collection change, where a system should be effective on the current corpus and remain us…
DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs re…
Parametric Memory Decoding for Zero-Shot Routing in LoRA-Based External Parametric Memory
With the rise of parametric memory, LoRA-based External Parametric Memory (EPM) has emerged as a modular solution, but existing routing met…
SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction
Reconstructing Computer-Aided Design (CAD) modeling sequences from images is crucial for preserving design intent and supporting parametric…
Conflict-Based Lazy Search for Fast Multi-Manipulator Planning
Employing multiple manipulators can boost efficiency and accomplish tasks that a single manipulator cannot do. However, real-time planning…
CSB: A Counting and Sampling tool for Bit-vectors
Satisfiability modulo theory (SMT) solvers have significantly advanced automated reasoning due to their effectiveness in solving problems a…
!Imperio, smolVLA: The Implications of Data Poisoning on Open Source Robotics
This work establishes that trigger-word data poisoning of vision language action models is practical, while at the same time the open-sourc…
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding
Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for L…
Mask-based Predictive Representations for Reinforcement Learning
Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information. It is crucial to abstract effe…
Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese)
This paper explains the principles and provides examples of a new method for distinguishing between FAKE human speech synthesized by genera…
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering
Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual qu…
Piercing Gilbreath's Conjecture: From Deep Number Theory Insights to Fintech and Cybersecurity
I propose a new methodology to attack the fascinating Gilbreath's conjecture about prime numbers, first posted in 1878 and unsolved to this…
CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving
End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often…
Detecting Hallucinations in Retrieval-Augmented Generation through Grounding-Aware Sensitivity by Perturbation (GASP)
Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level scor…
SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects
Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction,…
Hierarchical Multi-to-Single-Modal Knowledge Distillation for Disruption Prediction in EAST
Plasma disruption is a critical threat to tokamak safety. Existing data-driven predictors mainly rely on time-series diagnostic signals, wh…
Signal or Noise? Understanding Generative Models for Real-World Sensor Time Series
Generative models have changed how machine learning represents complex data distributions, especially in language and vision, yet many real…
HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models
World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to ca…
LBR: Towards Mitigating Length Bias in Large Language Models for Recommendation
Large language models (LLMs) have recently emerged as powerful backbones for recommender systems by reformulating recommendation as a token…
Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement
The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawi…
Risk-Constrained Freshness-Aware Semantic Caching for Open-Web Retrieval-Augmented LLMs
Semantic caching reduces the latency and cost of retrieval-augmented generation (RAG) by serving cached answers to semantically similar que…
Agentic-V2X: Small Language Model Agents for Deadline-Aware V2X Scheduling in 5G/6G Networks
Large Language Models (LLMs) are proposed as control interfaces for next-generation networks, but their latency, hallucinations, and lack o…
CausalGame: Benchmarking Causal Thinking of LLM Agents in Games
Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention. Since scientific discovery fundame…
HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM Inference
We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM i…
Fixed-Confidence Best-Arm Identification for Causal Mediation Analysis
This paper studies the problem of identifying the treatment that maximizes the expected natural direct potential outcome (NDPO), which capt…
One Framework for All: Cross-Modal Membership Inference for Generative Models
Large generative models across text-to-text, text-to-image, and image-to-text modalities have been shown to pose significant privacy risks.…
IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation
While Large Vision-Language Models (VLMs) demonstrate remarkable generic capabilities, their clinical reasoning in specialized domains like…
HASSL: Hierarchy-Aware Self-Supervised Learning Framework for Single Cell Microscopy
Hierarchical structure is common in image data, where fine-grained clusters often merge into larger, coarser semantic groups. In biological…
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound…
Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention
Recent NVFP4 pretraining methods mainly target transformer linear layers, leaving optimizer states, optimizer arithmetic and attention unde…
Transferability Between Understanding and Generation in Unified Multimodal Models
Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact…
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning
Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-p…
dOPSD: On-Policy Self-Distillation for Diffusion Language Models
Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autore…
evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations
The dominant practice in language model evaluation is to report a single accuracy number per model and declare the higher one better, witho…
On Pairwise Quantile Regression -- Statistical Guarantees and Applications
Quantile regression provides a powerful tool for summarizing the conditional distribution of a real valued random variable (r.v.) of intere…
Covert Trait Propagation Is Representation Alignment: Mechanistic Evidence from Hidden-Channel Distillation
A student model trained on pure uniform noise can still inherit its teacher's digit-classification ability, provided the two share initiali…
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their cap…
A Retrieval-Augmented Framework for Detecting and Resolving Pragmatic Ambiguities in Natural Language Requirements
Natural language requirements (NLRs) are essential for bridging communication gaps among diverse stakeholders in software development. Howe…
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Research dissemination, turning a paper into a poster, a talk video, and a blog post, is still a manual last mile. Prior automation treats…
Generative wave propagator
Seismic wavefield simulation is fundamental to seismology, but conventional finite-difference (FD) methods remain limited by numerical disp…
Wan-Streamer v0.2: Higher Resolution, Same Latency
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps t…
From Regulation to Requirements: An Automated Requirement Derivation and Explanation Pipeline
Ensuring software compliance with regulations such as the General Data Protection Regulation (GDPR) and the Artificial Intelligence Act (EU…
A Deep Learning-based surrogate model for Severe Accidents in nuclear reactors using ASTEC
Integral codes like the Accident Source Term Evaluation Code (ASTEC) are powerful tools to study the physics of Severe Accidents (SAs) in n…
Robustness Verification of an Autonomous Underwater Vehicle-based Plankton Classifier
The assessment of planktonic standing stocks and microorganism structures is critical for understanding upper ocean biological processes. C…
Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models
World-model evaluation for model-based reinforcement learning typically asks whether the learned model predicts reward and value well, whic…
Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning
Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent rei…
PulmoSight-XAI: An Explainable Multi-View Attention Ensemble with Gradient Boosting Meta-Learning for Multi-Label Chest X-Ray Classification
Automated chest X-ray classification remains challenging due to severe class imbalance, co-occurring pathologies, and the loss of localized…
Two Black Boxes, One Solver: Encoder Probing and Decoder Attribution for Neural Multi-Attribute VRP under Hard-Mask and Recourse Decoders
Neural autoregressive solvers for the Multi-Attribute Vehicle Routing Problem (MAVRP) reach competitive cost but offer no per-step justific…
Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5
Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Q…
Failures and Successes to Learn a Core Conceptual Distinction from the Statistics of Language
Generic statements like "tigers are striped" and "cars have radios" communicate information that is, in general, true. However, while the f…
Language Models Represent and Transform Concepts with Shared Geometry
How concepts are represented in neural networks is a fundamental question in machine learning. The dominant view treats concept representat…
Training-Free Model Selection and Domain-Aware Score Calibration for First-Shot Anomalous Sound Detection
First-shot anomalous sound detection in DCASE Challenge Task 2 must flag anomalies of unseen machine types with a single threshold, without…
Lyapunov-Guided Training for Hardware-Safe Neural Networks Under Fixed-Point Arithmetic
Low-precision neural networks are attractive for resource-constrained hardware, but fixed-point arithmetic introduces failure modes that ar…
Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse
Code language models are now trusted collaborators in production workflows for debugging, refactoring, and iterative repair, and every benc…
CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining
Camera-radar (CR) fusion is a practical sensing configuration for autonomous driving, but existing models are typically trained with task-s…
Auto: The AGI Compiler
Every LLM agent run re-derives its behavior token by token on a frontier model: brilliant, expensive, slow, and unbounded. We present Auto,…
Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interactio…
Explainable Novel Category Discovery in Semantic Concept Space
Novel category discovery aims to identify unseen classes from unlabeled data by transferring knowledge from labeled categories, but most ex…
Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption
We present a bidirectional framework for estimating the energy consumption of text-to-video (T2V) and text-to-video-audio (T2VA) models fro…
Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations
Accurate prediction of patient-specific therapeutic response from pre-treatment transcriptomes is hindered by the scarcity of matched clini…
EEG-SpikeAgent: Agentic Closed-Loop Program Synthesis for Automated EEG Spike Detection
Automated detection of interictal epileptiform discharges in scalp electroencephalography (EEG) is clinically important, but recent high-pe…
A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training
For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels co…
LLM-Driven CI-CD Workflow Intelligence for Cyber Systems Engineering
CI/CD workflows have become executable operational policy: they decide what gets built, tested, released, and deployed, and they mediate ho…
Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning
Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, languag…
TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models
Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated b…
LCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target Detection
Infrared small target detection (IRSTD) aims to identify long distance small targets from complex infrared backgrounds, and is a fundamenta…
Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-…
G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement
The rapid advancement of AI-generated videos poses increasing security risks and calls for robust detectors with strong cross-domain genera…
SILO: Simulation-in-the-Loop Sim-to-Real Transfer for Multi-Stage Cable Routing
Linear-deformable manipulation remains challenging due to the complex deformations of objects such as cables and ropes. Prior data-driven a…
Hierarchical Evidence-Driven Reasoning for Long Document Understanding
Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to…
Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations
Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment i…
Machine Learning for Depression Screening and Intervention: an Original Circadian Rhythm Score-based Methodology
Depression screening from large-scale behavioral data is challenged by fragmented circadian indicators, limited interpretability, and the l…
Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving
Reconstructing 3D scene structures from sparse, low-overlap observations remains a fundamental challenge in autonomous driving. Recent stat…
Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes
On-device LLM decoding is a hard-barriered CPU-SIMD computation that wants every core for milliseconds per token, while the rest of the OS…
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-La…
ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents
Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails. A model that never…
URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment
Synthesis planning aiming to find pathways of reactions for a target molecule is one of the most important and challenging tasks in drug di…
Strategic Buying Agents
Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide w…
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents
Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks. Howeve…
Geometry-Aware Motion Latents for Learning Robust Manipulation Policies
Learning motion latents for robotic manipulation heavily relies on extracting motion patterns from visual sequences, yet effective action a…
Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards
Automatic data visualization generation has advanced rapidly with multi-modal large language models, yet existing efforts largely focus on…
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment
Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which ine…
RustMizan: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities
LLM agents are increasingly applied to vulnerability analysis, but existing benchmarks have not kept pace. They typically rely on small non…
Wasserstein Residuals: Learning Gradient Flows from Population Dynamics
Reconstructing population dynamics is a central problem in the physical and data sciences. Often, the dynamics are modeled as a Wasserstein…
Trust Region Policy Distillation
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), wh…
Multi-Turn On-Policy Distillation with Prefix Replay
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student…
Predicting Drafted Deck Strength for "Magic: the Gathering"
Many real-world games do not admit a fixed, compact rule set: instead, their dynamics are defined by interactions among a large and often e…
An Exploration of Agentic Information Fusion for Test Maintenance Prediction
Test maintenance is a critical, yet costly, activity - particularly as codebases rapidly evolve. To assist, we present MAST, a multi-agent…
Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection…
HamQASBench: A Hamiltonian-Informed Diagnostic Benchmark for Evaluating Quantum Architecture Search
Quantum Architecture Search (QAS) automates the design of parameterized quantum circuits for variational quantum algorithms, yet existing b…
Pretraining Curricula Enable Selective Fine-tuning
Transformers follow implicit curricula whereby some tasks are learned before others. However, how explicit pretraining curricula influence…
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects.…
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, so high-…
Graph Representation Learning of Longitudinal Medical Imaging Trajectories for Treatment Response Prediction
In patients with breast cancer, pathological complete response (pCR) has been established as a clinically meaningful surrogate marker for l…
Efficient Perception in Automotive Detection and Tracking Using Neuromorphic Computing
Deep learning algorithms are notorious for their high carbon footprint and computational demands that limit their deployment on edge device…
Input Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study
How does the way information reaches a transformer -- as symbolic tokens, a clean per-factor "oracle" code, or an entangled perceptual vect…
DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as…
MemPose: Category-level Object Pose Estimation with Memory
In the pursuit of robust and generalizable category-level object pose estimation, most existing methods adopt parametric formulations that…
Multi-Robot Open Adaptive Teaming Across Unseen Environments, Partners, and Scales
Deploying robot teams in the real world requires simultaneous adaptation to unseen environments, unknown partners, and varying team sizes,…
Joint Velocity Slope Diffusion Prior for Structurally Constrained Velocity Model Building
High-resolution velocity models are crucial for reservoir characterization and subsurface delineation. However, the band limited nature of…
LLM for the development of FCM
This article is about the development of a fuzzy cognitive map using a local large language model. In the light of recent advances it is ev…
The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System
Many phenomena of deep learning are dynamical: they concern not only which minima exist, but how gradient descent reaches, avoids, or selec…
TACTIC-KG: Toward Small Agent Teams for Cyber Threat Intelligence Knowledge Graph Construction
Cyber Threat Intelligence (CTI) reports are predominantly unstructured, heterogeneous, and noisy, which limits their direct usability for a…
Comparison of Loss Functions for Robust Deep Learning-based Echocardiography Segmentation when Learning with Partially Labelled Data from Multiple Domains
Echocardiography is the first imaging modality used for assessing cardiac function, and accurate segmentation of cardiac structures is esse…
ImputeECG: Deep Learning Reconstruction of Complete 12-Lead Electrocardiograms from Incomplete Recordings for Cardiac Assessment
Complete digital 12-lead electrocardiograms (ECGs) are essential for AI-enabled cardiovascular assessment, yet many clinical ECG records, p…
Hyperparameter Transfer in Graph Neural Networks
The performance of deep learning models crucially depends on the settings of hyperparameters like learning rate, initialization scale, and…
Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses
Persistent memory has enabled large language model (LLM) agents to store factual knowledge, prior decisions, reasoning histories, tool usag…
LLM-Based Test Oracles: Source-of-Authority Taxonomy -- A Systematic Literature Review
Large language models (LLMs) are increasingly used to produce test oracles, the part of a test that decides whether observed behavior is co…
RUFNet: Query-Guided Support Mask Refinement and Uncertainty Fusion based on Hybrid Mamba for Few-Shot Brain Tumor Segmentation
Few-shot brain tumor segmentation remains challenging due to noisy support masks, inter-patient variations between support and query images…
Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection
Human value detection is commonly formulated as sentence-level multi-label classification over the 19 refined Schwartz values, typically pr…
AIFS-SUBS: Extending Data-Driven Forecasting to Sub-Seasonal Timescales
Data-driven models now rival numerical weather prediction in the medium range, but extending them to sub-seasonal lead times raises challen…
Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters
Grokking -- the delayed onset of generalization long after a network has fit its training set - -is usually studied in models too large to…
Localized LoRA-MoE: Block-wise Low-Rank Experts With Adaptive Routing
Large Language Models (LLMs) and high-dimensional perception networks increasingly rely on parameter-efficient fine-tuning (PEFT) to adapt…
Agent Data Injection Attacks are Realistic Threats to AI Agents
AI agents act on behalf of user prompts, consuming external data and taking actions based on the agent context. Prior research on AI agent…
Three-Phase Evaluation of AI-Assisted Software Development Life Cycle
This paper presents an exploratory evaluation of how increasing levels of AI autonomy affect software development productivity, requirement…
PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference
We present PDEFlow, an autonomous agentic framework that turns user-level ODE and PDE descriptions into solver-backed neural-operator pipel…
Open Problems in AI Incident Governance
AI systems may produce failures after deployment that pre-deployment safety assessments do not anticipate. Managing these failures requires…
Relational Multi-Agent Reinforcement Learning for Dynamic Pricing in High-Speed Railway Markets
In liberalised railway systems, operators must set prices dynamically in an environment with partial observability, as they retain private…
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistanc…
Unified Audio Intelligence Without Regressing on Text Intelligence
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-…
Noisy-Channel Minimum Bayes Risk Decoding
Minimum Bayes Risk (MBR) decoding yields more robust and higher-quality text generation than maximum a posteriori (MAP) decoding by selecti…
Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous Computing
Computing-in-Memory (CIM) accelerators execute Matrix-Vector Multiplications (MVMs) in memory, making them a compelling solution for Machin…
CanniUplift: A Holistic Framework for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling
Personalized incentive allocation is vital for e-commerce, where uplift modeling is the standard for estimating Individual Treatment Effect…
Privacy-Preserving Robustness Verification for Neural Networks
Neural network verification and data privacy are inherently in tension: verification demands full access to model parameters and input data…
Shifting from Discrete to Continuous Reference Data: QSM-Derived Horizontal Tree Biomass Distribution for Deep Learning Biomass Estimation
Conventional modeling approaches for LiDAR-based above-ground biomass (AGB) estimation rely on discrete plot-level inventory aggregates. Th…
Adaptive Inference Batching using Policy Gradients
Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains stat…
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extr…
Wavelet Scattering Transform for Interpretable Schizophrenia Biomarker Discovery and Classification from Resting-State EEG
Schizophrenia is a debilitating neuropsychiatric disorder characterized by profound cortical network dysregulation, for which objective, cl…
Air Quality Downscaling with Station-Guided Pseudo-Supervision
Super-resolving coarse atmospheric fields to local PM$_{2.5}$ variations is uniquely challenged by a mismatch in spatial support: while pix…
Topological Shape Representation for Aneurysm -- Bifurcation Detection
Automated detection of intracranial aneurysms (IAs) from CT angiography (CTA) is severely hindered by high false-positive rates. Convolutio…
Steering Optimisation Trajectories in Diffusion Representation Learning
We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace th…
TREK: Distill to Explore, Reinforce to Refine
Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls…
Multiplayer Interactive World Models with Representation Autoencoders
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-pl…
Selective Disclosure Watermarking for Large Language Models
Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs). Existing approaches in…
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or…
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text b…
GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks
For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable…
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon task…
What Does a Discrete Diffusion Model Learn?
What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor? At the level of jump rates, these are…
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deep…
Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty Quantification
Time-domain surveys generate many transient candidates, making Real-Bogus classification a critical step in automated discovery pipelines.…
Weak-to-Strong Generalization via Direct On-Policy Distillation
Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to r…
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting de…
DrugAgent: Reliable Multi-Agent Integration of Conflicting Biomedical Evidence for Drug-Target Interaction Assessment
Workflows in drug-target interaction (DTI) assessment require integrating heterogeneous data from predictive models, curated resources, and…
Neuro-symbolic Weak Supervision: Theory and Semantics
Weak supervision enables machine learning models to learn from limited or noisy labels, but it introduces challenges in reliability and sem…
Serious Games: Human-AI Interaction, Evolution, and Coevolution
The serious games between humans and AI have only just begun. Evolutionary Game Theory (EGT) models the competitive and cooperative strateg…
Shutdownable Agents through POST-Agency
Many fear that future artificial agents will resist shutdown. I present an idea - the POST-Agents Proposal - for ensuring that doesn't happ…
AI's Blind Spots: Geographic Knowledge and Diversity Deficit in Generated Urban Scenario
Diffusion-based text-to-image models are increasingly used for urban analysis and scenario generation, but their geographic knowledge and r…
Policy Improvement with Style-Specific Demonstrations
Proficient game agents with diverse play styles enrich the gaming experience and enhance the replay value of games. However, recent advance…
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
High-stakes decision-making involves navigating multiple competing objectives with expensive evaluations. For instance, in brachytherapy, c…
A Technical Survey of Reinforcement Learning Techniques for Large Language Models
This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Prox…
Interactive Learning for LLM Reasoning
Existing multi-agent learning approaches have developed interactive training environments to explicitly promote collaboration among multipl…
Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
Perturbation-based explainability methods face criticism due to their reliance on out-of-distribution mutants. This raises doubts about the…
RLIE: Rule Generation with Logistic Regression, Iterative Refinement, and Evaluation for Large Language Models
Large Language Models (LLMs) can propose rules in natural language, sidestepping the need for a predefined predicate space in traditional r…
トランスフォーマーベースのモデルと人間の脳ネットワーク間の位相調整のための統合幾何空間
これまでの脳と AI の連携研究は通常、特定の入力とタスクによって制約され、さまざまなモダリティを備えたモデル全体で組織特性を捕捉する能力が制限されていました。この研究では、Transformer ベースのモデルに焦点を当て、脳モデルのトポロジカル アライメント空間を導入します。神経メカニズムからアライメントを推測するのではなく、グラフベースの組織特性を通じてアライメントを調査し、モデルの固有の空間注意トポロジーを標準的な人間固有接続ネットワーク (ICN) にマッピングします。これにより、組織特性のレベルで視覚、言語、およびマルチモーダル システムにわたる、モダリティに依存せずタスクフリーの比較が可能になります。これらのモダリティとスケールにわたる 151 の Transformer ベースのモデルを分析すると、さまざまな程度のトポロジー アラインメントを反映する、連続的な円弧状の分布が観察されます。トレーニングの目的と一致して、グローバルなセマンティック抽象化に最適化されたモデルは高次の ICN とより密接に関連付けられ、ローカルの詳細に焦点を当てたモデルは低レベルの ICN と関連付けられました。さらに驚くべきことに、我々は非直観的な現象を発見しました。DINOv2 は以前のバージョンと比較してアライメントの低下を示し、蒸留された DeiT モデルは、より大きなモデルが高次の ICN とあまりうまくアライメントされない直観に反したスケーリング反転を示し、命令チューニングだけでなく微調整もアライメントに対する効果が限定的でした。さらに、トポロジカル アライメント スコアは、30 個のビジョン トランスフォーマーにおける ImageNet-1K Top-1 精度と有意でない相関関係を示しました (r=0.266、p=0.156)。この研究は、脳参照トポロジー マッピングを通じて、Transformer ベースのモデルの組織特性を比較するための新しい定量的観点を提供します。
原文 (English)
A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and Human Brain Networks
Whether artificial neural networks organize information comparably to the human brain remains unclear. Prior brain--AI alignment studies are constrained by specific inputs and tasks, limiting cross-modal comparison. Here we introduce a brain--model topological alignment space, mapping Transformer attention topology onto human intrinsic connectivity networks (ICNs) to enable task-free, modality-agnostic comparison. Analyzing 151 Transformer-based models with 62,480 attention head graphs, we observe a continuous arc-shaped distribution reflecting varying alignment. Models optimized for global semantics aligned with higher-order ICNs, while local-detail models aligned with sensory ICNs. Non-intuitive findings include reduced alignment in DINOv2 compared to its predecessors and a counterintuitive scaling inversion in distilled DeiT models, while fine-tuning and instruction tuning had limited effect. Alignment scores showed no significant correlation with ImageNet accuracy (r = 0.266, p = 0.156). This work offers a quantitative framework for comparing the organizational principles of artificial and biological systems.
Optimal-Agent-Selection: State-Aware Routing Framework for Efficient Multi-Agent Collaboration
The emergence of multi-agent systems powered by large language models (LLMs) has unlocked new frontiers in complex task-solving, enabling d…
CoT-X: クロスモデルの思考連鎖の転送と最適化のための適応フレームワーク
思考連鎖 (CoT) 推論は、大規模言語モデル (LLM) の問題解決能力を強化しますが、かなりの推論オーバーヘッドが発生し、リソースに制約のある設定での展開が制限されます。この論文では、適応推論要約フレームワークを通じて、さまざまなスケールとアーキテクチャのモデル間での効率的な CoT 転送を調査します。提案された方法は、重要度スコアリングによるセマンティック セグメンテーション、予算を考慮した動的圧縮、および一貫性の再構築を介して推論トレースを圧縮し、トークンの使用量を大幅に削減しながら重要な推論ステップを維持します。 10 の専門分野にわたる 7{,}501 の健康診断の質問に関する実験では、同じトークン予算の下で切り捨てよりも最大 40% 高い精度が示されました。 8 つの LLM (DeepSeek-R1 および Qwen3 を含む 1.5B ~ 32B パラメーター) からの 64 のモデル ペアの評価により、強力なモデル間移行性が確認されました。さらに、ガウス プロセス ベースのベイジアン最適化モジュールにより、評価コストが 84% 削減され、モデル サイズとクロスドメインの堅牢性の間のべき乗則の関係が明らかになります。これらの結果は、推論の要約が効率的な CoT 転送への実用的な道を提供し、厳しい計算制約下で高度な推論を可能にすることを示しています。コードは公開され次第公開されます。
原文 (English)
CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
Long Chain-of-Thought (CoT) traces can improve reasoning accuracy, but repeatedly generating them is costly for smaller or latency-constrained language models. This paper studies a practical alternative: produce a rich rationale once with a capable \emph{thinking} model, compress it, and reuse the compressed trace as context for a cheaper \emph{answering} model. We introduce CoT-X, an adaptive framework for cross-model CoT transfer. CoT-X segments reasoning traces into semantic units, scores their diagnostic and logical importance, selects budget-feasible evidence paths, and reconstructs a coherent compressed rationale for the answering model. On $7,501$ Japanese medical licensing questions spanning $10$ specialties, CoT-X improves accuracy over direct truncation by up to $40.5\%$ under the same token budget, with the largest gains at $64$--$256$ tokens. Across $64$ thinking--answering pairs from eight DeepSeek-R1 and Qwen3 models (1.5B--32B parameters), reasoning transfer is most reliable within a model family, yet remains effective across families once compression normalizes the trace. A Gaussian Process Bayesian optimization layer finds near-optimal model--budget configurations with $15$ evaluations rather than an exhaustive search over all $64$ pairs, reducing evaluation cost by $84\%$. These results show that reasoning quality, token budget, and model compatibility can be optimized jointly, making CoT-style reasoning more practical under realistic deployment constraints.
Turbo-Muon: Almost-Orthogonal Pre-Conditioning for Fast Muon Updates
Orthogonality-based optimizers, such as Muon, have recently shown strong performance across large-scale training and community-driven effic…
The Language of Bargaining: Linguistic Effects in LLM Negotiations
Negotiation is a core component of social intelligence, requiring agents to balance strategic reasoning, cooperation, and social norms. Rec…
OpenTinker: Separating Concerns in Agentic Reinforcement Learning
We introduce \textsc{OpenTinker}, an open infrastructure for training large language model (LLM) agents with many LoRA-backed policies over…
Programming over Thinking: Efficient and Robust Multi-Constraint Planning
Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting…
Graph Neural Networks are Heuristics
Graph neural networks are usually treated as auxiliaries for combinatorial optimization: they imitate algorithms, guide search, or supply s…
Toward Efficient Agents: Memory, Tool learning, and Planning
Recent years have witnessed increasing interest in extending large language models into agentic systems. While the effectiveness of agents…
Insect-inspired Visual Point-goal Navigation
Insect neuroethology provides a compelling biological template for efficient autonomous navigation. We draw an analogy between the formal e…
NEST: Nascent Encoded Steganographic Thoughts
Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is…
Framework of Thoughts: A Foundation Framework for Dynamic and Optimized Reasoning based on Chains, Trees, and Graphs
Prompting schemes such as Chain of Thought, Tree of Thoughts, and Graph of Thoughts can significantly enhance the reasoning capabilities of…
ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning
Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step…
HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis
While deep learning-based weather forecasting paradigms have made significant strides, addressing extreme weather diagnostics remains a for…
Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning
When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to fa…
Correlation-Weighted Multi-Reward Optimization for Compositional Generation
Text-to-image models produce images that align well with natural language prompts, but compositional generation has long been a central cha…
On the Ability of Transformers to Verify Plans
Transformers have shown inconsistent success in AI planning tasks, and theoretical understanding of when generalization should be expected…
Mecha-nudges for Machines
AI agents are becoming active decision-makers on the Internet. As they make decisions in the same environments as humans, the environments…
The Anatomy of Uncertainty in LLMs
Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment. Current approach…
Chronos: The AI Co-Historian
AI is increasingly supporting, accelerating, and automating scientific discovery across subjects. Yet, the adoption of AI in historical res…
TRACE: Capability-Targeted Agentic Training
Models often fail to complete agentic tasks because they lack core capabilities required by the target environment. However, mainstream app…
Gypscie: A Cross-Platform AI Artifact Management System
Artificial Intelligence (AI) models, encompassing both traditional machine learning (ML) and more advanced approaches such as deep learning…
Fun-TSG: A Function-Driven Multivariate Time Series Generator with Variable-Level Anomaly Labeling
Reliable evaluation of anomaly detection methods in multivariate time series remains an open challenge, largely due to the limitations of e…
CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning
Chain-of-Thought (CoT) prompting has emerged as a simple and effective way to elicit step-by-step solutions from large language models (LLM…
To Use AI as Dice of Possibilities with Timing Computation
The dominant noun-based modeling paradigm, grounded in probability theory and committed to pre-specified noun entities as primitive modelin…
Stop Automating Peer Review Without Rigorous Evaluation
Large language models offer a tempting solution to address the peer review crisis. This position paper argues that today's AI systems shoul…
Agentic Retrieval-Augmented Generation for Financial Document Question Answering
Financial document question answering (QA) demands complex multi-step numerical reasoning over heterogeneous evidence--structured tables, t…
Beyond the Black Box: Interpretability of Agentic AI Tool Use
AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because these tool-use decisions ar…
Attributing Emergence in Million-Agent Systems
Large language models (LLMs) can simulate human-like reasoning and decision-making in individual agents. LLM-powered multi-agent systems (M…
Sign-Separated Asymmetric Finite-Time Error Analysis of Q-Learning
Q-learning is known to suffer from overestimation bias: because the Bellman update maximizes noisy or imperfect action-value estimates, pos…
When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State
Outcome-only evaluation can certify economically unsafe agents: a policy can hit a business KPI while violating deployable behavioral disci…
HaorFloodAlert: A 72-Hour Machine Learning Early Warning System for Flash Floods in Bangladesh's Haor Wetlands
Every spring, flash floods strike the haor wetlands of northeast Bangladesh just before the boro rice harvest, and one flood can erase a fa…
MUSE-Autoskill: スキルの作成、記憶、管理、評価による自己進化エージェント
大規模言語モデル (LLM) エージェントは、再利用可能なスキルに依存して複雑なタスクを解決します。ただし、既存のスキル作成アプローチでは、スキルを孤立した静的な成果物として扱い、再利用性、信頼性、長期的な改善が制限されています。私たちは、統一されたライフサイクル (作成、記憶、管理、評価、洗練) の下でスキルを作成、再利用、洗練することにより、エージェントがタスク解決能力を継続的に向上できるようにする、スキル中心のエージェント フレームワークである MUSE-Autoskill Agent (Memory-Utilizing Skill Evolution) を提案します。当社のフレームワークにより、エージェントはオンデマンドでスキルを作成し、それらをタスク間で保存して再利用し、効率的に整理して選択し、単体テストや実行時のフィードバックを通じて評価して継続的に改善することができます。さらに、タスク全体にわたって各スキルの経験を蓄積するスキルレベルの記憶を導入し、時間の経過とともにより効果的な再利用と適応を可能にします。 SkillsBench の実験は、ライフサイクル管理されたスキルがタスクの成功、効率、再利用、およびエージェント間での移転を向上させることができるという最初の証拠を提供し、スキルを長命で経験を意識したテスト可能な資産として扱うことの重要性を強調しています。
原文 (English)
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
Large language model (LLM) agents rely on reusable skills to solve complex tasks, but existing skill creation approaches often treat skills as isolated, static artifacts, limiting reusability, reliability, and long-term improvement. We propose MUSE-Autoskill Agent (Memory-Utilizing Skill Evolution), a skill-centric agent framework that creates, reuses, and refines skills under a unified lifecycle: creation, memory, management, evaluation, and refinement. MUSE creates skills on demand, stores them across tasks, retrieves them through a skill catalog, and accumulates per-skill experience for later reuse and adaptation. Across the main reported settings on SkillsBench and SkillLearnBench, MUSE-Autoskill outperforms Hermes, Codex, and Claude Code. On SkillsBench, its self-created skills surpass human-authored skills on the successfully covered subset (85.24% vs. 81.17%), showing that lifecycle-managed skills can distill agent experience into highly effective reusable assets; MUSE-created skills also transfer to Hermes more effectively than Codex- or Claude-created skills, reaching 51.90% accuracy under transfer. These results highlight the importance of treating skills as long-lived, experience-aware, and testable assets.
言語モデルの文化的結合要素
LLM は、状況によって差別化が必要であるにもかかわらず、文化的グループ全体で平等に扱うことをデフォルトとすることがよくあります。これは違いに対する認識の欠如です。 Wang らによる N4 文化盗用ベンチマークの機構的解釈可能性と要因計画を使用します。 (2025) では、8 つのモデル (4 つのアーキテクチャ、ベースおよび命令) 全体の文化的結合に因果的に寄与する、モデルあたり 2 ~ 3 の中間層のアテンション ヘッドを特定します。文化的バインディングは、文化的アイテムを適切なアイデンティティに関連付けるプロセスです。これらのヘッドのアイデンティティとアイテムのエッジをノックアウトすると、結合強度が 9 ~ 23% 低下します。識別されたヘッドは、命令モデルからベースモデルに移行し、文化的結合が事前トレーニング時に作成されることを示唆しています。 $\alpha$ スケーリングは、段階的な用量反応と生成時の適度な増幅ステアリング ($\alpha = 2-3$) を示し、中立的な推論をほとんどそのままにしながら、文化的区別の精度を 1-3 pp 増加させます。知識調査タスクでは、モデルがそれに基づいて行動するよりも 3 ~ 5 倍多くのことを知っていることがわかり、ボトルネックが知識ではなくルーティングにあることを示しています。
原文 (English)
Cultural Binding Heads in Language Models
LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Using mechanistic interpretability and a factorial design on the N4 cultural appropriation benchmark from Wang et al. (2025), we identify 2-3 mid-layer attention heads per model that contribute causally to cultural binding across eight models (four architectures, base and instruct). Cultural binding is the process of associating cultural items with the appropriate identity. Knockout of the identity-to-item edges on these heads lowers the binding strength by 9-23%. The identified heads transfer from instruct to base models, suggesting that cultural binding is created at pre-training. An $\alpha$-scaling shows a graded dose-response and moderate amplification steering at generation ($\alpha = 2-3$) increases cultural differentiation accuracy by 1-3 pp while leaving neutral reasoning mostly intact. A knowledge probing task shows that models know 3-5 times more than they act upon it, indicating that the bottleneck lies in routing and not knowledge.
シーンの自己探索による視点をもとに計画を立てる
VLM は、各カメラの動きによってビューがどのように変化するかを予測し、事前にそのような動きを多数計画することができますか?私たちはこれを機能ビュー計画と呼びます。これには、(1) 単一のアクションがビューをどのように変換するかを理解すること、(2) ターゲット ビューを特定するために複数ターンの計画にわたってそのような変換を多数構成することが必要です。私たちは、実際の ScanNet シーン上の 3D ポイントクラウド環境である、私たちが提案する ViewSuite で両方の機能を調査します。 13 のフロンティア VLM にわたって、重大な計画のギャップが生じています。VLM は基本的なビューとアクションの知識を持っていますが、それを複数ターンの計画にわたって構成することができず、視点の距離が長くなるにつれてギャップが拡大します。このギャップを埋めるために、自己探索とビュー グラフの蒸留を交互に行う反復フレームワークを提案します。重要な洞察は、結果に関係なく、すべての探索軌跡が集合的にビュー グラフを形成し、シーン全体で視点がどのように接続されているかをコンパクトに捉えるということです。このグラフをさまざまな教師ありタスクに抽出すると、ポリシーの分布が再形成され、純粋な RL を遅らせる希薄な報酬が克服されます。これにより、インタラクティブ ビュー プランニングで Qwen2.5-VL-7B が 2.5% から 47.8% に向上し、GPT-5.4 Pro (18.5%) や Gemini 3.1 Pro (21.4%) を上回りました。自己探索は、3D 空間で積極的に推論して計画できる VLM への有望な道として浮上しています。
原文 (English)
Planning with the Views
Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1)understanding how a single action transforms the view, and (2)composing many such transformations across multi-turn plans to identify a target view. We probe both abilities in our proposed ViewSuite, a 3D point-cloud environment on real ScanNet scenes. Across 13 frontier VLMs, a critical planning gap emerges: they possess basic view-action knowledge but fail to compose it across multi-turn plans, with the gap widening as viewpoint distance grows. To close this gap, we propose an iterative framework that alternates self-exploration with view graph distillation. The key insight is that all exploration trajectories, regardless of their outcome, collectively form a view graph that compactly captures how viewpoints connect across a scene. Distilling this graph into diverse supervised tasks reshapes the policy distribution and overcomes the sparse rewards that stall pure RL. This improves Qwen2.5-VL-7B from 2.5% to 47.8% on interactive view planning, surpassing GPT-5.4 Pro (18.5%) and Gemini 3.1 Pro (21.4%). Self-exploration emerges as a promising path toward VLMs that can actively reason and plan in 3D space. Code and Data are at https://viewsuite.github.io.
Reasoning4Sciences: 推論言語モデルをすべての科学分野に橋渡しする
推論言語モデル (RLM) は科学研究のための強力なツールとして急速に台頭していますが、その影響は主に「ハード サイエンス」分野に集中しています。他の科学分野での RLM の導入が遅い、または導入されていないことが、研究の生産性の差の拡大を引き起こしています。この調査では、欧州研究評議会 (ERC) が使用する社会科学と人文科学、物理科学と工学、生命科学にわたる分類に従って、28 の科学分野にわたる RLM の採用に関する初めての包括的な分析を提供します。私たちは、RLM がどのように開発、評価され、分野全体に適用されるかを調査します。さらに、利用可能なドメイン固有の開発および評価リソースに基づいた成熟度指向の評価フレームワークを導入し、公開されているリソースのみを考慮した場合にさらに顕著になる RLM 成熟度の実質的な格差を明らかにします。最後に、分野を超えて普及しつつある現在の実装パラダイム、現在の課題、科学全体で RLM の導入を可能にする将来の方向性を強調します。
原文 (English)
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches
While Reasoning Language Models (RLMs) are rapidly emerging as powerful tools for scientific research, their impact is primarily concentrated in "hard science" fields. The slow -- or lack of -- adoption of RLMs in other branches of science is causing a widening gap in research productivity. In this survey, we provide the first comprehensive analysis of RLM adoption across 28 scientific disciplines following the classification used by the European Research Council (ERC), spanning the Social Sciences and Humanities, Physical Sciences and Engineering, and Life Sciences. We examine how RLMs are developed, evaluated, and applied across disciplines. Furthermore, we introduce a maturity-oriented assessment framework based on available domain-specific development and evaluation resources, revealing substantial disparities in RLM maturity that become even more pronounced when only publicly available resources are considered. Finally, we highlight current implementation paradigms that are gaining popularity across disciplines, current challenges, and future directions in enabling RLM adoption across science.
TokenMizer: Graph-Structured Session Memory for Long-Horizon LLM Context Management
Long-horizon LLM sessions outlive their context windows, and the standard mitigations - truncation, summarization, retrieval - share a stru…
チャットボットが問題解決主導の会話でどのように機能するかに関するいくつかの仮説。イノベーション幻想の裏付けとしての大規模言語モデル
この記事では、解決策に関連して問題について話し合うときの真の会話パートナーとしてのチャットボットの性質についての視点を提供します。チャットボットは何ができて、何ができないのか、そしてそれはどのように説明できるのでしょうか?私たちの議論は、集合力学、認知言語学、神経心理学、心理学に基づいています。私たちの議論は基本的なチャットボットに焦点を当てており、それによってより高度なチャットボットの中核機能についての意見を述べることができればと考えています。基本的なチャットボットは、シンプルなインターフェイスを備えたラージ言語モデル (LLM) で構成されていると想定されています。主な結果は次のとおりです。いわゆる比喩的な問題の伝播に基づいた人間の理解と思考の説明。 LLM のトレーニングに使用されるテキスト データセットには特定の特徴があり、これらのテキスト データセットは人間の思考と理解を部分的に模倣しているだけであるという仮説。 LLM トレーニング プロセスが、これらのデータセットから人為的な比喩的な問題の伝播を LLM にエンコードしているという仮説。基本的なチャットボットは人間に匹敵する思考パートナーにはなり得ないという私たちの結論。大規模言語モデルのさらなる開発もこれにはつながらないという私たちの結論です。 Yann LeCun 氏は、「動物と人間は、現在の AI や機械学習 (ML) システムの能力をはるかに超えた学習能力と世界の理解を示します。」と述べています。私たちの結論はこれと一致しています。ルカン氏のビジョンと私たちのビジョンは、ビッグテックの楽観主義とは相容れない。だからといって、チャットボットが存在し、個人と組織の両方で大規模に使用されており、したがってチャットボットを理解することが社会的および政治的に重要であるという事実は変わりません。私たちの記事は、チャットボットの機能、利点、欠点に関する議論に貢献することを目的としています。チャットボットがどのように機能するかについての研究で、結論に達するために使用したアプローチにはまだ出会っていません。
原文 (English)
Some hypotheses on how chatbots work in problem-solving-driven conversations. Large Language Models as confirmation of the Innovation Illusion
We discuss the nature of chatbots as conversation partners when discussing the solution of problems. What can chatbots do and what can't they do? We develop hypotheses on how this can this be explained. Our argument draws on insights from Aggregation Dynamics, Cognitive Linguistics, Neuropsychology and Psychology. We establish that chatbots are multifaceted and composite systems. Our argument focuses on basic chatbots in the hope of thereby making statements about the core functionality of more advanced chatbots. Basic chatbots are assumed to consist of a Large Language Model (LLM) with a simple interface. The main results of our research are: a description of human imagination, understanding and thinking based on so-called metaphorical problem propagations; the hypothesis that the texts in the text dataset used for training LLMs have specific characteristics and that these texts only partially imitate human thinking and understanding; the hypothesis that the LLM training process encodes artificial metaphorical problem propagations into an LLM from these text datasets. Our conclusions are that a basic chatbot cannot be a thinking partner capable of matching the cognitive flexibility of humans, and that further development of the Large Language Model will not lead to this either. But chatbots exist, that they are being used on a massive scale, by both individuals and organisations, and that it is therefore socially and politically important to understand them. Our article aims to contribute to the discussion on the functioning, benefits and drawbacks of chatbots. Cognitive Linguistics shows how the use of metaphor is an expression of our thinking. Aggregation Dynamics, is an attempt at a comprehensive systems theory. We believe that the concept of metaphorical problem propagation could provide an interesting addition for both. Chatbots a solution? For what?
数学的推論のための人工知能: 言語モデル、神経記号システム、および検証された発見の統合的調査
数学的推論は長い間、機械知能の厳しいテストとして機能してきました。過去 10 年間で、NLP 内のニッチな問題から、最も重要な AI フロンティアの 1 つに移行しました。この調査は、初期のルールベースの数学文章問題 (MWP) ソルバーとテンプレート駆動の幾何学システムから、神経式生成と LLM プロンプトを経て、現代の推論モデル、マルチエージェント システム、神経記号定理証明者、および検証済みの発見ワークフローに至るまで、この分野の進化に関する統一的な説明を提供します。私たちは 4 つの軸に沿ってランドスケープを整理します。(i) MWP 解決、マルチモーダル ジオメトリ、および VLM にわたる、テキストと図に関する非形式的な推論。 (ii) 自動形式化、戦術予測、コンパイラー主導の修復、および証明検索を含む、証明アシスタントにおける形式的推論。 (iii) 数学的発見。システムが構築を提案し、境界を改善し、未解決の問題への攻撃を支援します。 (iv) CoT プロンプト、ツールの使用、プロセス報酬モデル、RLVR など、生成と検証をますます結び付ける推論およびトレーニング時の手法。私たちは、小学校の算数、競技数学、幾何学、形式的証明、マルチモーダルおよび多言語推論、専門家の評価にわたる主要なベンチマークをカタログ化し、ベンチマークの飽和、汚染、レポートの不一致、および pass@1、多数決、検証者支援 pass@$k$ の区別を調べます。私たちは、摂動下での脆弱性、報酬ハッキング、マルチモーダル接地障害、脆弱な形式化、推論規模の推論のエネルギーコストなどの障害モードを批判的に評価します。現役の数学者からの最近の視点を活用して、検証された発見のワークフロー、推論の効率、AI 支援による形式化を広く利用できるようにするインフラストラクチャを中心とした将来の方向性を特定します。関連資料: https://github.com/Starscream-11813/awesome-AI4Math。
原文 (English)
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem within NLP to one of the most consequential AI frontiers. This survey provides a unified account of the field's evolution, from early rule-based math word problem (MWP) solvers and template-driven geometry systems, through neural expression generation and LLM prompting, to contemporary reasoning models, multi-agent systems, neuro-symbolic theorem provers, and verified discovery workflows. We organize the landscape along four axes: (i) informal reasoning over text and diagrams, spanning MWP solving, multimodal geometry, and VLMs; (ii) formal reasoning in proof assistants, including autoformalization, tactic prediction, compiler-guided repair, and proof search; (iii) mathematical discovery, where systems propose constructions, improve bounds, or assist attacks on open problems; and (iv) the inference and training-time techniques, including CoT prompting, tool use, process reward models, and RLVR, that increasingly connect generation with verification. We catalog major benchmarks across grade-school arithmetic, competition mathematics, geometry, formal proving, multimodal and multilingual reasoning, and expert evaluation, and we examine benchmark saturation, contamination, reporting mismatches, and the distinction between pass@1, majority voting, and verifier-assisted pass@$k$. We critically assess failure modes: brittleness under perturbation, reward hacking, multimodal grounding failures, fragile formalization, and the energy cost of reasoning-scale inference. Drawing on recent perspectives from working mathematicians, we identify future directions centered on verified-discovery workflows, reasoning efficiency, and infrastructure to make AI-assisted formalization broadly usable. Companion materials: https://github.com/Starscream-11813/awesome-AI4Math.
ComplexConstraints and Beyond: Expert Rubrics for RLVR
Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-w…
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, brow…
ウェアラブルデバイス上のEEG解析のための深層学習モデルの複雑さを軽減する
ウェアラブル ヘルスケア デバイスは、モノのインターネット (IoT) 分野で最も急速に成長しています。多くの自動ヘルスケア サービスは、2 つの重要な生物学的信号、つまり ECG と EEG に依存しており、それぞれ心臓と脳の活動を反映しています。ディープ ニューラル ネットワークは、これらの信号を処理および分析するための主な方法と考えられていますが、ウェアラブル デバイスのエネルギーと計算能力の非常に厳しい制約は、DNN モデルの計算、エネルギー、およびメモリ帯域幅の要求をはるかに下回っており、そのため、多くの実際のウェアラブル サービスでのディープ ラーニングの導入が妨げられています。この論文では、リソースに制約のあるウェアラブル デバイスに最先端の DNN モデルを展開する実現可能性を調査します。特に、パラメーターの量子化と電極削減法が使用される場合の DNN の精度と計算の複雑さの間のトレードオフを調査します。私たちの調査は、EEG 信号分析、特にてんかん発作の検出用に設計されたいくつかの最先端の DNN モデルに重点を置いています。私たちの調査結果は、これらの技術を慎重に適用すると、精度への悪影響を最小限に抑えながら、検討中の DNN の複雑さを大幅に軽減できることを示しています。これらの結果は、DNN ベースのオンライン EEG 分析をウェアラブル デバイスに適応させるときに遭遇する、精度と複雑さの軽減との間の明確なトレードオフを明らかにしています。
原文 (English)
Reducing the Complexity of Deep Learning Models for EEG Analysis on Wearable Devices
Wearable healthcare devices are the fastest-growing Internet of Things (IoT) sector. Many automated healthcare services rely on two crucial biological signals, namely ECG and EEG, which reflect the activity of the heart and brain, respectively. Although deep neural networks are considered the primary way to process and analyze these signals, the very tight energy and computational power constraints in wearable devices are far below the computational, energy, and memory bandwidth demands of DNN models, thereby impeding the deployment of deep learning in many practical wearable services. This paper investigates the feasibility of deploying state-of-the-art DNN models in resource-constrained wearable devices. Notably, we explore the trade-off between accuracy and computational complexity of DNNs when parameter quantization and electrode reduction methods are used. Our investigation centers on several state-of-the-art DNN models designed for EEG signal analysis, specifically for detecting epileptic seizures. Our findings demonstrate that, when applied judiciously, these techniques can significantly reduce the complexity of the DNNs under consideration with minimal adverse effects on accuracy. These results reveal the explicit trade-offs between accuracy and complexity reduction encountered when adapting DNN-based online EEG analysis for wearable devices.
医療ヒューリスティック学習: 解釈可能かつ監査可能な臨床意思決定ルールのための LLM 主導のフレームワーク
臨床表データの予測モデリングは臨床意思決定支援の中心であるため、強力な予測パフォーマンスだけでなく、透過的な意思決定ロジックも必要となります。ディープラーニングとツリーベースのアンサンブル手法は高精度を達成できますが、そのブラックボックス的な性質が臨床導入にとって依然として大きな障害となっています。この課題は、限られたサンプルサイズ、深刻なクラスの不均衡、診断基準や臨床文書の変更から生じる特徴の進化など、医療データの共通の特徴によってさらに悪化します。これらの問題に対処するために、臨床表予測のための勾配を超えた学習パラダイムのインスタンス化である医療ヒューリスティック学習 (MHL) を提案します。 MHL は、ニューラル ネットワークの重み更新に依存する代わりに、統計プローブ、医療知識プローブ、ルール合成、コード レベルの反復改良を統合する大規模言語モデル (LLM) 駆動のワークフローを使用して、決定論的で実行可能な意思決定システムを最適化します。結果として得られるモデルは、不透明なパラメーターとしてではなく、明示的に解釈可能で完全に監査可能で、臨床的に根拠のあるバージョン管理された純粋な Python 決定ルールとして表現されます。 MHL は、以前に検証されたルールから開始し、データ ドリフトまたは機能進化の下で更新された機能情報を使用してルールを繰り返し修正することにより、継続的な学習もサポートします。医療データセットに関する包括的な実験では、MHL がサンプルが少なく不均衡が非常に悪い設定でも強力な動作を維持しながら、最先端の手法に匹敵するパフォーマンスを達成することが示されています。この結果はさらに、この明示的なルール更新メカニズムが、機能の進化の下での壊滅的な忘却の軽減に役立つことを示しています。全体として、これらの発見は、非勾配ベースのヒューリスティック システムが、一か八かの臨床意思決定支援のための透明性と適応性のある代替手段を提供することを示唆しています。
原文 (English)
Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules
Predictive modeling for clinical decision support requires not only strong predictive performance but also transparent decision logic. Although deep learning and tree-based ensemble methods can achieve high accuracy, their black-box nature remains a major obstacle to clinical deployment. This challenge is further compounded by common characteristics of medical data, including limited sample sizes, severe class imbalance, and feature evolution arising from changes in diagnostic criteria and clinical documentation. To address these issues, we propose Medical Heuristic Learning (MHL), an instantiation of the learning beyond gradients paradigm for clinical prediction from structured medical data. Instead of relying on neural network weight updates, MHL uses a large language model (LLM) driven workflow that integrates statistical probes, medical knowledge probes, rule synthesis, and code-level iterative refinement to optimize a deterministic and executable rule-based expert system. The resulting model is expressed not as opaque parameters, but as versioned pure Python decision rules that are explicitly interpretable, fully auditable, and clinically grounded. MHL also supports continual learning by starting from previously validated rules and iteratively revising them using updated feature information under data drift or feature evolution. Comprehensive experiments on medical datasets show that MHL achieves performance comparable to state-of-the-art methods while maintaining strong behavior in small-sample and highly imbalanced settings. The results further indicate that this explicit rule-update mechanism can help alleviate catastrophic forgetting under feature evolution. Overall, these findings suggest that non-gradient-based heuristic systems offer a transparent and adaptable alternative for high-stakes clinical decision support.
Kairos: 物理 AI 用のネイティブ ワールド モデル スタック
世界モデルは、受動的なビジュアル ジェネレーターから物理 AI の基礎的な運用インフラストラクチャに移行しています。世界モデルは、異種混合の経験から世界の知識をネイティブに取得し、長期にわたって永続的な状態を維持し、実際の展開上の制約内で効率的に実行する必要があります。これらの要件に基づいて設計されたネイティブ ワールド モデル スタックである Kairos を紹介します。 (1) カイロスは、オープンワールドのビデオ、人間の行動データ、およびロボットの相互作用を漸進的な発達経路に編成する、クロスエンボディメント データ カリキュラムによって管理されるネイティブの事前トレーニング パラダイムを開拓することによって世界を学びます。 (2) Kairos は、ハイブリッド線形時間的注意を備えたネイティブ統合アーキテクチャ内で統一された世界の理解、生成、予測によって世界を維持します。スライディング ウィンドウの注意はローカル ダイナミクスを捕捉し、拡張されたスライディング ウィンドウは中間範囲の依存関係を捕捉し、ゲートされた線形注意は永続的なグローバル メモリを維持します。我々は、この時間因数分解が誤差の蓄積を厳密に制限し、拡張された範囲にわたる状態の伝播を数学的に保証することを実証する正式な理論的限界を確立します。 (3) Kairos は、実世界の観察、アクション、フィードバック ループのためのサーバーおよび消費者グレードのハードウェア上での低遅延ロールアウト生成をサポートする、展開を意識したシステム協調設計を組み込むことによって世界を運営します。具現化された世界モデル、長期計画、およびアクション ポリシーのベンチマークに関する実験では、Kairos が効率性と能力の強力なトレードオフを提供しながら、トップレベルのパフォーマンスを達成していることが示されています。これらの結果を総合すると、カイロスは将来の自己進化する物理的知性のための統合された運用基盤として位置づけられます。
原文 (English)
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
We introduce \textbf{Kairos}, a regret-aware native world-action model stack for Physical AI. Kairos is motivated by the view that a physical world model should not aim to fully simulate all future pixels, but should learn and maintain the information most relevant to embodiment control: object state, spatial relations, contact conditions, task progress, action consequences, failure boundaries, and deployment uncertainty. Kairos establishes three model-side prerequisites toward this goal. First, it \textbf{learns} control-relevant information through a \textbf{Cross-Embodiment Data Curriculum}, which organizes open-world videos, human behavioral data, and robot interactions into an intervention-strength progression from passive physical observation to intentional behavior and embodied action grounding. Second, it \textbf{maintains} control-sufficient states through a unified \textbf{understanding, generation, and prediction architecture} equipped with \textbf{Hybrid Linear Temporal Attention}, where local, mid-range, and global temporal pathways support multi-timescale state maintenance under efficient inference. Third, it \textbf{deploys} these states through a \textbf{Deployment-Aware System Co-Design}, treating latency, memory footprint, and hardware compatibility as first-order constraints for future observation, action, and feedback loops. Experiments on embodied world-model benchmarks, world-action benchmarks, long-horizon generation, and inference-efficiency evaluation show that Kairos achieves superior performance while offering a favorable efficiency to capability trade-off.
Skill Coverage: A Test Adequacy Metric for Agent Skills
Agent skills encode reusable procedural knowledge for large language model (LLM) agents, and existing benchmarks show that such skills can…
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation
Choreographic motion generation poses unique challenges for AI, demanding precise semantic control over complex, temporally structured, and…
BlockTrain を使用した分散型 AI トレーニングと推論
フロンティア AI トレーニングは、集中管理された高密度のアクセラレータ クラスターへのアクセスによってますます形作られています。これにより、ハイパースケーラーや大規模な集中研究所にとって構造的な利点が生まれ、オープンまたは独立した AI の取り組みが、希少な資本、特権的なインフラストラクチャ、およびデータセンターの地理に依存することになります。我々は、モデルが独立してトレーニング可能なブロックに分割され、それぞれが同じグローバル ターゲットから派生したローカル目標に基づいて最適化され、推論時に 1 つのモデルに構成される分散トレーニング プロトコルである Spheroid BlockTrain を紹介します。バイトレベルの WikiText では、同じセットアップのエンドツーエンド Transformer 参照の約 0.04 CE 以内で、BlockTrain はクロス エントロピー 1.359 (複雑度 3.89) に達しますが、アクティブな各ワーカーは 1 つのブロックのみをトレーニングし、フルモデル オプティマイザー状態を回避します。共有 6 ワーカー ブロックのトレーニング実行は、同じブロックの更新を 1 つの組み立てられたモデルに平均化することで CE 1.385 に達します。 HTTP/TCP トランスポート実験では、実際のシリアル化されたチェックポイントと更新を移動します。これには、15.22 GB を移動しながら CE を 5.580 から 1.811 に改善するパブリック IP 3 ホストの実行が含まれます。推論のために、現在の BlockTrain パスは完全な出力ごとに 1 つのブロック スタック トラバーサルを使用し、最大 75.80B パラメーターの論理 fp16 シェイプまでの 3 つのパブリック ネットワーク GPU ホストにわたる直接 TCP を介してサービスを提供します。これは、トラバーサルごとに 1 つのトークンではなく、WAN パイプライン トラバーサルごとに完全なシーケンスを出力するため、一致するプレーン自己回帰 TCP パイプライン ベースラインよりも優れたパフォーマンスを発揮します。
原文 (English)
Decentralised AI Training and Inference with BlockTrain
Frontier AI training is increasingly shaped by access to dense, centrally controlled accelerator clusters. This creates a structural advantage for hyperscalers and large centralized laboratories, and makes open or independent AI efforts depend on scarce capital, privileged infrastructure, and data-center geography. We present Spheroid BlockTrain, a decentralized training protocol in which a model is partitioned into independently trainable blocks, each optimized on a local objective derived from the same global target and composed at inference into one model. On byte-level WikiText, BlockTrain reaches cross entropy 1.359 (perplexity 3.89), within about 0.04 CE of a same-setup end-to-end Transformer reference, while each active worker trains only one block and avoids full-model optimizer state. A shared six-worker block training run reaches CE 1.385 by averaging same-block updates into one assembled model. HTTP/TCP transport experiments move real serialized checkpoints and updates, including a public-IP three-host run that improves CE from 5.580 to 1.811 while moving 15.22 GB. For inference, the current BlockTrain path uses one block-stack traversal per full output and serves over direct TCP across three public-network GPU hosts up to a 75.80B-parameter logical fp16 shape, outperforming a matched plain-autoregressive TCP pipeline baseline because it emits a full sequence per WAN pipeline traversal rather than one token per traversal.
BluTrain: AI システム用の C++/CUDA フレームワーク
大規模な深層学習の進歩は、モデリングよりもシステム エンジニアリングの問題です。トレーニング中のモデルの動作 (スループット、メモリ フットプリント、結果の数値的忠実度) は、アーキテクチャ自体によって決まるというよりは、そのアーキテクチャがハードウェア上でどのように表現されるかによって決まります。システムの複雑さを抽象化してモデリングをシームレスにし、反復的なオーケストレーション ロジックの必要性を排除しながら、このハードウェア表現に対する絶対的な制御を実現するために、BluTrain は、標準 C++ およびコア CUDA プログラミング モデルにおける堅牢で軽量なアーキテクチャ全般のトレーニング フレームワークとして第一原理から設計されました。リバースモード autograd を備えた型付きテンソル モジュール、線形代数ライブラリ、キャッシュ アロケータ、マルチモード分散実行モジュール、MLIR ベースの深層学習コンパイラなど、すべての層がネイティブに実装されています。 8 GPU 6000 Ada システム上の FP32 で 124M パラメータの GPT-2 ベースラインをトレーニングする正式な評価では、BluTrain は、スループット (平均 407K トークン/秒を維持するのに対し、PyTorch の 395K トークン/秒を維持) とメモリ効率 (最大 22% のフットプリント削減を達成) の両方で業界標準のベースラインを上回っています。数値的忠実度が向上し、最終的な検証損失がわずかに低下するように収束します。すべてのレイヤーがネイティブ チューニングに対して明示的にオープンであるため、パフォーマンスの上限はフレームワーク自身で引き上げることができます。
原文 (English)
BluTrain: A C++/CUDA Framework for AI Systems
Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that architecture is expressed on the hardware. To achieve absolute control over this hardware expression while abstracting away systems complexity to make modelling seamless and eliminating the need for repetitive orchestration logic, BluTrain was architected from first principles as a robust, lightweight, and architecture-general training framework in standard C++ and the core CUDA programming model. Every layer is implemented natively: a typed tensor module with reverse-mode autograd, a linear-algebra library, a caching allocator, a multi-mode distributed-execution module, and an MLIR-based deep-learning compiler. In formal evaluations training a 124M-parameter GPT-2 baseline in FP32 on an 8-GPU 6000 Ada system, BluTrain outperforms industry-standard baselines in both throughput (sustaining an average of 407K tokens/s versus PyTorch's 395K tokens/s) and memory efficiency (achieving up to a 22% footprint reduction), while strictly preserving numerical fidelity and converging to a marginally lower final validation loss. With every layer explicitly open to native tuning, the performance ceiling is the framework's own to raise.
Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval
Foundation-model agents in multi-step, open-ended environments frequently suffer from compounding errors, where early mistakes contaminate…
Autodata: An agentic data scientist to create high quality synthetic data
We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation dat…
グラフ ワールド モデルのロールアウト エラーについて
ワールド モデルは、学習したダイナミクスをロールフォワードすることによって計画を立てるためによく使用されます。ただし、多くの計画環境はベクターやイメージではありません。これらは、エージェント、ツール、スキル、ルート、依存関係のグラフです。これらの設定では、局所的な予測誤差が局所的にとどまるか、グラフ全体に広がる可能性があり、エッジが固定されるのではなく予測されると、故障モードが再び変化します。この論文では、グラフ ワールド モデル (GWM) における長期ロールアウト エラーを研究します。ノード、エッジ、およびグラフレベルの意思決定のためのアクションノードを備えた統合された固定エッジおよび動的エッジ GWM フレームワークを定式化します。トポロジに起因する増幅をモデルに起因する増幅から分離するグラフ値のロールアウト境界を開発し、動的エッジ ロールアウト用にノードとエッジの結合演算子を導入します。分析に基づいて、スペクトルの正則化、ロールアウトの一貫性、クリティカル ノードの重み付けを組み合わせたエラー認識 GWM を提案します。合成トポロジと異種エージェント グラフ テストベッドでは、ロールアウト エラーと計画の後悔が時間とともに増大し、構造が進化する場合には動的エッジ トレーニングが必要になります。また、エラー認識 GWM は、予測精度を維持しながら長期的な発散を防ぎます。現実世界のグラフ ベンチマークは、GWM の範囲を明確にします。GWM は動的なグラフ ロールアウトとエージェント プランニングに最も役立ちますが、特殊なグラフ モデルは静的または疎な予測タスクに引き続き強力です。
原文 (English)
Understanding Rollout Error in Graph World Models
World models are increasingly used for planning, yet most analyses of rollout error assume vector-valued states and scalar error amplification. Many planning environments, however, are naturally graph-structured: agents, tools, skills, routes, and dependencies interact through evolving relations. In this work, we study how prediction errors accumulate in Graph World Models (GWMs). We formulate fixed-edge and dynamic-edge GWM rollouts under a unified state-action transition framework and derive topology-aware error bounds. For fixed-edge rollouts, we show that long-horizon node error separates into a topology factor, governed by the graph spectral radius, and a model factor, governed by layer spectral norms. For dynamic-edge rollouts, we introduce a joint node-edge error operator that captures feedback between feature prediction and structure prediction, revealing when edge errors amplify future message passing. Motivated by these bounds, we propose Error-Aware GWM, a training objective that combines spectral regularization, rollout consistency, and critical-node weighting. Across synthetic graph topologies and heterogeneous agent-graph testbeds, we find that rollout error and planning regret grow with horizon, that dynamic-edge training is necessary when structure evolves, and that Error-Aware GWM improves long-horizon stability without sacrificing one-step accuracy. Our results characterize when graph world models remain reliable under autoregressive planning and when topology makes them fail.
グラウンディングされた反復言語計画: パラメーター化された世界モデルが LLM エージェントにおける幻覚の伝播をどのように軽減するか
言語エージェントの世界モデルには 2 つの便利な形式があります。エージェントベースの世界モデルは LLM API を呼び出し、言語で柔軟に推論しますが、そのエラーは幻覚的な状態変化として現れ、通常の回帰損失ではスコアを付けるのが困難です。パラメーター化された世界モデルは、トレーニングされた遷移予測子です。そのエラーは、NodeMSE、デルタ精度、妥当性精度などの量を使用して測定する方が簡単ですが、通常、スタンドアロン プランナーとしては弱いです。これら 2 つのファミリーを 4 つのグラフ構造の計画ベンチマークで比較し、エージェントベースのケースの操作上の幻覚測定基準を導入します。この比較により、\textbf{Grounded Iterative Language Planning} (GILP) が動機付けられ、小規模なパラメーター化されたバックボーンのみをトレーニングし、それを API ベースのエージェント推論と組み合わせます。バックボーンは、有効なアクション、予測された状態デルタ、リスク、および価値を提供します。 LLM はアクションと想像上のデルタを草案します。そして、この 2 つが一致しない場合には、整合性ゲートが修正を要求します。実際の GPT-4o-mini 通話では、GILP は幻覚状態の割合を 0.176 から 0.035 に削減します。調整されたシミュレーターアブレーションでは、最大 22% の余分な LLM コールを追加するだけで、成功率が 0.668 から 0.838 に上昇します。
原文 (English)
Agent vs. Parametric World Models: Hybrid Planning for Reliable Language Agents
Language agents plan by generating not only actions but also implicit predictions of how the world will change. These imagined state updates make agents flexible, but they also create a distinct failure mode: hallucinated state claims can be written into context and propagated across subsequent decisions. In contrast, parametric world models provide measurable transition errors but are often weaker semantic planners. We study this tradeoff in graph-structured planning environments and introduce metrics for agent-world-model error, including hallucinated-state rate, propagation depth, and long-horizon error growth. We then propose Hybrid World-Model Planning (Hybrid-WM), which keeps the language model as the planner while using a small parametric transition model to predict action validity, state deltas, risk, and value. A consistency gate compares the agent's imagined delta with the parametric prediction and triggers targeted revision only under disagreement. Across four graph-structured planning benchmarks, Hybrid-WM improves success while reducing hallucinated state propagation. In live GPT-4o-mini evaluations, it reduces hallucinated-state rate from 0.176 to 0.035; in calibrated simulator ablations, it improves success from 0.668 to 0.838 with modest additional inference. These results suggest that lightweight parametric transition models can serve as effective grounding mechanisms for language-agent planning without replacing semantic reasoning.
交通工学実践のためのカスタマイズされた生成 AI エージェント: 開発および継続的な事前トレーニング ガイドライン
生成人工知能 (AI) と大規模言語モデル (LLM) の最近の進歩により、複雑な推論、要約、質問応答タスクの自動化において大きな期待がもたれています。ただし、技術標準、エンジニアリング用語、およびドメイン固有のセマンティクスへの理解が不十分であるため、特殊なエンジニアリング ドメインにおける汎用 LLM の有効性は依然として限定的です。この研究では、輸送工学アプリケーション向けにカスタマイズされた生成 AI エージェントを開発するための体系的なアプローチを提案します。米国の輸送マニュアル、設計ガイドライン、規制文書の厳選されたコーパスを使用して、統合低ランク適応 (LoRA) フレームワークを通じて 6 つの最先端の LLM の継続的な事前トレーニングが実施されます。トレーニング プロセスは、収束とモデルの安定性を確保するために監視されます。パフォーマンスは、BLEU-4 や ROUGE などの標準的な自然言語処理メトリクスを使用して評価され、Qwen2.5-7B および LLaMA-3.1-8B が最高のドメイン アラインメントと応答品質を示しています。結果は、技術的な内容の解釈とコンテキスト固有の推論における LLM パフォーマンスの向上における LoRA ベースの適応の有効性を検証しました。この成果は、ドメインに特化した生成 AI エージェントを構築するための再現可能な開発フレームワークに貢献し、交通研究、設計、計画、政策分析における広範な展開をサポートします。
原文 (English)
Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline
Recent advancements in generative artificial intelligence (AI) and large language models (LLMs) have shown significant promise in automating complex reasoning, summarization, and question-answering tasks. However, the effectiveness of general-purpose LLMs in specialized engineering domains remains limited due to insufficient exposure to technical standards, engineering terminology, and domain-specific semantics. This study proposes a systematic approach to developing a customized generative AI agent for transportation engineering applications. A curated corpus of U.S. transportation manuals, design guidelines, and regulatory documents is used to conduct continued pretraining of six state-of-the-art LLMs through a unified low-rank adaptation (LoRA) framework. The training process is monitored to ensure convergence and model stability. Performance is evaluated using standard natural language processing metrics, including BLEU-4 and ROUGE, with Qwen2.5-7B and LLaMA-3.1-8B demonstrating the highest domain alignment and response quality. Results validate the effectiveness of LoRA-based adaptation in improving LLM performance on technical content interpretation and context-specific reasoning. This work contributes a reproducible development framework for constructing domain-specialized generative AI agents, supporting broader deployment in transportation research, design, planning, and policy analysis.
FADE: 大規模な視覚言語モデルにおける言語優先支配を軽減することによる幻覚の軽減
Large Vision-Language Model (LVLM) の優れた機能にもかかわらず、依然として幻覚の影響を受けやすく、入力画像と一致しないコンテンツが生成されます。最近の研究では、これは視覚入力に対する言語事前の優位性によるものであり、この優位性を緩和するために対照的なデコード方法が採用されていますが、そのメカニズムの起源は未解明のままです。各変換層を通る情報の流れを調査すると、アテンション モジュールが一貫して視覚的証拠を集約し、クリティカル層の FFN モジュールが言語事前情報のソースとして機能することがわかりました。これらの事前分布は視覚的な証拠を無効にする可能性があり、中間層での正しい予測が不正確な出力に向かってドリフトする原因となります。この洞察に基づいて、言語優先の優位性を減らすために FFN 出力を減衰するトレーニング不要の方法である FADE (FFN Attenuation for DEcoding) を提案します。 LLaVA-1.5、mPLUG-Owl2、および InstructBLIP にわたる POPE、CHAIR、および MME ベンチマークの評価では、FADE が推論効率を維持しながら幻覚を効果的に軽減することが示されています。
原文 (English)
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent studies attribute this to the dominance of language priors over visual inputs and employ contrastive decoding methods to mitigate this dominance, but the mechanistic origin remains unexplored. We investigate the information flow through each transformer layer and find that attention modules consistently aggregate visual evidence, while FFN modules at critical layers act as the source of language priors. These priors can override visual evidence, causing correct predictions in intermediate layers to drift toward incorrect outputs. Based on this insight, we propose FADE (FFN Attenuation for DEcoding), a training-free method that attenuates FFN outputs to reduce language-prior dominance. Evaluations on POPE, CHAIR, and MME benchmarks across LLaVA-1.5, mPLUG-Owl2, and InstructBLIP show that FADE effectively mitigates hallucinations while preserving inference efficiency.
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents colla…
やめることを学ぶことが役立つのはいつですか?推論モデルにおける早期終了に関するコストを意識した研究
推論モデルはインスタンスごとに異なる量の有用な計算を費やしますが、学習された停止ルールがいつ単純な信頼度や収束のしきい値を超えて改善するかは不明のままです。私たちは、推論言語モデル用の隠れステートフリー チェックポイント ストッパーである LearnStop を使用してこの疑問を研究します。固定予算チェックポイントで、LearnStop は現在の推論プレフィックスから短い回答を調査し、回答の信頼性、エントロピー、プレフィックス投票シェア、回答の安定性、バックトラッキング マーカー密度などのオンライン機能からプレフィックスの正しさを予測します。 GSM8K、MATH-500、MMLU-Pro、AIME-90、GPQA、Qwen3、DeepSeek-R1 蒸留にわたる 18 のタスク モデル設定全体にわたって、答えはタスクによって異なります。自由形式の計算では、学習された複数特徴停止により固定予算フロンティアが改善され、多くの場合スカラー出口を上回ります。Qwen3-32B を備えた GSM8K では、経験的フロンティアは +0.157 のポストホック ピーク適応ゲインに達し、検証で選択された操作点は正のゲインを保持し、最も強いスカラー ベースラインを超えるペアのゲインは +0.028 です。複数選択の非常に難しい設定では、スカラー信頼性、エントロピー、または安定性ルールが競合するか、より強力になります。したがって、学習停止をスカラー出口の普遍的な代替としてではなく、その値が軌道構造に依存するツールとして組み立てます。さらに、検証で選択された動作ポイント、ペア ブートストラップ テスト、有限グリッド ロストコレクト リスク校正、KV フォーク、プレフィックス キャッシュ、およびブラック ボックス レジームに基づくコスト計算、H100 サービング プロファイル、チェックポイント スケジュール スイープ、転送分析、および堅牢性チェックを提供します。主な実用的な発見は、学習された停止が、予算がいっぱいになる前に多くの問題が正解したが、信頼できるスカラー停止信号が 1 つも示されない場合に役立つということです。信頼度または答えの収束が停止問題をすでに解決している場合、その利点はほとんど失われます。
原文 (English)
When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models
Reasoning models spend test-time compute unevenly across instances, and a growing family of early-exit rules -- confidence thresholds, entropy monitors, answer-stability checks, and learned stoppers -- promises to reclaim the waste. These rules, however, are evaluated under heterogeneous protocols that leave the deployment question unanswered: at a fixed tolerance for losing correct answers, which policy saves more compute, and does the saving survive probe overhead? We answer this question with a controlled study across 18 task-model settings spanning GSM8K, MATH-500, MMLU-Pro, AIME-90, and GPQA on Qwen3 and DeepSeek-R1-distilled models, using LearnStop, a hidden-state-free logistic stopper over prefix-observable features, as the learned policy instrument. Under matched lost-correct risk at $\alpha$ = 0.15, with the scalar competitor selected on calibration data from confidence, entropy, confidence-leap, and run-stability exits, the answer forms three regimes. Learned stopping wins on all four primary Qwen3 free-form math settings (+3.2 to +21.2 pp additional total-token saving); calibrated scalar exits win on multiple-choice MMLU-Pro; and small hard benchmarks (AIME-90, GPQA) admit no certifiable aggressive policy at all. A trajectory decomposition predicts the regime: learning pays where answers oscillate and correctness evidence is spread across complementary signals, while a single confidence threshold suffices where most instances are already solved at the first checkpoint. Cost accounting sharpens the picture further -- the same policy that saves 32% of tokens under KV-cache forking costs 121% extra under black-box repeated prefilling. Together, these results replace the single-method race with a decision procedure for choosing a stopping rule from the trajectory structure and serving regime of the target workload.
税金を意識したパーソナライズされたポートフォリオ管理のための 3 段階の基礎モデル
我々は、これまでの金融 RL 作業すべてに共通する 3 つの制限、1) ティッカー ロックイン、2) モノリシック目標、3) 静的ユーザー モデルに対処する、パーソナライズされたポートフォリオ管理のための 3 段階の深層強化学習システムを紹介します。フェーズ 1 では、マルチアセット コーパスでの自己教師あり学習を介して、ティッカー ID のないクロスアセット エンコーダーを事前トレーニングします。学習されたゲート メカニズムを介して融合された、T5 ベースの時系列基盤モデルである Chronos を使用した凍結並列ブランチによって強化されます。私たちの知る限り、これはポートフォリオ管理 RL への時系列基礎モデルの最初の適用です。エンコーダーは、新しいティッカーの再トレーニングを必要としない、50 次元の観察可能なメタデータ ベクトルを介して、あらゆる公開取引資産に一般化します。フェーズ 2 では、エピソードごとにサンプリングされた 6 つの異なる投資目標 (短期アルファ、短期利益、長期利益、資本保全、税金損失の回収、および長期利益のみ) を同時に提供する目的条件付き報酬の下で、MoE (専門家混合) のポートフォリオ アクター評論家を PPO で微調整します。 MoE アーキテクチャは、各目標を専門のエキスパート ヘッド (勢い、成長、守備、税務意識) に割り当て、学習型インテント ルーターがアクティブな目標と現在の市場体制に基づいて専門家をブレンドすることで、目標間の勾配の競合を排除します。フェーズ 3 では、実際の証券取引履歴に基づいて微調整された 76 パラメーターの LoRA モジュールを介して推論時に各個人にさらに適合する軽量のパーソナライゼーション レイヤーを追加し、アンケートではなく明らかになった取引行動から投資目標を推測します。自然言語インテント パーサーは、自由形式の目標を構造化された投資目標パラメータに直接変換します。
原文 (English)
A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management
We present a three-phase deep reinforcement learning system for personalized portfolio management that addresses three limitations shared by all prior financial RL work: 1) ticker lock-in, 2) monolithic objectives , and 3) static user models. Phase 1 pretrains a ticker-identity-free cross asset encoder via self-supervised learning on a multi-asset corpus, augmented by a frozen parallel branch using Chronos, a T5-based time series foundation model, fused via a learned gating mechanism. To our knowledge, this is the first application of a time series foundation model to portfolio management RL. The encoder generalizes to any publicly traded asset via a 50-dimensional observable metadata vector that requires no retraining for new tickers. Phase 2 fine-tunes a MoE (Mixture of Experts) portfolio actor critic with PPO under an objective-conditioned reward that simultaneously serves six distinct investment goals sampled per episode: short-term alpha, short-term gain, long-term gain, capital preservation, tax-loss harvesting, and long-term-gains-only. A MoE architecture assigns each objective to a specialized expert head (momentum, growth, defensive, tax-aware), and a learned intent router blends experts based on the active objective and current market regime, which eliminates cross-objective gradient conflict. Phase 3 adds a lightweight personalization layer further adapted at inference time to each individual via a 76-parameter LoRA module fine-tuned on real brokerage transaction history, inferring investment objectives from revealed trading behavior rather than questionnaires. A natural language intent parser converts free-form goals directly into structured investment objective parameters.
相転移としての世界モデルの崩壊
水は温まるにつれて変化しないように見えますが、臨界点で沸騰します。私たちは、長期的な言語エージェントが暗黙の世界モデルにおいて同様の遷移を示すかどうかを尋ねます。一部のパラメーター設定では、ステート ロードを少量変更するか、ホライズンの 1 ステップを追加すると、動作はほとんど変わりません。臨界境界付近では、同じ小さな変化が突然の世界崩壊を引き起こします。この効果を、ステップごとの正確なゴールド状態を使用した決定論的タスク ファミリで研究します。状態カーディナリティ、依存関係密度、ホライズン、分岐、観察モード、突然変異率に関する大規模なグリッド検索により、解決されたプラトー、狭い遷移バンド、および崩壊フロアなどの状態図が明らかになります。ステップごとのトレースはメカニズムを示します。ワールド状態の忠実性はアクションの有効性の前に失敗するため、エージェントは単に間違ったアクションを選択しているだけではありません。それは腐敗した世界から行動しているのです。より強力なモデルは臨界境界を変換しますが、定性的な移行は除去しません。これらの結果により、世界モデルの崩壊が長期的なエージェントにとって目に見えるボトルネックとなっています。
原文 (English)
World-Model Collapse as a Phase Transition
Water looks unchanged as it warms, then at a critical point it boils. We ask whether long-horizon language agents show an analogous transition in their implicit world models. In some parameter settings, changing state load by a small amount, or adding a single step of horizon, leaves behavior nearly unchanged; near a critical boundary, the same small change causes a sudden world collapse. We study this effect in a deterministic task family with exact per-step gold state. A large grid search over state cardinality, dependency density, horizon, branching, observation mode, and mutation rate reveals a phase diagram: a solved plateau, a narrow transition band, and a collapse floor. Per-step traces show the mechanism: world-state fidelity fails before action validity, so the agent is not merely choosing a bad action; it is acting from a corrupted world. Stronger models translate the critical boundary but do not remove the qualitative transition. These results make world-model collapse a measurable bottleneck for long-horizon agents.
行動する前に世界に問う:世界モデルのキャリブレーションのための予算環境調査
長期的な言語エージェントは、アクションを選択するだけではありません。彼らは、ある決断から次の決断まで、世界のプライベートモデルを持ち続けています。モデルがドリフトすると、失敗したアクションが実行される前に、その後の失敗が決定される可能性があります。私たちは直接修復メカニズムを研究します。次のタスクのアクションにコミットする前に、エージェントは環境に 1 つの信念フィールドについて質問し、その答えをワールド モデルに書き戻すことができます。このため、環境との相互作用は、単にタスクを進めるための手段ではなく、希少なキャリブレーション リソースになります。構造化信念表の予算探索演算子である \method を紹介します。有用なプローブはどこでも同じではありません。ツールの依存関係などの手順上の信念は、多くの場合、対象を絞ったチェックによって修復できますが、それらのチェックにはタスクに必要なステップが費やされます。オブジェクトの位置やグラフの端などの空間的信念は、構造的な手がかりに大きく依存します。世界が画面外で変化するとき、エージェント自身の自信は不十分な指針になる可能性があります。タイプ階層化分析は、このプローブとアクションのフロンティアを形式化し、管理された実験により、プローブ ポリシーがタスクの構造に従う場合、計画途中の環境の証拠によって最終的な世界モデルのエラーが減少することが示されています。
原文 (English)
Ask the World Before Acting: Environment Probing for Calibrated Agent World Models
Language agents acting over long horizons must maintain beliefs about tool states, object locations, graph edges, and subgoal dependencies. When these beliefs drift, failures can be fixed neither by longer reasoning traces nor by ordinary self-reflection, since the missing evidence lies in the environment. We formulate environment probing as a budgeted decision problem for structured agent world models: before acting, the agent may query the current value of one belief field, update its table, and pay one interaction step. We introduce EnvProbe, a simple scoring policy that combines task criticality, staleness, verbalized uncertainty, and dependency role. A type-stratified analysis separates the benefit of belief repair from the cost of displaced task actions and predicts different behavior for procedural and spatial beliefs. In three controlled environments with gold belief states, EnvProbe improves terminal world-state accuracy over periodic probing by 11.76 percentage points on procedural tool-dependency tasks, 3.79 points on spatial tasks, and 6.45 points overall. Ablations show that task-structural terms are the main source of the gains, while self-reported uncertainty is unreliable under confident wrong beliefs. The results suggest that agent calibration should be treated as an action-selection problem over environment evidence, not only as a model-internal reasoning problem.
マルチモーダルの安全性のためにテキストによる拒否指示を活用する
大規模言語モデル (LLM) の安全性を向上させるために、トレーニング後の調整を実行するか、アクティベーション スペースでの拒否指示を利用できます。どちらの戦略もマルチモーダル LLM (MLLM) では実現可能性が低くなります。安全でないマルチモーダル データが必要であり、ユニモーダルの対応する戦略よりも収集が難しいからです。この研究では、この制約を緩和し、LLM バックボーンから直接抽出されたテキストの拒否指示がモダリティ (つまり、画像、ビデオ) 全体で一般化されるかどうかを調査します。予備的な調査結果はこの能力を裏付けていますが、有効性はレイヤーの選択、ステアリング強度、およびクロスモーダルアライメントによって条件付けされ、後者により安全なマルチモーダル入力が誤って拒否に向けて誘導されます。これに基づいて、マルチモーダル安全性データを必要とせずにマルチモーダル安全性を導入する、トレーニング不要の軽量アプローチである Modality-Agnostic Refusal Steering (MARS) を導入します。 MARS は、アクティベーションの再センタリングによってモダリティの不整合を修正し、幾何学的に定義された信頼領域内でステアリング強度を適応的にスケールし、最初に生成されたトークンで動作する最適な介入層を選択します。安全性、ユーティリティ、ビデオ ジェイルブレイク ベンチマークにわたる 5 つの SOTA MLLM で評価された MARS は、ユーティリティを維持しながら一貫した安全性の向上を実現します。これらの結果は、安全関連の構造がモダリティ間で共有されていること、およびテキストによる拒否の指示が、マルチモーダル調整のための強力だがまだ研究されていない基盤であることを明らかにしている。
原文 (English)
Harnessing Textual Refusal Directions for Multimodal Safety
To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feasible in Multimodal LLMs (MLLMs) as they require unsafe multimodal data, harder to collect than their unimodal counterpart. In this work, we relax this constraint and investigate whether textual refusal directions, extracted directly from the LLM backbone, generalize across modalities (i.e., image, video). Preliminary findings confirm this ability, though effectiveness is conditioned by layer selection, steering strength, and cross-modal alignment, with the latter causing safe multimodal inputs to be spuriously steered toward refusal. Building on this, we introduce Modality-Agnostic Refusal Steering (MARS), a light-weight training-free approach that injects multimodal safety without the need for multimodal safety data. MARS corrects modality misalignment via activation re-centering, adaptively scales steering strength within a geometrically defined trust region, and selects the optimal intervention layer, operating at the first generated token. Evaluated on five SOTA MLLMs across safety, utility, and video jailbreak benchmarks, MARS achieves consistent safety gains while preserving utility. These results reveal that safety-relevant structure is shared across modalities and that textual refusal directions are a powerful and underexplored foundation for multimodal alignment.
MMM データ モデル -- 分散型ナレッジ コモンズにおける知識の相互運用性の標準仕様
多くの情報システムはドキュメントを中心に構築されており、印刷物作成とリニア読み取り用に最適化された自己完結型ユニットです。ドキュメント中心の組織は大規模な普及には効果的ですが、知識を構造化し、更新し、共有し、再利用する方法に制約があります。形式的なアプローチはこれらの制限の一部に対処しますが、人間の使いやすさや範囲などの他のシステム特性よりも形式的な構造を優先するため、広範な貢献と採用を達成するのに苦労しています。 AI システムは文書作成を再構築していますが、人間による知識の表現と交換のための従来の文書に代わる統合されたポータブルな代替手段は提供されていません。この論文では、学際的な共同研究の実際的なニーズから生まれた知識文書化のためのデータ モデルである MMM を紹介し、情報システムの設計空間の比較分析の中に位置づけています。 MMM は、一連の規範的な制約とフリーテキスト ラベルの表現の自由を組み合わせたものです。セマンティックな収束を必要とせずに、分野、アプリケーション、展開全体で相互運用できるように設計されています。リファレンス実装とパイロット展開データは、実装可能性と初期の使いやすさを示しています。
原文 (English)
The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons
Many information systems are built around documents: self-contained units optimised for print production and linear reading. While effective for large-scale dissemination, the document-centric organisation constrains how knowledge can be structured, updated, shared, and reused. Formal approaches address some of these limitations but struggle to achieve widespread contribution and adoption due to their prioritisation of formal structure over other system properties such as human usability and scope. AI systems are reshaping document production, but without providing a unified portable alternative to traditional documents for humans' expression and exchange of knowledge. This paper presents MMM, a data model for knowledge documentation that emerged from the practical needs of interdisciplinary collaborative research, and positioned here within a comparative analysis of the design space of information systems. MMM combines a small set of normative constraints with the expressive freedom of free-text labels. It is designed for interoperability across disciplines, applications and deployments without requiring semantic convergence. A reference implementation and pilot deployment data demonstrate implementability and early usability.
Mnemosyne: AI 生成のワークフローを検証および修復するためのエージェント トランザクション処理
LLM、ソルバー、エージェント チームはワークフロー アクション、修復、計画を生成することが増えていますが、生成されたアクションは構文的には有効でも、古い、実行不可能、矛盾している、または修復を引き起こした証拠を破壊する可能性があります。エージェントティック トランザクション処理 (ATP) は、生成されたアクションが、宣言された実行可能な制約セット C の下で決定論的な承認を通過するまで、信頼できない提案として扱うトランザクション モデルです。原則は両面的です。提案は真実ではなく、すべての中断を予測する提案はありません。提案は何でも可能ですが、ランタイムのみが承認してコミットします。予期せぬ中断が発生した場合、新しい提案を信頼するのではなく、範囲内で反応的に修復します。 C と比較すると、コミットされた状態の正確性は、提案層の能力、誠実さ、学習とは無関係になります。私たちは、追加専用の遷移ログ、有効な状態の予測、依存関係に安全な補償、およびアクティブなコミットメント レコードを備えたランタイムである Mnemosyne で ATP を実現し、C に関連する 4 つの安全特性 (権限分離、シリアル等価生成許可、証拠保全修復、および義務の封じ込め) を、その局所的修復プロトコル (LCRP) の制限付きリアクティブ修復保証とともに証明します。再現可能なアーティファクトは、9 回の改ざんテストで対象となる違反を拒否しながら、有効な作業を依然として認め、投影と検証のオーバーヘッドは 6% 未満で、制限されたローカル修復編集では、グローバルな再計算よりも桁違いに少ない操作で済みます。 Mnemosyne はオープンソースです: https://github.com/eyuchang/Mnemosyne/tree/arxiv-atp-rq1-rq9b-r8-v2。
原文 (English)
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
LLMs increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidence that triggered a repair. We introduce Agentic Transaction Processing (ATP), a transaction model that treats generated actions as untrusted proposals until they pass deterministic admission under a declared, executable constraint set C. The governing principle is two-sided: a proposal is not truth, and no proposal foresees every disruption. Anything may propose, but only the runtime admits and commits; when an unforeseen disruption strikes, it repairs reactively within bounds rather than trusting a fresh proposal. Relative to C, committed-state correctness becomes independent of the competence, honesty, or learning of the proposing layer. We realize ATP in Mnemosyne, a runtime with an append-only transition log, effective-state projection, dependency-safe compensation, and active commitment records, and prove four safety properties relative to C (authority separation, serial-equivalent generative admission, evidence-preserving repair, and obligation containment) plus a bounded-reactive-repair guarantee (LCRP). A reproducible artifact rejects the targeted violations across nine falsification tests while still admitting valid work, at under 6% overhead, and local repair edits an order of magnitude fewer operations than global recompute. In live-proposer pilots, 80 static plan-entry and mid-execution repair proposals from four heterogeneous LLMs pass the same admission boundary, scored by an external cross-episode harness with zero invalid commits; the gate admits 24 of 40 live repair proposals and rejects 16, four as explicit safety rejections of over-broad rollback. Mnemosyne is open source: https://github.com/eyuchang/Mnemosyne/tree/mnemosyne-atp-postgres-rerun
AI ネイティブ ゲーム: 調査とロードマップ
生成 AI により、ゲームが実行時に対話、クエスト、キャラクター、画像、世界を生成できるようになりました。ただし、生成だけではゲームが AI ネイティブになるわけではなく、プレイアビリティが保証されるわけでもありません。この論文では、ランタイム生成 AI がコア ループを構成するかどうかによって AI ネイティブ ゲームを定義します。AI コンポーネントが削除されたり、些細な置き換えが行われたりすると、プレイの中心的な形式が崩壊するか、根本的に異なるものになります。この反事実基準は、AI ネイティブ ゲームを、AI 拡張ゲーム、境界アーティファクト、チャットボット、居酒屋スタイルのロールプレイ、手続き型コンテンツ生成、および AI 支援制作から区別します。この定義を使用して、候補となる成果物をスクリーニングし、公開されている 53 の AI ネイティブ ゲームとプロトタイプを分析します。二重軸の G/N 分類法を導入します。G 軸はプレイヤーが直面するゲームのタイプを捉え、N 軸は生成 AI をプレイに不可欠なものにする支配的な AI メカニズムを捉えます。このコーパスは言語優先デザイン、特に物語的冒険、認識論的相互作用、生成的物語を中心に集中していますが、意味論的判断、マルチエージェント シミュレーション、生成的構築、関係/仲間遊びなどのカテゴリーはあまり代表されていません。私たちは、設計上の中心的な問題は、セマンティックなオープン性を安定したゲームプレイに組み込むことであると主張します。 AI ネイティブの設計は、目標、ルール、状態、フィードバック、ペーシング、プレーヤーの主体性などの機械的不変条件に依存しており、これらにより、オープンエンドの AI 出力が解釈可能で結果的なものになります。最後に、制御可能な発電、機械としての AI 設計、マルチモーダルおよびマルチエージェント システム、推論の経済学、評価、安全性、規制のロードマップを示します。
原文 (English)
AI Native Games: A Survey and Roadmap
Generative AI now enables games to produce dialogue, quests, characters, images, and worlds at runtime. Yet generation alone does not make a game AI-native, nor does it guarantee playability. This paper defines AI-native games by whether runtime generative AI is constitutive of the core loop: if the AI component were removed or trivially replaced, the central form of play would collapse or become fundamentally different. This counterfactual criterion separates AI-native games from AI-augmented games and adjacent boundary artifacts. Using this definition, we screen candidate artifacts and analyze 53 publicly available AI-native games and prototypes. We introduce a dual-axis G/N taxonomy: the G-axis captures player-facing game type, while the N-axis captures the dominant AI mechanic that makes generative AI indispensable to play. The corpus is concentrated around language-forward designs, especially narrative adventure, epistemic interaction, and generative narrative, while categories such as semantic adjudication, multi-agent simulation, generative construction, and relationship/companion play remain less represented. We argue that the central design problem is organizing semantic openness into stable gameplay. AI-native design depends on mechanical invariants: goals, rules, state, feedback, pacing, and player agency that make open-ended AI outputs interpretable and consequential. We conclude with a roadmap for controllable generation, AI-as-mechanic design, multimodal and multi-agent systems, inference economics, evaluation, safety, and regulation.
PedNStream: 歩行者交通管理のためのスケーラブルなネットワーク フロー シミュレーション
大規模な群衆管理には、計算効率が高く、フィードバック ベースの制御と互換性のある歩行者シミュレーションが必要です。ただし、ほとんどのオープンソース ツールは微細なものであるか、ネットワーク規模の閉ループ評価用に設計されていません。このペーパーでは、リンク伝送モデル (LTM) に基づいた巨視的な歩行者ネットワーク負荷のためのオープンソースの Python ネイティブ シミュレーターである PedNStream (Pedestrian Network Flow Simulation) について説明します。このフレームワークは、拡散と活動に起因する変動を捉える確率的リンクダイナミクスを組み込むことで LTM ベースの歩行者モデルを拡張し、動的なユーザーの平衡ルート選択を、不確実で介入主導の設定に適したユーティリティベースの定式化に置き換えます。 PedNStream は、ゲート、フロー分離、ルート ガイダンスなどの介入用のコントローラー インターフェイスが組み込まれたモジュール式フレームワークとして実装されています。私たちはフレームワークを段階的に評価します。合成シナリオでは、キューの形成、スピルバック、輻輳の解消、適応型再ルーティングなどの主要なメカニズムを検証します。実際のネットワーク実験では、大規模な行動と観察された歩行者数との一貫性を評価します。閉ループのケーススタディではコントローラーの統合を実証し、ランタイム分析ではスケーラビリティを定量化します。これらの結果により、PedNStream は大規模な歩行者ネットワークのシミュレーションと制御のための効率的かつ実用的なテストベッドとして確立されます。
原文 (English)
PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management
Large-scale crowd management requires pedestrian simulations that are both computationally efficient and compatible with feedback-based control. However, most open-source tools are either microscopic or not designed for network-scale closed-loop evaluation. This paper presents PedNStream (Pedestrian Network Flow Simulation), an open-source, Python-native simulator for macroscopic pedestrian network loading based on the Link Transmission Model (LTM). The framework extends LTM-based pedestrian models by incorporating stochastic link dynamics that capture diffusion and activity-induced variability, and replaces dynamic user equilibrium route choice with a utility-based formulation suited to uncertain, intervention-driven settings. PedNStream is implemented as a modular framework with built-in controller interfaces for interventions such as gating, flow separation, and route guidance. We evaluate the framework in a staged manner. Synthetic scenarios verify key mechanisms, including queue formation, spillback, congestion dissipation, and adaptive rerouting. Real-network experiments assess large-scale behavior and consistency with observed pedestrian counts. A closed-loop case study demonstrates controller integration, and a runtime analysis quantifies scalability. These results establish PedNStream as an efficient and practical testbed for large-scale pedestrian network simulation and control.
決定論的で自己拡張的な反応分類のための検証可能なルールをエージェント的に生成
コンピューター支援合成計画では、各変換に決定論的で解釈可能なラベルを割り当てる反応ルールの大きなライブラリを使用して、標的分子をアクセス可能な前駆体に分割します。しかし、化学はロングテールであるため、手動エンコーディングが困難であり、既存のツールは新しい化学に適応できない固定ルールセットに依存しています。ここでは、大規模言語モデル (LLM) のマルチエージェント フレームワークが反応を分類し、665,901 件の米国特許反応にわたってルール自体を記述し、コーパスに対してテストする検証ループの下で各ルールを生成する、完全に自動化されたパイプラインを紹介します。人間によるキュレーションを行わずに、標準分類を 68 クラスから 14,073 クラスに拡張します。軽量の指紋分類器を使用して、目に見えない反応の 97.7% を分類し、主要な独自の分類器と一致しながら、化学をより細かく分解し、オンデマンドでトレーニング ディストリビューション外の化学に拡張します。その結果、生きた反応性データベースと、生成モデルを信頼性の高い自己拡張型のシンボリック システムに変えるための一般的なルートが得られます。
原文 (English)
Agentic generation of verifiable rules for deterministic, self-expanding reaction classification
Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign each transformation a deterministic, interpretable label. But chemistry is long-tailed, making manual encoding intractable, and existing tools rely on fixed rulesets that cannot adapt to new chemistries. Here we present a fully automated pipeline in which a multi-agent framework of large language models (LLMs) classifies reactions and writes the rules themselves across 665,901 US patent reactions, generating each rule under a verification loop that tests it against the corpus. It expands a standard taxonomy from 68 to 14,073 classes without human curation. With a lightweight fingerprint classifier, it classifies 97.7\% of unseen reactions, matching a leading proprietary classifier while resolving chemistry more finely and extending on demand to chemistry outside its training distribution. The result is a living reactivity database and a general route to turning generative models into reliable, self-expanding symbolic systems.
Agent4cs: 大規模な階層コードベースでのコード要約のためのマルチエージェント システム
大規模で複雑なコードベース、特に難読化された構造や不完全なドキュメントを持つコードベースを理解することは、依然として大きな課題です。既存のコード要約ソリューションは、単一の言語モデルやクロード コードのようなコーディング アシスタントに依存することが多く、ソース コードをフラット テキストとして扱い、リポジトリ内の豊富な相互依存関係や階層情報が十分に活用されていません。これらの欠点に対処するために、Agent4cs を提案します。Agent4cs は、大規模なコードベースをボトムアップ方式で要約するマルチエージェント フレームワークです。要約エージェントは、堅牢な要約を生成することに重点を置いています。キーワード抽出エージェントは、サブフォルダーから重要な情報を積極的に識別します。そして、品質保証エージェントは、読みやすさ、一貫性、完全性を高めるために出力を繰り返し改良します。 7 つのフロンティア モデルで評価された Agent4cs は、コード セグメントを使用した 2 つの構造化プロンプト ベースラインと比較して、すべてのフォルダー レベルにわたるセマンティックの一貫性を平均 8% 向上させます。さらに、現実世界のデータセットに対する広範な評価により、同じベースラインに対して正規化されたキーワード カバレッジ率が最大 38% 向上することが実証されています。
原文 (English)
Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases
Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge. Existing code summarization solutions often rely on a single language model or coding assistant like Claude Code, and treat source code as flat text, underutilizing the rich interdependencies and hierarchical information within a repository. To address these shortcomings, we propose Agent4cs - a multi-agent framework that summarizes large codebases in a bottom-up fashion, where a summarization agent focuses on producing robust summaries; a keyword-extraction agent proactively identifies critical information from subfolders; and a quality-assurance agent iteratively refines the outputs for readability, coherence, and completeness. Evaluated on 7 frontier models, Agent4cs improves semantic consistency across all folder levels by average 8% compared to two structured prompting baselines with code segments. Furthermore, extensive evaluation on real-world datasets demonstrates up to 38% gains in normalized keyword coverage rate over the same baselines.
Hawk: ハードウェアを意識した知識を活用して高性能 NPU カーネルを生成する
Neural Processing Unit (NPU) 用の高性能カーネルの開発は業界の重大なボトルネックであり、開発者は暗黙のハードウェア制約と厳密なメモリ階層を手動でナビゲートする必要があります。大規模な言語モデルには計り知れない自動化の可能性がありますが、ハードウェア固有の事前条件が根本的に欠如しているため、NPU では壊滅的に失敗します。同様の NPU カーネルから単純にコード スニペットを移植すると、コンパイラは通過する可能性がありますが、根底にあるハードウェア制約に盲目的に違反することにより、ランタイム クラッシュやパフォーマンスの低下が常に引き起こされます。これを克服するために、3 つのコア モジュールを通じてハードウェア認識の知識を活用する、トレーニング不要のフレームワークである Hawk を導入します。(1) 実行時知識合成モジュール。これは、3 部構成の実行可能知識表現を採用して、エラー コンテキストと実行可能セマンティクスを本質的に結合します。 (2) ボトルネック認識知識検索モジュール。クエリを直交構文およびハードウェアに合わせた意味空間に投影する 2D 検索パラダイムを実装します。 (3) エフェクト駆動型の知識抽出モジュール。LLM 駆動のセマンティック アービトレーションを利用して、経験的な実行フィードバックに基づいてエラーを枝刈りし、冗長性を統合することで、継続的に知識を抽出します。実際の NPU ワークロードに関する広範な評価により、Hawk が生成精度を 49.4% から 80.0% に向上させながら、最先端のベースラインと比較して最大 2.2 倍の実行速度向上を達成することが実証されました。
原文 (English)
Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation
Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually navigate implicit hardware constraints and strict memory hierarchies. While large language models offer immense automation potential, they fail catastrophically on NPUs due to a fundamental lack of hardware-specific priors. Naively transplanting code snippets from similar NPU kernels may pass the compiler, but it consistently triggers runtime crashes and performance degradation by blindly violating underlying hardware constraints. To overcome this, we introduce Hawk, a training-free framework that harnesses hardware-aware knowledge through three core modules: (1) Run-Time Knowledge Synthesis Module, which employs a Triple-Part Executable Knowledge Representation to inherently couple the error context with executable semantics; (2) Bottleneck-Aware Knowledge Retrieval Module, which implements a 2D-Retrieval paradigm to project queries into orthogonal syntactic and hardware-aligned semantic spaces; and (3) Effect-Driven Knowledge Distillation Module, which leverages LLM-driven semantic arbitration to continuously distill the knowledge by pruning errors and consolidating redundancies based on the empirical execution feedback. Extensive evaluations on real-world NPU workloads demonstrate that Hawk elevates generation accuracy from 49.4% to 80.0%, while achieving up to a 2.2x execution speedup over state-of-the-art baselines.
Raw-ECG-Replay-Free 継続的 ECG 導入における自律的なソース推論から専門家の保持を分離する
マルチソース ECG 展開では、以前の生の ECG を保持または再生できない場合、モデルに新しいデータ ソースを組み込む必要がある場合があります。事前トレーニングされたバックボーンをフリーズし、各ソースに分離された分類子を割り当てることでパラメータの干渉を防ぐことができますが、ソースのメタデータが利用できない場合でも展開には専門家を選択する必要があります。私たちは、凍結された 1024 次元の ECGFounder 機能に基づいて構築された増分エキスパート バンクである \ours{} を通じて、この違いを研究しています。到着するドメインごとにバランスの取れたソフトマックス線形エキスパートが追加されますが、軽量ルーターは、これまでに観察されたソースからの保持されたトレーニング特徴とドメイン ラベルにのみ適合します。検証によって調整されたマージン ルールは、単一のルーティングされたエキスパートにコミットするのではなく、最も可能性の高い 2 人のエキスパートを融合します。 CPSC、PTB-XL、ジョージア、およびチャップマン-紹興では、ソースを意識した専門家の選択は $0.7915\pm0.0036$ マクロ F1 に達し、一致するオフラインの独立ヘッド参照は $0.7885\pm0.0009$ に達し、強力なソースを認識した専門家の保持をサポートしています。ソース ID がない場合、MLP ルーターは $0.7756\pm0.0027$ に達し、トップ 2 マージン フュージョンは $0.7782\pm0.0022$ に達します。ハード MLP ルーティングに対する上位 2 のゲインは小さく ($+0.0026$)、ペアのブートストラップからの 95\% 信頼区間にはゼロが含まれます。 3 つのドメイン注文全体で、トップ 2 とオラクルの差は依然として 0.0111 ドルから 0.0133 ドルであり、自律的なソース推論が残りの主なボトルネックであることがわかります。生の ECG は再生されませんが、凍結されたトレーニング機能はルーターの更新のために保持されます。したがって、このメソッドはメモリフリーではありません。
原文 (English)
Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment
In multi-source ECG deployment, models may need to incorporate new data sources when earlier raw ECGs cannot be retained or replayed. Freezing a pretrained backbone and assigning each source an isolated classifier prevents parameter interference, but deployment still requires selecting an expert when source metadata are unavailable. We study this distinction through IRFE-ECG, an incremental expert bank built on frozen 1024-dimensional ECGFounder features. Each arriving domain adds a balanced-softmax linear expert, while a lightweight router is fitted only on retained training features and domain labels from sources observed so far. A validation-calibrated margin rule fuses the two most likely experts instead of committing to a single routed expert. On CPSC, PTB-XL, Georgia, and Chapman-Shaoxing, source-aware expert selection reaches $0.7915\pm0.0036$ Macro-F1 and a matched offline independent-head reference reaches $0.7885\pm0.0009$, supporting strong source-aware expert retention. Without source IDs, an MLP router reaches $0.7756\pm0.0027$ and top-2 margin fusion reaches $0.7782\pm0.0022$. The top-2 gain over hard MLP routing is small ($+0.0026$), with a 95\% confidence interval from paired bootstrap that includes zero. Across three domain orders, the top-2-to-oracle gap remains $0.0111$--$0.0133$, identifying autonomous source inference as the main remaining bottleneck. No raw ECGs are replayed, but frozen training features are retained for router updates; the method is therefore not memory-free.Code is available at https://github.com/yufanlu221/IRFE-ECG.
症状ではなくアンプを修復する: エージェント ロールアウトの安定したワールド モデルの修正
エージェントの計画が短いツール チェーンから数千または数万のステップを含む永続的なワークフローに移行すると、個別の予測ではなく、大規模な計画グラフ内で障害が発生します。間違いが発生するたびにグラフ全体を再計画することは、計算上現実的でも望ましいことでもありません。グラフ全体の再生は大量のコンテキスト バジェットを消費し、LLM を多くの無関係な症状にさらし、長いコンテキストの取得を低下させる可能性があります。この論文では、そのようなシステムに欠けているコンポーネント、つまり失敗した計画グラフを適切な場所に修復するワールド モデル コレクターについて研究します。 2 つの補正器ファミリーを比較します。 1 つ目は一般的なエンジニアリング アプローチです。ノードとエッジをスキャンし、疑わしいローカル領域を選択し、LLM に修復を依頼します。私たちは強力なエンジニアリング LLM コレクタを実装しており、特に非常に大規模なコンテキストが与えられた場合に役立つことがわかりました。 2 番目のファミリーは、私たちのアプローチである WM-SAR (World-Model Subgraph Amplification Repair) です。目に見える症状をスキャンするのではなく、サブグラフの増幅から逆算して、エラーを再増幅し続けるノードとエッジを特定し、その原因となるサブグラフのみを LLM に送信します。グラフ シミュレーションと LLM 修復実験全体で、WM-SAR は現実的なトークン バジェットの下でエンジニアリング コレクタを大幅に上回り、コンパクトな領域でほぼグラフ全体の安定化を達成し、LLM にクリーンな修復ターゲットを与えます。
原文 (English)
Repair the Amplifier, Not the Symptom: Stable World-Model Correction for Agent Rollouts
Long-horizon language agents increasingly maintain executable world models in the form of planning graphs, where tool calls, validators, memory updates, recovery branches, and final answers are connected by typed dependencies. When a rollout fails, repairing the most visible error can leave the underlying error-amplification path intact, while replaying the full graph is expensive and difficult for long-context models to use reliably. We study world-model correction: selecting a compact subgraph of a failed planning graph whose repair stabilizes subsequent rollouts. We first instantiate a strong family of engineering correctors, including pointwise error scans, TopK and window selection, local graph expansion, cascade repair, and full-context LLM repair. We then propose WM-SAR, a spectral subgraph repair method that estimates node-edge amplification, greedily grows a connected repair region by marginal residual-spectral relief, and sends only this region to an LLM for root-cause repair. Theoretically, we connect residual spectral radius to rollout error and planning regret, motivating repair as stabilization rather than attribution alone. Across synthetic calling-tree graphs, benchmark-inspired agent topologies, and cross-model LLM repair experiments, WM-SAR achieves stronger long-horizon stabilization and root-cause recovery under compact token budgets, matching much larger repair contexts while exposing the LLM to a cleaner causal subgraph.
LLM エージェントの大規模な安全性テスト: リスク発見から証拠に基づく検証まで
LLM エージェントは外部ツールを通じて自律的なアクションを実行することが増えており、複雑かつ進化する安全リスクにつながっています。しかし、既存の安全性テストは専門家が設計した安全性違反を対象としており、対応する結果はハードコードされたルールによって評価されるため、エージェントの進化に応じてテストを拡張するにはコストがかかります。この目的を達成するために、3 段階の自己強化パイプラインを通じて非決定性エージェントのソフトウェア エンジニアリング テスト原則をインスタンス化する、エンドツーエンドの自動安全性テスト フレームワークである Vera を紹介します。まず、文献に基づいた調査により、新たなリスクを継続的に発見し、安全リスク、攻撃手法、およびツール実行環境の分類に構造化します。第 2 に、分類次元全体にわたる組み合わせ構成により、実行可能な安全ケースが生成され、それぞれが具体的な安全目標、プログラムで構築された初期状態、および観察可能なアーティファクトに基づいた決定論的検証述語を指定します。第三に、適応的実行は、分離されたサンドボックスで異種エージェントを実行します。制御エージェントは実行時の観察に基づいてマルチターンの対話を制御しますが、証拠に基づいた検証者は、モデルの自己報告ではなく環境の状態とツール呼び出しの証拠から結果を判断します。 4 つの運用エージェント フレームワーク (OpenClaw、Hermes、Codex、Claude Code) で Vera を評価したところ、重大な安全上の弱点が明らかになり、マルチチャネル攻撃下での平均攻撃成功率は 93.9% に達しました。また、3 つの実行設定にわたる 124 のリスク カテゴリにわたる 1,600 の実行可能な安全性ケースで構成される Vera-Bench もリリースします。これらの結果は、急速に進化する大規模なエージェント システムの厳密で保守可能な安全性評価には、モジュール式の実行可能なテスト インフラストラクチャが不可欠であることを示しています。コードは https://github.com/Yunhao-Feng/Vera で公開されています。
原文 (English)
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and the corresponding outcomes are evaluated by hard-coded rules, making them costly to extend as agents evolve. To this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage, self-reinforcing pipeline. First, a literature-driven exploration continuously discovers and structures emerging risks into taxonomies of safety risks, attack methods, and tool execution environments. Second, combinatorial composition across taxonomy dimensions produces executable safety cases, each specifying a concrete safety goal, a programmatically constructed initial state, and a deterministic verification predicate grounded in observable artifacts. Third, adaptive execution runs heterogeneous agents in isolated sandboxes where a control agent steers multi-turn interaction based on runtime observations, while evidence-grounded verifiers judge outcomes from environment state and tool-call evidence rather than model self-report. We evaluate Vera on four production agent frameworks (OpenClaw, Hermes, Codex, Claude Code), revealing substantial safety weaknesses, with average attack success rates reaching 93.9\% under multi-channel attacks; we also release Vera-Bench, comprising 1600 executable safety cases spanning 124 risk categories across three execution settings. These results indicate that modular, executable testing infrastructure is essential for rigorous and maintainable safety evaluation of rapidly evolving agentic systems at scale. The code is publicly available at https://github.com/Yunhao-Feng/Vera.
ContextSniper: リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ
大規模な言語モデル エージェントは実際のリポジトリの問題を修復できますが、ファイル全体の読み取り、広範な検索、および有用な証拠が無関係なコードやログと混在する長いターミナル出力に多額のコンテキスト バジェットを費やすことがよくあります。このペーパーでは、リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ層である ContextSniper について説明します。 AntTrail の広範なエージェント メモリ エンジンのコーディングに特化したものとして、ContextSniper は正確な証拠選択のための Sniper 機能を実装しています。候補コードとランタイム証拠を取得し、ハイブリッド取得信号でランク付けし、意図を認識したコンテキスト ゲートを通じて長い出力をフィルタリングし、プロンプトの外で回復可能なソース コンテキストを保持しながらコンパクトな証拠パケットを返します。 OpenClaw と Claude Code を備えた SWE-bench Lite 上で、ホスト エージェント条件ごとに 50 タスクの実行を使用して ContextSniper を評価しました。 ContextSniper は、OpenClaw の場合、トークンの総使用量を 51.5%、ログに記録されたコストを 36.4% 削減し、Claude Code の場合、トークンの総使用量を 38.9%、推定コストを 27.3% 削減します。提出された解決率は、OpenClaw の場合は 26.0% から 24.0% に、Claude Code の場合は 32.0% から 30.0% にわずかに減少しました。 ContextSniper のパイロット テスト スクリプトは、https://github.com/Calluking/ContextSniper でオープンソース化されています。
原文 (English)
ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair
Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long terminal outputs where useful evidence is mixed with irrelevant code and logs. This paper presents ContextSniper, AntTrail's code-repair module for precision evidence selection in repository-level program repair, part of AntTrail's broader agent-memory engine. AntTrail is available at https://gitcode.com/datagallery/AntTrail. ContextSniper indexes code and action memory as three abstract levels, retrieves candidates with a hybrid ranker, filters long tool output through an intention-aware context gate, and returns compact evidence packets while keeping full source recoverable on demand. In a matched 50-task-per-condition comparison on SWE-bench Lite (same tasks, baseline vs.\ ContextSniper), ContextSniper reduces total token use by 51.5% and logged cost by 36.4% for OpenClaw, and by 38.9% and 27.3% for Claude Code, with submitted-resolution rates essentially unchanged in both host-agent settings. In a separate five-task comparison, ContextSniper beats existing memory- and RAG-style integrations on token efficiency. These results suggest ContextSniper can substantially cut token and cost overhead for repository-level repair agents without a measurable loss in repair quality. The evaluation harness for this study is available at https://gitcode.com/lukchiwang/ContextSniper.
PACE: エージェントの能力評価の代理
SWE-Bench や GAIA などのベンチマークで LLM エージェントを評価するには、費用と時間がかかり、複雑なインフラストラクチャが必要になる場合があります。 1 回の評価に数千ドルの費用がかかり、完了までに数日かかる場合があります。対照的に、個々の機能 (推論、コード生成など) をテストする非エージェント LLM ベンチマークは、高速かつ低コストで実行できます。このペーパーでは、高価なエージェント ベンチマークでのパフォーマンスが、慎重に選択されたアトミック評価インスタンスの少数のサブセットでのパフォーマンスによって正確に予測できるかどうかを調査します。 PACE は、既存の非エージェント評価からインスタンスを選択することでプロキシ ベンチマークを構築するフレームワークで、その集計スコアがエージェント ベンチマークでのモデル パフォーマンスを最も確実に予測するものを紹介します。アトミック機能にわたる候補インスタンスのプールを考慮すると、PACE は、ソース インスタンスのコンパクトなサブセットのモデルのスコアをターゲット エージェント ベンチマークのスコアにマッピングする回帰を当てはめます。サブセット自体は、2 つの相補的なインスタンス選択戦略、ターゲット関連性のローカル選択とグローバルに情報を提供するグローバル選択を組み合わせることによってキュレーションされます。このペーパーでは、PACE を 4 つのターゲット エージェント ベンチマークに適用し、このペーパーで評価する具体的なプロキシ ベンチマークである PACE-Bench を生成します。 14 のモデル、4 つのエージェント ベンチマーク、および 19 の非エージェント ベンチマークにわたる実験では、PACE-Bench が、リーブ ワンアウト相互検証 (LOOCV) 平均絶対誤差 (MAE) が 4% 未満、スピアマン相関が 0.80 以上、ペアワイズ モデル ランク付け精度が約 85% で、すべてエージェント評価コスト全体の 1% 未満でエージェント スコアを予測することが示されています。選択したプロキシ インスタンスをさらに分析し、各エージェント ベンチマークが独自に要求するスキルを明らかにします。 PACE を使用すると、担当者は、完全なエージェント評価のオーバーヘッドを発生させることなく、モデルの開発、選択、ルーティング中にエージェントのパフォーマンスの信頼できる推定値を取得できます。
原文 (English)
PACE: A Proxy for Agentic Capability Evaluation
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation) are fast and cheap to run. In this paper, we investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benchmarks by selecting instances from existing non-agentic evaluations whose aggregate scores most reliably predict model performances on agentic benchmarks. Given a pool of candidate instances spanning atomic capabilities, PACE fits a regression that maps a model's scores on a compact subset of source instances to its score on the target agentic benchmark. The subset itself is curated by combining two complementary instance-selection strategies, target-relevance local selection and globally informative global selection. We apply PACE to the 4 target agentic benchmarks in this paper, which yields PACE-Bench, the concrete proxy benchmark that we evaluate in the paper. Experiments across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks show that PACE-Bench predicts agentic scores with leave-one-out cross-validation (LOOCV) mean absolute error (MAE) under 4%, Spearman correlation above 0.80, and pairwise model-ranking accuracy around 85%, all at much less than 1% of the full agentic evaluation cost. We further analyze the selected proxy instances, revealing which skills each agentic benchmark uniquely demands. PACE enables practitioners to obtain reliable estimates of agentic performance during model development, selection, and routing, without the overhead of full agent evaluation.
ContextNest: 自律型 AI エージェントの検証可能なコンテキスト ガバナンス
自律型 AI エージェントは外部ナレッジ ストアへの依存度を高めていますが、ほとんどの検索パイプラインは、出所、バージョン ID、完全性、トレーサビリティ、またはポイントインタイムの再構成の永続的な保証なしで関連性を提供します。私たちはこれをコンテキスト ガバナンスとして形式化し、管理された AI 消費型ナレッジ ボールトのオープン仕様およびリファレンス実装である ContextNext を提示します。 ContextNext は、検索拡張生成 (RAG) を置き換えるものではありません。これは、検索の下にガバナンス層を提供し、検索システムが動作する前に、どの成果物が承認され、最新のもので、帰属可能であり、整合性が検証されているかを判断します。この仕様では、型指定された Markdown ドキュメントとメタデータ、決定論的な集合代数セレクター、contextnest:// URI 参照、SHA-256 ハッシュチェーンのバージョン履歴、グラフレベルのチェックポイント、モデル コンテキスト プロトコル (MCP) を介したライブ データのソース ノード、およびエージェント コンテキスト消費の監査トレースを組み合わせています。これらのメカニズムにより、組織はどのナレッジ バージョンがエージェント出力に情報を提供したか、またそれらのバージョンが使用時に AI に適格かどうかを再構築できます。我々は、2 つの対照実験からの最初の経験的結果を報告します。ガバナンス対取得の失敗モードを分離する古いバージョンの攻撃では、管理された選択により BM25 のスパース検索が厳密にパレート支配され、入力トークンの約 3 分の 1 のコストでより高い応答品質の合格率 (97% 対 93 ~ 90%) が得られます。 1,060 の文書コーパスに対する検索決定論実験では、決定論的セレクターと BM25 は同一のクエリを繰り返しても安定した文書セットを返します (Jaccard 1.0)。一方、密 + HNSW ベースラインはクエリの 80% で非決定的です (平均 Jaccard 0.611、最悪の場合 0.210)。これらの結果は、コンテキスト ガバナンスが障害モードに対処すること、取得品質だけでは解決できないことを示唆しています。コア エンジン、CLI、MCP サーバーをオープン ライセンスでリリースします。
原文 (English)
ContextNest: Verifiable Context Governance for Autonomous AI Agent
Autonomous AI agents increasingly depend on external knowledge stores, yet most retrieval pipelines provide relevance without durable guarantees of provenance, version identity, integrity, traceability, or point-in-time reconstruction. We formalize this as context governance and present ContextNest, an open specification and reference implementation for governed AI-consumable knowledge vaults. ContextNest does not replace Retrieval-Augmented Generation (RAG); it supplies the governance layer beneath retrieval, determining which artifacts are approved, current, attributable, and integrity-verified before retrieval systems operate over them. The specification combines typed Markdown documents with metadata, deterministic set-algebraic selectors, contextnest:// URI references, SHA-256 hash-chained version histories, graph-level checkpoints, source nodes for live data through the Model Context Protocol (MCP), and audit traces of agent context consumption. These mechanisms let organizations reconstruct which knowledge versions informed an agent output and whether those versions were AI-eligible when consumed. We report first empirical results from two controlled experiments. In a stale-version attack isolating the governance-versus-retrieval failure mode, governed selection strictly Pareto-dominates BM25 sparse retrieval, with higher answer-quality pass rate (97% versus 93-90%) at about one-third the input-token cost. In a retrieval-determinism experiment over a 1,060-document corpus, deterministic selectors and BM25 return stable document sets across repeated identical queries (Jaccard 1.0), while a dense+HNSW baseline is non-deterministic on 80% of queries (mean Jaccard 0.611, worst case 0.210). These results suggest that context governance addresses failure modes retrieval quality alone is not designed to resolve. We release a core engine, CLI, and MCP server under open licenses.
一貫性のない優先順位付けされたナレッジ ベースのクエリと修復: 複雑さの分析と抽象的な議論によるリンク
このペーパーでは、オントロジー、ファクトのセット、および矛盾するファクト間の優先関係で構成される優先順位付けされたナレッジ ベース (KB) における不整合処理の問題について検討します。データベース設定では、密接に関連するシナリオが研究され、優先順位が付けられた一貫性のないデータベースの最適な修復に関する 3 つの異なる概念 (グローバル、パレート、完了) が定義されました。グローバル最適修復、パレート最適修復、および完了最適修復の概念を私たちの設定に移した後、核となる推論タスクのデータ複雑性、つまり最適修復に基づく不整合耐性セマンティクスの下でのクエリ含意、固有の最適修復の存在、およびすべての最適修復の列挙を研究します。私たちの結果は、一般的な DL-Lite 方言で定式化されたオントロジーのこれらのタスクのデータ複雑さのほぼ完全な全体像を提供します。私たちの研究の 2 番目の貢献は、最適な修復と (セットベースの) 議論フレームワークの拡張のさまざまな概念との関係を明確にすることです。私たちの結果の中で、パレート最適修復が安定した拡張機能 (および多くの場合、優先拡張機能) に正確に対応していることを示し、根拠のある拡張機能からインスピレーションを得て有利な計算特性を享受できる、優先順位付き KB の新しいセマンティクスを提案します。私たちの研究では、好みに基づく議論の枠組みに関して独立して興味深い結果もいくつか得られています。
原文 (English)
Querying and Repairing Inconsistent Prioritized Knowledge Bases: Complexity Analysis and Links with Abstract Argumentation
In this paper, we explore the issue of inconsistency handling over prioritized knowledge bases (KBs), which consist of an ontology, a set of facts, and a priority relation between conflicting facts. In the database setting, a closely related scenario has been studied and led to the definition of three different notions of optimal repairs (global, Pareto, and completion) of a prioritized inconsistent database. After transferring the notions of globally-, Pareto- and completion-optimal repairs to our setting, we study the data complexity of the core reasoning tasks: query entailment under inconsistency-tolerant semantics based upon optimal repairs, existence of a unique optimal repair, and enumeration of all optimal repairs. Our results provide a nearly complete picture of the data complexity of these tasks for ontologies formulated in common DL-Lite dialects. The second contribution of our work is to clarify the relationship between optimal repairs and different notions of extensions for (set-based) argumentation frameworks. Among our results, we show that Pareto-optimal repairs correspond precisely to stable extensions (and often also to preferred extensions), and we propose a novel semantics for prioritized KBs which is inspired by grounded extensions and enjoys favourable computational properties. Our study also yields some results of independent interest concerning preference-based argumentation frameworks.
Double Fuzzy Probabilistic Interval Linguistic Term Set and a Dynamic Fuzzy Decision Making Model based on Markov Process with tts Application in Multiple Criteria Group Decision Making
The probabilistic linguistic term has been proposed to deal with probability distributions in provided linguistic evaluations. However, bec…
Domain Knowledge-Informed Self-Supervised Representations for Workout Form Assessment
Maintaining proper form while exercising is important for preventing injuries and maximizing muscle mass gains. Detecting errors in workout…
A Transformer-Based Contrastive Learning Approach for Few-Shot Sign Language Recognition
Sign language recognition from monocular video or 2D pose sequences is challenging, both because 3D information must be inferred from 2D ob…
Restricted Bernoulli Matrix Factorization: Balancing the trade-off between prediction accuracy and coverage in classification based collaborative filtering
Reliability measures associated with the prediction of the machine learning models are critical to strengthening user confidence in artific…
Unveiling the Unborn: Advancing Fetal Health Classification through Machine Learning
Fetal health classification is a critical task in obstetrics, enabling early identification and management of potential health problems. Ho…
Learning to Visually Connect Actions and their Effects
We introduce the novel concept of visually Connecting Actions and Their Effects (CATE) in video understanding. CATE can have applications i…
TERC: A Transfer Entropy Redundancy Criterion for State Variable Selection in Reinforcement Learning
Identifying the most suitable variables to represent the state is a fundamental challenge in Reinforcement Learning (RL). These variables m…
Graph Unitary Message Passing
Unitarity is a useful principle for stabilizing deep neural networks, but in graph neural networks (GNNs) instability is induced not only b…
CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes
Causal video question answering (QA) has garnered increasing interest, yet existing datasets often lack depth in causal reasoning. To addre…
MambaCapsule: Towards Transparent Cardiac Disease Diagnosis with Electrocardiography Using Mamba Capsule Network
Cardiac arrhythmia, a condition characterized by irregular heartbeats, often serves as an early indication of various heart ailments. With…
Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies
LLMs are increasingly used world-wide from daily tasks to agentic systems and data analytics, requiring significant GPU resources. While LL…
Stroke Prediction using Clinical and Social Features in Machine Learning
Every year in the United States, 800,000 individuals suffer a stroke - one person every 40 seconds, with a death occurring every four minut…
Robust Counterfactual Explanations under Model Multiplicity Using Multi-Objective Optimization
In recent years, explainability in machine learning has gained importance. In this context, counterfactual explanation (CE), which is an ex…
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the com…
Evaluating LLM-Based Regression Test Generation
Large Language Models (LLMs) have shown tremendous promise in automated software engineering. In this paper, we investigate LLMs for just-i…
Machine Unlearning via Information Theoretic Regularization
How can we effectively remove or ``unlearn'' undesirable information, such as specific features or the influence of individual data points,…
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
While guided decoding, especially value-guided methods, has emerged as a cost-effective alternative for controlling language model outputs…
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targ…
Explainable Bayesian deep learning through input-skip Latent Binary Bayesian Neural Networks
Modeling natural phenomena with artificial neural networks (ANNs) often provides highly accurate predictions. However, ANNs often suffer fr…
Empirical Computation: Prompting versus Programming
Large Language Model (LLM) agents can solve *any* computational problem *without* an algorithm in a runtime *independent* of the computatio…
Measuring the Robustness of Audio Deepfake Detection under Real-World Corruption
Deepfakes have emerged as a widespread and rapidly escalating concern in generative AI, spanning images, audio, and videos. Among these, au…
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low…
AgentDynEx: Nudging the Mechanics and Dynamics of Multi-Agent Simulations
Multi-agent large language model simulations have the potential to model complex human behaviors and interactions. If the mechanics are set…
GenShin: Guiding Rational Liposome Design by Ranking Liposomal Protein Corona through a Docking-Pose-Free GNN
Rational design of lipid nanoparticles (LNPs) for tissue-specific delivery critically depends on predicting the composition of the protein…
MOSAIC: Skill-Centric Manipulation Planning with Physics Simulation
Planning long-horizon manipulation motions using a set of predefined skills is a central challenge in robotics; solving it efficiently coul…
Exploring Context-aware and LLM-driven Locomotion for Immersive Virtual Reality
Locomotion plays a crucial role in shaping the user experience within virtual reality environments. In particular, hands-free locomotion of…
kAgent: An execution-guided crash resolution agent for the Linux kernel
Fuzzing frameworks like syzkaller have uncovered thousands of Linux kernel crashes, many of which are critical and security-sensitive. Howe…
Towards Understanding Deep Learning Model in Image Recognition via Coverage Test
Deep neural networks (DNNs) play a crucial role in the field of artificial intelligence, and their security-related testing has been a prom…
Quick ViTs: Speeding up Vision Transformers through Equivariance
Natural images exhibit strong geometric regularities: local structures, such as edges, corners, and textures, appear in many orientations a…
PDFBench: A Benchmark for De novo Protein Design from Function
Function-guided protein design is a crucial task with significant applications in drug discovery and enzyme engineering. However, the field…
Seven Security Challenges in Cross-domain Multi-agent LLM Systems
Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint di…
Boosting Automatic Exercise Evaluation Through Musculoskeletal Simulation-Based IMU Data Augmentation
Automated evaluation of movement quality can enhance physiotherapeutic treatment and sports training by providing objective, real-time feed…
Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
Natural Language Processing (NLP) has transformed various fields beyond linguistics by applying techniques originally developed for human l…
FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection
The rapid proliferation of surveillance cameras has increased the demand for automated violence detection. While CNNs and Transformers have…
Ensemble Elastic DQN: A Step Dependent Ensemble Approach for Reducing Overestimation in Deep Value-Based Reinforcement Learning
Deep Q-Networks (DQN) can suffer from overestimation bias because bootstrapped targets use a maximisation operation over noisy value estima…
InverseScope: Scalable Activation Inversion for Interpreting Large Language Models
Understanding the internal representations of large language models (LLMs) is a central challenge in interpretability research. Existing fe…
ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models
Diffusion models excel at image generation. Recent studies have shown that these models not only generate high-quality images but also enco…
Position: Use Sparse Autoencoders to Discover Unknowns
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their u…
Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues
Large language models (LLMs) may support counseling training, yet evidence from Japanese-language interactions and automated quality rating…
Interaction Techniques that Encourage Longer Prompts Can Improve Psychological Ownership when Writing with AI
Writing longer prompts for an AI assistant to generate a story increases psychological ownership, a user's feeling that the writing belongs…
Context Tuning for In-Context Optimization
We introduce Context Tuning, a simple and effective method to significantly enhance few-shot adaptation of large language models (LLMs) wit…
Last Layer Hamiltonian Monte Carlo
We explore the use of Hamiltonian Monte Carlo (HMC) sampling as a probabilistic last layer approach for deep neural networks (DNNs). While…
Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor
Empowered by large language models (LLMs), intelligent agents have become a popular paradigm for interacting with open environments to faci…
Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
Multimodal large-scale models have significantly advanced the development of web agents, enabling perception and interaction with digital e…
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
Tokenization is the first -- and often least scrutinized -- step of most NLP pipelines. Standard algorithms for learning tokenizers rely on…
Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images. Despite be…
Algorithmic Shortlisting in Participatory Budgeting
Participatory budgeting is a democratic innovation that allows citizens to propose and vote on public investment projects. To help organize…
VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
Vision Language Models (VLMs) are increasingly being used in a broad range of applications, bringing their security and behavioral control…
Rational Inverse Reasoning: Few-Shot Imitation by Inferring Intent through Planning
Humans can learn a new manipulation task from one or two demonstrations and then perform it in a new room, with new objects, under new cons…
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety…
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
As numerous instruction-tuning datasets continue to emerge, dynamically balancing and optimizing their mixtures has become a critical chall…
EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation
MultiModal Recommendation (MMR) systems have emerged as a promising solution for improving recommendation quality by leveraging rich item-s…
LLM-Assisted Semantic Alignment and Integration in Collaborative Model-Based Systems Engineering Using SysML v2
Cross-organizational collaboration in Model-Based Systems Engineering (MBSE) faces many challenges in achieving semantic alignment across i…
EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
Code Language Models (CodeLLMs) learn token importance from data correlations, whereas human developers attend selectively to semantically…
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
Existing Causal-Why Video Question Answering (VideoQA) models often struggle with higher-order reasoning, relying on opaque, monolithic pip…
A Systematic Survey on Large Language Models for Evolutionary Optimization: From Modeling to Solving
Large language models (LLMs) are increasingly integrated with evolutionary computation to support optimization tasks. This survey primarily…
Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
For groups of autonomous agents to achieve a particular goal, they must engage in coordination and long-horizon reasoning. Rather than rely…
ICR-RL: Deep Reinforcement Learning via In-Context Regression
Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling th…
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned…
Interpretable Nanoporous Materials Design with Symmetry-Aware Networks
Nanoporous materials hold promise for diverse sustainable applications, yet their vast chemical space poses challenges for efficient design…
Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator
We present the first language-model-driven agentic artificial intelligence (AI) system to autonomously execute multi-stage physics experime…
Polarity Detection of Sustainable Development Goals in News Text
The United Nations' Sustainable Development Goals (SDGs) provide a globally recognised framework for addressing major societal, environment…
Adaptive Margin RLHF via Preference over Preferences
Margin-based optimization is fundamental to improving generalization and robustness in classification tasks. In the context of reward model…
OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling
Pipeline parallelism is widely used to train large language models (LLMs). However, increasing heterogeneity in model architectures exacerb…
MAD-PINN: A Decentralized Physics-Informed Machine Learning Framework for Safe and Optimal Multi-Agent Control
Co-optimizing safety and performance in large-scale multi-agent systems remains a fundamental challenge. Existing approaches based on multi…
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
As Vision Language Models (VLMs) are deployed across safety-critical applications, understanding and controlling their behavioral patterns…
Quadratic Programming Approach for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games
There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player zero-sum imperfect-infor…
Panorama: Fast-Track Nearest Neighbors
Approximate Nearest-Neighbor Search (ANNS) pipelines for high-dimensional neural embeddings spend the bulk of their query time in candidate…
The Three Regimes of Offline-to-Online Reinforcement Learning
Offline-to-online reinforcement learning (RL) has emerged as a practical paradigm that leverages offline datasets for pretraining and onlin…
Verifier-free Test-Time Sampling for Vision-Language-Action Models
Vision-Language-Action models (VLAs) have demonstrated remarkable performance in robot control. However, they remain fundamentally limited…
TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance
Query-product relevance prediction is fundamental to e-commerce search and has become even more critical in the era of AI-powered shopping,…
SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
Interest in forestry automation is growing alongside rapid advances in deep learning. In particular, tree detection and taxonomic classific…
Conditional Clifford-Steerable CNNs for PDE Modeling
We introduce Conditional Clifford-Steerable CNNs (C-CSCNNs), a unified framework that incorporates equivariance to arbitrary pseudo-Euclide…
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
Large Language Models (LLMs) are increasingly used as interfaces to information, code, and real-world services, making prompt-level securit…
UNDREAM: Bridging Differentiable Rendering and Photorealistic Simulation for End-to-end Adversarial Attacks
Deep learning models deployed in safety critical applications like autonomous driving use simulations to test their robustness against adve…
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We s…
Neurosymbolic Characterization for Reliable Access Control Policy Analysis
Access control policies are reliability-critical configuration artifacts in cloud systems, yet administrators frequently struggle to verify…
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
Mixture-of-Experts (MoE) models have shown strong potential in scaling language models efficiently by activating only a small subset of exp…
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly importa…
Resilient by Design -- Active Inference for Distributed Continuum Intelligence
Failures are the norm in highly complex and heterogeneous devices spanning the distributed computing continuum (DCC), from resource-constra…
SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial…
When is a System Discoverable from Data? Discovery Requires Chaos
The deep learning revolution has spurred a rise in advances of using AI in sciences. Within physical sciences the main focus has been on di…
Walrus: A Cross-Domain Foundation Model for Continuum Dynamics
Foundation models have transformed machine learning for language and vision, but achieving comparable impact in physical simulation remains…
Score-Regularized Joint Sampling with Importance Weights for Flow Matching
Flow matching models effectively represent complex distributions, yet estimating expectations of functions of their outputs remains challen…
KAN vs LSTM Performance in Time Series Forecasting
This study presents a controlled comparison of baseline Kolmogorov-Arnold Networks (KAN), implemented via PyKAN, and Long Short-Term Memory…
Exploring the Rashomon Set for Concept-Based Models
In many machine learning problems, there may exist multiple models that achieve nearly identical predictive performance while relying on fu…
InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
Recent approaches in controllable novel view video generation often rely on fine-tuning pre-trained Video Diffusion Models (VDMs). This dom…
Developing an LLM-Based Feedback System Grounded in Evidence-Centered Design to Support Physics Problem Solving
Generative AI offers new opportunities for individualized and adaptive learning, e.g., through large language model (LLM)-based feedback sy…
AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression
Cloud-native systems have made operational work both more powerful and harder to automate: incidents unfold across microservices, logs and…
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the mul…
EvoXplain: When Machine Learning Models Agree on Predictions but Disagree on Why -- Measuring Mechanistic Multiplicity Across Training Runs
Machine learning models are primarily judged by predictive performance, especially in applied genomics, where explanations are read as biol…
JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
We consider problems of parameter estimation where design variables can be actively optimized to maximize information gain. To this end, we…
Resource-constrained Project Scheduling with Time-of-Use Energy Tariffs and Machine States: A Logic-based Benders Decomposition Approach
In this paper, we investigate the Resource-Constrained Project Scheduling Problem (RCPSP) with Time-of-Use (TOU) energy tariffs and machine…
Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework
Semantic Multi-Object Tracking (SMOT) is evolving from purely geometric localization toward comprehensive video understanding. However, exi…
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
The emergence of fine-grained numerical formats like NVFP4 presents new opportunities for efficient Large Language Model (LLM) inference. H…
Motion Attribution for Video Generation
Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTI…
One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization
Conventional audio equalization is a static process that requires manual and cumbersome adjustments to adapt to changing listening contexts…
Predicting Biased Human Decision-Making with Large Language Models in Conversational Settings
We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their prediction…
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning
Existing reinforcement learning methods for LLM reasoning implicitly assume that the policy generating training trajectories should coincid…
Tracing 3D Anatomy in 2D Strokes: A Multi-Stage Projection Driven Approach to Cervical Spine Fracture Identification
Cervical spine fractures require rapid and accurate diagnosis, yet automatic CT interpretation remains challenging as subtle injuries must…
No Reliable Evidence of Self-Reported Sentience in Small Large Language Models
Whether language models possess sentience has no empirical answer. But whether they believe themselves to be sentient can, in principle, be…
The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding
Federal research funding shapes the direction, diversity, and impact of the US scientific enterprise. Large language models (LLMs) are rapi…
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
Accurate evaluation of user satisfaction is critical for iterative development of conversational AI. However, for open-ended assistants, tr…
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging…
A Random Matrix Theory Perspective on the Consistency of Diffusion Models
Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same no…
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
Generating long-form content, such as minute-long videos and extended texts, is increasingly important for modern generative models. Block…
Multi-Way Representation Alignment
The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces.…
SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
We propose SHINE (Scalable Hyper In-context NEtwork), a scalable hypernetwork that can map diverse meaningful contexts into high-quality Lo…
Implementing Grassroots Logic Programs with Multiagent Transition Systems and AI (Full Version)
Grassroots Logic Programs (GLP) is a concurrent logic programming language in which logic variables are partitioned into paired readers and…
Endogenous Resistance to Activation Steering in Language Models
Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.g., ``wait…
Deriving Neural Scaling Laws from the statistics of natural language
Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no exi…
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
Process rewards have been widely used in deep reinforcement learning to improve training efficiency, reduce variance, and prevent reward ha…
Learning to Discover Iterative Spectral Algorithms
We introduce AutoSpec, a neural network framework for discovering iterative spectral algorithms for large-scale numerical linear algebra an…
pFedNavi: Structure-Aware Personalized Federated Vision-Language Navigation for Embodied AI
Vision-Language Navigation VLN requires large-scale trajectory instruction data from private indoor environments, raising significant priva…
NextCrystal: a Symmetry-Driven Generative Framework for Crystal Structure Prediction
Crystal structure prediction (CSP), which aims to predict the 3D atomic arrangement of a crystal from its composition, is central to materi…
SpectralGCD: Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
Generalized Category Discovery (GCD) aims to identify novel categories in unlabeled data while leveraging a small labeled subset of known c…
MARS: 報酬モデリングのためのマージンとセマンティックを意識したデータ拡張
報酬モデリングは、RLHF、RLAIF、PPO ベースのポリシー最適化などの調整パイプラインの中心ですが、その信頼性は、大規模に収集するには費用がかかる、限られた異種の人間の嗜好データによって制約されます。合成拡張は選好の監視を拡張できますが、既存の方法では、報酬モデルが不確実であったり、誤ったランキングが発生しやすい例をターゲットにすることなく、均一にまたは表現レベルで拡張することがよくあります。この論文では、マージンの低い嗜好ペアを優先し、選択された応答と拒否された応答の間のコントラストを強化するための洗練のための第 2 層としてセマンティック距離を使用する適応拡張フレームワークである MARS (Margin and Semantic-Aware Data Augmentation for Reward Modeling) を紹介します。 MARS は、複数の優先データセット、報酬モデルのバックボーン、下流のアライメント設定、および RewardBench や AlpacaEval を含むベンチマークにわたって、報酬モデルの品質とアライメントのパフォーマンスの両方を既存のベースラインよりも向上させます。私たちの結果は、報酬モデルの拡張は、モデルのマージンと意味構造の両方によって導かれる場合に最も効果的であることを示しています。
原文 (English)
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
Reward modeling is central to alignment pipelines such as RLHF, RLAIF, and PPO-based policy optimization, yet its reliability is constrained by limited and heterogeneous human preference data that are expensive to collect at scale. While synthetic augmentation can expand preference supervision, existing methods often augment uniformly or at the representation level, without targeting examples where the reward model is uncertain or prone to mis-ranking. In this paper, we introduce MARS (Margin and Semantic-Aware Data Augmentation for Reward Modeling), an adaptive augmentation framework that prioritizes low-margin preference pairs and uses semantic distance as a second layer for refinement to enhance the contrast between the chosen and rejected responses. Across multiple preference datasets, reward-model backbones, downstream alignment settings, and benchmarks including RewardBench and AlpacaEval, MARS improves both reward-model quality and alignment performance over existing baselines. Our results show that reward-model augmentation is most effective when guided by both model margins and semantic structure.
Inelastic Constitutive Kolmogorov-Arnold Networks: A generalized framework for automated discovery of interpretable inelastic material models
A key problem of solid mechanics is the identification of the constitutive law of a material, that is, the relation between strain history…
Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
Reinforcement Learning from Human Feedback (RLHF) or Verifiable Rewards (RLVR) are two key steps in the post-training of modern Language Mo…
SFL-Net: Source-Factorized Latent Representation Learning for Multi-Contrast MRI to Tau-PET Synthesis
Tau positron emission tomography supports Alzheimer's disease staging but is difficult to scale because of tracer, scanner, and radiation c…
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving
Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distri…
Causal Mechanism Reduction: Mechanism Replacement for Neural Network Pruning and Abstraction
Which internal mechanisms of a neural network can be replaced while preserving the computation it performs? Structured pruning asks for sma…
OSF: On Pre-training and Scaling of Sleep Foundation Models
Polysomnography (PSG) provides the gold standard for sleep assessment but suffers from substantial heterogeneity across recording devices a…
Efficient Flow Matching for Sparse-View CT Reconstruction
Generative models, particularly Diffusion Models (DM), have shown strong potential for Computed Tomography (CT) reconstruction serving as e…
GIPO: Gaussian Importance Sampling Policy Optimization
Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitatio…
LAW & ORDER: Adaptive Spatial Weighting for Medical Diffusion and Segmentation
Medical image analysis depends on accurate segmentation and controllable synthesis, but both tasks face severe spatial imbalance: lesions o…
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctne…
BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and seman…
Multi-Agent Reinforcement Learning for V2X Resource Allocation: Disentangling MARL Challenges Through Benchmarking
Radio resource allocation (RRA) is a critical function in cellular vehicle-to-everything (C-V2X) networks, where vehicles must share limite…
A Hybrid Quantum Circuit Born Machine Framework for Financial Volatility Forecasting: Quantum-Assisted Training and Classical Inference
Accurate financial volatility forecasting is crucial but challenged by the non-linear, highly correlated nature of market data. Recently, q…
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
We consider the task of end-to-end spreadsheet generation, where language models produce spreadsheet artifacts to satisfy users' explicit a…
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation cap…
Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation
Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise. Thus, like playing the slo…
How Transformers Reject Wrong Answers: Rotational Dynamics of Factual Constraint Processing
When a decoder-only transformer is forced to process matched correct and incorrect single-token continuations of a factual query, the two p…
Human-like Object Grouping in Self-supervised Vision Transformers
Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent objec…
Trust-Region Noise Search for Black-Box Alignment of Diffusion and Flow Models
Optimizing the noise samples of diffusion and flow models is an increasingly popular approach to align these models to target rewards at in…
CABTO: Context-Aware Behavior Tree Grounding for Robot Manipulation
Behavior Trees (BTs) offer a powerful paradigm for designing modular and reactive robot controllers. BT planning, an emerging field, provid…
Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation
LLM agents are becoming increasingly important in the security domain, but leading systems are often closed-source, cloud-based, hard to re…
A Two-stage Transformer Framework for Temporal Localization of Distracted Driver Behaviors
The identification of hazardous driving behaviors from in-cabin video streams is essential for enhancing road safety and supporting the det…
From Arithmetic to Logic: The Resilience of Logic and Lookup-Based Neural Networks Under Parameter Bit-Flips
The deployment of deep neural networks (DNNs) in safety-critical edge environments necessitates robustness against hardware-induced bit-fli…
Estimating Individual Tree Height and Species from UAV Imagery
Accurate estimation of forest biomass, a major carbon sink, relies heavily on tree-level traits such as height and species. Unoccupied Aeri…
MCLMR: A Model-Agnostic Causal Learning Framework for Multi-Behavior Recommendation
Multi-Behavior Recommendation (MBR) leverages multiple user interaction types (e.g., views, clicks, purchases) to enrich preference modelin…
SutureFormer: Learning Surgical Trajectories via Goal-conditioned Offline RL in Pixel Space
Predicting surgical needle trajectories from endoscopic video is critical for robot-assisted suturing, enabling anticipatory planning, real…
Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching
Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality…
コードと自然言語が出会う場所: LLM 統合アプリケーションのための分類ベースの情報フロー分析
LLM API 呼び出しはユビキタスなプログラム構造になりつつありますが、既存のプログラム分析が越えることのできない境界を作り出しています。ランタイム値は自然言語プロンプトに入り、LLM 内で不透明な処理を受け、プログラムが消費するコード、SQL、JSON、またはテキストとして再出現します。テイント分析、プログラムのスライス、依存関係の分析、変更の影響の分析など、関数の境界を越えてデータを追跡するすべての分析は、呼び出し先の動作のデータフロー概要に依存します。 LLM コールにはそのようなサマリーがないため、NL/PL 境界と呼ばれるものでこれらの分析はすべて中断されます。我々は、この境界を埋める最初の情報フロー手法を提案します。定量的情報フロー理論に基づいた私たちの分類法は、情報保存レベル (語彙的に保存されたものから完全にブロックされたものまで) と出力モダリティ (自然言語、構造化フォーマット、実行可能な成果物) という 2 つの直交する次元に沿って 24 のラベルを定義します。 4,154 個の実世界の Python ファイルから 9,083 個のプレースホルダーと出力のペアにラベルを付け、Cohen の $\kappa = 0.82$ およびほぼ完全なカバレッジ (0.01\% 分類不能) で信頼性を検証します。 2 つのダウンストリーム アプリケーションで分類の有用性を実証します。(1) 分類ベースのフィルタリングと LLM 検証を組み合わせた 2 段階のテイント伝播パイプラインは、353 の専門家による注釈付きペアで $F_1 = 0.923$ を達成し、6 つの現実世界の OpenClaw プロンプト インジェクション ケースでの言語間検証で有効性をさらに確認しました。 (2)~分類法に基づいた後方スライスにより、非伝播プレースホルダーを含むファイルのスライス サイズが平均 15\% 削減されます。ラベルごとの分析により、4 つのブロックされたラベルがほぼすべての非伝播ケースを説明していることが明らかになり、ツール ビルダーに実用的なフィルタリング基準が提供されます。
原文 (English)
Reachability Across the NL/PL Boundary: A Taxonomy-Driven Dataflow Model for LLM-Integrated Applications
LLM API calls have become a standard programming primitive, but they create a program boundary that disrupts traditional dataflow analysis. A runtime value may be inserted into a natural-language prompt through a template placeholder, transformed opaquely by the LLM, and returned as code, JSON, or text consumed by downstream logic. Existing analyses such as taint analysis and program slicing require a dataflow summary that describes how a callee maps inputs to outputs; an LLM call provides no such summary, breaking analysis at what we call the NL/PL boundary. We introduce PRISM, the first reachability model for this boundary. PRISM abstracts the missing dataflow summary of an LLM call as placeholder-to-output reachability. Because the LLM's internal transformation is opaque, the only observable signal is the input-output relationship, which spans an unbounded range of behaviors. PRISM therefore uses a finite taxonomy grounded in quantitative information flow theory. It classifies placeholder-output behavior into 25 labels along two dimensions: information preservation and output modality. Each label yields a reachability predicate for a placeholder. The model is sound with respect to its labeling, with residual error bounded empirically. PRISM is dependable and effective. Independent models and human annotators assign its labels consistently (Fleiss' kappa >= 0.72), and the labels cover 8,119 real-world pairs, leaving no pair unclassifiable; the Good-Turing discovery probability is 0.09%. For taint analysis, PRISM nearly doubles the conservative baseline and outperforms a direct LLM baseline, achieving F1 = 81.7%. Across six real OpenClaw CVEs, it detects every vulnerable flow and confirms every patch (F1 = 100%). In backward slicing, it removes about a quarter of irrelevant code without discarding any true dependency.
Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents
Agents that synthesize their own tools ship a second artifact alongside each answer: a software library that future tasks reuse, compose, a…
Streaming Model Cascades for Semantic SQL
Modern data warehouses extend SQL with semantic operators that invoke large language models on each qualifying row, making per-row inferenc…
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
Large Language Models (LLMs) have demonstrated advanced capabilities but often suffer from factual inaccuracies (hallucinations) and system…
AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection
Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible. Au…
Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types
Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks…
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been…
SeaAlert: Robust Severity Classification and LLM-Based Information Extraction for Noisy Maritime Distress Communications
Maritime distress communications transmitted over very high frequency (VHF) radio are safety-critical voice messages used to report emergen…
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the…
Don't Make Models Guess Security and Safety: Symbolic Guardrails for Domain-Specific AI Agents
There is increasing interest in integrating AI agents that invoke tools into domain-specific commercial software, where unintended tool cal…
Unveiling Stochasticity: Universal Multi-modal Probabilistic Modeling for Traffic Forecasting
Traffic forecasting is a challenging spatio-temporal modeling task and a critical component of urban transportation management. Current stu…
Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives
AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the…
Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts
Electroencephalography (EEG) foundation models have shown strong potential for learning generalizable representations from large-scale neur…
Towards Generalizable Deepfake Image Detection with Vision Transformers
In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the…
The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models
As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) a…
StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model
While generative models have set new benchmarks for Target Speaker Extraction (TSE), their inherent reliance on global context precludes de…
CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors
Vision--Language--Action (VLA) models often use intermediate representations to connect multimodal inputs with continuous control, yet spat…
Query2Diagram: Answering Developer Queries with UML Diagrams
Software documentation frequently becomes outdated or fails to exist entirely, yet developers need focused views of their codebase to under…
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
Frontier image generation has moved from artistic synthesis toward synthetic visual evidence. Systems such as GPT Image 2, Nano Banana Pro,…
Kwai Summary Attention Technical Report
Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in se…
Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity
Closed-source frontier labs do not disclose parameter counts. Storing F facts requires at least F/(bits per parameter) weights, so factual…
Fitting Horn DL Ontologies to ABox and Query Examples: A Tale of Simulation Quantifiers and Finite Models
We study the problem of fitting a description logic (DL) ontology to a given set of positive and negative examples that take the form of an…
Shao: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation
A common design pattern in high-quality music generation is to handle structure and fidelity in different representation spaces: a generato…
ARISE: A Repository-level Graph Representation and Toolset for Agentic Program Repair and Fault Localization
Automated program repair at repository scale requires an agent to locate a fault among thousands of files and synthesize a correct patch. E…
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
We present Sparse Backdoor, a supply-chain attack that plants a provably undetectable backdoor in pre-trained image classifiers, including…
IRC-Bench: Recognizing Entities from Contextual Cues in First-Person Reminiscences
When people recount personal memories, they often refer to people, places, and events indirectly, relying on con-textual cues rather than e…
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
Biological laboratory automation can reduce repetitive manual work and improve reproducibility, but reliable embodied execution in wet-lab…
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents
Persistent memory in LLM agents creates an attack surface that production safety classifiers do not observe: the payload enters via RAG ret…
Evolutionary Ensemble of Agents
We introduce Evolutionary Ensemble (EvE), a decentralized framework that organizes existing, highly capable coding agents into a live, co-e…
Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories
We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories. JE…
Active Sensing with Meta-Reinforcement Learning for Emitter Localization from RF Observations
Global navigation satellite system (GNSS) interference poses a serious threat to reliable positioning, especially in indoor and multipath-r…
CogAdapt: リード適応による臨床 ECG 基礎モデルのウェアラブル認知負荷評価への移行
リアルタイムの認知負荷評価は、人間とコンピューターの適応的なインタラクションに不可欠ですが、ラベル付けされたデータが限られており、被験者間の汎化が不十分であるため、依然として困難です。数百万の臨床記録で事前トレーニングされた最近の ECG 基礎モデルは豊富な表現を提供しますが、センサー構成の不一致やタスクの違いのため、ウェアラブル デバイスに直接適用することはできません。この論文では、臨床 ECG 基礎モデルをウェアラブル認知負荷評価に適応させるフレームワークである CogAdapt を提案します。 CogAdapt は、3 リードのウェアラブル信号を解剖学的に一貫した 12 リード表現に変換する学習可能なアダプターである LeadBridge と、壊滅的な忘却を防ぎながらエンコーダ層のフリーズを徐々に解除するプログレッシブ微調整戦略である ProFine を導入しています。 2 つの公開データセット (CLARE および CL-Drive) を 1 被験者除外相互検証で評価したところ、CogAdapt はゼロからトレーニングされたベースラインを大幅に上回り、マクロ F1 スコア 0.626 および 0.768 を達成したことが示されています。これらの結果は、ウェアラブル センサーによる被験者に依存しない認知負荷評価に対する基礎モデルの適応が期待できることを示しています。
原文 (English)
CogAdapt: Adapting Clinical ECG Foundation Models for Wearable Cognitive Load Assessment
Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labeled data are scarce and models generalize poorly across subjects. Recent ECG foundation models, pre-trained on millions of clinical diagnostic ECG recordings, yet they do not apply directly to wearable devices when the sensor configuration and the task both differ. We present CogAdapt, a framework that adapts a clinical ECG foundation model to wearable cognitive load assessment. CogAdapt has two parts. LeadBridge is a learnable adapter that maps 3-lead wearable signals to a 12-lead-compatible representation. ProFine is a progressive fine-tuning strategy that unfreezes encoder layers in stages while limiting representational drift in the pre-trained model. On two public datasets (CLARE and CL-Drive) under leave-one-subject-out cross-validation, CogAdapt reaches macro-F1 of 0.626 and 0.768, improving over from-scratch baselines by 11.2 and 16.1 percentage points. The results show that a clinical ECG pretraining can support subject-independent cognitive load assessment from wearable sensors.
高リスク AI システムと欧州 AI 法におけるアイデンティティの問題
EU 人工知能法 (AIA) は、事前の適合性評価、市販後のモニタリング、および「大幅な変更」時の再評価を中心に構築された高リスク AI システムのライフサイクル ガバナンス体制を確立しています。これらの義務は AI のアイデンティティ判断を前提としています。規制当局とプロバイダーは、更新されたシステムが長期間にわたって同じシステムのままであるかどうかを決定する必要があります。この研究では、このロジックがアーティファクト ID の機能 + フレームワークによってどのように明確化されるかを示します。このフレームワークは、「AI の信頼性」として捉えられる、適切な機能の状況依存の基準とともに、意図された機能によって AI システムを個別化します。さらに我々は、AIAは同期的同一性(規制上の目的で一度に2つのAIシステムが同一とみなされるべき場合)に関する内部の監査可能な基準を提供しておらず、代わりにそのような同一性の判断を分野別または調和化の手段に大きく委ねていると主張する。 function+ は、意図した機能と信頼性のプロファイルとレベルに基づいた同期アイデンティティ テストを提供し、調達、責任、市場監視などのガバナンス設定で同期アイデンティティの決定を検査可能にします。私たちの貢献は概念的なレンズと監査レンズです。私たちは、AIA ライフサイクル義務と機能 + アイデンティティ コンポーネント間の対応マップを提供し、監査と紛争のコンテキストに関する最小限の意思決定フローを通じて同期ケースを運用上判読できるようにします。最後に、実装に向けた 2 つの推奨事項を示します。(1) 意図された目的についての、より正確でテスト可能なレポート。(2) 経時的および導入間での比較可能性をサポートする、標準化された監査可能な信頼性レポート。
原文 (English)
High-Risk AI Systems and the Problem of Identity in the European AI Act
The EU Artificial Intelligence Act (AIA) establishes a lifecycle governance regime for high-risk AI systems built around ex-ante conformity assessment, post-market monitoring, and re-assessment upon "substantial modification." These obligations presuppose AI identity judgments: regulators and providers must decide when an updated system remains the same system over time. In this work, we show how this logic is clarified by the function+ framework of artifact identity, which individuates AI systems by their intended function together with context-sensitive criteria of appropriate functioning, captured as "AI trustworthiness." We further argue that the AIA does not provide an internal, auditable criterion for synchronic identity--when two AI systems at a given time should count as the same for regulatory purposes--and instead largely defers such sameness determinations to sectoral or harmonization instruments. function+ supplies a synchronic identity test anchored in intended function and trustworthiness profiles and levels, making synchronic identity decisions inspectable in governance settings such as procurement, liability, and market surveillance. Our contribution is a conceptual and auditing lens: we provide a correspondence map between AIA lifecycle obligations and function+ identity components, and we make the synchronic case operationally legible via a minimal decision flow for audit and dispute contexts. We conclude with two implementation-facing recommendations: (1) more precise, testable reporting of intended purpose, and (2) standardized, auditable trustworthiness reporting that supports comparability over time and across deployments.
隠された状態のプライバシーには空の中間がある
単層隠れ状態プライバシーについてテストした $1{,}536$ ガウス リリース共分散のうち、適応検索攻撃者に対して中程度の実用性と中程度のプライバシーの両方を達成するものはゼロです。補完的なフィッシャー ボールの下限を証明します。 $O(1)$ でのすべてのフルランク ガウス リリース フィッシャー ユーティリティは、マハラノビス信号が隠れ幅で線形に増加する方向を許可し、クラス内の均一なガウス安全性を除外し、経験的な空の中央と一致します。対角逆フィッシャー リリース $\Sigma^\star_{\mathrm{diag}}(\mathcal{K}) = (2\mathcal{K}/d)\,\mathrm{diag}(1/F_{ii})$ は、一次 KL 予算 $\mathcal{K}$ における独自のミニマックス最適対角メカニズムであり、ワースト攻撃者トップ 1 $\le を持つ唯一のリリースです。 32 モデル層グリッドの各ポイントで 0.001 ドルですが、中央を埋めるのではなく、プライバシー/ユーティリティの境界に位置します。ユークリッド検索下で $13\time$ パレート削減に達する一般化固有メカニズムは、適応マハラノビス攻撃者の下では $100\%$ トップ 1 に崩壊し、全軌道シーケンス インバーターは $94\%$ のクリーンな GPT-2 プレフィックスを回復しますが、$\Sigma_{\mathrm{diag}}$ の下では $0\%$ を回復します。スクラッチからトレーニングされた分割メモリ変換器は、90M で $G_{\mathrm{Mah}} \in [20, 33]$ に達し、固定トークン言語モデリングの損失ペナルティで 30M から 1B までの同じ予算の GPT ベースラインに対して $6$--$24\times$ の優位性を維持します。事前トレーニング済みモデルの最高値は 9.3 です。これらの結果は、隠れ状態のリリースを、ガウス クラス内のメカニズム設計からアーキテクチャまたはリリースの共同設計に再構成します。
原文 (English)
Hidden-State Privacy Has an Empty Middle
Of $1{,}536$ Gaussian release covariances we tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy against an adaptive retrieval attacker. We prove a complementary Fisher-ball lower bound: every full-rank Gaussian release at $O(1)$ Fisher utility admits a direction whose Mahalanobis signal grows linearly in hidden width, ruling out uniform Gaussian safety in the class and matching the empirical empty middle. The diagonal inverse-Fisher release $\Sigma^\star_{\mathrm{diag}}(\mathcal{K}) = (2\mathcal{K}/d)\,\mathrm{diag}(1/F_{ii})$ is the unique minimax-optimal diagonal mechanism at first-order KL budget $\mathcal{K}$ and the only release with worst-attacker top-1 $\le 0.001$ at every point of a 32 model-layer grid, but it sits on a privacy/utility edge rather than filling the middle. A generalized-eigen mechanism reaching $13\times$ Pareto reduction under Euclidean retrieval collapses to $100\%$ top-1 under the adaptive Mahalanobis attacker, and a full-trajectory sequence inverter recovers $94\%$ of clean GPT-2 prefixes but $0\%$ under $\Sigma_{\mathrm{diag}}$. A split-memory transformer trained from scratch reaches $G_{\mathrm{Mah}} \in [20, 33]$ at 90M and maintains a $6$--$24\times$ advantage over same-budget GPT baselines from 30M to 1B at a fixed-token language-modeling loss penalty; pretrained models top out at 9.3. These results reframe hidden-state release from mechanism-design within the Gaussian class to architecture or release co-design.
IndexMem: ロングコンテキスト LLM 推論のための潜在メモリを使用した KV キャッシュのエビクションを学習しました
大規模言語モデル (LLM) は長いコンテキストで動作することがますます期待されていますが、標準のソフトマックス アテンションではシーケンスの長さに応じて線形に増加する KV キャッシュが発生し、すぐに長いコンテキスト推論のボトルネックになります。実際的な解決策は、重要性の低い KV エントリを削除することです。ただし、既存の立ち退きポリシーは主にヒューリスティックであり、トークンの重要性の豊富で入力依存の分布を把握するのに苦労しています。この作業では、KV の重要性を予測する学習可能なインデクサーを導入し、重要なトークンをより正確に保持できるようにします。一方、単純にトークンを削除すると、その情報が永久に破棄され、回復不能な忘却や長距離にわたる取得の低下につながります。これに対処するために、追い出されたトークンをコンパクトなオンライン更新状態に圧縮し、KV 追い出しによって失われた注意の貢献を補うために残留読み出しを提供する軽量の潜在メモリ モジュールを提案します。まとめると、私たちの手法は、制限された KV 予算の下で正確なロングコンテキスト推論を可能にし、Qwen、Mistral、Llama モデル全体で RULER (4K/16K) の一貫した改善 (積極的なエビクション下で最大 25 ポイント)、著しく安定した Needle-in-a-Haystack の取得、既存のエビクション ポリシーと比較して優れた LongBench スコアと圧縮曲線を実現します。
原文 (English)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that grows linearly with sequence length, quickly becoming the bottleneck for long context inference. A practical remedy is to evict less important KV entries; however, existing eviction policies are largely heuristic and struggle to capture the rich, input-dependent distribution of token importance. In this work, we introduce a learnable indexer that predicts KV importance, enabling more accurate retention of critical tokens. Meanwhile, naively evicting tokens permanently discards their information, leading to irreversible forgetting and degraded retrieval over long ranges. To address this, we propose a lightweight latent memory module that compresses evicted tokens into a compact, online-updated state and provides residual readouts to compensate for the attention contributions lost through KV eviction. Collectively, our method enables accurate long-context inference under a bounded KV budget, delivering consistent improvements on RULER (4K/16K) across Qwen, Mistral, and Llama models (up to 25 points under aggressive eviction), markedly more stable Needle-in-a-Haystack retrieval, and superior LongBench scores and compression curves compared to existing eviction policies.
低次元部分空間での学習: 強化学習の直交ボトルネック
タスク関連の価値とポリシー構造が本質的に低次元である可能性があるという証拠が増えているにもかかわらず、深層強化学習 (RL) エージェントは一般に高次元のニューラル表現に依存しています。この研究では、固定正規直交投影を挿入してエンコーダの特徴を低次元の部分空間に制約する、シンプルかつ効果的な表現レベルのプリアを提示します。補助的な目的、事前トレーニング、基礎となる RL アルゴリズムへの変更は必要ありません。線形実現可能性の仮定の下では、ボトルネック次元が特徴空間における最適値関数の固有ランクを超える場合、ボトルネックは表現力を維持し、等価な低次元パラメータ化まで誘導された勾配ダイナミクスを変更しないままにすることを証明します。経験的に、単一タスクとマルチタスクのベンチマークの両方で、ボトルネックの次元がタスクに依存する小さなしきい値を超えると、ベースラインのパフォーマンスが同等か改善されることがわかりました。多くの場合、値の表現は損失なく非常に低い次元に圧縮でき、最小の十分な次元はエンコーダの幅よりも環境の複雑さに大きく依存します。さらに、表現幾何学を分析し、直交ボトルネックが特徴規範を安定させ、より高い有効ランクに関連していることを発見しました。これらの結果を総合すると、強化学習における多様体仮説の表現空間解釈がサポートされ、直交ボトルネックが RL 表現を形成するための軽量でアーキテクチャに依存しないメカニズムとして位置づけられます。
原文 (English)
Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning
Deep reinforcement learning (RL) agents commonly rely on high-dimensional neural representations, despite growing evidence that task-relevant value and policy structure may be intrinsically low-dimensional. In this work, we present a simple yet effective representation-level prior that inserts a fixed orthonormal projection to constrain encoder features to a low-dimensional subspace, requiring no auxiliary objectives, pretraining, or changes to the underlying RL algorithm. Under a linear realizability assumption, we prove that when the bottleneck dimension exceeds the intrinsic rank of the optimal value function in feature space, the bottleneck preserves expressivity and leaves the induced gradient dynamics unchanged up to an equivalent low-dimensional parameterization. Empirically, we find that across both single and multi-task benchmarks, baseline performance is either matched or improved once the bottleneck dimension exceeds a small task-dependent threshold; in many cases, value representations can be compressed to extremely low dimensions without loss, and the minimal sufficient dimension depends far more on environment complexity than encoder width. In addition, we analyze representation geometry and find that orthogonal bottlenecks stabilize feature norms and are associated with higher effective rank. Together, these results support a representation-space interpretation of the manifold hypothesis in reinforcement learning and position orthogonal bottlenecks as a lightweight, architecture-agnostic mechanism for shaping RL representations.
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
Memory is essential for enabling large language models to support long-horizon reasoning, yet existing memory systems remain unreliable and…
言語的監視なしの物理的相互作用を通じたワールドモデルにおける創発的な意味表現
世界モデルは、言語による監視なしに、物理的な探索から何を学ぶのでしょうか?私たちは、その答えは単一の原理、つまり物理世界の幾何学的構造によって整理されると主張します。 VAE ベースの世界モデルをランダムに具現化された探索でトレーニングすると、その潜在空間が物理幾何学を反映する空間意味構造を発達させることがわかりました。方向精度はランダムに初期化されたエンコーダーの場合は 0.677+-0.029 対 0.547、位置 RSA はランダム エンコーダーの場合は 0.192+-0.047 対 0.029 (6.6 倍の改善) であり、次のことがわかります。トレーニングは、CNN の帰納的バイアスを超えた真の構造的組織化を誘発します。 20 の時間チェックポイントにわたって、予測パフォーマンスとセマンティック整合性が同時に向上し (Spearman r=-0.61、p=0.004)、共有ドライバー アカウントと一致しています。これは二重ノックアウトによって確認されます。標準の KL 正則化 (ベータ = 0.1) により、エンコーダーが幾何学的構造から強制的に遠ざけられ、予測パフォーマンスとセマンティック アラインメントの両方が、共有ドライバー アカウントの予測どおり、ステップ 50,000 までにほぼ偶然に同時に崩壊します。ベータを 0.001 に下げると、幾何学的アクセスが復元され、両方の機能が一緒に回復します。これらの発見は、物理世界の幾何学を世界モデル表現の組織原理として確立し、意味論的に根拠のある身体化されたエージェントの設計に直接的な影響を及ぼします。
原文 (English)
Emergent Semantic Representations in World Models through Physical Interaction without Linguistic Supervision
What does a world model learn from physical exploration, without any linguistic supervision? We argue the answer is organized by a single principle: the geometric structure of the physical world. Training a VAE-based world model on random embodied exploration, we find that its latent space develops spatial semantic structure that mirrors physical geometry -- direction accuracy 0.677+-0.029 versus 0.547 for a randomly initialized encoder, and position RSA 0.192+-0.047 versus 0.029 for random encoders (6.6x improvement), showing that training induces genuine structural organization beyond CNN inductive bias. Across 20 temporal checkpoints, prediction performance and semantic alignment co-improve (Spearman r=-0.61, p=0.004), consistent with the shared-driver account. We confirm this through a double knockout: standard KL regularization (beta=0.1) forces the encoder away from geometric structure, and both prediction performance and semantic alignment collapse simultaneously to near-chance by step 50,000 -- exactly as the shared-driver account predicts. Reducing beta to 0.001 restores geometric access and recovers both capabilities together. These findings establish physical world geometry as the organizing principle of world model representations, with direct implications for the design of semantically grounded embodied agents.
Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations
We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for w…
Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms
As autonomous language model agents proliferate, forming an emerging agentic web with real-world consequences, what credibility signals can…
De-attribute to Forget for LLM Unlearning
The rapid development of large language models (LLMs) has raised concerns on the use of inappropriate data for training, which has led to a…
Automatically Differentiable Nonlinear Tensor Networks (ADNTNs) for Exponential Parameter Compression of Deep Neural Networks
Large deep neural networks are costly to store and deploy because inference must move and evaluate many parameters. This paper studies \emp…
ParetoPilot: Infer-Perturb-Guide 拡散によるゼロサロゲートオフライン多目的最適化
オフライン多目的最適化 (オフライン MOO) は、高価な環境との相互作用を行わずに、静的データセットに基づいた新しいパレート最適設計を発見することを目的としています。最近の生成手法は顕著な成功を収めていますが、主に外部サロゲート モデルに依存しています。この依存関係により、重大な計算オーバーヘッドが生じ、欺瞞的な評価に悩まされ、主流の生成モデルを条件付きで共同トレーニングするという一般的なパラダイムから逸脱します。これらのボトルネックに対処するために、オフライン MOO 用の新しいゼロ代理拡散フレームワークである ParetoPilot を提案します。 ParetoPilot は、事前トレーニングされた拡散モデルに本質的に組み込まれている条件付き事前確率を最大限に活用します。このフレームワークの核心として、Infer-Perturb-Guide (IPG) エンジンが導入されており、このエンジンは逆生成プロセスの無条件ノイズ除去ステップ内にシームレスにインターリーブされます。まず、条件付きおよび無条件のノイズ予測を照合することで、瞬間的な目標方向を暗黙的に推測します。次に、厳密な収束のために平行な重力場と相互多様性のためにエッジを意識した斥力を数学的に直交化し、動的にアニールされた摂動ベクトルを作成します。最後に、この摂動されたターゲットは、標準の分類子なしガイダンス (CFG) を介して生成プロセスをシームレスに制御します。 51 のタスクにわたる広範な実験により、ParetoPilot が 14 の最先端のサロゲートベースおよび逆生成ベースラインよりも優れたパフォーマンスを発揮することが実証されました。補助的なプロキシ トレーニングを排除することで、当社のアプローチはデータのプライバシーを保護しながら、ハイパーボリュームの改善と堅牢なパレート フロント カバレッジを実現します。
原文 (English)
ParetoPilot: Zero-Surrogate Offline Multi-Objective Optimization via Infer-Perturb-Guide Diffusion
Offline multi-objective optimization (Offline MOO) seeks Pareto-optimal designs from static datasets without additional environment interactions. Existing generative methods typically guide sampling with external surrogate or preference models, which adds training complexity and may provide unreliable guidance. We propose ParetoPilot, a plug-and-play method that guides designs to Pareto front at inference time using a pre-trained conditional diffusion model without any surrogate. ParetoPilot introduces an Infer-Perturb-Guide (IPG) engine within the reverse diffusion process. IPG first infers the individual conditional target for each sample in the batch by aligning its conditional and unconditional predictions. It then perturbs these targets collectively across the batch, balancing convergence toward the Pareto front and diversity among samples. Finally, the engine guides the generative trajectory toward the Pareto front by injecting these perturbed targets via standard Classifier-Free Guidance (CFG). Experiments on 51 tasks demonstrate that ParetoPilot achieves the best overall ranking among 16 methods and competitive hypervolume improvement.
必要なのは FP8 だけです (パート 1): HPC の聖杯としてのハードウェア FP64 の誤りを暴く
従来の HPC の定説では、ネイティブ ハードウェア FP64 シリコンは科学技術コンピューティングの還元不可能な基盤、つまり倍精度シミュレーションの「聖杯」であると考えられています。この論文では、この定説は間違っていると主張しています。B300 世代以降の AI に最適化された GPU では、豊富な FP8 テンソル スループットと中国剰余定理ベースの Ozaki Scheme II を組み合わせることで、正規の HPC カーネル スペクトル全体で完全な FP64 精度でメモリルーフ実行を回復します。 NVIDIA の Blackwell Ultra (B300) は、ネイティブ FP64 を約 1.3 TFLOPS (B200 から 31 倍) に低下させ、メモリに依存するカーネル (SpMV、GEMV、ステンシル) も計算に依存するようにレンダリングします。私たちは4つの貢献をしています。まず、統合分析モデルである Tensor-Memory Equilibrium (TME) モデルは、計算乗数アルファ、帯域幅乗数ベータ、および再構築レイテンシ ガンマでルーフラインを強化します。次に、ベータ -> 1 を駆動するメカニズムとしてレジスタレベルの融合を特定し、エミュレーションをメモリの壁の向こう側で本質的に自由にします。 3 番目に、Ozaki II ヴォールトは FP64 をネイティブの最大 1 TFLOPS から最大 500 TFLOPS (B300) および最大 400 TFLOPS (Rubin R200) までエミュレートし、帯域幅制限の領域ではメモリ上限に匹敵しながら、コンピューティング領域では B200 のネイティブ FP64 の上限を 1 桁以上上回ったと予測します。 4 番目に、H100 ベースラインに対して、Ozaki II は、B300 ネイティブ FP64 が課す最大 50 倍の回帰と比較して、調査したすべてのワークロードで H100 と一致またはそれを超えています。コンパニオン FFT 解析 (生き残った INT32 パイプでの Kulisch 固定点再構築) と、コンパニオン Part(2) 論文で報告されている FP32+Kahan 削減と組み合わせると、B300 で調査されたすべてのカーネル クラスがフル FP64 でメモリ ルーフに達します。証拠はタイトルの主張を裏付けています。Ozaki II と Kulisch のエスケープ ルートを備えた FP8 は、実稼働 HPC に必要なすべてです。ネイティブ FP64 シリコンは、もはやこれまで考えられてきた聖杯ではありません。
原文 (English)
FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (June 13th version)
Conventional HPC holds that native hardware FP64 is the irreducible foundation of scientific computing. On AI-optimized GPUs of the NVIDIA B300 generation and beyond, native FP64 throughput has collapsed to ~1.3 TFLOPS even as FP8 tensor throughput has grown to multiple PFLOPS. We argue something stronger than that this is survivable: the FP8 tensor-core matrix-multiply is the sole computational primitive on which double-precision scientific computing needs to be built. Every canonical kernel -- dense and sparse linear algebra, spectral transforms, stencils -- and every application composing them reduces, via the Chinese Remainder Theorem-based Ozaki Scheme II, to sequences of FP8 matrix operations; the only non-FP8 arithmetic is a bounded, fixed-width integer accumulation at reconstruction. Native FP64 is thereby demoted from a hardware requirement to a derived accuracy guarantee obtained by composition over the FP8 primitive. We organize the claim as a five-layer hierarchy -- the FP8 op, Ozaki II, the basic kernels or Berkeley "dwarfs", composite solvers, and full applications -- and, because the dwarf taxonomy already spans scientific computing, establish it by exhibiting the reduction for every dwarf rather than a sample. The claim is falsifiable, and we build the instrument that tests it: a Tensor-Memory Equilibrium (TME) model extending the Roofline with emulation parameters (alpha, beta, gamma). We identify register-level fusion as the mechanism that keeps emulation memory-bound, project recovered FP64 performance across B300 and Rubin against an H100 baseline, and close the kernel coverage with a companion FFT analysis and compensated reductions. The model could have returned a negative verdict; instead it passes across the dwarfs and their compositions. This is the analytical half of a two-part program, with a follow-on implementation to validate the thesis on real silicon.
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify…
Quickest Detection of Hallucination Onset: Delay Bounds and Learned CUSUM Statistics
Token-level hallucination detectors are evaluated as classifiers, by AUC over all tokens, yet a streaming monitor is judged by its reaction…
A Mathematical Theory of Value: a synthesis on goal-directed agency under resource constraints
We propose that value -- the quantity goal-directed agents create, destroy, and exchange -- is a lawful structural quantity in the same cat…
Iterative Visual Thinking and the Self-Correction Mirage in VLM Grounding
Letting a vision-language model (VLM) think longer at test time has driven much recent progress. A natural way to bring this to spatial gro…
Lect\=uraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching
Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but…
Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering
Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by…
Attention is Just Another Name for Coupling? A Fast-Slow ODE Perspective on Hierarchical Pretraining
We re-interpret Transformer pretraining as a fast-slow, singularly perturbed flow along depth, with untied weights as its non-autonomous fe…
A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning
Reinforcement learning (RL) systems often degrade when operating conditions differ from those previously encountered, reflecting distributi…
TransitNet: A Compact Attention-Augmented Deep Learning Framework for Low-SNR Transit Blind Searches
Motivated by the observational incompleteness of intermediate-to-long-period Earth-size planets, we present TransitNet, a compact attention…
OrthoReg: Orthogonal Regularization for Hybrid Symbolic-Neural Dynamical Systems
Dynamical systems are fundamental to modeling the natural world, yet modeling them involves a persistent trade-off: manually prescribed mec…
難しいですか、それとも到達していないだけですか?数学的推論の難易度推定におけるサンプリングの盲点を診断する
数学と科学の推論ベンチマークは、標準的な例ごとの難易度シグナルとして、ゴールドに達するサンプリングされたチェーンの割合である pass@k に依存します。同じシグナルが、検証可能な報酬、数学データのキュレーション、合成カリキュラム、検証者のトレーニングによる RL を推進します。このプロキシには、最も難しい層に永続的な盲点があることがわかります。テストした 8 つの自由形式数学セル (4 つのオープンウェイト モデルにわたる GSM8K および MATH) では、サンプリング シードが 6 回の試行で解決できなかった例の 10.3 ~ 22.9% が、代わりに 6 チェーンの決定論的体制による一致したコンピューティングで解決されました。これらは、貪欲なデコードにアクティベーション グラフティングを介して適用される 5 つの安価な残差ストリーム摂動を加えたものですが、貪欲だけではこれらの数学セルでは最大 6% しか解決できません。回復は追加の予算に応じて調整され、摂動全体でそのメカニズムの区別性が 12 個のセルすべてで検証されます (すべての設定で異種固定セット Jaccard <= 0.47)。アクティベーショングラフティングは、解読方法ではなく、内部表現への介入として使用されます。私たちはそれを純粋に診断および多様化ツールとして使用しており、復元されたアイテムは、変更されていないモデルが通常の推論で到達するのではなく、 pass@k= 0 % 層が残差ストリーム内で構造的に識別可能であることを示しています。
原文 (English)
Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation
Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficulty signal. The same signal drives RL with verifiable rewards, math data curation, synthetic curricula, and verifier training. We show this proxy has a persistent blind spot on its hardest stratum: on the eight free-form math cells we test (GSM8K and MATH across four open-weight models), 10.3-22.9% of the examples that no sampling seed solves in six tries are instead solved at matched compute by a six-chain deterministic regime. These are greedy decoding plus five cheap residual-stream perturbations applied via activation grafting, while greedy alone solves at most 6% on these math cells. Recovery scales with the additional budget, across perturbations whose mechanistic distinctness we verify across all twelve cells (cross-kind fix-set Jaccard <= 0.47 in every setup). Activation grafting is used as an intervention on internal representations, not a decoding method; we use it purely as a diagnostic and diversification tool, and our recovered items show that the pass@k= 0 % stratum is structurally identifiable in the residual stream rather than that the unmodified model reaches them under ordinary inference.
オプティカル フローを学習するための普遍的な制約としての三角形の一貫性
我々は、オプティカル フローの第一原理制約として三角整合性を提案します。これは、ネットワーク アーキテクチャ、監視タイプ、データセットに依存せず、画像ペアとマルチフレーム設定の両方に適用されます。このシンプルだが強力な制約は、2 つのフローを構成して 3 つ目のフローを誘導し、3 つのフロー間の一貫性を強制することです。合成されたフローは、(i) 画像ペアから生成され、サイクルの一貫性が得られます。 (ii) 複数のビデオ フレーム。時間的連鎖を通じてより長距離の動きを生成します。または (iii) 画像ペアを制御された合成変換と組み合わせて、データ拡張となります。この三角形の一貫性により、無視できるほどの計算オーバーヘッドが発生し、追加の注釈は必要ありません。オプティカル フローのジオメトリから直接導出されるため、モデル固有の仮定に依存せず、オプティカル フロー トレーニング用の「ユニバーサル」プラグ アンド プレイ コンポーネントとして機能します。実験では、教師あり、教師なし、転移学習の設定全体で一貫した改善が見られました。
原文 (English)
Triangular Consistency as a Universal Constraint for Learning Optical Flow
We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision type, and dataset, and applies to both image-pair and multi-frame settings. This simple but powerful constraint is to compose two flows to induce a third flow and enforce consistency among the three. The composed flows may arise from (i) image pairs, yielding cycle consistency; (ii) multiple video frames, producing longer-range motion through temporal chaining; or (iii) image pairs combined with controlled synthetic transformations, which becomes data augmentation. This triangular consistency introduces negligible computational overhead and requires no additional annotations. Since it is derived directly from the geometry of optical flow, it does not rely on model-specific assumptions and serves as a ``universal'' plug-and-play component for optical flow training. Experiments show consistent improvement across supervised, unsupervised, and transfer learning settings.
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance…
NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms
Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness und…
A Digital Twin Framework for Traffic-Aware UAV Pavement Monitoring in Open-Traffic Conditions
UAV-based pavement inspection can reduce the cost and risk of road-surface monitoring, but real-world deployment remains difficult when tra…
The Power of Light: Improving Synthetic-to-Real Domain Adaptation through Physically-Based Indirect Illumination
While synthetic data generation resolves the manual labeling bottleneck in computer vision, minimizing the syn-to-real domain gap requires…
VideoAgent: All-in-One Framework for Video Understanding and Editing
Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and…
ドラマディレクター: 幾何学に基づいた短編ドラマの生成
短いドラマは、素早いショットのリズム、対話による焦点の移動、映画的な基礎が要求されるため、プロンプト レベルまたはテキストのみのビデオ生成パイプラインでは対応するのが難しい課題が生じます。私たちは、グローバル プロットとローカル コンテキストが視覚的に根拠のあるマルチショット ビデオに変換される、プロットから短編ドラマへの生成を研究しています。私たちは、プランナーが深度とポーズによってインデックス付けされた実際の短編ドラマのショットのギャラリーから映画のジオメトリを借用できるようにする、ジオメトリに基づいたフレームワークである DramaDirector を提案します。 DramaDirector は、各ショットを静的なビジュアルと動的なナラティブ条件に分離し、学習されたテキストとビジュアルのアライメント報酬の下でスキーマに制約された SFT と GRPO を使用してプランナーをトレーニングし、最初のフレームの生成と画像からビデオへの合成をガイドする深度ポーズ参照を取得します。また、構造化されたストーリーボードと多次元の評価プロトコルを備えた、35 の実写ドラマ、2.8K のエピソード、81K のショットから構築されたベンチマークである DramaBoard も紹介します。実験の結果、DramaDirector は、忠実性、一貫性、および制御性に関して、代表的なマルチエージェントおよびビデオ生成のベースラインよりも向上していることが示されています。私たちのコードは https://github.com/iLearn-Lab/DramaDirector でリリースされています。
原文 (English)
DramaDirector: Geometry-Guided Short Drama Generation
Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation pipelines struggle to meet. We study plot-to-short-drama generation, where a global plot and local context are transformed into visually grounded multi-shot videos. We propose DramaDirector, a geometry-grounded framework that lets the planner borrow cinematographic geometry from a gallery of real short-drama shots indexed by depth and pose. DramaDirector decouples each shot into static visual and dynamic narrative conditions, trains the planner with schema-constrained SFT and GRPO under a learned text-visual alignment reward, and retrieves depth-pose references to guide first-frame generation and image-to-video synthesis. We also introduce DramaBoard, a benchmark built from 35 live-action dramas, 2.8K episodes, and 81K shots, with structured storyboards and multi-dimensional evaluation protocols. Experiments show that DramaDirector improves over representative multi-agent and video generation baselines on faithfulness, consistency, and controllability. Our code is released at: https://github.com/iLearn-Lab/DramaDirector
UC-Search: Risk-Aware Test-Time Search for Delayed Constrained Time-Series Control
Time-series deployments often need delayed feasible decisions, not only accurate forecasts. UC-Search is a trace-only retained-search layer…
Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
A language model's memory can be worse than no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behin…
ハイブリッド プライバシーを意識したセマンティック検索: SVD で切り詰められたドキュメント ジオメトリと、制限された脅威モデルに基づく CKKS 暗号化クエリの再ランキング
高密度埋め込みはセマンティック検索と検索拡張生成を強化しますが、埋め込み反転攻撃はベクトルからソース テキストを再構築する可能性があります。ベクトル データベースが漏洩すると、その背後にある文書も漏洩します。教科書的な防御策は極端です。検索全体を準同型的に暗号化するのは健全ですが、100 万文書規模では遅すぎます。その一方で、保護するずっと前にプライバシー ノイズによってランキングが低下します。静的コレクションと動的クエリの間の非対称性を利用した中間パスを研究します。コレクションは幾何学的に保護されています。各ベクトルは低次元の SVD 部分空間上で切り詰められ、所有者のみが知っている秘密の直交変換によって回転されます。クエリは暗号的に保護されています。クエリは CKKS 準同型暗号化の下で再ランク付けされるため、正直だが好奇心旺盛なサーバーはクエリやスコアを見ることはありません。 CKKS パラメータは、小規模なオフライン ベンチマークから取得されます。私たちは、保護された部分空間に限定された攻撃者の再構成エラーの厳しい下限を証明します。 100 万のドキュメントと 5 つのエンコーダでは、このスキームは 1 秒未満のレイテンシでランキングの品質を維持し (線形デノイザーとして強力なエンコーダでわずかに向上します)、保護されたスペースに対する既製の反転攻撃はノイズ フロアまで崩壊します。次に、より強力な敵対者をテストします。既知の平文攻撃者は、保持された次元とほぼ同じ数の漏洩ペアから直交プロクラステスによる回転を回復します。公開されている積量子化コードは、最近傍構造を保存します。ランダム投影、校正済みノイズ、および BEIR ベースラインは、切り捨てが無料のデノイザーではなく、エンコーダーに依存する精度コストであることを示しています。私たちは限界を述べています。クエリの機密性は暗号化されていますが、ドキュメントの保護は経験的な難読化レイヤー (SVD の切り捨てと秘密のローテーション) であり、暗号化のプリミティブではありません。また、各主張の脅威モデルを区切ります。
原文 (English)
Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model
Dense embeddings power semantic search and Retrieval-Augmented Generation, yet a leaked vector database leaks the text behind it, since embeddings invert with high fidelity. The textbook defences are extreme--homomorphic search is sound but far too slow at million-document scale, while privacy noise degrades ranking before it protects. We study a middle path built on an asymmetry: each static document vector is SVD-truncated and then rotated by a secret orthogonal transform held only by the data owner, while the dynamic query is protected cryptographically under CKKS, so an honest-but-curious server sees neither query values nor scores; the CKKS parameters are fixed by a small reproducible benchmark. We prove a tight lower bound on the reconstruction error of any decoder confined to the protected subspace. On a one-million-document, five-encoder corpus the wrapper preserves retrieval quality at sub-second latency--a mild linear denoiser on self-retrieval that reverses into a 2--8-point nDCG@10 cost on graded relevance--while an off-the-shelf inversion attack collapses to the floor. We then map the boundary: a known-plaintext attacker recovers the rotation by orthogonal Procrustes from about as many leaked pairs as the retained dimension, and the public quantization codes leak neighbour structure. The same geometry doubles as a privacy-preserving data-loss-prevention primitive for LLM firewalls, matching a plaintext detector at near parity. We state the limits plainly: query confidentiality is cryptographic, but document protection is an empirical obfuscation layer, not a cryptographic primitive.
細胞培養プロセス予測のためのラマンデータ融合を使用したマルチパス適応ゲート型ボトルネック潜在 ODE
哺乳類の細胞培養プロセスは多くのバイオ医薬品の製造を支えていますが、計画どおりに稼働し続けることは困難です。重要なプロセスパラメータは数日で変動し、規格外の傾向が確認されると介入するには遅すぎることがよくあります。初期段階の複数日にわたる予測により、供給、サンプリング、制御をタイムリーに調整できる可能性がありますが、バイオプロセスの予測は困難です。測定値がまばらで不規則にサンプリングされ、操作条件が細胞株や培地間で不均一であり、初期挙動がほぼ同じである実行が異なる将来に分岐する可能性があるためです。ゲート付きボトルネック潜在常微分方程式 (GB-Latent ODE) とマルチパス ジャストインタイム微調整 (MP-JIT-FT) を組み合わせた適応フレームワークを提案します。 GB-Latent ODE は、学習可能な変数ごとのゲーティングと、高次元のスパース入力を圧縮するマスク対応ボトルネックを備えた標準 Latent ODE を拡張し、限られたデータの下での学習を向上させます。部分的に観測された実行を考慮すると、MP-JIT-FT は同様の過去の軌跡を取得し、局所近傍を候補レジームにクラスター化し、レジームごとに個別のモデルを微調整して複数のもっともらしいパスを生成します。各パスには単一の平均予測ではなく、再構成ベースの信頼スコアが付けられます。さらに、ラマン分光データを融合します。機械学習ソフト センサーは、高密度のラマン スペクトルを疑似観測に変換し、まばらなオフライン測定を強化して、より堅牢なトレーニングを実現します。 14 条件にわたる 38 回のフェドバッチ 5L バイオリアクターの実行で、ラマン融合を使用した MP-JIT-FT は最高の平均ランクを達成し、9 つのターゲット変数のうち 8 つでグローバルな潜在 ODE ベースラインを上回りました。ローカルダイバージェンスメトリクスを使用して、局所的に類似したプレフィックスが発散する場合にマルチパスゲインが最大になるのに対し、初期のダイナミクスが後の動作を表す場合にはラマン融合が最も役立つことを示します。
原文 (English)
Multipath Adaptive Gated Bottleneck Latent ODE with Raman Data Fusion for Cell Culture Process Forecasting
Mammalian cell-culture processes underpin the manufacture of many biopharmaceuticals, yet keeping a run on track is hard: critical process parameters drift over days, and an off-specification trend is often confirmed too late to intervene. Early-stage, multi-day forecasts could enable timely adjustment of feeding, sampling, and control, but bioprocess forecasting is challenging because measurements are sparse and irregularly sampled, operating conditions are heterogeneous across cell lines and media, and runs with near-identical early behaviour can diverge into different futures. We propose an adaptive framework combining a Gated Bottleneck Latent Ordinary Differential Equation (GB-Latent ODE) with Multi-Path Just-In-Time Fine Tuning (MP-JIT-FT). The GB-Latent ODE augments the stan dard Latent ODE with learnable variable-wise gating and a mask-aware bottleneck that compress high-dimensional sparse inputs, improving learning under limited data. Given a partially observed run, MP-JIT-FT retrieves similar historical trajectories, clusters the local neighbourhood into candidate regimes, and fine-tunes a separate model per regime to produce multiple plausible paths, each with a reconstruction-based confidence score, not a single averaged forecast. We further fuse Raman spectroscopy data: a machine-learning soft sensor turns dense Raman spectra into pseudo-observations that enrich the sparse offline measurements for more robust training. On 38 fed-batch 5L bioreactor runs spanning 14 conditions, MP-JIT-FT with Raman fusion achieves the best average rank and outperforms a global Latent ODE baseline on 8 of 9 target variables. Using local-divergence metrics, we show the multi-path gains are largest when locally similar prefixes diverge, whereas Raman fusion helps most when early dynamics are representative of later behaviour.
残留重み付け補正を備えたヘビーボール Q ラーニング
本論文では、強化学習(RL)のための修正ヘビーボールQ学習法を提案し、その収束を確立する。また、この方法が標準の Q 学習よりも速く収束することが理論的に保証される条件も特定します。次に、同じ構造が線形関数近似を使用して Q ラーニングに拡張され、類似の収束ステートメントと加速ステートメントが導出されます。この分析は、Q 学習アルゴリズムのスイッチ線形システム (SLS) 表現と、関連するスイッチング ファミリの結合スペクトル半径 (JSR) に基づいています。この SLS の観点は、Q 学習の標準的な分析では一般的に使用されません。これは、補完的なフレームワークと、重いボールの勢いがどのように Q 学習を加速できるかについての新しい洞察を提供します。
原文 (English)
Heavy-Ball Q-Learning with Residual Weighting Correction
This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes convergence of its deterministic mean dynamics. It also identifies conditions under which the method is theoretically guaranteed to converge faster than standard Q-learning. The same construction is then extended to Q-learning with linear function approximation, where analogous convergence and acceleration statements are derived for the corresponding corrected fixed point. The sampled stochastic versions are treated through conditional-mean recursions and, in the stated linear-function-approximation setting, finite-time bounds. The analysis is based on a switched linear system (SLS) representation of Q-learning algorithms and on the joint spectral radius (JSR) of the associated switching families. This SLS viewpoint is not commonly used in standard analyses of Q-learning, and it provides a complementary framework and new insight into how heavy-ball momentum can accelerate Q-learning.
The Unverifiability of Artificial General Intelligence (AGI) Alignment, Static and Dynamic: From Trakhtenbrot's Wall to the Safety-Generality Tension
We establish the mathematical limits of AGI safety in two forms: verifying a fixed system, and verifying that a certified safety property p…
Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation
Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards.…
A Deep Multiscale Neural Network for Accurate Neurological Disorder Detection from MRI Scans and Real-Time Web Deployment
Neurological disorders involve diverse pathologies of the brain and nervous system, making early and accurate detection essential. While ma…
Automating the Design of Embodied Agent Architectures
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity expo…
Curvature-Guided Sheaf Diffusion for Unsupervised Community Detection on Heterophilic Graphs
Detecting communities in heterophilic graphs -- where connected nodes often belong to different classes -- is hard for unsupervised methods…
A Stochastic--Geometric Theory of Scaling Laws in Grokking
Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only…
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech
With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly impor…
細胞周期を意識した単細胞薬物摂動応答のモデル化
単細胞薬物摂動モデルは、転写反応の大きさだけでなく、治療によって細胞の増殖状態が変化するかどうかも予測する必要があります。細胞周期の変動は迷惑な変動として扱われることが多く、ベンチマーク パイプラインでは薬剤誘発性の相変化を主な予測ターゲットとして扱うことはほとんどないため、これは困難です。 scCycleMol は、標準化された分子アイデンティティ、用量と細胞株のメタデータ、および治療状態から得られる細胞周期監視による遺伝子発現を備えた厳選された 24 時間の SciPlex3 ベンチマークに基づいて構築された、細胞周期を意識した摂動予測フレームワークです。 scCycleMol は、入力共変量として細胞周期状態を使用する代わりに、予測された処理発現から監視を導き出し、それを環状 G1/S/G2M 期ターゲットを備えた学習可能な完全発現細胞周期ヘッドを通じて伝播させます。マーカーベースの監視、分子表現、および事前トレーニング戦略を評価して、改善の原因を特定します。 600,000 個を超える細胞、186 の摂動条件、複数のがん細胞株、および数千の遺伝子を含む SciPlex3 ベンチマーク全体で、scCycleMol は条件付き摂動ベースラインと比較して分布外発現予測を向上させます。 LINCS で事前トレーニングされた最良の循環モデルは、LINCS で事前トレーニングされた ChemCPA の 0.6800 および 0.5400 と比較して、予想される全遺伝子の r の 2 乗が 0.9093、差次的に発現される遺伝子の 2 乗が 0.6843 と予想されます。クローズドループの細胞周期監視により、ほぼ変化しない発現予測を維持しながら、位相精度が約 0.5 ~ 0.6 ポイント向上します。 Tahoe で事前学習されたバリアントは位相精度 0.9609 に達し、摂動モデリングにおける細胞周期を意識した明示的な監視の利点を強調しています。
原文 (English)
Modeling Cell-Cycle-Aware Single-Cell Drug Perturbation Responses
Single-cell drug perturbation models should capture transcriptional response magnitude and whether a treatment changes the proliferative state of the cell. This is difficult because cell-cycle variation is often treated as a nuisance factor, and benchmark processing rarely makes drug-induced phase changes a primary prediction target. We introduce scCycleMol, a cell-cycle-aware perturbation prediction framework built on a curated 24-hour SciPlex3 benchmark with standardized molecule identities, dose and cell-line metadata, modeled genes, and expression-derived cell-cycle supervision. scCycleMol derives cell-cycle supervision from the treated state and applies it to predicted treated expression without using phase as an input covariate. The model includes a learnable full-expression cell-cycle head with circular G1/S/G2M targets, and we evaluate readout-only supervision (with stop-gradient) versus closed-loop supervision (backpropagating through decoder, dose-response module, and drug representation). We also compare molecular representations and pretraining sources to isolate the effect of the cell-cycle objective. On a processed 24-hour SciPlex3 benchmark (635,541 cells, 186 perturbations, 188 compound embeddings, 3 cell lines, 4 doses plus DMSO, 5,080 genes), the best LINCS-pretrained circular variant reaches 0.9093 mean all-gene R-squared and 0.6843 mean DE-gene R-squared. Under matched preprocessing, closed-loop cell-cycle supervision improves phase accuracy by 0.54-0.62 points while keeping mean all-gene R-squared within 0.003 of matched chemCPA no-cell-cycle models; Tahoe-pretrained readout-only circular supervision achieves the strongest phase accuracy at 0.9609.
From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation
Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unpr…
CVE-TTP KG: Knowledge Graph Linking Software Vulnerabilities to Attack Behaviors
In the evolving threat landscape, adversaries exploit software vulnerabilities to launch sophisticated attacks, challenging traditional def…
WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignor…
GR2 Technical Report
Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking --…
TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands…
スペクトル幾何学とボソンブロッホプローブ: 量子学習の探求
この論文では、量子学習モデルでスペクトル幾何学がどのように現れるか、そしてそれを物理的に接地されたプローブでどのように診断できるかを研究します。グラフ正則化量子ネットワークでは、トレーニングにより出力類似度グラフが再編成され、有効スペクトル次元 デルタ S = +0.23 が増加し、ラプラシアン スペクトルが再形成されます。エッジ分解された 2 ボソン干渉は、この再構成を直接調査します。ボソン強化デルタ P_uv は、フィードラー エッジ スプリット |デルタ v_2| と相関します。 (r = -0.50)、学習されたスペクトル分割を干渉シグネチャにリンクします。位相図は、結合強度ガンマとノイズ デルタに対する性能の非単調な依存性を示しており、グラフの正則化により、制限された領域でのみ忠実度が向上します。ハードウェア実験により、ショットノイズの不確実性の範囲内で予測される干渉挙動が確認されます。また、ハイブリッド量子オートエンコーダーを分析し、その潜在表現の幾何学的診断としてブロッホ空間ドリフトを導入します。教師なし良性データしきい値を使用すると、モデルは高いランキング パフォーマンス (ROC-AUC 約 0.99) と無視できる程度の偽陰性率を達成します。絶対的なブロッホ ドリフトは異常を強く識別します (ROC-AUC 少なくとも約 0.9)。一方、連続的なドリフトはほぼランダムです (ROC-AUC 約 0.5)。これは、検出が局所的な変動ではなく永続的な状態空間の変位から生じることを示しています。これらの結果は、縮小単一量子ビット状態の幾何学と関連する量子フィッシャー情報を通じて、学習によって引き起こされるスペクトル組織化が測定可能な量子状態構造として現れ、ボソンプローブとブロッホプローブを使用して量子学習システムを診断するための統一されたスペクトル幾何学的フレームワークを確立することを示しています。
原文 (English)
Spectral Geometry and Bosonic-Bloch Probes: Explorations in Quantum Learning
This paper studies how spectral geometry emerges in quantum learning models and how it can be diagnosed with physically grounded probes. In graph-regularized quantum networks, training reorganizes the output similarity graph, increases the effective spectral dimension Delta S = +0.23, and reshapes the Laplacian spectrum. Edge-resolved two-boson interference directly probes this restructuring: the bosonic enhancement Delta P_uv correlates with the Fiedler edge split |Delta v_2| (r = -0.50), linking learned spectral partitions to interference signatures. A phase diagram shows a nonmonotonic dependence of performance on coupling strength gamma and noise delta, with graph regularization improving fidelity only in a restricted regime; hardware experiments confirm the predicted interference behavior within shot-noise uncertainty. We also analyze a hybrid quantum autoencoder and introduce Bloch-space drift as a geometric diagnostic of its latent representation. With an unsupervised benign-data threshold, the model achieves high ranking performance (ROC-AUC about 0.99) and negligible false-negative rates. Absolute Bloch drift strongly discriminates anomalies (ROC-AUC at least about 0.9), while consecutive drift is near random (ROC-AUC about 0.5), showing that detection arises from persistent state-space displacement rather than local fluctuations. Through the geometry of reduced single-qubit states and associated quantum Fisher information, these results show that learning-induced spectral organization appears as measurable quantum-state structure, establishing a unified spectral-geometric framework for diagnosing quantum learning systems with bosonic and Bloch probes.
トルコ語とアラビア語におけるヘイトスピーチの検出: 包括的な研究
オンラインのヘイトスピーチは、銃乱射事件、リンチ、民族浄化などの事件を含む、少数派に対する暴力の世界的な増加と関連している。この問題に取り組んでいる社会、特にヘイトスピーチが宗教、人種、民族、文化、国籍、移民ステータスに基づいて特定のグループをターゲットにしている場合、表現の自由と、広く使用されているオンラインプラットフォーム上で効果的なコンテンツモデレーションの必要性とのバランスをとるという課題に直面しています。この課題に応えて、私たちは、難民、イスラエル・パレスチナ紛争、トルコにおける反ギリシャ感情、民族または宗教コミュニティ(アレビ人、アルメニア人、アラブ人、ユダヤ人、クルド人)、LGBTI+という5つの異なるトピックをトルコ語でカバーし、アラビア語の1つのトピック(難民)をカバーする包括的なヘイトスピーチデータセットを導入します。さらに、ヘイト カテゴリ分類、ヘイト強度予測、ターゲット特定、ヘイト スピーチ スパン検出などのヘイト スピーチ分析の複数の側面に対処する最先端の BERT ベースのモデルを開発し、オンライン談話におけるヘイト コンテンツの包括的な理解を可能にします。
原文 (English)
Hate Speech Detection in Turkish and Arabic: A Comprehensive Study
Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech targets specific groups based on religion, race, ethnicity, culture, nationality, or migration status, face the challenge of balancing freedom of expression with the need for effective content moderation on widely used online platforms. In response to this challenge, we introduce a comprehensive hate speech dataset covering five distinct topics in Turkish: refugees, the Israel-Palestine conflict, anti-Greek sentiment in Turkey, ethnic or religious communities (Alevis, Armenians, Arabs, Jews, and Kurds), and LGBTI+, alongside one topic in Arabic (refugees). In addition, we develop state-of-the-art BERT-based models to address multiple dimensions of hate speech analysis, including hate category classification, hate intensity prediction, target identification, and hate speech span detection, enabling a comprehensive understanding of hateful content in online discourse.
大規模な 2 タワー検索のための LLM ベースのクラスタリングによるリアルタイムのハード ネガティブ サンプリング
2 タワー モデルは、大規模なレコメンデーション システム、特に検索段階で広く使用されています。 2 タワー モデルをトレーニングするための業界標準には、通常、バッチ内および/またはバッチ外のネガティブ サンプリングが含まれます。ただし、これらの方法では、モデルがすぐに学習できる簡単な否定的な結果が生成されることが多く、モデルに十分な挑戦を与えることができません。この問題に対処するために、大規模言語モデル (LLM) を利用してモデルのトレーニング中に同じクラスターからハード ネガティブを生成する、新しい自己教師ありハード ネガティブ サンプリング手法が提案されています。 LLM を利用してメディア表現を学習することにより、提案されたアプローチは、生成されるネガがより挑戦的で有益なものになることを保証します。このリアルタイム サンプリング フレームワークは、運用モデルにシームレスに統合できるように設計されており、最小限の計算複雑さで数十億のトレーニング データ ポイントを処理できます。公開データセットでの実験と大規模オンライン システムへの展開により、提案されたネガティブ サンプリング手法が業界で広く使用されている手法よりも優れていることが実証されました。さらに、産業用途での分析により、このサンプリング方法が推奨事項に固有のフィードバック ループを断ち切り、人気の偏りを大幅に軽減できることが明らかになりました。
原文 (English)
Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval
The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage. Industry standards for training two-tower models typically involve in-batch and/or out-of-batch negative sampling. However, these methods often produce easy negatives that models can quickly learn, failing to sufficiently challenge the model. To address this issue, a novel self-supervised hard negative sampling technique is proposed that leverages a large language model (LLM) to generate hard negatives from the same cluster during model training. By utilizing the LLM to learn media representations, the proposed approach ensures that the generated negatives are more challenging and informative. This real-time sampling framework is designed for seamless integration into production models, capable of handling billions of training data points with minimal computational complexity. Experiments on public datasets, along with deployment to a large-scale online system, demonstrate that the proposed negative sampling technique outperforms widely used industry methods. Furthermore, analysis in industrial applications reveals that this sampling method can help break inherent feedback loops in recommendations and significantly reduce popularity bias.
幻の文献: 一流の学会での査読に生き残る幻覚引用
大規模な言語モデルでは、裏付けのない主張を含む洗練された科学文書が生成され、幻覚がアーカイブ記録に残る可能性があります。技術的な記述によってこのリスクを評価することは難しく、多くの場合専門家の判断が必要ですが、引用はより監査可能な表面を提供します。つまり、参照は、互換性のある著者を持つ実際の学術著作物であるか、そうでないかのどちらかです。私たちは、査読手続きにおける引用幻覚を、存在しない著作物や著者と著者リストの大幅な不一致など、アイデンティティレベルの失敗に限定した保守的な定義を使用して測定します。私たちは、通常の書誌的差異(例: 開催地/年の違い、出版状況の更新、マイナーな名前のバリエーションなど)を明示的に除外します。引用を大規模に監査するために、複数の書誌ソースに対して書誌エントリを解決し、未解決のケースを Web 検索の再検証にエスカレートする検証パイプラインである RefChecker を構築します。 ICLR、ICML、NeurIPS、および USENIX Security から承認されたカメラ対応ペーパーに RefChecker を適用します。幻覚を起こした引用がアーカイブ記録に残っています。参考文献レベルの割合は通常 1% 未満ですが、論文レベルの失敗が目に見えるほど議事録は大きくなっています。2025 年には、NeurIPS および USENIX セキュリティ論文のおよそ 20 件に 1 件に、私たちの厳密な定義に基づくと、幻覚を起こしている可能性のある学術論文のような参考文献が少なくとも 2 件含まれています。また、ChatGPT 後の増加もいくつかの会場で観察されており、これには、単一の参考文献で 5 件以上の失敗を伴う論文の尾部や、受賞論文の中でも幻覚引用の可能性が高いことが含まれます。これらの結果は、査読だけでは引用の完全性を確実に強制することはできないが、監査は扱いやすい(会場規模の 1 回のスキャンで論文あたり約 0.04 ドル)ことを示唆しています。私たちは、出版前に日常的で再現可能な引用検証を行うための RefChecker をオープンソースにしています (https://github.com/markrussinovich/refchecker)。
原文 (English)
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival record. Assessing this risk via technical statements is difficult and often requires expert judgment, but citations provide a more auditable surface: a reference either resolves to a real scholarly work with compatible authorship, or it does not. We measure citation hallucination in peer-reviewed proceedings using a conservative definition limited to identity-level failures: non-existent works and substantial author-list mismatches. We explicitly exclude ordinary bibliographic drift (e.g., venue/year differences, publication-status updates, minor name variants). To audit citations at scale, we build RefChecker, a verification pipeline that resolves bibliography entries against multiple bibliographic sources and escalates unresolved cases to web-search re-verification. We apply RefChecker to accepted camera-ready papers from ICLR, ICML, NeurIPS, and USENIX Security. Hallucinated citations have entered the archival record. While reference-level rates are usually below 1%, proceedings are large enough that paper-level failures are visible: in 2025, roughly one in twenty NeurIPS and USENIX Security papers contains at least two likely hallucinated academic-paper-like references under our strict definition. We also observe post-ChatGPT increases in several venues, including a tail of papers with 5+ failures in a single bibliography, and likely hallucinated citations even among award-winning papers. These results suggest peer review alone does not reliably enforce citation integrity, yet auditing is tractable (about 0.04$ per paper in one venue-scale scan). We open-source RefChecker for routine, reproducible citation verification before publication (https://github.com/markrussinovich/refchecker).
Pano2World: 統合されたマルチビュー シーケンスによるエンドツーエンドの 3D 生成
単一のパノラマは、1 つのカメラ中心から視覚領域全体をキャプチャしますが、ユーザーは実際のシーンの探索を可能にすることなく、その場で見回すことに限定されます。単一のパノラマを、自由視点のナビゲーションのために永続的でレンダリング可能な 3D 表現に変換することへの関心が高まっています。既存の手法では、修復結果を伝播して基礎となるジオメトリを更新するビューごとの反復補完を採用するか、漸進的エラーの蓄積と煩雑なマルチステップ パイプラインを引き起こすか、ビデオ生成モデルの時間的一貫性事前分布を利用するかのいずれかですが、そのようなモデルに固有の連続軌道制約により、複数の方向からシーンを同時にカバーする際の柔軟性が制限されます。私たちは、単一の屋内パノラマを入力として受け取り、永続的で探索可能な 3D ガウス シーンを直接出力する Pano2World を紹介します。ソース パノラマが与えられると、Pano2World はまず粗い 3D ガウス プロキシを再構築し、適応的にサンプリングされた近くのポーズでレンダリングして、幾何学的に位置合わせされたガイダンス パノラマを取得します。次に、パノラマ拡散モデルは、ビューアウェア アテンション ルーティングを介してすべてのターゲット ビューを共同でノイズ除去します。各ターゲット ビューは、対応するガイダンス パノラマからの幾何学的制約と、ソース パノラマからのグローバル セマンティック ガイダンスを同時に受け取り、ビュー間の一貫性を自然に強化します。結合ノイズ除去中に形成されたマルチビューの隠れた特徴を VAE を介してピクセル ドメインにデコードすることで発生する情報損失を回避するために、これらの隠れた特徴をシーン潜在に直接蒸留し、その後最終的な 3D ガウス シーンにデコードするジオメトリ対応ブリッジ モジュールである潜在特徴アダプターを導入します。実験では、Pano2World が、マルチポジション パノラマ ノベルビュー合成ベンチマークにおいて、既存の方法よりも大幅に優れていることが実証されています。
原文 (English)
Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences
A single panorama captures the full visual sphere from one camera center, yet confines users to looking around in place without enabling true scene exploration. Converting a single panorama into a persistent, renderable 3D representation for free-viewpoint navigation has attracted growing interest; existing methods either adopt iterative per-view completion that propagates inpainting results to update the underlying geometry, leading to progressive error accumulation and cumbersome multi-step pipelines, or leverage the temporal consistency priors of video generation models, yet the continuous-trajectory constraint intrinsic to such models limits their flexibility in covering scenes from multiple directions simultaneously. We present Pano2World, which takes a single indoor panorama as input and directly outputs a persistent, explorable 3D Gaussian scene. Given the source panorama, Pano2World first reconstructs a coarse 3D Gaussian proxy and renders it at adaptively sampled nearby poses to obtain geometrically aligned guidance panoramas; a panoramic diffusion model then jointly denoises all target views via View-Aware Attention Routing, where each target view simultaneously receives geometric constraints from its corresponding guidance panorama and global semantic guidance from the source panorama, naturally enforcing cross-view consistency. To avoid the information loss incurred by decoding the multi-view hidden features formed during joint denoising back to the pixel domain via VAE, we introduce Latent Feature Adapter, a geometry-aware bridge module that directly distills these hidden features into a scene latent, subsequently decoded into the final 3D Gaussian scene. Experiments demonstrate that Pano2World significantly outperforms existing methods on the multi-position panoramic novel-view synthesis benchmark.
ワールド モデルからワールド アクション モデルへ: ロボット工学の簡潔なチュートリアル
世界モデルは、身体化されたインテリジェンスや生成シミュレーションでますます使用されていますが、その範囲はコミュニティ全体で依然として曖昧です。このチュートリアルでは、タスク関連の観測または状態の将来の進化を推定するアクション条件付き予測モデルとしてのワールド モデルの設計空間ビューを示します。私たちは既存の手法を観測空間世界モデルと状態空間世界モデルに分類し、視覚的な忠実度、空間構造、物理的解釈可能性、制御の使いやすさにおけるトレードオフを比較します。さらに、予測された未来と実行可能なロボットのアクションを結び付ける世界アクション モデルを紹介し、想像してから実行する、ビデオ機能条件付きアクション予測、ビデオとアクションの共同モデリング、およびポリシー学習のための補助ビデオ予測という 4 つの代表的なパラダイムを要約します。このチュートリアルの目標は、世界 (アクション) モデルの概念的範囲を明確にし、具体化された予測と制御のための構造化された分類法を提供することです。
原文 (English)
From World Models to World Action Models: A Concise Tutorial for Robotics
World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities. This tutorial presents a design-space view of world models as action-conditioned predictive models that estimate the future evolution of task-relevant observations or states. We categorize existing methods into observation-space and state-space world models, comparing their trade-offs in visual fidelity, spatial structure, physical interpretability, and control usability. We further introduce world action models, which connect predicted futures with executable robot actions, and summarize four representative paradigms: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and auxiliary video prediction for policy learning. The goal of this tutorial is to clarify the conceptual scope of world (action) models and provide a structured taxonomy for embodied prediction and control.
拡散変圧器のトレーニング後の枝刈り
拡散変換器 (DiT) は、画像生成において優れたパフォーマンスを示していますが、かなりの計算オーバーヘッドとリソース消費に悩まされています。トレーニング後の枝刈りは有望な解決策を提供します。ただし、DiT の独自のアーキテクチャ設計とパラメータ分布により、従来のプルーニング手法は適用できず、大幅なパフォーマンスの低下につながります。具体的には、LLM 用に開発された、一連の近似によってメトリクスを導出する従来の方法では、顕著性メトリクスにおける重みの相対的な寄与が増幅されます。さらに、DiT の重みは、LLM の重みよりも大幅に大きい値を示します。さらに、既存の枝刈り粒度では、モデル構造の変化が見落とされます。この論文では、カスタマイズされた顕著性基準と枝刈り粒度を導入することで枝刈りパフォーマンスを向上させる DiT-Pruning を提案します。私たちは、エネルギーベースの観点から重みと活性化の寄与のバランスを取る新しい指標を設計し、重要な要素をより効果的に特定できるようにします。さらに、2 次元の重み空間で明確なクラスタリング パターンが観察されます。したがって、クラスタリングを意識したプルーニング粒度を採用し、効果的なスパース割り当てを可能にします。さまざまな DiT に関する広範な評価により、特に高いスパース性の下で、私たちの方法が一貫して画質を維持することが示されています。 MJHQ 上の 512x512 解像度の FLUX.1-dev の場合、DiT-Pruning は 50% のスパース性で CLIP スコアの損失がわずか 0.001 であり、最近のプルーニング手法を劇的に上回っています。
原文 (English)
Post-Training Pruning for Diffusion Transformers
Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhead and resource consumption. Post-training pruning offers a promising solution; however, due to DiTs' unique architectural design and parameter distribution, traditional pruning methods are inapplicable, leading to significant performance degradation. Specifically, prior methods developed for LLMs, which derive metrics through a series of approximations, amplify the relative contribution of weights in the saliency metric. In addition, weights in DiTs exhibit significantly larger magnitudes than those in LLMs. Moreover, existing pruning granularity overlooks variations in model structures. In this paper, we propose DiT-Pruning, which improves pruning performance by introducing customized saliency criteria and pruning granularity. We design a novel metric that balances the contributions of weights and activations from an energy-based perspective, enabling more effective identification of important elements. Furthermore, we observe distinct clustering patterns in the two-dimensional weight space. Accordingly, we adopt a clustering-aware pruning granularity, enabling effective sparse allocation. Extensive evaluations on various DiTs show that our method consistently preserves image quality, especially under high sparsity. For FLUX.1-dev at 512x512 resolution on MJHQ, DiT-Pruning achieves only a 0.001 loss in CLIP score at 50% sparsity, dramatically outperforming recent pruning methods.
暗黙的な神経表現のための心臓運動事前分布の学習
Implicit Neural Representation (INR) は心臓の運動推定に適しており、運動フィールドの連続的でコンパクトな表現を提供します。ただし、INR を各画像シーケンスに適合させるのは時間がかかり、最適化の軌道に左右されます。学習された事前分布は、最適化を妥当な運動フィールドに向けて導き、より迅速な適応を可能にするのに役立ちますが、心臓の運動 INR の学習事前分布はまだ研究が進んでいません。この研究では、関節最適化によって学習された母集団事前学習、重み平均化によって取得されたコンセンサス事前学習、自動デコーダー、およびメタ学習を含む、心臓運動事前学習のための 4 つの戦略を比較します。英国バイオバンクからの短軸タグ付き心臓磁気共鳴画像を使用して、追跡精度、運動挙動、および適応軌道への影響を評価します。すべての学習された事前確率は、ランダムな初期化と比較して、早期適応パフォーマンスを大幅に向上させました。事前の単純なコンセンサスは効果的でしたが、自動デコーダは初期の適応中に大きな変形をより速く回復しました。メタ学習は初期に強力なパフォーマンスを達成し、50 回の反復にわたって最良の適応軌道を維持しました。
原文 (English)
Learning Cardiac Motion Priors for Implicit Neural Representations
Implicit neural representations (INRs) are well suited to cardiac motion estimation, providing continuous, compact representations of motion fields. However, fitting an INR to each image sequence is time-consuming and sensitive to the optimisation trajectory. Learned priors can help guide optimisation towards plausible motion fields and enable faster adaptation, but learning priors for cardiac motion INRs remains under-explored. In this work, we compare four strategies for learning cardiac motion priors, including a population prior learned by joint optimisation, a consensus prior obtained by weight averaging, auto-decoders, and meta-learning. Using short-axis tagged cardiac magnetic resonance images from the UK Biobank, we evaluate their impact on tracking accuracy, motion behaviour, and adaptation trajectory. All learned priors substantially improved early adaptation performance compared with random initialisation. While the simple consensus prior was effective, auto-decoders recovered large deformations faster during early adaptation. Meta-learning achieved strong early performance and maintained the best adaptation trajectory over 50 iterations.
安価なコード、コストのかかる判断: ガバナンス可能なエージェント ソフトウェア エンジニアリングに関するケース スタディ
生成 AI は、ソフトウェア エンジニアリングを、少ない実装労力を中心に組織された実践から、豊富で低コストのコード生成を中心に組織された実践へと移行させています。この変化は、エンジニアリングの中心となる問題を変えます。AI が有用なコードを生成できるかどうかではなく、AI を介した開発が検査可能、修正可能、保守可能であり続けるように、エンジニアがアーキテクチャ、ツール、証拠、フィードバック ループをどのように編成するかです。私たちは、一人称のケーススタディを通じてこの問題を研究します。この開発作業では、1 人の専門ソフトウェア エンジニアがフロンティア AI コーディング エージェントを使用してドキュメント アクセシビリティ修復システムを構築しました。12 週間の開発作業です。経験的な記録は、88 件の同時期のフィールド ノート、420 KLOC の製品コード、および 1.16 MLOC のテスト、lint、サポート ドキュメント、およびエージェント ツールで構成されています。この記録から、高速エージェントの実装がどのようにしてガバナンス可能になるかを説明するプロセス モデルとして表現される、ガバナンス変換の中範囲理論の候補を開発します。このモデルは、エージェント実装の速度が繰り返し起こる構造的故障クラスをどのように表面化するか、そしてエンジニアリングの判断がそれらの故障を耐久性のあるガバナンス メカニズムに変換することでどのように速度を維持するかを説明します。既知の義務から制御を導き出す既存のガバナンス モデルとは対照的に、ガバナンス変換では、エージェントの作業中にのみ表示される障害から制御がどのように発見されるかを説明します。私たちはモデルを使用してテスト可能な予測を行い、ソフトウェア エンジニアリングの研究と実践への影響を説明します。
原文 (English)
Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around abundant, low-cost code production. This shift changes the central engineering problem: not whether AI can generate useful code, but how engineers organize architectures, tools, evidence, and feedback loops so that AI-mediated development remains inspectable, correctable, and maintainable. We study this problem through a first-person case study: a 12-week development effort in which a single expert software engineer used frontier AI coding agents to build a document accessibility remediation system. The empirical record comprises 88 contemporaneous field notes, 420 KLOC of production code, and 1.16 MLOC of tests, lints, supporting documentation, and agent tooling. From this record, we develop a candidate middle-range theory of governance conversion, expressed as a process model explaining how high-velocity agentic implementation becomes governable. The model explains how agentic implementation velocity surfaces recurring structural failure classes, and how engineering judgment sustains velocity by converting those failures into durable governance mechanisms. In contrast to existing governance models that derive controls from known obligations, governance conversion explains how controls are discovered from failures that become visible only during agentic work. We use our model to make testable predictions and to describe implications for software engineering research and practice.
Diffusion-GR2: 拡散生成推論リランカー
生成推論の再ランカーは、候補リストを並べ替える前に思考の連鎖を発行することで強力な推奨精度を実現しますが、推論が遅くなります。自己回帰 (AR) デコーダは推論トークンごとに 1 回の連続した前方パスを費やし、推論トレースは生成されるランキングをはるかに上回ります。このコストを削減するために、ブロック拡散言語モデルは、いくつかのノイズ除去ステップで多くの位置を並行してデコードし、大幅に高速になりますが、単純に AR リランカーを 1 つに変換すると、2 つの精度ギャップが生じます。 (1) 構造的なギャップ: 回答位置は並行してノイズ除去され、独立してスコアリングされるため、デコーダーは無効なランキング (重複、欠落、またはセット外の識別子) を生成しますが、AR はこれを左から右のマスキングによって回避します。 (2) 分布ギャップ: 固定教師軌道上で変換されたモデルを微調整することは、推論時の独自のデコードと比較してポリシーから外れており、精度ギャップが残ります。高速化を維持しながら両方のギャップを埋めるために、AR 推論リランカー (GR2) をブロック拡散リランカーに変換するレシピである \textbf{Diffusion-GR2} を提案します。まず、変換微調整 (CFT) は、AR で初期化された拡散モデルを適応させて、外部の制約付きデコーダーを使用せずに、独自に答えを有効な置換にノイズ除去します。次に、オンポリシー蒸留 (OPD) が、AR 教師からの高密度のトークンごとのターゲットを使用して、独自のデコードされた軌道でモデルを監視します。最後に、OPD のポリシーに関するポリシーに加えて、再ランキング報酬に対して強化学習 (RL) ステージを適用します。 Amazon Beauty での実験では、Diffusion-GR2 が AR リランカーとほぼ同等に回復し、ブロック並列デコードにより、モデルの推論出力長でデコード スループットが $2.4$ ~ $3.5\times$ 向上することが実証されました。アブレーションにより、CFT がコンバージョン ギャップのほとんどを回復し、ポリシーに基づいた蒸留により AR リファレンスにさらに近づくことが示されています。
原文 (English)
Diffusion-GR2: Diffusion Generative Reasoning Re-ranker
Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trace far exceeds the ranking it produces. To reduce this cost, block-diffusion language models decode many positions in parallel over a few denoising steps and are substantially faster, yet naively converting an AR re-ranker into one opens two accuracy gaps: (1) a structural gap: answer positions are denoised in parallel and scored independently, so the decoder emits invalid rankings (duplicated, dropped, or out-of-set identifiers) that AR avoids through left-to-right masking; and (2) a distributional gap: fine-tuning the converted model on fixed teacher trajectories is off-policy relative to its own decoding at inference, leaving a residual accuracy gap. To close both gaps while keeping the speedup, we propose \textbf{Diffusion-GR2}, a recipe that converts our AR reasoning re-ranker (GR2) into a block-diffusion re-ranker. First, conversion fine-tuning (CFT) adapts the AR-initialized diffusion model to denoise the answer into a valid permutation on its own, without an external constrained decoder. Next, on-policy distillation (OPD) then supervises the model on its own decoded trajectories with dense per-token targets from the AR teacher. Finally, we apply a reinforcement-learning (RL) stage against a re-ranking reward on top of OPD's on-policy policy. Experiments on Amazon Beauty demonstrate that Diffusion-GR2 recovers to near-parity with the AR re-ranker, while block-parallel decoding raises decode throughput by $2.4$--$3.5\times$ at the model's reasoning output length. Ablations show that CFT recovers most of the conversion gap, and that on-policy distillation further closes it to the AR reference.
Kara: スライディング ウィンドウ KV キャッシュ圧縮を介した効率的な推論 LLM サービス
推論言語モデルでは長い思考連鎖 (CoT) が生成されることが多く、これによりデコード段階で大規模な KV キャッシュが蓄積され、デコードの待ち時間が長くなり、スループットが制限されます。これらの問題に対処するために、KV キャッシュ圧縮は、後続のデコードに有用な KV ペアを保持しながら、重要でない KV ペアを選択的に削除することでメモリのオーバーヘッドを削減する有望な技術として浮上しました。それにもかかわらず、既存の KV キャッシュ圧縮方法には 2 つの重要な制限があることがわかりました。1) しきい値でトリガーされる圧縮ポリシーでは、スループットの改善が限定的か、さらにはスループットが低下する可能性があり、シーケンスの特定のブロックから KV ペアが完全に削除される可能性があり、情報損失が悪化する可能性があります。 2) 通常、分離された KV ペアまたは厳格な境界を持つ固定サイズのチャンクのいずれかを保持し、重要な柔軟なサイズのチャンクを任意のトークン位置に保持できません。これらの制限を克服するために、最近生成されたコンテキストのみを操作してデコード時の圧縮を実行するスライディング ウィンドウ KV キャッシュ圧縮方法である Kara を提案します。 Kara は、双方向の注意を活用してウィンドウ内で有益な KV ペアをスコア化し、選択します。重要なセマンティック情報の柔軟な保存を可能にするために、選択した KV ペアのサブセットをチャンクに拡張する Token2Chunk モジュールを設計します。さらに、Kara を PagedAttendance に適応させ、vLLM に基づいて構築された推論フレームワークである KvLLM を開発します。これにより、KV キャッシュ メモリの使用量が削減され、出力スループットが効果的に向上します。広範な実験により、提案された Kara と KvLLM の一貫したパフォーマンスの向上が実証されました。
原文 (English)
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression
Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and incurs high decoding latency and limited throughput. To address these issues, KV cache compression has emerged as a promising technique for reducing memory overhead by selectively removing unimportant KV pairs while preserving useful ones for subsequent decoding. Nevertheless, we identify two key limitations in existing KV cache compression methods: 1) their threshold-triggered compression policy may provide limited throughput improvement or even reduce throughput, and may fully eliminate KV pairs from certain blocks of the sequence, potentially worsening information loss. 2) they typically retain either isolated KV pairs or fixed-size chunks with rigid boundaries, failing to preserve important flexible-sized chunks at arbitrary token positions. To overcome these limitations, we propose Kara, a sliding-window KV cache compression method that performs decoding-time compression by operating only on the recently generated context. Kara leverages bidirectional attention to score and select informative KV pairs in the window. To enable flexible preservation of important semantic information, we design a Token2Chunk module to expand a subset of selected KV pairs into chunks. Furthermore, we adapt Kara to PagedAttention and develop KvLLM, an inference framework built upon vLLM, which reduces KV cache memory usage and effectively improves output throughput. Extensive experiments demonstrate consistent performance improvements of proposed Kara and KvLLM.
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluatio…
Spin-Weighted Spherical Harmonics Enable Complete and Scalable $\mathrm{E}(3)$-Equivariant Networks
$\mathrm{E}(3)$-equivariant networks are promising for 3D atomistic system modeling, yet their scalability is limited by the $O(L^6)$ compl…
MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for use…
DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents
Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns requir…
Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander
We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone. Choo…
AI Virtue: What is "Good" Knowledge in the Age of Artificial Intelligence?
In the age of AI, what will be good knowledge? This article, which is accepted and forthcoming in a special issue of Modern Fiction Studies…
kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail
Large language models (LLMs) are increasingly deployed in domains requiring guardrails to detect unsafe, off-topic, or adversarial prompts.…
Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies
Flow-matching vision-language-action policies generate robot action chunks through an iterative transport process, creating an opportunity…
ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning
We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uni…
GAP-GDRNet: Geometry-Aware Monocular Visual Pose Sensing on a Single-Target Synthetic Spacecraft Dataset
Monocular relative pose sensing is a central perception problem in non-cooperative rendezvous and on-orbit servicing. In spacecraft images,…
Understanding Agent-Based Patching of Compiler Missed Optimizations
Compiler missed optimizations refer to cases in which compilers failed to optimize certain code. It takes many compiler developers' efforts…
DemoPSD: Disagreement-Modulated Policy Self-Distillation
On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single mo…