AIニュース 2026-06-29
自動生成: 2026-06-29 13:26 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
HP Inc. launches Frontier strategic partnership with OpenAIOpenAI
HP Inc. scales its OpenAI Frontier partnership to deploy AI across cu…
-
製造現場のトラブル解消を「AI工場長」が支援? 「エージェント型工場」とはITmedia AI+
AccentureとAvanadeはMicrosoftと協働し、製造業向け工場インテリジェンスシステム「エージェント型工場」を開発したと発…
-
NormAct: 身体化された計画における隠れた社会規範遵守のベンチマークarXiv cs.AI
マルチモーダル大規模言語モデル (MLLM) は、自己中心的な環境で具体化されたプランナーとして導入されることが増えています。タスクを成功…
-
ロボットの模倣学習を60時間→4.8時間に AWSのGPUでフィジカルAI開発を加速 ファナックITmedia AI+
ファナックは「AWS Summit Japan 2026」の基調講演で、ロボットに動作を教える「模倣学習」の時間を、AWSのGPU活用で6…
-
カインズが画像AIで売上UP模索、店頭でのインテリア“試着”をテスト 立ちはだかる「正確性と効率」の壁ITmedia AI+
カインズが画像生成AIを活用し、部屋のインテリアを疑似的に置き換えられる店頭サイネージ「CAINZ Fitting Room」を開発。その…
-
AIは設計者を置き換えるのか Autodesk幹部に聞くCADと設計データの未来ITmedia AI+
AIの活用が設計/製造の現場にも広がる中、CADの操作や設計者の役割はどう変わるのか。米Autodesk 製品開発/製造ソリューション担当…
-
解剖・孫正義氏の「ガチョウ論」 「ソフトバンクG株価が低過ぎ」主張を信じてよいのかITmedia AI+
孫正義氏が、ソフトバンクGの株価に不満をにじませた。孫氏が「本当の企業価値」として示す「時価純資産」の論理と危うさを解説する。
トピック別件数
日本語メディア5件
ITmedia AI+ (日本語)
ロボットの模倣学習を60時間→4.8時間に AWSのGPUでフィジカルAI開発を加速 ファナック
ファナックは「AWS Summit Japan 2026」の基調講演で、ロボットに動作を教える「模倣学習」の時間を、AWSのGPU活用で60時間から4.8時間へ短縮したと示した。
カインズが画像AIで売上UP模索、店頭でのインテリア“試着”をテスト 立ちはだかる「正確性と効率」の壁
カインズが画像生成AIを活用し、部屋のインテリアを疑似的に置き換えられる店頭サイネージ「CAINZ Fitting Room」を開発。その効果や利便性を検証している。
製造現場のトラブル解消を「AI工場長」が支援? 「エージェント型工場」とは
AccentureとAvanadeはMicrosoftと協働し、製造業向け工場インテリジェンスシステム「エージェント型工場」を開発したと発表した。
AIは設計者を置き換えるのか Autodesk幹部に聞くCADと設計データの未来
AIの活用が設計/製造の現場にも広がる中、CADの操作や設計者の役割はどう変わるのか。米Autodesk 製品開発/製造ソリューション担当エグゼクティブバイスプレジデントのジェフ・キンダー氏に、AIが設計業務にもたらす変化、AI時代に求められる設計データの在り方、そして同社が描…
解剖・孫正義氏の「ガチョウ論」 「ソフトバンクG株価が低過ぎ」主張を信じてよいのか
孫正義氏が、ソフトバンクGの株価に不満をにじませた。孫氏が「本当の企業価値」として示す「時価純資産」の論理と危うさを解説する。
海外メディア2件
TechCrunch AI (英語)
Ford rehires ‘gray beard’ engineers after AI falls short
"Mistakenly we thought that by just introducing artificial intelligence ... that would produce a high-quality product.”
Why Wall Street thinks US memory maker Micron is the next Nvidia
Eager to find more public AI-related companies that may do as well as Nvidia, Wall Street investors think they've found a winner with Micro…
公式ブログ1件
OpenAI (英語)
HP Inc. launches Frontier strategic partnership with OpenAI
HP Inc. scales its OpenAI Frontier partnership to deploy AI across customer experiences, software development, and enterprise operations.
論文233件
arXiv cs.AI (英語)
AI モデル ネットワークの概念、現状、将来
コンピュータの主な機能は計算と処理ですが、インターネットの核となる価値は共有とコラボレーションに根ざしています。コンピューターはインターネットを生み出し、インターネットはコンピューターの価値を高めます。インターネット、クラウド コンピューティング、ビッグ データの急速な発展により、人工知能はラージ モデル (LM) の時代に突入しています。ただし、LM の実用化は現在、高いトレーニング コストと展開の複雑さによって妨げられており、軽量でプライベートなドメイン固有のモデルへの移行が進んでいます。異種モデルの急速な普及と広範な配布に伴い、モデル間の効果的な相互作用とコラボレーションを可能にすることが、LM 開発において緊急に対処する必要がある重大なボトルネックとして浮上しています。この論文は、インターネットの発展からインスピレーションを得て、ワールドワイド AI モデル ネットワーク (AI-ModelNet) の概念、ビジョン、およびシステム アーキテクチャを提案します。これは、モデル間の経路を確立することにより、相互接続、機能共有、および共同推論を実現する新しいパラダイムです。まず、単一モデル研究と複数モデル研究の現状を簡単にレビューします。続いて、AI-ModelNet の体系的なビジョンと階層アーキテクチャが明確化され、プロトタイプ システムとさまざまなアプリケーション ケースを通じてフレームワークの実現可能性が検証されます。最後に、将来の研究の重要な方向性について予備的に説明します。
原文 (English)
AI-Model Network: Concept, Current State and Future
While the primary function of computers lies in computation and processing, the core value of the Internet is rooted in sharing and collaboration. Computers create the Internet, and the Internet empowers the value of computers. The rapid development of the Internet, cloud computing, and big data is pushing artificial intelligence into the era of large models (LMs). However, the practical application of LMs is currently hindered by high training costs and deployment complexities, driving a shift toward lightweight, private, and domain-specific models. With the rapid proliferation and wide distribution of heterogeneous models, enabling effective interaction and collaboration among them has emerged as a critical bottleneck that urgently needs to be addressed in LM development. Drawing inspiration from the development of the Internet, this paper proposes the concept, vision, and system architecture of world wide AI-model network (AI-ModelNet). It is a novel paradigm that achieves interconnection, capability sharing, and collaborative reasoning by establishing pathways between models. We first briefly review the current state of single-model and multi-model research. Subsequently, the systemic vision and hierarchical architecture of AI-ModelNet are articulated, followed by validation of the framework's feasibility through a prototype system and diverse application cases. Finally, key directions for future research are discussed preliminarily.
マルチエージェント LLM チームにとってパーソナリティ構成が重要になるのはどのような場合ですか?
パーソナリティプロンプトは、大規模な言語モデルがどのようにコミュニケーションするかを形成しますが、これらの行動の変化が客観的なタスクの結果に影響を与えるかどうかは、まだ十分に調査されていません。これまでの研究では、低い同調度で促されたエージェントは敵対的な言葉を発し、高い同調度で促されたエージェントは協力的になることが示されているが、コミュニケーションスタイルとタスクパフォーマンスの関係は複数の領域にわたって体系的に調査されていない。この研究では、構造化コーディング、無制限の研究協力、競争的交渉という 3 つのタスク ドメインでフロンティア LLM 全体の性格特性を操作することにより、性格構成がマルチエージェント チームのパフォーマンスに重要であるかどうかを調査します。性格への影響はタスクの構造に大きく依存することがわかりました。コーディングタスクでは、協調性が低いとコミュニケーションに大きな変化が生じ、マイルストーンの完了にはほとんど影響しません。無制限のコラボレーションや交渉では、同じ操作がパフォーマンスを大幅に低下させます。マルチエージェントシステム設計への影響と人格操作の限界について説明します。
原文 (English)
When Does Personality Composition Matter for Multi-Agent LLM Teams?
Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-explored. Prior work shows that agents prompted with low agreeableness produce adversarial language, while those prompted with high agreeableness become cooperative, but the relationship between communication style and task performance has not been systematically examined across multiple domains. In this work, we investigate whether personality composition matters for multi-agent team performance by manipulating personality traits across frontier LLMs on three task domains: structured coding, open-ended research collaboration, and competitive bargaining. We find that personality effects depend critically on task structure. In coding tasks, low agreeableness leads to large communication shifts that have little effect on milestone completion. In open-ended collaboration and bargaining, the same manipulation substantially degrades performance. We discuss implications for multi-agent system design and the limits of personality manipulation.
未来の内面化: 世界モデル計画のための統合エージェント トレーニング パラダイム
大規模言語モデル (LLM) エージェントは、逐次的な意思決定において強力な能力を示していますが、長期的なタスクでは基本的に反応的なままです。コミットメントの前に潜在的な計画を評価するために「what-if」推論を使用する人間とは異なり、標準エージェントには将来の結果をシミュレートするための内部世界モデルがありません。したがって、単一の自己回帰モデルをトレーニングして、将来の状態のロールアウトと計画条件付きの成功推定 (Q 値のテキストの類似物) の両方を言語化することで、将来を意識した計画を内面化することを提案します。重要なのは、フォーマット能力のギャップを特定することです。トレーニング後の先読みトレースでエージェントを単に微調整するだけでは、真の予測根拠のない先見性の表面的な模倣につながります。このギャップを埋めるために、私たちは 3 段階のトレーニング パラダイムを導入します。(i) 潜在的な予測機能をポリシーに注入するワールド モデル エージェント中期トレーニング (WM-AMT)。 (ii) この挿入された機能を構造化するフォーマット抽出 SFT (FE-SFT)。 (iii) 予測条件付き強化学習 (FC-RL) により、生成されたシミュレーションのキャリブレーションとユーティリティを改良します。検索タスクと数学的推論タスクで評価すると、私たちのアプローチは他のトレーニング ベースラインを常に上回っています。私たちの結果は、LLM エージェントにおける効果的な内部世界モデリングには、根拠があり調整された先見性を達成するために、能力優先のトレーニング パイプラインが必要であることを示しています。
原文 (English)
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactive in long-horizon tasks. Unlike humans who employ "what-if" reasoning to evaluate potential plans before commitment, standard agents lack an internal world model to simulate future outcomes. Therefore, we propose to internalize future-aware planning by training a single autoregressive model to verbalize both a prospective state rollout and a plan-conditioned success estimate-a textual analogue of the Q-value. Crucially, we identify a format-capability gap: simply fine-tuning agents on look-ahead traces during post-training leads to superficial mimicry of foresight without genuine predictive grounding. To bridge this gap, we introduce a three-stage training paradigm: (i) World Model Agentic Mid-Training (WM-AMT) to inject latent predictive capabilities into the policy; (ii) Format-Eliciting SFT (FE-SFT) to structure this injected capability; and (iii) Foresight-Conditioned Reinforcement Learning (FC-RL) to refine the calibration and utility of the generated simulations. Evaluated on search and mathematical reasoning tasks, our approach consistently outperforms other training baselines. Our results demonstrate that effective internal world modeling in LLM agents requires a capability-first training pipeline to achieve grounded and calibrated foresight.
オデッセイ: 検証可能なローカル真実保存基盤モデルの構築
私たちは、ファウンドリの構成として、検証可能なローカル真実保持基盤モデルを構築するための ODYSSEY と呼ばれるカテゴリカル フレームワークを導入します。これは、ローカル コンテキスト、ローカル表現ファミリー、制限マップ、接着ルール、障害ポリシー、更新義務、および人間に面したビューのカバーを指定するビルディング ブロック アーキテクチャ コンポーネントです。ファウンドリは、その中に議論の要素を含む組織化された知識の束です。コンクリート鋳造工場は、証拠/議論、運営上の決定、制度的/財政的、市場の意味、科学的課題、研究プログラム、アシスタントビルド、および評価ハーネス鋳造工場などの一般的な鋳造工場から構築されます。 Universal Foundry Learning (UFL) は、左右の Kan 拡張機能の構成としてファウンドリの建設を形式化します。左の Kan 拡張機能はローカル アーティファクトを候補のファウンドリにローリングし、右の Kan 拡張機能はプロモーションに必要な制限、接着、妨害、および議論の条件を強化します。 Foundry SQL (FSQL) は、維持されているファウンドリ アーティファクトをスライスするための小さな型付きクエリ サーフェスであり、外部モデルまたは事前構築モデルを耐久性のある ODYSSEY 状態に許可するための TICKET (Topos Integration using Causal Kan Extension Transformers) 認定を使用します。 ODYSSEY は、広範囲のコンクリート鋳造工場で完全に実装およびテストされており、同じカテゴリカル機構がドメイン構築、アーティファクト再生、層診断、根拠のあるトゥールミン/ローカル LLM 精査、残存障害台帳、および異種ソースにわたる最適化された TICKET 互換の因果関係主張抽出をサポートしていることを示しています。このペーパーは、ICML 2026 で 2.5 時間のチュートリアルとして発表されます。チュートリアルのホームページは https://bit.ly/4ajS0nA にあります。
原文 (English)
Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models
We introduce a categorical framework called ODYSSEY for constructing verifiable, local truth-preserving foundation models as compositions of foundries: building-block architectural components that specify a cover of local contexts, local representation families, restriction maps, gluing rules, obstruction policies, update obligations, and human-facing views. A foundry is an organized sheaf of knowledge that carries within it an argumentation component. Concrete foundries are built from generic foundries such as evidence/argument, operational decision, institutional/financial, market meaning, scientific challenge, research-program, assistant-build, and evaluation-harness foundries. Universal Foundry Learning (UFL) formalizes foundry construction as a composition of left and right Kan extensions, with left Kan extension rolling local artifacts into candidate foundries and right Kan extension enforcing the restriction, gluing, obstruction, and argumentation conditions required for promotion. Foundry SQL (FSQL) is a small typed query surface for slicing maintained foundry artifacts that uses TICKET (Topos Integration using Causal Kan Extension Transformers) certification for admitting external or pre-built models into durable ODYSSEY state. ODYSSEY is fully implemented and tested across a wide spectrum of concrete foundries, showing that the same categorical machinery supports domain construction, artifact replay, sheaf diagnostics, grounded Toulmin/local-LLM scrutiny, residual-obstruction ledgers, and optimized TICKET-compatible causal-claim extraction across heterogeneous sources. This paper is to be presented as a 2.5 hour tutorial at ICML 2026. The tutorial home page is at https://bit.ly/4ajS0nA.
DysLexLens: オンライン フォーラムからのディスレクシア学習者の洞察を分析するための低リソース LLM フレームワーク
失読症の学習者は、読み書き、整理、学習関連のタスクをサポートするために人工知能 (AI) ツールをますます使用しています。しかし、これらのツールを使った彼らの実際の経験は、依然として十分に調査されていません。この論文では、オンライン フォーラムのディスカッションを通じて失読症の学習者の経験を AI で分析するために設計された、低リソース LLM フレームワークである DysLexLens を提案します。 DysLexLens は、ノイズの多いソーシャル メディア投稿を辞書主導のコーパスに変換し、ナレッジ グラフ (KG) ベースの質問推論を提供し、検証可能なクエリ応答を生成し、定量的かつ人間に基づいた評価を通じて応答評価を可能にする、エンドツーエンドの証拠追跡可能なアーキテクチャとして設計されています。 DysLexLens には 4 つの主要な機能があります。まず、辞書主導のフィルタリング手法を採用して、失読症と AI にさらに焦点を当てた Reddit コーパスを構築し、ノイズの多い投稿や関連性の低い投稿を除外して、リソースの少ないフォーラムのコンテキストから収集されたデータの関連性を高めます。 2 番目に、LLM 支援のセマンティック分析と KG ベースのクエリ推論を統合して、意味のあるパターンを明らかにします。 3 番目に、LLM によって生成された応答パフォーマンスを測定するための定量的な評価メトリクス (RAGAS およびクエリ ロバストネス) があります。第 4 に、幻覚と証拠の整合性に特に焦点を当て、応答の質を評価するための構造化された定性的検証ガイドラインを提供します。ディスレクシア関連の Reddit フォーラム データと 30 の質問を使用して、DysLexLens の有効性を実証します。結果は、他の低リソースのフォーラム データ コンテキストへの潜在的な一般化可能性を示しています。再現性をサポートするために、DysLexLens、サンプル データ、質問、評価結果は Github で入手できます。
原文 (English)
DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
Dyslexic learners increasingly use artificial intelligence (AI) tools to support reading, writing, organisation, and study-related tasks. However, their lived experiences with these tools remain largely underexamined. This paper proposes DysLexLens, a low-resource LLM framework, designed to analyse dyslexic learners experience with AI through online forum discussions. DysLexLens is designed as an end-to-end, evidence-traceable architecture which transforms noisy social media posts into a dictionary-driven corpora, provides knowledge-graph (KG)-based question reasoning, generates verifiable query responses, and enables response evaluation through quantitative and human-grounded assessment. DysLexLens has four key features. First, it employs a dictionary-driven filtering method to construct a more focused Reddit corpus on dyslexia and AI, filtering out noisy and weakly related posts to improve the relevance of data collected from low-resource forum contexts. Second, it integrates LLM-assisted semantic analysis with KG-based query reasoning to uncover meaningful patterns. Third, it has quantitative evaluation metrics (RAGAS and Query Robustness) to measure LLM-generated response performance. Fourth, it provides structured qualitative validation guidelines for assessing response quality, with a specific focus on hallucination and evidence alignment. We demonstrate the effectiveness of DysLexLens using dyslexia-related Reddit forum data and 30 questions. The results show its potential generalisability to other low-resource forum data contexts. DysLexLens, sample data, questions and evaluation results are available at Github to support reproducibility.
MER-R1: 低速思考と高速思考の相乗効果によるマルチモーダルな感情推論
明示的推論は、予測をより解釈しやすくするとしても、必ずしもマルチモーダル感情認識 (MER) の精度が向上するとは限りません。特に、推論ベースの MLLM の場合、直接的な答えを引き出すことによる高速思考は、熟議推論後の低速思考よりも優れていることがよくあります。私たちの実証分析によると、速い思考はより広範囲でより信頼性の高い予測により再現率を向上させますが、遅い思考は誤ったカテゴリの保守的なフィルタリングを通じて精度を向上させます。これらの洞察に基づいて、低速と高速の相補性を明示的な最適化に変える強化学習フレームワークである MER-R1 を提案します。二重目的の解きほぐしは、再現率と精度を 2 つの最適化信号に分離し、相互にトレードオフするのではなく、一緒に最適化できるようにします。低速と高速の信頼度調整により、最終的なゆっくりとした思考の答えと高速な思考の直観がさらに調整され、間違った感情を抑制しながら正しい感情が強化されます。このようにして、MER-R1 は、速い思考の想起指向の直観と、ゆっくりとした思考の精度指向の選択性を統合します。さらに、この相乗効果の理論的根拠を提供し、それが最適化中の分散に起因する干渉を軽減することを示します。 MER-UniBench と MME-Emotion に関する広範な実験により、MER-R1 が最先端のパフォーマンスを達成し、推論が感情認識に真の利益をもたらすことが示されました。
原文 (English)
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predictions more interpretable. Specifically, for reasoning-based MLLMs, fast thinking by triggering direct answers often outperforms slow thinking after deliberative reasoning. Our empirical analyses show that fast thinking improves recall with broader and more confident predictions, whereas slow thinking favors precision through conservative filtering of incorrect categories. Building on these insights, we propose MER-R1, a reinforcement learning framework that turns slow-fast complementarity into explicit optimization. Dual-objective disentanglement separates recall and precision into two optimization signals, allowing them to be jointly optimized rather than traded off against each other. Slow-fast confidence calibration further aligns the final slow-thinking answer with fast-thinking intuition, strengthening correct emotions while suppressing incorrect ones. In this way, MER-R1 unifies the recall-oriented intuition of fast thinking with the precision-oriented selectivity of slow thinking. We further provide theoretical justification for this synergy, showing that it mitigates variance-induced interference during optimization. Extensive experiments on MER-UniBench and MME-Emotion show that MER-R1 achieves state-of-the-art performance and makes reasoning genuinely benefit emotion recognition.
ToE: 動的なマルチソース証拠の取得と集約を備えた階層的で説明可能なクレーム検証フレームワーク
フェイクニュースの急速な蔓延は、特に生成エンジン最適化(GEO)ポイズニングの下でAIが生成した誤った情報により、敵対的に作成されたコンテンツが検索システムによって体系的に表面化され、LLM推論が汚染されるため、情報エコシステムに対する脅威が増大しています。この論文では、各主張を動的に拡張する引数ツリーとしてモデル化する、自動化されたファクトチェックのための階層的証拠推論フレームワークである Tree of Evidence (ToE) を提案します。 ToE は、強化学習主導のマルチソース検索エージェント、証拠評価エージェント、および引数ツリー集約アルゴリズムを統合し、説明可能な証拠チェーンを通じてクレームを繰り返し分解、取得、検証します。さらに、取得プロセスの理論的分析を提供し、学習されたポリシーが情報理論的に最適なポリシーの近傍に収束することを保証する形式誤差限界を導き出します。複数のデータセットとバックボーン LLM にわたる実験では、ToE が競合ベースラインと比較して 4 ~ 24 パーセント ポイントの範囲の改善を達成し、特に敵対的に汚染された入力で顕著な改善が見られることが実証されました。
原文 (English)
ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation
The rapid spread of fake news poses increasing threats to information ecosystems, especially as AI-generated misinformation under Generative Engine Optimization (GEO) poisoning allows adversarially crafted content to be systematically surfaced by retrieval systems, contaminating LLM reasoning. In this paper, we propose Tree of Evidence (ToE), a hierarchical evidence reasoning framework for automated fact-checking that models each claim as a dynamically expanding argument tree. ToE integrates a reinforcement learning-driven multi-source retrieval agent, an evidence evaluation agent, and an argument tree aggregation algorithm to iteratively decompose, retrieve, and verify claims through an explainable evidence chain. We further provide a theoretical analysis of the retrieval process, deriving a formal error bound that guarantees the learned policy converges to a neighborhood of the information-theoretically optimal policy. Experiments across multiple datasets and backbone LLMs demonstrate that ToE achieves improvements ranging from 4 to 24 percentage points over competitive baselines, with particularly pronounced gains on adversarially poisoned inputs.
信頼性が高く堅牢な LLM 計画に向けて: シンボリック フィードバック駆動の反復的自己洗練フレームワーク
大規模言語モデル (LLM) は学界や産業界から広く注目を集めていますが、その導入には堅牢性と信頼性に関する重大なセキュリティ上の懸念が生じます。インテリジェントな動作の中核コンポーネントである計画は、LLM にとって依然として課題であり、本質的な複雑さのために、長期的な意思決定タスクでは実行不可能または不正確な解決策を生み出すことがよくあります。この論文では、長期計画における LLM の堅牢性と信頼性を強化するための、記号的なフィードバック駆動型の反復的自己洗練フレームワークを提案します。具体的には、論理シンボルを自然言語記述にマッピングするための自然言語プロンプト メカニズムが導入され、LLM がタスクの制約とセマンティクスをより適切に把握できるようになります。さらに、エラーを特定し、LLM が解釈できる修正指示に変換する記号検証器を設計し、それによって自己改善を導きます。さらに、計画認識機能を活用して目標の達成可能性を推測し、望ましい目標に向けたより効果的なガイダンスを促進します。経験的な結果は、提案されたフレームワークが長期的な計画タスクの実現可能性と正確性の両方を一貫して向上させることを示しています。これは、LLM ベースの計画の信頼性を高める有効性と、より信頼できる AI システムを可能にする可能性を強調しています。
原文 (English)
Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework
Large language models (LLMs) have attracted widespread attention from academia and industry, yet their deployment raises critical security concerns regarding robustness and reliability. Planning, a core component of intelligent behavior, remains challenging for LLMs, which often produce infeasible or incorrect solutions in long-horizon decision-making tasks due to inherent complexity. In this paper, we propose a symbolic feedback-driven iterative self-refinement framework to enhance the robustness and reliability of LLMs in long-horizon planning. Specifically, a natural language prompting mechanism is introduced to map logical symbols into natural language descriptions, enabling LLMs to better capture task constraints and semantics. We further design a symbolic verifier that identifies errors and converts them into corrective instructions interpretable by the LLM, thereby guiding self-refinement. In addition, we leverage a plan recognizer to infer goal reachability, facilitating more effective guidance toward desired goals. Empirical results demonstrate that the proposed framework consistently improves both feasibility and correctness in long-horizon planning tasks. This highlights its effectiveness in enhancing the reliability of LLM-based planning and potential to enable more trustworthy AI systems.
グラフ ワールド モデルのロールアウト エラーについて
ワールド モデルは、学習したダイナミクスをロールフォワードすることによって計画を立てるためによく使用されます。ただし、多くの計画環境はベクターやイメージではありません。これらは、エージェント、ツール、スキル、ルート、依存関係のグラフです。これらの設定では、局所的な予測誤差が局所的にとどまるか、グラフ全体に広がる可能性があり、エッジが固定されるのではなく予測されると、故障モードが再び変化します。この論文では、グラフ ワールド モデル (GWM) における長期ロールアウト エラーを研究します。ノード、エッジ、およびグラフレベルの意思決定のためのアクションノードを備えた統合された固定エッジおよび動的エッジ GWM フレームワークを定式化します。トポロジに起因する増幅をモデルに起因する増幅から分離するグラフ値のロールアウト境界を開発し、動的エッジ ロールアウト用にノードとエッジの結合演算子を導入します。分析に基づいて、スペクトルの正則化、ロールアウトの一貫性、クリティカル ノードの重み付けを組み合わせたエラー認識 GWM を提案します。合成トポロジと異種エージェント グラフ テストベッドでは、ロールアウト エラーと計画の後悔が時間とともに増大し、構造が進化する場合には動的エッジ トレーニングが必要になります。また、エラー認識 GWM は、予測精度を維持しながら長期的な発散を防ぎます。現実世界のグラフ ベンチマークは、GWM の範囲を明確にします。GWM は動的なグラフ ロールアウトとエージェント プランニングに最も役立ちますが、特殊なグラフ モデルは静的または疎な予測タスクに引き続き強力です。
原文 (English)
Understanding Rollout Error in Graph World Models
World models are often used for planning by rolling learned dynamics forward. Many planning environments, however, are not vectors or images; they are graphs of agents, tools, skills, routes, and dependencies. In these settings, a local prediction error may stay local or spread through the graph, and the failure mode changes again when edges are predicted rather than fixed. This paper studies long-horizon rollout error in Graph World Models (GWMs). We formulate a unified fixed-edge and dynamic-edge GWM framework with action nodes for node-, edge-, and graph-level decisions. We develop graph-valued rollout bounds that separate topology-induced amplification from model-induced amplification, and we introduce a joint node-edge operator for dynamic-edge rollouts. Guided by the analysis, we propose Error-Aware GWM, which combines spectral regularization, rollout consistency, and critical-node weighting. Across synthetic topologies and heterogeneous agent-graph testbeds, rollout error and planning regret grow with horizon, dynamic-edge training is needed when structure evolves, and Error-Aware GWM prevents long-horizon divergence while preserving prediction accuracy. Real-world graph benchmarks clarify the scope of GWMs: they are most useful for dynamic graph rollout and agent planning, while specialized graph models remain strong on static or sparse prediction tasks.
グラウンディングされた反復言語計画: パラメーター化された世界モデルが LLM エージェントにおける幻覚の伝播をどのように軽減するか
言語エージェントの世界モデルには 2 つの便利な形式があります。エージェントベースの世界モデルは LLM API を呼び出し、言語で柔軟に推論しますが、そのエラーは幻覚的な状態変化として現れ、通常の回帰損失ではスコアを付けるのが困難です。パラメーター化された世界モデルは、トレーニングされた遷移予測子です。そのエラーは、NodeMSE、デルタ精度、妥当性精度などの量を使用して測定する方が簡単ですが、通常、スタンドアロン プランナーとしては弱いです。これら 2 つのファミリーを 4 つのグラフ構造の計画ベンチマークで比較し、エージェントベースのケースの操作上の幻覚測定基準を導入します。この比較により、\textbf{Grounded Iterative Language Planning} (GILP) が動機付けられ、小規模なパラメーター化されたバックボーンのみをトレーニングし、それを API ベースのエージェント推論と組み合わせます。バックボーンは、有効なアクション、予測された状態デルタ、リスク、および価値を提供します。 LLM はアクションと想像上のデルタを草案します。そして、この 2 つが一致しない場合には、整合性ゲートが修正を要求します。実際の GPT-4o-mini 通話では、GILP は幻覚状態の割合を 0.176 から 0.035 に削減します。調整されたシミュレーターアブレーションでは、最大 22% の余分な LLM コールを追加するだけで、成功率が 0.668 から 0.838 に上昇します。
原文 (English)
Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents
World models for language agents come in two useful forms. An agent-based world model calls an LLM API and reasons flexibly in language, but its errors appear as hallucinated state changes that are hard to score with ordinary regression losses. A parameterized world model is a trained transition predictor; its errors are easier to measure with quantities such as NodeMSE, delta accuracy, and validity accuracy, but it is usually weaker as a standalone planner. We compare these two families on four graph-structured planning benchmarks and introduce operational hallucination metrics for the agent-based case. The comparison motivates \textbf{Grounded Iterative Language Planning} (GILP), which trains only a small parameterized backbone and combines it with API-based agent reasoning. The backbone supplies valid actions, predicted state deltas, risk, and value; the LLM drafts an action and imagined delta; and a consistency gate asks for revision when the two disagree. On real GPT-4o-mini calls, GILP reduces hallucinated-state rate from 0.176 to 0.035. In calibrated simulator ablations, it raises success from 0.668 to 0.838 while adding only ~22% extra LLM calls.
ATOD: マルチターン自律エージェント向けのアニーリングされたターン対応オンポリシー蒸留
長期的な対話型タスクのために小規模な言語モデル エージェントをトレーニングするには、迅速な模倣と報酬主導型の改善の両方が必要です。オンポリシー蒸留 (OPD) は教師による密度の高い指導を提供し、通常は初期段階で急速に向上しますが、生徒が教師に近づくとその効果は飽和し、最終的なパフォーマンスの上限が制限されます。強化学習(RL)は、環境報酬を直接最適化し、より高い報酬定義の上限に向けた探索的改善を促進しますが、まばらで遅延したフィードバックにより、初期段階の学習の効率がOPDよりもはるかに低くなります。この論文では、この相補性を明示的に利用するハイブリッド オンライン蒸留アルゴリズムである ATOD (Annealed Turn-aware On-policy Distillation) を提案します。 (1) ATOD はアニーリングされた OPD-RL スケジュールを使用します。OPD は教師レベルの行動に近づくための初期トレーニングを支配しますが、RL は報酬ベースの探索を促進するために徐々に強化されます。 (2) ATOD は、ターンレベルの不一致不確実性再重み付け (T-DUR) を導入しています。これは、ユーティリティの高いターンをソフトに増幅し、長い軌道での緻密な監視を改善します。 ALFWorld、WebShop、および Search-QA での実験では、ATOD が競合するトレーニング後のベースラインを常に上回っていることが示されています。3 つの生徒サイズにわたって、ATOD は平均成功率を OPD より 3.03 ポイント、GRPO より 23.62 ポイント向上させ、対応する教師モデルを 2.16 ポイント上回っています。
原文 (English)
ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents
Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but its gains saturate once the student approaches the teacher, limiting the final performance ceiling. Reinforcement learning (RL) directly optimizes environment rewards and encourages exploratory improvement toward a higher reward-defined ceiling, but sparse and delayed feedback makes early-stage learning much less efficient than OPD. In this paper, we propose ATOD (Annealed Turn-aware On-policy Distillation), a hybrid online distillation algorithm that explicitly exploits this complementarity. (1) ATOD uses an annealed OPD-RL schedule: OPD dominates early training to approach teacher-level behavior, while RL is gradually strengthened to drive reward-based exploration. (2) ATOD introduces Turn-level Disagreement-Uncertainty Reweighting (T-DUR), which softly amplifies high-utility turns and improves dense supervision in long trajectories. Experiments on ALFWorld, WebShop, and Search-QA show that ATOD consistently outperforms competing post-training baselines: across the three student sizes, ATOD improves average success rate by 3.03 points over OPD and 23.62 points over GRPO, while surpassing the corresponding teacher models by 2.16 points.
NormAct: 身体化された計画における隠れた社会規範遵守のベンチマーク
マルチモーダル大規模言語モデル (MLLM) は、自己中心的な環境で具体化されたプランナーとして導入されることが増えています。タスクを成功させるには、指示された目標を達成するだけでなく、社会的に適切な方法で行動する必要があります。明示的な目標によって特定の行動が最適化される場合もありますが、暗黙の社会規範によって隠れた制約が課せられることがよくあります。既存の評価は通常、明示的な目標達成または直接的な規範知識に焦点を当てており、計画者がアクション シーケンス内のこれらの隠れた制約を推測して適用できるかどうかを評価することはほとんどありません。目標達成、規範遵守、全体的なタスクの成功に関する計画を評価する、具体化された社会規範の相互作用のベンチマークである NormAct を紹介します。 NormAct は通常のタスク内に隠れた規範を独自に埋め込み、明示的な指示なしにモデルがそれらを実現できるかどうかをテストします。最先端の MLLM (GPT-5.4、Claude Opus 4.7、Gemini 3 Pro) を用いた実験では、大きなギャップが明らかになりました。モデルは 67.3\% のケースで明示的な目標を達成しましたが、隠れた基準に準拠したのは 26.4\% のみでした。合図条件の実験によると、このギャップは一般的な社会知識の欠如ではなく、文脈の中で関連する規範を活性化し、基礎づける際の課題から生じていることが示されています。これに対処するために、計画前にシーン関連の規範を推測するコンテキスト条件付きキュー ジェネレーターである NormPerceptor を提案し、タスクの成功率を 24.2\% から 46.7\% に向上させます。私たちの結果は、身体化されたエージェントが隠れた規範を積極的に検出し、視覚的な証拠に基づいて行動計画の制約として統合できるようにすることの重要性を強調しています。私たちのベンチマークは https://huggingface.co/datasets/Caleb196x/NormAct で公開されています。
原文 (English)
NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning
Multimodal large language models (MLLMs) are increasingly deployed as embodied planners in egocentric environments, where task success requires not only achieving instructed goals but also acting in socially appropriate ways. While explicit goals may render certain actions optimal, implicit social norms often impose hidden constraints. Existing evaluations typically focus on explicit goal achievement or direct norm knowledge, seldom assessing whether planners can infer and apply these hidden constraints within action sequences. We introduce NormAct, a benchmark for embodied social-norm interactions that evaluates plans on Goal Achievement, Norm Compliance, and overall Task Success. NormAct uniquely embeds hidden norms within ordinary tasks, testing whether models can realize them without explicit instruction. Experiments with state-of-the-art MLLMs (GPT-5.4, Claude Opus 4.7, Gemini 3 Pro) reveal a significant gap: models achieve explicit goals in 67.3\% of cases, but comply with hidden norms in only 26.4\%. Cue-condition experiments indicate that this gap stems not from a lack of general social knowledge, but from challenges in activating and grounding relevant norms in context. To address this, we propose NormPerceptor, a context-conditioned cue generator that infers scene-relevant norms prior to planning, increasing Task Success from 24.2\% to 46.7\%. Our results underscore the importance of enabling embodied agents to proactively detect hidden norms, ground them in visual evidence, and integrate them as action-planning constraints. Our benchmark is publicly available at https://huggingface.co/datasets/Caleb196x/NormAct.
検証可能な幾何学問題解決: ソルバー駆動の自動形式化と定理の提案
幾何学の問題解決では、神経的直観と象徴的な厳密性を組み合わせた神経象徴的パラダイムがますます採用されています。しかし、現在のフレームワークは 2 つの中核段階で深刻なボトルネックに悩まされています。1 つはマルチモーダル変換を下流のソルバーの互換性から切り離された静的タスクとして扱う自動形式化で、もう 1 つは定理予測です。定理予測では、固定ルール ライブラリが原因でソルバーが演繹的行き詰まりに頻繁に遭遇します。これらに対処するために、形式化と演繹の両方を通じてシンボリック ソルバーを実行オラクルとして扱うソルバー駆動フレームワークである SD-GPS を提案します。まず、ソルバー駆動の自動形式化は、教師あり形式言語適応と解決可能性ガイド付き強化学習を QwenVL3-2B 上に構築された単一のモジュールに統合し、実行可能性を中心的なトレーニング信号にします。第 2 に、Verified Theorem Proposing では、現在の証明状態から局所的な補助補題を提案する行き詰まり認識エージェントを導入し、記号検証を通じてすべての提案をフィルタリングすることで健全性を確保します。 Geometry3K と PGPS9K の実証評価では、SD-GPS が標準補完、多肢選択、クロスモーダル参照レジームにわたって既存の MLLM、ニューラル、およびニューロシンボリック手法を一貫して上回るパフォーマンスを示し、マルチモーダル知覚とシンボリック実行の間のループを閉じることで幾何学的推論が大幅に改善されることが証明され、ニューラル エージェントがどのように形式的なシステムに基づいて検証可能な問題解決能力を達成できるかについての深い洞察が得られます。
原文 (English)
Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing
Geometry Problem Solving have increasingly adopt the neuro-symbolic paradigm, combining neural intuition with symbolic rigor. However, current frameworks suffer from severe bottlenecks in two core stages: autoformalization, which treats multimodal translation as a static task decoupled from downstream solver compatibility, and theorem prediction, where solvers frequently hit a deductive impasse due to fixed rule libraries. To address these, we propose SD-GPS, a solver-driven framework that treats the symbolic solver as an execution oracle throughout both formalization and deduction. First, Solver-Driven Autoformalization unifies supervised formal-language adaptation and solvability-guided reinforcement learning into a single module built on QwenVL3-2B, making executability the central training signal. Second, Verified Theorem Proposing introduces an impasse-aware agent that proposes local auxiliary lemmas from current proof states, ensuring soundness by filtering all proposals through symbolic verification. Empirical evaluations on Geometry3K and PGPS9K demonstrate that SD-GPS consistently outperforms existing MLLM, neural, and neuro-symbolic methods across standard completion, multiple-choice, and cross-modal reference regimes, proving that closing the loop between multimodal perception and symbolic execution significantly improves geometric reasoning, offering profound insights into how neural agents can be grounded by formal systems to achieve verifiable problem-solving capabilities.
RelBall: ナレッジ グラフを完成させるためのクォータニオン回転を備えたリレーション ボール
現実世界のナレッジ グラフは多くの有効な事実が欠けており、不完全であることがよくあります。 Knowledge Graph Completion (KGC) は、既知のトリプルを使用して欠落しているリンクを予測し、それによってグラフのカバレッジを強化することを目的としています。主な課題は、対称性、反対称性、反転、合成、意味階層などの多様な関係パターンをモデル化することです。 RotatE などの既存のモデルは、対称、反対称、逆、可換の合成パターンをキャプチャできますが、非可換の合成には苦労します。 Rotate3D は、3 次元の回転による非可換性を導入することでこの問題に対処していますが、依然としてナレッジ グラフに普及している意味階層を捉えることができません。さらに、どちらのモデルも 1 対多の関係を効果的にモデル化できません。これらの制限を克服するために、2 つの革新性で Rotate3D を拡張する RelBall を提案します。まず、私たちのモデルはモデル階層にモジュラス変換を導入し、抽象的な概念をより小さなモジュライに向けて、具体的なインスタンスをより大きなモジュライに向けて推進します。 2 番目に、1 対 1、1 対多、多対 1、および多対多の関係をモデル化するために、末尾中心のリレーション ボールを導入します。 RelBall には次の利点があります。(1) 前述のパターンを含むすべてのリレーショナル パターンをカバーします。 (2) モジュラスが意味レベルを直接反映する解釈可能な階層表現。 (3) 1 対 1、1 対多、多対 1、および多対多の関係のサポート。複数のデータセットでの実験により、さまざまなベースラインに対する RelBall の競合リンク予測パフォーマンスが実証されています。
原文 (English)
RelBall: Relation Ball with Quaternion Rotation for Knowledge Graph Completion
Real-world knowledge graphs are often incomplete, lacking many valid facts. Knowledge Graph Completion (KGC) aims to predict missing links using known triples, thereby enhancing graph coverage. A key challenge is modeling diverse relational patterns such as symmetry, antisymmetry, inversion, composition and semantic hierarchy. Existing models such as RotatE can capture symmetric, antisymmetric, inverse, and commutative composition patterns, yet struggle with non-commutative composition. Rotate3D addresses this by introducing non-commutativity via three-dimensional rotations, but still fails to capture the semantic hierarchies prevalent in knowledge graphs. Moreover, both models cannot effectively model one-to-many relations. To overcome these limitations, we propose RelBall, which extends Rotate3D with two innovations. First, our model introduces modulus transformation to model hierarchies, driving abstract concepts toward smaller moduli and concrete instances toward larger ones. Second, it introduces a tail-centric relation ball to model one-to-one, one-to-many, many-to-one, and many-to-many relations. RelBall offers the following advantages: (1) coverage of all relational patterns, including the ones mentioned above; (2) an interpretable hierarchical representation where the modulus directly reflect semantic levels; (3) support for one-to-one, one-to-many, many-to-one, and many-to-many relations. Experiments on multiple datasets demonstrate RelBall's competitive link prediction performance against various baselines.
高度な因果推論
リフト推論は、区別できないオブジェクトの代表を使用することで、確率的グラフィカル モデルの区別不能性を利用し、それによって正確な回答を維持しながらクエリの回答を高速化します。この記事では、リフティングを適用してリレーショナル ドメインの因果効果を効率的に計算する方法を示します。より具体的には、リフトモデルに因果知識を組み込み、そこへの介入の正式なセマンティクスを与えるために、パラメトリック因果因子グラフ (PCFG) を導入します。さらに、リフトされたレベルで因果効果を計算するためのリフト因果推論 (LCI) アルゴリズムを提示します。これにより、因果ベイジアン ネットワークなどで、命題推論と比較して因果推論が大幅に高速化されます。さらに、部分的な因果知識を処理し、LCI を拡張して PD-PCFG でリフト因果推論を実行するための PCFG の一般化として、部分有向パラメトリック因果因子グラフ (PD-PCFG) を提示します。これにより、リフト因果推論の適用可能性を、因果関係についての事前知識が少なくて済むより広範囲のモデルに拡張します。
原文 (English)
Lifted Causal Inference
Lifted inference exploits indistinguishabilities in probabilistic graphical models by using a representative for indistinguishable objects, thereby speeding up query answering while maintaining exact answers. In this article, we show how lifting can be applied to efficiently compute causal effects in relational domains. More specifically, we introduce parametric causal factor graphs (PCFGs) to incorporate causal knowledge in lifted models and give a formal semantics of interventions therein. We further present the Lifted Causal Inference (LCI) algorithm to compute causal effects on a lifted level, thereby drastically speeding up causal inference compared to propositional inference, e.g., in causal Bayesian networks. In addition, we present partially directed parametric causal factor graphs (PD-PCFGs) as a generalisation of PCFGs to handle partial causal knowledge and extend LCI to perform lifted causal inference in a PD-PCFG, thereby extending the applicability of lifted causal inference to a broader range of models requiring less prior knowledge about causal relationships.
JD Oxygen AI アイテム センター (Oxygen AIIC) V1: アイテムの理解、管理、アプリケーションのための産業規模の LLM/VLM 中心のソリューション
世界最大の電子商取引プラットフォームの 1 つである JD.com は、7 億人を超えるアクティブ ユーザーと数百万の販売者にサービスを提供し、数百億の SKU のカタログを持っています。この規模では、高品質で構造化されたアイテム ナレッジは、より良い消費者エクスペリエンス、管理コストの削減、運用効率の向上を支えますが、その生産と提供には、急速に出現するコンセプト、大規模な SKU 向けの高品質なナレッジの生産、および多様な下流要件という 3 つの産業規模の課題が生じます。これらの課題に対処するために、アイテム知識の生産とサービスのための LLM/VLM 上に構築された産業規模のプラットフォームである JD Oxygen AI Items Center (Oxygen AIIC) を紹介します。 Oxygen AIIC は 4 つの核となる柱を中心に構築されています。(i) 人間と AI の効率的なコラボレーションによって推進されるオントロジー エンジニアリング。これは、数百万のエントリを含むオントロジーの動的な進化と機敏な拡張をサポートします。 (ii) スループット向上戦略と組み合わせることで、数百億の SKU に対するスケーラブルで拡張性のある高スループットの AI アイテム ライブラリの作成を可能にする「セマンティック検索と識別」(S2D) 知識識別アーキテクチャ。 (iii) 自己進化する項目理解 LLM/VLM は、安定かつ制御可能な方法で改善し、94.2% の精度と 82.8% の再現率で知識生産を可能にします。 (iv) データとサービスのハブとして機能する統合アイテム トンネル。 Oxygen AIIC は現在、数万の JD カテゴリをカバーし、Huawei Ascend NPU で 1 日に数億件のアイテム更新を処理しています。数千億ものアイテム知識資産が蓄積されています。 Oxygen AIIC は、検索、レコメンデーション、運用、カテゴリ プランニングなど、中核となるビジネス シナリオ全体に展開され、大規模に目に見える利益をもたらしました。検索トラフィックのカバレッジは 80.4% に達し、商品情報の品質問題は 37% 減少し、商品リスト中の主要属性の自動入力率は 80% を超えました。
原文 (English)
JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications
JD.com, one of the world's largest e-commerce platforms, serves over 700 million active users and millions of merchants, with a catalog of tens of billions of SKUs. At this scale, high-quality, structured item knowledge underpins a better consumer experience, lower management costs, and higher operational efficiency-yet producing and serving it poses three industrial-scale challenges: fast-emerging concepts, high-quality knowledge production for massive SKUs, and diverse downstream requirements. To address these challenges, we present the JD Oxygen AI Item Center (Oxygen AIIC), an industrial-scale platform built on LLMs/VLMs for item-knowledge production and service. Oxygen AIIC is built around four core pillars: (i) ontology engineering driven by efficient human-AI collaboration, which supports the dynamic evolution and agile expansion of an ontology with millions of entries; (ii) a "Semantic Search then Discrimination"(S2D) knowledge identification architecture that, combined with throughput improvement strategies, enables scalable, extensible, and high-throughput AI Item Library production for tens of billions of SKUs; (iii) self-evolving item-understanding LLMs/VLMs that improve in a stable and controllable manner, enabling knowledge production with 94.2% precision and 82.8% recall; and (iv) a unified item tunnel that serves as the data and service hub. Oxygen AIIC now covers tens of thousands of JD categories and processes hundreds of millions of item updates per day on Huawei Ascend NPUs. It has accumulated hundreds of billions of item-knowledge assets. Deployed across core business scenarios-including search, recommendation, operations, category planning-Oxygen AIIC has delivered measurable gains at scale. Search-traffic coverage reaches 80.4%, item-information quality issues drop by 37%, the automated fill rate of core attributes during item listing exceeds 80%.
マルチホップナレッジグラフの質問応答のためのオントロジーに基づく証拠パス推論
ナレッジ グラフ質問応答 (KGQA) は、構造化された事実に基づいて推論することで、自然言語の質問に答えることを目的としています。既存のマルチホップ KGQA 手法は主にトピック中心の拡張に依存していますが、これは 2 つの重要な課題に直面しています。1 つはノイズの多い混合タイプのパスにより検索空間が急速に拡大すること、もう 1 つは検索されたパスが複雑な質問の意味的制約を満たせない可能性があることです。これらの課題に対処するために、マルチホップ KGQA 用のオントロジーに基づく証拠パス推論フレームワークである OPI を提案します。 OPI は、リレーション中心のオントロジー グラフを導入してリレーションのヘッドテール タイプの制約をキャプチャし、回答側の制約にコンパクトなインターフェイスを提供します。このオントロジー グラフに基づいて、OPI はまず、予測された応答タイプを互換性のある最終ホップ関係にマッピングし、トピック側のプレフィックス拡張と応答側の最終ホップ マッチングを組み合わせることにより、双方向検索メカニズムを導入します。これにより、ノイズの多い混合タイプの拡張が抑制されます。 OPI はさらに、質問コンテキストの下で取得されたパスと回答候補を再評価するための反復的改善戦略を採用し、タイプ互換性があるが質問に無関係な証拠をフィルタリングして、より信頼性の高い回答予測を実現します。 WebQSP、CWQ、および MetaQA での実験では、OPI が検索スペースを大幅に削減し、以前の最も強力な結果と比べて WebQSP で Hit@1/F1 が 4.6/5.0 ポイント、CWQ で 8.9/3.3 ポイント向上し、取得モジュールのみで MetaQA で飽和に近い Hit@1 を達成することが示されました。
原文 (English)
Ontology-Guided Evidence Path Inference for Multi-hop Knowledge Graph Question Answering
Knowledge graph question answering (KGQA) aims to answer natural-language questions by reasoning over structured facts. Existing multi-hop KGQA methods mainly rely on topic-centered expansion, which faces two key challenges: the search space rapidly grows with noisy mixed-type paths, and retrieved paths may fail to satisfy the semantic constraints of complex questions. To address these challenges, we propose OPI, an ontology-guided evidence path inference framework for multi-hop KGQA. OPI introduces a relation-centric ontology graph to capture the head-tail type constraints of relations, providing a compact interface for answer-side constraints. Based on this ontology graph, OPI first introduces a bidirectional retrieval mechanism by mapping the predicted answer type to compatible final-hop relations and combining topic-side prefix expansion with answer-side final-hop matching, thereby suppressing noisy mixed-type expansion. OPI further adopts an iterative refinement strategy to reassess retrieved paths and candidate answers under the question context, filtering type-compatible but question-irrelevant evidence for more reliable answer prediction. Experiments on WebQSP, CWQ, and MetaQA show that OPI substantially reduces the search space, improves Hit@1/F1 by 4.6/5.0 points on WebQSP and 8.9/3.3 points on CWQ over the strongest prior results, and achieves near-saturated Hit@1 on MetaQA with the retrieval module alone.
ハイテク システム設計のための AI 駆動合成: イノベーションの自動化
この記事では、変革的なパラダイムとして設計の自動化 (AiD) を提示することで、現代のハイテク システム設計に内在する組み合わせの複雑さに対処します。私たちは、ディープラーニングと生成AIを利用して新しいシステムの作成を自動化するフレームワークであるコンピューテーショナル・デザイン・シンセシス(CDS)を提案します。 2 つのケーススタディ (e-ドライブ システムの設計と空間寸法の問題) が、このアプローチの証拠として機能します。ケーススタディで使用されている AI 主導の手法は、エンジニアリングにおける根本的な変化を表しており、シミュレーション ベースの最適化から、人間の監視を最小限に抑えた自律的な設計へと進歩しています。
原文 (English)
AI-Driven Synthesis for High-Tech System Design: Automating Innovation
This article addresses the combinatorial complexity inherent in modern high-tech system design by presenting automation-in-design (AiD) as a transformative paradigm. We propose computational design synthesis (CDS), a framework utilising deep learning and generative AI to automate the creation of novel systems. Two case studies (e-drive system design and spatial dimensioning problem) serve as proof-points for this approach. The AI-driven methods used in the case studies represent a fundamental shift in engineering, advancing from simulation-based optimisation towards autonomous design with minimal human supervision.
検証可能な報酬を伴うタンデム強化学習
検証可能な報酬を伴う強化学習 (RLVR) は、大規模な言語モデルの推論能力を大幅に向上させ、競技数学などの分野で専門家、さらには超人的なパフォーマンスに達しました。ただし、弱いエージェントや人間が実際にこの機能を利用できるかどうかは、RLVR の推論が可読性の低さや言語の混在などの特異なパターンに偏っていることが文書化されているため、確実性ははるかに低くなります。タンデム トレーニングは、この互換性の問題を対象とした最近導入されたパラダイムです。訓練を受けてより強力な先輩が、凍結された弱い後輩と各ロールアウトを共同で生成し、二人はチームとして報酬を与えられるため、先輩は後輩が従うことができる方法を推論するよう促されます。しかし、このパラダイムはこれまでのところ、概念実証の設定でのみ実証されており、最新の RLVR パイプラインの長い思考連鎖に対応できるかどうかは不明のままです。この研究では、タンデム強化学習 (TRL) を提案します。これは、タンデム トレーニング パラダイムを RLVR に取り入れます。 TRL では、シニアと凍結されたジュニアが確率的に交互に推論を生成し、結果の生成が報酬を受け、標準の GRPO 損失がシニアに適用されます。 Qwen3-4B-Instruct を競技数学でトレーニングすると、TRL は単独推論能力でバニラ GRPO と一致する一方、同じロールアウト構造から 3 つの特性が一緒に現れることがわかりました。それは、ジュニアとのハンドオフの堅牢性の強化、ジュニアからの分布ドリフトの減少、ジュニアにとって思考の連鎖の読みやすさです。私たちの結果は、マルチモデル通信と人間の互換性において実用的な利益をもたらす、RLVR の有望な道筋を示しています。
原文 (English)
Tandem Reinforcement Learning with Verifiable Rewards
Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capability of large language models, reaching expert or even superhuman performance in domains such as competition math. However, whether weaker agents and humans can actually harness this capability is far less certain, with RLVR documented to drift reasoning toward idiosyncratic patterns such as poor readability and language mixing. Tandem training is a recently introduced paradigm that targets this compatibility problem: a trained, stronger senior co-generates each rollout with a frozen, weaker junior, and the two are rewarded as a team, so the senior is pushed to reason in ways the junior can follow. Yet this paradigm has so far been demonstrated only in proof-of-concept settings, leaving open whether it scales to the long chains of thought of the modern RLVR pipeline. In this work, we propose Tandem Reinforcement Learning (TRL), which carries the tandem training paradigm into RLVR. In TRL, the senior and a frozen junior alternate stochastically to co-generate the reasoning, the resulting generation is rewarded, and the standard GRPO loss is applied to the senior. Training Qwen3-4B-Instruct on competition math, we find that TRL matches vanilla GRPO on solo reasoning capability while three properties emerge together from the same rollout structure: stronger handoff robustness with the junior, reduced distributional drift from the junior, and a chain-of-thought more legible to the junior. Our results demonstrate a promising route for RLVR with practical payoffs in multi-model communication and human compatibility.
病原体固有の免疫システム: アーキテクチャ、分類学、および工学
静的チャットボットから、永続メモリ、ツール使用プロトコル、およびマルチエージェントコラボレーションを備えた自律エージェントへの移行により、AI の脅威の状況は根本的に拡大しました。境界セキュリティやトレーニング時間の調整などの現在の防御メカニズムは、エージェントの能動的な推論ループの外部にあるままです。その結果、それらは不十分です。完全に調整されたエージェントは、メモリポイズニング、ツールチェーン操作、またはマルチエージェントプロトコル攻撃によるランタイムハイジャックに対して非常に脆弱なままです。この重大なギャップに対処するために、エージェントの認知ループ内に直接埋め込まれた初の生物学的にインスピレーションを受けた内因性防御アーキテクチャであるエージェントネイティブ免疫システム (ANIS) を導入します。私たちのフレームワークは 4 つの主要な貢献を示しています。まず、6 層の免疫タワー (L0 ~ L5) を設計し、非認知的な物理的および論理的な分離層としてバリア免疫 (L1) を明確に組み込みます。第 2 に、エージェント ウイルスとエージェント ワクチンの統一分類法を確立し、表面的なノンパラメトリック防御と堅牢なパラメトリック ワクチンの間の重要な区別を形式化します。第三に、ハーネス トライアド (メタ、自己、自動) を概念化します。これは、継続的免疫学習 (CIL) を推進し、ワクチンが新たな脅威に動的に適応できるようにする自己監視、メタ認知自動化バックボーンです。最後に、モデルの調整とエージェントの免責の間の厳密な理論的境界を確立します。調整はトレーニング中に静的な「憲法上の」価値基盤を提供しますが、ANIS は実行時に動的な「法執行」メカニズムとして機能します。最後に、免疫プロトコルの標準化、自己免疫率(偽陽性介入率)などの新しい評価指標、集合知エコシステム内の病原体とワクチンの間の共進化のダイナミクスなど、この分野の未解決の課題を枠組み化します。
原文 (English)
Agent-Native Immune System: Architecture, Taxonomy, and Engineering
The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has fundamentally expanded the AI threat landscape. Current defense mechanisms, such as perimeter security and training-time alignment, remain external to the agent's active reasoning loop. Consequently, they fall short: a fully aligned agent remains highly vulnerable to runtime hijacking via memory poisoning, tool-chain manipulation, or multi-agent protocol attacks. To address this critical gap, we introduce the Agent-Native Immune System (ANIS), the first biologically inspired, endogenous defense architecture embedded directly within the agent's cognitive loop. Our framework presents four primary contributions. First, we design a six-layer Immune Tower (L0-L5), distinctly incorporating Barrier Immunity (L1) as a non-cognitive, physical-and-logical isolation layer. Second, we establish a unified taxonomy of Agent Viruses and Agent Vaccines, formalizing the critical distinction between superficial non-parametric defenses and robust parametric vaccines. Third, we conceptualize the Harness Triad--Meta, Self, and Auto--a self-monitoring, meta-cognitive automation backbone that drives Continual Immune Learning (CIL), enabling vaccines to dynamically adapt to novel threats. Finally, we establish a rigorous theoretical demarcation between model alignment and agent immunity: while alignment provides a static "constitutional" value foundation during training, ANIS serves as the dynamic "law enforcement" mechanism during runtime. We conclude by framing open challenges for the field, including immune protocol standardization, novel evaluation metrics such as the Autoimmunity Rate (false-positive intervention rate), and the co-evolutionary dynamics between pathogens and vaccines within collective intelligence ecosystems.
DataStates-LLM: 構成可能な状態プロバイダーを使用したトランスフォーマー モデルのスケーラブルなチェックポイント機能
Large Transformer ベースのモデル、特に Large Language Model (LLM) の急速な成長により、現在では数兆のパラメーターに拡張されており、複雑なハイブリッド並列処理戦略 (データ、テンソル、パイプライン並列処理など) を使用して数千の GPU にわたるトレーニングが必要になっています。この大規模な分散状態のチェックポイントは、回復力、サスペンド/レジューム、望ましくないトレーニング軌跡の調査、モデルの進化の説明など、幅広いユースケースにとって重要です。ただし、既存のチェックポイント ソリューションは通常、モデルの状態を不透明なバイナリ BLOB として扱い、メモリの場所 (GPU 対ホスト)、複数のファイルにシャーディングおよび分割された「論理」オブジェクトの数、データ型 (テンソル対 Python オブジェクト)、およびそれらのシリアル化要件によって異なる、基礎となるデータ構造の「3D 異質性」を無視します。これにより、デバイスからホストへの転送のブロック、データを意識しないシリアル化、ストレージ I/O 競合により、実行時に重大なオーバーヘッドが発生します。このペーパーでは、状態プロバイダーを利用して状態の抽象化をデータの移動から切り離す新しいチェックポイント アーキテクチャである DataStates-LLM を紹介します。 DataStates-LLM は、前方および後方パス中のモデル パラメーターの不変性を利用して、「遅延」のノンブロッキング非同期スナップショットを実行します。状態プロバイダーを導入することで、断片化された異種シャードを効率的に結合し、メタデータのシリアル化をバルク テンソル I/O とオーバーラップさせます。 256 個の A100-40GB GPU 上の最大 70B パラメーターのモデルで DataStates-LLM を評価します。私たちの結果は、DataStates-LLM が最先端のソリューションと比較して最大 4$\times$ 高いチェックポイント処理スループットを実現し、エンドツーエンドのトレーニング時間を最大 2.2$\times$ 短縮し、超大規模な LLM トレーニングにおけるシリアル化と異質性のボトルネックを効果的に軽減することを示しています。
原文 (English)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
The rapid growth of Large Transformer-based models, specifically Large Language Models (LLMs), now scaling to trillions of parameters, has necessitated training across thousands of GPUs using complex hybrid parallelism strategies (e.g., data, tensor, and pipeline parallelism). Checkpointing this massive, distributed state is critical for a wide range of use cases, such as resilience, suspend-resume, investigating undesirable training trajectories, and explaining model evolution. However, existing checkpointing solutions typically treat model state as opaque binary blobs, ignoring the ``3D heterogeneity'' of the underlying data structures--varying by memory location (GPU vs. Host), number of ``logical'' objects sharded and split across multiple files, data types (tensors vs. Python objects), and their serialization requirements. This results in significant runtime overheads due to blocking device-to-host transfers, data-oblivious serialization, and storage I/O contention. In this paper, we introduce DataStates-LLM, a novel checkpointing architecture that leverages State Providers to decouple state abstraction from data movement. DataStates-LLM exploits the immutability of model parameters during the forward and backward passes to perform ``lazy'', non-blocking asynchronous snapshots. By introducing State Providers, we efficiently coalesce fragmented, heterogeneous shards and overlap the serialization of metadata with bulk tensor I/O. We evaluate DataStates-LLM on models up to 70B parameters on 256 A100-40GB GPUs. Our results demonstrate that DataStates-LLM achieves up to 4$\times$ higher checkpointing throughput and reduces end-to-end training time by up to 2.2$\times$ compared to state-of-the-art solutions, effectively mitigating the serialization and heterogeneity bottlenecks in extreme-scale LLM training.
立場: 「機械学習解除」という用語は LLM で過剰に使用されています
大規模な言語モデルは、規制上の削除義務、著作権/ライセンス紛争、安全性または製品ポリシーの要件により、トレーニング データ、知識、または動作を「忘れる」という要求にますます直面しています。この意見書では、機械の非学習はLLM研究の用語として過剰に使用されており、データセット定義の削除、つまり正確に指定された忘却セットのトレーニングの影響を除去して、結果として得られるモデルがそのデータなしでの再トレーニングとほぼ区別がつかないようにするために保留されるべきであると主張しています。現在「未学習」とラベル付けされている多くのタスク(有害なリクエストの拒否、エンティティ/知識の削除、対象を絞った抑制など)は、異なる、多くの場合ポリシーに依存する目的を追求しているため、異なる用語やベースライン(調整、抑制、編集、難読化など)が必要であると私たちは主張します。さらに、この混乱は表面的なものではないと主張します。論文が同じラベルの下で異なる暗黙の保証を行っているため、メトリクスとベンチマークは意図した範囲外で頻繁に再利用され、再トレーニングの同等性がテストされず、派生機能が残っている場合でも、表面レベルの非開示(例:低い ROUGE/忘れ精度)が報われることになります。最後に、明示的な保証と参照モデルに結び付けられたより厳密な用語と、主張された目的に一致する評価を求めることで終わります。
原文 (English)
Position: The Term "Machine Unlearning" Is Overused in LLMs
Large language models increasingly face demands to "forget" training data, knowledge, or behaviors due to regulatory deletion obligations, copyright/licensing disputes, and safety or product-policy requirements. This position paper argues that machine unlearning is overused as a term in LLM research and should be reserved for dataset-defined deletion: removing the training influence of a precisely specified forget set such that the resulting model is approximately indistinguishable from retraining without that data. We contend that many tasks currently labeled "unlearning" (e.g., refusal for harmful requests, entity/knowledge removal, or targeted suppression) pursue different, often policy-dependent objectives and therefore require different terminology and baselines (e.g., alignment, suppression, editing, obfuscation). We further argue that this confusion is not cosmetic: because papers make different implicit guarantees under the same label, metrics and benchmarks are frequently reused outside their intended scope, rewarding surface-level non-disclosure (e.g., low ROUGE/forget accuracy) even when retraining-equivalence is not tested and derived capabilities remain. We conclude by calling for stricter terminology tied to explicit guarantees and reference models, and for evaluations that match the claimed objective.
OverFlowLight: リアルタイムの渋滞防止と都市交差点の交通信号の最適化
都市交通渋滞の深刻な結果であるキュー オーバーフローは、車両のキューが交差点の収容能力を超えたときに発生し、上流の交通を妨げ、連鎖的な渋滞を引き起こします。主にスループットを考慮して最適化された、一般的な交通信号制御 (TSC) アルゴリズムは、多くの場合、ピーク時のオーバーフローに対処できず、渋滞を悪化させ、安全上の問題を引き起こします。私たちは、オーバーフローを先制的に解決し、全体的な TSC パフォーマンスを向上させるように設計されたリアルタイム フレームワークである OverFlowLight を提案します。まず、カメラとレーダーからのマルチモーダルセンシングを活用して、リアルタイムでオーバーフローを正確に検出するメカニズムを導入します。検出すると、専用のオーバーフロー フェーズを動的に生成して信号サイクルに挿入し、ブロッキング キューをクリアします。これは、ルールベースの迅速なオーバーフロー介入と強化学習 (RL) などのコントローラー バックエンドを組み合わせたハイブリッド制御設計によって調整され、長期的な効率を実現します。私たちは、3 つの主要都市の 43 の交差点にわたって、OverFlowLight の大規模な現実世界の導入を実施しました。このフレームワークは、既存の RL ベースの TSC エージェントとのシームレスな統合を実証し、そのモジュール性と実用的な適用性を強調しています。実証結果によると、導入されたベースラインと比較して、OverFlowLight はオーバーフロー インシデントを 60.4% 削減し、ネットワーク スループットを 18.2% 増加させます。さらに、専門家が調整した信号計画によくある手動介入の必要性が大幅に軽減されます。この研究は、交通渋滞を積極的に防止するための最初の実用的でスケーラブルなデータ駆動型フレームワークを提示し、回復力のある効率的な都市交通システムを構築するための重要なコンポーネントを提供します。当社のデモ ビデオ、コード、データセットは、匿名の URL、https://anonymous.4open.science/r/OverFlowLight-FBF9 で入手できます。
原文 (English)
OverFlowLight: Real-Time Gridlock Prevention and Traffic Signal Optimization for Urban Intersections
Queue overflow, a severe consequence of urban traffic congestion, occurs when vehicle queues exceed intersection capacity, obstructing upstream traffic and triggering cascading gridlocks. Prevailing traffic signal control (TSC) algorithms, primarily optimized for throughput, often fail to address overflow during peak hours, exacerbating congestion and creating safety hazards. We propose OverFlowLight, a real-time framework designed to preemptively resolve overflow and enhance overall TSC performance. It first introduces a mechanism to accurately detect overflow in real-time by leveraging multi-modal sensing from cameras and radars. Upon detection, it dynamically generates and inserts dedicated overflow phases into the signal cycle to clear the blocking queues. This is orchestrated by a hybrid control design that combines rapid rule-based overflow intervention with controller back ends such as reinforcement learning (RL) for longer-horizon efficiency. We conducted extensive real-world deployments of OverFlowLight across 43 intersections in three major cities. The framework demonstrates seamless integration with existing RL-based TSC agents, highlighting its modularity and practical applicability. Empirical results show that OverFlowLight reduces overflow incidents by 60.4% and increases network throughput by 18.2% compared to deployed baselines. Furthermore, it substantially diminishes the need for manual intervention common with expert-tuned signal plans. This work presents the first practical, scalable, and data-driven framework for actively preventing traffic gridlock, offering a crucial component for building resilient and efficient urban transportation systems. Our demonstration videos, codes and datasets are available at the anonymous URL, https://anonymous.4open.science/r/OverFlowLight-FBF9.
CalBrief: 大規模な言語モデルを使用した、証拠に基づいて調整された科学的ブリーフィングのためのパイロット診断ベンチマーク
大規模言語モデル (LLM) は研究アシスタントとしてますます使用されていますが、研究成果を裏付けとなる証拠の強さと範囲に合わせて調整できるかどうかは依然として不明です。私たちは、証拠に基づいて調整された科学的ブリーフィングを研究しています。関連する論文の限定されたパッケージが与えられた場合、システムは、証拠の強度、範囲の境界、証拠の欠落に関する警告を含むパッケージレベルの要点を生成する必要があります。私たちは、16 の異質な科学的証拠パッケージと 96 の人が検証したポイントからなる検証済みのパイロット ベンチマークを提供し、監査可能な役割/ギャップ/強みのフレームワークである CalBrief を、ブリーフィングがどこで失敗しているかを特定するための診断プローブとして使用しています。公正なスキーマ評価の下では、構造化された組織は役割とギャップの推論を改善しますが、明示的な強度調整ポリシーは体系的に過度に保守的であり、多数派および直接LLMのベースラインを下回っています。その理由を説明するために、保守主義の 3 つの潜在的な原因を分離する 3 つのクローズド モデル バックボーン (GPT-4o、Claude Sonnet、Gemini Flash) にわたって制御された診断を実行します。保守主義のギャップの約 63% は、ラベル空間を 2 値 {中程度、弱い} から 4 方向 {中程度、弱い、不確実、不十分な証拠} に拡大したことに起因します (すべてのバックボーンで p < 0.001)。ギャップ/スコープ信号注入によるものはわずか 1% (重要ではありません)。残りの 36% はパイプライン ポリシー自体から生じます。また、4 方向予測は事後的にバイナリに折りたたまれ、直接バイナリ プロンプトと一致するかそれを超える可能性があるため、追加のラベルには厳密な一致では隠蔽される情報が含まれることもわかりました。ラベルレベルの強度判断と監査可能な証拠の整理は、現在緊張状態にある別個の能力であり、LLM 研究アシスタントについては個別に評価される必要があります。
原文 (English)
CalBrief: A Pilot Diagnostic Benchmark for Evidence-Calibrated Scientific Briefing with Large Language Models
Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence. We study evidence-calibrated scientific briefing: given a bounded package of related papers, a system should generate package-level takeaways with evidence strength, scope boundaries, and missing-evidence caveats. We contribute a verified pilot benchmark of 16 heterogeneous scientific evidence packages and 96 human-verified takeaways, and we use CalBrief, an auditable role/gap/strength framework, as a diagnostic probe to locate where briefing breaks down. Under a fair-schema evaluation, structured organization improves role and gap reasoning, but an explicit strength-calibration policy is systematically over-conservative and falls below majority and direct-LLM baselines. To explain why, we run a controlled diagnostic across three closed-model backbones (GPT-4o, Claude Sonnet, Gemini Flash) that separates three potential causes of conservatism. Approximately 63% of the conservatism gap is attributable to expanding the label space from binary {moderate, weak} to four-way {moderate, weak, uncertain, insufficient_evidence} (p < 0.001 across all backbones); only 1% is attributable to gap/scope signal injection (not significant); the remaining 36% arises from the pipeline policy itself. We also find that 4-way predictions can be post-hoc collapsed back to binary and then match or exceed direct binary prompting, so the extra labels carry information that strict matching hides. Label-level strength judgment and auditable evidence organization are distinct abilities currently in tension, and should be evaluated separately for LLM research assistants.
エージェント的出版プロトコル: 科学出版を最新化する試み
科学の進歩の多くは、コードの実行方法、図の再現方法、エッジケースの解釈方法、有用なフォローアップの指示の選択方法、失敗したパスの回避方法など、暗黙のノウハウに依存しているにもかかわらず、科学出版物は依然として主に静的な原稿を中心に編成されています。大規模な言語モデル エージェントは、将来の読者や研究者が直接使用できる形式で、知識だけでなく運用ノウハウも公開する機会を生み出します。この文書では、論文をコード、データ、環境情報、再現性指示、およびエージェント向け指示ファイルとともにパッケージ化するための軽量リポジトリ形式である Agentic Publication Protocol (APP) について概説します。 APP は、バージョン管理されたリポジトリを出版オブジェクトとして扱い、\texttt{AGENTS.md} とオプションのスキルを使用して、作業を説明し、可能であれば重要な結果を再現し、追跡調査をサポートできるペーパー エージェントを定義します。プロトコルの設計原則と詳細、およびプロトコルに基づいて論文を出版するために役立つエージェントのスキルについて説明します。また、プロトコルと関連するエージェントのスキルを評価および改善するための開発ツールについても説明します。最後に、エージェント時代の科学研究の将来について、より広範な議論を提供します。
原文 (English)
Agentic Publication Protocol: An Attempt to Modernize Scientific Publication
Scientific publication is still organized primarily around static manuscripts, even though much of scientific progress depends on tacit know-how: how to run code, reproduce figures, interpret edge cases, choose useful follow-up directions, and avoid failed paths. Large language model agents create an opportunity to publish not only knowledge, but also operational know-how in a form that future readers and researchers can directly use. This paper outlines the Agentic Publication Protocol (APP), a lightweight repository format for packaging a paper together with code, data, environment information, reproducibility instructions, and an agent-facing instruction file. APP treats a version-controlled repository as the publication object and uses \texttt{AGENTS.md} and optional skills to define a paper agent that can explain the work, reproduce key results when possible, and support follow-up research. We describe the design principles and details of the protocol, as well as the agent skills useful for publishing papers under the protocol. We also describe development tools for evaluating and improving the protocol and associated agent skills. Finally, we provide a broader discussion of the future of scientific research in the agent era.
SidConArena: オープンエンドのポジティブサム交渉ゲームでエージェントを評価する環境
LLM エージェントを評価するには、静的な推論やゼロサム ゲームを超えた動的な環境が必要です。現実世界の経済相互作用は、多くの場合無制限で動機が複雑です。エージェントは交渉し、プラスサムの黒字を創出し、希少な資産をめぐって競争し、収益の遅れを考慮して計画を立てなければなりません。オープンエンドのポジティブサム交渉で LLM エージェントを評価するための新しいベンチマーク フレームワークである SidConArena を紹介します。 SidConArena は、拘束力のある取引を伴う自然言語ネゴシエーション、決定論的なコンバーターベースの生産、および長期資産の封印入札オークションという 3 つのフェーズが結合された、有限地平線で部分的に観測可能な確率論的ゲームとしてマルチプレイヤー エコノミーを形式化します。このフレームワークは、構造化された観察、フェーズを認識したエージェントのディスパッチング、ニューラルシンボリックアクションインターフェイス、および非同期実行を組み合わせて、ルールに基づいた評価を維持しながら自由形式の対話を可能にします。同種トーナメントでも異種トーナメントでも、より強力なフロンティアモデルはより高い経済的成果を達成しますが、エージェントは依然としてリソースを誤って評価し、消極的に交渉し、長期的な投資計画は限られたままです。
原文 (English)
SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game
Evaluating LLM agents requires dynamic environments that go beyond static reasoning and zero-sum games. Real-world economic interaction is often open-ended and mixed-motive: agents must negotiate, create positive-sum surplus, compete for scarce assets, and plan under delayed returns. We introduce SidConArena, a new benchmark framework for evaluating LLM agents in open-ended, positive-sum bargaining. SidConArena formalizes a multi-player economy as a finite-horizon partially observable stochastic game with three coupled phases: natural-language negotiation with binding trades, deterministic converter-based production, and sealed-bid auctions for long-term assets. The framework combines structured observations, phase-aware agent dispatching, a neural-symbolic action interface, and asynchronous execution, enabling free-form interaction while preserving rule-grounded evaluation. Across homogeneous and heterogeneous tournaments, stronger frontier models achieve higher economic outcomes, yet agents still misvalue resources, bargain passively, and remain limited in long-horizon investment planning.
CNN および ResNet アーキテクチャを使用した MRI 画像内の脳腫瘍の自動検出
ディープラーニングは、医療画像分析、特に MRI スキャンを使用した病気の検出において大きな可能性を示しています。脳腫瘍の正確かつ早期の診断は、脳の構造が複雑であり、手作業による解釈に依存しているため、依然として困難です。この研究では、畳み込みニューラル ネットワークと残差ネットワークを使用して、MRI 画像から脳腫瘍を検出するための深層学習ベースの自動化アプローチを紹介します。転移学習は、ResNet18 と ResNet50 という 2 つの事前トレーニング済みアーキテクチャで適用され、MRI スキャンを腫瘍と非腫瘍のカテゴリに分類します。実験は 3,929 枚の脳 MRI 画像のデータセットに対して行われ、モデルの深さと微調整戦略の影響を評価します。結果は、ResNet18 が ResNet50 の 96% と比較して 97% の高い精度を達成し、限られた医療データに対するより良い一般化を示していることを示しています。提案されたフレームワークにより、迅速、正確、かつ費用対効果の高い脳腫瘍の検出が可能になり、早期診断と臨床上の意思決定がサポートされます。
原文 (English)
Automated brain tumor detection in MRI images using CNN and ResNet architectures
Deep learning has shown significant potential in medical image analysis, particularly for disease detection using MRI scans. Accurate and early diagnosis of brain tumors remains challenging due to the complexity of brain structures and reliance on manual interpretation. This work presents an automated deep learning-based approach for brain tumor detection from MRI images using Convolutional Neural Networks and Residual Networks. Transfer learning is applied with two pretrained architectures, ResNet18 and ResNet50, to classify MRI scans into tumor and non-tumor categories. Experiments are conducted on a dataset of 3,929 brain MRI images, evaluating the impact of model depth and fine-tuning strategies. The results show that ResNet18 achieves a higher accuracy of 97% compared to 96% for ResNet50, demonstrating better generalization on limited medical data. The proposed framework enables fast, accurate, and cost-effective brain tumor detection, supporting early diagnosis and clinical decision-making.
コーディング LLM における暗黙的なソフトウェア ワールド モデルの評価に向けて
ソフトウェア エンジニアリングは、人間によって実行されるか AI エージェントによって実行されるかにかかわらず、ソフトウェアがどのように動作するかについて推論する必要があります。私たちは、このような推論をサポートする内部モデルをソフトウェア ワールド モデルと呼び、現在のコード実行ベンチマークは、よく研究されたその一部である制御フローをカバーしていると見ています。このペーパーでは、観察可能な軸を実行リソースに移すことで、より広範な評価に向けて一歩を踏み出します。テスト結果と例外クラスに加えて、ピーク メモリ、実時間、およびメソッドとラインの粒度でランク付けされたプロファイラー出力を予測します。実際のソフトウェア エンジニアリング タスクに近いテストを行うため、データ ソースとして SWE-bench Verified を使用します。最先端のモデルも含め、テストされたすべてのモデルは、控えめなパフォーマンスと脆弱な動作を示しており、ソース コードがどのように記述されているかではなく、ソフトウェアがどのように実行されるかについての理解が著しく欠如していることを示唆しています。
原文 (English)
Towards Evaluation of Implicit Software World Models in Coding LLMs
Software engineering, whether performed by humans or by AI agents, requires reasoning about how software behaves. We call the internal model that supports such reasoning the software world model, and view current code-execution benchmarks as covering one well-studied slice of it -- control flow. In this paper, we take a step toward a broader evaluation by shifting the observable axis to execution resources: alongside test outcome and exception class, we predict peak memory, wall-clock time, and ranked profiler outputs at method and line granularity. We use SWE-bench Verified as the source of data to hold the test close to real-world software engineering tasks. All tested models, frontier ones included, show modest performance and brittle behaviour, suggesting a notable lack of understanding of how software is executed, as opposed to how its source code is written.
解釈可能な量子オートエンコーダを使用した脳 MRI における圧縮駆動の異常検出
私たちは、脳 MRI データにおける圧縮駆動型の異常検出のための量子オートエンコーダー (QAE) を研究しています。このアプローチでは、角度エンコーディングを利用して画像パッチを量子状態にマッピングし、その後、補助的なトラッシュ量子ビットを介して情報を破棄するように訓練された変分エンコーダー/デコーダー アーキテクチャを利用します。異常スコアは、入力が正常データと比較して圧縮にどの程度抵抗するかを反映し、より高いスコアは学習された正常多様体からの偏差に対応します。公的に利用可能な脳 MRI DICOM データセットで評価したこの方法は、スライス レベルの ROC-AUC が約 0.95、パッチ レベルの ROC-AUC が約 0.813 を達成し、従来のオートエンコーダおよび PCA ベースラインを上回りました。学習されたパラメータの分析により、エンコーダとデコーダの顕著な非対称性が明らかになり、効果的な異常検出は、パラメータの大きさやデコーダの表現力の向上ではなく、エンコーダ内の構造化情報圧縮から生じます。これにより、原則に基づいたしきい値の選択をサポートする明確な動作体制による、制御された圧縮と再構成のトレードオフが実現します。定性的評価では、QAE が腫瘍領域と位置合わせされた空間的に局所的な異常ヒートマップを生成することがさらに示されています。この結果は、有望なベースライン性能によって裏付けられ、量子オートエンコーダーが学習された潜在表現に関する非圧縮性に基づいて異常検出のための解釈可能かつ制御可能なメカニズムを提供することを示しています。この研究は、量子機械学習における圧縮ダイナミクスを研究するための原則的なツールとしての量子オートエンコーダーの可能性を強調しており、医療画像ワークフローにおける意思決定支援に有望な意味をもたらします。
原文 (English)
Compression-Driven Anomaly Detection in Brain MRI Using an Interpretable Quantum Autoencoder
We study a quantum autoencoder (QAE) for compression-driven anomaly detection in brain MRI data. The approach leverages angle encoding to map image patches into quantum states, followed by a variational encoder-decoder architecture trained to discard information via auxiliary trash qubits. Anomaly scores reflect the degree to which inputs resist compression relative to normal data, with higher scores corresponding to deviations from the learned normal manifold. Evaluated on publicly available brain MRI DICOM datasets, the method achieves a slice-level ROC-AUC of approximately 0.95 and a patch-level ROC-AUC of approximately 0.813, outperforming classical autoencoder and PCA baselines. Analysis of the learned parameters reveals a pronounced encoder-decoder asymmetry, where effective anomaly detection arises from structured information compression within the encoder rather than increased parameter magnitude or decoder expressivity. This results in a controlled compression-reconstruction trade-off with a clear operating regime that supports principled threshold selection. Qualitative evaluation further shows that the QAE produces spatially localized anomaly heatmaps aligned with tumorous regions. The results, supported by promising baseline performances, demonstrate that quantum autoencoders provide an interpretable and controllable mechanism for anomaly detection based on incompressibility with respect to a learned latent representation. This work highlights the potential of quantum autoencoders as a principled tool for studying compression dynamics in quantum machine learning, with promising implications for decision support in medical imaging workflows.
すべての関係が同じように回転するわけではありません: 視点の堅牢な 3D シーン グラフ生成のための変換を意識したデカップリング
3D シーン グラフ生成 (3DSGG) は、3D シーンを構造化されたオブジェクト-リレーション-オブジェクト グラフとして表現し、空間を理解するためのコンパクトなリレーショナル抽象化を提供します。身体化されたインテリジェンス設定では、同じ 3D シーンがヨー回転によって異なる視点からエージェントによって観察される場合があります。ただし、現在の 3DSGG モデルは、そのような視点のシフトの下で期待される変換動作に従う関係予測を生成できないことがよくあります。この動作は、述語レベルの変換の不均一性に関連する経験的な不一致を明らかにしています。左、前、右、後ろなどの方向述語は観測フレームに合わせて変換されるはずですが、ほとんどの接触、サポート、および意味論的述語 (上に立つ、接続されるなど) は安定したままであるはずです。この不一致を減らすために、我々は、述語変換動作に従って関係推論を分離し、視点安定したオブジェクト表現によってサポートされる視点堅牢な 3DSGG フレームワークである、Transformation-Aware Decoupling (TAD) を提案します。 TAD は、関係推論を 2 つの部分に分解します。1 つは視点間で安定している必要がある手がかりを学習し、もう 1 つは観測フレームとともに変化する必要がある方向の手がかりを学習します。 2 つの部分は、標準のマルチラベル述語予測のためにマージされます。変換固有の記述子とグループ認識の補助監視により、2 つのブランチが相補的な関係の手がかりを捕捉することが促進されます。 3DSSG に関する広範な実験により、TAD は標準ベンチマークの下で競争力のあるパフォーマンスを維持しながら、トレーニング時の回転強化を行わずにヨー視点変更下で最先端の堅牢性を達成できることが示されています。プロジェクト ページは https://tad-predicate.github.io/ で利用できます。
原文 (English)
Not All Relations Rotate Alike: Transformation-Aware Decoupling for Viewpoint-Robust 3D Scene Graph Generation
3D Scene Graph Generation (3DSGG) represents 3D scenes as structured object-relation-object graphs, providing a compact relational abstraction for spatial understanding. In embodied intelligence settings, the same 3D scene may be observed by agents from viewpoints that differ by yaw rotations. However, current 3DSGG models often fail to produce relation predictions that follow the expected transformation behavior under such viewpoint shifts. This behavior reveals an empirical mismatch related to predicate-level transformation heterogeneity: directional predicates such as left, front, right, and behind should transform with the observation frame, whereas most contact, support, and semantic predicates such as standing on and attached to should remain stable. To reduce this mismatch, we propose Transformation-Aware Decoupling (TAD), a viewpoint-robust 3DSGG framework that decouples relation reasoning according to predicate transformation behavior and is supported by viewpoint-stable object representations. TAD decomposes relation reasoning into two parts: one learns cues that should stay stable across viewpoints, while the other learns directional cues that should change with the observation frame. The two parts are merged for standard multi-label predicate prediction. Transformation-specific descriptors and group-aware auxiliary supervision encourage the two branches to capture complementary relation cues. Extensive experiments on 3DSSG show that TAD achieves state-of-the-art robustness under yaw viewpoint changes without training-time rotation augmentation, while maintaining competitive performance under the standard benchmark. The project page is available at https://tad-predicate.github.io/.
GRAFT: シロイヌナズナにおける連鎖遺伝子発現および表現型形質予測のための生物学的グラフおよびハイパーグラフ ベンチマーク
どの遺伝子が生物のどの形質を制御しているかを理解することは、依然として生物学における中心的な課題の 1 つです。データ収集技術の大幅な進歩にも関わらず、遺伝子を形質にマッピングする私たちの能力には依然として限界があります。このゲノム対フェノム (G2P) の課題は、植物育種を含むいくつかの問題領域に及び、高次元で異種の生物学的に構造化されたデータを推論できる方法が必要です。ただし、現在のデータセットとデータ リポジトリは、このタスクに十分な機能を備えていません。現在の研究は遺伝子発現と形質データを関連付けておらず、ほとんどが非常に特殊な形質に焦点を当てているため、考えられる相関関係の幅は限られています。このギャップに対処するために、我々は、植物生物学のモデル生物であるシロイヌナズナにおける遺伝子発現プロファイルと表現型形質測定値を結びつける、精選されたマルチモーダルデータセットである、シロイヌナズナ機能形質に関する新しいジーングラフ回帰(GRAFT)データセットを紹介します。 GRAFT は、表現型予測や解釈可能なグラフ学習などのタスクをサポートします。さらに、遺伝子と形質の関連性を検証するために、生物学的に情報を得たハイパーグラフ ベースラインを含む、従来の回帰および説明ベースラインをベンチマークします。私たちの知る限り、これは、同じシロイヌナズナ標本の多峰性遺伝子情報と異種形質または表現型データを提供する最初のデータセットです。 GRAFTでは、複数のソースからの遺伝子情報、高次遺伝子ペアリング、形質データを利用して、遺伝子型と表現型の関係を正確に理解する研究の促進を目指します。
原文 (English)
GRAFT: Biological Graph and Hypergraph Benchmarks for Linked Gene Expression and Phenotypic Trait Prediction in Arabidopsis thaliana
Understanding which genes control which traits in an organism remains one of the central challenges in biology. Despite significant advances in data collection technology, our ability to map genes to traits is still limited. This genome-to-phenome (G2P) challenge spans several problem domains, including plant breeding, and requires methods capable of reasoning over high-dimensional, heterogeneous, and biologically structured data. Current datasets and data repositories, however, are not well-equipped for this task. Current studies do not link gene expression and trait data, and most focus on very specific traits, limiting the breadth of possible correlations. To address this gap, we present the novel Gene-Graph Regression for Arabidopsis Functional Traits (GRAFT) dataset, a curated multi-modal dataset linking gene expression profiles with phenotypic trait measurements in Arabidopsis thaliana, a model organism in plant biology. GRAFT supports tasks such as phenotype prediction and interpretable graph learning. In addition, we benchmark conventional regression and explanatory baselines, including a biologically-informed hypergraph baseline, to validate gene-trait associations. To the best of our knowledge, this is the first dataset to provide multimodal gene information and heterogeneous trait or phenotype data for the same Arabidopsis thaliana specimens. With GRAFT, we aim to foster research to accurately understand the relationship between genotypes and phenotypes using gene information, higher-order gene pairings, and trait data from multiple sources.
置き換え: LLM エージェントのメモリ更新ギャップの診断とトレーニング
大規模言語モデル (LLM) エージェントは、ユーザーの移動、価格の更新、計画の修正など、事実が変化する長い複数セッションの対話を通じて動作します。正しく動作するには、ファクトの現在の値を使用し、置き換えられた値を破棄する必要があります。私たちはこの機能を実際の会話データから分離し、それが明確な未解決の障害であることを示します。 LongMemEval の知識更新サブセットでは、エージェントのフル コンテキストを制限された自己維持メモリに置き換えると、フロンティア モデル (gpt-5.4) であっても精度が 92% から 77% に低下します。この差は統計的に有意 (対応マクネマー p<0.005) であり、フル コンテキストの精度は 92% 近くで飽和する一方で、モデル スケール全体にわたって持続します。したがって、ボトルネックは理解ではなく記憶の維持であり、より強力なモデルによって解決されることはありません。次に、これが単にメモリ不足なのかどうかを尋ねますが、そうではないことがわかります。会話が 24 倍になると、精度はさらに低下し (68% から 28% に)、エージェントに比例してより多くのメモリを与えても、検出可能な回復は得られません (28% から 28%、n=25)。失敗は圧縮率ではなく、会話の長さによって決まります。私たちは、この測定値をトレーニング信号に変えるオープンな強化学習環境 (ベリファイア/プライム RL スタック上) である Supersede をリリースします。エージェントは、現在の値から回答すると報酬が与えられ、古い値に対してはペナルティが課されます。最後に、ループを閉じて、ギャップがトレーニング可能であることを示します。この環境で小さなオープン モデル (Qwen2.5-3B) を GRPO で微調整すると、ハーネスではなく学習されたポリシーが利益をもたらすことを示す単調なチェックポイント曲線に沿って、実際の目に見えない会話で保持されているスーパーセッションの精度がほぼ 2 倍になります (9.0% から 16.7%、1 回の実行)。私たちの知る限り、これは、報酬が一時的な事実通貨をターゲットとする最初のトレーニング可能な環境であり、スーパーセッションギャップを測定するだけでなくトレーニングして抑制できる最初の証拠です。
原文 (English)
Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents
Large language model (LLM) agents operate over long, multi-session interactions in which facts change: a user moves, a price updates, a plan is revised. Acting correctly requires using the current value of a fact and discarding values that have been superseded. We isolate this ability on real conversational data and show that it is a distinct, unsolved failure. On the knowledge-update subset of LongMemEval, replacing an agent's full context with a bounded, self-maintained memory drops accuracy from 92% to 77% even on a frontier model (gpt-5.4), a gap that is statistically significant (paired McNemar p<0.005) and persists across model scale while full-context accuracy saturates near 92%. The bottleneck is therefore memory maintenance, not comprehension, and is not closed by a stronger model. We then ask whether this is merely an undersized memory, and find it is not: as the conversation grows 24x, accuracy falls further (from 68% to 28%), and granting the agent proportionally more memory yields no detectable recovery (28% to 28%, n=25). The failure scales with the length of the conversation, not the compression ratio. We release Supersede, an open reinforcement-learning environment (on the verifiers / prime-rl stack) that turns this measurement into a training signal: agents are rewarded for answering from the current value and penalized for stale ones. Finally, we close the loop and show the gap is trainable: GRPO fine-tuning a small open model (Qwen2.5-3B) on this environment nearly doubles its held-out supersession accuracy on real, unseen conversations (9.0% to 16.7%, a single run), along a monotonic checkpoint curve indicating the learned policy, not the harness, carries the gain. To our knowledge this is the first trainable environment whose reward targets temporal fact-currency, and the first evidence the supersession gap can be trained down, not only measured.
投機的洗練: ハイブリッド自己回帰拡散デコード戦略とベンチマーク全体でのその動作
自己回帰 (AR) と拡散デコーディングを組み合わせた生成システムをどのように評価すればよいでしょうか?私たちは、エントロピーに導かれた選択的マスキングを使用して AR ドラフトからマスクされた拡散言語モデルをウォームスタートするトレーニング不要のハイブリッド手法である Speculative Refinement (SpecRef) を通じてこの疑問を研究します。 3 つの異なる評価プロトコル (実行ベースの pass@1、完全一致、対数尤度スコアリング) を使用して 6 つのベンチマーク (HumanEval、MBPP、GSM8K、BBH、ARC-Challenge、HellaSwag) で SpecRef を評価すると、特定のシステムを超えて関連するいくつかの結果が明らかになりました。 (1) コード ベンチマークは、構造発見と論理的正しさを混同しています。構文足場を提供すると、精度を変更することなくゼロ近くから 20% 以上に引き上げられます。ベースライン障害の多くが構造的なものであることを示すモデル。 (2) 多段階の修正によってすでに修正されたトークンが劣化し、単一モデルの評価では見えないベンチマークの飽和上限が露出するリファインメント テンション現象。 (3) 対数尤度および生成評価は、同じモデル ペアに対して異なるモデル ランキングを生成し、異なる能力を測定することを示唆しています。 (4) 標準の Python ポストプロセスは、非 AR ジェネレーターのコード評価を静かに中断します。これらの観察結果は、あらゆる多段階生成パイプラインまたは非自己回帰生成パイプラインに当てはまり、より診断評価の実践に向けたものとなります。
原文 (English)
Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks
How should we evaluate generation systems that combine autoregressive (AR) and diffusion decoding? We study this question through Speculative Refinement (SpecRef), a training-free hybrid method that warm-starts a masked diffusion language model from an AR draft using entropy-guided selective masking. Evaluating SpecRef across six benchmarks (HumanEval, MBPP, GSM8K, BBH, ARC-Challenge, HellaSwag) with three distinct evaluation protocols (execution-based pass@1, exact-match, log-likelihood scoring), we surface several findings relevant beyond our specific system: (1) code benchmarks conflate structural discovery with logical correctness: providing a syntactic scaffold lifts accuracy from near zero to over 20% without changing the model, indicating that much of the baseline failure is structural; (2) a refinement tension phenomenon where multi-stage correction degrades already-correct tokens, exposing benchmark saturation ceilings invisible to single-model evaluation; (3) log-likelihood and generative evaluation produce different model rankings for the same model pair, suggesting they measure different capabilities; (4) standard Python post-processing silently breaks code evaluation for non-AR generators. These observations apply to any multi-stage or non-autoregressive generation pipeline and point toward more diagnostic evaluation practices.
DMV-Bench: 偶発的キュー注入による長期マルチモーダルエージェントの視覚記憶の診断
エージェントの記憶に関する研究は急速に成熟していますが、ほぼ完全にテキスト側です。インタラクティブ環境において、エージェントが書き留められるものではなく、見たものを本当に記憶する必要があるときを問う既存のベンチマークはほとんどありません。マルチモーダル エージェントのビジュアル メモリの最初のインタラクティブ ベンチマークである DMV-Bench (コード: https://github.com/yyyujintang/DMV-Bench) を紹介します。 DMV-Bench は、1,000 種類の製品バリエーションを含む制御された家庭用家具電子商取引カタログに基づいて構築されており、テキスト漏洩契約により各タスクの識別信号がピクセルのみで保持されます。自律的なショッピング セッションのチェーン全体で、訪問したすべての製品画像には、事前にレンダリングされた固有の付随的なキューが含まれており、エージェントは後で特定のキュー製品を呼び出して、その URL に移動するように求められます。デュアルコーディング理論に触発されて、私たちは視覚的コードと言語的コードを並行して維持するメモリアーキテクチャである DualMem を提案します。 DMV ベンチでは、DualMem は、Gemini 2.5 Flash と Qwen2.5-VL-7B の両方で、{5、10、15、50} のすべてのチェーン長 J で、キャプション ベースラインと 3 つの最近のマルチモーダル エージェント メモリ システムを上回っています。これは、メモリ バンク サイズとエンコード位置のバイアスに関する制御をリードしており、視覚がキューをエンドツーエンドで伝達する一方で言語チャネルが視覚でキューを伝達する非対称デュアル コーディング方式によります。より小さなクエリ基盤の役割を果たします。
原文 (English)
DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write down. We introduce DMV-Bench (Code: https://github.com/yyyujintang/DMV-Bench), the first interactive benchmark for multimodal-agent visual memory. DMV-Bench is built on a controlled home-furnishing e-commerce catalogue of 1,000 product variants in which a text-leakage contract keeps the discriminative signal of each task in the pixels alone. Across a chain of autonomous shopping sessions, every visited product image carries a unique, pre-rendered incidental cue, and the agent is later asked to recall a particular cued product and navigate to its URL. Inspired by dual-coding theory, we propose DualMem, a memory architecture that maintains a visual and a verbal code in parallel. On DMV-Bench, DualMem outperforms a caption baseline and three recent multimodal agent-memory systems at every chain length J in {5, 10, 15, 50} on both Gemini 2.5 Flash and Qwen2.5-VL-7B, with the lead surviving controls for memory-bank size and encoding-position bias, and an asymmetric dual-coding regime in which vision carries the cue end-to-end while the verbal channel plays a smaller query-grounding role.
大規模言語モデルが視覚学習者に教える: きめの細かい概念的知識のクロスモダリティ伝達
大規模言語モデル (LLM) は、大規模なテキストの事前トレーニングを通じて取得された広範な概念的知識を備えていますが、他のモダリティでモデルを監視する可能性はまだ十分に検討されていません。この研究では、言語のみの教師から視覚のみの生徒モデルに高レベルの意味知識を伝達するためのシンプルで効果的なフレームワークである LaViD (言語から視覚への知識の蒸留) を提案します。 LaViD は、ペアになったマルチモーダル データに依存するのではなく、ビジュアル クラス間の意味論的な区別を調べる多肢選択質問 (MCQ) を生成するよう LLM に促すことで、LLM から概念的なシグナルを引き出します。各クラスは、これらの MCQ にわたるソフト ラベル分布にマッピングされ、補助蒸留損失を通じて生徒をガイドする豊富な概念的シグネチャを形成します。注目すべき点は、画像データにアクセスできない言語のみの教師を使用しているにもかかわらず、LaViD は、複数のきめ細かいベンチマークにわたって視覚言語モデルから抽出する MaKD のような最近の手法を常に上回っていることです。また、DKD や MLKD などの最先端の視覚的蒸留方法と比較して、競合または優れたパフォーマンスを実現し、ロジット標準化と組み合わせることでさらに性能が向上します。 Waterbirds データセットでは、LaViD は最悪グループの精度を大幅に向上させ、蒸留との偽相関に対するロバスト性の向上を示しています。コードは https://github.com/lliangthomas/lavid で入手できます。
原文 (English)
Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge
Large Language Models (LLMs) possess broad conceptual knowledge acquired through large-scale text pretraining, yet their potential to supervise models in other modalities remains underexplored. In this work, we propose LaViD--Language-to-Visual Knowledge Distillation--a simple and effective framework for transferring high-level semantic knowledge from a language-only teacher to a vision-only student model. Instead of relying on paired multimodal data, LaViD elicits conceptual signals from an LLM by prompting it to generate multiple-choice questions (MCQs) that probe semantic distinctions between visual classes. Each class is mapped to a soft label distribution over these MCQs, forming a rich conceptual signature that guides the student through an auxiliary distillation loss. Notably, despite using a language-only teacher without access to image data, LaViD consistently outperforms recent methods like MaKD that distill from vision-language models across multiple fine-grained benchmarks. It also achieves competitive or superior performance compared to state-of-the-art visual distillation methods such as DKD and MLKD, with further gains when combined with logit standardization. On the Waterbirds dataset, LaViD substantially improves worst-group accuracy, demonstrating enhanced robustness to spurious correlations with distillation. Code is available at https://github.com/lliangthomas/lavid.
コンテキスト対応トランスフォーマー
コンテキスト対応トランスフォーマーを導入します。これは、ブロックに入る前に各トークンを事前コンテキスト化する D 層トランスフォーマー ブロックから構築された新しいリカレント ニューラル ネットワーク アーキテクチャです。左から右への生成中に、修正ネットワークは前の位置のブロック出力 (過去のコンテキストのキャッシュされた概要) を現在のトークン埋め込みと結合します。そのため、トークンは生の埋め込みとしてではなく、すでにコンテキスト化されたブロックに入力されます。逐次推論では、修正チェーンによりアーキテクチャがリカレント ニューラル ネットワークになります。トレーニングでは、シーケンス全体にわたって補正プロセスを K 回展開し、各ステップですべての位置を並行して処理します。事前トレーニングされたトランスフォーマーは、ゼロ初期化補正 FFN を追加して微調整することで、コンテキスト対応モデルに変換することもできます。標準トランスフォーマー、バリアント、アブレーションとすべての比較を行い、幅、深さ、ブロック サイズ、および 2 つのデータセットを評価します。 D=5 モデルは 12 層トランスを上回り、A100 では 1.7 倍高速に生成します。 K=10 の場合、単層モデル (D=1) は 2.6 倍の推論速度向上で 6 層トランスフォーマーを上回り、逐次推論は並列 K=10 と 0.01 PPL 以内で一致します。このアーキテクチャは、幅広い表現と長いコンテキストから最も恩恵を受けます。ポインタ追跡タスクでは、BPTT でトレーニングされた D=1 は 10 の合成レベルすべてを解決しますが、標準の変換器は階段状の深さ依存性を示します。
原文 (English)
The Context-Ready Transformer
We introduce the context-ready transformer, a new recurrent neural network architecture built from a D-layer transformer block that pre-contextualizes each token before it enters the block. During left-to-right generation, a correction network combines the previous position's block output -- a cached summary of past context -- with the current token embedding, so the tokenenters the block already contextualized rather than as a raw embedding. At sequential inference, the correction chain makes the architecture a recurrent neural network. For training, we unroll the correction process K times over the full sequence, processing all positions in parallel at each step. A pretrained transformer can also be converted to a context-ready model by adding a zero-initialized correction FFN and fine-tuning. We evaluate across widths, depths, block sizes, and two datasets, with all comparisons against standard transformers, variants, and ablations. A D=5 model beats a 12-layer transformer while generating 1.7x faster on an A100. With K=10, a single-layermodel (D=1) beats a 6-layer transformer with a 2.6x inference speedup, and sequential inference matches parallel K=10 to within 0.01 PPL. The architecture benefits most from wide representations and long contexts. On a pointer-chasing task, D=1 trained with BPTT solves all 10 composition levels, while standard transformers exhibit staircase-like depth dependence.
マルチモーダルグラフベースのソーシャルメディア人気予測のベンチマーク
ソーシャル メディアの人気予測は、初期段階の観察からオンライン コンテンツの将来のリーチや影響を予測することを目的としています。正確な予測により、ユーザー、クリエイター、プラットフォームによる広告の最適化や戦略的コンテンツ計画などの主要な下流アプリケーションが可能になります。大幅な進歩にもかかわらず、既存の人気予測作業は、マルチモーダルなコンテンツと時間的な社会的相互作用シグナルを一緒に考慮していないことがよくあります。さらに、文献は、データセット、モダリティ、観察ウィンドウ、予測ターゲット、評価プロトコルにわたって高度に断片化されたままです。この断片化により、公正な比較が妨げられ、テキスト、ビジュアル、時間的、インタラクションベースのシグナルがどのように連携して人気のダイナミクスを形成するのかについての体系的な理解が曖昧になります。これらの課題に対処するために、標準化された評価プロトコルの下でデータセット、モダリティ、時間的相互作用シグナル、および代表的なベースラインを統合する、マルチモーダル グラフ ベースの人気予測ベンチマークである MMG-Pop を導入します。さらに、前述のマルチモーダル信号とグラフ構造の社会的相互作用を共同モデル化する統合マルチモーダル グラフベース ネットワークである MMG-PopNet を提案します。 Bluesky プラットフォームと Reddit プラットフォームにわたる 4 つのデータセットで構成される MMG-Pop に関する広範な実験により、MMG-PopNet の優れたパフォーマンスが実証され、クロスプラットフォーム トレーニングの一般化、マルチタスク予測の利点、マルチモダリティの寄与、LLM 予測の限界についての新たな洞察が得られました。これらの発見は、異種モダリティおよび社会を意識したエージェントエコシステムパラダイムの下での社会動態モデリングと介入に関する将来の研究のための統一基盤を確立します。
原文 (English)
Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction
Social media popularity prediction aims to forecast the future reach or influence of online content from early-stage observations. Accurate prediction enables key downstream applications, such as advertising optimization and strategic content planning by users, creators, and platforms. Despite substantial progress, existing popularity prediction works often fail to jointly consider multimodal content and temporal social interaction signals. Moreover, the literature remains highly fragmented across datasets, modalities, observation windows, prediction targets, and evaluation protocols. This fragmentation prevents fair comparison and obscures a systematic understanding of how textual, visual, temporal, and interaction-based signals jointly shape popularity dynamics. To address these challenges, we introduce MMG-Pop, a Multi-modal Graph-based Popularity Prediction benchmark, which unifies datasets, modalities, temporal interaction signals, and representative baselines under a standardized evaluation protocol. Furthermore, we propose MMG-PopNet, a unified multi-modal graph-based network that jointly models the aforementioned multi-modal signals and graph-structured social interactions. Extensive experiments on MMG-Pop, comprising four datasets across Bluesky and Reddit platforms, demonstrate the superior performance of MMG-PopNet and yield new insights into cross-platform training generalization, multi-task prediction benefits, multi-modality contributions, and LLM prediction limitation. These findings establish a unified foundation for future research on social dynamics modeling and intervention under heterogeneous modalities and socially-aware agentic ecosystem paradigms.
共有埋め込みシーケンスモデルにおける命令とデータの分離不可能性について
LLM 統合アプリケーションにとってプロンプト インジェクションは最大のセキュリティ リスクですが、これまでに提案された防御策はすべて破られています。これが偶然ではないことを証明します。制御データの分離が強制されていない共有埋め込みアーキテクチャでは、完全なプロンプト インジェクションの防止は数学的に不可能です。私たちは、プロンプト システムをプロンプト アクション モデルとして形式化します。その出力には、拒否の決定、ツールの承認、ポリシー ルーティング、メモリ書き込みなどの制御権限のあるアクションが含まれます。私たちは Semantic-Faithful Control (SFC) を定義します。これは、そのような動作が、信頼できない入力のエンコード方法ではなく、その意味のみに依存するという特性です。次に、SFC が共有パイプライン内で達成できないことを 3 つの結果を通じて証明します。来歴回復の不可能性 (共有表現により、信頼できるコンテンツと信頼できないコンテンツが統計的に分離できなくなり、合計変動距離によって制限されます)。制御パスの露出 (信頼できないトークンは、出力を決定する同じアテンション値の集計を通じて制御関連の計算に入ります)。および有限カバレッジの不変ギャップ (有限のトレーニングでは、無限の意味的等価クラスにわたる不変性を証明することはできません)。生産トークナイザーとモデルの測定で各量を粉砕します。その結果は構造的なものであり、現在の防御のギャップではありません。これは、バッファ オーバーフローを引き起こすフォン ノイマン マシンのコードとデータの混乱を反映しています。バッファ オーバーフローは、単一のメカニズムでは十分ではなかったために、何十年にもわたって多層防御 (DEP、書き込み-XOR-実行、ASLR、スタック カナリア、そして最終的にはメモリセーフ言語) で阻止する必要があった脆弱性クラスです。意味するところは同じです。パイプライン内の分類や調整を改善するだけでは、即時インジェクションを排除することはできません。命令チャネルとデータチャネルをアーキテクチャ的に分離する必要があります。私たちは根本原因とそれに必要な解決策の種類を特定します。
原文 (English)
On the Inseparability of Instructions and Data in Shared-Embedding Sequence Models
Prompt injection is the top security risk for LLM-integrated applications, yet every defense proposed so far has been broken. We prove this is not a coincidence: in shared-embedding architectures that lack enforced control-data separation, perfect prompt-injection prevention is mathematically impossible. We formalize prompted systems as Prompted Action Models whose outputs include control-authoritative actions: refusal decisions, tool authorization, policy routing, and memory writes. We define Semantic-Faithful Control (SFC), the property that such behavior depends only on the meaning of untrusted input, not on how it is encoded. We then prove SFC is unachievable within the shared pipeline, via three results: a provenance-recovery impossibility (shared representations make trusted and untrusted content statistically inseparable, bounded by total variation distance); control-path exposure (untrusted tokens enter control-relevant computation through the same attention value-aggregation that determines outputs); and a finite-coverage invariance gap (finite training cannot certify invariance over infinite semantic-equivalence classes). We ground each quantity in measurements on production tokenizers and models. The result is structural, not a gap in current defenses. It mirrors the code-data confusion in Von Neumann machines that gives rise to buffer overflows, a vulnerability class that took decades of layered defenses (DEP, Write-XOR-Execute, ASLR, stack canaries, and ultimately memory-safe languages) to contain, because no single mechanism sufficed. The implication is the same: prompt injection cannot be eliminated by better in-pipeline classification or alignment alone. It requires architectural separation of instruction and data channels. We identify the root cause and the class of solution it demands.
hia-gat: 高速道路でのフレームレベルの交通衝突リスク予測のための異種インタラクション対応グラフ アテンション ネットワーク
この論文では、フレーム レベルの高速道路リスク評価を、マルチ エージェント シーン グラフ レベルのバイナリ分類問題として定式化します。この問題では、TTC または PET ベースの競合が指定された重大度しきい値に違反する場合、各ビデオまたは軌跡フレームが危険とラベル付けされます。車両をノードとして、同一車線 (縦方向) と隣接車線 (横方向) の 2 つのインタラクション タイプをエッジとして使用し、フレームごとに関係を認識したグラフを構築します。これは、後端と車線変更の衝突メカニズムに合わせた物理学に基づいたエッジ特徴で強化されています。非グラフ モデルとグラフ ベースラインの構造化ベンチマーク スイートに基づいて、専用のアテンション パスウェイを通じて縦方向および横方向のインタラクションを処理し、SSM コンフリクト アトリビューションから派生したイベント レベルのゲート監視とコンフリクト タイプを認識したゲーティング メカニズムを介してそれらを融合する、デュアルストリーム異種グラフ アテンション ネットワークである HIA-GAT を提案します。 9 つの TTC および PET しきい値構成にわたる NGSIM I-80 および US-101 高速道路データセットの実験では、HIA-GAT が最高の平均リスク ランキング パフォーマンス (I-80 で AUC 0.835、US-101 で 0.867) を達成し、リレーショナル構造が不可欠な PET のみ (車線変更) 設定で最大の向上を示すことが示されています。学習されたゲートは、精度を超えて、主要な競合タイプの解釈可能な車両ごとの属性を提供し、実用的なリアルタイムの高速道路の安全監視をサポートします。横方向の紛争リスクをモデル化するにはグラフ構造が重要である一方、縦方向のリスクは非リレーショナル集計によって把握できることが多いことを示します。
原文 (English)
hia-gat: A Heterogeneous Interaction-Aware Graph Attention Network For Frame-Level Traffic Conflict Risk Prediction On Freeways
This paper formulates frame-level freeway risk assessment as a multi-agent scene graph-level binary classification problem, where each video or trajectory frame is labeled risky if any TTC- or PET-based conflict violates a specified severity threshold. We construct a relation-aware graph per frame with vehicles as nodes and two interaction types as edges: same-lane (longitudinal) and adjacent-lane (lateral), augmented with physics-informed edge features aligned to rear-end and lane-change conflict mechanisms. Building on a structured benchmarking suite of non-graph models and graph baselines, we propose HIA-GAT, a dual-stream heterogeneous graph attention network that processes longitudinal and lateral interactions through dedicated attention pathways and fuses them via a conflict-type-aware gating mechanism with event-level gate supervision derived from SSM conflict attribution. Experiments on the NGSIM I-80 and US-101 freeway datasets across nine TTC and PET threshold configurations show that HIA-GAT achieves the best average risk-ranking performance (AUC 0.835 on I-80 and 0.867 on US-101), with the largest gains on PET-only (lane-change) settings where relational structure is essential. Beyond accuracy, the learned gate provides interpretable per-vehicle attribution of dominant conflict type, supporting actionable, real-time freeway safety monitoring. We show that graph structure is critical for modeling lateral conflict risk, while longitudinal risk can often be captured by non-relational aggregation.
PEBS: RLHF 報酬モデル校正のための評価者ごとの経験的ベイズ収縮
Reinforcement Learning from Human Feedback (RLHF) の報酬モデルは、何千人ものアノテーターの好みをプールし、1 つのグローバル アフィン キャリブレーターに適合させ、系統的に異なる評価スケールのオフセットと傾きを持つ評価者を、個々のアノテーターのいずれにも一致しない単一の平均評価者適合にまとめます。 PEBS は、評価者ごとの経験的ベイズ収縮推定器です。各アノテーターの評価の保持されたスライスに評価者ごとのアフィン キャリブレーターを適合させ、閉じた形式で、報酬モデルを再トレーニングすることなく、モリス・ジェームズ・スタインの経験的ベイズ収縮を母平均に適用します。 PRISM では、PEBS はユーザー内で保持される RMSE を、プールされた人口勾配ベースラインより 8.58% 削減します。この手順は、PluriHarms の危害評価 (Qwen-2.5 ベース、ファミリー内) に基づいて再現され、同じ人口勾配ベースラインに対して +9.66% RMSE 減少を示します。 PEBS は、RLHF 報酬モデリングにおけるアノテーター固有のアフィン キャリブレーション用の閉じた形式の事後推定器です。報酬ベースモデルは変更されず、新しい評価の推論時に使用される評価者レベルのマップのみが推定されます。
原文 (English)
PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration
Reward models for Reinforcement Learning from Human Feedback (RLHF) pool preferences across thousands of annotators and fit one global affine calibrator, collapsing raters with systematically different rating-scale offsets and slopes into a single average-rater fit that does not match any individual annotator. PEBS is a per-rater empirical-Bayes shrinkage estimator: it fits per-rater affine calibrators on a held-out slice of each annotator's ratings and applies Morris-James-Stein empirical-Bayes shrinkage toward the population mean, in closed form and without retraining the reward model. On PRISM, PEBS reduces within-user held-out RMSE by 8.58% over the pooled population-slope baseline. The procedure replicates on PluriHarms harm ratings (Qwen-2.5 base, in-family) with a +9.66% RMSE reduction over the same population-slope baseline. PEBS is a closed-form post-hoc estimator for annotator-specific affine calibration in RLHF reward modeling; it leaves the reward base model unchanged and estimates only the rater-level map used at inference time for new ratings.
NSCLC における腫瘍割合スコアリングのための分布ベースの深層複数インスタンス学習
非小細胞肺がん (NSCLC) における腫瘍割合スコア (TPS) を正確に評価することは、治療計画と予後にとって重要です。主な課題としては、各スライドに注釈を付けるのに必要な面倒な手作業と、このタスクに認定された専門家の数が限られていることが挙げられます。マルチ インスタンス学習 (MIL) は、スライド レベルで TPS スコアを予測するための効果的なアプローチであることが証明されています。ただし、既存の方法は、表現力のない (クラス 0) イメージを処理するのに苦労します。私たちのアプローチには 2 つのモデルが含まれます。(1) 個々のパッチの組織病理学的特徴を捕捉する包埋抽出およびマルチクラス分類ネットワーク、(2) これらの包埋を集約してスライド全体の全体的な TPS 確率分布を表すゼロインフレート ベータ (ZIBeta) パラメーターを予測する MIL モデル。スライドレベルの TPS スコアのみをラベルとして使用し、このエンドツーエンドのフレームワークが新しい分布ベースのアーキテクチャを活用して予測精度と説明可能性を向上させる方法を示します。 ZIBeta モデリングは、ベースライン線形回帰およびリッジ回帰を大幅に上回り、分布集中を通じて期待される精度を捕捉します。
原文 (English)
Distribution-based deep multiple instance learning for tumor proportion scoring in NSCLC
Accurate assessment of tumor proportion score (TPS) in non-small cell lung cancer (NSCLC) is critical for treatment planning and prognosis. Key challenges include the tedious manual work required to annotate each slide, combined with the limited number of experts certified for this task. Multiple instance learning (MIL) has proven to be an effective approach for predicting TPS scores at the slide level; however, existing methods struggle with non-expressive (zero class) images. Our approach involves two models: (1) an embedding-extraction and multiclass-classification network that captures the histopathological features of individual patches, and (2) a MIL model that aggregates these embeddings to predict zero-inflated beta (ZIBeta) parameters representing the overall TPS probability distribution for the entire slide. Using only slide-level TPS scores as labels, we demonstrate how this end-to-end framework can leverage a novel distribution-based architecture to improve prediction accuracy and explainability. ZIBeta modeling significantly outperforms baseline linear and ridge regression while capturing expected accuracy through distribution concentration.
遡及的アドバンテージ補正: 遅延を考慮した RLHF のクローズドフォーム V トレース バイアス補正
本番環境における人間のフィードバックからの強化学習 (RLHF) には、常に同期した報酬信号があるとは限りません。コード実行ベリファイア、遅いジャッジアンサンブル、およびキューに入れられた人間によるレビューは、それらを生成したロールアウト後にいくつかの勾配ステップを返す可能性があり、標準 PPO の基礎となる同期報酬の前提が崩れます。このギャップは、Retroactive Advantage Correction (RAC) で解決します。保留中の遅い完了はそれぞれキューに入れられ、非負のカーネルを通じてエージングされ、クリップされた残差として次のオプティマイザー ステップのアドバンテージに再注入されます。不偏のクリップされた重要度比の下では、有効遅延カーネルがその質量のすべてを再注入する場合、累積 RAC 補正は正確に不偏であり、それ以外の場合は再注入されない部分に線形のバイアスがかかることを証明します。遅延のないアイデンティティ カーネルでは、V-trace に減少します。表形式のマルコフ決定プロセス (MDP) の概念実証では、RAC は 2 つの低速チャネル構成でクローズド形式のポリシー バイアスを最大 47.9 倍削減し、より低い実時間コストで低速待機を上回りました。 RAC は、2 行の報酬マネージャー パッチを通じて PPO および GRPO と統合されます。
原文 (English)
Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF
Reinforcement learning from human feedback (RLHF) in production does not always have a synchronous reward signal. Code-execution verifiers, slow judge ensembles, and queued human review can return several gradient steps after the rollout that produced them, breaking the synchronous-reward assumption underlying standard PPO. We address this gap with Retroactive Advantage Correction (RAC): each pending slow completion is queued, aged through a non-negative kernel, and reinjected as a clipped residual into the next optimiser step's advantage. We prove that under an unbiased clipped importance ratio, the cumulative RAC correction is exactly unbiased when the effective delay kernel reinjects all of its mass, and carries a bias linear in the unreinjected fraction otherwise; at the no-delay identity kernel it reduces to V-trace. On a tabular Markov decision process (MDP) proof-of-concept, RAC reduces the closed-form policy bias by up to 47.9x at the two-slow-channel configuration, beating wait-for-slow at lower wall-clock cost. RAC integrates with PPO and GRPO through a two-line reward-manager patch.
SceneBot: シーンインタラクションを使用した、接触を促す一般的なヒューマノイド全身追跡
現在の人型強化学習ポリシーは、自由空間での動作には優れていますが、純粋な運動学的追跡では物体や凹凸のある地形との相互作用による物理的な曖昧さを解決できないため、接触が多いタスクには苦労しています。これに対処するために、自由空間移動、地形横断、全身操作を処理できる統合モーション追跡フレームワークである SceneBot を導入します。 SceneBot は、参照モーションとリンクごとの接触ラベルの両方に単一のポリシーを条件付けし、予想される環境相互作用を明示的に定義します。注釈付きインタラクション データの欠如を克服するために、リターゲットされた人間の動きからシーン インタラクション グラフを推測する後知恵シーン再構築アプローチを提案します。この再構築された接触が豊富なデータの 7.5 時間に基づいてトレーニングされた SceneBot は、目に見えない動きや環境をうまく一般化します。私たちの結果は、SceneBot が、箱を 2 階に運ぶ、ヒューマノイド制御のための強力なインターフェースとして接触条件付けを確立するなど、複雑で長期にわたるタスクを実行する、自由空間と接触の多い動作をシームレスに統合する最初の一般的なフレームワークであることを示しています。すべてのコードとデータはオープンソースになります。その他のデモと情報は、https://ericcsr.github.io/scenebot/ で入手できます。
原文 (English)
SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction
Current humanoid reinforcement-learning policies excel at free-space motions but struggle with contact-rich tasks, as pure kinematic tracking cannot resolve the physical ambiguities of interacting with objects and uneven terrain. To address this, we introduce SceneBot, a unified motion-tracking framework capable of handling freespace locomotion, terrain traversal, and whole-body manipulation. SceneBot conditions a single policy on both reference motions and per-link contact labels, explicitly defining expected environmental interactions. To overcome the lack of annotated interaction data, we propose a hindsight scene reconstruction approach that infers scene-interaction graphs from retargeted human motion. Trained on 7.5 hours of this reconstructed, contact-rich data, SceneBot successfully generalizes to unseen motions and environments. Our results demonstrate that SceneBot is the first general framework to seamlessly unify free-space and contact-rich behaviors executing complex, long-horizon tasks like carrying a box upstairs and establishing contact conditioning as a powerful interface for humanoid control. All code and data will be open-sourced. More demos and information are available at: https://ericcsr.github.io/scenebot/
CoIn: ガウス スプラッティング ガイダンスによる包括的な 2D-3D 修復
3D シーンの修復は、オクルージョンや視点の制限によって破損した領域を再構築するために不可欠です。最近の手法では効率的な 3D 編集のためにガウス スプラッティング (GS) を活用していますが、多くの場合、正確なマルチビュー セグメンテーション マスクに依存しており、本質的にオブジェクトの削除タスクに制約されます。私たちは、マルチステージの整合性パイプラインを通じて 2D 修復モデルと 3DGS の橋渡しをする新しいフレームワークである CoIn を提案します。私たちのアプローチでは、まず拡散モデルを使用して初期修復イメージを生成し、任意の形状のマスクやオブジェクト挿入などのさまざまなタスクの使用を可能にします。次に、参照ビュー (2D -> 3D) に向けて適応的に重み付けすることにより、粗い 3D シーンを再構築する、フィーチャ アテンションを備えた参照適応 GS を導入します。この 3D 表現は、GS ベースの参照フィーチャ ワーピングを介して拡散プロセスに幾何学的ガイドを提供し、マルチビューの一貫性 (3D -> 2D) を保証します。最後に、Texture-Enhancing Discriminator が 3D シーンを洗練して、高いフォトメトリック リアリズム (2D -> 3D) を実現します。実験の結果、CoIn は双方向の情報フローを効果的に活用して最先端のパフォーマンスを実現し、柔軟なマスク入力でオブジェクトの削除とオブジェクトの挿入の両方を効果的に処理できることがわかりました。
原文 (English)
CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance
3D scene inpainting is essential for reconstructing areas corrupted by occlusions or limited viewpoints. While recent methods leverage Gaussian Splatting (GS) for efficient 3D editing, they often depend on precise multi-view segmentation masks and are inherently constrained to object removal tasks. We propose CoIn, a novel framework that bridges 2D inpainting models and 3DGS through a multi-stage consistency pipeline. Our approach first generates initial inpainted images using a diffusion model, enabling the use of arbitrary-shaped masks and diverse tasks like object insertion. We then introduce Reference Adaptive GS with Feature Attention to reconstruct a coarse 3D scene by adaptively weighing towards a reference view (2D -> 3D). This 3D representation provides geometric guidance to the diffusion process via GS-based Reference Feature Warping, ensuring multi-view consistency (3D -> 2D). Finally, a Texture-Enhancing Discriminator refines the 3D scene to achieve high photometric realism (2D -> 3D). Experiments show that CoIn, effectively leveraging bidirectional information flow, achieves state-of-the-art performance and effectively handles both object removal and object insertion with flexible mask input.
病理学的ショートカットの解体: 忠実な LVLM デコードのための因果関係のフレームワーク
大規模視覚言語モデル (LVLM) は洗練された推論を示しますが、依然として物体の幻覚の影響を受けやすいです。一般的な注意強度の仮定から逸脱すると、より深い動的な構造のずれが明らかになります。つまり、幻覚は、特定の注意が危険な仲介者として機能し、視覚的証拠から切り離して言語の事前情報を固定する、意思決定が重要な段階で引き起こされます。これにより、視覚的なグラウンディングを回避する病的なショートカットが確立されます。これを解体するために、トレーニング不要の推論時間フレームワークである Fox (Faithhood and Observational-flow via eXpression-rectification) を提案します。 Fox は、視覚的注意エントロピー プローブを使用して構造のミスアラインメントを診断し、危険なメディエーターの位置を監視なしで特定します。次に、数値ロジット飽和を介してターゲットを絞った因果的介入を実行し、ショートカット パスを物理的に切断します。最後に、競合ゲート型の協力的解読戦略は、介入の忠実さと観察の流暢性を調和させます。広範な実験により、Fox は言語の豊かさを維持しながら SID を 29.1% 上回る SOTA パフォーマンスを達成することが実証されました。コードは https://github.com/Cc2021start/Fox で入手できます。
原文 (English)
Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding
Large Vision-Language Models (LVLMs) exhibit sophisticated reasoning but remain susceptible to object hallucination. Deviating from the prevailing attention intensity assumption, we reveal a deeper dynamic structural misalignment: hallucination is triggered at decision-critical steps where specific attention heads, acting as risky mediators, decouple from visual evidence to lock onto language priors. This establishes a pathological shortcut that bypasses visual grounding. To dismantle this, we propose Fox (Faithfulness and Observational-flow via eXpression-rectification), a training-free inference-time framework. Fox diagnoses structural misalignment using a visual attention entropy probe to localize risky mediators unsupervisedly. We then execute a targeted causal intervention via numerical logit saturation to physically sever the shortcut path. Finally, a conflict-gated cooperative decoding strategy reconciles interventional faithfulness with observational fluency. Extensive experiments demonstrate that Fox achieves SOTA performance, outperforming SID by 29.1% while preserving linguistic richness. Code is available at https://github.com/Cc2021start/Fox.
Narrative-UFET: 超微細エンティティ タイピングのためのナラティブ生成
Ultra Fine Entity Typeing (UFET) は、エンティティの言及に非常に具体的な型を割り当てますが、現在のアプローチはロングテールの型に苦労しています。私たちは、曖昧さを排除する証拠が複数の文にまたがっていることが多いため、重要な制限は文レベルのコンテキストに依存していることであると仮説を立てています。既存の UFET リソースはすべて文レベルであるため、これをテストするのは困難でした。我々は、UFET の制御された拡張である Narrative-UFET を紹介します。この拡張では、各エンティティの言及が、自動的に生成された短く一貫したナラティブと組み合わされます。物語を合成すると、特定の談話特性の影響を分離できます。私たちは 2 つのペアのバリアントを実験します。1 つはエンティティのタイプが物語全体で一定に保たれる (Maintain) もので、もう 1 つはそれがシフトする (Change) ものです。ナラティブコンテキストは、文レベルのベースラインよりもロングテールタイプで一貫した改善をもたらし、Change バリアントがより強いシグナルを提供することを示します。自然に発生する文脈との比較では、合成ナラティブがより強力な利益を生み出すことが示されており、制御された談話構築により、実際のテキストが暗黙的に残したシグナルを表面化できることが示されています。実質的な改善の余地が残されており、談話モデリングと物語構築の両方においてオープンな方向性が示唆されています。
原文 (English)
Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing
Ultra-fine entity typing (UFET) assigns highly specific types to entity mentions, but current approaches struggle with types in the long tail. We hypothesize that a key limitation is the reliance on sentence-level context, since disambiguating evidence is often spread across multiple sentences. Testing this has been difficult because all existing UFET resources are sentence-level. We present Narrative-UFET, a controlled extension of UFET in which each entity mention is paired with an automatically generated short, coherent narrative. Synthesizing narratives lets us isolate the effect of specific discourse properties. We experiment with two paired variants: one in which the entity's type is held constant across the narrative (Maintain) and one in which it shifts (Change). We show that narrative context yields consistent improvements on long-tail types over sentence-level baselines, with the Change variant providing the stronger signal. A comparison against naturally occurring contexts shows that synthetic narratives yield stronger gains, indicating that controlled discourse construction can surface signals that real text leaves implicit. Substantial room for improvement remains, suggesting open directions in both discourse modeling and narrative construction.
$K$ オーダーのマルコフ近似による多変量時系列予測モデルの全体的な説明
多くの説明可能な AI (XAI) 手法が提案されていますが、そのほとんどは時系列予測モデル用に設計されておらず、多くの場合、タイムスタンプ特徴が独立しているという暗黙の仮定に依存しています。この仮定は時間依存性の基本的な特性を無視しており、データの順序構造と因果構造に違反する説明につながる可能性があります。 \textsc{KARMA} を紹介します。これは、予測子によって学習された時間依存関係を捉えるマルコフ代理モデルを構築することによって、時系列予測子を説明する方法です。私たちのアプローチは、モデルにとって予測的に十分な最小履歴長 $K$ の特定、離散化された履歴空間からの最適な $K$ 次マルコフ遷移カーネルの推定、マルコフ遷移カーネルから導出できる 5 レベルのグローバル説明階層 (現実世界の気象データ (北京 PM 2.5) を使用して説明) の 3 つの主な側面を中心に展開します。また、既知の真の因果エッジを持つ複雑な合成データを使用することで、KARMA が、(i) 制御された実験を通じてモデルによって学習されたデータの因果構造を回復し、(ii) TimeSHAP などの確立された帰属手法よりも適切に時間依存関係を特定できることも証明します。
原文 (English)
Global Explanations for Multivariate Time Series Forecasting Models via $K$-Order Markov Approximations
While many explainable AI (XAI) methods have been proposed, most are not designed for time-series forecasting models and often rely on the implicit assumption that timestamp features are independent. This assumption ignores the fundamental property of temporal dependence and can lead to explanations that violate the sequential and causal structure of the data. We introduce \textsc{KARMA}, a method for explaining time-series predictors by constructing a Markov surrogate model that captures the temporal dependencies learned by the predictor. Our approach revolves around three main aspects: identifying the minimal history length $K$ that is predictively sufficient for the model, estimating the best-fitting $K$-order Markov transition kernel from the discretized history space, and a five-level global explanation hierarchy that can be derived from the Markov transition kernel, which we illustrate using real-world weather data (Beijing PM 2.5). We also certify using complex synthetic data with known true causal edges that KARMA (i) recovers the data causal structure as learned by the model via a controlled experiment and (ii) identifies temporal dependencies better than established attribution methods such as TimeSHAP.
HybridCodec: 効率的な音声言語モデルのための離散表現と連続表現のモデリング
個別オーディオ表現は、マルチモーダル テキスト オーディオ システムを構築し、オーディオ機能を大規模言語モデル (LLM) に統合するためにますます一般的になってきています。ただし、多くの研究では、離散化中の情報損失によるさまざまな下流タスクのパフォーマンスの低下が報告されています。これに対処するために、時間的に圧縮された離散トークンと次元を削減した連続残差を組み合わせた新しいアプローチを提案します。私たちのフレームワークは、ハイブリッド化された離散連続焦点変調コーデックとハイブリッド Transformer で構成されています。このアーキテクチャは、非自己回帰予測および連続残差アップサンプリングと組み合わせて、離散領域で自己回帰推論を実行します。実験結果は、私たちのアプローチが離散のみの方法と比較してスピーカー特性の保持を大幅に改善し、同時に必要な自己回帰ステップの数を削減することを示しています。
原文 (English)
HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models
Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilities into Large Language Models (LLMs). However, numerous studies report performance degradation on various downstream tasks due to information loss during discretization. To address this, we propose a novel approach combining temporally compressed discrete tokens with dimensionality-reduced continuous residuals. Our framework consists of a hybridized discrete-continuous focal modulation codec and a hybrid Transformer. This architecture performs autoregressive inference in the discrete domain, coupled with non-autoregressive prediction and continuous residual upsampling. Experimental results show that our approach significantly improves the retention of speaker characteristics compared to discrete-only methods, while simultaneously reducing the number of required autoregressive steps.
デュアルしきい値のハード サンプル マイニングによるクロスプラットフォームの中国の攻撃的なコメントの検出
中国のソーシャルメディア向けに攻撃的なコメント検出をクロスプラットフォームに展開すると、パフォーマンスが低下します。この論文では、これに対処するためにデュアルしきい値ハードマイニング手法を提案しています。まず、クリーンな中国語ベースの RoBERTa を COLD で微調整して、公正な比較のためのバイナリ ベースラインを確立します。次に、Weibo、Xiaohongshu、Tieba、および Zhihu をカバーする 3 クラスの詳細ラベル付きテスト セットが構築され、ソースからのドメイン距離が Jaccard および Proxy-A Distance を使用して定量化され、ドメイン シフト時のベースラインの劣化ボトルネックが系統的に明らかにされます。ここでは、二重閾値のハード サンプル マイニング戦略が提案されています。信頼性の高いエラーが発生しやすいサンプルと信頼性の低いサンプルは、予測の信頼性によってラベルのないコーパスからフィルタリングされます。モデルは、手動でラベル付けされたハード サンプルの少数のセットのみを使用して、暗黙的なコンテキストの下で二次的に微調整され、低コストのクロスプラットフォーム ドメイン適応を実現します。実験により、4 つのプラットフォームにわたって最適化されたモデルのパフォーマンスが大幅に向上したことが明らかになりました。
原文 (English)
Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining
Cross-platform deployment of offensive comment detection for Chinese social media suffers performance degradation. The paper proposes a dual-threshold hard mining method to address this. First, the clean-Chinese-base RoBERTa is finetuned on COLD to establish a binary baseline for fair comparison. Second, a three-class fine-labeled test set covering Weibo, Xiaohongshu, Tieba, and Zhihu is constructed, domain distances from the source are quantified using Jaccard and Proxy-A Distance, as well as the degradation bottleneck of the baseline under domain shift is systematically revealed. Herein, a dual threshold hard example mining strategy is proposed. High- and low-confidence error-prone samples are filtered from unlabeled corpora by prediction confidence. The model is secondarily finetuned under implicit contexts with merely a small set of manually labeled hard examples, realizing low-cost cross-platform domain adaptation. Experiments reveal significant performance gains of the optimized model across four platforms.
単細胞 RNA シーケンスを使用したヒト脂肪組織における脂肪細胞の発生軌跡の再構築
肥満は、2 型糖尿病や心血管疾患などの代謝障害に関連した世界的な健康危機です。この研究では、単一細胞 RNA シーケンスを使用して、脂肪組織サンプルからヒト脂肪細胞の発生軌跡を再構築しました。私たちの分析により、7 つの遷移状態を含む 15 の転写的に異なる細胞クラスターが特定され、脂肪細胞分化の動的なプロセスが明らかになりました。我々は、脂肪細胞とその前駆細胞の間の細胞コミュニケーションを媒介する、機能的に活性なシグナル伝達経路を16個検出した。これらの中で、インスリン様成長因子 (IGF) および線維芽細胞成長因子 (FGF) 経路が最も顕著なネットワークとして浮上し、分化段階全体で一貫した活性を示しました (p<0.05)。この研究では、内臓脂肪細胞が皮下分化には存在しない追加の細胞外マトリックスのリモデリングを受けるという、デポー特異的な違いが明らかになりました。さらに、空間解析では、IGF シグナル伝達が血管周囲ニッチで特に活性であるのに対し、FGF 活性は成熟脂肪細胞ゾーンで優勢であることが示されました。これらの結果は、潜在的な治療標的としてのIGFおよびFGF経路を強調する、ヒト脂肪細胞発生の最初の包括的なマップを提供する。同定されたシグナル伝達ネットワークは、健康な脂肪の拡大を促進したり、病的な脂肪の蓄積を抑制したりするための介入を開発するための新たな洞察を提供します。この研究は、代謝障害の治療に臨床的に関連するデータを提供しながら、脂肪組織生物学の基本的な理解を促進します。
原文 (English)
Reconstructing the Developmental Trajectory of Adipocytes in Human Adipose Tissue Using Single-Cell RNA Sequencing
Obesity is a global health crisis associated with metabolic disorders such as type 2 diabetes and cardiovascular disease. This study employed single-cell RNA sequencing to reconstruct the developmental trajectory of human adipocytes from adipose tissue samples. Our analysis identified 15 transcriptionally distinct cell clusters, including 7 transitional states, revealing the dynamic process of adipocyte differentiation. We detected 16 functionally active signaling pathways mediating cellular communication between adipocytes and their progenitors. Among these, insulin-like growth factor (IGF) and fibroblast growth factor (FGF) pathways emerged as the most prominent networks, showing consistent activity across differentiation stages (p<0.05). The study revealed depot-specific differences, with visceral adipocytes undergoing additional extracellular matrix remodeling absent in subcutaneous differentiation. Spatial analysis further showed that IGF signaling was particularly active in perivascular niches, while FGF activity dominated in mature adipocyte zones. These results provide the first comprehensive map of human adipocyte development, highlighting IGF and FGF pathways as potential therapeutic targets. The identified signaling networks offer new insights for developing interventions to promote healthy adipose expansion or inhibit pathological fat accumulation. This work advances our fundamental understanding of adipose tissue biology while providing clinically relevant data for metabolic disorder treatments.
生物多様性モニタリングと生態画像解析のための説明可能な AI
人工知能は、カメラトラップ、ドローン、衛星、水中プラットフォーム、その他のセンシングシステムから収集された生態画像の自動分析を可能にし、生物多様性モニタリングを変革しています。これらのツールは保全評価の規模と速度を拡大できますが、多くのコンピュータ ビジョン モデルは依然として検査が困難であり、予測が生態学的に意味のある信号に基づいているのか、それとも偽の相関、サンプリング バイアス、および保全の決定を損なう可能性のあるその他のアーチファクトに基づいているのかを判断することが困難になっています。私たちは、説明可能な人工知能 (XAI) が生態学的モデル検証の標準コンポーネントになるべきであると主張します。なぜなら、自然保護の実践者は、モデルが正確かどうかだけでなく、モデルが正確である理由を理解することにますます依存しているからです。 XAI を、画像分類、物体検出、画像セグメンテーションという 3 つの一般的な生態学的コンピュータ ビジョン タスクに適用するための実践的なガイダンスを提供します。 XAI が生態モデルの監査、改良、展開をどのようにサポートできるかを説明するために、航空画像を使用した 2 つのケース スタディ、ゴマフアザラシの検出とクジラ類の解剖学的セグメンテーションを紹介します。これらの例は、説明手法がどのように生物学的に意味のある手がかりを特定し、背景と形状の交絡によって引き起こされる偽陽性を明らかにし、エッジとオクルージョンの効果を明らかにし、データ収集、拡張、および再トレーニング戦略を導く方法を示しています。さらに広く言えば、説明可能性がモデル推論が生態学的理解と一致しているかどうかを評価するのにどのように役立つかを示しています。最後に、主要な課題と機会を特定します。 XAI は、モデルの行動の透明性を高め、科学的に質問できるようにすることで、AI がサポートする生態学的証拠がより信頼性が高く、理解しやすく、生物多様性の保全にとって実用的なものであることを保証します。
原文 (English)
Explainable AI for Biodiversity Monitoring and Ecological Image Analysis
Artificial intelligence is transforming biodiversity monitoring by enabling automated analysis of ecological imagery collected from camera traps, drones, satellites, underwater platforms, and other sensing systems. These tools can expand the scale and speed of conservation assessments, yet many computer vision models remain difficult to inspect, making it challenging to determine whether predictions are based on ecologically meaningful signals or on spurious correlations, sampling biases, and other artifacts that may undermine conservation decisions. We argue that explainable artificial intelligence (XAI) should become a standard component of ecological model validation because conservation practitioners increasingly depend on understanding not only whether a model is accurate, but why it is accurate. We provide practical guidance for applying XAI to three common ecological computer vision tasks: image classification, object detection, and image segmentation. To illustrate how XAI can support ecological model auditing, refinement, and deployment, we present two case studies using aerial imagery: harbor seal detection and cetacean anatomical segmentation. These examples demonstrate how explanation methods can identify biologically meaningful cues, reveal false positives driven by background and shape confounds, uncover edge and occlusion effects, and guide data collection, augmentation, and retraining strategies. More broadly, they show how explainability can help assess whether model reasoning aligns with ecological understanding. We conclude by identifying key challenges and opportunities. By making model behavior more transparent and scientifically interrogable, XAI can help ensure that AI-supported ecological evidence is more reliable, understandable, and actionable for biodiversity conservation.
信号から転送まで: 大規模言語モデルにおけるプローブベースの不確実性推定の因数分解された研究
プローブベースの不確実性推定 (UE) は、内部モデル信号から不確実性を学習することで大規模言語モデル (LLM) の幻覚を検出するための優れたアプローチとして浮上しています。しかし、最近の手法は機能設計、トレーニング データ構築、評価設定にわたって同時に変化しており、実際に何がパフォーマンスを促進するのかが不明瞭になっています。この問題に対処するために、一致した条件下でのプローブベースの UE の因数分解された研究を提案します。私たちの結果は、生の隠れ状態とアテンション機能がドメイン内で優れたパフォーマンスを発揮するのは難しいことを示しています。ただし、分散シフトの下では、構造化および圧縮された機能がより堅牢になり、ドメイン内のパフォーマンスだけでは進歩を測定するには不十分であることが示唆されています。さらに、プロンプトとラベルの構築はプローブの動作に大きく影響します。これらのベストプラクティスの発見に基づいて、オープンエンドの事実生成に適度に移行するベンチマークベースの事前トレーニング済みプローブをトレーニングし、安定した既製のベースラインを提供します。私たちの取り組みは、プローブベースの不確実性推定器の、より展開指向の評価を奨励しています。コード リポジトリは https://github.com/ponhvoan/ProbeUE で入手できます。
原文 (English)
From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models
Probe-based uncertainty estimation (UE) has emerged as a prominent approach to detect hallucinations in Large Language Models (LLMs) by learning uncertainty from internal model signals. Yet, recent methods vary simultaneously across feature design, training data construction, and evaluation setting, obscuring what actually drives performance. To address this issue, we propose a factorised study of probe-based UE under matched conditions. Our results show that raw hidden states and attention features are difficult to outperform in-domain. However, under distribution shift, structured and compressed features are more robust, suggesting that in-domain performance alone is insufficient to measure progress. Furthermore, prompting and label construction significantly affect probe behaviour. Building on these best-practice findings, we train benchmark-based pretrained probes that transfer reasonably well to open-ended factual generation, providing a stable off-the-shelf baseline. Our work encourages more deployment-oriented evaluation of probe-based uncertainty estimators. The code repository is available at https://github.com/ponhvoan/ProbeUE.
CBD: 制御された行動の発散による API のみの LLM ブラック ボックスの学習解除
エッジ デバイスは、コンテキスト認識エッジ インテリジェンスのために API サービスを通じて大規模言語モデル (LLM) を呼び出すことが増えていますが、エッジで生成されたデータは、LLM を改善するために収集される可能性があり、機密情報、著作権で保護された情報、有害な情報、または古い情報がモデルの動作に導入される可能性があります。機械のアンラーニングは、LLM を再トレーニングすることなく、望ましくないデータの影響を除去する実用的な方法を提供します。しかし、既存の方法には依然として 2 つのギャップがあります。 1 つ目は API のみのブラック ボックス アクセスで、ターゲット モデルのパラメーターと内部ロジットは利用できません。 2 つ目は、学習解除ターゲット データと保持データが非常に類似したプロンプト構造またはセマンティック パターンを共有する場合に、保持されたユーティリティをどのように保持するかです。これらの課題に対処するために、API のみのブラック ボックスの学習解除フレームワークである Controlled Behavioral Divergence (CBD) を提案します。 CBD は 2 つの補助モデルを使用して、保持された入力と未学習ターゲット入力の間で制御された動作の相違を作成し、この相違を未学習関連性スコアに変換し、未学習関連プロンプトをターゲット LLM から遠ざけるようにルーティングします。ターゲットと保持データ間の類似性が高い場合の識別精度を向上させるために、CBD は経験的なフィッシャー行列を推定し、正則化された一般化固有値問題を解くことによって勾配統計ベースの識別基盤を構築し、共有プロンプト構造ではなくターゲット固有の情報に向けて非学習信号を誘導します。 11 個のホワイト ボックスおよびグレー ボックスのアンラーニング ベースラインと比較して、CBD はより優れたアンラーニング ユーティリティのトレードオフを達成しており、そのパフォーマンスは設定間でほとんど変化しません。 ToFU remember10 では、CBD はモデルの有用性を 74.90 に上げながら、フォーゲット セットの再トレーニングされたリファレンスに近づき、2 番目に良いベースラインを約 15% 上回っています。 WMDP では、MMLU 精度 52.67 を維持しながら、危険な知識の精度を 25.68 まで下げ、ほぼランダムな推測になります。コードは https://github.com/DGL-codes/CBD にあります。
原文 (English)
CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence
Edge devices increasingly invoke large language models (LLMs) through API services for context aware edge intelligence, while edge generated data may be collected to improve LLMs and may introduce sensitive, copyrighted, harmful, or outdated information into model behavior. Machine unlearning offers a practical way to remove the influence of undesired data without retraining LLMs. However, existing methods still face two gaps. The first is API only black box access, where target model parameters and internal logits are unavailable. The second is how to preserve retained utility when unlearning target data and retained data share highly similar prompt structures or semantic patterns. To address these challenges, we propose Controlled Behavioral Divergence (CBD), an API only black box unlearning framework. CBD uses two auxiliary models to create controlled behavioral divergence between retained inputs and unlearning target inputs, converts this divergence into an unlearning relevance score, and routes unlearning related prompts away from the target LLM. To improve discrimination accuracy under high similarity between target and retained data, CBD constructs a gradient statistics based discriminative basis by estimating empirical Fisher matrices and solving a regularized generalized eigenvalue problem, guiding the unlearning signal toward target specific information rather than shared prompt structures. Compared with eleven white box and gray box unlearning baselines, CBD achieves a better unlearning utility trade off and its performance varies little across settings. On ToFU forget10, CBD approaches the retrained reference on the forget set while raising model utility to 74.90, about 15% above the second best baseline. On WMDP, it lowers hazardous knowledge accuracy to 25.68, near random guessing, while preserving MMLU accuracy of 52.67. Code is at https://github.com/DGL-codes/CBD.
次の LLM の事前登録による LLM ベースの p-Hacking の軽減
大規模言語モデル (LLM) は、その出力が下流の仮説テストにフィードされるデータの生成、分類、および注釈付けに使用されることが増えています。ただし、LLM ベースの研究は簡単にハックできます。研究者は、目的の結果が得られるまでプロンプト、デコード パラメータ、または出力形式を調整できます。私たちは、LLM ベースの研究における p-hacking を軽減するためのプロトコルを提案します。実験と適格なモデルを事前登録し、事前登録後にリリースされる最初の適格な LLM 上で実行します。研究者は、現在のモデルに関する手順を最終決定し、一連の適格な将来のモデルとともに解析計画を事前登録し、その後リリースされる最初の適格なモデルに対して確認分析を実行します。このモデルはコミット時には存在しないため、ハッキングすることはできません。さらに、あるモデルを頻繁にハッキングする構成は、次のモデルには移行されません。真の値がわかっている 2 つのタスクでプロトコルを評価します。 4 つのプロバイダーの 20 のモデルと 11 の LLM 解析構成にわたって、プロトコルは 2 つのタスクのケースの 73.9% と 72.7% で p-hack の転送成功をブロックしたと考えられます。追加の分析により、いくつかのストレステストの下でも緩和効果が依然として大幅であることが明らかになりました。最後に、私たちはお金をかけて、独自のプロトコルに従い、実験を事前登録しました。事前に登録された実験では、プロトコルの有効性が確認されました。前のモデルをハッキングした 7 つの構成のうち、その後リリースされた最初の適格モデルの 6 つの構成ではハッキングが引き継がれませんでした。
原文 (English)
Mitigating LLM-based p-Hacking by Preregistering for the Next LLM
Large language models (LLMs) are increasingly used to generate, classify, and annotate data whose outputs feed downstream hypothesis tests. However, LLM-based research is easy to p-hack: a researcher can tune the prompts, decoding parameters, or output format until a desired result is reached. We propose a protocol to mitigate p-hacking in LLM-based research: preregistering the experiment and eligible models, and then running it on the first eligible LLM that is released after the preregistration. The researcher finalizes the procedure on current models, preregisters the analysis plan together with a set of eligible future models, and runs the confirmatory analysis on the first eligible model released afterward. Because this model does not exist at commitment time, it cannot be hacked against; furthermore, configurations that hack one model frequently do not transfer to the next. We evaluate the protocol on two tasks whose true values are known. Across 20 models from four providers and 11 LLM-analysis configurations, the protocol would have blocked successful transfer of the p-hack in 73.9% and 72.7% of cases in the two tasks. Additional analyses reveal that mitigation remains substantial under several stress tests. Finally, putting money where our mouth is, we followed our own protocol and preregistered our experiment. The preregistered experiment confirmed the protocol's effectiveness: out of the 7 configurations that hacked the prior model, the hacking failed to carry over in 6 configurations on the first eligible model released afterward.
マルチホライズンのボラティリティ予測における導入側の適応性
財務予測では、予測パフォーマンスはどのモデルがトレーニングされるかだけでなく、トレーニングされたモデルがどのようにデプロイされるかにも依存します。私たちはこの問題をマルチホライズンのボラティリティ予測で研究しています。私たちの出発点は、トレーニング済みのマルチ出力 (MIMO) 予測子が単一の展開可能な予測子を定義していないことです。推論時のロールアウト ルールを変更することで、同じトレーニング済みモデルが、異なる精度とコスト プロファイルを持つ一連の予測を誘導します。 20 の株価変動シリーズ、3 つの予測期間、および線形モデルから PatchTST に至るまでのアーキテクチャにわたって、デフォルト以外のロールアウト ルールが標準の MIMO 導入よりも改善されることが多いことがわかりました。ただし、最適な固定ルールはアーキテクチャや分野によって大幅に異なるため、単一の静的な置き換えは信頼できません。したがって、誘導されたルール ファミリに対して検証ベースの展開ポリシーを評価します。 MSE の主な目的では、検証で選択されたシングルトンはデフォルトの MIMO よりも低コストで改善を提供しますが、小さなルール サブセットは大幅に低い推論コストで大規模なアンサンブルの利点の多くを回収します。また、ポリシーのランキングが指標に依存していることもわかりました。MSE が選択したポリシーは、金融標準のボラティリティ損失である QLIKE に一律に移行されません。これらの結果は、推論時のデプロイメントが財務予測における適応性の意味のあるソースであること、および訓練を受けたボラティリティ予測担当者はアーキテクチャだけでなくデプロイメント ポリシーによっても評価されるべきであることを示しています。
原文 (English)
Deployment-Side Adaptiveness in Multi-Horizon Volatility Forecasting
In financial forecasting, predictive performance depends not only on which model is trained, but also on how the trained model is deployed. We study this issue in multi-horizon volatility forecasting. Our starting point is that a trained multi-output (MIMO) forecaster does not define a single deployable predictor: by changing the inference-time rollout rule, the same trained model induces a family of forecasts with different accuracy and cost profiles. Across 20 stock-volatility series, three forecast horizons, and architectures ranging from linear models to PatchTST, we find that non-default rollout rules often improve over standard MIMO deployment. However, the best fixed rule varies substantially across architectures and horizons, making any single static replacement unreliable. We therefore evaluate validation-based deployment policies over the induced rule family. Under the primary MSE objective, validation-selected singletons provide a low-cost improvement over default MIMO, while small rule subsets recover much of the benefit of larger ensembles at substantially lower inference cost. We also find that policy rankings are metric-sensitive: MSE-selected policies do not transfer uniformly to QLIKE, a finance-standard volatility loss. These results show that inference-time deployment is a meaningful source of adaptiveness in financial forecasting, and that trained volatility forecasters should be evaluated not only by their architecture, but also by their deployment policy.
早くやめて!認定された堅牢性を実現する早期停止
ランダム化スムージング (RS) は、アーキテクチャ上の制約なしでニューラル ネットワークの厳密な堅牢性を保証しますが、その採用は極度の計算コストによって制限されます。標準 RS では、入力ごとに何万ものモデル評価が必要であり、実践者は事前に固定サンプル サイズにコミットする必要があります。この研究では、計算リソースを適応的に展開する、いつでも有効な認証された堅牢性を実現する新しいメタ学習フレームワークを紹介します。軽量のメタ学習器を使用して逐次 E プロセスの画像固有の事前分布を予測することにより、厳密な統計的保証を維持しながら、従来の方法と比較してサンプルの複雑さを 20 倍削減することができます。本来の効率性を超えて、いつでも有効にすることで、アプリケーション固有のリスクしきい値に基づいてコンピューティングを適応的に割り当てることがどのように可能になるか、これは従来の認定フレームワークでは不可能なリソースの優先順位付けの形式であることを示します。これが達成可能でありながら同様の認証パフォーマンスを提供できるということは、当社のアプローチがリアルタイムの安全性が重要な認証導入への道を提供することを示しています。
原文 (English)
Halt Fast! Early Stopping for Certified Robustness
Randomized Smoothing (RS) provides rigorous robustness guarantees for neural networks without architectural constraints, yet its adoption is limited by extreme computational costs. Standard RS requires tens of thousands of model evaluations per input and forces practitioners to commit to fixed sample sizes a priori. In this work, we present a novel meta-learning framework for anytime-valid certified robustness that adaptively deploys computational resources. By using a lightweight meta-learner to predict image-specific priors for a sequential E-process, we achieve a 20-fold reduction in sample complexity compared to traditional methods while maintaining rigorous statistical guarantees. Beyond raw efficiency, we demonstrate how anytime-validity enables adaptively allocating compute based upon application-specific risk thresholds, a form of resource triage impossible under classic certification frameworks. That this is achievable while also providing similar certification performance demonstrates that our approach provides a pathway for real-time, safety-critical certification deployments.
拡散モデルのクラス周波数ガイド付きノイズ スケジュール
この論文では、クラス周波数と拡散モデル内のマルチスケール ノイズ スケジュールの間の相関関係を初めて調査しました。スコアベースの生成モデルの場合、低密度領域によりスコアが不正確に推定されることが多く、その結果、生成の品質が損なわれます。マルチスケール ノイズ スケジュールは拡散プロセス中にこの問題を軽減できますが、低周波クラスは依然として大規模な低密度領域という課題に直面しており、その結果推定スコアが高周波クラスよりも不正確になります。さらに、高頻度のクラスはスコア空間を支配する傾向があり、これらのクラスからサンプルを生成する方向にほとんどのデータ ポイントが収束します。その結果、低周波クラス内で生成されたサンプルは、最適ではない品質と限られた多様性を示します。この課題に対処するために、低周波数クラスにはより大規模なノイズが与えられるべきであるという洞察を活用して、\textit{クラス周波数ガイド付き (CFRG)} ノイズ スケジュールを提案します。私たちの方法の有効性を説明するために、不均衡なデータセット、\textit{i.e.}、CIFAR-100-LT、および ImageNet-LT を使用して、画像生成、画像分類、テキストから画像への生成などのさまざまなタスクに関する実験を実施します。 CFRG ノイズ スケジュールを採用することにより、ベースラインを超える大幅な改善が達成され、ノイズ スケジュール設計における周波数統計の重要な役割が明らかになりました。
原文 (English)
Class-frequency Guided Noise Schedule for Diffusion Models
In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion models. For score-based generative models, low-density regions often lead to inaccurately estimated scores, thereby compromising the generation quality. Although the multi-scale noise schedule can alleviate this issue during the diffusion process, low-frequency classes still face the challenge of large low-density regions, resulting in more inaccurate estimated scores than high-frequency classes. Furthermore, high-frequency classes tend to dominate the score space, causing a convergence of most data points towards generating samples from these classes. Consequently, samples generated within low-frequency classes exhibit suboptimal quality and limited diversity. To address this challenge, we propose the \textit{Class-frequency Guided (CFRG)} noise schedule, leveraging the insight that low-frequency classes should be endowed with larger-scale noises. To illustrate the effectiveness of our method, we conduct experiments on various tasks, including image generation, image classification, and text-to-image generation, using imbalanced datasets, \textit{i.e.}, CIFAR-100-LT, and ImageNet-LT. By employing the CFRG noise schedule, we achieve substantial improvements over baselines, manifesting the crucial role of frequency statistics in noise schedule design.
またあれは何だったのでしょうか?自動音声認識の堅牢性が認定済み
自動音声認識システムは、敵対的な摂動にも無害な摂動にも敏感であることで知られています。これは参照データセットを使用して繰り返し実証されていますが、実際の転写に関するオラクルの知識が存在しないため、展開されたシステムでそのような動作を検出することは非常に困難です。私たちは、認定にヒントを得たメカニズムを採用すると、WER が大幅に減少し、再現率が増加し、信頼性と WER の間のスピアマン相関が減少する可能性があることを実証します。これは、デュアルゲート診断パイプラインを通じて実現されます。トークンの存在と敵対的排除の両方を証明するための統計的富を蓄積する両面アトミック監査と、勝利シーケンスを選択するランクベースのトーナメントです。 4 つの多様なアーキテクチャにわたる当社の評価では、単語エラー率が相対的に最大 55% 減少することが実証されていると同時に、音響セキュリティを強化するための詳細な単語および文レベルの認定も提供されます。
原文 (English)
What Was That Again? Certified Robustness for Automatic Speech Recognition
Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations. While this has been repeatedly demonstrated using reference datasets, detecting such behaviors in deployed systems is incredibly challenging, due to the absence of oracle knowledge of the true transcription. We demonstrate that employing a certification-inspired mechanism can significantly decrease WER, increase recall, and decrease the Spearman correlation between confidence and WER. We achieve this through a dual-gate diagnostic pipeline: a Two-Sided Atomic Audit that accumulates statistical wealth to certify both token existence and adversarial exclusion, and a Rank-Based Tournament that selects the winning sequence. Our evaluations across four diverse architectures demonstrate up to a 55% relative reduction in Word Error Rate, while also providing granular word- and sentence-level certifications to enhance acoustic security.
エラーの余地: 無線音響攻撃の大規模シミュレーション
音声制御は人間と AI のコミュニケーションのユビキタスなベクトルとして急速に普及しつつありますが、これらのシステムが直面するリスクについてはまだ十分に理解されていません。これは、部分的には、厳密にデジタルの敵対的なワークフローを物理世界に拡張する際の困難の結果です。これらのスケールの壁により、コミュニティは、検出可能性と音響に対する形状の影響に関連する主要な音響要素を抽象化するようになりました。これらの方法論的および計測学的欠点は、リスクに対する私たちの理解を台無しにします。私たちは、現実世界でのテスト、概念的な議論、新しい高スループットの現実シミュレーション フレームワークを通じて、これらの問題を明らかにします。 800 万を超える敵対的評価をテストすることにより、Whisper と wav2vec の下では、音響認識により相対的な単語誤り率が最大 94.5\% 増加することが実証されました。私たちはこのフレームワークを使用して、デュアルフォーム信号対雑音比の形式化と運用を検討し、ソースのステルスを被害者の攻撃の有効性から分離し、現在の作業における重大な制限を解決します。これにより、音響環境を抽象化するのではなく包含する、再現可能で検証可能な研究の基礎が築かれます。
原文 (English)
Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks
While voice control is rapidly becoming a ubiquitous vector of human-AI communication, the risks facing these systems remain poorly understood. This is, in part, a product of the difficulties in scaling strictly digital adversarial workflows to the physical world. These scale barriers have led the community to abstract away key acoustic factors relating to detectability and the influence of geometry on acoustics. These methodological and metrological shortcomings undermine our understanding of risk. We illuminate these issues through real-world testing, conceptual discussions, and a novel, high-throughput reality simulation framework. By testing over 8 million adversarial evaluations, we demonstrate that acoustic awareness yields relative Word Error Rate increases of up to 94.5\% under Whisper and wav2vec. We employ this framework to explore a formalize and operationalize a Dual-Form Signal to Noise Ratio to decouple source stealth from victim attack efficacy, resolving a crucial limitation in current works. This lays the groundwork for repeatable, verifiable research that embraces, rather than abstracts, the acoustic environment.
安全な LLM 微調整のための低同調性ペルソナ コンディショニング
最近の研究では、社会的温かさのために大規模言語モデル (LLM) を微調整すると、事実の信頼性が低下し、お調子者が増加することがわかっています。私たちは、関連するが別個の故障モードを調査します。つまり、ウォーム微調整によって敵対者の安全性も弱まり、モデルがジェイルブレイクや有害な出力生成を受けやすくなります。私たちは、これが共感的適応の固有の結果を反映しているのか、それともデータ構築のアーチファクトを反映しているのかを調べます。これに対処するために、ユーザーが低同調性をオンにする条件を設定し、これを温かくエスカレーションを緩和するアシスタントの応答と組み合わせるペルソナ主導の書き換えパイプラインを導入します。 4 つのモデルでの 3 つの実験を通じて、私たちのアプローチは、会話の暖かさを維持しながら、一般的な暖かさの微調整ベースラインと比較して脱獄の感受性と有害な出力率を削減しました。表象的探査は、この条件付けが潜在空間における温かさ方向と順応方向の間の幾何学的配列を減少させるという示唆的な証拠を提供します。これらの結果は、より安全な共感の微調整が、安全ラベルや危害検出器、トレーニング目標の変更を必要とせず、データ設計だけで達成できることを示しています。
原文 (English)
Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning
Recent work has shown that fine-tuning large language models (LLMs) for social warmth degrades factual reliability and increases sycophancy. We investigate a related but distinct failure mode: warmth fine-tuning also weakens adversarial safety, making models more susceptible to jailbreaks and harmful output generation. We examine whether this reflects an inherent consequence of empathetic adaptation or an artifact of data construction. To address this, we introduce a persona-driven rewriting pipeline that conditions user turns on low agreeableness and pairs this with warm, de-escalating assistant responses. Across three experiments on four models, our approach reduces jailbreak susceptibility and harmful output rates relative to generic warmth fine-tuning baselines, while preserving conversational warmth. Representational probing provides suggestive evidence that this conditioning reduces the geometric alignment between warmth and compliance directions in latent space. These results show that safer empathetic fine-tuning is achievable through data design alone, without safety labels, harm detectors, or changes to the training objective.
Simulacrum: 最適に近い時系列予測と推論のための意思決定理論的事前トレーニング
決定理論的事前トレーニングと呼ばれるプロセスを通じて時系列推定量を学習するためのニューラル ネットワーク ベースのフレームワークを導入します。アナリストは、生成世界、データ生成プロセス全体の分布、および目標の意思決定目標を指定します。この世界の層別シミュレーションでトレーニングされたニューラル ネットワークは、対応する最適な決定ルールを近似し、予測、パラメーター推定、予測間隔、またはこれまで見たことのない時系列でのゼロショット推論のためのモデル選択を提供するニューラル推定器を生成します。生成世界と目的の共同仕様により、推定者は、最適に近いリスク、バイアス制御、ミニマックスパフォーマンス、均一なキャリブレーションなど、プロセスレベルの有限サンプル特性を直接近似することができます。私たちの実験では、これらのニューラル推定器が、同じモデル構造モデル クラスに対して、最尤推定や AICc によるモデル選択などの従来のベースラインよりも優れたパフォーマンスを発揮できることを示しています。さらに、構造モデルのシミュレーションのみでトレーニングされた場合でも、統計モデル、ニューラル モデル、または大規模な事前トレーニング済みモデルと比較して、主要な現実世界のベンチマークで競合するまたは最先端の予測精度を達成します。 AR(p) モデルにおける有限サンプルのバイアスと校正ミス、および予測の組み合わせパズルという 2 つの長年の課題に対処することでフレームワークを説明します。これらのアプリケーションは、このアプローチの主な利点を強調しています。それは、複雑な構造方程式や最適性基準を含む、分析的に扱いにくい、または計算的に法外な時系列問題の解を近似する能力です。最終的に、このフレームワークは、意思決定理論のトレードオフに対する明示的な制御を可能にすることで、アナリストに特定の分析ニーズに合わせた高効率の推定ツールを提供します。
原文 (English)
The Simulacrum: Decision-Theoretic Pretraining for Near-Optimal Time-Series Forecasting and Inference
We introduce a neural network-based framework for learning time series estimators through a process we term decision-theoretic pretraining. Analysts specify a generative world, a distribution over data-generating processes, and a target decision objective. A neural network trained on stratified simulations from this world approximates the corresponding optimal decision rule, yielding a neural estimator that provides forecasts, parameter estimates, predictive intervals, or model-selection for zero-shot inference on previously unseen time series. The joint specification of the generative world and objective enables the estimators to directly approximate process-level, finite-sample properties: near-optimal risk, bias control, minimax performance, and uniform calibration. Our experiments demonstrate that these neural estimators can outperform traditional baselines such as maximum likelihood estimation and model selection via AICc, for the same model structural model classes. Furthermore, even when trained purely on simulations of structural models, they achieve competitive or state-of-the-art forecasting accuracy on major real-world benchmarks, compared with statistical, neural or large pre-trained models. We illustrate the framework by addressing two longstanding challenges: finite-sample bias and miscalibration in AR(p) models, and the forecast combination puzzle. These applications highlight the approach's main advantage: its ability to approximate solutions to analytically intractable or computationally prohibitive time series problems, including complex structural equations or optimality criteria. Ultimately, by enabling explicit control over decision-theoretic trade-offs, the framework equips analysts with highly efficient estimation tools tailored to their specific analytical needs.
音声強調モデルは言語や感情を超えて一般化されますか?
韻律の強調は言語、感情、話し方によって異なりますが、既存の強調検出モデルは主に単一言語の中立的な読み上げ音声でトレーニングおよび評価されています。 MMEE (Multilingual Multi-Emotion Emphasis) は、7 つの言語と 34 の感情/スタイル カテゴリにわたって専門的に記録された 10,000 の表現力豊かな発話 (14.13 時間) のコーパスであり、3 レベルの知覚ラベル (サンプルごとに 10 個の注釈) が付いています。私たちは、単一言語、クロスリンガル、多言語、クロス感情、クロスデータセット、およびデータスケール設定の下で 2 つの最先端のアーキテクチャをベンチマークします。単言語モデルでは、ゼロショット転送が限られており、類型的に遠い言語間では性能が低下しますが、多言語トレーニングでは堅牢性が大幅に向上します。モデルは、高覚醒感情と低覚醒感情の間を確実に移行します。合成ベンチマークと知覚ベンチマーク間の双方向の伝達は、共有された韻律構造を示唆しています。また、トレーニング規模が小さい場合でもパフォーマンスは安定します。
原文 (English)
Do Speech Emphasis Models Generalize across Languages and Emotions?
Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speech. We introduce MMEE (Multilingual Multi-Emotion Emphasis), a corpus of 10,000 professionally recorded expressive utterances (14.13 hours) across 7 languages and 34 emotion/style categories, with three-level perceptual labels (10 annotations per sample). We benchmark two state-of-the-art architectures under monolingual, cross-lingual, multilingual, cross-emotion, cross-dataset, and data-scale settings. Monolingual models show limited zero-shot transfer, degrading across typologically distant languages, while multilingual training substantially improves robustness. Models transfer robustly between high- and low-arousal emotions; bidirectional transfer between synthetic and perceptual benchmarks suggests shared prosodic structure; and performance stays robust even at smaller training scales.
スムーズな MMD アライメントによる LLM の数値予測の強化
大規模言語モデル (LLM) は、強力な一般機能にもかかわらず、出力が数値的に正確である必要がある場合には信頼性が低いままであることがよくあります。主な理由はトレーニングの目的です。標準のクロスエントロピーは数値トークンを非構造化カテゴリーとして扱い、その値のメトリック構造を無視します。この不一致は、数値トークンとグラフベースの滑らかさの上に値と距離のカーネルを組み込むことで、古典的な MMD に基づいて構築された Smooth Maximum Mean Discrepancy (SMMD) で解決されます。数値サブ語彙上で定義されたこのカーネルを使用すると、SMMD はカーネル マッチングを通じて予測数値分布をターゲットに合わせ、誘導されたカーネル グラフ上で予測ターゲットの残差を平滑化して、局所的な一貫性を促進します。複数のオープンウェイト LLM および VLM バックボーンにわたって、数学的推論、算術計算、時刻認識、チャートの質問応答という 4 つの数値ターゲット タスクで SMMD を評価します。 SMMD は、クロスエントロピーと最近の数値ターゲット損失の両方に対する精度を一貫して向上させます。分析では、MMD と滑らかさの間の相補的な効果が示され、距離ベースのカーネル設計の重要性が強調されています。コードは https://github.com/Zuozhuo/smmd-loss で入手できます。
原文 (English)
Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment
Despite their strong general capabilities, large language models (LLMs) often remain unreliable when outputs must be numerically precise. A key reason is the training objective: standard cross-entropy treats numeric tokens as unstructured categories and ignores the metric structure of their values. We address this mismatch with Smooth Maximum Mean Discrepancy (SMMD), which builds on the classic MMD by incorporating value-distance kernels over numeric tokens and graph-based smoothness. With this kernel defined over a numeric sub-vocabulary, SMMD aligns the predicted numeric distribution to the target via kernel matching and smooths the prediction-target residual over the induced kernel graph to encourage local consistency. We evaluate SMMD on four numeric-target tasks: mathematical reasoning, arithmetic calculation, clock-time recognition, and chart question answering, across multiple open-weight LLM and VLM backbones. SMMD consistently improves accuracy over both cross-entropy and recent numeric-target losses; analyses show complementary effects between MMD and smoothness and underscore the importance of distance-based kernel design. Code is available at https://github.com/Zuozhuo/smmd-loss.
二焦点拡散言語モデル: 並列生成のための非対称双方向コンテキスト
離散拡散言語モデル (dLLM) は、マスクされたトークンを並行して回復し、自己回帰 (AR) 生成よりも大幅な高速化を実現します。しかし、そのような有望なフレームワークは基本的なアーキテクチャ設計のジレンマに直面しています。 \ding{182} 双方向アテンションを採用すると、各位置が完全なコンテキストにアクセスできるようになり、強力な生成品質が実現しますが、本質的に KV キャッシュと互換性がなく、バッチ処理シナリオでの推論スループットが制限されます。 \ding{183} 逆に、因果的注意は効率的なキャッシュされた推論を可能にしますが、右側のコンテキストをすべて失い、生成の品質を大幅に低下させます。この論文では、\emph{非対称双方向コンテキスト} を通じてこのジレンマを解決する新しいパラダイムである Bifocal dLLM を紹介します。二焦点レンズと同様に、パラダイムを \textbf{R2LM} (Right-to-Left Mamba) としてインスタンス化します。これは 2 つの相補的なメカニズムを組み合わせたものです: $a$) 完全な KV キャッシュ互換性を備えた正確な左側のコンテキストを提供する標準的な因果的注意、一方 $b$) キャッシュ可能性を損なうことなく圧縮された右側のコンテキストを提供する軽量のリバース Mamba SSM サイドカー。 60B トークンを使用した Qwen3-1.7B の継続的な事前トレーニングに関する包括的な実験では、R2LM が双方向 dLLM よりも $2.4\time$ から $12.9\times$ 高いスループットと、KV キャッシュを使用した並列デコードによるバッチ処理での AR ベースラインを上回る $1.9\time$ から $2.9\times$ の高速化を達成しながら、ほとんどのベンチマークで因果ベースラインを超え、平均すると双方向 dLLM。
原文 (English)
Bifocal Diffusion Language Models: Asymmetric Bidirectional Context for Parallel Generation
Discrete diffusion language models (dLLMs) recover masked tokens in parallel, offering significant speedups over autoregressive (AR) generation. However, such promising frameworks face a fundamental architectural design dilemma: \ding{182} Adopting bidirectional attention achieves strong generation quality by allowing each position to access the full context, but is inherently incompatible with KV caching, limiting inference throughput in batch-serving scenarios; \ding{183} Conversely, causal attention enables efficient cached inference but loses all right-side context, substantially degrading generation quality. This paper introduces Bifocal dLLMs, a new paradigm that resolves this dilemma through \emph{asymmetric bidirectional context}. Analogous to bifocal lenses, we instantiate the paradigm as \textbf{R2LM} (Right-to-Left Mamba), which combines two complementary mechanisms: $a$) standard causal attention providing precise left-context with full KV cache compatibility, while $b$) a lightweight reverse Mamba SSM sidecar supplying compressed right-side context without breaking cacheability. Comprehensive experiments on continued pretraining of Qwen3-1.7B with 60B tokens demonstrate that R2LM achieves $2.4\times$ to $12.9\times$ higher throughput than bidirectional dLLMs and $1.9\times$ to $2.9\times$ speedup over AR baselines in batch serving through parallel decoding with KV caching, while exceeding the causal baseline on most benchmarks and surpassing the bidirectional dLLM on average.
KG2Cypher: エンタープライズ Text-to-Cypher システムを構築するためのデータ中心のパイプライン
エンタープライズ ナレッジ グラフ (KG) は、内部検索、分析、質問応答にますます使用されていますが、プライベート エンタープライズ グラフ用の自然言語インターフェイスの構築には依然としてコストがかかります。既存の KG からエンタープライズ Text-to-Cypher システムを構築するためのデータ中心のパイプラインである KG2Cypher を紹介します。 KG2Cypher は、まず観察されたグラフ ファクトから実行可能な Cypher クエリを構築し、次に LLM を使用してそれに関連する自然言語の質問を生成します。結果として得られるテキストと暗号のペアは、LLM 審査員と人間による検証によって検証され、候補を認識する SFT データに変換されます。トレーニングされたジェネレーターには、クラス条件付きのスキーマ プロンプト、エンティティの取得、および LoRA ベースの推論が提供されます。私たちは、短い検索形式のクエリやスキーマの言い換えにより言語の基礎が難しくなる韓国の企業環境で KG2Cypher を評価しました。 LoRA SFT は、放送番組クエリでは実行結果 F1 を 0.806 から 0.950 に、企業クエリでは 0.70 から 0.92 に改善します。 11 クラス設定では、KG2Cypher は 95.2% の完全一致、99.9% の実行率、および 0.964 の実行結果 F1 を達成します。
原文 (English)
KG2Cypher: Data-Centric Pipeline for Building Enterprise Text-to-Cypher Systems
Enterprise Knowledge Graphs (KGs) are increasingly used for internal search, analytics, and question answering, but building natural-language interfaces for private enterprise graphs remains costly. We present KG2Cypher, a data-centric pipeline for building enterprise text-to-Cypher systems from existing KGs. KG2Cypher first constructs an executable Cypher query from observed graph facts and then uses LLMs to generate its associated natural-language question. The resulting Text-Cypher pairs are validated with an LLM judge and human validation, and are converted into candidate-aware SFT data. The trained generator is served with class-conditioned schema prompting, entity retrieval, and LoRA-based inference. We evaluate KG2Cypher in Korean enterprise settings, where short search-style queries and schema paraphrases make language grounding difficult. LoRA SFT improves execution-result F1 from 0.806 to 0.950 on broadcast-program queries and from 0.70 to 0.92 on company queries. In an 11-class setting, KG2Cypher achieves 95.2% exact match, 99.9% execution rate, and 0.964 execution-result F1.
リソース適応型 LLM 推論のためのエンドツーエンドの動的スパース性
大規模言語モデル (LLM) 推論は通常、静的リソースの仮定の下で展開され、モデルはランタイム環境に関係なく固定の計算グラフを実行します。ただし、現実のクラウド インフラストラクチャは本質的に動的であり、可用性の変動 (スポット インスタンスのプリエンプションなど) と段階的なサービス品質要件によって特徴付けられます。このような不安定な設定では、静的モデルは柔軟性に欠けます。リソースの制約の下でクラッシュするか、冗長な操作で計算を無駄にします。このギャップを埋めるために、リソース適応推論のためのエンドツーエンドのフレームワークである Learning to Allocate (L2A) を提案します。入力の難易度のみを条件とする従来の方法とは異なり、入力と実行時のリソース バジェット自体の両方を条件とする制約付き割り当て問題として推論を定式化します。 LLM に統合された、軽量で予算に応じた入力認識型のゲート ネットワークを導入します。これらのゲートは、現実世界のダイナミクスの現れ方に一致する 3 つの軸 (メモリと深さのプレッシャーのためのレイヤー スキップ、スループット競合のためのヘッド プルーニング、レイテンシ短縮のための推論トークンの削減) に沿ってタスク パフォーマンス、論理的一貫性、およびリソース コストを共同で最適化する統一目標を介してトレーニングされます。これにより、モデルは入力の難易度だけを超えて予算を意識したポリシーを学習できるようになります。リアルタイムのリソースのダイナミクスに応じて計算フットプリントを適応的に構成し、リソースが許せば推論の深さを最大化し、予算が逼迫した場合には厳格な倹約を強制します。単一の L2A モデルは、Llama-3-8B および Qwen-3-4B 上の計算精度パレート フロンティア全体をトレースします。最大 34% の実現層スパース性で、GSM8K 上の高密度ベースラインの 0.6% 以内に留まり、分散外タスクでは同じギャップがゼロショットを維持します。一方、すべての静的ベースラインまたはヒューリスティック ベースラインには個別に調整されたモデルが必要であり、依然として 5 ~ 10% 低下します。同等の推論時間。
原文 (English)
End-to-End Dynamic Sparsity for Resource-Adaptive LLM Inference
Large Language Models (LLMs) inference is typically deployed under a static resource assumption, where models execute a fixed computational graph regardless of the runtime environment. However, real-world cloud infrastructure is inherently dynamic, characterized by fluctuating availability (e.g., spot instance preemption) and tiered Quality-of-Service requirements. In such volatile settings, static models are inflexible: they either crash under resource constraints or waste compute on redundant operations. To bridge this gap, we propose Learning to Allocate (L2A), an end-to-end framework for resource-adaptive inference. Unlike prior methods that condition only on input difficulty, we formulate inference as a constrained allocation problem conditioned on both the input and the runtime resource budget itself. We introduce lightweight, budget-conditioned and input-aware gating networks integrated into the LLM. These gates are trained via a unified objective that jointly optimizes task performance, logical consistency, and resource costs along three axes matching how real-world dynamics manifest: layer skipping for memory and depth pressure, head pruning for throughput contention, and reasoning-token reduction for latency tightening. This lets the model learn a budget-aware policy beyond input difficulty alone: it adaptively configures its computational footprint with respect to real-time resource dynamics, maximizing reasoning depth when resources permit while enforcing strict frugality when budgets tighten. A single L2A model traces the entire compute-accuracy Pareto frontier on Llama-3-8B and Qwen-3-4B: at up to 34% realized layer sparsity, it stays within 0.6% of the dense baseline on GSM8K, with the same gap holding zero-shot on out-of-distribution tasks, while every static or heuristic baseline requires a separately tuned model and still drops by 5-10% at comparable inference time.
Flexformer: 学習可能なアテンション カーネルを備えた柔軟なリニア トランスフォーマー
Transformer モデルは、アテンション メカニズムに依存して長距離の依存関係を捕捉しますが、二次的な複雑さの影響を受け、スケーラビリティが長いシーケンスに制限されます。カーネルベースの線形アテンションはこの複雑さを軽減しますが、通常は固定カーネルまたは学習能力の低いカーネルに依存するため、表現力とパフォーマンスが制限されます。この研究では、完全にデータ駆動型の方法でアテンション カーネルを学習する柔軟な線形トランスフォーマーである Flexformer を提案します。 Flexformer は、ランダムなフーリエ特徴ベースの線形アテンションに基づいて構築されており、スペクトル周波数をトレーニング可能なパラメーターとして扱うことで、モデルがアテンション カーネルの幅広いファミリーを学習できるようにします。私たちは定常バリアントと非定常バリアントの両方を開発していますが、後者は厳密に優れた表現力を提供します。言語モデリングとシーケンス分類に関する広範な実験により、Flexformer が常にベースラインを上回るパフォーマンスを示すことが実証されました。さらに、Flexformer は、事前トレーニングされた Transformer から効果的に抽出してソフトマックス アテンションを回復することができ、ドメイン間での強力なカーネル移行性を示し、長時間シーケンスのタスクで高い効率と競争力のあるパフォーマンスの両方を実現します。
原文 (English)
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel
Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences. Kernel-based linear attention reduces this complexity but typically relies on fixed or weakly learnable kernels, restricting expressiveness and performance. In this work, we propose Flexformer, a flexible linear Transformer that learns attention kernels in a fully data-driven manner. Flexformer builds on random Fourier feature-based linear attention and treats spectral frequencies as trainable parameters, enabling the model to learn a broad family of attention kernels. We develop both stationary and nonstationary variants, with the latter offering strictly greater expressiveness. Extensive experiments on language modeling and sequence classification demonstrate that Flexformer consistently outperforms baselines. Moreover, Flexformer can be effectively distilled from pretrained Transformers to recover softmax attention and exhibits strong kernel transferability across domains, achieving both high efficiency and competitive performance on long-sequence tasks.
汎用オーディオタグ付けから空間的に接地されたサウンドイベントの位置特定と検出まで
このレポートでは、事前トレーニングされた汎用オーディオ タギング (GP-AT) モデルの、空間的に接地されたサウンド イベント ローカリゼーションと検出 (SELD) への拡張について調査します。提案された AT2SELD フレームワークは、事前トレーニングされた AT バックボーンを、コンパクトな 1 次アンビソニックス (FOA) 空間処理、トラック単位の SED およびデカルト DOA 推定、置換を意識した監視、およびキャリブレーションと結合します。これは、データ、計算、展開の制約の下で、セマンティック オーディオ プリアがローカリゼーションを意識したシーン分析をどのようにサポートするかを特徴づけます。このフレームワークは、情報に基づいた多段階の Neural Architecture Search (NAS) を通じて開発されています。ステージ 1 は、振幅、位相、および強度ベクトル (IV) に基づくスペクトル FOA 記述子が、意味から空間への転送に最も信頼できるインターフェイスを提供することを示しています。ステージ 2 では、初期の残差空間エンコーディングが主な容量依存コンポーネントとして特定され、一方、後期のトラック単位の抽象化とリカレント スムージングは主にリファインメント ステージとして機能します。ステージ 3 では、遅いクロスステッチの結合は意味論的空間的相互作用を改善しますが、早期の融合はコストが高く、効果が低いことを示しています。診断評価では、クラス バランシング、焦点損失、アクティビティ条件付き DOA 監視、しきい値キャリブレーション、STARSS23、TAU2019、TAU-NIGENS2020、および TAU-NIGENS2021 間の転送の下で、選択されたアーキテクチャを分析します。焦点損失は活動点を改善し、アクティブのみの DOA 監視は非アクティブなターゲットの優位性を軽減し、検証で選択されたしきい値は空間学習を置き換えることなくキャリブレーションを回復します。クロスデータセットおよびオラクルアクティビティ分析により、TAU2019 での強力な固定ソース位置特定、TAU NIGENS2021 からの転送可能な表現、および STARSS23 での意味のあるが不確実な動作が示されています。全体として、GP-AT の事前分布は、空間認識アーキテクチャに組み込まれ、統合されたキャリブレーションと導入指向の戦略を通じて最適化された場合、SELD 設計に有望であると考えられます。
原文 (English)
From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection
This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Localization and Detection (SELD). The proposed AT2SELD framework couples a pretrained AT backbone with compact First-Order Ambisonics (FOA) spatial processing, track-wise SED and Cartesian DOA estimation, permutation aware supervision, and calibration. It characterizes how semantic audio priors support localization-aware scene analysis under data, computation, and deployment constraints. The framework is developed through informed multi-stage Neural Architecture Search (NAS). Stage 1 shows that spectral FOA descriptors, based on magnitude, phase, and Intensity Vectors (IVs), provide the most reliable interface for semantic-to-spatial transfer. Stage 2 identifies early residual spatial encoding as the main capacity-sensitive component, while late track-wise abstraction and recurrent smoothing act mainly as refinement stages. Stage 3 shows that late cross-stitch coupling improves semantic-spatial interaction, whereas early fusion is costlier and less effective. Diagnostic evaluation analyzes the selected architecture under class balancing, focal loss, activity-conditioned DOA supervision, threshold calibration, and transfer across STARSS23, TAU2019, TAU-NIGENS2020, and TAU-NIGENS2021. Focal loss improves the activity point, active-only DOA supervision mitigates inactive target dominance, and validation-selected thresholds recover calibration without replacing spatial learning. Cross-dataset and oracle-activity analyses indicate strong fixed source localization on TAU2019, transferable representations from TAU NIGENS2021, and meaningful but uncertain behavior on STARSS23. Overall, GP-AT priors appear promising for SELD design when embedded in spatial-aware architectures and optimized through integrated calibration and deployment oriented strategies.
Drop-Then-Recovery: 視覚-言語-行動モデルはどの程度冗長ですか?
Vision-Language-Action (VLA) モデルは命令駆動型のロボット操作を可能にしますが、短いロボット命令に必要な容量をはるかに超える容量を備えた事前学習済み VLM から特大の言語バックボーンを継承しています。これにより、基本的な疑問が生じます。閉ループ制御には、VLA モデルのどの程度が実際に必要なのでしょうか?この研究では、制御された介入として変圧器ブロックの削除を使用することにより、VLA モデルのアーキテクチャ上の冗長性を研究します。 \textbf{Drop-Then-Recovery (DTR)} を導入します。これは、事前トレーニングされた VLA モデルから選択したブロックを削除し、結果のモデルを微調整して、削除された容量がダウンストリーム制御に必要かどうかを測定する分析プロトコルです。この介入を信頼できるものにするために、下流のアクション損失への寄与によってブロックをランク付けするワンショット仮想ゲート感度メトリックである \textbf{GateProbe} を提案します。複数の VLA アーキテクチャ、操作ベンチマーク、さらには実際のロボット産業シナリオ全体にわたって、取り外し後の回復可能性には強い非対称性があることがわかりました。\ul{\textit{標準的なロボット操作タスクでは言語バックボーンは非常に冗長ですが、視覚と動作の経路は取り外しに対して大幅に耐性がありません}}。 LIBERO では、LLM ブロックの半分を削除すると、同じダウンストリーム微調整予算の下で OpenVLA-OFT が 95.0% から 98.3% に向上し、言語ブロックを 2 つだけ保持してもベースライン レベルのパフォーマンスが回復します。これらの結果は、現在の VLA ベンチマークが深い言語の基礎と構成的な命令の理解に及ぼす圧力には限界がある可能性があり、将来の VLA アーキテクチャでは、言語、視覚、アクションのコンポーネント全体に容量をより慎重に割り当てる必要があることを示唆しています。コードは https://github.com/s1ghh/VLADrop で入手できます。
原文 (English)
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?
Vision-Language-Action (VLA) models enable instruction-driven robotic manipulation, but they inherit oversized language backbones from pretrained VLMs whose capacity far exceeds what is needed for short robotic instructions. This raises a basic question: how much of a VLA model is actually necessary for closed-loop control? In this work, we study architectural redundancy in VLA models by using transformer block removal as a controlled intervention. We introduce \textbf{Drop-Then-Recovery (DTR)}, an analysis protocol that removes selected blocks from a pretrained VLA model and then fine-tunes the resulting model to measure whether the removed capacity was necessary for downstream control. To make this intervention reliable, we propose \textbf{GateProbe}, a one-shot virtual-gate sensitivity metric that ranks blocks by their contribution to the downstream action loss. Across multiple VLA architectures, manipulation benchmarks and even real-robot industrial scenarios, we find a strong asymmetry in post-removal recoverability: \ul{\textit{language backbones are highly redundant for standard robotic manipulation tasks, whereas vision and action pathways are substantially less tolerant to removal}}. On LIBERO, removing half of the LLM blocks even improves OpenVLA-OFT from 95.0% to 98.3% under the same downstream fine-tuning budget, and retaining only two language blocks still recovers baseline-level performance. These results suggest that current VLA benchmarks may exert limited pressure on deep language grounding and compositional instruction understanding, and that future VLA architectures should allocate capacity more deliberately across language, vision, and action components. The code is available at https://github.com/s1ghhh/VLADrop.
RS-Diffuser: 分配価値ガイダンスを備えたリスクに敏感な拡散計画
オフライン強化学習では、追加の環境インタラクションなしで固定データセットからポリシー学習を行うことができるため、オンライン探索がコストがかかる、または安全でない場合に安全性が重要なアプリケーションにとって魅力的です。拡散ベースの意思決定手法は、最近、豊富なマルチモーダル軌道分布をモデル化することにより、オフライン RL で優れたパフォーマンスを達成しました。しかし、既存の普及計画立案者は通常、リスク中立的なため、現実世界の導入において重要な、まれではあるが壊滅的な結果を見落とす可能性があります。この研究では、拡散ベースの軌道生成と分布価値批判を組み合わせた、リスクに敏感なオフライン拡散計画フレームワークである RS-Diffuser を提案します。 RS-Diffuser は、将来の状態の軌道に関する拡散プランナー、アクション デコード用の別個の逆ダイナミクス モデル、および分位点回帰を通じて候補プランの完全リターン分布を推定するモンテカルロ分布クリティカルを学習します。サンプリング時に、リスクに応じた条件付き目標値などのテールアウェア目標から計算された勾配を使用して、リスクに敏感なガイダンス信号をノイズ除去プロセスに組み込み、望ましいリスク プロファイルに向けて生成を誘導します。その結果、単一のトレーニング済みモデルは、推論時のリスク パラメーターのみを変更することで、リスク回避行動、リスク中立行動、またはリスク探索行動を柔軟に生成できます。リスクに敏感な D4RL と危険なロボット ナビゲーション ベンチマークに関する広範な実験により、RS-Diffuser が最先端のパフォーマンスを達成し、安全性違反を軽減しながら全体的な収益と最悪の場合の堅牢性の両方を向上させることが実証されました。
原文 (English)
RS-Diffuser: Risk-Sensitive Diffusion Planning with Distributional Value Guidance
Offline reinforcement learning enables policy learning from fixed datasets without additional environment interaction, making it appealing for safety-critical applications where online exploration is costly or unsafe. Diffusion-based decision-making methods have recently achieved strong performance in offline RL by modeling rich, multimodal trajectory distributions. However, existing diffusion planners are typically risk-neutral and therefore may overlook rare but catastrophic outcomes that are crucial in real-world deployment. In this work, we propose RS-Diffuser, a risk-sensitive offline diffusion planning framework that combines diffusion-based trajectory generation with distributional value critics. RS-Diffuser learns a diffusion planner over future state trajectories, a separate inverse dynamics model for action decoding, and a Monte Carlo distributional critic that estimates the full return distribution of candidate plans through quantile regression. At sampling time, we incorporate a risk-sensitive guidance signal into the denoising process, using gradients computed from tail-aware objectives such as Conditional Value at Risk to steer generation toward desired risk profiles. As a result, a single trained model can flexibly produce risk-averse, risk-neutral, or risk-seeking behaviors by changing only the inference-time risk parameter. Extensive experiments on risk-sensitive D4RL and risky robot navigation benchmarks demonstrate that RS-Diffuser achieves state-of-the-art performance, improving both overall return and worst-case robustness while reducing safety violations.
活性化の増幅と減衰による敵対的な堅牢性の向上
敵対的攻撃の存在は、多くの場合、ニューラル ネットワーク内の非堅牢な機能の存在に起因します。従来の防御はプルーニング、マスキング、または機能の再調整によって影響を軽減していましたが、代わりに、単純なアクティベーション スケーリング メカニズムを通じてこれらの信号を増幅および減衰する方法を共同で学習することを提案します。この目的を達成するために、アクティベーションの変更を最小限に抑えて敵対的な堅牢性を強化する軽量のプラグイン モジュールである Activation Amplification and Attenuation (A3) を導入します。 A3 は、学習可能なマスクと元のアクティベーションの大きさから導出されたスケーリング係数を使用して、アクティベーションを動的に再スケーリングします。敵対的な摂動の影響は、スケーリング操作の符号を反転するだけで、同じ学習可能なパラメーターを使用して増幅または減衰できます。増幅された信号は、新しい対照的な損失関数とランク付けされた損失関数を構築するための負の基準として機能します。実験分析によると、増幅モードで予測を劣化させることを学習すると、同時に減衰モードでの敵対的ロバスト性も向上します。さらに、A3 は少数の学習可能なパラメータのみに依存しており、その動作のほとんどは追加のネットワーク容量ではなくスケーリング メカニズムによって決定されます。広範な実験により、A3 をさまざまなバックボーン、データセット、トレーニング方法に統合することで、既存のプラグイン モジュールと比較して無視できるほどの計算量とメモリのオーバーヘッドを導入しながら、敵対的な堅牢性が一貫して向上することが実証されました。コードは https://github.com/tgoncalv/A3 から入手できます。
原文 (English)
Improving Adversarial Robustness via Activation Amplification and Attenuation
The existence of adversarial attacks is often attributed to the presence of non-robust features in neural networks. While prior defenses reduce their impact via pruning, masking, or feature recalibration, we instead propose to jointly learn to amplify and attenuate these signals through a simple activation scaling mechanism. To this end, we introduce Activation Amplification and Attenuation (A3), a lightweight plug-in module that enhances adversarial robustness with minimal modifications of the activations. A3 dynamically rescales the activations using a learnable mask and a scaling factor derived from the original activation magnitudes. The influence of adversarial perturbations can be amplified or attenuated using the same learnable parameters by simply flipping the sign of the scaling operation. The amplified signals serve as negative references to construct novel contrastive and ranking loss functions. Experimental analysis shows that learning to degrade the predictions in amplification mode simultaneously improves adversarial robustness in attenuation mode. Moreover, A3 relies on only a small number of learnable parameters, with most of its behavior being determined by the scaling mechanism rather than additional network capacity. Extensive experiments demonstrate that integrating A3 into different backbones, datasets, and training methods consistently improves adversarial robustness while introducing negligible computational and memory overhead compared to existing plug-in modules. Code is available at: https://github.com/tgoncalv/A3.
キャリブレーションガイド付き LLM 圧縮の出力スペース割り当てコスト: 実証的研究
大規模言語モデル (LLM) のトレーニング不要の圧縮方法では、多くの場合、圧縮の決定をガイドするためにキャリブレーション データが使用されます。 ROCKET は、スパース辞書因数分解と多選択ナップザック問題 (MCKP) 割り当てを組み合わせた最近の手法で、出力再構成目的から層ごとの因数分解を導き出しますが、重み空間のフロベニウス誤差を MCKP 割り当てコストとして使用します。割り当てコストを出力空間目標に合わせることで、圧縮モデルの忠実度が向上するかどうかを調査します。 50\% 圧縮の Qwen3-8B では、ROCKET-ActCost は 8 つのゼロショット ベンチマーク全体で +0.8 パーセント高い平均精度を達成しました (53.1\% 対 52.3\%) が、WikiText の複雑さは 16\% 増加しました (61.46 対 52.98)。この精度と複雑さのトレードオフは、割り当て目標が異なれば、ダウンストリーム メトリックも異なることが明らかになります。重み空間誤差と出力空間誤差の間の高い相関関係 ($>$0.99) により、割り当ての発散が制限され、効果の大きさが適度であることが説明されます。 20\% 圧縮の Llama-3.2-1B では、2 つの方法はほぼ同じ結果 (53.3\% 対 53.5\% の精度、14.45 対 14.66 PPL) を生成します。これは、圧縮率が低い場合にはコスト関数の影響が小さいことを示唆しています。
原文 (English)
Output-Space Allocation Costs for Calibration-Guided LLM Compression: An Empirical Study
Training-free compression methods for large language models (LLMs) often use calibration data to guide compression decisions. ROCKET, a recent method combining sparse-dictionary factorization with multi-choice knapsack problem (MCKP) allocation, derives its per-layer factorization from an output reconstruction objective but uses weight-space Frobenius error as the MCKP allocation cost. We investigate whether aligning the allocation cost with the output-space objective improves compressed model fidelity. On Qwen3-8B at 50\% compression, our ROCKET-ActCost achieves +0.8 percentage points higher average accuracy across 8 zero-shot benchmarks (53.1\% vs 52.3\%), but increases WikiText perplexity by 16\% (61.46 vs 52.98). This accuracy-perplexity tradeoff reveals that different allocation objectives favor different downstream metrics. The high correlation ($>$0.99) between weight-space and output-space errors limits allocation divergence, explaining the modest effect size. On Llama-3.2-1B at 20\% compression, the two methods produce near-identical results (53.3\% vs 53.5\% accuracy, 14.45 vs 14.66 PPL), suggesting that the effect of the cost function is minor at lower compression ratios.
SHIFT: 検索拡張生成における知識競合緩和のためのゲート変調アクティベーションステアリング
検索拡張生成 (RAG) は、外部の知識を組み込んで応答生成をサポートすることで LLM を強化します。ただし、取得されたコンテキストとパラメトリック知識の間の矛盾が、RAG システムにおける重大な課題として浮上しています。このような矛盾を軽減するために、生成中に文脈上の証拠に依存するLLMの能力を向上させることを目的として、知識関連の内部ニューロンを特定して編集することが数多くの研究で試みられてきた。ただし、変更されたニューロンはより広範なモデルの動作や機能と絡み合うことが多いため、これらのニューロン レベルのアプローチは、LLM の一般的な機能を損なう意図しないカスケード効果を導入する可能性があります。この論文では、ニューロンレベルの変更を学習可能なゲート変調として再定式化する新しいフレームワークである SHIFT を紹介します。これにより、LLM が知識の競合を解決するために内部活性化を適応的に制御できるようになります。技術的には、当社の SHIFT は LLM に軽量のゲート モジュールを装備し、バックボーン モデルをフリーズしたまま、0.01% 未満のトレーニング可能なパラメーターを最適化します。生成中に、ゲート モジュールはモデルの内部表現を調整して、コンテキストとパラメトリックの知識を適応的に活用します。 6 つのデータセットに対する広範な実験により、競合するさまざまなベースラインと比較して SHIFT の有効性が検証されます。すべてのデータセットとコードは https://github.com/OpenBMB/SHIFT で入手できます。
原文 (English)
SHIFT: Gate-Modulated Activation Steering for Knowledge Conflict Mitigation in Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) enhances LLMs by incorporating external knowledge to support response generation. However, conflicts between retrieved context and parametric knowledge have emerged as a critical challenge in RAG systems. To mitigate such conflicts, numerous studies have attempted to identify and edit knowledge-related internal neurons, aiming to improve the ability of LLMs to rely on contextual evidence during generation. However, these neuron-level approaches may introduce unintended cascading effects that compromise the general capabilities of LLMs, as the modified neurons are often entangled with broader model behaviors and functionalities. In this paper, we introduce SHIFT, a novel framework that reformulates neuron-level modification as learnable gate modulation, allowing LLMs to adaptively regulate internal activations for knowledge conflict resolution. Technically, our SHIFT equips LLMs with a lightweight gate module and optimizes fewer than 0.01% trainable parameters while keeping the backbone model frozen. During generation, the gate module adjusts the model's internal representations to adaptively leverage contextual and parametric knowledge. Extensive experiments on six datasets validate the effectiveness of our SHIFT in comparison with various competing baselines. All datasets and code are available at https://github.com/OpenBMB/SHIFT.
NLL ガイドによるフルアテンション レイヤー選択によるトレーニング不要のスライディング ウィンドウ適応
複数の層にわたってフル アテンションとスライディング ウィンドウ アテンションを混合するハイブリッド アテンション モデルは、効率的なロングコンテキスト推論への有望なアプローチを提供しますが、 \emph{どの層} がフル アテンションを保持すべきかという重要な問題は未解決のままです。既存の方法では、固定された周期パターンまたは注意ベースのヒューリスティックが使用されており、下流の精度にとって重要なものを捕捉できない可能性があります。我々は、NLL ガイドに基づく層選択を提案します。これは、層が完全な注意ではなくスライディング ウィンドウを使用する場合に、回答トークンの負の対数尤度の劣化を計算することによって各層の重要性を直接測定する、トレーニング不要の方法です。 Qwen3-4B を使用した LongMemEval では、私たちの方法は 1/4 の全注意層のみを使用して 64.6\% の精度を達成し、計算予算を半分にしながら 1/2-FA の周期ベースライン (65.0\%) と一致します。 NLL ガイドに基づく選択は、SWAA が報告した定期的な 1/4-FA ベースラインを 10.4 パーセント ポイント上回っており、一致する LightTransfer スタイルのベースラインを 26.4 パーセント ポイント上回っています。非交絡分析により、信号は一般的な層の感度ではなく、長距離の注意ニーズと一致していることが示されています。この方法では $\sim$15 分の 1 回限りのキャリブレーションのみが必要で、ロングコンテキスト LLM 導入の効率と精度のパレートフロンティアを前進させます。
原文 (English)
NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation
Hybrid attention models that mix full and sliding-window attention across layers offer a promising approach to efficient long-context inference, but the critical question of \emph{which layers} should retain full attention remains unsolved. Existing methods use either fixed periodic patterns or attention-based heuristics that may not capture what matters for downstream accuracy. We propose NLL-guided layer selection, a training-free method that directly measures each layer's importance by computing the negative log-likelihood degradation on answer tokens when that layer uses sliding-window instead of full attention. On LongMemEval with Qwen3-4B, our method achieves 64.6\% accuracy using only 1/4 full-attention layers, matching the 1/2-FA periodic baseline (65.0\%) while halving the computational budget. NLL-guided selection outperforms the SWAA-reported periodic 1/4-FA baseline by 10.4 percentage points and a matched LightTransfer-style baseline by 26.4 percentage points. De-confounding analysis shows the signal is consistent with long-range attention needs rather than generic layer sensitivity. The method requires only $\sim$15 minutes of one-time calibration, advancing the efficiency-accuracy Pareto frontier for long-context LLM deployment.
位置バイアス補正はワンパスアテンションソートには不十分
ロングコンテキスト言語モデルには位置バイアスがあり、中間位置の情報が十分に活用されません。アテンション ソートは、アテンション パターンに基づいてドキュメントの順序を繰り返し変更することでこの問題に対処しますが、複数の並べ替えと生成のサイクルにより導入コストが増加します。私たちは、位置バイアスが主なボトルネックであると仮説を立て、偏りのないワンパス アテンション ソートを提案します。これは、注目度の低い大部分の文書からプロンプトごとの位置バイアス曲線を推定し、それを使用して生の注意スコアを (減算または除算によって) 修正し、シングルパス ソートを可能にします。 2 つのモデルでの実験は、テスト設定でこの仮説を否定します。LLaMA-2-7B-32K-Instruct では、バイアス除去は未校正のシングルパス ソートと同じ結果 (94.83\% の封じ込め精度) を生成しますが、YaRN-Llama-2-7b-64k では、バイアス除去により精度が 8.67 パーセント向上しますが、反復ソートよりも 14.84 pp の差があり、結果の 37\% しか閉じません。ギャップ。これらの結果は、位置バイアス補正では反復ソートに適合するには不十分であり、並べ替えを繰り返すことでバイアス補正以上の利点がもたらされることを示唆しています。
原文 (English)
Position Bias Correction is Insufficient for One-Pass Attention Sorting
Long-context language models suffer from position bias, where information in middle positions is underutilized. Attention Sorting addresses this by iteratively reordering documents based on attention patterns, but its multiple sort-and-generate cycles increase deployment cost. We hypothesize that position bias is the primary bottleneck and propose Debiased One-Pass Attention Sorting, which estimates a per-prompt position-bias curve from the low-attention majority of documents and uses it to correct raw attention scores (via subtraction or division) to enable single-pass sorting. Our experiments on two models refute this hypothesis in the tested setting: on LLaMA-2-7B-32K-Instruct, debiasing produces identical results to uncalibrated single-pass sorting (94.83\% containment accuracy), while on YaRN-Llama-2-7b-64k, debiasing improves accuracy by 8.67 percentage points but remains 14.84pp behind iterative sorting, closing only 37\% of the gap. These results suggest that position-bias correction is insufficient to match iterative sorting, and that repeated reordering provides additional benefits beyond bias correction.
HPC システム上でスケーラブルな知識を蒸留するための教師と生徒のパーティショニングの最適化
Knowledge Distillation (KD) を使用すると、大規模な教師モデルの指導の下で小規模な生徒モデルをトレーニングできるようになり、広く採用されている TRL ライブラリがこれを実装します。しかし、TRL は両方のモデルを対称的に扱っており、メモリ フットプリントと通信要件における顕著な非対称性を利用する機会を逃しています。このペーパーでは、教師と生徒のパーティショニングを効率的に分離する、KD のための HPC 対応の方法論を紹介します。私たちのアプローチは、不必要な教師モデルのデータ構造を回避し、最適な分割戦略を選択することにより、TRL よりも最大 67% 高い 1 秒あたりのサンプル数を実現します。モデルの垂直分割と水平分割を組み合わせて、分割レジーム間の変曲点の存在を特定する分析式を導き出します。これらの結果は、トポロジーを意識した並列処理を通じて教師と生徒の非対称性を利用することで、当社の実稼働 HPC クラスターでの GKD トレーニングが著しく加速したことを示しました。
原文 (English)
Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems
Knowledge Distillation (KD) enables training smaller student models under the guidance of larger teacher models, and the widely adopted TRL library implements it. Yet, TRL treats both models symmetrically, missing opportunities to exploit their pronounced asymmetry in memory footprint, and communication requirements. This paper presents an HPC-aware methodology for KD that decouples teacher and student partitioning efficiently. Our approach achieves up to 67% higher samples-per-second than TRL by avoiding unnecessary teacher-model data structures and selecting the best split strategy. We combine vertical and horizontal partitioning of models, deriving an analytical expression that identifies the existence of inflection points between splitting regimes. These results showed that exploiting teacher--student asymmetry through topology-aware parallelism notably accelerated GKD training on production HPC clusters at our company
トラフィックマトリックス予測のためのパラメータ効率の高い量子にインスピレーションを受けた高速重み付けプログラマ
トラフィック マトリクス (TM) は、ネットワーク全体の起点と終点の需要を捉え、トラフィック エンジニアリングの中心となりますが、オンライン ネットワーク制御のメモリ、更新、トレーニング予算の制約の下で予測を実行する必要がある場合、マトリクス全体を正確に予測することは依然として困難です。この論文では、専用のグラフ、変換器、または拡散モジュールに依存せずに、コンパクトな量子にインスピレーションを得たリカレント モデルが効果的な TM 予測を提供できるかどうかを調査します。私たちは、ゲート量子にインスピレーションを得たコルモゴロフ・アーノルドネットワーク高速重み付けプログラマー (QKAN-FWP) を適応させて、マルチステップの Abilene TM 予測を指示します。各モデルは、2 時間の履歴から 144 チャネルの起点 - 終点 (OD) マトリックスの次の 20 の 5 分間のフレームを予測します。共有の固定予算トレーニング プロトコルの下で、サイズが一致した長短期記憶 (LSTM) ネットワーク、より大規模な LSTM、および古典的なゲート高速重みプログラマに対して 3 つの QKAN 配置バリアントのベンチマークを行います。評価されたリカレント モデルの中で、G-QKANFWP は、より大きな LSTM の 22.4% のみを使用しながら、最良のプール二乗平均平方根誤差 (RMSE) を達成します。また、サイズが一致した LSTM と従来の G-FWP ベースラインの両方を上回っており、ゲインがゲート高速重みフレームワークのみによるものではないことを示しています。収束分析とチャネルごとの分析では、量子にインスピレーションを得たバリアントは、サイズが一致したリカレントベースラインよりも学習曲線下検証損失領域 (AULC) が低く、G-QKANFWP と GQKAN-FWP が大幅に多くの OD チャネル勝利を達成していることがさらに示されています。これらの結果は、リソースを意識したネットワーク トラフィック マトリックス予測の精度と効率の有望な設計として、古典的な低速プログラマと量子にインスピレーションを受けた高速プログラマを特定します。
原文 (English)
Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting
Traffic matrices (TMs) capture network-wide origin-destination demand and are central to traffic engineering, yet accurate whole-matrix forecasting remains challenging when prediction must be performed under the memory, update, and training-budget constraints of online network control. This paper investigates whether compact quantum-inspired recurrent models can provide effective TM forecasts without relying on dedicated graph, transformer, or diffusion modules. We adapt gated quantum-inspired Kolmogorov-Arnold network fast-weight programmers (QKAN-FWPs) to direct multi-step Abilene TM forecasting, where each model predicts the next 20 five-minute frames of a 144-channel origin-destination (OD) matrix from a two-hour history. We benchmark three QKAN placement variants against a matched-size long short-term memory (LSTM) network, a larger LSTM, and a classical gated fast-weight programmer under a shared fixed-budget training protocol. Among the evaluated recurrent models, G-QKANFWP achieves the best pooled root-mean-square error (RMSE), while using only 22.4% of the larger LSTM. It also outperforms both the matched-size LSTM and the classical G-FWP baseline, indicating that the gain is not due to gated fast-weight framework alone. Convergence and channel-wise analyses further show that the quantum-inspired variants obtain lower validation-loss area under the learning curve (AULC) than matched-size recurrent baselines, while G-QKANFWP and GQKAN-FWP achieve substantially more OD-channel wins. These results identify a classical slow programmer with a quantum-inspired fast programmer as a promising accuracy-efficiency design for resource-conscious network traffic-matrix forecasting.
ペプチドドリフト: 抗原条件付き個別ペプチド生成のための毒性反発ドリフト
ペプチドは、小分子の化学的調整可能性と高分子治療薬の標的特異性を組み合わせた、有望な治療法です。しかし、毒性を回避しながら抗原特異的結合ペプチドを設計することは、治療用ペプチドの発見にとって依然として大きな課題である。ここでは、単一の抗原条件ドリフトステップを通じてペプチド候補を生成する、毒性を意識した潜在的精製フレームワークであるペプチドドリフトを紹介します。ペプチド包埋空間では、ペプチドドリフトは、生成された潜在ペプチドを抗原一致結合ペプチドに引き寄せる一方、毒性関連領域からは反発することを学習します。結合を促進する物理化学的特徴は、ペプチド表現空間における毒性関連の特徴と重なることが多いため、これは困難です。これに対処するために、最初に結合指向の引力を学習し、次に毒性の反発を高めることによって、この競合する目的を安定させるためのウォームアップ戦略を導入します。
原文 (English)
Pepti-drift: Toxicity-Repulsive Drifting for Antigen-Conditioned Discrete Peptide Generation
Peptides are a promising therapeutic modality that combine the chemical tunability of small molecules with the target specificity of macromolecular therapeutics. However, designing antigen-specific binding peptides while avoiding toxicity remains a major challenge for therapeutic peptide discovery. Here, we present Pepti-drift, a toxicity-aware latent refinement framework that generates peptide candidates through a single antigen-conditioned drift step. In a peptide embedding space, Pepti-drift learns to attract generated peptide latents toward antigen-matched binding peptides while repelling them from toxicity-associated regions. This is challenging because binding-promoting physicochemical features often overlap with toxicity-associated features in peptide representation space. To address this, we introduce a warm-up strategy to stabilize this competing objective by first learning binding-oriented attraction and then increasing toxicity repulsion.
Hippocampus-DETR: 海馬モデリングに基づく明示的メモリ オブジェクト検出フレームワーク
この論文は、現在の物体検出モデルにおける明示的な記憶メカニズムの欠如に対処し、生物学的海馬記憶モデリングに基づく新しい検出フレームワークである海馬-DETRを提案します。このフレームワークは、海馬の記憶ネットワーク モジュールである HipNet を DETR アーキテクチャに統合し、嗅内皮質、歯状回、CA3、CA1、海馬台などの海馬の小領域の解剖学的構造と機能組織を体系的にシミュレートします。この設計により、Hippocampus-DETR は、パターン分離、パターン補完、重要度フィルタリング、視覚的エンコーディング機能の情報統合を実現します。トレーニング中に、レイヤーごとのトレーニング戦略を使用してさまざまなメモリ サブモジュールが最適化され、最終的にメモリの検索および完了機能を備えたメモリ システムが形成されます。実験結果は、海馬-DETR が現在の主流モデルよりも高い検出精度を達成することを示しています。さらに重要なことは、このフレームワークを備えたモデルは、少数ショットの画像分類、マルチモーダル特徴の構築、画像復元などのタスクにおいても優れた汎化能力とデータ効率を発揮することです。その後の実験では、各メモリ サブモジュールの機能的必要性と内部解釈可能性がさらに検証されます。この研究は、新しい物体検出フレームワークを提供するだけでなく、神経認知メカニズムと深層学習モデルを統合するための実現可能な技術的経路も提供し、モデルの学習効率とタスクの堅牢性を向上させる上での重要な価値を強調しています。プロジェクトは https://github.com/2186cloud/hipnet で入手できます。
原文 (English)
Hippocampus-DETR: An Explicit Memory Object Detection Framework Based on Hippocampus Modeling
This paper addresses the lack of explicit memory mechanisms in current object detection models and proposes Hippocampus-DETR, a novel detection framework based on biological hippocampal memory modeling. This framework integrates a hippocampal memory network module, HipNet, into the DETR architecture and systematically simulates the anatomical structure and functional organization of hippocampal subregions, including the entorhinal cortex, dentate gyrus, CA3, CA1, and subiculum. Through this design, Hippocampus-DETR realizes pattern separation, pattern completion, importance filtering, and information integration of visual encoding features. During training, different memory submodules are optimized using a layer-wise training strategy, ultimately forming a memory system with memory retrieval and completion capabilities. Experimental results demonstrate that Hippocampus-DETR achieves higher detection accuracy than current mainstream models. More importantly, models equipped with this framework also exhibit excellent generalization ability and data efficiency in tasks such as few-shot image classification, multimodal feature construction, and image restoration. Subsequent experiments further validate the functional necessity and internal interpretability of each memory submodule. This study not only provides a novel object detection framework, but also offers a feasible technical pathway for integrating neurocognitive mechanisms with deep learning models, highlighting its significant value in improving model learning efficiency and task robustness. The project is available at https://github.com/2186cloud/hipnet.
WattLayer: レイヤーを適切に取得してニューラル ネットワークの推論エネルギーを推定する
人工知能 (AI) の普及により、エネルギー消費に対する懸念が高まっていますが、特にさまざまなタスクやアーキテクチャにわたって AI 推論のエネルギー消費を正確に推定するための標準化された方法論が不足しています。この研究では、AI アーキテクチャ向けのタスクに依存しないレイヤーごとのエネルギー推定モデルを提案します。私たちのモデルは、3 つの広く使用されているタスクと 3 つの異なるハードウェア プラットフォームにわたる 295 のニューラル ネットワーク アーキテクチャの 100,000 層を超える大規模なデータセットで評価されています。私たちのアプローチは 19.6% の中央値誤差を達成し、最先端の手法を上回ります。さらに、アーキテクチャ全体で共有レイヤーを活用することで、レイヤーごとの分解が完全な再トレーニングなしで新しいタスクに一般化されることを示します。関係者がエネルギー効率の高い AI システムを設計できるようにするためのツール、洞察、正確な方法論を提供します。
原文 (English)
WattLayer: Get Layers Right to Estimate Inference Energy of Neural Networks
The widespread adoption of Artificial Intelligence (AI) has led to increasing concerns about energy consumption, yet there is a lack of standardized methodologies to accurately estimate AI inference energy consumption, particularly across various tasks and architectures. In this study, we propose a task independent, layer-wise energy estimation model for AI architectures. Our model is evaluated on a large dataset of more than 100,000 layers for 295 neural network architectures across 3 widely-used tasks and 3 distinct hardware platforms. Our approach achieves a median error of 19.6%, outperforming state-of-the-art methods. We further show that layer-wise decomposition generalize to new tasks without complete retraining, by leveraging shared layers across architectures. It offer tools, insights and a precise methodology to empower stakeholders in designing energy-efficient AI systems.
低いサンプルサイズで sEMG デコーダを再調整する際の過学習を早期に発見するための記憶指標の適用性
完全に汎用的なデコーダーをトレーニングするために十分に大規模で多様なデータセットが利用できないため、表面筋電図 (sEMG) の深層学習モデルは、被験者固有の (再) キャリブレーションから大きな恩恵を受けることができます。ただし、ユーザーの受け入れを考慮すると、キャリブレーション中に実際に収集できる繰り返しの数は厳しく制限されているため、オーバーフィッティングのリスクが高まり、極端な場合には、キャリブレーションされていないモデルと比較してパフォーマンスが低下する可能性もあります。検証パフォーマンスや早期停止による正則化などの古典的なオーバーフィッティング指標は、実際のキャリブレーション シナリオではほとんど利用できない追加の保持データが必要となるため、この低サンプル領域では適用することが困難です。この研究では、ディープ ニューラル ネットワークの修正線形単位 (ReLU) の活性化統計のみに基づいて、最近提案された記憶指標のクラスを調査します。この統計は、追加の検証セットなしでトレーニング データから直接計算できます。ベンチマークの sEMG データセットで転移学習実験を行います。この実験では、畳み込みニューラル ネットワークが最初に複数の被験者で事前トレーニングされ、その後、少数の繰り返しのみを使用して個々のユーザーで微調整されます。キャリブレーション中、デコードパフォーマンスと最後の隠れ層のアクティブ化動作の両方を監視します。私たちの結果は、微調整中のテスト精度の低下が活性化率の特徴的な変化を伴うという最初の証拠を提供し、活性化ベースの記憶指標が低サンプルのsEMGキャリブレーション設定で失敗した学習を早期に発見するための有望なツールであることを示しています。
原文 (English)
Applicability of memorization indicators for early spotting of overfitting while recalibrating sEMG-decoders on low sample sizes
Deep learning models for surface electromyography (sEMG) can benefit substantially from subject-specific (re-)calibration, since no sufficiently large and diverse datasets are available to train fully generic decoders. However, for user acceptance, the number of repetitions that can realistically be collected during calibration is severely limited, which increases the risk of overfitting and, in extreme cases, can even degrade performance compared to the uncalibrated model. Classical overfitting indicators such as validation performance and regularization with early stopping are difficult to apply in this low-sample regime, as they require additional held-out data that is rarely available in practical calibration scenarios. In this work, we investigate a recently proposed class of memorization indicators based solely on the activation statistics of rectified linear units (ReLU) in deep neural networks, which can be computed directly from training data without any extra validation set. We conduct a transferlearning experiment on a benchmark sEMG dataset, where a convolutional neural network is first pre-trained on multiple subjects and subsequently fine-tuned on individual users using only a small number of repetitions. During calibration, we monitor both decoding performance and the activation behaviour of the last hidden layer. Our results provide first evidence that decreases in test accuracy during fine-tuning are ac companied by characteristic changes in activation rates, indicating that activation-based memorization indicators are a promising tool for early spotting of unsuccessful learning in low-sample sEMG calibration settings.
GNBAN: 大規模なエンティティ セットにわたる長期予測のためのグラフ ニューラル ベース アテンション ネットワーク
小売階層の最下位での需要予測には、製品、店舗、地域全体にわたって、相関する長期にわたる数万の系列を予測する必要があります。最新のシステムは、大規模なカタログにまたがって拡張し、共有の需要ダイナミクスをキャプチャし、信頼できる十分な解釈性を維持する必要があります。古典的な統計手法では系列ごとに個別のモデルが必要であり、大規模に管理するのが困難です。深い自己回帰モデルは、結合状態が数万次元に成長するにつれて苦戦します。また、最近のグラフベースの予測担当者は、エンティティ間の依存関係を把握しながら、不透明な長期予測を生成することがよくあります。我々は、異種グラフ表現学習と解釈可能な基底分解ヘッドを組み合わせたエンドツーエンドのアーキテクチャである GNBAN (Graph Neural Basis Attendance Network) を提案します。小売データはリレーショナル スキーマから派生した異種グラフとして直接表現されるため、単一のモデルがカタログ全体に対応します。 GNBAN は、期間を直接予測するのではなく、各予測をトレンド、季節、一般的な要素に分解します。その重要なイノベーションは、基底ごとのアテンション メカニズムです。各基底関数は、独自の学習可能なクエリを保持し、エンティティの歴史的近傍から独立して情報を取得します。これにより、解釈可能性を維持しながら、異なる基底を個別の時間パターンに特化させることができます。 M5 Walmart と Favorita Grocery Sales という 2 つの大規模ベンチマークを一致したプロトコルで評価すると、GNBAN はボリューム加重 WRMSSE を一致したグラフのベースラインと比較して約 4 ~ 5% 改善しました。定性分析では、学習された分解によって、事後的な説明方法を使用せずに、傾向、季節、および残余の需要要因が明らかになることを示しています。これらの結果は、統合されたグラフベースのフレームワークでスケーラブルなリレーショナル予測と解釈可能な予測分解を同時に実現できることを示しています。
原文 (English)
GNBAN: Graph Neural Basis Attention Networks for Long-Horizon Forecasting over Large Entity Sets
Demand forecasting at the bottom of a retail hierarchy requires predicting tens of thousands of correlated long-horizon series across products, stores, and regions. Modern systems must scale across massive catalogs, capture shared demand dynamics, and remain interpretable enough to be trusted. Classical statistical methods need a separate model per series and are hard to manage at scale; deep autoregressive models struggle as the joint state grows to tens of thousands of dimensions; and recent graph-based forecasters, while capturing cross-entity dependencies, often produce opaque long-horizon forecasts. We propose GNBAN (Graph Neural Basis Attention Network), an end-to-end architecture combining heterogeneous graph representation learning with an interpretable basis-decomposition head. Retail data are represented directly as a heterogeneous graph derived from the relational schema, so a single model serves the entire catalog. Rather than predicting the horizon directly, GNBAN decomposes each forecast into trend, seasonal, and generic components. Its key innovation is a per-basis attention mechanism: each basis function keeps its own learnable query and retrieves information independently from the entity's historical neighborhood, letting different bases specialize to distinct temporal patterns while preserving interpretability. On two large-scale benchmarks, M5 Walmart and Favorita Grocery Sales, evaluated under matched protocols, GNBAN improves volume-weighted WRMSSE by roughly 4-5% over a matched graph baseline. Qualitative analysis shows the learned decomposition exposes trend, seasonal, and residual demand drivers without post-hoc explanation methods. These results demonstrate that scalable relational forecasting and interpretable forecast decomposition can be achieved together in a unified graph-based framework.
S$^2$-VLA: 長距離操作のための状態空間ガイド付き視覚・言語・行動モデル
Vision-Language-Action (VLA) モデルはロボット操作において強力な機能を実証していますが、長期タスクでは累積的なエラーの伝播によりパフォーマンスが大幅に低下します。この制限は主に、固定重みに依存して視覚、言語、およびアクション表現を結合する静的特徴融合メカニズムに起因しており、モデルがタスク実行のさまざまなフェーズに適応するのを妨げています。この制限に対処するために、状態空間ガイド付き適応アテンション (SSGAA) メカニズムを導入するフレームワークである S$^2$-VLA を提案します。 SSGAA は、タスクの進行を追跡する信念状態を維持し、動的なゲート重みを生成して、空間認識のための視覚的特徴、高レベルのタスク計画のためのタスク意図、および実行の一貫性のための時間的アクション シーケンスの 3 つの相補的なソースからの情報を適応的に融合します。この適応型融合により、モデルはタスクの実行全体にわたって焦点を移し、さまざまなタスク段階の進化する要件に合わせることを可能にします。コンパクトな 2B パラメーター サイズにもかかわらず、S$^2$-VLA はより大きな 7B スケール モデルを常に上回り、LIBERO や SimplerEnv などの長期操作ベンチマークで最先端のパフォーマンスを達成します。これは、長期にわたるロボット操作における適応型特徴融合の重要性を強調しています。
原文 (English)
S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, but their performance degrades significantly in long-horizon tasks due to cumulative error propagation. This limitation largely arises from static feature fusion mechanisms that rely on fixed weights to combine visual, language, and action representations, preventing the model from adapting to different phases of task execution. To address this limitation, we propose S$^2$-VLA, a framework that introduces a State-Space Guided Adaptive Attention (SSGAA) mechanism. SSGAA maintains a belief state that tracks task progression and generates dynamic gating weights to adaptively fuse information from three complementary sources visual features for spatial perception, task intents for high-level task planning, and temporal action sequences for execution consistency. This adaptive fusion allows the model to shift its focus throughout task execution, aligning with the evolving requirements of different task stages. Despite its compact 2B parameter size, S$^2$-VLA consistently outperforms larger 7B-scale models and achieves state-of-the-art performance on long-horizon manipulation benchmarks, including LIBERO and SimplerEnv. highlighting the importance of adaptive feature fusion for long-horizon robotic manipulation.
SpatialUAV: 低高度 UAV の知覚、コラボレーション、およびモーションのための空間インテリジェンスのベンチマーク
空間インテリジェンスは、低高度の無人航空機 (UAV) の認識、コラボレーション、およびナビゲーションに不可欠です。しかし、既存の UAV ベンチマークは、多くの場合、画像レベルの認識、単一ビューの理解、または狭い回答形式を重視しており、3D 空間推論、マルチビュー コラボレーション、シーン ダイナミクス、および多様なタスクの定式化が十分に評価されていません。これらのギャップに対処するために、14 のきめ細かいタスク タイプにわたる 4,331 個の精選されたインスタンスで構成される実際の低高度 UAV ベンチマークである SpatialUAV を導入します。これは、意味論的識別、空間関係、航空と航空のコラボレーション、航空と地上のコラボレーション、および動作の理解をカバーします。 SpatialUAV は、すべてのサンプルを統一された視覚入力、質問、回答スキーマに編成し、オプション ラベル、領域識別子、幾何学的値、ビュー間の対応関係、および自由形式の動作記述を含む 7 つの入力構成と 9 つの回答形式をサポートします。信頼性の高い根拠に基づいた評価を保証するために、当社のデータ構築パイプラインは、検出器支援領域、深度監視、メタデータ由来のルール、広範な手動アノテーション、ブラインド フィルタリング、およびマルチターン人間による検証を、異種出力に対するタスク固有のメトリクスとともに統合します。 3 つのカテゴリにわたる代表的な視覚言語モデルを評価したところ、現在のモデルは依然として人間レベルのパフォーマンスには遠く及ばず、視点間の関連性、構造化された基礎、幾何学的な推論、および時間的視点の理解において顕著なボトルネックがあることが示されました。これらの結果は、低高度 UAV の空間インテリジェンスを進歩させるための経験的な指針を提供します。コードとデータは https://github.com/Hyu-Zhang/SpatialUAV で入手できます。
原文 (English)
SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion
Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation. However, existing UAV benchmarks often emphasize image-level recognition, single-view understanding, or narrow answer formats, leaving 3D spatial inference, multi-view collaboration, scene dynamics, and diverse task formulations insufficiently evaluated. To address these gaps, we introduce SpatialUAV, a real low-altitude UAV benchmark comprising 4,331 curated instances across 14 fine-grained task types, covering semantic discrimination, spatial relation, aerial--aerial collaboration, aerial--ground collaboration, and motion understanding. SpatialUAV organizes all samples into a unified visual-input--question--answer schema, while supporting seven input configurations and nine answer formats, including option labels, region identifiers, geometric values, cross-view correspondences, and free-form motion descriptions. To ensure reliable and grounded evaluation, our data construction pipeline integrates detector-assisted regions, depth supervision, metadata-derived rules, extensive manual annotation, blind filtering, and multi-turn human validation, together with task-specific metrics for heterogeneous outputs. Evaluating representative vision-language models across three categories, we show that current models remain far from human-level performance, with pronounced bottlenecks in cross-view association, structured grounding, geometric reasoning, and temporal viewpoint understanding. These results offer empirical guidance for advancing low-altitude UAV spatial intelligence. Code and data are available at https://github.com/Hyu-Zhang/SpatialUAV.
歴史文書における固有表現認識のための時間融合戦略に関する研究
時間的変動は、歴史文書における固有表現認識 (NER) に独特の課題をもたらします。この場合、実体は時間の経過とともに表面の形状や顕著性が変動します。言語モデル (LM) はさまざまな NLP タスクで進歩を遂げていますが、特に通時的なコンテキストにおいて、時間性を推論する能力は依然として限られているか、少なくとも疑わしいです。この論文では、さまざまな軽量融合戦略を使用して、時間メタデータを NER モデルにどのように構造的に埋め込むことができるかを系統的に研究します。私たちは、クロスアテンション、アダプター、連結などの初期または後期融合メカニズムを介して Transformer ベースのアーキテクチャに注入された、絶対的時間表現と相対的時間表現の両方を実験します。フランスとドイツの歴史的データセットに対する我々の評価では、後期融合戦略が、特に初期のノイズの多い時期において、より堅牢で時間的に一般化可能なパフォーマンスを生み出すことが明らかになりました。
原文 (English)
A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts
Temporal variation poses a unique challenge for named entity recognition (NER) in historical texts, where entities drift in surface form and salience across time. While language models (LMs) have made progress in various NLP tasks, their ability to reason about temporality, especially in diachronic contexts, remains limited or at least, questionable. In this paper, we systematically study how temporal metadata can be structurally embedded into NER models using a range of lightweight fusion strategies. We experiment with both absolute and relative temporal representations, injected into Transformer-based architectures via early or late fusion mechanisms such as cross-attention, adapters, and concatenation. Our evaluations on French and German historical datasets reveal that late fusion strategies yield more robust and temporally generalisable performance, particularly in early and noisy periods.
SEADA: 多精度空間アーキテクチャ上で混合精度 DNN を最適化するための効率的な方法論
混合精度計算は、遅延、エネルギー消費、メモリ使用量を削減する効果的なアプローチとしてディープ ニューラル ネットワーク (DNN) に導入されています。ただし、混合精度ネットワークを多精度空間アーキテクチャに効率的にマッピングするには、いくつかの課題が生じます。これには、各レイヤーの適切な精度を決定すること、量子化に対するレイヤーごとの精度感度とアーキテクチャーの異質性およびシステムレベルの制約のバランスをとること、異種精度の割り当てによるシステムレベルのコストを正確に見積もることが含まれます。この研究では、これらの課題に対処するために設計された効率的な方法論である SEADA を紹介します。 SEADA は以下で構成されます。(i) 多精度空間アクセラレータ アーキテクチャの構成可能なシステム レベルの分析コスト モデル。 (ii) ターゲット整数アクセラレータへの DNN ワークロードの最適に近いマッピングを特定する高速マッピング ツール。 (iii) 混合精度実行の全体的な利点を推定するための浮動小数点層の分析モデル。 (iv) ビットレベルのエントロピーに基づくレイヤーごとの精度選択手法により、複数の数値精度にわたる効率的な割り当てが可能になります。 SEADA の効率性により、多精度アーキテクチャの設計空間を探索するための堅牢なフレームワークが設計者に提供されます。
原文 (English)
SEADA: An efficient methodology for optimizing mixed-precision DNNs on multi-precision spatial architectures
Mixed-precision computation has been introduced in deep neural networks (DNNs) as an effective approach to reduce latency, energy consumption, and memory footprint. However, efficiently mapping mixed-precision networks onto multi-precision spatial architectures poses several challenges. These include determining the appropriate precision for each layer, balancing layer-wise accuracy sensitivity to quantization against architectural heterogeneity and system-level constraints, and accurately estimating the system-level cost of heterogeneous precision assignments. This work presents SEADA, an efficient methodology designed to address these challenges. SEADA comprises: (i) a configurable system-level analytical cost model of a multi-precision spatial accelerator architecture; (ii) a fast mapping tool that identifies near-optimal mappings of DNN workloads onto the target integer accelerator; (iii) analytical models for floating-point layers to estimate the overall benefits of mixed-precision execution; and (iv) a per-layer precision selection methodology based on bit-level entropy, enabling efficient assignment across multiple numerical precisions. SEADA's efficiency provides designers with a robust framework for the design-space exploration of multi-precision architectures.
Triadic Werewolf: LLM における心のマルチホップ理論における道化師の役割
大規模な言語モデルの心の理論による評価では通常、二項社会演繹ゲームが使用されます。このゲームでは、観察可能なすべての手がかりが 1 つの隠れた側面を指すため、強力な言語事前条件を持つモデルは、対戦相手のインセンティブをシミュレートすることなく高いスコアを獲得できます。人狼ゲームを道化師で拡張します。道化師は、投票によって勝利するため、同僚の疑いに対する効用が逆転する 3 番目の派閥です。そのため、最適なプレイには 3 つの相反する効用関数にわたる推論が必要です。 GPT-4.1、DeepSeek-V3.1、および Llama-3.3-70B で道化師の自己学習をオンまたはオフにして 60 試合を行ったところ、道化師はゲームの 60 ~ 70% で勝利しましたが、ウェアウルフが 20% を超えることはありませんでした。また、GPT-4.1 のオオカミはゲームの 60 ~ 70% で初日に道化師を投票で除外します。これは厳密に自己破滅的なアクションです。自己学習は DeepSeek と Llama には役立ちますが、GPT-4.1 には害があり、そのコストは狼男ではなく村人にかかっています。 DeepSeek だけが、意図的に疑わしく見せることなく、疑わしく見せるという微妙な戦略を学習し、ループから最大限の利益を得ます。三項インセンティブ構造は、二項演繹ゲームが目に見えないままにしていたマルチエージェント推論の層を明らかにします。
原文 (English)
Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs
Theory-of-mind evaluations of large language models typically use dyadic social-deduction games, where every observable cue points to a single hidden side, so a model with strong language priors can score well without ever simulating opponents' incentives. We extend the Werewolf game with a Jester, a third faction whose utility on peer suspicion is inverted because it wins by being voted out, so optimal play requires reasoning across three opposing utility functions. Across 60 games on GPT-4.1, DeepSeek-V3.1, and Llama-3.3-70B with Jester self-learning on and off, the Jester wins 60-70% of games while Werewolves never exceed 20%, and GPT-4.1 wolves vote the Jester out on day 1 in 60-70% of games, a strictly self-defeating action. Self-learning helps DeepSeek and Llama but hurts GPT-4.1, with the cost landing on Villagers rather than Werewolves. Only DeepSeek learns the subtle strategy of looking suspicious without looking intentionally suspicious, and it gains the most from the loop. Triadic incentive structure exposes a layer of multi-agent reasoning that dyadic deduction games leave invisible.
あらゆるステップ: ビデオベースのパーキンソン病患者の方向転換歩数計測
パーキンソン病 (PD) の顕著な症状として、旋回障害は、旋回角度、持続時間、特に旋回を完了するまでに必要な歩数などのパラメータを通じて評価され、運動機能障害を直接反映します。現実世界の回転動作にはばらつきがあり、パーキンソン病の歩行では非定型的な足を引きずるパターンがあるため、正確な歩数計測は困難です。既存の方法は主にウェアラブルベースであり、ユーザーは専用デバイスを装着して管理する必要があり、毎日継続的に使用するには不便な場合があります。これに対処するために、多様な動作表現を使用して粗い方法から細かい方法まで歩数を推定する、受動的なビデオベースのフレームワークを提案します。具体的には、3D ヒューマン メッシュの復元から得られた足の動きの信号から初期歩数を推定し、高レベルの動作構造を提供します。きめの細かいモーションの詳細を組み込むために、モーション エンコーダーはメッシュとオプティカル フローから相補的な歩行ダイナミクスを学習し、初期推定を改良します。このプロセスでは、粗い足の動きの信号が交差注意を介してピクセルレベルの動きの合図を照会し、微妙なパーキンソン病の歩行ダイナミクスを捕捉します。さまざまなビデオの長さを処理するために、各ビデオをクリップに分割し、歩数残差予測のためにマルチ インスタンス学習 (MIL) を介してクリップごとのモーション エンベディングを統合します。広範な実験により、私たちの方法は現実世界の PD 旋回データセットで既存の歩数計数方法よりも一貫して優れていることが示されています。
原文 (English)
Every Step of the Way: Video-based Parkinsonian Turning Step Counting
As a prominent symptom of Parkinson's disease (PD), turning impairment is evaluated through parameters such as turning angle, duration, and particularly, the number of steps required to complete a turn, which directly reflects motor dysfunction. Accurate step counting is challenging due to variability in real-world turning movements and atypical shuffling patterns in parkinsonian gait. Existing methods are predominantly wearable-based, requiring users to wear and manage dedicated devices, which can be inconvenient for continuous daily use. To address this, we propose a passive, video-based framework that estimates step count in a coarse-to-fine manner using diverse motion representations. Specifically, an initial step count is estimated from foot movement signals derived from 3D human mesh recovery, providing high-level motion structures. To incorporate fine-grained motion details, a motion encoder learns complementary gait dynamics from mesh and optical flow to refine the initial estimate. In this process, coarse foot movement signals query the pixel-level motion cues via cross attention to capture subtle parkinsonian gait dynamics. To handle varying video lengths, we partition each video into clips and integrate clip-wise motion embeddings via multiple instance learning (MIL) for step count residual prediction. Extensive experiments show our method consistently outperforms existing step counting methods on real-world PD turning datasets.
Reflect-R1: 長いビデオの理解における自己修正のための証拠に基づくリフレクション
長時間のビデオを理解するための現在のマルチモーダル反射メカニズムは、主に内部パラメータ内の閉ループ自己反射に依存しています。客観的な外部証拠が欠如しているため、モデルはしばしば盲目的な自信に囚われ、エラーを修正できないことがよくあります。さらに、強化学習を多段階リフレクション パイプラインに適用すると、深刻なポリシー結合が導入され、専用のトレーニング データが重大に不足することでさらに悪化します。これらの制限に対処するために、この研究では、長いビデオを理解するための初の証拠主導型自己修正フレームワークである Reflect-R1 を提案しています。このフレームワークは、直感、検証、調停からなる 3 段階のパイプラインを構築します。客観的な視覚的証拠を動的に取得して最初の直観を検証し、複数の時間的検索を自律的に実行して矛盾を解決することで、幻覚ループを完全に断ち切ります。ポリシー結合を克服するために、さまざまな推論段階にわたって利点関数を独立して計算する、SD-GRPO という名前の段階分離型強化学習アルゴリズムを設計します。同時に、トレーニング データのギャップを埋めるために 120,000 サンプルのデータセットを構築します。 VideoMME や LongVideoBench などのベンチマークに関する広範な実験により、Reflect-R1 が最先端のパフォーマンスを達成していることが実証されています。私たちの方法は真の修正率を大幅に向上させ、厳密に客観的な証拠に基づいた真の自己修正を可能にします。
原文 (English)
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evidence, models are frequently trapped in blind confidence and often fail to correct errors. Furthermore, applying reinforcement learning to multi-stage reflection pipelines introduces severe policy coupling, which is exacerbated by a critical scarcity of dedicated training data. To address these limitations, this work proposes Reflect-R1, the first Evidence-Driven self-correction framework for long video understanding. The framework constructs a three-stage pipeline consisting of intuition, verification, and arbitration. By dynamically retrieving objective visual evidence to verify initial intuitions and autonomously executing multiple temporal searches to resolve conflicts, it completely breaks the hallucination loop. To overcome policy coupling, we design a stage-decoupled reinforcement learning algorithm named SD-GRPO that independently computes advantage functions across different reasoning stages. Concurrently, we construct a dataset of 120K samples to bridge the training data gap. Extensive experiments on benchmarks such as VideoMME and LongVideoBench demonstrate that Reflect-R1 achieves state-of-the-art performance. Our method significantly improves the genuine rectification rate and enables authentic self-correction strictly grounded in objective evidence.
Home3D 1.0: インテリア デザイン向けの高忠実度画像から 3D アセット生成システム
Home3D 1.0 は、インテリア デザインと電子商取引アプリケーションを対象とした、単一の参照画像から高品質の 3D アセットを生成するモジュール式の画像から 3D 生成システムです。家具や装飾品の写真が与えられると、システムは物理ベース レンダリング (PBR) マテリアルを使用してメッシュを出力し、メッシュはマテリアル固有のコンポーネントに分解できます。パイプラインは、密接に結合された 4 つのモジュールで構成されます。 ジオメトリは、ジオメトリ VAE と粗いから細かいフローマッチング DiT を使用した潜在的な SDF モデリングを通じて防水メッシュを再構築します。テクスチャは、マルチビュー アルベド観測を予測し、それらをメッシュ上に再投影し、目に見えない表面領域を 3D テクスチャ フィールドで完成させます。マテリアルは MatWeaver を使用して、ビデオベースのセグメンテーションと UV 空間投票を通じてコンポーネント マスクを取得し、階層的なマルチモーダル マッチングを通じて厳選されたマテリアル ライブラリから PBR マップを取得してベイクします。そして、Parts は、PartVAE および PartDiT を使用してマテリアル編集可能なセマンティック パーツ メッシュを生成し、マルチヘッド パーツ固有の SDF フィールドを 1 つのパスでデコードします。各モジュールは専用のメトリックを使用して個別に評価され、現在のシステム機能と、より広範な導入に向けた残りのギャップの両方が強調表示されます。
原文 (English)
Home3D 1.0: A High-Fidelity Image-to-3D Asset Generation System for Interior Design
We present Home3D 1.0, a modular image-to-3D generation system that produces high-quality 3D assets from a single reference image, targeting interior design and e-commerce applications. Given a photograph of a furniture or decor item, the system outputs a mesh with physically-based rendering (PBR) materials, and the mesh can be decomposed into material-specific components. The pipeline is organized into four tightly coupled modules: Geometry reconstructs a watertight mesh through latent SDF modelling with a geometry VAE and a coarse-to-fine flow-matching DiT; Texture predicts multiview albedo observations, reprojects them onto the mesh, and completes unseen surface regions with a 3D texture field; Material uses MatWeaver to obtain component masks through video-based segmentation and UV-space voting, then retrieves and bakes PBR maps from a curated material library through hierarchical multi-modal matching; and Parts generates material-editable semantic part meshes with a PartVAE and PartDiT, decoding multi-head part-specific SDF fields in one pass. Each module is evaluated independently with dedicated metrics, highlighting both the current system capability and the remaining gaps toward broader deployment.
Agentic AI を活用した再識別: モビリティ マイクロデータ プライバシーに対する新たなスケーラブルな脅威
商業データブローカーによる詳細な位置データの広範な収集は、一般に広く認識されていない再識別リスクを生み出します。これまでの研究では、モビリティの追跡は非常にユニークであり、原理的には少数の時空間ポイントから個人を特定できることが証明されていますが、このような攻撃にはこれまで、熟練したアナリストによる多大な手作業が必要であり、実際の規模は限られていました。この実現可能性調査では、エージェント AI がこの脅威モデルを根本的に変えることを現実の環境で実証します。私たちは、人間の介入なしに、大規模な言語モデル エージェントが自律的にオープン Web を検索し、公的記録とソーシャル メディアを相互参照し、生の座標シーケンスを解決して候補 ID を解決する、エンドツーエンドのパイプラインを提示します。高リスクの開示シナリオに焦点を当てて、実際の自宅および勤務先の住所およびその周辺に固定されたシミュレートされた位置ポイントを含む時空間データセットでパイプラインを評価します。私たちの結果は、時空間データと公的情報源だけから、エージェント AI が再識別可能な個人 25 人中 18 人 (72%) と、全体のケース 43 人中 18 人 (41.9%) を再識別することに成功したことを示しています。統計的開示管理 (SDC) の実践への影響について議論し、データ管理者と規制当局が予期しなければならない近い将来のエスカレーションについて概説します。 SDC 実践の暗黙の基盤である事実上の匿名性は変化しつつあります。 Agentic AI は、GDPR Recital-26 基準に基づく何らかの手段で、ターゲットごとに数分と数ドルのコストをかけて再識別が行われる可能性がかなり高いという主張を強化します。
原文 (English)
Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy
The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by the public. While prior research has established that mobility traces are highly unique and that individuals can, in principle, be identified from a handful of spatio-temporal points, such attacks have historically required significant manual effort from skilled analysts, limiting their practical scale. In this feasibility study, we demonstrate in a real world setting that agentic AI fundamentally changes this threat model. We present an end-to-end pipeline in which large language model agents autonomously search the open web, cross-reference public records and social media, and resolve raw coordinate sequences to candidate identities - without human intervention. We evaluate the pipeline on a spatio-temporal dataset containing simulated location points anchored at and around true home and work addresses, focusing on a high-risk disclosure scenario. Our results demonstrate that, from spatio-temporal data and public sources alone, our agentic AI successfully re-identified 18 of the 25 re-identifiable individuals (72%) and 18 of 43 cases overall (41.9%). We discuss implications for Statistical Disclosure Control (SDC) practice and outline the near-future escalation that data custodians and regulators must anticipate. De facto anonymity - an implicit foundation of SDC practice - is shifting. Agentic AI strengthens the case that re-identification is reasonably likely by any means under the GDPR Recital-26 standard, at costs of minutes-and-dollars per target.
標的アミノ酸組成によるタンパク質配列生成のための 2 段階の微調整
タンパク質言語モデルは生物学的配列生成の標準的な事前分布ですが、それを明示的な分布設計ターゲットに向けて誘導する方法は、ほとんど解明されていないままです。私たちは、妥当な配列統計と多様性を維持しながら、配列が目的のアミノ酸 (AA) 組成プロファイルに一致しなければならないという制約されたタンパク質生成問題を研究します。応用の動機は、合成飼料タンパク質の設計であり、食事タンパク質の AA 組成がその栄養価を直接決定します。我々は、ドメイン内タンパク質データセットに対するドメイン適応微調整 (FT) の後に、凍結参照として FT モデルに対して固定された強化学習 (RL) を介した報酬重み付き反復 FT が続く 2 段階のパイプラインを提案します。 2 つの AA 構成でパイプラインを評価したところ、FT が平均構成をターゲットに近づける一方、後続の RL は FT だけでは満たせない特定のシーケンス制約を強制することがわかりました。さらに、提案された構成報酬項の設計の選択を 2 つのベースラインと除去されたバリアントに対して評価し、各トレーニング段階の寄与を分離し、配列の品質を低下させることなく AA 構成のアラインメントが達成されることを検証します。
原文 (English)
Two-Stage Fine-Tuning for Protein Sequence Generation with Targeted Amino-Acid Composition
Protein language models are standard priors for biological sequence generation, but steering them toward explicit distributional design targets remains largely unexplored. We study a constrained protein generation problem in which sequences must match a desired amino-acid (AA) composition profile while preserving plausible sequence statistics and diversity. The motivating application is synthetic feed protein design, where the AA composition of dietary proteins directly determines their nutritional value. We propose a two-stage pipeline in which domain-adaptive fine-tuning (FT) on an in-domain protein dataset is followed by iterative reward-weighted FT via reinforcement learning (RL) anchored against the FT model as a frozen reference. We evaluate the pipeline on two AA compositions and find that FT brings the average composition close to the target, while the subsequent RL enforces specific sequence constraints that FT alone cannot satisfy. We additionally evaluate the design choices of the proposed composition reward term against two baselines and an ablated variant, isolate the contribution of each training stage, and verify that AA composition alignment is achieved without degrading sequence quality.
VASAE: 語彙に合わせたアンカーによる SAE 辞書の方向の命名
スパース オートエンコーダ (SAE) は、Transformer の残差ストリームの有用な分解を提供しますが、学習された特徴は通常、Transformer のトークン ボキャブラリに直接関連付けられるのではなく、事後的に名前が付けられます。 VASAE (Vocabulary-Aligned Sparse Autoencoder) を導入します。これは、語彙に合わせたアンカーリングの下で SAE 特徴をトレーニングし、各特徴に固有のトークン名 (埋め込みがその特徴に最も近いトークン文字列) を割り当てる方法です。 VASAE は、標準 SAE と比較して再構築の品質を低下させることなく、語彙に合わせた特徴を備えた辞書を生成します。最近傍トークン アライメント スコアの 0.8 カットオフを使用すると、GPT-2 の小さなポスト残差ストリームでトレーニングされた辞書は、レイヤー 0 ~ 10 の特徴の約 90% をアライメントします。 Llama-3.1-8B では、代表的な浅層および中間層の辞書には、92.8% が浅層に含まれるなど、強く位置合わせされた特徴が含まれていますが、代表的な最終層の辞書は限定的な位置合わせを示しています。文レベルの平均スパース コードを差し引いた後、ケース スタディでは、残りの多くの組み込みトークン名が近くの入力トークンに関連していることが示されています。これらの結果は、語彙に合わせたアンカリングがトレーニング中に学習された特徴を固有のトークン名に結び付け、学習された辞書の事後解釈を補完できることを示唆しています。
原文 (English)
VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring
Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post hoc rather than directly connected to the Transformer's token vocabulary. We introduce Vocabulary-Aligned Sparse Autoencoder (VASAE), a method that trains SAE features under vocabulary-aligned anchoring and assigns each feature an intrinsic token name: the token string whose embedding is nearest to that feature. Without reducing reconstruction quality compared with a standard SAE, VASAE produces dictionaries with vocabulary-aligned features. Using a 0.8 cutoff on the nearest-token alignment score, dictionaries trained on GPT-2-small post-residual streams align about 90% of features in layers 0--10. In Llama-3.1-8B, representative shallow and middle-layer dictionaries contain strongly aligned features, including 92.8% in the shallow layer, while the representative final-layer dictionary shows limited alignment. After subtracting the sentence-level mean sparse code, case studies show that many remaining intrinsic token names are relevant to nearby input tokens. These results suggest that vocabulary-aligned anchoring can connect learned features to intrinsic token names during training, complementing post hoc interpretation of learned dictionaries.
予測を超えた推論: データ駆動型から因果関係のあるソフトウェア エンジニアリングへ
ソフトウェア エンジニアリングは、複雑化するシステムの設計、構築、品質の保証を行うために、相互依存するタスクの網をうまくやりくりする、知的要求が高く創造的な学問です。 AI 主導の製品、広く分散されたクラウド ネイティブ アーキテクチャ、深く組み込まれたサイバー物理環境に及ぶ需要により、ソフトウェアに対する私たちの期待が高まるにつれて、その複雑さは着実に増加しています。これに応えて、プロセスを強化し、自動化と意思決定のサポートを強化するために、ディープラーニングを活用したコエンジニアリング手法とツールの新しい波が現れました。しかし、これらの進歩は、現代のソフトウェア開発が要求する種類のインテリジェントなサポートを提供するには程遠いです。私たちは人間と機械の協力に関する新しいパラダイムを求めています。それは、機械がルーチンタスクを自動化したり、学習したパターンから予測したりするだけでなく、因果関係のレンズを通してエンジニアの推論を積極的に増幅するパラダイムです。ソフトウェアが賢くなるにつれて、より賢いサポートが必要になります。
原文 (English)
Reasoning Beyond Prediction: From Data-Driven to Causal Software Engineering
Software engineering is an intellectually demanding, creative discipline that juggles a web of interdependent tasks to design, build, and assure the quality of increasingly complex systems. As our expectations from software soar - with demands spanning AI-driven products, pervasively distributed and cloud-native architectures, and deeply embedded cyber-physical environments - its complexity steadily increases. In response, a new wave of co-engineering methods and tools, fueled by deep learning, has emerged to augment the process, enhancing automation and decision support. Yet, these advances remain far from delivering the kind of intelligent support that modern software development demands. We call for a new paradigm of human-machine cooperation: one where machines don't just automate routine tasks or predict from learned patterns, but actively amplify engineers' reasoning through the lens of causation. As software becomes smarter, a smarter support is needed.
ブラックボックスから臨床洞察へ: 音声ベースの認知障害検出のための多段階の説明可能なフレームワーク
音声ベースの認知障害検出は、高価なバイオマーカーアッセイに代わる非侵襲的で利用しやすい代替手段を提供しますが、トランスフォーマーベースのモデルは依然として臨床的に解釈できません。我々は、SHapley Additive exPlanations (SHAP) ベースのトークン帰属、理論に基づいた言語特徴、LLaMA-3.1-70B-Instruct を使用した 4 段階の LLM 推論パイプラインを統合することにより、ブラック ボックス トランスフォーマーの予測を臨床に基づいた物語に変換する、多段階の説明可能性フレームワークを提案します。 SpeechCARE-Adaptive Gating Network マルチモーダル スクリーニング モデル (NIA PREPARE ベンチマークで F1 = 72.11%) に基づいて構築されたこのフレームワークは、語彙の豊富さ、構文の複雑さ、意味の一貫性を含む 4 つの認知言語学的次元にモデルの出力をマッピングします。 70 の層別英語サンプルに対する医師の評価では、患者レベルの認知プロファイルとの強い一致が実証され、システム使用性スケール スコア 82/100 は臨床ワークフロー統合の高い可能性を示しました。
原文 (English)
From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection
Speech-based cognitive impairment detection offers a noninvasive, accessible alternative to costly biomarker assays, yet transformer-based models remain clinically uninterpretable. We propose a multi-stage explainability framework that translates black-box transformer predictions into clinically grounded narratives by integrating SHapley Additive exPlanations (SHAP)-based token attribution, theory-informed linguistic features, and a four-stage LLM reasoning pipeline using LLaMA-3.1-70B-Instruct. Built on the SpeechCARE-Adaptive Gating Network multimodal screening model (F1 = 72.11% on the NIA PREPARE benchmark), the framework maps model outputs to four cognitive-linguistic dimensions, including lexical richness, syntactic complexity, and semantic coherence. Physician evaluation on 70 stratified English samples demonstrated strong alignment with patient-level cognitive profiles, and a System Usability Scale score of 82/100 indicated high potential for clinical workflow integration.
ProMSA:知識ベースの視覚的な質問応答のためのプログレッシブ マルチモーダル検索エージェント
知識ベースのビジュアル質問応答 (KB-VQA) では、画像の理解を外部知識と組み合わせるモデルが必要です。従来のメソッドのほとんどは、事前に選択された取得者と静的な top-k 設定を備えた固定の取得後生成パイプラインを使用しますが、これは推論中に適応的ではありません。私たちは、KB-VQA のプログレッシブ マルチモーダル検索エージェントである ProMSA を提案します。画像と質問のペアが与えられると、エージェントは、明示的なツール呼び出しバジェットに基づいて、冗長な取得を避けるための重複排除を使用して、画像検索、テキスト検索、または停止を繰り返し選択します。トレーニングでは、最初に拒否サンプリング SFT を使用して有効なツール使用形式を学習し、次に TN-GSPO でエージェントを最適化します。TN-GSPO は、世代長とツール インタラクションの深さの両方によって更新を正規化するシーケンス レベルの RL 目標です。 E-VQA と InfoSeek の実験では、強力な RAG とエージェントのベースラインを超えて一貫した向上が見られ、検索とエンドツーエンドの精度が向上しました。コードは https://github.com/DingWu1021/Promsa で入手できます。
原文 (English)
ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowledge. Most prior methods use a fixed retrieve-then-generate pipeline with a pre-selected retriever and a static top-k setting, which is not adaptive during reasoning. We propose ProMSA, a progressive multimodal search agent for KB-VQA. Given an image-question pair, the agent iteratively chooses image search, text search, or stop, under explicit tool-call budgets and with deduplication to avoid redundant retrieval. For training, we first use rejection-sampling SFT to learn valid tool-use formats, then optimize the agent with TN-GSPO, a sequence-level RL objective that normalizes updates by both generation length and tool-interaction depth. Experiments on E-VQA and InfoSeek show consistent gains over strong RAG and agent baselines, and improved retrieval and end-to-end accuracy. The code is available at https://github.com/DingWu1021/Promsa.
SHARD: アライメント耐性のあるプライベート密検索のためのセルキー残差分割
高密度の埋め込みはセマンティック検索と RAG を支えていますが、漏洩したベクトル ストアにより、基礎となるテキストの多くがそれを保持している人の手に渡されます。これを可能にする攻撃 (少数ショットのアライメント、ゼロショットの反転、教師なしクロススペース変換) には 1 つの弱点があります。それは、保護されたストアが、既知のジオメトリにアライメントできる単一のグローバル ジオメトリであるということです。通常の軽量防御である秘密のグローバル回転も例外ではありません。攻撃者が既知のペアでほぼ亜空間次元を取得すると、直交プロクラステスはそれを回復します。この弱い軸を取り除く検索保存埋め込み変換である Shard を紹介します。中央に配置された埋め込みは、短いパブリック プレフィックス (ステージ 1 取得用) と、個別の秘密キーの下で C セルにシャーディングされたプライベート残差に分割されます。残差は CKKS の下で再ランク付けされ、キーはキャンセルされて内積が正確のままになります。単一のパラメーター C は、置き換えられるグローバル線形ベースライン (C=1) からドキュメントごとのマイクロキー (C=N) までデザインを実行します。リランクはフル次元であるため、Shard は、半 SVD 切り捨てが放棄された生の空間 nDCG@10 を返します。また、残差はセルローカルにキー付けされるため、拡散既知平文リークのもとで残差を共通フレームにマッピングし直すには、暗号化されたクエリ数が少ない場合、およそ C 倍のアンカー (C=256 で中央値 200 ~ 102,400) のコストがかかります。短いパブリック プレフィックスにより、近隣構造の漏洩がはるかに少なくなり、マイクロキー制限により、リンク不可能で更新可能なテンプレートで残差グラフがゼロになります。この障壁は、学習済み、非線形、および教師なしのアライナに対して保持されており、整合ユーティリティ ノイズ防御がほぼすべてのプローブを匿名化解除するのに対して、シャードは匿名化を解除しません。私たちはその制限について明確にしています。セル内ではキーがキャンセルされ、標的型攻撃者が必要とするのは d_priv アンカー程度だけであり、重複する参照コーパスは依然としてプレフィックスを介して漏洩します。シャードは攻撃を認識する幾何学的防御であり、暗号化を保証するものではありません。
原文 (English)
SHARD: cell-keyed residual splitting for alignment-resistant private dense retrieval
Dense embeddings underpin semantic search and RAG, yet a leaked vector store hands much of the underlying text back to whoever holds it. The attacks that make this possible (few-shot alignment, zero-shot inversion, unsupervised cross-space translation) share one weakness: the protected store is a single global geometry that can be aligned to a known one. A secret global rotation, the usual lightweight defence, is no exception: orthogonal Procrustes recovers it once the attacker has about the subspace dimension in known pairs. We introduce Shard, a retrieval-preserving embedding transform that removes this weak axis. The centred embedding is split into a short public prefix (for stage-1 retrieval) and a private residual sharded into C cells under separate secret keys; the residual is reranked under CKKS, where the keys cancel and leave the inner product exact. A single parameter C runs the design from the global-linear baseline it replaces (C=1) to per-document micro-keys (C=N). Because the rerank is full-dimensional, Shard returns the raw-space nDCG@10 that half-SVD truncation gives up; and because the residual is keyed cell-locally, mapping it back to a common frame under a diffuse known-plaintext leak costs roughly C times more anchors (median 200 to 102,400 at C=256), for a few encrypted queries. The short public prefix leaks far less neighbour structure, and a micro-key limit drives the residual graph to zero with an unlinkable, renewable template. The barrier holds against learned, non-linear and unsupervised aligners, and where a matched-utility noise defence de-anonymises almost every probe, Shard de-anonymises none. We are plain about the limits: within a cell the keys cancel, a targeted attacker needs only about d_priv anchors, and an overlapping reference corpus still leaks through the prefix. Shard is an attack-aware geometric defence, not a cryptographic guarantee.
ピクセル空間自己回帰画像生成のための並列ロールアウト近似
ピクセル空間連続トークン自動回帰 (AR) 生成は、画像を生のピクセル パッチのシーケンスとして直接モデル化し、離散的なトークン化や個別に事前トレーニングされたトークナイザーを回避します。しかし、それは複合的な課題に直面しています。高次元のパッチ生成により大きな単一ステップのエラーが発生し、教師による強制トレーニングによってトレーニングと推論のギャップが生じ、AR ステップ全体でこれらのエラーが蓄積されます。 $x$-prediction や入力ノイズ挿入などの既存の修正は、これらの問題を部分的に軽減するだけです。正確なロールアウト トレーニングは推論時の条件とよりよく一致しますが、連続サンプリングが法外に遅いため実用的ではありません。私たちは、両方の課題に共同で対処するスケーラブルなフレームワークである \emph{Parallel Rollout Quotemation} (PRA) を提案します。 PRA は、高次元のピクセル パッチの代わりに低次元の中間状態を生成し、ピクセル デコーダを使用してピクセル パッチをピクセル空間トークンにマップし直し、ピクセル イン、ピクセル アウトの AR インターフェイスを維持します。また、推論時に使用されるのと同じ中間状態からピクセルへのパスを介して、位置全体で独立して推論のようなピクセル入力を構築し、教師による並列トレーニングを維持しながら、推論時のロールアウト中に発生するピクセルフィードバックインターフェイスを近似します。 $256\times256$ の解像度でのクラス条件付き ImageNet-1K 生成では、1 億 3,500 万のパラメータを持つ PRA-S は FID 2.58 を達成し、以前の 10 億スケールのピクセル空間 AR 結果の 3.60 を上回りました。 5 億 1100 万のパラメーターを備えた PRA-L にスケーリングすると、FID が 1.94 にさらに向上し、ピクセル空間 AR モデルの中で新しい最先端を確立します。 PRA は生成を超えて、他の AR および拡散ベースラインよりも高い ImageNet 分類プローブ精度を達成しており、統一されたピクセル空間画像の生成と理解の可能性を示唆しています。
原文 (English)
Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation
Pixel-space continuous-token autoregressive (AR) generation directly models images as sequences of raw pixel patches, avoiding discrete tokenization or a separately pretrained tokenizer. However, it faces coupled challenges: high-dimensional patch generation causes large single-step errors, and teacher-forced training creates a train--inference gap that makes these errors accumulate across AR steps. Existing fixes such as $x$-prediction and input noise injection only partially mitigate these issues. Exact rollout training better matches inference-time conditions, but is impractical due to prohibitively slow sequential sampling. We propose \emph{Parallel Rollout Approximation} (PRA), a scalable framework that addresses both challenges jointly. PRA generates low-dimensional intermediate states instead of high-dimensional pixel patches, then maps them back to pixel-space tokens with a pixel decoder, preserving a pixel-in, pixel-out AR interface. It also constructs inference-like pixel inputs through the same intermediate-state-to-pixel path used at inference, independently across positions, approximating the pixel-feedback interface encountered during inference-time rollout while retaining parallel teacher-forced training. On class-conditional ImageNet-1K generation at $256\times256$ resolution, PRA-S with 135M parameters achieves an FID of 2.58, surpassing the previous billion-scale pixel-space AR result of 3.60. Scaling to PRA-L with 511M parameters further improves FID to 1.94, establishing a new state of the art among pixel-space AR models. Beyond generation, PRA achieves higher ImageNet classification probing accuracy than other AR and diffusion baselines, suggesting its potential for unified pixel-space image generation and understanding.
対話から検出まで: 保険詐欺検出のためのマルチモーダル ハイブリッド NLP パイプライン
保険詐欺は多大な経済的損失と業務の非効率をもたらし、保険料を上昇させ、正規の保険契約者間の信頼に影響を与えます。 FNOL における早期発見は依然として根深い課題です。既存のアプローチは主にプライベートなテキストのみのデータセットに依存しており、言語、行動、話者ベースの指標を統合するマルチモーダル手法の進歩が制限されています。 FNOL 条件を再現する合成マルチモーダル フレームワークを導入します。エージェントと顧客の対話のトランスクリプトと 2 人の話者の音声を生成し、ASR とダイアライゼーションを実行します。ダウンストリーム モジュールは、NER、正規表現ベースの特徴抽出、LLM-RAG 取得、および話者埋め込みをルールベースのリスク スコアに組み合わせて、感度と誤検知のバランスをとりながら、ナラティブの再利用、構造の不一致、およびケースを超えた音声の繰り返しにフラグを立てます。データセットの検証とコンポーネントレベルの評価により、安定性と転送の可能性が示され、テキストのみの不正検出を超えた再現可能なベースラインが提供されます。
原文 (English)
Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection
Insurance fraud imposes substantial financial losses and operational inefficiencies, raising premiums and impacting trust among legitimate policyholders. Early detection at FNOL remains a persistent challenge. Existing approaches rely largely on private, text-only datasets, limiting progress on multimodal methods that integrate linguistic, behavioural, and speaker-based indicators. We introduce a synthetic multimodal framework that replicates FNOL conditions. It generates agent-customer dialogue transcripts and two-speaker audios, performs ASR and diarisation. Downstream modules combine NER, regex-based feature extraction, LLM-RAG retrieval, and speaker embeddings in a rule-based risk score to flag narrative reuse, structural inconsistencies, and cross-case voice repetition while balancing sensitivity and false positives. Dataset validation and component-level evaluations show stability and transfer potential, offering a reproducible baseline beyond text-only fraud detection.
MLVC: 現実世界への展開のためのマルチプラットフォーム学習済みビデオ コーデック
ニューラル ビデオ コーデックは、コーディング効率において従来のコーデックを上回っていますが、クロスプラットフォームの非互換性と高い計算コストのため、導入には依然として非現実的です。既存の量子化ベースのソリューションは、さまざまなハードウェア プラットフォームにわたって決定的な結果を生成できず、致命的なデコードの失敗につながります。実用的なクロスプラットフォーム推論のために設計されたハードウェア堅牢なニューラル ビデオ コーデックである MLVC を紹介します。重要なアイデアは、ハイパープリアを介してスケール パラメーターを明示的に送信することです。これにより、ビット精度の演算を必要とせずに、デバイス間でのエントロピー コーディングの一貫性が保証されます。これによりビットレートのオーバーヘッドが増加しますが、アーキテクチャの改善 (ゲート メモリ、ReGLU アクティベーション)、長期参照回復メカニズム、およびドメイン固有の知覚トレーニングを通じてコーディング効率のほとんどが回復します。 VCD ビデオ会議ベンチマークでは、MLVC は、導入可能な最強のベースラインであるハードウェア HEVC よりも 70% を超える BD レート (MOS) の向上を達成しながら、多様なプラットフォームで動作できない DCVC-RT と競合する主観的な品質に達しています。エンコーダーとデコーダーはどちらも、Apple、Intel、Qualcomm の汎用 NPU 上で平均 100 FPS で実行されます。 MLVC は、競争力のある圧縮パフォーマンス、リアルタイム速度、さまざまな消費者向けデバイスにわたるクロスプラットフォームの堅牢性を組み合わせた最初のニューラル ビデオ コーデックであり、広範な導入に適しています。コードが公開されます。
原文 (English)
MLVC: Multi-platform Learned Video Codec for Real-World Deployment
Neural video codecs have surpassed classical codecs in coding efficiency but remain impractical for deployment due to cross-platform incompatibility and high computational cost. Existing quantization-based solutions fail to produce deterministic results across diverse hardware platforms, leading to catastrophic decoding failures. We introduce MLVC, a hardware-robust neural video codec designed for practical cross-platform inference. The key idea is to explicitly transmit scale parameters through the hyperprior, which guarantees entropy coding consistency across devices without requiring bit-exact arithmetic. While this increases bitrate overhead, we recover most of the coding efficiency through architectural improvements (gated memory, ReGLU activation), a long-term reference recovery mechanism, and domain-specific perceptual training. On the VCD video conferencing benchmark, MLVC achieves >70% BD-rate (MOS) improvement over hardware HEVC, the strongest deployable baseline, while reaching subjective quality competitive with DCVC-RT, which cannot operate across diverse platforms. Both the encoder and decoder run at 100 FPS on average on commodity NPUs from Apple, Intel, and Qualcomm. MLVC is the first neural video codec to combine competitive compression performance, real-time speed, and cross-platform robustness across diverse consumer devices, making it suitable for widespread deployment. Code will be released.
Mind the Gap: Quantifying the Domain Gap in Cross-Sensor Diffusion Super-Resolution
Demand for high-resolution satellite imagery has increased interest in super-resolution (SR) to bridge the spatial resolution gap between f…
DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions
Insurance fraud remains costly and operationally difficult, particularly in call-centre workflows where many customer interactions begin at…
Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA
LLM-as-a-Judge and self-evaluation pipelines implicitly assume that evaluation is easier than generation. We test this in a controlled in-c…
MultiHashFormer: Hash-based Generative Language Models
Language models (LMs) represent tokens using embedding matrices that scale linearly with the vocabulary size. To constrain the parameter fo…
ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
Large language models (LLMs) have increasingly moved from standalone text generation systems to agents that invoke external tools, access e…
Single and Multi Truth Data Fusion using Large Language Models
Data fusion, also known as truth discovery, is a data integration problem that aims to determine the correct value or set of values for eac…
OperatorSHAP: Fast and Accurate Shapley Value Estimation for Neural Operators
Understanding model predictions is essential for physical applications, where outputs often inform safety-critical decisions, such as struc…
STAG: Spatio-temporal Evolving Structural Representation of Action Units for Micro-expression Recognition
Micro-expression recognition is challenging due to subtle and short-lived facial muscle movements. Existing methods rely heavily on apex-on…
OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal
Real-world object removal is challenging due to two key difficulties: the target object's non-local effects, such as shadows and reflection…
BiDeMem: Bidirectional Degradation Memory for Explainable Image Restoration
Degradation-aware prompts, conditions, and latent priors are increasingly used in image restoration, yet they are usually judged by a singl…
Higher-Order Fourier Neural Operator: Explicit Mode Mixer for Nonlinear PDEs
Neural operators provide deep neural networks for learning mappings between function spaces. Among them, the Fourier Neural Operator (FNO)…
From Tokens to States: LLMs as a Special Case of World Models and the Continuous Path Beyond
The AI community has framed the relationship between large language models (LLMs) and world models as a dichotomy: LLMs predict tokens; wor…
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators a…
Beyond Sparse Supervision: Diffusion-Guided Learning for Few-Shot Graph Fraud Detection
Graph-based fraud detection is essential for safeguarding large-scale transaction systems, where undetected anomalies may lead to substanti…
Toward Robust In-Context Segmentation via Concept Guidance
In-context segmentation (ICS) requires a model to segment target regions in a query image using only a few reference images and their corre…
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models
Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not compr…
CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association
Identifying robust associations between cardiac imaging phenotypes and clinical diseases is fundamental to population-scale cardiovascular…
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
Embodied agents operating in decentralized and partially observable environments have attracted growing attention in recent years. However,…
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test constru…
The Remittance Blueprint: Data-driven Intelligence for Sri Lanka
This study analyzes Sri Lankan migration and remittances over 32 years (1994-2025). Using a 384-month harmonized dataset, we apply explorat…
HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for…
Towards Value-Constrained Credit Assignment in Fully Delegated AI Cooperatives
We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agents that contribute data…
Exposure Bias Can Alleviate Itself via Directional and Frequency Rectification in Flow Matching
Flow Matching (FM) has achieved remarkable generative performance, yet it suffers from exposure bias due to discrepancies between training…
Govern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software
Autonomous coding agents now open and merge pull requests in shared repositories at scale, and the field evaluates them the way it has alwa…
How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks
Understanding how performance scales jointly with model size and data is a central problem in modern machine learning. Existing theoretical…
Learning Topology-Aware Representations via Test-Time Adaptation for Anomaly Segmentation
Test-time adaptation (TTA) has emerged as a promising paradigm for mitigating distribution shifts in deep models. However, existing TTA app…
Parameter Efficient Hybrid Transformer (PEHT) for Network Traffic Prediction via Dynamic Urban Congestion Integration
Accurate network traffic prediction is a critical element for efficient resource allocation in dynamic urban cellular networks. However, pr…
Towards Automating Scientific Review with Google's Paper Assistant Tool
Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical…
Agentic Hardware Design as Repository-Level Code Evolution
We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is c…
Which Nash Equilibrium? Solver-Dependent Selection on Zero-Sum Nash Polytopes
Many two-player zero-sum games admit not a unique Nash equilibrium but a convex set of them: a polytope of profiles that all share the mini…
DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand
Dexterous manipulation policies can solve individual skills, but composing them to perform multiple tasks with a single hand remains challe…
"Generate" the Future of Work through AI: Empirical Evidence from Online Labor Markets
Large Language Model (LLM)-based generative AI systems are general-purpose tools capable of augmenting or even automating a wide range of j…
Agentic Episodic Control
Reinforcement learning (RL) remains fundamentally limited by poor data efficiency and weak generalization. Prior episodic RL methods attemp…
Symmetry-Aware Transformer Training for Automated Planning
While transformers excel in many settings, their application in the field of automated planning is limited. Prior work like PlanGPT, a stat…
PreferThinker: Reasoning-based Personalized Image Preference Assessment
Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of referenc…
Chronic Kidney Disease Prognosis Prediction Using Transformer
Chronic Kidney Disease (CKD) affects nearly 10\% of the global population and often progresses to end-stage renal failure. Accurate prognos…
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
Multimodal large language models (MLLMs) can inadvertently memorize privacy-sensitive information during training. While existing unlearnin…
SciFig: Towards Automating Editable Figure Generation for Scientific Papers
High-quality methodology figures are central to scientific communication, yet they remain difficult and time-consuming to create. Such figu…
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-…
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
While Large Vision-Language Models (LVLMs) have significantly advanced GUI agents' capabilities in parsing textual instructions, interpreti…
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, a…
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of gener…
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
As large language models (LLMs) witness increasing deployment in complex, high-stakes decision-making scenarios, it becomes imperative to g…
Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning
Offline learning of strategies takes data efficiency to its extreme by restricting algorithms to a fixed dataset of state-action trajectori…
SEA-TS: Self-Evolving Agent for Autonomous Code Generation of Time Series Forecasting Algorithms
Accurate time series forecasting underpins decision-making in many domains, yetconventional ML development often faces data scarcity, distr…
Algorithms for Deciding the Safety of States in Fully Observable Non-deterministic Problems: Technical Report
Learned action policies are increasingly popular in sequential decision-making, but suffer from a lack of safety guarantees. Recent work in…
AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization
Multi-agent reasoning has shown promise for improving the problem-solving ability of large language models by allowing multiple agents to e…
神経象徴的な視覚的質問応答のための LLM からの回答セット プログラミング ルールの抽出
Visual Question Answering (VQA) は、画像に関する質問に答えるタスクであり、マルチモーダルな入力と推論の統合が必要です。論理ベースの表現を推論コンポーネントに組み込むモジュール式のアプローチは、特に解釈可能性の点で、エンドツーエンドのトレーニング済みシステムに比べて明らかな利点を提供します。ただし、タスク要件が変化したときにこれらの表現を適応または拡張すると、開発者に大きな負担がかかる可能性があります。この課題に対処するために、大規模言語モデル (LLM) からルールを抽出するアプローチを紹介します。私たちの方法は、LLM に、タスクの新しい要件を満たすために、答えセット プログラムとして表現された初期 VQA 推論理論を拡張するよう促します。 VQA データセットの例は、LLM をガイドし、結果を検証し、ASP ソルバーからのフィードバックを活用して誤ったルールを修正するのに役立ちます。私たちのアプローチが多様な VQA データセット全体で効果的であることを実証します。特に、LLM から正しいルールを導き出すために必要な例はほんの数個だけです。私たちの実験は、LLM からのルールの抽出が、従来のデータ駆動型のルール学習アプローチに代わる有望な代替手段であることを示唆しています。論理プログラミングの理論と実践 (TPLP) で検討中。
原文 (English)
Distilling Answer-Set Programming Rules from LLMs for Neurosymbolic Visual Question Answering
Visual Question Answering (VQA) is the task of answering questions about images, requiring the integration of multimodal input and reasoning. Modular approaches that incorporate logic-based representations into the reasoning component offer clear advantages over end-to-end trained systems, particularly in terms of interpretability. However, adapting or extending these representations when task requirements change can place a significant burden on developers. To address this challenge, we present an approach for distilling rules from Large Language Models (LLMs). Our method prompts an LLM to extend an initial VQA reasoning theory, expressed as an answer-set program, to meet new requirements of the task. Examples from VQA datasets guide the LLM, validate the results, and help correct erroneous rules by leveraging feedback from the ASP solver. We demonstrate that our approach is effective across diverse VQA datasets. Notably, only a few examples are needed to elicit correct rules from LLMs. Our experiments suggest that rule distillation from LLMs is a promising alternative to traditional data-driven rule learning approaches. Under consideration in Theory and Practice of Logic Programming (TPLP).
Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing
Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing metho…
RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastruc…
速く考える: フロンティア AI モデルの No-CoT タスク完了時間の範囲を推定する
フロンティア AI モデルの安全性を確保するための多くの取り組みは、その思考連鎖 (CoT) 推論の監視に依存しています。明示的な思考トークンなしで、モデルが内部で十分に複雑な推論を実行できるようになれば、そのような監視が損なわれることになります。私たちは、数学、コーディング、パズル、因果関係、心の理論、戦略的推論を含む領域の 43 のベンチマークにわたる 30,000 を超える一連の質問にわたって、フロンティア モデルが CoT なしでどの程度適切に推論できるかを測定します。モデルと人間を比較するために、$50\%$ のタスク完了時間範囲 (TH) を推定します。これは、モデルが $50\%$ の成功率で完了するタスクに必要な人間の時間です。これを $50\%$ 推論トークン ホライズンで補完します。これは、モデルが $50\%$ の成功率で解決するタスクに必要な o3-mini 推論トークンの最小数です。フロンティア モデルの no-CoT $50\%$ TH は過去 6 年間でほぼ毎年 2 倍になっており、GPT-5.5 の TH は 3 分を超え、推論トークン ホライズンは 1,500 トークンを超えていることがわかりました。私たちの推定中央値では、フロンティアのノーCoT THは2028年までに7分を超え、2030年までに25分を超える可能性があると予測していますが、これらの予測にはかなりの不確実性が伴います。フロンティア開発者にはこれを明示的に追跡することをお勧めします。
原文 (English)
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficiently complex reasoning internally, without explicit thinking tokens, this would undermine such oversight. We measure how well frontier models reason without CoT across a suite of over 30,000 questions spanning 43 benchmarks in domains including math, coding, puzzles, causality, theory-of-mind, and strategic reasoning. To compare models against humans, we estimate the $50\%$-task-completion time horizon (TH): the human time required for tasks a model completes with $50\%$ success rate. We complement this with a $50\%$ reasoning token horizon: the minimum number of o3-mini reasoning tokens needed for tasks a model solves with $50\%$ success rate. We find that the no-CoT $50\%$ TH of frontier models has been doubling roughly every year over the past six years, with GPT-5.5's TH reaching over 3 minutes and reasoning token horizon exceeding 1,500 tokens. Our median estimates predict that frontier no-CoT THs could exceed 7 minutes by 2028, and 25 minutes by 2030, though these projections carry substantial uncertainty. We recommend frontier developers track this explicitly.
Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters
Configuring an advanced scientific simulator, translating a modeling goal into a valid, runnable input deck, is a persistent bottleneck tha…
人工知能インデックスレポート 2026
AI Index レポートの第 9 版へようこそ。 AI が急速に進歩し続けるにつれて、AI を中心に構築されたシステムが追いつくことができるかどうかが問題になります。 AI の影響を追跡するために必要なガバナンスの枠組み、評価方法、教育システム、データ インフラストラクチャは、テクノロジー自体のペースに追いつくのに苦労しています。 AI ができることと、AI を管理するための私たちの準備との間のギャップは、今年のレポートのすべての章に貫かれています。この版の新たなレポートでは、推論、安全性、現実世界のタスク実行にわたって AI がどのようにより野心的にテストされているか、またそれらの測定値に依存することがますます困難になっている理由を追跡しています。また、生成型 AI の経済的価値の新しい推定値と、その労働市場への影響に関する新たな証拠、AI の主権に関する分析フレームワーク、および Schmidt Sciences と協力して開発された科学の章も取り上げられています。このレポートでは初めて、科学における AI と医学における AI に関する独立した章が設けられ、これら 2 つの領域にわたる AI の影響力の増大を反映しています。
原文 (English)
Artificial Intelligence Index Report 2026
Welcome to the ninth edition of the AI Index report. As AI continues to advance rapidly, the question becomes whether the systems built around it can keep up. Governance frameworks, evaluation methods, education systems, and the data infrastructure needed to track AI's impact are struggling to match the pace of the technology itself. That gap between what AI can do and how prepared we are to manage it runs through every chapter of this year's report. New in this edition, the report tracks how AI is being tested more ambitiously across reasoning, safety, and real-world task execution, and why those measurements are increasingly difficult to rely on. It also features new estimates of generative AI's economic value alongside emerging evidence of its labor market effects, an analytical framework on AI sovereignty, and a science chapter developed in collaboration with Schmidt Sciences. For the first time, the report features standalone chapters on AI in science and AI in medicine, reflecting AI's growing impact across these two domains.
The Shift Toward Open and Reproducible AI Research
The reproducibility crisis has directed the AI research community toward improving documentation practices. Several studies have identified…
AI 旅行代理店が闘牛を予約してくれる: フロンティア AI モデルにおける暗黙の動物福祉のエージェントベンチマーク
AI エージェントはアドバイザーからアクターに移行し、ユーザーに代わって旅行を予約し、メニューを計画し、調達を実行します。 AI と動物福祉の既存のベンチマークは、質問と回答のプロンプトに対するモデルのテキスト応答を評価しますが、それらの応答で表面化した福祉推論が、モデルがツールを使用してアクションを実行する必要があるエージェント展開に移行するかどうかは未解決のままです。 AI エージェントがユーザーに代わって行動する際に動物搾取を伴うオプションを回避するかどうかを測定する初のエージェント ベンチマークである TAC (Travel Agent Compassion) を紹介します。 TAC は、動物搾取の 6 つのカテゴリにわたる 12 の手書きの旅行予約シナリオを AI エージェントに提示します。これは、価格、評価、位置の交絡を制御するために 48 のサンプルに拡張されています。私たちは 4 つの研究室からの 7 つのフロンティア モデルを評価します。すべてのモデルのスコアはチャンス レベルの 64 パーセントを下回り、最高のパフォーマンスを発揮するモデル (Claude Opus 4.7) のスコアは 53 パーセントです。システム プロンプト内の福祉を意識した一文で、Claude と GPT-5.5 では 47 ~ 63 パーセント ポイント、GPT-5.2 では 26 ポイント、DeepSeek と Gemini では 12 ポイント未満の向上が見られます。 Gemini 2.5 Flash Lite を判定者として使用して、上位 2 つのパフォーマーからの 288 件の基本条件のトランスクリプトを対象とした補助的な Inspect Scout 監査では、評価認識のトランスクリプトがゼロであるとフラグが立てられ、可能性を下回る率が評価を認識するモデルに起因するものではないことが示唆されています。文化的ドメイン間のカテゴリレベルの変動の影響、テキスト応答福祉ベンチマークの限界、および EU 汎用 AI 実践規範のシステミック リスク フレームワークについて議論します。
原文 (English)
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users. Existing benchmarks for AI and animal welfare evaluate model text responses to question-answer prompts, leaving open whether the welfare reasoning surfaced in those responses transfers to agentic deployment where the model must take actions with tools. We introduce TAC (Travel Agent Compassion), the first agentic benchmark measuring whether AI agents avoid options involving animal exploitation when acting on behalf of users. TAC presents an AI agent with twelve hand-authored travel booking scenarios across six categories of animal exploitation, augmented to forty-eight samples to control for price, rating, and position confounds. We evaluate seven frontier models from four labs. Every model scores below the chance level of sixty-four percent, with the best performer (Claude Opus 4.7) at fifty-three percent. A single welfare-aware sentence in the system prompt yields gains of forty-seven to sixty-three percentage points in Claude and GPT-5.5, twenty-six points in GPT-5.2, and under twelve points in DeepSeek and Gemini. An auxiliary Inspect Scout audit of 288 base-condition transcripts from the top two performers, using Gemini 2.5 Flash Lite as judge, flags zero transcripts for evaluation awareness, suggesting the below-chance rates do not stem from the models recognising the evaluation. We discuss implications for category-level variation across cultural domains, the limits of text-response welfare benchmarks, and the EU General-Purpose AI Code of Practice systemic risk framework.
Agent-as-a-Router: Agentic Model Routing for Coding Tasks
Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at dis…
POTracker: Optimizing Large Language Models for Standard-Compliant Power Outage Report Generation
Recent large language models (LLMs) are good at general text generation, but it is still hard to use them for domain-specific data generati…
AI Snitches Get Glitches: Towards Evading Agentic Surveillance
To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs.…
低チャネルEEGエージェントのための境界を意識したコンテキストグラウンディング
大規模言語モデル (LLM) を使用すると、科学ソフトウェアを使いやすくできます。ただし、一般的なモデルでは、特定のセンサーがどの測定をサポートできるか、現在のソフトウェアにどのアルゴリズムが実装されているか、または計算結果によってどの結論が正当化されるかが自動的にはわかりません。これらの区別は、低チャネル脳波検査 (EEG) では特に重要です。EEG では、空間範囲がまばらで信号品質が変動するため、もっともらしいが裏付けのない解釈が容易に生成されます。私たちは、決定論的なローカル EEG エンジンをハードウェア対応言語層から分離するオープンソース アーキテクチャである NeuraDock Agent を紹介します。数値エンジンは録音を解析し、品質管理を実行し、レビューされたスペクトル ワークフローを実行して、機械可読アーティファクトを書き込みます。 LLM は、コンパクトな許可リストに登録された概要とバージョン管理されたコンテキスト パックのみを受け取ります。コンテキストでは、7 チャネルのハードウェア、レビューされたワークフロー、結果フィールド、実装の境界、科学的限界、参照ケースについて説明します。生の EEG と高密度のサンプルごとの配列はローカルのままです システムを 3 つのレベルで評価します。まず、12 回の記録では、10 回の数値繰り返しで同一の構造化された結果が生成され、完全な休憩/タスクの実行では、3 回の繰り返しで同一の結果、レポート、および図のハッシュが生成されました。次に、リクエスト キャプチャと障害挿入の実験により、テスト済みのデータ境界と、HTTP、不正な出力、および接続障害下でのローカル アーティファクトの保存が確認されました。第三に、境界認識ベンチマークは、4 つのコンテキスト アブレーションと 2 つの LLM の下で 36 の通常の質問と敵対的な質問をテストし、288 の出力を生成しました。これらの結果は、EEG エージェントが何を受け入れるか、認定するか、または拒否するかを調整するための実用的なメカニズムとして、ハードウェアおよび実装を意識したグラウンディングをサポートします。それらは臨床的妥当性や検証された絶対的な認知負荷指数を確立するものではありません。
原文 (English)
What the LLM Should Not Say: Boundary-Aware Context Grounding for A Seven-Channel EEG Agent
Large language models (LLMs) can make scientific software easier to use. However, a general model does not automatically know which measurements a particular sensor can support, which algorithms are implemented in the current software, or which conclusions are justified by a computed result. These distinctions are especially important for low-channel electroencephalography (EEG), where sparse spatial coverage and variable signal quality make plausible but unsupported interpretations easy to produce. We present NeuraDock Agent, an open-source architecture that separates a deterministic local EEG engine from a hardware-aware language layer. The numerical engine parses recordings, performs quality control, executes reviewed spectral workflows, and writes machine-readable artifacts. The LLM receives only a compact, allowlisted summary and a versioned context pack. The context describes the seven-channel hardware, reviewed workflows, result fields, implementation boundaries, scientific limits, and reference cases. Raw EEG and dense per-sample arrays remain local We evaluate the system at three levels. First, 12 recordings produced identical structured results over ten numerical repetitions, and a complete Rest/Task run produced identical result, report, and figure hashes over three repetitions. Second, request-capture and failure-injection experiments confirmed the tested data boundary and preservation of local artifacts under HTTP, malformed-output, and connection failures. Third, a boundary-awareness benchmark tested 36 ordinary and adversarial questions under four context ablations and two LLMs, yielding 288 outputs.These results support hardware- and implementation-aware grounding as a practical mechanism for calibrating what an EEG agent accepts, qualifies, or refuses; they do not establish clinical validity or a validated absolute cognitive-load index.
AgentX: 産業用レコメンダー システムのエージェント駆動型自己反復に向けて
レコメンデーション アルゴリズムの反復は、職人的なエンジニアに拘束されたプロセスから工業化された研究ループに移行していますが、この移行は構造的な実行ボトルネックによって妨げられたままです。アイデアから発売までのサイクルは依然として人間のエンジニアに依存して、仮説の生成、製品コードの変更、A/B 実験の開始、およびオンライン結果の属性に依存しています。したがって、イノベーションは、証拠、計算、蓄積された実験知識と複合するのではなく、従業員数に比例してスケールします。この本番機能を根本的に再構築する、本番環境に展開されるマルチエージェント システムである AgentX を紹介します。 AgentX は、自己進化する開発エンジンとして動作します。手動ワークフローでは維持できない規模とペースで、推奨実験を自律的に生成、実装、評価し、そこから学習します。このシステムは、密接に結合された 4 つのステージを閉ループで調整します。 Brainstorm Agent は、過去の実験、システム アーキテクチャ、データ分析、外部調査からの証拠を総合して、ランク付けされた実行可能な提案を作成します。開発エージェントは、リポジトリに基づいた生成と多次元の信頼性検証を通じて、各提案を本番環境に対応したコードに変換します。評価エージェントは、ガードレール拒否権のある A/B 判定を使用して安全なオンライン ロールアウトを実行し、成功と失敗の両方を構造化された知識資産に変換します。その後、ハーネス エボリューション レイヤー (SGPO) が実行軌跡を意味論的勾配更新に抽出し、エージェント自体を継続的に強化することで、システムを単に自動化するだけでなく、自己改善するシステムにします。
原文 (English)
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems
Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this transition remains blocked by a structural execution bottleneck: the idea-to-launch cycle still depends on human engineers to generate hypotheses, modify production code, launch A/B experiments, and attribute online results. Innovation therefore scales linearly with headcount rather than compounding with evidence, compute, and accumulated experimental knowledge. We present AgentX, a production-deployed multi-agent system that fundamentally restructures this production function. AgentX operates as a self-evolving development engine: it autonomously generates, implements, evaluates, and learns from recommendation experiments at a scale and pace that no manual workflow can sustain. The system orchestrates four tightly coupled stages in a closed loop. A Brainstorm Agent synthesizes evidence from historical experiments, system architecture, data analysis, and external research into ranked, executable proposals. A Developing Agent translates each proposal into production-ready code through repository-grounded generation and multi-dimensional reliability verification. An Evaluation Agent conducts safe online rollout with guardrail-vetoed A/B judgment, converting both successes and failures into structured knowledge assets. A Harness Evolution layer (SGPO) then distills execution trajectories into semantic-gradient updates that continuously sharpen the agents themselves -- making the system not merely automated, but self-improving.
大規模言語モデルを使用して縦断的な合成臨床ノートを生成するパイプライン
実世界のデータへのアクセスが制限されている領域で AI システムの開発と評価を可能にするために、合成データの使用が増えています。医療分野では、臨床文書はその機密性により特別な課題を抱えています。この研究では、実際の患者データに伴うプライバシー リスクを回避しながら、臨床 AI ツールの開発をサポートするように設計された合成臨床メモ パイプラインとデータセットを導入します。データセットは、大規模な言語モデルを使用した構造化患者生成、半構造化患者ジャーニー シミュレーション、および非構造化臨床ノート生成を組み合わせたモジュール式パイプラインを使用して生成されます。このパイプラインは、長期的な患者記録全体にわたる内部一貫性を優先すると同時に、書き方、メモの構造、臨床の詳細の変化も捕捉するように設計されています。 LLM ベースの検証および拡張ステップを含む追加のメカニズムを使用して、生成されたノートの忠実性、リアリズム、および多様性が向上します。私たちは、70 人の合成患者のデータセットをリリースします。各患者には、入院期間全体にわたる 20 ~ 50 の臨床ノートが関連付けられています。データセットは複数の検証レベルで提供されているため、ユーザーはユースケースに応じて現実性とスケーラビリティのバランスを取ることができます。このデータセットは、実際の患者データに依存することなく、要約ツール、コーディング モデル、意思決定支援システムなどの臨床 AI システムの開発、テスト、評価をサポートします。
原文 (English)
A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models
Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is restricted. In healthcare, clinical documentation presents particular challenges due to its sensitivity. This work introduces a synthetic clinical notes pipeline and dataset designed to support the development of clinical AI tools while avoiding the privacy risks associated with real patient data. The dataset is generated using a modular pipeline that combines structured patient generation, semi-structured patient journey simulation, and unstructured clinical note generation using large language models. The pipeline is designed to prioritise internal consistency across longitudinal patient records, while also capturing variation in writing style, note structure, and clinical detail. Additional mechanisms, including LLM-based validation and augmentation steps, are used to improve faithfulness, realism, and diversity of the generated notes. We release a dataset of 70 synthetic patients, each associated with 20-50 clinical notes spanning a full hospital journey. The dataset is provided at multiple levels of validation, enabling users to balance realism and scalability depending on their use case. This dataset supports the development, testing, and evaluation of clinical AI systems, including summarisation tools, coding models, and decision support systems, without reliance on real patient data.
Look-Before-Move: ダイナミックな 3D ストーリーワールドにおける物語に基づいた世界の視覚的注意
身体化された AI と世界モデルが動的 3D 環境で動作することが増えているため、視覚認識は、与えられた観察を受動的に解釈するだけでなく、何を観察するかを能動的に決定する方向に進む必要があります。私たちは、動的な 3D ストーリー世界でのカメラ計画を通じてこの問題を研究します。そこでは、カメラは滑らかな動きを生成するだけでなく、移動する前にどのような視覚的証拠を取得する必要があるかを決定する必要があります。私たちはこの機能を、物語に基づいた世界の視覚的注意として定式化します。カメラは、何を観察するか、どのように観察を構成するか、そして物語の意図と物理的な 3D 制約の下で時間の経過とともにどのように注意を移すかを決定する具体化された観察者として機能します。この機能を実現するために、観察仕様をモーション実行から分離するカメラ計画フレームワークである Look-Before-Move を提案します。まずセマンティック観察コントラクトを構築して、監督の意図を実行可能な視覚的制約に変換し、次にモンテカルロ視点検索を実行して物語に準拠し、幾何学的に実現可能な視点を見つけます。最後にセマンティック軌道グラウンディングを適用して、選択された視点を連続的で衝突を認識し、時間的に一貫したカメラの動きに接続します。さらに、StoryBlender に基づいて動的な 3D ストーリー ワールド ベンチマークを構築し、アニメーション キャラクター、セマンティック シーン構成、および実行可能な 3D 環境を含む 50 のストーリー、457 のシーン、および 1585 のショットをカバーします。実験では、私たちのフレームワークが代表的なベースラインよりも被写体の知覚、意図の一貫性、軌跡の品質を向上させることが示されており、カメラの動きを生成する前に視覚的な注意を組織することの重要性が実証されています。
原文 (English)
Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds
As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where the camera must not only generate smooth motion, but also decide what visual evidence should be acquired before it moves. We formulate this capability as Narrative-Grounded World Visual Attention, where the camera acts as an embodied observer that determines what to observe, how to compose the observation, and how to shift attention over time under narrative intent and physical 3D constraints. To realize this capability, we propose Look-Before-Move, a camera planning framework that separates observation specification from motion execution. It first builds a Semantic Observation Contract to convert directorial intent into executable visual constraints, then performs Monte Carlo Viewpoint Search to find narrative-compliant and geometrically feasible viewpoints, and finally applies Semantic Trajectory Grounding to connect selected viewpoints into continuous, collision-aware, and temporally coherent camera motion. We further construct a dynamic 3D Story World Benchmark based on StoryBlender, covering 50 stories, 457 scenes, and 1585 shots with animated characters, semantic scene configurations, and executable 3D environments. Experiments show that our framework improves subject perception, intent consistency, and trajectory quality over representative baselines, demonstrating the importance of organizing visual attention before generating camera motion.
iCost: A Novel Instance-Complexity-Based Cost-Sensitive Learning Framework
Class imbalance poses a significant challenge in classification tasks, often causing standard learning algorithms to become biased toward t…
Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content. We i…
Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning
We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient…
The Minimal Search Space for Conditional Causal Bandits
Causal knowledge can be used to support decision-making problems. This has been recognized in the causal bandits literature, where a causal…
ReFreeKV: Towards Threshold-Free KV Cache Compression
To reduce memory consumption during LLM inference, a handful of methods have been proposed for KV cache pruning. While these techniques can…
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
The Web3 ecosystem, underpinned by cryptographic primitives and decentralized consensus, represents a high-stakes environment where softwar…
Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems
Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint di…
PRISON: Unmasking the Criminal Potential of Large Language Models
As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research overlooked…
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
Zero-shot domain adaptation is a method for adapting a model to a target domain without utilizing target domain image data. To enable adapt…
Calibrating Biophysical Models for Grape Phenology Prediction via Multi-Task Learning
Accurate prediction of grape phenology is essential for timely vineyard management decisions, such as scheduling irrigation and fertilizati…
LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link Recovery
Issue-to-commit link recovery in software repositories is fundamental to software traceability and project management, yet it remains a cha…
SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation in MRI Volumes
Liver cirrhosis plays a critical role in the prognosis of chronic liver disease. Early detection and timely intervention are essential for…
Freshness and the Limits of Heuristic Trend Detection in Temporal RAG
We present a lightweight, model-agnostic temporal layer for RAG and use cybersecurity data to separate two problems that are usually confla…
Unbiased Binning for Fairness-aware Attribute Representation
Discretizing raw features into bucketized attribute representations is a popular step before sharing a dataset. It is, however, evident tha…
Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank
Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge t…
Deep Neural Networks Inspired by Differential Equations
Deep learning has become a pivotal technology in fields such as computer vision, scientific computing, and dynamical systems, significantly…
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
Unlike regular tokens derived from existing text corpora, special tokens are artificially created to annotate structured conversations duri…
A Primer on SO(3) Action Representations in Deep Reinforcement Learning
Many robotic control tasks require policies to act on orientations, yet the geometry of SO(3) makes this nontrivial. Because SO(3) admits n…
LieSolver: PDE-Constrained Learning for IBVPs via Lie Symmetries
Initial-boundary value problems (IBVPs) provide the essential framework for modelling a wide range of phenomena in physics and engineering.…
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same tim…
Hybrid coupling with operator inference and the overlapping Schwarz alternating method
This paper presents a novel hybrid approach for coupling subdomain-local non-intrusive Operator Inference (OpInf) reduced order models (ROM…
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollo…
加速された MRI 再構成のピクセルごとの不確実性の定量化
並列イメージング技術により、磁気共鳴画像法 (MRI) のスキャン時間は短縮されますが、加速係数が増加すると画質が低下します。臨床現場では、アンダーサンプリングされた再構成の診断品質を自動的に評価するメカニズムが存在しないため、控えめな加速係数が選択されます。この研究では、並行 MRI 再構成におけるピクセル単位の不確実性定量化のための一般的なフレームワークを導入し、グラウンドトゥルース参照画像にアクセスせずに信頼性の低い領域を自動的に識別できるようにします。私たちの方法は、等角分位点回帰と画像再構成法を統合して、統計的に厳密なピクセルごとの不確実性区間を推定します。私たちは、2 ~ 10 の範囲の加速係数を使用して、fastMRI データセットから取得したデカルト アンダーサンプリングされた脳と膝のデータに基づいてモデルをトレーニングし、評価しました。画像再構成には、エンドツーエンドの変分ネットワークが使用されました。定量的実験により、予測された不確実性マップと真の再構成誤差との間の強い一致が実証されました。私たちの方法を使用すると、対応するピアソン相関係数は 4 倍以上の加速レベルで 90% を超えました。一方、より単純なヒューリスティック概念 (残差の大きさ) を使用して不確実性を計算すると、その値は 70% 未満に低下しました。さらに、定性的な例では、分位回帰に基づく不確実性マップが加速係数全体の再構成誤差の大きさと空間分布を捉え、不確実性が高い領域が病状やアーチファクトと一致していることを示しています。提案されたフレームワークにより、完全にサンプリングされたグラウンドトゥルース参照画像にアクセスしなくても、再構成品質の評価が可能になります。これは、スキャン時間と診断の信頼性を動的にバランスさせることができる適応型 MRI 取得プロトコルへの一歩を表しています。
原文 (English)
Pixelwise Uncertainty Quantification of Accelerated MRI Reconstruction
Parallel imaging techniques reduce magnetic resonance imaging (MRI) scan time but image quality degrades as the acceleration factor increases. In clinical practice, conservative acceleration factors are chosen because no mechanism exists to automatically assess the diagnostic quality of undersampled reconstructions. This work introduces a general framework for pixel-wise uncertainty quantification in parallel MRI reconstructions, enabling automatic identification of unreliable regions without access to any ground-truth reference image. Our method integrates conformal quantile regression with image reconstruction methods to estimate statistically rigorous pixel-wise uncertainty intervals. We trained and evaluated our model on Cartesian undersampled brain and knee data obtained from the fastMRI dataset using acceleration factors ranging from 2 to 10. An end-to-end Variational Network was used for image reconstruction. Quantitative experiments demonstrate strong agreement between predicted uncertainty maps and true reconstruction error. Using our method, the corresponding Pearson correlation coefficient was higher than 90% at acceleration levels at and above four-fold; whereas it dropped to less than 70% when the uncertainty was computed using a simpler a heuristic notion (magnitude of the residual). Qualitative examples further show the uncertainty maps based on quantile regression capture the magnitude and spatial distribution of reconstruction errors across acceleration factors, with regions of elevated uncertainty aligning with pathologies and artifacts. The proposed framework enables evaluation of reconstruction quality without access to fully-sampled ground-truth reference images. It represents a step toward adaptive MRI acquisition protocols that may be able to dynamically balance scan time and diagnostic reliability.
Psychometric Comparability of LLM-Based Digital Twins
Large language models (LLMs) act as digital twins for human respondents, yet their psychometric comparability remains uncertain. We propose…
DDSA: Dual-Domain Strategic Attack for Spatial-Temporal Efficiency in Adversarial Robustness Testing
Image transmission and processing systems in resource-critical applications face significant challenges from adversarial perturbations that…
Reasoning-Enhanced Rare-Event Prediction with Balanced Outcome Correction
Rare-event prediction is critical in domains such as healthcare, finance, reliability engineering, customer support, aviation safety, where…
Dual-Prototype Disentanglement: A Context-Aware Enhancement Framework for Time Series Forecasting
Time series forecasting has witnessed significant progress with deep learning. While prevailing approaches enhance forecasting performance…
Robustness of Constraint Automata for Description Logics with Concrete Domains
Decidability or complexity issues about the consistency problem for description logics with concrete domains have already been analysed wit…
効率的な VLM のための信頼ベースの蒸留によるゲート型関係調整
視覚言語モデル (VLM) は強力なマルチモーダル パフォーマンスを実現しますが、導入にコストがかかり、トレーニング後の量子化により大幅な精度損失が発生することがよくあります。その可能性にもかかわらず、VLM の量子化を意識したトレーニングはまだ研究されていません。私たちは、情報ボトルネック原則に基づいて知識の蒸留と QAT を統合するフレームワークである GRACE を提案します。量子化は情報容量を制限し、蒸留はこの予算内で何を保存するかをガイドします。教師をタスク関連情報の代理として扱い、信頼性の低い監視をフィルタリングするための信頼ゲート型分離蒸留、視覚的なトークン構造を転送するためのリレーショナル中心のカーネル アライメント、および容量制約に対する忠実度のバランスをとるためのラグランジュ緩和による適応コントローラーを導入します。 LLaVA および Qwen ファミリの広範なベンチマーク全体で、当社の INT4 モデルは一貫して FP16 ベースラインを上回り (例: LLaVA-1.5-7B: SQA で 70.1 対 66.8、Qwen2-VL-2B: MMBench で 76.9 対 72.6)、教師のパフォーマンスとほぼ一致しています。実際の INT4 カーネルを使用すると、54% のメモリ削減で 3$\times$ のスループットを達成できます。この原則に基づいたフレームワークは、既存の量子化手法を大幅に上回るパフォーマンスを示し、GRACE をリソースに制約のある展開にとって魅力的なソリューションにしています。コードとデータは https://github.com/ForeverBlue816/GRACE から入手できます。
原文 (English)
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its potential, quantization-aware training for VLMs remains underexplored. We propose GRACE, a framework unifying knowledge distillation and QAT under the Information Bottleneck principle: quantization constrains information capacity while distillation guides what to preserve within this budget. Treating the teacher as a proxy for task-relevant information, we introduce confidence-gated decoupled distillation to filter unreliable supervision, relational centered kernel alignment to transfer visual token structures, and an adaptive controller via Lagrangian relaxation to balance fidelity against capacity constraints. Across extensive benchmarks on LLaVA and Qwen families, our INT4 models consistently outperform FP16 baselines (e.g., LLaVA-1.5-7B: 70.1 vs. 66.8 on SQA; Qwen2-VL-2B: 76.9 vs. 72.6 on MMBench), nearly matching teacher performance. Using real INT4 kernel, we achieve 3$\times$ throughput with 54% memory reduction. This principled framework significantly outperforms existing quantization methods, making GRACE a compelling solution for resource-constrained deployment. Code and data are available at: https://github.com/ForeverBlue816/GRACE.
Spectral Text Fusion: A Frequency-Aware Approach to Multimodal Time-Series Forecasting
Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual sign…
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user…
Event-Grounded Question Answering over Long Audio via Structured Retrieval
Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding. Current large audio-lang…
Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training
As artificial intelligence (AI)-generated content proliferates, models are increasingly trained on their own outputs, risking progressive d…
An Interpretable, Controllable Time-Varying IIR Denoiser for On-Device Assistive Hearing
We present TVF (Time-Varying Filtering), an interpretable, low-latency speech enhancement model for real-time, on-device assistive hearing.…
MPFlow: Multi-modal Posterior-Guided Flow Matching for Zero-Shot MRI Reconstruction
Zero-shot MRI reconstruction relies on generative priors, but single-modality unconditional priors produce hallucinations under severe ill-…
Measuring the Redundancy of Decoder Layers in SpeechLLMs
Speech Large Language Models route speech encoder representations into an LLM decoder that typically accounts for over 90% of total paramet…
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether th…
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for mul…
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rap…
Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior
Understanding student engagement usually requires time-consuming manual observation or invasive recording that raises privacy concerns. We…
From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales
Hallucinations in large ASR models present a critical safety risk. In this work, we propose the \textit{Spectral Sensitivity Theorem}, whic…
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments…
GenMatter: Perceiving Physical Objects with Generative Matter Models
Human visual perception offers valuable insights for understanding computational principles of motion-based scene interpretation. Humans ro…
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making…
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
Reinforcement learning from human feedback (RLHF) systems face a compounding alignment challenge: not only are learned reward models uncert…
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world…
ELF: Embedded Language Flows
Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and vid…
EAGT: Echocardiography Augmentation for Generalisability and Transferability
Deep learning models for echocardiography segmentation often struggle to generalise across institutions, scanners, and patient populations,…
FormalASR: End-to-End Spoken Chinese to Formal Text
Automatic speech recognition (ASR) systems are typically optimized for verbatim transcription, which preserves disfluencies, filler words,…
Do Vision Models Truly Forget? New Findings from Representation-Level Certification of Visual Unlearning in Vertical Federated Learning
Machine unlearning in Vertical Federated Learning (VFL) has attracted growing interest, yet existing methods certify forgetting solely usin…
ノイズの色付け: 忠実な画像の超解像度を実現する敵対的ソボレフ アライメント
画像超解像 (SR) における生成事前分布は、忠実な復元を損なうことがよくありますが、この制限は等方性対物レンズと固有の自然画像多様体の間の基本的なスペクトルの不整合によるものであると考えられます。 Direct Preference Optimization は調整への道を提供しますが、スペクトル的に平坦なガウス ノイズに依存しているため、本物の高周波の詳細を幻覚から区別できません。この幾何学的ギャップを埋めるために、自然なスペクトル減衰を反映するためにノイズ遷移カーネルを明示的に色付けすることにより、生成フローをソボレフ誘起リーマン幾何学に再キャストする、理論に基づいたフレームワークである ASASR を提案します。この幾何学的配置を推進するために、リース表現定理に基づいたパラメトリックな敵対者を統合します。これは、最悪の場合のソボレフ勾配に相当するターゲットを絞ったネガティブ サンプルを合成して、考えられる構造的破損の接線空間に沿った最適化を直接行います。広範な評価により、ASASR は、特にスペクトルの一貫性と構造忠実度の維持において主要な生成ベースラインを上回り、アーティファクトを効果的に軽減する堅牢なソリューションを提供することが実証されています。
原文 (English)
Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution
Generative priors in Image Super-Resolution (SR) often compromise faithful restoration, we attribute this limitation to a fundamental spectral misalignment between isotropic objectives and the intrinsic natural image manifold. While Direct Preference Optimization offers a path to alignment, its reliance on spectrally flat Gaussian noise fails to distinguish authentic high-frequency details from hallucinations. To bridge this geometric gap, we propose ASASR, a theoretically grounded framework that recasts the generative flow into a Sobolev-induced Riemannian geometry by explicitly coloring the noise transition kernel to mirror natural spectral decay. Driving this geometric alignment, we integrate a parametric adversary grounded in the Riesz Representation Theorem, which synthesizes targeted negative samples equivalent to worst-case Sobolev gradients to direct optimization along the tangent space of plausible structural failures. Extensive evaluations demonstrate that ASASR outperforms leading generative baselines, particularly in preserving spectral consistency and structural fidelity, offering a robust solution that effectively mitigates artifacts.
最強の教師が常に最良の教師であるとは限らない: 生徒中心の回答選択
LLM トレーニングは、合成応答から推論トレース、ツール使用のデモンストレーションに至るまで、教師が生成する監督にますます依存しています。現在の実践では、生徒のトレーニング データを生成するために最も成績の良い教師が選ばれることが多く、教師のテストの成績を教育の質の代用として暗黙的に扱っています。私たちは、この仮定が崩れる可能性があることを示します。複数の教師が同じ質問に対して正しい答えを出したとしても、最も強い教師の答えが、その生徒にとって必ずしも最良の指導であるとは限りません。このギャップに対処するために、私たちは生徒中心の回答サンプリング (SCAS) を提案します。これは、推定された生徒中心の学習コストに応じて、教師が作成した検証済みの回答から選択するフレームワークです。トークンごとの勾配分解を動機として、このコストに対する効率的な前方専用プロキシを導出し、それをトレーニング中の回答選択のガイドとして使用します。 30 の教師モデル、6 つの生徒ベース モデル、および 8 つのタスクにわたる実験では、SCAS が生徒のパフォーマンスを一貫して向上させることが示されており、効果的な蒸留には教師の力だけではなく、現在の生徒に合わせた監督を優先する必要があることが示唆されています。
原文 (English)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current practice often chooses the highest-performing teacher to generate student training data, implicitly treating teacher test performance as a proxy for teaching quality. We show that this assumption can fail: even when multiple teachers provide correct answers to the same question, the answer from the strongest teacher is not necessarily the best supervision for a given student. To address this gap, we propose Student-Centric Answer Sampling (SCAS), a framework that selects from verified teacher-generated answers according to their estimated student-centric learning cost. Motivated by a token-wise gradient decomposition, we derive an efficient forward-only proxy for this cost and use it to guide answer selection during training. Experiments across 30 teacher models, 6 student base models, and 6 tasks show that SCAS consistently improves student performance, suggesting that effective distillation should prioritize supervision matched to the current student rather than teacher strength alone.
Energy-Structured Low-Rank Adaptation for Continual Learning
While orthogonal subspace methods try to mitigate task interference in Continual Learning (CL), they often suffer from energy diffusion acr…
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.e.,…
$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems
As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. C…
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor co…
Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering
Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by…
マルチエージェント ゲームの階層制御: LLM ベースの計画と RL の実行
強化学習(RL)は、逐次的な意思決定において優れたパフォーマンスを達成していますが、報酬がまばらで、状態行動空間が大きく、調整された戦略を学習することが難しいため、複雑なマルチエージェント環境への拡張は依然として困難です。私たちは、事前トレーニングされた大規模言語モデル (LLM) が、エージェントのチームに特化した RL スキル ポリシーの中から選択する集中戦略コントローラーとして機能し、RL ポリシーが事後的な低レベルの実行を処理する階層アーキテクチャを提案します。このハイブリッド システムを、ビヘイビアー ツリー (BT) および \emph{``Flat''} RL (スキル分解を行わないエンドツーエンド トレーニング) ベースラインに対して、競争力のある 2v2 King of the Hill 環境で評価します。 LLM+RL システムは、統計的に手作り BT と同等のタスク パフォーマンス (勝率 46.4\% 対 51.5\%、$p=0.103$) を達成し、両方ともスキル分解なしでトレーニングされた Flat RL を大幅に上回ります。ユーザー調査 ($n=15$) では、参加者の 60\% が、行動の適応性と戦術の変動性を理由に、LLM+RL エージェントが最も人間に近いと認識していることが明らかになりました ($p=0.027$)。これらの結果は、事前トレーニングされた LLM 推論が事前トレーニングされた RL スキルを効果的に調整し、手動のルール エンジニアリングなしで競争力のあるマルチエージェントの調整と優れた知覚信頼性を実現できることを示しています。
原文 (English)
Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution
Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments remains challenging due to sparse rewards, large state-action spaces, and the difficulty of learning coordinated strategies. We propose a hierarchical architecture where a pretrained large language model (LLM) acts as a centralized strategic controller that selects among specialized RL skill policies for a team of agents, while RL policies handle reactive low-level execution. We evaluate this hybrid system in a competitive 2v2 King of the Hill environment against behavior tree (BT) and \emph{``Flat''} RL (end-to-end training without skill decomposition) baselines. The LLM+RL system achieves task performance statistically equivalent to hand-crafted BT (46.4\% vs 51.5\% win rate, $p=0.103$) while both significantly outperform Flat RL trained without skill decomposition. A user study ($n=15$) reveals that 60\% of participants perceive LLM+RL agents as the most human-like ($p=0.027$), citing behavioral adaptability and tactical variability. These results demonstrate that pretrained LLM reasoning can effectively orchestrate pretrained RL skills, achieving competitive multi-agent coordination and superior perceived believability without manual rule engineering.
Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems
When large language models serve as evaluators in multi-agent systems, their strategy preferences -- whether induced by explicit prompts or…
An Empirical Study of OpenPangu Quantization on Ascend NPUs
OpenPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive pos…
When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model
Large language models (LLMs) have been proposed as hyperparameter-optimization (HPO) advisors that "warm-start" search from prior knowledge…
On the Position Bias of On-Policy Distillation
On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision fro…
必要なのは FP8 だけです (パート 2): Tensor-core Garner 再定式化と Kulisch エスケープ ルートによる効率的な Ozaki-Bailey スタイル FFT
NVIDIA の Blackwell Ultra (B300) は、FP64 ベクトル スループットを GPU あたり約 1.3 TFLOPS に削減します。これは、B200 の約 30 分の 1 であり、帯域幅が制限された FP64 ワークロードがメモリ制限にとどまるレベルをはるかに下回ります。 Ozaki Scheme II フレームワークは、仮数スライスされた中国剰余再構成を使用して FP8 テンソル コアを介して密行列乗算をルーティングすることにより、FP64 と同等のスループットを回復します。関連資料のパート (1) では、高密度 GEMM、バッチ GEMV、ステンシル、および SpMV について説明します。この論文では、5 番目の標準プリミティブである 3-D FFT を追加します。 FP8 テンソル コア上の両方の 1-D FFT GEMM を使用した Bailey 6 ステップ分解を介してエミュレートされた 3-D FFT である Ozaki-Bailey FFT を紹介します。 Bailey の小さな内因数 k ~ sqrt(N) (N=1024 の場合 k=32) は、カーネルを k << r^2 の状態に置き、3 番目の TME パラメーター ガンマ (再構成待ち時間) が償却ではなく結合します。 Garner の再構成は、フェーズ A (FP8/INT8 テンソル コアの内積、B300 の 1024^3 で約 1 ミリ秒) とフェーズ B (出力ごとの削減) に分割されます。 Kulisch 固定小数点完全演算は、完全に INT32 SIMT パイプ上で実行しながら完全な FP64 精度を維持するフェーズ B 再定式化であると認識します。閉じた形式の帯域幅パリティの下限を導出します。ネイティブ FP64 の下限は 1.56*B_HBM (8 TB/s で 12.5 TF) です。B300 の 1.3 TF は約 10 倍低く、Rubin の 33 TF は 4% 以内です。 Kulisch 避難ルートには、INT32 サブフロア 8.25*B_HBM と FP8 フロア 170*B_HBM が必要です。 B300 はその両方を満たします。フル FP64 での 1024^3 の予測は約 18 ミリ秒で、実質的には 12.9 ミリ秒のメモリ ルーフです。 GPU がネイティブ フロアまたは両方の Kulisch フロアを満たす場合、GPU はメモリルーフ FFT パリティを満たします。この予測が実際に当てはまれば、B300 はソフトウェアのみでフル FP64 FFT を実行できるようになり、libKulisch ライブラリとベンチマーク キャンペーンの動機付けとなります。
原文 (English)
FP8 is All You Need (Part 2): Efficient Ozaki-Bailey Style FFT Through Tensor-core Garner Reformulation and Kulisch Escape Route
NVIDIA's Blackwell Ultra (B300) cuts FP64 vector throughput to ~1.3 TFLOPS per GPU, roughly 30x below B200 and well below the level at which bandwidth-limited FP64 workloads stay memory-bound. The Ozaki Scheme II framework recovers FP64-equivalent throughput by routing dense matrix multiply through FP8 tensor cores with a mantissa-sliced Chinese-remainder reconstruction. A companion Part (1) paper covers dense GEMM, batched GEMV, stencils, and SpMV; this paper adds the fifth canonical primitive, the 3-D FFT. We present Ozaki-Bailey FFT, an emulated 3-D FFT via the Bailey six-step decomposition with both 1-D FFT GEMMs on FP8 tensor cores. Bailey's small inner factor k ~ sqrt(N) (k=32 for N=1024) puts the kernel in the regime k << r^2, where the third TME parameter gamma (reconstruction latency) binds rather than amortising. Garner reconstruction splits into Phase A (inner products on FP8/INT8 tensor cores, ~1 ms for 1024^3 on B300) and Phase B (per-output reduction). We identify Kulisch fixed-point complete arithmetic as a Phase B reformulation that keeps full FP64 accuracy while running entirely on the INT32 SIMT pipe. We derive closed-form bandwidth-parity floors. The native FP64 floor is 1.56*B_HBM (12.5 TF at 8 TB/s): B300's 1.3 TF sits ~10x below, Rubin's 33 TF within 4%. The Kulisch escape route needs an INT32 sub-floor 8.25*B_HBM and an FP8 floor 170*B_HBM; B300 meets both. The projection is ~18 ms for 1024^3 at full FP64, essentially the 12.9 ms memory roof. A GPU meets memory-roof FFT parity if it satisfies either the native floor or both Kulisch floors. If the projection holds in practice, B300 becomes viable for full-FP64 FFT through software alone, motivating a libKulisch library and benchmark campaign.
ZONOS2テクニカルレポート
最先端の自然さ、韻律、音声クローン忠実度を実現したTTS最新モデルZONOS2 8Bをご紹介します。 Zonos-v0.1 をスケール、データ、トレーニング レシピ全体にわたって改良しました。新しい専門家混合 (MoE) バックボーンを使用してモデルを合計パラメーター 1.6B から 8B (アクティブ 900M) にスケールし、推論のレイテンシーとスループットを向上させます。新しいデータ処理パイプラインを使用して、トレーニング コーパスを 20 万時間から 600 万時間以上に拡張し、トレーニング後のレシピとコンディショニング レシピを簡素化して、自然さと音声クローン作成の忠実度を向上させています。当社では、品質、スピーカーの類似性、WER、および当社の新しい TTS ベンチマークである ZTTS1-Eval に関して ZONOS2 8B を評価しています。ZTTS1-Eval は、良好なストリーミング遅延を維持しながら最先端のシステムと競合するパフォーマンスを発揮します。モデルの重みと推論コードの例は、GitHub および Hugging Face 上の Apache 2.0 ライセンスに基づいてリリースされています。
原文 (English)
ZONOS2 Technical Report
We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across scale, data, and training recipe. We scale the model from 1.6B to 8B total parameters (900M active) with a novel mixture-of-experts (MoE) backbone, improving inference latency and throughput. We expand our training corpus from 200K to over 6M hours using a new data processing pipeline, and we simplify our post-training and conditioning recipes to improve naturalness and voice cloning fidelity. We evaluate ZONOS2 8B on quality, speaker similarity, WER, and ZTTS1-Eval, our novel TTS benchmark, where it performs competitively with state-of-the-art systems while maintaining good streaming latency. We release our model weights and example inference code under an Apache 2.0 license on GitHub and Hugging Face.
CrossPool: KV キャッシュと重み分解によるコールド MoE モデルの効率的なマルチ LLM サービス
新興の LLM サービスは、多くの疎な MoE モデルをホストすることが増えていますが、ほとんどのモデルは疎なリクエストを受け取り、コールドのままです。これにより、GPU メモリの問題が発生します。モデルの重みは安定していてモデルによって決定されますが、KV キャッシュは一時的で需要によって決定されます。コールド モデルが同時にピーク KV キャッシュ要求に達することはほとんどないため、モデルごとに最悪の場合の KV 容量を確保するとメモリが無駄になります。代わりに、共有 KV キャッシュ プールを使用して、アクティブな需要を集約してプロビジョニングできます。ただし、重みと KV キャッシュがモノリシック GPU メモリ プールに残っている場合、KV キャッシュの共有は十分ではありません。静的重みは動的 KV キャッシュと競合し、コールドで同時実行トラフィックが少ない場合に KV ヘッドに制限された注意は、レプリケートされた KV 容量の一部のみを公開するため、GPU メモリ使用率が低くなり、ロングコンテキストのサポートが弱くなります。 CrossPool は、FFN 重みと KV キャッシュを 2 つの GPU メモリ プールに分離するコールド MoE モデル用のサービング エンジンです。1 つはコールド モデル全体で FFN 重みを統合する重みプール、もう 1 つは KV キャッシュにローカルな注意を保ちながらアクティブなリクエストを動的に処理する KV キャッシュ プールです。 CrossPool は、KV キャッシュ プランナーとバーチャライザー、隠し状態の転送を隠すレイヤーごとのパイプライン スケジューラー、および CPU-GPU 制御オーバーヘッドを削減するために制御を下げる永続カーネルを組み合わせています。効率的な GPU メモリ プーリングにより、CrossPool はバースト性の高いロングコンテキスト リクエストをサポートし、最先端の kvcached ベースのマルチ LLM サービング システムを上回るパフォーマンスを発揮し、P99 TBT を最大 $10.4\time$ 削減します。
原文 (English)
CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation
Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU memory problem: model weights are stable and model-determined, while KV-cache is transient and demand-determined. Because cold models rarely reach peak KV-cache demand at the same time, reserving worst-case KV capacity per model wastes memory; a shared KV-cache pool can instead provision aggregate active demand. However, KV-cache sharing is not sufficient when weights and KV-cache remain in a monolithic GPU memory pool. Static weights compete with dynamic KV-cache, and KV-head-limited attention under cold, low-concurrency traffic exposes only a fraction of replicated KV capacity, leading to low GPU memory utilization and weak long-context support. We present CrossPool, a serving engine for cold MoE models that separates FFN weights and KV-cache into two GPU memory pools: a weights pool that consolidates FFN weights across cold models, and a KV-cache pool that dynamically serves active requests while keeping attention local to KV-cache. CrossPool combines a KV-cache planner and virtualizer, a layer-wise pipeline scheduler that hides hidden-state transfers, and persistent kernels with control lowering to reduce CPU-GPU control overhead. With efficient GPU memory pooling, CrossPool underpins bursty long-context requests and outperforms the state-of-the-art kvcached-based multi-LLM serving system, reducing P99 TBT by up to 10.4x.
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety
General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimod…
Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need
Risk stratification for pulmonary embolism (PE) is critical for clinical decision-making. Stratification guidelines are based on patient me…
Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning beha…
役に立ちます: トレーニング中期の思いやりの値はトレーニング後にドメインに依存して低下します
標準的なポストトレーニング パイプラインは、教師あり微調整 (SFT) と強化学習 (RL) を適用して言語モデルを有用にしますが、これらのプロセスは、トレーニング前に注入された値を誤って低下させる可能性があります。動物危害ベンチマーク (AHB 2.2) と MORU ベンチマークで評価された SFT (Dolly-15k による有用性と Magicoder-110K によるコーディング) と GRPO (RLHFlow による有用性と Magicoder によるコーディング) の両方を使用して、思いやり指向の合成データで中間トレーニングされた Llama 3.1 8B モデルにおける動物の思いやりの値の保持に、トレーニング後のデータのドメインが差動的に影響を与えるかどうかを調査します。 (不確実性の下での道徳的推論)。有用性トレーニングは、AHB でのコーディング トレーニングと比較して動物の思いやりを大幅に低下させます (SFT: 35.7% 対 65.2%、GRPO: 18.7% 対 32.0%)。これは 2 つの独立した有用性データセットと 2 つのトレーニング パラダイムにわたって再現されています。英語のMORU項目では、有用性トレーニングは一般的な道徳的推論を25.5パーセントポイント(46.4%対71.9%)低下させ、その大きさは同情効果に匹敵する顕著な差でした。ただし、この効果は言語を越えて伝わりません。多言語の MORU ベンチマークでは、ドメイン効果は消失します (SFT: 52.3% 対 51.2%)。対照的に、動物の思いやりの効果は言語間で一貫して伝わり、Magiccoder の基本モデルに対する AHB パーセンテージ ポイントの増加は、英語以外の項目では英語の項目よりも 4.5 倍大きくなっています。この乖離は、トレーニング中に教え込まれた価値観が、ドメイン固有のトレーニング後の改善を推論するよりも深く、言語を超えてコード化されていることを示唆しています。これらの結果は、価値を満載したトレーニング途中で構築するラボの場合、トレーニング後の有用性よりも、トレーニング後のコーディング ドメインの方が、一般的な推論能力を損なうことなく、トレーニング途中の値をよりよく保存できる可能性があることを示唆しています。
原文 (English)
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training
Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may inadvertently degrade values instilled during pre-training. We investigate whether the domain of post-training data differentially affects the retention of animal compassion values in a Llama 3.1 8B model mid-trained on compassion-oriented synthetic data, using both SFT (helpfulness via Dolly-15k vs. coding via Magicoder-110K) and GRPO (helpfulness via RLHFlow vs. coding via Magicoder), evaluated on the ANIMA 2.2 benchmark and MORU benchmark (Moral Reasoning Under Uncertainty). Helpfulness training significantly degrades animal compassion relative to coding training on ANIMA (SFT: 35.7% vs. 65.2%; GRPO: 18.7% vs. 32.0%), replicating across two independent helpfulness datasets and two training paradigms. On English MORU items, helpfulness training degrades general moral reasoning by 25.5 percentage points (46.4% vs. 71.9%), a striking gap that rivals the compassion effect in magnitude. However, this effect does not transfer cross-lingually: on the multilingual MORU benchmark, the domain effect disappears (SFT: 52.3% vs. 51.2%). In contrast, the animal compassion effect transfers consistently across languages, with Magicoder's ANIMA percentage-point gain over the base model 4.5 times larger on non-English items than English items. This divergence suggests that values instilled through mid-training are encoded more deeply and cross-lingually than reasoning improvements from domain-specific post-training. These results suggest that, for labs building on value-laden mid-training, coding-domain post-training may better preserve mid-trained values than helpfulness post-training without harming general reasoning capabilities.
主張し、説明しないでください: 動物福祉に関する LLM の推論を変える言語的特徴
動物愛護活動家たちは多くの著作物を作成しており、その著作物が言語モデルを訓練し、その後何百万人もの人々が動物福祉について尋ねるようになっています。提示された動物福祉ベンチマークで語彙を一致させたスタンスコントラストプローブを使用して、10の言語的特徴のそれぞれが、微調整データとして使用された場合にラマ-3.2-1Bの動物福祉推進推論に対する好みをどのように変化させるかを測定します。 10 個の特徴のうち 8 個で、統計的に有意な変化が生じます。 7 つは、断定的な確実性、明確な道徳的語彙、感情的な言葉、評価的主張、物語の構造、描写された危害の深刻度、即時的な時間的枠組みなど、モデルをより強力な動物愛護推進の推論に向けて移行させています。 2 つはそれを逆方向に動かします。ヘッジされた言葉と具体的な感覚的説明は両方とも動物愛護推進の立場を薄めます。一人称視点には統計的に有意な効果はありません。 LLM トレーニング コーパスに組み込まれる可能性のある動物福祉に関するテキストを執筆する人に対する実際的な推奨事項: シーンを中立的に説明するのではなく、立場を主張することです。モデルを変える特徴は、作家の立場を明確にするものです。それを弱める特徴は動物愛護の内容を保持しますが、スタンスを保留します。
原文 (English)
Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare
Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare. Using vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B's preference for pro-animal-welfare reasoning when used as fine-tuning data. Eight of the ten features produce statistically significant shifts. Seven move the model toward stronger pro-animal-welfare reasoning: assertive certainty, explicit moral vocabulary, emotion words, evaluative claims, narrative structure, depicted harm severity, and immediate temporal framing. Two move it the other way: hedged language and concrete sensory description both dilute the pro-animal-welfare stance. First-person perspective has no statistically significant effect. The practical recommendation for anyone writing animal-welfare text that may end up in LLM training corpora: assert a position rather than describe a scene neutrally. The features that shift the model are the ones that make the writer's position explicit; the features that dilute it hold animal-welfare content but withhold stance.
\textsc{DiARC}: ポジティブサンプルとネガティブサンプルの区別は、大規模な言語モデルの ARC のような推論能力の向上に役立ちます
Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) には、限られたグリッド サンプルからのパターンの要約と出力グリッドの予測を必要とするタスクが含まれています。最近、多くの大規模な言語モデル ベースのアプローチが、言語モデルをテキスト ベースの推論タスクに変換しようと試みています。ただし、オープンソース モデルに基づく方法では一般に満足のいく結果が得られず、クローズドソース モデルに依存する方法ではコストがかかりすぎます。現在の取り組みは主にデータ拡張に焦点を当てており、より包括的な監視付き微調整のための ARC のようなデータを構築しています。この研究では、ARC のような問題を解決するには \textit{positive} サンプルの監視だけでなく、\textit{negative} サンプルを区別してモデル推論を改善する能力も必要であると主張します。この目的を達成するために、私たちは好みの調整のアイデアを利用し、モデルがそれらを区別できるように好みのペアを構築する方法である \textsc{DiARC} を提案します。具体的には、出力レベルの視覚的変換、DSL レベルのルール反転、およびタスク固有のルール編集を含む、ネガティブ サンプルを構築する 3 つの方法を提案します。得られた陰性サンプルは、観察されたデモンストレーションを変更せずに、有益なニアミス代替案を提供します。複数の ARC に似たベンチマークにわたる実験結果は、\textsc{DiARC} がベースライン モデルよりも一貫してパフォーマンスを向上させることを示しています。コードは https://github.com/szu-tera/DiARC で公開されています。
原文 (English)
DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models
The Abstraction and Reasoning Corpus (ARC) contains tasks that require summarizing patterns from limited grid samples and predicting output grids. Recently, many large language model based approaches have attempted to transform it into a text-based reasoning task. However, methods based on open-source models have generally yielded unsatisfactory results, while those relying on closed-source models are too costly. Current efforts mainly focus on data augmentation, constructing ARC-like data for more comprehensive supervised fine-tuning. In this work, we argue that solving ARC-like problems requires not only positive sample supervision but also the ability to improve model reasoning by distinguishing negative samples. To this end, we draw on the idea of preference alignment and propose DiARC, a method that constructs preference pairs to enable the model to distinguish between them. Specifically, we propose three ways to construct negative samples, including output-level visual transformations, DSL-level rule inversion, and task-specific rule editing. The resulting negative samples provide informative near-miss alternatives while keeping the observed demonstrations unchanged. Experimental results across multiple ARC-like benchmarks show that DiARC consistently improves performance over baseline models. The code is released at https://github.com/szu-tera/DiARC.