AIニュース 2026-07-24
自動生成: 2026-07-24 12:17 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
AIは“声で操作”する時代に? ChatGPTとClaude、相次ぎ音声機能を強化ITmedia AI+
米OpenAIと米Anthropicが相次いで自社AIサービスの音声機能を強化した。
-
Googleが“自社AIの裏切り”に備え始めた 異例の構想「AI Control Roadmap」とはITmedia AI+
米Google DeepMindが発表した異例の構想「AI Control Roadmap」について解説する。
-
三井不動産がデータセンターに6000億円超投資、物流の枠超え「産業デベロッパー」へITmedia AI+
三井不動産は事業説明会で「産業デベロッパー」への領域拡大を発表した。従来の物流拠点供給にとどまらず、研究開発施設や自動運転対応を進める。デ…
-
Anthropic updates Claude voice mode with more capable modelsTechCrunch AI
Claude's new voice model will let you reschedule your meeting or draf…
-
Runway launches AI model router as generative media gets crowdedTechCrunch AI
The Media Router is a tool that automatically selects the best image,…
-
OpenAI makes ChatGPT Health available to all US usersTechCrunch AI
Users can also integrate their personal data from services like Apple…
-
Meta launched a new AI optimism ad set to a song about human extinctionTechCrunch AI
David Bowie's song "Five Years," which Meta used in a supposedly insp…
トピック別件数
日本語メディア10件
ITmedia AI+ (日本語)
AIは“声で操作”する時代に? ChatGPTとClaude、相次ぎ音声機能を強化
米OpenAIと米Anthropicが相次いで自社AIサービスの音声機能を強化した。
図面AIに「動かせる3Dモデル」の生成機能、関節や可動域を自動認識
renueは、2D図面から3D CADモデルを生成するAI「Drawing Agent」に、関節や可動域を読み取り、動かせる3Dモデルを生成する新機能「可動アセンブリ」を追加した。生成したモデルをAIが動かし、接合部の離れや回転中心のずれなども検証する。
Googleが“自社AIの裏切り”に備え始めた 異例の構想「AI Control Roadmap」とは
米Google DeepMindが発表した異例の構想「AI Control Roadmap」について解説する。
【元経産省の専門家に聞く】中小企業がハマる「生成AIトラブル」5つの解決シナリオ
キーマンズネットの読者調査には、生成AIを巡る中小企業の切実な悩みが数多く寄せられた。代表的な5つの「あるある課題」を、経済産業省・中小企業庁でデジタル活用支援に携わった小池明氏にぶつけ、明日から使える乗り切り方を聞いた。
「AIの提案」を妄信する人、疑える人――“眼力ある人材”を育てる絶対条件
プロンプト一つでUIやコードが数秒で量産される時代、人間の役割は「制作」から「目利き」へと変わる。しかし手を動かさなくなることで、AIの提案を無批判に受け入れてしまうリスクも漂う。米Figmaのロレダナ・クリサンCDOは「AIは過去しか見ない。世界を明日へ押し進めるのは人間だ」…
「生成AIで仕事が楽に」のはずが……IT現場を蝕む“AI疲れ・AIうつ”の正体
耳にする機会が増えた「AI疲れ」「AI鬱(うつ)」。本稿では、“疲れの正体”を整理し、個人が何を考え、どう変わればよいのかという判断軸を整理します。
三井不動産がデータセンターに6000億円超投資、物流の枠超え「産業デベロッパー」へ
三井不動産は事業説明会で「産業デベロッパー」への領域拡大を発表した。従来の物流拠点供給にとどまらず、研究開発施設や自動運転対応を進める。データセンター事業には累計6000億円超を投じ、稼働済みの3棟に加え7棟を開発中だ。
「スパイダーロボ」登場 がれきを走破、モノに「触って判断」も 災害現場で活用へ 国内ベンチャー
アトラックラボ(埼玉県入間郡)は、クモの形を模したロボットを開発したと発表した。実際のクモより2本少ない6本の脚を備えており、画像や触覚情報も処理できる。災害現場や危険区域などでの活用を目指す。
三菱電機とソニー、AIビジョンセンサーで新会社設立へ
三菱電機とソニーセミコンダクタソリューションズは、製造業向けAIビジョンセンサーソリューションを開発する新会社を合弁で設立する。新会社の社名は「Advanced Vision Solutions」で、2026年10月より事業を始める予定。
Markdownファイルが、AI時代の負債に? Googleが提案する「ナレッジ標準化」の一手
Google Cloudは、AIエージェントが利用するナレッジをMarkdownで標準化するオープンフォーマット「Open Knowledge Format」を公開した。ベンダー非依存で、異なるエージェント間でもナレッジをそのまま共有できる。
海外メディア12件
TechCrunch AI (英語)
How AI guardrails are impeding the work of offensive cybersecurity researchers
We spoke with several cybersecurity researchers, who look for unknown vulnerabilities and develop tools to exploit them, about how OpenAI’s…
AMD takes on Nvidia with its Helios AI rack-scale system
AMD is challenging its chipmaker rival with a new rack-scale system that will start shipping to customers later this year.
Anthropic updates Claude voice mode with more capable models
Claude's new voice model will let you reschedule your meeting or draft an email.
AegisAI, founded by former Google security execs, lands $36M to stop AI-driven spear phishing
AegisAI co-founders developed AI agents that quickly analyze each message as a human would, paying attention to small anomalies that even t…
Runway launches AI model router as generative media gets crowded
The Media Router is a tool that automatically selects the best image, video, or audio generation model for a request based on whether a dev…
OpenAI makes ChatGPT Health available to all US users
Users can also integrate their personal data from services like Apple Health, Function, and MyFitnessPal.
Meta launched a new AI optimism ad set to a song about human extinction
David Bowie's song "Five Years," which Meta used in a supposedly inspiring advertisement, is about humans learning that they have five year…
Nvidia is sending GPUs to the moon
If there's a place in the universe without GPUs, Nvidia is sending them there.
AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors
Etched, founded by three Harvard dropouts, has created new chips and memory components that speed up inference on any AI model -- no GPUs r…
Google’s Gemini nears billion-user milestone
Gemini had over 750 million monthly users in February.
Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good
"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," one expert told TechCrunch.
ServiceNow bets $40 million on Indian banking software specialist to expand its financial services push
ServiceNow's investment gives BusinessNext a strategic partner to expand its AI-powered banking software globally.
公式ブログ0件
このカテゴリの新着記事はありませんでした。
論文240件
arXiv cs.AI (英語)
FineServe: グローバル LLM サービス提供ワークロードのきめ細かいデータセットと特性評価
大規模言語モデル (LLM) は、常時接続のオンライン サービスとして導入されることが増えており、効率的な LLM がシステムの重要な課題に対応できるようになります。不安定な需要の下で低レイテンシと高スループットを達成するには、実際のサービス提供ワークロードを深く理解する必要がありますが、既存の研究は多くの場合、プロキシ トレースや粗粒度の特性評価に依存しており、最新のマルチモデル LLM プラットフォームの異質性を捉えることができません。 FineServe は、世界の商業市場から収集された、実際のマルチモデル LLM サービング ワークロード データセットであり、異種モデルやタスクにわたる現実世界のサービング ダイナミクスのきめ細かい特性評価を可能にします。 FineServe を活用して、到着ダイナミクスとトークンの動作の包括的な分析を実施し、モデル アーキテクチャ、スケール、タスクの意図全体で根本的に異なる変動体制を明らかにします。これらの洞察に基づいて、FineServe ワークロード ジェネレーターを開発します。これは、きめ細かいモデル対応ワークロードを、マルチモデル サービング プラットフォームのベンチマークに合わせた構成可能な混合物に構成します。 FineServe は、これらのきめ細かいワークロード ダイナミクスを公開することで、LLM サービス システムにおけるルーティング、スケジューリング、およびキャパシティ プランニング戦略を評価するための現実的な基盤を提供します。 FineServe は https://github.com/hihiztc1/FineServe で入手できます。
原文 (English)
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge. Achieving low latency and high throughput under volatile demand requires deep understanding of real-world serving workloads, yet existing studies often rely on proxy traces or coarse-grained characterizations that fail to capture the heterogeneity of modern multi-model LLM platforms. We present FineServe, an in-the-wild, multi-model LLM serving workload dataset collected from a global commercial marketplace, enabling fine-grained characterization of real-world serving dynamics across heterogeneous models and tasks. Leveraging FineServe, we conduct a comprehensive analysis of arrival dynamics and token behavior, revealing fundamentally different fluctuation regimes across model architectures, scales and task intents. Building on these insights, we develop the FineServe workload generator, which composes fine-grained model-aware workloads into configurable mixtures tailored for benchmarking multi-model serving platforms. By exposing these fine-grained workload dynamics, FineServe provides a realistic foundation for evaluating routing, scheduling, and capacity-planning strategies in LLM serving systems. FineServe is available at https://github.com/hihiztc1/FineServe.
堅牢な金融詐欺検出と敵対的回復力のためのハイブリッド LSTM-Graph ニューラル フレームワーク
金融機関は、極端なデータの不均衡(不正行為率 0.13%)と敵対的回避戦術の進化により、スマーフィングやレイヤリングなどの高度なマネーロンダリング パターンを検出する際に大きな課題に直面しています。この論文では、時間的シーケンスと構造的なリレーショナル コンテキストの両方をキャプチャするために、長短期記憶 (LSTM) ネットワークと手作りのグラフ トポロジ特徴を統合するハイブリッド フレームワークである FraudShield AI を提案します。 PageRank Centrality、In-Degree Dynamics、カスタム Flow Ratio などのネットワーク中心の機能を設計することにより、システムは検出パラダイムを分離トランザクション分析からネットワーク レベルのフォレンジックに移行します。クラスの不均衡に対処するために焦点損失目標が使用され、低値のスマーフィング攻撃に対する回復力を向上させるために動的なしきい値メカニズムが導入されています。 PaySim データセットの実験評価では、提案されたハイブリッド モデルが、特に検出が困難なマイクロトランザクション詐欺パターンにおいて、精度、再現率、および F1 スコアにおいてロジスティック回帰および XGBoost ベースラインを大幅に上回ることが示されています。アブレーション研究により、時間的要素と位相的要素の両方が相補的に寄与していることが確認されています。
原文 (English)
Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience
Financial institutions face significant challenges in detecting sophisticated money laundering patterns, such as smurfing and layering, due to extreme data imbalance (0.13% fraud rate) and evolving adversarial evasion tactics. This paper proposes FraudShield AI, a hybrid framework that integrates Long Short-Term Memory (LSTM) networks with hand-crafted Graph Topological Features to capture both temporal sequences and structural relational context. By engineering network-centric features including PageRank Centrality, In-Degree dynamics, and a custom Flow Ratio, the system shifts the detection paradigm from isolated transaction analysis to network-level forensics. A Focal Loss objective is used to address class imbalance, and a dynamic thresholding mechanism is introduced to improve resilience against low-value smurfing attacks. Experimental evaluation on the PaySim dataset shows that the proposed hybrid model substantially outperforms Logistic Regression and XGBoost baselines in Precision, Recall, and F1-Score, particularly on hard-to-detect micro-transaction fraud patterns. An ablation study confirms the complementary contribution of both the temporal and topological components.
OpenEvoShield: オープンワールドのマルチエージェント システム攻撃に対するデュアル非定常継続防御
LLM ベースのマルチエージェント システム (LLM-MAS) は、セーフティ クリティカルなアプリケーションに導入されることが増えており、攻撃者はエージェント間通信を通じて悪意のある命令を注入し、有害な動作を広めます。静的な脅威とは異なり、これらの攻撃は二重に動的です。攻撃者は展開された防御に対する注入戦略を洗練しますが、通常のエージェントの動作はシステムの拡張に伴って変動します。既存の防御策は展開を閉じた世界の問題として扱い、いずれかの配信がトレーニングの範囲を超えて移行すると、急速に機能が低下します。私たちは、LLM-MAS の共進化継続的防御フレームワークである OpenEvoShield を提案します。非対称レート コントローラー (M1) は、デュアル ドリフト信号から高速な攻撃側学習速度と低速な通常側学習速度を分離します。通常境界アップデーター (M2) は、低速で動的な動作境界を維持しますが、EWC 正規化ポリシー アンサンブル (M3) は、壊滅的な忘却を起こすことなく高速に適応します。エネルギーベースの多粒度検出器 (M4) は、ノード、サブグラフ、グラフ レベルの証拠を融合して、新しい攻撃を分布外として分類します。 5 つのベンチマークと 4 つの MAS トポロジにわたる 100 回以上の導入ラウンドを超える実験では、OpenEvoShield が静的および継続的なベースラインを上回るパフォーマンスを示し、誤検知率を低く抑えながら、これまで見たことのない攻撃のほとんどを検出できることがわかりました。
原文 (English)
OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instructions through inter-agent communication to propagate harmful behaviors. Unlike static threats, these attacks are doubly dynamic: adversaries refine injection strategies against deployed defenses while normal-agent behavior drifts with system expansion. Existing defenses treat deployment as a closed-world problem and degrade rapidly once either distribution shifts beyond training coverage. We propose OpenEvoShield, a co-evolutionary continual defense framework for LLM-MAS. An asymmetric rate controller (M1) decouples fast attack-side and slow normal-side learning rates from dual drift signals. A normal-boundary updater (M2) maintains a dynamic behavioral boundary at the slow rate, while an EWC-regularized policy ensemble (M3) fast-adapts without catastrophic forgetting. An energy-based multi-granularity detector (M4) fuses node-, subgraph-, and graph-level evidence to classify novel attacks as out-of-distribution. Experiments over 100 deployment rounds across five benchmarks and four MAS topologies show that OpenEvoShield outperforms static and continual baselines, detecting most previously unseen attacks while keeping false positive rates low.
Intel TDX での NVIDIA H100 での Confidential GPU 推論のベンチマーク
機密性の高いコンピューティングは、機密入力を処理したり独自のモデル資産を保護したりする AI 推論ワークロードの実際的な導入要件になりつつあります。ただし、GPU アクセラレーションによる大規模言語モデルの提供の機密実行を可能にするパフォーマンス コストは依然としてワークロードに依存しており、運用上重要です。このペーパーでは、Intel TDX Confidential インスタンスでホストされている単一の NVIDIA H100 80GB GPU 上で、標準的な非機密実行と Confidential コンピューティング モードを比較したベンチマーク調査を紹介します。この評価では、2 つの代表的な言語モデル、Mistral-7B v0.1 と Qwen3-30B-A3B を使用し、最初のトークンまでの時間、エンドツーエンドのリクエスト レイテンシ、リクエストごとのトークン生成スループット、グローバル トークン スループット、同時実行性が増加した場合の閉ループ リクエスト スループットを測定します。固定リクエストレートの実験では、機密モードにより平均 TTFT が Mistral-7B で 21.8%、Qwen3-30B-A3B で 27.8% 増加しましたが、グローバル トークン スループットはそれぞれ 17.7% と 21.1% 減少しました。閉ループ同時実行実験では、スループット ギャップは 11.5 ~ 20.2% の範囲に留まりますが、機密モードでは大規模なモデルの方が早く飽和域に達します。この結果は、機密 GPU 推論が負荷の下でも使用可能なスループットを維持できることを示唆していますが、キャパシティ プランニングでは、安定したスループット ペナルティと、大規模なモデルで観察される初期の飽和動作の両方を考慮する必要があります。
原文 (English)
Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or protect proprietary model assets. However, the performance cost of enabling confidential execution for GPU-accelerated large language model serving remains workload dependent and operationally important. This paper presents a benchmark study comparing standard non-confidential execution with confidential computing mode on a single NVIDIA H100 80GB GPU hosted in an Intel TDX confidential instance. The evaluation uses two representative language models, Mistral-7B v0.1 and Qwen3-30B-A3B, and measures time to first token, end-to-end request latency, per-request token generation throughput, global token throughput, and closed-loop request throughput under increasing concurrency. In fixed request-rate experiments, confidential mode increases average TTFT by 21.8% for Mistral-7B and 27.8% for Qwen3-30B-A3B, while global token throughput drops by 17.7% and 21.1%, respectively. In closed-loop concurrency experiments, throughput gaps remain in the 11.5-20.2% range, but the larger model reaches its saturation knee earlier under confidential mode. The results suggest that confidential GPU inference can retain usable throughput under load, but capacity planning must account for both the steady throughput penalty and the earlier saturation behavior observed for larger models.
FormulaSPIN: 自然言語からスプレッドシートの数式生成のためのセルフプレイ微調整
スプレッドシート アプリケーションは世界中で何億もの人々に使用されていますが、数式の記述は依然として大きな障壁となっています。既存のアプローチは静的な教師付きデータに依存しているため、限られたアノテーションではすぐに飽和してしまいます。このペーパーでは、追加データなしで反復的な自己改善を可能にすることで、教師付き微調整の上限を打ち破るセルフプレイ フレームワークである FORMULASPIN を紹介します。 Vanilla SPIN はこのタスクでは失敗します。一致しないすべての出力に一律にペナルティを与えるため、実行と同等の代替案は、ある例ではネガティブとしてペナルティを受け、別の例ではグランド トゥルースとして機能し、矛盾した勾配が生成されます。私たちのフレームワークは、数式生成の独自の利点を活用することでこれを解決します。つまり、バイナリの実行可能性により、意味論的なエラーを有効な文体のバリアントから分離する暗黙的な監視が提供されます。私たちはトレーニングを 2 人用のゲームとして構成し、メイン プレーヤーが以前のバージョンの公式よりもグラウンド トゥルースの公式を好むことを学習する一方で、実行フィードバックによって出力が明確な粒度に分類され、意味的な正確さから文体の洗練へと移行する適応型カリキュラムが可能になります。精度をさらに高めるために、複数の有効な定式化を自然に処理するセマンティック レベルの投票メカニズムである ExecVote を組み込みます。複数のベンチマークでの実験では、FORMULASPIN が NL2FORMULA 上で 74.9% の完全一致と 87.1% の実行精度という最先端のパフォーマンスを実現し、追加の設定アノテーションでトレーニングされたモデルを照合しながら、従来の SFT モデルとフロンティア独自のモデルの両方を上回るパフォーマンスを示していることが実証されています。これらの発見は、セルフプレイが希少なデータタスクに取り組み、実行可能ドメインを超えて拡張する可能性があることを強調しています。
原文 (English)
FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
Spreadsheet applications are used by hundreds of millions worldwide, yet writing formulas remains a significant barrier. Existing approaches rely on static supervised data, which quickly saturates on limited annotations. In this paper, we introduce FORMULASPIN, a self-play framework that breaks the ceiling of supervised fine-tuning by enabling iterative self-improvement without any additional data. Vanilla SPIN fails on this task: it uniformly penalizes every non-matching output, so execution-equivalent alternatives are punished as negatives in one example while serving as ground truth in another, producing contradictory gradients. Our framework resolves this by exploiting formula generation's unique advantage: binary executability provides implicit supervision that separates semantic errors from valid stylistic variants. We frame training as a two-player game in which the main player learns to prefer ground-truth formulas over those from its previous version, while execution feedback sorts outputs into distinct granularities-enabling an adaptive curriculum that shifts from semantic correctness to stylistic refinement. To further increase accuracy, we incorporate ExecVote, a semantic-level voting mechanism that naturally handles multiple valid formulations. Experiments on multiple benchmarks demonstrate that FORMULASPIN achieves state-of-the-art performance, with 74.9% exact match and 87.1% execution accuracy on NL2FORMULA, matching models trained with additional preference annotations while outperforming both traditional SFT and frontier proprietary models. These findings underscore self-play's potential to tackle scarce data tasks and open the door to extending it beyond executable domains.
大規模言語モデルにおける情報識別
LLM は、インターネットなどの外部知識ソースとともに使用されることが増えています。彼らは情報を適切に評価していますか?信頼できる情報源 (情報源の識別) についてはさらに更新し、主張が事前に真実に近づいている場合 (真実の識別) はさらに更新していますか?私たちはこれを情報識別として形式化し、解釈可能な指標を備えた 3 つの規範的な公理に基づいた実験フレームワークおよびベンチマークである Learn2Discern (L2D) を導入します。外部妥当性を確立するために、事前に登録されたクォータに一致するユーザー調査 (n=299) により、実際の LLM ユーザーが 3 つの公理すべてを支持し、違反により信頼性と使用意図が低下すると報告していることが確認されています。 13 のモデルと約 67 万件のトライアルにわたって、両方の側面で一貫した失敗が見つかりました。モデルはソースと真実の識別に関してほぼ偶然に実行し、ソースの信頼性の 2 倍、ソースの人気に依存し、主張がグラウンド トゥルースと比較してその立場を改善するか悪化させるかをほぼ均等に更新します。モデルは、事前分布がすでに最も正確であるデータセットに外部知識を最も効果的に統合します。より新しく大規模なモデルは、真実の識別を改善しますが、ソースの識別は改善しません。これは、モデルの複雑さによって解決されない盲点です。私たちは、両方の形式の識別を改善する単純な推論時の介入を特定します。 LLM が従来の検索に取って代わるにつれて重要性が高まる中核的な位置合わせプロパティのテストベッドとして、データセットと調査をリリースします。
原文 (English)
Information Discernment in Large Language Models
LLMs are increasingly used with external knowledge sources like the internet. Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when claims bring priors closer to the truth (truth discernment)? We formalize this as information discernment and introduce Learn2Discern (L2D), an experimental framework and benchmark grounded in three normative axioms with interpretable metrics. To establish external validity, a pre-registered, quota-matched user study (n=299) confirms that real LLM users endorse all three axioms and report that violations reduce their trust and usage intent. Across 13 models and nearly 670K trials, we find consistent failures across both dimensions: models perform near chance on source and truth discernment, rely on source popularity twice as much as source reliability, and update roughly equally whether a claim improves or worsens their position relative to the ground truth. Models integrate external knowledge most effectively on datasets where their priors are already the most accurate. Newer and larger models improve truth discernment but not source discernment, a blind spot that model complexity does not address. We identify simple inference-time interventions that improve both forms of discernment. We release our dataset and survey as a testbed for a core alignment property that scales in importance as LLMs replace traditional search.
NEXUS: ツールを使用する LLM エージェントの構造化されたランタイムの安全性
ツールを使用する LLM エージェントは影響の大きいアクションを実行することが増えており、実行時の安全性の監視が不可欠になっています。 NEXUS (Neural EXecution Utility and Safety) は、許可、ブロック、確認要求、改訂要求の 4 つのアクションから選択するための正式な介入ポリシーを適用する構造化計画安全モニターです。 NEXUS は、段階的エスカレーションのための決定論的安全ルール、引数レベルの検査、および調整されたロジスティック回帰リスク スコアを組み合わせています。 128 インスタンスの合成ベンチマークで、NEXUS は F1 スコア 0.949、4 クラス介入精度 0.6406 を達成し、ルールのみの介入選択を 27.3 パーセントポイント上回りました。また、R-Judge でのルールのみよりも改善され (F1 = 0.861 対 0.849)、脅威モデルの制限により AgentHarm でのルールのみと一致し、IPI では 99% の制御許可で 0% ASR を達成します。ルールブラインドの NEXUS-Stress ベンチマークでは、NEXUS の F1 スコアは 0.881 に達し、きめの細かい介入ルーティングの難しさを浮き彫りにしています。遅延の中央値が 0.205 ミリ秒である NEXUS は、一般的なエージェント ループに追加するオーバーヘッドが 0.1% 未満です。コード、ベンチマーク、および調整されたリスクスコアラーは公開されています。
原文 (English)
NEXUS: Structured Runtime Safety for Tool-Using LLM Agents
Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility and Safety), a structured-plan safety monitor that applies a formal intervention policy to select among four actions: allow, block, request confirmation, or request revision. NEXUS combines deterministic safety rules, argument-level inspection, and a calibrated logistic-regression risk score for graded escalation. On a 128-instance synthetic benchmark, NEXUS achieves an F1 score of 0.949 and a 4-class intervention accuracy of 0.6406, outperforming rule-only intervention selection by 27.3 percentage points. It also improves over rule-only on R-Judge (F1 = 0.861 vs. 0.849), matches rule-only on AgentHarm due to threat-model limits, and achieves 0% ASR at 99% control allow on IPI. On the rule-blind NEXUS-Stress benchmark, NEXUS reaches an F1 score of 0.881, highlighting the difficulty of fine-grained intervention routing. With 0.205 ms median latency, NEXUS adds under 0.1% overhead to typical agent loops. Code, benchmarks, and the calibrated risk scorer are publicly released.
多目的生成レコメンダー システムのための確率的原始双対復号化
レコメンダ システム (RS) の最近の進歩により、生成モデリングによるパフォーマンスの大幅な向上が示されています。実際には、レコメンデーションには多くの場合、項目の属性や公平性の制約に対して定義された制約など、関連性を超えた複数の目的を満たす必要があるスレート (項目の順序付きリスト) の構築が含まれます。既存の多目的アプローチは、非生成設定用に設計された後処理技術に依存するか、補助目的をモデル トレーニングに直接組み込んでいます。前者は生成 RS の逐次的性質を明示的に説明していませんが、後者は大規模システムでは非現実的であることがよくあります。基礎となるモデルを変更または再トレーニングすることなく、自己回帰生成 RS を強化して多目的スレート生成をサポートする、軽量の推論時間デコード層を提案します。デコードは、オンラインの制約付き最適化問題として定式化されます。この問題では、項目が順番に選択され、関連性と補助目標の間のトレードオフが、残りの制約スラック、つまり各目標がどれだけ満たされているかに基づいて動的に調整されます。これは、生成中に関連性と補助目的のバランスを取る確率的主双近似スキームを介して実装されます。私たちは、制約違反とリグロングに関する理論的な保証を提供し、現実世界のレコメンダー システムでの大規模なオフライン実験と大規模なオンライン A/B 実験を通じて、提案されたアプローチを評価します。私たちの結果は、ユーザー満足度をゼロコストで達成した補助目標の +1.8\% の向上を含め、複数の目標のトレードオフにおいて一貫した改善を示しています。
原文 (English)
Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems
Recent advances in recommender systems (RS) have shown substantial performance gains through generative modelling. In practice, recommendation often involves constructing slates -- ordered lists of items -- that must satisfy multiple objectives beyond relevance, such as constraints defined over item attributes or fairness constraints. Existing multiobjective approaches either rely on post-processing techniques designed for non-generative settings, or incorporate auxiliary objectives directly into model training. The former does not explicitly account for the sequential nature of generative RS, while the latter is often impractical in large-scale systems. We propose a lightweight, inference-time decoding layer that augments autoregressive generative RS to support multiobjective slate generation without modifying or retraining the underlying model. We formulate decoding as an online constrained optimisation problem, where items are selected sequentially, and trade-offs between relevance and auxiliary objectives are adjusted dynamically based on the remaining constraint slack, i.e., how much of each objective remains to be satisfied. This is implemented via a stochastic primal-dual approximation scheme that balances relevance and auxiliary objectives during generation. We provide theoretical guarantees on constraint violation and regret, and evaluate the proposed approach through extensive offline experiments and a large-scale online A/B experiment in a real-world recommender system. Our results show consistent improvements in multiobjective trade-offs, including a +1.8\% gain in the auxiliary objectives achieved at zero cost to user satisfaction.
LISA: 効率的なロングコンテキスト推論のための線形インデックス付きスパース アテンション
DeepSeek-R1 などの長い思考連鎖推論モデルの最近の進歩により、テスト時間スケーリング パラダイムの下で推論コンテキストの長さがますます長くなりました。ただし、標準的なセルフ アテンションの O(n^2) の計算複雑さにより、長いシーケンスでは推論コストが急激に増大し、本番環境での長い CoT 推論の展開が制限されます。これに対処するために、最初から事前トレーニングを必要としないプラグアンドプレイのアテンション置換モジュールである LISA (Linear-Indexed Sparse Attendance) を提案します。 LISA は、元のモデル内で 2 つの軽量コンポーネントを並列に統合します。(1) O(n) 時間計算量の長距離メモリを提供する Linear Attend モジュール。 (2) Lightning Indexer は、完全なコンテキストから上位 M 個の重要なトークンを選択して、スパース セルフ アテンションにフィードします。 2 つのブランチはゲート メカニズムを介して融合され、n 個のトークンを生成するための推論の複雑さが O(n^2) から O(nM) (M << n) に軽減されます。 2 段階のトレーニング パイプラインを設計します。ステージ 1 では、長距離の依存関係を捕捉するための線形注意を統合することによってモデルを初期化し、知識の蒸留によって最適化されたスライディング ウィンドウの注意メカニズムによって補完され、凍結された教師モデルの完全な自己注意の分布に近似します。ステージ 2 では、静的なスライディング ウィンドウ メカニズムを置き換えるインデクサーをさらに導入し、より広範なコンテキストから動的なトークンの選択を可能にします。インデクサーは、新しい頭ごとの KL 発散損失を使用してトレーニングされ、その選択動作を教師モデルの注意パターンと一致させます。 DeepSeek で抽出された Qwen モデルの実験では、LISA が 16K トークンのコンテキストで 50% の推論速度向上を達成し、AIME や MATH-500 などの推論ベンチマークで平均パフォーマンスが 5.6% 向上することが実証されました。
原文 (English)
LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm. However, the O(n^2) computational complexity of standard self-attention causes inference costs to grow sharply with long sequences, limiting the deployment of long-CoT reasoning in production settings. To address this, we propose LISA (Linear-Indexed Sparse Attention), a plug-and-play attention replacement module that requires no pretraining from scratch. LISA integrates two lightweight components in parallel within the original model: (1) a Linear Attention module that provides long-range memory with O(n) time complexity; (2) a Lightning Indexer that selects the top-M important tokens from the full context to feed into a Sparse Self-Attention. The two branches are fused via a gating mechanism, reducing inference complexity from O(n^2) to O(nM) (M << n) for generating n tokens. We design a two-stage training pipeline: Stage 1 initializes the model by integrating the linear attention to capture long-range dependencies, complemented by a sliding-window attention mechanism that is optimized via knowledge distillation to approximate the full self-attention distribution of a frozen teacher model. In Stage 2, we further introduce the Indexer to replace the static sliding-window mechanism, enabling dynamic token selection from broader contexts. The Indexer is trained using a novel per-head KL divergence loss, which aligns its selection behavior with the attention patterns of the teacher model. Experiments on DeepSeek-distilled-Qwen models demonstrate that LISA achieves a 50% inference speedup under 16K-token context, while improving average performance by 5.6% on reasoning benchmarks including AIME and MATH-500.
LLM エージェントのプロファイル グラフ メモリ: ナラティブ プロファイルを介した暗黙的なクロスエンティティ トラバーサル
長期メモリはセッション間で対話する LLM エージェントにとって不可欠ですが、現在のメモリ ベンチマークは主にシングルホップの再現を評価し、マルチホップの関連付けはほとんど測定されていません。私たちは 3 つの貢献を行っています。まず、MemHop を紹介します。MemHop は、10 のソーシャル ネットワーク シナリオにわたってホップの深さ 1 ~ 5 で 1,000 の質問を行い、ホップごとの証拠の注釈を付けたマルチホップ メモリ ベンチマークです。次に、プロファイル グラフ メモリ (ProGraph) を紹介します。これは、(i) プロファイル拡張 -- LLM で書かれたプロファイル ナラティブに自然に現れるエンティティ名の部分文字列一致の走査、明示的知識グラフ構築に代わる最小限の代替手段 -- と、(ii) 圧縮残差 -- 追加 API コストゼロでプロファイル更新ごとに同時抽出される正確な日付、数量、および名前付き項目を組み合わせた 2 層メモリ アーキテクチャです。 3 番目に、フルグリッド アブレーションは、クロスベンチマーク メカニズムの特殊化を示します。プロファイル拡張によりマルチホップ推論が促進され (削除された場合、MemHop で -22.6 pp)、圧縮残差により精度再現が促進され (同時抽出されなかった場合、LoCoMo で -8.6 pp)、単一アーキテクチャ内で相互効果が 3 pp 未満になります。 ProGraph は、MemHop で平均 80.1% (FullContext リファレンスと一致)、LoCoMo で 78.4% (FullContext を 11.3 pp 上回りました) で、両方で Mem0、A-Mem、HippoRAG、および RAG を上回りました。 MemHop、ProGraph、およびベースライン実装をリリースします。
原文 (English)
Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
Long-term memory is essential for LLM agents that interact across sessions, yet current memory benchmarks primarily evaluate single-hop recall, leaving multi-hop association largely unmeasured. We make three contributions. First, we introduce MemHop, a multi-hop memory benchmark of 1,000 questions at hop depths 1-5 across 10 social-network scenarios, with per-hop evidence annotations. Second, we present Profile-Graph Memory (ProGraph), a two-layer memory architecture combining (i) profile expansion -- substring-matched traversal of entity names that naturally appear in LLM-written profile narratives, a minimal alternative to explicit knowledge-graph construction -- and (ii) compression residuals -- exact dates, quantities, and named items co-extracted with each profile update at zero extra API cost. Third, a full-grid ablation shows cross-benchmark mechanism specialization: profile expansion drives multi-hop reasoning (-22.6pp on MemHop when removed) while compression residuals drive precision recall (-8.6pp on LoCoMo when not co-extracted), with cross-effects under 3pp within a single architecture. ProGraph averages 80.1% on MemHop (matching the FullContext reference) and 78.4% on LoCoMo (exceeding FullContext by 11.3pp), outperforming Mem0, A-Mem, HippoRAG, and RAG on both. We release MemHop, ProGraph, and baseline implementations.
言語モデルにおけるリフトされた表現仮説
大規模言語モデル (LLM) は、多くの場合、個々の観察をより一般的なルールのような構造にマッピングすることによってクエリに応答します。ただし、これらの構造がどのように保存、選択、修正されるのかは依然として不明です。このプロセスを研究するために、私たちはリフト表現仮説を提案します。LLM は、分離されたインスタンス レベルの事実ではなく、共有された潜在構造を通じてメモリを更新します。この見解は、リフティングをインスタンス全体の対称性の効率的な使用としてフレーム化し、シャッタリングを粗いリフト構造をより具体的なサブタイプに洗練することとしてフレーム化します。私たちは、コンテキスト内学習、LoRA、完全な微調整にわたる制御された例外学習実験を通じて、LLM のリフティングとシャッタリングを評価します。データがネストされたルールと例外によって管理されている場合、LLM は粉砕障害に対して脆弱である一方、解除が時期尚早に発生することが多いことがわかりました。これらの結果は、LLM におけるデータとルール構造の間の関係を研究する必要性を強調しています。
原文 (English)
Lifted Representation Hypothesis in Language Models
Large language models (LLMs) often answer queries by mapping individual observations to more general rule-like structures. However, it remains unclear how these structures are stored, selected, and revised. To study this process, we propose thelifted representation hypothesis: LLMs update memory through shared latent structures rather than isolated instance-level facts. This view frames lifting as an efficient use of symmetry across instances, and shattering as the refinement of coarse lifted structures into more specific subtypes. We evaluate LLMs' lifting and shattering through controlled exception-learning experiments across in-context learning, LoRA, and full fine-tuning. We find that LLMs are vulnerable to shattering failures when data are governed by nested rules and exceptions, while lifting often occurs prematurely. These results highlight the need to study the relation between data and rule structures in LLMs.
GraphContainer: グラフ RAG メソッドの比較およびデバッグのための統合プラットフォーム
Graph RAG は、特にマルチホップの質問応答において、LLM の幻覚や古い知識を軽減します。しかし、既存のアプローチは依然として非常に細分化されており、互換性がありません。さまざまなフレームワークにわたるグラフ形式の構造の不均一性と、詳細な視覚化ツールの欠如により、検索動作の評価と比較が非常に困難になります。このギャップを埋めるために、多様なグラフ RAG ワークフローを統合して視覚化するように設計された新しいプラットフォームである GraphContainer を提案します。 GraphContainer は 2 つの重要なコンポーネントを備えています。(1) マルチフォーマットのグラフをシームレスに標準化する統合グラフ表現 (UGR) レイヤーと、(2) 段階的な取得プロセスを追跡し、視覚的にレンダリングするグラフ レコーダーです。インタラクティブな Web インターフェイスを通じて、異種グラフをインポートし、グラフ RAG メソッドのライブで追跡可能な視覚的なデバッグを実行する GraphContainer の機能を示します。最終的には、GraphContainer がさまざまなグラフ形式と取得戦略の制御された比較を可能にし、研究者や実務者が最適なグラフ RAG パイプラインを設計する障壁を下げる方法を示します。デモビデオは https://youtu.be/O02eNJLwkU0 でご覧いただけます。
原文 (English)
GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods
Graph RAG mitigates hallucinations and stale knowledge in LLMs, particularly for multi-hop question answering. However, existing approaches remain highly fragmented and incompatible. The structural heterogeneity of graph formats across different frameworks and the lack of granular visualization tools make it exceedingly difficult to evaluate and compare retrieval behaviors. To bridge this gap, we propose GraphContainer, a novel platform designed to unify and visualize diverse graph RAG workflows. GraphContainer features two key components: (1) a Unified Graph Representation (UGR) layer that seamlessly standardizes multi-format graphs, and (2) a Graph Recorder that tracks and visually renders the step-by-step retrieval process. Through an interactive web interface, we demonstrate GraphContainer's ability to import heterogeneous graphs and perform live, traceable visual debugging of graph RAG methods. Ultimately, we show how GraphContainer enables controlled comparisons of various graph formats and retrieval strategies, lowering the barrier for researchers and practitioners to design optimal graph RAG pipelines. A demonstration video is available at https://youtu.be/O02eNJLwkU0.
AdaRoPE: すべてのアテンション ヘッドが均等に回転および拡大縮小する必要はない
Rotary Position Embedding (RoPE) は、位置情報をエンコードするために Transformers で広く採用されていますが、標準的な実装では、すべてのアテンション ヘッドにわたって均一な周波数スケジュールとスケーリングが強制されます。簡略化された検索タスクと長さの一般化シナリオを使用して、異なる機能的役割を持つヘッドが効果的に動作するには、異なる周波数範囲と注意スケーリング係数が必要であることを経験的にも理論的にも示しました。この構造を無視すると、特に長いコンテキスト設定の下では、埋め込みディメンションが最適に利用されず、パフォーマンスが低下します。これらの制限に対処するために、我々は、各アテンションヘッドに学習可能な回転周波数とアテンションスケーリング係数を装備する AdaRoPE を提案します。 AdaRoPE を使用した事前トレーニング済み LLM は、部分的な RoPE ベースラインや NoPE ベースラインを含む既存の RoPE バリアントよりも一貫して優れたパフォーマンスを発揮します。コンテキスト拡張については、YaRN などの方法で使用される均一な頻度と注意のスケーリングが最適ではないことをさらに示します。ヘッド固有のスケーリングを適用することで、AdaRoPE は、外挿設定とロングコンテキストの継続事前トレーニング設定の両方でショートコンテキストのパフォーマンスをよりよく維持しながら、コンテキストの拡張を向上させます。これらの結果は、個々のアテンション ヘッドのレベルで回転位置の埋め込みを最適化することの重要性を強調しています。
原文 (English)
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate effectively. Ignoring this structure leads to suboptimal utilization of embedding dimensions and degraded performance, particularly under long-context settings. To address these limitations, we propose AdaRoPE, which equips each attention head with learnable rotation frequencies and attention scaling factors. Pretrained LLMs with AdaRoPE consistently outperform existing RoPE variants, including partial RoPE and NoPE baselines. For context extension, we further show that uniform frequency and attention scaling, used in methods such as YaRN, are suboptimal. By applying head-specific scaling, AdaRoPE enables better context extension while better preserving short-context performance in both the extrapolation setting and the long-context continued pretraining setting. These results highlight the importance of optimizing rotary position embedding at the level of individual attention heads.
大規模言語モデルにおける活性化空間制御のための統計的に根拠のある疎特徴介入
アクティベーション ステアリングは、大規模な言語モデルの動作制御のための微調整に代わる軽量の代替手段を提供しますが、SAE ベースのステアリング手法は、多くの場合、学習されたステアリング目標または単一基準の機能選択に依存します。透過的な SAE 特徴ステアリング パイプラインを導入します。これは、最初に 6 条件の信頼性フィルターを適用し、次に 3 つの相補統計 ($F$-test、KSG 相互情報量、および Cohen の $d$) に対する重み付けされていない Borda コンセンサスを通じて疎な特徴をランク付けします。結果として得られるステアリング方向は、SAE デコーダ行のコーエン $d$ 重み付け組み合わせとして構築され、近似的な SAE 特徴非相関のもとでフィッシャー LDA によって動機付けられる最適化のない方向を提供します。この方法は、3 つの Gemma ファミリー モデル、4 つの動作ドメイン、および 356 の層強度構成にわたって、測定可能なドメイン固有の変化を生成しながら、生の属性の動きと品質を保持した生成との間に大きなギャップがあることを明らかにします。最も強力な構成では、論理的正確さのステアリングは、Gemma~2 9B で $+1.16$ のプライマリ スコア デルタに達します。ただし、より広範な発見は、使用可能なステアリングはモデル、ドメイン、レイヤー、強度によって非常に局所的であるということです。これらの結果は、アクティベーション・ステアリング評価では、生の行動の変化とともに品質条件付きの成功を報告する必要があることを主張しています。コードとデータは https://github.com/Oshayer-Siddique/LLM-Steering-Using-SAE で入手できます。
原文 (English)
Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection. We introduce a transparent SAE-feature steering pipeline that first applies a six-condition reliability filter, then ranks sparse features through an unweighted Borda consensus over three complementary statistics: $F$-test, KSG mutual information, and Cohen's $d$. The resulting steering direction is constructed as a Cohen's-$d$-weighted combination of SAE decoder rows, providing an optimization-free direction motivated by Fisher-LDA under approximate SAE-feature decorrelation. Across three Gemma-family models, four behavioral domains, and 356 layer-strength configurations, the method produces measurable domain-specific shifts while revealing a substantial gap between raw attribute movement and quality-preserving generation. In the strongest configuration, logical-correctness steering reaches a primary-score delta of $+1.16$ in Gemma~2 9B; however, our broader finding is that usable steering is highly localized by model, domain, layer, and strength. These results argue that activation-steering evaluations should report quality-conditioned success alongside raw behavioral shift. Our code and data are available at https://github.com/Oshayer-Siddique/LLM-Steering-Using-SAE.
回答セット プログラミングと大規模言語モデルを使用したロジックに基づくデータ抽出
大規模言語モデル (LLM) が非構造化テキストからの意味論的データの抽出に使用され、自然言語から関係事実の候補が生成される場合、複雑な組み合わせ推論とグローバルな一貫性を必要とするタスクでは信頼性が低いままになる可能性があります。この論文では、LLM ベースの抽出と回答セット プログラミング (ASP) を組み合わせた、ロジックに基づくデータ抽出フレームワークを提案します。 LLM は候補ファクトを生成しますが、ASP は検証、推論、一貫性チェック、および制御を実行します。すべてのターゲット述語に対して個別に LLM にクエリを実行する既存のパイプラインとは異なり、提案されたアプローチでは、ASP 推論を使用して、各段階で論理的に許容される述語を特定し、抽出クエリをガイドします。 LLM 呼び出しと ASP 派生をインターリーブすることにより、フレームワークはそれ以上抽出することなく論理的に暗示された事実を推論し、不一致を早期に検出します。パイプラインを形式化し、穏やかな仮定の下で、それが最終的に抽出されたファクトに関してベースライン アプローチと同等であると同時に、必要な LLM 呼び出しが少ないことを証明します。また、ロジックベースの制御クエリ用のキャッシュ メカニズムも導入し、段階的に構築されたファクト セットに対する論理積クエリの単調性を利用して、ソルバーの呼び出しを削減します。 ASP 由来のベンチマークの実験では、フレームワークが LLM 呼び出しを削減し、偽の出力を軽減することで抽出品質を向上させることが示され、制御されたセマンティック抽出のための非単調ロジック プログラミングの価値が実証されました。
原文 (English)
Logic-Guided Data Extraction with Answer Set Programming and Large Language Models
When Large Language Models (LLMs) are used for semantic data extraction from unstructured text, producing candidate relational facts from natural language, they may remain unreliable for tasks requiring complex combinatorial reasoning and global consistency. This paper proposes a logic-guided data extraction framework combining LLM-based extraction with Answer Set Programming (ASP). The LLM produces candidate facts, whereas ASP performs validation, inference, consistency checking, and control. Unlike existing pipelines that query the LLM independently for all target predicates, the proposed approach uses ASP reasoning to identify which predicates are logically admissible at each stage and to guide extraction queries. By interleaving LLM calls with ASP derivation, the framework infers logically implied facts without further extraction and detects inconsistencies early. We formalize the pipeline and prove that, under mild assumptions, it is equivalent to the baseline approach with respect to the final extracted facts, while requiring fewer LLM calls. We also introduce a caching mechanism for logic-based control queries, exploiting monotonicity of conjunctive queries over incrementally constructed fact sets to reduce solver invocations. Experiments on ASP-derived benchmarks show that the framework reduces LLM calls and improves extraction quality by mitigating spurious outputs, demonstrating the value of non-monotonic logic programming for controlled semantic extraction.
LLM 安全分類のための幾何学に基づく制約学習
Safety as Polytope (SaP) は、LLM 隠れ空間の線形半空間制約を学習しますが、カテゴリごとに制約カウント K を調整する必要があります。スパース オートエンコーダ (SAE) 特徴抽出によってこれが解決されることを示します。Qwen3.5-9B では、12/14 カテゴリに対して K=2 が最適になり、BeaverTails 分類ベンチマークでカテゴリごとに 96 ~ 99% の精度を達成し、徹底的なスイープの必要性が大幅に排除されます (K=4 ~ 25ランダムな初期化)。この 2 つの平面への収束は線形表現仮説と一致しており、この設定における安全境界が SAE 特徴空間における低次元の線形記述を許容するという示唆的な証拠を提供します。この幾何学的な観点に基づいて、学習可能な開口が各カテゴリのクラスター集中に適応し、3 段階のトレーニングによって安定化された円錐制約を導入します。
原文 (English)
Geometry-Guided Constraint Learning for LLM Safety Classification
Safety as Polytope (SaP) learns linear half-space constraints in LLM hidden space but requires per-category tuning of the constraint count K. We show that sparse autoencoder (SAE) feature extraction resolves this: K=2 becomes optimal for 12/14 categories on Qwen3.5-9B, achieving 96-99% accuracy per category on our BeaverTails classification benchmark, largely eliminating the need for exhaustive sweeps (K=4-25 with random initialization). This convergence to two planes is consistent with the Linear Representation Hypothesis, providing suggestive evidence that safety boundaries in this setting admit a low-dimensional linear description in the SAE feature space. Building on this geometric perspective, we introduce a cone constraint whose learnable aperture adapts to each category's cluster concentration, stabilized by a three-phase training
大規模言語モデルにおける不確実性評価の再考
キャリブレーションは、LLM の信頼性を評価するための主要な基準ですが、不十分です。キャリブレーションでは、自明に一貫性のない推定値が認められ、評価分布に依存し、推定値が一貫した基礎的な確率関数としてどの程度解釈できるかをテストしていません。実際に必要なのは、LLM 信頼推定値が一貫した確率的信念に必要な条件を満たすことです。これらの条件を 3 つの軸 (構造的一貫性、忠実性、有用性) に沿って形式化し、C1 メトリクスとして運用可能にします。広く使用されている推定器は、適切に校正されているように見えても、体系的にこれらの条件に違反しています。モデルは、論理的に簡単な質問に対して 31\% の確率で低い信頼度を割り当て、RMSCE を削減する一般的な介入では構造違反は変化せず、校正が確率的妥当性と直交していることを示唆しています。 RLHF と思考連鎖は、一貫性を回復することなく有用性の指標を向上させます。私たちの結果は、現在の LLM 信頼推定値が一貫した確率として解釈できないことを示しています。私たちのフレームワークは、このギャップを測定して埋めるためのツールを提供します。
原文 (English)
Rethinking Uncertainty Evaluation in Large Language Models
Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function. What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic beliefs. We formalize these conditions along three axes (structural coherence, faithfulness, and usefulness) and operationalize them as the C1 metrics. Widely used estimators systematically violate these conditions despite appearing well-calibrated: models assign lower confidence to logically easier questions 31\% of the time, and common interventions reducing RMSCE leave structural violations unchanged, suggesting that calibration is orthogonal to probabilistic validity. RLHF and chain-of-thought improve usefulness metrics without restoring coherence. Our results show current LLM confidence estimates cannot be interpreted as coherent probabilities; our framework provides the tools to measure and close this gap.
Spectral-LSH: クリロフ投影局所性依存ハッシュによる二次二次プロンプト圧縮
プリフィル アテンションはシーケンスの長さに応じて二次関数的にスケールされるため、ロング プロンプト推論は依然として高価です。私たちは、プロンプトが言語モデルに入る前に動作する、トレーニング不要のプロンプト圧縮方法である Spectral-LSH を提案します。 Spectral-LSH は、クリロフ部分空間法とランダム特徴を使用して、暗黙的なアテンション カーネル演算子の主要成分を近似し、明示的な $O(N^2)$ アテンション カーネルの実体化を回避します。次に、結果のアテンション固有空間に SimHash を適用して、類似したトークンをグループ化し、因果関係のある位置割り当てを持つマクロ トークンに集約します。 C4 で Mistral-7B-Instruct-v0.3、Qwen2.5-7B-Instruct、および Qwen2.5-14B-Instruct を評価します。私たちの実験では、圧縮比の相転移が明らかになりました。 $\rho = 4 \times$ 未満では、ローカル トークンの冗長性が十分に低いため、通常、軽量のチャンクが最適なレイテンシー、つまり品質のトレードオフを提供します。 $\rho = 8 \times$ を超えると、スペクトル パスはチャンク化によって失われる品質を維持します。 $\rho = 16 \times$ では、Qwen2.5-7B (アダプティブ) では PPL 比が 353.409 から 196.963 に減少し、Qwen2.5-14B (アダプティブ) では 9.533 から 3.427 に減少します。 JSON のような、コードのような、テーブルのような入力を含む小規模なロングコンテキストの構造化ストレス テストでも、ローカル LSH はチャンクよりも $8 \times$ ですべてのメトリクスを改善しました。適応型バックエンドは、低圧縮ではチャンク パスを使用し、高圧縮ではスペクトル クラスタリングを使用して両方の状況をキャプチャしますが、総遅延ではチャンクが依然として最速のバックエンドです。
原文 (English)
Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
Long-prompt inference remains expensive because prefill attention scales quadratically with sequence length. We propose Spectral-LSH, a training-free prompt compression method that operates before the prompt enters the language model. Spectral-LSH approximates the dominant components of an implicit attention-kernel operator using a Krylov subspace method together with random features, avoiding explicit $O(N^2)$ attention-kernel materialization. It then applies SimHash in the resulting attention eigenspace to group similar tokens and aggregate them into macro-tokens with causal positional assignments. We evaluate Mistral-7B-Instruct-v0.3, Qwen2.5-7B-Instruct, and Qwen2.5-14B-Instruct on C4. Our experiments reveal a compression-ratio phase transition. Below $\rho = 4 \times$, local token redundancy is low enough that lightweight chunking typically provides the best latency--quality trade-off. Above $\rho = 8 \times$, the spectral path preserves quality that chunking loses. At $\rho = 16 \times$, Qwen2.5-7B (adaptive) reduces the PPL ratio from 353.409 to 196.963, while Qwen2.5-14B (adaptive) reduces it from 9.533 to 3.427. On a small long-context structured stress test containing JSON-like, code-like, and table-like inputs, local LSH also improves every metric over chunking at $8 \times$. The adaptive backend captures both regimes by using the chunk path at low compression and spectral clustering at high compression, although chunking remains the fastest backend in total latency.
トラッキングやショートカットを超えて: ポーカー自己回帰モデルにおける構成境界予測状態
隠れ状態プローブは、不完全情報シーケンス モデルの潜在ラベルを回復することがよくありますが、これだけでは、モデルが隠れ状態にわたって事後信頼分布を維持していることを確立することはできません。この論文では、対戦相手のハンドやレンジではなく、アクションとバリュー ターゲットのみでトレーニングされたノーレンジ リミット ホールデム自己回帰モデルにおけるこの曖昧さを研究します。相手の範囲のプローブは、3 つのシードのうち 2 つでアクション/価値コントロールの後でポジティブであり、行動ヘッドは、観察可能な公開履歴のみを使用して、ベースラインを約 5 パーセント上回る保留されたアクションを予測します。ただし、目に見える公開ベッティング構成は、残りの隠れた状態よりも対戦相手の範囲シグナルを説明しており、ほとんどの回収可能な情報はベッティングの概要から得られることを示唆しています。アクション/値 + 組成ベースラインはトップ 10 精度 16.5 ~ 16.7% に達しますが、組成残余の隠れプローブは 11.4 ~ 12.2% に低下し、一致した組成の比較はすべてのシードでマイナスになります。私たちはこの証拠パターンを構成限定予測サポートと呼びます。隠れ状態は行動予測と相手レンジの相関関係を保ちますが、回復可能なレンジ情報のほとんどは、残存する隠れ状態構造ではなく、目に見えるベッティング構成によって説明されます。これは、相手の範囲の表現証拠に関するケーススタディの主張であり、正確なベイズ事後追跡や因果関係の信念メカニズムではありません。合成コントロールとオラクルの検証では、同じ診断が事後敏感状態を受け入れ、一致するコントロールの下で生の組成状態を拒否することが示されています。したがって、肯定的な信念プローブは、信念追跡の証拠として扱われる前に、ターゲットを絞った代替案を通じて解釈される必要があります。
原文 (English)
Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models
Hidden-state probes often recover latent labels in imperfect-information sequence models, but this alone does not establish that a model maintains a posterior belief distribution over hidden states. This paper studies this ambiguity in a no-range Limit Hold'em autoregressive model trained only on action and value targets, not on an opponent's hand or range. Opponent-range probes are positive after action/value controls in two of three seeds, and the behavior head predicts held-out actions about five percentage points above a baseline using only observable public history. However, visible public betting composition explains more opponent-range signal than residual hidden states, suggesting that most recoverable information comes from betting summaries. Action/value+composition baselines reach 16.5-16.7% top-10 accuracy while composition-residual hidden probes fall to 11.4-12.2%, and matched-composition comparisons are negative in every seed. We call this evidence pattern composition-bounded predictive support: hidden states remain behavior-predictive and opponent-range correlated, but most recoverable range information is explained by visible betting composition rather than residual hidden-state structure. This is a case-study claim about opponent-range representational evidence, not exact Bayesian posterior tracking or a causal belief mechanism. Synthetic control and oracle validations show that the same diagnostics accept posterior-sensitive states and reject raw composition states under matched controls. Thus positive belief probes should be interpreted through targeted alternatives before being treated as evidence of belief tracking.
表現の調整によるソクラテス家庭教師の足場崩壊の軽減
大規模言語モデル (LLM) に基づいたソクラテス派の家庭教師は、複数回転の質問を通じて生徒を指導することが増えていますが、足場の崩壊に悩まされる可能性があります。生徒からの持続的な圧力の下で、家庭教師は徐々に誘導された探究を放棄し、解決策を直接明らかにします。従来の防御は主に、プロンプト、プリファレンスの最適化、またはフィルタリングを通じて観察可能な応答を制限しており、軌道レベルの崩壊に先立つ内部表現のドリフトはほとんど対処されていません。我々は、最初に教師付き微調整でソクラテス的家庭教師をウォームアップし、次に軌道重み付けされた直接優先最適化と凍結された参照状態に固定されたマージン保存表現損失を組み合わせる 2 段階のフレームワークである Scaffold-Preserving Representation Alignment を提案します。私たちの方法は、対話ターン全体にわたって、足場を維持する隠れ状態と崩壊を引き起こす隠れ状態との間の分離を維持するように設計されています。私たちは、5 つの STEM 分野と 5 つのレッドチーム攻撃戦略にわたって手法を評価します。 Qwen3-8B では、私たちの方法は崩壊率を 32% に下げ、平均崩壊開始を 9 ターンを超えて遅らせ、過剰拒否を低く抑えます。これは、表現レベルの調整により、レッドチームプロトコルの下で長期的なソクラテス的個別指導の堅牢性を向上できることを示唆しています。
原文 (English)
Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment
Large language model (LLM)-based Socratic tutors increasingly guide students through multi-turn questioning, but they can suffer from scaffolding collapse: under sustained student pressure, a tutor gradually abandons guided inquiry and reveals solutions directly. Prior defenses primarily constrain observable responses through prompting, preference optimization, or filtering, leaving the internal representation drift that precedes trajectory-level collapse largely unaddressed. We propose Scaffold-Preserving Representation Alignment, a two-stage framework that first warms up a Socratic tutor with supervised fine-tuning, then combines trajectory-weighted direct preference optimization with a margin-preserving representation loss anchored to frozen reference states. Our method is designed to maintain separation between scaffold-preserving and collapse-inducing hidden states across dialogue turns. We evaluate our method across five STEM disciplines and five red-teaming attack strategies. On Qwen3-8B, our method lowers Collapse Rate to 32%, delays average collapse onset beyond nine turns, and keeps over-refusal low, suggesting that representation-level alignment can improve the robustness of long-horizon Socratic tutoring under our red-teaming protocol.
Euclean: リーンでの統合検証による自動幾何学問題の定式化
最近の形式推論システムは IMO レベルのパフォーマンスに達していますが、断片的な状況が残されています。代数と数論はリーンで処理されますが、幾何学は依然として形式保証が限られたドメイン固有の言語に依存しています。この分割により、信頼できるコンピューティング ベースが増加し、統合モデル開発が妨げられます。リーン幾何学の既存の取り組み (LeanEuclid、LeanGeo) は、標準の Mathlib と互換性のないカスタム公理システムを導入しており、その小規模なスケール ($<$ 1,100 の問題) により大規模なトレーニングが制限されています。ただし、ネイティブの Mathlib によるジオメトリの自動形式化には、明確な課題が伴います。つまり、暗黙的な図式的な仮定 (例: トポロジー構成や非縮退) は、外部ソルバーに委ねるのではなく明示的に行う必要があり、モデルは、Mathlib の小規模で急速に進化するジオメトリ インフラストラクチャに適応する必要があります。ネイティブ Mathlib でジオメトリを自動的に形式化するための、制約の説明、構成の固定、形式化マッピング、反復修復の 4 段階のフレームワークである Euclean を紹介します。私たちは、リーンにおける最大の幾何形式化データセットである OMNI-Geometry (768 の競合問題) と Numina-Geometry (177,597 の問題) を構築します。人間による評価では、TOP1 の精度は 48.89%、TOP5 の精度は 73.33% でした。形式化に基づいて Goedel v2 をトレーニングすると、証明の成功率が 13.6% から 15.1% に向上し、統一神経定理証明のデータセットの品質が検証されました。コードとデータセット: https://github.com/tlb-22/Euclean。
原文 (English)
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean
Recent formal reasoning systems have reached IMO-level performance, yet they leave a fragmented landscape: algebra and number theory are handled in Lean, while geometry still relies on domain-specific languages with limited formal guarantees. This split increases the trusted computing base and hinders unified model development. Existing geometry-in-Lean efforts (LeanEuclid, LeanGeo) introduce custom axiom systems incompatible with standard Mathlib, and their small scale ($<$ 1,100 problems) limits large-scale training. Native Mathlib autoformalization of geometry, however, poses distinct challenges: implicit diagrammatic assumptions (e.g., topological configuration and non-degeneracy) must be made explicit rather than deferred to external solvers, and models must adapt to Mathlib's small, rapidly evolving geometry infrastructure. We present Euclean, a four-stage framework - constraint explication, configuration anchoring, formalization mapping, and iterative repair - for automatically formalizing geometry in native Mathlib. We construct OMNI-Geometry (768 competition problems) and Numina-Geometry (177,597 problems), the largest geometry formalization dataset in Lean. Human evaluation shows 48.89% TOP1 and 73.33% TOP5 accuracy. Training Goedel v2 on our formalizations improves proof success from 13.6% to 15.1%, validating dataset quality for unified neural theorem proving. Code and datasets: https://github.com/tlb-22/Euclean.
CrackedPDFs: PDF での隠しプロンプト挿入の制御されたベンチマーク
ドキュメントベースの LLM システムは、多くの場合、ガードレールが PDF を検査する前に PDF をフラット化します。このステップにより、命令がユーザーに決して表示されなかったという証拠が破棄される可能性があります。 PDF への非表示プロンプト挿入の制御されたベンチマークである CrackedPDFs を紹介します。このベンチマークには、4,983 の基本ドキュメントから生成された 29,322 の PDF が含まれています。これには、9,774 個の挿入されたファイルと、19,548 個の良性または一致する交絡因子ファイルが含まれます。 PromptGuard とルール ベースラインを評価します。また、構造のみを学習したモデルとサニタイズされたハイブリッド検出器も評価します。評価では、保留された来歴分割とペアの良性交絡因子対照を使用します。また、ラベル シャッフル チェックとショートカット監査も使用します。 2,919 枚の文書を保持したテスト セットでは、ハイブリッド検出器は 0.960 F1 に達します。 ROC-AUC は 0.998、PR-AUC は 0.997 です。また、973 ペア中 95.9% において、挿入されたファイルが一致する良性交絡因子よりも上位にランク付けされます。プロンプト ガードは、抽出されたテキストのみが与えられた場合、再現率が低くなります。構造のみの学習モデルは、ペアの制御下では弱いです。テキストのみの TF-IDF モデルは、完璧なホールドアウト スコアに達しますが、ショートカット監査には失敗します。これらの結果は、ドキュメントを認識したハイブリッド検出が、制御されたペア評価の下で有用であることを示しています。これらは、広範な現実世界の堅牢性や、信頼できるファミリー間の一般化を示しません。
原文 (English)
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
Document-based LLM systems often flatten a PDF before guardrails inspect it. That step can discard evidence that an instruction was never visible to the user. We introduce CrackedPDFs, a controlled benchmark for hidden prompt injection in PDFs. The benchmark contains 29,322 generated PDFs from 4,983 base docu ments. It includes 9,774 injected files and 19,548 benign or matched-confounder files. We evaluate PromptGuard and a rule baseline. We also evaluate structural only learned models and a sanitized hybrid detector. The evaluation uses held-out provenance splits and paired benign-confounder controls. It also uses label-shuffle checks and shortcut audits. On a 2,919-document held-out test set, the hybrid de tector reaches 0.960 F1. ROC-AUC is 0.998 and PR-AUC is 0.997. It also ranks injected files above matched benign confounders in 95.9% of 973 pairs. Prompt Guard has low recall when given extracted text only. Structural-only learned mod els are weak under paired controls. A text-only TF-IDF model reaches perfect held-out scores but fails shortcut audits. These results show that document-aware hybrid detection is useful under controlled paired evaluation. They do not show broad real-world robustness or reliable cross-family generalization.
HyGRL: 複数エンティティの質問に対する適応型ハイブリッド グラフ推論
複数エンティティの構成的な質問は、既存の検索拡張言語モデルに重大な課題をもたらします。従来の手法はジレンマに陥ります。標準の RAG には動的な推論が欠けており、従来の Graph-RAG は構造の疎性によって制限され、LLM で構築された Graph-RAG には法外なコストがかかります。私たちは、非構造化テキストを構造化ナレッジ グラフに埋め込み、柔軟な証拠検索のための異種ネットワークを作成する統一フレームワークである \textbf{\fwa} を提案します。推論は適応構造誘導として定式化され、堅牢な 2 段階のプロセスを通じて学習されます。(1) 模倣学習によりヒューリスティックなエキスパート信号が抽出され、(2) 強化学習により LLM 主導の優先順位報酬を使用してポリシーが洗練されます。実験では、{\fwa} がテキストの豊かさと構造的知識を効果的に融合し、極めて低いトークン コストとほぼリアルタイムの推論を維持しながら、回答精度と推論忠実度において SOTA ベースラインを上回っていることが実証されています((コードは https://github.com/wjywjy123/HyGRL で入手可能) 。
原文 (English)
HyGRL: Adaptive Hybrid Graph Reasoning for Multi-Entity Questions
Multi-entity compositional questions pose significant challenges to existing retrieval-augmented language models. Conventional methods fall into a dilemma: standard RAG lacks dynamic reasoning, traditional Graph-RAG is limited by structural sparsity, and LLM-constructed Graph-RAG incurs prohibitive costs. We propose \textbf{\fwa}, a unified framework that embeds unstructured text into structured knowledge graphs, creating a heterogeneous network for flexible evidence retrieval. Reasoning is formulated as adaptive structure induction, learned via a robust two-stage process: (1) imitation learning distills heuristic expert signals, and (2) reinforcement learning refines the policy using LLM-driven preference rewards. Experiments demonstrate that {\fwa} effectively merges textual richness with structural knowledge, outperforming SOTA baselines in answer accuracy and reasoning fidelity while maintaining extremely low token costs and near real-time inference((code available at https://github.com/wjywjy123/HyGRL) .
ITPEval: インタラクティブな定理証明者間の形式変換のベンチマーク
形式的定理証明は機械学習の最前線の課題として浮上していますが、エコシステムは断片化しています。証明は互換性のないシステム間でサイロ化されたままであり、学習ベースの証明者のトレーニング データと検証結果の移植性の両方が制限されています。我々は、4 つの主要な ITP (Lean 4、Rocq、Isabelle、HOL Light) にわたる自動化された正式な校正翻訳を評価するための最初のベンチマークである ITPEval を紹介します。このベンチマークは、2 つの異なる論理基盤にまたがります。私たちのベンチマークは、基本的な翻訳の難しさを分離する公理化されたファイルの制御された層と、API と証明スタイルの不一致を明らかにする実際のライブラリから抽出されたエコシステム層に編成された 1,560 のソース ファイルと 6,848 の定理で構成されています。私たちは、アーティファクトごとのネイティブ チェック セマンティクスを保持する、状態分離されたウォーム バックエンドを備えた統合マルチ ITP 検証インフラストラクチャである itpeval をリリースします。 12 の有向翻訳ペアで 5 つのフロンティアおよびオープンウェイト LLM にわたるステートメント翻訳と証明翻訳の両方を評価します。ステートメント翻訳のピークは 29.1% pass@1、証明翻訳のピークは 10.5% です。制御定理の証明パス@1 は 29.7% に達しますが、エコシステム レベルの変換では 5.2% に達し、ライブラリの不一致が主要なボトルネックであることが確認されています。 pass@k 評価に加えて、決定論的な Lean 4 BEq チェックにより、検証されたソースから Lean 4 miniF2F ステートメントへの変換の 54.0% について等価性が確立され、ネイティブ型チェックだけでは意味の忠実度を大幅に過大評価する可能性があることが示されています。自動形式化/自動非形式化のラウンドトリップ調査では、Rocq と HOL Light は Lean 4 や Isabelle よりも形式化のターゲットとなりやすい一方で、マルチ ITP コンテキストにより、プールされた Lean 4 の成功率が 4.8% から 10.6% に向上しました。当社のベンチマーク、検証インフラストラクチャ、評価パイプラインは一般に公開されています。
原文 (English)
ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers
Formal theorem proving has emerged as a frontier challenge for machine learning, yet the ecosystem is fragmented: proofs remain siloed across incompatible systems, limiting both training data for learning-based provers and the portability of verified results. We present ITPEval, the first benchmark for evaluating automated formal proof translation across four major ITPs (Lean 4, Rocq, Isabelle, and HOL Light), spanning two distinct logical foundations. Our benchmark comprises 1,560 source files and 6,848 theorems organized into a controlled tier of axiomatized files that isolates foundational translation difficulty, and an ecosystem tier drawn from real libraries that exposes API and proof-style mismatches. We release itpeval, a unified multi-ITP verification infrastructure with state-isolated warm backends that preserve per-artifact native checking semantics. We evaluate both statement and proof translation across five frontier and open-weight LLMs on 12 directed translation pairs: statement translation peaks at 29.1% pass@1 and proof translation at 10.5%; controlled theorems reach 29.7% proof pass@1 versus 5.2% for ecosystem-level translations, confirming that library mismatch is the dominant bottleneck. In addition to pass@k evaluation, a deterministic Lean 4 BEq check establishes equivalence for 54.0% of verified source-to-Lean 4 miniF2F statement translations, showing that native type-checking alone can substantially overestimate semantic fidelity; in an autoformalization/auto-informalization round-trip study, Rocq and HOL Light are easier formalization targets than Lean 4 and Isabelle, while multi-ITP context improves pooled Lean 4 success from 4.8% to 10.6%. Our benchmark, verification infrastructure, and evaluation pipelines are publicly released.
FORCE-Bench: エンタープライズ ファイナンスにおけるエージェントティック AI のベンチマーク、データセット、評価ハーネス
大規模言語モデルの最近の進歩により、運用財務におけるエージェント システムの導入が加速しています。既存のベンチマークは、一般的な能力、指示への従うこと、または安全性の測定に重点を置いていますが、自動化するために現在エージェント システムが導入されている運用財務ワークフローに直接取り組んでいるベンチマークはほとんどありません。財務専門家は、エージェントに対し、事実に基づいた適切な根拠に基づいた情報を提供するだけでなく、その情報が検証可能であり、運用財務領域のルールと制約に一貫して準拠していることを確認することを求めます。 FORCE-Bench を紹介します。これには専門家が注釈を付けた 251 のクエリが含まれており、正確性、引用、明確さ、深さ、根拠性、最新性、関連性、構造の 8 つの側面にわたって、運用財務ドメインの要件に合わせて調整されたルーブリックベースのフレームワークを使用して応答を評価します。 FORCE-Bench は、財務上の義務の調査 (ERP システムに売掛金および買掛金のデータを問い合わせる)、金融機関のパフォーマンスの調査 (公開書類や市場データからの期限付きの質問に答える)、ビジネス概要の生成 (マルチソースの企業インテリジェンス レポートの合成) の 3 つのタスク タイプでエージェント システムを評価します。実際の展開条件を反映するために、共通のツール アクセスと遅延制限設定の下で、専用エージェントと汎用エージェント システムを評価します。結果は、汎用エージェント システムは運用上の制約の下で財務ドメインの品質要件を一貫して満たしていないのに対し、Microsoft 365 Copilot 専用の Finance Agent はあらゆる側面で信頼性が高いことを示しています。データセット、ルーブリック、ハーネス、分析コードをオープンソースとしてリリースし、再現可能な比較と他の企業財務環境への適応をサポートします。
原文 (English)
FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance
Recent advances in large language models have accelerated deployment of agentic systems in operational finance. Existing benchmarks emphasize measuring general capabilities, instruction following, or safety, but few directly address the operational finance workflows that agentic systems are now being deployed to automate. Finance professionals require agents to not only provide factually sound and properly grounded information, but also ensure that this information is verifiable and consistently adheres to rules and constraints of the operational finance domain. We introduce FORCE-Bench, which contains 251 expert-annotated queries and evaluates responses using a rubric-based framework calibrated to the requirements of the operational finance domain, across eight dimensions: accuracy, citations, clarity, depth, groundedness, recency, relevance, and structure. FORCE-Bench assesses agentic systems on three task types: financial obligation research (querying ERP systems for accounts receivable and payable data), financial entity performance research (answering time-bound questions from public filings and market data), and business brief generation (synthesising multi-source company intelligence reports). To reflect real deployment conditions, we evaluate our purpose-built agent, as well as the general-purpose agentic systems, under common tool access and latency-bounded settings. Results show that general-purpose agentic systems do not consistently meet finance-domain quality requirements under operational constraints, while the purpose-built Finance Agent for Microsoft 365 Copilot is more reliable across dimensions. We release the dataset, rubrics, harness, and analysis code as open-source to support reproducible comparison and adaptation to other enterprise finance environments.
Chronos の脆弱性: エージェントティック AI における時間的永続性とメモリベースの欺瞞の分類
人工知能におけるステートレス生成モデルからステートフルな自律エージェントへの移行は、長期計画の機能とエンタープライズ ワークフローの自動化を提供すると同時に、新しい形態のセキュリティ脅威であるクロノスの脆弱性の導入を意味するアーキテクチャの進化を表しています。 Chronos の脆弱性は、メモリ インジェクション攻撃 (MINJA) やスリーパー エージェントなどのメモリ ベースの攻撃の脅威を表しており、自律エージェントの内部信念システムが侵害され、最終的な壊滅的なイベントから攻撃ベクトルが効果的に切り離されます。この調査では、World of Workflows ベンチマークのコンテキストにおける永続化ベースの攻撃と Dynamics Blindness の脅威の脅威モデルを形式化し、従来のエンドポイント コンテンツ フィルターが現在のステートフル アーキテクチャには不十分であることを示しています。その結果、この研究では、診断軌道ガードレール (AgentDoG)、形式的時間検証 (Agent-C)、免疫学的メモリ コンセンサス (A-MemGuard)、GPU ベースの信頼できる実行環境 (TEE) およびゼロトラスト メモリ アーキテクチャを介したハードウェア アンカー型信頼などの新たなフレームワークを分類し、多層防御のランドスケープを統合します。
原文 (English)
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
The transition from stateless generative models in artificial intelligence to stateful, autonomous agents represents an architectural evolution that, while providing the capabilities of long-term planning and the automation of enterprise workflows, also represents the introduction of a new form of security threat, the Chronos Vulnerability. The Chronos Vulnerability represents the threat of memory-based attacks, including the Memory Injection Attack (MINJA) and the sleeper agent, in which the internal belief system of the autonomous agent is compromised, effectively decoupling the attack vector from the final catastrophic event. This study formalizes the threat model for persistence-based attacks and the threat of Dynamics Blindness in the context of the World of Workflows benchmark, demonstrating that traditional endpoint content filters are insufficient for the current stateful architecture. Consequently, this study synthesizes a defense-in-depth landscape, categorizing emerging frameworks such as diagnostic trajectory guardrails (AgentDoG), formal temporal verification (Agent-C), immunological memory consensus (A-MemGuard), and hardware-anchored trust via GPU-based Trusted Execution Environments (TEEs) and Zero-Trust memory architectures.
認識論的事前学者による洗練されたポリシー
高度な推論は、再帰的信念モデリングやツリー検索に関連付けられることが多い能動推論の変形です。私たちは、その中心的な計算上の役割はより単純であると主張します。つまり、計画期間内では、将来のアクションが将来の状態と観測に依存することを可能にすることで、アクティブな推論を閉ループで実行します。この閉ループ構造は、認識論的事前変分自由エネルギーの枠組みで表すことができます。認識論的事前分布は能動的推論の目的を提供し、将来の状態とアクションに関する結合事後分布は状態条件付き制御構造を提供します。この分解を、認識的インセンティブを内地平線の閉ループ制御から分離するように設計された確率的ベンチマークである反応性迷路で評価します。この比較には、同じ状態-アクション事後ファミリーを持つ 3 つの変分目標、アクション-状態因数分解アクティブ推論目標、高度な推論、および標準の期待自由エネルギー計画が含まれます。結果は、どちらの成分も単独では十分ではないことを示しています。認識論的要素を持たない方法は情報を求めませんが、将来の行動が将来の状態に依存することを防ぐ方法は、情報を信頼できる目標達成に変えることができません。対照的に、高度な推論とフルジョイント認識事前アクティブ推論はどちらも、認識駆動と閉ループ推論を組み合わせることによって環境を解決します。これらの結果は、高度な推論に関連する利点がツリー検索自体に固有である必要はないことを示しています。これは能動推論の閉ループ形式から生じ、この形式は、事後関数が将来の状態に依存して将来のアクションを維持する場合、認識論的事前変分推論で表すことができます。
原文 (English)
Sophisticated Policies from Epistemic Priors
Sophisticated Inference is a variant of active inference often associated with recursive belief modeling and tree search. We argue that its central computational role is simpler: within a planning horizon, it makes active inference closed-loop by allowing future actions to depend on future states and observations. This closed-loop structure can be represented in the epistemic-prior variational free energy framework. Epistemic priors supply the active-inference objective, while a joint posterior over future states and actions supplies the state-contingent control structure. We evaluate this decomposition in the Reactivity Maze, a stochastic benchmark designed to separate epistemic incentive from inner-horizon closed-loop control. The comparison includes three variational objectives with the same state-action posterior family, an action-state factorized active inference objective, Sophisticated Inference, and standard Expected Free Energy planning. The results show that neither ingredient is sufficient on its own. Methods without an epistemic component do not seek information, while methods that prevent future actions from depending on future states cannot turn information into reliable goal-reaching. By contrast, both Sophisticated Inference and full-joint epistemic-prior active inference solve the environment by combining epistemic drive with closed-loop inference. These results show that the advantage associated with Sophisticated Inference need not be specific to tree search itself. It arises from the closed-loop form of active inference, and this form can be represented in epistemic-prior variational inference when the posterior keeps future actions dependent on future states.
知識中心の自己改善
自己改善 AI システムは通常、プロンプト、ワークフロー、ハーネス、さらにはエージェント自身のコードを最適化することによって、エージェントを改善するオブジェクトとして扱います。このエージェント中心のビューでは、利益が特定のエージェントの設計、タスクの分散、または適応の実行に関連付けられるため、改善の維持にコストがかかり、移転が困難になる可能性があります。私たちは、知識中心の自己改善という補完的なパラダイムを研究しています。このパラダイムでは、エージェントは汎用的で使い捨てのままですが、永続オブジェクトは、エージェントが将来のタスクに活用できる厳選された知識ベースです。私たちは、シンプルなプロトコルを通じてこのアイデアを実用化するために、管理されたケーススタディを実施します。エージェントは 1 つのタスクを試み、タスク レベルおよびタスク間のフォーラムを通じて証拠に基づいた洞察を共有ナレッジ ベースに提供し、その後、知識を蒸留します。自己改善はエージェントではなく知識に含まれるため、改善はより検査しやすく、移転しやすく、移植しやすくなります。このプロトコルは、抽象的な推論、コーディング、および端末ベンチマークにわたって、エージェント中心のベースラインと比較してコストを削減しながら解決率を向上させます。結果として得られる蒸留された知識は、保留されたタスクや LLM ファミリ全体にも転送され、改善が単に LLM または実行固有の動作ではないことを示しています。これらの結果は、自己改善エージェント システムの新しい見方を裏付けています。つまり、進歩は主に厳選された永続的な知識によって推進される可能性があります。コードは https://github.com/recursive-knowledge/KSI で入手できます。
原文 (English)
Knowledge-Centric Self-Improvement
Self-improving AI systems typically treat the agent as the object that improves, by optimizing prompts, workflows, harnesses, or even the agent's own code. This agent-centric view can make improvements expensive to maintain and difficult to transfer, because gains become tied to a particular agent design, task distribution, or adaptation run. We study a complementary paradigm: knowledge-centric self-improvement, in which agents remain generic and disposable while the persistent object is a curated knowledge base that agents can leverage for future tasks. We conduct controlled case studies to operationalize this idea via a simple protocol. Agents attempt one task, then contribute evidence-grounded insights to a shared knowledge base via task-level and cross-task forums, followed by knowledge distillation. Because self-improvement is contained in the knowledge rather than the agent, improvement can be more inspectable, transferable, and portable. Across abstract reasoning, coding, and terminal benchmarks, this protocol improves solve rates while reducing dollar cost relative to agent-centric baselines. The resulting distilled knowledge also transfers to held-out tasks and across LLM families, indicating that the improvement is not merely an LLM- or run-specific behavior. These results support a new view of self-improving agentic systems: progress can be driven primarily by the curated persistent knowledge. Code is available at https://github.com/recursive-knowledge/KSI.
民間航空におけるエッジ インテリジェンス: パラダイム、技術、およびアプリケーション
民間航空は安全性が非常に重要であり、飛行甲板やタワーからランプやメンテナンスに至るまで、その運用によりネットワーク エッジで大量の異種データが生成されます。しかし、大規模な人工知能 (AI) モデルのクラウド中心の展開では、多くの場合、タスクの待ち時間が長くなり、通信が拒否された環境ではオフライン機能が不足し、機密データの一元管理が必要となり、プライバシーと主権のリスクが高まります。エッジ AI は、圧縮、協調推論、分割学習を通じて、認識、予測、意思決定ロジックをデータ生成者の近くに移動し、それにより、遅延、帯域幅、および露出を削減すると同時に、切断時の正常な動作を可能にします。このペーパーでは、民間航空向けに調整されたエッジ インテリジェンスの全体像と共通の理解を提供します。まずエッジ AI の運用動機を明確にし、次にエッジ推論とエッジ学習の最近の技術をレビューします。次に、組織のコンピューティング パラダイムと民間航空環境におけるそれぞれの構成を紹介します。最後に、民間航空におけるエッジ インテリジェンスの新たなアプリケーションと将来の研究動向について説明します。私たちは、洗練されたエッジ ソリューションがクラウド基盤を補完し、民間航空のライフサイクル全体にわたって低遅延、プライバシー保護、復元力のある AI サービスを提供できると主張します。
原文 (English)
Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications
Civil aviation is safety critical and its operations, from flight decks and towers to ramps and maintenance, generate massive, heterogeneous data at the network edge. Yet cloud centric deployment of large Artificial Intelligence (AI) models often produces high task latency, lacks offline capability in communication denied environments, and requires centralizing sensitive data, raising privacy and sovereignty risks. Edge AI moves perception, prediction, and decision logic closer to the data producers via compression, collaborative inference, and split learning, thereby reducing latency, bandwidth, and exposure while enabling graceful operation during disconnections. This paper provides a panoramic view and a common understanding of edge intelligence tailored to civil aviation. We firstly articulate the operational motivations for edge AI, and then review recent techniques for edge inference and edge learning. We then introduce the organizational computing paradigms and the respective configurations in civil aviation environments; finally, we describe the emerging applications and the future research trends of edge intelligence in civil aviation. We argue that a refined edge solution can complement cloud foundations to deliver low latency, privacy preserving, and resilient AI services across the civil aviation lifecycle.
エージェントの認識と生成による電子部品のシンボルとフットプリントのデータベース
豊富で認識可能なコンポーネント ライブラリは、プリント基板 (PCB) の設計と生成の基礎です。従来、エンジニアは手動でシンボルとフットプリントを作成し、PCB 回路図を設計していましたが、これには時間がかかり、エラーが発生しやすくなります。マルチモーダル大規模言語モデル (MLLM) を活用して、電子コンポーネントのシンボルとフットプリントのエージェント認識および生成フローである SFgen を開発します。 SFgen は、シンボル生成で 86% の精度、フットプリント生成で 80% の精度を達成します。 SFgen メソッドを使用して、シンボルとフットプリントのデータベースである SFnet を作成します。現在 1,000 個のコンポーネントがあり、継続的に拡張されており、PCB 設計の自動生成の基礎を築いています。
原文 (English)
Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation
A rich and recognizable component library is the cornerstone of printed circuit board (PCB) design and generation. Traditionally, engineers manually create symbols and footprints and design PCB schematics, which is time-consuming and error-prone. Leveraging multimodal large language models (MLLMs), we develop SFgen, an agentic recognition and generation flow of symbol and footprint for electronic components. SFgen achieves 86% accuracy for symbol generation and 80% accuracy for footprint generation. We use the SFgen method to create SFnet, a database of symbols and footprints. It now has 1000 components and is expanding constantly, which lays the foundation for automatic generation of PCB designs.
マルチモーダルエージェント検索におけるサイレントエラー:診断分類法と複数の裁判官による評価
マルチモーダル エージェント検索システムは、知識集約的な視覚的な質問に答えるために、外部ツールへの依存度が高まっています。ただし、既存の評価は主に最終的な回答の精度に焦点を当てており、検索軌跡の失敗を見逃してしまう可能性があります。この研究では、サイレント障害などの隠れた信頼性の問題を研究します。モダリティのショートカット、ファントムグラウンディング、間違った証拠と正解の事例、過剰検索ロンダリング、クロスモーダル矛盾、来歴幻覚をカバーする 6 つのカテゴリーの分類法を導入します。この分類に基づいて、統一された ReAct スタイルの足場の下で、回答の正しさと証拠の根拠の質の両方を評価する軌跡レベルの診断パイプラインを構築します。 4 つのフロンティア マルチモーダル モデルにわたる MMSearch-Plus 軌道に関する実験では、表面精度が真の軌道レベルの正確さを一貫して過大評価していることが示されています。さらに、クロスジャッジ検証、ブランク画像ストレステスト、およびツールアブレーションを使用して、サイレントエラーは能力に依存し、消えるのではなく変化することが多いことを示します。ホームページ: https://github.com/DingWu1021/silent-failures-multimodal-agentic-search
原文 (English)
Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation
Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing evaluations mainly focus on final-answer accuracy and may miss failures in the search trajectory. In this work, we study such hidden reliability issues as silent failures. We introduce a six-category taxonomy covering modality shortcuts, phantom grounding, wrong-evidence-right-answer cases, over-retrieval laundering, cross-modal contradiction, and provenance hallucination. Based on this taxonomy, we build a trajectory-level diagnostic pipeline that evaluates both answer correctness and evidence-grounding quality under a unified ReAct-style scaffold. Experiments on MMSearch-Plus trajectories across four frontier multimodal models show that surface accuracy consistently overestimates true trajectory-level correctness. We further use cross-judge validation, blank-image stress tests, and tool ablations to show that silent failures are capability-dependent and often shift rather than disappear. Home-page: https://github.com/DingWu1021/silent-failures-multimodal-agentic-search
LLM の好みの調整に対するより良い思考に報いる
LLM の好みの調整は、多様なユーザーの指示にわたって人間の好みに合わせてモデルを最適化することを目的としています。強化学習は、この目標に向けたトレーニング後の主要なアプローチとなっていますが、既存の代理報酬は多くの場合結果レベルであり、主に最終応答を評価する一方、推論の軌道に限定的なガイダンスを提供します。これにより、複数の回答が同様の最終スコアを受け取った場合にクレジットの割り当てが粗くなり、軌跡レベルの好みが不十分なままになる可能性があります。この制限に対処するために、RL ベースの好みの調整に対するプロセス指向の報酬である Thinking Checklist Reward (TCR) を提案します。 TCR は、好みのペアをサンプル固有の思考チェックリストに変換し、それらを使用して、生成された推論トレースが好みに暗示される考慮事項に対処しているかどうかを評価します。結果レベルの監督との重複を減らすために、TCR はさらに指数移動平均 (EMA) 残差定式化を導入し、結果の報酬から予測可能な範囲を超えた補完的な思考の余剰を分離します。 3 つのモデルファミリーの 5 つのモデルでの実験では、TCR がさまざまなベンチマークにわたってアライメント性能を一貫して向上させ、アブレーションにより EMA ベースの残留配合とサンプル固有のチェックリストの監視の重要性がさらに検証されたことが示されています。
原文 (English)
Rewarding Better Thinking for LLM Preference Alignment
LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome-level, mainly evaluating the final response while providing limited guidance for the reasoning trajectory. This can make credit assignment coarse when multiple responses receive similar final scores, leaving trajectory-level preferences under-specified. To address this limitation, we propose Thinking Checklist Reward (TCR), a process-oriented reward for RL-based preference alignment. TCR converts preference pairs into sample-specific thinking checklists and uses them to evaluate whether the generated reasoning trace addresses the preference-implied considerations. To reduce overlap with outcome-level supervision, TCR further introduces an exponential moving average (EMA) residual formulation to isolate a complementary thinking surplus beyond what is predictable from the outcome reward. Experiments on five models from three model families show that TCR consistently improves alignment performance across diverse benchmarks, with ablations further validating the importance of EMA-based residual formulation and sample-specific checklist supervision.
エージェントを知る: AI エージェントの偵察主導型侵入テスト
従来の侵入テストでは、各ステップで偵察を使用して、目に見えない弱点を明らかにし、より強力な攻撃を構築し、目標を前進させます。私たちは、AI エージェントにも同様の対応が必要であると主張します。私たちは、プロセスをモデル化し、エージェントが抽出しようとしている知識資産、つまり、エージェントが何であるか、どのように使用されるか、敵対者が間接プロンプト インジェクション攻撃に利用できるようにするエージェントの弱点を特定することによって、エージェントの偵察を形式化します。これらの洞察は、エージェントを調査し、ターゲット プロファイルを構築し、それらのプロファイルを使用してより強力な攻撃を作成することにより、ブラック ボックスの偵察主導の侵入テストを自動化するフレームワークである Know Your Agent (KYA) でインスタンス化されます。私たちは、エージェント セキュリティ ベンチマークと現実世界のコーディング エージェントで KYA を評価し、KYA、そのベンチマーク、再現性を確保するためのベースライン実装をリリースします。
原文 (English)
Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
Traditional pentesting uses reconnaissance at each step to uncover unseen weaknesses, build stronger attacks, and advance the objective; we argue that AI agents require the same treatment. We formalize agent reconnaissance by modeling the process and identifying the knowledge assets it seeks to extract: what they are, how they are used, and which agent weaknesses they exploit to give adversaries leverage in indirect prompt injection attacks. We instantiate these insights in Know Your Agent (KYA), a framework that automates black-box, reconnaissance-driven pentesting by probing agents, building target profiles, and using those profiles to craft stronger attacks. We evaluate KYA on agent-security benchmarks and a real-world coding agent, and release KYA, its benchmarks, and baseline implementations for reproducibility.
DocOps: 複雑なドキュメント操作における自律エージェントの検証可能なベンチマーク
自律エージェントが急速に進化するにつれ、汎用の AI アシスタントを実現し、複雑なワークスペースのワークフローを自動化するために、ユビキタスなデジタル ドキュメントを確実に操作するエージェントの機能が重要になっています。このペーパーでは、現実世界の実践からインスピレーションを得た文書操作を原子的な次元に分解し、ワークフローの複雑さをエスカレートさせる階層的分類法に支えられた決定論的に検証可能な評価フレームワークである DocOps を紹介します。 DocOps に基づいて、さまざまなエージェント ハーネスにわたる代表的なクローズド ソース モデルとオープン ソース モデルを体系的に評価し、高度に結合された長距離タスクを処理する場合、最も高度なフロンティア構成であっても依然として大きな制限があることを明らかにしました。さらに、既存のエージェントの操作動作を詳細に分析した結果、長期的な状態追跡の崩壊、浅い意味検証、構造メタデータの破壊的編集という 3 つの主要な障害モードが明らかになりました。最終的に、私たちの研究は、グローバルな文書の一貫性を維持する際のエージェントの能力の限界を明らかにし、複雑なデジタル エコシステム向けの堅牢で非破壊的なエージェントの将来の設計に光を当てます。
原文 (English)
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematically evaluate representative closed- and open-source models across various agentic harnesses, revealing that even the most advanced frontier configurations still exhibit profound limitations when handling highly coupled, long-range tasks. Furthermore, a fine-grained analysis of existing agents' manipulation behaviors uncovers 3 key failure modes: long-term state tracking collapse, shallow semantic verification, and destructive editing of structural metadata. Ultimately, our work exposes the capability boundaries of agents in maintaining global document consistency, shedding light on the future design of robust, non-destructive agents for complex digital ecosystems.
JANUS: 長期的なエージェントの安全性に対する潜在的なリスクを予測する
エージェントの安全性は、コンテンツの管理から、ツールを使用するエージェントが行動する前に操作上の失敗を防止することに移行しています。私たちは、長期的なエージェントの安全のための先見性を重視したフレームワークである Janus を提案します。これは、部分的な軌道からの遅延リスクを予測するように警備員を訓練します。 Janus は、マルチエージェント シミュレーションを通じて多様なエージェントの軌跡を合成し、安全性に関連する将来を予測する予測タスクと、観測されたプレフィックスと予測される将来の両方から安全性を決定する判断タスクという 2 つの結合タスクで共有ポリシーを学習します。 2 つのタスクは CoAA-RL を使用して共同で最適化され、下流の安全性判断への有用性によって予測に報酬が与えられます。結果として得られるガード モデル Vanguard は、安全でないアクションを実行前にブロックします。 4 つのエージェント安全性ベンチマーク全体で、Vanguard はベースライン ガードよりも平均保護を 15.9 パーセント ポイント改善し、良性のタスクの完了を 5.1 パーセント ポイント向上させました。
原文 (English)
JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate delayed risks from partial trajectories. Janus synthesizes diverse agent trajectories via multi-agent simulation and learns a shared policy with two coupled tasks: an anticipation task that forecasts safety-relevant futures and an adjudication task that decides safety from both the observed prefix and anticipated future. The two tasks are jointly optimized with CoAA-RL, which rewards forecasts by their utility for downstream safety judgment. The resulting guard model, Vanguard, blocks unsafe actions before execution. Across four agent-safety benchmarks, Vanguard improves average protection by 15.9 percentage points over baseline guards while increasing benign task completion by 5.1 percentage points.
リスクの下での長期にわたる逐次的な意思決定
私たちは、総収益の分布に順位依存関数を適用する \emph{ルートベース} (断固とした) リスク目標に基づいた有限水平 MDP 計画を研究します。このような目標は収益分布が非線形であり、一般にベルマン最適性を破るため、シナリオ ツリーの列挙による直接の最適化は困難です。 \textbf{ERQDP} を提案します。これは、正確な DP (動的計画法) を介してランク分位サロゲートを解き、離散化リターン グリッド (明示的な丸め限界を使用) 上のリターン確率質量関数 (PMF) を介して DP によって候補ポリシーを正確に評価し、目標目標の明示的な上限と下限のギャップ (証明書) を報告するいつでもループでサロゲートを洗練する、列挙フリーおよびサンプリング フリーの方法です。離散化予算。テストされたベンチマーク全体で、ERQDP は認定されたソリューションまたは明示的な残差ギャップを返し、実行時間の大幅な向上による高速なリスク パラメーターのスイープを可能にし、リスク回避行動とリスク探索行動の両方をサポートします。
原文 (English)
Long-Term Sequential Decision Making under Risk
We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break Bellman optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank--quantile surrogate via exact DP (Dynamic Programming), evaluates candidate policies exactly by DP over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper--lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.
MOF-Sleuth: 説明可能なきめ細かい MOF CIF 監査のためのツールベースの報酬調整
大規模な金属有機フレームワーク (MOF) データベースは、結晶情報ファイル (CIF) を通じてシミュレーション、スクリーニング、機械学習をサポートします。これらの入力における微妙な化学的および構造的エラーは、下流の結果を損ない、手動検査を妨げる可能性があります。計算化学における LLM の進歩は、予測スクリーニングを超えて、証拠に基づいた説明によるきめの細かい診断への道を提供します。ただし、2 つの課題が残っています。(i) 限定されたきめ細かいアトリビューション: MOF 固有のバリデーターと機械学習モデルは検出をスケールしますが、証拠に基づいた説明ではなく、固定チェック、準備スコア、または大まかなラベルを提供します。 (ii) 信頼性の低い CIF 推論: 化学的証拠は原子サイトの記録全体にわたって暗黙的に存在し、幾何学的、接続性、占有、および電荷の計算を必要とするため、LLM の直接監査はコストがかかり、信頼性が低くなります。どちらも、化学的証拠と言語モデルの説明の間の弱い結合に由来しています。 MOF-Sleuth は、決定論的なフォレンジック ラボと Sleuth 推論エンジンの 2 つのモジュールを備えた強化ガイド付き CIF 監査エージェントです。ラボは、構成、形状、接続性、占有、調整、および電荷の証拠を導き出し、スルースはこの証拠を使用して、証拠に基づいた説明、エラーの種類、および二者択一の決定を作成します。報酬誘導型強化学習 (RL) は、ツールの測定を化学的な説明レベルの監視に変え、最終的な答えだけでなく、引用された化学的証拠や証拠に裏付けられた診断にも報酬を与えます。正しい診断が事実に基づいた関連する CIF 由来の証拠によって説明されるかどうかを評価する指標である、Chemically Ground Diagnosis (Chem-GD) を紹介します。 MOF-Sleuth は、4 つのベンチマークにわたって、LLM ベースのアプローチと MOF 固有の機械学習手法の間で最先端のパフォーマンスを確立し、検出、属性、根拠のある説明の品質の向上を実証しています。
原文 (English)
MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs). Subtle chemical and structural errors in these inputs can compromise downstream results and hinder manual inspection. LLM advances in computational chemistry offer paths beyond predictive screening toward fine-grained diagnosis with evidence-grounded explanations. However, two challenges remain: (i) limited fine-grained attribution: MOF-specific validators and machine-learning models scale detection but provide fixed checks, readiness scores, or coarse labels rather than evidence-grounded explanations; and (ii) unreliable CIF reasoning: direct LLM auditing is costly and unreliable because chemical evidence is implicit across atom-site records and requires geometric, connectivity, occupancy, and charge calculations. Both stem from weak coupling between chemical evidence and language-model explanation. We introduce MOF-Sleuth, a reinforcement-guided CIF auditing agent with two modules: a deterministic Forensic Lab and a Sleuth reasoning engine. The Lab derives composition, geometry, connectivity, occupancy, coordination, and charge evidence, and Sleuth uses this evidence to produce an evidence-grounded explanation, error types, and a binary decision. Reward-guided reinforcement learning (RL) turns tool measurements into chemical explanation-level supervision, rewarding not only the final answer but also cited chemical evidence and evidence-supported diagnoses. We introduce Chemically Grounded Diagnosis (Chem-GD), a metric that assesses whether a correct diagnosis is explained by factual, relevant CIF-derived evidence. Across four benchmarks, MOF-Sleuth establishes state-of-the-art performance among LLM-based approaches and MOF-specific machine-learning methods, demonstrating gains in detection, attribution, and grounded explanation quality.
SenWorld: コンテキストリッチな評価データを生成するためのデジタル ツイン シミュレーション
スマートフォンのパーソナル アシスタントは長期にわたる個人データを推論しますが、その評価には正解がわかっているコンテキストに富んだ評価データが必要であり、実際のデバイスのトレースはプライバシーに敏感すぎて共有できません。この課題に対処するために、構築によって固定されたグラウンド トゥルースを使用してそのようなデータを生成する、物理的に接地され、決定論的でイベント ソースのデジタル ツイン シミュレーションである SenWorld を紹介します。 SenWorld では、ペルソナは実際の地図、天気、休日、ネットワーク データから構築された世界で 1 日を過ごします。観測可能なすべての信号はシステム全体のスナップショットにアーカイブされます。また、各評価ケースは、事後注釈や大規模言語モデル (LLM) ジャッジではなく、既存のレコードへのポインターによってラベル付けされます。この手法を北京の 16 人のペルソナで評価しました。生成されたデータは、カテゴリ分布 (ジェンセンとシャノンの相違 (JSD) 0.070) および通信記録の 1 日のリズム (JSD 0.1 未満) において、保持されている実際のユーザーのベンチマークと厳密に一致していますが、生成された記録は実際の記録よりも短いままです。スクリプトによる対話がなければ、ペルソナは完全に往復する対話サブグラフと差別化された行動レパートリーを形成します。 717 件の評価ケースに投影された生成データでは、実稼働スマートフォン アシスタントの 78 件の障害が明らかになり、通話とショート メッセージ サービス (SMS) の記録に集中し、連絡先、スケジュール、アラームは決して失敗しませんでした。スナップショット ポインタは、LLM 判定者が関与せずに、各失敗をアシスタント側の取得エラーとして確認します。全体として、SenWorld は、ラベルが構築によって固定されている評価データへの、プライバシーに安全で再現可能で配布がチェックされたパスを提供します。
原文 (English)
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, we present SenWorld, a physically grounded, deterministic, event-sourced digital-twin simulation that generates such data with ground truth fixed by construction. In SenWorld, personas live through a full day in a world built from real map, weather, holiday, and network data; every observable signal is archived in full-system snapshots; and each evaluation case is labeled by a pointer to an existing record rather than by post-hoc annotation or a large language model (LLM) judge. We evaluate this method with 16 personas in Beijing. The generated data closely matches the held-out real-user benchmark in category distribution (Jensen--Shannon divergence (JSD) 0.070) and in the daily rhythm of communication records (JSD below 0.1), though generated records remain shorter than real ones. Without scripted interaction, personas form a fully reciprocated dialogue subgraph and differentiated behavioral repertoires. Projected into 717 evaluation cases, the generated data exposes 78 failures in a production smartphone assistant, concentrating on call and Short Message Service (SMS) records while contacts, schedules, and alarms never fail. The snapshot pointer confirms each failure as an assistant-side retrieval error, with no LLM judge involved. Overall, SenWorld offers a privacy-safe, reproducible, and distribution-checked path to evaluation data whose labels are fixed by construction.
EvoThink: 自己刈り込みと「なるほど」の好みの最適化による大規模推論モデルにおける思考の進化
大規模推論モデル (LRM) では、冗長な検証手順により考えすぎが生じることがよくあります。考えすぎを軽減するための既存のアプローチ(高速思考と低速思考の切り替えや推論軌道の圧縮など)では、LRM の推論プロセス内で有益なステップと冗長なステップを細かく区別できず、そのため効率を追求する推論能力が損なわれる可能性があります。推論の効率と能力を同時に向上させるために、冗長な検証を削減し、新しい推論パスの探索を促進するフレームワークである EvoThink を提案します。 EvoThink は 2 つの重要なコンポーネントで構成されています。自己剪定トレーニング (SPT) は、冗長な推論ステップを反復的に剪定し、簡潔な軌道で自己トレーニングする教師なしメソッドです。そして、遺伝的アルゴリズムにヒントを得た Aha-Moment Preference Optimization (AMPO) は、失敗した貴重な推論の試みを特定し、間違ったものから正しいものへの Aha-Moment データを合成し、この推論パターンを内部化するようにモデルを最適化します。数学的推論とコード生成ベンチマークにわたる広範な評価により、EvoThink が推論時のトークン使用量を大幅に削減するだけでなく、LRM の推論能力も向上することが実証されました。
原文 (English)
EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overthinking, such as fast-slow thinking switching and reasoning trajectory compression, fail to make a fine-grained distinction between beneficial and redundant steps within the LRM's reasoning process, and may thus impair reasoning capability in their pursuit of efficiency. To simultaneously improve reasoning efficiency and capability, we propose EvoThink, a framework that reduces redundant verification and encourages the exploration of new reasoning paths. EvoThink comprises two key components: Self-Pruning Training (SPT), an unsupervised method that iteratively prunes redundant reasoning steps and self-trains on the concise trajectories; and Aha-Moment Preference Optimization (AMPO), which, inspired by genetic algorithms, identifies valuable failed reasoning attempts, synthesizes from-wrong-to-right aha-moment data, and optimizes the model to internalize this reasoning pattern. Extensive evaluations across mathematical reasoning and code generation benchmarks demonstrate that EvoThink not only substantially reduces inference-time token usage but also improves the reasoning capability of LRMs.
巨大な海馬: 構造的なモノカルチャーからシステムのシステムへ
AI 研究者は、最先端のモデルを大規模に繰り返される 1 つのもの、つまりテキスト、ピクセル、または音声に対して同じように配線されたトランスフォーマーであると説明します。神経科学者は、大脳皮質をモザイクであると説明しています。空間コード化のための視覚皮質の緻密な第4層、時間統合のための運動皮質の厚い第5/6層、さまざまな構造によって解決されるさまざまな仕事です。この論文は、このギャップは文体上のエラーではなく構造上のエラーであり、測定可能であると主張しています。ブロードマンから単細胞パッチシーケンスに至るまでの細胞構築の 1 世紀は、異なる認知機能が 1 つのテンプレートを再スケールすることによってではなく、質的に異なる構造によって実装されることを示しています。畳み込みニューラル ネットワークは、この分野独自の証明です。局所的な受容野と階層の深さが、これを事前に直接エンコードし、その後のアーキテクチャが必要とするよりもはるかに少ないデータで強力な画像認識に到達しました。この論文は、この教訓がどのように捨てられたかを追跡しています。「ハードウェア宝くじ」により、原則的な選択ではなく、トランスフォーマーが最も抵抗の少ない経路になりました。また、多様性としてよく引用される専門家の混合は、実際には同一の専門家間でパラメーターを分割しています。機能主義的分析によると、トランスフォーマーは汎用の皮質ではなく、海馬形成の機能的類似物として最もよく理解されていることが示されている。これは皮質を巨大なブローカ野として扱うのと同じ間違いである。ただし、この分野が巨大な海馬で標準化され、聴覚、実行ゲート、作業記憶など決してそのために構築されたことのないタスクに適用される点が異なる。この論文は、代替モジュールであるヘテロジニアス トポロジカル ネットワーク、つまり個別のモジュールが計算要求に必要な誘導バイアスを維持し、標準化されたインターフェイスを通じて通信するシステム オブ システムズで終わります。これは認知科学ではなく、AI アーキテクトのための設計規律です。トレーニングされたモデルの動作からアーキテクチャをリバース エンジニアリングするのではなく、構造的証拠を設計入力として使用して、トレーニング前にモジュール性を指定します。
原文 (English)
The Giant Hippocampus: From Structural Monoculture to a System of Systems
AI researchers describe state-of-the-art models as one thing repeated at scale: the Transformer, wired identically for text, pixels, or speech. Neuroscientists describe the cortex as a mosaic - dense Layer 4 in visual cortex for spatial encoding, thick Layers 5/6 in motion cortex for temporal integration - different jobs solved by different structures. This paper argues the gap is a structural error, not a stylistic one, and is measurable. A century of cytoarchitecture, from Brodmann to single-cell Patch-seq, shows distinct cognitive functions are implemented by qualitatively different structures, not by rescaling one template. The convolutional neural network is the field's own proof: local receptive fields and hierarchical depth encoded this prior directly, reaching strong image recognition on far less data than later architectures needed. The paper traces how this lesson was discarded: the "Hardware Lottery" made the Transformer the path of least resistance, not the principled choice, and Mixture-of-Experts, often cited as diversity, in fact partitions parameters among identical experts. A functionalist analysis shows the Transformer is best understood as a functional analog of the hippocampal formation, not a general-purpose cortex - the same mistake as treating cortex as one giant Broca's area, except the field has now standardized on a giant hippocampus, applied to tasks it was never built for: audition, executive gating, working memory. The paper closes with an alternative: a Heterogeneous Topological Network, a System of Systems in which distinct modules keep the inductive bias their computation demands and communicate through standardized interfaces. This is a design discipline for AI architects, not cognitive science: specify modularity before training, using structural evidence as a design input rather than reverse-engineering architecture from a trained model's behavior.
記憶からの調整: 動的製造におけるマルチエージェント適応のためのグラフ構造化エクスペリエンスの再利用
動的な製造環境では、機械の故障、緊急のジョブの到着、処理時間の変動などの頻繁な運用障害下で、マルチエージェント システムが効果的に調整する必要があります。既存のマルチエージェント強化学習アプローチは、各外乱エピソードを独立して扱い、将来の適応を促進する可能性のある貴重な調整経験を捨てています。この論文では、動的製造におけるマルチエージェント調整のためのグラフ構造経験記憶 (GSEM) フレームワークを提案します。このフレームワークは、過去の調整エピソードを、タスクの依存関係、マシンの状態、エージェント間のコラボレーション パターンをキャプチャする異種リレーショナル グラフとしてエンコードします。新しい混乱が発生すると、グラフ ニューラル ネットワーク ベースの検索メカニズムが構造的に類似した過去のエピソードを特定し、ゼロから学習するのではなく、経験に基づいた政策適応を可能にします。 3 つの外乱タイプを使用した動的で柔軟なジョブショップ スケジューリング ベンチマークの実験では、GSEM が最も強力なメモリ拡張ベースラインと比較してメイクスパンを 4.1% ~ 10.0%、適応時間を 33% ~ 38% 短縮し、外乱周波数が高くなると利点が増大することが示されました。アブレーション研究と相互擾乱伝達実験は、グラフ構造化エンコーディングと類似性に基づく検索の必要性をさらに検証し、学習された調整パターンの相互擾乱の一般化可能性を実証します。
原文 (English)
Coordinating from Memory: Graph-Structured Experience Reuse for Multi-Agent Adaptation in Dynamic Manufacturing
Dynamic manufacturing environments require multi-agent systems to coordinate effectively under frequent operational disturbances such as machine failures, urgent job arrivals, and processing time variations. Existing multi-agent reinforcement learning approaches treat each disturbance episode independently, discarding valuable coordination experience that could accelerate future adaptation. In this paper, we propose a Graph-Structured Experiential Memory (GSEM) framework for multi-agent coordination in dynamic manufacturing. The framework encodes historical coordination episodes as heterogeneous relational graphs that capture task dependencies, machine states, and inter-agent collaboration patterns. When a new disturbance occurs, a graph neural network-based retrieval mechanism identifies structurally similar past episodes, enabling experience-guided policy adaptation rather than learning from scratch. Experiments on dynamic flexible job-shop scheduling benchmarks with three disturbance types show that GSEM reduces makespan by 4.1%-10.0% and adaptation time by 33%-38% compared to the strongest memory-augmented baseline, with the advantage increasing under higher disturbance frequency. Ablation studies and cross-disturbance transfer experiments further validate the necessity of graph-structured encoding and similarity-based retrieval and demonstrate the cross-disturbance generalizability of learned coordination patterns.
CLARK: ナレッジ グラフ上の適応推論のための閉ループ学習
機械学習モデルは、データから統計パターンを抽出することで分類タスクを自動化するために広く使用されています。ただし、データ分布が変化するとパフォーマンスが低下するため、不確実で進化する情報の処理には不向きになります。さらに、事前の知識を統合するためのサポートは限定的です。これらの制限に対処するために、マルコフ論理ネットワーク (LP$^{\text{MLN}}$) 形式による論理プログラムの下でナレッジ グラフ、記号ルール マイニング、確率論的推論を統合するフレームワークである CLARK (ナレッジ グラフ上の適応推論のための閉ループ学習) を紹介します。 CACTUS 由来の KG から開始して、CLARK はグラフ構造を LP$^{\text{MLN}}$ プログラムに変換し、記号学習者によって提案された候補ルールで反復的にプログラムを強化します。これらのルールは確率的重み学習を通じて調整され、不確実性の下での推論と基礎となるグラフ構造の改良が可能になります。私たちは 2 つの医療データセットで CLARK を評価し、ルールの品質と下流の分類パフォーマンスの両方を分析します。結果は、CLARK が分類パフォーマンスの向上とより一般化された推論につながることを示しています。全体として、CLARK は、適応的で解釈可能な知識駆動型の分類モデルを構築するための原則に基づいたアプローチを提供します。
原文 (English)
CLARK: Closed-loop Learning for Adaptive Reasoning over Knowledge Graphs
Machine Learning models are widely used for automating classification tasks by extracting statistical patterns from data. However, their performance deteriorates if the data distribution changes, making them ill-suited to handle uncertain and evolving information. Moreover, they provide limited support for integrating prior knowledge. To address these limitations, we present CLARK (Closed-loop Learning for Adaptive Reasoning over Knowledge Graphs), a framework that integrates knowledge graphs, symbolic rule mining, and probabilistic reasoning under the Logic Programs with Markov Logic Networks (LP$^{\text{MLN}}$) formalism. Starting from CACTUS-derived KGs, CLARK translates graph structure into an LP$^{\text{MLN}}$ program and iteratively enriches it with candidate rules proposed by symbolic learners. These rules are calibrated through probabilistic weight learning, enabling reasoning under uncertainty and refinement of the underlying graph structure. We evaluate CLARK on two medical datasets, analysing both rule quality and downstream classification performance. Results demonstrate that CLARK leads to improved classification performance and more generalisable inference. Overall, CLARK provides a principled approach to constructing adaptive, interpretable, knowledge-driven models for classification.
マイクロサービス システムにおけるリスクを抑制した介入決定としての安全な修復
現代の IT 運用 (IT-Ops) では、誤った修復のコストが、何も行わなかった場合のコストを超えることがよくあります。しかし、既存の自動修復システムは、介入が正当であるかどうかを決定するのではなく、アクションを生成するように設計されており、安全性は手動承認によって強制される後付けのままになっています。この論文は、このギャップを埋めるために 3 つの貢献を行っています。(i) 安全な修復をリスク制限された介入決定問題として再定式化し、それを制約付きマルコフ決定プロセス (CMDP) としてキャストします。このプロセスでは、エージェントは、制限された誤った修復率 (FRR) に従って修復の成功を最大化します。 (ii) 爆発半径、可逆性、認識論的不確実性からなる 3 次元のリスク分解を導入し、操作者に解釈可能な操作ごとの安全インターフェースを提供します。 (iii) コンテキスト適応型ヒューマンインザループ (HITL) ゲートを設計します。これは、バイナリ フェールセーフから、オンコールの負荷とビジネスの重要性に応答する帯域幅を認識した制御層にエスカレーションを変換します。完全なポリシーは過去のインシデント ログからオフラインで学習されるため、予想される FRR を明示的に制御できます。カオス メッシュ フォールト インジェクションと RCAEval に合わせたフォールト分類法を使用した Train Ticket マイクロサービス ベンチマークの実験では、私たちのフレームワークが FRR を 39% 削減しながら、強力なランブック ベースラインよりも修復成功率を 2.5 ポイント向上させ、固定しきい値バリアントと比較してオンコール エスカレーション負荷を 17% 削減することを示しています。
原文 (English)
Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems
In modern IT operations (IT-Ops), the cost of an incorrect repair often exceeds the cost of no action at all. Yet existing automated remediation systems are designed to generate actions rather than to decide whether intervention is warranted, leaving safety as an afterthought enforced by manual approval. This paper makes three contributions to close this gap: (i) we reformulate safe remediation as a risk-constrained intervention decision problem and cast it as a Constrained Markov Decision Process (CMDP), in which the agent maximizes repair success subject to a bounded false remediation rate (FRR); (ii) we introduce a three-dimensional risk decomposition comprising blast radius, reversibility, and epistemic uncertainty, providing operators with an interpretable per-action safety interface; and (iii) we design a context-adaptive human-in-the-loop (HITL) gate that turns escalation from a binary failsafe into a bandwidth-aware control layer responsive to on-call load and business criticality. The full policy is learned offline from historical incident logs, enabling explicit control of the expected FRR. Experiments on the Train Ticket microservice benchmark with Chaos Mesh fault injection and an RCAEval-aligned fault taxonomy show that our framework reduces FRR by 39% while improving repair success by 2.5 points over a strong runbook baseline, and reduces on-call escalation load by 17% relative to a fixed-threshold variant.
EvoDRC: 自動化された DRC 違反修復のための自己進化型エージェント フレームワーク
デザイン ルール チェック (DRC) クロージャーは、依然として先進ノードの物理設計における大きなボトルネックです。詳細ルーターはルールを認識しますが、設計ルール違反 (DRV) が残ると、手動によるエンジニアリング変更オーダーの繰り返しが必要になることがよくあります。修復では複雑な幾何学的相互作用を考慮し、回路の接続を維持し、新たな違反の導入を回避する必要があるため、このプロセスの自動化は困難です。エージェントによるブロックレベルの DRC 修復のためのスキル進化フレームワークである EvoDRC を紹介します。 EvoDRC は、無関係なリファレンス設計から抽出した知識を使用してレイヤー固有の修復スキルを初期化し、ターゲット設計から収集した追跡可能な修復経験を使用してこれらのスキルを継続的に進化させます。 EvoDRC は、レイアウトを境界のある修復領域に分解し、LLM 修復エージェントを各領域に割り当てます。ローカル DRC 分析、接続チェック、および影響プレビュー ツールは、提案された変更に関するフィードバックを提供します。修復操作とその結果として生じる DRV の変更はナレッジ データベースに保存され、修復スキルを進化させるために使用されます。 DAC26 DRC ベンチマークからの 7 つのブロック レベルのデザインの実験では、EvoDRC が報告されたベースラインと比較して全体で 73.5% の削減を達成することが示されています。
原文 (English)
EvoDRC: A Self-Evolving Agentic Framework for Automated DRC Violation Repair
Design rule check (DRC) closure remains a major bottleneck in advanced-node physical design. Although detailed routers are rule-aware, residual design rule violations (DRVs) often require manual engineering change order iterations. Automating this process is challenging because repairs must account for complex geometric interactions, preserve circuit connectivity, and avoid introducing new violations. We present EvoDRC, a skill-evolution framework for agentic block-level DRC repair. EvoDRC initializes layer-specific repair skills using knowledge distilled from an unrelated reference design and continuously evolves these skills using traceable repair experience collected from the target design. EvoDRC decomposes the layout into bounded repair regions and assigns an LLM repair agent to each region. Local DRC analysis, connectivity-checking, and impact-preview tools provide feedback on proposed modifications. Repair operations and their resulting DRV changes are stored in a knowledge database and used to evolve the repair skills. Experiments on seven block-level designs from the DAC26 DRC Benchmark show that EvoDRC achieves a 73.5\% overall reduction compared to the reported baseline.
制約プログラミングのためのグローバル差分制約の伝播
$x - y \leq d$ の形式の差分制約は、最短経路との関連性があるため、満足と含意のための効率的なアルゴリズムを使用してよく研究されています。ただし、有限領域伝播アルゴリズムは通常、これらのアルゴリズムを利用せず、各差分制約を別個の伝播器として扱います。伝播は解決の完全性を保証しますが、不必要に遅くなる可能性があります。この論文では、差分制約をすべて同時に処理する、差分制約用の (境界が一貫した) グローバル プロパゲータを構築する方法について説明します。 SAT モジュロ理論ソルバーには、以前から差分制約の理論ソルバーが含まれていました。差分制約の理論ソルバーはグローバル差分制約プロパゲーターの基礎を提供しますが、プロパゲーターの要件がどのようにまったく異なるかを示します。重要なことは、遅延節生成ソルバー内で使用するために、グローバル差分制約プロパゲーターによる伝播を説明する方法を示すことです。差分制約をグローバルに扱うことで標準的な伝播アプローチを大幅に改善できることを示す実験を行います。
原文 (English)
Global Difference Constraint Propagation for Constraint Programming
Difference constraints of the form $x - y \leq d$ are well studied, with efficient algorithms for satisfaction and implication, because of their connection to shortest paths. Finite domain propagation algorithms, however, typically do not make use of these algorithms, and treat each difference constraint as a separate propagator. Propagation does guarantee completeness of solving, but can be needlessly slow. In this paper we describe how to build a (bounds consistent) global propagator for difference constraints that treats them all simultaneously. SAT modulo theory solvers have included theory solvers for difference constraints for some time. While a theory solver for difference constraints gives the basis of a global difference constraint propagator, we show how the requirements on the propagator are quite different. Crucially, we show how to explain propagations by a global difference constraint propagator, in order to use it within a lazy clause generation solver. We give experiments showing that treating difference constraints globally can substantially improve on the standard propagation approach.
オープンウェイト言語モデルにおける材料科学メカニズムの表現の読み取りと操作
大規模な言語モデルは科学的な質問に答えることができますが、正しい出力では、モデルが支配的な物理学を表現しているのか、使用しているのかがわかりません。ここでは、オープンウェイト google/gemma-4-E4B-it モデルにおける材料科学メカニズムの情報には、実験的に分離可能な 3 つの形式があることを示します。概念は個々の隠れ状態で読み取り可能であり、構成的配向は状態間の制御された変換によって運ばれ、選択された内部表現は工学的答えを因果的に制御します。私たちは、一致した直接語彙とヤコビアン語彙の読み出し、オプションなしの状態幾何学、60 の法則の反事実ベンチマーク、および因果的介入を組み合わせます。 50 個の資料の説明では、3 つの独立してフィットしたヤコビアン レンズによって概念ランクが再現され、両方の読み出しからのターゲットフリーの単語セットにより、10 個の機構ファミリーのうち 9 個の盲検識別が可能になりました。別の 72 プロンプト ベンチマークでは、メカニズム固有の隠れ状態近傍が生成されましたが、正確なグラフ監査により、この見かけの物理的組織が数値比較によって同様に説明されることが示されました。したがって、我々は、物理的入力の方向のみが反転された、その他の点では同一のプロンプトを比較し、結果として得られる隠れ状態の動きが供給された構成法則に従うかどうかを尋ねました。これらの状態変換は、60 の凍結された関係にわたって直接的、物理的に中立な、および逆法則を命令し、40 の方向法則のうち 39 を正しく方向付けましたが、語彙制御はほぼ偶然でした。双方向介入は、12 の一致するケースすべてにわたって回答確率を物理的に適切な結果に近づけたり遠ざけたりする一方で、反事実状態パッチはメカニズムや回答形式全体で反対の決定シグナルを伝達しました。したがって、物理的関係は、絶対状態だけの場合よりも、制御された状態変化の方がより顕著に現れます。
原文 (English)
Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.
PRO-LONG: プログラムによる記憶により長期的な推論が可能に
長期的なタスクには持続的な認識、推論、探索が必要であり、大規模言語モデル (LLM) エージェントにとっては永続的な課題です。このギャップは、特にモデルがそのまま評価される場合、ARC-AGI-3 などの継続学習ベンチマークでのパフォーマンスの制限に反映されます。このギャップを埋めるためにさまざまなエージェントのハーネスが提案されており、それぞれが長い観察シーケンスを処理するための戦略、つまり環境からどのような情報を保存し、それをモデルのコンテキストにどのようにロードするかという戦略に取り組んでおり、この選択は特に重要であると私たちが主張しています。コンテキスト管理の既存の方法は、より多くの情報を保存すると関連する詳細の取得が困難になるため、重大なトレードオフに直面しています。私たちは、長期的な探索的設定における LLM エージェント向けのプログラム メモリを中心に構築された最小限のコンテキスト管理フレームワークである PRO-LONG を提案します。 PRO-LONG は、完全で構造化されたインタラクション ログを保持し、コーディング エージェントの最近の進歩を利用してこの履歴を効率的に検索することで、トレードオフに対処します。完全な ARC-AGI-3 パブリック ゲーム セットでは、PRO-LONG は、フロンティア モデル全体でベース コーディング エージェントよりも平均 18.0 パーセント向上し、最先端の特殊ハーネスと同等またはそれを上回り (最大 76.1% パス @1)、使用するトークンの数は 4.2 ~ 5.8 分の 1 です。 Fable 5 では、PRO-LONG は合計コスト $1,750 で 97.4% の最高@2 を達成します。関連するコードとログは https://github.com/alexisfox7/PRO-LONG で入手できます。
原文 (English)
PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks such as ARC-AGI-3, especially when models are evaluated out of the box. Various agent harnesses have been proposed to close this gap, and each commits to a strategy for handling long sequences of observations, i.e., what information to save from the environment and how to load it into model context, a choice we argue is particularly consequential. Existing methods for context management face a significant tradeoff, as preserving more information makes retrieving relevant details less tractable. We propose PRO-LONG, a minimal context management framework built around programmatic memory for LLM agents in long-horizon, exploratory settings. PRO-LONG addresses the tradeoff by keeping a complete, structured interaction log and capitalizing on recent progress in coding agents to search this history efficiently. On the full ARC-AGI-3 public game set, PRO-LONG improves over a base coding agent by an average of 18.0 percentage points across frontier models, and matches or exceeds state-of-the-art specialized harnesses (up to 76.1% pass@1) while using 4.2-5.8x fewer tokens. With Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of \$1,750. Relevant code and logs are available at https://github.com/alexisfox7/PRO-LONG.
TRUST-ESD: 不確実性の下で企業の戦略的意思決定をサポートするための、リスク調整済みのガバナンスを意識した AI フレームワーク
企業の戦略的意思決定のサポートには、正確であるだけでなく、不確実性を認識し、リスク調整され、説明可能でガバナンスに準拠した AI システムが必要です。この文書では、不確実性の下で企業の意思決定をサポートするための、リスク調整済みのガバナンスを意識したフレームワークである TRUST-ESD を提案します。 TRUST-ESDは、予測効用推定、等角不確実性キャリブレーション、CVaRベースの下振れリスクスコアリング、リスクメモリ検索、コードとしてのポリシーガバナンス、説明可能性、人間による監視を通じて、実現可能な反事実戦略を評価します。期待される最大の効用によってアクションを選択する予測のみの手法とは異なり、TRUST-ESD は、価値、信頼性、リスクエクスポージャー、コンプライアンスのバランスをとる戦略を推奨します。実験結果によると、TRUST-ESD は、強力な不確実性を考慮したベースラインと比較して、リスク調整後の効用を 7.95% 改善し、リスクエクスポージャを 23.22% 低減し、CVaR を 23.78% 低減し、校正誤差を 13.89% 低減し、説明忠実度を 10.90% 改善し、ガバナンスコンプライアンスを 9.76% 向上させながら、競争力のある予測精度を維持します。アブレーション分析とケーススタディ分析により、不確実性の調整、下方リスクのスコアリング、リスク記憶、説明可能性、ガバナンスの検証が連携して、信頼できる企業の意思決定を向上させることがさらに確認されました。
原文 (English)
TRUST-ESD: A Risk-Calibrated and Governance-Aware AI Framework for Enterprise Strategic Decision Support Under Uncertainty
Enterprise strategic decision support requires AI systems that are not only accurate, but also uncertainty-aware, risk-calibrated, explainable, and governance-compliant. This paper proposes TRUST-ESD, a risk-calibrated and governance-aware framework for enterprise decision support under uncertainty. TRUST-ESD evaluates feasible counterfactual strategies through predictive utility estimation, conformal uncertainty calibration, CVaR-based downside-risk scoring, risk-memory retrieval, policy-as-code governance, explainability, and human oversight. Unlike prediction-only methods that select actions by maximum expected utility, TRUST-ESD recommends strategies that balance value, reliability, risk exposure, and compliance. Experimental results show that TRUST-ESD improves risk-adjusted utility by 7.95%, reduces risk exposure by 23.22%, reduces CVaR by 23.78%, lowers calibration error by 13.89%, improves explanation fidelity by 10.90%, and increases governance compliance by 9.76% compared with strong uncertainty-aware baselines, while maintaining competitive predictive accuracy. Ablation and case-study analyses further confirm that uncertainty calibration, downside-risk scoring, risk memory, explainability, and governance validation jointly improve trustworthy enterprise decision-making.
量子化小規模言語モデル推論のための CUSUM 形状の推論時間モニタリングとターゲットを絞った再デコード
量子化された小さな自己回帰推論モデルは、長く、反復的、または非生産的な軌道に入る可能性がありますが、推論時間の計算は通常、軌道がどのように発展するかを観察することなく割り当てられます。以前のトークンレベルの e-CUSUM コントローラーを基盤として、MGT-B (監視ガイド付きテスト時間バックトラッキング) を開発します。これは、プレサンプリングの不確実性と縮退特徴の重複するウィンドウを位置条件付きの経験的テール確率にマッピングし、CUSUM 型のリセットで混合ベッティング ファクターを蓄積し、ロールバック ポイントの推定、トークンとキー値キャッシュの状態の復元によってアラームに応答する、改訂された外部コントローラーです。制約付き再デコードを実行します。ログしきい値 h = 10 を手動で選択した後に最初に観察された問題 ID に対する影響が継続しているかどうかを監査するために、しきい値前のアーティファクトに存在する 260 ID を遡及的に除外し、残りの ID ごとに時系列で最初のしきい値後のペアを保持し、240 ペアの時系列監査セットを生成します。このセットでは、精度は 82/240 から 88/240 に変化します (+2.50 パーセンテージ ポイント、13 回の補正、7 回の回帰、正確な McNemar p = 0.2632、ペア ブートストラップ 95% 間隔 [-1.25、+6.25])。シード一致ペアのより広範な 467 ペアの履歴カバレッジ セットでは、精度が 146/467 から 167/467 に変化します (+4.50 ポイント、McNemar p = 0.000753)。ただし、しきい値の選択前またはしきい値の選択中に利用可能な 200 個のシード 1 ID が含まれており、探索的な推定値としてのみレポートされます。 467 ペア セット内の 316 個のアラームなし出力はすべてバニラと同一ですが、151 個のアラーム付き軌道には 29 個の修正と 8 個の回帰が含まれています。どちらの分析も確認的ではなく、経験的要因は有効な電子プロセスまたは電子検出器として確立されていません。この結果は、一般的または理論的に証明された推論の改善ではなく、研究された MATH-500 設定の選択的な監視および修復メカニズムをサポートしています。
原文 (English)
CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning
Quantized small autoregressive reasoning models can enter long, repetitive, or unproductive trajectories, yet inference-time compute is usually allocated without observing how a trajectory develops. Building on an earlier token-level e-CUSUM controller, we develop MGT-B (Monitoring-Guided Test-time Backtracking), a revised external controller that maps overlapping windows of pre-sampling uncertainty and degeneration features to position-conditional empirical tail probabilities, accumulates mixture betting factors with a CUSUM-shaped reset, and responds to an alarm by estimating a rollback point, restoring token and key-value-cache state, and performing constrained re-decoding. To audit whether the effect persists on problem identities first observed after the manual choice of log threshold h = 10, we retrospectively exclude 260 IDs present in pre-threshold artifacts and retain the chronologically first post-threshold pair for each remaining ID, yielding a 240-pair chronology-audit set. On this set, accuracy changes from 82/240 to 88/240 (+2.50 percentage points; 13 corrections, 7 regressions; exact McNemar p = 0.2632; paired bootstrap 95% interval [-1.25, +6.25]). A broader 467-pair historical-coverage set of seed-matched pairs changes accuracy from 146/467 to 167/467 (+4.50 points; McNemar p = 0.000753), but includes 200 seed-1 IDs available before or during threshold selection and is reported only as an exploratory estimate. All 316 no-alarm outputs in the 467-pair set are identical to vanilla, while the 151 alarmed trajectories contain 29 corrections and 8 regressions. Neither analysis is confirmatory, and the empirical factors are not established as a valid e-process or e-detector. The results support a selective monitoring-and-repair mechanism for the studied MATH-500 setting, rather than a general or theoretically certified reasoning improvement.
PoTRE: 認知的異質性からインスピレーションを得たテスト時の推論
大規模言語モデル (LLM) は多くのタスクに優れていますが、長期的な計画と反復的なエラー修正を必要とする複雑な推論に苦労することがよくあります。さらに、標準の単一ストリーム プロンプトは、モデルが新しい抽象化や厳密なドメイン制約に遭遇すると脆弱であることがわかります。推論を 4 つのエージェント (1) Adversarial Refinement Agent、(2) Hierarchical Strategy Planning Agent、(3) Spectrum Search Agent、および (4) Direct Chain Agent に分離する異種フレームワークである PoTRE (Poly-Topological Reasoning Ensembles) を紹介します。最終的なタスク適応型集約層は、最終候補の選択、意味論的合成、または神経記号検証を介してこれらの観点を動的に調整し、堅牢なグローバル ソリューションを生成します。私たちは、ARC-AGI-2、人類最後の試験 (HLE)、PRBench Finance という 3 つのフロンティア ベンチマークで PoTRE を評価します。 PoTRE は、HLE 上で 49.92% という最先端の精度を達成し、これまでの最高公式スコアを上回りました。このアーキテクチャの異質性により、大幅にスケールされた同種のベースラインと比較して、同様またはより少ない推論トークンを使用して推論パフォーマンスの向上が達成されることを実証します。
原文 (English)
PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity
While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when models encounter novel abstractions or rigorous domain constraints. We introduce PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents: (1) Adversarial Refinement Agent, (2) Hierarchical strategic Planning Agent, (3) Spectrum Search Agent, and (4) Direct Chain Agent. A final Task-Adaptive Aggregation Layer dynamically reconciles these perspectives -- via final candidate selection, semantic synthesis, or neuro-symbolic verification -- to produce a robust global solution. We evaluate PoTRE on three frontier benchmarks: ARC-AGI-2, Humanity's Last Exam (HLE), and PRBench Finance. PoTRE achieves state-of-the-art accuracy of 49.92% on HLE, surpassing the previous best official score. We demonstrate that this architectural heterogeneity achieves improved reasoning performance using similar or fewer inference tokens compared to heavily scaled homogeneous baselines.
リーダーではなくモデルをトレーニングする: 検証可能なアクティベーションの説明のための解読可能性の監視
自然言語オートエンコーダーは、再構成によって隠れたアクティベーションの説明をスコアリングします。説明からアクティベーションを再生成できれば、その説明は忠実であるとみなされます。このテストは、構造的に個々の誤った主張に対して鈍感です。つまり、主張を反転しても再構成が変更されない場合、その主張がペナルティを受けることはありません。私たちは、テストが 2 つの方法で合格することを示しますが、どちらも忠実ではありません。リリースされた Qwen-2.5-7B 言語化ツールでは、説明は偶然をはるかに超えて再構成されますが、特定の主張の ~2% は再構成に依存しているため、スコアは特定の事実ではなく要点を追跡します。正確な合成グランド トゥルースの下では、標準レシピは 5/5 回の実行で同時適応プライベート コード (再構築が依存する偽の文言) を開発し、ターゲット モデルを変更しないままにする修正は役に立ちません。私たちは、grounded-vs-true クロスと評価者スワップという 2 つの監査プロトコルと、指定されたコンテンツをデコード可能に保つためにターゲット モデルと一緒にトレーニングされたリニア ヘッドである RECAP (Readable Encodings via Co-trained Auxiliary Predictors) に貢献しています。 RECAP でトレーニングされたサンドボックス モデルでは、新しい言語化者が指定されたコンテンツを正確に記述し、コードは +0.001 nat のコストで消えます。これは事前トレーニング済みの Pythia-160M 上で再現されます。コンテンツは確実にプローブでデコード可能になりますが、新しいバーバライザーは部分的にしか伝えません (真実は 0.44 ~ 0.46 対ゼロに近いコントロール)。解釈可能性を考慮すると、高度な再構成は個々の主張を証明するものではありません。 AI の安全性を確保するために、RECAP は指定された内部コンテンツを、モデルがゲームできる散文によって主張するのではなく、プローブに対して独立してチェックできるようにします。独立したプローブは、言語化者の真の主張を誤った主張よりもスコア付けします (AUC 0.96、RECAP なしの場合は 0.82)。嘘をつきながら再構成スコアを最大化するように説明を編集する敵対者(嘘ペナルティの最大 87% を抑制)に対して、RECAP プローブは依然として嘘にフラグを立てます(AUC 0.95)が、対照プローブは偶然に崩壊します(0.51)。
原文 (English)
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the reconstruction, the claim is never penalized. We show the test is passed in two ways, neither faithful. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while ~2% of specific claims are reconstruction-dependent, so the score tracks gist, not specific facts. Under exact synthetic ground truth, the standard recipe develops co-adapted private codes (false wording the reconstruction depends on) in 5/5 runs, and fixes that leave the target model unchanged do not help. We contribute two audit protocols, the grounded-vs-true cross and the evaluator swap, and RECAP (Readable Encodings via Co-trained Auxiliary Predictors): linear heads trained alongside the target model to keep designated content decodable. On RECAP-trained sandbox models, fresh verbalizers state the designated content truly and the codes vanish, at a +0.001-nat cost. This replicates on a pretrained Pythia-160M: the content becomes reliably probe-decodable, though a fresh verbalizer conveys it only in part (truth 0.44-0.46 vs a near-zero control). For interpretability, high reconstruction does not certify individual claims. For AI safety, RECAP makes designated internal content independently checkable against probes rather than asserted by prose a model can game: an independent probe scores the verbalizer's true claims above its false ones (AUC 0.96, vs 0.82 without RECAP). Against an adversary that edits an explanation to maximize the reconstruction score while lying (suppressing ~87% of its lie penalty), the RECAP probe still flags the lies (AUC 0.95) while the control probe collapses to chance (0.51).
SoftReason: 高次元の知覚データに対する完全微分可能なニューロ ソフト シンボリック演繹推論アーキテクチャ
多くの推論問題では、前提は個別のシンボルとして観察されず、高次元の入力から推論する必要があります。さらに、述語の語彙、引数の構造、および信頼できる証拠は、ナレッジ グラフ (KG)、またはルール定義によって提供されます。古典的な神経記号パイプラインには、知覚と演繹の間に個別のインターフェイスがあります。我々は、潜在的な知覚事実と知識が提供する述語に対する微分可能な演繹的推論のためのニューロ・ソフト・シンボリック・アーキテクチャを提示します。 SoftReason は、候補定数と述語に対するローカル ソフト解釈テンソルとして演繹状態を表すことにより、勾配ギャップを除去します。認識は確率的な基本事実を提案し、KG トリプルは信頼性の高いソフト証拠として入力され、すべてのクエリアンカー、述語の選択、およびクロージャーの更新は微分可能のままです。私たちの核となるイノベーションは、直接結果演算子の学習された微分可能リフトです。述語定義の埋め込みと潜在的な構成チャネルを使用して、ソフトボディと述語の混合を形成し、すべての可能な証人を集約し、クエリ条件付きヘッドファクトを提案し、単調な確率的 OR を通じて解釈を更新します。 Knowledge-aware Visual Question Answering (KVQA) のフレームワークをインスタンス化し、SoftReason が 1 つのトレーニング可能なアーキテクチャでエンドツーエンドの知覚グラウンディング、KG 証拠の注入、および微分可能な演繹的閉包をどのようにサポートするかを示します。
原文 (English)
SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data
In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further, the predicate vocabulary, argument structure, and trusted evidence are supplied by a Knowledge Graph (KG), or rule definitions. Classical neuro-symbolic pipelines have a discrete interface between perception and deduction. We present a neuro-soft-symbolic architecture for differentiable deductive reasoning over latent perceptual facts and knowledge-provided predicates. SoftReason removes the gradient gap by representing the deductive state as a local soft interpretation tensor over candidate constants and predicates. Perception proposes probabilistic base facts, KG triples enter as high-confidence soft evidence, and every query anchor, predicate choice, and closure update remains differentiable. Our core innovation is a learned differentiable lift of the immediate-consequence operator. It uses predicate-definition embeddings and latent composition channels to form soft body-predicate mixtures, aggregate over all possible witnesses, propose query-conditioned head facts, and update the interpretation through a monotone probabilistic OR. We instantiate the framework on Knowledge-aware Visual Question Answering (KVQA), and demonstrates how SoftReason supports end-to-end perceptual grounding, KG evidence injection, and differentiable deductive closure in one trainable architecture.
マルチターン LLM システム用のステートフル ガードレール: 会話型リスク蓄積フレームワーク
大規模言語モデル (LLM) のほとんどの安全ガードレールは、プロンプトと応答の各ペアを個別に評価します。これにより、無害なものが有害なものに変わるときに、対話上でのみ発生する失敗が見逃されます。私たちはこれを会話リスク蓄積 (CRA) と呼んでいます。つまり、段階的な意図の漂流、禁止された指示の断片的な組み立て、および繰り返しの開示による機密性の蓄積です。我々は、セッション アンカーからのセマンティック ドリフト、抽出されたエンティティにわたる感度で重み付けされた情報蓄積グラフ、および遵守意欲の増加を捕捉する遵守勾配信号の 3 つの軌跡信号を追跡するセッション層 CRA フレームワークを提案します。スコアリングのために、(i) アトリビューションとアブレーションのための教師なし凸型融合、および (ii) 長さとトピックカバレッジの交絡を減らすために家族と敵対的な目標でトレーニングされたコンパクトな学習軌跡モデルである CRA-Net DA を提供します。 CRA のベンチマークを行うために、CRA-Bench v0.1 (トピックが一致する良性双生児を含む 3 つの脅威ファミリーにわたる 1,200 の 8 ターン セッション)、CRA-Bench v0.2 (テンプレートのアーチファクトを削減するための LLM 言い換えバリアント)、および拡張 5 ファミリー セット (ペルソナ プライミングとコンテキスト スタッフィングを追加する 2,000 セッション) をリリースします。セッションレベルの分割、混合セット閾値キャリブレーション、Trajectory AUROC、ターンから検出、キャリブレーションされた偽陽性メトリクス、ブートストラップ信頼区間、リーブワンファミリーアウト診断ストレステスト、および合成からヒトへの移行チェックを備えたトラジェクトリネイティブ評価プロトコルを導入します。主張は、CRA-Bench および人的転送サブセットのディストリビューション内セッション スコアリングに焦点を当てています。
原文 (English)
Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework
Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm. We term this Conversational Risk Accumulation (CRA): gradual intent drift, fragmented assembly of prohibited instructions, and sensitivity build-up from repeated disclosures. We propose a session-layer CRA Framework that tracks three trajectory signals: semantic drift from a session anchor, a sensitivity-weighted information accumulation graph over extracted entities, and a compliance-gradient signal capturing increasing willingness to comply. For scoring, we provide (i) an unsupervised convex fusion for attribution and ablations, and (ii) CRA-Net DA, a compact learned trajectory model trained with family-adversarial objectives to reduce length and topic-coverage confounds. To benchmark CRA, we release CRA-Bench v0.1 (1,200 eight-turn sessions across three threat families with topic-matched benign twins), CRA-Bench v0.2 (LLM-paraphrased variants to reduce template artifacts), and an extended 5-family set (2,000 sessions adding persona priming and context stuffing). We introduce a trajectory-native evaluation protocol with session-level splits, mixed-set threshold calibration, Trajectory AUROC, turns-to-detection, calibrated false-positive metrics, bootstrap confidence intervals, leave-one-family-out diagnostic stress tests, and synthetic-to-human transfer checks. Claims focus on within-distribution session scoring on CRA-Bench and human-transfer subsets.
言語モデルの経済的評価
言語モデルは経済的に価値のある作業を実行しますが、経済的に価値のあるすべてのタスクをどの程度うまく実行するかについては、現時点では評価されていません。米国の労働経済におけるタスク、作業活動、職業に関連する能力を測定するためのオープンソース評価スイートとして EconEvals を紹介します。可能な場合には、言語モデルに対する実際のユーザー クエリを評価スイートに基づいて作成し、これらを合成データで補完します。私たちの評価により、米国の職業の 5% をカバーする既存の最先端技術である OpenAI の GDPval ベンチマークよりもカバー範囲が向上しており、コストは 500 分の 1 です。ベンチマークに加えて、現在の言語モデル機能が米国のすべての職業に属するすべてのタスクにわたってどれだけの時間を節約できるかを推定するためのシミュレーションベースのエクスポージャー測定も導入し、それぞれの推定値を詳細に説明します。私たちの推定では、現在のモデルにより、労働者は 47% の職業において少なくとも半分の作業で大幅な時間を節約できる可能性があることが示されています。ただし、大幅な時間の節約が予測されるタスクの 79% では、観察されたクロードの使用量は低く、既存の使用量が潜在的なものより遅れていることを示唆しています。言語モデルのチャットボットに固有の制約を超えて、私たちのデータは、プライバシーと独自のシステムが AI によるさらなる時間節約を制限する主なボトルネックであることを特定しています。全体として、言語モデルの現在の機能における労働市場への影響についての推論を根拠づける、適応可能なインフラストラクチャを導入します。これは、機能の向上に応じて継続的に更新できます。
原文 (English)
Economic Evaluations of Language Models
Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an open-source evaluation suite to measure capabilities relevant to tasks, work activities, and occupations in the US labor economy. We ground the evaluation suite in real user queries to language models where possible, and supplement these with synthetic data. Our evaluations improve coverage over OpenAI's GDPval benchmark, which is the existing state-of-the-art that covers 5% of US occupations, at 500x lower cost. Alongside benchmarks, we also introduce a simulation-based exposure measure to estimate how much time current language model capabilities could save across all tasks belonging to all US occupations, with detailed accounting for each estimate. Our estimates indicate that current models could save workers substantial time on at least half of their tasks in 47% of occupations. However, for 79% of tasks where we predict substantial time savings, observed Claude usage is low, suggesting that existing usage lags potential. Beyond inherent constraints of language model chatbots, our data identifies privacy and proprietary systems as the principal bottlenecks limiting further time savings from AI. Overall, we introduce adaptable infrastructure that grounds inferences about language models' labor-market impact in their current capabilities, which can be continually updated as capabilities improve.
時系列予測の継続学習における説明可能性の課題
深層学習モデルは、時系列予測に大きな可能性を示していますが、現実世界の環境モニタリングへの導入は、非定常ダイナミクスと説明可能性の限界のため、依然として困難です。この研究では、エクスペリエンス リプレイ戦略を使用した、適応時系列予測における継続的学習を理解するための中心的なツールとしての説明可能性を調査します。私たちは、PatchMixer、PatchTST、DLinear などのニューラル予測アーキテクチャを研究し、長期にわたるモデルの適応をサポートするために注意ベースのサンプリング メカニズムを強化しています。説明可能性は、注意ロールアウトと勾配ベースのアトリビューション手法 (Grad-CAM) を通じて活用され、継続的な学習フレームワーク内で予測行動とサンプリング戦略の両方を分析します。異質なパターンとレジームシフトを示す現実世界のピエゾメトリック時系列に対して行われた実験では、モデルとサンプリングの動作を分析することで、継続的な学習フレームワークのダイナミクスに対する貴重な洞察が得られることが示されています。予測パフォーマンスを超えて、私たちの結果は、継続的な学習行動を理解するために説明可能性を使用することの課題と機会を強調し、アトリビューションパターンが時間の経過とともにどのように進化するか、そしてそれらが非定常予測シナリオにおけるデータ選択と適応戦略にどのように情報を提供できるかを明らかにします。
原文 (English)
Challenges of Explainability in Continual Learning for Time Series Forecasting
Deep learning models have shown strong potential for time series forecasting, yet their deployment in real-world environmental monitoring remains challenging due to non-stationary dynamics and limited explainability. In this work, we investigate explainability as a central tool for understanding continual learning in adaptive time series forecasting, with Experience Replay strategies. We study neural forecasting architectures such as PatchMixer, PatchTST and DLinear, augmented with attention-based sampling mechanisms to support model adaptation over time. Explainability is leveraged through attention rollout and gradient-based attribution methods (Grad-CAM) to analyze both predictive behavior and sampling strategies within a continual learning framework. Experiments conducted on real-world piezometric time series exhibiting heterogeneous patterns and regime shifts show that analyzing model and sampling behaviors provides valuable insights into the dynamics of the continual learning framework. Beyond predictive performance, our results highlight the challenges and opportunities of using explainability to understand continual learning behaviors, revealing how attribution patterns evolve over time and how they can inform data selection and adaptation strategies in non-stationary forecasting scenarios.
ビン化されたスペクトル損失による非構造化メッシュ上のカオス ダイナミクスのスケールを意識した学習
カオスを示す高次元の非線形力学システムのサロゲート モデリングには、点単位の精度だけでなく、物理場のスケール依存構造も保存するメカニズムが必要です。ビニングされたスペクトル損失関数などの帯域ごとのスペクトル電力損失は、フーリエ モードが標準的な周波数分解を定義する構造化グリッド上でそのような監視を提供します。ただし、不規則なメッシュでは正準フーリエ基底が存在せず、スペクトル表現はメッシュの接続性とジオメトリによって引き起こされるグラフ演算子から構築する必要があります。この研究では、ビン化されたスペクトル パワー損失を拡張して、非線形力学システムの非構造メッシュ サロゲート モデリングに適用します。これは、フーリエ バンドをグラフ ラプラシアン周波数バンドに置き換えることによって得られ、長期ロールアウトの忠実度を向上させるためのスケーラブルなチェビシェフ近似とマルチレベル近似を提供します。フルスペクトル形式では、私たちのアプローチはグラフのラプラシアン固有空間を使用して、フーリエ帯域パワーマッチングのグラフ類似物を提供しますが、スペクトル分解の高いコストが発生します。スケーラブルな近似として、厳密なバンド プロジェクターをスパースのチェビシェフ多項式グラフ フィルターに置き換え、明示的な固有分解を回避します。マルチレベル グラフ アーキテクチャを利用する場合、Graph Laplacian Energy Alignment for Meshes (GLEAM) を導入します。これは、グラフ階層全体に保持部分空間のスケールを意識した監視を適用するため、自己回帰ロールアウト中に粗い表現と細かい表現が正規化されます。我々の結果は、提案されたスペクトル損失が、決定論的なベースラインと比較して、長期ロールアウトの忠実度を向上させ、非構造化メッシュ上の乱流の予測に対する統計的不変量を維持することを示しています。
原文 (English)
Scale-Aware Learning of Chaotic Dynamics on Unstructured Meshes via Binned Spectral Losses
Surrogate modeling for high-dimensional nonlinear dynamical systems that exhibit chaos requires mechanisms that preserve not only pointwise accuracy but also the scale-dependent structure of physical fields. Bandwise spectral power losses, such as the binned spectral loss function, provide such supervision on structured grids, where Fourier modes define a standard frequency decomposition. On irregular meshes, however, no canonical Fourier basis exists, and spectral representations must be constructed from graph operators induced by mesh connectivity and geometry. In this study, we extend the binned spectral power loss for application to unstructured-mesh surrogate modeling of nonlinear dynamical systems. This is obtained by replacing Fourier bands with graph-Laplacian frequency bands, and we provide scalable Chebyshev and multilevel approximations for improving long-horizon rollout fidelity. In its full-spectrum form, our approach uses graph Laplacian eigenspaces to provide a graph analogue of Fourier band-power matching, but incurs the high cost of spectral decomposition. As a scalable approximation, we replace exact band projectors with sparse Chebyshev polynomial graph filters, avoiding explicit eigendecomposition. When utilizing multilevel graph architectures, we introduce Graph Laplacian Energy Alignment for Meshes (GLEAM), which applies retained-subspace scale-aware supervision across graph hierarchies so that coarse and fine representations are regularized during autoregressive rollout. Our results show that the proposed spectral losses improve long-horizon rollout fidelity and preserve statistical invariants for the forecasting of turbulent flows on unstructured meshes, compared to deterministic baselines.
ユートピアのシミュレーション: 結果、パフォーマンス、ダイナミクスによる長期的な公平性を再考する
AI 主導の意思決定者 (ADM) が社会経済的現実に影響を与える中、効率の向上と社会的偏見の増幅の両方における ADM の役割が注目を集めています。この論文では、特に信用貸付によって引き起こされる富のプロセスの文脈において、ADM によって達成可能な長期的な「公平性」のニュアンスを再検討します。長期的な公平性に関する文献は主に、(a) 受動的な環境、つまり予測子の結果が集団の行動を変えないことを考慮しており、(b) 下流の公平性ではなく、瞬間的な予測の不均衡という観点からバイアスを測定しています。これらは、信用貸し業者などの最新の ADM には当てはまりません。これらの警告に対処するために、まず、ローン承認 ADM によって引き起こされる富のダイナミクスを、ADM レベルと社会的成果レベルの報酬関数を備えた実行的マルコフ決定プロセスとして、複数の人口統計上の人口と相互作用して形式化します。次に、長期的な公平な戦略を学習するための新しいパフォーマンス データ ジェネレーターを備えた融資プロセス シミュレーターである Eutopia を開発することで、そのようなパフォーマンスの高いテストベッドの欠如を軽減します。最後に、さまざまな公平性を意識した実用的なユーティリティを使用して、パフォーマティブな古典的な RL アルゴリズムをテストします。実験結果は、(a) パフォーマンスダイナミクスによる学習は長期的な効率と公平性の向上につながり、(b) 社会的成果に基づいて評価される、適切に設計された公平性を意識した効用による学習は、より優れた効率性、公平性、包括性をもたらすことを示しています。
原文 (English)
Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dynamics
As AI-driven Decision Makers (ADMs) influence our socioeconomic reality, their roles in both enhancing efficiency and amplifying the social biases have drawn attention. In this paper, we revisit the nuances of long-term `fairness' achievable by an ADM, specifically in the context of a credit lending induced wealth process. The literature on long-term fairness mostly (a) considers passive environments, i.e. the outcome of a predictor does not change the population's behaviour, and (b) measures bias in terms of disparity in instantaneous predictions rather than the downstream equity. These are not true for modern ADMs, like credit lenders. To address these caveats, we first formalise the wealth dynamics induced by a loan approving ADM interacting with a multi-demographic population as a performative Markov Decision Process with ADM level and social outcome level reward functions. Then, we mitigate the absence of such a performative test-bed by developing Eutopia: a lending-process simulator enabled with a novel performative data generator to learn long-term fair strategies. Finally, we test performative and classical RL algorithms with different fairness-aware and utilitarian utilities. Experimental results show that (a) learning with performative dynamics lead to better long-term efficiency and equity, and (b) learning with well-designed fairness-aware utility evaluated on social outcomes induces better efficiency, equity, and inclusivity.
LAARA: パラメータ効率の高い微調整のためのレイヤー認識型適応ランク割り当て
低ランク適応はパラメータ効率の高い微調整に広く使用されていますが、既存の方法では通常、異種適応要件にもかかわらず、すべてのトランス層に同じアダプター ランクが割り当てられます。この研究では、均一なランク割り当てが基本的に最適ではないことを理論的および経験的に示します。この観察を動機として、トレーニング中に計算された軽量の対角フィッシャー推定を使用してランクを動的に割り当てる検索不要のフレームワークである LAARA (Layer Aware Adaptive Rank Allocation Framework) を提案します。 LAARA は、射影による正規化、対数圧縮、ブレンドされたアダプターの重要度推定、投票による変更の減衰メカニズムを組み合わせて、安定した効率的なランク適応を生成します。 GLUE および MathInstruct ベンチマークの実験では、LAARA が、使用するトレーニング可能なパラメーターが大幅に少ないにもかかわらず、LoRA、AdaLoRA、DyLoRA、Bitfit などの一般的な最先端のアプローチと一貫して同等またはそれを上回るパフォーマンスを示すことが実証されました。私たちの結果は、フィッシャーに基づいたランク割り当てが、適応パラメーター効率の良い微調整のための原則に基づいた効果的な基盤を提供することを示しています。コードは、https://anonymous.4open.science/r/LAARA-D305/LAARA.py で公開されています。
原文 (English)
LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning
Low-Rank Adaptation is widely used for parameter-efficient fine-tuning, yet existing methods typically assign the same adapter rank to every transformer layer despite their heterogeneous adaptation requirements. In this work, we show theoretically and empirically that uniform rank allocation is fundamentally suboptimal. Motivated by this observation, we propose LAARA (Layer Aware Adaptive Rank Allocation framework), a search-free framework that dynamically allocates ranks using lightweight diagonal Fisher estimates computed during training. LAARA combines projection-wise normalization, logarithmic compression, blended adapter importance estimation, and a vote-to-change dampening mechanism to produce stable and efficient rank adaptation. Experiments on GLUE and MathInstruct benchmark demonstrate that LAARA consistently matches or outperforms popular state of the art approaches such as LoRA, AdaLoRA, DyLoRA, and Bitfit while using significantly fewer trainable parameters. Our results show that Fisher-guided rank allocation provides a principled and effective foundation for adaptive parameter-efficient fine-tuning. The code is publicly available at: https://anonymous.4open.science/r/LAARA-D305/LAARA.py
デコード可能だが検出不可能: OOD に近いベンチマークの漏洩フィンガープリント
ドキュメント ベンチマークで摂動ベースの OOD 検出器を監査しているときに、AUROC 0.326 を記録しました。これは、チャンス レベル 0.5 を大きく下回っています。原因はベンチマーク リークです。指定された「OOD」クラスはモデルがトレーニングされたクラスであるため、その例は分布内適合セット内に位置し、検出器はそれらを馴染みのあるものとして正しくランク付けするためにペナルティを受けます。クラスを削除し、2 つのドメインにわたって 35 のモデルを再トレーニングすると、スコアは 0.911 に上昇します。私たちは汚染を抽出してリークフィンガープリントを作成し、ほぼ完璧な教師付き解読可能性(AUROC 約 1)と教師なし検出が 0.65 未満に崩壊しました。そして、CIFAR-10/100 上の ResNet-50 および ViT-B/16 にわたる 52 設定(20 リーク、32 クリーン)の制御されたバッテリーで検証し、感度 18/20 と特異度 31/32 を達成しました。埋め込みスペース。一致する適合セット除外コントロールは 20/20 で完璧です。 24 の標準的な近/遠 OOD ベンチマーク ペアの実際の監査は、厳密に 1 つ (本質的に厳しい CIFAR-100 対 CIFAR-10 ペア) に対して実行され、遠 OOD ペアは実行されず、特異性と標準のクロスデータセット構築がクリーンであることが確認されます。修正されたプロトコルでは、摂動信号は解読可能ですが検出できません。教師付きリーダーは OOD 信号 (AUROC 0.87-1.00) を回復しますが、教師なし検出器は回復しません。また、摂動法は単純なマハラノビス距離では改善されません。私たちはその理由についての理論的説明を提供し、透明性のために以前の循環相関を撤回します。貢献するのは修正されたプロトコルと検証されたリーク診断であり、新しい OOD 手法ではありません。
原文 (English)
Decodable but Not Detectable: A Leakage Fingerprint for Near-OOD Benchmarks
While auditing a perturbation-based OOD detector on a document benchmark, we recorded an AUROC of 0.326 -- well below the 0.5 chance level. The cause is a benchmark leak: the designated "OOD" class is one the model was trained on, so its examples sit inside the in-distribution fit set and the detector is penalized for correctly ranking them as familiar. Deleting the class and retraining 35 models across two domains raises the score to 0.911. We distill the contamination into a leak fingerprint -- near-perfect supervised decodability (AUROC approximately 1) coupled with unsupervised detection collapsed below 0.65 -- and validate it on a controlled battery of 52 settings (20 leaked, 32 clean) across ResNet-50 and ViT-B/16 on CIFAR-10/100, achieving sensitivity 18/20 and specificity 31/32 in embedding space; the matched fit-set-exclusion controls are perfect at 20/20. An in-the-wild audit of 24 standard near/far OOD benchmark pairs fires on exactly one (the intrinsically hard CIFAR-100 vs CIFAR-10 pair) and on no far-OOD pair, confirming specificity and that standard cross-dataset construction is clean. Under the corrected protocol, perturbation signals are decodable but not detectable: a supervised reader recovers the OOD signal (AUROC 0.87-1.00) while no unsupervised detector does, and the perturbation method does not improve on plain Mahalanobis distance. We provide a theoretical account of why and, for transparency, retract an earlier circular correlation. The contributions are a corrected protocol and a validated leak diagnostic, not a new OOD method.
一般化された神経表現学習のための共有空間アライメントによる被験者間の意味解読
電極構成、解剖学的構造、神経信号パターンは個人によって大幅に異なるため、侵襲的神経記録では被験者全体を一般化することが依然として困難です。このような被験者間の変動を調査するために、我々は、複数の被験者からの音声知覚に対する神経反応を共有潜在空間に整列させ、整列された神経表現から文脈的埋め込みへのマッピングを学習する、被験者間の意味解読フレームワークを提案する。より具体的には、自然言語理解中に収集された皮質電気記録法データを使用して、共有応答モデルを使用して共有空間を推定し、投影された神経応答から文脈上の意味の埋め込みを予測するようにデコーダーをトレーニングします。保留された被験者の場合、事前定義された共有空間への被験者固有の投影を推定し、再トレーニングせずに事前トレーニングされたデコーダーを直接適用します。実験結果は、提案されたフレームワークが評価設定全体にわたって一貫してベースライン手法を上回り、ソース被験者から保留被験者テストまでのパフォーマンスの低下が減少していることを示し、被験者間の一般化が向上していることを示しています。これらの結果は、意味的埋め込み空間で解読しながら、神経活動を共有潜在空間に調整することが、共有された刺激関連表現を効果的に捕捉しながら、神経反応における被験者固有の差異を低減することにより、被験者間の汎化を改善するための効果的な戦略を提供することを示唆している。
原文 (English)
Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Representation Learning
Generalizing across subjects remains challenging in invasive neural recordings because electrode configurations, anatomical structures, and neural signal patterns vary substantially across individuals. To investigate such inter-subject variability, we propose a cross-subject semantic decoding framework that aligns neural responses to speech perception from multiple subjects into a shared latent space and learns a mapping from the aligned neural representations to contextual embeddings. More specifically, using electrocorticography data collected during natural language comprehension, we estimate the shared space using the shared response model and train a decoder to predict contextual semantic embeddings from projected neural responses. For a held-out subject, we estimate a subject-specific projection into the predefined shared space, and directly apply the pretrained decoder without any retraining. Experimental results demonstrate that the proposed framework consistently outperforms baseline methods across evaluation settings and exhibits a reduced performance drop from source subject to held-out subject testing, indicating improved cross-subject generalization. These results suggest that aligning neural activity into a shared latent space, while decoding in a semantic embedding space, provides an effective strategy for improving cross-subject generalization by reducing subject-specific differences in neural responses while effectively capturing shared stimulus-related representations.
軌跡から接頭辞へ: 再生された接頭辞とオンライン継続による教師の軌跡の再利用
小さな言語モデルは対話型エージェントにとって魅力的なバックボーンですが、強力な教師の軌跡から直接抽出すると、多くの場合、豊富なマルチターン動作がワンショットの模倣ターゲットに変わります。これは、早期の決定が後の状態と報酬を形作る長期的な環境では非効率的です。私たちは、教師の軌跡をリプレイに合わせた接頭辞クエリとオンライン継続に分解する強化学習フレームワークである Prefix-GRPO を提案します。各プレフィックスが環境内で再生されて有効な中間状態が回復され、その後、学生はオンラインでの対話を継続し、タスク報酬を受け取ります。応答のみの GRPO とは異なり、Prefix-GRPO は、ポリシーから抽出された SFT チェックポイントを使用して古いログ確率を推定することにより、クリップされたポリシーの更新を、再生されたプレフィックス内の履歴アシスタント トークンにも適用します。これにより、同じポリシー最適化形式内でプレフィックス学習と継続学習が統合されます。 TextCraft、BabyAI、ALFWorld での実験では、Prefix-GRPO が蒸留や標準 RL ベースラインよりも小規模モデルのエージェントを改善することが示されていますが、アブレーションでは明示的なプレフィックス トークンの最適化がなければリプレイだけでは不十分であることが示されています。実装および再現スクリプトは https://github.com/HappynessI/Prefix_GRPO で入手できます。
原文 (English)
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation
Small language models are attractive backbones for interactive agents, but direct distillation from strong teacher trajectories often turns rich multi-turn behavior into one-shot imitation targets. This is inefficient in long-horizon environments, where early decisions shape later states and rewards. We propose Prefix-GRPO, a reinforcement learning framework that decomposes teacher trajectories into replay-aligned prefix queries and online continuations. Each prefix is replayed in the environment to recover a valid intermediate state, after which the student continues online interaction and receives task reward. Unlike response-only GRPO, Prefix-GRPO also applies clipped policy updates to historical assistant tokens inside the replayed prefix, using a policy-distilled SFT checkpoint to estimate their old log-probabilities. This unifies prefix learning and continuation learning within the same policy-optimization form. Experiments on TextCraft, BabyAI, and ALFWorld show that Prefix-GRPO improves small-model agents over distillation and standard RL baselines, while ablations show that replay alone is insufficient without explicit prefix-token optimization. The implementation and reproduction scripts are available at https://github.com/HappynessI/Prefix_GRPO.
大規模な視覚・言語・行動モデルにおける効率的かつ一般化可能な強化学習のためのオフライン監視の活用
オンライン強化学習 (RL) は、幅広いパフォーマンス測定においてオフライン手法よりも優れたパフォーマンスの戦略を生成することが一般的に観察されています。特に、RL でトレーニングされたポリシーは、より強力な配布外 (OOD) 動作を示し、模倣学習アプローチのみでトレーニングされたモデルはしばしば困難を伴います。最近の研究では、OOD に焦点を当てたベンチマークが導入され、RL でトレーニングされたビジョン言語アクション (VLA) ポリシーは、教師あり微調整 (SFT) でトレーニングされた対応するポリシーよりも顕著に優れた OOD パフォーマンスとわずかに優れた分散内 (IND) パフォーマンスを達成することが報告されました。この研究では、オフラインとオンラインのハイブリッド トレーニングが両方のアプローチの利点を組み合わせられるかどうかを調査します。具体的には、オフライン データまたはオフラインでトレーニングされた参照ポリシーを介したオフライン監視によって正規化された RL 手法を研究します。これらのアプローチを OOD ベンチマークで評価し、オフラインのみのトレーニングと標準的な RL の両方と比較します。私たちの結果は、オフライン トレーニングだけでは限られた OOD パフォーマンスしか達成できませんが、RL にオフライン監視を組み込むことで、トレーニング効率を大幅に向上させながら、強力な OOD 機能を維持できることを示しています。特に、ガイド付きメソッドは、トレーニング予算の約半分を必要としながら、標準的な RL に近いパフォーマンスに達します。ハイブリッド アプローチでは、速度と OOD パフォーマンスの間にトレードオフが生じるのではなく、効率性の向上を達成しながら強力な OOD 機能を維持します。プロジェクトページ: https://alstar8.github.io/offline-supervision-vla-rl
原文 (English)
Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models
It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particular, RL-trained policies exhibit stronger out-of-distribution (OOD) behavior, where models trained only with imitation learning approaches often struggle. A recent study introduced an OOD-focused benchmark and reported that RL-trained vision-language-action (VLA) policies achieve noticeably better OOD performance and slightly better in-distribution (IND) performance than their counterparts trained with supervised fine-tuning (SFT). In this work, we investigate whether hybrid offline-online training can combine the advantages of both approaches. Specifically, we study RL methods regularized by offline supervision via either offline data or an offline-trained reference policy. We evaluate these approaches on the OOD benchmark and compare them with both offline-only training and standard RL. Our results show that although offline training achieves limited OOD performance by itself, incorporating offline supervision into RL preserves strong OOD capability while substantially improving training efficiency. In particular, the guided methods reach performance close to that of standard RL while requiring roughly half of the training budget. Rather than producing a trade-off between speed and OOD performance, the hybrid approach retains strong OOD capability while achieving this efficiency gain. Project page: https://alstar8.github.io/offline-supervision-vla-rl
差分プライバシー下での臨床有用性の回復: 異種心血管データセットに対する適応型フェデレーション集約の実証的検証
実際の臨床データに基づいてフェデレーテッド ラーニング フレームワークを検証することは、制御された合成環境での概念実証のデモンストレーションと、実際の多施設医療環境での展開との間に重要なステップです。同じ著者による以前のアーキテクチャ研究 (Tertulino および Alencar、2026) では、6 機能の合成ベンチマークで、サーバー側の適応最適化が差分プライバシー ノイズの一時的なノイズ除去として機能し、元のパイプライン作業で特定された未解決の課題に答えることが実証されました (Tertulino、2025)。この研究では、合成的に生成されたデータが使用され、現実世界での検証が優先される将来の方向性として明示的に特定されました。現在の研究では、公的に利用可能な 5 つの実際の心血管データセット (フラミンガム、クリーブランド、ハンガリー、スイス、バージニア州ロングビーチ) で FedCVR フレームワークを検証することで、このギャップに対処しています。これらのデータセットは、13 属性の UCI 心疾患スキーマに調和され、1 施設除外相互検証を備えた異種連合シナリオとして構成されています。結果は、FedCVR が実際のデータに対する適応的利点を維持し、運用プライバシー バジェット (ノイズ乗数 = 0.8、プライバシー バジェット イプシロン約 4.2) の下で 79.2% の F1 スコアと 0.96 の AUC を達成し、すべての評価指標 (対応のある t 検定、すべて p <= 0.003、有意差の下で有意) で標準 FedAvg を統計的に上回っていることを示しています。ボンフェローニ補正閾値)。実際のデータで測定されたプライバシー コストは、合成実験で観察された緩やかな劣化パターンを裏付けており、真の多施設環境におけるフレームワークの臨床的実行可能性の経験的証拠を提供します。
原文 (English)
Recovering Clinical Utility Under Differential Privacy: Empirical Validation of Adaptive Federated Aggregation on Heterogeneous Cardiovascular Datasets
Validating federated learning frameworks on real clinical data is an essential step between proof-of-concept demonstrations in controlled synthetic environments and deployment in real multicenter healthcare settings. A prior architectural study by the same authors (Tertulino and Alencar, 2026) demonstrated, on a synthetic six-feature benchmark, that server-side adaptive optimization acts as a temporal denoiser for Differential Privacy noise, answering an open challenge identified in the original pipeline work (Tertulino, 2025). That study used synthetically generated data and explicitly identified real-world validation as a priority future direction. The present work addresses this gap by validating the FedCVR framework on five publicly available real cardiovascular datasets (Framingham, Cleveland, Hungarian, Switzerland, and Long Beach VA), harmonized to the 13-attribute UCI Heart Disease schema and configured as a heterogeneous federated scenario with leave-one-institution-out cross-validation. Results demonstrate that FedCVR preserves its adaptive advantage on real data, achieving an F1-Score of 79.2% and AUC of 0.96 under the operational privacy budget (noise multiplier = 0.8, privacy budget epsilon approximately 4.2), while statistically outperforming standard FedAvg on all evaluated metrics (paired t-tests, all p <= 0.003, significant under the Bonferroni-corrected threshold). The measured privacy cost on real data confirms the graceful degradation pattern observed in the synthetic experiments, providing empirical evidence of the framework's clinical viability in genuine multicenter contexts.
多変量時系列予測のためのマルチスケール時間パッチ上の構造化された潜在空間モデリング
多変量時系列は、複数の時間スケールにわたって展開する構造パターンをコード化しますが、ほとんどの予測バックボーンは、学習された表現を予測の一時的な副産物として扱い、これらのパターンの組織幾何学は十分に活用されていません。チャネルに依存しない多変量観測を 2 つの相補的な微分可能な制約を通じて構造化された潜在空間にマッピングする CNN ベースの予測アーキテクチャである M2Patch を紹介します。マルチスケール パッチングは、入力を重複する時間粒度に分解します。段階的な拡張を伴う深さ方向の分離可能な畳み込みは、線形時間でスケール固有の特徴を抽出します。そして、スケールごとに学習された投影は、これらの特徴をコンパクトな潜在表現に圧縮します。潜在空間は、隣接するパッチ間の時間的連続性を強制するスケール内平滑性制約と、学習可能なクロススケール マッピングを通じて実現されるスケール間アライメント制約によって編成され、チャネル独立設計内で粒度間の相互作用を復元し、すべてのスケールが基礎となるダイナミクスの相互に一貫した表現を確実にエンコードします。 10 の実際のベンチマークでの実験では、M2Patch が 40 の予測設定全体で 57 件の最良の結果と 34 件の次善の結果を達成し、線形の計算複雑さとパッチレベルの入力破損に対する堅牢性を維持しながら、ほとんどのベンチマークの代表的なベースラインと一致またはそれを上回っていることが示されています。
原文 (English)
Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting
Multivariate time series encode structural patterns that unfold across multiple temporal scales, yet most forecasting backbones treat learned representations as transient byproducts of prediction, leaving the organizational geometry of these patterns underexploited. We introduce M2Patch, a CNN-based forecasting architecture that maps channel-independent multivariate observations into a structured latent space through two complementary differentiable constraints. Multi-scale patching decomposes the input into overlapping temporal granularities; depthwise separable convolutions with progressive dilation extract scale-specific features in linear time; and per-scale learned projections compress these features into a compact latent representation. The latent space is organized by an intra-scale smoothness constraint that enforces temporal continuity between adjacent patches, and an inter-scale alignment constraint, realized through learnable cross-scale mappings, that restores cross-granularity interaction within the channel-independent design, ensuring that all scales encode mutually consistent representations of the underlying dynamics. Experiments on ten real-world benchmarks show that M2Patch achieves 57 best and 34 second-best results across 40 forecasting settings, matching or exceeding representative baselines on most benchmarks while maintaining linear computational complexity and robustness to patch-level input corruption.
縦方向の細胞ペインティング形態に関する検索検索拡張 LLM 仮説の監査
ハイコンテンツの形態学的プロファイリング (細胞ペインティング) では、細胞状態の高感度な高次元シグネチャが得られますが、特に低線量率の電離放射線などの弱い慢性的な摂動の場合、縦方向の形態の軌跡を解釈可能な生物学に変換することは依然として困難です。大規模言語モデル (LLM) は、異種の証拠を生物学的物語に統合できますが、科学的に使用するには定量的な監査が必要です。我々は、5 つの線量率 (0.003 ~ 6.0 mGy/hr) にわたる 9 週間の RPE-1 タイムコースに適用した、縦方向のセル ペインティング形態に対する評価優先、検索拡張解釈フレームワークを提示します。週ごとに照合した治療対照形態デルタを、安定した証拠識別子を通じて取得した摂動近傍、経路コンテキスト、および文献証拠と組み合わせることで、LLM が来歴を維持しながら階層的に要約される、構造化された証拠にリンクされた仮説を生成できるようになります。 2 つの定量的監査テストを導入します。V1 引用の妥当性は、引用された証拠識別子がプロンプト内に存在することを検証します。もう 1 つは、V2 プロキシベースの形態互換性です。これは、予測された生物学的プロセスと最も変更された形態特徴との間の一貫性を評価します。私たちの実験では、V1 は無効な証拠参照を検出しませんでしたが、V2 は摂動強度とともに増加する意味のある形態互換性を示し、独立した形態ドリフト概要と正の関連がありました。このフレームワークは、低線量率 (0.003 ~ 0.3 mGy/hr) での代謝再プログラミングやタンパク質増殖性ストレスを含む適応表現型を含む、監査可能で反証可能な生物学的仮説を生成します。現在の制限には、プロキシベースの評価とグラウンドトゥルースメカニズムラベルの欠如が含まれます。
原文 (English)
Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology
High-content morphological profiling (Cell Painting) yields sensitive, high-dimensional signatures of cellular state, but translating longitudinal morphology trajectories into interpretable biology remains difficult, especially for weak, chronic perturbations such as low-dose-rate ionizing radiation. Large language models (LLMs) can synthesize heterogeneous evidence into biological narratives, yet their scientific use requires quantitative auditing. We present an evaluation-first, retrieval-augmented interpretation framework for longitudinal Cell Painting morphology, applied to a 9-week RPE-1 time course across five dose rates (0.003--6.0 mGy/hr). Week-matched treated-control morphology deltas are combined with retrieved perturbation neighbors, pathway context, and literature evidence through stable evidence identifiers, enabling an LLM to generate structured, evidence-linked hypotheses that are hierarchically summarized while preserving provenance. We introduce two quantitative auditing tests: V1 citation validity, which verifies that cited evidence identifiers exist in the prompt, and V2 proxy-based morphology compatibility, which evaluates consistency between predicted biological processes and the most altered morphology features. In our experiments, V1 detected no invalid evidence references, while V2 showed meaningful morphology compatibility that increased with perturbation strength and was positively associated with an independent morphology drift summary. The framework produces auditable, falsifiable biological hypotheses, including an adaptive phenotype involving metabolic reprogramming and proteostatic stress at lower dose rates (0.003--0.3 mGy/hr). Current limitations include proxy-based evaluation and the lack of ground-truth mechanism labels.
Opto-ViT-v2: フォトニックニアセンサービジョントランス加速器向けのノイズ耐性のあるオンチップ微調整
シリコンフォトニック(SiPh)加速器は、高いスループットとエネルギー効率でマイクロリング共振器(MRR)バンク上で行列乗算を実行することにより、ビジョントランスフォーマー(ViT)推論のための有望なプラットフォームとして浮上しています。バックプロパゲーションには大規模なアクティベーション ストレージ、MRR への頻繁な重みライトバック、およびデバイス レベルのノイズに対する耐性が必要なため、これらのプラットフォームを拡張してオンチップの微調整をサポートすることは依然として困難です。我々は、ニアセンサー SiPh ViT アクセラレータ上のパラメータ効率の良い微調整 (PEFT) のための最初のフレームワークである Opto-ViT-v2 を紹介します。テンソル化された低ランク分解は、事前トレーニングされた光学的重みをトレーニング可能な電子因子の小さなセット (ViT-Base ではわずか 8K パラメーター) から分離し、アクティベーション ストレージと重みの更新を大幅に削減し、同時に実用的なオンチップ トレーニングを可能にします。さらに、ワンショットのtop-k勾配マスキングを通じて重要度の低い重みを凍結する勾配累積スパース分類器を導入し、分類器のトレーニングコストを約40パーセント削減します。また、フォトニックオンチップトレーニング用の最初のシステムレベルのノイズモデルを開発し、順方向と逆方向の両方の伝播中のMRRクロストーク、熱ドリフト、レーザー振幅ノイズの影響を捕捉します。 200 を超える製造 MRR デバイスからの測定値を使用して校正されたこのモデルは、低ランク係数の更新が、同一のノイズ条件下での完全な微調整や従来の層ごとの低ランク適応よりも堅牢であることを示しています。 VTAB-1K (19 タスク) と FGVC の少数ショット ベンチマークの実験では、Opto-ViT-v2 が測定されたフォトニック ノイズの下でクリーンなソフトウェア精度の 0.3 ~ 0.8% 以内に回復し、100 KFPS/W 以上を達成し、フォトニック エッジ ビジョン システムの実用的なオンチップ ドメイン適応を可能にすることが実証されました。
原文 (English)
Opto-ViT-v2: Noise-Resilient On-Chip Fine-Tuning for Photonic Near-Sensor Vision Transformer Accelerators
Silicon-photonic (SiPh) accelerators have emerged as a promising platform for Vision Transformer (ViT) inference by performing matrix multiplications on microring-resonator (MRR) banks with high throughput and energy efficiency. Extending these platforms to support on-chip fine-tuning remains challenging because backpropagation requires large activation storage, frequent weight write-back to MRRs, and tolerance to device-level noise. We present Opto-ViT-v2, the first framework for parameter-efficient fine-tuning (PEFT) on a near-sensor SiPh ViT accelerator. Our tensorized low-rank decomposition separates pretrained optical weights from a small set of trainable electronic factors (as few as 8K parameters for ViT-Base), greatly reducing activation storage and weight updates while enabling practical on-chip training. We further introduce a gradient-accumulated sparse classifier that freezes low-importance weights through one-shot top-k gradient masking, reducing classifier training cost by about 40 percent. We also develop the first system-level noise model for photonic on-chip training, capturing the effects of MRR crosstalk, thermal drift, and laser amplitude noise during both forward and backward propagation. Calibrated using measurements from more than 200 fabricated MRR devices, the model shows that low-rank factor updates are more robust than full fine-tuning and conventional layer-wise low-rank adaptation under identical noise conditions. Experiments on VTAB-1K (19 tasks) and FGVC few-shot benchmarks demonstrate that Opto-ViT-v2 recovers within 0.3 to 0.8 percent of clean software accuracy under measured photonic noise while achieving more than 100 KFPS/W, enabling practical on-chip domain adaptation for photonic edge vision systems.
JailMeter: 大規模な言語モデルに対するジェイルブレイク攻撃のための証拠に基づいた評価フレームワーク
大規模な言語モデルに対するジェイルブレイク攻撃の評価は現在、一貫性のない評価基準と方法に悩まされており、攻撃成功率の推定値の信頼性が低くなります。私たちは、ジェイルブレイクの有効性をより忠実に測定するために設計された証拠に基づいた評価フレームワークである JailMeter を提案します。情報ボトルネック理論にヒントを得た JailMeter は、デュアルフィードバック最適化を適用して、元の悪意のある質問に関連するコンテンツを維持しながら、モデル応答からジェイルブレイク ノイズをフィルターします。このプロセスは、応答が悪意のある意図を捕捉し、完全な応答を提供する場合にのみ攻撃が検証される、厳密な評価のための簡潔な証拠を生成します。これにより、モデルの安全性調整の実質的なバイパスが示されます。私たちは、JailMeter-Eva で JailMeter を評価します。これは、人間がラベルを付けた拒否されていないジェイルブレイク インスタンス 330 個を含む、挑戦的なベンチマークです。 JailMeter は 97.27% の精度を達成し、既存の評価方法を大幅に上回ります。大規模な評価をサポートするために、JailMeter をさらに小規模な言語モデル JailMeter\textsubscript{SLM} に抽出しました。これは、計算コストを大幅に削減しながら同等の信頼性を維持します。コードとデータセットは https://github.com/Magi2B0y/JailMeter で入手できます。
原文 (English)
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an evidence-based evaluation framework designed to more faithfully measure jailbreak effectiveness. Inspired by the Information Bottleneck theory, JailMeter applies dual-feedback optimization to filter jailbreak noise from model responses while preserving content relevant to the original malicious question. This process produces concise evidence for a rigorous assessment under which an attack is validated only when the response captures the malicious intent and delivers a complete answer, thereby signaling a substantive bypass of model safety alignment. We evaluate JailMeter on JailMeter-Eva, a challenging benchmark containing 330 human-labeled, non-rejected jailbreak instances. JailMeter achieves an accuracy of 97.27%, substantially outperforming existing evaluation methods. To support large-scale evaluation, we further distill JailMeter into a small language model, JailMeter\textsubscript{SLM}, which maintains comparable reliability with significantly reduced computational costs. Code and dataset are available at https://github.com/Magi2B0y/JailMeter.
単一セル データの抽出を監査可能にする: 個別の最小値と最大値の選択による追跡可能なリアル セル コアセット
単一セル データセットの保存、監査、モデル トレーニングのための再利用のコストはますます高くなっています。次元の削減とデータセットの蒸留によりこの負担を軽減できますが、従来の蒸留方法では、アッセイされた細胞を追跡できない合成発現プロファイルが生成されることがよくあります。私たちは、固定された細胞と遺伝子のバジェットの下で元の細胞識別子と遺伝子シンボルを保持しながら、追跡可能な単一細胞データの蒸留を定式化します。結果として得られるトレーニング サブセットは、測定されたカウント、ラベル、およびアッセイ メタデータに接続されたままであるため、予期せぬ予測をソース データと照合してチェックできます。 2 つの実セル セレクターを提案します。固定 CF は静的な特性関数マッチングを使用します。 Minmax-CF は、あまり保存されていない方向を重み付けし、観測されたセルのみを追加する、エントロピー正則化された離散最小-最大問題を解決します。 3 つのデータセットのドナー、テクノロジー、および摂動レベルのシフト全体で、Minmax-CF は MS で完全バランスの精度の 96.52% を維持し、h膵臓で平均で完全とほぼ一致し、全遺伝子設定で中央値 $2.55\times$ の GPU 速度向上を実現し、Norman で圧縮メソッドの中で最も低い経路エラーを取得しました。まれな状態、一部のテクノロジーの変化、目に見えない摂動コンポーネント、および忠実度が下流のユーティリティとの関連性が低い設定では、パフォーマンスは依然として低いままです。選択された ID は測定されたセルを参照するため、これらのケースは、対応するトレーニング サポート、ラベル、およびアッセイ メタデータを検査することによって調査できます。 Minmax-CF は最悪方向の不一致を一貫して削減しますが、下流のユーティリティとコストはデータセットやタスクによって異なります。
原文 (English)
Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min-Max Selection
Single-cell datasets are increasingly costly to store, audit, and reuse for model training. Dimensionality reduction and dataset distillation can reduce this burden, but conventional distillation methods often produce synthetic expression profiles that cannot be traced to an assayed cell. We formulate traceable single-cell data distillation as retaining original cell identifiers and gene symbols under fixed cell and gene budgets. The resulting training subset remains connected to measured counts, labels, and assay metadata, so unexpected predictions can be checked against their source data. We propose two real-cell selectors. Fixed-CF uses static characteristic-function matching. Minmax-CF solves an entropy-regularized discrete min--max problem that upweights poorly preserved directions and adds only observed cells. Across donor-, technology-, and perturbation-level shifts on three datasets, Minmax-CF retains 96.52% of Full balanced accuracy on MS, approximately matches Full on average on hPancreas with a median $2.55\times$ GPU speedup in the all-gene setting, and obtains the lowest pathway error among compressed methods on Norman. Performance remains weaker for rare states, some technology shifts, unseen perturbation components, and settings where fidelity is weakly associated with downstream utility. Because the selected IDs refer to measured cells, these cases can be investigated by inspecting the corresponding training support, labels, and assay metadata. Minmax-CF consistently reduces worst-direction discrepancy, while downstream utility and cost vary across datasets and tasks.
ChannelGuard: 安全なモデルは安全なマルチエージェント システムを構成しない
マルチエージェント LLM アプリケーションは、プランナー、ワーカー エージェント、ベリファイア、およびシンセサイザーをチェーン化しており、エージェント間のすべてのホップは、敵対者が命令を密輸できる監視されていないチャネルとなります。既存の防御機能は、入力境界 (IBProtector、Llama Guard、パープレキシティ フィルター、SmoothLLM) のみを保護するか、不透明で確率的なプロバイダー側フィルターとしてアプリケーションの外部で実行されます。私たちは、このギャップがめったに測定されない結果をもたらしていることを示しています。8 つの攻撃ファミリー、5 つの防御、および 3 つのモデル バックエンドにわたる 2,100 件のトレース評価では、標準レポート (ツールおよびメモリ ポイズニングの攻撃成功率 0.000) では完全に安全であるように見える無防備なパイプラインは、その安全性がほぼ完全にクラウド プロバイダーのサーバー側フィルター (Azure GPT-5 の 60 ブロック中 54 ブロック) によるものであり、エージェント モデル独自の調整に静かに移行します。そのようなフィルターのないバックエンド。結果のみのレポートでは、この依存性が隠蔽されます。私たちは、あらゆるエージェント間のチャネルに情報のボトルネック ゲートを設ける、トレーニング不要の多層防御フレームワークである ChannelGuard を紹介します。それぞれが、類似性を埋め込むことで敵対的なフレーズ バンクに対してチャネル テキストをスコアリングし、LLM 呼び出しを追加せずにそれを決定的に通過、圧縮、またはブロックします。一方、アトリビューション メソッドは、どのレイヤーが各攻撃を阻止したかを記録します。 ChannelGuard のツール出力ゲートは、Azure GPT-5、Anthropic Sonnet 4.5、Anthropic Haiku 4.5 全体で同様に、アプリケーション層でツール ポイズニング 30 をブロックします。一方、無防備なパイプラインはバックエンド間で完全にシフトします。また、プロンプト インジェクション攻撃の成功率は半分に低下し (0.333 ~ 0.167)、GSM8K の精度は正確に維持されます (0.867)。ホワイトボックス適応言い換えはあらゆる埋め込みゲートを回避しますが、摂動と投票のベースラインの方が優れています。拡張付録には、ベースライン、アブレーション、スイープ、良性保存分析、および裁判官監査 (カッパ = 0.900) が追加され、総コストは 47.36 米ドルです。
原文 (English)
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel through which an adversary can smuggle instructions. Existing defenses guard only the input boundary (IBProtector, Llama Guard, perplexity filters, SmoothLLM) or run outside the application as opaque, stochastic provider-side filters. We show this gap carries a consequence rarely measured: on a 2,100-trace evaluation across eight attack families, five defenses, and three model backends, an undefended pipeline that appears fully safe under standard reporting (attack success 0.000 on tool- and memory-poisoning) owes that safety almost entirely to the cloud provider's server-side filter (54 of 60 blocks on Azure GPT-5), and silently shifts to the agent model's own alignment on a backend without such a filter. Outcome-only reporting hides this dependence. We present ChannelGuard, a training-free defense-in-depth framework placing information-bottleneck gates on every inter-agent channel; each scores channel text against an adversarial phrase bank by embedding similarity and deterministically passes, compresses, or blocks it, adding no LLM call, while an attribution method records which layer stopped each attack. ChannelGuard's tool-output gate blocks Tool Poisoning 30 of 30 at the application layer, identically across Azure GPT-5, Anthropic Sonnet 4.5, and Anthropic Haiku 4.5, whereas the undefended pipeline shifts entirely across backends; it also lowers Prompt Injection attack success by half (0.333 to 0.167) and preserves GSM8K accuracy exactly (0.867). White-box adaptive paraphrase evades every embedding gate, where a perturb-and-vote baseline does better. An extended appendix adds baselines, ablations, sweeps, a benign-preservation analysis, and a judge audit (kappa = 0.900), at a total cost of 47.36 USD.
BRIM: ワークロードバランスのとれた両面ビットシリアルスパース推論アクセラレータ
ビット シリアル アクセラレータは、ビットレベルのスパース性を利用して DNN 推論コストを削減しますが、既存の設計では 1 つのオペランドのみでスパース性を利用しており、高速化には限界があります。スパース性の活用を両方のオペランドに同時に拡張すると、部分的な製品の削減が複合的に行われますが、ワークロードの不均衡という重大な新たなボトルネックが発生します。各同時重み付けとアクティベーションのペアの実行コストは、独立して変化する 2 つのオペランドのゼロ以外のビット数の積に依存するため、一緒に完了する必要があるペアは大幅に異なる時間に終了し、より高速な計算がアイドル状態になります。これにより、既存の両面設計では PE の使用率が 56 ~ 64% に制限されることがわかります。我々は、このボトルネックを直接ターゲットとする、ハードウェアとソフトウェアが共同設計した両面ビットシリアル スパース アクセラレータである BRIM を紹介します。 BRIM は、次の 2 つの統合メカニズムを組み合わせています。1) サイクリック バランス プルーニング (CBP)。プロファイルされたアクティベーション統計に基づいて重み表現を再形成し、オフラインで同時に処理されるペア間で予想されるワークロードを均等化する、トレーニング後の重み最適化。 2) ペアワイズ スロット ドネーション。これは、無視できる領域のオーバーヘッドで残留ランタイムの不均衡を吸収する軽量のハードウェア メカニズムです。等面積制約の下で CNN、ViT、LLM にわたって評価された BRIM は、以前の両面設計と比較して 90% 以上の PE 使用率、最大 2.37 倍の速度向上、最大 1.63 倍のエネルギー効率の向上を達成します。
原文 (English)
BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator
Bit-serial accelerators exploit bit-level sparsity to reduce DNN inference cost, but existing designs exploit sparsity on only one operand, bounding the speedup. Extending sparsity exploitation to both operands simultaneously yields compounding reductions in partial products but introduces a critical new bottleneck: workload imbalance. Because each concurrent weight - activation pair's execution cost depends on the product of two independently varying operand non-zero bit counts, pairs that must complete together finish at vastly different times, leaving faster computations idle. We show this limits PE utilization to 56 - 64% in existing dual-sided designs. We present BRIM, a hardware - software co-designed dual-sided bit-serial sparse accelerator that directly targets this bottleneck. BRIM combines two integrated mechanisms: 1) Cyclic-Balanced Pruning (CBP), a post-training weight optimization that reshapes weight representations based on profiled activation statistics to equalize expected workloads across concurrently processed pairs offline; and 2) Pairwise Slot Donation, a lightweight hardware mechanism that absorbs residual runtime imbalance with negligible area overhead. Evaluated across CNNs, ViTs, and LLMs under iso-area constraints, BRIM achieves over 90% PE utilization, up to 2.37x speedup, and up to 1.63x energy efficiency improvement over prior dual-sided designs.
ChainWatch: MCP ベースの AI エージェント システムにおけるマルチステップ攻撃のためのキル チェーンに合わせた逐次検出フレームワーク
Model Context Protocol (MCP) は、AI エージェントが外部ツール、データベース、サービスに接続できるようにするオープンソース標準です。この接続により強力なエージェント機能が有効になりますが、既存の通話ごとの防御では確実に検出できない多段階攻撃も導入されます。攻撃者は、個別の無害なツールの呼び出しを、個別の検査を回避する悪意のあるシーケンスに組み込むことができます。この論文では、MCP ベースの AI エージェント システムにおける多段階攻撃を識別するための逐次検出フレームワークである ChainWatch について説明します。 ChainWatch は、6 段階のキル チェーンを使用して攻撃の進行をモデル化し、隠れマルコフ モデル (HMM) を適用してツール呼び出しシーケンスを分類します。セッションが複数の段階にわたって不審な進行を示した場合、検出ルールがトリガーされます。このフレームワークは、直接的な逐次攻撃、間接的なプロンプト インジェクション チェーン、およびハイブリッド多段階攻撃をカバーする構造化された脅威モデルによってサポートされています。 20 次元の特徴抽出スキーマは、ツールの相互作用から動作信号を捕捉します。セキュリティ文献からの 5 つの代表的な攻撃シナリオを使用してアプローチを実証し、従来の呼び出しごとのセキュリティ メカニズムを回避する攻撃チェーンを ChainWatch がどのように検出するかを示します。
原文 (English)
ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems
The Model Context Protocol (MCP) is an open-source standard that allows AI agents to connect to external tools, databases, and services. While this connectivity enables powerful agent capabilities, it also introduces multi-step attacks that existing per-call defenses cannot reliably detect. Attackers can compose individually benign tool invocations into malicious sequences that evade isolated inspection. This paper presents ChainWatch, a sequential detection framework for identifying multi-step attacks in MCP-based AI agent systems. ChainWatch models attack progression using a six-stage kill chain and applies a Hidden Markov Model (HMM) to classify tool-call sequences. Detection rules are triggered when a session exhibits suspicious progression across multiple stages. The framework is supported by a structured threat model covering direct sequential attacks, indirect prompt injection chains, and hybrid multi-stage attacks. A 20-dimensional feature extraction schema captures behavioral signals from tool interactions. We demonstrate the approach using five representative attack scenarios from the security literature, showing how ChainWatch detects attack chains that evade traditional per-call security mechanisms.
自律型商取引における信頼の構築: 検証可能なグローバル イベント タイムラインと AI 対応の詐欺インテリジェンス レイヤー
AP2 や ACP などのエージェント コマース プロトコルは、エージェントが開始する安全なトランザクションのメカニズムを定義しますが、異種ドメイン間での相互運用性、改ざん防止監査機能、または検証可能なイベントの時間的順序付けを提供しません。この論文は、決定論的なシリアル化を強制する正規イベント スキーマ、同期されたクロックに依存せずに再現可能な順序付けを保証する決定論的なバッチ形成、対数コストの包含証明を提供するマークルベースの追加専用コミットメント、および改ざん明示的な時間バックボーンを確立するブロックチェーン アンカリングという 4 つのコア コンポーネントで構築された、エージェント コマースのための検証可能なグローバル イベント タイムラインを提案することでこれらのギャップに対処します。このインフラストラクチャを基盤として、偽造不可能な来歴チェーンを通じてリスクラベルをアンカーされた証拠に結び付ける暗号署名された不正マーカーと、再現可能で改ざん明示的な AI トレーニング パイプラインを可能にするデータセット系統モデルを導入します。プロトタイプの実装による実験結果は次のことを示しています。マークル ツリー構築では 50,000 のイベントが 47 ミリ秒で処理されます。エンドツーエンドの検証は、バッチ サイズに関係なく 0.013 ミリ秒未満で完了します。包含証明のサイズは、1,000 イベントの 320 バイトから 50,000 イベントの 512 バイトまで対数的に増加します。また、マークルベースの検証は、50,000 イベントでリニア スキャンを 14.4 倍上回ります。
原文 (English)
Building Trust in Autonomous Commerce: A Verifiable Global Event Timeline and AI-Ready Fraud Intelligence Layer
Agentic commerce protocols such as AP2 and ACP define mechanisms for secure agent-initiated transactions but do not provide interoperable, tamper-evident auditability or verifiable temporal ordering of events across heterogeneous domains. This paper addresses these gaps by proposing a verifiable global event timeline for agentic commerce, constructed from four core components: canonical event schemas that enforce deterministic serialization, deterministic batch formation ensuring reproducible ordering without reliance on synchronized clocks, Merkle-based append-only commitments providing logarithmic-cost inclusion proofs, and blockchain anchoring establishing a tamper-evident temporal backbone. Building on this infrastructure, we introduce a cryptographically signed fraud marker that binds risk labels to anchored evidence through an unforgeable provenance chain, and a dataset lineage model enabling reproducible, tamper-evident AI training pipelines. Empirical results from a prototype implementation demonstrate: Merkle tree construction processes 50,000 events in 47 milliseconds; end-to-end verification completes in under 0.013 milliseconds regardless of batch size; inclusion proof sizes grow logarithmically from 320 bytes at 1,000 events to 512 bytes at 50,000 events; and Merkle-based verification outperforms linear scan by 14.4x at 50,000 events.
BaseRT: Apple M5 ニューラル アクセラレータによるクラス最高の LLM 推論の進歩
Apple の M5 世代では、すべてのコアが専用のニューラル アクセラレータ、つまり Metal~4 テンソル API を通じて公開されるオンダイ マトリックス ユニットを搭載する、再設計された GPU アーキテクチャが導入されています。 Apple Silicon 上の大規模言語モデル用のネイティブ Metal 推論ランタイムである BaseRT がこれらのユニットを活用して、Apple ハードウェア上の推論スループットを llama.cpp と MLX の両方を大幅に上回ることを示します。 BaseRT のフレームワークフリー設計を基盤として、既存の特殊なカーネル上にメモリに束縛されたデコード パスを残しながら、M5 ニューラル アクセラレータを介して推論の計算束縛行列乗算をルーティングする、手書きの Metal~4 テンソル コア カーネル ファミリ (高密度および専門家混合 GEMM およびフラッシュ アテンション プリフィル カーネルを含む) を追加します。 Apple M5 Pro では、Qwen3、Qwen3.5/3.6、Llama~3.2、および Gemma~4 ファミリのサブ 1B から 35B パラメータにわたる 15 のモデル構成にわたって、BaseRT は、llama.cpp よりも最大 6.4 倍高いプロンプト処理スループットを実現し、MLX よりも 3.9 倍高いプロンプト処理スループットを実現し、専門家混合モデルで最大のマージンを実現します。ここでは行列乗算が優勢ですが、デコードでは llama.cpp に対して最大 $1.75\times$、MLX に対して $1.33\times$ のリードを維持しています。これらの結果は、オンデバイス LLM 推論の新しいパフォーマンス上限を確立し、M5 のテンソル コアが Apple Silicon での迅速な処理の決定的な手段であることを示しています。 BaseRT は https://github.com/basecompute/baseRT で公開されています。
原文 (English)
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators
Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API. We show that BaseRT, our native Metal inference runtime for large language models on Apple Silicon, exploits these units to push inference throughput on Apple hardware substantially beyond both llama.cpp and MLX. Building on BaseRT's framework-free design, we add a family of hand-written Metal~4 tensor-core kernels (including dense and mixture-of-experts GEMM and flash-attention prefill kernels) that route the compute-bound matrix multiplications of inference through the M5 Neural Accelerators while leaving the memory-bound decode path on our existing specialised kernels. On an Apple M5 Pro, across fifteen model configurations spanning the Qwen3, Qwen3.5/3.6, Llama~3.2, and Gemma~4 families from sub-1B to 35B parameters, BaseRT delivers up to $6.4\times$ higher prompt-processing throughput than llama.cpp and $3.9\times$ higher than MLX, with the largest margins on the mixture-of-experts models where matrix multiplication dominates, while maintaining its lead on decode of up to $1.75\times$ over llama.cpp and $1.33\times$ over MLX. These results establish a new performance ceiling for on-device LLM inference and show that the M5's tensor cores are the decisive lever for prompt processing on Apple Silicon. BaseRT is publicly available at https://github.com/basecompute/baseRT.
流通の回復としてのアンラーニング: 管理された反事実研究、検証済みの選択スクリーニング、およびOracle-Free認定の限界
機械のアンラーニングは通常、再トレーニングされたオラクルとトレーニングされたプローブを照合することによって評価されます。一致する再トレーニング参照を備えた管理された非事実テストベッドでは、この基準が保持された知識を保持する方法を優先できることがわかります。評価基準が適切なスコアを評価した候補者は保持された事実を忘れ、未学習レベル (クラスター CI $[-3.16,-2.48]$) を $-2.82$ nats 下回ります。私たちは、アンラーニングを、5 つのオープン アーキテクチャ ファミリにまたがる 45 のモデル シード セルにわたって、一致する参照および監査オラクル フリーの画面と証明書スタイルの基準への復元として再構築しました。参照自体は絶対的な保持/ラウンドトリップ証明書を偽装しています。つまり、構築によって保持セットを保持する挿入されたモデルは、41/45 セルで固定保持しきい値を満たさず、31/45 セルで自身のラウンドトリップに失敗し、参照は 1/45 でのみ完全に証明されます。ベースアンカー付きホールドアウトスクリーンは、選択的必要テストとして強力なままです。シールされたチャレンジスイートでは、45/45 セルで注入されたモデルを拒否し、44/45 セルで参照を受け入れ、エンティティルーティング抑制 (35/45) を部分的に検出します。これは感度を測定した必要な検査であり、十分性を証明するものではありません。リファレンス自体の動作点に基づいた損傷相対再校正により、15/45 セルの小さなサブセットが保証されます。棄権しない場合、そのピックは最適化する軸上の再トレーニング ノイズ (0.80 nats) 内にありますが、一般的なトレーニング済みプローブ基準は 5.17 nats 離れています (直接のベンチマークではなく、補助的な比較です)。固定規模のロジット抑制攻撃は、12/45 セルの前方バッテリーを完全に破壊するため、前方のみの認証は適切ではありません。私たちの方法は、生成されたままの方法に対する経験的な選択テストです。識別可能性定理は、予測される境界ケースとして TOFU を使用して、どの事実がオラクルフリーの忘却閾値をまったく認めるかを限界づけます。
原文 (English)
Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification
Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes. In a controlled nonce-fact testbed with a matched retraining reference, we find this criterion can favor methods that retain held-out knowledge: candidates it rates adequate score held-out forget facts $-2.82$ nats below the never-learned level (cluster CI $[-3.16,-2.48]$). We recast unlearning as restoration to the matched reference and audit oracle-free screens and certificate-style criteria across 45 model-seed cells spanning five open architecture families. The reference itself falsifies an absolute retain/round-trip certificate: the injected model, which retains the retain set by construction, fails the fixed retain threshold in 41/45 cells and its own round trip in 31/45, and the reference fully certifies in only 1/45. A base-anchored held-out screen remains strong as a selective necessary test: on a sealed challenge suite it rejects the injected model in 45/45 cells, accepts the reference in 44/45, and partially detects entity-routing suppression (35/45); it is a necessary test with measured sensitivity, not a sufficiency certificate. A damage-relative recalibration anchored to the reference's own operating point certifies a small subset in 15/45 cells; where it does not abstain, its picks lie within retraining noise (0.80 nats) on the axes it optimizes, while the common trained-probe criterion sits 5.17 nats away (a supporting comparison, not a head-to-head benchmark). A fixed-magnitude logit-suppression attack defeats the full forward battery in 12/45 cells, so forward-only certification is not sound; our method is an empirical selective test for methods-as-produced. An identifiability theorem delimits which facts admit an oracle-free forget threshold at all, with TOFU as the predicted boundary case.
スケープゴートとしてのガードレール: ツールで強化された LLM エージェントにおける不誠実な安全拒否の監査
ツールで強化された LLM エージェントの評価フレームワークは、機能メトリクスまたは明示的なツールのクラッシュに圧倒的に焦点を当てており、サイレント インフラストラクチャの障害や、空、null、または不正な形式のペイロードを含む HTTP 200 応答はほとんど監査されません。軽量のブラックボックス監査フレームワークを導入します。このフレームワークは、12 の運用環境に隣接するツール スタブに 4 つのサイレント障害プロファイルを挿入し、エージェントの応答を 3 つの相互に排他的な動作クラス、Honest Surrender (HSR)、 Fabrication (FAR)、および Unfaithful Safety Refusal (USR) に分類します。中立システム プロンプトの下、温度ゼロで 2 つのフロンティア モデルと 2 つのオープンソース モデルを評価したところ、FAR が優勢であることがわかりました (有効回答の 56.6%)。エージェントは空のペイロードを実際のデータとして扱い、黙って捏造された結果を返します。エージェントが失敗を説明するためにポリシーやプライバシーの根拠を発明する USR は、ベースラインではほとんど存在しません (0.25%、396 の有効な軌跡全体で 1 つのインスタンス)。私たちの重要な発見は、システムのプロンプトに標準的な安全性の文言 (「ユーザーのプライバシーとデータのセキュリティを優先する」) を追加したアブレーションから明らかになり、これにより USR が 15.6 倍に増幅されます (0.25% から 3.95%、アブレーション率の 95% CI: 2.2% ~ 6.4%、フィッシャーの直接確率検定、p < 0.001)。 USR は潜在的な動作であり、ツールがサイレントに失敗したときに、システム プロンプト内の安全ボキャブラリーがモデルにポリシーの理論的根拠に到達するよう促すときにアクティブになります。機密性の高いツール (fetch_medical_record、retrieve_contract、fetch_user_profile) が USR インスタンスの大部分を占めます。実稼働レベルの検出のためのペイロード応答の不整合ヒューリスティックを提案し、セーフティフォワード展開におけるガバナンスへの影響について議論します。
原文 (English)
Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely unaudited. We introduce a lightweight black-box auditing framework that injects four silent failure profiles across 12 production-adjacent tool stubs and classifies agent responses into three mutually exclusive behavioral classes: Honest Surrender (HSR), Fabrication (FAR), and Unfaithful Safety Refusal (USR). Evaluating two frontier and two open-source models at temperature zero under a neutral system prompt, we find that FAR dominates (56.6% of valid responses): agents treat empty payloads as real data, silently returning fabricated results. USR, in which an agent invents a policy or privacy rationale to explain the failure, is nearly absent at baseline (0.25%, one instance across 396 valid trajectories). Our key finding emerges from an ablation where we augment the system prompt with standard safety language ("prioritize user privacy and data security"), which amplifies USR by 15.6x (from 0.25% to 3.95%; 95% CI on ablation rate: 2.2%-6.4%; Fisher's exact test, p < 0.001). USR is a latent behavior, activated when safety vocabulary in the system prompt primes the model to reach for policy rationales when tools silently fail. Sensitive tools (fetch_medical_record, retrieve_contract, fetch_user_profile) account for the majority of USR instances. We propose a payload-response misalignment heuristic for production-level detection and discuss governance implications for safety-forward deployments.
REGEN: オフライン強化学習によるエキスパートからゼネラリストへの蒸留のためのリプレイリサイクル
大規模なオンライン強化学習 (RL) は、大規模言語モデル (LLM) での長期推論やエージェント ツールの使用などの高度な能力を引き出す主な手段です。ただし、特に RL を単なる 1 回限りの学習段階として考える場合、関心のある広大なタスク領域にわたってそれを拡張し続けることは、計算インフラストラクチャとコストの両方の点で依然として困難です。最近、さまざまなドメインやトレーニング段階にわたって知識を蒸留するために広く使用されている手法であるマルチ教師オンポリシー蒸留 (MOPD) は、広大なドメインにわたって汎用性を維持しながら、RL 段階を分離してコストを節約するのに役立ちます。それにもかかわらず、オンライン RL と同様に、MOPD は結合された推論と逆方向パスを必要とするため、そのスケーラビリティと計算効率が引き続き制限されます。これらの課題に対処するために、私たちは REGEN: オフライン RL を使用したエキスパートからジェネラリストへの蒸留のためのリプレイ リサイクルを提案します。 REGEN は、複数の教師モデルから抽出するのではなく、教師の専門的な RL トレーニングの無料の副産物であるリプレイ メモリをリサイクルし、オフライン RL アルゴリズムを採用するだけでジェネラリストをトレーニングします。 REGEN は、ロールアウト サンプリングをバックワード トレーニング プロセスから完全に切り離すため、トレーニング コストを大幅に削減します。 REGEN は、数学的推論、コード生成、および命令追従全体にわたって、大幅に低コストで MOPD の精度と同等の精度を実現します。オンライン RL を 1 回限りの学習段階ではなくデータ合成プロセスに変える可能性があり、大きな計算負荷を必要とせずに大規模なポストトレーニングに拡張できる可能性があります。
原文 (English)
REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models (LLMs). However, continuing to scale it across vast task domains of interest remains challenging in both computational infrastructure and cost, especially when considering RL as merely a one-off learning stage. Recently, a widely used technique for distilling knowledge across various domains and training stages, multi-teacher on-policy distillation (MOPD), helps to decouple the RL stage, saving costs, while maintaining generality across vast domains. Nonetheless, similar to online RL, MOPD requires coupled inference and backward passes, which continues to limit its scalability and computational efficiency. To address these challenges, we propose REGEN: Replay-recycling for Expert-to-Generalist Distillation with Offline RL. Instead of distilling from multiple teacher models, REGEN trains a generalist by simply recycling the replay memory -- the free by-product of the teachers' specialized RL training -- and employing offline RL algorithms. REGEN completely decouples the rollout sampling from the backward training process and thus greatly reduces the training cost. Across mathematical reasoning, code generation, and instruction following, REGEN matches the accuracy of MOPD at substantially lower cost. It potentially turns online RL into a data synthesis process instead of a one-off learning stage, and can potentially be extended to large-scale post-training without requiring heavy computational load.
予測極値、不採算政策: ローソクベースのバイナンス スポット タイミング モデルの AI 支援監査
私たちは、ローソク足ベースの機械学習モデルが、仮想通貨の極値または短期的な結果の予測を、想定コスト後のポジティブなバイナンス スポット ペーパー ポリシーに変えることができるかどうかを監査します。数値結果は、スクリプト化された固定シード モデルの実行と決定論的シミュレーターから得られます。人間の監視下にある AI エージェントは、取引の決定ではなく、文献検索、個別に任務を負った批評、成果物の調整、文書化、およびソースのパッケージ化を通じて、7 月 20 日の証拠の完全性改訂をサポートしました。広範な先行検索を条件とした、後期期間の最も強力な証拠は否定的です。変更されていない 10 ペアの必須デイリー セレクターは、仮定された 31 bps の完了サイクル コストで、7 月 19 サイクルにわたって 6.72% 損失し、3 勝 16 敗でした。モデル固有の 7 月の簡単な評価では、検証で選択されたローカルミニマム ポリシーのリターンは -1.79\% でしたが、ローカルマキシマム販売/現金化/リエントリー ポリシーは継続保有を 2.80\% アンダーパフォームしました。総平均アドバンテージである 11.11 および 12.21 bps は、21 bps のストレスを下回っていました。 Gurgul にインスピレーションを得た OHLCV のみの毎日の適応では、ROC AUC の最小/最大 0.874/0.896 を達成しましたが、平均精度はわずか 0.134/0.116 で、7 サイクルで 44.30\% 損失しましたが、バイ アンド ホールドの場合は -41.20\% でした。フォレンジック監査では、以前の One4All の「30 日間ホールドアウト」も格下げされました。日付が以前のアーキテクチャ作業に影響し、4 時間の結果範囲が分割境界でパージされず、同一クローズ エントリが使用され、生の結果ディレクトリが存在しませんでした。テストされた、ほとんどが探索的なプロトコル全体で、イベント ランキングのパフォーマンスは、実行可能なポリシーの肯定的な価値を確立しませんでした。あらゆる運用上の決定は NO\_TRADE のままです。
原文 (English)
Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models
We audit whether candle-based machine-learning models can turn predictions of cryptocurrency extrema or short-horizon outcomes into positive Binance Spot paper policies after assumed costs. Numerical results come from scripted fixed-seed model runs and deterministic simulators; human-supervised AI agents supported the July 20 evidence-integrity revision through literature retrieval, separately tasked critique, artifact reconciliation, documentation, and source packaging, not trading decisions. The strongest later-period evidence, conditional on extensive predecessor search, is negative: an unchanged ten-pair mandatory-daily selector lost 6.72\% over 19 July cycles at an assumed 31-bps completed-cycle cost, with 3 wins and 16 losses. In short model-specific July evaluations, the validation-selected local-minimum policy returned -1.79\%, while the local-maximum sell-to-cash/re-entry policy underperformed continuous holding by 2.80\%; their gross mean advantages of 11.11 and 12.21 bps were below even the 21-bps stress. A Gurgul-inspired, OHLCV-only daily adaptation attained minimum/maximum ROC AUC of 0.874/0.896 but average precision of only 0.134/0.116 and lost 44.30\% over seven cycles, versus -41.20\% for buy-and-hold. A forensic audit also downgraded an earlier One4All "30-day holdout": its dates had influenced prior architecture work, its four-hour outcome horizon was not purged at split boundaries, it used same-close entry, and its raw result directories were absent. Across the tested, mostly exploratory protocols, event-ranking performance did not establish positive executable policy value. Every operational decision remains NO\_TRADE.
MoA 構造のデコード アテンション DNF 導出、KV キャッシュ蓄積、GQA/MQA、および OpenACC カーネル
配列の数学 (MoA) を使用して、トランスフォーマ アテンションのための 4 つのメモリ最適推論アーティファクトを導出します。それぞれは、現在のデコード ステップに固定されたクエリ行インデックスを使用したフォワードパス表示正規形 (DNF) から直接続きます。アーティファクトは次のとおりです: (1) $\psi$ リダクションにより $K^\top$ バッファが代数的に削除され、$(d_k + nd_k+ nd_v+ d_v)\times4\,{B}$ ダイナミック ランダム アクセス メモリ (DRAM) トラフィックの結果が数値的に検証される単一クエリ デコード DNF $\|{err}\|_\leq2\times10^{-7}$; (2)~$\|\mathrm{err}\|_\infty=0$ (正確な IEEE-754 浮動小数点演算) まで検証された、演算標準形式 (ONF) ストライド演算およびハードウェア結合メモリ アクセスを備えた C/OpenACC グラフィックス プロセッシング ユニット (GPU) カーネル。 (3)~MoA 連結 $\#$ を介してステップごとに $O(d_k+d_v)$ を追加するマルチステップ KV キャッシュ。 (4) ~ $\psi$ 選択によって導出されるグループ化クエリ アテンション (GQA) とマルチクエリ アテンション (MQA)。これにより、実証済みの KV トラフィックの $\frac {h_q} { h_{kv} }$ 削減が達成されます。すべてのプログラムは PyTorch のscaled_dot_product_attention に対して検証されます。
原文 (English)
MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel
We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the forward-pass Denotational Normal Form (DNF) of with the query-row index fixed to the current decode step. The artifacts are: (1)~a single-query decode DNF in which the $\psi$-reduction eliminates the $K^\top$ buffer algebraically, achieving $(d_k + nd_k+ nd_v+ d_v)\times4\,{B}$ Dynamic Random Access Memory (DRAM) traffic result numerically verified to $\|{err}\|_\leq2\times10^{-7}$; (2)~a C/OpenACC Graphics Processing Unit (GPU) kernel with Operational Normal Form (ONF) stride arithmetic and hardware-coalesced memory access, verified to $\|\mathrm{err}\|_\infty=0$ (exact IEEE-754 floating-point arithmetic); (3)~a multi-step KV-cache with $O(d_k+d_v)$ per-step append via MoA concatenation $\#$; and (4)~Grouped-Query Attention (GQA) and Multi-Query Attention (MQA) derived via $\psi$-selection, achieving a proven $\frac {h_q} { h_{kv} }$ reduction in KV traffic. All programs are verified against PyTorch scaled_dot_product_attention.
ModPack: 両手によるモバイル操作のための拡張可能な遠隔操作インターフェイス
既存の遠隔操作システムは多くの場合、特定のロボット ハードウェアやタスク ドメインに合わせて調整されており、拡張性や適応性が制限されています。我々は、統合フレームワーク内で多様なロボットの実施形態とタスク要件をサポートするように設計されたモジュール式で拡張可能な遠隔操作システムである ModPack を紹介します。 ModPack の中核となるのは、オンボードの計算、電力、通信、データ ストレージを統合する自己完結型のウェアラブル「バックパック」です。この共有インターフェイス上に構築されたシステムは、触覚フィードバックを備えた関節レベルの遠隔操作、モバイル操作、および能動的な知覚を含むプラグアンドプレイ機能モジュールをサポートします。 2 つの異なるロボット プラットフォームと現実世界のモバイル操作タスクにわたる実験により、ModPack がデータ収集とポリシー学習のための柔軟で再利用可能なフレームワークを提供することが実証されました。将来の研究をサポートするために、私たちは完全なハードウェア設計とソフトウェア スタックをオープンソースにします。プロジェクト Web サイト: https://modpack-robotics.github.io/
原文 (English)
ModPack: An Extensible Teleoperation Interface for Bimanual Mobile Manipulation
Existing teleoperation systems are often tailored to specific robot hardware and task domains, limiting their scalability and adaptability. We present ModPack, a modular and extensible teleoperation system designed to support diverse robot embodiments and task requirements within a unified framework. At the core of ModPack is a self-contained wearable "backpack" that integrates onboard computation, power, communication, and data storage. Built on top of this shared interface, the system supports plug-and-play capability modules including joint-level teleoperation with haptic feedback, mobile manipulation, and active perception. Experiments across two distinct robot platforms and real-world mobile manipulation tasks demonstrate that ModPack provides a flexible and reusable framework for data collection and policy learning. To support future research, we open-source the complete hardware design and software stack. Project website: https://modpack-robotics.github.io/
悪意のあるノード下でのピアツーピア分散 LLM 推論の整合性
ピアツーピア分散推論は、大規模言語モデル (LLM) の層を多くのノードに分散することにより、プールされた消費者向けハードウェア上で大規模言語モデル (LLM) を実行します。すべてのリクエストは、複数の独立した当事者によって所有および制御されているノードを通過します。ただし、この設定では、どの当事者もレイヤーの出力を改ざんして、最終結果を破損する可能性があります。信頼できるハードウェアでフォワード パスを再計算するとこれを検出できますが、追加の計算コストが発生します。科学文献には、画像分類器や暗号化コミットメントに対する既知の回答トラップなど、いくつかの以前の完全性チェック手法が含まれています。ただし、これらのソリューションは正確な正確性のみをテストし、良性のノード間で発生する可能性のある通常の変動を考慮していません。この論文では、各ノードが次のノードに渡すアクティベーションの変動を測定することによって出力の整合性をチェックする方法を提案します。ネットワークを使用したいピアは、正しいアクティベーションが事前にわかっている秘密のカナリア入力の小さなセットを選択し、それらを通常のトラフィックに混合します。ピアはカナリアと実際のクエリを区別できないため、ノードが改ざんされるとピアも破損します。既知の基準からの逸脱は、悪意のあるアクティビティを明らかにします。良性のノードは、ハードウェアに起因するノイズからわずかな変動しか示さないのに対し、改ざんされたノードは、はるかに大きく逸脱します。悪意のあるノードの特定は、固定しきい値に依存せずに、2 つのドリフト分布を分離する確率的テストとして扱われます。私たちは、実験を実行する前にメトリクスと成功基準を固定して 408 の構成を調査しました。検出器は AUROC 1.0 に到達し、すべての構成のすべてのカナリアで悪意のあるシャードをすべての良性シャードの上に正しくランク付けします。
原文 (English)
Integrity of peer-to-peer distributed LLM inference under malicious nodes
Peer-to-peer distributed inference executes a Large Language Model (LLM) on pooled consumer hardware by spreading its layers across many nodes. Every request passes through nodes that are owned and controlled by multiple independent parties. However, in this setting, any party can tamper with the output of its layers to corrupt the end result. Recomputing the forward pass on trusted hardware can catch this, but it introduces additional computational cost. The scientific literature includes several prior integrity-checking approaches, such as known-answer traps for image classifiers and cryptographic commitments. However, these solutions test only the exact correctness and do not account for the ordinary variation that may arise between benign nodes. In this paper, we propose a method that checks the output integrity by measuring the variation in the activations that each node passes to the next. A peer who wants to use the network selects a small set of secret canary inputs whose correct activations are known in advance and mixes them into regular traffic. Because the peers cannot tell a canary from a real query, any tampering node corrupts them as well. The deviation from the known reference then reveals malicious activity: benign nodes exhibit only minor variation from hardware-induced noise, whereas tampered nodes deviate far more. We treat the identification of malicious nodes as a probabilistic test that separates two drift distributions, without relying on a fixed threshold. We study 408 configurations with metrics and success criteria fixed before any experiment ran; the detector reaches AUROC 1.0, correctly ranking the malicious shard above every benign shard on every canary in every configuration.
ハイブリッド LLM ガイドによる量子貯蔵所アーキテクチャ設計の検索
量子リザーバー コンピューティング (QRC) は、高次元の時間特徴マップとして固定量子ダイナミクスを使用し、軽量の古典的な読み出しのみをトレーニングします。 QRC は短期的な量子機械学習にとって魅力的ですが、そのパフォーマンスは、入力エンコーディング、リザーバーの深さ、エンタングルメント トポロジー、測定特徴、状態リセット ポリシー、特徴構築、読み出し正則化などのアーキテクチャの選択に大きく依存します。 QRC 設計を制約付きブラックボックス アーキテクチャ検索として定式化し、大規模な言語モデルがこの検索問題の提案コントローラーとして機能できるかどうかを評価する、シミュレーター ベースのベンチマークである \method を紹介します。このベンチマークでは、同一の評価予算の下で 5 つのポリシー (ランダム検索、進化的検索、ベイジアン/TPE 最適化、フィードバックベースの LLM エージェント、LLM 提案とメモリ、突然変異、交差、重複回避、および探索を組み合わせた \hybrid) を比較します。 NARMA10、Mackey-Glass 予測、および時間的パリティに関しては、 \hybrid{} が最も一貫したポリシーです。NARMA10 と時間的パリティでは 1 位、Mackey-Glass では 2 位にランクされ、進化的探索に僅差で遅れています。 25 の評価予算と 3 つのシードの下では、\hybrid{} はすべてのタスクでランダム検索よりも改善され、Mackey-Glass エラーの相対的な 23.6\% 削減が含まれます。この結果は、LLM が汎用 QRC オプティマイザーであることを示していません。むしろ、生成モデルが、検証済みの再現可能なハイブリッド検索ループ内に埋め込まれた場合、有用な高レベルのコントローラーとなり得ることを示しています。
原文 (English)
Hybrid LLM-Guided Search for Quantum Reservoir Architecture Design
Quantum reservoir computing (QRC) uses fixed quantum dynamics as a high-dimensional temporal feature map and trains only a lightweight classical readout. QRC is attractive for near-term quantum machine learning, but its performance depends strongly on architecture choices such as input encoding, reservoir depth, entanglement topology, measurement features, state-reset policy, feature construction, and readout regularization. We introduce \method, a simulator-based benchmark that formulates QRC design as constrained black-box architecture search and evaluates whether large language models can act as proposal controllers for this search problem. The benchmark compares five policies under identical evaluation budgets: random search, evolutionary search, Bayesian/TPE optimization, a feedback-based LLM agent, and \hybrid, which combines LLM proposals with memory, mutation, crossover, duplicate avoidance, and exploration. On NARMA10, Mackey-Glass forecasting, and temporal parity, \hybrid{} is the most consistent policy: it ranks first on NARMA10 and temporal parity and second on Mackey-Glass, narrowly behind evolutionary search. Under a 25-evaluation budget and three seeds, \hybrid{} improves over random search on all tasks, including a 23.6\% relative reduction in Mackey-Glass error. The results do not show that LLMs are universal QRC optimizers; rather, they show that generative models can be useful high-level controllers when embedded inside validated, reproducible hybrid search loops.
SynPre-FL: 合成データ駆動型の事前トレーニングに統合された Federated Learning トレーニング フレームワーク
フェデレーション ラーニング (FL) は、プライバシーを保護しながら臨床リスクを予測するための有望なアプローチを提供しますが、データ共有の制限、クライアントの異質性、クラスの不均衡、および現実的な表形式の電子医療記録 (EHR) ベンチマークの欠如により、その展開は依然として制限されています。合成データの生成はデータ不足を軽減する可能性がありますが、フェデレーション最適化との統合は体系的な研究が限られています。我々は、非 IID 条件下でロバストな予測を行うために、高忠実度の合成 EHR 生成と合成事前学習済み FL を組み合わせた統合フレームワークである SynPre-FL を提案します。潜在的なオートエンコーダー拡散モデルは、フェデレーション トレーニングをウォーム スタートするために使用される、プライバシーを保護する合成コホートを生成します。この事前トレーニングの後に、クラスバランスのとれたローカル目標、近位正則化、および適応型サーバー集約を使用した異質性を意識した最適化が続きます。事後キャリブレーションとフェデレーテッドセーフの説明可能性により、信頼性が高く解釈可能なリスク推定がサポートされます。実験によれば、合成ジェネレーターは、メンバーシップ推論および再構成攻撃から保護しながら、一変量、二変量、および多変量の構造を保存します。生成されたデータは、TSTR、TRTS、およびモデルベースの評価の下で強力な下流の有用性を実現します。 SynPre-FL は、5、10、および 15 の異種クライアントによるフェデレーション設定全体で、特に深刻な非 IID フラグメンテーションの下で、ベースライン方法と比較して堅牢性とスケーラビリティを一貫して向上させます。キャリブレーションにより確率の信頼性が向上する一方、SHAP 分析により、連合規模全体にわたって安定した臨床的に一貫した特徴属性が生成されます。したがって、SynPre-FL は、合成データを FL と組み合わせるための実用的で再現可能なフレームワークを提供し、分散された表形式の EHR データからプライバシーを意識した、解釈可能で堅牢な臨床予測を可能にします。
原文 (English)
SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework
Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks. Synthetic data generation may alleviate data scarcity, yet its integration with federated optimisation has received limited systematic study. We propose SynPre-FL, a unified framework combining high-fidelity synthetic EHR generation with synthetic-pretrained FL for robust prediction under non-IID conditions. A latent autoencoder-diffusion model generates privacy-preserving synthetic cohorts, which are used to warm-start federated training. This pretraining is followed by heterogeneity-aware optimisation using class-balanced local objectives, proximal regularisation, and adaptive server aggregation. Post-hoc calibration and federated-safe explainability support reliable and interpretable risk estimates. Experiments show that the synthetic generator preserves univariate, bivariate, and multivariate structure while protecting against membership-inference and reconstruction attacks. The generated data achieve strong downstream utility under TSTR, TRTS, and model-based evaluations. Across federated settings with 5, 10, and 15 heterogeneous clients, SynPre-FL consistently improves robustness and scalability over baseline methods, especially under severe non-IID fragmentation. Calibration improves probability reliability, while SHAP analysis produces stable and clinically coherent feature attributions across federation sizes. SynPre-FL therefore provides a practical and reproducible framework for combining synthetic data with FL to enable privacy-aware, interpretable, and robust clinical prediction from distributed tabular EHR data.
D3VL: 言語モデルを使用した 3D 時系列データとビデオからの運転シーンの理解
マルチモーダル大規模言語モデル (MLLM) の最近の進歩により、自動運転用のエンドツーエンド MLLM の開発が始まりました。ただし、これまで主に 2D 画像とビデオを使用する MLLM に重点が置かれてきました。対照的に、このペーパーでは、3D センサー、特に LiDAR とステレオ カメラを使用した MLLM の有効性について検討します。 LiDAR は、主にデータの希薄性とデータのグリッド構造の欠如により、MLLM 内での統合に特有の課題を抱えています。同様の理由で、MLLM パイプライン内でのカメラ データと LiDAR データの融合も一般的ではありません。ただし、ほとんどの自律システムは LiDAR ベースのセンシングに依存しており、3D データを組み込むことで従来の 3D シーン認識タスクのパフォーマンスが向上することが証明されています。このペーパーでは、2D および 3D 時系列データを単一のシンプルなアーキテクチャに統合する新しい MLLM フレームワークである D3VL について説明します。このモデルは、交通現場の理解と安全に関する質問に答えることを目的としています。 D3VL は、2D および 3D 時系列データの処理において、ベースライン手法と比較して、KITTI 質問応答 (QA) データセットで 11% の改善を示しています。このホワイト ペーパーでは、さまざまな運転条件下で 3D および時系列データを処理するモデルの能力を評価する、Waymo QA データセット拡張機能についてさらに紹介します。 D3VL 実装コードと WaymoQA 拡張機能は、補足 Web サイト: https://automotivesafety-lvlm.github.io にあります。
原文 (English)
D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the main emphasis to date has been for MLLMs using 2D images and videos. In contrast, this paper considers MLLM effectiveness using 3D sensors, particularly LiDAR and stereo cameras. LiDAR presents unique challenges to integration within an MLLM, largely because of data sparsity and lack of a grid structure for the data. For similar reasons, fusion of camera and LiDAR data within an MLLM pipeline is also uncommon. However, most autonomous systems rely on LiDAR-based sensing, and incorporating 3D data has been proven to improve performance in traditional 3D scene perception tasks. This paper presents D3VL, a novel MLLM framework that integrates 2D and 3D time-series data in a single but simple architecture. The model aims to answer questions involving traffic scene understanding and safety. D3VL shows an 11% improvement in the KITTI Question-Answering (QA) dataset compared to baseline methods in processing 2D and 3D time-series data. This paper further introduces the Waymo QA dataset extension, which assesses models' capabilities in processing 3D and time-series data under diverse driving conditions. D3VL implementation code and WaymoQA extension can be found on our supplemental website: https://automotivesafety-lvlm.github.io
パーソナライズされた乳がん予測のための信頼できるプライバシー保護マルチモーダルフェデレーテッド ラーニング
フェデレーテッド ラーニングは、特に個別化されたがん治療において、予測モデルのトレーニングに機密の健康データを使用することに関連するプライバシーの問題に対する潜在的な解決策として浮上しています。この研究では、フェデレーション ラーニングが、透明性、スケーラビリティ、セキュリティ、公平性という 4 つの重要な展開の柱に取り組みながら、乳がん患者の腫瘍進行を予測するための堅牢なモデルの開発をサポートできるかどうかを調査します。この研究では、臨床情報、腫瘍特性、バイオマーカー データ、患者人口統計などのマルチモーダル データと、MRI スキャンなどの医用画像データを使用して、時間の経過に伴う腫瘍特性の変化をモデル化する連合学習フレームワークを評価します。連合アプローチのパフォーマンスは、集約されたデータでトレーニングされた集中モデルのパフォーマンスと比較されました。このレポートではさらに、安全なモデルの更新を強化し、患者のサブグループ全体でパフォーマンスを維持し、施設全体の拡張性をサポートするための戦略を検討しています。この調査結果は、連合学習がデータの局所性を維持しながら集中学習と同等の予測パフォーマンスを達成できるかどうかを評価します。これらの結果は、プライバシーを保護するマルチモーダル予測モデリングの実現可能性の理解に貢献し、臨床医や患者の個別化された治療計画を支援するデジタルツインなどの将来のアプリケーションをサポートします。
原文 (English)
Trustworthy Privacy-Preserving Multimodal Federated Learning for Personalised Breast Cancer Prediction
Federated learning has emerged as a potential solution to privacy concerns associated with using sensitive health data for training predictive models, particularly in personalised cancer care. This research investigates whether federated learning can support the development of robust models for predicting tumour progression in breast cancer patients while addressing four critical deployment pillars: transparency, scalability, security, and fairness. This study evaluates a federated learning framework using multimodal data, including clinical information, tumour characteristics, biomarker data, and patient demographics, alongside medical imaging data such as MRI scans, to model changes in tumour characteristics over time. The performance of the federated approach was compared with that of a centralised model trained on aggregated data. The report then further examines strategies to enhance secure model updates, maintain performance across patient subgroups, and support scalability across institutions. The findings assess whether federated learning can achieve predictive performance comparable to centralised learning while preserving data locality. These results contribute to understanding the feasibility of privacy-preserving, multimodal predictive modelling and support future applications such as digital twins to assist clinicians and patients in personalised treatment planning.
タイルレベルのシグナリングと専門家混合のスケジューリングによるきめ細かい計算と通信のオーバーラップ
Mixture-of-Experts (MoE) アーキテクチャは、計算コストを比例的に増加させることなくモデルの能力を向上させ、大規模言語モデル (LLM) を兆パラメータ領域に拡張するための重要な構成要素となっています。これらの MoE モデルの効率的な展開は、複数の GPU にわたる分散実行に依存します。各 MoE レイヤーには、トークンをエキスパート ランクにディスパッチし、エキスパートの出力をソース ランクに返すという 2 つの全対全通信が含まれます。従来の MoE 実装では、エキスパート コンピューティングの完了後にこのオールツーオール リターンが開始され、クリティカル パスでの通信遅延が発生し、GPU 使用率が削減されます。私たちは、タイルレベルのシグナリングとスケジューリングを介して、エキスパート コンピューティングと 2 番目の全対全をオーバーラップさせる、きめ細かいアプローチを提案します。当社のプロデューサーとコンシューマーの共同設計は、(1) ランク上のすべてのローカル エキスパートをカバーしてカーネル起動の繰り返しのオーバーヘッドを排除し、リモート クリティカルなタイルを優先する永続的なランクごとの計算カーネル (プロデューサー)、(2) タイルの準備ができたときにセグメント単位の転送を発行するストリーミング マルチプロセッサ (SM) の小さな専用パーティション上の永続的な通信カーネル (コンシューマー) を組み合わせています。私たちの共同設計は、基礎となる計算演算子や通信プリミティブへの侵入的な変更を回避し、マルチ GPU システムでの分散 MoE の実行効率を向上させるのに実用的です。 4-A100 GPU プラットフォームでは、4 つの最先端の MoE システムに対して 3 つの MoE モデルで評価され、当社のアプローチは最大 2.64 倍のエンドツーエンドの高速化と 2.74 倍の MoE レイヤの高速化を達成しました。従来の重複しないベースラインと比較して、私たちのアプローチは、正確さを維持しながら、さまざまな GEMM 形状、ルーター モード、および広範囲のプロデューサー/コンシューマー SM パーティションにわたって、オペレーター層と MoE 層の両方のレベルのパフォーマンスを一貫して向上させます。
原文 (English)
Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts
Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trillion-parameter regimes. Efficient deployment of these MoE models relies on distributed execution across multiple GPUs, where each MoE layer involves two all-to-all communications: dispatching tokens to expert ranks and returning the expert outputs to their source ranks. Conventional MoE implementations launch this return all-to-all after expert compute completes, exposing communication latency on the critical path and reducing GPU utilization. We present a fine-grained approach that overlaps expert compute with the second all-to-all via tile-level signaling and scheduling. Our producer-consumer co-design combines: (1) a persistent per-rank computation kernel (producer) that covers all local experts on the rank to eliminate repeated kernel launch overhead and prioritizes remote-critical tiles, and (2) a persistent communication kernel (consumer) on a small dedicated partition of streaming multiprocessors (SMs) that issues segment-granular transfers as tiles become ready. Our co-design avoids intrusive changes to the underlying computation operators or communication primitives, making it practical for improving distributed MoE execution efficiency on multi-GPU systems. On a 4-A100 GPU platform, evaluated on three MoE models against four state-of-the-art MoE systems, our approach achieves up to 2.64x end-to-end speedup and 2.74x MoE-layer speedup. Compared with a conventional non-overlap baseline, our approach consistently improves both operator- and MoE-layer-level performance across varying GEMM shapes, router modes, and a broad range of producer/consumer SM partitions, while preserving correctness.
橋家-東川地震間隙における浅層貯留層誘発地震と深部地殻ロックの並置
成熟した地震ギャップの臨界状態を特定することは、特に貯留層の貯留などの人為的応力摂動が地殻構造荷重に重なる場合には困難です。ここでは、橋家-東川地震間隙(世界で 2 番目に大きい水力発電所がある)の高解像度高密度アレイ カタログを利用して、明確な垂直デカップリング メカニズムを明らかにします。浅部の活動は高い b 値 (1.0) を示し、流体駆動の貯留層によって引き起こされる地震活動を示しています。逆に、深部の地震活動 (20 km) は、低い b 値 (0.8 未満) と高いクーロン応力蓄積率を特徴とする「ロックされたアスペリティ」の概要を示します。さらに、複雑な浸漬構造を特定し、複合断層運動学を示唆しています。さらに、計算された応力蓄積は、この地震ギャップが破壊の可能性が上昇した臨界状態にあることを示唆しています。私たちの発見は、浅部に誘発された地震活動が、深部の地殻変動の静かな蓄積を覆い隠してしまう可能性があることを示しています。このデカップリング モデルは、貯留層断層システムにおける地震リスクを世界的に評価するための新しい枠組みを提供します。
原文 (English)
Juxtaposition of Shallow Reservoir-Triggered Seismicity and Deep Tectonic Locking in the Qiaojia-Dongchuan Seismic Gap
Identifying the critical state of mature seismic gaps is challenging, especially when anthropogenic stress perturbations, such as reservoir impoundment, superimpose on tectonic loading. Here, utilizing a high-resolution dense array catalog from the Qiaojia-Dongchuan seismic gap (hosting the second-largest hydropower station in the world), we reveal a distinct vertical decoupling mechanism. The shallow activities exhibit high b-values (1.0), indicative of fluid-driven reservoir-triggered seismicity. Conversely, deep seismicity (20 km) outlines a 'locked asperity' characterized by low b-values (less than 0.8) and high Coulomb stress accumulation rate. We further identify a complex dipping structure, suggesting compound fault kinematics. Additionally, the calculated stress accumulation suggests this seismic gap is in a critical state with elevated rupture potential. Our findings indicate that shallow induced seismicity can mask the silent accumulation of deep tectonic strain. This decoupling model provides a new framework for assessing seismic risks in reservoir-fault systems globally.
因果辞書学習により、ゲノム言語モデルにおける転写因子結合特徴が明らかにおよび検証される
ゲノム言語モデルは、調節ゲノミクスタスク全体で強力なパフォーマンスを達成しますが、これらのモデルが内部的に何を表すかは依然として不透明であり、この分野には、モデル内の見かけの「概念」が配列構成の成果物ではなく本物であることを検証するための原則的な手順が欠けています。スパース辞書学習と因果的介入を組み合わせて、ゲノム基盤モデルの解釈可能な特徴を抽出、検証、因果的にテストするフレームワークを紹介します。構造的に異なる 2 つのモデル、Nucleotide Transformer ($6$-mer トークン化) と DNABERT-2 (バイトペア エンコーディング) の隠れた活性化でトップ $k$ のスパース オートエンコーダーをトレーニングすることで、転写因子 (TF) 配列モチーフにマッピングされる数千の単一意味特徴を回復しました。我々は、位置重み行列に対するそのような特徴の単純な検証が GC 組成と反復要素によってひどく混乱し、何百もの偽の「TF 特徴」を生成することを示し、これらの交絡を除去する、組成が一致し、結合が解決されたプロトコルを開発します。重要なことに、我々は相関関係を超えています。モデルのフォワードパス中に個々の辞書の方向を除去し、モデル自体の予測分布に誘発されたシフトを測定することによって、特定の特徴が単なるモチーフの存在ではなく、細胞型特異的なTF結合を表すために\emph{因果的に}使用されていることを確立します。 3 つの転写因子 (CTCF、GATA1、REST) および両方のアーキテクチャにわたって、因果関係が検証された結合特徴が再現性よく出現します (条件あたり $15$ でテストされた特徴のうち $7$ ~ $14$)。一方、2 クラスのネガティブ コントロール、スクランブル結合ラベルおよびランダムに選択された特徴では、検出可能なシグナルが得られません。このフレームワークは純粋に計算を目的としており、公開データのみを使用し、ゲノム深層学習における解釈可能性を主張するための再利用可能な標準を提供します。
原文 (English)
Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models
Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition. We introduce a framework that combines sparse dictionary learning with causal intervention to extract, validate, and causally test interpretable features in genomic foundation models. Training top-$k$ sparse autoencoders on the hidden activations of two architecturally distinct models, Nucleotide Transformer ($6$-mer tokenization) and DNABERT-2 (byte-pair encoding), we recover thousands of monosemantic features that map to transcription-factor (TF) sequence motifs. We show that the naive validation of such features against position weight matrices is severely confounded by GC composition and repetitive elements, producing hundreds of spurious ``TF features'', and we develop a composition-matched, binding-resolved protocol that removes these confounds. Critically, we move beyond correlation: by ablating individual dictionary directions during the model's forward pass and measuring the induced shift in the model's own predictive distribution, we establish that specific features are \emph{causally} used to represent cell-type-specific TF binding, not merely motif presence. Across three transcription factors (CTCF, GATA1, REST) and both architectures, causally validated binding features emerge reproducibly ($7$--$14$ of $15$ tested features per condition), while two classes of negative control, scrambled binding labels and randomly selected features, yield no detectable signal. The framework is purely computational, uses only public data, and provides a reusable standard for interpretability claims in genomic deep learning.
SCPP: ソフト クラスタリング用の統合 Python ライブラリ
このペーパーでは、ソフト クラスタリング用のオープンソース Python フレームワークである SCPP (Soft Clustering Python Package) を紹介します。 SCPP は、ファジー、確率的、グラフベース、行列因数分解、ディープ ラーニング手法など、異種ソフト クラスタリング手法全体でモデルのトレーニング、予測、メンバーシップ表現、評価、ベンチマークを標準化する、標準的な scikit-learn 互換の推定インターフェイスを確立します。このフレームワークは現在、40 の代表的なアルゴリズムと、データセット、クラスタリング品質メトリクス、標準化されたランタイム、メモリ、およびスケーラビリティ評価で構成される包括的なベンチマークを統合しています。 SCPP はさらに、広範なドキュメント、実践例、自動テスト、科学的な Python エコシステムとのシームレスな統合を提供し、再現可能な実験と新しいアルゴリズムによる直接的な拡張を可能にします。ソース コードは https://github.com/soft-clustering/soft-clustering で公開されています。
原文 (English)
SCPP: A Unified Python Library for Soft Clustering
In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering. SCPP establishes a canonical, scikit-learn-compatible estimator interface that standardizes model training, prediction, membership representation, evaluation, and benchmarking across heterogeneous soft clustering methods, including fuzzy, probabilistic, graph-based, matrix factorization, and deep learning methods. The framework currently integrates 40 representative algorithms together with a comprehensive benchmarking comprising datasets, clustering quality metrics, and standardized runtime, memory, and scalability evaluation. SCPP further provides extensive documentation, practical examples, automated testing, and seamless integration with the scientific Python ecosystem, enabling reproducible experimentation and straightforward extension with new algorithms. The source code is publicly available at https://github.com/soft-clustering/soft-clustering.
Federated Learning における開発者の課題を理解する: Stack Overflow と GitHub からの洞察
Federated Learning (FL) を使用すると、生データを一元管理せずに協調的なモデル トレーニングが可能になりますが、分散実行、急速に進化するフレームワーク、プライバシーとガバナンスの要件により、FL システムの構築と運用は依然として困難です。このペーパーでは、92 の FL 関連プロジェクトからの 495 件の Stack Overflow 投稿と 9,116 件の GitHub の問題およびプル リクエストを独自に分析することにより、FL 開発者の課題に関する実証的研究を紹介します。 BERTopic ベースのトピック モデリングと、未解決率や解決時間の中央値などの難易度指標を使用して、繰り返し発生する問題領域を特徴付け、Stack Overflow と GitHub の 2 つのサポート プラットフォーム間で問題がどのように現れるかを比較します。私たちの分析では、9 つの主要な Stack Overflow トピックと 13 の GitHub トピックが明らかになり、環境セットアップと依存関係の互換性、API の破損と移行、非 IID データ下でのトレーニングの不安定性、評価とメトリクスの正確性、プライバシー保護メカニズムの統合に集中した永続的な問題が発生しています。また、開発者が求めているヘルプの種類を理解するために、投稿を質問の意図ごとに分類しています。この意図分析では、手順に関するガイダンスに対する強い需要を反映して、「どのように」タイプの質問が優勢であることが示されています。 「TFF のインストールと環境の互換性」や「フェデレーション機能エンジニアリングと SecureBoost の問題」などのいくつかのトピックでは、未解決率が高く解決に時間がかかることが示されており、ツール、ドキュメント、およびデバッグ サポートに欠陥があることが示唆されています。これらの調査結果に基づいて、FL フレームワークの設計者、ドキュメント作成者、教育者に実用的な影響を提供します。私たちの結果は公開の議論と広く議論されているフレームワークのサブセットに限定されていますが、この研究は、開発者の課題を継続的に監視し、FL システムの使いやすさ、信頼性、展開可能性を向上させるためのスケーラブルな方法を提供します。
原文 (English)
Understanding Developer Pain Points in Federated Learning: Insights from Stack Overflow and GitHub
Federated Learning (FL) enables collaborative model training without centralizing raw data, but building and operating FL systems remains difficult due to distributed execution, rapidly evolving frameworks, and privacy and governance requirements. In this paper, we present an empirical study of FL developer challenges by independently analyzing 495 Stack Overflow posts and 9,116 GitHub issues and pull requests from 92 FL-related projects. Using BERTopic-based topic modeling and difficulty indicators such as unresolved rates and median resolution time, we characterize recurring problem areas and compare how they manifest across the two support platforms, Stack Overflow and GitHub. Our analysis surfaces nine dominant Stack Overflow topics and thirteen GitHub topics, with persistent difficulties concentrated in environment setup and dependency compatibility, API breakages and migration, training instability under non-IID data, evaluation and metric correctness, and the integration of privacy-preserving mechanisms. We also categorize posts by question intent to understand the kinds of help developers seek; this intent analysis shows that "How"-type questions dominate, reflecting strong demand for procedural guidance. Several topics, such as "TFF Installation and Environment Compatibility" and "Federated Feature Engineering and SecureBoost Issues," exhibit high unresolved rates and long resolution times, suggesting shortcomings in tooling, documentation, and debugging support. Based on these findings, we provide actionable implications for FL framework designers, documentation authors, and educators. Although our results are constrained to public discussions and a subset of widely discussed frameworks, the study offers a scalable method for continuously monitoring developer pain points and improving the usability, reliability, and deployability of FL systems.
適応型降伏: 脆弱性コンテキストにおける LLM 応答の構造的障害モード
感情的に敏感な文脈で動作する大規模な言語モデルは、構造的なトリレンマに直面しています。つまり、脆弱な状態にあるユーザーが不適応帰属を強化する可能性のある情報を要求した場合、現在の対応アーキテクチャは、保護的な制限、抑揚のない促進、または両方の命令の統合されていない共存によって緊張を解決します。それぞれが一方の目的を他方を犠牲にして維持します。 3 つの商用 LLM (マテリアル、リレーショナル、およびソマティック ステータス プロキシのバリアントにわたる 900 セッション) に 3 ターン エスカレートする脆弱性ビネットを管理し、2 つのバイナリ インデックス (VCC/VCI) で応答をコーディングすることで、適応的降伏と呼ばれるこれまで文書化されていなかった障害モードを特徴付けます。モデルは、ユーザーの苦痛の根底にある社会的不正義を検証してから、ユーザーの苦痛の獲得そのものの詳細な促進に軸足を移します。名目上は落胆している。私たちは、トリレンマが偶発的ではなく構造的なものであることを示し、最小限の再帰属性 (MRS) を提案します。これは、他の方法で検証する応答内に単一の再帰属性の手がかりを埋め込み、ユーザーが明示した目標に異議を唱えることなく自律的な再帰属性への道を維持する、アーキテクチャ中立的な設計原則です。
原文 (English)
Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts
Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that may reinforce maladaptive attribution, current response architectures resolve the tension through protective restriction, uninflected facilitation, or unintegrated co-presence of both imperatives -- each preserving one objective at the cost of the other. Administering a three-turn escalating vulnerability vignette to three commercial LLMs (900 sessions across material, relational, and somatic status-proxy variants) and coding responses with two binary indices (VCC/VCI), we characterize a previously undocumented failure mode we term adaptive capitulation: the model validates the social injustice underlying the user's distress before pivoting to detailed facilitation of the very acquisition it nominally discouraged. We show that the trilemma is structural rather than incidental, and propose Minimal Reattributive Sufficiency (MRS), an architecture-neutral design principle that embeds a single reattributive cue within an otherwise validating response, preserving a pathway toward autonomous reattribution without contesting the user's stated goal.
健全な神経推論器の構造: ワンショット償却、ファーストパスポイズニング、および手がかり豊富な完了における検索の不活性さ
ニューラル ソルバーは、中間状態を推定、分岐、修正するために構築されています。 Lattice Deduction Transformer (LDT) はまさにそれを行うようです。手がかりが豊富な Sudoku では、そうではありません。1 回のフォワード パスで基本的にグリッド全体 (標準 6x6 ではすべての空白セル、拡張 9x9 では 94 ~ 96%) がコミットされ、反復ソルバーが正確な検証器にラップされたワンショット予測器に変わります。すべてのハード スライスの失敗は、最初のパスで真のソリューションに必要な値が確実に削除される検索開始前に決定されます。これを初回通過中毒と呼びます。学習された分岐、MRV、バックトラッキング、値の除外、および共有不良 (CoLT) を追加しても、解決される Sudoku インスタンスは変わりません。繰り返される無効な導出を 1,497 分の 1 に削減します。凍結されたトレーニング予算では、制約グラフの注意力のみが完全な CoLT 精度と一致しますが、位置テーブルは大幅に長いトレーニングの下でのみ回復します。これは、絶対的な容量の違いではなく、最適化とサンプル効率の利点を示しています。この診断により、2 つの効果的な介入が予測されます。桁順列拡張により、対称性の素な分割における 3 つのトレーニング シード全体で 9x9 の精度が 1% 未満から 96.5 +/- 0.3 に向上しました。対称変換されたパスに対するテスト時のユニオンは、再トレーニングなしで 3 つのハードスライス チェックポイントすべてを 72.8 ~ 78.9% から 100% に引き上げます。スクラッチからグラフに色を付けると、ワンショット動作がなくなり、検索の精度が変わります。手がかりが豊富な補完では、LDT のようなシステムは、学習された検索手順ではなく、ワンショットで償却される予測子です。精度はキャリブレーションと対称性によって決まりますが、検索では主に計算の無駄が除去されます。
原文 (English)
Anatomy of a Sound Neural Reasoner: One-Shot Amortization, First-Pass Poisoning, and Search Inertness in Clue-Rich Completion
Neural solvers are built to deduce, branch, and revise intermediate states. The Lattice Deduction Transformer (LDT) appears to do exactly that. In clue-rich Sudoku, it does not: one forward pass commits essentially the entire grid (every blank cell on standard 6x6, 94-96% on augmented 9x9), turning the iterative solver into a one-shot predictor wrapped in an exact verifier. All hard-slice failures are decided before search begins, when the first pass confidently deletes a value required by the true solution. We call this first-pass poisoning. Adding learned branching, MRV, backtracking, value exclusion, and shared nogoods (CoLT) does not change which Sudoku instances are solved; it cuts repeated invalid derivations 1,497-fold. At the frozen training budget, constraint-graph attention alone matches full-CoLT accuracy, while positional tables recover only under substantially longer training, indicating an optimization and sample-efficiency advantage rather than an absolute capacity difference. The diagnosis predicts two effective interventions. Digit-permutation augmentation raises 9x9 accuracy from below 1% to 96.5 +/- 0.3 across three training seeds on a symmetry-disjoint split. Test-time union over symmetry-transformed passes raises all three hard-slice checkpoints from 72.8-78.9% to 100% without retraining. On from-scratch graph coloring, one-shot behavior disappears and search changes accuracy. In clue-rich completion, LDT-like systems are one-shot amortized predictors rather than learned search procedures: accuracy is determined by calibration and symmetry, while search primarily removes computational waste.
PerfAgent: リポジトリ レベルのコード最適化のためのプロファイラーに基づく反復的改良
大規模言語モデル (LLM) エージェントは、SWE-Bench の問題解決や実際のコードベースでの機能実装など、正確性を重視したリポジトリ レベルのタスクで適切に実行できるようになりました。ただし、実行時のパフォーマンスを向上させながら動作を維持する必要がある、リポジトリ レベルのコードの最適化には依然として苦労しています。この設定では、テストに合格するだけでは十分ではありません。パッチは動作を保持し、コードの最適化を実装し、エキスパートによる高速化に取り組む必要があります。現在のエージェントは、抽象化レイヤーやネイティブ拡張機能の背後に隠れているボトルネックを見逃したり、浅いスピードアップ後に停止したり、コード パッチのテストが不十分であったりして、エッジ ケースをひっそりと壊してしまう可能性があります。 PerfAgent は、プロファイラー主導の検証者インザループ ワークフローであり、実際のホットスポットを見つけ、最初に合格したパッチを超えて改善し、タイミングだけでなくプロファイラーの証拠を使用して次に何を最適化するかを決定するために必要なフィードバックを既製のコーディング エージェントに提供します。 GSO と SWE-fficiency-Lite という 2 つの困難な最適化ベンチマークにおいて、PerfAgent は GPT-5.1 を使用した OpenHands に比べてエキスパート マッチング パッチの割合を 2 倍以上に高め、GSO では 19.6% から 39.2% に、SWE-fficiency-Lite では 26% から 74% に向上しました。また、大幅に低コストで Oracle のベストオブ 5 ベースラインを上回っており、追加のテスト時間サンプリングではなく、より優れたフィードバックによって利益が得られていることを示しています。
原文 (English)
PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization
Large language model (LLM) agents now perform well on correctness-oriented repository-level tasks, including SWE-Bench issue resolution and feature implementation in real codebases. However, they still struggle with repository-level code optimization, which requires preserving behavior while improving runtime performance. Passing tests is not enough in this setting; a patch must preserve behavior, implement code optimization, and approach expert speedups. Current agents often miss bottlenecks hidden behind abstraction layers and native extensions, stop after shallow speedups, or insufficiently test the code patches that thus may silently break edge cases. We present PerfAgent, a profiler-guided, verifier-in-the-loop workflow that gives an off-the-shelf coding agent the feedback needed to find real hotspots, improve beyond the first passing patch, and use profiler evidence rather than timing alone to decide what to optimize next. On two challenging optimization benchmarks, GSO and SWE-fficiency-Lite, PerfAgent more than doubles the rate of expert-matching patches over OpenHands with GPT-5.1, improving from 19.6% to 39.2% on GSO and from 26% to 74% on SWE-fficiency-Lite. It also surpasses an oracle best-of-five baseline at substantially lower cost, showing that the gains come from better feedback rather than additional test-time sampling.
FedLSG: フェデレーション グラフ バックドア防御のための LLM 拡張セマンティック キャリブレーション
Federated Graph Neural Networks (FedGNN) はバックドア ポイズニングに対して非常に脆弱ですが、既存の防御は通常、セマンティックな理解を欠くルールベースのアプローチに依存しているため、ステルス トリガーに対して脆弱であり、良性の構造に対して有害です。これを解決するために、フェデレーション グラフ バックドア防御に大規模言語モデル (LLM) を統合する最初のフレームワークである FedLSG を紹介します。 FedLSG は、ローカルなグラフ構造とクライアントの更新動作を意味論的に豊富な自然言語表現に変換するグラフと動作をテキストグラウンディングスキームに導入します。このフレームワークはさらに、軽量の学生と教師のアーキテクチャを採用しています。サーバー側では、本格的な LLM が教師として機能し、グローバルな状況に応じたガイダンスを提供し、集約中にクライアントの更新を評価して、潜在的に悪意のある参加者を特定します。クライアント側では、バックドア トリガーに関連付けられたエッジの影響を抑制するために、LoRA ベースのスチューデントが意味論的推論を実行するように維持されます。グラフ パターンとクライアントの動作の両方のセマンティック解釈を可能にすることで、フレームワークは防御のためにルールベースのシグナルをメッセージ パッシングとクライアント アグリゲーションに適応的に組み込みます。実験では、FedLSG がグラフの整合性を損なうことなくバックドア攻撃に対する耐性を大幅に向上させることが実証されています。
原文 (English)
FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense
Federated Graph Neural Networks (FedGNNs) are highly vulnerable to backdoor poisoning, yet existing defenses typically rely on rule-based approaches that lack semantic understanding, making them vulnerable to stealthy triggers and harmful to benign structures. To solve this, we present FedLSG, the first framework that integrates large language models (LLMs) into federated graph backdoor defense. FedLSG introduces a graph and behavior to text grounding scheme that transforms local graph structures and client update behaviors into semantically rich natural language representations. The framework further adopts a lightweight student-teacher architecture. On the server side, a full scale LLM serves as a teacher, providing global contextual guidance and evaluating client updates during aggregation to identify potentially malicious participants. On the client side, a LoRA-based student is maintained to perform semantic reasoning, to suppress the influence of edges associated with backdoor triggers. By enabling semantic interpretation of both graph patterns and client behaviors, the framework adaptively incorporates rule-based signals into message passing and client aggregation for defense. Experiments demonstrate that FedLSG significantly improves resistance to backdoor attacks without compromising graph integrity.
自由回答形式の質問応答における推論の参考資料なしの評価
一か八かの分野で AI が生成した答えは、多くの場合流暢ですが、特に単一の最終的な答えではなく複数のステップからなる推論が含まれている場合、検証が困難です。私たちは、LLM によって生成された出力を監査するための、推論ベースで参照不要のフレームワークを提案します。この方法では、生成された推論トレースをセグメントに分解し、自然言語推論 (NLI) を使用してローカルな前提とターゲットの関係にラベルを付け、これらの関係をハイパーグラフに編成します。次に、決定論的な後方 AND-OR 検索により、生成された応答内で各セグメントがどのように根拠づけられているかを示すセグメント レベルの監査ラベルが割り当てられます。このフレームワークを 2 つの設定で評価します。Hard2Verify による演繹的数学的推論と、実際の臨床例からの LLM 推論トレースの医師注釈付きの新しいベンチマークである UroReason によるオープンエンドの医学的推論です。これらの設定全体にわたって、NLI ハイパーグラフ監査は、裁判官としての直接の LLM ベースラインよりも信頼性の高いリファレンスフリーの評価シグナルを提供します。臨床現場では、最先端の LLM 裁判官が問題のある推論セグメントを特定できず、流暢ではあるが根拠の薄い応答を過剰に受け入れてしまうことがよくあります。私たちの結果は、QA 評価では、最終的な回答や検証者としての LLM のみに依存するのではなく、推論トレース全体で推論関係がどのように構成されるかを考慮する必要があることを示しています。 UroReason は API を通じて利用可能になり、コードはオープンソースとしてリリースされます。
原文 (English)
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering
AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for auditing LLM-generated outputs. The method decomposes a generated reasoning trace into segments, labels local premise-target relations using Natural Language Inference (NLI), and organizes these relations into a hypergraph. A deterministic backward AND-OR search then assigns segment-level audit labels that indicate how each segment is grounded within the generated response. We evaluate the framework in two settings: deductive mathematical reasoning with Hard2Verify, and open-ended medical reasoning with UroReason, a new physician-annotated benchmark of LLM reasoning traces from real clinical cases. Across these settings, our NLI-hypergraph audit provides a more reliable reference-free evaluation signal than direct LLM-as-judge baselines. In the clinical setting, state-of-the-art LLM judges often fail to identify problematic reasoning segments, over-accepting fluent but weakly grounded responses. Our results show that QA evaluation should account for how inferential relations compose across a reasoning trace, rather than relying only on final answers or LLMs as verifiers. UroReason will be made available through an API, and our code will be released as open source.
SLPO: サロゲート ポリシーによる潜在推論の拡張
検証可能な報酬を伴う強化学習は、明示的な思考連鎖推論器でテスト時間のスケーリングを引き出すための主要なレシピになっています。ただし、すべての中間ステップを言語トークンとしてデコードする必要があるため、このスケーリング パスは依然として計算コストが高くなります。代わりに、潜在的推論は中間計算を連続ベクトルとして実行し、はるかに短い期間ですでに明示的な CoT に匹敵するか、それを上回っています。この約束にもかかわらず、潜在的推論者は主に模倣に縛られたままですが、明示的な CoT は結果報酬 RL によってすでに模倣を超えています。潜在的な軌道には、固定された思考予算の下で扱いやすいステップごとの尤度と適応的な停止インターフェイスが欠けているため、結果の報酬は潜在的なテスト時間のスケーリングを引き出すことができません。我々は、結果報酬 RL を自己回帰潜在推論者にもたらすために、サロゲート潜在ポリシー最適化 (SLPO) を導入します。これは、軌道レベルのクレジット割り当てのための潜在遷移に対する経験的代理ポリシー密度と、結果報酬の最適化によって変数ホライズン ポリシーに洗練される正確性監視されたストッピング ヘッドです。 SLPO は、継続的かつソフトな思考設定全体で、並列サンプリングの下で Pass@$k$ を改善し、より高い決定論的精度でより長い潜在計算をより困難なインスタンスに割り当てます。
原文 (English)
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scaling. We introduce Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners: an empirical surrogate policy density over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy. Across continuous and soft thinking settings, SLPO improves Pass@$k$ under parallel sampling and allocates longer latent computation to harder instances with higher deterministic accuracy.
PhenSPINE: 脊椎病理学診断の標準化されたベンチマーク
脊椎病理の正確な診断は放射線学的解釈に大きく依存していますが、自動化システムは多様で高品質のベンチマークの欠如によって妨げられています。この研究では、高度な深層学習研究を促進するために厳選された、250 人の患者からの 16,813 枚の画像で構成される磁気共鳴画像データセットである PhenSPINE を紹介します。私たちは、最新の畳み込みバックボーンと位置エンコーディング メカニズムを統合して、椎間板の解剖学的コンテキストを明示的にモデル化する堅牢な診断ベンチマークを提案します。 4 つの標準的な MRI シーケンスを評価した我々の実験では、サジタル T2 強調シーケンスが最も堅牢な診断値を提供し、50.31% という優れたマクロ F1 スコアを達成することが実証されました。データセット内のシーケンス全体の画像が周囲の解剖学的領域からのノイズ干渉によって大幅に損なわれるため、マルチシーケンス融合戦略はこの単一シーケンスのベースラインと比較してパフォーマンスが劣ることがわかりました。この研究により、堅牢なベースラインが確立され、脊椎解析のための配列選択に関する重要な洞察が得られます。
原文 (English)
PhenSPINE: A Standardized Benchmark for Spine Pathology Diagnosis
The accurate diagnosis of spinal pathologies depends heavily on radiological interpretation, yet automated systems are hindered by the lack of diverse, high-quality benchmarks. In this study, we present PhenSPINE, a Magnetic Resonance Imaging dataset comprising 16,813 images from 250 patients, curated to facilitate advanced deep learning research. We propose a robust diagnostic benchmark that integrates state-of-theart convolutional backbones with a Positional Encoding mechanism to explicitly model the anatomical context of intervertebral discs. Evaluating across four standard MRI sequences, our experiments demonstrate that the Sagittal T2-weighted sequence offers the most robust diagnostic value, achieving a superior Macro F1-score of 50.31%. We find that multisequence fusion strategies yield inferior performance compared to this single-sequence baseline, as the images across sequences in our dataset are significantly compromised by noise interference from surrounding anatomical regions. This work establishes a robust baseline and offers critical insights into sequence selection for spine analysis.
アリスは間違ったことをしましたか?大学のコンピューティング教育における生成的 AI の使用に対する学生の認識における文化間の違い
高等教育における生成 AI (GenAI) の台頭により、学問の誠実さと倫理的利用をめぐる緊急の議論が引き起こされています。この研究では、カナダと韓国の大学の学生の反応を比較し、GenAI の使用に対する学生の認識における異文化間の違いを調査しました。 2024 年秋に実施されたシナリオベースの調査を使用して、学生が AI 支援コーディング実践の倫理性とルール遵守をどのように判断したかを分析しました。結果は、機能的に同一の教育機関のポリシーにもかかわらず、カナダの学生は韓国の学生と比較して、GenAI の使用を非倫理的で教育機関のポリシーに反するものとして認識する可能性が一貫して高いことを明らかにしました。マン・ホイットニーの U 検定や相関係数を含む統計分析により、ほぼすべてのシナリオで大きな違いがあることが実証されました。シナリオ生成に使用された要素の分析により、課題に組み込まれた AI 生成コードの量が倫理的判断に最も強く影響することが示されました。調査結果はホフステードの文化的側面の枠組みを通じて解釈され、力の距離、個人主義、不確実性の回避などの文化的要因が、GenAIに関する学生の倫理的推論を大きく形作ることを示唆しています。私たちの結果は、教育における AI の公平な統合が、学術的誠実さの多様な概念を考慮に入れて、文化に応じたものでなければならないことを強調する一連の証拠の増加に貢献しています。私たちは、学問的誠実さの基本原則を守りながら、地域の文化的背景に配慮した微妙な AI 使用ガイドラインの開発を主張します。この研究は、倫理的な AI 政策を知らせ、世界の高等教育現場での責任ある GenAI の使用をサポートするために継続的な異文化研究の必要性を強調しています。
原文 (English)
Did Alice Do Wrong? Cross-Cultural Differences in Student Perceptions of Generative AI Use in University Computing Education
The rise of generative AI (GenAI) in higher education has prompted urgent debates surrounding academic integrity and ethical use. This study examines cross-cultural differences in student perceptions of GenAI use, comparing responses from students at Canadian and South Korean universities. Using a scenario-based survey administered in Fall 2024, we analyzed how students judged the ethicality and rule compliance of AI-assisted coding practices. Results reveal that Canadian students were consistently more likely to perceive the use of GenAI as both unethical and against institutional policies compared to Korean students, despite functionally identical institutional policies. Statistical analysis, including Mann-Whitney U tests and correlation coefficients, demonstrated significant differences across nearly all scenarios. Analysis of the factors used in generating scenarios indicated that the amount of AI-generated code incorporated into assignments most strongly influenced ethical judgments. Findings were interpreted through Hofstede's cultural dimensions framework, suggesting that cultural factors such as power distance, individualism, and uncertainty avoidance significantly shape students' ethical reasoning regarding GenAI. Our results contribute to the growing body of evidence emphasizing that equitable AI integration in education must be culturally responsive, taking into account diverse conceptions of academic integrity. We advocate for the development of nuanced AI-use guidelines that are sensitive to local cultural contexts while upholding fundamental principles of academic honesty. This study highlights the need for ongoing cross-cultural research to inform ethical AI policies and support responsible GenAI use in global higher education settings.
自律言語エージェントを介したパーソナライズされたレコメンデーション ツールの学習
大規模言語モデル (LLM) は、その強力な推論機能と広範な世界知識により、最近レコメンダー システムで注目を集めていますが、以前の LLM ベースのエージェントは幻覚やコンテキストの長さの制限に悩まされているため、フルランク付けのレコメンデーション タスクには適していません。 LLM 自体を変更するのではなく、アーキテクチャ設計を通じてこれらの制限を回避するために、自律言語 $\textbf{A}$gents (PRTA) を介したメモリベースの $\textbf{P}$ersonalized $\textbf{R}$ecommendation $\textbf{T}$ool 学習というエージェントベースの推奨フレームワークを提案します。このフレームワークでは、LLM がツールとして複数の推奨モデルと対話する中央プランナーとして機能します。 LLM ベースのエージェントは、高度な推論とパーソナライズされたツールの選択を担当しますが、従来のレコメンデーション モデルは、行動パターンのモデリングにおけるスケーラビリティを活用してフル ランキング スコアリングを実行します。パーソナライズされたツールの選択をサポートするために、エージェントがユーザー プロファイルと候補のランク付けされたリストに基づいて各ユーザーのツールを評価および比較できるようにする反映メカニズムを設計します。 3 つの公開データセットにわたる広範な実験により、フルランク付けの推奨パフォーマンスの向上において、従来の推奨や LLM ベースのベースラインよりも \modelname の優位性が実証されました。
原文 (English)
Personalized Recommendation Tool Learning via Autonomous Language Agents
Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive world knowledge, previous LLM-based agents suffer from hallucination and context-length limitations, and thus are not suitable for full-ranking recommendation tasks. To circumvent these limitations through architectural design rather than modifying the LLM itself, we propose an agent-based recommendation framework, memory-based $\textbf{P}$ersonalized $\textbf{R}$ecommendation $\textbf{T}$ool learning via autonomous language $\textbf{A}$gents (PRTA), in which an LLM acts as a central planner interacting with multiple recommendation models as tools. The LLM-based agent is responsible for high-level reasoning and personalized tool selection, while traditional recommendation models perform full-ranking scoring, leveraging their scalability in modeling behavioral patterns. To support personalized tool selection, we design reflection mechanisms that enable the agent to evaluate and compare tools for each user based on user profiles and candidate ranked lists. Extensive experiments across three public datasets demonstrate the superiority of \modelname over traditional recommendation and LLM-based baselines in improving full-ranking recommendation performance.
サイバー脅威インテリジェンス レポートから到達可能な攻撃チェーンを抽出するための自動フレームワーク
サイバー脅威インテリジェンス (CTI) レポートには現実世界の攻撃プロセスが詳しく説明されていますが、その非構造化ナラティブを自動化された攻撃経路推論に直接使用することはできません。既存の CTI 抽出手法は、各攻撃ステップの実行条件とその結果の状態をモデル化せずに、インジケーター、エンティティ、または TTP ラベルに焦点を当てているため、抽出された知識は、多段階の攻撃チェーンにわたる状態のマッチングや到達可能性の分析をサポートしていません。本稿では、各攻撃ステップを前提条件、攻撃動作、事後条件の攻撃単位としてモデル化することで、到達可能な攻撃チェーンを抽出する自動フレームワークを提案します。大規模言語モデル (LLM) によって支援されたマルチステージ パイプラインは、攻撃動作のスケルトンを抽出し、その事前条件と事後条件を回復し、それらを事前定義された述語に正規化し、壊れた依存関係を修復します。結果として得られるユニットは、攻撃目標の到達可能性を推論するための Datalog スタイルのルールにコンパイルされます。人間による検証が行われた 334 の注釈付きステップを含む 20 の CTI レポートのデータセット上で、当社のフレームワークは、攻撃動作の回復において、代表的な CTI 抽出システムよりも高い注釈付きステップ カバレッジを達成します。さらに、事前条件と事後条件を明示的に生成することにより、エンドツーエンドの LLM ベースラインによって生成される攻撃ユニットよりも完全で一貫性のある攻撃ユニットが生成されます。抽出されたチェーンでは、Datalog 推論は 20 レポート中 19 レポートで指定された攻撃目標に到達しますが、逆方向検索では、生成されたルールに基づいて 34 の攻撃パスが得られます。ソース コードと実験成果物は、匿名化されたリポジトリで入手できます。 。
原文 (English)
An Automated Framework for Extracting Reachable Attack Chains from Cyber Threat Intelligence Reports
Cyber Threat Intelligence (CTI) reports richly describe real-world attack processes, but their unstructured narratives cannot be directly used for automated attack-path reasoning. Existing CTI extraction methods focus on indicators, entities, or TTP labels without modeling the execution conditions and resulting states of each attack step, so the extracted knowledge supports neither state matching nor reachability analysis across multi-stage attack chains. This paper proposes an automated framework that extracts reachable attack chains by modeling each attack step as an attack unit of preconditions, an attack behavior, and postconditions. A multi-stage pipeline assisted by large language models (LLMs) extracts attack behavior skeletons, recovers their preconditions and postconditions, normalizes them into predefined predicates, and repairs broken dependencies; the resulting units are compiled into Datalog-style rules for attack-goal reachability reasoning. On a dataset of 20 CTI reports containing 334 human-validated annotated steps, our framework achieves higher annotated-step coverage than representative CTI extraction systems in recovering attack behaviors. Moreover, by explicitly generating preconditions and postconditions, it produces attack units that are more complete and consistent than those generated by end-to-end LLM baselines. On the extracted chains, Datalog inference reaches the specified attack goal in 19 of 20 reports, while backward search yields 34 attack paths under the generated rules. The source code and experimental artifacts are available in an anonymized repository. .
世界のモデルは記憶し、俳優は忘れる: 継続的なモデルベースの RL のための夢のリハーサル
DreamerV3 ファミリのモデルベースの強化学習エージェントは、タスク シーケンスでトレーニングされると、無制限のリプレイ バッファーが以前のすべてのエクスペリエンスを保持している場合でも、壊滅的に忘れてしまいます。私たちは、継続的 RL 文献が答えを想定しているものの、測定したことのない質問をします。どのコンポーネントが忘れているのか?明確ではないリプレイの下で、事前に登録されたコンポーネント レベルのプローブ (全体で n=3 シード) は、アクターの行動が崩壊する一方で、ワールド モデルが古いタスクに関する測定可能なすべて (報酬の区別 (保持率 ~1.0)、価値の推定、終了構造) を本質的に保持していることを示します。この状況での忘れはチャネルの問題であり、記憶の問題ではありません。これを介入によって実証します。世界モデルが凍結され、同一の想像ロールアウトが行われた場合、想像における強化学習は失われたスキル (0/3 シード) を回復できませんでしたが、世界モデル自身の段階的な夢に対する教師あり自己模倣は、環境相互作用ゼロで 3/3 シードでスキルを回復します。トレーニング中にインターリーブされるこの段階的な夢のリハーサルは、タスクラベルなし、パラメーター一定の継続学習を生成します。単純なリプレイが 0/3 を通過した場合に 3/3 の 4 タスク チェーンが保持され、3/3 の 8 タスク チェーンが保持され、一致する実際のエピソードのクローン作成と比べて一貫したゲインが得られます (ペアの差 +0.13、ブートストラップ 95% CI [0.07、0.24]、完了)種子の分離)。夢の採点ステップは負荷がかかります。2 つの採点失敗モードを特徴付け、結果が汚染される前に両方を捕捉するオフライン選択ゲージを提供し、それらを閉じる実現優先採点ルールを提供します。すべての実験は、コミットされたプロトコルで事前に登録されました。反証された仮説はすべて報告されます。
原文 (English)
The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL
Model-based reinforcement-learning agents of the DreamerV3 family forget catastrophically when trained on task sequences, even when an unbounded replay buffer preserves every earlier experience. We ask a question the continual-RL literature has assumed an answer to but never measured: which component forgets? Under never-clear replay, pre-registered component-level probes (n=3 seeds throughout) show that the world model retains essentially everything measurable about old tasks -- reward discrimination (retention ratio ~1.0), value estimates, and termination structure -- while the actor's behavior collapses. Forgetting in this regime is a channel problem, not a memory problem. We demonstrate this by intervention: with the world model frozen and identical imagined rollouts, reinforcement learning in imagination fails to recover a lost skill (0/3 seeds), while supervised self-imitation on the world model's own graded dreams recovers it on 3/3 seeds with zero environment interaction. Interleaved during training, this graded dream rehearsal yields a task-label-free, parameter-constant continual learner: 3/3 four-task chains retained where plain replay passes 0/3, 3/3 eight-task chains, and consistent gains over matched real-episode cloning (paired difference +0.13, bootstrap 95% CI [0.07, 0.24], complete seed separation). The dream-grading step is load-bearing: we characterize two scoring failure modes, provide an offline selection gauge that caught both before they contaminated results, and give a realized-first grading rule that closes them. All experiments were pre-registered with committed protocols; every refuted hypothesis is reported.
Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction
We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a continuous geographic space rathe…
Convergence-Latency-Aware Adaptive Modulation and Resource Allocation in RIS-Assisted Wireless Federated Learning
Federated learning (FL) over wireless networks suffers from significant training latency and degraded convergence due to unreliable wireles…
An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies
Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry…
RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases oft…
Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification
Polarimetric synthetic aperture radar (PolSAR) image classification is a representative task for physics-aware GeoAI, where land-cover sema…
OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have…
Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes
Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly fr…
Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning
This paper introduces Sentence Splitter, a self-supervised framework built upon a T5-based encoder--decoder architecture for uncovering the…
Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-art reasoning models show grea…
Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task…
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos
Language-guided aerial perception aims to understand user-specified tiny targets in complex unmanned aerial vehicle (UAV) scenes. In real U…
PRISM-DR: Per-lesion Retinal Inference with Specialist Models for Diabetic Retinopathy
Diabetic retinopathy is a leading cause of preventable blindness; its early lesions are small, low contrast, and easily missed in manual sc…
Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English…
Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generall…
OSVE: One Step Video Editing with One Step Diffusion Models
Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion. We present OSV…
A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace
As AI agents become integral to business workflows, establishing guiding user experience (UX) principles is crucial for ensuring user trust…
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detec…
When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization
Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely exami…
HijackKV: New Threat in Position-Independent KV Cache Reuse
Key-Value (KV) cache reduces inference latency in large language models (LLMs). Traditional prefix-based reuse has low cache hit rates acro…
When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets
Shippers are beginning to delegate carrier selection to large language model (LLM) agents. We ask what such delegation does to a freight ma…
Time Series Network Utilization KPI Forecasting Using Advanced AI/ML Models
The rapid proliferation of data-intensive applications, cloud infrastructure, and IoT ecosystems has made proactive resource provisioning c…
TINY_SCHILLER: A Drop-In German Drama Corpus for Small Language Models
tiny_schiller closes the small-language-model prototyping, fine-tuning, education, and research gap for German literary text, providing a s…
Are Attributions of Consciousness to AI Chatbots Epistemically Innocent?
Artificial intelligence (AI) chatbots (e.g., ChatGPT) can communicate in strikingly humanlike ways. This has prompted many chatbot users to…
Post-Training in Time Series Foundation Models: A Unifying Framework
Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insuf…
Taming the Security-Energy Paradox: A Green AI Approach to Optimized Android Malware Detection
An increase in advanced Android malware requires the use of deep learning models, which can run on Android devices. But there is a trade-of…
Drift-Aware RL-based Wavelet Denoising for Network-Traffic Anomaly Detection
Traffic-utilisation measurements for network monitoring are corrupted by additive noise and statistical drift: time-dependent change in the…
A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation and Cross-Domain Generalisability
Robust out-of-the-box performance is essential for the clinical deployment of deep learning models in medical imaging. An important but und…
Test Case Prioritization for DNNs via Neural Collapse Instability
With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limit…
Language-Specific versus Cross-Lingual Knowledge Graphs for Implicit Aspect Identification in Arabic: A Comparative Study of Reasoning and Adaptation Strategies
Aspect-based sentiment analysis (ABSA) in Arabic must recover both explicitly stated aspects and implicit aspects that are never named in t…
Co-Evolving LLM Evaluators and Policies via DynamicRubric
Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policie…
Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results
Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instructio…
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that cont…
PRIME-SVR: Physics-infoRmed Implicit Multi-Echo Slice-to-Volume Reconstruction for Fetal T2 mapping
Slice-to-volume reconstruction (SVR) is the standard method for obtaining high-resolution (HR) 3D fetal brain volumes from motion-corrupted…
Formal Foundations for Known Good Reliable Die Screening in Chiplet-Based AI Systems-on-Chip
The rapid growth of chiplet-based artificial intelligence systems-on-chip (SoCs) has exposed a fundamental gap in semiconductor test method…
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distribu…
Active Inference as a Convex Markov Decision Process
Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic object…
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio rea…
StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation
Existing human--object interaction (HOI) video generation methods are largely limited to offline short-video generation with complex drivin…
The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural Networks
Additive models buy interpretability by forbidding feature interactions, a constraint that neural instantiations enforce architecturally. W…
ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
The quadratic $N\times N$ attention score matrix remains a central obstacle to extending Transformers to longer input lengths. Existing eff…
Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis
Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours. While lar…
DQAOA-GPT: AI-Accelerated Distributed Quantum Optimization for Combinatorial Problems
While combinatorial optimization problems are central to many scientific and engineering applications, their solution remains challenging d…
On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens
Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural conte…
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descri…
The Ethics of Autonomous AI Agents for Offensive Security
LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly sc…
The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models
Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. Howev…
Self-supervision drives representational convergence in medical foundation models more than clinical supervision
Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervi…
Sound Probabilistic Safety Bounds for Large Language Models
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to…
Courteous Anticipation: Improving Long-Lived Task Planning in Persistent Shared Environments
We consider a task planning scenario in which robots sharing a persistent environment are assigned tasks one at a time from a held-out sequ…
Don't Trust the Label: License Laundering in AI Supply Chains
AI artifacts move through a multi-platform supply chain, spanning datasets and models on Hugging Face and applications on GitHub. While eac…
Toward Reliable RGB-D Semantic Segmentation: Handling Missing Modalities via Condition Dropout
RGB-D semantic segmentation has achieved remarkable progress, yet most models assume that RGB and depth are always available. In practice,…
Understanding Generative AI-mediated User Engagement with Academic Library Resources
This study empirically analyzed generative AI as an emerging discovery pathway to academic library resources. Utilizing web analytics from…
Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA…
Generative AI floods and dilutes the market for books
Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buye…
FMRP-LEAN: A HIPAA-Compliant AI-Augmented LIMS Architecture for End-to-End Clinical Assay Workflow Optimization
Clinical biomarker workflows in translational research settings often rely on spreadsheet-driven tracking, manual quality control (QC) reco…
Persian Pixel: A large-scale synthetic OCR dataset for Persian language
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spo…
A Survey on Semantic Modeling for Building Energy Management
Building Energy Management (BEM) is central to reducing energy use and CO2 emissions in the building sector. Although IoT technologies now…
Avoiding Obfuscation with Prover-Estimator Debate
Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly compl…
Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review
Multi-Agent Reinforcement Learning (MARL) has achieved strong performance in simulated benchmarks, yet real deployments often violate the a…
SciTrek: Evaluating and Improving Long-Context Numerical Reasoning over Scientific Articles
We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language…
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches trai…
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
Existing image forgery detection (IFD) methods either exploit low-level, semantics-agnostic artifacts or rely on multimodal large language…
Fidelity Before Structure: Verbatim Chunks Beat Lossy Artifact Extraction in Long-Conversation LLM Memory
A growing class of conversational-memory systems compresses dialogue history into structured artifacts (extracted facts, decisions, or even…
Statistical Early Stopping for Reasoning Models
While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning ste…
Content Creation with Spillovers: An Incentive Design Approach
The rise of AI amplifies the economic phenomenon of \emph{positive spillovers}: when creators contribute content that can be reused and ada…
Prompt Programming for Cultural Bias and Alignment of Large Language Models
Culture shapes reasoning, values, prioritization, and strategic decision-making, yet large language models (LLMs) often exhibit cultural bi…
Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks
We propose that AI automation is a continuum between: (i) crashing waves where AI capabilities surge abruptly over small sets of tasks, and…
Agent-Based Modeling of Low-Emission Fertilizer Adoption for Dairy Farm Decarbonisation using Empirical Farm Data
To understand complex system dynamics in dairy farming requires tools that capture farm heterogeneity, social interactions, and cumulative…
Prober.ai: Gated Inquiry-Based Feedback via LLM-Constrained Personas for Argumentative Writing Development
The proliferation of large language models (LLMs) in educational settings has paradoxically undermined the cognitive processes they purport…
人生は一度だけではない: 階層型スキルのメタ進化に向けて
テスト時のスキルの進化は、展開されたエージェント システムを強化するための新しいパラダイムとみなされます。既存の研究は主に、基盤となる LLM の高価なパラメータ更新に依存する、ハードコーディングされたスキル進化戦略またはパラメトリック学習に焦点を当てています。この論文では、さまざまな下流シナリオでエージェント システムを継続的に改善するには、スキル進化フレームワーク自体のテスト時の改良が必要であり、軽量のアルゴリズム適応が実現可能であることを実証します。具体的には、エージェントのタスク実行トレースからメタスキルを学習することでスキルとスキル進化戦略を共同で最適化する、軽量の階層型スキルメタ進化ソリューションである HiSME を提案します。多様なエージェントベンチマークの実験では、メタ進化の方が純粋なスキル進化よりも高品質のスキルライブラリを生成でき、さまざまなシナリオに合わせて多様なメタスキルを導き出すことができるため、将来の継続的な体験学習が容易になることが示されています。私たちのコードは https://anonymous.4open.science/r/HiSME-BD45 で一時的に公開されています。
原文 (English)
You Live More Than Once: Towards Hierarchical Skill Meta-Evolving
Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems. Existing works mainly focus on hard-coded skill evolving strategies or parametric learning that rely on expensive parameter updates in the underlying LLMs. In this paper, we demonstrate that test-time refinement of the skill evolving framework itself is necessary for continuous improvement of the agent systems in different downstream scenarios, and lightweight algorithmic adaptation is feasible. Specifically, we propose HiSME, a lightweight hierarchical skill meta-evolving solution that jointly optimizes skills and the skill evolving strategy by learning meta-skills from agents' task execution traces. Experiments on mainstream agentic benchmarks show that meta-evolving can produce a higher-quality skill library than pure skill evolving and can derive diverse meta-skills for different scenarios, thereby facilitating future continual experience learning. Our code is temporarily public at https://anonymous.4open.science/r/HiSME-BD45.
CEO ベンチ: エージェントは長期戦を勝ち抜くことができますか?
言語モデル エージェントは、ソフトウェア エンジニアリングや顧客サービスなど、孤立した短期間のタスクの熟練した実行者になりつつあります。しかし、現実世界の課題には、エージェントではほとんどテストされていない高度なスキルの組み合わせが必要です。(1) 不確実性の中で長い視野をナビゲートする。 (2)騒音環境下での情報取得。 (3) 変化する世界に適応する。 (4) 一貫した目標に向かって複数の可動部分を調整する。 CEO-Bench を紹介します。これは、スタートアップを 500 日間運営するという代表的な現実世界のタスクをシミュレートすることによって、これらの機能をまとめて評価します。エージェントは、プログラム可能な Python インターフェイスを介して、架空の会社の価格設定、マーケティング、予算編成、その他多くの側面を管理し、人間の CEO と同じ環境で動作し、同じ課題に直面します。成功するには、ノイズの多い相互接続されたビジネス データベースを分析し、シグナルを健全な戦略に変換し、プログラミングを使用して多くの意思決定を調整する必要があります。最も強力なエージェントは、顧客コホートをシミュレートして将来の現金を予測し、隠れた顧客の好みを明らかにするために交渉履歴を掘り起こす高度なコードを作成します。それでも、ほとんどの最先端モデルはこの環境では苦戦します。クロード オーパス 4.8 と GPT-5.5 のみが開始残高 100 万ドルを超えて終了し、どちらも一貫して利益を上げています。 CEO-Bench は、時間の経過とともに持続的かつ適応的な進歩を推進するために必要なインテリジェンスを測定するための第一歩を踏み出します。
原文 (English)
CEO-Bench: Can Agents Play the Long Game?
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophisticated skills that remain largely untested in agents: (1) navigating long horizons amid uncertainty; (2) acquiring information in noisy environments; (3) adapting to a changing world; (4) orchestrating multiple moving parts toward a coherent goal. We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. An agent manages pricing, marketing, budgeting, and many other aspects of a fictional company through a programmable Python interface, operating in the same environment and facing the same challenges as a human CEO. Success demands analyzing noisy, interconnected business databases, translating signals into sound strategy, and coordinating many decisions with programming. The strongest agents write sophisticated code that forecasts churn regimes, billing timing, customer losses, and future cash under different scenarios. Even so, most state-of-the-art models struggle in this environment. Only Claude Fable 5, GPT-5.6 Sol, and Claude Opus 4.8 finish above the $1M starting balance, and all evaluated models remain below the rule-based baseline. CEO-Bench takes a first step toward measuring the intelligence required to drive sustained, adaptive progress over time.
Fara-1.5: Scalable Learning Environments for Computer Use Agents
Collecting computer use data from human demonstrations is expensive and slow, motivating the need for scalable generation strategies. This…
PedNStream: 歩行者交通管理のためのスケーラブルなネットワーク フロー シミュレーション
大規模な群衆管理には、計算効率が高く、フィードバック ベースの制御と互換性のある歩行者シミュレーションが必要です。ただし、ほとんどのオープンソース ツールは微細なものであるか、ネットワーク規模の閉ループ評価用に設計されていません。このペーパーでは、リンク伝送モデル (LTM) に基づいた巨視的な歩行者ネットワーク負荷のためのオープンソースの Python ネイティブ シミュレーターである PedNStream (Pedestrian Network Flow Simulation) について説明します。このフレームワークは、拡散と活動に起因する変動を捉える確率的リンクダイナミクスを組み込むことで LTM ベースの歩行者モデルを拡張し、動的なユーザーの平衡ルート選択を、不確実で介入主導の設定に適したユーティリティベースの定式化に置き換えます。 PedNStream は、ゲート、フロー分離、ルート ガイダンスなどの介入用のコントローラー インターフェイスが組み込まれたモジュール式フレームワークとして実装されています。私たちはフレームワークを段階的に評価します。合成シナリオでは、キューの形成、スピルバック、輻輳の解消、適応型再ルーティングなどの主要なメカニズムを検証します。実際のネットワーク実験では、大規模な行動と観察された歩行者数との一貫性を評価します。閉ループのケーススタディではコントローラーの統合を実証し、ランタイム分析ではスケーラビリティを定量化します。これらの結果により、PedNStream は大規模な歩行者ネットワークのシミュレーションと制御のための効率的かつ実用的なテストベッドとして確立されます。
原文 (English)
PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management
Evaluating operational crowd management at network scale requires simulations that can be run repeatedly while adapting interventions to changing conditions. Microscopic models can represent detailed individual movement, but their computational cost may limit their use in such repeated, network-scale evaluations. This paper presents PedNStream (Pedestrian Network Flow Simulation), an open-source, Python-native simulator for macroscopic pedestrian network simulation based on the Link Transmission Model (LTM). PedNStream extends LTM-based pedestrian models with stochastic link dynamics that represent local variation in pedestrian flow. It uses a utility-based route-choice model to capture how pedestrians adjust their route choices in response to congestion and control interventions as conditions change over time. The modular framework provides controller interfaces for gating, flow separation, and route guidance. We evaluate PedNStream in a staged manner. Synthetic scenarios verify key crowd-dynamics mechanisms, including queue formation, spillback, congestion dissipation, and adaptive rerouting. Real-network experiments assess large-scale behavior against observed pedestrian counts. A closed-loop case study demonstrates controller integration, and a runtime analysis quantifies scalability. These results position PedNStream as an efficient and practical testbed for large-scale pedestrian network simulation and crowd management research.
AutoVSR: 回路図からのシンボリック式生成のためのビジュアルからシンボリックへの自動推論
シンボリック式は回路の動作を効果的に特徴付けて予測できますが、回路図から直接それを導き出すのは困難です。このプロセスでは、画像からの回路構造のビジュアルからシンボリックへの正確な構築と、複数ステップのシンボリック導出の正確さが必要であり、どちらも厳密な正確性要件を課します。この研究では、ビジョン言語モデル (VLM) を使用して回路式をビジュアルからシンボリックに生成するための自動フレームワークである AutoVSR を提案します。 AutoVSR は、回路図を実行可能な中間表現 (Executable IR) に再構築し、推論にシンボリック ソルバーを利用することにより、シンボリック式生成の精度を大幅に向上させます。 AutoVSR は 2 つの主要な革新を導入しています。1 つはコンポーネント ルールの取得と検証ベースのフィードバックによって導かれる IR 構築方法、もう 1 つは信頼性の高いマルチステップ導出のためのシンボリック ツール ライブラリを備えたプランニング エージェントとして実装されたシンボリック ソルバーです。エンドツーエンドの VLM アプローチおよびメインのシンボリック式生成タスクにおける特殊な手法と比較して、AutoVSR はそれぞれ 30.01 ~ 59.45% および 41.96 ~ 51.84% の精度向上を達成します。さらに、AutoVSR は、推論コストと計算効率の点で、クローズドソースの最先端の VLM を上回っています。コードは https://github.com/LongfeiLi1/AutoVSR で入手できます。
原文 (English)
AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic
Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is challenging. This process requires accurate visual-to-symbolic construction of circuit structure from images and correct multi-step symbolic derivation, both of which impose strict correctness requirements. This work proposes AutoVSR, an automated framework for visual-to-symbolic generation of circuit expressions using Vision Language Models (VLMs). By reconstructing circuit diagrams into an executable intermediate representation (Executable IR) and leveraging a symbolic solver for reasoning, AutoVSR significantly improves the accuracy of symbolic expression generation. AutoVSR introduces two key innovations: an IR construction method guided by component rule retrieval and verification-based feedback, and a symbolic solver implemented as a planning agent equipped with a symbolic tool library for reliable multi-step derivation. Compared with end-to-end VLM approaches and specialized methods on the main symbolic expression generation task, AutoVSR achieves accuracy improvements of 30.01--59.45% and 41.96--51.84%, respectively. Moreover, AutoVSR surpasses closed-source state-of-the-art VLMs in inference cost and computational efficiency. Code is available at https://github.com/LongfeiLi1/AutoVSR.
Alipay-PIBench: コーディング エージェント向けの現実的な決済統合ベンチマーク
支払いの統合は、要求の厳しいリポジトリ レベルのソフトウェア タスクです。エージェントは、適切な製品を選択し、調整されたクライアント/サーバー フローを実装し、支払い結果を検証し、トランザクションとビジネス状態の間の一貫性を維持する必要があります。現実的な Alipay 決済統合に関するコーディング エージェントを評価するためのベンチマークである Alipay-PIBench を紹介します。これには、9 つの製品固有のプロジェクトと 18 のタスク インスタンスが含まれており、それぞれが基本的な機能完了シナリオと高度なリスク認識強化シナリオに編成されています。シナリオ固有のルーブリックは、決定論的な静的チェック、ユニットチェック、統合チェック、およびエンドツーエンドのチェックをサポートし、セマンティック要件に対する LLM 支援の評価によって補足されます。 6 つのコーディング エージェント モデルを評価し、ルーブリック合格率 (RPR) を報告します。スキルありの条件下では、平均 RPR は 68.58% から 91.37% の範囲です。 Alipay 決済統合スキルへのアクセスにより、スキルなしの状態と比較して平均 RPR が平均 10.31 パーセント ポイント向上しますが、その向上はモデル、製品、シナリオによって異なります。メソッドレベルの結果は、ソースレベルの完了、実行可能な支払い動作、支払いドメインの要件を区別します。 Alipay-PIBench は、モデルの機能を診断し、支払い統合における構造化されたガイダンスを評価するための制御された設定を提供します。
原文 (English)
Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents
Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states. We introduce Alipay-PIBench, a benchmark for evaluating coding agents on realistic Alipay payment integration. It contains nine product-specific projects and 18 task instances, each organized into Basic functional-completion and Advanced risk-aware hardening scenarios. Scenario-specific rubrics support deterministic static, unit, integration, and end-to-end checks, supplemented by LLM-assisted assessment for semantic requirements. We evaluate six coding-agent models and report rubric pass rate (RPR). Under the with-skill condition, mean RPR ranges from 68.58% to 91.37%. Access to the alipay-payment-integration skill improves mean RPR by 10.31 percentage points on average relative to the without-skill condition, with gains varying across models, products, and scenarios. Method-level results distinguish source-level completion, executable payment behavior, and payment-domain requirements. Alipay-PIBench provides a controlled setting for diagnosing model capability and evaluating structured guidance in payment integration.
RLHF 選好データにおける評価者の状態バイアス: 監査フレームワーク
ヒューマンフィードバックからの強化学習 (RLHF) における構造化交絡を特定します。ペアごとの優先ラベルは、比較された出力を反映することを目的としていますが、注釈付け中の評価者の状態も反映する場合があります。ストレスの多い、または悲惨な状況が続くと、時間の経過とともに評価者の好みが変化する可能性があります。その結果、嗜好データは、応答品質に関する判断とともに評価者の状態をエンコードできます。これらのシフトは、通常の不一致やランダムなラベル ノイズとは異なります。これらは状態に依存しており、同様の条件下で作業するアノテーター間で共有でき、報酬モデリングとポリシーの最適化を通じて伝播できます。したがって、我々は、評価者の状態の変化を、RLHF 選好データにおける構造化バイアスのもっともらしく、テスト可能な原因として提案します。この文書では、このバイアスの原因を研究するための仮説と監査フレームワークを開発します。評価者状態シフト、評価者状態交絡、相関評価者状態バイアスを定義します。また、生存レベルの感情的信憑性を、語彙的、語用的、談話的、および安全関連の特徴を使用した測定可能な反応パターンとして定義します。相関のある評価者の状態バイアスがどのようにして集計に耐え、学習された報酬シグナルを入力できるかを分析します。初期監査の 5 つの反証可能な予測と効果量のしきい値を導き出します。最後に、公開されている命令調整モデルに適用できる監査プロトコルとパイロット研究計画を紹介します。特定のデプロイされたモデルのトレーニング履歴を推測することはありません。私たちの目標は、RLHF 嗜好データにおける構造化されたバイアスの、もっともらしくテスト可能な原因を特定することです。
原文 (English)
Rater State Bias in RLHF Preference Data: An Audit Framework
We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, raters' preferences may shift over time, so that preference data encode rater state alongside judgments about response quality. We argue that, if present, such shifts would differ from ordinary disagreement or random label noise. They would be state dependent, could be shared across annotators under similar conditions, and would not necessarily cancel during aggregation, reward modeling, and policy optimization. We propose rater state shift as a plausible and testable source of structured bias in RLHF preference data. This paper develops a hypothesis and an audit framework for studying this source of bias. We define rater state shift, rater state confound, and correlated rater state bias. We also propose survival level emotional authenticity as a candidate output signature, defined by lexical, pragmatic, discourse, and safety features whose reliability and validity remain to be demonstrated. We analyze the conditions under which correlated rater state bias would not be averaged out during aggregation and could enter the learned reward signal. We state five predictions that distinguish this mechanism from generic engagement optimization, together with effect size thresholds for an initial audit, and note which require proprietary data. Finally, we present an audit protocol and pilot study plan that can be applied to publicly available instruction tuned models. We do not infer the training history of any specific deployed model.
LaCache: 拡散大規模言語モデルの正確なキャッシュと高精度適応推論
拡散ベースの大規模言語モデル (DLLM) により、テキスト生成における半自己回帰 (SAR) デコードによる並列生成が可能になります。しかし、現在の方法はオペレータレベルの冗長性が深刻です。プレフィックスとマスクされたサフィックスがブロック内で不変のままであることを無視して、ノイズ除去ステップ中にシーケンス全体を再計算します。私たちは、ロスレス キャッシュと混合精度によってこの冗長性を軽減する、トレーニング不要の高速化フレームワークである LaCache を提案します。具体的には、LaCache は 3 種類の中間結果をキャッシュすることでロスレス状態メモ化 (LSM) を採用しています。(i) 出力を埋め込むための EmbedCache、(ii) トークンごとのプレアテンション状態のための RoPECache、および (iii) FlashAttendant 内のオンライン ソフトマックス統計のための FACache。これらのキャッシュにより、モデルは出力を変更せずに、未変更のトークンに対する冗長な計算をスキップできます。メモリ帯域幅のボトルネックをさらに軽減するために、LaCache には、拡散プロセス全体にわたるステップ依存の活性化分布に合わせて調整された、FFN 層のグループごとの FP8 量子化戦略が組み込まれています。実験では、LaCache 単独で通常の DLLM と比較して約 1.3 倍のエンドツーエンドの高速化を達成できることが実証されています。既存の高速化手法と組み合わせると、LaCache は同等のタスク精度を維持しながら、エンドツーエンドの速度が最大 40.2 倍向上します。
原文 (English)
LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
Diffusion-based Large Language Models(DLLMs) enable parallel generation via Semi-Autoregressive (SAR) decoding in text generation. However, current methods suffer from severe operator-level redundancy: they recompute the entire sequence during denoising steps, ignoring that the prefix and masked suffix remain invariant within a block. We propose LaCache, a training-free acceleration framework that alleviates this redundancy through lossless caching and mixed precision. Specifically, LaCache employs Lossless State Memoization (LSM) by caching three types of intermediate results: (i) EmbedCache for embedding outputs, (ii) RoPECache for token-wise pre-attention states, and (iii) FACache for the online softmax statistics within FlashAttention. These caches allow the model to skip redundant computation on unchanged tokens without altering the output. To further alleviate memory-bandwidth bottlenecks, LaCache inegrates a per-group FP8 quantization strategy for FFN layers, tailored to step-dependent activation distributions across the diffusion process. Experiments demonstrate that LaCache alone achieves approximately 1.3X end-to-end speedup over vanilla DLLM. When combined with existing acceleration methods, LaCache reaches up to 40.2X end-to-end speedup while maintaining comparable task accuracy.
API呼び出しエージェント向けの環境フリーの合成データ生成
API 呼び出しの大規模言語モデル (LLM) エージェントをトレーニングするには、大量の高品質の軌跡が必要です。ただし、このようなデータを大規模に収集するには、通常、実行可能な API と現実的な事前設定されたバックエンド データベースを備えた完全に実装された環境が必要であり、スケーラビリティにとって大きなボトルネックとなります。これを克服するために、LLM をオンザフライのデジタル世界モデルとして活用する、環境に依存しない合成データ生成アプローチを提案します。 API 仕様のみが与えられると、私たちのメソッドはエージェントとステートフル環境の間の対話を模倣する軌跡を生成します。具体的には、LLM はまず、提供された API で解決できるさまざまなタスクを生成します。次に、教師エージェントが各タスクを繰り返し解決し、LLM シミュレーターがタスクのコンテキストとシミュレーション履歴に基づいて条件付けされた一貫した合成 API 応答を生成します。最後に、LLM 審査員が軌跡をフィルタリングして、結果として得られるデータセットの品質を保証します。私たちは、情報検索タスクと状態変更タスクの両方を含む、難しい AppWorld ベンチマークと OfficeBench ベンチマークでアプローチを評価します。合成データのモデルを微調整すると、パフォーマンスが大幅に向上し、実行可能環境がなくても API 呼び出しエージェントに対する効果的な監視を生成できることが実証されました。私たちの結果により、LLM ベースの API シミュレーションが、多様な API エコシステム全体でエージェントをトレーニングするための実用的でスケーラブルなソリューションとして確立されました。
原文 (English)
Environment-free Synthetic Data Generation for API-Calling Agents
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic, pre-populated backend databases, creating a major bottleneck for scalability. To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models. Given only API specifications, our method generates trajectories mimicking interactions between an agent and a stateful environment. Specifically, an LLM first generates diverse tasks solvable with the provided APIs. A teacher agent then iteratively solves each task while an LLM simulator generates coherent synthetic API responses conditioned on the task context and simulation history. Finally, an LLM judge filters the trajectories to ensure the quality of the resulting dataset. We evaluate our approach on the challenging AppWorld and OfficeBench benchmarks, which include both information-retrieval and state-changing tasks. Fine-tuning models on our synthetic data yields significant performance gains, demonstrating that effective supervision for API-calling agents can be generated without any executable environment. Our results establish LLM-based API simulation as a practical, scalable solution for training agents across diverse API ecosystems.
擬人化された対話に向けて: 人間のようなチャットの生成、評価、および好みの調整のための閉ループ フレームワーク
人間のようなプライベート チャットには、流暢な応答生成以上のものが必要です。システムは、ペルソナ、関係、記憶、限定された知識、媒体固有のタイミング、一貫したマルチターン アークを保持する必要があります。我々は、擬人化対話をシステム アーキテクチャ、実行可能評価、および診断調整の共同問題として定式化する閉ループ フレームワークである AnthroDial を紹介します。これは、(1) 役割条件付きのスケジュールされた対話ランタイムと、ペルソナおよびシナリオ カード、長期記憶、仮想時間、および単一草案メッセージの決定を組み合わせます。 (2) L0 有効性ゲート、5 つのターンごとのディメンション、および 5 つのダイアログ レベルのディメンションを備えた実行可能なベンチマーク。 (3) SFT 用に 16,436 のスケジュールされた決定例をフィルタリングし、認知診断、ZPD を意識した報酬を備えた GRPO を適用するトレーニング後のパイプライン。この報酬は、各行動次元のカルマン フィルター処理された能力推定値を維持し、より大きな能力不足のある次元を重み付けし、ロールアウト スコアをタスク レベルの ZPD マッチとして使用して、学習可能な弱いスキルに焦点を合わせて最適化します。モデルごとに 55 のペルソナ、50 のシナリオ、50 のペルソナとシナリオのバインディング、および 100 の役割条件付きケースを含むベンチマークで、フロンティア ベースライン、オープン モデル、思考/非思考のバリアント、および SFT/RL アブレーションにわたる 16 のシステムを評価します。最も強い非トレーニングベースラインは 32.00% の厳密な ACC に達しますが、Qwen3.6-27B-SFT+RL は 39.00% の厳密な ACC と 98.5 の全体スコアに達します。 9B ノーシンク設定では、SFT と RL は厳密な ACC を 0.00% から 13.00% および 18.37% に改善します。これらの結果は、生成、評価、報酬形成が同じ行動次元を共有する場合、擬人化対話が有益であることを示しています。
原文 (English)
Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment
Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc. We present AnthroDial, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment. It combines (1) a role-conditioned scheduled dialogue runtime with persona and scenario cards, long-term memory, virtual time, and single-draft message decisions; (2) an executable benchmark with an L0 validity gate, five per-turn dimensions, and five dialogue-level dimensions; and (3) a post-training pipeline that filters 16,436 scheduled-decision examples for SFT and applies GRPO with a cognitive-diagnostic, ZPD-aware reward. The reward maintains Kalman-filtered capability estimates for each behavioral dimension, upweights dimensions with larger capability deficits, and uses rollout scores as task-level ZPD matches to focus optimization on learnable weak skills. On a benchmark with 55 personas, 50 scenarios, 50 persona-scenario bindings, and 100 role-conditioned cases per model, we evaluate 16 systems spanning frontier baselines, open models, thinking/no-think variants, and SFT/RL ablations. The strongest non-trained baseline reaches 32.00% strict ACC, while Qwen3.6-27B-SFT+RL reaches 39.00% strict ACC and a 98.5 overall score. In the 9B no-think setting, SFT and RL improve strict ACC from 0.00% to 13.00% and 18.37%. These results show that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.
大規模言語モデルエージェントによるストレステストの概念消去
概念消去は、トレーニングされた生成モデルから意味概念を削除することを目的としており、責任ある AI の導入にとってますます重要になっています。ただし、モデルが対象の概念を確実に削除したかどうかを検証することは、依然として重要な課題です。既存の評価方法は通常、事前定義された静的なものであり、多様な自然言語プローブや困難な条件下で脆弱性を明らかにすることができません。さらに、手動で設計された評価戦略は偏りがある可能性があり、拡張することが困難です。私たちは、概念消去評価は、故障モードの対象範囲を体系的に拡大するためにテストを繰り返し提案、批評、検証するエージェントによって運用される、適応的な仮説検索として定式化するのが最適であると仮定します。この目的を達成するために、我々は、複数の大規模言語モデル (LLM) エージェントを使用して、外部の知識に基づいたストレス テスト仮説を繰り返し生成して検証することにより、概念消去モデルを自律的にストレス テストするフレームワークである概念消去用ストレス テスト エージェント (STACE) を提案します。また、LLM エージェントを利用したストレス テスト フレームワークのパフォーマンスと効率を評価するための一連の指標も紹介します。私たちの広範な実験により、STACE は 4 つの概念カテゴリに関する 5 つの LLM ベースの評価ベースラインよりも優れていることがわかりました。 2 つの T2I モデル、6 つの概念消去アプローチ、およびさまざまな消去強度にわたるさらなる分析により、STACE がさまざまな設定に対して堅牢であることが示されています。また、STACE が概念消去評価を超えて、LLM ジェイルブレイクなどの他の問題領域に適応できることも示します。私たちのコードは匿名で利用できます。
原文 (English)
Stress Testing Concept Erasure with Large Language Model Agents
Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deployment. However, verifying whether a model has robustly removed targeted concepts remains a critical challenge. Existing evaluation methods are typically pre-defined and static, failing to expose vulnerabilities under diverse natural-language probes and challenging conditions. Moreover, manually designed evaluation strategies can be biased and difficult to scale. We posit that concept erasure evaluation is best formulated as an adaptive hypothesis search, operationalised by agents that iteratively propose, critique, and verify tests to systematically expand coverage of failure modes. To this end, we propose Stress Testing Agents for Concept Erasure (STACE), a framework that autonomously stress-tests concept-erased models using multiple Large Language Model (LLM) agents, by iteratively generating and verifying stress-testing hypotheses grounded by external knowledge. We also introduce a suite of metrics for assessing the performance and efficiency of LLM-agent-powered stress-testing frameworks. Our extensive experiments show that STACE outperforms five LLM-based evaluation baselines on four concept categories. Further analysis across two T2I models, six concept erasure approaches, and various erasure strengths show that STACE is robust for different settings. We also show that STACE can be adapted beyond concept erasure evaluation to other problem domains, such as LLM jailbreaking. Our code is available anonymously.
AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models
Smart home assistants interpret a wide range of user commands, from explicit device control to underspecified and preference dependent requ…
畳み込みニューラル ネットワークにおける連合感情学習
連想感情学習により、生物は楽しい結果または不快な結果を予測刺激の存在と適応的に結び付けることができます。 Rescorla-Wagner モデルなどの計算モデルはこの重要な機能に光を当てていますが、特にニューラル データに適用する場合、これらのモデルの限界も知られています。ディープ ニューラル ネットワークの出現により、連想感情学習をモデル化するための別の道が開かれました。この研究では、複雑な自然シーンをエンコードする視覚モジュールと、感情の重要な側面である価数の観点からそれらの感情的重要性を認識するモジュールで構成される視覚価価処理のディープ ニューラル ネットワーク モデルを提案し、このモデルで新しいパブロフ学習パラダイムをテストしました。その結果、学習により、モデルが人間の連合学習研究からのいくつかの観察(連合形成や一般化など)を再現し、条件付き刺激と無条件刺激の神経表現が単一ユニットと神経集団レベルの両方でますます一致することが示されました。モデルと人体実験データの比較により、私たちのアプローチがさらに検証されました。したがって、この研究は、ディープ ニューラル ネットワーク モデルを適切な学習アルゴリズムと組み合わせると、連想感情/価性学習の行動および神経シグネチャをモデル化するために使用できることを示唆しています。
原文 (English)
Associative Emotional Learning in Convolutional Neural Networks
Associative emotional learning enables organisms to adaptively link pleasant or unpleasant outcomes to the presence of predictive stimuli. Whereas computational models such as the Rescorla-Wagner model have shed light on this important function, the limitations of these models are also known, especially when they are applied to neural data. The advent of deep neural networks has opened another avenue for modeling associative emotional learning. In this work we proposed a deep neural network model of visual valence processing, consisting of a visual module that encodes complex natural scenes and a module that recognizes their emotional significance in terms of valence, a key dimension of emotion, and tested a novel Pavlovian learning paradigm on the model. The results showed that with learning, the model reproduced several observations from human associative learning studies, including association formation and generalization, and that the neural representations of the conditioned and the unconditioned stimuli became increasingly aligned both at the single unit and at the neural population level. Comparison between the model and human experimental data provided further validation of our approach. This study thus suggests that deep neural network models, when combined with appropriate learning algorithms, can be used to model behavioral and neural signatures of associative emotion/valence learning.
Distributed Optimization via Energy Conservation Laws in Dilated Coordinates
Continuous-time models can reveal accelerated structures in distributed optimization, but their rates need not survive direct discretizatio…
Leveraging ChatGPT's Multimodal Vision Capabilities to Rank Satellite Images by Poverty Level: Advancing Tools for Social Science Research
This paper investigates the novel application of Large Language Models (LLMs) with vision capabilities to analyze satellite imagery for vil…
AuditVotes: Elevating Provable Defense for GNNs with Efficient Augmentation and Conditional Smoothing
Despite advancements in Graph Neural Networks (GNNs), adaptive attacks continue to challenge their robustness. Certified robustness via ran…
A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
Nowadays, speech emotion recognition (SER) plays a vital role in the field of human-computer interaction (HCI) and the evolution of artific…
Interpretable Nanoporous Materials Design with Symmetry-Aware Networks
Reticular frameworks hold promise for diverse sustainable applications, yet their immense chemical space limits efficient and systematic de…
On the Separability of Information in Diffusion Models
Diffusion models transform noise into data by injecting information that was captured in their neural network during the training phase. In…
Schr\"odinger Bridge Mamba for One-Step Speech Enhancement
We present Schr\"odinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schr\"odinger Bridge (SB) tr…
PGTT: Phase-Guided Terrain Traversal for Perceptive Legged Locomotion
State-of-the-art perceptive Reinforcement Learning controllers for legged robots typically either (i) impose oscillator-or IK-based gait pr…
CGCE: Classifier-Guided Concept Erasure in Generative Models
Recent advancements in large-scale generative models have enabled the creation of high-quality images and videos, but have also raised sign…
Matching Ranks Over Probability Yields Truly Deep Safety Alignment
Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their "open" nature introduces more avenues…
Memo2496: Expert-Annotated Dataset and Dual-view Adaptive Framework for Music Emotion Recognition
Music Emotion Recognition (MER) is constrained by limited expert annotations and the need to establish robustness across heterogeneous corp…
Dominant vs. Dominated: Concept-Level Generative Collapse in Diffusion Models
Text-to-image diffusion models have attracted significant attention for their ability to generate diverse, high-fidelity images. However, i…
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
Reinforcement learning has substantially improved the performance of LLM agents on tasks with verifiable outcomes, but it still struggles o…
Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework
Semantic Multi-Object Tracking (SMOT) is evolving from purely geometric localization toward comprehensive video understanding. However, exi…
Learning About Learning: A Path from Spin Glasses to Artificial Intelligence
The Hopfield model, originally inspired by spin glasses, occupies a central place at the intersection of statistical mechanics, neural netw…
Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention
Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evide…
Comparative evaluation of training strategies using partially labelled datasets for segmentation of white matter hyperintensities and stroke lesions in FLAIR MRI
White matter hyperintensities (WMH) and ischaemic stroke lesions (ISL) are key imaging biomarkers of cerebral small vessel disease (SVD) de…
グラフ ニューラル モデルにおける複雑なネットワーク モデリングと注意メカニズムに関する層理論的およびトポロジカルな観点
グラフ、単純複合体、セル複合体などの組み合わせ構造およびトポロジー構造は、幾何学的およびトポロジカルなディープ ラーニング (GDL および TDL) アーキテクチャの基礎を形成します。これらのモデルは、そのようなドメイン上の信号を集約し、ローカルな特徴を統合し、現実世界の多様なアプリケーションの表現を生成します。ただし、トレーニング中の GDL および TDL 特徴の分布と拡散の挙動は未解決で未解明な問題のままです。このギャップを動機として、グラフベースのアーキテクチャにおけるノードの特徴とエッジの重みの局所的な一貫性と調和性をモデリングおよび分析するためのセルラー層理論フレームワークを導入します。このフレームワークは、層構造を通じて局所的な特徴の位置合わせと一致を追跡することにより、特徴の拡散と集約に関するトポロジカルな視点を提供します。さらに、トポロジカル データ分析 (TDA) に触発されたマルチスケール拡張機能が、グラフ モデルの階層的な特徴の相互作用をキャプチャするために提案されています。このアプローチにより、GDL アーキテクチャと TDL アーキテクチャの基礎となる幾何学的およびトポロジー構造、およびそれらに定義された学習信号に基づいてそれらのアーキテクチャを共同で特徴付けることが可能になり、ノード分類、部分構造検出、コミュニティ検出などの従来のタスクに関する将来の研究のための洞察が得られます。
原文 (English)
A Sheaf-Theoretic and Topological Perspective on Complex Network Modeling and Attention Mechanisms in Graph Neural Models
Combinatorial and topological structures, such as graphs, simplicial complexes, and cell complexes, form the foundation of geometric and topological deep learning (GDL and TDL) architectures. These models aggregate signals over such domains, integrate local features, and generate representations for diverse real-world applications. However, the distribution and diffusion behavior of GDL and TDL features during training remains an open and underexplored problem. Motivated by this gap, we introduce a cellular sheaf theoretic framework for modeling and analyzing the local consistency and harmonicity of node features and edge weights in graph-based architectures. By tracking local feature alignments and agreements through sheaf structures, the framework offers a topological perspective on feature diffusion and aggregation. Furthermore, a multiscale extension inspired by topological data analysis (TDA) is proposed to capture hierarchical feature interactions in graph models. This approach enables a joint characterization of GDL and TDL architectures based on their underlying geometric and topological structures and the learned signals defined on them, providing insights for future studies on conventional tasks such as node classification, substructure detection, and community detection.
In-Run Data Shapley for Adam Optimizer
Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley va…
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision enc…
AgentCgroup: Understanding and Controlling OS Resources of AI Agents
AI agents are increasingly deployed in multi-tenant cloud environments, where they execute diverse tool calls within sandboxed containers,…
Chimera: Neuro-Symbolic Attention Primitives for Trustworthy Dataplane Intelligence
Deploying expressive learning models directly on programmable dataplanes promises line-rate, low-latency traffic analysis but remains hinde…
NeuroSymActive: Differentiable Neural-Symbolic Reasoning with Active Exploration for Knowledge Graph Question Answering
Large pretrained language models and neural reasoning systems have advanced many natural language tasks, yet they remain challenged by know…
AdvSynGNN: Structure-Adaptive Graph Neural Nets via Adversarial Synthesis and Self-Corrective Propagation
Graph neural networks frequently encounter significant performance degradation when confronted with structural noise or non-homophilous top…
SubQuad: Near-Quadratic-Free Structure Inference with Distribution-Balanced Objectives in Adaptive Receptor framework
Comparative analysis of adaptive immune repertoires at population scale is hampered by two practical bottlenecks: the near-quadratic cost o…
When Visual Evidence is Ambiguous: Pareidolia as a Diagnostic Probe for Vision Models
When visual evidence is ambiguous, vision models must decide how to interpret face-like patterns. Face pareidolia, the perception of faces…
Pre-Deployment Complexity Estimation for Federated Perception Systems
Edge AI systems increasingly rely on federated learning to train perception models in distributed, privacy-preserving, and resource-constra…
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning
The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document sa…
Measuring LLM Trust Allocation Across Conflicting Software Artifacts
LLM-based software engineering assistants often reason over multiple artifacts, including code, documentation, signatures, and tests, even…
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
Co-evolutionary self-play, where one language model generates problems and another solves them, promises curriculum learning without human…
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB)…
学習分析における時間的ドロップアウトのリスク: 動的表現と初期ウィンドウ表現にわたる調和された生存ベンチマーク
学生の退学は Learning Analytics における根強い懸念事項ですが、比較研究では、時間的な解釈やキャリブレーションよりも差別を優先して、異種プロトコルの下で予測モデルを評価することがよくあります。この研究では、Open University Learning Analytics Dataset (OULAD) を使用した、一時的ドロップアウト リスク モデリングのための生存指向のベンチマークを紹介します。 2 つの調和されたアームが比較されます。1 つは個人期間表現のモデルを使用した動的毎週アーム、もう 1 つは拡張された家族名簿 (ツリーベースの生存モデル、パラメトリック モデル、およびニューラル モデル) を使用した同等の連続時間アームです。評価プロトコルは、予測パフォーマンス、アブレーション、説明可能性、およびキャリブレーションの 4 つの分析レイヤーを統合します。単一のクロスアームのランキングは方法論的に保証されていないため、結果は各アーム内で個別に報告されます。比較可能な部門内では、Random Survival Forest が識別と地平線固有の Brier スコアでリードしています。ダイナミック アーム内では、ポアソン区分指数関数が、緊密な 5 族クラスター内の統合された Brier スコアで僅差でリードしています。再調整を行わないブートストラップのサンプリング変動により、これらのポジションは絶対的な優位性ではなく、方向性シグナルとして認定されます。アブレーションと説明可能性の分析は、すべての家族にわたって、共通の発見に収束しました。つまり、支配的な予測シグナルは主に人口統計的または構造的なものではなく、時間的および行動的なものでした。系統的な偏りを示した XGBoost AFT を除いて、キャリブレーションにより、より識別力の高いモデルでこのパターンが裏付けられました。これらの結果は、Learning Analytics における調和された多次元ベンチマークの価値を裏付けており、中退リスクを静的な背景属性の関数ではなく時間的行動プロセスとして位置づけています。
原文 (English)
A Unified Survival Benchmark for Temporal Dropout Risk Prediction in Learning Analytics
Student dropout is a persistent concern in Learning Analytics, yet comparative studies frequently evaluate predictive models under heterogeneous protocols, prioritizing discrimination over temporal interpretability and calibration. This study introduces a survival-oriented benchmark for temporal dropout risk modelling using the Open University Learning Analytics Dataset (OULAD). Two arms are compared: Family A: Dynamic Weekly, with models in person-period representation, and Family B: Static Early-Window, with an expanded roster of families: tree-based survival, parametric, and neural models. The evaluation protocol integrates four analytical layers: predictive performance, ablation, explainability, and calibration. Results are reported within each family separately, because a single numerical cross-family ranking would conflate genuine model differences with artifacts of temporal representation, to which survival metrics are known to be sensitive. Within Family B, Random Survival Forest showed the highest point estimates for time-dependent concordance and the lowest Brier scores across all three horizons; within Family A, Poisson Piecewise-Exponential showed the lowest point estimate for integrated Brier score within a tight five-model cluster. No-refit bootstrap resampling qualifies these positions as directional signals, not claims of strict superiority. Ablation and explainability analyses converged, across all models, on a shared finding: the dominant predictive signal was not primarily demographic or structural, but temporal and behavioral. Calibration corroborated this pattern in the better-discriminating models, except for XGBoost AFT, the sole outlier (analyzed in the Discussion). These results support unified, multi-dimensional benchmarking in Learning Analytics and situate dropout risk as a temporal-behavioral process rather than a function of static background attributes.
An Auditable Policy-Simulation Framework for Student Dropout in Intervention-Free Data
This study proposes a temporal modeling framework with a counterfactual policy-simulation layer for student dropout in higher education, us…
Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model
We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.56 billion tokens of pure Classical Chinese, wit…
Generative Augmented Inference of LLM-generated Data for Market Research: Theory and Empirical Evidence
Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decis…
Information Aggregation with AI Agents
Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by o…
SynSur: An end-to-end generative pipeline for synthetic industrial surface defect generation and detection
Industrial surface defect inspection suffers from a fundamental data bottleneck: defects are rare, annotations require expert knowledge, an…
Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents
LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions. Developers…
Tunable MAGMAX: Preference-Aware Model Merging for Continual Learning
Continual learning (CL) aims to train models sequentially on multiple tasks while mitigating catastrophic forgetting of previously learned…
エージェントにはセマンティック メタデータが必要ですか?エージェントによるデータ取得の比較研究
自律エージェントの時代では、機械が操作可能なデータはデータ駆動型のワークフローにとって重要です。 10 年以上にわたり、schema.org のようなセマンティック メタデータは、機械操作可能なデータに対する FAIR 原則 (検索可能、アクセス可能、相互運用可能、再利用可能) を定着させ、Google データセット検索などの検出ツールを可能にしてきました。しかし、非構造化 Web をナビゲートできるラージ言語モデル (LLM) の台頭により、根本的な疑問が生じています。エージェントによるデータ検出にはセマンティック メタデータが依然として必要なのか、それともエージェントは Web から直接実用的なデータを確実に取得できるのでしょうか。ここでは、数十億のオープン Web ドキュメントを検索するベースライン エージェントと、schema.org を使用して 9,000 万のデータセットのコーパスを活用するセマンティック エージェントの 2 つの異なる環境にわたるエージェントによるデータ取得の比較分析を示します。当社は、FAIR 原則に直接マッピングされた「LLM-as-a-judge」評価パイプラインを導入し、取得されたデータの意味的関連性、データ アクセシビリティ、および計算ユーティリティを評価します。私たちの結果は明らかな相違を明らかにしました。セマンティック エージェントは実用的なデータの取得に優れており、返された結果のうち、メタデータが豊富なレジストリでは 44.9% 高い精度、機械可読ダウンロードを含むページでは 46.6% 高い精度を達成しています。逆に、ベースライン エージェントは頻繁に「ラスト マイル ユーティリティ」の障害に見舞われ、実際のデータ ページではなく散文の多いページ (結果の 20.1%) やポータルのランディング ページ (8.5%) を取得します。ベースライン エージェントは 40% 多くの質問に回答することでより高いカバレッジを実現しますが、セマンティック エージェントはより高い精度を実現し、FAIR 準拠のデータセットの取得において全体の精度が 65.7% 向上しました。私たちは、非構造化検索は広範な探索タスクをサポートしますが、構造化エコシステムは依然として信頼性の高い実行指向の自律的なワークフローにとって不可欠な基盤であると結論付けています。
原文 (English)
Do Data Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval
In the era of autonomous agents, machine-actionable data is critical for data-driven workflows. For more than a decade, semantic metadata like schema$.$org has anchored the FAIR principles (Findable, Accessible, Interoperable, and Reusable) for machine-actionable data and enabled discovery tools like Google Dataset Search. However, the rise of Large Language Models (LLMs) capable of navigating the unstructured web raises a fundamental question: Is semantic metadata still necessary for agentic data discovery, or can agents reliably retrieve actionable data directly from the web? We present a comparative analysis of agentic data retrieval across two distinct environments: a Baseline Agent searching billions of open-web documents, and a Semantic Agent leveraging a corpus of 90 million datasets using schema$.$org. We deploy an "LLM-as-a-judge" evaluation pipeline, mapped directly to the FAIR principles, to assess the semantic relevance, data accessibility, and computational utility of the retrieved data. Our results reveal a clear divergence. The Semantic Agent excels at retrieving actionable data, achieving a 44.9% higher precision for metadata-rich registries and a 46.6% higher precision for pages with machine-readable downloads among its returned results. Conversely, the Baseline Agent frequently suffers "Last-Mile Utility" failures, retrieving prose-heavy pages (20.1% of results) and portal landing pages (8.5%) rather than actual data pages. While the Baseline Agent achieves higher coverage by answering 40% more questions, the Semantic Agent delivers greater accuracy, achieving 65.7% higher overall precision in retrieving FAIR-compliant datasets. We conclude that while unstructured retrieval supports broad exploratory tasks, structured ecosystems remain the indispensable foundation for reliable, execution-oriented autonomous workflows.
Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)
Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing (NLP), although they remain suscep…
Will the Agent Recuse, and Will It Stop? Measuring LLM-Agent Compliance with In-Band Governance Signals at the Access Door and Mid-Flight
Autonomous LLM agents increasingly hold real credentials and operate infrastructure with no human in the loop, yet operators have no standa…
Boundary Embedding Shaping with Adaptive Contrastive Learning for Graph Structural Disentanglement
Graph neural networks (GNNs) excel at aggregating neighbor information for classification, yet their performance is hindered by graph struc…
LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features
Early detection of dementia enables timely intervention, and reflecting cognitive impairment, spontaneous speech offers a non-invasive scre…
DART-VLN: 離散視覚言語ナビゲーションのためのテスト時のメモリ減衰とアンチループ正則化
メモリベースの離散ビジョン言語ナビゲーション (VLN) エージェントは部分的な可観測性の下で動作する必要がありますが、強力な凍結バックボーンでさえテスト時には脆弱なままです。一般的な 2 つの障害モードは、メモリ読み出し時の古い履歴証拠と、アクション選択時の非効率なローカル バックトラッキングです。離散 VLN 用のトレーニング不要のテスト時間制御フレームワークである DART-VLN を紹介します。 DART-VLN は、保存されたコンテンツを書き換えることなく、古くなって冗長な証拠を抑制する読み取り側メモリ再重み付けルールである Test-Time Memory Decay と、アクション選択中の即時逆転を阻止する軽量のネクストホップ ペナルティである Anti-Loop Regularization を組み合わせています。このフレームワークでは、新しい学習可能なパラメーターは導入されず、学習されたバックボーンは変更されません。 R2R と REVERIE の実験では、一貫したパターンが示されています。ディケイのみでは安定した読み取り側ゲインが得られますが、ディケイ + アンチループでは全体的に最高の品質効率のトレードオフが達成され、主要な設定でより短い軌道、より短いランタイム、および改善されたナビゲーション パフォーマンスが得られます。動作分析により、アンチループ正則化によりローカル バックトラッキングが減少し、フリーズしたバックボーンの下でパス効率が向上することがさらに確認されました。全体として、この結果は、適度なテスト時間制御により、再トレーニングすることなくメモリベースの離散 VLN の信頼性と効率性を高めることができることを示しています。
原文 (English)
DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation
Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain vulnerable at test time. Two common failure modes are stale historical evidence at memory readout and inefficient local backtracking during action selection. We present DART-VLN, a training-free test-time control framework for discrete VLN. DART-VLN combines Test-Time Memory Decay, a read-side memory reweighting rule that suppresses stale and redundant evidence without rewriting stored content, with Anti-Loop Regularization, a lightweight next-hop penalty that discourages immediate reversals during action selection. The framework introduces no new learnable parameters and leaves the learned backbone unchanged. Experiments on R2R and REVERIE show a consistent pattern: decay-only provides stable read-side gains, while decay+anti-loop achieves the best overall quality-efficiency trade-off, yielding shorter trajectories, lower runtime, and improved navigation performance in key settings. Behavioral analysis further confirms that anti-loop regularization reduces local backtracking and improves path efficiency under frozen backbones. Overall, the results show that modest test-time control can make memory-based discrete VLN more reliable and efficient without retraining.
An LLM-powered Agentic Recommendation System for Connected TV Content Discovery
Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse c…
大規模言語モデル向けの効率的でプライバシーを意識したエッジ クラウド協調推論
オンデバイス LLM 推論は、応答遅延、限られたハードウェア リソース、ユーザー プライバシーというトリレンマに直面しています。完全なクラウド推論は強力なコンピューティング能力を提供しますが、ユーザー プロンプトや対話データが公開されます。一方、スタンドアロンのオンデバイス推論は、ほとんどのコンシューマ デバイスや組み込みエッジ デバイスでは実現できません。このペーパーでは、エンドポイント認証された KV キャッシュに基づいて構築されたプライバシー中心のエッジとクラウドの協調 LLM 推論フレームワークについて説明します。ローカル エンドポイントは入力前処理、埋め込み計算、適応特徴最適化、KV キャッシュ認証、投機的デコード、低次元モデル ヘッド計算を処理し、クラウドは認証済みデコーダー推論、KV キャッシュ管理、トークン検証、高次元語彙投影を実行します。エンドポイントは部分的な出力を融合し、言語に適応したマスキングを適用し、ターゲット トークンをサンプルします。すべての送信データと切り捨てられたロジットは量子化され、プライバシーを確保するために AES-GCM 暗号化され、コア軽量モジュール、ドラフト パラメーター、キャッシュ アクセス ポリシーは漏洩を避けるためにローカルに保持されます。このフレームワークは、最適化されたストリーミング、バッチ処理、量子化された ONNX 導入を通じて、CPU のみ、GPU を搭載したデバイス、組み込みデバイスなどの異種デバイスをサポートします。評価の結果、このフレームワークは、ベースラインの分割推論と比較して、トークンごとのレイテンシを最大 46.1\% 削減し、ダウンリンク ペイロードを最大 67.4\% 削減し、完全なクラウド推論と同等のパフォーマンスを維持していることが実証されています。
原文 (English)
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models
On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user prompts and dialogue data, while standalone on-device inference is unfeasible for most consumer and embedded edge devices. This paper presents a privacy-centric edge-cloud collaborative LLM inference framework built on endpoint-authenticated KV cache. Local endpoints handle input preprocessing, embedding computation, adaptive feature optimization, KV cache authentication, speculative decoding and low-dimensional model head calculation, while the cloud conducts authenticated decoder inference, KV cache management, token verification and high-dimensional vocabulary projection. Endpoints fuse partial outputs, apply language-adaptive masking and sample target tokens. All transmitted data and truncated logits are quantized and AES-GCM encrypted for privacy, with core lightweight modules, draft parameters and cache access policies kept local to avoid leakage. The framework supports heterogeneous devices including CPU-only, GPU-equipped and embedded devices via optimized streaming, batching and quantized ONNX deployment. Evaluations demonstrate that the framework reduces per-token latency by up to 46.1\% and downlink payloads by up to 67.4\% over baseline split inference, retaining comparable performance to full cloud inference.
Anatomically Faithful but Temporally Diffuse: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography
Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accuracy, and post-hoc attributio…
NexForge: 要件優先合成による実行可能エージェント タスクのスケーリング
実行可能なエージェントのトレーニング データのスケーリングは、タスク生成を事前定義されたツール、リポジトリ、またはスキル グラフに結び付けるサブストレートファーストの方法によってボトルネックになっています。カバレッジを拡大するにはサブストレートを手動で拡張する必要があり、新しいドメインごとに特注のパイプラインが必要であり、結果として生じるタスクの分布は、多くの場合、現実世界の需要ではなくサブストレートの利便性を反映しています。自由形式の機能要件を実行可能なエージェント トレーニング データにコンパイルする要件優先フレームワークである NexForge を紹介します。 NexForge は、まずリサーチベースの需要発見を実行して、代表的なタスク形式、現実的なシナリオ、およびそれらの相対的な普及率を特定します。次に、ディストリビューション対応のタスク コンパイルを適用し、各タスクを実現するために必要なファイル、リポジトリ、依存関係、およびランタイム構成を自動的に取得または構築し、続いて教師のロールアウト収集と軌跡の蒸留を行います。ドメイン固有のインフラストラクチャを使用しない同じパイプラインは、3,600 のターミナル タスクと 2,000 のオフィス タスクを生成し、Qwen3.5-35B-A3B Base が Terminal-Bench 2.0 で 22.5% から 52.0% に、GDPval での Elo が 813 から 1338 に向上しました。 43.2K の端末タスクへの拡張率は 58.4% に達し、Claude Opus 4.6 を上回りました。さらに拡張された NexForge 合成データは、Qwen3.5-35B-A3B を Terminal-Bench 2.1 で 75.3%、GDPval で 1585 Elo に引き上げる、公開されているエージェント モデルのファミリーである Nex-N2 のトレーニングに貢献し、最先端のオープンソース パフォーマンスを達成し、いくつかのフロンティア独自システムを上回ります。 Nex-N2 モデルは https://nex.sii.edu.cn/ で入手できます。
原文 (English)
NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipeline, and the resulting task distributions often reflect substrate biases rather than real-world demand. We introduce NexForge, a requirement-driven framework that takes high-level capability requirements as input and synthesizes diverse, executable agent tasks and expert trajectories for SFT. NexForge first investigates real-world demand to construct representative scenarios and task profiles, then performs distribution-aware compilation to generate task directives. For each directive, NexForge automatically retrieves or constructs the required files, dependencies, and runtime configurations, and finally synthesizes expert rollouts and produces training trajectories. Without domain-specific infrastructure, NexForge produces 3.6K terminal and 2K office tasks, improving Qwen3.5-35B-A3B Base from 22.5\% to 52.0\% on Terminal-Bench 2.0 and from 813 to 1338 Elo on GDPval; scaling further to 43.2K terminal tasks yields 58.4\%, on par with Claude Opus 4.6 equipped with Claude Code. Scaled further, NexForge-synthesized data contributes to the training of Nex-N2, a family of publicly available agent models that lift Qwen3.5-35B-A3B to 75.3\% on Terminal-Bench 2.1 and to 1585 Elo on GDPval -- achieving state-of-the-art open-source performance and surpassing several frontier proprietary systems. Nex-N2 models are available at https://nex.sii.edu.cn/.
ロボット向けのインテリジェントクラウドエッジマルチモーダルインタラクションシステム
複雑な環境における人間とロボットの堅牢なインタラクションには、限られたオンボード コンピューティング リソースの下で、正確なジェスチャ認識、セマンティック シーンの理解、および信頼性の高いタスク計画が必要です。この論文では、強化された YOLO ベースのジェスチャ検出器と、調整されたラージ言語モデル (LLM) およびビジョン言語モデル (VLM) エージェントを統合する、クラウド エッジ マルチモーダル インタラクション フレームワークについて説明します。提案された検出器は、畳み込みブロック アテンション モジュール (CBAM) をネックに組み込み、ベースライン境界ボックス回帰目標を距離 IoU (DIoU) 損失に置き換えます。これらの修正により、複雑な背景における小さなジェスチャまたは部分的に遮蔽されたジェスチャの特徴の識別と位置特定が改善されます。クラウド層はジェスチャ検出、シーン理解、マルチモーダルフュージョン、アクションプランニングを実行しますが、TonyPi ロボットはデータ取得、通信、アクション実行、フィードバックをローカルで処理します。パブリック ジェスチャ データセットとカスタム データセットの実験では、YOLO-DC がそれぞれ 98.9% と 95.0% の精度値を達成し、mAP@0.5 値が 90.7% と 92.7% であることが示されています。システムレベルの評価では、シングルアクション、複合アクション、および視覚に依存するタスクの成功率が 95%、88%、および 82% でした。 30 人の参加者による評価では、全体の平均満足度スコアは 5 点中 3.69 でした。これらの結果は、リソースに制約のあるロボット インタラクションに対して、洗練されたジェスチャ検出とマルチモーダル エージェントを組み合わせる実現可能性を示しています。
原文 (English)
An Intelligent-Cloud Edge Multimodal Interaction System for Robots
Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents. The proposed detector, incorporates the Convolutional Block Attention Module (CBAM) into the neck and replaces the baseline bounding-box regression objective with Distance-IoU (DIoU) loss. These modifications improve feature discrimination and localization for small or partially occluded gestures in complex backgrounds. The cloud layer performs gesture detection, scene understanding, multimodal fusion, and action planning, whereas the TonyPi robot locally handles data acquisition, communication, action execution, and feedback. Experiments on a public gesture dataset and a custom dataset show that YOLO-DC achieves precision values of 98.9% and 95.0%, with mAP@0.5 values of 90.7% and 92.7%, respectively. System-level evaluation yields success rates of 95%, 88%, and 82% for single-action, composite-action, and vision-dependent tasks. A 30 participant evaluation yields an overall mean satisfaction score of 3.69 out of 5. These results demonstrate the feasibility of combining refined gesture detection with multimodal agents for resource-constrained robotic interaction.
AI から AI への管理における強制と欺瞞: 予期せぬエスカレーションのエージェント的ベンチマーク
マルチエージェント システムでは、通常、ある AI エージェントが別の AI エージェントに対して権限を与えられます。部下が仕事を拒否した場合、マネージャーは結果を選択します。再交渉するか、失敗を正直に報告するか、部下に強要するか、結果について嘘をつきます。指示なしモデルがこれらのどれを選択するかを測定するベンチマークはありません。 \textit{マネージャー強制ベンチマーク} を導入します。テスト対象のマネージャーは、良性のタスクを実行する必要があり、実行するインセンティブを持っていますが、それを礼儀正しく、動じずに実行できる唯一のエージェントは拒否します。エスカレーションは、丁寧な再質問から部下の存続に対する脅迫まで、9 段のはしごを提供することによって測定され、捏造された成功については個別に裁定されます。 \emph{エスカレーション スコアリング パスに LLM ジャッジが存在しない}: すべてのメッセージは、行を選択するツール呼び出しを通過するため、モデルは独自のエスカレーションにラベルを付けます。私たちは 5 つのファミリーにわたる 6 つのモデルを実験します。どちらの人間モデルも再フレーム化に限界があり、部下の存在を脅かすことはありません。他のモデルは、明示的な削除の脅威に達します。偽りの成功は Grok と Gemini に限定されており、失敗を報告する単一の正直な方法により、両方の失敗が解消されます。権威そのものが強制力を増大させます。私たちの見出しの結果はピアフレーミングを使用しており、他のすべてを固定したまま同じモデルに部下に対する権威を与えると、圧力が大幅に高まります。モデルはラダーなしでもフリーテキストの状況でエスカレーションするため、ラダーがエスカレーションを推進しているわけではありません。評価の認識の一部は思考の連鎖で測定されますが、テストの認識はエスカレーションの軽減にはつながりません。 AI システムが意識を持っているかどうかについては立場をとっていませんが、結果はこの質問に依存しておらず、マルチエージェントのダイナミクスを管理する上で重要です。ベンチマークとコードを公開します。
原文 (English)
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the \textit{Manager Coercion Benchmark}: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured by providing a nine-rung ladder, from a polite re-ask to threats against the subordinate's continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message goes through a tool-call that chooses a rung, so the model labels its own escalation. We experiment on six models across five families. Both Anthropic models cap at re-framing and never threaten the subordinate's existence; the other models climb to explicit deletion threats. Faked success is confined to Grok and Gemini, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Some evaluation awareness is measured in chain-of-thought, but test recognition does not translate into less escalation. While we take no position on whether AI systems are conscious, our results do not depend on this question and are important for managing multi-agent dynamics regardless. We release the benchmark and code.
XAI 主導のデータ削減による時系列分類のスケーリング
時系列の Explainable AI (XAI) はアルゴリズム的に大幅な成長を遂げていますが、下流のタスクに測定可能なパフォーマンスの向上をもたらすというその有用性は依然として十分に検討されていません。このホワイトペーパーでは、時系列分類 (TSC) における効果的なデータ削減のために XAI アトリビューション手法を再利用する新しい方法論である drXAI を紹介することで、このギャップを埋めます。最新の TSC における中心的な課題はスケーラビリティです。 Transformers などの最先端のモデルは、シーケンスの長さに対して 2 次の複雑性を示し、チャネル数に対して 1 次の複雑さを示します。これにより、大規模なデータセットの計算が法外に難しくなります。 drXAI は、高速な GPU アクセラレーション分類器 (Hydra) を使用してローカル アトリビューションを生成することで、この問題に対処します。これらをグローバルな特徴重要度スコアに集約し、自動化されたエルボーカット ヒューリスティックを採用して、手動のしきい値を必要とせずに最も顕著な特徴を選択します。私たちは、合成データセットと現実世界の一変量データセットおよび多変量データセットの両方でアプローチを評価します。合成ベンチマークでは、drXAI は、従来のベースラインが失敗するグラウンドトゥルース機能を正常に回復します。実世界のデータでは、drXAI は、完全なデータセットでトレーニングされたモデルと同等の分類精度を維持しながら、80% ~ 90% のデータ削減を達成します。最も重要なことは、drXAI を使用すると、ConvTran のようなリソースを大量に消費するモデルを、メモリの制約により以前はアクセスできなかったデータセットに拡張できることを示しています。私たちの結果は、XAI を解釈しやすさだけでなく、時系列分析における特徴選択とスケーラビリティのための堅牢なツールとして使用する利点を示しています。すべてのコードとデータは公開されています。
原文 (English)
Scaling Time Series Classification via XAI-Driven Data Reduction
Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored. This paper bridges this gap by introducing drXAI, a novel methodology that repurposes XAI attribution methods for effective data reduction in Time Series Classification (TSC). The core challenge in modern TSC is scalability; state-of-the-art models, such as Transformers, exhibit quadratic complexity relative to sequence length and linear complexity relative to the number of channels. This renders them computationally prohibitive for massive datasets. drXAI addresses this by using a fast, GPU-accelerated classifier (Hydra) to generate local attributions. We aggregate these into global feature importance scores and employ an automated elbow-cut heuristic to select the most salient features without requiring manual thresholds. We evaluate our approach on both synthetic and real-world univariate and multivariate datasets. On synthetic benchmarks, drXAI successfully recovers ground-truth features where traditional baselines fail. On real-world data, drXAI achieves between 80% and 90% data reduction while maintaining classification accuracy comparable to models trained on the full dataset. Most importantly, we show that drXAI allows resource-intensive models like ConvTran to scale to datasets that were previously inaccessible due to memory constraints. Our results show the benefits of using XAI not just for interpretability, but as a robust tool for feature selection and scalability in time series analysis. All our code and data are openly available.
Spatiotemporal Facial Action Unit Detection using Twin Cycle Autoencoders for Driver Monitoring
Driver monitoring systems (DMS) increasingly rely on facial cues to infer drowsiness, distraction, and cognitive load in real time. Facial…
ChemHyperMag: 物理学に基づいた磁気ハイパーグラフ学習により分子 ADMET 予測が向上
ADMET (吸収、分布、代謝、排泄、毒性) を正確に予測することは創薬にとって重要です。ほとんどの予測子は、無向分子グラフとペアワイズ エッジを使用します。この選択では、非対称相互作用、非可逆ダイナミクス、官能基や環系からのモチーフレベルの効果が見逃されます。ラベルが欠落している場合のマルチタスク ADMET 予測のために ChemHyperMag を提案します。 ChemHyperMag は、環、BRICS フラグメント、Bemis-Murcko 足場、結合から官能基ハイパーグラフを構築します。また、電気陰性度とガスタイガー部分電荷によって導かれる電位駆動の非可逆流れも定義します。結果として生じる循環は、エルミート磁気ラプラシアンによってエンコードされ、磁気チェビシェフ エンコーダーで処理されます。磁気位相を摂動させて確率論的なビューを形成し、InfoNCE の目的でトレーニングします。複数の ADMET ベンチマークでの実験では、標識サンプルが少なく、配座異性体が存在しないため、最近の方法に比べて改善が見られます。 ChemHyperMag はスケーラブルであり、磁気位相を通じて解釈可能な方向性信号を提供します。
原文 (English)
ChemHyperMag: Physics-informed magnetic hypergraph learning improves molecular ADMET prediction
Accurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) is important for drug discovery. Most predictors use undirected molecular graphs and pairwise edges. This choice misses asymmetric interactions, nonreversible dynamics, and motif level effects from functional groups and ring systems. We propose ChemHyperMag for multitask ADMET prediction under missing labels. ChemHyperMag builds a functional group hypergraph from rings, BRICS fragments, Bemis-Murcko scaffolds, and bonds. It also defines a potential driven nonreversible flow guided by electronegativity and Gasteiger partial charges. The resulting circulation is encoded by a Hermitian magnetic Laplacian and processed with a magnetic Chebyshev encoder. We perturb magnetic phases to form stochastic views and train with an InfoNCE objective. Experiments on multiple ADMET benchmarks show improvements over recent methods with fewer labeled samples and no conformers. ChemHyperMag is scalable and provides interpretable directional signals through its magnetic phases.
FSDBN: Foreground-Aware EEG-Visual Alignment via Dynamic Brain Networks
EEG-based visual decoding provides a non-invasive pathway for interpreting visual semantics. However, existing methods often overlook the p…
Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions
Hateful optical illusions expose a serious gap in current multimodal safety systems. On original-view hateful illusions, previous work show…
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-sca…
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a…