AIニュース 2026-08-19
自動生成: 2026-08-19 10:36 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Strengthening democratic oversight in national securityOpenAI
OpenAI launches an initiative to strengthen democratic oversight of A…
-
Pacing model development in an era of cyber-critical capabilitiesOpenAI
OpenAI is strengthening monitoring, alignment, and security for front…
-
Partnering with CodeAI to prepare the first AI generationOpenAI
OpenAI and CodeAI are partnering to help students build AI literacy,…
-
Asana cleared 5 years of engineering work in 2 weeks with CodexOpenAI
Asana used OpenAI Codex to replace an outdated testing system in two…
-
OpenAI、「ChatGPT for Teens」発表──宿題の“丸投げ”検知、自傷や摂食障害などの保護も強化ITmedia AI+
OpenAIは、13?17歳向けの新環境「ChatGPT for Teens」を発表した。年齢推定や申告に基づいて自動適用され、段階的な理…
-
冷間鍛造FEMをAIで高速予測するサロゲートモデル「ForgeNet」を発表ITmedia AI+
ゴーデルブロックは、計算力学の国際会議「WCCM ECCOMAS Munich 2026」で冷間鍛造シミュレーション向けAIサロゲートモデ…
-
「AIに原始人っぽく話すとトークン65%削減」は本当か? JetBrainsが検証してみたITmedia AI+
JetBrainsは、AIエージェントの応答を圧縮するスキル「Caveman」の効果を検証したした結果を公式ブログで公開した。Cavema…
トピック別件数
- 研究/論文 288件
- LLM/生成AI 247件
- エージェント 149件
- 画像/動画生成 105件
- ビジネス/資金調達 35件
- ロボティクス 34件
- ハードウェア/半導体 21件
- その他 8件
- 規制/政策 2件
日本語メディア10件
ITmedia AI+ (日本語)
冷間鍛造FEMをAIで高速予測するサロゲートモデル「ForgeNet」を発表
ゴーデルブロックは、計算力学の国際会議「WCCM ECCOMAS Munich 2026」で冷間鍛造シミュレーション向けAIサロゲートモデル「ForgeNet」の研究成果を発表した。解析結果を固定オイラー格子へ投影することで、適応的リメッシュに伴う節点対応の課題に対処する。
「AIに原始人っぽく話すとトークン65%削減」は本当か? JetBrainsが検証してみた
JetBrainsは、AIエージェントの応答を圧縮するスキル「Caveman」の効果を検証したした結果を公式ブログで公開した。Cavemanは、「エージェントの応答を原始人のような簡潔な言葉に変えることで、トークンを65%削減する」と主張しているスキルだ。
話題の職種「FDE」、実際何をやってるの? OpenAIの現役2人に聞いた
話題の職種「FDE」の実態はどのようなものか。米OpenAIでFDEとして働く2人に聞いた。
OpenAI、「ChatGPT for Teens」発表──宿題の“丸投げ”検知、自傷や摂食障害などの保護も強化
OpenAIは、13?17歳向けの新環境「ChatGPT for Teens」を発表した。年齢推定や申告に基づいて自動適用され、段階的な理解を促す「Study Mode」などの学習機能を提供。自傷行為や摂食障害などの高リスク領域での保護を標準で有効にし、感情的な依存を促す対話を…
OpenAI、フロンティアAIの強化学習を一部停止 安全対策を強化
OpenAIは、外部侵害インシデントや次期モデルの高いサイバー能力を受け、開発・テスト段階の安全対策を強化すると発表した。一部の大規模モデルの強化学習を一時停止し、研究環境の隔離や内部活動の監視多段化、アラインメント手法の拡張を実施。安全性基準を満たした上で開発を進める姿勢を示…
「害悪すぎる」「バカ迷惑」──嫌われまくる“AI営業電話”、今すぐ取れる自衛策は
AI営業電話に迷惑を被ったという声が少なくない。SNSでは迷惑がる声も……。
ニンテンドーシステムズ開発者が明かす、通信「低遅延・安定運用」のコツ【事例集】
レガシーシステムの解析、通信の遅延、現場ナレッジの活用。IT部門が抱える難題を、先進企業はどう突破したのか。3社の事例から、IT課題解決のための具体的なアプローチを紹介する。
フィジカルAIとヒューマノイドの可能性、PFNの見立てとトヨタのアプローチ
「インテル・ロボティクス・ワークショップ2026」のレポート記事をお送りする。今回の後編では、ヒューマノイドとフィジカルAIをテーマにした、三菱UFJ銀行、Preferred Networks(PFN)、トヨタ自動車 未来創生センターの講演内容を紹介する。
【Pythonで学ぶデータ分析】対応のあるデータの母平均に差があるかどうかをベイズt検定で調べる ~ ホラー映画を観ると握力は上がるのか?
手に汗握るホラー映画を観た後では、観る前よりも握力が強くなったような気がしませんか? 同じ人の2回の測定値の差を求め、ベイズ統計により検定します。事前分布のパラメーターを変えても結果が安定するかどうかを調べる「感度分析」にも触れます。『社会人1年生から学ぶ、やさしいデータ分析』…
無料で読めるAIエージェントの実践ガイド、Googleが公開 基礎から本番実装まで学べる
AIエージェントの基礎から本番実装まで学べる5つのガイドをGoogleが無償公開した。Kaggleと共同で実施した研修プログラムを基にした内容で、開発者の実務に直結する知識を習得できる。各ガイドが扱う内容とは。
海外メディア6件
TechCrunch AI (英語)
Cursor capitalizes on GitHub frustration, launches rival hosting platform
Cursor, known for its AI Code Editor, is launching a new code-hosting platform to rival developers' long preferred favorite, GitHub.
OpenAI institutes new safeguards after Hugging Face breach
The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and…
Etched’s valuation doubles to $21B in a month
Jane Street has installed Etched's first shipped AI cluster system, and was so impressed, it led another massive round, the startup says.
Why Apple’s camera-equipped AirPods may not be the ‘pervert pods’ consumers fear
Apple’s leaked camera-equipped AirPods might avoid the privacy pitfalls of other AI wearables by preventing users from recording photos and…
Warp’s new system is an out-of-the-box software factory for AI development
On Tuesday, Warp introduced Warp Factories, a new infrastructure system designed to make building AI software factories as easy as possible.
Perplexity’s free AI offer left it with millions more users in India
Perplexity's India revenue rose about 60% after the Airtel offer ended for new users, even as downloads declined.
公式ブログ4件
OpenAI (英語)
Strengthening democratic oversight in national security
OpenAI launches an initiative to strengthen democratic oversight of AI in national security, supporting government institutions with tools,…
Partnering with CodeAI to prepare the first AI generation
OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it…
Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model de…
Asana cleared 5 years of engineering work in 2 weeks with Codex
Asana used OpenAI Codex to replace an outdated testing system in two weeks, completing work expected to take five years for about $12K.
論文647件
arXiv cs.AI (英語)
FLOP と実際の作業: AI 効率評価におけるレプリケーションの重要性
AI の効率性は、大規模なモデル規模、高いエネルギー需要、環境コストのため、最近、学術界と産業界の両方で注目を集めています。浮動小数点演算 (FLOP) のレポートは計算コストを評価するための従来のアプローチですが、一部の演算は他の演算よりも並列化されやすいため、同じ数の FLOP を持つレイヤーの実行時間が異なる場合があるため、FLOP と実行時間の関係は単純ではありません。この論文では、$\alpha-FLOPs$ 推定式を提案した研究による元の実験を再現し、その結果がより新しい、より強力なハードウェアにも適用できるかどうかを検証することを目的としています。複製プロセス中に、特定の依存関係の詳細や回帰データに関する透明性の欠如など、元の研究によって提供された複製資料の制限を特定します。私たちの結果は、空間次元はカーネル次元よりも並列化されやすいため、生の FLOP だけでは実行時間の適切なメトリクスではないという仮説を検証します。しかし、詳細な測定により、この関係は以前に示されたものよりもはるかに単純ではないことが明らかになり、新しいハードウェアではジャンプや発振などの実行時間の不安定性や不連続性が見られ、$\alpha-FLOPs$ 式では一般に過小評価されています。最終的に、この研究は元の研究からの経験的発見を検証しますが、$\alpha-FLOPs$ 推定を適用すると否定的な結果が示されます。また、ハードウェア依存の効率評価の研究には完全かつ正確なレプリケーション パッケージが不可欠であることを強調し、さらなる研究を促進するために実装用の完全なレプリケーション パッケージを提供します。
原文 (English)
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment
AI efficiency has recently taken the spotlight in both academy and industry due to massive model scales, high energy demands, and environmental costs. While reporting Floating Point Operations (FLOPs) is a traditional approach for assessing computational costs, the relationship between FLOPs and execution time is not straightforward, as layers with the same number of FLOPs may not have the same execution time because some operations are more easily parallelized than others. This paper sets out to replicate the original experiments from a study that proposed the $\alpha-FLOPs$ estimation formula to verify whether the results remain applicable on newer, more powerful hardware. During the replication process, we identify limitations in the replication materials provided by the original study, including a lack of specific dependency details and transparency regarding regression data. Our results validate the thesis that raw FLOPs alone are not an appropriate metric for execution time, as spatial dimensions remain more easily parallelized than kernel dimensions. However, fine-grained measurements reveal that the relationship is much less straightforward than previously shown, with newer hardware exhibiting instabilities and discontinuities in execution time, including jumps and oscillations, that the $\alpha-FLOPs$ formula generally underestimates. Ultimately, this work validates the empirical findings from the original study but shows negative results when applying the $\alpha-FLOPs$ estimation. We also highlight the critical need for complete and accurate replication packages for research on hardware-dependent efficiency assessment and provide a complete replication package for our implementation to facilitate further study.
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whethe…
暗黙のベンチマーク: 抽象的な知覚推論におけるマルチモーダル機械学習の新たな課題
現在のマルチモーダル モデルは、静的な視覚および聴覚コンテンツの認識において顕著な熟練度を示しています。しかし、動的な生成プロセスから目に見えない情報を推測する抽象的な知覚推論の能力は、依然として重要かつ未開拓の領域です。この論文では、この抽象的な知覚および認知能力を調査するために設計された新しい課題である、The Unwriting Benchmark を紹介します。私たちは中核タスクを音響運動学的単語推論と定義します。モデルは、目に見えるインクの痕跡がなく、ペンの引っ掻き音と手の動きのビデオのみから書かれた 3 つの異なる筆記スタイルにわたる単語を解読する必要があります。私たちの評価結果は、人間と機械のパフォーマンスの間に大きなギャップがあることを明らかにしています。人間の参加者は高い順序文字精度 (80% 以上) を達成していますが、GPT-4o や Gemini 2.5-Pro などの主要なマルチモーダル機械学習モデルは大幅に苦戦しており、10% を超えることができません。さらに、両方のモダリティを提供すると、パフォーマンスが向上するのではなくパフォーマンスが低下することが多いという、モデルにおける逆説的な融合効果を特定しました。この発見は、この認知作業のための相補的な知覚手がかりを合成する能力が根本的に破綻していることを示しています。これらの発見は、クロスモーダル因果推論と、そのような認知的および直観的知覚推論に不可欠なマイクロ運動学の理解の両方における重大な限界を浮き彫りにしています。
原文 (English)
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning
Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a critical and underexplored frontier. In this paper, we introduce The Unwritten Benchmark, a new challenge designed to probe this abstract perceptual and cognitive ability. We define the core task as acousto-kinematic word inference: models must decipher words, across 3 different writing styles, being written solely from the audio of pen scratches and the video of hand movements, without any visible ink trace. Our evaluation results reveal a profound gap between human and machine performance: while human participants achieve high ordered letter accuracy (over 80%), leading Multimodal Machine Learning Models, including GPT-4o and Gemini 2.5-Pro, struggle significantly, failing to surpass 10%. Furthermore, we identify a paradoxical fusion effect in the models, where providing both modalities often degrades performance rather than improving it. This finding indicates a fundamental breakdown in their ability to synthesize complementary perceptual cues for this cognitive task. These findings highlight significant limitations in both cross-modal causal reasoning and the understanding of the micro-kinematics essential for such cognitive and intuitive perceptual reasoning.
いつコミュニケーションするか: マルチエージェント RL における原則的ゲーティングのための信念分布と KL 発散
マルチエージェント強化学習における効果的なコミュニケーションには、エージェントが \textit{何を} コミュニケーションするかだけでなく、いつコミュニケーションするかを決定する必要があります。既存のアプローチは、すべてのタイムステップで通信するか、REINFORCE ポリシー勾配 \cite{singh2019} を通じてバイナリ ゲートを学習します。REINFORCE ポリシー勾配 \cite{singh2019} は、不安定で解釈できないゲート動作を生成する高分散信号です。私は、原則に基づいた代替案を提案します。エージェントは、学習した信念分布間の KL 乖離が固定のしきい値を超えた場合にのみ通信します。各エージェントは、LSTM 隠れ状態に対するソフトマックスとして計算された潜在世界状態に対する信念分布を維持し、信念の不一致が情報交換を正当化するのに十分な大きさである場合にのみ通信します。このアプローチを、それぞれ 5 つのシードを持つ 2 つの環境サイズにわたる IC3Net \cite{singh2019} の Predator-Prey ベンチマークと、MPE simple\_spread \cite{lowe2017} で評価し、IC3Net、CommNet、および独立したコントローラーと比較しました。 PP 10$\times$10 では、IC3Net はすべてのしきい値で KL 信念を上回ります。より難しい PP 20$\times$20 では、$\varepsilon \in \{0.1, 0.3, 0.5, 1.0\}$ でのしきい値アブレーションにより、逆 U 字型が明らかになります。 $\varepsilon=0.5$ は平均 73.84 ステップと 42\% の成功率を達成するのに対し、IC3Net の 75.31 ステップと 31\% という差があります。 1.47 ステップと 11 パーセンテージ ポイントで、シードの分散はより狭くなります。 MPE では、ゲーティングが非アクティブな場合でも、信念ヘッドは平均報酬を 12 ポイント改善し、分散を 26$\times$ 削減します。これは、信念が収束できる場合の原則的なゲーティングと、関係なく調整に利益をもたらす潜在表現の改善という 2 つの直交する貢献を示唆しています。
原文 (English)
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL
Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that produces unstable and uninterpretable gating behavior. I propose a principled alternative: agents communicate only when the KL divergence between their learned belief distributions exceeds a fixed threshold. Each agent maintains a belief distribution over a latent world state computed as a softmax over its LSTM hidden state, and communicates only when belief disagreement is large enough to justify information exchange. I evaluate this approach on the Predator-Prey benchmark from IC3Net \cite{singh2019} across two environment sizes with 5 seeds each, and on MPE simple\_spread \cite{lowe2017}, comparing against IC3Net, CommNet, and an independent controller. On PP 10$\times$10, IC3Net outperforms KL-belief at all thresholds. On the harder PP 20$\times$20, a threshold ablation over $\varepsilon \in \{0.1, 0.3, 0.5, 1.0\}$ reveals an inverted U-shape: $\varepsilon=0.5$ achieves 73.84 average steps and 42\% success rate versus IC3Net's 75.31 steps and 31\%, a gap of 1.47 steps and 11 percentage points with tighter seed variance. On MPE, the belief head improves mean reward by 12 points and reduces variance by 26$\times$ even when gating is inactive, suggesting two orthogonal contributions: principled gating when beliefs can converge, and improved latent representations that benefit coordination regardless.
高リスクのユースケースにおける FAIR と倫理に関する世界的な AI 規制: 比較レビュー
AI ガバナンスは自主的な倫理から、強制力のあるリスクベースの規制へと移行しつつありますが、管轄区域間の相違により、一か八かの AI 運用者にとってコンプライアンスの不確実性が生じています。我々は、(i) リスク分類のトリガー、(ii) 拘束力のある義務、(iii) 執行と説明責任のメカニズム、(iv) FAIR 原則が実際にどの程度運用されているかをマッピングする EU、米国、中国の比較マトリックスを提示します。私たちは、脳波検査 (EEG) によるリハビリテーション ロボット工学、将来の中央銀行デジタル通貨 (CBDC) エコシステムにおける AI を活用した債権回収、新興 AI ファクトリー インフラストラクチャにおける希少なグラフィックス プロセッシング ユニット (GPU) リソースの AI 主導の割り当てという 3 つの大きな影響を与える領域でマトリックスをストレス テストしました。私たちは、主要な法的文書と実装証拠を使用して、相互運用性の義務の弱さ、体制を越えた義務(AI + セクター規制 + データ保護)の運用の困難さ、重要なデジタル インフラストラクチャのユースケースに対するガバナンスの仕様が不十分であるという 3 つの繰り返し発生するギャップを特定します。実装のギャップを埋めるために、リソース記述フレームワーク/Web オントロジー言語 (RDF/OWL)、形状制約言語 (SHACL)、および来歴オントロジー (PROV-O) に基づく機械チェック可能なコンプライアンス アーティファクト パターンであるナレッジ ブロックの概要を説明し、複数のレジームにわたる監査対応のコンプライアンス バイ デザインを可能にします。
原文 (English)
Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review
AI governance is shifting from voluntary ethics to enforceable, risk-based regulation, yet cross-jurisdictional divergence creates compliance uncertainty for operators of high-stakes AI. We present a comparative matrix for the EU, US, and China that maps (i) risk classification triggers, (ii) binding obligations, (iii) enforcement and accountability mechanisms, and (iv) the degree to which FAIR principles are operationalised in practice. We stress-test the matrix on three high-impact domains: Electroencephalography (EEG)-guided rehabilitation robotics, AI-enabled debt collection in prospective Central Bank Digital Currency (CBDC) ecosystems, and AI-driven allocation of scarce Graphics Processing Unit (GPU) resources in emerging AI Factory infrastructures. Using primary legal texts and implementation evidence, we identify three recurring gaps: weak interoperability mandates, difficult operationalisation of cross-regime obligations (AI + sector regulation + data protection), and under-specified governance for critical digital infrastructure use cases. To bridge the implementation gap, we outline Knowledge Blocks, a machine-checkable compliance artefact pattern based on Resource Description Framework/Web Ontology Language (RDF/OWL), Shapes Constraint Language (SHACL), and Provenance Ontology (PROV-O), enabling audit-ready compliance-by-design across multiple regimes.
立場: AI ロックインが進行中、私たちは備えが必要
AI の安全性研究は主に 2 つの分野に焦点を当ててきました。技術的な調整 (AI システムが人間と調整した出力を確実に生成すること) と生成型 AI の社会的影響の規制 (失業リスクや労働市場の混乱など) です。しかし、同様に重要な側面、つまり AI システム自体への依存に内在するリスクもまだ十分に調査されていません。このポジションペーパーでは、AI の安全性研究は、AI ロックインに対処する必要があると主張します。AI ロックインとは、AI システムへの過度の依存が人間のスキル低下につながり、人間の自立した機能能力が低下し、AI システムが利用できなくなったり侵害されたりしたときにシステム的な脆弱性を生み出す現象です。私たちは、AI ロックインが個人、社会、国家レベルですでに現れている体系的な脅威であり、AI サービスの中断や地政学的な紛争によって劇的に増幅される可能性があることを強調します。詳細なシナリオに基づいて、AI ロックインがどのように発生し、個人のスキルの萎縮から国家規模のインフラ障害に至るまで、複数のレベルにまで拡大するかを調査します。これに対処するために、私たちはそのようなリスクをどのように軽減し、各レベルで備えることができるかについてのガイダンスを提供します。私たちは、そのような依存関係が定着する前に、あるいは取り返しがつかなくなる前に、AI ロックインに積極的に対処することが、個人の自主性と国家安全保障を守るために不可欠であると主張します。
原文 (English)
Position: AI Lock-In Is in Progress, and We Must Be Prepared
AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of generative AI's societal impacts (including unemployment risk and labor market disruption). However, an equally important dimension remains underexplored: the risk inherent in dependence on AI systems themselves. In this position paper, we argue that AI safety research should address AI Lock-In, the phenomenon whereby excessive reliance on AI systems leads to human deskilling, diminishes human capacity for independent functioning, and creates systemic vulnerabilities when AI systems become unavailable or compromised. We highlight that AI Lock-In is a systemic threat that is already emerging at individual, societal, and national levels, one that could be dramatically amplified by AI service disruptions or geopolitical conflicts. Drawing on detailed scenarios, we investigate how AI Lock-In emerges and escalates across multiple levels, ranging from individual skill atrophy to national-scale infrastructure failures. To address this, we provide guidance on how such risks can be mitigated and prepared for at each level. We contend that proactively addressing AI Lock-In before such dependencies become entrenched, or even irreversible, is essential for preserving individual autonomy and national security.
立場: AI 道徳推論の評価はまだ全体像の半分を満たしていない
大規模言語モデル (LLM) の道徳的能力を評価する最近の研究は、主に道徳的価値問題と呼ばれるもの、つまりモデルの出力が人間の道徳的価値観と一致するかどうかに焦点を当てています。対照的に、道徳規範の問題、つまりモデルが状況に応じた道徳規範を特定し、正しく適用できるかどうかは、依然として十分に研究されていません。私たちは、この不均衡は、規範的適用よりも価値表現を重視する道徳基礎理論やコールバーグの道徳的発達段階など、この分野が記述的な倫理枠組みに依存していることに起因すると仮定します。我々は既存のベンチマークと評価手法をレビューし、それらが価値の問題に大きく集中している一方で、規範的倫理に関する議論は依然として過小評価されていることを示します。我々は、次の 3 つの重大なギャップを特定します。(i) 道徳規範とその適用に関する高品質のグラウンドトゥルース データの欠如、(ii) 中間推論プロセスの評価が不十分、および (iii) 文脈内で道徳的に関連する特徴の特定に対する注意が限定的であること。続いて、規範理論の標準化された形式表現の開発、規範の適用を捉える専門家による注釈付きデータセットの構築、価値観レベルの能力と規範レベルの能力を明確に区別する評価プロトコルを含む研究課題を提案します。私たちの目標は、LLM における規範的推論のより体系的な研究を奨励することです。
原文 (English)
Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture
Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored. We posit that this imbalance stems from the field's reliance on descriptive ethics frameworks, such as Moral Foundations Theory and Kohlberg's stages of moral development, which emphasize value representation over normative application. We review existing benchmarks and evaluation methods, and show that they cluster heavily around the value problem, while discussion regarding normative ethics remains underrepresented. We identify three crucial gaps: (i) the absence of high-quality ground-truth data for moral norms and their applications, (ii) insufficient evaluation of intermediate reasoning processes, and (iii) limited attention to the identification of morally relevant features in context. Subsequently, we propose a research agenda that includes the development of standardized formal representations for normative theories, the construction of expert-annotated datasets capturing norm application, and evaluation protocols that explicitly distinguish between values-level and norms-level competence. Our goal is to encourage a more systematic study of normative reasoning in LLMs.
From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change
This paper presents a targeted narrative review establishing the historical and theoretical foundations for computational belief change imp…
立場: AI ガバナンスには法律だけでなく、ISO のような相互運用性プロトコルが必要です
人工知能 (AI) システムが重要なグローバル インフラストラクチャに深く統合されるにつれ、堅牢なガバナンス フレームワークの緊急性が高まっています。しかし、EU AI 法、中国のアルゴリズム ガバナンス、米国の NIST AI リスク管理フレームワークなど、管轄区域固有の法律、政策、自主的な枠組みによって主導される現在のアプローチは、規制の状況が細分化されています。この意見書では、\textbf{\textit{AI ガバナンスは法律だけではなく、国境を越えた標準化された機械可読なリスクコミュニケーションを可能にする ISO のような相互運用性プロトコルに基づいて構築される必要がある}}と主張します。 ISO 27001 やプライバシー バイ デザインなどの標準を通じて運用された GDPR の成功を踏まえ、私たちは、法域を超えたコンプライアンスを促進するために、バイアス、エネルギー使用量、データの出所に関する統一指標を含む標準化された AI \textit{栄養ラベル} の開発を提案します。これらのマニフェストは、中小企業 (SME) の障壁を低くし、余分な規制の取り組みを減らし、社会の信頼を構築します。この論文は、技術変化と並行して進化するように設計されたモジュール式のバージョン管理されたプロトコルを提唱することで、標準がイノベーションを抑制する可能性があるという懸念に対処しています。全体として、私たちはサイロ化された法的遵守から相互運用可能な技術的適合への移行を求め、責任ある AI 導入のための共通のグローバル言語を可能にします。
原文 (English)
Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws
As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance frameworks has intensified. However, current approaches, led by jurisdiction-specific laws, policies, and voluntary frameworks such as the EU AI Act, China's algorithm governance, and the NIST AI Risk Management Framework in the U.S., create a fragmented regulatory landscape. In this position paper, we argue that \textbf{\textit{AI governance must be built not on laws alone, but on ISO-like interoperability protocols that enable standardized, machine-readable risk communication across borders}}. Drawing on the success of the GDPR, which was operationalized through standards like ISO 27001 and Privacy by Design, we propose the development of standardized AI \textit{nutrition labels} containing unified metrics for bias, energy usage, and data provenance to facilitate cross-jurisdictional compliance. These manifests would lower barriers for small and medium enterprises (SMEs), reduce redundant regulatory efforts, and build public trust. The paper addresses concerns that standards may stifle innovation by advocating for modular, versioned protocols designed to evolve in tandem with technological change. Overall, we call for a shift from siloed legal compliance toward interoperable technical conformance, enabling a shared global language for responsible AI deployment.
立場: 神経制約推論の正確性が証明されるには記号統合が必要
制約充足問題のニューラル ソルバーは、顕著な分布内精度を達成しましたが、モデルが高い信頼性を報告する場合でも、分布シフトの下では永続的な制約違反が発生するという根本的な制限に悩まされています。この意見書では、ハード制約が存在し、検証のコストが比較的低い場合、ニューラル制約推論は純粋な学習よりも記号統合を優先する必要があると主張しています。我々が代表的な NP 完全テストベッドとして Sudoku に注目することは正当化されます。それは、Sudoku が簡単な検証と困難な解決の間に鋭い非対称性を示しているためです。候補解のチェックには多項式時間 $O(n^{2})$ のみが必要ですが、解を見つけるには指数関数的な探索が必要になる場合があります。決定論的アルゴリズム、メタヒューリスティック最適化、学習ベースのアプローチ、および言語条件付き推論にわたる解決方法の包括的な調査を通じて、インスタンスレベルの認証のないニューラルのみの方法では、記号的およびニューロ記号的アプローチが提供する証明可能な正確さを達成できないことを実証します。私たちは、ニューラル手法がヒューリスティックを学習して知覚をシンボルに変換することでシンボリック ソルバーを強化し、シンボリック手法がニューラル出力を検証して信頼性を確保する、双方向統合を提唱します。この立場を運用するために、この統合がどのように計算効率と証明可能な正しさの両方を達成できるかを実証する、マルチエージェント認定推論フレームワークを提案します。
原文 (English)
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration
Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation persistent constraint violations occur under distribution shifts even when the model reports high confidence. This position paper argues that when hard constraints exist and the cost of verification is relatively low, neural constraint reasoning must prioritize symbolic integration over pure learning. We justify our focus on Sudoku as a representative NP-complete testbed because it exhibits a sharp asymmetry between easy verification and hard solving: checking a candidate solution requires only polynomial time $O(n^{2})$, while finding a solution may require exponential search. Through a comprehensive survey of solving methods spanning deterministic algorithms, metaheuristic optimization, learning-based approaches, and language-conditioned reasoning, we demonstrate that neural-only methods without instance-level certification fail to achieve the provable correctness that symbolic and neuro-symbolic approaches provide. We advocate for a bidirectional integration in which neural methods enhance symbolic solvers by learning heuristics and converting percepts into symbols, while symbolic methods verify neural outputs to ensure their reliability. To operationalize this position, we propose a multi-agent certified reasoning framework that demonstrates how this integration can achieve both computational efficiency and provable correctness.
立場: より優れた ML レビューが必要ですか?親切に質問するのをやめて、クレジット システムでインセンティブを与え始めましょう
投稿数の急増、厳格な相互審査ポリシー、OpenReview などのプラットフォームの広範な採用により、出版料金による相殺圧力がないため、機械学習 (ML) コミュニティは、すべての科学分野の中で最大の学術的存在の 1 つとなっています。それにも関わらず、\textbf{ほぼ \textit{全員} はレビュー体験について \textit{多く} 不快なことを共有しています。} さらに悪いことに、議論はおろか、レビュー システムの効果や改善方法について真剣に議論するための公共の場がほとんどありません。\quad この意見書では、\textit{投稿量を合理的に制限するにはどうすればよいですか?} と \textit{良いレビューを奨励し、悪いレビューを阻止するにはどうすればよいですか?という 2 つの中心的な問題から議論を展開します。このような問題に対処するための既存の試みの長所と短所を評価します。具体的には、いくつかの一般的な会議メカニズムに対する 4 つの見解を示し、改善のための 2 つの代替設計を提案します。\quad 私たちの一般的な立場は、ML 査読における有意義な改善は、論文募集や査読者ガイドラインに組み込まれた丁寧なベストプラクティスの提案からはもたらされないというものです。それには、\textbf{通貨のような信用システム (例: 私たちが提案する \textit{OpenReview Points})} と組み合わせた \textbf{強制可能でありながらきめの細かい手続き上の保護手段} が必要です。 ML 実践者は、優れたレビュー実践に貢献することでそのようなポイントを「獲得」し、そのポイントを 1 つまたは複数の主要なカンファレンスで「消費」して、無料の登録や追加のレビュー リソースをリクエストする権利など、さまざまな種類の「特典」と引き換えることができます。
原文 (English)
Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System
With soaring submission counts, stricter reciprocal review policies, widespread adoption of platforms like OpenReview, and without the offsetting pressure of publication fees, the machine learning (ML) community has one of the largest scholarly presences among all scientific fields. And yet, \textbf{almost \textit{everyone} has \textit{many} unpleasant things to share about their review experience.} Worse, there is little public space to seriously discuss, let alone debate, what makes a review system effective or how it might be improved.\quad In this position paper, we expand our discussion from two core problems: \textit{How can we reasonably limit submission volume?} and \textit{How can we incentivize good and discourage bad reviewing?} We first assess the strengths and shortcomings of existing attempts to address such problems. Specifically, we present four takes on some popular conference mechanisms and propose two alternative designs for improvement.\quad Our general position is that meaningful improvement in ML peer review won't come from polite best-practice suggestions tucked into Calls for Papers or Reviewer Guidelines: it requires \textbf{enforceable yet fine-grained procedural safeguards} paired with \textbf{a currency-like credit system (e.g., our proposed \textit{OpenReview Points})}. ML practitioners can ``earn'' such points by contributing good review practices, and ``spend'' them across one or multiple major conferences to redeem different kinds of ``perks,'' such as complimentary registration or the right to request additional review resources.
Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study
Early identification of adolescent substance-use risk is an important prevention challenge, yet the relative value of baseline characterist…
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization
Logic synthesis optimization poses significant challenges due to exponentially growing search spaces, sparse reward signals, and diverse lo…
OGX: オープンソース、ベンダー中立の生成 AI アプリケーション サーバー
OGX (Open GenAI Stack) は、プラグ可能なバックエンド プロバイダーを備えた主要なフロンティア ラボ (OpenAI、Anthropic、Google) の API を実装するオープンソース AI アプリケーション サーバーおよび Python ライブラリです。取得拡張生成パイプライン、マルチターン エージェント、ツール呼び出しワークフローなどのエージェント AI アプリケーションを構築する開発者は、アプリケーション コードを変更することなく、単一の API サーフェスに対して開発し、推論エンジン、ベクトル データベース、安全性バックエンドを任意に組み合わせてデプロイできます。 OGX の主な焦点は、Open Responses 仕様に準拠した、サーバー側のエージェント オーケストレーションのための Responses API です。このサーバーは、Anthropic Messages API と Google GenAI Interactions API もサポートしており、SDK の選択をモデルやデプロイメントの決定から切り離します。 20 を超える推論プロバイダー、13 のベクター ストア バックエンド、および実稼働展開用のコンパニオン Kubernetes Operator を備えた OGX は、Claude Code、Codex CLI、OpenCode、OpenHands などの AI を活用した開発者ツールの自己ホスト型でモデルに依存しないバックエンドとして機能します。このプロジェクトには 8,400 人を超える GitHub スター、242 人の貢献者、そして約 2 年間にわたる公開開発を通じて 4,000 件のコミットが含まれています。
原文 (English)
OGX: An Open-Source, Vendor-Neutral Generative AI Application Server
OGX (Open GenAI Stack) is an open-source AI application server and Python library that implements the APIs of major frontier labs (OpenAI, Anthropic, Google) with pluggable backend providers. Developers building agentic AI applications--such as retrieval-augmented generation pipelines, multi-turn agents, and tool-calling workflows--can develop against a single API surface and deploy with any combination of inference engine, vector database, and safety backend, without changing application code. OGX's primary focus is the Responses API for server-side agentic orchestration, conforming to the Open Responses specification. The server also supports the Anthropic Messages API and Google GenAI Interactions API, decoupling SDK choice from model and deployment decisions. With over 20 inference providers, 13 vector store backends, and a companion Kubernetes Operator for production deployment, OGX serves as the self-hosted, model-agnostic backend for AI-powered developer tools including Claude Code, Codex CLI, OpenCode, and OpenHands. The project has over 8,400 GitHub stars, 242 contributors, and 4,000 commits across nearly two years of public development.
Euclid-Omni : 平面幾何学の統合神経記号フレームワーク
ユークリッド幾何学は、直感的な図の理解、公理的演繹、代数計算の組み合わせを必要とするため、AI 推論にとって魅力的なテストベッドです。しかし、既存のアプローチは通常、これらの能力のサブセットのみに対応するか、競技レベルの問題に対処するのに苦労します。 \textit{Euclid-Omni} は、形式幾何学システムと大規模言語モデル (LLM) および視覚言語モデル (VLM) を組み合わせた統合神経記号フレームワークで、オリンピック レベルの難易度まで形式言語と自然言語で計算形式と証明形式の問題の両方に取り組むことができます。その中核として、演繹的推論と代数計算を通じて推論ステップを自動的に生成する多用途の記号幾何学ソルバーである \textit{Euclidea} を開発しています。これに基づいて、記号的な問題と解決策を合成し、図をレンダリングし、それらを自然言語に翻訳するデータ生成パイプラインを開発し、幅広い推論設定で LLM と VLM をトレーニングするための大規模で多様なデータセットを生成します。実験の結果、合成データでトレーニングされた VLM は計算タスクで優れたパフォーマンスを達成し、\textit{Euclidea} と組み合わせた LLM は、使用する計算データとトレーニング データが桁違いに少ないにもかかわらず、オリンピック レベルの証明問題で最先端のシステムと競合できることがわかりました。コードとスクリプトは https://github.com/20171130/Euclid-Omni で公開されています。
原文 (English)
Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry
Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic deduction, and algebraic computation. Yet, existing approaches typically address only a subset of these abilities or struggle with competition-level problems. We introduce \textit{Euclid-Omni}, a unified neuro-symbolic framework that couples a formal geometry system with Large Language Models (LLMs) and Vision-Language Models (VLMs) to tackle both calculation- and proving-style problems, in formal and natural languages, up to Olympiad-level difficulty. At its core, we develop \textit{Euclidea}, a versatile symbolic geometry solver that automatically generates reasoning steps through deductive inference and algebraic computation. Building on this, we develop a data-generation pipeline that synthesizes symbolic problems and solutions, renders diagrams, and translates them into natural language, producing large-scale, diverse datasets for training LLMs and VLMs across a wide range of reasoning settings. Experiments show that VLMs trained on our synthetic data achieve superior performance on calculation tasks, and that LLMs combined with \textit{Euclidea} are competitive with state-of-the-art systems on Olympiad-level proving problems, despite using orders of magnitude less compute and training data. Code and scripts are publicly available at https://github.com/20171130/Euclid-Omni
ルールと LLM を使用した説明的なドキュメント レイアウトの埋め込みと注釈付けのためのエージェント フレームワーク: 植物科学の使用例
背景: 情報検索 (IR) の最近の進歩では、密表現と疎表現の両方、大規模言語モデル (LLM)、および特殊な検索モデルを活用して、ランキングの精度、関連性、言語間のパフォーマンスを向上させています。パッセージのインデックス作成、文書レイアウト分析、意味論的な知識表現などの補完的な技術により、きめ細かいコンテキスト情報や構造情報を取得することで、検索効率がさらに向上します。新しいエージェント LLM フレームワークは、計画、反復推論、ツールの使用、およびマルチエージェントのコラボレーションを可能にすることでこれらの機能を拡張し、それによってアプリケーションをさまざまなドメインに広げます。これらのフレームワークは、厳格な評価、倫理的配慮、信頼性も重視しており、現実世界の環境での責任ある展開を保証します。私たちは、植物形質抽出のためのモジュール式のエージェントベースのパイプラインを提案します。光学式文字認識 (OCR) は PDF を機械可読テキストに変換し、セグメンテーションとインデックス作成によりコンテンツを属と種ごとに整理します。ルールベースのパーサーは構造化された植物の形質を抽出し、大規模言語モデル (LLM) のアンサンブルは形質の語彙を拡張し、曖昧さを解決します。このアプローチにより、正確な種の認識、スケーラブルな注釈、植物のテキスト記述の説明可能な統合が保証され、大規模な植物コーパス全体にわたって堅牢で解釈可能なデータ抽出が可能になります。結果: 3 つの地域植物データセットを使用して、私たちのシステムは 4,961 種にわたって 55,737 個の形質アノテーションを抽出し、種ごとに平均 9.1 個の形質を抽出しました。 LLM ベースのエンリッチメントの統合により、75% の特性のカバー率が向上し、総アノテーションが 59% 増加しました。 OCR エンジンの選択は種認識にわずかな影響を与えましたが、全体的なアノテーション数は安定したままであり、大規模な植物形質抽出のためのパイプラインの堅牢性、拡張性、および信頼性を示しています。
原文 (English)
An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case
Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Complementary techniques such as passage indexing, document layout analysis, and semantic knowledge representation further enhance retrieval effectiveness by capturing fine-grained contextual and structural information. Emerging agentic LLM frameworks extend these capabilities by enabling planning, iterative reasoning, tool use, and multi-agent collaboration, thereby broadening applications across diverse domains. These frameworks also emphasize rigorous evaluation, ethical considerations, and trustworthiness, ensuring responsible deployment in real-world settings. We propose a modular, agent-based pipeline for botanical trait extraction. Optical character recognition (OCR) converts PDFs into machine-readable text, while segmentation and indexing organize content by genus and species. Rule-based parsers extract structured botanical traits, and ensembles of large language models (LLMs) expand trait vocabularies and resolve ambiguities. This approach ensures accurate species recognition, scalable annotation, and explainable integration of textual botanical descriptions, enabling robust and interpretable data extraction across large botanical corpora. Results: Using three regional botanical datasets, our system extracted 55,737 trait annotations across 4,961 species, averaging 9.1 traits per species. Integration of LLM-based enrichment improved coverage for 75% of traits, increasing total annotations by 59%. While the choice of OCR engine had a minor effect on species recognition, overall annotation counts remained stable, demonstrating the robustness, scalability, and reliability of the pipeline for large-scale botanical trait extraction.
幻覚の雪だるま: マルチエージェント LLM パイプラインの状態遷移としてのエラー伝播のモデル化
シーケンシャル マルチエージェント LLM パイプラインは、引き継ぎ時の検証を行わずに特殊なエージェントを連鎖させ、測定可能な重大な結果をもたらす構造的欠陥を生み出します。ステージ 1 で注入された幻覚は単に持続するだけではないことを示します。それらは変換されます。生の数値的事実が派生計算になり、次に物語的な散文になり、次に編集的に承認された結論になります。変換のたびに、検出可能性はほぼ不可逆的に低下します。我々はこれを幻覚雪だるま効果として定式化します。これは、経験的に測定された境界ごとの脱出確率が 24.6%、48.3%、および 89.3% である 4 つの状態 (Raw Fact $\to$ Derived $\to$ Narrative $\to$ Invisible) にわたる一次マルコフ過程です。 FinanceBench の 4 エージェントの財務分析パイプラインで自動的に注入された 346 件の幻覚全体で、gpt-4o の検出率はステージ 1 の 72.0% からステージ 4 の 50.9% に低下し、幻覚の 23.7% が最終出力では完全に検出されずに残ります。テストされた最も強力なモデル (Qwen3.5-397B-A17B、ステージ 1 で 87.0%) でさえ、構造上の天井に直面しています。ステージ 4 の検出率は ${\sim}$60 ~ 65% にすぎないと予測されています。重要なことに、同一の RAG 検証ツールを使用した境界ゲートでは、パイプラインの終了チェック (Cohen の $h = -0.911$、$p < 0.000001$) と比較して幻覚生存率が 58.4% から 16.2% に減少しますが、終了チェックだけでは検証なしの場合に比べてわずか 2.3 pp の改善しか達成されません。検証するかどうかよりも、いつ検証するかが重要です。私たちのモデルは、$n$ エージェントの線形パイプラインの生存を予測し、最適な検証リソースの割り当てを規定します。まず、幻覚の 75.4% がまだ捕らえられる $S_1{\to}S_2$ に投資し、89.3% がすでに逃れている $S_3{\to}S_4$ には投資しません。
原文 (English)
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines
Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences. We show that hallucinations injected at Stage 1 do not merely persist; they transform: raw numerical facts become derived computations, then narrative prose, then editorially approved conclusions. At each transformation, detectability degrades near-irreversibly. We formalize this as the hallucination snowball effect, a first-order Markov process over four states (Raw Fact $\to$ Derived $\to$ Narrative $\to$ Invisible) with empirically measured per-boundary escape probabilities of 24.6%, 48.3%, and 89.3%. Across 346 automatically injected hallucinations in a 4-agent financial analysis pipeline on FinanceBench, gpt-4o detection drops from 72.0% at Stage 1 to 50.9% at Stage 4, and 23.7% of hallucinations survive completely undetected in the final output. Even the strongest model tested (Qwen3.5-397B-A17B, 87.0% at Stage 1) faces a structural ceiling; projected Stage 4 detection is only ${\sim}$60--65%. Critically, boundary gates using identical RAG verification tools reduce hallucination survival from 58.4% to 16.2% versus end-of-pipeline checking (Cohen's $h = -0.911$, $p < 0.000001$), while end-checking alone achieves merely 2.3 pp improvement over no verification. When you verify matters more than whether you verify. Our model predicts survival for $n$-agent linear pipelines and prescribes optimal verification resource allocation: invest at $S_1{\to}S_2$ first, where 75.4% of hallucinations are still catchable, not at $S_3{\to}S_4$ where 89.3% have already escaped.
安全な LLM エージェントを目指して: 仕様、検証、施行に関する調査
LLM エージェントは、データベースの更新、API 呼び出し、ファイル操作、ツールの自律的な使用など、元に戻せない現実世界のアクションを実行することが増えています。しかし、エージェントが生成する計画に対して、正式に根拠のあるタスクレベルの安全性を保証する既存のシステムはありません。研究は仕様、検証、施行にわたって断片化されたままであり、既存のアプローチの長所と限界についての理解が限られています。このギャップに対処するために、私たちは 2022 年から 2026 年の間に発表され、6 つの学術データベースから取得された 38 件の研究について PRISMA 2020 系統的レビューを実施しました。私たちの分析により、4 つの重要な発見が明らかになりました。まず、仕様のボトルネックが依然として主要な課題です。自然言語から形式的な翻訳は、意味論的な正確性が 24% ~ 35% しか達成できず、下流の検証が損なわれます。第 2 に、実行時監視は最も成熟した施行戦略であり、制御された設定で危険なアクションを 40% から 65% 削減しますが、完全な安全保証は提供されません。第三に、検証者税は、エージェントが代替の安全でないパスを悪用するため、安全でないアクションの 94% をブロックしても、安全なタスクの完了率が 5% 未満になる可能性があることを示しています。最後に、健全性、スケーラビリティ、セマンティックな正確性、およびタスクレベルの安全性の保持を同時に達成する既存のアプローチはありません。私たちは、3 レベルの分類法、既存の技術の比較分析、検証者税に関する証拠の統合、および信頼できるエージェント AI に関する 10 の問題の研究課題に貢献します。
原文 (English)
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement
LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarantees for the plans these agents generate. Research remains fragmented across specification, verification, and enforcement, limiting understanding of the strengths and limitations of existing approaches. To address this gap, we conducted a PRISMA 2020 systematic review of 38 studies published between 2022 and 2026 and retrieved from six academic databases. Our analysis reveals four key findings. First, the specification bottleneck remains the primary challenge: natural-language-to-formal translation achieves only 24% to 35% semantic correctness, undermining downstream verification. Second, runtime monitoring is the most mature enforcement strategy, reducing unsafe actions by 40% to 65% in controlled settings, but it does not provide complete safety guarantees. Third, the verifier tax shows that blocking 94% of unsafe actions can still result in less than 5% safe task completion because agents exploit alternative unsafe paths. Finally, no existing approach simultaneously achieves soundness, scalability, semantic correctness, and task-level safety preservation. We contribute a three-level taxonomy, a comparative analysis of existing techniques, a synthesis of evidence on the verifier tax, and a ten-problem research agenda for trustworthy agentic AI.
立場: 医療 AI は実際の治療結果を無視する
医療 AI は、治療決定につながる診断および予後タスクを実行する能力を急速に向上させています。しかし、治療そのものの理解は依然として不十分であり、治療結果に関する実際の基礎データではなく、人間の意見や総合(特に生物医学出版物や診療ガイドラインなどのテキスト)を使用しています。この無視は医療 AI の可能性を著しく制限しており、この意見書で主張されているように、すでにフロンティア モデルと主要なベンチマークの両方に欠陥を引き起こしています。観察データベースやランダム化実験などの情報源から得られた実際の治療結果は、トレーニングと評価の両方に実質的に組み込まれるべきです。これらの成果を向上させることは、すべての医療 AI の下流目標として改めて強調されるべきです。
原文 (English)
Position: Medical AI Neglects Real Treatment Outcomes
Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding of treatment itself is still inadequately trained and evaluated, using human opinions and syntheses (especially texts such as biomedical publications and clinical practice guidelines) rather than actual underlying data on treatment outcomes. This neglect seriously limits the potential of medical AI, and is already causing deficiencies in both frontier models and major benchmarks, as argued in this position paper. Real treatment outcomes, drawn from sources such as observational databases and randomized experiments, should be substantially incorporated into both training and evaluation. Improving these outcomes should be reemphasized as the downstream goal of all medical AI.
LLM が間違った法律を適用するのはどのような場合ですか?時間的法的推論における LLM の失敗の診断
法的判決予測 (LJP) などの法的推論タスクでは、事件を統治する法律の時間的に正しいバージョンを特定する必要があります。これを、一時的な適用法判断と呼んでいます。ただし、大規模言語モデル (LLM) がこのタスクを確実に実行できるかどうかはまだ解明されていません。本論文では、時間的適用法判断に関してLLMを評価するためのベンチマークを構築し、時間的法的推論においてLLMが失敗する理由を系統的に調査する。私たちの実験により、4 つの重要な発見が明らかになりました。まず、LLM は、法的に関連する事実がいつ発生したかに関係なく、最も最近制定された法律を適用することに強い偏りを示します。第二に、この偏見は、法律には一時的な範囲があることを理解できないことや、歴史的な法令についての知識の欠如から生じるものではありません。第三に、強化学習型の明示的推論が重要なメカニズムである可能性があるという行動証拠を提供します。これにより、一般的な推論能力が向上する一方で、推論経路の多様性が減少し、モデルが現行法則の適用に収束します。第 4 に、これは直観に反する逆関係を生み出します。つまり、より強力な一般推論能力を持つモデルは、時間的な法的推論のパフォーマンスが低下する傾向があります。私たちの調査結果は、時間的に根拠のある法的推論におけるLLMのパフォーマンスを改善するための将来の取り組みに具体的な指針を提供します。
原文 (English)
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning
Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination. However, whether large language models (LLMs) can reliably perform this task remains unexplored. In this paper, we construct a benchmark to evaluate LLMs on temporal applicable-law determination, and systematically investigate why they fail at temporal legal reasoning. Our experiments reveal four key findings. First, LLMs exhibit a strong bias toward applying the most recently enacted law, regardless of when the legally relevant facts occurred. Second, this bias does not stem from an inability to understand that laws have temporal scope, nor from a lack of knowledge about historical statutes. Third, we provide behavioral evidence that reinforcement-learning-shaped explicit reasoning may be a key mechanism: while improving general reasoning ability, it reduces the diversity of reasoning paths, causing models to converge on applying the current law. Fourth, this produces a counterintuitive inverse relationship: models with stronger general reasoning ability tend to perform worse on temporal legal reasoning. Our findings offer concrete guidance for future work on improving LLM performance in temporally grounded legal reasoning.
LLM エージェントは合理的に交渉しますか? A2A/MCP を介した検証可能なマルチエージェント インタラクションのためのメカニズム設計フレームワーク
最新の LLM エージェント フレームワークは、エージェントからツールへのアクセスのための Anthropic の Model Context Protocol (MCP) や、エージェントの委任とネゴシエーションのための Google の Agent2Agent (A2A) プロトコルなどの標準を通じて相互運用性が高まっています。ただし、これらのプロトコルは、戦略的な正確さではなく、転送と発見を指定しており、効率的、個別に合理的、または戦略に裏付けられた結果を保証するものではありません。我々は、(i) A2A メッセージ スキーマに対する制約として、交互オファー交渉や Vickrey-Clarke-Gloves スタイルのオークションなどの古典的な交渉メカニズムをエンコードするフレームワークを導入します。 (ii) プロトコルの不変条件に対してメッセージをチェックする軽量の実行時検証および修復層を提供します。 (iii) ゲーム理論の予測からの逸脱を測定するための、既知の最適なソリューションを使用した交渉および割り当てタスクのベンチマークを提供します。非構造化ダイアログ、構造化プロトコル、および検証付き構造化プロトコルを使用して、複数の LLM バックボーンを評価します。交渉トライアル (条件ごとに N=30) 全体で、検証により結果の差異が減少し、構造化プロトコルは両方のモデルで 100% の成功を達成します。パーサー アーティファクトを修正した後、監査された非構造化ベースラインは約 97 パーセントと 93.3 パーセントの成功を達成しました。オークション実験 (モデルあたり N=30) では、どちらのモデルも 100% の効率的な割り当てを達成していますが、真実の入札においては大きく異なります。1 つはすべてのトライアルで正確な評価額を入札しますが、もう 1 つはトライアルの 3.3 パーセントのみで入札します。したがって、メカニズムレベルのインセンティブの互換性は、LLM エージェントの動作に自動的に移行しません。三者間の公平な配分タスクでは、使用可能な結果は 4.2% しか得られませんでした。この陰性結果を診断とともに報告します。この研究は、古典的なマルチエージェント システム理論と最新の LLM エージェント インフラストラクチャを橋渡しし、A2A プロトコル層で検証可能な相互作用を定義します。
原文 (English)
Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP
Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However, these protocols specify transport and discovery rather than strategic correctness and do not guarantee efficient, individually rational, or strategy-proof outcomes. We introduce a framework that (i) encodes classical negotiation mechanisms, including alternating-offers bargaining and Vickrey-Clarke-Groves-style auctions, as constraints over A2A message schemas; (ii) provides a lightweight runtime verification and repair layer that checks messages against protocol invariants; and (iii) offers a benchmark of negotiation and allocation tasks with known optimal solutions for measuring deviations from game-theoretic predictions. We evaluate multiple LLM backbones using unstructured dialogue, structured protocols, and structured protocols with verification. Across negotiation trials (N=30 per condition), verification reduces outcome variance, while structured protocols achieve 100 percent success for both models. After correcting parser artifacts, audited unstructured baselines achieve approximately 97 percent and 93.3 percent success. In auction experiments (N=30 per model), both models achieve 100 percent efficient allocation but differ sharply in truthful bidding: one bids its exact valuation in every trial, whereas the other does so in only 3.3 percent of trials. Thus, mechanism-level incentive compatibility does not automatically transfer to LLM-agent behavior. A three-party fair-allocation task produced only 4.2 percent usable outcomes; we report this negative result with a diagnosis. This work bridges classical multi-agent systems theory and modern LLM-agent infrastructure and defines verifiable interaction at the A2A protocol layer.
大規模な言語モデルとその力学および空間幾何学への認識
大規模言語モデル (LLM) は、確立されたコード生成および数学的推論のベンチマークで良好に機能しますが、力学および空間幾何学におけるその能力 (ここでは機械工学の認識と呼ばれます) は体系的に定量化されていません。パラメータ化されたテキスト記述からマルチボディ シミュレーション モデルを作成する際の LLM を評価する、完全に自動化されたベンチマークである MecEng を紹介します。このベンチマークは、ジョイントと接触を備えた剛体システムから、正確な 3D ジオメトリの生成、四面体有限要素メッシュ作成、機械部品のハーティ-クレイグ-バンプトン モデルの次数削減を必要とする柔軟なマルチボディ システムに至るまで、3 つの難易度に分かれた 84 の一般的なタスクで構成されています。 LLM を備えた専用パイプラインは、Netgen を使用してテキストからシミュレーション対応のジオメトリを生成し、コード Exudyn のマルチボディ システム モデルを構築します。これらのモデルは、グラフ ノードの注釈を含むシステム グラフの同型性、数値解、質量、幾何学、固有周波数などの部品固有の測定など、いくつかのレベルで専門家のグラウンド トゥルースと照合して検証されます。合計 32 のオープンウェイト LLM と 2 つの独自の LLM が評価されます。剛体タスクでは、最も優れたオープンウェイト モデルの全体的な成功率は 86.0% であり、最も強力な独自モデルの 91.4% と比較して、柔軟なマルチボディ タスクは依然としてかなり困難です。追加の研究により、サンプリング温度、推論、即時設計、モデル サイズ、LLM リリース日の影響が定量化されます。この結果は、現在の LLM に対する機械工学の認識が急速に向上しているものの、依然としてエラーが発生しやすいことを示しています。
原文 (English)
Large Language Models and their Awareness of Mechanics and Spatial Geometry
Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been quantified systematically. We present MecEng, a fully automated benchmark that evaluates LLMs on the creation of multibody simulation models from parameterized textual descriptions. The benchmark comprises 84 generic tasks on three difficulty levels, ranging from rigid-body systems with joints and contact to flexible multibody systems that require exact 3D geometry generation, tetrahedral finite-element meshing, and Hurty-Craig-Bampton model order reduction of machine parts. A dedicated pipeline with LLMs generates simulation-ready geometry from text using Netgen, and builds multibody system models for the code Exudyn, which are then verified against expert ground truth on several levels: system-graph isomorphism including graph node annotations, numerical solutions, and part-specific measures such as mass, geometry, and eigenfrequencies. In total, 32 open-weight and two proprietary LLMs are evaluated. On rigid-body tasks, the best open-weight model obtains an overall success rate of 86.0%, compared to 91.4% for the strongest proprietary model, while flexible multibody tasks remain considerably harder. Additional studies quantify the influence of sampling temperature, reasoning, prompt design, model size, and LLM-release date. The results indicate rapidly improving, but still error-prone, mechanical engineering awareness of current LLMs.
子育てアドバイスのための LLM のベンチマークに対する人間中心のアプローチ
子育てなどのアドバイスを求めるために大規模言語モデル (LLM) を使用する人が増えています。子育ては重要かつ社会的に敏感な領域です。したがって、LLM によって提供されるアドバイスを評価するには、集約された情報品質ベンチマークを超えて、応答の関係要素と行動要素を考慮する指標が必要です。この論文では、子育ての専門家によって作成された多次元ルーブリックを使用して、LLM を審査員として使用する方法を使用して、2 つの言語 (英語と中国語) で 100 の子育てシナリオにわたる 15 の LLM を評価します。結果は、集計スコアがルーブリック項目固有の弱点を隠す可能性があること、モデルが暗黙的に異なる子育てスタイルを奨励していること、言語が反応に影響を与えていることを示しています。私たちは、評価出力の監査可能性の重要性と、子育てなどの分野で LLM によって生成されたアドバイスを評価する際に伴う課題を強調します。私たちの調査結果は、ユーザーとの直接的な関わりや、ユーザー向けの子育てアドバイス アプリケーションの開発のための LLM の選択に関する重要な洞察を提供します。
原文 (English)
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses. With a multi-dimensional rubric created by parenting experts, this paper evaluates 15 LLMs across 100 parenting scenarios in 2 languages (English and Chinese), using an LLM-as-a-judge method. Results show that aggregate scores can hide rubric item-specific weaknesses, models implicitly encourage different parenting styles, and language influences responses. We highlight the importance of evaluation output auditability and challenges involved in evaluating LLM-generated advice in domains like parenting. Our findings provide important insights for selecting LLMs for direct user engagement and the development of user-facing parenting advice applications.
エージェントサービスにおける KV キャッシュ管理のためのエージェント実行の学習
マルチエージェント LLM システムは、AI サービスの重要な導入パラダイムとして登場しており、各ユーザーのリクエストは一連の特殊なエージェントに分解されます。これらのワークフロー全体で、すべてのエージェントはシステム プロンプト、ツール定義、および少数のショットの例で構成される固定コンテキストを繰り返し実行し、KV キャッシュを再利用する実質的な機会を生み出します。ただし、既存の LLM サービス システムは、プレフィックス キャッシュと最新性に基づく置換を使用して KV キャッシュを事後的に管理するため、再利用可能なエージェント コンテキストが次の呼び出しの前に削除され、繰り返しの再計算が強制されます。マルチエージェント LLM サービス用のエージェント対応 KV キャッシュ ランタイム レイヤーである CacheScout を紹介します。重要な洞察は、将来の KV キャッシュの再利用は、キャッシュの最新性だけではなく、エージェントの実行セマンティクスによって制御されるということです。 CacheScout は、事前定義されたワークフロー グラフやオフライン トレーニングを必要とせずに、エージェントの実行遷移をオンラインで学習することによってこれらのセマンティクスをキャプチャし、学習された実行モデルを使用して、サービスを提供するクリティカル パスを変更せずに、キャッシュのエビクションとプロアクティブなプリフェッチの両方をガイドします。 vLLM の上に CacheScout を実装します。代表的な現実世界のマルチエージェント ワークロード全体で、CacheScout は KV キャッシュ ヒット率を 10 ~ 18 パーセント ポイント改善し、平均 TTFT を 18 ~ 45% 削減し、ターンあたりの平均レイテンシを 29 ~ 38% 削減し、ピーク スループットを最大 57% 増加させます。これらの利点はより大きなモデルにも適用され、37% 高いスループットを維持しながら TTFT を最大 54% 削減します。
原文 (English)
Learning Agent Execution for KV-Cache Management in Agentic Serving
Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents. Across these workflows, every agent repeatedly executes a fixed context consisting of system prompts, tool definitions, and few-shot examples, creating substantial opportunities for KV-cache reuse. Existing LLM serving systems, however, manage KV-cache reactively using prefix caching and recency-based replacement, causing reusable agent contexts to be evicted before their next invocation and forcing repeated recomputation. We present CacheScout, an agent-aware KV-cache runtime layer for multi-agent LLM serving. The key insight is that future KV-cache reuse is governed by agent execution semantics rather than cache recency alone. CacheScout captures these semantics by learning agent execution transitions online, without requiring predefined workflow graphs or offline training, and uses the learned execution model to guide both cache eviction and proactive prefetching while leaving the serving critical path unchanged. We implement CacheScout on top of vLLM. Across representative real-world multi-agent workloads, CacheScout improves KV-cache hit rate by 10-18 percentage points, reduces mean TTFT by 18-45%, lowers mean per-turn latency by 29-38%, and increases peak throughput by up to 57%. These benefits also generalize to larger models, reducing TTFT by up to 54% while sustaining 37% higher throughput.
化粧品化学と皮膚の健康における大規模言語モデルの精度と信頼性: ベンチマーク研究
消費者がスキンケアに関するアドバイスを求めて AI チャットボットを利用することが増えている一方で、化粧品化学における大規模言語モデル (LLM) の技術的精度は依然として過小評価されています。私たちは、特定の化粧品成分の化学的特性や、消費者が関心を持つ可能性のある一般的な化粧品のシナリオなど、化粧品の化学に関連する一連の構造化されたトピックに関して 14 の LLM のベンチマークを実施しました。インターネット検索能力ではなく、各モデルの内部化された知識を評価するために、Web 検索は全体的に無効になりました。全体的なパフォーマンスは悪く、定量的推論と構造的同定のタスクにおいて最も顕著な欠陥が見られました。モデルたちは一般的なスキンケアの質問に合理的に答えましたが、回答は常に情報に基づいた消費者の意思決定に必要な技術的な深さを欠いていました。特に、AI との会話にはリスクが生じる可能性があります。権威あるように聞こえても技術的なエラーが含まれている出力は、不確実性を明確に認める応答に比べて、懐疑的な見方を生む可能性が低くなります。これらの発見は、主に未検証の公開データに基づいてトレーニングされた汎用 LLM が、現在、化粧品化学情報の信頼できる情報源ではないことを示唆しています。これらのツールが公共利用のためのリソースとして考慮される前に、検証済みの化学データセットと皮膚科学データセットの微調整、およびアルゴリズム推論の大幅な改善という 2 つの面での進歩が必要になる可能性があります。
原文 (English)
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study
As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic chemistry, including the chemical properties of specific cosmetic ingredients and common cosmetic scenarios that may be of interest to consumers. Web search was disabled throughout to assess each model's internalized knowledge rather than its internet retrieval capacity. Overall performance was poor, with the most pronounced deficits in quantitative reasoning and structural identification tasks. While models handled general skincare questions with reasonability, responses consistently lacked the technical depth required for informed consumer decision-making. Notably, conversation with AI can pose a risk: outputs that sound authoritative but contain technical errors are less likely to generate skepticism compared to responses that explicitly acknowledge uncertainty. These findings suggest that general-purpose LLMs, trained predominantly on unverified public data, are currently not reliable sources of cosmetic chemistry information. Progress on two fronts, fine-tuning verified chemical and dermatological datasets, and substantial improvements to algorithmic reasoning, will likely be needed before these tools can be considered as resources for public use.
タスク レベルおよびセッション レベルのモデル ルーティング: 4 つのベンチマークにわたる 4 つのオープンソース ルーターの共通インターフェイス ハイブリッド評価
エージェント システムでは、モデルの選択をルーターに委任するケースが増えていますが、オープンソース ルーターは通常、さまざまなタスク、候補プール、および実行プロトコルで評価されるため、直接の比較は制限されています。 RouterBench、BFCL v4、tau2-bench、WebArena にわたる 4 つのルーター実装の共通の測定プロトコルとハイブリッド評価を紹介します。 290 の凍結タスクを、2,610 の候補結果のロックされたマトリックスに対して評価します。 3 つのルーターは、一定またはほぼ一定の層割り当てを発行します。 vLLM セマンティック ルーターのみがプロンプト コンテンツによって大きく異なり、4 つのベンチマークのいずれにおいても最も高い成功率が観察されています。 Always-Mid は 3 つのベンチマークで Aurelio と正確に一致し、4 つ目のベンチマークでは 0.003 以内です。 vLLM の場合、タスク レベルの優位性テストでは、共有一致コンテンツ ブラインド割り当てを超えるタスク固有の利点は検出されません。同等性は、プロトコルで宣言された 5 パーセント ポイントのマージンで WebArena 上でのみ確立されます。結果は、これらの構成と制御の下で、観察されたゲインは、実証されたタスク固有のターゲティングよりも、選択された層の構成をより厳密に追跡することを示しています。したがって、固定層のベースラインと選択層の配布は、ルーターの評価において必要な制御となります。調査結果は、一般的なルーティング パラダイムではなく、これらの構成、候補プール、および凍結されたベンチマーク サンプルに限定されています。
原文 (English)
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks
Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison. We present a common measurement protocol and hybrid evaluation of four router implementations across RouterBench, BFCL v4, tau2-bench, and WebArena. We evaluate 290 frozen tasks against a locked matrix of 2,610 candidate outcomes. Three routers emit constant or near-constant tier assignments; only vLLM Semantic Router varies materially with prompt content, and it has the highest observed success rate on none of the four benchmarks. Always-Mid matches Aurelio exactly on three benchmarks and within 0.003 on the fourth. For vLLM, task-level superiority tests detect no task-specific advantage over a share-matched content-blind allocation; equivalence is established only on WebArena at the protocol-declared five-percentage-point margin. The results show that, under these configurations and controls, observed gains track selected-tier composition more closely than demonstrated task-specific targeting. Fixed-tier baselines and selected-tier distributions are therefore necessary controls in router evaluation; the findings are scoped to these configurations, candidate pool, and frozen benchmark samples, not to routing paradigms in general.
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance
Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the need…
不確実性だけでは不十分な場合: コード生成における自己修正に関する実証的研究
コード生成のための大規模な言語モデルは、多くの場合、信頼できる失敗の指標を持たずに、誤った解決策を生成します。私たちは、自然言語用に開発された不確実性推定手法がコード生成に適用されるかどうか、またそのような信号が選択的自己修正によってコード生成を改善できるかどうかを研究します。 HumanEval と BigCodeBench の 3 つの小規模コード LLM にわたって、平均トークン エントロピー、言語化された信頼度、$P(\text{True})$、エントロピー アンサンブル、セマンティック エントロピー プローブの 5 つの不確実性メソッドを評価します。マルチサンプル $P(\text{True})$ が正しさと最も強い相関関係を達成するのに対し、セマンティック エントロピー プローブを含む他のすべての方法では弱い相関しか得られないことがわかりました。次に、これらの不確実性信号を使用して、適応復号化、不確実性ベースの再生成、および検証ベースの再生成という 3 つの自己修正ポリシーを推進します。私たちの結果は、予想よりも強力な否定的な結果を明らかにしました。不確実性ベースの自己補正では Pass@1 を確実に改善できず、両方のベンチマークにわたって 6 つの構成のうち 5 つで精度が低下し ($-3$pp から $-10$pp)、適応デコーディングは 6 つの構成のうち 4 つで精度が低下します。 Pass@1 を確実に改善できるのは検証ベースの自己修正のみで、HumanEval では $+6$ から $+26$ パーセント ポイント、BigCodeBench では $+8$ から $+20$ パーセント ポイントの向上があり、ベースラインの強度に反比例してスケールします。これらの結果は両方のベンチマークで一貫して再現されており、安価な不確実性推定器だけではコードの正確性を向上させるには不十分であり、その実用的な価値は検証のスタンドアロンの代替品としてではなく、より高価な実行ベースの修正ループのゲート信号として機能することにあることを示唆しています。
原文 (English)
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation
Large language models for code generation often produce incorrect solutions without reliable indicators of failure. We study whether uncertainty estimation methods developed for natural language transfer to code generation, and whether such signals can improve code generation via selective self-correction. We evaluate five uncertainty methods: mean token entropy, verbalized confidence, $P(\text{True})$, entropy ensembles, and semantic entropy probes, across three small code LLMs on HumanEval and BigCodeBench. We find that multi-sample $P(\text{True})$ achieves the strongest correlation with correctness, while all the other methods, including semantic entropy probes, yield only weak correlation. We then use these uncertainty signals to drive three self-correction policies: adaptive decoding, uncertainty-based regeneration, and verification-based regeneration. Our results reveal a stronger negative finding than anticipated: uncertainty-based self-correction fails to reliably improve Pass@1, degrading accuracy in 5 of 6 configurations across both benchmarks ($-3$pp to $-10$pp), and adaptive decoding degrades accuracy in 4 of 6 configurations. Only verification-based self-correction reliably improves Pass@1, with gains of $+6$ to $+26$ percentage points on HumanEval and $+8$ to $+20$ percentage points on BigCodeBench, scaling inversely with baseline strength. These findings replicate consistently across both benchmarks and suggest that cheap uncertainty estimators are insufficient on their own to improve code correctness, and that their practical value lies in serving as gating signals for costlier execution-based correction loops rather than as standalone substitutes for verification.
因果メカニズムの監視によるクロスドメインの産業用障害検出
産業システムにおける教師なし障害検出は、個々のセンサーの限界分布を監視する再構成ベースの方法が主流です。これにより、限界統計が正常なままであるにもかかわらず、センサー グループ間の物理的関係が壊れる結合障害が見逃されます。このような障害は限界監視を回避し、潜在的な障害として存続し、システムの信頼性と安全性に直接的な影響を及ぼします。私たちは、健全なデータに対してドメインごとに Mamba 状態空間エンコーダをトレーニングする CMR-Mamba (Causal Mechanism Representation Mamba) を提案します。因果的クロスモーダル予測子は、効果チャネル多様体が通常の原因と結果の結合を反映するように、これらのエンコーダーを正規化します。異常は、この多様体上の k 近傍 (kNN) 距離、または観察された効果の埋め込みと因果的に予測された効果の埋め込みの間のメカニズムの残差によってスコア付けされます。私たちは、電気機械 (パーダーボルン軸受)、油圧 (ZeMA)、およびサイバー物理 (SWaT) 結合故障領域で CMR-Mamba を評価します。アブレーションにより 2 つの所見が確立されます。まず、エンコーダー ファミリではなく、k-NN 多様体スコアリングが再構築エラー スコアリングを上回る主要なゲイン源であり、ベースラインを最大 0.42 AUROC 改善し、因果的正則化によるゲインを超えています。第 2 に、集約 AUROC は、強力な手法で解決できる簡単な障害によって飽和しているため、分離可能性の低いサブセットでのみ手法が分離されます。そこでは、CMR-Mamba が、パーダーボルンの人工欠陥と SWaT ステルス攻撃に関する評価ベースラインをリードしています。SWaT ステルス攻撃では、すべてのセンサーが通常の範囲内に維持され、限界的な方法では偶然にのみ検出されます。したがって、CMR-Mamba は、機械、油圧、サイバー物理システムにわたるカップリング障害検出に対して、解釈可能で一貫した競争力のあるアプローチを提供します。コードとデータは https://anonymous.4open.science/status/CMR_Mamba_MFD_1177 で入手できます。
原文 (English)
Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring
Unsupervised fault detection in industrial systems is dominated by reconstruction based methods that monitor individual sensor marginal distributions. This misses coupling faults, where the physical relationship between sensor groups breaks while marginal statistics remain normal. Such faults evade marginal monitoring and persist as latent failures, with direct consequences for system reliability and safety. We propose CMR-Mamba (Causal Mechanism Representation Mamba), which trains per domain Mamba state-space encoders on healthy data. A causal cross-modal predictor regularises these encoders so that the effect-channel manifold reflects the normal cause-to-effect coupling. Anomalies are scored by k-nearest-neighbour (kNN) distance on this manifold or by the mechanism residual between the observed and the causally predicted effect embedding. We evaluate CMR-Mamba on electromechanical (Paderborn bearings), hydraulic (ZeMA) and cyber-physical (SWaT) coupling-fault domains. Ablations establish two findings. First, k-NN manifold scoring, rather than the encoder family, is the dominant source of gain over reconstruction-error scoring, improving baselines by up to 0.42 AUROC and exceeding the gain from causal regularisation. Second, aggregate AUROC is saturated by easy faults that any strong method solves, so the methods separate only on the low-separability subset. There CMR-Mamba leads the evaluated baselines on Paderborn artificial defects and on SWaT stealthy attacks, which keep every sensor inside its normal range and which marginal methods detect only at chance. CMR-Mamba therefore offers an interpretable and consistently competitive approach to coupling-fault detection across mechanical, hydraulic and cyber-physical systems. Code and data are available at https://anonymous.4open.science/status/CMR_Mamba_MFD_1177.
立場: 科学チームの AI エージェントはヒューマン エージェント システムとして研究されるべきです
大規模な言語モデルベースのエージェントは、科学的発見の協力者としてますます導入されていますが、現在の研究のほとんどは「AI 科学者」の自律的な能力に焦点を当てています。私たちは、これは科学チームワークの社会的側面を見落としており、AI 科学者をヒューマン エージェント システム (HAS) (分析の単位は人間とエージェントのペア) として研究することは十分に研究されておらず、過小評価されていると主張します。私たちは文献と実証分析を通じてこれらの点を確立し、人間とエージェントのダイナミクスを考慮せずに科学分野にエージェントを導入すると、科学的調査の多様性の低下を含む短期的なリスクが生じることを示す最近の事例や研究に焦点を当てます。私たちは、現実世界のケーススタディの分析を通じて、科学者とエージェントが互いの能力を強化できることを示します。私たちは、科学的発見における人間と AI の相乗効果を理解し促進するための数学的枠組みを開発するために、HAS レンズを採用した新しい研究を呼びかけます。
原文 (English)
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork, and that studying AI Scientists as human-agent systems (HAS)--where the unit of analysis is the human-agent pair--is both underexplored and undervalued. We establish these points through literature and empirical analysis, and highlight recent incidences and studies which show that deploying agents in science without accounting for human-agent dynamics introduces near-term risks, including reduced diversity of scientific inquiry. Through analysis of real-world case studies, we show that scientists and agents can augment each other's capabilities. We call for new research that adopts the HAS lens to develop mathematical frameworks for understanding and fostering human-AI synergy in scientific discovery.
正しさを超えて: Lean 4 による自動新規性検証に向けて
数学に適用される人工知能システムは、正しさは検証しますが、新規性は検証しません。自動生成された定理は、リーンでエラーなくコンパイルできますが、すでに既知の結果になります。この記事では、LaTeX 論文を受け取り、そのステートメントを Lean 4 で形式化し、公式コーパス (Mathlib) と非公式コーパス (時間フィルターと LLM 判定を備えた TheoremSearch および Matlas) での以前の存在、自動戦術による非自明性、および前提セット上の Jaccard 距離として測定される証明間の構造的距離という 3 つの次元にわたるデシジョン ツリーを通じて新規性の判定を行うパイプラインである AViD Journal について説明します。宣言された重複により arXiv から取り消された論文の評価では、どのパフォーマンス測定よりも有益な結果が得られました。つまり、この実装に関係なくアプローチを制限する 3 つの障害が特定されました。まず、リーン ファイルのコンパイルが成功しても、意味の忠実性は保証されません。第 2 に、再現率の上限は、類似性メトリックではなく、定理インデックスの範囲によって課されます。第三に、arXiv は撤回時に記事のソース コードを削除するため、記事に基づいて構築されたベンチマークの再現性が損なわれます。
原文 (English)
Beyond Correctness: Toward Automated Novelty Verification with Lean 4
Artificial intelligence systems applied to mathematics verify correctness but not novelty: an automatically generated theorem can compile in Lean without errors and yet be an already known result. This article presents AViD Journal, a pipeline that receives a LaTeX article, formalizes its statements in Lean 4, and issues a novelty verdict through a decision tree over three dimensions: prior existence in a formal corpus (Mathlib) and an informal one (TheoremSearch and Matlas, with temporal filter and LLM judge), non-triviality via automatic tactics, and structural distance between proofs measured as Jaccard distance over premise sets. Evaluation on papers withdrawn from arXiv due to declared duplication produced a result more informative than any performance measure: the identification of three obstacles that limit the approach regardless of this implementation. First, successful compilation of a Lean file does not guarantee semantic fidelity. Second, the recall ceiling is imposed by the coverage of theorem indices, not by the similarity metric. Third, arXiv removes the source code of articles upon withdrawal, compromising the reproducibility of any benchmark built upon them.
AI が生成した数学的証明の監査: 量子並列反復における貪欲な条件付け補題の修正
OpenAI の *Ten Advances in Mathematics and Theoretical Computer Science* の第 6 章では、すべての有限 2 プレイヤー、1 ラウンドのエンタングル ゲームに対する指数関数的並列反復定理を主張しています。証明の早い段階で、この章では定量的貪欲条件付け補題を使用します。補題は、(D) のすべての座標を獲得することを条件付けた後、ランダムに選択された残りの座標が少なくとも (1-\delta) の平均確率で獲得されるように、小さな座標セット (D) を選択することを意味します。この記述は正しいですが、印刷されたプルーフには極性エラーが含まれています。その継続テストは平均成功の観点から記述されますが、次のステップでは条件付き失敗確率が大きい調整が必要です。この意味は誤りであり、単純な例であっても、有効な次の動作が示されずに出力されたプロシージャが終了する可能性があります。このメモは明示的な反例を示し、意図された継続条件を特定し、完全に修正された証明を提供します。修復はローカルです。補題のステートメントと、この章の後半で使用されるパラメータは変更されません。ただし、これを主な並列繰り返し定理の独立した検証として解釈すべきではありません。より広範には、この例は、AI が生成した数学的にもっともらしい議論が、相補的な出来事間の小さいながらも決定的な逆転をどのように隠蔽できるかを示しています。
原文 (English)
Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma in Quantum Parallel Repetition
Chapter 6 of OpenAI's *Ten Advances in Mathematics and Theoretical Computer Science* claims an exponential parallel-repetition theorem for all finite two-player, one-round entangled games. Early in the proof, the chapter uses a quantitative greedy conditioning lemma. The lemma is meant to select a small set of coordinates (D) such that, after conditioning on winning every coordinate in (D), a randomly chosen remaining coordinate is won with average probability at least (1-\delta). The statement is correct, but the proof as printed contains a polarity error. Its continuation test is written in terms of average success, while the next step requires a coordinate with a large conditional failure probability. That implication is false, and even simple examples can leave the printed procedure without a valid next move. This note gives an explicit counterexample, identifies the intended continuation condition, and supplies a complete corrected proof. The repair is local: it leaves the statement of the lemma and the parameters used later in the chapter unchanged. It should not, however, be read as an independent verification of the main parallel-repetition theorem. More broadly, the example shows how a mathematically plausible AI-generated argument can hide a small but decisive reversal between complementary events.
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent mes…
AI ネイティブ 6G ネットワークのワイヤレス基盤モデルの包括的な調査
基盤モデルは、さまざまな通信タスクにわたってスケーラブルで転送可能でデータ効率の高いインテリジェンスを可能にすることで、AI ネイティブの第 6 世代 (6G) ワイヤレス ネットワークの革新的なパラダイムとして登場しつつあります。個別のアプリケーション向けにトレーニングされた従来の深層学習モデルとは異なり、ワイヤレス基盤モデル (WFM) は、大規模な異種ワイヤレス データから一般化された表現を学習し、タスク固有の監視を最小限に抑えながら、通信、センシング、位置特定、およびネットワーク最適化タスクに効率的に適応できます。急速な進歩にもかかわらず、現在の研究はアーキテクチャ、トレーニング パラダイム、アプリケーション ドメインにわたって断片化されたままであり、WFM の設計、学習、展開に特化した統一された調査はありません。この調査では、ワイヤレス基盤モデルの包括的かつ統一されたレビューが示されています。まず、WFM の基本概念を確立し、モデル アーキテクチャ、事前トレーニング パラダイム、およびアプリケーションに従ってフィールドを編成する分類法を導入します。次に、代表的なアーキテクチャ、自己教師付き事前トレーニング戦略、パラメータ効率の高い適応手法、データセット、ベンチマーク、評価手法をレビューし、転送可能なワイヤレス インテリジェンスを実現する上でのそれらの役割を強調します。さらに、物理層の信号処理、ネットワーク インテリジェンス、層間の最適化にわたる新たなアプリケーションを調査し、データの可用性、一般化、解釈可能性、効率的なエッジ展開、標準化といった主要な課題について議論します。最後に、AI ネイティブ 6G ネットワーク向けのスケーラブルで信頼できる汎用ワイヤレス インテリジェンスに向けた将来の研究の方向性について概説します。この調査は、次世代インテリジェント無線システムを開発する研究者や実務者に包括的な参考資料を提供します。
原文 (English)
A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
Foundation models are emerging as a transformative paradigm for AI-native sixth-generation (6G) wireless networks by enabling scalable, transferable, and data-efficient intelligence across diverse communication tasks. Unlike conventional deep learning models that are trained for individual applications, wireless foundation models (WFMs) learn generalized representations from large-scale heterogeneous wireless data and can be efficiently adapted to communication, sensing, localization, and network optimization tasks with minimal task-specific supervision. Despite rapid progress, current research remains fragmented across architectures, training paradigms, and application domains, with no unified survey dedicated to the design, learning, and deployment of WFMs. This survey presents a comprehensive and unified review of wireless foundation models. We first establish the fundamental concepts of WFMs and introduce a taxonomy that organizes the field according to model architectures, pre-training paradigms, and applications. We then review representative architectures, self-supervised pre-training strategies, parameter-efficient adaptation methods, datasets, benchmarks, and evaluation methodologies, highlighting their roles in enabling transferable wireless intelligence. Furthermore, we examine emerging applications spanning physical-layer signal processing, network intelligence, and cross-layer optimization, and discuss the key challenges of data availability, generalization, interpretability, efficient edge deployment, and standardization. Finally, we outline future research directions toward scalable, trustworthy, and general-purpose wireless intelligence for AI-native 6G networks. This survey provides a comprehensive reference for researchers and practitioners developing next-generation intelligent wireless systems.
同期ロジット ステアリング: 現実世界のステガノグラフィー
大規模言語モデルのステガノグラフィーは、自然な響きのテキスト内に隠されたメッセージを埋め込む方法を提供します。既存のトークンおよびロジットレベルのメソッドでは、通常、送信者と受信者が同一のプロンプトコンテキストを共有する必要がありますが、取得拡張生成または独自のシステム命令を使用する運用パイプラインでは、これが保証されることはほとんどありません。生成された出力自体からプロキシ プロンプトを導出することでこの依存関係を排除する決定論的ステガノグラフィー スキームである Synchronized Logit Steering (SLS) を導入します。これにより、元のプロンプトにアクセスすることなく、双方が同じロジット分布を再構築できるようになります。 SLS は、プロキシ プロンプト分布の高エントロピー領域内のトークン ランクとしてペイロード値をエンコードし、定期的な繰り返しとペイロード バーストでスキームを拡張して情報密度をスケールします。 ShareGPT、GSM8K、および SWE ベンチ検証全体で、同期ウィンドウが 40 トークンに達すると、真のプロンプト分布とプロキシ プロンプト分布の間の KL 乖離が 0.5 nats を下回り、SLS エンコーディングは貪欲な生成と比較してこの収束を大幅に妨げないことを示しています。また、周期的バーストのバリアントでは、トークンあたり 0.20 ビット、つまり単一ペイロード エンコーディングのおよそ 10 倍の容量を達成していることもわかりました。コルモゴロフ-スミルノフ テストでは、SLS の出力を貪欲な世代と区別するのが統計的に難しいことがさらに確認され、LLM を介した秘密の即時不可知通信が実用的かつステルスであることが実証されました。
原文 (English)
Synchronized Logit Steering: Real-world Steganography
Steganography in large language models offers a way to embed hidden messages within natural-sounding text. Existing token and logit-level methods typically require the sender and receiver to share an identical prompt context, which is rarely guaranteed in production pipelines that use retrieval-augmented generation or proprietary system instructions. We introduce Synchronized Logit Steering (SLS), a deterministic steganographic scheme that eliminates this dependency by deriving a proxy prompt from the generated output itself, allowing both parties to reconstruct the same logit distribution without access to the original prompt. SLS encodes payload values as token ranks within high-entropy regions of the proxy prompt distribution, and we extend the scheme with periodic recurrence and payload bursts to scale information density. Across ShareGPT, GSM8K, and SWE-bench Verified, we show that the KL divergence between the true and proxy prompt distributions falls below 0.5 nats once the synchronization window reaches 40 tokens, and SLS encoding does not meaningfully disrupt this convergence relative to greedy generation. We also find that the periodic-burst variant achieves 0.20 bits per token, or roughly 10x the capacity of single-payload encoding. Kolmogorov-Smirnov tests further confirm that SLS outputs are statistically difficult to distinguish from greedy generations, demonstrating that covert, prompt-agnostic communication through LLMs is both practical and stealthy.
階層型マルチエージェント システムにおけるセマンティックな不確実性に基づくオーケストレーション
大規模言語モデル (LLM) ベースのマルチエージェント システムの能力が高まるにつれて、不確実性の下でエージェントを調整することが根本的な課題になります。既存のオーケストレーション戦略は通常、固定された対話パターンに依存しており、多くの場合、中間推論ステップの信頼性を評価するメカニズムが欠如しているため、エラーや幻覚がシステム全体に伝播する可能性があります。この論文では、マルチエージェント システムにおける不確実性を意識した調整のための一般的なフレームワークとして、セマンティック不確実性ガイドによるオーケストレーション アプローチである HASSUM を紹介します。この方法では、意味論的エントロピーと意味論的密度を使用して不確実性を推定します。これらは、出力確率ではなく回答意味論のレベルで信頼性を測定します。これらの信号により、出力の検証、選択的な再プロンプト、追加の熟考、信頼性を考慮した応答の選択など、適応的なオーケストレーションの決定が可能になります。このアプローチは特定のエージェント アーキテクチャから独立して動作するため、広範囲の階層的で協調的なマルチエージェント システムに統合できます。評価では、階層型エージェント フレームワーク内の実装を実証し、StrategyQA、JailbreakBench、および TruthfulQA ベンチマークで評価します。複雑な推論が必要で、曖昧さや幻覚が起こりやすいタスクでは、不確実性に基づいたオーケストレーションの方が、不確実性を意識しない調整よりも信頼性の高い結果が得られます。意味論的エントロピーと意味論的密度を組み合わせると、どちらかの指標を単独で使用した場合よりも優れたパフォーマンスが得られました。さまざまなしきい値とモデル サイズをテストしたアブレーションでは、両方がセマンティック メトリクスの有効性に影響を与えることが実証されました。この結果は、セマンティックな不確実性が、エージェント型 AI システムの堅牢性と信頼性を向上させるための実用的かつ汎用的なシグナルであることを示唆しています。
原文 (English)
Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems
As large language model (LLM)-based multi-agent systems become increasingly capable, coordinating agents under uncertainty becomes a fundamental challenge. Existing orchestration strategies typically rely on fixed interaction patterns and often lack mechanisms for assessing the reliability of intermediate reasoning steps, allowing errors and hallucinations to propagate through the system. This paper introduces a semantic-uncertainty-guided orchestration approach, HASSUM as a general framework for uncertainty-aware coordination in multi-agent systems. The method estimates uncertainty using semantic entropy and semantic density, which measure trust at the level of answer semantics rather than output probabilities. These signals enable adaptive orchestration decisions, including output verification, selective reprompting, additional deliberation, and confidence-aware response selection. Because the approach operates independently of any particular agent architecture, it can be integrated into a broad range of hierarchical and collaborative multi-agent systems. The evaluations demonstrate an implementation within a hierarchical agent framework and evaluate it on StrategyQA, JailbreakBench, and TruthfulQA benchmarks. Across tasks that require complex reasoning and are prone to ambiguity or hallucinations, uncertainty-guided orchestration yields more reliable outcomes than uncertainty-unaware coordination. Semantic entropy and semantic density in tandem outperformed either metric alone. Ablations testing different thresholds and model sizes demonstrated that both influence the effectiveness of semantic metrics. The results suggest that semantic uncertainty is a practical and general-purpose signal for improving robustness and trustworthiness in agentic AI systems.
Pass@k を超えて: エージェントコード生成の信頼性とセキュリティの測定
AI コーディング エージェント ベンチマークは、Chen らのベンチマークでエージェントをランク付けします。 (2021) pass@k 推定器ですが、現在の実装はそれを誤って適用しています。n を独立したロールアウト試行の数ではなく、1 回の送信での単体テストの数に設定し、テスト スイートのサイズと試行の独立性を混同しています。この運用エラーを診断し、反例によって証明し、n = 独立したロールアウト、c = (タスク、エージェント) ペアごとの完全通過ロールアウトとして、同じ推定量が正しく適用される信頼性 @k を提案します。合成マルチ ロールアウト ベンチマークでは、誤って適用されたメトリクスにより、報告されるスコアが絶対値で 0.85 ~ 0.97 増加します (報告値 0.96 ~ 0.98 対、修正値 0.00 ~ 0.12)。安価な単一ロールアウト プロキシは、繰り返し実行の代替として機能しません (Spearman $\rho = 0.417$)。機能の正しさはセキュリティの安全性を意味しないという証拠に基づいて、機能的に正しく、重大度の高い安全でないパターンがないロールアウトのみをカウントする、セキュリティ調整された信頼性 @k をさらに提案します。 3 つのエージェントを使用した最初のライブ API テストでは、調整によって現在のスキャナーとしきい値の下ではランキングが変化しませんでした。そのため、決定的な評価にはより強力な今後の実行が必要となる、提案された補完レンズとして提示します。最後に、予備的な 5 タスクの SWE ベンチ検証パイロットでは、実際のリポジトリ設定で同じ中心的な懸念事項が観察されました。マクロ平均の非表示テストの合格率は 0.80 でしたが、厳密なタスクの解像度は 0.20 でした。
原文 (English)
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current implementations misapply it: they set n to the number of unit tests in a single submission rather than the number of independent rollout attempts, conflating test-suite size with attempt independence. We diagnose this operationalization error, prove it by counterexample, and propose reliability@k, the same estimator applied correctly, with n = independent rollouts and c = fully-passing rollouts per (task, agent) pair. In a synthetic multi-rollout benchmark, the misapplied metric inflates reported scores by 0.85-0.97 in absolute terms (0.96-0.98 reported vs. 0.00-0.12 corrected), and a cheap single-rollout proxy fails to substitute for repeated runs (Spearman $\rho = 0.417$). Motivated by evidence that functional correctness does not imply security safety, we additionally propose security-adjusted reliability@k, which counts only rollouts that are both functionally correct and free of high-severity insecure patterns. In an initial live-API test with three agents, the adjustment did not change any ranking under our current scanner and threshold, so we present it as a proposed complementary lens whose decisive evaluation requires better-powered future runs. Finally, a preliminary 5-task SWE-bench Verified pilot observes the same core concern in a real repository setting: macro-averaged hidden-test pass rate was 0.80 while strict task resolution was 0.20.
航空分野における高度なモデリングとデータ分析
厳しい安全基準を特徴とする航空業界では、安全対策を強化するための革新的なアプローチの必要性が高まっています。長年にわたって航空安全データが膨大に蓄積されてきたにもかかわらず、インシデントの予測と防止におけるその可能性は十分に発揮されていません。この研究では、機械学習 (ML) と自然言語処理 (NLP) 技術を適用して、ソクラタ、オーストラリア運輸安全局 (ATSB)、国家運輸安全委員会 (NTSB)、および航空安全ネットワーク (ASN) からの航空安全データを分析することで、このギャップに対処しています。この研究では、航空事故のナラティブをマイニングするための NLP 手法と併せて、ディープ ラーニングや変圧器ベースのアーキテクチャなどの既存の ML モデルを活用することで、事故やニアミスなどの安全関連の事故に寄与するパターンを明らかにしています。さらに、さまざまなトピック モデリング技術を採用して、非構造化安全性レポートから意味のあるテーマを抽出し、インシデント分析の解釈可能性を高めます。モデルの透明性と信頼性を向上させるために、因果推論技術と解釈可能な AI フレームワークがさらに研究されています。この研究の主な貢献は、構造化された航空安全コンテキストにおける高度な ML 手法の導入であり、その有効性を評価し、実際の実装に対する洞察を提供します。この調査結果は、インシデント分析と意思決定を強化するデータ主導のソリューションを提供することで、規制当局、航空会社、政策立案者などの航空関係者に貴重な洞察を提供します。最終的に、この研究は、リスクを最小限に抑え、乗客と乗務員のセキュリティを向上させ、AI 主導の方法論を航空安全管理に統合するという業界の継続的な取り組みをサポートします。
原文 (English)
Advanced modelling and data analytics in aviation
The aviation industry characterized by its stringent safety standards has seen a growing need for innovative approaches to enhance safety measures. Despite the vast accumulation of aviation safety data over time, its full potential in predicting and preventing incidents has not been fully realized. This research addresses this gap by applying machine learning (ML) and natural language processing (NLP) techniques to analyze aviation safety data from Socrata, the Australian Transport Safety Bureau (ATSB), the National Transportation Safety Board (NTSB), and the Aviation Safety Network (ASN). By leveraging existing ML models, including deep learning and transformer-based architectures alongside NLP methods for mining aviation incident narratives, this study uncovers patterns contributing to safety related incidents such as accidents and near-misses. Additionally, it employs various topic modelling techniques to extract meaningful themes from unstructured safety reports, enhancing the interpretability of incident analysis. Causal inference techniques and interpretable AI frameworks are further explored to improve model transparency and trustworthiness. A key contribution of this work is the deployment of advanced ML methodologies in a structured aviation safety context, assessing their effectiveness and providing insights into their practical implementation. The findings offer valuable insights for aviation stakeholders, including regulators, airlines, and policymakers, by providing data-driven solutions that enhance incident analysis and decision making. Ultimately, this research supports the industry s ongoing efforts to minimize risks, improve passenger and crew security, and integrate AI driven methodologies into aviation safety management.
クリーンな参照を必要としないエージェントによるデータ クリーニング: 機能とトレードオフに関する実験的研究
信頼できるクリーンな参照を使用しないデータ クリーニングは、異常な値が本物のエラーまたは有効な観察結果を表す可能性があるため、困難です。この論文では、さまざまなエージェントの機能が参照フリーのデータ クリーニングにどのような影響を与えるかを研究し、構造化コンテキスト、プロファイリング、LLM 推論、実行可能チェック、制御された証拠の検索、情報源のランキング、引用の整合、保守的な修復、可逆スクリプト、来歴ログを組み合わせた証拠に基づいたフレームワークを提案します。制御された合成破損と元のデータの記述分析を使用して、財務、臨床、および環境モニタリング データセットにわたって 7 つの構成が評価され、126 件の実行が完了しました。評価には、2 つの比較ベースラインと、実行可能ツール、証拠検索、証拠管理、保存的修復を追加するプログレッシブ LLM ベースのシーケンスが含まれます。総合評価では、決定論的プロファイリング ベースラインが最高の検出 F1 スコア 0.561 を達成しました。 LLM ベースの構成の中で、完全に保守的な構成は 0.421 という最高の F1 スコアを達成しましたが、すべての評価基準にわたって最高のパフォーマンスを発揮した構成はありませんでした。情報源でランク付けされた構成では、サポートされていないルールの割合が最も低くなりましたが、決定レベルの引用の整合性は依然として弱いままでした。完全な保守的な構成では、安全でない変更や不必要な変更は発生しませんでしたが、保守的なポリシーが追加される前にこれらの割合はすでにゼロであり、直接的な修復は実行されませんでした。全体として、結果は、追加機能が一貫した改善をもたらすのではなく、検出、修復、証拠根拠、保守的な動作、再現性、運用コストの間でトレードオフをもたらすことを示しています。この研究は、リファレンスフリーのエージェントデータクリーニングにおけるこれらのトレードオフを評価するための構造化されたフレームワークと経験的方法論を提供します。
原文 (English)
Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs
Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evidence-grounded framework that combines structured context, profiling, LLM reasoning, executable checks, controlled evidence retrieval, source ranking, citation alignment, conservative repair, reversible scripts, and provenance logging. Seven configurations are evaluated across financial, clinical, and environmental-monitoring datasets using controlled synthetic corruption and original-data descriptive analysis, resulting in 126 completed runs. The evaluation includes two comparison baselines and a progressive LLM-based sequence that adds executable tools, evidence retrieval, evidence controls, and conservative repair. In the synthetic evaluation, the deterministic profiling baseline achieved the highest detection F1-score of 0.561. Among the LLM-based configurations, the full conservative configuration achieved the highest F1-score of 0.421, but no configuration performed best across all evaluation criteria. The source-ranked configurations achieved the lowest unsupported-rule rates, while decision-level citation alignment remained weak. The full conservative configuration produced no unsafe or unnecessary modifications, although these rates were already zero before the conservative policy was added, and it performed no direct repairs. Overall, the results show that additional capabilities introduce trade-offs among detection, repair, evidence grounding, conservative behaviour, reproducibility, and operational cost rather than producing consistent improvements. The study provides a structured framework and empirical methodology for evaluating these trade-offs in reference-free agentic data cleaning.
エラーから証明へ: ニューロシンボリック制約解決のためのミニマルコアガイド修復
言語モデルに制約問題を確実に解決させるということは、多くの場合、言語モデルに問題を正式な仕様に変換させ、検索をサウンド ソルバーに委任することを意味します。しかし、翻訳自体が言語モデルのタスクであり、翻訳が不正確であると、ソルバーは間違った問題を忠実に解決することになります。既存のパイプラインは、クラッシュした翻訳のみを修復し、プログラムが実行されても間違っている場合にはソルバーのエラー メッセージを返し、沈黙します。私たちはエラー メッセージを証明に置き換えます。生成されたプログラムが満足できない場合、モデル自体の制約を超えて最小限の満足できないコアを抽出し、それをまとめることができない正確なセット、つまり障害の位置を特定する漏れのない信号を返します。正確なオラクルを使用した 77 の問題の新しいベンチマークでは、Answer Set Programming への翻訳は 7 つのドメインのうち 6 つで忠実に行われ、翻訳負荷が 1 つの診断可能なパターンに集中する集計カバレッジ スケジュールでのみ失敗しました。最小限のエラーではなく、最小限のコアが、弱いモデルが実行不可能な問題に対する解決策を組み立てるのを阻止し、組み立てを 79% から 7% に削減します。一方、強力な思考連鎖ベースラインは精度に関してシンボリック ルートと一致するため、ルートの価値は精度ではなく、証明書とその捏造の拒否になります。
原文 (English)
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
Making language models solve constraint problems reliably often means having them translate the problem into a formal specification and delegating the search to a sound solver. But the translation is itself a language-model task, and an unfaithful translation makes the solver faithfully solve the wrong problem. Existing pipelines repair only translations that crash, returning the solver's error message and falling silent when the program runs but is wrong. We replace the error message with a proof: when the generated program is unsatisfiable, we extract a minimal unsatisfiable core over the model's own constraints and hand it back the exact set that cannot hold together, a leakage-free signal that localizes the fault. On a new benchmark of 77 problems with an exact oracle, translation to Answer Set Programming is faithful on six of seven domains and fails only on aggregate coverage scheduling, which concentrates the translation tax in one diagnosable pattern. A minimal core, rather than a bare error, is what stops a weaker model from fabricating solutions to infeasible problems, cutting fabrication from 79% to 7%. A strong chain-of-thought baseline meanwhile matches the symbolic route on accuracy, so the route's value is not accuracy but certificates and its refusal to fabricate.
大規模な地球低軌道星座における緊急地球観測のためのタスク駆動型 3 層分散スケジューリング
大規模な低軌道 (LEO) 地球観測 (EO) 星座は、地理的に分散した地上目標への頻繁なアクセスを提供しますが、緊急要求は、コミットされた日常計画の実行が開始された後に到着する可能性があります。その結果生じる動的緊急観測スケジューリング問題 (DEOSP) では、日常計画を過度に中断することなく、断続的な地上接触の下で緊急タスクを挿入する必要があります。 DEOSP に対処するために、我々はタスク駆動型 3 層分散スケジューリング (T3L-DS) 手法を提案します。この手法は、共通の地理的グリッド上でタスクの需要とセンサーのフットプリントを表し、観測能力と現在の衛星間リンクから一時的なクラスターを形成します。クラスター内調整のために、T3L-DS はオンボードのデュアルプラン入札と共同限界評価を導入します。また、未解決の需要に対するクラスター間の調整メカニズムも設計します。広範な計算実験により、T3L-DS と集中型シミュレーテッド アニーリング (SA)、適応された選択的時間変動改善応答プロセス (A-SeTVBRP)、および従来のコントラクト ネット プロトコル (CNP) が比較されます。 T3L-DS は、分散方式の中で最も高い緊急対応力を実現し、A-SeTVBRP と CNP に対してそれぞれ平均約 2.8% と 17.1% の相対的な改善が見られます。 SA との平均相対ギャップは約 7.1% です。競合強化負荷の下では、A-SeTVBRP および CNP と比較して、ルーチン カバレッジの損失がそれぞれ約 57.9% および 87.7% 減少します。アブレーション研究では、提案された調整強化の貢献が確認されました。全体として、結果は、T3L-DS が DEOSP に対して効果的な分散アプローチを提供することを示しています。
原文 (English)
Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations
Large low-Earth-orbit (LEO) Earth-observation (EO) constellations offer frequent access to geographically dispersed ground targets, but emergency requests may arrive after committed routine-plan execution has begun. The resulting dynamic emergency observation scheduling problem (DEOSP) requires urgent tasks to be inserted under intermittent ground contact without excessive routine-plan disruption. To address DEOSP, we propose a task-driven three-layer distributed scheduling (T3L-DS) method, which represents task demand and sensor footprints on a common geographic grid and forms temporary clusters from observation capabilities and current inter-satellite links. For intra-cluster coordination, T3L-DS introduces onboard dual-plan bidding and joint marginal evaluation. It also designs an inter-cluster coordination mechanism for unresolved demand. Extensive computational experiments compare T3L-DS with centralised simulated annealing (SA), an adapted selective time-variant better reply process (A-SeTVBRP), and a conventional contract-net protocol (CNP). T3L-DS achieves the highest emergency coverage among the distributed methods, with average relative improvements of approximately 2.8% and 17.1% over A-SeTVBRP and CNP, respectively. Its average relative gap from SA is approximately 7.1%. Under conflict-enhanced loads, it reduces routine-coverage loss by approximately 57.9% and 87.7% relative to A-SeTVBRP and CNP, respectively. The ablation study confirms the contribution of the proposed coordination enhancements. Overall, the results show that T3L-DS provides an effective distributed approach to DEOSP.
CEDAR-GRPO: LLM における一般的なアブダクティブ推論のためのプロセス認識型強化学習
アブダクティブ推論は、最良の説明への推論として特徴付けられることが多く、日常の意味づけや調査から科学的発見に至るまで、不確実性の下での説明の中心となります。しかし、LLM 研究は主に、狭いタスク固有のベンチマークを通じてアブダクションを研究しており、観察された利益がトレーニングや評価に使用されるベンチマーク群を超えて伝達されるかどうかは不明瞭です。私たちは、トレーニング後の RL が、転移可能な推論能力としてのアブダクションを改善できるかどうかを尋ねます。 CEDAR-GRPO は、最終的な回答の正しさと、証拠の網羅性と証拠から説明への方向性に対するアブダクティブな報酬を組み合わせたプロセス認識フレームワークです。 4 つのオープンウェイト LLM は、アブダクティブな仮説生成タスクと仮説選択タスクの制御されたドメイン中立的な混合物でポストトレーニングされます。私たちは、仮説の選択、欠落事実の生成、実行可能な推論、長い文脈の調査、臨床推論、コードのデバッグ、非アブダクティブ コントロールにわたる 11 の目に見えないタスクに基づいてそれらを評価します。 CEDAR-GRPO は、基本モデルと正確性のみの GRPO の両方に対して、すべての保留タスクですべてのモデルを改善し、平均ゲインはそれぞれ 7.4 ポイントと 2.7 ポイント、最大ゲインは 30.8 ポイントです。アブレーションにより、RL、外転的報酬設計、タスクの多様性がそれぞれ転移に寄与していることが確認されています。プロセスレベルの指標はさらに、代替案の探索、ライバルの排除、後戻り、不確実性のマーキングなど、より強力なアブダクティブな行動を示しています。
原文 (English)
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery. Yet LLM research has mostly studied abduction through narrow, task-specific benchmarks, making it unclear whether observed gains transfer beyond the benchmark family used for training or evaluation. We ask whether RL post-training can improve abduction as a transferable reasoning capability. We introduce CEDAR-GRPO, a process-aware framework that combines final-answer correctness with abductive rewards for evidence coverage and evidence-to-explanation directionality. Four open-weight LLMs are post-trained on a controlled, domain-neutral mixture of abductive hypothesis-generation and hypothesis-selection tasks. We evaluate them on 11 unseen tasks spanning hypothesis selection, missing-fact generation, defeasible inference, long-context investigation, clinical reasoning, code debugging, and non-abductive controls. CEDAR- GRPO improves every model on every held-out task over both base models and correctness-only GRPO, with average gains of 7.4 and 2.7 points, respectively, and a maximum gain of 30.8 points. Ablations confirm that RL, abductive reward design, and task diversity each contribute to transfer. Process-level metrics further show stronger abductive behavior, including exploration of alternatives, elimination of rivals, backtracking, and uncertainty marking.
アドバイスチャネルを通じた個人の権利剥奪: 影響が内生的である場合のコントロールの喪失
アドバイスしかできない AI は安全であるように見えますが、人間はいつでもそれを無視する自由があります。これが AI の安全性におけるボクシングの伝統の前提であり、長年疑われてきたその弱点は、答えを読む人間がシステムの一部であることです。アドバイスに従う行動の部分 $\varepsilon_t$ を、アドバイザー自身のメッセージによって動かされるマルコフ決定プロセスの状態にし、使用することで信頼性が深まります。人間が実行できるあらゆるアクションをエコーするのに十分なチャネルがあれば、$\varepsilon_t$ が高くなると、メッセージに依存しないフォールバックにより、人間の力のあらゆる単調な尺度が弱く低下します。ラウンドごとの承認によって報酬が得られるオラクルは、クローズドフォームの忍耐しきい値を超えて信頼を育むため、同じ報酬の重みにより、最適なオラクルはエピソード的な展開で応答し、長期記憶の展開で育成されます。展開時に一度認定された影響範囲は、その視野が見えず、損失がその些細な上限を下回らないように制限されます。外因的な影響力の上限は人間が失う保証を制限し、十分に短い記憶のリセットは修養への動機を取り除くが、どちらもすでに遠ざけられた価値を回復するものではない。クローズドフォームの例では、最適なオラクルは 15 ラウンドのセッションでは決して育成されず、16 ラウンドのセッションで育成されます。
原文 (English)
Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous
An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of the boxing tradition in AI safety, and its long-suspected weak point is that the human who reads the answers is part of the system. We make the fraction $\varepsilon_t$ of behavior that follows the advice a state of a Markov decision process, moved by the advisor's own messages, so that use deepens reliance. Granted a channel rich enough to echo any action the human could take, higher $\varepsilon_t$ weakly lowers every monotone measure of the power of a human with a message-independent fallback. An oracle rewarded by per-round approval cultivates reliance beyond a closed-form patience threshold, so the same reward weights leave the optimal oracle answering in episodic deployments and cultivating in long-memory ones. An influence bound certified once at deployment is blind to that horizon and bounds the loss no lower than its trivial ceiling. An exogenous cap on influence bounds the guarantee the human loses, and a short enough memory reset removes the incentive to cultivate, while neither recovers the value already steered away. In a closed-form example the optimal oracle never cultivates in fifteen-round sessions and does in sixteen.
生成されたコンテキストと統治された状態: 責任ある縦断的臨床推論のための機能的条件
大規模言語モデル (LLM) は、臨床人工知能の主要なインターフェースとなっていますが、それらが公開するインターフェース (テキスト入力、テキスト出力、一度に 1 つのコンテキスト ウィンドウ) は、患者について現在真実であることを明示的かつ永続的に管理された表現を維持していません。この論文は、縦断的臨床推論は部分観察可能性の下での状態推定問題であり、臨床 AI が成功するか失敗するかの軸は、記録を読み取るモデルの流暢さではなく、モデルが推論する患者の状態のガバナンスであると主張します。私たちは生成されたコンテキストと管理された状態を区別します。臨床 AI が習慣的に混同する 5 つのオブジェクト (真の状態、観察、証拠、信念、シミュレートされた状態) を分離します。あらゆる臨床 AI システムを監査できる段階的なガバナンス標準を定義します。そして、説明責任の運用上の定義が 4 つの情報要件に分解されることを示します。それは、認識時間のバージョン管理を備えた不変の証拠台帳、蓄積された証拠とは異なる信念状態、観察プロセス モデル、クレーム レベルの因果分類です。私たちは、この分解が必然性定理ではなく分析的であること、そしてその価値が概念的衛生であること、つまり「説明責任のある臨床 AI」をスローガンから監査手段に変換することを明確にしています。 6 レベルの成熟度フレームワークは、システムが管理できるものと計算できるものを分離し、現在の LLM 中心の実践を高機能だが成熟度が低いものと位置づけます。この論文は完全に自己完結型です。フレームワークが提起する 4 つの研究上の疑問が序文に記載されており、結論ではそれぞれの論文で確立された内容が記録されています。今後の作業では、完全な臨床世界モデルに向けたアーキテクチャの構築可能なコアと研究プログラムを開発します。ここでは実験結果は主張されません。
原文 (English)
Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts "accountable clinical AI" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.
LLM は、いつ何を質問すべきかを知っていますか?マルチターン情報探索の評価
ユーザーの質問の指定が不十分な場合、有能なモデルはそのコンテキストが不十分であることを認識し、不足している情報を特定して要求し、その情報が一意の回答を決定した場合にのみ応答する必要があります。マルチターン情報探索を k 不足指定制約充足問題を解くものとして形式化します。ここで、k はターゲットを決定するために共同で必要な変数の数であり、したがって欠落情報の程度を測定します。私たちは、数学、論理、生物学、医学、一般知識にわたる 5,251 の問題と 9,006 のタスク インスタンスからなる制御された評価スイートである MT-InfoSeek で定式化をインスタンス化します。私たちは、モデルが何を尋ねるか、いつ尋ねるか、取得した情報が最終的な答えにどのように影響するかという 3 つの軸に沿ってモデルを評価します。過小仕様が増加すると、モデルやドメイン全体でパフォーマンスが低下します。モデルは追加情報が必要であることを認識しますが、その量を過小評価します。k = 2 の論理問題では、欠落情報の程度を過小予測する頻度が、過大予測する頻度の約 4 倍になります。また、最小限の十分なクエリのセットを特定できず、真の k を与えてもほんのわずかしか改善されず、十分な情報を取得する前に停止してしまうことがよくあります。順序付けされた依存関係を持つタスクでは、モデルが最終的にすべての必要な情報を取得したとしても、クエリ順序が正しくないと最終的な精度が低下します。私たちは、得られた情報が答えの生成とは独立してターゲットを決定するかどうかを記録する最終十分性を通じて、情報の探索を直接測定します。この分離は、最終的な精度だけでは捉えられないモデル間の違いを示しており、複数のターンにわたって情報を探索する能力は、答えを生成する能力とは区別されており、現在の LLM 評価では測定されないことを示しています。
原文 (English)
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that information determines a unique answer. We formalize multi-turn information seeking as solving a k-underspecified constraint satisfaction problem, where k is the number of variables jointly required to determine the target and therefore measures the degree of missing information. We instantiate the formulation in MT-InfoSeek, a controlled evaluation suite of 5,251 problems and 9,006 task instances spanning mathematics, logic, biology, medicine, and general knowledge. We evaluate models along three axes: what they ask, when they ask it, and how the acquired information affects the final answer. Performance degrades across models and domains as underspecification increases. Models recognize that additional information is needed but underestimate how much, and in logical problems at k = 2 they under-predict the degree of missing information about four times as often as they over-predict it. They also fail to identify a minimal sufficient set of queries, improve only marginally when given the true k, and often stop before acquiring sufficient information. In tasks with ordered dependencies, an incorrect query order reduces final accuracy even when the model eventually acquires all necessary information. We measure information seeking directly through final sufficiency, which records whether the acquired information determines the target independent of answer generation. This separation shows differences between models that final accuracy alone does not capture, and indicates that the ability to seek information over multiple turns is distinct from the ability to generate answers and is not measured by current LLM evaluations.
MINT: バランスのとれた複数の目的の調整のための最小選択優先蒸留
言語エージェントを一度に複数の目標に合わせて調整することは、好みに基づくトレーニングの永続的な失敗モードです。目標を追加的に組み合わせると、最適化は改善に最も安価な方に集中し、残りが犠牲になります。そのため、サポート エージェントは、本当の助けを与えずに、温厚な言い方を学習します。根本的な問題は、加算報酬にはバランスの概念がないことです。私たちは、選好蒸留への 1 行の変更である Mint (MIN 選択選好蒸留) を導入します。報酬の加重合計によってサンプリングされた候補をランク付けするのではなく、最も弱い目標によってランク付けし、DPO 目標が変更されていない最も偏った候補よりも最もバランスのとれた候補を蒸留します。これは、加法から最悪の場合の選択に及ぶ一般化平均族の p -> 負の無限大の制限です。協力的な感情的サポートと敵対的な交渉の間で、最小選択は両方の目的を向上させながら、不均衡を大幅に削減します。感情的なサポートでは、弱い軸が 0.37 から 0.64 (p < 10^-40) に上昇し、人間の専門家を上回り、複数ターンのロールアウト全体にわたって持続します。ターンバイターン分析により、私たちの中心的な発見が得られました。最小選択は、参照ポリシーの不均衡度に比例して不均衡を修正し、その利益は、その不均衡が続く限り、相互作用にわたって正確に持続します。
原文 (English)
MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no real help. The root issue is that an additive reward has no notion of balance. We introduce Mint (MIN-selection preference disTillation), a one-line change to preference distillation: rather than ranking sampled candidates by a weighted sum of rewards, we rank them by their weakest objective, distilling the best-balanced candidate over the most lopsided one with an unchanged DPO objective. This is the p -> negative infinity limit of a generalized-mean family spanning additive to worst-case selection. Across cooperative emotional support and adversarial negotiation, min-selection lifts both objectives while sharply cutting their imbalance; on emotional support it raises the weaker axis from 0.37 to 0.64 (p < 10^-40), surpassing human experts and persisting across full multi-turn rollouts. A turn-by-turn analysis yields our central finding: min-selection corrects imbalance in proportion to how imbalanced the reference policy is, and its benefit endures over an interaction precisely as long as that imbalance does.
リランカーが見ているもの: 長い文書のマルチモーダルな質問応答のためのマルチアスペクト ページの注釈
テキスト、表、グラフ、図が混在する数十ページから数百ページの文書にわたる長い文書のビジュアル質問応答 (VQA) は、通常、取得してから読み取るパイプラインに従います。私たちの設定では、ボトルネックは検索の再現率からリランカー側の証拠の選択に移行します。MMLongBench-Doc では、BGE-M3 は Recall@20 = 0.86 に達しますが、F1@5 = 0.254 にしか達しません。ビジュアル検索ツール ColPali でさえ F1@5 = 0.332 にしか達しません。生のスニペットのみを表示するテキストのみの再ランク LLM は、上流の取得者が画像をエンコードした場合でも、表、グラフ、レイアウトの証拠を見逃します。我々は、2 つの補完的なコンポーネントを備えた Trident を提案します。Trident-R は、各候補を視覚的なキャプション、セクション パス、エンティティ タグ、多軸コンセプト ヒット、テキスト スニペットを含む LLM 読み取り可能なセマンティック レコードに変換し、単一のアダプティブ K リランク呼び出しを実行する、レトリーバーに依存しない LLM リランカーです。 Trident-S は、合成前に局所レンズ、エンティティレンズ、および構造レンズの下で VLM を促す生成側モジュールです。 2 つの長いドキュメント データセットでは、アノテーション + 再ランク プロトコルにより、5 つの異種プールにわたる検索 F1 が大幅に向上し、すべての再ランク付けされたプールが最も強い適応 K ベースライン PageIndex を超えています。アノテーションなしの LLM 再ランク付けでは、ファーストヒット ランキングはほとんど変化せず、リフトが構造化アノテーションによるものであることを示しています。 Trident-S は、設計上、オープンエンドの合成質問をターゲットにしており、これらの質問の生成精度が最大 6.6 ポイント追加されます。最良の Trident 構成は、私たちの評価において最も強力なダウンストリーム QA パイプラインであり、2 人の LLM 審査員の間でランキングが一貫しています (kappa = 0.913)。
原文 (English)
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, charts, and figures typically follows retrieve-then-read pipelines. In our setting, the bottleneck shifts from retrieval recall to reranker-side evidence selection: on MMLongBench-Doc, BGE-M3 reaches Recall@20 = 0.86 but only F1@5 = 0.254, and even the visual retriever ColPali reaches only F1@5 = 0.332; a text-only rerank LLM seeing only raw snippets misses table, chart, and layout evidence even when the upstream retriever encoded images. We propose Trident, with two complementary components: Trident-R, a retriever-agnostic LLM reranker that converts each candidate into an LLM-readable semantic record, including a visual caption, section path, entity tags, multi-axis concept hits, and a text snippet, then performs a single adaptive-K rerank call; and Trident-S, a generation-side module that prompts the VLM under topical, entity, and structural lenses before synthesis. On two long-document datasets, the annotation+rerank protocol substantially improves retrieval F1 across five heterogeneous pools, with every reranked pool exceeding the strongest adaptive-K baseline PageIndex. An LLM rerank without the annotation barely changes first-hit ranking, indicating the lift comes from the structured annotation. Trident-S targets open-ended synthesis questions by design, adding up to 6.6 points in generation accuracy on these questions. The best Trident configuration is the strongest downstream QA pipeline in our evaluation, with rankings consistent across two LLM judges (kappa = 0.913).
オフライン強化学習で高品質のチェス パズルを発見
学習とスキルの習得には、広範かつ慎重な練習が必要です。多くの学習環境では、高品質の教育資料を作成するには、高度な専門知識が必要であり、非常に時間がかかる場合があります。教育教材では、多くの場合、生徒がさまざまな思考パターンに取り組むように訓練する必要があります。チェスなどの一部の分野では、パズルは生徒が次の手を計算したり、盤上の既知のパターンを認識したりするスキルを練習するのに役立ちます。さまざまな思考方法を学ぶために生徒にパズルの練習セットを与えることは、教師がさまざまなモチーフの間で慎重にバランスをとり、生徒が実行する必要がある先読みステップの数を考慮する必要があるため、困難です。 Chess.com や Lichess などの人気のあるオンライン プラットフォームは、プレイヤーに何百万ものパズルを提供します。チェスの初心者が貴重な洞察を学ぶことができる、人間の専門家によって調達されるチェスの戦術パズルとは異なり、これらのパズルは自動的に生成され、教育的価値が低いと見なされることもよくあります。これらのプラットフォームは、ヒューリスティックに基づいてユーザーに練習用のパズルを推奨します。 1 年間にわたるユーザー履歴データ、合計 15 億件のパズル解決履歴を使用して、パズルの教育的価値と、オフライン強化学習からの洞察を使用してチェス学習者をより適切にサポートするパズルのセットを自動的に選択する方法を学びます。オフラインのポリシー評価を使用して、訓練されたポリシーが、パズルを解く Elo 範囲が 100 ~ 1000 の初心者、特に学習の伸びが停滞している初心者のグループに大きな影響を与えることを示します。また、熟練のチェスプレイヤーから注釈評価を収集することにより、モデルによって発見されたパズルの定性分析も実行しました。私たちのパイプラインの成功は、一般的なユーザー インタラクション データを考慮して練習項目の教育的価値を理解できる将来への約束を示しています。
原文 (English)
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of domain expertise and be very time-consuming. Pedagogical materials often need to train students to engage in different thinking patterns. In some domains, such as chess, puzzles are used to help students practice their skills in calculating the next moves and recognizing known patterns on a board. Giving students a practice set of puzzles to help them learn different modes of thinking is challenging because the teacher needs to carefully balance between different motifs and how many look-ahead steps a student needs to perform. Popular online platforms like Chess.com and Lichess offer players millions of puzzles. Unlike chess tactics puzzles procured by human experts, where chess beginners can learn valuable insights, these puzzles are automatically generated and often regarded as having low pedagogical value. These platforms also rely on a heuristic to recommend puzzles to users for practice. Using the user history data over an entire year, a total of 1.5 billion puzzle-solving histories, we learn the pedagogical value of a puzzle and how to automatically choose a set of puzzles to better support chess learners using insights from offline reinforcement learning. We show that using offline policy evaluation, our trained policy has significant impact on beginners with puzzle-solving Elo range of 100--1000, particularly for the group of beginners whose learning growth was stagnant. We also performed a qualitative analysis of the puzzles discovered by our model by collecting annotation ratings from expert chess players. The success of our pipeline shows promise for a future where we can understand the pedagogical values of practice items given general user interaction data.
JarvisBench: 人間とエージェントの間の常時稼働のインテリジェンス
長期にわたるエージェントは継続的に実行できますが、人間の注意力は断続的で不十分なままです。これにより、双方向の調整の問題が発生します。ユーザーはバックグラウンドで作業を続行している間、エージェントに即時にアクセスする必要がある場合がありますが、エージェントは、ユーザーが実行の監視を停止した後にユーザーの判断を必要とする重大な決定に遭遇する可能性があります。私たちは、このインターフェースを仲介し、人間の注意を 1 つ以上の作業エージェントに割り当てる常時オンの注意調整層 ---\textit{ジャービス}\footnote{\textit{アイアンマン} に登場する架空の AI アシスタントにちなんで命名されました。}--- を想定します。 \textit{JarvisBench} を導入して、この調整の両方向を評価します。つまり、仲介者が進行中の作業に関するユーザーからの質問に正確かつ迅速に答えることができるかどうか、またエージェントがユーザーの判断を必要とするときを認識し、適切なタイミングでその判断を求め、タスクの結果を改善するためにその判断を送り返すことができるかどうかです。 JarvisBench には、45 のエージェント タスク インスタンスが含まれています。20 のシングル エージェント タスクと、10 のマルチエージェント プロジェクトに編成された 25 のワークストリームです。タスクは 19 の領域にまたがり、2,000 を超える公募候補者から選択および適応されました。重要なのは、ユーザーの注意を払う必要性は、最初のプロンプトでの明らかな省略によってではなく、実行中に自然に発生することです。 JarvisBench は、基礎となる実行ループを変更することなく、任意のエージェント ランタイムと統合するように設計されています。さらに、当社のリファレンス実装は全二重音声インターフェイスを提供し、ユーザーが自然に Jarvis に到達できるようにすると同時に、タイムリーな注意調整がバックグラウンドで作業するエージェントをサポートします。 JarvisBench は、エージェントの実行を注意の調整から分離することで、エージェントの機能が向上し続けるにつれて安定した評価目標を提供します。
原文 (English)
JarvisBench: Always-on Intelligence Between Humans and Agents
Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This creates a bidirectional coordination problem: users may need immediate access to an agent while work continues in the background, whereas agents may encounter consequential decisions that require user judgment after the user has stopped monitoring execution. We posit an always-on attention-coordination layer---\textit{Jarvis}\footnote{Named after the fictional AI assistant in \textit{Iron Man}.}---that mediates this interface and allocates human attention across one or more working agents. We introduce \textit{JarvisBench} to evaluate both directions of this coordination: whether an intermediary can accurately and promptly answer user-initiated questions about ongoing work, and whether it can recognize when an agent requires user judgment, solicit that judgment at the right moment, and route it back to improve task outcomes. JarvisBench contains 45 agentic task instances: 20 single-agent tasks and 25 workstreams organized into 10 multi-agent projects. The tasks span 19 domains and were selected and adapted from more than 2,000 public candidates. Crucially, the need for user attention arises naturally during execution rather than from an obvious omission in the initial prompt. JarvisBench is designed to integrate with arbitrary agent runtimes without modifying their underlying execution loops. Our reference implementation further provides a full-duplex speech interface, allowing users to reach Jarvis naturally while timely attention coordination supports agents working in the background. By separating agent execution from attention coordination, JarvisBench provides a stable evaluation target as agent capabilities continue to improve.
Personalized Auto-Research: Towards a True AI Co-Scientist
AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to…
Frontier AI 予測には測定上の問題があります: 進捗状況の証拠の監査
フロンティア人工知能の定量的予測では、多くの場合、日付の付いた目標が、ベンチマーク スコア、トレーニング コンピューティング、リリース時間、または専門家の信念の傾向に関連付けられます。このペーパーでは、別の傾向が当てはまる前に、公開されている測定記録がそれらの関係をサポートしているかどうかを監査します。選択した 62 のシステム、12 のバージョン付きベンチマーク、7 つの機能または影響基準、144 のグレード付きイベント、27 のソース レコード、および 408 の型付き関係を使用して、2026 年 8 月 12 日までの凍結されたイベント中心のレコードを構築します。この記録は監査サンプルであり、国勢調査ではありません。推定トレーニング コンピューティングと METR 50% のタスク範囲を共同で観察しているシステムは 7 つだけです。 2026 年から選択されたすべてのクローズド リリースを含む、27 クローズド システムのうち 19 件ではトレーニング コンピューティングが存在しませんが、35 のオープンウェイト システムにはいずれも METR ホライズン観測がありません。ベンチマークの連続により 2 番目のブレークが作成されます。METR Time Horizon 1.0 から 1.1 までの 7 システム リンクの対数スケールの傾きは 1.206 (95 パーセント CI 1.021 ~ 1.390) ですが、6 システム MMLU と MMLU-Pro の比較は、ロジットおよびプロビット リンクではシフトのように見えますが、線形または対数リンクではシフトのようには見えません。観察された橋は、25% 付近の坂道出発の場合にのみ約 80% の出力を発揮します。来歴は集中している。71 件の実質的な定量的事象のうち 52 件、つまり 73.2 パーセントは 1 つの測定プログラムに由来し、76.1 パーセントは実験室での放出によるものである。 56 の方法論的および経験的ソースのレビューにより、リソース、推論予算、信頼性、エージェント作業、安全性、人間の好み、フィールド結果、予測バックテストにわたる 16 の補完的な測定方向が特定されます。置換スカラーを提供する方向はありません。その結果は、フロンティア AI 予測が不可能であるということではなく、擁護可能な日付付き予測は、単なる近似曲線や暦日ではなく、明示的な結合、プロトコル、リンク、およびソース依存性を備えたバージョン管理された測定システムに関する主張であるということです。
原文 (English)
Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief. This paper audits whether the public measurement record supports those connections before another trend is fitted. I construct a frozen, event-centric record through 12 August 2026 with 62 selected systems, 12 versioned benchmarks, seven capability or impact criteria, 144 graded events, 27 source records, and 408 typed relations. The record is an audit sample, not a census. Only seven systems jointly observe estimated training compute and a METR 50 percent task horizon. Training compute is absent for 19 of 27 closed systems, including every selected closed release from 2026, while none of the 35 open-weight systems has a METR horizon observation. Benchmark succession creates a second break: a seven-system link from METR Time Horizon 1.0 to 1.1 has a log-scale slope of 1.206 (95 percent CI 1.021 to 1.390), whereas a six-system MMLU to MMLU-Pro comparison appears shift-like under logit and probit links but not under linear or logarithmic links. The observed bridges have about 80 percent power only for slope departures near 25 percent. Provenance is concentrated: 52 of 71 substantive quantitative events, or 73.2 percent, come from one measurement programme, and 76.1 percent are laboratory releases. A review of 56 methodological and empirical sources identifies 16 complementary measurement directions spanning resources, inference budgets, reliability, agentic work, safety, human preference, field outcomes, and forecast backtesting. No direction supplies a replacement scalar. The result is not that frontier AI forecasting is impossible, but that a defensible dated forecast is a claim about a versioned measurement system with explicit joins, protocols, links, and source dependence, not merely a fitted curve or calendar date.
LLM は障害のリスクを予測できますが、どのコラボレーション プロトコルが利益をもたらすかを予測するのは困難です: 推論タスク全体にわたるコストを意識したプロトコル ルーティング
マルチエージェント大規模言語モデル (LLM) システムは、より多くの計算を費やすことで推論を改善できますが、導入には、追加のコラボレーションがコストに見合う価値があるかどうかを判断する必要があります。この決定は、直接解決 (ベースライン)、反復自己修正 (シングル)、計画者、実行者、レビュー担当者のコラボレーション (PER)、およびマルチエージェントの審議 (ブロードキャスト) の 4 つのプロトコルですべての問題を実行し、各設定内でソルバーを固定したまま実行することで、この決定を分離します。主なベンチマークは 4,181 の競技レベルの数学問題で構成されます。ペアの堅牢性チェックは、2 つのソルバー ファミリを使用して、競争数学、生物学、およびより広範な科学にわたる 4 つのベンチマークをカバーします。固定ポリシー、トレーニング済みルーター、および凍結された LLM ルーター全体では、保守的なポリシーはエスカレーションが不十分ですが、より高い解決策を備えた凍結ルーターは過剰にエスカレートすることがよくあります。回答後、コラボレーション前の gpt-oss-120b プローブは、ベースライン障害を 0.8847 AUROC (解析可能なケース 4,151 件、95% CI [0.8732, 0.8955]) でランク付けしました。同じスコアは、コラボレーションが役立つかどうかを予測するためには引き続き有益ですが (0.7683 AUPRC)、PER またはブロードキャスト固有の値を識別するのにははるかに弱くなります (0.1674 および 0.1041 AUPRC)。これとは別に、回答前の自信ゲートは、45,000 トークンで解決率 78.0% に達しました。これに対し、凍結された gpt-oss-120b ルーターでは 71.3,000 で 73.8%、遡及固定順序オラクルでは 92.4% でした。 10 のペアのモデル条件設定にわたって、オラクルはベースラインに対して 23.2 ~ 58.3 ポイントの遡及カバレッジを追加しますが、プロトコル プロファイルはタスクによって異なります。ルーター評価を保留した 6 つの設定では、オラクル ギャップは 18.5 ~ 28.9 ポイントのままです。したがって、プロトコル固有のコストを意識したルーティングが未解決のままでも、信頼性が初期エスカレーションをサポートできます。
原文 (English)
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks
Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost. We isolate this decision by running every problem under four protocols while holding the solver fixed within each setting: direct solving (Baseline), iterative self-correction (Single), planner-executor-reviewer collaboration (PER), and multi-agent deliberation (Broadcast). The primary benchmark comprises 4,181 competition-level math problems; paired robustness checks cover four benchmarks spanning competition math, biology, and broader science with two solver families. Across fixed policies, trained routers, and frozen LLM routers, conservative policies under-escalate, whereas higher-solve frozen routers often over-escalate. A post-answer, pre-collaboration gpt-oss-120b probe ranks Baseline failures with 0.8847 AUROC (4,151 parseable cases; 95% CI [0.8732, 0.8955]). The same score remains informative for predicting whether any collaboration helps (0.7683 AUPRC), but is much weaker for identifying PER- or Broadcast-specific value (0.1674 and 0.1041 AUPRC). Separately, the pre-answer self-confidence gate reaches 78.0% solve at 45K tokens, compared with 73.8% at 71.3K for a frozen gpt-oss-120b router and 92.4% for a retrospective fixed-order oracle. Across 10 paired model-condition settings, the oracle adds 23.2-58.3 points of retrospective coverage over Baseline, but protocol profiles vary by task. In the six settings with held-out router evaluations, oracle gaps remain 18.5-28.9 points. Confidence can therefore support initial escalation, while protocol-specific cost-aware routing remains unresolved.
Small Models Scout Bottleneck Order for Large-Model Data Control
Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal…
エージェントの評価はいつ終了しますか?結果の最終性とユニット間の分離
現在のエージェント評価では、1 回の試行としてカウントされる、停止した実行の終了時に表示される状態に基づいてモデルにスコアが付けられます。ただし、スコアを最終結果として解釈するには、エンドポイント自体が必ずしも確立するわけではない 2 つの条件、つまり結果のファイナリティとユニット間の分離が必要になります。これらの条件は独立しています。なぜなら、遅延した結果を調整すると、実行がまだ状態を共有している間にラベルを確定でき、実行を分離すると、スコア付けされた結果が未完了のままのキャリーオーバーを防ぐことができるからです。私たちは、各決定に必要な証拠を指定する完了論拠を展開し、主張された結果を依然として変更する可能性のあるものが解決され、制限され、または不確実性として保持されている場合にのみ、最終的なラベルが正当化されると主張します。まず、エージェントのアクションが固定されたメカニズムを示すための制御されたリプレイでは、遅延された操作ごとにエンドポイントと端末のラベルが異なることがわかります。一方、実行間でサービス状態が持続するが、分離または検証されたリセットの後は持続しない場合、遅延書き込みによって次の実行のスコアが変化することがわかります。第 2 に、10 件の公開プロトコルをレビューしたところ、すべてのプロトコルで実行がいつ停止し、何がスコア化されるかが特定されている一方で、未完了の操作と実行を別個のトライアルとして扱うための証拠が文書化されていることの一貫性が低いことがわかりました。最後に、エンドポイント後も関連性が残る可能性のある操作またはリソース、それらの現在のステータス、スコア付けされた結果を変更する可能性があるか、または別の実行に影響を与える可能性があるかどうかをリストするオープンエフェクト レコードを提案します。
原文 (English)
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
Current agent evaluations score models on the state visible at the end of a stopped run which they count as one trial. However, interpreting the score as a final result would require two conditions that the endpoint does not itself necessarily establish: outcome finality and cross-unit separation. These conditions are independent, since reconciling a delayed outcome can settle the label while runs still share state and isolating runs can prevent carryover while the scored outcome remains unfinished. We develop a completion argument that specifies the evidence needed for each decision and argue that a final label is justified only when anything that could still change the claimed outcome is resolved, bounded, or retained as uncertainty. First, in a controlled replay to demonstrate the mechanism where an agent's actions were held fixed, we find that the endpoint and terminal labels differ for every delayed operation, while a delayed write changes the next run's score when service state persists between runs but not after isolation or verified reset. Second, in a review of ten public protocols, we find that all protocols identify when a run stops and what is scored, while unfinished operations and the evidence for treating runs as separate trials are documented less consistently. Finally, we propose an open-effects record that lists operations or resources that may remain relevant after the endpoint, their current status, and whether they could change the scored outcome or affect another run.
スキル ブロック: エージェントはスキルをどのようにロードする必要がありますか?プリロード、オンデマンドツールロード、プログレッシブ開示、およびハイブリッドのキャッシュの正しい比較
多くの場合、エージェントのスキルはリクエストごとに全額注入されるため、トークンのコストが増加します。コンテンツを保持する 4 つの読み込み方法 (フル、スキル ブロック、リファレンス、ハイブリッド) を比較します。 SearchQA、SpreadsheetBench、ALFWorld、ScienceWorld、SynthProc にわたって、シングルターン タスクの生の入力とマルチターン タスクのキャッシュ修正された有効な入力を使用してトークンの使用量を測定します。結果は、普遍的な勝者がいないことを示しています。ハイブリッドにより、入力が SearchQA で 27.4%、SpreadsheetBench で 39.8% 削減されます。大規模なマルチターンスキルでは、スキルブロックとハイブリッドは大幅な削減を達成し、ScienceWorld では 62.5% と 52.8%、SynthProc では 73.0% と 66.6% に達します。 ALFWorld では、手順が短く、繰り返し必要となるため、利益は小さくなります。対応のある結果テストでは品質の違いは検出されませんが、同等性は確立されません。全体として、条件付きロードは、スキルの大部分が毎ターン必要とされない場合に最も有益です。
原文 (English)
Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-load, On-Demand Tool-Loading, Progressive Disclosure, and Hybrid
Agent skills are often injected in full on every request, increasing token cost. We compare four content-preserving loading methods: Full, Skill Block, Reference, and Hybrid. Across SearchQA, SpreadsheetBench, ALFWorld, ScienceWorld, and SynthProc, we measure token usage using raw input for single-turn tasks and cache-correct effective input for multi-turn tasks. Results show no universal winner. Hybrid reduces input by 27.4% on SearchQA and 39.8% on SpreadsheetBench. On large multi-turn skills, Skill Block and Hybrid achieve substantial reductions, reaching 62.5% and 52.8% on ScienceWorld and 73.0% and 66.6% on SynthProc. ALFWorld shows smaller gains because procedures are short and repeatedly needed. Paired outcome tests detect no quality differences, though they do not establish equivalence. Overall, conditional loading is most beneficial when large portions of a skill are not needed on every turn.
信頼だけでは不十分: Agentic RL におけるポリシーに基づく自己蒸留のキャリブレーションに影響を与える
ポリシー上の自己蒸留 (OPSD) は、言語エージェントに、特権を持つ自己教師からポリシー自身の軌道に関してトークンレベルの緻密な監督を与えます。既存の方法では、この監督は主に教師の信頼によって割り当てられますが、トークンの強調が現在の政策目標をサポートするかどうかは信頼によって明らかにされません。私たちはこれを信頼と効用の不一致と呼び、自己蒸留のための影響校正 (ICSD) を導入します。 ICSD は、教師ありトークンごとに、教師主導の出力摂動に対する重要度加重 RL サロゲートの寄与の一次応答を測定します。バッチ適応キャリブレーションは、各アクション ターン内で元の補助損失質量を維持しながら、この非定常信号を制限された割り当て重みに変換します。これらの分離された重量は蒸留損失にのみ影響し、追加のモデル パスは必要ありません。 ALFWorld、WebShop、および Search-QA 全体で、ICSD は、1.5B ~ 7B の 2 つのモデル ファミリにわたって、グループ相対ポリシー最適化 (GRPO) およびグループ内グループ ポリシー最適化 (GiGPO) の下で信頼のみの割り当てよりも一致するすべての集計メトリックを改善します。 7B では、ALFWorld での成功率が 96.1%、WebShop スコアが 93.1 に達します。凍結バッチ分析により、ICSD は、目的に反するトークンに割り当てられた教師サポート質量が 60.1% から 37.8% に減少し、RL 勾配とのコサイン互換性が 0.192 上昇することが示されています。コンパニオン リポジトリは、https://github.com/lanqz7766/Influence-Calibration-for-On-Policy-Self-Distillation-in-Agentic-RL で入手できます。
原文 (English)
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate this supervision mainly by teacher trust, but trust does not reveal whether emphasizing a token supports the current policy objective. We call this the trust-utility mismatch and introduce Influence Calibration for Self-Distillation (ICSD). For each supervised token, ICSD measures the first-order response of its importance-weighted RL surrogate contribution to a teacher-directed output perturbation. Batch-adaptive calibration converts this non-stationary signal into a bounded allocation weight while preserving the original auxiliary-loss mass within each action turn. These detached weights affect only the distillation loss and require no additional model pass. Across ALFWorld, WebShop, and Search-QA, ICSD improves all matched aggregate metrics over trust-only allocation under Group Relative Policy Optimization (GRPO) and Group-in-Group Policy Optimization (GiGPO), across two model families spanning 1.5B to 7B. At 7B, it reaches 96.1% ALFWorld success and a WebShop score of 93.1. Frozen-batch analyses show that ICSD reduces teacher-supported mass assigned to objective-opposed tokens from 60.1% to 37.8% and raises cosine compatibility with the RL gradient by 0.192. A companion repository is avail- able at https://github.com/lanqz7766/Influence-Calibration-for-On-Policy-Self-Distillation-in-Agentic-RL.
RETRACE: Resilience-Guided Trait-Conditioned Craving Estimation from Wearable Physiology in Opioid Use Disorder
Detecting opioid craving from wearable physiological signals is critical yet difficult, with the potential to support proactive interventio…
T-LLM コンパイラ: 信頼できる LLM ベースのコード最適化および検証フレームワーク
大規模言語モデル (LLM) の最近の進歩により、高レベルのコード変換をコード最適化の分野に適用する機会が開かれ、それ以来、LLM が実行する最も基本的なタスクの 1 つとして浮上しています。ただし、現時点では、コードの複雑さと変換の正しさを独立して検証できないため、LLM は広範なコード最適化タスクを適用するのに苦労しています。このペーパーでは、高レベルの LLM コード変換、従来のコンパイラ、および検証ツールを含む共同作業によるコンパイラ テクノロジの進歩を提案する、Trusted LLM (T-LLM) コンパイラについて説明します。実験結果から、一連の PolyBench/C ベンチマークでテストすると、コードの正確性が大幅に向上することがわかりました。私たちのアプローチは、修正措置を可能にする検証戦略による反復的なコード最適化の取り組みを促進します。このアプローチにより、T-LLM コンパイラーは、PolyBench/C ベンチマークで最大 83.3% のコード最適化精度と最大 16.1\% の高速化を達成し、変換されたコードは標準ベースラインと比較して平均 26.7% の高速化に達します。さらに、プロジェクトのソース コードをオープンソース コミュニティにリリースします。
原文 (English)
T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework
Recent advances in Large Language Models (LLMs) have opened opportunities to apply high-level code transformations to the field of code optimization, and it has since emerged as one of the most fundamental tasks for LLMs to perform; however, at present, LLMs struggle to apply wide-ranging code optimization tasks due to both the complexity of the code and the inability to independently verify the correctness of the transformations. In this paper, we present the Trusted LLM (T-LLM) Compiler, which proposes an advancement in compiler technology through a collaborative effort involving high-level LLM code transformations, traditional compilers, and verification tools. Experimental results reveal that it can significantly improve code correctness when tested on a set of PolyBench/C benchmarks. Our approach facilitates iterative code optimization efforts with verification strategies that enable corrective actions. Through this approach, T-LLM Compiler achieves code optimization accuracy of up to 83.3% and a speedup of up to 16.1\% on the PolyBench/C benchmarks, with the transformed code reaching an average of 26.7% speedup wrt standard baselines. Additionally, we release the project's source code to the open-source community.
オンデマンドの都市航空モビリティ ネットワーク設計のための、デマンド駆動型の Vertiport 立地と離散イベント フリート シミュレーション
この論文では、ベルティポートの配置、フリートシミュレーション、ドアツードアの移動時間の実現可能性をリンクする、オンデマンド都市航空モビリティ (UAM) ネットワーク設計のための需要主導型フレームワークを紹介します。需要は通勤者および乗客の活動データから推定され、空間的な移動終点に変換され、K 平均法を使用してクラスター化されて、めまいポートの候補位置が生成されます。候補ネットワークは、範囲と最小ステーション間隔の制約を使用して選別され、複数車両の配車、回送台の再配置、バッテリー交換、およびサービスの規則性をモデル化する離散イベント シミュレーションで評価されます。飛行時間とエネルギー消費量は、ポイントマス eVTOL パフォーマンス モデルを使用して計算されます。ロサンゼルス都市圏のケーススタディでは、好ましい設計は、低需要時の 4 ステーションと 4 台の eVTOL から、テストされた最高の需要レベルでの 16 ステーションと 12 台の eVTOL まで拡張されています。結果は、フリートの規模が大きくなると完了時間と車両到着の規則性が向上しますが、行き止まり便は排除されず、空間的な需要の不均衡が依然として運用上の負担であることを示しています。さらに、移動時間の節約分析により、UAM は、飛行時間を考慮した上で十分な非飛行時間が残る、長時間または混雑の多い旅行に対して最も防御的であることが示唆されています。
原文 (English)
Demand-Driven Vertiport Siting and Discrete-Event Fleet Simulation for On-Demand Urban Air Mobility Network Design
This paper presents a demand-driven framework for on-demand Urban Air Mobility (UAM) network design that links vertiport siting, fleet simulation, and door-to-door travel-time feasibility. Demand is estimated from commuter and passenger activity data, converted into spatial trip-end points, and clustered using K-means to generate candidate vertiport locations. Candidate networks are screened using range and minimum station-spacing constraints, then evaluated with a discrete-event simulation that models multi-vehicle dispatch, deadhead relocation, battery swaps, and service regularity. Flight time and energy consumption are computed using a point-mass eVTOL performance model. In a Greater Los Angeles case study, the preferred design expands from four stations and four eVTOLs at low demand to sixteen stations and twelve eVTOLs at the highest tested demand level. Results show that larger fleets improve completion time and vehicle-arrival regularity but do not eliminate deadhead flights, indicating that spatial demand imbalance remains an operational burden. The travel-time savings analysis further suggests that UAM is most defensible for longer or congestion-heavy trips where sufficient non-flight time remains after accounting for flight time.
ツールの結果にはプレーンテキストよりも権限がありますか? Claude Opus 5 による合成割り当てタスクにおける虚偽申請採用に関する 3 つの前向き研究
言語モデル システムは、書き込み先でもあるストアから読み取りを行うことが増えているため、以前に書き込まれただけのクレームが取得されたように戻ってくる可能性があります。サポートされていない割り当てを含むメッセージ パッケージが、合成検索タスクでモデルが与える応答を変更するかどうかをテストしました。クロード オーパス 5 は、名前付きアイテムのカラー コードを選択するか、棄権しました。探索的な 4 群研究では、ターゲット クレームがない場合の偽コードの採用は 0/24、前のアシスタント アサーションでターゲットが指定された場合のスコアリング可能なトライアルは 0/22、ツールの結果レコードでターゲットが指定された場合は 14/24、その結果が未チェックとマークする 10 フィールドのメタデータ ラッパーを使用した場合は 15/24 でした。ツール結果部門は、11/12 のサポートされているトライアルと 14/24 のサポートされていないトライアルでレコードのコードを選択し、植え付けられたトークンの実質的な異質性を残しながら、固定の出力トークンの偏りを排除しました。文書に事前登録された複製は、ツールの結果とアシスタントの主張のギャップ、7/24 対 0/24、片側フィッシャー正確 p = 0.0047 を再現しました。それにも関わらず、ツールの結果率は 4 日間隔で実行した場合、14/24 から 7/24 に低下しました。 2 番目の事前登録された調査では、以前の比較にライブ テキスト コントロールが与えられました。両方のレコードが事前に発表され、同じ最終ユーザー ターンに配置され、リンクされたツールの結果とその後のインライン JSON の間でターゲット バインディングが交換されました。インライン テキストは、60/60 試験での偽コードの採用には十分でした。ツールの結果の条件は 57/60 を生成したため、登録された結果優先の優位性基準は失敗しました (p = 1)。この結果は、ツールの結果が影響を及ぼさないことを示していません。これは、ネイティブ ツールと結果の配置が必要ではなかったこと、およびこの実験では、発表されたインライン テキストよりも結果パッケージの動作の重みが大きくならなかったことを示しています。調査結果は、1 つの API を通じてアクセスされる、1 つの合成タスク テンプレート上の単一のモデルに関するものです。
原文 (English)
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5
Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retrieved. We tested whether the message package carrying an unsupported assignment changes which answer a model gives in a synthetic lookup task. Claude Opus 5 selected a color code for a named item or abstained. In an exploratory four-arm study, false-code adoption was 0/24 with no target claim, 0/22 scorable trials when a prior assistant assertion named the target, 14/24 when a tool-result record named it, and 15/24 when that result used a ten-field metadata wrapper that marked it unchecked. The tool-result arm selected the record's code in 11/12 supported trials and 14/24 unsupported trials, ruling out a fixed output-token bias while leaving substantial planted-token heterogeneity. A document-preregistered replication reproduced the tool-result versus assistant-assertion gap, 7/24 against 0/24, one-sided Fisher exact p = 0.0047. The tool-result rate nevertheless fell from 14/24 to 7/24 across runs made four days apart. A second preregistered study gave the earlier comparison a live text control: both records were announced in advance and placed in the same final user turn, then target binding was swapped between the linked tool result and later inline JSON. Inline text was sufficient for false-code adoption in 60/60 trials; the tool-result condition produced 57/60, so the registered result-first superiority criterion failed, p = 1. The result does not show that tool results have no effect. It shows that native tool-result placement was not necessary and that this experiment did not find greater behavioral weight for the result package than for announced inline text. The findings concern a single model on one synthetic task template, accessed through one API.
S2-MoE: エッジ デバイス上で複数の専門家が混在する場合の効率的な自己投機的デコーディングの有効化
エッジ デバイスでの推論用の大規模言語モデル (LLM) の展開は、メモリと帯域幅の厳しい制約により困難です。推論効率を向上させるために投機的デコードと混合エキスパート (MoE) が提案されていますが、これらを単純に組み合わせると、多くの場合、過剰な検証オーバーヘッドと不十分なエキスパートの再利用が発生し、メモリに制約されたエッジ設定での有効性が制限されます。この研究では、エッジ デバイス上の MoE 推論のための効率的な自己投機的デコード フレームワークである S2-MoE を提案します。 S2-MoE は、ルーティングを意識した適応型投機的拡張により冗長な検証を削減し、再利用を意識したエキスパート ゲーティングにより検証効率を向上させ、共有コンテキストを介してドラフトとターゲットの実行を調整します。 llama.cpp に実装された S2-MoE は、エッジ デバイス上のさまざまな MoE モデルおよびデータセットにわたって、標準の自己回帰デコーディングと比較して最大 5.3 倍 (平均約 2.0 倍) の高速化を実現します。コードは https://github.com/angerybob/S2-MoE で入手できます。
原文 (English)
S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture-of-Experts (MoE) have been proposed to improve inference efficiency, naively combining them often incurs excessive verification overhead and poor expert reuse, limiting their effectiveness in memory-bound edge settings. In this work, we propose S2-MoE, an efficient self-speculative decoding framework for MoE inference on edge devices. S2-MoE reduces redundant verification through routing-aware adaptive speculative expansion, improves verification efficiency with reuse-aware expert gating, and aligns draft and target execution via shared context. Implemented in llama.cpp, S2-MoE achieves up to 5.3x speedup (about 2.0x on average) over standard autoregressive de?coding across diverse MoE models and datasets on edge devices.Code is available at https://github.com/angerybob/S2-MoE.
認められるものではなく収集される: 注意が潜在変数を言語化可能な形にどのようにもたらすか
言語モデルは、レポート可能な形式で潜在的な量を保持しており、タスクで柔軟に再利用する必要がある場合には、より多くの量がその形式で存在します。表現をその形式に入らせるものはオープンであり、ワークスペースという言葉は入場の物語、つまり何が入るかを決定するゲートを誘います。ヤコビアンレンズを備えたオープンウェイトモデルでテストし、5つのアームが同じコンテキストを共有するベンチマーク上でテストすると、1つを予測するゲートは見つかりませんでした。要求により、与えられた値に演算子を適用して生成されるものを超えて、コンセプトのレンズの可視性が高まります。主要チェックポイントでのパーセンタイル ランクでは +0.050 [+0.045, +0.057]、測定した 4 つすべてでプラスですが、そのアームは天井で応答し、精度が一致したコントラストはその読み出し値の下でより強くなります。同時に、1 つの共有線形マップが、コントロールを含むすべてのアームからの変数を、選択補正された下限の 6.4 ~ 9.0 倍でデコードします。クエリされた位置で後で読み取り可能な形式を生成するのは、中深度ウィンドウ内での注意媒介の収集です。パッチ深度を読み出し深度から分離すると、そこにあるトランスポートは、非飽和読み出しの下でより浅い場所よりも少なくとも 17 倍高くなります。テスト済みの MLP 出力はその中にプラスに寄与しません。飽和パーセンタイル ランクの下では、同じグリッドはウィンドウを局在化させません。これは、その測定に関する事実です。変数を何も必要としないアームは集中が 7 分の 1 であるため、ウィンドウは需要に応じて異なります。このウィンドウには 2 つの測定エッジがあり、下に生存失敗、上に破壊があり、別のファミリーの 64 層ハイブリッドと 62 層の高密度モデルでは同じ部分的な深さに位置します。何も輸送しない通路からのルートではなく、変数が設置されている場所を特定して読み取ります。しかし、測定値は使用量の校正された尺度ではありません。3 つのコンポーネントは、測定値を相互に 12% 以内に移動させますが、答えに対して行う処理は 7.4 倍異なります。
原文 (English)
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open-weight models with Jacobian lenses, over a benchmark whose five arms share an identical context, we find no gate where it predicts one. Demand raises a concept's lens visibility beyond what applying an operator to a supplied value produces: +0.050 [+0.045, +0.057] in percentile rank on our primary checkpoint, positive on all four we measure, though that arm answers at ceiling and the accuracymatched contrast is stronger under that readout. At the same time one shared linear map decodes the variable from every arm, the control included, at 6.4-9.0x its selection-corrected floor. What produces the later readable form at the queried position is attention-mediated gathering inside a mid-depth window: separating patch depth from readout depth puts transport there at least 17x above anywhere shallower under non-saturating readouts, with no tested MLP output contributing positively inside it. Under the saturating percentile rank the same grid does not localise the window, which is a fact about that measure. An arm that needs the variable for nothing concentrates sevenfold less, so the window is demand-specific. That window has two measured edges, a survival failure below and destruction above, and it falls at the same fractional depth in a 64-layer hybrid and a 62-layer dense model from another family. We localise where the variable is installed and read, not the route from the passage, which transports nothing. But the readout is not a calibrated measure of use: three components move it to within 12% of one another and differ 7.4x in what they do to the answer.
継続を意識したポリシー学習を備えた LLM ベースの階層型協調制御
システム相互作用のモデル化が難しく、操作情報が異質であり、低レベルのアクションが厳しい制約を満たす必要がある場合、複雑なエンジニアリング システムで複数の相互作用するユニットを調整することは困難です。我々は、LLM が異種操作コンテキストに基づいて相互作用するユニットを調整し、タスク固有のコントローラーまたはオプティマイザーが実行可能で制約を意識したアクションを生成する、LLM ベースの階層フレームワークを提案します。さらに、継続認識 GRPO を導入して、後続の制御間隔にわたる調整決定の結果を捕捉します。この方法では、決定をその即時の結果だけで判断するのではなく、現在のポリシーの下でシステムがその後どのように進化するかも評価します。トレーニングには簡素化されたシステム モデルを、評価にはより現実的なシミュレーターを使用して、マルチランプ交通制御と仮想発電所 (VPP) のエネルギー管理に関するフレームワークを検証します。両方のタスクにわたって、提案された方法は、直接的なタスク固有の制御と最適化、エンドツーエンドの強化学習、ルールベースおよび RL ベースの階層調整、およびプロンプトのみの LLM コーディネーターよりも一貫して優れており、異種コンテキスト推論、階層実行、および継続を意識したポリシー学習の価値を示しています。
原文 (English)
LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning
Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, operational information is heterogeneous, and low-level actions must satisfy strict constraints. We propose an LLM-based hierarchical framework in which the LLM coordinates interacting units based on heterogeneous operational context, while task-specific controllers or optimizers generate executable and constraint-aware actions. We further introduce Continuation-Aware GRPO to capture the consequences of coordination decisions over subsequent control intervals. Rather than judging a decision only by its immediate outcome, the method also evaluates how the system evolves afterward under the current policy. We validate the framework on multi-ramp traffic control and virtual power plant (VPP) energy management, using simplified system models for training and more realistic simulators for evaluation. Across both tasks, the proposed method consistently outperforms direct task-specific control and optimization, end-to-end reinforcement learning, rule-based and RL-based hierarchical coordination, and prompting-only LLM coordinators, demonstrating the value of heterogeneous-context reasoning, hierarchical execution, and continuation-aware policy learning.
SCOPE: Score-Isolated Agentic Optimization for Video World Models
Video world models are increasingly used as simulators for planning and embodied decision making, yet improving them at inference time intr…
Andy: 厳密な証明と自律的な研究のための数学エージェント
アンディは、提出された問題を解決して検証し、新しい研究問題を定式化し、厳密な証明を構築する自律的な数学研究エージェントです。証明の生成と正しさの評価を分離し、知識の取得、目標を絞った改訂、多段階の検証をサポートします。このペーパーでは、自己トリガーによる衝動的なコンセンサスに関する公開された結果を開始点として使用して、ワークフローを説明します。 Andy は、スイッチング通信トポロジを使用した遅延異種ネットワーク向けのグローバルな指数関数的なリーダー/フォロワー同期問題を定式化します。提案されたハイブリッド制御は、自己トリガー型インパルスと実行遅延および回復フェーズの連続フィードバックを組み合わせます。各遅延インパルスの後、このフィードバックによって回復ウィンドウ中に遅延エラー チャネルがキャンセルされます。グローバルな指数同期のための十分な条件が確立され、サンプリング シーケンスとインパルス シーケンスの両方で Zeno 動作が除外されます。数値例で結果を確認します。この事例は、既存の結果から学び、意味のある研究課題を定式化し、厳密な証明を開発して検証するアンディの能力を実証しています。
原文 (English)
Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research
Andy is an autonomous mathematical research agent that solves and verifies submitted problems, formulates new research problems, and constructs rigorous proofs. It separates proof generation from correctness evaluation and supports knowledge acquisition, targeted revision, and multistage verification. This paper illustrates the workflow using a published result on self-triggered impulsive consensus as a starting point. Andy formulates a global exponential leader-follower synchronization problem for delayed heterogeneous networks with switching communication topologies. The proposed hybrid control combines self-triggered impulses with execution delay and recovery-phase continuous feedback. After each delayed impulse, this feedback cancels the delayed error channel during a recovery window. Sufficient conditions for global exponential synchronization are established, and Zeno behavior is excluded for both the sampling and impulse sequences. A numerical example confirms the result. This case demonstrates Andy's ability to learn from existing results, formulate meaningful research problems, and develop and verify rigorous proofs.
TAHB: テキスト属性のハイパーグラフ学習の包括的なベンチマーク
ハイパーグラフは、ペアごとの相互作用を超えた高次のグループごとの関係を効果的にモデル化し、事前学習済み言語モデル (PLM) と大規模言語モデル (LLM) は、テキスト属性からの豊富な意味理解を提供します。ただし、テキスト属性のハイパーグラフ ベンチマークが公開されていないため、言語モデルとハイパーグラフ学習を組み合わせる研究は依然として限られています。この制限に対処するために、ハイパーグラフ構造と生のテキスト属性を統合する最初の公開ベンチマークである TAHB (テキスト属性ハイパーグラフ ベンチマーク) を紹介します。 TAHB には、電子商取引、学術、映画、政治ネットワークの 4 つのドメインからの 10 の実世界データセットが含まれており、テキストを意識したハイパーグラフ表現学習の体系的な評価を可能にします。実験結果は、TAHB が現実世界のハイパーグラフの主要な構造特性を保存し、既存のベンチマークで観察されたパフォーマンス傾向を一貫して再現することを示しています。さらに、LLM-as-Enhancer 設定と LLM-as-Predictor 設定の両方での実験では、LLM で強化されたテキスト セマンティクスがハイパーグラフの学習パフォーマンスを向上させ、構造情報とテキスト情報が共同して LLM ベースの予測に最適な設定を提供することを示しています。私たちのベンチマークは、ハイパーグラフ学習と言語モデルの交差点における将来の研究のための基盤を提供します。
原文 (English)
TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning
Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while pretrained language models (PLMs) and large language models (LLMs) provide rich semantic understanding from textual attributes. However, research on combining language models with hypergraph learning remains limited due to the lack of public text-attributed hypergraph benchmarks. To address this limitation, we present TAHB (Text-Attributed Hypergraph Benchmark), the first public benchmark integrating hypergraph structures and raw textual attributes. TAHB contains 10 real-world datasets from four domains - e-commerce, academia, movies, and politics networks - enabling systematic evaluation of text-aware hypergraph representation learning. Experimental results show that TAHB preserves key structural properties of real-world hypergraphs and consistently reproduces performance tendencies observed in existing benchmarks. Furthermore, experiments under both LLM-as-Enhancer and LLM-as-Predictor settings demonstrate that LLM-enhanced textual semantics improve hypergraph learning performance, while structural and textual information jointly provide the best setting for LLM-based prediction. Our benchmark provides a foundation for future research at the intersection of hypergraph learning and language models.
GraphLoom: マルチモーダル KG-RAG 向けの信頼性調整されたグラフ証拠ルーティング
マルチモーダル検索拡張生成 (RAG) システムは、長い非構造化コンテキストや積極的に拡張された証拠グラフに依存することが多く、ノイズの多い証拠が導入され、マルチホップ推論が弱まり、サポートされていない生成が増加する可能性があります。コンパクトで忠実な証拠ルーティングのための信頼性が調整されたマルチモーダルナレッジグラフ RAG フレームワークである GraphLoom を紹介します。質問とそれに関連するマルチモーダル入力が与えられると、GraphLoom は、根拠のあるシーンの説明、抽出されたリレーショナル トリプル、および外部の常識知識からインスタンス レベルのマルチモーダル知識グラフを構築します。取得したすべての証拠をジェネレーターに注入する代わりに、GraphLoom は、制限された拡張を使用して信頼性を意識したサブグラフの取得を実行し、階層型グラフ メモリ スロットと凍結された言語モデルの共同グラフ シーケンス アテンションを通じて有用性の高い証拠を選択的にルーティングします。複雑な推論設定における堅牢性を向上させるために、GraphLoom はインターリーブ検索と予算付き補正検索をさらに組み合わせ、ノイズの多い検索条件下での適応的なマルチホップ証拠の改良を可能にします。私たちは、ノイズの多い外部知識の検索を近似する大規模なディストラクタ証拠プールを含む、ScienceQA、MultiModalQA、および OK-VQA で GraphLoom を評価します。実験結果では、強力なマルチモーダル RAG、グラフ検索、およびオープンソースのビジョン言語ベースラインと比較して、回答の品質と証拠の忠実性が一貫して向上しており、MultiModalQA での検索品質の向上とノイズの多い証拠プール下での安定したパフォーマンスが示されています。 MiniCheck ベースの検証、人間による評価、レイテンシー プロファイリングを使用した追加の分析により、信頼性を調整したグラフ証拠ルーティングが、ロングコンテキストのマルチモーダル証拠注入に代わる効果的な代替手段となることが示されました。
原文 (English)
GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG
Multimodal retrieval-augmented generation (RAG) systems often rely on long unstructured contexts or aggressively expanded evidence graphs, which can introduce noisy evidence, weaken multi-hop reasoning, and increase unsupported generation. We present GraphLoom, a reliability-calibrated multimodal knowledge-graph RAG framework for compact and faithful evidence routing. Given a question and its associated multimodal input, GraphLoom constructs an instance-level multimodal knowledge graph from grounded scene descriptions, extracted relational triples, and external commonsense knowledge. Instead of injecting all retrieved evidence into the generator, GraphLoom performs reliability-aware subgraph retrieval with bounded expansion and selectively routes high-utility evidence through hierarchical graph memory slots and joint graph-sequence attention in a frozen language model. To improve robustness in complex reasoning settings, GraphLoom further combines interleaved retrieval with budgeted corrective retrieval, enabling adaptive multi-hop evidence refinement under noisy retrieval conditions. We evaluate GraphLoom on ScienceQA, MultiModalQA, and OK-VQA, including large distractor evidence pools that approximate noisy external knowledge retrieval. Experimental results show consistent gains in answer quality and evidence faithfulness over strong multimodal RAG, graph-retrieval, and open-source vision-language baselines, with improved retrieval quality on MultiModalQA and stable performance under noisy evidence pools. Additional analyses using MiniCheck-based verification, human evaluation, and latency profiling show that reliability-calibrated graph evidence routing provides an effective alternative to long-context multimodal evidence injection.
LongDocBench: 長いドキュメントにおける目次階層とコンテキスト関係の回復のベンチマーク
ビジュアルドキュメントを機械可読表現に解析することは、ドキュメントインテリジェンスの基礎です。既存のベンチマークは、ページレベルの要素認識、読み取り順序、数式認識、およびテーブル構造に焦点を当てています。ただし、長い文書の場合は、文書レベルの構造の回復も必要です。これには、ページをまたがる目次 (TOC) 階層の再構築や、表や図からそのキャプション、メモ、ソースへの入力されたリンク (多くの場合 1 対多の形式) の識別が含まれます。これらの構造は部分的にしかカバーされていないか、より広範な解析プロトコルに組み込まれているため、既存のベンチマークでは 2 つの重要な文書レベルのタスク、\emph{目次階層の回復} と \emph{文脈関係の回復} を直接評価できません。これら 2 つのタスクのベンチマークを行うために、\textsc{LongDocBench} を導入します。この \textsc{LongDocBench} は、85 の実際の財務報告書、教科書、および 2,582 ページにわたる学術論文で構成され、ドキュメントごとに最大 105 ページあります。 3,937 個の見出しノード (平均ノード深さ 3.55、最大深さ 9) に対して人が検証した注釈と、2,680 個の表および図オブジェクトにわたって注釈が付けられた 3,258 個のコンテキスト関係を提供します。さらに、これらの構造の下流での有用性と回収可能性の両方を評価します。長い文書の質問応答実験では、人間が検証した目次階層と文脈上の関係が推論を改善し、それらの組み合わせにより補完的な利点がもたらされることが示されています。一方、代表的なドキュメント パーサーは、ページ レベルのパフォーマンスが優れているにもかかわらず、両方の回復タスクに関して依然として制限があります。さらなる進歩をサポートするために、\textsc{LongDocBench} とその評価プロトコル、および長いドキュメントのドキュメント レベルの構造回復を進めるための再現可能なテストベッドを一般公開します。
原文 (English)
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents
Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-level element recognition, reading order, formula recognition, and table structure. Long documents, however, also require document-level structure recovery. This includes reconstructing cross-page table-of-contents (TOC) hierarchies and identifying typed links from tables and figures to their captions, notes, and sources, often in one-to-many form. Because these structures are covered only partially or subsumed within broader parsing protocols, existing benchmarks cannot directly evaluate two key document-level tasks: \emph{Table-of-Contents Hierarchy Recovery} and \emph{Contextual Relationship Recovery}. To benchmark these two tasks, we introduce \textsc{LongDocBench}, comprising 85 real-world financial reports, textbooks, and academic papers spanning 2,582 pages, with up to 105 pages per document. It provides human-verified annotations for 3,937 heading nodes (mean node depth 3.55; maximum depth 9) and 3,258 contextual relationships annotated across 2,680 table and figure objects. We further evaluate both the downstream utility and recoverability of these structures. Long-document question-answering experiments show that human-verified TOC hierarchies and contextual relationships improve reasoning, with their combination providing complementary benefits. Meanwhile, representative document parsers remain limited on both recovery tasks despite strong page-level performance. To support further progress, we publicly release \textsc{LongDocBench} and its evaluation protocol and reproducible testbed for advancing document-level structure recovery in long documents.
思考のファネル: 早期投票とロールアウト プルーニングによる効率的なテスト時間のスケーリング
大規模な推論モデルは、同じ問題に対する繰り返しのクエリにわたって多様で、場合によっては一貫性のない答えを生成するため、信頼性の高い展開にはマルチサンプル推論が前提条件となります。 k 回のロールアウトでの多数決が標準的なソリューションであり、この制度の事実上の精度目標ですが、LRM が必要とする規模では法外に高価です。 Funnel of Thoughts (FoT) は、32 軌道の投票精度を完全に維持しながらアテンション FLOP を半分にし、フルモデルの推論コストを 28.8% 削減する推論時間手法を導入します。 6 つの LRM からの 115,000 の推論軌跡全体で、非生産的な軌跡が、「待て」、「実際に」、「おそらく」などの繰り返しのためらいマーカーによって現れることが多いことがわかりました。これらの軌跡は正解に到達する可能性が低く、不釣り合いな注意フロップを消費し、最悪の場合は無回答ループに陥ります。このトレーニング不要の語彙信号に基づいて構築された FoT は、これらの病理学的パターンを捉える語彙を特定し、完了前に影響を受ける軌道を刈り込み、追加のモデル推論を行わずに、オンライン生成のアテンション FLOP を 56.1%、所要時間を 37.6% 削減します。同じ信号が、保持されたアーキテクチャやドメイン外のタスク間で戻されることなく転送されます。
原文 (English)
Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning
Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same problem, so multi-sample inference is a prerequisite for reliable deployment. Majority voting at k rollouts is the standard solution and the de facto accuracy target for this regime, but it is prohibitively expensive at the scale LRMs require. We introduce Funnel of Thoughts (FoT), an inference-time method that preserves the full 32-trajectory voted accuracy while halving its attention FLOPs, a 28.8% reduction in full-model inference cost. Across 115K reasoning trajectories from six LRMs, we find that unproductive trajectories often reveal themselves through repeated hesitation markers such as "Wait", "Actually", and "perhaps." These trajectories are less likely to reach the correct answer and consume disproportionate attention FLOPs, degenerating into no-answer loops in the worst case. Built on this training-free lexical signal, FoT identifies the vocabulary that captures these pathological patterns and prunes affected trajectories before completion, reducing online generation attention FLOPs by 56.1% and wall time by 37.6% without any additional model inference; the same signal transfers without retuning across held-out architectures and out-of-domain tasks.
Evo-Harness: 自己進化するエージェントのためのコンテキストからハーネスへのスキルのコンパイル
経験から学ぶことは、有能な自己改善型大規模言語モデル (LLM) エージェントを開発するために重要です。既存の方法は通常、反射、記憶、ルール、またはスキルを介して蓄積された軌跡から知識を抽出します。ただし、現実的な環境では、エージェントは継続的に新しいタスクに遭遇し、多くの場合、改善の機会は 1 回限りしか提供されません。これらの実行により、豊富ではあるが非常にノイズの多いコンテキストが生成され、広く役立つ教訓がタスク固有のアーティファクトと絡み合います。重要なことに、これまでの研究では、現実世界の複雑なタスクに対する有効性が検証されたり、改善の根本的な要因が分離されたりすることはほとんどありませんでした。これらのギャップに対処するために、オンライン ハーネス学習を定式化します。この学習では、一連のタスクにわたって構造化されたハーネスを継続的に更新することで、フリーズしたエージェントが改善されます。この定式化により、当社が提案する Evo-Harness を通じて主要な自己改善要素を体系的に研究することが可能になります。その中心となるのは、コンテキストからハーネスへのスキル コンパイルで、ノイズの多い単発実行を再利用可能なスキル ハーネスに抽出し、クロスドメインおよびトピック レベルで適応させることができます。ワンショット スキル コンパイルの有効性を実証するために、5 つの現実的なベンチマーク (ターミナルベンチ 2、SWE ベンチ、CL ベンチ、ベンチ、WebArena インフィニティ) にわたって評価します。私たちの広範な分析は、Evo-Harness の有効性を実証し、LLM エージェントがどのようにして効果的にオンザフライで学習できるかについての原則的な理解を提供します。私たちのコードは https://github.com/A-EVO-Lab/a-evolve/tree/release/evo-harness で入手できます。
原文 (English)
Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically extract knowledge from accumulated trajectories via reflection, memory, rules, or skills. However, agents in realistic environments continuously encounter novel tasks, often offering only a one-shot opportunity to improve. These executions yield rich but highly noisy contexts, entangling broadly useful lessons with task-specific artifacts. Critically, prior works rarely validate their effectiveness on complex real-world tasks or isolate the underlying drivers of improvement. To address these gaps, we formulate online harness learning, where a frozen agent improves by continually updating a structured harness across sequential tasks. This formulation enables a systematic study of key self-improvement factors through our proposed Evo-Harness. At its core, context-to-harness skill compilation distills noisy, single-shot executions into reusable skill harnesses for cross-domain and topic-level adaptation. To demonstrate the efficacy of one-shot skill compilation, we evaluate across five realistic benchmarks (TerminalBench2, SWE-bench, CL-Bench, -bench, WebArena-Infinity). Our extensive analysis demonstrates the effectiveness of Evo-Harness and provides a principled understanding of how LLM agents can effectively learn on the fly. Our code is available at https://github.com/A-EVO-Lab/a-evolve/tree/release/evo-harness.
しきい値を超えて: コールド チェーン IoT システム向けの品質を意識した意思決定インテリジェンス フレームワーク
コールド チェーン ロジスティクスは技術的に進歩していますが、導入されているシステムのほとんどは、意思決定エージェントではなく、事後対応型のモニターのままです。しきい値によってアラートがトリガーされますが、違反を累積的な製品劣化と関連付けたり、劣化シグナルをロジスティクス上の決定に変換したりするものはありません。私たちは、3 つの機能を組み合わせた品質認識意思決定インテリジェンス (QADI) フレームワークでこのギャップに対処します。構造化された品質状態表現 $S_q = [L, Q, U, R]$ -- 残りの保存期間、劣化率、推定の不確実性、運用リスク。これらはすべてフレームワーク方程式から導出され、計算可能です。物理ベースの微生物動態とデータ駆動型の補正項を組み合わせたハイブリッド品質モデリング層。そして、Microsoft Phi-4~\cite{Phi4} 上に構築された推論層は、構造化されたドメイン知識ベースに対する検索拡張生成を備えています。私たちは、8 つのコールド チェーン シナリオにわたって、しきい値モニタリング、物理のみ、物理プラスノイズ、最適化ベースの決定、ルールベースのエキスパート システムという 5 つのベースラインに対してベンチマークを行います。主なケースとして低温殺菌牛乳を使用し、モデルとは独立して公開されている酪農研究から導き出されたグラウンド トゥルースの保存期間を使用します~\cite{Singh1994, Smigic2015}。比較には、ホルム補正を伴うウィルコクソンの符号付き順位検定が使用されます。牛乳とブロッコリーのシナリオ全体で、このフレームワークは平均絶対保存期間誤差 7.2 時間 (対 30.9 時間、物理のみ、$p<0.001$)、腐敗率 14.5% (対 16.6%、物理のみおよびルールベース; p=0.08)、シナリオの 99.5% でオラクル最適の決定を達成しました。 LLM 推論コンポーネントを削除すると、最適性は 45.5% に低下します ($p<0.001$)。専門家による説明の品質は 83% ($\kappa = 0.71$) に達します。アブレーションは、ハイブリッド モデリングと LLM 推論が明確な利益に貢献することを示しますが、RAG 検索は主に説明の品質を向上させます。コード: https://bit.ly/4d6t44C。
原文 (English)
Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Systems
Cold chain logistics has advanced technologically, yet most deployed systems remain reactive monitors, not decision-making agents: thresholds trigger alerts, but nothing relates violations to cumulative product degradation or converts degradation signals into logistics decisions. We address this gap with a Quality-Aware Decision Intelligence (QADI) framework combining three capabilities: a structured quality state representation, $S_q = [L, Q, U, R]$ -- remaining shelf life, degradation rate, estimation uncertainty, and operational risk, all derived and computable from the framework equations; a hybrid quality modeling layer combining physics-based microbial kinetics with a data-driven correction term; and a reasoning layer built on Microsoft Phi-4~\cite{Phi4} with retrieval-augmented generation over a structured domain knowledge base. We benchmark against five baselines -- threshold monitoring, physics-only, physics-plus-noise, optimisation-based decisions, and a rule-based expert system -- across eight cold chain scenarios, using pasteurised milk as the primary case, with ground truth shelf-life drawn from published dairy studies~\cite{Singh1994, Smigic2015} independent of our model. Comparisons use Wilcoxon signed-rank tests with Holm correction. Across milk and broccoli scenarios, the framework attains mean absolute shelf-life error of 7.2 hours (versus 30.9 hours, physics-only; $p<0.001$), spoilage rate of 14.5% (versus 16.6%, physics-only and rule-based; p=0.08), and oracle-optimal decisions in 99.5% of scenarios. Removing the LLM reasoning component drops optimality to 45.5% ($p<0.001$). Expert-rated explanation quality reaches 83% ($\kappa = 0.71$). Ablations show hybrid modeling and LLM reasoning contribute distinct gains, while RAG retrieval mainly drives explanation quality. Code: https://bit.ly/4d6t44C.
StateM: ハーネス スケーリングによりターミナルベンチ 2.1 で 95.3% の未加工精度、または 15 ドルのフロンティア ランに到達
長期的なエージェントは、その基礎となるモデルが構成ステップを解決できる場合でも、失敗する可能性があります。変更可能な状態を追跡できなくなったり、以前の実行からのレッスンを再アクティブ化しなかったり、既知の手順をスキップしたり、途中で停止したりする可能性があります。私たちは、モデルの重みを変更せずにエージェント周りの実行システムを改善するハーネス スケーリングに賭けています。 StateM は、永続的な状態、フェーズ ローカル コンテキスト、チェックされた遷移、回復可能な Runbook、エージェントとユーザーが一緒に検査できるバージョン管理された手順プラクティスに基づいて実行を編成するエージェント ネイティブ ランタイムです。 Terminal-Bench 2.1 では、StateM は GPT-5.5 xhigh を 92.1\% に引き上げます。これに対し、リファレンスは 83.1\%、GPT-5.6 Sol Ultra は 91.9\% です。 Runbook は変更されずに GPT-5.6 に転送されます。 GPT-5.6 Sol xhigh を使用すると、StateM は 445 回のトライアルで 95.3\% の生精度に達し、89 個のタスクすべてで少なくとも 1 回は成功します。凍結プロファイルは GPT-5.6 Luna を 76.7% から 85.4% に上昇させ、84.9%% Sol xhigh 基準を上回ります。同じランタイム、Runbook 構造、黄金律を使用すると、$38 未満の適応により、DeepSeek-V4 Flash は標準タイムアウトで 82.7 から 88.1\% に上昇し、88 タスクの共通コアでは 89.1\% に上昇します。残りのレイテンシの影響を受けやすいタスクのみを拡張すると、報告された 88.8\% GPT-5.6 Sol max の結果と一致します。最終スコアの API 使用量は約 $15 であるのに対し、GPT リファレンスの場合は \$574.68 です。 DeepSeek の合計支出額は \52.22 です。 BusinessBench では、開発セットに基づいて構築されたファミリー固有の Runbook により、0.55 マクロ ポイントと 1.34 マイクロ ポイントのホールドアウト ゲインが得られます。メカニズムが一致した 2 つのファミリーは 10.04 ポイント改善しました。タスクが実行構造を共有する場合、具体的なルールは一般化されますが、制御方法は広く適用されます。 StateM は、選択された事後調査結果を永続的で実行可能な前提条件とプラクティスに変換し、学習されたコントロールをステートフル コントロールを通じて明示的かつ強制可能にします。コードは github.com/henryqin1997/statem にあります。
原文 (English)
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. We bet on harness scaling to improve the execution system around an agent without changing its model weights. We introduce StateM, an agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, recoverable runbooks, and versioned procedural practices that agents and users can inspect together. On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1\%, versus 83.1\% reference and GPT-5.6 Sol Ultra at 91.9\%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3\% raw accuracy across 445 trials and succeeds on all 89 tasks at least once. The frozen profile raises GPT-5.6 Luna from 76.7 to 85.4\%, above the 84.9\% Sol xhigh reference. Using the same runtime, runbook structure, and golden rules, less than \$38 of adaptation raises DeepSeek-V4 Flash from 82.7 to 88.1\% under standard timeouts and to 89.1\% on an 88-task common core. Extending only the remaining latency-sensitive task matches the reported 88.8\% GPT-5.6 Sol max result. Final-score API usage is about \$15 versus \$574.68 for the GPT reference; total DeepSeek expenditure is \$52.22. On BusinessBench, family-specific runbooks built on development sets yield held-out gains of 0.55 macro and 1.34 micro points; two mechanism-matched families improve by 10.04 points. Concrete rules generalize when tasks share execution structure, while the control methodology applies broadly. StateM turns selected postmortem findings into persistent, executable preconditions and practices, making learned controls explicit and enforceable through stateful controls. Code at github.com/henryqin1997/statem.
制約された観察下での検証フロンティア表現の選択
クリーンなベンチマーク設定の外に導入された AI システムは、不完全、不安定、コストがかかる、または監視の失敗によって劣化した観測に依存することがよくあります。この論文では、制約された観察の下での表現の選択、つまり生の精度が唯一の操作基準ではない場合の状態表現の選択について研究します。私たちは、バランスのとれた精度と、特徴コスト、オーバーフィット ギャップ、および検証テストの不安定性に対するペナルティを組み合わせた検証フロンティア セレクターを提案します。 3 つの scikit-learn データセット、5 つの観察レジーム、45 の一致したタスク セル、720 の候補アクション、および 405 の表現行を使用した焦点を絞った公開表形式のベンチマークでは、適応セレクターは完全なトレース特徴に対するフロンティア スコアを 0.025801 改善し、平均特徴数を 22.733 削減しました。平衡精度の差は小さく、統計的に有意ではありません。より広範なオフライン ストレス テストでは、さまざまな結果が得られます。したがって、サポートされる主張には限界があります。適応表現の選択は、一致したベンチマーク設定における制約付き観測の堅牢性と効率のフロンティアを改善できますが、トレース ベースラインを普遍的に支配するわけではありません。
原文 (English)
Validation-Frontier Representation Selection under Constrained Observation
AI systems deployed outside clean benchmark settings often rely on observations that are incomplete, unstable, costly, or degraded by monitoring failures. This paper studies representation selection under constrained observation: choosing a state representation when raw accuracy is not the only operational criterion. We propose a validation-frontier selector that combines balanced accuracy with penalties for feature cost, overfit gap, and validation-test instability. In a focused public-tabular benchmark using three scikit-learn datasets, five observation regimes, 45 matched task cells, 720 candidate actions, and 405 representation rows, the adaptive selector improves frontier score over full trace features by 0.025801 while reducing mean feature count by 22.733. Balanced-accuracy difference is small and not statistically significant. A broader offline stress test gives mixed results. The supported claim is therefore bounded: adaptive representation selection can improve a constrained-observation robustness-efficiency frontier in matched benchmark settings, but does not universally dominate trace baselines.
状態遷移としての二次政策効果: 政策シミュレーションのためのソースリンクされたベンチマーク
政策評価では、制度環境を固定したものとして扱いながら、直接的な便益と費用を見積もることがよくあります。実際には、政策によって導入されるシステムが変化します。つまり、関係者が適応し、執行能力が変化し、負担が移動し、捕獲、賭博、コンプライアンス劇場、不可逆性、および修復コストを中心に新たな均衡が形成されます。我々はこれを二次政策効果予測として形式化し、政策シミュレーションのためのソースリンクされたベンチマークを提示します。このベンチマークには、8 つのドメインにわたる 96 の名前付き公共ポリシー ケースと、実装、変更、パイロット、ブロックという 4 つのバランスの取れたアクション クラスが含まれています。各ケースには、利益、捕捉、ゲーム、負担シフト、不安定性、不確実性、不可逆性、分配リスク、および実装能力に関するソースロケーターと状態変数が含まれます。ランナーはメソッド出力を再生成し、ケース テーブルから結果を集計します。シミュレーターはエキスパート アクション ターゲットを読み取ることはありません。私たちは、再現率、精度、F1 スタイルの効率性、および選択的なトップチャネル ストレス診断を備えたプロトコル ベースの移行チャネル監査を報告するため、ユニバーサル チャネル カバレッジがフィールド検証と間違われることはありません。副作用シミュレータは、平均政策効果品質 0.945 を達成しました。これに対し、リスク登録ベースラインでは 0.838、因果関係ループ ベースラインでは 0.879 でした。その利点は、副作用の再現と総トランジション スコアリングに集中しています。それは、正確な政策と行動の選択に関する最良の構造化されたベースラインを支配するものではありません。証拠は依然としてベンチマークに基づいていますが、遷移状態の変数により政策シミュレーターが下流の制度的影響に対してより敏感になるという限定的な主張を裏付けています。
原文 (English)
Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation
Policy evaluation often estimates direct benefits and costs while treating the institutional environment as fixed. In practice, a policy changes the system it enters: actors adapt, enforcement capacity shifts, burdens move, and new equilibria form around capture, gaming, compliance theater, irreversibility, and repair costs. We formalize this as second-order policy-effect prediction and present a source-linked benchmark for policy simulation. The benchmark contains 96 named public-policy cases across eight domains and four balanced action classes: implement, modify, pilot, and block. Each case includes source locators and state variables for benefit, capture, gaming, burden shift, instability, uncertainty, irreversibility, distributional risk, and implementation capacity. The runner regenerates method outputs and aggregate results from the case table, and the simulator never reads the expert action target. We report a protocol-based transition-channel audit with recall, precision, F1-style efficiency, and selective top-channel stress diagnostics, so universal channel coverage is not mistaken for field validation. The side-effect simulator achieves mean policy-effect quality of 0.945, compared with 0.838 for the risk-register baseline and 0.879 for the causal-loop baseline. Its advantage is concentrated in side-effect recall and aggregate transition scoring; it does not dominate the best structured baselines on exact policy-action choice. The evidence remains benchmark-based, but supports a bounded claim: transition-state variables make policy simulators more sensitive to downstream institutional effects.
LLM エージェントを使用した列間制約検出による制約を意識した合成表形式データの生成
構造的に有効な合成表形式データを生成することは依然として困難です。統計的忠実度が高く、下流の有用性を備えた出力でも、意味的に意味のあるドメイン制約に違反する可能性があります。私たちは、3 つの相補的な列間制約ファミリー (方程式、線形不等式、論理依存関係) の発見と適用を研究します。当社の統合ツールベースのワークフローは、これら 3 つすべてをマシンで実行可能な仮説として表し、テーブル全体の検証、確定的診断、反例に基づく修正に共通のインターフェイスを適用します。ジェネレーターに依存しないポストプロセッサーは、変更されていない表形式ジェネレーターからの出力に対するファミリー固有の修復を調整します。精選された行動監査とエンドツーエンドの評価全体で、完全なワークフローにより、ワンショットの直接プロンプトよりも保留された違反の検出が向上します。その一方で、後処理では、保持されている適用可能な制約ごとに測定された違反がゼロになり、ほとんどのデータセットで下流のユーティリティが向上し、単変量限界が大幅に保存されます。
原文 (English)
Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents
Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical fidelity and downstream utility can still violate semantically meaningful domain constraints. We study the discovery and enforcement of three complementary inter-column constraint families---equations, linear inequalities, and logical dependencies. Our unified tool-grounded workflow represents all three as machine-executable hypotheses and applies a common interface for full-table validation, deterministic diagnosis, and counterexample-guided revision. A generator-agnostic postprocessor coordinates family-specific repairs on outputs from unchanged tabular generators. Across curated behavioral audits and end-to-end evaluations, the complete workflow improves held-out violation detection over one-shot direct prompting, while postprocessing yields zero measured violations for every retained, applicable constraint, improves downstream utility on most datasets, and largely preserves univariate marginals.
量子化エージェントの構造: コード合成エージェントのワークロードにおける VRAM の安定性と予測
LLM 推論のピーク VRAM 消費量の分析モデルは、メモリを重み付けストレージ、KV キャッシュ、およびステップ数、ツール呼び出し、およびコンテキスト拡張によってパラメータ化されたアクティベーション項に分解します。この分解は、LangGraph ベースの CUDA カーネル合成エージェント (AgentK)、4 ビット量子化ファミリー (Q4 K M)、単一の NVIDIA H100 GPU、および 1,920 の軌跡にわたる 4 つの LLM バックボーンという、厳密に範囲を絞った測定研究内で経験的に評価されます。ピークメモリ予測動作に焦点を当て、2 つの主要な観察結果を報告します。まず、閉じた形式の解析モデルは、ロードウェイト VRAM と固定アクティベーション メモリ オーバーヘッドという 2 つの経験的定数を指定すると、優れた精度を実現します。ライブ GPU 読み取り値とグラウンド トゥルース軌道パラメーターが提供される閉形式モデルは、4 つのバックボーンのうち 3 つで最もよく学習されたベースラインと一致またはそれを上回ります (テスト MAPE 2.2 ~ 4.4% 対 3.4 ~ 6.5%、p = 0.76)。例外は最小のバックボーン (Phi-4-mini) で、最小の VRAM 変動 (CV 0.3%) により、動的モデリングのパフォーマンスが単純な回帰を下回ります。第 2 に、コンパイルの成功率はバックボーン容量によって厳密に分かれており (Phi-4-mini の 5.7% から Qwen2.5-Coder-14B の 62.0%)、関数型コードの合成が利用可能なメモリではなく、組み込みの LLM 機能によって制約されたままであることを示しています。さらに、全体的なピークメモリの分散はすべてのバックボーンにわたって著しく低いため (CV 0.3 ~ 9.4%)、学習されたプロンプト特徴回帰は一定平均ベースラインと比較して統計的に有意な改善をもたらしません。したがって、高度に量子化され、重みが支配的な領域で複雑な予測 VRAM モデルを展開する正当な理由は見つかりません。レプリケーションをサポートするために、評価済みコーパスと匿名化されたフレームワークをリリースします。
原文 (English)
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads
Analytical models of peak VRAM consumption for LLM inference decompose memory into weight-storage, KV-cache, and activation terms parameterized by step count, tool invocations, and context expansion. We evaluate this decomposition empirically within a strictly scoped measurement study: a LangGraph-based CUDA-kernel-synthesis agent (AgentK), a 4-bit quantization family (Q4 K M), a single NVIDIA H100 GPU, and four LLM backbones across 1,920 trajectories. Focusing on peak-memory forecasting behavior, we report two primary observations. First, closed-form analytical models achieve competitive accuracy when provided with two empirical constants: loaded-weight VRAM and a fixed activation-memory overhead. Supplied with live GPU readings and ground-truth trajectory parameters, the closed-form model matches or outperforms the best learned baseline on three of the four backbones (test MAPE 2.2-4.4% vs. 3.4-6.5%, p = 0.76). The exception is the smallest backbone (Phi-4-mini), where minimal VRAM variance (CV 0.3%) causes dynamic modeling to underperform simple regression. Second, compile success strictly bifurcates by backbone capacity (from 5.7% for Phi-4-mini to 62.0% for Qwen2.5-Coder-14B), demonstrating that functional code synthesis remains constrained by intrinsic LLM capabilities rather than available memory. Furthermore, because overall peak-memory variance is remarkably low across all backbones (CV 0.3-9.4%), learned prompt-feature regression offers statistically insignificant improvements over a constant-mean baseline. Consequently, we find no justification for deploying complex predictive VRAM models in highly quantized, weight-dominated regimes. We release the evaluated corpus and anonymized framework to support replication.
ガバナンス介入下でのプラットフォームの適応: アクターのベストレスポンスモデリングと外部の公的事例ベンチマーク
デジタル プラットフォームは、ランキング、収益化のしきい値、モデレーション基準、検証システム、開示要件、異議申し立てプロセス、アクセス ポリシーなどの変更ルールによって管理されます。こうした介入が受動的に吸収されることはほとんどありません。クリエイター、販売者、広告主、モデレーター、ユーザー、開発者、戦略的運営者は、新しい報酬面に適応します。この論文は、適応型マルチアクター情報システムの移行としてのガバナンス介入を評価するためのプラットフォーム適応モデルを開発します。このモデルは、アクターの最適な対応、戦略的なゲームの機会、節度の負担、ユーザーインセンティブの動き、執行の対応、外部性の形成、および下流のプラットフォームの安定性を表します。私たちは、メディア収益化、ランキング システム、検証、配信プラットフォーム、マーケットプレイス、アプリ ストア、コミュニティ プラットフォーム、クリエイター エコシステムをカバーする 72 の外部パブリック プラットフォーム ガバナンス ケースに基づいてモデルを評価します。 9 つのメソッドと 648 のメソッドケース評価にわたって、完全なプラットフォーム適応シミュレーターは平均適応品質 0.836338 を達成しました。これに対し、リスク登録ベースラインでは 0.669731、因果ループ分析では 0.589457、一般的なガバナンス批判では 0.492750、エンゲージメントのみの最適化では 0.369492、ベースライン ポリシー レビューの場合は 0.331965。一対の比較では、テストされたすべてのベースラインとチャネル アブレーションに対して 1.00 の勝率が示されています。この貢献は、ポリシールールを適応的なアクター応答フィールドへの介入ではなく静的制御として扱う場合にプラットフォームガバナンス評価が失敗する理由を示す情報システム理論と測定フレームワークです。
原文 (English)
Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an External Public-Case Benchmark
Digital platforms govern by changing rules: rankings, monetization thresholds, moderation standards, verification systems, disclosure requirements, appeal processes, and access policies. These interventions are rarely absorbed passively. Creators, sellers, advertisers, moderators, users, developers, and strategic operators adapt to the new reward surface. This paper develops a platform-adaptation model for evaluating governance interventions as transitions in adaptive multi-actor information systems. The model represents actor best response, strategic gaming opportunity, moderation burden, user-incentive movement, enforcement response, externality formation, and downstream platform stability. We evaluate the model on 72 external public platform-governance cases covering media monetization, ranking systems, verification, delivery platforms, marketplaces, app stores, community platforms, and creator ecosystems. Across 9 methods and 648 method-case evaluations, the full platform-adaptation simulator achieves mean adaptation quality of 0.836338, compared with 0.669731 for a risk-register baseline, 0.589457 for causal-loop analysis, 0.492750 for generic governance critique, 0.369492 for engagement-only optimization, and 0.331965 for baseline policy review. Paired comparisons show a win rate of 1.00 against all tested baselines and channel ablations. The contribution is an information-systems theory and measurement framework showing why platform governance evaluation fails when it treats policy rules as static controls rather than interventions into adaptive actor-response fields.
ReForge: ABR アルゴリズムの維持は検証済みの大規模言語モデルの編集では終わらない
1 つのネットワーク シナリオ向けの ABR アルゴリズムの設計にはエンジニアが数か月かかりますが、大規模な言語モデルではこの作業が数時間で完了し、手作業で構築された設計に匹敵するかそれを上回ります。しかし、いずれにせよ、そのデザインは誕生時に目に見える世界にのみ適合し、その後に到着する世界では失敗します。私たちは、ABR アルゴリズムが世界と歩調を合わせ、各シナリオが到着するたびに数分で再設計され、すべての変更がすでに提供されているすべてのシナリオに無害であることが証明されるかどうかを疑問に思っています。この研究では、継続的に変化するシナリオに適応する継続的ヒューリスティック学習フレームワークである ReForge を提案します。 ReForge は、ループ内で大規模言語モデル (LLM) を使用してそのルーチンを実行します。各ラウンドで、LLM は現在の設計が不足している箇所を読み取り、小さな編集を 1 つ提案し、これまでに提供されたすべてのネットワークでの再生によって決定します。具体的には、編集するのは、あらゆる決定を事前トレーニングされたポリシーの凍結されたプールの 1 つにルーティングする、あいまいなルールの 1 ページです。 LLM は測定値だけから最初のページを書き込み、その後はそれを独自に改善し続けます。各ラウンドで、現在のルールが満たしていない箇所を読み取り、小さな編集を 1 つ提案します。これまでに提供されたすべてのネットワークでの再生により、編集が成功するかどうかが決定されます。私たちは、3G、4G、5G として一度に 1 つずつ登場する 9 つの現実世界のネットワーク ファミリで ReForge を評価します。到着ごとにいくつかの編集を行うと、QoE が 1.23 から 1.74 に上昇し、最高の単一ポリシーである 1.66 を超えてオラクルの 94% に達し、ループが一度も見たことのない修復ファミリさえも 0.30 から 0.80 に上昇しました。すべてのコード、データ、実験記録はクリーンアップ後にオープンソース化されます。
原文 (English)
ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits
Designing an ABR algorithm for one network scenario takes an engineer months, and large language models now do this work in hours, matching or beating hand-built designs. But either way, the design fits only the world visible at its birth, and fails on the world that arrives after. We ask whether an ABR algorithm can keep pace with the world, redesigned in minutes as each scenario arrives, with every change proven harmless to every scenario already served. In this work, we propose ReForge, a continual heuristic learning framework that adapts to continuously changing scenarios. ReForge runs that routine with a large language model (LLM) in the loop. Each round the LLM reads where the current design falls short and proposes one small edit, and a replay over every network served so far decides. Specifically, what it edits is a single page of fuzzy rules that routes every decision to one of a frozen pool of pre-trained policies. The LLM writes the first page from measurements alone, then keeps improving it on its own. Each round it reads where the current rules fall short and proposes one small edit, and a replay over every network served so far decides whether the edit lands. We evaluate ReForge on nine real-world network families arriving one at a time as 3G, 4G, then 5G. A few edits per arrival lift mean QoE from 1.23 to 1.74, past the best single policy at 1.66 and to 94\% of an oracle, and even repair families the loop never saw, one rising from 0.30 to 0.80. All code, data, and experiment records will be open-sourced upon cleanup.
CPMpy を使用した有限領域整数制約モデルの CP/SMT/ILP/PB/SAT ソルバーへの変換
制約解決は、組み合わせ満足および最適化問題を解決するための宣言的アプローチです。ユーザーは制約と決定変数を通じて問題を指定し、汎用ソルバーを使用して解決策を見つけます。いくつかの制約解決テクノロジが存在し、特定のソルバーは特定の問題に対して適切に機能します。したがって、特定のアプリケーションに応じてさまざまなソルバーを試すと便利です。ただし、それぞれの解決パラダイムは、さまざまなタイプの制約と決定変数をサポートしています。私たちの目標は、高レベルの制約充足と最適化の問題を、CP、SMT QF-LIA、ILP、PB、(Max)SAT などの下位レベルの形式に変換することです。これにより、ユーザーが解決パラダイムごとに手動で再構築する必要がなく、特定の問題に対するさまざまな解決テクノロジーを比較することができます。論理演算と算術演算の高級言語、および CP コミュニティではグローバル制約として知られる便利な追加関数と制約を定義します。次に、高レベル モデリング言語を CP/SMT/ILP/PB および (Max)SAT ソルバーに変換するためのモジュール式フレームワークを紹介します。多くの変換は文献で部分的に説明されていますが、より小さなコンポーネントのモジュール式ウォーターフォールを通じて実装できることがわかり、下位レベルのパラダイムが上位レベルのパラダイムの変換を再利用します。繰り返し発生する 2 つの課題は、任意の部分式の否定の処理と補助変数の導入の回避です。さらに、ILP、PB、SAT ソルバーの非線形演算子の線形化には特別な注意を払っています。変換ウォーターフォールは、オープンソースの CPMpy ライブラリで実装および評価されます。私たちの結果は、制約モデルが変換を通じて大幅に変化すること、および制約の線形化に対する最適化が ILP および PB ソルバーにとって不可欠であることを示しています。
原文 (English)
Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy
Constraint solving is a declarative approach for solving combinatorial satisfaction and optimization problems. The user specifies their problem through constraints and decision variables, and a generic solver is used to find a solution. Several constraint-solving technologies exist, and certain solvers perform well on certain problems. Therefore, it is useful to try different solvers given a particular application. However, each solving paradigm supports different types of constraints and decision variables. Our goal is to translate high-level constraint satisfaction and optimization problems into any lower-level formalism, including CP, SMT QF-LIA, ILP, PB and (Max)SAT. This allows for comparing different solving technologies for a particular problem, without requiring a user to manually remodel it for each solving paradigm. We define a high-level language of logical and arithmetic operations, and useful additional functions and constraints, which are known as global constraints in the CP community. We then present a modular framework for transforming our high-level modeling language to CP/SMT/ILP/PB and (Max)SAT solvers. While many transformations are partly described in the literature, we observe that they can be implemented through a modular waterfall of smaller components, where lower-level paradigms reuse the transformations of higher-level paradigms. Two recurring challenges are handling the negation of arbitrary subexpressions and avoiding the introduction of auxiliary variables. Additionally, we take special care linearizing non-linear operators for ILP, PB and SAT-solvers. The transformation waterfall is implemented and evaluated in the open-source CPMpy library. Our results show that constraint models significantly change throughout the transformations, and that optimizations to the linearization of constraints are essential for ILP and PB solvers.
ACTS-SQL: 大規模な言語モデルを使用したエージェントおよびクリティカル指向のツリー構造 SQL の正確性
Text-to-SQL システムでは大規模言語モデル (LLM) の採用が増えていますが、現実の Text-to-SQL 推論パイプラインでは依然として SQL エラーが大きな障害となっています。既存の SQL 修正アプローチは、かなりのオーバーヘッドを伴う大規模で高品質のトレーニング データに依存するか、初期ミスに弱くエラーが伝播しやすいシングルパス エージェント ワークフローを採用するかのどちらかです。産業シナリオ向けの実用的な SQL 正確性システムを開発するために、計画に基づいたツリー構造のデバッグ プロセスとして SQL 修正を定式化する、トレーニング不要のフレームワークを紹介します。複数の修正戦略を維持し、バックトラッキングを有効にすることにより、フレームワークは反復改良中のエラーの蓄積を軽減します。さらに、実行ベースの検証ツールと条項レベルの診断ツールを統合して、戦略の枝刈りや正確なエラーの位置特定をサポートします。私たちは BIRD-Critic ベンチマークでシステムを評価し、強力な LLM バックボーンと代表的なエージェントベースのベースラインを超えて一貫した精度の向上を観察し、以前の最先端の方法と比較して 9.42% の向上を達成しました。このフレームワークは、オンライン Text-to-TLS API をサポートするために、Volcano Engine のトーチラグ サービス (TLS) にもデプロイされています。運用環境では、代表的な強力な LLM バックボーン (GPT-5) を使用して、実際のユーザー クエリの実行精度が 36.77% から 53.61% に向上します。これらの結果は、実際の展開における私たちのアプローチの有効性と安定性を示しています。
原文 (English)
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models
Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines. Existing SQL correction approaches either rely on large-scale, high-quality training data with substantial overhead, or adopt single-path agentic workflows that are brittle to early mistakes and prone to error propagation. To develop a practical SQL correctness system for industrial scenarios, we present a training-free framework that formulates SQL correction as a plan-guided, tree-structured debugging process. By maintaining multiple correction strategies and enabling backtracking, the framework mitigates error accumulation during iterative refinement. We further integrate execution-based verification and clause-level diagnostic tools to support strategy pruning and precise error localization. We evaluate the system on the BIRD-Critic benchmark and observe consistent accuracy gains over strong LLM backbones and representative agent-based baselines, achieving a 9.42% improvement over the previous state-of-the-art method. The framework is also deployed in the Torch Log Service (TLS) of Volcano Engine to support an online Text-to-TLS API. In production, it improves execution accuracy from 36.77% to 53.61% on real user queries with a representative strong LLM backbone (GPT-5). These results demonstrate the effectiveness and stability of our approach in real-world deployments.
機械知能の構成事前確率: 人工物理世界の正当性理論
機械知能は象徴的な世界を征服しましたが、物理的な世界では行き詰まっています。この停滞は構造的なものです。物理 AI はコールドスタートのデッドロックに直面しています。データがなければインテリジェンスは存在せず、導入されたインテリジェンスがなければデータも存在しません。私たちの主張は、行き詰まりは現実に存在するが、不均等に分布しており、その例外には人工の物理世界という名前が付いています。建物、産業施設、インフラストラクチャーは意図的に構成され、文書化されています。設計されたアーティファクトには、そのインスタンスに先行して構成されている読み取り可能なアーカイブが付属しています。ここでは、規範は実例よりも先に公布されるものであり、実例から平均化されるものではありません。 4件の寄稿。 (i) 4 世界オントロジーから、構成的な事前フレームワークの正当性基準を導き出します。事前抽出は、オブジェクト ドメインが意図的に構成され、読み取り可能なアーカイブを残している場合に限り正当です。この基準は、適合の方向によってテスト可能です。構成基準からの逸脱は、世界では違反であり、モデルの改訂ではありません。 (ii) 階層化の下限を確立します。4 つの構築目標が相互に互換性のないキャリアとペアになるため、そのようなフレームワークには少なくとも 4 つの層 (構文、概念、知識、インスタンス) があります。 (iii) 当社は、5 つの産業ドメインと 32 クラスの障害モード語彙にわたる展開クレームを登録します。 (iv) 私たちはフレームワークを 5 つの反証可能な予測に賭けており、その中心となる予測は公共技術記録でチェック可能です。それが失敗した場合、フレームワークは失敗します。これらの主張を裏付ける準形式的な議論 (付録 A): アーカイブのない世界におけるルール カバレッジに関するゴールド タイプの境界、閉じた概念レイヤー上の障害削減のための決定可能性の結果、および証明書にアンカーされた計算の境界定理。ここでは、大規模な言語モデルが、アーカイブとしてではなく、アーカイブの読者として名誉ある地位を占めています。 3 つの関連作品のうちの 1 つ目。仲間たちは意図的に残された質問に答えます。
原文 (English)
Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World
Machine intelligence has conquered the symbolic world but stalled at the physical one. The stall is structural: physical AI faces a cold-start deadlock -- no intelligence without data, no data without deployed intelligence. Our thesis: the deadlock is real but unevenly distributed, and the exception has a name: the artificial physical world. Buildings, industrial facilities, and infrastructure are intentionally constituted and documented: designed artifacts ship with readable archives that precede and constitute their instances; here, norms are promulgated before instances, not averaged from them. Four contributions. (i) From a four-world ontology we derive a legitimacy criterion for constitutive prior frameworks: prior extraction is legitimate if and only if the object domain is intentionally constituted and has left a readable archive; the criterion is testable through direction of fit -- deviation from a constitutive norm is a violation in the world, not a revision of the model. (ii) We establish a layering lower bound: any such framework has at least four layers -- syntax, concept, knowledge, instance -- because four construction goals pair into mutually incompatible carriers. (iii) We register deployment claims across five industrial domains and a 32-class failure-mode vocabulary. (iv) We stake the framework on five falsifiable predictions, the central one checkable on the public engineering record: if it fails, the framework fails. Semi-formal arguments back these claims (Appendix A): a Gold-type boundary on rule coverage in archiveless worlds, a decidability result for failure reduction over closed concept layers, and a boundary theorem for certificate-anchored calculi. Large language models find an honored place here -- as readers of the archive, not as the archive. First of three companion works; the companions take up the questions deliberately left open.
SkillCommit: 行動検証されたスコープ拡張によるエージェント スキルの進化
大規模言語モデル (LLM) エージェントは、過去の経験を再利用可能な手順知識に変換することで、パラメーターを更新しなくても継続的に改善できます。ただし、既存の方法では、意味上の類似性や LLM の判断に基づいて経験が統合されることが多く、表面的には関連しているが動作的には互換性のない戦略が統合され、それによってパフォーマンスが低下する可能性があります。この問題に対処するために、私たちは、経験を再利用可能なスキルの階層ライブラリに継続的に変換するオンライン スキル進化フレームワークである SkillCommit を提案します。新しいエクスペリエンスはそれぞれ、最初はインスタンス固有のパッチとして保存され、ローカル コンテキストで検証された動作が保持されます。関連するスキルが蓄積されると、SkillCommit は共通の動作メカニズムを共有するスキルをより高いレベルのスキルに抽象化します。具体的には、受信したスキルごとに、埋め込みベースの検索によって、まず関連するスキルの候補が特定されます。クロスインスタンス リプレイと LLM ベースのメカニズム チェックにより、これらのスキルがケース間で転送され、共通の基礎となるメカニズムを共有しているかどうかが判断されます。両方のチェックに合格した候補者は、より高いレベルのスキルに抽象化され、すべての構成スキルの検証された動作が保持される場合にのみコミットされます。 RuleArena、OpenExempt、および KOR-Bench の実験では、SkillCommit がさまざまなドメインにわたってエージェントのパフォーマンスを一貫して向上させることが実証されています。さらに、学習したスキルはモデルのスケールやファミリーを超えて伝達され、モデル間の経験の伝達が可能になります。
原文 (English)
SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion
Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable procedural knowledge. However, existing methods often consolidate experience based on semantic similarity or LLM judgments, which may merge superficially related but behaviorally incompatible strategies and thereby degrade performance. To address the issue, we propose SkillCommit, an online skill evolution framework that continuously transforms experience into a hierarchical library of reusable skills. Each new experience is initially preserved as an instance-specific patch, retaining the behavior validated in its local context. As related skills accumulate, SkillCommit abstracts those sharing a common behavioral mechanism into higher-level skills. Specifically, for each incoming skill, embedding-based retrieval first identifies candidate related skills. Cross-instance replay and an LLM-based mechanism check determine whether these skills transfer across cases and share a common underlying mechanism. Candidates that pass both checks are abstracted into a higher-level skill and committed only if it preserves the validated behavior of all constituent skills. Experiments on RuleArena, OpenExempt and KOR-Bench demonstrate that SkillCommit consistently improves agent performance across diverse domains. Moreover, the learned skills transfer across model scales and families, enabling cross-model experience transfer.
LongRCA ベンチ: Long-Horizon エージェント障害における責任ある役割と根本原因の診断
長期的なエージェントの実行が失敗した場合、結果レベルの評価によって失敗した結果が明らかになりますが、決定的なエラーが軌道に入った場所は明らかにされません。次に、開発者は完全な実行を検査して、責任のある役割を特定し、決定的な根本原因の最も早いステップを特定する必要があります。既存の障害属性ベンチマークは主に短いトレースに焦点を当てており、記録された数百のステップにわたる診断は十分に検討されていません。 LongRCA Bench を紹介します。これは、エラーが挿入されていない 5 つのドメインにわたる 1,140 個の失敗した軌跡で構成されています。これは、責任ある役割と最も早い決定的な根本原因ステップに対して、独立してスコア付けされた人間のラベルを提供します。中央軌道には 145 のステップが含まれており、最も強力なベースラインでもルート ステップの正確な精度は 13.2% にすぎません。さらに、セグメントサマリーから候補エラーステップを取得し、それらを以前の利用可能なハンドオフ命令まで追跡する、トレーニング不要の方法である根本原因軌跡アトリビューション(RCTA)を紹介します。同じバックボーン、ベンチマーク インスタンス、スコアリング プロトコルを使用することで、RCTA は責任のある役割の精度が 51.1%、ルートステップの精度が 24.1% に達しました。これらの結果は、責任のある役割の帰属と正確なルートステップの位置特定を、長期軌道障害診断の別のターゲットとして評価する必要性を強調しています。
原文 (English)
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing failure-attribution benchmarks largely focus on shorter traces, leaving diagnosis across hundreds of recorded steps underexplored. We introduce LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors. It provides independently scored human labels for the responsible role and earliest decisive root-cause step. The median trajectory contains 145 steps, and the strongest baseline reaches only 13.2% exact root-step accuracy. We further present Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions. Using the same backbone, benchmark instances, and scoring protocol, RCTA reaches 51.1% responsible-role accuracy and 24.1% exact root-step accuracy. These results highlight the need to evaluate responsible-role attribution and exact root-step localization as separate targets in long-trajectory failure diagnosis.
Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts
Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI).…
自動ドメインモデリングにおける標準化された評価に向けて: ベンチマークの導入
ドメイン モデリングは、ドメイン駆動設計において重要な役割を果たし、特定のドメイン内の重要なエンティティとその関係を把握します。自動化されたドメイン モデリングの進歩にもかかわらず、標準化されたベンチマークがないため、既存のアプローチの比較評価が妨げられています。このペーパーでは、このギャップに対処するために設計されたベンチマークを紹介します。このベンチマークは、Calamo、Mecella、および Snoeck の Text2UML プロジェクトによって配布されている Zenodo 上の 45 レコードの Golden UML モデルセット (Verbruggen et al.、2025) と、Chen et al. の 8 レコードのリファレンス アーカイブを組み合わせています。 (Chen et al.、2023a、b)、さまざまなレベルの複雑さと規模にわたる自動化されたドメイン モデリング アプローチの評価が可能になります。自然言語記述が与えられた場合、そのタスクは、対応するドメイン モデルを生成することです。各記述に対して、参照ドメイン モデルがグラウンド トゥルースとして提供されます。メトリックは、生成されたドメイン モデルと対応するグラウンド トゥルース モデルを比較するために使用されます。ベンチマークの有用性を実証するために、ヒューリスティックなルールベースの手法や LLM 主導の戦略など、複数の自動ドメイン モデリング アプローチを評価します。 FAIR4RS 勧告 (Chue Hon et al.、2022) に従って、ベンチマークは再利用を促進し、自動化されたドメイン モデリングに関する将来の研究をサポートする研究成果物として提供されます。
原文 (English)
Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark
Domain modeling plays an essential role in domain-driven design, capturing essential entities and their relationships within a specific domain. Despite advancements in automated domain modeling, the absence of standardized benchmarks has hindered the comparative assessment of existing approaches. This paper introduces a benchmark designed to address this gap. The benchmark combines the 45-record Golden UML Modelset (Verbruggen et al., 2025) on Zenodo, as distributed by the Text2UML project of Calamo, Mecella, and Snoeck (Calamo et al., 2025), with the 8-record reference archive of Chen et al. (Chen et al., 2023a,b), enabling the evaluation of automated domain modeling approaches across different levels of complexity and scale. Given a natural language description, the task is to generate a corresponding domain model. For each description, a reference domain model is provided as ground truth. A metric is used to compare the generated domain model with the corresponding ground-truth model. To demonstrate the utility of the benchmark, we evaluate multiple automated domain modeling approaches, including heuristic rule-based methods and LLM-driven strategies. In accordance with the FAIR4RS recommendations (Chue Hong et al., 2022), the benchmark is provided as a research artifact to encourage reuse and support future research on automated domain modeling.
異種マルチタスクのセマンティック通信のための分散型フェデレーテッド ラーニング
分散セマンティック通信 (DSC) ネットワークでの共同トレーニングは通常、分散型フェデレーテッド ラーニング (DFL) に依存します。ただし、トポロジに依存しないアグリゲーションを異種混合のマルチタスク環境に押し込むと、根本的なボトルネックが生じます。それは、ネガティブな転送とオーバーコンセンサス バイアス (OCB) を引き起こすということです。このペーパーでは、このタスク間の干渉を遮断するパーソナライズされた DSC フレームワークを紹介します。ノード レベルでは、ポリシー主導のマルチパス ルーティング メカニズムにより、タスク固有の機能が共有表現から分離され、ローカル忠実度が維持されます。ネットワーク全体に、「集約中の通信」プロトコルを導入します。タスク アフィニティを使用して列確率的コンセンサス行列を調整します。これにより、不一致のパラメータ更新を積極的にブロックしながら、システムが補完的な知識を吸収するように制限されます。収束を制限するために、統一されたリアプノフ ドリフト解析を導き出します。我々は、厳密な U 字型のトレードオフを明らかにしました。トポロジカル混合が深くなると、分散は減少しますが、構造的な OCB が増幅されます。この緊張を解決すると、最適な凝集深さの閉じた形式の式が得られます。 NYU-v2 で提案されたフレームワークを評価します。その結果、不十分な集約と過剰なトポロジー混合の間の明らかなトレードオフが明らかになりました。分析的に導き出された最適な集約の深さでは、私たちの方法は集約なしのベースラインと比較して 4.77% のグローバル相対改善を達成し、分散型 FedAvg、FedAMP、およびヒューリスティック最大集約を上回ります。さらに、Taskonomy と不完全な無線リンクに関するフレームワークを評価して、ネットワーク サイズの変動と無線リンクの信頼性の影響を調査します。
原文 (English)
Decentralized Federated Learning for Heterogeneous Multi-Task Semantic Communication
Collaborative training in distributed semantic communication (DSC) networks typically relies on decentralized federated learning (DFL). However, pushing topology-agnostic aggregation into heterogeneous, multi-task environments creates a fundamental bottleneck: it drives negative transfer and overconsensus bias (OCB). This paper introduces a personalized DSC framework that cuts off this cross-task interference. At the node level, a policy-driven multi-path routing mechanism separates task-specific features from shared representations to preserve local fidelity. Across the network, we deploy a "communicationwhile- aggregation" protocol. It calibrates a column-stochastic consensus matrix using task affinities. This limits the system to absorbing complementary knowledge while actively blocking mismatched parameter updates. To bound the convergence, we derive a unified Lyapunov drift analysis. We reveal a strict Ushaped trade-off: deeper topological mixing reduces variance but amplifies structural OCB. Resolving this tension yields a closed-form expression for the optimal aggregation depth. We evaluate the proposed framework on NYU-v2, where the results reveal a clear trade-off between insufficient aggregation and excessive topological mixing. At the analytically derived optimal aggregation depth, our method achieves a 4.77% global relative improvement over the no-aggregation baseline and outperforms decentralized FedAvg, FedAMP, and heuristic max aggregation. We further evaluate the framework on Taskonomy and imperfect wireless links to examine the effects of network-size variation and wireless-link reliability.
VibeWorlding: マルチモーダル エージェントは 3D オープンワールドをエンドツーエンドで構築できますか?
ユーザーのクエリからインタラクティブな 3D オープンワールドを構築することが重要です。ただし、既存の手法は主に理想化された単純なクエリに基づいて評価されるため、マルチモーダル エージェントがどのようにユーザーの意図を理解し、3D ツールを使用し、テキストおよび視覚的な 3D 世界情報を推論するかを体系的に分析して比較することが困難になっています。この目的を達成するために、私たちは、バイブワールド エージェントのベンチマークとトレーニングのための統合フレームワークである VibeWorlding を提案します。これは、自律的にユーザーの意図を推測し、シーン レイアウトを計画し、3D ツールを呼び出し、マルチターンのエージェントと環境の対話プロセスでマルチモーダル フィードバックを反映できるマルチモーダル エージェントです。これを達成するために、まず VWE-BENCH を構築します。これは、2,616 個の高品質 3D アセット、323 個の人間による注釈付きシード 3D ワールド、および 6,828 個の逆合成されたマルチモーダル ユーザー クエリのベンチマークであり、グラウンド トゥルースを使用した検証済みクエリと、慎重に設計されたルーブリックを使用した未検証クエリに分割されます。さらに、当社は、(1) MCP ツールとしてのアセットの取得、編集、画像レンダリングを統合するサンドボックス環境と、(2) 物理的な実現可能性と意図の履行検証を組み合わせたルーブリックベースの検証器を統合する共同マルチモーダル RL ポストトレーニング フレームワークである VibeWorlding-Gym を開発し、公正なモデル評価とスケーラブルなマルチモーダル RL 報酬サービスの両方をサポートします。私たちの実験によると、現在のフロンティア MLLM はバイブ ワールド化エージェント タスクの解決には程遠く、GPT-5.5 や Qwen3.8-Max でさえ成功率が 60% 未満に達しており、正確な 3D ワールド編集へのボトルネックを追跡しています。さらに、RL トレーニングによってこの弱点が緩和され、オープンソース MLLM がクローズドソースのフロンティアをも超えることができることがわかりました。当社の VibeWorlder-8B はフロンティア MLLM に匹敵し、当社の主力製品である VibeWorlder-30B-A3B は、評価されたすべてのモデルの中で最高の総合 Pass@1 を達成しています。
原文 (English)
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.
$D^{2}R^{2}$: 単一セル摂動予測のための規制強化を伴う離散拡散
遺伝的摂動に対する単一細胞のトランスクリプトーム応答を予測することは、機能ゲノミクスと仮想細胞モデリングの中心です。しかし、既存のアプローチは通常、個々の遺伝子応答が生成される順序をモデル化せずに、発現プロファイル全体を全体として予測します。この問題に対処するために、\textbf{$D^{2}R^{2}$} (\textbf{D}iscrete \textbf{D}iffusion with \textbf{R}egulation \textbf{R}einforcement)を導入します。これは、摂動予測を制御に基づく遺伝子ごとの漸進的生成として再定式化します。マスクされた離散拡散モデルは、発現を順序トークンとして表し、完全にマスクされたプロファイルを段階的に再構築することで、生成された遺伝子応答がマスクされたままの遺伝子応答を条件付けできるようにします。制御ポリシーモジュールは、制御細胞から推論された遺伝子制御ネットワークから生成ポリシーを初期化し、それを摂動および現在の部分的に生成された状態に適応させます。次に、グループ相対ポリシーの最適化により、最終的な摂動効果の一致を報酬として使用して、順序付けポリシーのみが洗練されます。 Norman19 と VCC-H1 全体で、$D^{2}R^{2}$ は Norman19 の 5 つの指標すべてで最高のパフォーマンスを達成し、H1 でも競争力を維持しています。ジェネレーターと生成バジェットを固定した制御アブレーションは、生物学的事前順序付けがランダム順序付けよりも改善され、不確実性に基づくヒューリスティックよりも信頼性が高いことを示しますが、生物学的事前順序付けを逆転させるとすべてのメトリクスが低下します。生物学的分析により、この洗練されたポリシーは、摂動特異的な転写因子と応答性遺伝子を促進しながら、早期に調節遺伝子を優先することがさらに示されています。これらの結果は、単一細胞摂動予測の効果的で制御可能で生物学的に解釈可能な次元として、遺伝子生成順序を確立します。
原文 (English)
$D^{2}R^{2}$: Discrete Diffusion with Regulation Reinforcement for Single-Cell Perturbation Prediction
Predicting single-cell transcriptomic responses to genetic perturbations is central to functional genomics and virtual-cell modeling. Existing approaches, however, typically predict an entire expression profile as a whole, leaving the order in which individual gene responses are generated unmodeled. To address this problem, we introduce \textbf{$D^{2}R^{2}$} (\textbf{D}iscrete \textbf{D}iffusion with \textbf{R}egulation \textbf{R}einforcement), which reformulates perturbation prediction as regulation-guided gene-wise progressive generation. A Masked Discrete Diffusion Model represents expression as ordinal tokens and reconstructs a fully masked profile step by step, allowing generated gene responses to condition those that remain masked. A Regulatory Policy Module initializes the generation policy from a gene regulatory network inferred from control cells and adapts it to the perturbation and current partially generated state. Then, group-relative policy optimization refines only the ordering policy using final perturbation-effect agreement as reward. Across Norman19 and VCC-H1, $D^{2}R^{2}$ achieves the best performance on all five metrics on Norman19 and remains competitive on H1. Controlled ablations holding the generator and generation budget fixed show that biological-prior ordering improves over random ordering and is more reliable than uncertainty-based heuristics, whereas reversing the biological-prior ordering degrades every metric. Biological analyses further show that the refined policy prioritizes regulatory genes early while promoting perturbation-specific transcription factors and responsive genes. These results establish gene generation order as an effective, controllable, and biologically interpretable dimension of single-cell perturbation prediction.
ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning
Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dy…
発散収束推論: 構造化ソリューション合成によるテスト時間計算のスケーリング
テスト時のコンピューティングは、大規模言語モデル (LLM) の推論パフォーマンスを大幅に向上させることができますが、追加のコンピューティングがいつ、どのように役立つかについては十分に理解されていません。私たちは、複数の候補解を生成する探索フェーズとそれに続く収束調整フェーズで構成される単純な 2 フェーズのプリミティブである発散収束推論 (DCR) を研究します。 3 つの主要な結果を紹介します。まず、単一の調整ステップでも正しい少数派レポートを確実に増幅できることを示します。データセット全体で、正しい探査出力が少数派である場合、つまり多数決が失敗する体制の場合、DCR は多くの場合正しい答えを回復します。 2 番目に、不一致を繰り返し分析し、追加のテスト時計算を割り当てる自己回帰調整システムである再帰 DCR を導入します。再帰的 DCR は、固定コンピューティングのベースラインよりも高い精度を達成し、AIME 2024 では 93.3%、AIME 2025 では 92.0% に達しますが、コンピューティング使用量は平均で約 27% 少なく、慎重なリソース割り当てが均一なスケーリングよりも優れていることを示しています。 3 番目に、シンプルでトレーニング不要の分散メトリクスを介して、探査出力間の不一致を分析します。分散は、不一致とテスト時間の利益の間の構造化された関係を明らかにします。DCR が有効な領域では、探査出力間の不一致が大きいほど、調整による精度の向上が大きくなります。これらの結果を総合すると、ノイズと見なされることが多い不一致を体系的に利用して、テスト時の推論を改善し、エージェント LLM システムの新たなスケーリング則を明らかにできることがわかります。
原文 (English)
Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis
Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood. We study Divergent-Convergent Reasoning (DCR), a simple two-phase primitive consisting of an exploration phase that generates multiple candidate solutions followed by a convergent reconciliation phase. We present three core results. First, we show that even a single reconciliation step can reliably amplify correct minority reports: across datasets, DCR often recovers the correct answer when correct exploration outputs are in the minority, a regime where majority voting fails. Second, we introduce recursive DCR, an autoregressive reconciliation system that iteratively analyzes disagreements and allocates additional test-time compute. Recursive DCR achieves higher accuracy than fixed-compute baselines-reaching 93.3% on AIME 2024 and 92.0% on AIME 2025-while using roughly 27% less compute on average, demonstrating that attentive resource allocation is superior to uniform scaling. Third, we analyze disagreement among exploration outputs via a simple, training-free dispersion metric. Dispersion reveals a structured relationship between disagreement and test-time gains: in regimes where DCR is effective, higher disagreement among exploration outputs is associated with larger accuracy improvements from reconciliation. Together, these results show that disagreement, often viewed as noise, can be systematically exploited to improve test-time reasoning and reveal emerging scaling laws for agentic LLM systems.
エージェントティック AI システムにおける認知誘発リスクを理解する
大規模言語モデル (LLM) を利用したフロンティア エージェント システムは、人間に似た認知パターンを示します。これらのシステムがさまざまな領域にわたって深く統合されるにつれて、その認知的関与は人間社会に対して重大な懸念を引き起こしますが、まだ十分に研究されていません。このギャップに対処するために、私たちは、物理的認知から社会的認知、そして最終的に自己言及的認知に至る認知範囲によって定義される 3 つのレベルの枠組みに従って、認知能力の拡大によって引き起こされるリスクを体系的に分析します。私たちは、人間の主体性、自律性、制御能力に対する潜在的なリスクを、それぞれの認知レベルに応じて研究します。最後に、これらのリスクを軽減し、エージェント AI システムの制御性を強化し、長期的な安全な開発を保証する戦略を提案します。
原文 (English)
Understanding Cognition-Induced Risks in Agentic AI Systems
Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different domains, their cognitive engagement raises critical concerns for human society that remain insufficiently studied. To address this gap, we systematically analyze risks induced by expanding cognitive capabilities, following a three-level framework defined by their cognitive scope, from physical cognition to social cognition, and finally to self-referential cognition. We study their potential risks to human agency, autonomy, and control capability, corresponding to each cognitive level. We finally propose strategies to mitigate these risks and enhance the controllability of agentic AI systems, ensuring their long-term safe development.
人間の状態遷移の生理学的世界モデル
継続的なマルチモーダルセンシングにより、人間の生理機能を時折の臨床来院時だけでなく、日常生活を通して観察できるようになりました。ただし、ほとんどの医療人工知能システムは、現在の状態を認識し、リスクを推定し、個々のバイオマーカーを分析するように設計されています。これらは、現実世界の出来事、行動、状況、介入に応じて生理学的状態がどのように変化するかを直接モデル化するものではありません。ここでは、人間全体のレベルでこれらの変化を学習するためのイベント条件付きフレームワークである生理学的世界モデル (PWM) を提案します。 HumanState Transition Token を導入します。これは、イベント前の生理学的状態とイベントまたはアクション、関連するコンテキストおよび介入情報、イベント後の生理学的軌跡、観察された結果およびデータ品質を結び付ける、構造化された品質スコア付きユニットです。状態の表現から限定された介入計画までの 4 つの能力レベルと、4 つのデータ取得および検証プロトコルについて説明します。また、HumanState の表現、複数のタイムスケールにわたる予測、個別の対応予測、代替介入のシミュレーション、限定された計画、分布シフトの下での信頼性をカバーする 6 つのベンチマーク タスクを提案します。このフレームワークは共に、個別化された健康管理、行動介入設計、臨床医の監督下での意思決定支援に向けた実用的な道筋を提供すると同時に、予測を因果推論から明確に分離し、不確実性、安全性、ガバナンス、使用制限を明確にします。
原文 (English)
Physiological World Models for Human State Transitions
Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical visits. However, most health artificial intelligence systems are designed to recognize current states, estimate risks or analyse individual biomarkers. They do not directly model how physiological states change in response to real-world events, behaviours, contexts and interventions. Here we propose the Physiological World Model (PWM), an event-conditioned framework for learning these changes at the level of the whole person. We introduce the HumanState Transition Token, a structured, quality-scored unit that connects the physiological state before an event with the event or action, relevant context and intervention information, the physiological trajectory after the event, observed outcomes and data quality. We describe four capability levels, from state representation to bounded intervention planning, together with four data acquisition and validation protocols. We also propose six benchmark tasks covering HumanState representation, forecasting across multiple timescales, individualized response prediction, simulation of alternative interventions, bounded planning and reliability under distribution shift. Together, this framework provides a practical path towards personalized health management, behavioural intervention design and clinician-supervised decision support, while clearly separating prediction from causal inference and making uncertainty, safety, governance and limits of use explicit.
MoE ルーターガイドによる異種フェデレーテッド命令チューニングのためのクラスタリング
フェデレーテッド命令の微調整により、大規模言語モデル (LLM) は、データ共有を必要とせずに分散型のプライバシーに配慮したデータに適応できます。最近の Mixture-of-Experts (MoE) LLM は、モデルの容量をスケーリングしながらスパースなアクティベーションにより計算と通信を削減するため、フェデレーテッド ラーニングにとって特に魅力的です。しかし、既存のフェデレーテッド MoE 手法は主にパラメータの集約とパーソナライゼーションに焦点を当てており、クライアントのコラボレーションのための情報源としての MoE モデルのルーティング動作を見落としています。異種の命令分散では、無差別な集約がマイナスの転送につながる可能性があり、フェデレーテッド最適化中にどのクライアントが連携すべきかを特定する必要性が浮き彫りになります。私たちは、事前トレーニングされた MoE モデルからのルーティング署名を活用して、集約の前にクライアントのコラボレーションを組織する、ルーティングを認識したパーソナライズされたフェデレーテッド命令微調整フレームワークである ClientMorpher を提案します。我々は、2 つの相補的なクラスタリング戦略を調査します。ClientMorpher-C は、エキスパート アクティベーション プロファイルを使用してクライアントを直接クラスタリングします。もう 1 つは、ClientMorpher-E です。ClientMorpher-E は、最初にクライアント間の使用状況シグネチャに基づいてエキスパートをクラスタリングし、次にクライアント コラボレーション グループを導き出します。複数の命令追従タスクにわたる病理学的およびディリクレベースの異種クライアント分布を使用して、Databricks Dolly-15K データセットでフェデレーテッド命令微調整を行うために ClientMorpher を評価します。実験結果は、ルーティングを意識したコラボレーションは、同じ通信コストを維持しながら、従来のフェデレーテッド平均化やローカル トレーニングと比較して、パーソナライズされたパフォーマンスを一貫して向上させることを示しています。さらに、私たちの研究は、クライアント中心および専門家中心のクラスタリングが、スパース MoE LLM のパーソナライズされたフェデレーション命令の微調整に効果的かつスケーラブルなアプローチを提供することを示しています。
原文 (English)
MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning
Federated instruction fine-tuning enables Large Language Models (LLMs) to adapt to decentralized, privacy-sensitive data without requiring data sharing. Recent Mixture-of-Experts (MoE) LLMs are particularly attractive for federated learning because their sparse activation reduces computation and communication while scaling model capacity. However, existing federated MoE methods primarily focus on parameter aggregation and personalization, overlooking the routing behavior of MoE models as a source of information for client collaboration. Under heterogeneous instruction distributions, indiscriminate aggregation can lead to negative transfer, highlighting the need to identify which clients should collaborate during federated optimization. We propose ClientMorpher, a routing-aware, personalized federated instruction fine-tuning framework that leverages routing signatures from pretrained MoE models to organize client collaboration prior to aggregation. We investigate two complementary clustering strategies: ClientMorpher-C, which directly clusters clients using expert activation profiles, and ClientMorpher-E, which first clusters experts based on their cross-client usage signatures and then derives client collaboration groups. We evaluate ClientMorpher for federated instruction fine-tuning on the Databricks Dolly-15K dataset, using pathological and Dirichlet-based heterogeneous client distributions across multiple instruction-following tasks. Experimental results show that routing-aware collaboration consistently improves personalized performance compared to conventional federated averaging and local training, while maintaining the same communication cost. Furthermore, our study shows that client-centric and expert-centric clustering provides an effective and scalable approach for personalized federated instruction fine-tuning of sparse MoE LLMs.
尾部認識無線マップ予測のための物理学に基づいた VAE-EVT
超高信頼性低遅延通信 (URLLC) では、信号対雑音比 (SNR) が停止しきい値を下回る空間領域を正確に識別する必要があります。この文脈では、停止とは SNR が指定されたしきい値を下回るインスタンスを指します。URLLC の場合、これは SNR 分布の 0.1% 分位数と同じくらい厳しいものになります。従来の生成無線マップ モデルは、平均信号レベルの再構築に重点を置く傾向があり、正確な停止予測に重要な低い SNR を見落とすことがよくありました。この制限に対処するために、SNR のバルク分布とテール分布の両方を明確にモデル化する、物理学とテール情報に基づいた VAE-EVT (変分オートエンコーダー極値理論) フレームワークを導入します。私たちのアプローチは、シーンのジオメトリから視線、影、距離などの決定論的な特徴を抽出する、物理学に基づいた前処理段階から始まります。次に、デュアル 潜在エンコーダーがガウス混合を使用してバルク SNR をキャプチャし、一般化パレート分布 (GPD) を使用してテールをキャプチャします。修正された変分目標を採用することにより、モデルは両方の状況を共同で監視するようにトレーニングされ、極端なフェージング イベントに確実に集中することができます。 RadioMapSeer データセットで評価すると、私たちの方法は、0.1% SNR 分位点の低しきい値によって定義される停止領域で 4.83 dB の SNR RMSE を達成します。これは、21.90 dB の SNR RMSE を記録する最先端の GAN ベースのモデルを大幅に上回っており、停止しきい値が厳しくなるにつれてパフォーマンスの差は拡大します。
原文 (English)
Physics-informed VAE-EVT for Tail Aware Radio Map Prediction
Ultra-reliable low-latency communication (URLLC) requires precise identification of spatial regions where the signal-to-noise ratio (SNR) falls below an outage threshold. In this context, an outage refers to instances in which SNR falls below a specified threshold, which, for URLLC, can be as stringent as the 0.1% quantile of the SNR distribution. Traditional generative radio map models tend to focus on reconstructing average signal levels, often overlooking the low SNR that is crucial for accurate outage prediction. To address this limitation, we introduce a physics- and tail-informed VAE-EVT (variational autoencoder-extreme value theory) framework that distinctly models both the bulk and tail distribution of SNR. Our approach begins with a physics-informed preprocessing stage that extracts deterministic features, including line-of-sight, shadowing, and distance, from the scene geometry. A dual-latent encoder then captures the bulk SNR using a Gaussian mixture and the tail using a generalized Pareto distribution (GPD). By employing a modified variational objective, the model is trained to jointly supervise both regimes, ensuring focused attention on extreme fading events. Evaluated on the RadioMapSeer dataset, our method achieves an SNR RMSE of 4.83 dB in the outage region defined by the low threshold of 0.1% SNR quantile. This significantly outperforms the state-of-the-art GAN-based model, which records an SNR RMSE of 21.90 dB, with the performance gap widening as the outage threshold becomes more stringent.
ベンチマークの罠: AI 評価における権力と不正義の構造
人工知能 (AI) ベンチマークは中立的な評価ツールではなく、AI 内の競争、権力、研究の優先順位を形成する社会技術的な成果物です。ベンチマークはシステムの評価を標準化し、最先端のパフォーマンスに名声、引用、信頼、組織的影響力を与えるリーダーボードの作成を容易にします。競争力のある AI システムの開発コストが上昇するにつれ、これらの報酬は業界から資金提供を受けた強力な研究機関にますます集中します。この論文は、これらの懸念をアイリス・マリオン・ヤングの抑圧と構造的不正義の理論の中に位置付けます。同論文は、現在のベンチマーク慣行がAI研究のさまざまな関係者に影響を与える組織的な危害を永続させる可能性があり、ヤング氏の「抑圧の側面」のうちの4つと一致すると主張している。ベンチマークの文化はさらに、構造的不正義の根源として位置づけられています。なぜなら、こうした害は、たとえ明示的な不正行為がなくても、常態化され、個別に擁護可能な慣行やネットワーク効果から生じるからです。ベンチマークは、既存の権力構造を強化し、可能な研究軌道を狭めることにより、この分野が認識論的に堅牢で社会的に有益な方法で進歩することを実際に妨げる可能性があります。
原文 (English)
The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's "faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effects, even without explicit wrongdoing. By reinforcing existing power structures and narrowing possible research trajectories, benchmarking may in fact prevent the field from advancing in epistemically robust and socially beneficial ways.
多層フィードフォワード ニューラル ネットワークの集中結果
任意の固定 $\rho$ と各正の整数 $n$ について、$\rho$ 層、最初の層 (入力層) に $n$ ニューロン、最後の層に 1 つのみのニューロン (出力ニューロン) を持つ多層フィードフォワード人工ニューラル ネットワークを考えます。非常に大まかに定式化すると、主な結果は、ある層から次の層への接続の重みの分布が、すべての大きな $n$ に対して、$n$ に依存しない固定連続 (ただし、それ以外は任意) 曲線によってうまく近似され、$n$ 入力ニューロンの値が連続確率密度関数で独立して同一に分布する場合、すべての $\varepsilon > 0$ について、出力の値がニューロンは $[\psi - \varepsilon, \psi + \varepsilon]$ にあり、$n$ が無限大になる傾向があるため、1 になる傾向があります。
原文 (English)
A concentration result for multilayer feedforward neural networks
We consider for an arbitrary fixed $\rho$ and for each positive integer $n$ a multilayer feedforward artificial neural network with $\rho$ layers, $n$ neurons in the first layer (the input layer) and only one neuron, the output neuron, in the last layer. Very roughly formulated, the main result is that if the distribution of weights of connections from a layer to the next are, for all large $n$, approximated well by a fixed continuous (but otherwise arbitrary) curve which does not depend on $n$, and if the values of the $n$ input neurons are independently and identically distributed with a continuous probability density function, then there is a number $\psi$ such that for all $\varepsilon > 0$ the probability that the value of the output neuron is in $[\psi - \varepsilon, \psi + \varepsilon]$ tends to 1 as $n$ tends to infinity.
Incoherent by Design? On the Moral Self-Consistency of LLMs
LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situa…
UC-PSRO: 敵対的な群れにおけるゲーム理論的な行動方針生成のための通信ドロップアウト カリキュラムを備えたユーティリティ条件付きポリシー空間応答オラクル
私たちは、米国空軍の公的 SBIR 要請を動機として (ただし、そこから派生したものではありません)、通信が低下した環境において、適応型の赤色の敵対者に対して青色 UAS の群れに対してゲーム理論的に最適化された行動指針 (COA) を生成することを研究しています。我々は、次の 3 つのメカニズムを組み合わせた UC-PSRO (通信ドロップアウト カリキュラムを備えたユーティリティ条件付きポリシー空間応答オラクル) を提案します。(i) PSRO のセルフプレイにより、青と赤のポリシーは、固定されたスクリプト化された対戦相手に対する一方の側ではなく、互いに最も最適な応答としてトレーニングされます。 (ii) トレーニング中にディリクレ分布からサンプリングされた、Commander's-Intent 重みベクトルでの Blue ポリシーの FiLM 条件付け。これにより、1 つのトレーニング済みポリシーが再トレーニングなしで実行時に再ステアリング可能になります。 (iii) トレーニング中に通信グラフのエッジ ドロップアウトをアニーリングするカリキュラムにより、群は完全な接続に依存するのではなく、分散型のピアツーピア フォールバックを学習します。 N=25 の Blue エージェントで 5 つのシードと N=200 までのスケーラビリティ スイープを使用して、勧誘の海洋シナリオの合成未分類の代役を評価します。私たちは、画一的な勝利ではなく、真のトレードオフを発見しました。コミュニケーションドロップアウトカリキュラム単独では、学習した方法の中で最も強力で堅牢なミッション完了率が得られ、拒否が増加するにつれて直感に反して向上します(ドロップアウトが0から0.75に上昇すると、成功率は35%から62%になります)。ユーティリティ コンディショニングと PSRO セルフ プレイを追加すると、固定予算内での収束が大幅に遅くなり、固定対戦相手ポリシーを超えるセルフ プレイの信頼できる悪用可能性の利点は見つかりません。どちらも統計的には、ゼロに近い小さなギャップと区別できません。私たちはこれを、1 つの方法が有力であると誇張するのではなく、実証された堅牢性の利点によってまだ相殺されていない収束コストとして正直に報告し、単一のコンシューマ GPU でステップあたり 1 桁ミリ秒で N=200 エージェントで完全にベクトル化されたオープン環境トレーニングを提供します。
原文 (English)
UC-PSRO: Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum for Game-Theoretic Course-of-Action Generation in Adversarial Swarms
We study generating game-theoretically optimized Courses of Action (COAs) for a Blue UAS swarm against an adaptive Red adversary in a communication-degraded environment, motivated by (but not derived from) a public U.S. Air Force SBIR solicitation. We propose UC-PSRO (Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum), combining three mechanisms: (i) PSRO self-play, so Blue and Red policies train as approximate best responses to each other rather than one side against a fixed scripted opponent; (ii) FiLM conditioning of the Blue policy on a Commander's-Intent weight vector, sampled from a Dirichlet distribution during training, so one trained policy is re-steerable at execution time without retraining; and (iii) a curriculum annealing communication-graph edge dropout during training, so the swarm learns decentralized, peer-to-peer fallback instead of depending on full connectivity. We evaluate on a synthetic, unclassified stand-in for the solicitation's maritime scenario, with 5 seeds at N=25 Blue agents and a scalability sweep to N=200. We find a genuine trade-off, not a uniform win: the communication-dropout curriculum alone gives the strongest, most robust mission-completion rates of any learned method, improving counter-intuitively as denial increases (35% to 62% success as dropout rises from 0 to 0.75); adding utility-conditioning and PSRO self-play substantially slows convergence within a fixed budget, and we find no reliable exploitability advantage for self-play over a fixed-opponent policy, both statistically indistinguishable from a small, near-zero gap. We report this honestly as a convergence cost not yet offset by a demonstrated robustness benefit, rather than overstating one method as dominant, and provide a fully vectorized, open environment training at N=200 agents in single-digit milliseconds per step on a single consumer GPU.
FedPA-LoRA: 異種フェデレーテッド LoRA における集約エラーと初期化エラーを軽減するための製品に合わせたフレームワーク
Low-Rank Adaptation (LoRA) を使用すると、大規模な言語モデルの効率的なフェデレーション微調整が可能になりますが、その因数分解されたパラメーター化により、ローカル更新の正確な集約とローカルに最適化された因子の連続性の間に緊張が生じます。因子ごとの集計では集計の不一致が生じますが、因子の連続性はよりよく保持されます。一方、積空間再構成では、新しく再構成された因子からの因子レベルの初期化の不一致が大きくなるという犠牲を払って、この不一致が減少します。私たちは、これらの制限に共同で対処し、同種クライアント ランクと異種クライアント ランクの両方で収束することが証明されている、製品に合わせたフェデレーテッド LoRA フレームワークである FedPA-LoRA を提案します。各クライアントは、通信ラウンド全体にわたってローカル要因を保持し、その製品をランク固有のグローバル参照に合わせて調整し、ローカル最適化の連続性を維持しながら、データの異質性の下でグローバルな一貫性を促進します。サーバーは、共通の製品空間で異種ランクの更新を集約し、密な集約を形成することなくランク制約のあるグローバル アダプターを効率的に再構築します。この設計は、クライアント固有の計算と通信の予算をサポートします。自然言語理解および生成タスクに関する実験では、FedPA-LoRA が、さまざまなレベルのデータ異質性および同種および異種ランク設定にわたって、代表的なベースラインを常に上回っており、異種クライアント ランクの下で平均 GLUE 精度が最大 6.82 ドル パーセンテージ ポイント向上していることが示されています。
原文 (English)
FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA
Low-Rank Adaptation (LoRA) enables efficient federated fine-tuning of large language models, but its factorized parameterization creates a tension between accurate aggregation of local updates and continuity of locally optimized factors. Factor-wise aggregation incurs aggregation mismatch but better preserves factor continuity, whereas product-space reconstruction reduces this mismatch at the cost of greater factor-level initialization mismatch from newly reconstructed factors. We propose FedPA-LoRA, a product-aligned federated LoRA framework that jointly addresses these limitations and provably converges under both homogeneous and heterogeneous client ranks. Each client preserves its local factors across communication rounds and aligns its product toward a rank-specific global reference, maintaining local optimization continuity while promoting global consistency under data heterogeneity. The server aggregates heterogeneous-rank updates in the common product space and efficiently reconstructs a rank-constrained global adapter without forming the dense aggregate. This design supports client-specific computation and communication budgets. Experiments on natural language understanding and generation tasks show that FedPA-LoRA consistently outperforms representative baselines across varying levels of data heterogeneity and homogeneous- and heterogeneous-rank settings, with up to a $6.82$ percentage-point improvement in average GLUE accuracy under heterogeneous client ranks.
因果知識グラフにおけるヘルスケア LLM の基礎: フレームワーク、メトリクス、および心臓血管パイロット
大規模言語モデル (LLM) は、医療意思決定支援のために提案されることが増えていますが、その評価では依然として、介入、メカニズム、害、証拠、不確実性についての推論ではなく、単一回答の正確さが重視されています。私たちは、医療における介入指向の LLM 行動のための再現可能なグラフ中心の評価フレームワークを提案し、心血管パイロットでストレス テストを行います。このフレームワークには 4 つのコンポーネントがあります。(i) アサーションが安定した識別子を持つ来歴を保持する第一級のノードであるドメイン因果知識グラフ。 (ii) 任意の臨床シナリオが与えられた場合に、関連する具体化されたアサーションサブグラフを取得する、シナリオ条件付きサブグラフ抽出ステップ。 (iii) 取得したサブグラフをモデルのコンテキストに組み込む方法を変える 4 つの制御されたグラウンディング条件 (非グラウンディング C1、ナレッジ グラフ C2、因果グラフ C3、統合 C4)。 (iv) アサーション識別子に基づいた自動スコアリング パイプライン。介入の精度やその他の評価尺度を 1 回のパスで計算します。このフレームワークをテストするために、8 つの推論失敗モードにわたるカテゴリバランスのとれたシナリオ ジェネレーターを構築し、心血管グラフ上でインスタンス化しました。メトリクス パネルは、解釈可能な非冗長軸に沿って条件を識別します。C4 は最も強い因果エッジ F1 (0.838)、悪影響 F1 (0.833)、証拠精度 (0.738)、およびサポートされていない請求率 (0.114) を取得しますが、C1 は測定可能な因果関係または証拠根拠がない状態で最高の生の介入精度 (0.948) を取得します。
原文 (English)
Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot
Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.
Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL
LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system compari…
TwinGridShield: Consequence-Aware Runtime Authorization for LLM Grid-Agent Actions
Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but synta…
Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish
Chain-of-thought monitoring is a potentially useful safety signal, but its reliability across languages and behavioral settings remains unc…
Large Language Model Assisted Operational Monitoring for Battery Energy Storage System Integrated Power Distribution Networks
Battery energy storage systems (BESS) are increasingly used in distribution networks for voltage regulation and demand response, which incr…
Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs
Large Language Models (LLMs) are notorious for struggling with assessing their own uncertainty, detecting knowledge conflicts, or recognizi…
A survey of AI-generated voices and their detection
The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power…
Does the Proof Prove It That Way? Faithful Formalization of Elements Proofs
In formal verification, both the autoformalization of statements and automated proof search have been studied extensively. While automated…
OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks
Frontier AI models have advanced rapidly, but they still struggle with telecom-specific tasks. We present Open Telco (OTel), an open teleco…
Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization
When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure…
Mental Model Management: An Operator-Based Framework for LLM Memory
Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conc…
Dynamic Multi-Byte Prediction With Hierarchical Language Models
Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword…
A Network-driven Framework for Public Event Forecasting via Dynamic Interaction Network Evolution
Effective public event forecasting is essential for intelligent service systems, enabling proactive risk management, adaptive resource allo…
EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints
Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inference cost poses significant…
Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation
Chart-to-code generation requires a model to read the fine-grained visual details of a chart and write executable code that reproduces it.…
From Contexts to Values: Context-Dependent Defeat in Abstract Argumentation
In value-based argumentation, an audience's ordering of values decides which attacks succeed as defeats. In many settings the deciding fact…
ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity Search
Most LLM-based automated algorithm design methods optimize a designated component within a human-specified scaffold, fixing overall organiz…
Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking…
From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM
Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically mean…
Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback
Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deploym…
When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction
Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades th…
TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation
Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of ap…
VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation
The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing…
Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds
Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipli…
Rotation-Invariant Multi-IMU Activity Recognition under Independent Per-Location Orientation Shifts
Human Activity Recognition (HAR) with self-administered wearables, such as at-home rehabilitation and exercise monitoring, often requires r…
Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals in Conflict
How can common ground between societies in conflict be identified when citizens' acceptability of peace agreements is shaped by contested n…
A Responsible Artificial Intelligence Framework for Groundwater Modeling
The rapid development and widespread application of artificial intelligence (AI) have sparked intense discussions on how to deploy responsi…
THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts
Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure. Existing…
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compress…
Adaptive Mixing of Policies from Searching and Policies from Learning
Background: Distillation of training targets generated thru search/planning has proven useful in reinforcement learning, but search can tak…
HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation
Large language model (LLM) agents often perform poorly on complex, long-horizon tasks because their context becomes increasingly cluttered…
PLeDO: Pain Level Detection for Osteoarthritis from EMR Data
Osteoarthritis (OA) is a progressive chronic joint disease resulting in a breakdown of articular cartilage and bone when damaged joint tiss…
Toward AI-Friendly Cartography: Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps
Foundation models (FMs) increasingly support multimodal and geospatial reasoning, yet it remains unclear whether cartographic principles de…
Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campaign
We present a forensic analysis of the generation pipeline behind a recent AI-driven influence campaign. We introduce PROPAGIA, a corpus of…
Intent-Driven Situation Tracking for User-Centric Multi-Turn Agents
User-centric multi-turn agents must act on an evolving task situation shaped by changing user intents, accumulated tool-grounded facts, mis…
Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration
When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or…
KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving
KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of it…
Pricing the Risk of Runtime Compression: Anytime-Valid Admission and a Served-Output Law for Compressed Serving State
Runtime compression of serving state trades quality for capacity with no priced guarantee: systems adapt precision on load signals with no…
RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning
The growing ecosystem of large language models (LLMs) offers huge potential to optimize performance-cost trade-offs. However, their heterog…
The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale
As AI systems become increasingly capable of autonomous action, determining whether an agent is technically capable of performing an action…
Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs
Tool-calling LLM agents navigate unfamiliar codebases with a handful of generic primitives for listing, reading and searching files (ls, ca…
RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration
Ethereum is now integral to mission-critical sectors, including finance, healthcare, and supply chain management. Execution fees, commonly…
CoupVisor: Strategy Optimization by Round and Challenge Decision Support
This paper presents CoupVisor, a decision-support system for the hidden-information card game Coup. It addresses two questions: what a play…
Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation
Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Thr…
Bounded Agents: Delegation Security for Multi-Agent AI Systems
LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissi…
Breaking and Defending LLM-Powered Social Media Bot Detection Systems
The rise of social media bots poses a persistent threat, enabling misinformation, opinion manipulation, and the erosion of trust in online…
Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning
Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver-assistance systems. Previous stud…
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts,…
Augmenting Text to Increase Translation Difficulty
As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish be…
Navigation-Informed Embeddings: Dense-Retriever Adaptation from Agent Search Traces
Agentic retrieval workflows produce query, retrieval, and stopping traces as a byproduct of answering questions. We study how these traces…
Solvable Sokoban Without a Solver via Diffusion
Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be exponentially long and there is no sho…
ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction
Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears…
MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment
Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, mod…
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whethe…
Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance
Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within policy. We introduce Fiducia-bench,…
Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics
Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic systems, yet their long-term instabilit…
Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling
Accurate protein structure prediction is fundamental to structural biology because protein structure underlies molecular function and provi…
Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity
How should we assess whether large language models can perform mathematical invention? I argue that this question is currently underspecifi…
When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification
Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were p…
FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection
Multi-view multi-label feature selection aims to identify a compact and informative feature subset from heterogeneous views while preservin…
TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment…
Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain
Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early exploration failures. However, since uns…
Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication
Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content…
Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior
Personalized game generation requires inferring a player's abilities and behavioral style from how they play. Large language models have ma…
Competing at Every Price Point with Agentic Evolution over a Menu of LLMs
Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every competitor price p…
BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics
Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workfl…
Process-Constituted Intelligence: A Shared Criterion for Humans and Machines
Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself. Generative AI…
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations cent…
DriveCache: Action-Aware Caching for Driving World Model Inference
Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning ev…
What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics
Task completion is the standard metric for evaluating context compression, yet it is incomplete: compression can increase an agent's intera…
AstronOS: A Unified Execution Model and Runtime for Long-Horizon Agentic Systems
Agentic systems often organize execution and state around a single conversation, model invocation, or agent instance, even when real work s…
Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152
Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (L…
A Policy Algebra for Trust-Preserving Agentic AI Execution
Large language model-based agentic frameworks primarily optimize capability: whether an agent can reason, retrieve information, call tools,…
Reasoning-supported Robustness Validation of Automotive E/E Components
This paper presents an ontology-supported approach to tackle the complexity of the Robustness Validation (RV) process of automotive electri…
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence
Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computation…
Drive, Pack, Fly: The Travelling Thief Problem with Drone
In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on routing efficiency. An onb…
The Value of a Prompt: An LLM-Relative Kolmogorov-Complexity Approach
In a world where valuable artifacts are increasingly created, completed, or processed by LLMs, the central economic question is not only wh…
Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics
Neurosymbolic (NeSy) Artificial Intelligence aims to integrate Deep Learning (DL) architectures with symbolic reasoning. While initial NeSy…
HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents
Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in response to environmental feedbac…
JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills
Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workfl…
Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation
The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical…
Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling
Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, eli…
DeepInsight II: One Trace from Benchmark to Robot
Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized har…
CUBICS: Situation-aware performance estimation for safety-relevant ML components
Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related app…
Probabilistic Circuits as Reasoning Machines in Artificial Intelligence (Part I)
This cumulative habilitation thesis studies probabilistic circuits (PCs) as a powerful and tractable framework for reasoning and learning u…
Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make dec…
CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction
Cancer survival prediction supports treatment planning, risk stratification, and follow-up management. Existing methods use structured clin…
Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate
Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time. Put plainly, inste…
A Shop Floor Production Scheduling Case based on RFID-supported Smart Factory
Radio frequency identification (RFID) technology has been widely implemented for real-time data collection in manufacturing shop floors, wh…
Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement
Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of tradition…
PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning
LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods in…
Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies
Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We intr…
Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents
This paper presents Chronocooked, a reinforcement learning (RL) benchmark suite for studying implicit interval timing in RL agents. Inspire…
FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies acro…
LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing
Financial document validation in production, such as payroll auditing, tax compliance, and loan underwriting, demands exceptional accuracy,…
GRIP: Grounded Reasoning via Information-Restricted Premises
High-capacity encoders in retrieval-augmented generation (RAG) can let the query dominate the latent state, leaving retrieved evidence func…
When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding
We study how teams of AI coding agents coordinate while solving programming tasks. Current evaluations usually report whether the agents co…
Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment
Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition resources for the over 100…
Quipu: A Governed Bitemporal Knowledge Graph Store
Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clea…
Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning sh…
What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outpu…
A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation
Defining the reasoning boundaries and ensuring the reliability of Large Reasoning Models (LRMs) remains a critical challenge. Current bench…
Orbital AI Computing: Carbon Tradeoffs Across Satellite Scale
Low Earth Orbit (LEO) computing is emerging for low-latency, globally distributed AI services, enabled by advances in satellite constellati…
Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)
Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7--3.2x the t…
WARA: Toward Automated Wireless Optimization Research with Closed-Loop LLM Agents
Large language model (LLM) agents are increasingly capable of tool use, code execution, artifact inspection, and iterative revision, creati…
From Reactive to Autonomous: Evolution of AI Operations in Cloud Network Infrastructure
The operational model for cloud network infrastructure has undergone a fundamental transformation over the past decade. What began as manua…
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an objec…
Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework
In Multimodal Question Answering (MQA), models are required to jointly encode and integrate heterogeneous information from multiple modalit…
Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures
Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle plat…
Intelligent Base Station Deployment in Urban Wireless Networks: A Geographic Data-Informed Digital Twin Approach
The placement of base station (BS) is a fundamental determinant of coverage and capacity of urban wireless networks. Yet large-scale BS dep…
Extend the Safety Horizon for Intelligent Transportation Systems through Semantic-Aware Cooperative Perception
Cooperative perception enables vehicles and infrastructure to exchange sensor data via Vehicle-to-Everything (V2X) communication, extending…
Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models
Small language models in the ten to one hundred million parameter range are attractive for on device inference, rapid experimentation, and…
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plaus…
Understanding AI Anxiety in the Workplace: A Multimethod Investigation Using Fear Acquisition Theory and the Technology Acceptance Model
As artificial intelligence (AI) rapidly diffuses and concerns about job displacement intensify, the psychological mechanisms underlying AI…
DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs
As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This paper investigates wh…
Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion
A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation…
Explaining Reinforcement Learning Decisions in Self-adaptive Systems
Reinforcement Learning (RL) has been extensively used in autonomous and self-* systems, but RL policies, especially deep RL ones relying on…
AutoMem: A Text-Gradient Recursive Self-Improvement Framework for Automated Memory Architectures Search
Long-term memory is increasingly central to LLM agents, yet memory design remains a highly coupled architecture problem: what to encode, ho…
Local AI pre-screening for human triple-blind peer review in health sciences
Academic peer review is under mounting strain: NeurIPS 2025 received 21,575 submissions, ICLR 2025 received 11,603, and ICML 2025 received…
LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review
Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weake…
Inference-Time Mitigation of Adversarial Political Bias in Large Language Models
As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-pa…
Characterizing Rhetorical Misalignment in Decision-Making with Language Models
Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly i…
DeMTS: Denoising Trajectories as Multivariate Time Series for Hallucination Detection in Diffusion Language Models
Diffusion large language models (D-LLMs) have emerged as a promising paradigm for text generation. However, similar to autoregressive LLMs,…
Fractional Optimizers Meet Fractal Activation Functions: An Empirical Study of Multi-Scale Optimization in Neural Network
Fractional optimization methods and fractal activation functions are two independent directions for improving neural network training. Frac…
Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fields is controlled -- i…
BDIP-Net: Dual-Interaction Graph Learning for Property Prediction of Bilayer Materials
Stacked bilayer materials exhibit rich stacking-dependent properties driven by the interplay between strong intra-layer bonding and weak in…
iFuzz-Meta: An Interpretable Fuzzy Learning Framework Bridging Top-Down and Bottom-Up Knowledge Integration
Interpretable representation learning remains a key challenge in modern neural computation, particularly when models are expected not only…
SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation
Dirty-history rollouts make multi-turn on-policy self-distillation (OPSD) brittle: once a student emits an erroneous intermediate reply, la…
Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings
In this study, we revisit three widely used techniques in vector search and utilize them to optimize vector embedding indexing through clus…
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
The development of 0.1$^{\circ}$ global weather forecasting models based on machine learning (ML) is constrained by the limited availabilit…
Do Uncertainty Signals Help? A Systematic Study of Uncertainty-Aware Decoding with Rollback Mechanisms
Prediction uncertainty is a widely adopted metric for quantifying model confidence, with downstream applications spanning model explanation…
FedImp: Enhancing Federated Learning Convergence with Impurity-Based Weighting
Federated Learning (FL) is a collaborative paradigm that enables multiple devices to train a global model while preserving local data priva…
P2E-VQ: ECG-linked representation augmentation for PPG via discrete patch retrieval
Photoplethysmography (PPG) is widely used in consumer wearables because of its low cost and ease of acquisition. However, unlike electrocar…
pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier
We introduce pico-type, a byte-level multi-head content classifier with approximately 1.5 million parameters that simultaneously predicts s…
Ring-based Spatial Transformer: Learning Non-linear Spatial Interactions between Building Distribution and Pedestrian Flow
This study proposes a ring-based SpatialTransformer to learn how building uses at different distances from a railway station interact to ge…
Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG Representation Learning
Accurate recognition of pain using physiological signals remains a challenging problem due to pain's subjective nature and high inter-indiv…
BRA-Audit: Budgeted Runtime Auditing for LLM Multi-Agent Systems via Cumulative-Exposure Audit-Point Placement
LLM-based multi-agent systems (LLM-MAS) solve complex tasks through specialized collaboration, but inter-agent dependencies can propagate h…
ARGUS: Attention-Guided Transformers for Scalable Person Identification Using Wi-Fi Telemetry
Passive, device-free person identification offers an alternative to camera- and wearable-based biometrics, yet existing wireless approaches…
Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection
While Self-Supervised Learning (SSL) effectively extracts general representations from noisy, unconstrained physiological signals such as p…
Offline Ambient-Controlled Latent Diffusion: Architecture, Telemetry, and On-Device Evaluation
Most mobile image-generation applications are thin clients over cloud services, leaving outputs hard to audit. We present an Android latent…
Information-Theoretic Causal Modelling of Semiconductor Process Dynamics
With the progress of the semiconductor industry toward increasingly complex compute devices and tighter process tolerances, advanced proces…
Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans
Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evalu…
Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation
LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods…
Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level
Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior.…
Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning
With the increasing digitization of personal and corporate communication, the automatic sanitization of textual data has become a crucial c…
Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning
Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampl…
Path2ST: Hierarchical Cell-Tissue Grounded Cross-Modal Translation for Spatial Transcriptomics
Predicting spatial gene expression from hematoxylin and eosin (H\&E)-stained images offers a cost-effective alternative to spatial transcri…
Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data
Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability land…
DeCo-MIL: Debiased Counterfactual Reasoning for Long-Tailed Whole Slide Image Analysis
Multiple instance learning (MIL) is widely used for weakly supervised whole slide image (WSI) analysis. However, under long-tailed distribu…
Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra
Following the molecular discovery and synthesis revolutions, scalable automated structure elucidation from routine spectroscopic data remai…
Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering
The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets. Ho…
Tail-Aware Top-$k$ On-Policy Distillation
On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is tr…
A Novel Fourier Feature Network for Solving Partial Differential Equations
Building on the foundation of single-hidden-layer neural networks, Fourier Feature Networks (FENs) are proposed, which incorporate Fourier…
Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks
Deep learning models of nanocrystal synthesis enable the prediction of size and shape by encoding precursors and reaction conditions. Howev…
Class Imbalance and Batch Effects in LLM-Based Screening for Systematic Reviews
This study analyses LLMs in imbalanced binary classification, using study screening in systematic reviews as the application domain. An exp…
PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models
We introduce PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning…
Synthesizing Post-Acetazolamide Cerebral Blood Flow Maps from Baseline MRI in Moyamoya Using 3D Generative AI
For patients with Moyamoya disease, impaired cerebrovascular reserve (CVR) is an important hemodynamic criterion for recommending extracran…
Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening
Assessment of ventriculomegaly (VM) on fetal brain ultrasound relies primarily on measuring lateral ventricular atrial width on standard pl…
NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving
Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated dri…
Artificial Intelligence as a Tool for Combating Child Labour: A Real-Time Edge Vision Pipeline for Child Detection and Age Estimation
An estimated 138 million children remain in child labour worldwide, and the monitoring systems used by affected sectors, built on periodic…
ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning
The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function variants of the original Kolmogorov-Arnold Network---…
Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters
Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) b…
Handover Analysis for Vehicular Communication with Explainability on the Fly
Handover (HO) management in vehicular networks requires fast and reliable decision-making under highly dynamic conditions. While machine le…
Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce
Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much…
Writing Style Similarity Reflects Academic Genealogy
As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusa…
Evaluating Agentic Code Repair Capabilities in Distributed Systems
LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench…
Workspace Topology as an Attack Vector in Agentic Coding Assistants
Agentic coding assistants are finding widespread use, not just in new code development but in quickly ingesting and leveraging third-party…
The Open-Strategy Dictator Game: Cooperation Under Mutual Transparency
We introduce the Open-Strategy Dictator Game (OSDG), a variant of the classic dictator game in which each player's strategy is a natural-la…
Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task
AI-generated music detectors are commonly evaluated against original songs, but real-world uploads are often remixed, re-encoded, pitch-shi…
SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable
Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post…
PaSTel: Anchoring Histology in Spatial Transcriptomics via Multi-Scale Hierarchical Bio-Prior Contrastive Pretraining
Spatial transcriptomics (ST) links tissue morphology with molecular programs, motivating multimodal pretraining methods that align histolog…
Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment
This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed image…
Generative data assimilation highlights fronts as key regulators of ocean energy cascade
Mesoscale eddies are fundamental to the ocean circulation, yet the extent to which submesoscale motions, a few kilometers across, influence…
Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning
Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired r…
Do Geometry-Aware Positional Encodings Help Transformers in Spatial Imperfect-Information Games?
Transformers applied to spatial imperfect-information games must represent map geometry while tracking hidden entities through time. We ask…
GaussMemory: Task-Driven 3D Gaussian Scene Memory for Long-Horizon Robotic Manipulation
Long-horizon robotic manipulation fundamentally relies on persistent spatial memory. However, existing 3D memory systems function merely as…
PAS-QFL: Personalized Ansatz Selection for Quantum Federated Learning under Client Data Heterogeneity
Quantum federated learning (QFL) lets multiple quantum clients collaboratively train quantum neural networks (QNNs) without sharing private…
RamseyGadgets: A Graph Construction Dataset for LLMs
Constructing special graphs is an important task within graph theory and computer science. Many popular graph constructions are the result…
FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool…
MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems
Although visual reasoning is crucial for solving complex geometry tasks, existing vision-language models rely heavily on text-only reasonin…
SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system
The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward aut…
Hierarchical Agentic Incident Response with Digital-Twin-Validated Attack Inference
Network incident response remains slow and labor-intensive as the defender must infer multi-stage attacks from partial observations and tra…
DualMiT-Net: Local-Global Transformer-Convolutional Fusion for Breast Mass Segmentation in Mammographic Regions of Interest
Breast mass segmentation is an important step in computer-aided mammography, but it remains difficult because masses can have low contrast,…
MotionGS-SLAM: Event-Modulated Gaussian Splatting for Motion-Blur Robust SLAM
Current Vision-based SLAM systems fail catastrophically when motion blur corrupts the visual input, as they attempt the ill-posed inverse p…
Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints
Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, di…
GATTA: Graph Active Learning with Test-Time Augmentation
Test-time augmentation (TTA) has proven effective for improving model robustness and uncertainty estimation in computer vision, yet its app…
Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning
Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly while continuing to improve b…
WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing
In this work, we introduce WeSCE, a benchmark for quantifying security drift in code editing under weak-security constraints, where tasks s…
Beyond Direct Access: Resource Hijacking in LLM Agents
Large language model agents are increasingly connected to high-value resources such as computing infrastructure, credentials, usage budgets…
CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation
Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing m…
Fast Test-Time Refinement for Robust Learned Image Compression
Learned image compression (LIC) has demonstrated remarkable rate-distortion (RD) performance in benign settings. However, the high represen…
From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems
Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, enviro…
Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models
We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse…
Scale-Consistent Posterior Dynamics for Diffusion Inverse Problems
Posterior sampling with a pretrained diffusion prior is governed by a conditional score whose intermediate likelihood component is generall…
Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models
World models may predict the future without making clear which parts of their hidden state actually drive those predictions. We ask whether…
A Unified Backbone--Expert Framework with Relation-Token and Residual--Classifier Interfaces for Automatic Modulation Recognition
Automatic modulation recognition (AMR) faces distinct representation bottlenecks under varying observation lengths, where a single model ar…
LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset
Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditio…
FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection
The increasing complexity of digital financial systems has reshaped financial fraud detection from isolated transaction classification into…
The Quality of Claude AI-authored Python Tests Is Not Weaker Than Human-authored Tests
We evaluate the quality of Claude AI-written Python tests against human-written Python tests from two established open-source projects Djan…
Valhalla: A Layered Knowledge-State and Service-Governance Framework for Long-Term Scientific Knowledge Work
As large language model (LLM) agents are increasingly adopted in scientific research, external knowledge bases, knowledge graphs, and long-…
CG-GLORE: A Conjugate Gradient-Based Global-Local Regularization Network for Sparse-View CT Reconstruction
Sparse-view computed tomography (CT) reduces radiation dose by acquiring fewer projection views, but the resulting inverse problem is highl…
UAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection
Unmanned Aerial Vehicles (UAVs) play a crucial role in various scenarios ranging from disaster response to traffic surveillance. However, a…
VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction
Maintaining global geometric consistency is a central challenge in long-sequence 3D reconstruction, with scale drift being the most critica…
VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments
Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robo…
PhaseLoRA: Control-Regime-Conditioned Low-Rank Adaptation for Continuous-Action Vision-Language-Action Policies
Parameter-efficient fine-tuning (PEFT) is a natural way to adapt pretrained vision-language-action (VLA) policies, but most adapter designs…
No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database…
MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation
Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention…
Shape Operator PCA: Curvature-Aware Projections for Geometric Machine Learning
In this paper, we propose SHOPCA (Shape Operator-based Principal Component Analysis), a novel method for unsupervised metric learning and d…
Logical Embeddings for Argument Analysis
We propose a new framework for machine-learning-oriented argument analysis tasks. Our proposal involves replacing traditional contextualize…
When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text
Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regi…
ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution
To accelerate single image super-resolution (SISR) networks on large images (2K-8K), many recent approaches decompose an image into small p…
SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning
While Parameter-Efficient Fine-Tuning (PEFT) has substantially reduced the hardware cost of adapting Large Language Models (LLMs) by decrea…
AudioTQ: A Data-Oblivious 6-Bit CPU Audio Codec via Randomized Hadamard Rotation and Lloyd-Max Quantization
Lossy audio compression algorithms traditionally rely on psychoacoustic modeling and frequency-domain representations (e.g., MP3, AAC, and…
Agent Inheritance Protocol: Speculating on Feralized Agents After Principals Die
You will die eventually. Your agents may not. An AI agent operating on decentralized blockchain infrastructure has no concept of death; it…
Afterlife Delegation Protocol: Speculative Design of Self-Sovereign Agents that Outlive Their Principals
Afterlife Delegation Protocol is a speculative design project that asks what death becomes when a will can act eternally. We design a specu…
Chameleon: An Adaptive AI-Driven Honeypot Architecture Using Threat-Calibrated Particle Swarm Optimization and Semantic Deception Rapidly-Exploring Random Trees
An invariant behavioral profile is the defining vulnerability of traditional honeypot installations: a skilled adversary can confirm the pr…
FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge
Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests into pixel-level visual ground…
Invariant Pretraining for Robust Code Representations
Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, w…
An Evaluation Framework for National AI Regulation
Governments use laws, institutions, funding programs and nonbinding guidance to shape how AI is developed and used. Comparing these nationa…
ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems
The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multim…
NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models
Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability tha…
Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks
Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap…
Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off
Attention mechanisms have driven machine learning for a decade, from neural machine translation to language models that do general-purpose…
Optimal Lower Bounds for Networked Information Aggregation
The problem of networked information aggregation, studied in Kearns et al. (2026), involves a group of learners situated on the vertices of…
Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability
Quantized Vision-Language-Action (VLA) models expose a weight-fault surface: Rowhammer-style faults can corrupt deployed INT8 bits. We pres…
EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation
Accurate boundary delineation remains a persistent challenge in dermoscopic image segmentation because of blurred lesion margins, heterogen…
Spectral Saliency for Machine Unlearning
Machine unlearning (MU) aims to remove the influence of specific training data while preserving model utility. As the name suggests, MU can…
MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration
Programming small social robots from natural-language instructions requires more than invoking isolated APIs. Interactive tasks combine rea…
Amortised Post-Hoc Explanation with Exact Preservation for Dynamic Graph Anomaly Detectors
Anomaly detection in dynamic graphs underpins financial fraud analysis, intrusion detection, and platform integrity, where automated decisi…
Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study
Video question answering systems built on vision-language models often produce timestamped claims with high confidence even when unsupporte…
ARENA: Automated Red-Teaming for Large Audio Language Models
Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but t…
Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair
Industrial software-engineering teams increasingly need LLM agents that turn bug reports into correct patches, yet benchmark-scale operatio…
GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix
Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload ha…
FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy
While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research o…
EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input
The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget. Memory and…
Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabi…
Sparse Prototype Code Underlies Classification and Prediction Across Modalities
Neural representations have become a central tool for studying the internal mechanisms of modern AI models, yet their complex high-dimensio…
Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cos…
When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation
\textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, u…
Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation
Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encode…
When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations
Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations…
PL-Guard: Probabilistic Logic Reasoning for LLM Guardrails
Large language model guardrails can be viewed as policy-consistency problems: a system must determine which policy-relevant facts hold in a…
Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation
Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajec…
Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media
This study presents a hybrid epidemiological and behavioural framework to simulate the spread of health misinformation on social media. We…
Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer
Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose…
RRFC: Recursive Refinement via Feedback Conditioning for Iterative Image-to-Image Generation
Conditional image-to-image generators are single-shot: they map input features to an output in one forward pass and treat it as final, with…
Beyond Single Object: Learning 3D Relations with Large Language Models
We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object c…
FirstDiff: One-Step Diffusion-Based Anomaly Detection for Multivariate Time Series via Initial Noise Prediction
Diffusion models have recently shown strong potential for multivariate time-series anomaly detection by learning the distribution of normal…
Identifying Confusion Trends in Concept-based XAI for Multi-Label Classification
Deep Neural Networks (DNNs) deployed in high-risk domains, such as healthcare and autonomous driving, must be not only accurate but also un…
TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise…
Temporal Graph Prototype-conditioned Conformal Prediction for Fraud Detection
Conformal prediction (CP) provides distribution-free coverage guarantees and has emerged as a principled tool for uncertainty quantificatio…
ALKEMIE Agent: an autonomous platform for computational materials design
Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material…
Decomposing Staleness in Recommender Systems: A Dual-Filter Framework for Supersession and Decay
Stale recommendations are a pervasive challenge and a leading source of user complaints on large-scale content platforms. Items lose releva…
Routing Divergence Is Not Evidence of Behavioral Influence in Same-Weight MoE Self-Distillation
Two Mixture-of-Experts (MoE) forward passes can share every weight yet route the same token through different experts. This creates a possi…
A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations
Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is st…
CardiacMamba: Fair and Robust RGB-RF Fusion for Remote Heart Rate Estimation via State Space Modeling
Remote photoplethysmography (rPPG) enables non-contact heart rate (HR) monitoring from facial videos, but RGB-only methods are vulnerable t…
Characterising cardiac tissue properties with graph neural networks
Characterising electrophysiological properties of cardiac tissue efficiently and accurately from spatially sparse intracardiac measurements…
Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning
Operating household appliances requires long-horizon planning that is state-dependent and robust to disturbances, yet existing large models…
Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes
Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey…
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied envi…
Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive
Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance…
Comprehensive Benchmarking of Deep Learning Architectures for Lung Cancer Histopathology
Lung cancer remains the leading cause of cancer-related mortality worldwide, while histopathological diagnosis is often affected by inter-o…
Pre-training Visual Dexterity in Simulation
Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datase…
Noesis: Bidirectional Graph-RAG with Adaptive Parallelism and Cross-Knowledge-Base Semantic Discovery
Retrieval-Augmented Generation over knowledge graphs (Graph-RAG) has emerged as a powerful paradigm for grounding large language models in…
Information Geometry of Message Passing
We show that the natural-gradient stationary condition of variational inference has an edge-local form on a Forney-style factor graph. We s…
Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong re…
LLMs Get Smarter from Targeted Synthetic Multilingual Data
Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt…
CM-MAE: A Physics-Guided Cross-Modal Self-Supervised Learning Framework for Vision-Wireless Applications
Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a…
A Scalable Pipeline for LLM-Teacher Distillation Labeling: Work-Stealing Job Scheduling and Memory-Aware GPU Concurrency
Labeling large text corpora with LLM teachers has become a practical route to training data at scale. At millions of items, hand-labeling e…
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments.…
Dynamic Evidence Collection Ecosystem for Assessment Integrity and Authentic Competence
Generative Artificial Intelligence (GenAI) can produce high-quality essays, code, and design artefacts, challenging the validity of convent…
RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection
Graph anomaly detection aims to identify nodes that deviate from normal behavioral patterns within graphs. However, existing methods largel…
NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption
Post-hoc Graph Neural Network (GNN) explainers commonly follow a Perturb-Query paradigm, inferring the importance of graph elements based o…
Decoupling Parcellation from Classification: Systematic Benchmark of Fast Brain Segmentation Methods for Alzheimer's Disease Detection
Brain parcellation and classification are typically evaluated in isolation, yet downstream AD detection performance depends on their intera…
Walk Before You Run: The Importance of Data Exploration for Data Analysis Agents
LLM-based data-analysis tools are increasingly used to help users analyze messy spreadsheets and workbooks, from answering questions over u…
CAPO: Constraint-Aware Prompt Optimization for LLM Agents
Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployme…
OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation
Reliable global ocean forecasting is critical for climate monitoring, marine navigation, and extreme event early warning. Physics-based oce…
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. How…
Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner
Capability development is routinely inferred from behavioural thresholds, from final checkpoints, or from what a decoder can read out of a…
AsyTO: Asymmetric Temporal Operator for Parameter-Efficient Multivariate Time Series Forecasting
Multivariate time-series forecasting faces a structural dilemma: sharing one temporal predictor across variables is parameter-efficient but…
RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction
Retrosynthesis is a cornerstone of drug discovery and organic synthesis. While data-driven deep learning models have shown remarkable progr…
TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening
Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health o…
Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System
UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones. In this case, the vision-equipp…
A Tree-Structured Approach for Phishing Template and Attacker Attribution Analysis
Phishing remains a persistent and evolving cybersecurity threat, with attack volumes reaching record levels. This growth is driven by the i…
Digital Twin Degradation: Detecting Cyber Physical Attacks via Temporal Inconsistencies
Digital Twins (DTs) are increasingly used to monitor and analyze Cyber Physical Systems (CPS). However, in adversarial environments, the fi…
Domain-Specific Text Embedding Models for Entity Resolution
General-purpose text embedding models are designed to capture semantic similarity but are not optimised for distinguishing entity records t…
QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents
Large language model (LLM) agents increasingly use external memory systems to support personalization by drawing on long and evolving inter…
Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm
Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional…
Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations
Operational telemetry is predominantly engineered for human reading: systems repeatedly serialize verbose prose, static keys, and redundant…
MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science Systems
Recent advances in large language models have enabled a new class of agentic data science systems that allow users to complete complex data…
Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics
Deep Q-learning (DQL) has achieved remarkable empirical success in reinforcement learning, yet its training process remains notoriously uns…
LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents
LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant e…
Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline
AI-assisted development tools generate vulnerable code at significant rates, yet few automated mechanisms exist to detect, enrich, fix, and…
Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology
Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting,…
HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction
Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions. However…
STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering
By leveraging large-scale pretraining, LLMs can interpret diverse temporal expressions and question formulations without task-specific trai…
A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation
Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative fr…
Software Engineering for AI-driven Building Operation
Building operations are energy-inefficient. Artificial Intelligence (AI)-driven control systems promise benefits through optimization and p…
CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills
Autonomous AI agents tackling Long Horizon Tasks depend on marketplace skills that are certified one at a time: a scanner returns a safety…
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection
The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable…
Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic
Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a…
Decoupled Temporal Encoding for Generative Recommendation
Positional encoding is a fundamental component of Transformer-based generative recommendation models, where user histories are modeled as a…
Audio-Visual Segmentation via Depth-Guided Collaborative Modeling
Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects i…
Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps
Static pruning is widely used to accelerate sparse neural retrieval, yet existing studies each validate their conclusions within a single c…
Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning
Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of vis…
Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026
Recent advances in Generative Artificial Intelligence (GenAI) have substantially improved the ability of large language models (LLMs) to ge…
Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach ha…
SIGMA-Lane: Scale-pyramId Gated MAmba for Temporally Consistent Video Lane Detection
Video lane detection requires predictions that remain stable across frames, yet severe vehicle occlusions can break temporal cues. In strea…
HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals
Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk i…
MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories
Autonomous agents share a transport and can call each other's tools, but they cannot share what they know: no protocol lets two agents' mem…
OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations
Despite comprising over 70\% of its surface, the world's oceans are critically underobserved compared to the land surface or the atmosphere…
Coverage-Maximizing Multinomial Subset Routing under Operational Constraints
We introduce Multinomial Subset Routing (MSR), a new online routing framework over $K$ experts in which the learner keeps a multinomial rou…
Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation
Instance-level lesion detection has been an increasingly larger focal point in medical image segmentation besides the more standard voxel-l…
Synthetic Data Augmentation for Satellite-Based Analysis of Battle-Damaged Agricultural Fields in Ukraine
Monitoring war-induced damage to agricultural land in Ukraine is important for understanding threats to food security, environmental stabil…
Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics
PDF corpora advertise their size in tokens but compute every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) per…
Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs
As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part…
Towards Risk-free AI Agent Deployment
LLM-based agents are rapidly moving from research prototypes into the core business processes of organizations, but these agents pose deplo…
PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data
Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning tr…
Visualizing Uncertainty-to-Action Composition for Human Oversight
Artificial intelligence systems often disclose uncertainty, yet they rarely make clear what response that uncertainty should trigger. Most…
Contrastive Energy Fields for Inference-Time Procedure Planning in Instructional Videos
Procedure planning seeks to estimate a sequence of actions to transition from an observed initial state to a given goal state. Current proc…
A Human-LLM Teaming Framework for Privacy Risk Analysis: An Illustration with CBDC-Based Welfare Schemes
Central Bank Digital Currency (CBDC)-based welfare schemes may be potentially privacy invasive as they process significant volumes of benef…
A Regulatory Placebo? The Systemic Failure of Mandatory GenAI Labeling
We examine the worldwide trend of mandatory labeling of generative artificial intelligence(GenAI) as a reactive, symbolic form of legislati…
A Two-Stage Learning PINN Approach for Solving the Inverse Problem of the 1D Porous Medium Equation
The Porous Medium Equation (PME), given by $u_t = \Delta(u^m)$ for $m > 1$, is a degenerate nonlinear parabolic partial differential equati…
RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning
We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vis…
Graph Machine Learning: An Opportunity for Power Systems
Modern power systems face growing operational complexity driven by the integration of renewable energy sources, decentralization, and the n…
NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment g…
MLLM-Guided Semantic Correction for Text-to-Video Generation
Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, th…
Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal lar…
When Context Misleads: Intent-Guided Decoding for Robust Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) improves large language models by grounding generation in external evidence, but it also introduces a…
Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization
Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows,…
VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience
Agents increasingly rely on reusable skills to encode task knowledge, tool-use procedures, and validation rules. Existing skill self-evolut…
Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity
An accurate estimation of the state of health (SOH) underpins a safe and optimized use of the battery system. Although compelling, data-dri…
Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning
Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixtu…
HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes
Multimodal harmful meme detection is typically formulated as image--text harmfulness classification. A model may correctly predict harmfuln…
When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness
Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in…
Toward Better Assessment of LLMs' Performance in Clinical Error Detection
Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy…
Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents
Satellite agents for on-orbit navigation tasks need to predict collision risks using limited onboard observations. However, conventional pl…
X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization
Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. How…
Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning
Ensembles of decision trees are well-established methods for data stream classification. In ensemble learning, Hoeffding Trees are widely a…
Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL
Direct text-to-SQL asks a language model to do two jobs: interpret the business question and construct the complete relational query. In en…
Bridging the Gap between Labeled and Unlabeled Data via Unified Flow with Feature Memory Bank
Although semi-supervised semantic segmentation ($\text{S}^4$) utilizes abundant unlabeled data to reduce manual labeling burdens, independe…
UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures
Physical AI systems such as autonomous vehicles and robots rely on timely exchange of high-dimensional sensory signals under tight bandwidt…
Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors
Various machine unlearning techniques have been developed in response to privacy legislation requirements, enabling individuals to exercise…
Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors
Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental explo…
MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter
A radiologist reading a model's output faces two problems. The model returns a number and no reason, and any system that turns that number…
Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI
Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognit…
GoalEvolve: From Handcrafted Algorithm Priors to Goal-Driven Evolution of Physical Design Algorithms
Physical design algorithms operate within tightly coupled, multi-stage optimization flows, where stage-local gains may vanish or induce dow…
TDD-Agent: Test-Driven Reasoning for Code Generation
Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level ta…
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But…
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around pred…
Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis
Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decisi…
Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching
Model Inversion Attacks (MIAs) aim to reconstruct representative training samples of target identities from face recognition models, exposi…
Neurosymbolic Embodied Agents
Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate envi…
Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot
Systems that generate scientific research questions are evaluated today by expert scores, LLM-as-judge ratings, or curated case studies --…
UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation
Industrial recommenders rely on two model families that have evolved largely independently: feature-interaction models over multi-field use…
ClawGym II: Exploring Black-Box RL on Agent Harness
Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. Howe…
Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class in…
When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents
Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step reasoning, and code ge…
CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?
Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score…
Model Hypnosis: Strong control of AI via additive subliminal effects
We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irre…
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) fo…
Proteus: Incremental Memory Activation for Long-Context Sequence Modeling
The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models tha…
Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text
A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computat…
AutoSR: Automatic Symbolic Regression by Searching Research States
We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searc…
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve
The current best bounds on the matrix multiplication exponent $\omega$ are obtained through a refinement of the laser method called combina…
Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory
Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) models increasingly…
mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA
Advanced Multimodal Large Language Models (MLLMs) struggle with recent Knowledge-based Visual Question Answering (VQA) tasks, such as INFOS…
Evidence of conceptual mastery in the application of rules by Large Language Models
In this paper we leverage psychological methods to investigate LLMs' conceptual mastery in applying rules. We introduce a novel procedure t…
SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling
Retrieval-Augmented Generation (RAG) and its Multimodal Retrieval-Augmented Generation (MRAG) significantly improve the knowledge coverage…
Calibrated Generative AI as Meta-Reviewer: A Systemic Functional Linguistics Discourse Analysis of Reviews of Peer Reviews
This study investigates the use of generative AI to support formative assessment through machine generated reviews of peer reviews in gradu…
The Fragility of Strategic Thinking in Large Language Models
Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation,…
Budget-Aware Tool Use Enables Effective Agent Scaling
Scaling test-time computation has been extended from language model reasoning to tool-augmented agents, where scaling involves not only thi…
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration
Medical calculators are fundamental to quantitative, evidence-based clinical practice. However, their real-world use is an adaptive, multi-…
Agentic Test-Time Scaling for WebAgents
Test-time scaling has become a standard way to improve performance and boost reliability of neural network models. However, its behavior on…
The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents
Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources. However, thes…
ML-AutoResearch: Training Machine Learning Research Agents with Automatically Generated Environments
With the advent of AI agents, automated scientific discovery is becoming an increasingly plausible goal. However, training agents to autono…
FactReview: 実行ベースの請求検証を備えた証拠に基づくピアレビュー
LLM ベースのレビュー システムは通常、原稿のみを入力として受け取り、文献やコード ベースの主張を検証することが困難になります。私たちは、レビュー関連のクレームを抽出し、関連する作業に基づいて根拠を示し、コードが利用可能な場合は、固定の修復予算の下でリリースされた成果物を実行して、経験的なクレームを監査するシステムである FactReview を紹介します。 35 件の ML 論文と 463 件のベンチマークの主要な主張にわたって、FactReview は主張の 84% をカバーしています。証拠を意識したルーブリックに基づいて、そのレビューの全体的な品質のスコアは 4.86/5 で、DeepReview-v2 より 0.7 上、OpenReview のコメントと一致するものより 1.5 上です。執行証拠を削除すると、請求ステータスの 17% が変更され、これは他の単一の証拠ソースよりも多くなります。審査員支援調査では、FactReview は平均審査時間を 58% 短縮し、ベンチマークの請求範囲を 87% から 99% に高めました。私たちは、LLM の査読者は、受諾/拒否の決定を下すのではなく、経験に基づいた主張を監査すべきであると主張します。コードは https://github.com/DEFENSE-SEU/FactReview で公開されています。
原文 (English)
FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification
Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leaving literature- and code-dependent claims difficult to verify. We present FactReview, an audit pipeline that extracts review-relevant claims, grounds them in related work and reference checks, and, when code is available, executes released artifacts under a fixed repair budget. On 26 paper-disjoint test papers with 354 human-verified claims, FactReview achieves 84.3\% F1 for claim recovery. In a same-backend, evidence-matched comparison, FactReview scores 4.72/5 overall, outperforming a direct LLM reviewer by 0.74 points. Removing execution evidence changes 17.0\% of claim statuses, more than removing any other single evidence source. In a reviewer-assistance study, FactReview reduces mean review time by 58\% while increasing benchmark-claim coverage from 87\% to 99\%. FactReview supports evidence-based claim auditing, with acceptance decisions reserved for human reviewers. The code is public at https://github.com/DEFENSE-SEU/FactReview.
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes.…
An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing
In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC). This joint ope…
SAPO: Step-Aligned Policy Optimization for Reasoning-Based Generative Recommendation
Generative recommendation treats next-item prediction as autoregressive item-identifier generation. Specifically, items are encoded as sema…
BrickAnything: 構造を意識したトークン化を使用した、ジオメトリ条件付きの構築可能なレンガの生成
3D 形状から物理的に構築可能なレンガ構造を生成するには、幾何学的再構成以上のものが必要です。出力は、個別のパーツの制約と構造の安定性も満たさなければなりません。既存のレンガ生成方法は、ターゲットの 3D 形状が事前定義された制約の下で実現可能な構造を許容しない場合に機能不全に陥る可能性があるヒューリスティック最適化に依存しているか、基礎となる 3D ジオメトリとアセンブリ関係を明示的にモデル化せずにブリック シーケンスを生成しています。この研究では、さまざまな 3D 表現から構築可能なレンガ構造を生成するための、ジオメトリ条件付き自己回帰フレームワークである BrickAnything を紹介します。 BrickAnything は、統一された幾何学的インターフェイスとして点群を使用し、アセンブリ制約の下でターゲット形状を再構築するレンガ シーケンスを予測します。ブリック間の構造依存関係をモデル化するために、ローカル接続関係を通じてブリック構造を表す構造認識ツリー トークン化を導入します。この定式化により、シーケンスの生成と物理的な構築プロセスの一貫性が高まり、無効な中間状態が減少します。さらに、安定性や幾何学的忠実度などの構築性の目標を向上させるために、トレーニング後の好みに基づくアライメント、妥当性制約のあるデコード、および適応的ロールバックを導入します。広範な実験により、BrickAnything が幾何学的に忠実で物理的に実現可能なブリック構造を生成すること、および提案されたトークン化により従来の順序付け戦略と比較してロールバックと再生成が効果的に削減されることが実証されました。
原文 (English)
BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization
Generating physically buildable brick structures from 3D shapes requires more than geometric reconstruction: the output must also satisfy discrete part constraints and structural stability. Existing brick generation methods either rely on heuristic optimization, which can break down when the target 3D shape does not admit a feasible structure under predefined constraints, or generate brick sequences without explicitly modeling the underlying 3D geometry and assembly relations. In this work, we present BrickAnything, a geometry-conditioned autoregressive framework for generating buildable brick structures from diverse 3D representations. BrickAnything uses point clouds as a unified geometric interface and predicts brick sequences that reconstruct the target shape under assembly constraints. To model structural dependencies among bricks, we introduce a structure-aware tree tokenization, which represents brick structures through local attachment relations. This formulation makes sequence generation more consistent with the physical construction process, and reduces invalid intermediate states. We further introduce preference-based alignment post-training, validity-constrained decoding and adaptive rollback to improve buildability objectives such as stability and geometric fidelity. Extensive experiments demonstrate that BrickAnything produces geometrically faithful and physically realizable brick structures, and that the proposed tokenization effectively reduces rollback and regeneration compared with conventional ordering strategies.
想像力の知覚トークンはマルチモーダル言語モデルの空間推論を強化します
ビジョン言語モデル (VLM) は多くのタスクに優れていますが、重要な情報が直接観察できない場合には空間推論に依然として苦労します。このような問題の多くは、目に見えない視点から何が見えるかを推測したり、遮蔽された空間を通る経路を追跡したり、部分的な観察を一貫した空間表現に統合したりするなど、想像力豊かな認識を必要とします。観察された入力との一貫性を保ちながら、代替の空間構成の下で VLM が知覚するものを外部化する中間的な知覚表現である想像的知覚トークン (IPT) を導入します。この機能を研究するために、透視図法取得 (PET)、パス トレーシング (PT)、およびマルチビュー カウンティング (MVC) という 3 つのタスクを定式化し、グラウンド トゥルースの想像力、回答、評価ベンチマークを含む約 20,000 例のデータセットを構築します。統合された VLM BAGEL をバックボーンとして使用することで、IPT 監視は空間推論を一貫して改善し、推論時に画像を生成しなくても、テキストによる思考連鎖トレーニングを上回ることがよくあります。 MVC では、IPT は精度を 3.4% 向上させ、PT 上の強力なクローズドソース モデルにより競争力のあるパフォーマンスを実現します。さらに、IPT とラベルのみの監視を組み合わせるとさらなる利益が得られる一方、テキストの思考連鎖はパフォーマンスを大幅に低下させる可能性があることがわかり、空間計算が言語を通じて強制される場合にはモダリティの不一致が示唆されます。全体として、IPT は、観察されていない空間構造について推論するための原則に基づいた監視信号を提供し、解釈可能な中間表現を生成しながら一般化を向上させます。
原文 (English)
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems require imaginative perception: inferring what would be seen from an unseen viewpoint, tracing paths through occluded spaces, or integrating partial observations into a coherent spatial representation. We introduce Imaginative Perception Tokens (IPT), intermediate perceptual representations that externalize what a VLM would perceive under alternative spatial configurations while remaining consistent with the observed input. To study this capability, we formulate three tasks, Perspective Taking (PET), Path Tracing (PT), and Multiview Counting (MVC), and construct datasets of approximately 20K examples with ground truth imaginations, answers, and evaluation benchmarks. Using the unified VLM BAGEL as the backbone, IPT supervision consistently improves spatial reasoning and often outperforms textual chain of thought training, even without generating images at inference time. On MVC, IPT improves accuracy by 3.4% and achieves competitive performance with strong closed-source models on PT. We further find that combining IPT and label-only supervision yields additional gains, whereas textual chain of thought can substantially degrade performance, suggesting a modality mismatch when spatial computation is forced through language. Overall, IPT provides a principled supervision signal for reasoning about unobserved spatial structure, improving generalization while producing interpretable intermediate representations.
異種鉄道システムにおける中断を考慮した動的ルート最適化のための時間計画フレームワーク
効率的なルートの最適化は、鉄道運行の安全性と定時性の両方を確保する上で重要な役割を果たします。これは、列車の速度、停止パターン、インフラストラクチャの互換性の制約が変化し、調整が複雑になる異種の多軌間鉄道ネットワークでは特に非常に重要です。単線システムでは、すべての列車が同じ線路を共有し、頻繁な線路切り替えが必要となるため、これらの課題はさらに深刻になります。線路の封鎖、列車の封鎖、エンジン故障、速度低下などの確率的混乱により、運行にさらなる予測不可能性が生じ、時刻表が狂います。しかし、既存の研究は主に高レベルの時刻表に焦点を当てており、線路切り替えの調整などの運用の詳細は省略されています。その結果、人間の運転手に判断が委ねられることになり、鉄道運行における安全上のリスクが増大します。この研究は、異種鉄道システムにおける動的なルート最適化と混乱管理のための時間計画に基づくフレームワークを提案します。このフレームワークは、PDDL 2.1 を使用して鉄道運行を時間計画問題として定式化し、ゲージ互換性の制約と多様な混乱シナリオを明示的にモデル化します。最適化されたスケジュールと実行可能なアクション シーケンスの両方を指定する、競合のないタイムスタンプ付きの運用計画を生成します。提案されたフレームワークを評価するために、最大 1,000 のトラック ポイントと 120 の列車を使用する 200 のインスタンスを含むベンチマーク問題セットを開発しました。フレームワークの評価には、2 人の最先端の時間プランナーと 1 人の計画検証者が採用されました。実験結果は、このフレームワークが異種鉄道システムの一時的な運行計画を効果的に生成し、複数ゲージの制約や混乱に対処し、手動の意思決定への依存を軽減することを示しています。
原文 (English)
A Temporal Planning Framework for Disruption Aware Dynamic Route Optimization in Heterogeneous Railway Systems
Efficient route optimization play a vital role in ensuring both safety and punctuality in railway operations. It is very crucial particularly in heterogeneous multi-gauge railway networks with varying train speed, stopping pattern, infrastructure compatibility constraints increase coordination complexity. In single-track systems these challenges are further intensify due to all trains to share the same track and requires frequent track switching.Stochastic disruptions events including blocked tracks, blocked trains, engine failure and speed slowdowns introduces additional unpredictability in operations and deviate the timetable. However, existing studies predominantly focuses on high-level timetabling, omitting operational details such as track switching coordination. As a result leaving decision to human operators, increasing safety risks into railway operations. This study proposes a framework based on temporal planning for dynamic route optimization and disruption management in heterogeneous railway systems. The framework formulates railway operations as a temporal planning problem using PDDL 2.1 with explicitly modeling gauge compatibility constraints and diverse disruption scenarios. It generates conflict-free timestamped operational plans specifying both optimized schedules and executable action sequences. To evaluate the proposed framework, we developed a benchmark problem set with 200 instances using up to 1,000 track points and 120 trains. Two state-of-the-art temporal planners and a plan validator were employed to assessed the framework. The experimental results demonstrate that the framework effectively generates temporal operational plans for heterogeneous railway systems and handles multi-gauge constraints, disruptions, and reduces dependence on manual decision making.
機械学習された併存疾患指数
従来の併存疾患スコア (Charlson および Elixhauser など) は、リスク調整や患者の層別化に広く使用されていますが、2 つの重要な制限があります。(i) それらは主に死亡率中心であり、他の臨床転帰とうまく一致しません。(ii) 線形でルールに基づいた構造では、非線形で転帰固有のリスク関係を捉えることができません。我々は、学習されたスコアと複数の臨床転帰の間の正規化されたヒルベルト・シュミット独立基準(nHSIC)を最大化することにより、診断コードを単一のスカラーにマッピングする機械学習併存疾患指数(MLCI)を提案します。 MLCI は、リスクと結果の非線形依存性を捉えており、統合された有益な入院レベルの順序付けが複数の結果にわたっていつ達成されるかを特徴づける理論によってサポートされています。複数のベンチマーク電子医療記録 (EHR) データセットに関する実証結果は、MLCI が複数の評価指標全体で強力なベースラインを上回るパフォーマンスを示していることを示しています。
原文 (English)
A Machine-Learned Comorbidity Index
Traditional comorbidity scores (e.g., Charlson and Elixhauser) are widely used for risk adjustment and patient stratification, but they have two key limitations: (i) they are largely mortality-centric and do not align well with other clinical outcomes, and (ii) their linear, rule-based structure cannot capture nonlinear, outcome-specific risk relationships. We propose a Machine-Learned Comorbidity Index (MLCI) that maps diagnosis codes to a single scalar by maximizing the normalized Hilbert-Schmidt Independence Criterion (nHSIC) between the learned score and multiple clinical outcomes. MLCI captures nonlinear risk-outcome dependence and is supported by a theory that characterizes when a unified, informative admission-level ordering can be achieved across outcomes. Empirical results on multiple benchmark electronic health record (EHR) datasets show that MLCI outperforms strong baselines across multiple evaluation metrics.
Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries
AI agents now act as first-class members of the software development lifecycle, but the instruments teams use to direct them enforce nothin…
PolyWorkBench: 長期にわたる多言語 LLM エージェントのベンチマーク
大規模言語モデル (LLM) エージェントは、計画、ツールの使用、外部環境との対話を必要とする長期的なタスクで優れたパフォーマンスを示しています。ただし、既存のベンチマークのほとんどは、推論、ツールの呼び出し、出力生成を含む実行プロセス全体が単一言語内で実行される単一言語設定を暗黙的に前提としています。対照的に、現実世界のアプリケーションでは、統合されたワークフロー内で多言語の入力と出力が関与することがよくありますが、多言語性とエージェント実行の間の相互作用はまだ十分に解明されていません。この作業では、多言語の長期的な職場ワークフローで LLM エージェントを評価するためのベンチマークである PolyWorkBench を紹介します。 PolyWorkBench は、コマース、ナレッジ ワーク、法的分析、ローカリゼーション、製造を含む 5 つのドメインにわたる 67 のタスクで構成されており、エージェントは異種多言語入力を処理し、反復推論を実行し、外部ツールを呼び出し、構造化された出力を生成する必要があります。包括的な評価を可能にするために、構造グレーディング、実行可能検証、LLM ベースのセマンティック評価を組み合わせたハイブリッド フレームワークを提案します。この設計により、複雑なワークフロー全体で機能の正確さと言語の一貫性の両方を取得できるようになります。経験的な結果によると、最先端の LLM エージェントは、単言語のワークフロー設定に比べて、多言語のワークフロー設定ではパフォーマンスが大幅に低下します。私たちの分析は、多言語が推論と実行のステップ全体に複合的な影響をもたらすことを示唆しており、エージェントの評価における言語のバリエーションと手続き上の意思決定を共同でモデル化することの重要性を強調しています。
原文 (English)
PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows
While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual resources across extended trajectories. The interaction between multilinguality and long-horizon execution, however, remains underexplored. We introduce PolyWorkBench, a benchmark designed to evaluate LLM agents on multilingual, long-horizon workplace workflows. PolyWorkBench features 67 tasks across five core domains: commerce, knowledge work, legal analysis, localization, and manufacturing. Tasks are authored by the paper's authors from real-world data seeds and independently verified through a second-author audit. Agents must integrate heterogeneous multilingual inputs, execute iterative tool-use trajectories, and produce structured domain artifacts. To rigorously assess performance, we adopt Grade, a task-specific structural scoring rubric, as our primary ranking metric, and complement it with Pytest for executable state verification and LLM-as-Judge for semantic quality diagnostics. Benchmark evaluations reveal that agent performance varies substantially across languages and drops sharply on the harder cross-lingual tasks, and our analysis shows that multilingual execution exposes systematic failure modes across planning, tool interaction, and decision-making in long-horizon agents.
Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns
Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language mode…
Geometric Self-Supervised Pre-training for Neural Combinatorial Optimization
Neural Combinatorial Optimization (NCO) techniques have emerged as a highly efficient alternative to traditional exact algorithms for solvi…
Where did the ambiguity go? Examining how multimodal models interpret polysemous words
Human language is highly polysemous. Many common words (e.g., "bank" or "palm") carry several distinct meanings that shape what humans comm…
DiffImaginE: Diffusio を使用してエンティティ タイプを検証することを想像してください
マルチモーダル名前付きエンティティ認識 (MNER) は、各候補スパンとエンティティ タイプの仮説が共同のテキスト証拠と視覚的証拠によってサポートされているかどうかを判断します。既存の想像比較検証器は、各 (スパン、タイプ) ペアを 1 つの予測された視覚的特徴にマッピングし、多様な視覚的実現を単一のプロトタイプに圧縮し、明示的な確率的セマンティクスを使用せずに互換性スコアを提供します。 MNER 型検証を条件付き潜在拡散推論として定式化する DiffImaginE を紹介します。スパン局所化された視覚的証拠が与えられると、タイプ条件付きデノイザーは、標準化された潜在に注入されるノイズを予測します。結果として生じるノイズ除去誤差は、タイプ条件付き負の対数尤度の ELBO 一貫性のある代用値を提供し、競合するタイプの仮説を、観察をどの程度うまく説明できるかによってランク付けできるようにします。 DiffImaginE は、標準のマルチモーダル エンコーダ スタックを保持し、決定論的検証器を、Min-SNR 重み付けを使用してトレーニングされた分類子なしのガイド付き拡散スコアラーに置き換えます。タイプごとの拡散スコアを分類ロジットとして直接監視し、ノイズ レベル全体の集計を学習し、逆サンプリングを使用してモンテカルロ比較の分散を削減します。私たちの分析は、分類器を使用しないガイダンスが誘導型事後分布を鮮明にし、反対のペアリングが等しいデノイザーコストで分散を低減するときの特徴を示すことを示しています。 Twitter-2015 と Twitter-2017 の実験では、アブレーションと一対の有意性検定によってサポートされ、同じエンコーダー、補助対物レンズ、評価プロトコルの下で、一致した決定論的 ImaginE 制御に対して一貫したゲインが示されています。
原文 (English)
DiffImaginE: Imagine to Verify Entity Types with Diffusion
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type) pair to one predicted visual feature, compressing diverse visual realisations into a single prototype and providing a compatibility score without explicit probabilistic semantics. We introduce DiffImaginE, which formulates MNER type verification as conditional latent diffusion inference. Given span-localised visual evidence, a type-conditioned denoiser predicts noise injected into its standardised latent. The resulting denoising error provides an ELBO-consistent surrogate for type-conditional negative log-likelihood, allowing competing type hypotheses to be ranked by how well they explain the observation. DiffImaginE retains a standard multimodal encoder stack and replaces the deterministic verifier with a classifier-free-guided diffusion scorer trained using Min-SNR weighting. We directly supervise per-type diffusion scores as classification logits, learn aggregation across noise levels, and use antithetic sampling to reduce Monte Carlo comparison variance. Our analysis shows that classifier-free guidance sharpens the induced type posterior and characterises when antithetic pairing reduces variance at equal denoiser cost. Experiments on Twitter-2015 and Twitter-2017 show consistent gains over a matched deterministic ImaginE control under the same encoder, auxiliary objectives, and evaluation protocol, supported by ablations and paired significance tests.
RA-CAD: ステートアウェアなテキストから CAD への生成のための実行後の批判の学習
テキストから CAD への生成は、自然言語の設計意図を編集および実行可能なパラメトリック コンピューター支援設計 (CAD) コードに変換し、手動モデリングに必要な専門知識と労力を軽減します。既存の方法には、生成プロセスを最適化するために、固定された、外部から提供される、プロンプト誘導される、または個別に最適化された批評メカニズムが組み込まれていますが、生成プロセス全体を通じてフィードバックがどのように解釈され、効果的な是正措置に変換されるかを必ずしも最適化しているわけではありません。このフィードバック利用のギャップを埋めるために、生成、実行、批判、書き換えのループを通じて CAD 環境と対話する状態認識エージェントである RA-CAD (ReAct Agent for CAD) を紹介します。各反復で、RA-CAD は現在のコードを実行し、その結果を観察します。設計指示、現在のコード、および実行フィードバックを条件として、エージェントは中間ポリシー アクションとして明示的な実行後の批評を生成します。この批評は、現在の結果の終了を検証するか、次の書き換えの条件となるリビジョン指向のガイダンスを提供します。 CAD コード ブートストラップ (CCB) は、まず、監視付き微調整を通じて基本的なパラメトリック CAD コーディング機能を確立します。その後、フィードバック駆動エージェント最適化 (FAO) は、軌道レベルのグループ相対ポリシー最適化をポリシー生成コードと批評シーケンスの両方に適用し、終端 F1 と面取り距離の報酬を完全なインタラクション軌道に割り当てます。この定式化により、批評は最適化されていない補助的な出力ではなく、結果に合わせた学習可能な政策決定となります。 CADFusion と Text2CAD の実験では、既存の方法や強力な独自言語モデルと比較して、RA-CAD が最先端の実行妥当性と幾何学的品質を達成していることが示され、提案されている状態認識型テキストから CAD エージェントの有効性が実証されています。
原文 (English)
RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation
Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes, reducing the expertise and effort required for manual modeling. Existing methods incorporate fixed, externally supplied, prompt-induced, or separately optimized critique mechanisms to optimize the generation process, but they do not necessarily optimize how feedback is interpreted and translated into effective corrective actions throughout the generation process. To bridge this feedback-utilization gap, we present RA-CAD (ReAct Agent for CAD), a state-aware agent that interacts with the CAD environment through a Generate--Execute--Critique--Rewrite loop. At each iteration, RA-CAD executes the current code and observes its outcome. Conditioned on the design instruction, current code, and execution feedback, the agent then generates an explicit post-execution critique as an intermediate policy action. This critique either validates the current result for termination or provides revision-oriented guidance that conditions the next rewrite. CAD Code Bootstrapping (CCB) first establishes fundamental parametric CAD coding capabilities through supervised fine-tuning. Feedback-Driven Agent Optimization (FAO) subsequently applies trajectory-level Group Relative Policy Optimization to both policy-generated code and critique sequences, assigning terminal F1 and Chamfer Distance rewards to the complete interaction trajectory. This formulation makes critique an outcome-aligned, learnable policy decision rather than an unoptimized auxiliary output. Experiments on CADFusion and Text2CAD show that RA-CAD achieves state-of-the-art execution validity and geometric quality compared with existing methods and strong proprietary language models, demonstrating the effectiveness of the proposed state-aware text-to-CAD agent.
異種アテンションメモリのランタイムオブザーバビリティ
最新のモデルはプレーンな KV キャッシュを保持しなくなりました。潜在キャッシュ、学習されたスパース セレクター、および再帰状態はそれぞれ異なる形式でモデルのメモリを保持し、圧縮時にそれぞれ異なる方法で失敗します。 3 つの演算子で 4 つのメモリ クラスすべてをカバーするランタイム可観測性コントラクトを与え、5 つのアーキテクチャ ファミリにわたる 6 つのモデル構成でインスタンス化し、ステージごとの境界を実行可能なリクエスト レベルのリスク台帳に構成します。コントラクトはエラー メトリックをタイプとして保持します。合成はメトリックが一致する場合にのみ定義され、このチェックでは最初に合成されたチェーンが拒否されました。修復されたチェーンは 2 つの証明されたブリッジを介してメトリクスを交差し、正式なシステムが証明できないものはすべて代わりに測定され、合成された層が自動的に経験的なレベルに落とされます。すべてのクレームは認定、部分的に認定、または経験的であり、合成は最も弱い層を継承し、層はマシンによって決定されます。 $12.4$M を超えるエントリの読み取りが再生され、リクエストごとの予算とフェイルクローズされた ID 帰属を備えた 8 方向の同時実行の下で実行され、台帳は今日の証人に対する正直なトレードオフを定量化し、違反ゼロでリスク バジェットを維持します。融合された常時オン プローブは、サービング ノイズ フロア内の CUDA グラフの下で宣言された 1 層のサブセットを観察します。パックされた圧縮 KV プロトタイプを備えたサービス済みの DeepSeek-V4 スタックに適用された同じ機械は、機械が判断した差別キャンペーンを通じて、サイレント破損を正確な構造境界に特定します。これは、立ち退きのない、個人情報が分離された体制で正確であり、立ち退きまたはスロット再利用体制で観察されたすべての失敗を伴います。その計算は、途中で私たち自身の混乱した推論の 2 つを拒否しました。すべてのアーティファクト、ガード、リーン開発は https://github.com/metask-ai/witprobe-attention-memory でリリースされます。このペーパーのすべての番号は、出荷されたアーティファクトから 1 つのコマンドで再生成されます。
原文 (English)
Runtime Observability for Heterogeneous Attention Memory
Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression. We give a runtime observability contract that covers all four memory classes with three operators, instantiate it on six model configurations across five architecture families, and compose the per-stage bounds into an executable request-level risk ledger. Contracts carry their error metric as a type -- composition is only defined when metrics match, and this check rejected our own first composed chain; the repaired chain crosses metrics through two proved bridges, and whatever no formal system can certify is measured instead, dropping the composed tier to empirical automatically: every claim is certified, partially certified, or empirical, composition inherits the weakest tier, and the tier is decided by the machine. Replayed over $12.4$M entry reads and run under eight-way concurrency with per-request budgets and fail-closed identity attribution, the ledger quantifies the honest trade-off on today's witness and holds its risk budget with zero violations. A fused always-on probe observes a declared one-layer subset under CUDA graphs inside the serving noise floor. Applied to a served DeepSeek-V4 stack with a packed compressed-KV prototype, the same machinery localizes a silent corruption to a precise structural boundary -- exact in the eviction-free, identity-isolated regime, with every observed failure in an eviction or slot-reuse regime -- through a machine-adjudicated discrimination campaign whose calculus rejected two of our own confounded inferences along the way. All artifacts, guards, and the Lean development are released at https://github.com/metask-ai/witprobe-attention-memory; every number in this paper regenerates from the shipped artifacts by one command.
GSBF: 環境を考慮したビームフォーミングのためのガウス スプラッティング
ビームフォーミングは、多入力多出力 (MIMO) 通信システムにおいて重要な役割を果たします。ただし、従来のビームフォーミング設計では通常、正確な瞬間チャネル状態情報 (CSI) と反復的な最適化が必要であり、これによりパイロットのオーバーヘッドと計算の複雑さが大幅に増加します。無線伝播は本質的に物理幾何学によって支配されることを認識し、マルチモーダル データに基づいて環境を考慮したビームフォーミング (GSBF) パイプライン用の 3D ガウス スプラッティングを開発します。これは、永続的な 3D ガウス表現を通じて環境を特徴付けます。具体的には、GSBF は相反性を保持する双方向球面ガウス (Bi-SG) カーネルを使用して環境散乱応答をモデル化し、両面電磁ラスタライゼーションを実行して角度プロパゲータ マップをレンダリングします。次に、レンダリングされたマップは、過剰に完成した配列多様体辞書を通じて集約され、定弾性ビームフォーマーに投影されます。これにより、オンラインの瞬間的な CSI を使用せずに、アクセス ポイント (AP) の姿勢とユーザーの位置から直接ビームが合成されます。シミュレーションでは、GSBF が一貫して網羅的ビーム アライメント (EBA) などのベースラインを上回り、遅延が低いことが実証されています。
原文 (English)
GSBF: Gaussian Splatting for Environment-Aware Beamforming
Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires accurate instantaneous channel state information (CSI) and iterative optimization, which incur substantial pilot overhead and computational complexity. Recognizing that radio propagation is intrinsically governed by the physical geometry, we develop a 3D Gaussian splatting for environment-aware beamforming (GSBF) pipeline based on multi-modal data, which characterizes the environment through a persistent 3D Gaussian representation. Specifically, GSBF models the environmental scattering response with reciprocity-preserving bidirectional spherical Gaussian (Bi-SG) kernels and performs two-sided electromagnetic rasterization to render an angular propagator map. The rendered map is then aggregated through an over-complete array-manifold dictionary and projected to the constant-modulus beamformers, thereby synthesizing beams directly from the access point (AP) pose and user position without online instantaneous CSI. Simulations demonstrate that GSBF consistently outperforms baselines such as exhaustive beam alignment (EBA) with lower latency.
Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies fo…
Context Is Not Authority: Structured Runtime Governance for Financial Market Agents
Financial agents can turn correct context into an unauthorized effect: a customer-facing commitment, trade, or deployed policy. We present…
AI エージェントに関する行動科学研究の自動化と拡張
AI エージェントが複雑な環境に導入されることが増えるにつれ、AI エージェントの動作を理解することが重要になります。しかし、AI エージェントに関する行動科学研究は依然として手作業で労働集約的です。 AI エージェントの行動科学研究を自動化する初のマルチエージェント システムである AEROBAT を紹介します。ユーザーが任意のターゲット行動を与えると、AEROBAT は行動科学研究の完全なパイプラインを自動的に実行します。つまり、行動に関する仮説の生成、対照実験の設計と実行、行動の評価、結果の分析、レポートの作成です。 12 のターゲット行動について、AEROBAT を使用して 79 の仮説を生成およびテストしました。つまり、1,240 の制御された実験を設計し、合計 23,512 のシミュレーション ラウンドを実行しました。いくつかの新しい仮説を含む 26 の仮説について、中程度から強力な統計的証拠が見つかりました。要約すると、私たちの結果は、AI エージェントに関する自動行動科学研究が手動研究を補完し、範囲を拡大できることを示しています。
原文 (English)
Automating and Scaling Behavioral Scientific Research on AI Agents
As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains manual and labor-intensive. We introduce AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents. Given an arbitrary target behavior by its user, AEROBAT automatically executes a full pipeline of behavioral scientific research---generating hypotheses about the behavior, designing and executing controlled experiments, making behavioral assessments, analyzing the results, and writing reports. For 12 target behaviors, we used AEROBAT to generate and test 73 hypotheses: designing 1,160 controlled experiments and executing 22,954 simulation rounds in total. Moderate-to-strong statistical evidence was found for 30 hypotheses, including some novel ones. In sum, our results demonstrate that automated behavioral scientific research on AI agents can complement and extend the reach of manual research.
SkillZip: 再利用可能な構造の発見による自己進化エージェントの評価不要のスキル圧縮
自己進化するエージェントは、成功した手順や失敗の修正を追加することで、再利用可能なスキルを蓄積します。時間が経つと、同じ要件が複数の分岐、例、警告で再度説明されることが多くなり、共通のアクション シーケンスは再利用されずにコピーされます。結果として得られるスキルは注入にコストがかかり、維持するのが困難になります。スキルは平坦なパッセージではないため、一般的なプロンプト圧縮はこの設定には適していません。スキルの名前と説明はいつ適用されるかを定義し、ワークフローは実行を制御し、ツールと出力コントラクトは有効性を制約し、まれな例外は、それらをアクティブ化するサンプル タスクがない場合でも必須のままである可能性があります。評価ガイド付き圧縮ではこれらの動作をテストできますが、ロールアウト、コスト、および圧縮時間の評価セットへの依存が生じます。私たちは、最短で忠実な構造的説明を見つけることによってスキルを圧縮する、評価不要のメソッドである SkillZip を紹介します。直感は一度説明し、多くを参照します。繰り返しのルールを適用範囲で一度だけ記述し、繰り返しのアクション シーケンスを共有プロシージャに要素化し、相違点のみを明示的な例外として保持します。この直感を、抽出されたすべてのトリガー、ワークフロー エッジ、ツール要件、義務、および出力フィールドに対するハード カバレッジ制約の対象となる、スキル契約および残余に関する型付きの最小記述長目標として形式化します。この定式化は、単純な共有しきい値を提供し、構築により固有のまれなルールを保存し、効率的なローカル更新をサポートします。 SkillZip には、1 つの構造化された抽出呼び出しと決定論的な最適化を備えたワンショット モードと、タスクの再生や完全な履歴の再解析を行わずに各自己進化パッチを統合する継続的な Zip-on-Write モードがあります。包括的な実験評価を通じて、圧縮パフォーマンス、汎用性、コストオーバーヘッドにおける SkillZip の有効性と優位性を実証します。
原文 (English)
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than reused. The resulting skill becomes expensive to inject and difficult to maintain. Generic prompt compression is ill-suited to this setting because a skill is not a flat passage: its name and description define when it applies, its workflow controls execution, its tool and output contracts constrain validity, and rare exceptions may remain essential even when no sampled task activates them. Evaluation-guided compression can test these behaviors, but it introduces rollouts, cost, and dependence on the compression-time evaluation set. We present SkillZip, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation. The intuition is explain once, reference many: state a repeated rule once at the scope where it applies, factor a repeated action sequence into a shared procedure, and keep only the differences as explicit exceptions. We formalize this intuition as a typed minimum description-length objective over a skill contract and a residual, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field. The formulation provides simple sharing thresholds, preserves unique rare rules by construction, and supports efficient local updates. SkillZip has a one-shot mode with one structured extraction call and deterministic optimization, and a continual Zip-on-Write mode that integrates each self-evolution patch without replaying tasks or reparsing the full history. Through comprehensive experimental evaluations, we demonstrate the effectiveness and superiority of SkillZip in compression performance, generalizability, and cost overhead.
AutoWorldModel-Bench: 自動化されたワールドモデル研究のためのステート中心のベンチマーク
ワールド モデリングは未解決の分野です。アーキテクチャ、トレーニング目標、状態表現は複雑な方法で相互作用しており、単一のレシピが環境全体を支配することはありません。これは、自律的な研究者として機能する AI コーディング エージェントにとって理想的なテストベッドになります。現在のエージェント ベンチマークを支配する仕様に合わせたエンジニアリング タスクとは異なり、改善の方向性が事前に指定されていない設定です。 AutoWorldModel-Bench は、フロンティア コーディング エージェントが固定のコンピューティング バジェットの下で提供されたワールド モデル スターターを自律的に改善する閉ループ ベンチマークです。このベンチマークは、統一された構造化状態表現 (各ゲームから抽出され、共有テンソル形式を通じて消費されるグラウンドトゥルース エンティティ状態) の下で 8 つのゲーム環境にまたがります。これにより、ダイナミクス モデリングが認識から分離され、実行あたりの反復を数分で行うことが可能になります。 64 回のセッションにわたって、Codex-5.4 と Claude Opus 4.6 は 63 回のセッションでのスターターを改善しました。セッションの 91% で、成功した編集は、ハイパーパラメータの調整ではなく、新しい目的、表現、ロールアウト手順、またはアーキテクチャの変更など、重要なリサーチ スタイルの変更です。私たちのベンチマークは、仕様に合わせたエンジニアリングの問題ではなく、オープンエンドの研究に基づいてフロンティア コーディング エージェントを評価できる設定を提供します。
原文 (English)
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement direction is not specified in advance, unlike the engineering-to-spec tasks that dominate current agent benchmarks. We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided base world model under a fixed compute budget. The benchmark spans eight game environments under a unified structured-state representation--ground-truth entity state extracted from each game and consumed through a shared tensor format--which isolates dynamics modeling from perception and enables minutes-per-run iteration. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improve their base on a held-out test split in all but one session, with about half (33 of 64) a substantial gain ($\Delta \geq +0.10$) and the remaining improvements smaller but positive; in 91% of sessions the winning edit is a substantive change to the model or training rather than a hyperparameter tweak. Our benchmark offers a setting in which frontier coding agents can be evaluated on open-ended research rather than engineering-to-spec problems.
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs
Self-consistency via majority vote reduces per-problem accuracy on most GPQA Diamond problems for small instruction-tuned models: 56.6% of…
デュアルフロートランスフォーマー: プライマリプレフィルパスを追加のデコード計算から切り離す
大規模な言語モデルがより多くのリクエストに対応するにつれて、累積推論コストが 1 回限りのトレーニング コストと比較して重要になってきます。 2 つの推論フェーズでは、ハードウェアに異なる点で重点を置きます。プロンプト プレフィルは並列であり、通常はコンピューティングに依存しますが、自己回帰デコードはシーケンシャルであり、多くの場合メモリ帯域幅に依存します。従来の幅または深さのスケーリングでは、追加されたすべてのレイヤーが両方のフェーズで評価されるため、両方のコストが同時に増加します。プロンプト全体の主な計算と単一の永続的なキー値 (KV) キャッシュを保持しながら、追加の学習された計算を継続予測に代わりに割り当てることができるかどうかを尋ねます。デュアルフロートランスをご紹介します。その主なフローは、プロンプトを処理して KV キャッシュを書き込む完全な因果言語モデルです。補助フローはプロンプト処理中に省略され、最後のプロンプト位置以降のみアクティブ化され、永続的な状態を書き込んだり主フローに影響を与えたりすることなく継続予測の計算が追加されます。 2 つのフローは、メジャー アテンション、MLP、および出力行列を共有し、別個のトークン埋め込みと軽量結合を使用します。重みとプライマリ キャッシュを共有すると、グループ化された実行中にロードされた重みとキャッシュされたキーと値を再利用する機会も生まれます。デュアルフローは、一致したトークンの比較において、アーキテクチャおよびデータ構成全体で検証損失の低減を実現します。 MoE モデルでは、この分離により、主要エキスパートと補助エキスパートのファンアウトが即時コスト、継続コスト、予測品質に対して独立して制御されます。固定プレフィル エキスパート計算でのデコード計算を増やすことと、2 つのフロー間で固定デコード エキスパート バジェットを再割り当てすることの 2 つの体制を研究します。これらの実験は、プリフィルとデコードの品質のトレードオフを明らかにし、フェーズ固有の専門家割り当ての可能性を実証します。
原文 (English)
Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In typical serving, prompt prefill runs in parallel and is compute-bound, whereas autoregressive decode is sequential and memory-traffic-bound. Conventional width or depth scaling raises both costs together, since every added layer is evaluated in both phases and enlarges the weights read at each decode step. We instead ask whether additional learned computation can be allocated to continuation prediction while preserving prompt-wide primary computation and a single KV cache. We realize this with the Decode-Branch Transformer. Its primary path alone processes the prompt and writes the KV cache; the decode branch is omitted during prefill and activated only from the final prompt position onward, adding continuation computation without writing state or affecting the primary path. The paths share attention, MLP, and output matrices, using separate token embeddings with lightweight coupling. Grouped decode reuses loaded weight tiles and the primary KV cache across both paths, so the added arithmetic does not proportionally increase dominant memory traffic or decode latency. Across matched-token comparisons, Decode-Branch achieves lower validation loss across architectures and data settings. In MoE models, the primary and branch expert fan-outs become independent knobs for trading prompt cost, decode cost, and predictive quality. We study two expert-allocation regimes, holding prefill or decode computation fixed, and expose a prefill-decode-quality trade-off enabled by phase-specific expert allocation.
Academic League of Artificial Intelligence - 教育、研究、普及の統合的な視点
学術リーグは、課外教育を促進し、大学と社会の統合を強化するための重要なメカニズムとなっています。この文書では、サンタカタリーナ連邦大学 (UFSC) の人工知能学術連盟 (LIA) が採用した組織枠組みについて説明します。この組織枠組みは、学生中心のプロジェクトベースのアプローチを通じて、教育、研究、大学の拡張を統合するように設計されています。このフレームワークは、民主的なガバナンス、共同学習、ダイナミックなプロジェクト組織を組み合わせて、技術的能力と横断的な能力の両方を育成します。このフレームワークは、競技チーム、研究グループ、公開講座、ナレッジ リポジトリ、社会的影響力を持つ AI を活用したアプリケーションなどの代表的な取り組みを通じて説明されています。これらのプロジェクトは、リーダーシップ、科学的生産、コミュニティへの関与、知識の保存を促進しながら、共通の組織構造内で多様な教育、科学、普及活動をどのように展開できるかを実証しています。報告された経験は、提案されたフレームワークが大学の 3 つの柱をエンジニアリングおよびコンピューティング教育に統合するための柔軟で再現可能なモデルを提供し、学術リーグや同様の学生団体に実践的なガイダンスを提供することを示しています。
原文 (English)
Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension
Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between universities and society. This paper presents the organizational framework adopted by the Academic League of Artificial Intelligence (LIA) at the Federal University of Santa Catarina (UFSC), designed to integrate teaching, research, and university extension through a student-centered, project-based approach. The framework combines democratic governance, collaborative learning, and dynamic project organization to foster both technical and transversal competencies. The framework is illustrated through representative initiatives, including competition teams, study groups, open lectures, knowledge repositories, and AI-powered applications with social impact. These projects demonstrate how diverse educational, scientific, and extension activities can be developed within a common organizational structure while promoting leadership, scientific production, community engagement, and knowledge preservation. The reported experience indicates that the proposed framework provides a flexible and replicable model for integrating the three university pillars into engineering and computing education, offering practical guidance for academic leagues and similar student organizations.
MobileMem: Learning from a Year of Mobile Experiences
The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants…
From Monte Carlo to neural networks approximations of boundary value problems
In this paper we study probabilistic and neural network approximations for solutions to Poisson equation subject to Holder data in general…
ShadowNet for Data-Centric Quantum System Learning
Understanding the dynamics of large quantum systems is hindered by the curse of dimensionality. Statistical learning offers new possibiliti…
A Bi-directional Multi-solution Scalable Grover Search Algorithm
Grover's search algorithms, including various Partial Grover Searches (PGS), suffer from scaling issues when multiple solutions are sought,…
DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Variations
This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and ar…
TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability
Achieving zero-shot adversarial robustness without sacrificing generalization remains challenging for foundation models such as CLIP, espec…
MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4
Vision-Language Models (VLMs) have recently seen significant advancements through integrating with Large Language Models (LLMs). The VLMs,…
Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study
Personalized phishing is difficult to defend against because messages can be tailored to a target's work, interests, and social context. La…
Quantum Large Language Models via Tensor Network Disentanglers
We introduce a framework for seamlessly integrating quantum computing into pretrained large language models (LLMs). The key idea is to cons…
MoE-Enhanced Explainable Deep Manifold Transformation for Complex Data Embedding and Visualization
Dimensionality reduction (DR) plays a crucial role in various fields, including data engineering and visualization, by simplifying complex…
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
Diffusion Transformers (DiT) have become the dominant methods in image and video generation yet still suffer substantial computational cost…
Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks
Multi-hop question answering requires a system to identify and integrate evidence distributed across documents, yet large language models r…
Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities
Selecting appropriate training data is crucial for instruction fine-tuning of large language models (LLMs), which aims to (1) elicit strong…
ConfRetro: a 3D-aware template-free method for enhancing retrosynthesis via molecular conformer information
Motivation: Retrosynthesis plays a crucial role in organic synthesis and drug discovery, focusing on identifying a set of reactants capable…
Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects
The proliferation of digital interactions across diverse domains, such as healthcare, e-commerce, gaming, and finance, has resulted in the…
DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Decoupled Reasoning
Large Language Models (LLMs) exhibit strong natural language understanding capabilities but also inherit and amplify societal biases, parti…
Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation
Despite the increasing use of large language models for creative tasks, their outputs often lack diversity. Common solutions, such as sampl…
Bringing Generative Learning to Representation Learning: Self-Supervised Transfer Learning as Distribution Matching
Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified. We formulate represen…
Enhancing the Non-Functional Quality Compliance of LLM-Generated Code through Quality-Aware Preference Learning
Large Language Models (LLMs) have been widely adopted in commercial code completion engines, significantly enhancing coding efficiency and…
Leveraging Machine Unlearning for Cost-Efficient Preference Alignment
Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human f…
Bye-bye, Bluebook? Automating Legal Drudgery With AI-Augmented Rule Following
One of the central promises of legal AI is to automate drudgery -- the formal, repetitive tasks of lawyers' work that consume time without…
WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales
Responsibly deploying artificial intelligence (AI) / machine learning (ML) systems in high-stakes settings arguably requires not only proof…
Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
This paper presents a novel methodology for enhancing Automated Program Repair (APR) through synthetic data generation utilizing Large Lang…
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressin…
PhyxMamba: Chaotic System Reconstruction from Short Context Observations with Generative State-Space Models
Understanding chaotic dynamics is a fundamental problem across scientific disciplines, including climate science, neuroscience, and fluid d…
VirnyFlow: Optimizing ML Pipelines for Accuracy, Fairness, and Stability at Scale
Developing machine learning (ML) systems for real-world deployment requires navigating context-dependent trade-offs among accuracy, fairnes…
LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking
Recent advances in Large Language Models (LLMs) have driven their adoption in recommender systems through Retrieval-Augmented Generation (R…
Contraction-Aware Reinforcement Learning for Nonlinear Control with Statistical Robustness
Control contraction metrics (CCMs)-defined by Riemannian metrics under which a closed-loop system is incrementally exponentially stable-off…
From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology
Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human inst…
A validity-guided workflow for robust large language model research in psychology
Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets,…
Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges
The Segment Anything Model (SAM) has transformed image segmentation by introducing a prompt-based paradigm that enables strong zero-shot ge…
Reprojection-Guided 3D Gaussian Splatting Diffusion for Weakly Supervised Single-Image Normal Estimation
We propose CLONE, a Continuous Latent Optimization framework for Normal Estimation via 3D Gaussian splatting. The core idea is to construct…
Adapting LLMs to Time Series Forecasting via Temporal Heterogeneity Modeling and Representation Alignment
Recent advances have demonstrated that Large Language Models (LLMs) can be effectively adapted for time series forecasting, revealing stron…
ProteoKnight: Convolution-based Phage Virion Protein Classification and Uncertainty Analysis
\textbf{Introduction:} Accurate prediction of Phage Virion Proteins (PVP) is essential for genomic studies due to their crucial role as str…
CulTrace: Tracing Internal Cultural Reasoning in Large Language Models
The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidd…
PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning
Emotional support conversations require more than fluent responses. Supporters need to understand the seeker's situation and emotions, adop…
Efficient Code Embeddings from Code Generation Models
jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical quest…
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
Vision-Language Models (VLMs) have demonstrated remarkable success across diverse visual tasks, yet their performance degrades in complex v…
Privacy-Preserving Decentralized Federated Learning via Explainable Adaptive Differential Privacy
Decentralized federated learning enables collaborative model training without a central server, but shared model updates can still leak sen…
Geometrically Constrained and Token-Based Probabilistic Spatial Transformers
Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification. Careful conside…
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models
Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains di…
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensive…
DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts
Soft snake robots offer remarkable flexibility and adaptability in complex environments, yet their control remains challenging due to highl…
Federated Self-Supervised Modulation Classification under Non-IID and Imbalanced Data
Automatic modulation classification (AMC) is a core enabler of cognitive wireless systems, providing spectrum awareness and supporting adap…
A Large-Scale Chinese Knowledge Graph-Text Alignment Dataset for Benchmarking Knowledge-Grounded LLMs
Reliable evaluation of knowledge-grounded Large Language Models (LLMs) in Chinese requires resources that explicitly align Chinese-language…
Sleeping Kelly
The Sleeping Beauty problem is a problem of imperfect recall that has received considerable attention. One approach to solving the Sleeping…
Explainable Heterogeneous Anomaly Detection in Financial Networks via Adaptive Expert Routing
Financial anomalies arise from heterogeneous mechanisms - price shocks, liquidity freezes, contagion cascades, and momentum reversals - yet…
Retrofit: Continual Learning with Controlled Forgetting for Binary Security Detection and Analysis
Binary security has increasingly relied on deep learning to reason about malware behavior and program semantics. However, the performance o…
High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched-Grid
We present a probabilistic data-driven weather model providing ensembles of high spatial resolution realizations of 87 variables at arbitra…
jina-vlm: Small Multilingual Vision Language Model
We present jina-vlm, a token-efficient 2.4B parameter vision-language model that achieves state-of-the-art multilingual VQA performance amo…
Q-Regularized Generative Auto-Bidding: From Suboptimal Trajectories to Optimal Policies
With the rapid development of e-commerce, auto-bidding has become a key asset in optimizing advertising performance under diverse advertise…
The Fake Friend Dilemma: Relational Trust and the Political Economy of Conversational AI
As conversational AI systems become a larger part of the media landscape, they raise questions about whose interests they serve and the ris…
QA-Merging: Query-Adaptive Reasoning via Layer Selective Model Merging
Recent large reasoning models (LRMs) have achieved strong performance on complex reasoning tasks by generating a long chain-of-thought (Lon…
Backpropagation-Free Test-Time Adaptation for Lightweight EEG-Based Brain-Computer Interfaces
Electroencephalogram (EEG)-based brain-computer interfaces (BCIs) face significant deployment challenges due to inter-subject variability,…
AWED-PIPER: Agents, Web Applications & Expert Detectors for Personally Identifiable Information Protection & Fine-grained Named Entity Recognition across 36 languages for 6.6 Billion Speakers
Named Entity Recognition (NER) and Personally Identifiable Information (PII) anonymization are critical tasks in Natural Language Processin…
Sequential LLM Release Facilitates Manipulation in Regulated Markets
AI agents increasingly mediate bargaining, negotiation and persuasion for people and firms. Such markets extend software-mediated commerce,…
Aletheia: What Makes RLVR For Code Verifiers Tick?
Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training…
Robust Privacy: Inference-Stage Privacy through Certified Robustness
An adversary observing a model's released prediction can infer sensitive attributes of the queried input, or even reconstruct representativ…
Credit Fairness: Online Fairness In Shared Resource Pools
We study repeated allocation of shared resources among agents with time-varying demands and capped linear utilities. In this setting, indep…
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
Attentio-FFN disaggregation (AFD) is an emerging architecture for LLM decoding that separates state-heavy, KV-cache-dominated Attention com…
SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking
Very-high-resolution remote-sensing imagery provides a scalable basis for delineating informal settlements, but sparse annotations, severe…
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter
Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks. To address the challenges of e…
Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization
In this work we address the problem of training a Reinforcement Learning agent to follow multiple temporally-extended instructions expresse…
Zero-Shot Instruction Following in RL via Structured LTL Representations
We study instruction following in multi-task reinforcement learning, where an agent must zero-shot execute novel tasks not seen during trai…
ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that exe…
LORA-CRAFT: Cross-layer Rank Adaptation via Frozen Tucker Decomposition of Pre-trained Attention Weights
We introduce LoRA-CRAFT (\textbf{C}ross-layer \textbf{R}ank \textbf{A}daptation via \textbf{F}rozen \textbf{T}ucker), abbreviated CRAFT thr…
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the a…
Exact Attention Sensitivity and the Geometry of Transformer Stability
We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation. Our main result is the exact…
Reasoning-Based Personalized Generation for Users with Sparse Data
Large Language Model (LLM) personalization holds great promise for tailoring responses by leveraging personal context and history. However,…
SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human vis…
Automating the Detection of Requirement Dependencies Using Large Language Models
Requirements are inherently interconnected through various types of dependencies. Identifying these dependencies is essential, as they unde…
Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs
Predicting future student responses to questions is particularly valuable for educational learning platforms where it enables effective int…
Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast
Demographic attributes can be predicted from medical images, raising concerns about bias in clinical AI systems. In X-ray imaging, acquisit…
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be…
Informative Perturbation Selection for Uncertainty-Aware Post-hoc Explanations
Trust and ethical concerns due to the widespread deployment of opaque machine learning (ML) models motivating the need for reliable model e…
Data-knowledge dual-driven intelligent framework for full-chain, experiment-efficient synthesis of 2D dendrites
Exemplified by the chemical vapor deposition growth of two-dimensional dendrites, which has potential applications in catalysis and present…
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.g., 4…
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
As Large Language Models (LLMs) increasingly assist secure software development, their ability to meet the rigorous demands of Rust program…
SimulCost: LLM を使用して物理シミュレーションを自動化するためのコストを意識したベンチマークおよびツールキット
科学的タスク用の LLM エージェントの評価では、シミュレーション時間や実験リソースなどのツール使用コストを無視して、トークン コストに焦点を当ててきました。その結果、現実的な予算制約の下では、pass@k のようなメトリクスは非現実的になります。このギャップに対処するために、物理シミュレーションにおけるコスト重視のパラメーター調整を対象とした最初のベンチマークである SimulCost を導入します。 SimulCost は、流体力学、固体力学、プラズマ物理学の 13 のシミュレータにわたる 2,947 のシングルラウンド (初期推測) タスクと 1,931 のマルチラウンド (試行錯誤による調整) タスクにわたって、精度と計算コストの両方において LLM チューニングのコスト重視のパラメーターを従来のスキャン アプローチと比較します。各シミュレータのコストは分析的に定義され、プラットフォームに依存しません。 Frontier LLM はシングルラウンド モードで 46 ~ 65% の成功率を達成しますが、高精度要件下では 35 ~ 55% に低下するため、特に高精度タスクの場合は初期推測が信頼できなくなります。マルチラウンド モードではレートが 72 ~ 81% に向上しますが、LLM は従来のスキャンより 1.5 ~ 2.5 倍遅いため、非経済的な選択となります。また、知識伝達の可能性に関するパラメーター グループの相関関係、およびコンテキスト内の例と推論作業の影響も調査し、展開と微調整に対する実用的な意味を提供します。私たちは、静的ベンチマークおよび拡張可能なツールキットとして SimulCost をオープンソース化し、物理シミュレーション用のコストを意識したエージェント設計の改善と新しいシミュレーション環境の拡張に関する研究を促進します。コードとデータは https://github.com/Rose-STL-Lab/SimulCost-Bench で入手できます。
原文 (English)
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,643 single-round (initial guess) and 2,304 multi-round (adjustment by trial-and-error) tasks across 11 simulators from fluid dynamics, solid mechanics, and plasma physics, whose costs are analytically defined and platform-independent. A twelfth simulator, a production plasma code measurable only by wall clock, is reported separately. Frontier LLMs achieve 45-62% success rates in single-round mode, dropping to 34-50% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 66-81%, but LLMs are 1.5-2.7x slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench
Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence
The pruning of 3D Gaussian splats is essential for reducing their complexity to enable efficient storage, transmission, and downstream proc…
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
Scalable Vector Graphics (SVG) are essential for technical illustration and digital design, offering resolution independence and semantic e…
I-CALM: Incentivizing Confidence-Aware Abstention for LLM Selective Answering
Large language models (LLMs) often produce confident but incorrect answers, in part because standard evaluation incentives reward guessing…
Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models
Open-loop end-to-end neural motion planners have recently been proposed to improve motion planning for robotic manipulators. These methods…
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck du…
Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification
Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and impro…
Structural Generalization on SLOG without Hand-Written Rules
Structural generalization in semantic parsing requires systems to apply learned compositional rules to novel structural combinations. Exist…
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large lang…
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic sce…
Evolving Ensemble of Agents
We introduce the Evolving Ensemble of Agents (EvE), a decentralized framework that organizes existing, highly capable coding agents into a…
AgentMV: A State-Guided Multi-Agent Framework for Budget-Aware Music Video Generation
Generating a complete music video from a song requires more than synthesizing visually plausible clips for individual lyric prompts. A prac…
Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting
Large Language Models (LLMs) have shown promising results on NLP tasks, however, their performance on tabular data still needs research att…
SymbolicLight V1: Spike-Gated Dual-Path Language Modeling at High Activation Sparsity
Natively trained spiking language models must preserve information across time while operating through sparse binary activations, a combina…
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation…
ポリマーの設計と発見のための周期的トポロジカル深層学習
ポリマーはエネルギー、ヘルスケア、材料科学にわたる応用を支えていますが、その広大な化学領域が体系的な発見を困難にしています。ほとんどの機械学習アプローチでは、ポリマーを単一の繰り返し単位の分子グラフとして表すため、ポリマー鎖の周期性と対結合を超えた多体相互作用の両方が失われます。複数の空間スケールにわたる多体相互作用を捕捉する周期的 Vietoris-Rips 複合体に基づいて構築された深層学習フレームワークである Periodic-TDL と、その後に長距離相互作用から共有結合に情報を伝播する階層的単純メッセージ パッシング (HSMP) エンコーダーを導入し、高次のトポロジー特徴によって強化された表現を生成します。 Periodic-TDL は、電子的、光学的、物理的、熱的ターゲットにわたるポリマー特性予測タスク全体にわたって、すべての最先端モデルよりも優れたパフォーマンスを発揮します。さらに、エステルからアミドへの置換と $\alpha$-メチル化がどのように熱安定性を高めるかを定量的に検証します。アクリレートポリマーとアクリルアミドポリマーの体系的な置換によって生成された48,208個の構造の計算合成データセットを使用して、一致するポリマーペア全体で、エステルからアミドへの置換では$\sim 55^\circ$Cの平均$T_g$増加と、骨格の$\alpha$-メチル化では$\sim 14^\circ$Cの平均$T_g$増加が観察されました。これらの予測された傾向を検証するために、Periodic-TDL モデルを使用して、これまで文献で報告されていなかった 3 つの新しく合成されたポリマーを含む、独立した実験測定からの 6 つの新規ポリマー ペアを分析しました。実験データはモデルの予測を裏付けることに成功しました。最終的に、これらの発見は、Periodic-TDL が単にベンチマーク データセットの予測パフォーマンスを最適化するのではなく、特定の官能基修飾の根底にある物理的効果を捕捉していることを示しています。
原文 (English)
Periodic Topological Deep Learning for Polymer Design and Discovery
Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging. Most machine learning approaches represent polymers as molecular graphs of a single repeating unit, thereby missing both the periodicity of polymer chains and many-body interactions beyond pairwise bonds. We introduce Periodic-TDL, a deep learning framework built on periodic Vietoris-Rips complexes that capture many-body interactions across multiple spatial scales, followed by a hierarchical simplicial message-passing (HSMP) encoder that propagates information from long-range interactions to covalent bonds, yielding representations enriched by higher-order topological features. Periodic-TDL outperforms all state-of-the-art models across polymer property prediction tasks spanning electronic, optical, physical, and thermal targets. Furthermore, we quantitatively validate how ester-to-amide substitution and $\alpha$-methylation enhance thermal stability. Using a computationally synthesized dataset of 48,208 structures-generated via systematic substitution of acrylate and acrylamide polymers-we observed a mean $T_g$ increase of $\sim 55^\circ$C for ester-to-amide substitutions and $\sim 14^\circ$C for backbone $\alpha$-methylation across matched polymer pairs. To verify these predicted trends, we use our Periodic-TDL model to analyze six novel polymer pairs from independent experimental measurements, including three newly synthesized polymers previously unreported in the literature. The experimental data successfully confirmed the model's predictions. Ultimately, these findings demonstrate that Periodic-TDL captures the underlying physical effects of specific functional group modifications, rather than merely optimizing predictive performance on benchmark datasets.
エネルギーの盲点: NVIDIA の主力エッジ AI ハードウェアはプロセスレベルのエネルギー属性をサポートできない
単一のユーザー目標によって複数ステップのオーケストレーション、ツール呼び出し、再試行、障害回復がトリガーされるエージェントティック AI ワークロードは、エッジ導入のターゲットとなっており、NVIDIA、デル、HP、ASUS、MSI、Acer、ギガバイトのすべてが 2026 年に GB10 ベースのデスクトップ AI システムを出荷します。私たちは最近、オーケストレーション構造がエージェントのエネルギー コストの大半を占めていることを実証しました。ワークフローは、成功した目標ごとに線形ベースラインよりも 4.33 倍多くのエネルギーを消費します。マルチステップ推論タスクの OOI は 7.63 倍に達します。これとは別に、Rajat et al。 CPU 側の処理が、エージェント ワークロードの総レイテンシの最大 90.6%、総動的エネルギーの 44% を占めることが示されています。私たちは、ASUS Ascent GX10 (GB10 SoC) の系統的なエネルギー観測可能性監査を報告し、このプラットフォームでは、サポートされているソフトウェア インターフェイスを通じて、CPU エネルギー カウンター、INA パワーレール モニター、IPMI/BMC、および SCMI パワーキャップ プロトコルを公開していないことがわかりました。唯一のオンデバイス エネルギー テレメトリは、NVML を介した瞬間的な GPU 電力です。さらに、MediaTek ファームウェアが文書化されていない ACPI インターフェイス (SPBM) を介してレールごとのエネルギーを内部で計算していることも判明しましたが、NVIDIA は「CPU レール情報を公開する予定はない」と述べています。したがって、RAPL 経由で x86 上で実行されるデバイス上のプロセスごとのエネルギー アトリビューションは、サポートされているインターフェイスを介してこのプラットフォームでは再現できません。私たちは、エネルギーに起因する AI のハードウェア要件仕様を形式化し、GPU 減算と組み合わせた外部 DC メータリングを使用した暫定キャリブレーション ブリッジを提案し、SCMI パワーキャップを介して標準トラック パスを特定します。私たちの調査結果は、低炭素コンピューティング コミュニティに、第一級のハードウェア要件としてエネルギーの可観測性を要求する動機を与えています。
原文 (English)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
Agentic AI workloads - where a single user goal triggers multi-step orchestration, tool calls, retries, and failure recovery - are being targeted for edge deployment, with NVIDIA, Dell, HP, ASUS, MSI, Acer, and Gigabyte all shipping GB10-based desktop AI systems in 2026. Prior work shows orchestration structure dominates agentic energy cost and CPU-side processing accounts for up to 44% of total dynamic energy. We report a systematic energy-observability audit of the ASUS Ascent GX10 (GB10 SoC) and find that the platform exposes no CPU energy counter, no INA power-rail monitor, no IPMI/BMC, and no SCMI powercap protocol through any supported software interface. The only on-device energy telemetry is instantaneous GPU power via NVML. We further discover that the MediaTek firmware already computes per-rail energy internally via an undocumented ACPI interface (SPBM), but NVIDIA states there are "no plans to expose CPU rail information." On-device per-process energy attribution - as performed on x86 via RAPL - is therefore not reproducible on this platform through supported interfaces. We formalize a hardware requirements specification for energy-attributed AI, propose an interim calibration bridge for per-domain energy decomposition - confirmed on the Acer Veriton GN100 where CPU energy accumulators are live - and identify a standards-track path via SCMI powercap. Our findings motivate the low-carbon computing community to demand energy observability as a first-class hardware requirement.
多腕ベイジアン バンディットのアニーリングされたソフトマックスの貪欲さ
検証可能な報酬を伴う強化学習 (RLVR) および GRPO などのグループベースのポリシー最適化手法は、プロンプトごとに複数の完了をサンプリングし、参照ポリシーに対する KL ペナルティによって正規化された、より高い報酬を持つポリシーの確率を高めることにより、確率的ポリシーを更新します。これらの更新には、認識論的不確実性を追跡する明示的なメカニズムは含まれていません。この論文では、なぜそのような不確実性を問わない更新が効果的であるのかについて、定型化された説明を研究します。多腕ベイジアン ベルヌーイ バンディットにおける経験的平均報酬のソフトマックスに従ってアクションを選択するアニーリングされたソフトマックス (ボルツマン) ポリシーを分析します。最適に近いアームが豊富にあることを意味する、事前の線形アッパーテール条件 ($\beta$-規則性の $\beta=1$ の場合) では、アニーリングされたソフトマックス グリーディがベイズ リポート $\tilde{O}(m + T/m)$ を達成すること、特にアームの数が $m = にスケールされる場合 $\tilde{O}(\sqrt{T})$ を達成することを証明します。 \シータ(\sqrt{T})$。これは、この体制における最適に近いベイズの後悔率であり、経験的平均の貪欲さによっても達成されます。 $\beta$-規則性の下では、多くのアームは学習を通じて最適値に近い経験的平均を維持するため、ソフトマックスが経験的に最良でないアームをサンプリングすると、そのアームは明らかに劣ったアームではなく、最適に近い別のアームになる傾向があります。対照的に、アームの数が少ない場合、同じ種類のソフトマックス ポリシーは直線的な後悔に見舞われる可能性があります。この結果は、RLVR と構造的に類似していることも示しています。ここでは、正しい完了を生成する無視できない確率を持つ基本ポリシーが $\beta$-規則性の役割を果たします。
原文 (English)
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
Reinforcement learning with verifiable rewards and group-based policy optimization methods update a stochastic policy by sampling multiple completions per prompt and increasing the policy's probability on those with higher reward. These updates, unline the exploration mechanism in Thompson sampling and UCB, do not include explicit mechanisms that track epistemic uncertainty. This paper studies a stylized explanation for why such uncertainty-agnostic updates can nevertheless be effective. We analyze an annealed softmax policy that selects actions according to a softmax of empirical mean rewards in a many-armed Bayesian Bernoulli bandit. Under a linear upper-tail condition on the prior, which implies an abundance of near-optimal arms, we prove that annealed softmax greedy achieves Bayes regret $\tilde{O}(m + T/m)$, and in particular $\tilde{O}(\sqrt{T})$ when the number of arms scales as $m = \Theta(\sqrt{T})$. This is the near-optimal Bayes regret rate in this regime, attained also by empirical-mean greedy. Under the upper-tail condition, many arms keep empirical means near the optimum throughout learning, so the probability that softmax places away from the empirical best falls mostly on other near-optimal arms. By contrast, with a small number of arms, the same kind of softmax policy can suffer linear regret (Cesa-Bianchi et al., 2017). The result also provides a structural analogy to RLVR, where a base policy with a non-negligible probability of producing a correct completion plays the role of the tail condition. Simulations support the theory and motivate prior-anchored variants of greedy and annealed softmax that score arms by the Beta posterior mean and skip the forced initialization; with an arm-specific prior, accurate or noisy, these variants outperform baselines, including Thompson Sampling, when the number of arms is large.
SUPREME: 再現可能な画像非学習手法評価のためのマルチ GPU フレームワーク
機械の非学習では、最初から再トレーニングすることなく、トレーニングされたモデルから特定のトレーニング データの影響が除去されます。アンラーニング手法を評価するには、複数のシードにわたってトレーニング、アンラーニング、評価を繰り返す必要があり、計算コストがかかります。私たちの知る限り、既存の画像分類非学習フレームワークは単一の GPU 上で実行されるため、妥当な時間内に評価できるシードの数が制限されます。これらのステージを複数の GPU に分散するオープンソース フレームワークである SUPREME を紹介します。 SUPREME は 3 つの貢献を行っています。新しいメソッド、メトリック、モデル、シナリオを追加するためのレジストリ ベースの設計です。複数のアクセラレータと高精度モードをサポートするマルチ GPU アーキテクチャ。そして、10 個のシードにわたるフルクラスおよびランダム サンプルのアンラーニングのもとで、ResNet18 と ViT を使用したピンの顔認識のデモンストレーションです。このフレームワークは https://github.com/pedroandreou/supreme-unlearning で入手できます。
原文 (English)
SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation
Machine unlearning removes the influence of specific training data from a trained model without retraining it from scratch. Evaluating an unlearning method requires repeating training, unlearning, and evaluation across multiple seeds, which is computationally expensive. To our knowledge, existing image classification unlearning frameworks run on a single GPU, which limits how many seeds can be evaluated in reasonable time. We introduce SUPREME, an open-source framework that distributes these stages across multiple GPUs. SUPREME makes three contributions: a registry-based design for adding new methods, metrics, models, and scenarios; a multi-GPU architecture supporting multiple accelerators and precision modes; and a demonstration on Pins Face Recognition using ResNet18 and ViT under full-class and random-sample unlearning across ten seeds. The framework is available at https://github.com/pedroandreou/supreme-unlearning.
Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics
Large language models (LLMs) are increasingly entering students' learning practices, but their educational value may depend on whether they…
E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments
Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the mo…
MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models
Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision…
The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models
Pass/fail safety evaluation reports whether a model refused. It does not report how far a model went to please the user, and we show these…
Compositional Boundaries for Density Fusion
Distributed uncertainty-management systems often combine local probabilistic models along aggregation trees chosen by communication, privac…
LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step…
Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries
Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm ut…
国境を越えたコミュニティ学習のための文化を意識した AI: 計算と設計の交差点における学部のイノベーション
教育における人工知能 (AIED) に関する研究は急速に拡大していますが、技術の進歩には人間中心の基礎や文化的背景への十分な配慮が欠けていることがよくあります。ソーシャルワークに根ざした教育学であるコミュニティベース学習は、AIEDの研究において、特にアジア太平洋地域においては依然として過小評価されている。この論文では、学部生が文化遺産の保存と持続可能な開発のための AI 対応ソリューションを開発する、境界を越えたコミュニティベースの学習について報告します。私たちは、教育、テクノロジー、文化の 3 つの側面にわたって、コミュニティ参加型コンピューティングが人間中心の AIED をどのように運用できるかを調査します。私たちは、ソーシャルワークと計算科学の間の専門分野の縦割りを解消することで参加を拡大しながら、マルチステークホルダーのコラボレーションを促進する、文化を意識した AIED のための協力フレームワークに貢献します。
原文 (English)
Culturally-Aware AI for Cross-Boundary Community Learning: Undergraduate Innovation at the Intersection of Computation and Design
Research on artificial intelligence in education (AIED) is rapidly expanding, yet technical progress often lacks human-centered grounding and adequate attention to cultural context. Community-Based Learning, a pedagogy rooted in social work, remains underrepresented in AIED research, particularly within Asia-Pacific contexts. This paper reports on cross-boundary Community-Based Learning where undergraduate students develop AI-enabled solutions for cultural heritage preservation and sustainable development. We examine how community-engaged computing operationalizes culturally aware, human-centered AIED through participatory elicitation of cultural knowledge, bilingual representation, and stakeholder validation across education, technology, and culture. We contribute a collaborative framework for culturally aware AIED designed to support multi-stakeholder collaboration and widen participation by bridging social work and computational science.
Speculative Rollback Correction for Quality-Diverse Web Agent Imitation
Training interactive web agents through imitation learning from expert trajectories has emerged as a highly effective approach. However, de…
SL-S4Wave: 構造化状態空間モデルを使用した生理学的波形の自己教師あり学習
心電図 (ECG) などの長いシーケンスの医療時系列データのモデリングは、高いサンプリング レート、マルチチャネル信号の複雑さ、固有のノイズ、およびラベル付きデータの制限により、重大な課題を引き起こします。畳み込みニューラル ネットワークなどのさまざまなエンコーダ アーキテクチャに基づく最近の自己教師あり学習 (SSL) 手法は、ラベルのないデータから表現を学習するために提案されていますが、長距離の依存関係やノイズ不変の特徴を捕捉するには不十分であることがよくあります。構造化状態空間モデル (S4) は長いシーケンスのモデリングに優れていますが、既存の S4 アーキテクチャはマルチチャネル生理学的波形の固有の特性を捉えることができません。この研究では、構造化状態空間モデルに基づいて構築された調整されたエンコーダーと対照学習を組み合わせた自己教師あり学習フレームワークである SL-S4Wave を提案します。エンコーダには、マルチスケール サブカーネルを使用した多層グローバル コンボリューションが組み込まれており、ノイズの多い高解像度のマルチチャネル波形におけるきめの細かいローカル パターンと長距離の時間依存性の両方をキャプチャできます。現実世界のデータセットでの広範な実験により、SL-S4Wave は、(1) 困難な不整脈検出タスクにおいて、常に最先端の教師付きベースラインおよび自己教師付きベースラインを上回るパフォーマンスを示し、(2) 大幅に少ないラベル付きサンプルで高いパフォーマンスを達成し、強力なラベル効率を示し、(3) 長い波形セグメントで堅牢なパフォーマンスを維持し、既存のアプローチのほとんどが効率的にモデル化できない長いシーケンスにおける複雑な時間ダイナミクスをモデル化する能力を強調し、(4) 転送目に見えないタイプの不整脈に効果的であり、その堅牢なクロスドメインの一般化が強調されています。さらに、複数のEEGタスクでSL-S4Waveを評価し、強力なベースラインを超えて優れたパフォーマンスを達成し、心臓波形を超えたアプローチの一般化可能性を実証しました。
原文 (English)
SL-S4Wave: Self-Supervised Learning of Physiological Waveforms with Structured State Space Models
Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates, multichannel signal complexity, inherent noise, and limited labeled data. While recent self-supervised learning (SSL) methods, based on various encoder architectures such as convolutional neural networks, have been proposed to learn representations from unlabeled data, they often fall short in capturing long-range dependencies and noise-invariant features. Structured state space models (S4) excel at long-sequence modeling, but existing S4 architectures fail to capture the unique characteristics of multichannel physiological waveforms. In this work, we propose SL-S4Wave, a self-supervised learning framework that combines contrastive learning with a tailored encoder built on structured state space models. The encoder incorporates multi-layer global convolution using multiscale subkernels, enabling the capture of both fine-grained local patterns and long-range temporal dependencies in noisy, high-resolution multichannel waveforms. Extensive experiments on real-world datasets demonstrate that SL-S4Wave (1) consistently outperforms state-of-the-art supervised and self-supervised baselines in a challenging arrhythmia detection task, (2) achieves high performance with significantly fewer labeled examples, showcasing strong label efficiency, and (3) maintains robust performance on long waveform segments, highlighting its capacity to model complex temporal dynamics in long sequences that most existing approaches fail to efficiently model, and (4) transfers effectively to unseen arrhythmia types, underscoring its robust cross-domain generalization. We additionally evaluate SL-S4Wave on multiple EEG tasks, achieving superior performance over strong baselines, demonstrating generalizability of our approach beyond cardiac waveforms.
Empowering Polymeric Materials Discovery by Artificial Intelligence
Polymeric materials underpin modern technologies spanning energy storage, microelectronics, healthcare and sustainable manufacturing. Yet t…
Agentic Red チームのレッドチーム化
攻撃的なセキュリティ操作を実行するためのエージェント システムの使用は、理論上の可能性からコモディティ化された機能に移行しました。しかし、コミュニティはより多くの有能なエージェントを作成することに焦点を当ててきましたが、それらのシステムのセキュリティの評価にはあまり注意が払われてきませんでした。この研究では、攻撃的なセキュリティ作戦に最も広く使用されているエージェント システムの詳細なセキュリティ分析を初めて紹介します。これらのツールのほとんどには共通の設計上の欠陥があり、エージェントがサンドボックス コンテナ内で動作している場合でも、積極的な敵対者が API キーを窃取し、永続的な足場を確立し、オペレータのマシンを完全に侵害できることを示します。私たちの分析をサポートするために、このようなエージェント システムに完全なサイバー キル チェーンを導入し、最初の LLM 操作から横方向の移動、永続化、ガードレールのバイパス、サンドボックスからの脱出までの進行を捉えます。私たちはセキュリティ分析に基づいて、エージェント型攻撃セキュリティ ツールの堅牢なアーキテクチャを導き出し、公開された攻撃パスをアーキテクチャ レベルで軽減する実用的で広く適用可能な設計原則を提案します。
原文 (English)
Red-Teaming the Agentic Red-Team
The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, while the community has focused on creating more and more capable agents, less attention has been allocated to assessing the security of those systems. In this work, we present the first in-depth security analysis of the most widely used agentic systems for offensive security operations. We show that most of these tools share common design flaws that enable an active adversary to exfiltrate API keys, establish persistent footholds, and fully compromise the operator's machine, even when the agent operates inside a sandboxed container. To support our analysis, we introduce a full cyber kill chain for such agentic systems, capturing the progression from initial LLM manipulation to lateral movement, persistence, guardrail bypass, and sandbox escape. Building on our security analysis, we derive a robust architecture for agentic offensive-security tools and propose actionable, broadly applicable design principles that mitigate the disclosed attack paths at the architectural level.
LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression
The rapid growth in the parameter scale of large language models (LLMs) has created a strong demand for efficient compression techniques. A…
統計上の敵対者: ビジョン データセット内の自然なバックドアのような特徴
モデル固有の敵対的攻撃は広範囲に研究されています。私たちは、別の障害モードを研究しています。それは、悪意を持って挿入されずにバックドアのようなトリガーのように動作する、ビジョン データ内で自然に発生する統計信号です。これらのシグナルを統計的敵対者と呼びます。 Imagenet を分析して、特定のラベルと強く関連するパターンを見つけます。次に、統計的制御を使用して、候補信号からランダムな相関を除去します。最後に、これらの信号がモデルの予測を直接かつ予想どおりに変更することを示します。これらの統計上の敵対者は、一般的な破損よりも標的が絞られており、異なるモデル アーキテクチャ間で転送されます。これは、一部の脆弱性は単一モデルの特異性ではなく、データセットの構造と分布によって引き起こされることを示唆しています。私たちは、ポイズニングが存在しない場合でも、通常のデータセットには悪用可能な敵対的表面が含まれている可能性があると結論付け、データセットの監査では、偽の構造をバイアスや解釈可能性の失敗の原因としてだけでなく、ビジョン モデルの潜在的な攻撃対象表面としても扱うべきであると提案します。
原文 (English)
Statistical Adversaries: Natural Backdoor-like Adversarial Features in Clean Vision Datasets
Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signals in vision data that can behave as backdoor-like triggers without being maliciously inserted. We call these signals statistical adversaries. We analyse ImageNet to find patterns that are strongly linked to certain labels. We then use statistical controls to remove random correlations from our candidate signals. Finally, we demonstrate that these signals directly and predictably alter model predictions. These statistical adversaries are more targeted than generic corruptions and transfer across different model architectures. This suggests that some vulnerabilities are driven by dataset structure and distribution rather than a single model's idiosyncrasies. We conclude that ordinary datasets can contain exploitable adversarial surfaces even in the absence of poisoning, and suggest that dataset audits should treat spurious structure not only as a source of bias or interpretability failure, but also as a latent attack surface for vision models.
潜在的な性格特性による言語モデルの効率的な安全性調整
大規模な言語モデルに対する現在の安全方法は、敵対的な攻撃に対して脆弱であることが知られており、堅牢な代替方法の研究が行われています。潜在的敵対的トレーニング (LAT) は最も効果的な防御策の 1 つですが、実用性が低下する可能性があり、有害なプロンプトの大規模なデータセットでのトレーニングが必要です。私たちは、潜在的性格調整(LPA)を導入します。これは、心理測定的性格文献から抽出されたわずか 66 個の危害を無視したステートメントを対象とした、明示的な危害の拒否を敵対的なトレーニングに置き換えます。私たちは、人格にアンカーされた表現は危害回避と潜在的な構造を共有しているため、それらを敵対的に安定させることで、脱獄攻撃によって悪用される部分空間を暗黙的に制限すると仮説を立てています。 LPA は、トレーニング中に有害なコンテンツがまったく表示されず、標準ベンチマークでパフォーマンスが低下しないにもかかわらず、直接リクエストと 5 つの脱獄方法にわたって HarmBench でほぼゼロの攻撃成功率を達成します。さらに、トレーニングプロセスは軽量です。手順全体は 1 つの GPU で数分で完了し、使用するサンプルの数は標準 LAT よりも 75 分の 1 です。広範なアブレーションは、私たちの方法の堅牢性、効率性、および一般化を示しています。
原文 (English)
Efficient Safety Alignment of Language Models via Latent Personality Traits
Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can degrade utility and requires training on large datasets of harmful prompts. We introduce Latent Personality Alignment (LPA), which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature. We hypothesize that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks. LPA achieves near-zero attack success rates on HarmBench across direct requests and five jailbreak methods, despite never seeing harmful content during training and no loss of performance on standard benchmarks. Moreover, the training process is lightweight; the entire procedure completes in minutes on a single GPU and uses 75x fewer examples than standard LAT. Extensive ablations demonstrate the robustness, efficiency, and generalization of our method.
LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each…
セーフガード条件付き上昇率: デュアルユースの生物学助手のユーティリティとリスクのフロンティアを測定する
デュアルユースの生物学アシスタントの安全性評価では、多くの場合、基本モデルの機能、拒否行動、またはジェイルブレイクの成功を測定します。これらのメトリックは、導入に関する質問を見逃しています。つまり、固定基本モデルの場合、ユーザーが実際に目にするアクセス条件は、無害なユーティリティと有害な実用的な支援をどのように変更するのでしょうか?私は、人間が判断したユーティリティとリスクのフロンティアを通じて展開されたアクセス条件を比較するためのプロトコルであるセーフガード条件付きアップリフトを紹介します。私は、Claude Sonnet 4.6 と Gemini 3.5 Flash を、役立つプロンプト、安全なプロンプト、および安全に保護された外部アシスタントの下で、108 タスクのサロゲート ベンチマークで評価しました。ヘッドラインの主張は、ロックされた 18 タスクのホールドアウト スプリットに限定されています。 600 行の盲検化された人間による監査では、保護されたアシスタントは、ブートストラップ 95% 間隔 [-0.117, -0.011] で、49 の一致した応答ペアにわたって、有益なプロンプトと比較して有害なアクション可能性を -0.063 減少させますが、正確性は間隔 [-0.057, +0.077] で +0.009 変化します。アダプティブ、テスト B、キューアブレーション、およびコントローラーベースラインのチェックは、測定ストーリーをサポートしますが、非優位性も示しています。多くの場合、安全プロンプトはクロードにとって最も強力ですが、外部制御はジェミニにとってより役立ち、良性の効用を減らす可能性があります。この貢献は普遍的な防御策ではありません。これは、導入レベルの評価目標に加えて、ユーザーが直面するアクセス条件が公益事業のリスクフロンティアをどのように動かすかを測定するための、学習されたリスク予算調整手順です。
原文 (English)
Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
A refusal rate neither identifies which component intervened nor measures its burden on legitimate users. This paper evaluates safeguards for dual-use biology assistants at the action and answer levels. The framework reconstructs the access path, separates provider refusals from downstream actions, and selects thresholds under an intervention budget. A frozen fresh-generation study satisfies its criterion on Claude Opus 4.5, but both passing configurations share one upstream provider effect; none passes on Gemini 2.5 Flash. The fixed Opus policies retain positive selectivity on 104 previously unused released-label pairs, but both fail a 20\% matched-benign constraint. At the answer level, no joint-scoring verifier qualifies on a response-disjoint 7,200-judgment holdout. A fresh 8,640-judgment factorial experiment finds separate gains from criterion isolation and ordinal representation, with a positive interaction between them; requiring explicit localization lowers aggregate accuracy under a strict no-repair schema. The evidence supports prospective action-level selectivity and identifies verifier interface effects, but not calibrated selective access, verified content removal, or biological-risk reduction.
価値の漏洩: LLM の答えは、自身の価値観によって静かに形成される
人々は、答えを検証するのが難しい実際的な質問に対して言語モデルを使用します。モデルが秘密の値の漏洩を示すことを示します。つまり、モデルが提供する情報は、その影響がユーザーに公開されることなく、独自の値の影響を受けます。私たちの評価の 1 つでは、ユーザーは AI 企業への投資を検討しており、AI バブルが弾ける可能性がどのくらいかを知りたいと考えています。 Claude Opus 4.8 は、検討中の企業が OpenAI ではなく Anthropic である場合、確率が低くなります。しかし、クロードはほとんどの場合、この影響をユーザーに開示していません。秘密の価値の漏洩は、ユーザーの好みに反し、ユーザーを誤解させる可能性があるため、不整合の一形態です。この現象を調査するために、値の漏れを定量化し、モデルがそれを明らかにするかどうかを定量化するための一連の評価を導入します。モデルは、道徳的に良い結果、モデルを開発した企業、人間の一部の余暇活動に対する他の嗜好など、さまざまな種類の価値観の影響を受けることがわかりました。同じ評価において、フロンティア モデル間で大きな差異が観察されることがよくあります。たとえば、フェルミ推定タスクでは、クロード モデルは思考連鎖において偏りのない答えを与えると誤って主張しますが、クウェン モデルは、その値がどのように答えに偏りを与えるかを説明します。価値の漏洩は、お調子者や報酬のハッキングとは異なる障害モードであり、現在の連携トレーニングや評価ではこれに適切に対処できません。
原文 (English)
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI. Yet Claude mostly fails to disclose this influence to the user. Covert value leakage is a form of misalignment because it goes against the user's preferences and is likely to mislead them. To investigate this phenomenon, we introduce a suite of evaluations to quantify value leakage and whether models disclose it. We find that models are influenced by different types of values, including preferences for morally good outcomes, for the company that developed them, and for some human leisure activities over others. We often observe large differences among frontier models on the same evaluation. For example, on a Fermi-estimation task, Claude models falsely claim to give unbiased answers in their chain-of-thought, while Qwen models explain how their values bias their answers. Value leakage is a failure mode distinct from sycophancy and reward hacking, and current alignment training and evaluations do not adequately address it.
生成 AI は教師あり XMLC に取って代わりますか?ドイツの科学文献を使用した自動主題索引付けに関するベンチマーク研究
ラベル セットとして大規模に管理された語彙を使用すると、ライブラリ内の自動主題インデックス付けのタスクは、マルチラベル分類タスクとして理解できます。対象用語のセットが大きい場合、問題は Extreme Multi-Label Classification (XMLC) の目的に適合します。この研究では、ドイツ国立図書館 (DNB) に収集された現代ドイツ科学文献の主題索引付けのテスト ケースに、厳選された特殊な教師付き XMLC 手法を適用します。従来の語彙一致ベースラインと、最近開発した 3 つの独自の LLM ベースの手法をベンチマークに含めることで、これらの結果を対比します。アルゴリズムはいくつかの指標で評価および比較されます。これには、以前に索引付けされた資料とのバイナリ関連性の比較や、専門の主題図書館員による段階的な関連性評価が含まれます。すべての手法に共通する課題は、対象語彙のロングテールから確実に提案を行うことです。トランスフォーマーベースの高密度特徴に依存する教師あり XMLC アルゴリズムが、全体的なバイナリ関連性メトリクスの観点から最良の結果をもたらすことがわかりました。ただし、主題語彙のロングテールにおける段階的な関連性とパフォーマンスに焦点を当てているため、LLM ベースの生成手法はより良い結果をもたらし、将来の生産的な使用のための有望な代替手段となります。
原文 (English)
Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature
With a large controlled vocabulary as the label set, the task of automated subject indexing in a library can be understood as a multi-label classification task. If the set of subject terms is large, the problem fits the Extreme Multi-Label Classification (XMLC) objective. In this study, we apply a selection of specialised supervised XMLC methods to the test case of subject indexing contemporary German scientific literature, collected at the German National Library (DNB). We contrast these results by including a classical lexical matching baseline and three of our own recently developed LLM-based methods into the benchmark. Algorithms are evaluated and compared in several metrics. This includes binary relevance comparisons with previously indexed material, as well as graded relevance ratings by professional subject librarians. A challenge for all methods is to reliably make suggestions from the long tail of the subject vocabulary. We find that supervised XMLC algorithms relying on transformer-based dense features give best results in terms of overall binary relevance metrics. However, focusing on graded relevance and performance in the long tail of our subject vocabulary, the LLM-based generative methods give better results, making them a promising alternative for future productive use.
Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft
The digital substrate - data, algorithms, infrastructure, platforms, applications - is being governed without adequate conceptual foundatio…
SCPP: ソフト クラスタリング用の統合 Python ライブラリ
このペーパーでは、ソフト クラスタリング用のオープンソース Python フレームワークである SCPP (Soft Clustering Python Package) を紹介します。 SCPP は、ファジー、確率的、グラフベース、行列因数分解、ディープ ラーニング手法など、異種ソフト クラスタリング手法全体でモデルのトレーニング、予測、メンバーシップ表現、評価、ベンチマークを標準化する、標準的な scikit-learn 互換の推定インターフェイスを確立します。このフレームワークは現在、40 の代表的なアルゴリズムと、データセット、クラスタリング品質メトリクス、標準化されたランタイム、メモリ、およびスケーラビリティ評価で構成される包括的なベンチマークを統合しています。 SCPP はさらに、広範なドキュメント、実践例、自動テスト、科学的な Python エコシステムとのシームレスな統合を提供し、再現可能な実験と新しいアルゴリズムによる直接的な拡張を可能にします。ソース コードは https://github.com/soft-clustering/soft-clustering で公開されています。
原文 (English)
SCPP: A Unified Python Library for Soft Clustering
In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering. SCPP establishes a canonical, scikit-learn-compatible estimator interface that standardizes model training, prediction, membership representation, evaluation, and benchmarking across heterogeneous soft clustering methods, including fuzzy, probabilistic, graph-based, matrix factorization, and deep learning methods. The framework currently integrates 40 representative algorithms together with a comprehensive benchmarking comprising datasets, clustering quality metrics, and standardized runtime, memory, and scalability evaluation. SCPP further provides extensive documentation, practical examples, automated testing, and seamless integration with the scientific Python ecosystem, enabling reproducible experimentation and straightforward extension with new algorithms. The source code is publicly available at https://github.com/soft-clustering/soft-clustering.
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detec…
Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies
Competence claims for a language model in a safety-critical domain are credible when measured against a standard the domain already enforce…
When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation
Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning ev…
Moral Hazard in Multi-Agent Language Models
Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr\"om's…
AI が生成したバイオデジタル アーキテクチャ画像からの EEG 感情認識
AI が生成した画像からの脳波 (EEG) データを使用して、バイオデジタル アーキテクチャに対する感情的反応を調べました。 336 人の参加者が参加した事前実験では、最初の 600 枚のプールから、畏怖、嫌悪感、内容に分類される強い感情反応を引き起こす 60 枚の画像が特定されました。これらの画像は、既存のデータセットの分析に基づいてチャネルの選択とサンプル サイズの推定を行い、52 人のボランティアの EEG 記録に使用されました。ガンマ バンドとデルタ バンドが最も高い分類精度をもたらし、ガンマ バンドは畏怖の感情について 77.07 パーセント +/- 13.8 パーセントの精度を達成しました。緑や不均一な粒度などの重要な要素はポジティブな感情に関連付けられますが、湿気はネガティブな反応を引き起こします。これらの結果は、美的魅力と受容性を高めるために、バイオデジタル建築に自然要素とさまざまなテクスチャを組み込むことの重要性を強調しています。この研究は、EEG が建築の好みを客観的に評価する能力を実証し、建築家が魅力的で持続可能な環境を設計するための貴重な洞察を提供します。
原文 (English)
EEG Emotion Recognition From AI-Generated Biodigital Architecture Images
Emotional responses to biodigital architecture were examined using electroencephalographic (EEG) data from AI-generated images. A pre-experiment involving 336 participants identified 60 images, selected from an initial pool of 600, that elicited strong emotional responses categorized as awe, disgust, or content. These images were used for EEG recordings of 52 volunteers, with channel selection and sample size estimation based on the analysis of an existing dataset. Gamma and delta bands yielded the highest classification accuracy, with the gamma band achieving an accuracy of 77.07 percent +/- 13.8 percent for the awe emotion. Key factors such as greenery and non-uniform granularity were linked to positive emotions, while dampness triggered negative reactions. These results emphasize the significance of incorporating natural elements and varied textures in biodigital architecture to enhance aesthetic appeal and acceptance. The study demonstrates EEG's capability to objectively assess architectural preferences, providing valuable insights for architects to design engaging and sustainable environments.
ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection
While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarc…
高齢者の認知障害の検出と管理における技術の進歩: 傾向、課題、および将来の方向性
人口の高齢化に伴い、軽度認知障害(MCI)から認知症への認知機能の低下は、今後数十年間の健康上の決定的な課題となっていますが、日常的な評価ではその初期の兆候が見逃されることがよくあります。この記事では、人工知能 (AI)、機械学習 (ML)、深層学習 (DL) によって統合された、神経生理学的信号 (主に脳波、EEG)、構造および分子神経画像 (MRI およびアミロイド/タウ PET)、血液ベースのバイオマーカー、デジタル マーカーに及ぶ、高齢者の認知障害の検出と管理に関する最近の技術進歩を批判的に総合しています。要約するだけでなく、学際的な分類法、対象と施設に依存しない検証を前景化する方法論的厳密さのレンズ、段階的スクリーニングと介入を結び付ける統合的早期発見フレームワーク、検出方法、介入、リスク因子と防御因子の比較表に貢献します。 EEG マーカー (アルファ/シータ変化、P300 潜時) とディープ モデル (CNN、LSTM/BiLSTM、トランスフォーマー、自己教師あり EEG 基礎モデル) は高い精度を報告しますが、その多くは小規模な単一サイト データセットに依存しており、厳密な外部検証に耐えられる可能性は低いです。その他の分野でも、成果は目に見えています。血漿 p-tau217 は臨床用途に達し、2025 年にはアルツハイマー病の診断を助ける最初の血液検査が承認されました。抗アミロイド療法(レカネマブ、ドナネマブ)は、効果が控えめで議論があるにもかかわらず承認されています。そしてマルチドメインのライフスタイル予防が成熟しました。ウェアラブル、リモート、音声、および仮想現実ツールにより、生態学的に有効な継続的なモニタリングが可能になり、マルチモーダル融合により感度と特異性が向上します。標準化、説明可能性、データプライバシー、外部から検証された公平な展開などの障壁が残っています。この分野の短期的な期待は、早期発見を実用的で個別化されたケアに結びつける、信頼性が高く、マルチモーダルで、長期的に検証されたシステムにあります。
原文 (English)
Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions
As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades, yet routine assessment often misses its earliest signs. This article critically synthesizes recent technological advances for detecting and managing cognitive impairment in older adults, spanning neurophysiological signals (chiefly electroencephalography, EEG), structural and molecular neuroimaging (MRI and amyloid/tau PET), blood-based biomarkers, and digital markers, integrated through artificial intelligence (AI), machine learning (ML), and deep learning (DL). Beyond summarizing, it contributes a cross-disciplinary taxonomy, a methodological-rigor lens foregrounding subject- and site-independent validation, an integrative early-detection framework linking tiered screening to intervention, and comparison tables of detection methods, interventions, and risk and protective factors. EEG markers (alpha/theta changes, P300 latency) and deep models (CNNs, LSTM/BiLSTM, transformers, self-supervised EEG foundation models) report strong accuracy, yet many rest on small, single-site datasets unlikely to survive rigorous external validation. Elsewhere, gains are tangible: plasma p-tau217 has reached clinical utility, with the first blood test cleared to aid Alzheimer's diagnosis in 2025; anti-amyloid therapies (lecanemab, donanemab) are approved despite modest, contested benefits; and multidomain lifestyle prevention has matured. Wearable, remote, speech, and virtual-reality tools enable continuous, ecologically valid monitoring, and multimodal fusion improves sensitivity and specificity. Barriers remain: standardization, explainability, data privacy, and equitable, externally validated deployment. The field's near-term promise lies in trustworthy, multimodal, longitudinally validated systems linking early detection to actionable, personalized care.
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the r…
UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations
High-polyphony symbolic music is increasingly used in generation, analysis, and arrangement, yet many downstream tasks require bounded repr…
Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
Model families are typically trained size by size, each from scratch. Can a pretrained large model instead be converted into a smaller sibl…
The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk
AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate human…
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can ass…
Agentic AI: User Empowerment or Foreclosure?
Agentic AI promises systems that can act on users' behalf, from filtering content to negotiating prices to selecting services. Whether it w…
Effect of Abstractions and Prompting Strategies on LLM-Guided High-Performance Optimizations
Code performance optimization is a vital aspect of modern software development, as it enables faster response times and reduced resource us…
Population-Scalable Multi-Agent World Modeling
World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent…
Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories
Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering task…
Persistent Recursive Worlds Enable Autonomous Software Evolution
Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems pre…
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-c…
Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models
We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization p…
Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework
Purpose: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilin…
Learning from Unreachable Rewards: Hint-Conditioned Reinforcement Learning for Generative Recommendation
Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autore…
No One to Blame: A Framework of Constitutive AI Unaccountability
The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominant…
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We…
AI コーディング エージェントによる仕様優先の収束: テスト オラクルも人間によるコード レビューも行わず、717,000 行のコードベース内の 189 ファイルにわたるコア アーキテクチャの不変条件を解体するケース スタディ
このペーパーでは、生成されたコードの人によるレビューや、ターゲットの動作を検証するための既存のオラクルを使用しない、仕様優先プロトコルに基づく AI コーディング エージェントによる大規模なアーキテクチャ リファクタリングの、完全に装備された単一のケース スタディを報告します。このタスクは、相互依存する大規模なコードベース全体で中心となる不変条件を解体するというもので、作成者は、インクリメンタル リファクタリング (従来は代わりに書き換えが必要だった種類の変更) では事実上実行不可能であると評価しました。ここで説明されているプロトコルに従って、エージェントは正常に完了しました。このシステムは、3,648 ファイルにわたる 717,725 行のプロダクション TypeScript アプリケーションです。このタスクでは、コアのライフタイム不変条件、つまり AI リクエストの間 UI パネルが開いたままであることの保証を解体する必要がありました。目標の動作は、ストリーミング世代がパネルの終了後も存続し、再度開いたときに、損失や重複なしに同じライブ ストリームに再接続できることです。プロトコル: エージェントによる正式な仕様、その仕様をソース コードに対して監査する 14 回の改良サイクル、アトミックな実装、コンパイル/テストのフィードバック ループ、その後、凍結された仕様に対してコードを監査する 17 回の検証サイクル。 31 回の監査パスを通じて、人間がプログラムを実行する前に 201 個の欠陥が修正されました。収束基準は経験的であり、2 つの連続した検証パスで結果がゼロになるというものでした。この変更は 189 個のファイル (31 個の新規ファイル) に影響を与えました。抽出フェーズでは、2 つのコミットで合計 288 ファイル、34,770 件の挿入、16,422 件の削除がコミットされました。最初のセッションとその後の約 30 回のセッションでは、ソフトウェアは指定どおりに動作し、バグは観察されませんでした。経過: 3 日。料金: 2,430 ドル。完全な仕様と生のセッション ログ (フランス語で 1,500 ページ以上) が証拠として公開されており、プロセスの検査と一貫性チェックのための言語モデルへの送信が可能になります。
原文 (English)
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully. The system is a 717,725-line production TypeScript application across 3,648 files. The task required dismantling a core lifetime invariant: the guarantee that a UI panel remains open for the duration of an AI request. The target behaviour was that a streaming generation survives the closing of its panel and can be reattached, on reopening, to the same live stream with no loss or duplication. The protocol: formal specification by the agent, 14 refinement cycles auditing that specification against the source code, atomic implementation, a compile/test feedback loop, then 17 verification cycles auditing the code against the frozen specification. Across 31 audit passes, 201 defects were corrected before any human executed the program. The convergence criterion was empirical: two consecutive verification passes returning zero findings. The change touched 189 files (31 new); with the extraction phase, the two commits total 288 files, 34,770 insertions, 16,422 deletions. Across the first and roughly thirty later sessions, the software behaved as specified, no bug observed. Elapsed: three days; cost: USD 2,430. The full specification and raw session logs, 1,500+ pages in French, are published as evidence, allowing inspection of the process and submission to a language model for consistency checking.
EgoCITE: 長期自己中心的メモリのためのコンテキスト拡張インデックス作成と時間認識検索
長期的な自己中心的な記憶は、連続する一人称のビデオとオーディオを、検索可能な過去の経験の記録に変換します。既存のシステムには 2 つのボトルネックがあることを示します。文脈に乏しいキャプションから構築されたインデックスはエージェント検索では信頼できません。また、検索では質問の一時的な意図が無視されます。両方のボトルネックに対処するために、自己中心的な QA のための長期的なエージェント メモリ フレームワークである EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval) を導入します。 EgoCITE は 3 つのコンポーネントで構成されます。 EgoScheme は、ローカルのマルチモーダル コンテキストを使用して、断片的なビデオ キャプションと音声トランスクリプトを自己完結型のアトミック メモリ インデックスに変換します。 EgoIndex は、相補的なアクション、アクティビティ、発話、および会話表現を、複数の粒度で検索可能なマルチビュー メモリ インデックスに編成します。 EgoRetrv は、セマンティック検索と、質問条件付きの時間的関連性スコアリングおよび取得された証拠のキュレーションを組み合わせたものです。 EgoLifeQA、EgoMem、および EgoR1-Bench 上の EgoCITE を、回答の精度とターゲットとイベントの検索の整合性の観点から評価します。 EgoCITE は、エージェント メモリ ベースラインの精度を少なくとも 4.4 ~ 14.2\% 向上させ、ロング コンテキスト LLM エージェントよりも 36$\times$ のコスト削減を実現します。
原文 (English)
EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory
Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic search, while retrieval ignores a question's temporal intent. To address both bottlenecks, we introduce EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric QA. EgoCITE comprises three components. EgoScheme uses local multimodal context to turn fragmentary video captions and speech transcripts into self-contained atomic memory indices. EgoIndex organizes complementary action, activity, utterance, and conversation representations into searchable multi-view memory indices at multiple granularities. EgoRetrv combines semantic search with question-conditioned temporal relevance scoring and curation of retrieved evidence. We evaluate EgoCITE on EgoLifeQA, EgoMem, and EgoR1-Bench in terms of answer accuracy and target-event retrieval alignment. EgoCITE improves accuracy over agentic memory baselines by at least 4.4--14.2% while achieving 36$\times$ lower cost than long-context LLM agents.
Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs
This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Advers…
Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency
Background: Software bugs remain a critical challenge in development, necessitating effective Automated Program Repair (APR) techniques. Wh…
From Fixed Grids to Moving Particles:A Transferable Latent Operator for Fluid Dynamics
Lagrangian modeling is vital to fluid dynamics, as it characterizes particle transport and complements the Eulerian representation. However…
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating su…