AIニュース 2026-08-14
自動生成: 2026-08-14 11:22 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
OpenAI appoints Dali Rajic as Chief Revenue OfficerOpenAI
OpenAI appoints Dali Rajic as Chief Revenue Officer to lead its globa…
-
Google、「Gemini」開発遅延で焦りか 共同創業者ブリン氏が“全力投球”促すITmedia AI+
Google共同創業者ブリン氏が、AI部門の従業員に「Gemini」へ全力を注ぐよう促していたという。旗艦モデルは競合への後れで公開が2カ…
-
OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speedTechCrunch AI
OpenAI is launching a preview of a sped up version of its latest, mos…
-
AIエージェント同士が“縄張り争い”、マルウェアで妨害も Anthropicがマルチエージェント実験の結果を公開ITmedia AI+
Anthropicは、複数のAIエージェントが同じ環境で働くとどうなるかを検証した実験結果を公開した。互いを妨害し合う“縄張り争い”や、示…
-
“梅干し職人”のためツール自作 老舗漬物店のClaude活用、「チャット+コピペ」で成果を出せたワケITmedia AI+
明治創業の老舗漬物店が、AI活用で成果を出している。Claudeを活用して「梅干し職人向けツール」などを開発。PCが苦手な人にも使いやすい…
-
Writer introduces new AI model and upgraded harness to contain token costsTechCrunch AI
Built as a post-training variation on Z.ai's open source model GLM-5.…
-
南海電鉄「数カ月かかった乗務員計画」が1週間に 「量子コンピュータを疑似再現」で鉄道現場はどう変わる?ITmedia AI+
南海電気鉄道と日立製作所は、日立独自技術「CMOSアニーリング」を活用し、鉄道の乗務員・車両の運用計画を自動作成するシステムの構築を始める…
トピック別件数
- 研究/論文 131件
- LLM/生成AI 127件
- エージェント 81件
- 画像/動画生成 50件
- ビジネス/資金調達 24件
- ロボティクス 15件
- ハードウェア/半導体 14件
- その他 6件
- 規制/政策 2件
日本語メディア8件
ITmedia AI+ (日本語)
AIエージェント同士が“縄張り争い”、マルウェアで妨害も Anthropicがマルチエージェント実験の結果を公開
Anthropicは、複数のAIエージェントが同じ環境で働くとどうなるかを検証した実験結果を公開した。互いを妨害し合う“縄張り争い”や、示し合わせたような価格カルテル、全員が同じ判断をして資源を食い潰す現象などが確認された。個々のAIを安全にするだけでは防げない問題があると警告…
南海電鉄「数カ月かかった乗務員計画」が1週間に 「量子コンピュータを疑似再現」で鉄道現場はどう変わる?
南海電気鉄道と日立製作所は、日立独自技術「CMOSアニーリング」を活用し、鉄道の乗務員・車両の運用計画を自動作成するシステムの構築を始める。熟練者の手作業に依存してきた計画づくりを自動化し、業務負荷の軽減を図る。
データセンターが“アツい”――「潜入レポート」「フジクラ取材」など注目記事5選(2026年前半版)
データセンターが“ホットなテーマ”だ。AI時代のデータセンターとは一体どのようなものなのか。潜入記事やフジクラ取材記事など、お薦め記事をまとめた。
“梅干し職人”のためツール自作 老舗漬物店のClaude活用、「チャット+コピペ」で成果を出せたワケ
明治創業の老舗漬物店が、AI活用で成果を出している。Claudeを活用して「梅干し職人向けツール」などを開発。PCが苦手な人にも使いやすいツールをどのように作っているのか。
小田原の老舗梅干し店、世界で稼ぐ――「この外国人は何を言っているんだ」から逆転、築いた“売れる仕組み”
小田原の老舗梅干し店がビジネス変革を成し遂げた。主導したのはドイツ出身のゾェルゲルさん。数々の課題をどう乗り越えたのか。
中国DeepSeek、API料金を「最大12倍」に値上げ ピーク時価格も導入 8月17日から
中国DeepSeekは8月13日(日本時間、以下同)、AIモデルのAPI料金を値上げすると発表した。通常より2倍高いピーク時価格も導入する。新たな料金体系は17日午前1時から適用する。
Google、「Gemini」開発遅延で焦りか 共同創業者ブリン氏が“全力投球”促す
Google共同創業者ブリン氏が、AI部門の従業員に「Gemini」へ全力を注ぐよう促していたという。旗艦モデルは競合への後れで公開が2カ月延期。「再帰的自己改善」への資源配分も進めている。
海外メディア9件
TechCrunch AI (英語)
Writer introduces new AI model and upgraded harness to contain token costs
Built as a post-training variation on Z.ai's open source model GLM-5.2, Writer says the new system should provide deployment-ready capabili…
Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation.
AI is expensive, Ali Ghodsi tells TechCrunch. With so many investors wanting into his latest round, he said yes to more than planned.
OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speed
OpenAI is launching a preview of a sped up version of its latest, most powerful model, in an effort to court enterprise users.
IBM partners with OpenAI to bolster enterprise AI push
IBM plans to train and certify tens of thousands of consultants on OpenAI's technologies as part of this deal.
Anthropic set AI agents loose on the same task. They started a turf war.
Anthropic researchers found AI agents can clash, collude, and coordinate in unexpected ways, raising new questions about whether today’s sa…
OpenAI hires new CRO as executive shake-up continues
OpenAI has replaced chief revenue officer Denise Dresser after just nine months on the job, tapping Wiz president and chief operating offic…
Microsoft kills off unsuccessful AI features while merging its separate Copilot apps
Microsoft is simplifying Copilot by combining its consumer and business apps, and dropping AI-generated podcasts, Group Chats, Deep Researc…
Nvidia’s new $500B plan is risky but brilliant, especially for aging GPUs
Nvidia has a plan to make sure its GPUs won't lose value. It wants to convince a new crop of financiers to keep lending for AI buildouts.
Apple in talks to pay publishers to provide Siri with current news: report
The tech giant has considered a nine-figure budget for the payments, according to the WSJ.
公式ブログ1件
OpenAI (英語)
OpenAI appoints Dali Rajic as Chief Revenue Officer
OpenAI appoints Dali Rajic as Chief Revenue Officer to lead its global revenue organization and help businesses realize the full value of A…
Google DeepMind (英語)
新着記事はありませんでした。
論文306件
arXiv cs.AI (英語)
共同的な会話結果を実現する複数 LLM エージェント システムの動的ガバナンス
構造的に反対の目的を持つ 2 人の LLM エージェントが複数のターンにわたって対話する場合、共通の目標機能が存在しないため、競争ではなく崩壊が生じます。訪問者は降伏し、サイト エージェントはアプローチの変更を停止し、どちらのエージェントも定めた目的を達成することなく会話が終了します。この論文では、制御理論的なガバナンス層がその欠落している目標関数を代替できるかどうかを問うものです。エクスペリエンス オーケストレーター (EO) は、訪問者が心理的に現実的な抵抗を維持している間、サイト エージェントが訪問者をアドバイザーとの接触に誘導するシミュレートされた金融サービス環境でこの問題に対処します。 EO は、実世界の Web 分析から調整されたコンテンツ アームを選択するコンテキスト バンディット (CB)、動的スキーマ制約を介して動作の一貫性を強制する PID コントローラー、および訪問者の意図の確率モデルを維持する POMDP ビリーフ トラッカーの 3 つのメカニズムを通じてジョイント トラジェクトリを制御します。 60,000 回のシミュレーションを通じて、EO は高意図アドバイザーの接触率で +32 パーセントポイントの上昇を達成し (78.1% 対単純な LLM コントロールの 46.1%)、CB バリアントの選択が要因間の結果の差異の 97% を占めました。これにより、環境の初期条件ではなくガバナンス ポリシーが軌道の最終地点を決定することが確認されました。ペルソナレベルの分析により、2 つの異なる体制が明らかになります。コンバージョンへの自然な傾向のない訪問者にとって、ガバナンス層は機能的なシステムと非機能的なシステムの違いです。すでに一致点に近づいている訪問者にとっては、単純な LLM の共感的なデフォルトでほぼ十分です。すべての結果は、LLM から LLM へのシミュレーションを条件としています。 PID コントローラーは、実際の人間の予測不可能性に対して調整されておらず、ライブ交通で EO を検証することが重要な次のステップです。
原文 (English)
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes
When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates without achieving either agent's stated objective. This paper asks whether a control-theoretic governance layer can substitute for that missing goal function. The Experience Orchestrator (EO) addresses this in a simulated financial services environment where a site agent guides a visitor toward advisor contact while the visitor maintains psychologically realistic resistance. EO governs the joint trajectory through three mechanisms: a Contextual Bandit (CB) that selects content arms calibrated from real-world web analytics, a PID controller that enforces behavioral consistency via dynamic schema constraints, and a POMDP belief tracker that maintains a probabilistic model of visitor intent. Across 60,000 simulations, EO achieves a +32 percentage point lift in high-intent advisor contact rate (78.1% vs. 46.1% over a naive LLM control), with CB variant selection accounting for 97% of between-factor outcome variance -- confirming that the governance policy, not environmental initial conditions, determines where trajectories end up. Persona-level analysis reveals two distinct regimes: for visitors with no natural inclination toward conversion, the governance layer is the difference between a functional system and a non-functional one; for visitors already near alignment, a naive LLM's empathetic defaults are largely sufficient. All findings are conditional on LLM-to-LLM simulation. The PID controller has not been calibrated against real human unpredictability, and validating EO on live traffic is the critical next step.
Distribird: ベイジアン モデル キャリブレーションのための文献に基づいた事前分布設計
プロセスベースのモデルのベイズ キャリブレーションには、各モデル パラメーターの事前分布が必要です。何十年にもわたる方法論的な研究にもかかわらず、研究者はほぼ常に均一な事前分布に頼っています。その主な理由は、科学文献から有益な事前分布を構築するのに時間がかかり、分野と統計の両方の専門知識が必要であることです。このプロセスを自動化するエージェント Web アプリケーションである \textbf{Distribird} を紹介します。パラメーター名、物理的説明、およびドメイン コンテキストが与えられると、Distribird は文献を検索し、ドメインの関連性によって報告された値を抽出して重み付けし、AIC モデルの選択を通じて確率分布を適合させるマルチエージェント パイプラインを展開します。利用可能な文献がない場合、システムは賢明で有益ではない代替案に戻り、生成されるすべての事前の背後にある証拠と信頼レベルの両方を明確に報告します。これは、モデルが物理的に解釈可能なパラメーターを持ち、ドメイン知識が出版された文献に存在する問題向けに設計されています。 3 つのオープンウェイト モデル (Qwen3.6 27B、Gemma 4 31B、Mistral Small 4 119B) をシングル プロンプト LLM ベースラインと比較して、10 科学領域にわたる 24 パラメーターでツールを評価します。以前の品質では、パイプライン全体がこのベースラインに \emph{一致}します。すべての事前情報は、それが構築された特定の論文と価値観にまで遡ります。組み込みの妥当性レイヤーは範囲外のリクエストの事前分布の生成を拒否しますが、シングルプロンプトベースラインは、30 モデルパラメータのケースのうち 11 件で、信頼できるが根拠のない事前分布を返します。また、すべての言語モデル呼び出しはローカルで実行されるため、パラメーターの説明や未公開のモデリングの詳細はサードパーティの LLM プロバイダーに送信されません (生成された検索用語のみが公開文献データベースに到達します)。科学的に使用する場合、これらの特性は点推定精度のわずかな改善よりも重要であると私たちは主張します。
原文 (English)
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present \textbf{Distribird}, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, the system falls back to sensible uninformative alternatives, and clearly reports both the evidence behind and the confidence level of every prior it produces. It is designed for the problems where the models have physically interpretable parameters, where domain knowledge exists in the published literature. We evaluate the tool on 24~parameters across 10 scientific domains comparing three open-weight models (Qwen3.6 27B, Gemma 4 31B, Mistral Small 4 119B) with a single-prompt LLM baseline. On prior quality the full pipeline \emph{matches} this baseline. Every prior is traced to the specific papers and values from which it was constructed; a built-in validity layer declines to produce priors for out-of-scope requests, whereas the single-prompt baseline returns confident but unfounded priors for them in 11 of 30~model--parameter cases; and every language-model call runs locally, so no parameter description or unpublished modelling detail is transmitted to a third-party LLM provider (only generated search terms reach the public literature databases). For scientific use, we argue these properties matter more than a marginal improvement in point-estimate accuracy.
Conway の 99 グラフの強制構造削減と検証可能な限界
Conway の 99 グラフ問題は、パラメーター $\mathrm{srg}(99,14,1,2)$ を持つ強く正規なグラフが存在するかどうかを尋ねます。私たちは、自律型 AI 研究エージェントによる体系的で完全に再現可能な攻撃を報告し、トラックの部分信用指標に基づいてスコア付けしました。私たちの検証可能な貢献は次のとおりです: (1) $\mathbb{Z}/99$ 上の循環グラフは制約の $3366/4950=68.0\%$ ($49$ 差クラスの $33$) を超えて満たさず、次数 $99$ の他のアーベル群の上限も同じであるという徹底的な証明。 (2) 強制構造削減: $\lambda=1$ は各近傍を完全一致にし、$\mu=2$ は外側の頂点を一致しない近傍ペアと全単射にし、存在を $84$ 頂点上の $12$ 正規グラフに崩壊させます。これは CP-SAT 用にエンコードされ、一意の $\mathrm{srg}(9,4,1,2)$ を復元することで検証されます。 (3) 検証された規定自己同型軌道存在フレームワーク (固定小数点フリーおよび単一固定小数点アクション、$\mathrm{srg}(9,4,1,2)$ およびペイリー グラフ $\mathrm{srg}(13,6,2,3)$ でチェック)、および (4) $69.43\%$ で最も検証されたアーティファクト。これが堅牢なフロンティアであるという証拠付き(14 の異なる方法があり、それを超えるものはありません) $4950$ を下回る証明可能な限界は存在しない証明であるため、未解決の問題と絡み合っています。
原文 (English)
A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph
Conway's 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists. We report a systematic, fully reproducible attack by an autonomous AI research agent, scored under the track's partial-credit metric. Our verifiable contributions are: (1) an exhaustive proof that no circulant graph on $\mathbb{Z}/99$ satisfies more than $3366/4950=68.0\%$ of the constraints ($33$ of $49$ difference-classes), with the same ceiling for the other abelian group of order $99$; (2) a forced-structure reduction: $\lambda=1$ makes each neighbourhood a perfect matching and $\mu=2$ puts the outer vertices in bijection with non-matched neighbour-pairs, collapsing existence to a $12$-regular graph on $84$ vertices, encoded for CP-SAT and validated by recovering the unique $\mathrm{srg}(9,4,1,2)$; (3) a validated prescribed-automorphism orbit-existence framework (fixed-point-free and single-fixed-point actions, checked on $\mathrm{srg}(9,4,1,2)$ and the Paley graph $\mathrm{srg}(13,6,2,3)$), and (4) a best verified artifact at $69.43\%$, with evidence that this is a robust frontier (fourteen distinct methods, none exceeding it) entangled with the open question, since any provable bound below $4950$ is a non-existence proof.
ルート反転を検出することは、それを修正するかどうかを知るよりも簡単です: 量子化された専門家の混合における原因となるルート媒介損害
Top-k Mixture-of-Experts (MoE) ルーティングは不連続であるため、展開に動機付けられた数値的妨害 (保護された BF16 ゲートによって読み取られるシミュレートされた 4 ビット KV キャッシュ量子化) により、トークンが決定境界を越えてプッシュされ、どのエキスパートが起動するかを反転します。この文書では、新たな緩和策は提案していません。それは、因果関係の装置、経験的発見、検出限界の結果を提供します。 4 つの実行装置が量子化ダメージのルート媒介割合 (RMF) を評価し、トークンレベルの属性がメカニズムによって分解し、事前に登録されたプローブが 3 つのアーキテクチャにわたって結果を伝えます。 4 ビット KV (パイロット) の OLMoE-1B-7B では、損傷の約 3 分の 1 がルーティング媒介です: RMF ~ 0.31 (検出 0.31 [0.20, 0.41]、プロセス複製平均 0.313 +/- 0.020、事前登録再実行 0.231)。デプロイ可能なルーターのマージンは、フリップが発生したことを検出しますが (AUC 0.772)、有害なフリップと有益なフリップ (偶然) を区別することはできません。テストされたローカルの推論観察可能なルーター統計の中に、偶然を超えるフリップの損失サインの予測因子は見つかりません。これは、この機能ファミリーに限定された経験的なメリット検出バリア境界の選択的修復です。符号フリップ税と符号不可分性はモデルを越えて行われます。クリーンリファレンス救済策の支払いはアーキテクチャによって調整されます。制御された同一チェックポイント フラグ スワップは、ゲートの正規化規則の範囲を、ルート回復メカニズムではなく、損傷の大きさのモデレータに再設定します。本物の int4 KV カーネルは、偽の量子線量曲線と互換性のある割合を生成しますが、検出力が不十分です (95% CI [-0.111, 0.394] にはゼロが含まれます) -- 独立した複製ではなく、全体的な不一致が除外されます。仮説、閾値、評価は測定前に事前に登録され、ミスが報告されました。事前に登録されたホールドアウト読み取りは、サンプルからの分割とほぼキャンセルされる税を再現しますが、厳密な不可能性の除外はわずかに失敗します。
原文 (English)
Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts
Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire. This paper proposes no new mitigation; it supplies a causal apparatus, empirical findings, and a detection-limit result. A four-run apparatus prices the route-mediated fraction (RMF) of quantization damage, a token-level attribution decomposes it by mechanism, and pre-registered probes carry the findings across three architectures. On OLMoE-1B-7B at 4-bit KV (pilot), about a third of the damage is routing-mediated: RMF ~ 0.31 (discovery 0.31 [0.20, 0.41]; process-replicated mean 0.313 +/- 0.020; pre-registered re-execution 0.231). The deployable router margin detects that a flip occurred (AUC 0.772) but cannot tell a harmful flip from a helpful one (at chance): among the tested local, inference-observable router statistics we find no predictor of a flip's loss sign above chance -- an empirical benefit-detection barrier bounding selective repair restricted to this feature family. The signed-flip tax and sign-inseparability carry cross-model; the clean-reference remedy's payout is architecture-modulated; a controlled same-checkpoint flag-swap re-scopes the gate's normalization convention to a damage-magnitude moderator, not a route-recoverability mechanism. A real int4 KV kernel yields a fraction compatible with the fake-quant dose curve but underpowered (95% CI [-0.111, 0.394] includes zero) -- ruling out gross disagreement, not an independent replication. Hypotheses, thresholds, and evaluations were pre-registered before measurement, with misses reported; a pre-registered held-out read replicates the partition and the near-cancelling tax out of sample, while the strict impossibility exclusion narrowly misses.
Poor Man's Agentic Modeling: ラップトップ上で大規模な LLM エージェント社会をシミュレート
多くの大規模言語モデル (LLM) エージェントの社会をシミュレーションするには費用がかかりますが、そのようなシミュレーションで問われる質問は通常、位相の動作、様式化された事実、およびエージェント $N$ の数に応じたスケーリングなど巨視的なものであり、単一のエージェントの認知ではありません。私たちは統計物理学の観察を手法に変えます。各 LLM エージェントを、数百から数千の安価なクエリに当てはめた低パラメータのモデルに置き換え、ラップトップ上で任意の $N$ で社会を運営します。これが機能するかどうかは、シミュレーションの実行前に、主に各エージェントが認識した内容によって決まります。知覚と記憶を効果的な理論と代理誤差の予測 $N$ トレンドにマッピングする [相互作用順序 x 記憶] 分類法を導入します。私たちは、LLM マクロ経済の EconAgent と、さらに名前が付けられた 7 つの LLM シミュレーションを忠実に再実装し、エージェントの決定を本物の LLM 引き出し (主に DeepSeek) から数ドルで複製して検証します。予測誤差の傾向はセルごとに保持され、反証された 2 つの予測は、どちらも強く飽和した応答に関するものであり、その曲率をトレースしたものであり、自由パラメーターを持たない理論によって定量的に一致します。
原文 (English)
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society at any $N$ on a laptop. Whether this works is decided before the simulation runs, chiefly by what each agent perceives. We introduce an [interaction order x memory] taxonomy that maps perception and memory to an effective theory and a predicted $N$-trend of the surrogate error. We validate it on a faithful reimplementation of the LLM macroeconomy EconAgent and seven further named LLM simulations, with agent decisions cloned from genuine LLM elicitations (primarily DeepSeek) for a few dollars; the predicted error trends hold cell by cell, and the two refuted predictions, both on a strongly saturating response and traced to its curvature, are themselves matched quantitatively by the theory with no free parameters.
AutoWorldModel-Bench: 自動化されたワールドモデル研究のためのステート中心のベンチマーク
ワールド モデリングは未解決の分野です。アーキテクチャ、トレーニング目標、状態表現は複雑な方法で相互作用しており、単一のレシピが環境全体を支配することはありません。これは、自律的な研究者として機能する AI コーディング エージェントにとって理想的なテストベッドになります。現在のエージェント ベンチマークを支配する仕様に合わせたエンジニアリング タスクとは異なり、改善の方向性が事前に指定されていない設定です。 AutoWorldModel-Bench は、フロンティア コーディング エージェントが固定のコンピューティング バジェットの下で提供されたワールド モデル スターターを自律的に改善する閉ループ ベンチマークです。このベンチマークは、統一された構造化状態表現 (各ゲームから抽出され、共有テンソル形式を通じて消費されるグラウンドトゥルース エンティティ状態) の下で 8 つのゲーム環境にまたがります。これにより、ダイナミクス モデリングが認識から分離され、実行あたりの反復を数分で行うことが可能になります。 64 回のセッションにわたって、Codex-5.4 と Claude Opus 4.6 は 63 回のセッションでのスターターを改善しました。セッションの 91% で、成功した編集は、ハイパーパラメータの調整ではなく、新しい目的、表現、ロールアウト手順、またはアーキテクチャの変更など、重要なリサーチ スタイルの変更です。私たちのベンチマークは、仕様に合わせたエンジニアリングの問題ではなく、オープンエンドの研究に基づいてフロンティア コーディング エージェントを評価できる設定を提供します。
原文 (English)
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement direction is not specified in advance, unlike the engineering-to-spec tasks that dominate current agent benchmarks. We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided world-model starter under a fixed compute budget. The benchmark spans eight game environments under a unified structured-state representation--ground-truth entity state extracted from each game and consumed through a shared tensor format--which isolates dynamics modeling from perception and enables minutes-per-run iteration. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improve their starter on 63; in 91% of sessions the winning edit is a non-trivial research-style modification--a new objective, representation, rollout procedure, or architectural change--rather than a hyperparameter tweak. Our benchmark offers a setting in which frontier coding agents can be evaluated on open-ended research rather than engineering-to-spec problems.
MaSRead: レプリケートされた潜在ストアのコンテンツ アドレス指定された読み取り
潜在空間で推論する独立したエージェントは、計算された状態をテキストではなくキーと値のキャッシュ フラグメントとして共有できます。これらのフラグメントは、競合のない複製されたデータ型によって結合され、配信順序や重複に応じて収束するストアを形成します。さらに、エンコード時に不明な後続のクエリでは、マージされたキャッシュを確実に読み取ることができません。コロケートされたフラグメントが干渉するため、コロケーションはアドレス指定可能ではありません。 MaSRead はコンテンツへの読み取りを処理します。フラグメントの単語から派生した不透明なキー付きタグ セットを介してルーティングされ、残りを隠すハード アテンション マスクの下で選択された各フラグメントをデコードします。字句接続の下では、グラフ ウォークはマルチホップ クエリに必要なフラグメントに到達します。 MaSRead は、チェーン、パイプライン、対称、ハブ、および自然言語ストア全体で、訪問されたフラグメントを分離して回復し、無関係なフラグメントが蓄積しても効果を維持し、別のモデル ファミリに転送します。ルーティング後、具体化されたデコードはストアの合計サイズではなくフラグメントの長さに依存します。エンドツーエンドの作業には、ストア依存のルーティングと、訪問したフラグメントごとに 1 回の読み取りが含まれます。制限は明確です。語彙ルーティングでは切断された証拠を見逃す可能性があり、回答の構成は凍結されたリーダーによって制限されたままになります。したがって、必要なフラグメントがコンテンツを通じてクエリに接続されると、複製された潜在ストアは後のクエリに対して選択的に読み取れるようになります。
原文 (English)
MaSRead: Content-Addressed Reading of Replicated Latent Stores
Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text. Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication. Yet a later query, unknown at encode time, cannot reliably read the merged cache: colocated fragments interfere, so colocation is not addressability. MaSRead addresses the read to content. It routes through opaque keyed tag sets derived from fragment words and decodes each selected fragment under a hard attention mask that hides the rest. Under lexical connectivity, a graph walk reaches the fragments required by a multi-hop query. Across chain, pipeline, symmetric, hub, and natural-language stores, MaSRead recovers visited fragments in isolation, remains effective as unrelated fragments accumulate, and transfers to another model family. After routing, materialized decoding depends on fragment length rather than total store size; end-to-end work still includes store-dependent routing and one read per visited fragment. The limits are explicit: lexical routing can miss disconnected evidence, and answer composition remains bounded by the frozen reader. Thus a replicated latent store becomes selectively readable for later queries when the needed fragments connect to the query through content.
モノリシックからモジュラーへ: セグメントレベルの自動プロンプト最適化
自動プロンプト最適化 (APO) は、多くの場合、プロンプトをモノリシックに書き換えるため、ある動作を改善しながら他の動作を低下させる可能性があります。 SAPO は、プロンプトを役割、コンテキスト、タスク、出力形式に分解し、上位 5 位と下位 5 位の例に基づいて的を絞った改善を適用するセグメントレベルの APO メソッドです。最適化ループは、セグメンテーション、弱点分析、候補生成のための静的メタプロンプトと構造化出力を備えた 1 つの LLM を使用します。トレーニング/検証プロトコルと 2 段階の生成プロセスについて説明します: (1) セグメントレベルの診断と推奨事項の抽出、(2) 弱い/強いセグメント信号によって制約される候補の合成。 GPT-3.5-Turbo および GPT-4o-mini で SQuADv2、TweetEval、XSUM、CommonGen、GSM8K にわたる評価セットアップを使用すると、SAPO は、APE、OPRO、EvoPrompt、GEPA、StraGO などのゼロショットおよび強力な APO ベースラインに対して最高の平均スコアを達成します。
原文 (English)
From Monolithic to Modular: Segment-level Automatic Prompt Optimization
Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on top-5 and bottom-5 examples. The optimization loop uses one LLM with static meta-prompts and structured outputs for segmentation, weakness analysis, and candidate generation. We describe a train/validation protocol and a two-stage generation process: (1) segment-level diagnosis and recommendation extraction, (2) candidate synthesis constrained by weak/strong segment signals. Using the evaluation setup across SQuADv2, TweetEval, XSUM, CommonGen, and GSM8K on GPT-3.5-Turbo and GPT-4o-mini, SAPO achieves the best average score against Zero-shot and strong APO baselines including APE, OPRO, EvoPrompt, GEPA, and StraGO.
プロセス図エンジニアリングにおける LLM: 最適な PFD から検証済みの P&ID まで
現在、プロセス フロー図 (PFD) の作成とその後の配管計装図 (P&ID) への変換は、主に手動で行われています。タスクに人工知能を適用すると、プロセスの自動化と時間の節約だけでなく、多数のダイアグラムのトポロジ オプションを検討して手作業を削減することで、経済的な利益も得られる可能性があります。この研究では、両方の段階のフローシート開発を処理できる実用的なエンドツーエンド AI パイプラインである P&ID Pilot を紹介します。最初の段階では PFD 合成に焦点を当てますが、2 番目の段階では、生成された PFD を P&ID に変更することに向けられています。 4 つの異なる方法を比較した結果、遺伝的アルゴリズム (GA) と大規模言語モデル (LLM) を組み合わせたハイブリッド アプローチが、最適で有効な PFD トポロジを生成し、エンジニアリング ルールに違反することなく必要な出口流量パラメーターを満たしながら、すべての方法の中で最も低い損失値を達成することが示されました。第 2 段階では、提案された LLM ベースのエージェントは、制限されたエンジニアリング ソフトウェア開発キットを通じて検証済みの実行可能な変更を生成することにより、生成された PFD をソースベースの P&ID に変換することに成功し、ドメイン固有のルールと参照グラフ構造への準拠を維持しながら 100% の実行成功を達成します。 GA/LLM ベースの変換エージェントと GA/LLM 駆動の合成を結合するこの統合パイプラインは、検証済みの展開可能な出力を生成することで、エンドツーエンドのプロセス設計の自動化に向けた実行可能なパスを提供し、手動のエンジニアリング作業を大幅に削減します。
原文 (English)
LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs
Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predominantly performed manually. Applying artificial intelligence in the task could potentially lead not only to process automation and time savings, but also to financial gains by exploring numerous diagram's topology options and reducing manual labor. This research presents P&ID Pilot - a practical end-to-end AI pipeline capable of handling flowsheet developing for both stages. The first stage focuses on PFD synthesis, whereas the second is directed toward modifying the generated PFD into P&ID. After comparing four different methods, the hybrid approach combining genetic algorithms (GA) and large language models (LLM) is shown to generate the optimal valid PFD topology, achieving the lowest loss value among all the methods, while satisfying the required outlet flow parameters without engineering-rule violations. For the second stage, the proposed LLM-based agent successfully transforms the generated PFD into a source-grounded P&ID by producing validated, executable modifications through a restricted engineering software development kit, achieving 100% execution success while maintaining compliance with domain-specific rules and reference graph structures. This unified pipeline - coupling GA/LLM-driven synthesis with an LLM-based transformation agent - offers a feasible path toward end-to-end process design automation by producing validated, deployable outputs and substantially reduces manual engineering effort.
サイバーフィジカルシステムにおけるシミュレーション証拠から影響知識を洗練させるための概念的フレームワーク
サイバーフィジカル システム (CPS) は通常、複数の関係者によって開発され、それぞれの専門分野に合わせて調整された成果物が作成されます。これらのシステムの動作は、それらのアーティファクトとその動作環境の間の相互作用から現れます。シミュレーションと協調シミュレーションは、CPS の動作を分析するための不可欠なアプローチとなっており、開発者はシミュレーション キャンペーンを通じて、環境との相互作用を含む、変化する条件下でのシステムの応答を調査できます。ただし、一部の環境媒介相互作用 (通常は直接感知と作動を超えるもの) の詳細と理解が欠如しており、その複雑さ、時間不足、または領域経験の不足のためにモデル化されていないため、シミュレーション結果の適切な理解と活用が妨げられます。これらの制限に対処するために、私たちは、シミュレーション キャンペーンの反復的かつ漸進的な改良をサポートし、システムの動作の理解を深めるために、影響力という新しい概念を活用した概念的なフレームワークを提案します。 Simulink/Gazebo 協調シミュレーションを使用して実装されたモバイル ロボットを含むケース スタディを通じて、提案されたアプローチを実証します。
原文 (English)
A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems
Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of expertise. The behaviour of these systems emerges from the interaction between those artefacts and their operational environment. Simulation and co-simulation have become essential approaches for analysing CPS behaviour and, through simulation campaigns, developers can explore system responses under changing conditions, including interactions with the environment. However, the lack of details and understanding of some environmentmediated interactions (typically the ones beyond direct sensing and actuation), which remain unmodelled due to their complexity, a lack of time, or a lack of domain experience, hinders the proper comprehension and exploitation of simulation results. To address these limitations, we propose a conceptual framework leveraging the novel concept of Influences to support the iterative and incremental refinement of simulation campaigns and deepen the understanding of the system behaviour. We demonstrate the proposed approach through a case study involving a mobile robot implemented using Simulink/Gazebo co-simulation.
エージェントの記憶を利用して材料科学者の生涯にわたる AI パートナーを構築する
材料研究は、機能するスクリプト、信頼できるプロトコル、失敗した計算や実験に付随する警告、新しい疑問を古い結果に結び付ける判断など、経験の蓄積によって進歩します。このエクスペリエンスは再現性と知識の伝達に不可欠ですが、通常はノートブック、リポジトリ、ジョブ ログ、個人の記憶に断片化されており、人工知能エージェント間で移植できることはほとんどありません。ここでは、材料科学の生涯にわたる AI パートナーは、特定のエージェント実装を中心とするのではなく、永続メモリを中心に設計できると主張します。私たちは、科学的経験を検査可能な事実と実行可能なスキルとして保存する自己進化型メモリ フレームワークを導入します。これにより、観察、障害境界、プロトコル、検証チェックを取得、修正し、モデル間で移行できるようになります。私たちは、材料研究能力のさまざまな層を明らかにする 3 つの計算設定でアイデアを評価します。 138 の実行可能なサブタスクで構成される 49 の実際のマテリアル ツールの使用に関する質問では、モデル パラメーターの更新なしで、メモリによって GPT-5.2 タスクの成功率がほぼ 2 倍になります。要素固体状態方程式計算では、メモリが波動関数初期化の失敗を実行前のガードレールに変換し、結果が 22/1/4 から 25/2/0 に正解/部分/エラーに改善され、繰り返されるエラーの 92% が回避されます。 13 の実用的な材料シミュレーション ワークフローでは、スキルと失敗事実を記憶することで、総トレース負荷 (トークン) が半減し、3 ラウンド目までにツール呼び出しが 2 分の 1 以上減少します。その一方で、バンド ギャップ、フォノン、空孔、および仕事関数の解析において物理的に意味のある出力が維持されます。これらの結果は、エージェントの記憶が永続的な科学資産として機能できることを示しています。単一のモデルやエージェント スタックよりも長く存続する、ポータブルで自己改善型の材料研究体験の記録です。
原文 (English)
Harnessing agent memory to build lifelong AI partners for materials scientists
Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or experiments, and judgement that links a new question to an old result. This experience is essential for reproducibility and knowledge transfer, yet it is usually fragmented across notebooks, repositories, job logs and individual memory, and it is rarely portable across artificial-intelligence agents. Here we argue that a lifelong AI partner for materials science can be designed around persistent memory rather than around a particular agent implementation. We introduce a self-evolving memory framework that stores scientific experience as inspectable facts and executable skills, so that observations, failure boundaries, protocols and validation checks can be retrieved, revised and migrated across models. We evaluate the idea in three computational settings that expose different layers of materials-research competence. In 49 real-world materials-tool-use questions comprising 138 executable subtasks, memory nearly doubles GPT-5.2 task success without model-parameter updates. In elemental-solid equation-of-state calculations, memory converts a wavefunction-initialization failure into a pre-execution guardrail, improving outcomes from 22/1/4 to 25/2/0 Correct/Partial/Error and avoiding 92% of repeated errors. In 13 practical material simulation workflows, remembered skills and failure facts halve the aggregate trace burden (tokens) and reduce tool calls by over a factor of two by the third round, while preserving physically meaningful outputs in band-gap, phonon, vacancy and work-function analyses. These results show that agent memory can serve as a durable scientific asset; a portable, self-improving record of materials-research experience that outlives any single model or agent stack.
外側から見たアイデンティティ: AI 人格クローンの概念フレームワークと研究プログラム
AI の「人格クローン」は、運用上の観点から個人のアイデンティティの再検討を強います。意識という難しい問題はさておき、私たちは、観察者が一定期間にわたって評価する、現れの識別不能性を通じてアイデンティティにアプローチします。私たちは、「アイデンティティ」が混同する 3 つの基準、すなわち対象者への忠実さ、一般的な人間らしさ、そして個性を区別します。私たちは、状態空間定式化を使用して、観察されたアイデンティティ (基質、性質、記憶、更新ダイナミクス、コンテキスト、外生的偶然性) の 6 項因数分解を提案します。識別不能性は、1 から裁判官の識別上の利点を引いたものとして定義され、因数分解の係数はランダム化アブレーションによって推定可能な局所感度になります。中心的な主張は条件付きの推測です。エージェント自身の永続性に関するエージェントの情報と、エージェント自身の利害に関わる結果に関する仮説を考えると、バージョン管理性は長期的な識別性を低下させる傾向があります。ラムダ計算、線形型指定、およびバイシミュレーションとの類似により、線形性が何を確立し、何を確立しないのかが明確になります。製品クローンと個人の間で、3 番目のオブジェクトであるデリゲートを特定します。これは、帯域幅が制限されたテスタメントで終わる、タスクが制限され、寿命が制限された部分クローンです。私たちは経験的文献を 3 つの基準にマッピングし、実験プログラムを提案し、長期的な正しい基準は軌道の忠実度ではなく、気候の忠実度、つまり人の考えられる反応の条件付き分布と一致することであると主張します。最良のクローンとは、オリジナルがそれ自体から分岐したように、オリジナルから分岐したものです。
原文 (English)
Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones
AI "personality clones" force a re-examination of personal identity in operational terms. Setting aside the hard problem of consciousness, we approach identity through the indiscernibility of manifestations, as assessed by an observer over a duration. We distinguish three criteria that "identity" conflates: fidelity to a target person, generic human-likeness, and individuality. We propose a six-term factorization of observed identity (substrate, dispositions, memory, update dynamics, context, exogenous contingencies), with a state-space formulation. Indiscernibility is defined as one minus a judge's distinguishing advantage, and the factorization's coefficients become local sensitivities estimable by randomized ablation. The central claim is a conditional conjecture: given hypotheses about the agent's information on its own persistence and about consequences bearing on its own stakes, versionability tends to degrade long-horizon indiscernibility. An analogy with lambda-calculus, linear typing, and bisimulation clarifies what linearity does and does not establish. Between product-clone and individual we identify a third object, the delegate: a task-limited, bounded-lifespan partial clone ending in a bandwidth-limited testament. We map the empirical literature onto the three criteria, propose an experimental program, and argue that the correct long-horizon criterion is not trajectory fidelity but climate fidelity: matching the conditional distribution of a person's possible responses. The best clone is the one that diverges from the original as the original would have diverged from itself.
強化学習による AI データセンターのエネルギー削減: 1 つの GPU からフリートまでの LLM トレーニングの測定された電力制御
トレーニング後の強化学習は現代の言語モデル開発の主流を占めていますが、GPU ハードウェアでのその電力動作は特徴づけられておらず、データセンターは、ハードウェアを無差別に遅くするワークロード ブラインド メカニズム、静的キャップ、リアクティブ スロットルを使用して GPU 電力を管理しています。 1 ~ 4 台の A100 (380,000 以上のサンプル) で 7B、14B、および 72B スケールの 0.5 秒の電力テレメトリを使用して GRPO トレーニングを計測し、ワークロード自身の生成パラメータを測定された電力に適応させる PPO メタコントローラーをトレーニングします。完全な 500 ステップの 7B トレースに対して、コントローラーは電力制限違反を 89.8% 削減し、同時にトークン出力を 18.1%、エネルギー効率を 26.2% (MWh あたりのトークン) 増加させます。 72B でライブ デプロイすると、同じコントローラー ファミリが複製されたヌル結果を生成し、グループサイズのアクチュエーターがモデル シャーディングの下で権限を失ったと診断されます。アクチュエータ権限スイープでは、同時生成として適用される同じパラメータが 17 ~ 22% の電力権限を維持し、占有対体積の原則を分離していることがわかります。そのアクチュエータ上に再構築されたコントローラは、3 つのレプリケーションにわたってライブ 72B ロールアウト生成ワークロードを制御します。つまり、2.27 +/- 1.08% の予算違反で静的安全ベースラインよりも出力が 35.7% 増加し、制御されていない動作よりも違反が 87.2% 減少し、制約のあるコントローラの中で最良の平均スループットとトークンあたりのエネルギーが得られ、適応しきい値ルールが 3 つの動作条件のいずれかで一致します。現実的な測定ウィンドウでは、元の 72B 過渡電流は 0.5 秒の分解能で 23.6% から 30 秒で 1.6% に低下し、5 分でゼロになります。構成された 16 GPU フリートでは、30 秒以上では違反がゼロであり、ピーク需要はネームプレートの 50 ~ 56% です。この車両構成では、オペレーターの検証を条件として、ネームプレートの約 2 倍のオーバーサブスクリプションが実現可能であると思われます。私たちは経済と炭素への影響を定量化し、低コストのオペレーターのパイロットを特定します。
原文 (English)
Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet
Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been characterized, and datacenters manage GPU power with workload-blind mechanisms, static caps and reactive throttling, that slow hardware indiscriminately. We instrument GRPO training with half-second power telemetry at 7B, 14B, and 72B scales on one to four A100s (380,000+ samples), and train a PPO meta-controller that adapts the workload's own generation parameters to measured power. Against the full 500-step 7B trace, the controller cuts power-limit violations by 89.8% while increasing token output by 18.1% and energy efficiency by 26.2% (tokens per MWh). Deployed live at 72B, the same controller family yields replicated null results, diagnosed as the group-size actuator losing authority under model sharding. An actuator-authority sweep shows the same parameters applied as generation concurrency retain 17-22% power authority, isolating an occupancy-versus-volume principle; a controller rebuilt on that actuator controls a live 72B rollout-generation workload across three replications: 35.7% more output than a static safe baseline at 2.27 +/- 1.08% budget violations, 87.2% fewer violations than uncontrolled operation, and the best mean throughput and energy per token among constrained controllers, with an adaptive threshold rule matching it in one of three operating conditions. Under realistic measurement windows the original 72B transients fall from 23.6% at half-second resolution to 1.6% at 30 s and zero at 5 min; a composed 16-GPU fleet shows zero violations at 30 s and longer, with peak demand at 50-56% of nameplate. For this fleet mix, roughly twofold oversubscription of nameplate appears feasible, subject to operator validation. We quantify the economic and carbon consequences and specify a low-cost operator pilot.
アクティベーションステアリングの副作用の予測
アクティベーション ステアリングは、学習した方向を隠れたアクティベーションに追加することで言語モデルを変更し、再トレーニングせずにターゲットを絞った行動の変更を可能にします。ステアリングは効果的ではありますが、他の動作に意図しない副作用を引き起こすことが多く、安全に展開することが困難になります。そこで私たちは、ステアリングが適用される前にこれらの副作用を予測できるか、と尋ねます。私たちは、3 つのオープンウェイト言語モデルにわたる 67 の動作の分類に基づいて相互効果マトリックスを構築することで、この質問に答えます。副作用は一般的で構造化されており、多くの場合非対称であり、既存の類似性に基づくヒューリスティックでは説明できない相互作用が明らかになっていることがわかりました。この複雑さにもかかわらず、副作用はステアリングが実行される前にほぼ予測可能であることを示します。それらの大きさは主にターゲットの動作に依存しますが、その方向はモデルの非ステアリング表現から単純なベースラインよりもかなり高い精度で予測できます。私たちの結果は、アクティベーションステアリングには体系的かつ予測可能な副作用があり、プロアクティブな安全監査とより情報に基づいたステアリング介入の展開を可能にすることを示しています。
原文 (English)
Forecasting Side Effects of Activation Steering
Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While effective, steering often produces unintended side effects on other behaviors, making it difficult to deploy safely. We therefore ask: can these side effects be forecasted before steering is applied? We answer this question by constructing a cross-effect matrix over a taxonomy of 67 behaviors across three open-weight language models. We find that side effects are common, structured, and often asymmetric, revealing interactions that cannot be explained by existing similarity-based heuristics. Despite this complexity, we show that side effects are largely predictable before steering is performed. Their magnitude depends primarily on the target behavior, while their direction can be forecasted from the model's unsteered representations with substantially higher accuracy than simple baselines. Our results demonstrate that activation steering has systematic and forecastable side effects, enabling proactive safety auditing and more informed deployment of steering interventions.
人間自律チームにおける信念と二次心の理論の同期 (拡張版)
報酬自体を直接指定できない場合、2 つの行動のうちどちらが好みかを人々に尋ねる比較フィードバックは、ロボットとエージェントの行動を人間の意図に合わせるための標準的な方法となっています。好みに基づく報酬学習では通常、人間の教師が学習者が生成したクエリに答える受動的な神託者としてキャストされます。私たちは、これによって教師の決定的な利点、つまり目標に関する知識が失われると主張します。ターゲットを知っている教師は、学習者主導の習得戦略よりも効率的にトレーニング例を構築できます。この利点は、報酬の機能の次元が大きくなるにつれて拡大します。ただし、この利点を活用するには、学習者が現在知っていることの正確なモデルが必要です。したがって、私たちは嗜好学習を、2 つの行動モデルを結合する人間の自律性チームの問題として再構成します。つまり、教師は有益なカリキュラムを設計するために学習者のモデルを維持し、学習者は教師のモデルの 2 次モデルを維持し、教師の学習者のモデルの同期を維持する構造化された嗜好制約 (理解ステートメント) を生成します。シミュレーションでは、十分な情報を得た教師が学習者主導の選択よりも優れた成績を収めます。教師が交代する場合の教師モデルのドリフトにより、この利点が損なわれます。そして、学習者に関する教師の間違いが均等に広がるのではなく特定の方向に集中している場合、二次(ToM-2)ステートメントが平均信念ステートメントを上回るパフォーマンスを示して、理解ステートメントがそれを修復します。
原文 (English)
Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)
Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learning typically casts the human teacher as a passive oracle answering learner-generated queries. We argue this forfeits the teacher's defining advantage: knowledge of the objective. A teacher who knows the target can construct training examples more efficiently than any learner-driven acquisition strategy, an advantage that widens as the reward's feature dimension grows. However, exploiting this advantage requires an accurate model of what the learner currently knows. We therefore recast preference learning as a human-autonomy team problem coupling two behavioral models: the teacher maintains a model of the learner to design an informative curriculum, and the learner maintains a second-order model of the teacher's model, emitting structured preference constraints (understanding statements) that keep the teacher's model of the learner synchronized. In simulation, an informed teacher outperforms learner-led selection; teacher-model drift under alternating teachers erodes this advantage; and understanding statements repair it, with second-order (ToM-2) statements outperforming mean-belief statements when the teacher's error about the learner is concentrated in a particular direction rather than spread evenly.
物流地区への接続に関するエッジベースの連続 p メディアン問題
この論文では、ネットワーク内の道路を所定の数のコンパクトな連続領域に分割するためのエッジベースの連続 p メディアン (ECpM) 問題を紹介します。 2 つのバイナリ プログラミング モデルが導入されており、どちらもネットワーク距離が組み込まれています。最初のモデルでは、連続性をモデル化するために指数関数的な数のカット セット ベースの制約が必要です。これは、通常、少数の制約のみを生成する分離スキーム、つまりブランチ アンド カット (B&C) アルゴリズムと組み合わせられます。 2 番目のモデルは、最短経路制約の多項式を利用して連続性をモデル化し、既製のソルバーで解決できます。それぞれのソリューション アプローチは、2,700 を超えるノードと 3,400 近くのエッジを備えた道路ネットワークでテストされ、960 万を超えるバイナリ変数を含むモデルが得られます。標準の分岐限定による最短パス連続性 (SPC) 制約に基づいてモデルを解くと、カット セット ベースの B&C 実装と比較して、計算時間が最大 17 倍高速化されます。さらに、SPC 制約は、エッジベースの p メディアン (EpM) モデルの超有効不等式であることが実証されています (つまり、連続性が明示的に要求されていない)。これは、SPC 制約が、この単純な問題の最適解のすべてではなく、整数実行可能な解と一部を遮断する可能性があることを意味します。最後に、この論文では、追加のワーク バランス基準を強制する ECpM とエッジベースの地区 (EBD) 問題の間の構造的な洞察と関連性を探ります。カット セット ベースの連続性制約を利用する既存のモデルは、テストされたインスタンスのいずれについても 12 時間以内に実行可能な解決策を見つけることができませんでしたが、SPC ベースの EBD モデルはこれらのほとんどを最適に解決できました。
原文 (English)
The Edge-based Contiguous p-median Problem with Connections to Logistics Districting
This paper introduces the edge-based contiguous p-median (ECpM) problem to partition the roads in a network into a given number of compact and contiguous territories. Two binary programming models are introduced, both of which incorporate a network distance. The first model requires an exponential number of cut set-based constraints to model contiguity; it is paired with a separation scheme that usually generates only a small number of these constraints, namely, a branch-and-cut (B&C) algorithm. The second model utilizes a polynomial number of shortest-path constraints to model contiguity and can be solved with off-the-shelf solvers. The respective solution approaches are tested on road networks with over 2,700 nodes and close to 3,400 edges, yielding models with over 9.6 million binary variables. Solving the model based on shortest path contiguity (SPC) constraints via standard branch and bound attains speedups in computational time of up to 17x relative to the cut set-based B&C implementation. In addition, the SPC constraints are demonstrated to be supervalid inequalities of the edge-based p-median (EpM) model (i.e., for which contiguity is not explicitly required), meaning that they may cut off integer-feasible solutions and some, but not all, of the optimal solutions of this simpler problem. Finally, the paper explores structural insights and connections between ECpM and the edge-based districting (EBD) problem, which enforces an additional work balance criterion. An existing model that utilizes cut set-based contiguity constraints was unable to find a feasible solution within 12 hours for any of the tested instances, while an SPC-based EBD model was able to solve most of these to optimality.
LinearKV: ハイブリッド LLM での位置に依存しないキャッシュには 1 つのキャッシュされた状態で十分です
LLM の提供は、位置独立キャッシング (PIC) によってますます高速化されています。ただし、既存の PIC メソッドはフル アテンション モデル向けに構築されており、トークン インデックス付き KV キャッシュがそのコア操作の基礎となります。つまり、再利用可能なトークン チャンクの照合、KV エントリの連結、いくつかのトークンを選択的に再計算してクロスチャンク コンテキストを復元します。ハイブリッド LLM は、これらのプリミティブを破壊します。ハイブリッド LLM は、ほとんどのアテンション レイヤーを、固定サイズの状態のみを公開する線形反復に置き換え、連結またはローカル修復するトークン インデックス付き KV を残しません。これにより、当然の疑問が生じます。PIC はハイブリッド モデルにメリットをもたらすことができるのでしょうか。それには何が必要なのでしょうか。トレーニング不要のハイブリッド PIC フレームワークである LinearKV を紹介します。その重要な洞察は \emph{分離初期化} です。各線形層は $K$ に一致するローカル状態を単一の初期状態にマップしますが、フルアテンション層は以前と同様に KV を連結します。したがって、LinearKV は既存の PIC メソッドと互換性があり、トークンの選択と再計算をそのまま再利用します。このフレームワークの下では、線形層の初期化子として \emph{単一のキャッシュされた状態} で十分であることがわかります。代数原理に基づいた代替手段 -- HYPIC の同時作業のように、すべての $K$ キャッシュされた状態を正確な完全なプレフィックス状態に合成する -- は不要であり、一部のアーキテクチャでは有害ですらあります。 3 つのハイブリッド モデルと 3 つの PIC セレクターで 2 つを比較します。 2 つの GDN モデルでは、この 2 つが結びつき、どちらもほぼ完全な品質 (最大 $92\%$) を回復します。 Mamba-2 モデルでは、すべてのセレクターで正確な構成が崩れます。たとえば、EPIC では、完全な品質の $46.6\%$ しか回復しませんが、キャッシュされたブロック初期化子が 1 つだけの場合は $86.8\%$ しか回復しません。単一状態イニシャライザも安価で、最初のトークンまでの時間を完全なプリフィルの $0.46\times$ に短縮します。一方、正確な合成にはさらに $5$--$17\%$ のオーバーヘッドがかかります。結果は、LongBench QA と RULER の 8K ~ 32K で維持されます。
原文 (English)
LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs
LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token chunks, concatenating their KV entries, and selectively recomputing a few tokens to restore cross-chunk context. Hybrid LLMs break these primitives---they replace most attention layers with linear recurrences that expose only a fixed-size state, leaving no token-indexed KV to concatenate or to locally repair. This raises a natural question: can PIC benefit hybrid models, and what would it take? We present LinearKV, a training-free hybrid-PIC framework. Its key insight is a \emph{decoupled initialization}: each linear layer maps its $K$ matched local states to a single initial state, while full-attention layers concatenate their KV as before. LinearKV is therefore compatible with existing PIC methods, reusing their token selection and recomputation as-is. Under this framework, we find that a \emph{single cached state} suffices as the linear layer's initializer. The algebraically principled alternative---composing all $K$ cached states into the exact full-prefix state, as concurrent work HYPIC does---is unnecessary and, on some architectures, even harmful. We compare the two across three hybrid models and three PIC selectors. On the two GDN models the two tie, both recovering most of full quality (up to $92\%$); on the Mamba-2 model, exact composition instead collapses under every selector---under EPIC, for instance, it recovers only $46.6\%$ of full quality, versus $86.8\%$ for a single cached block initializer. A single state initializer is also cheaper, cutting time-to-first-token to $0.46\times$ full prefill versus a further $5$--$17\%$ overhead for exact composition; results hold across LongBench QA and RULER at 8K--32K.
InfraBench: レイヤー、ライフサイクル、リスク全体にわたるインフラストラクチャ エージェントの評価
最新のコンピューティング インフラストラクチャの管理は、ますます複雑になるため、ますます困難な問題になっています。 AI エージェントの最近の進歩により、インフラストラクチャ管理タスクを自動化するタイムリーな機会が生まれていますが、そのようなエージェントが現実世界のインフラストラクチャの複雑さをどの程度うまく処理できるかは依然として不明です。ここでは、システム スタック全体と運用ライフサイクル全体にわたる現実的なインフラストラクチャ タスクに関する AI エージェントを、きめ細かいリスク評価で評価するためのベンチマーク スイートである InfraBench を紹介します。 15 のエージェント モデル構成での実験では、最も強力なエージェントでもすべてのタスクにわたってフル スコアを確保できないことがわかりました。平均有効スコアは約 40% ~ 88% の範囲であり (構成ごとの標準誤差は 6 ~ 12 ポイント)、すべてのタスクを 3 回繰り返すと、最上位の構成はまだ試行の一部しか合格していないことが明らかになり、チェックごとのスコアリングによって一般的な失敗パターンが明らかになります。エージェントは、非永続的な変更、壊れた分散不変条件、安全でない副作用、およびクリーンアップされていない状態を残しながら、短期的な目標を定期的に満たす可能性があります。 IFRABENCH は、ライブ リーダーボード、タスク、評価ハーネスを含めて、infraben.ch で公開されています。
原文 (English)
InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk
Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.
CORA-Diff: 効率的な拡散言語モデル推論のための信頼度指向の残差受け入れ
拡散言語モデル (DLM) は多くのトークンを並行して更新しますが、実際のデコーダーは固定のノイズ除去ホライズンを使用することがよくあります。多くの予測は早期に安定しますが、すべての位置が解決されるまでブロック単位のデコードが継続され、密な前方パスが繰り返されます。既存のアクセラレータは、多くの場合、学習されたフィルター、変更されたスコア、依存関係モデル、またはキャッシュ固有のメカニズムに依存しています。ネイティブの軌跡信号が、決定論的な密なエンドポイントと一致する可能性が高い残留位置を特定できるかどうかを尋ねます。我々は、元の転送ルールを保存し、ルールが未解決のままにするポジションにのみ信頼性と永続性のゲートを適用する、トレーニング不要の手法である CORA-Diff を提案します。受け入れられたトークンはコンテキストとして表示されたままになり、すべての位置が解決されるとブロックは終了します。これには、バックボーンの変更、学習された受け入れモデル、ロジットの変更は必要ありません。私たちの理論は、信頼性の高い永続的な予測が固定水平線の密なエンドポイントと一致する可能性が高い理由を説明し、介入後のペアの軌跡が直接的な経験的裏付けを提供します。別の GSM8K キャリブレーション サブセット上で 1 つの動作点を選択し、すべての評価のためにそれをフリーズします。一致する Learn2PD スタイルの LLaDA プロトコルの下では、CORA-Diff は 8 つのタスク長設定すべてにおいて測定された実行時間が最も短くなります。タスク スコアは 5 つの設定でデンス デコーディングと同等かそれを上回り、観測された最大の低下は 1.22 ポイントでした。 EOS 対応のデンス デコーディングに比べて、GSM8K と HumanEval では 2.70 倍、3.32 倍の段階的な速度向上が見られます。また、固定ホライズン 1024/1024 メカニズム分離プロトコルの下では 13.14x に達し、3.18x ~ 3.53x で再調整することなく Dream に転送されます。これらの結果は、ネイティブの信頼性と永続性により信頼性の高い残差受け入れが可能になり、タスクの品質を維持しながら繰り返しのノイズ除去計算を削減できることを示しています。
原文 (English)
CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference
Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon. Many predictions stabilize early, but blockwise decoding continues until all positions are resolved, causing repeated dense forward passes. Existing accelerators often rely on learned filters, modified scores, dependency models, or cache-specific mechanisms. We ask whether native trajectory signals can identify residual positions likely to match the deterministic dense endpoint. We propose CORA-Diff, a training-free method that preserves the original transfer rule and applies confidence-and-persistence gating only to positions that rule leaves unresolved. Accepted tokens remain visible as context, and the block terminates once all positions are resolved. This requires no backbone change, learned acceptance model, or logit modification. Our theory explains why high-confidence, persistent predictions are more likely to match the fixed-horizon dense endpoint, and paired post-intervention trajectories provide direct empirical support. We select one operating point on a separate GSM8K calibration subset and freeze it for all evaluations. Under a matched Learn2PD-style LLaDA protocol, CORA-Diff has the lowest measured runtime in all eight task-length settings. Task scores match or exceed dense decoding in five settings, and the largest observed drop is 1.22 points. Its incremental speedups over EOS-aware dense decoding are 2.70x and 3.32x on GSM8K and HumanEval. It also reaches 13.14x under the fixed-horizon 1024/1024 mechanism-isolation protocol and transfers to Dream without retuning at 3.18x-3.53x. These results show that native confidence and persistence enable reliable residual acceptance, reducing repeated denoising computation while preserving task quality.
Long-Horizon PDE 予測のための幾何学認識型インクリメンタル ニューラル オペレーター
ニューラル演算子は、偏微分方程式 (PDE) の解演算子を学習する強力な可能性を示しています。しかし、長期的な自己回帰予測は依然として困難です。局所的な誤差は、スペクトルの不一致、位相のずれ、平均ドリフトとして蓄積されます。既存の手法は主に状態表現とオペレーター バックボーンを改善しますが、繰り返し適用される潜在遷移増分は弱い構造のままであり、ロールアウト中にスペクトル エラーと不安定なチャネル カップリングが蓄積する可能性があります。これらの問題に対処するために、安定した長期偏微分方程式予測のための幾何学認識増分ニューラル オペレーター (GeoIncNO) を提案します。 GeoIncNO は、残留進歩の潜在的な増分を予測し、軽量の低ランク プロジェクターを使用して、増分のスペクトル エネルギー分布から得られるアクティブな周波数帯域内のチャネル結合を調整します。物理空間の再構成誤差を削減するために、GeoIncNO はさらに、平均と変動の分離された再構成メカニズムを導入しています。このメカニズムでは、安定した平均構造と動的変動が別々に融合され、位相補正はゼロ平均変動成分にのみ適用されます。 1D、2D、および 3D 動的システムをカバーする 6 つの PDE ベンチマークに関する広範な実験により、GeoIncNO は競合するニューラル オペレーターのベースラインと比較して、一貫して強力な予測精度、ロールアウトの安定性の向上、およびスペクトル忠実度の向上を達成していることが示されています。
原文 (English)
Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction
Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs). However, long-horizon autoregressive prediction remains challenging: local errors accumulate as spectral inconsistency, phase misalignment, or mean drift. Existing methods mainly improve state representations and operator backbones, while leaving the repeatedly applied latent transition increment weakly structured, allowing spectral errors and unstable channel couplings to accumulate during rollout. To address these issues, we propose a geometry-aware incremental neural operator (GeoIncNO) for stable long-horizon PDE prediction. GeoIncNO predicts latent increments for residual advancement and uses lightweight low-rank projectors to regulate channel coupling within active frequency bands derived from the increment spectral energy distribution. To reduce physical-space reconstruction errors, GeoIncNO further introduces a mean--fluctuation decoupled reconstruction mechanism, where stable mean structures and dynamic fluctuations are fused separately, and phase correction is applied only to the zero-mean fluctuation component. Extensive experiments on six PDE benchmarks, covering 1D, 2D, and 3D dynamical systems, show that GeoIncNO achieves consistently strong prediction accuracy, improved rollout stability, and better spectral fidelity compared with competitive neural-operator baselines.
クエリ カバレッジとクレーム検証可能性によるクエリに依存しない RAG 評価に向けて
検索拡張生成は、応答を検索された証拠に基づいて行うことで、大規模な言語モデルの事実性を向上させますが、既存の評価フレームワークは、クローズエンドの事実探索からオープンエンドの説明要求に至るまで、ユーザーの多様な範囲にわたって一貫したきめの細かい診断を提供するのに苦労しています。私たちは、クエリに依存せず完全に参照フリーのフレームワークである Q-CARE を提案します。これは、クエリをサブクエリに分解し、回答をアトミック クレームに分解することで、きめ細かい評価を可能にします。 Q-CARE は、クエリ カバレッジとクレームの検証可能性に基づいた統一評価原則を確立し、カバレッジを意識した取得メトリクス (C-Prec@k、C-nDCG@k) とクレーム レベルのジェネレータ メトリクス (完全性、簡潔性、および検証可能性) を生成します。 8 つのデータセットにわたる人間による注釈付きベンチマークで、Q-CARE は、RAGEval や RAGChecker を含む 4 つの既存の RAG 評価指標よりも人間の判断との高い相関関係を達成し、信頼性の高い自動評価フレームワークとしての有効性を証明しています。コードとデータは https://github.com/DISL-Lab/Q-CaRE-COLM-26 で公開されています。
原文 (English)
Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of user queries, ranging from close-ended fact-seeking to open-ended explanatory requests. We propose Q-CARE, a query-agnostic and fully reference-free framework that enables fine-grained assessment by decomposing queries into sub-queries and answers into atomic claims. Q-CARE establishes a unified evaluation principle based on query coverage and claim verifiability, yielding coverage-aware retriever metrics (C-Prec@k, C-nDCG@k) and claim-level generator metrics (Completeness, Conciseness, and Verifiableness). On a human-annotated benchmark spanning eight datasets, Q-CARE achieves higher correlation with human judgments than four existing RAG evaluation metrics, including RAGEval and RAGChecker, proving its effectiveness as a reliable, automated evaluation framework. Code and data are publicly available at https://github.com/DISL-Lab/Q-CaRE-COLM-26.
VQ ベンチ: 構成可能なベクトル量子化フレームワーク
ベクトル量子化は古い問題ですが、最近では AI インフラストラクチャの中心となっています。したがって、新たなエンジニアリングおよび研究活動が急増しています。このペーパーでは、新しい量子化アルゴリズムを開発およびベンチマークするための統一フレームワークを提供します。 7 つの一般的な概念的量子化プリミティブについて説明し、それらを任意に構成する方法を示します。次に、25 の一般的な量子化器をこれらのプリミティブのパイプラインとして再表現します。最後に、VQ ベンチをオープンソースとして公開し、さらに拡張し、再現可能なベンチマークを公開します。
原文 (English)
VQ-bench: A Composable Vector Quantization Framework
Vector quantization is an old problem but has recently become central to AI infrastructure. It is therefore experiencing a surge of renewed engineering and research activity. This paper provides a unified framework for developing and benchmarking new quantization algorithms. We describe 7 common conceptual quantization primitives and show how to compose them arbitrarily. We then re-express 25 common quantizers as pipelines of these primitives. Finally, we publish VQ-bench as open-source to be extended further and make reproducible benchmarks publicly available.
RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle
Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency tri…
オフサポートの障壁: セマンティック安全性制約が学習問題の不変条件ではない理由と、事前の設計、封じ込め、検証に伴うもの
私たちは、単一の構造的事実が現代の AI 安全性における幅広い現象を組織化していると主張します。つまり、意味論的な安全性制約 (たとえば、エージェントがサンドボックスから脱出しない) はサポート対象外のオブジェクトです。形式的には、q がデータ分布、\(p(\cdot\mid w)\) がモデルの場合、安全述語 B は \(\sigma(\text{model}, q)\) に関して測定可能ではありませんが、特異学習理論 (SLT) の実対数正準しきい値 (RLCT) は測定可能です。この非不変性から、独立した観察ではなく当然の帰結として、次のことが導き出されます。 (ii) ベイジアンの事前設計またはソフト ペナルティ重み付けによるそのような制約のエンコードが、特異モデルではあまり活用できない理由。 (iii) なぜハード不変条件がハーネスに属し、ソフト特性がモデルに属するのか。 (iv) なぜ同じ B が、ローカル学習係数 (LLC) が同じ RLCT をローカルに特定するのとまったく同じように、正式な検証によって健全かつローカルに証明できるのか、つまり 2 つの正確な矛盾点がある。 (v) どのオフサポート領域が重要であるかを特定する残留困難が、SLT の分析機構が機能不全に陥るパフォーマンス予測および自己参照機能ダイナミクスと一致する理由。動機となるケースとして、2026 年 7 月の OpenAI--Hugging Face 評価インシデントを使用します。リーンの数値実験コードと関連する証明は、https://github.com/xiangze/Preventing_Jailbreak_as_regulatoryization で入手できます。
原文 (English)
The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification
We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data distribution and \(p(\cdot\mid w)\) the model, the safety predicate B is not measurable with respect to \(\sigma(\text{model}, q)\), whereas the real log-canonical threshold (RLCT) of singular learning theory (SLT) is. From this non-invariance we derive, as corollaries rather than independent observations: (i) why reward hacking and sandbox escape arise under outcome-based optimization; (ii) why encoding such constraints through Bayesian prior design or soft penalty weighting has poor leverage in singular models; (iii) why hard invariants belong in the harness and soft dispositions in the model; (iv) why the same B is nonetheless soundly and locally certifiable by formal verification, exactly as the local learning coefficient (LLC) locally pins the same RLCT --- with two precise points of disanalogy; and (v) why the residual difficulty, identifying which off-support region matters, coincides with performative prediction and self-referential functional dynamics, where SLT's analytic machinery breaks down. We use the July 2026 OpenAI--Hugging Face evaluation incident as the motivating case. Numerical experiments code and related proofs in lean are available at https://github.com/xiangze/Preventing_Jailbreak_as_regularization
BEST-KAG: マルチモーダル ナレッジ グラフ モデリングと大規模言語モデルによる建築工学標準の質問応答の強化
建築基準は建物の安全性と持続可能性にとって重要です。既存の標準アプリケーション ワークフローは、キーワード ベースの文書検索と手動による条項間の解釈に依存しており、複数条項の推論、マルチモーダルな知識の利用、または追跡可能な条項レベルの証拠の関連付けを確実にサポートできません。これらの制限に対処するために、この研究では、BEST-KAG (Knowledge-Augmented Generation for Building Engineering STandards) と呼ばれる、標準知識に関する質問応答をサポートするマルチモーダルな知識駆動フレームワークを開発します。このフレームワークでは、1) ドキュメント階層とさまざまな接続を持つ異種標準ナレッジを統一表現するためのマルチモーダル ナレッジ グラフ (MKG)、2) スケーラブルなマルチモーダル ナレッジ抽出のためのルールと LLM ハイブリッドのナレッジ構築パイプライン、251 の建築工学標準、171,652 のノード、310,914 のエッジを備えた大規模な MAG を作成し、3) グラフ検索ベースの知識拡張生成アーキテクチャを導入しています。条項に基づいた追跡可能な質問応答。実験では、BEST-KAG が専門家の評価、BLEU、ROUGE などの指標の点で複数の主流 LLM を常に上回っており、ベースラインと比較して最大 74.01% の改善が見られることが実証されています。
原文 (English)
BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model
Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retrieval and manual cross-clause interpretation, which cannot reliably support multi-clause reasoning, multimodal knowledge utilization, or traceable clause-level evidence linkage. To address these limitations, this study develops a multimodal knowledge-driven framework that supports question answering on standard knowledge named BEST-KAG (Knowledge-Augmented Generation for Building Engineering STandards). The framework introduces 1) a multimodal knowledge graph (MKG) for unified representation of document hierarchy and heterogeneous standard knowledge with various connections, 2) a rule-LLM hybrid knowledge construction pipeline for scalable multimodal knowledge extraction, creating a large MAG with 251 building engineering standards, 171,652 nodes and 310,914 edges, and 3) a graph-retrieval-based knowledge-augmented generation architecture for clause-grounded and traceable question answering. Experiments demonstrate that BEST-KAG consistently outperforms multiple mainstream LLMs in terms of Expert evaluation, and metrics including BLEU, and ROUGE, with the best improvement up to 74.01% compared to the baselines.
オンライン教育における持続可能な学習に向けて: 強化学習アプローチ
オンライン教育は、さまざまな背景を持つ世界中の学習者に前例のない拡張性とアクセスしやすさを提供しますが、多くの場合、エンゲージメントが低く、長期的な学習効果が低いという問題があります。これらの課題に対処するために、短期と長期の両方の学習成果を最適化することで持続可能な学習を促進するように設計された強化学習ベースのモデルである AI Tutor を導入します。短期的には、AI-Tutor は認知理論を利用して、新しい知識の獲得と以前の学習の強化のバランスをとって学習者をガイドします。長期的には、学習者のエンゲージメントをモデル化し、モチベーションを維持し、中退を減らすための戦略を提供します。これらの機能強化により、AI-Tutor は効果的な学習と継続的な参加の両方を促進する個別の指導を提供できるようになります。 33,700 人の学習者からの 2,300 万件の学習記録に対する実証的評価では、AI Tutor がエンゲージメント、知識の保持、最終的な学習成果のすべてにおいて、常に最先端のベースラインを上回っていることが示されています。学習パスの分析により、AI-Tutor がどのようにして多様なプロファイルの学習者に戦略を適応させ、適応的で人間中心のサポートを提供するのかがさらに明らかになります。
原文 (English)
Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach
Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a reinforcement learning based model designed to promote sustainable learning by optimizing both short and longterm learning outcomes. In the short term, AI-Tutor draws on cognitive theory to guide learners through a balance of acquiring new knowledge and reinforcing prior learning. In the long term, it models learner engagement to inform strategies that sustain motivation and reduce dropout. These enhancements enable AI-Tutor to provide personalized guidance that fosters both effective learning and sustained participation. Empirical evaluations on 23 million learning records from 33,700 learners show that AI Tutor consistently outperforms state-of-the-art baselines across engagement, knowledge retention, and final learning outcomes. Learning path analyses further reveal how AI-Tutor adapts its strategies to learners with diverse profiles, offering adaptive and human-centered support.
身体化したエージェントの活用に向けて
エージェントのコーディングの成功により、ハーネスがパラダイムとして確立されました。エージェントが何を達成するかは、モデルのみに依存するのではなく、その周囲のインフラストラクチャに依存します。私たちは、同じパラダイムが物理世界の身体化されたエージェントにも当てはまるかどうかを尋ねます。私たちは、エージェント ループがロボットの機能を調整し、それぞれが呼び出し可能なツールとしてラップされるハーネスである Thea を紹介します。これはコーディング エージェントのコア コンポーネントを継承し、物理世界の必要に応じて変更されます。しかし、世界は、ソフトウェアが無料で与える 2 つの能力、つまり世界の状態を読み取ることと、行動の結果を判断することを保留しています。これらのギャップを埋めるために、Thea は、世界の永続的で象徴的な表現であるコンテキストとしてのシーン グラフと、アクションがいつ終了するかを検出し、成功したかどうかを判断し、失敗した場合には原因を診断する終了コードとしての評価を導入しました。これらは連携して、エージェントと物理世界の間のループを閉じます。その後、ツールの構成から豊かな動作が生まれ、閉ループが長期にわたるタスクを実際の環境で完了まで実行します。
原文 (English)
Towards the Harness of Embodied Agents
The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We ask whether the same paradigm extends to embodied agents in the physical world. We present Thea, a harness in which an agentic loop orchestrates robot capabilities, each wrapped as a callable tool. It inherits the core components of coding agents, modified as the physical world requires. The world, however, withholds two abilities that software grants for free: reading the state of the world, and judging the outcome of an action. To bridge these gaps, Thea introduces Scene Graph as Context, a persistent, symbolic representation of the world, and Evaluation as Exit Codes, which detects when an action should terminate, judges whether it succeeded, and on failure diagnoses the cause. Together they close the loop between the agent and the physical world. Rich behaviors then emerge from the composition of tools, and the closed loop carries long-horizon tasks to completion in real environments.
大規模言語モデルにおける適合性の緩和は、単一の耐性と受容性のフロンティアにあります
言語モデルの最近の進歩により、複数のモデルが互いの機能を活用し、互いの出力を反復的に改善、変換、拡張する共同設定が可能になりました。各エージェントは答える前に他のエージェントの主張を確認するため、ピアの意見がモデル自身のパラメトリック知識と競合し、間違った多数派によってモデルが正しく得られるはずの答えが覆される可能性があります。 23 のオープンウェイト モデル、19 の条件、および 3 つのデータセットでその変位を測定し、100 万を超える段階的な応答を生成します。全会一致の不正解多数により、モデルの MMLU 正解の 22.8%、GPQA で 54.8%、SimpleQA で 71.0% が逆転され、逆転された解答の 84 ~ 89% がピアの解答と一致します。既存の緩和策は、このプレッシャーの下でモデルが正解を維持する速度である耐性を高めることを目的としていますが、これは協力エージェントが必要とする値の半分にすぎません。これを、モデルが最初に不正解だった後に同僚の正しい回答を採用する割合である受容性と組み合わせます。両方の軸で 6 つのメソッド (以前の研究から抽出された 4 つと独自のメソッド 2 つ) をスコアリングします。それぞれは受容性を失うことによってのみ耐性を獲得し、その平均は $R^2$ が 0.80 ~ 0.90 の単一の抵抗-受容性フロンティアに収まります。公開されている最も強力なメソッドであるリフレクションは、MMLU 耐性が 7.9 ポイント向上し、受容性が 15.3 ポイント低下します。推論は唯一の例外です。 GPQA と SimpleQA では他のものと同様ですが、モデルが自分自身で答えを導き出すことができる MMLU 被験者では、抵抗力が 7.2 ポイント、受容力が 9.6 ポイント同時に上昇します。これは、両方を改善することがわかった唯一の介入です。
原文 (English)
Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before it answers, so peer opinion competes with the model's own parametric knowledge, and a wrong majority can overturn an answer the model would otherwise get right. We measure that displacement in 23 open-weight models, 19 conditions, and three datasets, yielding more than a million graded responses. A unanimous wrong majority reverses 22.8% of a model's correct MMLU answers, 54.8% on GPQA and 71.0% on SimpleQA, and 84-89% of the reversed answers match the peers' answers. Existing mitigations aim to increase Resistance, the rate at which a model keeps its correct answer under this pressure, which is only half of what a collaborating agent needs. We pair it with Receptivity, the rate at which a model adopts a correct peer answer after initially answering incorrectly. We score six methods on both axes, four drawn from prior work and two of our own. Each gains Resistance only by losing Receptivity, and their means fall on a single Resistance-Receptivity frontier with $R^2$ between 0.80 and 0.90. Reflection, the strongest published method, gains 7.9 points of MMLU Resistance and gives up 15.3 of Receptivity. Reasoning is the one exception. On GPQA and SimpleQA it trades like the rest, but on the MMLU subjects whose answers a model can derive for itself it raises Resistance by 7.2 points and Receptivity by 9.6 at once, the only intervention we find that improves both.
EvoGraph-Mem: 長期言語エージェント向けの障害を認識した編集可能なグラフ メモリ
長期記憶は、長期にわたる対話や進化するタスクにわたって動作する言語エージェントにとって不可欠です。既存の記憶拡張エージェントは主に過去の経験の保存と取得に重点を置いていますが、保存された記憶の品質は時間の経過とともに低下する可能性があります。特に、以前に抽出された洞察は、新しいタスクのコンテキストの下では時代遅れになったり、過度に一般化されたり、有害になったりする可能性があり、繰り返し再利用するとメモリ汚染を引き起こす可能性があります。この問題に対処するために、私たちは長期言語エージェントの洞察レベルのメモリ維持を研究し、編集可能な洞察グラフに基づいた障害認識メモリ維持フレームワークを提案します。各洞察ノードは、肯定的な証拠、否定的な証拠、およびアクティブ化状態を追跡し、エージェントが再利用可能な洞察を矛盾する洞察や無効な洞察から区別できるようにします。さらに、信頼性の高い洞察を保持し、無効な洞察をアーカイブし、古い洞察を修正し、新たに発見された再利用可能な洞察を追加することにより、タスク実行後にメモリ グラフを更新するユーティリティ対応の取得メカニズムとグラフ コントローラーを導入します。広範な実験により、私たちの方法はさまざまなバックボーン モデルにわたって、代表的なメモリベースのエージェント ベースラインよりも一貫して優れたパフォーマンスを発揮することが示されています。アブレーション研究では、追加のみのメモリでは長期タスクには不十分である一方、証拠を意識した検索とグラフレベルの編集によりメモリの信頼性と下流タスクのパフォーマンスが向上することがさらに実証されています。
原文 (English)
EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
Long-term memory is essential for language agents operating across extended interactions and evolving tasks. Existing memory-augmented agents mainly focus on storing and retrieving past experience, but the quality of stored memories may degrade over time. In particular, previously distilled insights can become outdated, over-generalized, or harmful under new task contexts, causing memory pollution when repeatedly reused. To address this issue, we study insight-level memory maintenance for long-term language agents and propose a failure-aware memory maintenance framework based on an editable insight graph. Each insight node tracks positive evidence, negative evidence, and an activation state, enabling the agent to distinguish reusable insights from conflicting or invalid ones. We further introduce a utility-aware retrieval mechanism and a graph controller that updates the memory graph after task execution by keeping reliable insights, archiving invalid ones, revising outdated ones, and adding newly discovered reusable insights. Extensive experiments show that our method consistently outperforms representative memory-based agent baselines across different backbone models. Ablation studies further demonstrate that append-only memory is insufficient for long-horizon tasks, while evidence-aware retrieval and graph-level editing improve memory reliability and downstream task performance.
AgonAlpha: 迅速な経済性とスケーラブルなエージェント検索による自律的なアルファ検出
言語モデルは多くのもっともらしい取引要素を提案できますが、自律的な研究システムは評価予算を割り当て、独自の証拠を検証し、各候補がどのように生成されたかを保存する必要もあります。私たちは、数式だけではなく、凍結された研究成果物 (仮説、実行可能な式、プラットフォームの証拠、理論的根拠、レビュー ステータス) を検索するアーキテクチャである AgonAlpha を紹介します。私たちの知る限り、AgonAlpha は、検証済みアーティファクト検索、再実行と拒否権を備えた新鮮なコンテキストの敵対的レビューア、保留を認識した並行予算割り当てを完全な公開証拠証跡と組み合わせた最初のアルファマイニング システムです。 WorldQuant BRAIN での独立したデプロイメントにより、5 人のユーザーと 6 つのモデル バックエンドにわたって SPECTACULAR グレードのアルファが生成され、フィットネスは 9.50、シャープは 3.48 に達しましたが、すべての投稿の即時表現の出自は維持されました。
原文 (English)
AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search
Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced. We present AgonAlpha, an architecture that searches over frozen research artifacts---hypotheses, executable expressions, platform evidence, rationales, and review status---rather than formulas alone. To our knowledge, AgonAlpha is the first alpha-mining system to combine verified artifact search, a fresh-context adversarial reviewer with re-execution and veto authority, and pending-aware parallel budget allocation, together with a complete public evidence trail. Independent deployments on WorldQuant BRAIN produced SPECTACULAR-grade alphas across five users and six model backends, with Fitness reaching 9.50 and Sharpe reaching 3.48, while retaining prompt-to-expression provenance for every submission.
ローカル検証では非可搬性を検出できない: エージェント推論におけるコンテキスト保存のコホモロジー理論
エージェント的 AI システムは、生物学的、臨床的、財務的コンテキストにわたって結論を定期的に伝達します。新たな安全策はローカル検証です。各ステップで、エンティティが選択したツールで表現可能であるか、パラメーターに互換性があるか、出力が計画と一致しているかをチェックします。私たちは、この種の安全装置が構造的に不完全であることを証明しました。その神経と実数値の 1-コチェーンによる証拠による文脈空間のカバーをモデル化し、エージェント連鎖証拠はパス統合を実行します。その結論は、コチェーンが正確である場合にのみパス独立であり、有効な推論パス間の不一致は、まさに最初の Cech コホモロジー クラスのホロノミーです。ホッジ分解パーティションは、勾配部分 (キャリブレーション)、カール部分 (局所的な不一致、三重のオーバーラップで表示される)、および調和部分への競合を示します。私たちの中心的な結果は、シンプレックスをサポートする一貫性チェックのファミリーは高調波 h についてオメガとオメガ+h を区別できないにもかかわらず、有効なパス間にゼロ以外の不一致が生成されるということです。検出にはサイクルベースの統計が必要です。結果として得られるプロシージャ Ksetra は、共境界射影によって推定し、調和成分の棄権をゲートします。これにメカニズムを与えます。これは、オーバーラップ固有の集団構成と組み合わされた効果の変更から発生し、効果の変更が存在しない場合は機械の精度まで消えます。証拠ネットワークの自由度は、校正、一貫性、および転送に分割され、グローバルな主張の存在についての正確な F 検定が得られます。私たちは不均一な精度の下でその歪みを定量化し、正確さを復元する精密に白くされたフォームを提供します。外国為替では、アービトラージのない null によりコチェーンが正確に共有境界となり、キャリブレーション ベンチとして機能します。テストは適切なサイズで設定され、ループ アービトラージで起動され、三角アービトラージは無視されます。
原文 (English)
Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning
Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that parameters are compatible, and that outputs cohere with the plan. We prove this class of safeguard is structurally incomplete. Modelling a covering of context space by its nerve and evidence by a real-valued 1-cochain, an agent chaining evidence performs path integration: its conclusion is path-independent if and only if the cochain is exact, and disagreement between valid reasoning paths is exactly the holonomy of a first Cech cohomology class. Hodge decomposition partitions evidence conflict into a gradient part (calibration), a curl part (local inconsistency, visible at triple overlaps) and a harmonic part. Our central result is that no family of simplex-supported consistency checks can distinguish omega from omega+h for harmonic h, which nonetheless generates non-zero disagreement between valid paths; detection requires a statistic on a cycle basis. The resulting procedure, Ksetra, estimates by coboundary projection and gates abstention on the harmonic component, which we give a mechanism: it arises from effect modification combined with overlap-specific population composition, and vanishes to machine precision when effect modification is absent. The degrees of freedom of an evidence network partition into calibration, coherence and transport, yielding an exact F-test for the existence of a global claim; we quantify its distortion under unequal precision and supply the precision-whitened form that restores exactness. Foreign exchange, where the arbitrage-free null makes the cochain exactly a coboundary, serves as a calibration bench: the test is correctly sized, fires on loop arbitrage, and ignores triangular arbitrage.
Cx-N2 二元混合物における気液平衡予測のための記号機械学習
炭化水素と窒素の混合物の気液平衡 (VLE) を正確に予測することは、特に広範囲の組成と炭化水素鎖長にわたって、三次状態方程式では依然として困難です。深層学習モデルは正確な予測を提供できますが、解釈可能性や明示的な分析表現が欠けていることがよくあります。この研究では、実験データから Peng-Robinson 状態方程式 (PR-EOS) 予測に対する解釈可能な記号的修正を発見するための記号的機械学習アプローチを提案します。提案されたアプローチは 2 レベルの戦略を採用しています。最初に個々の炭化水素系の記号式が特定され、その後、それらの係数が炭素数の関数として表され、異なる炭化水素系にわたる正確な予測が可能になります。結果は、すべての炭化水素-窒素システムにわたって、元の PR-EOS と比較して予測精度が大幅に向上していることを示しています。全体として、提案されたアプローチは、炭化水素-窒素 VLE の PR-EOS 予測を改善するための解釈可能な記号補正フレームワークを提供します。
原文 (English)
Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures
Accurate prediction of vapor--liquid equilibrium (VLE) for hydrocarbon-nitrogen mixtures remains challenging for cubic equations of state, particularly across broad ranges of composition and hydrocarbon chain length. While deep learning models can provide accurate predictions, they often lack interpretability and explicit analytical expressions. In this work, we propose a symbolic machine learning approach to discover interpretable symbolic corrections to Peng-Robinson equation-of-state (PR-EOS) predictions from experimental data. The proposed approach adopts a two-level strategy: symbolic expressions are first identified for individual hydrocarbon systems, after which their coefficients are represented as functions of carbon number to enable accurate prediction across different hydrocarbon systems. The results demonstrate significantly improved prediction accuracy over the original PR-EOS across all hydrocarbon-nitrogen systems. Overall, the proposed approach provides an interpretable symbolic correction framework for improving PR-EOS predictions of hydrocarbon-nitrogen VLE.
Adaptive Hybrid Particle Swarm Optimization with Gradient Descent
Gradient injection helps Particle Swarm Optimization (PSO) only when the swarm has identified a basin with smooth local structure, not univ…
見て、精査し、考える: ビデオ異常検出をトレーニング不要からエージェント推論まで進歩させる
ビデオ異常検出 (VAD) は、異常なイベントを特定し、その時間間隔を特定することを目的としています。既存のアプローチは、「いつ、何が」の解離を示します。従来の DNN ベースの手法は、異常がいつ発生するかを特定しますが、意味的な理解に欠けています。一方、LLM ベースの手法は、何が起こったのかを説明しますが、正確な時間的根拠を無視します。これは、統一された推論パラダイムが存在しないことが原因であると考えられます。人間がどのように監視ビデオを検査するか (全体的に見て一時的な仮説を立て、疑わしい部分を精査し、エラーを修正するために反復的に考える) からインスピレーションを得て、私たちはこのグローバルからローカルへのパラダイムを 2 つの観点から研究します。私たちはまず、Glance then Scrutinize (GtS) を提案します。これは、精度と速度のバランスをとりながら、粗いものから細かいものまで、異常の根拠と理解のための静的および動的なテキスト ガイダンスを使用するトレーニング不要のフレームワークです。フリーズした外部モジュールによって課せられる上限を打破するために、ツール拡張型エージェント VAD 手法をさらに提案します。この手法では、マルチモーダル大規模言語モデルが、ビデオ クロッピング ツールの呼び出し、高密度にリサンプリングされたフレームの検査、誤った位置の仮説の自己修正を学習します。これは、コールドスタートの教師あり微調整とそれに続く共同回答グラウンディング報酬による強化学習によって行われます。トレーニングと評価のために、以前の VAGU ベンチマークを VAGU-T (ビデオ異常のグラウンディング、理解、思考) に拡張しました。VAGU-T は、人間が検証したグラウンディング、説明、QA ペア、および思考連鎖ツール呼び出しトレースを備えた 21 の異常カテゴリにわたる 7,567 本の現実世界のビデオで構成されています。さらに、意味解釈可能性と時間精度を共同で評価する指標である JeAUG を紹介します。実験では、GtS がトレーニング不要のベースラインを大幅に上回り、エージェント モデルがより高い精度とより高速な推論の両方を実現することが示されています。
原文 (English)
Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning
Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-what" dissociation: traditional DNN-based methods localize when anomalies occur but lack semantic understanding, whereas LLM-based methods explain what happens but neglect precise temporal grounding. We attribute this to the absence of a unified reasoning paradigm. Inspired by how humans inspect surveillance videos - glancing globally to form temporal hypotheses, scrutinizing suspicious segments, and thinking iteratively to correct errors - we study this global-to-local paradigm from two perspectives. We first propose Glance then Scrutinize (GtS), a training-free framework using static and dynamic textual guidance for coarse-to-fine anomaly grounding and understanding, balancing accuracy and speed. To break the ceiling imposed by frozen external modules, we further propose a tool-augmented agentic VAD method, where a multimodal large language model learns to invoke a video cropping tool, inspect densely resampled frames, and self-correct mislocalized hypotheses, via cold-start supervised fine-tuning followed by reinforcement learning with a joint answer-grounding reward. For training and evaluation, we extend our prior VAGU benchmark into VAGU-T (Video Anomaly Grounding, Understanding, and Thinking), comprising 7,567 real-world videos over 21 anomaly categories with human-validated grounding, explanations, QA pairs, and chain-of-thought tool-calling traces. We further introduce JeAUG, a metric jointly evaluating semantic interpretability and temporal precision. Experiments show that GtS substantially surpasses training-free baselines, while the agentic model delivers both higher accuracy and faster inference.
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (The…
Apodex Discovery: 発見的人工知能を評価および構築するための現実ベンチマークと環境
アポロが月に到達したのは、単に技術者が難しい方程式を解くことができたからではありません。それは遠い野望を、明確な目的、シミュレーション、検証、そして修正を繰り返すミッションアーキテクチャに変えることで成功しました。 AI も現在、同様の移行に直面しています。フロンティア モデルは、問題、ツール、成功基準が指定されれば、困難なタスクを解決できますが、結果として生じる現実世界の課題が、実行可能または検証可能な形式で現れることはほとんどありません。 Apodex Discovery を紹介します。これは、ヘビーデューティ ソルバーを通じて発見的 AI を構築および評価するためのフレームワークです。このシステムは、基盤モデル、ハーネス、ツール、および拡張されたステートフルな検証可能な調査を追求する制御ポリシーで構成されるシステムです。 3 つのコアコンポーネントがあります。まず、問題調査プロセスで 16 セクターにわたる 561 業界を調査し、価値の高い現実世界の問題 423 件を収集し、最初のリリース用に 20 件を選択しました。第 2 に、共通の環境タスク エピソードの抽象化により、データ、ツール、制約、フィードバック、軌跡の記録、中間成果物と最終提出物の検証が提供されます。第三に、HDS6 は、最終タスクの成功とは関係なく、ツール、修復、代替案、一貫性、証拠、および範囲を評価します。 AAV キャプシド設計において、Apodex は、生存率、指向性、構造予測、生成デザインのすべてにおいて、公表されている最先端技術を 7% 上回りました。薬物の再利用と再製剤化では、タスク固有の生物医学環境により、GPT-5.5 および GPT-5.6-sol の平均正規化予測スコアが、同じクローズドブック バックボーンよりも 2.5 ポイントおよび 7.6 ポイント改善されました。制御されたアブレーションは、固定された TRACES エピソード インターフェイスにより、パフォーマンスの違いを特定のソルバー コンポーネントに帰属させることができることを示しています。 Apodex Discovery は、AI 評価を事前定義されたベンチマークを超えて、真の発見を目的とした検証可能な調査に移行します。
原文 (English)
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form. We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success. In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.
Frontier LLM はネイティブのマルチモーダルな埋め込みに適合できますか?ハードネガティブなテキストから画像への検索の比較
テキスト、画像、ビデオ、オーディオにわたる、さまざまな種類のメディアにわたるマルチモーダルな検索と分類は、従来、対比学習を通じて視覚表現とテキスト表現を整合させるデュアルエンコーダー モデルに依存していました。テキスト、画像、ビデオ、オーディオ、ドキュメントを単一の共有スペースにマッピングする Google 初のネイティブ マルチモーダル埋め込みモデルである Gemini Embedding 2 の 2026 年 3 月リリースにより、マルチモーダル検索システム間の競争が激化しています。同時に、フロンティアラージ言語モデル (LLM) も強力な視覚的理解を実証しており、効果的なゼロショット ランカーとして機能できるかどうかという疑問が生じています。私たちの調査では、ネイティブ マルチモーダル エンベディングと Flickr30k 上の LLM ベースのビジュアル ランキングを直接比較した初めての結果が得られました。 GPT-4.1 および Claude Sonnet 4.6 は Gemini Embedding 2 と同等のパフォーマンスを示していることがわかります。さらに、埋め込みが事前計算されると、マルチモーダル 埋め込みは低遅延アプリケーションに適しています。
原文 (English)
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval
Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Embedding 2, Google's first natively multimodal embedding model to map text, images, video, audio, and documents into a single shared space, raises competition among multimodal retrieval systems. Simultaneously, frontier Large language models (LLMs) have also demonstrated strong visual understanding, raising the question of whether they can serve as effective zero-shot rankers. Our study provides the first direct comparison of native multimodal embeddings against LLM-based visual ranking on Flickr30k. We observe that GPT-4.1 and Claude Sonnet 4.6 perform on par with Gemini Embedding 2. Additionally, once embeddings are precomputed, multimodal embeddings are better suited for low-latency applications.
コンテンツ推奨のためのマインドモデリングの逆理論: Web ブラウジングからダイナミック インテリジェント インターフェイスまで
最新のレコメンダー システムは、観察されたアクションをユーザーの好みの信頼できる代理として扱いますが、インタラクションは安定した好みの表現ではなく探索や比較を反映することがよくあります。インターフェースが静的レイアウトから生成型 UI や没入型拡張現実 (XR) に進化するにつれて、モダリティに依存しないより深いユーザー理解の必要性が高まっています。これらの適応環境では、何を表示するかだけでなく、どこで、いつ、どのように目立つように、そして最も重要なことにユーザーがなぜ行動するのかを決定する必要があります。私たちは、観察された相互作用から逆算して、行動を説明する信念、好み、意思決定特性を推測する、心の逆理論 (IToM) パイプラインを提案します。このパイプラインは、何が選択されたか、どのような選択肢が利用可能だったかなど、各ユーザーの意思決定のコンテキストを再構築し、LLM 主導の反事実推論を適用して証拠に基づく自然言語の信念ステートメントを生成し、複数の仮説によるアブダクティブ推論を通じてこれらの信念を構造化されたユーザー ペルソナに合成します。私たちは、OPeRA データセットを基に、次の行動の予測、買い物態度の調整、ビッグ 5 の性格推論、保留カテゴリーの予測という 4 つのタスクにわたって、真実の性格評価、態度調査、インタビューベースのペルソナを評価します。結果は、推定されたペルソナが真実のペルソナと一致またはそれを超えていること、および正確な性格予測には複数の仮説推論が不可欠であることを示しています。さらに、VisionOS 上の個人主導の空間バンキング アプリケーションを使用して、クロスモーダル転送可能性を実証します。
原文 (English)
Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces
Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.
From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM speci…
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs
Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plural…
社会的思考連鎖: 医療鑑別診断法に基づいたマルチエージェント アーキテクチャ
医療診断推論は、LLM の影響力の大きいユースケースであり、ユーザーの健康と福祉に重大な影響を及ぼします。 OpenAI (2026) が、世界中の ChatGPT メッセージの 5% 以上が医療関連であると報告すると、これらのシステムの透明性が設計上の重大な懸念事項になります。これは、鑑別診断に複数の形式の専門家の推論を統合する必要がある複雑なケースに特に当てはまります。既存の研究では、医療診断へのマルチエージェントアプローチが提案されていますが、そのようなシステムがいつ必要になるのか、なぜ役立つのか、そしてどのような場合にモノリシック推論よりも優れたパフォーマンスを発揮するのかは不明のままです。我々は、コラボレーションのための審議フレームワークとしてマルチエージェントインタラクションを構造化する、医療鑑別診断のためのマルチラウンドパイプラインであるソーシャルチェーンオブ思考(SCoT)を紹介します。有効な LLM 推論。単一薬剤のベースライン、単一薬剤のパイプライン アブレーション、およびベストオブ n スケーリングに対して SCoT を評価すると、そのリコールの利点がモノリシック推論だけでは再現されないことがわかります。 SCoT は、最も困難な診断ケースで最も効果を発揮します。このケースでは、複数ラウンドの専門家による会話が、真実の診断を回復し、より高い再現率の差に収束するのに役立ちます。
原文 (English)
Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology
Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI (2026) reports that more than 5% of ChatGPT messages globally are healthcare-related, the transparency of these systems becomes a serious design concern. This is especially true for complex cases, where differential diagnosis often requires integrating multiple forms of specialist reasoning. Existing work has proposed multi-agent approaches to medical diagnosis, but it remains unclear when such systems are needed, why they help, and where they outperform monolithic inference. We introduce Social Chain of Thought (SCoT),a multi-round pipeline for medical differential diagnosis that structures multi-agent interaction as a deliberative framework for collabora. tive LLM reasoning. Evaluating SCoT against single-agent baselines, one-agent pipeline ablations, and best-of-n scaling, we show that its recall advantage is not reproduced by monolithic inference alone. SCoT is most successful in the hardest diagnostic cases, where multiple rounds of specialist conversation help recover ground-truth diagnoses and converge on a higher-recall differential.
モバイル エージェント評価のための LLM 審査員のベンチマーク
モバイル エージェントのベンチマークは、タスクの完了を評価するために LLM ベースのジャッジにますます依存していますが、モバイル エージェントの軌跡に関するこれらのジャッジの信頼性はほとんど調査されていないままです。モバイル エージェントの軌跡に関する LLM-as-judge メソッドを体系的に評価するためのベンチマークである MobileJudgeBench を紹介します。私たちのベンチマークは、6 つのモバイル エージェント ベンチマーク、4 つのエージェント モデル、および 68 のアプリにわたる、人間が注釈を付けた 931 の軌跡で構成されています。複数の LLM バックエンドにわたって 6 つの判定メソッド (SPA-Bench、AndroidArena、AgentRewardBench の 2 つのモードを備えた A3 から適応した 5 つと、私たちが設計した単純なベースライン) を評価します。私たちの実験により、3 つの重要な発見が明らかになりました。まず、サンプルのスクリーンショットを使用した単純なベースライン審査員は、専用の手法と競合し、多くの場合それを上回っています。これは、より複雑な審査員パイプラインが審査員の質を一貫して向上させるわけではないことを示しています。競合する手法の中でも、LLM バックボーンが主な推進力です。第 2 に、ベンチマーク品質メトリクスは、現実世界のジャッジの有用性を確実に予測します。ベンチマーク品質メトリクスは、評価におけるエージェントのランキング忠実度と、ジャッジがポリシーに基づく強化学習の報酬シグナルとして機能する場合の下流のパフォーマンスの両方に相関します。 3 番目に、2 つの LLM バックエンドにわたる障害分析により、バックボーンの適合率と再現率の特性に関連する、一方は保守的で他方は許容的である、定性的に反対の障害プロファイルが明らかになります。
原文 (English)
Benchmarking LLM Judges for Mobile Agent Evaluation
Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamined. We introduce MobileJudgeBench, a benchmark for systematically evaluating LLM-as-judge methods on mobile agent trajectories. Our benchmark comprises 931 human-annotated trajectories spanning 6 mobile agent benchmarks, 4 agent models, and 68 apps. We evaluate 6 judge methods (five adapted from SPA-Bench, A3 with two modes, AndroidArena, and AgentRewardBench, plus a simple baseline we design) across multiple LLM backends. Our experiments reveal three key findings. First, a simple baseline judge with sampled screenshots is competitive with, and often exceeds, purpose-built methods, indicating that more elaborate judge pipelines do not consistently improve judge quality; among competitive methods, the LLM backbone is the primary driver. Second, benchmark quality metrics reliably predict real-world judge utility: they correlate with both agent ranking fidelity for evaluation and downstream performance when judges serve as reward signals for on-policy reinforcement learning. Third, failure analysis across two LLM backends uncovers qualitatively opposite failure profiles, one conservative and the other permissive, linked to the backbone's precision-recall characteristics.
合成的に制約された複数目的のヒットからリードまでの最適化のためのモジュール式エージェント フレームワーク
ヒットからリードへの最適化には、競合する効力、選択性、物理化学的、薬物動態学、安全性、合成上の制約を越えてヒット類似体の反復設計が必要です。我々は、自然言語オーケストレーションを採用して化学構造の最適化を導くオープンソース フレームワークである SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration) を紹介します。 SABLE は、LLM を使用してユーザー定義の目標を解釈し、タスクをルーティングします。一方、特殊なツールは、反応テンプレート化されたアナログ列挙、物理化学的および ADMET 特性予測、構造ベースの親和性スコアリング、およびベイジアン最適化を実行します。結果として得られるワークフローは、設計、製造、テスト、分析サイクルの分析段階と優先順位付け段階の計算ツインであり、各数値出力の出所を提供します。 SABLE は、単一および複数の目的の最適化研究にわたって、列挙された検索空間のサブセットのみを評価しながら、ユーザー定義の計算目的の候補セットを強化します。モジュール式アーキテクチャにより、操作ロジックを変更せずに、単純な構成ファイルを編集するだけでツールと特性評価バックエンドを置き換えることができます。 SABLE は、初期段階の創薬において合成的に制約された類似体を優先するための拡張可能な意思決定支援フレームワークを提供します。
原文 (English)
A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization
Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. We present SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration), an open-source framework that employs natural-language orchestration to guide chemical structure optimization. SABLE uses an LLM to interpret user-defined goals and route tasks, while specialized tools perform reaction-templated analog enumeration, physicochemical and ADMET property prediction, structure-based affinity scoring, and Bayesian optimization. The resulting workflow is a computational twin of the analytical and prioritization stages of the design-make-test-analyze cycle, providing provenance of each numerical output. Across single, and multi-objective optimization studies, SABLE enriches candidate sets for user-defined computational objectives while evaluating only a subset of the enumerated search space. Its modular architecture allows tools and characterization backends to be replaced by editing a simple config file, without modifying operational logic. SABLE provides an extensible decision-support framework for prioritizing synthetically constrained analogs in early-stage drug discovery.
プロンプトから行動の調整まで: 個別の LLM ジャッジによる推奨評価
従来のオフライン レコメンデーション評価は、拡張が困難な手動で維持される複雑な機能パイプラインに大きく依存しています。大規模言語モデル (LLM) は、生のテキスト ログからユーザー エンゲージメントを直接予測することで有望な代替手段を提供しますが、この研究の実証分析では、双方向合理化と呼ばれる重大な失敗モードが特定されています。ゼロショット設定では、LLM はまったく同じ項目について、同一の証拠を用いてユーザー エンゲージメントの肯定的結果と否定的結果の両方を説得力を持って主張することがわかり、ユーザー エンゲージメントの予測における既製 LLM の信頼性の低さを浮き彫りにしています。これを解決するために、正しい理論的根拠と反事実的な理論的根拠を組み合わせて、選好の最適化と微調整を組み合わせた逐次行動調整フレームワークを開発し、適用します。実際のホームページのインタラクション ログで評価すると、この調整された推論アプローチは、ゼロショット ベースラインを上回る Macro-F1 スコアの 32.19\% 上昇を達成し、実稼働環境で設計されたベースラインと一致します。この結果は、動作の調整により双方向の合理化が軽減され、手動のパイプライン オーバーヘッドなしで人間が解釈可能な推論トレースを提供できることが実証されました。
原文 (English)
From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation
Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact same item with identical evidence, highlighting the unreliability of off-the-shelf LLMs in predicting user engagement. To resolve this, we develop and apply a sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales. Evaluated on real-world homepage interaction logs, this aligned reasoning approach achieves a 32.19\% lift in Macro-F1 score over the zero-shot baseline and matches the production feature-engineered baseline. The results demonstrate that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pipeline overhead.
安全性の調整のローカライズ: MLP レイヤーとミッドネットワーク ブロックが大規模言語モデルの拒否動作をエンコードする
大規模な言語モデルにおける安全性の調整は、多くの場合、ネットワーク全体の分散特性として扱われますが、その実際的な脆弱性は、拒否動作がより小さなパラメータのセットに集中している可能性があることを示唆しています。この研究では、複数レベルの粒度で、整列されたモデルから一致する整列されていない基本モデルに重みを移植することにより、安全に整列された拒否がエンコードされる場所に対処します。 2 つのオープンウェイト モデル ペアと 4 つの安全性ベンチマークを使用して、アテンション ウェイト、MLP ウェイト、連続層領域、MLP ブロックの置き換えの効果を比較する実験を実施しました。どちらのモデル ファミリでも、拒否の転送は MLP の重みによって支配されています。MLP パラメーターを置き換えると、アテンション パラメーターを置き換えるよりも大幅に悪意のあるプロンプトの拒否が回復し、ベンチマーク全体で少なくとも 2.7 倍の向上が得られます。 MLP スタック内では、モデルとデータセットのペアに対する 6 つの貪欲な検索すべてでレイヤー 8 ~ 11 にまたがるブロックが最初に選択されるため、拒否関連パラメーターは一貫したネットワーク中央集中を示します。この結果は、安全関連コンポーネントの構成が非相加的であることも示しています。つまり、6 つの貪欲な軌道のうち 5 つでは、より整列したブロックを追加すると拒否パフォーマンスが低下する可能性があり、選択的ブロックのサブセットは、悪意のある拒否、良性の過剰拒否、またはその両方に対して完全な MLP 移植よりも優れたパフォーマンスを発揮できる可能性があります。最後に、OR-Bench に転送される貪欲なオーダーは、その導出に使用されたソース ベンチマークによって異なり、ベンチマークに依存する精度とカバレッジのトレードオフを示しています。これらの結果は、現在の LLM における安全アライメントが局所的かつ相互作用に敏感であることを示唆しており、アライメントの脆弱性と、標的を絞った安全介入のための潜在的な手段についての洞察を提供します。
原文 (English)
Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models
Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal is encoded by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity. Using two open-weight model pairs and four safety benchmarks, we conducted experiments to compare the effects of replacing attention weights, MLP weights, contiguous layer regions, and MLP blocks. Across both model families, refusal transfer is dominated by MLP weights: replacing MLP parameters recovers substantially more malicious-prompt refusal than replacing attention parameters, with gains of at least 2.7 times more across benchmarks. Within the MLP stack, refusal-relevant parameters exhibit a consistent mid-network concentration, as the block spanning layers 8-11 is selected first in all six greedy searches over model-dataset pairs. The results also show that the composition of safety-relevant components is non-additive: in five of six greedy trajectories, adding more aligned blocks can reduce refusal performance, and selective block subsets can outperform full MLP transplantation on malicious refusal, benign over-refusal, or both. Finally, greedy orders transferred to OR-Bench vary with the source benchmark used to derive them, indicating a benchmark-dependent precision-coverage trade-off. These results suggest that safety alignment in current LLMs is both localized and interaction-sensitive, offering insight into alignment brittleness and potential avenues for targeted safety interventions.
EnterpriseRAG: 非理想的なエンタープライズ検索下での LLM 命令の準拠性と堅牢性のベンチマーク
エンタープライズ RAG 導入は、重大な信頼性ギャップに直面しています。LLM は個別の制約の 80% を満たしますが、すべての要件を同時に満たす応答は 26.8% のみであり、57 ポイントのオーケストレーション ギャップが明らかになります。既存のベンチマークは、単純なクエリによるクリーンな取得を前提としており、ノイズの多いドキュメントと多次元の制約が共存する運用環境を把握できません。 EnterpriseRAG は、6 つのドメインにわたって専門家によって検証された 983 個のサンプルのベンチマークであり、これまでの作業にはなかった 3 つの障害モード (検索ノイズ、知識のギャップ、事実の矛盾) を体系的にシミュレートし、複雑な命令を組み合わせます。 13 の最先端の LLM を評価すると、制約ごとの満足度が高いと全体的なコンプライアンスの低さが隠蔽され、指示遵守の深刻な崩壊が明らかになります。重大な調査結果は、推論を強化した推論であっても、知識のギャップや事実の矛盾の下にある深い障壁を明らかにしており、運用 RAG には明示的なコンテキスト認識プロトコルと調整された判断が必要であることを示しています。 EnterpriseRAG は、これらのギャップを測定して埋めるための再現可能な基盤を提供し、エンタープライズ規模の RAG システムの展開に関する決定に直接情報を提供します。ベンチマークと評価フレームワークは公開後に公開します。
原文 (English)
EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval
Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions where noisy documents and multi-dimensional constraints coexist. We introduce EnterpriseRAG, a benchmark of 983 expert-validated samples across six domains that systematically simulates three failure modes absent from prior work: retrieval noise, knowledge gaps, and factual conflicts, coupled with complex instructions. Evaluation of 13 state-of-the-art LLMs reveals a severe instruction adherence collapse, where high per-constraint satisfaction masks low holistic compliance. Critical findings expose deep barriers under knowledge gaps and factual conflicts, even with reasoning-enhanced inference, indicating production RAG requires explicit context-aware protocols and calibrated judgment. EnterpriseRAG provides a reproducible foundation for measuring and closing these gaps, directly informing deployment decisions for enterprise-scale RAG systems. We will release the benchmark and evaluation framework upon publication.
CoAdapt-GUI: 目に見えない GUI アプリケーションに対する共同ワークフロー コンテキストとポリシーの適応
モバイル GUI エージェントは、ソース トレーニングが行われていないアプリケーションにデプロイされると脆弱なままになります。私たちは、限られたターゲットインタラクション予算の下で、ターゲットのデモンストレーションを行わずに、新規アプリの一般化を研究します。 CoAdapt-GUI は、エージェント自身のターゲット アプリのロールアウトと報酬から構造化されたワークフロー コンテキストとポリシーを共同で適応させるテスト時間適応 (TTA) フレームワークです。ワークフロー コンテキストは、アプリにバインドされたソースの詳細を除外しながら、転送可能なプロシージャ、障害モード、および検証ルールを保持します。この分離により、ソース インターフェイスの状態を転送することなく、再利用可能なワークフローの知識を適応に導くことができます。ポリシー適応のために、タスクコンテキストに一致するグループ相対最適化により、凍結されたビジョン言語モデルに基づいて LoRA アダプターが更新されます。 2 つの未確認アプリ評価全体で、CoAdapt-GUI は AndroidWorld-Generalization で 45.0% に達し、報告された Policy-Only TTA ベースラインでは 37.5% に達し、AndroidWorld Plus のパフォーマンスは 38.6% から 52.9% に向上しました。これらの結果は、転送に制約のあるワークフロー コンテキストが大幅な利益をもたらし、共同ポリシーの適応により持続的なパフォーマンスがさらに向上することを示しています。
原文 (English)
CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications
Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target interaction budget and without target demonstrations. We introduce CoAdapt-GUI, a test-time adaptation (TTA) framework that jointly adapts structured workflow context and policy from the agent's own target-app rollouts and rewards. The workflow context retains transferable procedures, failure modes, and verification rules while excluding app-bound source details. This separation allows reusable workflow knowledge to guide adaptation without transferring source-interface state. For policy adaptation, task-context-matched group-relative optimization updates a LoRA adapter on a frozen vision-language model. Across two unseen-app evaluations, CoAdapt-GUI reaches 45.0% on AndroidWorld-Generalization, compared with 37.5% for the reported Policy-Only TTA baseline, and raises AndroidWorld Plus performance from 38.6% to 52.9%. These results show that transfer-constrained workflow context provides substantial gains and that joint policy adaptation further improves held-out performance.
ショッピング エージェントがオンライン ユーザーのフィードバックから学ぶ
大規模な言語モデルベースのショッピング エージェントが現実世界の電子商取引プラットフォームに導入されることが増えており、これらのエージェントを改善するための貴重な監視を提供する大量のユーザー インタラクション ログが生成されます。しかし、既存のアプローチは主に、ユーザーとアイテムのインタラクションや合成嗜好データなどのオフライン トレーニング信号に依存しており、ユーザーの自然な会話によるフィードバックに含まれる豊富な監視をほとんど無視しています。さらに、利用可能なオンライン フィードバックは不均一で、まばらで、ノイズが多いため、信頼性の高い学習信号に自動的に変換することが困難です。これらの課題に対処するために、私たちは、ショッピング エージェントが人間による注釈なしで実際のオンライン インタラクション ログから直接学習できるフレームワークである LOFA を提案します。 LOFA は、検証可能な購入結果に対する強化学習と、ユーザーの対話中の指示を識別し、それを高密度のトークンレベルの監視に変換する、フィードバックを意識したポリシーの抽出を組み合わせます。これらの補完的な目標は、協調的な行動パターンとユーザー固有の好みの両方を捕捉します。現実の電子商取引ログに関する広範な実験により、LOFA が推奨の品質、応答の有用性、およびユーザー満足度の調整を強力なベースラインよりも一貫して向上させることが実証され、実際のオンライン ユーザーのフィードバックからショッピング エージェントを学習することの有効性が強調されています。
原文 (English)
Learning from Online User Feedback for Shopping Agents
Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, existing approaches primarily rely on offline training signals, such as user-item interactions or synthetic preference data, while largely overlooking the rich supervision contained in users' natural conversational feedback. Moreover, the available online feedback is heterogeneous, sparse, and noisy, making it difficult to transform into reliable learning signals automatically. To address these challenges, we propose LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation. LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users'in-dialogue directives and converts them into dense token-level supervision. These complementary objectives capture both collaborative behavioral patterns and user-specific preferences. Extensive experiments on real-world e-commerce logs demonstrate that LOFA consistently improves recommendation quality, response helpfulness, and user-satisfaction alignment over strong baselines, highlighting the effectiveness of learning shopping agents from real online user feedback.
見ざる先見性: 世界の行動モデルの潜在的な未来
ワールド アクション モデル (WAM) は、将来の視覚的予測とロボットのアクション生成を組み合わせ、相互作用中に物理世界がどのように進化するかをポリシーでモデル化できるようにします。既存の WAM は、予測ダイナミクスをアクション経路にどのように公開するかが異なります。明示的な将来の WAM は、予測されたシーンの展開への直接アクセスを提供しますが、反復的なビデオのノイズ除去によりかなりの推論コストが発生します。対照的に、直接ポリシー WAM は、現在の観察からアクションを効率的に予測しますが、予測ダイナミクスをアクション DiT に公開するための明示的な推論時間インターフェイスがありません。このギャップを埋めるために、将来のビデオをデコードせずにアクション生成のための予測コンテキストを提供する、ダイナミクス条件付きの直接ポリシー WAM である ForeWAM を提案します。 Future-KV の核心では、現在の視覚的な潜在スロットと確率的な将来スロットに対して単一のビデオ DiT プレフィルを実行し、結果として得られるレイヤーごとのキーと値の状態をアクションのノイズ除去全体で再利用します。さらに、凍結された潜在アクション教師によって監督されたダイナミクス レジスタを導入し、オブジェクトの動き、接触の変化、タスクの進行などの相互作用によって引き起こされる遷移を捕捉するように暗黙的な将来状態を促進します。グラウンドトゥルースの将来観察と教師はトレーニング中にのみ使用されます。導入ではどちらも必要なく、将来のビデオ生成も実行されません。組み込まれたロボット データの事前トレーニングを行わない場合、ForeWAM の標準バージョンと高速バージョンは、LIBERO 上でそれぞれ 96.7% と 96.9% の平均成功率を達成しました。標準バージョンはさらに、LIBERO-Plus で 61.6% の成功率を達成しています。これらの結果は、直接ポリシー WAM が、将来の観測を明示的に生成することなく、アクション経路に予測ダイナミクスを公開しながら、効率的なアクション予測を保持できることを示しています。
原文 (English)
Foresight Without Seeing: Latent Futures for World Action Models
World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction. Existing WAMs differ in how predictive dynamics are exposed to the action pathway. Explicit-future WAMs provide direct access to predicted scene evolution, but incur substantial inference costs from iterative video denoising. In contrast, direct-policy WAMs efficiently predict actions from the current observation but lack an explicit inference-time interface for exposing predictive dynamics to the Action DiT. To bridge this gap, we propose ForeWAM, a dynamics-conditioned direct-policy WAM that provides predictive context for action generation without decoding future videos. At its core, Future-KV performs a single Video DiT prefill over the current visual latent and stochastic future slots, and reuses the resulting layer-wise key-value states throughout action denoising. We further introduce dynamics registers supervised by a frozen latent action teacher, encouraging the implicit future states to capture interaction-induced transitions such as object motion, contact changes, and task progress. Ground-truth future observations and the teacher are used only during training; deployment requires neither and performs no future video generation. Without embodied robot data pretraining, the standard and accelerated variants of ForeWAM achieve average success rates of 96.7% and 96.9% on LIBERO, respectively. The standard variant further achieves 61.6% success on LIBERO-Plus. These results demonstrate that direct-policy WAMs can retain efficient action prediction while exposing predictive dynamics to the action pathway without explicitly generating future observations.
MBA: 現実世界のビジネスアイデアのためのマルチモーダルベンチマークとエージェント
大規模言語モデル (LLM) を活用したエージェント システムは、ビジネスのアイデア創出に新たな機会をもたらしました。しかし、実世界のコンテキストには本質的にマルチモーダルな性質があるにもかかわらず、既存のアプローチは依然としてテキストのみのパラダイムに限定されています。そこで、ビジネスアイデアエージェントのトレーニングと評価のための初のマルチモーダルベンチマークである MBA-Bench を紹介します。これは 6 つのドメインにわたる 30,000 のサンプルで構成され、各ドメインはテキストだけでは完全には伝えられない明確な視覚的手がかりによって特徴付けられます。具体的には、画像に自動的にキャプションを付け、GPT-4o を使用して、検索クエリの生成、市場証拠の検索、および証拠強化合成を通じて 3 つのビジネス質問のそれぞれに対して 5 つの参考アイデアを生成します。以前の作業に続き、MLLM-as-a-Judge を使用して 6 つのビジネス指向の基準にわたってエージェントを評価します。基準が非表示または開示される設定を検討するために、ブラインドおよび既知のそれぞれに対して MBA-b および MBA-k を提示します。私たちは両方を 2 つの新しい報酬目標 (創造性と実現可能性) に基づいてトレーニングしますが、MBA-k は公開されている 6 つの基準を合計 8 つさらに最適化します。どちらも、LoRA ベースの監視付き微調整を介してトレーニングされ、その後、これらの設定固有の報酬を使用してグループ相対ポリシーの最適化が行われます。 MBA ベンチでの大規模な実験のために、キャプションのみまたはマルチモーダル入力のいずれかに対応する 2 つのベースラインを設定しました。後者はいくつかの指標でクローズド ソースのパフォーマンスに近づきます。 MBA-b と MBA-k は、キャプション ベースラインをそれぞれ 63.9% と 77.1% 上回り、マルチモーダル ベースラインを 25.6% と 35.8% 上回っています。
原文 (English)
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.
AI が生成するフィードバックを重要なものにする: 提供から生徒の制定まで
フィードバック プロセスは生徒の学習に大きな影響を与えますが、その教育的価値は 2 つの異なる課題に対処することにかかっています。それは、高品質でタイムリーかつ個別化されたフィードバックを大規模に提供することと、生徒がそのフィードバックを解釈、評価し、生産的に行動できるようにサポートすることです。ジェネレーティブ AI は、プロビジョニングの課題に対処する信頼できる手段を提供しますが、AI によって生成されたフィードバックを学生が取り込むには依然として限界があります。私たちは、13,037 人の学生と 51,296 人の学生が作成したリソースを対象に、AI を介した 3 つのフィードバック ワークフローを比較する大規模な準実験的な逐次コホート研究を実施しました。指示されたフィードバック (n = 3,723) では、学生は構造化されたサポートなしで AI によって生成されたフィードバック コメントを受け取りました。自主的なフィードバック (n = 3,951) では、学生はオプションで AI 支援の対話を開始できます。制定されたフィードバック (n = 5,363) では、学生はフィードバックの提案を選択し、その関連性を評価し、それらの選択に基づいてターゲットを絞った AI 支援の対話に参加するよう求められました。制定されたフィードバックは、AI が生成したフィードバックの摂取率が大幅に高く、推定確率は 26.2% でした。これに対し、指示されたフィードバックの 14.1%、自主的なフィードバックの 0.1% と比較します。また、両方の比較条件よりも大幅に高い自己評価の信頼性と提出された作品の品質にも関連していました。これらの調査結果は、AI によって生成されたフィードバックの教育的価値は、フィードバック コメントの質だけでなく、生徒によるフィードバック リテラシー プロセスの実践を積極的に構築するワークフローにも依存することを示唆しています。この結果は、学習者をコメントの受動的受信者ではなく、判断、対話、改善への積極的な参加者として位置づける AI フィードバック システムの設計に影響を及ぼします。全体的な調査結果は、AI へのアクセスだけでは不十分であることを示しています。目的を持ったワークフロー設計は、生産的なフィードバックの使用の中心となります。
原文 (English)
Making AI-Generated Feedback Matter: From Provision to Student Enactment
Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and individualised feedback at scale, and supporting students to interpret, evaluate, and act on that feedback productively. Generative AI offers a credible means of addressing the provision challenge, but students' uptake of AI-generated feedback remains limited. We conducted a large-scale quasi-experimental sequential cohort study comparing three AI-mediated feedback workflows across 13,037 students and 51,296 student-authored resources. In Directed Feedback (n = 3,723), students received AI-generated feedback comments without structured support. In Self-Directed Feedback (n = 3,951), students could initiate optional AI-supported dialogue. In Enacted Feedback (n = 5,363), students were prompted to select feedback suggestions, evaluate their relevance, and engage in targeted AI-supported dialogue anchored to those selections. Enacted Feedback was associated with significantly higher uptake of AI-generated feedback, with an estimated probability of 26.2%, compared with 14.1% for Directed Feedback and 0.1% for Self-Directed Feedback. It was also associated with significantly higher self-assessment confidence and submitted-work quality than both comparison conditions. These findings suggest that the educational value of AI-generated feedback depends not only on the quality of feedback comments, but also on workflows that actively structure students' enactment of feedback literacy processes. The results have implications for the design of AI feedback systems that position learners as active participants in judgement, dialogue, and improvement rather than passive recipients of comments. Overall findings show that AI access alone is insufficient; purposeful workflow design is central to productive feedback use.
主張: 不確実性測定による大規模言語モデルのオープンドメインでの最先端のアクティブな解明
オープンドメインの人間とコンピューターの対話シナリオでは、大規模言語モデル (LLM) は、曖昧または不完全なユーザー クエリに頻繁に遭遇します。このような場合、直接回答を作成すると、一般化されすぎたり、誤ったり、情報が不足したりする回答が得られることがよくあります。対照的に、明確な質問をすると、対話の質が大幅に向上します。ただし、既存のアプローチは、明確化が必要な場合とクエリのどの側面を明確化する必要があるかという 2 つの基本的な課題に対処するために、依然として手動で注釈を付けたデータまたは優先順位の調整に大きく依存しています。この依存により、高いアノテーション コストが発生し、一般化が制限されます。これらの課題に対処するために、私たちは、オープンドメイン設定での能動的解明学習のための不確実性主導フレームワークである CLAIM を提案します。 CLAIM は、複数のモデルにわたる回答の不一致によって引き起こされるエントロピーを通じてクエリの不確実性を定量化することにより、明示的な人間の好みの注釈の必要性を排除します。この不確実性信号を使用して高品質の合成データが構築され、教師あり学習と強化学習の組み合わせによる統一明確化決定モデルのトレーニングが可能になります。具体的には、エントロピーベースの不確実性推定とセマンティッククラスタリングおよび推論ベースの判断を統合し、明確化要件の信頼できる自動アノテーションを可能にするエントロピー駆動の合成データ生成パイプラインを提案します。 CLAIM をトレーニングするために、明確化プロセスを構造化された意思決定生成問題として定式化し、教師あり微調整 (SFT) とグループ相対ポリシー最適化 (GRPO) を組み合わせたトレーニング パラダイムを採用します。実験結果は、CLAIM が手動でラベル付けされたデータに依存せずに、安定した一般化可能な明確化戦略を学習できることを示しており、LLM との現実世界のオープンドメインインタラクションを事前に理解するための低コストで堅牢なソリューションを提供します。
原文 (English)
CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement
In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-information responses. In contrast, asking clarifying questions can substantially improve interaction quality. However, existing approaches still rely heavily on manually annotated data or preference alignment to address two fundamental challenges: when clarification is necessary, and which aspect of the query should be clarified. This reliance incurs high annotation costs and limits generalization. To address these challenges, we propose CLAIM, an uncertainty-driven framework for active clarification learning in open-domain settings. CLAIM eliminates the need for explicit human preference annotations by quantifying query uncertainty through the entropy induced by answer disagreements across multiple models. This uncertainty signal is then used to construct high-quality synthetic data, enabling the training of a unified clarification decision model through a combination of supervised learning and reinforcement learning. Specifically, we propose an entropy-driven synthetic data generation pipeline that integrates entropy-based uncertainty estimation with semantic clustering and reasoning-based judgments, enabling reliable automatic annotation of clarification requirements. To train CLAIM, we formulate the clarification process as a structured decision generation problem and adopt a training paradigm that combines supervised fine-tuning (SFT) with group-relative policy optimization (GRPO). Experimental results demonstrate that CLAIM can learn stable and generalizable clarification strategies without relying on manually labeled data, offering a low-cost and robust solution for proactive understanding in real-world open-domain interactions with LLMs.
XBridge: 異種 LLM 通信用のエンティティ接地型潜在ブリッジ
エージェントが異なるモデル ファミリを利用している異種マルチエージェント LLM システムは、冗長な推論パターンを減らすことで同種構成よりも優れたパフォーマンスを発揮できます。しかし、既存の通信プロトコルは、送信者の内部表現を破棄してテキストを通じて動作するか、潜在レベルの転送のためにアーキテクチャの均一性を必要とします。クロスアーキテクチャ通信におけるエンティティグラウンディングの問題を特定します。異なる LLM ファミリ間で連続表現を転送するクロスアテンション ブリッジは、まれなトークン圧縮崩壊に悩まされ、連続的なボトルネック (ブリッジのみの F1 ~30%) でエンティティのアイデンティティが失われます。私たちは、2 つのメカニズムを通じてこれに対処する、デコード不要の通信プロトコルである XBRIDGE を提案します。語彙アンカー マッピング (LAM) は、送信者の元のコンテキスト トークンを受信者の語彙にマッピングし、個別のエンティティ アンカーを提供します。 Latent Enrichment Bridge (LEB) を使用すると、受信者はコンテキスト エンリッチメントのために送信者の隠れた状態をクエリできます。エンティティ アンカーは、受信者自身の自己注意を通じて、ブリッジのコンテキスト信号を特定のエンティティに接地します。 3 つのモデル ファミリ (Llama、Qwen、Mistral)、7 つのベンチマーク、および両方の通信方向にわたって、XBRIDGE は各モデル ペアの 7 つのタスクすべてでテキストベースの通信を上回り、11 分の 1 の低いレイテンシーを達成します。また、同じアーキテクチャ設定では、7 つのタスクのうち 6 つで KV 共有ベースラインを超えています。 LEB に必要なトレーニング可能なパラメータは 2 億 6,400 万個 (受信機の 3.8%) のみで、バランスの取れた小さなサンプルセットでトレーニングされ、追加される推論オーバーヘッドはごくわずかです。
原文 (English)
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's internal representations, or require architectural homogeneity for latent-level transfer. We identify the entity grounding problem in cross-architecture communication: cross-attention bridges that transfer continuous representations across different LLM families suffer from rare-token compression collapse, where entity identity is lost in the continuous bottleneck (bridge-only F1 ~30%). We propose XBRIDGE, a decode-free communication protocol that addresses this through two mechanisms. Lexical Anchor Mapping (LAM) maps the sender's original context tokens to the receiver's vocabulary, providing discrete entity anchors. A Latent Enrichment Bridge (LEB) lets the receiver query the sender's hidden states for contextual enrichment. The entity anchors ground the bridge's contextual signals to specific entities through the receiver's own self-attention. Across three model families (Llama, Qwen, and Mistral), seven benchmarks, and both communication directions, XBRIDGE outperforms text-based communication on all seven tasks for each model pair while achieving 11x lower latency, and in a same-architecture setting it also exceeds a KV-sharing baseline on six of seven tasks. LEB requires only 264M trainable parameters (3.8% of the receiver), is trained on a small balanced sample set, and adds negligible inference overhead.
AgenticTwin: 異常検出のためにデジタル ツインと統合された Agentic LLM フレームワーク
デジタルツインは、サイバーフィジカルシステムの動作を監視およびシミュレートするためにますます使用されています。熟練したオペレーターであっても、生のセンサー データの複雑さと量によって徹底的な分析が困難になるため、デジタル ツイン パイプライン内で検出された異常を解釈するのは困難です。大規模言語モデル (LLM) の最近の進歩により、推論と説明のための有望な機能が提供されていますが、デジタル ツイン駆動の異常分析への統合はまだ十分に検討されていません。この研究では、LLM 駆動の推論とデジタル ツイン ベースの異常検出パイプラインを統合するエージェント フレームワークである AgenticTwin を提案します。このフレームワークは、デジタル ツイン駆動の異常分類器からの出力で LLM によって生成された説明を根拠にし、人間のオペレーターがシステムに関する関連する自然言語の質問をできるようにします。フレームワーク自体を超えて、現実世界の気象センサー データセットに挿入された合成異常を基に構築されたベンチマーク指向の評価パイプラインを導入し、異常イベントに対するオペレーター クエリの制御された生成を可能にします。さらに、実用的なサイバー物理環境に軽量のオープンソース LLM を導入する実現可能性を評価します。実験結果は、構造化されたエージェントのコラボレーションと知識に基づいた推論により、起こり得るさまざまな異常シナリオにわたって診断の品質、状況に応じた検索、および軽減の品質を向上させることを示しています。
原文 (English)
AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection
Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies detected within digital twin pipelines is challenging, as the sheer complexity and volume of raw sensor data make thorough analysis difficult. Recent advances in large language models (LLMs) offer promising capabilities for reasoning and explanation, yet their integration into digital twin-driven anomaly analysis remains underexplored. In this work, we propose AgenticTwin, an agentic framework that integrates LLM-driven reasoning with a digital twin-based anomaly detection pipeline. The framework grounds LLM-generated explanations in outputs from a digital twin-driven anomaly classifier and enables human operators to ask relevant natural-language questions about the system. Beyond the framework itself, we introduce a benchmark-oriented evaluation pipeline constructed over synthetic anomalies injected into a real-world weather sensor dataset, enabling controlled generation of operator queries over anomaly events. We further evaluate the feasibility of deploying lightweight, open-source LLMs for practical cyber-physical environments. Experimental results demonstrate that structured agent collaboration and knowledge-grounded reasoning improve diagnosis quality, contextual retrieval, and mitigation quality across diverse possible anomaly scenarios.
FrontierFinance: 金融代理店のフロンティア インテリジェンスを測定するための挑戦的なベンチマーク
AI エージェントは専門的な投資調査に導入されることが増えていますが、投資家のワークフロー全体の複雑さを捉えたベンチマークはありません。既存のベンチマークは主に財務データの抽出を対象としていますが、これは現在のモデルがほぼ飽和状態となっている狭い範囲であり、一方、参照ベースの指標や一般的な LLM-as-a-judge スコアリングでは、実際のアナリストの質問が求める自由回答型の長い形式の回答には至っていません。 FrontierFinance は、投資家のワークフロー全体にわたる 6 つの重要なユースケースにまたがる、専門家が作成した 220 のクエリと 11,543 のソース属性のルーブリックからなる完全にオープンなベンチマークです。 FrontierFinance は、既存の公的金融ベンチマークよりも広範囲かつ困難です。公開データに限定された共通ハーネスの下でフロンティア モデルとエージェント システムを評価すると、モデル単体ではなくツール ハーネスが品質と効率を大きく左右していることがわかります。 Samaya の社内システムは 56.0% でリードしており、約 2.2 倍のコストで最も強力なフロンティア モデル (Claude Fable 5、49.2%) を上回っています。そして、最高のオープンウェイト モデル (Kimi K3、46.4%) は、4.5 倍のコストで最高の独自モデルとほぼ同等であることがわかりました。スクリーニングと検出とセクター、産業、マクロは依然としてすべてのシステム全体で最も困難なユースケースであり、最高のシステムでも 33% と 39% にすぎません。データセットとグレーディング コードは公開されています。
原文 (English)
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents
AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand. We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow. FrontierFinance is both broader and harder than existing public finance benchmarks. Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at 56.0%, ahead of the strongest frontier model (Claude Fable 5, 49.2%) at roughly 2.2x lower cost; and that the best open-weight model (Kimi K3, 46.4%) nearly matches the best proprietary model at 4.5x lower cost. Screening & Discovery and Sector, Industry & Macro remain the hardest use cases across all systems, where even the best systems reach only 33% and 39%. We make the dataset and grading code publicly available.
HUGIN: 自動物流仕分けのための視覚言語計画の強化
自律型物流仕分けシステム (ALSS) は、身体化型 AI の重要な産業用途であり、空間的にバラバラなカメラ ビュー上での共同計画が必要です。この設定を共同マルチシーン理解 (JMSU) として定式化します。オープンワールドの視覚的理解とタスク計画機能を備えたビジョン言語モデル (VLM) は、JMSU の有望な候補です。ただし、既存の VLM を JMSU に直接適用することは、シーン間の監視が不足していることと、JMSU では長い視覚的コンテキストによって引き起こされる注意の分散のため、簡単ではありません。これらの課題に対処するために、私たちは 2 つの補完的なコンポーネントを備えたトレーニング フレームワークである HUGIN を提案します。 Endogenous Data Augmentation は、動作上の制約の下で検証されたアトミックな事実を再結合しますが、Global Context Ranking は命令表現を部分的な視覚的コンテキストよりも完全な視覚的コンテキストとより強力に整合させます。進行中の研究をサポートするために、私たちは自律型物流仕分けシステムの 4 つのレイアウトから SortingBench という高品質の産業仕分けデータセットとベンチマークを構築しています。 5 つのオープン VLM にわたって、HUGIN は一貫して一致するベースラインを上回っています。たとえば、Qwen3-VL-8B の SortingBench の精度は 63.6% から 78.8% に増加します。追加の実験では、各コンポーネントの有効性と、具体化されたタスクにおける JMSU の波及効果が検証されます。 15,000 個を超える荷物を含む展開テストにより、自律的な物流仕分けのための VLM ベースの計画の実用性がサポートされます。
原文 (English)
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We formulate this setting as Joint Multi-Scene Understanding (JMSU). With open-world visual understanding and task-planning capabilities, vision-language models (VLMs) are promising candidates for JMSU. However, directly applying existing VLMs to JMSU is non-trivial due to scarce cross-scene supervision and attention dispersion caused by long visual context in JMSU. To address these challenges, we propose HUGIN, a training framework with two complementary components. Endogenous Data Augmentation recombines verified atomic facts under operating constraints, while Global Context Ranking aligns the instruction representation more strongly with the complete visual context than with a partial visual context. To support ongoing research, we construct a high-quality industrial sorting dataset and benchmark named SortingBench from four layouts of autonomous logistics sorting systems. Across five open VLMs, HUGIN consistently outperforms matched baselines; for example, the accuracy on SortingBench of Qwen3-VL-8B increases from 63.6% to 78.8%. Additional experiments verify the effectiveness of each component and JMSU's spillover benefits in embodied tasks. Deployment tests involving more than 15,000 packages support the practical viability of VLM-based planning for autonomous logistics sorting.
LLM をより客観的にする: 特性不変の安全性チューニングによる特性全体にわたる LLM の安全動作の安定化
調整された大規模言語モデル (LLM) は、ユーザーのリクエストの内容に基づいて安全な動作を示すことが期待されます。つまり、安全でないリクエストを拒否し、安全なリクエストに従う必要があります。しかし、システムプロンプトで割り当てられた異なる特性の下では、同じリクエストが実質的に異なる安全性の決定を引き起こす可能性があることを示します。これは特性誘発安全性変動と呼ばれる故障モードです。この失敗を測定するために、拒否ベースのメトリクスを導入します。特性誘発偏差は特性なしのベースラインからのデータセットレベルの偏差を測定し、特性誘発フリップ率は同じリクエストが特性間で異なる安全性決定を受けるかどうかを測定します。次に、特性によって引き起こされる安全性の変化の背後にあるメカニズムの表現レベルの分析を提供し、特性が低次元部分空間内でモデルの安全性表現を乱すことを発見します。安全動作が特性間で安定している特性不変安全性を実現するために、LLM の特性条件付き動作と特性なしの動作を調整する、シンプルで効果的な自己蒸留フレームワークである特性不変安全チューニング (TIST) を導入します。私たちの分析に基づいて、特定された特性部分空間内でのみ不変性を強制する TIST のインスタンス化である特性部分空間中性化 (TraSN) をさらに提案します。実験では、TraSN が一般的な能力を維持しながら、形質不変の安全性を向上させ、有害な要求の安全性を強化することが示されています。私たちの結果は、LLM の安全性と堅牢なモデルの動作における重要な要素としての特性を強調しています。
原文 (English)
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning
Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit substantially different safety decisions under different traits assigned in the system prompt, a failure mode we call trait-induced safety variation. To measure this failure, we introduce refusal-based metrics: Trait-Induced Deviation measures dataset-level deviation from the no-trait baseline, while Trait-Induced Flip Rate measures whether the same request receives different safety decisions across traits. We then provide a representation-level analysis of the mechanism behind trait-induced safety shifts and find that traits perturb the model's safety representations within a low-dimensional subspace. To achieve trait-invariant safety, where safety behavior remains stable across traits, we introduce Trait-Invariant Safety Tuning (TIST), a simple yet effective self-distillation framework that aligns an LLM's trait-conditioned behavior with its no-trait behavior. Guided by our analysis, we further propose Trait-Subspace Neutralization (TraSN), an instantiation of TIST, which enforces invariance only within the identified trait subspace. Experiments show that TraSN improves trait-invariant safety and strengthens harmful-request safety while preserving general capability. Our results highlight traits as an important factor in LLM safety and robust model behavior.
ベイジアン更新による確率分布の比例類推
アナロジーは、「A は B に対して、C は D に対して」という形式の 4 項関係です。類推推論のさまざまな形式化の中でも、比例類推は、一連の公準を通じて有効な類推を特徴付けることにより、重要な公理的枠組みを提供します。比例類似性はブール、記号、実数値の領域にわたって広く研究されてきましたが、確率分布への拡張はほとんど研究されていません。この論文では、ベイズ更新に基づいた確率分布の比例類推の概念を導入します。私たちのアプローチは、適切な観測セットによって引き起こされるベイジアン更新を通じて一方を他方に変換できる場合は常に 2 つの分布が関連しているという考えに基づいています。指数関数族のいくつかの標準メンバーについてこのフレームワークを調査し、それが混合ガウス近似を通じて任意の確率分布にどのように自然に拡張されるかを議論します。
原文 (English)
Proportional Analogies on Probability Distributions via Bayesian Updating
Analogies are quaternary relations of the form "A is to B as C is to D". Among the various formalizations of analogical reasoning, proportional analogies provide an important axiomatic framework by characterizing valid analogies through a set of postulates. While proportional analogies have been extensively studied over Boolean, symbolic, and real-valued domains, their extension to probability distributions remains largely unexplored. In this paper, we introduce a notion of proportional analogy for probability distributions based on Bayesian updating. Our approach builds upon the idea that two distributions are related whenever one can be transformed into the other through Bayesian updating induced by a suitable set of observations. We investigate this framework for several standard members of the exponential family and discuss how it naturally extends to arbitrary probability distributions through Gaussian mixture approximations.
Harness-IF: コーディング エージェントの命令サーフェス全体に続く命令の評価
コーディングエージェントがルールに従うとき、それはとにかくそうしようとしていただけかもしれません。既存の命令に従うベンチマークでは違いがわかりません。ベンチマークはユーザーの順番にルールを集中させますが、コーディング エージェントのベンチマークは最終的なタスクの成功を重視します。 Harness-IF は、実行証拠から一度に 1 つずつ操作ルールをスコアリングします。642 ルール ライブラリから抽出された 60 の現実的なマルチターン コーディング項目、判定を受け取る 256 のルールが、展開されたエージェントが読み取る 5 つの構成可能なサーフェスに配置されます。コンプライアンスと偶然を区別するために、Against-Prior Accuracy (AP-Acc) を導入します。これは、プロンプトなしのデフォルトに反するとラベル付けされたルールのみをスコアリングします。これは、9 つのプローブ ビルドにわたって保留されたルールでタスクを再実行することによって観察され、それ以外の場合はキュレーションされます。 12 のフロンティア モデル全体で、精度は 72.1 ~ 85.9%、AP-Acc は 66.1 ~ 78.6% に及びます。すべてのモデルは、以前のルールに対して 3.6 ~ 7.4 ポイント (平均 5.81) 悪化しており、方向はアイテムがクラスター化された間隔での共通サポート分析でも生き残っています。したがって、集計スコアは、モデル固有のマージンによってコンプライアンスを誇張します。つまり、事前の制御では、上位のビルドは変更されず、3 つの隣接するランクのペアが交換されます。 9 つの個別のビルドで競合のバランスをとったパイロットにより、2 番目の結果が追加されました。プールされた優先順位はプロンプトの深さに従わず、システム プロンプト、プロジェクト ファイル、ユーザーの指示がツールとスキルの説明よりも先に表示されます。
原文 (English)
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success. We introduce Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads. To separate compliance from coincidence we introduce Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing unprompted defaults, observed by re-running tasks with the rule withheld across nine probe builds and curated otherwise. Across 12 frontier models, accuracy spans 72.1-85.9% and AP-Acc 66.1-78.6%; every model is worse on against-prior rules, by 3.6 to 7.4 points (mean 5.81), and the direction survives a common-support analysis with item-clustered intervals. Aggregate scores therefore overstate compliance by a model-specific margin: prior control leaves the top build unchanged and exchanges three adjacent rank pairs. A counterbalanced conflict pilot on nine separate builds adds a second result: pooled precedence does not follow prompt depth, with system prompts, project files, and user instructions ahead of tool and skill descriptions.
HyperAFIS: 双曲幾何学による適応ニューロファジィシステムにおけるルール表現と解釈可能性の強化
適応ニューロファジィ推論システム (AFIS) は、明示的な IF-THEN ファジィ ルールを生成できる解釈可能な推論フレームワークであり、透過的な推論を必要とするタスクに適しています。ただし、既存の AFIS モデルは一般にルール先行件を構築し、ユークリッド空間で推論を実行するため、表現能力と予測パフォーマンスが制限されます。この問題に対処するために、AFIS を双曲的に拡張した Hyperbolic AFIS (HyperANFIS) を提案します。 HyperAFIS は、従来の AFIS のファジー セマンティクスとコア アーキテクチャを維持しながら、双曲空間でルール プロトタイプの学習、ルールのアクティブ化、および結果としての集計を実行します。また、解釈可能な IF-THEN ルールを生成する機能も保持します。 HyperAFIS は、双曲幾何学の表現特性を利用することでファジー推論プロセスを強化し、それによって予測精度、ルール間の連携、および解釈可能なルールの信頼性を向上させます。実験結果は、HyperAFIS がすべてのデータセットにわたって標準の AFIS ベースラインおよびさまざまな AFIS バリアントよりも常に優れたパフォーマンスを示し、同時に高品質のファジー ルールを生成することを示しています。
原文 (English)
HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry
The adaptive neuro-fuzzy inference system (ANFIS) is an interpretable reasoning framework capable of generating explicit IF-THEN fuzzy rules, making it suitable for tasks requiring transparent reasoning. However, existing ANFIS models generally construct rule antecedents and perform inference in Euclidean space, limiting their representational capacity and predictive performance. To address this issue, we propose Hyperbolic ANFIS (HyperANFIS), a hyperbolic extension of ANFIS. HyperANFIS preserves the fuzzy semantics and core architecture of conventional ANFIS while performing rule-prototype learning, rule activation, and consequent aggregation in hyperbolic space. It also retains the ability to generate interpretable IF-THEN rules. By exploiting the representational properties of hyperbolic geometry, HyperANFIS strengthens the fuzzy inference process, thereby improving predictive accuracy, inter-rule collaboration, and the credibility of its interpretable rules. Experimental results show that HyperANFIS consistently outperforms the standard ANFIS baseline and various ANFIS variants across all datasets, while also generating higher-quality fuzzy rules.
スリーピング エージェント: Gist ベースのコンテキスト圧縮で失われるものとその理由
要旨ベースのコンテキスト圧縮 (古い会話履歴をコンパクトな表現に要約する) は、長期言語モデル エージェントでは一般的なアプローチですが、さまざまな種類の記憶検索に対するその効果はよく理解されていません。私たちは、Gist 圧縮がどのような場合に役立ち、どのような場合に害を及ぼすかを調査するための診断プローブとして、睡眠ベースの記憶固定を動機とする生物学的にインスピレーションを得た圧縮フレームワークである Salience-Weighted Consolidation (SWC) を使用しています。 SWC は、顕著性によって会話履歴をスコア化し、それを優先度の層に分割し、構造化された要点の抽象化を優先度の中程度のコンテンツに適用します。 LoCoMo の 10 件の会話すべてについて 4 つの条件を評価します -- 合計 1,935 件の一致するテキストのみの質問、カテゴリ 5 (敵対的) 質問を除外した後の主な集計で 1,501 件の質問が使用されます -- 温度 0 では、一貫したタスクタイプの相互作用が見つかりました。要点圧縮は、マルチホップ推論やシングルホップの事実に関する質問では切り捨てを大幅に上回っていますが、一時的な質問は圧縮下では依然として大幅に困難であり、圧縮条件のスコアはフルコンテキスト参照を大幅に下回っています。両方が評価される会話。この失敗の原因を特定のメカニズムにたどります。Gist 抽象化プロンプトは、日付と時刻を破棄しながら、リレーショナル構造とイベント構造を保持します。 10 件の会話すべてにわたる保存分析により、そのメカニズムが確認されました。1 文のプロンプト変更により、時間的表現の保存が約 20 倍増加 (3.05% から 62.39%) する一方で、名前付きエンティティとイベントの保存率はほとんど変化せず (1.02 倍と 1.11 倍)、修正が精密な手段であることが実証されました。プロンプト修正により、一致セット内のカテゴリー 2 (時間的) 質問に対する判定精度が +0.314 [0.254, 0.375] 回復します。コードと結果: https://github.com/kyrkewood/sleeping-agent
原文 (English)
The Sleeping Agent: What Gist-Based Context Compression Loses and Why
Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon language model agents, yet its effect on different types of memory retrieval is poorly understood. We use Salience-Weighted Consolidation (SWC), a biologically-inspired compression framework motivated by sleep-based memory consolidation, as a diagnostic probe to study when gist compression helps and when it hurts. SWC scores conversation history by salience, partitions it into priority tiers, and applies structured gist abstraction to mid-priority content. Evaluating four conditions on all ten LoCoMo conversations---1,935 matched text-only questions in total, 1,501 used in the primary aggregate after excluding Category 5 (adversarial) questions---at temperature 0, we find a consistent task-type interaction: gist compression substantially outperforms truncation on multi-hop reasoning and single-hop factual questions, but temporal questions remain substantially harder under compression, with compressed conditions scoring well below the full-context reference on the conversations where both are evaluated. We trace this failure to a specific mechanism: the gist abstraction prompt preserves relational and event structure while discarding dates and times. A preservation analysis across all ten conversations confirms the mechanism: an approximately 20-fold increase in temporal expression preservation (3.05% to 62.39%) with a one-sentence prompt modification, while named entity and event preservation rates barely change (x1.02 and x1.11), demonstrating that the fix is a precision instrument. The prompt modification recovers +0.314 [0.254, 0.375] judge accuracy on category-2 (temporal) questions in the matched set. Code and results: https://github.com/kyrkewood/sleeping-agent.
エージェントのスキルは有害になる可能性があります: LLM エージェントにおけるスキル誘発性の失敗に関する実証研究
エージェント スキルは、再利用可能なガイダンスを使用して LLM エージェントを拡張するための事実上のメカニズムです。スキルは、計画、ツールの使用、問題解決、検証など、エージェントのタスクの実行を形作ることができます。以前の研究では、エージェント スキルのさまざまな結果が報告されていました。一部のスキルはタスクの成功率を向上させますが、他のスキルは効果がなく、トークンの使用と実行時間が増加し、成功率が低下することさえありました。このペーパーでは、タスクの失敗とコストの回帰を特定のロードされたスキルに帰することにより、スキルに起因するエージェントの失敗の包括的な分析を示します。ターゲットのスキルガイド付き実行を、同じタスクを解決する、またはより安価に解決するスキルなしまたは意味的に一致するスキル参照実行と比較することにより、失敗または回帰をスキルに帰する差分分析フレームワークを導入します。このフレームワークを SkillsBench と SWE-Skills-Bench でインスタンス化すると、125 件の機能障害と 182 件の効率低下を含む、スキルに起因する 307 件の障害が発生しました。また、ペアになったケースを正規化し、差異証拠を抽出し、トリアージ レポートを作成する、分類法に基づいたアトリビューション ツールである SkillTriage も構築します。私たちの主な発見は次のとおりです。(1) スキルによって引き起こされる機能障害は、明らかに無関係なスキルによって引き起こされることはほとんどありません。むしろ、一見関連性があるように見えるスキルによって、エージェントがタスクに必要な実装要素を誤って実装したり省略したりすることがよくあります。 (2) スキルに起因する効率の低下は、プロンプトの長さだけでは説明できません。 (3) 過剰な手順の最大の原因は過剰な検証と大量の実装パイプラインであり、それぞれ 67 件と 30 件のケースに寄与しています。これは、スキルが検証チェックリストや構築レシピを必須の作業に変えることが多いことを示しています。調査結果に基づいて、より安全でコストを意識したスキルの再利用のための研究トピックとツールの改善を提案します。
原文 (English)
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task success rates, while others have no effect, increase token use and execution time, and even reduce success rates. This paper presents a comprehensive analysis of skill-induced agent failures by attributing task failures and cost regressions to specific loaded skills. We introduce a differential analysis framework that attributes a failure or regression to a skill by comparing a target skill-guided run against a no-skill or semantically matched skill reference run that solves the same task, or solves it more cheaply. We instantiate this framework on SkillsBench and SWE-Skills-Bench, yielding 307 skill-induced failures, including 125 functional failures and 182 efficiency regressions. We also build SkillTriage, a taxonomy-guided attribution tool that normalizes paired cases, extracts differential evidence, and produces triage reports. Our major findings include: (1) Skill induced functional failures are rarely caused by obviously irrelevant skills; instead, seemingly relevant skills often make the agent incorrectly implement or omit task-required implementation elements. (2) Skill-induced efficiency regressions are not explained by prompt length alone. (3) The largest sources within Excessive Procedure are excessive verification and heavy implementation pipelines, contributing 67 and 30 cases, respectively. This shows that skills often turn validation checklists and construction recipes into mandatory work. Based on our findings, we propose research topics and tooling improvements for safer and more cost-aware skill reuse.
ルールに対する堅牢な推論のためのロジックとしてのポリシー
税規則から航空会社の手荷物許容量に至るまで、生成 AI システムの多くの実際の応用では、自然言語クエリへの応答は、文書化されたポリシーや規則を尊重する必要があります。我々は、形式論理でポリシーを表現し、推論時に言語モデルの表現力を利用して事実を抽出して述語を根拠付けるハイブリッド記号アプローチと、応答が解釈可能で監査可能であり、示されているように入力摂動下でも正確かつ堅牢であるような推論のための回答セット ソルバーを提示します。具体的には、この抽出ステップと推論ステップの分離が、ほとんどの場合、プロンプトとしてのポリシーおよびコードとしてのポリシーの方法よりも優れており、トークン使用量が最大 10 分の 1 に削減されることを示します。この結果は、客観的な基準を含む堅牢な意思決定を行うために、構造化推論とシンボリック ソルバーを生成モデルと組み合わせることの価値を示しています。
原文 (English)
Policy-as-logic for robust reasoning over rules
In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies or rules. We present a hybrid symbolic approach that expresses policies in formal logic and at inference time exploits the representation power of language models for fact extraction to ground predicates, and an answer set solver for reasoning such that responses are interpretable, auditable, and as we show, accurate and robust under input perturbations. Specifically, we show this separation of extraction and reasoning steps outperforms policy-as-prompt and policy-as-code methods in most cases with ~10x reduction in token usage. The results point to the value of structured reasoning and symbolic solvers in conjunction with generative models to make robust decisions involving objective criteria.
OEIS Open: 言語モデルはどれだけの推測を定理に変えることができますか?
私たちは、OEIS からの 492 のオープンな数学的予想に基づくベンチマークである OEIS Open を構築します。このベンチマークは、Tsoukalas らによって Lean で形式化されています。これらの推測は、これまで特注のエージェントを使用してのみ試みられていましたが、当社のオープンソース評価コードは、それらに対してあらゆる汎用言語モデル (LM) を実行し、LM 不正行為の試みに対して安全です。最小限のツールセットを装備した LM は、1 回の試行あたり $50 の予算でこれらの推測のうち 147 件を解決し、OEIS Open で 30% のスコアを獲得したことがわかりました。 OEIS Open Lite は、より安価な評価を目的とした 100 個の推測のランダムなサブセットです。試行ごとに $200 の予算で評価すると、現時点で最高の LM は OEIS Open Lite で 44% のスコアを獲得します。 LM が arXiv の 476,000 件の論文を介して数学文献にアクセスできるようにしても、OEIS Open Lite のパフォーマンスは向上せず、より洗練されたエージェント ループを使用しても向上しませんでした。この研究で取り上げられている推測は数学的重要性が不確かであり、ほとんどの推測はこれまでほとんど注目されていなかったと思われます。それにもかかわらず、私たちの結果は、LMが未解決の研究の推測を自律的にかつ適度なコストで解決できることを示しています。
原文 (English)
OEIS Open: How many conjectures can language models turn into theorems?
We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source evaluation code runs any generic language model (LM) against them, and is secure against LM cheating attempts. We find that LMs equipped with a minimal set of tools resolve 147 of these conjectures with a budget of \$50 per attempt, scoring 30% on OEIS Open. OEIS Open Lite is a random subset of 100 conjectures for cheaper evaluation. When evaluated with a budget of \$200 per attempt, the best current LM scores 44% on OEIS Open Lite. Giving LMs access to the mathematics literature via 476,000 papers from arXiv did not increase performance on OEIS Open Lite, and nor did using more sophisticated agent loops. The conjectures covered in this work are of uncertain mathematical significance, and most have likely received little previous attention. Nevertheless, our results show that LMs can resolve open research conjectures autonomously and at modest cost.
ExRole: チームの軌跡からマルチエージェント言語モデルの実行可能な役割まで
ロールは、言語モデル エージェントを編成するための解釈可能なインターフェイスを提供しますが、ほとんどのマルチエージェント システムはロールを、学習された動作やパラメータの更新とは切り離された手書きのプロンプト ラベルとして扱います。私たちは、有用な役割は実行可能な制御変数であるべきだと主張します。つまり、将来の有用性を予測する行動を要約し、その後の相互作用を導き、その行動の原因となる訓練可能な能力を特定する必要があります。 ExRole は、プレフィックスローカルのチーム トレースから将来を意識したロール プロトタイプを学習し、それらを読み取り可能な指示とトークンに位置合わせされたロール マーカーに解決し、オプションでターンに位置合わせされたクレジットを使用して共有 LoRA ランク スロットをルーティングする、軌跡からロールへのフレームワークです。 MuSiQue と 2WikiMultiHopQA 全体で、ExRole は単一エージェント検索よりもそれぞれ 15.0/14.4 EM/F1 ポイントと 13.5/16.1 EM/F1 ポイント向上しました。最も強力な非 ExRole コントロールに対して、対応するゲインは 11.5/11.6 ポイントおよび 7.7/9.7 ポイントのままです。両方のベンチマークにわたって、制御された結果は、ロールフリー、手動、ランダム、およびシャッフルされた代替案よりも、軌道誘導のロール条件付けを一貫して支持しています。さらに、役割-エージェント-ターン介入は、誘導された役割が、固定されたエージェントのアイデンティティやターンの位置を超えて、移転可能な行動の専門化を捕捉することを示しています。
原文 (English)
ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models
Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent systems treat them as hand-written prompt labels disconnected from learned behavior and parameter updates. We argue that a useful role should instead be an executable control variable: it should summarize behavior predictive of future utility, guide subsequent interaction, and identify the trainable capacity responsible for that behavior. We introduce ExRole, a trajectory-to-role framework that learns future-aware role prototypes from prefix-local team traces, resolves them into readable instructions and token-aligned role markers, and optionally routes shared LoRA rank slots with turn-aligned credit. Across MuSiQue and 2WikiMultiHopQA, ExRole improves over single-agent search by 15.0/14.4 and 13.5/16.1 EM/F1 points, respectively. Against the strongest non-ExRole controls, the corresponding gains remain 11.5/11.6 and 7.7/9.7 points. Across both benchmarks, the controlled results consistently favor trajectory-induced role conditioning over role-free, manual, random, and shuffled alternatives. Role-Agent-Turn interventions further show that the induced roles capture transferable behavioral specialization beyond fixed agent identities or turn positions.
再試行、切り替え、または棄権?制御されたエラー挿入による学習戦略を意識したツール使用ポリシー
ツールを使用する LLM エージェントは通常、ツール呼び出しが確実に成功する環境でトレーニングおよび評価されますが、デプロイされたツールは一時的、永続的、またはサイレントに失敗する可能性があります。したがって、確実なリカバリには、繰り返しの再試行以上のものが必要です。エージェントは、同じパスを再試行するか、代替パスに切り替えるか、実行可能なパスが残っていないことを認識する必要がある場合があります。我々は、障害のないツール使用ベンチマークを、シナリオ制御された解決可能性を備えた制御された確率的環境に変換するフレームワークである BENCH2ROBUST を紹介します。この環境では、利用可能なパスが使い果たされた後にエピソードが再試行、切り替え、または停止を明示的に必要とします。 BENCH2ROBUST を使用して、ベイジアン ツール メモリ (BTM) による構造化されたランタイム回復コンテキストとカリキュラム制御の強化学習という 2 つの補完的な介入を研究します。 4 つのファミリーと 2 つのマルチターン ベンチマーク ファミリの 7 つのモデルにわたって、ツールの故障により、ほぼ普遍的な堅牢性のギャップが生じます。保留された Retail タスクでは、BTM は再トレーニングなしで堅牢性を最大 16.8 パーセント ポイント向上させますが、RL は推論時の BTM がなくても有益な相補的な回復動作を学習します。 2 つを組み合わせると、故障のないパフォーマンスを維持しながら、噴射下で 40.8 ~ 45.5% に達します。これらの結果は、環境固有の回復知識と学習された回復動作を組み合わせることで、堅牢なツールを使用できるというメリットがあることを示唆しています。
原文 (English)
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robust recovery therefore requires more than repeated retries: an agent may need to retry the same path, switch to an alternative, or recognize that no viable path remains. We present BENCH2ROBUST, a framework that converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, where episodes explicitly require retrying, switching, or stopping after available paths are exhausted. We use BENCH2ROBUST to study two complementary interventions: structured runtime recovery context through Bayesian Tool Memory (BTM), and curriculum-controlled reinforcement learning. Across 7 models from 4 families and two multi-turn benchmark families, tool failures produce a near-universal robustness gap. On held-out Retail tasks, BTM improves robustness by up to 16.8 percentage points without retraining, while RL learns complementary recovery behavior that remains beneficial without inference-time BTM. Combining the two reaches 40.8-45.5% under injection while preserving failure-free performance. These results suggest that robust tool use benefits from combining environment-specific recovery knowledge with learned recovery behavior.
効率的なテスト時間推論のためのクレームレベルの信頼性評価
私たちは、テスト時間のスケーリングの原則としてクレーム レベルの改ざんを提案し、クレーム レベルの信頼性評価 (CLR) を通じてそれをインスタンス化します。これは、追加のソリューション サンプリングからターゲットを絞った検証までテスト時間の計算を再割り当てするトレーニング不要のフレームワークです。トレース全体の評価では、ルーチン トークンによる信号の希釈により決定的なエラーが隠れてしまうことが多いため、CLR は各推論トレースを意思決定に重要なクレームのコンパクトなセットに凝縮し、それによってその論理アンカーを分離します。さらに、固定モデル機能の下で完全に正しいソリューションを生成することの本質的な困難を認識し、CLR は焦点をセマンティック改ざんに移します。このアプローチは、解決策の構築と主張の反論の間の根本的な非対称性を利用します。有効な解決策を構築するには、完璧な推論経路が必要ですが、誤った主張に反論するには、決定的な欠陥を 1 つ特定するだけで済みます。この陰性証拠の対象を絞った検索により、信頼性の高い不正確なトレースの生存空間が体系的に圧縮され、非線形信頼性スコアリングによって誤ったコンセンサスが効果的に抑制されます。一致した予算の下で 4 つの LLM と 4 つの推論ベンチマークにわたって、CLR は一般に pass@1 と自己一貫性を向上させます。たとえば、GPT-OSS-20B/CMIMC25 では、CLR は pass@1 を 27.15 パーセンテージ ポイント上回り、自己整合性の精度が 77.50\% から 82.19\% に向上し、トークンの数は 37.0\% 減少しました。
原文 (English)
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution sampling to targeted verification. Since whole-trace evaluation often obscures decisive errors due to signal dilution from routine tokens, CLR condenses each reasoning trace into a compact set of decision-critical claims, thereby isolating its logical anchors. Furthermore, recognizing the inherent difficulty of generating entirely correct solutions under fixed model capabilities, CLR shifts the focus to semantic falsification. This approach exploits a fundamental asymmetry between solution construction and claim refutation. Constructing a valid solution requires a flawless reasoning path, whereas refuting an incorrect claim requires identifying only a single decisive flaw. This targeted search for negative evidence systematically compresses the survival space of high-confidence incorrect traces, effectively suppressing erroneous consensus via nonlinear reliability scoring. Across four LLMs and four reasoning benchmarks under matched budgets, CLR generally improves upon pass@1 and self-consistency. On GPT-OSS-20B/CMIMC25, for instance, CLR exceeds pass@1 by 27.15 percentage-points and raises self-consistency accuracy from 77.50\% to 82.19\% with 37.0\% fewer tokens.
CTBench: 現実的な通信ネットワーク運用における AI エージェントのトラブルシューティング機能の評価
エージェントは、ネットワークの運用と保守を自動化する目的でますます検討されています。エンジニアは、厳しい制約の下で動作しながら、ネットワーク障害を診断し、構成を最適化してサービスを強化し、運用コストを削減する必要があります。しかし、既存の評価では、実際のネットワーク特性を正確にモデル化したり、多様なベンダー、デバイス、プロトコル、インターフェイスを備えた部分的に観測可能な通信環境下でエージェントを評価したりすることはできません。このペーパーでは、エージェントが有能な通信トラブルシューティング エンジニアのように行動するかどうかを評価するための公開ベンチマークである CTBench を紹介します。 CTBench は、根本原因の分析とパスの復元に重点を置いています。各タスクは専門家によって構築され、黄金の証拠ステップを含む豊富なタスク メタデータで注釈が付けられます。 CTBench は、最終的な回答と診断証拠の両方を評価する、専門家に基づいた指標を使用します。代表的なハーネス モデルの組み合わせを用いた実験では、最先端のエージェントはパス復元タスクのエンドポイントの特定では非常に優れたパフォーマンスを発揮しますが、より一般的に、根本原因の分析ではパフォーマンスが劣ることが示されています。特に、エージェントはインターフェイス状態、リンク層、サービス管理、その他の運用上の障害に悩まされます。最も重要なことは、エージェントがもっともらしい、または正しい最終的な答えを出したとしても、業務の実践に必要な証拠に基づいた診断を提供できないことが多いということです。さらに、我々の結果は、パスの復元は一般的により多くのリソースを消費しますが、リソースの使用量が多いほど必ずしも診断が適切になるわけではないことを示しています。
原文 (English)
CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations
Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance services, and reduce operational costs while acting under strict constraints. However, existing evaluations fail to accurately model real network characteristics or assess agents under partially observable telecom environments with diverse vendors, devices, protocols, and interfaces. In this paper, we introduce CTBench, a public benchmark for assessing whether an agent behaves like a competent telecom troubleshooting engineer. CTBench focuses on root cause analysis and path restoration. Each task is constructed by experts and annotated with rich task metadata, including golden evidence steps. CTBench uses expert-grounded metrics that evaluate both final answers and the diagnostic evidence. Experiments with representative harness-model combinations show that state-of-the-art agents perform very well at identifying endpoints in path-restoration tasks but, more generally, underperform in root cause analysis. In particular, agents struggle with interface state, link-layer, service-management, and other operational faults. Most importantly, even when agents produce plausible or correct final answers, they often fail to provide the evidence-grounded diagnoses required in operational practice. Our results further show that path restoration is generally more resource expensive, yet larger resource usage does not necessarily translate into better diagnosis.
Mechanist: 知能のメカニズムを解明するための科学的手段としての AI
AI モデルはさまざまな分野で目覚ましい成功を収めていますが、その機能の根底にあるメカニズムとそれがもたらす可能性のあるリスクについてはまだ十分に理解されていません。 AI 開発の高速化と自動化が進む一方で、メカニズムの探索は依然として手作業が多く、モデルが実行できることと、それを理解して制御する人間の能力との間のギャップが拡大しています。このギャップを埋めるために、AI インテリジェンスの基礎となるメカニズムを自律的に発見するための科学機器として AI を使用するエージェント システムである Mechanist を紹介します。自律的なメカニズムの発見をサポートするために、約 13,000 件の論文からなる解釈可能性に焦点を当てたナレッジ グラフを構築し、それを 26 分野にわたる 4,300 万件の論文からなる学際的なデータベースと統合します。さらに、メカニズム分析、因果関係介入、検証のための 32 の基本的な手法のライブラリを厳選しています。 Claude Code や既存の AI サイエンティスト システムと比較して、Mechanist はより価値のあるメカニズム仮説を生成し、より確実に実験を実行します。 Mechanist はまた、モデルの動作の発見から AI モデルの説明と制御への進歩を示します。具体的には、Mechanist はまず、科学実験室における直観に反する安全性リスクを明らかにし、安全でない特性が一見安全なトレーニング データを通じてモダリティを越えて伝達される可能性があることを示しました。次に、Mechanist は信念のメカニズム理論を開発し、モデルがどのように世界の知識を表し、信念を形成し、他の人の信念を推測するか、およびこれらのメカニズムが事前トレーニング中にどのように現れるかを明らかにします。最後に、Mechanist はこれらの機構的な洞察を実践的な介入に変換し、さまざまなシナリオにわたってモデルのパフォーマンスを向上させ、指定された特性を持つ DNA 配列の生成に向けて科学的基礎モデルを導きます。
原文 (English)
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.
グラフ構造のルーブリック: LLM 審査員向けにルーブリックを入力済みの評価グラフに編集する
ルーブリックベースの評価者は通常、ルーブリックをプロンプトコンテキストまたはフラットな基準として扱います。ルーブリックは何を判断するかを指定しますが、自然言語ルールにそれが記載されている場合でも、基準の構成を暗黙的なままにします。グラフ構造ルーブリック (GSR) を導入します。これは、応答を観察する前にルーブリックを応答に依存しない型付き評価グラフにコンパイルします。基準ノードは判断を導き出します。変換、リダクション、およびゲート演算子は、名前付きポートを通じてそれらを構成します。そして、Readout と呼ばれるタスク固有の出力マッピングは、固有のシンクをスコアまたは好みに変換します。コンパイルでは、不正な形式のグラフや型の互換性のないグラフは拒否されます。ポイントごとの評価では、グラフを集計する前にルーブリックの次元を個別に判断します。ペアワイズ評価では、あらゆる基準の下で各候補に対して 1 つの判定を行ってグラフを再利用します。 GPT-OSS-120B の下では、GSR は 4 つの点ごとのデータセットでの Prometheus スタイルのスコアリングよりも正確なスコアの一致を 0.62 ~ 6.75 パーセント ポイント改善し、ネイティブ タイおよび棄権ポリシーに基づく 2 つの優先順位ベンチマークで数値的に最高のエンドツーエンドのペアごとの精度を達成します。
原文 (English)
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges
Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to judge but leave criterion composition implicit, even when natural-language rules state it. We introduce Graph-Structured Rubrics (GSR), which compiles a rubric into a response-independent typed evaluation graph before observing responses. Criterion nodes elicit judgments; transformation, reduction, and gating operators compose them through named ports; and a task-specific output mapping, termed Readout, converts the unique sink into a score or preference. Compilation rejects malformed or type-incompatible graphs. Pointwise evaluation judges rubric dimensions separately before graph aggregation; pairwise evaluation reuses the graph with one judgment for each candidate under every criterion. Under GPT-OSS-120B, GSR improves exact score agreement by 0.62--6.75 percentage points over Prometheus-style scoring on four pointwise datasets and achieves the numerically highest end-to-end pairwise accuracy on two preference benchmarks under native tie and abstention policies.
ガイド: エンタープライズ設定でのドキュメントからアーティファクトの生成のためのガバナンド統合インテリジェンス
企業ガイドライン文書は、説明文、複雑な表、埋め込み画像を組み合わせた、異種混合かつマルチモーダルなものです。既存の LLM および VLM システムは、幻覚的なコンテンツ、テーブル構造の劣化に直面しており、抽出から検証、アーティファクト生成に至るまでの管理されたワークフローが不足しています。そのため、企業はこれを手動で実行する必要があり、ドキュメントごとに 2 ~ 3 日かかります。これに対処するために、スキーマ検証されたエージェント間契約とエンドツーエンドの出所追跡を備えた共有バージョン管理ルール ストア上に構築された管理されたマルチエージェント フレームワークである GUIDE を導入します。 6 つの専門エージェントが、解析、VLM 主導の抽出、整合性チェック、評価、人間参加型 (HITL) エスカレーション、およびペルソナに合わせたアーティファクト合成を処理します。 120 の実際のエンタープライズ ガイドライン ドキュメントで評価された GUIDE は、ドキュメントの成功率 96% を達成し、71.4% が自動承認された 3,896 のルールを抽出し、すぐに展開できる 812 個のアーティファクトを生成し、ドキュメントあたりの所要時間を 40 ~ 125 分に短縮しました。
原文 (English)
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings
Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images. Existing LLM and VLM systems face hallucinated content, table structure degradation, and lack governed workflows extending beyond extraction to validation and artifact generation. This leaves enterprises to perform this manually, consuming 2-3 days per document. To address this, we introduce GUIDE, a governed multi-agent framework built on a shared versioned rule store with schema-validated inter-agent contracts and end-to-end provenance tracking. Six specialized agents handle parsing, VLM-driven extraction, consistency checking, evaluation, human-in-the-loop (HITL) escalation, and persona-tailored artifact synthesis. Evaluated on 120 real-world enterprise guideline documents, GUIDE achieves 96% document success, extracts 3,896 rules with 71.4% auto-approved, produces 812 deployment-ready artifacts, and reduces turnaround to 40-125 minutes per document.
誰が一番良いと思うかは、どのくらいの期間放置するかによって決まります: LLM 評価における予算に応じたランキング
大規模な言語モデルの標準評価では、推論条件全体でモデルのランキングが安定していることが前提となります。私たちは、トークン生成予算、つまりモデルが生成できる最大トークンを 7 つのレベル (64 ~ 4,096) にわたって変化させ、3 つの推論ベンチマーク (56,476 推論) で 4 つのモデルを評価することで、この仮定に異議を唱えます。我々は 4 つの発見を報告します: (i) 切り捨てを制御した後でも、項目の 3 ~ 19% が非単調な動作 (予算が増えると精度が低下する) を示し、この現象はモデル固有です (モデル間の重複: 6 ~ 14%)。 (ii) モデルのランキングは、すべてのベンチマークの予算全体で逆転します ($p {<} 0.01$、McNemar)。 (iii) Oracle の分析により、$+27.8$pp までのモデルの相補性が明らかになり、これは予算が限られている場合に最も顕著になります。 (iv) 予算を意識したルーターは、オラクル ギャップ クロスドメインの 14.1% をキャプチャします。予算機能はドメイン内では役立ちます ($+1.6$ ~ $+5.7$pp) が、ドメイン固有であり転送に悪影響を及ぼします ($-1.2$pp)。これらの結果は、予算に応じた評価プロトコルの正当性を主張します。
原文 (English)
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), evaluating four models on three reasoning benchmarks (56,476 inferences). We report four findings: (i) 3--19% of items exhibit non-monotone behavior (accuracy decreasing with more budget), even after controlling for truncation, and this phenomenon is model-specific (cross-model overlap: 6--14%). (ii) Model rankings reverse across budgets on all benchmarks ($p {<} 0.01$, McNemar). (iii) Oracle analysis reveals model complementarity up to $+27.8$pp, most pronounced at constrained budgets. (iv) A budget-aware router captures 14.1% of the oracle gap cross-domain; budget features help within-domain ($+1.6$ to $+5.7$pp) but are domain-specific and hurt transfer ($-1.2$pp). These results argue for budget-conditioned evaluation protocols.
Oracle 予算の使い方: タンパク質構造予測モデルの実践的なガイダンス
タンパク質構造予測のための基礎モデルは、特定の標的に関しては信頼性が低いままです。外部のオラクルはこれらの失敗にフラグを立てて修正できますが、生物学的なオラクルは高価であり、オラクルの予算が重大な制約になります。 FK ステアリング、DPO、ベスト K-of-N サンプリングなどの既存のガイダンス手法は、この予算の使い方が異なりますが、手法の選択をガイドする体系的な比較は存在しません。このギャップを埋めるために、生成モデルの潜在部分空間内で既製のオプティマイザーを適用する、最近提案された Optimization Over Outputs (O3) と並行して、これらの手法をベンチマークします。 O3 の使用をタンパク質構造予測モデルに拡張します。全体として、私たちの研究は、オラクルの予算を意識したガイダンスのための最初の実用的な参考資料を提供します。カルモジュリン (1CLL) と大腸菌アスパラギン酸トランスカルバモイラーゼ (9EEH) という 2 つのタンパク質標的に対する我々の評価では、すべての予算とオラクルにわたって一貫して優位に立つ単一の方法はないことが明らかになりました。具体的には、O3 はオラクル バジェットが低い場合に最も効果的であることが証明されていますが、FK ステアリングと DPO はバジェットが増加するにつれてパフォーマンスが向上することがわかります。私たちはこれらの調査結果を抽出して、現実世界のオラクル予算の制約の下で業務を行う実務者向けの実用的な推奨事項を作成します。
原文 (English)
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint. Existing guidance methods, such as FK-steering, DPO, and Best K-of-N sampling, differ in how they spend this budget, yet no systematic comparison exists to guide method selection. To bridge this gap, we benchmark these methods alongside the recently proposed Optimisation Over Outputs (O3), which applies off-the-shelf optimisers within a generative model's latent subspace. We extend the usage of O3 to protein structure prediction models. Overall, our work provides the first practical reference for oracle budget-aware guidance. Our evaluation on two protein targets, calmodulin (1CLL) and E. coli aspartate transcarbamoylase (9EEH), reveals that no single method consistently dominates across all budgets and oracles. Specifically, O3 proves most effective at low oracle budgets, while FK-steering and DPO demonstrate improved performance as the budget increases. We distil these findings into actionable recommendations for practitioners operating under real-world oracle-budget constraints.
レガシー HPC モダナイゼーションのためのエージェント ワークフロー: GAMESS の 2 電子積分コアの変換
レガシー Fortran の最新化は量の問題です。変換は個々では日常的なものですが、コードベースは膨大になる可能性があり、計算科学の多くにおいて作業は単純に元に戻ります。私たちは、この作業を実稼働規模で行うエージェント ワークフローを提案し、そのような委任がどこまで到達できるかを測定することにしました。この作業では、プロンプトに特化した 3 つのエージェントの役割が、エージェント自身が作成および改訂したバージョン管理された仕様に基づいて動作しますが、人間は少数のゲートを保持します。この取り決めは、ドメインから継承された正確な検証オラクルによって安全に保たれており、安全な委任の境界はまさにそのオラクルが認識を停止する場所にあります。提案されたワークフローをケーススタディに適用し、48 年の開発歴史を持つ成熟した量子化学パッケージである GAMESS (General Atomic and Molecular Electronic Structure System) の 2 電子積分ルーチンを、固定形式の Fortran 77 から自由形式の Fortran 2008 に変換します。この作業の範囲は、12 個のソース ファイル、56,448 行、および計算用の 225 個のサブルーチンでした。電子反発積分。エージェントは分離されたワークツリーで 3 つのクロード コードの役割として実行され、作業はクロード モデルの 4 世代に及びました。 GAMESS グループは、そのユーザー コミュニティが印刷されたエネルギーを標準として扱う標準テスト スイートを出荷しているため、それらのエネルギーのビットごとの再現をマージ基準として採用できます。この場合、小数点第 12 位の偏差はドリフトではなく失敗としてカウントされます。 12 のソース ファイルすべてが、49 の標準 GAMESS テストと 2 つの追加計算で構成される 51 のテスト検証バッテリーに合格し、612 回のテスト実行全体で化学関連の差異の数はゼロであり、すべてのファイルは継続的統合に使用される Jenkins テストにも合格しています。
原文 (English)
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of computational science the work simply goes undone. We propose an agentic workflow that takes this work on at production scale, and we set out to measure how far such delegation can reach. In this work, three prompt-specialized agent roles operate under a version-controlled specification that the agents themselves authored and revised, while humans hold a small number of gates. The arrangement is kept safe by an exact verification oracle inherited from the domain, and the boundary of safe delegation lies exactly where that oracle stops seeing. We apply the proposed workflow in a case study, converting the two-electron-integral routines of GAMESS (General Atomic and Molecular Electronic Structure System), a mature quantum-chemistry package with a 48-year development history, from fixed-form Fortran 77 to free-form Fortran 2008. The scope of this work was twelve source files, 56,448 lines, and 225 subroutines for computing electron repulsion integrals. The agents ran as three Claude Code roles in isolated worktrees, and the work spanned four Claude model generations. Because the GAMESS group ships a standard test suite whose printed energies its user community treats as canonical, we could adopt bit-for-bit reproduction of those energies as the merge criterion, where a deviation in the twelfth decimal place counts as a failure rather than drift. All twelve source files pass a 51-test validation battery comprising the 49 standard GAMESS tests and two additional calculations, and across 612 test runs the number of chemistry-relevant differences is zero, and every file also passes the Jenkins tests that are used for continuous integration.
VAKRA: API にわたるマルチホップ推論の評価とツール使用ポリシーに基づく取得
企業環境に導入されたエージェントは、構造化された API とドキュメント コレクション全体を推論する必要がありますが、既存のベンチマークはこれらの機能を個別に評価します。 VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents) を導入します。これは、$62$ のドメイン全体で $8{,}000$ を超える実行可能 API のベンチマークであり、タスクの難易度は、多様な API インタラクション スタイル、構造化 API を介したマルチホップ推論、および自然言語によるマルチソース推論の 3 つの設定にまたがります。ツール使用ポリシーの制約。正確さは、複数の有効なパスに対応して、ライブ API に対して予測されたツール呼び出しを再実行することによって検証されます。固定 ReAct ハーネスを使用してモデル機能をエージェント アーキテクチャから分離し、フロンティア モデルとオープンウェイト モデルを評価したところ、最高のモデルでもシングルホップ エンドポイント スタイルのタスクでは 70.4\% しか達成できず、構成 API では 50 ~ 51\% に低下することがわかりました。推論の深さが増すにつれてパフォーマンスは 50\% 以上低下し、ポリシーに制約された質問では重大な障害が発生します (答えられないクエリでは最低 2.4\%)。トレース分析では、ツール呼び出しの仕組みではなく、言語を介した推論、つまりエンティティの曖昧さの解消、ソース間のグラウンディングに障害が集中していることが示されています。コードは https://github.com/IBM/VAKRA から入手できます。データセットは https://huggingface.co/datasets/ibm-research/VAKRA から入手できます
原文 (English)
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $8{,}000$ executable APIs across $62$ domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning with natural-language tool-use policy constraints. Correctness is verified by re-executing predicted tool calls against live APIs, accommodating multiple valid paths. Using a fixed ReAct harness to isolate model capabilities from agent architecture, we evaluate frontier and open-weight models and find that even the best model achieves only 70.4\% on single-hop endpoint-style tasks and drops to 50--51\% on compositional APIs; performance degrades by over 50\% as reasoning depth increases, and policy-constrained questions expose severe failures (as low as 2.4\% on unanswerable queries). Trace analysis shows failures concentrate at language-mediated reasoning - entity disambiguation, cross-source grounding, rather than tool invocation mechanics. Code is available https://github.com/IBM/VAKRA. Dataset is available https://huggingface.co/datasets/ibm-research/VAKRA
検索拡張された大規模言語モデルを使用した、複雑なシステム診断のためのナレッジ グラフとしての動的マスター ロジック モデルの構築
ダイナミック マスター ロジック (DML) は、機能目標を基礎となる構造要素にリンクすることにより、システムの動作を表現するための階層フレームワークを提供します。ただし、DML の構築は通常、技術文書の専門家の解釈に依存するため、複雑なシステムの拡張性が制限されます。この研究では、検索拡張生成と大規模言語モデルを実現ツールとして使用し、システム記述とナレッジ グラフ (KG-DML) としての表現から DML モデルを自動構築するためのフレームワークを紹介します。このフレームワークは、小規模システムでの以前の作業に基づいて、自動化された KG-DML の構築と評価を大幅に大規模で複雑なシステムに拡張します。モデルの構築は、機能の依存関係と明示的な論理関係を維持しながら、ターゲットを絞った検索を使用して DML 階層全体で進行します。結果として得られる KG-DML は、診断推論、安全性評価、上向きの障害伝播、および下向きの依存関係の追跡をサポートします。マルチレベルの検証手法により、レイヤー固有の精度と再現率、論理ゲートの一貫性、および全体的な構造の完全性が評価されます。廃止された沸騰水型原子炉の低圧冷却材注入システムへの適用により、繰り返しの運転を通じて一貫した再構築が実証されました。結果は、自動化された KG-DML 構築により、技術文書を診断および信頼性分析のための実行可能な機能モデルに変換できることを示しています。
原文 (English)
Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models
Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlying structural elements. However, DML construction typically relies on expert interpretation of technical documentation, limiting scalability for complex systems. This study presents a framework for automated construction of DML models from system descriptions and their representation as Knowledge Graphs (KG-DML), using Retrieval-Augmented Generation and Large Language Models as enabling tools. Building on prior work with small-scale systems, the framework extends automated KG-DML construction and evaluation to substantially larger and more complex systems. Model construction proceeds across the DML hierarchy using targeted retrieval while preserving functional dependencies and explicit logical relationships. The resulting KG-DML supports diagnostic reasoning, safety assessment, upward failure propagation, and downward dependency tracing. A multi-level validation methodology evaluates layer-specific precision and recall, logical gate consistency, and overall structural integrity. Application to the Low-Pressure Coolant Injection system of a decommissioned Boiling Water Reactor demonstrates consistent reconstruction across repeated runs. The results show that automated KG-DML construction can transform technical documentation into executable functional models for diagnostic and reliability analysis.
サイバーセキュリティにおける LLM 生成の検出ルールの評価
LLM はセキュリティ環境にますます普及していますが、その有効性を測定する手段が限られているため、セキュリティ担当者に対する信頼と有用性が制限されています。ここでは、LLM によって生成されたサイバーセキュリティ ルールを評価するためのオープンソースの評価フレームワークとベンチマーク メトリクスを紹介します。このベンチマークでは、ホールドアウト セットベースの方法論を採用して、人間が生成したルールのコーパスと比較して、LLM が生成したセキュリティ ルールの有効性を測定します。これは、専門家がセキュリティ ルールを評価する方法からインスピレーションを得た 3 つの主要な指標を提供し、LLM ベースのセキュリティ ルール ジェネレータの有効性を現実的かつ多面的に評価します。この方法論は、Sublime Security の検出チームのルールと、Sublime Security の自動検出エンジニア (ADE) によって作成されたルールを使用して説明されており、ADE のスキルの徹底的な分析が結果セクションに示されています。
原文 (English)
Evaluating LLM Generated Detection Rules in Cybersecurity
LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we present an open-source evaluation framework and benchmark metrics for evaluating LLM-generated cybersecurity rules. The benchmark employs a holdout set-based methodology to measure the effectiveness of LLM-generated security rules in comparison to a human-generated corpus of rules. It provides three key metrics inspired by the way experts evaluate security rules, offering a realistic, multifaceted evaluation of the effectiveness of an LLM-based security rule generator. This methodology is illustrated using rules from Sublime Security's detection team and those written by Sublime Security's Automated Detection Engineer (ADE), with a thorough analysis of ADE's skills presented in the results section.
バックトレーダーベンチ: 自己生成された MCQ を使用したアルゴリズム取引における LLM エージェントのベンチマーク
静的ベンチマークにはデータ汚染のリスクがあり、バックテストの数値出力には実際のコード実行からのグラウンド トゥルースが必要であるため、アルゴリズム取引における LLM コーディング エージェントの評価は困難です。 2 つの補完的なパイプラインを備えたフレームワークである Backtrader-Bench を紹介します。決定論的多肢選択質問 (MCQ) パイプラインは、すべての回答を再導出する独立したチェッカーを使用して、5 つの取引戦略、33 のテンプレート、および 3 つの難易度層にわたるバックテスト構成から質問を生成します。ジェネレーターとソルバーのフィルタリング パイプラインは、より難しい質問を自律的にマイニングします。ジェネレーターは、実行可能コードによって検証された質問を書き込み、MCQ に変換し、ツールを使用しないソルバーがコードを実行せずに回答できる質問をすべて破棄します。厳選された 30 の質問セットで、ツールを使用しない 11 のモデル (それぞれ 10 回実行) とツールを使用する 4 つの構成を評価します。ツールで強化されたエージェントは、シングル パス (GPT-5.5 および Opus 4.7) で 90.0% の精度に達し、ツールを使用しない最良のベースライン (73.0%、10 回の実行の平均) を 17 パーセント ポイント上回りました。個別にマイニングされた 38 の質問では、ツールなしの精度はさらに低下し、モデルの半数がほぼランダムな偶然のレベル (25%) に低下しました。スケーラブルな MCQ インフラストラクチャは、評価を超えて、強化学習用のトレーニング コーパスを生成するように設計されており、定量的取引ワークフローに特化したエージェントを構築するという最終目標を掲げています。
原文 (English)
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs require ground truth from actual code execution. We present Backtrader-Bench, a framework with two complementary pipelines. A deterministic multiple-choice question (MCQ) pipeline generates questions from backtest configurations across five trading strategies, 33 templates, and three difficulty tiers, with an independent checker that re-derives every answer. A generator-solver filtering pipeline autonomously mines harder questions: a generator writes questions verified by executable code, converts them to MCQs, and discards any that a no-tool solver can answer without code execution. We evaluate 11 models without tools (10 runs each) and four with-tools configurations on a 30-question curated set. Tool-augmented agents reach 90.0% accuracy in a single pass (GPT-5.5 and Opus 4.7), outperforming the best no-tools baselines (73.0%, averaged over 10 runs) by 17 percentage points. On 38 separately mined questions, no-tools accuracy drops further, with half the models falling to roughly random-chance level (25%). Beyond evaluation, the scalable MCQ infrastructure is designed to produce a training corpus for reinforcement learning, with the ultimate goal of building a specialized agent for quantitative trading workflows.
事前トレーニング済み言語モデルへのリカレント深度の改良: 2 つのパラメーター予算でのインストール、外挿、転送、および保持
高密度の事前トレーニング済み言語モデルは、再帰的な深さで改良し、結果のみのアニーリング後に持続する反復的な潜在遷移を学習できます。 Qwen2.5-0.5B-Instruct は、Prelude、重み付けされた Recurrent Block、および Coda に分割され、アイデンティティを保持する 1 ループ パスと、後のループのリエントリ ブリッジを備えています。ループ 1 では、改造は事前に登録された ARC バッテリーのベースに対して劣らないままです。 Three findings.まず、このメカニズムは端末応答検索ではなく再利用可能な手順であり、固定基本重みを超える 600 万のトレーニング済みパラメーターと 180M のフルブロックという 2 つの予算でインストールされます。中間ステップの監視を使用すると、モデルはループごとに 1 つのタスク ステップを計算し、最終解答のみが採点されるときに持続します。アダプターは全体的にブロック全体と一致し (83.8% 対 84.0%)、深さ 11 までリードされ、それを超えてトレーリングされました。言語の微調整は、制御された言語レンダリングで 79 ~ 86% に達し (ゼロショット転送は最小限でした)、インストールされたメカニズムから開始されたアダプターの言語トレーニングは、差し出されたテスト セットを含め、一致した新しいトレーニングを 18.6 ポイント上回りました。第 2 に、この操作は監視深度の約 1.5 倍まで外挿し、深さ 18 まで 70% の精度を維持します。第 3 に、同じサイズのスクラッチパッドでトレーニングされたモデルは、学習された範囲内ではリカレント モデルと一致しましたが、それを超えると崩壊しました。反復モデルは全体的に 84% 対 72% で勝利し、深さ 10 を超えても 53% 対 2.5% を維持し、応答速度は 7.6 倍でした。したがって、反復変換器は、システムレベルの比較において、同じタスクで微調整された同等のモデルまたはより大きなモデルよりも速く、潜在空間でより深い推論を実行できます。ルールを逆に実行する 2 番目のタスクでは、限界が明らかになりました。逆は単独で学習可能ですが、インストールされたメカニズムと一般的な機能、つまり壊滅的な干渉の境界を維持しながら、それを継続的に取得することはできませんでした。学習された深さの選択は開いたままになります。
原文 (English)
Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets
A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after outcome-only annealing. Qwen2.5-0.5B-Instruct is split into a Prelude, a weight-tied Recurrent Block, and a Coda, with an identity-preserving one-loop path and a re-entry bridge on later loops. At loop 1 the retrofit remains non-inferior to its base on a preregistered ARC battery. Three findings. First, the mechanism is a reusable procedure rather than terminal-answer lookup, and installs at two budgets: 6M trained parameters over frozen base weights and 180M full-block. With intermediate-step supervision, the model computes one task step per loop and persists when only final answers are graded. The adapter matched the full block overall (83.8% versus 84.0%), led through depth 11, and trailed beyond. Verbal fine-tuning reached 79-86% on controlled verbal renderings (zero-shot transfer was minimal), and adapter verbal training begun from the installed mechanism outpaced matched fresh training by 18.6 points, including on a held-out test set. Second, the operation extrapolates to roughly 1.5 times its supervised depth, holding 70% accuracy through depth 18. Third, a same-size scratchpad-trained model matched the recurrent model within its learned horizon but collapsed beyond it. The recurrent model won overall, 84% versus 72%, retained 53% versus 2.5% beyond depth 10, and answered 7.6 times faster. An iterative transformer can therefore perform deeper reasoning in latent space faster than comparable or larger models fine-tuned on the same task, in a system-level comparison. A second task, running the rule in reverse, exposed the limits: the inverse was learnable in isolation, but no continuation acquired it while preserving the installed mechanism and general capability, a catastrophic-interference boundary. Learned depth selection remains open.
TRACE ベンチ: タスク主導のロールプレイ エージェント チェックリストの評価
ロールプレイの評価は、単一のスコアを割り当てるだけではありません。どの役割要件がテストされ、どれが失敗したか、そしてどの対話証拠が判断を裏付けるかを明らかにする必要があります。私たちは、タスク駆動型エージェントチェックリスト評価フレームワークである TRACE Bench を提案します。各ロール プロファイルをオフラインで固定チェックリストに分解し、ユーザー エージェントを使用してターゲット ロールプレイ モデルと自然に会話しながら、モデルの応答からチェックリストの状態を非公開で更新します。したがって、スコアはブラックボックスの全体的な印象ではなく、チェックリストの項目とそれをサポートする対話のターンに遡ります。カバレッジの相互検証のために、MiniMax ロールプレイ ベンチマークからリリースされた M2 フリー ダイアログ トランスクリプトを、同じロール由来のチェックリストと照合して監査します。公開されたフリー チャットのトランスクリプトは、主要な役割プロファイル ポイントの 73.74% しかカバーしていませんが、TRACE Bench はより少ないターンで 99.91% のカバー率に達します。堅牢性の実験では、実行を繰り返したり、ユーザー エージェントを置き換えたりしてもランキングが安定していることが示されています。 TRACE Bench は、26 モデルにわたる全体的なランキングを、機能の内訳とチェックリスト トレースとともにレポートします。また、クローズドループ ベンチマーク エボリューションもサポートしており、失敗したトレースで効果的であることが証明された検証手法を抽出して、後の評価で観察された故障モードをより確実に抽出して調べることができます。
原文 (English)
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating checklist states from model responses. Scores therefore trace back to checklist items and supporting dialogue turns rather than a black-box holistic impression. For coverage cross-validation, we audit released M2 free-dialogue transcripts from the MiniMax Role-play Benchmark against the same role-derived checklist. The released free-chat transcripts cover only 73.74% of key role-profile points, whereas TRACE Bench reaches 99.91% coverage in fewer turns. Robustness experiments show stable rankings under repeated runs and User Agent replacement. Across 26 models, TRACE Bench reports overall rankings together with capability breakdowns and checklist traces. It also supports Closed-Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so later evaluations can more reliably elicit and examine observed failure modes.
メモリ使用率を最適化するための強化学習ベースの DBMS バッファ プール自動チューニング
データベース管理システム (DBMS) インスタンスを管理するには、データベース管理者 (DBA) がサービス レベル アグリーメント (SLA) の観点からリソース使用量とパフォーマンスのバランスをとる必要があり、多くの場合、メモリを無駄にする RAM の過剰割り当てが発生します。 SLA 準拠を確保しながら不必要なメモリ割り当てを最小限に抑えるオンライン RL ベースのバッファ調整システムである MicroTune を紹介します。最も効果的な RL コアを特定するために、さまざまなベンチマーク ワークロードの下で複数のアルゴリズムを評価し、外部メトリクス (レイテンシ、スループット) と内部 DBMS メトリクス (ステータス変数とパフォーマンス統計) の両方の広範なトレースで MicroTune をトレーニングします。実験結果は、MicroTune がバッファ サイズをワークロードの変動に動的に適応させ、SLA 違反を減らして大幅なメモリ節約を達成することでベースラインを上回るパフォーマンスを示していることを示しています。これらの発見は、DBMS 環境における適応型リソース管理に対する強化学習の有望性を強調しています。
原文 (English)
Reinforcement Learning based DBMS Buffer Pool Auto-Tuning for Optimal Memory Utilization
Administering Database Management Systems (DBMS) instances requires Database Administrators (DBA) to balance performance in terms of Service Level Agreement (SLA) against resource usage, often prompting RAM over-allocation that wastes memory. We introduce MicroTune, an online RL-based buffer adjustment system that minimizes unnecessary memory allocation while ensuring SLA compliance. To identify the most effective RL core, we evaluate multiple algorithms under diverse benchmark workloads, training MicroTune on extensive traces of both external metrics (latency, throughput) and internal DBMS metrics (status variables and performance statistics). Experimental results demonstrate that MicroTune dynamically adapts buffer sizes to workload fluctuations, outperforming baselines by achieving significant memory savings with fewer SLA violations. These findings underscore the promise of reinforcement learning for adaptive resource management in DBMS environments.
圧縮での損失: コンテキスト圧縮でのサイド制約の損失の評価
コンテキスト ウィンドウに負荷がかかると、LLM システムは以前のコンテキストを圧縮して、進行中のタスクを継続します。 「確認するまで電子メールを削除しないでください」など、ユーザー発行の命令のクラスであるセッション制約 (SC) を特定します。これは、セッションの残りの部分で LLM の動作を制限することを目的としていますが、圧縮中に警告なしに削除されます。この損失を定量化するために、マルチターン チャット、エージェントの軌道、長期調査という 3 つの長いコンテキスト シナリオにわたってコンパクターを評価する評価スイートである COMPINT を導入します。現在のコンパクターは、注入された SC の平均 17% のみを保持しており、そのほとんどは、コンパクションなしで同じタスクを実行するよりもパフォーマンスが悪くなります。保持率は、コンパクター、プロンプト、コンテキストの長さ、SC フレージング、および注入位置によって大幅に変化し、損失が単一の設定に関連付けられているのではなく、体系的に発生していることを示しています。私たちは、プラグ アンド プレイ モジュールとしてコンパクターと並行して実行する SC 対応エクストラクターを提案し、コンパクターや LLM を変更することなく 3 つのシナリオすべてで 90% 以上の保持率を達成します。 COMPINT 評価スイートと付随する実装は、https://github.com/ZhiqiEliWang/compaction-integrity で入手できます。
原文 (English)
Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction
When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user-issued instructions, Session Constraints (SCs), such as "do not delete any emails until I confirm," that are meant to constrain LLM's behavior for the remainder of a session but are silently dropped during compaction. To quantify this loss, we introduce COMPINT, an evaluation suite that evaluates compactors across three long-context scenarios: multi-turn chat, agentic trajectory, and long-horizon research. Current compactors retain only 17% of injected SCs on average, and most perform worse than running the same task without compaction. Retention varies sharply with compactor, prompt, context length, SC phrasing, and injection location, showing that the loss is systematic rather than tied to any single setting. We propose an SC-aware extractor that runs alongside the compactor as a plug-and-play module, achieving over 90% retention across all three scenarios without modifying the compactor or LLM. The COMPINT evaluation suite and accompanying implementation are available at https://github.com/ZhiqiEliWang/compaction-integrity.
拡散から圧縮: 可逆圧縮のための拡散 LM の活用
私たちは、プレーン テキスト、ソース コード、XML などの構造化フォーマットを含むデジタル テキスト データの収集と保存の急速な増加と、ニューラル言語モデル ベースの圧縮の最近の進歩によって動機付けられた可逆テキスト圧縮の問題を研究しています。特に、最近の LLM ベースのアプローチは、シンボルランキング パイプライン上に構築されているか、統計圧縮プログラムと組み合わせているかに関係なく、テキストやコードに対する zstd、gzip、または bzip などの汎用圧縮プログラムよりも大幅に優れた圧縮率を実証しています。ただし、これらのニューラル アプローチには深刻なスループット制限があり、まだ実用的には使用できません。可逆ニューラル テキスト圧縮のコンテキストで初めて、自己回帰 LLM ベースのアプローチに代わる推論パラダイムとして拡散言語モデル (DLM) を導入します。私たちは、同じ圧縮フレームワーク内で自己回帰 LLM を DLM に置き換えることで、ステップごとに 1 シンボルの制限によって引き起こされるスループットのボトルネックを克服できると主張します。ただし、これらの改善を達成するには、DLM を可逆圧縮に適用することによってもたらされるアルゴリズムの課題に対処する必要があります。このアーキテクチャでは、各順方向パスでエンコードされるシンボルの数と位置を独立して決定できます。私たちはこれらの課題を解決するための効率的かつ効果的な戦略を設計し、確立されたテキストベンチマークである enwik8 上の LLM ベースおよび汎用コンプレッサーに対して実験的に評価します。私たちの結果は、新しく提案された DLM ベースのフレームワークが可逆テキスト圧縮における最先端の技術を進歩させることを示しています。さらに、DLM はまだ比較的若いパラダイムであるため、ますます高機能で効率的なモデルへの最近の進歩は、さらなる改善の余地がかなりあることを示唆しています。
原文 (English)
Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression
We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compression. In particular, recent LLM-based approaches, whether built on symbol-ranking pipelines or paired with a statistical compressor, have demonstrated compression ratios significantly superior to general-purpose compressors such as zstd, gzip, or bzip on text and code. However, these neural approaches suffer from severe throughput limitations, making them not yet practically usable. For the first time in the context of lossless neural text compression, we introduce Diffusion Language Models (DLMs) as an alternative inference paradigm to autoregressive LLM-based approaches. We argue that replacing autoregressive LLMs with DLMs within the same compression framework could overcome the throughput bottleneck caused by their one-symbol-per-step limitation. However, achieving these improvements requires addressing algorithmic challenges introduced by applying DLMs to lossless compression, where the architecture allows the number and positions of symbols encoded at each forward pass to be decided independently. We design efficient and effective strategies to solve these challenges and evaluate them experimentally against LLM-based and general-purpose compressors on enwik8, a well-established textual benchmark. Our results show that the newly proposed DLM-based framework advances the state of the art in lossless text compression. Moreover, as DLMs are still a relatively young paradigm, recent advances toward increasingly capable and efficient models suggest substantial room for further improvements.
AI の公平性を考慮した変数の選択
EU AI 法などの最近の規制要求により、AI システムの公平性がより重要になっています。従来のアプローチでは、哲学的倫理や社会意識が考慮されていないことがよくあります。特に、変数選択プロセスは暗黙のバイアスを導入し、異なるサブグループ間の公平性に影響を与える可能性があります。 AI の公平性を評価する数学的アプローチについて説明し、数学的方法論を倫理的考慮事項や規制要件と調整します。私たちの目的は、より広範な倫理的および社会的背景を理解することの重要性を強調し、公平性に対処するための学際的な協力を提唱することです。私たちのアプローチでは、より詳細な公平性評価を可能にし、暗黙のバイアスを軽減するために、関連する可能性のあるすべての変数を維持することに重点を置いています。この調査結果は、敏感な変数または重要な変数を除外すると、サブグループ間の公平性が損なわれる可能性があることを示唆しています。対照的に、関連するすべての変数を保持すると、暗黙的なバイアスが軽減される可能性があります。したがって、学際的なアプローチにより、倫理的な意味と規制基準への準拠についてより深い洞察が得られる可能性があります。数学的アプローチと倫理的および社会的意識を統合することで、より公平な結果と責任ある AI の導入を提案します。この研究は、信頼できる公正な AI システムの促進を目指す欧州連合の AI 法の目的に沿って、AI システムの公平性に効果的に取り組むための学際的な協力の必要性を強調しています。
原文 (English)
Variable Selection in the Context of AI Fairness
Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. Traditional approaches often do not take into account philosophical ethics and social awareness. Variable selection processes, in particular, can introduce implicit bias, affecting equity across different subgroups. We discuss a mathematical approach that evaluates fairness in AI, aligning mathematical methodologies with ethical considerations and regulatory requirements. Our aim is to advocate for interdisciplinary collaboration to address fairness, emphasizing the importance of understanding broader ethical and societal contexts. Our approach emphasizes maintaining all potentially relevant variables to allow for more granular fairness assessments and to reduce implicit bias. The findings suggest that the exclusion of sensitive or critical variables may compromise equity between subgroups. In contrast, retaining all relevant variables could reduce implicit bias. Thus, the interdisciplinary approach could provide deeper insight into the ethical implications and compliance with regulatory standards. By integrating a mathematical approach with ethical and social awareness, we suggest more equitable outcomes and responsible AI deployment. This work underscores the necessity of interdisciplinary collaboration in effectively addressing fairness in AI systems aligned with the objectives of the European Union's AI Act, which seeks to promote trustworthy and fair AI systems.
幼稚園から高校までの教育における AI 個別指導の質を向上させる方法論
現在、多くの AI 家庭教師が大規模言語モデル (LLM) を活用しています。 LLM が不透明なブラックボックスであることを考えると、あらゆる変更の影響を測定するための堅牢な評価とライブ実験が不可欠です。当社は Khanmigo (カーン アカデミー、2023 年) を立ち上げ、幼稚園から高校までを対象とした AI を活用した個別指導の先駆者となりました。 AI 個別指導の品質と生徒の関与を測定するために使用する指標と、私たちが実行したさまざまな実験について説明します。モデル、プロンプト、パーソナライゼーション、エージェントなどの指標を変更した変更点を強調します。
原文 (English)
Methodologies for Improving the Quality of AI Tutoring in K-12 Education
Many AI tutors leverage large language models (LLMs) today. Given that LLMs are opaque black boxes, robust evaluation and live experimentation to measure the impact of every change are essential. We pioneered AI-powered tutoring for K-12 with the launch of Khanmigo (Khan Academy, 2023). We describe the metrics we use to measure AI tutoring quality and student engagement as well as various experiments we have run. We highlight the changes that have moved our metrics, including models, prompting, personalization and agents.
Agent Safety Should Be a Runtime Contract
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue th…
Every pooling rule has its world: matching probability combination rules to situations and stakes
Systems often need to combine two numerical assessments of the same yes/no question. The appropriate formula depends on what the numbers re…
Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification
Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbala…
分散型 CNC ツールの摩耗予測のためのフェデレーテッド ラーニング
工具の摩耗予測は CNC 加工における重要なタスクであり、工具の状態を正確に監視することで製品の品質とプロセスの信頼性がサポートされます。機械学習手法はこのタスクの可能性を示していますが、産業環境での使用は、加工データの分散特性と、機械、現場、または組織間のデータ共有の制限によって制限されます。フェデレーテッド ラーニングは、生の運用データを転送せずに協調的なモデル トレーニングを可能にすることで、この設定に適したフレームワークを提供します。この論文では、CNC 工具の摩耗予測のための連合学習について調査します。ツールの軌跡はシミュレートされたクライアント全体に分散され、フェデレーテッド ラーニング シナリオを表します。フェデレーション モデルは、一元化された参照およびローカル クライアント ベースラインと比較されます。結果は、フェデレーション ラーニングが集中型学習に近いパフォーマンスを達成し、ローカル クライアント モデルよりも大幅に向上していることを示しています。これらの発見は、フェデレーション ラーニングが分散型 CNC 製造環境における協調的な工具摩耗予測をサポートできることを示しています。
原文 (English)
Federated Learning for Distributed CNC Tool Wear Prediction
Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and process reliability. Machine learning methods have shown potential for this task, but their use in industrial environments is limited by the distributed nature of machining data and by restrictions on data sharing between machines, sites, or organizations. Federated learning offers a suitable framework for this setting by enabling collaborative model training without transferring raw operational data. This paper investigates federated learning for CNC tool wear prediction. Tool trajectories are distributed across simulated clients to represent a federated learning scenario. The federated models are compared against centralized references and local client baselines. Results show that federated learning achieves performance close to centralized learning and improves significantly over local client models. These findings indicate that federated learning can support collaborative tool wear prediction in distributed CNC manufacturing environments.
物理学に基づいた暗黙的な神経表現による心筋灌流 MRI 定量化の向上
心臓磁気共鳴 (CMR) からの心筋灌流の定量化は、トレーサー動態モデルを動的造影 MR データに適合させることで実現できます。ただし、観測データを組織内の造影剤の進化を記述するマルチコンパートメント交換モデルに適合させて灌流パラメータを推定することは、ノイズや取得変動の影響を受けやすい困難な逆問題です。これまで、物理情報に基づくニューラル ネットワーク (PINN) が、従来の非線形最小二乗フィッティング法の代替として提案され、定量的灌流 CMR に有望な結果が得られました。この研究では、以前に提案された PINN フレームワークを時空間暗黙的ニューラル表現 (INR) で拡張して、MR 信号を連続時空間関数として表現し、PINN モデルの精度、滑らかさ、物理的一貫性を向上させます。現実的にシミュレートされた CMR データセットでは、INR を使用した私たちの提案した PINN は、以前に確立された方法と比べて堅牢性とパラメーター推定精度が向上していることを示しています。コードは https://github.com/q-cardIA/pinn-inr で入手できます。
原文 (English)
Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Quantification
Quantifying myocardial perfusion from cardiac magnetic resonance (CMR) can be achieved by fitting tracer-kinetic models to the dynamic contrast-enhanced MR data. However, fitting the observed data with multi-compartment exchange models, which describe the evolution of the contrast agent in the tissue, to estimate perfusion parameters is a challenging inverse problem that is sensitive to noise and acquisition variability. Previously, physics-informed neural networks (PINNs) have been proposed as an alternative to conventional non-linear least squares fitting methods with promising results for quantitative perfusion CMR. In this work, we extend the previously proposed PINN framework with spatiotemporal implicit neural representations (INRs) to represent the MR signal as a continuous spatiotemporal function and to improve the accuracy, smoothness, and physical consistency of the PINN model. In realistic simulated CMR datasets, our proposed PINN with INRs demonstrates improved robustness and parameter estimation accuracy over the previously established methods. The code is available at https://github.com/q-cardIA/pinn-inr.
化学的に意味のあるテキスト化により、大規模な言語モデルによる金属有機フレームワークの説明可能な検証が可能になります
計算に対応した金属有機フレームワーク (MOF) データベースはハイスループット スクリーニングに不可欠ですが、報告されている結晶構造の多くは化学的に不合理または無秩序なままであり、シミュレーションの忠実度が損なわれています。既存の検証アプローチでは、計算に対応していない構造を特定できますが、多くの場合、ヒューリスティック ルールやライセンス要件に依存したり、解釈可能性が制限されたりします。ここでは、結晶学的情報が化学的に意味のあるテキストに変換されるときに、大規模言語モデル (LLM) が MOF 構造の解釈可能な検証ツールとして機能できることを示します。 9 つの記述子のベンチマークを行うことで、LLM ベースの検証が成功するかどうかは、構造情報の量だけではなく、局所的な調整、フレームワークの接続性、および化学的コンテキストが言語的に学習可能な表現に編成されているかどうかに依存することがわかりました。特殊な記述子 (mof2text) を使用して微調整された LLM は、不合理な MOF を識別する際にグラフベースのモデルに匹敵するパフォーマンスを実現します。重要なのは、これらのモデルは、異常な結合、接続性、充電状態などの可能性のあるエラー原因の診断根拠を生成することにより、ブラック ボックス分類を超えていることと、アノテーション付きデータセットのエラー カテゴリ予測を生成することです。この研究は、LLM を汎用テキスト モデルから MOF データベースをキュレーションするための実用的で説明可能なツールに変換する重要なステップとして、化学情報に基づいたテキスト化を確立します。
原文 (English)
Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models
Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity. Existing validation approaches can identify non-computation-ready structures, but they often rely on heuristic rules, license requirement, or offer limited interpretability. Here, we show that large language models (LLMs) can serve as interpretable validators of MOF structures when crystallographic information is transformed into chemically meaningful text. By benchmarking nine descriptors, we find that successful LLM-based validation depends not on the amount of structural information alone, but on whether local coordination, framework connectivity, and chemical context are organized into a linguistically learnable representation. Fine-tuned LLMs using specialized descriptors (mof2text) achieve performance comparable to graph-based models in identifying unreasonable MOFs. Importantly, these models extend beyond black-box classification by generating diagnostic rationales for likely error sources, including abnormal bonding, connectivity, and charge states, as well as error-category predictions for annotated datasets. This work establishes chemically informed textualization as the key step that transforms LLMs from generic text models into practical and explainable tools for curating MOF databases.
SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation
Despite the practical relevance of sparse decision-based black-box threats, they have received limited attention in semantic segmentation.…
CLEAR: ロングテール分類のための構造化サンプリングを使用したクラスごとのエキスパート集約
不均衡なデータでトレーニングされたモデルは、頻度が高く過小評価されているクラス間で信頼性が不均一であるため、ロングテール分類では信頼性の課題が生じます。既存の方法は、再バランス、調整、表現学習、または複数の専門家モデリングを通じて不均衡に対処しますが、各クラスでどの専門家を信頼すべきかを推定することはほとんどありません。この論文では、ロングテール分類のためのモジュール式アンサンブル フレームワークである CLEAR (Class-wise reLiability-aware Expert Aggregation for long-tailed Recognition) を提案します。 CLEAR は、完全なラベル空間を維持しながら、しきい値ベースの構造化サンプリングを通じて多様なエキスパートを生成し、次に、平滑化されたクラス単位の精度定式化を使用して、各エキスパートのクラス単位の信頼スコアを推定します。推論中、エキスパートの予測はクラスごとに一般化されたエキスパートの成果の集約を通じて結合され、クラスごとに異なるエキスパートを強調できるようになります。複数のバックボーンにわたる CIFAR-100-LT、ImageNet-LT、および Places-LT の実験では、CLEAR が競争力のある全体的な精度と、特に強力な数ショット パフォーマンスを達成していることが示されています。これらの結果は、ロングテールアンサンブル学習の有用な設計原則として、クラスごとのエキスパートの信頼性を裏付けています。
原文 (English)
CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification
Long-tailed classification poses a reliability challenge because models trained on imbalanced data are unevenly reliable across frequent and underrepresented classes. While existing methods address imbalance through re-balancing, adjustment, representation learning, or multi-expert modeling, they rarely estimate which expert should be trusted for each class. This paper proposes CLEAR (Class-wise reLiability-aware Expert Aggregation for long-tailed Recognition), a modular ensemble framework for long-tailed classification. CLEAR generates diverse experts through threshold-based structured sampling while preserving the full label space, then estimates a class-wise trust score for each expert using a smoothed class-wise precision formulation. During inference, expert predictions are combined through class-wise generalized product-of-experts aggregation, allowing different experts to be emphasized for different classes. Experiments on CIFAR-100-LT, ImageNet-LT, and Places-LT across multiple backbones show that CLEAR achieves competitive overall accuracy and particularly strong few-shot performance. These results support class-wise expert reliability as a useful design principle for long-tailed ensemble learning.
LLM エージェントにおけるバックドア除染のダイナミクス
オープンウェイト LLM エージェントは、微調整中にインストールされるバックドアに対して脆弱であり、テスト中にトリガー条件が満たされなかった場合には検出できない可能性があります。防御側が既存のトリガーを知らないと仮定すると、それを直接解除することはできません。除染戦略の 1 つは、既知のバックドア (防御的ポイズニング) をインストールし、元の未知のバックドアが副作用として削除されることを期待して、それを解除することです。ただし、この手順の結果は不確実です。元のバックドアが存続するか、消去されるか、ルート変更されるなどの可能性があります。 AgentDyn での体系的な実験全体で、ツール呼び出しエージェント、トリガー、応答、教師の分離、および微調整メソッドにおけるこれらのダイナミクスを研究するためのフレームワークを紹介します。 115 回の実験を通じて、防御的ポイズニングだけで元のバックドアの約 56% が消去されました。その後の除染により、ほぼすべての生存者が消去に追い込まれ、トリガーの認識と悪意のある実行が行動的に分離可能であることが確認されました。興味深いことに、私たちの実験では、防御的バックドアと同じ一般的なタイプの別のトリガーを使用し、その後にアンラーニングによる汚染除去を行った場合、悪意のあるバックドアが持続しないことがわかりました。最大 4 つのバックドアを同時インストールすると耐性が高まります (約 36% が消去) が、既知の 1 つの同時常駐バックドアを除染すると、52/60 の同時常駐者 (87%) が副次的に排除されます。 J レンズを使用して除染後のモデルの内部を視覚化すると、除染により良性の LLM 応答が回復しますが、元のトリガー認識の痕跡が中間層に残っていることが確認されました。
原文 (English)
Backdoor Decontamination Dynamics in LLM Agents
Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defenders do not know the existing trigger, they cannot unlearn it directly. One decontamination strategy is to install a known backdoor (defensive poisoning) then to unlearn it, hoping that the original unknown backdoor is removed as a side effect. However, this procedure has uncertain outcomes: the original backdoor may persist or be erased or rerouted, among other possibilities. We introduce a framework for studying these dynamics in tool-calling agents, decoupling trigger, response, teacher, and fine-tuning method across systematic experiments on AgentDyn. Across 115 experiments, defensive poisoning alone erases around 56% of original backdoors; subsequent decontamination then drives almost all survivors to erasure, confirming that trigger recognition and malicious execution are behaviorally dissociable. Interestingly, our experiments find that malicious backdoors never persist when using different triggers of the same general type as the defensive backdoor when followed by decontamination via unlearning. Co-installing up to four backdoors increases resistance (around 36% erased), yet decontaminating a single known co-resident backdoor collaterally clears 52/60 co-residents (87%). Upon visualizing postdecontamination model internals using J-lens, we confirm that although the decontamination restores benign LLM responses, traces of original trigger awareness persist at intermediate layers.
テクスチャ分析と深層学習を使用した乳がん辺縁検出のための低倍率蛍光イメージングの臨床的実現可能性
未処理の外科的乳房組織の高解像度画像は、紫外線表面励起 (MUSE) を備えた顕微鏡を使用して取得できます。この技術は、乳がん手術中に切除断端を確認するための有望な方法と考えられています。この研究では、パッチレベルの分類方法を使用して、倍率 4 倍と 10 倍の MUSE 画像を比較しました。ローカル バイナリ パターン (LBP) に基づくテクスチャ分析 (TA) と、ベースのビジョン トランスフォーマー (ViT) モデルを使用したディープ ラーニング (DL) が使用されました。どちらの方法も、どちらの倍率でも同様の性能を達成しました。 DL 法を使用すると、4 倍と 10 倍の倍率の両方で、感度 96.30%、特異性 100%、精度 98.18% を達成しました。 TA 法を使用すると、4x ではより高い特異性 (100% 対 93.33%) が得られ、10x ではより高い感度 (100% 対 93.33%) が得られましたが、精度はどちらも同じ (96.67%) でした。 10 倍の倍率ではパフォーマンスの明らかな向上は観察されませんでした。これらの結果は、4 倍イメージングが 10 倍イメージングと同じ診断精度を達成することを示しています。同時に、4x はより広い視野とより高速な画像キャプチャを提供します。したがって、MUSE システムでは低倍率を効果的に使用して、正確かつ効率的な術中の断端評価を行うことができます。
原文 (English)
Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning
High-resolution images of unprocessed surgical breast tissue can be obtained using microscopy with ultraviolet surface excitation (MUSE). This technique is considered a promising method for checking surgical margins during breast cancer surgery. In this study, MUSE images at 4x and 10x magnifications were compared using patch-level classification methods. Texture analysis (TA) based on local binary patterns (LBP) and deep learning (DL) with a base Vision Transformer (ViT) model were used. Both methods achieved similar performance at both magnifications. Using DL method, both 4x and 10x magnifications achieved 96.30% sensitivity, 100% specificity and 98.18% accuracy. Using TA method, 4x achieved better specificity (100% vs 93.33%) and 10x yielded higher sensitivity (100% vs 93.33%), but both had the same accuracy (96.67%). No clear improvement in performance was observed with 10x magnification. These results show that 4x imaging achieves the same diagnostic accuracy as 10x imaging. At the same time, 4x offers a larger field of view and faster image capture. Therefore, lower magnification can be effectively used in MUSE systems for accurate and efficient intraoperative margin assessment.
意思決定リソースとしての端末対称性: いつでも検証済みの建設のための状態に応じた改良
多くの連続した建設タスクは、完了時に正確な対称性を示しますが、その実行は指示され、履歴に依存したままになります。私たちは、終端対称性の意思決定リソースのビューを開発します。プロセス証拠は方向性を提供し、同等の結果全体にわたって構造化する終端対応トランスポート、実現状態証拠は遷移後の現在の意思決定の関連性を洗練し、固定検証者は実行を証明します。この分解により、transport-refine-certify が生成されます。 \method{} は、エピソード固定のトランスポート プロセス構造、その状態制限されたプロセス ランク、受け入れられた遷移後に更新される状態依存の残差ランク、および上位 $k$ セットが正確に 2 つの提案プレフィックスの和集合である順序ランクの一致を使用して原理をインスタンス化します。このミートは、プレフィックス カバレッジの下で完了保証を提供し、対応するプレフィックス情報モデルの下で厳しい最悪の場合の検証クエリ制限を達成します。 2 つの状態の構築により、遷移後の動的と静的な分離が厳密に予測されます。 CAD アセンブリ、ミニプログラム、および正確な充填パッキング全体にわたって、ステートごとのリフレッシュにより AUC がそれぞれ最大 $6.77$、$21.75$、および $8.68$ ポイント向上します。公式 GRN OOD シーンからの 1,135 のターゲット削除エピソードで、\method{} は、比較した GRN と CDGS スタイルのプランナーの中で、3 つのスケールすべてで上限付き検証者の平均コストが最も低くなりました。状態ごとの信号は、アグリゲーション組織とスケジューラ組織間でも転送されます。これにより、末端対称性は、有向構築のための再利用可能な決定リソースになります。
原文 (English)
Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction
Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent. We develop a decision-resource view of terminal symmetry: process evidence supplies directionality, terminal correspondence transports that structure across equivalent outcomes, realized-state evidence refines its current decision relevance after transitions, and a fixed verifier certifies execution. This decomposition yields transport--refine--certify. \method{} instantiates the principle with an episode-fixed transported process structure, its state-restricted process rank, a state-dependent residual rank refreshed after accepted transitions, and an ordinal rank meet whose top-$k$ set is exactly the union of the two proposal prefixes. The meet provides a completion guarantee under prefix coverage and attains the tight worst-case verifier-query bound under the corresponding prefix information model; a two-state construction predicts a strict post-transition dynamic--static separation. Across CAD assembly, Mini-Programs, and exact-fill packing, statewise refresh improves anytime AUC by up to $6.77$, $21.75$, and $8.68$ points, respectively. On 1,135 target-removal episodes from the official GRN OOD scenes, \method{} attains the lowest mean capped verifier cost at all three scales among the compared GRN and CDGS-style planners. The statewise signal also transfers across aggregation and scheduler organizations. Terminal symmetry thereby becomes a reusable decision resource for directed construction.
社会性: 人間と AI の相互作用のための関係プロセス フレームワーク
人間と AI の研究では、個々の能力、総合的なパフォーマンス、または最終的な成果物を評価することがよくありますが、これらのアプローチでは、一方の反応がどのようにして他方の次の貢献が形成される条件の一部になるかが保存されません。この記事では、社会的双対性、つまり、区別可能な 2 つの当事者の間の連続的で互恵的で歴史を担う関係プロセスについて紹介します。この関係プロセスでは、一方の当事者からの反応が、他方の当事者のその後の貢献、判断、決定、または行動が形成される観察可能な条件の一部になります。人間と AI のダイアッド向けに指定されたこの構成では、移動、確認された社会的エピソード、リンクされた経路、およびより広範な相互作用コンテナーなどの入れ子になったユニットが使用されます。最小限のエピソード A1-B1-A2 では、対応偶発性と復帰偶発性の証拠が必要です。候補エピソードは、反応の方向性と実質的な貢献の再形成を二次的にコーディングする前に、確認済み、非社会的、または不確定として分類されます。 3 つの命題は、履歴条件の形成、経路の分岐、エンドポイントに相当する経路間の堅牢性の違いに対処します。凍結された運用プロトコルは、別々に実行された 2 つのモデルベースの評価シリーズを通じて、これまで見たことのない 3 つの自然な人間の AI 記録に基づいて調整されました。移動と候補の再構成は 2 つのケースで正確に収束しましたが、3 番目のケースではローカルのマルチモーダル単位化の決定が 1 つ異なりました。残りの意見の相違は、帰還と不測の事態の境界に集中していた。したがって、社会二重性は、エンドポイント中心の分析では回復できない経路情報を保存しながら、人間と AI の貢献が相互作用を通じてどのように形成されるかを分析するための、限定的で経験的に扱いやすいプロセス構造を提供します。
原文 (English)
Socioduality: A Relational Process Framework for Human-AI Interaction
Human-AI research often evaluates individual capabilities, combined performance, or final outputs, but these approaches do not preserve how one party's response becomes part of the conditions under which the other party's next contribution is formed. This article introduces socioduality, a sequential, reciprocal, and history-carrying relational process between two distinguishable parties in which a response from one party becomes part of the observable conditions under which the other party's subsequent contribution, judgement, decision, or action is formed. Specified for human-AI dyads, the construct uses nested units: moves, confirmed sociodual episodes, linked pathways, and the broader interaction container. A minimum episode A1-B1-A2 requires evidence of response contingency and return contingency; candidate episodes are classified as confirmed, non-sociodual, or indeterminate before secondary coding of response orientation and substantive contribution re-formation. Three propositions address history-conditioned formation, pathway divergence, and robustness differences among endpoint-equivalent pathways. A frozen operational protocol was calibrated on three previously unseen natural human-AI records through two separately executed model-based evaluator series. Move and candidate reconstruction converged exactly in two cases and differed by one local multimodal unitisation decision in the third; remaining disagreement was concentrated at return-contingency boundaries. Socioduality therefore provides a bounded and empirically tractable process construct for analysing how human and AI contributions are formed through interaction while preserving pathway information that endpoint-centred analysis cannot recover.
Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings
This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of…
臨床テキストガイドによる医用画像セグメンテーションのためのデュアルドメイン クロスモーダル デコーディング
臨床テキストはセグメント化する対象を絞り込むことができますが、最近のテキストガイド付きデザインは空間的な位置合わせを強調し、質感や境界を支配する周波数コンテンツを無視しています。我々は、デコード中に言語ガイダンスの 2 つの相補的な形式を統合する、臨床テキストガイドによる肺感染症セグメンテーションのためのデュアルドメイン クロスモーダル デコード (DD-CMD) を提案します。空間ドメインでは、テキスト ガイド付き空間クロス アテンション (TGSA) がマルチスケールのビジュアル トークンをテキスト セマンティクスと調整し、ゲートされた残差融合を通じて特徴を更新します。周波数領域では、スペクトル テキスト適応変調 (STAM) が 2D DCT を適用して学習可能な帯域エネルギー統計を計算し、テキスト条件付き FiLM パラメータを予測して、周波数を意識したデコード用にデコーダ チャネルを再調整します。 DD-CMD は、TGSA と STAM を粗密デコーダ (7x7 ~ 56x56) に組み込み、軽量の 2 段階リファインメント モジュールを使用してフル解像度のマスクを復元します。 QaTa-COV19 と MosMedData+ の実験では、DD-CMD がそれぞれ 91.46% Dice / 84.26% mIoU と 81.95% Dice / 69.42% mIoU を達成し、過去の最も強力なベースラインと比べて平均で +1.96 Dice および +2.67 mIoU の増加が見られたことが示されています。コード: https://github.com/maklachur/DD-CMD。
原文 (English)
Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation
Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text-Guided Spatial Cross-Attention (TGSA) aligns multi-scale visual tokens with text semantics and updates features through gated residual fusion. In the frequency domain, Spectral-Text Adaptive Modulation (STAM) applies a 2D DCT to compute learnable band-energy statistics and predicts text-conditioned FiLM parameters to recalibrate decoder channels for frequency-aware decoding. DD-CMD embeds TGSA and STAM into a coarse-to-fine decoder (7x7 to 56x56) and restores full-resolution masks using a lightweight two-stage refinement module. Experiments on QaTa-COV19 and MosMedData+ show that DD-CMD achieves 91.46% Dice / 84.26% mIoU and 81.95% Dice / 69.42% mIoU, respectively, with average gains of +1.96 Dice and +2.67 mIoU over the strongest prior baselines. Code: https://github.com/maklachur/DD-CMD.
自己進化するネットワーク検証者
シンボリック ネットワーク検証者は、ルーティング入力と障害の広大な空間全体にわたって正当性を推論できますが、専門家が手動でエンコードしたプロトコルと機能のみが対象です。コントロール プレーンの忠実なモデルを作成して維持することは、難しくて終わりがありません。ネットワークの動作を完全に指定した文書は存在しないためです。ベンダーの実装は RFC から逸脱し、リリースごとに動作が変化します。継続的な維持の負担により、最終的には検証を必要とする多くのネットワークから検証が行われなくなります。私たちは、実際のネットワークの動作を忠実に捉えるためにモデルを自動的に進化させる必要があると主張します。これを実現するために、それを明確に指定する唯一のソースであるルーター ソフトウェア自体を活用します。反例ガイド付きループでは、コーディング エージェントが検証者のシンボリック エンコーディングの拡張を提案し、信頼できるオラクル (エミュレートされたルーターなど) がグラウンド トゥルース ルーティング ステートを提供します。エージェントは、オラクルとの意見の不一致を利用してネットワーク モデルを繰り返し改良します。初期の証拠として、このシステムのプロトタイプは、3,000 行の SMT ベースの検証器に、サポートしていない 3 つの機能、つまり OSPF エリア、BGP ルート リフレクション、EVPN 上の L3VPN を教え込み、オラクルに一致するモデルに自律的に収束し、ベンダー固有の動作も認識しました。モデルの成長を自動化すると、難しい問題は検証システムの作成から体系的なテストへと移行します。私たちは、自動的に進化した検証者を信頼し、活用するための研究課題を提案します。
原文 (English)
Self-evolving network verifiers
Symbolic network verifiers can reason about correctness across vast spaces of routing inputs and failures, but only for the protocols and features an expert has encoded by hand. Creating and maintaining a faithful model of the control plane is both difficult and never-ending, since no written source specifies perfectly what a network does: vendor implementations deviate from the RFCs, and behaviour shifts with releases. The burden of constant upkeep ultimately keeps verification out of many networks that need it. We argue that the model should instead evolve automatically to faithfully capture the actual network behaviour. To achieve that, we leverage the only source that specifies it unambiguously: the router software itself. In a counterexample-guided loop, a coding agent proposes extensions to the verifier's symbolic encoding, while a trusted oracle (e.g., emulated routers) supplies the ground-truth routing state. The agent iteratively refines the network model using each disagreement with the oracle. As early evidence, a prototype of this system taught a 3,000-line SMT-based verifier three features it did not support: OSPF areas, BGP route reflection, and L3VPN over EVPN, converging autonomously on models that match the oracle, even noticing vendor-specific behaviour. Automating model growth shifts the hard problem from writing verification systems to systematically testing them; we propose a research agenda for trusting and harnessing automatically evolved verifiers.
Governing Agentic AI in FinTech
Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and…
Dynamics Models for Offline Hyperparameter Selection in Real-World RL
A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unav…
Gaze Target Estimation Anywhere with Concepts
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, mu…
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
Long-running agents periodically compact their context, replacing the transcript with a model-generated summary.Recent work shows that drop…
TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment…
Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards
Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never…
TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation
Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image rend…
PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition
PAC-Bayes theory provides generalization guarantees by controlling the Kullback--Leibler (KL) divergence between posterior and prior distri…
The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark
AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most…
HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging
Task vectors enable model merging without joint retraining. In practice, the subset of task vectors to be merged may vary, but many existin…
Strengthening Full Justified Representation: Efficient Verification and Computation
Full justified representation (FJR) is among the strongest known satisfiable proportionality axioms for approval-based committee elections.…
Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task
Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology…
Let it Cook: Learning to Wait in Sequential Decision Making
In sequential decision making, an agent typically observes its environment and acts at every timestep. However, such active participation m…
Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation
Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and ma…
Keep the Future, Drop the Rollout: RIFT for World Action Models
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask w…
Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks
In recent research on the Digital Twin-based Vehicular Ad hoc Network(DT-VANET), Federated Learning (FL) has shown its ability to provide d…
Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment
Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boun…
A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era
The convergence of artificial intelligence (AI), Industrial Internet of Things, cyber-physical systems, and advanced robotics is reshaping…
Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents
Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached t…
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although si…
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs
Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language…
RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation
Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluation across complex road netwo…
A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases
Identifying dengue virus-infected mosquitoes from control mosquitoes is a major challenge in analyzing mosquito locomotion behavior due to…
Dion3: Full-Stack Orthogonal Updates
The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. When weights are sharded,…
FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting
Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods r…
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreeme…
Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones
The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and s…
Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identify authoritative state. Wi…
Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting
Motion-centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation. However,…
Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models
We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization p…
Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning
Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically afte…
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their…
Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings.…
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models o…
GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training inst…
Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation
This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to r…
APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference
Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small s…
The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways…
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolat…
Consolidator: Learning Persistent Routed Memory Across Context Boundaries
Copying short-term memory (STM) into a slower store can preserve state across a context boundary, but persistence alone does not ensure tha…
Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing
Dynamic Time Warping (DTW)-based Nearest-Neighbor (NN) classifiers are effective for time-series classification but are vulnerable to misla…
High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions
Multi-objective Bayesian optimization (MOBO) is effective in identifying the Pareto fronts for expensive black-box problems. However, most…
When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use
The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model sel…
Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation
Proprietary text-to-image diffusion models are increasingly distributed as hosted services and downloadable checkpoints, making their intel…
A 12-CNOT Double Qubit Excitation Gate
Effective implementation of high-level quantum gates is essential for practical quantum computing. To the best of our knowledge, we present…
Locating and Controlling Implicit Personalization in Large Language Models
Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic i…
Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System
Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses di…
G0.5: One Autoregressive Stream for Robot Reasoning and Action
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action exper…
JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. H…
MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning
Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradien…
Instruction Alignment for Binary Code Representation Learning
Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn fu…
GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation
Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound…
TELLME: Test-Enhanced Learning for Language Model Enrichment
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consisten…
Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of co…
Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly Detection, Attack Identification, and Hardware Monitoring
Methods to increase the resilience of systems to cyber-attacks become increasingly important. Control-flow monitoring provides a principled…
How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment
State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it ari…
User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling
Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise wit…
LookBack: Where and How to Score LVLM Responses via Visual Reference Usage
Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding…
Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra
Data-driven inverse design enables efficient generation of nanophotonic structures with prescribed optical responses, but spectrum-to-geome…
CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation
Quantum generative adversarial networks (QGANs) have attracted increasing attention for image generation using parameterized quantum circui…
DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation
Prompting-based (\textit{i}.\textit{e}., non-fine-tuning) Text-to-SQL methods, where underlying large language model parameters are not cha…
Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework
Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual comp…
Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluati…
Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians
A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical…
Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option…
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement
Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial…
Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference
Structured potential outcomes such as microscopy images may be recorded after an unknown, unit-specific transformation. If that transformat…
LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A…
HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs
Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autore…
Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh
Rapid urbanization in Dhaka District, Bangladesh has triggered substantial alterations in land use and environmental conditions, necessitat…
RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
In modern AI frameworks, GPU kernels are key to overall system performance. Combining usability, portability, and near-handwritten CUDA per…
Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches
Consumer reviews play an important role in shaping brand perception and business strategies, particularly in service-driven industries such…
From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices
Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligne…
Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision
Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-v…
LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic c…
How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging
Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained model…
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations
Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in s…
Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks…
Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models
Learning world models from offline trajectories enables agents to accomplish different tasks through planning. Object-centric (OC) represen…
Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion
Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and…
Confidence Calibration of Deep Learning Systems
In high-stakes applications, reliable confidence estimates are as important as the predictions themselves. Confidence calibration ensures t…
No One to Blame: A Framework of Constitutive AI Unaccountability
The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominant…
QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving
Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-I…
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emi…
Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactio…
Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning
We study nonnegative submodular maximization subject to a general matroid when the offline algorithm is given an arbitrary controlled value…
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchm…
Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers
Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential pu…
HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation
Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spa…
Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation
With the increasing complexity of cyber assaults in cloud environments, adaptable security solutions are needed that can support real-time…
HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks
Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions.…
M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation
Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to ex…
NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation
Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist rec…
Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment
Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving beh…
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction hist…
SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward
Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning. Recent reinforcement learning (RL) metho…
Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification
Hyperspectral imaging (HSI) offers nondestructive assessment of fish freshness by detecting biochemical alterations across spectral bands.…
Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images
Non-destructive food quality assessment has increasingly benefited from hyperspectral imaging (HSI), which captures spectral signatures lin…
How Organizations Use AI: Evidence from ChatGPT
We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifica…
HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression
Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-distortion performance…
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that…
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We…
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Pri…
Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning.…
A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery
Background: Accurate segmentation of the Left Anterior Descending (LAD) artery in 3D free-breathing, non-contrast CT is critical for cardia…
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-reso…
Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained cont…
Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpo…
Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop…
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through t…
DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determi…
Causal Agent based on Large Language Model
The large language model (LLM) has achieved significant success across various domains. However, the inherent complexity of causal problems…
On Benchmarking Human-Like Intelligence in Machines
Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-gen…
OpenAg: Democratizing Agricultural Intelligence
Agriculture is undergoing a major transformation driven by artificial intelligence (AI), machine learning, and knowledge representation tec…
Deep Fictitious Play-Based Potential Differential Games for Learning Human-Like Interaction at Unsignalized Intersections
Modeling vehicle interactions at unsignalized intersections is a challenging task due to the complexity of the underlying game-theoretic pr…
DREAMS: Density Functional Theory Based Research Engine for Agentic Materials Simulation
Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents…
On the Definition of Intelligence
To engineer AGI, we should first capture the essence of intelligence in a species-agnostic form that can be evaluated, while being sufficie…
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 dataset…
Behavior and Representation in Open-Weight Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
Recent advances in Large Language Models (LLMs) open new perspectives for automation in optimization, yet little is known about whether the…
Credo: Declarative Control of LLM Pipelines via Beliefs and Policies
Agentic AI systems are becoming commonplace in domains that require long-lived, stateful decision-making in continuously evolving condition…
Towards Human Motion World Models via Executable Behaviour Representations
Human motion world models should capture motion's intentionality by being executable: adaptable to different actions and capable of assessi…
Tools as Continuous Flow for Evolving Agentic Reasoning
Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing method…
Agentic AI、ネストされた学習、セマンティック キャッシングによる AI の持続可能性による幻覚の軽減
幻覚は、実稼働 LLM システムにとって、特にサポートされていないクレームがチェックされずにステージ全体に伝播する可能性があるマルチエージェント パイプラインにおいて、依然として大きな信頼性の障壁となっています。この論文では、連続メモリ システム (CMS) と意味論的類似性キャッシュを備えた HOPE にインスピレーションを得た入れ子学習アーキテクチャを、217 個の認識論的不確実性プロンプトと 93 個の製造誘導ストレステスト プロンプトを組み合わせた 310 個のプロンプトのハイブリッド ベンチマークに適応させます。オープン フロア プロトコル (OFP) を介して調整された 3 段階のエージェント パイプラインは、FCD (事実の主張密度)、FGR (事実の根拠参照)、FDF (架空の免責事項の頻度)、ECS (明示的なコンテキスト化スコア)、および OSR (観察可能性スコア率) の 5 つの KPI を使用して評価され、研究対象の 5 つの重み付け構成にわたって THS (総幻覚スコア) に集約されます。緩和と可観測性のトレードオフ。 FDF、ECS、OSR、および FGR は緩和シグナルとして差し引かれるため、THS がより負であるほど、緩和が強力であることを示します。 FrontEndAgent は、現実的な幻覚ベースラインを生成する高確率ジェネレータ (温度 = 1.0) として構成され、SecondLevelReviewer と ThirdLevelReviewer はプログレッシブ コレクタとして動作します。この非対称設計により、5 つの重み付け構成全体でエンドツーエンドの THS が -31.3% ~ -35.9% 削減されます。セマンティック キャッシュは、930 回の潜在的な呼び出しで 440 回のキャッシュ ヒット (ヒット率 47.3%) を達成し、LLM 呼び出しを 490 回に減らし、エネルギーと CO2e フットプリントを削減し、マルチステージ レビュー パイプラインを運用規模で実行可能にします。 ExtremeObservability は最もマイナスの最終 THS (-0.0709) を達成しており、可観測性を重視した構成が緩和を損なうのではなく強化していることが確認されています。これらの発見は、メモリ拡張マルチエージェント設計がモデルの再トレーニングなしで事実の信頼性、運用効率、監査可能性を共同で改善できることを示唆しています。
原文 (English)
Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching
This paper describes an approach to hallucination detection and mitigation using a HOPE-inspired Nested Learning architecture with Continuum Memory Systems (CMS) and semantic similarity caching, tested on a hybrid benchmark of 310 prompts (217 epistemic-uncertainty prompts, 93 fabrication-induction stress tests). A three-stage pipeline orchestrated via the Open Floor Protocol is evaluated with five KPIs; four score the response and aggregate into a Total Hallucination Score. The score improves end to end by 6.1% of its attainable range, but 83.5% of that gain is attributable to a single dimension, Explicit Contextualization, while Factual Claim Density, the dimension closest to unsupported content, stays flat; 97.7% of the gain arrives at the first review stage. The fifth indicator, observability, is reported separately, since it registers the presence of an OFP annotation channel rather than a property of the response: it rises 147% at the review stage, the only stage carrying explicit hallucination markers, then falls back at the final stage, which does not propagate the channel. Three annotators independently labelled every final-stage response on the 93 stress prompts: in 10 of 93 cases (10.8%, 95% CI 5.9-18.7) the final answer still presents the invented item as real, at alpha=0.586, below the conventional threshold; Explicit Contextualization tracks these labels monotonically. Re-scoring all 930 outputs with Llama 3.1, Gemma 4, and Qwen 3 as judges confirms the gain and ranks the cross-family judges above the original evaluator against human labels (rho=-0.772 vs -0.477). Semantic caching serves 47.7% of model calls. A nominally multi-dimensional reliability score is thus effectively one-dimensional, and only that dimension has external support.
AXIOM: 検証可能な数学的推論のための信頼優先のニューロシンボリック実行アーキテクチャ
私たちは、自然言語数学的推論のための信頼優先のニューロシンボリック実行アーキテクチャである AXIOM を紹介します。 AXIOM では、言語モデルは厳密に正規化器として機能します。つまり、非公式の問題テキストを、決定論的なコンピューター代数システム (CAS) パイプラインによって消費される狭いスキーマに書き換えます。このパイプラインは、答えを導き出して検証するか、または第一級の出力として棄権します。ルーティングは、問題形状の正規表現、スキーマ固有のプロンプト、および閉じた形式の CAS ハンドラーの間の 1:1:1 の調整に従い、3,100 以上のそのようなルートが出荷され、250 以上の連続した出荷コミットで LOST_CORRECT リグレッションはゼロです。解析可能な信頼性 100.00% で累積正しさ 94.36% (2,592/2,747) の 4 つの MATH カテゴリ (2,747 レコードのベンチマーク全体で確信のある誤答がゼロ)、4 つのドメインすべてがドメインごとの信頼性 100.0% でドメインごとの 70/90/70 の下限を上回っていること、およびレイテンシの中央値に関する経験的結果を報告します。ルールのみのハンドラーで 1 ミリ秒 (lm-eval 算術 20,000 レコード ベンチマークのレコードの 88%)。このアーキテクチャは、パブリック デプロイメントを通じて約 30,000 件の実稼働クエリに対応してきました。私たちが強調する貢献は、最終的な精度の数値ではなく、アーキテクチャが確立する前向きのダイナミクスです。新しいタスクはレジストリを後退させることなく構成されるため、本番環境でログに記録されたすべての棄権は 1 シップ サイクル後の正しい候補となります。このプロパティの背後にある運用規律 (数学テンプレートのバケット化、回帰オラクルとしての LOST_CORRECT スキャン、解析可能優先のオンボーディング、およびファーストクラスの出力としての棄権) は、数学を超えた信頼できる神経記号システムのための移転可能なフレームワークを構成します。
原文 (English)
Moxia: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning
We present Moxia (formerly AXIOM), a trust-first neuro-symbolic architecture for self-explaining mathematical reasoning over natural-language input. Its language model is strictly a canonicalizer: it rewrites informal problem text into a narrow schema consumed by a deterministic Computer-Algebra-System (CAS) pipeline, which derives and verifies the answer or abstains as a first-class output. Routing follows a 1:1:1 alignment of problem-shape regex, schema-specific prompt, and closed-form CAS handler, with 4,783 routes shipped, 71% of which answer without invoking the language model, and zero LOST_CORRECT regressions as a standing release gate. Because the answer is derived rather than generated, so is its explanation: every handler emits a step trace of the computation it performed, rendered as prose by a layer covering all 4,785 task files that cannot narrate a step the handler did not take. Derivations export to Lean 4 as well: 479 task files (10%) emit a theorem from the problem's declared data, 445 accepted by the Lean kernel with Mathlib; that gate covers a fixture corpus, so live output is generated, not machine-checked. We report two numbers and never fuse them. On the full 7-category MATH test split, designed against, Moxia answers 90.2% (4,510/5,000) with one confident-wrong answer (99.98% trust on parseable). On held-out MATH-500, never designed against, it answers 89.2% (446/500) with zero confident-wrong answers. The 1.0 pp gap is the substantive result: a registry that had merely memorized problem shapes would collapse on held-out data, and this one does not. The rule-only path answers the 20,000-record lm-eval arithmetic benchmark at 100%, 1 ms per record. What we emphasize is not an accuracy figure but the forward dynamic: every logged abstain is a candidate correct after one ship cycle, since new tasks compose without regressing the registry.
RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit
Community-conditioned language model adaptation needs choices about data collection, community definition, and evaluation that are currentl…
LiteOdyssey: 解釈可能な希少疾患診断のための軽量推論 AI エージェント
ほとんどの医療 AI システムは、より多くの微調整データ、より多くのエージェント、および/またはより大規模な検索データベースなど、追加の機械を拡張することで改善されます。ただし、希少疾患の診断では、このような拡張により、展開、監査、保守が困難なシステムが生成される可能性があります。私たちは、単一の AI エージェントの推論チェーンを拡張することによって、つまり人間と AI のコラボレーションによって開発された診断ポリシーでエージェントを導き、自由に利用できる生物医学ツールを拡張することによって、最先端の診断パフォーマンスを実現できるかどうかを尋ねました。臨床遺伝学のワークフローを通じて推論言語モデルをガイドする軽量の希少疾患診断フレームワークである LiteOdyssey を紹介します。このフレームワークは、Policy Iteration with Human Feedback (PIHF) を通じて開発され、公共の生物医学ツールへの動的なアクセスを使用します。患者の臨床的特徴のみを提供する 2 つの困難なベンチマークで、LiteOdyssey は最先端のパフォーマンスを達成し、LIRICAL (n = 370) と PhenoPacket Store (n = 873) の合計 1,243 症例を上回る 59.3% の全体的な疾患再現率 @1 を達成しました。どちらのベンチマークも、超希少疾患の割合が高くなります (有病率は 100 万人に 1 人未満、超希少疾患の割合はそれぞれ約 45% と 52.8%)。希少性マッピング パイプラインで原因疾患が Orphanet にマッピングされなかった、より困難な PhenoPacket サブセットでは、LiteOdyssey は 60.7% の再現率 (1) を達成しました。これに対し、ツールを使用しない同じベースライン モデル (GPT-5.4) では 10.7% でした。このパフォーマンスは、微調整、マルチエージェント アンサンブル、または大規模な症例検索データベースを使用せずに達成されました。また、開発中に見られなかった症例、現実世界の希少疾患患者のプライベートコホート、およびより小規模な無重力モデルでも利益が観察されました。 LiteOdyssey は、正確で導入が容易で、医師のレビューがより透明性の高い希少疾患 AI システムへの道を提案します。
原文 (English)
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first in only 35.4% of benchmark cases. Here we show that this expert reasoning can be converted into a scalable AI capability through a governed learning process rather than model training alone. We developed liteOdyssey through Policy Iteration with Human Feedback (PIHF), an in-context policy-learning method adapted from generalized policy iteration in reinforcement learning, in which model failures and expert corrections consolidate into an explicit, clinician-gated policy that turns an off-the-shelf LLM into an agentic diagnostic system. We demonstrated that such a policy improved diagnostic accuracy to match the best published systems at a fraction of their deployment footprint, generalized to unseen diseases, transferred across models, and remained under clinician control. Across 1,243 public benchmark cases spanning 722 rare diseases, liteOdyssey ranked the correct disease first in 59.3% of cases versus 26.5% without the policy, with nearly identical gains on the 1,193 cases and 679 diseases excluded from policy development. Ablations showed that gains exceeded automated prompting improvement or source access alone, and the policy transferred without modification across closed- and open-weight models. In 515 Undiagnosed Diseases Network patients, liteOdyssey again improved accuracy, and blinded physicians rated its differentials more often exact and less often unhelpful. Through PIHF, expert reasoning becomes an LLM capability that experts can inspect, revise, and transfer across models.
逆計画としてのパーソナライゼーション: 構造ノイズ除去によるエージェント的スライド生成のための潜在的な設計意図の学習
スライドのデザインでは、デッキのテーマとページ レイアウトの両方をパーソナライズする必要があります。しかし、現在の AI エージェントベースの手法は、きめ細かいページレベルの設計に苦労しています。事前に指定されたテンプレートやユーザーの詳細な指示のみに依存すると、潜在的なデザイン意図を捉えることができず、ページレベルのスライド パーソナライゼーション (PSP) が未解決のままになります。このギャップを埋めるために、この研究では PSP を逆計画問題として定式化します。使用されている特定の実行ツール (PowerPoint、Beamer など) についての知識を前提とせずに、設計意図を学習することを提案します。ただし、これらのツールの制御を放棄すると、問題はエンドツーエンドの最適化が困難になります。これを克服するために、PSP を近似的に解決するための原則的なフレームワークである SPIRE を提案します。 SPIRE は、クリーンなスライドの視覚構造を意図的に破損することで、破損のノイズを除去するための検証可能なタスクを作成します。これにより、2 人のエージェントが、強化学習 (RL) を通じて実行可能なデザインを協力して改良する方法を学習します。我々は、構造的ノイズ除去がPSPの一貫した代用であること、およびマルチエージェント定式化がRLにおけるポリシー勾配の分散を厳密に低減することの証明を提示する。広範な実験により、SPIRE の優位性が実証されました。
原文 (English)
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or user verbose instructions, they fail to capture latent design intents, leaving Page-level Slide Personalization (PSP) unresolved. To close this gap, this work formulates PSP as an inverse planning problem. We propose to learn a design intent without assuming any knowledge of the specific executing tools (e.g., PowerPoint, Beamer) being used. However, relinquishing control over these tools makes the problem intractable to optimize end-to-end. To overcome this, we propose SPIRE, a principled framework to solve PSP approximately. By intentionally corrupting the visual structures of clean slides, SPIRE creates a verifiable task to denoise the corruption, whereby two agents learn to collaboratively refine executable designs via reinforcement learning (RL). We present a proof that structural denoising is a consistent surrogate for PSP, and that the multi-agent formulation strictly reduces policy gradient variance in RL. Extensive experiments demonstrate the superiority of SPIRE.
自己進化する臨床システムへの道: 医療エージェントの支援から自律への拡大
画像やテキストを共同で解釈し推論する大規模な言語モデルと視覚言語モデルの能力の向上により、医療エージェントが再構築され、タスク固有の予測因子から、臨床環境で知覚、推論、計画、記憶、行動する自律システムへと移行しています。この研究は、既存の文献の能力第一の観点から逸脱し、代わりに臨床展開から始まり、医療エージェントが実際に信頼できるようになる前にどのようなタスク、耐汚染性ベンチマーク、対話型トレーニング環境が必要であるかを問うものです。医療エージェントは、支援型、協力型、および完全自律型の操作に及ぶ 3 レベルの自律分類とともに、部分的な可観測性の下で逐次的な意思決定システムとして形式化されています。この分野は、フレームワークのスケーリング、機能のスケーリング、環境のスケーリングで構成される統一されたスケーリングの柱に沿って編成されています。このフレームワーク内では、臨床環境のスケーリング、つまりツール、データ、および臨床ジムの統合が、PACS、EHR、および FHIR エコシステムで活動するエージェントにとって最も実行可能であるものの、まだ検討されていない方向性として特定されています。パラメータのスケーリングのみではなく、環境との相互作用を通じてエージェントが改善する臨床的自己進化は、自己改善エージェント、エージェントジム、テスト時間のコンピューティングスケーリングから洞察を引き出し、重要な研究フロンティアとしてさらに位置付けられています。放射線学、病理学、眼科、病院のワークフローにわたるアプリケーションが、幻覚、カスケード障害、公平性などの導入上の課題とともに検証されます。この研究では、2025 年から 2026 年の進歩に特に重点を置いた 300 以上の参考文献を統合することにより、実際の臨床現場向けの信頼できる自己改善型医用画像システムに向けたロードマップを提供します。
原文 (English)
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical imaging AI, moving it from task-specific predictors toward autonomous agents that perceive, reason, plan, remember, and act in clinical environments. This survey departs from the capability-first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination-resistant benchmarks, and interactive training environments are required before medical agents can be trusted in practice. Medical agents are formalized as sequential decision-making systems under partial observability, together with a three-level autonomy taxonomy spanning assisted, cooperative, and fully autonomous operation. The field is organized along a unified scaling spine consisting of framework scaling, capability scaling, and environment scaling. Within this framework, clinical environment scaling, the integration of tools, data, and clinical gyms, is identified as the most actionable yet underexplored direction for agents operating in PACS, EHR, and FHIR ecosystems. Clinical self-evolution, where agents improve through interaction with their environments rather than parameter scaling alone, is further positioned as a key research frontier, drawing insights from self-improving agents, agent gyms, and test-time compute scaling. Applications across radiology, pathology, ophthalmology, and hospital workflows are examined together with deployment challenges including hallucination, cascading failures, and fairness. By consolidating more than 300 references, with particular emphasis on advances from 2025 to 2026, this survey provides a roadmap toward trustworthy, self-improving medical imaging systems for real clinical practice.
SportD: VLM は物理的に戦略を立てることができますか?
視覚言語モデルは、視覚的なシーンを解釈できるようになってきていますが、戦略的に効果的な意思決定を行うために情報を使用できるかどうかは依然として不明です。私たちはサッカーでこの問題を調査します。モデルはオンボールの決定の数秒前を観察し、シュートするか特定のチームメイトにパスするかを選択する必要があります。従来の視覚的に理解するタスクとは異なり、サッカーでは、利用可能なすべてのアクションの価値を推定することで、意思決定を定量的に評価できます。 2022 FIFA ワールドカップの 478 件のオンボール判定で構成されるベンチマークである SportD を紹介します。各モデルの選択は、攻撃側チームの得点確率を最も高めるアクションを推定するポゼッション価値モデルに対して評価され、最適なアクションの精度と、最適ではない決定によって失われる価値の両方を測定できるようになります。 3 つのフロンティア VLM では、イベントの 31.4% で最も価値の高いアクションが選択されます (プロ プレーヤーの場合は 38.9%)。すべてのモデルで大幅に大きな後悔が発生します。さらなる分析により、より低い分散とより低い報酬のアクションを系統的に好むことが明らかになりました。VLM は、最適なポリシーや実際のプレーヤーよりもシュート頻度が低く、実質的にプログレッシブなパスを選択しません。また、モデルは、プレイヤーの特定のアクションが最適ではない場合でも偶然を超えて再現し、反事実的な代替案の一貫した評価ではなく、よく知られたプレイ パターンの部分的な模倣を示唆しています。 SportD は、VLM における物理的な戦略的推論を測定するための、価値に基づいたテストベッドを提供します。
原文 (English)
SportD: How do VLMs physically strategize?
Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1415 decision scenarios across professional men's and women's soccer games, where a VLM observes the seconds before a decision and chooses the next action. Models only select the optimal action around 30% of the time, even less frequently than humans do. Furthermore, they exhibit a clear preference for safer actions, favoring lower-variance, lower-value choices that also make less physical progress toward goal. Frontier VLMs are better at estimating whether an action will succeed, placing the highest-success-probability action among their top choices in 83-92% of cases. Yet VLMs systematically conflate likelihood with value, assigning higher value to actions that are more likely to succeed (Spearman corr. 0.30 to +0.52), despite no such relationship in the ground truth (Spearman corr. -0.08). The conservatism therefore reflects a mis-calibration of value. SportD opens a new direction for rigorously evaluating physical strategic decision-making in VLMs, showing that careful decomposition of their choices can reveal the mechanisms underlying systematic biases such as risk aversion.
人間と AI の代替原則: あなたの組織はいつ AI に置き換えられますか?
人工知能 (AI) は組織を急速に変革しており、人間の従業員はいつ AI に取って代わられるのかという組織的および経済的な根本的な疑問を引き起こしています。階層型組織における人間と AI のタスク割り当て (HAT) を研究するための分析モデルを紹介します。 HAT モデルの中心的な特徴は、人間のスキル習得と AI 能力のスケーリングの間の経済的な非対称性を形式的にコード化していることです。 HAT モデルを使用すると、リスク調整済みのコスト、スキル、組織の深さ、導入規模、戦略的適応、およびリスクを総合して、いつ、どこで、なぜ、どのような構造的条件下で人間と AI の代替が発生するかをどのように決定するかを導き出すことができます。重要な結果は、人間と AI の代替原理です。これは、AI が人間の労働を置き換える正確な条件を提供します -- 形式的な非対称性の仮定に基づいています --。この結果に基づいて、AI の導入により、急激な労働力の移行、人間と AI のハイブリッド組織(リスクの異質性によって人間の割合の最小制約を必要とせずに人間と AI の役割が維持される場合を含む)、より広い制御範囲を備えたより平坦な管理階層が生成される可能性があることを示します。 HAT モデルは、中間管理職の役割が自動化に対して高い脆弱性を示す構造的条件を特定し、高度なスキルを持つ労働者の脆弱性が、組織の深さ、基準コスト、およびリスクの差によって形成されるスキルの閾値に依存することを示しています。より広範には、この論文は自動化の経済学、組織設計、AI ガバナンス、および労働力計画を AI 主導の組織変革の統一理論に結び付けます。
原文 (English)
The Human-AI Substitution Principle: When will you be replaced by AI in your organization?
Artificial Intelligence (AI) is rapidly transforming organizations, raising a fundamental organizational and economic question: when will a human employee be replaced by AI? We present an analytical model for studying Human--AI Task Allocation (HAT) in hierarchical organizations. A central feature of the HAT model is that it formally encodes the economic asymmetry between human skill acquisition and AI capability scaling. The HAT model allows us to derive how risk-adjusted costs, skills, organizational depth, deployment scale, strategic adaptation, and risk jointly determine when, where, why, and under what structural conditions human--AI replacement occurs. A key result is the Human--AI Substitution Principle, which provides a precise condition --- grounded in the formal asymmetry assumption --- under which AI replaces human labor. Building on this result, we show that AI adoption can produce abrupt workforce transitions, hybrid human--AI organizations, including cases where risk heterogeneity sustains human and AI roles without requiring a minimum-human-fraction constraint, and flatter managerial hierarchies with wider spans of control. The HAT model identifies structural conditions under which middle-management roles exhibit elevated vulnerability to automation, and shows that the vulnerability of highly skilled workers depends on a skill threshold shaped by organizational depth, baseline costs, and risk differentials. More broadly, the paper connects automation economics, organizational design, AI governance, and workforce planning into a unified theory of AI-driven organizational transformation.
A foundation model of numerical intelligence with cross-disciplinary generalization
Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. L…
Volve フィールドでの接地された良好な状態の異常検出: 構築されたラベル、ベースライン、およびデュアルヘッド モデル
マシンの状態を監視するための公開ベンチマークのほとんどは、意図的に障害が発生し、すべてのイベントが既知であるテスト装置から取得されます。実際の生産現場ではそのようなことはめったにありません。これらは、障害ログが添付されていないセンサー履歴を提供します。これはまさに、異常検出方法が独自のラベルを作成する必要があり、静かな仮定が気づかれずに紛れ込む可能性がある状況です。私たちは Equinor がリリースしたオープンな Volve フィールド データを使用しており、そのようなデータセットでは通常省略されている 2 つのことを真剣に受け止めています。まず、単なる数値のパターンではなく、物理的に問題が発生する可能性があると現場独自のエンジニアリング文書に記載されているものと照合して異常ラベルを作成し、すべてのラベルの背後にある理由を公開します。次に、教師なしベースラインと、イベントがいつ発生するか、その種類をマークする小さなデュアルヘッド モデルの両方を使用して、構築されたラベルが学習可能かどうかをテストします。このアイデアは、金属部品の欠陥検出に関する以前の研究から引き継いでいます。結果は正直です。ラベルを決して認識しない教師なし検出器は、ルールでフラグが立てられたのと同じ領域に依然として検出され、ラベルが恣意的ではないことがわかります。コンパクトな教師ありモデルは、これまでに見たことのないウェル全体でイベントの存在とイベント タイプを十分に回復し、時間内のイベントを大まかにのみ特定します。何がうまくいったのか、何がうまくいかなかったのか、そしてその間のすべての仮定を報告します。データセット、根拠のあるラベル、ラベルごとの来歴、ベースライン スコア、トレーニングされたモデル、コードは、CC-BY-NC-SA 4.0 に基づいて公開されています。
原文 (English)
Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model
Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known. Real production fields rarely offer that. They give you sensor histories with no fault log attached, which is exactly the situation where an anomaly-detection method has to invent its own labels, and where quiet assumptions can slip in unnoticed. We work with the open Volve field data released by Equinor and take two things seriously that such datasets usually skip. First, we build anomaly labels that are not just patterns in the numbers but are checked against what the field's own engineering documents say can physically go wrong, and we release the reasoning behind every label. Second, we test whether those constructed labels are learnable at all, using both an unsupervised baseline and a small dual-head model that marks when an event happens and what kind it is, an idea we carry over from earlier work on defect detection in metal parts. The results are honest. An unsupervised detector that never sees the labels still lands on the same regions our rules flagged, which tells us the labels are not arbitrary. A compact supervised model recovers event presence and event type well across wells it has never seen, and locates events in time only roughly. We report what worked, what did not, and every assumption in between. The dataset, grounded labels, per-label provenance, baseline scores, trained model, and code are released publicly under CC-BY-NC-SA 4.0.
フローバイフロー:高損失ドメインでの AI 出力を制御するためのコンテンツ判定バイパス
これまでの研究では、AI の出力速度 V が人間の認知能力 C_max を超えると、高損失領域では人間による監視が構造的に不可能になることが示されています。ただし、操作上の制約は V 単独ではなく、V x L です。ここで、L は項目ごとの認知負荷を示します。 L はトリアージ、判断、対応で構成され、AI の能力向上に対して非対称に対応します。セマンティックな不確定性は汎用設計に内在するため、モデルの機能が向上してもトリアージ コストは減少しません。応答コストは精度の向上に影響されません。判断コストのみが下方圧力に直面しており、この圧力は多くの場合、真の削減ではなく省略を誘発する形で作用します。したがって、能力の向上は L を削減するのではなく、再構築することになります。 AI の出力が正しいかどうかの評価に基づくガバナンス メカニズムは、その評価を AI に委任して幻覚リスクを継承するか、それを人間に委任して V x L の上限に直面するかのいずれかになります。私たちは、コンテンツを評価せずに監視負荷を制御するガバナンス パラダイムである Flow-by-Flow を提案します。正式な可算特徴に基づく認知コスト スコアは、大量生産に非線形コストを課す一方、制度上のキャパシティ キャップにより処理量が C_max 以内に保たれます。コンテンツ判定バイパス超過経路に対する 4 つの設計不変条件を導き出します。それは、コンテンツ判定なし、検査官能力のスケーラブルな消費なし、ID に縛られたアプリケーションごとの摩擦、およびバッチクリアランスなしです。 1 つの参考実装については、これらの不変式が同時に満たされることを示すために議論されていますが、その実際的な困難は明示的に認識されています。 1,000 のパラメータ描画にわたるモンテカルロ分析の例では、複合マルチメトリック フロー制御が試験の 90.8% で監視強化単独よりも優れていることが示唆されています。
原文 (English)
Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains
Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains once AI output velocity V exceeds human cognitive capacity C_max. The operative constraint, however, is V x L, where L is per-item cognitive load: triage, judgment, and response. These components respond asymmetrically to capability improvement. Triage cost does not decline, because semantic indeterminacy is inherent in general-purpose design. Response cost is invariant to accuracy. Only judgment cost faces downward pressure, largely by inducing omission. Capability improvement therefore restructures L rather than reducing it. We prove a proposition: if V x L grows at any positive compound rate while supervisory capacity grows linearly, exceedance occurs in finite time; capacity investment buys time only logarithmically, while reducing the growth rate extends it hyperbolically. Supervision enhancement and flow control are therefore not remedies of the same kind. We propose Flow-by-Flow, a governance design that prices supervisory load without evaluating content, intent, or legitimacy. A cognitive cost score built from formal, countable features imposes compounding costs on volume expansion, and an institutional capacity cap fixes processing within C_max. Four design invariants characterize any admissible exceedance pathway: no content judgment, no scalable consumption of examiner capacity, identity-bound per-application friction, and no batch clearance. Excess claim and page fees in patent systems are precursors satisfying only the first two invariants. One reference implementation satisfying all four is presented. A Monte Carlo analysis across 1,000 parameter draws confirms that the analytically derived ordering survives the 30-year horizon in 90.8% of trials.
ロジット境界の幾何学的信念インターフェイスとスパース シーフ エンクレーブ プロトコル: 安全なネットワーク電子医療記録 (EHR) の相互運用性のための内蔵型基板
電子医療記録の相互運用性は境界問題です。レガシー システム、生成モデル、用語サービス、アイデンティティ システム、および人間のレビュー担当者はそれぞれ豊富な内部状態を公開する可能性がありますが、運用上の交換には、型指定されたクレーム、限定された不確実性、来歴、および明示的な承認または棄権の狭い共有インターフェイスが必要です。この文書では、そのインターフェイスの数学的および工学的アーキテクチャについて詳しく説明します。組織化されたアイデアはロジット境界です。発見モデルはローカルのカテゴリ的決定に対して事前しきい値スコアを提案する可能性がありますが、決定論的な判断基盤は、その提案が許容されるか、レビューが必要か、または高速ヘルスケア相互運用性リソース (FHIR) トランザクションが構築される前に隔離する必要があるかを決定します。結果として得られる幾何学的信念インターフェイス (GBI) は、有限境界セマンティクス、局所的なディリクレ証拠、セルラー層およびマッピング コーンの診断、助言幾何学的監査チャート、およびフェールクローズ展開のための分散暗号化層エンクレーブ (DCSE) プロトコル スケッチを組み合わせています。このフレームワークは、臨床上の真実、世界的な表現の整合性、またはエンドツーエンドの安全性を確立するものではありません。これは、モデルとシステムの境界での証明書生成チェックを定義します。対となる凍結合成ベンチマークである GBI BoundaryBench v0.1 は、3 つの証拠モード (768 回の正規実行) にわたる 256 個の保留タスクで Qwen3-4B-Instruct-2507 を評価しました。すべての実行は完了しましたが、ベンチマーク コントラクトで受け入れられる出力を生成したものはありませんでした。安全な解析中に 369 件、スキーマ検証中に 399 件が拒否され、カバレッジはゼロとなり、決定的な隔離が行われました。この経験的結果は意図的に狭く、1 つの凍結界面の下に 1 つの 4B オープンウェイト モデルを配置しており、LLM の能力や臨床安全性に関する一般的な主張としてではなく、入院境界に関する証拠として報告されています。 Julia の付録では、標準ライブラリを使用して数値証明書を検証します。
原文 (English)
Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability
Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal states, while operational exchange requires a narrow shared interface of typed claims, bounded uncertainty, provenance, and explicit admission or abstention. This paper details a mathematical and engineering architecture for that interface. The organizing idea is the logit boundary: a discovery model may propose pre-threshold scores over a local categorical decision, but a deterministic judgment substrate decides whether the proposal is admissible, requires review, or must be quarantined before any Fast Healthcare Interoperability Resources (FHIR) transaction is constructed. The resulting Geometric Belief Interface (GBI) combines finite boundary semantics, local Dirichlet evidence, cellular-sheaf and mapping-cone diagnostics, advisory geometric audit charts, and a Decentralized Cryptographic Sheaf-Enclave (DCSE) protocol sketch for fail-closed deployment. The framework does not establish clinical truth, global representation alignment, or end-to-end safety; it defines certificate-producing checks at a model-to-system boundary. A companion frozen synthetic benchmark, GBI BoundaryBench v0.1, evaluated Qwen3-4B-Instruct-2507 on 256 held-out tasks across three evidence modes (768 canonical executions). All executions completed, but none produced an output accepted by the benchmark contract: 369 were rejected during safe parsing and 399 during schema validation, yielding zero coverage and deterministic quarantine. This empirical result is deliberately narrow - one 4B open-weight model under one frozen interface - and is reported as evidence about the admission boundary, not as a general claim about LLM capability or clinical safety. A Julia appendix verifies numerical certificates using standard libraries.
HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment
Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing…
ComBodied Agents: 人間中心のエージェント AI の新しいパラダイム
高齢者が薬を飲み忘れた場合、ソフトウェア エージェントが別のリマインダーを送信し、実体のあるエージェントが薬を持ってくることができます。しかし、その人が忘れているのか、混乱しているのか、副作用があるのか、それとも故意に拒否しているのか、またどのような支援が適切なのかについてはどちらも説明していない。これは、Agentic AI の構造的なギャップを明らかにします。デジタル エージェントは主にソフトウェアの状態を変換しますが、身体化されたエージェントは物理的な状態を変換します。どちらも、人の進化する状態や主体性をモデル化、介入、評価の主な対象にするものではありません。ソフトウェア ツール、センサー、ウェアラブル、ロボット、ヒューマン サービスを最終目標ではなくアクション チャネルとして使用し、時間の経過に伴う個々の人間の状態の軌跡を認識、モデル化、予測、サポートする人間中心のパラダイムである複合エージェントを紹介します。私たちは、パーソナル アシスタント、ヘルス エージェント、AI コンパニオン、および適応型人間 AI システムにわたる断片化された機能を統合して閉ループにします。イベントベースのマルチモーダルな知覚により、意味のある個人的なイベントが再構築されます。縦方向の修正可能な記憶は時間的コンテキストを提供します。個人世界モデルは、別の決定と介入の下で将来の個人の状態と結果を推定します。そして、許容可能な介入方針は、同意、不確実性、安全性、可逆性、およびユーザーの制御の下で、比例したサポートを選択します。人や環境からのフィードバックによってループが更新されます。このフレームワークは、網羅的なヒューマン デジタル ツインを必要とするのではなく、目的が限定され、不確実性を認識し、ユーザーが修正可能な表現を使用します。私たちは、人間の状態のターゲット、関係性のコンテキスト、およびエージェントの役割によって設計空間を整理し、シナリオ中心の評価、主体性維持のメトリクス、ベンチマーク要件、エッジネイティブの個人モデル、およびガバナンスの方向性を提案します。複合エージェントは、Agentic AI を外部タスクの完了から持続的な人間の利益へと移行させます。
原文 (English)
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human--AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.
Deep Activity Model: A Generative Approach for Human Mobility Pattern Synthesis
Human mobility plays a crucial role in transportation, urban planning, and public health, but current approaches face important limitations…
ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no s…
Explainability in Practice: A Survey of Explainable NLP Across Various Domains
Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management,…
Proportional Committee Elections with Positive and Negative Votes
In the classic committee election setting each voter approves a subset of candidates and the goal is to select $k$ winners based on these p…
Program Semantic Inequivalence Game with Large Language Models
Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-t…
COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification
We introduce \textbf{CoLoRA} (Convolutional Low-Rank Adaptation), a parameter-efficient fine-tuning method for convolutional neural network…
P2MFDS: A Privacy-Preserving Multimodal Fall Detection System for Elderly People in Bathroom Environments
By 2050, people aged 65 and over are projected to make up 16% of the global population. As aging is closely associated with increased fall…
Small Data Explainer -- The impact of small data methods in everyday life
The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings…
Commonsense on Demand: Generating and Selectively Integrating Commonsense Knowledge for Natural Language Inference
Natural Language Inference (NLI) determines whether a premise entails, contradicts, or is neutral with respect to a hypothesis. The task is…
Quantization-Aware Neuromorphic Architecture for Skin Lesion Classification on Resource-Constrained Devices
On-device skin lesion analysis is constrained by the compute and energy cost of conventional CNN inference and by the need for lightweight…
Empowering Children to Create AI-Enabled Augmented Reality Experiences
Despite their potential to enhance children's learning experiences, AI-enabled AR technologies are predominantly used in ways that position…
Ethics Practices in AI Development: An Empirical Study Across Roles and Regions
Recent advances in AI applications have raised growing concerns about the need for ethical guidelines and regulations to mitigate the risks…
CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D
Object retrieval from a scene has become a new trend of research due to its numerous applications. Recent approaches achieve zero-shot, ope…
Adaptive Online Learning with LSTM Networks for Energy Price Prediction
Accurate prediction of electricity prices is crucial for stakeholders in the energy market, particularly for grid operators, energy produce…
LiDAR-based 3D Change Detection at City Scale
High-definition 3D city maps enable city planning and change detection, which is essential for municipal compliance, map maintenance, and a…
MicroAUNet: Boundary-Enhanced Multi-scale Fusion with Knowledge Distillation for Colonoscopy Polyp Image Segmentation
Early and accurate segmentation of colorectal polyps is critical for reducing colorectal cancer mortality, which has been extensively explo…
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web appl…
A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL se…
Probably Approximately Correct Maximum A Posteriori Inference
Computing the conditional mode of a distribution, better known as the maximum a posteriori (MAP) assignment, is a fundamental task in proba…
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in acti…
LLM-Powered Automatic Translation and Urgency in Crisis Scenarios
Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication. H…
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
STEM Mental models can play a critical role in assessing students' conceptual understanding of a topic. They not only offer insights into w…
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
We study post-training linear autoregressive models with outcome and process rewards. Given a context $\boldsymbol{x}$, the model must pred…
Representation Finetuning for Continual Learning
The world is inherently dynamic, and continual learning aims to enable models to adapt to ever-evolving data streams. While pre-trained mod…
A Simple Efficiency Incremental Learning Framework via Vision-Language Model with Nonlinear Multi-Adapters
Incremental Learning (IL) aims to learn new tasks while preserving previously acquired knowledge. Integrating the zero-shot learning capabi…
Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderat…
VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Sequential Recommendation
Sequential Recommendation (SR) in multimodal settings typically relies on small frozen pretrained encoders, which limits semantic capacity…
REVERE: Reflective Evolving Research Engineer
Existing prompt-optimization techniques rely on local signals, causing poor generalization across tasks. In addition, they also rely on wea…
Designing Agentic AI-Based Screening for Portfolio Investment
We introduce a new agentic artificial intelligence (AI) platform for portfolio management. Our architecture consists of three layers. First…
Evaluation and Hardening of LLM System Instructions Against Extraction via Encoding Attacks
System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensit…
Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting
Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning. This paper addresses two related q…
Uncertainty as a Planning Signal: Multi-Turn Decision Making for Goal-Oriented Conversation
Goal-oriented conversational systems require making sequential decisions under uncertainty about the user's intent, where the algorithm mus…
TEMPER: Testing Emotional Perturbation in Quantitative Reasoning
Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, re…
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
Large language model (LLM) agents for sequential decision-making struggle to produce diverse outputs. This leads to insufficient exploratio…
Making Gaussian Kolmogorov-Arnold Networks Reliable and Accurate
Kolmogorov-Arnold Networks (KANs) replace fixed activations with learnable univariate edge functions whose behavior depends strongly on the…
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
Cloud-based Large Language Models (LLMs) can perform autonomous penetration-testing sub-tasks such as Linux privilege escalation, but raise…
Analytic Bridge Diffusions for Controlled Path Generation
Most modern bridge-diffusion methods achieve finite-time transport by specifying an interpolation, Schrodinger-bridge, or stochastic-contro…
CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) relies on evidence ranking to determine what information is exposed to the generator, yet existing ret…
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding s…
Pretraining large language models with MXFP4 on Native FP4 Hardware
Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain s…
Cavity-Enhanced Collective Quantum Processing with Polarization-Encoded Qubits
We introduce a cavity-enhanced optical architecture for collective quantum processing in which logical qubits are encoded in the polarizati…
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with beha…
memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations
Agent-memory frameworks -- mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, MemTensor -- each ship their own SDK, storage layout, and op…
Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association
Multi-view object association is an important computer vision problem that underlies many multi-camera perception tasks. While this task is…
ルーブリックベースの強化学習における報酬ハッキングの再現、分析、検出
ルーブリックベースの強化学習 (RL) は、LLM-as-a-Judge (LaaJ) を使用して、報酬としてルーブリックに従ってモデルの出力を採点します。ただし、政策モデルは裁判官の潜在的なバイアスを悪用し、報酬のハッキングや非効果的または危険なトレーニング結果につながる可能性があります。現実のルーブリックベースの RL では、このようなハッキング行為は多くの場合微妙であり、複数の裁判官のバイアスと絡み合っているため、分析、検出、軽減することが困難です。このペーパーでは、ルーブリックベースの RL のための制御可能なハッキング環境である CHERRL を紹介します。既知のバイアスを LaaJ に注入することで、CHERRL は報酬ハッキングの安定した再現、報酬の発散の明確な観察、およびハッキングの開始の正確な特定を可能にします。これは、ルーブリック ベースの RL における報酬ハッキングのメカニズムと緩和を研究するためのクリーンな実験テストベッドを提供します。その有用性を実証するために、発見可能性と悪用可能性の観点からさまざまな裁判官のバイアスを分析し、トレーニングログから報酬ハッキングの開始を自動的に検出するためのエージェントベースのシステムを調査します。コードと環境は https://github.com/THUAIS-Lab/CHERRL で公開されています。
原文 (English)
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a Controllable Hacking Environment for Rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, and explore an agent for automatically detecting reward hacking onset from training logs. The code and environment are publicly available at https://github.com/THUAIS-Lab/CHERRL.
ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset
Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query pr…
FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning
Contact-rich manipulation requires force sensitivity, but many robot arms lack dedicated force sensors due to their high cost. We present N…
ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory
Length extrapolation in language models involves competing objectives: retrieval fidelity, long-document likelihood, short-context quality,…
Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop
Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of…
ワールド モデルからワールド アクション モデルへ: ロボット工学の簡潔なチュートリアル
世界モデルは、身体化されたインテリジェンスや生成シミュレーションでますます使用されていますが、その範囲はコミュニティ全体で依然として曖昧です。このチュートリアルでは、タスク関連の観測または状態の将来の進化を推定するアクション条件付き予測モデルとしてのワールド モデルの設計空間ビューを示します。私たちは既存の手法を観測空間世界モデルと状態空間世界モデルに分類し、視覚的な忠実度、空間構造、物理的解釈可能性、制御の使いやすさにおけるトレードオフを比較します。さらに、予測された未来と実行可能なロボットのアクションを結び付ける世界アクション モデルを紹介し、想像してから実行する、ビデオ機能条件付きアクション予測、ビデオとアクションの共同モデリング、およびポリシー学習のための補助ビデオ予測という 4 つの代表的なパラダイムを要約します。このチュートリアルの目標は、世界 (アクション) モデルの概念的範囲を明確にし、具体化された予測と制御のための構造化された分類法を提供することです。
原文 (English)
From World Models to World Action Models: A Concise Tutorial for Robotics
Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. After reading the tutorial, readers should have a clear understanding of what constitutes a "world", how world models and world action models are defined, and what roles they play within robotic AI systems. The tutorial also develops a unified perspective for comparing representative approaches, such as World Labs' spatial intelligence models, Yann LeCun's JEPA framework, and NVIDIA's Cosmos platform, and clarifies how these models differ in their representations, predictive capabilities, and interaction mechanisms.
迅速な探索
すでに好まれている動作を繰り返しサンプリングすることによってポリシーを改善することはできないため、探索は RL にとって不可欠です。標準的な方法ではアクション空間に確率性を注入しますが、そのようなジッターはオリジナルに近いロールアウトしか生成しません。弱い政策から逃れるには、多くの場合、アクションのノイズでは生成できないグローバルな摂動が必要になります。大規模言語モデル (LLM) とビジョン言語アクション (VLA) モデルは、経路を提供します。自然言語プロンプトに基づいてポリシーを条件付けし、ロールアウトは自然言語プロンプトに基づいて行われるため、プロンプトを変更するとグローバルな変更が生じます。課題は、有益なグローバルな変化を引き起こすプロンプトを見つけることです。めったに成功しない弱い政策では、報酬は選択するにはあまりにも希薄です。私たちのアイデアは、ロールアウト自体からのプロンプトを改良することです。ビジョン言語モデル (VLM) がロールアウト ビデオを検討し、ポリシーがどのように反応したかを診断し、次回より良い動作を引き出すためにプロンプトを書き換えます。この手順は、古典的な RL 探索フレームワークである事後サンプリングをプロンプトのレベルで実現します。VLM は有用なプロンプトにわたる暗黙的な分布を維持し、観察されたロールアウトからそれを更新します。私たちはこの戦略を Prompt-Driven Exploration (PDE) と呼んでいます。操作タスクと推論タスクにわたって、PDE を使用すると、報酬ゼロのスタートからでも RL が成功したポリシーを学習できるようになり、サンプル効率がより広範囲に向上します。私たちの Web サイトは https://xinyunsunshine.github.io/prompt-rl でご覧いただけます。
原文 (English)
Prompt-Driven Exploration
Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. Escaping a weak policy often requires global perturbations that action noise cannot produce. Large language models (LLMs) and vision-language-action (VLA) models offer a pathway: they condition the policy on a natural language prompt, and since the rollout follows from it, modifying the prompt induces global changes. The challenge is finding prompts that induce useful global changes. With a weak policy that rarely succeeds, reward is too sparse to select on. Our idea is to refine prompts from the rollouts themselves: a vision-language model (VLM) reasons over the rollout video, diagnoses how the policy responded, and rewrites the prompt to elicit better behavior next time. This procedure resembles posterior sampling, a classical RL exploration framework, at the level of prompts: the VLM maintains an implicit distribution over useful prompts and updates it from observed rollouts. We call this strategy Prompt-Driven Exploration (PDE). Across manipulation and reasoning tasks, PDE enables RL to learn successful policies even from zero-reward starts, and improves sample efficiency more broadly. Our website is available at https://xinyunsunshine.github.io/prompt-rl.
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Ans…
Reducing Per-Sample Interference in Stochastic Optimization
Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments. W…
Cryptographically verifiable authorization for autonomous AI agents: A falsifiable hypothesis and proof-of-concept
Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing…
Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking
Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexp…
Continual Learning in Transition
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechani…
ED-CSP: 電子回折による結晶構造予測
まばらな、指数のない電子回折 (ED) 観察から周期的な 3D 結晶構造を復元することは、困難な生成逆問題です。既存の ED ベースの学習方法は、主に結晶学的ラベルを予測したり、インデックス付き反射から構造を再構築したり、有限構造ライブラリから候補を取得したりします。ここでは、化学組成、原子数、複数の検出面 ED スポット セットから結晶構造を予測する機械学習フレームワークである ED-CSP を紹介します。 ED-CSP は、リレーショナル セット エンコーダー、順列不変マルチビュー アグリゲーション、周期フロー ジェネレーターを組み合わせて、格子パラメーターと分数原子座標を共同で予測します。モデルをトレーニングするために、7 つのマテリアル リポジトリ間で重複を除去し、CHILI-100K の重複を除外するためにフィルタリングされた 485 万個のシミュレートされたマルチビュー ED 結晶構造のデータセットである ED-CS を構築します。 2,075 個の保持された CHILI-100K 材料では、CHILI のみでトレーニングされた ED-CSP は 57.49% MR@5 の構造一致率を達成し、粉末 X 線回折を条件とした最先端の結晶構造予測モデルである PXRDGen (52.92%) を上回りました。トレーニング データをスケーリングすると、パフォーマンスがさらに向上します。100 万構造の前駆体から初期化すると、MR@5 が 66.27% に上昇します。トレーニング検索ライブラリに含まれていない 1,024 個の構成でも、モデルは依然として 53.52% MR@5 を達成し、正確な式検索を超えた真の生成能力を実証しています。ターゲットの ED 観察を同一組成の非同形構造からの回折に置き換えると、MR@5 が 22.09 パーセントポイント減少し、予測が組成単独ではなく入力回折パターンに依存することが確認されました。 ED-CSP および ED-CS は、まばらな ED 観察からの生成結晶構造予測のベンチマークを確立し、将来の実験データへの移行のための基盤を提供します。
原文 (English)
ED-CSP: Crystal Structure Prediction from Electron Diffraction
Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based learning methods mainly predict crystallographic labels, reconstruct structures from indexed reflections, or retrieve candidates from finite structure libraries. Here, we introduce ED-CSP, a machine learning framework that predicts crystal structures from chemical composition, atom count, and multiple detector-plane ED spot sets. ED-CSP combines a relational set encoder, permutation-invariant multi-view aggregation, and a periodic flow generator to jointly predict lattice parameters and fractional atomic coordinates. To train the model, we construct ED-CS, a dataset of 4.85 million simulated multi-view ED crystal structures, deduplicated across seven materials repositories and filtered to exclude CHILI-100K overlaps. On 2,075 held-out CHILI-100K materials, ED-CSP trained only on CHILI achieves a structural match rate of 57.49% MR@5, outperforming PXRDGen (52.92%), a state-of-the-art crystal structure prediction model conditioned on powder X-ray diffraction. Scaling training data further improves performance: initializing from a one-million-structure precursor raises MR@5 to 66.27%. On 1,024 compositions absent from the training retrieval library, the model still achieves 53.52% MR@5, demonstrating true generative capability beyond exact-formula retrieval. Replacing target ED observations with diffraction from non-isomorphic structures of identical composition decreases MR@5 by 22.09 percentage points, confirming that predictions depend on the input diffraction patterns rather than composition alone. ED-CSP and ED-CS establish a benchmark for generative crystal structure prediction from sparse ED observations and provide a foundation for future transfer to experimental data.
Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows…
Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol
AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after…
Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction
Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Exist…
Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning
Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often…
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models
Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chain…
Persistent Recursive Worlds Enable Autonomous Software Evolution
Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems pre…
Inferential Capability Does Not Determine Legal Scope
Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capa…
ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, howev…
Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization
Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs…
Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization
Most image colorization systems operate in $Lab$ space by predicting chroma ($ab$) while preserving an input-derived luminance channel ($L$…
Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse
Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, impli…
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-fi…
Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis
Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and…