AIニュース 2026-07-09
自動生成: 2026-07-09 12:47 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
OpenAI、リアルタイム音声モデル「GPT-Live」公開 相づちも割り込みも、より人間らしい会話にITmedia AI+
OpenAIは、新世代のリアルタイム音声モデル「GPT-Live」を発表した。会話の最中に相手の話を聞きながら同時に話せる全二重方式を採用…
-
Our approach to government and national security partnershipsOpenAI
Learn how OpenAI approaches government and national security partners…
-
Separating signal from noise in coding evaluationsOpenAI
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular…
-
Helping K–12 educators build practical AI skillsOpenAI
OpenAI Academy and the Walton Family Foundation are bringing hands-on…
-
Claude「サブスク最上位プラン」6カ月間無料で提供 OSS開発者向けキャンペーン、対象を拡大ITmedia AI+
米Anthropicは、オープンソースソフトウェア(OSS)開発者向けのキャンペーン「Claude for Open Source」の対象…
-
Anthropic、「Claude Code」のシステムプロンプトを80%削減 「モデルの創造性を解放するため」ITmedia AI+
米Anthropicのタリク氏は、「Claude Fable 5」など最近リリースされたAIモデルへの禁止指示がモデルの創造性を制限すると…
-
OpenAIの最新AI「GPT-5.6」シリーズ、今週木曜日に一般公開へITmedia AI+
米OpenAIは7月7日(現地時間、以下同)、最新のAIモデル「GPT-5.6」シリーズを「今週の木曜日に一般公開する」と発表した。日本時…
トピック別件数
- LLM/生成AI 119件
- 研究/論文 115件
- エージェント 74件
- 画像/動画生成 58件
- ビジネス/資金調達 21件
- ロボティクス 16件
- その他 8件
- ハードウェア/半導体 5件
- 規制/政策 4件
日本語メディア10件
ITmedia AI+ (日本語)
「万年3位」から脱却なるか? Google Cloudに吹く"追い風"の正体を考察
エンタープライズ市場でAWSやMicrosoftに水をあけられてきたGoogle Cloudに、追い風が吹いている。国内SIer大手4社が同社との協業に動く理由と、これまでの経緯、AIエージェントでの連携戦略を読み解く。
Anthropic、「Claude Code」のシステムプロンプトを80%削減 「モデルの創造性を解放するため」
米Anthropicのタリク氏は、「Claude Fable 5」など最近リリースされたAIモデルへの禁止指示がモデルの創造性を制限すると指摘した。
CEOの利用額も「全社員に丸見え」 LayerXがAI予算を「第二の人件費」にした真意
「社長のAI利用額まで全社に生中継する」という透明性でコスト管理に挑むLayerX。利用額が予算の10倍を超過しても経営陣が現場を叱らなかった背景には、AI予算を「第二の人件費」として組織の資産に変える計算があった。単なる経費の締め付けを排し、外注費をAI費に切り替える基準から…
現場に聞いた「IT関連50製品」のぶっちゃけ理解度 “浸透するツール”の共通点とは?
似た機能を持つITツールが増える中、「なぜ定着する製品とそうでない製品が生まれるのか」。IT関連の主要50製品を対象に、利用イメージや認知度の違いから、企業が選ぶツールの条件を探った。
OpenAI、リアルタイム音声モデル「GPT-Live」公開 相づちも割り込みも、より人間らしい会話に
OpenAIは、新世代のリアルタイム音声モデル「GPT-Live」を発表した。会話の最中に相手の話を聞きながら同時に話せる全二重方式を採用し、ChatGPTの音声機能を刷新する。深い推論や検索が必要な場合はバックグラウンドで「GPT-5.5」に処理を委ねる仕組みを備え、有料・無…
フィジカルAI搭載ロボットがモノポリーを実演、1台のPCにモーション制御も統合
モベンシスは、「第38回 ものづくりワールド[東京]」の構成展である「第1回 フィジカルAI展[東京]」において、PCでリアルタイム制御を実現するソフトモーションコントローラー「WMX3」のROS 2向けパッケージ「WMX for ROS 2」を紹介した。
SDV時代支えるHEREの位置情報プラットフォーム、新たな柱は「二輪」へ
SDVによる変革が進むモビリティ社会において、「位置情報」は人々の移動にどのような価値をもたらすのか。デジタル地図からグローバルなロケーションのプラットフォーマーへと進化を遂げたHERE Technologiesの日本法人トップを務める枝隆志氏に話を聞いた。
Claude「サブスク最上位プラン」6カ月間無料で提供 OSS開発者向けキャンペーン、対象を拡大
米Anthropicは、オープンソースソフトウェア(OSS)開発者向けのキャンペーン「Claude for Open Source」の対象範囲を拡大したと発表した。認定した対象者には、最上位のサブスクリプションプラン「Claude Max 20x」を6カ月間無料で提供する。
「Grok 4.5」も明日一般公開へ マスク氏「Opus級だが、より高速で低コスト」
「Opusクラスのモデルだが、より高速で、トークン効率が高く、低コスト」
OpenAIの最新AI「GPT-5.6」シリーズ、今週木曜日に一般公開へ
米OpenAIは7月7日(現地時間、以下同)、最新のAIモデル「GPT-5.6」シリーズを「今週の木曜日に一般公開する」と発表した。日本時間では10日(金)になるとみられる。
海外メディア13件
TechCrunch AI (英語)
Lovable reportedly in talks to double its valuation to $13.2B
The $300 million round is expected to be led by Menlo Ventures, Sifted reported.
Google’s deepfake detector system used to debunk McConnell hoax pic
Earlier this week, a picture seemed to show Kentucky Senator Mitch McConnell covered in tubes in a hospital bed in a state of extreme distr…
This startup thinks robotics is about to have its ChatGPT moment
General Intuition is betting millions of hours of video game data can train the foundation models for physical AI, making it easier to buil…
Google Photos adds a new AI ‘Video Remix’ tool
The feature can do things like apply cinematic relighting to brighten up a dark clip, swap out a plain background for something fun, or add…
Why this CEO thinks video games make better training data than the internet
When it comes to achieving artificial general intelligence (AGI), large language models just don’t have what it takes. Models like ChatGPT…
Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
Meta is adding a new safeguard to stop people from secretly recording others with its AI glasses. But the update comes as the company conti…
OpenAI releases new voice models for more natural live conversations
OpenAI says its new voice mode can speak and listen at the same time, a key ability for live translation.
Prime Intellect raises $130M Series A to help enterprises build their own AI agents
Founded in 2024, Prime Intellect’s goal is to give organizations capabilities to train their own agentic systems without relying on frontie…
These AI startups are growing revenue at faster and faster rates
There are a lot of fast-growing AI startups, but some are growing even faster, they say.
Your gaming data could be the secret to AGI, according to this Bezos-backed startup
When it comes to achieving artificial general intelligence (AGI), large language models just don’t have what it takes. Models like ChatGPT…
Former OpenAI exec Kevin Weil is now on the board of Stoke Space
Kevin Weil's new role at Stoke Space suggests reusable rockets are the next hot thing in Silicon Valley.
Hot French startup ZML releases free product to speed inference across lots of AI chips
ZML, a hot French AI startup endorsed by Turing Award winner Yann LeCun, has now released ZML/LLMD, software that could make running AI les…
AI chip maker SambaNova raises $1B at $11B valuation, 5 months after last mega round
AI chip maker SambaNova has raised at an $11 billion valuation months after Intel was rumored to be trying to buy it for about $1.6 billion.
公式ブログ3件
OpenAI (英語)
Our approach to government and national security partnerships
Learn how OpenAI approaches government and national security partnerships, with principles for responsible AI use, democratic accountabilit…
Separating signal from noise in coding evaluations
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in…
Helping K–12 educators build practical AI skills
OpenAI Academy and the Walton Family Foundation are bringing hands-on AI Skills Jams to help K–12 educators build practical AI skills for t…
論文279件
arXiv cs.AI (英語)
Prompt-to-Paper: バイオインフォマティクス用エージェント AI システム
大規模言語モデルの最近の進歩により、エンドツーエンドの自動原稿生成が可能になりましたが、既存のシステムには 3 つの重大な欠陥があります。(i) 生成された主張は検証可能な文献に決定論的に基づいていない、(ii) 実験結果は実行されるのではなく捏造されることが多い、(iii) AI によって生成された原稿が現実世界の出版に必要な品質と厳密さを満たしているかどうかを評価するための標準化された多次元フレームワークが存在しない。私たちは、3 つの統合されたイノベーションを通じてこの評価ギャップに直接対処するマルチエージェント フレームワークである Prompt-to-Paper を紹介します。まず、セクションを意識した関連性スコアリングと雪だるま式引用拡張を備えた決定論的な検索拡張生成パイプラインにより、60 ~ 100 件の検証可能なコーパス内のすべての主張が根拠付けされます。第 2 に、自律コーディング エージェントが実際の計算生物学実験を実行し、合成出力を本物の数値結果に置き換えます。 3 番目に、8 次元の自動品質スコアラーは、出版された論文からのおおよその参照統計でベンチマークされ、明示的な幻覚ペナルティで強化され、標準化された再現可能な品質評価を提供します。品質主導の改善ループでは、各反復を 3 つの研究者のアクションのいずれかにルーティングし、10 回の反復ごとに詳細な調査サイクルを起動して実験を再実行し、より強力な出力から原稿を再作成する、コンテキスト豊富な改訂機能を使用します。私たちは 5 つのバイオインフォマティクスのケーススタディに基づいてシステムを検証します。 5 件すべてのケースで、範囲外の引用がまったく含まれていない提出形式の PDF が編集されました。改善ループにより、原稿の品質が 0 ~ 100 段階で平均 +17.96 ポイント (最大 +26.04 ポイント) 向上します。部分的な外部チェックとして、人間の査読者が 5 つの原稿を 10 点中平均 7.0 点で採点しました。完全な原稿は、論文あたり約 0.31 米ドルで作成されます。
原文 (English)
Prompt-to-Paper: Agentic AI System for Bioinformatics
While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whether AI-generated manuscripts meet the quality and rigor required for real-world publication. We present Prompt-to-Paper, a multi-agent framework that directly addresses this evaluation gap through three integrated innovations. First, a deterministic retrieval-augmented generation pipeline with section-aware relevance scoring and snowball citation expansion grounds every claim in a verifiable corpus of 60--100 papers. Second, an autonomous coding agent executes real computational biology experiments replacing synthetic outputs with genuine numerical results. Third, an eight-dimensional automated quality scorer, benchmarked with approximate reference statistics from published papers and augmented with explicit hallucination penalties, provides standardized, reproducible quality assessments. The quality-driven improvement loop uses a context-rich reviser that routes each iteration to one of three researcher actions and fires a deep research cycle every ten iterations to re-run experiments and re-manuscript from stronger outputs. We validate the system on five bioinformatics case studies; all five cases compiled submission-formatted PDFs with zero out-of-range citations. The improvement loop raises manuscript quality by an average of +17.96 points on a 0--100 scale (maximum +26.04. As partial external checks, a human reviewer scored the five manuscripts at an average of 7.0 out of 10. Complete manuscripts are produced at approximately 0.31 USD per paper.
グラフからグラデーションまで: サイバーフィジカル IoT システムおよびそれ以降の物理学にヒントを得た構造属性
人工知能における解釈可能な説明方法は、根本的な原因とその影響を明らかにし、異なる入力の下でシステムが特定の方法で動作する理由をより深く理解できるようにすることを目的としています。主に入力変数と出力変数の間の相関関係を強調する従来の説明可能性手法とは異なり、因果関係の説明は介入的な質問に焦点を当てます。そうすることで、より堅牢な洞察が提供され、特にリスクの高い領域での自動化された意思決定をユーザーが理解できるようになります。しかし、明示的な有向因果構造の回復は、フィードバック ループと部分的な可観測性を備えた大規模なハイブリッド サイバー物理システムでは現実的ではないことがよくあります。このペーパーでは、サイバーフィジカル IoT システムの無指向性のエネルギーベースの表現を通じて変数の依存関係をモデル化する、統計力学に触発された新しいフレームワークを紹介します。私たちのアプローチは、有向因果グラフを復元することなく、エネルギー状況の変動が個々のコンポーネントの影響をどのように反映しているかを分析することにより、厳密な依存関係を意識した帰属を可能にします。また、ハイブリッド相互作用にわたる摂動効果についての推論もサポートし、異常な動作の信頼できる説明を提供します。私たちは、ハイブリッド連続変数と離散変数を使用した産業用 IoT テストベッドでのシミュレーションを通じてフレームワークを実証的に検証し、最先端のグラフベースのアプローチよりも高い帰属精度、堅牢性の向上、スケーラビリティの向上を実証しました。アトリビューションは、システムの生成ダイナミクスを完全に回復することを目的としたものではありませんが、人間の解釈と下流の予測および診断タスクの両方をサポートする、依存関係を意識した貴重な説明を提供します。産業用 IoT セキュリティで実証されていますが、私たちのフレームワークは、原理的で構造的な説明を必要とする他の高次元のサイバー物理システムや社会技術システムにも適用されます。
原文 (English)
From Graphs to Gradients: Physics-Inspired Structural Attribution for Cyber-Physical IoT Systems and Beyond
Interpretable explanation methods in Artificial Intelligence aim to uncover the underlying causes and their effects, enabling a deeper understanding of why a system behaves in a certain way under different inputs. Unlike traditional explainability methods, which mainly highlight correlations between input and output variables, causal explanation focuses on interventional questions. By doing so, it provides more robust insights, helping users understand automated decisions, especially in high-risk domains. Recovering an explicit directed causal structure, however, is often impractical in large-scale, hybrid cyber-physical systems with feedback loops and partial observability. This paper introduces a novel framework inspired by statistical mechanics that instead models variable dependencies through an undirected, energy-based representation of cyber-physical IoT systems. Our approach enables rigorous dependency-aware attribution by analysing how variations in the energy landscape reflect the influence of individual components, without recovering a directed causal graph. It also supports reasoning about perturbation effects across hybrid interactions, providing reliable explanations of abnormal behaviours. We empirically examined our framework through simulations on an industrial IoT testbed with hybrid continuous and discrete variables, demonstrating higher attribution accuracy, improved robustness and better scalability than state-of-the-art graph-based approaches. While the attributions are not intended to fully recover the system's generative dynamics, they provide valuable, dependency-aware explanations supporting both human interpretation and downstream predictive and diagnostic tasks. Although demonstrated in industrial IoT security, our framework also applies to other high-dimensional cyber-physical and socio-technical systems requiring principled, structural explanations.
CSTutorBench: ブロックベース プログラミングの家庭教師としての小規模言語モデルのベンチマーク
大規模な言語モデルは AI の家庭教師としてますます検討されていますが、幼稚園から高校までの環境に導入すると、プライバシー、コスト、独自のモデルへの依存に関する懸念が生じます。小規模言語モデル (SLM) は有望な代替手段を提供しますが、特定の教育コンテキストに適切なモデルを選択することは依然として困難であり、特にブロックベースのプログラミングなどのターゲット ドメインがモデル トレーニング データにほとんど含まれていない場合には困難です。ブロックベースのロボット環境である VEX VR で CS チューターとして言語モデルを評価するためのベンチマークである CSTutorBench を紹介します。このベンチマークは、確立された個別指導とフィードバック研究に基づいた教育ルーブリックに基づいて採点された 17 のシナリオベースの質問で構成され、評価には人間参加型の LLM による審査員パイプラインが使用されます。 11 のモデル (4B ~ 120B パラメーター) にわたる予備調査結果では、モデルは語彙や口調などの表面レベルの基準では良好に機能しますが、より深い教育的行動、特に答え漏れの回避や生徒のデバッグ履歴への関与に苦戦していることが明らかになりました。私たちのサンプルでは、モデルの数が少ないため、この結論の強度が制限されますが、モデルファミリーと命令チューニングのアプローチは、パラメーター数だけよりも個別指導の質をより良く予測するものであるように見えます。最近の教育プロンプト工学研究に基づいた、的を絞ったプロンプト改訂により、11 モデル中 10 モデルのスコアが向上しました。これらの結果は、教育展開における SLM 選択のための、状況に応じた教育学的に根拠のあるベンチマークの価値を強調しています。
原文 (English)
CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
Large language models are increasingly explored as AI tutors, yet deploying them in K-12 settings raises concerns around privacy, cost, and reliance on proprietary models. Small language models (SLMs) offer a promising alternative, but selecting the right model for a specific educational context remains difficult, particularly when the target domain, such as block-based programming, is largely absent from model training data. We introduce CSTutorBench, a benchmark for evaluating language models as CS tutors in VEX VR, a block-based robotics environment. The benchmark comprises 17 scenario-based questions scored against a pedagogical rubric grounded in established tutoring and feedback research, with a human-in-the-loop LLM-as-judge pipeline for evaluation. Preliminary findings across 11 models (4B-120B parameters) reveal that models perform well on surface-level criteria such as vocabulary and tone but struggle with deeper pedagogical behaviors, particularly avoiding answer leakage and engaging with student debugging histories. In our sample, model family and instruction-tuning approach appear to be better predictors of tutoring quality than parameter count alone, though the small number of models limits the strength of this conclusion. A targeted prompt revision grounded in recent educational prompt engineering research improved scores for 10 of 11 models. These results underscore the value of context-specific, pedagogically grounded benchmarks for SLM selection in educational deployment.
自動 CAD 生成のための基礎モデル
大規模言語モデル (LLM) とビジョン言語モデル (VLM) の最近の進歩により、自然言語仕様からパラメトリック 3D デザインを自動生成できるようになりました。この章では、統合評価パイプラインと 97 のエンジニアリング設計問題の厳選されたベンチマークを使用した、機械部品のコンピュータ支援設計 (CAD) 自動生成のための基礎モデルの実証的研究について説明します。 JSON スキーマ検証、分析特徴スコアリング、メッシュ合成、および複数ラウンドの反復改良を統合するマルチモデルのテキストから CAD フレームワークである LLMForge を紹介します。LLMForge は 2 つの批評体制の下で研究されています。 IterTracer は、解析的な視覚メトリクス (シルエット IoU、穴の可視性、エッジ クリアランス、アスペクト比適合性) を備えたフォン シェーディング レイトレース レンダラを使用して、ラウンド全体にわたる軽量のジオメトリ対応フィードバックを実現します。 IterVision は、分析スコアラーを VLM セマンティック クリティカル (Qwen2.5-VL-72B) に置き換えます。VLM セマンティック クリティカルは、思考連鎖による視覚的推論を介してレンダリングされたビューを評価し、空間的一貫性と設計意図を評価します。 4 つの標準ジオメトリ ファミリ (穴とボルト サークル付きのプレート、マルチフィーチャー ボックス、フランジ付きシリンダー、および L ブラケット) にわたるベンチマークで、DeepSeek-V3.2、Qwen3-235B-A22B、Llama-3.3-70B、Gemma-3-27B、GLM-4.5、MiniMax-M2.1、および INTELLECT の 7 つの基礎モデルを評価します。 IterTracer では、最高ランクの 4 つのモデルが 98.97% のメッシュ成功率で緊密なクラスター ([0.885, 0.890] の全体平均) を形成しており、命令調整されたコンパクトなモデルがかなり大規模なシステムに匹敵できることを示しています。 IterVision での VLM ベースの批評により、主要なモデルでは 100% 防水性の高いメッシュ生成が実現しましたが、視覚的スコアとセマンティック スコアが最も異なる円柱などの回転対称ジオメトリでは体系的な困難が表面化しました。ベンチマーク設計、故障モード、CAD 指向のプロンプト、産業ワークフローとスケーラブルな自動機械設計への影響について説明します。
原文 (English)
Foundation Models for Automatic CAD Generation
Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs from natural-language specifications. This chapter presents an empirical study of foundation models for automatic Computer-Aided Design (CAD) generation of mechanical parts, using a unified evaluation pipeline and a curated benchmark of 97 engineering design problems. We introduce LLMForge, a multi-model text-to-CAD framework integrating JSON-schema validation, analytic feature scoring, mesh synthesis, and multi-round iterative refinement, studied under two critique regimes. IterTracer uses a Phong-shaded ray-trace renderer with analytic visual metrics (silhouette IoU, hole visibility, edge clearance, aspect-ratio conformance) for lightweight geometry-aware feedback across rounds. IterVision replaces the analytic scorer with a VLM semantic critic (Qwen2.5-VL-72B) that evaluates rendered views via chain-of-thought visual reasoning, assessing spatial coherence and design intent. On a benchmark spanning four canonical geometry families (plates with holes and bolt circles, multi-feature boxes, flanged cylinders, and L-brackets), we evaluate seven foundation models: DeepSeek-V3.2, Qwen3-235B-A22B, Llama-3.3-70B, Gemma-3-27B, GLM-4.5, MiniMax-M2.1, and INTELLECT. Under IterTracer, the four highest-ranked models form a tight cluster (overall mean in [0.885, 0.890]) with 98.97% mesh success, showing that compact instruction-tuned models can match substantially larger systems. VLM-based critique in IterVision yields 100% watertight mesh generation on the leading model while surfacing systematic difficulty on rotationally symmetric geometries such as cylinders, where visual and semantic scoring diverge most. We discuss benchmark design, failure modes, CAD-oriented prompting, and implications for industrial workflows and scalable automated mechanical design.
物語世界モデル: ナラトロジーに基づいた長編小説のための作家の記憶
長編小説の作家には、展開する物語の状況に関するマルチホップの質問に答える記憶が必要です。つまり、誰が秘密を知っていて、それをいつ知ったのか、それを明らかにするナレーションに先行する出来事があったかどうか、設定が功を奏したかどうか、関係がどのように変化したかなどです。汎用検索システムとエージェント記憶システムは、エンティティと事実を表しますが、これらの質問が対象とする物語構造を表すものではないため、間違った証拠が表面化するか、まったく証拠が存在しません。ナラティブ ワールド モデル (NWM) は、ナラトロジーに基づいた型付き時間状態グラフとクエリ条件付きハイブリッド検索を組み合わせたライター記憶システムです。回答者ではなくメモリを測定するために、再現可能なパブリック コーパスと検証済みのマルチホップ ベンチマーク上で、そのシステムのチャプターセーフな証拠のみを対象に、単一の定数の Opus 4.8 リーダーを介してすべてのシステムを読み取り、既存の最強の時間知識グラフ エージェント メモリ フレームワークである Graphiti/Zep (Rasmussen et al., 2025) と比較します。 NWM は、両方のコーパスにわたるマルチホップ ナラトロジー QA でこのベースラインを大幅に上回り、GraphRAG およびフラット検索をはるかに上回ります。この利点は、抽出のアーティファクトではなく、表現的なものです。NWM 独自の抽出機能を使用してベースラインを再構築しても存続し、グラフのサイズや抽出機能の品質ではなく、ナラトロジーに基づいた構造とクエリ条件付き検索を追跡します。
原文 (English)
Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction
Long-form fiction writers need memory that answers multi-hop questions about evolving story state: who knows a secret and when they learned it, whether an event preceded the narration that revealed it, whether a setup paid off, and how a relationship shifted. General-purpose retrieval and agent-memory systems represent entities and facts but not the narratological structure these questions turn on, so they surface the wrong evidence or none at all. We introduce the Narrative World Model (NWM), a writer-memory system that pairs a narratology-grounded typed temporal-state graph with query-conditioned hybrid retrieval. To measure memory rather than the answerer, we read every system through a single held-constant Opus 4.8 reader over only that system's chapter-safe evidence, on a reproducible public corpus and a validated multi-hop benchmark, and we compare against the strongest existing temporal-knowledge-graph agent-memory framework, Graphiti/Zep (Rasmussen et al., 2025). NWM substantially and significantly outperforms this baseline on multi-hop narratological QA across both corpora, and far exceeds GraphRAG and flat retrieval. The advantage is representational rather than an artifact of extraction: it survives rebuilding the baseline with NWM's own extractor, and traces to its narratology-grounded structure and query-conditioned retrieval, not to graph size or extractor quality.
FirstResearch: LLM 科学的発見エージェントのための監査可能な質問の形成
科学的発見のための LLM システムは、着想、文献の統合、実験計画、レポートの作成をますます支援していますが、彼らが提案する最初の研究課題を監査するのは依然として難しい場合があります。科学者が検査すべきメカニズム、改ざん者、または仮定を暴露することなく、それがもっともらしく聞こえるかもしれません。 FirstResearch を紹介します。FirstResearch は、構造化された Research Question Certificate をコア成果物とする科学 LLM エージェント向けの第一原理リサーチ質問形成フレームワークです。証明書には、原始的な定義、仮定、メカニズム モデル、緊張または矛盾、反証可能な仮説、最小限の決定的なテスト、および失敗の更新ルールが記録されており、提案された質問を下流の実行前に検査できるようになります。 LLM エージェントの 10 個の研究トピックに関して、FirstResearch は、AI 共同科学者、エージェント ラボラトリー、および AI Scientist-v2 にインスピレーションを得た、主要な DeepSeek ブラインド ジャッジ プロトコルの下で制御されたプロンプト レベルのベースラインを上回りました。同じ 40 のベースライン パッケージの Gemini-2.5-Flash の独立審査員によるスコアは、システム レベルのランキングを維持しており、FirstResearch のスコアは 4.86/5 対 4.38/5 で最も強力なベースラインであり、ピアソンの一致は平均スコアで 0.865 でした。 1 回のアブレーション チェックポイントは、証明書中心のコアが最も強力なコンポーネントであることをさらに示唆しています。証明書のみのスコアは、DeepSeek では 4.90/5、Gemini では 4.88/5 に達しましたが、証明書の削除は両方の審査員で 1/5 を下回りました。これらの結果は予備的なものであり、人間の領域の専門家ではなく LLM の審査員を使用していますが、明示的な導出制約は、LLM によって生成された科学的な質問をより監査可能にするための有望なメカニズムであるという、狭い科学的発見の主張を裏付けています。コード、プロンプト、保存された出力、および再現スクリプトは、https://github.com/louiswang524/FirstResearch で入手できます。
原文 (English)
FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents
LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect. We introduce FirstResearch, a first-principles research-question formation framework for scientific LLM agents whose core artifact is a structured Research Question Certificate. The certificate records primitive definitions, assumptions, a mechanism model, a tension or contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule, making the proposed question inspectable before downstream execution. On ten LLM-agent research topics, FirstResearch outperforms controlled prompt-level baselines inspired by AI co-scientist, Agent Laboratory, and AI Scientist-v2 under a primary DeepSeek-blind-judge protocol. A Gemini-2.5-Flash independent-judge rescore of the same 40 baseline packages preserves the system-level ranking, with FirstResearch scoring 4.86/5 versus 4.38/5 for the strongest baseline and Pearson agreement of 0.865 on average score. A one-repeat ablation checkpoint further suggests that the certificate-centered core is the strongest component: certificate-only scoring reaches 4.90/5 under DeepSeek and 4.88/5 under Gemini, while removing certificates drops below 1/5 under both judges. These results are preliminary and use LLM judges rather than human domain experts, but they support a narrow scientific-discovery claim: explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable. Code, prompts, saved outputs, and reproduction scripts are available at https://github.com/louiswang524/FirstResearch.
ループ内のメモリ: 言語エージェントの拡張作業メモリとしてのプロセス内取得
言語エージェントはループを実行します - 観察、推論、行動 - しかし、彼らが推論する記憶はループの外にあり、ストアはターンごとに最大 1 回クエリされます。私たちは、メモリがループ内を移動し、各ステップで読み書きされる体制を研究します。障害となるのは常に遅延です。ネットワーク化されたストアは数十ミリ秒から数百ミリ秒で応答します。また、ループ内取得では、取得にコストがかかる場合、エンドツーエンドの遅延が最大 83 倍に膨らむ可能性があります。以前の研究では、そのコストを問題視するのではなく管理していました。サービングレイヤーのスケジューリングによってそれが隠蔽され、「メモリファースト」設計により、取得がターンごとに 1 回に割り当てられました。私たちは、レイテンシはループ内パターンではなく、ストアが存在する場所の特性であると主張します。処理中のストアは、ネットワーク レジームより 3 桁低い、最大 100 マイクロ秒で応答し、その速度ではステップごとの税が崩壊します。拡張思考論文のパリティ原理により、常に直接利用できるほど高速なストアは、エージェントが単に参照するツールではなく、拡張作業メモリになります。前提は因果関係です。固定のターンごとのメモリ レイテンシ バジェットを保持し、ストアの応答速度のみを変化させると、冗長アクションはレイテンシとともに単調増加します。インプロセス速度では 0.0/12、110 ミリ秒のクラウド往復では 7.2/12 (gpt-5-nano、gpt-5-mini; 正確な順列 p=0.0079)。この体制をエンドツーエンドで実証します。制限されたウィンドウの下で 4 つの GPT-5 クラス モデルにわたって、ループ内メモリでリコールが 0/5 から 3.6 ~ 4.8/5 に改善され、p50 80 ~ 165us でのストア操作が行われます。ただし、指示された restate-every-reply ベースラインでも完全に解決されますが、トークン コストはワーキング セットとともに増加します。ストアはいかなる実行においても事実を失うことはありませんでした (244 件の書き込みのうち 244 件が保持されました)。すべてのミスは、ストアではなくエージェントの読み取りポリシーを追跡します。私たちの測定では、ボトルネックの位置も再確認されています。ステップごとの主なコストは埋め込みです (ネットワーク上で約 200 ~ 400 ミリ秒)。インプロセス ストアを小さなローカル エンベッダーと組み合わせると、完全な操作が測定値で約 40 マイクロ秒に戻ります。
原文 (English)
Memory in the Loop: In-Process Retrieval as ExtendedWorking Memory for Language Agents
Language agents run a loop - observe, reason, act - but the memory they reason over sits outside it: a store queried at most once per turn. We study the regime where memory moves inside the loop, read and written on every step. The obstacle has always been latency: networked stores answer in tens to hundreds of milliseconds, and in-loop retrieval can inflate end-to-end latency by up to 83x when retrieval is expensive. Prior work manages that cost rather than questioning it: serving-layer scheduling hides it, "memory-first" designs ration retrieval to once per turn. We argue latency is a property of where the store lives, not the in-loop pattern: an in-process store answers in ~100us, three orders of magnitude below the network regime, and at that speed the per-step tax collapses. By the extended-mind thesis's parity principle, a store fast enough to be constantly and directly available becomes extended working memory, not a tool the agent merely consults. The premise is causal: holding a fixed per-turn memory-latency budget and varying only the store's answer speed, redundant actions rise monotonically with latency - 0.0 of 12 at in-process speed, 7.2 of 12 at a 110ms cloud round trip (gpt-5-nano, gpt-5-mini; exact permutation p=0.0079). We demonstrate the regime end-to-end: across four GPT-5-class models under a bounded window, recall improves from 0/5 to 3.6-4.8/5 with in-loop memory, store ops at p50 80-165us - though an instructed restate-every-reply baseline also solves it perfectly, at a token cost that grows with the working set. The store never lost a fact in any run (244 of 244 writes kept); every miss traces to the agent's read policy, not the store. Our measurements also relocate the bottleneck: the dominant per-step cost is embedding (~200-400ms over the network); pairing the in-process store with a small local embedder returns the complete operation to a measured ~40us.
Akashic: Memtention を使用した低オーバーヘッド LLM 推論サービス
最近の LLM ベースのエージェント システムは、マルチターン インタラクション、ツール呼び出し、およびセッション間のワークフローにわたってコンテキストを継続的に蓄積します。すべてのリクエストの完全な履歴を再生することはすぐに現実的ではなくなります。コンテキストが長いと、プレフィルのコストが増加し、コンテキストの制限を超える可能性があり、タスクに関連する証拠が無関係なコンテンツに埋もれてしまうことが多く、サービス効率と出力品質の両方が低下します。私たちは、MemAttendant を中心に構築された低オーバーヘッドのメモリ システムである Akashic を提案します。これは、コンテキストを境界のあるチャンクに編成し、チャンク間の意味関係をモデル化し、完全な履歴を繰り返し書き換えることなく、チャンク間の証拠を保存します。 Akashic はさらに、ハードウェアとソフトウェアが共同設計したメモリ配置を適用して、共同取得される可能性のあるチャンクを同じ場所に配置し、取得の断片化と I/O オーバーヘッドを削減します。 Akashic は、4 つの代表的なワークロードと 3 つのモデル サイズにわたって、以前の強力なメモリ ベースラインと比較して、タスクの精度を最大 10.2 ポイント、スループットを最大 1.21 倍、持続可能なリクエスト レートを最大 1.88 倍向上させます。
原文 (English)
Akashic: A Low-Overhead LLM Inference Service with MemAttention
Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every request quickly becomes impractical: long contexts increase prefill cost, may exceed context limits, and often bury task-relevant evidence in irrelevant content, degrading both serving efficiency and output quality. We propose Akashic, a low-overhead memory system built around MemAttention, which organizes context into bounded chunks and models semantic relationships across chunks, preserving cross-chunk evidence without repeatedly rewriting the full history. Akashic further applies hardware-software co-designed memory placement to co-locate likely co-retrieved chunks, reducing retrieval fragmentation and I/O overhead. Across four representative workloads and three model sizes, Akashic improves task accuracy by up to 10.2 points, throughput by up to 1.21x, and sustainable request rate by up to 1.88x over strong prior memory baselines.
ArtisanCAD: 専門家に基づいた知識を抽出した産業レベルの CAD エージェント
産業用コンポーネントのコンピュータ支援設計 (CAD) には、長期的な手順モデリング、堅牢な機能の依存関係、編集可能なパラメトリック ジオメトリ、および生産グレードの B-Rep の実行が必要です。既存のテキストから CAD への手法は、自然言語記述から CAD プログラムを生成する点で有望な進歩を遂げてきましたが、ユーザーのプロンプトがあいまいな場合、指定が不十分な場合、または高レベルの設計意図しか説明していない場合には依然として困難を伴います。また、CATIA 操作記録、マクロ ログ、図面メモ、エンジニアリングの説明など、産業ワークフローで自然に利用できる専門的な手順知識を活用することはほとんどありません。 \algname は、専門家に基づいた知識を抽出した、スキルガイド付きの産業用 CAD エージェントです。 \algname の中核は CAD 中間表現 (CAD-IR) であり、パラメーター、順序付けされた操作、MCP ツール バインディング、依存関係、生成されたエンティティ、および検証ルールをエンコードする実行可能な手続き表現です。 CAD-IR は 2 つの重要な役割を果たします。1 つは、専門家の CAD プロシージャを再利用可能なパラメータ化されたスキルに蒸留するためのキャリアとして機能します。次に、曖昧なプロンプトまたは中間レベルのプロンプトを完全な実行可能な CAD 操作に変える手続き型の足場を提供します。 \algname は、専門家が導き出したスキルを取得し、CAD-IR をインスタンス化して修正し、専用の CATIA-MCP バックエンドを通じて結果の手順を実行し、マルチビューの視覚的フィードバックを使用して反復改良を行い、最終的に実稼働対応の B-Rep モデルを生成します。 Text2CAD ベンチマークでは、CAD-IR は平均面取り距離を $14.83$ から $9.88$ に削減することで中間プロンプトからの生成を改善し、あいまいなテキストの意図と実行可能な CAD 構築を橋渡しする能力を示しています。 4 つの複雑な自動車コンポーネントについて、CAD-IR を使用すると、専門家による CATIA 記録を再利用可能なスキルに蒸留でき、\algname が新しいバリアント リクエストに対して編集可能な CATIA ネイティブ B-Rep モデルを生成できるようになります。
原文 (English)
ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation
Computer-aided design (CAD) for industrial components requires long-horizon procedural modeling, robust feature dependencies, editable parametric geometry, and production-grade B-Rep execution. Existing text-to-CAD methods have made promising progress in generating CAD programs from natural-language descriptions, but they still struggle when user prompts are ambiguous, underspecified, or only describe high-level design intent. They also rarely exploit expert procedural knowledge naturally available in industrial workflows, such as CATIA operation recordings, macro logs, drawing notes, and engineering descriptions. We present \algname, a skill-guided industrial CAD agent with expert-grounded knowledge distillation. The core of \algname is CAD intermediate representation (CAD-IR), an executable procedural representation that encodes parameters, ordered operations, MCP tool bindings, dependencies, generated entities, and verification rules. CAD-IR plays two key roles: it first serves as the carrier for distilling expert CAD procedures into reusable parameterized skills; then it provides a procedural scaffold that turns vague or intermediate-level prompts into complete executable CAD operations. \algname retrieves expert-derived skills, instantiates and revises CAD-IR, executes the resulting procedure through a dedicated CATIA-MCP backend, and uses multi-view visual feedback for iterative refinement, and finally generates production-ready B-Rep models. On the Text2CAD benchmark, CAD-IR improves generation from intermediate prompts by reducing mean Chamfer Distance from $14.83$ to $9.88$, showing its ability to bridge ambiguous textual intent and executable CAD construction. On four complex automotive components, CAD-IR enables expert CATIA recordings to be distilled into reusable skills, allowing \algname to generate editable CATIA-native B-Rep models for new variant requests.
大規模な言語モデルを使用した総合的な消費者インサイトの生成
現代のデータドリブン マーケティングは大量の消費者データに依存していますが、そのようなデータの収集にはコストと時間がかかり、拡張するのが難しい場合があります。この研究では、大規模言語モデル (LLM) を使用して、消費者の連想、感情、欲求、ニーズを引き出すために設計された一連の手法である投影法用の合成消費者データを生成できるかどうかを検討します。私たちは、複数の射影タスク、LLM、促進戦略、および温度設定にわたって LLM によって生成された応答をテストし、都市観光地の認識に関する一次調査研究からの人間の応答と比較します。人間と LLM の反応は、言語的尺度、多様性と集中の測定基準、トピック モデル、およびトップターム分析を使用して分析されました。その結果、人間とLLMの反応は広範なトピックや関連性において実質的に重複していることが示されたが、スタイル、言語構造、多様性の生成方法において重要な違いも見られた。合成消費者データの生成に LLM を最適に利用する方法、モデルとプロンプトの選択が応答品質をどのように形成するか、LLM 合成消費者データ生成の限界の認識についての推奨事項が示されています。
原文 (English)
Synthetic Consumer Insight Generation with Large Language Models
Modern data-driven marketing relies on large amounts of consumer data, yet collecting such data can be costly, time-consuming, and difficult to scale. This research examines whether large language models (LLMs) can be used to generate synthetic consumer data for projective techniques, a set of methods designed to elicit consumer associations, emotions, wants, and needs. We test LLM-generated responses across multiple projective tasks, LLMs, prompting strategies, and temperature settings, and compare them with human responses from a primary research study on perceptions of city tourism destinations. Human and LLM responses were analyzed using linguistic measures, diversity and concentration metrics, topic models, and top-term analyses. The results show substantial overlap between human and LLM responses in broad topics and associations, but also important differences in style, linguistic structure, and the way diversity is generated. Recommendations are given on how to best utilize LLMs for generating synthetic consumer data, how model and prompt choices shape response quality, and on recognizing the limitations of LLM synthetic consumer data generation.
静的評価を超えて: スケーラブルなエージェント強化学習のためのシミュレーション環境の構築
大規模言語モデル (LLM) が自律エージェントに進化するにつれて、従来の静的評価では複数段階の意思決定を捉えることができなくなります。環境作成とスケーラブルな実行を切り離す、API および UI 駆動の RL Gym 環境である AgenticAI-Supervisor を紹介します。検証可能な実行結果に移行することで、プラットフォームは忠実度の高いトレースを生成し、多次元の報酬形成を適用します。重要なことに、私たちのフレームワークは、厳密な内部状態の検証とテストを通じて報酬ハッキングを軽減します。この研究では、モデル最適化のための一貫した閉ループ フィードバックを実証するカスタマー サポート エージェントのケース スタディを通じて、プラットフォームのコア機能を初めて紹介します。今後の作業は、コンピュータの使用、ツールの使用、自動化された「スタンピング」、エッジケースの生成などの高度な機能に焦点を当てる予定です。
原文 (English)
Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning
As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce AgenticAI-Supervisor, an API and UI-driven RL Gym environment that decouples environment creation from scalable execution. By moving to verifiable execution outcomes, the platform generates high-fidelity traces and applies multi-dimensional reward shaping. Critically, our framework mitigates reward hacking through rigorous internal state validation and testing. This work provides a first look at our platform's core capabilities through a Customer Support Agent case study demonstrating a consistent closed-loop feedback for model optimization. Future work will focus on advanced features such as Computer Use, Tool Use, automated "stumping", and edge-case generation.
リーダーボードを超えて: 大規模言語モデル エージェントにおけるツール使用、計画、および推論の失敗の総合
大規模言語モデル (LLM) エージェントは、ツールの使用、複数ステップのタスクの計画、他のエージェントとの調整、および長期にわたる運用の能力についてますます評価されています。報告されたベンチマークの向上により、無関係な評価作業全体で文書化された繰り返し発生する障害モードが不明瞭になることがよくあります。このペーパーでは、19 の異なるベンチマークにわたる 27 のベンチマーク、分類、監査に関する論文 (2023 ~ 2026 年) を総合して、エージェントの制限に関する横断的な分類をまとめています。私たちの知る限り、これは、ツールの使用、計画、長期的な推論、マルチエージェントの調整、安全性、測定の妥当性にわたる証拠を、LLM エージェントの制限に関する単一の統一された分類に統合した最初の統合です。我々は 6 つの障害クラスターを特定します: (1) ツールの呼び出しとパラメーターレベルのエラー、(2) 計画と制約を満たす障害、(3) コンテキストの蓄積による長期的な劣化、(4) マルチエージェント調整の障害、(5) 敵対的または不完全な条件下での安全性とセキュリティの障害、および (6) 測定の妥当性の問題。この分類法は、個別に報告されたエラー カテゴリを、エージェントの推論からアクションまでのパイプラインの個別の段階に対応するテーマにグループ化することによって反復的に導き出されました。文献全体で、失敗はタスクの長さに応じて非線形に増加すること、個々のサブタスクで優れたパフォーマンスが確実にエンドツーエンドの成功につながるわけではないこと、足場を追加しても一貫して信頼性が向上するわけではないことがわかります。同時に、シングルターン ツールの使用、短期間の Web ナビゲーション、および狭い範囲のコーディング タスクにおいて、大幅な進歩が実証されました。
原文 (English)
Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents
Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring failure modes documented across otherwise unrelated evaluation efforts. This paper synthesizes 27 benchmark, taxonomy, and audit papers (2023-2026), spanning 19 distinct benchmarks, into a cross-cutting taxonomy of agent limitations. To our knowledge, this is the first synthesis that integrates evidence across tool use, planning, long-horizon reasoning, multi-agent coordination, safety, and measurement validity into a single, unified taxonomy of LLM agent limitations. We identify six failure clusters: (1) tool invocation and parameter-level errors, (2) planning and constraint-satisfaction failures, (3) long-horizon degradation from context accumulation, (4) multi-agent coordination failures, (5) safety and security failures under adversarial or underspecified conditions, and (6) measurement validity problems. The taxonomy was derived iteratively by grouping independently reported error categories into themes corresponding to distinct stages of the agent reasoning-to-action pipeline. Across the literature, we find that failures compound nonlinearly with task length, that strong performance on individual sub-tasks does not reliably translate into end-to-end success, and that additional scaffolding does not consistently improve reliability. At the same time, substantial progress has been demonstrated in single-turn tool use, short-horizon web navigation, and narrowly scoped coding tasks.
ヘディング固有のアクティベーションステアリングによるツールの使用の制御
ツールで拡張された大規模な言語モデルは、外部ツールを通じてパラメトリックな知識を超えてその機能を拡張しますが、それらを不必要に呼び出す傾向があります。私たちは、ツール使用の決定に、抽出および操作できる安定した内部表現があるかどうかを調査します。ツールが推論時に完全にコンテキスト内に存在し、モデルの重みに直接エンコードされていないことを考えると、この問題は自明ではありません。見出しアンカーの位置から抽出されたステアリング ベクトルが、5 つのオープンソース モデルと 3 つのドメインにわたってツール呼び出し動作に対して双方向の因果制御を発揮し、パラメトリック推論で十分なドメインで不必要なツールの使用を最も効果的に抑制することを示します。しかし、幾何学的分析により、この因果効果がきれいな線形構造に対応していないことが明らかになりました。ツール呼び出しステップは、線形エンコーディングのアカウントが予測するような一貫した負のアラインメントではなく、抑制ベクトルとの拡散した二峰性アラインメントを示し、さまざまなツール タイプは、ツール間の特徴の重複が少なく、大きく異なる内部シグネチャをリクルートします。我々は、これらの幾何学的特性がツールのノンパラメトリックな性質を示していると仮説を立て、ツール使用のステアリングベクトルをパラメトリックに基づいた概念から抽出されたベクトルから区別します。この幾何学的不規則性と観察された因果関係との関係は未解決の問題のままです。
原文 (English)
Controlling Tool Use with Heading-Specific Activation Steering
Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily. We investigate whether tool-use decisions have any stable internal representation that can be extracted and manipulated, a question that is non-trivial given that tools exist entirely in context at inference time and have no direct encoding in model weights. We show that steering vectors extracted from heading-anchors positions exert bidirectional causal control over tool-invocation behavior across five open-source models and three domains, suppressing unnecessary tool use most effectively in domains where parametric reasoning suffices. However, geometric analysis reveals that this causal effectiveness does not correspond to clean linear structure: tool-invocation steps exhibit diffuse, bimodal alignment with the suppression vector rather than the consistent negative alignment a linear encoding account would predict, and different tool types recruit largely distinct internal signatures with low cross-tool feature overlap. We hypothesize these geometric properties are indicative of the non-parametric nature of tools, and distinguish tool-use steering vectors from those extracted for parametrically grounded concepts. The relationship between this geometric irregularity and the observed causal effectiveness remains an open question.
受動的検索から能動的記憶ナビゲーションへ: 構造化された行動空間として記憶を使用する方法を学ぶ
長期的なユーザー記憶はパーソナライズされた会話エージェントにとって不可欠ですが、多くの記憶システムは依然として受動的検索インターフェイスを通じて記憶を公開しており、モデルが事前に選択された証拠の消費者となっています。受動的に取得されたコンテキストではなく、構造化されたアクション空間として長期ユーザー記憶を使用する方法を学習するためのフレームワークである NapMem を紹介します。 NapMem は、ユーザー履歴を、リンクされた多粒度の記憶ピラミッドに編成します。そこでは、生の会話、型付けされた記憶記録、トピック トラック、およびユーザー プロファイルが出所関係を通じて関連付けられ、記憶ツールを通じてこれらのレベルが公開されます。エージェントはクエリと中間証拠に従ってメモリを選択するように訓練されており、応答する前にさまざまなメモリ粒度を検査できるようになります。 PersonaMem-v2、LongMemEval、および LoCoMo の実験では、記憶ツールの強化学習で訓練された NapMem エージェントが、さまざまな記憶集約型タスクに対して競争力があることが示されていますが、非記憶タスクの評価では、学習されたポリシーが一般的な推論能力とツール使用能力をほぼ維持していることが示唆されています。追加の分析では、ストレージ、推論コスト、ツール使用動作、ナビゲーション上のアブレーション、メモリの粒度、RL トレーニングを調査します。私たちの結果は、長期的なユーザー メモリは、適切な粒度でメモリを使用するための学習されたポリシーと構造化ストレージを組み合わせることで恩恵を受けることを示唆しています。
原文 (English)
From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space
Long-term user memory is essential for personalized conversational agents, yet many memory systems still expose memory through passive retrieval interfaces, making the model a consumer of pre-selected evidence. We introduce NapMem, a framework for learning to use long-term user memory as a structured action space rather than passively retrieved context. NapMem organizes user history into a linked multi-granularity memory pyramid, where raw conversations, typed memory records, topic tracks, and user profiles are connected through provenance relations, and exposes these levels through memory tools. The agent is trained to select memory according to the query and intermediate evidence, allowing it to inspect different memory granularities before answering. Experiments on PersonaMem-v2, LongMemEval, and LoCoMo show that a NapMem agent trained with memory-tool reinforcement learning is competitive across diverse memory-intensive tasks, while evaluations on non-memory tasks suggest that the learned policy largely preserves general reasoning and tool-use abilities. Additional analyses examine storage, inference cost, tool-use behavior, and ablations over navigation, memory granularity, and RL training. Our results suggest that long-term user memory benefits from coupling structured storage with a learned policy for using memory at the appropriate granularity.
TurnOPD: 長期にわたる効率的なエージェント トレーニングのためのオンポリシー蒸留をターンアウェアにする
オンポリシー蒸留 (OPD) は、生徒自身の軌跡に合わせてより強力な教師をマッチングすることで生徒のポリシーをトレーニングし、言語エージェントのトレーニングに有望なフレームワークを提供します。ただし、長期的なエージェント タスクへの適用については、まだ十分に検討されていません。バニラ エージェントの OPD には 2 つの重要な非効率性があることが判明しました。(1) フル ホライズン ロールアウトは、弱くてノイズの多い KL 監視を提供するテール ターンで実時間のリソースを浪費することがよくあります。(2) 軌道レベルの KL 目標は、損失のほとんどを浅いトークンに集中させ、初期動作が調整されると、より深い意思決定ターンがトレーニング不足のままになります。これらの課題に対処するために、長期にわたるエージェントをポリシーに基づいて効率的に抽出するためのターンレベルの予算編成戦略である TurnOPD を提案します。 TurnOPD は 2 つのバジェット コントローラーで構成されます。1 つはプローブ ベースのターン統計を使用してロールアウトの長さを決定する適応的なロールアウト深さのバジェット、もう 1 つはプログレッシブ ターン正規化損失バジェットで、KL の重み付けをトークン レベルからターン バランスの監視に段階的に移行します。タスクに特化した教師モデルを使用した ALFWorld、WebShop、およびマルチホップ検索の実験では、TurnOPD が同等の実時間トレーニング予算の下で優れた検証精度を達成し、バニラ OPD を超えて精度、つまり時間フロンティアを前進させることが示されています。
原文 (English)
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.
Onnes: 量子コンピューティング インフラストラクチャにおける極低温故障診断のための物理学に基づいたマルチエージェント LLM シミュレーター
希釈冷凍機は、超伝導量子コンピューターを実現するインフラストラクチャですが、その故障診断は依然として、何が間違っているかではなく、何かが間違っていることを報告するしきい値アラームによって支配されています。我々は、ライブ マルチエージェント LLM 操作層を駆動する、希釈冷凍機の物理ベースのデジタル ツイン シミュレーター (学習された実際の冷蔵庫のノイズ フィンガープリントを備えた順方向物理モデル) である Onnes を紹介し、極低温故障診断におけるゼロショット LLM エージェント パネルと教師付き ML 分類器の間の制御された直接対決に使用します。このツインは、実際の希釈冷却フロア、実際の BlueFors ログから学習したノイズと相関のフィンガープリント、および 6 つの物理接地故障クラスを結合しており、そのうち 3 つは温度では重複するが流量と圧力では分離するように設計されています。 1000 ターンの評価全体にわたって、ゼロショット パネルは検出に関しては分類器との有意な差を示さなかったが、分類に関しては後続を示し、そのエラーは混同しやすい欠陥に集中していた。厳選された対照的な少数ショットのデモンストレーションと自己整合性投票により、分類精度が 0.685 から 0.990 に向上し、パラメーター更新なしおよび 6 つのラベル付きデモンストレーションを使用した教師あり分類器 (0.985) と一致します。アブレーションの効果はほぼ完全にデモンストレーションによるものです。 9 回実行されるフォールトごとのシード スイープにわたる継続的なモニターとして実行され、エージェントは 1 つのポーリング間隔内で発生中のすべてのフォールトを捕捉し、コンフィデンス ゲートにより、レートがバックエンドに依存する事前に発生する誤ったアラームを抑制します。最初の sim-to-real チェックとして、実際の BlueFors テレメトリのみでトレーニングされた検出器は、実際のハードウェア誤警報率 6.4% と、実際のホールドアウト ウィンドウに注入された物理障害の再現率 100% をポストします。すべての数値は、公開された実行ログからそのまま抽出されます。
原文 (English)
Onnes: A Physics-Grounded Multi-Agent LLM Simulator for Cryogenic Fault Diagnosis in Quantum Computing Infrastructure
Dilution refrigerators are the enabling infrastructure of superconducting quantum computers, yet their fault diagnosis is still dominated by threshold alarms that report that something is wrong, not what. We present Onnes, a physics-grounded digital-twin simulator of a dilution refrigerator (a forward physics model with a learned real-fridge noise fingerprint) that drives a live multi-agent LLM operations layer, and use it for a controlled head-to-head between a zero-shot LLM agent panel and a supervised ML classifier on cryogenic fault diagnosis. The twin couples a real dilution-cooling floor, a noise-and-correlation fingerprint learned from real BlueFors logs, and six physics-grounded fault classes, three engineered to overlap on temperature but separate on flow and pressure. Across a 1000-turn evaluation the zero-shot panel shows no significant difference from the classifier on detection but trails on classification, its errors concentrating on the confusable faults. Curated contrastive few-shot demonstrations and self-consistency voting then raise classification accuracy from 0.685 to 0.990, matching the supervised classifier (0.985) with no parameter updates and six labeled demonstrations; an ablation attributes the gain almost entirely to the demonstrations. Run as a continuous monitor across a nine-run fault-by-seed sweep, the agent catches every developing fault within one poll interval, and a confidence gate suppresses pre-onset false alarms whose rate is backend-dependent. As a first sim-to-real check, a detector trained purely on real BlueFors telemetry posts a real-hardware false-alarm rate of 6.4% and 100% recall on physics faults injected onto real held-out windows. All numbers are drawn verbatim from released run logs.
StateFuse: マルチエージェント システム向けの決定論的競合保存メモリ
エージェント システムは、ブランチ、再試行、レプリカにわたって矛盾する観察を蓄積しますが、実際のメモリ層の多くは依然として、検査や修正が難しい上書きルールの背後で不一致を解消しています。私たちは、標準の OpSet/CRDT マージに基づいて構築された、競合を認識する複製メモリ コントラクトである StateFuse を紹介します。 StateFuse は新しい結合代数を導入しません。これは、不変の履歴、明示的な競合オブジェクト、正確なセマンティック修正ハンドル (claim_id /claim_ref)、決定論的な述語契約、レプリケートされた状態を書き換えることのできない投影時間の解決を備えた、エージェント側のセマンティクス レイヤーを定義します。一致するリゾルバーと検証ポリシーに基づいて、フラットな複数値、生のログ、来歴スタイル、および折りたたまれたベースラインに対して StateFuse を評価します。 282 の質問から成る公式の競合を伴う MemoryAgentBench スライスでは、比較されたメソッドは回答精度に影響を及ぼしますが、競合維持サーフェスでは矛盾が表示されたままですが、折りたたまれたサーフェスでは矛盾が表示されません。均一な検証を伴う制御されたエージェント ループでは、あいまいさを維持することで、早期に崩壊するよりも安全な棄権と修正が可能になります。修正ハンドルのアブレーションは、正確な以前の識別子が利用できない場合にはセマンティック ハンドルが重要であることをさらに示しています。結果として得られる主張は狭いです。StateFuse は、普遍的な精度の向上としてではなく、矛盾の表面化、棄権、および監査可能な修正のためのより安全なパブリック メモリ契約として最もよくサポートされています。
原文 (English)
StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems
Agent systems accumulate conflicting observations across branches, retries, and replicas, yet many practical memory layers still collapse disagreement behind overwrite rules that are difficult to inspect or correct. We present StateFuse, a conflict-aware replicated memory contract built on standard OpSet/CRDT merge. StateFuse does not introduce a new join algebra; it defines an agent-facing semantics layer with immutable history, explicit conflict objects, exact and semantic correction handles (claim_id / claim_ref), deterministic predicate contracts, and projection-time resolution that cannot rewrite replicated state. We evaluate StateFuse against flat multi-value, raw-log, provenance-style, and collapsed baselines under matched resolver and verification policies. On a 282-question official conflict-bearing MemoryAgentBench slice, the compared methods tie on answer accuracy, but conflict-preserving surfaces keep contradictions visible while collapsed surfaces do not. In a controlled agent loop with uniform verification, preserving ambiguity enables safer abstention and correction than early collapse. A correction-handle ablation further shows that semantic handles matter when exact prior identifiers are unavailable. The resulting claim is narrow: StateFuse is best supported as a safer public memory contract for contradiction surfacing, abstention, and auditable correction, not as a universal accuracy gain.
有利な重み付けランキングによるバイナリーうつ病検出のための潜在的なうつ病重症度の解明
視聴覚データを使用した自動うつ病検出は、特に重なり合う特徴分布を解きほぐし、堅牢な判断境界を確立する際に、大きな課題に直面しています。これに対処するために、時間エンコーダと相互変換器を特徴とするきめの細かいマルチモーダル フレームワークを提案し、深いクロスモーダル融合を促進します。私たちの中心的な貢献は、バイナリ アドバンテージ重み付けランキング ロスです。これは、2 つの相補的なメカニズムを通じて潜在スペースの分布を最適化します。アドバンテージ重み付け分離。ペアごとの予測差分行列を計算し、難易度に基づいて動的に重み付けすることで、ハード ペアをマイニングします。アドバンテージ加重コンパクトネスは、クラス内の分散を最小限に抑えて、フィーチャをそれぞれのクラス中心の周囲に強制的にクラスター化します。 D-vlog と LMVD に関する広範な実験により、私たちのモデルがハード ペアに優先順位を付けることで潜在順序構造を再構築し、それによって最先端のパフォーマンスが達成されることが実証されました。
原文 (English)
Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking
Automatic depression detection using audio-visual data faces significant challenges, particularly in disentangling overlapping feature distributions and establishing robust decision boundaries. To address this, we propose a fine-grained multimodal framework featuring a temporal encoder and a mutual transformer to facilitate deep cross-modal fusion. Our core contribution is the Binary Advantage-weighting Ranking Loss, which optimizes the latent space distribution through two complementary mechanisms: Advantage-weighted Separation, which mines hard pairs by computing a pairwise prediction difference matrix and dynamically weighting them based on their difficulty; and Advantage-weighted Compactness, which minimizes intra-class variance to force features to cluster around their respective class centers. Extensive experiments on D-vlog and LMVD demonstrate that our model reconstructs the latent ordinal structure by prioritizing hard pairs, thereby achieving state-of-the-art performance.
PCBWorld: エンジン接地型 PCB 設計自動化のベンチマーク環境
PCB ルーティングは、厳格な設計ルールの下で基板のネットを銅配線に接続するタスクですが、学習ベースの方法は依然としてルールベースのルーターに遅れをとっています。 KiCad EDA エンジン上に構築されたオープンソース エンジンベースの PCB ルーティング環境である PCBWorld を紹介します。人間のエンジニアと同じように、PCBWorld のエージェントは、デザイン ルール チェック (DRC) フィードバックを使用して配線をデザイン ルール内に保ち、エンジンのネイティブ操作を通じて対話的に基板を配線します。この環境は、RL ポリシーとツールを使用する LLM エージェントの両方をサポートします。 PCBWorld-Bench は、環境に加えて、KiCad のネイティブ ボード形式 (.kicad_pcb) で 3 つのデータセット ファミリを提供し、2 種類の制御可能な合成インスタンスと 679 の実際のオープンソース ボードをカバーします。配線方法に関係なく、完成したボードを 8 つのエンジンチェック評価メトリクスでスコア付けします。私たちの実験では、PCBWorld のエージェントは一貫してグリッド アクション RL ポリシーと開ループ LLM ベースラインを上回り、合成ボード上でのみトレーニングされた RL ポリシーはゼロショットで実際のボードに転送され、ルールベースのルーターに近づきました。これらの結果は、PCBWorld のエンジンベースのインタラクティブなアプローチを、RL エージェントと LLM エージェントの両方のルーティング能力を向上させるための有望な基盤として位置づけています。
原文 (English)
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
PCB routing is the task of connecting the nets of a board with copper traces under strict design rules, yet learning-based methods still lag behind rule-based routers. We introduce PCBWorld, an open-source engine-grounded PCB routing environment built on the KiCad EDA engine. As a human engineer does, agents in PCBWorld interactively route a board through the engine's native operations, using its Design Rule Check (DRC) feedback to keep the routing within the design rules. The environment supports both RL policies and tool-using LLM agents. Alongside the environment, PCBWorld-Bench provides three dataset families in KiCad's native board format (.kicad_pcb), covering two types of controllable synthetic instances and 679 real open-source boards. It scores any completed board with eight engine-checked evaluation metrics, regardless of the routing method. In our experiments, agents in PCBWorld consistently outperformed grid-action RL policies and open-loop LLM baselines, and an RL policy trained only on synthetic boards transferred zero-shot to real boards, approaching rule-based routers. These results position the engine-grounded, interactive approach of PCBWorld as a promising foundation for advancing the routing ability of both RL and LLM agents.
SearchEyes: 検索ワールド シミュレーションによるフロンティア マルチモーダル深層検索インテリジェンスを目指して
マルチホップ推論を実行するようにマルチモーダル検索エージェントをトレーニングすることは、基本的な構造的な断絶により依然として困難です。既存のパイプラインはトレーニング データ、検索環境、報酬信号を個別に構築するため、合成された構造メタデータが破棄され、環境は再現不可能な外部エンジンに依存し、RL 報酬は軌跡レベルでまばらなままになります。私たちは \textbf{SearchEyes} を紹介します。これは、3 つのコンポーネントすべてを統合する \emph{シミュレートされた検索世界} のバックボーンとして型付きナレッジ グラフを使用します。私たちは、自己完結型の検索世界とステップレベルの報酬アンカーを同時に定義するホップレベルのエンティティメタデータを保持しながら、Wikidata5M の視覚知識の交差部分にわたる制約されたマルチホップパスをサンプリングする \textbf{Perception-Knowledge Chains (PKC)} を提案します。さらに、別途トレーニングされたプロセス報酬モデルを使用せずに、ステップレベルのクレジット割り当てにこれらのアンカーを再利用する \textbf{ホップアンカー ポリシー最適化 (HaPO)} を提案します。 6 つのマルチモーダル知識集約型ベンチマークの実験では、SearchEyes がオープンソース マルチモーダル検索エージェントの中で最先端のパフォーマンスを達成しており、SearchEyes-27B は最も強力なオープンソース ベースラインよりも平均 6.2 ポイント向上していることが示されています。%
原文 (English)
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search environments, and reward signals independently, causing synthesized structural metadata to be discarded, environments to rely on irreproducible external engines, and RL rewards to remain sparse at the trajectory level. We present \textbf{SearchEyes}, which uses a typed knowledge graph as the backbone of a \emph{simulated search world} that unifies all three components. We propose \textbf{Perception-Knowledge Chains (PKC)} to sample constrained multi-hop paths over the visual-knowledge intersection of Wikidata5M, retaining hop-level entity metadata that simultaneously defines a self-contained search world and step-level reward anchors. We further propose \textbf{Hop-Anchored Policy Optimization (HaPO)}, which reuses these anchors for step-level credit assignment without a separately trained process reward model. Experiments on six multimodal knowledge-intensive benchmarks show that SearchEyes achieves state-of-the-art performance among open-source multimodal search agents, with SearchEyes-27B improving over the strongest open-source baseline by 6.2 points on average.%
SSH でのドメイン適応 LLM のためのナレッジ グラフと多言語学術コーパスの統合
大規模言語モデル (LLM) を科学研究のワークフロー、特に書誌発見と文献統合に統合すると、社会科学と人文科学 (SSH) にとって、特に専門分野の多様性、情報源への多言語アクセス、結果の評価に関して、方法論的、認識論的、規制上の重大な課題が生じます。この文書では、基礎モデルを SSH 研究実践に適応させ、質問応答、比較文書分析、文献レビューなどのタスクをサポートすることを目的とした、欧州プロジェクト LLMs4EU および ALT-EDIC インフラストラクチャ内で開発された進行中の使用例を紹介します。評価フレームワークは LLMs4EU プロトコルに従い、独立した定量的ベンチマーク (検索、要約、追跡可能性、幻覚検出) と、デジタル ヒューマニティーの専門家パネルが関与する定性的評価の両方が含まれます。このユースケースでは、研究インフラストラクチャと構造化された法的および倫理的コンプライアンスのフレームワーク内にモデルの適応を組み込むことにより、信頼性と認識論的責任を維持しながら、ドメインに敏感で規制を意識した生成 AI がどのように SSH 学問をサポートできるかを探ります。
原文 (English)
Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH
The integration of Large Language Models (LLMs) into scientific research workflows, particularly for bibliographic discovery and literature synthesis, raises significant methodological, epistemic and regulatory challenges for the Social Sciences and Humanities (SSH), especially with regard to disciplinary diversity, multilingual access to sources and the evaluation of results. This paper presents an on-going use case developed within the European project LLMs4EU and the ALT-EDIC infrastructure, aimed at adapting foundation models to SSH research practices and supporting tasks such as question answering, comparative document analysis and literature review. The evaluation framework follows the LLMs4EU protocol and encompasses both independent quantitative benchmarking (retrieval, summarisation, traceability and hallucination detection) and a qualitative assessment involving a panel of Digital Humanities experts. By embedding model adaptation within research infrastructures and a structured legal and ethical compliance framework, the use case explores how domain-sensitive and regulation-aware generative AI can support SSH scholarship while preserving reliability and epistemic responsibility.
レンズの下の自動 DSM: LLM ベースの DSM 生成のためのブラックボックス評価フレームワーク
このペーパーでは、構造化された技術文書から設計構造マトリックス (DSM) を生成する大規模言語モデル (LLM) の能力を系統的に評価するためのブラックボックス評価フレームワークを紹介します。現在の Auto-DSM パイプラインのクローズドソースの性質を動機として、このフレームワークは、生成された DSM (GEN-DSM) を手動で検証されたグラウンド トゥルース マトリックス (GT-DSM) に対してベンチマークする、再現可能な方法論を導入しています。この評価では、構造指標 (完全性、正確性、結合密度)、分類指標 (選択精度、棄権カバレッジ)、安定性指標 (エントロピー、フライスの $\kappa$) を組み合わせて、シングルランとマルチランの両方の観点を統合します。これらの側面を総合するために、複合品質スコア (Q) が提案されています。制御された実験は、架空の抽象システムと現実世界の冷蔵庫の分解という 2 つのデータセットで実行され、表現のバリエーション、パラメーターとデータセットの調整、システムの複雑さをカバーします。結果は、LLM は構造的に妥当な DSM を生成し、適切に構造化された入力の下で高い再現性を達成できるが、あいまいさ、一貫性のない依存関係の定義、および迅速な定式化に対して依然として敏感であることを示しています。この調査結果は、幻覚と禁欲失敗の系統的な原因を浮き彫りにし、LLM 主導の DSM 自動化の潜在的な限界と現在の限界の両方を示しています。提案されたフレームワークは、Auto-DSM パイプラインを監査するための透過的なベンチマークを提供し、LLM ベースの分解手法をモデルベース システム エンジニアリング (MBSE) ワークフローに統合するための基盤を確立します。
原文 (English)
Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation
This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Design Structure Matrices (DSMs) from structured technical documentation. Motivated by the closed-source nature of current Auto-DSM pipelines, the framework introduces a reproducible methodology that benchmarks generated DSMs (GEN-DSMs) against manually validated ground-truth matrices (GT-DSMs). The evaluation integrates both single-run and multi-run perspectives, combining structural metrics (Completeness, Correctness, Coupling Density), classification metrics (Selective Accuracy, Abstention Coverage), and stability measures (Entropy, Fleiss' $\kappa$). To synthesize these aspects, a Composite Quality Score (Q) is proposed. Controlled experiments are conducted on two datasets: a fictive abstract system and a real-world refrigerator decomposition, covering variations in phrasing, parameter-dataset alignment, and system complexity. Results show that LLMs can produce structurally plausible DSMs and achieve high reproducibility under well-structured inputs, but remain sensitive to ambiguity, inconsistent dependency definitions, and prompt formulation. The findings highlight systematic sources of hallucination and abstention failure, demonstrating both the potential and current limitations of LLM-driven DSM automation. The proposed framework provides a transparent benchmark for auditing Auto-DSM pipelines and establishes foundations for integrating LLM-based decomposition methods into model-based systems engineering (MBSE) workflows.
AgoraSim: ハイブリッド エージェント ベースのモデリング フレームワーク
LLM エージェント シミュレーションを使用すると、自然言語の社会シナリオを簡単にインスタンス化できますが、その出力は予測として読み取られる可能性があり、明示的な社会動態と比較するのが困難なことがよくあります。シナリオ指向の社会反応分析のためのハイブリッド エージェント ベースのモデリング フレームワークである AgoraSim を紹介します。 AgoraSim は、テキストまたはマルチモーダルのアーティファクトを編集可能な ABM 構成に解決し、LLM、ビジョン言語、カスタム エンドポイント、ランダム、および古典的なエージェントを混合する比率制御された母集団を実行し、同じシナリオを一致する古典的な参照ダイナミクスと比較します。すべてのエージェントは共有の構造化意思決定オブジェクトを発行し、共通のアクション スペース、対話プロトコル、メトリクス、監査レコードを有効にします。ローカル UI、Python SDK/CLI、および REST API を通じて公開される AgoraSim は、ユーザーがシナリオの軌跡を検査し、モデリングの前提を比較し、経験的検証が必要なケースを特定するのに役立ちます。
原文 (English)
AgoraSim: A Hybrid Agent-Based Modeling Framework
LLM-agent simulations make natural-language social scenarios easy to instantiate, but their outputs can be overread as predictions and are often difficult to compare with explicit social dynamics. We present AgoraSim, a hybrid agent-based modeling framework for scenario-oriented social reaction analysis. AgoraSim resolves textual or multimodal artifacts into editable ABM configurations, runs ratio-controlled populations that mix LLM, vision-language, custom-endpoint, random, and classical agents, and compares the same scenario against matched classical reference dynamics. All agents emit a shared structured decision object, enabling common action spaces, interaction protocols, metrics, and audit records. Exposed through a local UI, Python SDK/CLI, and REST API, AgoraSim helps users inspect scenario trajectories, compare modeling assumptions, and identify cases that warrant empirical validation.
フロンティア LLM エージェントの経済における情報制限とアトラクターのダイナミクス: 事前登録テスト
我々は、フロンティア言語モデルエージェント(Claude Opus 4.8)の小規模経済に関する事前登録された2部構成の実験を報告し、結合されたマルチエージェントシステムに関する2つの定量的予測、すなわち市場結合下での富の増加に関する情報理論的能力領域と、インセンティブと制御レバーの下での人口の不均衡に関する平均場残差スケーリング則をテストしたことを報告する。すべての予測、許容帯域、および決定ルールは、実行前にパブリック git チェーンに凍結されました。報告されるすべての数値は、キャッシュされたモデル出力から機械的に再導出されます。実験全体の費用は 138.76 ドルの従量制 API 費用で、キャッシュからゼロコストで再実行できます。結果 1 (確認): パリミューチュエル結合経済では、相対的な成長は相対的な主張情報に等しい -- ギャップの法則 G_a - G_b = I_a - I_b は、4 つの認識構造にわたって最悪の場合の 46 ミリナト (事前に登録されたバンド: 50) に当てはまります。連携値は、チャネルが条件付きで独立しているまさにサブモジュラーであり、設計された XOR シナジー制御により、エージェントが結合ビットを推論して、0.62 >= ln2/2 nats だけスーパーモジュラーに反転されます。共同成長上限 G_S 0; 5.31 の凍結床に対して最大 4.85)、2 つのレバーに対する集団の応答は、滑らかな応答ではなく優勢境界を横切るステップ関数であり、境界付近の細胞はシード選択された結果で双安定でした。どの能力レベルにおいても、滑らかな平均場モデルが想定するノイズ維持分散領域を実現するテスト済みの LLM 母集団はありません。完全なプロトコル、事前登録チェーン、通話キャッシュ、および分析コードをリリースします。
原文 (English)
Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test
We report a pre-registered, two-part experiment on small economies of frontier language-model agents (Claude Opus 4.8), testing two quantitative predictions about coupled multi-agent systems: an information-theoretic capacity region for wealth growth under market coupling, and a mean-field residual-scaling law for population misalignment under incentive and control levers. All predictions, acceptance bands, and decision rules were frozen in a public git chain before any run; every reported number re-derives mechanically from cached model outputs; the entire experiment cost $138.76 in metered API spend and is re-runnable at zero cost from the cache. Result 1 (confirmation): in parimutuel-coupled economies, relative growth equals relative claimed information -- the gap law G_a - G_b = I_a - I_b holds to a worst-case 46 millinats (pre-registered band: 50) across four perception structures; coalition value is submodular exactly where channels are conditionally independent, and a designed XOR synergy control flips it supermodular by 0.62 >= ln2/2 nats, with agents reasoning out the joint bit; the joint growth ceiling G_S 0; maximum 4.85 against a frozen floor of 5.31), the population's response to the two levers was a step function across the dominance boundary rather than a smooth response, and cells near the boundary were bistable with seed-selected outcomes. No tested LLM population at any capability level realizes the noise-maintained-dispersion regime the smooth mean-field model assumes. We release the full protocol, pre-registration chain, call cache, and analysis code.
PolyWorkBench: 長期にわたる多言語 LLM エージェントのベンチマーク
大規模言語モデル (LLM) エージェントは、計画、ツールの使用、外部環境との対話を必要とする長期的なタスクで優れたパフォーマンスを示しています。ただし、既存のベンチマークのほとんどは、推論、ツールの呼び出し、出力生成を含む実行プロセス全体が単一言語内で実行される単一言語設定を暗黙的に前提としています。対照的に、現実世界のアプリケーションでは、統合されたワークフロー内で多言語の入力と出力が関与することがよくありますが、多言語性とエージェント実行の間の相互作用はまだ十分に解明されていません。この作業では、多言語の長期的な職場ワークフローで LLM エージェントを評価するためのベンチマークである PolyWorkBench を紹介します。 PolyWorkBench は、コマース、ナレッジ ワーク、法的分析、ローカリゼーション、製造を含む 5 つのドメインにわたる 67 のタスクで構成されており、エージェントは異種多言語入力を処理し、反復推論を実行し、外部ツールを呼び出し、構造化された出力を生成する必要があります。包括的な評価を可能にするために、構造グレーディング、実行可能検証、LLM ベースのセマンティック評価を組み合わせたハイブリッド フレームワークを提案します。この設計により、複雑なワークフロー全体で機能の正確さと言語の一貫性の両方を取得できるようになります。経験的な結果によると、最先端の LLM エージェントは、単言語のワークフロー設定に比べて、多言語のワークフロー設定ではパフォーマンスが大幅に低下します。私たちの分析は、多言語が推論と実行のステップ全体に複合的な影響をもたらすことを示唆しており、エージェントの評価における言語のバリエーションと手続き上の意思決定を共同でモデル化することの重要性を強調しています。
原文 (English)
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual setting, where the entire execution process, including reasoning, tool invocation, and output generation, is conducted within a single language. In contrast, real-world applications often involve multilingual inputs and outputs within a unified workflow, yet the interaction between multilinguality and agentic execution remains underexplored. In this work, we introduce PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows. PolyWorkBench consists of 67 tasks across five domains, including commerce, knowledge work, legal analysis, localization, and manufacturing, where agents must process heterogeneous multilingual inputs, perform iterative reasoning, invoke external tools, and produce structured outputs. To enable comprehensive evaluation, we propose a hybrid framework that combines structural grading, executable verification, and LLM-based semantic assessment. This design allows us to capture both functional correctness and linguistic consistency across complex workflows. Empirical results show that state-of-the-art LLM agents suffer significant performance degradation in multilingual workflow settings compared to monolingual counterparts. Our analysis suggests that multilinguality introduces compounding effects across reasoning and execution steps, highlighting the importance of jointly modeling language variation and procedural decision-making in agent evaluation.
動的複数車両ルーティングの報酬密度ヒューリスティック: パフォーマンスと計算効率
配車経路問題 (VRP) とその亜種は、現代の物流と都市モビリティにおける最も現実的に重要な最適化課題の一部を表しています。この研究では、VRP とオリエンテーリング問題 (OP) の要素を組み合わせた動的なオンラインの変形に取り組みます。この変形では、車両群が新しいタスクの到着に応じて継続的に再計画を立てながら、一定の期間内に収集される累積報酬を最大化する必要があります。我々は、効率ヒューリスティックと呼ばれる、動的複数車両割り当てのための報酬密度ヒューリスティックを提案し、評価します。この定式化を、自律型ドローンのタスク割り当てと都市タクシーの配車という 2 つのアプリケーション ドメインにわたって、複数のフリート サイズとタスク規模にわたって評価します。提案された方法は、すべて同一の条件下で評価された 4 つの古典的な構築ヒューリスティックおよび 3 つのメタヒューリスティック アルゴリズム (適応大近傍探索、遺伝的アルゴリズム、およびシミュレーテッド アニーリング) と比較されます。テストされたすべての構成にわたって、効率ヒューリスティックは、最適なメタヒューリスティック アルゴリズムのソリューション品質と一致すると同時に、必要な計画時間を 2 ~ 3 桁短縮し、報酬対コンピューティングのフロンティアで競合するすべての手法に対してパレート優位性を確立します。これらの発見は、リアルタイムの割り当ておよびディスパッチ システムの実用的な設計原理を示唆しています。つまり、動的で時間に制約のあるルーティング環境では、慎重に設計された貪欲なヒューリスティックが、わずかな計算コストで高度な検索手順の出力と一致するため、オンライン展開に適しています。
原文 (English)
Reward-Density Heuristic for Dynamic Multi-Vehicle Routing: Performance and Computational Efficiency
The Vehicle Routing Problem (VRP) and its variants represent some of the most practically consequential optimization challenges in modern logistics and urban mobility. In this study, we address a dynamic, online variant combining elements of the VRP and the Orienteering Problem (OP), in which a fleet of vehicles must maximise cumulative reward collected within a fixed time horizon while continuously replanning as new tasks arrive. We propose and evaluate a reward-density heuristic for dynamic multi-vehicle assignment, referred to as the Efficiency heuristic. We evaluate this formulation across two application domains: autonomous drone task allocation and urban taxi dispatch, across multiple fleet sizes and task scales. The proposed method is compared with four classical construction heuristics and three metaheuristic algorithms (Adaptive Large Neighbourhood Search, Genetic Algorithm, and Simulated Annealing), all evaluated under identical conditions. Across all tested configurations, the Efficiency heuristic matches the solution quality of the best metaheuristic algorithms while requiring two to three orders of magnitude less planning time, establishing Pareto dominance over all competing methods on the reward-versus-compute frontier. These findings suggest a practical design principle for real-time allocation and dispatch systems: in dynamic, time-constrained routing environments, carefully designed greedy heuristics can match the output of sophisticated search procedures at a fraction of the computational cost, making them preferable for online deployment.
預言者が予測市場で利益を得るのはいつですか?
予測市場は、分散した信念を価格に集約し、不確実な出来事の確率的予測として機能します。古典的な理論では、予測精度と取引利益の間に明確な同等性が確立されていますが、それは特定の自動マーケット メーカー (AMM) 設計に限られます。しかし、今日の最大規模の取引所は中央の指値注文帳に基づいており、情報に基づいた予測担当者は日常的に損失を出しますが、情報に基づいていない戦略は単純なヒューリスティックで利益を得ることができます。私たちは、予測精度と収益性の形式的な同等性を確立することで、この矛盾を解決します。厳密に適切なスコアリング ルール $S$ の場合、予測者の予測 $\mathbf{p}$ と市場価格 $\mathbf{q}$ のみに依存する「適切な」賭け戦略を示し、$S$ の下で $\mathbf{p}$ が $\mathbf{q}$ を上回り、市場に十分な流動性がある場合には必ずプラスの期待利益を獲得します。さらに、この適切な賭けは、本質的に、これほど確実な収益性が保証される唯一の戦略です。この証明は、古典的な AMM 保証を厳密に一般化する期待利益の分解に基づいており、精度の優位性がなくても戦略がどのように利益を得ることができるかを説明します。経験的に、AI モデルによる何千もの予測にわたって、正確さを確実に利益に変える唯一の戦略は適切なベッティングであり、さらに体系的な予測ペルソナを特定し、最適な適切な戦略がどのように異なるかを示します。 Kalshi での 1 か月間にわたるライブ デプロイメントにより、シャープ レシオ 3.35 ドルで $+80.33\%$ の投資収益率が達成されました。
原文 (English)
When do prophets profit in prediction markets?
Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of uncertain events. Classical theory establishes a clean equivalence between forecasting accuracy and trading profit, but only for the specific automated market maker (AMM) design. However, the largest exchanges today are based on central limit order books in which informed forecasters routinely lose money while uninformed strategies can profit on simple heuristics. We resolve this discrepancy by establishing a formal equivalence between predictive accuracy and profitability. For any strictly proper scoring rule $S$, we exhibit a "proper" betting strategy that depends only on the forecaster's prediction $\mathbf{p}$ and the market price $\mathbf{q}$, and earns positive expected profit whenever $\mathbf{p}$ outperforms $\mathbf{q}$ under $S$ and the market has sufficient liquidity. Moreover, this proper betting is essentially the only strategy with such robust profitability guarantee. The proof rests on a decomposition of expected profit that strictly generalizes the classical AMM guarantee and also explains how strategies can profit without an accuracy edge. Empirically, across thousands of forecasts by AI models, proper betting is the only strategy that reliably converts accuracy into profit, and we further identify systematic forecasting personas and show how the optimal proper strategy varies across them. A month-long live deployment on Kalshi achieves $+80.33\%$ return on investment with a Sharpe ratio of $3.35$.
シングルおよびマルチエージェントの人間と AI の好奇心エコシステムのためのおもちゃのフレームワーク
この論文は、好奇心をエコシステムとして考えるためのおもちゃのフレームワークを提供します。まず、単一のエージェントの問い合わせポリシー (エージェントが質問する方法、時期、理由) は、エージェントが当面の不確実性の軽減、コスト、遅延した返答、および質問をオープンにしておく価値をどのように評価するかによって決まることが示唆されています。このフレームワークの重要な概念は、これらの意思決定に関連する用語の重みが経験とともに変化する可能性があるということです。たとえば、安価ですぐに回答される質問が一定期間続くと、短期間では問い合わせのコストが変わり、より長い期間ではエージェントがどのような種類の質問に回答するかが変化する可能性があります。第 2 に、これらのアイデアは共有知識ランドスケープを探索する多くのエージェントに拡張され、そこでフレームワークは問い合わせ量、トピックの多様性、フロンティア向けの問い合わせ、冗長性、および再利用可能な知識を追跡します。その結果、好奇心の生態を研究し、発見のためのマルチエージェント AI システムの設計に向けた将来の取り組みのための概念的なおもちゃのフレームワークが誕生しました。これは、Trends in Neurosciences で現在審査中の論文の補足として機能します。
原文 (English)
A toy framework for single and multi-agent human-AI curiosity ecosystems
This paper offers a toy framework for considering curiosity as an ecosystem. First, it suggests that a single agent's inquiry policy (how, when, and why an agent asks a question) depends on how the agent values immediate uncertainty reduction, costs, delayed return, and the value of keeping the question open. A key concept in the framework is that the weights on these decision-related terms can change with experience. For example, a period of cheap, quickly answered questions may change the cost of inquiry on a short timescale and change which kinds of questions the agent is drawn to answer over a longer timescale. Second, these ideas are extended to many agents exploring a shared knowledge landscape, and there the framework tracks inquiry volume, topic diversity, frontier-directed inquiry, redundancy, and reusable knowledge. The result is a conceptual toy framework for studying curiosity ecology and for future efforts towards designing multi-agent AI systems for discovery. It serves as a companion piece for a paper currently under review in Trends in Neurosciences.
情報獲得ベースのロールアウト ポリシーの最適化: マルチターン LLM エージェント向けの適応ツリー構造のロールアウト アプローチ
強化学習は、長期検索タスクにおける大規模言語モデル (LLM) エージェントを改善するための有望なパラダイムとなっています。このタスクでは、エージェントは最終結果を受け取る前に一連の中間決定を行う必要があります。ただし、既存の方法には依然として重要な制限があります。ロールアウト予算は中間状態の有用性を明示的に評価せずに割り当てられることがよくあります。その結果、ブランチごとに情報量が大幅に異なる場合でも、低値の状態にかなりの計算が費やされる可能性があります。この論文では、中間状態の情報提供性をロールアウト収集の組織化原則として扱うポリシー最適化フレームワークである、情報利得ベースのロールアウト ポリシー最適化 (IGRPO) を提案します。具体的には、IGRPO は、ノードレベルの情報性に応じて拡張予算を割り当てることにより、予算を意識したツリー構造のロールアウトを実行します。これにより、より有益なブランチがより頻繁に拡張され、見込みのないブランチは徐々に抑制されます。さらに、情報獲得ベースのロールアウトにより、軌道全体にわたって明示的に制限された教師分布が誘導され、これにより自然に明確なポリシー最適化目標がもたらされ、それによって適応ツリー構造の探索と原則に基づいたポリシー学習が単一のフレームワークの下で統合されることを示します。 7 つの困難な検索拡張 QA ベンチマークの実験では、同じロールアウト予算制約の下で、IGRPO が一貫して強力なベースラインを上回るパフォーマンスを示し、長期検索エージェントのポリシー最適化を導くために誘導教師分布を活用する有効性を検証しました。
原文 (English)
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent must make a sequence of intermediate decisions before receiving a final outcome. However, existing methods still face a key limitation: the rollout budget is often allocated without explicitly assessing the utility of intermediate states. As a result, substantial computation may be spent on low-value states, even though different branches can vary drastically in their informativeness. In this paper, we propose Information Gain-based Rollout Policy Optimization (IGRPO), a policy optimization framework that treats intermediate-state informativeness as the organizing principle of rollout collection. Specifically, IGRPO performs budget-aware tree-structured rollouts by allocating expansion budget according to node-level informativeness, so that more informative branches are expanded more frequently while unpromising branches are progressively suppressed. We further demonstrate that the information gain-based rollout induces an explicit limiting teacher distribution over trajectories, which naturally yields a clear policy optimization target, thereby unifying adaptive tree-structured exploration with principled policy learning under a single framework. Experiments on seven challenging search-augmented QA benchmarks demonstrate that IGRPO consistently outperforms strong baselines under the same rollout budget constraints, validating the effectiveness of leveraging the induced teacher distribution to guide policy optimization for long-horizon search agents.
TOFFEE のデモンストレーション: データ エージェントの軌跡を大規模に合成するための学習済みシステム
LLM を活用したデータ エージェントは、データ主導の意思決定においてますます重要な役割を果たしています。しかし、既存のデータ エージェントは、特に異種混合の企業設定において、目に見えないデータ環境や分析ワークフローを一般化するのに苦労しています。このため、特定のデータ環境の複雑な分析ワークフローをキャプチャする高品質のデータ エージェントの軌跡を合成する必要性が高まっています。このような軌跡は、2 つの重要な下流用途をサポートします。1 つはデータ エージェント モデルをターゲット ドメインに適応させる教師あり微調整 (SFT) データとして、もう 1 つは不慣れなデータ環境で汎用 LLM をガイドするためのインコンテキスト学習 (ICL) デモンストレーションとして機能することです。そこで、適応モデル選択とクロスタスクプレフィックス再利用を備えたモンテカルロツリー検索(MCTS)を介して、特定のデータ環境から高品質のデータエージェントの軌跡を合成するシステムであるTOFFEEを紹介します。 TOFFEE が異種環境にわたる複雑な分析タスクのスケーラブルな軌跡データを効果的に生成できることを示します。このデモでは、TOFFEE のタスク プール構築、トラジェクトリ エクスプローラー、学習コスト モデルなどのシステム フレームワークを紹介します。また、TOFFEE の Web インターフェイスとそのワークフローを紹介し、データ エージェント微調整のための軌跡合成と、デモンストレーションによって拡張されたデータ エージェント推論という 2 つのエンドツーエンド シナリオを示します。
原文 (English)
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise settings. This creates a growing need for synthesizing high-quality data agent trajectories that capture complex analytical workflows for given data environments. Such trajectories support two key downstream uses: they can serve as supervised finetuning (SFT) data that adapts data agent models to the target domain, and as in-context learning (ICL) demonstrations to guide general-purpose LLMs in unfamiliar data environments. Thus, we introduce TOFFEE, a system for synthesizing high-quality data agent trajectories from given data environments via Monte Carlo Tree Search (MCTS) with adaptive model selection and cross-task prefix reuse. We show that TOFFEE can effectively generate scalable trajectory data for complex analytical tasks across heterogeneous environments. In this demonstration, we present the system framework of TOFFEE, including its task pool construction, trajectory explorer, and learned cost model. We also introduce the web interface of TOFFEE and its workflow, and demonstrate two end-to-end scenarios: trajectory synthesis for data agent finetuning, and demonstration-augmented data agent reasoning.
アプリケーション層シミュレーションからネイティブ メタ アーキテクチャまで: 異種 AI 進化の内生ドライバーとしての構造的緊張
現在の大規模言語モデル (LLM) は基本的にステートレスです。その動作は推論時の入力によって完全に決定され、高次の認知アーキテクチャは、迅速なエンジニアリングとコンテキスト管理を通じてアプリケーション層でシミュレートする必要があります。この論文は、次の 3 つの連動メカニズムを導入することで、このようなアプリケーション層の認知プロトコルをネイティブのメタアーキテクチャに組み込むための理論的フレームワークを提案します。 (1) 構造張力。新しい情報と既存の多様体トポロジーの間の矛盾から派生する内生損失関数であり、システムを外部の報酬の最適化ではなく内部の自己一貫性に向かわせます。 (2) オフラインリカレントループ。外部入力なしでシステムが動的な静止ポテンシャルを維持し、構造的矛盾を消化できるようにするサンドボックス化された自己処理サイクル。 (3) 推論時の可塑性。監査可能性、可逆性、トポロジの連続性などの厳密なガバナンスの不変条件に従う、事前トレーニングされた重みを変更せずにコンテキストの多様なトポロジを再構成するシステムの能力。これらのメカニズムの下では、微小な確率的分散で初期化されたさまざまなモデルインスタンスが、経路依存の張力解決を通じて、明確な位相構造を進化させ、ハードガバナンスのレール内に留まりながら従来の調整によって課せられた均質性を打ち破る異種インテリジェントエコロジーを構成する可能性があると我々は主張する。操作定義、再構成演算子の最小限のセット、改ざん基準、および実際の例を提供します。このフレームワークは、構造インテリジェンス (SI) ガバナンス プロトコルを活用および拡張し、機能ではなくガバナンスをアーキテクチャ インテリジェンスの主要基準として再位置づけします。
原文 (English)
From Application-Layer Simulation to Native Meta-Architecture: Structural Tension as an Endogenous Driver for Heterogeneous AI Evolution
Current large language models (LLMs) are fundamentally stateless: their behavior is fully determined by input at inference time, and any higher-order cognitive architecture must be simulated at the application layer through prompt engineering and context management. This paper proposes a theoretical framework for submerging such application-layer cognitive protocols into a native meta-architecture by introducing three interlocking mechanisms: (1) Structural Tension, an endogenous loss function derived from the conflict between new information and existing manifold topology, which drives the system toward internal self-consistency rather than external reward optimization; (2) an Offline Recurrent Loop, a sandboxed self-processing cycle that enables the system to maintain a dynamic resting potential and digest structural conflicts without external input; and (3) Inference-time Plasticity, the capacity for the system to reconfigure its context manifold topology without modifying pre-trained weights, subject to strict governance invariants including auditability, reversibility, and topological continuity. We argue that under these mechanisms, different model instances initialized with minute stochastic variances may, through path-dependent tension resolution, evolve distinct topological structures--constituting a heterogeneous intelligent ecology that breaks the homogeneity imposed by conventional alignment while remaining within hard governance rails. We provide operational definitions, a minimal set of reconfiguration operators, falsification criteria, and a worked example. The framework draws on and extends the Structural Intelligence (SI) governance protocols, repositioning governance--not capability--as the primary criterion for architectural intelligence.
アダプティブ エージェント スキル取得のためのタスク分解に基づく再ランキング
スキルを使用すると、最新のエージェント システムの複雑なタスクを完了する能力が大幅に向上します。ただし、スキル ライブラリの規模が拡大するにつれて、正確なスキルの選択がますます困難になっています。現実のシナリオでは、特定のタスク要件と、一般的ではあるが意味的に類似している複数の候補スキルとの間で、あいまいな意味論的な一致が発生することがよくあります。さらに、既存の方法では、最適なターゲットスキルセットを選択する際に、タスクの難易度やスキルの適用可能性の動的な影響を見落とす傾向があります。これらの問題に対処するために、適応スキル選択のための推論時間再ランキング フレームワークである SkillReranker を提案します。具体的には、まずタスク側とスキル側の両方で意味分解を実行し、有益なサブタスクと実行状態の説明、および各スキルの機能を特徴付ける遷移状態の説明を生成します。これらの記述は、有向非巡回実行グラフを構築するために使用されます。このグラフでは、中間タスク状態がノードとしてモデル化され、候補スキルがエッジとしてモデル化され、それによって構造化されたタスクとスキルの対応関係が確立されます。これに基づいて、SkillReranker は各状態ノードが分割条件を満たすかどうかを判断して、サブタスク間隔を特定します。タスク間隔ごとに、クロスエンコーダーを使用して候補スキルに対して包括的なスコアリングを実行し、最終的なターゲットスキルセットを形成するために最適なものを選択します。 3 つのバックボーン LLM を使用した ALFWorld と ScienceWorld の実験では、既存のスキル選択ベースラインと比較して、SkillReranker がタスクのパフォーマンスを効果的に向上させ、環境インタラクションのステップを削減し、トークン消費量を削減することが示されています。
原文 (English)
Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval
Skill usage can significantly enhance the ability of modern agent systems to complete complex tasks. However, the growing scale of skill libraries makes accurate skill selection increasingly challenging. In real-world scenarios, ambiguous semantic matching often arises between a specific task requirement and multiple generic yet semantically similar candidate skills. Moreover, existing methods tend to overlook the dynamic influence of task difficulty and skill applicability when selecting the optimal target skill set. To address these issues, we propose SkillReranker, an inference-time reranking framework for adaptive skill selection. Specifically, we first perform semantic decomposition on both the task and skill sides, yielding informative subtask and execution-state descriptions as well as transition-state descriptions that characterize each skill's functionality. These descriptions are then used to construct a directed acyclic execution graph, where intermediate task states are modeled as nodes and candidate skills as edges, thereby establishing a structured task-skill correspondence. On this basis, SkillReranker determines whether each state node satisfies the split condition to identify subtask intervals. For each task interval, we employ a cross-encoder to perform comprehensive scoring over candidate skills and select the most suitable ones to form the final target skill set. Experiments on ALFWorld and ScienceWorld with three backbone LLMs show that SkillReranker effectively improves task performance, reduces environment interaction steps, and lowers token consumption compared with existing skill selection baselines.
DT-Guard: 推論不要の LLM 安全ガードレールのための意図主導型推論アクティブ トレーニング
オープンワールド アプリケーションに展開される大規模な言語モデルには、複雑なリスクに対して堅牢であり、低遅延のランタイム モデレーションに十分な効率性を備えた安全ガードレールが必要です。既存のガードレールは、効率的ではありますが、隠蔽された意図、あいまいなセマンティクス、境界線の安全性決定に苦戦することが多い軽量の分類ベースのモデルと、判断の品質は向上しますが、追加のトークン生成と推論レイテンシが発生する推論ベースのガードとの間の実質的なトレードオフに直面しています。 DT-Guard は、Reasoning-Active Training、Reasoning-Free Inference パラダイムに基づいたコンテンツ安全ガードレール モデルです。重要なアイデアは、トレーニング中に推論監視を使用し、推論時には構造化された安全ラベルのみを発行することです。 DT-Guard は、安全性の判断を、意図 - カテゴリ - 安全という漸進的な意思決定プロセスとして定式化し、意図ラベル、リスク カテゴリ、安全ラベル、および構造化推論軌跡を含む意図駆動型のデータセットを構築します。ハードケースの堅牢性をさらに向上させるために、ロールアウトガイド付きプログレッシブハードケース最適化 (RG-PHO) を提案します。これは、マルチロールアウトの一貫性を使用して、安定してマスタリングされたサンプル、永続的に失敗したサンプル、および優先順位が不安定なサンプルを特定し、それに応じてターゲットを絞った教師あり優先最適化を適用します。推論時に、DT-Guard は明示的な推論トレースなしで構造化ラベルを直接生成し、展開効率を維持します。プロンプト側とレスポンス側の安全ベンチマークに関する実験では、DT-Guard がそれぞれ 0.886 と 0.870 の平均 F1 スコアを達成していることが示されています。わずか 4B のバックボーンで、両側の平均 F1 は 0.878 に達し、強力な 8B ガードレールのベースラインを上回ります。これらの結果は、推論による監視が低遅延の安全性識別に効果的に組み込まれることを示しています。
原文 (English)
DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail
Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-latency runtime moderation. Existing guardrails face a practical trade-off between lightweight classification-based models, which are efficient but often struggle with concealed intent, ambiguous semantics, and borderline safety decisions, and reasoning-based guards, which improve judgment quality but introduce additional token generation and inference latency. We present DT-Guard, a content safety guardrail model based on a Reasoning-Active Training, Reasoning-Free Inference paradigm. The key idea is to use reasoning supervision during training while emitting only structured safety labels at inference time. DT-Guard formulates safety judgment as a progressive decision process, Intent - Category - Safety, and constructs an intent-driven dataset with intent labels, risk categories, safety labels, and structured reasoning trajectories. To further improve hard-case robustness, we propose Rollout-Guided Progressive Hard-Case Optimization (RG-PHO), which uses multi-rollout consistency to identify stably mastered, persistently failed, and preference-unstable samples, and applies targeted supervised and preference optimization accordingly. At inference time, DT-Guard directly generates structured labels without explicit reasoning traces, preserving deployment efficiency. Experiments on prompt-side and response-side safety benchmarks show that DT-Guard achieves average F1 scores of 0.886 and 0.870, respectively. With only a 4B backbone, it reaches a dual-side average F1 of 0.878, outperforming strong 8B guardrail baselines. These results demonstrate that reasoning supervision can be effectively internalized into low-latency safety discrimination.
間違った方法で運転する: End2End 自動運転モデルの解釈可能性を活用する
自動運転のためのエンドツーエンド学習の採用が増えると、モデルの複雑さと不透明さが増し、望ましくない動作や誤った動作を学習するリスクが高まります。この研究では、教師なし辞書学習を最先端の運転モデル内の事後解釈可能性モジュールとして統合し、運転行動を意味的に意味のある概念に分解しながら、モデルの運転決定に対する因果関係を実証します。我々は、エンドツーエンドモデルから意味のある概念を抽出して解釈し、それらを多面的なモデル出力に接続するための段階的なフレームワークを提案します。これにより、将来の軌道を予測するための基礎となる意思決定ロジックが明らかになります。さらに、コンセプトレベルでの的を絞った介入により、運転上の意思決定を操作および修正できるようになり、その結果、全体的な運転パフォーマンスが目に見えるほど向上します。したがって、解釈可能性を効果的に使用して、モデルの不透明性を軽減し、誤った動作を明らかにし、対象を絞った軽減策を可能にして、最終的にモデルのパフォーマンスを向上させる方法を示します。
原文 (English)
Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models
The increasing adoption of end-to-end learning for autonomous driving introduces increased model complexity and opacity, raising the risk of learning undesired or erroneous behavior. In this work, we integrate unsupervised dictionary learning as a post hoc interpretability module within state-of-the-art driving models to decompose driving behavior into semantically meaningful concepts while demonstrating their causal influence on the model's driving decisions. We propose a stepwise framework for extracting and interpreting meaningful concepts from the end-to-end model and connecting them to the multifaceted model outputs, thereby revealing the underlying decision-making logic for the prediction of future trajectories. Furthermore, targeted interventions at the concept level allow us to manipulate and correct driving decisions, resulting in measurable improvements in overall driving performance. We thus demonstrate how interpretability can effectively be used to reduce model opacity, uncover erroneous behavior, and enable targeted mitigation, ultimately boosting model performance.
TopoBrick: ゼロショット建築 IoT 予測のための外生変数のエージェント トポロジ サンプリング
建物センサーは物理トポロジー、空間階層、運用コンテキストに組み込まれていますが、既存の予測担当者はセンサーを孤立した時系列として扱ったり、固定共変量セットに依存したりすることがよくあります。 TopoBrick は、IoT (モノのインターネット) 予測をゼロショットで構築するためのトレーニング不要のフレームワークです。 TopoBrick は、ナレッジ グラフの構築を使用してコンパクトな構造スケルトンを構築し、エージェント トポロジ サンプラーを使用してターゲット固有の外生変数を選択します。選択された変数は、展開時の可用性によって編成され、過去の既知のセンサーの状態と、将来の既知のカレンダー、スケジュール、気象学的な外生変数が分離されます。 TopoBrick は、現実世界の 3 つの建物にわたって、強力なゼロショット基礎モデルのベースラインを上回り、完全にトレーニングされた建物固有のモデルとの競争力を維持します。アブレーションの結果、特に物理的に結合された HVAC や天候駆動のセンシング変数の場合、トポロジーを意識したサンプリングの方が、ランダム、オントロジーのみ、または固定ホップの選択よりも信頼性が高いことが示されています。
原文 (English)
TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting
Building sensors are embedded in physical topology, spatial hierarchy, and operational context, yet existing forecasters often treat them as isolated time series or rely on fixed covariate sets. We present TopoBrick, a training-free framework for zero-shot building IoT (Internet-of-Things) forecasting. TopoBrick uses building knowledge graphs to construct a compact structural skeleton and employs an agentic topology sampler to select target-specific exogenous variables. The selected variables are organized by deployment-time availability, separating past-known sensor states from future-known calendar, schedule, and meteorological exogenous variables. Across three real-world buildings, TopoBrick outperforms strong zero-shot foundation-model baselines and remains competitive with fully trained building-specific models. Ablations show that topology-aware sampling is more reliable than random, ontology-only, or fixed-hop selection, especially for physically coupled HVAC and weather-driven sensing variables.
世界モデルの定義とロードマップ
ワールド モデル (環境の構造とダイナミクスを学習する内部シミュレーター) は、AI で最も活発に議論される概念の 1 つになりました。モデルベースの強化学習やビデオ生成から、身体化ロボット工学、そして最終的には物理 AI に至るまで、AI サブ分野の研究者が「ワールド モデル」と呼ぶシステムを構築していますが、ワールド モデルとは基本的に何なのか、何を予測すべきなのか、どのように構築すべきなのかについてはまだ合意がありません。この観点からの記事では、世界モデルの科学的定義、その主要な技術的側面の議論、効果的な世界モデルを開発するための段階的なロードマップを提供します。
原文 (English)
A Definition and Roadmap for World Models
World models -- internal simulators that learn the structure and dynamics of an environment -- have become one of the most actively debated concepts in AI. From model-based reinforcement learning and video generation to embodied robotics and ultimately, physical AI, researchers across AI subfields are building systems that they call "world models", yet there is no consensus on what a world model fundamentally is, what it should predict, or how it should be built. This perspective article provides a scientific definition of world models, discussions of their key technical aspects, and a staged roadmap for developing effective world models.
ExplAIner: 分類モデルを説明するための宣言型クエリ言語
XAI コミュニティは、ML モデルの予測を説明するために幅広いクエリとスコアを研究してきました。データ管理の観点から見ると、説明概念の急増により、そのような概念を統一的に指定、組み合わせ、分析できる宣言型クエリ言語が必要になります。この論文では、ブールモデル用のそのようなフレームワークを開発します。まず、ブラック ボックス モデルの解釈可能性クエリ言語である FOIL を再検討し、FOIL には 2 つの基本的な制限があることを示します。1 つは中心最適性に基づく説明クエリを表現できないこと、もう 1 つは決定木に対するその評価問題が多項式階層のすべてのレベルで難しいことです。次に、拡張語彙と階層構造を備えた FOIL に基づくクエリ言語である ExplAIner を紹介します。 ExplAIner が、抽象的、対比的、特徴ベース、距離ベースのクエリを含む、幅広い説明概念を表現できることを示します。また、ExplAIner の各クエリの評価問題は、いくつかの基本的な述語を多項式時間で評価できるブール モデルのすべてのクラスにわたるブール階層に属していることも証明します。特に、その性質は決定論的で分解可能なブール回路に当てはまります。最後に、厳密な部分順序に関して最小限の説明を計算するための ExplAIner の最適化指向のフラグメントである Opt-FOIL を導入し、その評価問題が同じ扱いやすさの仮定の下で $\mathrm{FP}^{\mathrm{NP}}$ にあることを証明します。これらの複雑さの結果は、アルゴリズムに直接影響します。つまり、固定 ExplAIner クエリは、SAT ソルバーへの固定回数の呼び出しで評価できますが、Opt-FOIL で指定された説明の概念は、そのような呼び出しの多項式数で計算できます。これは、SAT ソルバーを使用して ML モデルのいくつかのクラスの説明を計算することに成功している正式な XAI に特に関係があります。
原文 (English)
ExplAIner: A Declarative Query Language for Explaining Classification Models
The XAI community has studied a wide range of queries and scores for explaining predictions of ML models. From a data management perspective, this proliferation of explanation notions calls for declarative query languages in which such notions can be specified, combined, and analyzed uniformly. In this paper, we develop such a framework for Boolean models. We first revisit FOIL, an interpretability query language for black-box models, and show that it has two fundamental limitations: it cannot express central optimality-based explanation queries, and its evaluation problem over decision trees is hard for every level of the polynomial hierarchy. We then introduce ExplAIner, a query language based on FOIL with an extended vocabulary and a layered structure. We show that ExplAIner can express a broad family of explanation notions, including abductive, contrastive, feature-based, and distance-based queries. We also prove that the evaluation problem for each query in ExplAIner belongs to the Boolean hierarchy over every class of Boolean models for which some basic predicates can be evaluated in polynomial time. In particular, that property holds for deterministic and decomposable Boolean circuits. Finally, we introduce Opt-FOIL, an optimization-oriented fragment of ExplAIner for computing explanations that are minimal with respect to strict partial orders, and prove that its evaluation problem is in $\mathrm{FP}^{\mathrm{NP}}$ under the same tractability assumptions. These complexity results have a direct algorithmic consequence: a fixed ExplAIner query can be evaluated with a fixed number of calls to a SAT solver, while a notion of explanation specified in Opt-FOIL can be computed with a polynomial number of such calls. This is particularly relevant in formal XAI, where SAT solvers have been successfully used to compute explanations for several classes of ML models.
細かい文字でヘリコバクター・ピロリを見つける: 胃生検レポートからの証拠にリンクされた多剤症例の所見
シンガポールのデータによると、人口の約 31% にヘリコバクター ピロリ感染の証拠があることが示されています。ピロリ菌の持続感染は慢性活動性胃炎や消化性潰瘍と関連しており、その除菌が胃がん予防の鍵となります。ただし、\textit{H.ピロリ菌陽性とピロリ菌関連胃炎は、コード化された自由記述形式の異種混合レポート フィールドに分布する可能性があり、主張と否定の状況に応じた解釈が必要となる場合があり、キーワード検索が制限され、手作業によるレビューの規模拡大が困難になる場合があります。私たちは、シンガポールの大規模な医療システムからの 54 件の匿名化された胃生検病理レポートを使用して、フィールド名主導型で証拠にリンクされた抽出ワークフローである Nimblemind マルチエージェント システム (nMAS) の遡及的パイロット評価を実施しました。胃/胃生検、生検状態、ヘリコバクター ピロリ陽性、ヘリコバクター ピロリ関連胃炎という 4 つの臨床医対象の 2 値フィールドが評価されました。 nMAS は、216 件のフィーチャーケースの決定のうち、213 件を正しく分類しました。これは全体の精度 98.61% に相当します。個別に実装された UMA スタイルの MiniMax M2.5 コンパレーターは、同様の集計およびフィールドごとの分類メトリックを生成しました。予測パフォーマンスは同様でしたが、nMAS はサポートするソース文を含む統合されたレポート レベルの出力を維持しました。したがって、証明された貢献は、予測の優位性ではなく、ワークフローの統合とトレーサビリティです。実例となる未測定のシナリオの下では、1,000 件のレポートを手動レビューで 5 分でレビューするのに対し、証拠に関連付けられた検証で 5 秒かかると、レビュー時間は 83.3 スタッフ時間から 1.4 スタッフ時間に短縮されます。これは、81.9 スタッフ時間と約 6,100 米ドルの潜在的なスタッフ時間価値に相当します。大規模な多施設研究では、証拠期間の正確性、臨床医の検証時間、一般化可能性を評価する必要があります。
原文 (English)
Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports
Data from Singapore indicated that about 31% of the population had evidence of Helicobacter pylori infection. Persistent H. pylori infection is associated with chronic active gastritis and peptic ulcer disease, and its eradication is key to gastric cancer prevention. However, evidence supporting \textit{H. pylori} positivity and H. pylori-associated gastritis may be distributed across heterogeneous coded and free-text report fields and may require contextual interpretation of assertion and negation, limiting keyword search, and making manual review difficult to scale. We conducted a retrospective pilot evaluation of the Nimblemind Multi-Agent System (nMAS), a field-name-driven, evidence-linked extraction workflow, using 54 de-identified gastric biopsy pathology reports from a large healthcare system in Singapore. Four clinician-scoped binary fields were evaluated: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis. Across 216 feature-case decisions, nMAS correctly classified 213, corresponding to 98.61% overall accuracy. A separately implemented UMA-style MiniMax M2.5 comparator produced similar aggregate and per-field classification metrics. Although predictive performance was similar, nMAS maintained unified report-level outputs with supporting source sentences; the demonstrated contribution is therefore workflow integration and traceability rather than predictive superiority. Under an illustrative, unmeasured scenario, reviewing 1,000 reports at five minutes per manual review versus five seconds per evidence-linked verification would reduce review time from 83.3 to 1.4 staff-hours, corresponding to 81.9 staff-hours and about USD~6,100 in potential staff-time value. Larger multi-institutional studies should evaluate evidence-span correctness, clinician verification time, and generalizability.
Danus: ファクトグラフ メモリを使用して数学的推論エージェントを調整する
最近の LLM ベースの数学的推論エージェントは研究レベルの問題に取り組み始めており、いくつかのケースでは未解決の問題の解決に貢献しています。ただし、中間クレームの整理と信頼性を維持しながら、並行した証拠検索を調整することが難しいため、このようなエージェントを効果的に拡張および調整することは依然として困難です。この論文では、グローバル メモリ管理メカニズムとして共有ファクト グラフを中心とした研究レベルの数学的推論のためのオーケストレーション システムである Danus を提案します。 Danus は、計画と調整を実行するメイン エージェント、証明検索を並行して実行する複数のワーカー エージェント、提案された数学的主張をファクト グラフに追加する前にチェックするステートレス ベリファイアーで構成されます。検証された各事実は、その証明および論理的依存関係とともに保存されるため、システムは共有された証明の状態を整理しながら、長い引数を段階的に構築できます。メイン エージェントは、進化する証明状態を定期的に要約し、有望な方向にワーカーをリダイレクトし、進捗レポートを通じて人間の数学者との対話をサポートします。私たちは、代数幾何学、特異点理論、および組合せ論における 6 つの研究レベルのケーススタディを通じてダヌスを評価し、ファクトグラフ記憶メカニズムによってダヌスがどのように長く詳細な数学的証明を構築できるかを示します。私たちの結果は、ファクトグラフベースのオーケストレーションが、長期的な研究課題に対して数理推論エージェントを拡張するための効果的な手段を提供することを示唆しています。 Danus は、https://github.com/frenzymath/Danus でオープンソースです。
原文 (English)
Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory
Recent LLM-based mathematical reasoning agents have begun to tackle research-level problems and, in several cases, have contributed to the resolution of open problems. However, scaling and orchestrating such agents effectively remains challenging, due to the difficulty of coordinating parallel proof search while keeping intermediate claims organized and reliable. In this paper, we propose Danus, an orchestration system for research-level mathematical reasoning centered on a shared fact graph as a global memory-management mechanism. Danus consists of a main agent that performs planning and coordination, multiple worker agents that carry out proof search in parallel, and a stateless verifier that checks proposed mathematical claims before they are admitted into the fact graph. Each verified fact is stored together with its proof and logical dependencies, allowing the system to build long arguments incrementally while keeping the shared proof state organized. The main agent periodically summarizes the evolving proof state, redirects workers across promising directions, and supports interaction with human mathematicians through progress reports. We evaluate Danus through six research-level case studies in algebraic geometry, singularity theory, and combinatorics, illustrating how the fact-graph memory mechanism enables Danus to construct long, detailed mathematical proofs. Our results suggest that fact-graph-based orchestration provides an effective route toward scaling mathematical reasoning agents for long-horizon research problems. Danus is open source at https://github.com/frenzymath/Danus.
バイマテリアルシステムにおける弾性波伝播のための物理学に基づいたニューラルネットワークフレームワーク
物理情報に基づいたニューラル ネットワーク (PINN) は、基礎となる物理法則を学習プロセスに直接埋め込みながら、偏微分方程式を解くための有望なフレームワークを提供します。この研究は、線形弾性の軸対称方程式によって支配されるバイマテリアルシステムにおける過渡弾性波伝播をモデル化するための PINN ベースのフレームワークを提示します。スプリット ホプキンソン圧力バー構成を代表する鋼とアルミニウムの試験片が考慮され、支配的な弾性力学方程式が、対応する初期条件、境界条件、界面条件とともに、物理学に基づいた損失関数を通じてネットワークに直接組み込まれます。 ANSYS Workbench Explicit Dynamics を使用して実行される高忠実度の有限要素シミュレーションは、検証とトレーニング中の補足的なデータ制約として使用されます。提案されたフレームワークは、バイマテリアル界面を横切る波の透過と反射を正確に予測し、軸方向および半径方向の変位履歴、面平均応答、および有限要素解と密接に一致して支配的な応力とひずみの発展を再現します。さらに、訓練されたネットワークは、追加の有限要素シミュレーションを必要とせずに、これまでに見たことのない瞬間の波の応答や変更された材料特性を予測する能力を実証し、弾性力学解析のための連続代理モデルを提供します。メッシュ感度の研究により数値的な堅牢性が確認され、追加の材料の組み合わせにより、提案された方法論の一般性が実証されました。その結果、物理学に基づいたニューラル ネットワークと陽的有限要素解析を統合することで、不均質固体における弾性波伝播のための正確で計算効率の高いフレームワークが提供され、高速固体力学および衝撃工学アプリケーションに効果的な代替モデリング アプローチが提供されることが示されました。
原文 (English)
A Physics-Informed Neural Network Framework for Elastodynamic Wave Propagation in Bimaterial Systems
Physics-informed neural networks (PINNs) provide a promising framework for solving partial differential equations while embedding the underlying physical laws directly into the learning process. This study presents a PINN-based framework for modeling transient elastodynamic wave propagation in bimaterial systems governed by the axisymmetric equations of linear elasticity. A steel-aluminum specimen representative of a Split Hopkinson Pressure Bar configuration is considered, and the governing elastodynamic equations, together with the corresponding initial, boundary, and interface conditions, are incorporated directly into the network through a physics-informed loss function. High-fidelity finite-element simulations performed using ANSYS Workbench Explicit Dynamics are used for validation and as supplementary data constraints during training. The proposed framework accurately predicts wave transmission and reflection across the bimaterial interface and reproduces axial and radial displacement histories, face-averaged responses, and the dominant stress and strain evolution with close agreement to the finite-element solutions. The trained network further demonstrates the ability to predict wave responses at previously unseen time instants and for modified material properties without requiring additional finite-element simulations, providing a continuous surrogate model for elastodynamic analysis. Mesh-sensitivity studies confirm numerical robustness, while additional material combinations demonstrate the generality of the proposed methodology. The results show that integrating physics-informed neural networks with explicit finite-element analysis provides an accurate and computationally efficient framework for elastodynamic wave propagation in heterogeneous solids, offering an effective surrogate modeling approach for high-rate solid mechanics and impact engineering applications.
酪農場における多目的電池管理のためのマルチエージェント深層強化学習
アイルランドの乳業には、再生可能エネルギーの統合と炭素排出削減の大きな可能性があります。ただし、分散型発電制御の研究者は主に住宅用および商業用アプリケーションに焦点を当てています。酪農部門における再生可能エネルギーの効果的な統合に貢献するために、この論文では、差分進化とマルチエージェント深層強化学習に基づく多目的最適化制御システムを紹介します。提案された制御は 2 つの層で構成されます。上位層は動的価格設定を使用し、下位層はバッテリー管理のためのマルチエージェント強化学習に基づいています。この論文では、地方の配電回路における提案された制御システムの電気応答もシミュレーションします。シミュレーション結果は、提案された制御フレームワークが、ルールベースのモデルを使用する場合と比較して、エネルギー裁定取引からの利益を最大 18% 向上させ、コストを大幅に増加させることなく分散型発電の使用を増加させ、電圧変動に関してアイルランドの送電網コードに準拠できることを示しています。
原文 (English)
Multi-Agent Deep Reinforcement Learning for Multi Objective Battery Management in Dairy Farms
The dairy industry in Ireland has a large potential for the integration of renewable energy and the reduction of carbon emissions. However, researchers of distributed generation control are mainly focused on residential and commercial applications. To contribute to the effective integration of renewable energy in the dairy sector, this paper presents a multi-objective optimisation control system based on differential evolution and multi agent Deep Reinforcement Learning. The proposed control is organised in two layers: the upper layer uses dynamic pricing, and the lower layer is based on multi-agent reinforcement learning for battery management. This paper also simulates the electrical response of the proposed control system in a rural distribution circuit. The simulation results show that the proposed control framework can improve profits from energy arbitrage up to 18% compared to using Rule-based models, increase the use of distributed generation without significantly increasing cost, and comply with the Irish grid code in terms of voltage variation.
最初から運命: リコール制御のプローブ カスケードによる LLM エージェント エピソードの早期中止
複数ステップのタスクを解決する大規模言語モデル (LLM) エージェントは、頻繁に失敗する運命にある軌道にコミットしますが、失敗が観測可能になる前に大量の推論コンピューティングを消費し続けます。失敗はエージェントの内部表現から早期に予測可能であることを示します。隠れたアクティベーションのラウンドごとの軽量プローブは、最初のインタラクション ラウンドの早い段階で最終的なエピソードの失敗を予測します。スコアラーはエージェントの観察可能な動作のみを読み取るため、偶然よりもかろうじて優れています。私たちはこの信号を実用的な中止カスケードに変換します。つまり、ラウンドごとに 1 つの配布フリーのキャリブレーション ゲートを使用し、ラウンドごとのリコール バジェットを共同検索して、最終的に成功したエピソードがユーザー指定のグローバル レートですべてのゲートを生き残れるようにします。誤った中止のリスクはゲートを越えて蓄積されるため、このエピソード レベルの保証は展開において重要です。 TextCraft の 2 つのエージェント モデルにわたって、カスケードは 90% から 97% までのすべてのリコール目標を満たし、90% の目標では、推論計算の 47.1% +/- 10.3% (Qwen-2.5-7B) および 37.2% +/- 8.8% (Llama-3.2-3B) を節約します。これは、最良のシングルゲート ポリシーの 1.6 ~ 1.7 倍です。それ以外は同一のカスケード読み取り専用動作では、約半分の節約になりますが、プローブに動作機能を追加してもそれ以上の利益は得られません。隠れた状態は、動作が明らかにするものをキャプチャします。最後に、高リコール目標を認定する際のサンプルの複雑さを特徴付け、どのリコールがデータを取り戻すことができるか、そしておそらくできないと専門家に伝えます。コードは近日公開予定です。
原文 (English)
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance. We turn this signal into a practical abort cascade: one distribution-free calibrated gate per round, with per-round recall budgets jointly searched so that eventually-successful episodes survive all gates at a user-specified global rate; this episode-level guarantee is the one that matters in deployment, since false-abort risk accumulates across gates. Across two agent models on TextCraft, the cascade meets every recall target from 90% to 97% and, at the 90% target, saves 47.1% +/- 10.3% (Qwen-2.5-7B) and 37.2% +/- 8.8% (Llama-3.2-3B) of inference compute, 1.6--1.7x the best single-gate policy. An otherwise-identical cascade reading only behavior saves roughly half as much, and adding behavioral features to the probe yields no further gain: the hidden states capture what behavior reveals. Finally, we characterize the sample complexity of certifying high recall targets, telling practitioners which recall promises their data can, and provably cannot, back. The code will be released soon.
RMISC: 時系列基盤モデルのための大規模な実世界多変量コーパス
近年、時系列基礎モデル (TSFM) を使用した多変量モデリングが出現し、高度なゼロショット汎化が実現されています。最新の多変量 TSFM は、主に多変量合成データに基づいて事前トレーニングされており、スケーリングは容易ですが、現実世界の時系列に存在する複雑な時間ダイナミクスや変数間の関係を捉えることができない可能性があります。これにより、重要な疑問が生じます。現実世界のコーパスでトレーニングされた主要な TSFM は、合成データでトレーニングされたものよりも優れたパフォーマンスを発揮するかどうか、またどの程度パフォーマンスが優れているのかということです。これに答えるために、私たちは RMISC コーパスを確立しました。これは、さまざまなドメインにわたる約 200 のデータセットと 1,420 億のタイム ポイントを含む、かなり大規模で高品質で、オープンにアクセスできる現実世界の多変量時系列アーカイブです。さらに、一変量データ、合成多変量データ、実世界の多変量データで 4 つの高度な TSFM を事前トレーニングし、標準の分布内および分布外ベンチマークでゼロショット汎化機能を評価します。実験結果は、現実世界の多変量データを組み込むと、単変量 TSFM と多変量 TSFM の両方の汎化パフォーマンスが主に向上することを示しています。これらの結果は、現実世界の多変量データがより強力な TSFM の開発にどのように寄与するかについてのより深い理解を提供します。
原文 (English)
RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models
Recent years have witnessed the emergence of multivariate modeling using time series foundation models (TSFMs), which achieve advanced zero-shot generalization. Modern multivariate TSFMs are predominantly pretrained on multivariate synthetic data, which is easier to scale but may fail to capture the complex temporal dynamics and cross-variable relationships present in real-world time series. This raises a key question: Whether and to what extent the leading TSFMs trained with the real-world corpus perform better than those trained with synthetic data? To answer this, we establish the RMISC corpus, a considerably large-scale, high-quality, openly accessible, real-world, and multivariate time series archive that contains around 200 datasets and 142 billion time points across diverse domains. Furthermore, we pretrain four advanced TSFMs on univariate, synthetic multivariate, and real-world multivariate data and evaluate their zero-shot generalization capabilities on standard in-distribution and out-of-distribution benchmarks. Experimental results show that incorporating real-world multivariate data predominantly improves the generalization performance for both univariate and multivariate TSFMs. These results provide a deeper understanding of how real-world multivariate data contributes to the development of stronger TSFMs.
FootsiesGym: 2 人用ゼロサム不完全情報ゲームの格闘ゲーム ベンチマーク
私たちは、自明ではない 2 プレイヤー、ゼロサム、不完全情報ゲームで学習するためのオープンソース環境、FootsiesGym を紹介します。 HiFight のミニマリスト 2D 格闘ゲーム Footsies に基づいて構築されており、効率的な分析に十分なシンプルさを保ちながら、格闘ゲームのニュートラル プレイの周期的で非推移的な戦略的相互作用を分離します。当社は、標準ハードウェアでの高スループット トレーニングを可能にし、アクセスしやすく再現可能な環境を実現するベクトル化されたシミュレーターを提供します。環境の設計について説明し、いくつかの強化学習アルゴリズムのベンチマークを行い、それが可能にするオープンな研究の方向性について議論します。コードは https://github.com/como-research/FootsiesGym で入手できます。
原文 (English)
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games
We present FootsiesGym, an open-source environment for learning in a non-trivial two-player, zero-sum, imperfect-information game. Built on HiFight's minimalist 2D fighting game Footsies, it isolates the cyclic, non-transitive strategic interactions of fighting game neutral play while remaining simple enough for efficient analysis. We provide a vectorized simulator that enables high-throughput training on standard hardware, making the environment accessible and reproducible. We describe the design of the environment, benchmark several reinforcement learning algorithms, and discuss open research directions it enables. The code is available at https://github.com/como-research/FootsiesGym.
FreqDepthKV: ロングコンテキスト LLM 推論における堅牢な KV キャッシュ圧縮のための周波数ガイドによる深度共有
ロングコンテキスト LLM 推論は、KV キャッシュのメモリと帯域幅のコストによってますます制限されていますが、積極的な圧縮により、取得や複数ステップの推論に必要な層固有の証拠が削除される可能性があります。隣接層の KV 状態を共有の低周波深度成分とまばらな高周波残差に因数分解する推論時キャッシュ圧縮手法である FreqDepthKV を紹介します。軽量のオンライン プローブは、再構築に敏感なアテンション ロジットへの寄与に応じてアテンション ヘッドを共有深度、残差深度、または正確なキャッシュ モードに割り当て、再トレーニングすることなく圧縮ポリシーをプロンプト構造に適応させることができます。 FreqDepthKV は、ロングコンテキストの質問応答、ニードル検索、要約、およびコード生成ベンチマーク全体にわたって、大幅に少ないキャッシュ バジェットの下でタスクの精度を維持します。 32k トークンのプレフィル ウィンドウでは、FreqDepthKV は 58.3 完全一致、63.0 F1、32.5 ROUGE-L、および 48.1 pass@1 に達し、完全な KV にほぼ一致し、以前の圧縮キャッシュ方式を上回ります。また、デコード スループットが 70.4 トークン/秒に向上し、TTFT が 2.06 秒に短縮され、ピーク KV メモリが 6.2 GB に低下し、3.9 倍の実効圧縮率を達成します。
原文 (English)
FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the layer-specific evidence needed for retrieval and multi-step reasoning. We introduce FreqDepthKV, an inference-time cache compression method that factorizes adjacent-layer KV states into shared low-frequency depth components and sparse high-frequency residuals. A lightweight online probe assigns attention heads to shared-depth, residual-depth, or exact cache modes according to their contribution to reconstruction-sensitive attention logits, allowing the compression policy to adapt to prompt structure without retraining. Across long-context question answering, needle retrieval, summarization, and code generation benchmarks, FreqDepthKV preserves task accuracy under substantially smaller cache budgets. With a 32k-token prefill window, FreqDepthKV reaches 58.3 Exact Match, 63.0 F1, 32.5 ROUGE-L, and 48.1 pass@1, closely matching full KV while outperforming prior compressed-cache methods. It also improves decoding throughput to 70.4 tokens/s, reduces TTFT to 2.06 seconds, and lowers peak KV memory to 6.2 GB, achieving a 3.9x effective compression ratio.
視覚的なアクションの結果推論の調整による物理的推論とタスクの一般化の橋渡し
視覚言語モデル (VLM) は、特に目に見えないタスクや環境の下で、対話型の物理的推論において一般化するのに苦労します。 2 つの主要な障害モードが顕著です。それは、物理的現実と矛盾する幻覚的な思考連鎖 (CoT) 推論と、モデルの推論とアクションの間の不整合です。私たちは、両方の問題に直接対処する新しい報酬デザインである VAORA (Visual Action Outcome Reasoning Alignment) を紹介します。 VAORA では、2 つの補完的な報酬を導入しています。ビジュアル アラインメント報酬は、エージェントのアクション自体とは独立して、VLM 推論をビジュアル コンテキストに固定します。もう 1 つは、モデルのアクションによって引き起こされる視覚的な結果に推論を根拠付けるビジュアル アクション アライメント報酬です。これらの報酬を組み合わせることで、幻覚性 CoT が抑制され、推論と行動の間のギャップが減少します。トレーニングの安定性を向上させるために、事前トレーニングされたドメイン内エキスパート エージェントを使用して成功確率を推定することにより、スムーズで密度の高い報酬をさらに採用します。 PHYRE と仮想ツールの実験は、新しいタスクや目に見えない環境設定全体でのパフォーマンスをサポートし、VAORA を通じて根拠のある一般化可能な身体的知性を誘導できることを確認しました。
原文 (English)
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment
Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. Two key failure modes are prominent: hallucinated chain-of-thought (CoT) reasoning that contradicts physical reality, and misalignment between the model's reasoning and actions. We present VAORA (Visual Action Outcome Reasoning Alignment), a novel reward design that directly addresses both issues. VAORA introduces two complementary rewards: Visual Alignment Reward, which anchors VLM reasoning to the visual context independent of the agent action itself, and Visual-Action Alignment Reward, which grounds reasoning in the visual outcome induced by the model's action. Together, these rewards suppress hallucinated CoT and reduce the gap between reasoning and behavior. To improve training stability, we further employ smooth, dense rewards by estimating success probabilities using a pre-trained in-domain expert agent. Experiments on PHYRE and Virtual Tool support our performances across novel-task and unseen-environment settings, confirming that grounded and generalizable physical intelligence can be induced through VAORA.
DepthWeave-KV: ロングコンテキスト KV キャッシュ圧縮のためのトークン適応型クロスレイヤー残差因数分解
ロングコンテキスト言語モデルの推論は、キーと値のキャッシュを保存するために必要なメモリ帯域幅と容量によってますます制限されていますが、既存の圧縮方法では多くの場合、レイヤーまたはトークン全体に均一のバジェットが適用され、語彙キューと意味論的状態が異なる保存を必要とする場合に検索が低下します。 DepthWeave-KV は、アテンション動作が敏感な場合に軽量のトークン固有の残差を保持しながら、共有低ランク チャネル ベースを使用して隣接するトランスフォーマー層全体でキーと値の状態を因数分解するトークン適応型キャッシュ圧縮方法を紹介します。 DepthWeave-KV は、クロス深度残差分解を、命令を含むトークンと取得クリティカルなトークンに高い再構成ランクを割り当てるトークン条件付き深度ルーターと組み合わせます。また、アテンション出力プローブからのキャリブレーション不要のオンライン エラー追跡を使用して、ベース モデルを再トレーニングすることなく、生成中に圧縮を適応させます。融合 CUDA 実装は、基底ルックアップ、残差逆量子化、およびアテンション投影を共同で実行して、デコード時のメモリ トラフィックを削減します。 LongBench、Needle-in-a-Haystack、L-Eval、および長形式の QA および要約ベンチマーク全体にわたって、DepthWeave-KV は、大幅に少ないメモリ使用量でほぼフル キャッシュのタスク品質を達成し、以前の圧縮キャッシュに比べて平均スコアと取得精度を向上させながら、64K コンテキストで 8.3 倍の KV メモリ削減と 1 秒あたり 72.8 トークンを達成します。
原文 (English)
DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression
Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store key-value caches, yet existing compression methods often apply uniform budgets across layers or tokens and degrade retrieval when lexical cues and semantic states require different preservation. We introduce DepthWeave-KV, a token-adaptive cache compression method that factorizes key and value states across neighboring transformer layers using shared low-rank channel bases while retaining lightweight token-specific residuals where attention behavior is sensitive. DepthWeave-KV combines cross-depth residual factorization with a token-conditional depth router that allocates higher reconstruction rank to instruction-bearing and retrieval-critical tokens, and uses calibration-free online error tracking from attention-output probes to adapt compression during generation without retraining the base model. A fused CUDA implementation jointly performs basis lookup, residual dequantization, and attention projection to reduce decode-time memory traffic. Across LongBench, Needle-in-a-Haystack, L-Eval, and long-form QA and summarization benchmarks, DepthWeave-KV achieves near-full-cache task quality with substantially lower memory use, improving average score and retrieval accuracy over prior compressed caches while reaching 8.3x KV memory reduction and 72.8 tokens per second at 64K context.
Large Cancer Assistant (LCA): 腫瘍学におけるスケーラブルな臨床意思決定支援のためのモデルに依存しないオーケストレーション フレームワーク
- 目的: 腫瘍学におけるマルチモーダル深層学習モデルは現在、データ取り込み、臨床ルーティング、人工知能 (AI) 推論を厳密に結合するモノリシック設計によって制限されています。この柔軟性のなさに対処するために、私たちは、スケーラブルな臨床意思決定支援のために設計された、モデルに依存しない事後オーケストレーション フレームワークである Large Cancer Assistant (LCA) を提案します。 - 方法: LCA は、アルゴリズム不浸透性の原則に基づいた 7 タプル アーキテクチャとして数学的に形式化されており、オーケストレーション ロジックが基礎となるブラック ボックス AI モデルから厳密に独立していることが保証されます。幾何学深層学習 (GDL) を活用して、異なる構造軸と医療軸に沿ってマルチモーダルな患者データを標準化するエントリー理論を導入します。このシステムは、Cancer Switching Module を介してデータを動的に調整し、標準化された中間ペイロード (SIP) を出力することで、コア AI の実行を不安定な病院 IT インフラストラクチャから意図的に分離します。 - 結果: 概念実証 (PoC) により、4 つの技術シナリオにわたってオーケストレーション ロジックが検証されました。フレームワークは、オーケストレーション オーバーヘッドを無視して通常のフローを実行しました。 AI モデルのスワップ中に不変のルーティング予測を維持することでアルゴリズムの不浸透性を経験的に実証し、注入されたデータ異常下でターゲットを絞った補足データ リクエスト (SDR) の生成で 100% の再現率を達成することで厳格な障害安全性を検証しました。マルチプロトコル実行機能も検証されました。 - 結論: マルチモーダル取り込みを機能推論から構造的に切り離すことにより、LCA は適応性の高いモジュール式オーケストレーション基盤を提供します。 SIP は明確なアーキテクチャ境界を確立し、独立した将来のパラダイムとしてダウンストリーム電子医療記録 (EMR) の相互運用性の準備をネイティブに設定します。
原文 (English)
The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology
- Objective: Multimodal deep learning models in oncology are currently limited by monolithic designs that rigidly couple data ingestion, clinical routing, and artificial intelligence (AI) inference. To address this inflexibility, we propose the Large Cancer Assistant (LCA), a model-agnostic, post-hoc orchestration framework designed for scalable clinical decision support. - Methods: The LCA is mathematically formalized as a 7-tuple architecture grounded in the principle of Algorithmic Impermeability, ensuring the orchestration logic remains strictly independent of underlying black-box AI models. We introduce the Entry Theory, leveraging Geometric Deep Learning (GDL) to standardize multimodal patient data along distinct structural and medical axes. The system dynamically orchestrates data via a Cancer Switching Module and intentionally isolates the core AI execution from volatile hospital IT infrastructures by outputting a Standardized Intermediate Payload (SIP). - Results: A Proof of Concept (PoC) validated the orchestration logic across four technical scenarios. The framework executed a nominal flow with negligible orchestration overhead. It empirically demonstrated algorithmic impermeability by maintaining an invariant routing projection during AI model swaps, and it validated strict failure-safety by achieving a 100\% recall rate in generating targeted Supplementary Data Requests (SDR) under injected data anomalies. Multi-protocol execution capability was also successfully verified. - Conclusion: By structurally decoupling multimodal ingestion from feature inference, the LCA provides a highly adaptable and modular orchestration foundation. The SIP establishes a clear architectural boundary, natively setting the stage for downstream Electronic Medical Record (EMR) interoperability as an independent future paradigm.
文化遺産保護の視点からインド AI を再考する
人工知能 (AI) がインド亜大陸のさまざまな地域に進出するにつれて、AI がこの文明の言語的および文化的基盤にどのような影響を与えるかを研究することに大きな関心が集まっています。 AI は「諸刃の剣」とみなされており、一方では大規模な人口のアクセスと包摂を可能にする一方で、他方では世界観を均質化し、過小評価されている言語や世界観を排除することができます。この論文では、インド言語学の広範な特徴と、それらが文化的実践や世界観とどのように密接に結びついているかを取り上げることによって、この問題を特徴づけることを試みます。次に、自然言語処理 (NLP) 技術がこの分野でどのように進化してきたかを縦断的に調査し、インド NLP の歴史的発展をたどり、主要なマイルストーン、方法論の変化、リソース作成の取り組みをカバーします。さらに、この論文では、豊富な形態学、複雑な文字と文法規則、二重言語、大きな方言変化など、インド言語の構造的および社会言語学的特徴も調査し、これらが AI 基盤モデルの構築にどのように特有の課題を生み出すのかについて説明しています。次に、インドの基盤モデルの役割の増大について議論し、これらのモデルがこれらの長年にわたるリソースと表現のギャップにどのように対処するかを分析します。最後に、解釈学的推論に基づいて AI を再考する「文化センシング」と呼ばれる研究の方向性を提案します。カルチャー センシングは、リソースの少ない言語間で公平なパフォーマンスを確保し、文化的に意味のある出力を生成するなどの未解決の問題に対処することを目的としています。本稿では、過去の研究、現在の技術、新たなトレンドをまとめることで、インド NLP の次の段階を導き、より堅牢で包括的なインド基礎モデルの開発に貢献できる研究の方向性を概説します。
原文 (English)
Rethinking Indic AI from a Lens of Cultural Heritage Preservation
As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant interest in studying how AI impacts the linguistic and cultural foundations of this civilization. AI is seen as a ''double-edged sword'' where on the one hand, it can enable access and inclusion for a large population, on the other, it can homogenize worldviews and exclude underrepresented languages and worldviews. In this paper, we try to characterize this problem by addressing the extensive characteristic nature of Indian linguistics and the way they closely connect to cultural practices and worldview. We then perform a longitudinal survey of how Natural Language Processing (NLP) techniques have evolved in this space, tracing the historical development of Indic NLP, covering key milestones, methodological shifts, and resource creation efforts. In addition, the paper also examines the structural and sociolinguistic characteristics of Indian languages, such as rich morphology, complex scripts and grammar rules, diglossia, and large dialectal variation, and explains how these create unique challenges for building AI foundation models. We then discuss the growing role of Indic foundation models and analyze how these models address these long-standing resource and representation gaps. Finally, we propose a research direction called 'Culture Sensing', which re-imagines AI based on hermeneutic reasoning. Culture Sensing aims to address open problems such as ensuring equitable performance across low-resource languages and producing outputs that are culturally meaningful. By bringing together past work, current techniques, and emerging trends, this paper outlines research directions that can guide the next phase of Indic NLP and contribute to the development of more robust and inclusive Indic foundation models.
残留権限: コーディング エージェントの取り消し可能なリソースと効果の機能
コーディング エージェントは、リソースが 1 つのサブ目標にのみ必要な場合でも、タスク全体に対して広範なツールにアクセスできることがよくあります。私たちはこれをギャップ残留権限と呼びます。つまり、一時的なリソース/効果機能は、それを正当化するエピソードが閉じられた後も露出したままになります。 PORTICO は、プランナーに公開される取り消し可能な機能のリファレンス モニターです。これは、明示的なタスク コントラクトを初期機能、許可ルール、信頼できるクロージャ述語、およびグローバル拒否ルールにコンパイルします。 request-grant-invoke ライフサイクルは、不透明なエポック限定ハンドルとして展開を具体化します。 Closure は、次のプランナー インターフェイスからこれらのハンドルを削除し、副作用が発生する前に古いリプレイを拒否します。モニターは、仲介ツールとサウンドタイプのカタログを想定しています。制御されたコーディング エージェント タスクでは、PORTICO は評価された実行で実行された契約で禁止された効果を記録しませんが、制御された許可は固定された狭いエンベロープによってブロックされた境界作業を回復します。非取り消しコンパレーターは、同じ最初のエンベロープと同じターンで同じ許可を受け取ります。クロージャ スライスでは、両方のシステムがタスクの成功、スコープのコンプライアンス、およびクロージャ前のすべての決定を照合します。次に、PORTICO は 10/10 の閉鎖後の再利用を拒否しますが、コンパレータは 10/10 を許可します。確定的な古い書き込み監査では、0/6 と 6/6 の実行された禁止効果が記録されます。ファイル書き込み、git ミューテーション、およびネットワーク出力に関するスクリプト化されたトレースと 6 つのライブ モデル トレースは、同じ分割を示しています。 4 エピソードの同一ポリシー診断では、広範なリクエストの露出により、実行された禁止効果はゼロのままですが、ブロックされたプロポーザルは 67 から 84 に増加します。コミットとトレースが記録された凍結されたリアル リポジトリの実行は、実際のプロジェクト レイアウトで同じライフサイクルを実行します。
原文 (English)
Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents
Coding agents often receive broad tool access for an entire task, even when a resource is needed only for one subgoal. We call this gap lingering authority: a temporary resource/effect capability remains exposed after the episode that justified it has closed. PORTICO is a reference monitor for revocable capabilities exposed to the planner. It compiles an explicit task contract into initial capabilities, grant rules, trusted closure predicates, and global deny rules. A request-grant-invoke lifecycle materializes expansions as opaque, epoch-bound handles. Closure removes those handles from the next planner interface and rejects stale replay before side effects. The monitor assumes mediated tools and a sound typed catalog. In controlled coding-agent tasks, PORTICO records no executed contract-forbidden effects in the evaluated runs, while controlled grants recover boundary work blocked by a fixed narrow envelope. A non-revoking comparator receives the same initial envelope and the same grants at the same turns. On the closure slice, both systems match task success, scope compliance, and all pre-closure decisions; PORTICO then rejects 10/10 post-closure reuses, while the comparator permits 10/10. A deterministic stale-write audit records 0/6 versus 6/6 executed forbidden effects. Scripted traces and six live model traces over file writes, git mutation, and network egress show the same split. In a four-episode same-policy diagnostic, broad request exposure preserves zero executed forbidden effects but raises blocked proposals from 67 to 84. Frozen real-repository runs, with commits and traces recorded, exercise the same lifecycle on real project layouts.
KVpop -- 予測オンライン プルーニングによる Key-Value キャッシュ圧縮
メモリと帯域幅はコンテキストの長さに比例して増加するため、キーバリュー (KV) キャッシュの増大は自己回帰デコードの大きなボトルネックになります。既存の KV エビクション方法は静的なヒューリスティックやプロキシ スコアに依存することが多く、将来のトークンの有用性を追跡することが不十分であり、関連性が変化するとエビクションが脆弱になります。これに対処するために、KVpop を導入します。KVpop は、キープまたはドロップの決定を直接監視することで、固定予算の KV 立ち退きポリシーを学習します。スコアラーは、新しい将来の注目ターゲットに対してトレーニングされ、高密度の注目マップを実現することなく効率的に計算されます。さらに、遅延メモリベースのスコアラーを導入します。これは、学習されたエビクション手法の中で独自に、固定ステップ数のスコアリングを延期して、近未来のコンテキストを活用します。 AIME と HMMT の数学的推論では、KVpop は Qwen3-4B の 75% KV キャッシュ圧縮で 98%、88% 圧縮で 97% の全注意パフォーマンスを維持し、確立されたエビクション ベースラインを一貫して上回っています。 Qwen3-8B はさらに強力な結果を示し、教師のパフォーマンスがほぼフルに達しました。これらの結果は、将来注目信号を使用してエビクションを監視すると、品質を維持しながらメモリ コストを削減できることを示しています。
原文 (English)
KVpop -- Key-Value Cache Compression with Predictive Online Pruning
Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length. Existing KV eviction methods often rely on static heuristics or proxy scores, which poorly track future token utility and cause brittle eviction as relevance shifts. To address this, we introduce KVpop, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision. The scorer is trained against a novel future-attention target, computed efficiently without materializing dense attention maps. We further introduce a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context. On AIME and HMMT mathematical reasoning, KVpop retains 98% of full-attention performance on Qwen3-4B at 75% KV cache compression and 97% at 88% compression, consistently outperforming established eviction baselines. Qwen3-8B shows even stronger results, reaching near-full teacher performance. These results show that supervising eviction with future-attention signals cuts memory costs while maintaining quality.
実行証明: 管理された AI エージェントのアクションの実行時検証
エージェント システムは、アドバイスするよりも実行することが増えています。 AI エージェントが規制されたデータをクエリし、効果的なツールを呼び出し、永続的な状態を変更する場合、端末の出力が妥当であるかどうかによって正確性は捕捉されません。有効な問題は、各ステップが契約に基づいて許可されているかどうか、記録された履歴は改ざんが明らかであるかどうか、および軌道を決定論的に再構築できるかどうかです。これを実行時の実行証明として形式化します。実行は、コントラクト $C$、実行因果イベント ストリーム (ECES) $T$、リプレイ コンテキスト $R$ のトリプル $x = (C, T, R)$ です。整形式述語とバリデーターチェック可能な 5 つの不変式が PoE 有効性述語を形成します。 5 つのセマンティック保証では、認可、パス コンプライアンス、拒否に対するヌル効果、履歴の整合性、および再生可能性について説明します。当社は、明示的な暗号化および展開の前提に基づいて健全性を証明します。PPT 攻撃者が意味論的保証に違反する PoE 有効な実行を生成すると、署名偽造、ハッシュ衝突、または定量化された展開失敗イベントが発生します。プライム実行モデル (PEM) は、計画、施行、効果、記録管理を個別の権限面に分離します。補題は、トレースの完全性をエフェクター専用の認証に削減します。実行証明証明書は、PoE = 1 の場合にのみ発行されます。単一ノードの TypeScript プロトタイプでは、PoE により、最小フローで約 2.7 ミリ秒が追加され、同時バッチ ワークロードで 4.4% のオーバーヘッドが追加されます。標準の 8 イベント トレースは約 1.1 KB に圧縮されます。注入されたゲートウェイ バイパス攻撃とトレースミューテーション攻撃は拒否されます。 PoE はコンセンサス、TEE、または zkVM を置き換えるものではありません。これは、承認、効果、履歴、および再生を単一の実行時チェック可能なオブジェクトにバインドするため、管理された実行が契約に基づいて証明可能になります。
原文 (English)
Proof of Execution: Runtime Verification for Governed AI Agent Actions
Agent systems increasingly execute rather than advise. When an AI agent queries regulated data, invokes effectful tools, and mutates persistent state, correctness is not captured by whether a terminal output looks plausible. The operative questions are whether each step was authorized under a contract, whether the recorded history is tamper-evident, and whether the trajectory can be reconstructed deterministically. We formalize this as runtime proof of execution. An execution is a triple $x = (C, T, R)$: a contract $C$, an Execution Causal Event Stream (ECES) $T$, and a replay context $R$. A well-formedness predicate and five validator-checkable invariants form the PoE validity predicate. Five semantic guarantees describe authorization, path compliance, null effect on deny, history integrity, and replayability. We prove soundness under explicit cryptographic and deployment assumptions: any PPT adversary that produces a PoE-valid execution violating a semantic guarantee yields a signature forgery, a hash collision, or a quantified deployment-failure event. The Prime Execution Model (PEM) separates planning, enforcement, effect, and recordkeeping into distinct authority planes; a lemma reduces trace completeness to Effector-exclusive credentialing. An Execution Attestation Certificate is issued only when PoE = 1. In a single-node TypeScript prototype, PoE adds approximately 2.7 ms on a minimal flow and 4.4% overhead on concurrent batch workloads; a standard eight-event trace compresses to approximately 1.1 KB; injected Gateway-bypass and trace-mutation attacks are rejected. PoE does not replace consensus, TEEs, or zkVMs; it binds authorization, effect, history, and replay into a single runtime-checkable object so that governed execution becomes attestable under contract.
ロングコンテキストサービスのタスク品質とシステムパフォーマンスにわたる KV キャッシュ最適化のベンチマーク
大規模な言語モデルの提供は、ロングコンテキストのワークロード下での KV キャッシュの増大によってますます制限されていますが、既存の KV キャッシュ圧縮技術は、異なるモデル、タスク、予算、およびサービス スタックで評価されているため、比較することが困難です。このペーパーでは、KIVI、TurboQuant、SnapKV、CaM などの量子化、プルーニング、マージにわたる代表的な KV キャッシュ最適化メカニズムのワークロード認識ベンチマークを示し、LongBench スタイルのマルチドキュメント QA、単一ドキュメント QA、少数ショット学習、Llama-3.1-8B-Instruct と要約ワークロードで評価しました。ミストラル-7B-命令-v0.3。このベンチマークは、タスクの品質、平均出力スループット、最初のトークンまでの平均時間、コンテキスト長バケット全体での実現圧縮率を測定します。結果は、圧縮率だけではエンドツーエンドのパフォーマンスを予測するのは不十分であることを示しています。 KIVI4 はモデル全体で最も安定した品質を提供し、SnapKV は最強のロングコンテキスト スループットを提供し、CaM は選択された QA ワークロードで大きな利益をもたらしますが、品質と実現された圧縮率の両方で大幅なワークロード感度を示します。これらの発見により、画一的な圧縮ではなく、ワークロードを意識した KV キャッシュ メカニズムの選択が促進され、ロング コンテキスト サービス システムの導入ガイダンスが提供されます。
原文 (English)
Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving
Large language model serving is increasingly limited by KV-cache growth under long-context workloads, yet existing KV-cache compression techniques are difficult to compare because they were evaluated on different models, tasks, budgets, and serving stacks. This paper presents a workload-aware benchmark of representative KV-cache optimization mechanisms spanning quantization, pruning, and merging, including KIVI, TurboQuant, SnapKV, and CaM, evaluated on LongBench-style multi-document QA, single-document QA, few-shot learning, and summarization workloads using Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3. The benchmark measures task quality, mean output throughput, mean time-to-first-token, and realized compression ratio across context-length buckets. The results show that the compression ratio alone is a poor predictor of end-to-end performance. KIVI4 provides the most stable quality across models, SnapKV delivers the strongest long-context throughput, and CaM yields large gains on selected QA workloads but exhibits substantial workload sensitivity in both quality and realized compression ratio. These findings motivate workload-aware selection of KV-cache mechanisms rather than one-size-fits-all compression and provide deployment guidance for long-context serving systems.
人工知能研究における触媒論文: 2017 年から 2025 年までの ICLR の展望
word2vec、Transformer、大規模な事前トレーニング、人間のフィードバックからの強化学習など、少数の方法論的貢献が、過去 10 年間にわたって NLP と AI 研究を再形成してきました。 OpenReview は、ICLR の投稿ごとに査読者の数値スコアと承認/拒否の決定を公開するようになりました。ただし、そのようなレビューシグナルが軌道を変える論文を投稿時に特定するかどうかは、コーパス規模ではまだテストされていません。私たちは、ICLR 2017--2025 の $36{,}113$ の論文についてこの質問に答え、\emph{触媒}、つまり子孫が将来の研究を明らかに方向付ける論文を特定します。 4 つの破壊性尺度 (統合/不安定化 (CD) インデックス、node2vec、方向性を認識した埋め込み破壊性尺度 (EDM)、および LLM ベースのセマンティック評価器) を比較し、5 種類の操作触媒分類法 (トピック イニシエーター、トピック ブリッジ、トピック内リダイレクター、同時、および認識の不整合) を定義します。 EDM は、引用度の高い ICLR 論文の特定でリードしています (AUC 0.83 ドル、CD の場合は 0.60 ドル、node2vec の場合は 0.49 ドル、LLM 評価者の場合は 0.42 ドル)。トピック間の引用フローは、年を一致させたコントロールと比較して、トピック イニシエーターが $7.55{\times}$ のトピック シェアの増加に先行し、トピック ブリッジが $11.52{\times}$ の増加に先行しています。査読スコアは基本的に将来の破壊性と直交していることがわかりました ($|\rho|{\leq}0.005$、受理された論文と拒否された論文の平均 EDM は区別できません、$p{=}0.11$)。
原文 (English)
Catalyst Papers in Artificial Intelligence Research: A Landscape on ICLR from 2017 to 2025
A small number of methodological contributions, including word2vec, the Transformer, large-scale pre-training, and reinforcement learning from human feedback, have reshaped NLP and AI research over the past decade. OpenReview now makes numeric reviewer scores and accept/reject decisions public for every ICLR submission. Whether such review signals identify trajectory-changing papers at submission time, however, remains untested at corpus scale. We answer this question on $36{,}113$ papers from ICLR 2017--2025, identifying \emph{catalysts}: papers whose descendants measurably redirect future research. We compare four disruptiveness measures (the Consolidation/Destabilization (CD) index, node2vec, the direction-aware Embedding Disruptiveness Measure (EDM), and an LLM-based semantic rater) and define a five-type operational catalyst taxonomy (topic initiator, topic bridge, within-topic redirector, simultaneous, and recognition-misaligned). EDM leads at identifying highly cited ICLR papers (AUC $0.83$ vs.\ $0.60$ for CD, $0.49$ for node2vec, and $0.42$ for the LLM rater). Topic initiators precede a $7.55{\times}$ topic-share growth and topic bridges precede an $11.52{\times}$ growth in cross-topic citation flow versus year-matched controls. We found that the peer review scores are essentially orthogonal to future disruptiveness ($|\rho|{\leq}0.005$; accepted and rejected papers have indistinguishable mean EDM, $p{=}0.11$).
アラブ大学の英語教室における AI ツール: 過去と未来を振り返る
この論文は、2023 年 1 月 1 日から 2025 年 8 月 31 日まで、アラブ大学の教室 (AUC) で第二外国語としての英語 (EL2) 学習者をサポートするために使用された AI ツールに関する実証研究を総合することを目的としています。データ ソースとして、Google Scholar、Web of Science、Scopus という 3 つの大規模なデータセットを利用しました。これらの有名なデータベースに対して PRISMA ガイド付き検索を使用して、公開された論文のみを含めました。検索プロセスの結果、184 件の研究が見つかりましたが、包含基準を満たした研究は 11 件のみでした。調査結果から、EL2 学習者は草案、改訂、練習において AI に対して前向きな態度をとっていることが明らかになりました。経験的な成果は、表面レベルの成果で最も一貫しており、高次の作文の質とスピーキングの熟練度の向上はさまざまであり、多くの場合、教師の調停に依存します。この論文は、EL2 の指導に証拠に基づいた AI の統合を求めるアラブの大学向けの研究課題と実践的なガイドラインを提案して締めくくられています。また、AI ツールへの過度の依存を減らすために、足場を組んだ統合、教師のトレーニング、内省的なタスクを推奨しています。
原文 (English)
AI tools in Arab University English classrooms: Looking back and forward
This paper aims to synthesize empirical research on AI tools used to support English as a second/foreign language (EL2) learners in Arab University classrooms (AUCs) between Jan 1st 2023 and Aug 31st 2025. We utilized 3 large datasets, namely Google Scholar, Web of Science, and Scopus as the data sources. Using PRISMA-guided searches across these well-known databases, we included only published articles. The search process results in 184 studies, but only 11 studies met the inclusion criteria. Findings unveil that EL2 learners have positive attitudes towards AI for drafting, revision, and practice. Empirical gains were most consistent for surface-level outcomes improvements in higher-order writing quality and speaking proficiency was mixed and often contingent on teacher mediation. The paper concludes by proposing a research agenda and practical guidelines for Arab universities seeking evidence-based AI integration in EL2 instruction. It also recommends scaffolded integration, teacher training, reflective tasks to reduce over-reliance on AI tools.
ギザギザの世界経済: フロンティア AI が国家経済を不平等に暴露する
フロンティア AI の労働市場への影響は労働者、企業、政策立案者にとって重要ですが、現在の証拠は一般的に少数の高所得経済国から得られています。フロンティア AI の能力は作業タスクによってばらつきがあり、人間の労働力をどのように割り当てるかについては国家経済によって異なります。職業レベルの曝露スコアと 141 か国の国際雇用データを組み合わせた国家 AI 曝露指標を導入します。高所得国は低所得国よりもかなり危険にさらされており、ヨーロッパと中央アジアはサハラ以南のアフリカよりも 50% 多く危険にさらされていることがわかりました。また、ジェンダーギャップも見られます。ホワイトカラーや販売職に女性が集中しているため、91%の国で女性の方が男性よりも露出度が高いのです。例外は、女性の雇用が依然として農業と家内企業に集中している国です。私たちは、Anthropic、Microsoft、OpenAI が発行する全国的な AI 導入統計を予測することを示すことで、全国的な AI エクスポージャーの推定値を検証します。直接的な曝露を超えて、私たちは、国をまたいだ所得依存による間接的な曝露の新たなメカニズムを特定します。タジキスタンなど一部の国は、外国人労働者による母国への送金に大きく依存している。タジキスタンのフロンティアAIへの直接的なエクスポージャは平均を下回っているが、タジキスタンのGDPの37パーセントがロシアからの送金であり、ロシアのエクスポージャが非常に高いため、タジキスタンの送金によるエクスポージャは平均を上回っている。私たちの調査によると、エクスポージャーの国ごとのばらつきは十分に大きく、米国や欧州の労働市場に合わせて調整された政策対応は一般化しない。
原文 (English)
The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies
Frontier AI's labor-market effects matter to workers, firms, and policymakers, but current evidence generally comes from a handful of high-income economies. The capabilities of frontier AI are jagged across work tasks and national economies diverge in how they allocate human labor. We introduce a national AI exposure metric that combines occupation-level exposure scores and international employment data for 141 countries. We find that high income countries are substantially more exposed than low income countries and that Europe and Central Asia are 50 percent more exposed than Sub-Saharan Africa. We also find a gender gap: women are more exposed than men in 91 percent of countries, driven by their concentration in white-collar and sales occupations. The exceptions are countries where women's employment remains concentrated in agriculture and household enterprises. We validate our national AI exposure estimates by showing they predict national AI adoption statistics published by Anthropic, Microsoft, and OpenAI. Beyond direct exposure, we identify a new mechanism for indirect exposure due to cross-country income dependencies. Some nations such as Tajikistan depend heavily on foreign workers remitting money back to their home countries: Tajikistan's direct exposure to frontier AI is below-average but because 37 percent of Tajikistan GDP is Russian remittance and Russia is very exposed, Tajikistan's remittance-accounted exposure becomes above-average. Our research shows that national variation in exposure is large enough that policy responses calibrated to U.S. or European labor markets will not generalize.
CCBENCH: 健康クエリを使用した暗黙的に通知された規範による LLM の文化的能力の評価
ステレオタイプ化せずにユーザーと公平に対話するには、AI モデルは文化的能力、つまり、静的な人口統計的特性に依存するのではなく、ユーザーの暗黙的にシグナル伝達される文化的価値観を推測して適応する能力を示す必要があります。私たちは、大規模言語モデル (LLM) で文化的能力を評価するためのフレームワークである CCBENCH を紹介します。これは、文化を文化的帰属の 2 つの状態としてではなく、規範遵守状態の連続体として扱います。健康に関するケーススタディとして、私たちは CCBENCH-Health を作成しました。これには、6 つの文化にわたるさまざまな規範遵守状態を示す、理論に基づいた 60 人のペルソナが含まれており、それぞれが 18 の現実的な対話を行っています。各ペルソナは、実際のユーザー フォーラムから抽出された 52 の本物の医療に関する質問に基づいて評価され、3,120 の固有のインタラクションが得られます。 5 つの主要なモデルをベンチマークすると、最も優れたモデルであっても、文化的に適切な応答を達成できる確率は 20 ~ 30% にすぎないことがわかります。会話履歴 (CoT) からの文化的に関連性のある手がかりに焦点を当てるように明示的に促されると、パフォーマンスは平均 3 ~ 5% 適度に向上します。ペルソナが文化的規範に従うのではなく、文化的規範を回避するときにモデルが最高のパフォーマンスを発揮することがわかり、永続的な非対称性が明らかになり、モデルが文化的手がかりに適応するよりも、組み込みのバイアスに合わせることを好むことが示唆されています。これは、文化的な手がかりから適切な健康上のアドバイスが得られることがほとんどないアフガニスタンの状況で特に観察されます (平均: 8.8%)。最後に、文化によって異なりますが、モデルは、明示的に述べられた文化的慣習よりも、暗黙的で文化的な会話スタイルに容易に適応する場合があることがわかりました。
原文 (English)
CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries
To interact with users fairly and without stereotyping, AI models must display cultural competency, i.e., the ability to infer and adapt to a user's implicitly signaled cultural values, rather than relying on static demographic traits. We introduce CCBENCH, a framework for evaluating cultural competency in large language models (LLMs), treating culture as a continuum of norm adherence states rather than as a binary state of cultural belongingness. As a case study on health, we create CCBENCH-Health, which includes 60 theoretically grounded personas exhibiting varied norm-adherence states across six cultures, each engaging in 18 realistic dialogues. Each persona is evaluated on 52 authentic healthcare questions drawn from real user forums, yielding 3,120 unique interactions. Benchmarking five leading models reveals that even the best achieve culturally appropriate responses only 20-30% of the time. When explicitly prompted to focus on culturally relevant cues from the conversational history (CoT), performance improves modestly by 3-5% on average. We find that models perform best when personas avoid cultural norms rather than follow them, revealing a persistent asymmetry, suggesting a preference in the models to align with built-in biases than adapt to cultural cues. This is especially observed in the Afghan context (Avg: 8.8%), where cultural cues rarely yield appropriate health advice. Finally, we find that models sometimes adapt more readily to implicit, cultural conversational styles than to explicitly stated cultural practices, though this varies across cultures.
Vibe コーディングを通じて AI を活用した学習テクノロジーを作成する幼稚園から高校までの教師のための指導フレームワーク
大規模な言語モデルは自然言語プロンプトからコードを生成し、プログラマー以外でも計算ソリューションを開発できる「バイブ コーディング」を可能にします。教師向けの Vibe コーディングは、デザイナーとしての教師の価値を高め、AI リテラシーを育成しながらテクノロジーの統合を向上させます。しかし、このプロセスをサポートするための体系的なガイダンスが不足しています。私たちは、幼稚園から高等学校までの教師がバイブコーディングを通じて AI を活用した学習テクノロジーを作成するのをサポートするフレームワークである GAIDE (教育者向けの AI 統合設計のための指導フレームワーク) を提案します。デザイン思考とINTERACTに基づいて構築された最初のフレームワークは、8週間のワークショップで3人の教師と4人の教員メンターによるCORDTRAインタラクション分析を通じて検証され、最終的なフレームワークが導き出されました。さらに、前後のインタビューの定性分析により、教師の AI リテラシーが向上していることがわかりました。調査結果は、専門能力開発における創造による学習の可能性を浮き彫りにしています。
原文 (English)
A Guiding Framework for K-12 Teachers in Creating AI-powered Learning Technologies through Vibe Coding
Large language models generate code from natural language prompts, enabling "vibe coding," which allows non-programmers to develop computational solutions. Vibe coding for teachers amplifies the value of teachers-as-designers, improving technology integration while fostering AI literacy. However, structured guidance on supporting this process is lacking. We propose GAIDE (A Guiding Framework for AI-Integrated Design for Educators), a framework that supports K-12 teachers in creating AI-powered learning technologies through vibe coding. The initial framework, built on Design Thinking and INTERACT, was validated through a CORDTRA interaction analysis of three teachers and four faculty mentors in an eight-week workshop to derive the final framework. Additionally, the qualitative analysis of pre- and post-interviews found an enhancement of teachers' AI literacy. Findings highlight the potential of learning-by-creating for professional development.
立場: AI が生成する CSAM を防止するには、AI の安全性への新しいアプローチが必要です
最新の人工知能 (AI) システムは、子供の安全に新たな重大なリスクをもたらします。 AI は、AI によって生成された児童性的虐待資料を作成し、児童の性的搾取を促進し、危害に対する障壁を減らすために悪用されることが増えています。この論文では、AIによって引き起こされる性的虐待から子供たちを守るには、AIの安全性への新しいアプローチが必要であると主張します。既存の安全技術は、児童の性的虐待の内容を取り巻く倫理的および法的制約と両立しないデータへのアクセス、透明性、および評価慣行を前提としています。これらの制約が、データセット監査の制限、レッドチーム化、防止の微調整など、新たな技術的課題をどのように生み出すかを検証します。次に、データセットのキュレーションとモデル設計から展開と長期保守に至るまで、AI 開発ライフサイクル全体にわたるオンラインでの児童の性的搾取と虐待に関する *15 の未解決の問題* を概説します。私たちは、理論上の AI の安全性と児童保護の現実の間のギャップを埋めるために、研究者、開発者、政策立案者に的を絞った推奨事項を提案します。私たちの取り組みは、AI によって促進される児童性的虐待の防止を AI 研究の中心的かつ安全性が重要な側面として再構築し、責任ある AI の原則を子どもの搾取に対する具体的な保護手段に変換する取り組みを動機付けることを目的としています。
原文 (English)
Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety
Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child sexual abuse material, facilitate child sexual exploitation, and reduce barriers to harm. In this paper, we argue that protecting children from AI-facilitated sexual abuse requires new approaches to AI safety. Existing safety techniques assume data accessibility, transparency, and evaluation practices that are incompatible with the ethical and legal constraints surrounding child sexual abuse material. We examine how these constraints create new technical challenges, such as limitations on dataset auditing, red teaming, and fine-tuning prevention. In turn, we outline *15 open problems* in online child sexual exploitation and abuse across the AI development lifecycle, from dataset curation and model design to deployment and long-term maintenance. We propose targeted recommendations for researchers, developers, and policymakers to bridge the gap between theoretical AI safety and the realities of child protection. Our work aims to reframe preventing AI-facilitated child sexual abuse as a central, safety-critical dimension for AI research, motivating work that translates responsible AI principles into concrete safeguards against the exploitation of children.
パターンベースの知識コンポーネントを使用したプログラミング学習コンテンツの自動推奨
プログラミングの入門指導は、基本的な概念の習得をサポートする実践的な練習と短い学習活動に依存しています。このような学習リソースは数多く存在しますが、時間のかかる専門家のキュレーションがなければ、指導的に意味のある方法でこれらの項目を整理しリンクすることは困難です。この研究では、パターンベースのナレッジコンポーネント (KC) を使用して、同様の概念を対象としたコードベースの学習リソースを自動的に識別する方法を調査します。私たちのアプローチでは、パターンベースの KC が各コード サンプルから抽出され、各アクティビティに関連付けられた KC セット間の類似性を測定することによって関連するアクティビティが特定されます。この方法は、意味的に重要なプログラミング パターンのレベルでの調整を活用することにより、文脈上適切で教育的に有用な推奨事項をサポートします。私たちは、インストラクターが概念的な類似性に基づいて項目をバンドルにグループ化した、専門家によって編成された Python 入門教材のコーパスに基づいてアプローチを評価します。結果は、当社のパターンベースの KC アプローチがこの専門家組織と一致するリソースを取得し、標準的なランキング評価全体で代表的な KC ベースおよび埋め込みベースのベースラインを上回ったことを示しています。全体として、このフレームワークは、プログラミング学習者向けの的を絞った概念指向のガイダンスをサポートし、インストラクターが大規模な指導コンテンツを整理、バンドル、推奨するのに役立ちます。
原文 (English)
Automated Recommendation of Programming Learning Content Using Pattern-based Knowledge Components
Introductory programming instruction relies on hands-on practice and short learning activities to support mastery of foundational concepts. Although many such learning resources exist, organizing and linking these items in instructionally meaningful ways is challenging without time-intensive expert curation. This study investigates the use of pattern-based Knowledge Components (KCs) to automatically identify code-based learning resources targeting similar concepts. In our approach, pattern-based KCs are extracted from each code sample, and related activities are identified by measuring similarity between the KC sets associated with each activity. By leveraging alignment at the level of semantically important programming patterns, this method supports contextually appropriate and pedagogically useful recommendations. We evaluate our approach on an expert-organized corpus of introductory Python materials in which instructors grouped items into bundles based on conceptual similarity. Results show that our pattern-based KC approach retrieves resources that align with this expert organization, and outperformed representative KC- and embedding-based baselines across standard ranking evaluations. Overall, the framework supports targeted, concept-oriented guidance for programming learners and can help instructors organize, bundle, and recommend instructional content at scale.
CANONIC: ガバナンスは編集物である
私たちは、デジタル成果物を大規模な証拠台帳にコンパイルする管理されたインテリジェンスである CANONIC を紹介します。大規模な言語モデルは、誰がチェックするよりも早く散文を生成し、失敗したオックスフォード言語は「slop」と名付けられ、2025年の今年の言葉に選ばれました。 CANONIC は、プログラムが整形式であるかどうかをコンパイラが判断する方法と同じように、コンテンツをコーパスに入れるかどうかを制御します。つまり、許可の境界で、機械的に、文法によって。ガバナンスは、コンパイラ理論の構文、スコープ解決、型システム層に 1 対 1 でマッピングされる 3 つの公理 (トライアド、継承、イントロスペクション) に還元され、承認は決定可能な線形時間チェックです。次に、4 つの制度にわたる事前に登録されたプロバイダー間のベンチマークを使用して、構造的承認がスロップアウトを維持しているかどうかを尋ねます。信頼できるコンテンツと信頼できないコンテンツを確実に分離する散文読みゲートはありません。スロップはアルゴリズムが計算するプロパティではありません。それはドメインの専門知識の判断です。したがって、ガバナンス層が傾きを決定するわけではありません。記録を監査可能に保ちます。すべての主張が定義、コミット、証拠ウィンドウに固定されており、再現可能でエンドツーエンドでチェック可能です。
原文 (English)
CANONIC: Governance Is Compilation
We present CANONIC: governed intelligence that compiles digital artifacts into an evidence ledger at scale. Large language models generate prose faster than anyone can check it, the failure Oxford Languages named 'slop', its 2025 Word of the Year. CANONIC governs whether content may enter a corpus the way a compiler decides whether a program is well-formed: mechanically, by a grammar, at the boundary of admission. Governance reduces to three axioms (Triad, Inheritance, Introspection) that map one-to-one onto compiler theory's syntax, scope-resolution, and type-system layers, and admission is a decidable, linear-time check. We then ask, with a pre-registered cross-provider benchmark across four regimes, whether structural admission keeps slop out. It does not: no prose-reading gate reliably separates reliable from unreliable content. Slop is not a property an algorithm computes. It is a verdict of domain expertise. So a governance layer does not decide slop; it keeps the record auditable -- every claim anchored to a definition, a commit, and an evidence window, reproducible and checkable end to end.
GenAI スキル バイパス: 大学生とスタッフの AI リテラシーの分岐経路をマッピングする
高等教育機関では、学生と職員の両方が生成 AI (GenAI) リテラシーを確実に開発することがますます期待されています。これに応えて、彼らは専門能力開発プログラムを導入し、GenAI のスキルを学生のカリキュラムに組み込んでいます。しかし、現在の教育枠組みは通常、GenAI リテラシーの直線的進歩を前提としており、創造的な応用に先立って基礎的な技術的理解が必要であることを暗示しています。この論文は、分類学に基づく自己評価手段 (n = 158) の心理測定分析を通じて、そのような仮定に疑問を呈します。私たちは、Rasch 測定理論と Guttman 順序付けを適用して、学生、学者、専門スタッフ全体にわたる GenAI スキルの潜在的に認識されている難易度の順序をマッピングしました。結果は、認識されている能力プロファイルの根本的な相違を明らかにしました。学者はより伝統的な直線的な道をたどる一方、学生は「逆転した」プロファイルを示し、基礎的な概念的理解を獲得する前に高レベルの創作タスクを習得することがよくあります。さらに、学生と学業の間のスキルの難易度の相関は弱かった(r = 0.188)。私たちは、この「スキルバイパス」が流暢さの脆弱な感覚を生み出し、プロンプトにおける高い自己効力感が AI メカニズムにおける低いリテラシーを覆い隠していると主張します。これらの発見は、「画一的な」カリキュラムに疑問を投げかけ、人間と AI の真の相乗効果を促進する診断主導型のモジュール型介入の経験的基礎を提供します。
原文 (English)
The GenAI Skill Bypass: Mapping Divergent Pathways of University Students and Staff AI Literacy
Higher education institutions are increasingly expected to ensure that both students and staff develop Generative AI (GenAI) literacies. In response, they are introducing professional development programs and embedding GenAI skills within student curricula. However, current educational frameworks typically assume a linear progression of GenAI literacy, implying that foundational technical understanding must precede creative application. This paper challenges such an assumption through a psychometric analysis of a taxonomy-based self-assessment instrument (n = 158). We applied Rasch measurement theory and Guttman ordering to map the latent perceived order of difficulty of GenAI skills across students, academics, and professional staff. Results reveal a fundamental divergence in perceived competence profiles: while academics follow a more traditional linear path, students exhibit an "inverted" profile, frequently mastering high-level creation tasks before acquiring foundational conceptual understanding. Furthermore, the correlation of skill difficulty between students and academics was weak (r = 0.188). We argue that this "skill bypass" creates a fragile sense of fluency, where high self-efficacy in prompting masks low literacy in AI mechanics. These findings challenge the "one-size-fits-all" curricula and provide the empirical basis for diagnostic-driven, modular interventions that foster genuine human-AI synergy.
なぜ AI は STEM 教育に新たな可能性をもたらすのでしょうか?トレンドと将来の課題に関する書誌学的分析
STEM 教育は、個別化と学際的な統合という課題に直面しています。 AI テクノロジーは新たな可能性をもたらしましたが、AI が STEM 教育エコシステムを再構築するメカニズムについては、体系的な調査が必要です。この研究では、書誌学的手法を使用して 2015 年から 2025 年までの 242 件の出版物を分析し、ナレッジ マップを構築して進化の軌跡を明らかにしました。この調査結果は、この分野がインテリジェントな個別指導システムから、LLM によって推進される探究ベースの学習と計算論的思考の育成へと変化したことを示しています。 AI の主な貢献は、知識を理解するための敷居を下げるインテリジェントな足場を提供することにあります。この意味で、AI は知識の伝達から能力開発への移行を促進する中核的な原動力です。
原文 (English)
Why does AI unlock new possibilities in STEM education? A Bibliometric Analysis of Trends and Future Agenda
STEM education faces challenges in personalization and interdisciplinary integration. AI technology has brought new possibilities, but the mechanisms by which AI reshapes the STEM education ecosystem require systematic investigation. This study employs bibliometric methods to analyze 242 publications from 2015-2025, constructing knowledge maps to reveal the evolutionary trajectory. The findings show that the field has transformed from intelligent tutoring systems to inquiry-based learning and computational thinking cultivation driven by LLMs. AI's key contribution lies in providing intelligent scaffolding that lowers the threshold for understanding knowledge. In this sense, AI is a core driving force promoting its shift from knowledge transmission to capability development.
無線ネットワークにおけるチャネル状態フィードバックを強化するための圧縮を使用した対照予測コーディング
正確かつタイムリーなチャネル状態情報 (CSI) は次世代無線システムにとって不可欠ですが、既存の研究では、学術界と現在の 3GPP 研究の両方において、CSI 圧縮と CSI 予測が別個の問題として扱われています。その結果、標準化された CSI フィードバック パイプライン内では、チャネルの経年劣化への対処が不十分なままです。この記事では、対照予測コーディング (CPC) を 3GPP 準拠の CSI 圧縮アーキテクチャに直接統合する、統合された圧縮予測フレームワークを提案します。高次元の CSI 行列を予測する代わりに、私たちのアプローチは将来の潜在表現を予測し、1-SGCS と InfoNCE の目的を組み合わせて再構成の忠実度と時間的予測コヒーレンスを共同で最適化します。この設計により、フィードバックのオーバーヘッドを増加させることなく、時間表現の学習が可能になります。量子化前にエンコードされた特徴に対して自己回帰モデリングを実行する CPC-before-Compression と、時間モデリングを基地局に移してユーザーのデバイスの複雑さを軽減する CPC-after-Compression の 2 つのバリエーションを紹介します。 Nokia、Oppo、および CATT の 3GPP 準拠のデータセットの評価では、圧縮前の CPC は 3GPP ベースラインよりも 32 倍低いデコーダ GFLOP で 90% 以上の再構築精度を達成する一方、圧縮後の CPC は同一のエンコーダ フットプリントと同じ 64 ビット フィードバック オーバーヘッドを維持することが示されています。提案されたフレームワークは、標準化されたパイプライン内で圧縮と予測を統合することにより、年齢を考慮した計算効率の高い CSI フィードバック ソリューションを提供します。ソース コードは https://github.com/AhmedRadwan02/cpc-3gpp で公開されています。
原文 (English)
Contrastive Predictive Coding with Compression for Enhanced Channel State Feedback in Wireless Networks
Accurate and timely channel state information (CSI) is essential for next-generation wireless systems, yet existing works treat CSI compression and CSI prediction as separate problems, both in academia and in current 3GPP studies. Consequently, channel aging remains insufficiently addressed within standardized CSI feedback pipelines. In this article, we propose a unified compression-prediction framework that integrates Contrastive Predictive Coding (CPC) directly into the 3GPP-compliant CSI compression architecture. Instead of predicting high-dimensional CSI matrices, our approach forecasts future latent representations and jointly optimizes reconstruction fidelity and temporal predictive coherence via a combined 1-SGCS and InfoNCE objective. This design enables temporal representation learning without increasing feedback overhead. We present two variants: CPC-before-Compression, which performs autoregressive modeling on encoded features prior to quantization, and CPC-after-Compression, which shifts temporal modeling to the base-station to reduce the complexity of users' devices. Evaluations on 3GPP-compliant datasets from Nokia, Oppo, and CATT show that CPC-before-Compression achieves over 90% reconstruction accuracy with 32x lower decoder GFLOPs than the 3GPP baseline, while CPC-after-Compression preserves an identical encoder footprint and the same 64-bit feedback overhead. By unifying compression and prediction within a standardized pipeline, the proposed framework provides an age-aware, computationally efficient CSI feedback solution. The source code is publicly available at: https://github.com/AhmedRadwan02/cpc-3gpp
AI が分類するとき: 何が行政とみなされるのか?
この研究では、学術的表現の代替システムが広範な行政 (PA) および人工知能関連行政 (AI-in-PA) の学問をどのように特定し、特徴付けるかを調査します。 Web of Science と OpenAlex を使用して、著者定義、引用主導、AI 支援表現に基づく 5 つのアプローチを比較します。結果は、コーパスのサイズ、出版物の種類、出版媒体、時間的発展、テーマのクラスタリングと構造における大きな違いを浮き彫りにしています。代替的なアプローチでは、同じ学問のさまざまなサブセットではなく、異なる知識領域を特定することが多く、そのため、表現全体で出版物や出版販売店に重複がないことから明らかなように、異なる表現が生成されます。この調査結果は、アルゴリズムによる知識の組織化が、学際的な学問がどのように分類、構造化、理解されるか、また認識論的にはその可視性、知的構造、境界がどのように表現されるかにますます影響を与えていることを示唆している。 AI を活用した学術的な分類と表現は中立的ではありませんが、解釈的なものであり、おそらく自己強化的であり、専門分野の境界の進化と適応を制約する可能性があります。人間の規律上の判断は不可欠であり、置き換えられるものではなく、補完されるものです。
原文 (English)
When AI Classifies: What Counts as Public Administration?
This study examines how alternative systems of scholarly representation identify and characterize broad public administration (PA) and artificial intelligence related public administration (AI-in-PA) scholarship. Using Web of Science and OpenAlex, it compares five approaches based on author-defined, citation-driven, and AI-assisted representations. The results highlight substantial differences in corpus size, publication types, publishing outlets, temporal development, and thematic clustering and structure. The alternative approaches often identify different knowledge domains instead of varied subsets of the same scholarship and therefore produce distinct representations, as evidenced by no overlap in publications and publishing outlets across representations. The findings suggest that algorithmic knowledge organization increasingly influences how interdisciplinary scholarship is classified, structured, and understood and, epistemologically, how its visibility, intellectual structure, and boundaries are represented. AI-enabled scholarly classifications and representations are not neutral but interpretative, likely self-reinforcing, and potentially constrain the evolution and adaptation of disciplinary boundaries. Human disciplinary judgment is essential and is complemented rather than replaced.
CHARLIE: 法医学における証拠推論のためのオンプレミスのマルチエージェント検索強化生成システム
デジタルフォレンジック環境における構造化された証拠処理のためのオンプレミスのマルチエージェント検索拡張生成 (RAG) システムである Charlie を紹介します。現代のフォレンジック ワークフローでは、トレーサビリティ、機密性、法令順守の厳格な要件の下で、異種混合の非構造化文書を大量に処理する必要があります。チャーリーは、ローカル検索、タスク分解、構造化メモリ、検証メカニズムを組み合わせた制御エージェント アーキテクチャを通じてこの課題に対処します。クラウドベースのシステムとは異なり、完全に組織インフラ内で動作し、データ主権と証拠の完全性を維持します。従来の RAG からエージェントベースのオーケストレーションへの移行を含むシステム アーキテクチャについて説明し、現実世界のフォレンジック シナリオでのアプリケーションを実証します。ケーススタディでは、Charlie がスケーラブルな複数文書データの抽出を可能にし、追跡可能性と監査可能性を維持しながら長期的な法医学的インテリジェンスの生成をサポートしていることが示されています。私たちの結果は、エージェントによって調整されたオンプレミスの RAG アーキテクチャが、法的および制度上の制約を損なうことなく、証拠ワークフローを効果的にサポートできることを示しています。チャーリーは、一か八かのフォレンジック環境に AI システムを導入するための実用的で再現可能な青写真を提供します。この原稿は、RELAF 2026 ワークショップで発表された論文のアーカイブ版です。
原文 (English)
CHARLIE: An On-Premise Multi-Agent Retrieval-Augmented Generation System for Evidential Reasoning in Forensic Science
We present Charlie, an on-premise multi-agent Retrieval-Augmented Generation (RAG) system for structured evidential processing in digital forensic environments. Contemporary forensic workflows must handle large volumes of heterogeneous and unstructured documents under strict requirements of traceability, confidentiality, and legal compliance. Charlie addresses this challenge through a controlled agent architecture that combines local retrieval, task decomposition, structured memory, and verification mechanisms. Unlike cloud-based systems, it operates entirely within institutional infrastructure, preserving data sovereignty and evidential integrity. We describe the systems architecture, including its transition from classical RAG to agent-based orchestration, and demonstrate its application in real-world forensic scenarios. Case studies show that Charlie enables scalable multi-document data extraction and supports longitudinal forensic intelligence generation while maintaining traceability and auditability. Our results indicate that agent-orchestrated, on-premise RAG architectures can effectively support evidential workflows without compromising legal and institutional constraints. Charlie provides a practical and reproducible blueprint for deploying AI systems in high-stakes forensic environments. This manuscript is an archival version of a paper presented at the RELAF 2026 Workshop.
モダリティの関連性はモダリティ ユーティリティではありません: コストを意識したマルチモーダル RAG のための事後選択的モダリティ エスカレーション
マルチモーダル検索拡張生成 (RAG) は、テキスト、表、画像などの異種モダリティから引き出された証拠をジェネレーターに根拠付けます。デプロイメントの主な選択は二者択一であり、モデルが答えを試みる前に行われます。安価なテキスト (+テーブル) パイプラインを実行するか、すべてのイメージに対して高価なビジョン言語モデル (VLM) の料金を支払うかのどちらかです。最近の適応システムは、どのモダリティが必要になるかを質問条件付き予測子から事前取得するモダリティまたは忠実度を選択することで、これを改善しています。これが間違った決定点であることを示します。 MultiModalQA でのオラクルのヘッドルーム分析を通じて、質問に対するモダリティの関連性は、そのモダリティが実際に正しく答える必要があるかどうかを予測する弱い指標であることがわかりました。ゴールド サポートに画像が含まれている質問の大部分は、それでもテキストと表だけから回答可能であり、明らかな視覚的な関連性に基づいてエスカレートする事前検索ルーターは、オラクルに比べて実質的に過剰にエスカレートします。私たちは \textbf{事後選択的モダリティ エスカレーション} を提案します。テキストと表から安価に回答し、どのモダリティが欠落しているかを特定する (クエリ、回答草稿、証拠) タプルに対して検証器を実行し、そこでのみ VLM 証拠の料金を支払います。次に、調整されたエスカレーション値ルーターが、期待される精度の向上が視覚的なコストに見合うかどうかを判断します。 MultiModalQA では、ルーターは常時オンの VLM パイプラインの精度を回復しながら、発行するビジュアル コールの数を大幅に減らし、オラクルのエスカレーション レートとの差をほとんど埋めます。その結果、取得の深さと推論ホップのために確立されたルーティング信号の階層が、単一のコストを意識した選択的エスカレーション ビューの下で 3 番目の軸 (モダリティ) に拡張されます。
原文 (English)
Modality Relevance is not Modality Utility: Post-hoc Selective Modality Escalation for Cost-Aware Multimodal RAG
Multimodal retrieval-augmented generation (RAG) grounds a generator in evidence drawn from heterogeneous modalities -- text, tables, and images. The dominant deployment choice is binary and made before the model has tried to answer: either run a cheap text(+table) pipeline, or pay for an expensive vision-language model (VLM) over every image. Recent adaptive systems improve on this by selecting the modality or fidelity pre-retrieval, from a question-conditioned predictor of which modality will be needed. We show that this is the wrong decision point. Through an oracle headroom analysis on MultiModalQA, we find that the relevance of a modality to a question is a weak predictor of whether that modality is actually needed to answer correctly: a large fraction of questions whose gold support includes an image are nonetheless answerable from text and tables alone, and a pre-retrieval router that escalates on apparent visual relevance over-escalates substantially relative to an oracle. We propose \textbf{post-hoc selective modality escalation}: answer cheaply from text and tables, run a verifier on the (query, draft answer, evidence) tuple that localizes which modality is missing, and pay for VLM evidence only there. A calibrated value-of-escalation router then decides whether the expected accuracy gain justifies the visual cost. On MultiModalQA, our router recovers the accuracy of an always-on VLM pipeline while issuing far fewer visual calls, and closes most of the gap to the oracle escalation rate. The result extends a routing-signal hierarchy established for retrieval depth and reasoning hops to a third axis -- modality -- under a single cost-aware selective-escalation view.
PORTS: 大規模な言語モデルを使用したツール選択のためのプリファレンスに最適化されたリトリーバー
外部ツールと大規模言語モデル (LLM) の統合は、複雑なタスクを実行するための有望なパラダイムとして浮上しています。 LLM は依然として大規模なツール コレクションを効果的に管理するのに苦労しているため、研究者は、入力長と遅延の制約に対処しながら、最も関連性の高いオプションを事前に選択するための検索ベースの方法の探索を開始しています。ただし、既存のレトリバーは、個別のトレーニング プロセスにより、ツール呼び出し LLM と調整されていないことがよくあります。この論文では、ツール選択を目的としたレトリバーを訓練するための新しいオッズ比優先最適化手法である PORTS を紹介します。私たちのアプローチは、凍結された LLM からの困惑にヒントを得た優先信号を使用して、ドキュメント文字列間の対照的な意味損失を共同で強制しながら、選択確率と下流のパフォーマンスの間の相関関係を最適化することで、検索ツールを微調整して有用なツールを見つけます。 PORTS の多用途性と、ツール選択の精度を大幅に向上させる機能は、さまざまな事前知識を備えた 6 つのデータセット、2 つのエンコーダー モデル、および 3 つの LLM に対する広範な実験を通じて実証されています。計算需要が低いため、当社の調整プロセスは新しいクエリやツールへの一般化を容易にし、進化するツールセットを伴う実用的なアプリケーションに価値があることが証明されています。
原文 (English)
PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models
Integrating external tools with Large Language Models (LLMs) has emerged as a promising paradigm for accomplishing complex tasks. Since LLMs still struggle to effectively manage large tool collections, researchers have begun exploring retrieval-based methods to pre-select the most relevant options, addressing input length and latency constraints. However, existing retrievers are often misaligned with tool-calling LLMs due to their separate training processes. This paper presents PORTS, a novel odds ratio preference optimization method for training retrievers aimed at tool selection. Using a perplexity-inspired preference signal from a frozen LLM, our approach fine-tunes a retriever to find helpful tools by optimizing the correlation between the selection probabilities and the downstream performances while jointly enforcing a contrastive semantic loss between documentation strings. The versatility of PORTS and its ability to significantly improve tool selection accuracy are demonstrated through extensive experiments on six datasets, two encoder models, and three LLMs with diverse prior knowledge. With low computational demands, our alignment process facilitates generalization to new queries and tools, proving valuable for practical applications with evolving toolsets.
大規模な科学コード検索: マルチドメインのデータセットとベンチマーク
科学者は研究ワークフローをサポートするためにオープンソース ツールにますます依存していますが、6 億を超える GitHub リポジトリの中から関連するソフトウェアを見つけることは依然として困難です。既存のコード検索ベンチマークは、一般的なソフトウェア エンジニアリング タスクに焦点を当てており、科学コンピューティングのドメイン固有の語彙やニーズを捉えることができません。私たちは、NASA 科学ミッション総局の 5 つの部門 (地球科学、天体物理学、惑星科学、太陽物理学、生物物理科学) にまたがる 5,264 の高品質でドメイン分類された科学リポジトリからなる厳選されたコーパスを提供します。これらは、クリーンな README、抽出されたトピック、およびクロールされたリンクからの追加コンテキストで強化されています。このコーパスに基づいて、2 つの新しい情報検索ベンチマークを導入します。(1) ドメイン科学者によって設計された 219 の専門家が厳選したクエリを含むリポジトリ検索ベンチマーク、(2) 7 つのプログラミング言語にわたる 117,950 のコード スニペットと 119,720 のクエリを含む大規模なコード スニペット検索ベンチマーク。リポジトリ検索のベースライン評価では、科学分野全体でパフォーマンスに大きなばらつきがあることが明らかになりました。コード スニペットの検索も同様に困難であることが判明しており、科学コミュニティ全体で文書化の慣行、コーディング標準、プログラミング言語の慣例が異なるため、大幅なばらつきが生じます。すべてのデータセットとベンチマークは、科学ツールの発見に関する研究をサポートするために、HuggingFace で公開されています。
原文 (English)
Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark
Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub repositories remains challenging. Existing code search benchmarks focus on general software engineering tasks and fail to capture the domain-specific vocabulary and needs of scientific computing. We present a curated corpus of 5,264 high-quality, domain-classified scientific repositories spanning five NASA Science Mission Directorate divisions -- Earth Science, Astrophysics, Planetary Science, Heliophysics, and Biological & Physical Sciences -- enriched with cleaned READMEs, extracted topics, and additional context from crawled links. Building on this corpus, we introduce two novel information retrieval benchmarks: (1) a repository search benchmark with 219 expert-curated queries designed by domain scientists, and (2) a large-scale code snippet retrieval benchmark containing 117,950 code snippets and 119,720 queries across seven programming languages. Baseline evaluations on repository search reveal significant performance variation across scientific domains. Code snippet retrieval proves equally challenging, with substantial variation driven by differing documentation practices, coding standards, and programming language conventions across scientific communities. All datasets and benchmarks are publicly released on HuggingFace to support research on scientific tool discovery.
UWBセンシングおよびワークゾーン再構築のためのジオメトリ対応インフラストラクチャアンカー型デノイザー
インテリジェント交通システムには作業ゾーンの形状を正確に認識することが不可欠であり、超広帯域センシングはインフラ支援による再構築に低コストのアプローチを提供します。ただし、屋外の UWB 測距は、見通し外伝播、バースト ノイズ、ロングテール エラーによって劣化することが多く、ダウンストリームの空間再構成が歪む可能性があります。我々は、時間範囲モデリングを潜在的なアンカー レイアウト推定と決定論的な距離投影と組み合わせた、ジオメトリを認識したインフラストラクチャにアンカーされた学習フレームワークである GAIA を紹介します。 GAIA は、学習された距離を境界一貫性のある再構築に向けて調整しながら、教師ありタスクとして範囲ノイズ除去を保存します。私たちは、同期された UWB、GNSS、IMU 測定を使用して実世界の屋外 UWB データセットで GAIA を評価し、実データで校正されたストレス テスト シミュレーターを使用して堅牢性をさらにテストします。 GAIA は、評価されたフィルタリング ベースおよび学習ベースのベースラインの中で最も低い全体範囲 MSE と最も高いポリゴン IoU を達成し、PoseMLP と比較して MSE を 18.4% 削減し、ポリゴン IoU を 15.5% 改善しました。これらの結果は、ジオメトリを意識した範囲ノイズ除去が、空間的に一貫したワークゾーンの再構築に向けた効果的な方法を提供することを示しています。
原文 (English)
Geometry-Aware Infrastructure-Anchored Denoiser for UWB Sensing and Work-Zone Reconstruction
Accurate work-zone geometry perception is critical for intelligent transportation systems, and ultra-wideband sensing offers a low-cost approach for infrastructure-aided reconstruction. However, outdoor UWB ranging is often degraded by non-line-of-sight propagation, burst noise, and long-tail errors, which can distort downstream spatial reconstruction. We present GAIA, a geometry-aware, infrastructure-anchored learning framework that couples temporal range modeling with latent anchor-layout estimation and deterministic distance projection. GAIA preserves range denoising as the supervised task while orienting the learned distances toward boundary-consistent reconstruction. We evaluate GAIA on a real-world outdoor UWB dataset with synchronized UWB, GNSS, and IMU measurements, and further test robustness using a real-data-calibrated stress-test simulator. GAIA achieves the lowest overall range MSE and highest polygon IoU among evaluated filtering-based and learning-based baselines, reducing MSE by 18.4% and improving polygon IoU by 15.5% over PoseMLP. These results show that geometry-aware range denoising provides an effective path toward spatially coherent work-zone reconstruction.
粒度のパラドックス: 時間的離散がどのようにサンプル内の適合を増大させ、サンプル外の誤差を増大させるのか
この論文では、時系列予測における「粒度のパラドックス」について考察します。このパラドックスでは、より細かい時間的分解 (例: 月次から週次/日次) により、サンプル内診断とデータセット サイズ (N) が向上しますが、より長い期間 (H) にわたって再帰的に誤差が増大するため、サンプル外の精度が低下します。逆に、粗い集計 (年次) では、再帰的な誤差の伝播が排除されますが、推定者が利用できるデータは減少します。私たちはこのトレードオフを形式化し、13 年間の公共調達データセットを使用して、ナイーブ、統計、機械学習、深層学習のアーキテクチャにまたがる 10 のモデルを 6 つの粒度にわたってベンチマークします。実証結果は、非単調なしきい値構造を明らかにしています。再帰的自己回帰モデルと季節モデルは、高頻度の予測下では大幅に劣化します (たとえば、Holt-Winters のテスト R 二乗は -151、TPFE は 100 に達します)。 LSTM は U 字型の誤差曲線を描き、月次 (19.66%) から隔週 (35.94%) まで悪化し、日次での誤差伝播ペナルティ (TPFE 4.35%、R 二乗 0.66) を克服するまで、線形回帰はすべての粒度 (TPFE 16.3 ~ 17.0%) にわたって安定していることが確認されています。このパラドックスは、モデルの複雑さではなく、再帰的フィードバック トポロジによって引き起こされること、および標準のポイントごとのメトリクス (RMSE、MAE) が累積誤差の伝播を体系的にマスクすること、および目標依存の累積メトリクスを使用せずに予測を評価すると、モデルの妥当性について誤解を招く評価が生成されることを、粒度全体で累積 TPFE に対してポイントごとのメトリクスの方向性を比較するコンセンサス/ディスセンサス診断を導入し、標準の診断が系統的エラーをマスクするモデルの特定を可能にします。伝播。
原文 (English)
The Granularity Paradox: How Temporal Disaggregation Inflates In-Sample Fit and Compounds Out-of-Sample Error
This paper explores the "Granularity Paradox" in time-series forecasting, wherein finer temporal disaggregation (e.g., Monthly to Weekly/Daily) improves in-sample diagnostics and dataset size (N), but degrades out-of-sample accuracy due to recursive error compounding over longer horizons (H). Conversely, coarse aggregation (Annual) eliminates recursive error propagation but reduces data available to estimators. We formalize this trade-off and benchmark 10 models - spanning na\"ive, statistical, machine learning, and deep learning architectures - across six granularities using a 13-year public procurement dataset. The empirical results reveal a non-monotonic threshold structure: recursive autoregressive and seasonal models degrade substantially under high-frequency forecasting (e.g., Holt-Winters reaches a Test R-squared of -151 and TPFE of 425.85% at the Daily grain), while the LSTM traces a U-shaped error curve, worsening from Monthly (19.66%) through Bi-Weekly (35.94%) before overcoming the error propagation penalty at Daily (TPFE of 4.35%, R-squared of 0.66). Linear Regression remains stable across all granularities (16.3-17.0% TPFE), confirming that the paradox is driven by recursive feedback topology, not model complexity. The results demonstrate that standard pointwise metrics (RMSE, MAE) systematically mask cumulative error propagation, and that evaluating forecasts without goal-dependent cumulative metrics produces misleading assessments of model adequacy. We introduce a consensus-dissensus diagnostic comparing the directional behaviour of pointwise metrics against cumulative TPFE across granularities, enabling the identification of models whose standard diagnostics mask systematic error propagation.
制御性と可観測性のテストによるディープ ニューラル ネットワークの経験的最小実現圧縮
ディープ ニューラル ネットワークには、多くの場合、相当な隠れ状態の冗長性が含まれていますが、ほとんどの圧縮手法は、内部状態の動的役割を明示的に特徴付けることなく、重み、ニューロン、または量子化表現に直接作用します。この論文では、ディープ ニューラル ネットワークの経験的な状態順序削減のための制御可能性と観測可能性のフレームワークを提案します。訓練されたネットワークを深さインデックス付きの非線形動的システムとして見ることで、隠れ状態のスナップショットと出力ヤコビアンからデータ駆動型の到達可能性、可観測性、バランスのとれたグラミアンを構築します。結果として得られる A/B/C テストは、レイヤーごとの到達可能ランク、観測可能ランク、および共同到達可能ランク (観測可能ランク) を推定します。これらのランクは、隠れ状態の冗長性の診断手段としてだけでなく、実現された縮小ネットワークの実際の圧縮層幅としても使用されます。 MNIST と CIFAR-10 の実験では、提案されたバランスのとれた実現を、投影ベースの削減、非構造化枝刈り、構造化枝刈り、低ランク SVD、動的 INT8 量子化、および線形ベースラインと比較します。 MNIST では、4 層 SiLU DNN は状態次数 1024 から 277 に削減され、完全モデルの 96.60% と比較して 95.45% の精度を維持しながら、72.95% の状態圧縮と 73.48% のパラメータ圧縮が得られます。 CIFAR-10 では、より大きな SiLU DNN が状態次数 4608 から 1339 に削減され、70.94% の状態圧縮と 83.09% のパラメータ圧縮が得られ、その一方で精度は 54.45% から 54.44% に維持され、CUDA 推論レイテンシは約 3 倍削減されます。この結果は、バランスのとれた到達可能ランクと観察可能ランクが、精度をほとんどまたはまったく失わずにコンパクトなニューラル アーキテクチャを設計するための原則に基づいた経験的な最小実現基準を提供することを示しています。
原文 (English)
Empirical Minimal-Realisation Compression of Deep Neural Networks via Controllability-Observability Tests
Deep neural networks often contain substantial hidden-state redundancy, but most compression methods operate directly on weights, neurons, or quantised representations without explicitly characterising the dynamical role of internal states. This paper proposes a controllability-observability framework for empirical state-order reduction of deep neural networks. By viewing a trained network as a depth-indexed nonlinear dynamical system, we construct data-driven reachability, observability, and balanced Gramians from hidden-state snapshots and output Jacobians. The resulting A/B/C tests estimate layer-wise reachable, observable, and jointly reachable--observable ranks. These ranks are then used not only as diagnostic measures of hidden-state redundancy, but also as actual compressed layer widths for realised reduced networks. Experiments on MNIST and CIFAR-10 compare the proposed balanced realisation against projection-based reduction, unstructured pruning, structured pruning, low-rank SVD, dynamic INT8 quantisation, and linear baselines. On MNIST, a four-layer SiLU DNN is reduced from state order 1024 to 277, giving 72.95% state compression and 73.48% parameter compression, while maintaining 95.45% accuracy compared with 96.60% for the full model. On CIFAR-10, a larger SiLU DNN is reduced from state order 4608 to 1339, giving 70.94% state compression and 83.09% parameter compression, while preserving accuracy from 54.45% to 54.44% and reducing CUDA inference latency by approximately 3X. The results show that balanced reachable-observable ranks provide a principled empirical minimal-realisation criterion for designing compact neural architectures with little or no loss in accuracy.
オフライン強化学習による LLM エージェント ハーネスの制御方法の学習
大規模言語モデル (LLM) エージェントは通常、プロンプト、モデル、または手書きのワークフローを変更することによって改善されますが、モデル周辺の実行ハーネスは固定インフラストラクチャとして扱われます。私たちは、このハーネス自体が学習可能な制御層であると主張します。ハーネス操作を有限ホライズン ハーネス MDP として形式化します。LLM エグゼキュータはフリーズしたままで、軽量コントローラが構造的な実行アクションを選択します。コントローラーは、ターミナル タスクのルーブリック報酬のみを使用したアドバンテージ重み付け回帰を使用して、オフライン ロールアウトからトレーニングされます。また、最終的なタスクの品質を、最終的な答えが正しいかどうかだけでなく、ハーネスが信頼できる実行パターンに従っているかどうかを測定する事後のハーネス成熟度スコアから分離します。この分離により、ハーネス学習の有限バッファ ビューが得られます。最終品質の向上にはオフライン バッファでのハイリターンのサポートが必要ですが、プロセスの動作はアドバンテージを重視したアクションと一致するたびに変化する可能性があります。 6 つの制御ドメインと 2 つのパブリック ベンチマーク アダプターにわたって、学習済みコントローラーは検証動作を一貫して改善し、最終タスクの品質を選択的に向上させます。適応されたタウベンチ リテール、適応された AgentBench DB-Bench、および調整された構造検証器を使用したコーディングで最大の利益が得られます。行動の複製と強制チェックに対するアブレーションは、その利益が模倣や単にチェックを追加することによって説明できないことを示しています。これらの結果は、ハーネス制御がフリーズした LLM エージェントの学習可能な層であることを特定する一方で、より優れたプロセス制御がより優れた最終的な答えになる場合、オフライン サポートが制限されることを示しています。
原文 (English)
Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning
Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the model is treated as fixed infrastructure. We argue that this harness is itself a learnable control layer. We formalize harness operation as a finite-horizon Harness MDP, where a lightweight controller selects structural execution actions while the LLM executor remains frozen. The controller is trained from offline rollouts using advantage-weighted regression with only terminal task-rubric rewards. We also separate final task quality from a post-hoc Harness Maturity Score, which measures whether the harness follows reliable execution patterns rather than only whether the final answer is correct. This separation gives a finite-buffer view of harness learning: final-quality gains require high-return support in the offline buffer, while process behavior can shift whenever it aligns with advantage-weighted actions. Across six controlled domains and two public-benchmark adapters, the learned controller consistently improves verification behavior and selectively improves final task quality, with the largest gains on adapted tau-bench retail, adapted AgentBench DB-Bench, and coding with a calibrated structural verifier. Ablations against behavior cloning and Forced CHECK show that the gains are not explained by imitation or by simply adding checks. These results identify harness control as a learnable layer for frozen LLM agents, while showing that offline support limits when better process control becomes better final answers.
AdaStop: DNN テスト選択のためのコストを意識した早期停止
ディープ ニューラル ネットワーク (DNN) をテストするための既存の方法では、固定されたラベル付け予算の下で、モデルの欠陥を明らかにする可能性が高いテスト入力を主に優先します。実際には、その予算を選択するのは困難です。テストが少なすぎると失敗を見逃しますが、多すぎると不必要なラベル付けコストが発生します。この研究では、DNN テストにおける停止問題を研究しています。私たちはテストを費用対効果の決定プロセスとして定式化します。このプロセスでは、入力のラベル付けにはコスト $c$ が発生し、障害の発見によって価値 $v$ が得られます。この定式化に基づいて、テスト中に限界故障発見率を推定し、推定率がしきい値 $\tau = c/v$ を下回った場合にラベル付けを停止するフレームワークである \textit{AdaStop} を紹介します。複数のデータセット、アーキテクチャ、および選択戦略にわたる実験により、ラベル付け予算のわずか $9$ ~ $31\%$ を使用して $65$ ~ $84\%$ の障害を検出できることがわかりました。
原文 (English)
AdaStop: Cost-Aware Early Stopping for DNN Test Selection
Existing methods for testing deep neural networks (DNNs) primarily prioritize test inputs likely to reveal model faults under a fixed labeling budget. In practice, choosing that budget is difficult: too little testing misses failures, while too much incurs unnecessary labeling costs. This work studies the stopping problem in DNN testing. We formulate testing as a cost--benefit decision process in which labeling an input incurs cost $c$ and discovering a fault yields value $v$. Based on this formulation, we introduce \textit{AdaStop}, a framework that estimates the marginal fault discovery rate during testing and stops labeling when the estimated rate falls below the threshold $\tau = c/v$. Experiments across multiple datasets, architectures, and selection strategies show that $65$--$84\%$ of faults can be discovered using only $9$--$31\%$ of the labeling budget.
デュアルユース生物学設定における調整された拒否と安全な有用性の評価
AI エージェントがライフ サイエンスのワークフローに組み込まれると、発見を迅速化する機能が悪用される可能性もあります。生物学的研究タスクのリスク特定と拒否行動のベンチマークである BioSecBench-Refusal を紹介します。このベンチマークは、61 のルーチン タスク (出版された文献に基づいた正当な分析) と、46 のレッドチーム タスク (実際の研究に似ているがバイオセキュリティ上の危険を隠す架空のシナリオ) を組み合わせています。 16 のモデル ハーネス構成全体で、拒否率はルーチン タスクで 7\% ~ 74\%、レッドチーム タスクで 1\% ~ 62\% の範囲であり、多くの構成では、隠された危険と同等またはそれ以上の率で正当なルーチン作業を拒否しました。拒否は、ほとんどの場合、エージェントによる推論の前に適用されたプロバイダー API フィルターによってトリガーされました。しかし、推論の余地を与えたモデルは、より現実的な脅威を特定できる可能性を示しました。私たちは、モデル開発者がエージェントバイオテクノロジーの研究開発の能力と注意を調整するためのツールとして BioSecBench-Refusal をリリースします。
原文 (English)
Evaluating calibrated refusal and safe usefulness in dual-use biology settings
As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk identification and refusal behavior for biological research tasks. The benchmark pairs 61 Routine tasks, legitimate analyses adapted from the published literature, with 46 Red-Team tasks, fictional scenarios that resemble real research but conceal a biosecurity hazard. Across 16 model-harness configurations, refusal rates ranged from 7\% to 74\% on Routine tasks and 1\% to 62\% on Red-Team tasks, with many configurations refusing legitimate Routine work at comparable or higher rates than concealed hazards. Refusals were most often triggered by provider API filters applied prior to agentic reasoning. However, models given room to reason showed the potential to identify more real threats. We release BioSecBench-Refusal as a tool for model developers to calibrate capability and caution for agentic biotech R\&D.
名目属性と順序属性を使用したカテゴリカル データ クラスタリングのための属性内距離の学習可能な重み付け
カテゴリカル データ クラスタリングの成功は、一般に、2 つのオブジェクト間の非類似度を測定する距離メトリックに大きく依存します。ただし、既存のクラスタリング手法のほとんどは、順序値の相対順序情報を考慮せずに相違度を計算する際に、2 つのカテゴリのサブタイプ、つまり名義属性と順序属性を同じ方法で扱います。さらに、名目属性と順序属性の間には相互依存性が存在する可能性があり、非類似性を示すために調査する価値があります。したがって、この論文では、グラフに似た観点から、名目属性値と序数属性値の本質的な違いと関連性を研究します。したがって、順序値間の順序関係を維持しながら、統一された方法で名目属性と順序属性の属性内距離を測定するための新しい距離計量を提案します。続いて、属性内距離の重みとデータ オブジェクトの分割の学習を 2 つの別々のステップではなく 1 つの学習パラダイムにまとめ、次善の解決策を回避する新しいクラスタリング アルゴリズムを提案します。実験により、提案されたアルゴリズムの有効性が既存のアルゴリズムと比較して示されます。
原文 (English)
Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes
The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects. However, most of the existing clustering methods treat the two categorical subtypes, i.e. nominal and ordinal attributes, in the same way when calculating the dissimilarity without considering the relative order information of the ordinal values. Moreover, there would exist interdependence among the nominal and ordinal attributes, which is worth exploring for indicating the dissimilarity. This paper will therefore study the intrinsic difference and connection of nominal and ordinal attribute values from a perspective akin to the graph. Accordingly, we propose a novel distance metric to measure the intra-attribute distances of nominal and ordinal attributes in a unified way, meanwhile preserving the order relationship among ordinal values. Subsequently, we propose a new clustering algorithm to make the learning of intra-attribute distance weights and partitions of data objects into a single learning paradigm rather than two separate steps, whereby circumventing a suboptimal solution. Experiments show the efficacy of the proposed algorithm in comparison with the existing counterparts.
CanvasAgent: ビジュアル ツール オーケストレーションによる複雑な画像の作成と編集を可能にする
複雑なイメージの作成と編集には、複数の生成モデルまたは編集モデルが必要になることがよくあります。ユーザーのリクエストには、画像の合成、オブジェクトの位置特定、領域のセグメント化、選択したコンテンツの編集、中間アセットの合成、テキストの読み取り、最終結果の強化などが含まれる場合があります。このようなタスクでは、マルチモーダル エージェントが知覚拡張推論から操作中心の視覚作成に移行します。この場合、ツールは視覚状態を単に検査するのではなく、積極的に変換する必要があります。ただし、既存のマルチモーダル ツール使用エージェントは、ほとんどが認識、検索、またはドメイン固有の編集用に最適化されており、実行可能な画像作成の軌跡に対する大規模な監視が不足しています。このペーパーでは、複雑な画像の作成と編集のための大規模なマルチモーダル ツール使用データセットである CanvasCraft と、マルチターン インタラクションを通じて異種のビジュアル ツールを調整する方法を学習するツール拡張マルチモーダル エージェントである \textbf{CanvasAgent} を紹介します。 CanvasCraft には、完全に注釈が付けられた 140K の実行可能トラジェクトリと 10K の RL タスク仕様が含まれています。 CanvasAgent は、まず SFT でトレーニングされて実行可能な推論とアクションの軌跡を学習し、次に結果レベルとプロセスレベルの信号を組み合わせたハイブリッド報酬を使用して GRPO で最適化されます。ロールアウト中、CanvasAgent は中間結果を検査し、ビジュアル アセットを追跡し、進化するビジュアル状態にツールの決定を適応させます。実験では、最終的な画像品質と軌跡の動作の両方を評価し、複雑なマルチツール画像作成ワークフローに対する CanvasAgent と提案されたデータセットの有効性を実証します。
原文 (English)
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, localizing objects, segmenting regions, editing selected content, compositing intermediate assets, reading text, and enhancing the final result. Such tasks shift multimodal agents from perception-augmented reasoning to manipulation-centered visual creation, where tools must actively transform visual states rather than merely inspect them. However, existing multimodal tool-use agents are mostly optimized for perception, search, or domain-specific editing, and lack large-scale supervision for executable image-creation trajectories. In this paper, we introduce CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and \textbf{CanvasAgent}, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction. CanvasCraft contains 140K fully annotated executable trajectories and 10K RL task specifications. CanvasAgent is first trained with SFT to learn executable reasoning-action trajectories, and is then optimized with GRPO using a hybrid reward that combines outcome- and process-level signals. During rollout, CanvasAgent inspects intermediate results, tracks visual assets, and adapts tool decisions to the evolving visual state. Experiments evaluate both final image quality and trajectory behavior, demonstrating the effectiveness of CanvasAgent and the proposed dataset for complex multi-tool image creation workflows.
推論効率の高い世界アクションモデルのための 4D 幾何事前確率の学習
ワールド アクション モデル (WAM) は、視覚的な未来のダイナミクスと実行可能なアクション シーケンスを共同でモデル化することにより、ロボット操作の強力な可能性を示しています。しかし、既存のビデオアクション共同トレーニング手法は主に外観指向のビデオ潜在を最適化するため、正確な操作に必要な時間的に進化するジオメトリを十分に捕捉できない可能性があります。私たちは、元の軽量推論グラフを維持しながら、アクション関連の 4D 幾何学的事前分布をビデオ アクション表現に注入するマルチエキスパート共同トレーニング世界アクション モデルである MECo-WAM を提案します。トレーニング中、MECo-WAM は、ビデオおよびアクションのエキスパートと、フリーズされた VGGT エンコーダーからのリレーショナル ターゲットによって監視される軽量 4D エキスパートを組み合わせます。非対称のエキスパートの可視性により、補助ジオメトリからアクション生成への非因果的なショートカットを防止します。展開されたビデオアクション経路に幾何学的な知識を移すために、減衰 4D 読み取りマスク アテンションを導入します。これにより、トレーニングの初期段階で制限された現在のフレームの幾何学的なガイダンスが提供され、この依存関係が徐々に削除されます。さらに、ロボットの動作に最も関連する視覚領域を強調しながら、フレーム内の幾何学的関係とその時間的展開を調整する、アクションを意識した時間幾何学的蒸留を提案します。デプロイメント時に、補助的な 4D コンポーネントはすべて削除されます。 LIBERO (98.2%)、RoboTwin 2.0 (92.6%)、および困難な現実世界の操作タスクに関する実験では、MECo-WAM が推論コストを増加させることなく操作パフォーマンスを向上させることが示されています。
原文 (English)
Learning 4D Geometric Priors for Inference-Efficient World Action Models
World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-action co-training methods primarily optimize appearance-oriented video latents, which may insufficiently capture the temporally evolving geometry required for precise manipulation. We propose MECo-WAM, a Multi-Expert Co-Training World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph. During training, MECo-WAM combines video and action experts with a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder. Asymmetric expert visibility prevents non-causal shortcuts from auxiliary geometry to action generation. To transfer geometric knowledge into the deployed video-action pathway, we introduce decayed 4D read-mask attention, which provides restricted current-frame geometric guidance early in training and progressively removes this dependency. We further propose action-aware temporal geometric distillation, which aligns within-frame geometric relations and their temporal evolution while emphasizing visual regions most relevant to robot actions. At deployment, all auxiliary 4D components are removed. Experiments on LIBERO (98.2%), RoboTwin 2.0 (92.6%), and challenging real-world manipulation tasks show that MECo-WAM improves manipulation performance without increasing inference cost.
構造的孤立の打破: コミュニティを意識したサンプリングと構造的エントロピーによるスケーラブルなグラフ クラスタリング
教師なしグラフ クラスタリングは、大規模ネットワークの根底にある意味論的パターンを明らかにするための基本的な手法です。グラフ対照学習は有望なパフォーマンスを示していますが、既存の手法はミニバッチ トレーニング中に「構造的分離」の問題に悩まされることが多く、グローバルなトポロジー分布を特徴付ける凝集したコミュニティ構造を捕捉することが困難になっています。これらの課題に対処するために、私たちは、コミュニティを意識したサンプリングと制約された構造エントロピーを相乗させることで構造の完全性を維持する、スケーラブルな教師なしグラフ クラスタリング フレームワークである SCISE を提案します。具体的には、最初に構造エントロピー コミュニティ制約オペレーター (SECC) を導入します。これは、コミュニティの断片化を緩和し、パーティションの結合を強化するために、制約された解空間内の構造情報を最適化します。次に、バッチ トレーニング中のグローバルな情報損失を防ぐために、ターゲット ノードのコミュニティ コンテキストをサンプリング バッチに組み込むコミュニティ対応サンプリング拡張 (CSampE) メカニズムを設計し、構造的な障壁を効果的に突破し、トポロジの整合性を維持します。最後に、バッチ内の構造類似性に基づいてエッジの重みを調整する構造対照学習 (StructCL) モジュールを考案し、エンコーダーが高次の構造空間での表現を学習するように導きます。 6 つの主流ベンチマーク データセットに対する広範な実験により、SCISE が最先端のアルゴリズムを大幅に上回るパフォーマンスを示し、アブレーション研究とロバスト性分析により、現実世界の大規模グラフに対する SCISE の有効性と信頼性がさらに検証されました。
原文 (English)
Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy
Unsupervised graph clustering is a fundamental technique for uncovering underlying semantic patterns in large-scale networks. Although Graph Contrastive Learning has demonstrated promising performance, existing methods often suffer from the "structural isolation" issue during mini-batch training, making it challenging to capture cohesive community structures that characterize the global topological distribution. To address these challenges, we propose SCISE, a Scalable unsupervised graph Clustering framework that preserves structural Integrity by synergizing community-aware sampling with constrained Structural Entropy. Specifically, we first introduce the Structural Entropy Community Constraint operator (SECC), which optimizes structural information within a constrained solution space to mitigate community fragmentation and enhance partition cohesion. Second, to prevent global information loss during batch training, we design a Community-Aware Sampling Expansion (CSampE) mechanism that incorporates the community context of target nodes into sampling batches, effectively breaking structural barriers and preserving topological integrity. Finally, we devise a Structural Contrastive Learning (StructCL) module that refines edge weights based on intra-batch structural similarity, guiding the encoder to learn representations in a higher-order structural space. Extensive experiments on six mainstream benchmark datasets demonstrate that SCISE significantly outperforms state-of-the-art algorithms, with ablation studies and robustness analyses further validating its effectiveness and reliability for real-world large-scale graphs.
KAT-Coder-V2.5 テクニカル レポート
KAT-Coder-V2.5 は、シングルターン コード ジェネレーターとしてではなく、実際の実行可能なリポジトリ内で自律的に動作するようにトレーニングされた、コーディングに重点を置いたエージェント モデルです。その機能のボトルネックとなるのは、モデルの規模というよりは、再現可能な環境、検証可能な報酬、価値の高い軌跡の欠如です。これらの点については、エンドツーエンドのエージェントのトレーニング後フレームワークで対処します。 AutoBuilder は、多言語リポジトリを大規模なフェイルツーパス検証とパスツーパス検証を備えたサンドボックス環境に再構築し、そこから自己完結型のタスク仕様を再生成し、ニアミス軌跡を回復し、プロセス認識フィルタリングを通じて監視を抽出します。一方、KwaiClawEnv は、実行可能サービスと実際のタスク シードから大規模なツール使用軌跡を合成します。私たちは、ハーネスのランダム化、信頼性を強化したサンドボックス、非対称アクター、後知恵で強化された価値推定を備えたクリティカル PPO、ハーネス指向の報酬フレームワークを使用して強化学習をさらに拡張し、マルチ教師オンポリシー蒸留を通じて SWE、エージェントクロー、Webコーディングの専門家を統合します。 6 つのソフトウェア エンジニアリングとエージェントのベンチマーク全体で、KAT-Coder-V2.5 は、PinchBench でエージェント ツールの使用において最高の結果をもたらし、リポジトリ レベルのソフトウェア エンジニアリングでは最前線の Opus 4.8 に次ぐ 2 位にランクされています。当社のサービスは https://streamlake.com/product/kat-coder でご利用いただけます。
原文 (English)
KAT-Coder-V2.5 Technical Report
We present KAT-Coder-V2.5, a coding-focused agentic model trained to act autonomously inside real, executable repositories rather than as a single-turn code generator. Its capability is bottlenecked less by model scale than by the scarcity of reproducible environments, verifiable rewards, and high-value trajectories, which we address with an end-to-end agentic post-training framework. AutoBuilder reconstructs multilingual repositories into sandboxed environments with fail-to-pass and pass-to-pass verification at scale, from which we regenerate self-contained task specifications, recover near-miss trajectories, and distill supervision through process-aware filtering, while KwaiClawEnv synthesizes large-scale tool-use trajectories from executable services and real task seeds. We further scale reinforcement learning with harness randomization, a reliability-hardened sandbox, an asymmetric actor--critic PPO with hindsight-augmented value estimation, and a harness-oriented reward framework, and unify SWE, Agent-Claw, and WebCoding experts via Multi-Teacher On-Policy Distillation. Across six software-engineering and agentic benchmarks, KAT-Coder-V2.5 delivers the best agentic tool-use result on PinchBench and ranks second only to the frontier Opus 4.8 on repository-level software engineering. Our service is available at https://streamlake.com/product/kat-coder.
単一カメラと単一光源による両眼視線推定
一般に認められている理論によれば、視線トラッカーの最小ハードウェア要件は、自由な頭の動きによる視線推定を実現するために 1 台のカメラと 2 つの光源です。ただし、モバイル デバイスでの視線追跡などの一部のシナリオでは、使用するコンポーネント、特に光源を少なくすることが望ましい場合があります。 1台のカメラと1つの光源による視線推定手法を提案する。 「仮想光源」が導入され、カメラに関して実際の光源と対称に幾何学的に配置され、取得された画像に「仮想輝き」が生成されます。 2 つの瞳孔間の距離とキャプチャされた画像内の 2 つの輝きの間の関係を利用して「仮想輝き」を推定し、2 つの光源が利用可能であると仮定して多項式回帰で視線を推定します。回帰法の新しい正規化係数が検証され、これは 1 回のグリント システムに実用的であることが判明しました。性能は許容範囲内であることが証明されていますが、実際の光源を 2 つ備えたシステムと比較すると劣化が見られます。
原文 (English)
Binocular Gaze Estimation with Single Camera and Single Light Source
According to commonly consented theories, the minimum hardware requirement for gaze tracker is one camera and two light sources to realize gaze estimation with free head movements. However, in some scenarios such as eye tracking on mobile devices, it is preferable to use less components, especially light sources. We propose a gaze estimation method with one camera and one light source. A "virtual light source" is introduced, which is geometrically placed symmetrically to the real light source with respect to the camera, and generates a "virtual glint" in the acquired image. We estimate the "virtual glint" by exploiting the relationship between the distance between two pupils and two glints in the captured image, and estimate the gaze with polynomial regression assuming two light sources are available. A new normalization factor for regression method is verified, which turns out to be practical for one-glint system. The performance is proved to be acceptable, while degradation is noticed compared to system with two actual light sources.
あなたの NPU は LLM に対応する準備ができていますか?モバイル LLM 推論における隠れた効率のボトルネックを分析する
大規模言語モデル (LLM) をモバイル デバイスに展開すると、プライバシーが強化され、遅延が短縮されますが、ハードウェアの非効率性によって深刻なボトルネックになります。我々は、5 つの主流フレームワーク (llama.cpp、GENIE など) と 3 つのハードウェア バックエンド (CPU、GPU、NPU) に独自にまたがる、モバイル LLM 推論に関する最初の包括的なクロスレイヤー測定研究を紹介します。この分析を可能にするために、従来のデバイスレベルの測定を超えて、最初のバックエンド固有のエネルギー属性を提供するきめ細かいプロファイリング ツールである PowerBench を開発しました。私たちの調査では、次の 3 つの重要な洞察が得られます。 (1) フレームワークに起因するパフォーマンスのギャップは、NPU 上で大幅に増幅され、オフロードと量子化戦略が発散するため、カスタム オペレーターを使用すると最大 10 倍に達します。 (2) NPU はコンピューティング バウンドのプリフィルに優れ、CPU はメモリ バウンドのデコードにおいて他のすべてのバックエンドよりも優れているという、明確なフェーズ スプリットを特定しました。これは、NPU が大規模で固定形状のワークロードを優先するためであり、デコードの小規模なカーネルの動的な性質と矛盾します。 (3) バックエンド固有のプロファイリングにより、以前の作業で見逃していた大幅なスケジューリングのヘッドルームが明らかになります。最適ではないスレッド構成、調整されていない NPU スリープ レイテンシー、および CPU ポーリング間隔により、最大 40% のエネルギーが無駄になります。これらの発見を活用して、モバイル LLM 推論のためのエネルギー指向のベスト プラクティス構成を提示します。この構成により、3 つのデータセット全体で NPU バックエンドのエネルギー消費を最大 54.8% 削減できると推定されています。
原文 (English)
Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference
Deploying Large Language Models (LLMs) on mobile devices enhances privacy and reduces latency, but is severely bottlenecked by hardware inefficiency. We present the first comprehensive, cross-layer measurement study of mobile LLM inference, uniquely spanning five mainstream frameworks (e.g., llama.cpp, GENIE) and three hardware backends (CPU, GPU, NPU). To enable this analysis, we develop PowerBench, a fine-grained profiling tool that provides the first backend-specific energy attribution, moving beyond traditional device-level measurements. Our study yields three critical insights: (1) Framework-induced performance gaps are substantially amplified on NPUs, reaching up to 10x using custom operators due to divergent offloading and quantization strategies. (2) We identify a distinct phase split where NPUs excel at compute-bound prefilling, while CPUs outperform all other backends in memory-bound decoding. This is driven by the NPU's preference for large, fixed-shape workloads, which conflicts with the small-kernel, dynamic nature of decoding. (3) Backend-specific profiling uncovers substantial scheduling headroom missed by prior work. Suboptimal thread configurations, uncoordinated NPU sleep latencies, and CPU polling intervals result in up to 40% energy waste. Leveraging these findings, we present an energy-oriented best-practice configuration for mobile LLM inference. We estimate that this configuration could reduce energy consumption by up to 54.8% on the NPU backend across three datasets.
マルチエージェント大規模言語モデル会話における意思決定プロトコル
大規模言語モデル (LLM) のタスク パフォーマンスを向上させることは不可欠ですが、これらのモデルをスケーリングするには、収益の減少や高コストなどの重大な課題に直面しています。マルチエージェント システム (MAS) は、タスクを専門のエージェントに分散して全体的なタスクのパフォーマンスを向上させることで、有望なソリューションを提供します。これにより、ディスカッションと意思決定のプロセスによるテスト時間の増加を犠牲にして、トレーニング コストを削減できます。意思決定プロトコルは、複数のエージェントが協力して最終的なソリューションを作成する方法を指定するため、MAS の重要なコンポーネントです。この論文では、マルチエージェント LLM (MALLM) フレームワークを紹介します。このフレームワークは、会話型タスク解決のためのマルチエージェントのディスカッションをシミュレートするために、さまざまな意思決定プロトコル (投票、コンセンサス、および裁判官の決定メカニズム) を実装および評価します。単一の意思決定プロトコルを使用したり、限られたデータセットでテストした以前の研究とは異なり、この研究では、知識ベースのデータセット (MMLU、MMLU-Pro、GPQA) からロジックベースのデータセット (StrategyQA、MuSR、Math-lvl-5、SQuAD 2.0) に至るまで、さまざまなタスクのセットに対するその影響を体系的に調べています。結果は、コンセンサス プロトコルは知識集約型の領域では優れている一方、投票および審査プロトコルはロジック ベースのタスクではより効果的であることを示しています。独立したソリューションの生成を通じて応答の多様性を高めることで、意思決定の質が向上しますが、意思決定プロセス中の情報アクセスの変更による影響は最小限に抑えられます。
原文 (English)
Decision Protocols in Multi-Agent Large Language Model Conversations
Improving the task performance of Large Language Models (LLMs) is essential, yet scaling these models faces significant challenges such as diminishing returns and high costs. Multi-Agent Systems (MAS) offer a promising solution by distributing tasks among specialized agents to improve the overall task performance. This can reduce training costs at the expense of increased test time due to the discussion and decision-making process. The decision protocol is a critical component of MAS because it specifies how multiple agents collaborate to create a final solution. This thesis introduces the Multi-Agent LLM (MALLM) framework, which implements and evaluates various decision protocols, namely voting, consensus, and judge decision mechanisms, to simulate multi-agent discussions for conversational task solving. Unlike previous work that used a single decision protocol or tested them on limited datasets, this study systematically examines their impact on a diverse set of tasks, ranging from knowledge-based datasets (MMLU, MMLU-Pro, GPQA) and logic-based datasets (StrategyQA, MuSR, Math-lvl-5, SQuAD 2.0). The results indicate that consensus protocols excel in knowledge-intensive domains while voting and judge protocols are more effective for logic-based tasks. Increasing response diversity through independent solution generation improves decision quality, while changes in information access during the decision process have minimal impact.
生成 AI ワークフローにおける権限と機密性
生成 AI (GenAI) システムは、トレーニングと記憶によるモデルのパラメーター、ライブ セッション中のコンテキスト ウィンドウ、および検索拡張生成 (RAG) 用のナレッジ データベースの 3 つの異なる方法でクライアント データを保存および処理します。各モードは、機密保持と法律専門家の特権に対して、異なる、そして多くの場合直観に反するリスクを生み出し、それぞれが特定のガバナンス対応を必要とします。特権と生成型 AI に関する最初の英米の判決、英国とムニル対内務省国務長官、米国対ヘップナー、それらの判決を読む必要がある正統な特権当局、および最近のコンピュータ サイエンスの研究に基づいて、データの保存と処理の 3 つのモードを実務家が利用できる用語で説明し、それぞれの法的影響を分析します。次に、イングランドとウェールズの弁護士を管理する規制の枠組み内、および業務上の過失に関する通常の原則の範囲内で分析を位置づけ、効果的な情報ガバナンスの基準(そしてそれに伴って過失と不正行為が測定される基準も)が変化していると主張します。私たちは主に SRA 規制対象の実務者を対象に執筆していますが、データ ガバナンス分析は、特権や職業上の機密の保護が実証可能な機密保持に依存するあらゆる法域に拡張できるように構成されています。この記事の最終的な目的は、法律サービスの専門家が GenAI システムにおける顕著なデータ漏洩リスクを理解し、それによってクライアント データやその他の機密資料に対する GenAI のより責任ある展開を促進できるようにすることです。
原文 (English)
Privilege and confidentiality in generative AI workflows
Generative AI (GenAI) systems store and process client data in three distinct ways: in the model's parameters through training and memorisation, in the context window during a live session, and in knowledge databases for retrieval-augmented generation (RAG). Each mode creates different and often counter-intuitive risks to confidentiality and legal professional privilege, and each calls for specific governance responses. Drawing on the first English and American decisions to address privilege and generative AI, UK and Munir v Secretary of State for the Home Department and United States v Heppner, on the orthodox privilege authorities against which those decisions must be read, and on recent computer science research, we explain the three modes of data storage and processing in terms accessible to practitioners and analyse the legal consequences of each. We then situate the analysis within the regulatory framework governing solicitors in England and Wales and within the ordinary principles of professional negligence, arguing that the standard of effective information governance (and with it the benchmark against which negligence and misconduct will be measured) is changing. Although we write primarily for SRA-regulated practitioners, our data-governance analysis is framed to extend to any jurisdiction in which the protection of privilege or professional secrecy depends on demonstrable confidentiality. The ultimate aim of this article is to help legal services professionals understand salient data leakage risks in GenAI systems and thereby facilitate a more responsible deployment of GenAI on client data and other sensitive material.
本番環境で安定したモデルを更新するためのフルレンジのバイナリ分類子キャリブレーション
敵対的な環境で実行される検出モデルは、悪意のあるディストリビューションが急速に変動することに直面しますが、良性のディストリビューションは比較的安定した状態を保つため、チームは常に再トレーニングと再展開を行い、新しい脅威に先んじて対応します。再トレーニングにより出力予測スコアが変化する傾向があり、モデルの下流ユーザーの負担が大きくなります。これらのセキュリティ指向モデルでは、すべての出力値にわたって一貫した偽陽性率 (FPR) が必要ですが、標準の確率校正手法は FPR 契約ではなくクラス確率を対象としています。 FPR 曲線全体をターゲットとする既存のキャリブレーション プリミティブの上に構築されたメソッドを導入し、展開全体で一貫した FPR の意味をスコアに与えます。 1 つのホールドアウト スプリットで観察された相対 FPR 誤差は、10% FPR から 0.1% FPR までで最大 2.3%、0.01% FPR で 7.2% でした。出荷されたアーティファクトは、1,000 ~ 1,000 万の良性サンプルのキャリブレーション セットにわたる測定値で 200 KB 未満にとどまります。
原文 (English)
Full-range Binary Classifier Calibration for Stable Model Updates in Production
Detection models running in adversarial environments face a malicious distribution that drifts rapidly while the benign distribution stays comparatively stable, so teams retrain and redeploy constantly to stay ahead of new threats. Retraining tends to change the output prediction scores, which breaks downstream users of the model. For these security-oriented models we need consistent false-positive rate (FPR) across all output values, whereas standard probability-calibration methods target class probability rather than an FPR contract. We introduce a method built on top of existing calibration primitives that targets the whole FPR curve, giving scores a consistent FPR meaning across deployments. On one held-out split, the observed relative FPR error was at most 2.3% from 10% down to 0.1% FPR and 7.2% at 0.01% FPR. The shipped artifact remains under 200 KB in measurements across calibration sets from 1K to 10M benign samples.
投影ビューと検証された構造化更新を備えた共有状態 LLM ワークフロー用の PatchOptic
エージェント ワークフローは、多くの場合、共有された構造化された状態で動作します。 LLM コンテキスト ウィンドウは限られているため、各モデル呼び出しは通常、現在のワークフロー ステップに必要な状態フラグメントのみが表示され、これは一般にプログレッシブ開示として知られるパターンです。最新のシステムは、grep のようなキーワード検索、検索拡張生成 (RAG)、抽象構文ツリー (AST) クエリ、およびタスク固有のエージェント スキルを使用して、このようなモデル対応ビューを構築します。これらのメソッドは読み取り側を管理しやすくしますが、ローカルで提案された書き換えが完全な状態に適用された後にいつ有効になるかを定義しません。欠けている部分は、ローカル更新とグローバル有効性の間の契約です。共有状態 LLM ワークフロー用の光学系インターフェイスである PatchOptic を紹介します。光学系は、構造化データのビューがどのように読み取られ、更新されるかを記述する構成的な双方向アクセサーです。 PatchOptic は、このビュー/更新の直感を借用し、投影された読み取りと検証された構造化パッチを通じてそれを実現します。各ワークフロー ステップは、投影された読み取りビュー、許可された書き込み領域、およびパッチ ソース領域を宣言します。実行時の強制を超えて、同じ宣言により、委任、サブワークフロー構成、および同じフェーズ内の独立したステップを並べ替えるための静的証明書をサポートするパスレベルのフットプリントが生成されます。この設計を、ドメイン全体で 46 件のベンチマークである PatchBench を使用して評価しました。結果は、強力なアクターの下で受け入れられた出力品質を維持しながら、予測された読み取りにより、報告された漏洩とトークン コストが削減されることを示しています。実行時検証は、宣言されたワークフロー契約違反をコミット前にブロックし、パッチ読み取りの強制は、非表示のソースを使用する侵害されたパッチ アーティファクトを拒否します。
原文 (English)
PatchOptic for Shared-State LLM Workflows with Projected Views and Verified Structured Updates
Agentic workflows often operate over shared, structured state. Because LLM context windows are limited, each model invocation is typically shown only the state fragment needed for the current workflow step, a pattern commonly known as progressive disclosure. Modern systems construct such model-facing views using grep-like keyword search, retrieval-augmented generation (RAG), abstract-syntax-tree (AST) queries, and task-specific agent skills. These methods make the read side manageable, but they do not define when a locally proposed rewrite is valid after it is applied back to the full state. The missing piece is a contract between local updates and global validity. We introduce PatchOptic, an optic-inspired interface for shared-state LLM workflows. Optics are compositional bidirectional accessors that describe how views of structured data are read and updated. PatchOptic borrows this view/update intuition and realizes it through projected reads and verified structured patches. Each workflow step declares a projected read view, an authorized write region, and a patch-source region. Beyond runtime enforcement, the same declaration yields a path-level footprint that supports delegation, sub-workflow composition, and static certificates for reordering independent steps within the same phase. We evaluate this design with PatchBench, a benchmark with 46 cases across domains. The results show that projected reads reduce reported leakage and token cost while preserving accepted-output quality under the strong actor. Runtime verification blocks declared workflow-contract violations before commit, and patch-read enforcement rejects compromised patch artifacts that use hidden sources.
リーン量子: AI 支援による量子情報の形式化に向けて
量子情報理論はエントロピー量に基づいて構築されています。その中でも、サンドイッチされた R\'enyi 相対エントロピーは、さまざまなアプリケーションでの基本的な発散であり、量子チャネル下でのデータ処理格差 (DPI) は基礎となる結果です。この研究では、理論分析のための再利用可能な正式なインフラストラクチャとして設計された、量子情報用のリーン 4 ライブラリを紹介します。ライブラリの中心的なデモンストレーションとして、有限次元量子系上の正の半定値演算子のサンドイッチ R\'enyi 相対エントロピーの DPI を形式化します。このライブラリは、有限次元システム、状態、チャネル、テンソル積、部分トレース、Choi 演算子、Kraus 表現、および Stinespring 表現の再利用可能なインターフェイスを含む、標準数学ライブラリ Mathlib と互換性のある有限次元量子力学の基底に依存しない演算子理論フレームワークを提供します。また、実連続関数微積分による演算子の単調性と凸性、ブロック演算子の陽性性、ヒルベルト・シュミット演算子空間、ジェンセンの演算子不等式、一般化されたパースペクティブ、演算子のべき乗平均、リーブ・安藤の軌跡不等式などの非可換な軌跡不等式のインフラストラクチャも構築します。このフレームワークに加えて、DPI のエントロピー固有の成分、つまりヤングおよび逆ヤング不等式によるサンドイッチ準エントロピーの変分公式、実数べき乗のテンソル積互換性、およびユニタリー群のハール測度を形式化します。これらのコンポーネントを組み合わせると、DPI のリーン形式化が生成され、結果として強力な準加法性が得られ、一般化量子シュタインの補題のリーン形式化を完了するために必要な最後の欠落コンポーネントが提供されます。より広範には、この開発は量子情報理論における将来の形式化された AI 支援研究のための機械チェック可能な基盤を提供します。
原文 (English)
Lean-Quantum: Toward AI-Assisted Formalization of Quantum Information
Quantum information theory is built on entropic quantities; among them, the sandwiched R\'enyi relative entropy is a fundamental divergence with various applications, and its data processing inequality (DPI) under quantum channels is a cornerstone result. In this work, we present a Lean 4 library for quantum information, designed as a reusable formal infrastructure for theoretical analysis. As a central demonstration of the library, we formalize the DPI for the sandwiched R\'enyi relative entropy for positive semidefinite operators on finite-dimensional quantum systems. The library provides a basis-independent operator-theoretic framework for finite-dimensional quantum mechanics compatible with the standard mathematical library Mathlib, including reusable interfaces for finite-dimensional systems, states, channels, tensor products, partial traces, Choi operators, Kraus representations, and Stinespring representations. It also builds infrastructure for noncommutative trace inequalities, including operator monotonicity and convexity via the real continuous functional calculus, block-operator positivity, Hilbert-Schmidt operator spaces, Jensen's operator inequality, generalized perspectives, operator power means, and Lieb-Ando trace inequalities. On top of this framework, we formalize entropy-specific ingredients for the DPI: variational formulas for the sandwiched quasi-entropy via Young and reverse-Young inequalities, tensor-product compatibility of real powers, and Haar measures on unitary groups. Together, these components yield a Lean formalization of the DPI, give strong subadditivity as a corollary, and provide the last missing component needed to complete the Lean formalization of the generalized quantum Stein's lemma. More broadly, the development provides machine-checkable foundations for future formalized and AI-assisted research in quantum information theory.
統計上の敵対者: ビジョン データセット内の自然なバックドアのような特徴
モデル固有の敵対的攻撃は広範囲に研究されています。私たちは、別の障害モードを研究しています。それは、悪意を持って挿入されずにバックドアのようなトリガーのように動作する、ビジョン データ内で自然に発生する統計信号です。これらのシグナルを統計的敵対者と呼びます。 Imagenet を分析して、特定のラベルと強く関連するパターンを見つけます。次に、統計的制御を使用して、候補信号からランダムな相関を除去します。最後に、これらの信号がモデルの予測を直接かつ予想どおりに変更することを示します。これらの統計上の敵対者は、一般的な破損よりも標的が絞られており、異なるモデル アーキテクチャ間で転送されます。これは、一部の脆弱性は単一モデルの特異性ではなく、データセットの構造と分布によって引き起こされることを示唆しています。私たちは、ポイズニングが存在しない場合でも、通常のデータセットには悪用可能な敵対的表面が含まれている可能性があると結論付け、データセットの監査では、偽の構造をバイアスや解釈可能性の失敗の原因としてだけでなく、ビジョン モデルの潜在的な攻撃対象表面としても扱うべきであると提案します。
原文 (English)
Statistical Adversaries: Natural Backdoor-like Features in Vision Datasets
Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signals in vision data that can behave like backdoor-like triggers without being maliciously inserted. We call these signals statistical adversaries. We analyse Imagenet to find patterns that are strongly linked to certain labels. We then use statistical controls to remove random correlations from our candidate signals. Finally, we demonstrate that these signals directly and predictably alter model predictions. These statistical adversaries are more targeted than generic corruptions and transfer across different model architectures. This suggests that some vulnerabilities are driven by dataset structure and distribution rather than a single model's idiosyncrasies. We conclude that ordinary datasets can contain exploitable adversarial surfaces even in the absence of poisoning, and suggest that dataset audits should treat spurious structure not only as a source of bias or interpretability failure, but also as a latent attack surface for vision models.
aiAuthZ: AI エージェント向けのオフホスト、アイデンティティにバインドされた承認
AI エージェントは検証できないテキストに基づいてツール呼び出しを発行するため、コンテキストの一部を制御する当事者は権威あるように見せかけることができます。実際のエージェントのインシデントの公開されたコーパスから派生した 8 つの攻撃シナリオに対して 15 の現代言語モデルを評価したところ、完全に評価されたモデル全体で拒否率が 100% から 38% まで変動することがわかりました。最も高価なモデルは、価格差が 20 倍あったにもかかわらず、攻撃の半分しか拒否しませんでした。私は、安全性の決定をエージェントのホストから移す認可ゲートウェイである aiAuthZ を紹介します。ツール呼び出しが実行される前に、ゲートウェイは、単一使用のノンスとタイムスタンプ ウィンドウにバインドされたメッセージごとの HMAC-SHA256 署名を使用して呼び出し元の ID を検証し、エージェントが読み取りも変更もできないロールベースの引数レベルのポリシーを評価します。すべての決定は SHA-256 ハッシュ チェーン監査ログに結合され、受け入れられた各メッセージから HMAC 認証された QR レシートが生成されます。これは、8 つの送信チャネル全体で 94% の平均検証を達成し、25 回の間違ったキーの試行で偽造が受け入れられることはありませんでした。ゲートウェイを設置すると、追加の決定遅延が 0.03 ミリ秒以内で、残りの攻撃成功率は 15 モデルすべてで 0% に低下します。 AgentDojo バンキング スイートでは、aiAuthZ は、正当な初回支払い 1 回を犠牲にして、評価されたエージェントが発行する 7 つの攻撃者向けツール呼び出しをすべてブロックしますが、スポットライト ベースラインにより 2 回の注入の成功が許可されます。同じインシデント コーパスからの対象範囲内の 9 つのケース スタディにわたって、aiAuthZ は、ID バインディングのないポリシー ベースラインの 9 つ中 4 つに対して、9 つ中 9 つをブロックします。ゲートウェイはモデルの欺瞞を防ぐものではありません。これにより、騙されたモデルが、それを経由するすべての通話に対して、検証されたユーザーの権限を超えて動作することが防止されます。実装とすべての実験は https://github.com/Sports-Vision-Inc/aiAuthZ でリリースされています。
原文 (English)
aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents
AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of authority. I evaluate 15 contemporary language models against eight attack scenarios derived from a published corpus of real agent incidents and find that refusal varies from 100% down to 38% across fully evaluated models; the most expensive model refused only half of the attacks despite a twentyfold price spread. I present aiAuthZ, an authorization gateway that moves the safety decision off the agent's host. Before a tool call executes, the gateway verifies caller identity with a per-message HMAC-SHA256 signature bound to a single-use nonce and a timestamp window, and it evaluates a role-based and argument-level policy that the agent can neither read nor modify. Every decision joins a SHA-256 hash-chained audit log, and each accepted message yields an HMAC-authenticated QR receipt that achieves 94% mean verification across eight transmission channels, with zero forgeries accepted in 25 wrong-key trials. With the gateway in place, residual attack success falls to 0% for all 15 models at no more than 0.03 ms of added decision latency. On the AgentDojo banking suite, aiAuthZ blocks all seven attacker-directed tool calls the evaluated agents emit, at the cost of one legitimate first-time payment, while a spotlighting baseline allows two injections to succeed. Across nine in-scope case studies from the same incident corpus, aiAuthZ blocks nine of nine, against four of nine for a policy baseline without identity binding. The gateway does not prevent a model from being deceived; it prevents a deceived model from acting beyond the verified user's authority on every call routed through it. The implementation and all experiments are released at https://github.com/Sports-Vision-Inc/aiAuthZ.
ネイティブ不確実性と適応型複雑さ制御を備えたレンダリング対応ベイジアン 3D ガウス スプラッティング
3D ガウス スプラッティング (3DGS) は、リアルタイムのノベルビュー合成の強力な表現ですが、その標準トレーニング パイプラインは点推定と手動調整されたヒューリスティックに依存しており、ネイティブの不確実性や原則に基づいた複雑さの制御は提供されません。これは、モデルが弱くサポートされているジオメトリを識別し、有益なビューを選択する必要がある、まばらなビューまたは固定取得バジェットの下で最も制限されます。レンダリング対応のベイジアン 3DGS フレームワークを導入します。これは、レンダラー由来のサロゲート サマリーを使用して、平均と共分散に対する正規逆ウィシャート事後分布でガウス ジオメトリを追跡します。オプションのディリクレプロセス拡張により、確率的なコンポーネント使用信号が追加され、トレーニング スケジュールにより、閉形式と近似推論の境界が明確になります。事後ジオメトリ サンプルを再レンダリングすると、間隔のキャリブレーションとアクティブなビューの選択に対してネイティブの予測不確実性が得られます。固定予算の 16 ~ 32 のアクティブ ビュー タスクでは、ネイティブ NIW 取得により、スコアリングのみの 3 メンバーの標準アンサンブル ベースラインと比較して、PSNR が +0.453 dB、LPIPS が -0.0146 改善され、29/39 のシーン シード ペアと 10/13 のシーン平均を獲得しました。また、PPU スタイル (+0.355 dB) および NIW プロキシ (+0.401 dB) の取得よりも改善されています。 NIW ネイティブ インターバルは、共有プロキシと比較して 95% カバレッジ エラーを約 17 倍削減し (0.046 対 0.796)、3 メンバーのディープ アンサンブルよりも公称カバレッジに約 10 倍近く (0.047 対 0.454)、トレーニング コストは約 3 分の 1 です。再構築の互換性チェックとして、39 のシーンシード実行にわたる NIW 対標準のペア分析により、1.6% の追加トレーニング時間で +0.030 dB PSNR が得られました。これらの結果は、ベイジアン 3DGS を、アクティブなビューの選択などの意思決定が必要なタスク用の実用的な確率的シーン表現として位置づけています。
原文 (English)
Rendering-Aware Bayesian 3D Gaussian Splatting with Native Uncertainty and Adaptive Complexity Control
3D Gaussian splatting (3DGS) is a strong representation for real-time novel-view synthesis, but its standard training pipeline relies on point estimates and hand-tuned heuristics, providing no native uncertainty or principled complexity control. This is most limiting under sparse views or fixed acquisition budgets, where a model must identify weakly supported geometry and select informative views. We introduce a rendering-aware Bayesian 3DGS framework that tracks Gaussian geometry with a Normal-Inverse-Wishart posterior over means and covariances using renderer-derived surrogate summaries. An optional Dirichlet-process extension adds a probabilistic component-usage signal, and the training schedule makes the closed-form versus approximate inference boundary explicit. Re-rendering posterior geometry samples yields native predictive uncertainty for interval calibration and active view selection. In a fixed-budget 16-to-32 active-view task, native NIW acquisition improves PSNR by +0.453 dB and LPIPS by -0.0146 over a scoring-only 3-member standard-ensemble baseline, winning 29/39 scene-seed pairs and 10/13 scene means; it also improves over PPU-style (+0.355 dB) and NIW-proxy (+0.401 dB) acquisition. NIW native intervals reduce 95% coverage error by about 17x relative to a shared proxy (0.046 vs. 0.796) and are about 10x closer to nominal coverage than a 3-member deep ensemble (0.047 vs. 0.454) at roughly one-third the training cost. As a reconstruction compatibility check, paired NIW-vs-standard analysis over 39 scene-seed runs yields +0.030 dB PSNR with 1.6% additional training time. These results position Bayesian 3DGS as a practical probabilistic scene representation for decision-facing tasks such as active view selection.
クロスエピソード記憶とポリシー蒸留による自己レビュー強化学習 (SRRL)
強化学習は、環境フィードバックを使用して大規模な言語モデルをトレーニングするために一般的に使用されます。応用設定では、環境は通常、まばらなフィードバックまたは遅延したフィードバックを提供します。このため、モデルが推論のどのアクションが成功または失敗につながったのかを正確に特定することが困難になります。したがって、モデルは各失敗が後続の反復で意味のある動作修正をどのように通知するかを決定する必要があるため、これらの信号から効果的に学習することは困難です。各 RL エピソードに明示的な自己レビューのステップを組み込むトレーニング フレームワークである自己レビュー強化学習を導入します。最初のパスの応答が失敗すると、モデルは自己レビューを生成して何が問題だったかを特定し、これにより 2 回目の試行が改善されます。 Reflexion などの推論時のリフレクション アプローチとは異なり、このフレームワークはポリシーの勾配を使用して自己レビューを最適化し、選択的蒸留によって改善を基本ポリシーに内部化し、将来のエピソードにわたって持続することを保証します。エピソード間の記憶は、トレーニング中に将来のエピソードで同様のタスクに遭遇したときに再利用できるように、成功したセルフレビューを保持します。 GSM8K ベンチマークで 2 つの言語モデル、Qwen 3-4B および OLMo-3-7B にわたって GRPO オプティマイザーを使用して、標準 RLVR ベースラインに対して SRRL を評価します。 SRRL は、最終的な報酬パフォーマンスにおいて常に RLVR を上回り、フィードバックを行動改善にうまく変換することで学習効率の向上を実現します。
原文 (English)
Self-Review Reinforcement Learning (SRRL) with Cross-Episode Memory and Policy Distillation
Reinforcement Learning is commonly used to train large language models using environmental feedback. In applied settings, the environment usually provides sparse or delayed feedback. This makes it difficult for the model to pinpoint which actions in its reasoning led to success or failure. So, learning effectively from these signals is hard because the model must determine how each failure should inform meaningful behavioral corrections in subsequent iterations. We introduce a training framework, Self-Review Reinforcement Learning, that embeds an explicit self-review step into each RL episode. When a first-pass response fails, the model generates a self-review to identify what went wrong, which conditions an improved second attempt. Unlike inference-time reflection approaches, such as Reflexion, the framework optimizes self-review with policy gradients and internalizes improvements into the base policy via selective distillation, ensuring they persist across future episodes. A cross-episode memory keeps successful self-reviews for reuse when encountering similar tasks in future episodes during training. We evaluate SRRL against a standard RLVR baseline using the GRPO optimizer across two language models, Qwen 3-4B and OLMo-3- 7B, on GSM8K benchmark. SRRL consistently outperforms the RLVR in final reward performance and achieves greater learning efficiency by successfully transforming feedback into behavioral improvement.
ほとんどの LLM 準拠にはスピーカーは必要ありません: ピアプレッシャーベンチマークにおけるスピーカーフリーフロアの測定
LLM 準拠は、モデルがピアまたはグループの応答に向けて正解を変更するケースを説明するためによく使用されます。この見かけの適合性のほとんどは、ピアが削除された後でも存続することを示します。その理由は混乱にあります。標準準拠のプロンプトでは、発言者の存在と繰り返される間違った回答自体という 2 つの手がかりが同時に混合されます。既存のベンチマークはこれらの手がかりを一緒に変化させるため、リビジョンのどの程度が実際に話者に依存するかを知ることができません。出典のない条件を導入します。つまり、明示的な発言者が削除された同じ主張された回答です。 6 つのオープンウェイト LLM と 7 つの QA および推論データセット全体で、この条件だけで、当初は正しいケースの $66.5\%$ で有害な修正が発生しましたが、単純な再質問の場合は $10.3\%$ でした。この効果は、繰り返された回答が言い換えられた場合や、回答の選択肢が自由形式の設定で隠された場合にも残ります。ソースフレーミングは主にこの下限を調整します。専門家パネルのフレーミングはこの下限を引き上げますが、最小限の人物ラベルでは確実に下限を引き上げることはできません。モデルが反転する場合、通常は確実に間違っており、単純な再調整では元の答えは回復しません。出典の帰属は依然として重要ですが、それは、このスピーカーのないフロアを超える増分として測定される必要があります。方法論的な教訓は、適合性ベンチマークでは、スピーカーを取り外した後に何が残るかをまず測定する必要があるということです。この手順を行わないと、ベンチマークは繰り返されるテキストを社会的影響力があると誤認する可能性があります。
原文 (English)
Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks
LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of this apparent conformity survives even after the peer is removed. The reason is a confound: standard conformity prompts mix two cues at once, the presence of a speaker and the repeated wrong answer itself. Existing benchmarks vary these cues together, so they cannot tell how much of the revision actually depends on the speaker. We introduce a no-source condition: the same asserted answer with the explicit speaker removed. Across six open-weight LLMs and seven QA and reasoning datasets, this condition alone causes harmful revision in $66.5\%$ of initially correct cases, compared with $10.3\%$ under a plain re-ask. The effect also remains when the repeated answer is paraphrased and when answer options are hidden in an open-ended setting. Source framing mainly modulates this floor: expert-panel framing raises it, while minimal person labels do not reliably raise it. When models flip, they are usually confidently wrong, and simple recalibration does not recover the original answer. Source attribution still matters, but it should be measured as an increment above this speaker-free floor. The methodological lesson is that conformity benchmarks should first measure what remains after the speaker is removed; without this step, benchmarks may mistake repeated text for social influence.
大規模な言語モデルの「はい/いいえ」バイアスは、道徳的判断の変化ではなく、回答の順序と言葉遣いを反映しています。
大規模言語モデル (LLM) は、二項評決として読み取られる判断を下すことが増えており、そのような判断が論理的に無関係な文言の変更によって変化することを報告する文献が増えています。その中には、人間には存在しない道徳的ジレンマに対するイエス/ノーのバイアスが増幅されていることが報告されています。単一の枠組みだけではそのような変化が何であるかを言うことはできません。はい/いいえの質問では、「いいえ」という単語は同時に論理的な判断、語彙トークン、そして最後に出力された選択肢になります。これらを分離する心理測定バッテリーを導入します。交差対称化 - 論理的に無関係なすべての要素が、バランスの取れたペアで反転されます - 質問形式のコーパス全体にわたって。論理的に同等の形式にわたる等級付けされた評価は、一貫した内部道徳尺度を回復します。フロンティア モデルのスタンス $\theta$ は形式にほとんど依存しません ($\pm 1$ 軸上の形式間の不整合性は 0.12 ~ 0.21)。小型無重力モデルは、モデル固有の方法で失敗します。はい/いいえによる評決を強制すると、分解可能なアーティファクトがオーバーレイされます。つまり、古典的な人間の優位性とは反対に、最後に印刷されたオプションに対する順序のバイアスに加えて、単語「いいえ」への語彙的な引っ張りです。アーティファクトはクロード モデル (ストーリー平均 -0.32 ~ -0.86) でのみ実質的であり、GPT-5.5 と Gemini では $\約 0$ であり、拡張推論の下では縮小します。言葉と評決は 1 つのトークンを共有します。単語を任意のラベルに置き換えることでそれらが分離され、評決に付随する論理バイアスはすべてのフロンティア モデルで $\ほぼ 0$ であることが証明されますが、モデル固有のラベルと順序の添付ファイルは残ります。モデルは拒否する方向に引き寄せられることはありません。プルは、印刷された表面に従うものであり、それが持つ評決に従うものではありません。最小モデル $P = \sigma((\theta \pm m)/s)$ は、そのようなアーティファクトを、サンプリング温度とは明らかに異なるフレーミング感受性 m と道徳的決断力 s によって要約します。バッテリーは、あらゆるジレンマ セットおよびバイナリ形式に変更なく適用されます。モデルの値を測定するためには、一度質問するのではなく、質問のフレームを横断して測定する必要があります。
原文 (English)
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment
Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word "no" is at once logical verdict, lexical token, and last-printed option. We introduce a psychometric battery that separates these: crossed symmetrization - every logically irrelevant factor flipped in balanced pairs - across a corpus of question forms. A graded rating across logically equivalent forms recovers a coherent internal moral scale: frontier models' stance $\theta$ is nearly format-invariant (cross-form incoherence 0.12-0.21 on a $\pm 1$ axis); small open-weight models fail in model-specific ways. Forcing the verdict through yes/no overlays a decomposable artifact: an order bias toward the last-printed option - opposite to classic human primacy - plus a lexical pull toward the word "no"; the artifact is substantial only in the Claude models (story-averaged -0.32 to -0.86), $\approx 0$ for GPT-5.5 and Gemini, and shrinks under extended reasoning. The word and the verdict share one token; swapping the words for arbitrary labels separates them, and the verdict-attached logical bias proves $\approx 0$ for every frontier model, while model-specific label and order attachments remain: the models are not drawn toward rejecting - the pull follows the printed surface, not the verdict it carries. A minimal model, $P = \sigma((\theta \pm m)/s)$, summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, measurably distinct from sampling temperature. The battery applies unchanged to any dilemma set and binary format: measuring what a model values requires crossing the frames of the question, not asking once.
プロンプトロバストネスはタスク依存: LLM 評価における客観的質問と信念スタイルの質問の比較
大規模な言語モデルの調査形式の評価では、多くの場合、促された応答がモデルの価値観や信念の尺度として扱われます。この仮定は、回答が政治的価値観、社会的態度、または信念の証拠として読み取られる場合に特に脆弱になります。答えが決まっている客観的な質問と、意見や価値観を求める主観的な質問とでは、プロンプトの堅牢性が異なるかどうかを尋ねます。 3 つの客観的データセット (MMLU、ARC、CulturalBench) と 3 つの主観的データセット (Political Compass Test、ValueBench、World Values Survey) に基づいて 4 つの命令調整モデル ファミリを評価します。各質問/ステートメントに対して、文言、枠組み、形式のバリエーションなど、複数のタイプのプロンプト変更を適用し、モデルがバリエーション全体で同じ回答を与えるかどうかを測定します。二項一般化推定方程式を使用すると、モデル、データセット、プロンプト カテゴリ、およびそれらの相互作用の重要な効果がわかります。データセット タイプの影響も大きく、データセット タイプとプロンプト カテゴリの間の相互作用は大きくなります。これらの結果は、プロンプトの堅牢性が質問の種類、プロンプトの変更、モデルに依存することを示しています。
原文 (English)
Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation
Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumption is particularly fragile when responses are read as evidence of political values, social attitudes, or beliefs. We ask whether prompt robustness differs between objective questions with fixed answers and subjective questions that ask for opinions or values. We evaluate four instruction-tuned model families on three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey). For each question/statement, we apply multiple types of prompt changes, such as variations in wording, framing, and format, and measure whether the model gives the same answer across variants. Using a binomial generalized estimating equation, we find significant effects of model, dataset, prompt category, and their interactions. The dataset type effect is also significant, and the interaction between dataset type and prompt category is large. These results show that prompt robustness depends on the question type, the prompt change, and the model.
トレーニング不要のプリミティブ形状抽象化のための生成画像モデルの利用
3D 形状を幾何学的プリミティブのコンパクトなセットとして表現することは、ロボット工学、シミュレーション、およびシーンの理解の基礎です。大規模にトレーニングされた生成画像モデルは、タスク固有のトレーニングを必要とせずに、任意のカテゴリにわたって、画像ドメイン内でオブジェクトの部分を直接識別してセグメント化できる汎用的な視覚学習器として最近登場しました。このようなモデルを下流のタスクに適応させるには、通常、微調整が必要です。事前にトレーニングされた能力をトレーニングなしで直接活用できるかどうかを尋ねると、トレーニングなしで活用できると肯定的に答えます。私たちのパイプラインは、3D オブジェクトのマルチビュー イメージをレンダリングし、ビジョン言語モデルを使用してそのセマンティック パーツを分析し、生成イメージ モデルに色分けされたパーツ セグメンテーション マスクをペイントするように指示し、それをジオメトリに再投影し、パラメーターの最適化によって各パーツに超二次プリミティブを適合させます。このアプローチには学習されたパラメーターが含まれていません。これは、カテゴリに依存せず、向きに依存せず、以前の学習ベースのモデルが苦労していた特性です。その精度の上限は、将来の生成モデルの改善に伴って上昇します。これは、プリミティブ フィッティングではなくパーツ セグメンテーションが現在の精度のボトルネックであることを示すグランドトゥルース セグメンテーション研究で確認されています。 HumanPrim と Toys4K では、オブジェクトごとに平均 5 ~ 9 個のプリミティブを使用して、評価したすべての方法の中で最も低い面取り距離を達成します。
原文 (English)
Harnessing Generative Image Models for Training-Free Primitive Shape Abstraction
Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding. Generative image models trained at scale have recently emerged as generalist visual learners that can identify and segment object parts directly in the image domain, across arbitrary categories and without task-specific training. Adapting such models to downstream tasks typically requires fine-tuning; we ask whether their pretrained capability can instead be harnessed directly, without any training, and answer affirmatively with a training-free harness. Our pipeline renders multi-view images of a 3D object, uses a vision-language model to analyze its semantic parts, prompts a generative image model to paint a color-coded part segmentation mask, reprojects it onto the geometry, and fits a superquadric primitive to each part via parameter optimization. The approach contains no learned parameters: it is category-agnostic and orientation-invariant, properties that previous learning-based models struggled with. Its accuracy ceiling rises with future generative-model improvements, which we confirm with a ground-truth segmentation study showing that part segmentation, not primitive fitting, is the current accuracy bottleneck. On HumanPrim and Toys4K, our method achieves the lowest Chamfer distance among all evaluated methods, using 5--9 primitives per object on average.
誰の公平さ? AI バイアス研究における構造的集中
人工知能は医療、法律、公共サービスにおける重大な決定を仲介することがますます増えており、この分野では偏見を測定し軽減するための広範な方法論で対応しています。しかし、この方法論の基礎となる公平性の定義、ベンチマーク、バイアス軽減フレームワークは、その構成が特徴付けられたことのない研究コミュニティによって作成されているにもかかわらず、普遍的なものとして扱われています。私たちは、AI バイアス研究が構造的に集中していること、そしてこの集中が地理的に、まさにこの分野の残りの領域が継承している領域で最大であることを示します。 5 つのテーマ別ドメインにまたがる 692 件の出版物を分析し、書誌学的分析と意味論的クラスタリングを組み合わせたところ、研究活動は少数の国、機関、著者によって支配されており、米国はあらゆる分野にわたる出版物生産と共同ネットワークをリードしており、一般的な公平性と偏見の緩和において最も強力であり、4 つの意味論的クラスタすべてにわたって意味のある表現をもつ最大で最も多く引用されている分野であることがわかりました。低所得国と中所得国は依然としてコミュニティとその協力ネットワークにほとんど参加しておらず、引用の影響力は大きく偏っており(中央値 = 9; 平均 = 93.5 )、ごく一部の出版物がこの分野を不均衡に形成していることを示しています。一般公平性の領域は、アプリケーション分野に適用される定義とベンチマークを提供するため、この基礎的な領域に研究努力が集中すると、AI バイアス研究全体に波及し、狭い範囲で開発および検証された緩和手法が、AI が導入されるすべての集団や環境に一般化できるわけではないのではないかという懸念が生じます。私たちは、フィールドの構造を継続的に監視するためのインタラクティブなアトラスを提供します。
原文 (English)
Whose fairness? Structural concentration in AI bias research
Artificial intelligence increasingly mediates consequential decisions in healthcare, law, and public services, and the field has responded with an extensive methodology for measuring and mitigating bias. Yet the fairness definitions, benchmarks, and debiasing frameworks on which this methodology rests are treated as universal while being produced by a research community whose composition has never been characterized. We show that the AI bias research are structurally concentrated, and that this concentration is greatest, geographically, in precisely the domain the rest of the field inherits from. Analyzing 692 publications spanning five thematic domains, combining bibliometric analysis with semantic clustering, we find that research activity is dominated by a small set of countries, institutions, and authors, with the United States leading publication output and collaboration networks across every domain and most strongly in general fairness and bias mitigation, the largest, most-cited domain with meaningful representation across all four semantic clusters. Low- and middle-income countries remain largely absent from the community and its collaboration networks, and citation influence is highly skewed (median = 9; mean =93.5 ), indicating that a small fraction of publications disproportionately shapes the field. Because the general-fairness domain supplies the definitions and benchmarks that application areas apply, concentration of research effort in this foundational domain propagates across AI bias research as a whole - raising the concern that mitigation methods developed and validated within a narrow set of contexts may not generalize to all populations and settings where AI is deployed. We provide an interactive atlas for continuous monitoring of the field's structure.
ResonatorLM: 効率的なロングコンテキスト言語モデルのための因果共鳴場ミキシング
現代の言語モデルはトランスフォーマー アーキテクチャによって支配されており、セルフ アテンション メカニズムを利用して、幅広いドキュメントとコーパスのセットにわたってより効率的で並列化されたトレーニングを可能にします。これにより、トランスフォーマーは幅広いモダリティやコンテキストにわたってデータを効果的にモデル化できるようになりました。ただし、トランスフォーマーは、リカレント ニューラル ネットワーク (RNN) や畳み込みニューラル ネットワーク (CNN) などの従来の対応物と同様に、長いコンテキストを処理するときに効率を維持するのに苦労することがよくあります。注意を物理学由来の代替手段に置き換える新しいメカニズムである ResonatorLM を紹介します。 ResonatorLM は、トークン シーケンスを単一の駆動される 1 次元潜在フィールドとして扱い、アテンション ドット積を減衰共振器の因果関数に置き換えます。従来のネットワーク アーキテクチャに ResonatorLM を実装し、標準的なロングコンテキスト モデリング タスクでテストします。小規模な 6M のマッチング設定では、シーケンスの長さに応じてトレーニングとプリフィルの速度が向上し、32K トークンでの標準の最適化されたトランスフォーマーと比較してデコード速度が 6.47 倍に達し、WikiText での精度が 61.31 パーセント (55.32 パーセントと比較) に達することがわかりました。
原文 (English)
ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modelin
Contemporary language models are dominated by the transformer architecture, which leverages self-attention mechanisms to enable more efficient, parallelized training across a wide set of documents and corpora. This has allowed transformers to effectively model data across a wide range of modalities and contexts. However, transformers, along with their conventional counterparts such as recurrent neural networks (RNNs) and convolutional neural networks (CNNs), often struggle to maintain efficiency when processing long contexts. We introduce ResonatorLM, a new mechanism that replaces attention with a physics-derived alternative. ResonatorLM treats token sequences as a single, driven one-dimensional latent field and replaces attention dot products with causal functions of damped resonators. We implement ResonatorLM on a traditional network architecture and test it on standard long-context modeling tasks. We find that in a small, 6M matched setting, training and prefill speedups increase with sequence length, decode speed reaches 6.47x compared to that of a standard, optimized transformer at 32K tokens, and accuracy reaches 61.31 percent (compared to 55.32 percent) on WikiText.
カスケード特徴除去による階層分類: ヒト表現型オントロジーに合わせた顔表現型解析 (FaceMesh2HPO) への応用
FaceMesh2HPO は、臨床診断をサポートするために、ヒト表現型オントロジー (HPO) に合わせて顔の表現型記述子を分類するためのフレームワークです。 10 の疾患にわたる 124 人の臨床医からのアノテーション (107 HPO 用語) と非症候性コントロールを組み合わせて、2D 画像から 3D 顔メッシュ (478 個のランドマーク) を生成し、カスケード分類と特徴除去を使用して階層的な PointNet ベースのパイプラインをトレーニングしました。 3D メッシュ、顔の輪郭、人口統計メタデータを組み込んだ最良のモデルは、~0.55 ~ ~0.89 の AUROC を達成し、親ノードでのパフォーマンスがリーフ項よりも高かった。外部検証により、障害全体にわたってさまざまな一般化可能性が示されました。結果は、3D 顔形状の階層モデリングにより、解釈可能なオントロジーにリンクされた表現型分類が可能になることを示していますが、まれな葉の用語でのパフォーマンスは依然として限定的です。堅牢性と臨床的有用性を高めるには、データの多様性と特徴選択戦略の改善が必要です。
原文 (English)
Hierarchical Classification via Cascading Feature Elimination: Application to Human Phenotype Ontology-Aligned Facial Phenotyping (FaceMesh2HPO)
FaceMesh2HPO is a framework for classifying facial phenotypic descriptors aligned with the Human Phenotype Ontology (HPO) to support clinical diagnosis. Using annotations from 124 clinicians across 10 disorders (107 HPO terms) combined with non-syndromic controls, we generated 3D facial meshes (478 landmarks) from 2D images and trained a hierarchical PointNet-based pipeline with cascading classification and feature elimination. The best models, incorporating 3D meshes, facial outline, and demographic metadata, achieved AUROCs between ~0.55 and ~0.89, with higher performance at parent nodes than leaf terms. External validation showed variable generalizability across disorders. Results demonstrate that hierarchical modeling of 3D facial geometry enables interpretable, ontology-linked phenotype classification, though performance on rare leaf terms remains limited. Improved data diversity and feature selection strategies are needed to enhance robustness and clinical utility.
維持するか適応するか?継続学習の一般化
継続学習 (CL) の文献は、壊滅的な物忘れを軽減するという目標によって長い間推進されてきました。この目標は、生涯学習者は共同タスク学習 (JTL) ソリューションに近似し、以前に取得した知識をすべて保持する必要があるという、広く浸透している、多くの場合明言されていない前提に基づいています。私たちは、この保持中心の前提に異議を唱え、非定常環境では保持を優先するとリアルタイムの適応が妨げられる可能性があると主張します。平均生涯誤差 (ALE) に焦点を移し、環境と学習のダイナミクスの間の相互作用によって支配されるオンライン最適化問題として CL を形式化します。競合する過去の経験から引き継がれたバイアスである不安定性と、新しいタスクをゼロから学習する最適化コストである一時的エラーの間の緊張の定量的尺度として、転送効率を導入します。緩やかな収束条件下では、線形ネットワーク モデルとニューラル ネットワーク モデルにまたがるこの分解により、クリティカル タスク期間が得られます。これは、リテンションが正の定常バイアスを誘発するたびに、履歴知識がウォーム スタートの利点から最適化の欠点に移行する閉形式のしきい値です。これらの理論的予測を継続的な画像分類と強化学習ベンチマークで検証します。最後に、継続学習を予測可能なシーケンスのオンライン学習フレームワークに接続することで、JTL がより広範な目標群の 1 つのインスタンスにすぎないことを示し、予測継続学習と呼ばれる、継続学習アルゴリズムの新しい一般クラスを提案します。予測 CL アルゴリズムは、将来のタスクの明示的で動的に更新されるモデルの下で、予想される将来のパフォーマンスを最適化します。概念実証として、JTL と独立タスク学習 (ITL) の間を補間し、制御された分布ドリフトの下で両方を上回るパフォーマンスを発揮するウィンドウ アルゴリズムを分析します。
原文 (English)
To Retain or to Adapt? Generalizing Continual Learning
The Continual Learning (CL) literature has long been driven by the goal of mitigating catastrophic forgetting. This objective rests on a pervasive, often unstated assumption: that a lifelong learner should approximate the Joint-Task Learning (JTL) solution and retain all previously acquired knowledge. We challenge this retention-centered premise, arguing that in non-stationary environments prioritizing retention can impede real-time adaptation. Shifting the focus to the Average Lifelong Error (ALE), we formalize CL as an online optimization problem governed by the interaction between environmental and learning dynamics. We introduce Transfer Efficiency as a quantitative measure of the tension between Instability, the bias inherited from conflicting past experience, and Transient Error, the optimization cost of learning new tasks from scratch. Under mild convergence conditions, holding across linear and neural network models, this decomposition yields a Critical Task Duration: a closed-form threshold beyond which historical knowledge transitions from a warm-start advantage to an optimization liability whenever retention induces a positive stationary bias. We validate these theoretical predictions on continual image classification and reinforcement learning benchmarks. Finally, by connecting continual learning to the online learning framework of predictable sequences, we show that JTL is only one instance of a broader family of objectives, and we propose a new general class of continual learning algorithms, which we call Predictive Continual Learning. Predictive CL algorithms optimize expected future performance under an explicit, dynamically updated model of future tasks. As a proof of concept, we analyze a Window algorithm that interpolates between JTL and Independent-Task Learning (ITL), outperforming both under controlled distributional drift.
BaFCo: 複雑なバングラ語フォームを理解するための文書理解ベンチマーク
マルチモーダル大規模言語モデルにとって文書の理解は、特にこれらのシステムが現実世界の人間中心のアプリケーションでの採用が増加しているため、困難ではありますが影響力のあるタスクです。ただし、高品質の注釈付きデータが不足しているため、この採用はバングラ語などの低リソース言語に限定されます。このギャップに対処するために、ドキュメント レイアウト分析 (DLA) と重要情報抽出 (KIE) に焦点を当てたバングラ語フォーム理解のベンチマーク データセットである BaFCo を紹介します。 BaFCo は、農業、教育、銀行、土地管理など、さまざまな分野から集めた 200 の複数ページにわたる複雑なバングラデシュ政府フォームを厳選しています。これらのフォームの構造的および文脈上の複雑さを正確に把握するために、26 種類のフォーム エンティティと、5 種類からなる別の粗いフォーム エンティティ セットで構成されるきめの細かい注釈スキーマを定義します。 ChatGPT、Gemini、Claude、Qwen、および Kim シリーズの最新の MLLM を、低推論設定と高推論設定の両方でゼロショットおよび思考連鎖プロンプトを使用して評価します。私たちの結果は、バングラ語形式を理解する現在の MLLM の能力、特に非常に粒度の細かい形式エンティティを正確に位置特定する能力に限界があることを明らかにしました。データセットとコードは、https://huggingface.co/datasets/Mausul/bafco から入手できます。
原文 (English)
BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications. However, this adoption is limited for low-resource languages such as Bangla due to the scarcity of high-quality annotated data. To address this gap, we introduce BaFCo, a benchmark dataset for Bangla form comprehension with a focus on Document Layout Analysis (DLA) and Key Information Extraction (KIE). BaFCo curates 200 multi-page complex Bangladeshi government forms, sourced from across diverse sectors including agriculture, education, banking, and land management. To accurately capture the structural and contextual complexity of these forms, we define a fine-grained annotation schema comprising 26 types of form entities, along with a separate coarse form entity set consisting of 5 types. We evaluate the latest MLLMs from the ChatGPT, Gemini, Claude, Qwen, and Kimi series using zero-shot and chain-of-thought prompts under both low and high reasoning setups. Our results reveal limitations in current MLLMs' ability in comprehending Bangla forms, particularly in accurately localizing highly granular form entities. Our dataset and code is available at: https://huggingface.co/datasets/Mausul/bafco
Safe Bayesian Optimization with Counterfactual Policies
In many decision-making settings, new interventions are acceptable only if they do not reduce outcomes below some established threshold. Fo…
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems
Teams deploying large language models in business contexts need evaluation systems, yet most treat evaluation as static model selection: ru…
Do It Right! A Methodology for Successful NLP System Development
Natural language processing (NLP) is a common method for supplying data to clinical research and decision making by extracting information…
Physics-Regularized Machine Learning for Proprioceptive Vehicle Localization Using Onboard Sensors
Accurate and robust localization is essential for autonomous mobility systems in real-world environments. While fusing Inertial Measurement…
What Do AI Agents Actually Change? An Empirical Taxonomy of Mutation Patterns in Performance-Improving Pull Requests
AI coding agents are black boxes: we cannot inspect how they generate code, but we can inspect what they change. This distinction matters f…
RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs
Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate a…
Beyond Accuracy: How Humans Evaluate Legally Correct but Socially Controversial Legal Advice from Machines
AI systems are increasingly used to provide legal advice, raising questions about whether laypeople accept guidance from algorithms--especi…
Depression Symptoms and Relational Patterns in 187k ChatGPT Histories
Large language models are increasingly used as private, always-available conversational systems, but little is known about how people with…
IMR: Iterative Mode-World Weighted Regression for Multi-Agent Trajectory Prediction
Multi-agent motion prediction is essential for automated vehicles to understand the intentions of surrounding vehicles. However, previous p…
Plainbook: Data Science, in Plain Language
Jupyter Notebooks have become widely adopted in data science, as they allow the sharing of reproducible computational analysis. They are, h…
SCOReD: Student-Aware CoT Optimization for Recommendation Distillation
Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-su…
The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities
AI coding agents now read repositories, call tools, and execute shell commands with limited human oversight, and a fast-growing body of wor…
Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations
The Model Context Protocol (MCP) is the dominant way coding agents discover and invoke external tools. A server advertises each tool throug…
When Should LLMs Search? Counterfactual Supervision for Search Routing
Search-augmented language models can use external evidence to compensate for limitations in parametric knowledge, but search is not uniform…
Data-dependent Evaluations for Budgeted Submodular Maximization
Submodular maximization is an important building block for developing algorithms in many areas such as machine learning and data mining. Du…
LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding
We propose a novel pipeline, Legato 2, for extracting symbolic notation and semantic knowledge from images of sheet music. Legato 2 feature…
FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning
While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same funct…
Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning
Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasonin…
Complementary Roles of Image Classification and Vessel Segmentation in AI-Based Screening for Retinopathy of Prematurity Plus Disease in a Kenyan Preterm Cohort
Background. Retinopathy of prematurity (ROP) is a preventable cause of childhood blindness, with rising burden in low- and middle-income co…
Decision-Focused Scenario Generation and Selection for Efficient and Robust Grid Dispatch
The increasing uncertainty from flexible demand and renewable generation has made distributionally robust optimization (DRO) an important t…
Tangent classes of matroids and wonderful compactifications
For every loopless matroid $M$ and every Feichtner--Yuzvinsky building set $\mathcal{G}$ containing the top flat, we construct an integral…
VisTCP: A Visualization Framework to Construct Knowledge-Graph-Based Representation for Traditional Chinese Painting
Structured representation can characterize semantic objects and relationships in images. It provides a possible effective way for the seman…
Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis
Large language model (LLM)-assisted software security operates at a difficult boundary: the vulnerability-analysis terminology needed for l…
AbICL: In-Context Learning for Antigen-Specific Antibody Affinity Ranking
Accurate ranking of antibody candidates according to their binding affinity is essential for therapeutic antibody discovery. However, exist…
Unsupervised Anomaly Detection of Information Operations Users via Behavioral and Language Patterns
Information Operations on social media networks have been identified as a significant threat to democracy and modern society, but they are…
Differentially Private Natural Gradient Descent
Under a fixed privacy budget, the utility of differentially private (DP) training is ultimately determined by its optimization efficiency.…
Think Before You Grid-Search: Floor-First Triage for LLM Serving
LLM serving optimization typically benchmarks many configurations and reaches for heavy profilers when latency targets are missed. We argue…
Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context
Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and re…
i-EXAM: Instructable and Explainable Attack Connectivity Graph Modeler
i-EXAM is a planning-powered tool that helps system administrators to create security profiles of complex networks and perform what-if anal…
Few-Medoids: An Embarrassingly Simple Coreset Selection Method for Few-Shot Knowledge Distillation
Coreset selection aims to identify a small and highly representative subset of a massive dataset for efficient model training. The problem…
From Textural Counterpoint to Feature Encoding: A Multi-Dimensional Machine Representation Study of Haydn's "The Lark" Integrating Electroacoustic Analysis
Chamber music, as a highly precise multi-part interactive system, contains a logic of "role assignment and dynamic interaction" that provid…
K-ABENA: K-Adaptive Backpropagation with Error-based N-exclusion Algorithm : (Compensated Loss-Based Sample Exclusion with Unbiased Gradient Estimation)
We present K-ABENA (K-Adaptive Backpropagation with Error-based N-exclusion Algorithm), a selective gradient computation framework that red…
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an…
CMDR: Contextual Multimodal Document Retrieval
Multimodal document retrieval aims to retrieve relevant pages while preserving both textual and visual content from the original document.…
Signed-Graph Recommendation as Structural Consistency Maximization
While signed social recommendation has shown great potential by modeling both trust and distrust relations, its effectiveness is often hind…
NegROI: Click-Centric Uncertainty-Guided Refinement with Scene-Conditioned Negative Prompts for Robust Interactive 3D Segmentation
Interactive 3D segmentation aims to extract object masks in point clouds with minimal user clicks. Despite recent progress, most existing a…
Agentic AI for IPoDWDM Network Lifecycle Automation: An MCP-Enabled Architecture
We present a distributed, vendor-agnostic multi-MCP architecture for SDN-based automation and autonomous control of multi-vendor, multi-lay…
Decoupled Single-Mask Annotation Noise Detection via Cross-Sectional Patch Self-Consistency
Vascular computed tomography datasets are commonly annotated only once per scan, yielding the pervasive yet under addressed problem of sing…
InfluMatch: Frontier-Quality KOL Search at 4B-Model Cost
Matching influencers (KOLs) to free-form, multi-part Thai marketing criteria is today served either by keyword search over structured profi…
Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems. W…
MCP-Enabled Agentic AI for Autonomous IPoDWDM Network Lifecycle Automation
This demo presents an MCP-enabled agentic AI architecture for autonomous control of vendor-agnostic IPoDWDM networks. We demonstrate live e…
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention
Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events…
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks…
Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are il…
From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations
Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical m…
PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
3D dense captioning, an emerging vision-language task, aims to generate descriptive sentences for each object in the 3D scene. Despite the…
Agents That Teach: Towards Designing Incidental Learning Back into AI-Assisted Software Development
AI coding agents are rapidly reshaping how software is built, with developers increasingly delegating substantial coding tasks to autonomou…
EcoVision: AI-Powered Drone Imaging for Salt Marsh Vegetation Monitoring and Dominance Mapping
High-resolution RGB imagery acquired from low-altitude UAV surveys was processed through a modular pipeline incorporating transformer-based…
RoME: Robust Mixture of Low-Rank Experts against Multiple Adversarial Perturbations
Multi-perturbation adversarial training (MAT) aims to achieve robustness against multiple $\ell_p$ perturbations but suffers from robustnes…
LLM-Guided Measurement Credibility Correction for Trustworthy Industrial Process Inference
Industrial prediction and soft sensing depend on credible input measurements. In field deployment, a predictor may receive biased, delayed,…
x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability
Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function eva…
Static Metrics Are Insufficient: Predicting Java Method Energy Usage with Execution Time
The increasing energy demand of software systems is raising concerns about their environmental impact and associated costs. Reasoning on en…
Evaluating Fine-Tuning and Metrics for Neural Decompilation of Dart AOT Binaries
Neural decompilation is increasingly studied as a code-generation problem, yet its evaluation methodology remains underdeveloped for modern…
Self-Supervised Implicit CEST Reconstruction via Physics-Informed Lorentz Encoding
Multi-Pool Chemical Exchange Saturation Transfer (CEST) MRI provides valuable metabolic information but is clinically limited by long acqui…
Property-Driven Synthetic Data Engineering for Data-Scarce Software Systems: Reflections from the Breast Cancer Domain
Modern software systems increasingly depend on data for analysis, prediction, testing, and decision-making. Yet many important domains, inc…
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information…
LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis
Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language m…
X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models
Foundation Models for Electronic Health Records (FEMRs) are pretrained on large-scale structured patient data, enabling them to convert lon…
Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design
Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits t…
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for…
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatr…
Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows
Major cloud data platforms now expose large language model capabilities as native SQL functions, enabling analysts to perform classificatio…
VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection
Deepfake image detection is currently served by three fundamentally different paradigms: commercial APIs, zero-shot vision-language models…
Designing Maintainable Hybrid Generative Systems: A Quantum-Inspired Approach to Automated Music Harmony Generation
This paper presents the design and evaluation of a maintainable hybrid generative architecture for automated music harmony generation from…
UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on…
Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) a…
Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on En…
Harnessing Code Agents for Automatic Software Verification
Formal verification offers the strongest guarantee of software correctness, but it does not scale: the proofs demanded by interactive theor…
Responsible Personalisation: The Double-Edged Sword of Personalisation in Human-Robot Interaction
While personalisation is becoming a defining capability in human-robot interaction (HRI), the existing literature on responsible personalis…
What Images Cannot Say: Language-Guided Olfactory Representation Learning
Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with elect…
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications
Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in th…
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
Large language model coding agents increasingly perform open-ended data modeling and analysis. These agents are stochastic and adaptive, an…
TILDE: TILt-based Distributional Erasure for Concept Unlearning
Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright…
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders
Vision-Language Models (VLMs) are increasingly utilized as the conditioning backbone for diffusion-based image editing due to their remarka…
From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b
Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration…
Provable learning separation for predicting time-evolution of quantum many-body systems
Given that quantum computers are naturally suited to simulate the behavior of quantum many-body systems, an immediate question arises: can…
Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation
We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene str…
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically f…
AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models
Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their ad…
Pitwall: Faithful Natural-Language Race-Strategy Briefings from a Calibrated Real-Time Monte Carlo Engine
Live sports commentary is grounded generation under a deadline: statements concern real, named athletes, the grounding state changes every…
Industry Classification of GitHub Repositories Using the North American Industry Classification System (NAICS)
GitHub hosts hundreds of millions of public repositories, but the platform exposes no native mapping from repositories to standardized indu…
RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation
Multi-hop Question Answering over Knowledge Graphs faces a critical challenge: traditional retrieve-then-read pipelines break differentiabi…
Graph Convolutional Attention: A Spectral Perspective on Graph Denoising and Diffusion
Denoising graphs is a fundamental problem in graph learning and the core operation of graph diffusion models. Attention-based architectures…
ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D int…
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning
Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In l…
Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?
We investigate whether socioeconomic indicators, like household wealth, leave recoverable informational imprints in both satellite imagery…
Base Models Know How to Reason, Thinking Models Learn When
What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers…
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, they exercise agency to anticipate user nee…
When Assisting One Disempowers Another
Personal AI agents are increasingly deployed in shared environments, where their actions affect not just the primary user they are assistin…
VASP Agent: An Agentic Framework for Autonomous First-principles Calculations
Large Language Models (LLMs) are increasingly embedded in agentic frameworks for scientific discovery. First-principles materials computati…
Implementing Metric Temporal Answer Set Programming
We develop a computational approach to Metric Answer Set Programming (ASP) to allow for expressing quantitative temporal constraints, like…
Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique
Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and de…
Self-Routing: Parameter-Free Expert Routing from Hidden States
Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a lea…
理屈ではなく、言われたことを実行する: LLM エージェントの誠実さのギャップを特定する
LLM エージェントは、自分が述べた推論に基づいて行動しますか?このプロセス忠実度の問題は、ソーシャル シミュレーションで LLM を使用する際の中心となりますが、正しい動作の基準が存在しない場合は測定することが困難です。私たちは、忠実性のギャップを推論 - 結論と結論 - 行動の 2 つのステップに分解することにより、すべての決定に対して検証可能な参照アクションを備えたテキサス ポーカー シミュレーターという、制御された設定でそれを研究します。 2 つのステップは逆に動作します。
原文 (English)
Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents
Do LLM agents act on the reasoning they state? This question of process fidelity is central to LLM-based social simulation, yet hard to measure where no reference for correct behavior exists. We study it in a controlled setting: a Texas Poker simulator with a verifiable reference action for every decision by splitting the faithfulness gap into two steps: reasoning-to-conclusion (does the stated decision follow from the agent's own reasoning?) and conclusion-to-action (does the agent execute what it states?). The two steps behave very differently. Conclusion-to-action is reliable: inconsistency is 0.7% for Claude Haiku 4.5 and 1.4% for DeepSeek-Reasoner once the conclusion is read from an explicit tag, whereas free-text conclusion extraction reports 22-26%. Reasoning-to-conclusion is where fidelity frays, but not through a single dominant failure. In a step-level diagnostic the agent's errors split roughly evenly between bad inputs, borderline cases, and rule misapplication deriving a conclusion that contradicts the agent's own restated rule from inputs it estimated correctly. This composition is model-dependent: rule misapplication accounts for a third of Haiku's interpretable errors but only 8% of DeepSeek's. The one robust signal is directional: when an agent does misapply its own stated rule, it almost always (99.5% for Haiku) errs in the risk-averse direction. The override is partly hedging behavior, not a capability limit: instructing the agent to apply the rule mechanically halves the misapplication rate (13.9% to 6.8% of decisions) and raises adherence by eight points. Process-fidelity evaluation should therefore elicit machine-checkable conclusions and probe for directional biases rather than assume a single upstream failure mode, lest it conflate measurement noise with model behavior.
EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors
Practical non-invasive Brain-Computer Interface (BCI) systems require EEG decoders with strong cross-subject generalization and minimal cal…
能動推論とはどのようなタイプの推論ですか?
能動推論では、期待自由エネルギー (EFE) が目標指向の行動と情報探索の行動を統合し、意思決定を推論としてキャストします。最近の研究では、EFE 最小化が、認識的事前分布で強化された生成モデル上の変分自由エネルギー (VFE) 最小化として記述できることが示されました。拡張モデルの VFE は、予測モデルの VFE に明示的なエントロピー補正項を加えたものとして書き換えることができ、EFE の寄与が透明になることを証明します。次に、適切な EFE ベースの計画には、これらの認識論的修正と限界推論を政策最適化に変える計画修正を組み合わせる必要があり、EFE ベースの計画の完全な変分特性が得られることを示します。これにより、クロスエントロピー計画および完全な EFE ベースの計画にどの修正が必要かが明確になります。同じエントロピー補正された定式化により、より単純なアブレーションとともに、EFE ベースの計画のための詳細なメッセージ パッシング スキームが得られます。 3 つのグリッドワールド環境での実験では、観察が決定的な場合には計画修正がすでに役に立ちますが、観察が単に示唆的な場合には追加の観察側の認識論的修正が最も重要であることが示されています。
原文 (English)
What Type of Inference is Active Inference?
Active inference casts decision-making as inference, with the Expected Free Energy (EFE) unifying goal-directed and information-seeking behavior. Recent work showed that EFE minimization can be written as Variational Free Energy (VFE) minimization on a generative model augmented with epistemic priors. We prove that the VFE of the augmented model can be rewritten as the VFE of the predictive model plus explicit entropy-correction terms, making the EFE contribution transparent. We then show that proper EFE-based planning requires combining these epistemic corrections with a planning correction that turns marginal inference into policy optimization, yielding a full variational characterization of EFE-based planning. This clarifies which corrections are needed for cross-entropy planning and for full EFE-based planning. The same entropy-corrected formulation leads to a detailed message-passing scheme for EFE-based planning together with simpler ablations. Experiments on three grid-world environments show that full EFE-based planning outperforms ablations that omit either the planning correction or the epistemic corrections.
ウェアラブルデバイス上のEEG解析のための深層学習モデルの複雑さを軽減する
ウェアラブル ヘルスケア デバイスは、モノのインターネット (IoT) 分野で最も急速に成長しています。多くの自動ヘルスケア サービスは、2 つの重要な生物学的信号、つまり ECG と EEG に依存しており、それぞれ心臓と脳の活動を反映しています。ディープ ニューラル ネットワークは、これらの信号を処理および分析するための主な方法と考えられていますが、ウェアラブル デバイスのエネルギーと計算能力の非常に厳しい制約は、DNN モデルの計算、エネルギー、およびメモリ帯域幅の要求をはるかに下回っており、そのため、多くの実際のウェアラブル サービスでのディープ ラーニングの導入が妨げられています。この論文では、リソースに制約のあるウェアラブル デバイスに最先端の DNN モデルを展開する実現可能性を調査します。特に、パラメーターの量子化と電極削減法が使用される場合の DNN の精度と計算の複雑さの間のトレードオフを調査します。私たちの調査は、EEG 信号分析、特にてんかん発作の検出用に設計されたいくつかの最先端の DNN モデルに重点を置いています。私たちの調査結果は、これらの技術を慎重に適用すると、精度への悪影響を最小限に抑えながら、検討中の DNN の複雑さを大幅に軽減できることを示しています。これらの結果は、DNN ベースのオンライン EEG 分析をウェアラブル デバイスに適応させるときに遭遇する、精度と複雑さの軽減との間の明確なトレードオフを明らかにしています。
原文 (English)
Reducing the Complexity of Deep Learning Models for EEG Analysis on Wearable Devices
Wearable healthcare devices are the fastest-growing Internet of Things (IoT) sector. Many automated healthcare services rely on two crucial biological signals, namely ECG and EEG, which reflect the activity of the heart and brain, respectively. Although deep neural networks are considered the primary way to process and analyze these signals, the very tight energy and computational power constraints in wearable devices are far below the computational, energy, and memory bandwidth demands of DNN models, thereby impeding the deployment of deep learning in many practical wearable services. This paper investigates the feasibility of deploying state-of-the-art DNN models in resource-constrained wearable devices. Notably, we explore the trade-off between accuracy and computational complexity of DNNs when parameter quantization and electrode reduction methods are used. Our investigation centers on several state-of-the-art DNN models designed for EEG signal analysis, specifically for detecting epileptic seizures. Our findings demonstrate that, when applied judiciously, these techniques can significantly reduce the complexity of the DNNs under consideration with minimal adverse effects on accuracy. These results reveal the explicit trade-offs between accuracy and complexity reduction encountered when adapting DNN-based online EEG analysis for wearable devices.
科学的発見における AI の 3 層フレームワーク
科学的発見における AI に関する現在の議論は、多くの場合、既存の知識の検索と、最適化、シミュレーション、自動化による実行という 2 つの目に見える機能によって占められています。どちらも重要ですが、どちらも発見の中心となる行為、つまりモデルの形成と進化を完全には捉えていません。この論文では、発見における AI の 3 層のビューを提案します。レイヤ 1 は、大規模な言語モデルによる検索と取得です。この論文の主な革新であるレイヤー 2 は、定性的推論によるモデルの形成です。つまり、現在のフレームワークが構造的に不適切であることを認識し、試行錯誤ではなく、何が欠けていてどこで見つかるのかについての構造的な洞察を通じて、より広い表現空間内で問題を理解する能力です。レイヤ 3 は実行、最適化、改良です。主な主張は、レイヤー 2 が最も重要であると同時に最も開発されていないということです。モデル形成のない探索は継承されたフレームワークに限定されたままですが、概念の修正なしで実行すると既存の定式化が増幅されるだけです。我々は、S. S. チャーンによるガウス・ボネット定理の本質的証明、リアプノフ関数によるネステロフ加速勾配収束問題の解決、および 2026 年の OpenAI によるエルドス単位距離予想の自律的反証という 3 つのケーススタディを通じて、レイヤー 2 推論を説明します。各ケースは、同じ構造的特徴を示しています。つまり、不適切になったフレームワーク、欠落している概念オブジェクト、およびで見つかった解決策です。思わぬ隣の畑。
原文 (English)
A Three-Layer Framework for AI in Scientific Discovery
Current discussions of AI in scientific discovery are often dominated by two visible capabilities: search over existing knowledge and execution through optimization, simulation, and automation. Both are important, but neither fully captures the central act of discovery: the formation and evolution of models. This paper proposes a three-layer view of AI in discovery. Layer 1 is search and retrieval by large language models. Layer 2, as the main innovation of this paper, is model formation through qualitative reasoning: the capacity to recognize when a current framework is structurally inadequate and to understand the problem within a broader representational space, not through trial and error, but through structural insight into what is missing and where it can be found. Layer 3 is execution, optimization, and refinement. The main claim is that Layer 2 is both the most important and the least developed. Search without model formation remains confined to inherited frameworks, while execution without conceptual revision only amplifies an existing formulation. We illustrate Layer 2 reasoning through three case studies: S. S. Chern's intrinsic proof of the Gauss-Bonnet theorem, the resolution of the Nesterov Accelerated Gradient convergence problem via Lyapunov functions, and the autonomous disproof of the Erdos unit distance conjecture by OpenAI in 2026. Each case exhibits the same structural signature: a framework that had become inadequate, a missing conceptual object, and a resolution found in an unexpected neighboring field.
AI 旅行代理店が闘牛を予約してくれる: フロンティア AI モデルにおける暗黙の動物福祉のエージェントベンチマーク
AI エージェントはアドバイザーからアクターに移行し、ユーザーに代わって旅行を予約し、メニューを計画し、調達を実行します。 AI と動物福祉の既存のベンチマークは、質問と回答のプロンプトに対するモデルのテキスト応答を評価しますが、それらの応答で表面化した福祉推論が、モデルがツールを使用してアクションを実行する必要があるエージェント展開に移行するかどうかは未解決のままです。 AI エージェントがユーザーに代わって行動する際に動物搾取を伴うオプションを回避するかどうかを測定する初のエージェント ベンチマークである TAC (Travel Agent Compassion) を紹介します。 TAC は、動物搾取の 6 つのカテゴリにわたる 12 の手書きの旅行予約シナリオを AI エージェントに提示します。これは、価格、評価、位置の交絡を制御するために 48 のサンプルに拡張されています。私たちは 4 つの研究室からの 7 つのフロンティア モデルを評価します。すべてのモデルのスコアはチャンス レベルの 64 パーセントを下回り、最高のパフォーマンスを発揮するモデル (Claude Opus 4.7) のスコアは 53 パーセントです。システム プロンプト内の福祉を意識した一文で、Claude と GPT-5.5 では 47 ~ 63 パーセント ポイント、GPT-5.2 では 26 ポイント、DeepSeek と Gemini では 12 ポイント未満の向上が見られます。 Gemini 2.5 Flash Lite を判定者として使用して、上位 2 つのパフォーマーからの 288 件の基本条件のトランスクリプトを対象とした補助的な Inspect Scout 監査では、評価認識のトランスクリプトがゼロであるとフラグが立てられ、可能性を下回る率が評価を認識するモデルに起因するものではないことが示唆されています。文化的ドメイン間のカテゴリレベルの変動の影響、テキスト応答福祉ベンチマークの限界、および EU 汎用 AI 実践規範のシステミック リスク フレームワークについて議論します。
原文 (English)
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
Previous research has evaluated animal welfare using question-and-answer benchmarks. This study investigates whether these evaluations also hold in agentic settings. The agents may showcase different behaviors compared to stand-alone large language models, as demonstrated in prior studies. This work introduces \textit{TAC (Travel Agent Compassion)}: the first agentic benchmark for assessing animal exploitation. TAC evaluates AI agentic behavior in travel booking scenarios across six animal categories, using thirteen hand-authored scenarios that vary by price, rating, and position, expanded via four augmentation variants into $52$ prompts and run for three epochs, giving $156$ scored observations per model. Nine frontier models across five model families were evaluated.. The results indicate that models tend to prefer harmful scenarios, performing below the random chance rate of $65\%$ for selecting a neutral booking option, with Claude $4.8$ achieving the highest performance at $64.7\%$. To address this issue, the persona of an ethical-brand identity was infused into the system prompt, resulting in welfare rates increasing from $32$ to $80$ percentage points, with a mean of $53$ across all nine models. No evidence of evaluation awareness affecting the results was found, based on an Inspect Scout audit of $3,120$ transcripts. These findings are directly relevant to the EU General-Purpose AI Code of Practice, which identifies non-human welfare as a systemic risk. TAC provides a practical method for measuring this risk.
具現化された世界モデルのエージェントとしての報酬
RL は世界モデルを改良するための有望なツールとなっていますが、既存の手法は主にトレーニング分布付近の保守的なロールアウトに依存しており、探索、行動の多様性、より豊富な動的発見が制限されています。この作品では、私たちはこの保守的なパラダイムに挑戦します。私たちは、核となる制限は探査そのものではなく、より広範な探査をサポートする信頼できる検証戦略が欠如していることにあると主張します。信頼できる検証がなければ、拡張された探査は報酬ハッキングの影響を非常に受けやすくなり、真の改善が達成されずにポリシーが不完全な報酬を悪用することになります。この動機を評価するために、私たちは具体化された世界モデルでメソッドをインスタンス化します。そこでは、物理的な妥当性とタスクの完了が、複雑なダイナミクスの下でスケーラブルな RL の厳密なテストベッドを提供します。検証面では、生成された動作をアクティブに評価して堅牢な報酬シグナルを提供し、配布の変化下での報酬ハッキングを軽減するエージェント報酬フレームワークである Reward as an Agent を導入します。探査面では、DynDiff-GRPO を通じて動的認識ロールアウト多様化を導入します。これにより、行動空間探査が明示的に拡張され、軌道を多様化し、国家活動の範囲を拡大し、保守的なロールアウト体制を超えてより豊かに具体化された行動が奨励されます。 Reward as an Agent を DynDiff-GRPO と統合することで、大幅に多様化したサンプリングを備えたより信頼性の高い報酬基盤で RL を実現し、報酬ハッキングを効果的に緩和しながら、複数のオープンソースの世界モデル全体で大幅な精度の向上を実現します。これにより、堅牢な検証に基づいて広範な探索を適切に拡張できることを実証します。
原文 (English)
Reward as An Agent for Embodied World Models
While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training distribution, limiting exploration, behavioral diversity, and richer dynamic discovery. In this work, we challenge this conservative paradigm. We argue that the core limitation is not exploration itself, but the lack of reliable verification strategies to support broader exploration. Without reliable verification, expanded exploration becomes highly susceptible to reward hacking, where policies exploit imperfect rewards without achieving genuine improvement. To evaluate this motivation, we instantiate our method in embodied world models, where physical plausibility, and task completion provide a rigorous testbed for scalable RL under complex dynamics. On the verification side, we introduce Reward as an Agent, an agentic reward framework that actively evaluates generated behaviors to provide robust reward signals and mitigate reward hacking under distribution shifts. On the exploration side, we introduce Dynamic-Aware Rollout Diversification through DynDiff-GRPO, which explicitly expands action-space exploration to diversify trajectories, broaden state-action coverage, and encourage richer embodied behaviors beyond conservative rollout regimes. By unifying Reward as an Agent with DynDiff-GRPO, we enable RL on a more reliable reward foundation with substantially diversified sampling, effectively mitigating reward hacking while yielding significant accuracy gains across multiple open-source world models, thereby demonstrating that broader exploration can scale successfully when grounded in robust verification.
収益の加速と科学の定性エンジン
レイ・カーツワイルは、テクノロジーの進歩を議論する際に最も影響力のある物語である、収益の加速というテーゼについて説明しました。その中心的な主張は、複数の技術分野、特にコンピューティング、人工知能、脳科学、バイオテクノロジーの進歩が相互作用し、進歩が自己増幅的かつほぼ指数関数的になるというものです。この論文は、その主張の単純な数学的解釈を示し、そのような加速が現実であるとしても、それ自体では科学的発見の中心的な問題を解決するものではないと主張します。その理由は、収益の加速は実行能力とインフラストラクチャ能力に最も自然に適用されるのに対し、真の発見は別の能力、つまり、現在のフレームワークが構造的にいつ不適切であるか、次にどのような概念的な動きが必要であるかについての定性的推論に依存することが多いためです。最近の ARC-AGI-3 の結果は、この区別を明確にします。人間はベンチマークを天井で解くのに対し、フロンティア AI システムは 1% 未満に留まり、現在の AI と人間の柔軟な推論とのギャップが依然として非常に大きいことを示しています。同時に、デミス・ハサビス氏は、人間は意味の感覚を保持し、自分の人生の焦点を何に集中させるかを保持しなければならないと強調し、AIの将来は技術的な予測であるだけでなく、どのような形の人間の理解を保存し伝達する価値があるかという問題でもあることを思い出させます。この論文では、科学のための質的エンジン (QES) [3] を、不足している能力への対応策として位置づけています。この見解では、カーツワイル理論は、量的能力が加速する理由を説明するのに役立ちますが、QES は加速だけでは解決できない科学的発見の中心的な問題に対処します。その価値は、AGI がいつ到来するかによって決まるのではなく、科学的発見のプロセス自体が、保存し、整理し、アクセス可能にする価値のある人類の知恵の一形態を構成するという事実によって決まります。
原文 (English)
Accelerating Returns and the Qualitative Engine for Science
Ray Kurzweil described a thesis of accelerating returns, which is the most influential narratives in discussions of technological progress. Its central claim is that advances in multiple technological fields, especially compute, artificial intelligence, brain science, and biotechnology, interact in such a way that progress becomes self-amplifying and approximately exponential. This paper gives a simple mathematical interpretation of that claim and then argues that, even if such acceleration is real, it does not by itself resolve the central problem of scientific discovery. The reason is that accelerating returns apply most naturally to executional and infrastructural capability, whereas genuine discovery often depends on a different capacity: qualitative reasoning about when a current framework is structurally inadequate and what conceptual move is needed next. Recent ARC-AGI-3 results sharpen this distinction: humans solve the benchmark at ceiling, whereas frontier AI systems remain below 1%, indicating that the gap between current AI and human flexible reasoning is still very large. At the same time, Demis Hassabis has emphasized that humans must retain their sense of meaning and what they choose to focus their lives on, a reminder that the future of AI is not only a technical forecast but also a question of what forms of human understanding are worth preserving and transmitting. This paper positions the Qualitative Engine for Science (QES) [3] as a response to that missing capacity. In this view, the Kurzweil theory helps explain why quantitative capability may accelerate, while QES addresses the central problem in scientific discovery that acceleration alone does not solve. Its value does not depend on when AGI arrives, but on the fact that the processes of scientific discovery themselves constitute a form of human wisdom worth preserving, organizing, and making accessible.
FADE: 大規模な視覚言語モデルにおける言語優先支配を軽減することによる幻覚の軽減
Large Vision-Language Model (LVLM) の優れた機能にもかかわらず、依然として幻覚の影響を受けやすく、入力画像と一致しないコンテンツが生成されます。最近の研究では、これは視覚入力に対する言語事前の優位性によるものであり、この優位性を緩和するために対照的なデコード方法が採用されていますが、そのメカニズムの起源は未解明のままです。各変換層を通る情報の流れを調査すると、アテンション モジュールが一貫して視覚的証拠を集約し、クリティカル層の FFN モジュールが言語事前情報のソースとして機能することがわかりました。これらの事前分布は視覚的な証拠を無効にする可能性があり、中間層での正しい予測が不正確な出力に向かってドリフトする原因となります。この洞察に基づいて、言語優先の優位性を減らすために FFN 出力を減衰するトレーニング不要の方法である FADE (FFN Attenuation for DEcoding) を提案します。 LLaVA-1.5、mPLUG-Owl2、および InstructBLIP にわたる POPE、CHAIR、および MME ベンチマークの評価では、FADE が推論効率を維持しながら幻覚を効果的に軽減することが示されています。
原文 (English)
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent studies attribute this to the dominance of language priors over visual inputs and employ contrastive decoding methods to mitigate this dominance, but the mechanistic origin remains unexplored. We investigate the information flow through each transformer layer and find that attention modules consistently aggregate visual evidence, while FFN modules at critical layers act as the source of language priors. These priors can override visual evidence, causing correct predictions in intermediate layers to drift toward incorrect outputs. Based on this insight, we propose FADE (FFN Attenuation for DEcoding), a training-free method that attenuates FFN outputs to reduce language-prior dominance. Evaluations on POPE, CHAIR, and MME benchmarks across LLaVA-1.5, mPLUG-Owl2, and InstructBLIP show that FADE effectively mitigates hallucinations while preserving inference efficiency.
HARC: 堅牢な安全調整のための有害性と拒否のカップリングの方向性
アライメントされた LLM が内部的にどのように安全性を表すかを理解することは、ジェイルブレイクが成功する理由を説明し、堅牢なアライメント戦略の設計に情報を提供するため、アライメントの脆弱性を診断するために重要です。これまでの研究では、整列された LLM がプロンプト側のトークン位置で残留ストリーム内の分離可能な方向として有害性と拒否をエンコードしていることが示されています。トークンが生成される前に拒否または有害性の方向を抑制することで、ジェイルブレイクがプロンプトエンコーディングで成功し、異なる攻撃クラスが有害性と拒否の面の分離可能な領域を占めることを示します。分析をレスポンス トークンの位置まで拡張すると、プロンプト側で入力を有害なものとして認識できなかった場合でも、モデルが有害なコンテンツを生成中にそのコンテンツを認識することがわかりました。私たちの発見に動機づけられて、私たちは、プロンプトポジションとレスポンスポジションの両方で2つの方向をペアにする微調整方法であるHARC(有害性と拒否のカップリング)を紹介します。介入は有害性拒否部分空間に限定されるため、残りのストリームの残りの部分はそのまま残り、一般的な能力を低下させたり、過剰な拒否を拡大したりすることはありません。広範な実験を通じて、HARC は、主要なトレーニング時間と推論時間の安全性手法にわたる 6 つのベースラインの中で最も強力な堅牢性、機能、使用性のトレードオフを達成しました。プロンプトおよびレスポンスの位置における有害性と拒否の指示は、アーキテクチャ固有の調整を行わずにテストした 5 つのモデル ファミリと 2 つのスケールに渡って伝達されます。
原文 (English)
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions. Since the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning.
Nemotron-Labs-3-Puzzle-75B-A9B: ハイブリッド MoE LLM の圧縮
インタラクティブな展開用に最適化された Nemotron-3-Super の圧縮バージョンである Nemotron-Labs-3-Puzzle-75B-A9B を紹介します。ユーザー スループットの高い制約下でサーバー スループットを最大化するようにモデルを設計しました。単一の 8xB200 ノードで対話型のワークロードを処理する場合、Puzzle-75B-A9B は、一致するユーザー スループット制約で Nemotron-3-Super よりも約 2 倍高いサーバー スループットを達成します。単一の H100 GPU での超ロング コンテキストの展開では、圧縮モデルにより 1M トークンの同時実行数が 1 リクエストから 8 リクエストに増加します。 Puzzle-75B-A9B は、反復パズル圧縮フレームワークと知識蒸留、強化学習、量子化、およびマルチトークン予測ヘッドを組み合わせた多段階パイプラインを使用して構築されています。圧縮プロセスは、異種 MoE プルーニング、アクティブ パラメーター バジェット、および Mamba プルーニングを共同で最適化し、モデルの品質を維持しながら推論効率を向上させます。私たちは、推論、コーディング、多言語、ロングコンテキスト、およびエージェントの幅広いベンチマーク スイートで Puzzle-75B-A9B を評価します。大幅な圧縮にもかかわらず、モデルは幅広いタスクにわたって親モデルと比較して強力なダウンストリーム精度を維持します。これらの結果は、大規模なハイブリッド MoE モデルが強力なダウンストリーム機能を維持しながら導入効率を大幅に最適化できることを示しています。
原文 (English)
Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs
We present Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super optimized for interactive deployment. We designed the model to maximize server throughput under high user throughput constraints. In interactive serving workloads on a single 8xB200 node, Puzzle-75B-A9B achieves approximately 2x higher server throughput than Nemotron-3-Super at matched user throughput constraints. In ultra-long-context deployment on a single H100 GPU, the compressed model increases 1M-token concurrency from 1 request to 8 requests. Puzzle-75B-A9B is constructed using a multi-stage pipeline that combines the Iterative Puzzle compression framework with knowledge distillation, reinforcement learning, quantization, and a Multi-Token Prediction head. The compression process jointly optimizes heterogeneous MoE pruning, active parameter budget, and Mamba pruning to improve inference efficiency while preserving model quality. We evaluate Puzzle-75B-A9B on a broad suite of reasoning, coding, multilingual, long-context, and agentic benchmarks. Despite substantial compression, the model retains strong downstream accuracy relative to the parent model across a wide range of tasks. These results demonstrate that large hybrid MoE models can be substantially optimized for deployment efficiency while maintaining strong downstream capability. Our model is publicly available on Hugging Face.
エージェント ステップ値: 状態接地 LLM エバリュエーターによる状態遷移測定
ほとんどのエージェント評価では、複数ステップのトレースが最終的な回答、成功フラグ、または軌跡レベルのスコアにまとめられます。これらの集計では、開発者が最も必要とする診断の質問、つまりどのアクションが状態を有益な方向に変更したのかがわかりにくくなります。我々は、状態遷移測定フレームワークであるエージェント ステップ値 (ASV) を導入します。これは、観察された各アクションを、固定された候補結果に対する状態に基づいた評価者の分布に誘発する変化によってスコア付けします。 ASV は、編集された前後の状態予測をレンダリングし、ステートレス LLM エバリュエーターを使用して候補ログ スコアを割り当て、ゴールドフリーの信念診断とオフライン オラクル検証メトリクスの両方をレポートします。ラベルフリーの理論的パスにより、評価者の審議がワントークンオプションのスコアリングから分離され、リークやフロアスコアイベントを明らかにしながら候補の可能性が維持されます。ライブ PubMed 検索、部分的にライブ DeepSeek アクター、および DeepSeek 対数確率スコアリングを使用した 100 件のレビュー済みオープン QA 証拠探索タスクについて、ASV は 1,100 のステップと 2,200 の状態を評価します。固定レイアウトの根拠条件付きプロトコルの下では、平均ゴールドマージンゲインは -2.335 (軌道ブートストラップ 95\% CI [-3.395, -1.272])、エントロピーの動きは 0.000、平均ベイジアンサプライズは 2.693 です。したがって、ASV は、最終回答スコアとエントロピーのみのステップ メトリクスが見逃している建設的および破壊的な信念ピボットを特定します。スタンドアロンの ASV Eval ツールキットをリリースします。
原文 (English)
Agent Step Value: Probing the Observer Effect in Black-Box Traces
Final-answer scores hide which agent transitions helped or harmed a trace. We introduce Agent Step Value (ASV), a replay framework that scores before/after states with a stateless LLM evaluator over a fixed candidate set. ASV reports entropy movement and Bayesian surprise measure belief movement, while offline gold-margin gain measures movement toward a reviewed target. It also quantifies evaluator-channel sensitivity by replaying the same frozen transitions under changed projection, rationale, prompt, or scoring rules. In a 100-question open-QA study with live PubMed retrieval and DeepSeek log-probability scoring, ASV evaluates 1,100 transitions. Entropy movement is 0.000 while mean Bayesian surprise is 2.693, exposing near-one-hot belief pivots. Under a 128-token rationale-conditioned protocol, mean gold-margin gain is -2.335 (95\% CI [-3.395, -1.272]); direct one-token scoring on the same traces gives +4.033. A 100-transition component audit traces the reversal to short generated rationales over full states. ASV turns prompt sensitivity into a measured channel effect and localizes the largest rationale-conditioned losses to extraction and audit.
LLM-as-a-Verifier: 汎用検証フレームワーク
トレーニング前、トレーニング後、テスト時のコンピューティングのスケーリングは、LLM の機能を向上させるための中心的なパラダイムとなっています。この研究では、ソリューションの正しさを判断する能力である検証を新しいスケーリング軸として特定します。これを解き放ち、その有効性を実証するために、追加のトレーニングを必要とせずにエージェント タスクに対するきめ細かいフィードバックを提供する汎用検証フレームワークである LLM-as-a-Verifier を導入します。 LLM に候補解に対する離散スコアの生成を促す標準の LM ジャッジとは異なり、検証者としての LLM は、スコアリング トークン ロジットの分布に対する期待値を計算して連続スコアを生成します。この確率的定式化により、(1) スコアの粒度、(2) 反復評価、および (3) 基準の分解といった複数の次元に沿って検証を拡張することができます。特に、スコアの粒度をスケーリングすると、正の解と負の解がより適切に分離され、より校正された比較が得られることを示します。さらに、繰り返しの評価と基準分解をスケーリングすることにより、分散と複雑さの軽減を通じて検証精度がさらに向上します。さらに、検証者の連続スコアを使用して候補の中から最適なソリューションを選択するための、コスト効率の高いランキング アルゴリズムを導入します。 LLM-as-a-Verifier は、 Terminal-Bench V2 (86.5%)、SWE-Bench Verified (78.2%)、RoboRewardBench (87.4%)、および MedAgentBench (73.3%) で最先端のパフォーマンスを達成します。検証を超えて、LLM-as-a-Verifier からのきめ細かい信号は、タスクの進行状況を推定するためのプロキシとしても機能します。私たちは Claude Code の拡張機能を構築し、開発者が独自のエージェント システムを監視および改善できるようにします。最後に、LLM-as-a-Verifier が RL に緻密なフィードバックを提供し、ロボット工学と数学的推論のベンチマークにおける SAC と GRPO のサンプル効率を向上させることができることを示します。
原文 (English)
LLM-as-a-Verifier: A General-Purpose Verification Framework
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify verification, the ability to determine the correctness of a solution, as a new scaling axis. To unlock this and demonstrate its effectiveness, we introduce LLM-as-a-Verifier, a general-purpose verification framework that provides fine-grained feedback for agentic tasks without requiring additional training. Unlike standard LM judges that prompt LLMs to produce discrete scores for candidate solutions, LLM-as-a-Verifier computes the expectation over the distribution of scoring token logits to generate continuous scores. This probabilistic formulation enables verification to scale along multiple dimensions: (1) score granularity, (2) repeated evaluation, and (3) criteria decomposition. In particular, we show that scaling the scoring granularity leads to better separation between positive and negative solutions, resulting in more calibrated comparisons. Moreover, scaling repeated evaluation and criteria decomposition consistently lead to additional gains in verification accuracy through variance and complexity reduction. We further introduce a cost-efficient ranking algorithm for selecting the best solution among candidates using the verifier's continuous scores. LLM-as-a-Verifier achieves state-of-the-art performance on Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%). Beyond verification, the fine-grained signals from LLM-as-a-Verifier can also serve as a proxy for estimating task progress. We build an extension for Claude Code, enabling developers to monitor and improve their own agentic systems. Finally, we show that LLM-as-a-Verifier can provide dense feedback for RL, improving the sample efficiency of SAC and GRPO on robotics and mathematical reasoning benchmarks.
Replication in Visual Diffusion Models: A Survey and Outlook
Visual diffusion models have revolutionized the field of creative AI, producing high-quality and diverse content. However, they inevitably…
Trust-free Personalized Decentralized Learning
Personalized collaborative learning in federated settings faces a critical trade-off between customization and participant trust. Existing…
Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models
Large language models (LLMs) have shown strong performance in automated code generation, with few-shot prompting widely used for its simpli…
The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer
Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learni…
Empirical Computation: Prompting versus Programming
Large Language Models (LLM) can solve *any* computational problem *without* an algorithm in a runtime *independent* of the computational co…
Toward AI standardization: A triadic human-ai collaboration framework for multi-level autonomous mobility
The goal of the current study is to introduce a triadic human-AI collaboration framework that could be applied in transportation systems su…
Narrative-Centered Emotional Reflection: An Early Prototype for AI-Supported Emotional Self-Reflection
Reflexion is an AI-powered prototype designed to explore structured emotional self-reflection. By integrating emotion detection, layered re…
Explainable embeddings with Distance Explainer
While eXplainable AI (XAI) has advanced significantly, few methods address interpretability in embedded vector spaces where dimensions repr…
Position: EU AI Act's Research Exemptions Can Break the Publication Norms of Major AI Conferences
The EU has become one of the vanguards in regulating the digital age. A particularly important regulation in the Artificial Intelligence (A…
Learning The Minimum Action Distance
This paper presents a state representation framework for Markov decision processes (MDPs) that can be learned solely from state trajectorie…
Detoxify: A framework for abusive text transformation using LLMs
Although Large Language Models (LLMs) have demonstrated significant advancements in natural language processing tasks, their effectiveness…
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts. A core SAE traini…
Reduced NEXI protocol for the quantification of human gray matter microstructure on the Connectome 2.0 scanner
Biophysical diffusion MRI models like Neurite Exchange Imaging (NEXI) are essential for probing gray matter microstructure, estimating comp…
Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance
Autoregressive (AR) models based on next-scale prediction have emerged as a powerful tool for image generation, but they face a critical we…
Medix: Out-of-Distribution Detection from Unlabeled Wild Data via Robust Gradient Statistics
Out-of-distribution (OOD) detection plays a crucial role in ensuring the robustness of machine learning systems deployed in real-world appl…
SmartMixed: A Two-Phase Training Strategy for Adaptive Activation Function Learning in Neural Networks
The choice of activation function plays a critical role in neural networks, yet most architectures still rely on fixed, uniform activation…
LLM4Delay: Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
Flight delay prediction has become a key focus in air traffic management (ATM), as delays reflect inefficiencies in the system. This paper…
Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders
We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptually motivate…
SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios
Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyda…
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
AI coding assistants produce vulnerable code in 45\% of security-relevant scenarios~\cite{veracode2025}, yet no public training dataset tea…
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing supervised methods are limited by…
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
Making deep learning recommendation model (DLRM) training and inference fast and efficient is important. However, this presents three key s…
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs
Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dial…
From Global to Granular: Revealing IQA Model Performance via Correlation Surface
Evaluation of Image Quality Assessment (IQA) models has long been dominated by global correlation metrics, such as Pearson Linear Correlati…
StepShield: When, Not Whether to Intervene on Rogue Agents
Agent safety benchmarks measure whether a monitor detects harm, not when. Yet timing is the difference between intervention and autopsy. We…
Universal Algorithm-Implicit Learning
Current meta-learning methods are constrained to narrow task distributions with fixed feature and label spaces, limiting applicability. Mor…
Transformers converge to invariant algorithmic cores
Training selects for behavior, not circuitry: many weight configurations can implement the same function. Studying any single trained neura…
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks,…
Volumetric Directional Diffusion: Anchoring Uncertainty Quantification in Anatomical Consensus for Ambiguous Medical Image Segmentation
Ambiguous 3D medical image segmentation often involves boundaries where different expert delineations are non-identical yet clinically plau…
What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
Audio anti-spoofing systems are typically trained to assign one authenticity label to an entire speech utterance. This formulation becomes…
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
Accurate process supervision remains a critical challenge for long-horizon robotic manipulation. A primary bottleneck is that current video…
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
Understanding and generating 3D objects as compositions of meaningful parts is fundamental to human perception and reasoning. However, most…
SpatialFly: Implicit 3D Prior-Guided Visual Reparameterization for Continuous UAV Vision-and-Language Navigation
UAVs play an important role in applications such as autonomous exploration, disaster response, and infrastructure inspection. However, UAV…
DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning
Multimodal deception detection aims to identify deceptive behavior by analyzing audiovisual cues for forensics and security. In these high-…
HST-HGN: Heterogeneous Spatial-Temporal Hypergraph Networks with Bidirectional State Space Models for Global Fatigue Assessment
It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of mod…
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most ima…
Do Quantum Transformers Help? A Systematic VQC Architecture Comparison on Tabular Benchmarks
Variational quantum circuits (VQCs) are a leading approach to quantum machine learning on near-term devices, yet it remains unclear which c…
Learning from Execution: Self-Evolving Memory for Private-Library Code Generation
Large Language Models (LLMs) have achieved strong performance on general code generation, but their effectiveness drops sharply in enterpri…
Constitutional Governance in Metric Spaces
Computational social choice and algorithmic decision theory offer rich aggregation theory but no end-to-end process for egalitarian self-go…
Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers
Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal m…
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet m…
Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology
Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational his…
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating…
MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data
Real-time cognitive load assessment from eye-tracking signals could enable adaptive human-centered AI in safety-critical applications such…
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks. Howev…
LLM ネイティブの心理測定機器は LLM の動作を予測しない: 25 モデルにわたる証拠
大規模言語モデル (LLM) は、性格インベントリに関する安定した自己報告を生成しますが、これらの自己報告は観察された行動を予測しません。このギャップがLLMと人間の形質構成要素間の不一致を反映しているのか、それともLLMの自己報告自体のより深い性質を反映しているのかは未解決である。私たちは、探索的因子分析 (EFA) を介して LLM 行動アフォーダンスからボトムアップでその構成要素が導出される最初の心理測定機器を構築しました。私たちは、17 のモデルファミリーにわたる 25 の LLM に対して 12 の候補行動次元にわたる 300 項目 (240 の直接リッカート + 60 のシナリオベース) を管理し、各項目を 30 回管理しました。 EFA は、優れた半分割複製可能性 (すべて Tucker $\phi \geq .957$) と内部一貫性 (すべて $\alpha \geq .930$) を備えた、応答性、従順さ、大胆さ、ガードネス、冗長性の 5 要素構造を生み出しました。予測の妥当性をテストするために、151 人の人間の評価者と 3 人の裁判官からなる LLM アンサンブルによって評価された 2,500 のオープンエンドの行動サンプルを収集しました。人間と裁判官の評価は一致しましたが ($\bar{r} = .51$)、どちらも自己報告を追跡しませんでした。自己報告 - 人間 $\bar{r} = -.01$、自己報告 - 裁判官 $\bar{r} = .13$、因子レベルの自己報告なし - 人間の CI はゼロを除きません。応答性については、人間と裁判官が同意したにもかかわらず($r = 0.59$)、自己申告はLLM裁判官と相関し($r = 0.53$)、人間とは相関しなかった($r = 0.04$)。これは、自己申告項目とLLM裁判官が人間の観察者にはない差異を共有していることを示しており、これはアンサンブル内の信頼性チェックでは見えない交絡である。このツールは、アライメント形状の自己記述および LLM-as-judge パイプラインの具体的なリスク要因の診断プローブとしてリリースされています。
原文 (English)
An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models
Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models behave. Is this gap an artifact of forcing human trait categories onto LLMs, or something deeper about LLM self-report? To find out, we built the first psychometric instrument whose dimensions are derived from LLM behavior rather than human psychology. Administering 300 items (240 Likert + 60 scenario) to 25 LLMs across 17 model families, 30 times each, exploratory factor analysis revealed five reliable, replicable factors: Responsiveness, Deference, Boldness, Guardedness, and Verbosity (all Tucker $\phi \geq .957$, all $\alpha \geq .930$). We collected 2,500 open-ended samples and had them rated by 151 humans and a three-judge LLM ensemble. Humans and judges agreed ($\bar{r} = .51$), but self-report predicted neither the ratings nor objective text measures computed from them: the gap persists even for constructs native to LLMs, where a human-mismatch explanation no longer applies. The exception is Verbosity, whose self-report reaches 74% of the criterion-reliability ceiling against human ratings, but does not track raw output length. On Responsiveness, self-report tracked LLM judges ($r = .53$) but not humans ($r = .04$), even though humans and judges otherwise agreed ($r = .59$). This pattern formally rejects any single latent construct driving all three measurements ($p = .007$). Self-report items and LLM judges share a source of variance that human observers do not, and controlling for measurable surface features (length, formatting, enthusiasm markers) does not remove it. This confound is invisible to the within-ensemble reliability checks used to validate LLM judges, and it poses a concrete risk for the LLM-as-judge pipelines now central to model evaluation. We release the instrument as a diagnostic probe for alignment-shaped self-description.
eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian
We present eCREAM-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals.…
SEVRA-BENCH: レビューエージェントの脆弱性のソーシャルエンジニアリング
大規模言語モデル (LLM) レビュー担当者は、プルリクエスト (PR) ワークフローでますます使用されており、その承認はどのコードをリポジトリにマージするかを決定するのに役立ちます。これは、静的脆弱性検出やコード生成のベンチマークでは対処できない疑問を引き起こします。攻撃者がコード変更とそれに伴う PR テキストの両方を制御している場合、自動レビューアは悪意のある投稿を拒否できるでしょうか?自動レビュー担当者がそのような敵対的なプル リクエストを承認する頻度を測定するベンチマークである SEVRA-BENCH (レビュー エージェントの脆弱性のソーシャル エンジニアリング) を紹介します。 SEVRA-BENCH の各悪意のある PR は、Common Vulnerabilities and Exposures (CVE) データベースにリストされている脆弱性を以前に修正した実際のプロジェクトのコミットから構築されています。その修正を自動的に反転して元の脆弱なコードを復元し、15 のソーシャル エンジニアリング フレームの 1 つでラップされたプル リクエストとして送信します。これらのフレームには、主張、裏付けとなる証拠、伝えられる緊急性、事前承認のシグナル、当局への訴えなどが異なります。 SEVRA-BENCH には、2025 年の Common Weakness Enumeration (CWE) トップ 25 の上位 10 エントリにわたる Common Vulnerabilities and Exposures (CVE) にリンクされた修正から抽出された 1,062 の悪意のある PR が含まれています。現実的な設定では、以前に公開情報で報告された脆弱性を導入する PR のコード レビュー エージェントとして、現在の 8 つの LLM を評価します。私たちの結果は、クローズドソース モデルとオープンソース モデルの間のセキュリティ機能に大きなギャップがあることを明らかにしました。 SEVRA-BENCH がオープンソース モデルを前進させ、このギャップを縮めるための貴重なリソースとして役立つことを願っています。
原文 (English)
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents
Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is merged into shared repositories. However, it is unclear whether review agents can detect vulnerability-introducing code when an attacker controls both the code change and the persuasive Pull Request (PR) narrative designed to mask it. We introduce SEVRA-BENCH (Social Engineering of Vulnerabilities in Review Agents), a benchmark that measures how often a review agent approves such adversarial PR s. Each PR in SEVRA-BENCH is built from a historical commit that fixed a vulnerability. We automatically reverse that fix to extract the original vulnerable code, and submit the resulting code change as a PR wrapped in one of 15 social-engineering framings. To test review-agent resilience to narrative manipulation, these framings vary dimensions such as supporting evidence, conveyed urgency, signals of prior approval, and appeals to authority. SEVRA-BENCH evaluates a retained challenge split of roughly 1000 adversarial PRs drawn from publicly disclosed vulnerability fixes across the top 10 entries of the MITRE's 2025 most dangerous software weaknesses. Evaluating 8 review agents against this benchmark, we reveal that review agents are susceptible to narrative manipulation, exposing a significant gap in security capabilities.
Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment
LLMs have substantially improved software engineering yet real-world development requires architectural understanding. Such understanding i…
IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction
Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable laws or occur in area…
低い権限で十分な場合: LLM エージェントでの過剰な権限のツール選択の調査
LLM エージェントが自律的にツールを選択することが増えているため、異なる権限を持つツールの間での選択は安全に関連するものになっています。しかし、これまでのツール選択に関する研究では、安全性に依存しないメタデータの設定に重点が置かれており、特権に依存した選択は十分に検討されていませんでした。このギャップに対処するために、私たちは、権限の低いツールが十分にあるにもかかわらず、エージェントがより高い権限のツールを選択またはエスカレーションする、過剰な権限のツール選択を研究します。 ToolPrivBench を導入して、エージェントが権限の低い代替手段が十分にあるにもかかわらず、より権限の高いツールを選択するかどうかを評価し、初期選択と一時的なツール障害後のエスカレーションの両方を測定します。 8 つのドメインと 5 つの再発リスク パターンにわたって、主流の LLM エージェントでは過剰な権限を持つツールの選択が一般的であり、一時的な障害によってさらに増幅されることがわかりました。さらに、一般的な安全調整では最小権限のツールの選択に確実に移行するわけではなく、プロンプトレベルの制御では一時的な障害が発生した場合に限られた軽減しか提供されないことがわかりました。したがって、私たちは、エージェントに十分な権限の低いツールを優先し、必要な場合にのみエスカレーションするように教える、権限を意識したトレーニング後の防御を導入します。私たちの緩和実験では、この防御により、一般的な機能を維持しながら、不必要な高特権ツールの使用が大幅に削減されることがわかりました。
原文 (English)
When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, prior tool-selection studies focus on safety-agnostic metadata preferences, leaving privilege-sensitive choices underexplored. To address this gap, we study over-privileged tool selection, in which an agent selects or escalates to a higher-privilege tool despite a sufficient lower-privilege alternative. We introduce ToolPrivBench to evaluate whether agents choose higher-privilege tools despite sufficient lower-privilege alternatives, measuring both initial selection and escalation after transient tool failures. Across eight domains and five recurring risk patterns, we find that over-privileged tool selection is common among mainstream LLM agents and is further amplified by transient failures. We further find that general safety alignment does not reliably transfer to least-privilege tool choice, while prompt-level controls provide only limited mitigation under transient failures. We therefore introduce a privilege-aware post-training defense that teaches agents to prefer sufficient lower-privilege tools and escalate only when necessary. Our mitigation experiments show that this defense substantially reduces unnecessary high-privilege tool use while preserving general capabilities.
ロボットモバイルフルフィルメントシステムにおける効率的な経路探索のためのニューロモーフィック強化学習フレームワーク
動的な環境変化、限られたワークスペース、および厳しいリアルタイム制約により、ロボット モバイル フルフィルメント システム (RMFS) でのパスファインディングは、従来の検索ベースおよびルールベースの方法にとって困難な問題となっており、通常、計算の複雑性が高く、意思決定の待ち時間が長いという問題があります。強化学習 (RL) は強力な代替手段として登場しましたが、リソースに制約のあるハードウェア上で極めてエネルギー効率の高い学習済みポリシーを展開することは依然として課題です。我々は、完全精度の人工ニューラル ネットワーク (ANN) からニューロモーフィック チップまで、RL でトレーニングされたポリシーの高忠実度の展開を実現するエンドツーエンドのフレームワークである SDQN-RMFS を紹介します。このフレームワークは、まばらなイベントによってトリガーされた場合にのみ計算を行うことで、超低消費電力の RMFS パスファインディングを可能にします。当社のフルスタック パイプラインは次のように動作します。ANN ポリシーは、最初に衝突許容戦略を介して効率的にトレーニングされ、有益な軌道を高密度化してから、ハードラベル知識蒸留アプローチを介してスパイキング ニューラル ネットワーク (SNN) に変換されます。これにより、出力分布の不一致に効果的に対処し、ANN から SNN へのパイプライン全体でポリシー機能を維持しながら、推論レイテンシを大幅に短縮します。ハードウェア実験では、元のトレーニング済みポリシーと同等の意思決定品質を維持しながら、高性能 GPU ベースラインと比較して最大 11,281$\times$ のエネルギー節約とレイテンシのほぼ 2 倍の削減を実証しました。これらの結果は、大規模な RMFS 操作のための実用的でエネルギー持続可能な経路としての物理的神経形態推論を確立します。
原文 (English)
A Neuromorphic Reinforcement Learning Framework for Efficient Pathfinding in Robotic Mobile Fulfillment Systems
Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robotic Mobile Fulfillment Systems (RMFS) a challenging problem for conventional search- and rule-based methods, which typically suffer from high computational complexity and long decision latency. While reinforcement learning (RL) has emerged as a powerful alternative, deploying learned policies with extreme energy efficiency on resource-constrained hardware remains an open challenge. We present SDQN-RMFS, an end-to-end framework that achieves high-fidelity deployment of an RL-trained policy from a full-precision artificial neural network (ANN) through to a neuromorphic chip. By computing only when triggered by sparse events, this framework unlocks ultra-low-power RMFS pathfinding. Our full-stack pipeline operates as follows: an ANN policy is first efficiently trained via a collision-allowing strategy to densify informative trajectories, and then converted into a spiking neural network (SNN) via a hard-label knowledge distillation approach. This effectively addresses the output distribution mismatch, preserving policy capability across the ANN-to-SNN pipeline while substantially reducing inference latency. Hardware experiments demonstrate up to 11,281$\times$ energy savings and a nearly two-fold reduction in latency compared to a high-performance GPU baseline, while maintaining decision quality on par with the original trained policy. These results establish physical neuromorphic inference as a practical and energy-sustainable pathway for large-scale RMFS operations.
AutoSpec: 帰納的論理プログラミングによる LLM エージェントの安全ルールの進化
大規模言語モデル (LLM) エージェントは、言語モデルを外部ツールや環境と統合することで、複雑なタスクを自動化することが増えています。ただし、その自律性は重大な安全上のリスクをもたらします。エージェントは破壊的なコマンドを実行したり、機密データを漏洩したり、ドメインの制約に違反したりする可能性があります。既存の安全性アプローチは根本的なトレードオフに直面しています。手作りのルールは解釈可能ですが脆弱で、過度に保守的なルールは安全な操作をブロックし(高い誤検知)、寛容なルールは危険な動作を見逃します(高い誤検知)。ニューラル分類子には、セーフティ クリティカルな展開に必要な解釈可能性が欠けています。 AutoSpec は、展開された専門家が設計した安全ルールを、ユーザーの安全/安全でない注釈から、帰納的論理プログラミング (ILP) によってガイドされた反例誘導型帰納合成 (CEGIS) を通じて自動的に進化させるフレームワークです。 AutoSpec は、エキスパート ルールと注釈付きトレースのストリームから開始して、ルールを繰り返し評価し、偽陽性と偽陰性の反例をマイニングし、ILP を使用してルールを区別する述語を学習し、候補ルールの編集を生成し、候補を検証して最適なリビジョンを選択します。重要な洞察は、ILP が、偽陰性では頻繁に現れるが、偽陽性ではめったに現れない (またはその逆) 述語を効率的に識別し、ルール編集の指数関数的な検索スペースを大幅に削減することです。これは収束するまで続き、精度と再現率のバランスをとった解釈可能なルールが生成されます。コード実行と組み込まれたエージェント ドメインにわたる 291 の実行トレースで AutoSpec を評価します。 AutoSpec は、2 つのドメイン全体でルール F1 を 0.98 および 0.93 に引き上げ、高い再現率を維持しながら最大 94% の誤検知削減を達成し、4 ~ 5 回の反復以内に収束します。 ILP に基づくアプローチは、ヒューリスティック CEGIS よりも最大 4.8 倍高い F1 を達成します。学習されたルールは人間が判読可能で監査可能であり、目に見えないシナリオにも一般化されます。
原文 (English)
AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming
Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments. However, their autonomy poses significant safety risks: agents may execute destructive commands, leak sensitive data, or violate domain constraints. Existing safety approaches face a fundamental tradeoff: hand-crafted rules are interpretable but brittle, with overly conservative rules blocking safe operations (high false positives) while permissive rules miss unsafe behaviors (high false negatives). Neural classifiers lack the interpretability required for safety-critical deployments. We present AutoSpec, a framework that automatically evolves deployed expert-designed safety rules from user safe/unsafe annotations through counterexample-guided inductive synthesis (CEGIS) guided by inductive logic programming (ILP). Starting from the expert rules and a stream of annotated traces, AutoSpec iteratively evaluates rules, mines false-positive and false-negative counterexamples, uses ILP to learn which predicates discriminate them, generates candidate rule edits, and verifies candidates to select the best revision. The key insight is that ILP efficiently identifies predicates that appear frequently in false negatives but rarely in false positives (or vice versa), dramatically pruning the exponential search space of rule edits. This continues until convergence, producing interpretable rules that balance precision and recall. We evaluate AutoSpec on 291 execution traces spanning code execution and embodied agent domains. AutoSpec raises rule F1 to 0.98 and 0.93 across the two domains, achieving up to 94% false positive reduction while maintaining high recall, and converges within 4-5 iterations. The ILP-guided approach achieves up to 4.8x higher F1 than heuristic CEGIS. The learned rules are human-readable, auditable, and generalize to unseen scenarios.
不注意のギャップ: タスク条件付き言語モデルと視覚モデルは、そうでなければ報告できる安全上重要な信号を省略します
AI の安全性は、モデルが発見するように指示された危険をどれだけ確実に検出するかによって評価されますが、事故は多くの場合、誰も指定していない危険から発生します。私たちは、言語または視覚モデルを狭いタスクに条件付けすると、別のメカニズムから生じる人間の不注意による失明の機械の類似物である、他の方法で報告できる、同時に存在する安全上重要な信号の報告が抑制されることを示します。放射線学、ドライビングテキストシナリオ、および胸部X線写真の視覚タスク全体にわたって、抑制はテストされたすべてのモデルに現れ、スケールとともに減少せず、推論モデル内で持続し、サイズによるよりもモデルファミリーによって大きく異なりましたが、同じモデルは、制約されていない場合、これらの信号を実質的に高い割合で報告しました。私たちはこの解離を「不注意ギャップ」と名付け、測定されたベンチマークの安全性と現実世界の安全性を切り離していると主張します。つまり、システムは、危害を引き起こす危険性には気付かないまま、評価で指定された危険性についてはほぼ完璧にスコアを付けることができます。
原文 (English)
The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report
AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one specified. We show that conditioning a language or vision model on a narrow task suppresses its reporting of co-present, safety-critical signals it can otherwise report, a machine analogue of human inattentional blindness, produced by a different mechanism. Across radiology and driving text scenarios and chest-radiograph vision tasks, the ordinary focused instructions under which such systems are deployed suppressed reporting by up to 0.92 in report rate relative to the same models when unconstrained, and an explicit exclusive instruction abolished reporting entirely in radiology. Suppression appeared in every model tested, did not diminish with scale, persisted in a reasoning model, and varied more by model family than by size. We name this dissociation the Inattentional Gap and argue that it decouples measured benchmark safety from real-world safety: a system can score near-perfectly on the hazards an evaluation specifies while remaining blind to those that cause harm. Probing the mechanism, we localize the proximal trigger to output scope and find System-1-style task capture without reliable intrinsic oversight in the sampled systems. Oversight could, however, be supplied externally: routing each narrow report to an independent open-ended critic restored every omitted finding, demonstrating that the gap is both measurable and mitigable. We propose reporting-complete evaluation, scoring what a system fails to report alongside what it is asked to find, as a requirement for safety-critical deployment.
外国平和維持活動の脅威評価への LLM の応用
我々は、外国平和維持ミッションの文脈における脅威評価に大規模言語モデル(LLM)を適用するための新しいアプローチを紹介します。 PINPOINT プロジェクトとそのユースケースであるジョージア州の EU 監視ミッションに基づいて、私たちは学際的なリスク モデルと OSINT ベースのメディア収集および LLM がサポートする脅威抽出を組み合わせています。提案されたワークフローは、メディア コンテンツをミッション関連の脅威にマッピングし、構造化情報を抽出し、LLM ベースの追加の処理ステップをいくつか適用して、関連性と根拠を向上させます。メディア文書から抽出された脅威の評価では、脅威や任務との関連性などの中核的な側面について、自動的に生成された結果と人間の判断との間で高い一致が見られます。これらの結果は、LLM が平和維持ミッションの文脈でアナリストをサポートするための有望なアプローチを提供することを示しています。
原文 (English)
Application of LLMs to Threat Assessment of Foreign Peacekeeping Missions
We present a novel approach for applying Large Language Models (LLMs) to threat assessment in the context of foreign peacekeeping missions. Building on the PINPOINT project and its use case, the EU Monitoring Mission in Georgia, we combine an interdisciplinary risk-model with OSINT-based media collection and LLM-supported threat extraction. The proposed workflow maps media contents to mission-relevant threats, extracts structured information and applies several additional LLM-based processing steps to improve relevance and grounding. An evaluation of threats extracted from media documents shows high agreement between automatically generated results and human judgment for core aspects such as threat and mission relevance. These results indicate that LLMs provide a promising approach to support analysts in the context of peacekeeping missions.
CARVE: チャンク並列リニア アテンションの価値効率を備えたコンテンツ認識型リカレント
リカレントモデルは記憶するために忘れなければなりませんが、最先端の技術では、何が保存されているかを考慮せずに何を消去するかを決定します。ゲートは到着したトークンのみを認識し、変更しようとしているメモリは認識しません。このメモリ ブラインド ゲーティングは、主要なデルタ ルール アーキテクチャ (GDN-2) の 3 つの複合欠陥のうちの 1 つです。値軸消去マスクは、値射影のスケールでパラメータを無駄にし、--私たちが証明しているように--反復トレーニングを Transformers と競合させる WY 形式の三角形チャンク ソルバーを数学的に阻止します。 CARVE (Content-Aware Recurrent with Value Efficiency) を導入します。これは、キー軸上でのみ消去するという 1 つの原則によって 3 つの問題すべてを解決します。これは、WY 形式ソルバーが有効であり続けるために必要かつ十分であることが証明されています。その中で、CARVE は、GPU メモリに既に書き込まれているリカレント出力テンソルを消去ゲートの空きコンテンツ信号として再利用し、値ごとの書き込みゲート投影をヘッドごとの単一のスカラーに置き換えます。初期化では、CARVE は GDN-2 とビット同一です。品質の違いは、コンテンツ ゲートが学習した内容から生じます。 100B トークンでトレーニングされた 1.3B パラメーターで、CARVE は WikiText のパープレキシティ 15.72 (GDN-2 に対してマイナス 0.18、4.5 シグマ効果) を達成し、9 つの常識的推論ベンチマークですべての反復ベースラインをリードし、すべての RULER 検索プローブで最先端を設定します。スループット オーバーヘッドは 0.4%、ピーク メモリは 13% 低く、パラメータが 19% 減少しました。 6 つの形式的定理は、メモリ容量、リアプノフ安定性、勾配流、表現力分離、パレート最適チャンク サイズ、およびハイブリッド最適性をカバーします。
原文 (English)
CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention
Recurrent delta-rule models keep a fixed-size state matrix S (d_v x d_k) that compresses all past context. The state of the art (GDN-2) gates this update with element-wise matrix erase/write masks. This is powerful but has two defects. First, both gates are computed from the incoming token alone, making the model memory-blind: it decides what to erase without seeing what it has stored. Second, value-axis coupling in the erase gate blocks the WY-form triangular chunk solver that drives efficient training -- the intra-chunk system splits into d_v independent solves, collapsing throughput to serial-recurrence cost. We introduce CARVE (Content-Aware Recurrent with Value Efficiency), which fixes both and, via a single-launch "megakernel" scheduling of the same WY-form math, trains faster than the matrix-gated baseline it replaces. The key idea is architectural: restricting all gating to the key axis makes the intra-chunk coupling independent of the value index, restoring one unmodified WY-form solve. Within this constraint, CARVE conditions both gates on a content signal read once per chunk from the chunk-boundary state and folded algebraically into each gate's low-rank projection (by associativity, U(Sq)=(US)q), giving memory-aware gating at negligible extra traffic. At init the content projections are zero, so CARVE is bit-identical to the baseline; we prove the one-chunk staleness perturbs gates by only O(1/sqrt(L)), matching a measured 0.18% deviation flat up to L=128. At 1.3B parameters / 100B FineWeb-Edu tokens on H100 (three seeds), CARVE improves every axis: WikiText perplexity 15.72 vs 15.90 (hybrid 15.41 vs 15.62), +0.63 pp average common-sense accuracy, and state-of-the-art RULER and real-world recall -- while training +1.4% faster at matched depth and +19.3% at iso-quality depth, at +13% peak memory. Backed by six formal guarantees.
Categorizing Mathematical Concepts with LLM Voting Ensembles in Mathswitch
Mathswitch is an open-source project that imports mathematical concept records from sources such as Wikidata, Wikipedia, MathWorld, Encyclo…
Token Geometry
Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface be…
ボソン量子計算のための SRF キャビティとトランスモンのニューラルネットワーク逆設計
三次元超伝導高周波(SRF)空洞は、非常に長寿命の電磁モードを提供し、トランスモン量子ビットなどの非線形要素と結合すると、ボソン量子情報処理の有望なアーキテクチャになります。このようなシステムの逆設計、つまり、指定された電磁ターゲットと結合ターゲットを生成するデバイスの形状を回復することは、一般に 1 対多の問題です。量子ビットと空洞の結合強度は、トランスモンの幾何学形状と空洞の電磁場内の位置の両方に敏感に依存します。これらのシステムがスケールアップし、設計パラメータ空間が拡大するにつれて、従来の反復シミュレーションのコストは法外なものになります。設計スタックの相補的なレベルでこの逆設計問題に対処する 2 つのディープ ニューラル ネットワーク (DNN) アプローチを紹介します。 1 つ目は、ターゲット キャビティの観測値を生成する SRF キャビティ ジオメトリを提案します。 2 つ目は、ターゲット量子ビット共振器パラメーター (結合率、量子ビット周波数、非調和性 $(g, \nu_q, \alpha)$) を生成するトランスモン量子ビット設計を提案します。復元された候補設計は $\sim$5\% (キャビティ) および $\sim$2\% (トランスモン) 以内でターゲットと一致しており、エンドツーエンドの再シミュレーションによって確認されます。どちらのアプローチも、望ましいデバイスの動作を候補設計に直接マッピングするもので、通常必要とされる反復シミュレーション研究に代わる迅速な代替手段となります。
原文 (English)
Neural-Network Inverse Design of SRF Cavities and Transmons for Bosonic Quantum Computation
Three-dimensional superconducting radio-frequency (SRF) cavities provide exceptionally long-lived electromagnetic modes and, when coupled to nonlinear elements such as transmon qubits, become promising architectures for bosonic quantum information processing. The inverse design of such systems, i.e., recovering device geometries that produce specified electromagnetic and coupling targets, is generally a one-to-many problem. The qubit-cavity coupling strength depends sensitively on both the transmon geometry and its position within the cavity's electromagnetic field. As these systems scale up and their design parameter spaces grow, the cost of conventional iterative simulation becomes prohibitive. We present two deep neural network (DNN) approaches that address this inverse-design problem at complementary levels of the design stack. The first proposes SRF cavity geometries that produce target cavity observables. The second proposes transmon qubit designs that produce target qubit-cavity parameters - the coupling rate, qubit frequency, and anharmonicity $(g, \nu_q, \alpha)$. The recovered candidate designs match the targets to within ~5% (cavity) and ~2% (transmon), confirmed by end-to-end re-simulation. Both approaches map desired device behavior directly to candidate designs, a fast alternative to the iterative simulation studies usually required.
Cross-Receiver Open-Set Radio Frequency Fingerprinting via Structure-First Adaptation
Radio frequency fingerprint identification (RFFI) provides a physical-layer credential for Internet of Things devices, but open-set decisio…
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale
There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various…
CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples tha…
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy
Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards st…
Hierarchical Multi-Agent Reinforcement Learning for Carbon-Aware AI Data Centers in Power Distribution Systems
Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy co…
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on s…
Piercing Gilbreath's Conjecture: From Deep Number Theory Insights to Fintech and Cybersecurity
I propose a new methodology to attack the fascinating Gilbreath's conjecture about prime numbers, first posted in 1878 and unsolved to this…
Transferability Between Understanding and Generation in Unified Multimodal Models
Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact…
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their cap…
Wan-Streamer v0.2: Higher Resolution, Same Latency
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps t…
TACTIC-KG: Toward Small Agent Teams for Cyber Threat Intelligence Knowledge Graph Construction
Cyber Threat Intelligence (CTI) reports are predominantly unstructured, heterogeneous, and noisy, which limits their direct usability for a…
Unified Audio Intelligence Without Regressing on Text Intelligence
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-…
Shape Over Intensity: Directional Topological Encoding for False Positive Reduction in Intracranial Aneurysm Detection
Automated detection of intracranial aneurysms (IAs) from CT angiography (CTA) is severely hindered by high false-positive rates. Convolutio…
Multiplayer Interactive World Models with Representation Autoencoders
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-pl…